For Teams with strict data-residency needs · Self-Hosted A/B Testing Platform

ExperimentHub

Self-hosted A/B testing with deterministic variant assignment.

DecisionLet the statistics engine say no: return a sequential verdict such as keep running, not a bare p-value.

Instead of a plain t-test endpoint that calls a winner at p < 0.05.Cost: early looks need a higher bar, so a significant p-value can still mean keep running.

Self-hosted · public repo, no licence file yet
Status
Local build · not deployed
Ownership
Solo · product and architecture
Evidence
Built: works locally, lightly tested
Checked
Last verified

Self-hosted, multi-tenant A/B testing for teams whose users' data cannot leave their infrastructure. Deterministic assignment, buffered ingestion, and a statistics engine that can say keep running. Solo build.

Role
Solo Builder · Product & Architecture
Timeline
Mar 2026 – Aug 2026 (last commit 22 Aug 2026)
Team
Solo
My part
Everything: product definition, architecture and the code. Solo build.
Stack and tools8

Elixir · Phoenix · Rust · Python · FastAPI · React · TypeScript · Kafka

ExperimentHub command centre on the synthetic demo tenant: running experiments, draft queue, assignments today, enabled flags, recent experiments and a lifecycle snapshot.

Command centre on the demo tenant; every number is synthetic seed data.

TL;DR

Teams with strict data-residency needs can't ship their users' behaviour to SaaS experimentation vendors. I built ExperimentHub solo: self-hosted, multi-tenant A/B testing where the same user always lands in the same variant, ingestion never slows the product being measured, and the statistics engine can say keep running when a naive p-value says ship. Evidence so far comes from a synthetic demo tenant, not production use.

ContinueSequential verdict at p = 0.0337 (synthetic demo)
Chapters

Problem

Experimentation is the closest thing product management has to a scientific instrument, and it is exactly the tooling most teams outsource. Hosted platforms such as Optimizely and LaunchDarkly are excellent, but they carry two structural costs: your users' behavioural data leaves your infrastructure, and assignment depends on a vendor's SDK, rules and uptime.

The gap: a self-hosted platform covering the full lifecycle (draft → running → paused → concluded) with assignment that is fast, deterministic and entirely local, and statistics that stop a team from calling a winner too early.

Users

The PM or analyst at a privacy-sensitive team, or anyone building an internal experimentation platform, who wants an experiment they can trust on infrastructure their own company controls.

There was no interview programme behind this one, and adoption is not measured. Three known failure modes of experimentation shaped the plan:

  • Assignment is where trust leaks first. If the variant depends on a database row or a network call, the same person can land in different buckets, and every downstream result inherits that doubt.
  • Peeking is the quiet killer. Stopping at the first p-value under 0.05 inflates false positives, and a t-test endpoint encourages the habit. The question a PM actually asks is "can I stop this experiment today?"
  • Analytics must never slow the product being measured, or analysis load becomes a regression in the thing under test.

Decision

Three supporting principles followed. Assignment is a pure function: same experiment, same user, same variant, forever. Ingestion and analysis are decoupled. And the platform is multi-tenant from day one, with roles (viewer/editor/admin), API keys and audit logging as platform features rather than bolt-ons.

Trade-offs

Alternatives considered, and where each landed
OptionStatusWhy
Sequential alpha-spending boundaries (O'Brien-Fleming, Pocock)ChosenPeeking inflates false positives; early looks need a higher bar.
Plain t-test endpoint onlyRejectedWould have called the demo experiment a winner at a look the boundary refused.
Pure-hash assignment, no lookupChosenStateless; reproducible from any service with the same hash.
Assignment by database lookupRejectedA round trip per decision; results depend on stored state, not a contract.
Buffered ingestion, analysis pulled on a scheduleChosenBuffers through outages; analysis never touches the measured product.
Synchronous analysis on event writeRejectedTies product latency to whatever the statistics engine is doing.
Schema-level tenant isolationDeferredA harder boundary than row-level, at the cost of migration convenience.
Bayesian exposure, CUPED, flag-creation UI, persistent results cacheDeferredDeliberately deferred; the results cache is still in-memory.

What shipped

An experiment begins as a draft. The PM defines variants and the metrics that decide it, and the power and sample-size calculator says how many users each variant needs, so the experiment is sized before it starts rather than argued about afterwards.

Starting it hands assignment to the platform. Each user is placed by a deterministic calculation over the experiment and the user's ID, so the same person sees the same variant on every visit, from every service. Events land in a buffer first, so a traffic spike or a slow analysis never touches the product's own speed.

Experiments move through explicit states, with scheduled starts and ends, guardrails and notifications handled in the background. The results view puts p-value, power and effect size above the variant table, and under sequential monitoring the engine says whether the boundary is crossed and recommends continue or stop.

  • 4 statistical methods: z-test, Welch's t-test, O'Brien-Fleming and Pocock sequential boundaries, plus power analysis.
  • Multi-tenant with roles, API keys, feature flags, audit and GDPR export/erase.
  • 50+ API routes and 11 background worker modules for lifecycle automation, guardrails, retention and partitioning.

The command centre at the top of this page shows portfolio counters plus the next action each experiment needs, on the synthetic demo tenant.

Results for a concluded demo experiment on synthetic seed data: p-value, power and effect size sit above the variant table, so the read is a decision.
Under the hoodArchitecture and implementation

Deterministic assignment in Rust

The assignment core is a small Rust crate, 122 lines across its src/ files (inline unit tests included; the separate test and benchmark files are not counted), exposed to Elixir as a Rustler NIF with a WASM build target. It runs MurmurHash3 (128-bit) over "{experiment_key}:{user_id}", maps the result into a 10,000-slot basis-points bucket space, then walks the cumulative traffic allocations to pick the variant. A pure-Elixir fallback keeps development working where the NIF can't compile.

// Simplified from assignment_core/src
pub fn assign_variant(user_id: &str, experiment_key: &str, allocations: &[u32]) -> usize {
    let bucket = hash_to_bucket(user_id, experiment_key); // murmur3_x64_128 % 10_000
    let mut cumulative = 0;
    for (i, &alloc) in allocations.iter().enumerate() {
        cumulative += alloc;
        if bucket < cumulative { return i; }
    }
    allocations.len() - 1
}

Event ingestion built to absorb bursts

An event_collector app receives single and batch events, publishes to Kafka through Broadway, and buffers through Kafka outages. Analysis is pulled by Oban-scheduled workers, never pushed synchronously, so the statistical engine reads aggregates on a schedule instead of chasing a stream.

A statistical engine that respects peeking

The Python/FastAPI engine implements z-tests for proportions and Welch's t-test, O'Brien-Fleming and Pocock alpha-spending boundaries for sequential monitoring, and a power/sample-size calculator. Services authenticate with an internal key and propagate W3C trace context.

Lifecycle, tenancy and components

Explicit state machines with optimistic locking. Background workers handle scheduled starts and ends, analysis triggers, guardrail monitoring, notifications, data retention and partition management. Every state change lands in an audit log.

Four Elixir umbrella apps (domain core, Phoenix web layer, event collector, assignment engine wrapper) plus the Rust core, the Python engine and a React 19 + TypeScript dashboard with Phoenix Channels pushing live updates. JWT sessions and API keys are tenant-scoped; tenancy is enforced at row level.

Validation

The evidence comes from a seeded demo tenant whose data is synthetic, generated by the repo's demo seed module. It is not production use.

The clearest proof is the engine's readout on the synthetic checkout-copy experiment at an interim look, 5,500 of 7,678 required samples in. The z-test gave p = 0.0337 and a 17.9% relative lift, which looks shippable. The O'Brien-Fleming boundary at 71.6% information demanded an observed z above 2.316; the observed z was 2.127. Verdict: continue.

Artifactanalyze --experiment checkout-copy-demo (synthetic seed data)
$ analyze --experiment checkout-copy-demo --metric checkout_conversion
ExperimentHub statistical engine · frequentist + sequential readout
──────────────────────────────────────────────────────────────────
variant                  n  conversions     rate
control              2,700          270   10.00%
reassurance-copy     2,800          330   11.79%

two-proportion z-test (z_test_proportions)
  lift          +1.79 pp absolute   +17.9% relative
  p-value       0.0337            significant at α=0.05: yes
  95% CI (diff) [+0.14 pp, +3.43 pp]
  cohen's h     0.0574            power achieved: 0.57

sample size (MDE 2 pp, power 0.80)
  required 3,839/variant · collected 5,500 of 7,678 · sufficient: no

sequential monitoring · obrien_fleming α-spending
  information fraction   0.716
  nominal α spent        0.0206
  boundary |z| >         2.316
  observed z             2.127
  verdict                CONTINUE — boundary not crossed

recommendation  insufficient_data: Only 5500 of 7678 required samples collected. Continue running.
overall_status  insufficient_data · computed in 6 ms

The engine's output against the synthetic seeded demo experiment at an interim look. The naive p-value clears α=0.05; the O'Brien-Fleming boundary does not. These are seed numbers, not results from a real product.

Assignment is verifiable by construction: the bucket math is small enough to re-implement in any SDK and check against the crate's golden vectors.

Limits and next

  • No production tenant. Adoption, assignment latency under load and ingestion throughput are not measured.
  • Deferred on purpose: feature-flag creation UI, Bayesian analysis exposure and CUPED; the results cache is still in-memory.
  • Licence: not yet chosen. The repository has no licence file, so the code is public to read but not licensed for reuse.
  • Next bet: work the deferred list in the order a PM would feel it: results cache out of memory, then the feature-flag creation UI, then Bayesian exposure.
  • Open question: whether row-level tenancy is enough. It keeps deployment simple but pushes discipline into every query.
  • What I'd do differently: build the continue-or-stop recommendation (now in the results view) earlier in the build order. A PM's real question is "can I stop this experiment today?", and it should have shaped the dashboard from the first screen.

Credits

Solo build: product definition, architecture and code. Source: github.com/atavisticrystal6888/A-B-Testing-Platform (public; returned HTTP 200 on 30 Sep 2026). The design lessons are written up in Designing a self-hosted experimentation platform.

  • Product

    DeskTasks

    Your task list, pinned behind every window on the desktop.

    Decision: Pin the widget behind every window, not on top of them, and refuse attention: accountless, local-first, no dock icon.

    Solo · product, design and engineering

    LiveHosted v1.3.1 line is frozen; v2 alpha runs locallyEvidence: Tested

    0Failures · release gate run, 15 Sep 2026

  • Product

    Better-Half

    Cycle-aware daily guidance for long-distance couples.

    Decision: Enforce consent in the database, not the interface: every sharing choice maps to a row-level security policy.

    Solo · product and engineering

    PrivateEvidence: Tested

    575RLS policy checks (test suite)