Problem
Experimentation is the closest thing product management has to a scientific instrument, and it is exactly the tooling most teams outsource. Hosted platforms such as Optimizely and LaunchDarkly are excellent, but they carry two structural costs: your users' behavioural data leaves your infrastructure, and assignment depends on a vendor's SDK, rules and uptime.
The gap: a self-hosted platform covering the full lifecycle (draft → running → paused → concluded) with assignment that is fast, deterministic and entirely local, and statistics that stop a team from calling a winner too early.
Users
The PM or analyst at a privacy-sensitive team, or anyone building an internal experimentation platform, who wants an experiment they can trust on infrastructure their own company controls.
There was no interview programme behind this one, and adoption is not measured. Three known failure modes of experimentation shaped the plan:
- Assignment is where trust leaks first. If the variant depends on a database row or a network call, the same person can land in different buckets, and every downstream result inherits that doubt.
- Peeking is the quiet killer. Stopping at the first p-value under 0.05 inflates false positives, and a t-test endpoint encourages the habit. The question a PM actually asks is "can I stop this experiment today?"
- Analytics must never slow the product being measured, or analysis load becomes a regression in the thing under test.
Decision
Three supporting principles followed. Assignment is a pure function: same experiment, same user, same variant, forever. Ingestion and analysis are decoupled. And the platform is multi-tenant from day one, with roles (viewer/editor/admin), API keys and audit logging as platform features rather than bolt-ons.
Trade-offs
What shipped
An experiment begins as a draft. The PM defines variants and the metrics that decide it, and the power and sample-size calculator says how many users each variant needs, so the experiment is sized before it starts rather than argued about afterwards.
Starting it hands assignment to the platform. Each user is placed by a deterministic calculation over the experiment and the user's ID, so the same person sees the same variant on every visit, from every service. Events land in a buffer first, so a traffic spike or a slow analysis never touches the product's own speed.
Experiments move through explicit states, with scheduled starts and ends, guardrails and notifications handled in the background. The results view puts p-value, power and effect size above the variant table, and under sequential monitoring the engine says whether the boundary is crossed and recommends continue or stop.
- 4 statistical methods: z-test, Welch's t-test, O'Brien-Fleming and Pocock sequential boundaries, plus power analysis.
- Multi-tenant with roles, API keys, feature flags, audit and GDPR export/erase.
- 50+ API routes and 11 background worker modules for lifecycle automation, guardrails, retention and partitioning.
The command centre at the top of this page shows portfolio counters plus the next action each experiment needs, on the synthetic demo tenant.
Under the hoodArchitecture and implementation
Deterministic assignment in Rust
The assignment core is a small Rust crate, 122 lines across its src/ files (inline unit tests included; the separate test and benchmark files are not counted), exposed to Elixir as a Rustler NIF with a WASM build target. It runs MurmurHash3 (128-bit) over "{experiment_key}:{user_id}", maps the result into a 10,000-slot basis-points bucket space, then walks the cumulative traffic allocations to pick the variant. A pure-Elixir fallback keeps development working where the NIF can't compile.
// Simplified from assignment_core/src
pub fn assign_variant(user_id: &str, experiment_key: &str, allocations: &[u32]) -> usize {
let bucket = hash_to_bucket(user_id, experiment_key); // murmur3_x64_128 % 10_000
let mut cumulative = 0;
for (i, &alloc) in allocations.iter().enumerate() {
cumulative += alloc;
if bucket < cumulative { return i; }
}
allocations.len() - 1
}
Event ingestion built to absorb bursts
An event_collector app receives single and batch events, publishes to Kafka through Broadway, and buffers through Kafka outages. Analysis is pulled by Oban-scheduled workers, never pushed synchronously, so the statistical engine reads aggregates on a schedule instead of chasing a stream.
A statistical engine that respects peeking
The Python/FastAPI engine implements z-tests for proportions and Welch's t-test, O'Brien-Fleming and Pocock alpha-spending boundaries for sequential monitoring, and a power/sample-size calculator. Services authenticate with an internal key and propagate W3C trace context.
Lifecycle, tenancy and components
Explicit state machines with optimistic locking. Background workers handle scheduled starts and ends, analysis triggers, guardrail monitoring, notifications, data retention and partition management. Every state change lands in an audit log.
Four Elixir umbrella apps (domain core, Phoenix web layer, event collector, assignment engine wrapper) plus the Rust core, the Python engine and a React 19 + TypeScript dashboard with Phoenix Channels pushing live updates. JWT sessions and API keys are tenant-scoped; tenancy is enforced at row level.
Validation
The evidence comes from a seeded demo tenant whose data is synthetic, generated by the repo's demo seed module. It is not production use.
The clearest proof is the engine's readout on the synthetic checkout-copy experiment at an interim look, 5,500 of 7,678 required samples in. The z-test gave p = 0.0337 and a 17.9% relative lift, which looks shippable. The O'Brien-Fleming boundary at 71.6% information demanded an observed z above 2.316; the observed z was 2.127. Verdict: continue.
$ analyze --experiment checkout-copy-demo --metric checkout_conversion
ExperimentHub statistical engine · frequentist + sequential readout
──────────────────────────────────────────────────────────────────
variant n conversions rate
control 2,700 270 10.00%
reassurance-copy 2,800 330 11.79%
two-proportion z-test (z_test_proportions)
lift +1.79 pp absolute +17.9% relative
p-value 0.0337 significant at α=0.05: yes
95% CI (diff) [+0.14 pp, +3.43 pp]
cohen's h 0.0574 power achieved: 0.57
sample size (MDE 2 pp, power 0.80)
required 3,839/variant · collected 5,500 of 7,678 · sufficient: no
sequential monitoring · obrien_fleming α-spending
information fraction 0.716
nominal α spent 0.0206
boundary |z| > 2.316
observed z 2.127
verdict CONTINUE — boundary not crossed
recommendation insufficient_data: Only 5500 of 7678 required samples collected. Continue running.
overall_status insufficient_data · computed in 6 ms
The engine's output against the synthetic seeded demo experiment at an interim look. The naive p-value clears α=0.05; the O'Brien-Fleming boundary does not. These are seed numbers, not results from a real product.
Assignment is verifiable by construction: the bucket math is small enough to re-implement in any SDK and check against the crate's golden vectors.
Limits and next
- No production tenant. Adoption, assignment latency under load and ingestion throughput are not measured.
- Deferred on purpose: feature-flag creation UI, Bayesian analysis exposure and CUPED; the results cache is still in-memory.
- Licence: not yet chosen. The repository has no licence file, so the code is public to read but not licensed for reuse.
- Next bet: work the deferred list in the order a PM would feel it: results cache out of memory, then the feature-flag creation UI, then Bayesian exposure.
- Open question: whether row-level tenancy is enough. It keeps deployment simple but pushes discipline into every query.
- What I'd do differently: build the continue-or-stop recommendation (now in the results view) earlier in the build order. A PM's real question is "can I stop this experiment today?", and it should have shaped the dashboard from the first screen.
Credits
Solo build: product definition, architecture and code. Source: github.com/atavisticrystal6888/A-B-Testing-Platform (public; returned HTTP 200 on 30 Sep 2026). The design lessons are written up in Designing a self-hosted experimentation platform.

