A tool that audits A/B test conclusions before they become shipping decisions.
It does not replace your testing platform. It catches the reasoning errors that platforms do not stop: teams that ran the test outside a formal framework, stopped it early because the dashboard turned green, ran five variants without correcting for multiple comparisons, or are about to declare a winner on three days of data that does not include a weekend.
The model never touches a number. All statistical computation is deterministic Python. The model classifies the situation and interprets the results in plain language. This is a deliberate architectural decision, documented below.
| Flag | What it means |
|---|---|
| EARLY_STOPPING | Test stopped at first significance without a sequential testing framework |
| INSUFFICIENT_DURATION | Test ran fewer than 7 days, missing a full weekly cycle |
| NOVELTY_EFFECT_RISK | New feature, short duration; initial lift may be curiosity-driven |
| MULTIPLE_COMPARISONS_VARIANTS | More than 2 variants; Bonferroni correction required |
| MULTIPLE_COMPARISONS_METRICS | More than 1 primary metric; family-wise error rate inflated |
| SAMPLE_RATIO_MISMATCH | Assignment split deviates significantly from planned split; randomization may be broken |
| INSUFFICIENT_POWER | Sample too small relative to a detectable effect size |
| LOW_BASE_RATE | Sub-1 percent base rate with inadequate sample; result very likely noise |
| INSUFFICIENT_SAMPLE_VS_PLAN | Collected less than 50 percent of planned sample size |
| EXTERNAL_EVENT_CONTAMINATION | Concurrent event reported; lift may not be attributable to the variant |
| NO_RANDOMIZED_CONTROL | Pre-post comparison; no concurrent control group |
| CAUSAL_OVERCLAIM | Causal language used on non-randomized data |
| SEGMENT_REVERSAL | Aggregate result reverses within subgroups; Simpson's paradox |
| INSUFFICIENT_WASHOUT | Prior test ended too recently; carryover behaviour may contaminate control |
| BORDERLINE_RESULT | P-value between 0.04 and 0.05; result should be treated with extra caution |
| PRACTICAL_SIGNIFICANCE_WARNING | Large sample, statistically significant, but effect size is very small |
Why the model never computes statistics. Language models hallucinate arithmetic. A model given raw numbers will sometimes return plausible-looking but wrong p-values. The only reliable way to compute statistical results is with deterministic code. The architecture enforces this: compute.py runs all arithmetic and returns a structured result object. The model receives that result object and reasons about it in natural language. The model cannot alter the numbers.
Why two-sided tests by default. One-sided tests are only appropriate when the direction of the effect was pre-registered before the experiment ran. Most teams do not pre-register direction and then look at results. Defaulting to two-sided tests is the conservative and appropriate choice for after-the-fact audits.
Why 20 percent relative MDE for the power check. A 10 percent relative MDE requires very large samples at low base rates and would flag most real-world tests as underpowered. A 20 percent relative lift is a more realistic threshold for a product change to be worth shipping. Teams targeting smaller effects should run their own power analysis before the test, not rely on this tool.
Why Bonferroni and not a more powerful correction. Bonferroni is the most conservative correction and the easiest to explain to a stakeholder. For an audit tool used after the fact, erring toward caution is the right default. Teams running planned experiments with multiple variants should use Benjamini-Hochberg or a Bayesian approach, which are better choices pre-registration.
The SRM threshold of p < 0.001. A sample ratio mismatch at p < 0.05 may be random variation. At p < 0.001, the chance of observing that split by chance is very low and the randomization mechanism should be investigated before any conclusion is drawn.
experiment-auditor/
├── src/
│ ├── compute.py # All statistics. Deterministic. No model.
│ ├── gate.py # LLM: classify the situation and plan the analysis
│ └── interpreter.py # LLM: turn computed results into a plain-language verdict
├── evals/
│ └── test_cases.json # 20 labelled cases covering the full failure taxonomy
├── tests/
│ └── test_compute.py # Pytest tests for the computation layer
├── main.py # CLI entry point
└── requirements.txt
The tool is evaluated against 20 hand-labelled cases in evals/test_cases.json. Each case has a ground truth verdict (VALID, INVALID, or INCONCLUSIVE) and a set of expected flags. The cases cover every failure mode in the taxonomy, including compound failures where multiple violations occur simultaneously.
Gate accuracy and flag precision are reported separately. A tool that refuses everything is not useful; the evaluation measures both false positives and false negatives.
- Computation layer (compute.py) — complete, 10 of 10 tests passing
- Evaluation suite (test_cases.json) — 20 cases complete
- Gate layer (gate.py) — in progress
- Interpreter layer (interpreter.py) — in progress
- CLI (main.py) — in progress
- Deployment — planned