Skip to content

About

Audits A/B test conclusions before they become shipping decisions. Catches peeking, underpowering, SRM, and multiple comparisons errors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

experiment-auditor

A tool that audits A/B test conclusions before they become shipping decisions.

It does not replace your testing platform. It catches the reasoning errors that platforms do not stop: teams that ran the test outside a formal framework, stopped it early because the dashboard turned green, ran five variants without correcting for multiple comparisons, or are about to declare a winner on three days of data that does not include a weekend.

The model never touches a number. All statistical computation is deterministic Python. The model classifies the situation and interprets the results in plain language. This is a deliberate architectural decision, documented below.


What it catches

Flag What it means
EARLY_STOPPING Test stopped at first significance without a sequential testing framework
INSUFFICIENT_DURATION Test ran fewer than 7 days, missing a full weekly cycle
NOVELTY_EFFECT_RISK New feature, short duration; initial lift may be curiosity-driven
MULTIPLE_COMPARISONS_VARIANTS More than 2 variants; Bonferroni correction required
MULTIPLE_COMPARISONS_METRICS More than 1 primary metric; family-wise error rate inflated
SAMPLE_RATIO_MISMATCH Assignment split deviates significantly from planned split; randomization may be broken
INSUFFICIENT_POWER Sample too small relative to a detectable effect size
LOW_BASE_RATE Sub-1 percent base rate with inadequate sample; result very likely noise
INSUFFICIENT_SAMPLE_VS_PLAN Collected less than 50 percent of planned sample size
EXTERNAL_EVENT_CONTAMINATION Concurrent event reported; lift may not be attributable to the variant
NO_RANDOMIZED_CONTROL Pre-post comparison; no concurrent control group
CAUSAL_OVERCLAIM Causal language used on non-randomized data
SEGMENT_REVERSAL Aggregate result reverses within subgroups; Simpson's paradox
INSUFFICIENT_WASHOUT Prior test ended too recently; carryover behaviour may contaminate control
BORDERLINE_RESULT P-value between 0.04 and 0.05; result should be treated with extra caution
PRACTICAL_SIGNIFICANCE_WARNING Large sample, statistically significant, but effect size is very small

Architecture and design decisions

Why the model never computes statistics. Language models hallucinate arithmetic. A model given raw numbers will sometimes return plausible-looking but wrong p-values. The only reliable way to compute statistical results is with deterministic code. The architecture enforces this: compute.py runs all arithmetic and returns a structured result object. The model receives that result object and reasons about it in natural language. The model cannot alter the numbers.

Why two-sided tests by default. One-sided tests are only appropriate when the direction of the effect was pre-registered before the experiment ran. Most teams do not pre-register direction and then look at results. Defaulting to two-sided tests is the conservative and appropriate choice for after-the-fact audits.

Why 20 percent relative MDE for the power check. A 10 percent relative MDE requires very large samples at low base rates and would flag most real-world tests as underpowered. A 20 percent relative lift is a more realistic threshold for a product change to be worth shipping. Teams targeting smaller effects should run their own power analysis before the test, not rely on this tool.

Why Bonferroni and not a more powerful correction. Bonferroni is the most conservative correction and the easiest to explain to a stakeholder. For an audit tool used after the fact, erring toward caution is the right default. Teams running planned experiments with multiple variants should use Benjamini-Hochberg or a Bayesian approach, which are better choices pre-registration.

The SRM threshold of p < 0.001. A sample ratio mismatch at p < 0.05 may be random variation. At p < 0.001, the chance of observing that split by chance is very low and the randomization mechanism should be investigated before any conclusion is drawn.


Project structure

experiment-auditor/
├── src/
│   ├── compute.py        # All statistics. Deterministic. No model.
│   ├── gate.py           # LLM: classify the situation and plan the analysis
│   └── interpreter.py    # LLM: turn computed results into a plain-language verdict
├── evals/
│   └── test_cases.json   # 20 labelled cases covering the full failure taxonomy
├── tests/
│   └── test_compute.py   # Pytest tests for the computation layer
├── main.py               # CLI entry point
└── requirements.txt

Evaluation

The tool is evaluated against 20 hand-labelled cases in evals/test_cases.json. Each case has a ground truth verdict (VALID, INVALID, or INCONCLUSIVE) and a set of expected flags. The cases cover every failure mode in the taxonomy, including compound failures where multiple violations occur simultaneously.

Gate accuracy and flag precision are reported separately. A tool that refuses everything is not useful; the evaluation measures both false positives and false negatives.


Status

  • Computation layer (compute.py) — complete, 10 of 10 tests passing
  • Evaluation suite (test_cases.json) — 20 cases complete
  • Gate layer (gate.py) — in progress
  • Interpreter layer (interpreter.py) — in progress
  • CLI (main.py) — in progress
  • Deployment — planned

About

Audits A/B test conclusions before they become shipping decisions. Catches peeking, underpowering, SRM, and multiple comparisons errors.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages