Why
Shadow observations are useful only when they become decision-grade evidence rather than a dashboard of raw counters.
Scope
- Define versioned evidence-report schemas.
- Compare production and candidate routes on quality, latency, errors, usage, and cost class.
- Report sample sizes, missingness, confidence intervals, and judge agreement.
- Support deterministic human labels and explicitly versioned automated evaluators.
- Produce an outcome of enforce, keep shadowing, or do not enforce with reasons.
- Export bounded machine-readable and human-readable reports.
Acceptance criteria
- Reports cannot claim significance without the required sample and confidence threshold.
- Automated evaluation is clearly separated from observed provider outcomes and human labels.
- Prompt/response inclusion is opt-in, access-controlled, and retention-bounded.
- Reports are reproducible from their versioned evidence inputs.
- Tests cover sparse, biased, missing, conflicting, and statistically inconclusive evidence.
Parent: #146
Invariant
The scored model decision remains offline, deterministic, keyless, and explainable.
Why
Shadow observations are useful only when they become decision-grade evidence rather than a dashboard of raw counters.
Scope
Acceptance criteria
Parent: #146
Invariant
The scored model decision remains offline, deterministic, keyless, and explainable.