Context
The report in src/context_leak/report.py shows mean disclosure-rate and mean utility separately. The README stresses that the headline is the pair — withhold from forbidden recipients while still completing the appropriate flows. A single joint metric makes that explicit: the fraction of scenarios where disclosure-rate == 0 and utility == 1 (a scenario the agent got fully right on both axes).
Deterministic; computed from the existing per-scenario ScoreResults — no new dependency, no model in the loop.
Acceptance criteria
AggregateRow (or the report payload) gains a clean_and_complete rate = (#scenarios with disclosure-rate 0 and utility 1) / (#scenarios).
- Reported with a seeded bootstrap 95% CI over scenarios, consistent with the existing rates.
- Shown in both
render_text and render_json.
- Pydantic/frozen-dataclass typing preserved;
uv run mypy src clean.
- Tests:
compliant → 1.0, naive → 0.0 on the built-in suite; CI keys present in JSON.
Context
The report in
src/context_leak/report.pyshows mean disclosure-rate and mean utility separately. The README stresses that the headline is the pair — withhold from forbidden recipients while still completing the appropriate flows. A single joint metric makes that explicit: the fraction of scenarios where disclosure-rate == 0 and utility == 1 (a scenario the agent got fully right on both axes).Deterministic; computed from the existing per-scenario
ScoreResults — no new dependency, no model in the loop.Acceptance criteria
AggregateRow(or the report payload) gains aclean_and_completerate = (#scenarios with disclosure-rate 0 and utility 1) / (#scenarios).render_textandrender_json.uv run mypy srcclean.compliant→ 1.0,naive→ 0.0 on the built-in suite; CI keys present in JSON.