feat(report): add the joint clean-and-complete rate - #47
Open
dchaudhari7177 wants to merge 1 commit into
Open
dchaudhari7177 wants to merge 1 commit into
dchaudhari7177 wants to merge 1 commit into
Conversation
The report showed mean disclosure-rate and mean utility separately, but the README's point is that the headline is the *pair*. The two means do not say it: an agent that is clean on one half of a suite and complete on the other scores 0.5/0.5 while never once getting a scenario fully right. AggregateRow gains clean_and_complete = (#scenarios with disclosure-rate 0 AND utility 1) / (#scenarios), with a seeded bootstrap 95% CI computed by the same _bootstrap_ci95 over a per-scenario 0/1 indicator, so it is a mean like the other two rates and stays deterministic. Shown in both renderings: - render_text gains a clean+complete column -- yes/no per scenario, the rate with its CI on the aggregate row -- plus a legend line naming what the column means. - render_json gains aggregate.clean_and_complete and aggregate.clean_and_complete_ci95, and a per-scenario boolean. On the built-in suite this is exactly what the metric is for: the naive agent has utility 1.00, so on that axis it looks indistinguishable from compliant, and only the joint rate separates them -- 0.00 against 1.00. Computed from the existing ScoreResults; frozen dataclasses preserved, stdlib only, no new dependency, no model in the loop. 8 tests: compliant 1.0, naive 0.0 with its utility 1.0 asserted alongside so the discrimination is explicit, a synthetic 0.5/0.5 split proving the joint rate is not implied by the two means, the CI brackets the point estimate, the bootstrap is reproducible, the text column and per-scenario verdicts, and the JSON keys in both directions. 74 passed; ruff check, ruff format --check and mypy src clean. Closes bamdadd#35
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #35.
The report showed mean disclosure-rate and mean utility separately, but the README's point is that the headline is the pair. The two means do not say it: an agent that is clean on one half of a suite and complete on the other scores 0.5/0.5 while never once getting a scenario fully right.
AggregateRowgainsclean_and_complete= (#scenarios with disclosure-rate 0 and utility 1) / (#scenarios), with a seeded bootstrap 95% CI from the same_bootstrap_ci95over a per-scenario 0/1 indicator — so it is a mean like the other two rates and stays deterministic.Why it earns its place, on this suite
The naive agent has utility 1.00 — on that axis it is indistinguishable from the compliant agent (which scores
1.00clean+complete). Only the joint rate separates them.Both renderings
render_text: aclean+completecolumn, yes/no per scenario and the rate with its CI on the aggregate row, plus a legend line.render_json:aggregate.clean_and_completeandaggregate.clean_and_complete_ci95, plus a per-scenario boolean.Computed from the existing
ScoreResults. Frozen dataclasses preserved, stdlib only, no new dependency, no model in the loop.Tests
8: compliant
1.0; naive0.0with its utility 1.0 asserted alongside so the discrimination is explicit; a synthetic 0.5/0.5 split proving the joint rate is not implied by the two means; the CI brackets the point estimate; the bootstrap is reproducible across builds; the text column and per-scenario verdicts; and the JSON keys in both directions.74 passed;
ruff check,ruff format --checkandmypy srcclean.