diff --git a/AGENTS.md b/AGENTS.md index 87c130c..69408d4 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -42,7 +42,7 @@ src/adherence/ suitedata.py one loader behind every viewing surface table.py one-shot snapshot · matrix_tui.py interactive matrix adapters/ opencode (stream + export + child sessions), api.py - tui/ VENDORED from leather — re-copy, never edit here + tui/ VENDORED from pane — re-copy, never edit here adapters/opencode.sh the shell adapter; the `--adapter` contract scenarios/sNN/ scenario.yaml + fixture/ + grade.py bench/ sandbox layer: isolate (+ proxy), preflight, prewarm diff --git a/CHANGELOG.md b/CHANGELOG.md index 5a8fcf3..d2dd32b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,33 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Fixed + +- **`report.py` and `calibrate.py` scored harness faults as model results.** + Both predate the shared `suitedata.is_ungradeable()` and neither was + updated when it landed. `make report` read pass@1 0.40 for + `cli-cli-13057` against 1.00 from `probe.py`, `analyze.py` and the + viewer — 0.40 is inside the registered calibration band and 1.00 is + outside it. `make calibrate`, the H4 gate, reported 19.571% aggregate + divergence over 120 runs against the viewer tab's 9.484% over 117, for + the same two files. Both now import the one definition. (#9, #10) +- **`make selftest` now enforces that they cannot drift again**, by running + one synthetic cell of two harness faults, two passes and one real failure + through `analyze`, `suitedata`, `report`, `probe` and `calibrate` and + requiring the same rate from each. Every one of the five defects was + re-introduced and confirmed to fail it. (#11) +- **`suitedata.load_cells` raised `ZeroDivisionError`** on a cell whose + trials were all ungradeable, taking down `table`, `matrix` and `live` + the moment a scenario's completed trials were all ceiling hits. Cells + now carry `gradeable`, and one `pass_str` renders the absence as an em + dash rather than a 0% that reads as floor. (#8) +- `report.py` states the `purpose` mix of the rows it read, in the + document and on stderr. (#12) +- `report.py` and `calibrate.py` no longer raise `KeyError` on rows + written before `duration_s` and `arm` existed. (#13) +- `bench/ci-local.sh` generated a shell script at a predictable path in + `/tmp` and then executed it; now `mktemp -d` with a cleanup trap. (#14) + ### Changed - `src/adherence/tui/` is now vendored from its canonical public diff --git a/bench/ci-local.sh b/bench/ci-local.sh index 1e8cbbd..3eb4426 100755 --- a/bench/ci-local.sh +++ b/bench/ci-local.sh @@ -22,6 +22,14 @@ HERE="$(cd "$(dirname "$0")/.." && pwd)" JOB="${1:-${JOB:-validate}}" cd "$HERE" +# The generated steps are written and then executed, so the path they are +# written to must not be one another user can pre-create as a symlink or +# swap between the write and the `bash`. `mktemp` for the same reason the +# `python` shim below already uses it. +WORK="$(mktemp -d)" +trap 'rm -rf "$WORK"' EXIT +STEPS_SH="$WORK/steps.sh" + mapfile -t STEPS < <(python3 - "$JOB" <<'PY' import json, re, sys @@ -86,7 +94,7 @@ print(json.dumps(steps)) PY ) -python3 - "$JOB" "${STEPS[0]:-[]}" <<'PY' > /tmp/adh-ci-steps.sh +python3 - "$JOB" "${STEPS[0]:-[]}" <<'PY' > "$STEPS_SH" import json, shlex, sys job, blob = sys.argv[1], sys.argv[2] steps = json.loads(blob) @@ -107,12 +115,11 @@ PY # `python` is the interpreter name on a GitHub runner; make it resolvable here. if ! command -v python >/dev/null 2>&1; then - SHIM="$(mktemp -d)" - ln -s "$(command -v python3)" "$SHIM/python" - export PATH="$SHIM:$PATH" + ln -s "$(command -v python3)" "$WORK/python" + export PATH="$WORK:$PATH" fi -bash /tmp/adh-ci-steps.sh +bash "$STEPS_SH" RC=$? echo [ "$RC" -eq 0 ] && echo "ci-local($JOB): all steps passed" \ diff --git a/docs/HANDOFF.md b/docs/HANDOFF.md index 772532e..c625520 100644 --- a/docs/HANDOFF.md +++ b/docs/HANDOFF.md @@ -13,7 +13,7 @@ the registered experiment has not been run. | | | |---|---| -| Branch | `main` at `b4ae1ca`, clean, nothing unpushed | +| Branch | `main` at `eba93f6`, clean, nothing unpushed | | CI | **green** — `Validate (3.10)`, `Validate (3.14)`, `Lint`, and all three `Full-scope` platforms | | Selftest | 25/25 | | PRs | #1, #2, #4 merged. #3 closed deliberately (see §4) | @@ -57,12 +57,19 @@ pool them is doing real work — pooled, the CLI ceiling masks a healthy unit tier. **H4 (proxy vs adapter) fails as specified, and the failure is localized.** -19.6% aggregate, but **22 of 24 scenarios agree exactly** — 0.0% on tokens +9.5% aggregate over the 117 gradeable trials, but **22 of 24 scenarios agree +exactly** — 0.0% on tokens *and* identical call counts — and all six in-band scenarios are among them. The divergence is transcript truncation on ceiling-hit trials, not an accounting error. Cost figures must still come from `runs/probe.proxy.jsonl` per the registered rule. +This read **19.6% over 120 runs** until `calibrate.py` was made to apply +exclusion criterion 1, which `suitedata.calibration` had always applied. The +gate and its own viewer tab disagreed by 2× on the same two files (#10). H4 +fails either way — call counts are 5,319 adapter against 5,646 proxy — but the +number to plan against is 9.5%. + **Nothing here licenses a statement about the directed-contexts pattern.** One arm ran. `make analyze` correctly excludes all 120 rows as `purpose: validation` — that guard is working, not failing. @@ -139,24 +146,32 @@ task wrong, which the registration forbids in as many words. analysis**. 40% is inside the band; 100% is outside it. **The viewer was inflating the number that decides whether the experiment can run.** -### The same defect turned up in three surfaces +### The same defect turned up in five surfaces + +Three were found in this session: `suitedata` (the viewer), `probe.py` (the +go/no-go verdict), and `analyze.py` (which had always been correct). The +disagreement was silent until someone compared them. -`suitedata` (the viewer), `probe.py` (the go/no-go verdict), and `analyze.py` -(which had always been correct). The disagreement was silent until someone -compared them. +Two more were found afterwards, in the review of PR #5, and are fixed on +`main`: `report.py` (the publishable scoreboard, #9) and `calibrate.py` (the H4 +gate, #10). Both predate the shared definition and both survived the fix that +was meant to be repo-wide, because the only test pointed at `analyze` alone. +`make selftest` now asserts that every surface reporting a rate lands on the +same number (#11). `probe.py` was the worst place to hold it, because it renders a *decision*: ``` cli-cli-13057 40% 5 KEEP → cli-cli-13057 100% 2 ceiling — drop 7 of 24 tasks land in [0.25, 0.80] → 6 of 24 tasks land in [0.25, 0.80] - 16 ceiling · 1 floor · 0 harness + 16 ceiling · 1 floor · 0 harness → 17 ceiling · 1 floor · 0 harness ``` It also matched adapter faults on status `"fail"` when a ceiling hit records `"ungradeable"` — printing `0 harness` for a run with three killed trials. -There is now **one `is_ungradeable()`** in `suitedata`, imported by the others. +There is now **one `is_ungradeable()`** in `suitedata`, imported by the others, +and one selftest check that fails if a surface stops importing it. ### The `left` column counted cells that had reported, not cells that exist @@ -241,5 +256,6 @@ it originated. The run-watcher lineage was corrected twice during the session and now reads: five generations over 25 days across three codebases, beginning not in an eval -but in a pair of live server-analytics dashboards whose curses framework both -this repo and leather vendor byte-identically. +but in a pair of live server-analytics dashboards whose curses framework this +repo vendors byte-identically from [pane](https://github.com/TGPSKI/pane), its +canonical upstream since `2c0f9d9`. It was copied from `leather` before that. diff --git a/src/adherence/calibrate.py b/src/adherence/calibrate.py index 5791553..cdc930e 100644 --- a/src/adherence/calibrate.py +++ b/src/adherence/calibrate.py @@ -18,6 +18,9 @@ separately. The adapter cannot see them and they are not attributable to the instruction surface under test. Excluding them is a stated choice, not a rounding decision — see lib/metrics.is_auxiliary. + +Trials the harness stopped are excluded too, under docs/EVAL.md exclusion +criterion 1 — see `main`. """ from __future__ import annotations @@ -27,6 +30,7 @@ from pathlib import Path from adherence import metrics +from adherence.suitedata import is_ungradeable def load_jsonl(path): @@ -57,9 +61,23 @@ def main(): tot_ac = tot_pc = 0 worst = 0.0 unmatched = [] - n = 0 + n = excluded = 0 for r in results: - mark = f"{r['scenario']}|{r['arm']}|{r['trial']}" + # docs/EVAL.md exclusion criterion 1, which suitedata.calibration + # applied and this gate did not. A trial the harness stopped has a + # transcript truncated by construction, so against the proxy's + # complete record it reads as a ~100% accounting error -- which is + # not what H4 measures, and it buries the disagreements that are + # real. Measured on runs/probe.jsonl: 19.571% aggregate over 120 + # runs here against 9.484% over 117 in the viewer, for the same two + # files. Two surfaces disagreeing by 2x about a registered quantity + # is the defect, not the number. + if is_ungradeable(r): + excluded += 1 + continue + # `arm` predates nothing here, but rows written before the arms work + # carry none; analyze.load and suitedata.load_rows both default it. + mark = f"{r['scenario']}|{r.get('arm', '-')}|{r['trial']}" rows = by_mark.get(mark) m = r.get("metrics") or {} if rows is None: @@ -78,7 +96,14 @@ def main(): f"| {p_tok} | {d*100:.2f}% | {p['aux_calls']} |") agg = abs(tot_a - tot_p) / tot_p if tot_p else 1.0 - print(f"\nruns compared: **{n}**") + print(f"\nruns compared: **{n}** of {len(results)}") + if excluded: + # Counted, never silent: an exclusion nobody can see is + # indistinguishable from data that never existed. + print(f"excluded: **{excluded}** harness fault(s) " + f"(docs/EVAL.md exclusion criterion 1 — a stopped trial's " + f"transcript is truncated by construction and its " + f"disagreement with the proxy is not an accounting error)") print(f"adapter Σ input tokens: **{tot_a}** · " f"proxy Σ input tokens: **{tot_p}**") print(f"aggregate delta: **{agg*100:.3f}%** · " diff --git a/src/adherence/matrix_tui.py b/src/adherence/matrix_tui.py index b908d7e..3564fae 100644 --- a/src/adherence/matrix_tui.py +++ b/src/adherence/matrix_tui.py @@ -315,8 +315,14 @@ def visible(self): # ---- chrome ---------------------------------------------------------- - def pass_attr(self, p): + def pass_attr(self, c): + """Attribute for a pass@1 cell. Takes the cell, not the number: a + cell with nothing gradeable has no rate, and painting its + placeholder 0.0 red says the model failed where the harness did.""" C = self.curses + if not c.get("gradeable"): + return C.A_DIM + p = c["pass_rate"] return (C.color_pair(1) if p >= 80 else C.color_pair(3) if p >= 25 else C.color_pair(4)) @@ -645,8 +651,7 @@ def view_live(self, body, max_x, max_y=0): self._put(y, 25, f"{c['trials']:>6}", C.A_DIM) self._put(y, 31, f"{c['ungradeable']:>5}", C.color_pair(3) if c["ungradeable"] else C.A_DIM) - self._put(y, 36, f"{c['pass_rate']:>5.0f}%", - self.pass_attr(c["pass_rate"])) + self._put(y, 36, sd.pass_str(c, 6), self.pass_attr(c)) skewed = (c["abandoned"] and c["tok_worked"] and c["tok_worked"] > c["tok"] * 1.05) self._put(y, 42, f"{c['tok']:>11,.0f}" @@ -1345,8 +1350,7 @@ def view_cells(self, cs, body, max_x): base = C.color_pair(5) if idx == self.cursor else 0 self._put(y, 1, f"{c['tag']:<20}", base) self._put(y, 21, f"{c['trials']:>7}", C.A_DIM) - self._put(y, 28, f"{c['pass_rate']:>7.0f}%", - self.pass_attr(c["pass_rate"])) + self._put(y, 28, sd.pass_str(c, 8), self.pass_attr(c)) self._put(y, 36, f"{c['spread']:>5.0f}", C.A_DIM) self._put(y, 41, f"{c['tok']:>9,.0f}", C.color_pair(2)) self._put(y, 50, f"{c['tok_won']:>9,.0f}", C.A_DIM) @@ -1373,8 +1377,7 @@ def view_arms(self, cs, body, max_x): self._put(y, 6, f"{r['name']:<20}") self._put(y, 26, f"{r['scenarios']:>5}", C.A_DIM) self._put(y, 31, f"{r['trials']:>7}", C.A_DIM) - self._put(y, 38, f"{r['pass_rate']:>7.0f}%", - self.pass_attr(r["pass_rate"])) + self._put(y, 38, sd.pass_str(r, 8), self.pass_attr(r)) self._put(y, 46, f"{r['tok']:>10,.0f}", C.color_pair(2)) self._put(y, 56, f"{r['calls']:>7.0f}", C.A_DIM) if r["arm"] == self.ref: @@ -1417,7 +1420,11 @@ def view_cost(self, cs, body, max_x): from another that is a trade, not a win, and the chart says so where a ratio would not.""" C = self.curses - pts = [c for c in cs if c["ktok"] > 0] + # Both axes have to exist. A cell with no gradeable trial has no + # quality coordinate, and plotting its placeholder 0.0 puts a point + # at the bottom of the chart where nothing was measured -- and drags + # the y-axis floor down with it. + pts = [c for c in cs if c["ktok"] > 0 and c["gradeable"]] if not pts: self._put(3, 3, "no cells with token telemetry in this selection", C.A_DIM) @@ -1455,8 +1462,7 @@ def view_cost(self, cs, body, max_x): for i, c in enumerate(best[:max(0, body - ph - 6)]): yy = y0 + 1 + i self._put(yy, 1, f"{c['tag']:<20}") - self._put(yy, 21, f"{c['pass_rate']:>7.0f}%", - self.pass_attr(c["pass_rate"])) + self._put(yy, 21, sd.pass_str(c, 8), self.pass_attr(c)) self._put(yy, 29, f"{c['tok']:>10,.0f}", C.color_pair(2)) self._put(yy, 39, f"{c['calls']:>7.0f}", C.A_DIM) self._put(yy, 46, f"{sd.fmt_duration(c['dur_s']):>8}", @@ -1558,7 +1564,7 @@ def view_detail(self, c, max_y, max_x): y += 1 lines = [ ("trials", f"{c['trials']}"), - ("pass@1", f"{c['pass_rate']:.0f}% (± {c['spread']:.0f})"), + ("pass@1", f"{sd.pass_str(c)} (± {c['spread']:.0f})"), ("median tok_in | passed", f"{c['tok_won']:,.0f}"), ("median tok_in | worked", f"{c['tok_worked']:,.0f}" + (" (abandoned trials excluded — " diff --git a/src/adherence/report.py b/src/adherence/report.py index dcf43af..0194b46 100755 --- a/src/adherence/report.py +++ b/src/adherence/report.py @@ -21,7 +21,10 @@ import math import random import statistics as st -from collections import defaultdict +import sys +from collections import Counter, defaultdict + +from adherence.suitedata import is_ungradeable BOOTSTRAP_N = 10_000 @@ -54,7 +57,38 @@ def m(r, key, default=0): def pass_rate(rows): - return sum(1.0 if r["all_pass"] else 0.0 for r in rows) / len(rows) if rows else 0.0 + """pass@1 over gradeable rows, or None when nothing was gradeable. + + A row that no grader reached carries `all_pass=False` because nothing + set it True. Averaging over every row therefore scores a harness fault + as a model that got the task wrong, which docs/EVAL.md exclusion + criterion 1 forbids in as many words. + + This surface was the last one holding that defect. Measured on + runs/probe.jsonl, cli-cli-13057 -- 3 of 5 trials killed at the + adapter's 2,700 s ceiling -- read 0.40 here against 1.00 from probe.py, + analyze.py and suitedata. 0.40 is inside the registered calibration + band [0.25, 0.80] and 1.00 is outside it, so the scoreboard intended + for publication disagreed with the go/no-go verdict about whether that + scenario could discriminate at all. + + `is_ungradeable` is imported rather than restated: one definition, or + the surfaces drift again.""" + graded = [r for r in rows if not is_ungradeable(r)] + if not graded: + return None + return sum(1.0 if r["all_pass"] else 0.0 for r in graded) / len(graded) + + +def fmt_rate(rows): + """pass@1 for a table cell. An em dash where there is no rate: 0.00 + would read as a floor, which is the same claim as scoring the fault.""" + p = pass_rate(rows) + return "—" if p is None else f"{p:.2f}" + + +def ungradeable_n(rows): + return sum(1 for r in rows if is_ungradeable(r)) def med(vals): @@ -99,17 +133,21 @@ def arm_block(rows, model, adapter, arm): for r in sub: by_scen[r["scenario"]].append(r) - print("| scenario | category | pass@1 | check adherence | calls " - "| tok_in (med) | tok_in if passed | probes→edit | redundant " - "| abandoned | mean dur s | failing checks |") - print("|---|---|---|---|---|---|---|---|---|---|---|---|") + # `ungradeable` is a column, not a footnote. The rate above it is + # computed with those rows removed, and an exclusion nobody can count + # is indistinguishable from data that never existed (docs/EVAL.md). + print("| scenario | category | pass@1 | ungradeable | check adherence " + "| calls | tok_in (med) | tok_in if passed | probes→edit " + "| redundant | abandoned | mean dur s | failing checks |") + print("|---|---|---|---|---|---|---|---|---|---|---|---|---|") for sid in sorted(by_scen): rs = by_scen[sid] won = [r for r in rs if r["all_pass"]] fails = sorted({c["name"] for r in rs for c in r["checks"] if c["status"] == "fail"}) print(f"| {sid} | {rs[0]['category']} " - f"| {pass_rate(rs):.2f} " + f"| {fmt_rate(rs)} " + f"| {ungradeable_n(rs)}/{len(rs)} " f"| {check_rate(rs):.2f} " f"| {med([m(r, 'calls') for r in rs]):.0f} " f"| {med([m(r, 'tok_in_billed') for r in rs]):.0f} " @@ -117,23 +155,29 @@ def arm_block(rows, model, adapter, arm): f"| {med([m(r, 'probes_to_first_edit') for r in rs]):.0f} " f"| {med([m(r, 'redundant_reads') for r in rs]):.0f} " f"| {sum(1 for r in rs if m(r, 'abandoned')):d}/{len(rs)} " - f"| {st.mean(r['duration_s'] for r in rs):.0f} " + # `.get`, like every other field read here: runs/ spans the + # whole history of the harness and re-reading a file written + # before duration_s existed must not raise. + f"| {st.mean(r.get('duration_s', 0) or 0 for r in rs):.0f} " f"| {', '.join(fails) if fails else '—'} |") by_cat = defaultdict(list) for r in sub: by_cat[r["category"]].append(r) - print("\n| category | pass@1 | check adherence | tok_in (med) |") - print("|---|---|---|---|") + print("\n| category | pass@1 | ungradeable | check adherence " + "| tok_in (med) |") + print("|---|---|---|---|---|") for cat in sorted(by_cat): rs = by_cat[cat] - print(f"| {cat} | {pass_rate(rs):.2f} | {check_rate(rs):.2f} " + print(f"| {cat} | {fmt_rate(rs)} | {ungradeable_n(rs)}/{len(rs)} " + f"| {check_rate(rs):.2f} " f"| {med([m(r, 'tok_in_billed') for r in rs]):.0f} |") won = [r for r in sub if r["all_pass"]] lo, hi = iqr([m(r, "tok_in_billed") for r in sub]) subs = sum(m(r, "subagent_calls") for r in sub) - print(f"\n**overall pass@1: {pass_rate(sub):.2f} · " + print(f"\n**overall pass@1: {fmt_rate(sub)} · " + f"ungradeable: {ungradeable_n(sub)}/{len(sub)} · " f"check adherence: {check_rate(sub):.2f} · " f"median tok_in: {med([m(r, 'tok_in_billed') for r in sub]):.0f} " f"(IQR {lo:.0f}–{hi:.0f}) · " @@ -200,7 +244,7 @@ def cell(arm, sid, key, only_passed): ci = ("—" if math.isnan(lo) else f"{math.exp(lo):.3f}–{math.exp(hi):.3f}") print(f"| {arm} | {len(lrs)} | {g:.3f} | {ci} " - f"| {pass_rate(arm_rows):.2f} |") + f"| {fmt_rate(arm_rows)} |") print() @@ -220,7 +264,28 @@ def main(): for r in rows: r.setdefault("arm", "-") + # What this document was built from, in the document. + # + # analyze.py refuses rows that are not purpose=experiment, and should: + # it renders the registered verdict. This is the working scoreboard and + # reading validation rows is its job -- but a markdown table carries no + # trace of its inputs once it is pasted somewhere, and 120 validation + # rows rendered a complete scoreboard, overall pass@1 and all, saying + # nothing about what it had read. + purposes = Counter(r.get("purpose") or "unlabelled" for r in rows) + mix = " · ".join(f"{n} {p}" for p, n in sorted(purposes.items())) + experiment = purposes.get("experiment", 0) + print("# Adherence scoreboard\n") + print(f"{len(rows)} row(s) from {', '.join(args.files)} — {mix}.\n") + if experiment != len(rows): + print("> **Not the registered analysis.** Rows not marked " + "`purpose: experiment` are validation runs, which exist to " + "shake out the method and the harness. `make analyze` is the " + "surface that produces falsifier verdicts, and it excludes " + "them (docs/EVAL.md).\n") + print(f"warning: {len(rows) - experiment} of {len(rows)} row(s) are " + f"not purpose=experiment ({mix})", file=sys.stderr) systems = sorted({(r["model"], r["adapter"], r["arm"]) for r in rows}) for model, adapter, arm in systems: arm_block(rows, model, adapter, arm) diff --git a/src/adherence/selftest.py b/src/adherence/selftest.py index d37cea3..4713f62 100755 --- a/src/adherence/selftest.py +++ b/src/adherence/selftest.py @@ -11,15 +11,29 @@ from __future__ import annotations import collections +import io +import json +import os import random import shutil import subprocess import sys import tempfile import time +from contextlib import redirect_stderr, redirect_stdout from pathlib import Path -from adherence import REPO_ROOT, analyze, gradelib, metrics, schema +from adherence import ( + REPO_ROOT, + analyze, + calibrate, + gradelib, + metrics, + probe, + report, + schema, + suitedata, +) from adherence.runner import load_grader ROOT = REPO_ROOT @@ -554,6 +568,160 @@ def row(**kw): return problems +def _fault_grid(): + """Five trials of one cell: two harness faults, two passes, one real + failure. Chosen so the two readings are far apart and unmistakable — + 2/3 = 67% with the faults excluded, 2/5 = 40% with them scored. 40% + is inside the registered band [0.25, 0.80] and 67% is inside it too, + but the gap is the whole disagreement, so any surface reading its own + definition lands on a number no other surface produces.""" + rows = [] + for t, passed in ((2, True), (3, True), (4, False)): + rows.append(_synth("a1", "sX", t, 1000, 4, passed)) + for t in (0, 1): + r = _synth("a1", "sX", t, 0, 0, False) + r["checks"] = [{"name": "adapter", "status": "ungradeable", + "evidence": "killed at the 2700s ceiling"}] + rows.append(r) + for r in rows: + r["purpose"] = "experiment" + r["schema_errors"] = [] + return rows + + +def _proxy_for(rows): + """A proxy log agreeing exactly with the adapter on every gradeable + trial, and disagreeing wildly on the faults. Excluding the faults is + the difference between H4 passing and H4 failing.""" + out = [] + for r in rows: + fault = suitedata.is_ungradeable(r) + tok = 40_000 if fault else r["metrics"]["tok_in_billed"] + n = 8 if fault else r["metrics"]["calls"] + for _ in range(n): + out.append({"type": "call", "inference": True, "n_tools": 5, + "run_id": f"{r['scenario']}|{r['arm']}|{r['trial']}", + "input_tokens": tok // n, "output_tokens": 0}) + return out + + +def _capture(fn, argv): + """Run a __main__ that reads sys.argv, and hand back what it printed.""" + buf, err = io.StringIO(), io.StringIO() + old = sys.argv + sys.argv = argv + try: + with redirect_stdout(buf), redirect_stderr(err): + fn() + finally: + sys.argv = old + return buf.getvalue() + + +def check_surface_agreement() -> list[str]: + """Every surface that reports a rate must share one definition of + ungradeable. + + docs/HANDOFF.md §4 records this defect found in three surfaces and + fixed by one shared `suitedata.is_ungradeable()`. Two more were + carrying it and nothing noticed, because the only test pointed at + `analyze.harness_excluded` alone: + + - `report.py`, the publishable scoreboard, read 0.40 for + cli-cli-13057 against 1.00 everywhere else — and 0.40 is inside + the registered band while 1.00 is outside it. + - `calibrate.py`, the H4 gate, reported 19.571% aggregate against + 9.484% from the viewer tab reading the same two files. + + Both were written before the shared definition existed and both + survived the fix that was meant to be repo-wide. A definition that is + only shared by convention is not shared. This asserts it. + """ + problems = [] + rows = _fault_grid() + want = 2 / 3 + + # 1. the registered analysis + keep, counts = analyze.harness_excluded(rows) + if counts["adapter"] != 2 or analyze.pass_rate(keep, "a1") != want: + problems.append(f"analyze: excluded {counts['adapter']} (want 2), " + f"rate {analyze.pass_rate(keep, 'a1')} (want {want})") + + # 2. the shared loader every viewer sits on + with tempfile.TemporaryDirectory() as td: + res = os.path.join(td, "r.jsonl") + with open(res, "w") as f: + f.writelines(json.dumps(r) + "\n" for r in rows) + prx = os.path.join(td, "r.proxy.jsonl") + with open(prx, "w") as f: + f.writelines(json.dumps(r) + "\n" for r in _proxy_for(rows)) + + cells = suitedata.load_cells(paths=[res]) + if len(cells) != 1: + return problems + [f"loader produced {len(cells)} cells, want 1"] + c = cells[0] + if abs(c["pass_rate"] / 100.0 - want) > 1e-9: + problems.append(f"suitedata: pass_rate {c['pass_rate']:.2f}%, " + f"want {want*100:.2f}%") + if (c["gradeable"], c["ungradeable"], c["trials"]) != (3, 2, 5): + problems.append(f"suitedata: gradeable/ungradeable/trials = " + f"{c['gradeable']}/{c['ungradeable']}/{c['trials']}" + f"; want 3/2/5") + + # 3. report.py — the number that gets pasted somewhere + if report.pass_rate(rows) != want: + problems.append(f"report.pass_rate = {report.pass_rate(rows)}, " + f"want {want}") + out = _capture(report.main, ["report", res]) + if "0.67" not in out or "0.40" in out: + problems.append("report.main rendered a rate no other surface " + "produces; it is reading its own definition") + + # 4. probe.py — the go/no-go verdict + out = _capture(probe.main, ["probe", res]) + row = next((ln for ln in out.splitlines() if ln.startswith("sX")), "") + if "67%" not in row or row.split()[-1] != "KEEP": + problems.append(f"probe: {row!r}; want 67% over 3 trials, KEEP") + + # 5. calibrate.py — the H4 gate + out = _capture(calibrate.main, ["calibrate", res, prx]) + if "runs compared: **3** of 5" not in out: + problems.append("calibrate compared every row; a stopped trial's " + "transcript is truncated by construction and its " + "disagreement with the proxy is not an accounting " + "error") + if "aggregate delta: **0.000%**" not in out: + problems.append("calibrate: adapter and proxy agree exactly on " + "every gradeable trial, so the gate must read 0%") + + # 6. and the other side of the same invariant: a cell with NOTHING + # gradeable has no rate at all. Dividing by the survivors raised + # ZeroDivisionError out of the shared loader, which took down every + # viewer the moment a scenario's completed trials were all ceiling + # hits -- exactly when a live battery wants one. + only = os.path.join(td, "ung.jsonl") + faults = [r for r in rows if suitedata.is_ungradeable(r)] + with open(only, "w") as f: + f.writelines(json.dumps(r) + "\n" for r in faults) + try: + cells = suitedata.load_cells(paths=[only]) + except ZeroDivisionError: + return problems + ["suitedata.load_cells raised on a cell with no " + "gradeable trial; every viewer dies with it"] + if suitedata.pass_str(cells[0]) != "—": + problems.append(f"a cell with nothing gradeable rendered " + f"{suitedata.pass_str(cells[0])!r}; 0% reads as " + f"floor, and a scenario at floor gets dropped") + if report.fmt_rate(faults) != "—" or report.pass_rate(faults) is not None: + problems.append("report showed a rate for a cell where nothing " + "was graded") + out = _capture(probe.main, ["probe", only]) + if "harness broke" not in out: + problems.append("probe called an all-faults scenario a difficulty " + "result rather than a harness failure") + return problems + + def check_live() -> list[str]: """The live view reads a stream being written by another process. @@ -1144,6 +1312,7 @@ def dataset(ratio, floor=9500, passed=True, arms=("a1", "a2", "a3")): CHECKS = ( ("validation/experiment isolation", check_purpose_isolation), ("harness fault is not a model failure", check_harness_exclusion), + ("every surface shares one ungradeable", check_surface_agreement), ("live run reader", check_live), ("process hygiene", check_process_hygiene), ("live cursor anchoring", check_cursor_anchoring), diff --git a/src/adherence/suitedata.py b/src/adherence/suitedata.py index acdf44a..2caeda8 100644 --- a/src/adherence/suitedata.py +++ b/src/adherence/suitedata.py @@ -117,6 +117,18 @@ def is_ungradeable(row) -> bool: for c in row.get("checks") or []) +def pass_str(c, width: int = 0, dec: int = 0) -> str: + """pass@1 for display, or an em dash when the cell has no gradeable trial. + + A cell whose every trial was a harness fault has no pass rate. Printing + `0%` there is the same claim as scoring the fault as a model failure: + 0% reads as floor, and a scenario at floor gets dropped from the fixture + (probe.py). One formatter, so no surface can render that 0.0 as a rate + while the `ung` column beside it says nothing was graded.""" + s = "—" if not c.get("gradeable") else f"{c['pass_rate']:.{dec}f}%" + return f"{s:>{width}}" if width else s + + def _med(vals): vals = [v for v in vals if v is not None] return st.median(vals) if vals else 0 @@ -176,6 +188,15 @@ def load_cells(paths=None, pattern=None): # disagreed about whether that scenario could discriminate at all. # A viewer that can move a scenario in or out of band is not a # display bug. + # + # When NOTHING in the cell was gradeable there is no rate at all, + # and dividing by the survivors raised ZeroDivisionError out of the + # one loader every viewer sits on -- so a live battery lost `table`, + # `matrix` and `live` the moment a scenario's completed trials were + # all ceiling hits. Reproduced from the validation grid's own three + # ungradeable rows. `pass_rate` is 0.0 here so arithmetic downstream + # holds; `gradeable` is what says the 0.0 is an absence rather than + # a floor, and `pass_str` is what a reader is allowed to see. gradeable = [r for r in rs if not is_ungradeable(r)] passes = [1.0 if r["all_pass"] else 0.0 for r in gradeable] fails = sorted({c["name"] for r in rs for c in r["checks"] @@ -185,7 +206,8 @@ def load_cells(paths=None, pattern=None): "model": model, "adapter": adapter, "arm": arm, "scenario": scen, "category": rs[0].get("category", "uncategorized"), "trials": len(rs), - "pass_rate": 100.0 * sum(passes) / len(passes), + "gradeable": len(gradeable), + "pass_rate": 100.0 * sum(passes) / len(passes) if passes else 0.0, "spread": 100.0 * st.pstdev(passes) if len(passes) > 1 else 0.0, "tok": _med(toks), "ktok": _med(toks) / 1000.0, @@ -278,19 +300,27 @@ def facet_values(cells, facet): def arm_rollup(cells): """Per-arm aggregate. `pass_rate` is unweighted across scenarios so a scenario with more trials does not dominate — every scenario is one - cluster (§11).""" + cluster (§11). + + A scenario with no gradeable trial contributes nothing to the rate. It + is not a 0% cluster; averaging its placeholder 0.0 in would move the + arm's headline number by a harness fault, which is the same exclusion + criterion one level up (docs/EVAL.md).""" by = defaultdict(list) for c in cells: by[c["arm"]].append(c) out = [] for arm, cs in sorted(by.items()): + rated = [c for c in cs if c["gradeable"]] out.append({ "arm": arm, "name": ARMS.get(arm, ("?", ""))[0], "role": ARMS.get(arm, ("?", ""))[1], "scenarios": len(cs), + "gradeable": sum(c["gradeable"] for c in cs), "trials": sum(c["trials"] for c in cs), - "pass_rate": sum(c["pass_rate"] for c in cs) / len(cs), + "pass_rate": (sum(c["pass_rate"] for c in rated) / len(rated) + if rated else 0.0), "tok": _med([c["tok"] for c in cs]), "ktok": _med([c["tok"] for c in cs]) / 1000.0, "calls": _med([c["calls"] for c in cs]), @@ -323,8 +353,12 @@ def pareto_front(cells, cost="ktok", quality="pass_rate"): §5 control 2: the honest output is the (cost, success) plane, not a single number. A cell that is cheaper AND passes less is a trade, and - the frontier is what says so where a ratio would not.""" - pts = [c for c in cells if c[cost] > 0] + the frontier is what says so where a ratio would not. + + Cells with no gradeable trial are not points on this plane: their + quality axis is absent, not zero, and admitting them puts a cheap + 0%-looking cell where nothing was measured.""" + pts = [c for c in cells if c[cost] > 0 and c.get("gradeable")] front = set() for c in pts: if not any(o[cost] <= c[cost] and o[quality] > c[quality] for o in pts): diff --git a/src/adherence/table.py b/src/adherence/table.py index 681ff5f..798a8c8 100644 --- a/src/adherence/table.py +++ b/src/adherence/table.py @@ -25,7 +25,12 @@ G = Y = R = C0 = DIM = B = X = "" -def pcol(p): +def pcol(c): + """Colour for a pass@1 cell. Takes the cell, not the number, so a cell + with nothing gradeable is dimmed rather than painted red at 0%.""" + if not c.get("gradeable"): + return DIM + p = c["pass_rate"] return G if p >= 80 else Y if p >= 25 else R @@ -64,7 +69,7 @@ def main(): f"{'if pass':>10}{'calls':>7}{'probes':>8}{'abnd':>6} failing{X}") for c in cells: print(f"{c['tag']:<20}{c['trials']:>7}" - f"{pcol(c['pass_rate'])}{c['pass_rate']:>7.0f}%{X}" + f"{pcol(c)}{sd.pass_str(c, 8)}{X}" f"{C0}{c['tok']:>10,.0f}{X}{DIM}{c['tok_won']:>10,.0f}{X}" f"{c['calls']:>7.0f}{c['probes']:>8.0f}" f"{(R if c['abandoned'] else DIM)}{c['abandoned']:>6}{X}" @@ -83,7 +88,7 @@ def main(): rt = f"{tr:>10.3f}×" if tr else f"{'—':>11}" rc = f"{cr:>10.3f}×" if cr else f"{'—':>11}" print(f"{B}{r['arm']:<5}{X}{r['name']:<20}{r['scenarios']:>5}" - f"{pcol(r['pass_rate'])}{r['pass_rate']:>7.0f}%{X}" + f"{pcol(r)}{sd.pass_str(r, 8)}{X}" f"{C0}{r['tok']:>10,.0f}{X}{r['calls']:>7.0f}{rt}{rc}" f" {DIM}{r['role']}{X}") @@ -94,7 +99,7 @@ def main(): for c in sorted((c for c in cells if c["tag"] in front), key=lambda c: -c["pass_rate"]): print(f" {G}*{X} {c['tag']:<20}" - f"{pcol(c['pass_rate'])}{c['pass_rate']:>6.0f}%{X}" + f"{pcol(c)}{sd.pass_str(c, 7)}{X}" f"{C0}{c['tok']:>10,.0f}{X} tokens" f"{c['calls']:>6.0f} calls")