Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ src/adherence/
suitedata.py one loader behind every viewing surface
table.py one-shot snapshot · matrix_tui.py interactive matrix
adapters/ opencode (stream + export + child sessions), api.py
tui/ VENDORED from leather — re-copy, never edit here
tui/ VENDORED from pane — re-copy, never edit here
adapters/opencode.sh the shell adapter; the `--adapter` contract
scenarios/sNN/ scenario.yaml + fixture/ + grade.py
bench/ sandbox layer: isolate (+ proxy), preflight, prewarm
Expand Down
27 changes: 27 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,33 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed

- **`report.py` and `calibrate.py` scored harness faults as model results.**
Both predate the shared `suitedata.is_ungradeable()` and neither was
updated when it landed. `make report` read pass@1 0.40 for
`cli-cli-13057` against 1.00 from `probe.py`, `analyze.py` and the
viewer — 0.40 is inside the registered calibration band and 1.00 is
outside it. `make calibrate`, the H4 gate, reported 19.571% aggregate
divergence over 120 runs against the viewer tab's 9.484% over 117, for
the same two files. Both now import the one definition. (#9, #10)
- **`make selftest` now enforces that they cannot drift again**, by running
one synthetic cell of two harness faults, two passes and one real failure
through `analyze`, `suitedata`, `report`, `probe` and `calibrate` and
requiring the same rate from each. Every one of the five defects was
re-introduced and confirmed to fail it. (#11)
- **`suitedata.load_cells` raised `ZeroDivisionError`** on a cell whose
trials were all ungradeable, taking down `table`, `matrix` and `live`
the moment a scenario's completed trials were all ceiling hits. Cells
now carry `gradeable`, and one `pass_str` renders the absence as an em
dash rather than a 0% that reads as floor. (#8)
- `report.py` states the `purpose` mix of the rows it read, in the
document and on stderr. (#12)
- `report.py` and `calibrate.py` no longer raise `KeyError` on rows
written before `duration_s` and `arm` existed. (#13)
- `bench/ci-local.sh` generated a shell script at a predictable path in
`/tmp` and then executed it; now `mktemp -d` with a cleanup trap. (#14)

### Changed

- `src/adherence/tui/` is now vendored from its canonical public
Expand Down
17 changes: 12 additions & 5 deletions bench/ci-local.sh
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,14 @@ HERE="$(cd "$(dirname "$0")/.." && pwd)"
JOB="${1:-${JOB:-validate}}"
cd "$HERE"

# The generated steps are written and then executed, so the path they are
# written to must not be one another user can pre-create as a symlink or
# swap between the write and the `bash`. `mktemp` for the same reason the
# `python` shim below already uses it.
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
STEPS_SH="$WORK/steps.sh"

mapfile -t STEPS < <(python3 - "$JOB" <<'PY'
import json, re, sys

Expand Down Expand Up @@ -86,7 +94,7 @@ print(json.dumps(steps))
PY
)

python3 - "$JOB" "${STEPS[0]:-[]}" <<'PY' > /tmp/adh-ci-steps.sh
python3 - "$JOB" "${STEPS[0]:-[]}" <<'PY' > "$STEPS_SH"
import json, shlex, sys
job, blob = sys.argv[1], sys.argv[2]
steps = json.loads(blob)
Expand All @@ -107,12 +115,11 @@ PY

# `python` is the interpreter name on a GitHub runner; make it resolvable here.
if ! command -v python >/dev/null 2>&1; then
SHIM="$(mktemp -d)"
ln -s "$(command -v python3)" "$SHIM/python"
export PATH="$SHIM:$PATH"
ln -s "$(command -v python3)" "$WORK/python"
export PATH="$WORK:$PATH"
fi

bash /tmp/adh-ci-steps.sh
bash "$STEPS_SH"
RC=$?
echo
[ "$RC" -eq 0 ] && echo "ci-local($JOB): all steps passed" \
Expand Down
36 changes: 26 additions & 10 deletions docs/HANDOFF.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ the registered experiment has not been run.

| | |
|---|---|
| Branch | `main` at `b4ae1ca`, clean, nothing unpushed |
| Branch | `main` at `eba93f6`, clean, nothing unpushed |
| CI | **green** — `Validate (3.10)`, `Validate (3.14)`, `Lint`, and all three `Full-scope` platforms |
| Selftest | 25/25 |
| PRs | #1, #2, #4 merged. #3 closed deliberately (see §4) |
Expand Down Expand Up @@ -57,12 +57,19 @@ pool them is doing real work — pooled, the CLI ceiling masks a healthy unit
tier.

**H4 (proxy vs adapter) fails as specified, and the failure is localized.**
19.6% aggregate, but **22 of 24 scenarios agree exactly** — 0.0% on tokens
9.5% aggregate over the 117 gradeable trials, but **22 of 24 scenarios agree
exactly** — 0.0% on tokens
*and* identical call counts — and all six in-band scenarios are among them. The
divergence is transcript truncation on ceiling-hit trials, not an accounting
error. Cost figures must still come from `runs/probe.proxy.jsonl` per the
registered rule.

This read **19.6% over 120 runs** until `calibrate.py` was made to apply
exclusion criterion 1, which `suitedata.calibration` had always applied. The
gate and its own viewer tab disagreed by 2× on the same two files (#10). H4
fails either way — call counts are 5,319 adapter against 5,646 proxy — but the
number to plan against is 9.5%.

**Nothing here licenses a statement about the directed-contexts pattern.** One
arm ran. `make analyze` correctly excludes all 120 rows as `purpose: validation`
— that guard is working, not failing.
Expand Down Expand Up @@ -139,24 +146,32 @@ task wrong, which the registration forbids in as many words.
analysis**. 40% is inside the band; 100% is outside it. **The viewer was
inflating the number that decides whether the experiment can run.**

### The same defect turned up in three surfaces
### The same defect turned up in five surfaces

Three were found in this session: `suitedata` (the viewer), `probe.py` (the
go/no-go verdict), and `analyze.py` (which had always been correct). The
disagreement was silent until someone compared them.

`suitedata` (the viewer), `probe.py` (the go/no-go verdict), and `analyze.py`
(which had always been correct). The disagreement was silent until someone
compared them.
Two more were found afterwards, in the review of PR #5, and are fixed on
`main`: `report.py` (the publishable scoreboard, #9) and `calibrate.py` (the H4
gate, #10). Both predate the shared definition and both survived the fix that
was meant to be repo-wide, because the only test pointed at `analyze` alone.
`make selftest` now asserts that every surface reporting a rate lands on the
same number (#11).

`probe.py` was the worst place to hold it, because it renders a *decision*:

```
cli-cli-13057 40% 5 KEEP → cli-cli-13057 100% 2 ceiling — drop
7 of 24 tasks land in [0.25, 0.80] → 6 of 24 tasks land in [0.25, 0.80]
16 ceiling · 1 floor · 0 harness
16 ceiling · 1 floor · 0 harness → 17 ceiling · 1 floor · 0 harness
```

It also matched adapter faults on status `"fail"` when a ceiling hit records
`"ungradeable"` — printing `0 harness` for a run with three killed trials.

There is now **one `is_ungradeable()`** in `suitedata`, imported by the others.
There is now **one `is_ungradeable()`** in `suitedata`, imported by the others,
and one selftest check that fails if a surface stops importing it.

### The `left` column counted cells that had reported, not cells that exist

Expand Down Expand Up @@ -241,5 +256,6 @@ it originated.

The run-watcher lineage was corrected twice during the session and now reads:
five generations over 25 days across three codebases, beginning not in an eval
but in a pair of live server-analytics dashboards whose curses framework both
this repo and leather vendor byte-identically.
but in a pair of live server-analytics dashboards whose curses framework this
repo vendors byte-identically from [pane](https://github.com/TGPSKI/pane), its
canonical upstream since `2c0f9d9`. It was copied from `leather` before that.
31 changes: 28 additions & 3 deletions src/adherence/calibrate.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,9 @@
separately. The adapter cannot see them and they are not attributable to
the instruction surface under test. Excluding them is a stated choice,
not a rounding decision — see lib/metrics.is_auxiliary.

Trials the harness stopped are excluded too, under docs/EVAL.md exclusion
criterion 1 — see `main`.
"""
from __future__ import annotations

Expand All @@ -27,6 +30,7 @@
from pathlib import Path

from adherence import metrics
from adherence.suitedata import is_ungradeable


def load_jsonl(path):
Expand Down Expand Up @@ -57,9 +61,23 @@ def main():
tot_ac = tot_pc = 0
worst = 0.0
unmatched = []
n = 0
n = excluded = 0
for r in results:
mark = f"{r['scenario']}|{r['arm']}|{r['trial']}"
# docs/EVAL.md exclusion criterion 1, which suitedata.calibration
# applied and this gate did not. A trial the harness stopped has a
# transcript truncated by construction, so against the proxy's
# complete record it reads as a ~100% accounting error -- which is
# not what H4 measures, and it buries the disagreements that are
# real. Measured on runs/probe.jsonl: 19.571% aggregate over 120
# runs here against 9.484% over 117 in the viewer, for the same two
# files. Two surfaces disagreeing by 2x about a registered quantity
# is the defect, not the number.
if is_ungradeable(r):
excluded += 1
continue
# `arm` predates nothing here, but rows written before the arms work
# carry none; analyze.load and suitedata.load_rows both default it.
mark = f"{r['scenario']}|{r.get('arm', '-')}|{r['trial']}"
rows = by_mark.get(mark)
m = r.get("metrics") or {}
if rows is None:
Expand All @@ -78,7 +96,14 @@ def main():
f"| {p_tok} | {d*100:.2f}% | {p['aux_calls']} |")

agg = abs(tot_a - tot_p) / tot_p if tot_p else 1.0
print(f"\nruns compared: **{n}**")
print(f"\nruns compared: **{n}** of {len(results)}")
if excluded:
# Counted, never silent: an exclusion nobody can see is
# indistinguishable from data that never existed.
print(f"excluded: **{excluded}** harness fault(s) "
f"(docs/EVAL.md exclusion criterion 1 — a stopped trial's "
f"transcript is truncated by construction and its "
f"disagreement with the proxy is not an accounting error)")
print(f"adapter Σ input tokens: **{tot_a}** · "
f"proxy Σ input tokens: **{tot_p}**")
print(f"aggregate delta: **{agg*100:.3f}%** · "
Expand Down
28 changes: 17 additions & 11 deletions src/adherence/matrix_tui.py
Original file line number Diff line number Diff line change
Expand Up @@ -315,8 +315,14 @@ def visible(self):

# ---- chrome ----------------------------------------------------------

def pass_attr(self, p):
def pass_attr(self, c):
"""Attribute for a pass@1 cell. Takes the cell, not the number: a
cell with nothing gradeable has no rate, and painting its
placeholder 0.0 red says the model failed where the harness did."""
C = self.curses
if not c.get("gradeable"):
return C.A_DIM
p = c["pass_rate"]
return (C.color_pair(1) if p >= 80 else
C.color_pair(3) if p >= 25 else C.color_pair(4))

Expand Down Expand Up @@ -645,8 +651,7 @@ def view_live(self, body, max_x, max_y=0):
self._put(y, 25, f"{c['trials']:>6}", C.A_DIM)
self._put(y, 31, f"{c['ungradeable']:>5}",
C.color_pair(3) if c["ungradeable"] else C.A_DIM)
self._put(y, 36, f"{c['pass_rate']:>5.0f}%",
self.pass_attr(c["pass_rate"]))
self._put(y, 36, sd.pass_str(c, 6), self.pass_attr(c))
skewed = (c["abandoned"] and c["tok_worked"]
and c["tok_worked"] > c["tok"] * 1.05)
self._put(y, 42, f"{c['tok']:>11,.0f}"
Expand Down Expand Up @@ -1345,8 +1350,7 @@ def view_cells(self, cs, body, max_x):
base = C.color_pair(5) if idx == self.cursor else 0
self._put(y, 1, f"{c['tag']:<20}", base)
self._put(y, 21, f"{c['trials']:>7}", C.A_DIM)
self._put(y, 28, f"{c['pass_rate']:>7.0f}%",
self.pass_attr(c["pass_rate"]))
self._put(y, 28, sd.pass_str(c, 8), self.pass_attr(c))
self._put(y, 36, f"{c['spread']:>5.0f}", C.A_DIM)
self._put(y, 41, f"{c['tok']:>9,.0f}", C.color_pair(2))
self._put(y, 50, f"{c['tok_won']:>9,.0f}", C.A_DIM)
Expand All @@ -1373,8 +1377,7 @@ def view_arms(self, cs, body, max_x):
self._put(y, 6, f"{r['name']:<20}")
self._put(y, 26, f"{r['scenarios']:>5}", C.A_DIM)
self._put(y, 31, f"{r['trials']:>7}", C.A_DIM)
self._put(y, 38, f"{r['pass_rate']:>7.0f}%",
self.pass_attr(r["pass_rate"]))
self._put(y, 38, sd.pass_str(r, 8), self.pass_attr(r))
self._put(y, 46, f"{r['tok']:>10,.0f}", C.color_pair(2))
self._put(y, 56, f"{r['calls']:>7.0f}", C.A_DIM)
if r["arm"] == self.ref:
Expand Down Expand Up @@ -1417,7 +1420,11 @@ def view_cost(self, cs, body, max_x):
from another that is a trade, not a win, and the chart says so
where a ratio would not."""
C = self.curses
pts = [c for c in cs if c["ktok"] > 0]
# Both axes have to exist. A cell with no gradeable trial has no
# quality coordinate, and plotting its placeholder 0.0 puts a point
# at the bottom of the chart where nothing was measured -- and drags
# the y-axis floor down with it.
pts = [c for c in cs if c["ktok"] > 0 and c["gradeable"]]
if not pts:
self._put(3, 3, "no cells with token telemetry in this selection",
C.A_DIM)
Expand Down Expand Up @@ -1455,8 +1462,7 @@ def view_cost(self, cs, body, max_x):
for i, c in enumerate(best[:max(0, body - ph - 6)]):
yy = y0 + 1 + i
self._put(yy, 1, f"{c['tag']:<20}")
self._put(yy, 21, f"{c['pass_rate']:>7.0f}%",
self.pass_attr(c["pass_rate"]))
self._put(yy, 21, sd.pass_str(c, 8), self.pass_attr(c))
self._put(yy, 29, f"{c['tok']:>10,.0f}", C.color_pair(2))
self._put(yy, 39, f"{c['calls']:>7.0f}", C.A_DIM)
self._put(yy, 46, f"{sd.fmt_duration(c['dur_s']):>8}",
Expand Down Expand Up @@ -1558,7 +1564,7 @@ def view_detail(self, c, max_y, max_x):
y += 1
lines = [
("trials", f"{c['trials']}"),
("pass@1", f"{c['pass_rate']:.0f}% (± {c['spread']:.0f})"),
("pass@1", f"{sd.pass_str(c)} (± {c['spread']:.0f})"),
("median tok_in | passed", f"{c['tok_won']:,.0f}"),
("median tok_in | worked", f"{c['tok_worked']:,.0f}"
+ (" (abandoned trials excluded — "
Expand Down
Loading