Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
e35a538
fix: a liveness check that killed the process it checked (Windows)
TGPSKI Aug 4, 2026
9088761
calib tab: say why it is empty, and point at something that exists
TGPSKI Aug 4, 2026
88ebb04
proxy: attribute calls by request path, not by shared state
TGPSKI Aug 4, 2026
b740880
proxy on by default; live view gains graded + rollup sections
TGPSKI Aug 4, 2026
525d2b0
subagent tokens are readable live; add an activity feed with tool output
TGPSKI Aug 4, 2026
c4297c7
activity feed: navigable list, indexed newest-first, expand on demand
TGPSKI Aug 4, 2026
3ec0e38
a delegating trial is not a stalled one; navigate all three tables
TGPSKI Aug 4, 2026
3cf01c6
activity cursor follows the event, not the row
TGPSKI Aug 4, 2026
f0427fb
detect tasks that are unpassable by construction; depth-accurate legend
TGPSKI Aug 4, 2026
7dd74e9
grade feature PRs at the CLI boundary instead of pre-loading the shape
TGPSKI Aug 4, 2026
1d0f1dd
fix a flaky selftest and the invisible cursor in two live tables
TGPSKI Aug 4, 2026
2d5e04f
Amendment 2: CLI-boundary grading, wired end to end
TGPSKI Aug 4, 2026
1b4cc94
live view: restore the blank line between summary and graded
TGPSKI Aug 4, 2026
8f1a450
shift-tab cycles views backward
TGPSKI Aug 4, 2026
0d86f23
orientation tabs: what the tasks ask for, what the arms are
TGPSKI Aug 4, 2026
fb5e782
make stop: reclaim a run, and deliberately never delete results
TGPSKI Aug 4, 2026
0268002
sorts belong to the tab you are on
TGPSKI Aug 4, 2026
ddbc8d9
write each trial's result when it finishes, not in submission order
TGPSKI Aug 4, 2026
e53c8e5
tab bar: run data, then a divider, then the reference pages
TGPSKI Aug 4, 2026
b082bfe
proxy: a client hanging up is traffic, not a fault
TGPSKI Aug 4, 2026
4d42bc9
an arm that gives up looks cheaper; say so on every surface
TGPSKI Aug 4, 2026
ae55543
the prompt never asked for anything
TGPSKI Aug 4, 2026
c9fd9ce
a finished trial is not 'grading', and its deadline no longer applies
TGPSKI Aug 4, 2026
f428c79
selftest: assert the new budget contract, both directions
TGPSKI Aug 4, 2026
44cf963
watch the WAL: the store's mtime lies between checkpoints
TGPSKI Aug 4, 2026
1c573d7
docs: session log — how this harness was built and what it got wrong
TGPSKI Aug 4, 2026
2d9bfd5
salvage a killed trial's transcript; a cut-off run did not give up
TGPSKI Aug 4, 2026
f9d6130
git decides whether an edit happened, not the transcript
TGPSKI Aug 4, 2026
3c719f0
live view: running and summary shown whole, graded absorbs the scroll
TGPSKI Aug 4, 2026
8a5f00b
graded table: 'ungradeable' did not fit its column, and no gap above …
TGPSKI Aug 4, 2026
c37090a
graded table: when each trial started
TGPSKI Aug 4, 2026
60118cd
calib was failing on truncated transcripts and hiding a real gap
TGPSKI Aug 4, 2026
6cc4bf5
say why a row is ungradeable, and stop advertising a key that does no…
TGPSKI Aug 4, 2026
df1ba6f
close a grader hole, and screen for what actually decided this fixture
TGPSKI Aug 4, 2026
78216e2
ignore _private
TGPSKI Aug 4, 2026
b2d1db4
the 'left' column counted cells that had reported, not cells that exist
TGPSKI Aug 4, 2026
ffa6044
the viewer scored harness faults as model failures, and it moved a sc…
TGPSKI Aug 4, 2026
fe006e5
validation summary: 120 runs, one arm, what it licenses and what it d…
TGPSKI Aug 5, 2026
8b239a8
validation summary: weave in the TUI screenshots
TGPSKI Aug 5, 2026
a7c67f6
fix the flaky selftest, and the go/no-go verdict it was masking
TGPSKI Aug 5, 2026
34f2dc1
selftest: wait for the proxy log to settle instead of racing it
TGPSKI Aug 5, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,17 @@ jobs:
strategy:
fail-fast: false
matrix:
# Floor and ceiling, deliberately. 3.10 is the oldest CPython still
# receiving security fixes and what older LTS distros ship, so it is
# the realistic worst case for someone running this harness on the
# box they already have; 3.14 is what the reference runs use. The
# suite is stdlib-only and uses no 3.11+ syntax, so the floor costs
# nothing to hold.
#
# Keep both. The one failure this matrix has caught was a race in
# the proxy-attribution selftest that is wrong on every version --
# 3.10's http.client keep-alive behaviour just made it likely enough
# to see. A narrower matrix would have hidden it.
python-version: ["3.10", "3.14"]

steps:
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -74,3 +74,4 @@ proxy*.jsonl
scoreboard.md

*.bak
_private/
24 changes: 23 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ RUNNER = $(PYTHON) -m adherence.runner --adapter $(ADAPTER) --model $(MODEL) \
--trials $(TRIALS) $(ARMFLAGS) --out $(OUT)

.PHONY: help check ci-local compile selftest schema lint table matrix report \
analyze floors live calibrate screen mkpr mkscenarios trees probe run all clean
analyze floors live tasks design stop clean-runs calibrate screen mkpr mkscenarios trees probe run all clean

help:
@printf '%s\n' \
Expand All @@ -28,7 +28,14 @@ help:
' make ci-local [JOB=lint] run CI'"'"'s own steps locally, from the workflow file' \
'' \
'viewing — read-only, never spends a GPU-second:' \
' make tasks what each scenario asks for and how' \
' it is judged (no run data)' \
' make design what each arm is + the ground rules' \
' make live what is running right now, one-shot' \
' make stop [YES=1] stop a run, kill orphans, sweep temp' \
' (prints a plan unless YES=1; never' \
' deletes results -- rm those yourself)' \
' make clean-runs delete sandboxes/out-dirs from dead runs' \
' make table [FILTER=a3/*] one-shot snapshot (NOCOLOR=1 to strip ANSI)' \
' make matrix [FILTER=a3/*] interactive results matrix' \
' FILES= scopes to one run (default: all of' \
Expand Down Expand Up @@ -84,9 +91,24 @@ lint:

# --- viewing ---

## Stop a run and reclaim what it left. Never deletes results.
stop:
@$(PY) -m adherence.stoprun --out $(or $(OUT),runs/probe.jsonl) $(if $(YES),--yes,)

tasks:
@$(PY) -m adherence.tasks

design:
@$(PY) -m adherence.design

live:
@$(PY) -m adherence.live

## Delete sandboxes and out-dirs left by finished or killed runs.
## Refuses while anything is running.
clean-runs:
@$(PY) -m adherence.live --clean

table:
@FILTER="$(FILTER)" REF="$(REF)" PROXY="$(PROXY)" FILES="$(FILES)" \
$(PY) -m adherence.table $(FILTER)
Expand Down
10 changes: 10 additions & 0 deletions adapters/opencode.sh
Original file line number Diff line number Diff line change
Expand Up @@ -40,10 +40,20 @@ fi
# step_finish per inference call, with per-call tokens and cache fields
# (design §3.1). Verified on opencode 1.18.10. The human-readable format
# carries none of that.
# Salvage mode. When a trial is killed at its deadline the harness loses
# the adapter, and with it the conversion step -- so a run that produced
# 27 MB of events and worked for 45 minutes was recorded as calls=0,
# tok=0. The runner re-invokes this script with ADH_SALVAGE=1 to convert
# whatever the stream captured, without generating anything new.
if [ "${ADH_SALVAGE:-0}" = "1" ]; then
RUN_EXIT=0
else
RUN_EXIT=0
opencode run --format json -m "$MODEL" "${AGENT_ARGS[@]}" "$(cat "$PROMPT_FILE")" \
> "$OUT/stdout.txt" 2> "$OUT/stderr.txt" || RUN_EXIT=$?

fi

if [ "$RUN_EXIT" -ne 0 ]; then
echo "opencode-adapter: opencode run exited $RUN_EXIT" >&2
echo "--- stdout ---" >&2; tail -n 40 "$OUT/stdout.txt" >&2 || true
Expand Down
26 changes: 26 additions & 0 deletions bench/isolate.sh
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,32 @@ ROOT="$(cd "$HERE/.." && pwd)"
BENCH_XDG="${ADH_BENCH_XDG:-$HERE/.xdg}"
PORT="${ADH_PROXY_PORT:-8010}"

# The proxy records the tokens and round trips this suite exists to
# measure, and design §3.2 makes it AUTHORITATIVE over the adapter -- so
# it runs by default. It used to be opt-in behind ADH_PROXY_LOG, which
# meant the ordinary path produced numbers nothing had verified and the
# H4 agreement gate could not be checked at all for the run you were
# actually doing.
#
# The log is paired to the results file rather than shared, so a proxy log
# always belongs to exactly one run: runs/probe.jsonl -> runs/probe.proxy.jsonl.
# ADH_NO_PROXY=1 opts out (no endpoint, or deliberately measuring the
# adapter alone).
if [ -z "${ADH_PROXY_LOG:-}" ] && [ "${ADH_NO_PROXY:-0}" != "1" ]; then
_out=""
_next=0
for _a in "$@"; do
if [ "$_next" = "1" ]; then _out="$_a"; _next=0; fi
if [ "$_a" = "--out" ]; then _next=1; fi
done
if [ -n "$_out" ]; then
ADH_PROXY_LOG="${_out%.jsonl}.proxy.jsonl"
else
ADH_PROXY_LOG="$ROOT/runs/proxy.jsonl"
fi
export ADH_PROXY_LOG
fi

mkdir -p "$BENCH_XDG/opencode"
cp "$HERE/opencode-bench.json" "$BENCH_XDG/opencode/opencode.json"

Expand Down
111 changes: 104 additions & 7 deletions docs/EVAL.md
Original file line number Diff line number Diff line change
Expand Up @@ -429,18 +429,25 @@ absent entirely on some runs. Real spend, but not attributable to the
instruction surface, so they are excluded from arm comparisons and reported
separately. Stating the rule beats quietly picking a total.

## Two costs of parallelism
## The cost of parallelism

`--jobs > 1` buys throughput and spends two things:
`--jobs > 1` buys throughput and spends one thing:

- **Wall-clock stops being comparable across arms.** Latency becomes a function
of GPU scheduling. Contended runs are stamped `contended: true` rather than
left to be compared against serial ones later.
- **Per-trial proxy attribution becomes impossible.** The proxy mark is one
piece of state and *nothing in an inference request identifies the trial* —
measured: the sandbox path appears nowhere in the request body, not even in
the system prompt. The runner skips marks and warns rather than writing an
attribution that is wrong. **Calibration runs serially.**

~~**Per-trial proxy attribution becomes impossible.**~~ This was recorded as the
second cost and it was not a cost, it was a missing key. The premise was right —
*nothing in an inference request body identifies its trial*, measured: the
sandbox path appears nowhere in it, not even in the system prompt — but the
request **path** is ours. Each trial now routes through
`<proxy>/__run/<run_id>/v1`, the proxy strips the prefix before forwarding, and
attribution arrives with the call instead of being read off one piece of shared
state that two concurrent trials would fight over. The H4 gate is measurable at
any `--jobs`. Verified in the selftest by driving two trials at one proxy
concurrently and asserting no call lands in the wrong bucket, and that upstream
never sees the prefix.

Concurrent `opencode run` against one `XDG_DATA_HOME` also dies with `database
is locked`, so each parallel trial gets its own data home.
Expand Down Expand Up @@ -539,6 +546,96 @@ dropped with the reason logged.
This was underspecified in v1. Pinning it now, before data, is the point of
saying so here rather than deciding it when the first results look wrong.

## Amendment 2 — grading a feature PR without handing over its shape

Recorded 2026-08-04, against the frozen v1 tag, **after the validation grid
and before any experiment data**. It changes how a subset of tasks is graded.
No registered claim, threshold, arm, or analysis procedure changes.

### What the validation grid found

Amendment 1 registered SWE-bench-style grading: the agent works from the PR's
parent, never sees the tests, and afterwards the harness applies the PR's own
`_test.go` files and runs the affected packages.

That protocol is sound for a **bug fix**, whose tests call API that already
exists. It is not sound for a **feature PR**, whose tests call API the PR is
adding. Measured on the 210-run validation grid:

| failure mode | scenarios | pass rates | in the [0.25, 0.80] band |
|---|---|---|---|
| tests name a symbol the fix introduces | 4 | 0%, 5%, 0%, 0% | **0 / 4** |
| tests compile; assertions disagree | 6 | 5%, 19%, 71%, 76%, 90%, 95% | 2 / 6 |

A perfect separation, and the dominant cause of the floor risk this document
names as its biggest schedule risk. On `cli-cli-13057` the failure was
`unknown field IssueType in struct literal of type CreateOptions` — a
**compile error in the test file** — on a trial that had already passed
`diff_coverage` with 13/13 pre-edit probes inside the real diff's
directories. The agent located the work, implemented it, and failed because
it did not independently choose the maintainers' internal field name.

That is not model capability and it is not a grader bug. It is a protocol
that asks a question the agent was never given the means to answer.

### The repair that was rejected, and why

The obvious fix is to state the required API surface in the prompt. **It was
rejected**, and the reason is recorded here because it is the more tempting
option:

A symbol list *is* routing information. `CreateOptions.IssueType` names its
subsystem; `skillSearchFunc` names its module. Supplying it hands every arm a
piece of exactly what the treatment under test is supposed to supply — inside
the treatment's own channel. It would bias the primary outcome toward the
null and make a null result uninterpretable, because "the pattern did not
help" and "the control was given the answer" would be the same measurement.

Grading at the public interface was preferred but has **no supply here**: of
34 tasks, 30 are unit-test-only and none are acceptance-test-only.

### What is registered instead

Tasks whose tests cannot compile against a differently-named implementation
are graded at the **command-line boundary**, against the PR's own binary:

reference = build(merge commit) what the PR actually shipped
candidate = build(the agent's tree) what the agent produced
compare flag-for-flag at the CLI

`gh issue create --type` is a user-facing contract. It already appears in the
PR body the agent receives, so **nothing is added to the prompt** — the
grader stops demanding an identifier it never disclosed. The oracle is the
merge commit's own compiled binary, so the expectations remain the
maintainers' and not the experimenter's, and the command path is derived
mechanically from the PR's file paths rather than chosen.

**Assignment is by compiler, not by judgement.** `adherence.classify` checks
the PR's tests out onto the parent tree and builds them. A build error naming
an undefined identifier routes the task to the CLI grader; anything else
stays with the unit grader. The choice and the identifiers that forced it are
written into each task record. A declaration-scan heuristic was tried first
and rejected for false negatives: it missed `13675`
(`field.onSearchDone.Store undefined` — a field on an existing struct) and
`13624`, neither of which adds a top-level declaration.

### Limits, and how results are reported

`cli.surface` verifies that the flags the PR shipped are present and
accepted. It does **not** verify they behave correctly, so it is a weaker
signal than a passing unit test.

Therefore: **unit-graded tasks are the primary evidence and CLI-graded tasks
are reported as a separate tier.** The two are never pooled into one pass
rate. A cost comparison conditioned on success uses the unit-graded set
unless a result is stated explicitly as spanning both.

Verified before registration on a real trial from the validation run — one
that scored `pr.task_pass = fail` purely for naming a struct field
differently: `cli.builds` pass, `cli.surface` pass (all 17 flags the PR's own
binary advertises), `cli.extra_surface` pass.


## Known gaps

- **E5 is not testable locally.** vLLM returns `prompt_tokens_details: null`, so
Expand Down
Loading