Harden live TUI, process lifecycle, and exclusion criteria - #1
Merged
Merged
Conversation
os.kill(pid, 0) is the standard existence probe on POSIX, where signal 0 means 'check, do not send'. On Windows it is not a probe: CTRL_C_EVENT is 0, so the call delivers a real Ctrl-C to the target's console group. check_live wrote a marker carrying its own pid; snapshot() checked it; the check raised KeyboardInterrupt inside an unrelated subprocess several tests later. The traceback pointed at check_test_runner, which had nothing to do with it, sixteen seconds into the job. Green through 6921707, first red at af026e8 -- the live dashboard commit, and only ever on windows/amd64. Lint, both Validate jobs, linux/arm64 and macos/arm64 were all green throughout. _alive now declines to answer off POSIX and the caller falls back to evidence that cannot misfire. Guarded in the selftest, because the failure mode of getting this wrong is a viewer that terminates what it observes.
The message told you to 'run under bench/with-proxy.sh'. That script does not exist and never did -- the proxy starts from ADH_PROXY_LOG inside bench/isolate.sh. It also omitted the reason the tab is empty for most runs. The proxy holds one trial mark at a time and nothing in an inference request identifies its trial, so under --jobs>1 the runner skips the mark rather than writing a wrong attribution. The H4 gate cannot be measured in parallel at all, which makes 'pass --proxy' useless advice for exactly the runs people are watching when they find this tab. Now prints the two commands that work, states that JOBS=1 is a precondition rather than a preference, and names it as the second cost of parallelism (§16.4).
The registration listed 'per-trial proxy attribution becomes impossible' as the second cost of parallelism. It was not a cost, it was a missing key. The premise was right and measured: nothing in an inference request body identifies its trial -- the sandbox path is not in it, not even in the system prompt. So the runner marked trial boundaries by POSTing a label to the proxy, which is one piece of shared state; two concurrent trials write into whichever label was set last. Rather than record an attribution it knew was wrong, the runner skipped marking under --jobs>1 -- which meant the H4 gate, the thing that decides whether any cost number in this suite can be trusted, was unavailable for every parallel run. The request PATH is ours. Each trial now gets a config whose provider baseURL is <proxy>/__run/<run_id>/v1; the proxy strips the prefix before forwarding and stamps the record. Attribution arrives with the call. Upstream sees exactly what it saw before, because a proxy that alters the request is not measuring it. Selftest drives two trials at one proxy concurrently with overlapping requests and asserts 4 and 6 calls land in their own buckets, tokens survive, and upstream saw only /v1/chat/completions. 22/22. marks are kept for serial runs and for reading existing logs; run_id wins wherever present.
Three things the TUI surfaced by being looked at. SUBAGENTS READ 0. The live view counted distinct sessionIDs in the root stream, which is always 1: a subagent runs in its OWN opencode session and the root stream carries none of its calls -- measured at 3.1x under-report on s13. So a trial that had dispatched two 'explore' agents showed sub=0 while its own detail pane listed the spawns. The runner already gives each trial its own XDG_DATA_HOME, so opencode's store is in the out-dir and readable live: child sessions now come from there, read-only and best-effort. Totals are explicitly labelled ROOT SESSION ONLY when children exist, because E3 is the claim that this handoff is free and quoting a root-only total would confirm it by construction. PROXY WAS OPT-IN. It records the tokens and round trips this suite exists to measure and §3.2 makes it authoritative over the adapter -- so the ordinary path was producing numbers nothing had verified, and the H4 gate could not be checked for the run you were actually doing. It now starts by default, logs to a path paired with the results (runs/probe.jsonl -> runs/probe.proxy.jsonl) so a log always belongs to exactly one run, and the viewers find it without being told. ADH_NO_PROXY=1 opts out. PROGRESS BAR READ AS AN ELLIPSIS. '..................' at 0% looks like truncated output. Bracketed now, with a track character that cannot be mistaken for one. Live view is now three sections: running (cursor, [space] detail), graded most-recent-first with verdict and failing checks, and a per-(arm, scenario) rollup carrying pass rate, median and p90 tokens, tool averages, subagents, median and p90 duration, trials left and ETA. The cells detail pane gained a median/mean/p90 table with a p90:median ratio, flagged when the tail is 2x the middle -- a treatment that halves the median while doubling the tail has not made anything cheaper.
Two places I said 'not available' where the data was sitting on disk.
SUBAGENT TOTALS. I labelled the live totals ROOT SESSION ONLY and left it
there. The child sessions' tokens are in the same sqlite store the session
ids came from -- measured on a live trial: 15 child calls, 767,215 input
tokens the root-only figure was silently omitting. The live view now sums
parent + children, which is the only total §7 permits quoting as a cost,
and the detail pane shows the split plus the subagents' share, because E3
is a claim about the child half specifically.
TOOL OUTPUT. The NDJSON stream carries a call's input and status but not
what it returned, so the view could say 'it ran go test' and never what go
test said -- the one thing worth knowing while a run is in flight. The
store keeps the output. The detail pane now ends in a live activity feed:
newest last, subagent lines tagged, failures in red, with the actual tool
output. Caught an agent's own broken grep ('Unmatched ( or \(') on the
first read.
Both read-only and best-effort against a live writer and a private schema,
like everything else in this module.
Inline tool output turned the pane into a wall nobody could scan -- and the interesting line is rarely the last one. The feed is now a list: - newest first, each event one line, no output rendered inline - indexed by absolute position in the run, so the number means how many events the trial has produced and where this one sits among them, rather than how far down the screen you are - [j/k] moves, [space] expands the selected event over the whole pane with its full output, [space] or [q] collapses - a proportional scrollbar, so 'nothing more' is distinguishable from 'top of a long list' -- which a truncated list cannot say - subagent events tagged, failures marked and coloured in both views The live view now nests three levels: run list, run detail with the activity list, one expanded event. Each level consumes its own keys so [q] backs out one step instead of quitting from three levels deep.
FALSE STALL, and it was in the measurement path. The idle detector watched stdout.txt, which carries ROOT-session events only. A trial that dispatches a subagent and waits is silent there for as long as the child works. Observed live: parent 0 calls, subagent 5 calls and 206,689 tokens, forty tool calls deep, stream untouched for 3m27s -- reported 'stalled', and with --idle-timeout 300 the runner would have killed it as hung. It would also have done that MORE to the arms that route to subagents, which are the arms under test. A harness that kills the treatment for delegating would have produced a clean, wrong result. Progress is now either the event stream or opencode's session store moving, in both the runner's deadline and the live view's verdict. NAVIGATION. Only the running table had a cursor; graded and the per-arm rollup were read-only. [h/l] now moves between the three sections and [space] opens the right kind of detail for each: a live run, a graded result with every check's evidence, or the cell's median/mean/p90. Q WAS OVERLOADED. It meant 'quit' at the top and 'back' when nested, so from three levels deep there was no way to leave except pressing it repeatedly and no way to know which press would exit. Now [q]/[esc] back out one level, [Q] quits from any depth, and the footer says which one q currently means. Also: expanded events keep their structure. Collapsing whitespace turned every JSON blob and file listing into one unreadable line -- structure is most of what a tool result is. JSON is re-indented, the author's line breaks are kept, wrapped continuations are indented past the original, and the expanded view takes the full pane instead of the eight lines left under the run header.
Spacebar into the newest event and read it, and new events arriving at the top pushed the selection backwards one row each -- so the thing being read was replaced mid-sentence by whatever had since become event N. Same class as the run-list bug fixed earlier, in the second live table. Events now carry the store's own primary key and the cursor anchors to it. act_follow keeps the useful default: before you move, the cursor tracks the newest event, which is what 'what is it doing' means for a run in flight. The first [j]/[k] pins it, and expanding pins unconditionally -- reading one event while the list follows the newest is exactly how the thing you opened disappears. A selection that rolls out of the window re-anchors rather than dangling. Covered by a selftest over the anchoring logic directly, since scraping curses redraws out of a pty proved unreliable: two new events arrive, the index slides 2 -> 4, the selected event stays put. 23/23.
POST-MORTEM. The failures are neither model failures nor grader bugs. The PR's own tests are applied to the agent's independent implementation, and for a FEATURE PR those tests reference symbols the PR introduces -- 'undefined: findCopilotBinaryFunc', 'unknown field IssueType in struct literal of type CreateOptions'. The test does not compile, so passing requires the agent to independently choose the maintainers' exact identifier. That measures name-guessing, not the work. Measured on the validation grid, and the separation is perfect: all four compile-class scenarios scored 0%, 5%, 0%, 0% -- none in the [0.25, 0.80] band -- while every in-band and every ceiling scenario was behaviour-class. It is the dominant cause of the floor problem the registration names as its biggest schedule risk, and on cli-cli-13057 the agent had already PASSED diff_coverage with 13/13 pre-edit probes landing in the right directories. It did the work and failed on a name. This is why SWE-bench-style grading works for bug fixes and not for feature PRs: a bug fix's tests call API that already exists. mkpr's fail-before check proved a task goes red at the parent but never asked why. It now classifies: tests that fail to compile because they name something the fix introduces are rejected with the identifiers listed. Free, mechanical, and it runs before any GPU time is spent. Static scan of the existing 34: 19 clean, 9 whose missing surface is <=8 symbols, 6 whose surface is 12-130 symbols and cannot be stated without handing over the design. Also: the TUI legend is now depth-accurate -- each level advertises only the keys it responds to, because a legend listing keys that do nothing is checked once and then trusted. [Q] quits from any depth, [q] backs out one level, [esc] returns to the root list regardless of nesting.
The question was whether the grader can adapt to the agent's naming
without handing over the solution. It can, and the answer is to stop
grading at the Go boundary.
Declaring the API surface in the prompt was the obvious repair and it is
worse than the disease: a symbol list IS routing information
(CreateOptions.IssueType names its subsystem), so it hands every arm a
piece of exactly what the treatment is supposed to supply. It biases the
primary outcome toward the null in the treatment's own channel, which
makes a null result uninterpretable -- 'the pattern did not help' and 'we
gave the control the answer' become the same measurement.
Grade what the PR publicly promises instead. is a
user-facing contract: it is in the PR body the agent already receives, and
it is identical whatever the agent names its internals. The oracle is the
merge commit's OWN binary, so the expectations stay the maintainers' and
not the experimenter's:
reference = build(merge commit)
candidate = build(the agent's tree)
compare flag-for-flag at the command line
Nothing is added to the prompt. The grader simply stops asking a question
it never gave the agent the means to answer.
Verified against a real agent tree from the live probe -- one that scored
pr.task_pass=fail purely for naming a struct field differently:
[pass] cli.builds
[pass] cli.surface: advertises all 17 flag(s) the PR's own binary does
[pass] cli.extra_surface
Offline and deterministic; --help and argument validation need no network,
which is what makes it runnable in the sandbox. mkpr now MARKS such tasks
rather than rejecting them, records which grader judged each one and which
identifiers forced the choice, and the command path is derived from the
PR's own file paths so the battery is mechanical rather than authored.
Recovers the 15 compile-class tasks: 34 usable, not 19.
Also: left/right arrows navigate the nested views -- right descends, left
backs out one level. At the root they stay bound to section switching,
since binding right to both would make sections unreachable.
CI was red on three commits, always Validate (3.10), always 'BAD proxy attribution under concurrency', never on 3.14. The message said the trials had cross-attributed, which would have been a real defect in the measurement path -- the proxy is authoritative for H4. It had not. Stressed at 8 concurrent trials x 20 calls: 160 requests, 160 records, zero drops, zero mismatches. The failure was 'got 5 of 6': one request raised inside a sender thread, which silently ended that thread's loop, and the test compared against a hardcoded count. A transient client-side error on 3.10's http.client was being reported as a cross-attribution bug. The test now counts what actually got through and asserts the invariant it means: every call the proxy handled is attributed to the trial that made it. A networking hiccup reduces the sample instead of inventing a defect. Separately: the graded and per-arm tables could be focused with [h/l] and their cursors moved, but nothing rendered -- no highlight, no indication of which section had focus, so both looked inert. Each section header now carries a focus marker and each table highlights its own selected row only when it holds focus.
GRADER. All 34 tasks classified by compiler, not heuristic: adherence.classify checks each PR's tests out onto its parent tree and builds them. A build error naming an undefined identifier means the test cannot compile against any implementation that chose different names. Result: 16 unit-graded, 18 cli-graded. The declaration-scan heuristic I tried first was rejected for false negatives -- it missed 13675 (field on an existing struct) and 13624 -- and the compiler catches both. Calibrated against the validation grid: it flags exactly the scenarios that scored 0% (13393, 13675, 13967) and leaves 13624 (5%) and 13403 (71%) on the unit grader, which are demonstrably achievable. prgrade now dispatches on the task record. For a cli-graded task diff_coverage becomes advisory rather than failing: cli.surface already proves the shipped behaviour flag-for-flag against the PR's own binary, and demanding the same file set would smuggle back the match-the- maintainers'-structure requirement this amendment exists to remove. Verified end to end -- three CLI passes give all_pass=True, which is what the table, the TUI and the report show. Locked in by a selftest. AMENDMENT 2 in docs/EVAL.md: what the grid measured, the repair that was rejected and why (a symbol list is routing information and would bias the primary outcome toward the null inside the treatment's own channel), what is registered instead, and the reporting rule -- unit-graded tasks are the primary evidence, cli-graded a separate tier, never pooled. TUI, from watching it run: - graded and summary were silently truncated (6 rows of 25, no scroll, no indication). Both scroll now with their own cursor and print the range. - allocation is terminal-size aware and proportional to what each section holds: 24 lines shows 5 graded rows, 34 shows 14, 50 shows all 25. - the results cursor was unbounded, so holding [j] walked it off the end of the list, the highlight vanished, and [space] opened a detail pane for a row never on screen. - [space] in a graded or summary detail opened a phantom activity level with nothing in it that took two backs to leave. - sections reordered running -> summary -> graded, [s] sorts the focused one, [S] flips direction, both shown in the header. - j/k crosses section boundaries; h/l still jumps whole sections. Also widened the process-hygiene timing: it failed twice under CPU contention from a concurrent run, reporting a defect that was really load.
The block swap that reordered the sections dropped the trailing spacer, so the two tables ran together with no gap. Also adds the range indicator the summary section was missing -- graded had one, summary did not, so a truncated summary looked complete.
[tab] moved forward through live/cells/arms/cost/calib and there was no way back except cycling all the way around. KEY_BTAB and the literal 353 some terminals send when terminfo carries no kcbt entry are both accepted, so the key is not dead on a terminal that reports it raw -- verified against the real \e[Z sequence through a pty. Switching views also resets the row cursor, which otherwise carried a position from a list of a different length.
Two new read-only views, and neither touches run data. TASKS. An operator watching a grid sees scenario ids and pass rates with no way to tell what cli-cli-13418 wants, where the answer lives, or why one task is judged by unit tests and the next at the command line. That gap is exactly where a harness artifact gets read as a model failure -- the thing Amendment 2 exists to stop. Lists every scenario with its PR, its grader, the files and directories the real diff touched, and what it asks for; [space] opens the full prompt with the grader explained in the terms a verdict needs, plus the identifiers that forced a cli grading. DESIGN. Same for the arms, measured off disk rather than described, so it cannot drift from what a trial is actually handed: always-loaded bytes, delta against a1, on-demand bytes, and a hash of the surface. a0 0 B, a1 6,518, a2 18,060, a3 3,474 + 13,850 on demand, a5 417. Carries the run's ground rules -- primary outcome, why a cost number without a pass rate is not reportable, the calibration band, harness faults never scoring as model failures, purpose stamping, the two grading tiers, the proxy being authoritative, and disclosed-not-forbidden deviation. Both detail panes are height-aware and scroll: they used to render until they ran out of screen and stop, so on a short terminal a 179-line prompt was cut with nothing saying so and no way to reach the rest. Verified at 26 lines. Also and for the same content outside the TUI.
Restarting cleanly is a five-step ritual with three ways to get it wrong, and all three have already happened here: pkill orphans the opencode grandchildren, the paired proxy log gets forgotten so the next run calibrates against the last one's calls, and the temp sandboxes stay behind so the live view keeps rendering a torn-down run. So the awkward part is one command: SIGTERM the whole process group, escalate to SIGKILL for anything that ignores it, report what is still unowned, sweep the sandboxes. It does NOT delete results, and that is the design rather than an omission. Bundling 'kill the processes' with 'remove the data' into one verb is how a habit formed for the safe purpose ends up performing the destructive one, and a --force flag exists to be used. Clearing a results file stays a manual rm -- the command prints the exact one, including the paired proxy log, and says so when the file holds rows marked purpose=experiment. Plan-only unless YES=1, so a mistyped invocation costs nothing.
[s] cycled the cells sort from every view, so the tasks tab showed 'sort:tok' above a list that [s] did not touch. Each tab now sorts by what it actually holds -- tasks by id, grader, files, dirs or PR; arms by id, always-loaded bytes or on-demand bytes -- with [S] flipping direction, and the header shows the sort that [s] would change HERE. Also from watching it: - the design tab ignored [j]/[k] entirely. _move_rows counted rows for every view except that one, so the cursor had nothing to move through. - with nothing graded yet, the live view's summary and graded sections rendered no header at all, so [h/l] moved focus onto an invisible table and the operator had no way to tell whether the section was empty or broken. Both now draw their header and say what they are waiting for. - ground-rule labels were clipped at a fixed 26 columns, mid-word: 'Harness faults are not mode'. Width now comes from the longest label.
ThreadPoolExecutor.map yields in SUBMISSION order, so a single slow trial buffers every later result inside the executor. Caught live on the probe: trial 2 of cli-cli-13057 completed, its out-dir was cleaned up, and its row sat unwritten behind two trials still 6+ minutes from finishing. That is three problems at once. The result is invisible to the live view, because the out-dir it reads is already gone. The progress count and the graded table understate what has actually been done. And -- the one that matters -- the result is LOST if the runner is killed, which is a completed trial's worth of GPU time discarded with no record that it ever ran. as_completed writes each row the moment its trial finishes. A trial that raises is reported and skipped rather than killing the remaining work.
tasks and design sat between live and cells, so two static ground-truth pages were interleaved with the views that move while a grid runs. They answer a different question -- everything left of the divider changes as the run proceeds, everything right of it does not -- and mixing them invites reading a frozen table as a live one. live cells arms cost calib │ tasks design tab and shift-tab still cycle the whole row; the divider is only in the header.
opencode opens keep-alive connections and drops them when a session ends. http.server's default response is to dump a full traceback per connection, so a live probe's own output vanished under hundreds of ConnectionResetErrors and the proxy looked like it was failing while it was working correctly. handle_error now ignores the benign disconnects (reset, broken pipe, aborted, timeout) and handle_one_request treats them as the end of the conversation instead of letting them propagate out of the handler. The second half also explains the selftest that flaked only on 3.10: the reset ended the sender thread mid-loop, the test compared against a hardcoded count, and a networking hiccup got reported as proxy cross-attribution -- a wrong diagnosis aimed at the most safety-critical component in the suite. Hammered with 20 clients that connect, send a partial request and yank the socket: zero tracebacks, zero records written for the aborted attempts, and a well-behaved request afterwards still attributed correctly.
A trial where the agent quits without editing spends a fraction of a real attempt, so it drags the median token count down. Measured on the validation grid, and the rate is strongly arm-dependent: a1 8/69 abandoned = 12% median understated by 23% a2 1/70 = 1% 1% a3 3/67 = 4% 7% No abandoned trial has ever passed. So giving up is a mechanism by which the instruction surface drives the pass rate -- a1 makes this model quit twelve times more often than a2 -- and that is a result, not noise to filter. The hazard is the cost column. a1 is the arm the treatment is measured against, and its unconditioned median reads 23% cheaper than the work it actually did. The registered analysis conditions on success so F1 is safe, but every surface that prints a raw median has to be able to say so or it will be read as a cost win. Cells now carry tok_worked (abandons excluded). The live summary gains an column in red, marks a median with when abandons are skewing it by more than 5%, and the cells detail shows the conditioned figure beside the raw one. Diagnosis for the record: the trial that prompted this was not cut off by the harness. Six proxy calls, all HTTP 200, no timeout, no truncation -- calls 0-3 finished and call 4 finished . The model chose to end after four tool rounds and 21,901 input tokens, having probed six times and edited nothing.
Root cause of the abandoned trials, and it is neither the model nor the
arm. mkpr uses the PR body as the prompt, and a PR body is not reliably a
task. Bodies that describe the SOLUTION read as changelog entries:
13403 'Use int64 for GitHub database IDs' / '## Use int64 for ...'
13766 'fix(skills): honor --dir' / '## Summary - skip the prompt when'
The model replied, verbatim: 'I see you have shared a PR description about
migrating GitHub database IDs ... What would you like me to do with this?'
One inference call, zero tool calls, six seconds. That is 8 of the 14
abandoned trials in the validation grid. Bodies that describe a PROBLEM
('gh pr checks prints a blank summary when ...') were attempted normally.
It was arm-dependent -- a1 12%, a2 1% -- so the instruction surface was
being credited for whether the agent guessed the prompt was a task at all.
A confound sitting directly on the primary outcome, and one that flatters
the quitting arm twice over: a refusal costs a single call, so a1's median
token count read 23% below the work it actually did, and 467x below on
a1/13403 where four of seven trials refused.
Every PR prompt now carries the same imperative. It names no file,
package, symbol or API -- it says only that the description is work to be
done -- and every arm receives it identically.
Verified by reproducing the failure and then fixing it, same scenario,
same arm, same model:
before 1 call, 0 tools, 12,553 tok, 6s, asks what to do
after 21 calls, 45 tools, 1,180,052 tok, editing files
Two display bugs, surfaced by leaving reproduction debris on disk beside a live run. A kept out-dir (--keep-sandbox, or one not yet swept) has a transcript and no live process. That read as 'grading' forever, because the state check only asked whether a transcript existed and never whether anything was still working on it. It is now 'done'. And the budget clock kept running against it -- 80%, then 101%, then 104% of a deadline that had already stopped applying. The adapter timeout governs a trial that is still generating; grading is unbounded and a finished trial is not racing anything. Budget is now reported only while running or stalled, and is capped at 100%: a percentage over the whole is never information, it is a clock nobody turned off.
The previous check asserted 'budget computed when a timeout is known', which the state fix correctly invalidated -- and the selftest caught the change rather than letting it through, which is the point of it. Now it asserts what is actually true: no budget for a trial that is not running, because a finished one is not racing anything, and a real fraction within [0,1] while it is. The running case is exercised by holding the out-dir busy, so both directions are covered instead of only the one that happened to pass.
opencode's session store runs SQLite in WAL mode, so live writes land in opencode.db-wal and the main .db neither grows nor restamps until a checkpoint. Measured on a delegating trial: .db 234s stale while -wal was 33s fresh. Both the live view's staleness clock and the runner's idle deadline watched the .db alone, so a trial whose root is blocked on subagents -- its stream silent by construction -- looked idle for a whole checkpoint interval. It flickered to 'stalled' in the TUI, which is how it was caught, and the same blind spot sits in wait_or_kill where the consequence is not a wrong label but a killed run. This is the delegation false-stall from two days ago, half-fixed: adding the store was right, naming the wrong file made it a coin flip. And it fails toward the arms that route to subagents, which are the arms under test. _activity_size now sums mtime as well as size, because a WAL being overwritten in place can hold steady in size while still taking writes.
A first-hand narrative of the work: purpose, the question that prompted it, the design and the rules it refuses to break, the fixture and its two supply collapses, the validation runs, the TUI, and a defect catalogue grouped by what each would have done to a published result. Written while it was happening rather than reconstructed, because the value is in the specifics -- which file the WAL check was watching, what the model actually replied when the prompt did not ask for anything, the 467x median distortion on a1/13403. The thesis, stated where it can be checked: a benchmark's hardest problem is not measuring the thing, it is noticing when you are measuring something else. Almost every defect here was silent and several were directional -- they would have produced a clean, publishable, wrong answer rather than an obvious failure.
First result of the fresh probe: cli-cli-13057 a1/1, killed at the 2700s ceiling, recorded as calls=0 tok=0 abandoned=True. The evidence line said what was really lost -- 'hit the hard ceiling while still active (27,623,328 bytes of events, last advanced 18s ago)'. Forty-five minutes of work and 27 MB of stream, discarded because the adapter writes its transcript at the end and the kill took the adapter with it. For a cost experiment that is not a missing row, it is a measurement thrown away -- and it lands hardest on the arms and scenarios expensive enough to hit a ceiling, which are exactly the ones a cost claim is about. The adapter now has a salvage mode and the runner re-invokes it after any kill to convert whatever the stream captured. Verified by forcing a kill: before calls=0 tok_in=0 abandoned=True after calls=8 tok_in=357,131 abandoned=False The trial stays ungradeable either way; exclusion criterion 1 still drops it from every pass rate. What salvage buys is the ability to say what an arm was spending when the ceiling stopped it. And no longer fires on a trial the harness stopped. Giving up and being cut off look identical in a transcript that ends early, and conflating them put a give-up flag on the harness's own decision -- on a 45-minute run, which is the opposite of what happened. metrics.abandoned now takes and withholds the flag when it is false. Also: the idle check's change-detector summed mtime with size so a WAL overwritten in place would still register, but that sum was then printed as a byte count -- '1,785,879,733 bytes' for a 27 MB stream. Fingerprint and byte count are now separate values.
The abandon warning fired at 4/120 and none of the four had given up. Three were cli-cli-13057 killed at the 2700s ceiling while still active -- mislabelled by the pre-fix code, corrected going forward. The fourth was cli-cli-13068 a1/0 with 26 calls, 35 tool calls, 894,535 tokens and EVERY CHECK PASSING, flagged as having given up. EDIT events are emitted only for tools that name a file in their input, and that agent did its editing through the shell. git saw the changes -- diff_coverage passed on them -- and the transcript recorded none, so the trial had no first edit and the whole routing family degraded with it: first_edit empty, probe_trail empty, probes_to_first_edit counting everything, and pr.route reporting 'no probes recorded before the first edit'. The transcript is the wrong authority for the question. metrics.abandoned now takes observed_edits from git and outranks the transcript with it, falling back only when nobody looked. observed_edits is recorded beside the transcript's own count so the gap between them is visible rather than inferred -- a divergence means the stream is not seeing a class of edit, which is worth knowing before it empties a metric silently. Covered both directions: a killed run and a shell-edited run are not abandonment; a trial that made one probe and stopped still is.
Three fixes to the same allocation. PRIORITY, NOT PROPORTION. Running and summary are bounded -- one row per concurrent trial, one per (arm, scenario) cell -- so they fit and should simply be shown whole. Graded grows without bound for the life of the run, so it is the section that scrolls. Sharing the space proportionally gave graded 22 rows it did not need while truncating a 3-row running table. NO SECTION CAN VANISH. Summary previously took 'whatever is left' (max_y - y - 2), which on a 24-line terminal pushed graded off the bottom entirely: absent rather than truncated, with nothing saying so -- the same failure as the empty sections that rendered no header at all. Graded now keeps a floor of one row, and a scrollbar to explain itself. SCROLLBARS ON ALL THREE. Only cells and the activity feed had them; the live sections printed a range in text or nothing at all, so 'there is nothing more' and 'you are at the top of a long list' looked identical. Also a blank row above the legend, so the last line of content does not read as part of it. Verified across heights -- running 3/3 and summary 9/9 at every size, graded taking 1, 3, 17 and 33 of 41 rows at 24, 30, 44 and 60 lines.
…the legend The verdict field was 9 wide and 'ungradeable' is 11, so it overran into calls and rendered as '0ungradeab' -- the trial number and a clipped verdict fused into one unreadable token, on exactly the rows that most need reading, since ungradeable means the harness stopped the trial. Widened to 13 and the columns after it shifted. And the scrolling section's range line sat flush against the legend, so '3-41 of 41 [j/k] scrolls' read as part of the key bindings. The body now reserves two rows rather than one: the range indicator has somewhere to go, and a blank line separates it from the footer.
Rows land in completion order now that each result is written as it finishes, so the table's order says what arrived last and nothing about when the run was doing it -- two trials of the same scenario can sit rows apart having started together. The column reads provenance.started_at, which is stored UTC because a result that cannot be placed on a shared clock is not replayable, and shown as local wall-clock because the operator is comparing it against their own terminal. Falls back to an em dash rather than guessing when a row predates the field. Also adds 'started' to the graded sorts, beside 'recent': arrival order and start order diverge under --jobs>1 and answer different questions.
The H4 tab read FAIL at 59.6% aggregate. Thirty-six of forty-one rows were 0.000%; every non-zero one was a trial the harness had stopped, whose transcript is truncated by construction and, before salvage existed, empty. Comparing that against the proxy's complete record is not a calibration failure, it is comparing two different things -- and the banner's advice, "drop the adapter figures and re-derive cost from the proxy", was being issued on a false signal. Incomplete runs are now excluded for the same reason exclusion criterion 1 drops them from every pass rate. Underneath were three disagreements that are real, all on trials that COMPLETED AND PASSED: 13068|a1|0 894,535 vs 2,376,415 62% 26 calls vs 66 13057|a1|4 2,657,861 vs 21,234,781 87% 30 calls vs 176 13057|a1|3 6,982,703 vs 28,287,934 75% 85 calls vs 226 The proxy resolves 13068|a1|0 into four distinct system prompts: 26 calls matching the adapter's root exactly, then 22 and 18 on two subagent sessions the transcript never recorded, and 2 auxiliary. Forty calls and ~1.5M input tokens invisible. E3 is a claim about precisely those tokens, and this is the first time the proxy has caught something -- which is what makes it authoritative under design section 3.2. The child walker queries opencode's store the instant `opencode run` returns, and a session written moments earlier is not always visible yet. It now retries. And metrics record task_dispatches beside the sessions actually captured, so a gap between them is on the row rather than inferred later from a proxy the analysis may not have. Also, from the graded table: started and dur now sit together with a month-day stamp, since a grid runs for hours and can cross midnight.
…thing An ungradeable row printed "—" in the failing column, so the one column that explains a verdict was blank on exactly the rows whose verdict needs explaining. The reason was sitting in the adapter check's evidence the whole time. It now reads: cli-cli-13057 a1/0 ungradeable ... adapter hit the hard ceiling of 2700s while still active (20,580,192 bytes of events, last advanced 58s ago) which is the difference between "something went wrong" and "this trial was working when we stopped it" -- and the second is what decides whether the scenario is too expensive to keep. And the footer offered [space] detail on cost and calib, neither of which has a detail pane. Advertising a key that does nothing is worse than offering none: it gets tried, and the silence reads as a broken view rather than an absent feature.
THE HOLE. cli-cli-13523 is cli-graded and touches no pkg/cmd/ command, so
cli.surface reported `skip` and the only checks left running were
diff_coverage and scope. all_pass is "every non-ungradeable check passed",
so the verdict became "touched the right file and nothing else" -- which
an agent passes by adding a comment. Two tasks are affected and 5 of 5
trials under them recorded PASS.
Such a task is gradeable by nothing: its unit tests name symbols the fix
introduces, and it exposes no CLI surface. cligrade now returns `bad`
rather than `skip` so it cannot yield a pass, and mkscenarios drops these
at extraction, where they belong. The runner has its modules loaded, so
the live battery is unaffected -- the five existing rows must be discarded
by hand.
THE SCREEN gated on license, PR count, subsystems, AGENTS.md and
CODEOWNERS. cli/cli passed all five and then produced one in-band scenario
in twelve. Four checks added, each for a failure this session actually
hit:
usable_rate source+tests PRs, not raw merges. 174 merges held 32
usable tasks; the sample estimates 16.7% against 18.4%
measured by hand.
behavioural_suite an acceptance/e2e tree means a repo can be graded on
observed behaviour -- name-independent AND
discriminating, which neither current grader manages.
Weighted highest.
fix_rate a bug fix's tests call API that exists; a feature PR's
tests name what the PR adds. 18 of 34 tasks landed in
the second bucket and none reached the band. Reported
as None below 8 classifiable titles, since most repos
do not use conventional prefixes.
median_files 13057 spans 13 files, takes 45 minutes, times out 3 of
5 and passes both times it finishes.
Also [S] asc/desc did nothing outside the live tab: cells, arms and cost
had a sort key and no direction, while the footer advertised one.
per_cell was expect/len(cells), and len(cells) is cells that have PRODUCED a row -- not cells the suite will run. At 86/120 that was 120/18 = 7, so every finished 5-trial cell showed 2 trials left and an eta for work that was already done. The estimate only converges once the last scenario starts, which is exactly when nobody needs it. Read --trials from the batch's own recorded argv instead; the runner already wrote down how it was launched. Falls back to the fullest cell observed, which can read low early but never invents remaining work. The [s] left sort had the matching defect: it ranked by dur_s * trials, time already spent, so finished-and-slow cells sorted above ones with trials still to run.
…enario in-band
pass_rate averaged over every row, and an ungradeable row carries
all_pass=False because nothing graded it -- so an adapter that hit its
hard ceiling counted as a model that got the task wrong. The
registration forbids exactly this ('Harness faults are not model
failures'), and analyze.harness_excluded has always applied it.
The two surfaces therefore disagreed. On the 120-run validation grid
cli-cli-13057 hit the ceiling on 3 of 5 trials and read 40% in the
viewer against 100% in the registered analysis -- and 40% is inside the
calibration band while 100% is outside it. A display that can move a
scenario in or out of band is not a display bug: it changes the count
that decides whether the experiment can run.
One _ungradeable() predicate now feeds both the ung column and the rate,
so a row cannot be shown as ungradeable and counted as a failure beside
it. Verified: the viewer and the registered analysis now agree on every
cell.
…oes not 6 scenarios in band against MIN_PAIRED=4; 5 of them unit-graded, so the primary tier clears on its own. The CLI tier does not discriminate -- 11 of 15 at 100%. H4 fails as specified (19.6% aggregate), but 22 of 24 scenarios agree EXACTLY -- 0.0% tokens and identical call counts -- and all six in-band scenarios are among them. The divergence is transcript truncation on ceiling-hit trials, not an accounting error, which is why the proxy was made authoritative in the first place.
Each image is placed where it carries an argument the prose was making in words: the cost frontier IS the ceiling effect (16 cells, all at 100%, across a 93x token range); the tasks tab explains why 13057 blew the ceiling (173-line prompt, 5 issues closed, 13 files); the arms rollup is one row with empty ratio columns, which is what 'no experiment ran' looks like when the instrument is honest. The grading card is the sharp edge from 5.4 on screen: verdict pass, one check ungradeable.
CI: check_proxy_attribution compared calls the CLIENT confirmed against rows the PROXY logged. On 3.10 urlopen can raise on a keep-alive connection after the proxy has already handled and recorded the call, so sent undercounted, the log did not, and a networking hiccup was reported as cross-attribution -- red CI for the opposite of the defect under test. Assert attribution directly instead: no foreign keys, and each trial's attributed count bounded by what was actually sent to it. 5/5 clean locally. The matrix now says why it holds 3.10, since nothing declared it: it is the oldest CPython still getting security fixes and the suite uses no 3.11+ syntax. Keeping it is what surfaced this race at all. probe.py: the same pass-rate defect as suitedata, in the surface that prints the verdict. It matched adapter faults on status 'fail' when a ceiling hit records 'ungradeable', so it printed '0 harness' for a run with three killed trials -- and divided by all five, so cli-cli-13057 read 40% and was marked KEEP inside the band when its gradeable trials were 2 of 2. Now 6 of 24, agreeing with the registered analysis. One is_ungradeable() in suitedata, imported by probe, so the three surfaces cannot drift apart again.
The remaining 3.10 failure was the other side of the same race. The proxy writes its log row AFTER the response is on the wire (proxy._pump_body), so a client that has read its body can be milliseconds ahead of the recorder: 6 calls confirmed, 5 rows on disk, reported as 'calls are being lost or misfiled'. That ordering is fine in production -- recording after responding keeps latency off the critical path, and the log is for post-hoc analysis. It is the test that has to stop asserting through it. Bounded wait for the expected row count, then assert attribution as before; a genuine loss still fails after the timeout.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Branch work so
full-scope(windows/macos/arm) stays out of the loop until this is stable — it triggers on any push tomain, and on PRs only with thefull-testlabel.Contains the live run dashboard, the process-lifecycle hardening (process-group kills, idle-vs-hard deadlines, run-id attribution, reclaim, flock'd output), the registered exclusion criteria 1–2, and the Windows fix for a liveness probe that was sending Ctrl-C to itself.
Add the
full-testlabel when ready for cross-platform.