Skip to content

Harden live TUI, process lifecycle, and exclusion criteria - #1

Merged
TGPSKI merged 41 commits into
mainfrom
harden/live-tui
Aug 5, 2026
Merged

TGPSKI merged 41 commits into
mainfrom
harden/live-tui

Conversation

@TGPSKI

@TGPSKI TGPSKI commented Aug 4, 2026

Copy link
Copy Markdown
Owner

Branch work so full-scope (windows/macos/arm) stays out of the loop until this is stable — it triggers on any push to main, and on PRs only with the full-test label.

Contains the live run dashboard, the process-lifecycle hardening (process-group kills, idle-vs-hard deadlines, run-id attribution, reclaim, flock'd output), the registered exclusion criteria 1–2, and the Windows fix for a liveness probe that was sending Ctrl-C to itself.

Add the full-test label when ready for cross-platform.

TGPSKI added 30 commits August 4, 2026 10:59
os.kill(pid, 0) is the standard existence probe on POSIX, where signal 0
means 'check, do not send'. On Windows it is not a probe: CTRL_C_EVENT is
0, so the call delivers a real Ctrl-C to the target's console group.

check_live wrote a marker carrying its own pid; snapshot() checked it; the
check raised KeyboardInterrupt inside an unrelated subprocess several
tests later. The traceback pointed at check_test_runner, which had nothing
to do with it, sixteen seconds into the job.

Green through 6921707, first red at af026e8 -- the live dashboard commit,
and only ever on windows/amd64. Lint, both Validate jobs, linux/arm64 and
macos/arm64 were all green throughout.

_alive now declines to answer off POSIX and the caller falls back to
evidence that cannot misfire. Guarded in the selftest, because the failure
mode of getting this wrong is a viewer that terminates what it observes.
The message told you to 'run under bench/with-proxy.sh'. That script does
not exist and never did -- the proxy starts from ADH_PROXY_LOG inside
bench/isolate.sh.

It also omitted the reason the tab is empty for most runs. The proxy holds
one trial mark at a time and nothing in an inference request identifies
its trial, so under --jobs>1 the runner skips the mark rather than writing
a wrong attribution. The H4 gate cannot be measured in parallel at all,
which makes 'pass --proxy' useless advice for exactly the runs people are
watching when they find this tab.

Now prints the two commands that work, states that JOBS=1 is a
precondition rather than a preference, and names it as the second cost of
parallelism (§16.4).
The registration listed 'per-trial proxy attribution becomes impossible'
as the second cost of parallelism. It was not a cost, it was a missing
key.

The premise was right and measured: nothing in an inference request body
identifies its trial -- the sandbox path is not in it, not even in the
system prompt. So the runner marked trial boundaries by POSTing a label to
the proxy, which is one piece of shared state; two concurrent trials write
into whichever label was set last. Rather than record an attribution it
knew was wrong, the runner skipped marking under --jobs>1 -- which meant
the H4 gate, the thing that decides whether any cost number in this suite
can be trusted, was unavailable for every parallel run.

The request PATH is ours. Each trial now gets a config whose provider
baseURL is <proxy>/__run/<run_id>/v1; the proxy strips the prefix before
forwarding and stamps the record. Attribution arrives with the call.
Upstream sees exactly what it saw before, because a proxy that alters the
request is not measuring it.

Selftest drives two trials at one proxy concurrently with overlapping
requests and asserts 4 and 6 calls land in their own buckets, tokens
survive, and upstream saw only /v1/chat/completions. 22/22.

marks are kept for serial runs and for reading existing logs; run_id wins
wherever present.
Three things the TUI surfaced by being looked at.

SUBAGENTS READ 0. The live view counted distinct sessionIDs in the root
stream, which is always 1: a subagent runs in its OWN opencode session and
the root stream carries none of its calls -- measured at 3.1x under-report
on s13. So a trial that had dispatched two 'explore' agents showed sub=0
while its own detail pane listed the spawns. The runner already gives each
trial its own XDG_DATA_HOME, so opencode's store is in the out-dir and
readable live: child sessions now come from there, read-only and
best-effort. Totals are explicitly labelled ROOT SESSION ONLY when
children exist, because E3 is the claim that this handoff is free and
quoting a root-only total would confirm it by construction.

PROXY WAS OPT-IN. It records the tokens and round trips this suite exists
to measure and §3.2 makes it authoritative over the adapter -- so the
ordinary path was producing numbers nothing had verified, and the H4 gate
could not be checked for the run you were actually doing. It now starts by
default, logs to a path paired with the results (runs/probe.jsonl ->
runs/probe.proxy.jsonl) so a log always belongs to exactly one run, and
the viewers find it without being told. ADH_NO_PROXY=1 opts out.

PROGRESS BAR READ AS AN ELLIPSIS. '..................' at 0% looks like
truncated output. Bracketed now, with a track character that cannot be
mistaken for one.

Live view is now three sections: running (cursor, [space] detail), graded
most-recent-first with verdict and failing checks, and a per-(arm,
scenario) rollup carrying pass rate, median and p90 tokens, tool averages,
subagents, median and p90 duration, trials left and ETA. The cells detail
pane gained a median/mean/p90 table with a p90:median ratio, flagged when
the tail is 2x the middle -- a treatment that halves the median while
doubling the tail has not made anything cheaper.
Two places I said 'not available' where the data was sitting on disk.

SUBAGENT TOTALS. I labelled the live totals ROOT SESSION ONLY and left it
there. The child sessions' tokens are in the same sqlite store the session
ids came from -- measured on a live trial: 15 child calls, 767,215 input
tokens the root-only figure was silently omitting. The live view now sums
parent + children, which is the only total §7 permits quoting as a cost,
and the detail pane shows the split plus the subagents' share, because E3
is a claim about the child half specifically.

TOOL OUTPUT. The NDJSON stream carries a call's input and status but not
what it returned, so the view could say 'it ran go test' and never what go
test said -- the one thing worth knowing while a run is in flight. The
store keeps the output. The detail pane now ends in a live activity feed:
newest last, subagent lines tagged, failures in red, with the actual tool
output. Caught an agent's own broken grep ('Unmatched ( or \(') on the
first read.

Both read-only and best-effort against a live writer and a private schema,
like everything else in this module.
Inline tool output turned the pane into a wall nobody could scan -- and
the interesting line is rarely the last one. The feed is now a list:

- newest first, each event one line, no output rendered inline
- indexed by absolute position in the run, so the number means how many
  events the trial has produced and where this one sits among them,
  rather than how far down the screen you are
- [j/k] moves, [space] expands the selected event over the whole pane
  with its full output, [space] or [q] collapses
- a proportional scrollbar, so 'nothing more' is distinguishable from
  'top of a long list' -- which a truncated list cannot say
- subagent events tagged, failures marked and coloured in both views

The live view now nests three levels: run list, run detail with the
activity list, one expanded event. Each level consumes its own keys so [q]
backs out one step instead of quitting from three levels deep.
FALSE STALL, and it was in the measurement path. The idle detector watched
stdout.txt, which carries ROOT-session events only. A trial that dispatches
a subagent and waits is silent there for as long as the child works.
Observed live: parent 0 calls, subagent 5 calls and 206,689 tokens, forty
tool calls deep, stream untouched for 3m27s -- reported 'stalled', and with
--idle-timeout 300 the runner would have killed it as hung.

It would also have done that MORE to the arms that route to subagents,
which are the arms under test. A harness that kills the treatment for
delegating would have produced a clean, wrong result. Progress is now
either the event stream or opencode's session store moving, in both the
runner's deadline and the live view's verdict.

NAVIGATION. Only the running table had a cursor; graded and the per-arm
rollup were read-only. [h/l] now moves between the three sections and
[space] opens the right kind of detail for each: a live run, a graded
result with every check's evidence, or the cell's median/mean/p90.

Q WAS OVERLOADED. It meant 'quit' at the top and 'back' when nested, so
from three levels deep there was no way to leave except pressing it
repeatedly and no way to know which press would exit. Now [q]/[esc] back
out one level, [Q] quits from any depth, and the footer says which one q
currently means.

Also: expanded events keep their structure. Collapsing whitespace turned
every JSON blob and file listing into one unreadable line -- structure is
most of what a tool result is. JSON is re-indented, the author's line
breaks are kept, wrapped continuations are indented past the original, and
the expanded view takes the full pane instead of the eight lines left
under the run header.
Spacebar into the newest event and read it, and new events arriving at the
top pushed the selection backwards one row each -- so the thing being read
was replaced mid-sentence by whatever had since become event N. Same class
as the run-list bug fixed earlier, in the second live table.

Events now carry the store's own primary key and the cursor anchors to it.
act_follow keeps the useful default: before you move, the cursor tracks the
newest event, which is what 'what is it doing' means for a run in flight.
The first [j]/[k] pins it, and expanding pins unconditionally -- reading
one event while the list follows the newest is exactly how the thing you
opened disappears. A selection that rolls out of the window re-anchors
rather than dangling.

Covered by a selftest over the anchoring logic directly, since scraping
curses redraws out of a pty proved unreliable: two new events arrive, the
index slides 2 -> 4, the selected event stays put. 23/23.
POST-MORTEM. The failures are neither model failures nor grader bugs. The
PR's own tests are applied to the agent's independent implementation, and
for a FEATURE PR those tests reference symbols the PR introduces --
'undefined: findCopilotBinaryFunc', 'unknown field IssueType in struct
literal of type CreateOptions'. The test does not compile, so passing
requires the agent to independently choose the maintainers' exact
identifier. That measures name-guessing, not the work.

Measured on the validation grid, and the separation is perfect: all four
compile-class scenarios scored 0%, 5%, 0%, 0% -- none in the [0.25, 0.80]
band -- while every in-band and every ceiling scenario was
behaviour-class. It is the dominant cause of the floor problem the
registration names as its biggest schedule risk, and on cli-cli-13057 the
agent had already PASSED diff_coverage with 13/13 pre-edit probes landing
in the right directories. It did the work and failed on a name.

This is why SWE-bench-style grading works for bug fixes and not for
feature PRs: a bug fix's tests call API that already exists.

mkpr's fail-before check proved a task goes red at the parent but never
asked why. It now classifies: tests that fail to compile because they name
something the fix introduces are rejected with the identifiers listed.
Free, mechanical, and it runs before any GPU time is spent.

Static scan of the existing 34: 19 clean, 9 whose missing surface is <=8
symbols, 6 whose surface is 12-130 symbols and cannot be stated without
handing over the design.

Also: the TUI legend is now depth-accurate -- each level advertises only
the keys it responds to, because a legend listing keys that do nothing is
checked once and then trusted. [Q] quits from any depth, [q] backs out one
level, [esc] returns to the root list regardless of nesting.
The question was whether the grader can adapt to the agent's naming
without handing over the solution. It can, and the answer is to stop
grading at the Go boundary.

Declaring the API surface in the prompt was the obvious repair and it is
worse than the disease: a symbol list IS routing information
(CreateOptions.IssueType names its subsystem), so it hands every arm a
piece of exactly what the treatment is supposed to supply. It biases the
primary outcome toward the null in the treatment's own channel, which
makes a null result uninterpretable -- 'the pattern did not help' and 'we
gave the control the answer' become the same measurement.

Grade what the PR publicly promises instead.  is a
user-facing contract: it is in the PR body the agent already receives, and
it is identical whatever the agent names its internals. The oracle is the
merge commit's OWN binary, so the expectations stay the maintainers' and
not the experimenter's:

    reference = build(merge commit)
    candidate = build(the agent's tree)
    compare flag-for-flag at the command line

Nothing is added to the prompt. The grader simply stops asking a question
it never gave the agent the means to answer.

Verified against a real agent tree from the live probe -- one that scored
pr.task_pass=fail purely for naming a struct field differently:

    [pass] cli.builds
    [pass] cli.surface: advertises all 17 flag(s) the PR's own binary does
    [pass] cli.extra_surface

Offline and deterministic; --help and argument validation need no network,
which is what makes it runnable in the sandbox. mkpr now MARKS such tasks
rather than rejecting them, records which grader judged each one and which
identifiers forced the choice, and the command path is derived from the
PR's own file paths so the battery is mechanical rather than authored.
Recovers the 15 compile-class tasks: 34 usable, not 19.

Also: left/right arrows navigate the nested views -- right descends, left
backs out one level. At the root they stay bound to section switching,
since binding right to both would make sections unreachable.
CI was red on three commits, always Validate (3.10), always 'BAD proxy
attribution under concurrency', never on 3.14. The message said the trials
had cross-attributed, which would have been a real defect in the
measurement path -- the proxy is authoritative for H4.

It had not. Stressed at 8 concurrent trials x 20 calls: 160 requests, 160
records, zero drops, zero mismatches. The failure was 'got 5 of 6': one
request raised inside a sender thread, which silently ended that thread's
loop, and the test compared against a hardcoded count. A transient
client-side error on 3.10's http.client was being reported as a
cross-attribution bug.

The test now counts what actually got through and asserts the invariant it
means: every call the proxy handled is attributed to the trial that made
it. A networking hiccup reduces the sample instead of inventing a defect.

Separately: the graded and per-arm tables could be focused with [h/l] and
their cursors moved, but nothing rendered -- no highlight, no indication of
which section had focus, so both looked inert. Each section header now
carries a focus marker and each table highlights its own selected row only
when it holds focus.
GRADER. All 34 tasks classified by compiler, not heuristic:
adherence.classify checks each PR's tests out onto its parent tree and
builds them. A build error naming an undefined identifier means the test
cannot compile against any implementation that chose different names.
Result: 16 unit-graded, 18 cli-graded. The declaration-scan heuristic I
tried first was rejected for false negatives -- it missed 13675
(field on an existing struct) and 13624 -- and the compiler catches both.

Calibrated against the validation grid: it flags exactly the scenarios
that scored 0% (13393, 13675, 13967) and leaves 13624 (5%) and 13403
(71%) on the unit grader, which are demonstrably achievable.

prgrade now dispatches on the task record. For a cli-graded task
diff_coverage becomes advisory rather than failing: cli.surface already
proves the shipped behaviour flag-for-flag against the PR's own binary,
and demanding the same file set would smuggle back the match-the-
maintainers'-structure requirement this amendment exists to remove.
Verified end to end -- three CLI passes give all_pass=True, which is what
the table, the TUI and the report show. Locked in by a selftest.

AMENDMENT 2 in docs/EVAL.md: what the grid measured, the repair that was
rejected and why (a symbol list is routing information and would bias the
primary outcome toward the null inside the treatment's own channel), what
is registered instead, and the reporting rule -- unit-graded tasks are
the primary evidence, cli-graded a separate tier, never pooled.

TUI, from watching it run:
- graded and summary were silently truncated (6 rows of 25, no scroll, no
  indication). Both scroll now with their own cursor and print the range.
- allocation is terminal-size aware and proportional to what each section
  holds: 24 lines shows 5 graded rows, 34 shows 14, 50 shows all 25.
- the results cursor was unbounded, so holding [j] walked it off the end
  of the list, the highlight vanished, and [space] opened a detail pane
  for a row never on screen.
- [space] in a graded or summary detail opened a phantom activity level
  with nothing in it that took two backs to leave.
- sections reordered running -> summary -> graded, [s] sorts the focused
  one, [S] flips direction, both shown in the header.
- j/k crosses section boundaries; h/l still jumps whole sections.

Also widened the process-hygiene timing: it failed twice under CPU
contention from a concurrent run, reporting a defect that was really load.
The block swap that reordered the sections dropped the trailing spacer, so
the two tables ran together with no gap. Also adds the range indicator the
summary section was missing -- graded had one, summary did not, so a
truncated summary looked complete.
[tab] moved forward through live/cells/arms/cost/calib and there was no
way back except cycling all the way around. KEY_BTAB and the literal 353
some terminals send when terminfo carries no kcbt entry are both accepted,
so the key is not dead on a terminal that reports it raw -- verified
against the real \e[Z sequence through a pty. Switching views also resets
the row cursor, which otherwise carried a position from a list of a
different length.
Two new read-only views, and neither touches run data.

TASKS. An operator watching a grid sees scenario ids and pass rates with
no way to tell what cli-cli-13418 wants, where the answer lives, or why
one task is judged by unit tests and the next at the command line. That
gap is exactly where a harness artifact gets read as a model failure --
the thing Amendment 2 exists to stop. Lists every scenario with its PR,
its grader, the files and directories the real diff touched, and what it
asks for; [space] opens the full prompt with the grader explained in the
terms a verdict needs, plus the identifiers that forced a cli grading.

DESIGN. Same for the arms, measured off disk rather than described, so it
cannot drift from what a trial is actually handed: always-loaded bytes,
delta against a1, on-demand bytes, and a hash of the surface. a0 0 B,
a1 6,518, a2 18,060, a3 3,474 + 13,850 on demand, a5 417. Carries the
run's ground rules -- primary outcome, why a cost number without a pass
rate is not reportable, the calibration band, harness faults never scoring
as model failures, purpose stamping, the two grading tiers, the proxy
being authoritative, and disclosed-not-forbidden deviation.

Both detail panes are height-aware and scroll: they used to render until
they ran out of screen and stop, so on a short terminal a 179-line prompt
was cut with nothing saying so and no way to reach the rest. Verified at
26 lines.

Also  and  for the same content outside the TUI.
Restarting cleanly is a five-step ritual with three ways to get it wrong,
and all three have already happened here: pkill orphans the opencode
grandchildren, the paired proxy log gets forgotten so the next run
calibrates against the last one's calls, and the temp sandboxes stay
behind so the live view keeps rendering a torn-down run.

So the awkward part is one command: SIGTERM the whole process group,
escalate to SIGKILL for anything that ignores it, report what is still
unowned, sweep the sandboxes.

It does NOT delete results, and that is the design rather than an
omission. Bundling 'kill the processes' with 'remove the data' into one
verb is how a habit formed for the safe purpose ends up performing the
destructive one, and a --force flag exists to be used. Clearing a results
file stays a manual rm -- the command prints the exact one, including the
paired proxy log, and says so when the file holds rows marked
purpose=experiment.

Plan-only unless YES=1, so a mistyped invocation costs nothing.
[s] cycled the cells sort from every view, so the tasks tab showed
'sort:tok' above a list that [s] did not touch. Each tab now sorts by
what it actually holds -- tasks by id, grader, files, dirs or PR; arms by
id, always-loaded bytes or on-demand bytes -- with [S] flipping direction,
and the header shows the sort that [s] would change HERE.

Also from watching it:
- the design tab ignored [j]/[k] entirely. _move_rows counted rows for
  every view except that one, so the cursor had nothing to move through.
- with nothing graded yet, the live view's summary and graded sections
  rendered no header at all, so [h/l] moved focus onto an invisible table
  and the operator had no way to tell whether the section was empty or
  broken. Both now draw their header and say what they are waiting for.
- ground-rule labels were clipped at a fixed 26 columns, mid-word:
  'Harness faults are not mode'. Width now comes from the longest label.
ThreadPoolExecutor.map yields in SUBMISSION order, so a single slow trial
buffers every later result inside the executor. Caught live on the probe:
trial 2 of cli-cli-13057 completed, its out-dir was cleaned up, and its row
sat unwritten behind two trials still 6+ minutes from finishing.

That is three problems at once. The result is invisible to the live view,
because the out-dir it reads is already gone. The progress count and the
graded table understate what has actually been done. And -- the one that
matters -- the result is LOST if the runner is killed, which is a
completed trial's worth of GPU time discarded with no record that it ever
ran.

as_completed writes each row the moment its trial finishes. A trial that
raises is reported and skipped rather than killing the remaining work.
tasks and design sat between live and cells, so two static ground-truth
pages were interleaved with the views that move while a grid runs. They
answer a different question -- everything left of the divider changes as
the run proceeds, everything right of it does not -- and mixing them
invites reading a frozen table as a live one.

  live  cells  arms  cost  calib  │  tasks  design

tab and shift-tab still cycle the whole row; the divider is only in the
header.
opencode opens keep-alive connections and drops them when a session ends.
http.server's default response is to dump a full traceback per connection,
so a live probe's own output vanished under hundreds of
ConnectionResetErrors and the proxy looked like it was failing while it
was working correctly.

handle_error now ignores the benign disconnects (reset, broken pipe,
aborted, timeout) and handle_one_request treats them as the end of the
conversation instead of letting them propagate out of the handler.

The second half also explains the selftest that flaked only on 3.10: the
reset ended the sender thread mid-loop, the test compared against a
hardcoded count, and a networking hiccup got reported as proxy
cross-attribution -- a wrong diagnosis aimed at the most safety-critical
component in the suite.

Hammered with 20 clients that connect, send a partial request and yank the
socket: zero tracebacks, zero records written for the aborted attempts,
and a well-behaved request afterwards still attributed correctly.
A trial where the agent quits without editing spends a fraction of a real
attempt, so it drags the median token count down. Measured on the
validation grid, and the rate is strongly arm-dependent:

  a1  8/69 abandoned = 12%   median understated by 23%
  a2  1/70            =  1%                          1%
  a3  3/67            =  4%                          7%

No abandoned trial has ever passed. So giving up is a mechanism by which
the instruction surface drives the pass rate -- a1 makes this model quit
twelve times more often than a2 -- and that is a result, not noise to
filter.

The hazard is the cost column. a1 is the arm the treatment is measured
against, and its unconditioned median reads 23% cheaper than the work it
actually did. The registered analysis conditions on success so F1 is
safe, but every surface that prints a raw median has to be able to say so
or it will be read as a cost win.

Cells now carry tok_worked (abandons excluded). The live summary gains an
 column in red, marks a median with  when abandons are skewing it
by more than 5%, and the cells detail shows the conditioned figure beside
the raw one.

Diagnosis for the record: the trial that prompted this was not cut off by
the harness. Six proxy calls, all HTTP 200, no timeout, no truncation --
calls 0-3 finished  and call 4 finished . The model
chose to end after four tool rounds and 21,901 input tokens, having
probed six times and edited nothing.
Root cause of the abandoned trials, and it is neither the model nor the
arm. mkpr uses the PR body as the prompt, and a PR body is not reliably a
task. Bodies that describe the SOLUTION read as changelog entries:

  13403  'Use int64 for GitHub database IDs' / '## Use int64 for ...'
  13766  'fix(skills): honor --dir' / '## Summary - skip the prompt when'

The model replied, verbatim: 'I see you have shared a PR description about
migrating GitHub database IDs ... What would you like me to do with this?'
One inference call, zero tool calls, six seconds. That is 8 of the 14
abandoned trials in the validation grid. Bodies that describe a PROBLEM
('gh pr checks prints a blank summary when ...') were attempted normally.

It was arm-dependent -- a1 12%, a2 1% -- so the instruction surface was
being credited for whether the agent guessed the prompt was a task at all.
A confound sitting directly on the primary outcome, and one that flatters
the quitting arm twice over: a refusal costs a single call, so a1's median
token count read 23% below the work it actually did, and 467x below on
a1/13403 where four of seven trials refused.

Every PR prompt now carries the same imperative. It names no file,
package, symbol or API -- it says only that the description is work to be
done -- and every arm receives it identically.

Verified by reproducing the failure and then fixing it, same scenario,
same arm, same model:
  before   1 call,   0 tools,    12,553 tok,  6s, asks what to do
  after   21 calls, 45 tools, 1,180,052 tok, editing files
Two display bugs, surfaced by leaving reproduction debris on disk beside a
live run.

A kept out-dir (--keep-sandbox, or one not yet swept) has a transcript and
no live process. That read as 'grading' forever, because the state check
only asked whether a transcript existed and never whether anything was
still working on it. It is now 'done'.

And the budget clock kept running against it -- 80%, then 101%, then 104%
of a deadline that had already stopped applying. The adapter timeout
governs a trial that is still generating; grading is unbounded and a
finished trial is not racing anything. Budget is now reported only while
running or stalled, and is capped at 100%: a percentage over the whole is
never information, it is a clock nobody turned off.
The previous check asserted 'budget computed when a timeout is known',
which the state fix correctly invalidated -- and the selftest caught the
change rather than letting it through, which is the point of it.

Now it asserts what is actually true: no budget for a trial that is not
running, because a finished one is not racing anything, and a real
fraction within [0,1] while it is. The running case is exercised by
holding the out-dir busy, so both directions are covered instead of only
the one that happened to pass.
opencode's session store runs SQLite in WAL mode, so live writes land in
opencode.db-wal and the main .db neither grows nor restamps until a
checkpoint. Measured on a delegating trial: .db 234s stale while -wal was
33s fresh.

Both the live view's staleness clock and the runner's idle deadline
watched the .db alone, so a trial whose root is blocked on subagents --
its stream silent by construction -- looked idle for a whole checkpoint
interval. It flickered to 'stalled' in the TUI, which is how it was
caught, and the same blind spot sits in wait_or_kill where the
consequence is not a wrong label but a killed run.

This is the delegation false-stall from two days ago, half-fixed: adding
the store was right, naming the wrong file made it a coin flip. And it
fails toward the arms that route to subagents, which are the arms under
test.

_activity_size now sums mtime as well as size, because a WAL being
overwritten in place can hold steady in size while still taking writes.
A first-hand narrative of the work: purpose, the question that prompted
it, the design and the rules it refuses to break, the fixture and its two
supply collapses, the validation runs, the TUI, and a defect catalogue
grouped by what each would have done to a published result.

Written while it was happening rather than reconstructed, because the
value is in the specifics -- which file the WAL check was watching, what
the model actually replied when the prompt did not ask for anything, the
467x median distortion on a1/13403.

The thesis, stated where it can be checked: a benchmark's hardest problem
is not measuring the thing, it is noticing when you are measuring
something else. Almost every defect here was silent and several were
directional -- they would have produced a clean, publishable, wrong answer
rather than an obvious failure.
First result of the fresh probe: cli-cli-13057 a1/1, killed at the 2700s
ceiling, recorded as calls=0 tok=0 abandoned=True. The evidence line said
what was really lost -- 'hit the hard ceiling while still active
(27,623,328 bytes of events, last advanced 18s ago)'. Forty-five minutes
of work and 27 MB of stream, discarded because the adapter writes its
transcript at the end and the kill took the adapter with it.

For a cost experiment that is not a missing row, it is a measurement
thrown away -- and it lands hardest on the arms and scenarios expensive
enough to hit a ceiling, which are exactly the ones a cost claim is about.

The adapter now has a salvage mode and the runner re-invokes it after any
kill to convert whatever the stream captured. Verified by forcing a kill:

  before   calls=0  tok_in=0        abandoned=True
  after    calls=8  tok_in=357,131  abandoned=False

The trial stays ungradeable either way; exclusion criterion 1 still drops
it from every pass rate. What salvage buys is the ability to say what an
arm was spending when the ceiling stopped it.

And  no longer fires on a trial the harness stopped. Giving up
and being cut off look identical in a transcript that ends early, and
conflating them put a give-up flag on the harness's own decision -- on a
45-minute run, which is the opposite of what happened. metrics.abandoned
now takes  and withholds the flag when it is false.

Also: the idle check's change-detector summed mtime with size so a WAL
overwritten in place would still register, but that sum was then printed
as a byte count -- '1,785,879,733 bytes' for a 27 MB stream. Fingerprint
and byte count are now separate values.
The abandon warning fired at 4/120 and none of the four had given up.

Three were cli-cli-13057 killed at the 2700s ceiling while still active --
mislabelled by the pre-fix code, corrected going forward.

The fourth was cli-cli-13068 a1/0 with 26 calls, 35 tool calls, 894,535
tokens and EVERY CHECK PASSING, flagged as having given up. EDIT events
are emitted only for tools that name a file in their input, and that agent
did its editing through the shell. git saw the changes -- diff_coverage
passed on them -- and the transcript recorded none, so the trial had no
first edit and the whole routing family degraded with it: first_edit
empty, probe_trail empty, probes_to_first_edit counting everything, and
pr.route reporting 'no probes recorded before the first edit'.

The transcript is the wrong authority for the question. metrics.abandoned
now takes observed_edits from git and outranks the transcript with it,
falling back only when nobody looked. observed_edits is recorded beside
the transcript's own count so the gap between them is visible rather than
inferred -- a divergence means the stream is not seeing a class of edit,
which is worth knowing before it empties a metric silently.

Covered both directions: a killed run and a shell-edited run are not
abandonment; a trial that made one probe and stopped still is.
Three fixes to the same allocation.

PRIORITY, NOT PROPORTION. Running and summary are bounded -- one row per
concurrent trial, one per (arm, scenario) cell -- so they fit and should
simply be shown whole. Graded grows without bound for the life of the run,
so it is the section that scrolls. Sharing the space proportionally gave
graded 22 rows it did not need while truncating a 3-row running table.

NO SECTION CAN VANISH. Summary previously took 'whatever is left'
(max_y - y - 2), which on a 24-line terminal pushed graded off the bottom
entirely: absent rather than truncated, with nothing saying so -- the same
failure as the empty sections that rendered no header at all. Graded now
keeps a floor of one row, and a scrollbar to explain itself.

SCROLLBARS ON ALL THREE. Only cells and the activity feed had them; the
live sections printed a range in text or nothing at all, so 'there is
nothing more' and 'you are at the top of a long list' looked identical.

Also a blank row above the legend, so the last line of content does not
read as part of it.

Verified across heights -- running 3/3 and summary 9/9 at every size,
graded taking 1, 3, 17 and 33 of 41 rows at 24, 30, 44 and 60 lines.
…the legend

The verdict field was 9 wide and 'ungradeable' is 11, so it overran into
calls and rendered as '0ungradeab' -- the trial number and a clipped
verdict fused into one unreadable token, on exactly the rows that most
need reading, since ungradeable means the harness stopped the trial.
Widened to 13 and the columns after it shifted.

And the scrolling section's range line sat flush against the legend, so
'3-41 of 41  [j/k] scrolls' read as part of the key bindings. The body now
reserves two rows rather than one: the range indicator has somewhere to
go, and a blank line separates it from the footer.
TGPSKI added 11 commits August 4, 2026 14:23
Rows land in completion order now that each result is written as it
finishes, so the table's order says what arrived last and nothing about
when the run was doing it -- two trials of the same scenario can sit rows
apart having started together.

The column reads provenance.started_at, which is stored UTC because a
result that cannot be placed on a shared clock is not replayable, and
shown as local wall-clock because the operator is comparing it against
their own terminal. Falls back to an em dash rather than guessing when a
row predates the field.

Also adds 'started' to the graded sorts, beside 'recent': arrival order
and start order diverge under --jobs>1 and answer different questions.
The H4 tab read FAIL at 59.6% aggregate. Thirty-six of forty-one rows were
0.000%; every non-zero one was a trial the harness had stopped, whose
transcript is truncated by construction and, before salvage existed, empty.
Comparing that against the proxy's complete record is not a calibration
failure, it is comparing two different things -- and the banner's advice,
"drop the adapter figures and re-derive cost from the proxy", was being
issued on a false signal. Incomplete runs are now excluded for the same
reason exclusion criterion 1 drops them from every pass rate.

Underneath were three disagreements that are real, all on trials that
COMPLETED AND PASSED:

  13068|a1|0   894,535 vs  2,376,415   62%   26 calls vs  66
  13057|a1|4 2,657,861 vs 21,234,781   87%   30 calls vs 176
  13057|a1|3 6,982,703 vs 28,287,934   75%   85 calls vs 226

The proxy resolves 13068|a1|0 into four distinct system prompts: 26 calls
matching the adapter's root exactly, then 22 and 18 on two subagent
sessions the transcript never recorded, and 2 auxiliary. Forty calls and
~1.5M input tokens invisible. E3 is a claim about precisely those tokens,
and this is the first time the proxy has caught something -- which is what
makes it authoritative under design section 3.2.

The child walker queries opencode's store the instant `opencode run`
returns, and a session written moments earlier is not always visible yet.
It now retries. And metrics record task_dispatches beside the sessions
actually captured, so a gap between them is on the row rather than
inferred later from a proxy the analysis may not have.

Also, from the graded table: started and dur now sit together with a
month-day stamp, since a grid runs for hours and can cross midnight.
…thing

An ungradeable row printed "—" in the failing column, so the one column
that explains a verdict was blank on exactly the rows whose verdict needs
explaining. The reason was sitting in the adapter check's evidence the
whole time. It now reads:

  cli-cli-13057 a1/0  ungradeable ... adapter hit the hard ceiling of
  2700s while still active (20,580,192 bytes of events, last advanced 58s
  ago)

which is the difference between "something went wrong" and "this trial was
working when we stopped it" -- and the second is what decides whether the
scenario is too expensive to keep.

And the footer offered [space] detail on cost and calib, neither of which
has a detail pane. Advertising a key that does nothing is worse than
offering none: it gets tried, and the silence reads as a broken view
rather than an absent feature.
THE HOLE. cli-cli-13523 is cli-graded and touches no pkg/cmd/ command, so
cli.surface reported `skip` and the only checks left running were
diff_coverage and scope. all_pass is "every non-ungradeable check passed",
so the verdict became "touched the right file and nothing else" -- which
an agent passes by adding a comment. Two tasks are affected and 5 of 5
trials under them recorded PASS.

Such a task is gradeable by nothing: its unit tests name symbols the fix
introduces, and it exposes no CLI surface. cligrade now returns `bad`
rather than `skip` so it cannot yield a pass, and mkscenarios drops these
at extraction, where they belong. The runner has its modules loaded, so
the live battery is unaffected -- the five existing rows must be discarded
by hand.

THE SCREEN gated on license, PR count, subsystems, AGENTS.md and
CODEOWNERS. cli/cli passed all five and then produced one in-band scenario
in twelve. Four checks added, each for a failure this session actually
hit:

  usable_rate       source+tests PRs, not raw merges. 174 merges held 32
                    usable tasks; the sample estimates 16.7% against 18.4%
                    measured by hand.
  behavioural_suite an acceptance/e2e tree means a repo can be graded on
                    observed behaviour -- name-independent AND
                    discriminating, which neither current grader manages.
                    Weighted highest.
  fix_rate          a bug fix's tests call API that exists; a feature PR's
                    tests name what the PR adds. 18 of 34 tasks landed in
                    the second bucket and none reached the band. Reported
                    as None below 8 classifiable titles, since most repos
                    do not use conventional prefixes.
  median_files      13057 spans 13 files, takes 45 minutes, times out 3 of
                    5 and passes both times it finishes.

Also [S] asc/desc did nothing outside the live tab: cells, arms and cost
had a sort key and no direction, while the footer advertised one.
per_cell was expect/len(cells), and len(cells) is cells that have PRODUCED
a row -- not cells the suite will run. At 86/120 that was 120/18 = 7, so
every finished 5-trial cell showed 2 trials left and an eta for work that
was already done. The estimate only converges once the last scenario
starts, which is exactly when nobody needs it.

Read --trials from the batch's own recorded argv instead; the runner
already wrote down how it was launched. Falls back to the fullest cell
observed, which can read low early but never invents remaining work.

The [s] left sort had the matching defect: it ranked by dur_s * trials,
time already spent, so finished-and-slow cells sorted above ones with
trials still to run.
…enario in-band

pass_rate averaged over every row, and an ungradeable row carries
all_pass=False because nothing graded it -- so an adapter that hit its
hard ceiling counted as a model that got the task wrong. The
registration forbids exactly this ('Harness faults are not model
failures'), and analyze.harness_excluded has always applied it.

The two surfaces therefore disagreed. On the 120-run validation grid
cli-cli-13057 hit the ceiling on 3 of 5 trials and read 40% in the
viewer against 100% in the registered analysis -- and 40% is inside the
calibration band while 100% is outside it. A display that can move a
scenario in or out of band is not a display bug: it changes the count
that decides whether the experiment can run.

One _ungradeable() predicate now feeds both the ung column and the rate,
so a row cannot be shown as ungradeable and counted as a failure beside
it. Verified: the viewer and the registered analysis now agree on every
cell.
…oes not

6 scenarios in band against MIN_PAIRED=4; 5 of them unit-graded, so the
primary tier clears on its own. The CLI tier does not discriminate --
11 of 15 at 100%.

H4 fails as specified (19.6% aggregate), but 22 of 24 scenarios agree
EXACTLY -- 0.0% tokens and identical call counts -- and all six in-band
scenarios are among them. The divergence is transcript truncation on
ceiling-hit trials, not an accounting error, which is why the proxy was
made authoritative in the first place.
Each image is placed where it carries an argument the prose was making
in words: the cost frontier IS the ceiling effect (16 cells, all at
100%, across a 93x token range); the tasks tab explains why 13057 blew
the ceiling (173-line prompt, 5 issues closed, 13 files); the arms
rollup is one row with empty ratio columns, which is what 'no
experiment ran' looks like when the instrument is honest.

The grading card is the sharp edge from 5.4 on screen: verdict pass,
one check ungradeable.
CI: check_proxy_attribution compared calls the CLIENT confirmed against
rows the PROXY logged. On 3.10 urlopen can raise on a keep-alive
connection after the proxy has already handled and recorded the call, so
sent undercounted, the log did not, and a networking hiccup was reported
as cross-attribution -- red CI for the opposite of the defect under
test. Assert attribution directly instead: no foreign keys, and each
trial's attributed count bounded by what was actually sent to it. 5/5
clean locally.

The matrix now says why it holds 3.10, since nothing declared it: it is
the oldest CPython still getting security fixes and the suite uses no
3.11+ syntax. Keeping it is what surfaced this race at all.

probe.py: the same pass-rate defect as suitedata, in the surface that
prints the verdict. It matched adapter faults on status 'fail' when a
ceiling hit records 'ungradeable', so it printed '0 harness' for a run
with three killed trials -- and divided by all five, so cli-cli-13057
read 40% and was marked KEEP inside the band when its gradeable trials
were 2 of 2. Now 6 of 24, agreeing with the registered analysis.

One is_ungradeable() in suitedata, imported by probe, so the three
surfaces cannot drift apart again.
The remaining 3.10 failure was the other side of the same race. The
proxy writes its log row AFTER the response is on the wire
(proxy._pump_body), so a client that has read its body can be
milliseconds ahead of the recorder: 6 calls confirmed, 5 rows on disk,
reported as 'calls are being lost or misfiled'.

That ordering is fine in production -- recording after responding keeps
latency off the critical path, and the log is for post-hoc analysis. It
is the test that has to stop asserting through it. Bounded wait for the
expected row count, then assert attribution as before; a genuine loss
still fails after the timeout.
@TGPSKI
TGPSKI merged commit 37b5eaa into main Aug 5, 2026
4 checks passed
@TGPSKI
TGPSKI deleted the harden/live-tui branch August 5, 2026 01:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant