Skip to content

Vendor the TUI layer from pane, its new canonical upstream - #6

Merged
TGPSKI merged 2 commits into
mainfrom
vendor-pane
Aug 5, 2026
Merged

TGPSKI merged 2 commits into
mainfrom
vendor-pane

Conversation

@TGPSKI

@TGPSKI TGPSKI commented Aug 5, 2026

Copy link
Copy Markdown
Owner

Summary

  • src/adherence/tui/ is now vendored from pane (pane@a51e682) instead of leather — pane is the canonical upstream the whole portfolio re-vendors from
  • picks up bar_chart's aggregation binning, median/p90-guarded outlier clipping, and top-row label collision handling
  • adds interact.py (prompt_search, filter_rows, hbar, cycle); README/ruff.toml provenance notes and CHANGELOG updated

Test plan

  • make check green with the new vendored copy (selftest 25/25)
  • pane upstream: 37 tests green incl. grid-level chart regressions
  • pane check-vendor.sh reports byte-identity

TGPSKI added 2 commits August 4, 2026 21:47
…rong

Written to be read cold. The status is that the harness is validated and
the fixture is the open question: 6 of 24 scenarios discriminate against
a threshold of 4, and the registered experiment has not been run.

Next steps are ordered and blocking-first -- drop the two unusable
scenarios, materialize the remaining arms so the floors cross-check can
run at all, then decide the grading tier.

The session record leads with reversals rather than fixes, because the
useful part is what was believed and then disproved: a headline that
read 7 before the viewer stopped scoring harness faults as model
failures, the same defect found in three surfaces including the one that
prints the go/no-go verdict, and four consecutive red pushes that were
not four bugs but the only outcome the CI trigger allowed.

It also records two process failures -- bytecode committed to a public
repo, and ruleset changes applied to four repositories when one was in
scope -- with the rule that follows: unclear means stop and ask.
src/adherence/tui/ now comes from github.com/TGPSKI/pane (tools/vendor.sh,
pane@a51e682) instead of leather: charts.py gains aggregation binning,
outlier clipping and top-row label collision handling; interact.py joins
the vendor set; __init__.py carries the provenance stamp. README, ruff
exclusion comment and CHANGELOG updated.
@TGPSKI TGPSKI added the full-test Run the cross-platform full-scope CI matrix on this PR label Aug 5, 2026
@TGPSKI
TGPSKI merged commit 2c0f9d9 into main Aug 5, 2026
10 checks passed
TGPSKI added a commit that referenced this pull request Aug 9, 2026
* a cell with nothing gradeable has no pass rate, not a pass rate of zero

suitedata.load_cells divided by the count of gradeable rows without
checking it was non-zero, so a cell whose completed trials were all
harness faults raised ZeroDivisionError out of the one loader that
table, matrix and live all sit on. Reproduced from the validation
grid's own three ungradeable rows:

  FILES=<3 rows from runs/probe.jsonl> make table
  suitedata.py:188 in load_cells -> ZeroDivisionError

During a live battery that is the first scenario whose completed
trials are ceiling hits, which is exactly when the viewer is wanted.
probe.py already guarded the same case; the shared loader did not.

0.0 alone would have been the other half of the same defect -- 0%
reads as floor, and a scenario at floor gets dropped from the fixture.
So the cell carries `gradeable`, one `pass_str` formatter renders the
absence as an em dash, and `pass_attr`/`pcol` take the cell rather
than the number so it is dimmed instead of painted red.

The same exclusion applies one level up: arm_rollup averages over
scenarios that have a rate, and the Pareto frontier and the cost/
quality scatter drop cells whose quality axis is absent.

Refs #8

* report.py was the fourth surface scoring harness faults as model failures

docs/HANDOFF.md records this defect found in three surfaces and fixed by
one shared suitedata.is_ungradeable(). report.py -- the publishable
markdown scoreboard -- was not one of them, and still averaged all_pass
over every row.

Measured on runs/probe.jsonl, cli-cli-13057 (3 of 5 trials killed at the
adapter's 2,700 s ceiling):

  probe.py / analyze.py / suitedata   1.00
  make report                         0.40

0.40 is inside the registered band [0.25, 0.80] and 1.00 is outside it,
so the surface meant for publication disagreed with the go/no-go verdict
about whether that scenario discriminates. Overall pass@1 for the grid
moves 0.86 -> 0.88.

is_ungradeable is imported, not restated. An `ungradeable` column now
sits beside every rate it was computed without, because an exclusion
nobody can count is indistinguishable from data that never existed, and
a cell with no gradeable trial prints an em dash rather than 0.00.

Also here, same file:

- The document now states its own inputs. analyze.py refuses rows that
  are not purpose=experiment; this is the working scoreboard so refusing
  is wrong, but 120 validation rows rendered a complete scoreboard with
  no trace of what it had read. Header line, blockquote, stderr warning.
- st.mean(r['duration_s'] ...) raised KeyError on rows predating the
  field, alone among the reads in that row.

Closes #9, closes #12

* the H4 gate read twice the number its own viewer tab showed

suitedata.calibration() excludes adapter-fault trials, and says why: a
trial the harness stopped has a transcript truncated by construction, so
against the proxy's complete record it reads as a ~100% accounting error,
which is not what H4 measures and buries the disagreements that are real.

calibrate.py -- the registered gate, the one that decides whether P2
opens -- iterated every result row. Same two files, runs/probe.jsonl and
runs/probe.proxy.jsonl:

                        runs   aggregate   worst
  make calibrate         120     19.571%   100.000%
  suitedata.calibration  117      9.484%    87.483%

docs/HANDOFF.md reports "19.6% aggregate" as the H4 result. That is the
un-excluded figure. H4 fails either way -- call counts are 5319 adapter
against 5646 proxy after exclusion -- but two surfaces disagreeing by 2x
about a registered quantity is the defect, not the number.

They now agree exactly. The excluded count is printed rather than
dropped, per exclusion criterion 1.

Also: r['arm'] read directly where analyze.load and suitedata.load_rows
both default it, so a row written before the arms work raised KeyError.

Closes #10, closes #13

* a definition shared only by convention is not shared

check_harness_exclusion pointed at analyze.harness_excluded alone. The
invariant it was standing in for is repo-wide -- "there is now one
is_ungradeable() in suitedata, imported by the others" -- and two
surfaces drifted from it without CI noticing, both fixed in the two
commits before this one.

The new check builds one cell of five trials: two killed by the harness,
two passes, one real failure. Excluded, the rate is 67%; scored, it is
40%. Every surface that reports a rate is then run against it --
analyze, suitedata, report, probe, calibrate -- and must land on the same
number. The all-faults case is asserted from the other side: no rate at
all, an em dash, "harness broke", and no ZeroDivisionError.

Each of the five defects was re-introduced and confirmed to fail it:

  report.pass_rate ignores ungradeable      -> 0.40, want 0.667
  calibrate compares every row              -> 5 of 5, aggregate non-zero
  probe scores harness faults               -> 40% over 5 trials, KEEP
  pass_str renders an absent rate as 0%     -> '0%'
  suitedata divides by the survivors        -> ZeroDivisionError

selftest: 26/26.

Closes #11

* ci-local wrote a shell script to a predictable path, then ran it

/tmp/adh-ci-steps.sh is guessable and world-writable-adjacent: on a
shared box it can be pre-created as a symlink, or replaced between the
write and the bash that executes it with the invoking user's privileges.

mktemp -d, and the python shim moves into the same directory so one trap
cleans up both. The script already used mktemp for the shim, so the
pattern was there; this path was missed.

make ci-local: all steps passed, including the A2 == A3 byte assertion.

Closes #14

* docs: correct the upstream, the SHA, the ceiling count and the H4 figure

Found reviewing PR #5, whose content had already landed via #6.

- AGENTS.md said the TUI is vendored from leather. Upstream moved to pane
  in 2c0f9d9; README.md and CHANGELOG.md already said so, the routing
  table agents actually read did not. Same correction in HANDOFF §6.
- HANDOFF §1 pinned main at b4ae1ca, two re-vendor commits behind.
- HANDOFF §4's probe block put "16 ceiling" under the corrected column
  with no arrow. cli-cli-13057 moved into ceiling, so 16 is the
  before-value; the run prints 17.
- HANDOFF §2 quoted 19.6% for H4. That was calibrate.py reading rows that
  exclusion criterion 1 drops. It is 9.5% over 117 gradeable trials.
- HANDOFF §4 said the ungradeable defect turned up in three surfaces. It
  was five; the other two are fixed earlier in this branch.

Closes #15

* CHANGELOG: the ungradeable surface sweep
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-test Run the cross-platform full-scope CI matrix on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant