Chat seat performance, RSI coach, provider hub, and sentinel - #63
Merged
Conversation
Tools-only ACP end_turn no longer clears "Jevons is working…"; chrome stays through agent_note re-prompts until visible text seals (or silent residual / vacuous settle / cancel). Hermetic lifecycle + chrome policy tests in chat_events_test.js.
History replay sat at scrollTop≈0 so tall near-end bubbles were treated as below the fold and left collapsed after pin. Suppress off-screen collapse mid-replay; after pin / stick-to-bottom expand all tall messages still in the viewport (not only latest). Pure helpers + hermetic oracles.
Owner-corrected model: frontier = ready set (not next-ticket queue); worker-per-leaf + policy-on-set; map T155/T193/T198/T222 vs T254.1–6; ship-J/ship-B/covered/defer/oos table; T262.4 draft needs-owner; no Beads dual-write; T254 stays parked until owner accept.
POST /api/asides with text registers, rehydrates, and sends the owner opening prompt in the same turn; deliver failure is loud HTTP. UI passes opening body on aside: and shows RHS working chrome until first reply. Hermetic: go test ./internal/server -run TestHandleCreateAside*; node web/scripts/agent_transcript_test.js (T263). Co-Authored-By: Grok <noreply@x.ai>
Pin flash-class never-paint for aside wires: prefix detection (incl. image-prefixed bodies), isMainAsideWireUserText local fallback, and addMsg last-line defense so a main bubble cannot flash then self-clear. Hermetics cover incident fixture + history coalesce skip.
Single-graph contain scale-to-fill helpers + hermetic 10+ node oracle. Multi-component pack residual (bin-pack later). Co-Authored-By: Grok <noreply@x.ai>
Pure TargetContextChrome helpers + top-edge message tab so owner-facing 🎯 asks paint ledger context (repo · PO). Hermetic fixture covers jevons-vs-bullseye disambiguation. Coordinates with T267 via shared extract/resolve; no PR (T104). Co-Authored-By: Grok <noreply@x.ai>
Promote RHS agent/aside Transcript toward main-chat microcosm: wire merge keeps in-flight working, sidebar send opens optimistic bubble+working, and inspect chrome uses agent-named working label. Hermetic T265 pure helpers plus index wiring greps. Conversation-only pane (no nested fleet/frontier). Local master only (T104). Residual: VirtualList optional for short pane; T269 aside dismiss chrome may co-land in index CSS/handlers.
Aside purpose rows paint a right-edge × visible on hover/focus; activate calls dismissFleetAside (DELETE /api/asides). Work/PO rows never show it. Hermetic: fleet_row_test T269 suite + index wire contract. Co-Authored-By: Grok <noreply@x.ai>
When an owner ask surfaces for a specific 🎯 target (__TARGET_ASK__:Tn or needs-owner prose), auto-select the owning product PO and emphasize the matching Frontier row (ft-highlight, scroll into view). Pure plan helpers are hermetic; seal path only (no mid-stream thrash). Coordinates with T266 context chrome; static web hard-reload residual (T188). Local master T104.
Archive dismissed asides to stateDir/asides-history.json on DELETE; expose GET /api/asides/history. Create path records kind (side|capture|target) in open meta so filing vs side-chat stay distinguishable after the live 💡 row is gone. RHS "Closed" affordance loads the durable list (not session-only). Oracles: go test ./internal/server -run Aside|Closed|NormalizeAside; node web/scripts/aside_history_test.js; freeform create opts carry kind.
Re-sample in-flight level on history_meta/status instead of edge-only send memory. Server emits history_meta.working from waiting/stream/ PromptInFlight; client applyWorkingLevelSample re-arms or stays idle. Hermetic: workingLevelFromSample open-turn vs sealed; Go history_meta working true/false; index wiring.
…T271) Tall AABB(card∪hosts) kept frontier cards open when the pointer moved up/down over other table rows still within the card height. Product path now uses card ∪ hosts ∪ horizontal corridor only (HIDE_GRACE_MS=0). Oracle: node web/scripts/instant_tip_test.js (all green, incl T271 cases).
…T273) Main bubble context tab (T266) and shared speaker helpers: never paint overseer/jevons/jevons-po; non-overseer identity is bold purple 〈name〉 with no middle-dot and no Jevons prefix. Hermetics cover both omit and paint fixtures. Co-Authored-By: Grok <noreply@x.ai>
Split oversized connected components (cap 24 nodes/diagram) so Mermaid can render large external ledgers like orthograph (was one 64-node diagram → hung empty panel). Pause earlier-history hydrate while the large graph overlay is open so "Loading earlier…" does not flash. Add render timeout + hermetic oracles (Go + Node). Co-Authored-By: Grok <noreply@x.ai>
…er only (T273) c200de1 over-omitted by treating jevons-po like bare Jevons speaker and gating the whole tab on speaker-omit. Split paths: speaker-omit is only overseer/product Jevons; context-paint always shows product·owning-PO (including jevons-po); non-overseer other products still paint 〈agent〉. Hermetics cover all three fixtures. Co-Authored-By: Grok <noreply@x.ai>
Coalesce same-worker idle/agent notes on the overseer queue, drain owner turns first and alone, interrupt fleet-only chews so owner is not deferred, and bind working level to owner turns only. Hermetic flood test bounds queue depth and proves drain-to-zero + owner priority. Co-Authored-By: Grok <noreply@x.ai>
Empty ModelFeed during warm-up is unknown (no owner notice); after DefaultProviderFeedWarmUp still empty → stall. Hermetics cover both paths.
… (T219) Extend staffops pure policy with deliberate-stop ignore, max actions/hour budget, and observe surfaces (overseer/fleet/eventlog/frontier). Add a continuous daemon sentinel loop that samples product APIs/logs, classifies harness-ok|repair|file+PO|ignore, performs control-plane repairs, and delivers residual missions to jevons-po — no product implement, no Ship. Compose T325.4 staff ops; mechanical floor remains T204/T207/T85. Co-Authored-By: Grok <noreply@x.ai>
…epth (T346) Stall:frontier file+PO only when ≥1 unattended-ready (LeafReady) leaf after design/deferred/parked/needs-owner filters; all-gated hubs are harness-ok. Co-Authored-By: Grok <noreply@x.ai>
…d family words (T348) The FAMILIES list was load-bearing: a Claude model id whose family word was not in it (fable before T312, any future family) condensed to '' and the fleet row painted the bare splat despite /api/agents reporting a concrete model — the regression class the owner shot on jv-t347-reload-pin. Now a 'claude-…' id with an unlisted family word sniffs the first alphabetic token after 'claude' as the family (initial from its first letter), and a family-less id (claude-2) paints its bare version. Grok/GPT keep their deliberately letter-free badges; non-Claude unknown ids are untouched. Oracle: node --test web/scripts/model_prefix_test.js green (new T348 block: unlisted family → Z6, bedrock spelling → N4.5, family-less → 2, production-id sweep asserts condenseModel/label never empty when a version-bearing model string is present, paint path emits <sub> after the icon). make test-go + test-web green. collapse-test failure is pre-existing on a clean tree (T261 fixture, concurrent T347 work). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015SDy7MhKDjbbZ8GH4o36hz
…el (T348) Live probe after the T348 web fix caught two Claude rows reporting model='<synthetic>' on /api/agents — Claude Code's stamp on frames it wrote itself (API errors/cancellations, clustered around daemon restarts). The session-log parser has filtered it since T311, but the live wire did not: one synthetic frame pinned '<synthetic>' into the sticky hub and the badge painted a bare mark with a '<synthetic>' tooltip until the next real turn. modelFromEvent now treats synthetic as "frame names none" (stickiness keeps the last real model), and the feed drops + clears any pre-fix hub residue so the pin/log chain stands in. Oracle: go test ./internal/server -count=1 green (new TestSyntheticFrameNeverEvictsTheLearnedModel, TestSyntheticHubResidueNeverReachesTheFeed, synthetic cases in TestModelFromEvent); go vet clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015SDy7MhKDjbbZ8GH4o36hz
…parade (T347) Owner regression: hard-reload of a long chat painted every replayed turn old-to-new (full markdown parse per frame) and, once per-frame work opened >150ms gaps, the idle fallback ended pin suppression mid-burst — every remaining frame then stick-to-bottom pinned: an upward scroll parade with a ~60s settle. - Replay-burst appends are lazy shells (estimate-height geometry, zero parse); the single history_meta pin materializes only the end band, viewport-first, rAF-capped — end-first fill-upward, far-above stays shells. - New _awaitingHistoryMeta guard: pre-meta frames re-enter suppression if the idle fallback fired early (VirtualList.shouldReenterReplay, 30s cap so a meta-less server degrades to live pinning). - virtualizeMessages + flushRematerializeFrame gate on the burst; stream render/seal keep replay shells unpainted (estimate only); post-band materialize runs the T66/T261 tall auto-expand pass. - Oracles: replayHydrateTrace proves zero mid-hydrate scrollTop climb, zero replay paints, band-bounded materialize, dist-to-end 0 — and catches both regressions (per-message pin parade; eager full-list materialize). collapse-test T261 fixture retuned for lazy replay shells. Oracle evidence: node web/scripts/virtual_list_test.js PASS; make test-web PASS; make test-ui PASS (collapse, replay-scroll, virtual-list, t341 included). Hardens T119/T119.1 against T336/T341-era paths reintroducing the parade. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015bvCHd3fNxCmaJKYLqTN4e
…²) reflow (T349) Owner P0: typing froze for tens of seconds (blind-typed blobs landing as one chunk) with repeated Firefox slow-script dialogs on a long transcript. Root cause: virtualizeMessages interleaved layout reads (offsetTop/ offsetHeight of the next row) with dematerialize writes (class/style/ innerHTML of the previous row). Every write invalidated layout, so each subsequent read forced a full reflow of the flow container — O(rows²) sync layout when many rows dematerialize at once. Secondary: rematerialize was capped at 6 items/frame but unbounded in time — a few giant markdown bodies still blocked a frame for seconds. Fleet paint and the 3s cost poll profiled non-causal (ms-scale; already coalesced + hidden-gated since T289). - virtualizeMessages: read phase collects all geometry before any write; planVirtualizePass caps demat at DEMATERIALIZE_PER_FRAME (40) and remat at the T336 cap, viewport-first; leftovers re-arm the next frame. Landed with the planner core in e07b1b5 (co-land window with T347); this commit adds the remaining wiring. - dematerializeMsg(el, knownHeight): reuses the pre-read height — no layout read inside the write phase. - flushRematerializeFrame + virtualize remat batch: FRAME_BUDGET_MS (10ms) time budget; unpainted items spill to the next frame. - refreshAgents: descriptor build + tree paint defer to a frame boundary (scheduleFleetPaint, latest fetch wins) so an agents_changed burst cannot stretch an input task. - Hermetic stress oracle virtualizePassTrace proves both directions: phased passes bound writes/reflows per frame and converge in ~n/cap frames; the legacy interleaved pass shows ~n forced reflows (the freeze). Wire tests pin the index.html plumbing; chat-ui virtual-list test drives capped passes to convergence. Oracles: make test green (Go + node hermetic incl. new T349 tests + Playwright UI). Journey exception: composer input-latency under live Grok is not deterministically measurable; the mechanism is covered hermetically and the DOM path in real-render Playwright. Residual (class-3): owner confirms typing feels live in daily use; T349 converging, not achieved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U4a4dJ64zreSqS4y9LPP5a
…gated expansion pins (T350) Owner (post-T341/T347/T349): main chat text still randomly jiggles ~1px. Root cause (measured, not guessed): the virtualize read phase, the demat fallback, and the remat settle all measured rows with integer offsetHeight, while natural row heights are fractional (.msg line-height 1.6 x 14px = 22.4px lines; Chromium 1/64px layout units). A shell frozen at the rounded height differs from its material height by the fraction (~0.2-0.5px), so every band demat/remat cycle shifted all content below by that amount -- the random sub-pixel/1px jiggle. Probe: shell 44 / 89 px vs material 44.390625 / 89.171875 px on seeded rows. Second bypass: refreshLatestExpansion (runs on every message event) and expandInViewNearEnd wrote scrollTop = scrollHeight unconditionally while tracking, skipping the T341 shouldPinScroll gate. - virtualizeMessages read phase, flushRematerializeFrame pending reads, dematerializeMsg fallback, rematerializeMsg settle: measure with getBoundingClientRect().height (fractional, exactly representable as CSS minHeight) instead of offsetHeight. Shell and material heights are now identical: zero scrollHeight movement across demat/remat cycles. - pinToEndGated(): sync pin for expand/collapse geometry that consults shouldPinScroll (scrollTop-distance branch); refreshLatestExpansion and expandInViewNearEnd route through it. Real expands (>= one 22.4px line) still pin; sub-threshold noise no longer rewrites scrollTop. - t341-jiggle-thrash-test de-greenwashed: 120-row seed so real shells exist; rotating-shell demat/remat churn, phase-alternated across sampler ticks; refreshLatestExpansion in the stress loop; bottom-of-viewport probe (owner-visible line position); steady-state maxDelta assertions are exactly 0; churn>=10 guard so the stress can never silently no-op. Pre-fix the extended oracle FAILS (halfMaxBotDelta=0.171875); post-fix green with churn=20, all deltas 0. - virtual_list_test: T350 pure cases for the shouldPinScroll equal-height branch + wire pins (fractional reads present, no ungated scrollTop writes in refreshLatestExpansion/expandInViewNearEnd, pinToEndGated wiring). No T347/T349 regression: their planners/budgets untouched; full make test green (Go + node hermetic + Playwright UI incl. replay/virtualize suites). Residual (class-3): owner confirms jiggle gone in daily use. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019eCjingzL6QecQGgVWRtRe
Thin client does not need to hold the screen awake; voice/chat can reattach on unlock. Owner: no need to prevent sleep.
…ins, rect-exact hydrate (T351) Owner (post-T350 b60ab7c): main chat still randomly shifts ~1px in daily use. Forensics on the DAILY page (scrollTop-setter spy + per-frame bottom probe), not hermetics: the bottom-visible line moved every ~420ms, paired with progressiveLoadRemainingHistory's loadEarlier writes, for minutes after every hard reload; live turn appends moved it too. Root cause (measured): scroll offsets are integer-quantized by the engine — a scrollTop write cannot express a fraction (assigning 100.5 reads back 101; over-assign clamps to an integer) — while chat content is fractional: 22.4px text lines, 13.1875px turn-markers, 19.2px (1.2rem) bubble margins, 790.703px container. Every write path re-derived scrollTop from integer geometry (finalPinScrollTop = sh − ch; loadEarlier compensated with an integer scrollHeight delta), so whenever frac(total content height) drifted — every hydrate page, every live append — the pinned viewport landed at a different sub-pixel offset from the true content bottom. Sub-pixel phase change rasterizes as the owner's random ~1px jiggle. Why T341/T350 hermetics greenwashed (maxDelta=0 while product moved): (1) no fraction-drift stimulus — the oracle never ran loadEarlier/hydrate or live appends, and its seeded transcript's total-height fraction never changed, so integer pins sat at a constant sub-pixel offset; (2) the hand-rolled seed bypassed product paint (raw textContent, up to ~9px off renderBody truth) and the virtualize min-height freeze held every row at the stale seeded height forever. Fix (content layer + write layer): - Whole-pixel row snap: every #messages row locks min-height at the ceiling of its natural height (VirtualList.snappedRowLockPx; ≤1px invisible slack). MutationObserver seam snaps appended rows (bubbles, turn-markers, asides); renderBody/applyClipState re-snap repaints and clip toggles (3-phase clear→measure→lock flush, T349 discipline; shrink-safe); ResizeObserver re-snaps async growers (images, mermaid); remat settle locks at the snapped ceiling and caches the snapped box so shell == material stays exact (T350) on the pixel grid; virtualize drains pending snaps before its read phase. - Whole-pixel row margins: margins escape the box snap — .msg bubbles 1.2rem (19.2px) → 19px; msg-has-context-tab 0.65rem (10.4px) → 10px. - Clamp-exact pins: all pin writes over-assign scrollHeight (VirtualList.pinWriteScrollTop) instead of integer sh − ch; fraction-tolerant pinned check (isPinnedAtEnd, eps 1px); forced pins bypass the skip fast-path so the post-snap corrective pin always lands. - Rect-exact hydrate compensation: loadEarlier measures the anchor row's rect before/after the page insert (hydrateCompensatedScrollTop) instead of trusting the integer scrollHeight delta. Oracle (fails pre-fix): t351-fractional-pin-test pins the owner-visible truth — the fractional gap between the container's viewport bottom edge and the last message's rendered bottom edge must be EXACTLY constant across live fraction-drift appends (pre-fix: 11 distinct values, 0.64px max delta) and real loadEarlier hydrate pages (pre-fix: 9 distinct, 0.58px). Post-fix both legs hold one constant value, maxDelta 0. t341 re-run green with scrollDownPinned=0 (was ≤12 budget); full make test green (Go + node hermetic + Playwright UI). Product-path evidence (daily UI, real journal, fleet active): 90s scrollTop-spy observation pre-fix showed a bottom-probe move every ~420ms; post-fix 10,677 frame samples, ZERO bottom-probe moves, while the same hydrate writes (471) and fleet/cost activity ran. Residual (class-3): owner confirms jiggle gone after hard reload. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AfBzhoyHPZ5Nqh4SnyWTkF
The 2026-08-09 RSI ops live drill appended two synthetic error rows (source=rsi-drill, component=rsi_drill, msg rsi_ops_live_drill…) to exercise the coach loop (T243/T333). The sentinel clustered them as a daemon_error and fired file+PO on event:error:rsi_drill — a false positive against deliberate stimulus. ClusterEventAnomalies now drops rows matching the drill markers before any clustering, via a pure IsSyntheticDrillRow: source or component exactly rsi-drill/rsi_drill (case- and hyphen-insensitive), msg containing rsi_ops_live_drill, decision=live_drill from a drill origin, or an explicit fields.drill marker. Exact token match only, so a real error from rsi / rsi_coach still classifies as before. EventRow gains Source and Drill; both sentinel tail paths now project through one eventRowFromEvent helper that carries the markers (and parses RFC3339 as well as RFC3339Nano, which the StateDir path missed). Hermetic: drill rows alone yield zero anomalies and never reach file+PO; real daemon errors beside them cluster byte-identically to the baseline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TXSdQkTsaPUuxAJYaHbVy1
Fine sensors, coarse conclusions. The drip cursor (T243) seeds at EOF, so everything that happened before the coach booted was invisible to it. Add a bounded backward pass over history on its own slow cadence: - internal/rsi/retro.go: git mine (repair churn per commit+scope, reverts), eventlog tail reader bounded by window, chat/session turn filters, the retro value bar (ClassifyRetroValue), durable retro state. - internal/rsi/coach_retro.go: RunRetroOnce + a separate retro schedule (default 6h) that never advances the drip cursor and never files bullseye. - Judgments carry Mode=retrospective and commit SHAs / session ids as evidence pointers; delivery stays sparse — retro rate cap (2), quality bar, and T333 disposition suppressions all apply. - Dials documented and retunable: retro_enabled, retro_interval_sec, retro_lookback_hours (default 168), retro_rate_cap, retro_min_count, retro_max_commits/event_rows/sessions, retro_workdir. No full-history re-scan on the 15m drip cadence. - MCP: jevons_rsi_coach_cycle mode=drip|retro|both; configure gains the retro dials; status reports the last pass. Hermetics: history-derived judgment delivered with git evidence, low-value cluster suppressed (retro_weak_phrase_low_count / retro_one_off_git_noise), window floor honoured, drip cursor untouched, repeat pass re-delivers nothing. go test ./internal/rsi ./internal/mcpserver ./scripts/... -count=1 PASS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0191mDHYiQ2L2ygEgMwGMnPx
The first live retro pass showed a judgment about cmd/jevonsd citing commits that touched internal/mcpserver and web/scripts: evidence pointers matched on shared kind, and every git_rework row has the same kind. Carry the cluster key on the candidate and match evidence exactly; the loose kind fallback stays for hand-built candidates. go test ./internal/rsi ./internal/mcpserver -count=1 PASS (new: TestJudgmentEvidenceCitesOwnCluster — two scopes sharing a commit) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0191mDHYiQ2L2ygEgMwGMnPx
Release prep for chat seat performance, RSI coach, provider hub, and sentinel.
GET /api/rsi/dispositions reads the durable T333 disposition store and returns the judgment list with disposition state, so the owner sees what the coach judged and what the overseer did with it — not only MCP chat dumps inside the overseer session. - Entries carry an owner-readable evidence summary (SummarizeEvidence, pure) and T353 retro provenance (Mode), both recorded at delivery. - Blank disposition normalizes to "pending" on the wire. - Empty store is an honest empty list; an unknown ?disposition= filter is a 400 rather than a silently empty list that reads as "nothing found". Hermetic: internal/server/rsi_dispositions_test.go covers empty, and the pending + ignore_with_reason + file shapes plus the filter. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AwjoJNpsxEnBGup8hiamyC
CI has no /Users/marcelo/work/... path; use published claudia pseudo-version.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ship the local master stack since v0.11.0: chat seat performance, RSI coach/sentinel, provider hub, frontier table layout, fleet parks.
Highlights
Test plan
go test -tags "sqlite_preupdate_hook sqlite_fts5" ./...greenNotes