Skip to content

v0.3.0: token savings (wait_task, brief view, auto-fix/escalation, usage report, opt-in savers) — stacked on #4 - #5

Merged
mrchatam merged 7 commits into
mainfrom
feat/v0.3-token-savings
Sep 26, 2026
Merged

mrchatam merged 7 commits into
mainfrom
feat/v0.3-token-savings

Conversation

@mrchatam

@mrchatam mrchatam commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #4. This PR's base is now feat/task-handoff, so the diff shows only the v0.3 changes. Merge #4 first; then retarget this PR to main (GitHub may do it automatically when #4's branch is merged) and merge it. Do not merge before #4.

The branch was rebased onto #4's review fixes (force-pushed).

Why

Grok Workhorse exists to reduce supervisor (e.g. Grok) usage: smaller, user-chosen models do bounded tasks. v0.3.0 (Unreleased) cuts what the supervisor reads and how often it has to act, and adds opt-in token savers for workers. Details: docs/token-savings.md.

What's in it

  1. wait_task long-poll: one id or many (any/all), capped at 55 s so each call stays under the MCP TS SDK's default 60 s request timeout. Also workhorse wait. The delegation skill now prefers it.
  2. Compact JSON from the MCP shim, plus task_result view: "brief" (~0.5-1 KB). full stays the default and is unchanged apart from compact JSON (the complete handoff is kept for compatibility).
  3. auto_fix_rounds plus escalate along the escalate_to chains (cheap -> mid -> strong), with hard caps (fix rounds ≤ 3, max_auto_runs ≤ 6, optional max_tokens and max_cost_usd). Fix rounds count per profile; the total is capped by max_auto_runs. Budgets include auto-review tasks, 0 = explicit zero budget, checked between runs. The trail is in result.auto.
  4. auto_review on a cheap profile: advisory verdict, new transient status reviewing; request_changes sets the handoff to needs_review. The verdict is read from the reviewer's first line (negations handled, otherwise unclear). Review tasks are hidden from list_tasks by default and recovered/cancelled after a crash.
  5. usage_report (MCP/RPC) and workhorse stats: by profile and by day, with per-run tokens. The supervisor number is labelled ESTIMATE and is conservative (only successful workers' output tokens minus the supervisor's own I/O; worker input not counted); the formula is documented.
  6. Presets and size routing in profiles.json, list_models shows them, new tiered example config, tightened skill draft.
  7. delegate_tasks batch (≤ 10), then wait_task over many ids.
  8. Test-only stub backend. It is registered only with WH_ENABLE_STUB_BACKEND=1 in the daemon env and only runs its bundled script; workhorse health warns if the flag is on. It drives the approval, continue, retry/fallback, restart (SIGKILL and SIGTERM), handoff, auto-fix, escalation and review flows. CI now runs npm run test:stub.
  9. approvals.require_operator: MCP approve only records a request (with an id). A human confirms with sudo workhorse approve, which prints the pending request and sends a separate operator token (only its SHA-256 is stored) plus the displayed request id; the daemon refuses if the request changed. A parked task keeps a persistent operator gate until the operator answers, also after update_handoff state=closed, cancel_task or reject; reject with instructions needs the operator. This gates the parked-task flow; it is not a capability boundary (a supervisor can still delegate a new task asking the same thing, visible in the audit log).
  10. Operations: audit.jsonl rotation, per-profile stall_minutes, and the installer installs the pinned Kilo/OpenCode from committed lockfiles (npm ci), falling back to npm install -g. The executable link comes from the package's bin field (opencode-ai 1.18.32 ships bin/opencode.exe); a new CI job installs each pin with npm ci --ignore-scripts and checks the link target exists.

Worker token savers (all opt-in, token_savers in daemon.json or per profile, workhorse token-savers, installer --token-savers):

  • terse and minimal_code fragments, in our own wording, inspired by Caveman and Ponytail (both MIT); attribution in NOTICE.
  • rtk: the guard plugin rewrites worker bash commands via rtk rewrite (RTK, Apache-2.0; Kilo/OpenCode only), with a minimal environment (no keys, telemetry disabled), a 1 s timeout and a fallback to the original command; only the rtk binary is bound into the sandbox.
  • Headroom was evaluated and not integrated: it needs a proxy that sees code and keys, plus heavy ML dependencies.
  • The daemon's own test run is never compressed.

Benchmarks (small samples; token counts are estimates via the o200k tokenizer unless provider-reported)

  • Supervisor, stub backend:
    • A successful task drops from about 3,060 tokens read in 8 calls (v0.2 flow, assuming 5 status polls) to about 360 tokens in 2 calls. The best case for v0.2 (no polls) is about 1,940.
    • A task that needs one fix round drops from about 6,900 tokens in about 16 calls to about 400 tokens in 2 calls, with the fix done by the cheap worker.
  • RTK on this repo: −35% over a mix of 8 commands (git status −76%, git log −61%, git diff −13%, test output not rewritten). The "60-90%" claim is not reproduced overall.
  • Prompt savers, real model (Nemotron 3 Ultra, 3 A/B pairs): output tokens −10%, same diff sizes, input dominated by turn-count noise. Indicative only.

Review fixes (independent review of #4 and #5)

All 13 findings addressed; #4's parts (restart socket race, closed-on-parked semantics, approval source attribution) are on feat/task-handoff. README has a CI badge and a "Measured savings" section (all figures labelled estimates, with method).

Tests

Local, on b18f9e0:

  • Full suite (npm test) run twice: 114 tests, 112 pass, 0 fail, 2 skipped (the live test and the stub-prerequisites placeholder), both runs.
  • npm run test:unit: 63/63. (The old handoff-dedupe test was removed because the full view keeps the complete handoff again.)
  • npm run test:stub run 10 times in a row: 10/10 green, each 20 pass + 1 placeholder skip. The restart tests run on their own: 4/4 green.
  • CI on this PR: unit + stub + remaining files, and the new installer-pins job, all green.

- wait_task long-poll (single/many ids, any/all, capped at 55 s) and delegate_tasks batches
- compact MCP JSON, task_result view brief|full, handoff context deduplicated in the full view
- presets and size routing; auto_fix_rounds and escalate_to chains with hard caps and a trail
- optional advisory auto_review on a cheap profile (transient status reviewing)
- usage_report / workhorse stats with per-run tokens and a labelled supervisor ESTIMATE
- opt-in worker token_savers: terse and minimal_code fragments, RTK rewrite in the guard plugin
- approvals.require_operator with a separate operator token confirmed via the CLI
- per-profile stall_minutes, audit.jsonl rotation
- TEST-ONLY stub backend, registered only with WH_ENABLE_STUB_BACKEND=1
Runs approval, continue, retry/fallback, restart, handoff, auto-fix, escalation, auto-review,
wait_task, delegate_tasks, presets, require_operator, token savers and usage_report against the
real daemon with a scripted worker (git + python3 only).
… token-saver flags

npm ci from scripts/pins/<cli> (integrity-checked) when the pinned version is requested, else fall
back to npm install -g with a warning. New --token-savers and --rtk-bin options.
…fers wait_task

docs/token-savings.md (features, RTK/Headroom/Caveman/Ponytail evaluation, benchmarks, usage
estimate formula), configuration/handoff/architecture/security/troubleshooting updates, README,
delegation skill draft, NOTICE attributions, CHANGELOG 0.3.0 (Unreleased), version 0.3.0.
…gets, auto-review children, rtk env

- approvals.require_operator: a task that parks keeps a persistent operator_gate until the operator
  answers with the token; continue_task is refused on gated tasks whatever their status, closing via
  update_handoff/cancel_task/reject keeps the gate, reject with instructions needs the operator token
- approval requests get an id; the operator CLI prints the pending request and sends back the
  displayed id; the daemon refuses the decision if the request changed
- reviewVerdict reads the first line only, handles negations, else unclear
- max_tokens/max_cost_usd 0 is an explicit zero budget (invalid values fail validate); auto-review
  child tokens/cost count against the parent budget; continue_task restarts run counters only
- auto-review children: recover orphans after a crash, hidden from list_tasks unless
  include_auto_reviews, never needs_attention
- full view keeps the complete handoff (backcompat); brief stays small
- rtk rewrite: minimal env, 1 s timeout, falls back to the original command; bind only the rtk file
- wait_task stops when the client disconnects; usage_report estimate counts worker output only
- docs: operator gate is a parked-task flow gate, not a capability boundary; caps, trail, formula
… package.json bin

opencode-ai@1.18.32 maps opencode to ./bin/opencode.exe; scripts/pin-bin.mjs reads the bin field and
install.sh refuses a dangling link. CI: npm ci --ignore-scripts on each pin and assert the target exists.
…GELOG, recovery and budget tests

- README: CI status badge; 'Measured savings' table, every figure labelled an estimate with how it was
  measured, linking docs/token-savings.md
- CHANGELOG 0.3.0: operator gate and request ids, review verdict parsing, budgets (0 = explicit zero,
  reviews counted, per-run overshoot, continue semantics), rtk hardening, installer bin resolution,
  auto-review list/recovery, full view unchanged, conservative usage estimate
- stub e2e: restart recovery re-links a review child and cancels an orphan review
- unit: auto.max_tokens / max_cost_usd 0 kept as explicit zero budgets, invalid values reported
@mrchatam
mrchatam force-pushed the feat/v0.3-token-savings branch from 6890984 to b18f9e0 Compare September 26, 2026 22:03
@mrchatam
mrchatam changed the base branch from main to feat/task-handoff September 26, 2026 22:03
@mrchatam
mrchatam changed the base branch from feat/task-handoff to main September 26, 2026 22:34
@mrchatam
mrchatam merged commit de8adff into main Sep 26, 2026
2 checks passed
@mrchatam
mrchatam deleted the feat/v0.3-token-savings branch September 26, 2026 22:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant