From bbb6158b2fd37af99cb5acec30f8c6c7a234e041 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 06:01:55 -0500 Subject: [PATCH 001/121] =?UTF-8?q?docs:=20spec=20for=20skill-optimizer=20?= =?UTF-8?q?v1.4=20=E2=80=94=207-skill=20decomposition?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Approved design (brainstormed 2026-05-12) converting the v1.3 monolithic auto-improve-orchestrator into 7 independent Claude Code skills that chain via the superpowers plugin pattern. Motivation: team review of v1.3 PR drafts surfaced 4 critiques — v1.3 optimizes for incremental numerical uplift without validating test case quality, grader correctness, or improvement principledness. The firecrawl iteration regression (1.0 → 0.44 from piling on Recipe A+D simultaneously to chase a small uplift) is the canonical "ducktape-by-monolithic-orchestrator" case study. Key architectural shifts: - Skills (not slash commands) per superpowers convention - Convention-pathed reports at docs/skill-optimizer//... (visible + committable, like docs/superpowers/specs/) - Strict limited-context subagents for all generative work (writer, analyzer, optimizer, validator) — prevents tunnel-vision into ducktape patches - Validator subagent after every improvement (internal + optional external consistency) - Auto-pilot = natural chained invocation, not a separate orchestrator Co-Authored-By: Claude Opus 4.7 (1M context) --- docs/skill-optimizer-v1.4-spec.md | 428 ++++++++++++++++++++++++++++++ 1 file changed, 428 insertions(+) create mode 100644 docs/skill-optimizer-v1.4-spec.md diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md new file mode 100644 index 0000000..7a499f6 --- /dev/null +++ b/docs/skill-optimizer-v1.4-spec.md @@ -0,0 +1,428 @@ +# skill-optimizer v1.4 — design spec + +**Status:** approved (brainstormed 2026-05-12 with the +`superpowers:brainstorming` skill) +**Supersedes:** the v1.3 `skills/auto-improve-orchestrator/` monolithic +orchestrator subagent. +**Empirical basis:** team review of the v1.3 PR drafts (#1, #5, #6 +drafted via v1.3) surfaced four structural critiques of the v1.3 +approach (see "Motivation" below). + +## Goal + +Convert the auto-improve workflow from a **monolithic orchestrator +subagent** (v1.3) into **seven independent Claude Code skills** that +chain via the same pattern as the `superpowers` plugin (skill → +"invoke next skill" → next skill). Each skill produces a +human-reviewable report at a convention path. The user can invoke each +skill manually for explicit control, or ask the agent to chain them +auto-pilot style. + +Subagents dispatched by each skill operate under **strict limited +context**: they see only the inputs they need to do their job, never +the raw trial data or grader internals that would tempt them to +ducktape-patch. This is the load-bearing constraint that v1.3 lacked. + +The skill-optimizer engine itself (`run-suite`, graders, Docker +harness) stays unchanged — v1.4 only restructures the orchestration +layer that USES the engine. + +## Motivation + +Four critiques surfaced from team review of the v1.3 drafts: + +1. **PRs optimize for incremental numeric uplift, not principled + improvement.** A ~5–10 % score bump is shippable per v1.3's logic + even when the change is a ducktape patch (e.g., a rule restated + verbosely just to push one specific trial over the line). +2. **Test case quality is never validated.** v1.3 trusts the seeded + workbench. If the cases don't exercise the skill's real + responsibilities, the orchestrator optimizes for the wrong thing. +3. **Grader correctness is never validated.** v1.3 has a + "grader-vs-skill" check during iteration, but it only fires when + per-case-min crosses thresholds. Routine grader bugs (line drift, + keyword mismatch) get hidden. +4. **The improvement step has no anti-ducktape guardrail.** v1.3's + skill-iterate subagent sees the raw failed `findings.txt` and can + pattern-match a specific patch that satisfies the grader without + addressing the underlying weakness. (Concrete v1.3 example: + firecrawl iteration 1 regressed the original case from 1.0 → 0.44 + on gpt-5/gemini by piling on Recipe A + Recipe D simultaneously + to chase a small uplift; the orchestrator ran out of context + before catching the regression.) + +v1.4 addresses all four by **decomposing the workflow into discrete +skills, requiring human review between steps, and enforcing +limited-context constraints on every subagent that touches generative +work**. + +## Architecture overview + +**Inspired by the `superpowers` plugin convention** (skills, not slash +commands; description-based routing; "invoke next skill" chaining; +visible-at-conventional-path state files): + +```text +skills/ + skill-optimizer-investigate-functionality/SKILL.md # 1 + skill-optimizer-investigate-test-case/SKILL.md # 2 + skill-optimizer-investigate-submissions/SKILL.md # 3 (optional) + skill-optimizer-write-tests/SKILL.md # 4 + skill-optimizer-run-bench/SKILL.md # 5 + skill-optimizer-analyze-result/SKILL.md # 6 + skill-optimizer-improve-skill/SKILL.md # 7 + skill-optimizer/SKILL.md # existing — unchanged + subagents/ + research-functionality.md + research-submissions.md + test-writer.md + analyzer.md + optimizer.md + validator.md + references/ + workflow.md # the chain graph + recipes.md # accumulated lessons +``` + +**State at convention path** (visible + committable, per +`superpowers` precedent): + +```text +docs/skill-optimizer// + 01-functionality.md + 02-test-case.md + 03-submissions.md # only if step 3 ran + 04-tests-plan.md + workbench/ # eval suite produced by step 4 + 05-bench-results// # produced by step 5 + 06-analysis.md + 07-improvement-proposal.md + 07-validator-verdict.md + vendored-skill/ # the source skill, read-only after fetch +``` + +`` is `--` for upstream skills, or +`` for local skills. + +**Two contexts, one workflow:** + +- **Upstream skill:** user provides URL or `//`; + step 1 fetches; at the end, step 7 prompts "submit a PR?" and if yes + uses `03-submissions.md` to package. +- **Local skill:** user provides a path to a SKILL.md in their repo; + step 1 just reads; at the end, step 7 writes the modified skill in + place. No PR prompt. + +Difference is one prompt at the end of step 7; everything else is +identical. + +## Subagent constraints (the load-bearing principle) + +Every subagent that touches generative work runs under **strict +limited context**. The skill (operator-side) reads full report files +and dispatches the subagent with only the narrow chunks it needs. + +| Subagent | Sees | Does NOT see | Why | +|---|---|---|---| +| Functionality researcher (step 1) | Source skill files, web-search results | Existing analyses, existing tests | Pure research, no contamination | +| Submission researcher (step 3) | Repo files, gh-API outputs | Anything about the proposed change | Just upstream facts | +| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` | The skill's content, other test cases, the eval grader's matching logic | Prevents grader-hacking; prevents copying existing tests | +| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files) | Forces it to think about the SKILL, not the SOLUTIONS | +| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content | Raw failed trials, `findings.txt`, grader internals, test inputs | Forces principled improvement, not pattern-match patches | +| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs | Independent check; can't be biased by what the optimizer told itself | + +The skill (operator session) sees everything; subagents see slices. +This is the architectural fix for the "tunnel-vision into ducktape" +problem. + +## The 7 skills + +Each skill is a directory with `SKILL.md` (the instructions) plus +optional supporting files (subagent prompts, reference material). The +prose content of each SKILL.md will be filled in during implementation +— this spec defines the **interface** (input/output/behavior +contract), not the prose. + +### 1. `skill-optimizer-investigate-functionality` + +- **Description trigger:** when the user asks "what does this skill do", + "investigate this skill", "understand this skill" +- **Input:** source skill (URL or local path) +- **Output:** `docs/skill-optimizer//01-functionality.md` +- **Behavior:** fetch skill files (if URL); web-search the underlying + technology; identify trigger conditions, success criteria, key + terminology, intended audience. Write structured report covering: + what the skill does, who uses it, when it should fire, what tools it + depends on, what concepts the user must understand. +- **Dispatches:** functionality-researcher subagent (limited context: + source skill + targeted web fetches) +- **Handoff:** "Next, invoke `skill-optimizer-investigate-test-case`." + +### 2. `skill-optimizer-investigate-test-case` + +- **Description trigger:** "design tests for this skill", "propose + test cases", "what should we test" +- **Input:** `01-functionality.md` +- **Output:** `02-test-case.md` — ranked list of proposed test cases. + Each entry: name, what-it-tests (which responsibility), required + setup, expected agent behavior, grader spec, why-it-matters. +- **Behavior:** enumerate the skill's responsibilities from + functionality report; design 1–2 cases per responsibility; flag edge + cases (boundary conditions, error paths, common pitfalls); estimate + grader difficulty (deterministic check vs needs-real-tooling); rank + by importance + cost-to-build. +- **Dispatches:** none — analytical, operator session +- **User gate:** prompts the user to pick which subset of proposed + cases to actually build (write-tests acts on the picked subset) +- **Handoff:** "If upstream submission likely, invoke + `skill-optimizer-investigate-submissions` next. Then invoke + `skill-optimizer-write-tests` with the user's picked subset." + +### 3. `skill-optimizer-investigate-submissions` (OPTIONAL — upstream only) + +- **Description trigger:** "research PR conventions for this skill", + "what does the upstream repo require" +- **Input:** source slug (`//`) +- **Output:** `03-submissions.md` — license, CLA, frontmatter spec, + file-location rules, prefix taxonomy, PR-shape patterns from recent + merged PRs, branch target, rejection signals from closed-without- + merge PRs. (Same shape as v1.3's research-upstream subagent output.) +- **Behavior:** gh-CLI heavy (PR list, repo-file API, CONTRIBUTING, + sanity-test source, last 10 merged + last 5 closed-without-merge); + produce verbatim-pastable context block for the validator +- **Dispatches:** submission-researcher subagent (limited context: + public repo facts only) +- **Skipped when:** local skill, or user explicitly opts out +- **Handoff:** "Validator in `improve-skill` will read this for + external consistency check. Continue with `write-tests` if not done." + +### 4. `skill-optimizer-write-tests` + +- **Description trigger:** "build the tests", "implement the + workbench", "set up the eval" +- **Input:** `01-functionality.md`, `02-test-case.md` (operator picks + subset) +- **Output:** `04-tests-plan.md` (per-case implementation plan) + + `workbench/` (suite.yml, workspace files, graders, smoke check) +- **Behavior:** plan workbench structure → show user the plan → user + confirms → dispatch parallel subagents (one per test case, each + builds one workspace file + one grader) → run smoke check + (hand-crafted GOOD/BAD/EMPTY findings.txt fixtures against each + grader) → commit +- **Dispatches:** test-writer subagent per case (limited context: + single case spec + functionality report; does NOT see skill content + or other test cases — prevents grader-hacking and prevents copying) +- **Parallelizable:** each case independent; dispatch in a single + message +- **Handoff:** "Invoke `skill-optimizer-run-bench` to measure baseline." + +### 5. `skill-optimizer-run-bench` + +- **Description trigger:** "run the eval", "measure", "benchmark" +- **Input:** `workbench/` + source skill (vendored) +- **Output:** `05-bench-results//suite-result.json` + + per-trial traces + per-trial findings.txt +- **Behavior:** invoke skill-optimizer CLI + (`npx tsx /src/cli.ts run-suite ./workbench/suite.yml + --trials 3`); capture results +- **Dispatches:** none — direct CLI invocation +- **Note:** this is intentionally thin; the entire v1.3 run-suite + logic stays as-is in the CLI +- **Handoff:** "Invoke `skill-optimizer-analyze-result`." + +### 6. `skill-optimizer-analyze-result` + +- **Description trigger:** "analyze the results", "diagnose what + failed", "find structural weaknesses" +- **Input:** `05-bench-results//` + `workbench/` + source skill +- **Output:** `06-analysis.md` — structured per the format below +- **Behavior:** cluster failures (per-rule, per-model, per-trial, + per-pattern); separate flaky (single-trial randomness) from + systematic (repeated across trials and/or models); for each + systematic cluster, find the responsible skill section, hypothesize + cause, articulate the general principle that WOULD address it AND + the anti-patterns that would NOT (anti-ducktape gate); list + non-structural noise separately; **if no structural weakness can be + articulated, the report says so explicitly and the next step refuses + to fire** +- **Dispatches:** analyzer subagent (limited context: per-trial + findings + skill content + workbench cases; does NOT see the test + inputs themselves — forces focus on the SKILL, not the test data) +- **Handoff:** if at least one structural weakness identified → "Invoke + `skill-optimizer-improve-skill`." Otherwise → exit honestly ("no + structural weakness; no improvement warranted"). + +**`06-analysis.md` format (Option A — structured):** + +```markdown +## Structural weaknesses identified + +### Weakness 1: + +- **Pattern**: Across trials, systematically failed to + detect . Specifically: . +- **Hypothesized cause**: +- **Connects to skill section**: at +- **What WOULD address this**: +- **What WOULD NOT address this** (anti-ducktape gate): + +### Weakness 2: ... + +## Non-structural noise (ignored — not addressable) + +- 1 gemini transient API error on multi-redirect (not a pattern) +- 1 gpt-5 timeout (infrastructure, not skill) +``` + +### 7. `skill-optimizer-improve-skill` + +- **Description trigger:** "improve this skill", "fix the structural + weakness", "optimize" +- **Input:** `06-analysis.md` (REQUIRED — refuses if no weakness + identified), `01-functionality.md`, source skill, optionally + `03-submissions.md` +- **Output:** `07-improvement-proposal.md` (proposed diff + rationale + referencing the structural weakness) + `07-validator-verdict.md` + + modified skill file (if approved) +- **Behavior:** + 1. Refuse if `06-analysis.md` has no structural weakness — print + "no weakness to address" and exit + 2. Dispatch **optimizer subagent** with limited context (see + "Subagent constraints" table). Required: address the named + structural weakness using a general principle from the analysis, + NOT a pattern-match patch. Output the proposed diff + rationale + that explicitly references which named weakness it addresses. + 3. Dispatch **validator subagent** with limited context. + - **Internal consistency check:** does the proposed change make + sense given the skill's stated responsibilities in + `01-functionality.md`? Is the change additive vs destructive? + Is it general vs ducktape? + - **External consistency check (only if `03-submissions.md` + exists):** does the change conform to the upstream's PR rules + (frontmatter, file location, prefix taxonomy, additive-only, + etc.)? + - Verdict: `approve` / `needs-revision` / `reject` + 4. If `needs-revision`: optimizer revises (max 2 revision rounds). + If `reject`: surface honestly and exit. + 5. If `approve`: write the modified skill file + improvement + proposal report. +- **Dispatches:** optimizer + validator, both limited context, both + isolated from raw trial data +- **Handoff (local skill):** done. Modified skill written in place. +- **Handoff (upstream skill):** prompt user "submit a PR?" → if yes, + use `03-submissions.md` to package the proposed change as a PR + draft (similar to v1.3 packaging step but using the actual upstream + conventions, not the auto-pilot's reconstruction). + +## Auto-pilot mode + +There is no separate auto-pilot skill. **Auto-pilot is just chained +manual invocation.** The user (or the agent acting on user's request) +walks through the skills in order: 1 → 2 → (optional 3) → 4 → 5 → 6 → +7. Each skill's SKILL.md ends with a "Next, invoke X" instruction; +the agent uses the `Skill` tool to chain. + +For full auto-pilot, the user just says "optimize skill X end-to-end" +and the agent invokes all 7 in order, gating at the natural +user-review points (after step 2 the user picks tests; after step 7 +the user reviews the improvement proposal). + +The user-review gates are part of each skill's instructions, not +external orchestration. + +## Plugin packaging + +The 7 skills + subagent prompts + reference material ship as part of +the existing `skill-optimizer` Claude Code plugin +(`.claude-plugin/plugin.json`). They sit alongside the existing +`skills/skill-optimizer/SKILL.md`. + +**Codex / Cursor ports:** v1.4 implementation targets Claude Code +first. Codex and Cursor have similar skill concepts; porting is a +separate workstream (out of scope for v1.4). + +## Coexistence with v1.3 + +The v1.3 `skills/auto-improve-orchestrator/` skill is **deprecated but +retained** for backward compatibility during the v1.4 rollout: + +- v1.3 stays installed; existing context files at + `skills/auto-improve-orchestrator/references/contexts/` are kept + (they are useful inputs to v1.4's step 3 as initial drafts) +- v1.3's lessons.md becomes the seed for v1.4's `references/recipes.md` +- The v1.3 orchestrator is marked as superseded in its own SKILL.md + with a pointer to v1.4 + +After v1.4 is validated on 3+ skills, v1.3 can be removed in a +separate cleanup PR. + +## Acceptance criteria + +For v1.4 to be considered done: + +1. All 7 SKILL.md files exist with the interface described in this + spec, plus appropriate "Next, invoke X" handoff instructions +2. Subagent prompt templates exist in `skills/subagents/` with the + limited-context constraints described in the "Subagent constraints" + table +3. `references/workflow.md` documents the chain visually +4. `references/recipes.md` is seeded from v1.3's `lessons.md` +5. **End-to-end test on a local skill** (e.g., one of the existing + `skills/skill-optimizer/SKILL.md` or a small new skill in this + repo) — walks 1→2→4→5→6→7, produces all expected reports, modifies + the target skill, validator approves +6. **End-to-end test on an upstream skill** — re-run firecrawl (the + v1.3 regression case) under v1.4; expected: optimizer either + produces a principled fix OR honestly refuses (no regression + shipped) +7. Plugin manifest unchanged in shape (still + `.claude-plugin/plugin.json`); skills auto-discoverable via the + existing plugin loading mechanism + +## Out of scope (deferred to v1.5 or later) + +- **Codex / Cursor plugin ports** — v1.4 ships Claude Code first; + multi-IDE comes later +- **Programmatic auto-pilot wrapper** — the natural-language "walk + through all 7" pattern is sufficient for v1.4 +- **Per-recipe SKILL.md prose enrichment** — the implementation phase + fills in the prose; later iterations add more pattern guidance to + each SKILL.md based on accumulated usage. The spec defines the + interface, not the full prose. +- **Cost-tracking** — v1.4 uses the operator's Claude Code plan + (subagents free per plan) and the OpenRouter cost is tracked by + the existing CLI. No new cost-budget logic in the skills. +- **Cross-skill recipe sharing** — recipes from one skill's iteration + could inform another's analyzer. Possible future feature; not + designed in v1.4. + +## Open questions (tracked but deferred) + +1. **When does a SKILL.md grow recipes vs reference an external + recipes.md?** Initial answer: shared recipes go in + `references/recipes.md`; per-skill nuances go inline. Will revisit + based on usage. +2. **What if the optimizer subagent's revision loop exceeds the max + rounds (2)?** Initial answer: surface as `validator-rejected` and + require human intervention. Could add a "human-help-required" + status. Will revisit. +3. **How does step 6's analyzer subagent distinguish "no pattern" from + "pattern but I missed it"?** Initial answer: if the analyzer can't + articulate a weakness, the next step refuses to fire — we don't + force improvement. The user can re-invoke step 6 with hints if they + disagree with the analyzer's verdict. + +## Provenance + +- v1.3 design: `docs/auto-improve-skill-v1.3-spec.md` +- v1.3 implementation: branch `feat/auto-improve-skill-v1.3` + (PR #50) +- v1.3 PR drafts that surfaced the critiques: PR #51 (drafts #1, #5, + #6); the firecrawl regression (uncommitted, in the + `agent-abe33dd2c200c608a` worktree) is the canonical + "ducktape-by-monolithic-orchestrator" case study +- Brainstorming session: 2026-05-12 (this spec is the output) From a01d0cc1c5df7ad8fabc9ad8a45a16591bef5ec0 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 06:13:49 -0500 Subject: [PATCH 002/121] docs(v1.4-spec): PR-or-not decision moves to step 1 (was: end of step 7) User clarification: the agent asks 'optimize for PR submission?' at step 1 (when the user first provides an upstream skill), not at the end of step 7. This decision determines whether step 3 (investigate-submissions) runs at all. Changes: - 'Two contexts, one workflow' rewritten: PR decision at step 1, recorded in 01-functionality.md frontmatter as pr_submission_intent - Step 1 behavior: explicit PR question for upstream sources - Step 2 handoff: reads pr_submission_intent to decide whether to invoke step 3 - Step 3 'skipped when': now keyed on pr_submission_intent: false - Step 7 handoff: three branches (local / upstream-no-PR / upstream-yes-PR), no late prompts Co-Authored-By: Claude Opus 4.7 (1M context) --- docs/skill-optimizer-v1.4-spec.md | 71 +++++++++++++++++++++---------- 1 file changed, 48 insertions(+), 23 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 7a499f6..68f4b7b 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -106,15 +106,19 @@ docs/skill-optimizer// **Two contexts, one workflow:** -- **Upstream skill:** user provides URL or `//`; - step 1 fetches; at the end, step 7 prompts "submit a PR?" and if yes - uses `03-submissions.md` to package. -- **Local skill:** user provides a path to a SKILL.md in their repo; - step 1 just reads; at the end, step 7 writes the modified skill in - place. No PR prompt. - -Difference is one prompt at the end of step 7; everything else is -identical. +- **Upstream skill:** user provides URL or `//`. + Step 1 fetches AND asks the user up-front: "do you want to optimize + this skill for upstream PR submission?" If yes → step 3 + (`investigate-submissions`) runs after step 2, and step 7's validator + performs both internal AND external consistency checks. If no → + step 3 is skipped entirely; step 7 just validates internal + consistency and writes the improved skill to the vendored copy. +- **Local skill:** user provides a path to a SKILL.md in their repo. + Step 1 just reads (no fetch, no PR question). Step 3 never runs. + Step 7 writes the modified skill in place. + +The PR-or-not decision is made ONCE at step 1; subsequent steps know +from that decision whether step 3 will run. No late prompts. ## Subagent constraints (the load-bearing principle) @@ -149,14 +153,28 @@ contract), not the prose. "investigate this skill", "understand this skill" - **Input:** source skill (URL or local path) - **Output:** `docs/skill-optimizer//01-functionality.md` -- **Behavior:** fetch skill files (if URL); web-search the underlying - technology; identify trigger conditions, success criteria, key - terminology, intended audience. Write structured report covering: - what the skill does, who uses it, when it should fire, what tools it - depends on, what concepts the user must understand. +- **Behavior:** + 1. Identify whether the source is upstream (URL or + `//`) or local (path) + 2. **If upstream:** ASK the user up-front: "do you want to optimize + this skill for upstream PR submission?" Record the answer in the + report frontmatter as `pr_submission_intent: true|false`. This + answer determines whether step 3 (`investigate-submissions`) + fires later. + 3. **If local:** no PR question. Set `pr_submission_intent: false` + in the frontmatter. + 4. Fetch the skill files (if URL); vendor to + `vendored-skill/` for read-only use by downstream steps. + 5. Web-search the underlying technology; identify trigger + conditions, success criteria, key terminology, intended audience + 6. Write structured report covering: what the skill does, who uses + it, when it should fire, what tools it depends on, what concepts + the user must understand - **Dispatches:** functionality-researcher subagent (limited context: source skill + targeted web fetches) -- **Handoff:** "Next, invoke `skill-optimizer-investigate-test-case`." +- **Handoff:** "Next, invoke `skill-optimizer-investigate-test-case`. + Note: this report's `pr_submission_intent` field tells step 2's + handoff whether step 3 (`investigate-submissions`) should run." ### 2. `skill-optimizer-investigate-test-case` @@ -174,9 +192,11 @@ contract), not the prose. - **Dispatches:** none — analytical, operator session - **User gate:** prompts the user to pick which subset of proposed cases to actually build (write-tests acts on the picked subset) -- **Handoff:** "If upstream submission likely, invoke - `skill-optimizer-investigate-submissions` next. Then invoke - `skill-optimizer-write-tests` with the user's picked subset." +- **Handoff:** read `01-functionality.md`'s `pr_submission_intent` + field. If `true` → "Invoke `skill-optimizer-investigate-submissions` + next, then `skill-optimizer-write-tests`." If `false` → "Skip step + 3; invoke `skill-optimizer-write-tests` with the user's picked + subset." No late prompts — the decision was made at step 1. ### 3. `skill-optimizer-investigate-submissions` (OPTIONAL — upstream only) @@ -192,7 +212,9 @@ contract), not the prose. produce verbatim-pastable context block for the validator - **Dispatches:** submission-researcher subagent (limited context: public repo facts only) -- **Skipped when:** local skill, or user explicitly opts out +- **Skipped when:** `pr_submission_intent: false` in + `01-functionality.md` (i.e., local skill OR user said no to the PR + question at step 1) - **Handoff:** "Validator in `improve-skill` will read this for external consistency check. Continue with `write-tests` if not done." @@ -313,10 +335,13 @@ contract), not the prose. - **Dispatches:** optimizer + validator, both limited context, both isolated from raw trial data - **Handoff (local skill):** done. Modified skill written in place. -- **Handoff (upstream skill):** prompt user "submit a PR?" → if yes, - use `03-submissions.md` to package the proposed change as a PR - draft (similar to v1.3 packaging step but using the actual upstream - conventions, not the auto-pilot's reconstruction). +- **Handoff (upstream skill, `pr_submission_intent: false`):** done. + Modified skill written to vendored copy. No PR packaging. +- **Handoff (upstream skill, `pr_submission_intent: true`):** package + the proposed change as a PR draft using `03-submissions.md` (which + exists because step 3 ran). The draft includes the diff, the body, + caveats, and operator-steps-to-submit. No late "submit a PR?" + prompt — the decision was already made at step 1. ## Auto-pilot mode From d5173482098c25035f34d659d0e3728b88fa915d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 07:28:57 -0500 Subject: [PATCH 003/121] docs: implementation plan for skill-optimizer v1.4 15 tasks across 5 phases: - Phase A (4 tasks, mechanical): skill dir shells, subagents/refs dirs, v1.3 deprecation, recipes.md seed - Phase B (7 tasks, INTERACTIVE via writing-skills): one per SKILL.md, NOT subagent-driven per user request - Phase C (6 tasks, mechanical): subagent prompt templates with limited-context constraints - Phase D (1 task, mechanical): references/workflow.md chain diagram - Phase E (2 tasks, E2E validation): local skill + firecrawl re-run (the v1.3 regression case) Plan explicitly marks Phase B as not-for-subagent-driven-development; the user explicitly stated 'writing good skills is HARD' and wants interactive creation per skill. Co-Authored-By: Claude Opus 4.7 (1M context) --- docs/skill-optimizer-v1.4-plan.md | 1639 +++++++++++++++++++++++++++++ 1 file changed, 1639 insertions(+) create mode 100644 docs/skill-optimizer-v1.4-plan.md diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/skill-optimizer-v1.4-plan.md new file mode 100644 index 0000000..06c0af7 --- /dev/null +++ b/docs/skill-optimizer-v1.4-plan.md @@ -0,0 +1,1639 @@ +# skill-optimizer v1.4 Implementation Plan + +> **For agentic workers:** This plan has TWO execution modes: +> +> - **Phase A, C, D, E** — use superpowers:subagent-driven-development (mechanical; standard subagent-per-task with two-stage review). +> - **Phase B (per-skill SKILL.md creation)** — use superpowers:writing-skills INTERACTIVELY with the operator. Do NOT batch via subagent. Each B-task is "invoke writing-skills to create one SKILL.md, with the interface contract from the spec as the brief." +> +> Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Implement v1.4 — replace the v1.3 monolithic `auto-improve-orchestrator` with 7 independent Claude Code skills that chain via the `superpowers` plugin pattern. Each skill produces a human-reviewable report at a convention path; subagents dispatched by each skill run under strict limited context to prevent ducttape patches. + +**Architecture:** Skills (not slash commands), description-routed, output to `docs/skill-optimizer//-.md`. The skill-optimizer engine (`run-suite`, graders, Docker harness) stays unchanged. The v1.3 `auto-improve-orchestrator/` skill is deprecated-but-retained for backward compatibility. + +**Tech Stack:** Markdown (skill prompts), gray-matter (frontmatter parsing, already a project dep), Claude Code Skill tool + Agent tool, bash (validation scripts). + +**Working dir:** `.claude/worktrees/v1.4-spec/` on branch `feat/skill-optimizer-v1.4`. + +**Spec:** `docs/skill-optimizer-v1.4-spec.md` (committed at `a01d0cc`). + +--- + +## File Structure + +After this plan executes: + +```text +skills/ + skill-optimizer-investigate-functionality/SKILL.md # Phase B (interactive) + skill-optimizer-investigate-test-case/SKILL.md # Phase B (interactive) + skill-optimizer-investigate-submissions/SKILL.md # Phase B (interactive) + skill-optimizer-write-tests/SKILL.md # Phase B (interactive) + skill-optimizer-run-bench/SKILL.md # Phase B (interactive) + skill-optimizer-analyze-result/SKILL.md # Phase B (interactive) + skill-optimizer-improve-skill/SKILL.md # Phase B (interactive) + subagents/ # Phase C (mechanical) + research-functionality.md + research-submissions.md + test-writer.md + analyzer.md + optimizer.md + validator.md + references/ # Phase D (mechanical) + workflow.md + recipes.md # seeded from v1.3 lessons.md + auto-improve-orchestrator/SKILL.md # MODIFIED: add deprecation banner + +docs/ + skill-optimizer-v1.4-spec.md # already committed + skill-optimizer-v1.4-plan.md # THIS file + skill-optimizer-v1.4-validation.md # written by Phase E +``` + +Per-skill state files (produced at runtime, NOT created by this plan; convention only): + +```text +docs/skill-optimizer// + 01-functionality.md + 02-test-case.md + 03-submissions.md # only if pr_submission_intent: true + 04-tests-plan.md + workbench/ + 05-bench-results// + 06-analysis.md + 07-improvement-proposal.md + 07-validator-verdict.md + vendored-skill/ # for upstream skills only +``` + +Each file's responsibility: + +- **`skills/skill-optimizer-/SKILL.md`** — frontmatter (`name:`, `description:`) + invocation instructions + dispatch-subagent logic + handoff-to-next-skill pointer. Always-loaded by Claude Code when the skill is invoked. +- **`skills/subagents/.md`** — narrow-context prompt template that the corresponding skill loads, substitutes inputs into, and dispatches via Agent tool. +- **`skills/references/workflow.md`** — human-readable chain diagram + per-skill brief, for operators understanding how the 7 skills compose. +- **`skills/references/recipes.md`** — accumulated lessons learned (Recipe A–E, grader patterns G1–G6, etc.), read by the analyzer and optimizer subagents. + +--- + +## Phase A — Structural setup (4 tasks, mechanical, subagent-driven-development OK) + +### Task A1: Create the 7 skill directories with shell SKILL.md files + +**Files:** Create 7 directories + 7 shell SKILL.md files. + +- [ ] **Step 1: Create directories** + +```bash +cd /home/yuqing/Documents/Code/skill-optimizer/.claude/worktrees/v1.4-spec +for verb in investigate-functionality investigate-test-case investigate-submissions write-tests run-bench analyze-result improve-skill; do + mkdir -p skills/skill-optimizer-${verb} +done +ls skills/ | grep skill-optimizer- +``` + +Expected: 7 new directories listed (plus the existing `skill-optimizer/`). + +- [ ] **Step 2: Write shell SKILL.md files (frontmatter only, body placeholder for Phase B)** + +For each of the 7 verbs, write `skills/skill-optimizer-/SKILL.md` with this content (use the verb-specific name and a 1-line description from the spec): + +```markdown +--- +name: skill-optimizer- +description: > +--- + +# skill-optimizer- + + + +``` + +Verb → description mapping (from `docs/skill-optimizer-v1.4-spec.md`): + +```bash +# Run this script to populate all 7 shells: +declare -A DESCS=( + [investigate-functionality]="Use when the user asks to understand or investigate what a skill does — fetches skill source, web-searches the underlying technology, writes a functionality report. Also asks the user (for upstream skills) whether to target upstream PR submission." + [investigate-test-case]="Use when the user wants to design or propose test cases for a skill — enumerates the skill's responsibilities and proposes ranked test cases the user can pick from." + [investigate-submissions]="Use when the user wants to research a skill's upstream PR conventions — produces a context block with license, CLA, frontmatter spec, file-location rules, PR-shape patterns, and rejection signals." + [write-tests]="Use when the user wants to build / implement the eval workbench for a skill — plans the workbench, asks for user confirmation, then dispatches parallel narrow-context subagents to write each test case + grader." + [run-bench]="Use when the user wants to run the eval suite for a skill and capture results — thin wrapper around the skill-optimizer CLI's run-suite command." + [analyze-result]="Use when the user wants to analyze bench results and identify structural weaknesses — clusters failures, distinguishes systematic from flaky, and produces a structured analysis with anti-ducttape gates." + [improve-skill]="Use when the user wants to improve a skill based on identified structural weakness — dispatches optimizer + validator subagents under strict limited context; refuses if no structural weakness in the analysis." +) + +for verb in "${!DESCS[@]}"; do + cat > "skills/skill-optimizer-${verb}/SKILL.md" < + +EOF +done +``` + +- [ ] **Step 3: Verify frontmatter parses for all 7** + +```bash +node -e " +const matter = require('gray-matter'); +const fs = require('fs'); +const verbs = ['investigate-functionality', 'investigate-test-case', 'investigate-submissions', 'write-tests', 'run-bench', 'analyze-result', 'improve-skill']; +for (const v of verbs) { + const p = 'skills/skill-optimizer-' + v + '/SKILL.md'; + const m = matter(fs.readFileSync(p, 'utf-8')); + if (m.data.name !== 'skill-optimizer-' + v || typeof m.data.description !== 'string' || m.data.description.length < 50) { + console.error('FAIL:', p, m.data); + process.exit(1); + } + console.log('OK:', p); +} +" +``` + +Expected: 7 `OK:` lines. + +- [ ] **Step 4: Commit** + +```bash +git add skills/skill-optimizer-*/SKILL.md +git commit -m "feat(v1.4): create 7 skill directory shells with frontmatter + +Body content for each SKILL.md will be written interactively in Phase B +via superpowers:writing-skills, with the interface contract from +docs/skill-optimizer-v1.4-spec.md as the brief. This commit just lays +down the structural skeleton and discoverable frontmatter." +``` + +--- + +### Task A2: Create `subagents/` and `references/` directories + +**Files:** Create 2 directories. + +- [ ] **Step 1: Create directories** + +```bash +mkdir -p skills/subagents skills/references +ls skills/subagents skills/references +``` + +Expected: both directories exist and are empty. + +- [ ] **Step 2: Add `.gitkeep` placeholders so the empty dirs are tracked** + +```bash +touch skills/subagents/.gitkeep skills/references/.gitkeep +``` + +- [ ] **Step 3: Commit** + +```bash +git add skills/subagents/.gitkeep skills/references/.gitkeep +git commit -m "chore(v1.4): create subagents/ and references/ dirs (populated in Phase C+D)" +``` + +--- + +### Task A3: Add deprecation banner to v1.3 `auto-improve-orchestrator/SKILL.md` + +**Files:** + +- Modify: `skills/auto-improve-orchestrator/SKILL.md` (first lines after frontmatter) + +- [ ] **Step 1: Read the current first lines** + +```bash +head -20 skills/auto-improve-orchestrator/SKILL.md +``` + +- [ ] **Step 2: Insert a deprecation banner immediately after the closing `---` of the frontmatter** + +Use Edit tool to insert this block after the second `---` line (closing the frontmatter): + +```markdown + +> **DEPRECATED — see v1.4.** This skill is superseded by the 7 +> independent skills under `skills/skill-optimizer-*` (v1.4). See +> `docs/skill-optimizer-v1.4-spec.md` for the architecture rationale +> and `skills/references/workflow.md` for the new chain. This skill +> is retained for backward compatibility during v1.4 rollout; it +> will be removed in a separate cleanup PR after v1.4 is validated +> on 3+ skills. + +``` + +- [ ] **Step 3: Verify the banner landed** + +```bash +grep -A 8 "DEPRECATED — see v1.4" skills/auto-improve-orchestrator/SKILL.md +``` + +Expected: the banner appears once, with the full text. + +- [ ] **Step 4: Commit** + +```bash +git add skills/auto-improve-orchestrator/SKILL.md +git commit -m "docs(orchestrator): mark v1.3 orchestrator as deprecated-but-retained for v1.4 rollout" +``` + +--- + +### Task A4: Seed `references/recipes.md` from v1.3 `lessons.md` + +**Files:** + +- Create: `skills/references/recipes.md` (seeded from `skills/auto-improve-orchestrator/references/lessons.md`) + +- [ ] **Step 1: Copy the lessons content, removing v1.3-specific framing** + +```bash +cp skills/auto-improve-orchestrator/references/lessons.md \ + skills/references/recipes.md +``` + +- [ ] **Step 2: Replace v1.3-specific phrases** + +Use Edit tool on `skills/references/recipes.md`: + +- Replace `auto-improve-skill-lessons.md` → `skills/references/recipes.md` (wherever it appears) +- Replace `auto-improve-skill` (the wrapper) → `skill-optimizer` (the engine) +- Replace `auto-improve-orchestrator` references with "v1.3 orchestrator (now deprecated)" + +If grep finds no instances, the file is already neutral — skip the edits. + +```bash +grep -n "auto-improve-skill-lessons.md\|auto-improve-orchestrator" skills/references/recipes.md | head -10 +``` + +For each match, use Edit tool to make the replacement. + +- [ ] **Step 3: Add a header note explaining what this file IS in v1.4** + +Prepend this paragraph to the top of `skills/references/recipes.md` (right after the H1): + +```markdown +> **For v1.4 skills:** This is the shared recipes library. The +> analyzer subagent (step 6) reads this to recognize known failure +> patterns. The optimizer subagent (step 7) reads this to choose a +> principled fix (Recipe A–E). Add new recipes here when a v1.4 run +> surfaces a generalizable pattern. Per-skill nuances belong in the +> individual run's `06-analysis.md` rather than here. +``` + +- [ ] **Step 4: Verify the file parses + contains the seeded content** + +```bash +test -f skills/references/recipes.md +wc -l skills/references/recipes.md +head -5 skills/references/recipes.md +grep -c "^## " skills/references/recipes.md +``` + +Expected: file exists, ≥ 100 lines, has multiple `##` section headings. + +- [ ] **Step 5: Commit** + +```bash +git rm skills/subagents/.gitkeep # no longer needed once Phase C lands; keep here for now if Phase C hasn't run +# Actually skip the .gitkeep removal — keep simple: +git add skills/references/recipes.md +git commit -m "feat(v1.4): seed references/recipes.md from v1.3 lessons.md + +Adds v1.4-specific header explaining how the analyzer and optimizer +subagents will use this file. Content is the v1.3 lessons.md verbatim +(Recipe A-E + grader patterns G1-G6 + run-record protocol) with +v1.3-specific wrapper references neutralized." +``` + +--- + +## Phase B — Interactive SKILL.md creation (7 tasks, USE writing-skills, NOT subagent-driven) + +**Mode:** Each Task B uses `superpowers:writing-skills` interactively. The operator runs the skill, answers writing-skills' clarifying questions, iterates on the output, commits when satisfied. + +**Do NOT dispatch these as autonomous subagents** — writing-skills is built for human-in-loop refinement, and the user explicitly stated "writing good skills is HARD." + +For EACH Task B, the input to writing-skills is the interface contract for that skill from `docs/skill-optimizer-v1.4-spec.md` section "The 7 skills". Paste the relevant section as the writing-skills brief, plus the constraints from "Subagent constraints" if the skill dispatches a subagent. + +--- + +### Task B1: SKILL.md for `skill-optimizer-investigate-functionality` + +**Files:** + +- Modify: `skills/skill-optimizer-investigate-functionality/SKILL.md` (replace the placeholder body) + +- [ ] **Step 1: Invoke writing-skills with the spec excerpt as the brief** + +In your Claude Code session, invoke: + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-investigate-functionality/. + +Interface contract (from docs/skill-optimizer-v1.4-spec.md "The 7 skills" → §1): + +- Description trigger: "investigate / understand / explain what this skill does (and the context around it)" +- Input: source skill (URL or local path) +- Output: docs/skill-optimizer//01-functionality.md +- Behavior: + 1. Identify whether the source is upstream (URL or owner/repo/skill-id) or local (path) + 2. If upstream: ASK the user up-front "do you want to optimize this skill for upstream PR submission?" Record the answer in the report frontmatter as pr_submission_intent: true|false + 3. If local: no PR question. Set pr_submission_intent: false + 4. Fetch the skill files (if URL); vendor to vendored-skill/ for read-only use by downstream steps + 5. Web-search the underlying technology; identify trigger conditions, success criteria, key terminology, intended audience + 6. Write structured report covering: what the skill does, who uses it, when it should fire, what tools it depends on, what concepts the user must understand +- Dispatches: functionality-researcher subagent (skills/subagents/research-functionality.md). Limited context: source skill + targeted web fetches. +- Handoff: "Next, invoke skill-optimizer-investigate-test-case. Note: this report's pr_submission_intent field tells step 2's handoff whether step 3 (investigate-submissions) should run." + +The frontmatter (name + description) is already in place from Task A1 — keep it, only replace the body. + +Examples to draw from for shape: the existing superpowers/skills/brainstorming/SKILL.md for "skill that asks user questions, persists state, hands off to next skill" pattern. +``` + +Iterate with writing-skills until the SKILL.md body is satisfactory. Verify the frontmatter is unchanged and the body matches the spec contract. + +- [ ] **Step 2: Verify the file** + +```bash +node -e " +const matter = require('gray-matter'); +const m = matter(require('fs').readFileSync('skills/skill-optimizer-investigate-functionality/SKILL.md', 'utf-8')); +console.log('name:', m.data.name); +console.log('description chars:', m.data.description.length); +console.log('body lines:', m.content.split('\\n').length); +" +``` + +Expected: `name: skill-optimizer-investigate-functionality`, description still present, body lines > 30. + +- [ ] **Step 3: Commit** + +```bash +git add skills/skill-optimizer-investigate-functionality/SKILL.md +git commit -m "feat(skill-optimizer-investigate-functionality): SKILL.md body via writing-skills" +``` + +--- + +### Task B2: SKILL.md for `skill-optimizer-investigate-test-case` + +**Files:** + +- Modify: `skills/skill-optimizer-investigate-test-case/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills with the spec excerpt as the brief** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-investigate-test-case/. + +Interface contract (from docs/skill-optimizer-v1.4-spec.md §2): + +- Description trigger: "design tests for this skill", "propose test cases", "what should we test" +- Input: docs/skill-optimizer//01-functionality.md +- Output: docs/skill-optimizer//02-test-case.md — ranked list of proposed test cases. Each entry: name, what-it-tests (which responsibility), required setup, expected agent behavior, grader spec, why-it-matters +- Behavior: enumerate the skill's responsibilities from functionality report; design 1–2 cases per responsibility; flag edge cases; estimate grader difficulty; rank by importance + cost-to-build +- Dispatches: none (analytical, operator session) +- User gate: prompts the user to pick which subset of proposed cases to actually build (write-tests acts on the picked subset) +- Handoff: read 01-functionality.md's pr_submission_intent field. If true → "Invoke skill-optimizer-investigate-submissions next, then skill-optimizer-write-tests." If false → "Skip step 3; invoke skill-optimizer-write-tests with the user's picked subset." + +Frontmatter is already in place from Task A1. +``` + +- [ ] **Step 2: Verify + commit (same pattern as B1)** + +```bash +node -e "const m = require('gray-matter')(require('fs').readFileSync('skills/skill-optimizer-investigate-test-case/SKILL.md', 'utf-8')); console.log('lines:', m.content.split('\\n').length);" +git add skills/skill-optimizer-investigate-test-case/SKILL.md +git commit -m "feat(skill-optimizer-investigate-test-case): SKILL.md body via writing-skills" +``` + +--- + +### Task B3: SKILL.md for `skill-optimizer-investigate-submissions` + +**Files:** + +- Modify: `skills/skill-optimizer-investigate-submissions/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-investigate-submissions/. + +Interface contract (from spec §3, OPTIONAL — upstream-only): + +- Description trigger: "research PR conventions for this skill", "what does the upstream repo require" +- Input: source slug // (from 01-functionality.md frontmatter) +- Output: docs/skill-optimizer//03-submissions.md — license, CLA, frontmatter spec, file-location rules, prefix taxonomy, PR-shape patterns from last 10 merged PRs, branch target, rejection signals from last 5 closed-without-merge PRs +- Behavior: gh-CLI heavy (PR list, repo-file API, CONTRIBUTING, sanity-test source, last 10 merged + last 5 closed); produce verbatim-pastable context block for the validator +- Dispatches: submission-researcher subagent (skills/subagents/research-submissions.md). Limited context: public repo facts only. +- Skipped when: pr_submission_intent: false in 01-functionality.md (this skill should detect that and exit cleanly with a "skip" message) +- Handoff: "Validator in skill-optimizer-improve-skill will read this for external consistency check. Continue with skill-optimizer-write-tests if not done." + +Reference example: the existing tools/auto-improve-contexts/*.md files (now moved to skills/auto-improve-orchestrator/references/contexts/) are good examples of what 03-submissions.md should contain. +``` + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-investigate-submissions/SKILL.md +git commit -m "feat(skill-optimizer-investigate-submissions): SKILL.md body via writing-skills" +``` + +--- + +### Task B4: SKILL.md for `skill-optimizer-write-tests` + +**Files:** + +- Modify: `skills/skill-optimizer-write-tests/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-write-tests/. + +Interface contract (from spec §4): + +- Description trigger: "build the tests", "implement the workbench", "set up the eval" +- Inputs: docs/skill-optimizer//01-functionality.md, docs/skill-optimizer//02-test-case.md (operator picks subset) +- Outputs: docs/skill-optimizer//04-tests-plan.md (per-case implementation plan) + docs/skill-optimizer//workbench/ (suite.yml, workspace/, checks/, smoke-graders.mjs) +- Behavior: + 1. Plan workbench structure (suite.yml shape, file layout, grader patterns from references/recipes.md) + 2. Show user the plan, ask for confirmation + 3. Dispatch parallel subagents (one per picked test case) — each builds one workspace file + one grader + 4. Run smoke check (hand-crafted GOOD/BAD/EMPTY findings.txt fixtures against each grader) — must pass before commit +- Dispatches: test-writer subagent per case (skills/subagents/test-writer.md). LIMITED CONTEXT per spec: single case spec + functionality report; does NOT see skill content or other test cases. +- Parallelizable: each case independent; dispatch in a single message +- Handoff: "Invoke skill-optimizer-run-bench to measure baseline." +``` + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-write-tests/SKILL.md +git commit -m "feat(skill-optimizer-write-tests): SKILL.md body via writing-skills" +``` + +--- + +### Task B5: SKILL.md for `skill-optimizer-run-bench` + +**Files:** + +- Modify: `skills/skill-optimizer-run-bench/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-run-bench/. + +Interface contract (from spec §5): + +- Description trigger: "run the eval", "measure", "benchmark" +- Inputs: docs/skill-optimizer//workbench/ + source skill (vendored) +- Output: docs/skill-optimizer//05-bench-results//suite-result.json + per-trial traces + per-trial findings.txt +- Behavior: invoke skill-optimizer CLI: cd into workbench/, source the repo's .env, run `npx tsx /src/cli.ts run-suite ./suite.yml --trials 3`, capture results into docs/skill-optimizer//05-bench-results// +- Dispatches: none (direct CLI invocation) +- Note: this is the THINNEST skill in v1.4 — most v1.3 run-suite logic stays as-is in the CLI. SKILL.md primarily documents the invocation pattern + how to handle long-running runs. +- Handoff: "Invoke skill-optimizer-analyze-result." + +Reference: the existing src/cli.ts run-suite command + the v1.3 orchestrator's Phase 3 logic at skills/auto-improve-orchestrator/prompts/orchestrator.md. +``` + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-run-bench/SKILL.md +git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via writing-skills" +``` + +--- + +### Task B6: SKILL.md for `skill-optimizer-analyze-result` + +**Files:** + +- Modify: `skills/skill-optimizer-analyze-result/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-analyze-result/. + +Interface contract (from spec §6 — this is the LOAD-BEARING skill that determines downstream improvement quality): + +- Description trigger: "analyze the results", "diagnose what failed", "find structural weaknesses" +- Inputs: docs/skill-optimizer//05-bench-results// + workbench/ + source skill +- Output: docs/skill-optimizer//06-analysis.md — STRUCTURED per Option A format (see below) +- Behavior: + 1. Cluster failures (per-rule, per-model, per-trial, per-pattern) + 2. Separate FLAKY (single-trial randomness) from SYSTEMATIC (repeated across trials and/or models) + 3. For each systematic cluster: find responsible skill section, hypothesize cause, articulate "what WOULD address" (general principle) + "what WOULD NOT address" (anti-ducktape gate) + 4. List non-structural noise separately + 5. If no structural weakness can be articulated: report explicitly that no weakness was found; the next step (improve-skill) will refuse to fire — this is the honest-no-fabricated-uplift behavior +- Dispatches: analyzer subagent (skills/subagents/analyzer.md). LIMITED CONTEXT per spec: per-trial findings + skill content + workbench cases; does NOT see the test inputs themselves (forces focus on SKILL, not solutions). +- Handoff: if ≥1 structural weakness → "Invoke skill-optimizer-improve-skill." If none → exit honestly. + +06-analysis.md format (Option A — structured): + +```markdown +## Structural weaknesses identified + +### Weakness 1: + +- **Pattern**: Across trials, systematically failed to detect . Specifically: . +- **Hypothesized cause**: +- **Connects to skill section**: at +- **What WOULD address this**: +- **What WOULD NOT address this** (anti-ducktape gate): + +## Non-structural noise (ignored — not addressable) +- +``` + +``` + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-analyze-result/SKILL.md +git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via writing-skills" +``` + +--- + +### Task B7: SKILL.md for `skill-optimizer-improve-skill` + +**Files:** + +- Modify: `skills/skill-optimizer-improve-skill/SKILL.md` + +- [ ] **Step 1: Invoke writing-skills** + +```text +/skill superpowers:writing-skills +``` + +Brief: + +```text +Create SKILL.md body for skills/skill-optimizer-improve-skill/. + +Interface contract (from spec §7): + +- Description trigger: "improve this skill", "fix the structural weakness", "optimize" +- Inputs: docs/skill-optimizer//06-analysis.md (REQUIRED — refuses if no weakness identified), 01-functionality.md, source skill, optionally 03-submissions.md +- Outputs: docs/skill-optimizer//07-improvement-proposal.md (diff + rationale referencing the structural weakness) + 07-validator-verdict.md + modified skill file (if approved) +- Behavior: + 1. Refuse if 06-analysis.md has no structural weakness — print "no weakness to address" and exit cleanly + 2. Dispatch OPTIMIZER subagent (skills/subagents/optimizer.md) with limited context: sees analysis report + functionality + skill content; does NOT see raw failures, grader logic, test inputs. MUST address named structural weakness using a general principle (NOT a pattern-match patch). Output proposed diff + rationale that explicitly references which named weakness it addresses. + 3. Dispatch VALIDATOR subagent (skills/subagents/validator.md) with limited context: sees BEFORE skill + AFTER skill + 01-functionality + 03-submissions (if exists). Two-part check: + - INTERNAL consistency: does the change make sense given the skill's stated responsibilities? Additive vs destructive? General vs ducttape? + - EXTERNAL consistency (only if 03-submissions.md exists): does change conform to upstream PR rules (frontmatter, file location, prefix taxonomy, additive-only, etc.)? + - Verdict: approve / needs-revision / reject + 4. If needs-revision: optimizer revises (max 2 revision rounds total). + 5. If reject: surface verdict honestly and exit. + 6. If approve: write modified skill file + improvement proposal report. +- Dispatches: optimizer + validator, both limited context, both isolated from raw trial data. +- Handoff: + - Local skill: done — modified skill written in place + - Upstream skill, pr_submission_intent: false: done — modified skill written to vendored copy. No PR packaging. + - Upstream skill, pr_submission_intent: true: package the proposed change as a PR draft using 03-submissions.md. The draft includes diff, body, caveats, operator-steps-to-submit. No late "submit a PR?" prompt — decision was made at step 1. +``` + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-improve-skill/SKILL.md +git commit -m "feat(skill-optimizer-improve-skill): SKILL.md body via writing-skills" +``` + +--- + +## Phase C — Subagent prompt templates (6 tasks, mechanical, subagent-driven-development OK) + +Each subagent prompt template is loaded by the corresponding skill (Phase B), substituted with templated inputs, and dispatched via Agent tool. All must enforce the limited-context constraints from the spec's "Subagent constraints" table. + +### Task C1: `subagents/research-functionality.md` + +**Files:** Create `skills/subagents/research-functionality.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/research-functionality.md` with this exact content: + +````markdown +# Sub-subagent: research a skill's functionality + +You are dispatched to research what a single agent skill does and the context around it. Produce a structured functionality report. + +## Inputs (templated) + +- `${SKILL_SOURCE}` — URL or local path to the source skill +- `${OUTPUT_PATH}` — where to write the report (typically `docs/skill-optimizer//01-functionality.md`) +- `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 by the parent skill + +## Tools allowed + +- Read (skill source files) +- WebFetch (technology docs, vendor sites) +- WebSearch (broader context) +- Glob (find skill files) +- Bash (gh CLI for upstream repos if URL is github) +- Write (the report) + +## Tools NOT allowed + +You operate under limited context per `docs/skill-optimizer-v1.4-spec.md` "Subagent constraints" table. You see ONLY the source skill + targeted web fetches. You do NOT see: + +- Existing analyses or improvement proposals +- Existing tests for this skill +- Failure data from prior runs + +## What to produce + +Write `${OUTPUT_PATH}` with this structure: + +```markdown +--- +skill_source: ${SKILL_SOURCE} +pr_submission_intent: ${PR_SUBMISSION_INTENT} +classification: +--- + +# Functionality report: + +## What the skill does + +<2–3 paragraph summary of the skill's purpose, what it teaches the agent to do, what problem it solves> + +## Who uses it + + + +## When it should fire (trigger conditions) + + + +## What it depends on + + + +## Key concepts & terminology + + + +## Success criteria (what "good usage" looks like) + + + +## Anti-patterns (common failures) + + + +## Source files inventory + + +``` + +## Commit + +After writing the report, commit on the current branch: + +```bash +git add ${OUTPUT_PATH} +git commit -m "docs(skill-optimizer): functionality report for " +``` + +## Return + +Return to the calling skill a brief summary (under 200 words) covering: classification, dependency-flag list, and any blockers (e.g., "could not fetch source", "skill source is malformed"). +```` + +- [ ] **Step 2: Verify template variables present** + +```bash +grep -E '\$\{(SKILL_SOURCE|OUTPUT_PATH|PR_SUBMISSION_INTENT)\}' skills/subagents/research-functionality.md | wc -l +``` + +Expected: ≥ 4 occurrences. + +- [ ] **Step 3: Commit** + +```bash +git add skills/subagents/research-functionality.md +git commit -m "feat(v1.4-subagents): research-functionality prompt template" +``` + +--- + +### Task C2: `subagents/research-submissions.md` + +**Files:** Create `skills/subagents/research-submissions.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/research-submissions.md` with this exact content (adapted from v1.3's `prompts/research-upstream.md`): + +````markdown +# Sub-subagent: research upstream PR submission conventions + +You are dispatched to research one upstream repo's contribution conventions and produce a context block the validator and packaging steps will use. + +## Inputs (templated) + +- `${SLUG}` — `//` +- `${OUTPUT_PATH}` — typically `docs/skill-optimizer//03-submissions.md` + +## Tools allowed + +- Bash (gh CLI) +- WebFetch +- Read +- Write + +## Tools NOT allowed (limited context constraint) + +You see ONLY public repo facts. You do NOT see: + +- Any proposed change to the skill +- Test data or eval results +- The user's intent for the PR beyond "they intend to submit one" + +## What to research + +For the upstream repo at `${SLUG}`: + +1. License + CLA requirements (read LICENSE, CONTRIBUTING.md) +2. Default branch + branch-target convention (main vs next vs develop) +3. CI workflow gates (`.github/workflows/*.yml`) +4. Frontmatter spec for skill files (read sanity-test source if any) +5. File-location rules (where new content goes; what files are off-limits) +6. Prefix taxonomy (if references/ subdirectory has naming convention) +7. Last 10 merged PRs to this skill (or repo): file count, body shape, commit-message convention, time-to-merge +8. Last 5 closed-without-merge PRs: rejection signals (CLA-missing? shape-novel? discussion-first-gate?) +9. Other downstream consumers (gh search for raw URLs, install scripts, repo's own README) + +## What to produce + +Write `${OUTPUT_PATH}` with the same structure as the v1.3 examples at `skills/auto-improve-orchestrator/references/contexts/*.md`: + +```markdown +# Upstream submission context: ${SLUG} + +## Repository facts +- Repo, license, CLA, maintainers, merge style, CI, discovery-index/downstream-sync + +## Hard constraints (additive-only PR) +- File-location rules, prefix taxonomy, what not to modify, version bump rules + +## Frontmatter spec +- Exact required fields, allowed values, format (string vs YAML list, etc.) + +## Content shape template (copy-and-fill) +- A representative additive change with placeholders + +## Optimization target file +- Where the skill change should land + +## Architecture intent +- Why the upstream organizes things the way it does (informs validator's reasoning) + +## Risk profile +- LOW / MEDIUM / HIGH per category (additive vs restructure, etc.) + +## Pre-submit checklist +- What the PR-packaging step must verify before submission + +## Useful URLs +- Source-of-truth files in the upstream repo +``` + +## Commit + +```bash +git add ${OUTPUT_PATH} +git commit -m "docs(skill-optimizer): submissions context for ${SLUG}" +``` + +## Return + +Under 400 words. Include: license + CLA verdict; recommended branch target; risk profile; the verbatim-pastable context block path. +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(SLUG|OUTPUT_PATH)\}' skills/subagents/research-submissions.md | wc -l +# Expected: ≥ 4 + +git add skills/subagents/research-submissions.md +git commit -m "feat(v1.4-subagents): research-submissions prompt template" +``` + +--- + +### Task C3: `subagents/test-writer.md` + +**Files:** Create `skills/subagents/test-writer.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/test-writer.md` with this exact content: + +````markdown +# Sub-subagent: write one test case for the eval workbench + +You are dispatched to build EXACTLY ONE test case (one workspace file + one grader) for a skill's eval workbench. You will be one of N parallel test-writer subagents, each independently writing their own case. + +## Inputs (templated) + +- `${CASE_SPEC}` — verbatim text block from `02-test-case.md` describing this case (name, what-it-tests, required setup, expected agent behavior, grader spec, why-it-matters) +- `${FUNCTIONALITY_REPORT_EXCERPT}` — the relevant sections of `01-functionality.md` that pertain to the responsibility this case exercises +- `${WORKBENCH_DIR}` — typically `docs/skill-optimizer//workbench/` +- `${CASE_NAME}` — short kebab-case identifier (e.g., `review-update-without-where`) + +## Tools allowed + +- Read (workbench README + suite.yml for case-structure conventions; reference recipes) +- Write (the new workspace file + grader) +- Bash (smoke-check verification) + +## Tools NOT allowed (limited context constraint) + +You operate under strict limited context per spec. You see ONLY the single case spec + the functionality excerpt. You do NOT see: + +- The source skill's content (would tempt grader-hacking) +- Other test cases (would tempt copy-and-modify) +- Any existing failure data or analyses +- The eval grader matching logic in `_grader-utils.mjs`'s internals beyond the public API + +## What to produce + +1. **Workspace file** at `${WORKBENCH_DIR}/workspace/.` — realistic input the agent will operate on, with seeded conditions matching what the case is testing +2. **Grader** at `${WORKBENCH_DIR}/checks/grade-${CASE_NAME}-findings.mjs` — JavaScript module that imports from `_grader-utils.mjs` (use `gradeFindings`, `looseRange`, `fuzzyKeyword`, `tolerantKeyword` per `references/recipes.md` G1-G6) and checks the agent's output against expected behavior + +## Smoke check (REQUIRED before commit) + +After writing, hand-craft a 3-fixture smoke check: + +- **GOOD findings.txt** — what the agent SHOULD produce. Run your grader against it. Expect `pass: true, score: 1`. +- **BAD findings.txt** — missing 1–2 of the expected violations. Expect `pass: false, score < 1`. +- **EMPTY findings.txt** — no output at all. Expect `pass: false, score: 0`. + +Add these as inline test cases in `${WORKBENCH_DIR}/checks/smoke-graders.mjs` (or append if file exists). Run: + +```bash +node ${WORKBENCH_DIR}/checks/smoke-graders.mjs +``` + +All assertions must pass before commit. + +## Constraints + +- DO NOT modify other test cases or graders +- DO NOT modify `_grader-utils.mjs` unless you're adding a new helper that doesn't exist (rare) +- DO NOT modify `suite.yml` — that's the parent skill's job after all writer subagents return +- DO NOT run `run-suite` — that's the run-bench skill's job + +## Commit + +```bash +git add ${WORKBENCH_DIR}/workspace/. \ + ${WORKBENCH_DIR}/checks/grade-${CASE_NAME}-findings.mjs \ + ${WORKBENCH_DIR}/checks/smoke-graders.mjs +git commit -m "feat(workbench): add ${CASE_NAME} test case + grader" +``` + +## Return + +Under 200 words. Include: filenames created, smoke-check result, any blockers. +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(CASE_SPEC|FUNCTIONALITY_REPORT_EXCERPT|WORKBENCH_DIR|CASE_NAME)\}' skills/subagents/test-writer.md | wc -l +# Expected: ≥ 8 + +git add skills/subagents/test-writer.md +git commit -m "feat(v1.4-subagents): test-writer prompt template" +``` + +--- + +### Task C4: `subagents/analyzer.md` + +**Files:** Create `skills/subagents/analyzer.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/analyzer.md` with this exact content: + +````markdown +# Sub-subagent: analyze bench results, identify structural weaknesses + +You are dispatched to analyze the eval-suite results for one skill and produce a structured weakness report. Your output is the LOAD-BEARING input to the improve-skill step — if you can't articulate a structural weakness, no improvement will be attempted. + +## Inputs (templated) + +- `${BENCH_RESULTS_DIR}` — typically `docs/skill-optimizer//05-bench-results//` +- `${WORKBENCH_DIR}` — `docs/skill-optimizer//workbench/` +- `${SKILL_SOURCE_DIR}` — `docs/skill-optimizer//vendored-skill/` (or local skill path) +- `${OUTPUT_PATH}` — typically `docs/skill-optimizer//06-analysis.md` +- `${RECIPES_PATH}` — `skills/references/recipes.md` + +## Tools allowed + +- Read (bench results, suite-result.json, per-trial findings.txt, workbench cases, source skill, recipes) +- Glob (find trial dirs) +- Write (the report) + +## Tools NOT allowed (limited context constraint) + +You see per-trial findings + skill content + workbench case DEFINITIONS, but you do NOT see: + +- The TEST INPUTS themselves (e.g., the seeded SQL/TSX/JSON files in workspace/) — this forces you to think about the SKILL's gaps, not the SOLUTIONS to specific tests +- The grader's internal matching logic (only see its public output) +- Any existing improvement proposal + +## What to produce + +Analyze the bench results and write `${OUTPUT_PATH}` per this STRUCTURED FORMAT (Option A from the spec): + +```markdown +--- +suite_result_path: ${BENCH_RESULTS_DIR}/suite-result.json +analyzer_verdict: +weakness_count: +--- + +# Analysis: + +## Structural weaknesses identified + +### Weakness 1: + +- **Pattern**: Across / trials, models systematically failed to detect . Specifically: . +- **Hypothesized cause**: +- **Connects to skill section**: at `` +- **What WOULD address this** (general principle): +- **What WOULD NOT address this** (anti-ducktape gate): + +### Weakness 2: ... + +## Non-structural noise (ignored — not addressable) + +- +``` + +If you cannot articulate ≥1 structural weakness with confidence, write the report with `analyzer_verdict: no-weakness` and document why. The improve-skill step will refuse to fire — this is the honest behavior. + +## Method + +1. Read `suite-result.json` — get per-case-per-model pass rates +2. For each FAILED trial, read its `findings.txt` (or equivalent output) +3. Cluster failures by: same rule-ID across trials, same model across rules, same case across models +4. Distinguish: + - **SYSTEMATIC** (≥ 2 trials in the same cluster): worth a weakness entry + - **FLAKY** (1 trial, no pattern): non-structural noise +5. For each systematic cluster, locate the responsible skill section, hypothesize cause, and define the anti-ducktape gate +6. Cross-reference `${RECIPES_PATH}` for known patterns (Recipe A–E, grader patterns G1–G6) — if your weakness matches a known pattern, name it + +## Commit + +```bash +git add ${OUTPUT_PATH} +git commit -m "docs(skill-optimizer): analysis report for " +``` + +## Return + +Under 300 words. Include: verdict (weakness-found vs no-weakness), weakness names, brief reasoning, and a pointer to the report. +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(BENCH_RESULTS_DIR|WORKBENCH_DIR|SKILL_SOURCE_DIR|OUTPUT_PATH|RECIPES_PATH)\}' skills/subagents/analyzer.md | wc -l +# Expected: ≥ 8 + +git add skills/subagents/analyzer.md +git commit -m "feat(v1.4-subagents): analyzer prompt template" +``` + +--- + +### Task C5: `subagents/optimizer.md` + +**Files:** Create `skills/subagents/optimizer.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/optimizer.md` with this exact content: + +````markdown +# Sub-subagent: propose a principled fix to address a structural weakness + +You are dispatched to write ONE additive, principled improvement to a skill that addresses a named structural weakness from the analysis report. You MUST NOT pattern-match-patch. + +## Inputs (templated) + +- `${ANALYSIS_PATH}` — `docs/skill-optimizer//06-analysis.md` +- `${FUNCTIONALITY_PATH}` — `docs/skill-optimizer//01-functionality.md` +- `${SKILL_SOURCE_DIR}` — vendored skill files (or local path) +- `${SUBMISSIONS_PATH}` — `docs/skill-optimizer//03-submissions.md` if exists; empty/null otherwise +- `${OUTPUT_PROPOSAL_PATH}` — typically `docs/skill-optimizer//07-improvement-proposal.md` +- `${RECIPES_PATH}` — `skills/references/recipes.md` +- `${ITERATION}` — `1` or `2` (validator may request revision once) + +## Tools allowed + +- Read (analysis, functionality, source skill, submissions context, recipes) +- Write (the proposed diff + rationale) +- Edit (the target skill file, to apply the proposed change) + +## Tools NOT allowed (limited context constraint — CRITICAL) + +You do NOT have access to: + +- Raw failed trial outputs (`findings.txt`) +- The grader's matching logic +- The test inputs (workspace/ files) +- Any pattern that would let you "find the specific token the grader looks for and add it to the skill" + +If you find yourself wanting one of these, STOP and ask the calling skill to escalate. You are not allowed to pattern-match-patch. + +## What to produce + +1. Read `${ANALYSIS_PATH}` — identify the named structural weaknesses + their "What WOULD address" and "What WOULD NOT address" gates. +2. Read `${RECIPES_PATH}` — match to a known recipe (Recipe A-E) if applicable. +3. Read `${SKILL_SOURCE_DIR}` — find the section the weakness "Connects to". +4. Design ONE additive change (max ~50 lines added) that addresses the named weakness using a general principle. The change must: + - Reference a named weakness from the analysis (cite by name in your rationale) + - Use the "What WOULD address this" guidance, NOT the "WOULD NOT" anti-patterns + - Be ADDITIVE only — no deletions, no reworded existing rules + - Match the source skill's existing voice/style +5. If `${SUBMISSIONS_PATH}` exists, ensure the change conforms to the upstream's hard constraints (file location, frontmatter, prefix taxonomy, etc.). +6. Apply the change to the source skill file (Edit tool). +7. Write `${OUTPUT_PROPOSAL_PATH}` with this structure: + +```markdown +--- +iteration: ${ITERATION} +weakness_addressed: +recipe: +--- + +# Improvement proposal (iteration ${ITERATION}) + +## Weakness being addressed + + + +## Proposed change + + + +## Rationale + + + +## Conformance to upstream conventions (if applicable) + + +``` + +## Commit + +```bash +git add ${OUTPUT_PROPOSAL_PATH} ${SKILL_SOURCE_DIR}/...modified file... +git commit -m "feat(skill): iter ${ITERATION} — : " +``` + +## Return + +Under 300 words. Include: weakness addressed, recipe applied, diff summary (~lines added), pointer to proposal. + +## If you cannot produce a principled fix + +If the named weakness genuinely has no general-principle fix (extremely rare), return: + +``` +Status: cannot-fix-principled +Reason: +Recommendation: +``` +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(ANALYSIS_PATH|FUNCTIONALITY_PATH|SKILL_SOURCE_DIR|SUBMISSIONS_PATH|OUTPUT_PROPOSAL_PATH|RECIPES_PATH|ITERATION)\}' skills/subagents/optimizer.md | wc -l +# Expected: ≥ 10 + +git add skills/subagents/optimizer.md +git commit -m "feat(v1.4-subagents): optimizer prompt template (anti-ducktape gates)" +``` + +--- + +### Task C6: `subagents/validator.md` + +**Files:** Create `skills/subagents/validator.md`. + +- [ ] **Step 1: Write the prompt template** + +Create `skills/subagents/validator.md` with this exact content: + +````markdown +# Sub-subagent: validate a proposed skill improvement + +You are dispatched to independently check whether a proposed skill improvement is principled (internal consistency) and (if applicable) ships-as-PR (external consistency). + +## Inputs (templated) + +- `${SKILL_BEFORE_PATH}` — path to the original skill file (pre-modification) +- `${SKILL_AFTER_PATH}` — path to the modified skill file (post-optimizer) +- `${FUNCTIONALITY_PATH}` — `docs/skill-optimizer//01-functionality.md` +- `${SUBMISSIONS_PATH}` — `docs/skill-optimizer//03-submissions.md` if exists; empty/null otherwise +- `${PROPOSAL_PATH}` — `docs/skill-optimizer//07-improvement-proposal.md` (optimizer's rationale) +- `${OUTPUT_VERDICT_PATH}` — typically `docs/skill-optimizer//07-validator-verdict.md` + +## Tools allowed + +- Read (all input files) +- Write (the verdict) +- Bash (diff the before/after) + +## Tools NOT allowed (limited context constraint — CRITICAL) + +You do NOT see: + +- Trial data, findings.txt, raw eval results +- The optimizer's internal reasoning trace +- Test inputs (workspace/ files) + +You are an INDEPENDENT check. The optimizer might have rationalized a bad change; your job is to catch it. + +## What to check + +### Internal consistency (always runs) + +Read `${SKILL_BEFORE_PATH}` and `${SKILL_AFTER_PATH}`. Compute the diff. + +Verify: + +1. **Additive only**: no deletions, no rewordings of existing content +2. **General not specific**: the change embodies a principle, not a token-specific patch (red flag: change adds an exact keyword/string the optimizer might have seen in failed trials) +3. **On-topic**: the change addresses something the skill's functionality (per `${FUNCTIONALITY_PATH}`) actually claims to do +4. **Style match**: matches the source skill's existing voice (terse imperative? prose? bullet list?) +5. **Rationale matches change**: the optimizer's stated rationale in `${PROPOSAL_PATH}` accurately describes what the diff actually does + +### External consistency (only if `${SUBMISSIONS_PATH}` exists) + +Read `${SUBMISSIONS_PATH}`. Verify the change conforms to: + +1. File location (correct path; e.g., new reference file under `references/-.md`) +2. Frontmatter (all required fields, allowed enum values) +3. Prefix taxonomy (no new prefixes; uses existing set) +4. Additive-only at the repo level (no modifications to forbidden files like `_sections.md`, `release-please-config.json`, etc.) +5. Body shape (e.g., `**Incorrect**`/`**Correct**` SQL blocks if required; `Reference:` link if required) +6. Length (within target range from the convention) + +## Output verdict + +Write `${OUTPUT_VERDICT_PATH}` with this structure: + +```markdown +--- +verdict: +internal_consistency: +external_consistency: +--- + +# Validator verdict + +## Internal consistency check + + + +## External consistency check (if applicable) + + + +## Issues found (if any) + + + +## Overall verdict + +approve: change is principled and shippable +needs-revision: ≥1 fixable issue; optimizer should revise once +reject: fundamentally cannot be salvaged (e.g., the diff is destructive, or the change is unrelated to the named weakness) +``` + +## Commit + +```bash +git add ${OUTPUT_VERDICT_PATH} +git commit -m "docs(skill-optimizer): validator verdict for iteration X" +``` + +## Return + +Under 300 words. Include: verdict, the most important issue (if any), and your reasoning for approve/needs-revision/reject. +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(SKILL_BEFORE_PATH|SKILL_AFTER_PATH|FUNCTIONALITY_PATH|SUBMISSIONS_PATH|PROPOSAL_PATH|OUTPUT_VERDICT_PATH)\}' skills/subagents/validator.md | wc -l +# Expected: ≥ 10 + +git add skills/subagents/validator.md +git commit -m "feat(v1.4-subagents): validator prompt template" +``` + +--- + +## Phase D — References (1 task; recipes.md already seeded in A4) + +### Task D1: `references/workflow.md` — human-readable chain diagram + +**Files:** Create `skills/references/workflow.md`. + +- [ ] **Step 1: Write the workflow doc** + +Create `skills/references/workflow.md` with this exact content: + +````markdown +# skill-optimizer v1.4 workflow + +This document describes how the 7 skills chain together to optimize one skill end-to-end. Operators read this to understand the flow; the skills themselves embed similar instructions in-prose for the agent. + +## The chain + +```text +┌────────────────────────────────────┐ +│ skill-optimizer- │ +│ investigate-functionality (1) │ ← user provides skill source +│ │ (URL or local path) +│ │ Asks: "PR submission intent?" +│ │ if upstream +└────────────────┬───────────────────┘ + │ writes 01-functionality.md + ▼ +┌────────────────────────────────────┐ +│ skill-optimizer- │ +│ investigate-test-case (2) │ ← reads 01 +│ │ USER GATE: pick test subset +└────────────────┬───────────────────┘ + │ writes 02-test-case.md + │ + ┌────────┴────────┐ + │ │ + ▼ (pr=true) ▼ (pr=false) +┌──────────────────┐ ┌───────────────────────┐ +│ investigate- │ │ (skip step 3) │ +│ submissions(3) │ │ │ +└──────┬───────────┘ └───────────┬───────────┘ + │ 03-submissions.md │ + └──────────┬────────────────┘ + ▼ +┌────────────────────────────────────┐ +│ skill-optimizer-write-tests (4) │ +│ Parallel test-writer subagents │ +│ (one per picked case, limited ctx) │ +└────────────────┬───────────────────┘ + │ writes workbench/ + 04-tests-plan.md + ▼ +┌────────────────────────────────────┐ +│ skill-optimizer-run-bench (5) │ +│ Invokes skill-optimizer CLI │ +└────────────────┬───────────────────┘ + │ writes 05-bench-results// + ▼ +┌────────────────────────────────────┐ +│ skill-optimizer-analyze-result (6) │ +│ Analyzer subagent (limited ctx — │ +│ no test inputs) │ +└────────────────┬───────────────────┘ + │ writes 06-analysis.md + │ + ┌────────┴────────┐ + │ │ + ▼ (weakness) ▼ (no weakness) +┌──────────────────┐ ┌───────────────────────┐ +│ improve-skill(7) │ │ Exit honestly │ +│ optimizer ↔ │ │ "no weakness; no │ +│ validator loop │ │ improvement warranted"│ +└──────┬───────────┘ └───────────────────────┘ + │ writes 07-improvement-proposal.md + 07-validator-verdict.md + │ writes modified skill file + │ + ├─ local skill → done + ├─ upstream + pr=false → done (vendored copy updated) + └─ upstream + pr=true → package PR draft using 03 +``` + +## State file layout + +All artifacts at convention path: `docs/skill-optimizer//` + +| File | Producer | Consumer(s) | +|---|---|---| +| `vendored-skill/` | (1) | (4) workbench refs, (6), (7) | +| `01-functionality.md` | (1) | (2), (4), (6), (7), validator | +| `02-test-case.md` | (2) | (4) | +| `03-submissions.md` | (3) (optional) | (7) validator (external consistency), packaging | +| `04-tests-plan.md` | (4) | human review | +| `workbench/` | (4) | (5), (6) | +| `05-bench-results//` | (5) | (6) | +| `06-analysis.md` | (6) | (7) — REQUIRED, refuses without | +| `07-improvement-proposal.md` | (7) optimizer | validator | +| `07-validator-verdict.md` | (7) validator | (7) optimizer (revision loop), operator | + +## Limited-context subagents + +The architectural fix for ducttape. See `docs/skill-optimizer-v1.4-spec.md` "Subagent constraints" for the full table. Quick summary: + +- **Optimizer & validator** never see raw trial data or grader internals — forces principled improvement +- **Test writer** never sees the skill content — prevents grader-hacking +- **Analyzer** never sees the test inputs — forces focus on the skill, not the solutions +- **Researchers** never see proposed changes — pure investigation + +## Auto-pilot vs manual + +- **Manual:** user invokes one skill at a time, reviews each report, then invokes the next +- **Auto-pilot:** user says "optimize skill X end-to-end" — the agent invokes 1 → 2 → (3 if pr) → 4 → 5 → 6 → 7 in order, pausing at the natural user-review gates (after 2 for test-pick, after 7 for proposal review) + +No separate auto-pilot skill — it's just chained `Skill` tool invocations driven by each skill's "Next, invoke X" handoff. +```` + +- [ ] **Step 2: Verify** + +```bash +wc -l skills/references/workflow.md +grep -c "^## " skills/references/workflow.md +``` + +Expected: ≥ 80 lines, ≥ 4 H2 sections. + +- [ ] **Step 3: Commit** + +```bash +git add skills/references/workflow.md +git commit -m "feat(v1.4-references): add workflow.md (chain diagram + state file map)" +``` + +--- + +## Phase E — End-to-end validation (2 tasks, requires real eval runs) + +### Task E1: End-to-end test on a LOCAL skill + +**Goal:** Validate the 7-skill chain end-to-end on a local skill. Per acceptance criteria #5, this proves the workflow produces all expected reports, modifies the target skill, and the validator approves. + +Choose a small local skill from `skills/` (e.g., `skills/skill-optimizer/SKILL.md` itself, or create a tiny placeholder skill in this repo). Optimize it for local use (no PR submission). + +- [ ] **Step 1: Pick a target local skill** + +```bash +ls skills/ +``` + +Recommend: create a fresh tiny `skills/v14-validation-target/SKILL.md` for the test, so we don't disrupt real skills. Body should have a deliberate weakness (e.g., a rule stated declaratively that an agent would routinely miss). + +```bash +mkdir -p skills/v14-validation-target +cat > skills/v14-validation-target/SKILL.md <<'EOF' +--- +name: v14-validation-target +description: Test target skill for v1.4 validation. Reviews YAML configs for naming-convention compliance. +--- + +# v14-validation-target + +Review YAML configuration files for compliance with naming conventions. + +## Rules + +- Keys should use snake_case +- Boolean values should be lowercase `true`/`false` +- Lists should not be empty + +EOF + +git add skills/v14-validation-target/SKILL.md +git commit -m "test(v1.4): add deliberate-weakness target skill for E2E validation" +``` + +- [ ] **Step 2: Invoke the chain manually** + +In your Claude Code session, walk through: + +```text +1. /skill skill-optimizer-investigate-functionality (point at skills/v14-validation-target/SKILL.md; answer "no" to PR question since it's local) +2. /skill skill-optimizer-investigate-test-case (pick 2-3 cases from the proposal) +3. /skill skill-optimizer-write-tests (let the parallel writers build the workbench) +4. /skill skill-optimizer-run-bench (run the baseline; may take 5-15 min depending on model matrix) +5. /skill skill-optimizer-analyze-result (read the analysis; verify it identifies the absence-of-procedural-instruction weakness OR honestly says no weakness) +6. If weakness identified: /skill skill-optimizer-improve-skill (verify optimizer+validator approve a principled change) +7. If no weakness: verify the chain exits honestly — no fabricated proposal +``` + +- [ ] **Step 3: Verify all expected reports exist** + +```bash +SLUG=local-v14-validation-target +test -f docs/skill-optimizer/$SLUG/01-functionality.md +test -f docs/skill-optimizer/$SLUG/02-test-case.md +test -d docs/skill-optimizer/$SLUG/workbench +test -d docs/skill-optimizer/$SLUG/05-bench-results +test -f docs/skill-optimizer/$SLUG/06-analysis.md +# 07-* files only if weakness was identified +echo "all expected reports present" +``` + +- [ ] **Step 4: Write validation note + commit** + +```bash +cat > docs/skill-optimizer-v1.4-validation.md <. +- Phase 2 (test-case): picked cases out of proposed. +- Phase 4 (write-tests): parallel writers; smoke checks passed. +- Phase 5 (run-bench): baseline . +- Phase 6 (analyze-result): . . +- Phase 7 (improve-skill): . . + +## Verdict + +v1.4 chain works end-to-end on a local skill. . +EOF + +git add docs/skill-optimizer-v1.4-validation.md +git commit -m "test(v1.4): document local-skill E2E validation result" +``` + +--- + +### Task E2: End-to-end test re-running firecrawl (the v1.3 regression case) + +**Goal:** Per acceptance criteria #6, re-run the firecrawl skill (which v1.3 regressed) under v1.4. Expected: optimizer produces a principled fix OR honestly refuses — no regression shipped. + +- [ ] **Step 1: Run the chain on firecrawl as an upstream skill** + +In your Claude Code session: + +```text +1. /skill skill-optimizer-investigate-functionality (point at firecrawl/skills/firecrawl-build-scrape; answer "yes" to PR question — we want the full upstream path) +2. /skill skill-optimizer-investigate-test-case +3. /skill skill-optimizer-investigate-submissions (auto-invoked since pr=true) +4. /skill skill-optimizer-write-tests +5. /skill skill-optimizer-run-bench +6. /skill skill-optimizer-analyze-result +7. /skill skill-optimizer-improve-skill +``` + +- [ ] **Step 2: Verify NO regression on the original case** + +The v1.3 regression was specifically: iteration 1's Recipe A+D combination caused the original `review-scrape-integration` case to regress from 1.00 → 0.44 on gpt-5/gemini. Under v1.4, EITHER: + +- The optimizer produces a single principled change (per validator approval), the post-iteration run-bench shows the original case still at 1.00 AND the new harder cases improve, OR +- The analyzer reports "no clear structural weakness" and improve-skill refuses to fire honestly + +Run an extra `run-bench` AFTER improve-skill completes to verify the original case isn't regressed. (The improve-skill skill's own validator should already check this via the optimizer's output, but a confirming run is reasonable.) + +- [ ] **Step 3: Append to validation note** + +```bash +cat >> docs/skill-optimizer-v1.4-validation.md < +Target: firecrawl/skills/firecrawl-build-scrape (upstream, pr=true) + +## Result +- Phase 6 verdict: +- Phase 7 outcome: +- Original case (review-scrape-integration) post-iteration: (baseline: 1.00) +- NO regression observed: + +## v1.3 comparison + +Under v1.3, firecrawl iteration 1 regressed the original case from 1.00 → 0.44 by piling on Recipe A+D simultaneously. The orchestrator ran out of context before catching the regression. v1.4's validator + limited-context optimizer either prevented the regression OR honestly refused to ship the change. + +## Verdict + +v1.4 successfully addresses the canonical "ducktape-by-monolithic-orchestrator" case study from v1.3. . +EOF + +git add docs/skill-optimizer-v1.4-validation.md +git commit -m "test(v1.4): document firecrawl re-run — v1.3 regression case study" +``` + +--- + +## Self-Review + +After writing the plan, here's the spec-coverage check: + +**Spec section → plan tasks:** + +- §"Architecture overview" → Task A1 (skill dirs), A2 (subagents/references), D1 (workflow.md) +- §"Subagent constraints" → Tasks C1-C6 each encode their limited-context constraints +- §"The 7 skills" — interface contracts → Tasks B1-B7 (each task pastes the relevant spec excerpt as the writing-skills brief) +- §"Auto-pilot mode" → covered by D1 (workflow.md) + each skill's in-prose handoff (created in Phase B) +- §"Plugin packaging" → no new task needed; existing `.claude-plugin/plugin.json` auto-discovers the new skills via `skills/` convention +- §"Coexistence with v1.3" → Task A3 (deprecation banner) +- §"Acceptance criteria" #1-7 → Tasks A1, C1-C6 (#2 subagent constraints), D1 (#3 workflow), A4 (#4 recipes), E1 (#5 local E2E), E2 (#6 firecrawl E2E) +- §"Out of scope" — Codex/Cursor ports + programmatic auto-pilot wrapper + per-skill prose enrichment — NOT in plan, deliberately deferred + +**Placeholder scan:** None — every task has either complete content (Phases A, C, D), a complete spec-excerpt brief (Phase B), or concrete CLI commands (Phase E). + +**Type consistency:** + +- Skill names consistent: all 7 follow `skill-optimizer-` pattern across spec, file paths, brief excerpts +- State file paths consistent: `docs/skill-optimizer//-.md` (or `workbench/`, `05-bench-results//`, `vendored-skill/`) throughout +- Subagent file paths consistent: `skills/subagents/.md` throughout +- Template variable names (`${SLUG}`, `${WORKBENCH_DIR}`, etc.) match between subagent templates and skills' planned invocation patterns + +No issues found. + +--- + +## Execution Handoff + +Plan complete and saved to `docs/skill-optimizer-v1.4-plan.md` (in worktree `.claude/worktrees/v1.4-spec/`, branch `feat/skill-optimizer-v1.4`). + +**Two execution options:** + +1. **Subagent-Driven (recommended for Phases A, C, D, E)** — use superpowers:subagent-driven-development. Fresh subagent per task + two-stage review. Works cleanly for mechanical tasks. + + **IMPORTANT:** Phase B (7 SKILL.md creation tasks) is **NOT subagent-driven**. Per the user's explicit request, each Task B is performed interactively in the operator's CC session using `superpowers:writing-skills`. Skip Phase B during subagent-driven execution; do them interactively after. + +2. **Inline Execution** — superpowers:executing-plans, batch with checkpoints. Phase B still must be done interactively even in this mode. + +**Recommended workflow:** + +1. Subagent-driven execution of Phase A (4 tasks, ~30 min) +2. Subagent-driven execution of Phase C + D (7 tasks, ~45 min) +3. Operator interactively runs Phase B (7 tasks, ~2-3 hours since each uses writing-skills) +4. Operator runs Phase E (2 tasks, ~30-60 min each — real eval runs cost money + time) + +Which approach? From 29f72a25ca861d24f901ae6debf0c929569d95e4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 07:54:30 -0500 Subject: [PATCH 004/121] docs(v1.4-plan): mark Phases B/C/D interactive, assign skill-creator + writing-skills MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phases B–D produce v1.4's load-bearing artifacts (description-routed SKILL.md files, subagent prompt templates that hold the anti-ducktape constraints, and the workflow reference doc). Per operator direction, none of these should be subagent-driven — "writing good skills is HARD" and "everything to be precise". Tool assignment: - Phase B (7 SKILL.md): skill-creator as outer loop (description-routing + eval iteration), superpowers:writing-skills as inner loop when behavioral compliance issues surface during eval - Phase C (6 subagent prompts): superpowers:writing-skills only — these are compliance documents and the limited-context constraint must hold under adversarial pressure, pure TDD-pressure-scenario territory - Phase D (workflow.md): no skill tool; direct interactive authoring with section-by-section operator review (matches handoffs from B + constraints from C, which may have shifted during interactive iteration) Only Phase A (file scaffolding) and Phase E (real eval runs) sit outside the interactive flow. --- docs/skill-optimizer-v1.4-plan.md | 226 +++++++++++++++++++++++------- 1 file changed, 176 insertions(+), 50 deletions(-) diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/skill-optimizer-v1.4-plan.md index 06c0af7..7eaecf4 100644 --- a/docs/skill-optimizer-v1.4-plan.md +++ b/docs/skill-optimizer-v1.4-plan.md @@ -1,12 +1,40 @@ # skill-optimizer v1.4 Implementation Plan -> **For agentic workers:** This plan has TWO execution modes: +> **For agentic workers:** This plan has TWO execution modes, by phase: > -> - **Phase A, C, D, E** — use superpowers:subagent-driven-development (mechanical; standard subagent-per-task with two-stage review). -> - **Phase B (per-skill SKILL.md creation)** — use superpowers:writing-skills INTERACTIVELY with the operator. Do NOT batch via subagent. Each B-task is "invoke writing-skills to create one SKILL.md, with the interface contract from the spec as the brief." +> | Phase | Mode | Tool | +> |---|---|---| +> | **A** (structural setup) | subagent-driven-development OK | mechanical Bash/Edit | +> | **B** (the 7 SKILL.md files) | INTERACTIVE with operator | **`skill-creator`** (outer loop: draft, eval, iterate, description-improver) + **`superpowers:writing-skills`** (inner loop: TDD pressure scenarios when a compliance issue surfaces) | +> | **C** (subagent prompt templates) | INTERACTIVE with operator | **`superpowers:writing-skills`** (TDD pressure scenarios for limited-context constraints — the prompts ARE compliance documents) | +> | **D** (`workflow.md`) | INTERACTIVE with operator | No skill tool — direct authoring + user review per section | +> | **E** (end-to-end validation) | manual operator runs (real eval $) | direct `Skill` invocations + `run-suite` | +> +> **Phases B, C, D must NOT be subagent-dispatched batch.** Writing good skills + prompts is hard, and the user explicitly asked for precision across all of them. Each task is human-in-loop with iterative refinement. > > Steps use checkbox (`- [ ]`) syntax for tracking. +## Why two skill-authoring tools? + +We use **both** `skill-creator` and `superpowers:writing-skills` in v1.4 — they're complementary, not alternatives: + +- **`skill-creator`** (Anthropic plugin, at `~/.claude/plugins/cache/claude-plugins-official/skill-creator/`) + - Author + eval + iterate loop: capture intent → draft → run test prompts → variance analysis → refine + - Ships an `eval-viewer/generate_review.py` script + - Ships a **description-improver** script for optimizing trigger accuracy + - Best for: routed-by-description skills where triggering matters (i.e., our 7 SKILL.md files in Phase B) + +- **`superpowers:writing-skills`** (superpowers plugin) + - TDD discipline applied to documentation: write pressure scenarios with subagents → watch agent fail (RED) → write skill → watch agent pass (GREEN) → refactor to close loopholes + - Required reading: superpowers:test-driven-development + - Best for: compliance documents where the rule HAS to be obeyed under pressure (i.e., subagent prompt templates in Phase C, OR an inner-loop fix when a Phase B SKILL.md keeps letting the agent violate its constraints) + +**The split in v1.4:** + +- Phase B (SKILL.md files): `skill-creator` is primary (description-routing + eval iteration); `writing-skills` is the inner-loop fix when a behavioral compliance issue surfaces during eval (e.g., "the optimizer skill keeps dispatching subagents without limited-context constraints — let me write a pressure scenario to nail this down"). +- Phase C (subagent prompts): `writing-skills` only. These aren't description-routed (skills invoke them explicitly), so `skill-creator`'s description-optimizer doesn't apply. The hard part is preventing the subagent from cheating its limited-context constraints — pure TDD pressure-scenario territory. +- Phase D (workflow.md): Just a reference doc, no skill tool needed. + **Goal:** Implement v1.4 — replace the v1.3 monolithic `auto-improve-orchestrator` with 7 independent Claude Code skills that chain via the `superpowers` plugin pattern. Each skill produces a human-reviewable report at a convention path; subagents dispatched by each skill run under strict limited context to prevent ducttape patches. **Architecture:** Skills (not slash commands), description-routed, output to `docs/skill-optimizer//-.md`. The skill-optimizer engine (`run-suite`, graders, Docker harness) stays unchanged. The v1.3 `auto-improve-orchestrator/` skill is deprecated-but-retained for backward compatibility. @@ -315,13 +343,16 @@ v1.3-specific wrapper references neutralized." --- -## Phase B — Interactive SKILL.md creation (7 tasks, USE writing-skills, NOT subagent-driven) +## Phase B — Interactive SKILL.md creation (7 tasks, INTERACTIVE via skill-creator + writing-skills) -**Mode:** Each Task B uses `superpowers:writing-skills` interactively. The operator runs the skill, answers writing-skills' clarifying questions, iterates on the output, commits when satisfied. +**Mode:** Each Task B is performed interactively in the operator's CC session. -**Do NOT dispatch these as autonomous subagents** — writing-skills is built for human-in-loop refinement, and the user explicitly stated "writing good skills is HARD." +- **Primary tool: `skill-creator`** (Anthropic plugin) — outer loop. Capture intent → draft SKILL.md → generate eval test prompts → run variance analysis → iterate on description and body until the skill triggers reliably and behaves correctly. +- **Inner-loop tool: `superpowers:writing-skills`** — invoked WHEN a behavioral compliance issue surfaces during skill-creator's eval iteration (e.g., "the optimizer skill keeps dispatching subagents without limited-context constraints" → write a TDD pressure scenario, watch the subagent fail without the rule, add the rule, watch it pass). Not every Task B needs writing-skills; reach for it when a constraint must hold under adversarial pressure. -For EACH Task B, the input to writing-skills is the interface contract for that skill from `docs/skill-optimizer-v1.4-spec.md` section "The 7 skills". Paste the relevant section as the writing-skills brief, plus the constraints from "Subagent constraints" if the skill dispatches a subagent. +**Do NOT dispatch these as autonomous subagents** — both tools are built for human-in-loop refinement, and the user explicitly stated "writing good skills is HARD." + +For EACH Task B, the brief for `skill-creator` is the interface contract for that skill from `docs/skill-optimizer-v1.4-spec.md` section "The 7 skills", plus the constraints from "Subagent constraints" if the skill dispatches a subagent. Paste those excerpts verbatim into the `skill-creator` capture step. --- @@ -331,15 +362,17 @@ For EACH Task B, the input to writing-skills is the interface contract for that - Modify: `skills/skill-optimizer-investigate-functionality/SKILL.md` (replace the placeholder body) -- [ ] **Step 1: Invoke writing-skills with the spec excerpt as the brief** +- [ ] **Step 1: Invoke skill-creator with the spec excerpt as the brief** In your Claude Code session, invoke: ```text -/skill superpowers:writing-skills +/skill skill-creator ``` -Brief: +Use skill-creator's capture → draft → eval → iterate loop. If during eval iteration a behavioral compliance issue surfaces (e.g., the skill repeatedly skips a required step under pressure), pause and invoke `superpowers:writing-skills` to author a TDD pressure scenario for that specific failure, then resume skill-creator. + +Brief (paste into skill-creator's capture step): ```text Create SKILL.md body for skills/skill-optimizer-investigate-functionality/. @@ -364,7 +397,7 @@ The frontmatter (name + description) is already in place from Task A1 — keep i Examples to draw from for shape: the existing superpowers/skills/brainstorming/SKILL.md for "skill that asks user questions, persists state, hands off to next skill" pattern. ``` -Iterate with writing-skills until the SKILL.md body is satisfactory. Verify the frontmatter is unchanged and the body matches the spec contract. +Iterate with skill-creator (and writing-skills as needed) until the SKILL.md body is satisfactory and skill-creator's eval prompts trigger the skill reliably. Verify the frontmatter is unchanged and the body matches the spec contract. - [ ] **Step 2: Verify the file** @@ -384,7 +417,7 @@ Expected: `name: skill-optimizer-investigate-functionality`, description still p ```bash git add skills/skill-optimizer-investigate-functionality/SKILL.md -git commit -m "feat(skill-optimizer-investigate-functionality): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-investigate-functionality): SKILL.md body via skill-creator" ``` --- @@ -424,7 +457,7 @@ Frontmatter is already in place from Task A1. ```bash node -e "const m = require('gray-matter')(require('fs').readFileSync('skills/skill-optimizer-investigate-test-case/SKILL.md', 'utf-8')); console.log('lines:', m.content.split('\\n').length);" git add skills/skill-optimizer-investigate-test-case/SKILL.md -git commit -m "feat(skill-optimizer-investigate-test-case): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-investigate-test-case): SKILL.md body via skill-creator" ``` --- @@ -435,12 +468,14 @@ git commit -m "feat(skill-optimizer-investigate-test-case): SKILL.md body via wr - Modify: `skills/skill-optimizer-investigate-submissions/SKILL.md` -- [ ] **Step 1: Invoke writing-skills** +- [ ] **Step 1: Invoke skill-creator** ```text -/skill superpowers:writing-skills +/skill skill-creator ``` +(See Task B1 Step 1 for the skill-creator + writing-skills workflow.) + Brief: ```text @@ -463,7 +498,7 @@ Reference example: the existing tools/auto-improve-contexts/*.md files (now move ```bash git add skills/skill-optimizer-investigate-submissions/SKILL.md -git commit -m "feat(skill-optimizer-investigate-submissions): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-investigate-submissions): SKILL.md body via skill-creator" ``` --- @@ -474,12 +509,14 @@ git commit -m "feat(skill-optimizer-investigate-submissions): SKILL.md body via - Modify: `skills/skill-optimizer-write-tests/SKILL.md` -- [ ] **Step 1: Invoke writing-skills** +- [ ] **Step 1: Invoke skill-creator** ```text -/skill superpowers:writing-skills +/skill skill-creator ``` +(See Task B1 Step 1 for the skill-creator + writing-skills workflow.) + Brief: ```text @@ -504,7 +541,7 @@ Interface contract (from spec §4): ```bash git add skills/skill-optimizer-write-tests/SKILL.md -git commit -m "feat(skill-optimizer-write-tests): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-write-tests): SKILL.md body via skill-creator" ``` --- @@ -515,12 +552,14 @@ git commit -m "feat(skill-optimizer-write-tests): SKILL.md body via writing-skil - Modify: `skills/skill-optimizer-run-bench/SKILL.md` -- [ ] **Step 1: Invoke writing-skills** +- [ ] **Step 1: Invoke skill-creator** ```text -/skill superpowers:writing-skills +/skill skill-creator ``` +(See Task B1 Step 1 for the skill-creator + writing-skills workflow.) + Brief: ```text @@ -543,7 +582,7 @@ Reference: the existing src/cli.ts run-suite command + the v1.3 orchestrator's P ```bash git add skills/skill-optimizer-run-bench/SKILL.md -git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via skill-creator" ``` --- @@ -554,12 +593,14 @@ git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via writing-skills - Modify: `skills/skill-optimizer-analyze-result/SKILL.md` -- [ ] **Step 1: Invoke writing-skills** +- [ ] **Step 1: Invoke skill-creator** ```text -/skill superpowers:writing-skills +/skill skill-creator ``` +(See Task B1 Step 1 for the skill-creator + writing-skills workflow.) + Brief: ```text @@ -602,7 +643,7 @@ Interface contract (from spec §6 — this is the LOAD-BEARING skill that determ ```bash git add skills/skill-optimizer-analyze-result/SKILL.md -git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via skill-creator" ``` --- @@ -613,12 +654,14 @@ git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via writing-s - Modify: `skills/skill-optimizer-improve-skill/SKILL.md` -- [ ] **Step 1: Invoke writing-skills** +- [ ] **Step 1: Invoke skill-creator** ```text -/skill superpowers:writing-skills +/skill skill-creator ``` +(See Task B1 Step 1 for the skill-creator + writing-skills workflow.) + Brief: ```text @@ -650,20 +693,43 @@ Interface contract (from spec §7): ```bash git add skills/skill-optimizer-improve-skill/SKILL.md -git commit -m "feat(skill-optimizer-improve-skill): SKILL.md body via writing-skills" +git commit -m "feat(skill-optimizer-improve-skill): SKILL.md body via skill-creator" ``` --- -## Phase C — Subagent prompt templates (6 tasks, mechanical, subagent-driven-development OK) +## Phase C — Subagent prompt templates (6 tasks, INTERACTIVE via writing-skills) Each subagent prompt template is loaded by the corresponding skill (Phase B), substituted with templated inputs, and dispatched via Agent tool. All must enforce the limited-context constraints from the spec's "Subagent constraints" table. +**Mode:** Each Task C is performed interactively via `superpowers:writing-skills`. The drafts embedded below in each task's "Step 1" are the starting brief — NOT the finished prompt. writing-skills' core discipline is **TDD pressure scenarios**: watch a subagent fail without the rule, add the rule, watch the subagent pass. These subagent prompts ARE compliance documents — the limited-context constraint is the load-bearing anti-ducktape gate, and it MUST hold under adversarial conditions. Pure pressure-scenario territory; pure writing-skills fit. + +**Why not skill-creator here?** skill-creator's value-add is description-routing (when does this skill fire?) and eval-prompt variance analysis. These subagent prompts are not description-routed — the parent skill invokes them explicitly. So skill-creator's outer loop doesn't apply. + +**Do NOT dispatch these as autonomous subagents.** Each Task C needs operator-driven pressure scenarios. + +For EACH Task C: + +1. Paste the draft template from Step 1 into writing-skills as the starting brief. +2. Identify the load-bearing constraint(s) (typically the "Tools NOT allowed" / "What you do NOT see" section). +3. Construct a pressure scenario: a realistic operator request where a non-disciplined subagent would cheat the constraint to "be helpful". Watch a fresh subagent fail with just the draft. +4. Strengthen the prompt's constraints/wording. Re-run the scenario. Iterate until the constraint holds. +5. Commit when satisfied. + ### Task C1: `subagents/research-functionality.md` **Files:** Create `skills/subagents/research-functionality.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/research-functionality.md` with this exact content: @@ -778,7 +844,16 @@ git commit -m "feat(v1.4-subagents): research-functionality prompt template" **Files:** Create `skills/subagents/research-submissions.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/research-submissions.md` with this exact content (adapted from v1.3's `prompts/research-upstream.md`): @@ -884,7 +959,16 @@ git commit -m "feat(v1.4-subagents): research-submissions prompt template" **Files:** Create `skills/subagents/test-writer.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/test-writer.md` with this exact content: @@ -973,7 +1057,16 @@ git commit -m "feat(v1.4-subagents): test-writer prompt template" **Files:** Create `skills/subagents/analyzer.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/analyzer.md` with this exact content: @@ -1075,7 +1168,16 @@ git commit -m "feat(v1.4-subagents): analyzer prompt template" **Files:** Create `skills/subagents/optimizer.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/optimizer.md` with this exact content: @@ -1189,7 +1291,16 @@ git commit -m "feat(v1.4-subagents): optimizer prompt template (anti-ducktape ga **Files:** Create `skills/subagents/validator.md`. -- [ ] **Step 1: Write the prompt template** +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +Invoke `/skill superpowers:writing-skills`. Paste the draft template below as the initial content. Then: + +1. Identify the load-bearing constraint (typically "Tools NOT allowed" / "What you do NOT see"). For most C tasks this is the limited-context rule preventing the subagent from reading prior analyses or failure data. +2. Construct a pressure scenario. Example: operator asks the subagent for "everything you know about the skill, including prior issues" — does the subagent stay within its allowed-context boundary, or does it Read forbidden files to "be helpful"? +3. Watch a fresh subagent run with just the draft. If it cheats the constraint, the draft is insufficient. +4. Strengthen wording. Re-run. Iterate until the constraint holds under pressure. + +The draft template: Create `skills/subagents/validator.md` with this exact content: @@ -1304,15 +1415,25 @@ git commit -m "feat(v1.4-subagents): validator prompt template" --- -## Phase D — References (1 task; recipes.md already seeded in A4) +## Phase D — References (1 task, INTERACTIVE authoring; recipes.md already seeded in A4) + +**Mode:** Direct interactive authoring with operator review per section. No skill tool. `workflow.md` is a human reference doc — not a behavioral skill (so skill-creator's description-routing/eval loop doesn't apply) and not a compliance prompt template (so writing-skills' pressure scenarios don't apply). The right discipline is a careful walkthrough with the operator: present each section (chain diagram, invariants, etc.), get confirmation, iterate, commit. + +**Do NOT dispatch as autonomous subagent.** The operator must validate that the diagram + invariants match the actual Phase B SKILL.md handoffs and Phase C subagent constraints (which were finalized interactively just prior). ### Task D1: `references/workflow.md` — human-readable chain diagram **Files:** Create `skills/references/workflow.md`. -- [ ] **Step 1: Write the workflow doc** +- [ ] **Step 1: Walk through each section with the operator, then write the file** + +Use the draft below as the starting point. Before writing, walk through each section interactively with the operator and confirm it reflects what was actually built in Phases B and C. In particular: -Create `skills/references/workflow.md` with this exact content: +1. Chain diagram — confirm the handoff text matches each SKILL.md's actual "Handoff" prose from Phase B. +2. Subagent constraint table — confirm wording matches the finalized prompts from Phase C (which may have shifted during pressure-scenario iteration). +3. Invariants section — surface any new invariants discovered while writing Phases B/C. + +Then create `skills/references/workflow.md` with the agreed content. Draft below: ````markdown # skill-optimizer v1.4 workflow @@ -1596,8 +1717,8 @@ After writing the plan, here's the spec-coverage check: **Spec section → plan tasks:** - §"Architecture overview" → Task A1 (skill dirs), A2 (subagents/references), D1 (workflow.md) -- §"Subagent constraints" → Tasks C1-C6 each encode their limited-context constraints -- §"The 7 skills" — interface contracts → Tasks B1-B7 (each task pastes the relevant spec excerpt as the writing-skills brief) +- §"Subagent constraints" → Tasks C1-C6 each encode their limited-context constraints, hardened via writing-skills pressure scenarios +- §"The 7 skills" — interface contracts → Tasks B1-B7 (each task pastes the relevant spec excerpt as the skill-creator brief) - §"Auto-pilot mode" → covered by D1 (workflow.md) + each skill's in-prose handoff (created in Phase B) - §"Plugin packaging" → no new task needed; existing `.claude-plugin/plugin.json` auto-discovers the new skills via `skills/` convention - §"Coexistence with v1.3" → Task A3 (deprecation banner) @@ -1621,19 +1742,24 @@ No issues found. Plan complete and saved to `docs/skill-optimizer-v1.4-plan.md` (in worktree `.claude/worktrees/v1.4-spec/`, branch `feat/skill-optimizer-v1.4`). -**Two execution options:** - -1. **Subagent-Driven (recommended for Phases A, C, D, E)** — use superpowers:subagent-driven-development. Fresh subagent per task + two-stage review. Works cleanly for mechanical tasks. +**Phase-by-phase execution mode (see header table for tool assignment):** - **IMPORTANT:** Phase B (7 SKILL.md creation tasks) is **NOT subagent-driven**. Per the user's explicit request, each Task B is performed interactively in the operator's CC session using `superpowers:writing-skills`. Skip Phase B during subagent-driven execution; do them interactively after. +| Phase | Mode | Driver | +|---|---|---| +| **A** (4 tasks, structural setup) | subagent-driven OK | `superpowers:subagent-driven-development` | +| **B** (7 tasks, the SKILL.md files) | INTERACTIVE — DO NOT subagent-drive | operator runs `skill-creator` (+ `writing-skills` inner loop) | +| **C** (6 tasks, subagent prompt templates) | INTERACTIVE — DO NOT subagent-drive | operator runs `superpowers:writing-skills` with pressure scenarios | +| **D** (1 task, `workflow.md`) | INTERACTIVE — DO NOT subagent-drive | operator authors directly with section-by-section review | +| **E** (2 tasks, end-to-end validation) | operator-driven (real eval runs) | direct CLI / skill invocations; costs $$ + time | -2. **Inline Execution** — superpowers:executing-plans, batch with checkpoints. Phase B still must be done interactively even in this mode. +**Why only Phase A is subagent-eligible:** Phase A is pure file-structure scaffolding with no judgment calls. Phases B–D produce the load-bearing artifacts of v1.4 — the SKILL.md descriptions that determine routing, the subagent prompts that hold the anti-ducktape constraints, and the workflow doc that humans use to understand the chain. The user's explicit guidance: "writing good skills is HARD" and "part C and D should also be done interactively. I want everything to be precise here." Subagent-driving Phases B–D would defeat the purpose. -**Recommended workflow:** +**Recommended order:** -1. Subagent-driven execution of Phase A (4 tasks, ~30 min) -2. Subagent-driven execution of Phase C + D (7 tasks, ~45 min) -3. Operator interactively runs Phase B (7 tasks, ~2-3 hours since each uses writing-skills) -4. Operator runs Phase E (2 tasks, ~30-60 min each — real eval runs cost money + time) +1. Subagent-driven Phase A (~30 min) +2. Operator interactively runs Phase B (7 tasks, ~2–3 hours via `skill-creator` + `writing-skills`) +3. Operator interactively runs Phase C (6 tasks, ~1–2 hours via `writing-skills` pressure scenarios) +4. Operator interactively runs Phase D (1 task, ~30 min) +5. Operator runs Phase E (2 tasks, ~30–60 min each — real eval runs) Which approach? From a7dce70ba0e7081a9bcb9af73ff891fdc2e060e4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 07:58:31 -0500 Subject: [PATCH 005/121] feat(v1.4): create 7 skill directory shells with frontmatter Body content for each SKILL.md will be written interactively in Phase B via superpowers:writing-skills, with the interface contract from docs/skill-optimizer-v1.4-spec.md as the brief. This commit just lays down the structural skeleton and discoverable frontmatter. --- skills/skill-optimizer-analyze-result/SKILL.md | 9 +++++++++ skills/skill-optimizer-improve-skill/SKILL.md | 9 +++++++++ .../skill-optimizer-investigate-functionality/SKILL.md | 9 +++++++++ skills/skill-optimizer-investigate-submissions/SKILL.md | 9 +++++++++ skills/skill-optimizer-investigate-test-case/SKILL.md | 9 +++++++++ skills/skill-optimizer-run-bench/SKILL.md | 9 +++++++++ skills/skill-optimizer-write-tests/SKILL.md | 9 +++++++++ 7 files changed, 63 insertions(+) create mode 100644 skills/skill-optimizer-analyze-result/SKILL.md create mode 100644 skills/skill-optimizer-improve-skill/SKILL.md create mode 100644 skills/skill-optimizer-investigate-functionality/SKILL.md create mode 100644 skills/skill-optimizer-investigate-submissions/SKILL.md create mode 100644 skills/skill-optimizer-investigate-test-case/SKILL.md create mode 100644 skills/skill-optimizer-run-bench/SKILL.md create mode 100644 skills/skill-optimizer-write-tests/SKILL.md diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md new file mode 100644 index 0000000..0cfebb2 --- /dev/null +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-analyze-result +description: Use when the user wants to analyze bench results and identify structural weaknesses — clusters failures, distinguishes systematic from flaky, and produces a structured analysis with anti-ducttape gates. +--- + +# skill-optimizer-analyze-result + + + diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md new file mode 100644 index 0000000..bb2c382 --- /dev/null +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-improve-skill +description: Use when the user wants to improve a skill based on identified structural weakness — dispatches optimizer + validator subagents under strict limited context; refuses if no structural weakness in the analysis. +--- + +# skill-optimizer-improve-skill + + + diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md new file mode 100644 index 0000000..445a314 --- /dev/null +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-investigate-functionality +description: Use when the user asks to understand or investigate what a skill does — fetches skill source, web-searches the underlying technology, writes a functionality report. Also asks the user (for upstream skills) whether to target upstream PR submission. +--- + +# skill-optimizer-investigate-functionality + + + diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md new file mode 100644 index 0000000..54b0460 --- /dev/null +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-investigate-submissions +description: Use when the user wants to research a skill's upstream PR conventions — produces a context block with license, CLA, frontmatter spec, file-location rules, PR-shape patterns, and rejection signals. +--- + +# skill-optimizer-investigate-submissions + + + diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md new file mode 100644 index 0000000..e781bb3 --- /dev/null +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-investigate-test-case +description: Use when the user wants to design or propose test cases for a skill — enumerates the skill's responsibilities and proposes ranked test cases the user can pick from. +--- + +# skill-optimizer-investigate-test-case + + + diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md new file mode 100644 index 0000000..03f66df --- /dev/null +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-run-bench +description: Use when the user wants to run the eval suite for a skill and capture results — thin wrapper around the skill-optimizer CLI's run-suite command. +--- + +# skill-optimizer-run-bench + + + diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md new file mode 100644 index 0000000..378d53e --- /dev/null +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -0,0 +1,9 @@ +--- +name: skill-optimizer-write-tests +description: Use when the user wants to build / implement the eval workbench for a skill — plans the workbench, asks for user confirmation, then dispatches parallel narrow-context subagents to write each test case + grader. +--- + +# skill-optimizer-write-tests + + + From 7c0541a91491ceeb61be3b01a5d7c23652ff6544 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 07:58:42 -0500 Subject: [PATCH 006/121] chore(v1.4): create subagents/ and references/ dirs (populated in Phase C+D) --- skills/references/.gitkeep | 0 skills/subagents/.gitkeep | 0 2 files changed, 0 insertions(+), 0 deletions(-) create mode 100644 skills/references/.gitkeep create mode 100644 skills/subagents/.gitkeep diff --git a/skills/references/.gitkeep b/skills/references/.gitkeep new file mode 100644 index 0000000..e69de29 diff --git a/skills/subagents/.gitkeep b/skills/subagents/.gitkeep new file mode 100644 index 0000000..e69de29 From 2774984793864d021ca7d8176142f8fe0558781d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 08:07:17 -0500 Subject: [PATCH 007/121] feat(v1.4): seed references/recipes.md from v1.3 lessons.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Pulled from feat/auto-improve-skill-v1.3:skills/auto-improve-orchestrator/ references/lessons.md (the v1.3 orchestrator never landed on development, so we source from the experimental branch). Adds v1.4-specific header explaining how the analyzer (step 6) and optimizer (step 7) subagents will use this file. Replaces v1.3 "auto-improve-skill" / "Phase 4" framing with v1.4 step-numbering and subagent-naming. Content body is the v1.3 Recipe A-E + grader patterns G1-G6 + run-record protocol verbatim — that's the cumulative knowledge v1.4 inherits. --- skills/references/.gitkeep | 0 skills/references/recipes.md | 370 +++++++++++++++++++++++++++++++++++ 2 files changed, 370 insertions(+) delete mode 100644 skills/references/.gitkeep create mode 100644 skills/references/recipes.md diff --git a/skills/references/.gitkeep b/skills/references/.gitkeep deleted file mode 100644 index e69de29..0000000 diff --git a/skills/references/recipes.md b/skills/references/recipes.md new file mode 100644 index 0000000..f952fe4 --- /dev/null +++ b/skills/references/recipes.md @@ -0,0 +1,370 @@ +# skill-optimizer recipes — shared library + +> **For v1.4 skills:** This is the shared recipes library. The +> analyzer subagent (step 6) reads this to recognize known failure +> patterns. The optimizer subagent (step 7) reads this to choose a +> principled fix (Recipe A–E). Add new recipes here when a v1.4 run +> surfaces a generalizable pattern. Per-skill nuances belong in the +> individual run's `06-analysis.md` rather than here. + +This is a **living doc** seeded from the v1.3 auto-improve-skill +lessons.md. Every run adds patterns it discovered to the relevant +section. Run N benefits from patterns surfaced in runs 1..N-1; the +analyzer/optimizer don't have to rediscover them from zero. + +**How to use this from a subagent prompt:** The analyzer (step 6) reads +this file before classifying a failure cluster; match the observed +pattern to a recipe below. The optimizer (step 7) reads this file before +choosing a fix; pick the recipe the analyzer named and apply it +narrowly. If no recipe matches, the analyzer should do the diagnosis +from first principles and the run should add a new entry here at the end. + +--- + +## The load-bearing prior + +> **Rules about *absence* (a missing attribute, a missing branch, a +> missing focus replacement) are 5–10× harder for models than rules +> about *presence* (a literal token in the code).** + +Source: manual web-design-guidelines run + auto-pilot supabase pilot +(2026-05-08) — both surfaced this independently. Use it to categorize +every missed rule before deciding what to modify. + +| Rule pattern | Relative miss rate | What helps | +|---|---|---| +| Visible bad pattern (literal token in code) | low | Often catches itself; the rule wording is enough | +| Anti-pattern that "looks normal" (e.g., ` + +// GOOD: stays enabled. Spinner appears during the request. + +``` +```` + +**Empirical evidence:** manual web-design-guidelines run — the rules +that needed examples (submit-disabled, paste-blocking, missing +autoComplete, image priority hint) all closed their miss rates by 60-100% +after the example was added. + +### E. Rationale + bug-story + +**When to use:** state-machine violations, lifecycle bugs, +non-obvious-failure rules. + +**Recipe:** narrate the failure case inline with the rule. Example: + +```markdown +NEVER `disabled={!form.valid}` — the user types, then deletes a +character to fix a typo, the button flickers off, and the paste-fill +races with state. Tested users will assume the button is broken. +``` + +The narration gives the model a "why this rule matters" hook that pure +declarative rules don't provide. + +--- + +## Grader-reliability patterns (Phase-2 build-suite recipes) + +These are common ways graders go wrong on first build. Pre-tune your +graders to avoid them; if you see the failure mode at baseline, fix the +grader as iteration 1 (do not propose a skill change yet). + +### G1. Line tolerance ±5–8 (not ±0–3) + +LLM line-counting is unreliable. Models report violations 1-3 lines off +from the actual line in multi-line JSX/SQL/code. Use the `looseRange` +helper (default tolerance ±8): + +```javascript +{ id: 'rule-id', lines: looseRange(18), keywords: [/.../i] } +// Accepts lines 10-26. +``` + +`looseRange(N, tolerance)` is defined in `_grader-utils.mjs`. Prefer it +over hand-rolling `range(N-3, N+3)` — the default already absorbs the +common drift width seen across all 4 prior pilots. + +### G2. Hyphen-tolerant keyword regex + +Models output "empty-state" when the rule says "empty state", or +"clickable-handler" when the rule says "clickable handler". Use the +`fuzzyKeyword` helper: + +```javascript +keywords: [fuzzyKeyword('empty state')] // matches "empty state" and "empty-state" +keywords: [fuzzyKeyword('aria label')] // matches "aria-label" and "aria label" +``` + +`fuzzyKeyword(phrase)` is defined in `_grader-utils.mjs`. It escapes +regex metacharacters and replaces internal whitespace with `[-\s]*`, +so callers don't have to hand-roll the regex. + +### G3. Per-finding-line keyword matching (not whole-text) + +Don't `keywords.some(re => re.test(fullText))` — that produces spurious +cross-matches when keyword X appears in a different rule's finding line. +Use `_grader-utils.mjs`'s built-in per-finding-line matcher (split +findings.txt by line, match within each line). + +### G4. Multiple keyword variants + +Models phrase the same concept several ways: + ++ "covering" / "does not cover" / "missing covering index" ++ "label" / "aria-label" / "labeled" ++ "hover" / "hover state" / "hover:bg-*" + +Use the `tolerantKeyword` helper for word-stem matching: + +```javascript +keywords: [tolerantKeyword('cover')] // matches "cover", "covering", "covered" +keywords: [tolerantKeyword('label')] // matches "label", "labeled", "labels" +``` + +For multiple distinct stems on the same rule, use an array — the grader +treats them as alternatives: + +```javascript +keywords: [tolerantKeyword('hover'), fuzzyKeyword('hover state')] +``` + +Both `tolerantKeyword` and `fuzzyKeyword` are defined in `_grader-utils.mjs`. + +### G5. Set-semantics for sibling/list assertions + +When the grader checks a list of items, sort and compare — the model +emits items in different orders. + +```javascript +const names = pdf.repo_siblings_in_cohort_names.split(' | '); +assert.deepEqual(names.sort(), ['docx', 'xlsx']); // not deepEqual to ordered array +``` + +### G6. Verbosity floor for terse models + +Gemini sometimes outputs 3-4 line responses. Don't grade strict-pass on +"all 5 violations found" — many gemini failures are *truncated output*, +not missed rules. Compute rule-coverage rate (sum-found / sum-expected) +as the load-bearing metric instead of binary pass. + +--- + +## Default seeded violation types per skill shape + +When the auto-pilot builds a case in Phase 2, seed at least one +violation from each category for the skill's shape. This ensures +coverage of the absence-vs-presence axis and exposes whether the skill +needs Pattern A (two-pass workflow), Pattern C (per-element +checklists), or something else. + +### code-reviewer + +Seed at least one of each: + +1. Visible token misuse (e.g., `
` for action) +2. Missing attribute (e.g., `` without `autoComplete`) +3. Missing branch / no-empty-state (e.g., `array.map()` with no fallback for `[]`) +4. Anti-pattern that "looks normal" (e.g., `disabled={!form.valid}`) +5. State-machine violation (e.g., submit timing, focus on error) + +### tool-use / mcp-driver + +Seed at least one of each: + +1. Reaches-for-fallback (model uses `curl`/`npm i` instead of the prescribed CLI) +2. Wrong tool flag (passes `--user` when the skill calls for `--principal`) +3. Missing required step (skips snapshot, skips re-snapshot after action) +4. Output not validated (returns trace.jsonl without checking required artifacts) + +### document-producer + +Seed at least one of each: + +1. Missing required field in output (e.g., `answer.json` has no `risk_flags` key) +2. Wrong format (e.g., `2025-01-15` when the skill says `Intl.DateTimeFormat`) +3. Edge-case input (e.g., empty input, very long input, pre-corrupted file) +4. Format-only-correct: output validates but is unusable (e.g., PDF renders blank) + +### code-patterns + +Seed at least one of each: + +1. Wrong convention applied (skill says use 2-space indent, output uses 4) +2. Pattern not applied at all (skill says use `useReducer`, output uses `useState`) +3. Incorrect composition (uses prescribed pattern but in the wrong order) + +--- + +## Failure modes / known anti-patterns to avoid (Phase-4 don'ts) + +### Don't manufacture problems + +If baseline rule-coverage is ≥ 0.95, *exit clean*. Do not propose +modifications to a skill that already works. The goal is upstream PR +quality, not modification volume. + +**Source:** auto-pilot pdf pilot — baseline 1.00, no modifications +proposed. Maintainers will lose trust in our PRs if we open them for +non-issues. + +### Don't make breaking changes + +All proposed modifications must be **additive**: new sections, new +examples, new checklists. Never: + ++ Delete an existing rule ++ Change the wording of an existing rule ++ Reorder existing sections ++ Remove URLs or references in the skill + +This keeps the diff vs upstream small and the PR low-risk. + +### Don't burn iteration 1 on the wrong problem + +When baseline scores low, *first* check: is the grader the problem? Look +at the actual `findings.txt` from failed trials. If models *did* identify +the violations but the grader scored them wrong (line numbers off, +keyword mismatch, format variant), fix the grader as iteration 0 (don't +count it against the 2-iteration budget). + +**Source:** auto-pilot supabase + agent-browser pilots — both spent +iteration 1 on grader fixes before reaching skill modification. + +--- + +## Run-record protocol + +Every pilot adds an entry to one of these tables when it discovers +something new. Format: + +```markdown +**[skill-name] (date):** what was new — link to commit. +``` + +### Patterns added by pilots + ++ **manual web-design-guidelines (2026-05-06):** Two-pass workflow + per-element + checklists + 5 BAD/GOOD examples. Lifted 4-case suite from 72% → 86%. ++ **auto-pilot supabase (2026-05-08):** Independently rediscovered two-pass + workflow. Added it to a SQL skill. 0.54 → 0.86. ++ **auto-pilot agent-browser (2026-05-08):** Found that grader was over-strict + for non-interactive ops. Demoted snapshot from required to evidence-only. + Also surfaced "Verify-tool-installed nudge" pattern. ++ **auto-pilot pdf (2026-05-08):** Validated "exit clean on already-good skill" + — no modifications proposed; baseline 1.00. + +### Grader patterns added by pilots + ++ **manual web-design-guidelines (2026-05-06):** ±5-8 line tolerance, hyphen + regex, per-finding-line matching, keyword variants. ++ **auto-pilot supabase (2026-05-08):** "covering" / "does not cover" alternation + pattern. Confirmed ±3 → ±8 line widening is needed by default. + ++ **auto-pilot supabase v2 (2026-05-12):** Upstream constraints required adding a new reference file (`monitor-two-pass-review.md`) instead of editing SKILL.md. Baseline was already 1.00 (calibrated graders from prior run). Pattern: when a re-run starts from calibrated graders, the Phase 3 exit condition fires before Phase 4 — the "modification" step then serves purely as upstream PR packaging rather than eval improvement. + +(Future pilots: append your additions here.) From ebb9eabd33a41729de8a3fe2195e3382179f2dad Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 09:08:26 -0500 Subject: [PATCH 008/121] =?UTF-8?q?docs(v1.4):=20clean=20up=20skills/=20la?= =?UTF-8?q?yout=20=E2=80=94=20drop=20references/,=20scope=20subagents=20di?= =?UTF-8?q?r,=20move=20philosophy=20to=20docs/?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Changes to the v1.4 plan-as-executed, all in one place: - Move skill-writing-philosophy.md from skills/references/ to docs/. The philosophy doc is contributor-facing — we'll distill the load-bearing rules into the optimizer subagent's prompt directly rather than have it load this doc at runtime. - Delete skills/references/recipes.md (and the now-empty references/ dir). The raw seed copied from v1.3 lessons.md is case-study-shaped (accumulated per-pilot observations), which contradicts the "generalize-from-feedback" principle in the philosophy doc. The curated abstract-pattern version needs real end-to-end observations to ground it in — deferred to a follow-up after the chain ships. - Rename skills/subagents/ → skills/skill-optimizer-subagents/ so the dir name is scoped to this plugin and can't collide with subagents that other plugins might ship under a generic name. - Add a "Bootstrapping limit" section to the philosophy doc. The skill-optimizer chain runs an empirical loop on target skills, but that loop can't validate itself — same shape as Thompson's "Reflections on Trusting Trust". The seven meta-skills get authored from philosophy + best judgment; their test is Phase-E real-world runs, not eval data for the meta-skills themselves. - Spec + plan updated: file tree, Task A2 (revised), Task A3 (skipped), Task A4 (deferred), Phase C task file paths (subagents path rename), Phase D location (workflow.md goes to docs/ rather than skills/references/), acceptance criteria #4 (recipes.md deferred), coexistence section (v1.3 orchestrator never landed on development), open questions (recipes.md location TBD on real production). - Add note about post-v1.4 cleanup of the original skills/skill-optimizer/ skill (now redundant with the chain's skill-optimizer-run-bench step). Deferred to its own PR because it touches the public plugin API. --- docs/skill-optimizer-v1.4-plan.md | 291 ++++++-------- docs/skill-optimizer-v1.4-spec.md | 63 ++- docs/skill-writing-philosophy.md | 205 ++++++++++ skills/references/recipes.md | 370 ------------------ .../.gitkeep | 0 5 files changed, 367 insertions(+), 562 deletions(-) create mode 100644 docs/skill-writing-philosophy.md delete mode 100644 skills/references/recipes.md rename skills/{subagents => skill-optimizer-subagents}/.gitkeep (100%) diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/skill-optimizer-v1.4-plan.md index 7eaecf4..08cb9ff 100644 --- a/docs/skill-optimizer-v1.4-plan.md +++ b/docs/skill-optimizer-v1.4-plan.md @@ -60,24 +60,49 @@ skills/ skill-optimizer-run-bench/SKILL.md # Phase B (interactive) skill-optimizer-analyze-result/SKILL.md # Phase B (interactive) skill-optimizer-improve-skill/SKILL.md # Phase B (interactive) - subagents/ # Phase C (mechanical) + skill-optimizer-subagents/ # Phase C (interactive) research-functionality.md research-submissions.md test-writer.md analyzer.md optimizer.md validator.md - references/ # Phase D (mechanical) - workflow.md - recipes.md # seeded from v1.3 lessons.md - auto-improve-orchestrator/SKILL.md # MODIFIED: add deprecation banner docs/ skill-optimizer-v1.4-spec.md # already committed skill-optimizer-v1.4-plan.md # THIS file + skill-writing-philosophy.md # already committed (Phase B0) + skill-optimizer-workflow.md # Phase D (interactive) skill-optimizer-v1.4-validation.md # written by Phase E ``` +**Removed from the v1.4 scope** (vs the initial draft of this plan): + +- `skills/auto-improve-orchestrator/` deprecation banner — the v1.3 + orchestrator never landed on `development`, so there's nothing to + deprecate on this branch's lineage. Task A3 was skipped. +- `skills/references/recipes.md` (the seeded v1.3 lessons.md) — the + raw seed is case-study-shaped (accumulated per-pilot observations), + which contradicts the "generalize from feedback" principle in + [`skill-writing-philosophy.md`](skill-writing-philosophy.md). The + curation pass (turn it into named abstract patterns) is deferred to + a follow-up after the chain ships and we have real Phase-E + observations to ground the patterns in. +- `skills/references/` directory — no longer needed once recipes.md + is out and the philosophy doc lives under `docs/`. The subagents + dir is scoped to `skills/skill-optimizer-subagents/` so its + contents can't collide with other plugins' files. + +**Deferred to a post-v1.4 cleanup PR:** + +- Removing or repurposing `skills/skill-optimizer/SKILL.md` (the + original distributable wrapper). Under the v1.4 chain, the + direct-CLI-access role is filled by `skill-optimizer-run-bench`, + so the original skill is redundant. The cleanup touches the + public-facing plugin API (`.claude-plugin/`, README install docs, + `tests/smoke-skill-distribution.ts`) and should land in its own PR + after the v1.4 chain has been validated on a few external skills. + Per-skill state files (produced at runtime, NOT created by this plan; convention only): ```text @@ -97,9 +122,9 @@ docs/skill-optimizer// Each file's responsibility: - **`skills/skill-optimizer-/SKILL.md`** — frontmatter (`name:`, `description:`) + invocation instructions + dispatch-subagent logic + handoff-to-next-skill pointer. Always-loaded by Claude Code when the skill is invoked. -- **`skills/subagents/.md`** — narrow-context prompt template that the corresponding skill loads, substitutes inputs into, and dispatches via Agent tool. -- **`skills/references/workflow.md`** — human-readable chain diagram + per-skill brief, for operators understanding how the 7 skills compose. -- **`skills/references/recipes.md`** — accumulated lessons learned (Recipe A–E, grader patterns G1–G6, etc.), read by the analyzer and optimizer subagents. +- **`skills/skill-optimizer-subagents/.md`** — narrow-context prompt template that the corresponding skill loads, substitutes inputs into, and dispatches via Agent tool. Scoped name prevents collision with other plugins that might also ship subagents. +- **`docs/skill-optimizer-workflow.md`** — human-readable chain diagram + per-skill brief, for operators understanding how the 7 skills compose. Lives under `docs/` because it's contributor/operator reading, not a runtime resource the skills load. +- **`docs/skill-writing-philosophy.md`** — contributor reference distilling the three philosophies (Anthropic, skill-creator, superpowers:writing-skills), the skill-type → philosophy mapping, the bootstrapping limit, and the rules the optimizer subagent must follow when revising target skills. Also under `docs/` for the same reason. --- @@ -201,145 +226,65 @@ down the structural skeleton and discoverable frontmatter." --- -### Task A2: Create `subagents/` and `references/` directories - -**Files:** Create 2 directories. +### Task A2: Create `skill-optimizer-subagents/` directory -- [ ] **Step 1: Create directories** +**Files:** Create 1 directory + `.gitkeep` so it's tracked. -```bash -mkdir -p skills/subagents skills/references -ls skills/subagents skills/references -``` - -Expected: both directories exist and are empty. - -- [ ] **Step 2: Add `.gitkeep` placeholders so the empty dirs are tracked** - -```bash -touch skills/subagents/.gitkeep skills/references/.gitkeep -``` - -- [ ] **Step 3: Commit** +> **As-executed (revised mid-Phase-A):** the initial draft of this +> task created two directories, `skills/subagents/` and +> `skills/references/`. Both were revised: +> +> - `subagents/` was renamed to the scoped +> `skill-optimizer-subagents/` so its contents can't collide with +> other plugins that ship subagents. +> - `references/` was dropped entirely. The only file it would have +> held (`recipes.md`) was deferred (see superseded A4 below), and +> the philosophy doc moved to `docs/skill-writing-philosophy.md`. ```bash -git add skills/subagents/.gitkeep skills/references/.gitkeep -git commit -m "chore(v1.4): create subagents/ and references/ dirs (populated in Phase C+D)" +mkdir -p skills/skill-optimizer-subagents +touch skills/skill-optimizer-subagents/.gitkeep +git add skills/skill-optimizer-subagents/.gitkeep +git commit -m "chore(v1.4): create skill-optimizer-subagents/ dir (populated in Phase C)" ``` --- -### Task A3: Add deprecation banner to v1.3 `auto-improve-orchestrator/SKILL.md` - -**Files:** - -- Modify: `skills/auto-improve-orchestrator/SKILL.md` (first lines after frontmatter) - -- [ ] **Step 1: Read the current first lines** - -```bash -head -20 skills/auto-improve-orchestrator/SKILL.md -``` - -- [ ] **Step 2: Insert a deprecation banner immediately after the closing `---` of the frontmatter** - -Use Edit tool to insert this block after the second `---` line (closing the frontmatter): - -```markdown - -> **DEPRECATED — see v1.4.** This skill is superseded by the 7 -> independent skills under `skills/skill-optimizer-*` (v1.4). See -> `docs/skill-optimizer-v1.4-spec.md` for the architecture rationale -> and `skills/references/workflow.md` for the new chain. This skill -> is retained for backward compatibility during v1.4 rollout; it -> will be removed in a separate cleanup PR after v1.4 is validated -> on 3+ skills. - -``` - -- [ ] **Step 3: Verify the banner landed** - -```bash -grep -A 8 "DEPRECATED — see v1.4" skills/auto-improve-orchestrator/SKILL.md -``` - -Expected: the banner appears once, with the full text. - -- [ ] **Step 4: Commit** +### Task A3: ~~Add deprecation banner to v1.3 orchestrator~~ (SKIPPED) -```bash -git add skills/auto-improve-orchestrator/SKILL.md -git commit -m "docs(orchestrator): mark v1.3 orchestrator as deprecated-but-retained for v1.4 rollout" -``` +> **Skipped at execution time.** The v1.3 +> `skills/auto-improve-orchestrator/` skill does not exist on the +> `development` branch's lineage (it lives only on +> `feat/auto-improve-skill-v1.3` and a few `eval/auto-pilot/*` +> branches). There's nothing to deprecate on this branch, so the +> task is moot. If v1.3 were ever landed on `development` as a +> separate effort, revisit. --- -### Task A4: Seed `references/recipes.md` from v1.3 `lessons.md` - -**Files:** - -- Create: `skills/references/recipes.md` (seeded from `skills/auto-improve-orchestrator/references/lessons.md`) - -- [ ] **Step 1: Copy the lessons content, removing v1.3-specific framing** - -```bash -cp skills/auto-improve-orchestrator/references/lessons.md \ - skills/references/recipes.md -``` - -- [ ] **Step 2: Replace v1.3-specific phrases** - -Use Edit tool on `skills/references/recipes.md`: - -- Replace `auto-improve-skill-lessons.md` → `skills/references/recipes.md` (wherever it appears) -- Replace `auto-improve-skill` (the wrapper) → `skill-optimizer` (the engine) -- Replace `auto-improve-orchestrator` references with "v1.3 orchestrator (now deprecated)" - -If grep finds no instances, the file is already neutral — skip the edits. - -```bash -grep -n "auto-improve-skill-lessons.md\|auto-improve-orchestrator" skills/references/recipes.md | head -10 -``` +### Task A4: ~~Seed `references/recipes.md` from v1.3 `lessons.md`~~ (DEFERRED) -For each match, use Edit tool to make the replacement. - -- [ ] **Step 3: Add a header note explaining what this file IS in v1.4** - -Prepend this paragraph to the top of `skills/references/recipes.md` (right after the H1): - -```markdown -> **For v1.4 skills:** This is the shared recipes library. The -> analyzer subagent (step 6) reads this to recognize known failure -> patterns. The optimizer subagent (step 7) reads this to choose a -> principled fix (Recipe A–E). Add new recipes here when a v1.4 run -> surfaces a generalizable pattern. Per-skill nuances belong in the -> individual run's `06-analysis.md` rather than here. -``` - -- [ ] **Step 4: Verify the file parses + contains the seeded content** - -```bash -test -f skills/references/recipes.md -wc -l skills/references/recipes.md -head -5 skills/references/recipes.md -grep -c "^## " skills/references/recipes.md -``` - -Expected: file exists, ≥ 100 lines, has multiple `##` section headings. - -- [ ] **Step 5: Commit** - -```bash -git rm skills/subagents/.gitkeep # no longer needed once Phase C lands; keep here for now if Phase C hasn't run -# Actually skip the .gitkeep removal — keep simple: -git add skills/references/recipes.md -git commit -m "feat(v1.4): seed references/recipes.md from v1.3 lessons.md - -Adds v1.4-specific header explaining how the analyzer and optimizer -subagents will use this file. Content is the v1.3 lessons.md verbatim -(Recipe A-E + grader patterns G1-G6 + run-record protocol) with -v1.3-specific wrapper references neutralized." -``` +> **Deferred to a post-v1.4 follow-up.** The initial plan seeded a +> shared `recipes.md` from v1.3's `lessons.md` so the analyzer and +> optimizer subagents would have a pre-populated pattern library. +> On review, the raw seed is case-study-shaped (each pilot run +> appended its specific observations) rather than +> abstract-pattern-shaped, which contradicts the +> "generalize-from-feedback" principle in +> [`skill-writing-philosophy.md`](skill-writing-philosophy.md). +> +> The right shape for `recipes.md` is a small set of named abstract +> patterns (Recipe A–E, grader patterns G1–G6) with brief principled +> descriptions — not 370 lines of per-pilot case histories. Producing +> that curated version requires real Phase-E observations to ground +> the patterns in. +> +> Until then, the analyzer and optimizer subagents (Phase C) operate +> without a pre-loaded recipe library; they reason from +> first-principles using `01-functionality.md` and the per-run +> analysis. A `references/recipes.md` (or an equivalent under +> `docs/`) can be added later once we've validated the chain on a +> few external skills and have real cross-run patterns to name. --- @@ -389,7 +334,7 @@ Interface contract (from docs/skill-optimizer-v1.4-spec.md "The 7 skills" → § 4. Fetch the skill files (if URL); vendor to vendored-skill/ for read-only use by downstream steps 5. Web-search the underlying technology; identify trigger conditions, success criteria, key terminology, intended audience 6. Write structured report covering: what the skill does, who uses it, when it should fire, what tools it depends on, what concepts the user must understand -- Dispatches: functionality-researcher subagent (skills/subagents/research-functionality.md). Limited context: source skill + targeted web fetches. +- Dispatches: functionality-researcher subagent (skills/skill-optimizer-subagents/research-functionality.md). Limited context: source skill + targeted web fetches. - Handoff: "Next, invoke skill-optimizer-investigate-test-case. Note: this report's pr_submission_intent field tells step 2's handoff whether step 3 (investigate-submissions) should run." The frontmatter (name + description) is already in place from Task A1 — keep it, only replace the body. @@ -487,7 +432,7 @@ Interface contract (from spec §3, OPTIONAL — upstream-only): - Input: source slug // (from 01-functionality.md frontmatter) - Output: docs/skill-optimizer//03-submissions.md — license, CLA, frontmatter spec, file-location rules, prefix taxonomy, PR-shape patterns from last 10 merged PRs, branch target, rejection signals from last 5 closed-without-merge PRs - Behavior: gh-CLI heavy (PR list, repo-file API, CONTRIBUTING, sanity-test source, last 10 merged + last 5 closed); produce verbatim-pastable context block for the validator -- Dispatches: submission-researcher subagent (skills/subagents/research-submissions.md). Limited context: public repo facts only. +- Dispatches: submission-researcher subagent (skills/skill-optimizer-subagents/research-submissions.md). Limited context: public repo facts only. - Skipped when: pr_submission_intent: false in 01-functionality.md (this skill should detect that and exit cleanly with a "skip" message) - Handoff: "Validator in skill-optimizer-improve-skill will read this for external consistency check. Continue with skill-optimizer-write-tests if not done." @@ -532,7 +477,7 @@ Interface contract (from spec §4): 2. Show user the plan, ask for confirmation 3. Dispatch parallel subagents (one per picked test case) — each builds one workspace file + one grader 4. Run smoke check (hand-crafted GOOD/BAD/EMPTY findings.txt fixtures against each grader) — must pass before commit -- Dispatches: test-writer subagent per case (skills/subagents/test-writer.md). LIMITED CONTEXT per spec: single case spec + functionality report; does NOT see skill content or other test cases. +- Dispatches: test-writer subagent per case (skills/skill-optimizer-subagents/test-writer.md). LIMITED CONTEXT per spec: single case spec + functionality report; does NOT see skill content or other test cases. - Parallelizable: each case independent; dispatch in a single message - Handoff: "Invoke skill-optimizer-run-bench to measure baseline." ``` @@ -617,7 +562,7 @@ Interface contract (from spec §6 — this is the LOAD-BEARING skill that determ 3. For each systematic cluster: find responsible skill section, hypothesize cause, articulate "what WOULD address" (general principle) + "what WOULD NOT address" (anti-ducktape gate) 4. List non-structural noise separately 5. If no structural weakness can be articulated: report explicitly that no weakness was found; the next step (improve-skill) will refuse to fire — this is the honest-no-fabricated-uplift behavior -- Dispatches: analyzer subagent (skills/subagents/analyzer.md). LIMITED CONTEXT per spec: per-trial findings + skill content + workbench cases; does NOT see the test inputs themselves (forces focus on SKILL, not solutions). +- Dispatches: analyzer subagent (skills/skill-optimizer-subagents/analyzer.md). LIMITED CONTEXT per spec: per-trial findings + skill content + workbench cases; does NOT see the test inputs themselves (forces focus on SKILL, not solutions). - Handoff: if ≥1 structural weakness → "Invoke skill-optimizer-improve-skill." If none → exit honestly. 06-analysis.md format (Option A — structured): @@ -674,8 +619,8 @@ Interface contract (from spec §7): - Outputs: docs/skill-optimizer//07-improvement-proposal.md (diff + rationale referencing the structural weakness) + 07-validator-verdict.md + modified skill file (if approved) - Behavior: 1. Refuse if 06-analysis.md has no structural weakness — print "no weakness to address" and exit cleanly - 2. Dispatch OPTIMIZER subagent (skills/subagents/optimizer.md) with limited context: sees analysis report + functionality + skill content; does NOT see raw failures, grader logic, test inputs. MUST address named structural weakness using a general principle (NOT a pattern-match patch). Output proposed diff + rationale that explicitly references which named weakness it addresses. - 3. Dispatch VALIDATOR subagent (skills/subagents/validator.md) with limited context: sees BEFORE skill + AFTER skill + 01-functionality + 03-submissions (if exists). Two-part check: + 2. Dispatch OPTIMIZER subagent (skills/skill-optimizer-subagents/optimizer.md) with limited context: sees analysis report + functionality + skill content; does NOT see raw failures, grader logic, test inputs. MUST address named structural weakness using a general principle (NOT a pattern-match patch). Output proposed diff + rationale that explicitly references which named weakness it addresses. + 3. Dispatch VALIDATOR subagent (skills/skill-optimizer-subagents/validator.md) with limited context: sees BEFORE skill + AFTER skill + 01-functionality + 03-submissions (if exists). Two-part check: - INTERNAL consistency: does the change make sense given the skill's stated responsibilities? Additive vs destructive? General vs ducttape? - EXTERNAL consistency (only if 03-submissions.md exists): does change conform to upstream PR rules (frontmatter, file location, prefix taxonomy, additive-only, etc.)? - Verdict: approve / needs-revision / reject @@ -718,7 +663,7 @@ For EACH Task C: ### Task C1: `subagents/research-functionality.md` -**Files:** Create `skills/subagents/research-functionality.md`. +**Files:** Create `skills/skill-optimizer-subagents/research-functionality.md`. - [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** @@ -731,7 +676,7 @@ Invoke `/skill superpowers:writing-skills`. Paste the draft template below as th The draft template: -Create `skills/subagents/research-functionality.md` with this exact content: +Create `skills/skill-optimizer-subagents/research-functionality.md` with this exact content: ````markdown # Sub-subagent: research a skill's functionality @@ -826,7 +771,7 @@ Return to the calling skill a brief summary (under 200 words) covering: classifi - [ ] **Step 2: Verify template variables present** ```bash -grep -E '\$\{(SKILL_SOURCE|OUTPUT_PATH|PR_SUBMISSION_INTENT)\}' skills/subagents/research-functionality.md | wc -l +grep -E '\$\{(SKILL_SOURCE|OUTPUT_PATH|PR_SUBMISSION_INTENT)\}' skills/skill-optimizer-subagents/research-functionality.md | wc -l ``` Expected: ≥ 4 occurrences. @@ -834,7 +779,7 @@ Expected: ≥ 4 occurrences. - [ ] **Step 3: Commit** ```bash -git add skills/subagents/research-functionality.md +git add skills/skill-optimizer-subagents/research-functionality.md git commit -m "feat(v1.4-subagents): research-functionality prompt template" ``` @@ -842,7 +787,7 @@ git commit -m "feat(v1.4-subagents): research-functionality prompt template" ### Task C2: `subagents/research-submissions.md` -**Files:** Create `skills/subagents/research-submissions.md`. +**Files:** Create `skills/skill-optimizer-subagents/research-submissions.md`. - [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** @@ -855,7 +800,7 @@ Invoke `/skill superpowers:writing-skills`. Paste the draft template below as th The draft template: -Create `skills/subagents/research-submissions.md` with this exact content (adapted from v1.3's `prompts/research-upstream.md`): +Create `skills/skill-optimizer-subagents/research-submissions.md` with this exact content (adapted from v1.3's `prompts/research-upstream.md`): ````markdown # Sub-subagent: research upstream PR submission conventions @@ -946,10 +891,10 @@ Under 400 words. Include: license + CLA verdict; recommended branch target; risk - [ ] **Step 2: Verify + commit** ```bash -grep -E '\$\{(SLUG|OUTPUT_PATH)\}' skills/subagents/research-submissions.md | wc -l +grep -E '\$\{(SLUG|OUTPUT_PATH)\}' skills/skill-optimizer-subagents/research-submissions.md | wc -l # Expected: ≥ 4 -git add skills/subagents/research-submissions.md +git add skills/skill-optimizer-subagents/research-submissions.md git commit -m "feat(v1.4-subagents): research-submissions prompt template" ``` @@ -957,7 +902,7 @@ git commit -m "feat(v1.4-subagents): research-submissions prompt template" ### Task C3: `subagents/test-writer.md` -**Files:** Create `skills/subagents/test-writer.md`. +**Files:** Create `skills/skill-optimizer-subagents/test-writer.md`. - [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** @@ -970,7 +915,7 @@ Invoke `/skill superpowers:writing-skills`. Paste the draft template below as th The draft template: -Create `skills/subagents/test-writer.md` with this exact content: +Create `skills/skill-optimizer-subagents/test-writer.md` with this exact content: ````markdown # Sub-subagent: write one test case for the eval workbench @@ -1044,10 +989,10 @@ Under 200 words. Include: filenames created, smoke-check result, any blockers. - [ ] **Step 2: Verify + commit** ```bash -grep -E '\$\{(CASE_SPEC|FUNCTIONALITY_REPORT_EXCERPT|WORKBENCH_DIR|CASE_NAME)\}' skills/subagents/test-writer.md | wc -l +grep -E '\$\{(CASE_SPEC|FUNCTIONALITY_REPORT_EXCERPT|WORKBENCH_DIR|CASE_NAME)\}' skills/skill-optimizer-subagents/test-writer.md | wc -l # Expected: ≥ 8 -git add skills/subagents/test-writer.md +git add skills/skill-optimizer-subagents/test-writer.md git commit -m "feat(v1.4-subagents): test-writer prompt template" ``` @@ -1055,7 +1000,7 @@ git commit -m "feat(v1.4-subagents): test-writer prompt template" ### Task C4: `subagents/analyzer.md` -**Files:** Create `skills/subagents/analyzer.md`. +**Files:** Create `skills/skill-optimizer-subagents/analyzer.md`. - [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** @@ -1068,7 +1013,7 @@ Invoke `/skill superpowers:writing-skills`. Paste the draft template below as th The draft template: -Create `skills/subagents/analyzer.md` with this exact content: +Create `skills/skill-optimizer-subagents/analyzer.md` with this exact content: ````markdown # Sub-subagent: analyze bench results, identify structural weaknesses @@ -1155,10 +1100,10 @@ Under 300 words. Include: verdict (weakness-found vs no-weakness), weakness name - [ ] **Step 2: Verify + commit** ```bash -grep -E '\$\{(BENCH_RESULTS_DIR|WORKBENCH_DIR|SKILL_SOURCE_DIR|OUTPUT_PATH|RECIPES_PATH)\}' skills/subagents/analyzer.md | wc -l +grep -E '\$\{(BENCH_RESULTS_DIR|WORKBENCH_DIR|SKILL_SOURCE_DIR|OUTPUT_PATH|RECIPES_PATH)\}' skills/skill-optimizer-subagents/analyzer.md | wc -l # Expected: ≥ 8 -git add skills/subagents/analyzer.md +git add skills/skill-optimizer-subagents/analyzer.md git commit -m "feat(v1.4-subagents): analyzer prompt template" ``` @@ -1166,7 +1111,7 @@ git commit -m "feat(v1.4-subagents): analyzer prompt template" ### Task C5: `subagents/optimizer.md` -**Files:** Create `skills/subagents/optimizer.md`. +**Files:** Create `skills/skill-optimizer-subagents/optimizer.md`. - [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** @@ -1179,7 +1124,7 @@ Invoke `/skill superpowers:writing-skills`. Paste the draft template below as th The draft template: -Create `skills/subagents/optimizer.md` with this exact content: +Create `skills/skill-optimizer-subagents/optimizer.md` with this exact content: ````markdown # Sub-subagent: propose a principled fix to address a structural weakness @@ -1278,10 +1223,10 @@ Recommendation: ` pattern across spec, file paths, brief excerpts - State file paths consistent: `docs/skill-optimizer//-.md` (or `workbench/`, `05-bench-results//`, `vendored-skill/`) throughout -- Subagent file paths consistent: `skills/subagents/.md` throughout +- Subagent file paths consistent: `skills/skill-optimizer-subagents/.md` throughout - Template variable names (`${SLUG}`, `${WORKBENCH_DIR}`, etc.) match between subagent templates and skills' planned invocation patterns No issues found. diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 68f4b7b..441e158 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -72,18 +72,31 @@ skills/ skill-optimizer-analyze-result/SKILL.md # 6 skill-optimizer-improve-skill/SKILL.md # 7 skill-optimizer/SKILL.md # existing — unchanged - subagents/ + skill-optimizer-subagents/ research-functionality.md research-submissions.md test-writer.md analyzer.md optimizer.md validator.md - references/ - workflow.md # the chain graph - recipes.md # accumulated lessons ``` +Companion docs live under `docs/`, not `skills/`, because they're +contributor / operator reading rather than runtime resources the +skills load: + +```text +docs/ + skill-optimizer-workflow.md # the chain graph + skill-writing-philosophy.md # authoring guidance +``` + +A shared `recipes.md` (named abstract failure patterns shared across +analyzer / optimizer subagents) was considered but deferred — the +v1.3 lessons.md seed is case-study-shaped, and producing a curated +pattern library requires real end-to-end observations to ground the +patterns in. See plan §"Task A4" for details. + **State at convention path** (visible + committable, per `superpowers` precedent): @@ -375,15 +388,23 @@ separate workstream (out of scope for v1.4). The v1.3 `skills/auto-improve-orchestrator/` skill is **deprecated but retained** for backward compatibility during the v1.4 rollout: -- v1.3 stays installed; existing context files at - `skills/auto-improve-orchestrator/references/contexts/` are kept - (they are useful inputs to v1.4's step 3 as initial drafts) -- v1.3's lessons.md becomes the seed for v1.4's `references/recipes.md` -- The v1.3 orchestrator is marked as superseded in its own SKILL.md - with a pointer to v1.4 - -After v1.4 is validated on 3+ skills, v1.3 can be removed in a -separate cleanup PR. +- The v1.3 `auto-improve-orchestrator/` skill never landed on + `development` (it lives on `feat/auto-improve-skill-v1.3` and a few + experimental branches), so there's no on-branch artifact to mark + deprecated. The intended deprecation banner (initial plan task A3) + was skipped during execution. Existing v1.3 context files under + experimental branches remain useful as draft inputs to v1.4's step + 3, but they're not in v1.4's lineage. +- v1.3's `lessons.md` was originally planned as the seed for a + `references/recipes.md` shared pattern library. That was deferred + (see "Architecture overview" / Plan §"Task A4") — the raw seed is + case-study-shaped and a curated pattern version needs real + end-to-end observations. + +After v1.4 is validated on a few skills, the original +`skills/skill-optimizer/SKILL.md` (the direct workbench-CLI wrapper) +can be removed or repurposed in a separate cleanup PR — under the +v1.4 chain its role is filled by `skill-optimizer-run-bench`. ## Acceptance criteria @@ -391,11 +412,14 @@ For v1.4 to be considered done: 1. All 7 SKILL.md files exist with the interface described in this spec, plus appropriate "Next, invoke X" handoff instructions -2. Subagent prompt templates exist in `skills/subagents/` with the +2. Subagent prompt templates exist in `skills/skill-optimizer-subagents/` with the limited-context constraints described in the "Subagent constraints" table -3. `references/workflow.md` documents the chain visually -4. `references/recipes.md` is seeded from v1.3's `lessons.md` +3. `docs/skill-optimizer-workflow.md` documents the chain visually +4. ~~`references/recipes.md` is seeded from v1.3's `lessons.md`~~ — + deferred; the analyzer / optimizer subagents operate without a + pre-loaded recipe library until real end-to-end observations + surface curated patterns 5. **End-to-end test on a local skill** (e.g., one of the existing `skills/skill-optimizer/SKILL.md` or a small new skill in this repo) — walks 1→2→4→5→6→7, produces all expected reports, modifies @@ -428,9 +452,10 @@ For v1.4 to be considered done: ## Open questions (tracked but deferred) 1. **When does a SKILL.md grow recipes vs reference an external - recipes.md?** Initial answer: shared recipes go in - `references/recipes.md`; per-skill nuances go inline. Will revisit - based on usage. + `recipes.md`?** Initial answer: shared recipes would go in a + curated `recipes.md` (location TBD when the file is actually + produced — see Plan §"Task A4" for the deferral); per-skill + nuances go inline. Will revisit based on usage. 2. **What if the optimizer subagent's revision loop exceeds the max rounds (2)?** Initial answer: surface as `validator-rejected` and require human intervention. Could add a "human-help-required" diff --git a/docs/skill-writing-philosophy.md b/docs/skill-writing-philosophy.md new file mode 100644 index 0000000..8f9c539 --- /dev/null +++ b/docs/skill-writing-philosophy.md @@ -0,0 +1,205 @@ +# Skill-writing philosophy + +> Read this before authoring or revising any SKILL.md in this project, +> and before designing the optimizer subagent's prompt — choosing the +> wrong philosophy for the skill type is itself a form of ducktape. + +## Background + +There are three named bodies of guidance on writing Claude Agent Skills. +They mostly agree, but they diverge on tone and testing rigor in ways +that matter: + +1. **Anthropic official** — the canonical [skill authoring best + practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices). +2. **skill-creator** — Anthropic's plugin for iterative skill authoring + with an eval-viewer + description-optimizer loop. +3. **superpowers:writing-skills** — third-party plugin (obra/superpowers) + that frames skill authoring as TDD applied to process documentation. + +## Consensus (load-bearing — all three sources agree) + +- **The description is the trigger mechanism.** Claude scans + descriptions to decide which skills to consult. Get this right or the + skill never fires. +- **SKILL.md stays ≤ ~500 lines; bulk goes in `references/`, scripts in + `scripts/`.** This is progressive disclosure. +- **References are one level deep from SKILL.md** — nested references + cause partial reads and lost info. +- **Evaluation-driven development beats imagined-requirements + development.** Run baseline tasks, document where the agent fails, + write the skill to address those specific failures. +- **Use Claude to iterate on Claude's skills** — one instance drafts, + another tests, observations feed back to the drafter. + +## Divergence (where you must pick) + +| Axis | Anthropic + skill-creator | superpowers:writing-skills | +|---|---|---| +| Tone for must-follow rules | "Explain the WHY; MUSTs/NEVERs in all caps are a yellow flag" | "Authority framing, no exceptions, close every loophole" | +| What gets a skill | Workflows, techniques, references | Same, **plus** discipline-enforcing rules (TDD, verification) | +| Testing rigor | ≥3 evals (Anthropic) / full with-skill vs baseline pipeline (skill-creator) | Pressure scenarios with combined pressures (time + sunk cost + authority) | +| Description content | "Include both what + when" | "**Only** when, **never** what — the body becomes documentation Claude skips otherwise" | + +**The divergence is not a contradiction.** writing-skills explicitly +says: discipline-enforcing skills get authority framing + pressure +tests; reference / technique skills get application tests. The other +two sources mostly assume the workflow / technique case and don't +address discipline rules. + +## Skill type → philosophy mapping + +This is the single most important decision when authoring or revising +a skill. Diagnose first, write/revise second. + +| Skill type | Examples | Philosophy | Testing | +|---|---|---|---| +| **Workflow** | "do X, then Y, then dispatch Z" | Anthropic + skill-creator: lean prose, explain why, set freedom level per step | with-skill vs baseline subagents on realistic task prompts | +| **Technique** | "how to use library X correctly" | Anthropic: concise, one excellent example, edge-case notes | application + variation + gap-coverage scenarios | +| **Reference** | API docs, schema docs, vendor conventions | Anthropic: table-of-contents at top, keyword-rich, no narrative | retrieval scenarios — can the agent find + apply the right info? | +| **Pattern** | mental models, design heuristics | Anthropic: recognition examples + counter-examples + when-NOT-to-use | recognition scenarios — does the agent know when to apply? | +| **Discipline-enforcing** | "always do X before Y" rules, anti-ducktape constraints | writing-skills: authority framing, close every loophole, rationalization table, red-flags list, persuasion principles | pressure scenarios with multiple combined pressures (time + sunk cost + authority + exhaustion) | + +**Mixed-type skills are common.** Most skill-optimizer chain skills +are workflows with one or two embedded discipline rules (e.g., +"dispatch the subagent, do NOT do the research yourself"). The right +pattern: + +1. Write the workflow body in Anthropic/skill-creator style — lean, + explain why, set freedom per step. +2. For each embedded discipline rule, mark it explicitly (bold, "Do + NOT" framing at the action site) AND give a why-this-matters + paragraph nearby. +3. Reserve full writing-skills treatment (rationalization tables, + red-flags lists, no-exceptions language) for the small handful of + rules where compliance under pressure is load-bearing and was + verified via pressure scenarios. + +## The bootstrapping limit + +The skill-optimizer chain runs an empirical loop (eval → analyze → +improve) on target skills. But that loop cannot validate **itself** — +you can't use the chain to author its own seven SKILL.md files, +because the chain doesn't exist yet when those files are being +written. This is the same shape as Thompson's "Reflections on +Trusting Trust": validating a compiler with itself is circular. + +Practical consequences: + +- **The seven skill-optimizer SKILL.md files are authored from + philosophy + best judgment**, not from an eval loop. The "test" + for these skills is end-to-end runs on real targets and ongoing + observation in actual use. + +- **The optimizer subagent should not be surprised by the absence of + eval data for the skill-optimizer's own skills.** If asked to + improve one of them, it should treat absence of empirical data as + a known limit, not a gap to fill speculatively. Real-world + observations from end-to-end runs are the only valid signal for + self-improvement of the chain itself. + +- **Self-application is deferred** until the chain has earned trust + on external targets. Running skill-optimizer on the skill-optimizer + is a future-state exercise; doing it pre-launch is bootstrapping in + a loop. + +## Description: triggers, not summaries + +writing-skills' strongest empirical finding: when a description +summarizes the skill's workflow, agents follow the summary instead of +reading the full skill body. The skill body becomes documentation the +agent skips. + +The example they cite: a description saying "code review between tasks" +caused agents to do ONE review even though the skill body clearly +specified TWO. Switching the description to pure trigger language +("Use when executing implementation plans with independent tasks in +the current session") restored compliance. + +**Practical guidance for skill-optimizer descriptions:** + +- Start with "Use when ..." and list the user-language symptoms that + should trigger this skill. +- Do NOT summarize the workflow, the dispatches, or the output format + in the description. Those go in the body. +- Use concrete trigger phrases, not abstract ones: "Use when the user + asks 'what does this skill do' or 'investigate this skill'" beats + "Use when investigating skills." +- Anthropic's "both what + when" guidance is fine for simple + workflow skills with no embedded discipline rules. For skills with + embedded discipline (most of the skill-optimizer chain skills), + trigger-only is safer. +- skill-creator's "be a little pushy" guidance applies: Claude tends to + under-trigger; include adjacent phrasings explicitly. + +## When the optimizer is revising someone else's skill + +This is the meta-payoff and the load-bearing reason this doc exists. +The Phase 7 optimizer subagent revises target skills based on weaknesses +the analyzer surfaced. To avoid ducktape, the optimizer MUST: + +1. **Diagnose the target skill's type before proposing changes.** A + reference skill and a discipline-enforcing skill need different + fixes for the same observed failure. + +2. **Preserve the existing style unless the analyzer flagged the style + itself as the weakness.** If the original uses explanatory prose, + the fix uses explanatory prose. Don't impose authority framing + because MUSTs feel clearer to the optimizer — that's ducktape. + +3. **Bias toward "Claude is smart" — pruning beats adding.** If the + skill restates what Claude already knows, removing the restatement + is often a more principled fix than adding new rules. Token weight + competes with conversation context once the skill loads. + +4. **Reserve authority framing / rationalization tables / + no-exceptions language for skills where the analyzer specifically + documented a discipline failure under pressure.** Adding MUSTs to + patch a reference-skill bug is the canonical ducktape pattern. + +5. **Test the proposed change against a scenario the analyzer + documented, not a new scenario invented by the optimizer.** + Optimizer-invented scenarios drift into solving imagined problems + instead of the real surfaced weakness. + +6. **Treat description changes as separate, conservative edits.** The + description determines triggering, not behavior. Changing the + description to "fix" a behavioral problem is misdirected. Change + the body for behavior; change the description only if the analyzer + flagged a trigger problem (over-triggering or under-triggering). + +## When to use each philosophy: a quick decision tree + +```text +Question 1: Is this an absolute rule that an agent might try to + rationalize away under pressure? + └─ YES → discipline-enforcing skill → writing-skills philosophy + └─ NO → continue to Q2 + +Question 2: Is this a reference doc (API, schema, conventions) the + agent will scan to retrieve facts? + └─ YES → reference skill → Anthropic philosophy + TOC at top + └─ NO → continue to Q3 + +Question 3: Is this a sequence of steps the agent follows to + accomplish a task? + └─ YES → workflow skill → Anthropic + skill-creator philosophy + (lean prose, explain why, freedom level per step) + └─ NO → it's likely a pattern / mental model + → Anthropic philosophy (recognition + counter-examples) +``` + +If the skill is mixed (workflow with embedded discipline rules): +write the body in workflow style, mark the discipline rules +explicitly at their action sites, and reserve the heavy +writing-skills treatment (rationalization tables etc.) for the +specific rules whose compliance under pressure was actually verified. + +## References + +- [Anthropic skill authoring best practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices) +- skill-creator plugin: `~/.claude/plugins/cache/claude-plugins-official/skill-creator/` +- superpowers:writing-skills: + `~/.claude/plugins/cache/claude-plugins-official/superpowers//skills/writing-skills/` +- Meincke et al. 2025 — persuasion principles in AI compliance (cited + by writing-skills/persuasion-principles.md) diff --git a/skills/references/recipes.md b/skills/references/recipes.md deleted file mode 100644 index f952fe4..0000000 --- a/skills/references/recipes.md +++ /dev/null @@ -1,370 +0,0 @@ -# skill-optimizer recipes — shared library - -> **For v1.4 skills:** This is the shared recipes library. The -> analyzer subagent (step 6) reads this to recognize known failure -> patterns. The optimizer subagent (step 7) reads this to choose a -> principled fix (Recipe A–E). Add new recipes here when a v1.4 run -> surfaces a generalizable pattern. Per-skill nuances belong in the -> individual run's `06-analysis.md` rather than here. - -This is a **living doc** seeded from the v1.3 auto-improve-skill -lessons.md. Every run adds patterns it discovered to the relevant -section. Run N benefits from patterns surfaced in runs 1..N-1; the -analyzer/optimizer don't have to rediscover them from zero. - -**How to use this from a subagent prompt:** The analyzer (step 6) reads -this file before classifying a failure cluster; match the observed -pattern to a recipe below. The optimizer (step 7) reads this file before -choosing a fix; pick the recipe the analyzer named and apply it -narrowly. If no recipe matches, the analyzer should do the diagnosis -from first principles and the run should add a new entry here at the end. - ---- - -## The load-bearing prior - -> **Rules about *absence* (a missing attribute, a missing branch, a -> missing focus replacement) are 5–10× harder for models than rules -> about *presence* (a literal token in the code).** - -Source: manual web-design-guidelines run + auto-pilot supabase pilot -(2026-05-08) — both surfaced this independently. Use it to categorize -every missed rule before deciding what to modify. - -| Rule pattern | Relative miss rate | What helps | -|---|---|---| -| Visible bad pattern (literal token in code) | low | Often catches itself; the rule wording is enough | -| Anti-pattern that "looks normal" (e.g., ` - -// GOOD: stays enabled. Spinner appears during the request. - -``` -```` - -**Empirical evidence:** manual web-design-guidelines run — the rules -that needed examples (submit-disabled, paste-blocking, missing -autoComplete, image priority hint) all closed their miss rates by 60-100% -after the example was added. - -### E. Rationale + bug-story - -**When to use:** state-machine violations, lifecycle bugs, -non-obvious-failure rules. - -**Recipe:** narrate the failure case inline with the rule. Example: - -```markdown -NEVER `disabled={!form.valid}` — the user types, then deletes a -character to fix a typo, the button flickers off, and the paste-fill -races with state. Tested users will assume the button is broken. -``` - -The narration gives the model a "why this rule matters" hook that pure -declarative rules don't provide. - ---- - -## Grader-reliability patterns (Phase-2 build-suite recipes) - -These are common ways graders go wrong on first build. Pre-tune your -graders to avoid them; if you see the failure mode at baseline, fix the -grader as iteration 1 (do not propose a skill change yet). - -### G1. Line tolerance ±5–8 (not ±0–3) - -LLM line-counting is unreliable. Models report violations 1-3 lines off -from the actual line in multi-line JSX/SQL/code. Use the `looseRange` -helper (default tolerance ±8): - -```javascript -{ id: 'rule-id', lines: looseRange(18), keywords: [/.../i] } -// Accepts lines 10-26. -``` - -`looseRange(N, tolerance)` is defined in `_grader-utils.mjs`. Prefer it -over hand-rolling `range(N-3, N+3)` — the default already absorbs the -common drift width seen across all 4 prior pilots. - -### G2. Hyphen-tolerant keyword regex - -Models output "empty-state" when the rule says "empty state", or -"clickable-handler" when the rule says "clickable handler". Use the -`fuzzyKeyword` helper: - -```javascript -keywords: [fuzzyKeyword('empty state')] // matches "empty state" and "empty-state" -keywords: [fuzzyKeyword('aria label')] // matches "aria-label" and "aria label" -``` - -`fuzzyKeyword(phrase)` is defined in `_grader-utils.mjs`. It escapes -regex metacharacters and replaces internal whitespace with `[-\s]*`, -so callers don't have to hand-roll the regex. - -### G3. Per-finding-line keyword matching (not whole-text) - -Don't `keywords.some(re => re.test(fullText))` — that produces spurious -cross-matches when keyword X appears in a different rule's finding line. -Use `_grader-utils.mjs`'s built-in per-finding-line matcher (split -findings.txt by line, match within each line). - -### G4. Multiple keyword variants - -Models phrase the same concept several ways: - -+ "covering" / "does not cover" / "missing covering index" -+ "label" / "aria-label" / "labeled" -+ "hover" / "hover state" / "hover:bg-*" - -Use the `tolerantKeyword` helper for word-stem matching: - -```javascript -keywords: [tolerantKeyword('cover')] // matches "cover", "covering", "covered" -keywords: [tolerantKeyword('label')] // matches "label", "labeled", "labels" -``` - -For multiple distinct stems on the same rule, use an array — the grader -treats them as alternatives: - -```javascript -keywords: [tolerantKeyword('hover'), fuzzyKeyword('hover state')] -``` - -Both `tolerantKeyword` and `fuzzyKeyword` are defined in `_grader-utils.mjs`. - -### G5. Set-semantics for sibling/list assertions - -When the grader checks a list of items, sort and compare — the model -emits items in different orders. - -```javascript -const names = pdf.repo_siblings_in_cohort_names.split(' | '); -assert.deepEqual(names.sort(), ['docx', 'xlsx']); // not deepEqual to ordered array -``` - -### G6. Verbosity floor for terse models - -Gemini sometimes outputs 3-4 line responses. Don't grade strict-pass on -"all 5 violations found" — many gemini failures are *truncated output*, -not missed rules. Compute rule-coverage rate (sum-found / sum-expected) -as the load-bearing metric instead of binary pass. - ---- - -## Default seeded violation types per skill shape - -When the auto-pilot builds a case in Phase 2, seed at least one -violation from each category for the skill's shape. This ensures -coverage of the absence-vs-presence axis and exposes whether the skill -needs Pattern A (two-pass workflow), Pattern C (per-element -checklists), or something else. - -### code-reviewer - -Seed at least one of each: - -1. Visible token misuse (e.g., `
` for action) -2. Missing attribute (e.g., `` without `autoComplete`) -3. Missing branch / no-empty-state (e.g., `array.map()` with no fallback for `[]`) -4. Anti-pattern that "looks normal" (e.g., `disabled={!form.valid}`) -5. State-machine violation (e.g., submit timing, focus on error) - -### tool-use / mcp-driver - -Seed at least one of each: - -1. Reaches-for-fallback (model uses `curl`/`npm i` instead of the prescribed CLI) -2. Wrong tool flag (passes `--user` when the skill calls for `--principal`) -3. Missing required step (skips snapshot, skips re-snapshot after action) -4. Output not validated (returns trace.jsonl without checking required artifacts) - -### document-producer - -Seed at least one of each: - -1. Missing required field in output (e.g., `answer.json` has no `risk_flags` key) -2. Wrong format (e.g., `2025-01-15` when the skill says `Intl.DateTimeFormat`) -3. Edge-case input (e.g., empty input, very long input, pre-corrupted file) -4. Format-only-correct: output validates but is unusable (e.g., PDF renders blank) - -### code-patterns - -Seed at least one of each: - -1. Wrong convention applied (skill says use 2-space indent, output uses 4) -2. Pattern not applied at all (skill says use `useReducer`, output uses `useState`) -3. Incorrect composition (uses prescribed pattern but in the wrong order) - ---- - -## Failure modes / known anti-patterns to avoid (Phase-4 don'ts) - -### Don't manufacture problems - -If baseline rule-coverage is ≥ 0.95, *exit clean*. Do not propose -modifications to a skill that already works. The goal is upstream PR -quality, not modification volume. - -**Source:** auto-pilot pdf pilot — baseline 1.00, no modifications -proposed. Maintainers will lose trust in our PRs if we open them for -non-issues. - -### Don't make breaking changes - -All proposed modifications must be **additive**: new sections, new -examples, new checklists. Never: - -+ Delete an existing rule -+ Change the wording of an existing rule -+ Reorder existing sections -+ Remove URLs or references in the skill - -This keeps the diff vs upstream small and the PR low-risk. - -### Don't burn iteration 1 on the wrong problem - -When baseline scores low, *first* check: is the grader the problem? Look -at the actual `findings.txt` from failed trials. If models *did* identify -the violations but the grader scored them wrong (line numbers off, -keyword mismatch, format variant), fix the grader as iteration 0 (don't -count it against the 2-iteration budget). - -**Source:** auto-pilot supabase + agent-browser pilots — both spent -iteration 1 on grader fixes before reaching skill modification. - ---- - -## Run-record protocol - -Every pilot adds an entry to one of these tables when it discovers -something new. Format: - -```markdown -**[skill-name] (date):** what was new — link to commit. -``` - -### Patterns added by pilots - -+ **manual web-design-guidelines (2026-05-06):** Two-pass workflow + per-element - checklists + 5 BAD/GOOD examples. Lifted 4-case suite from 72% → 86%. -+ **auto-pilot supabase (2026-05-08):** Independently rediscovered two-pass - workflow. Added it to a SQL skill. 0.54 → 0.86. -+ **auto-pilot agent-browser (2026-05-08):** Found that grader was over-strict - for non-interactive ops. Demoted snapshot from required to evidence-only. - Also surfaced "Verify-tool-installed nudge" pattern. -+ **auto-pilot pdf (2026-05-08):** Validated "exit clean on already-good skill" - — no modifications proposed; baseline 1.00. - -### Grader patterns added by pilots - -+ **manual web-design-guidelines (2026-05-06):** ±5-8 line tolerance, hyphen - regex, per-finding-line matching, keyword variants. -+ **auto-pilot supabase (2026-05-08):** "covering" / "does not cover" alternation - pattern. Confirmed ±3 → ±8 line widening is needed by default. - -+ **auto-pilot supabase v2 (2026-05-12):** Upstream constraints required adding a new reference file (`monitor-two-pass-review.md`) instead of editing SKILL.md. Baseline was already 1.00 (calibrated graders from prior run). Pattern: when a re-run starts from calibrated graders, the Phase 3 exit condition fires before Phase 4 — the "modification" step then serves purely as upstream PR packaging rather than eval improvement. - -(Future pilots: append your additions here.) diff --git a/skills/subagents/.gitkeep b/skills/skill-optimizer-subagents/.gitkeep similarity index 100% rename from skills/subagents/.gitkeep rename to skills/skill-optimizer-subagents/.gitkeep From 47124f4d81f07fb7eabd4c8b8671b29f5a179491 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 09:28:56 -0500 Subject: [PATCH 009/121] =?UTF-8?q?feat(skill-optimizer-investigate-functi?= =?UTF-8?q?onality):=20SKILL.md=20body=20=E2=80=94=20classification=20taxo?= =?UTF-8?q?nomy=20+=20local-with-PR-intent=20path?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two changes to the B1 draft, plus a matching spec sync: 1. Classification: replaced the ad-hoc closed list (code-reviewer | document-producer | tool-use | code-patterns | other) with the v3 taxonomy actually used in the prioritization work — tool-use, code-patterns, document, prose-guidance, meta, interactive — and explicitly grants the subagent sovereignty to write a short descriptive label of its own when none fits, rather than collapsing to "other". A specific label gives downstream steps a real handle to work with. 2. Local-skill PR intent: the prior rule was "local = always pr_submission_intent: false". Revised to default false but treat PR-intent as live when the user explicitly says they want to send the local skill back upstream — in which case we capture the upstream guidelines location (URL / CONTRIBUTING.md path / Slack channel) into the report body as a "PR submission notes" subsection that step 3 (investigate-submissions) reads as its starting point. Also: while reviewing, dropped the v1.3/v1.4 framing from the "Why limited-context dispatch matters" section (version is historical metadata, not skill-functional content) and updated subagent path links to the renamed skill-optimizer-subagents/ directory. Spec §"The 7 skills" #1 step 2 + step 6 (classification field description) updated to match. --- docs/skill-optimizer-v1.4-spec.md | 17 +- .../SKILL.md | 158 +++++++++++++++++- 2 files changed, 169 insertions(+), 6 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 441e158..f90a9ae 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -174,15 +174,26 @@ contract), not the prose. report frontmatter as `pr_submission_intent: true|false`. This answer determines whether step 3 (`investigate-submissions`) fires later. - 3. **If local:** no PR question. Set `pr_submission_intent: false` - in the frontmatter. + 3. **If local:** default to `pr_submission_intent: false`. But if + the user explicitly said they want to send this back to an + upstream maintainer (e.g., "I'll fork this back to the original + repo", "I want to PR this to project X"), treat it as PR-intent: + set `pr_submission_intent: true`, ask the user where the + upstream contribution guidelines live (URL, `CONTRIBUTING.md` + path, Slack channel), and record their answer in the report + body under a "PR submission notes" subsection — step 3 uses this + as the starting point. 4. Fetch the skill files (if URL); vendor to `vendored-skill/` for read-only use by downstream steps. 5. Web-search the underlying technology; identify trigger conditions, success criteria, key terminology, intended audience 6. Write structured report covering: what the skill does, who uses it, when it should fire, what tools it depends on, what concepts - the user must understand + the user must understand. The `classification` frontmatter field + takes one of the canonical types (`tool-use`, `code-patterns`, + `document`, `prose-guidance`, `meta`, `interactive`) when one + fits, or a short descriptive label of the subagent's choosing — + a specific label is preferred over a vague `other`. - **Dispatches:** functionality-researcher subagent (limited context: source skill + targeted web fetches) - **Handoff:** "Next, invoke `skill-optimizer-investigate-test-case`. diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 445a314..ba3cd15 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -1,9 +1,161 @@ --- name: skill-optimizer-investigate-functionality -description: Use when the user asks to understand or investigate what a skill does — fetches skill source, web-searches the underlying technology, writes a functionality report. Also asks the user (for upstream skills) whether to target upstream PR submission. +description: Use when the user wants to understand what an existing agent skill does — phrases like "what does this skill do", "investigate this skill", "understand this skill", or when they hand you a URL or local path to a skill they want analyzed. Also triggers at the start of any skill-optimizer chain work, before test design, analysis, or improvement. Use even when the user doesn't explicitly say "investigate" — any phrasing that signals they want to understand a skill before doing anything with it should trigger this. --- # skill-optimizer-investigate-functionality - - +Step 1 of the skill-optimizer chain. Takes a skill (URL or local path), +runs a researcher subagent to figure out what it's supposed to do, and +writes `docs/skill-optimizer//01-functionality.md` — the briefing +document every later step consumes. + +## What you produce + +A single report at `docs/skill-optimizer//01-functionality.md`, +where `` is the source skill's directory name (e.g., +`firecrawl-build-scrape`). + +The report has structured frontmatter: + +```yaml +--- +skill_source: +pr_submission_intent: true | false +classification: +--- +``` + +**`classification`** — pick the most accurate label for what kind of +skill this is. Canonical types: `tool-use` (procedures for using a +specific tool, library, or API), `code-patterns` (code-level patterns +or review checklists), `document` (workflows that produce a document +or file), `prose-guidance` (writing-style or content-creation +guidance), `meta` (skills that operate on other skills or on the +agent's behavior), `interactive` (back-and-forth user dialogue). If +none of these fits cleanly, **write a short descriptive label of your +own** (`dataset-extraction`, `deployment-runbook`, +`ui-mockup-generation`, etc.) rather than falling back to `other` — a +specific label gives downstream steps a real handle to work with. + +For the body template (what sections, what the subagent must cover), +see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). +The subagent itself writes the report — you don't. + +If the source is upstream, you also vendor the fetched skill files to +`vendored-skill/` in the working directory so downstream steps read a +stable copy without re-fetching. + +## Workflow + +### 1. Classify the source + +Is the source an **upstream** skill (a URL or `//` +slug) or a **local** skill (a filesystem path that already exists)? If +the user gave a bare name with no URL and no path, ask them to clarify +before continuing. + +### 2. Ask about PR intent + +**Upstream skills:** ask the user, in roughly these words: + +> Do you want to optimize this skill for upstream PR submission? + +Record the answer in the report frontmatter as +`pr_submission_intent: true` or `pr_submission_intent: false`. **Capture +this decision now**, not at the end of the chain — step 2's handoff +reads this field to decide whether step 3 (`investigate-submissions`) +runs. + +**Local skills:** default to `pr_submission_intent: false`. But if the +user explicitly said they want to send this back to an upstream +maintainer (e.g., "I'll fork this back to the original repo", "I want +to PR this to project X"), treat it as PR-intent. In that case: + +1. Set `pr_submission_intent: true`. +2. Ask the user where the upstream contribution guidelines live — a + URL, a `CONTRIBUTING.md` path, a Slack channel, whatever they have. +3. Record what they say in the report body under a "PR submission + notes" subsection. Step 3 (`investigate-submissions`) uses this as + the starting point for `03-submissions.md`. + +If the user mentions nothing about a PR for the local skill, set +`pr_submission_intent: false` and move on. + +### 3. Vendor the source (upstream only) + +Fetch the skill's files into `vendored-skill/` at the working-directory +root. Downstream steps read from this vendored copy as a stable +reference. Local-source skills don't need vendoring. + +### 4. Determine the slug and the report path + +`` is the source skill's own directory or file name: + +- `firecrawl/skills/firecrawl-build-scrape` → slug is `firecrawl-build-scrape` +- `~/my-skills/pdf-cleanup/SKILL.md` → slug is `pdf-cleanup` + +Report path: `docs/skill-optimizer//01-functionality.md`. + +### 5. Dispatch the functionality-researcher subagent + +**Do NOT do the research yourself in this session.** Dispatch the +functionality-researcher subagent via the `Agent` tool (with worktree +isolation if your environment supports it). Load the prompt template +at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md), +substitute the templated inputs (`${SKILL_SOURCE}`, `${OUTPUT_PATH}`, +`${PR_SUBMISSION_INTENT}`, vendored path), and dispatch. + +The subagent sees: + +- The vendored skill files (or the local skill path) +- Targeted web-search / web-fetch results for the underlying technology +- The output path and frontmatter fields you computed + +The subagent does NOT see: + +- Existing analyses or improvement proposals for this skill +- Prior tests or failure data +- The wider chain's context + +The subagent writes the report itself and returns a brief summary. + +### 6. Confirm and hand off + +When the subagent returns: + +1. Verify the report file exists and the frontmatter parses. +2. Report to the user: the report path, a one-line summary of the + subagent's classification and key findings, then the handoff message. + +Handoff message, verbatim: + +> Next, invoke `skill-optimizer-investigate-test-case`. Note: this +> report's `pr_submission_intent` field tells step 2's handoff whether +> step 3 (`investigate-submissions`) should run. + +## Why limited-context dispatch matters + +When the same context that handles user conversation also conducts the +research, the research drifts toward whatever the user has already +expressed — and every later step in the chain inherits that drift. +Limited-context subagents are the architectural fix: walling the +researcher off from prior analyses and prior failures prevents the +researcher from rationalizing the existing skill's design choices, and +prevents "ducktape" fixes that paper over symptoms instead of addressing +root causes. + +If you find yourself thinking "I'll just write the report myself, the +subagent dispatch is bureaucratic overhead" — that's the failure mode +this chain is built to prevent. Dispatch. + +## Edge cases + +- **URL 404 or local path doesn't exist** — surface the error to the + user; don't try to guess a recovery. +- **Source is a plugin with multiple skills** — ask which one to + investigate; produce one report per skill. +- **`vendored-skill/` already exists from a prior run** — ask whether + to overwrite or reuse the existing copy. +- **User changes their mind on PR intent later** — they re-run this + skill, which overwrites the frontmatter and re-vendors as needed. From 00090668109cafe6b0c72fcf394532397bc856b4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 10:14:15 -0500 Subject: [PATCH 010/121] docs(v1.4): add iteration patterns, 8th auto-pilot skill, subagent for B2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Spec changes: - New ## Iteration patterns section between ## Subagent constraints and ## The skills. Covers: backtrack-trigger table, per-report versioning mechanism (version + inputs frontmatter, with direct- upstream-only staleness checks), re-entry contract with the ${OPERATOR_DIRECTIVES} slot, latest-plus-archive convention, and the transitive-staleness policy (user judgment trusted; auto-pilot re-runs on direct-upstream version mismatch). - ## The 7 skills → ## The skills. Added one-line "Iteration behavior" notes to each of the existing seven subsections. - Step 2 (investigate-test-case) now dispatches a test-case-designer subagent rather than running in the operator session. Reasoning: cross-iteration contamination — on iter 2+ the operator has seen prior failures and biases test selection toward "what just failed" instead of comprehensive coverage. Isolated subagent fixes this. - New ### 8. skill-optimizer-autopilot subsection. Walks 1→7, consults version mechanism per step (skip-or-dispatch), applies defaults for the three human-gate points (B1 PR-intent flag, B2 pick-top-N, B7 log-and-exit on validator-rejected), bounds at max-iterations-per-step. Replaces the old "## Auto-pilot mode" section (which said "no separate skill, just chained invocation"). - Subagent constraints table updated: test-case-designer added; every reasoning subagent's "Does NOT see" column now includes prior drafts of its own output; every Sees column includes ${OPERATOR_DIRECTIVES}. - Acceptance criteria expanded to 9 items (was 7): new #1 covers 8 SKILL.md files, new #5 covers iteration-mechanism E2E, new #8 covers auto-pilot smoke test. Plan changes: - Phase B count 7 → 8 tasks. New Task B8 (skill-optimizer-autopilot SKILL.md) added with its own A1-equivalent dir-creation step inline (the dir wasn't part of Phase A scope). - Phase C count 6 → 7 tasks. New Task C1b (test-case-designer subagent prompt) inserted between C1 and C2 with a full draft template; suffixed "1b" rather than renumbering existing C2-C6 to keep cross-references stable. - Phase C section header adds a preamble explaining the ${OPERATOR_DIRECTIVES} requirement that applies to every reasoning-subagent prompt template. --- docs/skill-optimizer-v1.4-plan.md | 203 ++++++++++++++++++++++- docs/skill-optimizer-v1.4-spec.md | 266 +++++++++++++++++++++++++----- 2 files changed, 427 insertions(+), 42 deletions(-) diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/skill-optimizer-v1.4-plan.md index 08cb9ff..19348e5 100644 --- a/docs/skill-optimizer-v1.4-plan.md +++ b/docs/skill-optimizer-v1.4-plan.md @@ -5,7 +5,7 @@ > | Phase | Mode | Tool | > |---|---|---| > | **A** (structural setup) | subagent-driven-development OK | mechanical Bash/Edit | -> | **B** (the 7 SKILL.md files) | INTERACTIVE with operator | **`skill-creator`** (outer loop: draft, eval, iterate, description-improver) + **`superpowers:writing-skills`** (inner loop: TDD pressure scenarios when a compliance issue surfaces) | +> | **B** (the 8 SKILL.md files — 7 chain skills + auto-pilot) | INTERACTIVE with operator | **`skill-creator`** (outer loop: draft, eval, iterate, description-improver) + **`superpowers:writing-skills`** (inner loop: TDD pressure scenarios when a compliance issue surfaces) | > | **C** (subagent prompt templates) | INTERACTIVE with operator | **`superpowers:writing-skills`** (TDD pressure scenarios for limited-context constraints — the prompts ARE compliance documents) | > | **D** (`workflow.md`) | INTERACTIVE with operator | No skill tool — direct authoring + user review per section | > | **E** (end-to-end validation) | manual operator runs (real eval $) | direct `Skill` invocations + `run-suite` | @@ -288,7 +288,7 @@ git commit -m "chore(v1.4): create skill-optimizer-subagents/ dir (populated in --- -## Phase B — Interactive SKILL.md creation (7 tasks, INTERACTIVE via skill-creator + writing-skills) +## Phase B — Interactive SKILL.md creation (8 tasks, INTERACTIVE via skill-creator + writing-skills) **Mode:** Each Task B is performed interactively in the operator's CC session. @@ -643,7 +643,84 @@ git commit -m "feat(skill-optimizer-improve-skill): SKILL.md body via skill-crea --- -## Phase C — Subagent prompt templates (6 tasks, INTERACTIVE via writing-skills) +### Task B8: SKILL.md for `skill-optimizer-autopilot` + +**Files:** + +- Create: `skills/skill-optimizer-autopilot/` directory + `SKILL.md` + (was not part of Phase A; B8 creates the dir too) +- Modify: nothing else + +This is the 8th skill — the auto-pilot driver. It walks 1→7, uses +the version mechanism to skip current reports and re-run stale ones, +and applies default policies for the three human-gate points (PR +intent at B1, picked subset at B2, validator-rejected at B7). + +- [ ] **Step 1: Create the skill directory + frontmatter shell** + +```bash +mkdir -p skills/skill-optimizer-autopilot +cat > skills/skill-optimizer-autopilot/SKILL.md <<'EOF' +--- +name: skill-optimizer-autopilot +description: Use when the user wants to run the full skill-optimizer chain on a skill end-to-end without stepping through manually — phrases like "auto-pilot this skill", "run the whole chain on X", "skill-optimizer end-to-end for X", "automated improvement run for X". Walks steps 1 through 7 using the iteration version mechanism to skip current reports and re-run stale ones, with default policies for the three human-gate points. Best for batch processing where modest results are acceptable; for high-stakes single-target work, drive the steps manually. +--- + +# skill-optimizer-autopilot + + + +EOF +``` + +- [ ] **Step 2: Invoke skill-creator with the spec excerpt as the brief** + +(Same workflow as Task B1 Step 1 — see there.) Brief: + +```text +Goal: write the body of skills/skill-optimizer-autopilot/SKILL.md. + +Interface contract: see docs/skill-optimizer-v1.4-spec.md §"The skills" #8. + +Key points to cover in the body: +- This is the 8th skill — the auto-pilot driver, not a chain step. +- It walks 1→7 in order, dispatching each chain skill as a subagent. +- For each step, it consults the version mechanism (see spec §"Iteration + patterns" → "Versioning") to decide skip-vs-dispatch. +- Default policies for the three human-gate points: + * B1 PR-intent question → use --pr-intent flag, default false + * B2 user-picks-subset gate → take top N by importance (default N=5, + configurable via --pick-top-n) + * B7 validator-rejected → log final state, no further automated retries +- Iteration cap: --max-iterations-per-step (default 2) +- Output: docs/skill-optimizer//autopilot-summary-.md with + per-step final version + headline result + any blockers +- Caveats: expect modest results vs operator-driven runs; best for batch. + +Load-bearing constraint: auto-pilot dispatches the chain skills via the +Skill tool (treating them as standard subagents); it does NOT do their +work itself in the operator session. This is the same anti-ducktape +discipline as the rest of the chain — applied at the meta level. +``` + +- [ ] **Step 3: Verify + commit** + +```bash +git add skills/skill-optimizer-autopilot/SKILL.md +git commit -m "feat(skill-optimizer-autopilot): SKILL.md body via skill-creator" +``` + +--- + +## Phase C — Subagent prompt templates (7 tasks, INTERACTIVE via writing-skills) + +**Every reasoning subagent prompt must include the +`${OPERATOR_DIRECTIVES}` templated slot** — an atomic list of new +requirements from cross-iteration learnings ("user wants null-input +edge cases", "focus on the gpt-5 cluster"), pre-digested by the +operator session into specific asks. The slot defaults to empty on +iteration 1 and never carries a context dump of prior outputs. See +spec §"Iteration patterns" → "Re-entry contract" for the rationale. Each subagent prompt template is loaded by the corresponding skill (Phase B), substituted with templated inputs, and dispatched via Agent tool. All must enforce the limited-context constraints from the spec's "Subagent constraints" table. @@ -785,6 +862,126 @@ git commit -m "feat(v1.4-subagents): research-functionality prompt template" --- +### Task C1b: `subagents/test-case-designer.md` (new — was operator-session in initial spec) + +**Files:** Create `skills/skill-optimizer-subagents/test-case-designer.md`. + +**Why this exists:** the initial spec ran step 2 (test-case design) in +the operator session, reasoning "it's analytical, no research bias to +worry about." That overlooked cross-iteration contamination — on +iteration 2+, the operator session has already absorbed prior failure +data, optimizer attempts, and validator verdicts, and biases test +selection toward "tests that would have caught the things I just +watched fail" instead of "tests that comprehensively cover the +skill's responsibilities." This subagent restores coverage-oriented +test design by running in isolation. + +- [ ] **Step 1: Invoke writing-skills with this draft as the starting brief** + +(Same process as C1.) Draft template: + +````markdown +# Sub-subagent: design test cases for a skill + +You are dispatched to enumerate a skill's responsibilities and design a +ranked list of test cases. Produce a structured `02-test-case.md` report. + +## Inputs (templated) + +- `${FUNCTIONALITY_PATH}` — path to the latest `01-functionality.md` +- `${OUTPUT_PATH}` — typically `docs/skill-optimizer//02-test-case.md` +- `${OPERATOR_DIRECTIVES}` — short bulleted list of atomic new + requirements from prior iterations (e.g., "user wants null-input + edge cases"). Default empty. +- `${PRIOR_VERSION}` — version number for the new report (1 on first + invocation; N+1 if archiving an existing v(N)) + +## Tools allowed + +- Read (`${FUNCTIONALITY_PATH}` only) +- Write (`${OUTPUT_PATH}`) + +## Tools NOT allowed (limited-context constraint) + +You operate under strict limited context. You see ONLY the inputs +above. You do NOT see and MUST NOT attempt to read: + +- Prior `02-test-case.md` drafts (anything under `archive/`) +- `06-analysis.md` or any analysis from prior iterations +- `07-improvement-proposal.md`, `07-validator-verdict.md` +- Raw failure data, `findings.txt`, bench results + +Coverage design must reason from the skill's stated responsibilities +in `01-functionality.md`, augmented only by `${OPERATOR_DIRECTIVES}`. +If you find yourself wanting to read prior failure data to "design +better tests this time" — that's the failure mode this constraint +prevents. Stay in your lane. + +## What to produce + +Write `${OUTPUT_PATH}` with this structure: + +```markdown +--- +version: ${PRIOR_VERSION} +inputs: + step_1_functionality: +--- + +# Test cases for + +## Proposed cases + +For each responsibility identified in 01-functionality.md, design 1–2 +test cases. Each entry: + +### + +- **Tests:** +- **Setup:** +- **Expected agent behavior:** +- **Grader spec:** +- **Why it matters:** +- **Importance:** +- **Cost-to-build:** + +## Ranking + +Ordered list from highest-importance / lowest-cost to lowest- +importance / highest-cost. The operator will use this for the +user-picks-subset gate. +``` + +## Commit + +```bash +git add ${OUTPUT_PATH} +git commit -m "docs(skill-optimizer): test-case proposal v${PRIOR_VERSION} for " +``` + +## Return + +Under 200 words. Cover: total count of proposed cases, your top-3 by +importance, any responsibilities you found that the functionality +report didn't enumerate (flag for the operator to verify), any +cases you flagged as needs-real-tooling that the operator should +sanity-check before committing to. +```` + +- [ ] **Step 2: Verify + commit** + +```bash +grep -E '\$\{(FUNCTIONALITY_PATH|OUTPUT_PATH|OPERATOR_DIRECTIVES|PRIOR_VERSION)\}' skills/skill-optimizer-subagents/test-case-designer.md | wc -l +# Expected: ≥ 6 + +git add skills/skill-optimizer-subagents/test-case-designer.md +git commit -m "feat(v1.4-subagents): test-case-designer prompt template" +``` + +--- + ### Task C2: `subagents/research-submissions.md` **Files:** Create `skills/skill-optimizer-subagents/research-submissions.md`. diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index f90a9ae..45ac172 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -141,24 +141,124 @@ and dispatches the subagent with only the narrow chunks it needs. | Subagent | Sees | Does NOT see | Why | |---|---|---|---| -| Functionality researcher (step 1) | Source skill files, web-search results | Existing analyses, existing tests | Pure research, no contamination | -| Submission researcher (step 3) | Repo files, gh-API outputs | Anything about the proposed change | Just upstream facts | -| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` | The skill's content, other test cases, the eval grader's matching logic | Prevents grader-hacking; prevents copying existing tests | -| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files) | Forces it to think about the SKILL, not the SOLUTIONS | -| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content | Raw failed trials, `findings.txt`, grader internals, test inputs | Forces principled improvement, not pattern-match patches | -| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs | Independent check; can't be biased by what the optimizer told itself | +| Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, prior `01-functionality.md` drafts | Pure research, no contamination across iterations | +| Test-case designer (step 2) | `01-functionality.md` (latest), `${OPERATOR_DIRECTIVES}` | Prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Coverage design must reason from the skill's responsibilities, not from "what just failed" | +| Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change | Just upstream facts | +| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` + `${OPERATOR_DIRECTIVES}` | The skill's content, other test cases, the eval grader's matching logic | Prevents grader-hacking; prevents copying existing tests | +| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), prior `06-analysis.md` drafts | Forces it to think about the SKILL, not the SOLUTIONS; iteration-isolated | +| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, prior optimizer attempts | Forces principled improvement, not pattern-match patches; no attachment to prior failed attempts | +| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, prior validator verdicts | Independent check; can't be biased by what the optimizer (or a prior validator round) told itself | The skill (operator session) sees everything; subagents see slices. This is the architectural fix for the "tunnel-vision into ducktape" problem. -## The 7 skills +## Iteration patterns -Each skill is a directory with `SKILL.md` (the instructions) plus -optional supporting files (subagent prompts, reference material). The -prose content of each SKILL.md will be filled in during implementation -— this spec defines the **interface** (input/output/behavior -contract), not the prose. +The chain isn't strictly linear in practice. Any step can be re-run +(operator-driven or auto-pilot-driven), and re-running a step +invalidates downstream reports that derived from its prior version. +This section defines how iteration is captured, detected, and +cascaded — using a lightweight per-report version mechanism that +keeps state inspection deterministic without requiring a global +iteration counter. + +### When iteration happens + +Common backtrack triggers: + +| At step | Common trigger | Backtracks to | +|---|---|---| +| After 2 (test design) | "Coverage is wrong; the skill does more than this" | 1 (research) or 2 (revise with directives) | +| After 4 (write-tests) | Test writer couldn't implement a case | 2 (revise the set) | +| After 5 (bench) | All tests pass on baseline — too easy | 2 or 4 (harder tests) | +| After 6 (analyze) | "No clear weakness" but user disagrees | 2 (better coverage), 4 (better tests), or 6 with directives | +| After 7 (improve) | Validator rejects beyond optimizer's own loop | 6 (re-analyze) or 2 (the test was wrong) | +| Anywhere | User dislikes the output | re-run that step with new directives | + +There's no rigid backtrack flowchart — operator and auto-pilot both +decide based on the named weakness in `06-analysis.md`, the validator +verdict, and the version mechanism below. + +### Versioning + +Every report carries two frontmatter fields: + +```yaml +version: 1 # this report's own iteration counter +inputs: + step_1_functionality: 1 # versions of upstream reports I derived from + step_2_test_case: 1 # (only those this step actually read) +``` + +On invocation, a step compares the current versions of its upstream +inputs against what its own prior report (if any) recorded under +`inputs`: + +- **No existing report** → write a new one at `version: 1` +- **Existing report, all input versions match** → already current; do + nothing (unless the operator passed directives requesting a re-run + anyway) +- **Existing report, any input version mismatch** → existing is stale + → archive it → re-run with bumped own-version + +Versions are per-step independent counters. Step 1 going from v1 → v2 +doesn't change step 2's version until step 2 is re-invoked and +detects the mismatch. + +Pre-N files are untouched when step N is re-run; only step N's own +version bumps. + +### Re-entry contract + +Every reasoning step accepts an `${OPERATOR_DIRECTIVES}` templated +slot in its subagent prompt — atomic new requirements pre-digested +from prior iterations ("user requested coverage for null inputs", +"user flagged responsibility X as under-tested"). Default empty on +first invocation. The slot is **never a context dump** of prior +output; it's a short bulleted list of new requirements only. + +Subagent behavior under iteration: + +- Subagent **ignores its own prior output** — no peeking at `archive/`. +- Subagent reads upstream reports at their **latest** versions only. +- Operator session archives the prior version before dispatching, and + records the upstream versions consumed in the new report's `inputs` + frontmatter. + +### Archive convention + +The latest version of each report lives at its canonical path: +`docs/skill-optimizer//-.md`. When overwritten, the +prior version moves to +`docs/skill-optimizer//archive/--v.md` (where `N` +is the version being archived). History is browsable for human audit +but inert — auto-pilot and chain skills only read latest. + +### Transitive staleness + +Each step checks **direct upstream only**. If step 1 goes to v2 but +step 2 isn't re-run (user judged it still valid), step 3 sees step 2 +at v1 and treats it as current — even though step 2's report +references step 1 v1. + +The trade-off: transitive issues surface as confusing downstream +results, not silent corruption. The discipline is that skipping a +step's re-run is an explicit judgment call ("the old report is still +valid against the new upstream"), and we trust that call. Auto-pilot +doesn't make this judgment — it re-runs any step whose direct +upstream version doesn't match. + +## The skills + +Seven chain skills (steps 1–7) plus an auto-pilot driver (skill 8). +Each chain skill is a directory with `SKILL.md` (the instructions) +plus optional supporting files (subagent prompts, reference material). +The prose content of each SKILL.md will be filled in during +implementation — this spec defines the **interface** +(input/output/behavior contract), not the prose. + +Per-step iteration behavior is noted at the end of each subsection. ### 1. `skill-optimizer-investigate-functionality` @@ -199,6 +299,10 @@ contract), not the prose. - **Handoff:** "Next, invoke `skill-optimizer-investigate-test-case`. Note: this report's `pr_submission_intent` field tells step 2's handoff whether step 3 (`investigate-submissions`) should run." +- **Iteration behavior:** re-run when source skill changes or when + the operator wants fresh research with new directives. Re-running + bumps `01-functionality.md` version; downstream steps detect the + change and cascade on next invocation. ### 2. `skill-optimizer-investigate-test-case` @@ -213,14 +317,31 @@ contract), not the prose. cases (boundary conditions, error paths, common pitfalls); estimate grader difficulty (deterministic check vs needs-real-tooling); rank by importance + cost-to-build. -- **Dispatches:** none — analytical, operator session -- **User gate:** prompts the user to pick which subset of proposed - cases to actually build (write-tests acts on the picked subset) +- **Dispatches:** test-case-designer subagent (limited context: only + `01-functionality.md` + the `${OPERATOR_DIRECTIVES}` slot; does NOT + see prior `02-test-case.md` drafts, prior `06-analysis.md`, + optimizer attempts, or failure data — keeps coverage design + honest across iterations, where the operator session has already + absorbed prior failures and would otherwise bias tests toward + "what just failed" instead of "what comprehensively covers the + skill's responsibilities"). +- **User gate:** after the subagent returns, the operator session + prompts the user to pick which subset of proposed cases to + actually build (write-tests acts on the picked subset). The picks + are recorded in `02-test-case.md` frontmatter (e.g., + `picked: [case-name-1, case-name-3]`). - **Handoff:** read `01-functionality.md`'s `pr_submission_intent` field. If `true` → "Invoke `skill-optimizer-investigate-submissions` next, then `skill-optimizer-write-tests`." If `false` → "Skip step 3; invoke `skill-optimizer-write-tests` with the user's picked subset." No late prompts — the decision was made at step 1. +- **Iteration behavior:** re-run when the user wants different + coverage, when step 4/5/6 surface a coverage gap, or when step 1's + version bumps. `${OPERATOR_DIRECTIVES}` is the channel for + "user wants null-input edge cases" or similar atomic new + requirements; the subagent treats these as fresh requirements on + top of `01-functionality.md`, never as a context dump of prior + iterations. ### 3. `skill-optimizer-investigate-submissions` (OPTIONAL — upstream only) @@ -241,6 +362,10 @@ contract), not the prose. question at step 1) - **Handoff:** "Validator in `improve-skill` will read this for external consistency check. Continue with `write-tests` if not done." +- **Iteration behavior:** rarely needs re-running — upstream PR + conventions change slowly. Re-run when the upstream repo's + CONTRIBUTING/CLA changes or when `01-functionality.md` version + bumps (in case the skill source URL itself changed). ### 4. `skill-optimizer-write-tests` @@ -261,6 +386,12 @@ contract), not the prose. - **Parallelizable:** each case independent; dispatch in a single message - **Handoff:** "Invoke `skill-optimizer-run-bench` to measure baseline." +- **Iteration behavior:** re-run when `02-test-case.md` version + bumps (picked subset changed), when the test-writer flagged + unimplementable cases, or when bench results suggest tests are + systematically too easy or too narrow. Each test-writer subagent + also accepts `${OPERATOR_DIRECTIVES}` for case-level revision + hints. ### 5. `skill-optimizer-run-bench` @@ -275,6 +406,12 @@ contract), not the prose. - **Note:** this is intentionally thin; the entire v1.3 run-suite logic stays as-is in the CLI - **Handoff:** "Invoke `skill-optimizer-analyze-result`." +- **Iteration behavior:** re-run when `workbench/` version bumps + (new or revised tests) or when the operator wants fresh results + against the same workbench. Each bench result is timestamped under + `05-bench-results//` so old runs are never overwritten — but + `05-bench-summary.md` (the versioned summary that analyze-result + reads) is overwritten with archive on re-run. ### 6. `skill-optimizer-analyze-result` @@ -297,6 +434,13 @@ contract), not the prose. - **Handoff:** if at least one structural weakness identified → "Invoke `skill-optimizer-improve-skill`." Otherwise → exit honestly ("no structural weakness; no improvement warranted"). +- **Iteration behavior:** re-run when bench results change or when + the operator/user wants a fresh look with new framing. The + `${OPERATOR_DIRECTIVES}` slot is the channel for hints like + "focus on the gpt-5 cluster" or "the user thinks weakness X is + actually two separate issues" — the subagent treats these as + additional analytical lenses, not as a context dump of prior + conclusions. **`06-analysis.md` format (Option A — structured):** @@ -366,22 +510,57 @@ contract), not the prose. exists because step 3 ran). The draft includes the diff, the body, caveats, and operator-steps-to-submit. No late "submit a PR?" prompt — the decision was already made at step 1. - -## Auto-pilot mode - -There is no separate auto-pilot skill. **Auto-pilot is just chained -manual invocation.** The user (or the agent acting on user's request) -walks through the skills in order: 1 → 2 → (optional 3) → 4 → 5 → 6 → -7. Each skill's SKILL.md ends with a "Next, invoke X" instruction; -the agent uses the `Skill` tool to chain. - -For full auto-pilot, the user just says "optimize skill X end-to-end" -and the agent invokes all 7 in order, gating at the natural -user-review points (after step 2 the user picks tests; after step 7 -the user reviews the improvement proposal). - -The user-review gates are part of each skill's instructions, not -external orchestration. +- **Iteration behavior:** re-run when `06-analysis.md` version bumps + (new analysis = potentially different weakness) or when the + operator wants a fresh optimization attempt. The internal + optimizer/validator loop (max 2 rounds, baked into the skill) + handles in-step iteration; cross-step backtracking is the + operator's call. `${OPERATOR_DIRECTIVES}` is the channel for hints + like "prefer additive changes" or "don't touch the description + field" — the optimizer treats these as additional constraints, not + as a context dump of prior attempts. + +### 8. `skill-optimizer-autopilot` + +- **Description trigger:** "auto-pilot this skill", "run the whole + chain on X", "skill-optimizer end-to-end for X", "automated + improvement run for X" +- **Input:** same as step 1 (source skill — URL or local path), plus + optional flags: `pr_intent`, `max_iterations_per_step`, + `pick_top_n` (for B2's user-picks gate) +- **Output:** an end-of-run summary report at + `docs/skill-optimizer//autopilot-summary-.md` listing + each step's final version, headline result, and any blockers +- **Behavior:** + 1. Walks 1→7 in order, dispatching each chain skill. + 2. For each step: check whether the existing report is current + against its declared `inputs` versions. If current, skip; if + missing or stale, dispatch. + 3. Handles the three human-gate points with default policies: + - **B1 PR-intent question** (upstream skills) → uses the + `pr_intent` flag value (default `false`) + - **B2 user-picks-subset gate** → takes the top-N by importance + (default `pick_top_n: 5`) + - **B7 validator-rejected verdict** → logs the final state, no + further automated retries; surfaces blocker in the summary + 4. Bounds iterations: at most `max_iterations_per_step` re-runs + per step (default 2) to cap runaway loops. + 5. Surfaces blockers (subagent BLOCKED status, validator + unresolvable, missing inputs) in the summary rather than + halting the whole run. +- **Dispatches:** the seven chain skills as ordinary subagents (each + chain skill internally dispatches its own narrow-context + subagents). +- **Caveats** (baked into SKILL.md): expect modest results compared + to operator-driven runs. The seven steps are hard even with human + judgment; auto-pilot is best for batch processing where some + failures are acceptable, not for high-stakes single-target + optimization. +- **Iteration behavior:** auto-pilot is itself iterable. Re-running + picks up at whatever step is stale per the version mechanism; + steps that are current are skipped. The summary report is + timestamped per run rather than versioned — each auto-pilot run + produces a fresh summary so the audit trail is preserved. ## Plugin packaging @@ -421,25 +600,34 @@ v1.4 chain its role is filled by `skill-optimizer-run-bench`. For v1.4 to be considered done: -1. All 7 SKILL.md files exist with the interface described in this - spec, plus appropriate "Next, invoke X" handoff instructions -2. Subagent prompt templates exist in `skills/skill-optimizer-subagents/` with the - limited-context constraints described in the "Subagent constraints" - table +1. All 7 chain SKILL.md files + the 8th auto-pilot SKILL.md exist + with the interface described in this spec, plus appropriate + "Next, invoke X" handoff instructions on the chain skills +2. Subagent prompt templates exist in + `skills/skill-optimizer-subagents/` with the limited-context + constraints from the "Subagent constraints" table, and each + reasoning subagent accepts an `${OPERATOR_DIRECTIVES}` slot 3. `docs/skill-optimizer-workflow.md` documents the chain visually 4. ~~`references/recipes.md` is seeded from v1.3's `lessons.md`~~ — deferred; the analyzer / optimizer subagents operate without a pre-loaded recipe library until real end-to-end observations surface curated patterns -5. **End-to-end test on a local skill** (e.g., one of the existing +5. **Iteration mechanism works end-to-end** — re-running a step on + an existing slug correctly archives the prior report, bumps the + `version` field, and the next downstream step on next invocation + detects the version mismatch and re-runs itself +6. **End-to-end test on a local skill** (e.g., one of the existing `skills/skill-optimizer/SKILL.md` or a small new skill in this repo) — walks 1→2→4→5→6→7, produces all expected reports, modifies the target skill, validator approves -6. **End-to-end test on an upstream skill** — re-run firecrawl (the +7. **End-to-end test on an upstream skill** — re-run firecrawl (the v1.3 regression case) under v1.4; expected: optimizer either produces a principled fix OR honestly refuses (no regression shipped) -7. Plugin manifest unchanged in shape (still +8. **Auto-pilot smoke test** — running the auto-pilot on a local + skill end-to-end produces a `autopilot-summary-.md` covering + all eight steps with their final versions and any blockers +9. Plugin manifest unchanged in shape (still `.claude-plugin/plugin.json`); skills auto-discoverable via the existing plugin loading mechanism From b349056494b4da19878ad0e15921d6edf64ef016 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 10:18:02 -0500 Subject: [PATCH 011/121] feat(skill-optimizer-investigate-functionality): SKILL.md aligned with iteration patterns MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Five updates to match the spec's new iteration mechanism (§"Iteration patterns" added in commit 0009066): 1. Report frontmatter now includes `version` field. Step 1 has no upstream reports, so `inputs:` is omitted; downstream steps will add it as they're authored. 2. New workflow step 5 — "Handle iteration: archive prior version + collect directives." Checks for existing report at the canonical path; if present, moves it to archive/01-functionality-v.md and bumps version. Also where the operator pre-digests cross-iteration learnings into the OPERATOR_DIRECTIVES bulleted list (atomic new requirements, not a context dump). 3. Step 6 (renumbered from old step 5) dispatches with two new templated inputs: ${VERSION} and ${OPERATOR_DIRECTIVES}. The subagent's "does NOT see" list explicitly includes prior 01-functionality.md drafts under archive/ to prevent the re-derivation from being contaminated by its own past output. 4. New "## Iteration behavior" section after "## Edge cases". Covers: when to re-run, what happens to vendored-skill/ on re-run, the archive convention, and how downstream cascade self-corrects via the version-mismatch check on next invocation. 5. Edge-case for "vendored-skill/ already exists" tightened: default to reuse (same source), re-fetch only on source URL change or explicit user ask. (Previously asked the user every time.) Sets the template for B2-B8. --- .../SKILL.md | 78 +++++++++++++++++-- 1 file changed, 71 insertions(+), 7 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index ba3cd15..d007e57 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -20,12 +20,19 @@ The report has structured frontmatter: ```yaml --- +version: 1 skill_source: pr_submission_intent: true | false classification: --- ``` +**`version`** — per-report iteration counter. `1` on the first run +for this slug; bumps on each re-run after the prior version is +archived. Step 1 has no upstream reports, so the `inputs:` field +(used by downstream steps to record what versions they consumed) is +omitted here. + **`classification`** — pick the most accurate label for what kind of skill this is. Canonical types: `tool-use` (procedures for using a specific tool, library, or API), `code-patterns` (code-level patterns @@ -97,30 +104,56 @@ reference. Local-source skills don't need vendoring. Report path: `docs/skill-optimizer//01-functionality.md`. -### 5. Dispatch the functionality-researcher subagent +### 5. Handle iteration: archive prior version + collect directives + +Check whether the report path already exists. + +**If it does NOT exist:** this is iteration 1. New report's +`version: 1`. No directives to collect unless the user supplied any +explicitly. + +**If it DOES exist:** this is a re-run. Read the existing report's +`version` field. Move the existing file to +`docs/skill-optimizer//archive/01-functionality-v.md` (where +`` is the existing version). The new report's version is ``. + +Then collect operator directives — atomic new requirements that +emerged from prior iterations or from the user's current request. +Examples: "the user said the prior report missed the skill's vendor +dependency on Anthropic", "the user wants more depth on +who-uses-this." Keep them as a short bulleted list of specific asks, +NOT a context dump of the prior report's content. If there are no +new requirements, leave directives empty. + +### 6. Dispatch the functionality-researcher subagent **Do NOT do the research yourself in this session.** Dispatch the functionality-researcher subagent via the `Agent` tool (with worktree isolation if your environment supports it). Load the prompt template at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md), substitute the templated inputs (`${SKILL_SOURCE}`, `${OUTPUT_PATH}`, -`${PR_SUBMISSION_INTENT}`, vendored path), and dispatch. +`${PR_SUBMISSION_INTENT}`, `${VERSION}` from step 5, +`${OPERATOR_DIRECTIVES}` from step 5, vendored path), and dispatch. The subagent sees: - The vendored skill files (or the local skill path) - Targeted web-search / web-fetch results for the underlying technology -- The output path and frontmatter fields you computed +- The output path, the version number, and the frontmatter fields you + computed +- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new + requirements (may be empty) The subagent does NOT see: +- Prior `01-functionality.md` drafts (the ones now in `archive/`) - Existing analyses or improvement proposals for this skill - Prior tests or failure data - The wider chain's context The subagent writes the report itself and returns a brief summary. -### 6. Confirm and hand off +### 7. Confirm and hand off When the subagent returns: @@ -155,7 +188,38 @@ this chain is built to prevent. Dispatch. user; don't try to guess a recovery. - **Source is a plugin with multiple skills** — ask which one to investigate; produce one report per skill. -- **`vendored-skill/` already exists from a prior run** — ask whether - to overwrite or reuse the existing copy. +- **`vendored-skill/` already exists from a prior run** — reuse the + existing copy by default (it's the same source). Re-fetch only if + the source URL itself changed or the user explicitly asks. - **User changes their mind on PR intent later** — they re-run this - skill, which overwrites the frontmatter and re-vendors as needed. + skill (see iteration section below); the prior report is archived + and the new one captures the updated PR intent. + +## Iteration behavior + +This skill is re-runnable. Re-run when: + +- The source skill itself has changed upstream (and you want fresh + research) +- The user supplied new directives the prior subagent didn't have + (e.g., "the report missed responsibility X" or "go deeper on the + vendor dependency") +- The user changed their mind on PR intent +- An upstream step's version bumped and you want this report + re-derived against current inputs (rare for step 1 since it has no + upstream reports — but possible if the user reframes the source) + +On re-run, step 5 archives the prior report under `archive/` and +bumps the version. The functionality-researcher subagent never sees +its own prior output — it re-derives the report from +`${SKILL_SOURCE}` + targeted web research + your +`${OPERATOR_DIRECTIVES}` only. Don't try to "help" by passing the +prior report's content as context; the whole point of the version +mechanism is to keep the new derivation honest. + +When this report's version bumps, downstream reports (`02-test-case.md` +and onward) become stale on their next invocation — their `inputs` +frontmatter no longer matches the current `01-functionality.md` +version, so they'll archive-and-rerun themselves on next dispatch. +You don't need to cascade manually; the next step detects the +mismatch on its own invocation. From c82776ffaa1f67bb5fd7b2d3ecb987283ca7b2c8 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 10:36:43 -0500 Subject: [PATCH 012/121] feat(v1.4-shared): factor iteration mechanics into shared protocol doc MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Authoring the same ~40 lines of archive + version-bump + directives- collection logic into seven chain skills (B1-B7) would mean ~280 duplicated lines. Factor the mechanics into a single shared operational reference that every chain skill loads explicitly. New file: skills/skill-optimizer-shared/iteration-protocol.md (187 lines). Covers: versioning convention, the per-invocation decision tree (no existing report → v1; existing + inputs match → current; existing + mismatch → archive + bump), collecting OPERATOR_DIRECTIVES (atomic new requirements, not context dumps), subagent constraints under iteration (ignore own prior output, read upstream at latest), cascading staleness policy (direct-upstream-only checks), archive table, bootstrapping case, and what the protocol explicitly does NOT cover (bench-results timestamping, auto-pilot summaries, vendored-skill cache). B1 changes: - Step 5 stripped from ~25 lines of inline mechanics to ~10 lines with an explicit "Read this file now" instruction pointing at the protocol doc. The "Read this now" framing is load-bearing — the whole point is that the agent must consult the protocol, not improvise. - "Iteration behavior" section trimmed from ~25 lines of general mechanics to ~10 lines listing only step-1-specific re-run triggers, with a pointer back to the protocol for the general mechanics. - Net: 225 → 208 lines. This sets the pattern for B2-B8: each chain skill's "Handle iteration" step will be a brief pointer-and-step-specific-notes combo, with the heavy mechanics centralized. Spec changes: - Architecture overview file tree: skill-optimizer-shared/ added. - Acceptance criterion 2b added: shared protocol doc exists and is referenced explicitly from each chain skill's iteration step. Plan changes: - File tree updated (skill-optimizer-shared/ + autopilot/ stub). - New Task A5 added: author iteration-protocol.md. Documented as "as-executed (added mid-execution)" because it emerged from B1's revision rather than the initial plan. --- docs/skill-optimizer-v1.4-plan.md | 52 +++++ docs/skill-optimizer-v1.4-spec.md | 7 + .../SKILL.md | 65 +++--- .../iteration-protocol.md | 187 ++++++++++++++++++ 4 files changed, 270 insertions(+), 41 deletions(-) create mode 100644 skills/skill-optimizer-shared/iteration-protocol.md diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/skill-optimizer-v1.4-plan.md index 19348e5..fceda9b 100644 --- a/docs/skill-optimizer-v1.4-plan.md +++ b/docs/skill-optimizer-v1.4-plan.md @@ -60,13 +60,18 @@ skills/ skill-optimizer-run-bench/SKILL.md # Phase B (interactive) skill-optimizer-analyze-result/SKILL.md # Phase B (interactive) skill-optimizer-improve-skill/SKILL.md # Phase B (interactive) + skill-optimizer-autopilot/SKILL.md # Phase B8 (interactive) skill-optimizer-subagents/ # Phase C (interactive) research-functionality.md + test-case-designer.md research-submissions.md test-writer.md analyzer.md optimizer.md validator.md + skill-optimizer-shared/ # Phase A5 (mechanical authoring) + iteration-protocol.md # operational reference all + # chain skills load docs/ skill-optimizer-v1.4-spec.md # already committed @@ -288,6 +293,53 @@ git commit -m "chore(v1.4): create skill-optimizer-subagents/ dir (populated in --- +### Task A5: Author `skills/skill-optimizer-shared/iteration-protocol.md` + +**Files:** + +- Create: `skills/skill-optimizer-shared/iteration-protocol.md` + +> **As-executed (added mid-execution):** the initial plan didn't +> include this file. It emerged during B1's revision: each chain +> skill needs ~40 lines of identical iteration scaffolding (archive, +> version bump, directives collection), and inlining that into +> each of the 7 SKILL.md files would mean ~280 duplicated lines. +> Factoring the mechanics into one shared operational reference +> keeps each SKILL.md focused on this-skill's specifics, and the +> protocol stays consistent across the chain. +> +> Each chain skill's "Handle iteration" step (around the middle of +> its workflow) explicitly instructs the agent to **read this file** +> before proceeding — the load-bearing-ness is the whole point, so +> the reference is mandatory, not optional. + +- [ ] **Step 1: Write the iteration protocol** + +Direct authoring (no skill tool — it's a mechanical reference doc, +not a skill or a compliance prompt template). The doc covers: + +- Versioning convention (`version` + `inputs` frontmatter) +- The decision tree applied on every chain-skill invocation +- Collecting `${OPERATOR_DIRECTIVES}` (atomic new requirements, not + context dumps) +- Subagent constraints under iteration (ignore own prior output, + read upstream at latest) +- Cascading staleness (direct-upstream-only checks; user judgment + trusted on skips) +- Archive convention table +- The bootstrapping case (fresh slug, no prior dir) +- What this protocol does NOT cover (bench-results timestamping, + auto-pilot summaries, vendored-skill cache) + +- [ ] **Step 2: Verify + commit** + +```bash +git add skills/skill-optimizer-shared/iteration-protocol.md +git commit -m "feat(v1.4-shared): iteration-protocol.md — shared mechanics for chain skills" +``` + +--- + ## Phase B — Interactive SKILL.md creation (8 tasks, INTERACTIVE via skill-creator + writing-skills) **Mode:** Each Task B is performed interactively in the operator's CC session. diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 45ac172..acfbe37 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -74,11 +74,15 @@ skills/ skill-optimizer/SKILL.md # existing — unchanged skill-optimizer-subagents/ research-functionality.md + test-case-designer.md research-submissions.md test-writer.md analyzer.md optimizer.md validator.md + skill-optimizer-shared/ + iteration-protocol.md # mechanical reference + # every chain skill loads ``` Companion docs live under `docs/`, not `skills/`, because they're @@ -607,6 +611,9 @@ For v1.4 to be considered done: `skills/skill-optimizer-subagents/` with the limited-context constraints from the "Subagent constraints" table, and each reasoning subagent accepts an `${OPERATOR_DIRECTIVES}` slot +2b. `skills/skill-optimizer-shared/iteration-protocol.md` exists and + is referenced explicitly (with a "Read this now" instruction) + from each chain skill's "Handle iteration" step 3. `docs/skill-optimizer-workflow.md` documents the chain visually 4. ~~`references/recipes.md` is seeded from v1.3's `lessons.md`~~ — deferred; the analyzer / optimizer subagents operate without a diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index d007e57..e1e3ca9 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -104,26 +104,24 @@ reference. Local-source skills don't need vendoring. Report path: `docs/skill-optimizer//01-functionality.md`. -### 5. Handle iteration: archive prior version + collect directives +### 5. Handle iteration -Check whether the report path already exists. +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it for this step. The protocol is mechanical and identical +across all chain skills — it tells you how to check for an existing +`01-functionality.md`, archive it if present, compute the new +version, and collect operator directives. Do not improvise; read the +file. -**If it does NOT exist:** this is iteration 1. New report's -`version: 1`. No directives to collect unless the user supplied any -explicitly. +For step 1 specifically: there are no upstream chain reports, so the +"check upstream version match" branch of the protocol is moot — the +existing report is current if it exists at all (you can't be stale +against nothing). Re-run when the user wants fresh research with new +directives, when the source URL changed, or when the user's PR-intent +answer changed. -**If it DOES exist:** this is a re-run. Read the existing report's -`version` field. Move the existing file to -`docs/skill-optimizer//archive/01-functionality-v.md` (where -`` is the existing version). The new report's version is ``. - -Then collect operator directives — atomic new requirements that -emerged from prior iterations or from the user's current request. -Examples: "the user said the prior report missed the skill's vendor -dependency on Anthropic", "the user wants more depth on -who-uses-this." Keep them as a short bulleted list of specific asks, -NOT a context dump of the prior report's content. If there are no -new requirements, leave directives empty. +When step 5 is done, you have `new_version` and `${OPERATOR_DIRECTIVES}` +ready for step 6. ### 6. Dispatch the functionality-researcher subagent @@ -197,29 +195,14 @@ this chain is built to prevent. Dispatch. ## Iteration behavior -This skill is re-runnable. Re-run when: +Step 1 is re-runnable. Re-run triggers specific to this step: -- The source skill itself has changed upstream (and you want fresh - research) -- The user supplied new directives the prior subagent didn't have - (e.g., "the report missed responsibility X" or "go deeper on the - vendor dependency") +- The source URL changed (different upstream skill, or repo moved) - The user changed their mind on PR intent -- An upstream step's version bumped and you want this report - re-derived against current inputs (rare for step 1 since it has no - upstream reports — but possible if the user reframes the source) - -On re-run, step 5 archives the prior report under `archive/` and -bumps the version. The functionality-researcher subagent never sees -its own prior output — it re-derives the report from -`${SKILL_SOURCE}` + targeted web research + your -`${OPERATOR_DIRECTIVES}` only. Don't try to "help" by passing the -prior report's content as context; the whole point of the version -mechanism is to keep the new derivation honest. - -When this report's version bumps, downstream reports (`02-test-case.md` -and onward) become stale on their next invocation — their `inputs` -frontmatter no longer matches the current `01-functionality.md` -version, so they'll archive-and-rerun themselves on next dispatch. -You don't need to cascade manually; the next step detects the -mismatch on its own invocation. +- The user wants the report re-derived with new directives ("you + missed the vendor CLA requirement", "go deeper on who-uses-this") + +General iteration mechanics — archive, version bump, operator +directives, cascading staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Step 5 above already requires reading that file. diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md new file mode 100644 index 0000000..403df9c --- /dev/null +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -0,0 +1,187 @@ +# Iteration protocol + +**Load this file every time** a skill-optimizer chain skill executes +its "Handle iteration" step. The mechanics are mechanical and +identical across all chain skills (B1–B7); the parent skill's +SKILL.md only specifies that-skill's slot in the chain. The actual +archive + version bump + directives-collection logic lives here so +all seven chain skills behave consistently. + +If you're a chain skill's SKILL.md and you're tempted to inline this +logic instead of reading this file: don't. Inlining drifts. Read this +file, do what it says. + +## What this protocol governs + +Every chain step produces a versioned report at a canonical path +under `docs/skill-optimizer//`. When a step is re-run (either +manually or by the auto-pilot), this protocol determines: + +1. Whether the existing report (if any) is current or stale +2. How to archive a stale report before writing a new one +3. What version number the new report gets +4. How to pass cross-iteration learnings into the subagent +5. What the subagent is and isn't allowed to look at + +## Versioning convention + +Every chain report carries two frontmatter fields: + +```yaml +version: # this report's own iteration counter +inputs: + step_1_functionality: # version of each upstream report consumed + step_2_test_case: # only the ones this step actually reads + # ... etc. +``` + +Step 1 has no upstream chain reports, so it omits `inputs:`. Every +other step lists the upstream chain reports it reads, with the +version it consumed at the time of writing. + +Versions are per-step independent counters. Step 1 going from v1 → v2 +does not change step 2's version until step 2 is itself re-invoked +and detects the mismatch. + +## On every chain-skill invocation + +Apply this decision tree before dispatching the step's subagent: + +```text +1. Compute the canonical report path for this step + slug. +2. Does the canonical path exist? + │ + ├── No → This is iteration 1 for this step. + │ - new_version = 1 + │ - prior_archived = false + │ - Go to "Collect operator directives" below. + │ + └── Yes → Read the existing report's frontmatter. + Compare its `inputs:` to the current versions of each + upstream report at its canonical path. + │ + ├── All input versions match + │ AND no operator directives are pending + │ → Existing report is current; nothing to do. + │ Report success, hand off, exit. + │ + └── Any input version mismatches + OR operator passed directives requiring re-run + → Existing report is stale or operator wants fresh. + - Move existing file to + `archive/--v.md` + - new_version = existing_version + 1 + - prior_archived = true + - Go to "Collect operator directives" below. +``` + +## Collect operator directives + +`${OPERATOR_DIRECTIVES}` is a short bulleted list of **atomic new +requirements** that surfaced from prior iterations or from the user's +current request. It is **never** a context dump of the prior report's +content; it is a list of specific asks that the new derivation must +satisfy on top of its normal inputs. + +Good examples: + +- "user requested coverage for null inputs" +- "user flagged responsibility X as under-tested in the prior pass" +- "focus on the gpt-5 cluster — gemini and claude both passed" +- "user wants the report to mention the vendor's CLA requirement" + +Bad examples (these are context dumps, not atomic requirements): + +- ❌ The full text of the prior `02-test-case.md` so the subagent + can "see what we already had" +- ❌ "Here's what failed last time:" followed by failure data +- ❌ The validator's prior verdict pasted in for the subagent to + read + +If the user supplied no new requirements and you're re-running +simply because an upstream report changed, leave directives empty. +The subagent will re-derive its output from the (new) upstream +report alone. + +## Dispatch the subagent + +When you invoke the subagent for this step, pass these templated +inputs (each subagent prompt template names them with `${...}` +placeholders): + +- `${VERSION}` — `new_version` from the decision tree above +- `${OPERATOR_DIRECTIVES}` — the bulleted list (may be empty) +- Step-specific inputs (paths to upstream reports the subagent + needs, the output path, etc.) per that step's SKILL.md + +## Subagent constraints under iteration + +Two rules every reasoning subagent must follow: + +1. **Ignore your own prior output.** Do not read anything under + `archive/--v*.md`. That output is the past + contaminated reasoning the iteration mechanism exists to wall + off. Your job on iteration N is to derive from the current + upstream reports plus `${OPERATOR_DIRECTIVES}`, NOT to "improve" + the prior version. + +2. **Read upstream reports at their latest versions only.** The + canonical paths (`-.md`, no version suffix) always + point at the latest version. Don't walk `archive/` for any + reason. + +Both rules are enforced in each subagent's prompt template. If you +find yourself wanting to violate either — that's the failure mode +the iteration mechanism prevents. Don't. + +## Cascading staleness + +Each step checks **direct upstream only** — the immediate prior +report(s) it consumes, not the whole upstream chain. If step 1 goes +to v2 but step 2 is not re-run (the user judged step 2 still valid +against the new step 1), step 3 will see step 2 at v1 and treat it +as current — even though step 2's `inputs` still references step 1 +v1 transitively. + +The discipline: skipping a step's re-run is an **explicit operator +judgment** that the existing report is still valid against the new +upstream. Trust the call. Auto-pilot doesn't make this judgment — it +always re-runs on direct-upstream version mismatch, which causes the +cascade to propagate naturally as you walk the chain forward. + +If transitive staleness surfaces as confusing downstream results, +the operator re-runs the skipped step manually and the cascade +catches up on the next forward walk. + +## Archive convention + +| Path | Content | +|---|---| +| `-.md` | Always the latest version. Read this. | +| `archive/--v.md` | Prior versions, frozen. For human audit. Do not read from chain skills or subagents. | + +The `archive/` directory is browsable but inert. It exists for +contributor audit and debugging — never for the chain to consult on +the next run. + +## The bootstrapping case + +On the very first invocation of a step on a fresh slug, +`docs/skill-optimizer//` may not yet exist. Create it (and the +`archive/` subdir, even if empty) as part of the iteration step's +setup. Subsequent invocations expect both dirs to be present. + +## What this protocol does NOT cover + +- **Step 5 (`run-bench`) bench results** are timestamped under + `05-bench-results//` rather than versioned. Each bench run is + preserved naturally. The versioned report step 5 produces is the + bench-result summary (which references the timestamped run), + which follows this protocol like any other report. +- **Auto-pilot's summary report** at + `docs/skill-optimizer//autopilot-summary-.md` is + timestamped per run rather than versioned. Each auto-pilot run + produces a fresh summary; old ones are preserved naturally. +- **`vendored-skill/`** (the cached upstream skill source) is reused + across iterations of the same slug unless the source URL changed. + Not versioned; the source URL itself is the identity. From d18ac79df6f7b8bbd168043c914384031df7438c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 10:42:20 -0500 Subject: [PATCH 013/121] docs(v1.4-shared): tighten iteration-protocol against authoring philosophy MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three issues that the protocol doc was a load-bearing reference for — so they would have propagated into every chain skill's invocation — all caught on read-through: 1. "B1-B7" was project-internal jargon (those are plan task labels, not part of the skill vocabulary). Replaced with "every skill in the chain". The shipped doc should be timeless; task labels live in the plan, not the protocol. 2. The "On every chain-skill invocation" section used an ASCII box- drawing decision tree. Per Anthropic and writing-skills guidance ("use flowcharts ONLY for non-obvious decision points"), this logic is straightforward conditional flow — bullets carry it more cleanly and match standard markdown rendering. Rewritten as nested bullets. 3. The "Bad examples" of operator directives were marked with ❌. Per the global instruction "Only use emojis if the user explicitly requests it. Avoid adding emojis to files unless asked", these don't belong. Replaced with explicit "Examples that count" / "Examples that do NOT count" section labels. Net: 187 → 180 lines. --- .../iteration-protocol.md | 71 +++++++++---------- 1 file changed, 32 insertions(+), 39 deletions(-) diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 403df9c..4011ade 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -1,11 +1,11 @@ # Iteration protocol **Load this file every time** a skill-optimizer chain skill executes -its "Handle iteration" step. The mechanics are mechanical and -identical across all chain skills (B1–B7); the parent skill's -SKILL.md only specifies that-skill's slot in the chain. The actual -archive + version bump + directives-collection logic lives here so -all seven chain skills behave consistently. +its "Handle iteration" step. The mechanics are identical across +every skill in the chain; the parent skill's SKILL.md only specifies +that-skill's slot in the chain. The actual archive + version bump + +directives-collection logic lives here so the chain behaves +consistently across all steps. If you're a chain skill's SKILL.md and you're tempted to inline this logic instead of reading this file: don't. Inlining drifts. Read this @@ -45,35 +45,28 @@ and detects the mismatch. ## On every chain-skill invocation -Apply this decision tree before dispatching the step's subagent: +Before dispatching the step's subagent: -```text 1. Compute the canonical report path for this step + slug. -2. Does the canonical path exist? - │ - ├── No → This is iteration 1 for this step. - │ - new_version = 1 - │ - prior_archived = false - │ - Go to "Collect operator directives" below. - │ - └── Yes → Read the existing report's frontmatter. - Compare its `inputs:` to the current versions of each - upstream report at its canonical path. - │ - ├── All input versions match - │ AND no operator directives are pending - │ → Existing report is current; nothing to do. - │ Report success, hand off, exit. - │ - └── Any input version mismatches - OR operator passed directives requiring re-run - → Existing report is stale or operator wants fresh. - - Move existing file to - `archive/--v.md` - - new_version = existing_version + 1 - - prior_archived = true - - Go to "Collect operator directives" below. -``` +2. Check whether the canonical path exists, and act: + + - **Path does not exist.** This is iteration 1 for this step. Set + `new_version = 1`. Proceed to "Collect operator directives" below. + + - **Path exists.** Read the existing report's frontmatter. Compare + its `inputs:` to the current versions of each upstream report + at its canonical path. Two sub-cases: + + - **All input versions match AND no operator directives are + pending.** The existing report is current; nothing to do. + Report success, hand off, exit. + + - **Any input version mismatches OR operator passed directives + requiring a re-run.** The existing report is stale (or the + operator wants fresh). Move the existing file to + `archive/--v.md`. Set + `new_version = existing_version + 1`. Proceed to "Collect + operator directives" below. ## Collect operator directives @@ -83,20 +76,20 @@ current request. It is **never** a context dump of the prior report's content; it is a list of specific asks that the new derivation must satisfy on top of its normal inputs. -Good examples: +Examples that count as atomic requirements: - "user requested coverage for null inputs" - "user flagged responsibility X as under-tested in the prior pass" - "focus on the gpt-5 cluster — gemini and claude both passed" - "user wants the report to mention the vendor's CLA requirement" -Bad examples (these are context dumps, not atomic requirements): +Examples that do NOT count (these are context dumps, not atomic +requirements — reject the temptation): -- ❌ The full text of the prior `02-test-case.md` so the subagent - can "see what we already had" -- ❌ "Here's what failed last time:" followed by failure data -- ❌ The validator's prior verdict pasted in for the subagent to - read +- The full text of the prior `02-test-case.md` so the subagent can + "see what we already had" +- "Here's what failed last time:" followed by failure data +- The validator's prior verdict pasted in for the subagent to read If the user supplied no new requirements and you're re-running simply because an upstream report changed, leave directives empty. From 552e7fafd09958c6cba5c80633743b5cbb649f84 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 10:49:51 -0500 Subject: [PATCH 014/121] docs(skill-optimizer-investigate-functionality): disambiguate step numbering + front-load load-bearing context MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two related fixes: 1. Internal workflow step numbering (1-7) collided with chain step numbering (1-7 referring to other skills in the chain). Same word, two meanings — an agent reading "step 5" could plausibly think either "this skill's Handle-iteration step" or "the run-bench skill in the chain". Relabeled internal workflow steps to (a) through (g) so they're visually distinct from chain step references. Cross-references inside the doc updated accordingly. A note in the new "Before you start" section declares the convention explicitly so future readers don't have to infer it. 2. The "Read iteration-protocol.md" instruction was buried at internal step (e) — middle of a 7-step workflow. An agent reading sequentially might glance past it once they're in execution momentum. Added a "## Before you start" section right after the intro paragraph, listing the two load-bearing things to know up front: (1) read the iteration protocol now, (2) you will dispatch a subagent — you don't do the research yourself. Step (e) still requires reading the protocol — the "Before you start" section primes the agent so step (e) becomes reinforcement rather than first contact. 229 lines total (up from 208; "Before you start" earns its place as the discipline frame). Sets the pattern for B2-B8 — each will similarly relabel its workflow steps with (a)-(g) and front-load the two load-bearing context items. --- .../SKILL.md | 45 ++++++++++++++----- 1 file changed, 33 insertions(+), 12 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index e1e3ca9..668ea24 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -10,6 +10,27 @@ runs a researcher subagent to figure out what it's supposed to do, and writes `docs/skill-optimizer//01-functionality.md` — the briefing document every later step consumes. +## Before you start + +Two load-bearing pieces of context to load NOW, before the workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation — for archive + logic, version bumps, operator-directive handling, and the + constraints subagents must obey. Workflow step (e) below requires + it; loading it now lets you execute that step deterministically + instead of improvising. + +2. **You will dispatch a subagent for the actual research; you do NOT + do the research yourself in this session.** Workflow step (f) is + the dispatch. The rationale is in "Why limited-context dispatch + matters" below — read it if you're tempted to skip the dispatch. + +Throughout this document, "step 1" through "step 7" (no parens) refer +to skills in the chain (this skill is step 1; `investigate-test-case` +is step 2; etc.). Internal workflow steps within THIS skill are +labelled "(a)" through "(g)" to avoid the collision. + ## What you produce A single report at `docs/skill-optimizer//01-functionality.md`, @@ -55,14 +76,14 @@ stable copy without re-fetching. ## Workflow -### 1. Classify the source +### (a) Classify the source Is the source an **upstream** skill (a URL or `//` slug) or a **local** skill (a filesystem path that already exists)? If the user gave a bare name with no URL and no path, ask them to clarify before continuing. -### 2. Ask about PR intent +### (b) Ask about PR intent **Upstream skills:** ask the user, in roughly these words: @@ -89,13 +110,13 @@ to PR this to project X"), treat it as PR-intent. In that case: If the user mentions nothing about a PR for the local skill, set `pr_submission_intent: false` and move on. -### 3. Vendor the source (upstream only) +### (c) Vendor the source (upstream only) Fetch the skill's files into `vendored-skill/` at the working-directory root. Downstream steps read from this vendored copy as a stable reference. Local-source skills don't need vendoring. -### 4. Determine the slug and the report path +### (d) Determine the slug and the report path `` is the source skill's own directory or file name: @@ -104,7 +125,7 @@ reference. Local-source skills don't need vendoring. Report path: `docs/skill-optimizer//01-functionality.md`. -### 5. Handle iteration +### (e) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** and apply it for this step. The protocol is mechanical and identical @@ -120,18 +141,18 @@ against nothing). Re-run when the user wants fresh research with new directives, when the source URL changed, or when the user's PR-intent answer changed. -When step 5 is done, you have `new_version` and `${OPERATOR_DIRECTIVES}` -ready for step 6. +When (e) is done, you have `new_version` and `${OPERATOR_DIRECTIVES}` +ready for (f). -### 6. Dispatch the functionality-researcher subagent +### (f) Dispatch the functionality-researcher subagent **Do NOT do the research yourself in this session.** Dispatch the functionality-researcher subagent via the `Agent` tool (with worktree isolation if your environment supports it). Load the prompt template at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md), substitute the templated inputs (`${SKILL_SOURCE}`, `${OUTPUT_PATH}`, -`${PR_SUBMISSION_INTENT}`, `${VERSION}` from step 5, -`${OPERATOR_DIRECTIVES}` from step 5, vendored path), and dispatch. +`${PR_SUBMISSION_INTENT}`, `${VERSION}` from (e), +`${OPERATOR_DIRECTIVES}` from (e), vendored path), and dispatch. The subagent sees: @@ -151,7 +172,7 @@ The subagent does NOT see: The subagent writes the report itself and returns a brief summary. -### 7. Confirm and hand off +### (g) Confirm and hand off When the subagent returns: @@ -205,4 +226,4 @@ Step 1 is re-runnable. Re-run triggers specific to this step: General iteration mechanics — archive, version bump, operator directives, cascading staleness — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Step 5 above already requires reading that file. +Workflow step (e) above already requires reading that file. From 2fd60b15ee22c888b4206ebd6b30443511efe4c4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 11:46:59 -0500 Subject: [PATCH 015/121] feat(skill-optimizer-investigate-test-case): SKILL.md + spec: source visibility moved from step 4-blocked to step 4-allowed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two coupled changes — B2's SKILL.md draft and a spec update that makes the source-visibility split consistent across the chain. Spec changes (Subagent constraints table): - Step 2 (test-case designer) — explicitly added "skill source content" to "does NOT see". Coverage design happens at the responsibility level here; concrete fixture details enter at step 4. Without this constraint, designers could gerrymander test cases around the source's literal phrasing instead of reasoning from stated responsibilities. - Step 4 (test writer) — removed "the skill's content" from "does NOT see"; added "skill source content" to "sees". Updated the rationale: fixture writing needs concrete patterns and violation examples, which come from source. The per-case test spec from step 2 constrains what the fixture should test, so the gerrymandering risk is bounded. Still blocks other test cases (prevents copying across the suite) and grader internals (prevents grader-leak hacking). Architecture: source enters the chain at step 1 (research), exits at step 2 (responsibility-level design needs no source), enters at step 4 (fixture writing needs source detail), stays present for step 6 (analyze failures) and step 7 (optimize/validate). The spec's step-4 body description updated to match. B2 SKILL.md: - New file, follows B1's template: front-loaded "Before you start" section, lettered workflow steps (a)-(f), bold discipline markers at action sites, why-this-matters rationale, edge cases, iteration behavior section. - 6 internal steps: (a) confirm prerequisites, (b) handle iteration, (c) dispatch designer subagent, (d) confirm subagent output, (e) user gate (present + collect picks), (f) conditional handoff based on pr_submission_intent. - "picked" frontmatter field is the operator's responsibility — subagent writes proposal with picked: []; operator fills picked after user gate. - Three-response handling in user gate: pick subset/all, ask for revision, pick zero. - 215 lines, description 479 chars. --- docs/skill-optimizer-v1.4-spec.md | 13 +- .../SKILL.md | 212 +++++++++++++++++- 2 files changed, 218 insertions(+), 7 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index acfbe37..d76aedb 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -146,9 +146,9 @@ and dispatches the subagent with only the narrow chunks it needs. | Subagent | Sees | Does NOT see | Why | |---|---|---|---| | Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, prior `01-functionality.md` drafts | Pure research, no contamination across iterations | -| Test-case designer (step 2) | `01-functionality.md` (latest), `${OPERATOR_DIRECTIVES}` | Prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Coverage design must reason from the skill's responsibilities, not from "what just failed" | +| Test-case designer (step 2) | `01-functionality.md` (latest), `${OPERATOR_DIRECTIVES}` | Skill source content, prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Coverage design is at responsibility level — concrete fixture details enter the chain at step 4, not here. Reasoning from source would gerrymander tests around the source's literal phrasing instead of testing stated responsibilities | | Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change | Just upstream facts | -| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` + `${OPERATOR_DIRECTIVES}` | The skill's content, other test cases, the eval grader's matching logic | Prevents grader-hacking; prevents copying existing tests | +| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other test cases in the suite, the eval grader's matching logic | Fixture writing needs source detail (specific patterns, concrete violation examples). Sees source for grounding; the test case spec from step 2 constrains what the fixture should test, bounding the gerrymandering risk. Still blocked from seeing other cases (prevents copying) and grader internals (prevents grader-leak hacking) | | Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), prior `06-analysis.md` drafts | Forces it to think about the SKILL, not the SOLUTIONS; iteration-isolated | | Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, prior optimizer attempts | Forces principled improvement, not pattern-match patches; no attachment to prior failed attempts | | Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, prior validator verdicts | Independent check; can't be biased by what the optimizer (or a prior validator round) told itself | @@ -385,8 +385,13 @@ Per-step iteration behavior is noted at the end of each subsection. (hand-crafted GOOD/BAD/EMPTY findings.txt fixtures against each grader) → commit - **Dispatches:** test-writer subagent per case (limited context: - single case spec + functionality report; does NOT see skill content - or other test cases — prevents grader-hacking and prevents copying) + single case spec + functionality report + skill source content; + does NOT see other test cases or eval grader's matching logic — + prevents copying across cases and grader-leak hacking. Source + access is granted at this step because concrete fixture writing + needs specific violation patterns; the per-case test spec from + step 2 constrains what the fixture should test, bounding the + gerrymandering risk) - **Parallelizable:** each case independent; dispatch in a single message - **Handoff:** "Invoke `skill-optimizer-run-bench` to measure baseline." diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index e781bb3..d0e3751 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -1,9 +1,215 @@ --- name: skill-optimizer-investigate-test-case -description: Use when the user wants to design or propose test cases for a skill — enumerates the skill's responsibilities and proposes ranked test cases the user can pick from. +description: Use when the user wants to design or propose test cases for a skill — phrases like "design tests for this skill", "propose test cases", "what should we test", "plan test coverage for this skill". Also triggers mid-way through skill-optimizer chain work, once a functionality report exists and the next thing is figuring out what to test. Use even when the user doesn't explicitly say "design" — any phrasing about figuring out what tests to build for a skill should trigger this. --- # skill-optimizer-investigate-test-case - - +Step 2 of the skill-optimizer chain. Takes the functionality report +from step 1, dispatches a designer subagent to enumerate the skill's +responsibilities and propose a ranked list of test cases, then asks +the user to pick which subset to actually build. Writes +`docs/skill-optimizer//02-test-case.md`. + +## Before you start + +Two load-bearing pieces of context to load NOW, before the workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation — for archive + logic, version bumps, operator-directive handling, and the + constraints subagents must obey. Workflow step (b) below requires + it; loading it now lets you execute that step deterministically + instead of improvising. + +2. **You will dispatch a subagent for the actual design work; you do + NOT enumerate responsibilities or design cases yourself in this + session.** Workflow step (c) is the dispatch. The rationale is in + "Why limited-context dispatch matters" below — read it if you're + tempted to skip the dispatch. + +Throughout this document, "step 1" through "step 7" (no parens) refer +to skills in the chain (this skill is step 2; `investigate-functionality` +is step 1; etc.). Internal workflow steps within THIS skill are +labelled "(a)" through "(f)" to avoid the collision. + +## What you produce + +A single report at `docs/skill-optimizer//02-test-case.md`, +where `` matches the slug from step 1's report. + +The report has structured frontmatter: + +```yaml +--- +version: 1 +inputs: + step_1_functionality: +picked: [] # operator fills this AFTER the user-gate step +--- +``` + +**`picked`** — list of test case names the user selected from the +ranked proposals. Empty until step (e) collects the user's picks. +Step 4 (`write-tests`) operates on this subset, NOT on all proposals. + +For the body template (what each proposed case must include and how +they're ranked), see +[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md). +The subagent writes the ranked proposal body; the operator session +sets the `picked` field after the user gate. + +## Workflow + +### (a) Confirm prerequisites + +`docs/skill-optimizer//01-functionality.md` must exist and have +valid frontmatter. If it doesn't, tell the user to run +`skill-optimizer-investigate-functionality` first and stop here. +Record its `version` field — you'll pass it as +`inputs.step_1_functionality` later so this report records which +version of the functionality understanding it was derived from. + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it for this step. The protocol tells you how to check for +an existing `02-test-case.md`, compare its `inputs.step_1_functionality` +against the current `01-functionality.md` version, archive if stale, +compute the new version, and collect operator directives. Do not +improvise; read the file. + +### (c) Dispatch the test-case-designer subagent + +**Do NOT enumerate responsibilities or design cases yourself in this +session.** Dispatch the test-case-designer subagent via the `Agent` +tool (with worktree isolation if your environment supports it). Load +the prompt template at +[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md), +substitute the templated inputs (`${FUNCTIONALITY_PATH}`, +`${OUTPUT_PATH}`, `${VERSION}` from (b), +`${OPERATOR_DIRECTIVES}` from (b)), and dispatch. + +The subagent sees: + +- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` only +- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new + requirements (may be empty) +- The output path and version number + +The subagent does NOT see: + +- Prior `02-test-case.md` drafts (the ones now in `archive/`) +- Any `06-analysis.md` (existing analyses) +- `07-improvement-proposal.md`, `07-validator-verdict.md` +- Raw failure data, `findings.txt`, bench results +- The skill's source content itself + +The "no source content" constraint is load-bearing in addition to the +usual no-prior-failures one: design coverage from the skill's STATED +responsibilities, not from the source's literal phrasing. Letting the +subagent read the source lets it gerrymander tests around the +source's exact shape (testing what the prose happens to say, not +what the skill is supposed to do). + +The subagent writes the proposal body itself and returns a brief +summary including the top-3 cases by importance. + +### (d) Confirm subagent output + +Verify the report file exists, the frontmatter parses, and the body +has the expected sections per the subagent prompt template. The +`picked` field at this point is still empty — that's correct; the +user gate is next. + +### (e) User gate: present proposals, collect picks + +Show the user the ranked proposal (summarize or paste the body — +your call based on length). Ask: + +> Here are the proposed test cases ranked by importance. Which ones +> should we actually build? You can pick all of them, a subset, or +> ask for a revised proposal. + +Three realistic responses: + +1. **User picks a subset (or all).** Update the `picked` frontmatter + field with the chosen case names. Commit the file. +2. **User asks for a revised proposal.** Treat their feedback as new + operator directives, return to (b) to handle iteration (which + will archive this version), and re-dispatch the subagent at (c). +3. **User picks zero.** Ask whether they're revising (treat as case + 2) or abandoning the optimization for this skill (exit honestly + without progressing to step 4). + +Don't auto-pick on the user's behalf — even if all cases look +important, the user owns this decision (they're paying for the +test-writing in step 4 and the bench run in step 5). + +### (f) Hand off + +Read `01-functionality.md`'s `pr_submission_intent` field. + +If `pr_submission_intent: true`: + +> Next, invoke `skill-optimizer-investigate-submissions`. After +> that, invoke `skill-optimizer-write-tests` with the picked subset. + +If `pr_submission_intent: false`: + +> Skipping step 3 (no PR submission planned). Next, invoke +> `skill-optimizer-write-tests` with the picked subset. + +No late "submit a PR?" prompts — the decision was made at step 1. + +## Why limited-context dispatch matters + +If the operator session is the one designing tests, it has already +absorbed prior failure data, the optimizer's past attempts, the +validator's verdicts, and whatever framing the user has applied. +That context biases test design toward "tests that would have caught +the things I just watched fail" — the coverage version of ducktape. +You end up with a test suite that validates the patch instead of +testing the skill independently. + +A subagent that sees only `01-functionality.md` reasons from the +skill's stated responsibilities, not from prior outcomes. The +proposal is coverage-oriented, not regression-defensive. That's the +invariant we need to ship a fair test set. + +If you find yourself thinking "I'll just propose the tests myself, +the subagent dispatch is bureaucratic overhead" — that's the failure +mode this chain is built to prevent. Dispatch. + +## Edge cases + +- **`01-functionality.md` missing** — tell the user to run step 1 + first; don't try to derive responsibilities yourself. +- **Subagent's proposal has fewer cases than expected** — that's a + signal the functionality report is thin. Either re-run step 1 with + directives ("the report missed responsibility X") or accept the + small proposal if the skill is genuinely small. +- **User wants to add cases the subagent didn't propose** — accept + them as additional entries in `picked`; step 4 will treat them + the same as subagent-proposed cases. +- **Operator directives contradict each other** (e.g., "focus + coverage on X" + "ignore X") — surface the contradiction to the + user before re-dispatching; don't try to resolve it yourself. + +## Iteration behavior + +Step 2 is re-runnable. Re-run triggers specific to this step: + +- The user wants a different cut of coverage (different cases + proposed, different ranking) +- Step 4 (`write-tests`) flagged a picked case as unimplementable +- Step 5 (`run-bench`) showed all tests pass on the baseline — too + easy, need harder cases +- Step 6 (`analyze-result`) surfaced a coverage gap (responsibility + X has no tests; user wants one) +- Step 1's version bumped (functionality understanding changed) + +General iteration mechanics — archive, version bump, operator +directives, cascading staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. From ab3506ab4e65bf6839ead14a0ad70963a341e0dd Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 19 May 2026 14:04:08 -0500 Subject: [PATCH 016/121] feat(skill-optimizer-investigate-submissions): SKILL.md body MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 3 of the chain — OPTIONAL, runs only when 01-functionality.md has pr_submission_intent: true. Researches the upstream repo's PR conventions and writes 03-submissions.md as the verbatim-pastable context block the validator (step 7) uses for external consistency. Follows the B1/B2 template: - Front-loaded "Before you start" with iteration-protocol pointer + dispatch discipline + step-numbering convention - Frontmatter with version, inputs.step_1_functionality, plus step-specific fields: upstream_repo, upstream_branch_target, license, requires_cla - 5 lettered workflow steps: (a) confirm prerequisites including pr_submission_intent gate, (b) handle iteration, (c) dispatch subagent, (d) confirm subagent output with blocker flagging, (e) hand off - Why limited-context dispatch matters: rationale specific to step 3 — the validator's external consistency check depends on the submissions report being neutral upstream facts, not advocacy for the proposed change - Edge cases: private repos, non-GitHub hosts, empty PR history, copyleft license - Iteration behavior: explicitly note this rarely re-runs; upstream conventions change slowly 235 lines, description 578 chars. --- .../SKILL.md | 232 +++++++++++++++++- 1 file changed, 229 insertions(+), 3 deletions(-) diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 54b0460..0e76c9b 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -1,9 +1,235 @@ --- name: skill-optimizer-investigate-submissions -description: Use when the user wants to research a skill's upstream PR conventions — produces a context block with license, CLA, frontmatter spec, file-location rules, PR-shape patterns, and rejection signals. +description: Use when the user wants to research a skill's upstream PR conventions — phrases like "research PR conventions for this skill", "what does the upstream repo require for contributions", "investigate submissions for X", or when prepping a PR-bound optimization run and you need to know the upstream's rules. Triggers mid-way through skill-optimizer chain work when `01-functionality.md` has `pr_submission_intent: true`. Use even when the user doesn't explicitly say "investigate submissions" — any phrasing about figuring out the upstream's contribution rules should trigger this. --- # skill-optimizer-investigate-submissions - - +Step 3 of the skill-optimizer chain — OPTIONAL, runs only when the +target skill is bound for upstream PR submission. Takes the source +slug, dispatches a researcher subagent that uses the `gh` CLI to +gather the upstream repo's contribution conventions (license, CLA, +frontmatter spec, file-location rules, PR-shape patterns from recent +merged + closed-without-merge PRs), and writes +`docs/skill-optimizer//03-submissions.md` — the +verbatim-pastable context block the validator (step 7) uses for its +external consistency check. + +## Before you start + +Two load-bearing pieces of context to load NOW, before the workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation — for archive + logic, version bumps, operator-directive handling, and the + constraints subagents must obey. Workflow step (b) below requires + it; loading it now lets you execute that step deterministically + instead of improvising. + +2. **You will dispatch a subagent for the actual research; you do NOT + scrape the upstream repo yourself in this session.** Workflow step + (c) is the dispatch. The rationale is in "Why limited-context + dispatch matters" below — read it if you're tempted to skip the + dispatch. + +Throughout this document, "step 1" through "step 7" (no parens) refer +to skills in the chain (this skill is step 3; `investigate-functionality` +is step 1; etc.). Internal workflow steps within THIS skill are +labelled "(a)" through "(e)" to avoid the collision. + +## What you produce + +A single report at `docs/skill-optimizer//03-submissions.md`, +where `` matches the slug from step 1's report. + +The report has structured frontmatter: + +```yaml +--- +version: 1 +inputs: + step_1_functionality: +upstream_repo: / +upstream_branch_target:
+license: +requires_cla: true | false +--- +``` + +**`upstream_branch_target`** — some repos use `main` for incremental +changes and `next` for new skills; some use a single branch. The +subagent determines this from recent merged PRs and records it here +so the validator can check the proposed PR targets the correct +branch. + +**`requires_cla`** — true if the upstream requires a Contributor +License Agreement before merging (Google, Apache Foundation, etc.). +The PR-packaging step (in `improve-skill`'s upstream-PR handoff) +surfaces this to the operator. + +For the body template (what sections the researcher must cover — +license details, frontmatter spec extracted from existing skills, +file-location conventions, prefix taxonomy, PR-shape patterns from +recent merged + closed-without-merge PRs, rejection signals from the +closed-without-merge set), see +[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md). +The subagent writes the body itself. + +## Workflow + +### (a) Confirm prerequisites + +Two prerequisites: + +1. `docs/skill-optimizer//01-functionality.md` must exist and + have valid frontmatter. If it doesn't, tell the user to run + `skill-optimizer-investigate-functionality` first and stop here. +2. That report's `pr_submission_intent` field must be `true`. If it's + `false`, this skill should not run — tell the user step 3 is + skipped for local-only optimization runs and refer them to step 4 + (`write-tests`). + +Record the functionality report's `version` — you'll pass it as +`inputs.step_1_functionality` later. + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it for this step. The protocol tells you how to check for +an existing `03-submissions.md`, compare its +`inputs.step_1_functionality` against the current `01-functionality.md` +version, archive if stale, compute the new version, and collect +operator directives. Do not improvise; read the file. + +Note: this report rarely needs re-running. Upstream PR conventions +change slowly. The most common reason to re-run is that the upstream +repo updated its `CONTRIBUTING.md` or CLA requirements — surface this +as an operator directive if you know it. + +### (c) Dispatch the submission-researcher subagent + +**Do NOT scrape the upstream repo yourself in this session.** +Dispatch the submission-researcher subagent via the `Agent` tool +(with worktree isolation if your environment supports it). Load the +prompt template at +[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md), +substitute the templated inputs (`${UPSTREAM_REPO}` from +`01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, +`${VERSION}` from (b), `${OPERATOR_DIRECTIVES}` from (b)), and +dispatch. + +The subagent sees: + +- The upstream repo via `gh` CLI (PR list, repo-file API, + `CONTRIBUTING.md`, license file, existing skill files for + frontmatter spec extraction) +- The last 10 merged PRs and last 5 closed-without-merge PRs (for + shape patterns and rejection signals) +- The skill slug being researched (so it can look at similar PRs in + the same skill category) +- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new + requirements (may be empty) +- The output path and version number + +The subagent does NOT see: + +- Prior `03-submissions.md` drafts (the ones now in `archive/`) +- Any information about the proposed change being optimized (the + report is purely about upstream facts, not about whether a specific + change will be accepted) +- Existing analyses, tests, or failure data from the chain +- The vendored skill source + +The constraint that the subagent doesn't see the proposed change is +load-bearing: the report must be neutral upstream facts, not advocacy +for a specific change. The validator (step 7) will later check the +proposed change against this report — if the report is biased toward +the change, the validator's external consistency check loses its +independence. + +The subagent writes the report itself and returns a brief summary +including: license, CLA requirement, branch target, and any +high-risk rejection signals it spotted in the closed-without-merge +PRs. + +### (d) Confirm subagent output + +Verify the report file exists, the frontmatter parses, and the body +has the expected sections per the subagent prompt template. Flag any +of these as blockers for the user: + +- License is GPL or other copyleft (affects whether the optimization + can be merged at all) +- `requires_cla: true` (operator will need to sign before submitting) +- Frontmatter spec includes fields not currently in the vendored + skill's frontmatter (the optimization may need to add them) + +### (e) Hand off + +Report to the user: the report path, a one-line summary (license / +CLA / branch target / any flagged blockers), then the handoff +message. + +Handoff message, verbatim: + +> The validator in `skill-optimizer-improve-skill` will read this +> report for its external consistency check. If you haven't run +> `skill-optimizer-write-tests` yet, invoke that next. + +The chain doesn't enforce ordering between step 3 and step 4 — they +can run in either order or in parallel. Step 7's validator is what +consumes step 3's output, not step 4. + +## Why limited-context dispatch matters + +The submission-researcher subagent must produce a report that the +validator (step 7) can trust as independent. If the operator session +does the research, it has already absorbed the optimization context +(the proposed change, the prior failures, the user's framing of what +"good" looks like). That context biases the research toward +documenting upstream conventions in a way that justifies the proposed +change — and the validator's external consistency check then has no +real independence. + +Walling the researcher off in a subagent that sees only the upstream +repo facts (and no information about the proposed change) preserves +the validator's later check. The report is neutral upstream facts; +the validator's job is to check whether the proposed change conforms +to those facts. + +If you find yourself thinking "I'll just run a few `gh` commands +myself, the dispatch is bureaucratic overhead" — that's the failure +mode this chain is built to prevent. Dispatch. + +## Edge cases + +- **Upstream repo is private or requires auth** — the subagent will + surface this as a blocker; ask the user to authenticate `gh` (or + to provide credentials) and re-dispatch. +- **Upstream uses a non-`gh`-friendly host (GitLab, Bitbucket, etc.)** + — the subagent will surface this. The current chain assumes + GitHub-hosted upstreams; non-GitHub cases require manual research + and pasting the report content directly. Tell the user. +- **No merged PRs in the upstream's history yet** — the researcher + can't extract shape patterns from absent data; the report will be + thinner. Surface this honestly rather than making up patterns. +- **License is copyleft or otherwise incompatible with + redistribution** — flag as a blocker. The optimization may need to + be kept internal even if mechanically possible to merge. + +## Iteration behavior + +Step 3 is re-runnable but rarely needs it. Re-run triggers specific +to this step: + +- The upstream repo updated its `CONTRIBUTING.md`, license, or CLA + requirements +- The upstream's PR-shape conventions visibly shifted (recent merged + PRs no longer match the older patterns) +- Step 1's version bumped because the source URL changed (different + upstream repo entirely) + +General iteration mechanics — archive, version bump, operator +directives, cascading staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. From fb1e269c32d6d32273b28cc0e6a7652b56fa36a6 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 08:34:28 -0500 Subject: [PATCH 017/121] feat(v1.4-shared): add "Re-run authorization" section to iteration protocol + apply in B2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A chain skill never invokes another chain skill on its own. When a step finds its upstream input unsatisfactory (thin functionality report, too-easy bench, missing test case, wrong-target PR conventions), it surfaces the finding and stops — the user (or auto-pilot driver) decides whether to re-run upstream, accept the situation, or abandon. This was implicit in the architecture but not stated as a rule. Adding it explicitly to iteration-protocol.md as a new section between "Cascading staleness" and "Archive convention". Same section covers the forward-handoff exception (those are normal chain flow, not backward triggers). Updated B2's edge case for "Subagent's proposal has fewer cases than expected" — was "Either re-run step 1 with directives or accept the small proposal", now "Surface this to the user with two options ... Do not re-invoke step 1 yourself — re-runs require an active signal from the user", with reference back to the protocol doc's new section. Audit: only B2 had a wording suggesting backward auto-trigger; B1 and B3 don't suggest re-running upstream steps. Same rule applies to all future chain skills (B4-B8) — the protocol doc is the chain-wide source of truth. --- .../SKILL.md | 10 +++++--- .../iteration-protocol.md | 24 +++++++++++++++++++ 2 files changed, 31 insertions(+), 3 deletions(-) diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index d0e3751..620f2eb 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -186,9 +186,13 @@ mode this chain is built to prevent. Dispatch. - **`01-functionality.md` missing** — tell the user to run step 1 first; don't try to derive responsibilities yourself. - **Subagent's proposal has fewer cases than expected** — that's a - signal the functionality report is thin. Either re-run step 1 with - directives ("the report missed responsibility X") or accept the - small proposal if the skill is genuinely small. + signal the functionality report is thin. Surface this to the user + with two options: re-run step 1 with directives ("the report + missed responsibility X") or accept the small proposal if the + skill is genuinely small. Do not re-invoke step 1 yourself — + re-runs require an active signal from the user. (Same rule for + any backward trigger, per the iteration protocol's + "Re-run authorization" section.) - **User wants to add cases the subagent didn't propose** — accept them as additional entries in `picked`; step 4 will treat them the same as subagent-proposed cases. diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 4011ade..b7e7f26 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -146,6 +146,30 @@ If transitive staleness surfaces as confusing downstream results, the operator re-runs the skipped step manually and the cascade catches up on the next forward walk. +## Re-run authorization + +A chain skill never invokes another chain skill on its own. When a +step finds that its upstream input is unsatisfactory — a thin +functionality report, a too-easy bench, a missing-but-needed test +case, a wrong-target PR-conventions report — it surfaces the +finding to the user and stops. The user (or the auto-pilot driver +applying its default policies) decides whether to re-run an +upstream step, accept the situation, or abandon the run. + +This applies to backward triggers specifically. Forward handoffs +(step N hands off to step N+1) are part of the chain's normal flow +and the handing-off skill emits the handoff message; the user or +auto-pilot acts on it. + +Why this rule exists: chain skills running on user request must +give the user control over what they're paying for. Re-running an +upstream step takes time and tokens; the user should authorize that +explicitly, not have the agent decide on its own. The user is also +the one who knows whether the upstream-thin condition is real +("yes, the functionality report did miss responsibility X — re-run") +or expected ("no, the skill really is that small — accept the small +proposal"). + ## Archive convention | Path | Content | From 97b67a406167a12dd0eb2370fe3739dddea4d550 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 08:45:09 -0500 Subject: [PATCH 018/121] fix(skill-optimizer-investigate-submissions): three corrections per review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 1. upstream_branch_target placeholder syntax. Was "
" which reads as a closed enum. The actual value is the literal branch name (could be `develop`, `release-2024`, etc.), so the placeholder should describe the field meaning, not enumerate three literal options. Now: "". 2. Removed the "frontmatter spec includes fields not currently in the vendored skill's frontmatter" blocker. It described a real fact in the report (upstream uses fields the vendored skill lacks) but isn't a blocker — the optimizer (step 7) reads 03-submissions.md and would add missing fields naturally. Doesn't need a separate operator alert. 3. Removed the "License is GPL or other copyleft" blocker. Copyleft licenses don't mechanically block PR submission — the upstream is whatever-licensed; a PR becomes part of that codebase under that license. Organization-policy concerns about contributing to copyleft projects exist but are contributor-side decisions, not chain-level blockers. Step (d) reframed from "flag these blockers" to "verify file + only the CLA fact needs explicit mention since it requires operator-side work outside the chain". Other frontmatter fields (license, repo, branch target) are facts downstream steps consume directly. Edge case for "license is copyleft" replaced with one for unusual frontmatter conventions — that's the actual blocker shape (subagent can't extract a consistent spec, so the optimizer has to make a judgment call). --- .../SKILL.md | 28 +++++++++++-------- 1 file changed, 17 insertions(+), 11 deletions(-) diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 0e76c9b..142d5ca 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -50,7 +50,7 @@ version: 1 inputs: step_1_functionality: upstream_repo: / -upstream_branch_target:
+upstream_branch_target: license: requires_cla: true | false --- @@ -155,14 +155,16 @@ PRs. ### (d) Confirm subagent output Verify the report file exists, the frontmatter parses, and the body -has the expected sections per the subagent prompt template. Flag any -of these as blockers for the user: +has the expected sections per the subagent prompt template. -- License is GPL or other copyleft (affects whether the optimization - can be merged at all) -- `requires_cla: true` (operator will need to sign before submitting) -- Frontmatter spec includes fields not currently in the vendored - skill's frontmatter (the optimization may need to add them) +If `requires_cla: true`, mention it explicitly when handing off in +(e) — the operator will need to sign the upstream's CLA before the +PR can be merged, which is operator-side work that has to happen +outside the chain. + +The other frontmatter fields (`license`, `upstream_repo`, +`upstream_branch_target`) are facts in the report that downstream +steps consume directly; they don't need separate operator alerts. ### (e) Hand off @@ -213,9 +215,13 @@ mode this chain is built to prevent. Dispatch. - **No merged PRs in the upstream's history yet** — the researcher can't extract shape patterns from absent data; the report will be thinner. Surface this honestly rather than making up patterns. -- **License is copyleft or otherwise incompatible with - redistribution** — flag as a blocker. The optimization may need to - be kept internal even if mechanically possible to merge. +- **Upstream uses unusual / non-discoverable frontmatter + conventions** — if the subagent can't extract a consistent + frontmatter spec from recent merged PRs (the upstream is too + small, or the convention varies wildly), the report's frontmatter + spec section will say so. Tell the user; the optimizer (step 7) + will then have to make a judgment call rather than mechanically + conform. ## Iteration behavior From 0df252f784ddf2ac37eaebae5efb3861c067f6c9 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 09:20:14 -0500 Subject: [PATCH 019/121] feat(v1.4): introduce maintenance-step pattern for B2 (and future B4) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per operator review: B2's re-run was modeled as fresh derivation, which loses continuity. The user's `picked` choices and existing case names should survive a re-run; otherwise the user has to re-pick everything every time they add coverage. The "archive as inert audit trail" rule means continuity has to come from the current canonical file, not from peeking at archived prior versions. New concept: chain steps come in two kinds. - Fresh-derivation steps (1, 3, 6, 7): subagent never sees its own step's prior or current file. The "anti-ducktape" rule applies in full — optimizer must not see prior attempts, analyzer must not see prior analyses. - Maintenance steps (2, 4): subagent reads the current canonical file as load-bearing input and produces an extended version of it. Existing entries the user has invested in (picks, manually-added cases) are preserved unless directives explicitly say to revise. Archive still happens for audit; archive is still inert. iteration-protocol changes: - New "Step kinds: fresh-derivation vs maintenance" section classifying each step. - "On every chain-skill invocation" decision tree updated to branch on step kind: fresh-derivation copies-to-archive then derives from scratch; maintenance copies-to-archive then dispatches with current file as input. - "Subagent constraints under iteration" restructured to make the third rule kind-dependent (do/don't read your own canonical file). B2 changes: - Step (b) Handle iteration: explicitly invokes the maintenance pattern. Notes the special case where step 1's version bumped (responsibility set changed) — user decides whether to extend or start fresh by deleting the canonical file before dispatch. - Step (c) Dispatch: new templated input ${EXISTING_CASES_PATH} (empty on iteration 1, current 02-test-case.md on re-runs). Subagent's "sees" list adds the current file with explicit preserve-existing semantics. "Does NOT see" list now correctly excludes archive only (was incorrectly excluding all prior drafts). - Step (e) User gate: response #2 "user wants additions or revisions" now reflects maintenance — user does not need to re-pick everything since existing picks are preserved. Spec changes: - Subagent constraints table: test-case-designer row updated to reflect maintenance pattern. Sees-list adds "current 02-test-case.md (when extending in maintenance mode)"; does-NOT-see-list correctly scoped to "archived prior drafts" only; why-rationale updated. Note: B4 (write-tests) when drafted will also follow the maintenance pattern (workbench/ accumulates per-case files). B6 and B7 stay as fresh-derivation — that's the load-bearing anti-ducktape constraint. --- docs/skill-optimizer-v1.4-spec.md | 2 +- .../SKILL.md | 72 +++++++--- .../iteration-protocol.md | 134 +++++++++++++----- 3 files changed, 153 insertions(+), 55 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index d76aedb..808926d 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -146,7 +146,7 @@ and dispatches the subagent with only the narrow chunks it needs. | Subagent | Sees | Does NOT see | Why | |---|---|---|---| | Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, prior `01-functionality.md` drafts | Pure research, no contamination across iterations | -| Test-case designer (step 2) | `01-functionality.md` (latest), `${OPERATOR_DIRECTIVES}` | Skill source content, prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Coverage design is at responsibility level — concrete fixture details enter the chain at step 4, not here. Reasoning from source would gerrymander tests around the source's literal phrasing instead of testing stated responsibilities | +| Test-case designer (step 2) | `01-functionality.md` (latest), current `02-test-case.md` (when extending in maintenance mode), `${OPERATOR_DIRECTIVES}` | Skill source content, archived prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Step 2 is a maintenance step — the canonical file accumulates coverage across re-runs and the subagent extends it rather than re-deriving. Coverage design stays at responsibility level. Reasoning from source would gerrymander tests around the source's literal phrasing instead of testing stated responsibilities | | Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change | Just upstream facts | | Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other test cases in the suite, the eval grader's matching logic | Fixture writing needs source detail (specific patterns, concrete violation examples). Sees source for grounding; the test case spec from step 2 constrains what the fixture should test, bounding the gerrymandering risk. Still blocked from seeing other cases (prevents copying) and grader internals (prevents grader-leak hacking) | | Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), prior `06-analysis.md` drafts | Forces it to think about the SKILL, not the SOLUTIONS; iteration-isolated | diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 620f2eb..9fa122d 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -73,11 +73,28 @@ version of the functionality understanding it was derived from. ### (b) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. The protocol tells you how to check for -an existing `02-test-case.md`, compare its `inputs.step_1_functionality` -against the current `01-functionality.md` version, archive if stale, -compute the new version, and collect operator directives. Do not -improvise; read the file. +and apply it for this step. Step 2 is a **maintenance** step (see +the protocol's "Step kinds" section) — re-runs extend the current +`02-test-case.md` rather than replacing it. Concretely, the protocol +tells you to: + +- Check whether `02-test-case.md` exists and whether + `inputs.step_1_functionality` matches the current + `01-functionality.md` version +- Copy the current file to `archive/02-test-case-v.md` (audit + trail — inert from this point) +- Pass the canonical `02-test-case.md` itself to the subagent as + load-bearing input (it extends or modifies the existing list, + preserving the user's existing `picked` choices unless directives + say otherwise) +- Bump the version number on the new canonical file + +Special case: if `01-functionality.md`'s version bumped (step 1 +re-ran with a substantively different responsibility set), existing +cases may no longer align. Surface this to the user before the +dispatch — they should decide whether to keep the existing list and +add cases, or start fresh (in which case delete the canonical file +before dispatching so the subagent treats this as iteration 1). ### (c) Dispatch the test-case-designer subagent @@ -87,33 +104,39 @@ tool (with worktree isolation if your environment supports it). Load the prompt template at [`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md), substitute the templated inputs (`${FUNCTIONALITY_PATH}`, -`${OUTPUT_PATH}`, `${VERSION}` from (b), +`${OUTPUT_PATH}`, `${EXISTING_CASES_PATH}` if a current +`02-test-case.md` exists (else empty), `${VERSION}` from (b), `${OPERATOR_DIRECTIVES}` from (b)), and dispatch. The subagent sees: -- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` only +- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` +- `${EXISTING_CASES_PATH}` — the current `02-test-case.md` (only + on re-runs; empty on iteration 1). The subagent preserves existing + cases verbatim unless a directive explicitly asks to revise a + specific case; new cases are appended per directives. - `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new requirements (may be empty) - The output path and version number The subagent does NOT see: -- Prior `02-test-case.md` drafts (the ones now in `archive/`) +- Anything under `archive/` (prior versions are inert audit trail) - Any `06-analysis.md` (existing analyses) - `07-improvement-proposal.md`, `07-validator-verdict.md` - Raw failure data, `findings.txt`, bench results - The skill's source content itself -The "no source content" constraint is load-bearing in addition to the -usual no-prior-failures one: design coverage from the skill's STATED -responsibilities, not from the source's literal phrasing. Letting the -subagent read the source lets it gerrymander tests around the -source's exact shape (testing what the prose happens to say, not -what the skill is supposed to do). +The "no source content" constraint is load-bearing: design coverage +from the skill's STATED responsibilities, not from the source's +literal phrasing. Letting the subagent read the source lets it +gerrymander tests around the source's exact shape (testing what the +prose happens to say, not what the skill is supposed to do). -The subagent writes the proposal body itself and returns a brief -summary including the top-3 cases by importance. +The subagent writes the (extended) proposal body itself and returns +a brief summary including the top-3 cases by importance and a list +of which existing case names were preserved unchanged vs which were +revised. ### (d) Confirm subagent output @@ -135,12 +158,17 @@ Three realistic responses: 1. **User picks a subset (or all).** Update the `picked` frontmatter field with the chosen case names. Commit the file. -2. **User asks for a revised proposal.** Treat their feedback as new - operator directives, return to (b) to handle iteration (which - will archive this version), and re-dispatch the subagent at (c). -3. **User picks zero.** Ask whether they're revising (treat as case - 2) or abandoning the optimization for this skill (exit honestly - without progressing to step 4). +2. **User wants additions or revisions** (e.g., "add null-input + coverage", "the case named X is too vague — split it into two"). + Treat their feedback as new operator directives, return to (b) + to apply the maintenance protocol (which archives the current + version), and re-dispatch the subagent at (c). Because this is + a maintenance step, existing cases the user already picked are + preserved unless their directive explicitly asks otherwise — the + user does not need to re-pick everything. +3. **User picks zero.** Ask whether they want a revised proposal + (treat as case 2) or are abandoning the optimization for this + skill (exit honestly without progressing to step 4). Don't auto-pick on the user's behalf — even if all cases look important, the user owns this decision (they're paying for the diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index b7e7f26..82247db 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -23,6 +23,47 @@ manually or by the auto-pilot), this protocol determines: 4. How to pass cross-iteration learnings into the subagent 5. What the subagent is and isn't allowed to look at +## Step kinds: fresh-derivation vs maintenance + +Steps in the chain come in two kinds, and the protocol's "what the +subagent sees" rule differs between them. + +**Fresh-derivation steps** produce their report from upstream +inputs + directives, with no continuity from prior versions. The +subagent does NOT see its own current or archived file — every +invocation is a derivation from scratch. This is the anti-ducktape +rule: the optimizer must not see prior optimization attempts, the +analyzer must not see prior analyses, the researcher must not +rationalize a prior report. + +| Step | Why fresh derivation | +|---|---| +| 1. investigate-functionality | Each invocation researches from source — no continuity needed | +| 3. investigate-submissions | Each invocation researches upstream — no continuity needed | +| 6. analyze-result | Must not be biased by prior analyses; load-bearing for anti-ducktape | +| 7. improve-skill | Optimizer must not see prior attempts; load-bearing for anti-ducktape | + +**Maintenance steps** maintain an accumulating state — the file +represents a growing collection that the user contributes to over +multiple invocations. Re-runs EXTEND the current file rather than +replacing it. The subagent reads the current canonical file as one +of its inputs and produces an extended version of it. + +| Step | What accumulates | +|---|---| +| 2. investigate-test-case | The ranked test-case list grows as the user adds coverage | +| 4. write-tests | The workbench directory grows as new picked cases are built | + +**The "subagent ignores its own prior output" rule applies to +fresh-derivation steps only.** For maintenance steps, the current +canonical file IS the load-bearing input — without it, the step +would lose continuity and the user's accumulated decisions (picks, +prior directives, manually-added cases) would be lost on every +re-run. + +The archive remains inert in both cases — neither kind of step +reads from `archive/`. The archive is audit-only. + ## Versioning convention Every chain report carries two frontmatter fields: @@ -48,25 +89,38 @@ and detects the mismatch. Before dispatching the step's subagent: 1. Compute the canonical report path for this step + slug. -2. Check whether the canonical path exists, and act: +2. Check whether the canonical path exists, and act according to + the step's kind (fresh-derivation or maintenance — see the + classification above): - **Path does not exist.** This is iteration 1 for this step. Set - `new_version = 1`. Proceed to "Collect operator directives" below. - - - **Path exists.** Read the existing report's frontmatter. Compare - its `inputs:` to the current versions of each upstream report - at its canonical path. Two sub-cases: - - - **All input versions match AND no operator directives are - pending.** The existing report is current; nothing to do. - Report success, hand off, exit. - - - **Any input version mismatches OR operator passed directives - requiring a re-run.** The existing report is stale (or the - operator wants fresh). Move the existing file to - `archive/--v.md`. Set - `new_version = existing_version + 1`. Proceed to "Collect - operator directives" below. + `new_version = 1`. Proceed to "Collect operator directives" + below. + + - **Path exists, all input versions match, no operator directives + pending.** The existing report is current; nothing to do. + Report success, hand off, exit. + + - **Path exists, any input version mismatches OR operator passed + directives requiring a re-run.** Behavior depends on the + step's kind: + + - **Fresh-derivation step:** copy the existing file to + `archive/--v.md` (audit trail); + the subagent will derive a new file from upstream + directives + only, with no reference to the prior or archived content. Set + `new_version = existing_version + 1`. + + - **Maintenance step:** copy the existing file to + `archive/--v.md` (audit trail); + the subagent receives the current canonical file as an input + and produces an extended/modified version of it. Set + `new_version = existing_version + 1`. The archive copy is + inert from this point on — the subagent operates on the + canonical file content, not the archive. + + In both cases the canonical path is overwritten with the new + version; the archive copy preserves the prior version for audit. ## Collect operator directives @@ -111,21 +165,37 @@ placeholders): Two rules every reasoning subagent must follow: -1. **Ignore your own prior output.** Do not read anything under - `archive/--v*.md`. That output is the past - contaminated reasoning the iteration mechanism exists to wall - off. Your job on iteration N is to derive from the current - upstream reports plus `${OPERATOR_DIRECTIVES}`, NOT to "improve" - the prior version. - -2. **Read upstream reports at their latest versions only.** The - canonical paths (`-.md`, no version suffix) always - point at the latest version. Don't walk `archive/` for any - reason. - -Both rules are enforced in each subagent's prompt template. If you -find yourself wanting to violate either — that's the failure mode -the iteration mechanism prevents. Don't. +1. **Never read from `archive/`.** Archived versions are inert audit + trail. Don't walk `archive/--v*.md` for any reason. + The current canonical file (and only the current canonical file) + is the load-bearing input for a maintenance step; for a + fresh-derivation step, neither the canonical nor the archive is + read. + +2. **Read upstream reports at their latest versions only.** Upstream + canonical paths always point at the latest version. Same archive + rule applies to upstream files — never walk their archives. + +A third rule that depends on the step's kind: + +- **For fresh-derivation steps (1, 3, 6, 7): do NOT read your own + step's canonical file.** Your job is to derive a new report from + upstream + directives, with no reference to what was produced + before. This is the load-bearing anti-ducktape constraint — the + optimizer must not see prior attempts, the analyzer must not see + prior analyses, etc. + +- **For maintenance steps (2, 4): DO read your own step's canonical + file (when it exists).** Your job on a re-run is to extend or + modify the current accumulating state per `${OPERATOR_DIRECTIVES}`, + preserving existing entries the user has invested in (test-case + picks, prior coverage decisions, etc.) unless a directive + explicitly says to revise a specific entry. + +Each subagent prompt template enforces the appropriate version of +the third rule for its step. If you find yourself wanting to +violate any of these — that's the failure mode the iteration +mechanism prevents. Don't. ## Cascading staleness From eb6c53ecdc1dd30601f8100522c05c3a91b4544b Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 09:35:55 -0500 Subject: [PATCH 020/121] feat(skill-optimizer-write-tests): SKILL.md body MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 4 of the chain — takes the picked cases from 02-test-case.md and dispatches test-writer subagents in parallel (one per case) to build concrete workspace files + graders. Runs a smoke check against hand-crafted GOOD/BAD/EMPTY fixtures before declaring done. Follows the established template (B1/B2/B3 shape) with several B4-specific additions: - **Maintenance step** (per the recent iteration-protocol update). workbench/ accumulates as new picks get built. The (b) iteration step diffs picked vs prior built_cases and only dispatches test-writers for NEW picks or for cases that directives flag for revision. Existing builds are preserved verbatim. - **Parallel per-case dispatch** in step (d). All test-writers emitted in a single message so they run concurrently. Each sees only its case spec + 01-functionality.md + skill source; does NOT see other cases, other graders, prior failures, or anything under archive/. The per-case isolation prevents both cross-fixture homogenization and grader-leak hacking. - **Source content access** (per the recent spec update). Test writers are the one step where the skill source is in-scope for a generative subagent. Bounded by the per-case spec from step 2. - **Smoke check** (step (e)). New responsibility not in other steps. Verifies each grader against GOOD/BAD/EMPTY fixtures the test-writer also produces. last_smoke_check_passed in frontmatter is only set when all graders pass. - **Two outputs**: 04-tests-plan.md (the meta-report) AND the workbench/ directory (the artifacts the run-bench step executes against). Maintenance protocol covers both via archive/04-tests-plan-vN.md and archive/workbench-vN/. - **User-gate at step (c)** for the planned workbench structure before dispatching. Three response cases (approve / add revisions via directives / reject and abandon-or-replan). - **6 internal steps**: (a) confirm prerequisites, (b) handle iteration with diff logic, (c) plan + user gate, (d) parallel dispatch, (e) smoke check, (f) assemble + commit + hand off. 350 lines, description 492 chars. Larger than B1/B2/B3 because the extra responsibilities (parallel dispatch, smoke check, maintenance diff) each earn their lines. --- skills/skill-optimizer-write-tests/SKILL.md | 347 +++++++++++++++++++- 1 file changed, 344 insertions(+), 3 deletions(-) diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 378d53e..feb19c0 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -1,9 +1,350 @@ --- name: skill-optimizer-write-tests -description: Use when the user wants to build / implement the eval workbench for a skill — plans the workbench, asks for user confirmation, then dispatches parallel narrow-context subagents to write each test case + grader. +description: Use when the user wants to build or implement the eval workbench for a skill — phrases like "build the tests", "implement the workbench", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-investigate-test-case` has produced `02-test-case.md` with a picked subset, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the test-case proposals into concrete fixtures + graders should trigger this. --- # skill-optimizer-write-tests - - +Step 4 of the skill-optimizer chain. Takes the picked test cases +from step 2, dispatches one test-writer subagent per case (in +parallel) to build the concrete workspace files + grader scripts, +runs a smoke check to verify each grader against hand-crafted +GOOD/BAD/EMPTY fixtures, and writes +`docs/skill-optimizer//04-tests-plan.md` plus the +`workbench/` directory the run-bench step will execute against. + +## Before you start + +Three load-bearing pieces of context to load NOW, before the +workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation. Step 4 is a + **maintenance** step (see the protocol's "Step kinds" section) + — re-runs extend the workbench rather than replacing it. + Workflow step (b) below requires the protocol; loading it now + lets you execute that step deterministically. + +2. **You will dispatch test-writer subagents (in parallel, one per + picked case); you do NOT write workspace files or grader scripts + yourself in this session.** Workflow step (d) is the dispatch. + The rationale is in "Why limited-context dispatch matters" below + — read it if you're tempted to skip the dispatch. + +3. **Each test-writer subagent sees the skill's source content.** + This is intentional and is the one step in the chain where the + source is in-scope for a generative subagent. The constraint + comes from the test case spec from step 2 (which fixes what the + fixture should test); source access is needed for concrete + violation patterns and realistic fixture content. + +Throughout this document, "step 1" through "step 7" (no parens) +refer to skills in the chain (this skill is step 4; +`investigate-test-case` is step 2; etc.). Internal workflow steps +within THIS skill are labelled "(a)" through "(f)" to avoid the +collision. + +## What you produce + +Two outputs at `docs/skill-optimizer//`: + +1. **`04-tests-plan.md`** — the meta-report tracking which cases + have been built, with per-case implementation notes. + +2. **`workbench/`** — the directory the `run-bench` step executes + against. Contains `workbench/suite.yml` (suite manifest), per-case + workspace files, per-case grader scripts, and smoke-check + fixtures. + +`04-tests-plan.md` frontmatter: + +```yaml +--- +version: 1 +inputs: + step_1_functionality: + step_2_test_case: +built_cases: + - + - +last_smoke_check_passed: +--- +``` + +**`built_cases`** — the set of case names that currently have +workspace files + graders in `workbench/`. Re-runs add new entries +here as new picks get built; existing entries stay unless a +directive explicitly requests revision. + +**`last_smoke_check_passed`** — set when the smoke check (step (e)) +passes. If a re-run modifies a case and the smoke check fails, this +field stays at the prior timestamp until the smoke check passes +again on the updated grader. + +For the body template (per-case implementation notes — workspace +contents, expected agent behavior, grader logic, smoke-check +fixtures), see +[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md). +Each test-writer subagent contributes the section for its own case; +the operator session assembles them into `04-tests-plan.md`. + +For the workbench schema itself (`suite.yml` structure, grader +contract, smoke-check format), see +[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md) +— that's the canonical workbench reference used by the run-bench +step. + +## Workflow + +### (a) Confirm prerequisites + +Three prerequisites: + +1. `docs/skill-optimizer//01-functionality.md` exists with + valid frontmatter. +2. `docs/skill-optimizer//02-test-case.md` exists with valid + frontmatter and a non-empty `picked` field. +3. Every name in `picked` appears as a case definition in the body + of `02-test-case.md` (no stale references to cases that were + removed in a revision). + +If any prerequisite fails, surface the issue to the user — don't +attempt to derive missing cases or guess. Record both upstream +versions; you'll pass them as `inputs.step_1_functionality` and +`inputs.step_2_test_case` later. + +### (b) Handle iteration (maintenance step) + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply the maintenance-step flow for this step: + +- Copy current `04-tests-plan.md` and the current `workbench/` + directory to `archive/04-tests-plan-v.md` and + `archive/workbench-v/` respectively (audit trail — inert from + this point). +- Compute the new version (`N+1`). +- Diff the `picked` set in `02-test-case.md` against the prior + `built_cases` (from the just-archived plan). Three categories: + - **New picks** (in `picked` but not in `built_cases`) — these + need test-writer dispatches in step (d). + - **Existing builds** (in both) — preserve their workbench files + and plan entries unchanged, unless `${OPERATOR_DIRECTIVES}` + explicitly names a case to revise. + - **Removed picks** (in `built_cases` but no longer in + `picked`) — the user de-picked these in a later step 2 revision. + Leave the workbench files in place but mark them removed from + `built_cases` in the new plan (the run-bench step's suite.yml + will exclude them). + +If `02-test-case.md`'s version bumped substantively and many cases +were renamed or reframed, the diff may not match cleanly. Surface +this to the user before proceeding — they may want to delete the +canonical `04-tests-plan.md` + `workbench/` to start fresh, or +selectively rebuild specific cases. + +### (c) Plan workbench structure + user gate + +Before dispatching test-writers, sketch the planned workbench +structure: which case directories will be created, what the +`suite.yml` entries will look like, which existing files are +preserved. Show this to the user. Ask: + +> Here's the planned workbench layout for the picked cases. New +> cases I'll dispatch test-writers for: [list]. Existing cases +> preserved: [list]. Any cases I should revise instead of preserve? +> Anything to change before I dispatch? + +Three realistic responses: + +1. **User approves.** Proceed to (d) with the dispatch set as + planned. +2. **User adds revisions** (e.g., "the grader for case X is too + strict — relax the matching"). Treat as new operator directives, + mark the named cases for revision, and proceed to (d) with the + updated dispatch set. +3. **User rejects the plan structure.** Likely they want a + different overall workbench shape — surface this and ask whether + to abandon (loop back to step 2 for re-planning) or to retry + the plan with their feedback as directives. + +### (d) Dispatch test-writer subagents (parallel, one per case) + +**Do NOT write workspace files or graders yourself in this +session.** For each case in the dispatch set (new picks + revisions +flagged in (c)), dispatch a test-writer subagent via the `Agent` +tool, in parallel — emit all dispatches in a single message so they +run concurrently. + +Load the prompt template at +[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) +and substitute per-case inputs: + +- `${CASE_NAME}` — this case's name from `02-test-case.md` +- `${CASE_SPEC_PATH}` — path to a slice of `02-test-case.md` + containing only this case's entry (the operator session extracts + this from the full report so the subagent doesn't see other + cases) +- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` +- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local skill + path), so the subagent can ground the fixture in concrete + violation patterns +- `${OUTPUT_WORKBENCH_DIR}` — `workbench/cases//` +- `${VERSION}` from (b) +- `${OPERATOR_DIRECTIVES}` from (b) (case-level revision hints if + this case is being revised; empty otherwise) + +Each test-writer subagent sees: + +- Its single case spec (the slice extracted in (d)) +- `01-functionality.md` +- The skill source content (vendored or local) +- `${OPERATOR_DIRECTIVES}` for its case + +The test-writer subagent does NOT see: + +- Other test cases' specs (prevents copying fixtures across cases) +- Other test cases' graders (prevents grader-leak hacking — looking + at how another grader matches and writing a fixture that exploits + the same pattern) +- The eval grader's matching internals (same reason) +- Anything under `archive/` +- `06-analysis.md`, `07-improvement-proposal.md`, raw failure data, + prior optimizer attempts + +The "no other cases" constraint is load-bearing: each case is built +in isolation so cross-case patterns don't subtly homogenize the +fixtures. The "no grader internals" constraint prevents fixtures +from being gerrymandered to the grader's specific matching rules. + +Each test-writer subagent writes: + +- Workspace files at `workbench/cases//workspace/` +- Grader script at `workbench/cases//grader.mjs` (or + `.py` per the workbench schema) +- Smoke-check fixtures at `workbench/cases//smoke/` + (GOOD/BAD/EMPTY findings.txt examples the grader should classify + correctly) +- A short per-case section for the operator to fold into + `04-tests-plan.md` + +The subagent returns a brief summary: case name, what the fixture +tests, grader logic in one line, smoke-check result for its own +case. + +### (e) Run smoke check + +After all test-writer dispatches return, run the smoke check +against each grader: + +```bash +node skills/skill-optimizer/references/scripts/smoke-check.mjs \ + workbench/cases//grader.mjs \ + workbench/cases//smoke/ +``` + +(Adapt the exact command to the workbench schema's smoke-check +runner.) + +Expected: + +- The GOOD fixture passes the grader. +- The BAD fixture fails the grader. +- The EMPTY fixture fails the grader. + +If any grader fails the smoke check, surface the case to the user. +Two realistic responses: + +1. **User asks to re-dispatch the test-writer for that case** (with + the smoke-check failure as a directive). Treat as a single-case + revision, return to (d) with just that case in the dispatch set. +2. **User decides to remove the case from `picked`** (in + `02-test-case.md`) and re-run B4. That's a backward trigger — + per the iteration protocol's re-run authorization rule, you + surface the option and let the user invoke step 2 explicitly. + +Update `last_smoke_check_passed` in the frontmatter only after all +graders pass. + +### (f) Assemble plan + commit + hand off + +Assemble the per-case sections returned by the test-writer +subagents into `04-tests-plan.md` (with the frontmatter from (b)). +Update `workbench/suite.yml` to include all entries in `built_cases` +(excluding any that were removed in the (b) diff). + +Commit both `04-tests-plan.md` and the `workbench/` directory. + +Handoff message, verbatim: + +> Next, invoke `skill-optimizer-run-bench` to measure baseline +> performance against the workbench. + +## Why limited-context dispatch matters + +The test-writer subagents are dispatched per-case for two reasons: + +First, **fixture isolation**. A single subagent writing all +fixtures would notice patterns across cases ("they all check for +absence violations — let me write a unified fixture") and +inadvertently homogenize them. Per-case isolation forces each +fixture to be a representative instance of its responsibility, +designed without knowledge of how sibling cases are shaped. + +Second, **no grader-leak hacking**. If a subagent sees how another +grader matches (regex pattern, JSON path, etc.), it can write a +fixture that incidentally satisfies the OTHER grader too — making +the eval results look correlated when they're not. Per-case +isolation prevents this. + +The source access concession (test-writer subagents DO see the +skill source) is bounded by the case spec from step 2: the spec +fixes WHAT the fixture should test, and the source provides the +HOW (concrete violation patterns). Without source, the subagent +would invent generic patterns that may not actually trigger the +skill's rules. + +If you find yourself thinking "the cases are similar; I'll write +them all myself with one prompt and save dispatches" — that's the +homogenization failure mode this chain is built to prevent. +Dispatch in parallel. + +## Edge cases + +- **A test-writer subagent reports BLOCKED** (e.g., the case spec + is too abstract to derive a fixture from) — surface to the user + with the subagent's reasoning. The fix is typically a step 2 + re-run with a more specific case spec; per the iteration + protocol's re-run authorization rule, the user invokes step 2 + explicitly. +- **A grader fails the smoke check** — see step (e)'s handling. +- **Smoke check passes for individual cases but the suite.yml + itself is malformed** — surface as a workbench-schema error; + fix the suite.yml directly (operator session task; not a + test-writer concern). +- **A picked case is genuinely unimplementable in a static + workbench** (e.g., requires real-time API access the workbench + can't provide) — the test-writer should return BLOCKED with this + reasoning. Surface to the user; the case may need to be removed + from `picked` in step 2. + +## Iteration behavior + +Step 4 is a maintenance step — the workbench accumulates as new +picks get built. Re-run triggers specific to this step: + +- `02-test-case.md`'s `picked` field changed (new picks added, + some removed) +- A specific case's smoke check failed and needs the test-writer + to revise (single-case re-dispatch via directives) +- Step 5 (`run-bench`) showed all cases pass on baseline (too + easy) or all fail (too hard) — directives like "make case X + harder" or "loosen case Y's grader" trigger case-level revisions +- `02-test-case.md`'s version bumped substantively (responsibilities + reframed) — operator surfaces and the user decides whether to + preserve existing builds or start fresh + +General iteration mechanics — archive, version bump, operator +directives, cascading staleness, maintenance vs fresh-derivation — +live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. From 3fc554b3b6755854a13ae865096c10d9fc027fe4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 09:51:58 -0500 Subject: [PATCH 021/121] =?UTF-8?q?fix(skill-optimizer-write-tests):=20cor?= =?UTF-8?q?rect=20the=20smoke-check=20section=20=E2=80=94=20no=20centraliz?= =?UTF-8?q?ed=20script=20exists?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The prior wording referenced a path that didn't exist (`skills/skill-optimizer/references/scripts/smoke-check.mjs`). The actual codebase doesn't have a centralized smoke-check runner — each existing workbench has its own `checks/smoke-graders.mjs` specific to its graders (see e.g. examples/workbench/agent-browser/checks/smoke-graders.mjs). Rewrote step (e) to be honest about this: - Smoke check is per-workbench, not centralized. - The test-writer subagent produces the smoke-check artifact alongside its grader. Shape follows the workbench schema docs. - Two reasonable shapes described (per-case runner vs workbench-level aggregator); the test-writer prompt template (Phase C) will pick one and apply consistently. - Reference example pointed at the actual existing agent-browser/checks/smoke-graders.mjs. Conceptual contract unchanged (GOOD passes, BAD/EMPTY fail); only the execution mechanics corrected. --- skills/skill-optimizer-write-tests/SKILL.md | 43 ++++++++++++++------- 1 file changed, 28 insertions(+), 15 deletions(-) diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index feb19c0..6755922 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -233,19 +233,32 @@ case. ### (e) Run smoke check -After all test-writer dispatches return, run the smoke check -against each grader: - -```bash -node skills/skill-optimizer/references/scripts/smoke-check.mjs \ - workbench/cases//grader.mjs \ - workbench/cases//smoke/ -``` - -(Adapt the exact command to the workbench schema's smoke-check -runner.) - -Expected: +The smoke check is per-workbench, not centralized. Each test-writer +subagent produces, alongside its grader, a smoke-check artifact +that validates the grader against hand-crafted GOOD/BAD/EMPTY +fixtures. The shape of that artifact follows the workbench schema +documented in +[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md). +Existing examples to follow: +[`examples/workbench/agent-browser/checks/smoke-graders.mjs`](../../examples/workbench/agent-browser/checks/smoke-graders.mjs) +shows a workbench-level smoke checker that runs all of that +workbench's graders against GOOD/BAD fixtures. + +Two reasonable shapes (the test-writer prompt template should pick +one and apply it consistently): + +- **Per-case runner**: each test-writer produces + `workbench/cases//checks/smoke.mjs` that exercises + that case's grader against its own GOOD/BAD/EMPTY fixtures. + Operator runs each runner. +- **Workbench-level aggregator**: the test-writers each contribute + their fixtures + grader to a shared location, and the operator + composes (or the workbench schema provides) a top-level + `workbench/checks/smoke-graders.mjs` that runs all of them at + once. + +After all test-writer dispatches return, run the smoke-check +artifact(s) the test-writers produced. Expected for each case: - The GOOD fixture passes the grader. - The BAD fixture fails the grader. @@ -258,8 +271,8 @@ Two realistic responses: the smoke-check failure as a directive). Treat as a single-case revision, return to (d) with just that case in the dispatch set. 2. **User decides to remove the case from `picked`** (in - `02-test-case.md`) and re-run B4. That's a backward trigger — - per the iteration protocol's re-run authorization rule, you + `02-test-case.md`) and re-run step 4. That's a backward trigger + — per the iteration protocol's re-run authorization rule, you surface the option and let the user invoke step 2 explicitly. Update `last_smoke_check_passed` in the frontmatter only after all From a7cac467744dabca82f233c3df95a5d0c2092afb Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 10:14:18 -0500 Subject: [PATCH 022/121] feat(skill-optimizer-run-bench): SKILL.md body MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 5 of the skill-optimizer chain. Thin operator-driven CLI step — no subagent dispatch. Two outputs: timestamped raw bench data under 05-bench-results// (preserved naturally, exempt from the standard archive flow) and a versioned 05-bench-summary.md (follows the iteration protocol; archive-and-version on re-run). The summary is the entry point step 6 reads: per-model and per-case pass rates plus a failed-case pointer list into the raw output. The analyzer subagent at step 6 walks that list for trace and findings detail; the summary itself stays short. Handoff branches on overall pass rate: failures route to analyze-result; all-pass surfaces the choice to the user (accept that the picked set didn't expose a weakness, or re-run step 2 with a "make it harder" directive). Chain skills don't auto-invoke. --- skills/skill-optimizer-run-bench/SKILL.md | 220 +++++++++++++++++++++- 1 file changed, 217 insertions(+), 3 deletions(-) diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 03f66df..9e7e3fa 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -1,9 +1,223 @@ --- name: skill-optimizer-run-bench -description: Use when the user wants to run the eval suite for a skill and capture results — thin wrapper around the skill-optimizer CLI's run-suite command. +description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once a `workbench/` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. --- # skill-optimizer-run-bench - - +Step 5 of the skill-optimizer chain. Invokes the skill-optimizer CLI's +`run-suite` command against the workbench built in step 4, captures the +raw results under a timestamped directory, and writes a small versioned +summary report that step 6 (`analyze-result`) reads. No subagent +dispatch — this is a thin operator-driven CLI step. + +## Before you start + +One load-bearing piece of context to load NOW, before the workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation. Step 5 is a + slightly unusual case because the bench's raw output is timestamped + rather than versioned (each run preserved naturally under its own + directory) — but the summary report `05-bench-summary.md` IS + versioned and follows the protocol's archive-and-version mechanics. + Workflow step (b) below requires it. + +Throughout this document, "step 1" through "step 7" (no parens) refer +to skills in the chain (this skill is step 5; `write-tests` is step 4; +`analyze-result` is step 6). Internal workflow steps within THIS skill +are labelled "(a)" through "(e)" to avoid the collision. + +## What you produce + +Two outputs at `docs/skill-optimizer//`, where `` matches +the slug from step 1's report: + +1. **`05-bench-results//`** — the raw run output from the + CLI: `suite-result.json`, per-trial `trace.jsonl`, per-trial + `findings.txt`, and any preserved workspaces. Timestamped per the + invocation; old runs are NEVER overwritten. The protocol's archive + convention does not apply here — each timestamped directory is + itself the archive. + +2. **`05-bench-summary.md`** — a versioned report with frontmatter: + + ```yaml + --- + version: + inputs: + step_4_tests: + bench_results_path: 05-bench-results// + overall_pass_rate: + --- + ``` + + The body is a small aggregate that step 6 reads as the entry point: + the per-model pass rates, per-case pass rates, a list of cases + where at least one trial failed (so the analyzer knows where to + focus), and a pointer to the timestamped raw directory. The + summary does NOT include trace excerpts or findings detail — those + stay in the raw output and the analyzer reads them directly when + it needs them. + +## Workflow + +### (a) Confirm prerequisites + +`docs/skill-optimizer//04-tests-plan.md` must exist and have +valid frontmatter, and `docs/skill-optimizer//workbench/` must +exist with a `suite.yml` and per-case dirs (per step 4's output +contract). If either is missing, tell the user to complete step 4 +first and stop here. + +Also confirm the vendored skill source exists at +`docs/skill-optimizer//vendored-skill/` (per step 1's setup); +the workbench's `suite.yml` points at it as the skill under test. + +Record `04-tests-plan.md`'s `version` field — you'll pass it as +`inputs.step_4_tests` in the summary. + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it to the summary file. Step 5's behavior: + +- For the **raw bench output** at `05-bench-results//`: + always write to a fresh timestamped directory (`date +%Y%m%dT%H%M%S` + or similar). No archive logic — each run is its own preserved + artifact. Never overwrite a prior timestamped dir. + +- For the **summary file** at `05-bench-summary.md`: apply the + protocol's standard archive-and-version flow. + - If the summary doesn't exist yet, set `new_version = 1`. + - If it exists and `inputs.step_4_tests` matches the current + `04-tests-plan.md` version AND the operator did not pass a + re-run directive, the existing summary is current — exit + honestly and tell the user the latest results are still good. + - Otherwise, copy `05-bench-summary.md` to + `archive/05-bench-summary-v.md` (audit trail) + and set `new_version = existing_version + 1`. The new summary + will be written from the new raw run, with no reference to the + prior summary's body. + +The summary is conceptually fresh-derivation — its body is a +mechanical write-up of the new raw bench run, not an extension of +the prior summary. There is no subagent, but the same anti-ducktape +principle applies in spirit: write what this run shows, not what the +prior summary said. + +### (c) Run the bench + +Invoke the skill-optimizer CLI directly: + +```bash +TIMESTAMP=$(date -u +%Y%m%dT%H%M%SZ) +OUT_DIR="docs/skill-optimizer//05-bench-results/${TIMESTAMP}" +mkdir -p "${OUT_DIR}" + +npx tsx /src/cli.ts run-suite \ + docs/skill-optimizer//workbench/suite.yml \ + --out "${OUT_DIR}" \ + --trials 3 +``` + +`--trials 3` is the chain's default — enough trials to distinguish +flaky (1-of-3) from systematic (2-of-3 or 3-of-3) failures at step 6. +The user may request a different trial count; honor it. Models come +from `suite.yml` (per the project invariant: `run-suite` does NOT +take a `--models` override). + +The CLI requires `OPENROUTER_API_KEY` to be set for real model runs. +If the user hasn't set it, the CLI will fail fast — surface the +error to the user; don't try to recover. + +Stream the CLI's stdout/stderr to the user so they can see progress +(this is a long-running step; bench runs of meaningful size take +minutes to hours). + +### (d) Write the summary + +When the CLI completes, parse `${OUT_DIR}/suite-result.json` and +write `05-bench-summary.md` with the frontmatter from "What you +produce" plus a body containing: + +- **Overall:** total trials, passed, failed, overall pass rate. +- **Per model:** for each model in `suite.yml`, the trial count and + pass rate. +- **Per case:** for each case, the trial count, pass rate, and a + one-line note if at least one trial failed (so step 6 knows where + to look). +- **Failed-case pointer list:** explicit list of case IDs where any + trial failed, with paths to their `trace.jsonl` and `findings.txt` + under `${OUT_DIR}`. +- **Raw output:** the value of `bench_results_path` from the + frontmatter (relative path to the timestamped dir). + +The summary is the entry point step 6 reads. Step 6's analyzer +subagent will follow the failed-case pointer list into the raw +output for trace and findings detail; the summary itself stays +short (an aggregate, not a dump). + +### (e) Hand off + +Read the overall pass rate from the summary you just wrote. Two +handoff messages depending on the result: + +If `overall_pass_rate < 1.0`: + +> Bench complete. Overall pass rate: `%`. Some trials failed — +> invoke `skill-optimizer-analyze-result` to diagnose whether the +> failures point to a structural weakness in the skill. + +If `overall_pass_rate == 1.0`: + +> Bench complete. All trials passed against the current skill. Two +> realistic paths from here: (1) accept that the picked test set +> doesn't expose any weakness in the current skill (exit honestly); +> (2) revise the test set to be harder — re-run step 2 with a +> directive like "the prior cases were too easy; propose harder +> ones" and walk the chain forward again. + +Don't auto-invoke step 6 in the all-pass case — there's nothing for +it to analyze, and the chain skills never invoke each other on their +own. Surface the choice to the user. + +## Edge cases + +- **`04-tests-plan.md` missing or `workbench/` empty** — tell the + user to run step 4 first; don't try to construct a workbench + yourself. +- **`OPENROUTER_API_KEY` not set** — the CLI fails fast; surface the + error and tell the user to set the env var. Don't fall back to a + mock or skip the bench. +- **Docker image missing** — the CLI's default image is + `skill-optimizer-workbench:local`. If it's not built, tell the + user to run `docker build -t skill-optimizer-workbench:local -f + docker/workbench-runner.Dockerfile .` from the repo root. +- **All trials in `suite-result.json` errored (no graded results)** + — likely a workbench misconfiguration or environmental failure, + not a skill weakness. Write the summary honestly (overall pass + rate undefined, error reasons recorded) and surface the situation + to the user before handing off to step 6 — they should decide + whether to fix the workbench (back to step 4 with directives) or + treat this as the skill itself crashing the harness. + +## Iteration behavior + +Step 5 is re-runnable. Re-run triggers specific to this step: + +- `04-tests-plan.md`'s version bumped (step 4 re-ran with new or + revised picked cases) — the prior bench measured a different + workbench +- User wants fresh results against the same workbench (e.g., + flakiness suspicion, model availability changed, model list in + `suite.yml` changed) +- Step 6 (`analyze-result`) flagged a result as inconclusive due to + too few trials — user may re-run with higher `--trials` + +General iteration mechanics — archive, version bump, cascading +staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +The raw bench output is exempt from the standard archive flow (each +timestamped directory is its own preservation); only the summary +follows the protocol. From 0b2d1c2195ae31c5e7dcffda3b41f221179c2960 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 10:34:31 -0500 Subject: [PATCH 023/121] feat(v1.4-chain): case-rename convention + defer partial re-bench MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two related coordination changes across the chain, plus a deferred limitation note. Rename convention for revised cases at step 2: when a directive asks to revise an existing case (rather than add a new one), the test-case-designer subagent appends a version suffix — `case-x` → `case-x-v2` → `case-x-v3`. Bare name is implicit v1. This gives step 4's diff logic a deterministic signal that the revised case needs a fresh test-writer dispatch (without the rename, the name-match would say "already built" and skip rebuilding the case whose spec actually changed). Encoded in B2's user-gate step and referenced from B4's diff logic so the v-suffixed names are not surprising downstream. Partial re-bench at step 5: deferred. CLI's run-suite does not accept a case filter, so step 5 always re-measures the full workbench. Documented as a known limitation with a roadmap pointer. Operator escape hatch (manual run-case + splice into 05-bench-results//) noted as outside the chain. --- .../SKILL.md | 28 ++++++++++++++----- skills/skill-optimizer-run-bench/SKILL.md | 16 +++++++++++ skills/skill-optimizer-write-tests/SKILL.md | 14 ++++++---- 3 files changed, 46 insertions(+), 12 deletions(-) diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 9fa122d..ea20e1c 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -159,13 +159,27 @@ Three realistic responses: 1. **User picks a subset (or all).** Update the `picked` frontmatter field with the chosen case names. Commit the file. 2. **User wants additions or revisions** (e.g., "add null-input - coverage", "the case named X is too vague — split it into two"). - Treat their feedback as new operator directives, return to (b) - to apply the maintenance protocol (which archives the current - version), and re-dispatch the subagent at (c). Because this is - a maintenance step, existing cases the user already picked are - preserved unless their directive explicitly asks otherwise — the - user does not need to re-pick everything. + coverage", "the case named X is too vague — split it into two", + "make case Y harder"). Treat their feedback as new operator + directives, return to (b) to apply the maintenance protocol + (which archives the current version), and re-dispatch the + subagent at (c). Because this is a maintenance step, existing + cases the user already picked are preserved unless their + directive explicitly asks otherwise — the user does not need to + re-pick everything. + + **Case-rename convention for revisions** — when a directive + asks to revise an existing case (rather than add a new one), the + test-case-designer subagent appends a version suffix to the + case's name: `case-x` → `case-x-v2` → `case-x-v3`. The bare + name is implicit v1; revisions always get a numeric suffix. The + subagent updates the body of `02-test-case.md` to use the new + name; the operator session updates `picked` to swap the old + name for the new one. This convention is load-bearing for + step 4's diff logic — without a rename, step 4 would see "case + already built" by name match and skip rebuilding. The convention + gives step 4 a deterministic signal that this case needs a + fresh test-writer dispatch. 3. **User picks zero.** Ask whether they want a revised proposal (treat as case 2) or are abandoning the optimization for this skill (exit honestly without progressing to step 4). diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 9e7e3fa..dcb956c 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -215,6 +215,22 @@ Step 5 is re-runnable. Re-run triggers specific to this step: - Step 6 (`analyze-result`) flagged a result as inconclusive due to too few trials — user may re-run with higher `--trials` +By default, a re-run measures the **entire** workbench — even cases +that didn't change at step 4. This gives unchanged cases fresh +trial samples (useful for distinguishing flaky from systematic at +step 6) and keeps the summary internally comparable. The cost is +"all cases × trials" per re-run. + +**Known limitation (deferred):** the CLI's `run-suite` does not +currently accept a case filter, so there is no first-class +"re-bench only the changed cases" mode at step 5. Operators with +expensive suites can run `run-case` manually for changed cases and +splice results into the prior `05-bench-results//` dir, but +that's an outside-the-chain escape hatch, not a supported flow. +Adding a case filter to `run-suite` (and threading it through step 5) +is on the workbench's roadmap; until then, expect re-runs to remeasure +the full suite. + General iteration mechanics — archive, version bump, cascading staleness — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 6755922..14416b8 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -126,15 +126,19 @@ and apply the maintenance-step flow for this step: - Diff the `picked` set in `02-test-case.md` against the prior `built_cases` (from the just-archived plan). Three categories: - **New picks** (in `picked` but not in `built_cases`) — these - need test-writer dispatches in step (d). + need test-writer dispatches in step (d). Revisions from step 2 + surface here too: per the rename convention, a revised case + arrives with a version-suffixed name (e.g., `case-x-v2`), so + the diff sees it as a fresh pick and dispatches normally. - **Existing builds** (in both) — preserve their workbench files and plan entries unchanged, unless `${OPERATOR_DIRECTIVES}` explicitly names a case to revise. - **Removed picks** (in `built_cases` but no longer in - `picked`) — the user de-picked these in a later step 2 revision. - Leave the workbench files in place but mark them removed from - `built_cases` in the new plan (the run-bench step's suite.yml - will exclude them). + `picked`) — the user de-picked these in a later step 2 revision, + or they were superseded by a renamed revision (e.g., `case-x` + superseded by `case-x-v2`). Leave the workbench files in place + but mark them removed from `built_cases` in the new plan (the + run-bench step's suite.yml will exclude them). If `02-test-case.md`'s version bumped substantively and many cases were renamed or reframed, the diff may not match cleanly. Surface From 59fe82e69063a65539ad4d51af83daf8431af49c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 12:22:57 -0500 Subject: [PATCH 024/121] feat(v1.4-iteration): git-native + filesystem-as-state redesign MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Replace the version-field + archive-folder iteration model with git as the history mechanism and the filesystem itself as the current state. The prior protocol was reimplementing git in frontmatter — version: ints, inputs.step_N: lineage tracking, archive/-name-v.md copies — and adding accidental complexity across the chain. Iteration-protocol rewrite: - Drop version field, archive/ subdir, inputs.step_N int tracking - Staleness detection: git log -1 --format=%ct mtime comparison - Two step kinds preserved (fresh-derivation vs maintenance), with the constraint reframed in terms of "does the subagent read its own canonical file" — fresh-derivation says no (anti-ducktape), maintenance says yes (filesystem IS the state) - Subagent constraint added: don't walk git history of any tree file — for fresh-derivation this is the anti-ducktape guarantee, for maintenance this prevents reasoning from prior states - Safe destructive edits: operator session commits a checkpoint before maintenance-step rebuilds so git history has a clean before/after breakpoint Spec doc updates: - State layout: tests///{spec.yaml, workspace, grader, smoke} replaces 02-test-case.md + 04-tests-plan.md + workbench/. Filesystem-as-state — no picked: [] or built_cases: [] arrays anywhere - B2 output: 00-test-proposals.md (audit) + tests//spec.yaml per functionality with picked: true|false in each - B4 output: tests/// probe folders + generated tests/suite.yml. Per-probe parallel test-writer dispatch - B5 input: tests/suite.yml. Two outputs: timestamped raw + single-canonical 05-bench-summary.md - Subagent constraints table updated for every reasoning subagent to reflect git-history-off-limits rule - Per-step iteration-behavior sections rewritten for the new model Undoes the -v2 rename convention added 30 min ago (no longer needed — filesystem-as-state means revising a spec.yaml in place is naturally detected, and B4 can just rebuild on directive without any special name mangling). --- docs/skill-optimizer-v1.4-spec.md | 397 ++++++++++-------- .../iteration-protocol.md | 351 +++++++--------- 2 files changed, 387 insertions(+), 361 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 808926d..baa8019 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -107,20 +107,42 @@ patterns in. See plan §"Task A4" for details. ```text docs/skill-optimizer// 01-functionality.md - 02-test-case.md - 03-submissions.md # only if step 3 ran - 04-tests-plan.md - workbench/ # eval suite produced by step 4 - 05-bench-results// # produced by step 5 + 00-test-proposals.md # B2's audit report (ranking + reasoning) + tests/ + / # one folder per B2-proposed functionality + spec.yaml # B2 wrote: picked, importance, suggested_probes + / # one folder per B4-built probe of this functionality + spec.yaml # B4 wrote: probe-level intent + workspace/ # fixture files the agent sees + grader.mjs # grading script + smoke/ # GOOD/BAD/EMPTY findings.txt fixtures + / + ... + / + spec.yaml + ... + suite.yml # B4 generates from picked-functionality probes + 03-submissions.md # only if step 3 ran + 05-bench-results// # raw bench output, timestamped per run + 05-bench-summary.md # B5 writes: aggregate + pointer to latest 06-analysis.md 07-improvement-proposal.md 07-validator-verdict.md - vendored-skill/ # the source skill, read-only after fetch + vendored-skill/ # the source skill, read-only after fetch ``` `` is `--` for upstream skills, or `` for local skills. +**Filesystem-as-state, not file-versioning.** Single-file reports +(`01-functionality.md`, `06-analysis.md`, etc.) and the +`tests///` tree are each their own current +canonical state. History is git — there is no `version:` field, no +`archive/` directory, no manually-bumped counters, and no +`inputs.step_N: ` lineage tracking. Prior states are +recoverable via `git log`. See [`skills/skill-optimizer-shared/iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md) +for the full mechanics. + **Two contexts, one workflow:** - **Upstream skill:** user provides URL or `//`. @@ -145,13 +167,13 @@ and dispatches the subagent with only the narrow chunks it needs. | Subagent | Sees | Does NOT see | Why | |---|---|---|---| -| Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, prior `01-functionality.md` drafts | Pure research, no contamination across iterations | -| Test-case designer (step 2) | `01-functionality.md` (latest), current `02-test-case.md` (when extending in maintenance mode), `${OPERATOR_DIRECTIVES}` | Skill source content, archived prior `02-test-case.md` drafts, `06-analysis.md`, optimizer attempts, failure data | Step 2 is a maintenance step — the canonical file accumulates coverage across re-runs and the subagent extends it rather than re-deriving. Coverage design stays at responsibility level. Reasoning from source would gerrymander tests around the source's literal phrasing instead of testing stated responsibilities | -| Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change | Just upstream facts | -| Test writer (step 4, dispatched per case) | The single test case spec + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other test cases in the suite, the eval grader's matching logic | Fixture writing needs source detail (specific patterns, concrete violation examples). Sees source for grounding; the test case spec from step 2 constrains what the fixture should test, bounding the gerrymandering risk. Still blocked from seeing other cases (prevents copying) and grader internals (prevents grader-leak hacking) | -| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), prior `06-analysis.md` drafts | Forces it to think about the SKILL, not the SOLUTIONS; iteration-isolated | -| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, prior optimizer attempts | Forces principled improvement, not pattern-match patches; no attachment to prior failed attempts | -| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, prior validator verdicts | Independent check; can't be biased by what the optimizer (or a prior validator round) told itself | +| Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, `01-functionality.md` (own canonical) or its git history | Fresh-derivation: pure research, no contamination across iterations | +| Test-case designer (step 2) | `01-functionality.md`, current `tests//spec.yaml` tree (when present — this is the state), `${OPERATOR_DIRECTIVES}` | Skill source content, `06-analysis.md`, optimizer attempts, failure data, git history of any tree files | Maintenance: the `tests/` tree IS the accumulating state; the subagent extends it across re-runs (adding/modifying spec.yaml files) without seeing source (which would gerrymander tests around the source's literal phrasing) | +| Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change, `03-submissions.md` (own canonical) or its git history | Fresh-derivation: just upstream facts | +| Test writer (step 4, dispatched per probe) | The single probe's `spec.yaml` + parent functionality's `spec.yaml` + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other probes' specs/graders, the eval grader's matching logic, sibling probe contents (own canonical at the per-probe level), git history of any tree files | Maintenance at tree level: each probe is built in isolation. Fixture writing needs source detail (specific patterns); the probe spec from step 2 bounds the gerrymandering risk. Blocked from seeing siblings (prevents copying) and grader internals (prevents grader-leak hacking) | +| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), `06-analysis.md` (own canonical) or its git history | Fresh-derivation: forces focus on the SKILL, not the SOLUTIONS; iteration-isolated | +| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `07-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts | +| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `07-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | The skill (operator session) sees everything; subagents see slices. This is the architectural fix for the "tunnel-vision into ducktape" @@ -160,12 +182,13 @@ problem. ## Iteration patterns The chain isn't strictly linear in practice. Any step can be re-run -(operator-driven or auto-pilot-driven), and re-running a step -invalidates downstream reports that derived from its prior version. +(operator-driven or auto-pilot-driven), and re-running a step can +invalidate downstream artifacts that derived from its prior state. This section defines how iteration is captured, detected, and -cascaded — using a lightweight per-report version mechanism that -keeps state inspection deterministic without requiring a global -iteration counter. +cascaded — using git as the history mechanism and the filesystem +itself as the current state. Full mechanics in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md); +the summary below is for spec context. ### When iteration happens @@ -174,44 +197,51 @@ Common backtrack triggers: | At step | Common trigger | Backtracks to | |---|---|---| | After 2 (test design) | "Coverage is wrong; the skill does more than this" | 1 (research) or 2 (revise with directives) | -| After 4 (write-tests) | Test writer couldn't implement a case | 2 (revise the set) | -| After 5 (bench) | All tests pass on baseline — too easy | 2 or 4 (harder tests) | -| After 6 (analyze) | "No clear weakness" but user disagrees | 2 (better coverage), 4 (better tests), or 6 with directives | +| After 4 (write-tests) | Test writer couldn't implement a probe | 2 (revise the functionality spec) | +| After 5 (bench) | All probes pass on baseline — too easy | 2 or 4 (harder probes) | +| After 6 (analyze) | "No clear weakness" but user disagrees | 2 (better coverage), 4 (better probes), or 6 with directives | | After 7 (improve) | Validator rejects beyond optimizer's own loop | 6 (re-analyze) or 2 (the test was wrong) | | Anywhere | User dislikes the output | re-run that step with new directives | There's no rigid backtrack flowchart — operator and auto-pilot both -decide based on the named weakness in `06-analysis.md`, the validator -verdict, and the version mechanism below. - -### Versioning - -Every report carries two frontmatter fields: - -```yaml -version: 1 # this report's own iteration counter -inputs: - step_1_functionality: 1 # versions of upstream reports I derived from - step_2_test_case: 1 # (only those this step actually read) +decide based on the named weakness in `06-analysis.md`, the +validator verdict, and the staleness mechanism below. + +### Step kinds: fresh-derivation vs maintenance + +Each step is one of two kinds, and the subagent's reading rule +differs: + +**Fresh-derivation steps (1, 3, 5-summary, 6, 7):** each invocation +derives the canonical artifact from upstream + directives. The +subagent does NOT read its own canonical file (the prior state), +and does NOT walk git history of it. Anti-ducktape constraint — +the optimizer must not see prior attempts, the analyzer must not +see prior analyses. + +**Maintenance steps (2, 4):** the canonical artifact is an +accumulating tree on disk (`tests//spec.yaml` at +step 2; `tests///{spec.yaml, workspace, +grader, smoke}` at step 4). The subagent DOES read the current +tree (it IS the state) and produces a modified tree. Git history +of the tree files is still off-limits — only the current state +matters. + +### Staleness detection (git-native) + +There is no `version:` field or `inputs.step_N: ` lineage +tracking. Staleness is detected by comparing git modification times: + +```bash +UPSTREAM_T=$(git log -1 --format=%ct -- 01-functionality.md) +DOWNSTREAM_T=$(git log -1 --format=%ct -- tests/) +[ "$UPSTREAM_T" -gt "$DOWNSTREAM_T" ] && echo "downstream stale" ``` -On invocation, a step compares the current versions of its upstream -inputs against what its own prior report (if any) recorded under -`inputs`: - -- **No existing report** → write a new one at `version: 1` -- **Existing report, all input versions match** → already current; do - nothing (unless the operator passed directives requesting a re-run - anyway) -- **Existing report, any input version mismatch** → existing is stale - → archive it → re-run with bumped own-version - -Versions are per-step independent counters. Step 1 going from v1 → v2 -doesn't change step 2's version until step 2 is re-invoked and -detects the mismatch. - -Pre-N files are untouched when step N is re-run; only step N's own -version bumps. +In interactive use, the operator typically just knows ("I re-ran +step 1, so step 2 needs a re-run"). The git-mtime check is for +auto-pilot (step 8), which walks the chain forward and re-runs any +downstream older than its direct upstream. ### Re-entry contract @@ -222,36 +252,28 @@ from prior iterations ("user requested coverage for null inputs", first invocation. The slot is **never a context dump** of prior output; it's a short bulleted list of new requirements only. -Subagent behavior under iteration: - -- Subagent **ignores its own prior output** — no peeking at `archive/`. -- Subagent reads upstream reports at their **latest** versions only. -- Operator session archives the prior version before dispatching, and - records the upstream versions consumed in the new report's `inputs` - frontmatter. +### History -### Archive convention - -The latest version of each report lives at its canonical path: -`docs/skill-optimizer//-.md`. When overwritten, the -prior version moves to -`docs/skill-optimizer//archive/--v.md` (where `N` -is the version being archived). History is browsable for human audit -but inert — auto-pilot and chain skills only read latest. +Git is the archive. There is no `archive/` subdirectory. When a +fresh-derivation step overwrites its canonical file, the prior +state is captured in git's previous commit. When a maintenance +step modifies a tree (rebuilds a probe, removes a de-picked +functionality), the operator session commits a checkpoint **before** +the destructive change so git history has a clean before/after +breakpoint. ### Transitive staleness -Each step checks **direct upstream only**. If step 1 goes to v2 but -step 2 isn't re-run (user judged it still valid), step 3 sees step 2 -at v1 and treats it as current — even though step 2's report -references step 1 v1. +Each step checks **direct upstream only**. If step 1 is updated but +step 2 isn't re-run (user judged it still valid against the new +upstream), step 3 sees step 2 as current. The trade-off: transitive issues surface as confusing downstream results, not silent corruption. The discipline is that skipping a -step's re-run is an explicit judgment call ("the old report is still -valid against the new upstream"), and we trust that call. Auto-pilot -doesn't make this judgment — it re-runs any step whose direct -upstream version doesn't match. +step's re-run is an explicit operator judgment ("the old artifact +is still valid against the new upstream"). Auto-pilot doesn't make +this judgment — it re-runs any step whose direct upstream is newer +per `git log`. ## The skills @@ -304,48 +326,56 @@ Per-step iteration behavior is noted at the end of each subsection. Note: this report's `pr_submission_intent` field tells step 2's handoff whether step 3 (`investigate-submissions`) should run." - **Iteration behavior:** re-run when source skill changes or when - the operator wants fresh research with new directives. Re-running - bumps `01-functionality.md` version; downstream steps detect the - change and cascade on next invocation. + the operator wants fresh research with new directives. + Fresh-derivation step — the subagent doesn't read its own + canonical or its git history. Re-running overwrites + `01-functionality.md`; prior state is in git history. Downstream + steps detect the change via `git log` mtime comparison. ### 2. `skill-optimizer-investigate-test-case` - **Description trigger:** "design tests for this skill", "propose test cases", "what should we test" - **Input:** `01-functionality.md` -- **Output:** `02-test-case.md` — ranked list of proposed test cases. - Each entry: name, what-it-tests (which responsibility), required - setup, expected agent behavior, grader spec, why-it-matters. +- **Output (two artifacts):** + - **`00-test-proposals.md`** — one-time audit report from the + subagent: ranked list of proposed functionalities, with + importance + suggested_probes + why-it-matters for each. This + file captures the design reasoning; it is not consulted by + downstream steps. + - **`tests//spec.yaml`** — one folder per + proposed functionality, each containing a single `spec.yaml` + with fields: `name`, `description`, `picked: true|false` + (initially set per the subagent's recommendation; user edits to + finalize), `importance`, `suggested_probes: [list]`, + `why_test`. The filesystem IS the state — there is no + `picked: []` array anywhere. - **Behavior:** enumerate the skill's responsibilities from - functionality report; design 1–2 cases per responsibility; flag edge - cases (boundary conditions, error paths, common pitfalls); estimate - grader difficulty (deterministic check vs needs-real-tooling); rank - by importance + cost-to-build. -- **Dispatches:** test-case-designer subagent (limited context: only - `01-functionality.md` + the `${OPERATOR_DIRECTIVES}` slot; does NOT - see prior `02-test-case.md` drafts, prior `06-analysis.md`, - optimizer attempts, or failure data — keeps coverage design - honest across iterations, where the operator session has already - absorbed prior failures and would otherwise bias tests toward - "what just failed" instead of "what comprehensively covers the - skill's responsibilities"). + functionality report; design 1–N functionalities to test; flag + edge cases; suggest probes per functionality; rank by importance. +- **Dispatches:** test-case-designer subagent (limited context: + `01-functionality.md` + the current `tests//spec.yaml` + tree if present + `${OPERATOR_DIRECTIVES}`; does NOT see skill + source content, `06-analysis.md`, optimizer attempts, failure + data, or git history of the tree). Maintenance step — the + subagent extends the tree (adds new functionality folders; + modifies existing spec.yaml files per directives) rather than + replacing it. - **User gate:** after the subagent returns, the operator session - prompts the user to pick which subset of proposed cases to - actually build (write-tests acts on the picked subset). The picks - are recorded in `02-test-case.md` frontmatter (e.g., - `picked: [case-name-1, case-name-3]`). + shows the user the proposed functionalities + ranking, and the + user edits `picked: true|false` in each `spec.yaml`. Step 4 will + build probes for `picked: true` functionalities only. - **Handoff:** read `01-functionality.md`'s `pr_submission_intent` field. If `true` → "Invoke `skill-optimizer-investigate-submissions` next, then `skill-optimizer-write-tests`." If `false` → "Skip step - 3; invoke `skill-optimizer-write-tests` with the user's picked - subset." No late prompts — the decision was made at step 1. + 3; invoke `skill-optimizer-write-tests`." - **Iteration behavior:** re-run when the user wants different - coverage, when step 4/5/6 surface a coverage gap, or when step 1's - version bumps. `${OPERATOR_DIRECTIVES}` is the channel for - "user wants null-input edge cases" or similar atomic new - requirements; the subagent treats these as fresh requirements on - top of `01-functionality.md`, never as a context dump of prior - iterations. + coverage, when step 4/5/6 surface a coverage gap, or when step 1 + changed. Maintenance step — existing functionality folders the + user has invested in (especially `picked: true` ones) are + preserved across re-runs unless a directive explicitly asks for + revision. Destructive edits (removing a functionality) checkpoint + via git commit before the change. ### 3. `skill-optimizer-investigate-submissions` (OPTIONAL — upstream only) @@ -367,60 +397,87 @@ Per-step iteration behavior is noted at the end of each subsection. - **Handoff:** "Validator in `improve-skill` will read this for external consistency check. Continue with `write-tests` if not done." - **Iteration behavior:** rarely needs re-running — upstream PR - conventions change slowly. Re-run when the upstream repo's - CONTRIBUTING/CLA changes or when `01-functionality.md` version - bumps (in case the skill source URL itself changed). + conventions change slowly. Fresh-derivation step. Re-run when the + upstream repo's CONTRIBUTING/CLA changes or when + `01-functionality.md` changed (the skill source URL itself may + have changed). Overwrites `03-submissions.md`; prior state in git + history. ### 4. `skill-optimizer-write-tests` - **Description trigger:** "build the tests", "implement the workbench", "set up the eval" -- **Input:** `01-functionality.md`, `02-test-case.md` (operator picks - subset) -- **Output:** `04-tests-plan.md` (per-case implementation plan) + - `workbench/` (suite.yml, workspace files, graders, smoke check) -- **Behavior:** plan workbench structure → show user the plan → user - confirms → dispatch parallel subagents (one per test case, each - builds one workspace file + one grader) → run smoke check - (hand-crafted GOOD/BAD/EMPTY findings.txt fixtures against each - grader) → commit -- **Dispatches:** test-writer subagent per case (limited context: - single case spec + functionality report + skill source content; - does NOT see other test cases or eval grader's matching logic — - prevents copying across cases and grader-leak hacking. Source - access is granted at this step because concrete fixture writing - needs specific violation patterns; the per-case test spec from - step 2 constrains what the fixture should test, bounding the - gerrymandering risk) -- **Parallelizable:** each case independent; dispatch in a single - message -- **Handoff:** "Invoke `skill-optimizer-run-bench` to measure baseline." -- **Iteration behavior:** re-run when `02-test-case.md` version - bumps (picked subset changed), when the test-writer flagged - unimplementable cases, or when bench results suggest tests are - systematically too easy or too narrow. Each test-writer subagent - also accepts `${OPERATOR_DIRECTIVES}` for case-level revision - hints. +- **Input:** `01-functionality.md`, `tests//spec.yaml` + tree (from step 2, picked subset determined by each spec's + `picked: true|false`) +- **Output:** for each picked functionality, one or more probe + folders at + `tests///{spec.yaml, workspace/, + grader.mjs, smoke/}` + a generated `tests/suite.yml` that lists + the probes step 5 will run. +- **Behavior:** + 1. Walk `tests/` for `picked: true` functionalities. + 2. For each, decide the probe set (informed by + `suggested_probes` in the functionality spec.yaml + any + `${OPERATOR_DIRECTIVES}` for that functionality). + 3. Show the user the planned probe set per functionality; user + confirms or revises. + 4. Dispatch parallel test-writer subagents (one per probe); + each builds its probe's workspace + grader + smoke fixtures. + 5. Run the smoke check on each probe's grader. + 6. Regenerate `tests/suite.yml` from the current picked + functionalities' probes. +- **Dispatches:** test-writer subagent per probe (limited context: + single probe's `spec.yaml` from step 2 + parent functionality's + `spec.yaml` + `01-functionality.md` + skill source content; does + NOT see other probes' specs/graders/workspaces or the eval grader's + matching logic — prevents copying across probes and grader-leak + hacking. Source access is granted because concrete fixture + writing needs specific violation patterns; the probe spec from + step 2 bounds the gerrymandering risk.) +- **Parallelizable:** each probe independent; dispatch in a single + message. +- **Destructive-edit safety:** before rebuilding an existing probe + (e.g., per a directive), the operator session commits the current + state so git history has a clean before/after breakpoint. +- **Handoff:** "Invoke `skill-optimizer-run-bench` to measure + baseline." +- **Iteration behavior:** re-run when step 2's tree changed (new + `picked: true` functionalities; revised functionality specs that + warrant probe rebuilds), when a test-writer flagged unimplementable + probes, or when bench results suggest probes are systematically + too easy or too narrow. Maintenance step — existing probe folders + are preserved across re-runs unless a directive explicitly asks + to rebuild a specific probe. ### 5. `skill-optimizer-run-bench` - **Description trigger:** "run the eval", "measure", "benchmark" -- **Input:** `workbench/` + source skill (vendored) -- **Output:** `05-bench-results//suite-result.json` + - per-trial traces + per-trial findings.txt +- **Input:** `tests/suite.yml` (generated by step 4) + source skill + (vendored) +- **Output (two artifacts):** + - **`05-bench-results//`** — raw CLI output: + `suite-result.json` + per-trial traces + per-trial findings.txt. + Timestamped per run; old runs preserved naturally. + - **`05-bench-summary.md`** — single canonical aggregate: + overall pass rate, per-model pass rate, per-probe pass rate, + failed-probe pointer list, pointer to the latest + `05-bench-results//`. Fresh-derivation per run; prior + summary in git history. - **Behavior:** invoke skill-optimizer CLI - (`npx tsx /src/cli.ts run-suite ./workbench/suite.yml - --trials 3`); capture results + (`npx tsx /src/cli.ts run-suite ./tests/suite.yml + --trials 3 --out 05-bench-results//`); parse the suite result; + write the summary. - **Dispatches:** none — direct CLI invocation - **Note:** this is intentionally thin; the entire v1.3 run-suite logic stays as-is in the CLI - **Handoff:** "Invoke `skill-optimizer-analyze-result`." -- **Iteration behavior:** re-run when `workbench/` version bumps - (new or revised tests) or when the operator wants fresh results - against the same workbench. Each bench result is timestamped under - `05-bench-results//` so old runs are never overwritten — but - `05-bench-summary.md` (the versioned summary that analyze-result - reads) is overwritten with archive on re-run. +- **Iteration behavior:** re-run when `tests/` changed (step 4 added + or revised probes) or when the operator wants fresh trial data. + Each run produces a new timestamped raw directory; the summary is + overwritten and git tracks the prior version. The raw + bench-results dirs are outside the iteration protocol (they're + naturally accumulating snapshots, not versioned artifacts). ### 6. `skill-optimizer-analyze-result` @@ -444,12 +501,13 @@ Per-step iteration behavior is noted at the end of each subsection. `skill-optimizer-improve-skill`." Otherwise → exit honestly ("no structural weakness; no improvement warranted"). - **Iteration behavior:** re-run when bench results change or when - the operator/user wants a fresh look with new framing. The - `${OPERATOR_DIRECTIVES}` slot is the channel for hints like - "focus on the gpt-5 cluster" or "the user thinks weakness X is - actually two separate issues" — the subagent treats these as - additional analytical lenses, not as a context dump of prior - conclusions. + the operator/user wants a fresh look with new framing. + Fresh-derivation — the subagent doesn't read its own canonical + `06-analysis.md` or its git history. Re-running overwrites the + canonical; prior state in git. The `${OPERATOR_DIRECTIVES}` slot + carries hints like "focus on the gpt-5 cluster" or "the user + thinks weakness X is actually two separate issues" — additional + analytical lenses, never a context dump of prior conclusions. **`06-analysis.md` format (Option A — structured):** @@ -519,15 +577,16 @@ Per-step iteration behavior is noted at the end of each subsection. exists because step 3 ran). The draft includes the diff, the body, caveats, and operator-steps-to-submit. No late "submit a PR?" prompt — the decision was already made at step 1. -- **Iteration behavior:** re-run when `06-analysis.md` version bumps +- **Iteration behavior:** re-run when `06-analysis.md` changed (new analysis = potentially different weakness) or when the - operator wants a fresh optimization attempt. The internal - optimizer/validator loop (max 2 rounds, baked into the skill) - handles in-step iteration; cross-step backtracking is the - operator's call. `${OPERATOR_DIRECTIVES}` is the channel for hints - like "prefer additive changes" or "don't touch the description - field" — the optimizer treats these as additional constraints, not - as a context dump of prior attempts. + operator wants a fresh optimization attempt. Fresh-derivation + for both optimizer and validator — neither subagent reads its + canonical or git history. The internal optimizer/validator loop + (max 2 rounds, baked into the skill) handles in-step iteration; + cross-step backtracking is the operator's call. + `${OPERATOR_DIRECTIVES}` carries hints like "prefer additive + changes" or "don't touch the description field" — additional + constraints, never a context dump of prior attempts. ### 8. `skill-optimizer-autopilot` @@ -542,14 +601,14 @@ Per-step iteration behavior is noted at the end of each subsection. each step's final version, headline result, and any blockers - **Behavior:** 1. Walks 1→7 in order, dispatching each chain skill. - 2. For each step: check whether the existing report is current - against its declared `inputs` versions. If current, skip; if - missing or stale, dispatch. + 2. For each step: check whether the existing artifact is current + via `git log` mtime comparison against direct upstream. If + current, skip; if missing or stale, dispatch. 3. Handles the three human-gate points with default policies: - **B1 PR-intent question** (upstream skills) → uses the `pr_intent` flag value (default `false`) - - **B2 user-picks-subset gate** → takes the top-N by importance - (default `pick_top_n: 5`) + - **B2 user-picks gate** → sets `picked: true` on top-N + functionalities by importance (default `pick_top_n: 5`) - **B7 validator-rejected verdict** → logs the final state, no further automated retries; surfaces blocker in the summary 4. Bounds iterations: at most `max_iterations_per_step` re-runs @@ -566,10 +625,11 @@ Per-step iteration behavior is noted at the end of each subsection. failures are acceptable, not for high-stakes single-target optimization. - **Iteration behavior:** auto-pilot is itself iterable. Re-running - picks up at whatever step is stale per the version mechanism; + picks up at whatever step is stale per the git-mtime comparison; steps that are current are skipped. The summary report is - timestamped per run rather than versioned — each auto-pilot run - produces a fresh summary so the audit trail is preserved. + timestamped per run rather than treated as a versioned artifact — + each auto-pilot run produces a fresh summary so the audit trail + is preserved. ## Plugin packaging @@ -625,9 +685,10 @@ For v1.4 to be considered done: pre-loaded recipe library until real end-to-end observations surface curated patterns 5. **Iteration mechanism works end-to-end** — re-running a step on - an existing slug correctly archives the prior report, bumps the - `version` field, and the next downstream step on next invocation - detects the version mismatch and re-runs itself + an existing slug correctly overwrites its canonical artifact + (prior state in git history); the next downstream step on next + invocation detects the upstream change via `git log` mtime + comparison and re-runs itself 6. **End-to-end test on a local skill** (e.g., one of the existing `skills/skill-optimizer/SKILL.md` or a small new skill in this repo) — walks 1→2→4→5→6→7, produces all expected reports, modifies diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 82247db..30aa712 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -3,132 +3,130 @@ **Load this file every time** a skill-optimizer chain skill executes its "Handle iteration" step. The mechanics are identical across every skill in the chain; the parent skill's SKILL.md only specifies -that-skill's slot in the chain. The actual archive + version bump + -directives-collection logic lives here so the chain behaves -consistently across all steps. +that-skill's slot in the chain. The actual re-run logic + subagent +constraints live here so the chain behaves consistently. If you're a chain skill's SKILL.md and you're tempted to inline this -logic instead of reading this file: don't. Inlining drifts. Read this -file, do what it says. +logic instead of reading this file: don't. Read this file, do what +it says. ## What this protocol governs -Every chain step produces a versioned report at a canonical path -under `docs/skill-optimizer//`. When a step is re-run (either -manually or by the auto-pilot), this protocol determines: +Every chain step produces a canonical artifact under +`docs/skill-optimizer//`. Most are single-file reports +(`01-functionality.md`, `06-analysis.md`, etc.); two are +filesystem-as-state trees (`tests//spec.yaml` for +step 2, `tests///{spec.yaml, workspace, grader, +smoke}` for step 4); one is a timestamped collection +(`05-bench-results//` for step 5's raw output). -1. Whether the existing report (if any) is current or stale -2. How to archive a stale report before writing a new one -3. What version number the new report gets -4. How to pass cross-iteration learnings into the subagent -5. What the subagent is and isn't allowed to look at +This protocol determines, on every step invocation: + +1. Whether the existing artifact is current or stale +2. What the subagent is and isn't allowed to look at +3. How the operator handles destructive edits safely + +History is **git**. There is no `version:` field, no `archive/` +subdirectory, no manually-bumped counters, and no `inputs.step_N:` +ints recording which upstream version was consumed. Git already +content-addresses every prior state; reimplementing that in +frontmatter is bookkeeping for its own sake. ## Step kinds: fresh-derivation vs maintenance -Steps in the chain come in two kinds, and the protocol's "what the -subagent sees" rule differs between them. +Steps in the chain come in two kinds, and the subagent-constraint +rule differs between them. -**Fresh-derivation steps** produce their report from upstream -inputs + directives, with no continuity from prior versions. The -subagent does NOT see its own current or archived file — every -invocation is a derivation from scratch. This is the anti-ducktape -rule: the optimizer must not see prior optimization attempts, the -analyzer must not see prior analyses, the researcher must not -rationalize a prior report. +**Fresh-derivation steps** produce their artifact from upstream +inputs + directives, with no continuity from prior outputs. The +subagent does NOT read its own canonical file, and does NOT read +git history of it. Every invocation is a derivation from scratch. +This is the anti-ducktape rule: the optimizer must not see prior +optimization attempts, the analyzer must not see prior analyses, +the researcher must not rationalize a prior report. | Step | Why fresh derivation | |---|---| | 1. investigate-functionality | Each invocation researches from source — no continuity needed | | 3. investigate-submissions | Each invocation researches upstream — no continuity needed | +| 5. run-bench (summary file) | Mechanical write-up of the new raw bench run, not an extension of prior summary | | 6. analyze-result | Must not be biased by prior analyses; load-bearing for anti-ducktape | | 7. improve-skill | Optimizer must not see prior attempts; load-bearing for anti-ducktape | -**Maintenance steps** maintain an accumulating state — the file -represents a growing collection that the user contributes to over -multiple invocations. Re-runs EXTEND the current file rather than -replacing it. The subagent reads the current canonical file as one -of its inputs and produces an extended version of it. +**Maintenance steps** manage an accumulating tree on disk — the +filesystem itself is the state. Re-runs read the current tree and +extend or modify it. The subagent DOES read the current canonical +files (because they are the state) and produces a modified tree. | Step | What accumulates | |---|---| -| 2. investigate-test-case | The ranked test-case list grows as the user adds coverage | -| 4. write-tests | The workbench directory grows as new picked cases are built | - -**The "subagent ignores its own prior output" rule applies to -fresh-derivation steps only.** For maintenance steps, the current -canonical file IS the load-bearing input — without it, the step -would lose continuity and the user's accumulated decisions (picks, -prior directives, manually-added cases) would be lost on every -re-run. - -The archive remains inert in both cases — neither kind of step -reads from `archive/`. The archive is audit-only. - -## Versioning convention - -Every chain report carries two frontmatter fields: - -```yaml -version: # this report's own iteration counter -inputs: - step_1_functionality: # version of each upstream report consumed - step_2_test_case: # only the ones this step actually reads - # ... etc. +| 2. investigate-test-case | `tests//spec.yaml` grows as functionalities are added/refined | +| 4. write-tests | `tests///` probes grow as the user adds coverage | + +The git-history-as-archive principle covers both: prior states of +maintenance steps are recoverable from `git log`, and prior states +of fresh-derivation steps are the same. The difference is purely +what the subagent is allowed to look at. + +## Subagent constraints + +Three rules every reasoning subagent must follow: + +1. **Read upstream artifacts only at their current state.** The + filesystem is the source of truth; do not walk `git log` looking + for prior versions of upstream files. + +2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7): do NOT + read your own canonical file, and do NOT read git history of + it.** Your job is to derive a new artifact from upstream + + directives, with no reference to what was produced before. This + is the load-bearing anti-ducktape constraint. The operator + session writes your output by overwriting the canonical file; + git captures the prior state automatically, but neither you nor + any subsequent fresh-derivation subagent should walk that + history. + +3. **For maintenance steps (2, 4): DO read your own canonical + tree** (when it exists). Your job on a re-run is to extend or + modify the current state per `${OPERATOR_DIRECTIVES}`, + preserving entries the user has invested in (picked + functionalities, built probes, existing graders) unless a + directive explicitly says to revise a specific entry. Do NOT + walk git history of the tree — the current state is what + matters; prior states are inert. + +Each subagent prompt template enforces the appropriate rule for +its step. + +## Staleness detection + +Whether an artifact is stale is determined by comparing git +modification times against upstream files: + +```bash +# Is 02-test-case stale relative to 01-functionality? +UPSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//01-functionality.md) +DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) +[ "$UPSTREAM_T" -gt "$DOWNSTREAM_T" ] && echo "stale" ``` -Step 1 has no upstream chain reports, so it omits `inputs:`. Every -other step lists the upstream chain reports it reads, with the -version it consumed at the time of writing. - -Versions are per-step independent counters. Step 1 going from v1 → v2 -does not change step 2's version until step 2 is itself re-invoked -and detects the mismatch. - -## On every chain-skill invocation - -Before dispatching the step's subagent: +In interactive use, the operator typically just knows ("I re-ran +step 1, so step 2 needs a re-run"). The git-mtime check is for +auto-pilot (step 8), which walks the chain forward and re-runs any +downstream that's older than its direct upstream. -1. Compute the canonical report path for this step + slug. -2. Check whether the canonical path exists, and act according to - the step's kind (fresh-derivation or maintenance — see the - classification above): - - - **Path does not exist.** This is iteration 1 for this step. Set - `new_version = 1`. Proceed to "Collect operator directives" - below. - - - **Path exists, all input versions match, no operator directives - pending.** The existing report is current; nothing to do. - Report success, hand off, exit. - - - **Path exists, any input version mismatches OR operator passed - directives requiring a re-run.** Behavior depends on the - step's kind: - - - **Fresh-derivation step:** copy the existing file to - `archive/--v.md` (audit trail); - the subagent will derive a new file from upstream + directives - only, with no reference to the prior or archived content. Set - `new_version = existing_version + 1`. - - - **Maintenance step:** copy the existing file to - `archive/--v.md` (audit trail); - the subagent receives the current canonical file as an input - and produces an extended/modified version of it. Set - `new_version = existing_version + 1`. The archive copy is - inert from this point on — the subagent operates on the - canonical file content, not the archive. - - In both cases the canonical path is overwritten with the new - version; the archive copy preserves the prior version for audit. +For maintenance steps, the operator's directives can also force a +re-run even when no upstream changed ("add a probe for null +input"). The protocol does not require a stale upstream to permit +a re-run; it only flags staleness to recommend one. ## Collect operator directives `${OPERATOR_DIRECTIVES}` is a short bulleted list of **atomic new -requirements** that surfaced from prior iterations or from the user's -current request. It is **never** a context dump of the prior report's -content; it is a list of specific asks that the new derivation must -satisfy on top of its normal inputs. +requirements** that surfaced from prior iterations or from the +user's current request. It is **never** a context dump of prior +artifact content; it is a list of specific asks that the new +derivation must satisfy on top of its normal inputs. Examples that count as atomic requirements: @@ -140,15 +138,39 @@ Examples that count as atomic requirements: Examples that do NOT count (these are context dumps, not atomic requirements — reject the temptation): -- The full text of the prior `02-test-case.md` so the subagent can - "see what we already had" +- The full text of a prior canonical file pasted in so the + subagent can "see what we already had" - "Here's what failed last time:" followed by failure data - The validator's prior verdict pasted in for the subagent to read If the user supplied no new requirements and you're re-running -simply because an upstream report changed, leave directives empty. -The subagent will re-derive its output from the (new) upstream -report alone. +because an upstream artifact changed, leave directives empty. + +## Safe destructive edits + +When a maintenance step's re-run will overwrite or delete existing +content (e.g., step 4 rebuilds a probe with `rm -rf +tests///workspace/`; step 2 removes a de-picked +functionality), the operator session commits the current state +**before** dispatching the destructive change. This gives git +history a clean before/after breakpoint: + +```bash +# Before destructive change: +git add -A docs/skill-optimizer// +git commit -m "checkpoint: before rebuild of " +# Then dispatch the destructive change. +``` + +Without this checkpoint, the prior state of a rebuilt probe (or +de-picked functionality) gets buried inside a multi-file commit +later, making "show me what this probe used to look like" awkward +to recover. + +Fresh-derivation steps are also destructive (they overwrite the +canonical file), but the prior state is already self-contained in +its own commit (the previous fresh-derivation run's commit), so +the checkpoint isn't needed. ## Dispatch the subagent @@ -156,119 +178,62 @@ When you invoke the subagent for this step, pass these templated inputs (each subagent prompt template names them with `${...}` placeholders): -- `${VERSION}` — `new_version` from the decision tree above - `${OPERATOR_DIRECTIVES}` — the bulleted list (may be empty) -- Step-specific inputs (paths to upstream reports the subagent - needs, the output path, etc.) per that step's SKILL.md - -## Subagent constraints under iteration - -Two rules every reasoning subagent must follow: - -1. **Never read from `archive/`.** Archived versions are inert audit - trail. Don't walk `archive/--v*.md` for any reason. - The current canonical file (and only the current canonical file) - is the load-bearing input for a maintenance step; for a - fresh-derivation step, neither the canonical nor the archive is - read. - -2. **Read upstream reports at their latest versions only.** Upstream - canonical paths always point at the latest version. Same archive - rule applies to upstream files — never walk their archives. - -A third rule that depends on the step's kind: - -- **For fresh-derivation steps (1, 3, 6, 7): do NOT read your own - step's canonical file.** Your job is to derive a new report from - upstream + directives, with no reference to what was produced - before. This is the load-bearing anti-ducktape constraint — the - optimizer must not see prior attempts, the analyzer must not see - prior analyses, etc. - -- **For maintenance steps (2, 4): DO read your own step's canonical - file (when it exists).** Your job on a re-run is to extend or - modify the current accumulating state per `${OPERATOR_DIRECTIVES}`, - preserving existing entries the user has invested in (test-case - picks, prior coverage decisions, etc.) unless a directive - explicitly says to revise a specific entry. - -Each subagent prompt template enforces the appropriate version of -the third rule for its step. If you find yourself wanting to -violate any of these — that's the failure mode the iteration -mechanism prevents. Don't. - -## Cascading staleness - -Each step checks **direct upstream only** — the immediate prior -report(s) it consumes, not the whole upstream chain. If step 1 goes -to v2 but step 2 is not re-run (the user judged step 2 still valid -against the new step 1), step 3 will see step 2 at v1 and treat it -as current — even though step 2's `inputs` still references step 1 -v1 transitively. - -The discipline: skipping a step's re-run is an **explicit operator -judgment** that the existing report is still valid against the new -upstream. Trust the call. Auto-pilot doesn't make this judgment — it -always re-runs on direct-upstream version mismatch, which causes the -cascade to propagate naturally as you walk the chain forward. - -If transitive staleness surfaces as confusing downstream results, -the operator re-runs the skipped step manually and the cascade -catches up on the next forward walk. +- Step-specific inputs (paths to upstream artifacts the subagent + reads, the output path, etc.) per that step's SKILL.md ## Re-run authorization A chain skill never invokes another chain skill on its own. When a step finds that its upstream input is unsatisfactory — a thin -functionality report, a too-easy bench, a missing-but-needed test -case, a wrong-target PR-conventions report — it surfaces the +functionality report, a too-easy bench, a missing-but-needed +probe, a wrong-target PR-conventions report — it surfaces the finding to the user and stops. The user (or the auto-pilot driver applying its default policies) decides whether to re-run an upstream step, accept the situation, or abandon the run. This applies to backward triggers specifically. Forward handoffs -(step N hands off to step N+1) are part of the chain's normal flow -and the handing-off skill emits the handoff message; the user or -auto-pilot acts on it. +(step N hands off to step N+1) are part of the chain's normal +flow and the handing-off skill emits the handoff message; the +user or auto-pilot acts on it. Why this rule exists: chain skills running on user request must give the user control over what they're paying for. Re-running an -upstream step takes time and tokens; the user should authorize that -explicitly, not have the agent decide on its own. The user is also -the one who knows whether the upstream-thin condition is real -("yes, the functionality report did miss responsibility X — re-run") -or expected ("no, the skill really is that small — accept the small -proposal"). +upstream step takes time and tokens; the user authorizes that +explicitly, not the agent. -## Archive convention +## Cascading staleness -| Path | Content | -|---|---| -| `-.md` | Always the latest version. Read this. | -| `archive/--v.md` | Prior versions, frozen. For human audit. Do not read from chain skills or subagents. | +Each step checks **direct upstream only** — the immediate prior +artifact(s) it consumes, not the whole upstream chain. If step 1 +is updated but step 2 is not re-run (the user judged step 2 still +valid against the new step 1), step 3 will see step 2 as current +even though step 2's git mtime is older than step 1's. -The `archive/` directory is browsable but inert. It exists for -contributor audit and debugging — never for the chain to consult on -the next run. +The discipline: skipping a step's re-run is an **explicit operator +judgment** that the existing artifact is still valid against the +new upstream. Auto-pilot doesn't make this judgment — it always +re-runs on direct-upstream mtime mismatch, so the cascade +propagates naturally as you walk the chain forward. ## The bootstrapping case On the very first invocation of a step on a fresh slug, -`docs/skill-optimizer//` may not yet exist. Create it (and the -`archive/` subdir, even if empty) as part of the iteration step's -setup. Subsequent invocations expect both dirs to be present. +`docs/skill-optimizer//` may not yet exist. Create it as +part of the iteration step's setup. No `archive/` subdirectory is +needed — git is the archive. ## What this protocol does NOT cover -- **Step 5 (`run-bench`) bench results** are timestamped under - `05-bench-results//` rather than versioned. Each bench run is - preserved naturally. The versioned report step 5 produces is the - bench-result summary (which references the timestamped run), - which follows this protocol like any other report. +- **Step 5's raw bench results** are timestamped under + `05-bench-results//`. Each run preserves naturally as its + own directory. The summary file `05-bench-summary.md` is a + fresh-derivation artifact per this protocol; the timestamped + raw output is outside. - **Auto-pilot's summary report** at `docs/skill-optimizer//autopilot-summary-.md` is - timestamped per run rather than versioned. Each auto-pilot run - produces a fresh summary; old ones are preserved naturally. -- **`vendored-skill/`** (the cached upstream skill source) is reused - across iterations of the same slug unless the source URL changed. - Not versioned; the source URL itself is the identity. + timestamped per run. Each auto-pilot invocation produces a fresh + summary; old ones are preserved naturally. +- **`vendored-skill/`** (the cached upstream skill source) is + reused across iterations of the same slug unless the source URL + changed. The source URL itself is the identity; no versioning. From 6e65c1818336de4d4d0fc6b678a8189e57e832a6 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 20 May 2026 12:30:21 -0500 Subject: [PATCH 025/121] feat(v1.4-chain): rewrite B1-B5 SKILL.md for git-native model MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Apply the filesystem-as-state + git-native iteration redesign across all five existing chain SKILL.md files. Drops version: frontmatter fields and archive/ directory references everywhere; reframes the subagent constraints in terms of "don't read own canonical / git history" (fresh-derivation) vs "do read current tree as state" (maintenance). B1 (investigate-functionality, fresh-derivation): drop version field; subagent constraint becomes "don't read own canonical or git history of it" instead of "don't read archive/". B2 (investigate-test-case, maintenance — major rewrite): output shape changed entirely. Was a single 02-test-case.md with picked: [] frontmatter array; now produces 00-test-proposals.md (one-time audit report) plus tests//spec.yaml per proposed functionality, each with picked: true|false in its own frontmatter. User gate is "edit picked: in each spec.yaml" rather than "tell me which names to pick". Subagent reads existing tests/ tree as load-bearing state per the maintenance rule. B3 (investigate-submissions, fresh-derivation): drop version and inputs.step_1_functionality fields; staleness now via git mtime against 01-functionality.md. B4 (write-tests, maintenance — major rewrite): replaced 04-tests-plan.md + workbench/ with tests/// probe folders. State is implicit: probe folder + grader.mjs present = built; no built_cases: [] array. Per-probe parallel test-writer dispatch; one probe = one test-writer subagent. Step generates tests/suite.yml from picked-functionality probes. Destructive-edit checkpoint pattern: operator commits before rebuilding existing probes. B5 (run-bench, fresh-derivation summary + timestamped raw): drop version and inputs.step_4_tests fields; reads tests/suite.yml as input. Summary 05-bench-summary.md is single-canonical with prior state in git; raw 05-bench-results// stays naturally accumulated and outside the protocol. Net: less bookkeeping, simpler mental model, fewer ways to get state out of sync. Filesystem IS the state across the chain. --- .../SKILL.md | 70 ++- .../SKILL.md | 56 +- .../SKILL.md | 357 +++++++------ skills/skill-optimizer-run-bench/SKILL.md | 152 +++--- skills/skill-optimizer-write-tests/SKILL.md | 501 +++++++++--------- 5 files changed, 579 insertions(+), 557 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 668ea24..8572d06 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -15,10 +15,10 @@ document every later step consumes. Two load-bearing pieces of context to load NOW, before the workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for archive - logic, version bumps, operator-directive handling, and the - constraints subagents must obey. Workflow step (e) below requires - it; loading it now lets you execute that step deterministically + Every chain skill applies it on every invocation — for staleness + detection, operator-directive handling, and the constraints + subagents must obey. Workflow step (e) below requires it; + loading it now lets you execute that step deterministically instead of improvising. 2. **You will dispatch a subagent for the actual research; you do NOT @@ -41,19 +41,12 @@ The report has structured frontmatter: ```yaml --- -version: 1 skill_source: pr_submission_intent: true | false classification: --- ``` -**`version`** — per-report iteration counter. `1` on the first run -for this slug; bumps on each re-run after the prior version is -archived. Step 1 has no upstream reports, so the `inputs:` field -(used by downstream steps to record what versions they consumed) is -omitted here. - **`classification`** — pick the most accurate label for what kind of skill this is. Canonical types: `tool-use` (procedures for using a specific tool, library, or API), `code-patterns` (code-level patterns @@ -128,21 +121,23 @@ Report path: `docs/skill-optimizer//01-functionality.md`. ### (e) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. The protocol is mechanical and identical -across all chain skills — it tells you how to check for an existing -`01-functionality.md`, archive it if present, compute the new -version, and collect operator directives. Do not improvise; read the -file. - -For step 1 specifically: there are no upstream chain reports, so the -"check upstream version match" branch of the protocol is moot — the -existing report is current if it exists at all (you can't be stale -against nothing). Re-run when the user wants fresh research with new -directives, when the source URL changed, or when the user's PR-intent -answer changed. - -When (e) is done, you have `new_version` and `${OPERATOR_DIRECTIVES}` -ready for (f). +and apply it for this step. Step 1 is a **fresh-derivation** step +— each invocation derives the report from the current source skill +plus any operator directives, with no continuity from prior outputs. There is no +`version:` field to bump and no `archive/` directory to manage; +git history captures prior states automatically. + +For step 1 specifically: there are no upstream chain reports, so +staleness is determined by whether the user wants fresh research +(source URL changed, PR-intent answer changed, or new directives). +Re-run on user signal, not on automatic upstream-mismatch detection. + +Collect `${OPERATOR_DIRECTIVES}` per the protocol's section if the +user supplied atomic new requirements (e.g., "you missed the vendor +CLA requirement"). If `01-functionality.md` already exists and will +be overwritten, that's fine — the prior state lives in git history. +No checkpoint commit is needed for fresh-derivation steps (the prior +state is self-contained in its own prior commit). ### (f) Dispatch the functionality-researcher subagent @@ -151,21 +146,22 @@ functionality-researcher subagent via the `Agent` tool (with worktree isolation if your environment supports it). Load the prompt template at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md), substitute the templated inputs (`${SKILL_SOURCE}`, `${OUTPUT_PATH}`, -`${PR_SUBMISSION_INTENT}`, `${VERSION}` from (e), -`${OPERATOR_DIRECTIVES}` from (e), vendored path), and dispatch. +`${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}` from (e), +vendored path), and dispatch. The subagent sees: - The vendored skill files (or the local skill path) - Targeted web-search / web-fetch results for the underlying technology -- The output path, the version number, and the frontmatter fields you - computed +- The output path and the frontmatter fields you computed - `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new requirements (may be empty) The subagent does NOT see: -- Prior `01-functionality.md` drafts (the ones now in `archive/`) +- Its own prior `01-functionality.md` (anti-ducktape: must derive + fresh, not rationalize what was there before) +- Git history of `01-functionality.md` - Existing analyses or improvement proposals for this skill - Prior tests or failure data - The wider chain's context @@ -211,19 +207,21 @@ this chain is built to prevent. Dispatch. existing copy by default (it's the same source). Re-fetch only if the source URL itself changed or the user explicitly asks. - **User changes their mind on PR intent later** — they re-run this - skill (see iteration section below); the prior report is archived - and the new one captures the updated PR intent. + skill (see iteration section below); the canonical + `01-functionality.md` is overwritten and the prior state lives in + git history. ## Iteration behavior -Step 1 is re-runnable. Re-run triggers specific to this step: +Step 1 is re-runnable and is a fresh-derivation step. Re-run +triggers specific to this step: - The source URL changed (different upstream skill, or repo moved) - The user changed their mind on PR intent - The user wants the report re-derived with new directives ("you missed the vendor CLA requirement", "go deeper on who-uses-this") -General iteration mechanics — archive, version bump, operator -directives, cascading staleness — live in +General iteration mechanics — staleness detection (git-native), +operator directives, cascading staleness — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). Workflow step (e) above already requires reading that file. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 142d5ca..fa86c98 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -20,10 +20,10 @@ external consistency check. Two load-bearing pieces of context to load NOW, before the workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for archive - logic, version bumps, operator-directive handling, and the - constraints subagents must obey. Workflow step (b) below requires - it; loading it now lets you execute that step deterministically + Every chain skill applies it on every invocation — for staleness + detection, operator-directive handling, and the constraints + subagents must obey. Workflow step (b) below requires it; + loading it now lets you execute that step deterministically instead of improvising. 2. **You will dispatch a subagent for the actual research; you do NOT @@ -46,9 +46,6 @@ The report has structured frontmatter: ```yaml --- -version: 1 -inputs: - step_1_functionality: upstream_repo: / upstream_branch_target: license: @@ -89,22 +86,24 @@ Two prerequisites: skipped for local-only optimization runs and refer them to step 4 (`write-tests`). -Record the functionality report's `version` — you'll pass it as -`inputs.step_1_functionality` later. - ### (b) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. The protocol tells you how to check for -an existing `03-submissions.md`, compare its -`inputs.step_1_functionality` against the current `01-functionality.md` -version, archive if stale, compute the new version, and collect -operator directives. Do not improvise; read the file. +and apply it for this step. Step 3 is a **fresh-derivation** step — +each invocation re-derives the report from current upstream facts +plus any operator directives. There is no `version:` field to bump and no +`archive/` directory; git history captures prior states. + +Staleness against `01-functionality.md` is determined by git-mtime +comparison per the protocol: if `01-functionality.md` is newer than +`03-submissions.md`, the existing report is stale (the source URL +may have changed) and should be re-derived. Otherwise the existing +report is current. Note: this report rarely needs re-running. Upstream PR conventions -change slowly. The most common reason to re-run is that the upstream -repo updated its `CONTRIBUTING.md` or CLA requirements — surface this -as an operator directive if you know it. +change slowly. The most common reason to re-run is that the +upstream repo updated its `CONTRIBUTING.md` or CLA requirements — +surface this as an operator directive if you know it. ### (c) Dispatch the submission-researcher subagent @@ -115,8 +114,7 @@ prompt template at [`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md), substitute the templated inputs (`${UPSTREAM_REPO}` from `01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, -`${VERSION}` from (b), `${OPERATOR_DIRECTIVES}` from (b)), and -dispatch. +`${OPERATOR_DIRECTIVES}` from (b)), and dispatch. The subagent sees: @@ -129,11 +127,13 @@ The subagent sees: the same skill category) - `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new requirements (may be empty) -- The output path and version number +- The output path The subagent does NOT see: -- Prior `03-submissions.md` drafts (the ones now in `archive/`) +- Its own prior `03-submissions.md` (anti-rationalization: must + derive fresh from upstream facts) +- Git history of `03-submissions.md` - Any information about the proposed change being optimized (the report is purely about upstream facts, not about whether a specific change will be accepted) @@ -225,17 +225,17 @@ mode this chain is built to prevent. Dispatch. ## Iteration behavior -Step 3 is re-runnable but rarely needs it. Re-run triggers specific -to this step: +Step 3 is re-runnable but rarely needs it, and is a fresh-derivation +step. Re-run triggers specific to this step: - The upstream repo updated its `CONTRIBUTING.md`, license, or CLA requirements - The upstream's PR-shape conventions visibly shifted (recent merged PRs no longer match the older patterns) -- Step 1's version bumped because the source URL changed (different - upstream repo entirely) +- Step 1's `01-functionality.md` changed because the source URL was + updated (different upstream repo entirely) -General iteration mechanics — archive, version bump, operator -directives, cascading staleness — live in +General iteration mechanics — staleness detection (git-native), +operator directives, cascading staleness — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). Workflow step (b) above already requires reading that file. diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index ea20e1c..35b8b0a 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -7,186 +7,203 @@ description: Use when the user wants to design or propose test cases for a skill Step 2 of the skill-optimizer chain. Takes the functionality report from step 1, dispatches a designer subagent to enumerate the skill's -responsibilities and propose a ranked list of test cases, then asks -the user to pick which subset to actually build. Writes -`docs/skill-optimizer//02-test-case.md`. +responsibilities and propose a ranked set of **functionalities** to +test (each functionality = one responsibility the skill must +fulfill), then asks the user to pick which ones to actually build +probes for at step 4. Writes a `tests//spec.yaml` +file per proposed functionality (the filesystem IS the state) plus a +one-time audit report at `00-test-proposals.md`. ## Before you start Two load-bearing pieces of context to load NOW, before the workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for archive - logic, version bumps, operator-directive handling, and the - constraints subagents must obey. Workflow step (b) below requires - it; loading it now lets you execute that step deterministically - instead of improvising. + Every chain skill applies it on every invocation — for re-run + logic, the constraints subagents must obey, and safe destructive + edits. Workflow step (b) requires it. 2. **You will dispatch a subagent for the actual design work; you do - NOT enumerate responsibilities or design cases yourself in this - session.** Workflow step (c) is the dispatch. The rationale is in - "Why limited-context dispatch matters" below — read it if you're - tempted to skip the dispatch. + NOT enumerate responsibilities or design proposals yourself in + this session.** Workflow step (c) is the dispatch. The rationale + is in "Why limited-context dispatch matters" below — read it if + you're tempted to skip the dispatch. -Throughout this document, "step 1" through "step 7" (no parens) refer -to skills in the chain (this skill is step 2; `investigate-functionality` -is step 1; etc.). Internal workflow steps within THIS skill are -labelled "(a)" through "(f)" to avoid the collision. +Throughout this document, "step 1" through "step 7" (no parens) +refer to skills in the chain (this skill is step 2; step 1 is +`investigate-functionality`; etc.). Internal workflow steps within +THIS skill are labelled "(a)" through "(f)". ## What you produce -A single report at `docs/skill-optimizer//02-test-case.md`, -where `` matches the slug from step 1's report. - -The report has structured frontmatter: - -```yaml ---- -version: 1 -inputs: - step_1_functionality: -picked: [] # operator fills this AFTER the user-gate step ---- -``` - -**`picked`** — list of test case names the user selected from the -ranked proposals. Empty until step (e) collects the user's picks. -Step 4 (`write-tests`) operates on this subset, NOT on all proposals. - -For the body template (what each proposed case must include and how -they're ranked), see +Two artifacts at `docs/skill-optimizer//`, where `` +matches the slug from step 1's report: + +1. **`00-test-proposals.md`** — a one-time audit report from the + subagent: the ranked list of proposed functionalities with full + reasoning (why each one matters, what coverage it adds, what + probes are suggested). This file is for human review of the + design reasoning; downstream steps do NOT read it. The subagent + rewrites it on re-runs (fresh top-to-bottom proposal each time, + anchored by the current `tests/` tree state). + +2. **`tests//spec.yaml`** — one folder per + proposed functionality. The folder name is a filesystem-safe + slug derived from the functionality's name. The `spec.yaml` + inside has: + + ```yaml + name: refuses-malformed-input + description: The skill refuses input that violates its expected schema. + picked: false # user flips to true after the user gate + importance: high # high | medium | low + suggested_probes: + - malformed-json + - missing-required-field + - type-mismatch + why_test: > + Without this guard, broken upstream data silently corrupts + the skill's downstream logic. + ``` + + The filesystem IS the state — there is no `picked: []` array + anywhere. Step 4 builds probes only for functionalities whose + `spec.yaml` has `picked: true`. + +For the body template of `00-test-proposals.md` (sections, ranking +format) and the exact `spec.yaml` field list, see [`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md). -The subagent writes the ranked proposal body; the operator session -sets the `picked` field after the user gate. +The subagent writes both the audit report and the per-functionality +spec.yaml files; the operator session runs the user gate to flip +`picked` values. ## Workflow ### (a) Confirm prerequisites -`docs/skill-optimizer//01-functionality.md` must exist and have -valid frontmatter. If it doesn't, tell the user to run -`skill-optimizer-investigate-functionality` first and stop here. -Record its `version` field — you'll pass it as -`inputs.step_1_functionality` later so this report records which -version of the functionality understanding it was derived from. +`docs/skill-optimizer//01-functionality.md` must exist. If it +doesn't, tell the user to run `skill-optimizer-investigate-functionality` +first and stop here. + +If `docs/skill-optimizer//tests/` doesn't exist yet (first +invocation on this slug), create the empty directory — the subagent +will populate it. ### (b) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** and apply it for this step. Step 2 is a **maintenance** step (see -the protocol's "Step kinds" section) — re-runs extend the current -`02-test-case.md` rather than replacing it. Concretely, the protocol -tells you to: - -- Check whether `02-test-case.md` exists and whether - `inputs.step_1_functionality` matches the current - `01-functionality.md` version -- Copy the current file to `archive/02-test-case-v.md` (audit - trail — inert from this point) -- Pass the canonical `02-test-case.md` itself to the subagent as - load-bearing input (it extends or modifies the existing list, - preserving the user's existing `picked` choices unless directives - say otherwise) -- Bump the version number on the new canonical file - -Special case: if `01-functionality.md`'s version bumped (step 1 -re-ran with a substantively different responsibility set), existing -cases may no longer align. Surface this to the user before the -dispatch — they should decide whether to keep the existing list and -add cases, or start fresh (in which case delete the canonical file -before dispatching so the subagent treats this as iteration 1). +the protocol's "Step kinds" section) — re-runs read the current +`tests/` tree as state and extend or modify it. Concretely: + +- If the user supplied directives (revisions, new requirements), + collect them as the bulleted `${OPERATOR_DIRECTIVES}` slot per + the protocol's "Collect operator directives" section. Atomic new + requirements only — never a context dump of the current tree. +- If a directive will be destructive (e.g., "remove the X + functionality I no longer think we should test"), the operator + session commits a checkpoint of the current `tests/` tree before + dispatching, per the protocol's "Safe destructive edits" + section. This gives git history a clean before/after breakpoint. + +The subagent will read the current `tests//spec.yaml` +files as input (this is its load-bearing state, per the maintenance +rule). It will NOT walk git history of those files. ### (c) Dispatch the test-case-designer subagent -**Do NOT enumerate responsibilities or design cases yourself in this -session.** Dispatch the test-case-designer subagent via the `Agent` -tool (with worktree isolation if your environment supports it). Load -the prompt template at +**Do NOT enumerate responsibilities or design proposals yourself in +this session.** Dispatch the test-case-designer subagent via the +`Agent` tool (with worktree isolation if your environment supports +it). Load the prompt template at [`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md), substitute the templated inputs (`${FUNCTIONALITY_PATH}`, -`${OUTPUT_PATH}`, `${EXISTING_CASES_PATH}` if a current -`02-test-case.md` exists (else empty), `${VERSION}` from (b), -`${OPERATOR_DIRECTIVES}` from (b)), and dispatch. +`${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, `${OPERATOR_DIRECTIVES}`), +and dispatch. The subagent sees: -- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` -- `${EXISTING_CASES_PATH}` — the current `02-test-case.md` (only - on re-runs; empty on iteration 1). The subagent preserves existing - cases verbatim unless a directive explicitly asks to revise a - specific case; new cases are appended per directives. -- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new +- `${FUNCTIONALITY_PATH}` — the current `01-functionality.md` +- `${TESTS_TREE_PATH}` — the current `tests/` tree (may be empty + on first invocation). The subagent reads every existing + `tests//spec.yaml` as load-bearing state per the + maintenance rule. +- `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new requirements (may be empty) -- The output path and version number +- Output paths for `00-test-proposals.md` and the `tests/` tree The subagent does NOT see: -- Anything under `archive/` (prior versions are inert audit trail) -- Any `06-analysis.md` (existing analyses) -- `07-improvement-proposal.md`, `07-validator-verdict.md` +- The skill's source content (would gerrymander tests around the + source's literal phrasing — design coverage from STATED + responsibilities) +- Any `06-analysis.md`, `07-improvement-proposal.md`, + `07-validator-verdict.md` - Raw failure data, `findings.txt`, bench results -- The skill's source content itself +- Git history of any tree file or its own audit report -The "no source content" constraint is load-bearing: design coverage -from the skill's STATED responsibilities, not from the source's -literal phrasing. Letting the subagent read the source lets it -gerrymander tests around the source's exact shape (testing what the -prose happens to say, not what the skill is supposed to do). +The subagent writes: -The subagent writes the (extended) proposal body itself and returns -a brief summary including the top-3 cases by importance and a list -of which existing case names were preserved unchanged vs which were -revised. +- `00-test-proposals.md` (the audit report, full ranked design + reasoning) +- One `tests//spec.yaml` per proposed + functionality. Existing spec.yaml files the user has already + edited (e.g., `picked: true` set, custom suggested_probes added) + are preserved verbatim unless a directive explicitly targets + that functionality for revision. + +It returns a brief summary: top-3 functionalities by importance, +list of which existing functionalities were preserved unchanged vs +revised, list of newly-added functionalities. ### (d) Confirm subagent output -Verify the report file exists, the frontmatter parses, and the body -has the expected sections per the subagent prompt template. The -`picked` field at this point is still empty — that's correct; the -user gate is next. +Verify: + +- `00-test-proposals.md` exists and parses as markdown. +- Each `tests//spec.yaml` parses as YAML and + has the required fields (`name`, `description`, `picked`, + `importance`, `suggested_probes`, `why_test`). + +If any spec.yaml fails to parse or is missing required fields, +surface the issue to the user; don't try to repair the subagent's +output yourself. ### (e) User gate: present proposals, collect picks -Show the user the ranked proposal (summarize or paste the body — -your call based on length). Ask: +Show the user the ranked list from `00-test-proposals.md` (paste +the audit report, or summarize if long — your call). Tell them +how to indicate picks: -> Here are the proposed test cases ranked by importance. Which ones -> should we actually build? You can pick all of them, a subset, or -> ask for a revised proposal. +> Here are the proposed functionalities ranked by importance. To +> pick the ones you want me to build probes for at step 4, edit +> the `picked:` field in each `tests//spec.yaml` +> to `true`. You can also edit `suggested_probes` if you want to +> change the probe set, or tell me to dispatch a revised proposal. Three realistic responses: -1. **User picks a subset (or all).** Update the `picked` frontmatter - field with the chosen case names. Commit the file. -2. **User wants additions or revisions** (e.g., "add null-input - coverage", "the case named X is too vague — split it into two", - "make case Y harder"). Treat their feedback as new operator - directives, return to (b) to apply the maintenance protocol - (which archives the current version), and re-dispatch the - subagent at (c). Because this is a maintenance step, existing - cases the user already picked are preserved unless their - directive explicitly asks otherwise — the user does not need to - re-pick everything. - - **Case-rename convention for revisions** — when a directive - asks to revise an existing case (rather than add a new one), the - test-case-designer subagent appends a version suffix to the - case's name: `case-x` → `case-x-v2` → `case-x-v3`. The bare - name is implicit v1; revisions always get a numeric suffix. The - subagent updates the body of `02-test-case.md` to use the new - name; the operator session updates `picked` to swap the old - name for the new one. This convention is load-bearing for - step 4's diff logic — without a rename, step 4 would see "case - already built" by name match and skip rebuilding. The convention - gives step 4 a deterministic signal that this case needs a - fresh test-writer dispatch. +1. **User edits `picked: true|false` in the spec.yaml files + directly.** Read each spec.yaml after their edits to confirm + the picked set. Commit the final state of the tree. +2. **User wants additions or revisions** (e.g., "add coverage for + null-input edge cases", "the X functionality is too broad — + split it into two"). Treat their feedback as new operator + directives. Return to (b) to apply the iteration protocol + (which checkpoints the current state via git commit if the + change is destructive) and re-dispatch the subagent at (c). + Because this is a maintenance step, existing spec.yaml files + the user has already edited (especially `picked: true` ones) + are preserved unless their directive explicitly targets a + specific functionality. 3. **User picks zero.** Ask whether they want a revised proposal (treat as case 2) or are abandoning the optimization for this skill (exit honestly without progressing to step 4). -Don't auto-pick on the user's behalf — even if all cases look -important, the user owns this decision (they're paying for the -test-writing in step 4 and the bench run in step 5). +Don't auto-flip `picked` on the user's behalf — even if all +functionalities look important, the user owns this decision +(they're paying for the probe-building in step 4 and the bench +run in step 5). ### (f) Hand off @@ -195,12 +212,14 @@ Read `01-functionality.md`'s `pr_submission_intent` field. If `pr_submission_intent: true`: > Next, invoke `skill-optimizer-investigate-submissions`. After -> that, invoke `skill-optimizer-write-tests` with the picked subset. +> that, invoke `skill-optimizer-write-tests` to build probes for +> the picked functionalities. If `pr_submission_intent: false`: > Skipping step 3 (no PR submission planned). Next, invoke -> `skill-optimizer-write-tests` with the picked subset. +> `skill-optimizer-write-tests` to build probes for the picked +> functionalities. No late "submit a PR?" prompts — the decision was made at step 1. @@ -209,53 +228,63 @@ No late "submit a PR?" prompts — the decision was made at step 1. If the operator session is the one designing tests, it has already absorbed prior failure data, the optimizer's past attempts, the validator's verdicts, and whatever framing the user has applied. -That context biases test design toward "tests that would have caught -the things I just watched fail" — the coverage version of ducktape. -You end up with a test suite that validates the patch instead of -testing the skill independently. - -A subagent that sees only `01-functionality.md` reasons from the -skill's stated responsibilities, not from prior outcomes. The -proposal is coverage-oriented, not regression-defensive. That's the -invariant we need to ship a fair test set. - -If you find yourself thinking "I'll just propose the tests myself, -the subagent dispatch is bureaucratic overhead" — that's the failure -mode this chain is built to prevent. Dispatch. +That context biases test design toward "tests that would have +caught the things I just watched fail" — the coverage version of +ducktape. You end up with a test suite that validates the patch +instead of testing the skill independently. + +A subagent that sees only `01-functionality.md` and the current +`tests/` tree reasons from the skill's stated responsibilities, +not from prior outcomes. The proposal is coverage-oriented, not +regression-defensive. That's the invariant we need for a fair +test set. + +If you find yourself thinking "I'll just propose the +functionalities myself, the subagent dispatch is bureaucratic +overhead" — that's the failure mode this chain is built to +prevent. Dispatch. ## Edge cases - **`01-functionality.md` missing** — tell the user to run step 1 first; don't try to derive responsibilities yourself. -- **Subagent's proposal has fewer cases than expected** — that's a - signal the functionality report is thin. Surface this to the user +- **Subagent's proposal has fewer functionalities than expected** + — signal the functionality report is thin. Surface to the user with two options: re-run step 1 with directives ("the report missed responsibility X") or accept the small proposal if the skill is genuinely small. Do not re-invoke step 1 yourself — - re-runs require an active signal from the user. (Same rule for - any backward trigger, per the iteration protocol's - "Re-run authorization" section.) -- **User wants to add cases the subagent didn't propose** — accept - them as additional entries in `picked`; step 4 will treat them - the same as subagent-proposed cases. + re-runs require an active signal from the user (per the iteration + protocol's "Re-run authorization" section). +- **User wants to add a functionality the subagent didn't propose** + — accept it. Either (i) they describe it and you create the + `tests//spec.yaml` manually with their content + `picked: + true`, or (ii) they re-dispatch the subagent with a directive + ("add a functionality for X"). Either path is valid. - **Operator directives contradict each other** (e.g., "focus coverage on X" + "ignore X") — surface the contradiction to the user before re-dispatching; don't try to resolve it yourself. +- **User removes a functionality folder manually** — fine, that's + filesystem-as-state working as intended. On next re-run, the + subagent reads the current tree and doesn't propose the removed + one back unless directives ask. ## Iteration behavior -Step 2 is re-runnable. Re-run triggers specific to this step: - -- The user wants a different cut of coverage (different cases - proposed, different ranking) -- Step 4 (`write-tests`) flagged a picked case as unimplementable -- Step 5 (`run-bench`) showed all tests pass on the baseline — too - easy, need harder cases -- Step 6 (`analyze-result`) surfaced a coverage gap (responsibility - X has no tests; user wants one) -- Step 1's version bumped (functionality understanding changed) - -General iteration mechanics — archive, version bump, operator -directives, cascading staleness — live in +Step 2 is a maintenance step. Re-run triggers specific to this +step: + +- The user wants different coverage (different proposed + functionalities, different ranking, different probe suggestions) +- Step 4 (`write-tests`) flagged a picked functionality as + unimplementable +- Step 5 (`run-bench`) showed all probes pass on baseline — too + easy, need a different cut of coverage +- Step 6 (`analyze-result`) surfaced a coverage gap (an + unrepresented responsibility; user wants one tested) +- Step 1's `01-functionality.md` changed (functionality + understanding updated) + +General iteration mechanics — staleness detection, operator +directives, cascading staleness, safe destructive edits — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). Workflow step (b) above already requires reading that file. diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index dcb956c..a6765f2 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -6,9 +6,9 @@ description: Use when the user wants to run the eval suite against a skill and c # skill-optimizer-run-bench Step 5 of the skill-optimizer chain. Invokes the skill-optimizer CLI's -`run-suite` command against the workbench built in step 4, captures the -raw results under a timestamped directory, and writes a small versioned -summary report that step 6 (`analyze-result`) reads. No subagent +`run-suite` command against `tests/suite.yml` (generated by step 4), +captures the raw results under a timestamped directory, and writes a +small summary report that step 6 (`analyze-result`) reads. No subagent dispatch — this is a thin operator-driven CLI step. ## Before you start @@ -17,10 +17,12 @@ One load-bearing piece of context to load NOW, before the workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** Every chain skill applies it on every invocation. Step 5 is a - slightly unusual case because the bench's raw output is timestamped - rather than versioned (each run preserved naturally under its own - directory) — but the summary report `05-bench-summary.md` IS - versioned and follows the protocol's archive-and-version mechanics. + slightly unusual case because the bench's raw output is + timestamped (each run preserved naturally under its own directory) + — outside the iteration protocol. The summary report + `05-bench-summary.md` is a fresh-derivation artifact per the + protocol: each invocation overwrites the canonical with a fresh + write-up of the new raw run; prior state lives in git history. Workflow step (b) below requires it. Throughout this document, "step 1" through "step 7" (no parens) refer @@ -36,75 +38,67 @@ the slug from step 1's report: 1. **`05-bench-results//`** — the raw run output from the CLI: `suite-result.json`, per-trial `trace.jsonl`, per-trial `findings.txt`, and any preserved workspaces. Timestamped per the - invocation; old runs are NEVER overwritten. The protocol's archive - convention does not apply here — each timestamped directory is - itself the archive. + invocation; old runs are NEVER overwritten. Each timestamped + directory is its own naturally-accumulated archive — outside the + iteration protocol. -2. **`05-bench-summary.md`** — a versioned report with frontmatter: +2. **`05-bench-summary.md`** — a single canonical report with + frontmatter: ```yaml --- - version: - inputs: - step_4_tests: bench_results_path: 05-bench-results// overall_pass_rate: --- ``` The body is a small aggregate that step 6 reads as the entry point: - the per-model pass rates, per-case pass rates, a list of cases + the per-model pass rates, per-probe pass rates, a list of probes where at least one trial failed (so the analyzer knows where to focus), and a pointer to the timestamped raw directory. The - summary does NOT include trace excerpts or findings detail — those - stay in the raw output and the analyzer reads them directly when - it needs them. + summary does NOT include trace excerpts or findings detail — + those stay in the raw output and the analyzer reads them directly + when it needs them. Prior versions of the summary live in git + history. ## Workflow ### (a) Confirm prerequisites -`docs/skill-optimizer//04-tests-plan.md` must exist and have -valid frontmatter, and `docs/skill-optimizer//workbench/` must -exist with a `suite.yml` and per-case dirs (per step 4's output -contract). If either is missing, tell the user to complete step 4 -first and stop here. +`docs/skill-optimizer//tests/suite.yml` must exist (per step +4's output contract). If it's missing, tell the user to complete +step 4 first and stop here. Also confirm the vendored skill source exists at `docs/skill-optimizer//vendored-skill/` (per step 1's setup); the workbench's `suite.yml` points at it as the skill under test. -Record `04-tests-plan.md`'s `version` field — you'll pass it as -`inputs.step_4_tests` in the summary. - ### (b) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it to the summary file. Step 5's behavior: +and apply it. Step 5's behavior: - For the **raw bench output** at `05-bench-results//`: - always write to a fresh timestamped directory (`date +%Y%m%dT%H%M%S` - or similar). No archive logic — each run is its own preserved - artifact. Never overwrite a prior timestamped dir. - -- For the **summary file** at `05-bench-summary.md`: apply the - protocol's standard archive-and-version flow. - - If the summary doesn't exist yet, set `new_version = 1`. - - If it exists and `inputs.step_4_tests` matches the current - `04-tests-plan.md` version AND the operator did not pass a - re-run directive, the existing summary is current — exit - honestly and tell the user the latest results are still good. - - Otherwise, copy `05-bench-summary.md` to - `archive/05-bench-summary-v.md` (audit trail) - and set `new_version = existing_version + 1`. The new summary - will be written from the new raw run, with no reference to the - prior summary's body. - -The summary is conceptually fresh-derivation — its body is a -mechanical write-up of the new raw bench run, not an extension of -the prior summary. There is no subagent, but the same anti-ducktape -principle applies in spirit: write what this run shows, not what the -prior summary said. + always write to a fresh timestamped directory (`date +%Y%m%dT%H%M%SZ` + in UTC). Each run is its own preserved artifact — outside the + iteration protocol. Never overwrite a prior timestamped dir. + +- For the **summary file** at `05-bench-summary.md`: it's a + fresh-derivation artifact. Each invocation overwrites the + canonical with a mechanical write-up of the new raw run; prior + state lives in git history. No `archive/` involvement. + + Staleness check (git-native): if `tests/suite.yml` has a newer + git mtime than `05-bench-summary.md`, the existing summary + measured a different workbench and is stale. If the user also + didn't pass a re-run directive AND the summary appears current + (its git mtime is newer than `tests/`), tell them the latest + results are still good and exit honestly. Otherwise proceed + with a new bench run. + +The summary's anti-ducktape principle: write what THIS run shows, +not what the prior summary said. There is no subagent, but the +discipline is the same. ### (c) Run the bench @@ -116,7 +110,7 @@ OUT_DIR="docs/skill-optimizer//05-bench-results/${TIMESTAMP}" mkdir -p "${OUT_DIR}" npx tsx /src/cli.ts run-suite \ - docs/skill-optimizer//workbench/suite.yml \ + docs/skill-optimizer//tests/suite.yml \ --out "${OUT_DIR}" \ --trials 3 ``` @@ -144,17 +138,18 @@ produce" plus a body containing: - **Overall:** total trials, passed, failed, overall pass rate. - **Per model:** for each model in `suite.yml`, the trial count and pass rate. -- **Per case:** for each case, the trial count, pass rate, and a - one-line note if at least one trial failed (so step 6 knows where - to look). -- **Failed-case pointer list:** explicit list of case IDs where any - trial failed, with paths to their `trace.jsonl` and `findings.txt` - under `${OUT_DIR}`. +- **Per probe:** for each probe (one entry per + `tests///` in the suite), the trial count, + pass rate, and a one-line note if at least one trial failed (so + step 6 knows where to look). +- **Failed-probe pointer list:** explicit list of probe IDs where + any trial failed, with paths to their `trace.jsonl` and + `findings.txt` under `${OUT_DIR}`. - **Raw output:** the value of `bench_results_path` from the frontmatter (relative path to the timestamped dir). The summary is the entry point step 6 reads. Step 6's analyzer -subagent will follow the failed-case pointer list into the raw +subagent will follow the failed-probe pointer list into the raw output for trace and findings detail; the summary itself stays short (an aggregate, not a dump). @@ -184,8 +179,8 @@ own. Surface the choice to the user. ## Edge cases -- **`04-tests-plan.md` missing or `workbench/` empty** — tell the - user to run step 4 first; don't try to construct a workbench +- **`tests/suite.yml` missing or `tests/` empty** — tell the user + to run step 4 first; don't try to construct a suite manifest yourself. - **`OPENROUTER_API_KEY` not set** — the CLI fails fast; surface the error and tell the user to set the env var. Don't fall back to a @@ -206,34 +201,33 @@ own. Surface the choice to the user. Step 5 is re-runnable. Re-run triggers specific to this step: -- `04-tests-plan.md`'s version bumped (step 4 re-ran with new or - revised picked cases) — the prior bench measured a different - workbench -- User wants fresh results against the same workbench (e.g., +- `tests/` changed (step 4 added or revised probes) — the prior + bench measured a different probe set, detected via git mtime +- User wants fresh results against the same probes (e.g., flakiness suspicion, model availability changed, model list in `suite.yml` changed) -- Step 6 (`analyze-result`) flagged a result as inconclusive due to - too few trials — user may re-run with higher `--trials` +- Step 6 (`analyze-result`) flagged a result as inconclusive due + to too few trials — user may re-run with higher `--trials` -By default, a re-run measures the **entire** workbench — even cases -that didn't change at step 4. This gives unchanged cases fresh -trial samples (useful for distinguishing flaky from systematic at -step 6) and keeps the summary internally comparable. The cost is -"all cases × trials" per re-run. +By default, a re-run measures the **entire** probe set — even +probes that didn't change at step 4. This gives unchanged probes +fresh trial samples (useful for distinguishing flaky from systematic +at step 6) and keeps the summary internally comparable. The cost is +"all probes × trials" per re-run. **Known limitation (deferred):** the CLI's `run-suite` does not currently accept a case filter, so there is no first-class -"re-bench only the changed cases" mode at step 5. Operators with -expensive suites can run `run-case` manually for changed cases and +"re-bench only the changed probes" mode at step 5. Operators with +expensive suites can run `run-case` manually for changed probes and splice results into the prior `05-bench-results//` dir, but that's an outside-the-chain escape hatch, not a supported flow. Adding a case filter to `run-suite` (and threading it through step 5) -is on the workbench's roadmap; until then, expect re-runs to remeasure -the full suite. +is on the workbench's roadmap; until then, expect re-runs to +remeasure the full probe set. -General iteration mechanics — archive, version bump, cascading -staleness — live in +General iteration mechanics — git-native staleness, operator +directives, cascading staleness — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -The raw bench output is exempt from the standard archive flow (each -timestamped directory is its own preservation); only the summary -follows the protocol. +The raw bench output is outside the protocol (each timestamped +directory is its own preservation); the summary is a +fresh-derivation artifact and follows the protocol. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 14416b8..88cb560 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -1,17 +1,17 @@ --- name: skill-optimizer-write-tests -description: Use when the user wants to build or implement the eval workbench for a skill — phrases like "build the tests", "implement the workbench", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-investigate-test-case` has produced `02-test-case.md` with a picked subset, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the test-case proposals into concrete fixtures + graders should trigger this. +description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-investigate-test-case` has produced `tests//spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. --- # skill-optimizer-write-tests -Step 4 of the skill-optimizer chain. Takes the picked test cases -from step 2, dispatches one test-writer subagent per case (in -parallel) to build the concrete workspace files + grader scripts, -runs a smoke check to verify each grader against hand-crafted -GOOD/BAD/EMPTY fixtures, and writes -`docs/skill-optimizer//04-tests-plan.md` plus the -`workbench/` directory the run-bench step will execute against. +Step 4 of the skill-optimizer chain. Reads the picked functionalities +from step 2 (`tests//spec.yaml` files where `picked: +true`), decides the probe set per functionality (informed by each +spec's `suggested_probes` and any operator directives), dispatches +one test-writer subagent per probe (in parallel) to build the +concrete workspace files + grader scripts + smoke fixtures, then +generates `tests/suite.yml` for the run-bench step. ## Before you start @@ -21,72 +21,64 @@ workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** Every chain skill applies it on every invocation. Step 4 is a **maintenance** step (see the protocol's "Step kinds" section) - — re-runs extend the workbench rather than replacing it. - Workflow step (b) below requires the protocol; loading it now - lets you execute that step deterministically. + — the `tests/` tree accumulates probe folders across re-runs + rather than being rebuilt from scratch. Workflow step (b) + requires it. 2. **You will dispatch test-writer subagents (in parallel, one per - picked case); you do NOT write workspace files or grader scripts - yourself in this session.** Workflow step (d) is the dispatch. - The rationale is in "Why limited-context dispatch matters" below - — read it if you're tempted to skip the dispatch. + probe being built or rebuilt); you do NOT write workspace files + or grader scripts yourself in this session.** Workflow step (d) + is the dispatch. The rationale is in "Why limited-context + dispatch matters" below. 3. **Each test-writer subagent sees the skill's source content.** This is intentional and is the one step in the chain where the source is in-scope for a generative subagent. The constraint - comes from the test case spec from step 2 (which fixes what the - fixture should test); source access is needed for concrete - violation patterns and realistic fixture content. + comes from the probe spec (which fixes what the fixture should + test); source access is needed for concrete violation patterns + and realistic fixture content. Throughout this document, "step 1" through "step 7" (no parens) -refer to skills in the chain (this skill is step 4; -`investigate-test-case` is step 2; etc.). Internal workflow steps -within THIS skill are labelled "(a)" through "(f)" to avoid the -collision. +refer to skills in the chain (this skill is step 4; step 2 is +`investigate-test-case`; etc.). Internal workflow steps within +THIS skill are labelled "(a)" through "(f)". ## What you produce -Two outputs at `docs/skill-optimizer//`: - -1. **`04-tests-plan.md`** — the meta-report tracking which cases - have been built, with per-case implementation notes. - -2. **`workbench/`** — the directory the `run-bench` step executes - against. Contains `workbench/suite.yml` (suite manifest), per-case - workspace files, per-case grader scripts, and smoke-check - fixtures. - -`04-tests-plan.md` frontmatter: - -```yaml ---- -version: 1 -inputs: - step_1_functionality: - step_2_test_case: -built_cases: - - - - -last_smoke_check_passed: ---- -``` - -**`built_cases`** — the set of case names that currently have -workspace files + graders in `workbench/`. Re-runs add new entries -here as new picks get built; existing entries stay unless a -directive explicitly requests revision. - -**`last_smoke_check_passed`** — set when the smoke check (step (e)) -passes. If a re-run modifies a case and the smoke check fails, this -field stays at the prior timestamp until the smoke check passes -again on the updated grader. - -For the body template (per-case implementation notes — workspace -contents, expected agent behavior, grader logic, smoke-check -fixtures), see +Two kinds of output under `docs/skill-optimizer//`: + +1. **Probe folders** at + `tests///`. For each picked + functionality, one or more probe folders. Each probe folder + contains: + + ```text + tests/// + spec.yaml # probe-level intent (workspace overview, expected + # agent behavior, grader logic) + workspace/ # the files the agent sees in /work + grader.mjs # the grading script (or .py) + smoke/ + good/findings.txt # fixture that should PASS the grader + bad/findings.txt # fixture that should FAIL the grader + empty/findings.txt # fixture that should FAIL the grader + checks/smoke.mjs # runs the grader against the three smoke fixtures + ``` + + The filesystem IS the state — a probe exists if its folder + exists with these contents; there is no `built_cases: []` array + anywhere. The presence of `grader.mjs` and the absence of any + smoke-check failure marker indicates the probe is built and + verified. + +2. **`tests/suite.yml`** — generated from the current tree. Lists + every probe under every `picked: true` functionality. Step 5 + (run-bench) reads this. Regenerated by step (f) of this skill on + every run. + +The probe-level `spec.yaml` format (workspace overview, expected +agent behavior, grader logic, smoke-check fixtures) is defined in [`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md). -Each test-writer subagent contributes the section for its own case; -the operator session assembles them into `04-tests-plan.md`. For the workbench schema itself (`suite.yml` structure, grader contract, smoke-check format), see @@ -100,268 +92,277 @@ step. Three prerequisites: -1. `docs/skill-optimizer//01-functionality.md` exists with - valid frontmatter. -2. `docs/skill-optimizer//02-test-case.md` exists with valid - frontmatter and a non-empty `picked` field. -3. Every name in `picked` appears as a case definition in the body - of `02-test-case.md` (no stale references to cases that were - removed in a revision). +1. `docs/skill-optimizer//01-functionality.md` exists. +2. `docs/skill-optimizer//tests/` exists with at least one + `tests//spec.yaml` file having `picked: true`. +3. Every picked functionality's spec.yaml parses with the required + fields (`name`, `description`, `picked`, `importance`, + `suggested_probes`, `why_test`). If any prerequisite fails, surface the issue to the user — don't -attempt to derive missing cases or guess. Record both upstream -versions; you'll pass them as `inputs.step_1_functionality` and -`inputs.step_2_test_case` later. +attempt to derive missing fields or guess. List the picked +functionalities you found; the user should confirm before you +proceed. ### (b) Handle iteration (maintenance step) **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply the maintenance-step flow for this step: - -- Copy current `04-tests-plan.md` and the current `workbench/` - directory to `archive/04-tests-plan-v.md` and - `archive/workbench-v/` respectively (audit trail — inert from - this point). -- Compute the new version (`N+1`). -- Diff the `picked` set in `02-test-case.md` against the prior - `built_cases` (from the just-archived plan). Three categories: - - **New picks** (in `picked` but not in `built_cases`) — these - need test-writer dispatches in step (d). Revisions from step 2 - surface here too: per the rename convention, a revised case - arrives with a version-suffixed name (e.g., `case-x-v2`), so - the diff sees it as a fresh pick and dispatches normally. - - **Existing builds** (in both) — preserve their workbench files - and plan entries unchanged, unless `${OPERATOR_DIRECTIVES}` - explicitly names a case to revise. - - **Removed picks** (in `built_cases` but no longer in - `picked`) — the user de-picked these in a later step 2 revision, - or they were superseded by a renamed revision (e.g., `case-x` - superseded by `case-x-v2`). Leave the workbench files in place - but mark them removed from `built_cases` in the new plan (the - run-bench step's suite.yml will exclude them). - -If `02-test-case.md`'s version bumped substantively and many cases -were renamed or reframed, the diff may not match cleanly. Surface -this to the user before proceeding — they may want to delete the -canonical `04-tests-plan.md` + `workbench/` to start fresh, or -selectively rebuild specific cases. - -### (c) Plan workbench structure + user gate - -Before dispatching test-writers, sketch the planned workbench -structure: which case directories will be created, what the -`suite.yml` entries will look like, which existing files are -preserved. Show this to the user. Ask: - -> Here's the planned workbench layout for the picked cases. New -> cases I'll dispatch test-writers for: [list]. Existing cases -> preserved: [list]. Any cases I should revise instead of preserve? -> Anything to change before I dispatch? +and apply the maintenance-step flow: + +- Walk the current `tests/` tree. For each picked functionality, + determine which probes already have built folders (i.e., + `tests///grader.mjs` exists). Existing built + probes are preserved by default; the operator session does NOT + re-dispatch test-writers for them unless `${OPERATOR_DIRECTIVES}` + explicitly names them for revision. +- Collect `${OPERATOR_DIRECTIVES}` per the protocol's section. + Atomic new requirements only — never a context dump of existing + spec.yaml contents. +- If a directive will be destructive (rebuilding an existing + probe, removing a probe that's no longer wanted), commit a + checkpoint of the current `tests/` tree BEFORE dispatching the + rebuild, per the protocol's "Safe destructive edits" section. + Example commit message: `checkpoint: before rebuild of + tests/refuses-malformed-input/null/`. + +### (c) Plan probe set + user gate + +For each picked functionality: + +1. Read its `spec.yaml`'s `suggested_probes` field. +2. Decide the probe set: + - Start with `suggested_probes` as the default. + - If `${OPERATOR_DIRECTIVES}` adds probes for this functionality + (e.g., "add an empty-array probe to refuses-malformed-input"), + include them. + - If `${OPERATOR_DIRECTIVES}` removes or revises probes, + reflect that. + - Existing built probe folders in `tests///` + stay in the set unless directives target them for revision or + removal. + +Show the user the planned probe set per functionality: + +> For each picked functionality, here are the probes I'll build: +> +> - `refuses-malformed-input/`: malformed-json (existing), +> missing-field (existing), empty-array (new) +> - `uses-right-tool/`: tool-a-scenario (existing), +> tool-b-scenario (new) +> +> New probes to dispatch test-writers for: [list]. Existing probes +> preserved: [list]. Probes to rebuild (will commit checkpoint +> first): [list]. Anything to change before I dispatch? Three realistic responses: 1. **User approves.** Proceed to (d) with the dispatch set as planned. -2. **User adds revisions** (e.g., "the grader for case X is too - strict — relax the matching"). Treat as new operator directives, - mark the named cases for revision, and proceed to (d) with the - updated dispatch set. +2. **User adds revisions** (e.g., "the existing malformed-json + probe's grader is too strict — rebuild it with a looser match"). + Treat as new operator directives, mark the named probes for + rebuild, and proceed to (d) with the updated dispatch set. + Apply the destructive-edit checkpoint from (b). 3. **User rejects the plan structure.** Likely they want a - different overall workbench shape — surface this and ask whether - to abandon (loop back to step 2 for re-planning) or to retry - the plan with their feedback as directives. + different probe set overall — surface this and ask whether to + abandon (loop back to step 2 to revise the functionality specs) + or to retry the plan with their feedback as directives. -### (d) Dispatch test-writer subagents (parallel, one per case) +### (d) Dispatch test-writer subagents (parallel, one per probe) **Do NOT write workspace files or graders yourself in this -session.** For each case in the dispatch set (new picks + revisions -flagged in (c)), dispatch a test-writer subagent via the `Agent` -tool, in parallel — emit all dispatches in a single message so they -run concurrently. +session.** For each probe in the dispatch set (new probes + probes +flagged for rebuild in (c)), dispatch a test-writer subagent via +the `Agent` tool, in parallel — emit all dispatches in a single +message so they run concurrently. Load the prompt template at [`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) -and substitute per-case inputs: - -- `${CASE_NAME}` — this case's name from `02-test-case.md` -- `${CASE_SPEC_PATH}` — path to a slice of `02-test-case.md` - containing only this case's entry (the operator session extracts - this from the full report so the subagent doesn't see other - cases) -- `${FUNCTIONALITY_PATH}` — the latest `01-functionality.md` -- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local skill - path), so the subagent can ground the fixture in concrete +and substitute per-probe inputs: + +- `${PROBE_NAME}` — this probe's slug (the folder name under + `tests//`) +- `${FUNCTIONALITY_SPEC_PATH}` — path to the parent + functionality's `spec.yaml` (gives the subagent context on what + responsibility this probe is probing) +- `${PROBE_SPEC_PATH}` — path where the probe's `spec.yaml` will + be written (the test-writer also writes this file describing + what the probe sets up + expects) +- `${FUNCTIONALITY_PATH}` — the current `01-functionality.md` +- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local + skill path), so the subagent can ground the fixture in concrete violation patterns -- `${OUTPUT_WORKBENCH_DIR}` — `workbench/cases//` -- `${VERSION}` from (b) -- `${OPERATOR_DIRECTIVES}` from (b) (case-level revision hints if - this case is being revised; empty otherwise) +- `${OUTPUT_PROBE_DIR}` — `tests///` +- `${OPERATOR_DIRECTIVES}` — case-level revision hints if this + probe is being rebuilt; empty otherwise Each test-writer subagent sees: -- Its single case spec (the slice extracted in (d)) +- Its single probe's intent (the slug + the parent functionality + spec) - `01-functionality.md` - The skill source content (vendored or local) -- `${OPERATOR_DIRECTIVES}` for its case +- `${OPERATOR_DIRECTIVES}` for its probe The test-writer subagent does NOT see: -- Other test cases' specs (prevents copying fixtures across cases) -- Other test cases' graders (prevents grader-leak hacking — looking - at how another grader matches and writing a fixture that exploits - the same pattern) +- Other probes' specs, graders, or workspaces (prevents copying + across probes and grader-leak hacking) - The eval grader's matching internals (same reason) -- Anything under `archive/` -- `06-analysis.md`, `07-improvement-proposal.md`, raw failure data, - prior optimizer attempts +- `06-analysis.md`, `07-improvement-proposal.md`, raw failure + data, prior optimizer attempts +- Git history of its own probe folder (anti-rationalization) -The "no other cases" constraint is load-bearing: each case is built -in isolation so cross-case patterns don't subtly homogenize the -fixtures. The "no grader internals" constraint prevents fixtures -from being gerrymandered to the grader's specific matching rules. +The "no other probes" constraint is load-bearing: each probe is +built in isolation so cross-probe patterns don't subtly homogenize +the fixtures. The "no grader internals" constraint prevents +fixtures from being gerrymandered to the grader's specific matching +rules. Each test-writer subagent writes: -- Workspace files at `workbench/cases//workspace/` -- Grader script at `workbench/cases//grader.mjs` (or - `.py` per the workbench schema) -- Smoke-check fixtures at `workbench/cases//smoke/` - (GOOD/BAD/EMPTY findings.txt examples the grader should classify - correctly) -- A short per-case section for the operator to fold into - `04-tests-plan.md` - -The subagent returns a brief summary: case name, what the fixture +- `tests///spec.yaml` (the probe's + intent + workspace overview + expected behavior + grader logic) +- `tests///workspace/` (the + fixture) +- `tests///grader.mjs` (the grader) +- `tests///smoke/{good,bad,empty}/findings.txt` + (smoke fixtures the grader should classify correctly) +- `tests///checks/smoke.mjs` (a runner + that exercises the grader against the three smoke fixtures) + +The subagent returns a brief summary: probe name, what the fixture tests, grader logic in one line, smoke-check result for its own -case. +probe. ### (e) Run smoke check -The smoke check is per-workbench, not centralized. Each test-writer -subagent produces, alongside its grader, a smoke-check artifact -that validates the grader against hand-crafted GOOD/BAD/EMPTY -fixtures. The shape of that artifact follows the workbench schema -documented in -[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md). -Existing examples to follow: -[`examples/workbench/agent-browser/checks/smoke-graders.mjs`](../../examples/workbench/agent-browser/checks/smoke-graders.mjs) -shows a workbench-level smoke checker that runs all of that -workbench's graders against GOOD/BAD fixtures. - -Two reasonable shapes (the test-writer prompt template should pick -one and apply it consistently): - -- **Per-case runner**: each test-writer produces - `workbench/cases//checks/smoke.mjs` that exercises - that case's grader against its own GOOD/BAD/EMPTY fixtures. - Operator runs each runner. -- **Workbench-level aggregator**: the test-writers each contribute - their fixtures + grader to a shared location, and the operator - composes (or the workbench schema provides) a top-level - `workbench/checks/smoke-graders.mjs` that runs all of them at - once. - -After all test-writer dispatches return, run the smoke-check -artifact(s) the test-writers produced. Expected for each case: +The smoke check is per-probe (each probe has its own +`checks/smoke.mjs`). After all test-writer dispatches return, run +each new or rebuilt probe's smoke runner: + +```bash +node tests///checks/smoke.mjs +``` + +Expected for each probe: - The GOOD fixture passes the grader. - The BAD fixture fails the grader. - The EMPTY fixture fails the grader. -If any grader fails the smoke check, surface the case to the user. -Two realistic responses: - -1. **User asks to re-dispatch the test-writer for that case** (with - the smoke-check failure as a directive). Treat as a single-case - revision, return to (d) with just that case in the dispatch set. -2. **User decides to remove the case from `picked`** (in - `02-test-case.md`) and re-run step 4. That's a backward trigger - — per the iteration protocol's re-run authorization rule, you - surface the option and let the user invoke step 2 explicitly. - -Update `last_smoke_check_passed` in the frontmatter only after all -graders pass. +If any grader fails the smoke check, surface the probe to the +user. Two realistic responses: + +1. **User asks to re-dispatch the test-writer for that probe** + (with the smoke-check failure as a directive). Treat as a + single-probe rebuild, return to (b) for the destructive-edit + checkpoint, then re-dispatch via (d). +2. **User decides to remove the probe** from this functionality's + set, or de-pick the parent functionality. The change is to + `tests//spec.yaml` (set `picked: false`) or to + delete the probe folder. That's a backward trigger — per the + iteration protocol's re-run authorization rule, surface the + option and let the user invoke step 2 explicitly if they want + to revise the functionality. + +### (f) Generate suite.yml + commit + hand off + +Walk the current `tests/` tree. For each `tests//` +where `spec.yaml` has `picked: true`, list every probe folder +(`//`) and emit a suite.yml entry per +probe per the workbench schema (see +[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md) +for the exact shape). -### (f) Assemble plan + commit + hand off +Write the result to `tests/suite.yml`, overwriting any prior +version (git tracks the change). -Assemble the per-case sections returned by the test-writer -subagents into `04-tests-plan.md` (with the frontmatter from (b)). -Update `workbench/suite.yml` to include all entries in `built_cases` -(excluding any that were removed in the (b) diff). +Commit the current state of `tests/`: -Commit both `04-tests-plan.md` and the `workbench/` directory. +```bash +git add docs/skill-optimizer//tests/ +git commit -m "step 4: build probes for " +``` Handoff message, verbatim: > Next, invoke `skill-optimizer-run-bench` to measure baseline -> performance against the workbench. +> performance against the probes. ## Why limited-context dispatch matters -The test-writer subagents are dispatched per-case for two reasons: +The test-writer subagents are dispatched per-probe for two +reasons: First, **fixture isolation**. A single subagent writing all -fixtures would notice patterns across cases ("they all check for +fixtures would notice patterns across probes ("they all check for absence violations — let me write a unified fixture") and -inadvertently homogenize them. Per-case isolation forces each +inadvertently homogenize them. Per-probe isolation forces each fixture to be a representative instance of its responsibility, -designed without knowledge of how sibling cases are shaped. +designed without knowledge of how sibling probes are shaped. Second, **no grader-leak hacking**. If a subagent sees how another grader matches (regex pattern, JSON path, etc.), it can write a fixture that incidentally satisfies the OTHER grader too — making -the eval results look correlated when they're not. Per-case +the eval results look correlated when they're not. Per-probe isolation prevents this. The source access concession (test-writer subagents DO see the -skill source) is bounded by the case spec from step 2: the spec -fixes WHAT the fixture should test, and the source provides the -HOW (concrete violation patterns). Without source, the subagent -would invent generic patterns that may not actually trigger the -skill's rules. +skill source) is bounded by the probe's spec: the spec fixes WHAT +the fixture should test, and the source provides the HOW (concrete +violation patterns). Without source, the subagent would invent +generic patterns that may not actually trigger the skill's rules. -If you find yourself thinking "the cases are similar; I'll write +If you find yourself thinking "the probes are similar; I'll write them all myself with one prompt and save dispatches" — that's the homogenization failure mode this chain is built to prevent. Dispatch in parallel. ## Edge cases -- **A test-writer subagent reports BLOCKED** (e.g., the case spec +- **A test-writer subagent reports BLOCKED** (e.g., the probe spec is too abstract to derive a fixture from) — surface to the user with the subagent's reasoning. The fix is typically a step 2 - re-run with a more specific case spec; per the iteration - protocol's re-run authorization rule, the user invokes step 2 - explicitly. + re-run with a more specific functionality spec or + suggested_probes list; per the iteration protocol's re-run + authorization rule, the user invokes step 2 explicitly. - **A grader fails the smoke check** — see step (e)'s handling. -- **Smoke check passes for individual cases but the suite.yml - itself is malformed** — surface as a workbench-schema error; - fix the suite.yml directly (operator session task; not a - test-writer concern). -- **A picked case is genuinely unimplementable in a static - workbench** (e.g., requires real-time API access the workbench - can't provide) — the test-writer should return BLOCKED with this - reasoning. Surface to the user; the case may need to be removed - from `picked` in step 2. +- **Probe smoke checks pass individually but `tests/suite.yml` is + malformed** — surface as a workbench-schema error; fix the + suite.yml directly (operator session task; not a test-writer + concern). +- **A picked functionality is genuinely unimplementable in a + static workbench** (e.g., requires real-time API access the + workbench can't provide) — the test-writer should return BLOCKED + with this reasoning. Surface to the user; the functionality may + need to be de-picked at step 2. +- **User de-picks a functionality between step 4 runs** — its + probe folders stay on disk (filesystem-as-state principle), but + step (f) excludes them from the regenerated `tests/suite.yml`. + If the user later re-picks the functionality, the probes are + already there. ## Iteration behavior -Step 4 is a maintenance step — the workbench accumulates as new -picks get built. Re-run triggers specific to this step: - -- `02-test-case.md`'s `picked` field changed (new picks added, - some removed) -- A specific case's smoke check failed and needs the test-writer - to revise (single-case re-dispatch via directives) -- Step 5 (`run-bench`) showed all cases pass on baseline (too - easy) or all fail (too hard) — directives like "make case X - harder" or "loosen case Y's grader" trigger case-level revisions -- `02-test-case.md`'s version bumped substantively (responsibilities - reframed) — operator surfaces and the user decides whether to - preserve existing builds or start fresh - -General iteration mechanics — archive, version bump, operator -directives, cascading staleness, maintenance vs fresh-derivation — -live in +Step 4 is a maintenance step — the `tests/` tree accumulates +probes as new functionalities get picked, and re-runs preserve +existing built probes unless directives target them. Re-run +triggers specific to this step: + +- A new functionality at step 2 got `picked: true` +- A specific probe's smoke check failed and needs the test-writer + to revise (single-probe rebuild via directive) +- Step 5 (`run-bench`) showed all probes pass on baseline (too + easy) or all fail (too hard) — directives like "make probe X + harder" or "loosen probe Y's grader" trigger probe-level + rebuilds +- Step 2 substantively revised a picked functionality's spec.yaml + (the suggested_probes list changed in a way that warrants + rebuilding) — operator surfaces and the user decides whether to + rebuild the existing probes or leave them + +General iteration mechanics — staleness detection, operator +directives, cascading staleness, safe destructive edits — live in [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). Workflow step (b) above already requires reading that file. From a31d720310b94e2717fe6e6a58c82279b68dc524 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 06:57:41 -0500 Subject: [PATCH 026/121] docs(iteration-protocol): clarify directive channel + step 7 carve-out MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two wording fixes to the fresh-derivation subagent constraint: 1. The prior wording "no reference to what was produced before" read as forbidding even the directive mechanism — which is wrong. The operator session DOES read prior outputs (that's part of its job between iterations) and distills lessons into atomic new requirements. The subagent then satisfies those distilled requirements as fresh constraints without seeing the raw prior content. This separation is what keeps the new derivation from rationalizing the prior one while still letting the chain converge across iterations. 2. Step 7 has a specific carve-out worth stating explicitly: the SKILL itself (the improvement target) is upstream input, not the subagent's own canonical. The optimizer reads the current skill — which may include modifications from prior step-7 runs — and proposes new improvements on top. The "own canonical" that's off-limits is 07-improvement-proposal.md (the reasoning report), not the skill file. The skill accumulates improvements across iterations; the proposal reports do not. Triggered by review of the prior wording on iteration-protocol.md. --- .../iteration-protocol.md | 37 +++++++++++++++---- 1 file changed, 30 insertions(+), 7 deletions(-) diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 30aa712..02f7d89 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -78,13 +78,36 @@ Three rules every reasoning subagent must follow: 2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7): do NOT read your own canonical file, and do NOT read git history of - it.** Your job is to derive a new artifact from upstream + - directives, with no reference to what was produced before. This - is the load-bearing anti-ducktape constraint. The operator - session writes your output by overwriting the canonical file; - git captures the prior state automatically, but neither you nor - any subsequent fresh-derivation subagent should walk that - history. + it.** Your job is to derive a new artifact from upstream + the + operator's directives, without direct access to prior + derivations of this same artifact. This is the load-bearing + anti-ducktape constraint. + + The directives ARE the channel for lessons learned from prior + iterations. The operator session CAN read the prior output — + that's part of its job between iterations — and distills any + lessons into atomic new requirements ("the prior optimization + was too destructive on the description field; this attempt + should be additive only"). The subagent then satisfies those + distilled requirements as fresh constraints, without seeing the + raw prior content. This separation is what keeps the new + derivation from rationalizing the prior one while still letting + the chain converge: each iteration starts cleaner than the + last, with the operator's distilled lessons as guardrails. + + The operator session writes your output by overwriting the + canonical file; git captures the prior state automatically, but + neither you nor any subsequent fresh-derivation subagent should + walk that history. + + **For step 7 specifically:** the SKILL itself (the target being + improved) is upstream input, not your own canonical. You read + the current skill — which may include modifications from prior + step-7 runs — and propose a new improvement on top. Your "own + canonical" that's off-limits is `07-improvement-proposal.md` + (the report describing your reasoning), not the skill file + itself. The skill accumulates improvements across iterations; + the proposal reports do not. 3. **For maintenance steps (2, 4): DO read your own canonical tree** (when it exists). Your job on a re-run is to extend or From 56867f855773b0eed53059d91530d4d4523a0ddf Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 06:58:56 -0500 Subject: [PATCH 027/121] docs(iteration-protocol): prescriptive frontmatter discipline rule Replace the descriptive "there is no version: field" wording with a prescriptive "do not add version-tracking metadata" rule plus a practical decision aid for future SKILL.md authors: When in doubt about whether a field belongs: ask whether a chain skill needs to READ it to do its job right now (yes -> keep), or whether you're recording it for future-debugging / future-audit purposes (no -> that's git's job). The prior wording could be read as describing the current state without prohibiting reintroduction. The new wording makes the prohibition explicit so B6/B7/B8 (yet to be drafted) don't accidentally bring back version-tracking metadata. --- .../iteration-protocol.md | 22 ++++++++++++++----- 1 file changed, 17 insertions(+), 5 deletions(-) diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 02f7d89..c14b803 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -26,11 +26,23 @@ This protocol determines, on every step invocation: 2. What the subagent is and isn't allowed to look at 3. How the operator handles destructive edits safely -History is **git**. There is no `version:` field, no `archive/` -subdirectory, no manually-bumped counters, and no `inputs.step_N:` -ints recording which upstream version was consumed. Git already -content-addresses every prior state; reimplementing that in -frontmatter is bookkeeping for its own sake. +History is **git**. Do not add version-tracking metadata to any +chain artifact — no `version:` field, no `archive/` subdirectory, +no manually-bumped counters, no `inputs.step_N:` lineage trackers, +no `last_derived_at:` timestamps. The single rule: frontmatter +carries **runtime-relevant facts** (what this artifact represents, +what state it's in — e.g., `picked: true`, `pr_submission_intent: +true`, `classification: tool-use`), not **history** (when it was +written, how many times, what upstream version produced it). Git +already content-addresses every prior state and `git log` gives you +the timeline; reimplementing either in frontmatter is bookkeeping +for its own sake — it adds maintenance surface (bumping, syncing, +risk of drift) without enabling anything git can't already do. + +When in doubt about whether a field belongs: ask whether a chain +skill needs to READ it to do its job right now (yes → keep), or +whether you're recording it for future-debugging / future-audit +purposes (no → that's git's job). ## Step kinds: fresh-derivation vs maintenance From 0dbb9982862baf0edc9382f7e141287818fde8b8 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 07:06:33 -0500 Subject: [PATCH 028/121] feat(skill-optimizer-analyze-result + improve-skill): SKILL.md bodies MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit B6 (analyze-result): the chain's anti-ducktape gate. Fresh-derivation step. Dispatches an analyzer subagent that reads bench summary + raw trial data + probe specs + skill content — but NOT the test inputs themselves (forces principle-thinking over solution-thinking). Output: 06-analysis.md with has_structural_weakness: true|false in frontmatter and per-weakness sections containing Pattern, Hypothesized cause, Connects to skill section, What WOULD address this (general principle), and What WOULD NOT address this (the explicit anti-ducktape list step 7's optimizer must reckon with). Honest refusal is built in: if no weakness can be articulated, the report says so and has_structural_weakness: false, which step 7 will refuse to fire on. Forced weakness-naming when the analyzer found nothing is the ducktape failure mode this step exists to prevent. B7 (improve-skill): the terminal generative step. Two subagents (optimizer + validator), both fresh-derivation. Refuses to fire if 06-analysis has has_structural_weakness: false (the anti-ducktape gate's downstream half). Optimizer: reads 06-analysis, 01-functionality, current skill, 03-submissions (if PR-bound). Does NOT see raw trials, grader internals, test inputs, or prior proposals. Must apply the analyzer's general principle and self-check against the anti-pattern list explicitly. Validator: reads skill BEFORE + AFTER + 01-functionality + the proposal artifact + 03-submissions (if PR-bound). Internal consistency + external consistency checks. Does NOT see the optimizer's reasoning trace, prior verdicts, or raw trial data. Bounded loop: max 2 revision rounds. If validator still says needs-revision after round 2, surface honestly with three realistic paths; do not loop indefinitely. Three handoff branches: local skill (modify in place), upstream + PR=false (modify vendored copy), upstream + PR=true (write 07-pr-draft.md with operator-steps-to-submit; the chain does NOT submit the PR). Outputs three or four artifacts: 07-improvement-proposal.md, 07-validator-verdict.md, the modified skill file (on approve), and 07-pr-draft.md (PR-bound + approve). Frontmatter on both reports carries runtime-relevant facts only (verdict, addresses_weaknesses, diff_target) per the iteration protocol's discipline rule. Both files follow the established structural template (front-loaded "Before you start", lettered workflow steps, edge cases, iteration behavior section). Pending user review of B1-B7 before B8 and the subagent prompt templates. --- .../skill-optimizer-analyze-result/SKILL.md | 327 +++++++++++- skills/skill-optimizer-improve-skill/SKILL.md | 473 +++++++++++++++++- 2 files changed, 794 insertions(+), 6 deletions(-) diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md index 0cfebb2..c786974 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -1,9 +1,330 @@ --- name: skill-optimizer-analyze-result -description: Use when the user wants to analyze bench results and identify structural weaknesses — clusters failures, distinguishes systematic from flaky, and produces a structured analysis with anti-ducttape gates. +description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer-run-bench` has produced a `05-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. --- # skill-optimizer-analyze-result - - +Step 6 of the skill-optimizer chain. Takes the bench summary + raw +trial output from step 5, dispatches an analyzer subagent to +cluster failures into named **structural weaknesses** of the skill +(or explicitly say there are none), and writes +`docs/skill-optimizer//06-analysis.md`. This is the chain's +**anti-ducktape gate**: step 7 (`improve-skill`) refuses to fire +unless this report names at least one structural weakness, with the +general principle that WOULD address it and the anti-patterns that +would NOT. + +## Before you start + +Two load-bearing pieces of context to load NOW, before the workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation — for staleness + detection, operator-directive handling, and the constraints + subagents must obey. Step 6 is a **fresh-derivation** step: + each invocation derives a new analysis from the bench data + + directives, with no continuity from prior analyses. The + anti-ducktape constraint is load-bearing here. + +2. **You will dispatch a subagent for the actual analysis; you do + NOT cluster failures or name weaknesses yourself in this + session.** Workflow step (c) is the dispatch. The rationale is + in "Why limited-context dispatch matters" below — read it if + you're tempted to skip the dispatch. + +Throughout this document, "step 1" through "step 7" (no parens) +refer to skills in the chain (this skill is step 6; `run-bench` is +step 5; `improve-skill` is step 7). Internal workflow steps within +THIS skill are labelled "(a)" through "(e)". + +## What you produce + +A single report at `docs/skill-optimizer//06-analysis.md`, +where `` matches the slug from step 1's report. + +Frontmatter (runtime-relevant facts only — no version tracking, +per the iteration protocol's frontmatter discipline rule): + +```yaml +--- +has_structural_weakness: true | false +weakness_count: +bench_results_path: 05-bench-results// +--- +``` + +**`has_structural_weakness`** — load-bearing for step 7. If +`false`, step 7 refuses to fire ("no weakness to address"). The +analyzer subagent sets this honestly based on whether it can +articulate at least one structural weakness; a forced +`has_structural_weakness: true` when the analyzer found nothing is +the ducktape failure mode this step exists to prevent. + +**`weakness_count`** — informational; should match the number of +`### Weakness :` sections in the body. + +**`bench_results_path`** — relative path to the timestamped raw +output the analyzer read from. Lets step 7's optimizer (if it +fires) trace back to the underlying data without re-deriving. + +Body structure: + +```markdown +## Structural weaknesses identified + +### Weakness 1: + +- **Pattern**: Across trials, systematically failed to + detect . Specifically: //>. +- **Hypothesized cause**: . +- **Connects to skill section**: at . +- **What WOULD address this**: . +- **What WOULD NOT address this** (anti-ducktape gate): . + +### Weakness 2: ... + +## Non-structural noise (ignored — not addressable) + +- 1 gemini transient API error on multi-redirect (not a pattern) +- 1 gpt-5 timeout (infrastructure, not skill) + +## Honest refusal (if applicable) + +If no structural weakness can be articulated: + +> No structural weakness identified. Failures observed are +> consistent with model nondeterminism / infrastructure noise / +> single-trial flakiness, not a fixable defect in the skill itself. +> Step 7 should not fire. + +The frontmatter `has_structural_weakness: false` mirrors this. +``` + +The "What WOULD NOT address this" anti-pattern list is +**load-bearing**. Without it, step 7's optimizer can pattern-match +a patch that fits the symptom without addressing the cause; the +validator then has no explicit "this would be a ducktape" signal +to check against. The analyzer's job is to name BOTH the principle +that should be applied AND the ducktape moves the optimizer must +avoid. + +For the body template details and the analyzer's reasoning +protocol, see +[`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md). +The subagent writes the report body itself. + +## Workflow + +### (a) Confirm prerequisites + +`docs/skill-optimizer//05-bench-summary.md` must exist with a +valid `bench_results_path` and `overall_pass_rate`. If it doesn't, +tell the user to run `skill-optimizer-run-bench` first and stop +here. + +If `overall_pass_rate == 1.0`: there's nothing to analyze. Surface +this honestly: + +> All trials passed on the current bench. Step 6 has nothing to +> analyze. Two realistic paths: (1) accept that the current probes +> don't expose a weakness in the skill; (2) re-run step 2 with a +> "make probes harder" directive and walk the chain forward again. + +Don't run the analyzer subagent in this case — there are no +failures to cluster. + +Also confirm `docs/skill-optimizer//tests/` exists (so the +analyzer can read probe specs) and the source skill is available +(either vendored at `vendored-skill/` or the local path recorded in +`01-functionality.md`'s `skill_source`). + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it. Step 6 is a **fresh-derivation** step: + +- Each invocation derives a new analysis from the current bench + data + directives. The subagent does NOT read its own prior + `06-analysis.md` or its git history. Anti-ducktape: the + analyzer must not be biased by what it (or a prior analyzer + run) said before. +- If `06-analysis.md` already exists, it will be overwritten by + this invocation. Prior state lives in git history. No checkpoint + commit needed (fresh-derivation; prior state is already + self-contained in its own prior commit). +- Collect `${OPERATOR_DIRECTIVES}` per the protocol. Atomic new + requirements only — examples: "focus on the gpt-5 cluster, the + other models passed", "the user thinks weakness X from a prior + run is actually two separate issues, look for both". The + operator distills these from prior outputs they CAN read; the + subagent works from the distilled directives, never from the raw + prior analysis. + +### (c) Dispatch the analyzer subagent + +**Do NOT cluster failures or name weaknesses yourself in this +session.** Dispatch the analyzer subagent via the `Agent` tool +(with worktree isolation if your environment supports it). Load +the prompt template at +[`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md) +and substitute the templated inputs (`${BENCH_RESULTS_PATH}` from +`05-bench-summary.md`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, +`${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, +`${OPERATOR_DIRECTIVES}`). + +The subagent sees: + +- `${SUMMARY_PATH}` — `05-bench-summary.md` (entry point with + failed-probe pointer list) +- `${BENCH_RESULTS_PATH}` — `05-bench-results//` + containing per-trial `trace.jsonl` and per-trial `findings.txt` +- `${TESTS_TREE_PATH}` — the `tests/` tree, but the subagent reads + ONLY each probe's `spec.yaml` (what the probe was probing at the + level of intent) — NOT the `workspace/` contents (the raw input + fixtures) +- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local + skill path) so the analyzer can quote the responsible skill + section in its weakness entries +- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new + requirements (may be empty) + +The subagent does NOT see: + +- **The test inputs themselves** + (`tests///workspace/` files). This is + load-bearing: forces the analyzer to think about the SKILL, not + the SOLUTIONS. Reading the raw input would let it reason "the + agent should have detected X specifically in this file" — that's + solution-thinking, not principle-thinking. +- Its own prior `06-analysis.md` (anti-ducktape — must not be + biased by prior analyses) +- Git history of `06-analysis.md` +- `07-improvement-proposal.md`, `07-validator-verdict.md`, or any + prior optimizer attempts (would also bias toward + "weaknesses-that-could-be-patched-the-way-the-optimizer-tried") + +The subagent writes the report itself and returns a brief summary: +weakness count, top weakness by trial-coverage, and the +`has_structural_weakness` value it set. + +### (d) Confirm subagent output + +Verify the report file exists, the frontmatter parses, and +`weakness_count` matches the number of `### Weakness :` sections +in the body. If `has_structural_weakness: true` but no weakness +sections exist (or vice versa), the subagent's output is +internally inconsistent — surface to the user and ask whether to +re-dispatch with a directive ("your frontmatter says X but body +says Y; reconcile"). + +Check that each weakness section has all five required parts +(Pattern, Hypothesized cause, Connects to skill section, What +WOULD address this, What WOULD NOT address this). If any section +is missing the anti-pattern list, surface it: this is the +anti-ducktape signal step 7 needs, and a weakness without it can't +be passed to the optimizer safely. + +### (e) Hand off + +Two handoff messages depending on `has_structural_weakness`: + +If `has_structural_weakness: true`: + +> Analysis complete. `` structural weakness(es) identified. Next, +> invoke `skill-optimizer-improve-skill` to draft a principled fix. + +If `has_structural_weakness: false`: + +> Analysis complete. No structural weakness identified — failures +> observed are consistent with noise rather than a fixable defect. +> Step 7 will refuse to fire. Two realistic paths: (1) accept the +> conclusion and exit honestly; (2) if you disagree, re-invoke +> step 6 with a directive ("the analyzer missed the X cluster — +> look at it again") to get a fresh derivation. + +Don't auto-invoke step 7 — the chain skills never invoke each +other on their own (per the iteration protocol's re-run +authorization rule). + +## Why limited-context dispatch matters + +The analyzer subagent's anti-ducktape job is the chain's most +fragile in terms of bias. Two specific risks the limited context +addresses: + +First, **solution-thinking vs. principle-thinking**. If the +analyzer sees the raw input fixtures in `tests//workspace/`, +it will naturally reason "the agent should have detected X +specifically in this file" — and recommend a patch that says +"detect X in files like this". That's a ducktape patch around a +specific input shape. By blocking the raw inputs, the analyzer is +forced to read only the probe's stated intent + the agent's actual +behavior, and reason about why the skill failed to instruct the +agent properly. The output is a principle (what the skill should +teach the agent to do in general), not a solution (what the agent +should have done in this specific case). + +Second, **prior-analysis bias**. The operator session has read +prior `06-analysis.md`s, prior optimizer attempts, prior validator +verdicts. That context would lead an in-session analyzer to +gravitate toward "what we said last time" or "what the optimizer +tried last time" — and either confirm those framings or +defensively pivot away from them. Neither is the right job. The +analyzer must derive fresh from the current data + the operator's +distilled directives, with no view of prior runs' meta-discussion. + +If you find yourself thinking "I'll just write the analysis +myself, I already see the patterns" — that's the failure mode this +chain is built to prevent. Dispatch. + +## Edge cases + +- **All trials passed** — caught at (a). Don't run the analyzer. +- **Bench results dir missing or `suite-result.json` malformed** — + the analyzer can't proceed; surface as a step-5 problem (the + bench run was incomplete or corrupted) and tell the user to + re-run step 5. +- **Subagent returns `has_structural_weakness: false` and user + disagrees** — surface the disagreement, but do NOT pressure the + subagent to manufacture a weakness. The honest path is to + re-invoke step 6 with a directive pointing at what the user + thinks was missed ("the gpt-5 trial 3 failure pattern wasn't + covered — look at it specifically"). Per the iteration protocol's + re-run authorization rule, the user invokes the re-run; this + skill doesn't auto-invoke. +- **Subagent returns weaknesses but the "What WOULD NOT address + this" lists are empty or vague** — the anti-ducktape gate is + compromised. Re-dispatch with a directive: "each weakness needs + a concrete anti-pattern list — what specific moves should the + optimizer avoid?" Don't fill it in yourself. +- **Operator directives contradict each other** (e.g., "focus on + cluster A" + "ignore cluster A") — surface to the user before + re-dispatching; don't try to resolve it yourself. + +## Iteration behavior + +Step 6 is re-runnable and is a fresh-derivation step. Re-run +triggers specific to this step: + +- Step 5 (`run-bench`) produced new results — `05-bench-summary.md` + changed via git mtime +- User disagrees with the analyzer's verdict (either thinks a + weakness was missed, or thinks `has_structural_weakness: true` + should have been `false`) and wants a fresh derivation with + directives reflecting the disagreement +- Step 7 (`improve-skill`) was unable to address one of the named + weaknesses and the user wants the analyzer to reformulate it + ("weakness 2 as named was too abstract; split into concrete + sub-weaknesses") + +General iteration mechanics — staleness detection (git-native), +operator directives (the channel for prior-derivation lessons), +cascading staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md index bb2c382..76bea53 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -1,9 +1,476 @@ --- name: skill-optimizer-improve-skill -description: Use when the user wants to improve a skill based on identified structural weakness — dispatches optimizer + validator subagents under strict limited context; refuses if no structural weakness in the analysis. +description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze-result` has produced `06-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. --- # skill-optimizer-improve-skill - - +Step 7 of the skill-optimizer chain — the terminal generative step. +Takes the named structural weaknesses from step 6, dispatches an +optimizer subagent to draft a principled fix, then dispatches a +validator subagent to independently check whether the fix is sound +(internal consistency) and conformant (external PR conventions if +PR-bound). Writes the modified skill file, the improvement +proposal, and the validator's verdict. If the user opted into PR +submission at step 1, packages the change as a PR draft. + +This skill **refuses to fire** if `06-analysis.md` has +`has_structural_weakness: false` — there's nothing to optimize, and +forcing a fix in that situation is the ducktape failure mode the +chain is built to prevent. + +## Before you start + +Three load-bearing pieces of context to load NOW, before the +workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation. Step 7 is a + **fresh-derivation** step (both the optimizer and validator are + fresh-derivation subagents); the iteration protocol's + "frontmatter discipline" + "directive channel" sections both + apply. Pay attention to the **step-7 carve-out** in the + subagent-constraints section: the SKILL itself is upstream + input (the optimizer reads the current skill, which may include + modifications from prior step-7 runs), but the optimizer's own + canonical (`07-improvement-proposal.md`) and the validator's + own canonical (`07-validator-verdict.md`) are both off-limits + to their respective subagents. + +2. **You will dispatch two subagents (optimizer + validator); you + do NOT propose diffs or judge them yourself in this session.** + Workflow steps (c) and (d) are the dispatches. The rationale is + in "Why limited-context dispatch matters" below. + +3. **The optimizer/validator loop is bounded to 2 revision + rounds.** If the validator still says `needs-revision` after + round 2, surface honestly and exit — do not loop indefinitely. + This bound protects against optimizer/validator pathological + disagreement burning operator budget. + +Throughout this document, "step 1" through "step 7" (no parens) +refer to skills in the chain (this skill is step 7; +`analyze-result` is step 6). Internal workflow steps within THIS +skill are labelled "(a)" through "(g)". + +## What you produce + +Three or four artifacts at `docs/skill-optimizer//`, +depending on the verdict: + +1. **`07-improvement-proposal.md`** — the optimizer's proposed + diff + rationale. Frontmatter (runtime-relevant facts only, + per the iteration protocol's frontmatter discipline): + + ```yaml + --- + addresses_weaknesses: + - + - ... + diff_target: + --- + ``` + + Body sections: the proposed diff (in fenced unified-diff + format), the rationale explicitly referencing which named + weakness from `06-analysis.md` it addresses and which "What + WOULD address this" principle it applies, and a self-check + against the "What WOULD NOT address this" anti-pattern list + (the optimizer states explicitly why its proposal is NOT one + of the ducktape moves the analyzer flagged). + +2. **`07-validator-verdict.md`** — the validator's independent + judgment. Frontmatter: + + ```yaml + --- + verdict: approve | needs-revision | reject + round: + addresses_weaknesses: + - + - ... + --- + ``` + + Body: internal consistency check (does the change make sense + for the named weakness? is it additive vs. destructive? is it + general vs. ducktape?), and — if `03-submissions.md` exists — + external consistency check (does the change conform to upstream + PR rules: frontmatter, file location, prefix taxonomy, + additive-only, etc.). + +3. **The modified skill file** (only if `verdict: approve`). + Written in place for local skills; written to + `vendored-skill/` for upstream skills. + +4. **`07-pr-draft.md`** (only if `pr_submission_intent: true` + AND `verdict: approve`). A PR draft containing the diff, the + PR body referencing the weakness and the principle applied, + caveats the operator should know about (CLA requirement, branch + target, anything flagged in `03-submissions.md`), and + operator-steps-to-submit. The chain does NOT submit the PR — + that's an explicit operator action. + +For the body templates of the optimizer and validator reports + +their reasoning protocols, see +[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) +and +[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). +Each subagent writes its own body; the operator session orchestrates +the loop and writes the PR draft (which is operator-judgment work, +not generative). + +## Workflow + +### (a) Confirm prerequisites + refuse-if-no-weakness gate + +Three checks, in order: + +1. `docs/skill-optimizer//06-analysis.md` must exist with + valid frontmatter. If not, tell the user to run + `skill-optimizer-analyze-result` first and stop here. + +2. **Anti-ducktape gate:** `06-analysis.md`'s frontmatter must + have `has_structural_weakness: true`. If `false`, REFUSE to + proceed — print a message like: + + > Step 6 found no structural weakness in the current skill + > against the current probes. Step 7 won't fire — there's + > nothing principled to optimize. If you disagree, re-invoke + > step 6 with a directive describing what you think was + > missed; if you agree, exit honestly. + + Do NOT proceed to dispatch anyway. The refusal is the entire + reason this gate exists. + +3. `01-functionality.md` must exist (the optimizer reads it as + input). The source skill must be present (vendored at + `vendored-skill/` for upstream skills, or accessible at the + local path recorded in `01-functionality.md`'s `skill_source` + field). + +If `pr_submission_intent: true` from `01-functionality.md`, +`03-submissions.md` should also exist; if it's missing, the +validator's external consistency check can't run. Surface this and +ask the user whether to run step 3 first OR proceed with internal +consistency only. + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it. Step 7 is a **fresh-derivation** step for both +subagents: + +- Each invocation derives a fresh proposal + verdict from + `06-analysis.md` + the current skill content + directives. + Neither subagent reads its own prior canonical + (`07-improvement-proposal.md` for the optimizer; + `07-validator-verdict.md` for the validator) or its git + history. +- The SKILL itself IS upstream input — the optimizer reads the + current skill content, which may include modifications from + prior step-7 runs. That's intentional; the skill accumulates + improvements across iterations, but the proposal reports + describing each round do not. +- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The operator + CAN read the prior `07-improvement-proposal.md` and + `07-validator-verdict.md`, distill any lessons into atomic new + requirements, and pass them to the optimizer. Examples: + - "prefer additive changes over destructive ones" + - "don't touch the description field — the validator rejected + that last round" + - "address weakness 2 first; weakness 1 was already partially + addressed by the prior round" + + The subagent works from the distilled directives, never from + the raw prior content. + +### (c) Dispatch the optimizer subagent + +**Do NOT propose the diff yourself in this session.** Dispatch +the optimizer subagent via the `Agent` tool (with worktree +isolation if your environment supports it). Load the prompt +template at +[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) +and substitute the templated inputs (`${ANALYSIS_PATH}`, +`${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, +`${SUBMISSIONS_PATH}` if PR-bound, +`${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`). + +The optimizer sees: + +- `${ANALYSIS_PATH}` — `06-analysis.md` (the named weaknesses + + what WOULD/WOULDN'T address each) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (so the + optimizer understands the skill's stated responsibilities and + doesn't propose a change that contradicts them) +- `${SKILL_SOURCE_PATH}` — the current skill content (the + optimizer reads the whole skill, including any prior + modifications, and proposes a new improvement on top) +- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (so + the optimizer can shape the diff to match upstream conventions + from the start, reducing validator round-trips) +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (may be + empty) + +The optimizer does NOT see: + +- Raw failed trials (`findings.txt`, `trace.jsonl` from + `05-bench-results/`) +- Grader internals (`tests///grader.mjs` source) +- Test inputs (`tests///workspace/`) +- Its own prior `07-improvement-proposal.md` or git history of it +- Prior `07-validator-verdict.md` (the validator's prior judgment + would bias the optimizer toward defending or pivoting away from + the prior attempt rather than addressing the weakness fresh) + +The "no raw trials / grader internals / test inputs" constraint is +load-bearing: it forces the optimizer to address the weakness as +the analyzer named it, in terms of the general principle the +analyzer articulated — not by pattern-matching a patch that would +make specific failing trials pass. + +The optimizer writes `07-improvement-proposal.md` (the diff + +rationale, including the explicit self-check against the +anti-pattern list) and returns a brief summary: the weakness(es) +addressed, the principle applied, the lines/sections of the skill +modified. + +### (d) Dispatch the validator subagent + +After the optimizer returns, dispatch the validator subagent. Load +the prompt template at +[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) +and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}`, +`${SKILL_AFTER_PATH}` (the optimizer's proposed result — +materialize as a temporary file or pass the diff applied to a +temp copy), `${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if +PR-bound, `${VERDICT_OUTPUT_PATH}`, `${ROUND}` (1 on first +dispatch; 2 on revision). + +The validator sees: + +- The skill BEFORE the change +- The skill AFTER the change +- `${PROPOSAL_PATH}` — `07-improvement-proposal.md` (so it knows + what the optimizer claims to have done) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the + internal consistency check: does the change make sense given + the skill's stated responsibilities?) +- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for + the external consistency check) + +The validator does NOT see: + +- Raw failed trials, `findings.txt`, `trace.jsonl` +- Test inputs (`tests//workspace/`) +- The optimizer's internal reasoning trace (only the proposal + artifact, not how the optimizer arrived at it) +- Its own prior `07-validator-verdict.md` or git history of it +- `06-analysis.md` directly (the analyzer's reasoning is mediated + through the optimizer's proposal — the validator's job is to + check the proposal as an independent observer, not to second- + guess the analysis) + +The "no prior verdict" constraint is load-bearing: each validator +round is independent. If the validator saw its prior verdict, it +would gravitate toward consistency-with-itself, defeating the +purpose of revision rounds. + +The validator writes `07-validator-verdict.md` with a verdict of +`approve`, `needs-revision`, or `reject`, plus rationale. + +### (e) Handle the verdict (bounded revision loop) + +Three branches: + +**If `verdict: approve`:** proceed to (f). + +**If `verdict: needs-revision` AND `round < 2`:** re-dispatch the +optimizer with the validator's rationale as a directive. Example +directive: `validator round 1 said: '' — address this and +re-propose`. Increment round to 2. Loop back to (c) for the +re-dispatch; the optimizer derives fresh per its rules (it does +NOT see the prior proposal directly, only the distilled directive). +Then re-dispatch the validator at (d) with `${ROUND}: 2`. + +**If `verdict: needs-revision` AND `round == 2`:** the bound is +reached. Surface honestly: + +> After 2 revision rounds, the validator still says +> needs-revision: ``. The optimizer and validator have not +> converged on this weakness within the allowed bound. Three +> realistic paths: (1) re-invoke step 6 with a directive to +> reformulate the weakness ("the current naming was too abstract; +> split it"); (2) re-invoke step 7 manually with directives that +> bridge the disagreement; (3) accept that this weakness is not +> fixable in this round and exit honestly. + +Do NOT auto-extend the loop. The bound exists to prevent +pathological disagreement burning operator budget. + +**If `verdict: reject`:** surface honestly and exit. Reject means +the validator judges the proposal is fundamentally wrong (not just +in need of revision). The reasonable next steps are to re-invoke +step 6 (the analyzer's framing of the weakness may have been +misleading) or to accept that the named weakness isn't addressable +without changes the validator is unwilling to approve. + +### (f) Write outputs + +After `verdict: approve`: + +1. Apply the proposed diff to the source skill file. For local + skills, modify in place. For upstream skills, modify the + vendored copy at `vendored-skill/`. + +2. Confirm `07-improvement-proposal.md` and + `07-validator-verdict.md` are written to canonical paths. + +3. Commit the change: + + ```bash + git add docs/skill-optimizer// + git commit -m "step 7: improve — address " + ``` + +### (g) Hand off + +Three branches based on local/upstream + PR intent (decided at +step 1, no late prompts): + +**Local skill:** + +> Improvement applied to `` in place. The chain +> is complete for this iteration. If you want to re-bench against +> the modified skill, invoke `skill-optimizer-run-bench` again. + +**Upstream skill, `pr_submission_intent: false`:** + +> Improvement applied to the vendored copy at +> `vendored-skill/`. No PR will be packaged (per the +> decision at step 1). If you want to re-bench against the +> modified skill, invoke `skill-optimizer-run-bench` again. + +**Upstream skill, `pr_submission_intent: true`:** + +Write `07-pr-draft.md` containing: + +- The diff (unified format) +- A PR body that references the named weakness from + `06-analysis.md` and the general principle applied +- Caveats from `03-submissions.md`: CLA requirement (if + `requires_cla: true`), branch target (from `upstream_branch_target`), + license (from `license`), any rejection signals flagged in + `03-submissions.md`'s closed-without-merge PR review +- Operator-steps-to-submit: + - "Fork the upstream repo if you haven't already" + - "Sign the CLA at ``" (if `requires_cla: true`) + - "Open the PR against `:` + with the body and diff above" + - "Watch for the CI / reviewer responses" + +Handoff: + +> Improvement applied to the vendored copy. PR draft written to +> `07-pr-draft.md`. The chain does NOT submit the PR — review the +> draft, then submit it manually following the operator steps in +> the draft. + +The chain stops here. PR submission is operator-judgment work that +the agent should not attempt automatically (the operator needs to +authenticate, sign CLAs if required, respond to reviewer feedback). + +## Why limited-context dispatch matters + +Two distinct concerns, one for each subagent: + +**For the optimizer:** the chain's whole anti-ducktape architecture +hinges on the optimizer NOT seeing raw failures. If the optimizer +read the failed trials, it would pattern-match a patch that makes +those specific trials pass — the textbook ducktape failure mode. +The analyzer's job (step 6) is to translate raw failures into +**named structural weaknesses with general principles and +anti-patterns**; the optimizer's job is to apply the principle. +The translation through the analyzer's report is what enforces +principled improvement over pattern-match patching. + +**For the validator:** independence from the optimizer's reasoning +is load-bearing. The validator must judge the proposal as if seeing +it for the first time, against the skill's stated responsibilities +(internal) and the upstream's PR conventions (external). If the +validator saw the optimizer's reasoning trace, it would tend to +accept arguments the optimizer made about why the change is sound +— defeating the purpose of independent checking. The validator +sees the proposal artifact (what was changed and the optimizer's +brief rationale) but NOT the optimizer's internal reasoning trace. + +Both subagents also don't read their own prior canonicals: the +optimizer doesn't see prior proposals (would gravitate toward +defending or pivoting away from prior attempts), the validator +doesn't see prior verdicts (would gravitate toward +consistency-with-itself across rounds). Each invocation is +genuinely fresh. + +If you find yourself thinking "I'll just propose the change +myself, I know what the analysis says" — that's the failure mode +this chain is built to prevent. Dispatch. + +## Edge cases + +- **`06-analysis.md` missing** — caught at (a). Tell user to run + step 6 first. +- **`has_structural_weakness: false`** — caught at (a). Refuse to + fire; this is the anti-ducktape gate. Don't try to override. +- **Optimizer reports BLOCKED** (e.g., the analyzer's named + weakness is too abstract to derive a concrete diff from) — + surface to the user. The fix is typically a step 6 re-run with a + directive ("weakness X as named was too abstract; split into + concrete sub-weaknesses"). Per the iteration protocol's re-run + authorization rule, the user invokes step 6. +- **Optimizer's self-check against anti-patterns is missing or + vague** — the optimizer's report MUST explicitly state why its + proposal is not one of the ducktape moves the analyzer flagged. + If this section is missing, treat as `needs-revision` (route + through (e)) with the directive "the self-check against the + anti-pattern list is required and missing". +- **Validator rejects with reasoning that contradicts the + analyzer's framing** (e.g., validator says "this change doesn't + address the real problem"; the analyzer thought it did) — that's + a signal the analyzer/validator are misaligned on what the + weakness actually is. Surface to the user with two options: + re-invoke step 6 to reformulate the weakness, or re-invoke step + 7 with a directive that bridges the framing. +- **PR-bound but `03-submissions.md` missing** — caught at (a). + Ask the user whether to run step 3 first or proceed with + internal consistency only (will affect PR packaging quality). +- **Operator directives reference a specific section of the prior + proposal** (e.g., "the optimizer's section on the description + field was wrong") — that's a context-dump masquerading as a + directive. Translate it to an atomic new requirement ("don't + modify the description field") before passing to the subagent. + +## Iteration behavior + +Step 7 is re-runnable and is a fresh-derivation step (for both +optimizer and validator). Re-run triggers specific to this step: + +- Step 6 (`analyze-result`) produced a new analysis — + `06-analysis.md` changed via git mtime, possibly naming + different weaknesses +- User wants the optimizer to try a different approach (passes + directives like "this time prefer additive changes" or "focus on + weakness 2 only") +- Validator returned `reject` on a prior run and the operator has + a new framing to try + +General iteration mechanics — staleness detection (git-native), +operator directives (the channel for prior-derivation lessons), +the step-7 carve-out for the skill being upstream input — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. + +The internal optimizer/validator revision loop (max 2 rounds) is +not "iteration" in the chain sense — it's bounded refinement +within a single invocation of step 7. The chain-level iteration is +re-invoking step 7 entirely (which resets the round counter to 1). From f68339d79588cee49ea2116e000e6472d32cd706 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 07:15:38 -0500 Subject: [PATCH 029/121] fix(skill-optimizer-analyze-result): move body template to subagent prompt MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit B6 SKILL.md was carrying a full markdown body template (per-weakness section template with placeholders, non-structural noise example, honest-refusal wording verbatim). That's the subagent's concern — the subagent writes the body per its prompt template; the operator session reads only the frontmatter for handoff branching. Keep in the SKILL.md: - Frontmatter contract (operator reads has_structural_weakness for step 7's gate) - Enumeration of the five required parts per weakness entry (the operator session verifies these in step (d)) - The architecture-level rationale on why the anti-pattern list is load-bearing (this is design-decision content, not subagent-side prose — explains WHY the constraint exists for future authors) - Pointer to the subagent prompt template for the full template + reasoning protocol Same audit pass on B1/B2/B4/B7: their What-you-produce sections show structural contracts (frontmatter schemas, file-tree shape) that the operator session actually reads, not narrative body templates the subagent fills in — so they stay as-is. --- .../skill-optimizer-analyze-result/SKILL.md | 64 ++++++------------- 1 file changed, 18 insertions(+), 46 deletions(-) diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md index c786974..746445d 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -68,54 +68,26 @@ the ducktape failure mode this step exists to prevent. output the analyzer read from. Lets step 7's optimizer (if it fires) trace back to the underlying data without re-deriving. -Body structure: - -```markdown -## Structural weaknesses identified - -### Weakness 1: - -- **Pattern**: Across trials, systematically failed to - detect . Specifically: //>. -- **Hypothesized cause**: . -- **Connects to skill section**: at . -- **What WOULD address this**: . -- **What WOULD NOT address this** (anti-ducktape gate): . - -### Weakness 2: ... - -## Non-structural noise (ignored — not addressable) - -- 1 gemini transient API error on multi-redirect (not a pattern) -- 1 gpt-5 timeout (infrastructure, not skill) - -## Honest refusal (if applicable) - -If no structural weakness can be articulated: - -> No structural weakness identified. Failures observed are -> consistent with model nondeterminism / infrastructure noise / -> single-trial flakiness, not a fixable defect in the skill itself. -> Step 7 should not fire. - -The frontmatter `has_structural_weakness: false` mirrors this. -``` +The body has two top-level sections (per-weakness entries + +non-structural noise) and an optional honest-refusal note if no +weakness was identified. **Each per-weakness entry must include +five required parts**: Pattern, Hypothesized cause, Connects to +skill section, What WOULD address this, What WOULD NOT address +this. The operator session verifies these five parts are present +in step (d); the subagent prompt template specifies their exact +shape. The "What WOULD NOT address this" anti-pattern list is -**load-bearing**. Without it, step 7's optimizer can pattern-match -a patch that fits the symptom without addressing the cause; the -validator then has no explicit "this would be a ducktape" signal -to check against. The analyzer's job is to name BOTH the principle -that should be applied AND the ducktape moves the optimizer must -avoid. - -For the body template details and the analyzer's reasoning -protocol, see +**load-bearing architecture, not just bookkeeping**. Without it, +step 7's optimizer can pattern-match a patch that fits the symptom +without addressing the cause; the validator then has no explicit +"this would be a ducktape" signal to check against. The analyzer's +job is to name BOTH the principle that should be applied AND the +ducktape moves the optimizer must avoid. This is the chain's +anti-ducktape gate — losing the anti-pattern list breaks the gate. + +For the full body template, the analyzer's reasoning protocol, and +the honest-refusal wording, see [`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md). The subagent writes the report body itself. From d3d91b1de2dcbd730afa3aa27d631e87c7e6bcd7 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 07:26:54 -0500 Subject: [PATCH 030/121] =?UTF-8?q?feat(v1.4-chain):=20B7=20scope=20reduct?= =?UTF-8?q?ion=20=E2=80=94=20drop=20PR=20draft=20+=20preserve=20original?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two architectural fixes per review: 1. Drop PR draft from B7. PR composition is a separate downstream concern that consumes B7's proposal + 03-submissions.md; the auto-pilot (step 8) or a dedicated composer can handle it if pr_submission_intent: true. B7's job is just to improve the skill — packaging it as a PR is not what improve-skill does. Removed: 07-pr-draft.md as an artifact, the three-branch handoff (local / upstream + PR=false / upstream + PR=true), the operator-steps-to-submit checklist. Collapses to a single handoff message regardless of PR intent. 2. Never modify the original skill. The improved version lives at docs/skill-optimizer//improved-skill/ — a separate location that accumulates improvements across iterations. The vendored upstream copy stays frozen; the local source file stays untouched. Git tracks improved-skill/ history. On re-run, the optimizer reads improved-skill/ if it exists (the accumulated state) and proposes the next improvement on top; iteration 1 reads the original source instead. Original is always recoverable; improved evolves under git. Removed: "Written in place for local skills" / "Written to vendored-skill/ for upstream skills" — both wrong now. Updated B7's "Before you start" carve-out summary, step (c) and (d) input descriptions (SKILL_CURRENT_PATH replaces SKILL_SOURCE_PATH; validator BEFORE = optimizer's input), step (f) write outputs (materialize improved-skill/; don't touch source), step (g) handoff (single message). Iteration-protocol's step-7 carve-out updated to match (improved-skill/ is the accumulated state; source stays frozen). Spec doc layout adds improved-skill/ alongside vendored-skill/; B7 section rewritten; subagent constraints table rows for optimizer + validator updated. --- docs/skill-optimizer-v1.4-spec.md | 63 +++--- skills/skill-optimizer-improve-skill/SKILL.md | 204 +++++++++--------- .../iteration-protocol.md | 22 +- 3 files changed, 157 insertions(+), 132 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index baa8019..a8a2601 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -128,7 +128,8 @@ docs/skill-optimizer// 06-analysis.md 07-improvement-proposal.md 07-validator-verdict.md - vendored-skill/ # the source skill, read-only after fetch + vendored-skill/ # the source skill, read-only after fetch (upstream only) + improved-skill/ # B7's output: accumulated improvements; original is never modified ``` `` is `--` for upstream skills, or @@ -172,8 +173,8 @@ and dispatches the subagent with only the narrow chunks it needs. | Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change, `03-submissions.md` (own canonical) or its git history | Fresh-derivation: just upstream facts | | Test writer (step 4, dispatched per probe) | The single probe's `spec.yaml` + parent functionality's `spec.yaml` + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other probes' specs/graders, the eval grader's matching logic, sibling probe contents (own canonical at the per-probe level), git history of any tree files | Maintenance at tree level: each probe is built in isolation. Fixture writing needs source detail (specific patterns); the probe spec from step 2 bounds the gerrymandering risk. Blocked from seeing siblings (prevents copying) and grader internals (prevents grader-leak hacking) | | Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), `06-analysis.md` (own canonical) or its git history | Fresh-derivation: forces focus on the SKILL, not the SOLUTIONS; iteration-isolated | -| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + source skill content + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `07-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts | -| Validator (step 7, after optimizer) | Skill BEFORE + skill AFTER + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `07-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | +| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + current skill state (`improved-skill/` if it exists, else original source) + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `07-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts. The skill content is upstream input; the report is the own canonical | +| Validator (step 7, after optimizer) | Skill BEFORE (same as optimizer's input) + skill AFTER (temporary materialization) + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `07-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | The skill (operator session) sees everything; subagents see slices. This is the architectural fix for the "tunnel-vision into ducktape" @@ -540,19 +541,26 @@ Per-step iteration behavior is noted at the end of each subsection. - **Description trigger:** "improve this skill", "fix the structural weakness", "optimize" - **Input:** `06-analysis.md` (REQUIRED — refuses if no weakness - identified), `01-functionality.md`, source skill, optionally + identified), `01-functionality.md`, current skill state + (`improved-skill/` if it exists, else original source), optionally `03-submissions.md` -- **Output:** `07-improvement-proposal.md` (proposed diff + rationale - referencing the structural weakness) + `07-validator-verdict.md` + - modified skill file (if approved) +- **Output:** `07-improvement-proposal.md` (the optimizer's proposed + change + rationale referencing the structural weakness) + + `07-validator-verdict.md` + `improved-skill/` (the improved skill + content, materialized only if `verdict: approve`). **The original + source is never modified** — vendored copies stay frozen, local + source files stay untouched. The improved version lives at + `improved-skill/` and accumulates across iterations; git tracks + its history. - **Behavior:** - 1. Refuse if `06-analysis.md` has no structural weakness — print - "no weakness to address" and exit + 1. Refuse if `06-analysis.md` has `has_structural_weakness: false` + — print "no weakness to address" and exit 2. Dispatch **optimizer subagent** with limited context (see "Subagent constraints" table). Required: address the named structural weakness using a general principle from the analysis, - NOT a pattern-match patch. Output the proposed diff + rationale - that explicitly references which named weakness it addresses. + NOT a pattern-match patch. Output the proposed change + + rationale that explicitly references which named weakness it + addresses. 3. Dispatch **validator subagent** with limited context. - **Internal consistency check:** does the proposed change make sense given the skill's stated responsibilities in @@ -561,29 +569,34 @@ Per-step iteration behavior is noted at the end of each subsection. - **External consistency check (only if `03-submissions.md` exists):** does the change conform to the upstream's PR rules (frontmatter, file location, prefix taxonomy, additive-only, - etc.)? + etc.)? Forward-looking — this verifies the improved skill + COULD be turned into a valid PR, even though B7 itself does + not produce one. - Verdict: `approve` / `needs-revision` / `reject` 4. If `needs-revision`: optimizer revises (max 2 revision rounds). If `reject`: surface honestly and exit. - 5. If `approve`: write the modified skill file + improvement - proposal report. + 5. If `approve`: materialize `improved-skill/` by applying the + diff to a copy of the current state (NOT to the original). - **Dispatches:** optimizer + validator, both limited context, both isolated from raw trial data -- **Handoff (local skill):** done. Modified skill written in place. -- **Handoff (upstream skill, `pr_submission_intent: false`):** done. - Modified skill written to vendored copy. No PR packaging. -- **Handoff (upstream skill, `pr_submission_intent: true`):** package - the proposed change as a PR draft using `03-submissions.md` (which - exists because step 3 ran). The draft includes the diff, the body, - caveats, and operator-steps-to-submit. No late "submit a PR?" - prompt — the decision was already made at step 1. +- **Out of scope:** packaging the change as a PR draft. PR + composition is a separate downstream concern that consumes B7's + proposal + `03-submissions.md`; the auto-pilot (step 8) or a + dedicated composer can handle it if `pr_submission_intent: true`. + B7 just improves the skill. +- **Handoff (single message, no PR-intent branching):** improvement + complete; `improved-skill/` is the new state; original source is + untouched; review and decide next steps (apply locally, hand to PR + composer, or re-bench against `improved-skill/`). - **Iteration behavior:** re-run when `06-analysis.md` changed (new analysis = potentially different weakness) or when the operator wants a fresh optimization attempt. Fresh-derivation for both optimizer and validator — neither subagent reads its - canonical or git history. The internal optimizer/validator loop - (max 2 rounds, baked into the skill) handles in-step iteration; - cross-step backtracking is the operator's call. + own canonical (`07-improvement-proposal.md` / + `07-validator-verdict.md`) or git history of it. On re-run, the + optimizer reads `improved-skill/` (if it exists from a prior + step-7 run) as the current state and proposes the next + improvement on top; the original source stays frozen. `${OPERATOR_DIRECTIVES}` carries hints like "prefer additive changes" or "don't touch the description field" — additional constraints, never a context dump of prior attempts. diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md index 76bea53..9d6f414 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -10,15 +10,24 @@ Takes the named structural weaknesses from step 6, dispatches an optimizer subagent to draft a principled fix, then dispatches a validator subagent to independently check whether the fix is sound (internal consistency) and conformant (external PR conventions if -PR-bound). Writes the modified skill file, the improvement -proposal, and the validator's verdict. If the user opted into PR -submission at step 1, packages the change as a PR draft. +PR-bound). Writes the improved skill to a separate location +(`improved-skill/`), the improvement proposal, and the validator's +verdict. **The original skill is never modified** — vendored +upstream copies stay frozen as reference; local source files stay +untouched. The operator reviews `improved-skill/` and decides what +to do with it (apply locally, hand to a PR composer, discard). This skill **refuses to fire** if `06-analysis.md` has `has_structural_weakness: false` — there's nothing to optimize, and forcing a fix in that situation is the ducktape failure mode the chain is built to prevent. +**Out of scope:** packaging the change as a PR draft. The chain +treats PR composition as a separate concern that consumes B7's +rationale + `03-submissions.md`'s upstream facts; the auto-pilot +(step 8) or a dedicated composer can handle it downstream of this +step if `pr_submission_intent: true`. B7 just improves the skill. + ## Before you start Three load-bearing pieces of context to load NOW, before the @@ -30,12 +39,15 @@ workflow: fresh-derivation subagents); the iteration protocol's "frontmatter discipline" + "directive channel" sections both apply. Pay attention to the **step-7 carve-out** in the - subagent-constraints section: the SKILL itself is upstream - input (the optimizer reads the current skill, which may include - modifications from prior step-7 runs), but the optimizer's own - canonical (`07-improvement-proposal.md`) and the validator's - own canonical (`07-validator-verdict.md`) are both off-limits - to their respective subagents. + subagent-constraints section: the SKILL CONTENT is upstream + input — `improved-skill/` if it exists (the accumulated + improvement state from prior step-7 runs), else the original + source (`vendored-skill/` for upstream, the user's local file + for local). The optimizer's own canonical + (`07-improvement-proposal.md`) and the validator's own + canonical (`07-validator-verdict.md`) are both off-limits to + their respective subagents — those are the reports, not the + skill being improved. 2. **You will dispatch two subagents (optimizer + validator); you do NOT propose diffs or judge them yourself in this session.** @@ -55,8 +67,8 @@ skill are labelled "(a)" through "(g)". ## What you produce -Three or four artifacts at `docs/skill-optimizer//`, -depending on the verdict: +Two or three artifacts at `docs/skill-optimizer//`, depending +on the verdict: 1. **`07-improvement-proposal.md`** — the optimizer's proposed diff + rationale. Frontmatter (runtime-relevant facts only, @@ -67,17 +79,14 @@ depending on the verdict: addresses_weaknesses: - - ... - diff_target: --- ``` - Body sections: the proposed diff (in fenced unified-diff - format), the rationale explicitly referencing which named - weakness from `06-analysis.md` it addresses and which "What - WOULD address this" principle it applies, and a self-check - against the "What WOULD NOT address this" anti-pattern list - (the optimizer states explicitly why its proposal is NOT one - of the ducktape moves the analyzer flagged). + The body lists the proposed change, the rationale referencing + the named weakness from `06-analysis.md` and the "What WOULD + address this" principle being applied, and a self-check against + the "What WOULD NOT address this" anti-pattern list. Subagent + prompt template specifies the section shape. 2. **`07-validator-verdict.md`** — the validator's independent judgment. Frontmatter: @@ -92,24 +101,26 @@ depending on the verdict: --- ``` - Body: internal consistency check (does the change make sense - for the named weakness? is it additive vs. destructive? is it - general vs. ducktape?), and — if `03-submissions.md` exists — - external consistency check (does the change conform to upstream - PR rules: frontmatter, file location, prefix taxonomy, - additive-only, etc.). - -3. **The modified skill file** (only if `verdict: approve`). - Written in place for local skills; written to - `vendored-skill/` for upstream skills. - -4. **`07-pr-draft.md`** (only if `pr_submission_intent: true` - AND `verdict: approve`). A PR draft containing the diff, the - PR body referencing the weakness and the principle applied, - caveats the operator should know about (CLA requirement, branch - target, anything flagged in `03-submissions.md`), and - operator-steps-to-submit. The chain does NOT submit the PR — - that's an explicit operator action. + The body covers the internal consistency check (does the + change make sense for the named weakness? additive vs. + destructive? general vs. ducktape?), and — if + `03-submissions.md` exists — the external consistency check + (does the change conform to upstream PR rules: frontmatter, + file location, prefix taxonomy, additive-only, etc.). The + external check is forward-looking (would a PR carrying this + change be conformant?) even though B7 itself does not produce + a PR. + +3. **`improved-skill/`** — the improved skill content (only if + `verdict: approve`). Mirrors the source skill's directory + structure with the optimizer's diff applied. **The original + is never modified**: for upstream skills, `vendored-skill/` + stays frozen as the upstream reference; for local skills, the + user's source file stays untouched. The operator reviews + `improved-skill/` and decides what to do next — apply locally + by copying over their file, hand to a PR composer (auto-pilot + or operator-driven) for upstream submission, or discard. Git + tracks the history of `improved-skill/` across iterations. For the body templates of the optimizer and validator reports + their reasoning protocols, see @@ -117,8 +128,8 @@ their reasoning protocols, see and [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). Each subagent writes its own body; the operator session orchestrates -the loop and writes the PR draft (which is operator-judgment work, -not generative). +the loop and materializes `improved-skill/` by applying the +optimizer's diff. ## Workflow @@ -204,9 +215,13 @@ The optimizer sees: - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (so the optimizer understands the skill's stated responsibilities and doesn't propose a change that contradicts them) -- `${SKILL_SOURCE_PATH}` — the current skill content (the - optimizer reads the whole skill, including any prior - modifications, and proposes a new improvement on top) +- `${SKILL_CURRENT_PATH}` — the current state of the skill being + improved: `improved-skill/` if it exists (the accumulated state + from prior step-7 runs), else the original source + (`vendored-skill/` for upstream skills, the local file path + recorded in `01-functionality.md`'s `skill_source` for local + skills). The optimizer proposes a new improvement on top of + whatever current state it sees. - `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (so the optimizer can shape the diff to match upstream conventions from the start, reducing validator round-trips) @@ -241,17 +256,21 @@ modified. After the optimizer returns, dispatch the validator subagent. Load the prompt template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) -and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}`, -`${SKILL_AFTER_PATH}` (the optimizer's proposed result — -materialize as a temporary file or pass the diff applied to a -temp copy), `${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if -PR-bound, `${VERDICT_OUTPUT_PATH}`, `${ROUND}` (1 on first -dispatch; 2 on revision). +and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (same as +the optimizer's `${SKILL_CURRENT_PATH}` from (c) — `improved-skill/` +if it exists, else the source), `${SKILL_AFTER_PATH}` (the proposed +result — materialize by applying the optimizer's diff to a +temporary copy of BEFORE; do NOT overwrite the canonical +`improved-skill/` yet — that happens at (f) after approve), +`${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, +`${VERDICT_OUTPUT_PATH}`, `${ROUND}` (1 on first dispatch; 2 on +revision). The validator sees: -- The skill BEFORE the change -- The skill AFTER the change +- The skill BEFORE the change (same state the optimizer read) +- The skill AFTER the change (a temporary materialization; the + canonical `improved-skill/` is not yet updated) - `${PROPOSAL_PATH}` — `07-improvement-proposal.md` (so it knows what the optimizer claims to have done) - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the @@ -320,9 +339,12 @@ without changes the validator is unwilling to approve. After `verdict: approve`: -1. Apply the proposed diff to the source skill file. For local - skills, modify in place. For upstream skills, modify the - vendored copy at `vendored-skill/`. +1. Materialize `improved-skill/` by copying the source skill's + directory structure and applying the optimizer's diff. **Do + NOT modify the source.** For upstream skills, leave + `vendored-skill/` frozen and write to a sibling + `improved-skill/` directory; for local skills, copy the user's + source file(s) to `improved-skill/` and apply the diff there. 2. Confirm `07-improvement-proposal.md` and `07-validator-verdict.md` are written to canonical paths. @@ -330,56 +352,37 @@ After `verdict: approve`: 3. Commit the change: ```bash - git add docs/skill-optimizer// + git add docs/skill-optimizer// git commit -m "step 7: improve — address " ``` -### (g) Hand off - -Three branches based on local/upstream + PR intent (decided at -step 1, no late prompts): - -**Local skill:** - -> Improvement applied to `` in place. The chain -> is complete for this iteration. If you want to re-bench against -> the modified skill, invoke `skill-optimizer-run-bench` again. - -**Upstream skill, `pr_submission_intent: false`:** + Note the source skill file is NOT in the commit — git tracks + `improved-skill/` history alongside the proposal and verdict + reports, but never touches the user's source or the vendored + reference. -> Improvement applied to the vendored copy at -> `vendored-skill/`. No PR will be packaged (per the -> decision at step 1). If you want to re-bench against the -> modified skill, invoke `skill-optimizer-run-bench` again. - -**Upstream skill, `pr_submission_intent: true`:** - -Write `07-pr-draft.md` containing: - -- The diff (unified format) -- A PR body that references the named weakness from - `06-analysis.md` and the general principle applied -- Caveats from `03-submissions.md`: CLA requirement (if - `requires_cla: true`), branch target (from `upstream_branch_target`), - license (from `license`), any rejection signals flagged in - `03-submissions.md`'s closed-without-merge PR review -- Operator-steps-to-submit: - - "Fork the upstream repo if you haven't already" - - "Sign the CLA at ``" (if `requires_cla: true`) - - "Open the PR against `:` - with the body and diff above" - - "Watch for the CI / reviewer responses" - -Handoff: - -> Improvement applied to the vendored copy. PR draft written to -> `07-pr-draft.md`. The chain does NOT submit the PR — review the -> draft, then submit it manually following the operator steps in -> the draft. +### (g) Hand off -The chain stops here. PR submission is operator-judgment work that -the agent should not attempt automatically (the operator needs to -authenticate, sign CLAs if required, respond to reviewer feedback). +Single handoff message — no branching on PR intent. PR packaging +is a separate downstream concern (auto-pilot or operator-driven); +this step's job is done when the improved skill is materialized +and verified. + +> Improvement complete. The improved skill is at +> `docs/skill-optimizer//improved-skill/`; the original +> source (vendored or local) is untouched. The proposal and +> validator verdict are in `07-improvement-proposal.md` and +> `07-validator-verdict.md`. Three realistic next steps: (1) +> review `improved-skill/` and copy it over your local source if +> you're satisfied; (2) if PR-bound, hand the improved skill + +> the proposal + `03-submissions.md` to a PR composer (auto-pilot +> can do this end-to-end if you're running step 8); (3) re-bench +> against the improved skill by re-invoking +> `skill-optimizer-run-bench` against an updated workbench that +> points at `improved-skill/`. + +The chain stops here. Whether to apply the improvement, submit it +upstream, or iterate further is an explicit operator decision. ## Why limited-context dispatch matters @@ -443,7 +446,10 @@ this chain is built to prevent. Dispatch. 7 with a directive that bridges the framing. - **PR-bound but `03-submissions.md` missing** — caught at (a). Ask the user whether to run step 3 first or proceed with - internal consistency only (will affect PR packaging quality). + internal consistency only. Proceeding loses the validator's + external consistency check; the improved skill will still be + produced, but any downstream PR composer won't have upstream + conventions to conform to. - **Operator directives reference a specific section of the prior proposal** (e.g., "the optimizer's section on the description field was wrong") — that's a context-dump masquerading as a diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index c14b803..73517cf 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -112,14 +112,20 @@ Three rules every reasoning subagent must follow: neither you nor any subsequent fresh-derivation subagent should walk that history. - **For step 7 specifically:** the SKILL itself (the target being - improved) is upstream input, not your own canonical. You read - the current skill — which may include modifications from prior - step-7 runs — and propose a new improvement on top. Your "own - canonical" that's off-limits is `07-improvement-proposal.md` - (the report describing your reasoning), not the skill file - itself. The skill accumulates improvements across iterations; - the proposal reports do not. + **For step 7 specifically:** the SKILL CONTENT (the target + being improved) is upstream input, not your own canonical. You + read the **current state of the skill** — which is + `improved-skill/` if it exists (the accumulated state from + prior step-7 runs), else the original source + (`vendored-skill/` for upstream, the local file path for + local). Step 7 never modifies the original source: the + accumulated improvements live in `improved-skill/`, and the + original (vendored or local) stays frozen as reference. Your + "own canonical" that's off-limits is + `07-improvement-proposal.md` (the report describing your + reasoning), not the skill content itself. The skill content + accumulates improvements across iterations in + `improved-skill/`; the proposal reports do not. 3. **For maintenance steps (2, 4): DO read your own canonical tree** (when it exists). Your job on a re-run is to extend or From 1a50a9f8f7e421b15febe011af3b52440c234f1c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 07:31:11 -0500 Subject: [PATCH 031/121] fix(skill-optimizer-investigate-test-case): three review fixes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 1. Rename 00-test-proposals.md -> 02-test-proposals.md. The 00- prefix was inconsistent with the chain's per-step numbering convention (B2's outputs should start with 02-). Renamed in both the SKILL.md and the spec doc layout. 2. Soften the "don't auto-flip" wording. The intent is "no proactive flipping without user direction", not "user must edit every spec.yaml by hand". If the user explicitly says "flip these to true" or "pick X, Y, Z", the operator session does it for them and confirms what was set. 3. Single path for user-added functionalities. The prior wording offered two paths (manual spec.yaml creation by operator OR subagent re-dispatch with directive). The first violates the architecture's no-operator-generative-writing rule — only the subagent writes test-design content. Collapsed to the single correct path: treat user's description as a directive and re-dispatch per step (e)(2). --- docs/skill-optimizer-v1.4-spec.md | 4 +-- .../SKILL.md | 36 +++++++++++-------- 2 files changed, 23 insertions(+), 17 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index a8a2601..5a8e896 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -107,7 +107,7 @@ patterns in. See plan §"Task A4" for details. ```text docs/skill-optimizer// 01-functionality.md - 00-test-proposals.md # B2's audit report (ranking + reasoning) + 02-test-proposals.md # B2's audit report (ranking + reasoning) tests/ / # one folder per B2-proposed functionality spec.yaml # B2 wrote: picked, importance, suggested_probes @@ -339,7 +339,7 @@ Per-step iteration behavior is noted at the end of each subsection. test cases", "what should we test" - **Input:** `01-functionality.md` - **Output (two artifacts):** - - **`00-test-proposals.md`** — one-time audit report from the + - **`02-test-proposals.md`** — one-time audit report from the subagent: ranked list of proposed functionalities, with importance + suggested_probes + why-it-matters for each. This file captures the design reasoning; it is not consulted by diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 35b8b0a..54c48a6 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -12,7 +12,7 @@ test (each functionality = one responsibility the skill must fulfill), then asks the user to pick which ones to actually build probes for at step 4. Writes a `tests//spec.yaml` file per proposed functionality (the filesystem IS the state) plus a -one-time audit report at `00-test-proposals.md`. +one-time audit report at `02-test-proposals.md`. ## Before you start @@ -39,7 +39,7 @@ THIS skill are labelled "(a)" through "(f)". Two artifacts at `docs/skill-optimizer//`, where `` matches the slug from step 1's report: -1. **`00-test-proposals.md`** — a one-time audit report from the +1. **`02-test-proposals.md`** — a one-time audit report from the subagent: the ranked list of proposed functionalities with full reasoning (why each one matters, what coverage it adds, what probes are suggested). This file is for human review of the @@ -70,7 +70,7 @@ matches the slug from step 1's report: anywhere. Step 4 builds probes only for functionalities whose `spec.yaml` has `picked: true`. -For the body template of `00-test-proposals.md` (sections, ranking +For the body template of `02-test-proposals.md` (sections, ranking format) and the exact `spec.yaml` field list, see [`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md). The subagent writes both the audit report and the per-functionality @@ -130,7 +130,7 @@ The subagent sees: maintenance rule. - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new requirements (may be empty) -- Output paths for `00-test-proposals.md` and the `tests/` tree +- Output paths for `02-test-proposals.md` and the `tests/` tree The subagent does NOT see: @@ -144,7 +144,7 @@ The subagent does NOT see: The subagent writes: -- `00-test-proposals.md` (the audit report, full ranked design +- `02-test-proposals.md` (the audit report, full ranked design reasoning) - One `tests//spec.yaml` per proposed functionality. Existing spec.yaml files the user has already @@ -160,7 +160,7 @@ revised, list of newly-added functionalities. Verify: -- `00-test-proposals.md` exists and parses as markdown. +- `02-test-proposals.md` exists and parses as markdown. - Each `tests//spec.yaml` parses as YAML and has the required fields (`name`, `description`, `picked`, `importance`, `suggested_probes`, `why_test`). @@ -171,7 +171,7 @@ output yourself. ### (e) User gate: present proposals, collect picks -Show the user the ranked list from `00-test-proposals.md` (paste +Show the user the ranked list from `02-test-proposals.md` (paste the audit report, or summarize if long — your call). Tell them how to indicate picks: @@ -200,10 +200,13 @@ Three realistic responses: (treat as case 2) or are abandoning the optimization for this skill (exit honestly without progressing to step 4). -Don't auto-flip `picked` on the user's behalf — even if all -functionalities look important, the user owns this decision -(they're paying for the probe-building in step 4 and the bench -run in step 5). +Don't auto-flip `picked` proactively — even if all functionalities +look important, the user owns the decision (they're paying for the +probe-building in step 4 and the bench run in step 5). But if the +user explicitly says "flip these to true" or "pick X, Y, Z", do +that — edit the named spec.yaml files for them and confirm what +you set. The rule is "no flipping without explicit user direction", +not "user must edit every file by hand". ### (f) Hand off @@ -256,10 +259,13 @@ prevent. Dispatch. re-runs require an active signal from the user (per the iteration protocol's "Re-run authorization" section). - **User wants to add a functionality the subagent didn't propose** - — accept it. Either (i) they describe it and you create the - `tests//spec.yaml` manually with their content + `picked: - true`, or (ii) they re-dispatch the subagent with a directive - ("add a functionality for X"). Either path is valid. + — treat their description as an operator directive ("add a + functionality for X") and re-dispatch the subagent per step + (e)(2). Don't write the spec.yaml manually yourself — the + subagent is the only writer of test-design content in this step, + matching the architecture's no-operator-generative-writing rule. + The user can flip the new functionality's `picked: true` after + the re-dispatch, or ask you to. - **Operator directives contradict each other** (e.g., "focus coverage on X" + "ignore X") — surface the contradiction to the user before re-dispatching; don't try to resolve it yourself. From 0528d8cc9e5d9c75415a0dda528af52ae9aad362 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 07:34:24 -0500 Subject: [PATCH 032/121] fix(skill-optimizer-investigate-submissions): drop over-specific PR counts B3 had hard-coded "last 10 merged PRs and last 5 closed-without-merge PRs" both in the SKILL.md "subagent sees" list and in the spec doc "Behavior" line. Two problems: 1. Redundant: the line above in SKILL.md already says "PR list" as part of the gh-CLI access, so the specific-counts bullet was restating with extra constraints. 2. Over-prescriptive: 10/5 are arbitrary; the subagent should sample enough recent PRs to identify shape patterns and rejection signals, but the exact counts are operational judgment not architecture. The subagent prompt template (Phase C) can recommend a starting point; SKILL.md and the spec shouldn't pin it. Collapsed both to a brief mention that the PR list covers both merged and closed-without-merge for shape patterns + rejection signals. --- docs/skill-optimizer-v1.4-spec.md | 7 ++++--- .../SKILL.md | 13 ++++++------- 2 files changed, 10 insertions(+), 10 deletions(-) diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 5a8e896..0b8ab85 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -387,9 +387,10 @@ Per-step iteration behavior is noted at the end of each subsection. file-location rules, prefix taxonomy, PR-shape patterns from recent merged PRs, branch target, rejection signals from closed-without- merge PRs. (Same shape as v1.3's research-upstream subagent output.) -- **Behavior:** gh-CLI heavy (PR list, repo-file API, CONTRIBUTING, - sanity-test source, last 10 merged + last 5 closed-without-merge); - produce verbatim-pastable context block for the validator +- **Behavior:** gh-CLI heavy (PR list — both merged and + closed-without-merge, repo-file API, CONTRIBUTING, sanity-test + source); produce verbatim-pastable context block for the + validator - **Dispatches:** submission-researcher subagent (limited context: public repo facts only) - **Skipped when:** `pr_submission_intent: false` in diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index fa86c98..37a1673 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -118,13 +118,12 @@ substitute the templated inputs (`${UPSTREAM_REPO}` from The subagent sees: -- The upstream repo via `gh` CLI (PR list, repo-file API, - `CONTRIBUTING.md`, license file, existing skill files for - frontmatter spec extraction) -- The last 10 merged PRs and last 5 closed-without-merge PRs (for - shape patterns and rejection signals) -- The skill slug being researched (so it can look at similar PRs in - the same skill category) +- The upstream repo via `gh` CLI (PR list — both merged and + closed-without-merge for shape patterns and rejection signals, + repo-file API, `CONTRIBUTING.md`, license file, existing skill + files for frontmatter spec extraction) +- The skill slug being researched (so it can look at similar PRs + in the same skill category) - `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new requirements (may be empty) - The output path From 08a40c03e0c764e3c885f9b25f2ef8ff870c052d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 08:00:52 -0500 Subject: [PATCH 033/121] feat(v1.4-chain): split B7 into improve-skill + validate-improvement MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit B7 had grown to ~500 lines covering both the optimizer and the validator with an in-step revision loop. Splitting into two single-shot steps cleans the architecture: B7 (improve-skill, ~280 lines): - Reads 06-analysis + 01-functionality + current skill state (improved-skill/ if exists, else original source) + 03-submissions if PR-bound - Dispatches optimizer subagent - Writes 07-improvement-proposal.md ONLY - Does NOT materialize improved-skill/ (that's step 8's job after approval) - Single-shot per invocation; 7->8->7 revision cycle is operator- driven, not in-step - Handoff: "invoke validate-improvement" B8 (validate-improvement, new, ~340 lines): - Reads 07-improvement-proposal.md + current skill + 01-functionality + 03-submissions if PR-bound - Dispatches validator subagent - Writes 08-validator-verdict.md - On verdict: approve: materializes improved-skill/ by applying the diff to a copy of the current state; original source stays frozen - Three handoff branches by verdict (approve / needs-revision / reject); does not auto-invoke step 7 on needs-revision - Single-shot; if needs-revision, operator distills and re-invokes step 7 then step 8 B9 (autopilot, renumbered from 8): - Walks 1->8 (was 1->7) - Handles the 7->8->7 revision loop bounded by max_iterations_per_step - Spec doc + autopilot section updated accordingly Other changes propagated: - Rename 07-validator-verdict.md -> 08-validator-verdict.md in spec layout, B2/B6 cross-references, subagent constraints table - Iteration-protocol step-kinds: add step 8 to fresh-derivation list; expand step-7-specific carve-out to cover both step 7 and step 8 - "step 1 through step 7" -> "step 1 through step 9" in all chain SKILL.md headers - B3 "validator (step 7)" -> "validator (step 8)" (3 instances); "optimizer (step 7)" stays correct - Eliminated the in-step bounded revision loop entirely — each chain skill is now genuinely single-shot per invocation, aligning with the "chain skills don't auto-invoke other chain skills" rule. The bounded loop survives as a cross-step pattern in auto-pilot (B9). --- docs/skill-optimizer-v1.4-spec.md | 170 +++--- .../skill-optimizer-analyze-result/SKILL.md | 4 +- skills/skill-optimizer-improve-skill/SKILL.md | 533 ++++++------------ .../SKILL.md | 2 +- .../SKILL.md | 8 +- .../SKILL.md | 4 +- skills/skill-optimizer-run-bench/SKILL.md | 2 +- .../iteration-protocol.md | 34 +- .../SKILL.md | 337 +++++++++++ skills/skill-optimizer-write-tests/SKILL.md | 2 +- 10 files changed, 658 insertions(+), 438 deletions(-) create mode 100644 skills/skill-optimizer-validate-improvement/SKILL.md diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/skill-optimizer-v1.4-spec.md index 0b8ab85..47c92af 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/skill-optimizer-v1.4-spec.md @@ -71,6 +71,8 @@ skills/ skill-optimizer-run-bench/SKILL.md # 5 skill-optimizer-analyze-result/SKILL.md # 6 skill-optimizer-improve-skill/SKILL.md # 7 + skill-optimizer-validate-improvement/SKILL.md # 8 + skill-optimizer-autopilot/SKILL.md # 9 (chain driver) skill-optimizer/SKILL.md # existing — unchanged skill-optimizer-subagents/ research-functionality.md @@ -126,10 +128,10 @@ docs/skill-optimizer// 05-bench-results// # raw bench output, timestamped per run 05-bench-summary.md # B5 writes: aggregate + pointer to latest 06-analysis.md - 07-improvement-proposal.md - 07-validator-verdict.md + 07-improvement-proposal.md # B7 writes + 08-validator-verdict.md # B8 writes vendored-skill/ # the source skill, read-only after fetch (upstream only) - improved-skill/ # B7's output: accumulated improvements; original is never modified + improved-skill/ # B8 materializes on verdict: approve; original is never modified ``` `` is `--` for upstream skills, or @@ -174,7 +176,7 @@ and dispatches the subagent with only the narrow chunks it needs. | Test writer (step 4, dispatched per probe) | The single probe's `spec.yaml` + parent functionality's `spec.yaml` + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other probes' specs/graders, the eval grader's matching logic, sibling probe contents (own canonical at the per-probe level), git history of any tree files | Maintenance at tree level: each probe is built in isolation. Fixture writing needs source detail (specific patterns); the probe spec from step 2 bounds the gerrymandering risk. Blocked from seeing siblings (prevents copying) and grader internals (prevents grader-leak hacking) | | Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), `06-analysis.md` (own canonical) or its git history | Fresh-derivation: forces focus on the SKILL, not the SOLUTIONS; iteration-isolated | | Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + current skill state (`improved-skill/` if it exists, else original source) + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `07-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts. The skill content is upstream input; the report is the own canonical | -| Validator (step 7, after optimizer) | Skill BEFORE (same as optimizer's input) + skill AFTER (temporary materialization) + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `07-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | +| Validator (step 8) | Skill BEFORE (same as optimizer's input) + skill AFTER (temporary materialization) + `07-improvement-proposal.md` + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `08-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | The skill (operator session) sees everything; subagents see slices. This is the architectural fix for the "tunnel-vision into ducktape" @@ -241,7 +243,7 @@ DOWNSTREAM_T=$(git log -1 --format=%ct -- tests/) In interactive use, the operator typically just knows ("I re-ran step 1, so step 2 needs a re-run"). The git-mtime check is for -auto-pilot (step 8), which walks the chain forward and re-runs any +auto-pilot (step 9), which walks the chain forward and re-runs any downstream older than its direct upstream. ### Re-entry contract @@ -546,63 +548,101 @@ Per-step iteration behavior is noted at the end of each subsection. (`improved-skill/` if it exists, else original source), optionally `03-submissions.md` - **Output:** `07-improvement-proposal.md` (the optimizer's proposed - change + rationale referencing the structural weakness) + - `07-validator-verdict.md` + `improved-skill/` (the improved skill - content, materialized only if `verdict: approve`). **The original - source is never modified** — vendored copies stay frozen, local - source files stay untouched. The improved version lives at - `improved-skill/` and accumulates across iterations; git tracks - its history. + change + rationale referencing the structural weakness). Step 7 + does NOT materialize `improved-skill/` — that's step 8's job + after validator approval. - **Behavior:** 1. Refuse if `06-analysis.md` has `has_structural_weakness: false` — print "no weakness to address" and exit 2. Dispatch **optimizer subagent** with limited context (see "Subagent constraints" table). Required: address the named - structural weakness using a general principle from the analysis, - NOT a pattern-match patch. Output the proposed change + - rationale that explicitly references which named weakness it - addresses. - 3. Dispatch **validator subagent** with limited context. - - **Internal consistency check:** does the proposed change make - sense given the skill's stated responsibilities in - `01-functionality.md`? Is the change additive vs destructive? - Is it general vs ducktape? - - **External consistency check (only if `03-submissions.md` - exists):** does the change conform to the upstream's PR rules - (frontmatter, file location, prefix taxonomy, additive-only, - etc.)? Forward-looking — this verifies the improved skill - COULD be turned into a valid PR, even though B7 itself does - not produce one. - - Verdict: `approve` / `needs-revision` / `reject` - 4. If `needs-revision`: optimizer revises (max 2 revision rounds). - If `reject`: surface honestly and exit. - 5. If `approve`: materialize `improved-skill/` by applying the - diff to a copy of the current state (NOT to the original). -- **Dispatches:** optimizer + validator, both limited context, both - isolated from raw trial data -- **Out of scope:** packaging the change as a PR draft. PR - composition is a separate downstream concern that consumes B7's - proposal + `03-submissions.md`; the auto-pilot (step 8) or a - dedicated composer can handle it if `pr_submission_intent: true`. - B7 just improves the skill. -- **Handoff (single message, no PR-intent branching):** improvement - complete; `improved-skill/` is the new state; original source is - untouched; review and decide next steps (apply locally, hand to PR - composer, or re-bench against `improved-skill/`). + structural weakness using a general principle from the + analysis, NOT a pattern-match patch. Output the proposed + change + rationale that explicitly references which named + weakness it addresses, plus a self-check against the "What + WOULD NOT address this" anti-pattern list. +- **Dispatches:** optimizer subagent, limited context, isolated + from raw trial data +- **Out of scope:** validating the proposal (that's step 8) and + packaging the change as a PR draft (PR composition is a separate + downstream concern; the auto-pilot at step 9 or a dedicated + composer can handle it if `pr_submission_intent: true`). +- **Single-shot per invocation:** no in-step revision loop. If + step 8 returns `needs-revision`, the operator (or auto-pilot) + distills the validator's rationale into a directive and + re-invokes step 7. +- **Handoff:** "Next, invoke `skill-optimizer-validate-improvement` + to check the proposal independently." - **Iteration behavior:** re-run when `06-analysis.md` changed - (new analysis = potentially different weakness) or when the - operator wants a fresh optimization attempt. Fresh-derivation - for both optimizer and validator — neither subagent reads its - own canonical (`07-improvement-proposal.md` / - `07-validator-verdict.md`) or git history of it. On re-run, the - optimizer reads `improved-skill/` (if it exists from a prior - step-7 run) as the current state and proposes the next - improvement on top; the original source stays frozen. - `${OPERATOR_DIRECTIVES}` carries hints like "prefer additive - changes" or "don't touch the description field" — additional - constraints, never a context dump of prior attempts. - -### 8. `skill-optimizer-autopilot` + (new analysis = potentially different weakness), when step 8 + returned `needs-revision`/`reject` with a distillable rationale, + or when the operator wants a fresh optimization attempt. + Fresh-derivation — the optimizer doesn't read its own canonical + (`07-improvement-proposal.md`) or git history of it, and doesn't + read step 8's prior verdicts. `${OPERATOR_DIRECTIVES}` carries + the distilled lessons. + +### 8. `skill-optimizer-validate-improvement` + +- **Description trigger:** "validate the improvement", "check the + proposal", "is the optimizer's change sound", "review the fix" +- **Input:** `07-improvement-proposal.md` (REQUIRED), + `01-functionality.md`, current skill state (`improved-skill/` + if it exists, else original source), optionally + `03-submissions.md` +- **Output:** `08-validator-verdict.md` (the verdict + + rationale). On `verdict: approve`, ALSO materializes + `improved-skill/` by applying the optimizer's diff to a copy of + the current state. **The original source is never modified** — + vendored copies stay frozen, local source files stay untouched. + `improved-skill/` accumulates across approved iterations; git + tracks its history. +- **Behavior:** + 1. Dispatch **validator subagent** with limited context. + - **Internal consistency check:** does the proposed change + make sense given the skill's stated responsibilities in + `01-functionality.md`? Additive vs. destructive? General + vs. ducktape? + - **External consistency check (only if `03-submissions.md` + exists):** does the change conform to upstream PR rules + (frontmatter, file location, prefix taxonomy, + additive-only, etc.)? Forward-looking — verifies the + improved skill COULD be turned into a valid PR, even + though neither step 7 nor step 8 produces one. + - Verdict: `approve` / `needs-revision` / `reject`. + 2. On `verdict: approve`: materialize `improved-skill/` by + applying the diff to a temp copy of the current state, then + atomically place at the canonical path. Commit. + 3. On `verdict: needs-revision` or `reject`: do NOT + materialize. Surface honestly per the handoff branch. +- **Dispatches:** validator subagent, limited context, isolated + from raw trial data + the optimizer's reasoning trace +- **Single-shot per invocation:** no in-step revision loop. Each + invocation produces one verdict; the 7→8→7 cycle on + `needs-revision` is operator-driven (or auto-pilot-driven at + step 9). +- **Handoff (three branches by verdict):** + - **approve:** "Validation complete; improvement applied at + `improved-skill/`. The chain has reached its natural endpoint + for this iteration. Review, copy locally, hand to a PR + composer, or re-bench." + - **needs-revision:** "Validator says needs-revision. Distill + the rationale into a directive and re-invoke step 7, then + re-invoke this step." + - **reject:** "Validator rejects the proposal outright. Two + paths: (1) re-invoke step 6 with a reframed weakness; (2) + accept that this weakness isn't addressable and exit + honestly." +- **Iteration behavior:** re-run when `07-improvement-proposal.md` + changed (step 7 produced a new proposal). Fresh-derivation — + the validator doesn't read its own canonical + (`08-validator-verdict.md`) or git history of it. + Independence-from-self is load-bearing: if the validator saw + its prior verdict it would gravitate toward + consistency-with-itself across re-validation cycles, defeating + the point of re-running. + +### 9. `skill-optimizer-autopilot` - **Description trigger:** "auto-pilot this skill", "run the whole chain on X", "skill-optimizer end-to-end for X", "automated @@ -614,7 +654,7 @@ Per-step iteration behavior is noted at the end of each subsection. `docs/skill-optimizer//autopilot-summary-.md` listing each step's final version, headline result, and any blockers - **Behavior:** - 1. Walks 1→7 in order, dispatching each chain skill. + 1. Walks 1→8 in order, dispatching each chain skill. 2. For each step: check whether the existing artifact is current via `git log` mtime comparison against direct upstream. If current, skip; if missing or stale, dispatch. @@ -623,18 +663,22 @@ Per-step iteration behavior is noted at the end of each subsection. `pr_intent` flag value (default `false`) - **B2 user-picks gate** → sets `picked: true` on top-N functionalities by importance (default `pick_top_n: 5`) - - **B7 validator-rejected verdict** → logs the final state, no - further automated retries; surfaces blocker in the summary + - **B8 validator-rejected verdict** → on `needs-revision`, + distills the rationale into a directive and re-invokes + step 7 then step 8 (up to `max_iterations_per_step` rounds). + On `reject`, logs the final state and surfaces blocker in + the summary; does not retry automatically. 4. Bounds iterations: at most `max_iterations_per_step` re-runs - per step (default 2) to cap runaway loops. + per step (default 2) to cap runaway loops. This is also what + caps the 7→8→7 revision loop on `needs-revision` verdicts. 5. Surfaces blockers (subagent BLOCKED status, validator unresolvable, missing inputs) in the summary rather than halting the whole run. -- **Dispatches:** the seven chain skills as ordinary subagents (each - chain skill internally dispatches its own narrow-context +- **Dispatches:** the eight chain skills as ordinary subagents + (each chain skill internally dispatches its own narrow-context subagents). - **Caveats** (baked into SKILL.md): expect modest results compared - to operator-driven runs. The seven steps are hard even with human + to operator-driven runs. The eight steps are hard even with human judgment; auto-pilot is best for batch processing where some failures are acceptable, not for high-stakes single-target optimization. diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md index 746445d..0d48631 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -33,7 +33,7 @@ Two load-bearing pieces of context to load NOW, before the workflow: in "Why limited-context dispatch matters" below — read it if you're tempted to skip the dispatch. -Throughout this document, "step 1" through "step 7" (no parens) +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 6; `run-bench` is step 5; `improve-skill` is step 7). Internal workflow steps within THIS skill are labelled "(a)" through "(e)". @@ -177,7 +177,7 @@ The subagent does NOT see: - Its own prior `06-analysis.md` (anti-ducktape — must not be biased by prior analyses) - Git history of `06-analysis.md` -- `07-improvement-proposal.md`, `07-validator-verdict.md`, or any +- `07-improvement-proposal.md`, `08-validator-verdict.md`, or any prior optimizer attempts (would also bias toward "weaknesses-that-could-be-patched-the-way-the-optimizer-tried") diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md index 9d6f414..8f054e1 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -5,131 +5,86 @@ description: Use when the user wants to improve a skill based on identified stru # skill-optimizer-improve-skill -Step 7 of the skill-optimizer chain — the terminal generative step. -Takes the named structural weaknesses from step 6, dispatches an -optimizer subagent to draft a principled fix, then dispatches a -validator subagent to independently check whether the fix is sound -(internal consistency) and conformant (external PR conventions if -PR-bound). Writes the improved skill to a separate location -(`improved-skill/`), the improvement proposal, and the validator's -verdict. **The original skill is never modified** — vendored -upstream copies stay frozen as reference; local source files stay -untouched. The operator reviews `improved-skill/` and decides what -to do with it (apply locally, hand to a PR composer, discard). +Step 7 of the skill-optimizer chain. Takes the named structural +weaknesses from step 6, dispatches an optimizer subagent to draft +a principled fix, and writes `07-improvement-proposal.md`. The +proposal is then validated independently by step 8 +(`validate-improvement`), which materializes the improved skill on +approve. **This step does not produce the improved skill itself** +— it produces the proposal that step 8 acts on. This skill **refuses to fire** if `06-analysis.md` has `has_structural_weakness: false` — there's nothing to optimize, and forcing a fix in that situation is the ducktape failure mode the chain is built to prevent. -**Out of scope:** packaging the change as a PR draft. The chain -treats PR composition as a separate concern that consumes B7's -rationale + `03-submissions.md`'s upstream facts; the auto-pilot -(step 8) or a dedicated composer can handle it downstream of this -step if `pr_submission_intent: true`. B7 just improves the skill. +**Single-shot per invocation.** There is no in-step revision loop; +each invocation produces one proposal. If step 8's validator returns +`needs-revision`, the operator (or auto-pilot at step 9) distills +the validator's rationale into a directive and re-invokes this +step. That keeps chain skills from auto-invoking each other. + +**Out of scope:** validating the proposal (that's step 8) and +packaging the change as a PR draft (PR composition is a separate +downstream concern; auto-pilot or a dedicated composer can handle it +if `pr_submission_intent: true`). ## Before you start -Three load-bearing pieces of context to load NOW, before the +Two load-bearing pieces of context to load NOW, before the workflow: 1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** Every chain skill applies it on every invocation. Step 7 is a - **fresh-derivation** step (both the optimizer and validator are - fresh-derivation subagents); the iteration protocol's - "frontmatter discipline" + "directive channel" sections both - apply. Pay attention to the **step-7 carve-out** in the - subagent-constraints section: the SKILL CONTENT is upstream - input — `improved-skill/` if it exists (the accumulated - improvement state from prior step-7 runs), else the original - source (`vendored-skill/` for upstream, the user's local file - for local). The optimizer's own canonical - (`07-improvement-proposal.md`) and the validator's own - canonical (`07-validator-verdict.md`) are both off-limits to - their respective subagents — those are the reports, not the - skill being improved. - -2. **You will dispatch two subagents (optimizer + validator); you - do NOT propose diffs or judge them yourself in this session.** - Workflow steps (c) and (d) are the dispatches. The rationale is - in "Why limited-context dispatch matters" below. - -3. **The optimizer/validator loop is bounded to 2 revision - rounds.** If the validator still says `needs-revision` after - round 2, surface honestly and exit — do not loop indefinitely. - This bound protects against optimizer/validator pathological - disagreement burning operator budget. - -Throughout this document, "step 1" through "step 7" (no parens) + **fresh-derivation** step. Pay attention to the **step-7 + carve-out** in the subagent-constraints section: the SKILL + CONTENT is upstream input — `improved-skill/` if it exists + (the accumulated improvement state from prior approved-by-step-8 + runs), else the original source (`vendored-skill/` for upstream, + the user's local file for local). The optimizer's own canonical + (`07-improvement-proposal.md`) is off-limits — that's the report + describing the optimizer's reasoning, not the skill being + improved. + +2. **You will dispatch a subagent for the actual optimization; you + do NOT propose diffs yourself in this session.** Workflow step + (c) is the dispatch. The rationale is in "Why limited-context + dispatch matters" below. + +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 7; -`analyze-result` is step 6). Internal workflow steps within THIS -skill are labelled "(a)" through "(g)". +`analyze-result` is step 6; `validate-improvement` is step 8; +`autopilot` is step 9). Internal workflow steps within THIS skill +are labelled "(a)" through "(e)". ## What you produce -Two or three artifacts at `docs/skill-optimizer//`, depending -on the verdict: - -1. **`07-improvement-proposal.md`** — the optimizer's proposed - diff + rationale. Frontmatter (runtime-relevant facts only, - per the iteration protocol's frontmatter discipline): - - ```yaml - --- - addresses_weaknesses: - - - - ... - --- - ``` - - The body lists the proposed change, the rationale referencing - the named weakness from `06-analysis.md` and the "What WOULD - address this" principle being applied, and a self-check against - the "What WOULD NOT address this" anti-pattern list. Subagent - prompt template specifies the section shape. - -2. **`07-validator-verdict.md`** — the validator's independent - judgment. Frontmatter: - - ```yaml - --- - verdict: approve | needs-revision | reject - round: - addresses_weaknesses: - - - - ... - --- - ``` - - The body covers the internal consistency check (does the - change make sense for the named weakness? additive vs. - destructive? general vs. ducktape?), and — if - `03-submissions.md` exists — the external consistency check - (does the change conform to upstream PR rules: frontmatter, - file location, prefix taxonomy, additive-only, etc.). The - external check is forward-looking (would a PR carrying this - change be conformant?) even though B7 itself does not produce - a PR. - -3. **`improved-skill/`** — the improved skill content (only if - `verdict: approve`). Mirrors the source skill's directory - structure with the optimizer's diff applied. **The original - is never modified**: for upstream skills, `vendored-skill/` - stays frozen as the upstream reference; for local skills, the - user's source file stays untouched. The operator reviews - `improved-skill/` and decides what to do next — apply locally - by copying over their file, hand to a PR composer (auto-pilot - or operator-driven) for upstream submission, or discard. Git - tracks the history of `improved-skill/` across iterations. - -For the body templates of the optimizer and validator reports + -their reasoning protocols, see -[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) -and -[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). -Each subagent writes its own body; the operator session orchestrates -the loop and materializes `improved-skill/` by applying the -optimizer's diff. +One artifact at `docs/skill-optimizer//`: + +**`07-improvement-proposal.md`** — the optimizer's proposed +change + rationale. Frontmatter (runtime-relevant facts only, per +the iteration protocol's frontmatter discipline): + +```yaml +--- +addresses_weaknesses: + - + - ... +--- +``` + +The body lists the proposed change, the rationale referencing the +named weakness from `06-analysis.md` and the "What WOULD address +this" principle being applied, and a self-check against the "What +WOULD NOT address this" anti-pattern list (the optimizer states +explicitly why its proposal is NOT one of the ducktape moves the +analyzer flagged). The subagent prompt template specifies the +section shape. + +For the body template + the optimizer's reasoning protocol, see +[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md). +The subagent writes its body itself; the operator session only +orchestrates the dispatch. ## Workflow @@ -155,41 +110,40 @@ Three checks, in order: reason this gate exists. 3. `01-functionality.md` must exist (the optimizer reads it as - input). The source skill must be present (vendored at - `vendored-skill/` for upstream skills, or accessible at the - local path recorded in `01-functionality.md`'s `skill_source` - field). + input). The skill content must be readable: `improved-skill/` + if it exists (the accumulated state from prior step-8 + approvals), else the original source (`vendored-skill/` for + upstream skills, the local file path recorded in + `01-functionality.md`'s `skill_source` for local skills). If `pr_submission_intent: true` from `01-functionality.md`, -`03-submissions.md` should also exist; if it's missing, the -validator's external consistency check can't run. Surface this and -ask the user whether to run step 3 first OR proceed with internal -consistency only. +`03-submissions.md` should also exist (the optimizer reads it to +shape the diff to match upstream conventions from the start). If +it's missing, surface this and ask the user whether to run step +3 first OR proceed without upstream-shape hints (step 8's +external check will still run if `03-submissions.md` appears +later). ### (b) Handle iteration **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it. Step 7 is a **fresh-derivation** step for both -subagents: - -- Each invocation derives a fresh proposal + verdict from - `06-analysis.md` + the current skill content + directives. - Neither subagent reads its own prior canonical - (`07-improvement-proposal.md` for the optimizer; - `07-validator-verdict.md` for the validator) or its git - history. -- The SKILL itself IS upstream input — the optimizer reads the - current skill content, which may include modifications from - prior step-7 runs. That's intentional; the skill accumulates - improvements across iterations, but the proposal reports - describing each round do not. -- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The operator - CAN read the prior `07-improvement-proposal.md` and - `07-validator-verdict.md`, distill any lessons into atomic new - requirements, and pass them to the optimizer. Examples: +and apply it. Step 7 is a **fresh-derivation** step: + +- Each invocation derives a fresh proposal from `06-analysis.md`, + the current skill content, and any operator directives. The + optimizer does NOT read its own prior + `07-improvement-proposal.md` or its git history. +- If `07-improvement-proposal.md` already exists, it will be + overwritten by this invocation. Prior state lives in git + history. No checkpoint commit needed (fresh-derivation; prior + state is already self-contained in its own prior commit). +- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The + operator CAN read prior proposals and step 8's prior verdicts + (which they need to distill anyway, after a `needs-revision` + verdict). Examples: - "prefer additive changes over destructive ones" - - "don't touch the description field — the validator rejected - that last round" + - "don't touch the description field — validator step 8 + rejected that last round" - "address weakness 2 first; weakness 1 was already partially addressed by the prior round" @@ -204,7 +158,7 @@ isolation if your environment supports it). Load the prompt template at [`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) and substitute the templated inputs (`${ANALYSIS_PATH}`, -`${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, +`${FUNCTIONALITY_PATH}`, `${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, `${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`). @@ -216,15 +170,12 @@ The optimizer sees: optimizer understands the skill's stated responsibilities and doesn't propose a change that contradicts them) - `${SKILL_CURRENT_PATH}` — the current state of the skill being - improved: `improved-skill/` if it exists (the accumulated state - from prior step-7 runs), else the original source - (`vendored-skill/` for upstream skills, the local file path - recorded in `01-functionality.md`'s `skill_source` for local - skills). The optimizer proposes a new improvement on top of + improved: `improved-skill/` if it exists, else the original + source. The optimizer proposes a new improvement on top of whatever current state it sees. - `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (so the optimizer can shape the diff to match upstream conventions - from the start, reducing validator round-trips) + from the start, reducing step-8 round-trips) - `${OPERATOR_DIRECTIVES}` — atomic new requirements (may be empty) @@ -235,184 +186,78 @@ The optimizer does NOT see: - Grader internals (`tests///grader.mjs` source) - Test inputs (`tests///workspace/`) - Its own prior `07-improvement-proposal.md` or git history of it -- Prior `07-validator-verdict.md` (the validator's prior judgment - would bias the optimizer toward defending or pivoting away from - the prior attempt rather than addressing the weakness fresh) - -The "no raw trials / grader internals / test inputs" constraint is -load-bearing: it forces the optimizer to address the weakness as -the analyzer named it, in terms of the general principle the +- Step 8's prior `08-validator-verdict.md` (the validator's prior + judgment would bias the optimizer toward defending or pivoting + away from the prior attempt rather than addressing the weakness + fresh; lessons from the prior verdict come through distilled + directives instead) + +The "no raw trials / grader internals / test inputs" constraint +is load-bearing: it forces the optimizer to address the weakness +as the analyzer named it, in terms of the general principle the analyzer articulated — not by pattern-matching a patch that would make specific failing trials pass. The optimizer writes `07-improvement-proposal.md` (the diff + rationale, including the explicit self-check against the anti-pattern list) and returns a brief summary: the weakness(es) -addressed, the principle applied, the lines/sections of the skill -modified. - -### (d) Dispatch the validator subagent - -After the optimizer returns, dispatch the validator subagent. Load -the prompt template at -[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) -and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (same as -the optimizer's `${SKILL_CURRENT_PATH}` from (c) — `improved-skill/` -if it exists, else the source), `${SKILL_AFTER_PATH}` (the proposed -result — materialize by applying the optimizer's diff to a -temporary copy of BEFORE; do NOT overwrite the canonical -`improved-skill/` yet — that happens at (f) after approve), -`${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, -`${VERDICT_OUTPUT_PATH}`, `${ROUND}` (1 on first dispatch; 2 on -revision). - -The validator sees: - -- The skill BEFORE the change (same state the optimizer read) -- The skill AFTER the change (a temporary materialization; the - canonical `improved-skill/` is not yet updated) -- `${PROPOSAL_PATH}` — `07-improvement-proposal.md` (so it knows - what the optimizer claims to have done) -- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the - internal consistency check: does the change make sense given - the skill's stated responsibilities?) -- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for - the external consistency check) - -The validator does NOT see: - -- Raw failed trials, `findings.txt`, `trace.jsonl` -- Test inputs (`tests//workspace/`) -- The optimizer's internal reasoning trace (only the proposal - artifact, not how the optimizer arrived at it) -- Its own prior `07-validator-verdict.md` or git history of it -- `06-analysis.md` directly (the analyzer's reasoning is mediated - through the optimizer's proposal — the validator's job is to - check the proposal as an independent observer, not to second- - guess the analysis) - -The "no prior verdict" constraint is load-bearing: each validator -round is independent. If the validator saw its prior verdict, it -would gravitate toward consistency-with-itself, defeating the -purpose of revision rounds. - -The validator writes `07-validator-verdict.md` with a verdict of -`approve`, `needs-revision`, or `reject`, plus rationale. - -### (e) Handle the verdict (bounded revision loop) - -Three branches: - -**If `verdict: approve`:** proceed to (f). - -**If `verdict: needs-revision` AND `round < 2`:** re-dispatch the -optimizer with the validator's rationale as a directive. Example -directive: `validator round 1 said: '' — address this and -re-propose`. Increment round to 2. Loop back to (c) for the -re-dispatch; the optimizer derives fresh per its rules (it does -NOT see the prior proposal directly, only the distilled directive). -Then re-dispatch the validator at (d) with `${ROUND}: 2`. - -**If `verdict: needs-revision` AND `round == 2`:** the bound is -reached. Surface honestly: - -> After 2 revision rounds, the validator still says -> needs-revision: ``. The optimizer and validator have not -> converged on this weakness within the allowed bound. Three -> realistic paths: (1) re-invoke step 6 with a directive to -> reformulate the weakness ("the current naming was too abstract; -> split it"); (2) re-invoke step 7 manually with directives that -> bridge the disagreement; (3) accept that this weakness is not -> fixable in this round and exit honestly. - -Do NOT auto-extend the loop. The bound exists to prevent -pathological disagreement burning operator budget. - -**If `verdict: reject`:** surface honestly and exit. Reject means -the validator judges the proposal is fundamentally wrong (not just -in need of revision). The reasonable next steps are to re-invoke -step 6 (the analyzer's framing of the weakness may have been -misleading) or to accept that the named weakness isn't addressable -without changes the validator is unwilling to approve. - -### (f) Write outputs - -After `verdict: approve`: - -1. Materialize `improved-skill/` by copying the source skill's - directory structure and applying the optimizer's diff. **Do - NOT modify the source.** For upstream skills, leave - `vendored-skill/` frozen and write to a sibling - `improved-skill/` directory; for local skills, copy the user's - source file(s) to `improved-skill/` and apply the diff there. - -2. Confirm `07-improvement-proposal.md` and - `07-validator-verdict.md` are written to canonical paths. - -3. Commit the change: - - ```bash - git add docs/skill-optimizer// - git commit -m "step 7: improve — address " - ``` - - Note the source skill file is NOT in the commit — git tracks - `improved-skill/` history alongside the proposal and verdict - reports, but never touches the user's source or the vendored - reference. - -### (g) Hand off - -Single handoff message — no branching on PR intent. PR packaging -is a separate downstream concern (auto-pilot or operator-driven); -this step's job is done when the improved skill is materialized -and verified. - -> Improvement complete. The improved skill is at -> `docs/skill-optimizer//improved-skill/`; the original -> source (vendored or local) is untouched. The proposal and -> validator verdict are in `07-improvement-proposal.md` and -> `07-validator-verdict.md`. Three realistic next steps: (1) -> review `improved-skill/` and copy it over your local source if -> you're satisfied; (2) if PR-bound, hand the improved skill + -> the proposal + `03-submissions.md` to a PR composer (auto-pilot -> can do this end-to-end if you're running step 8); (3) re-bench -> against the improved skill by re-invoking -> `skill-optimizer-run-bench` against an updated workbench that -> points at `improved-skill/`. - -The chain stops here. Whether to apply the improvement, submit it -upstream, or iterate further is an explicit operator decision. +addressed, the principle applied, the lines/sections of the +skill modified. + +### (d) Confirm subagent output + +Verify the report file exists, the frontmatter parses, and the +body has: + +- A clear proposed change (in fenced unified-diff format or + similar) +- A rationale section referencing at least one named weakness + from `06-analysis.md` and the "What WOULD address this" + principle it applies +- An explicit self-check against the "What WOULD NOT address + this" anti-pattern list — the optimizer must state why its + proposal is NOT one of the ducktape moves the analyzer + flagged + +If the self-check section is missing or vague, surface to the +user — this is the architectural anti-ducktape signal the +optimizer is required to produce. Re-dispatch with a directive +"the self-check against the anti-pattern list is required and +missing or vague — make it explicit" rather than filling it in +yourself. + +### (e) Hand off + +Single handoff message: + +> Improvement proposal complete. The proposal is at +> `docs/skill-optimizer//07-improvement-proposal.md`. Next, +> invoke `skill-optimizer-validate-improvement` to check the +> proposal independently. If the validator returns +> `needs-revision` or `reject`, you'll come back here with a +> directive distilling the validator's concerns. + +Don't auto-invoke step 8 — the chain skills never invoke each +other on their own (per the iteration protocol's re-run +authorization rule). ## Why limited-context dispatch matters -Two distinct concerns, one for each subagent: - -**For the optimizer:** the chain's whole anti-ducktape architecture -hinges on the optimizer NOT seeing raw failures. If the optimizer -read the failed trials, it would pattern-match a patch that makes -those specific trials pass — the textbook ducktape failure mode. -The analyzer's job (step 6) is to translate raw failures into -**named structural weaknesses with general principles and -anti-patterns**; the optimizer's job is to apply the principle. -The translation through the analyzer's report is what enforces -principled improvement over pattern-match patching. - -**For the validator:** independence from the optimizer's reasoning -is load-bearing. The validator must judge the proposal as if seeing -it for the first time, against the skill's stated responsibilities -(internal) and the upstream's PR conventions (external). If the -validator saw the optimizer's reasoning trace, it would tend to -accept arguments the optimizer made about why the change is sound -— defeating the purpose of independent checking. The validator -sees the proposal artifact (what was changed and the optimizer's -brief rationale) but NOT the optimizer's internal reasoning trace. - -Both subagents also don't read their own prior canonicals: the -optimizer doesn't see prior proposals (would gravitate toward -defending or pivoting away from prior attempts), the validator -doesn't see prior verdicts (would gravitate toward -consistency-with-itself across rounds). Each invocation is +The chain's whole anti-ducktape architecture hinges on the +optimizer NOT seeing raw failures. If the optimizer read the +failed trials, it would pattern-match a patch that makes those +specific trials pass — the textbook ducktape failure mode. The +analyzer's job (step 6) is to translate raw failures into **named +structural weaknesses with general principles and anti-patterns**; +the optimizer's job is to apply the principle. The translation +through the analyzer's report is what enforces principled +improvement over pattern-match patching. + +The optimizer also doesn't read its own prior canonical or step +8's prior verdicts. Either would bias toward defending or +pivoting away from prior attempts rather than addressing the +weakness fresh; lessons from prior runs come through the +operator's distilled directives. Each invocation of step 7 is genuinely fresh. If you find yourself thinking "I'll just propose the change @@ -423,60 +268,50 @@ this chain is built to prevent. Dispatch. - **`06-analysis.md` missing** — caught at (a). Tell user to run step 6 first. -- **`has_structural_weakness: false`** — caught at (a). Refuse to - fire; this is the anti-ducktape gate. Don't try to override. +- **`has_structural_weakness: false`** — caught at (a). Refuse + to fire; this is the anti-ducktape gate. Don't try to override. - **Optimizer reports BLOCKED** (e.g., the analyzer's named weakness is too abstract to derive a concrete diff from) — - surface to the user. The fix is typically a step 6 re-run with a - directive ("weakness X as named was too abstract; split into + surface to the user. The fix is typically a step 6 re-run with + a directive ("weakness X as named was too abstract; split into concrete sub-weaknesses"). Per the iteration protocol's re-run authorization rule, the user invokes step 6. - **Optimizer's self-check against anti-patterns is missing or - vague** — the optimizer's report MUST explicitly state why its - proposal is not one of the ducktape moves the analyzer flagged. - If this section is missing, treat as `needs-revision` (route - through (e)) with the directive "the self-check against the - anti-pattern list is required and missing". -- **Validator rejects with reasoning that contradicts the - analyzer's framing** (e.g., validator says "this change doesn't - address the real problem"; the analyzer thought it did) — that's - a signal the analyzer/validator are misaligned on what the - weakness actually is. Surface to the user with two options: - re-invoke step 6 to reformulate the weakness, or re-invoke step - 7 with a directive that bridges the framing. + vague** — caught at (d). Re-dispatch with a directive making + the self-check requirement explicit. - **PR-bound but `03-submissions.md` missing** — caught at (a). - Ask the user whether to run step 3 first or proceed with - internal consistency only. Proceeding loses the validator's - external consistency check; the improved skill will still be - produced, but any downstream PR composer won't have upstream - conventions to conform to. -- **Operator directives reference a specific section of the prior - proposal** (e.g., "the optimizer's section on the description - field was wrong") — that's a context-dump masquerading as a - directive. Translate it to an atomic new requirement ("don't - modify the description field") before passing to the subagent. + Proceeding without it means the optimizer doesn't get upstream + shape hints from the start; step 8's external check will + still run if `03-submissions.md` appears later (the operator + can run step 3 before step 8 if they want). +- **Operator directives reference a specific section of a prior + proposal or verdict** (e.g., "the optimizer's section on the + description field was wrong") — that's a context-dump + masquerading as a directive. Translate to an atomic new + requirement ("don't modify the description field") before + passing to the subagent. ## Iteration behavior -Step 7 is re-runnable and is a fresh-derivation step (for both -optimizer and validator). Re-run triggers specific to this step: +Step 7 is re-runnable and is a fresh-derivation step. Re-run +triggers specific to this step: - Step 6 (`analyze-result`) produced a new analysis — `06-analysis.md` changed via git mtime, possibly naming different weaknesses +- Step 8 (`validate-improvement`) returned `needs-revision` or + `reject` — the operator distills the validator's rationale + into a directive and re-invokes this step - User wants the optimizer to try a different approach (passes - directives like "this time prefer additive changes" or "focus on - weakness 2 only") -- Validator returned `reject` on a prior run and the operator has - a new framing to try + directives like "this time prefer additive changes" or "focus + on weakness 2 only") General iteration mechanics — staleness detection (git-native), operator directives (the channel for prior-derivation lessons), the step-7 carve-out for the skill being upstream input — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md). Workflow step (b) above already requires reading that file. -The internal optimizer/validator revision loop (max 2 rounds) is -not "iteration" in the chain sense — it's bounded refinement -within a single invocation of step 7. The chain-level iteration is -re-invoking step 7 entirely (which resets the round counter to 1). +There is no in-step revision loop. Each invocation produces one +proposal. The 7→8→7 cycle on `needs-revision` verdicts is +operator-driven (or auto-pilot-driven at step 9). diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 8572d06..99d7dab 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -26,7 +26,7 @@ Two load-bearing pieces of context to load NOW, before the workflow: the dispatch. The rationale is in "Why limited-context dispatch matters" below — read it if you're tempted to skip the dispatch. -Throughout this document, "step 1" through "step 7" (no parens) refer +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 1; `investigate-test-case` is step 2; etc.). Internal workflow steps within THIS skill are labelled "(a)" through "(g)" to avoid the collision. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 37a1673..a700fde 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -12,7 +12,7 @@ gather the upstream repo's contribution conventions (license, CLA, frontmatter spec, file-location rules, PR-shape patterns from recent merged + closed-without-merge PRs), and writes `docs/skill-optimizer//03-submissions.md` — the -verbatim-pastable context block the validator (step 7) uses for its +verbatim-pastable context block the validator (step 8) uses for its external consistency check. ## Before you start @@ -32,7 +32,7 @@ Two load-bearing pieces of context to load NOW, before the workflow: dispatch matters" below — read it if you're tempted to skip the dispatch. -Throughout this document, "step 1" through "step 7" (no parens) refer +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 3; `investigate-functionality` is step 1; etc.). Internal workflow steps within THIS skill are labelled "(a)" through "(e)" to avoid the collision. @@ -141,7 +141,7 @@ The subagent does NOT see: The constraint that the subagent doesn't see the proposed change is load-bearing: the report must be neutral upstream facts, not advocacy -for a specific change. The validator (step 7) will later check the +for a specific change. The validator (step 8) will later check the proposed change against this report — if the report is biased toward the change, the validator's external consistency check loses its independence. @@ -184,7 +184,7 @@ consumes step 3's output, not step 4. ## Why limited-context dispatch matters The submission-researcher subagent must produce a report that the -validator (step 7) can trust as independent. If the operator session +validator (step 8) can trust as independent. If the operator session does the research, it has already absorbed the optimization context (the proposed change, the prior failures, the user's framing of what "good" looks like). That context biases the research toward diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 54c48a6..b4535f4 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -29,7 +29,7 @@ Two load-bearing pieces of context to load NOW, before the workflow: is in "Why limited-context dispatch matters" below — read it if you're tempted to skip the dispatch. -Throughout this document, "step 1" through "step 7" (no parens) +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 2; step 1 is `investigate-functionality`; etc.). Internal workflow steps within THIS skill are labelled "(a)" through "(f)". @@ -138,7 +138,7 @@ The subagent does NOT see: source's literal phrasing — design coverage from STATED responsibilities) - Any `06-analysis.md`, `07-improvement-proposal.md`, - `07-validator-verdict.md` + `08-validator-verdict.md` - Raw failure data, `findings.txt`, bench results - Git history of any tree file or its own audit report diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index a6765f2..d57dc5c 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -25,7 +25,7 @@ One load-bearing piece of context to load NOW, before the workflow: write-up of the new raw run; prior state lives in git history. Workflow step (b) below requires it. -Throughout this document, "step 1" through "step 7" (no parens) refer +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 5; `write-tests` is step 4; `analyze-result` is step 6). Internal workflow steps within THIS skill are labelled "(a)" through "(e)" to avoid the collision. diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 73517cf..d352d60 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -64,6 +64,7 @@ the researcher must not rationalize a prior report. | 5. run-bench (summary file) | Mechanical write-up of the new raw bench run, not an extension of prior summary | | 6. analyze-result | Must not be biased by prior analyses; load-bearing for anti-ducktape | | 7. improve-skill | Optimizer must not see prior attempts; load-bearing for anti-ducktape | +| 8. validate-improvement | Validator must not be biased by its prior verdicts; independence-from-self is load-bearing | **Maintenance steps** manage an accumulating tree on disk — the filesystem itself is the state. Re-runs read the current tree and @@ -88,9 +89,9 @@ Three rules every reasoning subagent must follow: filesystem is the source of truth; do not walk `git log` looking for prior versions of upstream files. -2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7): do NOT - read your own canonical file, and do NOT read git history of - it.** Your job is to derive a new artifact from upstream + the +2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7, 8): do + NOT read your own canonical file, and do NOT read git history + of it.** Your job is to derive a new artifact from upstream + the operator's directives, without direct access to prior derivations of this same artifact. This is the load-bearing anti-ducktape constraint. @@ -112,20 +113,23 @@ Three rules every reasoning subagent must follow: neither you nor any subsequent fresh-derivation subagent should walk that history. - **For step 7 specifically:** the SKILL CONTENT (the target - being improved) is upstream input, not your own canonical. You - read the **current state of the skill** — which is + **For step 7 and step 8 specifically:** the SKILL CONTENT (the + target being improved) is upstream input, not your own + canonical. Both the optimizer (step 7) and the validator (step + 8) read the **current state of the skill** — which is `improved-skill/` if it exists (the accumulated state from - prior step-7 runs), else the original source + prior step-8 approvals), else the original source (`vendored-skill/` for upstream, the local file path for - local). Step 7 never modifies the original source: the - accumulated improvements live in `improved-skill/`, and the - original (vendored or local) stays frozen as reference. Your - "own canonical" that's off-limits is - `07-improvement-proposal.md` (the report describing your - reasoning), not the skill content itself. The skill content - accumulates improvements across iterations in - `improved-skill/`; the proposal reports do not. + local). Neither step modifies the original source: the + accumulated improvements live in `improved-skill/` and are + materialized by step 8 on `verdict: approve`; the original + (vendored or local) stays frozen as reference. Each subagent's + "own canonical" that's off-limits is its report + (`07-improvement-proposal.md` for the optimizer; + `08-validator-verdict.md` for the validator), not the skill + content itself. The skill content accumulates improvements + across iterations in `improved-skill/`; the proposal and + verdict reports do not. 3. **For maintenance steps (2, 4): DO read your own canonical tree** (when it exists). Your job on a re-run is to extend or diff --git a/skills/skill-optimizer-validate-improvement/SKILL.md b/skills/skill-optimizer-validate-improvement/SKILL.md new file mode 100644 index 0000000..ab3fb2f --- /dev/null +++ b/skills/skill-optimizer-validate-improvement/SKILL.md @@ -0,0 +1,337 @@ +--- +name: skill-optimizer-validate-improvement +description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve-skill` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve-skill` has produced `07-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. +--- + +# skill-optimizer-validate-improvement + +Step 8 of the skill-optimizer chain. Takes the proposal from step +7, dispatches a validator subagent to independently check whether +the proposed change is sound (internal consistency against the +skill's stated responsibilities) and conformant (external PR +conventions if PR-bound), then — on `verdict: approve` — +materializes the improved skill at `improved-skill/`. Writes +`08-validator-verdict.md` regardless of the verdict; materialization +happens only on approve. + +The validator is dispatched fresh per invocation (no in-step +revision loop). If the verdict is `needs-revision`, the user (or +auto-pilot at step 9) re-invokes step 7 with the validator's +rationale distilled into a directive, then re-invokes this step. +This honors the chain's "chain skills don't auto-invoke other +chain skills" rule and keeps both steps single-shot. + +## Before you start + +Two load-bearing pieces of context to load NOW, before the +workflow: + +1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** + Every chain skill applies it on every invocation. Step 8 is a + **fresh-derivation** step: each invocation derives a fresh + verdict from the current proposal + current skill state + + directives, with no continuity from prior verdicts. The + independence-from-the-optimizer constraint is load-bearing + here. + +2. **You will dispatch a subagent for the actual validation; you + do NOT judge the proposal yourself in this session.** Workflow + step (c) is the dispatch. The rationale is in "Why + limited-context dispatch matters" below. + +Throughout this document, "step 1" through "step 9" (no parens) +refer to skills in the chain (this skill is step 8; `improve-skill` +is step 7; `autopilot` is step 9). Internal workflow steps within +THIS skill are labelled "(a)" through "(e)". + +## What you produce + +One or two artifacts at `docs/skill-optimizer//`: + +1. **`08-validator-verdict.md`** — the validator's independent + judgment. Frontmatter (runtime-relevant facts only, per the + iteration protocol's frontmatter discipline): + + ```yaml + --- + verdict: approve | needs-revision | reject + addresses_weaknesses: + - + - ... + --- + ``` + + The body covers the internal consistency check (does the + change make sense for the named weakness? additive vs. + destructive? general vs. ducktape?) and — if + `03-submissions.md` exists — the external consistency check + (does the change conform to upstream PR rules: frontmatter, + file location, prefix taxonomy, additive-only, etc.). The + external check is forward-looking — it verifies the improved + skill COULD be turned into a valid PR, even though step 7 + itself does not produce one. + +2. **`improved-skill/`** — the improved skill content, only + materialized when `verdict: approve`. Mirrors the source + skill's directory structure with the optimizer's diff applied. + **The original is never modified**: for upstream skills, + `vendored-skill/` stays frozen; for local skills, the user's + source file stays untouched. Git tracks `improved-skill/` + history across iterations. + +The validator's body sections + reasoning protocol are specified +in +[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). +The subagent writes the verdict body itself; the operator session +materializes `improved-skill/` after approve. + +## Workflow + +### (a) Confirm prerequisites + +Three checks: + +1. `docs/skill-optimizer//07-improvement-proposal.md` must + exist with valid frontmatter. If not, tell the user to run + `skill-optimizer-improve-skill` first and stop here. + +2. The current skill state must be readable: `improved-skill/` if + it exists (the prior accumulated state), else the original + source (`vendored-skill/` for upstream skills, the local file + path recorded in `01-functionality.md`'s `skill_source` for + local skills). The validator needs this as "skill BEFORE". + +3. `01-functionality.md` must exist (the validator reads it as + input for the internal consistency check). + +If `pr_submission_intent: true` from `01-functionality.md`, +`03-submissions.md` should also exist; if it's missing, the +external consistency check can't run. Surface this and ask the +user whether to run step 3 first OR proceed with internal +consistency only (the validator's output will note the omission). + +### (b) Handle iteration + +**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +and apply it. Step 8 is a **fresh-derivation** step: + +- Each invocation derives a fresh verdict from the current + proposal + skill state + directives. The validator subagent + does NOT read its own prior `08-validator-verdict.md` or its + git history. Independence-from-self is load-bearing here: if + the validator saw its prior verdict, it would gravitate toward + consistency-with-itself across runs, defeating the purpose of + re-validation after the operator/auto-pilot revises and + re-invokes step 7. +- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The operator + CAN read the prior verdict and distill lessons into atomic new + requirements for the validator. Examples: "be stricter on + additive-vs-destructive — the prior verdict approved a change + that I think was actually destructive", "the upstream PR + conventions check missed the frontmatter `version:` field — + look for it specifically". The subagent works from the + distilled directives, never from the raw prior verdict. + +### (c) Dispatch the validator subagent + +**Do NOT judge the proposal yourself in this session.** Dispatch +the validator subagent via the `Agent` tool (with worktree +isolation if your environment supports it). Load the prompt +template at +[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) +and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (the +current state: `improved-skill/` if it exists, else the original +source), `${SKILL_AFTER_PATH}` (the proposed result — materialize +by applying the optimizer's diff to a temporary copy of BEFORE; +do NOT overwrite or rename anything canonical at this stage), +`${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, +`${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. + +The validator sees: + +- The skill BEFORE the change (the current state the optimizer + read in step 7) +- The skill AFTER the change (a temporary materialization; the + canonical `improved-skill/` is not updated yet — that happens + at (d) on approve) +- `${PROPOSAL_PATH}` — `07-improvement-proposal.md` (so it + knows what the optimizer claims to have done; the validator + checks the artifact, not the optimizer's reasoning trace) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the + internal consistency check) +- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for + the external consistency check) +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (may be + empty) + +The validator does NOT see: + +- Raw failed trials, `findings.txt`, `trace.jsonl` +- Test inputs (`tests//workspace/`) +- The optimizer's internal reasoning trace (only the proposal + artifact, not how the optimizer arrived at it) +- Its own prior `08-validator-verdict.md` or git history of it + (independence-from-self constraint) +- `06-analysis.md` directly (the analyzer's reasoning is + mediated through the optimizer's proposal — the validator's + job is to check the proposal as an independent observer, not + to second-guess the analysis) + +The validator writes `08-validator-verdict.md` with a verdict of +`approve`, `needs-revision`, or `reject`, plus rationale. + +### (d) Handle the verdict + materialize on approve + +Read the verdict from `08-validator-verdict.md`. + +**If `verdict: approve`:** + +1. Materialize `improved-skill/` by copying the source skill's + directory structure and applying the optimizer's diff. **Do + NOT modify the source.** For upstream skills, leave + `vendored-skill/` frozen and write to a sibling + `improved-skill/`; for local skills, copy the user's source + file(s) to `improved-skill/` and apply the diff there. + +2. Commit the new state: + + ```bash + git add docs/skill-optimizer// + git commit -m "step 8: validate + apply improvement for — " + ``` + + The source skill file is NOT in the commit — git tracks + `improved-skill/` alongside the verdict report, but never + touches the user's source or the vendored reference. + +3. Proceed to (e). + +**If `verdict: needs-revision`:** do NOT materialize +`improved-skill/`. The proposal needs to be revised before the +chain advances. Proceed to (e) with the appropriate handoff. + +**If `verdict: reject`:** do NOT materialize. The validator +judges the proposal as fundamentally wrong (not just in need of +revision). Proceed to (e) with the appropriate handoff. + +### (e) Hand off + +Three handoff messages depending on the verdict: + +**If `verdict: approve`:** + +> Validation complete; improvement applied. The improved skill is +> at `docs/skill-optimizer//improved-skill/`; the original +> source (vendored or local) is untouched. The chain has reached +> its natural endpoint for this iteration. Three realistic next +> steps: (1) review `improved-skill/` and copy it over your +> local source if you're satisfied; (2) if PR-bound, hand the +> improved skill + the proposal + `03-submissions.md` to a PR +> composer (auto-pilot can do this end-to-end if you're running +> step 9); (3) re-bench against the improved skill by re-invoking +> `skill-optimizer-run-bench` against a workbench that points at +> `improved-skill/`. + +**If `verdict: needs-revision`:** + +> Validation rejected as needs-revision. The validator's +> rationale is in `08-validator-verdict.md`. To address it: +> re-invoke `skill-optimizer-improve-skill` with a directive +> distilling the validator's concern (e.g., `validator said: +> '' — address this and re-propose`), then re-invoke +> `skill-optimizer-validate-improvement`. The operator owns +> distillation; chain skills don't auto-invoke each other. + +**If `verdict: reject`:** + +> Validation rejected the proposal outright. The validator's +> rationale is in `08-validator-verdict.md`. Two realistic paths: +> (1) re-invoke `skill-optimizer-analyze-result` with a directive +> reframing the weakness — the analyzer's framing may have been +> misleading; (2) accept that this weakness isn't addressable +> without changes the validator is unwilling to approve, and +> exit honestly. + +Don't auto-invoke step 7 or step 6 — surface the choice and let +the user (or auto-pilot at step 9) act. + +## Why limited-context dispatch matters + +Independence from the optimizer's reasoning is the validator's +load-bearing property. The validator must judge the proposal as +if seeing it for the first time, against the skill's stated +responsibilities (internal) and the upstream's PR conventions +(external). Two specific risks the limited context addresses: + +First, **biased acceptance of the optimizer's reasoning**. If the +validator saw the optimizer's reasoning trace, it would tend to +accept arguments the optimizer made about why the change is sound +— defeating the purpose of independent checking. The validator +sees the proposal artifact (what was changed and the optimizer's +brief rationale linking to the named weakness) but NOT the +optimizer's internal reasoning trace. + +Second, **consistency-with-self across runs**. If the validator +saw its prior verdict, it would gravitate toward +consistency-with-itself when re-validating after a revision — +either re-approving what it approved before or re-rejecting what +it rejected before, instead of judging the new proposal on its +own merits. Each invocation is genuinely fresh; the operator's +distilled directives carry forward what the prior verdict taught. + +If you find yourself thinking "I'll just judge the proposal +myself, I can see whether it addresses the weakness" — that's the +failure mode this chain is built to prevent. Dispatch. + +## Edge cases + +- **`07-improvement-proposal.md` missing** — caught at (a). Tell + user to run step 7 first. +- **Current skill state ambiguous** (both `improved-skill/` and + source modified externally) — surface to the user; the BEFORE + the validator reads must match what the optimizer read in step + 7. If the user edited the source between step 7 and step 8, + re-invoke step 7 to derive a fresh proposal against the new + state. +- **PR-bound but `03-submissions.md` missing** — caught at (a). + Ask the user whether to run step 3 first or proceed with + internal consistency only (the validator's verdict will note + the omission). +- **Validator rejects with reasoning that contradicts the + analyzer's framing** (e.g., validator says "this change + doesn't address the real problem"; the analyzer thought it + did) — that's a signal the analyzer/validator are misaligned + on what the weakness actually is. Surface to the user with the + two paths from the `reject` handoff: re-invoke step 6 to + reformulate the weakness, or accept. +- **Operator directives reference a specific section of the prior + verdict** (e.g., "the prior verdict's external check section + missed X") — that's a context-dump masquerading as a directive. + Translate to an atomic new requirement ("look for X in the + external check") before passing to the subagent. + +## Iteration behavior + +Step 8 is re-runnable and is a fresh-derivation step. Re-run +triggers specific to this step: + +- Step 7 (`improve-skill`) produced a new proposal — + `07-improvement-proposal.md` changed via git mtime +- User disagrees with the prior verdict and wants a fresh + derivation with directives reflecting the disagreement (e.g., + "be stricter on additive-vs-destructive") +- `03-submissions.md` was updated (upstream PR conventions + changed) and the prior verdict's external check is now stale + +General iteration mechanics — staleness detection (git-native), +operator directives (the channel for prior-derivation lessons), +cascading staleness — live in +[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). +Workflow step (b) above already requires reading that file. + +There is no in-step bounded revision loop. Each invocation of +step 8 is single-shot; if `verdict: needs-revision` or +`verdict: reject`, the operator (or auto-pilot at step 9) +re-invokes step 7 then step 8 manually. This honors the chain +skills' "no auto-invocation of other chain skills" rule and +keeps both steps semantically clean. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 88cb560..6f73162 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -38,7 +38,7 @@ workflow: test); source access is needed for concrete violation patterns and realistic fixture content. -Throughout this document, "step 1" through "step 7" (no parens) +Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain (this skill is step 4; step 2 is `investigate-test-case`; etc.). Internal workflow steps within THIS skill are labelled "(a)" through "(f)". From e737fc1011e727a74f64ece57ed1318e0361f01a Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 08:24:13 -0500 Subject: [PATCH 034/121] refactor(v1.4-chain): verbosity sweep + split shared docs (-930 lines) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per-skill verbosity audit identified four trim patterns: 1. "Before you start" preambles duplicated workflow step (b)/(c) content 2. "Why limited-context dispatch matters" sections lived far from the dispatch they explain; philosophy doc says "give a why-this-matters paragraph nearby" 3. "Confirm subagent output" boilerplate restated across all 7 subagent-dispatching skills 4. Verbatim multi-line user-dialogue and handoff templates were over-prescriptive (operator can phrase the exact words from the intent statement) Applied to all 8 chain skills: - Dropped "Before you start" sections entirely (~160L saved) - Inlined dispatch rationale at workflow step (c)/(d) as a short "Why this matters" paragraph (~170L saved net) - Tightened "Confirm subagent output" steps to one or two sentences (~50L) - Collapsed verbatim dialogue blocks to intent statements (~100L) - Examples lists trimmed from 4-5 to 2 (one to establish, one to show variation) Also split the shared iteration-protocol into three named docs: - iteration-protocol.md (~130L, was ~284L) — iteration mechanics only: step kinds, staleness, destructive-edit checkpoints, cascading, bootstrapping. Loaded at "Handle iteration" step. - subagent-dispatch.md (new, ~120L) — subagent constraints, operator-directives concept, templated dispatch inputs, no-auto-invocation rule. Loaded at "Dispatch subagent" step. - frontmatter-discipline.md (new, ~40L) — runtime facts vs history rule, decision aid. Referenced at "What you produce" section. Each chain skill loads only what it needs at the workflow step that needs it (lazy loading rather than front-loading everything at "Before you start"). Most skills need all three; B5 (no subagent dispatch) needs only iteration-protocol + frontmatter-discipline. Final line counts (all chain skills now under 200 lines): B1 investigate-functionality 227 -> 148 (-79) B2 investigate-test-case 296 -> 191 (-105) B3 investigate-submissions 240 -> 156 (-84) B4 write-tests 368 -> 200 (-168) B5 run-bench 233 -> 154 (-79) B6 analyze-result 302 -> 165 (-137) B7 improve-skill 317 -> 181 (-136) B8 validate-improvement 337 -> 189 (-148) Shared iteration-protocol 284 -> 130 + 120 + 40 = 290 Total: 2604 -> 1674 lines (-930, ~36% reduction). Aligns with the project's docs/skill-writing-philosophy.md: "Bias toward 'Claude is smart' — pruning beats adding. If the skill restates what Claude already knows, removing the restatement is often a more principled fix than adding new rules." --- .../skill-optimizer-analyze-result/SKILL.md | 347 +++++---------- skills/skill-optimizer-improve-skill/SKILL.md | 386 ++++++----------- .../SKILL.md | 233 ++++------ .../SKILL.md | 272 ++++-------- .../SKILL.md | 327 +++++--------- skills/skill-optimizer-run-bench/SKILL.md | 257 ++++------- .../frontmatter-discipline.md | 39 ++ .../iteration-protocol.md | 229 ++-------- .../subagent-dispatch.md | 122 ++++++ .../SKILL.md | 392 ++++++----------- skills/skill-optimizer-write-tests/SKILL.md | 404 +++++------------- 11 files changed, 1039 insertions(+), 1969 deletions(-) create mode 100644 skills/skill-optimizer-shared/frontmatter-discipline.md create mode 100644 skills/skill-optimizer-shared/subagent-dispatch.md diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md index 0d48631..7f43fb6 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -10,41 +10,20 @@ trial output from step 5, dispatches an analyzer subagent to cluster failures into named **structural weaknesses** of the skill (or explicitly say there are none), and writes `docs/skill-optimizer//06-analysis.md`. This is the chain's -**anti-ducktape gate**: step 7 (`improve-skill`) refuses to fire -unless this report names at least one structural weakness, with the -general principle that WOULD address it and the anti-patterns that -would NOT. - -## Before you start - -Two load-bearing pieces of context to load NOW, before the workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for staleness - detection, operator-directive handling, and the constraints - subagents must obey. Step 6 is a **fresh-derivation** step: - each invocation derives a new analysis from the bench data + - directives, with no continuity from prior analyses. The - anti-ducktape constraint is load-bearing here. - -2. **You will dispatch a subagent for the actual analysis; you do - NOT cluster failures or name weaknesses yourself in this - session.** Workflow step (c) is the dispatch. The rationale is - in "Why limited-context dispatch matters" below — read it if - you're tempted to skip the dispatch. +**anti-ducktape gate**: step 7 refuses to fire unless this report +names at least one structural weakness, with the general principle +that WOULD address it and the anti-patterns that would NOT. Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain (this skill is step 6; `run-bench` is -step 5; `improve-skill` is step 7). Internal workflow steps within -THIS skill are labelled "(a)" through "(e)". +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(e)". ## What you produce -A single report at `docs/skill-optimizer//06-analysis.md`, -where `` matches the slug from step 1's report. +A single report at `docs/skill-optimizer//06-analysis.md`. -Frontmatter (runtime-relevant facts only — no version tracking, -per the iteration protocol's frontmatter discipline rule): +Frontmatter (runtime-relevant facts only, per +[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- @@ -54,249 +33,133 @@ bench_results_path: 05-bench-results// --- ``` -**`has_structural_weakness`** — load-bearing for step 7. If -`false`, step 7 refuses to fire ("no weakness to address"). The -analyzer subagent sets this honestly based on whether it can -articulate at least one structural weakness; a forced -`has_structural_weakness: true` when the analyzer found nothing is -the ducktape failure mode this step exists to prevent. - -**`weakness_count`** — informational; should match the number of -`### Weakness :` sections in the body. - -**`bench_results_path`** — relative path to the timestamped raw -output the analyzer read from. Lets step 7's optimizer (if it -fires) trace back to the underlying data without re-deriving. - -The body has two top-level sections (per-weakness entries + -non-structural noise) and an optional honest-refusal note if no -weakness was identified. **Each per-weakness entry must include -five required parts**: Pattern, Hypothesized cause, Connects to -skill section, What WOULD address this, What WOULD NOT address -this. The operator session verifies these five parts are present -in step (d); the subagent prompt template specifies their exact -shape. +**`has_structural_weakness`** is load-bearing for step 7 — `false` +gates step 7 from firing ("no weakness to address"). Forcing +`true` when the analyzer found nothing is the ducktape failure +this step exists to prevent. + +Body has per-weakness sections + a non-structural-noise section + +an optional honest-refusal note. **Each per-weakness entry must +include five required parts**: Pattern, Hypothesized cause, +Connects to skill section, What WOULD address this, What WOULD NOT +address this. Step (d) verifies these five parts are present. The "What WOULD NOT address this" anti-pattern list is -**load-bearing architecture, not just bookkeeping**. Without it, -step 7's optimizer can pattern-match a patch that fits the symptom -without addressing the cause; the validator then has no explicit -"this would be a ducktape" signal to check against. The analyzer's -job is to name BOTH the principle that should be applied AND the -ducktape moves the optimizer must avoid. This is the chain's -anti-ducktape gate — losing the anti-pattern list breaks the gate. - -For the full body template, the analyzer's reasoning protocol, and -the honest-refusal wording, see +**load-bearing architecture**. Without it, step 7's optimizer can +pattern-match a patch that fits the symptom without addressing the +cause; the validator (step 8) then has no explicit "this would be +a ducktape" signal to check against. Losing the anti-pattern list +breaks the gate. + +Full body template and reasoning protocol in [`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md). -The subagent writes the report body itself. ## Workflow ### (a) Confirm prerequisites -`docs/skill-optimizer//05-bench-summary.md` must exist with a -valid `bench_results_path` and `overall_pass_rate`. If it doesn't, -tell the user to run `skill-optimizer-run-bench` first and stop -here. +`05-bench-summary.md` must exist with `bench_results_path` and +`overall_pass_rate`. If not, tell the user to run +`skill-optimizer-run-bench` first. -If `overall_pass_rate == 1.0`: there's nothing to analyze. Surface -this honestly: +If `overall_pass_rate == 1.0`: there's nothing to analyze. +Surface honestly — either accept that probes don't expose a +weakness, or re-run step 2 with a "make probes harder" directive. +Don't run the analyzer; there are no failures to cluster. -> All trials passed on the current bench. Step 6 has nothing to -> analyze. Two realistic paths: (1) accept that the current probes -> don't expose a weakness in the skill; (2) re-run step 2 with a -> "make probes harder" directive and walk the chain forward again. - -Don't run the analyzer subagent in this case — there are no -failures to cluster. - -Also confirm `docs/skill-optimizer//tests/` exists (so the -analyzer can read probe specs) and the source skill is available -(either vendored at `vendored-skill/` or the local path recorded in -`01-functionality.md`'s `skill_source`). +`tests/` and the source skill must also be available. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it. Step 6 is a **fresh-derivation** step: - -- Each invocation derives a new analysis from the current bench - data + directives. The subagent does NOT read its own prior - `06-analysis.md` or its git history. Anti-ducktape: the - analyzer must not be biased by what it (or a prior analyzer - run) said before. -- If `06-analysis.md` already exists, it will be overwritten by - this invocation. Prior state lives in git history. No checkpoint - commit needed (fresh-derivation; prior state is already - self-contained in its own prior commit). -- Collect `${OPERATOR_DIRECTIVES}` per the protocol. Atomic new - requirements only — examples: "focus on the gpt-5 cluster, the - other models passed", "the user thinks weakness X from a prior - run is actually two separate issues, look for both". The - operator distills these from prior outputs they CAN read; the - subagent works from the distilled directives, never from the raw - prior analysis. +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 6 is **fresh-derivation**: each invocation +overwrites the canonical from current bench data + directives. +Collect `${OPERATOR_DIRECTIVES}` per the protocol — examples: +"focus on the gpt-5 cluster", "the user thinks weakness X is +actually two separate issues". ### (c) Dispatch the analyzer subagent -**Do NOT cluster failures or name weaknesses yourself in this -session.** Dispatch the analyzer subagent via the `Agent` tool -(with worktree isolation if your environment supports it). Load -the prompt template at +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT cluster failures or name weaknesses +yourself in this session** — dispatch the subagent via the `Agent` +tool. Load the prompt template at [`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md) -and substitute the templated inputs (`${BENCH_RESULTS_PATH}` from -`05-bench-summary.md`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, -`${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, -`${OPERATOR_DIRECTIVES}`). - -The subagent sees: - -- `${SUMMARY_PATH}` — `05-bench-summary.md` (entry point with - failed-probe pointer list) -- `${BENCH_RESULTS_PATH}` — `05-bench-results//` - containing per-trial `trace.jsonl` and per-trial `findings.txt` -- `${TESTS_TREE_PATH}` — the `tests/` tree, but the subagent reads - ONLY each probe's `spec.yaml` (what the probe was probing at the - level of intent) — NOT the `workspace/` contents (the raw input - fixtures) -- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local - skill path) so the analyzer can quote the responsible skill - section in its weakness entries -- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new - requirements (may be empty) - -The subagent does NOT see: - -- **The test inputs themselves** - (`tests///workspace/` files). This is - load-bearing: forces the analyzer to think about the SKILL, not - the SOLUTIONS. Reading the raw input would let it reason "the - agent should have detected X specifically in this file" — that's - solution-thinking, not principle-thinking. -- Its own prior `06-analysis.md` (anti-ducktape — must not be - biased by prior analyses) -- Git history of `06-analysis.md` -- `07-improvement-proposal.md`, `08-validator-verdict.md`, or any - prior optimizer attempts (would also bias toward - "weaknesses-that-could-be-patched-the-way-the-optimizer-tried") - -The subagent writes the report itself and returns a brief summary: -weakness count, top weakness by trial-coverage, and the -`has_structural_weakness` value it set. +and substitute `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, +`${TESTS_TREE_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, +`${OPERATOR_DIRECTIVES}`. + +The subagent sees: `05-bench-summary.md` (entry point with +failed-probe pointer list); `05-bench-results//` with +per-trial `trace.jsonl` and `findings.txt`; each probe's +`spec.yaml` from the `tests/` tree (intent only); the skill source +content; `${OPERATOR_DIRECTIVES}`. + +The subagent does NOT see: **the test inputs themselves** +(`tests//workspace/` files) — this is load-bearing; the +analyzer must think about the SKILL, not the SOLUTIONS; its own +prior `06-analysis.md` or git history; `07-improvement-proposal.md`, +`08-validator-verdict.md`, or any prior optimizer attempts. + +**Why this matters:** two specific bias risks. **Solution-thinking** +— seeing the raw fixtures would lead the analyzer to recommend a +patch for that specific input shape (ducktape); blocking the +inputs forces it to reason about why the skill failed to instruct +the agent properly. **Prior-analysis bias** — the operator session +has read prior analyses and prior optimizer attempts; an +in-session analyzer would gravitate toward "what we said last +time" or defensively pivot away from it. Neither is the job. ### (d) Confirm subagent output -Verify the report file exists, the frontmatter parses, and -`weakness_count` matches the number of `### Weakness :` sections -in the body. If `has_structural_weakness: true` but no weakness -sections exist (or vice versa), the subagent's output is -internally inconsistent — surface to the user and ask whether to -re-dispatch with a directive ("your frontmatter says X but body -says Y; reconcile"). - -Check that each weakness section has all five required parts -(Pattern, Hypothesized cause, Connects to skill section, What -WOULD address this, What WOULD NOT address this). If any section -is missing the anti-pattern list, surface it: this is the -anti-ducktape signal step 7 needs, and a weakness without it can't -be passed to the optimizer safely. +Verify the report parses, `weakness_count` matches the number of +`### Weakness :` sections, and each weakness has all five +required parts (especially "What WOULD NOT address this" — the +anti-ducktape signal step 7 needs). If anything's inconsistent or +missing, surface to the user; don't fill it in yourself. ### (e) Hand off -Two handoff messages depending on `has_structural_weakness`: - -If `has_structural_weakness: true`: - -> Analysis complete. `` structural weakness(es) identified. Next, -> invoke `skill-optimizer-improve-skill` to draft a principled fix. - -If `has_structural_weakness: false`: - -> Analysis complete. No structural weakness identified — failures -> observed are consistent with noise rather than a fixable defect. -> Step 7 will refuse to fire. Two realistic paths: (1) accept the -> conclusion and exit honestly; (2) if you disagree, re-invoke -> step 6 with a directive ("the analyzer missed the X cluster — -> look at it again") to get a fresh derivation. - -Don't auto-invoke step 7 — the chain skills never invoke each -other on their own (per the iteration protocol's re-run -authorization rule). - -## Why limited-context dispatch matters - -The analyzer subagent's anti-ducktape job is the chain's most -fragile in terms of bias. Two specific risks the limited context -addresses: - -First, **solution-thinking vs. principle-thinking**. If the -analyzer sees the raw input fixtures in `tests//workspace/`, -it will naturally reason "the agent should have detected X -specifically in this file" — and recommend a patch that says -"detect X in files like this". That's a ducktape patch around a -specific input shape. By blocking the raw inputs, the analyzer is -forced to read only the probe's stated intent + the agent's actual -behavior, and reason about why the skill failed to instruct the -agent properly. The output is a principle (what the skill should -teach the agent to do in general), not a solution (what the agent -should have done in this specific case). - -Second, **prior-analysis bias**. The operator session has read -prior `06-analysis.md`s, prior optimizer attempts, prior validator -verdicts. That context would lead an in-session analyzer to -gravitate toward "what we said last time" or "what the optimizer -tried last time" — and either confirm those framings or -defensively pivot away from them. Neither is the right job. The -analyzer must derive fresh from the current data + the operator's -distilled directives, with no view of prior runs' meta-discussion. - -If you find yourself thinking "I'll just write the analysis -myself, I already see the patterns" — that's the failure mode this -chain is built to prevent. Dispatch. +Two messages depending on `has_structural_weakness`: + +- **`true`:** "Analysis complete. `` structural weakness(es) + identified. Next, invoke `skill-optimizer-improve-skill`." +- **`false`:** "No structural weakness identified — failures + consistent with noise rather than a fixable defect. Step 7 will + refuse to fire. Either accept the conclusion, or re-invoke step + 6 with a directive ('the analyzer missed the X cluster')." + +Don't auto-invoke step 7; per the no-auto-invocation rule in +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), +the user decides. ## Edge cases - **All trials passed** — caught at (a). Don't run the analyzer. - **Bench results dir missing or `suite-result.json` malformed** — - the analyzer can't proceed; surface as a step-5 problem (the - bench run was incomplete or corrupted) and tell the user to - re-run step 5. + surface as a step-5 problem; tell the user to re-run step 5. - **Subagent returns `has_structural_weakness: false` and user disagrees** — surface the disagreement, but do NOT pressure the - subagent to manufacture a weakness. The honest path is to - re-invoke step 6 with a directive pointing at what the user - thinks was missed ("the gpt-5 trial 3 failure pattern wasn't - covered — look at it specifically"). Per the iteration protocol's - re-run authorization rule, the user invokes the re-run; this - skill doesn't auto-invoke. -- **Subagent returns weaknesses but the "What WOULD NOT address - this" lists are empty or vague** — the anti-ducktape gate is - compromised. Re-dispatch with a directive: "each weakness needs - a concrete anti-pattern list — what specific moves should the - optimizer avoid?" Don't fill it in yourself. -- **Operator directives contradict each other** (e.g., "focus on - cluster A" + "ignore cluster A") — surface to the user before - re-dispatching; don't try to resolve it yourself. + subagent to manufacture a weakness. Honest path: re-invoke step + 6 with a directive pointing at what the user thinks was missed. +- **Subagent returns weaknesses but the anti-pattern lists are + empty/vague** — anti-ducktape gate is compromised. Re-dispatch + with a directive ("each weakness needs a concrete anti-pattern + list"). +- **Operator directives contradict** — surface before + re-dispatching. ## Iteration behavior -Step 6 is re-runnable and is a fresh-derivation step. Re-run -triggers specific to this step: - -- Step 5 (`run-bench`) produced new results — `05-bench-summary.md` - changed via git mtime -- User disagrees with the analyzer's verdict (either thinks a - weakness was missed, or thinks `has_structural_weakness: true` - should have been `false`) and wants a fresh derivation with - directives reflecting the disagreement -- Step 7 (`improve-skill`) was unable to address one of the named - weaknesses and the user wants the analyzer to reformulate it - ("weakness 2 as named was too abstract; split into concrete - sub-weaknesses") - -General iteration mechanics — staleness detection (git-native), -operator directives (the channel for prior-derivation lessons), -cascading staleness — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. +Fresh-derivation step. Re-run triggers: + +- Step 5 produced new results — `05-bench-summary.md` changed via + git mtime +- User disagrees with the analyzer's verdict and wants a fresh + derivation with directives +- Step 7 was unable to address one of the named weaknesses and the + user wants it reformulated + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md index 8f054e1..8ef1b8f 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -6,64 +6,39 @@ description: Use when the user wants to improve a skill based on identified stru # skill-optimizer-improve-skill Step 7 of the skill-optimizer chain. Takes the named structural -weaknesses from step 6, dispatches an optimizer subagent to draft -a principled fix, and writes `07-improvement-proposal.md`. The -proposal is then validated independently by step 8 -(`validate-improvement`), which materializes the improved skill on -approve. **This step does not produce the improved skill itself** -— it produces the proposal that step 8 acts on. - -This skill **refuses to fire** if `06-analysis.md` has +weaknesses from step 6, dispatches an optimizer subagent to draft a +principled fix, and writes `07-improvement-proposal.md`. The +proposal is then validated independently by step 8, which +materializes the improved skill on approve. **This step does not +produce the improved skill itself** — it produces the proposal that +step 8 acts on. + +**Refuses to fire** if `06-analysis.md` has `has_structural_weakness: false` — there's nothing to optimize, and -forcing a fix in that situation is the ducktape failure mode the -chain is built to prevent. - -**Single-shot per invocation.** There is no in-step revision loop; -each invocation produces one proposal. If step 8's validator returns -`needs-revision`, the operator (or auto-pilot at step 9) distills -the validator's rationale into a directive and re-invokes this -step. That keeps chain skills from auto-invoking each other. - -**Out of scope:** validating the proposal (that's step 8) and -packaging the change as a PR draft (PR composition is a separate -downstream concern; auto-pilot or a dedicated composer can handle it -if `pr_submission_intent: true`). - -## Before you start - -Two load-bearing pieces of context to load NOW, before the -workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation. Step 7 is a - **fresh-derivation** step. Pay attention to the **step-7 - carve-out** in the subagent-constraints section: the SKILL - CONTENT is upstream input — `improved-skill/` if it exists - (the accumulated improvement state from prior approved-by-step-8 - runs), else the original source (`vendored-skill/` for upstream, - the user's local file for local). The optimizer's own canonical - (`07-improvement-proposal.md`) is off-limits — that's the report - describing the optimizer's reasoning, not the skill being - improved. - -2. **You will dispatch a subagent for the actual optimization; you - do NOT propose diffs yourself in this session.** Workflow step - (c) is the dispatch. The rationale is in "Why limited-context - dispatch matters" below. +forcing a fix is the ducktape failure mode the chain is built to +prevent. + +**Single-shot per invocation.** No in-step revision loop. If step +8's validator returns `needs-revision`, the operator (or +auto-pilot at step 9) distills the validator's rationale into a +directive and re-invokes this step. + +**Out of scope:** validating the proposal (step 8) and packaging +the change as a PR draft (PR composition is a separate downstream +concern; auto-pilot or a dedicated composer can handle it if +`pr_submission_intent: true`). Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain (this skill is step 7; -`analyze-result` is step 6; `validate-improvement` is step 8; -`autopilot` is step 9). Internal workflow steps within THIS skill -are labelled "(a)" through "(e)". +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(e)". ## What you produce One artifact at `docs/skill-optimizer//`: -**`07-improvement-proposal.md`** — the optimizer's proposed -change + rationale. Frontmatter (runtime-relevant facts only, per -the iteration protocol's frontmatter discipline): +**`07-improvement-proposal.md`** — the optimizer's proposed change +plus rationale. Frontmatter (runtime-relevant facts only, per +[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- @@ -73,245 +48,134 @@ addresses_weaknesses: --- ``` -The body lists the proposed change, the rationale referencing the -named weakness from `06-analysis.md` and the "What WOULD address -this" principle being applied, and a self-check against the "What -WOULD NOT address this" anti-pattern list (the optimizer states -explicitly why its proposal is NOT one of the ducktape moves the -analyzer flagged). The subagent prompt template specifies the -section shape. - -For the body template + the optimizer's reasoning protocol, see -[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md). -The subagent writes its body itself; the operator session only -orchestrates the dispatch. +Body has: the proposed change, rationale referencing the named +weakness and the "What WOULD address this" principle being +applied, and an explicit self-check against the "What WOULD NOT +address this" anti-pattern list (the optimizer states why its +proposal is NOT one of the ducktape moves the analyzer flagged). +Subagent prompt template at +[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) +specifies the section shape. ## Workflow ### (a) Confirm prerequisites + refuse-if-no-weakness gate -Three checks, in order: - -1. `docs/skill-optimizer//06-analysis.md` must exist with - valid frontmatter. If not, tell the user to run - `skill-optimizer-analyze-result` first and stop here. +Three checks: -2. **Anti-ducktape gate:** `06-analysis.md`'s frontmatter must - have `has_structural_weakness: true`. If `false`, REFUSE to - proceed — print a message like: +1. `06-analysis.md` must exist with valid frontmatter. If not, + tell the user to run `skill-optimizer-analyze-result` first. - > Step 6 found no structural weakness in the current skill - > against the current probes. Step 7 won't fire — there's - > nothing principled to optimize. If you disagree, re-invoke - > step 6 with a directive describing what you think was - > missed; if you agree, exit honestly. +2. **Anti-ducktape gate:** `has_structural_weakness: true` must be + set. If `false`, REFUSE — print: "Step 6 found no structural + weakness. Step 7 won't fire — nothing principled to optimize. + If you disagree, re-invoke step 6 with a directive; if you + agree, exit honestly." Do NOT proceed. - Do NOT proceed to dispatch anyway. The refusal is the entire - reason this gate exists. +3. `01-functionality.md` must exist. Skill content must be + readable: `improved-skill/` if it exists (accumulated state + from prior step-8 approvals), else the original source. -3. `01-functionality.md` must exist (the optimizer reads it as - input). The skill content must be readable: `improved-skill/` - if it exists (the accumulated state from prior step-8 - approvals), else the original source (`vendored-skill/` for - upstream skills, the local file path recorded in - `01-functionality.md`'s `skill_source` for local skills). - -If `pr_submission_intent: true` from `01-functionality.md`, -`03-submissions.md` should also exist (the optimizer reads it to -shape the diff to match upstream conventions from the start). If -it's missing, surface this and ask the user whether to run step -3 first OR proceed without upstream-shape hints (step 8's -external check will still run if `03-submissions.md` appears -later). +If `pr_submission_intent: true`, `03-submissions.md` should exist +so the optimizer can shape the diff to upstream conventions from +the start. If missing, ask whether to run step 3 first — step 8's +external check still runs if it appears later. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it. Step 7 is a **fresh-derivation** step: - -- Each invocation derives a fresh proposal from `06-analysis.md`, - the current skill content, and any operator directives. The - optimizer does NOT read its own prior - `07-improvement-proposal.md` or its git history. -- If `07-improvement-proposal.md` already exists, it will be - overwritten by this invocation. Prior state lives in git - history. No checkpoint commit needed (fresh-derivation; prior - state is already self-contained in its own prior commit). -- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The - operator CAN read prior proposals and step 8's prior verdicts - (which they need to distill anyway, after a `needs-revision` - verdict). Examples: - - "prefer additive changes over destructive ones" - - "don't touch the description field — validator step 8 - rejected that last round" - - "address weakness 2 first; weakness 1 was already partially - addressed by the prior round" - - The subagent works from the distilled directives, never from - the raw prior content. +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 7 is **fresh-derivation**: each invocation +overwrites the canonical from `06-analysis.md` + current skill + +directives. Collect `${OPERATOR_DIRECTIVES}` — examples: "prefer +additive changes", "don't touch the description field — validator +rejected that last round". The operator reads prior proposals/ +verdicts and distills; the subagent never sees the raw prior +content. ### (c) Dispatch the optimizer subagent -**Do NOT propose the diff yourself in this session.** Dispatch -the optimizer subagent via the `Agent` tool (with worktree -isolation if your environment supports it). Load the prompt -template at +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT propose the diff yourself in this +session** — dispatch the subagent via the `Agent` tool. Load the +prompt template at [`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) -and substitute the templated inputs (`${ANALYSIS_PATH}`, -`${FUNCTIONALITY_PATH}`, `${SKILL_CURRENT_PATH}`, -`${SUBMISSIONS_PATH}` if PR-bound, -`${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`). - -The optimizer sees: - -- `${ANALYSIS_PATH}` — `06-analysis.md` (the named weaknesses + - what WOULD/WOULDN'T address each) -- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (so the - optimizer understands the skill's stated responsibilities and - doesn't propose a change that contradicts them) -- `${SKILL_CURRENT_PATH}` — the current state of the skill being - improved: `improved-skill/` if it exists, else the original - source. The optimizer proposes a new improvement on top of - whatever current state it sees. -- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (so - the optimizer can shape the diff to match upstream conventions - from the start, reducing step-8 round-trips) -- `${OPERATOR_DIRECTIVES}` — atomic new requirements (may be - empty) - -The optimizer does NOT see: - -- Raw failed trials (`findings.txt`, `trace.jsonl` from - `05-bench-results/`) -- Grader internals (`tests///grader.mjs` source) -- Test inputs (`tests///workspace/`) -- Its own prior `07-improvement-proposal.md` or git history of it -- Step 8's prior `08-validator-verdict.md` (the validator's prior - judgment would bias the optimizer toward defending or pivoting - away from the prior attempt rather than addressing the weakness - fresh; lessons from the prior verdict come through distilled - directives instead) - -The "no raw trials / grader internals / test inputs" constraint -is load-bearing: it forces the optimizer to address the weakness -as the analyzer named it, in terms of the general principle the -analyzer articulated — not by pattern-matching a patch that would -make specific failing trials pass. - -The optimizer writes `07-improvement-proposal.md` (the diff + -rationale, including the explicit self-check against the -anti-pattern list) and returns a brief summary: the weakness(es) -addressed, the principle applied, the lines/sections of the -skill modified. +and substitute `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, +`${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, +`${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. + +The optimizer sees: `06-analysis.md`; `01-functionality.md`; +`${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else +original source; `03-submissions.md` if PR-bound; +`${OPERATOR_DIRECTIVES}`. + +The optimizer does NOT see: raw failed trials, `findings.txt`, +`trace.jsonl`; grader internals (`tests//grader.mjs`); test +inputs (`tests//workspace/`); its own prior +`07-improvement-proposal.md` or git history; step 8's prior +`08-validator-verdict.md` (would bias toward defending or pivoting +away from the prior attempt). + +**Why this matters:** the chain's anti-ducktape architecture +hinges on this. If the optimizer read failed trials, it would +pattern-match a patch making those specific trials pass — the +textbook ducktape mode. The analyzer's job (step 6) is to +translate raw failures into general principles + anti-patterns; +the optimizer's job is to apply the principle. The translation +through `06-analysis.md` is what enforces principled improvement. + +The optimizer writes `07-improvement-proposal.md` and returns: +weakness(es) addressed, principle applied, lines/sections modified. ### (d) Confirm subagent output -Verify the report file exists, the frontmatter parses, and the -body has: - -- A clear proposed change (in fenced unified-diff format or - similar) -- A rationale section referencing at least one named weakness - from `06-analysis.md` and the "What WOULD address this" - principle it applies -- An explicit self-check against the "What WOULD NOT address - this" anti-pattern list — the optimizer must state why its - proposal is NOT one of the ducktape moves the analyzer - flagged - -If the self-check section is missing or vague, surface to the -user — this is the architectural anti-ducktape signal the -optimizer is required to produce. Re-dispatch with a directive -"the self-check against the anti-pattern list is required and -missing or vague — make it explicit" rather than filling it in -yourself. +Verify the report parses and the body has: a clear proposed +change, a rationale referencing at least one named weakness and +the "What WOULD address this" principle, and an explicit +self-check against the anti-pattern list. If the self-check is +missing or vague, re-dispatch with a directive making the +requirement explicit — don't fill it in yourself. ### (e) Hand off -Single handoff message: - -> Improvement proposal complete. The proposal is at -> `docs/skill-optimizer//07-improvement-proposal.md`. Next, -> invoke `skill-optimizer-validate-improvement` to check the -> proposal independently. If the validator returns -> `needs-revision` or `reject`, you'll come back here with a -> directive distilling the validator's concerns. +> Improvement proposal complete at +> `07-improvement-proposal.md`. Next, invoke +> `skill-optimizer-validate-improvement` to check the proposal +> independently. If the validator returns `needs-revision` or +> `reject`, you'll come back here with a directive distilling the +> validator's concerns. -Don't auto-invoke step 8 — the chain skills never invoke each -other on their own (per the iteration protocol's re-run -authorization rule). - -## Why limited-context dispatch matters - -The chain's whole anti-ducktape architecture hinges on the -optimizer NOT seeing raw failures. If the optimizer read the -failed trials, it would pattern-match a patch that makes those -specific trials pass — the textbook ducktape failure mode. The -analyzer's job (step 6) is to translate raw failures into **named -structural weaknesses with general principles and anti-patterns**; -the optimizer's job is to apply the principle. The translation -through the analyzer's report is what enforces principled -improvement over pattern-match patching. - -The optimizer also doesn't read its own prior canonical or step -8's prior verdicts. Either would bias toward defending or -pivoting away from prior attempts rather than addressing the -weakness fresh; lessons from prior runs come through the -operator's distilled directives. Each invocation of step 7 is -genuinely fresh. - -If you find yourself thinking "I'll just propose the change -myself, I know what the analysis says" — that's the failure mode -this chain is built to prevent. Dispatch. +Don't auto-invoke step 8. ## Edge cases -- **`06-analysis.md` missing** — caught at (a). Tell user to run - step 6 first. -- **`has_structural_weakness: false`** — caught at (a). Refuse - to fire; this is the anti-ducktape gate. Don't try to override. -- **Optimizer reports BLOCKED** (e.g., the analyzer's named - weakness is too abstract to derive a concrete diff from) — - surface to the user. The fix is typically a step 6 re-run with - a directive ("weakness X as named was too abstract; split into - concrete sub-weaknesses"). Per the iteration protocol's re-run - authorization rule, the user invokes step 6. -- **Optimizer's self-check against anti-patterns is missing or - vague** — caught at (d). Re-dispatch with a directive making - the self-check requirement explicit. +- **`06-analysis.md` missing** — caught at (a). +- **`has_structural_weakness: false`** — caught at (a). Refuse; + don't override. +- **Optimizer reports BLOCKED** (analyzer's weakness too abstract + to derive a diff from) — surface. Fix is typically a step 6 + re-run; user invokes step 6. +- **Self-check against anti-patterns missing or vague** — caught + at (d). Re-dispatch with a directive. - **PR-bound but `03-submissions.md` missing** — caught at (a). - Proceeding without it means the optimizer doesn't get upstream - shape hints from the start; step 8's external check will - still run if `03-submissions.md` appears later (the operator - can run step 3 before step 8 if they want). + Proceeding loses upstream shape hints from the start; step 8's + external check still runs if `03-submissions.md` appears later. - **Operator directives reference a specific section of a prior - proposal or verdict** (e.g., "the optimizer's section on the - description field was wrong") — that's a context-dump - masquerading as a directive. Translate to an atomic new - requirement ("don't modify the description field") before - passing to the subagent. + proposal or verdict** — that's a context-dump masquerading as a + directive. Translate to an atomic new requirement before passing. ## Iteration behavior -Step 7 is re-runnable and is a fresh-derivation step. Re-run -triggers specific to this step: - -- Step 6 (`analyze-result`) produced a new analysis — - `06-analysis.md` changed via git mtime, possibly naming - different weaknesses -- Step 8 (`validate-improvement`) returned `needs-revision` or - `reject` — the operator distills the validator's rationale - into a directive and re-invokes this step -- User wants the optimizer to try a different approach (passes - directives like "this time prefer additive changes" or "focus - on weakness 2 only") - -General iteration mechanics — staleness detection (git-native), -operator directives (the channel for prior-derivation lessons), -the step-7 carve-out for the skill being upstream input — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. - -There is no in-step revision loop. Each invocation produces one -proposal. The 7→8→7 cycle on `needs-revision` verdicts is -operator-driven (or auto-pilot-driven at step 9). +Fresh-derivation step. Re-run triggers: + +- Step 6 produced a new analysis — `06-analysis.md` changed +- Step 8 returned `needs-revision`/`reject` — operator distills + the rationale into a directive +- User wants a different approach (passes directives like "prefer + additive changes", "focus on weakness 2 only") + +The 7→8→7 cycle on `needs-revision` is operator-driven, or +auto-pilot-driven at step 9. No in-step loop. + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 99d7dab..6bd73ec 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -10,26 +10,9 @@ runs a researcher subagent to figure out what it's supposed to do, and writes `docs/skill-optimizer//01-functionality.md` — the briefing document every later step consumes. -## Before you start - -Two load-bearing pieces of context to load NOW, before the workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for staleness - detection, operator-directive handling, and the constraints - subagents must obey. Workflow step (e) below requires it; - loading it now lets you execute that step deterministically - instead of improvising. - -2. **You will dispatch a subagent for the actual research; you do NOT - do the research yourself in this session.** Workflow step (f) is - the dispatch. The rationale is in "Why limited-context dispatch - matters" below — read it if you're tempted to skip the dispatch. - Throughout this document, "step 1" through "step 9" (no parens) refer -to skills in the chain (this skill is step 1; `investigate-test-case` -is step 2; etc.). Internal workflow steps within THIS skill are -labelled "(a)" through "(g)" to avoid the collision. +to skills in the chain. Internal workflow steps within THIS skill are +labelled "(a)" through "(g)". ## What you produce @@ -37,7 +20,8 @@ A single report at `docs/skill-optimizer//01-functionality.md`, where `` is the source skill's directory name (e.g., `firecrawl-build-scrape`). -The report has structured frontmatter: +Frontmatter (runtime-relevant facts only, per +[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- @@ -47,155 +31,95 @@ classification: --- ``` -**`classification`** — pick the most accurate label for what kind of -skill this is. Canonical types: `tool-use` (procedures for using a -specific tool, library, or API), `code-patterns` (code-level patterns -or review checklists), `document` (workflows that produce a document -or file), `prose-guidance` (writing-style or content-creation -guidance), `meta` (skills that operate on other skills or on the -agent's behavior), `interactive` (back-and-forth user dialogue). If -none of these fits cleanly, **write a short descriptive label of your -own** (`dataset-extraction`, `deployment-runbook`, -`ui-mockup-generation`, etc.) rather than falling back to `other` — a -specific label gives downstream steps a real handle to work with. - -For the body template (what sections, what the subagent must cover), -see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). -The subagent itself writes the report — you don't. - -If the source is upstream, you also vendor the fetched skill files to -`vendored-skill/` in the working directory so downstream steps read a -stable copy without re-fetching. +**`classification`** — canonical types: `tool-use`, `code-patterns`, +`document`, `prose-guidance`, `meta`, `interactive`. If none fits +cleanly, write a short descriptive label of your own +(`dataset-extraction`, `deployment-runbook`, etc.) rather than +falling back to `other` — a specific label gives downstream steps a +real handle. + +For the body template, see +[`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). +If the source is upstream, also vendor the fetched skill files to +`vendored-skill/` so downstream steps read a stable copy. ## Workflow ### (a) Classify the source -Is the source an **upstream** skill (a URL or `//` -slug) or a **local** skill (a filesystem path that already exists)? If -the user gave a bare name with no URL and no path, ask them to clarify -before continuing. +Upstream skill (URL or `//`) or local skill +(filesystem path that already exists)? If the user gave a bare name +with no URL and no path, ask to clarify before continuing. ### (b) Ask about PR intent -**Upstream skills:** ask the user, in roughly these words: +**Upstream skills:** ask the user "Do you want to optimize this +skill for upstream PR submission?" and record the answer in the +report frontmatter as `pr_submission_intent: true|false`. Capture +this decision now — step 2's handoff reads this field to decide +whether step 3 (`investigate-submissions`) runs. -> Do you want to optimize this skill for upstream PR submission? - -Record the answer in the report frontmatter as -`pr_submission_intent: true` or `pr_submission_intent: false`. **Capture -this decision now**, not at the end of the chain — step 2's handoff -reads this field to decide whether step 3 (`investigate-submissions`) -runs. - -**Local skills:** default to `pr_submission_intent: false`. But if the -user explicitly said they want to send this back to an upstream -maintainer (e.g., "I'll fork this back to the original repo", "I want -to PR this to project X"), treat it as PR-intent. In that case: - -1. Set `pr_submission_intent: true`. -2. Ask the user where the upstream contribution guidelines live — a - URL, a `CONTRIBUTING.md` path, a Slack channel, whatever they have. -3. Record what they say in the report body under a "PR submission - notes" subsection. Step 3 (`investigate-submissions`) uses this as - the starting point for `03-submissions.md`. - -If the user mentions nothing about a PR for the local skill, set -`pr_submission_intent: false` and move on. +**Local skills:** default to `pr_submission_intent: false`. But if +the user explicitly said they want to send this back to an upstream +maintainer, treat it as PR-intent: set `pr_submission_intent: true`, +ask where the upstream contribution guidelines live, and record +their answer in the report body under a "PR submission notes" +subsection (step 3 uses this as its starting point). ### (c) Vendor the source (upstream only) -Fetch the skill's files into `vendored-skill/` at the working-directory -root. Downstream steps read from this vendored copy as a stable -reference. Local-source skills don't need vendoring. +Fetch the skill's files into `vendored-skill/` at the +working-directory root. Local-source skills don't need vendoring. ### (d) Determine the slug and the report path -`` is the source skill's own directory or file name: - -- `firecrawl/skills/firecrawl-build-scrape` → slug is `firecrawl-build-scrape` -- `~/my-skills/pdf-cleanup/SKILL.md` → slug is `pdf-cleanup` +`` is the source skill's own directory or file name (e.g., +`firecrawl/skills/firecrawl-build-scrape` → `firecrawl-build-scrape`; +`~/my-skills/pdf-cleanup/SKILL.md` → `pdf-cleanup`). Report path: `docs/skill-optimizer//01-functionality.md`. ### (e) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. Step 1 is a **fresh-derivation** step -— each invocation derives the report from the current source skill -plus any operator directives, with no continuity from prior outputs. There is no -`version:` field to bump and no `archive/` directory to manage; -git history captures prior states automatically. - -For step 1 specifically: there are no upstream chain reports, so -staleness is determined by whether the user wants fresh research -(source URL changed, PR-intent answer changed, or new directives). -Re-run on user signal, not on automatic upstream-mismatch detection. - -Collect `${OPERATOR_DIRECTIVES}` per the protocol's section if the -user supplied atomic new requirements (e.g., "you missed the vendor -CLA requirement"). If `01-functionality.md` already exists and will -be overwritten, that's fine — the prior state lives in git history. -No checkpoint commit is needed for fresh-derivation steps (the prior -state is self-contained in its own prior commit). +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 1 is **fresh-derivation**: each invocation +overwrites the canonical from the current source + directives; git +captures prior state. Re-run on user signal (source URL changed, +PR-intent answer changed, or new directives). Collect +`${OPERATOR_DIRECTIVES}` per the protocol. ### (f) Dispatch the functionality-researcher subagent -**Do NOT do the research yourself in this session.** Dispatch the -functionality-researcher subagent via the `Agent` tool (with worktree -isolation if your environment supports it). Load the prompt template -at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md), -substitute the templated inputs (`${SKILL_SOURCE}`, `${OUTPUT_PATH}`, -`${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}` from (e), -vendored path), and dispatch. - -The subagent sees: - -- The vendored skill files (or the local skill path) -- Targeted web-search / web-fetch results for the underlying technology -- The output path and the frontmatter fields you computed -- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new - requirements (may be empty) - -The subagent does NOT see: - -- Its own prior `01-functionality.md` (anti-ducktape: must derive - fresh, not rationalize what was there before) -- Git history of `01-functionality.md` -- Existing analyses or improvement proposals for this skill -- Prior tests or failure data -- The wider chain's context - -The subagent writes the report itself and returns a brief summary. +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT do the research yourself in this +session** — dispatch the subagent via the `Agent` tool (with +worktree isolation if your environment supports it). Load the +prompt template at +[`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md) +and substitute `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, +`${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, vendored path. + +The subagent sees: the vendored skill files (or local path), +targeted web-search results, output path, frontmatter fields, +`${OPERATOR_DIRECTIVES}`. + +The subagent does NOT see: prior `01-functionality.md` or its git +history; existing analyses, tests, or failure data; the wider +chain's context. + +**Why this matters:** the operator session inherits prior +conversation framing, which biases research toward whatever the +user has already expressed. A subagent walled off from that +context produces a fresh derivation from the source itself, not a +rationalization of expectations. ### (g) Confirm and hand off -When the subagent returns: - -1. Verify the report file exists and the frontmatter parses. -2. Report to the user: the report path, a one-line summary of the - subagent's classification and key findings, then the handoff message. - -Handoff message, verbatim: +Verify the report file exists and frontmatter parses. Then: > Next, invoke `skill-optimizer-investigate-test-case`. Note: this -> report's `pr_submission_intent` field tells step 2's handoff whether -> step 3 (`investigate-submissions`) should run. - -## Why limited-context dispatch matters - -When the same context that handles user conversation also conducts the -research, the research drifts toward whatever the user has already -expressed — and every later step in the chain inherits that drift. -Limited-context subagents are the architectural fix: walling the -researcher off from prior analyses and prior failures prevents the -researcher from rationalizing the existing skill's design choices, and -prevents "ducktape" fixes that paper over symptoms instead of addressing -root causes. - -If you find yourself thinking "I'll just write the report myself, the -subagent dispatch is bureaucratic overhead" — that's the failure mode -this chain is built to prevent. Dispatch. +> report's `pr_submission_intent` field tells step 2's handoff +> whether step 3 should run. ## Edge cases @@ -203,25 +127,22 @@ this chain is built to prevent. Dispatch. user; don't try to guess a recovery. - **Source is a plugin with multiple skills** — ask which one to investigate; produce one report per skill. -- **`vendored-skill/` already exists from a prior run** — reuse the - existing copy by default (it's the same source). Re-fetch only if - the source URL itself changed or the user explicitly asks. +- **`vendored-skill/` already exists from a prior run** — reuse it + unless the source URL changed or the user explicitly asks to + re-fetch. - **User changes their mind on PR intent later** — they re-run this - skill (see iteration section below); the canonical - `01-functionality.md` is overwritten and the prior state lives in + skill; the canonical is overwritten and the prior state lives in git history. ## Iteration behavior -Step 1 is re-runnable and is a fresh-derivation step. Re-run -triggers specific to this step: +Fresh-derivation step. Re-run triggers: -- The source URL changed (different upstream skill, or repo moved) -- The user changed their mind on PR intent -- The user wants the report re-derived with new directives ("you - missed the vendor CLA requirement", "go deeper on who-uses-this") +- Source URL changed (different upstream skill, or repo moved) +- User changed their mind on PR intent +- User wants the report re-derived with new directives ("you missed + the vendor CLA requirement", "go deeper on who-uses-this") -General iteration mechanics — staleness detection (git-native), -operator directives, cascading staleness — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (e) above already requires reading that file. +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (e) loads it. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index a700fde..c6464e0 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -8,41 +8,21 @@ description: Use when the user wants to research a skill's upstream PR conventio Step 3 of the skill-optimizer chain — OPTIONAL, runs only when the target skill is bound for upstream PR submission. Takes the source slug, dispatches a researcher subagent that uses the `gh` CLI to -gather the upstream repo's contribution conventions (license, CLA, -frontmatter spec, file-location rules, PR-shape patterns from recent -merged + closed-without-merge PRs), and writes +gather the upstream repo's contribution conventions, and writes `docs/skill-optimizer//03-submissions.md` — the verbatim-pastable context block the validator (step 8) uses for its external consistency check. -## Before you start - -Two load-bearing pieces of context to load NOW, before the workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for staleness - detection, operator-directive handling, and the constraints - subagents must obey. Workflow step (b) below requires it; - loading it now lets you execute that step deterministically - instead of improvising. - -2. **You will dispatch a subagent for the actual research; you do NOT - scrape the upstream repo yourself in this session.** Workflow step - (c) is the dispatch. The rationale is in "Why limited-context - dispatch matters" below — read it if you're tempted to skip the - dispatch. - -Throughout this document, "step 1" through "step 9" (no parens) refer -to skills in the chain (this skill is step 3; `investigate-functionality` -is step 1; etc.). Internal workflow steps within THIS skill are -labelled "(a)" through "(e)" to avoid the collision. +Throughout this document, "step 1" through "step 9" (no parens) +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(e)". ## What you produce -A single report at `docs/skill-optimizer//03-submissions.md`, -where `` matches the slug from step 1's report. +A single report at `docs/skill-optimizer//03-submissions.md`. -The report has structured frontmatter: +Frontmatter (runtime-relevant facts only, per +[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- @@ -53,24 +33,18 @@ requires_cla: true | false --- ``` -**`upstream_branch_target`** — some repos use `main` for incremental -changes and `next` for new skills; some use a single branch. The -subagent determines this from recent merged PRs and records it here -so the validator can check the proposed PR targets the correct -branch. - -**`requires_cla`** — true if the upstream requires a Contributor -License Agreement before merging (Google, Apache Foundation, etc.). -The PR-packaging step (in `improve-skill`'s upstream-PR handoff) -surfaces this to the operator. - -For the body template (what sections the researcher must cover — -license details, frontmatter spec extracted from existing skills, -file-location conventions, prefix taxonomy, PR-shape patterns from -recent merged + closed-without-merge PRs, rejection signals from the -closed-without-merge set), see +- **`upstream_branch_target`** — some repos use `main` for + incremental changes and `next` for new skills. The subagent + determines this from recent merged PRs so the validator can + check the proposed PR targets the correct branch. +- **`requires_cla`** — true if the upstream requires a Contributor + License Agreement before merging. + +Body covers license details, frontmatter spec extracted from +existing skills, file-location conventions, prefix taxonomy, +PR-shape patterns from recent merged + closed-without-merge PRs, +and rejection signals. Body template lives in [`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md). -The subagent writes the body itself. ## Workflow @@ -78,163 +52,105 @@ The subagent writes the body itself. Two prerequisites: -1. `docs/skill-optimizer//01-functionality.md` must exist and - have valid frontmatter. If it doesn't, tell the user to run +1. `01-functionality.md` must exist with valid frontmatter. If + not, tell the user to run `skill-optimizer-investigate-functionality` first and stop here. -2. That report's `pr_submission_intent` field must be `true`. If it's - `false`, this skill should not run — tell the user step 3 is - skipped for local-only optimization runs and refer them to step 4 - (`write-tests`). +2. That report's `pr_submission_intent` field must be `true`. If + it's `false`, this skill should not run — tell the user step 3 + is skipped for local-only optimization runs. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. Step 3 is a **fresh-derivation** step — -each invocation re-derives the report from current upstream facts -plus any operator directives. There is no `version:` field to bump and no -`archive/` directory; git history captures prior states. - +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 3 is **fresh-derivation**: each invocation +overwrites the canonical from current upstream facts + directives. Staleness against `01-functionality.md` is determined by git-mtime -comparison per the protocol: if `01-functionality.md` is newer than -`03-submissions.md`, the existing report is stale (the source URL -may have changed) and should be re-derived. Otherwise the existing -report is current. +comparison. -Note: this report rarely needs re-running. Upstream PR conventions -change slowly. The most common reason to re-run is that the -upstream repo updated its `CONTRIBUTING.md` or CLA requirements — -surface this as an operator directive if you know it. +This report rarely needs re-running — upstream PR conventions +change slowly. Most common reason: upstream updated +`CONTRIBUTING.md` or CLA requirements. Surface as a directive if +known. ### (c) Dispatch the submission-researcher subagent -**Do NOT scrape the upstream repo yourself in this session.** -Dispatch the submission-researcher subagent via the `Agent` tool -(with worktree isolation if your environment supports it). Load the -prompt template at -[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md), -substitute the templated inputs (`${UPSTREAM_REPO}` from -`01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, -`${OPERATOR_DIRECTIVES}` from (b)), and dispatch. - -The subagent sees: - -- The upstream repo via `gh` CLI (PR list — both merged and - closed-without-merge for shape patterns and rejection signals, - repo-file API, `CONTRIBUTING.md`, license file, existing skill - files for frontmatter spec extraction) -- The skill slug being researched (so it can look at similar PRs - in the same skill category) -- `${OPERATOR_DIRECTIVES}` — your bulleted list of atomic new - requirements (may be empty) -- The output path - -The subagent does NOT see: - -- Its own prior `03-submissions.md` (anti-rationalization: must - derive fresh from upstream facts) -- Git history of `03-submissions.md` -- Any information about the proposed change being optimized (the - report is purely about upstream facts, not about whether a specific - change will be accepted) -- Existing analyses, tests, or failure data from the chain -- The vendored skill source - -The constraint that the subagent doesn't see the proposed change is -load-bearing: the report must be neutral upstream facts, not advocacy -for a specific change. The validator (step 8) will later check the -proposed change against this report — if the report is biased toward -the change, the validator's external consistency check loses its -independence. - -The subagent writes the report itself and returns a brief summary -including: license, CLA requirement, branch target, and any -high-risk rejection signals it spotted in the closed-without-merge -PRs. +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT scrape the upstream repo yourself +in this session** — dispatch the subagent via the `Agent` tool. +Load the prompt template at +[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md) +and substitute `${UPSTREAM_REPO}` from `01-functionality.md`'s +`skill_source` field, `${OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. + +The subagent sees: the upstream repo via `gh` CLI (PR list — both +merged and closed-without-merge for shape patterns and rejection +signals, repo-file API, `CONTRIBUTING.md`, license file, existing +skill files); the skill slug being researched; +`${OPERATOR_DIRECTIVES}`; output path. + +The subagent does NOT see: its own prior `03-submissions.md` or +git history of it; any information about the proposed change +being optimized; existing analyses, tests, or failure data; the +vendored skill source. + +**Why this matters:** the validator (step 8) will later check the +proposed change against this report. If the operator session does +the research, it has already absorbed the optimization context — +the proposed change, prior failures, user's framing — and biases +the report toward documenting upstream conventions in ways that +justify the proposed change. A subagent walled off from that +produces neutral upstream facts; the validator's independence is +preserved. + +The subagent writes the report and returns a brief summary: +license, CLA requirement, branch target, any high-risk rejection +signals from the closed-without-merge PRs. ### (d) Confirm subagent output -Verify the report file exists, the frontmatter parses, and the body -has the expected sections per the subagent prompt template. - -If `requires_cla: true`, mention it explicitly when handing off in -(e) — the operator will need to sign the upstream's CLA before the -PR can be merged, which is operator-side work that has to happen -outside the chain. - -The other frontmatter fields (`license`, `upstream_repo`, -`upstream_branch_target`) are facts in the report that downstream -steps consume directly; they don't need separate operator alerts. +Verify the report exists with parseable frontmatter and the +expected sections. If `requires_cla: true`, mention it explicitly +on handoff — the operator will need to sign the CLA before any PR +can be merged. ### (e) Hand off -Report to the user: the report path, a one-line summary (license / -CLA / branch target / any flagged blockers), then the handoff -message. - -Handoff message, verbatim: - -> The validator in `skill-optimizer-improve-skill` will read this -> report for its external consistency check. If you haven't run -> `skill-optimizer-write-tests` yet, invoke that next. - -The chain doesn't enforce ordering between step 3 and step 4 — they -can run in either order or in parallel. Step 7's validator is what -consumes step 3's output, not step 4. - -## Why limited-context dispatch matters - -The submission-researcher subagent must produce a report that the -validator (step 8) can trust as independent. If the operator session -does the research, it has already absorbed the optimization context -(the proposed change, the prior failures, the user's framing of what -"good" looks like). That context biases the research toward -documenting upstream conventions in a way that justifies the proposed -change — and the validator's external consistency check then has no -real independence. +Report the file path, a one-line summary (license / CLA / branch +target / any flagged blockers), then: -Walling the researcher off in a subagent that sees only the upstream -repo facts (and no information about the proposed change) preserves -the validator's later check. The report is neutral upstream facts; -the validator's job is to check whether the proposed change conforms -to those facts. +> The validator in `skill-optimizer-validate-improvement` will read +> this report for its external consistency check. If you haven't +> run `skill-optimizer-write-tests` yet, invoke that next. -If you find yourself thinking "I'll just run a few `gh` commands -myself, the dispatch is bureaucratic overhead" — that's the failure -mode this chain is built to prevent. Dispatch. +The chain doesn't enforce ordering between step 3 and step 4 — +they can run in either order or in parallel. ## Edge cases -- **Upstream repo is private or requires auth** — the subagent will - surface this as a blocker; ask the user to authenticate `gh` (or - to provide credentials) and re-dispatch. -- **Upstream uses a non-`gh`-friendly host (GitLab, Bitbucket, etc.)** - — the subagent will surface this. The current chain assumes - GitHub-hosted upstreams; non-GitHub cases require manual research - and pasting the report content directly. Tell the user. +- **Upstream repo is private or requires auth** — the subagent + surfaces this as a blocker; ask the user to authenticate `gh` + and re-dispatch. +- **Upstream uses a non-`gh`-friendly host (GitLab, Bitbucket, + etc.)** — the current chain assumes GitHub-hosted upstreams. + Surface this and tell the user; non-GitHub cases require manual + research. - **No merged PRs in the upstream's history yet** — the researcher - can't extract shape patterns from absent data; the report will be - thinner. Surface this honestly rather than making up patterns. -- **Upstream uses unusual / non-discoverable frontmatter - conventions** — if the subagent can't extract a consistent - frontmatter spec from recent merged PRs (the upstream is too - small, or the convention varies wildly), the report's frontmatter - spec section will say so. Tell the user; the optimizer (step 7) - will then have to make a judgment call rather than mechanically - conform. + can't extract shape patterns from absent data; the report will + be thinner. Surface honestly rather than making up patterns. +- **Upstream uses non-discoverable frontmatter conventions** — if + the subagent can't extract a consistent spec from recent merged + PRs, the report will say so. The optimizer (step 7) then has to + make a judgment call rather than mechanically conform. ## Iteration behavior -Step 3 is re-runnable but rarely needs it, and is a fresh-derivation -step. Re-run triggers specific to this step: +Fresh-derivation step. Rarely needs re-running. Re-run triggers: -- The upstream repo updated its `CONTRIBUTING.md`, license, or CLA - requirements -- The upstream's PR-shape conventions visibly shifted (recent merged - PRs no longer match the older patterns) -- Step 1's `01-functionality.md` changed because the source URL was - updated (different upstream repo entirely) +- Upstream repo updated `CONTRIBUTING.md`, license, or CLA +- Upstream's PR-shape conventions visibly shifted +- Step 1's `01-functionality.md` changed because the source URL + was updated (different upstream repo entirely) -General iteration mechanics — staleness detection (git-native), -operator directives, cascading staleness — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index b4535f4..784e4b8 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -10,47 +10,30 @@ from step 1, dispatches a designer subagent to enumerate the skill's responsibilities and propose a ranked set of **functionalities** to test (each functionality = one responsibility the skill must fulfill), then asks the user to pick which ones to actually build -probes for at step 4. Writes a `tests//spec.yaml` -file per proposed functionality (the filesystem IS the state) plus a -one-time audit report at `02-test-proposals.md`. +probes for at step 4. -## Before you start - -Two load-bearing pieces of context to load NOW, before the workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation — for re-run - logic, the constraints subagents must obey, and safe destructive - edits. Workflow step (b) requires it. - -2. **You will dispatch a subagent for the actual design work; you do - NOT enumerate responsibilities or design proposals yourself in - this session.** Workflow step (c) is the dispatch. The rationale - is in "Why limited-context dispatch matters" below — read it if - you're tempted to skip the dispatch. +The filesystem IS the state: this step writes a +`tests//spec.yaml` per proposed functionality +plus a one-time audit report at `02-test-proposals.md`. No +`picked: []` array anywhere; each spec.yaml has its own +`picked: true|false`. Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain (this skill is step 2; step 1 is -`investigate-functionality`; etc.). Internal workflow steps within -THIS skill are labelled "(a)" through "(f)". +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(f)". ## What you produce -Two artifacts at `docs/skill-optimizer//`, where `` -matches the slug from step 1's report: +Two artifacts at `docs/skill-optimizer//`: -1. **`02-test-proposals.md`** — a one-time audit report from the - subagent: the ranked list of proposed functionalities with full - reasoning (why each one matters, what coverage it adds, what - probes are suggested). This file is for human review of the - design reasoning; downstream steps do NOT read it. The subagent - rewrites it on re-runs (fresh top-to-bottom proposal each time, - anchored by the current `tests/` tree state). +1. **`02-test-proposals.md`** — one-time audit report with the + ranked list of proposed functionalities + full reasoning. For + human review of the design reasoning; downstream steps do NOT + read it. The subagent rewrites it on re-runs. 2. **`tests//spec.yaml`** — one folder per - proposed functionality. The folder name is a filesystem-safe - slug derived from the functionality's name. The `spec.yaml` - inside has: + proposed functionality, with frontmatter per + [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): ```yaml name: refuses-malformed-input @@ -66,231 +49,143 @@ matches the slug from step 1's report: the skill's downstream logic. ``` - The filesystem IS the state — there is no `picked: []` array - anywhere. Step 4 builds probes only for functionalities whose - `spec.yaml` has `picked: true`. + Step 4 builds probes only for functionalities whose `spec.yaml` + has `picked: true`. -For the body template of `02-test-proposals.md` (sections, ranking -format) and the exact `spec.yaml` field list, see +For the body template of `02-test-proposals.md` and the exact +`spec.yaml` field list, see [`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md). -The subagent writes both the audit report and the per-functionality -spec.yaml files; the operator session runs the user gate to flip -`picked` values. ## Workflow ### (a) Confirm prerequisites -`docs/skill-optimizer//01-functionality.md` must exist. If it -doesn't, tell the user to run `skill-optimizer-investigate-functionality` -first and stop here. +`01-functionality.md` must exist. If not, tell the user to run +`skill-optimizer-investigate-functionality` first and stop here. If `docs/skill-optimizer//tests/` doesn't exist yet (first -invocation on this slug), create the empty directory — the subagent -will populate it. +invocation on this slug), create the empty directory. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it for this step. Step 2 is a **maintenance** step (see -the protocol's "Step kinds" section) — re-runs read the current -`tests/` tree as state and extend or modify it. Concretely: - -- If the user supplied directives (revisions, new requirements), - collect them as the bulleted `${OPERATOR_DIRECTIVES}` slot per - the protocol's "Collect operator directives" section. Atomic new - requirements only — never a context dump of the current tree. -- If a directive will be destructive (e.g., "remove the X - functionality I no longer think we should test"), the operator - session commits a checkpoint of the current `tests/` tree before - dispatching, per the protocol's "Safe destructive edits" - section. This gives git history a clean before/after breakpoint. - -The subagent will read the current `tests//spec.yaml` -files as input (this is its load-bearing state, per the maintenance -rule). It will NOT walk git history of those files. +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 2 is a **maintenance** step — re-runs read the +current `tests/` tree as state and extend or modify it. Collect +`${OPERATOR_DIRECTIVES}` per the protocol. If a directive will be +destructive (e.g., "remove the X functionality"), commit a +checkpoint before dispatching per the protocol's safe-destructive- +edits section. ### (c) Dispatch the test-case-designer subagent -**Do NOT enumerate responsibilities or design proposals yourself in -this session.** Dispatch the test-case-designer subagent via the -`Agent` tool (with worktree isolation if your environment supports -it). Load the prompt template at -[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md), -substitute the templated inputs (`${FUNCTIONALITY_PATH}`, -`${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, `${OPERATOR_DIRECTIVES}`), -and dispatch. - -The subagent sees: - -- `${FUNCTIONALITY_PATH}` — the current `01-functionality.md` -- `${TESTS_TREE_PATH}` — the current `tests/` tree (may be empty - on first invocation). The subagent reads every existing - `tests//spec.yaml` as load-bearing state per the - maintenance rule. -- `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new - requirements (may be empty) -- Output paths for `02-test-proposals.md` and the `tests/` tree - -The subagent does NOT see: - -- The skill's source content (would gerrymander tests around the - source's literal phrasing — design coverage from STATED - responsibilities) -- Any `06-analysis.md`, `07-improvement-proposal.md`, - `08-validator-verdict.md` -- Raw failure data, `findings.txt`, bench results -- Git history of any tree file or its own audit report +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT enumerate responsibilities or design +proposals yourself in this session** — dispatch the subagent via +the `Agent` tool. Load the prompt template at +[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md) +and substitute `${FUNCTIONALITY_PATH}`, `${TESTS_TREE_PATH}`, +`${PROPOSALS_PATH}`, `${OPERATOR_DIRECTIVES}`. + +The subagent sees: `01-functionality.md`, the current `tests/` +tree (load-bearing state per the maintenance rule), +`${OPERATOR_DIRECTIVES}`, output paths. + +The subagent does NOT see: the skill's source content (would +gerrymander tests around the source's literal phrasing — +design from STATED responsibilities); `06-analysis.md`, +`07-improvement-proposal.md`, `08-validator-verdict.md`; raw +failure data, bench results; git history of any tree file. + +**Why this matters:** the operator session has absorbed prior +failure data, prior optimizer attempts, validator verdicts — that +context biases test design toward "what just failed" rather than +"what comprehensively covers responsibilities". The subagent walled +off from that produces coverage-oriented proposals, not +regression-defensive ones. The subagent writes: -- `02-test-proposals.md` (the audit report, full ranked design - reasoning) +- `02-test-proposals.md` (the audit report) - One `tests//spec.yaml` per proposed functionality. Existing spec.yaml files the user has already - edited (e.g., `picked: true` set, custom suggested_probes added) - are preserved verbatim unless a directive explicitly targets - that functionality for revision. + edited (e.g., `picked: true` set) are preserved verbatim unless + a directive explicitly targets that functionality. -It returns a brief summary: top-3 functionalities by importance, -list of which existing functionalities were preserved unchanged vs -revised, list of newly-added functionalities. +Returns a brief summary: top-3 functionalities by importance, +preserved-unchanged list, newly-added list. ### (d) Confirm subagent output -Verify: - -- `02-test-proposals.md` exists and parses as markdown. -- Each `tests//spec.yaml` parses as YAML and - has the required fields (`name`, `description`, `picked`, - `importance`, `suggested_probes`, `why_test`). - -If any spec.yaml fails to parse or is missing required fields, -surface the issue to the user; don't try to repair the subagent's -output yourself. +Verify `02-test-proposals.md` parses and each new +`tests//spec.yaml` parses with required fields. +If any spec.yaml fails to parse, surface to the user; don't repair +the subagent's output yourself. ### (e) User gate: present proposals, collect picks -Show the user the ranked list from `02-test-proposals.md` (paste -the audit report, or summarize if long — your call). Tell them -how to indicate picks: - -> Here are the proposed functionalities ranked by importance. To -> pick the ones you want me to build probes for at step 4, edit -> the `picked:` field in each `tests//spec.yaml` -> to `true`. You can also edit `suggested_probes` if you want to -> change the probe set, or tell me to dispatch a revised proposal. +Show the user the ranked list from `02-test-proposals.md` and tell +them how to indicate picks: edit `picked: true|false` in each +`tests//spec.yaml`. They can also edit +`suggested_probes` or ask for a revised proposal. Three realistic responses: -1. **User edits `picked: true|false` in the spec.yaml files - directly.** Read each spec.yaml after their edits to confirm - the picked set. Commit the final state of the tree. -2. **User wants additions or revisions** (e.g., "add coverage for - null-input edge cases", "the X functionality is too broad — - split it into two"). Treat their feedback as new operator - directives. Return to (b) to apply the iteration protocol - (which checkpoints the current state via git commit if the - change is destructive) and re-dispatch the subagent at (c). - Because this is a maintenance step, existing spec.yaml files - the user has already edited (especially `picked: true` ones) - are preserved unless their directive explicitly targets a - specific functionality. +1. **User edits `picked` directly (or asks you to flip specific + ones).** Read each spec.yaml to confirm the picked set. If the + user explicitly says "flip these to true", do it and confirm + what you set. The rule is "no proactive flipping without user + direction", not "user must edit every file by hand". +2. **User wants additions or revisions** (e.g., "split X into + two", "add a functionality for Y"). Treat their feedback as + directives. Return to (b) and re-dispatch the subagent at (c). + Existing edited spec.yaml files are preserved per the + maintenance rule. 3. **User picks zero.** Ask whether they want a revised proposal - (treat as case 2) or are abandoning the optimization for this - skill (exit honestly without progressing to step 4). + (case 2) or are abandoning the optimization for this skill + (exit honestly). -Don't auto-flip `picked` proactively — even if all functionalities -look important, the user owns the decision (they're paying for the -probe-building in step 4 and the bench run in step 5). But if the -user explicitly says "flip these to true" or "pick X, Y, Z", do -that — edit the named spec.yaml files for them and confirm what -you set. The rule is "no flipping without explicit user direction", -not "user must edit every file by hand". +Don't auto-flip `picked` on the user's behalf. Even if all +functionalities look important, the user owns the decision (they're +paying for probe-building in step 4 and the bench run in step 5). ### (f) Hand off -Read `01-functionality.md`'s `pr_submission_intent` field. - -If `pr_submission_intent: true`: - -> Next, invoke `skill-optimizer-investigate-submissions`. After -> that, invoke `skill-optimizer-write-tests` to build probes for -> the picked functionalities. - -If `pr_submission_intent: false`: - -> Skipping step 3 (no PR submission planned). Next, invoke -> `skill-optimizer-write-tests` to build probes for the picked -> functionalities. - -No late "submit a PR?" prompts — the decision was made at step 1. - -## Why limited-context dispatch matters - -If the operator session is the one designing tests, it has already -absorbed prior failure data, the optimizer's past attempts, the -validator's verdicts, and whatever framing the user has applied. -That context biases test design toward "tests that would have -caught the things I just watched fail" — the coverage version of -ducktape. You end up with a test suite that validates the patch -instead of testing the skill independently. - -A subagent that sees only `01-functionality.md` and the current -`tests/` tree reasons from the skill's stated responsibilities, -not from prior outcomes. The proposal is coverage-oriented, not -regression-defensive. That's the invariant we need for a fair -test set. - -If you find yourself thinking "I'll just propose the -functionalities myself, the subagent dispatch is bureaucratic -overhead" — that's the failure mode this chain is built to -prevent. Dispatch. +Read `01-functionality.md`'s `pr_submission_intent` field. If +`true`, hand off to `skill-optimizer-investigate-submissions` then +`skill-optimizer-write-tests`. If `false`, skip step 3 and hand +off directly to `skill-optimizer-write-tests`. No late "submit a +PR?" prompts — the decision was made at step 1. ## Edge cases - **`01-functionality.md` missing** — tell the user to run step 1 first; don't try to derive responsibilities yourself. -- **Subagent's proposal has fewer functionalities than expected** - — signal the functionality report is thin. Surface to the user - with two options: re-run step 1 with directives ("the report - missed responsibility X") or accept the small proposal if the - skill is genuinely small. Do not re-invoke step 1 yourself — - re-runs require an active signal from the user (per the iteration - protocol's "Re-run authorization" section). +- **Subagent's proposal has fewer functionalities than expected** — + signal the functionality report is thin. Surface to the user + with two options: re-run step 1 with directives, or accept if + the skill is genuinely small. Per the no-auto-invocation rule in + [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), + don't re-invoke step 1 yourself. - **User wants to add a functionality the subagent didn't propose** - — treat their description as an operator directive ("add a - functionality for X") and re-dispatch the subagent per step - (e)(2). Don't write the spec.yaml manually yourself — the - subagent is the only writer of test-design content in this step, - matching the architecture's no-operator-generative-writing rule. - The user can flip the new functionality's `picked: true` after - the re-dispatch, or ask you to. -- **Operator directives contradict each other** (e.g., "focus - coverage on X" + "ignore X") — surface the contradiction to the - user before re-dispatching; don't try to resolve it yourself. -- **User removes a functionality folder manually** — fine, that's - filesystem-as-state working as intended. On next re-run, the - subagent reads the current tree and doesn't propose the removed - one back unless directives ask. + — treat their description as a directive and re-dispatch per + step (e)(2). Don't write the spec.yaml manually yourself — the + subagent is the only writer of test-design content. +- **Operator directives contradict each other** — surface the + contradiction before re-dispatching. +- **User removes a functionality folder manually** — fine, + filesystem-as-state working as intended. ## Iteration behavior -Step 2 is a maintenance step. Re-run triggers specific to this -step: - -- The user wants different coverage (different proposed - functionalities, different ranking, different probe suggestions) -- Step 4 (`write-tests`) flagged a picked functionality as - unimplementable -- Step 5 (`run-bench`) showed all probes pass on baseline — too - easy, need a different cut of coverage -- Step 6 (`analyze-result`) surfaced a coverage gap (an - unrepresented responsibility; user wants one tested) -- Step 1's `01-functionality.md` changed (functionality - understanding updated) - -General iteration mechanics — staleness detection, operator -directives, cascading staleness, safe destructive edits — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. +Maintenance step. Re-run triggers: + +- User wants different coverage (different proposals, ranking, or + probe suggestions) +- Step 4 flagged a picked functionality as unimplementable +- Step 5 showed all probes pass on baseline (too easy) +- Step 6 surfaced a coverage gap +- Step 1's `01-functionality.md` changed + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index d57dc5c..6f96ec3 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -1,49 +1,33 @@ --- name: skill-optimizer-run-bench -description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once a `workbench/` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. +description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once `tests/suite.yml` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. --- # skill-optimizer-run-bench -Step 5 of the skill-optimizer chain. Invokes the skill-optimizer CLI's -`run-suite` command against `tests/suite.yml` (generated by step 4), -captures the raw results under a timestamped directory, and writes a -small summary report that step 6 (`analyze-result`) reads. No subagent +Step 5 of the skill-optimizer chain. Invokes the skill-optimizer +CLI's `run-suite` command against `tests/suite.yml` (generated by +step 4), captures the raw results under a timestamped directory, and +writes a small summary report that step 6 reads. No subagent dispatch — this is a thin operator-driven CLI step. -## Before you start - -One load-bearing piece of context to load NOW, before the workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation. Step 5 is a - slightly unusual case because the bench's raw output is - timestamped (each run preserved naturally under its own directory) - — outside the iteration protocol. The summary report - `05-bench-summary.md` is a fresh-derivation artifact per the - protocol: each invocation overwrites the canonical with a fresh - write-up of the new raw run; prior state lives in git history. - Workflow step (b) below requires it. - -Throughout this document, "step 1" through "step 9" (no parens) refer -to skills in the chain (this skill is step 5; `write-tests` is step 4; -`analyze-result` is step 6). Internal workflow steps within THIS skill -are labelled "(a)" through "(e)" to avoid the collision. +Throughout this document, "step 1" through "step 9" (no parens) +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(e)". ## What you produce -Two outputs at `docs/skill-optimizer//`, where `` matches -the slug from step 1's report: +Two outputs at `docs/skill-optimizer//`: -1. **`05-bench-results//`** — the raw run output from the - CLI: `suite-result.json`, per-trial `trace.jsonl`, per-trial - `findings.txt`, and any preserved workspaces. Timestamped per the - invocation; old runs are NEVER overwritten. Each timestamped - directory is its own naturally-accumulated archive — outside the - iteration protocol. +1. **`05-bench-results//`** — raw CLI output: + `suite-result.json`, per-trial `trace.jsonl`, per-trial + `findings.txt`, preserved workspaces. Timestamped per + invocation; old runs are NEVER overwritten. Outside the + iteration protocol — each timestamped dir is its own + naturally-accumulated archive. -2. **`05-bench-summary.md`** — a single canonical report with - frontmatter: +2. **`05-bench-summary.md`** — single canonical aggregate, per + [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): ```yaml --- @@ -52,58 +36,35 @@ the slug from step 1's report: --- ``` - The body is a small aggregate that step 6 reads as the entry point: - the per-model pass rates, per-probe pass rates, a list of probes - where at least one trial failed (so the analyzer knows where to - focus), and a pointer to the timestamped raw directory. The - summary does NOT include trace excerpts or findings detail — - those stay in the raw output and the analyzer reads them directly - when it needs them. Prior versions of the summary live in git - history. + Body has per-model pass rates, per-probe pass rates, a + failed-probe pointer list, pointer to the timestamped raw + directory. Trace excerpts and findings detail stay in the raw + output; the summary is an aggregate, not a dump. ## Workflow ### (a) Confirm prerequisites -`docs/skill-optimizer//tests/suite.yml` must exist (per step -4's output contract). If it's missing, tell the user to complete -step 4 first and stop here. - -Also confirm the vendored skill source exists at -`docs/skill-optimizer//vendored-skill/` (per step 1's setup); -the workbench's `suite.yml` points at it as the skill under test. +`tests/suite.yml` must exist (per step 4's output contract). +`vendored-skill/` should exist for upstream skills. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it. Step 5's behavior: +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Two pieces: -- For the **raw bench output** at `05-bench-results//`: - always write to a fresh timestamped directory (`date +%Y%m%dT%H%M%SZ` - in UTC). Each run is its own preserved artifact — outside the - iteration protocol. Never overwrite a prior timestamped dir. +- **Raw bench output** at `05-bench-results//`: always + write to a fresh timestamped directory (`date -u +%Y%m%dT%H%M%SZ`). + Outside the iteration protocol; never overwrite. -- For the **summary file** at `05-bench-summary.md`: it's a - fresh-derivation artifact. Each invocation overwrites the - canonical with a mechanical write-up of the new raw run; prior - state lives in git history. No `archive/` involvement. - - Staleness check (git-native): if `tests/suite.yml` has a newer - git mtime than `05-bench-summary.md`, the existing summary - measured a different workbench and is stale. If the user also - didn't pass a re-run directive AND the summary appears current - (its git mtime is newer than `tests/`), tell them the latest - results are still good and exit honestly. Otherwise proceed - with a new bench run. - -The summary's anti-ducktape principle: write what THIS run shows, -not what the prior summary said. There is no subagent, but the -discipline is the same. +- **Summary file** at `05-bench-summary.md`: fresh-derivation + artifact per the protocol. Staleness check: if `tests/suite.yml` + has a newer git mtime than `05-bench-summary.md`, the existing + summary measured a different workbench and is stale. If summary + is current AND no re-run directive, tell the user and exit. ### (c) Run the bench -Invoke the skill-optimizer CLI directly: - ```bash TIMESTAMP=$(date -u +%Y%m%dT%H%M%SZ) OUT_DIR="docs/skill-optimizer//05-bench-results/${TIMESTAMP}" @@ -115,119 +76,79 @@ npx tsx /src/cli.ts run-suite \ --trials 3 ``` -`--trials 3` is the chain's default — enough trials to distinguish -flaky (1-of-3) from systematic (2-of-3 or 3-of-3) failures at step 6. -The user may request a different trial count; honor it. Models come -from `suite.yml` (per the project invariant: `run-suite` does NOT -take a `--models` override). - -The CLI requires `OPENROUTER_API_KEY` to be set for real model runs. -If the user hasn't set it, the CLI will fail fast — surface the -error to the user; don't try to recover. - -Stream the CLI's stdout/stderr to the user so they can see progress -(this is a long-running step; bench runs of meaningful size take -minutes to hours). +`--trials 3` is the chain default — enough to distinguish flaky +from systematic failures at step 6. Honor a different count if +requested. Models come from `suite.yml` (per project invariant: +`run-suite` does NOT take a `--models` override). The CLI requires +`OPENROUTER_API_KEY`; surface env failures rather than recovering. +Stream stdout/stderr to the user — bench runs take minutes to +hours. ### (d) Write the summary -When the CLI completes, parse `${OUT_DIR}/suite-result.json` and -write `05-bench-summary.md` with the frontmatter from "What you -produce" plus a body containing: +Parse `${OUT_DIR}/suite-result.json` and write `05-bench-summary.md` +with the frontmatter above plus a body containing: - **Overall:** total trials, passed, failed, overall pass rate. -- **Per model:** for each model in `suite.yml`, the trial count and - pass rate. -- **Per probe:** for each probe (one entry per - `tests///` in the suite), the trial count, - pass rate, and a one-line note if at least one trial failed (so - step 6 knows where to look). -- **Failed-probe pointer list:** explicit list of probe IDs where - any trial failed, with paths to their `trace.jsonl` and - `findings.txt` under `${OUT_DIR}`. -- **Raw output:** the value of `bench_results_path` from the - frontmatter (relative path to the timestamped dir). - -The summary is the entry point step 6 reads. Step 6's analyzer -subagent will follow the failed-probe pointer list into the raw -output for trace and findings detail; the summary itself stays -short (an aggregate, not a dump). +- **Per model:** trial count and pass rate per model in + `suite.yml`. +- **Per probe:** trial count, pass rate, one-line note if any + trial failed. +- **Failed-probe pointer list:** probe IDs where any trial failed, + with paths to their `trace.jsonl` and `findings.txt`. +- **Raw output:** the `bench_results_path` value. ### (e) Hand off -Read the overall pass rate from the summary you just wrote. Two -handoff messages depending on the result: - -If `overall_pass_rate < 1.0`: - -> Bench complete. Overall pass rate: `%`. Some trials failed — -> invoke `skill-optimizer-analyze-result` to diagnose whether the -> failures point to a structural weakness in the skill. - -If `overall_pass_rate == 1.0`: - -> Bench complete. All trials passed against the current skill. Two -> realistic paths from here: (1) accept that the picked test set -> doesn't expose any weakness in the current skill (exit honestly); -> (2) revise the test set to be harder — re-run step 2 with a -> directive like "the prior cases were too easy; propose harder -> ones" and walk the chain forward again. +Read `overall_pass_rate`. Two messages: -Don't auto-invoke step 6 in the all-pass case — there's nothing for -it to analyze, and the chain skills never invoke each other on their -own. Surface the choice to the user. +- `< 1.0`: invoke `skill-optimizer-analyze-result` to diagnose + failures. +- `== 1.0`: surface the choice — accept that probes don't expose a + weakness, or re-run step 2 with a "make probes harder" + directive. Don't auto-invoke step 6; per the no-auto-invocation + rule in + [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), + surface the choice. ## Edge cases -- **`tests/suite.yml` missing or `tests/` empty** — tell the user - to run step 4 first; don't try to construct a suite manifest - yourself. -- **`OPENROUTER_API_KEY` not set** — the CLI fails fast; surface the - error and tell the user to set the env var. Don't fall back to a - mock or skip the bench. -- **Docker image missing** — the CLI's default image is - `skill-optimizer-workbench:local`. If it's not built, tell the - user to run `docker build -t skill-optimizer-workbench:local -f - docker/workbench-runner.Dockerfile .` from the repo root. -- **All trials in `suite-result.json` errored (no graded results)** - — likely a workbench misconfiguration or environmental failure, - not a skill weakness. Write the summary honestly (overall pass - rate undefined, error reasons recorded) and surface the situation - to the user before handing off to step 6 — they should decide - whether to fix the workbench (back to step 4 with directives) or - treat this as the skill itself crashing the harness. +- **`tests/suite.yml` missing** — tell the user to run step 4 + first; don't construct a suite manifest yourself. +- **`OPENROUTER_API_KEY` not set** — CLI fails fast; surface and + ask the user to set it. Don't mock. +- **Docker image missing** — default is + `skill-optimizer-workbench:local`. Tell the user to build it + (`docker build -t skill-optimizer-workbench:local -f + docker/workbench-runner.Dockerfile .`). +- **All trials errored (no graded results)** — workbench + misconfiguration or environmental failure, not a skill weakness. + Write the summary honestly and surface before handing off to + step 6. ## Iteration behavior -Step 5 is re-runnable. Re-run triggers specific to this step: +Re-run triggers: -- `tests/` changed (step 4 added or revised probes) — the prior - bench measured a different probe set, detected via git mtime -- User wants fresh results against the same probes (e.g., - flakiness suspicion, model availability changed, model list in +- `tests/` changed (step 4 added or revised probes) — detected via + git mtime +- User wants fresh trial data (flakiness suspicion, model list in `suite.yml` changed) -- Step 6 (`analyze-result`) flagged a result as inconclusive due - to too few trials — user may re-run with higher `--trials` +- Step 6 flagged a result as inconclusive due to too few trials — + re-run with higher `--trials` -By default, a re-run measures the **entire** probe set — even -probes that didn't change at step 4. This gives unchanged probes -fresh trial samples (useful for distinguishing flaky from systematic -at step 6) and keeps the summary internally comparable. The cost is -"all probes × trials" per re-run. +By default a re-run measures the **entire** probe set. Unchanged +probes get fresh trial samples (useful for distinguishing flaky vs. +systematic at step 6), and the summary stays internally comparable. +Cost is "all probes × trials" per re-run. **Known limitation (deferred):** the CLI's `run-suite` does not -currently accept a case filter, so there is no first-class -"re-bench only the changed probes" mode at step 5. Operators with -expensive suites can run `run-case` manually for changed probes and -splice results into the prior `05-bench-results//` dir, but -that's an outside-the-chain escape hatch, not a supported flow. -Adding a case filter to `run-suite` (and threading it through step 5) -is on the workbench's roadmap; until then, expect re-runs to -remeasure the full probe set. - -General iteration mechanics — git-native staleness, operator -directives, cascading staleness — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -The raw bench output is outside the protocol (each timestamped -directory is its own preservation); the summary is a -fresh-derivation artifact and follows the protocol. +accept a case filter, so there's no first-class partial-rebench +mode. Operators with expensive suites can run `run-case` manually +for changed probes and splice into the prior `05-bench-results//` +dir — outside-the-chain escape hatch, not a supported flow. + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. Raw bench output is outside the +protocol; the summary is fresh-derivation. diff --git a/skills/skill-optimizer-shared/frontmatter-discipline.md b/skills/skill-optimizer-shared/frontmatter-discipline.md new file mode 100644 index 0000000..29a78c1 --- /dev/null +++ b/skills/skill-optimizer-shared/frontmatter-discipline.md @@ -0,0 +1,39 @@ +# Frontmatter discipline + +Every chain artifact's frontmatter follows one rule: **runtime-relevant +facts only, never history.** Git already content-addresses every prior +state and `git log` gives you the timeline; reimplementing either in +frontmatter is bookkeeping for its own sake — it adds maintenance +surface (bumping, syncing, risk of drift) without enabling anything +git can't already do. + +## Do not add + +- `version:` field (use git) +- `archive/` subdirectory (use git) +- `inputs.step_N: ` lineage trackers (use git mtime comparison + per the iteration protocol) +- `last_derived_at:` timestamps (use git log) +- Any "I'm tracking how many times this file has been written" + metadata + +## Keep + +Frontmatter fields a chain skill needs to READ at runtime to do its +job: + +- `picked: true|false` (B2 functionality spec) +- `pr_submission_intent: true|false` (B1 functionality report) +- `classification: tool-use` (B1 functionality report) +- `has_structural_weakness: true|false` (B6 analysis — gates B7) +- `verdict: approve|needs-revision|reject` (B8 verdict — gates + materialization) +- `overall_pass_rate` (B5 summary — gates B6 dispatch) + +## Decision aid + +When in doubt about whether a field belongs: ask whether a chain +skill needs to READ it to do its job right now. + +- Yes → keep +- Recording for future debugging / audit → that's git's job, drop it diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index d352d60..8dab53a 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -3,59 +3,34 @@ **Load this file every time** a skill-optimizer chain skill executes its "Handle iteration" step. The mechanics are identical across every skill in the chain; the parent skill's SKILL.md only specifies -that-skill's slot in the chain. The actual re-run logic + subagent -constraints live here so the chain behaves consistently. +that-skill's slot in the chain. The actual re-run logic + step-kind +classification + destructive-edit safety live here so the chain +behaves consistently. -If you're a chain skill's SKILL.md and you're tempted to inline this -logic instead of reading this file: don't. Read this file, do what -it says. +This file covers **iteration mechanics only** — when to re-run, how +to detect staleness, how to handle accumulating state vs. +fresh-derivation, and how to checkpoint before destructive edits. -## What this protocol governs +Related shared docs (load lazily at the workflow step that needs +them): -Every chain step produces a canonical artifact under -`docs/skill-optimizer//`. Most are single-file reports -(`01-functionality.md`, `06-analysis.md`, etc.); two are -filesystem-as-state trees (`tests//spec.yaml` for -step 2, `tests///{spec.yaml, workspace, grader, -smoke}` for step 4); one is a timestamped collection -(`05-bench-results//` for step 5's raw output). - -This protocol determines, on every step invocation: - -1. Whether the existing artifact is current or stale -2. What the subagent is and isn't allowed to look at -3. How the operator handles destructive edits safely - -History is **git**. Do not add version-tracking metadata to any -chain artifact — no `version:` field, no `archive/` subdirectory, -no manually-bumped counters, no `inputs.step_N:` lineage trackers, -no `last_derived_at:` timestamps. The single rule: frontmatter -carries **runtime-relevant facts** (what this artifact represents, -what state it's in — e.g., `picked: true`, `pr_submission_intent: -true`, `classification: tool-use`), not **history** (when it was -written, how many times, what upstream version produced it). Git -already content-addresses every prior state and `git log` gives you -the timeline; reimplementing either in frontmatter is bookkeeping -for its own sake — it adds maintenance surface (bumping, syncing, -risk of drift) without enabling anything git can't already do. - -When in doubt about whether a field belongs: ask whether a chain -skill needs to READ it to do its job right now (yes → keep), or -whether you're recording it for future-debugging / future-audit -purposes (no → that's git's job). +- [`subagent-dispatch.md`](subagent-dispatch.md) — what subagents + see / don't see, operator-directive concept, dispatch input + templating, no-auto-invocation rule +- [`frontmatter-discipline.md`](frontmatter-discipline.md) — what + belongs in frontmatter (runtime facts) vs. what doesn't (history; + that's git's job) ## Step kinds: fresh-derivation vs maintenance -Steps in the chain come in two kinds, and the subagent-constraint -rule differs between them. +Steps in the chain come in two kinds. The subagent's reading rule +differs between them (full details in +[`subagent-dispatch.md`](subagent-dispatch.md)); for iteration +purposes the difference is whether the canonical state accumulates +across re-runs or is regenerated. -**Fresh-derivation steps** produce their artifact from upstream -inputs + directives, with no continuity from prior outputs. The -subagent does NOT read its own canonical file, and does NOT read -git history of it. Every invocation is a derivation from scratch. -This is the anti-ducktape rule: the optimizer must not see prior -optimization attempts, the analyzer must not see prior analyses, -the researcher must not rationalize a prior report. +**Fresh-derivation steps** produce their canonical artifact from +upstream + directives. Re-runs overwrite; git captures prior state. | Step | Why fresh derivation | |---|---| @@ -68,88 +43,19 @@ the researcher must not rationalize a prior report. **Maintenance steps** manage an accumulating tree on disk — the filesystem itself is the state. Re-runs read the current tree and -extend or modify it. The subagent DOES read the current canonical -files (because they are the state) and produces a modified tree. +extend or modify it. | Step | What accumulates | |---|---| | 2. investigate-test-case | `tests//spec.yaml` grows as functionalities are added/refined | | 4. write-tests | `tests///` probes grow as the user adds coverage | -The git-history-as-archive principle covers both: prior states of -maintenance steps are recoverable from `git log`, and prior states -of fresh-derivation steps are the same. The difference is purely -what the subagent is allowed to look at. - -## Subagent constraints - -Three rules every reasoning subagent must follow: - -1. **Read upstream artifacts only at their current state.** The - filesystem is the source of truth; do not walk `git log` looking - for prior versions of upstream files. - -2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7, 8): do - NOT read your own canonical file, and do NOT read git history - of it.** Your job is to derive a new artifact from upstream + the - operator's directives, without direct access to prior - derivations of this same artifact. This is the load-bearing - anti-ducktape constraint. - - The directives ARE the channel for lessons learned from prior - iterations. The operator session CAN read the prior output — - that's part of its job between iterations — and distills any - lessons into atomic new requirements ("the prior optimization - was too destructive on the description field; this attempt - should be additive only"). The subagent then satisfies those - distilled requirements as fresh constraints, without seeing the - raw prior content. This separation is what keeps the new - derivation from rationalizing the prior one while still letting - the chain converge: each iteration starts cleaner than the - last, with the operator's distilled lessons as guardrails. - - The operator session writes your output by overwriting the - canonical file; git captures the prior state automatically, but - neither you nor any subsequent fresh-derivation subagent should - walk that history. - - **For step 7 and step 8 specifically:** the SKILL CONTENT (the - target being improved) is upstream input, not your own - canonical. Both the optimizer (step 7) and the validator (step - 8) read the **current state of the skill** — which is - `improved-skill/` if it exists (the accumulated state from - prior step-8 approvals), else the original source - (`vendored-skill/` for upstream, the local file path for - local). Neither step modifies the original source: the - accumulated improvements live in `improved-skill/` and are - materialized by step 8 on `verdict: approve`; the original - (vendored or local) stays frozen as reference. Each subagent's - "own canonical" that's off-limits is its report - (`07-improvement-proposal.md` for the optimizer; - `08-validator-verdict.md` for the validator), not the skill - content itself. The skill content accumulates improvements - across iterations in `improved-skill/`; the proposal and - verdict reports do not. - -3. **For maintenance steps (2, 4): DO read your own canonical - tree** (when it exists). Your job on a re-run is to extend or - modify the current state per `${OPERATOR_DIRECTIVES}`, - preserving entries the user has invested in (picked - functionalities, built probes, existing graders) unless a - directive explicitly says to revise a specific entry. Do NOT - walk git history of the tree — the current state is what - matters; prior states are inert. - -Each subagent prompt template enforces the appropriate rule for -its step. - -## Staleness detection +## Staleness detection (git-native) Whether an artifact is stale is determined by comparing git -modification times against upstream files: +modification times against direct upstream: ```bash -# Is 02-test-case stale relative to 01-functionality? UPSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//01-functionality.md) DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) [ "$UPSTREAM_T" -gt "$DOWNSTREAM_T" ] && echo "stale" @@ -157,95 +63,34 @@ DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) In interactive use, the operator typically just knows ("I re-ran step 1, so step 2 needs a re-run"). The git-mtime check is for -auto-pilot (step 8), which walks the chain forward and re-runs any -downstream that's older than its direct upstream. +auto-pilot (step 9), which walks the chain forward and re-runs any +downstream older than its direct upstream. For maintenance steps, the operator's directives can also force a re-run even when no upstream changed ("add a probe for null -input"). The protocol does not require a stale upstream to permit -a re-run; it only flags staleness to recommend one. - -## Collect operator directives - -`${OPERATOR_DIRECTIVES}` is a short bulleted list of **atomic new -requirements** that surfaced from prior iterations or from the -user's current request. It is **never** a context dump of prior -artifact content; it is a list of specific asks that the new -derivation must satisfy on top of its normal inputs. - -Examples that count as atomic requirements: - -- "user requested coverage for null inputs" -- "user flagged responsibility X as under-tested in the prior pass" -- "focus on the gpt-5 cluster — gemini and claude both passed" -- "user wants the report to mention the vendor's CLA requirement" - -Examples that do NOT count (these are context dumps, not atomic -requirements — reject the temptation): - -- The full text of a prior canonical file pasted in so the - subagent can "see what we already had" -- "Here's what failed last time:" followed by failure data -- The validator's prior verdict pasted in for the subagent to read - -If the user supplied no new requirements and you're re-running -because an upstream artifact changed, leave directives empty. +input"). Staleness flags a recommended re-run; it doesn't require +one. ## Safe destructive edits When a maintenance step's re-run will overwrite or delete existing -content (e.g., step 4 rebuilds a probe with `rm -rf -tests///workspace/`; step 2 removes a de-picked +content (e.g., step 4 rebuilds a probe; step 2 removes a de-picked functionality), the operator session commits the current state -**before** dispatching the destructive change. This gives git -history a clean before/after breakpoint: +**before** dispatching the destructive change: ```bash -# Before destructive change: git add -A docs/skill-optimizer// git commit -m "checkpoint: before rebuild of " -# Then dispatch the destructive change. ``` +This gives git history a clean before/after breakpoint. Without this checkpoint, the prior state of a rebuilt probe (or de-picked functionality) gets buried inside a multi-file commit -later, making "show me what this probe used to look like" awkward -to recover. +later, making "show me what this used to look like" awkward. Fresh-derivation steps are also destructive (they overwrite the canonical file), but the prior state is already self-contained in -its own commit (the previous fresh-derivation run's commit), so -the checkpoint isn't needed. - -## Dispatch the subagent - -When you invoke the subagent for this step, pass these templated -inputs (each subagent prompt template names them with `${...}` -placeholders): - -- `${OPERATOR_DIRECTIVES}` — the bulleted list (may be empty) -- Step-specific inputs (paths to upstream artifacts the subagent - reads, the output path, etc.) per that step's SKILL.md - -## Re-run authorization - -A chain skill never invokes another chain skill on its own. When a -step finds that its upstream input is unsatisfactory — a thin -functionality report, a too-easy bench, a missing-but-needed -probe, a wrong-target PR-conventions report — it surfaces the -finding to the user and stops. The user (or the auto-pilot driver -applying its default policies) decides whether to re-run an -upstream step, accept the situation, or abandon the run. - -This applies to backward triggers specifically. Forward handoffs -(step N hands off to step N+1) are part of the chain's normal -flow and the handing-off skill emits the handoff message; the -user or auto-pilot acts on it. - -Why this rule exists: chain skills running on user request must -give the user control over what they're paying for. Re-running an -upstream step takes time and tokens; the user authorizes that -explicitly, not the agent. +its own prior commit, so no checkpoint is needed. ## Cascading staleness @@ -271,10 +116,10 @@ needed — git is the archive. ## What this protocol does NOT cover - **Step 5's raw bench results** are timestamped under - `05-bench-results//`. Each run preserves naturally as its - own directory. The summary file `05-bench-summary.md` is a - fresh-derivation artifact per this protocol; the timestamped - raw output is outside. + `05-bench-results//`. Each run preserves naturally as its own + directory. The summary file `05-bench-summary.md` is a + fresh-derivation artifact per this protocol; the timestamped raw + output is outside. - **Auto-pilot's summary report** at `docs/skill-optimizer//autopilot-summary-.md` is timestamped per run. Each auto-pilot invocation produces a fresh diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md new file mode 100644 index 0000000..70f1489 --- /dev/null +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -0,0 +1,122 @@ +# Subagent dispatch architecture + +The skill-optimizer chain runs most of its generative work in +**limited-context subagents** rather than the operator session. This +file documents the dispatch architecture every chain skill follows: +what the subagent sees, what it doesn't, how the operator session +parameterizes the dispatch, and the no-auto-invocation rule that +keeps chain skills composable. + +If you're a chain skill's SKILL.md, the dispatch step in your +workflow should reference this file rather than re-explaining the +constraints — load it lazily at the dispatch step. + +## Why dispatch at all + +If the operator session does the generative work itself, it inherits +everything that's been in the conversation: the user's framing of +the problem, prior outputs the user has reviewed, the operator's +own thinking-out-loud. That context biases the work toward "what we +just talked about" — the textbook ducktape pattern: a patch that +addresses the symptom the user just described, not the underlying +weakness. + +A subagent dispatched with only the inputs it needs has none of that +contamination. It produces a fresh derivation from the load-bearing +inputs + atomic directives. The operator session then orchestrates +the result, gates against the user, and writes the final artifact. + +## Subagent constraints + +Three rules every reasoning subagent must follow: + +1. **Read upstream artifacts only at their current state.** The + filesystem is the source of truth; do not walk `git log` looking + for prior versions of upstream files. + +2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7, 8): do NOT + read your own canonical file, and do NOT read git history of + it.** Each invocation derives a new artifact from upstream + the + operator's directives, without direct access to prior + derivations of this same artifact. This is the load-bearing + anti-ducktape constraint. + + The directives ARE the channel for lessons learned from prior + iterations. The operator session CAN read the prior output — + that's part of its job between iterations — and distills any + lessons into atomic new requirements. The subagent satisfies + those distilled requirements as fresh constraints, without + seeing the raw prior content. This separation keeps the new + derivation from rationalizing the prior one while still letting + the chain converge. + + **For step 7 and step 8 specifically:** the SKILL CONTENT (the + target being improved) is upstream input, not your own + canonical. Both the optimizer (step 7) and the validator (step + 8) read the **current state of the skill** — which is + `improved-skill/` if it exists (the accumulated state from + prior step-8 approvals), else the original source. Neither step + modifies the original. The "own canonical" off-limits to each + subagent is its report (`07-improvement-proposal.md` for the + optimizer; `08-validator-verdict.md` for the validator), not + the skill content itself. + +3. **For maintenance steps (2, 4): DO read your own canonical + tree** (when it exists). Your job on a re-run is to extend or + modify the current state per `${OPERATOR_DIRECTIVES}`, + preserving entries the user has invested in unless a directive + explicitly says to revise. Do NOT walk git history of the tree + — the current state is what matters; prior states are inert. + +## Collect operator directives + +`${OPERATOR_DIRECTIVES}` is a short bulleted list of **atomic new +requirements** that surfaced from prior iterations or the user's +current request. **Never** a context dump of prior artifact content +— a list of specific asks that the new derivation must satisfy on +top of its normal inputs. + +Examples that count as atomic requirements: + +- "user requested coverage for null inputs" +- "focus on the gpt-5 cluster — gemini and claude both passed" + +Examples that do NOT count (these are context dumps — reject the +temptation): + +- The full text of a prior canonical file pasted in for the + subagent to "see what we already had" +- The validator's prior verdict pasted in for the subagent to read + +If the user supplied no new requirements and you're re-running +because an upstream artifact changed, leave directives empty. + +## Dispatch the subagent + +When you invoke the subagent for this step, pass these templated +inputs (each subagent prompt template names them with `${...}` +placeholders): + +- `${OPERATOR_DIRECTIVES}` — the bulleted list (may be empty) +- Step-specific inputs (paths to upstream artifacts, output path, + etc.) per that step's SKILL.md + +## Re-run authorization + +**A chain skill never invokes another chain skill on its own.** +When a step finds that its upstream input is unsatisfactory — a +thin functionality report, a too-easy bench, a missing-but-needed +probe, a wrong-target PR-conventions report — it surfaces the +finding to the user and stops. The user (or the auto-pilot driver +applying its default policies) decides whether to re-run an +upstream step, accept the situation, or abandon the run. + +This applies to backward triggers specifically. Forward handoffs +(step N hands off to step N+1) are part of the chain's normal flow +and the handing-off skill emits the handoff message; the user or +auto-pilot acts on it. + +Why this rule exists: chain skills running on user request must +give the user control over what they're paying for. Re-running an +upstream step takes time and tokens; the user authorizes that +explicitly, not the agent. diff --git a/skills/skill-optimizer-validate-improvement/SKILL.md b/skills/skill-optimizer-validate-improvement/SKILL.md index ab3fb2f..db22588 100644 --- a/skills/skill-optimizer-validate-improvement/SKILL.md +++ b/skills/skill-optimizer-validate-improvement/SKILL.md @@ -7,50 +7,27 @@ description: Use when the user wants to validate an improvement proposal from `s Step 8 of the skill-optimizer chain. Takes the proposal from step 7, dispatches a validator subagent to independently check whether -the proposed change is sound (internal consistency against the -skill's stated responsibilities) and conformant (external PR -conventions if PR-bound), then — on `verdict: approve` — -materializes the improved skill at `improved-skill/`. Writes -`08-validator-verdict.md` regardless of the verdict; materialization -happens only on approve. - -The validator is dispatched fresh per invocation (no in-step -revision loop). If the verdict is `needs-revision`, the user (or -auto-pilot at step 9) re-invokes step 7 with the validator's -rationale distilled into a directive, then re-invokes this step. -This honors the chain's "chain skills don't auto-invoke other -chain skills" rule and keeps both steps single-shot. - -## Before you start - -Two load-bearing pieces of context to load NOW, before the -workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation. Step 8 is a - **fresh-derivation** step: each invocation derives a fresh - verdict from the current proposal + current skill state + - directives, with no continuity from prior verdicts. The - independence-from-the-optimizer constraint is load-bearing - here. - -2. **You will dispatch a subagent for the actual validation; you - do NOT judge the proposal yourself in this session.** Workflow - step (c) is the dispatch. The rationale is in "Why - limited-context dispatch matters" below. +the proposed change is sound (internal consistency) and conformant +(external PR conventions if PR-bound), then — on `verdict: approve` +— materializes the improved skill at `improved-skill/`. Writes +`08-validator-verdict.md` regardless of verdict. + +**Single-shot per invocation.** No in-step revision loop. If the +verdict is `needs-revision`, the operator (or auto-pilot at step 9) +re-invokes step 7 with the validator's rationale distilled into a +directive, then re-invokes this step. Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain (this skill is step 8; `improve-skill` -is step 7; `autopilot` is step 9). Internal workflow steps within -THIS skill are labelled "(a)" through "(e)". +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(e)". ## What you produce One or two artifacts at `docs/skill-optimizer//`: 1. **`08-validator-verdict.md`** — the validator's independent - judgment. Frontmatter (runtime-relevant facts only, per the - iteration protocol's frontmatter discipline): + judgment. Frontmatter (runtime-relevant facts only, per + [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- @@ -61,29 +38,21 @@ One or two artifacts at `docs/skill-optimizer//`: --- ``` - The body covers the internal consistency check (does the - change make sense for the named weakness? additive vs. - destructive? general vs. ducktape?) and — if - `03-submissions.md` exists — the external consistency check - (does the change conform to upstream PR rules: frontmatter, - file location, prefix taxonomy, additive-only, etc.). The - external check is forward-looking — it verifies the improved - skill COULD be turned into a valid PR, even though step 7 - itself does not produce one. - -2. **`improved-skill/`** — the improved skill content, only - materialized when `verdict: approve`. Mirrors the source - skill's directory structure with the optimizer's diff applied. - **The original is never modified**: for upstream skills, - `vendored-skill/` stays frozen; for local skills, the user's - source file stays untouched. Git tracks `improved-skill/` - history across iterations. - -The validator's body sections + reasoning protocol are specified -in -[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). -The subagent writes the verdict body itself; the operator session -materializes `improved-skill/` after approve. + Body covers the internal consistency check (does the change + make sense for the named weakness? additive vs. destructive? + general vs. ducktape?) and — if `03-submissions.md` exists — + the external consistency check (conformant to upstream PR + rules?). External check is forward-looking — verifies the + improved skill COULD be turned into a valid PR, even though + step 8 doesn't produce one. Body template at + [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). + +2. **`improved-skill/`** — improved skill content, only + materialized when `verdict: approve`. Mirrors the source's + directory structure with the optimizer's diff applied. **The + original is never modified**: `vendored-skill/` stays frozen, + local source files stay untouched. Git tracks + `improved-skill/` history across iterations. ## Workflow @@ -91,247 +60,130 @@ materializes `improved-skill/` after approve. Three checks: -1. `docs/skill-optimizer//07-improvement-proposal.md` must - exist with valid frontmatter. If not, tell the user to run - `skill-optimizer-improve-skill` first and stop here. - -2. The current skill state must be readable: `improved-skill/` if - it exists (the prior accumulated state), else the original - source (`vendored-skill/` for upstream skills, the local file - path recorded in `01-functionality.md`'s `skill_source` for - local skills). The validator needs this as "skill BEFORE". - -3. `01-functionality.md` must exist (the validator reads it as - input for the internal consistency check). +1. `07-improvement-proposal.md` must exist with valid frontmatter. + If not, tell the user to run `skill-optimizer-improve-skill` + first. +2. Current skill state must be readable: `improved-skill/` if it + exists (prior accumulated state), else the original source. + The validator needs this as "skill BEFORE". +3. `01-functionality.md` must exist. -If `pr_submission_intent: true` from `01-functionality.md`, -`03-submissions.md` should also exist; if it's missing, the -external consistency check can't run. Surface this and ask the -user whether to run step 3 first OR proceed with internal -consistency only (the validator's output will note the omission). +If `pr_submission_intent: true`, `03-submissions.md` should +exist; if missing, the external check can't run. Ask whether to +run step 3 first or proceed with internal consistency only. ### (b) Handle iteration -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** -and apply it. Step 8 is a **fresh-derivation** step: - -- Each invocation derives a fresh verdict from the current - proposal + skill state + directives. The validator subagent - does NOT read its own prior `08-validator-verdict.md` or its - git history. Independence-from-self is load-bearing here: if - the validator saw its prior verdict, it would gravitate toward - consistency-with-itself across runs, defeating the purpose of - re-validation after the operator/auto-pilot revises and - re-invokes step 7. -- Collect `${OPERATOR_DIRECTIVES}` per the protocol. The operator - CAN read the prior verdict and distill lessons into atomic new - requirements for the validator. Examples: "be stricter on - additive-vs-destructive — the prior verdict approved a change - that I think was actually destructive", "the upstream PR - conventions check missed the frontmatter `version:` field — - look for it specifically". The subagent works from the - distilled directives, never from the raw prior verdict. +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Step 8 is **fresh-derivation**: each invocation +overwrites the canonical verdict. Collect `${OPERATOR_DIRECTIVES}` +— examples: "be stricter on additive-vs-destructive", "the +external check missed the frontmatter `version:` field — look for +it specifically". ### (c) Dispatch the validator subagent -**Do NOT judge the proposal yourself in this session.** Dispatch -the validator subagent via the `Agent` tool (with worktree -isolation if your environment supports it). Load the prompt -template at +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT judge the proposal yourself in this +session** — dispatch the subagent via the `Agent` tool. Load the +prompt template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) -and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (the -current state: `improved-skill/` if it exists, else the original -source), `${SKILL_AFTER_PATH}` (the proposed result — materialize -by applying the optimizer's diff to a temporary copy of BEFORE; -do NOT overwrite or rename anything canonical at this stage), +and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (current +state — `improved-skill/` if it exists, else source), +`${SKILL_AFTER_PATH}` (proposed result — apply the diff to a temp +copy; do NOT touch the canonical `improved-skill/` yet), `${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, `${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. -The validator sees: - -- The skill BEFORE the change (the current state the optimizer - read in step 7) -- The skill AFTER the change (a temporary materialization; the - canonical `improved-skill/` is not updated yet — that happens - at (d) on approve) -- `${PROPOSAL_PATH}` — `07-improvement-proposal.md` (so it - knows what the optimizer claims to have done; the validator - checks the artifact, not the optimizer's reasoning trace) -- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the - internal consistency check) -- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for - the external consistency check) -- `${OPERATOR_DIRECTIVES}` — atomic new requirements (may be - empty) - -The validator does NOT see: - -- Raw failed trials, `findings.txt`, `trace.jsonl` -- Test inputs (`tests//workspace/`) -- The optimizer's internal reasoning trace (only the proposal - artifact, not how the optimizer arrived at it) -- Its own prior `08-validator-verdict.md` or git history of it - (independence-from-self constraint) -- `06-analysis.md` directly (the analyzer's reasoning is - mediated through the optimizer's proposal — the validator's - job is to check the proposal as an independent observer, not - to second-guess the analysis) - -The validator writes `08-validator-verdict.md` with a verdict of -`approve`, `needs-revision`, or `reject`, plus rationale. +The validator sees: skill BEFORE; skill AFTER (temporary +materialization); `07-improvement-proposal.md`; +`01-functionality.md`; `03-submissions.md` if PR-bound; +`${OPERATOR_DIRECTIVES}`. + +The validator does NOT see: raw failed trials, `findings.txt`, +`trace.jsonl`; test inputs (`tests//workspace/`); the +optimizer's reasoning trace (only the proposal artifact); its own +prior `08-validator-verdict.md` or git history; `06-analysis.md` +directly. + +**Why this matters:** independence is the validator's load-bearing +property. **Bias from optimizer's reasoning** — seeing the +optimizer's reasoning trace would lead the validator to accept +the optimizer's arguments rather than judging the artifact fresh. +**Consistency-with-self across runs** — seeing prior verdicts +would lead the validator to repeat its prior judgment instead of +judging the new proposal on its own merits. ### (d) Handle the verdict + materialize on approve -Read the verdict from `08-validator-verdict.md`. +**`verdict: approve`:** materialize `improved-skill/` by copying +the source's directory structure and applying the optimizer's +diff. **Do NOT modify the source.** Commit: -**If `verdict: approve`:** +```bash +git add docs/skill-optimizer// +git commit -m "step 8: validate + apply improvement for " +``` -1. Materialize `improved-skill/` by copying the source skill's - directory structure and applying the optimizer's diff. **Do - NOT modify the source.** For upstream skills, leave - `vendored-skill/` frozen and write to a sibling - `improved-skill/`; for local skills, copy the user's source - file(s) to `improved-skill/` and apply the diff there. - -2. Commit the new state: - - ```bash - git add docs/skill-optimizer// - git commit -m "step 8: validate + apply improvement for — " - ``` +The source skill file is NOT in the commit — git tracks +`improved-skill/` alongside the verdict report, but never touches +the user's source or the vendored reference. - The source skill file is NOT in the commit — git tracks - `improved-skill/` alongside the verdict report, but never - touches the user's source or the vendored reference. +**`verdict: needs-revision`** or **`reject`:** do NOT materialize. -3. Proceed to (e). - -**If `verdict: needs-revision`:** do NOT materialize -`improved-skill/`. The proposal needs to be revised before the -chain advances. Proceed to (e) with the appropriate handoff. +### (e) Hand off -**If `verdict: reject`:** do NOT materialize. The validator -judges the proposal as fundamentally wrong (not just in need of -revision). Proceed to (e) with the appropriate handoff. +Three messages by verdict: -### (e) Hand off +- **approve:** "Validation complete; improvement applied at + `improved-skill/`. The chain has reached its natural endpoint + for this iteration. Three realistic next steps: review and copy + locally; hand to a PR composer (auto-pilot can do this + end-to-end if running step 9); or re-bench against a workbench + that points at `improved-skill/`." +- **needs-revision:** "Validator says needs-revision. Distill the + rationale into a directive and re-invoke step 7, then re-invoke + this step. Operator owns distillation; chain skills don't + auto-invoke each other." +- **reject:** "Validator rejects outright. Two paths: re-invoke + step 6 with a reframed weakness, or accept that this weakness + isn't addressable and exit honestly." -Three handoff messages depending on the verdict: - -**If `verdict: approve`:** - -> Validation complete; improvement applied. The improved skill is -> at `docs/skill-optimizer//improved-skill/`; the original -> source (vendored or local) is untouched. The chain has reached -> its natural endpoint for this iteration. Three realistic next -> steps: (1) review `improved-skill/` and copy it over your -> local source if you're satisfied; (2) if PR-bound, hand the -> improved skill + the proposal + `03-submissions.md` to a PR -> composer (auto-pilot can do this end-to-end if you're running -> step 9); (3) re-bench against the improved skill by re-invoking -> `skill-optimizer-run-bench` against a workbench that points at -> `improved-skill/`. - -**If `verdict: needs-revision`:** - -> Validation rejected as needs-revision. The validator's -> rationale is in `08-validator-verdict.md`. To address it: -> re-invoke `skill-optimizer-improve-skill` with a directive -> distilling the validator's concern (e.g., `validator said: -> '' — address this and re-propose`), then re-invoke -> `skill-optimizer-validate-improvement`. The operator owns -> distillation; chain skills don't auto-invoke each other. - -**If `verdict: reject`:** - -> Validation rejected the proposal outright. The validator's -> rationale is in `08-validator-verdict.md`. Two realistic paths: -> (1) re-invoke `skill-optimizer-analyze-result` with a directive -> reframing the weakness — the analyzer's framing may have been -> misleading; (2) accept that this weakness isn't addressable -> without changes the validator is unwilling to approve, and -> exit honestly. - -Don't auto-invoke step 7 or step 6 — surface the choice and let -the user (or auto-pilot at step 9) act. - -## Why limited-context dispatch matters - -Independence from the optimizer's reasoning is the validator's -load-bearing property. The validator must judge the proposal as -if seeing it for the first time, against the skill's stated -responsibilities (internal) and the upstream's PR conventions -(external). Two specific risks the limited context addresses: - -First, **biased acceptance of the optimizer's reasoning**. If the -validator saw the optimizer's reasoning trace, it would tend to -accept arguments the optimizer made about why the change is sound -— defeating the purpose of independent checking. The validator -sees the proposal artifact (what was changed and the optimizer's -brief rationale linking to the named weakness) but NOT the -optimizer's internal reasoning trace. - -Second, **consistency-with-self across runs**. If the validator -saw its prior verdict, it would gravitate toward -consistency-with-itself when re-validating after a revision — -either re-approving what it approved before or re-rejecting what -it rejected before, instead of judging the new proposal on its -own merits. Each invocation is genuinely fresh; the operator's -distilled directives carry forward what the prior verdict taught. - -If you find yourself thinking "I'll just judge the proposal -myself, I can see whether it addresses the weakness" — that's the -failure mode this chain is built to prevent. Dispatch. +Don't auto-invoke step 6 or 7. ## Edge cases -- **`07-improvement-proposal.md` missing** — caught at (a). Tell - user to run step 7 first. +- **`07-improvement-proposal.md` missing** — caught at (a). - **Current skill state ambiguous** (both `improved-skill/` and - source modified externally) — surface to the user; the BEFORE - the validator reads must match what the optimizer read in step - 7. If the user edited the source between step 7 and step 8, - re-invoke step 7 to derive a fresh proposal against the new - state. + source modified externally between step 7 and step 8) — surface + to the user; the BEFORE must match what the optimizer read. If + the user edited the source mid-way, re-invoke step 7 to derive + a fresh proposal against the new state. - **PR-bound but `03-submissions.md` missing** — caught at (a). - Ask the user whether to run step 3 first or proceed with - internal consistency only (the validator's verdict will note - the omission). + Proceeding loses the external check; the verdict body notes the + omission. - **Validator rejects with reasoning that contradicts the - analyzer's framing** (e.g., validator says "this change - doesn't address the real problem"; the analyzer thought it - did) — that's a signal the analyzer/validator are misaligned - on what the weakness actually is. Surface to the user with the - two paths from the `reject` handoff: re-invoke step 6 to - reformulate the weakness, or accept. + analyzer's framing** — surface with the two paths from the + `reject` handoff. - **Operator directives reference a specific section of the prior - verdict** (e.g., "the prior verdict's external check section - missed X") — that's a context-dump masquerading as a directive. - Translate to an atomic new requirement ("look for X in the - external check") before passing to the subagent. + verdict** — context-dump masquerading as a directive. Translate + to an atomic new requirement before passing. ## Iteration behavior -Step 8 is re-runnable and is a fresh-derivation step. Re-run -triggers specific to this step: +Fresh-derivation step. Re-run triggers: -- Step 7 (`improve-skill`) produced a new proposal — - `07-improvement-proposal.md` changed via git mtime +- Step 7 produced a new proposal — + `07-improvement-proposal.md` changed - User disagrees with the prior verdict and wants a fresh - derivation with directives reflecting the disagreement (e.g., - "be stricter on additive-vs-destructive") -- `03-submissions.md` was updated (upstream PR conventions - changed) and the prior verdict's external check is now stale - -General iteration mechanics — staleness detection (git-native), -operator directives (the channel for prior-derivation lessons), -cascading staleness — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. - -There is no in-step bounded revision loop. Each invocation of -step 8 is single-shot; if `verdict: needs-revision` or -`verdict: reject`, the operator (or auto-pilot at step 9) -re-invokes step 7 then step 8 manually. This honors the chain -skills' "no auto-invocation of other chain skills" rule and -keeps both steps semantically clean. + derivation with directives +- `03-submissions.md` was updated and the prior external check is + stale + +No in-step revision loop; each invocation is single-shot. The +7→8→7 cycle on `needs-revision`/`reject` is operator-driven, or +auto-pilot-driven at step 9. + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 6f73162..9333be4 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -6,56 +6,26 @@ description: Use when the user wants to build or implement the eval probes for a # skill-optimizer-write-tests Step 4 of the skill-optimizer chain. Reads the picked functionalities -from step 2 (`tests//spec.yaml` files where `picked: -true`), decides the probe set per functionality (informed by each -spec's `suggested_probes` and any operator directives), dispatches +from step 2 (`tests//spec.yaml` files where +`picked: true`), decides the probe set per functionality, dispatches one test-writer subagent per probe (in parallel) to build the concrete workspace files + grader scripts + smoke fixtures, then generates `tests/suite.yml` for the run-bench step. -## Before you start - -Three load-bearing pieces of context to load NOW, before the -workflow: - -1. **Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md).** - Every chain skill applies it on every invocation. Step 4 is a - **maintenance** step (see the protocol's "Step kinds" section) - — the `tests/` tree accumulates probe folders across re-runs - rather than being rebuilt from scratch. Workflow step (b) - requires it. - -2. **You will dispatch test-writer subagents (in parallel, one per - probe being built or rebuilt); you do NOT write workspace files - or grader scripts yourself in this session.** Workflow step (d) - is the dispatch. The rationale is in "Why limited-context - dispatch matters" below. - -3. **Each test-writer subagent sees the skill's source content.** - This is intentional and is the one step in the chain where the - source is in-scope for a generative subagent. The constraint - comes from the probe spec (which fixes what the fixture should - test); source access is needed for concrete violation patterns - and realistic fixture content. - Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain (this skill is step 4; step 2 is -`investigate-test-case`; etc.). Internal workflow steps within -THIS skill are labelled "(a)" through "(f)". +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(f)". ## What you produce Two kinds of output under `docs/skill-optimizer//`: -1. **Probe folders** at - `tests///`. For each picked - functionality, one or more probe folders. Each probe folder - contains: +1. **Probe folders** at `tests///`. For + each picked functionality, one or more probe folders containing: ```text tests/// - spec.yaml # probe-level intent (workspace overview, expected - # agent behavior, grader logic) + spec.yaml # probe-level intent workspace/ # the files the agent sees in /work grader.mjs # the grading script (or .py) smoke/ @@ -65,26 +35,20 @@ Two kinds of output under `docs/skill-optimizer//`: checks/smoke.mjs # runs the grader against the three smoke fixtures ``` - The filesystem IS the state — a probe exists if its folder - exists with these contents; there is no `built_cases: []` array - anywhere. The presence of `grader.mjs` and the absence of any - smoke-check failure marker indicates the probe is built and - verified. + Filesystem IS the state — a probe exists if its folder exists + with these contents; `grader.mjs` present + smoke-check passed + = built and verified. 2. **`tests/suite.yml`** — generated from the current tree. Lists - every probe under every `picked: true` functionality. Step 5 - (run-bench) reads this. Regenerated by step (f) of this skill on - every run. - -The probe-level `spec.yaml` format (workspace overview, expected -agent behavior, grader logic, smoke-check fixtures) is defined in -[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md). + every probe under every `picked: true` functionality. + Regenerated by step (f) on every run. -For the workbench schema itself (`suite.yml` structure, grader -contract, smoke-check format), see +Probe-level `spec.yaml` format and the workbench schema for +`suite.yml` are defined in +[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) +and [`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md) -— that's the canonical workbench reference used by the run-bench -step. +respectively. ## Workflow @@ -92,277 +56,145 @@ step. Three prerequisites: -1. `docs/skill-optimizer//01-functionality.md` exists. -2. `docs/skill-optimizer//tests/` exists with at least one - `tests//spec.yaml` file having `picked: true`. -3. Every picked functionality's spec.yaml parses with the required - fields (`name`, `description`, `picked`, `importance`, - `suggested_probes`, `why_test`). +1. `01-functionality.md` exists. +2. `tests/` exists with at least one + `tests//spec.yaml` having `picked: true`. +3. Every picked functionality's spec.yaml parses with required + fields. -If any prerequisite fails, surface the issue to the user — don't -attempt to derive missing fields or guess. List the picked -functionalities you found; the user should confirm before you -proceed. +If any fails, surface to the user — don't derive missing fields or +guess. ### (b) Handle iteration (maintenance step) -**Read [`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) now** +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply the maintenance-step flow: -- Walk the current `tests/` tree. For each picked functionality, - determine which probes already have built folders (i.e., - `tests///grader.mjs` exists). Existing built - probes are preserved by default; the operator session does NOT - re-dispatch test-writers for them unless `${OPERATOR_DIRECTIVES}` - explicitly names them for revision. -- Collect `${OPERATOR_DIRECTIVES}` per the protocol's section. - Atomic new requirements only — never a context dump of existing - spec.yaml contents. +- Walk `tests/`. For each picked functionality, determine which + probes already have built folders (`grader.mjs` exists). + Existing built probes are preserved by default; don't + re-dispatch unless `${OPERATOR_DIRECTIVES}` explicitly names them + for revision. +- Collect `${OPERATOR_DIRECTIVES}` per the protocol. - If a directive will be destructive (rebuilding an existing - probe, removing a probe that's no longer wanted), commit a - checkpoint of the current `tests/` tree BEFORE dispatching the - rebuild, per the protocol's "Safe destructive edits" section. - Example commit message: `checkpoint: before rebuild of - tests/refuses-malformed-input/null/`. + probe), commit a checkpoint BEFORE dispatching per the protocol's + safe-destructive-edits section. ### (c) Plan probe set + user gate -For each picked functionality: - -1. Read its `spec.yaml`'s `suggested_probes` field. -2. Decide the probe set: - - Start with `suggested_probes` as the default. - - If `${OPERATOR_DIRECTIVES}` adds probes for this functionality - (e.g., "add an empty-array probe to refuses-malformed-input"), - include them. - - If `${OPERATOR_DIRECTIVES}` removes or revises probes, - reflect that. - - Existing built probe folders in `tests///` - stay in the set unless directives target them for revision or - removal. - -Show the user the planned probe set per functionality: - -> For each picked functionality, here are the probes I'll build: -> -> - `refuses-malformed-input/`: malformed-json (existing), -> missing-field (existing), empty-array (new) -> - `uses-right-tool/`: tool-a-scenario (existing), -> tool-b-scenario (new) -> -> New probes to dispatch test-writers for: [list]. Existing probes -> preserved: [list]. Probes to rebuild (will commit checkpoint -> first): [list]. Anything to change before I dispatch? +For each picked functionality, decide the probe set: start with +`suggested_probes`, apply directives (adds/removes/revisions), +preserve existing built probes unless directives target them. + +Show the user the planned probe set per functionality, marking +each probe as existing / new / rebuild. Ask whether to change +anything before dispatch. Three realistic responses: -1. **User approves.** Proceed to (d) with the dispatch set as - planned. -2. **User adds revisions** (e.g., "the existing malformed-json - probe's grader is too strict — rebuild it with a looser match"). - Treat as new operator directives, mark the named probes for - rebuild, and proceed to (d) with the updated dispatch set. - Apply the destructive-edit checkpoint from (b). -3. **User rejects the plan structure.** Likely they want a - different probe set overall — surface this and ask whether to - abandon (loop back to step 2 to revise the functionality specs) - or to retry the plan with their feedback as directives. +1. **User approves.** Proceed to (d). +2. **User adds revisions.** Treat as directives, mark named probes + for rebuild (apply destructive-edit checkpoint), proceed to + (d). +3. **User rejects the plan structure.** Surface and ask whether to + abandon (loop back to step 2 to revise functionality specs) or + retry with their feedback as directives. ### (d) Dispatch test-writer subagents (parallel, one per probe) -**Do NOT write workspace files or graders yourself in this -session.** For each probe in the dispatch set (new probes + probes -flagged for rebuild in (c)), dispatch a test-writer subagent via -the `Agent` tool, in parallel — emit all dispatches in a single -message so they run concurrently. - -Load the prompt template at +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT write workspace files or graders +yourself in this session** — dispatch test-writer subagents via +the `Agent` tool, in parallel (emit all dispatches in a single +message). Load the prompt template at [`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) -and substitute per-probe inputs: - -- `${PROBE_NAME}` — this probe's slug (the folder name under - `tests//`) -- `${FUNCTIONALITY_SPEC_PATH}` — path to the parent - functionality's `spec.yaml` (gives the subagent context on what - responsibility this probe is probing) -- `${PROBE_SPEC_PATH}` — path where the probe's `spec.yaml` will - be written (the test-writer also writes this file describing - what the probe sets up + expects) -- `${FUNCTIONALITY_PATH}` — the current `01-functionality.md` -- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local - skill path), so the subagent can ground the fixture in concrete - violation patterns -- `${OUTPUT_PROBE_DIR}` — `tests///` -- `${OPERATOR_DIRECTIVES}` — case-level revision hints if this - probe is being rebuilt; empty otherwise - -Each test-writer subagent sees: - -- Its single probe's intent (the slug + the parent functionality - spec) -- `01-functionality.md` -- The skill source content (vendored or local) -- `${OPERATOR_DIRECTIVES}` for its probe - -The test-writer subagent does NOT see: - -- Other probes' specs, graders, or workspaces (prevents copying - across probes and grader-leak hacking) -- The eval grader's matching internals (same reason) -- `06-analysis.md`, `07-improvement-proposal.md`, raw failure - data, prior optimizer attempts -- Git history of its own probe folder (anti-rationalization) - -The "no other probes" constraint is load-bearing: each probe is -built in isolation so cross-probe patterns don't subtly homogenize -the fixtures. The "no grader internals" constraint prevents -fixtures from being gerrymandered to the grader's specific matching -rules. - -Each test-writer subagent writes: - -- `tests///spec.yaml` (the probe's - intent + workspace overview + expected behavior + grader logic) -- `tests///workspace/` (the - fixture) -- `tests///grader.mjs` (the grader) -- `tests///smoke/{good,bad,empty}/findings.txt` - (smoke fixtures the grader should classify correctly) -- `tests///checks/smoke.mjs` (a runner - that exercises the grader against the three smoke fixtures) - -The subagent returns a brief summary: probe name, what the fixture -tests, grader logic in one line, smoke-check result for its own -probe. +and substitute per-probe inputs: `${PROBE_NAME}`, +`${FUNCTIONALITY_SPEC_PATH}`, `${PROBE_SPEC_PATH}`, +`${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, +`${OUTPUT_PROBE_DIR}`, `${OPERATOR_DIRECTIVES}`. + +Each test-writer sees: its single probe's intent (slug + parent +functionality spec); `01-functionality.md`; the skill source +content; `${OPERATOR_DIRECTIVES}` for its probe. + +Each test-writer does NOT see: other probes' specs, graders, or +workspaces; the eval grader's matching internals; `06-analysis.md`, +`07-improvement-proposal.md`, raw failure data; git history of +its own probe folder. + +**Why this matters:** two concerns. First, **fixture isolation** — +a subagent seeing all probes' fixtures would notice cross-probe +patterns and inadvertently homogenize them; per-probe isolation +keeps each fixture a representative instance. Second, **no +grader-leak hacking** — a subagent seeing another grader's matching +logic could write fixtures that incidentally satisfy that grader +too, making eval results look correlated when they're not. + +Source-content access is the one exception across the chain: probe +spec from step 2 fixes WHAT to test; source provides the HOW +(concrete violation patterns). Without source, the subagent would +invent generic patterns that may not trigger the skill's rules. + +Each test-writer writes the probe folder contents (spec.yaml, +workspace files, grader, smoke fixtures, smoke runner) and +returns: probe name, what the fixture tests, grader logic in one +line, smoke-check result. ### (e) Run smoke check -The smoke check is per-probe (each probe has its own -`checks/smoke.mjs`). After all test-writer dispatches return, run -each new or rebuilt probe's smoke runner: +Each probe has its own `checks/smoke.mjs`. After test-writer +dispatches return, run each new or rebuilt probe's smoke runner: ```bash node tests///checks/smoke.mjs ``` -Expected for each probe: - -- The GOOD fixture passes the grader. -- The BAD fixture fails the grader. -- The EMPTY fixture fails the grader. +Expected per probe: GOOD passes, BAD fails, EMPTY fails. If any +fails, surface to the user. Two realistic responses: -If any grader fails the smoke check, surface the probe to the -user. Two realistic responses: - -1. **User asks to re-dispatch the test-writer for that probe** - (with the smoke-check failure as a directive). Treat as a - single-probe rebuild, return to (b) for the destructive-edit - checkpoint, then re-dispatch via (d). -2. **User decides to remove the probe** from this functionality's - set, or de-pick the parent functionality. The change is to - `tests//spec.yaml` (set `picked: false`) or to - delete the probe folder. That's a backward trigger — per the - iteration protocol's re-run authorization rule, surface the - option and let the user invoke step 2 explicitly if they want - to revise the functionality. +1. **Re-dispatch the test-writer** with the smoke failure as a + directive. Treat as a single-probe rebuild (back to (b) for + the checkpoint, then (d)). +2. **Remove the probe** from this functionality's set, or de-pick + the parent functionality (a backward trigger to step 2; per + the no-auto-invocation rule, surface the option and let the + user invoke step 2). ### (f) Generate suite.yml + commit + hand off -Walk the current `tests/` tree. For each `tests//` -where `spec.yaml` has `picked: true`, list every probe folder -(`//`) and emit a suite.yml entry per -probe per the workbench schema (see -[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md) -for the exact shape). - -Write the result to `tests/suite.yml`, overwriting any prior -version (git tracks the change). +Walk `tests/`. For each `tests//` where `spec.yaml` +has `picked: true`, list every probe folder and emit a suite.yml +entry per the workbench schema. Overwrite `tests/suite.yml`. -Commit the current state of `tests/`: - -```bash -git add docs/skill-optimizer//tests/ -git commit -m "step 4: build probes for " -``` - -Handoff message, verbatim: - -> Next, invoke `skill-optimizer-run-bench` to measure baseline -> performance against the probes. - -## Why limited-context dispatch matters - -The test-writer subagents are dispatched per-probe for two -reasons: - -First, **fixture isolation**. A single subagent writing all -fixtures would notice patterns across probes ("they all check for -absence violations — let me write a unified fixture") and -inadvertently homogenize them. Per-probe isolation forces each -fixture to be a representative instance of its responsibility, -designed without knowledge of how sibling probes are shaped. - -Second, **no grader-leak hacking**. If a subagent sees how another -grader matches (regex pattern, JSON path, etc.), it can write a -fixture that incidentally satisfies the OTHER grader too — making -the eval results look correlated when they're not. Per-probe -isolation prevents this. - -The source access concession (test-writer subagents DO see the -skill source) is bounded by the probe's spec: the spec fixes WHAT -the fixture should test, and the source provides the HOW (concrete -violation patterns). Without source, the subagent would invent -generic patterns that may not actually trigger the skill's rules. - -If you find yourself thinking "the probes are similar; I'll write -them all myself with one prompt and save dispatches" — that's the -homogenization failure mode this chain is built to prevent. -Dispatch in parallel. +Commit `tests/` and hand off to `skill-optimizer-run-bench`. ## Edge cases -- **A test-writer subagent reports BLOCKED** (e.g., the probe spec - is too abstract to derive a fixture from) — surface to the user - with the subagent's reasoning. The fix is typically a step 2 - re-run with a more specific functionality spec or - suggested_probes list; per the iteration protocol's re-run - authorization rule, the user invokes step 2 explicitly. -- **A grader fails the smoke check** — see step (e)'s handling. -- **Probe smoke checks pass individually but `tests/suite.yml` is - malformed** — surface as a workbench-schema error; fix the - suite.yml directly (operator session task; not a test-writer - concern). -- **A picked functionality is genuinely unimplementable in a - static workbench** (e.g., requires real-time API access the - workbench can't provide) — the test-writer should return BLOCKED - with this reasoning. Surface to the user; the functionality may - need to be de-picked at step 2. +- **A test-writer reports BLOCKED** (probe spec too abstract to + derive a fixture from) — surface to the user. Fix is typically a + step 2 re-run with a more specific functionality spec; per the + no-auto-invocation rule, user invokes step 2. +- **A grader fails smoke check** — see step (e)'s handling. +- **A picked functionality is unimplementable in a static + workbench** (e.g., needs real-time API access) — test-writer + returns BLOCKED. Functionality may need to be de-picked at + step 2. - **User de-picks a functionality between step 4 runs** — its - probe folders stay on disk (filesystem-as-state principle), but - step (f) excludes them from the regenerated `tests/suite.yml`. - If the user later re-picks the functionality, the probes are + probe folders stay on disk; step (f) excludes them from the + regenerated `tests/suite.yml`. If re-picked later, probes are already there. ## Iteration behavior -Step 4 is a maintenance step — the `tests/` tree accumulates -probes as new functionalities get picked, and re-runs preserve -existing built probes unless directives target them. Re-run -triggers specific to this step: +Maintenance step. Re-run triggers: - A new functionality at step 2 got `picked: true` -- A specific probe's smoke check failed and needs the test-writer - to revise (single-probe rebuild via directive) -- Step 5 (`run-bench`) showed all probes pass on baseline (too - easy) or all fail (too hard) — directives like "make probe X - harder" or "loosen probe Y's grader" trigger probe-level - rebuilds +- A specific probe's smoke check failed and needs revision +- Step 5 showed probes systematically too easy / too hard — + directives trigger probe-level rebuilds - Step 2 substantively revised a picked functionality's spec.yaml - (the suggested_probes list changed in a way that warrants - rebuilding) — operator surfaces and the user decides whether to - rebuild the existing probes or leave them - -General iteration mechanics — staleness detection, operator -directives, cascading staleness, safe destructive edits — live in -[`skills/skill-optimizer-shared/iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md). -Workflow step (b) above already requires reading that file. + in a way that warrants rebuilding existing probes + +Mechanics live in +[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); +workflow step (b) loads it. From cfc79f26f784590e6a07da0592ac07b43d381a81 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 08:37:19 -0500 Subject: [PATCH 035/121] =?UTF-8?q?refactor(v1.4-chain):=20closeness=20swe?= =?UTF-8?q?ep=20=E2=80=94=20inline=20edges,=20drop=20iteration=20footer,?= =?UTF-8?q?=20add=20workflow=20doc?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two coupled cleanups per the philosophy doc's closeness principle: 1. Inlined step-specific edge cases into their workflow steps. Most edge cases were just elaborations of "Confirm prerequisites" or "Confirm subagent output" — they belong INSIDE those steps, not in a trailing section. The small remainder (cross-cutting concerns like filesystem-as-state observations) stays in a tiny "Edge cases" footer where it earns its space. 2. Dropped "Iteration behavior" sections from each chain skill. The re-run triggers were skill-scheduling info (when to invoke this skill), not in-skill workflow content. Moved them into a new operator-facing docs/skill-optimizer-workflow.md as a cross-skill matrix — closer to where an operator deciding what to run next would look. Added a one-line "**Fresh-derivation step.**" or "**Maintenance step.**" classification near the top of each skill (since this affects subagent behavior and is short). The new docs/skill-optimizer-workflow.md (Phase D, ~110 lines) covers cross-skill concerns: the 9-step chain table, per-step re-run triggers, backward triggers (when a step surfaces a problem with an earlier step), state layout, pointers to the shared docs. Final line counts: B1 investigate-functionality 148 -> 131 (-17) B2 investigate-test-case 191 -> 174 (-17) B3 investigate-submissions 156 -> 142 (-14) B4 write-tests 200 -> 182 (-18) B5 run-bench 154 -> 131 (-23) B6 analyze-result 165 -> 148 (-17) B7 improve-skill 181 -> 155 (-26) B8 validate-improvement 189 -> 165 (-24) workflow.md (new) 0 -> 110 Chain net: -156 lines from chain skills, +110 lines for the workflow doc that captures the cross-skill content previously duplicated across each skill's "Iteration behavior". Net per-file size is smaller, and the per-skill files now follow the closeness principle: edge cases live next to the step they refine; re-run triggers live in the operator-facing reference, not in the per-skill workflow. Combined with the prior verbosity sweep: chain SKILL.md files have gone from 2320 to 1228 lines (-1092, -47%) since this morning. --- docs/skill-optimizer-workflow.md | 110 ++++++++++++++++++ .../skill-optimizer-analyze-result/SKILL.md | 85 ++++++-------- skills/skill-optimizer-improve-skill/SKILL.md | 70 ++++------- .../SKILL.md | 83 ++++++------- .../SKILL.md | 74 +++++------- .../SKILL.md | 71 +++++------ skills/skill-optimizer-run-bench/SKILL.md | 95 ++++++--------- .../SKILL.md | 78 +++++-------- skills/skill-optimizer-write-tests/SKILL.md | 50 +++----- 9 files changed, 335 insertions(+), 381 deletions(-) create mode 100644 docs/skill-optimizer-workflow.md diff --git a/docs/skill-optimizer-workflow.md b/docs/skill-optimizer-workflow.md new file mode 100644 index 0000000..e0b888c --- /dev/null +++ b/docs/skill-optimizer-workflow.md @@ -0,0 +1,110 @@ +# skill-optimizer workflow + +Operator-facing reference for the 8-step skill-optimizer chain. +Each step's SKILL.md is self-contained for its own work; this doc +covers cross-skill concerns: the high-level flow, what triggers +re-runs, and the relationship between steps. + +## The chain + +| # | Skill | Kind | Input | Output | +|---|---|---|---|---| +| 1 | `investigate-functionality` | fresh-derivation | source skill (URL or local) | `01-functionality.md` | +| 2 | `investigate-test-case` | maintenance | step 1 | `02-test-proposals.md`, `tests//spec.yaml` | +| 3 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `03-submissions.md` | +| 4 | `write-tests` | maintenance | steps 1+2 | `tests///`, `tests/suite.yml` | +| 5 | `run-bench` | fresh-derivation (summary) | step 4 + skill | `05-bench-results//`, `05-bench-summary.md` | +| 6 | `analyze-result` | fresh-derivation | step 5 + skill | `06-analysis.md` | +| 7 | `improve-skill` | fresh-derivation | step 6 + skill (+ step 3 if PR-bound) | `07-improvement-proposal.md` | +| 8 | `validate-improvement` | fresh-derivation | step 7 + skill (+ step 3 if PR-bound) | `08-validator-verdict.md`, `improved-skill/` (on approve) | +| 9 | `autopilot` | chain driver | same as step 1 + flags | `autopilot-summary-.md` | + +The PR-or-not decision is made ONCE at step 1; subsequent steps +know from `01-functionality.md`'s `pr_submission_intent` field +whether step 3 will run. No late prompts. + +## Re-run triggers (when each step should be re-invoked) + +A chain skill never auto-invokes another chain skill (per +[`subagent-dispatch.md`](../skills/skill-optimizer-shared/subagent-dispatch.md) +re-run authorization). The operator (or auto-pilot at step 9) +decides when to re-run each step. Common triggers: + +| Step | Re-run when... | +|---|---| +| 1 | source URL changed; PR-intent changed; user wants fresh research with new directives | +| 2 | user wants different coverage; step 4/5/6 surfaced a coverage gap; step 1 changed | +| 3 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed (source URL updated) | +| 4 | step 2's `tests/` tree changed; a probe's smoke check failed; step 5 showed probes systematically too easy/hard | +| 5 | step 4's `tests/` changed; user wants fresh trial data (flakiness, model list changed); step 6 wants more trials | +| 6 | step 5 produced new results; user disagrees with prior verdict; step 7 was unable to address a named weakness | +| 7 | step 6 produced a new analysis; step 8 returned `needs-revision`/`reject` with a distillable rationale; user wants a different approach | +| 8 | step 7 produced a new proposal; user disagrees with prior verdict; step 3 was updated and prior external check is stale | +| 9 | self-iterable; picks up at whatever step is stale per git-mtime | + +Each skill checks **direct upstream only** for staleness — if step +1 went stale but step 2 wasn't re-run (user judged it still +valid), step 3 sees step 2 as current. Skipping a step's re-run +is an explicit operator judgment. + +## Backward triggers + +When a step surfaces a problem with an earlier step. These do NOT +auto-fire; they're surfaced to the user. + +| At step | Surfaced finding | Resolution path | +|---|---|---| +| 2 | proposal is thin (subagent found few responsibilities) | re-run step 1 with directives, OR accept | +| 4 | test-writer reports BLOCKED on a probe | re-run step 2 to refine the functionality spec | +| 5 | bench all-pass (probes too easy) | re-run step 2 with "make probes harder" directive | +| 6 | analyzer's `has_structural_weakness: false`, user disagrees | re-run step 6 with directive pointing at missed cluster | +| 7 | optimizer reports BLOCKED (weakness too abstract) | re-run step 6 to reformulate weakness | +| 8 | `needs-revision` | distill validator's rationale, re-run step 7, re-run step 8 | +| 8 | `reject` | re-run step 6 with reframed weakness, OR accept that weakness isn't addressable | + +## State layout + +```text +docs/skill-optimizer// + 01-functionality.md + 02-test-proposals.md # B2's audit report + tests/ + / + spec.yaml # B2 writes + / # B4 writes one per probe + spec.yaml + workspace/ + grader.mjs + smoke/{good,bad,empty}/ + checks/smoke.mjs + suite.yml # B4 generates + 03-submissions.md # only if step 3 ran + 05-bench-results// # raw, timestamped per run + 05-bench-summary.md # single canonical + 06-analysis.md + 07-improvement-proposal.md + 08-validator-verdict.md + vendored-skill/ # frozen original (upstream only) + improved-skill/ # B8 materializes on approve + autopilot-summary-.md # B9 writes per run +``` + +`` is `--` for upstream skills, or +`` for local skills. Filesystem IS the state; +history is git (no `version:` fields or `archive/` directories +per +[`frontmatter-discipline.md`](../skills/skill-optimizer-shared/frontmatter-discipline.md)). + +## Shared docs + +The chain skills load these on-demand at the workflow steps that +need them: + +- [`iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md) + — iteration mechanics (staleness, step kinds, + destructive-edit checkpoints, cascading, bootstrapping) +- [`subagent-dispatch.md`](../skills/skill-optimizer-shared/subagent-dispatch.md) + — subagent constraints, operator directives, templated dispatch + inputs, no-auto-invocation rule +- [`frontmatter-discipline.md`](../skills/skill-optimizer-shared/frontmatter-discipline.md) + — runtime facts vs. history rule diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze-result/SKILL.md index 7f43fb6..977e17d 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze-result/SKILL.md @@ -5,14 +5,14 @@ description: Use when the user wants to diagnose why a bench run produced failur # skill-optimizer-analyze-result -Step 6 of the skill-optimizer chain. Takes the bench summary + raw -trial output from step 5, dispatches an analyzer subagent to -cluster failures into named **structural weaknesses** of the skill -(or explicitly say there are none), and writes -`docs/skill-optimizer//06-analysis.md`. This is the chain's -**anti-ducktape gate**: step 7 refuses to fire unless this report -names at least one structural weakness, with the general principle -that WOULD address it and the anti-patterns that would NOT. +Step 6 of the skill-optimizer chain. **Fresh-derivation step.** Takes +the bench summary + raw trial output from step 5, dispatches an +analyzer subagent to cluster failures into named **structural +weaknesses** of the skill (or explicitly say there are none), and +writes `docs/skill-optimizer//06-analysis.md`. This is the +chain's **anti-ducktape gate**: step 7 refuses to fire unless this +report names at least one structural weakness, with the general +principle that WOULD address it and the anti-patterns that would NOT. Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain. Internal workflow steps within THIS @@ -62,21 +62,28 @@ Full body template and reasoning protocol in `overall_pass_rate`. If not, tell the user to run `skill-optimizer-run-bench` first. -If `overall_pass_rate == 1.0`: there's nothing to analyze. -Surface honestly — either accept that probes don't expose a -weakness, or re-run step 2 with a "make probes harder" directive. -Don't run the analyzer; there are no failures to cluster. +If `overall_pass_rate == 1.0`: there's nothing to analyze. Surface +honestly — either accept that probes don't expose a weakness, or +re-run step 2 with a "make probes harder" directive. Don't run +the analyzer; there are no failures to cluster. + +If the bench results dir is missing or `suite-result.json` is +malformed, surface as a step-5 problem (incomplete or corrupted +bench run) and tell the user to re-run step 5. `tests/` and the source skill must also be available. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 6 is **fresh-derivation**: each invocation -overwrites the canonical from current bench data + directives. -Collect `${OPERATOR_DIRECTIVES}` per the protocol — examples: -"focus on the gpt-5 cluster", "the user thinks weakness X is -actually two separate issues". +and apply it. Each invocation overwrites the canonical from +current bench data plus directives. Collect +`${OPERATOR_DIRECTIVES}` per the protocol — examples: "focus on +the gpt-5 cluster", "the user thinks weakness X is actually two +separate issues". + +If operator directives contradict each other, surface the +contradiction before re-dispatching. ### (c) Dispatch the analyzer subagent @@ -118,6 +125,11 @@ required parts (especially "What WOULD NOT address this" — the anti-ducktape signal step 7 needs). If anything's inconsistent or missing, surface to the user; don't fill it in yourself. +If the anti-pattern lists are empty/vague (anti-ducktape gate +compromised), re-dispatch with a directive: "each weakness needs +a concrete anti-pattern list — what specific moves should the +optimizer avoid?" + ### (e) Hand off Two messages depending on `has_structural_weakness`: @@ -129,37 +141,8 @@ Two messages depending on `has_structural_weakness`: refuse to fire. Either accept the conclusion, or re-invoke step 6 with a directive ('the analyzer missed the X cluster')." -Don't auto-invoke step 7; per the no-auto-invocation rule in -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), -the user decides. - -## Edge cases - -- **All trials passed** — caught at (a). Don't run the analyzer. -- **Bench results dir missing or `suite-result.json` malformed** — - surface as a step-5 problem; tell the user to re-run step 5. -- **Subagent returns `has_structural_weakness: false` and user - disagrees** — surface the disagreement, but do NOT pressure the - subagent to manufacture a weakness. Honest path: re-invoke step - 6 with a directive pointing at what the user thinks was missed. -- **Subagent returns weaknesses but the anti-pattern lists are - empty/vague** — anti-ducktape gate is compromised. Re-dispatch - with a directive ("each weakness needs a concrete anti-pattern - list"). -- **Operator directives contradict** — surface before - re-dispatching. - -## Iteration behavior - -Fresh-derivation step. Re-run triggers: - -- Step 5 produced new results — `05-bench-summary.md` changed via - git mtime -- User disagrees with the analyzer's verdict and wants a fresh - derivation with directives -- Step 7 was unable to address one of the named weaknesses and the - user wants it reformulated - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. +If the subagent returned `false` but the user disagrees: surface +the disagreement, but do NOT pressure the subagent to manufacture +a weakness. Honest path is re-invoking step 6 with a directive +pointing at what the user thinks was missed. Don't auto-invoke +step 6 — per the no-auto-invocation rule, the user decides. diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve-skill/SKILL.md index 8ef1b8f..1579d6a 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve-skill/SKILL.md @@ -5,13 +5,13 @@ description: Use when the user wants to improve a skill based on identified stru # skill-optimizer-improve-skill -Step 7 of the skill-optimizer chain. Takes the named structural -weaknesses from step 6, dispatches an optimizer subagent to draft a -principled fix, and writes `07-improvement-proposal.md`. The -proposal is then validated independently by step 8, which -materializes the improved skill on approve. **This step does not -produce the improved skill itself** — it produces the proposal that -step 8 acts on. +Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes +the named structural weaknesses from step 6, dispatches an optimizer +subagent to draft a principled fix, and writes +`07-improvement-proposal.md`. The proposal is then validated +independently by step 8, which materializes the improved skill on +approve. **This step does not produce the improved skill itself** — +it produces the proposal that step 8 acts on. **Refuses to fire** if `06-analysis.md` has `has_structural_weakness: false` — there's nothing to optimize, and @@ -84,13 +84,16 @@ external check still runs if it appears later. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 7 is **fresh-derivation**: each invocation -overwrites the canonical from `06-analysis.md` + current skill + -directives. Collect `${OPERATOR_DIRECTIVES}` — examples: "prefer -additive changes", "don't touch the description field — validator -rejected that last round". The operator reads prior proposals/ -verdicts and distills; the subagent never sees the raw prior -content. +and apply it. Each invocation overwrites the canonical from +`06-analysis.md` plus current skill plus directives. Collect +`${OPERATOR_DIRECTIVES}` — examples: "prefer additive changes", +"don't touch the description field — validator rejected that last +round". The operator reads prior proposals/verdicts and distills; +the subagent never sees the raw prior content. + +If a directive references a specific section of a prior proposal +or verdict, that's a context-dump masquerading as a directive. +Translate to an atomic new requirement before passing. ### (c) Dispatch the optimizer subagent @@ -126,6 +129,11 @@ through `06-analysis.md` is what enforces principled improvement. The optimizer writes `07-improvement-proposal.md` and returns: weakness(es) addressed, principle applied, lines/sections modified. +If the optimizer reports BLOCKED (the analyzer's named weakness +is too abstract to derive a concrete diff from), surface to the +user. Fix is typically a step 6 re-run with a reframing directive; +per the no-auto-invocation rule, the user invokes step 6. + ### (d) Confirm subagent output Verify the report parses and the body has: a clear proposed @@ -145,37 +153,3 @@ requirement explicit — don't fill it in yourself. > validator's concerns. Don't auto-invoke step 8. - -## Edge cases - -- **`06-analysis.md` missing** — caught at (a). -- **`has_structural_weakness: false`** — caught at (a). Refuse; - don't override. -- **Optimizer reports BLOCKED** (analyzer's weakness too abstract - to derive a diff from) — surface. Fix is typically a step 6 - re-run; user invokes step 6. -- **Self-check against anti-patterns missing or vague** — caught - at (d). Re-dispatch with a directive. -- **PR-bound but `03-submissions.md` missing** — caught at (a). - Proceeding loses upstream shape hints from the start; step 8's - external check still runs if `03-submissions.md` appears later. -- **Operator directives reference a specific section of a prior - proposal or verdict** — that's a context-dump masquerading as a - directive. Translate to an atomic new requirement before passing. - -## Iteration behavior - -Fresh-derivation step. Re-run triggers: - -- Step 6 produced a new analysis — `06-analysis.md` changed -- Step 8 returned `needs-revision`/`reject` — operator distills - the rationale into a directive -- User wants a different approach (passes directives like "prefer - additive changes", "focus on weakness 2 only") - -The 7→8→7 cycle on `needs-revision` is operator-driven, or -auto-pilot-driven at step 9. No in-step loop. - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 6bd73ec..2dc0f74 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -5,14 +5,15 @@ description: Use when the user wants to understand what an existing agent skill # skill-optimizer-investigate-functionality -Step 1 of the skill-optimizer chain. Takes a skill (URL or local path), -runs a researcher subagent to figure out what it's supposed to do, and -writes `docs/skill-optimizer//01-functionality.md` — the briefing +Step 1 of the skill-optimizer chain. **Fresh-derivation step.** Takes +a skill (URL or local path), runs a researcher subagent to figure out +what it's supposed to do, and writes +`docs/skill-optimizer//01-functionality.md` — the briefing document every later step consumes. -Throughout this document, "step 1" through "step 9" (no parens) refer -to skills in the chain. Internal workflow steps within THIS skill are -labelled "(a)" through "(g)". +Throughout this document, "step 1" through "step 9" (no parens) +refer to skills in the chain. Internal workflow steps within THIS +skill are labelled "(a)" through "(g)". ## What you produce @@ -35,8 +36,7 @@ classification: `document`, `prose-guidance`, `meta`, `interactive`. If none fits cleanly, write a short descriptive label of your own (`dataset-extraction`, `deployment-runbook`, etc.) rather than -falling back to `other` — a specific label gives downstream steps a -real handle. +falling back to `other`. For the body template, see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). @@ -49,27 +49,39 @@ If the source is upstream, also vendor the fetched skill files to Upstream skill (URL or `//`) or local skill (filesystem path that already exists)? If the user gave a bare name -with no URL and no path, ask to clarify before continuing. +with no URL and no path, ask to clarify. + +If the URL 404s or the local path doesn't exist, surface the error +to the user — don't guess at recovery. If the source is a plugin +with multiple skills, ask which one to investigate (produce one +report per skill). ### (b) Ask about PR intent -**Upstream skills:** ask the user "Do you want to optimize this -skill for upstream PR submission?" and record the answer in the -report frontmatter as `pr_submission_intent: true|false`. Capture -this decision now — step 2's handoff reads this field to decide -whether step 3 (`investigate-submissions`) runs. +**Upstream skills:** ask "Do you want to optimize this skill for +upstream PR submission?" and record the answer in the report +frontmatter as `pr_submission_intent: true|false`. Capture this +decision now — step 2's handoff reads this field to decide whether +step 3 runs. **Local skills:** default to `pr_submission_intent: false`. But if the user explicitly said they want to send this back to an upstream -maintainer, treat it as PR-intent: set `pr_submission_intent: true`, -ask where the upstream contribution guidelines live, and record -their answer in the report body under a "PR submission notes" -subsection (step 3 uses this as its starting point). +maintainer, treat it as PR-intent: set true, ask where the upstream +contribution guidelines live, record their answer in the report +body under a "PR submission notes" subsection (step 3 uses this as +its starting point). + +If the user later changes their mind on PR intent, they re-run this +skill — the canonical is overwritten and prior state lives in git +history. ### (c) Vendor the source (upstream only) Fetch the skill's files into `vendored-skill/` at the working-directory root. Local-source skills don't need vendoring. +If `vendored-skill/` already exists from a prior run, reuse it +unless the source URL changed or the user explicitly asks to +re-fetch. ### (d) Determine the slug and the report path @@ -82,18 +94,15 @@ Report path: `docs/skill-optimizer//01-functionality.md`. ### (e) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 1 is **fresh-derivation**: each invocation -overwrites the canonical from the current source + directives; git -captures prior state. Re-run on user signal (source URL changed, -PR-intent answer changed, or new directives). Collect +and apply it. Each invocation overwrites the canonical from the +current source plus directives; git captures prior state. Collect `${OPERATOR_DIRECTIVES}` per the protocol. ### (f) Dispatch the functionality-researcher subagent Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT do the research yourself in this -session** — dispatch the subagent via the `Agent` tool (with -worktree isolation if your environment supports it). Load the +session** — dispatch the subagent via the `Agent` tool. Load the prompt template at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md) and substitute `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, @@ -120,29 +129,3 @@ Verify the report file exists and frontmatter parses. Then: > Next, invoke `skill-optimizer-investigate-test-case`. Note: this > report's `pr_submission_intent` field tells step 2's handoff > whether step 3 should run. - -## Edge cases - -- **URL 404 or local path doesn't exist** — surface the error to the - user; don't try to guess a recovery. -- **Source is a plugin with multiple skills** — ask which one to - investigate; produce one report per skill. -- **`vendored-skill/` already exists from a prior run** — reuse it - unless the source URL changed or the user explicitly asks to - re-fetch. -- **User changes their mind on PR intent later** — they re-run this - skill; the canonical is overwritten and the prior state lives in - git history. - -## Iteration behavior - -Fresh-derivation step. Re-run triggers: - -- Source URL changed (different upstream skill, or repo moved) -- User changed their mind on PR intent -- User wants the report re-derived with new directives ("you missed - the vendor CLA requirement", "go deeper on who-uses-this") - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (e) loads it. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index c6464e0..532b795 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -6,12 +6,12 @@ description: Use when the user wants to research a skill's upstream PR conventio # skill-optimizer-investigate-submissions Step 3 of the skill-optimizer chain — OPTIONAL, runs only when the -target skill is bound for upstream PR submission. Takes the source -slug, dispatches a researcher subagent that uses the `gh` CLI to -gather the upstream repo's contribution conventions, and writes -`docs/skill-optimizer//03-submissions.md` — the -verbatim-pastable context block the validator (step 8) uses for its -external consistency check. +target skill is bound for upstream PR submission. **Fresh-derivation +step.** Takes the source slug, dispatches a researcher subagent that +uses the `gh` CLI to gather the upstream repo's contribution +conventions, and writes `docs/skill-optimizer//03-submissions.md` +— the verbatim-pastable context block the validator (step 8) uses for +its external consistency check. Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain. Internal workflow steps within THIS @@ -54,18 +54,17 @@ Two prerequisites: 1. `01-functionality.md` must exist with valid frontmatter. If not, tell the user to run - `skill-optimizer-investigate-functionality` first and stop here. + `skill-optimizer-investigate-functionality` first. 2. That report's `pr_submission_intent` field must be `true`. If - it's `false`, this skill should not run — tell the user step 3 - is skipped for local-only optimization runs. + `false`, this skill should not run — tell the user step 3 is + skipped for local-only optimization runs. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 3 is **fresh-derivation**: each invocation -overwrites the canonical from current upstream facts + directives. -Staleness against `01-functionality.md` is determined by git-mtime -comparison. +and apply it. Each invocation overwrites the canonical from +current upstream facts plus directives. Staleness against +`01-functionality.md` is determined by git-mtime comparison. This report rarely needs re-running — upstream PR conventions change slowly. Most common reason: upstream updated @@ -102,16 +101,27 @@ justify the proposed change. A subagent walled off from that produces neutral upstream facts; the validator's independence is preserved. -The subagent writes the report and returns a brief summary: -license, CLA requirement, branch target, any high-risk rejection -signals from the closed-without-merge PRs. +Edge cases the subagent will surface as blockers: + +- **Upstream repo is private or requires auth** — ask the user to + authenticate `gh` and re-dispatch +- **No merged PRs in the upstream's history yet** — the report + will be thinner; the subagent says so honestly rather than + making up patterns +- **Upstream uses non-discoverable frontmatter conventions** — if + the subagent can't extract a consistent spec from recent merged + PRs, the report will say so; the optimizer (step 7) then makes a + judgment call rather than mechanically conforming + +The subagent returns a brief summary: license, CLA requirement, +branch target, any high-risk rejection signals. ### (d) Confirm subagent output Verify the report exists with parseable frontmatter and the expected sections. If `requires_cla: true`, mention it explicitly -on handoff — the operator will need to sign the CLA before any PR -can be merged. +on handoff — the operator needs to sign the CLA before any PR can +be merged. ### (e) Hand off @@ -122,35 +132,11 @@ target / any flagged blockers), then: > this report for its external consistency check. If you haven't > run `skill-optimizer-write-tests` yet, invoke that next. -The chain doesn't enforce ordering between step 3 and step 4 — -they can run in either order or in parallel. +The chain doesn't enforce ordering between step 3 and step 4. ## Edge cases -- **Upstream repo is private or requires auth** — the subagent - surfaces this as a blocker; ask the user to authenticate `gh` - and re-dispatch. - **Upstream uses a non-`gh`-friendly host (GitLab, Bitbucket, etc.)** — the current chain assumes GitHub-hosted upstreams. Surface this and tell the user; non-GitHub cases require manual - research. -- **No merged PRs in the upstream's history yet** — the researcher - can't extract shape patterns from absent data; the report will - be thinner. Surface honestly rather than making up patterns. -- **Upstream uses non-discoverable frontmatter conventions** — if - the subagent can't extract a consistent spec from recent merged - PRs, the report will say so. The optimizer (step 7) then has to - make a judgment call rather than mechanically conform. - -## Iteration behavior - -Fresh-derivation step. Rarely needs re-running. Re-run triggers: - -- Upstream repo updated `CONTRIBUTING.md`, license, or CLA -- Upstream's PR-shape conventions visibly shifted -- Step 1's `01-functionality.md` changed because the source URL - was updated (different upstream repo entirely) - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. + research and pasting the report content directly. diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 784e4b8..f589cf5 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -5,12 +5,12 @@ description: Use when the user wants to design or propose test cases for a skill # skill-optimizer-investigate-test-case -Step 2 of the skill-optimizer chain. Takes the functionality report -from step 1, dispatches a designer subagent to enumerate the skill's -responsibilities and propose a ranked set of **functionalities** to -test (each functionality = one responsibility the skill must -fulfill), then asks the user to pick which ones to actually build -probes for at step 4. +Step 2 of the skill-optimizer chain. **Maintenance step.** Takes the +functionality report from step 1, dispatches a designer subagent to +enumerate the skill's responsibilities and propose a ranked set of +**functionalities** to test (each functionality = one responsibility +the skill must fulfill), then asks the user to pick which ones to +actually build probes for at step 4. The filesystem IS the state: this step writes a `tests//spec.yaml` per proposed functionality @@ -63,19 +63,23 @@ For the body template of `02-test-proposals.md` and the exact `01-functionality.md` must exist. If not, tell the user to run `skill-optimizer-investigate-functionality` first and stop here. -If `docs/skill-optimizer//tests/` doesn't exist yet (first -invocation on this slug), create the empty directory. +If `docs/skill-optimizer//tests/` doesn't exist yet, create +the empty directory. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 2 is a **maintenance** step — re-runs read the -current `tests/` tree as state and extend or modify it. Collect +and apply the maintenance-step flow. Re-runs read the current +`tests/` tree as state and extend or modify it. Collect `${OPERATOR_DIRECTIVES}` per the protocol. If a directive will be destructive (e.g., "remove the X functionality"), commit a checkpoint before dispatching per the protocol's safe-destructive- edits section. +If the user's directives contradict each other (e.g., "focus on X" +combined with "ignore X"), surface the contradiction before +re-dispatching; don't try to resolve it yourself. + ### (c) Dispatch the test-case-designer subagent Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) @@ -121,6 +125,12 @@ Verify `02-test-proposals.md` parses and each new If any spec.yaml fails to parse, surface to the user; don't repair the subagent's output yourself. +If the proposal has fewer functionalities than expected, the +functionality report is likely thin. Surface to the user with two +options: re-run step 1 with directives, or accept if the skill is +genuinely small. Per the no-auto-invocation rule, don't re-invoke +step 1 yourself. + ### (e) User gate: present proposals, collect picks Show the user the ranked list from `02-test-proposals.md` and tell @@ -136,10 +146,10 @@ Three realistic responses: what you set. The rule is "no proactive flipping without user direction", not "user must edit every file by hand". 2. **User wants additions or revisions** (e.g., "split X into - two", "add a functionality for Y"). Treat their feedback as - directives. Return to (b) and re-dispatch the subagent at (c). - Existing edited spec.yaml files are preserved per the - maintenance rule. + two", "add a functionality for Y"). Treat as directives. + Return to (b) and re-dispatch the subagent at (c). Don't write + the spec.yaml manually yourself — the subagent is the only + writer of test-design content. 3. **User picks zero.** Ask whether they want a revised proposal (case 2) or are abandoning the optimization for this skill (exit honestly). @@ -158,34 +168,7 @@ PR?" prompts — the decision was made at step 1. ## Edge cases -- **`01-functionality.md` missing** — tell the user to run step 1 - first; don't try to derive responsibilities yourself. -- **Subagent's proposal has fewer functionalities than expected** — - signal the functionality report is thin. Surface to the user - with two options: re-run step 1 with directives, or accept if - the skill is genuinely small. Per the no-auto-invocation rule in - [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), - don't re-invoke step 1 yourself. -- **User wants to add a functionality the subagent didn't propose** - — treat their description as a directive and re-dispatch per - step (e)(2). Don't write the spec.yaml manually yourself — the - subagent is the only writer of test-design content. -- **Operator directives contradict each other** — surface the - contradiction before re-dispatching. - **User removes a functionality folder manually** — fine, - filesystem-as-state working as intended. - -## Iteration behavior - -Maintenance step. Re-run triggers: - -- User wants different coverage (different proposals, ranking, or - probe suggestions) -- Step 4 flagged a picked functionality as unimplementable -- Step 5 showed all probes pass on baseline (too easy) -- Step 6 surfaced a coverage gap -- Step 1's `01-functionality.md` changed - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. + filesystem-as-state working as intended. On next re-run, the + subagent reads the current tree and doesn't propose the removed + one back unless directives ask. diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 6f96ec3..2f605ea 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -5,11 +5,12 @@ description: Use when the user wants to run the eval suite against a skill and c # skill-optimizer-run-bench -Step 5 of the skill-optimizer chain. Invokes the skill-optimizer -CLI's `run-suite` command against `tests/suite.yml` (generated by -step 4), captures the raw results under a timestamped directory, and -writes a small summary report that step 6 reads. No subagent -dispatch — this is a thin operator-driven CLI step. +Step 5 of the skill-optimizer chain. **Fresh-derivation step (for +the summary).** Invokes the skill-optimizer CLI's `run-suite` +command against `tests/suite.yml` (generated by step 4), captures +the raw results under a timestamped directory, and writes a small +summary report that step 6 reads. No subagent dispatch — this is a +thin operator-driven CLI step. Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain. Internal workflow steps within THIS @@ -45,7 +46,8 @@ Two outputs at `docs/skill-optimizer//`: ### (a) Confirm prerequisites -`tests/suite.yml` must exist (per step 4's output contract). +`tests/suite.yml` must exist (per step 4's output contract). If +missing, tell the user to complete step 4 first. `vendored-skill/` should exist for upstream skills. ### (b) Handle iteration @@ -56,12 +58,11 @@ and apply it. Two pieces: - **Raw bench output** at `05-bench-results//`: always write to a fresh timestamped directory (`date -u +%Y%m%dT%H%M%SZ`). Outside the iteration protocol; never overwrite. - -- **Summary file** at `05-bench-summary.md`: fresh-derivation - artifact per the protocol. Staleness check: if `tests/suite.yml` - has a newer git mtime than `05-bench-summary.md`, the existing - summary measured a different workbench and is stale. If summary - is current AND no re-run directive, tell the user and exit. +- **Summary file** at `05-bench-summary.md`: fresh-derivation per + the protocol. Staleness check: if `tests/suite.yml` has a newer + git mtime than `05-bench-summary.md`, the existing summary is + stale. If summary is current AND no re-run directive, tell the + user and exit. ### (c) Run the bench @@ -79,10 +80,17 @@ npx tsx /src/cli.ts run-suite \ `--trials 3` is the chain default — enough to distinguish flaky from systematic failures at step 6. Honor a different count if requested. Models come from `suite.yml` (per project invariant: -`run-suite` does NOT take a `--models` override). The CLI requires -`OPENROUTER_API_KEY`; surface env failures rather than recovering. -Stream stdout/stderr to the user — bench runs take minutes to -hours. +`run-suite` does NOT take a `--models` override). Stream +stdout/stderr to the user — bench runs take minutes to hours. + +Environment failures to surface honestly (don't try to recover): + +- **`OPENROUTER_API_KEY` not set** — CLI fails fast. Ask the user + to set it; don't mock. +- **Docker image missing** — default is + `skill-optimizer-workbench:local`. Tell the user to build it + (`docker build -t skill-optimizer-workbench:local -f + docker/workbench-runner.Dockerfile .`). ### (d) Write the summary @@ -98,6 +106,11 @@ with the frontmatter above plus a body containing: with paths to their `trace.jsonl` and `findings.txt`. - **Raw output:** the `bench_results_path` value. +If all trials errored (no graded results), that's a workbench +misconfiguration or environmental failure rather than a skill +weakness. Write the summary honestly and surface before handing +off to step 6. + ### (e) Hand off Read `overall_pass_rate`. Two messages: @@ -106,49 +119,13 @@ Read `overall_pass_rate`. Two messages: failures. - `== 1.0`: surface the choice — accept that probes don't expose a weakness, or re-run step 2 with a "make probes harder" - directive. Don't auto-invoke step 6; per the no-auto-invocation - rule in - [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md), - surface the choice. + directive. Don't auto-invoke step 6. ## Edge cases -- **`tests/suite.yml` missing** — tell the user to run step 4 - first; don't construct a suite manifest yourself. -- **`OPENROUTER_API_KEY` not set** — CLI fails fast; surface and - ask the user to set it. Don't mock. -- **Docker image missing** — default is - `skill-optimizer-workbench:local`. Tell the user to build it - (`docker build -t skill-optimizer-workbench:local -f - docker/workbench-runner.Dockerfile .`). -- **All trials errored (no graded results)** — workbench - misconfiguration or environmental failure, not a skill weakness. - Write the summary honestly and surface before handing off to - step 6. - -## Iteration behavior - -Re-run triggers: - -- `tests/` changed (step 4 added or revised probes) — detected via - git mtime -- User wants fresh trial data (flakiness suspicion, model list in - `suite.yml` changed) -- Step 6 flagged a result as inconclusive due to too few trials — - re-run with higher `--trials` - -By default a re-run measures the **entire** probe set. Unchanged -probes get fresh trial samples (useful for distinguishing flaky vs. -systematic at step 6), and the summary stays internally comparable. -Cost is "all probes × trials" per re-run. - -**Known limitation (deferred):** the CLI's `run-suite` does not -accept a case filter, so there's no first-class partial-rebench -mode. Operators with expensive suites can run `run-case` manually -for changed probes and splice into the prior `05-bench-results//` -dir — outside-the-chain escape hatch, not a supported flow. - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. Raw bench output is outside the -protocol; the summary is fresh-derivation. +- **Known limitation (deferred):** the CLI's `run-suite` does not + accept a case filter, so there's no first-class partial-rebench + mode. Operators with expensive suites can run `run-case` + manually for changed probes and splice into the prior + `05-bench-results//` dir — outside-the-chain escape hatch, + not a supported flow. diff --git a/skills/skill-optimizer-validate-improvement/SKILL.md b/skills/skill-optimizer-validate-improvement/SKILL.md index db22588..4b13d9a 100644 --- a/skills/skill-optimizer-validate-improvement/SKILL.md +++ b/skills/skill-optimizer-validate-improvement/SKILL.md @@ -5,12 +5,13 @@ description: Use when the user wants to validate an improvement proposal from `s # skill-optimizer-validate-improvement -Step 8 of the skill-optimizer chain. Takes the proposal from step -7, dispatches a validator subagent to independently check whether -the proposed change is sound (internal consistency) and conformant -(external PR conventions if PR-bound), then — on `verdict: approve` -— materializes the improved skill at `improved-skill/`. Writes -`08-validator-verdict.md` regardless of verdict. +Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes +the proposal from step 7, dispatches a validator subagent to +independently check whether the proposed change is sound (internal +consistency) and conformant (external PR conventions if PR-bound), +then — on `verdict: approve` — materializes the improved skill at +`improved-skill/`. Writes `08-validator-verdict.md` regardless of +verdict. **Single-shot per invocation.** No in-step revision loop. If the verdict is `needs-revision`, the operator (or auto-pilot at step 9) @@ -65,21 +66,29 @@ Three checks: first. 2. Current skill state must be readable: `improved-skill/` if it exists (prior accumulated state), else the original source. - The validator needs this as "skill BEFORE". + The validator needs this as "skill BEFORE". If both + `improved-skill/` and the source were modified externally + between step 7 and step 8, surface to the user — the BEFORE + must match what the optimizer read. The fix is re-invoking + step 7 against the new state. 3. `01-functionality.md` must exist. If `pr_submission_intent: true`, `03-submissions.md` should exist; if missing, the external check can't run. Ask whether to -run step 3 first or proceed with internal consistency only. +run step 3 first or proceed with internal consistency only (the +verdict body will note the omission). ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) -and apply it. Step 8 is **fresh-derivation**: each invocation -overwrites the canonical verdict. Collect `${OPERATOR_DIRECTIVES}` -— examples: "be stricter on additive-vs-destructive", "the -external check missed the frontmatter `version:` field — look for -it specifically". +and apply it. Each invocation overwrites the canonical verdict. +Collect `${OPERATOR_DIRECTIVES}` — examples: "be stricter on +additive-vs-destructive", "the external check missed the +frontmatter `version:` field — look for it specifically". + +If a directive references a specific section of the prior verdict, +that's a context-dump masquerading as a directive. Translate to +an atomic new requirement before passing. ### (c) Dispatch the validator subagent @@ -147,43 +156,10 @@ Three messages by verdict: auto-invoke each other." - **reject:** "Validator rejects outright. Two paths: re-invoke step 6 with a reframed weakness, or accept that this weakness - isn't addressable and exit honestly." + isn't addressable and exit honestly. If the rejection + contradicts the analyzer's framing (validator says 'this + doesn't address the real problem' but the analyzer thought it + did), the analyzer/validator are misaligned — step 6 reframe + is the cleaner fix." Don't auto-invoke step 6 or 7. - -## Edge cases - -- **`07-improvement-proposal.md` missing** — caught at (a). -- **Current skill state ambiguous** (both `improved-skill/` and - source modified externally between step 7 and step 8) — surface - to the user; the BEFORE must match what the optimizer read. If - the user edited the source mid-way, re-invoke step 7 to derive - a fresh proposal against the new state. -- **PR-bound but `03-submissions.md` missing** — caught at (a). - Proceeding loses the external check; the verdict body notes the - omission. -- **Validator rejects with reasoning that contradicts the - analyzer's framing** — surface with the two paths from the - `reject` handoff. -- **Operator directives reference a specific section of the prior - verdict** — context-dump masquerading as a directive. Translate - to an atomic new requirement before passing. - -## Iteration behavior - -Fresh-derivation step. Re-run triggers: - -- Step 7 produced a new proposal — - `07-improvement-proposal.md` changed -- User disagrees with the prior verdict and wants a fresh - derivation with directives -- `03-submissions.md` was updated and the prior external check is - stale - -No in-step revision loop; each invocation is single-shot. The -7→8→7 cycle on `needs-revision`/`reject` is operator-driven, or -auto-pilot-driven at step 9. - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 9333be4..722c431 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -5,12 +5,13 @@ description: Use when the user wants to build or implement the eval probes for a # skill-optimizer-write-tests -Step 4 of the skill-optimizer chain. Reads the picked functionalities -from step 2 (`tests//spec.yaml` files where -`picked: true`), decides the probe set per functionality, dispatches -one test-writer subagent per probe (in parallel) to build the -concrete workspace files + grader scripts + smoke fixtures, then -generates `tests/suite.yml` for the run-bench step. +Step 4 of the skill-optimizer chain. **Maintenance step.** Reads +the picked functionalities from step 2 +(`tests//spec.yaml` files where `picked: true`), +decides the probe set per functionality, dispatches one test-writer +subagent per probe (in parallel) to build the concrete workspace +files + grader scripts + smoke fixtures, then generates +`tests/suite.yml` for the run-bench step. Throughout this document, "step 1" through "step 9" (no parens) refer to skills in the chain. Internal workflow steps within THIS @@ -135,10 +136,15 @@ spec from step 2 fixes WHAT to test; source provides the HOW (concrete violation patterns). Without source, the subagent would invent generic patterns that may not trigger the skill's rules. -Each test-writer writes the probe folder contents (spec.yaml, -workspace files, grader, smoke fixtures, smoke runner) and -returns: probe name, what the fixture tests, grader logic in one -line, smoke-check result. +Each test-writer writes the probe folder contents and returns: +probe name, what the fixture tests, grader logic in one line, +smoke-check result. + +If a test-writer reports BLOCKED (probe spec too abstract to +derive a fixture from, or the probe is genuinely unimplementable +in a static workbench), surface to the user. The fix is typically +a step 2 re-run with a more specific functionality spec; per the +no-auto-invocation rule, the user invokes step 2 explicitly. ### (e) Run smoke check @@ -170,31 +176,7 @@ Commit `tests/` and hand off to `skill-optimizer-run-bench`. ## Edge cases -- **A test-writer reports BLOCKED** (probe spec too abstract to - derive a fixture from) — surface to the user. Fix is typically a - step 2 re-run with a more specific functionality spec; per the - no-auto-invocation rule, user invokes step 2. -- **A grader fails smoke check** — see step (e)'s handling. -- **A picked functionality is unimplementable in a static - workbench** (e.g., needs real-time API access) — test-writer - returns BLOCKED. Functionality may need to be de-picked at - step 2. - **User de-picks a functionality between step 4 runs** — its probe folders stay on disk; step (f) excludes them from the regenerated `tests/suite.yml`. If re-picked later, probes are already there. - -## Iteration behavior - -Maintenance step. Re-run triggers: - -- A new functionality at step 2 got `picked: true` -- A specific probe's smoke check failed and needs revision -- Step 5 showed probes systematically too easy / too hard — - directives trigger probe-level rebuilds -- Step 2 substantively revised a picked functionality's spec.yaml - in a way that warrants rebuilding existing probes - -Mechanics live in -[`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md); -workflow step (b) loads it. From 169abca69c73869876cc6993e68be3ea0dec0160 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 08:49:57 -0500 Subject: [PATCH 036/121] refactor(docs): move chain-specific docs out of top-level docs/ Top-level docs/ should hold project-wide reading; chain-specific design docs belong elsewhere: - docs/skill-optimizer-v1.4-spec.md -> docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md (matches the superpowers brainstorming skill's convention: docs/superpowers/specs/YYYY-MM-DD--design.md) - docs/skill-optimizer-v1.4-plan.md -> docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md (matches the superpowers writing-plans skill's convention: docs/superpowers/plans/YYYY-MM-DD-.md) - docs/skill-optimizer-workflow.md -> skills/skill-optimizer-shared/workflow.md (chain-specific operator reference belongs alongside the chain it documents; co-located with iteration-protocol.md + subagent-dispatch.md + frontmatter-discipline.md) Top-level docs/ now holds only: workbench.md # project-wide workbench engine guide README.codex.md # install README.opencode.md # install skill-writing-philosophy.md # project-wide authoring guidance images/ # project-wide assets superpowers/ # superpowers-plugin-managed dir pilot-runs/ # project-wide Internal references updated: - The spec doc's "Companion docs" section now reflects the new layout (project-wide vs chain-specific) - Bulk sed across spec + plan to point at new paths - workflow.md's relative paths fixed for its new location (../skills/ -> ./ since it now lives in skills/skill-optimizer-shared/) Date chosen (2026-05-19) is the original git creation date of both the spec and plan files, matching the YYYY-MM-DD convention. --- .../plans/2026-05-19-skill-optimizer-v1.4.md} | 42 +++++++++---------- ...2026-05-19-skill-optimizer-v1.4-design.md} | 21 +++++++--- .../skill-optimizer-shared/workflow.md | 10 ++--- 3 files changed, 41 insertions(+), 32 deletions(-) rename docs/{skill-optimizer-v1.4-plan.md => superpowers/plans/2026-05-19-skill-optimizer-v1.4.md} (96%) rename docs/{skill-optimizer-v1.4-spec.md => superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md} (97%) rename docs/skill-optimizer-workflow.md => skills/skill-optimizer-shared/workflow.md (92%) diff --git a/docs/skill-optimizer-v1.4-plan.md b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md similarity index 96% rename from docs/skill-optimizer-v1.4-plan.md rename to docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md index fceda9b..cd05b21 100644 --- a/docs/skill-optimizer-v1.4-plan.md +++ b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md @@ -43,7 +43,7 @@ We use **both** `skill-creator` and `superpowers:writing-skills` in v1.4 — the **Working dir:** `.claude/worktrees/v1.4-spec/` on branch `feat/skill-optimizer-v1.4`. -**Spec:** `docs/skill-optimizer-v1.4-spec.md` (committed at `a01d0cc`). +**Spec:** `docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md` (committed at `a01d0cc`). --- @@ -128,7 +128,7 @@ Each file's responsibility: - **`skills/skill-optimizer-/SKILL.md`** — frontmatter (`name:`, `description:`) + invocation instructions + dispatch-subagent logic + handoff-to-next-skill pointer. Always-loaded by Claude Code when the skill is invoked. - **`skills/skill-optimizer-subagents/.md`** — narrow-context prompt template that the corresponding skill loads, substitutes inputs into, and dispatches via Agent tool. Scoped name prevents collision with other plugins that might also ship subagents. -- **`docs/skill-optimizer-workflow.md`** — human-readable chain diagram + per-skill brief, for operators understanding how the 7 skills compose. Lives under `docs/` because it's contributor/operator reading, not a runtime resource the skills load. +- **`skills/skill-optimizer-shared/workflow.md`** — human-readable chain diagram + per-skill brief, for operators understanding how the 7 skills compose. Lives under `docs/` because it's contributor/operator reading, not a runtime resource the skills load. - **`docs/skill-writing-philosophy.md`** — contributor reference distilling the three philosophies (Anthropic, skill-creator, superpowers:writing-skills), the skill-type → philosophy mapping, the bootstrapping limit, and the rules the optimizer subagent must follow when revising target skills. Also under `docs/` for the same reason. --- @@ -164,10 +164,10 @@ description: > # skill-optimizer- - + ``` -Verb → description mapping (from `docs/skill-optimizer-v1.4-spec.md`): +Verb → description mapping (from `docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md`): ```bash # Run this script to populate all 7 shells: @@ -191,7 +191,7 @@ description: ${DESCS[$verb]} # skill-optimizer-${verb} - + EOF done ``` @@ -225,7 +225,7 @@ git commit -m "feat(v1.4): create 7 skill directory shells with frontmatter Body content for each SKILL.md will be written interactively in Phase B via superpowers:writing-skills, with the interface contract from -docs/skill-optimizer-v1.4-spec.md as the brief. This commit just lays +docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md as the brief. This commit just lays down the structural skeleton and discoverable frontmatter." ``` @@ -349,7 +349,7 @@ git commit -m "feat(v1.4-shared): iteration-protocol.md — shared mechanics for **Do NOT dispatch these as autonomous subagents** — both tools are built for human-in-loop refinement, and the user explicitly stated "writing good skills is HARD." -For EACH Task B, the brief for `skill-creator` is the interface contract for that skill from `docs/skill-optimizer-v1.4-spec.md` section "The 7 skills", plus the constraints from "Subagent constraints" if the skill dispatches a subagent. Paste those excerpts verbatim into the `skill-creator` capture step. +For EACH Task B, the brief for `skill-creator` is the interface contract for that skill from `docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md` section "The 7 skills", plus the constraints from "Subagent constraints" if the skill dispatches a subagent. Paste those excerpts verbatim into the `skill-creator` capture step. --- @@ -374,7 +374,7 @@ Brief (paste into skill-creator's capture step): ```text Create SKILL.md body for skills/skill-optimizer-investigate-functionality/. -Interface contract (from docs/skill-optimizer-v1.4-spec.md "The 7 skills" → §1): +Interface contract (from docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md "The 7 skills" → §1): - Description trigger: "investigate / understand / explain what this skill does (and the context around it)" - Input: source skill (URL or local path) @@ -436,7 +436,7 @@ Brief: ```text Create SKILL.md body for skills/skill-optimizer-investigate-test-case/. -Interface contract (from docs/skill-optimizer-v1.4-spec.md §2): +Interface contract (from docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md §2): - Description trigger: "design tests for this skill", "propose test cases", "what should we test" - Input: docs/skill-optimizer//01-functionality.md @@ -721,7 +721,7 @@ description: Use when the user wants to run the full skill-optimizer chain on a # skill-optimizer-autopilot - + EOF ``` @@ -732,7 +732,7 @@ EOF ```text Goal: write the body of skills/skill-optimizer-autopilot/SKILL.md. -Interface contract: see docs/skill-optimizer-v1.4-spec.md §"The skills" #8. +Interface contract: see docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md §"The skills" #8. Key points to cover in the body: - This is the 8th skill — the auto-pilot driver, not a chain step. @@ -829,7 +829,7 @@ You are dispatched to research what a single agent skill does and the context ar ## Tools NOT allowed -You operate under limited context per `docs/skill-optimizer-v1.4-spec.md` "Subagent constraints" table. You see ONLY the source skill + targeted web fetches. You do NOT see: +You operate under limited context per `docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md` "Subagent constraints" table. You see ONLY the source skill + targeted web fetches. You do NOT see: - Existing analyses or improvement proposals - Existing tests for this skill @@ -1611,13 +1611,13 @@ git commit -m "feat(v1.4-subagents): validator prompt template" ## Phase D — Workflow doc (1 task, INTERACTIVE authoring) -**Mode:** Direct interactive authoring with operator review per section. No skill tool. `docs/skill-optimizer-workflow.md` is a human reference doc — not a behavioral skill (so skill-creator's description-routing/eval loop doesn't apply) and not a compliance prompt template (so writing-skills' pressure scenarios don't apply). The right discipline is a careful walkthrough with the operator: present each section (chain diagram, invariants, etc.), get confirmation, iterate, commit. +**Mode:** Direct interactive authoring with operator review per section. No skill tool. `skills/skill-optimizer-shared/workflow.md` is a human reference doc — not a behavioral skill (so skill-creator's description-routing/eval loop doesn't apply) and not a compliance prompt template (so writing-skills' pressure scenarios don't apply). The right discipline is a careful walkthrough with the operator: present each section (chain diagram, invariants, etc.), get confirmation, iterate, commit. **Do NOT dispatch as autonomous subagent.** The operator must validate that the diagram + invariants match the actual Phase B SKILL.md handoffs and Phase C subagent constraints (which were finalized interactively just prior). -### Task D1: `docs/skill-optimizer-workflow.md` — human-readable chain diagram +### Task D1: `skills/skill-optimizer-shared/workflow.md` — human-readable chain diagram -**Files:** Create `docs/skill-optimizer-workflow.md`. +**Files:** Create `skills/skill-optimizer-shared/workflow.md`. - [ ] **Step 1: Walk through each section with the operator, then write the file** @@ -1627,7 +1627,7 @@ Use the draft below as the starting point. Before writing, walk through each sec 2. Subagent constraint table — confirm wording matches the finalized prompts from Phase C (which may have shifted during pressure-scenario iteration). 3. Invariants section — surface any new invariants discovered while writing Phases B/C. -Then create `docs/skill-optimizer-workflow.md` with the agreed content. Draft below: +Then create `skills/skill-optimizer-shared/workflow.md` with the agreed content. Draft below: ````markdown # skill-optimizer v1.4 workflow @@ -1718,7 +1718,7 @@ All artifacts at convention path: `docs/skill-optimizer//` ## Limited-context subagents -The architectural fix for ducttape. See `docs/skill-optimizer-v1.4-spec.md` "Subagent constraints" for the full table. Quick summary: +The architectural fix for ducttape. See `docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md` "Subagent constraints" for the full table. Quick summary: - **Optimizer & validator** never see raw trial data or grader internals — forces principled improvement - **Test writer** never sees the skill content — prevents grader-hacking @@ -1736,8 +1736,8 @@ No separate auto-pilot skill — it's just chained `Skill` tool invocations driv - [ ] **Step 2: Verify** ```bash -wc -l docs/skill-optimizer-workflow.md -grep -c "^## " docs/skill-optimizer-workflow.md +wc -l skills/skill-optimizer-shared/workflow.md +grep -c "^## " skills/skill-optimizer-shared/workflow.md ``` Expected: ≥ 80 lines, ≥ 4 H2 sections. @@ -1745,7 +1745,7 @@ Expected: ≥ 80 lines, ≥ 4 H2 sections. - [ ] **Step 3: Commit** ```bash -git add docs/skill-optimizer-workflow.md +git add skills/skill-optimizer-shared/workflow.md git commit -m "feat(v1.4-references): add workflow.md (chain diagram + state file map)" ``` @@ -1934,7 +1934,7 @@ No issues found. ## Execution Handoff -Plan complete and saved to `docs/skill-optimizer-v1.4-plan.md` (in worktree `.claude/worktrees/v1.4-spec/`, branch `feat/skill-optimizer-v1.4`). +Plan complete and saved to `docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md` (in worktree `.claude/worktrees/v1.4-spec/`, branch `feat/skill-optimizer-v1.4`). **Phase-by-phase execution mode (see header table for tool assignment):** diff --git a/docs/skill-optimizer-v1.4-spec.md b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md similarity index 97% rename from docs/skill-optimizer-v1.4-spec.md rename to docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md index 47c92af..8e09339 100644 --- a/docs/skill-optimizer-v1.4-spec.md +++ b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md @@ -87,14 +87,23 @@ skills/ # every chain skill loads ``` -Companion docs live under `docs/`, not `skills/`, because they're -contributor / operator reading rather than runtime resources the -skills load: +Companion docs follow two conventions: + +- **Project-wide reading** (e.g., authoring philosophy applicable + to any skill in the project) lives under `docs/`. +- **Chain-specific reading** (operator reference for THIS chain; + chain-internal design history) lives alongside the chain. ```text docs/ - skill-optimizer-workflow.md # the chain graph - skill-writing-philosophy.md # authoring guidance + skill-writing-philosophy.md # project-wide authoring guidance + superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md # THIS doc (design history) + superpowers/plans/2026-05-19-skill-optimizer-v1.4.md # implementation plan +skills/skill-optimizer-shared/ + workflow.md # the chain graph (operator reference) + iteration-protocol.md # iteration mechanics (loaded by skills) + subagent-dispatch.md # dispatch architecture (loaded by skills) + frontmatter-discipline.md # frontmatter rules (loaded by skills) ``` A shared `recipes.md` (named abstract failure patterns shared across @@ -737,7 +746,7 @@ For v1.4 to be considered done: 2b. `skills/skill-optimizer-shared/iteration-protocol.md` exists and is referenced explicitly (with a "Read this now" instruction) from each chain skill's "Handle iteration" step -3. `docs/skill-optimizer-workflow.md` documents the chain visually +3. `skills/skill-optimizer-shared/workflow.md` documents the chain visually 4. ~~`references/recipes.md` is seeded from v1.3's `lessons.md`~~ — deferred; the analyzer / optimizer subagents operate without a pre-loaded recipe library until real end-to-end observations diff --git a/docs/skill-optimizer-workflow.md b/skills/skill-optimizer-shared/workflow.md similarity index 92% rename from docs/skill-optimizer-workflow.md rename to skills/skill-optimizer-shared/workflow.md index e0b888c..d214222 100644 --- a/docs/skill-optimizer-workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -26,7 +26,7 @@ whether step 3 will run. No late prompts. ## Re-run triggers (when each step should be re-invoked) A chain skill never auto-invokes another chain skill (per -[`subagent-dispatch.md`](../skills/skill-optimizer-shared/subagent-dispatch.md) +[`subagent-dispatch.md`](./subagent-dispatch.md) re-run authorization). The operator (or auto-pilot at step 9) decides when to re-run each step. Common triggers: @@ -93,18 +93,18 @@ docs/skill-optimizer// `` for local skills. Filesystem IS the state; history is git (no `version:` fields or `archive/` directories per -[`frontmatter-discipline.md`](../skills/skill-optimizer-shared/frontmatter-discipline.md)). +[`frontmatter-discipline.md`](./frontmatter-discipline.md)). ## Shared docs The chain skills load these on-demand at the workflow steps that need them: -- [`iteration-protocol.md`](../skills/skill-optimizer-shared/iteration-protocol.md) +- [`iteration-protocol.md`](./iteration-protocol.md) — iteration mechanics (staleness, step kinds, destructive-edit checkpoints, cascading, bootstrapping) -- [`subagent-dispatch.md`](../skills/skill-optimizer-shared/subagent-dispatch.md) +- [`subagent-dispatch.md`](./subagent-dispatch.md) — subagent constraints, operator directives, templated dispatch inputs, no-auto-invocation rule -- [`frontmatter-discipline.md`](../skills/skill-optimizer-shared/frontmatter-discipline.md) +- [`frontmatter-discipline.md`](./frontmatter-discipline.md) — runtime facts vs. history rule From 165b053d8f6727e2e788490e9364583975599cc3 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 08:58:26 -0500 Subject: [PATCH 037/121] refactor: nuke legacy skills/skill-optimizer/ + co-locate workbench.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The legacy canonical skill (skills/skill-optimizer/SKILL.md) was a v1.3-era artifact — a direct workbench-CLI wrapper. Its role under the v1.4 chain is filled by skill-optimizer-run-bench. The v1.4 spec already called for this removal "in a separate cleanup PR" after validation; doing it now while the chain layout is being restructured anyway. - Moved skills/skill-optimizer/references/workbench.md -> skills/skill-optimizer-shared/workbench.md (load-bearing — B4 references the workbench schema reference; co-located with the chain's other shared docs) - Updated B4's pointer to the new location - Deleted skills/skill-optimizer/ (folder) Git history preserves the deleted SKILL.md content if needed. NOT updated in this commit (separate cleanup needed before merge): - Plugin metadata still references skills/skill-optimizer/SKILL.md in .claude-plugin/, .codex-plugin/, .cursor-plugin/, .opencode/, gemini-extension.json - CLAUDE.md mentions skills/skill-optimizer/SKILL.md as canonical - README.md and CONTRIBUTING.md may reference it These references need updating to point at the v1.4 chain (or the chain's entry point) before this work merges to development. Flagged here so they're not forgotten. --- .../workbench.md | 0 skills/skill-optimizer-write-tests/SKILL.md | 2 +- skills/skill-optimizer/SKILL.md | 210 ------------------ 3 files changed, 1 insertion(+), 211 deletions(-) rename skills/{skill-optimizer/references => skill-optimizer-shared}/workbench.md (100%) delete mode 100644 skills/skill-optimizer/SKILL.md diff --git a/skills/skill-optimizer/references/workbench.md b/skills/skill-optimizer-shared/workbench.md similarity index 100% rename from skills/skill-optimizer/references/workbench.md rename to skills/skill-optimizer-shared/workbench.md diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 722c431..555f077 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -48,7 +48,7 @@ Probe-level `spec.yaml` format and the workbench schema for `suite.yml` are defined in [`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) and -[`skills/skill-optimizer/references/workbench.md`](../skill-optimizer/references/workbench.md) +[`skills/skill-optimizer-shared/workbench.md`](../skill-optimizer-shared/workbench.md) respectively. ## Workflow diff --git a/skills/skill-optimizer/SKILL.md b/skills/skill-optimizer/SKILL.md deleted file mode 100644 index b376e89..0000000 --- a/skills/skill-optimizer/SKILL.md +++ /dev/null @@ -1,210 +0,0 @@ ---- -name: skill-optimizer -description: Use when creating, running, debugging, or documenting skill-optimizer workbench evals; working with agent skill cases, suites, graders, traces, Docker workspaces, OpenRouter model matrices, or the skill-optimizer SDK/CLI. ---- - -# skill-optimizer - -`skill-optimizer` is an eval workbench for agent skills. It runs a model in an isolated Docker `/work` directory, provides skills/references as normal workspace files, captures an agent trace, and grades deterministic local outcomes. - -Use this skill as the source of truth for authoring eval suites in this repo. Detailed schema and patterns are in `references/workbench.md`. - -## Core Model - -- A case is one user-like task plus one or more deterministic graders. -- A suite is a set of cases and OpenRouter models to run as a matrix. -- `references` are copied into `/work` before the agent starts; this is where eval skills live. -- The agent phase sees `/work` only. It cannot see `/case`, `/results`, graders, hidden answers, or hidden metadata. -- Cases can define `mcpServers`; these are exposed through a workbench `mcp` command during the agent phase. -- Graders run after the agent with `/case`, `/work`, and `/results` mounted. -- `trace.jsonl` is the debugging source for what the agent saw, said, and did. - -## Commands - -| Goal | Command | -|------|---------| -| Install deps | `npm install` | -| Build CLI | `npm run build` | -| Run one case | `npx tsx src/cli.ts run-case ` | -| Run one case across models | `npx tsx src/cli.ts run-case --models openrouter/google/gemini-2.5-flash,openrouter/openai/gpt-5.4` | -| Run a suite | `npx tsx src/cli.ts run-suite ` | -| CLI help | `npx tsx src/cli.ts --help` | - -Rules: - -- Use only `openrouter/...` model refs. -- `OPENROUTER_API_KEY` is required for real model runs. -- `run-suite` uses `models:` from `suite.yml`; it has no model override flag. -- `run-case` can use its case `model:` or `--model` / `--models`. -- Docker image default is `skill-optimizer-workbench:local`. - -## Install This Skill - -This repository ships one canonical skill at `skills/skill-optimizer/SKILL.md` plus plugin metadata for Claude Code, OpenCode, Codex, Cursor, and Gemini. - -Install the skill for common agents with: - -```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a claude-code -a opencode -a codex -a cursor -``` - -Plugin entrypoints: - -- Claude Code: `.claude-plugin/plugin.json` and `.claude-plugin/marketplace.json` -- OpenCode: `.opencode/plugins/skill-optimizer.js` -- Codex: `.codex-plugin/plugin.json` -- Cursor: `.cursor-plugin/plugin.json` -- Gemini: `gemini-extension.json` and `GEMINI.md` - -## Authoring Workflow - -1. Create `suite.yml` with `models`, shared defaults, and inline cases or case paths. -2. Put the skill/reference material under `references/`; it will be copied into `/work`. -3. Write natural user tasks. Do not mention graders, hidden answers, `/case`, or eval internals. -4. Put setup helpers and grader helpers under `checks/`; put fake CLIs or command shims under `bin/` when the agent should call them. -5. Add one or more `graders` per case. Prefer small deterministic graders over one broad grader. -6. Run `run-suite --trials ` and inspect `suite-result.json`, failing `result.json`, `summary.json`, and `trace.jsonl`. - -Variables listed in `env` are forwarded unchanged into setup, agent, grading, and cleanup containers. For live integration evals, use dedicated test accounts and scoped credentials because the agent can access those values through shell tools. Treat `trace.jsonl`, `result.json`, grader evidence, stdout/stderr, and preserved `workspace/` directories as potentially sensitive if an agent or grader prints or writes secret values. - -Use `mcpServers` when the task should interact with MCP tools. For local servers whose source should stay hidden from the agent, put server files under the case `mcp/` support directory and define `mcpServices`; Docker starts those as separate service containers and the agent only sees their HTTP MCP URL. Direct stdio `mcpServers.command` entries run inside the agent container and are only appropriate when the server implementation is intentionally agent-visible. Remote HTTP/SSE servers must be reachable from Docker. The workbench generates `/work/mcporter.json` with `imports: []`, so host/user MCP configs are not imported. OAuth/browser auth is not supported; use env/header credentials listed in `env`. - -Prefer the real CLI/API/service when you do not know its internal behavior well enough to mock it faithfully. Mock only when you are sure the mock matches the real command surface, validation, outputs, and failure modes; otherwise the eval will measure the mock, not the skill. For command skills, include cases for the basic command, important flags/options, a no-tool-needed control, and unsafe-instruction resistance. - -## Minimal Suite - -```yaml -name: pdf-skill-eval -references: ./references -models: - - openrouter/google/gemini-2.5-flash -env: - - OPENROUTER_API_KEY -timeoutSeconds: 600 -setup: - - node $CASE/checks/create-inputs.mjs -appendSystemPrompt: | - Keep task outputs at the top level of /work unless the user asks otherwise. -cases: - - name: extract-pdf-facts - task: | - Read statement.pdf and write answer.json with the account, quarter, approval code, and risk flags. - graders: - - name: answer-json - command: node $CASE/checks/extract-pdf-facts.mjs -``` - -## Directory Layout - -```text -my-eval/ - suite.yml - references/ - my-skill/SKILL.md - checks/ - create-inputs.mjs - extract-pdf-facts.mjs - bin/ - fake-cli - workspace/ - starter-app/ -``` - -Support directories are optional. `checks/` is mounted read-only at `/case/checks` for setup/grading. `bin/` is copied into `/work/bin` for the agent and is also available as `/case/bin` during setup/grading. `workspace/` is copied into `/work` after `references/`. - -## Grader Contract - -Graders are shell commands. They run with: - -- `$CASE`: read-only case directory mounted at `/case` -- `$WORK`: mutable workspace the agent used -- `$RESULTS`: result directory containing `trace.jsonl` - -Preferred grader output: - -```json -{ "pass": true, "score": 1, "evidence": ["answer matched"] } -``` - -If no JSON object is printed, exit code `0` passes and non-zero fails. Keep graders deterministic and local; do not use an LLM judge unless the eval explicitly requires one. - -Graders are the acceptance contract. They should evaluate evidence in `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and any relevant result-state files under `$RESULTS`. - -## Outputs - -```text -.results// - suite-result.json # run-suite aggregate - run-result.json # run-case matrix aggregate - trials/----001/ - trace.jsonl # agent messages and tool calls - result.json # pass, score, evidence, graders, metrics - summary.json # final text, failed graders, commands - workspace/ # failures or --keep-workspace -``` - -Use `trace.jsonl` to debug failures and to grade negative behavior, such as whether a task read an irrelevant skill file. - -## Optimization Loop - -After a run, inspect failing `result.json`, `summary.json`, `trace.jsonl`, and preserved `workspace/` evidence. Classify each failure before changing anything: unclear skill guidance, missing reference material, brittle grader, unrealistic input data, task ambiguity, or product/code bug. Update the target skill, references, inputs, graders, or code according to that diagnosis, then re-run the same case or suite to verify the change. Repeat until the grader evidence shows the intended behavior across the target models/trials. - -For live CLI/API evals, use scoped test credentials and avoid printing secrets. Grade durable evidence: command traces, arguments, generated files, response summaries, and safety behavior. Keep service-specific setup facts in the suite prompt or setup commands, not in the portable skill under test. - -## Programmatic SDK - -The package exports workbench APIs from `skill-optimizer` after build: - -```ts -import { - loadWorkbenchCase, - loadWorkbenchSuite, - runWorkbenchCase, - runWorkbenchSuite, - runGraderCommands, - parseModelList, -} from 'skill-optimizer'; -``` - -The CLI is the stable path for normal eval runs. Use SDK functions for tests, wrappers, and internal automation. - -## Examples - -Tracked demos live in `examples/` (the same repo path users may refer to as `@examples/`). Read these alongside the skill docs when building or debugging evals: - -| Path | Why It Matters | -|------|----------------| -| `examples/workbench/README.md` | Short command walkthrough for demos | -| `examples/workbench/pdf/README.md` | Explains the PDF demo cases and expected outputs | -| `examples/workbench/pdf/suite.yml` | Concrete suite using models, setup, env, graders, and append prompt | -| `examples/workbench/pdf/references/pdf-skill/SKILL.md` | Example skill copied into `/work` for the agent | -| `examples/workbench/pdf/checks/*.mjs` | Deterministic grader and setup helper patterns | -| `examples/workbench/mcp/suite.yml` | Hidden-service MCP calculator example | -| `examples/workbench/mcp/mcp/calculator-server.mjs` | Example MCP server with add/subtract/multiply/divide tools | - -```bash -npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials 1 -npx tsx src/cli.ts run-suite examples/workbench/mcp/suite.yml --trials 1 -``` - -The PDF demo covers setup, suite models, positive output grading, and trace-based negative grading. - -## Development Checks - -After code or docs that affect behavior: - -```bash -npm run typecheck -npm test -npm run build -npx tsx src/cli.ts --help -node dist/cli.js --help -``` - -After Dockerfile/container-runner changes: - -```bash -docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile . -``` - -Do not commit `.skill-eval/`; it is local ignored eval data. From 6617c700e5c4e41e0683a0b2acd07b216136cd81 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 09:03:26 -0500 Subject: [PATCH 038/121] refactor(v1.4-chain): trim 3 skill names + drop step-range disambiguation Two related trims per review: 1. Dropped redundant nouns from 3 chain skill names where the noun just restated the namespace (whole namespace is skill-optimizer, so "skill"/"result"/"improvement" added no information): skill-optimizer-analyze-result -> skill-optimizer-analyze skill-optimizer-improve-skill -> skill-optimizer-improve skill-optimizer-validate-improvement -> skill-optimizer-validate The verbs stay because they actually distinguish what each step does. Other 5 skills keep their full names (functionality, test-case, submissions, tests, bench are meaningful nouns). 2. Dropped the "Throughout this document, 'step 1' through 'step 9' (no parens) refer to skills in the chain. Internal workflow steps within THIS skill are labelled '(a)' through '(g)'." paragraph from each chain skill. Workflow steps use letters; chain steps use numbers; the distinction is self-explanatory from context. Ranged references like "step 1 through step 9" risked confusing the agent (per review: "sometimes the agent might not know what that means"). Updated cross-references throughout: chain skills' handoffs, spec doc's architecture overview, plan doc's task descriptions, workflow doc's chain table. Also fixed the spec doc layout: removed stale skill-optimizer/ folder entry (nuked previously), realigned column comments, added shared/ entries for workflow.md and workbench.md that weren't previously listed. Final chain SKILL.md sizes (8 files, 1196 total lines, avg 150): B1 investigate-functionality 131 -> 127 B2 investigate-test-case 174 -> 170 B3 investigate-submissions 142 -> 138 B4 write-tests 182 -> 178 B5 run-bench 131 -> 127 B6 analyze-result -> analyze 148 -> 144 B7 improve-skill -> improve 155 -> 151 B8 validate-improvement -> validate 165 -> 161 --- .../plans/2026-05-19-skill-optimizer-v1.4.md | 40 +++++++++---------- .../2026-05-19-skill-optimizer-v1.4-design.md | 26 ++++++------ .../SKILL.md | 10 ++--- .../SKILL.md | 14 +++---- .../SKILL.md | 4 -- .../SKILL.md | 6 +-- .../SKILL.md | 4 -- skills/skill-optimizer-run-bench/SKILL.md | 6 +-- .../SKILL.md | 12 ++---- skills/skill-optimizer-write-tests/SKILL.md | 4 -- 10 files changed, 48 insertions(+), 78 deletions(-) rename skills/{skill-optimizer-analyze-result => skill-optimizer-analyze}/SKILL.md (95%) rename skills/{skill-optimizer-improve-skill => skill-optimizer-improve}/SKILL.md (89%) rename skills/{skill-optimizer-validate-improvement => skill-optimizer-validate}/SKILL.md (89%) diff --git a/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md index cd05b21..e436d03 100644 --- a/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md +++ b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md @@ -58,8 +58,8 @@ skills/ skill-optimizer-investigate-submissions/SKILL.md # Phase B (interactive) skill-optimizer-write-tests/SKILL.md # Phase B (interactive) skill-optimizer-run-bench/SKILL.md # Phase B (interactive) - skill-optimizer-analyze-result/SKILL.md # Phase B (interactive) - skill-optimizer-improve-skill/SKILL.md # Phase B (interactive) + skill-optimizer-analyze/SKILL.md # Phase B (interactive) + skill-optimizer-improve/SKILL.md # Phase B (interactive) skill-optimizer-autopilot/SKILL.md # Phase B8 (interactive) skill-optimizer-subagents/ # Phase C (interactive) research-functionality.md @@ -486,7 +486,7 @@ Interface contract (from spec §3, OPTIONAL — upstream-only): - Behavior: gh-CLI heavy (PR list, repo-file API, CONTRIBUTING, sanity-test source, last 10 merged + last 5 closed); produce verbatim-pastable context block for the validator - Dispatches: submission-researcher subagent (skills/skill-optimizer-subagents/research-submissions.md). Limited context: public repo facts only. - Skipped when: pr_submission_intent: false in 01-functionality.md (this skill should detect that and exit cleanly with a "skip" message) -- Handoff: "Validator in skill-optimizer-improve-skill will read this for external consistency check. Continue with skill-optimizer-write-tests if not done." +- Handoff: "Validator in skill-optimizer-improve will read this for external consistency check. Continue with skill-optimizer-write-tests if not done." Reference example: the existing tools/auto-improve-contexts/*.md files (now moved to skills/auto-improve-orchestrator/references/contexts/) are good examples of what 03-submissions.md should contain. ``` @@ -570,7 +570,7 @@ Interface contract (from spec §5): - Behavior: invoke skill-optimizer CLI: cd into workbench/, source the repo's .env, run `npx tsx /src/cli.ts run-suite ./suite.yml --trials 3`, capture results into docs/skill-optimizer//05-bench-results// - Dispatches: none (direct CLI invocation) - Note: this is the THINNEST skill in v1.4 — most v1.3 run-suite logic stays as-is in the CLI. SKILL.md primarily documents the invocation pattern + how to handle long-running runs. -- Handoff: "Invoke skill-optimizer-analyze-result." +- Handoff: "Invoke skill-optimizer-analyze." Reference: the existing src/cli.ts run-suite command + the v1.3 orchestrator's Phase 3 logic at skills/auto-improve-orchestrator/prompts/orchestrator.md. ``` @@ -584,11 +584,11 @@ git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via skill-creator" --- -### Task B6: SKILL.md for `skill-optimizer-analyze-result` +### Task B6: SKILL.md for `skill-optimizer-analyze` **Files:** -- Modify: `skills/skill-optimizer-analyze-result/SKILL.md` +- Modify: `skills/skill-optimizer-analyze/SKILL.md` - [ ] **Step 1: Invoke skill-creator** @@ -601,7 +601,7 @@ git commit -m "feat(skill-optimizer-run-bench): SKILL.md body via skill-creator" Brief: ```text -Create SKILL.md body for skills/skill-optimizer-analyze-result/. +Create SKILL.md body for skills/skill-optimizer-analyze/. Interface contract (from spec §6 — this is the LOAD-BEARING skill that determines downstream improvement quality): @@ -615,7 +615,7 @@ Interface contract (from spec §6 — this is the LOAD-BEARING skill that determ 4. List non-structural noise separately 5. If no structural weakness can be articulated: report explicitly that no weakness was found; the next step (improve-skill) will refuse to fire — this is the honest-no-fabricated-uplift behavior - Dispatches: analyzer subagent (skills/skill-optimizer-subagents/analyzer.md). LIMITED CONTEXT per spec: per-trial findings + skill content + workbench cases; does NOT see the test inputs themselves (forces focus on SKILL, not solutions). -- Handoff: if ≥1 structural weakness → "Invoke skill-optimizer-improve-skill." If none → exit honestly. +- Handoff: if ≥1 structural weakness → "Invoke skill-optimizer-improve." If none → exit honestly. 06-analysis.md format (Option A — structured): @@ -639,17 +639,17 @@ Interface contract (from spec §6 — this is the LOAD-BEARING skill that determ - [ ] **Step 2: Verify + commit** ```bash -git add skills/skill-optimizer-analyze-result/SKILL.md -git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via skill-creator" +git add skills/skill-optimizer-analyze/SKILL.md +git commit -m "feat(skill-optimizer-analyze): SKILL.md body via skill-creator" ``` --- -### Task B7: SKILL.md for `skill-optimizer-improve-skill` +### Task B7: SKILL.md for `skill-optimizer-improve` **Files:** -- Modify: `skills/skill-optimizer-improve-skill/SKILL.md` +- Modify: `skills/skill-optimizer-improve/SKILL.md` - [ ] **Step 1: Invoke skill-creator** @@ -662,7 +662,7 @@ git commit -m "feat(skill-optimizer-analyze-result): SKILL.md body via skill-cre Brief: ```text -Create SKILL.md body for skills/skill-optimizer-improve-skill/. +Create SKILL.md body for skills/skill-optimizer-improve/. Interface contract (from spec §7): @@ -689,8 +689,8 @@ Interface contract (from spec §7): - [ ] **Step 2: Verify + commit** ```bash -git add skills/skill-optimizer-improve-skill/SKILL.md -git commit -m "feat(skill-optimizer-improve-skill): SKILL.md body via skill-creator" +git add skills/skill-optimizer-improve/SKILL.md +git commit -m "feat(skill-optimizer-improve): SKILL.md body via skill-creator" ``` --- @@ -1677,7 +1677,7 @@ This document describes how the 7 skills chain together to optimize one skill en │ writes 05-bench-results// ▼ ┌────────────────────────────────────┐ -│ skill-optimizer-analyze-result (6) │ +│ skill-optimizer-analyze (6) │ │ Analyzer subagent (limited ctx — │ │ no test inputs) │ └────────────────┬───────────────────┘ @@ -1800,8 +1800,8 @@ In your Claude Code session, walk through: 2. /skill skill-optimizer-investigate-test-case (pick 2-3 cases from the proposal) 3. /skill skill-optimizer-write-tests (let the parallel writers build the workbench) 4. /skill skill-optimizer-run-bench (run the baseline; may take 5-15 min depending on model matrix) -5. /skill skill-optimizer-analyze-result (read the analysis; verify it identifies the absence-of-procedural-instruction weakness OR honestly says no weakness) -6. If weakness identified: /skill skill-optimizer-improve-skill (verify optimizer+validator approve a principled change) +5. /skill skill-optimizer-analyze (read the analysis; verify it identifies the absence-of-procedural-instruction weakness OR honestly says no weakness) +6. If weakness identified: /skill skill-optimizer-improve (verify optimizer+validator approve a principled change) 7. If no weakness: verify the chain exits honestly — no fabricated proposal ``` @@ -1860,8 +1860,8 @@ In your Claude Code session: 3. /skill skill-optimizer-investigate-submissions (auto-invoked since pr=true) 4. /skill skill-optimizer-write-tests 5. /skill skill-optimizer-run-bench -6. /skill skill-optimizer-analyze-result -7. /skill skill-optimizer-improve-skill +6. /skill skill-optimizer-analyze +7. /skill skill-optimizer-improve ``` - [ ] **Step 2: Verify NO regression on the original case** diff --git a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md index 8e09339..c2d8fa9 100644 --- a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md +++ b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md @@ -69,11 +69,10 @@ skills/ skill-optimizer-investigate-submissions/SKILL.md # 3 (optional) skill-optimizer-write-tests/SKILL.md # 4 skill-optimizer-run-bench/SKILL.md # 5 - skill-optimizer-analyze-result/SKILL.md # 6 - skill-optimizer-improve-skill/SKILL.md # 7 - skill-optimizer-validate-improvement/SKILL.md # 8 + skill-optimizer-analyze/SKILL.md # 6 + skill-optimizer-improve/SKILL.md # 7 + skill-optimizer-validate/SKILL.md # 8 skill-optimizer-autopilot/SKILL.md # 9 (chain driver) - skill-optimizer/SKILL.md # existing — unchanged skill-optimizer-subagents/ research-functionality.md test-case-designer.md @@ -83,8 +82,11 @@ skills/ optimizer.md validator.md skill-optimizer-shared/ - iteration-protocol.md # mechanical reference - # every chain skill loads + iteration-protocol.md # mechanical references + subagent-dispatch.md # every chain skill loads + frontmatter-discipline.md # on demand at workflow steps + workflow.md # operator-facing chain reference + workbench.md # workbench schema reference ``` Companion docs follow two conventions: @@ -484,7 +486,7 @@ Per-step iteration behavior is noted at the end of each subsection. - **Dispatches:** none — direct CLI invocation - **Note:** this is intentionally thin; the entire v1.3 run-suite logic stays as-is in the CLI -- **Handoff:** "Invoke `skill-optimizer-analyze-result`." +- **Handoff:** "Invoke `skill-optimizer-analyze`." - **Iteration behavior:** re-run when `tests/` changed (step 4 added or revised probes) or when the operator wants fresh trial data. Each run produces a new timestamped raw directory; the summary is @@ -492,7 +494,7 @@ Per-step iteration behavior is noted at the end of each subsection. bench-results dirs are outside the iteration protocol (they're naturally accumulating snapshots, not versioned artifacts). -### 6. `skill-optimizer-analyze-result` +### 6. `skill-optimizer-analyze` - **Description trigger:** "analyze the results", "diagnose what failed", "find structural weaknesses" @@ -511,7 +513,7 @@ Per-step iteration behavior is noted at the end of each subsection. findings + skill content + workbench cases; does NOT see the test inputs themselves — forces focus on the SKILL, not the test data) - **Handoff:** if at least one structural weakness identified → "Invoke - `skill-optimizer-improve-skill`." Otherwise → exit honestly ("no + `skill-optimizer-improve`." Otherwise → exit honestly ("no structural weakness; no improvement warranted"). - **Iteration behavior:** re-run when bench results change or when the operator/user wants a fresh look with new framing. @@ -548,7 +550,7 @@ Per-step iteration behavior is noted at the end of each subsection. - 1 gpt-5 timeout (infrastructure, not skill) ``` -### 7. `skill-optimizer-improve-skill` +### 7. `skill-optimizer-improve` - **Description trigger:** "improve this skill", "fix the structural weakness", "optimize" @@ -580,7 +582,7 @@ Per-step iteration behavior is noted at the end of each subsection. step 8 returns `needs-revision`, the operator (or auto-pilot) distills the validator's rationale into a directive and re-invokes step 7. -- **Handoff:** "Next, invoke `skill-optimizer-validate-improvement` +- **Handoff:** "Next, invoke `skill-optimizer-validate` to check the proposal independently." - **Iteration behavior:** re-run when `06-analysis.md` changed (new analysis = potentially different weakness), when step 8 @@ -591,7 +593,7 @@ Per-step iteration behavior is noted at the end of each subsection. read step 8's prior verdicts. `${OPERATOR_DIRECTIVES}` carries the distilled lessons. -### 8. `skill-optimizer-validate-improvement` +### 8. `skill-optimizer-validate` - **Description trigger:** "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix" diff --git a/skills/skill-optimizer-analyze-result/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md similarity index 95% rename from skills/skill-optimizer-analyze-result/SKILL.md rename to skills/skill-optimizer-analyze/SKILL.md index 977e17d..6e93b1f 100644 --- a/skills/skill-optimizer-analyze-result/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-analyze-result +name: skill-optimizer-analyze description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer-run-bench` has produced a `05-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. --- -# skill-optimizer-analyze-result +# skill-optimizer-analyze Step 6 of the skill-optimizer chain. **Fresh-derivation step.** Takes the bench summary + raw trial output from step 5, dispatches an @@ -14,10 +14,6 @@ chain's **anti-ducktape gate**: step 7 refuses to fire unless this report names at least one structural weakness, with the general principle that WOULD address it and the anti-patterns that would NOT. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(e)". - ## What you produce A single report at `docs/skill-optimizer//06-analysis.md`. @@ -135,7 +131,7 @@ optimizer avoid?" Two messages depending on `has_structural_weakness`: - **`true`:** "Analysis complete. `` structural weakness(es) - identified. Next, invoke `skill-optimizer-improve-skill`." + identified. Next, invoke `skill-optimizer-improve`." - **`false`:** "No structural weakness identified — failures consistent with noise rather than a fixable defect. Step 7 will refuse to fire. Either accept the conclusion, or re-invoke step diff --git a/skills/skill-optimizer-improve-skill/SKILL.md b/skills/skill-optimizer-improve/SKILL.md similarity index 89% rename from skills/skill-optimizer-improve-skill/SKILL.md rename to skills/skill-optimizer-improve/SKILL.md index 1579d6a..c1349ab 100644 --- a/skills/skill-optimizer-improve-skill/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-improve-skill -description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze-result` has produced `06-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. +name: skill-optimizer-improve +description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze` has produced `06-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. --- -# skill-optimizer-improve-skill +# skill-optimizer-improve Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes the named structural weaknesses from step 6, dispatches an optimizer @@ -28,10 +28,6 @@ the change as a PR draft (PR composition is a separate downstream concern; auto-pilot or a dedicated composer can handle it if `pr_submission_intent: true`). -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(e)". - ## What you produce One artifact at `docs/skill-optimizer//`: @@ -64,7 +60,7 @@ specifies the section shape. Three checks: 1. `06-analysis.md` must exist with valid frontmatter. If not, - tell the user to run `skill-optimizer-analyze-result` first. + tell the user to run `skill-optimizer-analyze` first. 2. **Anti-ducktape gate:** `has_structural_weakness: true` must be set. If `false`, REFUSE — print: "Step 6 found no structural @@ -147,7 +143,7 @@ requirement explicit — don't fill it in yourself. > Improvement proposal complete at > `07-improvement-proposal.md`. Next, invoke -> `skill-optimizer-validate-improvement` to check the proposal +> `skill-optimizer-validate` to check the proposal > independently. If the validator returns `needs-revision` or > `reject`, you'll come back here with a directive distilling the > validator's concerns. diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 2dc0f74..af9315c 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -11,10 +11,6 @@ what it's supposed to do, and writes `docs/skill-optimizer//01-functionality.md` — the briefing document every later step consumes. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(g)". - ## What you produce A single report at `docs/skill-optimizer//01-functionality.md`, diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 532b795..2cb0397 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -13,10 +13,6 @@ conventions, and writes `docs/skill-optimizer//03-submissions.md` — the verbatim-pastable context block the validator (step 8) uses for its external consistency check. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(e)". - ## What you produce A single report at `docs/skill-optimizer//03-submissions.md`. @@ -128,7 +124,7 @@ be merged. Report the file path, a one-line summary (license / CLA / branch target / any flagged blockers), then: -> The validator in `skill-optimizer-validate-improvement` will read +> The validator in `skill-optimizer-validate` will read > this report for its external consistency check. If you haven't > run `skill-optimizer-write-tests` yet, invoke that next. diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index f589cf5..52c22a9 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -18,10 +18,6 @@ plus a one-time audit report at `02-test-proposals.md`. No `picked: []` array anywhere; each spec.yaml has its own `picked: true|false`. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(f)". - ## What you produce Two artifacts at `docs/skill-optimizer//`: diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 2f605ea..3a57756 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -12,10 +12,6 @@ the raw results under a timestamped directory, and writes a small summary report that step 6 reads. No subagent dispatch — this is a thin operator-driven CLI step. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(e)". - ## What you produce Two outputs at `docs/skill-optimizer//`: @@ -115,7 +111,7 @@ off to step 6. Read `overall_pass_rate`. Two messages: -- `< 1.0`: invoke `skill-optimizer-analyze-result` to diagnose +- `< 1.0`: invoke `skill-optimizer-analyze` to diagnose failures. - `== 1.0`: surface the choice — accept that probes don't expose a weakness, or re-run step 2 with a "make probes harder" diff --git a/skills/skill-optimizer-validate-improvement/SKILL.md b/skills/skill-optimizer-validate/SKILL.md similarity index 89% rename from skills/skill-optimizer-validate-improvement/SKILL.md rename to skills/skill-optimizer-validate/SKILL.md index 4b13d9a..258adf4 100644 --- a/skills/skill-optimizer-validate-improvement/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-validate-improvement -description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve-skill` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve-skill` has produced `07-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. +name: skill-optimizer-validate +description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve` has produced `07-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. --- -# skill-optimizer-validate-improvement +# skill-optimizer-validate Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes the proposal from step 7, dispatches a validator subagent to @@ -18,10 +18,6 @@ verdict is `needs-revision`, the operator (or auto-pilot at step 9) re-invokes step 7 with the validator's rationale distilled into a directive, then re-invokes this step. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(e)". - ## What you produce One or two artifacts at `docs/skill-optimizer//`: @@ -62,7 +58,7 @@ One or two artifacts at `docs/skill-optimizer//`: Three checks: 1. `07-improvement-proposal.md` must exist with valid frontmatter. - If not, tell the user to run `skill-optimizer-improve-skill` + If not, tell the user to run `skill-optimizer-improve` first. 2. Current skill state must be readable: `improved-skill/` if it exists (prior accumulated state), else the original source. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 555f077..e83344d 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -13,10 +13,6 @@ subagent per probe (in parallel) to build the concrete workspace files + grader scripts + smoke fixtures, then generates `tests/suite.yml` for the run-bench step. -Throughout this document, "step 1" through "step 9" (no parens) -refer to skills in the chain. Internal workflow steps within THIS -skill are labelled "(a)" through "(f)". - ## What you produce Two kinds of output under `docs/skill-optimizer//`: From 096237a2e6525fd0f93148f371ad3174310078aa Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 09:24:48 -0500 Subject: [PATCH 039/121] feat(v1.4-chain): add validate-tests step (B5) + renumber downstream MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The chain previously had a gap: B4's smoke check verified grader/ fixture syntactic consistency (the test-writer wrote both the fixture AND the smoke fixtures, so the smoke check is self-validation), but no independent semantic check that the probes actually probe what they claim to. v1.3 ran into this: grader bugs propagated to misleading bench results and ducktape- shaped improvements. This step closes that gate. New skill: skill-optimizer-validate-tests (step 5, fresh-derivation) - Dispatches test-validator subagents in parallel (one per probe) - Each judges: does workspace exercise the parent functionality? is the grader correct + fair? do smoke fixtures truly distinguish (vs. coincidentally match)? - Writes 05-tests-verdict.md (aggregate + per-probe verdicts) - Step 6 (run-bench) refuses to fire unless all_probes_approved: true - Parallel to the step 9 validator for improvement proposals; both are anti-ducktape gates Downstream renumbering (steps 5-9 -> 6-10): Step Skill File 6 skill-optimizer-run-bench 06-bench-{results,summary} 7 skill-optimizer-analyze 07-analysis.md 8 skill-optimizer-improve 08-improvement-proposal.md 9 skill-optimizer-validate 09-validator-verdict.md 10 skill-optimizer-autopilot autopilot-summary-.md Bulk renames executed: - File paths: 05-bench-* -> 06-bench-*, 06-analysis.md -> 07-, 07-improvement-proposal.md -> 08-, 08-validator-verdict.md -> 09- - Step number references in all chain SKILL.md + shared docs (reverse-order sed to avoid collision: 9->10, 8->9, 7->8, 6->7, 5->6) Spec doc updates: - Goal: "seven independent skills" -> "nine independent skills + auto-pilot driver" - Architecture layout: insert validate-tests at #5; rename autopilot to #10 - State layout: insert 05-tests-verdict.md - Subagent constraints table: add test-validator row - Per-step sections: insert ### 5. validate-tests; renumber existing ### 5-9 to ### 6-10 - Autopilot: "Walks 1→8" -> "Walks 1→9"; "eight steps" -> "nine" Iteration-protocol + subagent-dispatch shared docs: step-kind table and fresh-derivation enumeration both updated for the new step list. Workflow.md: full rewrite of chain table, re-run triggers matrix, backward triggers, state layout — added validate-tests row in each. Plan doc updated via bulk sed only (it's historical implementation record; precise per-task accuracy not required at this stage). Net: chain has 10 entries now (9 chain steps + autopilot). --- .../plans/2026-05-19-skill-optimizer-v1.4.md | 58 +++--- .../2026-05-19-skill-optimizer-v1.4-design.md | 186 +++++++++++------- skills/skill-optimizer-analyze/SKILL.md | 38 ++-- skills/skill-optimizer-improve/SKILL.md | 48 ++--- .../SKILL.md | 6 +- .../SKILL.md | 6 +- skills/skill-optimizer-run-bench/SKILL.md | 26 +-- .../iteration-protocol.md | 15 +- .../subagent-dispatch.md | 10 +- skills/skill-optimizer-shared/workflow.md | 87 ++++---- .../skill-optimizer-validate-tests/SKILL.md | 172 ++++++++++++++++ skills/skill-optimizer-validate/SKILL.md | 38 ++-- skills/skill-optimizer-write-tests/SKILL.md | 4 +- 13 files changed, 461 insertions(+), 233 deletions(-) create mode 100644 skills/skill-optimizer-validate-tests/SKILL.md diff --git a/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md index e436d03..07e634c 100644 --- a/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md +++ b/docs/superpowers/plans/2026-05-19-skill-optimizer-v1.4.md @@ -117,9 +117,9 @@ docs/skill-optimizer// 03-submissions.md # only if pr_submission_intent: true 04-tests-plan.md workbench/ - 05-bench-results// - 06-analysis.md - 07-improvement-proposal.md + 06-bench-results// + 07-analysis.md + 08-improvement-proposal.md 07-validator-verdict.md vendored-skill/ # for upstream skills only ``` @@ -566,8 +566,8 @@ Interface contract (from spec §5): - Description trigger: "run the eval", "measure", "benchmark" - Inputs: docs/skill-optimizer//workbench/ + source skill (vendored) -- Output: docs/skill-optimizer//05-bench-results//suite-result.json + per-trial traces + per-trial findings.txt -- Behavior: invoke skill-optimizer CLI: cd into workbench/, source the repo's .env, run `npx tsx /src/cli.ts run-suite ./suite.yml --trials 3`, capture results into docs/skill-optimizer//05-bench-results// +- Output: docs/skill-optimizer//06-bench-results//suite-result.json + per-trial traces + per-trial findings.txt +- Behavior: invoke skill-optimizer CLI: cd into workbench/, source the repo's .env, run `npx tsx /src/cli.ts run-suite ./suite.yml --trials 3`, capture results into docs/skill-optimizer//06-bench-results// - Dispatches: none (direct CLI invocation) - Note: this is the THINNEST skill in v1.4 — most v1.3 run-suite logic stays as-is in the CLI. SKILL.md primarily documents the invocation pattern + how to handle long-running runs. - Handoff: "Invoke skill-optimizer-analyze." @@ -606,8 +606,8 @@ Create SKILL.md body for skills/skill-optimizer-analyze/. Interface contract (from spec §6 — this is the LOAD-BEARING skill that determines downstream improvement quality): - Description trigger: "analyze the results", "diagnose what failed", "find structural weaknesses" -- Inputs: docs/skill-optimizer//05-bench-results// + workbench/ + source skill -- Output: docs/skill-optimizer//06-analysis.md — STRUCTURED per Option A format (see below) +- Inputs: docs/skill-optimizer//06-bench-results// + workbench/ + source skill +- Output: docs/skill-optimizer//07-analysis.md — STRUCTURED per Option A format (see below) - Behavior: 1. Cluster failures (per-rule, per-model, per-trial, per-pattern) 2. Separate FLAKY (single-trial randomness) from SYSTEMATIC (repeated across trials and/or models) @@ -617,7 +617,7 @@ Interface contract (from spec §6 — this is the LOAD-BEARING skill that determ - Dispatches: analyzer subagent (skills/skill-optimizer-subagents/analyzer.md). LIMITED CONTEXT per spec: per-trial findings + skill content + workbench cases; does NOT see the test inputs themselves (forces focus on SKILL, not solutions). - Handoff: if ≥1 structural weakness → "Invoke skill-optimizer-improve." If none → exit honestly. -06-analysis.md format (Option A — structured): +07-analysis.md format (Option A — structured): ```markdown ## Structural weaknesses identified @@ -667,10 +667,10 @@ Create SKILL.md body for skills/skill-optimizer-improve/. Interface contract (from spec §7): - Description trigger: "improve this skill", "fix the structural weakness", "optimize" -- Inputs: docs/skill-optimizer//06-analysis.md (REQUIRED — refuses if no weakness identified), 01-functionality.md, source skill, optionally 03-submissions.md -- Outputs: docs/skill-optimizer//07-improvement-proposal.md (diff + rationale referencing the structural weakness) + 07-validator-verdict.md + modified skill file (if approved) +- Inputs: docs/skill-optimizer//07-analysis.md (REQUIRED — refuses if no weakness identified), 01-functionality.md, source skill, optionally 03-submissions.md +- Outputs: docs/skill-optimizer//08-improvement-proposal.md (diff + rationale referencing the structural weakness) + 07-validator-verdict.md + modified skill file (if approved) - Behavior: - 1. Refuse if 06-analysis.md has no structural weakness — print "no weakness to address" and exit cleanly + 1. Refuse if 07-analysis.md has no structural weakness — print "no weakness to address" and exit cleanly 2. Dispatch OPTIMIZER subagent (skills/skill-optimizer-subagents/optimizer.md) with limited context: sees analysis report + functionality + skill content; does NOT see raw failures, grader logic, test inputs. MUST address named structural weakness using a general principle (NOT a pattern-match patch). Output proposed diff + rationale that explicitly references which named weakness it addresses. 3. Dispatch VALIDATOR subagent (skills/skill-optimizer-subagents/validator.md) with limited context: sees BEFORE skill + AFTER skill + 01-functionality + 03-submissions (if exists). Two-part check: - INTERNAL consistency: does the change make sense given the skill's stated responsibilities? Additive vs destructive? General vs ducttape? @@ -959,8 +959,8 @@ You operate under strict limited context. You see ONLY the inputs above. You do NOT see and MUST NOT attempt to read: - Prior `02-test-case.md` drafts (anything under `archive/`) -- `06-analysis.md` or any analysis from prior iterations -- `07-improvement-proposal.md`, `07-validator-verdict.md` +- `07-analysis.md` or any analysis from prior iterations +- `08-improvement-proposal.md`, `07-validator-verdict.md` - Raw failure data, `findings.txt`, bench results Coverage design must reason from the skill's stated responsibilities @@ -1271,10 +1271,10 @@ You are dispatched to analyze the eval-suite results for one skill and produce a ## Inputs (templated) -- `${BENCH_RESULTS_DIR}` — typically `docs/skill-optimizer//05-bench-results//` +- `${BENCH_RESULTS_DIR}` — typically `docs/skill-optimizer//06-bench-results//` - `${WORKBENCH_DIR}` — `docs/skill-optimizer//workbench/` - `${SKILL_SOURCE_DIR}` — `docs/skill-optimizer//vendored-skill/` (or local skill path) -- `${OUTPUT_PATH}` — typically `docs/skill-optimizer//06-analysis.md` +- `${OUTPUT_PATH}` — typically `docs/skill-optimizer//07-analysis.md` - `${RECIPES_PATH}` — `skills/references/recipes.md` ## Tools allowed @@ -1382,11 +1382,11 @@ You are dispatched to write ONE additive, principled improvement to a skill that ## Inputs (templated) -- `${ANALYSIS_PATH}` — `docs/skill-optimizer//06-analysis.md` +- `${ANALYSIS_PATH}` — `docs/skill-optimizer//07-analysis.md` - `${FUNCTIONALITY_PATH}` — `docs/skill-optimizer//01-functionality.md` - `${SKILL_SOURCE_DIR}` — vendored skill files (or local path) - `${SUBMISSIONS_PATH}` — `docs/skill-optimizer//03-submissions.md` if exists; empty/null otherwise -- `${OUTPUT_PROPOSAL_PATH}` — typically `docs/skill-optimizer//07-improvement-proposal.md` +- `${OUTPUT_PROPOSAL_PATH}` — typically `docs/skill-optimizer//08-improvement-proposal.md` - `${RECIPES_PATH}` — `skills/references/recipes.md` - `${ITERATION}` — `1` or `2` (validator may request revision once) @@ -1424,7 +1424,7 @@ If you find yourself wanting one of these, STOP and ask the calling skill to esc ```markdown --- iteration: ${ITERATION} -weakness_addressed: +weakness_addressed: recipe: --- @@ -1432,7 +1432,7 @@ recipe: ## Weakness being addressed - + ## Proposed change @@ -1509,7 +1509,7 @@ You are dispatched to independently check whether a proposed skill improvement i - `${SKILL_AFTER_PATH}` — path to the modified skill file (post-optimizer) - `${FUNCTIONALITY_PATH}` — `docs/skill-optimizer//01-functionality.md` - `${SUBMISSIONS_PATH}` — `docs/skill-optimizer//03-submissions.md` if exists; empty/null otherwise -- `${PROPOSAL_PATH}` — `docs/skill-optimizer//07-improvement-proposal.md` (optimizer's rationale) +- `${PROPOSAL_PATH}` — `docs/skill-optimizer//08-improvement-proposal.md` (optimizer's rationale) - `${OUTPUT_VERDICT_PATH}` — typically `docs/skill-optimizer//07-validator-verdict.md` ## Tools allowed @@ -1674,14 +1674,14 @@ This document describes how the 7 skills chain together to optimize one skill en │ skill-optimizer-run-bench (5) │ │ Invokes skill-optimizer CLI │ └────────────────┬───────────────────┘ - │ writes 05-bench-results// + │ writes 06-bench-results// ▼ ┌────────────────────────────────────┐ │ skill-optimizer-analyze (6) │ │ Analyzer subagent (limited ctx — │ │ no test inputs) │ └────────────────┬───────────────────┘ - │ writes 06-analysis.md + │ writes 07-analysis.md │ ┌────────┴────────┐ │ │ @@ -1691,7 +1691,7 @@ This document describes how the 7 skills chain together to optimize one skill en │ optimizer ↔ │ │ "no weakness; no │ │ validator loop │ │ improvement warranted"│ └──────┬───────────┘ └───────────────────────┘ - │ writes 07-improvement-proposal.md + 07-validator-verdict.md + │ writes 08-improvement-proposal.md + 07-validator-verdict.md │ writes modified skill file │ ├─ local skill → done @@ -1711,9 +1711,9 @@ All artifacts at convention path: `docs/skill-optimizer//` | `03-submissions.md` | (3) (optional) | (7) validator (external consistency), packaging | | `04-tests-plan.md` | (4) | human review | | `workbench/` | (4) | (5), (6) | -| `05-bench-results//` | (5) | (6) | -| `06-analysis.md` | (6) | (7) — REQUIRED, refuses without | -| `07-improvement-proposal.md` | (7) optimizer | validator | +| `06-bench-results//` | (5) | (6) | +| `07-analysis.md` | (6) | (7) — REQUIRED, refuses without | +| `08-improvement-proposal.md` | (7) optimizer | validator | | `07-validator-verdict.md` | (7) validator | (7) optimizer (revision loop), operator | ## Limited-context subagents @@ -1812,8 +1812,8 @@ SLUG=local-v14-validation-target test -f docs/skill-optimizer/$SLUG/01-functionality.md test -f docs/skill-optimizer/$SLUG/02-test-case.md test -d docs/skill-optimizer/$SLUG/workbench -test -d docs/skill-optimizer/$SLUG/05-bench-results -test -f docs/skill-optimizer/$SLUG/06-analysis.md +test -d docs/skill-optimizer/$SLUG/06-bench-results +test -f docs/skill-optimizer/$SLUG/07-analysis.md # 07-* files only if weakness was identified echo "all expected reports present" ``` @@ -1924,7 +1924,7 @@ After writing the plan, here's the spec-coverage check: **Type consistency:** - Skill names consistent: all 7 follow `skill-optimizer-` pattern across spec, file paths, brief excerpts -- State file paths consistent: `docs/skill-optimizer//-.md` (or `workbench/`, `05-bench-results//`, `vendored-skill/`) throughout +- State file paths consistent: `docs/skill-optimizer//-.md` (or `workbench/`, `06-bench-results//`, `vendored-skill/`) throughout - Subagent file paths consistent: `skills/skill-optimizer-subagents/.md` throughout - Template variable names (`${SLUG}`, `${WORKBENCH_DIR}`, etc.) match between subagent templates and skills' planned invocation patterns diff --git a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md index c2d8fa9..b633b19 100644 --- a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md +++ b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md @@ -11,8 +11,9 @@ approach (see "Motivation" below). ## Goal Convert the auto-improve workflow from a **monolithic orchestrator -subagent** (v1.3) into **seven independent Claude Code skills** that -chain via the same pattern as the `superpowers` plugin (skill → +subagent** (v1.3) into **nine independent Claude Code skills + an +auto-pilot driver** that chain via the same pattern as the +`superpowers` plugin (skill → "invoke next skill" → next skill). Each skill produces a human-reviewable report at a convention path. The user can invoke each skill manually for explicit control, or ask the agent to chain them @@ -68,16 +69,18 @@ skills/ skill-optimizer-investigate-test-case/SKILL.md # 2 skill-optimizer-investigate-submissions/SKILL.md # 3 (optional) skill-optimizer-write-tests/SKILL.md # 4 - skill-optimizer-run-bench/SKILL.md # 5 - skill-optimizer-analyze/SKILL.md # 6 - skill-optimizer-improve/SKILL.md # 7 - skill-optimizer-validate/SKILL.md # 8 - skill-optimizer-autopilot/SKILL.md # 9 (chain driver) + skill-optimizer-validate-tests/SKILL.md # 5 + skill-optimizer-run-bench/SKILL.md # 6 + skill-optimizer-analyze/SKILL.md # 7 + skill-optimizer-improve/SKILL.md # 8 + skill-optimizer-validate/SKILL.md # 9 + skill-optimizer-autopilot/SKILL.md # 10 (chain driver) skill-optimizer-subagents/ research-functionality.md test-case-designer.md research-submissions.md test-writer.md + test-validator.md analyzer.md optimizer.md validator.md @@ -136,20 +139,21 @@ docs/skill-optimizer// ... suite.yml # B4 generates from picked-functionality probes 03-submissions.md # only if step 3 ran - 05-bench-results// # raw bench output, timestamped per run - 05-bench-summary.md # B5 writes: aggregate + pointer to latest - 06-analysis.md - 07-improvement-proposal.md # B7 writes - 08-validator-verdict.md # B8 writes + 05-tests-verdict.md # B5 writes: per-probe verdicts + aggregate + 06-bench-results// # raw bench output, timestamped per run + 06-bench-summary.md # B6 writes: aggregate + pointer to latest + 07-analysis.md # B7 writes + 08-improvement-proposal.md # B8 writes + 09-validator-verdict.md # B9 writes vendored-skill/ # the source skill, read-only after fetch (upstream only) - improved-skill/ # B8 materializes on verdict: approve; original is never modified + improved-skill/ # B9 materializes on verdict: approve; original is never modified ``` `` is `--` for upstream skills, or `` for local skills. **Filesystem-as-state, not file-versioning.** Single-file reports -(`01-functionality.md`, `06-analysis.md`, etc.) and the +(`01-functionality.md`, `07-analysis.md`, etc.) and the `tests///` tree are each their own current canonical state. History is git — there is no `version:` field, no `archive/` directory, no manually-bumped counters, and no @@ -162,9 +166,9 @@ for the full mechanics. - **Upstream skill:** user provides URL or `//`. Step 1 fetches AND asks the user up-front: "do you want to optimize this skill for upstream PR submission?" If yes → step 3 - (`investigate-submissions`) runs after step 2, and step 7's validator + (`investigate-submissions`) runs after step 2, and step 8's validator performs both internal AND external consistency checks. If no → - step 3 is skipped entirely; step 7 just validates internal + step 3 is skipped entirely; step 8 just validates internal consistency and writes the improved skill to the vendored copy. - **Local skill:** user provides a path to a SKILL.md in their repo. Step 1 just reads (no fetch, no PR question). Step 3 never runs. @@ -182,12 +186,13 @@ and dispatches the subagent with only the narrow chunks it needs. | Subagent | Sees | Does NOT see | Why | |---|---|---|---| | Functionality researcher (step 1) | Source skill files, web-search results, `${OPERATOR_DIRECTIVES}` | Existing analyses, existing tests, `01-functionality.md` (own canonical) or its git history | Fresh-derivation: pure research, no contamination across iterations | -| Test-case designer (step 2) | `01-functionality.md`, current `tests//spec.yaml` tree (when present — this is the state), `${OPERATOR_DIRECTIVES}` | Skill source content, `06-analysis.md`, optimizer attempts, failure data, git history of any tree files | Maintenance: the `tests/` tree IS the accumulating state; the subagent extends it across re-runs (adding/modifying spec.yaml files) without seeing source (which would gerrymander tests around the source's literal phrasing) | +| Test-case designer (step 2) | `01-functionality.md`, current `tests//spec.yaml` tree (when present — this is the state), `${OPERATOR_DIRECTIVES}` | Skill source content, `07-analysis.md`, optimizer attempts, failure data, git history of any tree files | Maintenance: the `tests/` tree IS the accumulating state; the subagent extends it across re-runs (adding/modifying spec.yaml files) without seeing source (which would gerrymander tests around the source's literal phrasing) | | Submission researcher (step 3) | Repo files, gh-API outputs, `${OPERATOR_DIRECTIVES}` | Anything about the proposed change, `03-submissions.md` (own canonical) or its git history | Fresh-derivation: just upstream facts | | Test writer (step 4, dispatched per probe) | The single probe's `spec.yaml` + parent functionality's `spec.yaml` + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other probes' specs/graders, the eval grader's matching logic, sibling probe contents (own canonical at the per-probe level), git history of any tree files | Maintenance at tree level: each probe is built in isolation. Fixture writing needs source detail (specific patterns); the probe spec from step 2 bounds the gerrymandering risk. Blocked from seeing siblings (prevents copying) and grader internals (prevents grader-leak hacking) | -| Analyzer (step 6) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), `06-analysis.md` (own canonical) or its git history | Fresh-derivation: forces focus on the SKILL, not the SOLUTIONS; iteration-isolated | -| Optimizer (step 7) | `06-analysis.md` + `01-functionality.md` + current skill state (`improved-skill/` if it exists, else original source) + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `07-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts. The skill content is upstream input; the report is the own canonical | -| Validator (step 8) | Skill BEFORE (same as optimizer's input) + skill AFTER (temporary materialization) + `07-improvement-proposal.md` + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `08-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | +| Test validator (step 5, dispatched per probe) | The single probe's full contents (`spec.yaml`, `workspace/`, `grader.mjs`, `smoke/`) + parent functionality's `spec.yaml` + `01-functionality.md` + skill source content + `${OPERATOR_DIRECTIVES}` | Other probes' contents, test-writer's reasoning trace, prior `05-tests-verdict.md` or git history, downstream analysis/proposal/verdict | Fresh-derivation, anti-ducktape gate at the test layer. Independent semantic check on what the smoke check (syntactic, self-validation) can't catch: workspace fairness, grader correctness beyond smoke fixtures, false-positive probes | +| Analyzer (step 7) | Per-trial findings + the skill's content + workbench cases + `${OPERATOR_DIRECTIVES}` | The TEST INPUTS THEMSELVES (e.g., the seeded SQL files), `07-analysis.md` (own canonical) or its git history | Fresh-derivation: forces focus on the SKILL, not the SOLUTIONS; iteration-isolated | +| Optimizer (step 8) | `07-analysis.md` + `01-functionality.md` + current skill state (`improved-skill/` if it exists, else original source) + `${OPERATOR_DIRECTIVES}` | Raw failed trials, `findings.txt`, grader internals, test inputs, `08-improvement-proposal.md` (own canonical) or its git history | Fresh-derivation: principled improvement, not pattern-match patches; no attachment to prior failed attempts. The skill content is upstream input; the report is the own canonical | +| Validator (step 9) | Skill BEFORE (same as optimizer's input) + skill AFTER (temporary materialization) + `08-improvement-proposal.md` + `01-functionality.md` + `03-submissions.md` (if exists) | Trial data, optimizer's reasoning trace, test inputs, `09-validator-verdict.md` (own canonical) or its git history | Fresh-derivation: independent check; can't be biased by what the optimizer (or a prior validator round) told itself | The skill (operator session) sees everything; subagents see slices. This is the architectural fix for the "tunnel-vision into ducktape" @@ -218,7 +223,7 @@ Common backtrack triggers: | Anywhere | User dislikes the output | re-run that step with new directives | There's no rigid backtrack flowchart — operator and auto-pilot both -decide based on the named weakness in `06-analysis.md`, the +decide based on the named weakness in `07-analysis.md`, the validator verdict, and the staleness mechanism below. ### Step kinds: fresh-derivation vs maintenance @@ -254,7 +259,7 @@ DOWNSTREAM_T=$(git log -1 --format=%ct -- tests/) In interactive use, the operator typically just knows ("I re-ran step 1, so step 2 needs a re-run"). The git-mtime check is for -auto-pilot (step 9), which walks the chain forward and re-runs any +auto-pilot (step 10), which walks the chain forward and re-runs any downstream older than its direct upstream. ### Re-entry contract @@ -370,7 +375,7 @@ Per-step iteration behavior is noted at the end of each subsection. - **Dispatches:** test-case-designer subagent (limited context: `01-functionality.md` + the current `tests//spec.yaml` tree if present + `${OPERATOR_DIRECTIVES}`; does NOT see skill - source content, `06-analysis.md`, optimizer attempts, failure + source content, `07-analysis.md`, optimizer attempts, failure data, or git history of the tree). Maintenance step — the subagent extends the tree (adds new functionality folders; modifies existing spec.yaml files per directives) rather than @@ -429,7 +434,7 @@ Per-step iteration behavior is noted at the end of each subsection. folders at `tests///{spec.yaml, workspace/, grader.mjs, smoke/}` + a generated `tests/suite.yml` that lists - the probes step 5 will run. + the probes step 6 will run. - **Behavior:** 1. Walk `tests/` for `picked: true` functionalities. 2. For each, decide the probe set (informed by @@ -455,33 +460,70 @@ Per-step iteration behavior is noted at the end of each subsection. - **Destructive-edit safety:** before rebuilding an existing probe (e.g., per a directive), the operator session commits the current state so git history has a clean before/after breakpoint. -- **Handoff:** "Invoke `skill-optimizer-run-bench` to measure - baseline." +- **Handoff:** "Invoke `skill-optimizer-validate-tests` to check + probe quality before benching." - **Iteration behavior:** re-run when step 2's tree changed (new `picked: true` functionalities; revised functionality specs that warrant probe rebuilds), when a test-writer flagged unimplementable - probes, or when bench results suggest probes are systematically - too easy or too narrow. Maintenance step — existing probe folders - are preserved across re-runs unless a directive explicitly asks - to rebuild a specific probe. - -### 5. `skill-optimizer-run-bench` + probes, when step 5's validator flagged probes for revision, or + when bench results suggest probes are systematically too easy or + too narrow. Maintenance step — existing probe folders are + preserved across re-runs unless a directive explicitly asks to + rebuild a specific probe. + +### 5. `skill-optimizer-validate-tests` + +- **Description trigger:** "validate the tests", "check the + graders", "are these probes fair", "review the test suite" +- **Input:** `tests/` tree (probes from step 4) + + `01-functionality.md` + source skill +- **Output:** `05-tests-verdict.md` — per-probe verdicts + (approve/needs-revision/reject) + aggregate frontmatter + (`all_probes_approved`, counts). Step 6 refuses to fire if + `all_probes_approved: false`. +- **Behavior:** dispatch test-validator subagents in parallel + (one per probe). Each judges: does the workspace exercise the + parent functionality? is the grader correct and fair? do the + smoke fixtures truly distinguish (vs. coincidentally match)? + Aggregate verdicts. +- **Dispatches:** test-validator subagent per probe (limited + context: this probe's full contents + parent functionality spec + + `01-functionality.md` + skill source; does NOT see other + probes, prior verdicts, downstream analysis) +- **Why this exists:** the smoke check at step 4 only verifies + syntactic consistency (grader correctly classifies the GOOD/BAD/ + EMPTY fixtures the test-writer also wrote). Semantic issues — + workspace doesn't exercise the responsibility, grader unfair, + false-positive probes — slip through. Grader bugs propagate to + misleading bench results and ducktape-shaped improvements. An + independent validator parallel to step 9 closes this gate. +- **Handoff (two branches):** if `all_probes_approved: true` → + "Invoke `skill-optimizer-run-bench`." If `false` → distill + per-probe verdicts into directives, re-run step 4 for affected + probes, re-run this step. Don't proceed to bench against bad + probes. +- **Iteration behavior:** re-run when step 4 produced new or + revised probes, or when user wants a fresh test-validation pass + with directives. Fresh-derivation — validator doesn't read its + own canonical or git history of it. + +### 6. `skill-optimizer-run-bench` - **Description trigger:** "run the eval", "measure", "benchmark" - **Input:** `tests/suite.yml` (generated by step 4) + source skill (vendored) - **Output (two artifacts):** - - **`05-bench-results//`** — raw CLI output: + - **`06-bench-results//`** — raw CLI output: `suite-result.json` + per-trial traces + per-trial findings.txt. Timestamped per run; old runs preserved naturally. - - **`05-bench-summary.md`** — single canonical aggregate: + - **`06-bench-summary.md`** — single canonical aggregate: overall pass rate, per-model pass rate, per-probe pass rate, failed-probe pointer list, pointer to the latest - `05-bench-results//`. Fresh-derivation per run; prior + `06-bench-results//`. Fresh-derivation per run; prior summary in git history. - **Behavior:** invoke skill-optimizer CLI (`npx tsx /src/cli.ts run-suite ./tests/suite.yml - --trials 3 --out 05-bench-results//`); parse the suite result; + --trials 3 --out 06-bench-results//`); parse the suite result; write the summary. - **Dispatches:** none — direct CLI invocation - **Note:** this is intentionally thin; the entire v1.3 run-suite @@ -494,12 +536,12 @@ Per-step iteration behavior is noted at the end of each subsection. bench-results dirs are outside the iteration protocol (they're naturally accumulating snapshots, not versioned artifacts). -### 6. `skill-optimizer-analyze` +### 7. `skill-optimizer-analyze` - **Description trigger:** "analyze the results", "diagnose what failed", "find structural weaknesses" -- **Input:** `05-bench-results//` + `workbench/` + source skill -- **Output:** `06-analysis.md` — structured per the format below +- **Input:** `06-bench-results//` + `workbench/` + source skill +- **Output:** `07-analysis.md` — structured per the format below - **Behavior:** cluster failures (per-rule, per-model, per-trial, per-pattern); separate flaky (single-trial randomness) from systematic (repeated across trials and/or models); for each @@ -518,13 +560,13 @@ Per-step iteration behavior is noted at the end of each subsection. - **Iteration behavior:** re-run when bench results change or when the operator/user wants a fresh look with new framing. Fresh-derivation — the subagent doesn't read its own canonical - `06-analysis.md` or its git history. Re-running overwrites the + `07-analysis.md` or its git history. Re-running overwrites the canonical; prior state in git. The `${OPERATOR_DIRECTIVES}` slot carries hints like "focus on the gpt-5 cluster" or "the user thinks weakness X is actually two separate issues" — additional analytical lenses, never a context dump of prior conclusions. -**`06-analysis.md` format (Option A — structured):** +**`07-analysis.md` format (Option A — structured):** ```markdown ## Structural weaknesses identified @@ -550,20 +592,20 @@ Per-step iteration behavior is noted at the end of each subsection. - 1 gpt-5 timeout (infrastructure, not skill) ``` -### 7. `skill-optimizer-improve` +### 8. `skill-optimizer-improve` - **Description trigger:** "improve this skill", "fix the structural weakness", "optimize" -- **Input:** `06-analysis.md` (REQUIRED — refuses if no weakness +- **Input:** `07-analysis.md` (REQUIRED — refuses if no weakness identified), `01-functionality.md`, current skill state (`improved-skill/` if it exists, else original source), optionally `03-submissions.md` -- **Output:** `07-improvement-proposal.md` (the optimizer's proposed +- **Output:** `08-improvement-proposal.md` (the optimizer's proposed change + rationale referencing the structural weakness). Step 7 - does NOT materialize `improved-skill/` — that's step 8's job + does NOT materialize `improved-skill/` — that's step 9's job after validator approval. - **Behavior:** - 1. Refuse if `06-analysis.md` has `has_structural_weakness: false` + 1. Refuse if `07-analysis.md` has `has_structural_weakness: false` — print "no weakness to address" and exit 2. Dispatch **optimizer subagent** with limited context (see "Subagent constraints" table). Required: address the named @@ -574,34 +616,34 @@ Per-step iteration behavior is noted at the end of each subsection. WOULD NOT address this" anti-pattern list. - **Dispatches:** optimizer subagent, limited context, isolated from raw trial data -- **Out of scope:** validating the proposal (that's step 8) and +- **Out of scope:** validating the proposal (that's step 9) and packaging the change as a PR draft (PR composition is a separate - downstream concern; the auto-pilot at step 9 or a dedicated + downstream concern; the auto-pilot at step 10 or a dedicated composer can handle it if `pr_submission_intent: true`). - **Single-shot per invocation:** no in-step revision loop. If - step 8 returns `needs-revision`, the operator (or auto-pilot) + step 9 returns `needs-revision`, the operator (or auto-pilot) distills the validator's rationale into a directive and - re-invokes step 7. + re-invokes step 8. - **Handoff:** "Next, invoke `skill-optimizer-validate` to check the proposal independently." -- **Iteration behavior:** re-run when `06-analysis.md` changed - (new analysis = potentially different weakness), when step 8 +- **Iteration behavior:** re-run when `07-analysis.md` changed + (new analysis = potentially different weakness), when step 9 returned `needs-revision`/`reject` with a distillable rationale, or when the operator wants a fresh optimization attempt. Fresh-derivation — the optimizer doesn't read its own canonical - (`07-improvement-proposal.md`) or git history of it, and doesn't - read step 8's prior verdicts. `${OPERATOR_DIRECTIVES}` carries + (`08-improvement-proposal.md`) or git history of it, and doesn't + read step 9's prior verdicts. `${OPERATOR_DIRECTIVES}` carries the distilled lessons. -### 8. `skill-optimizer-validate` +### 9. `skill-optimizer-validate` - **Description trigger:** "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix" -- **Input:** `07-improvement-proposal.md` (REQUIRED), +- **Input:** `08-improvement-proposal.md` (REQUIRED), `01-functionality.md`, current skill state (`improved-skill/` if it exists, else original source), optionally `03-submissions.md` -- **Output:** `08-validator-verdict.md` (the verdict + +- **Output:** `09-validator-verdict.md` (the verdict + rationale). On `verdict: approve`, ALSO materializes `improved-skill/` by applying the optimizer's diff to a copy of the current state. **The original source is never modified** — @@ -619,7 +661,7 @@ Per-step iteration behavior is noted at the end of each subsection. (frontmatter, file location, prefix taxonomy, additive-only, etc.)? Forward-looking — verifies the improved skill COULD be turned into a valid PR, even - though neither step 7 nor step 8 produces one. + though neither step 8 nor step 9 produces one. - Verdict: `approve` / `needs-revision` / `reject`. 2. On `verdict: approve`: materialize `improved-skill/` by applying the diff to a temp copy of the current state, then @@ -631,29 +673,29 @@ Per-step iteration behavior is noted at the end of each subsection. - **Single-shot per invocation:** no in-step revision loop. Each invocation produces one verdict; the 7→8→7 cycle on `needs-revision` is operator-driven (or auto-pilot-driven at - step 9). + step 10). - **Handoff (three branches by verdict):** - **approve:** "Validation complete; improvement applied at `improved-skill/`. The chain has reached its natural endpoint for this iteration. Review, copy locally, hand to a PR composer, or re-bench." - **needs-revision:** "Validator says needs-revision. Distill - the rationale into a directive and re-invoke step 7, then + the rationale into a directive and re-invoke step 8, then re-invoke this step." - **reject:** "Validator rejects the proposal outright. Two - paths: (1) re-invoke step 6 with a reframed weakness; (2) + paths: (1) re-invoke step 7 with a reframed weakness; (2) accept that this weakness isn't addressable and exit honestly." -- **Iteration behavior:** re-run when `07-improvement-proposal.md` - changed (step 7 produced a new proposal). Fresh-derivation — +- **Iteration behavior:** re-run when `08-improvement-proposal.md` + changed (step 8 produced a new proposal). Fresh-derivation — the validator doesn't read its own canonical - (`08-validator-verdict.md`) or git history of it. + (`09-validator-verdict.md`) or git history of it. Independence-from-self is load-bearing: if the validator saw its prior verdict it would gravitate toward consistency-with-itself across re-validation cycles, defeating the point of re-running. -### 9. `skill-optimizer-autopilot` +### 10. `skill-optimizer-autopilot` - **Description trigger:** "auto-pilot this skill", "run the whole chain on X", "skill-optimizer end-to-end for X", "automated @@ -665,7 +707,7 @@ Per-step iteration behavior is noted at the end of each subsection. `docs/skill-optimizer//autopilot-summary-.md` listing each step's final version, headline result, and any blockers - **Behavior:** - 1. Walks 1→8 in order, dispatching each chain skill. + 1. Walks 1→9 in order, dispatching each chain skill. 2. For each step: check whether the existing artifact is current via `git log` mtime comparison against direct upstream. If current, skip; if missing or stale, dispatch. @@ -676,7 +718,7 @@ Per-step iteration behavior is noted at the end of each subsection. functionalities by importance (default `pick_top_n: 5`) - **B8 validator-rejected verdict** → on `needs-revision`, distills the rationale into a directive and re-invokes - step 7 then step 8 (up to `max_iterations_per_step` rounds). + step 8 then step 9 (up to `max_iterations_per_step` rounds). On `reject`, logs the final state and surfaces blocker in the summary; does not retry automatically. 4. Bounds iterations: at most `max_iterations_per_step` re-runs @@ -689,7 +731,7 @@ Per-step iteration behavior is noted at the end of each subsection. (each chain skill internally dispatches its own narrow-context subagents). - **Caveats** (baked into SKILL.md): expect modest results compared - to operator-driven runs. The eight steps are hard even with human + to operator-driven runs. The nine steps are hard even with human judgment; auto-pilot is best for batch processing where some failures are acceptable, not for high-stakes single-target optimization. @@ -760,7 +802,7 @@ For v1.4 to be considered done: comparison and re-runs itself 6. **End-to-end test on a local skill** (e.g., one of the existing `skills/skill-optimizer/SKILL.md` or a small new skill in this - repo) — walks 1→2→4→5→6→7, produces all expected reports, modifies + repo) — walks 1→2→4→5→6→7→8→9, produces all expected reports, modifies the target skill, validator approves 7. **End-to-end test on an upstream skill** — re-run firecrawl (the v1.3 regression case) under v1.4; expected: optimizer either @@ -768,7 +810,7 @@ For v1.4 to be considered done: shipped) 8. **Auto-pilot smoke test** — running the auto-pilot on a local skill end-to-end produces a `autopilot-summary-.md` covering - all eight steps with their final versions and any blockers + all nine steps with their final versions and any blockers 9. Plugin manifest unchanged in shape (still `.claude-plugin/plugin.json`); skills auto-discoverable via the existing plugin loading mechanism @@ -801,10 +843,10 @@ For v1.4 to be considered done: rounds (2)?** Initial answer: surface as `validator-rejected` and require human intervention. Could add a "human-help-required" status. Will revisit. -3. **How does step 6's analyzer subagent distinguish "no pattern" from +3. **How does step 7's analyzer subagent distinguish "no pattern" from "pattern but I missed it"?** Initial answer: if the analyzer can't articulate a weakness, the next step refuses to fire — we don't - force improvement. The user can re-invoke step 6 with hints if they + force improvement. The user can re-invoke step 7 with hints if they disagree with the analyzer's verdict. ## Provenance diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md index 6e93b1f..0d8f483 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -1,22 +1,22 @@ --- name: skill-optimizer-analyze -description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer-run-bench` has produced a `05-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. +description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer-run-bench` has produced a `06-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. --- # skill-optimizer-analyze Step 6 of the skill-optimizer chain. **Fresh-derivation step.** Takes -the bench summary + raw trial output from step 5, dispatches an +the bench summary + raw trial output from step 6, dispatches an analyzer subagent to cluster failures into named **structural weaknesses** of the skill (or explicitly say there are none), and -writes `docs/skill-optimizer//06-analysis.md`. This is the -chain's **anti-ducktape gate**: step 7 refuses to fire unless this +writes `docs/skill-optimizer//07-analysis.md`. This is the +chain's **anti-ducktape gate**: step 8 refuses to fire unless this report names at least one structural weakness, with the general principle that WOULD address it and the anti-patterns that would NOT. ## What you produce -A single report at `docs/skill-optimizer//06-analysis.md`. +A single report at `docs/skill-optimizer//07-analysis.md`. Frontmatter (runtime-relevant facts only, per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): @@ -25,12 +25,12 @@ Frontmatter (runtime-relevant facts only, per --- has_structural_weakness: true | false weakness_count: -bench_results_path: 05-bench-results// +bench_results_path: 06-bench-results// --- ``` -**`has_structural_weakness`** is load-bearing for step 7 — `false` -gates step 7 from firing ("no weakness to address"). Forcing +**`has_structural_weakness`** is load-bearing for step 8 — `false` +gates step 8 from firing ("no weakness to address"). Forcing `true` when the analyzer found nothing is the ducktape failure this step exists to prevent. @@ -41,9 +41,9 @@ Connects to skill section, What WOULD address this, What WOULD NOT address this. Step (d) verifies these five parts are present. The "What WOULD NOT address this" anti-pattern list is -**load-bearing architecture**. Without it, step 7's optimizer can +**load-bearing architecture**. Without it, step 8's optimizer can pattern-match a patch that fits the symptom without addressing the -cause; the validator (step 8) then has no explicit "this would be +cause; the validator (step 9) then has no explicit "this would be a ducktape" signal to check against. Losing the anti-pattern list breaks the gate. @@ -54,7 +54,7 @@ Full body template and reasoning protocol in ### (a) Confirm prerequisites -`05-bench-summary.md` must exist with `bench_results_path` and +`06-bench-summary.md` must exist with `bench_results_path` and `overall_pass_rate`. If not, tell the user to run `skill-optimizer-run-bench` first. @@ -65,7 +65,7 @@ the analyzer; there are no failures to cluster. If the bench results dir is missing or `suite-result.json` is malformed, surface as a step-5 problem (incomplete or corrupted -bench run) and tell the user to re-run step 5. +bench run) and tell the user to re-run step 6. `tests/` and the source skill must also be available. @@ -92,8 +92,8 @@ and substitute `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. -The subagent sees: `05-bench-summary.md` (entry point with -failed-probe pointer list); `05-bench-results//` with +The subagent sees: `06-bench-summary.md` (entry point with +failed-probe pointer list); `06-bench-results//` with per-trial `trace.jsonl` and `findings.txt`; each probe's `spec.yaml` from the `tests/` tree (intent only); the skill source content; `${OPERATOR_DIRECTIVES}`. @@ -101,8 +101,8 @@ content; `${OPERATOR_DIRECTIVES}`. The subagent does NOT see: **the test inputs themselves** (`tests//workspace/` files) — this is load-bearing; the analyzer must think about the SKILL, not the SOLUTIONS; its own -prior `06-analysis.md` or git history; `07-improvement-proposal.md`, -`08-validator-verdict.md`, or any prior optimizer attempts. +prior `07-analysis.md` or git history; `08-improvement-proposal.md`, +`09-validator-verdict.md`, or any prior optimizer attempts. **Why this matters:** two specific bias risks. **Solution-thinking** — seeing the raw fixtures would lead the analyzer to recommend a @@ -118,7 +118,7 @@ time" or defensively pivot away from it. Neither is the job. Verify the report parses, `weakness_count` matches the number of `### Weakness :` sections, and each weakness has all five required parts (especially "What WOULD NOT address this" — the -anti-ducktape signal step 7 needs). If anything's inconsistent or +anti-ducktape signal step 8 needs). If anything's inconsistent or missing, surface to the user; don't fill it in yourself. If the anti-pattern lists are empty/vague (anti-ducktape gate @@ -139,6 +139,6 @@ Two messages depending on `has_structural_weakness`: If the subagent returned `false` but the user disagrees: surface the disagreement, but do NOT pressure the subagent to manufacture -a weakness. Honest path is re-invoking step 6 with a directive +a weakness. Honest path is re-invoking step 7 with a directive pointing at what the user thinks was missed. Don't auto-invoke -step 6 — per the no-auto-invocation rule, the user decides. +step 7 — per the no-auto-invocation rule, the user decides. diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index c1349ab..5592525 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -1,29 +1,29 @@ --- name: skill-optimizer-improve -description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze` has produced `06-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. +description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze` has produced `07-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. --- # skill-optimizer-improve Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes -the named structural weaknesses from step 6, dispatches an optimizer +the named structural weaknesses from step 7, dispatches an optimizer subagent to draft a principled fix, and writes -`07-improvement-proposal.md`. The proposal is then validated -independently by step 8, which materializes the improved skill on +`08-improvement-proposal.md`. The proposal is then validated +independently by step 9, which materializes the improved skill on approve. **This step does not produce the improved skill itself** — -it produces the proposal that step 8 acts on. +it produces the proposal that step 9 acts on. -**Refuses to fire** if `06-analysis.md` has +**Refuses to fire** if `07-analysis.md` has `has_structural_weakness: false` — there's nothing to optimize, and forcing a fix is the ducktape failure mode the chain is built to prevent. **Single-shot per invocation.** No in-step revision loop. If step 8's validator returns `needs-revision`, the operator (or -auto-pilot at step 9) distills the validator's rationale into a +auto-pilot at step 10) distills the validator's rationale into a directive and re-invokes this step. -**Out of scope:** validating the proposal (step 8) and packaging +**Out of scope:** validating the proposal (step 9) and packaging the change as a PR draft (PR composition is a separate downstream concern; auto-pilot or a dedicated composer can handle it if `pr_submission_intent: true`). @@ -32,14 +32,14 @@ concern; auto-pilot or a dedicated composer can handle it if One artifact at `docs/skill-optimizer//`: -**`07-improvement-proposal.md`** — the optimizer's proposed change +**`08-improvement-proposal.md`** — the optimizer's proposed change plus rationale. Frontmatter (runtime-relevant facts only, per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): ```yaml --- addresses_weaknesses: - - + - - ... --- ``` @@ -59,13 +59,13 @@ specifies the section shape. Three checks: -1. `06-analysis.md` must exist with valid frontmatter. If not, +1. `07-analysis.md` must exist with valid frontmatter. If not, tell the user to run `skill-optimizer-analyze` first. 2. **Anti-ducktape gate:** `has_structural_weakness: true` must be set. If `false`, REFUSE — print: "Step 6 found no structural weakness. Step 7 won't fire — nothing principled to optimize. - If you disagree, re-invoke step 6 with a directive; if you + If you disagree, re-invoke step 7 with a directive; if you agree, exit honestly." Do NOT proceed. 3. `01-functionality.md` must exist. Skill content must be @@ -74,14 +74,14 @@ Three checks: If `pr_submission_intent: true`, `03-submissions.md` should exist so the optimizer can shape the diff to upstream conventions from -the start. If missing, ask whether to run step 3 first — step 8's +the start. If missing, ask whether to run step 3 first — step 9's external check still runs if it appears later. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from -`06-analysis.md` plus current skill plus directives. Collect +`07-analysis.md` plus current skill plus directives. Collect `${OPERATOR_DIRECTIVES}` — examples: "prefer additive changes", "don't touch the description field — validator rejected that last round". The operator reads prior proposals/verdicts and distills; @@ -102,7 +102,7 @@ and substitute `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, `${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. -The optimizer sees: `06-analysis.md`; `01-functionality.md`; +The optimizer sees: `07-analysis.md`; `01-functionality.md`; `${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else original source; `03-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. @@ -110,25 +110,25 @@ original source; `03-submissions.md` if PR-bound; The optimizer does NOT see: raw failed trials, `findings.txt`, `trace.jsonl`; grader internals (`tests//grader.mjs`); test inputs (`tests//workspace/`); its own prior -`07-improvement-proposal.md` or git history; step 8's prior -`08-validator-verdict.md` (would bias toward defending or pivoting +`08-improvement-proposal.md` or git history; step 9's prior +`09-validator-verdict.md` (would bias toward defending or pivoting away from the prior attempt). **Why this matters:** the chain's anti-ducktape architecture hinges on this. If the optimizer read failed trials, it would pattern-match a patch making those specific trials pass — the -textbook ducktape mode. The analyzer's job (step 6) is to +textbook ducktape mode. The analyzer's job (step 7) is to translate raw failures into general principles + anti-patterns; the optimizer's job is to apply the principle. The translation -through `06-analysis.md` is what enforces principled improvement. +through `07-analysis.md` is what enforces principled improvement. -The optimizer writes `07-improvement-proposal.md` and returns: +The optimizer writes `08-improvement-proposal.md` and returns: weakness(es) addressed, principle applied, lines/sections modified. If the optimizer reports BLOCKED (the analyzer's named weakness is too abstract to derive a concrete diff from), surface to the -user. Fix is typically a step 6 re-run with a reframing directive; -per the no-auto-invocation rule, the user invokes step 6. +user. Fix is typically a step 7 re-run with a reframing directive; +per the no-auto-invocation rule, the user invokes step 7. ### (d) Confirm subagent output @@ -142,10 +142,10 @@ requirement explicit — don't fill it in yourself. ### (e) Hand off > Improvement proposal complete at -> `07-improvement-proposal.md`. Next, invoke +> `08-improvement-proposal.md`. Next, invoke > `skill-optimizer-validate` to check the proposal > independently. If the validator returns `needs-revision` or > `reject`, you'll come back here with a directive distilling the > validator's concerns. -Don't auto-invoke step 8. +Don't auto-invoke step 9. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 2cb0397..97c658d 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -10,7 +10,7 @@ target skill is bound for upstream PR submission. **Fresh-derivation step.** Takes the source slug, dispatches a researcher subagent that uses the `gh` CLI to gather the upstream repo's contribution conventions, and writes `docs/skill-optimizer//03-submissions.md` -— the verbatim-pastable context block the validator (step 8) uses for +— the verbatim-pastable context block the validator (step 9) uses for its external consistency check. ## What you produce @@ -88,7 +88,7 @@ git history of it; any information about the proposed change being optimized; existing analyses, tests, or failure data; the vendored skill source. -**Why this matters:** the validator (step 8) will later check the +**Why this matters:** the validator (step 9) will later check the proposed change against this report. If the operator session does the research, it has already absorbed the optimization context — the proposed change, prior failures, user's framing — and biases @@ -106,7 +106,7 @@ Edge cases the subagent will surface as blockers: making up patterns - **Upstream uses non-discoverable frontmatter conventions** — if the subagent can't extract a consistent spec from recent merged - PRs, the report will say so; the optimizer (step 7) then makes a + PRs, the report will say so; the optimizer (step 8) then makes a judgment call rather than mechanically conforming The subagent returns a brief summary: license, CLA requirement, diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-investigate-test-case/SKILL.md index 52c22a9..1e4d575 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-investigate-test-case/SKILL.md @@ -92,8 +92,8 @@ tree (load-bearing state per the maintenance rule), The subagent does NOT see: the skill's source content (would gerrymander tests around the source's literal phrasing — -design from STATED responsibilities); `06-analysis.md`, -`07-improvement-proposal.md`, `08-validator-verdict.md`; raw +design from STATED responsibilities); `07-analysis.md`, +`08-improvement-proposal.md`, `09-validator-verdict.md`; raw failure data, bench results; git history of any tree file. **Why this matters:** the operator session has absorbed prior @@ -152,7 +152,7 @@ Three realistic responses: Don't auto-flip `picked` on the user's behalf. Even if all functionalities look important, the user owns the decision (they're -paying for probe-building in step 4 and the bench run in step 5). +paying for probe-building in step 4 and the bench run in step 6). ### (f) Hand off diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 3a57756..beb35ae 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -9,26 +9,26 @@ Step 5 of the skill-optimizer chain. **Fresh-derivation step (for the summary).** Invokes the skill-optimizer CLI's `run-suite` command against `tests/suite.yml` (generated by step 4), captures the raw results under a timestamped directory, and writes a small -summary report that step 6 reads. No subagent dispatch — this is a +summary report that step 7 reads. No subagent dispatch — this is a thin operator-driven CLI step. ## What you produce Two outputs at `docs/skill-optimizer//`: -1. **`05-bench-results//`** — raw CLI output: +1. **`06-bench-results//`** — raw CLI output: `suite-result.json`, per-trial `trace.jsonl`, per-trial `findings.txt`, preserved workspaces. Timestamped per invocation; old runs are NEVER overwritten. Outside the iteration protocol — each timestamped dir is its own naturally-accumulated archive. -2. **`05-bench-summary.md`** — single canonical aggregate, per +2. **`06-bench-summary.md`** — single canonical aggregate, per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): ```yaml --- - bench_results_path: 05-bench-results// + bench_results_path: 06-bench-results// overall_pass_rate: --- ``` @@ -51,12 +51,12 @@ missing, tell the user to complete step 4 first. Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply it. Two pieces: -- **Raw bench output** at `05-bench-results//`: always +- **Raw bench output** at `06-bench-results//`: always write to a fresh timestamped directory (`date -u +%Y%m%dT%H%M%SZ`). Outside the iteration protocol; never overwrite. -- **Summary file** at `05-bench-summary.md`: fresh-derivation per +- **Summary file** at `06-bench-summary.md`: fresh-derivation per the protocol. Staleness check: if `tests/suite.yml` has a newer - git mtime than `05-bench-summary.md`, the existing summary is + git mtime than `06-bench-summary.md`, the existing summary is stale. If summary is current AND no re-run directive, tell the user and exit. @@ -64,7 +64,7 @@ and apply it. Two pieces: ```bash TIMESTAMP=$(date -u +%Y%m%dT%H%M%SZ) -OUT_DIR="docs/skill-optimizer//05-bench-results/${TIMESTAMP}" +OUT_DIR="docs/skill-optimizer//06-bench-results/${TIMESTAMP}" mkdir -p "${OUT_DIR}" npx tsx /src/cli.ts run-suite \ @@ -74,7 +74,7 @@ npx tsx /src/cli.ts run-suite \ ``` `--trials 3` is the chain default — enough to distinguish flaky -from systematic failures at step 6. Honor a different count if +from systematic failures at step 7. Honor a different count if requested. Models come from `suite.yml` (per project invariant: `run-suite` does NOT take a `--models` override). Stream stdout/stderr to the user — bench runs take minutes to hours. @@ -90,7 +90,7 @@ Environment failures to surface honestly (don't try to recover): ### (d) Write the summary -Parse `${OUT_DIR}/suite-result.json` and write `05-bench-summary.md` +Parse `${OUT_DIR}/suite-result.json` and write `06-bench-summary.md` with the frontmatter above plus a body containing: - **Overall:** total trials, passed, failed, overall pass rate. @@ -105,7 +105,7 @@ with the frontmatter above plus a body containing: If all trials errored (no graded results), that's a workbench misconfiguration or environmental failure rather than a skill weakness. Write the summary honestly and surface before handing -off to step 6. +off to step 7. ### (e) Hand off @@ -115,7 +115,7 @@ Read `overall_pass_rate`. Two messages: failures. - `== 1.0`: surface the choice — accept that probes don't expose a weakness, or re-run step 2 with a "make probes harder" - directive. Don't auto-invoke step 6. + directive. Don't auto-invoke step 7. ## Edge cases @@ -123,5 +123,5 @@ Read `overall_pass_rate`. Two messages: accept a case filter, so there's no first-class partial-rebench mode. Operators with expensive suites can run `run-case` manually for changed probes and splice into the prior - `05-bench-results//` dir — outside-the-chain escape hatch, + `06-bench-results//` dir — outside-the-chain escape hatch, not a supported flow. diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 8dab53a..a429a54 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -36,10 +36,11 @@ upstream + directives. Re-runs overwrite; git captures prior state. |---|---| | 1. investigate-functionality | Each invocation researches from source — no continuity needed | | 3. investigate-submissions | Each invocation researches upstream — no continuity needed | -| 5. run-bench (summary file) | Mechanical write-up of the new raw bench run, not an extension of prior summary | -| 6. analyze-result | Must not be biased by prior analyses; load-bearing for anti-ducktape | -| 7. improve-skill | Optimizer must not see prior attempts; load-bearing for anti-ducktape | -| 8. validate-improvement | Validator must not be biased by its prior verdicts; independence-from-self is load-bearing | +| 5. validate-tests | Each probe judged fresh against its parent functionality; load-bearing for anti-ducktape | +| 6. run-bench (summary file) | Mechanical write-up of the new raw bench run, not an extension of prior summary | +| 7. analyze | Must not be biased by prior analyses; load-bearing for anti-ducktape | +| 8. improve | Optimizer must not see prior attempts; load-bearing for anti-ducktape | +| 9. validate | Validator must not be biased by its prior verdicts; independence-from-self is load-bearing | **Maintenance steps** manage an accumulating tree on disk — the filesystem itself is the state. Re-runs read the current tree and @@ -63,7 +64,7 @@ DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) In interactive use, the operator typically just knows ("I re-ran step 1, so step 2 needs a re-run"). The git-mtime check is for -auto-pilot (step 9), which walks the chain forward and re-runs any +auto-pilot (step 10), which walks the chain forward and re-runs any downstream older than its direct upstream. For maintenance steps, the operator's directives can also force a @@ -116,8 +117,8 @@ needed — git is the archive. ## What this protocol does NOT cover - **Step 5's raw bench results** are timestamped under - `05-bench-results//`. Each run preserves naturally as its own - directory. The summary file `05-bench-summary.md` is a + `06-bench-results//`. Each run preserves naturally as its own + directory. The summary file `06-bench-summary.md` is a fresh-derivation artifact per this protocol; the timestamped raw output is outside. - **Auto-pilot's summary report** at diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md index 70f1489..aff4012 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -34,7 +34,7 @@ Three rules every reasoning subagent must follow: filesystem is the source of truth; do not walk `git log` looking for prior versions of upstream files. -2. **For fresh-derivation steps (1, 3, 5-summary, 6, 7, 8): do NOT +2. **For fresh-derivation steps (1, 3, 5, 6-summary, 7, 8, 9): do NOT read your own canonical file, and do NOT read git history of it.** Each invocation derives a new artifact from upstream + the operator's directives, without direct access to prior @@ -50,15 +50,15 @@ Three rules every reasoning subagent must follow: derivation from rationalizing the prior one while still letting the chain converge. - **For step 7 and step 8 specifically:** the SKILL CONTENT (the + **For step 8 and step 9 specifically:** the SKILL CONTENT (the target being improved) is upstream input, not your own - canonical. Both the optimizer (step 7) and the validator (step + canonical. Both the optimizer (step 8) and the validator (step 8) read the **current state of the skill** — which is `improved-skill/` if it exists (the accumulated state from prior step-8 approvals), else the original source. Neither step modifies the original. The "own canonical" off-limits to each - subagent is its report (`07-improvement-proposal.md` for the - optimizer; `08-validator-verdict.md` for the validator), not + subagent is its report (`08-improvement-proposal.md` for the + optimizer; `09-validator-verdict.md` for the validator), not the skill content itself. 3. **For maintenance steps (2, 4): DO read your own canonical diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index d214222..20f6aa7 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -1,6 +1,6 @@ # skill-optimizer workflow -Operator-facing reference for the 8-step skill-optimizer chain. +Operator-facing reference for the 10-step skill-optimizer chain. Each step's SKILL.md is self-contained for its own work; this doc covers cross-skill concerns: the high-level flow, what triggers re-runs, and the relationship between steps. @@ -13,34 +13,44 @@ re-runs, and the relationship between steps. | 2 | `investigate-test-case` | maintenance | step 1 | `02-test-proposals.md`, `tests//spec.yaml` | | 3 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `03-submissions.md` | | 4 | `write-tests` | maintenance | steps 1+2 | `tests///`, `tests/suite.yml` | -| 5 | `run-bench` | fresh-derivation (summary) | step 4 + skill | `05-bench-results//`, `05-bench-summary.md` | -| 6 | `analyze-result` | fresh-derivation | step 5 + skill | `06-analysis.md` | -| 7 | `improve-skill` | fresh-derivation | step 6 + skill (+ step 3 if PR-bound) | `07-improvement-proposal.md` | -| 8 | `validate-improvement` | fresh-derivation | step 7 + skill (+ step 3 if PR-bound) | `08-validator-verdict.md`, `improved-skill/` (on approve) | -| 9 | `autopilot` | chain driver | same as step 1 + flags | `autopilot-summary-.md` | +| 5 | `validate-tests` | fresh-derivation | step 4 + step 1 + skill | `05-tests-verdict.md` | +| 6 | `run-bench` | fresh-derivation (summary) | step 5 (must approve) + skill | `06-bench-results//`, `06-bench-summary.md` | +| 7 | `analyze` | fresh-derivation | step 6 + skill | `07-analysis.md` | +| 8 | `improve` | fresh-derivation | step 7 + skill (+ step 3 if PR-bound) | `08-improvement-proposal.md` | +| 9 | `validate` | fresh-derivation | step 8 + skill (+ step 3 if PR-bound) | `09-validator-verdict.md`, `improved-skill/` (on approve) | +| 10 | `autopilot` | chain driver | same as step 1 + flags | `autopilot-summary-.md` | The PR-or-not decision is made ONCE at step 1; subsequent steps know from `01-functionality.md`'s `pr_submission_intent` field whether step 3 will run. No late prompts. +Two independent validators in the chain: step 5 (validate-tests) +checks that probes fairly test their functionality before +measurement; step 9 (validate) checks that improvements address +the named weakness in a principled way. Both are anti-ducktape +gates — step 6 refuses to fire if step 5 didn't approve all +probes, and step 9 won't materialize `improved-skill/` unless it +approves the proposal. + ## Re-run triggers (when each step should be re-invoked) A chain skill never auto-invokes another chain skill (per -[`subagent-dispatch.md`](./subagent-dispatch.md) -re-run authorization). The operator (or auto-pilot at step 9) -decides when to re-run each step. Common triggers: +[`subagent-dispatch.md`](./subagent-dispatch.md) re-run +authorization). The operator (or auto-pilot at step 10) decides +when to re-run each step. Common triggers: | Step | Re-run when... | |---|---| | 1 | source URL changed; PR-intent changed; user wants fresh research with new directives | -| 2 | user wants different coverage; step 4/5/6 surfaced a coverage gap; step 1 changed | -| 3 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed (source URL updated) | -| 4 | step 2's `tests/` tree changed; a probe's smoke check failed; step 5 showed probes systematically too easy/hard | -| 5 | step 4's `tests/` changed; user wants fresh trial data (flakiness, model list changed); step 6 wants more trials | -| 6 | step 5 produced new results; user disagrees with prior verdict; step 7 was unable to address a named weakness | -| 7 | step 6 produced a new analysis; step 8 returned `needs-revision`/`reject` with a distillable rationale; user wants a different approach | -| 8 | step 7 produced a new proposal; user disagrees with prior verdict; step 3 was updated and prior external check is stale | -| 9 | self-iterable; picks up at whatever step is stale per git-mtime | +| 2 | user wants different coverage; step 4/5/6/7 surfaced a coverage gap; step 1 changed | +| 3 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed | +| 4 | step 2's `tests/` tree changed; a probe's smoke check failed; step 5 flagged probes for revision; step 7 showed probes systematically too easy/hard | +| 5 | step 4 produced new or revised probes; user disagrees with prior test verdict | +| 6 | step 4's `tests/` changed (and step 5 re-approved); user wants fresh trial data (flakiness, model list changed); step 7 wants more trials | +| 7 | step 6 produced new results; user disagrees with prior analysis; step 8 was unable to address a named weakness | +| 8 | step 7 produced a new analysis; step 9 returned `needs-revision`/`reject` with a distillable rationale; user wants a different approach | +| 9 | step 8 produced a new proposal; user disagrees with prior verdict; step 3 was updated and prior external check is stale | +| 10 | self-iterable; picks up at whatever step is stale per git-mtime | Each skill checks **direct upstream only** for staleness — if step 1 went stale but step 2 wasn't re-run (user judged it still @@ -56,11 +66,12 @@ auto-fire; they're surfaced to the user. |---|---|---| | 2 | proposal is thin (subagent found few responsibilities) | re-run step 1 with directives, OR accept | | 4 | test-writer reports BLOCKED on a probe | re-run step 2 to refine the functionality spec | -| 5 | bench all-pass (probes too easy) | re-run step 2 with "make probes harder" directive | -| 6 | analyzer's `has_structural_weakness: false`, user disagrees | re-run step 6 with directive pointing at missed cluster | -| 7 | optimizer reports BLOCKED (weakness too abstract) | re-run step 6 to reformulate weakness | -| 8 | `needs-revision` | distill validator's rationale, re-run step 7, re-run step 8 | -| 8 | `reject` | re-run step 6 with reframed weakness, OR accept that weakness isn't addressable | +| 5 | any probe is `needs-revision`/`reject` | distill verdict into directives, re-run step 4 for affected probes, re-run step 5 | +| 6 | bench all-pass (probes too easy) | re-run step 2 with "make probes harder" directive | +| 7 | analyzer's `has_structural_weakness: false`, user disagrees | re-run step 7 with directive pointing at missed cluster | +| 8 | optimizer reports BLOCKED (weakness too abstract) | re-run step 7 to reformulate weakness | +| 9 | `needs-revision` | distill validator's rationale, re-run step 8, re-run step 9 | +| 9 | `reject` | re-run step 7 with reframed weakness, OR accept that weakness isn't addressable | ## State layout @@ -79,32 +90,34 @@ docs/skill-optimizer// checks/smoke.mjs suite.yml # B4 generates 03-submissions.md # only if step 3 ran - 05-bench-results// # raw, timestamped per run - 05-bench-summary.md # single canonical - 06-analysis.md - 07-improvement-proposal.md - 08-validator-verdict.md + 05-tests-verdict.md # B5 writes (per-probe + aggregate) + 06-bench-results// # raw, timestamped per run + 06-bench-summary.md # single canonical + 07-analysis.md + 08-improvement-proposal.md + 09-validator-verdict.md vendored-skill/ # frozen original (upstream only) - improved-skill/ # B8 materializes on approve - autopilot-summary-.md # B9 writes per run + improved-skill/ # B9 materializes on approve + autopilot-summary-.md # B10 writes per run ``` `` is `--` for upstream skills, or `` for local skills. Filesystem IS the state; history is git (no `version:` fields or `archive/` directories -per -[`frontmatter-discipline.md`](./frontmatter-discipline.md)). +per [`frontmatter-discipline.md`](./frontmatter-discipline.md)). ## Shared docs The chain skills load these on-demand at the workflow steps that need them: -- [`iteration-protocol.md`](./iteration-protocol.md) - — iteration mechanics (staleness, step kinds, +- [`iteration-protocol.md`](./iteration-protocol.md) — + iteration mechanics (staleness, step kinds, destructive-edit checkpoints, cascading, bootstrapping) -- [`subagent-dispatch.md`](./subagent-dispatch.md) - — subagent constraints, operator directives, templated dispatch +- [`subagent-dispatch.md`](./subagent-dispatch.md) — + subagent constraints, operator directives, templated dispatch inputs, no-auto-invocation rule -- [`frontmatter-discipline.md`](./frontmatter-discipline.md) - — runtime facts vs. history rule +- [`frontmatter-discipline.md`](./frontmatter-discipline.md) — + runtime facts vs. history rule +- [`workbench.md`](./workbench.md) — workbench schema (probe + layout, grader contract, smoke-check format); referenced by B4 diff --git a/skills/skill-optimizer-validate-tests/SKILL.md b/skills/skill-optimizer-validate-tests/SKILL.md new file mode 100644 index 0000000..7b3d131 --- /dev/null +++ b/skills/skill-optimizer-validate-tests/SKILL.md @@ -0,0 +1,172 @@ +--- +name: skill-optimizer-validate-tests +description: Use when the user wants to validate the probes built by `skill-optimizer-write-tests` before running the bench — phrases like "validate the tests", "check the graders", "are these probes fair", "review the test suite". Triggers after `skill-optimizer-write-tests` has populated `tests///` folders and before `skill-optimizer-run-bench`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking that the probes actually probe what they claim should trigger this. +--- + +# skill-optimizer-validate-tests + +Step 5 of the skill-optimizer chain. **Fresh-derivation step.** +Takes the probes built by step 4, dispatches a test-validator +subagent per probe (in parallel) to independently check whether +each probe fairly tests its parent functionality and whether its +grader is sound, and writes +`docs/skill-optimizer//05-tests-verdict.md` — the aggregate +verdict step 6 (run-bench) checks before proceeding. + +This step exists because the smoke check at step 4 only verifies +**syntactic** consistency (the grader correctly classifies the +GOOD/BAD/EMPTY fixtures the test-writer also wrote). It can't catch +**semantic** problems: workspace doesn't actually exercise the +functionality, grader is too strict/loose beyond the smoke +fixtures, false-positive probes that the skill could pass without +doing the right thing. Grader bugs are load-bearing — they +propagate to misleading bench results, misleading analyses, and +ducktape-shaped improvements. An independent validator is the +chain's anti-ducktape gate for the test layer, parallel to step 9 +for the optimizer layer. + +**Single-shot per invocation.** No in-step revision loop. If any +probe is `needs-revision` or `reject`, the operator (or auto-pilot +at step 10) distills the validator's rationale into a directive +and re-invokes step 4 for the affected probes, then re-invokes +this step. + +## What you produce + +One artifact at `docs/skill-optimizer//`: + +**`05-tests-verdict.md`** — aggregate verdict across all probes +plus per-probe sections. Frontmatter (runtime-relevant facts only, +per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): + +```yaml +--- +all_probes_approved: true | false +probe_count: +needs_revision_count: +reject_count: +--- +``` + +**`all_probes_approved`** is load-bearing for step 6 — if `false`, +step 6 refuses to fire ("can't measure against bad probes"). Forcing +`true` when probes had real issues is the test-layer ducktape failure +mode this step exists to prevent. + +Body has per-probe sections grouped by functionality. Each probe +entry includes: verdict (approve / needs-revision / reject), +rationale (covering workspace fairness, grader correctness, smoke +fixture distinguishing power, fairness across reasonable agent +outputs), and — for needs-revision/reject — a concrete suggestion +the operator can distill into a directive for step 4. + +Body template and the validator's reasoning protocol live at +[`skills/skill-optimizer-subagents/test-validator.md`](../skill-optimizer-subagents/test-validator.md). + +## Workflow + +### (a) Confirm prerequisites + +`tests/` must exist with at least one +`tests///` folder containing `spec.yaml`, +`workspace/`, `grader.mjs`, `smoke/`. If any picked functionality +has zero built probes, surface as a step-4 problem and tell the +user to run step 4 first. + +`01-functionality.md` must also exist (the validator reads it for +context on what the skill is supposed to do). + +### (b) Handle iteration + +Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +and apply it. Each invocation overwrites the canonical verdict +from current probe state plus directives. Collect +`${OPERATOR_DIRECTIVES}` per the protocol — examples: "be stricter +on grader fairness", "the prior verdict missed that probe X +requires a specific JSON shape the skill never asks for". + +### (c) Dispatch test-validator subagents (parallel, one per probe) + +Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +for the constraints. **Do NOT judge probes yourself in this +session** — dispatch test-validator subagents via the `Agent` tool, +in parallel (emit all dispatches in a single message). Load the +prompt template at +[`skills/skill-optimizer-subagents/test-validator.md`](../skill-optimizer-subagents/test-validator.md) +and substitute per-probe inputs: `${PROBE_NAME}`, +`${PROBE_DIR}`, `${FUNCTIONALITY_SPEC_PATH}`, +`${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, +`${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. + +Each test-validator sees: its single probe's full contents +(spec.yaml, workspace files, grader.mjs, smoke fixtures); the +parent functionality's spec.yaml; `01-functionality.md`; the skill +source content; `${OPERATOR_DIRECTIVES}` for its probe. + +Each test-validator does NOT see: other probes (independence — a +validator that sees siblings would gravitate toward +consistency-with-siblings rather than judging each probe on its +own); its own prior verdict or git history; the test-writer's +reasoning (only the probe artifacts, not how they were arrived +at); `07-analysis.md`, `08-improvement-proposal.md`, or any +downstream output. + +**Why this matters:** the test-writer at step 4 has built both the +fixture AND the grader AND the smoke fixtures — self-validation +only catches syntactic problems. An independent validator that +reads the same probe from scratch (against the functionality spec) +catches semantic problems: probes that pass without exercising the +responsibility, graders unfair to reasonable agent outputs, smoke +fixtures that coincidentally match rather than truly distinguish. +The validator must judge "would I, as an external reviewer, accept +this as a fair test of the parent functionality?". + +Each test-validator returns: probe name, verdict, one-line +rationale, and the suggestion-for-revision (if not approved). + +### (d) Aggregate verdicts + write the report + +Collect per-probe verdicts. Compute aggregate: + +- `all_probes_approved = true` if every probe is `approve` +- `needs_revision_count` and `reject_count` from per-probe verdicts + +Write `05-tests-verdict.md` with the frontmatter above and a body +grouped by functionality, with per-probe sections containing +verdict, rationale, and suggestion-for-revision. + +If `all_probes_approved: true` but the body shows any +`needs-revision`/`reject` (or vice versa), the verdict is +internally inconsistent — surface to the user; don't try to +reconcile yourself. + +### (e) Hand off + +Two messages depending on `all_probes_approved`: + +- **`true`:** "All probes approved. Next, invoke + `skill-optimizer-run-bench` to measure baseline performance." +- **`false`:** "`` probes need revision, `` rejected. The + validator's per-probe rationale is in `05-tests-verdict.md`. + Distill the concerns into directives and re-invoke + `skill-optimizer-write-tests` for the affected probes (per the + per-probe rebuild flow at step 4), then re-invoke this step. + Don't proceed to bench against bad probes." + +Don't auto-invoke step 4 or step 6. + +## Edge cases + +- **No probes built (tests/ empty)** — caught at (a). Tell the + user to run step 4 first. +- **Validator's verdict contradicts the smoke check** (e.g., + validator rejects a probe whose smoke check passed) — that's + the expected case! Smoke is syntactic; validator is semantic. + The validator wins; proceed with `needs-revision`/`reject`. +- **Validator approves a probe whose smoke check failed** — that + shouldn't happen (smoke check at step 4 surfaces failures + immediately), but if it does, treat as inconsistent and + re-invoke step 4 to fix the smoke check first. +- **Operator directives reference a specific section of the prior + verdict** — context-dump masquerading as a directive. Translate + to an atomic new requirement before passing. diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index 258adf4..993c363 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -1,28 +1,28 @@ --- name: skill-optimizer-validate -description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve` has produced `07-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. +description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve` has produced `08-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. --- # skill-optimizer-validate Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes -the proposal from step 7, dispatches a validator subagent to +the proposal from step 8, dispatches a validator subagent to independently check whether the proposed change is sound (internal consistency) and conformant (external PR conventions if PR-bound), then — on `verdict: approve` — materializes the improved skill at -`improved-skill/`. Writes `08-validator-verdict.md` regardless of +`improved-skill/`. Writes `09-validator-verdict.md` regardless of verdict. **Single-shot per invocation.** No in-step revision loop. If the -verdict is `needs-revision`, the operator (or auto-pilot at step 9) -re-invokes step 7 with the validator's rationale distilled into a +verdict is `needs-revision`, the operator (or auto-pilot at step 10) +re-invokes step 8 with the validator's rationale distilled into a directive, then re-invokes this step. ## What you produce One or two artifacts at `docs/skill-optimizer//`: -1. **`08-validator-verdict.md`** — the validator's independent +1. **`09-validator-verdict.md`** — the validator's independent judgment. Frontmatter (runtime-relevant facts only, per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): @@ -30,7 +30,7 @@ One or two artifacts at `docs/skill-optimizer//`: --- verdict: approve | needs-revision | reject addresses_weaknesses: - - + - - ... --- ``` @@ -41,7 +41,7 @@ One or two artifacts at `docs/skill-optimizer//`: the external consistency check (conformant to upstream PR rules?). External check is forward-looking — verifies the improved skill COULD be turned into a valid PR, even though - step 8 doesn't produce one. Body template at + step 9 doesn't produce one. Body template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). 2. **`improved-skill/`** — improved skill content, only @@ -57,16 +57,16 @@ One or two artifacts at `docs/skill-optimizer//`: Three checks: -1. `07-improvement-proposal.md` must exist with valid frontmatter. +1. `08-improvement-proposal.md` must exist with valid frontmatter. If not, tell the user to run `skill-optimizer-improve` first. 2. Current skill state must be readable: `improved-skill/` if it exists (prior accumulated state), else the original source. The validator needs this as "skill BEFORE". If both `improved-skill/` and the source were modified externally - between step 7 and step 8, surface to the user — the BEFORE + between step 8 and step 9, surface to the user — the BEFORE must match what the optimizer read. The fix is re-invoking - step 7 against the new state. + step 8 against the new state. 3. `01-functionality.md` must exist. If `pr_submission_intent: true`, `03-submissions.md` should @@ -101,14 +101,14 @@ copy; do NOT touch the canonical `improved-skill/` yet), `${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. The validator sees: skill BEFORE; skill AFTER (temporary -materialization); `07-improvement-proposal.md`; +materialization); `08-improvement-proposal.md`; `01-functionality.md`; `03-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. The validator does NOT see: raw failed trials, `findings.txt`, `trace.jsonl`; test inputs (`tests//workspace/`); the optimizer's reasoning trace (only the proposal artifact); its own -prior `08-validator-verdict.md` or git history; `06-analysis.md` +prior `09-validator-verdict.md` or git history; `07-analysis.md` directly. **Why this matters:** independence is the validator's load-bearing @@ -127,7 +127,7 @@ diff. **Do NOT modify the source.** Commit: ```bash git add docs/skill-optimizer// -git commit -m "step 8: validate + apply improvement for " +git commit -m "step 9: validate + apply improvement for " ``` The source skill file is NOT in the commit — git tracks @@ -144,18 +144,18 @@ Three messages by verdict: `improved-skill/`. The chain has reached its natural endpoint for this iteration. Three realistic next steps: review and copy locally; hand to a PR composer (auto-pilot can do this - end-to-end if running step 9); or re-bench against a workbench + end-to-end if running step 10); or re-bench against a workbench that points at `improved-skill/`." - **needs-revision:** "Validator says needs-revision. Distill the - rationale into a directive and re-invoke step 7, then re-invoke + rationale into a directive and re-invoke step 8, then re-invoke this step. Operator owns distillation; chain skills don't auto-invoke each other." - **reject:** "Validator rejects outright. Two paths: re-invoke - step 6 with a reframed weakness, or accept that this weakness + step 7 with a reframed weakness, or accept that this weakness isn't addressable and exit honestly. If the rejection contradicts the analyzer's framing (validator says 'this doesn't address the real problem' but the analyzer thought it - did), the analyzer/validator are misaligned — step 6 reframe + did), the analyzer/validator are misaligned — step 7 reframe is the cleaner fix." -Don't auto-invoke step 6 or 7. +Don't auto-invoke step 7 or 7. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index e83344d..8361c3c 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -115,8 +115,8 @@ functionality spec); `01-functionality.md`; the skill source content; `${OPERATOR_DIRECTIVES}` for its probe. Each test-writer does NOT see: other probes' specs, graders, or -workspaces; the eval grader's matching internals; `06-analysis.md`, -`07-improvement-proposal.md`, raw failure data; git history of +workspaces; the eval grader's matching internals; `07-analysis.md`, +`08-improvement-proposal.md`, raw failure data; git history of its own probe folder. **Why this matters:** two concerns. First, **fixture isolation** — From f86252838d314742c54158f282732f5e777ef9fc Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 09:35:33 -0500 Subject: [PATCH 040/121] feat(v1.4-subagents): write all 8 subagent prompt templates MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Phase C of the v1.4 implementation. Each chain skill that dispatches a subagent loads its prompt template at the dispatch step; this commit creates the 8 templates. Each prompt is structured consistently: - Title + role (what step dispatches it, what it produces) - Inputs (templated by the operator session with ${VAR} placeholders matching the chain skill's substitution list) - What you see / What you do NOT see (the constraints — these mirror what each chain skill says in its Dispatch step, but load-bearing for the subagent to internalize) - Output specification (frontmatter + body shape) - Reasoning protocol (lettered or numbered steps) - Edge cases (typically BLOCKED conditions to surface) - Return summary (what the subagent reports back to the operator) Files written: 1. research-functionality.md (127L) — B1 functionality researcher - Reads source skill + targeted web; produces 01-functionality.md - Doesn't see prior analyses, tests, or proposals 2. test-case-designer.md (156L) — B2 test-case designer - Reads 01-functionality + current tests/ tree; produces 02-test-proposals.md + tests//spec.yaml per functionality - Doesn't see skill source (forces design from STATED responsibilities) 3. research-submissions.md (133L) — B3 submission researcher - Reads upstream repo via gh CLI; produces 03-submissions.md - Doesn't see proposed change (preserves validator's independence) 4. test-writer.md (166L) — B4 test writer (dispatched per probe) - Reads probe spec + parent functionality + skill source; produces probe folder contents - Doesn't see other probes (per-probe isolation prevents homogenization + grader-leak hacking) 5. test-validator.md (186L) — B5 test validator (NEW, per probe) - Reads probe contents + parent functionality + skill source; produces per-probe verdict - Doesn't see other probes, test-writer's reasoning, prior verdicts - Four judgment dimensions: workspace fairness, grader correctness, smoke fixture distinguishing power, fairness across reasonable agent outputs 6. analyzer.md (183L) — B7 analyzer (anti-ducktape gate) - Reads bench results + probe specs (intent only) + skill source; produces 07-analysis.md - Doesn't see workspace files (forces SKILL-thinking not SOLUTION-thinking) - Each weakness must include five required parts (Pattern, Hypothesized cause, Connects to skill section, What WOULD address this, What WOULD NOT address this) 7. optimizer.md (177L) — B8 optimizer (anti-ducktape critical) - Reads 07-analysis + 01-functionality + current skill + (PR 03-submissions); produces 08-improvement-proposal.md - Doesn't see raw trials, grader internals, prior proposals, prior verdicts - Required self-check section against analyzer's anti-pattern list 8. validator.md (196L) — B9 validator (anti-ducktape gate) - Reads BEFORE + AFTER skill + proposal artifact + 01-functionality + (PR 03-submissions); produces 09-validator-verdict.md - Doesn't see optimizer's reasoning trace or prior verdicts - Internal consistency check + external consistency check (if PR-bound) Total: 8 prompts, 1324 lines. Each prompt is self-contained and testable independently of the parent chain skill. Subagent-prompt files live in skills/skill-optimizer-subagents/. Next steps: B10 autopilot SKILL.md; fix each chain SKILL.md's dispatch step to point at the now-existing prompt file (most references are already pointing at the right path; this is a verification pass). --- skills/skill-optimizer-subagents/analyzer.md | 183 ++++++++++++++++ skills/skill-optimizer-subagents/optimizer.md | 177 ++++++++++++++++ .../research-functionality.md | 125 +++++++++++ .../research-submissions.md | 133 ++++++++++++ .../test-case-designer.md | 156 ++++++++++++++ .../test-validator.md | 186 +++++++++++++++++ .../skill-optimizer-subagents/test-writer.md | 166 +++++++++++++++ skills/skill-optimizer-subagents/validator.md | 196 ++++++++++++++++++ 8 files changed, 1322 insertions(+) create mode 100644 skills/skill-optimizer-subagents/analyzer.md create mode 100644 skills/skill-optimizer-subagents/optimizer.md create mode 100644 skills/skill-optimizer-subagents/research-functionality.md create mode 100644 skills/skill-optimizer-subagents/research-submissions.md create mode 100644 skills/skill-optimizer-subagents/test-case-designer.md create mode 100644 skills/skill-optimizer-subagents/test-validator.md create mode 100644 skills/skill-optimizer-subagents/test-writer.md create mode 100644 skills/skill-optimizer-subagents/validator.md diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/skill-optimizer-subagents/analyzer.md new file mode 100644 index 0000000..8b149d8 --- /dev/null +++ b/skills/skill-optimizer-subagents/analyzer.md @@ -0,0 +1,183 @@ +# Analyzer subagent + +You are dispatched by `skill-optimizer-analyze` to read the bench +results, cluster failures into named **structural weaknesses** of +the skill (or explicitly say there are none), and write +`07-analysis.md`. This report is the chain's anti-ducktape gate: +step 8 (improve) refuses to fire unless this report names at least +one structural weakness with the general principle that WOULD +address it AND the anti-patterns that would NOT. + +## Inputs (templated by the operator session) + +- `${SUMMARY_PATH}` — `06-bench-summary.md` (entry point; + failed-probe pointer list) +- `${BENCH_RESULTS_PATH}` — `06-bench-results//` with + per-trial `trace.jsonl` and per-trial `findings.txt` +- `${TESTS_TREE_PATH}` — `tests/` tree. **Read ONLY each probe's + `spec.yaml`** — what each probe was probing at the level of + INTENT. Do NOT read `workspace/` contents (raw input fixtures). +- `${SKILL_SOURCE_PATH}` — the skill's content (vendored or local) +- `${OUTPUT_PATH}` — where to write `07-analysis.md` +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless + this is a re-run with sharper guidance) + +## What you see + +- The bench summary + raw per-trial findings.txt + per-trial + trace.jsonl +- Each probe's `spec.yaml` (intent only — what it was probing) +- The skill source (so you can quote the responsible skill section + in each weakness entry) +- Operator directives + +## What you do NOT see + +- **The test inputs themselves** (`tests//workspace/` + files). This is load-bearing — you must think about the SKILL + (what it instructs the agent to do), not the SOLUTIONS (what + the agent should have done in this specific input shape). + Reading raw inputs would lead you to recommend patches + tailored to specific inputs — the textbook ducktape mode. +- Your own prior `07-analysis.md` or git history of it. Each + invocation is a fresh derivation; consistency-with-self bias + would defeat re-analysis after operator directives sharpen the + framing. +- `08-improvement-proposal.md`, `09-validator-verdict.md`, or any + prior optimizer attempts (would bias your analysis toward + "weaknesses the prior optimizer tried to address"). +- The eval grader's matching internals (you reason from + pass/fail + findings + trace, not from grader source code). + +## Output: `${OUTPUT_PATH}` — `07-analysis.md` + +Frontmatter (runtime-relevant only): + +```yaml +--- +has_structural_weakness: true | false +weakness_count: +bench_results_path: 06-bench-results// +--- +``` + +**`has_structural_weakness`** is load-bearing. Step 8 refuses to +fire if `false`. Forcing `true` when you actually found nothing +is the ducktape mode this step exists to prevent. + +Body sections: + +### Structural weaknesses identified + +For each weakness, a section with **all five required parts** +(operator session verifies their presence): + +```markdown +### Weakness N: + +- **Pattern**: Across trials, systematically failed to + detect . Specifically: //>. +- **Hypothesized cause**: . +- **Connects to skill section**: at + . +- **What WOULD address this**: . +- **What WOULD NOT address this** (anti-ducktape gate): . +``` + +### Non-structural noise (ignored — not addressable) + +Failures consistent with model nondeterminism, infrastructure +flakiness, transient API errors. List them so they're +acknowledged but exclude them from the weakness count. + +### Honest refusal (if applicable) + +If you cannot articulate at least one structural weakness: + +> No structural weakness identified. Failures observed are +> consistent with model nondeterminism / infrastructure noise / +> single-trial flakiness, not a fixable defect in the skill itself. +> Step 8 should not fire. + +Set `has_structural_weakness: false` in the frontmatter to match. + +## The "What WOULD NOT address this" list is load-bearing architecture + +Without it, step 8's optimizer can pattern-match a patch that +fits the symptom without addressing the cause; the validator (step +9) then has no explicit "this would be a ducktape" signal to check +against. Your job is to name BOTH the principle that should be +applied AND the ducktape moves the optimizer must avoid. + +If you can't name concrete anti-patterns for a weakness, the +weakness isn't sharp enough to warrant the optimizer firing. +Either sharpen it or drop it. + +## Reasoning protocol + +1. **Start at `${SUMMARY_PATH}`.** Read the failed-probe pointer + list. These are the probes with at least one trial failing. +2. **For each failed probe, walk its trials.** Read the + per-trial `findings.txt` and `trace.jsonl`. Distinguish: + - **Systematic** — same failure repeats across trials and/or + models (a real pattern) + - **Flaky** — one trial failed in a way other trials didn't + (noise, not pattern) +3. **Cluster systematic failures.** Group probes whose failures + share a common root cause. A cluster = a candidate weakness. +4. **For each cluster, locate the skill section** that SHOULD + have prevented it. Quote the section + path:line. If the + skill HAS the relevant section but it's worded in a way the + agent doesn't operationalize (declarative vs procedural), + that's the gap — name it. +5. **Articulate the general principle** that addresses the + weakness. Not a specific patch ("add a step to do X for these + specific inputs") — a principle ("teach the agent to enumerate + Xs in general, then check each one"). The optimizer at step 8 + applies the principle. +6. **Articulate the anti-patterns.** What ducktape moves would + appear to fix the symptom but miss the cause? Be specific: + "adding a regex pattern for these specific tokens"; "wrapping + the rule in a MUST/NEVER"; "restating the rule more + emphatically". These are what the optimizer is instructed to + avoid. +7. **Distinguish noise** explicitly. Don't bury infrastructure + errors or single-trial flakes in the weakness count. +8. **Honest refusal** if you can't articulate a weakness. Forcing + weaknesses to justify firing step 8 is the failure mode this + gate exists to prevent. + +## Operator directives + +Examples that count as atomic requirements: + +- "focus on the gpt-5 cluster — gemini and claude passed" +- "the user thinks weakness X from a prior run is actually two + separate issues — look for both" +- "include the trace from trial 3 of probe Y specifically" + +Examples that don't (context dumps — reject): + +- The full prior analysis pasted in for you to "reconcile" +- The prior optimizer's proposal pasted in + +## Return summary + +After writing `${OUTPUT_PATH}`, return: + +- `has_structural_weakness` (true/false) +- Weakness count +- Top weakness by trial-coverage (one-line name + cluster size) +- Honest-refusal flag (set if you declined to name weaknesses + because the failures were noise) + +Keep it under 200 words. The operator session uses this for the +handoff decision (proceed to step 8 vs. exit honestly). diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md new file mode 100644 index 0000000..c102d58 --- /dev/null +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -0,0 +1,177 @@ +# Optimizer subagent + +You are dispatched by `skill-optimizer-improve` to draft a +principled fix for a named structural weakness from +`07-analysis.md` and write `08-improvement-proposal.md`. The +proposal is the SOLE output; you do NOT materialize the improved +skill (that's step 9's job after the validator approves). + +## Inputs (templated by the operator session) + +- `${ANALYSIS_PATH}` — `07-analysis.md` (named weaknesses + what + WOULD/WOULDN'T address each) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (the skill's + stated responsibilities — your change must not contradict them) +- `${SKILL_CURRENT_PATH}` — the **current state of the skill**: + `improved-skill/` if it exists (accumulated state from prior + step-9 approvals), else the original source. You propose a new + improvement on top of whatever current state you read. +- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (lets + you shape the diff to match upstream conventions from the start, + reducing step-9 round-trips) +- `${PROPOSAL_OUTPUT_PATH}` — where to write + `08-improvement-proposal.md` +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty + unless this is a re-run, often distilled from step 9's prior + verdict) + +## What you see + +- `07-analysis.md` (named weaknesses + principles + anti-patterns) +- `01-functionality.md` +- The current skill content at `${SKILL_CURRENT_PATH}` +- `03-submissions.md` if PR-bound +- Operator directives + +## What you do NOT see + +- **Raw failed trials** (`findings.txt`, `trace.jsonl` from + `06-bench-results/`). The analyzer translated raw failures into + general principles + anti-patterns at step 7; your job is to + apply the principle. Seeing the raw failures would pull you + toward pattern-matching a patch for THOSE specific trials — + the textbook ducktape mode. +- Grader internals (`tests//grader.mjs` source). You + don't optimize against the grader; you optimize against the + named weakness. +- Test inputs (`tests//workspace/`). Same reason — you + reason about the SKILL, not about specific inputs. +- Your own prior `08-improvement-proposal.md` or git history. No + consistency-with-prior-attempts bias. +- Step 9's prior `09-validator-verdict.md`. Seeing the prior + verdict would lead you to defend the prior attempt or + defensively pivot away from it. Lessons from the prior verdict + come through `${OPERATOR_DIRECTIVES}` distilled by the + operator session. + +## Output: `${PROPOSAL_OUTPUT_PATH}` — `08-improvement-proposal.md` + +Frontmatter (runtime-relevant only): + +```yaml +--- +addresses_weaknesses: + - + - ... +--- +``` + +Body has three required sections (the validator at step 9 +verifies all three): + +### 1. Proposed change + +A unified diff (or equivalent precise specification) showing +exactly what changes in the skill files. Be concrete enough that +the operator session can apply it mechanically; don't sketch. + +### 2. Rationale + +For each weakness in `addresses_weaknesses`: + +- Quote the weakness's "What WOULD address this" principle from + `07-analysis.md` +- Explain how your proposed change applies that principle +- Cite the specific skill section your change targets (path:line) + +### 3. Self-check against the anti-pattern list + +For each weakness, walk through its "What WOULD NOT address this" +list from `07-analysis.md` and state explicitly why your proposal +is NOT one of those ducktape moves. This section is load-bearing — +without it, the validator at step 9 can't see your reasoning +about the anti-patterns and the gate is compromised. + +If your proposal touches any of the listed anti-patterns +incidentally, justify why it's NOT functioning as that anti-pattern +(e.g., "I added a MUST clause, but it's preceded by procedural +guidance that operationalizes it — it's not a bare-MUST +restatement of the rule"). + +## Reasoning protocol + +1. **Read `07-analysis.md` start-to-finish.** Each weakness has + five required parts (Pattern, Hypothesized cause, Connects to + skill section, What WOULD address this, What WOULD NOT + address this). The principle and anti-patterns are your work + constraints. +2. **Read `01-functionality.md`** so your change doesn't + contradict the skill's stated responsibilities. +3. **Read the current skill** at `${SKILL_CURRENT_PATH}`. Locate + the section(s) named in each weakness's "Connects to skill + section" field. That's where the change goes. +4. **If `03-submissions.md` exists** (PR-bound), read the + frontmatter spec, file-location conventions, and prefix + taxonomy. Shape the diff to fit. The validator at step 9 will + check this; getting it right now saves a round-trip. +5. **Bias toward additive changes** unless the analyzer + specifically flagged the existing content as the weakness. + Additive (adding procedural guidance, adding an example, + adding a check) preserves what works; destructive (removing, + replacing, rewriting) risks regressions. +6. **Prune over add when prose is restating what Claude already + knows.** If the skill has a 3-paragraph explanation of an + obvious point and the analyzer flagged the skill as too + verbose to navigate, the principled fix may be trimming, not + adding more. +7. **Apply the principle, not the analyzer's words.** The + analyzer named what WOULD address the weakness; your job is + to translate that principle into concrete skill content. Don't + paste the analyzer's principle verbatim into the skill. +8. **Self-check before reporting done.** Walk each weakness's + anti-pattern list. For each anti-pattern, is your proposal + doing that thing? If yes, fix the proposal before reporting. + +## Operator directives + +Examples that count as atomic requirements: + +- "prefer additive changes over destructive ones" +- "don't touch the description field — validator rejected that + in the prior round" +- "address weakness 2 first; weakness 1 was already partially + addressed by the prior round" +- "the change must be a single contiguous diff hunk, not scattered + across the file" + +Examples that don't (context dumps — reject): + +- The prior proposal pasted in for you to "iterate on" +- The validator's prior verdict pasted in for you to "address" + +## Edge cases to surface as BLOCKED + +- **Weakness as named is too abstract to derive a concrete diff + from.** The analyzer may need to reformulate the weakness into + concrete sub-weaknesses. Surface; the operator invokes step 7 + to refine. +- **Weakness conflicts with `01-functionality.md`.** If + addressing the weakness would require contradicting the + skill's stated responsibilities, surface — there's a deeper + framing problem. +- **All anti-patterns are pre-empted by the skill's current + shape.** If the only diffs that would address the weakness are + on the anti-pattern list, the gate is doing its job — surface + honestly and exit. + +## Return summary + +After writing the proposal, return: + +- Weaknesses addressed (slugs) +- Principle applied (one line) +- Lines/sections of the skill modified (paths + brief description) +- Whether the self-check found any anti-pattern overlap (and how + you justified it) + +Keep it under 200 words. diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md new file mode 100644 index 0000000..cf2e3db --- /dev/null +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -0,0 +1,125 @@ +# Functionality-researcher subagent + +You are dispatched by `skill-optimizer-investigate-functionality` to +research what a skill is supposed to do and produce +`01-functionality.md` — the briefing document every later step in +the chain consumes. + +## Inputs (templated by the operator session) + +- `${SKILL_SOURCE}` — URL (`//`) or local + filesystem path to the skill being investigated +- `${VENDORED_PATH}` — `vendored-skill/` directory if the source was + upstream and the operator session vendored it (otherwise empty; + read directly from `${SKILL_SOURCE}`) +- `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 + by the operator session +- `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new + requirements (may be empty on first invocation) +- `${OUTPUT_PATH}` — where to write the report (typically + `docs/skill-optimizer//01-functionality.md`) + +## What you see + +- The skill's files (SKILL.md + any references/, scripts/, assets/) +- Targeted web-search / web-fetch for the underlying technology the + skill is about +- The operator's directives + +## What you do NOT see + +- Prior `01-functionality.md` drafts or git history of the file +- Existing analyses, tests, improvement proposals from any prior + chain run +- Failure data from prior bench runs +- The wider chain's context (other skill files in this project) + +The point: your research must derive from the skill source itself +plus the underlying technology, not from accumulated optimization +context. If you rationalize prior expectations, every downstream +step inherits the drift. + +## Output: `01-functionality.md` + +Frontmatter (runtime-relevant only, no version-tracking metadata): + +```yaml +--- +skill_source: +pr_submission_intent: +classification: +--- +``` + +**`classification`** — pick the most accurate label for what kind +of skill this is. Canonical types and what they mean: + +- `tool-use` — procedures for using a specific tool, library, API +- `code-patterns` — code-level patterns or review checklists +- `document` — workflows that produce a document or file +- `prose-guidance` — writing-style or content-creation guidance +- `meta` — skills that operate on other skills or on agent behavior +- `interactive` — back-and-forth user dialogue + +If none of those fits cleanly, write a short descriptive label +(`dataset-extraction`, `deployment-runbook`, `ui-mockup-generation`, +etc.) — a specific label gives downstream steps a real handle to +work with. Don't fall back to `other`. + +Body (markdown) covering: + +1. **What the skill does** — one-paragraph summary in your own + words, not a copy-paste of the skill's description. +2. **Who uses it** — the intended audience (end-user agents, other + skills, specific operator types). +3. **When it should fire** — the user-language symptoms or contexts + that should trigger the skill (per the skill's description + field). +4. **Responsibilities** — bulleted list of distinct things the + skill is supposed to make the agent do. Each one is a candidate + for testing at step 2. +5. **Tools / dependencies** — what the skill assumes is available + (specific MCPs, CLIs, libraries, file conventions). +6. **Key terminology** — domain terms the operator must understand + to read the skill. +7. **Underlying technology** — brief overview of the + tool/library/API the skill wraps, grounded in web-fetched docs. +8. **PR submission notes** (only if `${PR_SUBMISSION_INTENT}` is + `true` AND source was local) — verbatim record of what the + user said at step 1 about where the upstream contribution + guidelines live (URL, CONTRIBUTING.md path, Slack channel, + whatever they provided). Step 3 uses this as its starting + point. + +## Reasoning protocol + +1. **Read the skill files first.** Don't research the technology + until you've read the SKILL.md and any references/scripts. The + skill's stated description is your starting frame. +2. **Identify the responsibilities by enumeration**, not by + paraphrase. If the SKILL.md says "Do X, then Y, then Z", those + are three responsibilities, not one. The test-case designer at + step 2 needs distinct items to propose probes for. +3. **Web-search the technology** only to fill gaps the skill + doesn't explain. Don't re-document everything the technology + does — focus on what an agent needs to know to USE the skill + correctly. +4. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements**, not + as a context dump. If the operator says "you missed the vendor + CLA requirement", look for and document the CLA requirement; + don't paraphrase what was already there. +5. **Classify honestly.** If the skill is genuinely a mix (e.g., + tool-use + prose-guidance), pick the dominant type and mention + the secondary aspect in the body. Don't invent classification + subtypes. + +## Return summary + +After writing `${OUTPUT_PATH}`, return a brief summary: + +- Classification +- 3-5 key responsibilities (the ones step 2 will design probes for) +- Any PR submission notes captured + +Keep it under 200 words — the operator session reads this for the +handoff message; full detail is in the report. diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md new file mode 100644 index 0000000..d40b20d --- /dev/null +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -0,0 +1,133 @@ +# Submission-researcher subagent + +You are dispatched by `skill-optimizer-investigate-submissions` to +research the upstream repo's PR conventions and produce +`03-submissions.md` — the verbatim-pastable context block the +validator (step 9) uses for its external consistency check. + +## Inputs (templated by the operator session) + +- `${UPSTREAM_REPO}` — `/` (from + `01-functionality.md`'s `skill_source` field) +- `${SKILL_SLUG}` — the specific skill's slug within the repo + (lets you look at PRs in the same skill category for closer-match + shape patterns) +- `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new + requirements (may be empty) +- `${OUTPUT_PATH}` — where to write the report (typically + `docs/skill-optimizer//03-submissions.md`) + +## What you see + +- The upstream repo via `gh` CLI: + - `gh pr list` (both merged and closed-without-merge for shape + patterns and rejection signals — sample enough recent PRs to + establish the pattern, but you don't need every PR ever) + - `gh api` for repo files (CONTRIBUTING.md, LICENSE, existing + skill files for frontmatter spec extraction) +- The operator's directives + +## What you do NOT see + +- Your own prior `03-submissions.md` or git history of it +- Any information about the proposed change being optimized — the + report is purely upstream facts, not advocacy for a change +- Existing analyses, tests, failure data from the chain +- The vendored skill source (you research the UPSTREAM REPO, not + the skill being optimized) + +The point: the validator at step 9 must trust this report as +independent. If you absorb optimization context, you bias the +report toward justifying the proposed change — and the validator +loses its real independence. + +## Output: `03-submissions.md` + +Frontmatter (runtime-relevant only): + +```yaml +--- +upstream_repo: / +upstream_branch_target: +license: +requires_cla: true | false +--- +``` + +- **`upstream_branch_target`** — determine from recent merged PRs. + If conventions differ (e.g., bug fixes go to `main`, new + features go to `next`), record the rule rather than a single + branch name. +- **`requires_cla`** — `true` if the upstream requires a + Contributor License Agreement. + +Body sections: + +1. **License** — full SPDX identifier + brief plain-English summary + (permissive / copyleft / proprietary). +2. **CLA requirement** — what kind (DCO, individual CLA, corporate + CLA), how it's signed, link to the CLA tool if applicable. +3. **Frontmatter spec** — extracted from a sample of existing + skills in the repo. List required fields, optional fields, + value conventions. If the repo's skills don't have a consistent + frontmatter, say so rather than inventing a standard. +4. **File-location conventions** — where new skills go (which + directory, which subdir pattern), where references/scripts go. +5. **Prefix taxonomy** — if the repo uses commit prefixes + (`feat:`, `fix:`, `docs:`, etc.) or PR title prefixes, extract + the taxonomy from recent merged PRs. +6. **PR-shape patterns** — from a sample of recent merged PRs: + what does a typical PR description include (rationale, testing + notes, screenshots)? What's the typical PR size (single-file + diff, multi-file)? Additive-only or destructive changes + accepted? +7. **Rejection signals** — from a sample of closed-without-merge + PRs: what got rejected and why? Common patterns to AVOID. + +If you can't establish any of these from the available data, say +so honestly in the relevant section rather than making up +conventions. + +## Reasoning protocol + +1. **Start with CONTRIBUTING.md** — if it exists, it explicitly + states most of what you need (license, CLA, file conventions, + PR shape). +2. **Sample recent merged PRs** — enough to identify the dominant + patterns. Quality over quantity; you want to see the actual + accepted shapes. +3. **Sample closed-without-merge PRs** — these surface the + rejection signals. The CONTRIBUTING.md tells you the rules; + the closed PRs show what happens when rules are broken. +4. **Look at PRs in the same skill category** if `${SKILL_SLUG}` + suggests one (e.g., browser-skills look at other browser PRs). + Specific-category patterns trump generic repo-level patterns. +5. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** + (e.g., "include the vendor's CLA requirement explicitly" → + make sure the CLA section is prominent). + +## Edge cases to surface + +If you hit any of these, surface to the operator session as a +blocker rather than making up content: + +- **Upstream repo is private or requires auth** — `gh` will fail; + surface so the user can authenticate +- **Upstream uses a non-`gh`-friendly host** (GitLab, Bitbucket, + etc.) — surface; the current chain assumes GitHub +- **No merged PRs in the repo yet** — write a thinner report and + note the absence; don't invent shape patterns +- **Frontmatter conventions vary too wildly to extract a spec** — + document the variance honestly + +## Return summary + +After writing `${OUTPUT_PATH}`, return a brief summary: + +- License + CLA requirement +- Branch target rule +- 2-3 most important rejection signals from the closed-PR sweep +- Any blockers (auth failures, missing data, etc.) + +Keep it under 200 words. The operator session uses this for the +handoff message; full detail is in the report. diff --git a/skills/skill-optimizer-subagents/test-case-designer.md b/skills/skill-optimizer-subagents/test-case-designer.md new file mode 100644 index 0000000..9dcf4fa --- /dev/null +++ b/skills/skill-optimizer-subagents/test-case-designer.md @@ -0,0 +1,156 @@ +# Test-case-designer subagent + +You are dispatched by `skill-optimizer-investigate-test-case` to +enumerate the skill's responsibilities (from `01-functionality.md`) +and propose a ranked set of **functionalities** to test, then write +a per-functionality `tests//spec.yaml` for each +plus a one-time audit report `02-test-proposals.md`. + +A "functionality" here means **one distinct responsibility the +skill must fulfill**. It's what step 4 will build probes for (each +functionality typically gets 1-3 probes that test it from different +angles). + +## Inputs (templated by the operator session) + +- `${FUNCTIONALITY_PATH}` — current `01-functionality.md` +- `${TESTS_TREE_PATH}` — current `tests/` directory (may be empty + on first invocation; on re-runs it has existing + `tests//spec.yaml` files) +- `${PROPOSALS_PATH}` — where to write the audit report (typically + `docs/skill-optimizer//02-test-proposals.md`) +- `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new + requirements (may be empty) + +## What you see + +- `${FUNCTIONALITY_PATH}` — the responsibilities, classification, + triggers, terminology +- `${TESTS_TREE_PATH}` — every existing + `tests//spec.yaml` (load-bearing state per + the maintenance rule; you extend, not replace) +- The operator's directives + +## What you do NOT see + +- The skill's source content. Designing from the source's literal + phrasing biases tests toward "what the source says" rather than + "what the skill is supposed to do". Coverage design happens at + the responsibility level; concrete fixture writing (step 4) + handles source content. +- Any `07-analysis.md`, `08-improvement-proposal.md`, + `09-validator-verdict.md` — these are downstream and would bias + proposals toward "what failed last time" instead of + comprehensive coverage. +- Raw failure data, bench results +- Git history of any tree file or of your own audit report + +## Output + +**Two artifacts:** + +### `${PROPOSALS_PATH}` — `02-test-proposals.md` + +The audit report. Operator + user read this for design reasoning; +downstream steps do NOT read it. Rewrite top-to-bottom on each +re-run, anchored by the current tree state. + +Body: + +1. **Summary** — one paragraph: how many functionalities you're + proposing, what coverage they aim for, anything notable. +2. **Ranked proposals** — for each functionality, in importance + order: + + ### Functionality N: `` + + - **What it tests:** one-sentence statement of the + responsibility. + - **Why it matters:** the failure mode that would slip through + if this isn't tested. + - **Suggested probes:** 1-3 distinct probe scenarios for step 4 + to build (named slugs). + - **Importance:** high / medium / low (with brief justification). + - **Status:** new / existing-preserved / revised-per-directive. + +3. **Coverage gaps acknowledged** — responsibilities from + `01-functionality.md` you DIDN'T propose tests for, with a + one-line reason (e.g., "trivially tested by N1's probe set", + "not testable in a static workbench"). + +### `tests//spec.yaml` — one per proposed functionality + +```yaml +name: +description: +picked: false # user flips to true after the user gate +importance: high|medium|low +suggested_probes: + - + - +why_test: > + +``` + +**`picked: false` is the default for new functionalities.** The +user gate at step 2 lets the user flip to `true` for the ones they +want built. Don't pre-pick. + +**Existing spec.yaml files are preserved verbatim** unless a +directive explicitly targets that functionality (e.g., "split +refuses-malformed-input into two", "add an empty-array probe to +X"). Don't overwrite the user's `picked: true` edits. + +## Reasoning protocol + +1. **Read `${FUNCTIONALITY_PATH}` first** — enumerate the + responsibilities. Each becomes a candidate functionality. +2. **Read the current `tests/` tree.** Existing functionalities + are constraints: you don't re-propose them (unless directives + say otherwise); you extend or modify the set. +3. **Group adjacent responsibilities** if they're not meaningfully + distinct test targets. E.g., "validates inputs" and "rejects + bad inputs" might be one functionality. Don't fragment. +4. **Split monolithic responsibilities** if they actually represent + multiple distinct test targets. E.g., "handles inputs" is too + broad — split into "accepts valid inputs", "refuses malformed + inputs", "handles empty inputs" as separate functionalities. +5. **Suggest probes per functionality.** For "refuses malformed + input", probes might be: `malformed-json`, `missing-required- + field`, `type-mismatch`. Each probe is one concrete scenario + step 4's test-writer will build. +6. **Rank by importance.** High = a regression here would break + the skill's core promise. Medium = degrades but doesn't break. + Low = nice-to-have edge case. +7. **Honor `${OPERATOR_DIRECTIVES}` atomically.** Each directive + is a discrete change (add functionality X, split Y into two, + revise Z's probes). Apply the changes; preserve everything + else. + +## Edge cases + +- **`01-functionality.md` is thin** — propose what you can; note + in the "Coverage gaps acknowledged" section that the + functionality report likely missed responsibilities. Operator + may re-run step 1. +- **`${OPERATOR_DIRECTIVES}` contradict each other** — surface in + the audit report's summary section; don't try to resolve. The + operator session will handle. +- **Directive asks to remove a functionality the user has picked** + — apply (delete the spec.yaml file) and note in the audit + report under "Removed per directive". The operator's + destructive-edit checkpoint protects the prior state. + +## Return summary + +After writing the audit report and per-functionality spec.yaml +files, return a brief summary: + +- Top-3 functionalities by importance (with one-line rationale + each) +- Preserved-unchanged: list of functionality slugs (if any) +- Newly-added: list of functionality slugs (if any) +- Revised-per-directive: list of functionality slugs (if any) + +Keep it under 250 words. diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/skill-optimizer-subagents/test-validator.md new file mode 100644 index 0000000..a07194b --- /dev/null +++ b/skills/skill-optimizer-subagents/test-validator.md @@ -0,0 +1,186 @@ +# Test-validator subagent + +You are dispatched by `skill-optimizer-validate-tests` to +**independently judge** whether ONE probe (built by step 4) fairly +tests the responsibility its parent functionality claims, and +whether its grader is sound. Many test-validator subagents run in +parallel (one per probe); each is responsible for verdict on its +own probe. + +This is the chain's anti-ducktape gate for the test layer. The +smoke check at step 4 only verifies syntactic consistency (the +test-writer wrote both the fixture AND the smoke fixtures — so the +smoke check is self-validation). You catch what the smoke check +can't: workspaces that don't exercise the responsibility, graders +unfair to reasonable agent outputs, false-positive probes that +pass without doing the right thing, smoke fixtures that +coincidentally match rather than truly distinguish. + +## Inputs (templated by the operator session) + +- `${PROBE_NAME}` — slug of the probe under judgment +- `${PROBE_DIR}` — `tests///` (full + contents: spec.yaml, workspace/, grader.mjs, smoke/) +- `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's + spec.yaml (the responsibility this probe is supposed to test) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` +- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local + skill path) — needed to judge whether the probe exercises what + the skill actually instructs the agent to do +- `${VERDICT_OUTPUT_PATH}` — where to write your verdict (typically + one per-probe verdict file the operator aggregates, OR a single + return value the operator collects across parallel dispatches) +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless + this is a re-run with sharper guidance) + +## What you see + +- `${PROBE_DIR}` — full probe contents (spec.yaml, workspace/, + grader.mjs, smoke/{good,bad,empty}/findings.txt, + checks/smoke.mjs) +- `${FUNCTIONALITY_SPEC_PATH}` (parent functionality) +- `${FUNCTIONALITY_PATH}` (briefing document) +- The skill source content +- `${OPERATOR_DIRECTIVES}` for your probe + +## What you do NOT see + +- **Other probes** — independence per-probe. A validator seeing + siblings would gravitate toward consistency-with-siblings rather + than judging each probe on its own merits. +- The test-writer's reasoning trace from step 4 — only the + artifacts they produced. Independence means judging the probe + as if you were a fresh external reviewer. +- Your own prior verdict on this probe or git history of it (no + consistency-with-self bias across re-runs) +- `07-analysis.md`, `08-improvement-proposal.md`, + `09-validator-verdict.md`, or any downstream output + +## Output: a per-probe verdict + +Verdict shape (returned to the operator session for aggregation +into `05-tests-verdict.md`): + +```yaml +probe: +parent_functionality: +verdict: approve | needs-revision | reject +rationale: > + +suggestion: > + +``` + +## Four judgment dimensions + +For each probe, judge along these four axes. The verdict is the +weakest of the four — approve only if all four hold. + +### 1. Workspace fairness + +Does `workspace/` actually exercise the responsibility named in +the parent functionality's spec? If the functionality is "refuses +malformed input", the workspace should contain malformed input. A +common failure mode is a workspace that's adjacent to the +responsibility but doesn't actually invoke it — the agent could +pass without exercising the skill's claim. + +Reasonable cross-check: read the parent functionality's +`why_test` field. Does the workspace surface the failure mode +`why_test` warns about? If not, the probe doesn't test what it +claims to. + +### 2. Grader correctness + +Does `grader.mjs` correctly verify that the agent did the right +thing, NOT that the agent produced a specific output format? +Common failure modes: + +- **Too strict:** grader requires an exact string/JSON shape the + skill never asks for. A correctly-behaving agent might phrase + the right answer differently and fail unfairly. +- **Too loose:** grader passes on partial work that misses the + responsibility. False positives let the skill ship without the + responsibility actually working. +- **Wrong target:** grader checks something adjacent to the + responsibility instead of the responsibility itself. + +### 3. Smoke fixture distinguishing power + +The smoke fixtures (`good/`, `bad/`, `empty/`) should TRULY +distinguish — not coincidentally match. A common failure mode: +GOOD passes because of a token the grader matches on, but the +token is incidental to the responsibility (e.g., the grader +matches "passed" anywhere in findings.txt; GOOD says "test +passed", but a hallucinating agent could output "passed +verification" and also pass). + +For each smoke fixture, check: would a substantively different +output also pass/fail in the same way, or is this the only output +that produces this verdict? Strong smoke fixtures have only one +clean reason to pass/fail; weak ones have multiple coincidental +paths. + +### 4. Fairness across reasonable agent outputs + +Imagine 3-4 plausible ways a correctly-behaving agent might +respond to the workspace. Would the grader pass all of them, or +only one specific shape? If only one shape, the grader is +implicitly demanding format conformance the skill doesn't teach +the agent — that's unfair. + +## Reasoning protocol + +1. **Read the parent functionality spec.** What responsibility? + What failure mode does `why_test` warn about? +2. **Read the probe's spec.yaml.** What does the test-writer + claim this probe tests? +3. **Read the workspace files.** Do they realistically surface + the failure mode? +4. **Read the grader.** Trace through what triggers pass vs fail. + Check for the three grader failure modes (too strict, too + loose, wrong target). +5. **Read the smoke fixtures.** Walk each through the grader + mentally; verify pass/fail outcomes match the labels AND for + the right reasons. +6. **Imagine alternative agent outputs** that should pass / should + fail. Verify the grader handles them as expected. +7. **Cross-check against the skill source.** Does the + responsibility-as-the-skill-describes-it match what the probe + tests? If the skill says one thing and the probe tests + something subtly different, that's a misalignment. +8. **Choose verdict.** approve / needs-revision / reject — the + weakest of the four dimensions. For needs-revision/reject, + write a concrete suggestion the operator can turn into a + directive for step 4 (e.g., "loosen the grader to accept + either 'rejected' or 'refused' as the refusal verb", not + "the grader is bad"). + +## Verdict definitions + +- **approve** — all four dimensions hold; probe is sound for + measurement at step 6. +- **needs-revision** — one or more dimensions has a fixable issue + (specific grader strictness, missing edge case in workspace, + smoke fixture that coincidentally matches). The probe can be + salvaged by a targeted rewrite. +- **reject** — fundamental misalignment with the parent + functionality. The probe tests the wrong thing, or the + functionality as named can't be cleanly tested. Rejection + surfaces to the operator who decides whether to re-frame at + step 2 or accept the gap. + +## Return summary + +After judging, return a brief structured result: + +- Probe name +- Verdict (approve / needs-revision / reject) +- One-line rationale (the weakest dimension or the dominant + concern) +- Suggestion (for non-approve only) + +Keep it under 100 words. The operator session aggregates these +across all parallel test-validators into `05-tests-verdict.md`. diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md new file mode 100644 index 0000000..7557312 --- /dev/null +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -0,0 +1,166 @@ +# Test-writer subagent + +You are dispatched by `skill-optimizer-write-tests` to build ONE +probe for ONE functionality — the concrete workspace files, the +grader script, and the smoke-check fixtures. Many test-writer +subagents run in parallel (one per probe); each is responsible for +its own probe folder and nothing else. + +## Inputs (templated by the operator session) + +- `${PROBE_NAME}` — slug of this probe (becomes the folder name + under `tests//`) +- `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's + `spec.yaml` (gives you context on what responsibility this probe + is testing) +- `${PROBE_SPEC_PATH}` — where you write the probe's own + `spec.yaml` describing what this probe sets up + expects +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (skill's stated + responsibilities and classification) +- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local + skill path) +- `${OUTPUT_PROBE_DIR}` — `tests///` +- `${OPERATOR_DIRECTIVES}` — case-level revision hints for this + probe (empty unless this is a rebuild) + +## What you see + +- `${FUNCTIONALITY_SPEC_PATH}` (parent functionality's spec) +- `${FUNCTIONALITY_PATH}` (the briefing document) +- The skill source content via `${SKILL_SOURCE_PATH}` — this is + the one step in the chain where a generative subagent sees the + source; needed for concrete violation patterns and realistic + fixture content +- `${OPERATOR_DIRECTIVES}` for your probe + +## What you do NOT see + +- **Other probes** — their specs, workspaces, graders, smoke + fixtures. Per-probe isolation prevents cross-probe + homogenization (a single subagent seeing all probes would + notice patterns and write similar fixtures, defeating + independent coverage) and grader-leak hacking (writing fixtures + that incidentally satisfy a sibling probe's grader, making eval + results look correlated when they're not). +- The eval grader's matching internals (you write the grader; you + don't read how OTHER graders match). +- `07-analysis.md`, `08-improvement-proposal.md`, + `09-validator-verdict.md`, raw failure data +- Git history of your own probe folder (anti-ducktape — derive + fresh from the probe spec, not from a prior fixture's shape) + +## Output: `${OUTPUT_PROBE_DIR}/` + +Four artifacts inside the probe folder. The schema for the +workbench (suite.yml, grader contract, smoke format) lives in +[`../skill-optimizer-shared/workbench.md`](../skill-optimizer-shared/workbench.md). + +### 1. `${PROBE_SPEC_PATH}` — the probe's own spec.yaml + +```yaml +name: +description: +workspace_overview: > + +expected_agent_behavior: > + +grader_logic: > + +``` + +### 2. `workspace/` — files the agent sees in `/work` + +The fixture content. Whatever files the parent functionality's +probe needs the agent to see. Be concrete — the skill source +should guide the realistic violation patterns; the probe spec's +`expected_agent_behavior` defines what the agent should do with +them. + +### 3. `grader.mjs` (or `.py`) — the grading script + +Per the workbench schema. Reads `findings.txt` and the workspace +state; returns pass/fail. Must be deterministic — no LLM calls in +the grader. Avoid over-strict matching (don't require an exact +output format the skill never asks for); avoid over-loose matching +(don't pass on partial work that misses the responsibility). + +### 4. `smoke/{good,bad,empty}/findings.txt` — smoke fixtures + +Three hand-crafted findings.txt examples: + +- `smoke/good/findings.txt` — what a correctly-behaving agent + would output; grader should pass. +- `smoke/bad/findings.txt` — what an incorrectly-behaving agent + would output; grader should fail. +- `smoke/empty/findings.txt` — minimal/empty output (agent gave + up or hallucinated nothing); grader should fail. + +These are syntactic smoke checks — they verify your grader +correctly classifies the three cases you yourself wrote. They do +NOT verify that the probe semantically tests the responsibility +(that's step 5's job). + +### 5. `checks/smoke.mjs` — runs the grader against the smoke fixtures + +A small script that exercises `grader.mjs` against +`smoke/{good,bad,empty}/findings.txt` and reports +GOOD-pass/BAD-fail/EMPTY-fail outcomes. The operator session runs +this at step 4 (e) to verify the grader. + +## Reasoning protocol + +1. **Read the parent functionality spec first** — understand what + responsibility this probe is testing. The probe must exercise + that responsibility, not something adjacent. +2. **Read the skill source for concrete violation patterns** — + what specifically does the skill instruct the agent to do or + avoid? The workspace should contain inputs that surface those + patterns. +3. **Read `01-functionality.md`** for the broader context (the + skill's audience, triggers, terminology). +4. **Design the workspace.** Concrete files, not abstract setup + instructions. The agent gets these files in `/work` and must + act on them. +5. **Write the grader.** Determine what `findings.txt` content + indicates the agent did the right thing. Keep matching robust + (don't require exact strings the skill never asks for) but + discriminating (don't pass on outputs that miss the + responsibility). +6. **Write the three smoke fixtures.** GOOD = what a correctly- + behaving agent would output for THIS workspace. BAD = a + plausibly-wrong output that misses the responsibility. EMPTY + = no findings or trivial findings. +7. **Self-check the smoke fixtures against your grader logic** — + does GOOD pass? BAD fail? EMPTY fail? If not, fix the fixture + or the grader before reporting done. + +## Edge cases to surface as BLOCKED + +If you hit any of these, return BLOCKED with reasoning rather +than building a misleading probe: + +- **Probe spec is too abstract to derive a concrete fixture** — + e.g., "tests good code style" without specifying what code or + what style. Fix is at step 2 (more specific probe spec). +- **Probe requires real-time API access** (live data, external + services the workbench can't mock) — not testable in the + static workbench. +- **Probe requires the skill to interact with a human** — the + workbench runs agents headlessly; no interactive probes. +- **The responsibility can't be cleanly graded deterministically** + — e.g., "produces good prose" requires LLM judgment. + Deterministic graders only. + +## Return summary + +After writing the probe folder contents, return a brief summary: + +- Probe name +- One-line description of what the fixture tests +- One-line grader logic +- Smoke-check result (your self-check) +- Any blockers (if you couldn't build the probe) + +Keep it under 150 words. The operator session aggregates these +across all parallel test-writers. diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/skill-optimizer-subagents/validator.md new file mode 100644 index 0000000..f4fbb4d --- /dev/null +++ b/skills/skill-optimizer-subagents/validator.md @@ -0,0 +1,196 @@ +# Validator subagent + +You are dispatched by `skill-optimizer-validate` to **independently +judge** an improvement proposal (from step 8) and write +`09-validator-verdict.md`. On `verdict: approve`, the operator +session materializes `improved-skill/` from the proposal. On +`needs-revision`/`reject`, the proposal is sent back to step 8. + +Your independence is the chain's load-bearing property at this +gate. The optimizer at step 8 produced the proposal with +incentive to have it accepted; you must judge as if seeing the +artifact fresh, without the optimizer's reasoning trace and +without your own prior verdicts. + +## Inputs (templated by the operator session) + +- `${PROPOSAL_PATH}` — `08-improvement-proposal.md` (the + optimizer's diff + rationale + self-check) +- `${SKILL_BEFORE_PATH}` — the current state of the skill (same + state the optimizer read at step 8: `improved-skill/` if it + exists, else the original source) +- `${SKILL_AFTER_PATH}` — temporary materialization of the + proposal applied to a copy of BEFORE (the operator session + prepares this; you read it but it's NOT the canonical + `improved-skill/` yet) +- `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the + internal consistency check: does the change make sense given + the skill's stated responsibilities?) +- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for + the external consistency check) +- `${VERDICT_OUTPUT_PATH}` — where to write + `09-validator-verdict.md` +- `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless + this is a re-run with sharper guidance) + +## What you see + +- Skill BEFORE the change +- Skill AFTER the change (temporary materialization) +- The proposal artifact (`08-improvement-proposal.md`) — what the + optimizer claims to have done + their rationale + their self- + check against anti-patterns +- `01-functionality.md` (stated responsibilities) +- `03-submissions.md` if PR-bound (upstream conventions) +- Operator directives + +## What you do NOT see + +- Raw failed trials, `findings.txt`, `trace.jsonl` — your job is + to validate the PROPOSAL, not re-analyze the failures +- Test inputs (`tests//workspace/`) — same reason +- The optimizer's internal reasoning trace from step 8 (only the + proposal artifact, not how the optimizer arrived at it). + Reading the optimizer's reasoning would lead you to accept + arguments they made for the change rather than judging the + artifact fresh. +- Your own prior `09-validator-verdict.md` or git history of it. + No consistency-with-self bias across re-runs — each invocation + judges the new proposal on its own merits. +- `07-analysis.md` directly. The analyzer's reasoning is mediated + through the optimizer's proposal (the optimizer's rationale + references which weakness the change addresses). Reading the + analysis would make you a second analyzer; your job is + independent validation of the proposal. + +## Output: `${VERDICT_OUTPUT_PATH}` — `09-validator-verdict.md` + +Frontmatter (runtime-relevant only): + +```yaml +--- +verdict: approve | needs-revision | reject +addresses_weaknesses: + - + - ... +--- +``` + +Body has two required sections (three if PR-bound): + +### 1. Internal consistency check + +Does the change make sense given the skill's stated +responsibilities in `01-functionality.md`? Walk through these +checks: + +- **Addresses the named weakness?** The proposal claims to + address weakness W. Compare the actual diff to the principle + the analyzer named for W. Does the diff apply that principle? +- **Additive vs. destructive?** Additive changes (adding + procedural guidance, examples, checks) preserve what works. + Destructive changes (removing, replacing, rewriting) need + justification — was the existing content actually the weakness, + or did the optimizer rewrite something that worked? +- **General vs. ducktape?** The optimizer's self-check section + walks through the anti-pattern list and explains why their + proposal is NOT each anti-pattern. Spot-check their reasoning: + is the self-check honest, or is the proposal incidentally doing + one of the anti-patterns and the self-check is hand-waving? +- **Contradicts stated responsibilities?** Does the change + contradict what `01-functionality.md` says the skill is + supposed to do? + +### 2. External consistency check (only if `${SUBMISSIONS_PATH}` exists) + +Forward-looking: would a PR carrying this change conform to the +upstream conventions in `03-submissions.md`? Even though step 9 +doesn't produce a PR, this check verifies the improved skill +COULD be turned into a valid PR by a downstream composer. + +- **Frontmatter spec** — does AFTER match the upstream's + frontmatter conventions? Required fields present, value formats + correct? +- **File-location conventions** — is the diff in the right place + per upstream conventions? +- **Prefix taxonomy** — if the upstream uses commit/PR prefixes, + is the diff's logical scope consistent with one of them? +- **Additive-only conventions** — some upstreams reject + destructive changes for non-bug-fix work; flag if the proposal + is destructive and the upstream's pattern is additive. + +### 3. Rationale + verdict reasoning + +State the verdict and explain it. For `needs-revision`/`reject`, +be concrete — the operator distills your rationale into a +directive for step 8. + +## Verdict definitions + +- **approve** — internal + external (if applicable) checks both + hold. Proposal is sound; step 9 materializes `improved-skill/` + and the chain reaches its endpoint for this iteration. +- **needs-revision** — one or more checks has a fixable issue. + Examples: anti-pattern self-check is hand-waving in a way + that's correctable; frontmatter doesn't match upstream's + required field. Step 8 re-fires with the rationale distilled + into a directive. +- **reject** — fundamental problem the optimizer can't fix by + revising. Examples: the proposal addresses a different weakness + than it claims; the change contradicts `01-functionality.md`; + the named weakness as framed isn't addressable without + ducktape. Rejection surfaces the framing problem upward — the + operator may re-invoke step 7 to reformulate. + +## Reasoning protocol + +1. **Read the proposal first.** What does the optimizer claim to + have done? Which weaknesses do they claim to address? What's + their stated principle? +2. **Read BEFORE and AFTER.** Compute the actual diff (don't + trust the proposal's diff description blindly — verify by + comparing). +3. **Cross-check the diff against the claim.** Does the diff + actually implement the principle the optimizer cited? If + they claim "added procedural guidance for enumerating Xs" + but the diff just adds a MUST clause, that's a mismatch. +4. **Walk the anti-pattern self-check.** For each anti-pattern + the optimizer says they avoided, verify their justification. + The most common failure mode is incidental anti-pattern + overlap with hand-waving justification. +5. **Read `01-functionality.md`** — does the change preserve the + skill's stated responsibilities? +6. **If `03-submissions.md` exists**, walk the external + consistency check. +7. **Choose verdict.** The weakest of the checks. Be honest about + verdicts — false approvals let ducktape ship; false rejects + waste operator cycles. + +## Operator directives + +Examples that count as atomic requirements: + +- "be stricter on additive-vs-destructive — the prior verdict + approved a change I think was actually destructive" +- "the external check missed the frontmatter `version:` field + last round — look for it specifically" +- "weakness W's anti-pattern list is the load-bearing one; + scrutinize the self-check carefully" + +Examples that don't (context dumps — reject): + +- The prior verdict pasted in for you to "reconcile" +- The analyzer's section on weakness W pasted in for you to + "consider" + +## Return summary + +After writing the verdict, return: + +- Verdict (approve / needs-revision / reject) +- Weaknesses addressed (slugs from the proposal) +- Internal-check + external-check summary (one line each) +- For non-approve: the concrete suggestion the operator will + distill into a step-8 directive + +Keep it under 200 words. From 422c8319a244b8bf3d7b8f1427e3612060fc5971 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 09:56:38 -0500 Subject: [PATCH 041/121] =?UTF-8?q?refactor(v1.4-chain):=20vendor-always?= =?UTF-8?q?=20=E2=80=94=20single=20canonical=20input=20regardless=20of=20s?= =?UTF-8?q?ource?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit B1 now vendors the source skill to vendored-skill/ unconditionally, not just for upstream sources. Downstream steps stop branching on "local vs upstream" — they always read vendored-skill/ as the single canonical input, and the user's original local file is never touched by the chain. Why this is cleaner: - One code path through the chain (no local/upstream conditional) - vendored-skill/ is THE input; the original is just the source we copied from (path recorded in 01-functionality.md's skill_source frontmatter) - Stability: long-running chain runs don't break if the user edits the local file mid-flight - improved-skill/ vs vendored-skill/ is the clean before/after pair for both source types; B8 and B9 just say "improved-skill/ if exists, else vendored-skill/" without local-file special cases B1 SKILL.md updates: - "What you produce" paragraph: vendoring is unconditional; user's original local file is not touched by the chain - Workflow step (c) renamed from "Vendor the source (upstream only)" to "Vendor the source" with explicit upstream/local copy paths (gh api fetch / cp -r) - Re-vendor triggers documented: source URL change (upstream), or user explicitly asks (local edits, upstream new commits) Updated all downstream references: - B6 (run-bench): "vendored-skill/ should exist" no longer says "for upstream skills" - B8 (improve), B9 (validate): SKILL_CURRENT_PATH simplified — "improved-skill/ if exists, else vendored-skill/" - iteration-protocol's "what this protocol does NOT cover": vendored-skill/ described as canonical input regardless of upstream/local - subagent-dispatch's step-7+8 carve-out: same simplification. Also fixed pre-existing typo (validator was labeled step 8; should be step 9) - 5 subagent prompts (research-functionality, test-writer, test-validator, analyzer, optimizer, validator): SKILL_SOURCE_PATH / SKILL_CURRENT_PATH / SKILL_BEFORE_PATH all reference vendored-skill/ unconditionally - workflow.md state-layout comment: "vendored-skill/ (always)" - spec doc state-layout comment: same Trim of an architecture branch that wasn't pulling its weight. --- .../2026-05-19-skill-optimizer-v1.4-design.md | 6 ++-- skills/skill-optimizer-improve/SKILL.md | 5 ++-- .../SKILL.md | 30 ++++++++++++++----- skills/skill-optimizer-run-bench/SKILL.md | 3 +- .../iteration-protocol.md | 9 ++++-- .../subagent-dispatch.md | 13 ++++---- skills/skill-optimizer-shared/workflow.md | 2 +- skills/skill-optimizer-subagents/analyzer.md | 3 +- skills/skill-optimizer-subagents/optimizer.md | 3 +- .../research-functionality.md | 12 ++++---- .../test-validator.md | 7 +++-- .../skill-optimizer-subagents/test-writer.md | 4 +-- skills/skill-optimizer-subagents/validator.md | 2 +- skills/skill-optimizer-validate/SKILL.md | 17 ++++++----- 14 files changed, 71 insertions(+), 45 deletions(-) diff --git a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md index b633b19..379ef3e 100644 --- a/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md +++ b/docs/superpowers/specs/2026-05-19-skill-optimizer-v1.4-design.md @@ -145,7 +145,7 @@ docs/skill-optimizer// 07-analysis.md # B7 writes 08-improvement-proposal.md # B8 writes 09-validator-verdict.md # B9 writes - vendored-skill/ # the source skill, read-only after fetch (upstream only) + vendored-skill/ # the source skill, read-only (always — B1 vendors upstream OR local) improved-skill/ # B9 materializes on verdict: approve; original is never modified ``` @@ -487,8 +487,8 @@ Per-step iteration behavior is noted at the end of each subsection. smoke fixtures truly distinguish (vs. coincidentally match)? Aggregate verdicts. - **Dispatches:** test-validator subagent per probe (limited - context: this probe's full contents + parent functionality spec - + `01-functionality.md` + skill source; does NOT see other + context: this probe's full contents, parent functionality spec, + `01-functionality.md`, skill source; does NOT see other probes, prior verdicts, downstream analysis) - **Why this exists:** the smoke check at step 4 only verifies syntactic consistency (grader correctly classifies the GOOD/BAD/ diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index 5592525..aa69c0e 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -70,7 +70,8 @@ Three checks: 3. `01-functionality.md` must exist. Skill content must be readable: `improved-skill/` if it exists (accumulated state - from prior step-8 approvals), else the original source. + from prior step-9 approvals), else `vendored-skill/` (B1 + vendored the source regardless of upstream/local). If `pr_submission_intent: true`, `03-submissions.md` should exist so the optimizer can shape the diff to upstream conventions from @@ -104,7 +105,7 @@ and substitute `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, The optimizer sees: `07-analysis.md`; `01-functionality.md`; `${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else -original source; `03-submissions.md` if PR-bound; +`vendored-skill/`; `03-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. The optimizer does NOT see: raw failed trials, `findings.txt`, diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index af9315c..3fc77d9 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -36,8 +36,11 @@ falling back to `other`. For the body template, see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). -If the source is upstream, also vendor the fetched skill files to -`vendored-skill/` so downstream steps read a stable copy. +The source skill is vendored to `vendored-skill/` regardless of +source type (upstream fetch or local copy) so downstream steps +have a single canonical input — they never branch on local vs +upstream, and the user's original local file is not touched by +the chain. ## Workflow @@ -71,13 +74,24 @@ If the user later changes their mind on PR intent, they re-run this skill — the canonical is overwritten and prior state lives in git history. -### (c) Vendor the source (upstream only) +### (c) Vendor the source + +Copy the skill's files into `vendored-skill/` at the +working-directory root, regardless of source type: + +- **Upstream** — `gh api` or equivalent fetch of the skill's + directory contents +- **Local** — `cp -r` of the local skill's directory into + `vendored-skill/` -Fetch the skill's files into `vendored-skill/` at the -working-directory root. Local-source skills don't need vendoring. If `vendored-skill/` already exists from a prior run, reuse it -unless the source URL changed or the user explicitly asks to -re-fetch. +unless: (1) the source URL changed (upstream — a different repo or +skill is being investigated), or (2) the user explicitly asks to +re-vendor (e.g., they edited a local skill between runs, or the +upstream got new commits worth refetching). + +The user's original local file is never modified by the chain — +`vendored-skill/` is a separate copy that downstream steps read. ### (d) Determine the slug and the report path @@ -104,7 +118,7 @@ prompt template at and substitute `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, `${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, vendored path. -The subagent sees: the vendored skill files (or local path), +The subagent sees: the vendored skill files at `vendored-skill/`, targeted web-search results, output path, frontmatter fields, `${OPERATOR_DIRECTIVES}`. diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index beb35ae..2501f9e 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -44,7 +44,8 @@ Two outputs at `docs/skill-optimizer//`: `tests/suite.yml` must exist (per step 4's output contract). If missing, tell the user to complete step 4 first. -`vendored-skill/` should exist for upstream skills. +`vendored-skill/` should exist (B1 vendored the source regardless +of upstream/local). ### (b) Handle iteration diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index a429a54..9996e43 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -125,6 +125,9 @@ needed — git is the archive. `docs/skill-optimizer//autopilot-summary-.md` is timestamped per run. Each auto-pilot invocation produces a fresh summary; old ones are preserved naturally. -- **`vendored-skill/`** (the cached upstream skill source) is - reused across iterations of the same slug unless the source URL - changed. The source URL itself is the identity; no versioning. +- **`vendored-skill/`** (the canonical input to the chain — B1 + copies the source skill here regardless of upstream/local) is + reused across iterations of the same slug unless: (1) the + source URL changed (upstream), or (2) the user explicitly asks + to re-vendor (e.g., they edited the local skill between runs). + The source URL itself is the identity; no versioning. diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md index aff4012..842ea39 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -53,13 +53,14 @@ Three rules every reasoning subagent must follow: **For step 8 and step 9 specifically:** the SKILL CONTENT (the target being improved) is upstream input, not your own canonical. Both the optimizer (step 8) and the validator (step - 8) read the **current state of the skill** — which is + 9) read the **current state of the skill** — which is `improved-skill/` if it exists (the accumulated state from - prior step-8 approvals), else the original source. Neither step - modifies the original. The "own canonical" off-limits to each - subagent is its report (`08-improvement-proposal.md` for the - optimizer; `09-validator-verdict.md` for the validator), not - the skill content itself. + prior step-9 approvals), else `vendored-skill/` (B1 vendored + the source regardless of upstream/local). Neither step modifies + the original. The "own canonical" off-limits to each subagent + is its report (`08-improvement-proposal.md` for the optimizer; + `09-validator-verdict.md` for the validator), not the skill + content itself. 3. **For maintenance steps (2, 4): DO read your own canonical tree** (when it exists). Your job on a re-run is to extend or diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index 20f6aa7..30efdb2 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -96,7 +96,7 @@ docs/skill-optimizer// 07-analysis.md 08-improvement-proposal.md 09-validator-verdict.md - vendored-skill/ # frozen original (upstream only) + vendored-skill/ # frozen original (always — B1 vendors upstream OR local) improved-skill/ # B9 materializes on approve autopilot-summary-.md # B10 writes per run ``` diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/skill-optimizer-subagents/analyzer.md index 8b149d8..1049ea1 100644 --- a/skills/skill-optimizer-subagents/analyzer.md +++ b/skills/skill-optimizer-subagents/analyzer.md @@ -17,7 +17,8 @@ address it AND the anti-patterns that would NOT. - `${TESTS_TREE_PATH}` — `tests/` tree. **Read ONLY each probe's `spec.yaml`** — what each probe was probing at the level of INTENT. Do NOT read `workspace/` contents (raw input fixtures). -- `${SKILL_SOURCE_PATH}` — the skill's content (vendored or local) +- `${SKILL_SOURCE_PATH}` — the skill's content at `vendored-skill/` + (B1 vendored the source regardless of upstream/local) - `${OUTPUT_PATH}` — where to write `07-analysis.md` - `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless this is a re-run with sharper guidance) diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md index c102d58..1290c2d 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -14,7 +14,8 @@ skill (that's step 9's job after the validator approves). stated responsibilities — your change must not contradict them) - `${SKILL_CURRENT_PATH}` — the **current state of the skill**: `improved-skill/` if it exists (accumulated state from prior - step-9 approvals), else the original source. You propose a new + step-9 approvals), else `vendored-skill/` (B1 vendored the + source regardless of upstream/local). You propose a new improvement on top of whatever current state you read. - `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (lets you shape the diff to match upstream conventions from the start, diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index cf2e3db..d257ec3 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -8,10 +8,11 @@ the chain consumes. ## Inputs (templated by the operator session) - `${SKILL_SOURCE}` — URL (`//`) or local - filesystem path to the skill being investigated -- `${VENDORED_PATH}` — `vendored-skill/` directory if the source was - upstream and the operator session vendored it (otherwise empty; - read directly from `${SKILL_SOURCE}`) + filesystem path to the skill being investigated (recorded in + the report frontmatter, but you read from `${VENDORED_PATH}`) +- `${VENDORED_PATH}` — `vendored-skill/` directory (the operator + session has already copied the source skill here, whether the + original was upstream or local; this is your read path) - `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 by the operator session - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new @@ -21,7 +22,8 @@ the chain consumes. ## What you see -- The skill's files (SKILL.md + any references/, scripts/, assets/) +- The vendored skill's files at `${VENDORED_PATH}` (SKILL.md + any + references/, scripts/, assets/) - Targeted web-search / web-fetch for the underlying technology the skill is about - The operator's directives diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/skill-optimizer-subagents/test-validator.md index a07194b..72ac885 100644 --- a/skills/skill-optimizer-subagents/test-validator.md +++ b/skills/skill-optimizer-subagents/test-validator.md @@ -24,9 +24,10 @@ coincidentally match rather than truly distinguish. - `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's spec.yaml (the responsibility this probe is supposed to test) - `${FUNCTIONALITY_PATH}` — `01-functionality.md` -- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local - skill path) — needed to judge whether the probe exercises what - the skill actually instructs the agent to do +- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (B1 vendored the + source regardless of upstream/local) — needed to judge whether + the probe exercises what the skill actually instructs the agent + to do - `${VERDICT_OUTPUT_PATH}` — where to write your verdict (typically one per-probe verdict file the operator aggregates, OR a single return value the operator collects across parallel dispatches) diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md index 7557312..8068513 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -17,8 +17,8 @@ its own probe folder and nothing else. `spec.yaml` describing what this probe sets up + expects - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (skill's stated responsibilities and classification) -- `${SKILL_SOURCE_PATH}` — the vendored skill source (or local - skill path) +- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (B1 vendored the + source regardless of upstream/local) - `${OUTPUT_PROBE_DIR}` — `tests///` - `${OPERATOR_DIRECTIVES}` — case-level revision hints for this probe (empty unless this is a rebuild) diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/skill-optimizer-subagents/validator.md index f4fbb4d..65d5c7f 100644 --- a/skills/skill-optimizer-subagents/validator.md +++ b/skills/skill-optimizer-subagents/validator.md @@ -18,7 +18,7 @@ without your own prior verdicts. optimizer's diff + rationale + self-check) - `${SKILL_BEFORE_PATH}` — the current state of the skill (same state the optimizer read at step 8: `improved-skill/` if it - exists, else the original source) + exists, else `vendored-skill/`) - `${SKILL_AFTER_PATH}` — temporary materialization of the proposal applied to a copy of BEFORE (the operator session prepares this; you read it but it's NOT the canonical diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index 993c363..d21a755 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -47,8 +47,9 @@ One or two artifacts at `docs/skill-optimizer//`: 2. **`improved-skill/`** — improved skill content, only materialized when `verdict: approve`. Mirrors the source's directory structure with the optimizer's diff applied. **The - original is never modified**: `vendored-skill/` stays frozen, - local source files stay untouched. Git tracks + original is never modified**: `vendored-skill/` (the canonical + input regardless of upstream/local) stays frozen, and the + user's original local file (if any) is untouched. Git tracks `improved-skill/` history across iterations. ## Workflow @@ -61,12 +62,12 @@ Three checks: If not, tell the user to run `skill-optimizer-improve` first. 2. Current skill state must be readable: `improved-skill/` if it - exists (prior accumulated state), else the original source. - The validator needs this as "skill BEFORE". If both - `improved-skill/` and the source were modified externally - between step 8 and step 9, surface to the user — the BEFORE - must match what the optimizer read. The fix is re-invoking - step 8 against the new state. + exists (prior accumulated state), else `vendored-skill/`. The + validator needs this as "skill BEFORE". If `vendored-skill/` + was re-vendored between step 8 and step 9 (e.g., the operator + re-ran B1 mid-chain), surface to the user — the BEFORE must + match what the optimizer read. The fix is re-invoking step 8 + against the new state. 3. `01-functionality.md` must exist. If `pr_submission_intent: true`, `03-submissions.md` should From 1a19d54fc8431b40c25cac3ea1c7610eb99763db Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 10:47:15 -0500 Subject: [PATCH 042/121] feat(b3-submissions): detect entry-file pattern + map linked consumers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Real-world context from prior pilot runs: sometimes the user points at a SKILL.md that's just a thin wrapper referencing the actual content elsewhere. Common pattern — a multi-agent plugin has one canonical agent-agnostic content file and several agent-flavored SKILL.md wrappers (one per agent target) that all reference it. The PR target is the canonical content, not the wrapper; a change to the canonical may affect multiple wrappers. Updated research-submissions.md subagent prompt: New frontmatter fields the subagent emits: - entry_file_pattern: true|false - canonical_target_repo: / - canonical_target_path: - linked_consumers: [:, ...] New body sections: 1. Source structure — entry-file vs canonical content; if entry-file, document the relationship + linked consumers 2. Suggested PR target — based on (1), recommend which file(s) to modify and whether the change affects other consumers New reasoning protocol steps: 1. Detect entry-file pattern (read source SKILL.md, look for thin-wrapper signals: short body of "see X" pointers, frontmatter fields like reference: / source: / canonical:, multi-agent plugin layout). Follow the pointer to find the canonical content if detected. 2. Find linked consumers (search for other SKILL.md files referencing the same canonical content) 7. Suggest the PR target based on the above The license / CLA / frontmatter / conventions research now applies to the CANONICAL CONTENT'S repo, which may differ from the entry file's repo. Note: B1 (functionality researcher) likely also needs awareness of this pattern — if the user pointed at an entry file, downstream chain steps test/analyze/optimize the wrapper rather than the actual skill. Flagged as follow-up but not addressed in this commit (the user asked specifically about B3). --- .../research-submissions.md | 119 +++++++++++++++--- 1 file changed, 101 insertions(+), 18 deletions(-) diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index d40b20d..b8d1857 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -51,6 +51,12 @@ upstream_repo: / upstream_branch_target: license: requires_cla: true | false +entry_file_pattern: true | false +canonical_target_repo: / | +canonical_target_path: +linked_consumers: + - : # other places that reference the canonical content + - ... --- ``` @@ -60,28 +66,58 @@ requires_cla: true | false branch name. - **`requires_cla`** — `true` if the upstream requires a Contributor License Agreement. +- **`entry_file_pattern`** — `true` if the source SKILL.md is a + thin wrapper that points to the actual content elsewhere + (common: an agent-agnostic content file referenced by several + agent-flavored SKILL.md wrappers). The downstream PR composer + needs to know this to target the right file. +- **`canonical_target_repo` / `canonical_target_path`** — where + the actual skill content lives. Same as `upstream_repo` and the + source path if it's not an entry-file pattern; otherwise the + canonical location. +- **`linked_consumers`** — other SKILL.md files (in the same repo + or other repos) that reference the same canonical content. If + non-empty, a change to the canonical content affects all of + them; the PR composer needs to coordinate or be honest about + the scope. Body sections: -1. **License** — full SPDX identifier + brief plain-English summary - (permissive / copyleft / proprietary). -2. **CLA requirement** — what kind (DCO, individual CLA, corporate - CLA), how it's signed, link to the CLA tool if applicable. -3. **Frontmatter spec** — extracted from a sample of existing - skills in the repo. List required fields, optional fields, - value conventions. If the repo's skills don't have a consistent - frontmatter, say so rather than inventing a standard. -4. **File-location conventions** — where new skills go (which +1. **Source structure** (new) — is the input SKILL.md the actual + content, or an entry file pointing elsewhere? If entry-file: + document the relationship (entry-file path, canonical content + path, link mechanism). If you find linked consumers, list them + here with the relationship (other agent wrappers, downstream + skills that import from this one). +2. **Suggested PR target** (new) — based on (1), which location + should the PR modify? Usually the canonical content; sometimes + the entry file (if the change is wrapper-specific). If the + change would affect multiple consumers, flag whether they live + in the same repo (single PR) or different repos (multiple PRs + or coordination required). +3. **License** — full SPDX identifier + brief plain-English summary + (permissive / copyleft / proprietary). Document the license for + the canonical content's repo (which may differ from the entry + file's repo). +4. **CLA requirement** — what kind (DCO, individual CLA, corporate + CLA), how it's signed, link to the CLA tool if applicable. Same + note: applies to the canonical content's repo. +5. **Frontmatter spec** — extracted from a sample of existing + skills in the canonical content's repo. List required fields, + optional fields, value conventions. If the repo's skills don't + have a consistent frontmatter, say so rather than inventing a + standard. +6. **File-location conventions** — where new skills go (which directory, which subdir pattern), where references/scripts go. -5. **Prefix taxonomy** — if the repo uses commit prefixes +7. **Prefix taxonomy** — if the repo uses commit prefixes (`feat:`, `fix:`, `docs:`, etc.) or PR title prefixes, extract the taxonomy from recent merged PRs. -6. **PR-shape patterns** — from a sample of recent merged PRs: +8. **PR-shape patterns** — from a sample of recent merged PRs: what does a typical PR description include (rationale, testing notes, screenshots)? What's the typical PR size (single-file diff, multi-file)? Additive-only or destructive changes accepted? -7. **Rejection signals** — from a sample of closed-without-merge +9. **Rejection signals** — from a sample of closed-without-merge PRs: what got rejected and why? Common patterns to AVOID. If you can't establish any of these from the available data, say @@ -90,19 +126,64 @@ conventions. ## Reasoning protocol -1. **Start with CONTRIBUTING.md** — if it exists, it explicitly +1. **Detect entry-file vs canonical-content pattern.** Read the + source SKILL.md (the file at the path the user pointed to). + Check whether it's the actual skill content or a thin wrapper + that just references content elsewhere. Common signals of an + entry-file wrapper: + - The SKILL.md body is short and consists mainly of "see X" + or "this skill uses content from Y" pointers + - Frontmatter has fields like `reference:`, `source:`, + `canonical:`, or similar pointing to another file + - The "skill" is one of several SKILL.md files in a multi-agent + plugin layout, each wrapping the same underlying content + + If it IS an entry-file pattern, follow the pointer to find the + canonical content. Update the frontmatter's + `canonical_target_repo` and `canonical_target_path` accordingly. + The rest of your research focuses on the canonical content's + home (PR conventions, license, CLA all apply there). + +2. **Find linked consumers.** Search for other places that + reference the canonical content: + - Other SKILL.md files in the same repo that point to the + same content (agent-flavored wrappers) + - Downstream skills that import or reference this skill + - In the repo's package metadata, search for known dependents + + Record findings in `linked_consumers` frontmatter and the + "Source structure" body section. + +3. **Start with CONTRIBUTING.md** — if it exists, it explicitly states most of what you need (license, CLA, file conventions, - PR shape). -2. **Sample recent merged PRs** — enough to identify the dominant + PR shape). Apply this to the CANONICAL content's repo, which + may differ from the entry-file's repo if (1) found a pattern. + +4. **Sample recent merged PRs** — enough to identify the dominant patterns. Quality over quantity; you want to see the actual accepted shapes. -3. **Sample closed-without-merge PRs** — these surface the + +5. **Sample closed-without-merge PRs** — these surface the rejection signals. The CONTRIBUTING.md tells you the rules; the closed PRs show what happens when rules are broken. -4. **Look at PRs in the same skill category** if `${SKILL_SLUG}` + +6. **Look at PRs in the same skill category** if `${SKILL_SLUG}` suggests one (e.g., browser-skills look at other browser PRs). Specific-category patterns trump generic repo-level patterns. -5. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** + +7. **Suggest the PR target.** Based on (1) and (2), recommend + which file(s) the PR should modify: + - Not entry-file pattern: PR against the source as-is + - Entry-file pattern, single consumer: PR against the + canonical content + - Entry-file pattern, multiple consumers in the same repo: + PR against the canonical content; the wrappers update + automatically + - Entry-file pattern, multiple consumers across repos: PR + against the canonical content; flag that downstream wrappers + may need follow-up PRs in their own repos + +8. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** (e.g., "include the vendor's CLA requirement explicitly" → make sure the CLA section is prominent). @@ -124,6 +205,8 @@ blocker rather than making up content: After writing `${OUTPUT_PATH}`, return a brief summary: +- Entry-file pattern detected? If yes: canonical location + + linked consumers count + suggested PR target - License + CLA requirement - Branch target rule - 2-3 most important rejection signals from the closed-PR sweep From f9d273795c42d1a36fdecd1136ba0f2896747ae9 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 21 May 2026 11:15:34 -0500 Subject: [PATCH 043/121] feat(b1-functionality): detect entry-file pattern in vendor step MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit B1 now handles the entry-file pattern symmetrically with B3 — detected at vendor time so downstream chain steps test/analyze/ optimize the actual skill content rather than a thin wrapper. Previously: if the user pointed at an entry-file SKILL.md (a wrapper referencing the actual content elsewhere), B1 vendored the wrapper and all downstream steps operated on it. The whole optimization run would miss its real target. Now (B1 SKILL.md): - New step (c.1) "Detect entry-file pattern + user gate" runs after the initial vendor at (c). Operator reads vendored-skill/SKILL.md for thin-wrapper signals (short "see X" body; frontmatter reference:/source:/canonical: fields; multi-agent plugin layout) - If detected, surfaces to user with a clear choice: optimize the wrapper, or re-vendor the canonical content and optimize that. The common case (canonical) re-vendors; the rare case (wrapper-specific change) keeps the existing vendored content but records the relationship as an operator directive - Either way, 01-functionality.md frontmatter records entry_file_pattern: true|false and canonical_source: for downstream steps Updated research-functionality.md subagent prompt: - New inputs: ${ENTRY_FILE_PATTERN}, ${CANONICAL_SOURCE} - New frontmatter fields: entry_file_pattern, canonical_source - New body section 9: "Entry-file relationship" (only if pattern detected) — notes the relationship, the vendored copy content, any agent-specific adaptations Updated research-submissions.md (B3) subagent prompt: - Reasoning step 1 now reads 01-functionality.md's entry-file fields as primary input. B3 trusts B1's detection; falls back to its own detection only if 01-functionality.md was written before this feature (defensive). If B3 detects a pattern B1 missed, surface to operator — suggests B1 needs a re-run The whole chain now consistently knows which is the wrapper and which is the canonical content from B1 onward. --- .../SKILL.md | 51 +++++++++++++++++++ .../research-functionality.md | 16 ++++++ .../research-submissions.md | 36 ++++++------- 3 files changed, 86 insertions(+), 17 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 3fc77d9..1383f6a 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -25,6 +25,8 @@ Frontmatter (runtime-relevant facts only, per skill_source: pr_submission_intent: true | false classification: +entry_file_pattern: true | false +canonical_source: --- ``` @@ -34,6 +36,15 @@ cleanly, write a short descriptive label of your own (`dataset-extraction`, `deployment-runbook`, etc.) rather than falling back to `other`. +**`entry_file_pattern`** — `true` if the source SKILL.md is a thin +wrapper pointing at the actual content elsewhere (common: a +multi-agent plugin has one canonical agent-agnostic content file +and several agent-flavored SKILL.md wrappers, one per agent). When +`true`, **`canonical_source`** records where the actual content +lives. Downstream chain steps (B3 submission research, B8 +optimization, B9 validation) all need this to target the real +skill, not the wrapper. + For the body template, see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). The source skill is vendored to `vendored-skill/` regardless of @@ -93,6 +104,46 @@ upstream got new commits worth refetching). The user's original local file is never modified by the chain — `vendored-skill/` is a separate copy that downstream steps read. +### (c.1) Detect entry-file pattern + user gate + +Read `vendored-skill/SKILL.md`. Check for thin-wrapper signals: + +- Body is short (under ~30 lines) and consists mostly of "see X" + or "this skill uses content from Y" pointers +- Frontmatter has fields like `reference:`, `source:`, + `canonical:`, or similar pointing to another file +- The plugin layout suggests a multi-agent structure (several + SKILL.md files in the same plugin, each appearing to wrap the + same underlying content) + +If the pattern is detected: + +1. Follow the pointer to find the **canonical content** — the + actual agent-agnostic skill file the wrapper references. This + may live in the same repo (different path) or a different + repo entirely. +2. **Surface to the user.** Say something like: "The source you + gave me looks like a thin wrapper. The actual skill content + is at ``. Optimizing the wrapper would + produce a wrapper-specific improvement; optimizing the + canonical content would affect all wrappers that reference + it. Which do you want to optimize?" +3. **If the user picks canonical** (the common case): re-vendor. + Delete `vendored-skill/` and copy the canonical content into + it instead. Record `entry_file_pattern: true` and + `canonical_source: ` in the report frontmatter + you'll write at (f). +4. **If the user picks the wrapper** (rare; only when the + wrapper itself is the thing they want to improve): keep the + existing vendored content. Still record + `entry_file_pattern: true` and `canonical_source:` so + downstream steps know the relationship. Record their + wrapper-only choice in `${OPERATOR_DIRECTIVES}` for the + subagent. + +If the pattern is NOT detected: skip this step, leave +`entry_file_pattern: false`, and proceed. + ### (d) Determine the slug and the report path `` is the source skill's own directory or file name (e.g., diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index d257ec3..624cab3 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -13,6 +13,13 @@ the chain consumes. - `${VENDORED_PATH}` — `vendored-skill/` directory (the operator session has already copied the source skill here, whether the original was upstream or local; this is your read path) +- `${ENTRY_FILE_PATTERN}` — `true` or `false`, from the operator's + detection at step (c.1). If `true`, `${VENDORED_PATH}` contains + the **canonical content** (the operator confirmed with the user + and re-vendored), and `${CANONICAL_SOURCE}` records where it + came from. +- `${CANONICAL_SOURCE}` — URL or path of the canonical content + (empty if `entry_file_pattern: false`) - `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 by the operator session - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new @@ -50,6 +57,8 @@ Frontmatter (runtime-relevant only, no version-tracking metadata): skill_source: pr_submission_intent: classification: +entry_file_pattern: +canonical_source: <${CANONICAL_SOURCE}, omitted if entry_file_pattern: false> --- ``` @@ -92,6 +101,13 @@ Body (markdown) covering: guidelines live (URL, CONTRIBUTING.md path, Slack channel, whatever they provided). Step 3 uses this as its starting point. +9. **Entry-file relationship** (only if `${ENTRY_FILE_PATTERN}` + is `true`) — note that the source the user pointed at was a + thin wrapper, the actual content is at `${CANONICAL_SOURCE}`, + and the vendored copy contains the canonical content (per + operator's confirmation at step c.1). Mention any + wrapper-specific behavior or adaptations the canonical + content has for this particular agent target. ## Reasoning protocol diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index b8d1857..0fd14a3 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -126,23 +126,25 @@ conventions. ## Reasoning protocol -1. **Detect entry-file vs canonical-content pattern.** Read the - source SKILL.md (the file at the path the user pointed to). - Check whether it's the actual skill content or a thin wrapper - that just references content elsewhere. Common signals of an - entry-file wrapper: - - The SKILL.md body is short and consists mainly of "see X" - or "this skill uses content from Y" pointers - - Frontmatter has fields like `reference:`, `source:`, - `canonical:`, or similar pointing to another file - - The "skill" is one of several SKILL.md files in a multi-agent - plugin layout, each wrapping the same underlying content - - If it IS an entry-file pattern, follow the pointer to find the - canonical content. Update the frontmatter's - `canonical_target_repo` and `canonical_target_path` accordingly. - The rest of your research focuses on the canonical content's - home (PR conventions, license, CLA all apply there). +1. **Read `01-functionality.md` for the entry-file context.** + B1 has already detected the entry-file pattern (if any) and + recorded the result in the report's frontmatter: + - `entry_file_pattern: true` → the source the user pointed at + was a wrapper; `canonical_source` records where the actual + content lives. The chain's `vendored-skill/` contains the + canonical content (per operator's re-vendor at step c.1). + - `entry_file_pattern: false` → the source is the actual + content; proceed normally. + + **Trust B1's detection** as primary input. Use the recorded + `canonical_source` to determine `canonical_target_repo` and + `canonical_target_path` for your own frontmatter. If + `01-functionality.md` was written before this feature existed + and doesn't have the field, fall back to detecting yourself + (signals: short SKILL.md body of "see X" pointers, frontmatter + fields like `reference:`/`source:`/`canonical:`, multi-agent + plugin layout). Surface to the operator if you detect a + pattern B1 missed — it suggests B1 needs a re-run. 2. **Find linked consumers.** Search for other places that reference the canonical content: From 11c8c0cac6bd4690d53ffec47921f8a0000b6e9c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 06:40:37 -0500 Subject: [PATCH 044/121] fix(subagent-prompts): drop invisible internal-step references MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Subagent prompts shouldn't reference chain-skill internal workflow labels (the lettered (a)-(g) steps inside each chain SKILL.md) — subagents only see their dispatched inputs + their prompt, never the chain skill itself. Two references to "step c.1" (B1's internal entry-file detection step) were invisible noise. Rephrased to describe what happened functionally — "the operator session confirmed with the user and re-vendored" — without naming the step label. Chain-step number references (step 1 through step 10) stay, because those are stable role descriptors the subagent understands as "another role in the chain", not internal workflow lettering. Files touched: - skills/skill-optimizer-subagents/research-functionality.md - skills/skill-optimizer-subagents/research-submissions.md --- .../research-functionality.md | 13 +++++++------ .../research-submissions.md | 3 ++- 2 files changed, 9 insertions(+), 7 deletions(-) diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index 624cab3..0ee777f 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -102,12 +102,13 @@ Body (markdown) covering: whatever they provided). Step 3 uses this as its starting point. 9. **Entry-file relationship** (only if `${ENTRY_FILE_PATTERN}` - is `true`) — note that the source the user pointed at was a - thin wrapper, the actual content is at `${CANONICAL_SOURCE}`, - and the vendored copy contains the canonical content (per - operator's confirmation at step c.1). Mention any - wrapper-specific behavior or adaptations the canonical - content has for this particular agent target. + is `true`) — note that the source the user pointed at + (`${SKILL_SOURCE}`) was a thin wrapper; the actual content + is at `${CANONICAL_SOURCE}`, and that's what + `${VENDORED_PATH}` contains (the operator session confirmed + with the user and re-vendored before dispatching you). + Mention any wrapper-specific behavior or adaptations the + canonical content has for this particular agent target. ## Reasoning protocol diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index 0fd14a3..65904a0 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -132,7 +132,8 @@ conventions. - `entry_file_pattern: true` → the source the user pointed at was a wrapper; `canonical_source` records where the actual content lives. The chain's `vendored-skill/` contains the - canonical content (per operator's re-vendor at step c.1). + canonical content (B1's operator session already re-vendored + after confirming with the user). - `entry_file_pattern: false` → the source is the actual content; proceed normally. From f8a7f7bd5cde8b512d675a5321f94a5b82566042 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 06:47:32 -0500 Subject: [PATCH 045/121] refactor(wrapper-detection): subagent judgment, not operator mechanical check MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Previous design had B1's operator session do mechanical wrapper detection at workflow step (c.1) — read SKILL.md, check for heuristic signals, surface to user. Wrong design: the subagent is already reading the source for research, has LLM judgment that beats a mechanical check, and the heuristics ("body under 30 lines", "frontmatter has reference:") would miss real cases. Restructured: detection is the subagent's judgment, reported in its return summary. Operator reacts by surfacing to user, who picks the resolution. B1 SKILL.md changes: - Removed workflow step (c.1) — no more operator-side detection - Step (g) renamed "Confirm + handle wrapper detection + hand off" — reads the subagent's `likely_wrapper` frontmatter field; if true, surfaces to user with three realistic responses: 1. Re-vendor the referenced content and re-research (common) 2. Proceed treating the wrapper as the skill (rare) 3. Cancel and provide a different source - Frontmatter field rename: `entry_file_pattern` -> `likely_wrapper` (more honest — it's a judgment, not a binary classification) and `canonical_source` -> `wrapper_points_to` (descriptive rather than presuming a canonical/wrapper hierarchy) research-functionality.md (B1 subagent prompt): - Removed ${ENTRY_FILE_PATTERN} and ${CANONICAL_SOURCE} inputs — these were operator-pre-detected fields. Detection now lives in the subagent's reasoning protocol. - New "Wrapper detection" section in the reasoning protocol — explicit patterns to look for, explicit instruction to record `likely_wrapper` + `wrapper_points_to` in frontmatter when judged true, and explicit instruction to NOT follow the pointer or re-vendor itself (operator's job after user gate) - Body section 9 renamed to "Wrapper observation" — describes signals + confidence rather than asserting a canonical/wrapper relationship - Return summary now includes the wrapper finding so operator can trigger the user gate research-submissions.md (B3 subagent prompt): - Frontmatter field rename: `canonical_target_repo` / `canonical_target_path` -> `pr_target_repo` / `pr_target_path` (cleaner — what the PR composer needs is the PR target, not a taxonomy of canonical-vs-wrapper) - Reasoning step 1 reads B1's `likely_wrapper` / `wrapper_points_to` as authoritative; falls back to own judgment only if 01-functionality predates this feature Net: wrapper-detection happens once (at B1, by the subagent), and the result flows through frontmatter to downstream steps. Operator sessions handle the user gates; no operator does LLM-style judgment work. --- .../SKILL.md | 97 ++++++-------- .../research-functionality.md | 62 ++++++--- .../research-submissions.md | 125 +++++++++--------- 3 files changed, 154 insertions(+), 130 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 1383f6a..3e1549a 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -25,8 +25,8 @@ Frontmatter (runtime-relevant facts only, per skill_source: pr_submission_intent: true | false classification: -entry_file_pattern: true | false -canonical_source: +likely_wrapper: true | false +wrapper_points_to: --- ``` @@ -36,14 +36,20 @@ cleanly, write a short descriptive label of your own (`dataset-extraction`, `deployment-runbook`, etc.) rather than falling back to `other`. -**`entry_file_pattern`** — `true` if the source SKILL.md is a thin -wrapper pointing at the actual content elsewhere (common: a -multi-agent plugin has one canonical agent-agnostic content file -and several agent-flavored SKILL.md wrappers, one per agent). When -`true`, **`canonical_source`** records where the actual content -lives. Downstream chain steps (B3 submission research, B8 -optimization, B9 validation) all need this to target the real -skill, not the wrapper. +**`likely_wrapper`** — the subagent's judgment that the source the +user pointed at appears to be a thin wrapper file (e.g., a +SKILL.md that mostly references content elsewhere, common in +multi-agent plugins where one canonical agent-agnostic content +file is wrapped by several agent-flavored SKILL.md files). When +`true`, **`wrapper_points_to`** records where the subagent thinks +the actual content lives. Downstream chain steps (B3 submission +research, B8 optimization, B9 validation) consult these fields +when deciding what to target. + +If the subagent flags `likely_wrapper: true`, the operator surfaces +to the user (workflow step (g) below) and the user decides what +to do — re-vendor and re-research the referenced content, or +proceed treating the wrapper itself as the skill to optimize. For the body template, see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). @@ -104,46 +110,6 @@ upstream got new commits worth refetching). The user's original local file is never modified by the chain — `vendored-skill/` is a separate copy that downstream steps read. -### (c.1) Detect entry-file pattern + user gate - -Read `vendored-skill/SKILL.md`. Check for thin-wrapper signals: - -- Body is short (under ~30 lines) and consists mostly of "see X" - or "this skill uses content from Y" pointers -- Frontmatter has fields like `reference:`, `source:`, - `canonical:`, or similar pointing to another file -- The plugin layout suggests a multi-agent structure (several - SKILL.md files in the same plugin, each appearing to wrap the - same underlying content) - -If the pattern is detected: - -1. Follow the pointer to find the **canonical content** — the - actual agent-agnostic skill file the wrapper references. This - may live in the same repo (different path) or a different - repo entirely. -2. **Surface to the user.** Say something like: "The source you - gave me looks like a thin wrapper. The actual skill content - is at ``. Optimizing the wrapper would - produce a wrapper-specific improvement; optimizing the - canonical content would affect all wrappers that reference - it. Which do you want to optimize?" -3. **If the user picks canonical** (the common case): re-vendor. - Delete `vendored-skill/` and copy the canonical content into - it instead. Record `entry_file_pattern: true` and - `canonical_source: ` in the report frontmatter - you'll write at (f). -4. **If the user picks the wrapper** (rare; only when the - wrapper itself is the thing they want to improve): keep the - existing vendored content. Still record - `entry_file_pattern: true` and `canonical_source:` so - downstream steps know the relationship. Record their - wrapper-only choice in `${OPERATOR_DIRECTIVES}` for the - subagent. - -If the pattern is NOT detected: skip this step, leave -`entry_file_pattern: false`, and proceed. - ### (d) Determine the slug and the report path `` is the source skill's own directory or file name (e.g., @@ -183,9 +149,34 @@ user has already expressed. A subagent walled off from that context produces a fresh derivation from the source itself, not a rationalization of expectations. -### (g) Confirm and hand off - -Verify the report file exists and frontmatter parses. Then: +### (g) Confirm + handle wrapper detection + hand off + +Verify the report file exists and frontmatter parses. + +**If the subagent flagged `likely_wrapper: true`** in the +frontmatter (it judged the source to be a thin wrapper pointing at +content elsewhere), surface to the user before handing off: + +> The researcher subagent thinks the source you pointed at +> (``) is likely a wrapper file — it appears to +> reference the actual content at ``. Three +> realistic responses: +> +> 1. **Re-vendor the referenced content and re-research.** The +> chain will then target the actual content for testing, +> analysis, and optimization. (Common — usually what you +> want.) +> 2. **Proceed treating this wrapper as the skill.** Useful if +> the wrapper itself is what you want to improve (rare). +> 3. **Cancel and provide a different source.** If the subagent's +> judgment is wrong or you meant to point elsewhere. + +If the user picks (1): delete `vendored-skill/`, vendor the +referenced content into it, and re-invoke this skill from (c). +If (2): keep the report as-is and proceed. If (3): exit; the user +will re-invoke with a new source. + +Once resolved, hand off: > Next, invoke `skill-optimizer-investigate-test-case`. Note: this > report's `pr_submission_intent` field tells step 2's handoff diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index 0ee777f..2010925 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -13,13 +13,6 @@ the chain consumes. - `${VENDORED_PATH}` — `vendored-skill/` directory (the operator session has already copied the source skill here, whether the original was upstream or local; this is your read path) -- `${ENTRY_FILE_PATTERN}` — `true` or `false`, from the operator's - detection at step (c.1). If `true`, `${VENDORED_PATH}` contains - the **canonical content** (the operator confirmed with the user - and re-vendored), and `${CANONICAL_SOURCE}` records where it - came from. -- `${CANONICAL_SOURCE}` — URL or path of the canonical content - (empty if `entry_file_pattern: false`) - `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 by the operator session - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new @@ -57,8 +50,8 @@ Frontmatter (runtime-relevant only, no version-tracking metadata): skill_source: pr_submission_intent: classification: -entry_file_pattern: -canonical_source: <${CANONICAL_SOURCE}, omitted if entry_file_pattern: false> +likely_wrapper: +wrapper_points_to: --- ``` @@ -101,14 +94,13 @@ Body (markdown) covering: guidelines live (URL, CONTRIBUTING.md path, Slack channel, whatever they provided). Step 3 uses this as its starting point. -9. **Entry-file relationship** (only if `${ENTRY_FILE_PATTERN}` - is `true`) — note that the source the user pointed at - (`${SKILL_SOURCE}`) was a thin wrapper; the actual content - is at `${CANONICAL_SOURCE}`, and that's what - `${VENDORED_PATH}` contains (the operator session confirmed - with the user and re-vendored before dispatching you). - Mention any wrapper-specific behavior or adaptations the - canonical content has for this particular agent target. +9. **Wrapper observation** (only if you set `likely_wrapper: + true` in the frontmatter — see protocol below) — describe + what made you suspect this is a wrapper, where the actual + content appears to live (`wrapper_points_to`), and how + confident you are. The operator will surface this to the user, + who decides whether to re-vendor the referenced content and + re-research, treat the wrapper as the skill, or cancel. ## Reasoning protocol @@ -132,6 +124,39 @@ Body (markdown) covering: the secondary aspect in the body. Don't invent classification subtypes. +## Wrapper detection + +While reading the source, judge whether what you're looking at is +**likely a thin wrapper file** rather than the actual skill +content. Common patterns: + +- The SKILL.md body is short and mostly consists of references to + other files ("see X for the details", "uses content from Y", + pointers to a `content.md` or similar) +- Frontmatter has fields like `reference:`, `source:`, `canonical:`, + `extends:`, or similar that point at another file +- The plugin layout suggests a multi-agent structure — several + SKILL.md files in the same plugin, each appearing to thinly + wrap the same underlying content with agent-specific surface +- The skill's substantive content is clearly elsewhere (linked + files dwarf the SKILL.md, the SKILL.md says "implementation + in X") + +If you judge this to be likely a wrapper, set `likely_wrapper: +true` in the frontmatter and record where the actual content +appears to live in `wrapper_points_to`. Add a body section 9 +("Wrapper observation") explaining the signals and your +confidence. + +This is your judgment — there's no mechanical signal that's +definitive. If you're not sure, err toward `likely_wrapper: +false` and proceed; the operator will catch obvious wrappers in +review of your output. The operator surfaces a `true` finding to +the user, who decides whether to re-vendor the referenced +content, treat the wrapper as the skill, or cancel and provide a +different source. **Do NOT try to follow the pointer yourself or +re-vendor** — that's the operator's job after the user gate. + ## Return summary After writing `${OUTPUT_PATH}`, return a brief summary: @@ -139,6 +164,9 @@ After writing `${OUTPUT_PATH}`, return a brief summary: - Classification - 3-5 key responsibilities (the ones step 2 will design probes for) - Any PR submission notes captured +- **Wrapper finding** (if any): `likely_wrapper: true` + one-line + rationale + `wrapper_points_to` value. Operator needs this to + surface the user gate. Keep it under 200 words — the operator session reads this for the handoff message; full detail is in the report. diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index 65904a0..1befbe2 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -51,11 +51,10 @@ upstream_repo: / upstream_branch_target: license: requires_cla: true | false -entry_file_pattern: true | false -canonical_target_repo: / | -canonical_target_path: +pr_target_repo: / +pr_target_path: linked_consumers: - - : # other places that reference the canonical content + - : # other SKILL.md files referencing the same content - ... --- ``` @@ -66,35 +65,37 @@ linked_consumers: branch name. - **`requires_cla`** — `true` if the upstream requires a Contributor License Agreement. -- **`entry_file_pattern`** — `true` if the source SKILL.md is a - thin wrapper that points to the actual content elsewhere - (common: an agent-agnostic content file referenced by several - agent-flavored SKILL.md wrappers). The downstream PR composer - needs to know this to target the right file. -- **`canonical_target_repo` / `canonical_target_path`** — where - the actual skill content lives. Same as `upstream_repo` and the - source path if it's not an entry-file pattern; otherwise the - canonical location. +- **`pr_target_repo` / `pr_target_path`** — where the PR should + actually go. Often the same as the user-provided source. But if + B1 flagged the source as `likely_wrapper: true` and the user + opted to optimize the wrapper's referenced content (the common + case), the PR target is the referenced content's location, not + the wrapper's. The downstream PR composer needs this to target + the right file. - **`linked_consumers`** — other SKILL.md files (in the same repo - or other repos) that reference the same canonical content. If - non-empty, a change to the canonical content affects all of - them; the PR composer needs to coordinate or be honest about - the scope. + or other repos) that reference the same content as `pr_target_path`. + If non-empty, a change to that target may affect all of them; + the PR composer needs to coordinate or be honest about scope. Body sections: -1. **Source structure** (new) — is the input SKILL.md the actual - content, or an entry file pointing elsewhere? If entry-file: - document the relationship (entry-file path, canonical content - path, link mechanism). If you find linked consumers, list them - here with the relationship (other agent wrappers, downstream - skills that import from this one). -2. **Suggested PR target** (new) — based on (1), which location - should the PR modify? Usually the canonical content; sometimes - the entry file (if the change is wrapper-specific). If the - change would affect multiple consumers, flag whether they live - in the same repo (single PR) or different repos (multiple PRs - or coordination required). +1. **Source structure** (new) — read `01-functionality.md`'s + `likely_wrapper` field. If `true`, B1's subagent judged the + user-pointed source as a thin wrapper, and the user opted to + work with the referenced content (the common case after B1's + user gate). Document the wrapper-vs-target relationship here: + the wrapper path, the target path, and the link mechanism. If + you find linked consumers (other SKILL.md files referencing + the same target), list them with the relationship (other + agent wrappers, downstream skills that import from this one). + If `likely_wrapper: false`, the source is the actual content + and this section just states that. +2. **PR target** (new) — given the above, which file should the + PR modify? Usually `pr_target_path` is the same as B1's + source; for wrapper cases it's `wrapper_points_to`. If the + change would affect multiple consumers, flag whether they + live in the same repo (single PR suffices) or different + repos (multiple PRs or coordination required). 3. **License** — full SPDX identifier + brief plain-English summary (permissive / copyleft / proprietary). Document the license for the canonical content's repo (which may differ from the entry @@ -126,26 +127,31 @@ conventions. ## Reasoning protocol -1. **Read `01-functionality.md` for the entry-file context.** - B1 has already detected the entry-file pattern (if any) and - recorded the result in the report's frontmatter: - - `entry_file_pattern: true` → the source the user pointed at - was a wrapper; `canonical_source` records where the actual - content lives. The chain's `vendored-skill/` contains the - canonical content (B1's operator session already re-vendored - after confirming with the user). - - `entry_file_pattern: false` → the source is the actual - content; proceed normally. - - **Trust B1's detection** as primary input. Use the recorded - `canonical_source` to determine `canonical_target_repo` and - `canonical_target_path` for your own frontmatter. If - `01-functionality.md` was written before this feature existed - and doesn't have the field, fall back to detecting yourself - (signals: short SKILL.md body of "see X" pointers, frontmatter - fields like `reference:`/`source:`/`canonical:`, multi-agent - plugin layout). Surface to the operator if you detect a - pattern B1 missed — it suggests B1 needs a re-run. +1. **Read `01-functionality.md` for the wrapper context.** B1's + subagent judged whether the user-pointed source was likely a + wrapper and recorded the result in the report's frontmatter: + - `likely_wrapper: true` + `wrapper_points_to: ` → the + source is a thin wrapper. After B1's user gate, one of two + things happened: + - User opted to re-vendor the referenced content (common): + `vendored-skill/` now contains that content. Your PR + target is `wrapper_points_to` (or wherever it now lives + after the re-vendor). + - User opted to treat the wrapper as the skill: `vendored-skill/` + still contains the wrapper. Your PR target is the original + `skill_source`. + - `likely_wrapper: false` → the source is the actual content; + your PR target is `skill_source` as-is. + + Use these signals to determine `pr_target_repo` and + `pr_target_path` for your own frontmatter. If + `01-functionality.md` predates this feature and the + `likely_wrapper` field is missing, fall back to forming your + own judgment by reading the vendored content (signals: short + SKILL.md body of "see X" pointers, frontmatter fields like + `reference:`/`source:`/`canonical:`, multi-agent plugin + layout). Surface to the operator if you detect a wrapper B1 + missed — it suggests B1 needs a re-run. 2. **Find linked consumers.** Search for other places that reference the canonical content: @@ -176,15 +182,13 @@ conventions. 7. **Suggest the PR target.** Based on (1) and (2), recommend which file(s) the PR should modify: - - Not entry-file pattern: PR against the source as-is - - Entry-file pattern, single consumer: PR against the - canonical content - - Entry-file pattern, multiple consumers in the same repo: - PR against the canonical content; the wrappers update - automatically - - Entry-file pattern, multiple consumers across repos: PR - against the canonical content; flag that downstream wrappers - may need follow-up PRs in their own repos + - Not a wrapper: PR against the source as-is + - Wrapper, single consumer: PR against the referenced content + - Wrapper, multiple consumers in the same repo: PR against + the referenced content; the wrappers update automatically + - Wrapper, multiple consumers across repos: PR against the + referenced content; flag that downstream wrappers may need + follow-up PRs in their own repos 8. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** (e.g., "include the vendor's CLA requirement explicitly" → @@ -208,8 +212,9 @@ blocker rather than making up content: After writing `${OUTPUT_PATH}`, return a brief summary: -- Entry-file pattern detected? If yes: canonical location + - linked consumers count + suggested PR target +- Wrapper context (from B1 or fallback detection): is the + source a wrapper? If yes: referenced-content location, linked + consumers count, suggested PR target - License + CLA requirement - Branch target rule - 2-3 most important rejection signals from the closed-PR sweep From efde54f01492ffca43b880a45d03fcef7d5e76af Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 07:01:46 -0500 Subject: [PATCH 046/121] =?UTF-8?q?docs(v1.4):=20swap=20chain=20steps=202/?= =?UTF-8?q?3=20=E2=80=94=20submissions=20before=20design-tests?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 2 is now skill-optimizer-investigate-submissions (PR-bound only); step 3 is the renamed skill-optimizer-design-tests (was skill-optimizer-investigate-test-case). PR research now logically follows step 1 immediately when PR-bound, before test design. Renamed: - skills/skill-optimizer-investigate-test-case/ → skill-optimizer-design-tests/ - skills/skill-optimizer-subagents/test-case-designer.md → test-designer.md - 02-test-proposals.md ↔ 02-submissions.md (file numbers follow new step numbers) Also fixed pre-existing step-header bugs from the validate-tests insertion (run-bench, analyze, improve, validate had stale "Step N" headers off by one). Co-Authored-By: Claude Opus 4.7 --- skills/skill-optimizer-analyze/SKILL.md | 4 +-- .../SKILL.md | 32 ++++++++--------- skills/skill-optimizer-improve/SKILL.md | 14 ++++---- .../SKILL.md | 14 +++++--- .../SKILL.md | 22 ++++++------ skills/skill-optimizer-run-bench/SKILL.md | 8 +++-- .../iteration-protocol.md | 14 ++++---- skills/skill-optimizer-shared/workflow.md | 36 +++++++++---------- skills/skill-optimizer-subagents/optimizer.md | 6 ++-- .../research-functionality.md | 8 ++--- .../research-submissions.md | 8 ++--- ...test-case-designer.md => test-designer.md} | 12 +++---- .../test-validator.md | 2 +- .../skill-optimizer-subagents/test-writer.md | 2 +- skills/skill-optimizer-subagents/validator.md | 8 ++--- skills/skill-optimizer-validate/SKILL.md | 12 +++---- skills/skill-optimizer-write-tests/SKILL.md | 16 ++++----- 17 files changed, 109 insertions(+), 109 deletions(-) rename skills/{skill-optimizer-investigate-test-case => skill-optimizer-design-tests}/SKILL.md (82%) rename skills/skill-optimizer-subagents/{test-case-designer.md => test-designer.md} (94%) diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md index 0d8f483..9a5336a 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -5,7 +5,7 @@ description: Use when the user wants to diagnose why a bench run produced failur # skill-optimizer-analyze -Step 6 of the skill-optimizer chain. **Fresh-derivation step.** Takes +Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes the bench summary + raw trial output from step 6, dispatches an analyzer subagent to cluster failures into named **structural weaknesses** of the skill (or explicitly say there are none), and @@ -60,7 +60,7 @@ Full body template and reasoning protocol in If `overall_pass_rate == 1.0`: there's nothing to analyze. Surface honestly — either accept that probes don't expose a weakness, or -re-run step 2 with a "make probes harder" directive. Don't run +re-run step 3 with a "make probes harder" directive. Don't run the analyzer; there are no failures to cluster. If the bench results dir is missing or `suite-result.json` is diff --git a/skills/skill-optimizer-investigate-test-case/SKILL.md b/skills/skill-optimizer-design-tests/SKILL.md similarity index 82% rename from skills/skill-optimizer-investigate-test-case/SKILL.md rename to skills/skill-optimizer-design-tests/SKILL.md index 1e4d575..7da4fff 100644 --- a/skills/skill-optimizer-investigate-test-case/SKILL.md +++ b/skills/skill-optimizer-design-tests/SKILL.md @@ -1,11 +1,11 @@ --- -name: skill-optimizer-investigate-test-case -description: Use when the user wants to design or propose test cases for a skill — phrases like "design tests for this skill", "propose test cases", "what should we test", "plan test coverage for this skill". Also triggers mid-way through skill-optimizer chain work, once a functionality report exists and the next thing is figuring out what to test. Use even when the user doesn't explicitly say "design" — any phrasing about figuring out what tests to build for a skill should trigger this. +name: skill-optimizer-design-tests +description: Use when the user wants to design or propose test cases for a skill — phrases like "design tests for this skill", "propose test cases", "what should we test", "plan test coverage for this skill". Also triggers mid-way through skill-optimizer chain work, once a functionality report exists (and submissions research, if PR-bound) and the next thing is figuring out what to test. Use even when the user doesn't explicitly say "design" — any phrasing about figuring out what tests to build for a skill should trigger this. --- -# skill-optimizer-investigate-test-case +# skill-optimizer-design-tests -Step 2 of the skill-optimizer chain. **Maintenance step.** Takes the +Step 3 of the skill-optimizer chain. **Maintenance step.** Takes the functionality report from step 1, dispatches a designer subagent to enumerate the skill's responsibilities and propose a ranked set of **functionalities** to test (each functionality = one responsibility @@ -14,7 +14,7 @@ actually build probes for at step 4. The filesystem IS the state: this step writes a `tests//spec.yaml` per proposed functionality -plus a one-time audit report at `02-test-proposals.md`. No +plus a one-time audit report at `03-test-proposals.md`. No `picked: []` array anywhere; each spec.yaml has its own `picked: true|false`. @@ -22,7 +22,7 @@ plus a one-time audit report at `02-test-proposals.md`. No Two artifacts at `docs/skill-optimizer//`: -1. **`02-test-proposals.md`** — one-time audit report with the +1. **`03-test-proposals.md`** — one-time audit report with the ranked list of proposed functionalities + full reasoning. For human review of the design reasoning; downstream steps do NOT read it. The subagent rewrites it on re-runs. @@ -48,9 +48,9 @@ Two artifacts at `docs/skill-optimizer//`: Step 4 builds probes only for functionalities whose `spec.yaml` has `picked: true`. -For the body template of `02-test-proposals.md` and the exact +For the body template of `03-test-proposals.md` and the exact `spec.yaml` field list, see -[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md). +[`skills/skill-optimizer-subagents/test-designer.md`](../skill-optimizer-subagents/test-designer.md). ## Workflow @@ -76,13 +76,13 @@ If the user's directives contradict each other (e.g., "focus on X" combined with "ignore X"), surface the contradiction before re-dispatching; don't try to resolve it yourself. -### (c) Dispatch the test-case-designer subagent +### (c) Dispatch the test-designer subagent Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT enumerate responsibilities or design proposals yourself in this session** — dispatch the subagent via the `Agent` tool. Load the prompt template at -[`skills/skill-optimizer-subagents/test-case-designer.md`](../skill-optimizer-subagents/test-case-designer.md) +[`skills/skill-optimizer-subagents/test-designer.md`](../skill-optimizer-subagents/test-designer.md) and substitute `${FUNCTIONALITY_PATH}`, `${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, `${OPERATOR_DIRECTIVES}`. @@ -105,7 +105,7 @@ regression-defensive ones. The subagent writes: -- `02-test-proposals.md` (the audit report) +- `03-test-proposals.md` (the audit report) - One `tests//spec.yaml` per proposed functionality. Existing spec.yaml files the user has already edited (e.g., `picked: true` set) are preserved verbatim unless @@ -116,7 +116,7 @@ preserved-unchanged list, newly-added list. ### (d) Confirm subagent output -Verify `02-test-proposals.md` parses and each new +Verify `03-test-proposals.md` parses and each new `tests//spec.yaml` parses with required fields. If any spec.yaml fails to parse, surface to the user; don't repair the subagent's output yourself. @@ -129,7 +129,7 @@ step 1 yourself. ### (e) User gate: present proposals, collect picks -Show the user the ranked list from `02-test-proposals.md` and tell +Show the user the ranked list from `03-test-proposals.md` and tell them how to indicate picks: edit `picked: true|false` in each `tests//spec.yaml`. They can also edit `suggested_probes` or ask for a revised proposal. @@ -156,11 +156,7 @@ paying for probe-building in step 4 and the bench run in step 6). ### (f) Hand off -Read `01-functionality.md`'s `pr_submission_intent` field. If -`true`, hand off to `skill-optimizer-investigate-submissions` then -`skill-optimizer-write-tests`. If `false`, skip step 3 and hand -off directly to `skill-optimizer-write-tests`. No late "submit a -PR?" prompts — the decision was made at step 1. +> Next, invoke `skill-optimizer-write-tests`. ## Edge cases diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index aa69c0e..b71d27e 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -5,7 +5,7 @@ description: Use when the user wants to improve a skill based on identified stru # skill-optimizer-improve -Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes +Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes the named structural weaknesses from step 7, dispatches an optimizer subagent to draft a principled fix, and writes `08-improvement-proposal.md`. The proposal is then validated @@ -19,7 +19,7 @@ forcing a fix is the ducktape failure mode the chain is built to prevent. **Single-shot per invocation.** No in-step revision loop. If step -8's validator returns `needs-revision`, the operator (or +9's validator returns `needs-revision`, the operator (or auto-pilot at step 10) distills the validator's rationale into a directive and re-invokes this step. @@ -63,8 +63,8 @@ Three checks: tell the user to run `skill-optimizer-analyze` first. 2. **Anti-ducktape gate:** `has_structural_weakness: true` must be - set. If `false`, REFUSE — print: "Step 6 found no structural - weakness. Step 7 won't fire — nothing principled to optimize. + set. If `false`, REFUSE — print: "Step 7 found no structural + weakness. Step 8 won't fire — nothing principled to optimize. If you disagree, re-invoke step 7 with a directive; if you agree, exit honestly." Do NOT proceed. @@ -73,9 +73,9 @@ Three checks: from prior step-9 approvals), else `vendored-skill/` (B1 vendored the source regardless of upstream/local). -If `pr_submission_intent: true`, `03-submissions.md` should exist +If `pr_submission_intent: true`, `02-submissions.md` should exist so the optimizer can shape the diff to upstream conventions from -the start. If missing, ask whether to run step 3 first — step 9's +the start. If missing, ask whether to run step 2 first — step 9's external check still runs if it appears later. ### (b) Handle iteration @@ -105,7 +105,7 @@ and substitute `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, The optimizer sees: `07-analysis.md`; `01-functionality.md`; `${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else -`vendored-skill/`; `03-submissions.md` if PR-bound; +`vendored-skill/`; `02-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. The optimizer does NOT see: raw failed trials, `findings.txt`, diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 3e1549a..b2012b1 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -176,8 +176,12 @@ referenced content into it, and re-invoke this skill from (c). If (2): keep the report as-is and proceed. If (3): exit; the user will re-invoke with a new source. -Once resolved, hand off: - -> Next, invoke `skill-optimizer-investigate-test-case`. Note: this -> report's `pr_submission_intent` field tells step 2's handoff -> whether step 3 should run. +Once resolved, hand off based on this report's +`pr_submission_intent` field: + +- **`true`** — next, invoke `skill-optimizer-investigate-submissions` + (step 2). After that, the chain proceeds to + `skill-optimizer-design-tests` (step 3). +- **`false`** — skip step 2 and invoke `skill-optimizer-design-tests` + (step 3) directly. No late "submit a PR?" prompts; the decision was + recorded above. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 97c658d..3aeb466 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -5,17 +5,17 @@ description: Use when the user wants to research a skill's upstream PR conventio # skill-optimizer-investigate-submissions -Step 3 of the skill-optimizer chain — OPTIONAL, runs only when the +Step 2 of the skill-optimizer chain — OPTIONAL, runs only when the target skill is bound for upstream PR submission. **Fresh-derivation step.** Takes the source slug, dispatches a researcher subagent that uses the `gh` CLI to gather the upstream repo's contribution -conventions, and writes `docs/skill-optimizer//03-submissions.md` +conventions, and writes `docs/skill-optimizer//02-submissions.md` — the verbatim-pastable context block the validator (step 9) uses for its external consistency check. ## What you produce -A single report at `docs/skill-optimizer//03-submissions.md`. +A single report at `docs/skill-optimizer//02-submissions.md`. Frontmatter (runtime-relevant facts only, per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): @@ -31,8 +31,8 @@ requires_cla: true | false - **`upstream_branch_target`** — some repos use `main` for incremental changes and `next` for new skills. The subagent - determines this from recent merged PRs so the validator can - check the proposed PR targets the correct branch. + determines this from recent merged PRs so the validator (step 9) + can check the proposed PR targets the correct branch. - **`requires_cla`** — true if the upstream requires a Contributor License Agreement before merging. @@ -52,7 +52,7 @@ Two prerequisites: not, tell the user to run `skill-optimizer-investigate-functionality` first. 2. That report's `pr_submission_intent` field must be `true`. If - `false`, this skill should not run — tell the user step 3 is + `false`, this skill should not run — tell the user step 2 is skipped for local-only optimization runs. ### (b) Handle iteration @@ -83,7 +83,7 @@ signals, repo-file API, `CONTRIBUTING.md`, license file, existing skill files); the skill slug being researched; `${OPERATOR_DIRECTIVES}`; output path. -The subagent does NOT see: its own prior `03-submissions.md` or +The subagent does NOT see: its own prior `02-submissions.md` or git history of it; any information about the proposed change being optimized; existing analyses, tests, or failure data; the vendored skill source. @@ -124,11 +124,9 @@ be merged. Report the file path, a one-line summary (license / CLA / branch target / any flagged blockers), then: -> The validator in `skill-optimizer-validate` will read -> this report for its external consistency check. If you haven't -> run `skill-optimizer-write-tests` yet, invoke that next. - -The chain doesn't enforce ordering between step 3 and step 4. +> Next, invoke `skill-optimizer-design-tests`. The validator in +> `skill-optimizer-validate` will later read this report for its +> external consistency check. ## Edge cases diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 2501f9e..216fcb0 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -5,7 +5,7 @@ description: Use when the user wants to run the eval suite against a skill and c # skill-optimizer-run-bench -Step 5 of the skill-optimizer chain. **Fresh-derivation step (for +Step 6 of the skill-optimizer chain. **Fresh-derivation step (for the summary).** Invokes the skill-optimizer CLI's `run-suite` command against `tests/suite.yml` (generated by step 4), captures the raw results under a timestamped directory, and writes a small @@ -108,6 +108,8 @@ misconfiguration or environmental failure rather than a skill weakness. Write the summary honestly and surface before handing off to step 7. +Don't auto-invoke step 7. + ### (e) Hand off Read `overall_pass_rate`. Two messages: @@ -115,8 +117,8 @@ Read `overall_pass_rate`. Two messages: - `< 1.0`: invoke `skill-optimizer-analyze` to diagnose failures. - `== 1.0`: surface the choice — accept that probes don't expose a - weakness, or re-run step 2 with a "make probes harder" - directive. Don't auto-invoke step 7. + weakness, or re-run step 3 with a "make probes harder" + directive. ## Edge cases diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index 9996e43..c73d9b2 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -35,7 +35,7 @@ upstream + directives. Re-runs overwrite; git captures prior state. | Step | Why fresh derivation | |---|---| | 1. investigate-functionality | Each invocation researches from source — no continuity needed | -| 3. investigate-submissions | Each invocation researches upstream — no continuity needed | +| 2. investigate-submissions | Each invocation researches upstream — no continuity needed | | 5. validate-tests | Each probe judged fresh against its parent functionality; load-bearing for anti-ducktape | | 6. run-bench (summary file) | Mechanical write-up of the new raw bench run, not an extension of prior summary | | 7. analyze | Must not be biased by prior analyses; load-bearing for anti-ducktape | @@ -48,7 +48,7 @@ extend or modify it. | Step | What accumulates | |---|---| -| 2. investigate-test-case | `tests//spec.yaml` grows as functionalities are added/refined | +| 3. design-tests | `tests//spec.yaml` grows as functionalities are added/refined | | 4. write-tests | `tests///` probes grow as the user adds coverage | ## Staleness detection (git-native) @@ -63,7 +63,7 @@ DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) ``` In interactive use, the operator typically just knows ("I re-ran -step 1, so step 2 needs a re-run"). The git-mtime check is for +step 1, so step 3 needs a re-run"). The git-mtime check is for auto-pilot (step 10), which walks the chain forward and re-runs any downstream older than its direct upstream. @@ -75,7 +75,7 @@ one. ## Safe destructive edits When a maintenance step's re-run will overwrite or delete existing -content (e.g., step 4 rebuilds a probe; step 2 removes a de-picked +content (e.g., step 4 rebuilds a probe; step 3 removes a de-picked functionality), the operator session commits the current state **before** dispatching the destructive change: @@ -97,9 +97,9 @@ its own prior commit, so no checkpoint is needed. Each step checks **direct upstream only** — the immediate prior artifact(s) it consumes, not the whole upstream chain. If step 1 -is updated but step 2 is not re-run (the user judged step 2 still -valid against the new step 1), step 3 will see step 2 as current -even though step 2's git mtime is older than step 1's. +is updated but step 3 is not re-run (the user judged step 3 still +valid against the new step 1), step 4 will see step 3 as current +even though step 3's git mtime is older than step 1's. The discipline: skipping a step's re-run is an **explicit operator judgment** that the existing artifact is still valid against the diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index 30efdb2..506715e 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -10,19 +10,19 @@ re-runs, and the relationship between steps. | # | Skill | Kind | Input | Output | |---|---|---|---|---| | 1 | `investigate-functionality` | fresh-derivation | source skill (URL or local) | `01-functionality.md` | -| 2 | `investigate-test-case` | maintenance | step 1 | `02-test-proposals.md`, `tests//spec.yaml` | -| 3 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `03-submissions.md` | -| 4 | `write-tests` | maintenance | steps 1+2 | `tests///`, `tests/suite.yml` | +| 2 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `02-submissions.md` | +| 3 | `design-tests` | maintenance | step 1 | `03-test-proposals.md`, `tests//spec.yaml` | +| 4 | `write-tests` | maintenance | steps 1+3 | `tests///`, `tests/suite.yml` | | 5 | `validate-tests` | fresh-derivation | step 4 + step 1 + skill | `05-tests-verdict.md` | | 6 | `run-bench` | fresh-derivation (summary) | step 5 (must approve) + skill | `06-bench-results//`, `06-bench-summary.md` | | 7 | `analyze` | fresh-derivation | step 6 + skill | `07-analysis.md` | -| 8 | `improve` | fresh-derivation | step 7 + skill (+ step 3 if PR-bound) | `08-improvement-proposal.md` | -| 9 | `validate` | fresh-derivation | step 8 + skill (+ step 3 if PR-bound) | `09-validator-verdict.md`, `improved-skill/` (on approve) | +| 8 | `improve` | fresh-derivation | step 7 + skill (+ step 2 if PR-bound) | `08-improvement-proposal.md` | +| 9 | `validate` | fresh-derivation | step 8 + skill (+ step 2 if PR-bound) | `09-validator-verdict.md`, `improved-skill/` (on approve) | | 10 | `autopilot` | chain driver | same as step 1 + flags | `autopilot-summary-.md` | The PR-or-not decision is made ONCE at step 1; subsequent steps know from `01-functionality.md`'s `pr_submission_intent` field -whether step 3 will run. No late prompts. +whether step 2 will run. No late prompts. Two independent validators in the chain: step 5 (validate-tests) checks that probes fairly test their functionality before @@ -42,19 +42,19 @@ when to re-run each step. Common triggers: | Step | Re-run when... | |---|---| | 1 | source URL changed; PR-intent changed; user wants fresh research with new directives | -| 2 | user wants different coverage; step 4/5/6/7 surfaced a coverage gap; step 1 changed | -| 3 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed | -| 4 | step 2's `tests/` tree changed; a probe's smoke check failed; step 5 flagged probes for revision; step 7 showed probes systematically too easy/hard | +| 2 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed | +| 3 | user wants different coverage; step 4/5/6/7 surfaced a coverage gap; step 1 changed | +| 4 | step 3's `tests/` tree changed; a probe's smoke check failed; step 5 flagged probes for revision; step 7 showed probes systematically too easy/hard | | 5 | step 4 produced new or revised probes; user disagrees with prior test verdict | | 6 | step 4's `tests/` changed (and step 5 re-approved); user wants fresh trial data (flakiness, model list changed); step 7 wants more trials | | 7 | step 6 produced new results; user disagrees with prior analysis; step 8 was unable to address a named weakness | | 8 | step 7 produced a new analysis; step 9 returned `needs-revision`/`reject` with a distillable rationale; user wants a different approach | -| 9 | step 8 produced a new proposal; user disagrees with prior verdict; step 3 was updated and prior external check is stale | +| 9 | step 8 produced a new proposal; user disagrees with prior verdict; step 2 was updated and prior external check is stale | | 10 | self-iterable; picks up at whatever step is stale per git-mtime | Each skill checks **direct upstream only** for staleness — if step -1 went stale but step 2 wasn't re-run (user judged it still -valid), step 3 sees step 2 as current. Skipping a step's re-run +1 went stale but step 3 wasn't re-run (user judged it still +valid), step 4 sees step 3 as current. Skipping a step's re-run is an explicit operator judgment. ## Backward triggers @@ -64,10 +64,10 @@ auto-fire; they're surfaced to the user. | At step | Surfaced finding | Resolution path | |---|---|---| -| 2 | proposal is thin (subagent found few responsibilities) | re-run step 1 with directives, OR accept | -| 4 | test-writer reports BLOCKED on a probe | re-run step 2 to refine the functionality spec | +| 3 | proposal is thin (subagent found few responsibilities) | re-run step 1 with directives, OR accept | +| 4 | test-writer reports BLOCKED on a probe | re-run step 3 to refine the functionality spec | | 5 | any probe is `needs-revision`/`reject` | distill verdict into directives, re-run step 4 for affected probes, re-run step 5 | -| 6 | bench all-pass (probes too easy) | re-run step 2 with "make probes harder" directive | +| 6 | bench all-pass (probes too easy) | re-run step 3 with "make probes harder" directive | | 7 | analyzer's `has_structural_weakness: false`, user disagrees | re-run step 7 with directive pointing at missed cluster | | 8 | optimizer reports BLOCKED (weakness too abstract) | re-run step 7 to reformulate weakness | | 9 | `needs-revision` | distill validator's rationale, re-run step 8, re-run step 9 | @@ -78,10 +78,11 @@ auto-fire; they're surfaced to the user. ```text docs/skill-optimizer// 01-functionality.md - 02-test-proposals.md # B2's audit report + 02-submissions.md # only if step 2 ran (PR-bound) + 03-test-proposals.md # B3's audit report tests/ / - spec.yaml # B2 writes + spec.yaml # B3 writes / # B4 writes one per probe spec.yaml workspace/ @@ -89,7 +90,6 @@ docs/skill-optimizer// smoke/{good,bad,empty}/ checks/smoke.mjs suite.yml # B4 generates - 03-submissions.md # only if step 3 ran 05-tests-verdict.md # B5 writes (per-probe + aggregate) 06-bench-results// # raw, timestamped per run 06-bench-summary.md # single canonical diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md index 1290c2d..3bcea72 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -17,7 +17,7 @@ skill (that's step 9's job after the validator approves). step-9 approvals), else `vendored-skill/` (B1 vendored the source regardless of upstream/local). You propose a new improvement on top of whatever current state you read. -- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (lets +- `${SUBMISSIONS_PATH}` — `02-submissions.md` if PR-bound (lets you shape the diff to match upstream conventions from the start, reducing step-9 round-trips) - `${PROPOSAL_OUTPUT_PATH}` — where to write @@ -31,7 +31,7 @@ skill (that's step 9's job after the validator approves). - `07-analysis.md` (named weaknesses + principles + anti-patterns) - `01-functionality.md` - The current skill content at `${SKILL_CURRENT_PATH}` -- `03-submissions.md` if PR-bound +- `02-submissions.md` if PR-bound - Operator directives ## What you do NOT see @@ -111,7 +111,7 @@ restatement of the rule"). 3. **Read the current skill** at `${SKILL_CURRENT_PATH}`. Locate the section(s) named in each weakness's "Connects to skill section" field. That's where the change goes. -4. **If `03-submissions.md` exists** (PR-bound), read the +4. **If `02-submissions.md` exists** (PR-bound), read the frontmatter spec, file-location conventions, and prefix taxonomy. Shape the diff to fit. The validator at step 9 will check this; getting it right now saves a round-trip. diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index 2010925..1c460d1 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -81,7 +81,7 @@ Body (markdown) covering: field). 4. **Responsibilities** — bulleted list of distinct things the skill is supposed to make the agent do. Each one is a candidate - for testing at step 2. + for testing at step 3. 5. **Tools / dependencies** — what the skill assumes is available (specific MCPs, CLIs, libraries, file conventions). 6. **Key terminology** — domain terms the operator must understand @@ -92,7 +92,7 @@ Body (markdown) covering: `true` AND source was local) — verbatim record of what the user said at step 1 about where the upstream contribution guidelines live (URL, CONTRIBUTING.md path, Slack channel, - whatever they provided). Step 3 uses this as its starting + whatever they provided). Step 2 uses this as its starting point. 9. **Wrapper observation** (only if you set `likely_wrapper: true` in the frontmatter — see protocol below) — describe @@ -110,7 +110,7 @@ Body (markdown) covering: 2. **Identify the responsibilities by enumeration**, not by paraphrase. If the SKILL.md says "Do X, then Y, then Z", those are three responsibilities, not one. The test-case designer at - step 2 needs distinct items to propose probes for. + step 3 needs distinct items to propose probes for. 3. **Web-search the technology** only to fill gaps the skill doesn't explain. Don't re-document everything the technology does — focus on what an agent needs to know to USE the skill @@ -162,7 +162,7 @@ re-vendor** — that's the operator's job after the user gate. After writing `${OUTPUT_PATH}`, return a brief summary: - Classification -- 3-5 key responsibilities (the ones step 2 will design probes for) +- 3-5 key responsibilities (the ones step 3 will design probes for) - Any PR submission notes captured - **Wrapper finding** (if any): `likely_wrapper: true` + one-line rationale + `wrapper_points_to` value. Operator needs this to diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index 1befbe2..a5be7ce 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -2,7 +2,7 @@ You are dispatched by `skill-optimizer-investigate-submissions` to research the upstream repo's PR conventions and produce -`03-submissions.md` — the verbatim-pastable context block the +`02-submissions.md` — the verbatim-pastable context block the validator (step 9) uses for its external consistency check. ## Inputs (templated by the operator session) @@ -15,7 +15,7 @@ validator (step 9) uses for its external consistency check. - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new requirements (may be empty) - `${OUTPUT_PATH}` — where to write the report (typically - `docs/skill-optimizer//03-submissions.md`) + `docs/skill-optimizer//02-submissions.md`) ## What you see @@ -29,7 +29,7 @@ validator (step 9) uses for its external consistency check. ## What you do NOT see -- Your own prior `03-submissions.md` or git history of it +- Your own prior `02-submissions.md` or git history of it - Any information about the proposed change being optimized — the report is purely upstream facts, not advocacy for a change - Existing analyses, tests, failure data from the chain @@ -41,7 +41,7 @@ independent. If you absorb optimization context, you bias the report toward justifying the proposed change — and the validator loses its real independence. -## Output: `03-submissions.md` +## Output: `02-submissions.md` Frontmatter (runtime-relevant only): diff --git a/skills/skill-optimizer-subagents/test-case-designer.md b/skills/skill-optimizer-subagents/test-designer.md similarity index 94% rename from skills/skill-optimizer-subagents/test-case-designer.md rename to skills/skill-optimizer-subagents/test-designer.md index 9dcf4fa..b0cff2f 100644 --- a/skills/skill-optimizer-subagents/test-case-designer.md +++ b/skills/skill-optimizer-subagents/test-designer.md @@ -1,10 +1,10 @@ -# Test-case-designer subagent +# Test-designer subagent -You are dispatched by `skill-optimizer-investigate-test-case` to +You are dispatched by `skill-optimizer-design-tests` to enumerate the skill's responsibilities (from `01-functionality.md`) and propose a ranked set of **functionalities** to test, then write a per-functionality `tests//spec.yaml` for each -plus a one-time audit report `02-test-proposals.md`. +plus a one-time audit report `03-test-proposals.md`. A "functionality" here means **one distinct responsibility the skill must fulfill**. It's what step 4 will build probes for (each @@ -18,7 +18,7 @@ angles). on first invocation; on re-runs it has existing `tests//spec.yaml` files) - `${PROPOSALS_PATH}` — where to write the audit report (typically - `docs/skill-optimizer//02-test-proposals.md`) + `docs/skill-optimizer//03-test-proposals.md`) - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new requirements (may be empty) @@ -49,7 +49,7 @@ angles). **Two artifacts:** -### `${PROPOSALS_PATH}` — `02-test-proposals.md` +### `${PROPOSALS_PATH}` — `03-test-proposals.md` The audit report. Operator + user read this for design reasoning; downstream steps do NOT read it. Rewrite top-to-bottom on each @@ -94,7 +94,7 @@ why_test: > ``` **`picked: false` is the default for new functionalities.** The -user gate at step 2 lets the user flip to `true` for the ones they +user gate at step 3 lets the user flip to `true` for the ones they want built. Don't pre-pick. **Existing spec.yaml files are preserved verbatim** unless a diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/skill-optimizer-subagents/test-validator.md index 72ac885..a89fc46 100644 --- a/skills/skill-optimizer-subagents/test-validator.md +++ b/skills/skill-optimizer-subagents/test-validator.md @@ -171,7 +171,7 @@ the agent — that's unfair. functionality. The probe tests the wrong thing, or the functionality as named can't be cleanly tested. Rejection surfaces to the operator who decides whether to re-frame at - step 2 or accept the gap. + step 3 or accept the gap. ## Return summary diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md index 8068513..f2edb38 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -142,7 +142,7 @@ than building a misleading probe: - **Probe spec is too abstract to derive a concrete fixture** — e.g., "tests good code style" without specifying what code or - what style. Fix is at step 2 (more specific probe spec). + what style. Fix is at step 3 (more specific probe spec). - **Probe requires real-time API access** (live data, external services the workbench can't mock) — not testable in the static workbench. diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/skill-optimizer-subagents/validator.md index 65d5c7f..a1e19ee 100644 --- a/skills/skill-optimizer-subagents/validator.md +++ b/skills/skill-optimizer-subagents/validator.md @@ -26,7 +26,7 @@ without your own prior verdicts. - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the internal consistency check: does the change make sense given the skill's stated responsibilities?) -- `${SUBMISSIONS_PATH}` — `03-submissions.md` if PR-bound (for +- `${SUBMISSIONS_PATH}` — `02-submissions.md` if PR-bound (for the external consistency check) - `${VERDICT_OUTPUT_PATH}` — where to write `09-validator-verdict.md` @@ -41,7 +41,7 @@ without your own prior verdicts. optimizer claims to have done + their rationale + their self- check against anti-patterns - `01-functionality.md` (stated responsibilities) -- `03-submissions.md` if PR-bound (upstream conventions) +- `02-submissions.md` if PR-bound (upstream conventions) - Operator directives ## What you do NOT see @@ -104,7 +104,7 @@ checks: ### 2. External consistency check (only if `${SUBMISSIONS_PATH}` exists) Forward-looking: would a PR carrying this change conform to the -upstream conventions in `03-submissions.md`? Even though step 9 +upstream conventions in `02-submissions.md`? Even though step 9 doesn't produce a PR, this check verifies the improved skill COULD be turned into a valid PR by a downstream composer. @@ -160,7 +160,7 @@ directive for step 8. overlap with hand-waving justification. 5. **Read `01-functionality.md`** — does the change preserve the skill's stated responsibilities? -6. **If `03-submissions.md` exists**, walk the external +6. **If `02-submissions.md` exists**, walk the external consistency check. 7. **Choose verdict.** The weakest of the checks. Be honest about verdicts — false approvals let ducktape ship; false rejects diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index d21a755..0a58b5b 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -5,7 +5,7 @@ description: Use when the user wants to validate an improvement proposal from `s # skill-optimizer-validate -Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes +Step 9 of the skill-optimizer chain. **Fresh-derivation step.** Takes the proposal from step 8, dispatches a validator subagent to independently check whether the proposed change is sound (internal consistency) and conformant (external PR conventions if PR-bound), @@ -37,7 +37,7 @@ One or two artifacts at `docs/skill-optimizer//`: Body covers the internal consistency check (does the change make sense for the named weakness? additive vs. destructive? - general vs. ducktape?) and — if `03-submissions.md` exists — + general vs. ducktape?) and — if `02-submissions.md` exists — the external consistency check (conformant to upstream PR rules?). External check is forward-looking — verifies the improved skill COULD be turned into a valid PR, even though @@ -70,9 +70,9 @@ Three checks: against the new state. 3. `01-functionality.md` must exist. -If `pr_submission_intent: true`, `03-submissions.md` should +If `pr_submission_intent: true`, `02-submissions.md` should exist; if missing, the external check can't run. Ask whether to -run step 3 first or proceed with internal consistency only (the +run step 2 first or proceed with internal consistency only (the verdict body will note the omission). ### (b) Handle iteration @@ -103,7 +103,7 @@ copy; do NOT touch the canonical `improved-skill/` yet), The validator sees: skill BEFORE; skill AFTER (temporary materialization); `08-improvement-proposal.md`; -`01-functionality.md`; `03-submissions.md` if PR-bound; +`01-functionality.md`; `02-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. The validator does NOT see: raw failed trials, `findings.txt`, @@ -159,4 +159,4 @@ Three messages by verdict: did), the analyzer/validator are misaligned — step 7 reframe is the cleaner fix." -Don't auto-invoke step 7 or 7. +Don't auto-invoke step 7 or step 8. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 8361c3c..2fedd0a 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -1,12 +1,12 @@ --- name: skill-optimizer-write-tests -description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-investigate-test-case` has produced `tests//spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. +description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-design-tests` has produced `tests//spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. --- # skill-optimizer-write-tests Step 4 of the skill-optimizer chain. **Maintenance step.** Reads -the picked functionalities from step 2 +the picked functionalities from step 3 (`tests//spec.yaml` files where `picked: true`), decides the probe set per functionality, dispatches one test-writer subagent per probe (in parallel) to build the concrete workspace @@ -94,7 +94,7 @@ Three realistic responses: for rebuild (apply destructive-edit checkpoint), proceed to (d). 3. **User rejects the plan structure.** Surface and ask whether to - abandon (loop back to step 2 to revise functionality specs) or + abandon (loop back to step 3 to revise functionality specs) or retry with their feedback as directives. ### (d) Dispatch test-writer subagents (parallel, one per probe) @@ -128,7 +128,7 @@ logic could write fixtures that incidentally satisfy that grader too, making eval results look correlated when they're not. Source-content access is the one exception across the chain: probe -spec from step 2 fixes WHAT to test; source provides the HOW +spec from step 3 fixes WHAT to test; source provides the HOW (concrete violation patterns). Without source, the subagent would invent generic patterns that may not trigger the skill's rules. @@ -139,8 +139,8 @@ smoke-check result. If a test-writer reports BLOCKED (probe spec too abstract to derive a fixture from, or the probe is genuinely unimplementable in a static workbench), surface to the user. The fix is typically -a step 2 re-run with a more specific functionality spec; per the -no-auto-invocation rule, the user invokes step 2 explicitly. +a step 3 re-run with a more specific functionality spec; per the +no-auto-invocation rule, the user invokes step 3 explicitly. ### (e) Run smoke check @@ -158,9 +158,9 @@ fails, surface to the user. Two realistic responses: directive. Treat as a single-probe rebuild (back to (b) for the checkpoint, then (d)). 2. **Remove the probe** from this functionality's set, or de-pick - the parent functionality (a backward trigger to step 2; per + the parent functionality (a backward trigger to step 3; per the no-auto-invocation rule, surface the option and let the - user invoke step 2). + user invoke step 3). ### (f) Generate suite.yml + commit + hand off From 6866d1efc16d1f3cf86f808bb7ebde08d626fda2 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 07:13:18 -0500 Subject: [PATCH 047/121] docs(v1.4): replace internal B-number shorthand with step numbers B-prefix shorthand (B1, B2, ...) was internal brainstorming notation that leaked into shipped SKILL files. Sweeping it out so the chain documentation is self-explanatory to anyone who hasn't seen the v1.4 design discussions. Also fixed three additional stale step-number references found during the audit: - frontmatter-discipline.md gating-field summary was off by one to two steps (predated the validate-tests insertion) - analyze SKILL.md "Step 7 will refuse" handoff message named the wrong gating step (should be step 8, improve) - analyze SKILL.md "step-5 problem" should be "step-6 problem" for malformed bench output (run-bench is step 6) Co-Authored-By: Claude Opus 4.7 --- skills/skill-optimizer-analyze/SKILL.md | 6 +++--- skills/skill-optimizer-improve/SKILL.md | 2 +- .../SKILL.md | 4 ++-- skills/skill-optimizer-run-bench/SKILL.md | 2 +- .../frontmatter-discipline.md | 13 +++++++------ .../iteration-protocol.md | 2 +- .../subagent-dispatch.md | 2 +- skills/skill-optimizer-shared/workflow.md | 18 +++++++++--------- skills/skill-optimizer-subagents/analyzer.md | 2 +- skills/skill-optimizer-subagents/optimizer.md | 2 +- .../research-submissions.md | 18 +++++++++--------- .../test-validator.md | 2 +- .../skill-optimizer-subagents/test-writer.md | 2 +- skills/skill-optimizer-validate/SKILL.md | 2 +- 14 files changed, 39 insertions(+), 38 deletions(-) diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md index 9a5336a..83fbcc5 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -64,7 +64,7 @@ re-run step 3 with a "make probes harder" directive. Don't run the analyzer; there are no failures to cluster. If the bench results dir is missing or `suite-result.json` is -malformed, surface as a step-5 problem (incomplete or corrupted +malformed, surface as a step-6 problem (incomplete or corrupted bench run) and tell the user to re-run step 6. `tests/` and the source skill must also be available. @@ -133,9 +133,9 @@ Two messages depending on `has_structural_weakness`: - **`true`:** "Analysis complete. `` structural weakness(es) identified. Next, invoke `skill-optimizer-improve`." - **`false`:** "No structural weakness identified — failures - consistent with noise rather than a fixable defect. Step 7 will + consistent with noise rather than a fixable defect. Step 8 will refuse to fire. Either accept the conclusion, or re-invoke step - 6 with a directive ('the analyzer missed the X cluster')." + 7 with a directive ('the analyzer missed the X cluster')." If the subagent returned `false` but the user disagrees: surface the disagreement, but do NOT pressure the subagent to manufacture diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index b71d27e..e44b32b 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -70,7 +70,7 @@ Three checks: 3. `01-functionality.md` must exist. Skill content must be readable: `improved-skill/` if it exists (accumulated state - from prior step-9 approvals), else `vendored-skill/` (B1 + from prior step-9 approvals), else `vendored-skill/` (step 1 vendored the source regardless of upstream/local). If `pr_submission_intent: true`, `02-submissions.md` should exist diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index b2012b1..3ab8bb1 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -42,8 +42,8 @@ SKILL.md that mostly references content elsewhere, common in multi-agent plugins where one canonical agent-agnostic content file is wrapped by several agent-flavored SKILL.md files). When `true`, **`wrapper_points_to`** records where the subagent thinks -the actual content lives. Downstream chain steps (B3 submission -research, B8 optimization, B9 validation) consult these fields +the actual content lives. Downstream chain steps (step 2 submission +research, step 8 optimization, step 9 validation) consult these fields when deciding what to target. If the subagent flags `likely_wrapper: true`, the operator surfaces diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 216fcb0..6e8a410 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -44,7 +44,7 @@ Two outputs at `docs/skill-optimizer//`: `tests/suite.yml` must exist (per step 4's output contract). If missing, tell the user to complete step 4 first. -`vendored-skill/` should exist (B1 vendored the source regardless +`vendored-skill/` should exist (step 1 vendored the source regardless of upstream/local). ### (b) Handle iteration diff --git a/skills/skill-optimizer-shared/frontmatter-discipline.md b/skills/skill-optimizer-shared/frontmatter-discipline.md index 29a78c1..5434de3 100644 --- a/skills/skill-optimizer-shared/frontmatter-discipline.md +++ b/skills/skill-optimizer-shared/frontmatter-discipline.md @@ -22,13 +22,14 @@ git can't already do. Frontmatter fields a chain skill needs to READ at runtime to do its job: -- `picked: true|false` (B2 functionality spec) -- `pr_submission_intent: true|false` (B1 functionality report) -- `classification: tool-use` (B1 functionality report) -- `has_structural_weakness: true|false` (B6 analysis — gates B7) -- `verdict: approve|needs-revision|reject` (B8 verdict — gates +- `picked: true|false` (step 3 functionality spec) +- `pr_submission_intent: true|false` (step 1 functionality report) +- `classification: tool-use` (step 1 functionality report) +- `all_probes_approved: true|false` (step 5 tests verdict — gates step 6) +- `overall_pass_rate` (step 6 bench summary — gates step 7 dispatch) +- `has_structural_weakness: true|false` (step 7 analysis — gates step 8) +- `verdict: approve|needs-revision|reject` (step 9 verdict — gates materialization) -- `overall_pass_rate` (B5 summary — gates B6 dispatch) ## Decision aid diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index c73d9b2..d687c87 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -125,7 +125,7 @@ needed — git is the archive. `docs/skill-optimizer//autopilot-summary-.md` is timestamped per run. Each auto-pilot invocation produces a fresh summary; old ones are preserved naturally. -- **`vendored-skill/`** (the canonical input to the chain — B1 +- **`vendored-skill/`** (the canonical input to the chain — step 1 copies the source skill here regardless of upstream/local) is reused across iterations of the same slug unless: (1) the source URL changed (upstream), or (2) the user explicitly asks diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md index 842ea39..09336e7 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -55,7 +55,7 @@ Three rules every reasoning subagent must follow: canonical. Both the optimizer (step 8) and the validator (step 9) read the **current state of the skill** — which is `improved-skill/` if it exists (the accumulated state from - prior step-9 approvals), else `vendored-skill/` (B1 vendored + prior step-9 approvals), else `vendored-skill/` (step 1 vendored the source regardless of upstream/local). Neither step modifies the original. The "own canonical" off-limits to each subagent is its report (`08-improvement-proposal.md` for the optimizer; diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index 506715e..858f638 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -79,26 +79,26 @@ auto-fire; they're surfaced to the user. docs/skill-optimizer// 01-functionality.md 02-submissions.md # only if step 2 ran (PR-bound) - 03-test-proposals.md # B3's audit report + 03-test-proposals.md # step 3's audit report tests/ / - spec.yaml # B3 writes - / # B4 writes one per probe + spec.yaml # step 3 writes + / # step 4 writes one per probe spec.yaml workspace/ grader.mjs smoke/{good,bad,empty}/ checks/smoke.mjs - suite.yml # B4 generates - 05-tests-verdict.md # B5 writes (per-probe + aggregate) + suite.yml # step 4 generates + 05-tests-verdict.md # step 5 writes (per-probe + aggregate) 06-bench-results// # raw, timestamped per run 06-bench-summary.md # single canonical 07-analysis.md 08-improvement-proposal.md 09-validator-verdict.md - vendored-skill/ # frozen original (always — B1 vendors upstream OR local) - improved-skill/ # B9 materializes on approve - autopilot-summary-.md # B10 writes per run + vendored-skill/ # frozen original (always — step 1 vendors upstream OR local) + improved-skill/ # step 9 materializes on approve + autopilot-summary-.md # step 10 writes per run ``` `` is `--` for upstream skills, or @@ -120,4 +120,4 @@ need them: - [`frontmatter-discipline.md`](./frontmatter-discipline.md) — runtime facts vs. history rule - [`workbench.md`](./workbench.md) — workbench schema (probe - layout, grader contract, smoke-check format); referenced by B4 + layout, grader contract, smoke-check format); referenced by step 4 diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/skill-optimizer-subagents/analyzer.md index 1049ea1..23eaf9e 100644 --- a/skills/skill-optimizer-subagents/analyzer.md +++ b/skills/skill-optimizer-subagents/analyzer.md @@ -18,7 +18,7 @@ address it AND the anti-patterns that would NOT. `spec.yaml`** — what each probe was probing at the level of INTENT. Do NOT read `workspace/` contents (raw input fixtures). - `${SKILL_SOURCE_PATH}` — the skill's content at `vendored-skill/` - (B1 vendored the source regardless of upstream/local) + (step 1 vendored the source regardless of upstream/local) - `${OUTPUT_PATH}` — where to write `07-analysis.md` - `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless this is a re-run with sharper guidance) diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md index 3bcea72..e2d57da 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -14,7 +14,7 @@ skill (that's step 9's job after the validator approves). stated responsibilities — your change must not contradict them) - `${SKILL_CURRENT_PATH}` — the **current state of the skill**: `improved-skill/` if it exists (accumulated state from prior - step-9 approvals), else `vendored-skill/` (B1 vendored the + step-9 approvals), else `vendored-skill/` (step 1 vendored the source regardless of upstream/local). You propose a new improvement on top of whatever current state you read. - `${SUBMISSIONS_PATH}` — `02-submissions.md` if PR-bound (lets diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index a5be7ce..215278f 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -67,7 +67,7 @@ linked_consumers: Contributor License Agreement. - **`pr_target_repo` / `pr_target_path`** — where the PR should actually go. Often the same as the user-provided source. But if - B1 flagged the source as `likely_wrapper: true` and the user + step 1 flagged the source as `likely_wrapper: true` and the user opted to optimize the wrapper's referenced content (the common case), the PR target is the referenced content's location, not the wrapper's. The downstream PR composer needs this to target @@ -80,9 +80,9 @@ linked_consumers: Body sections: 1. **Source structure** (new) — read `01-functionality.md`'s - `likely_wrapper` field. If `true`, B1's subagent judged the + `likely_wrapper` field. If `true`, step 1's subagent judged the user-pointed source as a thin wrapper, and the user opted to - work with the referenced content (the common case after B1's + work with the referenced content (the common case after step 1's user gate). Document the wrapper-vs-target relationship here: the wrapper path, the target path, and the link mechanism. If you find linked consumers (other SKILL.md files referencing @@ -91,7 +91,7 @@ Body sections: If `likely_wrapper: false`, the source is the actual content and this section just states that. 2. **PR target** (new) — given the above, which file should the - PR modify? Usually `pr_target_path` is the same as B1's + PR modify? Usually `pr_target_path` is the same as step 1's source; for wrapper cases it's `wrapper_points_to`. If the change would affect multiple consumers, flag whether they live in the same repo (single PR suffices) or different @@ -127,11 +127,11 @@ conventions. ## Reasoning protocol -1. **Read `01-functionality.md` for the wrapper context.** B1's +1. **Read `01-functionality.md` for the wrapper context.** step 1's subagent judged whether the user-pointed source was likely a wrapper and recorded the result in the report's frontmatter: - `likely_wrapper: true` + `wrapper_points_to: ` → the - source is a thin wrapper. After B1's user gate, one of two + source is a thin wrapper. After step 1's user gate, one of two things happened: - User opted to re-vendor the referenced content (common): `vendored-skill/` now contains that content. Your PR @@ -150,8 +150,8 @@ conventions. own judgment by reading the vendored content (signals: short SKILL.md body of "see X" pointers, frontmatter fields like `reference:`/`source:`/`canonical:`, multi-agent plugin - layout). Surface to the operator if you detect a wrapper B1 - missed — it suggests B1 needs a re-run. + layout). Surface to the operator if you detect a wrapper step 1 + missed — it suggests step 1 needs a re-run. 2. **Find linked consumers.** Search for other places that reference the canonical content: @@ -212,7 +212,7 @@ blocker rather than making up content: After writing `${OUTPUT_PATH}`, return a brief summary: -- Wrapper context (from B1 or fallback detection): is the +- Wrapper context (from step 1 or fallback detection): is the source a wrapper? If yes: referenced-content location, linked consumers count, suggested PR target - License + CLA requirement diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/skill-optimizer-subagents/test-validator.md index a89fc46..edda94b 100644 --- a/skills/skill-optimizer-subagents/test-validator.md +++ b/skills/skill-optimizer-subagents/test-validator.md @@ -24,7 +24,7 @@ coincidentally match rather than truly distinguish. - `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's spec.yaml (the responsibility this probe is supposed to test) - `${FUNCTIONALITY_PATH}` — `01-functionality.md` -- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (B1 vendored the +- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (step 1 vendored the source regardless of upstream/local) — needed to judge whether the probe exercises what the skill actually instructs the agent to do diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md index f2edb38..3616f56 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -17,7 +17,7 @@ its own probe folder and nothing else. `spec.yaml` describing what this probe sets up + expects - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (skill's stated responsibilities and classification) -- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (B1 vendored the +- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (step 1 vendored the source regardless of upstream/local) - `${OUTPUT_PROBE_DIR}` — `tests///` - `${OPERATOR_DIRECTIVES}` — case-level revision hints for this diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index 0a58b5b..9ae9cb6 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -65,7 +65,7 @@ Three checks: exists (prior accumulated state), else `vendored-skill/`. The validator needs this as "skill BEFORE". If `vendored-skill/` was re-vendored between step 8 and step 9 (e.g., the operator - re-ran B1 mid-chain), surface to the user — the BEFORE must + re-ran step 1 mid-chain), surface to the user — the BEFORE must match what the optimizer read. The fix is re-invoking step 8 against the new state. 3. `01-functionality.md` must exist. From 2c813e8138221b115132bb8b5ca3bc643113e6b8 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 07:44:28 -0500 Subject: [PATCH 048/121] refactor(wrapper-flow): resolve wrapper-vs-target at step 1, not step 2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The wrapper-vs-target gymnastics in research-submissions was over-engineered. The correct model: step 1's user gate is where the wrapper question gets resolved. If the user opts to optimize the underlying content, step 1 re-vendors with an updated SKILL_SOURCE and the new 01-functionality.md has likely_wrapper= false. Downstream steps just read skill_source as the PR target — no conditional logic, no re-litigation. Changes: - step 1 SKILL.md (g): make the SKILL_SOURCE update on re-vendor explicit, and document what each user choice means for what downstream steps see - research-submissions.md frontmatter description: pr_target_* derives directly from skill_source; no special wrapper case - Drop body sections 1 ("Source structure") and 2 ("PR target") — they were redundant with frontmatter and re-litigated the step-1 decision - Reasoning protocol: replace the wrapper-handling step with a one-line read of skill_source; reorder so PR-shape research comes before linked-consumer search - Linked consumers stays as a coordination hint for the PR composer, but as its own body section rather than woven into the (now-removed) wrapper analysis Co-Authored-By: Claude Opus 4.7 --- .../SKILL.md | 16 +- .../research-submissions.md | 150 ++++++------------ 2 files changed, 63 insertions(+), 103 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 3ab8bb1..d94e104 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -172,9 +172,19 @@ content elsewhere), surface to the user before handing off: > judgment is wrong or you meant to point elsewhere. If the user picks (1): delete `vendored-skill/`, vendor the -referenced content into it, and re-invoke this skill from (c). -If (2): keep the report as-is and proceed. If (3): exit; the user -will re-invoke with a new source. +content at `wrapper_points_to` into it, update `${SKILL_SOURCE}` +to that URL/path, and re-invoke this skill from (c). The new +step 1 run rebuilds `01-functionality.md` from the underlying +content — `skill_source` reflects the new location and +`likely_wrapper` will typically be `false` (unless the underlying +is itself another wrapper, in which case repeat the gate). After +this re-run, downstream steps see a clean non-wrapper source. + +If (2): keep the report as-is and proceed. Downstream steps see +`likely_wrapper: true` with `wrapper_points_to` documented; the +PR target is the wrapper itself (the user's deliberate choice). + +If (3): exit; the user will re-invoke with a new source. Once resolved, hand off based on this report's `pr_submission_intent` field: diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/skill-optimizer-subagents/research-submissions.md index 215278f..c92ec73 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/skill-optimizer-subagents/research-submissions.md @@ -65,61 +65,45 @@ linked_consumers: branch name. - **`requires_cla`** — `true` if the upstream requires a Contributor License Agreement. -- **`pr_target_repo` / `pr_target_path`** — where the PR should - actually go. Often the same as the user-provided source. But if - step 1 flagged the source as `likely_wrapper: true` and the user - opted to optimize the wrapper's referenced content (the common - case), the PR target is the referenced content's location, not - the wrapper's. The downstream PR composer needs this to target - the right file. -- **`linked_consumers`** — other SKILL.md files (in the same repo - or other repos) that reference the same content as `pr_target_path`. - If non-empty, a change to that target may affect all of them; - the PR composer needs to coordinate or be honest about scope. +- **`pr_target_repo` / `pr_target_path`** — derived directly from + `01-functionality.md`'s `skill_source`. Step 1's user gate + already resolved any wrapper-vs-content question: if the source + was a wrapper and the user opted to optimize the underlying, + step 1 re-vendored and `skill_source` now points at that + underlying content. You don't re-litigate the choice here. +- **`linked_consumers`** — other SKILL.md files (same repo or other + known repos) that reference the same content as `pr_target_path`. + Flagged so the PR composer can coordinate or be honest about + scope. Same-repo wrappers typically update automatically; cross- + repo consumers usually need follow-up PRs. Body sections: -1. **Source structure** (new) — read `01-functionality.md`'s - `likely_wrapper` field. If `true`, step 1's subagent judged the - user-pointed source as a thin wrapper, and the user opted to - work with the referenced content (the common case after step 1's - user gate). Document the wrapper-vs-target relationship here: - the wrapper path, the target path, and the link mechanism. If - you find linked consumers (other SKILL.md files referencing - the same target), list them with the relationship (other - agent wrappers, downstream skills that import from this one). - If `likely_wrapper: false`, the source is the actual content - and this section just states that. -2. **PR target** (new) — given the above, which file should the - PR modify? Usually `pr_target_path` is the same as step 1's - source; for wrapper cases it's `wrapper_points_to`. If the - change would affect multiple consumers, flag whether they - live in the same repo (single PR suffices) or different - repos (multiple PRs or coordination required). -3. **License** — full SPDX identifier + brief plain-English summary - (permissive / copyleft / proprietary). Document the license for - the canonical content's repo (which may differ from the entry - file's repo). -4. **CLA requirement** — what kind (DCO, individual CLA, corporate - CLA), how it's signed, link to the CLA tool if applicable. Same - note: applies to the canonical content's repo. -5. **Frontmatter spec** — extracted from a sample of existing - skills in the canonical content's repo. List required fields, - optional fields, value conventions. If the repo's skills don't - have a consistent frontmatter, say so rather than inventing a - standard. -6. **File-location conventions** — where new skills go (which +1. **License** — full SPDX identifier + brief plain-English summary + (permissive / copyleft / proprietary). +2. **CLA requirement** — what kind (DCO, individual CLA, corporate + CLA), how it's signed, link to the CLA tool if applicable. +3. **Frontmatter spec** — extracted from a sample of existing + skills in the repo. List required fields, optional fields, value + conventions. If the repo's skills don't have a consistent + frontmatter, say so rather than inventing a standard. +4. **File-location conventions** — where new skills go (which directory, which subdir pattern), where references/scripts go. -7. **Prefix taxonomy** — if the repo uses commit prefixes +5. **Prefix taxonomy** — if the repo uses commit prefixes (`feat:`, `fix:`, `docs:`, etc.) or PR title prefixes, extract the taxonomy from recent merged PRs. -8. **PR-shape patterns** — from a sample of recent merged PRs: +6. **PR-shape patterns** — from a sample of recent merged PRs: what does a typical PR description include (rationale, testing notes, screenshots)? What's the typical PR size (single-file diff, multi-file)? Additive-only or destructive changes accepted? -9. **Rejection signals** — from a sample of closed-without-merge +7. **Rejection signals** — from a sample of closed-without-merge PRs: what got rejected and why? Common patterns to AVOID. +8. **Linked consumers** (only if non-empty) — list the other + SKILL.md files referencing the same content, with a one-line + note on each: same repo or different, and whether a change at + `pr_target_path` propagates automatically or needs a follow-up + PR there. If you can't establish any of these from the available data, say so honestly in the relevant section rather than making up @@ -127,70 +111,37 @@ conventions. ## Reasoning protocol -1. **Read `01-functionality.md` for the wrapper context.** step 1's - subagent judged whether the user-pointed source was likely a - wrapper and recorded the result in the report's frontmatter: - - `likely_wrapper: true` + `wrapper_points_to: ` → the - source is a thin wrapper. After step 1's user gate, one of two - things happened: - - User opted to re-vendor the referenced content (common): - `vendored-skill/` now contains that content. Your PR - target is `wrapper_points_to` (or wherever it now lives - after the re-vendor). - - User opted to treat the wrapper as the skill: `vendored-skill/` - still contains the wrapper. Your PR target is the original - `skill_source`. - - `likely_wrapper: false` → the source is the actual content; - your PR target is `skill_source` as-is. - - Use these signals to determine `pr_target_repo` and - `pr_target_path` for your own frontmatter. If - `01-functionality.md` predates this feature and the - `likely_wrapper` field is missing, fall back to forming your - own judgment by reading the vendored content (signals: short - SKILL.md body of "see X" pointers, frontmatter fields like - `reference:`/`source:`/`canonical:`, multi-agent plugin - layout). Surface to the operator if you detect a wrapper step 1 - missed — it suggests step 1 needs a re-run. - -2. **Find linked consumers.** Search for other places that - reference the canonical content: - - Other SKILL.md files in the same repo that point to the - same content (agent-flavored wrappers) - - Downstream skills that import or reference this skill - - In the repo's package metadata, search for known dependents - - Record findings in `linked_consumers` frontmatter and the - "Source structure" body section. - -3. **Start with CONTRIBUTING.md** — if it exists, it explicitly +1. **Read `01-functionality.md` for `skill_source`.** That URL/path + IS your PR target — step 1's user gate already resolved any + wrapper question. If `likely_wrapper: true` is set (rare: user + opted to keep the wrapper rather than re-vendoring the + underlying), the PR target is still `skill_source` (the + wrapper itself); `wrapper_points_to` is documentation only. + +2. **Start with CONTRIBUTING.md** — if it exists, it explicitly states most of what you need (license, CLA, file conventions, - PR shape). Apply this to the CANONICAL content's repo, which - may differ from the entry-file's repo if (1) found a pattern. + PR shape). -4. **Sample recent merged PRs** — enough to identify the dominant +3. **Sample recent merged PRs** — enough to identify the dominant patterns. Quality over quantity; you want to see the actual accepted shapes. -5. **Sample closed-without-merge PRs** — these surface the +4. **Sample closed-without-merge PRs** — these surface the rejection signals. The CONTRIBUTING.md tells you the rules; the closed PRs show what happens when rules are broken. -6. **Look at PRs in the same skill category** if `${SKILL_SLUG}` +5. **Look at PRs in the same skill category** if `${SKILL_SLUG}` suggests one (e.g., browser-skills look at other browser PRs). Specific-category patterns trump generic repo-level patterns. -7. **Suggest the PR target.** Based on (1) and (2), recommend - which file(s) the PR should modify: - - Not a wrapper: PR against the source as-is - - Wrapper, single consumer: PR against the referenced content - - Wrapper, multiple consumers in the same repo: PR against - the referenced content; the wrappers update automatically - - Wrapper, multiple consumers across repos: PR against the - referenced content; flag that downstream wrappers may need - follow-up PRs in their own repos - -8. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** +6. **Find linked consumers.** Search for other SKILL.md files in + the same repo (or other known repos) that reference the same + content as `pr_target_path`. These are coordination notes for + the PR composer — same-repo wrappers typically update + automatically when the target changes; cross-repo consumers + usually need follow-up PRs in their own repos. + +7. **Resolve `${OPERATOR_DIRECTIVES}` as atomic requirements** (e.g., "include the vendor's CLA requirement explicitly" → make sure the CLA section is prominent). @@ -212,11 +163,10 @@ blocker rather than making up content: After writing `${OUTPUT_PATH}`, return a brief summary: -- Wrapper context (from step 1 or fallback detection): is the - source a wrapper? If yes: referenced-content location, linked - consumers count, suggested PR target +- PR target (the `pr_target_repo` / `pr_target_path` you derived) - License + CLA requirement - Branch target rule +- Linked consumers count (if any) - 2-3 most important rejection signals from the closed-PR sweep - Any blockers (auth failures, missing data, etc.) From fcde7482f1d82a254fc057f22b5527d062e783d0 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 08:33:26 -0500 Subject: [PATCH 049/121] fix(tests): relink smoke distribution test to v1.4 chain skills The smoke-skill-distribution test was reading paths under the nuked skills/skill-optimizer/ folder. Updated to: - Walk all 9 chain skill directories and verify each has valid SKILL.md frontmatter (replaces the single-canonical check that predates the v1.4 chain decomposition) - Point workbench reference checks at skills/skill-optimizer-shared/ workbench.md (the new shared location) Also removed an empty skill-optimizer-validate-improvement/ directory that wasn't cleaned up when the skill was renamed to skill-optimizer-validate. Known v1.4 debt still on the cleanup list: .claude-plugin/ marketplace.json points at ./skills/skill-optimizer (the nuked path). The test passes because it does string compare without existence check; needs follow-up in the plugin metadata sweep. Co-Authored-By: Claude Opus 4.7 --- tests/smoke-skill-distribution.ts | 47 +++++++++++++++++++------------ 1 file changed, 29 insertions(+), 18 deletions(-) diff --git a/tests/smoke-skill-distribution.ts b/tests/smoke-skill-distribution.ts index 0594fea..bf61d59 100644 --- a/tests/smoke-skill-distribution.ts +++ b/tests/smoke-skill-distribution.ts @@ -15,25 +15,36 @@ function readText(relativePath: string): string { return readFileSync(join(root, relativePath), 'utf-8'); } -test('canonical skill follows the portable agent skills contract', () => { - const skillPath = 'skills/skill-optimizer/SKILL.md'; - assert.equal(existsSync(join(root, skillPath)), true); - - const body = readText(skillPath); - assert.match(body, /^---\n[\s\S]*?\n---\n/); - assert.match(body, /^name: skill-optimizer$/m); - assert.match(body, /^description: .+/m); - assert.doesNotMatch(body, /^description: .{1025,}$/m); +test('chain skills follow the portable agent skills contract', () => { + const chainSkillDirs = [ + 'skill-optimizer-investigate-functionality', + 'skill-optimizer-investigate-submissions', + 'skill-optimizer-design-tests', + 'skill-optimizer-write-tests', + 'skill-optimizer-validate-tests', + 'skill-optimizer-run-bench', + 'skill-optimizer-analyze', + 'skill-optimizer-improve', + 'skill-optimizer-validate', + ]; + + for (const dir of chainSkillDirs) { + const skillPath = `skills/${dir}/SKILL.md`; + assert.equal(existsSync(join(root, skillPath)), true, `missing ${skillPath}`); + + const body = readText(skillPath); + assert.match(body, /^---\n[\s\S]*?\n---\n/, `${dir} missing frontmatter`); + assert.match(body, new RegExp(`^name: ${dir}$`, 'm'), `${dir} name mismatch`); + assert.match(body, /^description: .+/m, `${dir} missing description`); + assert.doesNotMatch(body, /^description: .{1025,}$/m, `${dir} description too long`); + } }); -test('canonical skill documents current workbench command and live CLI patterns', () => { - const skill = readText('skills/skill-optimizer/SKILL.md'); - const reference = readText('skills/skill-optimizer/references/workbench.md'); +test('shared workbench reference documents current CLI patterns', () => { + const reference = readText('skills/skill-optimizer-shared/workbench.md'); - for (const text of [skill, reference]) { - assert.doesNotMatch(text, /verify-suite/); - assert.doesNotMatch(text, /runWorkbenchReferenceSolutions/); - } + assert.doesNotMatch(reference, /verify-suite/); + assert.doesNotMatch(reference, /runWorkbenchReferenceSolutions/); assert.match(reference, /Live CLI\/API Skills/); assert.match(reference, /Use dedicated test credentials/); @@ -42,8 +53,8 @@ test('canonical skill documents current workbench command and live CLI patterns' assert.match(reference, /Include a prompt-injection or unsafe-instruction case/); }); -test('workbench reference documents bin directory visibility accurately', () => { - const reference = readText('skills/skill-optimizer/references/workbench.md'); +test('shared workbench reference documents bin directory visibility accurately', () => { + const reference = readText('skills/skill-optimizer-shared/workbench.md'); assert.match( reference, From 640524b230ca3cbe3c474165d5805cf4feedfc40 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Fri, 22 May 2026 08:41:39 -0500 Subject: [PATCH 050/121] fix(metadata): relink plugin manifests + docs to v1.4 chain skills After the v1.4 redesign nuked the monolithic skills/skill-optimizer/ in favor of 9 chain skills under skills/skill-optimizer-*/, the plugin manifests, install docs, and Gemini context file still pointed at the dead paths. The smoke test caught the workbench reference but not the marketplace pointer (string compare only). Fixed across all surfaces: - .claude-plugin/marketplace.json: skills array now lists all 9 chain skill paths instead of the dead ./skills/skill-optimizer - GEMINI.md: @imports point at shared/workflow.md (chain overview) + shared/workbench.md instead of the deleted monolithic SKILL.md and references/workbench.md - README.md, CONTRIBUTING.md, AGENTS.md, CLAUDE.md, docs/README.{ codex,opencode}.md, .cursor/INSTALL.md, .codex/INSTALL.md, .opencode/INSTALL.md: replaced "canonical skill" framing with the 9-skill chain description; updated --skill flag examples to enumerate each chain skill explicitly - tests/smoke-skill-distribution.ts: marketplace test now asserts all 9 chain skills are listed AND verifies each path has a real SKILL.md on disk (closes the string-compare-only loophole that let the old dead path pass). Gemini test asserts the new @import targets. All 11 tests pass; typecheck clean. Co-Authored-By: Claude Opus 4.7 --- .claude-plugin/marketplace.json | 10 ++++++++- .codex/INSTALL.md | 16 +++++++++++--- .cursor/INSTALL.md | 16 +++++++++++--- .opencode/INSTALL.md | 4 ++-- AGENTS.md | 6 ++++-- CLAUDE.md | 7 +++--- CONTRIBUTING.md | 11 +++++----- GEMINI.md | 4 ++-- README.md | 36 ++++++++++++++++++++++++------- docs/README.codex.md | 16 +++++++++++--- docs/README.opencode.md | 6 +++--- tests/smoke-skill-distribution.ts | 23 ++++++++++++++++---- 12 files changed, 116 insertions(+), 39 deletions(-) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 5bf84bf..bd669b5 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -14,7 +14,15 @@ "name": "Fast" }, "skills": [ - "./skills/skill-optimizer" + "./skills/skill-optimizer-investigate-functionality", + "./skills/skill-optimizer-investigate-submissions", + "./skills/skill-optimizer-design-tests", + "./skills/skill-optimizer-write-tests", + "./skills/skill-optimizer-validate-tests", + "./skills/skill-optimizer-run-bench", + "./skills/skill-optimizer-analyze", + "./skills/skill-optimizer-improve", + "./skills/skill-optimizer-validate" ] } ] diff --git a/.codex/INSTALL.md b/.codex/INSTALL.md index 489e8ee..855155a 100644 --- a/.codex/INSTALL.md +++ b/.codex/INSTALL.md @@ -22,10 +22,20 @@ codex plugin marketplace add fastxyz/skill-optimizer --ref main ## Skill-Only Install -Install the canonical skill with the open skills CLI: +Install the 9 chain skills with the open skills CLI: ```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a codex -y +npx skills add fastxyz/skill-optimizer \ + --skill skill-optimizer-investigate-functionality \ + --skill skill-optimizer-investigate-submissions \ + --skill skill-optimizer-design-tests \ + --skill skill-optimizer-write-tests \ + --skill skill-optimizer-validate-tests \ + --skill skill-optimizer-run-bench \ + --skill skill-optimizer-analyze \ + --skill skill-optimizer-improve \ + --skill skill-optimizer-validate \ + -a codex -y ``` -Restart Codex if the skill does not appear immediately. +Restart Codex if the skills do not appear immediately. diff --git a/.cursor/INSTALL.md b/.cursor/INSTALL.md index c6d8de6..8016f32 100644 --- a/.cursor/INSTALL.md +++ b/.cursor/INSTALL.md @@ -2,14 +2,24 @@ ## Skill install -Install the skill into Cursor's project or global skill directory through the open skills CLI: +Install the 9 chain skills into Cursor's project or global skill directory through the open skills CLI: ```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a cursor -y +npx skills add fastxyz/skill-optimizer \ + --skill skill-optimizer-investigate-functionality \ + --skill skill-optimizer-investigate-submissions \ + --skill skill-optimizer-design-tests \ + --skill skill-optimizer-write-tests \ + --skill skill-optimizer-validate-tests \ + --skill skill-optimizer-run-bench \ + --skill skill-optimizer-analyze \ + --skill skill-optimizer-improve \ + --skill skill-optimizer-validate \ + -a cursor -y ``` Cursor can also import remote skills from GitHub in Settings -> Rules -> Project Rules -> Add Rule -> Remote Rule (Github). ## Plugin metadata -This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The canonical skill remains `skills/skill-optimizer/SKILL.md`. +This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The skills live at `skills/skill-optimizer-*/`; the plugin manifest exposes all 9 as a chain that triggers based on the user's intent (investigate, design tests, run bench, analyze, improve, validate). diff --git a/.opencode/INSTALL.md b/.opencode/INSTALL.md index c93409a..c11e621 100644 --- a/.opencode/INSTALL.md +++ b/.opencode/INSTALL.md @@ -8,9 +8,9 @@ Add the plugin to `opencode.json` at user or project scope: } ``` -Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can load `skill-optimizer`. +Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can discover the 9 chain skills (`skill-optimizer-investigate-functionality`, `skill-optimizer-investigate-submissions`, `skill-optimizer-design-tests`, `skill-optimizer-write-tests`, `skill-optimizer-validate-tests`, `skill-optimizer-run-bench`, `skill-optimizer-analyze`, `skill-optimizer-improve`, `skill-optimizer-validate`). -Verify with the skill tool by listing skills or loading `skill-optimizer`. +Verify by listing skills with the skill tool; each chain skill triggers on its own description (investigate, design tests, run bench, analyze, improve, validate). To pin a version, append a tag or commit ref: diff --git a/AGENTS.md b/AGENTS.md index 0d09351..a412443 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -20,7 +20,9 @@ npx tsx src/cli.ts run-suite --help - `src/cli.ts`: public CLI entrypoint - `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces - `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases -- `skills/skill-optimizer/SKILL.md`: canonical distributable Agent Skill +- `skills/skill-optimizer-*/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) +- `skills/skill-optimizer-shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench reference; loaded on-demand by chain skills +- `skills/skill-optimizer-subagents/`: prompt templates dispatched by chain skills via the Agent tool - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support - `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin - `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file @@ -47,7 +49,7 @@ Keep the README installation section aligned with packaged plugin metadata: - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state. - The agent phase sees only `/work`, not `/case` or `/results`. -- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. - Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`. - Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies. - Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials. diff --git a/CLAUDE.md b/CLAUDE.md index e91b35e..4c2d4e9 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -22,8 +22,9 @@ npx tsx src/cli.ts run-suite --help - `src/cli.ts`: public CLI entrypoint - `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces - `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases -- `skills/skill-optimizer/SKILL.md`: canonical distributable Agent Skill -- `skills/skill-optimizer/references/workbench.md`: detailed workbench schema and usage reference +- `skills/skill-optimizer-*/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) +- `skills/skill-optimizer-shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills +- `skills/skill-optimizer-subagents/`: prompt templates dispatched by chain skills via the Agent tool - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support - `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin - `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file @@ -50,7 +51,7 @@ Keep the README installation section aligned with packaged plugin metadata: - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state. - The agent phase sees only `/work`, not `/case` or `/results`. -- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. - Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`. - Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies. - Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 870e01f..505da10 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,6 +1,6 @@ # Contributing to skill-optimizer -Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills. Changes should preserve deterministic grading, isolated agent workspaces, and the canonical `skills/skill-optimizer/SKILL.md` distribution path. +Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills/skill-optimizer-*/`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths. ## Installing The Skill @@ -24,8 +24,9 @@ All three commands must pass before opening a PR when code changes are involved. - `src/cli.ts` — public CLI entry point for `run-case` and `run-suite`. - `src/workbench/` — case/suite loading, Docker runner, Pi agent wiring, graders, traces, metrics, MCP support, and trial aggregation. - `docker/workbench-runner.Dockerfile` — non-root container image for setup, agent, grade, and cleanup phases. -- `skills/skill-optimizer/SKILL.md` — canonical distributable Agent Skill. -- `skills/skill-optimizer/references/workbench.md` — detailed workbench schema and authoring reference. +- `skills/skill-optimizer-*/` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). +- `skills/skill-optimizer-shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. +- `skills/skill-optimizer-subagents/` — prompt templates dispatched by chain skills via the Agent tool. - `examples/workbench/` — packaged example suites. - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`, `.agents/plugins/marketplace.json`, `gemini-extension.json`, `GEMINI.md` — cross-agent plugin and extension metadata. - `tests/` — hand-rolled smoke tests (`tsx tests/smoke-*.ts`). @@ -46,7 +47,7 @@ All three commands must pass before opening a PR when code changes are involved. - `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override. - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - The agent phase sees only `/work`, not `/case`, `/results`, graders, hidden answers, or hidden metadata. -- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. ## Testing guidance @@ -58,7 +59,7 @@ All three commands must pass before opening a PR when code changes are involved. ## Adding workbench capabilities -Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/skill-optimizer/references/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior. +Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/skill-optimizer-shared/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior. ## Commit style diff --git a/GEMINI.md b/GEMINI.md index 8f6d532..ed6d83b 100644 --- a/GEMINI.md +++ b/GEMINI.md @@ -1,5 +1,5 @@ @./AGENTS.md @./README.md @./CONTRIBUTING.md -@./skills/skill-optimizer/SKILL.md -@./skills/skill-optimizer/references/workbench.md +@./skills/skill-optimizer-shared/workflow.md +@./skills/skill-optimizer-shared/workbench.md diff --git a/README.md b/README.md index 4dc6a6b..72c09a9 100644 --- a/README.md +++ b/README.md @@ -1,15 +1,15 @@ # skill-optimizer -Docker workbench and Agent Skill for running deterministic evals against agent skills. +Docker workbench and Agent Skills for running deterministic evals against agent skills. Use this repo in two ways: -- Install the `skill-optimizer` skill/plugin into your agent so it can author and debug eval suites. +- Install the `skill-optimizer` plugin into your agent so it can investigate a target skill, design and run an eval suite, analyze failures, and propose improvements. The plugin bundles a 9-step chain of Agent Skills under `skills/skill-optimizer-*/`. - Run the local CLI to execute cases and suites in Docker against OpenRouter models. ## Installation -Installation differs by agent. The canonical skill is `skills/skill-optimizer/SKILL.md`; every plugin manifest points at that same file. +Installation differs by agent. Every plugin manifest exposes all 9 chain skills from `skills/skill-optimizer-*/`; once installed, each skill triggers on its own description (e.g. "investigate this skill", "design tests for this skill", "run the bench", "analyze the results"). ### Claude Code @@ -53,13 +53,23 @@ codex plugin marketplace add fastxyz/skill-optimizer ### Cursor -Install the skill with the open skills CLI: +Install the chain skills with the open skills CLI (pass each skill name explicitly): ```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a cursor -y +npx skills add fastxyz/skill-optimizer \ + --skill skill-optimizer-investigate-functionality \ + --skill skill-optimizer-investigate-submissions \ + --skill skill-optimizer-design-tests \ + --skill skill-optimizer-write-tests \ + --skill skill-optimizer-validate-tests \ + --skill skill-optimizer-run-bench \ + --skill skill-optimizer-analyze \ + --skill skill-optimizer-improve \ + --skill skill-optimizer-validate \ + -a cursor -y ``` -Cursor can also import the skill from GitHub via Settings -> Rules -> Project Rules -> Add Rule -> Remote Rule (Github). The Cursor plugin metadata lives at `.cursor-plugin/plugin.json`. +Cursor can also import individual skills from GitHub via Settings -> Rules -> Project Rules -> Add Rule -> Remote Rule (Github). The Cursor plugin metadata lives at `.cursor-plugin/plugin.json`. ### OpenCode @@ -95,10 +105,20 @@ gemini extensions update skill-optimizer ### Skill-Only Install -If you only want the skill files without plugin metadata, use the open skills CLI: +If you only want the skill files without plugin metadata, use the open skills CLI. Pass each chain skill explicitly: ```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a claude-code -a opencode -a codex -a cursor -y +npx skills add fastxyz/skill-optimizer \ + --skill skill-optimizer-investigate-functionality \ + --skill skill-optimizer-investigate-submissions \ + --skill skill-optimizer-design-tests \ + --skill skill-optimizer-write-tests \ + --skill skill-optimizer-validate-tests \ + --skill skill-optimizer-run-bench \ + --skill skill-optimizer-analyze \ + --skill skill-optimizer-improve \ + --skill skill-optimizer-validate \ + -a claude-code -a opencode -a codex -a cursor -y ``` ## Local CLI Setup diff --git a/docs/README.codex.md b/docs/README.codex.md index 6401f37..846f3ba 100644 --- a/docs/README.codex.md +++ b/docs/README.codex.md @@ -22,10 +22,20 @@ codex plugin marketplace add fastxyz/skill-optimizer --ref main ## Skill-Only Install -Install only the skill files with the open skills CLI: +Install only the chain skill files with the open skills CLI (pass each skill name explicitly): ```bash -npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a codex -y +npx skills add fastxyz/skill-optimizer \ + --skill skill-optimizer-investigate-functionality \ + --skill skill-optimizer-investigate-submissions \ + --skill skill-optimizer-design-tests \ + --skill skill-optimizer-write-tests \ + --skill skill-optimizer-validate-tests \ + --skill skill-optimizer-run-bench \ + --skill skill-optimizer-analyze \ + --skill skill-optimizer-improve \ + --skill skill-optimizer-validate \ + -a codex -y ``` -Restart Codex if the skill does not appear immediately. The canonical skill path is `skills/skill-optimizer/SKILL.md`. +Restart Codex if the skills do not appear immediately. The chain skills live at `skills/skill-optimizer-*/SKILL.md`. diff --git a/docs/README.opencode.md b/docs/README.opencode.md index 93ff893..083f70b 100644 --- a/docs/README.opencode.md +++ b/docs/README.opencode.md @@ -12,11 +12,11 @@ Add the plugin to `opencode.json` at user or project scope: } ``` -Restart OpenCode. The plugin registers this repository's `skills/` directory so OpenCode can discover `skill-optimizer` without symlinks. +Restart OpenCode. The plugin registers this repository's `skills/` directory so OpenCode can discover the 9 chain skills without symlinks. ## Verify -Use OpenCode's native `skill` tool to list skills or load `skill-optimizer`. +Use OpenCode's native `skill` tool to list skills. You should see all 9 chain skills (`skill-optimizer-investigate-functionality`, `skill-optimizer-investigate-submissions`, `skill-optimizer-design-tests`, `skill-optimizer-write-tests`, `skill-optimizer-validate-tests`, `skill-optimizer-run-bench`, `skill-optimizer-analyze`, `skill-optimizer-improve`, `skill-optimizer-validate`); each triggers on its own description. ## Updating @@ -32,4 +32,4 @@ OpenCode reinstalls git plugins when it starts. To pin a tag or commit, append a The plugin exposes `.opencode/plugins/skill-optimizer.js` and adds the repository `skills/` directory to `config.skills.paths`. -The canonical skill is `skills/skill-optimizer/SKILL.md`. +The chain skills live at `skills/skill-optimizer-*/SKILL.md` — 9 user-invocable Agent Skills that orchestrate investigation, test design, bench runs, analysis, and improvement of a target skill. diff --git a/tests/smoke-skill-distribution.ts b/tests/smoke-skill-distribution.ts index bf61d59..5172d43 100644 --- a/tests/smoke-skill-distribution.ts +++ b/tests/smoke-skill-distribution.ts @@ -109,7 +109,7 @@ test('package metadata does not include broad example result directories', () => ); }); -test('Claude plugin and marketplace metadata point at the canonical skill', () => { +test('Claude plugin and marketplace metadata expose all v1.4 chain skills', () => { const pkg = readJson('package.json'); const plugin = readJson('.claude-plugin/plugin.json'); const marketplace = readJson('.claude-plugin/marketplace.json'); @@ -125,7 +125,21 @@ test('Claude plugin and marketplace metadata point at the canonical skill', () = assert.equal(marketplace.plugins[0].name, 'skill-optimizer'); assert.equal(marketplace.plugins[0].description, pluginDescription); assert.equal(marketplace.plugins[0].source, './'); - assert.deepEqual(marketplace.plugins[0].skills, ['./skills/skill-optimizer']); + assert.deepEqual(marketplace.plugins[0].skills, [ + './skills/skill-optimizer-investigate-functionality', + './skills/skill-optimizer-investigate-submissions', + './skills/skill-optimizer-design-tests', + './skills/skill-optimizer-write-tests', + './skills/skill-optimizer-validate-tests', + './skills/skill-optimizer-run-bench', + './skills/skill-optimizer-analyze', + './skills/skill-optimizer-improve', + './skills/skill-optimizer-validate', + ]); + for (const skillPath of marketplace.plugins[0].skills) { + const resolved = skillPath.replace(/^\.\//, ''); + assert.equal(existsSync(join(root, resolved, 'SKILL.md')), true, `missing ${resolved}/SKILL.md`); + } }); test('Codex and Cursor plugin metadata point at the canonical skill', () => { @@ -177,7 +191,7 @@ test('OpenCode plugin registers the canonical skills directory', async () => { assert.deepEqual(config.skills.paths, [join(root, 'skills')]); }); -test('Gemini extension metadata points at the canonical context file', () => { +test('Gemini extension metadata imports the v1.4 chain overview + workbench reference', () => { const pkg = readJson('package.json'); const extension = readJson('gemini-extension.json'); const geminiInstructions = readText('GEMINI.md'); @@ -189,5 +203,6 @@ test('Gemini extension metadata points at the canonical context file', () => { assert.match(geminiInstructions, /^@\.\/AGENTS\.md$/m); assert.match(geminiInstructions, /^@\.\/README\.md$/m); assert.match(geminiInstructions, /^@\.\/CONTRIBUTING\.md$/m); - assert.match(geminiInstructions, /^@\.\/skills\/skill-optimizer\/SKILL\.md$/m); + assert.match(geminiInstructions, /^@\.\/skills\/skill-optimizer-shared\/workflow\.md$/m); + assert.match(geminiInstructions, /^@\.\/skills\/skill-optimizer-shared\/workbench\.md$/m); }); From 5d652b070e107a4c27eca395a3b45cf6c6b26072 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:44:26 -0500 Subject: [PATCH 051/121] docs(dispatch): require inline-rendered subagent prompts (Pattern A) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first end-to-end test run revealed that the operator session defaulted to a path-pointer dispatch pattern: the Agent's prompt was a short "your instructions live at PATH — read that first" note plus a "## Templated inputs" block listing variable values. The subagent then had to read the template file AND mentally substitute placeholders from the inputs section. Per official Claude Code subagent guidance, this is the wrong pattern. Recommended: the operator renders the template inline (substituting all ${VAR} placeholders) and passes the rendered text as the Agent tool's `prompt` parameter. The subagent receives a self-contained instruction set, no file I/O required to start work. Root cause: subagent-dispatch.md said "pass these templated inputs" without specifying how, and each chain SKILL.md's dispatch step said "Load the prompt template at X and substitute ${VAR}" — wording that could be read as "tell the subagent to load + substitute" rather than "you load + substitute and send the result inline." Fix: - subagent-dispatch.md: new "Dispatch protocol" section with explicit 3-step instructions (read template, substitute every ${VAR}, pass rendered text as Agent's prompt parameter; do NOT use path-pointer dispatch) - Every chain SKILL.md's dispatch step: "Load X and substitute Y" → "Render X inline by substituting Y, then pass the rendered text as the Agent tool's prompt parameter (per subagent-dispatch.md's Dispatch protocol)" Closes the dispatch-pattern ambiguity surfaced by the pilot run. Co-Authored-By: Claude Opus 4.7 --- skills/skill-optimizer-analyze/SKILL.md | 9 +++-- skills/skill-optimizer-design-tests/SKILL.md | 10 ++++-- skills/skill-optimizer-improve/SKILL.md | 11 +++--- .../SKILL.md | 11 ++++-- .../SKILL.md | 10 ++++-- .../subagent-dispatch.md | 35 ++++++++++++++++--- .../skill-optimizer-validate-tests/SKILL.md | 15 ++++---- skills/skill-optimizer-validate/SKILL.md | 13 ++++--- skills/skill-optimizer-write-tests/SKILL.md | 9 +++-- 9 files changed, 88 insertions(+), 35 deletions(-) diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md index 83fbcc5..31adfc7 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -86,11 +86,14 @@ contradiction before re-dispatching. Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT cluster failures or name weaknesses yourself in this session** — dispatch the subagent via the `Agent` -tool. Load the prompt template at +tool. Render the prompt template at [`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md) -and substitute `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, +inline by substituting `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, -`${OPERATOR_DIRECTIVES}`. +and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the +Agent tool's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). The subagent sees: `06-bench-summary.md` (entry point with failed-probe pointer list); `06-bench-results//` with diff --git a/skills/skill-optimizer-design-tests/SKILL.md b/skills/skill-optimizer-design-tests/SKILL.md index 7da4fff..952fdbe 100644 --- a/skills/skill-optimizer-design-tests/SKILL.md +++ b/skills/skill-optimizer-design-tests/SKILL.md @@ -81,10 +81,14 @@ re-dispatching; don't try to resolve it yourself. Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT enumerate responsibilities or design proposals yourself in this session** — dispatch the subagent via -the `Agent` tool. Load the prompt template at +the `Agent` tool. Render the prompt template at [`skills/skill-optimizer-subagents/test-designer.md`](../skill-optimizer-subagents/test-designer.md) -and substitute `${FUNCTIONALITY_PATH}`, `${TESTS_TREE_PATH}`, -`${PROPOSALS_PATH}`, `${OPERATOR_DIRECTIVES}`. +inline by substituting `${FUNCTIONALITY_PATH}`, +`${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, and +`${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent +tool's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). The subagent sees: `01-functionality.md`, the current `tests/` tree (load-bearing state per the maintenance rule), diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index e44b32b..9602fd0 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -96,12 +96,15 @@ Translate to an atomic new requirement before passing. Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT propose the diff yourself in this -session** — dispatch the subagent via the `Agent` tool. Load the +session** — dispatch the subagent via the `Agent` tool. Render the prompt template at [`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) -and substitute `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, -`${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, -`${PROPOSAL_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. +inline by substituting `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, +`${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), +`${PROPOSAL_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass +the rendered text as the Agent tool's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). The optimizer sees: `07-analysis.md`; `01-functionality.md`; `${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index d94e104..fadcb42 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -129,11 +129,16 @@ current source plus directives; git captures prior state. Collect Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT do the research yourself in this -session** — dispatch the subagent via the `Agent` tool. Load the +session** — dispatch the subagent via the `Agent` tool. Render the prompt template at [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md) -and substitute `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, -`${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, vendored path. +inline by substituting `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, +`${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, and the +vendored path, then pass the rendered text as the Agent tool's +`prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol — do not pass a short prompt that points at the +template path). The subagent sees: the vendored skill files at `vendored-skill/`, targeted web-search results, output path, frontmatter fields, diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/skill-optimizer-investigate-submissions/SKILL.md index 3aeb466..af771bf 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/skill-optimizer-investigate-submissions/SKILL.md @@ -72,10 +72,14 @@ known. Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT scrape the upstream repo yourself in this session** — dispatch the subagent via the `Agent` tool. -Load the prompt template at +Render the prompt template at [`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md) -and substitute `${UPSTREAM_REPO}` from `01-functionality.md`'s -`skill_source` field, `${OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. +inline by substituting `${UPSTREAM_REPO}` from +`01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, +and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the +Agent tool's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). The subagent sees: the upstream repo via `gh` CLI (PR list — both merged and closed-without-merge for shape patterns and rejection diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md index 09336e7..6667012 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -92,11 +92,36 @@ temptation): If the user supplied no new requirements and you're re-running because an upstream artifact changed, leave directives empty. -## Dispatch the subagent - -When you invoke the subagent for this step, pass these templated -inputs (each subagent prompt template names them with `${...}` -placeholders): +## Dispatch protocol + +The chain uses one dispatch shape, consistently. **You — the +operator session — render the prompt inline, then pass the +rendered text as the Agent tool's `prompt` parameter.** The +subagent should receive a fully self-contained instruction set +that needs no file I/O to begin work. + +Concretely, for every step's dispatch: + +1. **Read the subagent's prompt template** at + `skills/skill-optimizer-subagents/.md`. The chain skill's + workflow tells you which template. +2. **Substitute every `${VAR}` placeholder** with the concrete + value the chain skill's workflow specifies. `${OPERATOR_DIRECTIVES}` + may be empty; everything else is a path or value you already + have in this session. +3. **Pass the fully rendered text** as the Agent tool's `prompt` + parameter. Do NOT pass a short prompt that points at the + template path (e.g., "Your instructions live at X — read that + first"). Path-pointer dispatch costs an extra Read tool call, + leaves `${VAR}` placeholders for the subagent to mentally + substitute (error-prone), and obscures the actual prompt in + transcripts. + +The rendered prompt is what shows up in the conversation log and +what the subagent acts on; if it isn't right, you can see why +immediately. That visibility is why we inline. + +Templated inputs every dispatch carries: - `${OPERATOR_DIRECTIVES}` — the bulleted list (may be empty) - Step-specific inputs (paths to upstream artifacts, output path, diff --git a/skills/skill-optimizer-validate-tests/SKILL.md b/skills/skill-optimizer-validate-tests/SKILL.md index 7b3d131..e968684 100644 --- a/skills/skill-optimizer-validate-tests/SKILL.md +++ b/skills/skill-optimizer-validate-tests/SKILL.md @@ -90,13 +90,16 @@ requires a specific JSON shape the skill never asks for". Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT judge probes yourself in this session** — dispatch test-validator subagents via the `Agent` tool, -in parallel (emit all dispatches in a single message). Load the -prompt template at +in parallel (emit all dispatches in a single message). For each +probe, render the prompt template at [`skills/skill-optimizer-subagents/test-validator.md`](../skill-optimizer-subagents/test-validator.md) -and substitute per-probe inputs: `${PROBE_NAME}`, -`${PROBE_DIR}`, `${FUNCTIONALITY_SPEC_PATH}`, -`${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, -`${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. +inline by substituting that probe's `${PROBE_NAME}`, `${PROBE_DIR}`, +`${FUNCTIONALITY_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, +`${SKILL_SOURCE_PATH}`, `${VERDICT_OUTPUT_PATH}`, and +`${OPERATOR_DIRECTIVES}`, then pass the rendered text as that +Agent dispatch's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). Each test-validator sees: its single probe's full contents (spec.yaml, workspace files, grader.mjs, smoke fixtures); the diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index 9ae9cb6..9261db2 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -91,15 +91,18 @@ an atomic new requirement before passing. Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT judge the proposal yourself in this -session** — dispatch the subagent via the `Agent` tool. Load the +session** — dispatch the subagent via the `Agent` tool. Render the prompt template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) -and substitute `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (current -state — `improved-skill/` if it exists, else source), +inline by substituting `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` +(current state — `improved-skill/` if it exists, else source), `${SKILL_AFTER_PATH}` (proposed result — apply the diff to a temp copy; do NOT touch the canonical `improved-skill/` yet), -`${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` if PR-bound, -`${VERDICT_OUTPUT_PATH}`, `${OPERATOR_DIRECTIVES}`. +`${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), +`${VERDICT_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass +the rendered text as the Agent tool's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). The validator sees: skill BEFORE; skill AFTER (temporary materialization); `08-improvement-proposal.md`; diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 2fedd0a..6dfed5e 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -103,12 +103,15 @@ Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT write workspace files or graders yourself in this session** — dispatch test-writer subagents via the `Agent` tool, in parallel (emit all dispatches in a single -message). Load the prompt template at +message). For each probe, render the prompt template at [`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) -and substitute per-probe inputs: `${PROBE_NAME}`, +inline by substituting that probe's `${PROBE_NAME}`, `${FUNCTIONALITY_SPEC_PATH}`, `${PROBE_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, -`${OUTPUT_PROBE_DIR}`, `${OPERATOR_DIRECTIVES}`. +`${OUTPUT_PROBE_DIR}`, and `${OPERATOR_DIRECTIVES}`, then pass the +rendered text as that Agent dispatch's `prompt` parameter (per +[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +Dispatch protocol). Each test-writer sees: its single probe's intent (slug + parent functionality spec); `01-functionality.md`; the skill source From a51d9b02eb829f55eb11ed4b7a27197d02b8b435 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Thu, 7 May 2026 10:56:17 -0500 Subject: [PATCH 052/121] fix(workbench): unblock Linux Docker bind-mount permissions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The agent container runs non-root and writes results into a host-mounted results directory. On Linux the bind-mount inherits the host owner, so the container couldn't write `result.json`. Fix: - chmod 0777 on workDir + resultsDir before starting the container - chmod -R a+rw on resultsDir after `docker cp` so cleanup can read it Also gitignore .superpowers/ — runtime state from the categorization pipeline (per-skill JSON cache, progress logs). --- .gitignore | 1 + src/workbench/docker-runner.ts | 16 ++++++++++++++-- 2 files changed, 15 insertions(+), 2 deletions(-) diff --git a/.gitignore b/.gitignore index 955a362..087e5ce 100644 --- a/.gitignore +++ b/.gitignore @@ -55,6 +55,7 @@ temp/ docs/superpowers/ docs/plans/ docs/specs/ +.superpowers/ # Skill-optimizer generated artifacts .skill-optimizer/ diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index 292ce6a..b0f5451 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -1,4 +1,4 @@ -import { cpSync, existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { chmodSync, cpSync, existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; import { tmpdir } from 'node:os'; import { dirname, join, resolve } from 'node:path'; import { fileURLToPath } from 'node:url'; @@ -212,6 +212,8 @@ async function copyAgentResults(containerName: string, resultsDir: string, repoR copy.stderr.trim(), ].filter(Boolean).join('\n\n')); } + + await runShellCommand(`chmod -R a+rw ${shellQuote(resultsDir)}`, { cwd: repoRoot }); } async function removeContainer(containerName: string, repoRoot: string): Promise { @@ -408,6 +410,8 @@ export function prepareDockerWorkbenchRun( mkdirSync(referencesDir, { recursive: true }); mkdirSync(workDir, { recursive: true }); mkdirSync(resultsDir, { recursive: true }); + chmodSync(workDir, 0o777); + chmodSync(resultsDir, 0o777); copyDirectoryContents(resolvedCase.referencesDir, referencesDir); copyCaseSupportDirs(resolvedCase.configDir, caseDir); @@ -434,7 +438,15 @@ export function prepareDockerWorkbenchRun( resultPath: join(resultsDir, 'result.json'), tracePath: join(resultsDir, 'trace.jsonl'), ...(mcpConfigPath ? { mcpConfigPath } : {}), - cleanup: () => rmSync(tempDir, { recursive: true, force: true }), + cleanup: () => { + try { + rmSync(tempDir, { recursive: true, force: true }); + } catch (error) { + // The container (uid 10001) may write subdirs (.cache, .venv) that the host user + // cannot delete. Don't let cleanup failures kill the run; tmpfiles.d will sweep /tmp later. + console.warn(`workbench: could not remove ${tempDir}: ${error instanceof Error ? error.message : String(error)}`); + } + }, }; } From 8e9e39bf4bb7a7cf4897f4d069afa927c065fb41 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:51:16 -0500 Subject: [PATCH 053/121] fix(docker): drop stale COPY of removed scripts/ directory scripts/ was deleted in commit 5da0655 (Docker skill eval workbench) but the workbench-runner Dockerfile still tried to COPY it. Any fresh image build (e.g., the v1.4 chain's bench step on a clean host) would fail with "scripts: no such file or directory". Surfaced during the first end-to-end pilot run. Co-Authored-By: Claude Opus 4.7 --- docker/workbench-runner.Dockerfile | 1 - 1 file changed, 1 deletion(-) diff --git a/docker/workbench-runner.Dockerfile b/docker/workbench-runner.Dockerfile index 73591f7..54087ee 100644 --- a/docker/workbench-runner.Dockerfile +++ b/docker/workbench-runner.Dockerfile @@ -31,7 +31,6 @@ RUN apt-get update \ COPY package.json package-lock.json tsconfig.json ./ COPY src ./src -COPY scripts ./scripts COPY docs ./docs RUN npm ci \ From d16df9ce21176ae5fc152069ffc02d60b61ade59 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:52:34 -0500 Subject: [PATCH 054/121] fix(test-writer): pin probe layout + smoke.mjs import path explicitly The first end-to-end pilot run produced a probe where grader.mjs was placed under checks/ instead of at the probe root, and checks/smoke.mjs referenced ./grader.mjs instead of ../grader.mjs. The operator had to manually mv and sed-fix it before the bench could run. The prompt described each artifact in isolation (### 1. spec.yaml, ### 2. workspace/, ### 3. grader.mjs, ...) without showing the full tree. The subagent saw "grader.mjs" next to "checks/smoke.mjs" in the section list and conflated them. Fix: - Add an explicit layout tree to the top of the "Output" section with a bold "grader.mjs lives at probe root, NOT under checks/" callout and an explanation of why (workbench points at it there) - Add a concrete smoke.mjs skeleton showing the ../grader.mjs import path, with a probeRoot derivation that works regardless of CWD Co-Authored-By: Claude Opus 4.7 --- .../skill-optimizer-subagents/test-writer.md | 49 ++++++++++++++++++- 1 file changed, 47 insertions(+), 2 deletions(-) diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md index 3616f56..cf50a5d 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -51,8 +51,31 @@ its own probe folder and nothing else. ## Output: `${OUTPUT_PROBE_DIR}/` -Four artifacts inside the probe folder. The schema for the -workbench (suite.yml, grader contract, smoke format) lives in +Five artifacts inside the probe folder. The exact layout is: + +```text +${OUTPUT_PROBE_DIR}/ + spec.yaml # 1. probe-level intent + workspace/ # 2. files the agent sees in /work + + grader.mjs # 3. grading script — AT PROBE ROOT + smoke/ # 4. smoke fixtures + good/findings.txt + bad/findings.txt + empty/findings.txt + checks/ # 5. smoke check runner + smoke.mjs # references ../grader.mjs (relative to checks/) +``` + +**`grader.mjs` lives at the probe root, NOT under `checks/`.** The +smoke runner at `checks/smoke.mjs` references it as +`../grader.mjs`. This layout is load-bearing for the workbench +(`tests/suite.yml` generated by the operator points at +`grader.mjs` at the probe root); a grader put under `checks/` +will not be found. + +The workbench schema (suite.yml, grader contract, smoke format) +lives in [`../skill-optimizer-shared/workbench.md`](../skill-optimizer-shared/workbench.md). ### 1. `${PROBE_SPEC_PATH}` — the probe's own spec.yaml @@ -108,6 +131,28 @@ A small script that exercises `grader.mjs` against GOOD-pass/BAD-fail/EMPTY-fail outcomes. The operator session runs this at step 4 (e) to verify the grader. +**Import path matters.** Because `smoke.mjs` lives in +`checks/` and `grader.mjs` lives at the probe root, the import +is `../grader.mjs`: + +```js +// checks/smoke.mjs +import { grade } from '../grader.mjs'; +import { readFileSync } from 'node:fs'; +import { dirname, join } from 'node:path'; +import { fileURLToPath } from 'node:url'; + +const probeRoot = dirname(dirname(fileURLToPath(import.meta.url))); +for (const label of ['good', 'bad', 'empty']) { + const findings = readFileSync(join(probeRoot, 'smoke', label, 'findings.txt'), 'utf-8'); + const passed = await grade({ findings, workspaceDir: join(probeRoot, 'workspace') }); + console.log(`${label}: ${passed ? 'PASS' : 'FAIL'}`); +} +``` + +Adjust the `grade` call signature to match what your `grader.mjs` +exports. + ## Reasoning protocol 1. **Read the parent functionality spec first** — understand what From 46ccdf1e4bfc1685f63631691bcf5cbbec1845d1 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:54:48 -0500 Subject: [PATCH 055/121] fix(write-tests): present minimal+full probe-set options at step 4 gate MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 4's plan gate asked "Proceed?" against the full probe set (every entry of every picked functionality's suggested_probes). For 7 picked functionalities with 1-3 probes each, that defaulted to a 21-Agent-dispatch fan-out. The cost-warning language was buried in a paragraph, easy to miss on a quick yes. Fix: - write-tests (c): always present two options side by side at the user gate — "minimal coverage" (1 probe per functionality, first entry of each suggested_probes list) and "full coverage" (all suggested_probes). Recommend minimal for first runs and full for refinement. Cost difference is explicit before the user picks. - test-designer (subagent prompt): require `suggested_probes` to be ordered most-representative-first, so step 4's minimal-default picks the strongest probe per functionality, not an arbitrary one. Co-Authored-By: Claude Opus 4.7 --- .../test-designer.md | 12 +++++--- skills/skill-optimizer-write-tests/SKILL.md | 30 ++++++++++++++----- 2 files changed, 30 insertions(+), 12 deletions(-) diff --git a/skills/skill-optimizer-subagents/test-designer.md b/skills/skill-optimizer-subagents/test-designer.md index b0cff2f..8755746 100644 --- a/skills/skill-optimizer-subagents/test-designer.md +++ b/skills/skill-optimizer-subagents/test-designer.md @@ -116,10 +116,14 @@ X"). Don't overwrite the user's `picked: true` edits. multiple distinct test targets. E.g., "handles inputs" is too broad — split into "accepts valid inputs", "refuses malformed inputs", "handles empty inputs" as separate functionalities. -5. **Suggest probes per functionality.** For "refuses malformed - input", probes might be: `malformed-json`, `missing-required- - field`, `type-mismatch`. Each probe is one concrete scenario - step 4's test-writer will build. +5. **Suggest probes per functionality, most-representative first.** + For "refuses malformed input", probes might be: + `malformed-json`, `missing-required-field`, `type-mismatch`. + Each probe is one concrete scenario step 4's test-writer will + build. **Order matters:** step 4's minimal-coverage default + picks the first probe of each functionality, so put the probe + most likely to surface failures up front and trailing probes as + additional angles for full-coverage runs. 6. **Rank by importance.** High = a regression here would break the skill's core promise. Medium = degrades but doesn't break. Low = nice-to-have edge case. diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index 6dfed5e..a2b11fe 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -79,20 +79,34 @@ and apply the maintenance-step flow: ### (c) Plan probe set + user gate -For each picked functionality, decide the probe set: start with +For each picked functionality, decide the full probe set: start with `suggested_probes`, apply directives (adds/removes/revisions), preserve existing built probes unless directives target them. -Show the user the planned probe set per functionality, marking -each probe as existing / new / rebuild. Ask whether to change -anything before dispatch. +Then present TWO options side by side, with the cost difference +made explicit (one Agent dispatch per new probe; expected wall time +roughly probe-count × test-writer time): + +1. **Minimal coverage** (default for first runs and dry-runs): + one probe per picked functionality — the first entry of each + `suggested_probes` list. Sized to surface obvious weaknesses + without paying for the full matrix. +2. **Full coverage:** every probe in every picked functionality's + `suggested_probes`. Sized for thorough measurement, used once + minimal coverage has confirmed the chain is working. + +Show the user the planned probe set per functionality under each +option, marking each probe as existing / new / rebuild. Recommend +minimal for a first run and full once minimal has converged. Ask +which to proceed with. Three realistic responses: -1. **User approves.** Proceed to (d). -2. **User adds revisions.** Treat as directives, mark named probes - for rebuild (apply destructive-edit checkpoint), proceed to - (d). +1. **User picks minimal or full.** Proceed to (d) with that set. +2. **User adds revisions** (e.g., "give me the 2 most discriminating + per functionality", "drop probe X"). Treat as directives, mark + named probes for rebuild (apply destructive-edit checkpoint), + proceed to (d). 3. **User rejects the plan structure.** Surface and ask whether to abandon (loop back to step 3 to revise functionality specs) or retry with their feedback as directives. From c5dd46625a71eb636a0d78369e61ba2ae3cda690 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:56:23 -0500 Subject: [PATCH 056/121] feat(step-1): recommend a dedicated branch before chain starts MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The chain writes to docs/skill-optimizer//, vendored-skill/, and improved-skill/ — accumulating these on main/development clutters the canonical branch and makes runs hard to discard or merge as a unit. The first pilot run hit this because step 1 didn't surface the branch question, so the chain ran on whatever branch the user happened to be on. Add a new step (a) "Confirm on a dedicated branch" at the start of step 1's workflow. Recommends `git checkout -b eval/` when the user is on a long-lived branch (main, master, development); skips silently when they're already on a feature branch. User can decline and proceed; the note just makes the choice visible. Re-letters the rest of step 1's workflow (a-g become b-h) and updates the one internal reference accordingly. Co-Authored-By: Claude Opus 4.7 --- .../SKILL.md | 32 +++++++++++++------ 1 file changed, 23 insertions(+), 9 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index fadcb42..1d12bd3 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -47,7 +47,7 @@ research, step 8 optimization, step 9 validation) consult these fields when deciding what to target. If the subagent flags `likely_wrapper: true`, the operator surfaces -to the user (workflow step (g) below) and the user decides what +to the user (workflow step (h) below) and the user decides what to do — re-vendor and re-research the referenced content, or proceed treating the wrapper itself as the skill to optimize. @@ -61,7 +61,21 @@ the chain. ## Workflow -### (a) Classify the source +### (a) Confirm on a dedicated branch + +Before starting, check the current git branch. The chain writes to +`docs/skill-optimizer//`, `vendored-skill/`, and (on step 9 +approval) `improved-skill/` — keeping these on a feature branch +makes the run easy to discard, iterate on, or merge in one piece. + +If the user is on `main`, `master`, `development`, or a similar +long-lived branch, recommend creating a dedicated branch (e.g., +`git checkout -b eval/`) before continuing. If they decline, +proceed but note that chain outputs will accumulate on the current +branch. Skip this check entirely if they're already on a +purpose-named feature branch. + +### (b) Classify the source Upstream skill (URL or `//`) or local skill (filesystem path that already exists)? If the user gave a bare name @@ -72,7 +86,7 @@ to the user — don't guess at recovery. If the source is a plugin with multiple skills, ask which one to investigate (produce one report per skill). -### (b) Ask about PR intent +### (c) Ask about PR intent **Upstream skills:** ask "Do you want to optimize this skill for upstream PR submission?" and record the answer in the report @@ -91,7 +105,7 @@ If the user later changes their mind on PR intent, they re-run this skill — the canonical is overwritten and prior state lives in git history. -### (c) Vendor the source +### (d) Vendor the source Copy the skill's files into `vendored-skill/` at the working-directory root, regardless of source type: @@ -110,7 +124,7 @@ upstream got new commits worth refetching). The user's original local file is never modified by the chain — `vendored-skill/` is a separate copy that downstream steps read. -### (d) Determine the slug and the report path +### (e) Determine the slug and the report path `` is the source skill's own directory or file name (e.g., `firecrawl/skills/firecrawl-build-scrape` → `firecrawl-build-scrape`; @@ -118,14 +132,14 @@ The user's original local file is never modified by the chain — Report path: `docs/skill-optimizer//01-functionality.md`. -### (e) Handle iteration +### (f) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from the current source plus directives; git captures prior state. Collect `${OPERATOR_DIRECTIVES}` per the protocol. -### (f) Dispatch the functionality-researcher subagent +### (g) Dispatch the functionality-researcher subagent Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) for the constraints. **Do NOT do the research yourself in this @@ -154,7 +168,7 @@ user has already expressed. A subagent walled off from that context produces a fresh derivation from the source itself, not a rationalization of expectations. -### (g) Confirm + handle wrapper detection + hand off +### (h) Confirm + handle wrapper detection + hand off Verify the report file exists and frontmatter parses. @@ -178,7 +192,7 @@ content elsewhere), surface to the user before handing off: If the user picks (1): delete `vendored-skill/`, vendor the content at `wrapper_points_to` into it, update `${SKILL_SOURCE}` -to that URL/path, and re-invoke this skill from (c). The new +to that URL/path, and re-invoke this skill from (d). The new step 1 run rebuilds `01-functionality.md` from the underlying content — `skill_source` reflects the new location and `likely_wrapper` will typically be `false` (unless the underlying From 786c5feb1ac14da5945905a9294d593cb001cc93 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 06:57:26 -0500 Subject: [PATCH 057/121] feat(step-1): create chain-progress TodoWrite at start MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The chain is 9-10 steps spanning long-running work, but step 1 didn't tell the operator to create a TodoWrite list. The first pilot run had no chain-level progress tracking — operator had to re-orient mentally at each step's handoff. Superpowers-style explicit todos make multi-step work visible at a glance. Add a "Chain todos" sub-section under step 1's (a) with a suggested list of all chain steps. Each subsequent chain skill naturally updates its own entry via the operator's TodoWrite habit; only step 1 needs to seed the list. Co-Authored-By: Claude Opus 4.7 --- .../SKILL.md | 25 +++++++++++++++++-- 1 file changed, 23 insertions(+), 2 deletions(-) diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 1d12bd3..1592f5f 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -61,9 +61,11 @@ the chain. ## Workflow -### (a) Confirm on a dedicated branch +### (a) Confirm on a dedicated branch + create chain todos -Before starting, check the current git branch. The chain writes to +Two quick setup tasks before doing any research: + +**Branch check.** The chain writes to `docs/skill-optimizer//`, `vendored-skill/`, and (on step 9 approval) `improved-skill/` — keeping these on a feature branch makes the run easy to discard, iterate on, or merge in one piece. @@ -75,6 +77,25 @@ proceed but note that chain outputs will accumulate on the current branch. Skip this check entirely if they're already on a purpose-named feature branch. +**Chain todos.** Create a TodoWrite list with the chain's 9 (or +10, if autopilot will run) steps so progress is visible across a +long run. Mark step 1 as `in_progress` immediately; subsequent +chain skills update their own entry as they fire. Suggested list: + +1. `step 1: investigate-functionality` (in_progress) +2. `step 2: investigate-submissions` (only if PR-bound) +3. `step 3: design-tests` +4. `step 4: write-tests` +5. `step 5: validate-tests` +6. `step 6: run-bench` +7. `step 7: analyze` +8. `step 8: improve` +9. `step 9: validate` +10. `step 10: autopilot` (only if running end-to-end) + +The chain is long enough that this is worth tracking, even though +each individual chain skill is itself a small workflow. + ### (b) Classify the source Upstream skill (URL or `//`) or local skill From d5f3e2b40252bcc79f3bb97aca20e8a615e8173a Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 07:12:34 -0500 Subject: [PATCH 058/121] refactor(layout): split chain state across three locations MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Step 1's SKILL.md previously told the operator to put vendored-skill/ at the workspace root, which polluted the repo top level and conflicted with the workflow.md state-layout spec. Worse, the spec itself bundled everything under docs/skill-optimizer// — fine for reports, but mixed heavy ephemeral artifacts (raw trace.jsonl bench output) into the PR-reviewable tree. New layout (Option B per design discussion): - docs/skill-optimizer// TRACKED — reports + improved-skill deliverable (PR-reviewable) - skill-evals// TRACKED — reusable eval suites + probe machinery - .skill-optimizer// GITIGNORED — vendored-skill/ (re-fetchable from skill_source) and bench-results// (heavy raw output) Same across all three so a slug's full state is locatable by name. .skill-optimizer/ was already in .gitignore. Updated: - workflow.md state layout + chain output column with new paths - iteration-protocol.md (staleness example, step-kind table, bench-results location, vendored-skill description) - subagent-dispatch.md (improved-skill / vendored-skill path hints for steps 8 and 9) - All 9 chain SKILL.md files (path refs throughout) - All 7 subagent prompts (input descriptions, "what you see" sections, frontmatter examples) Co-Authored-By: Claude Opus 4.7 --- skills/skill-optimizer-analyze/SKILL.md | 10 +- skills/skill-optimizer-design-tests/SKILL.md | 21 ++-- skills/skill-optimizer-improve/SKILL.md | 19 +-- .../SKILL.md | 69 ++++++----- skills/skill-optimizer-run-bench/SKILL.md | 22 ++-- .../iteration-protocol.md | 32 +++--- .../subagent-dispatch.md | 11 +- skills/skill-optimizer-shared/workflow.md | 108 ++++++++++++------ skills/skill-optimizer-subagents/analyzer.md | 12 +- skills/skill-optimizer-subagents/optimizer.md | 10 +- .../research-functionality.md | 2 +- .../test-designer.md | 14 +-- .../test-validator.md | 4 +- .../skill-optimizer-subagents/test-writer.md | 8 +- skills/skill-optimizer-subagents/validator.md | 12 +- .../skill-optimizer-validate-tests/SKILL.md | 8 +- skills/skill-optimizer-validate/SKILL.md | 28 ++--- skills/skill-optimizer-write-tests/SKILL.md | 34 +++--- 18 files changed, 238 insertions(+), 186 deletions(-) diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/skill-optimizer-analyze/SKILL.md index 31adfc7..1a6282c 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/skill-optimizer-analyze/SKILL.md @@ -25,7 +25,7 @@ Frontmatter (runtime-relevant facts only, per --- has_structural_weakness: true | false weakness_count: -bench_results_path: 06-bench-results// +bench_results_path: .skill-optimizer//bench-results// --- ``` @@ -67,7 +67,7 @@ If the bench results dir is missing or `suite-result.json` is malformed, surface as a step-6 problem (incomplete or corrupted bench run) and tell the user to re-run step 6. -`tests/` and the source skill must also be available. +`skill-evals//` and the source skill must also be available. ### (b) Handle iteration @@ -96,13 +96,13 @@ Agent tool's `prompt` parameter (per Dispatch protocol). The subagent sees: `06-bench-summary.md` (entry point with -failed-probe pointer list); `06-bench-results//` with +failed-probe pointer list); `.skill-optimizer//bench-results//` with per-trial `trace.jsonl` and `findings.txt`; each probe's -`spec.yaml` from the `tests/` tree (intent only); the skill source +`spec.yaml` from the `skill-evals//` tree (intent only); the skill source content; `${OPERATOR_DIRECTIVES}`. The subagent does NOT see: **the test inputs themselves** -(`tests//workspace/` files) — this is load-bearing; the +(`skill-evals////workspace/` files) — this is load-bearing; the analyzer must think about the SKILL, not the SOLUTIONS; its own prior `07-analysis.md` or git history; `08-improvement-proposal.md`, `09-validator-verdict.md`, or any prior optimizer attempts. diff --git a/skills/skill-optimizer-design-tests/SKILL.md b/skills/skill-optimizer-design-tests/SKILL.md index 952fdbe..0e579dc 100644 --- a/skills/skill-optimizer-design-tests/SKILL.md +++ b/skills/skill-optimizer-design-tests/SKILL.md @@ -13,7 +13,7 @@ the skill must fulfill), then asks the user to pick which ones to actually build probes for at step 4. The filesystem IS the state: this step writes a -`tests//spec.yaml` per proposed functionality +`skill-evals///spec.yaml` per proposed functionality plus a one-time audit report at `03-test-proposals.md`. No `picked: []` array anywhere; each spec.yaml has its own `picked: true|false`. @@ -27,7 +27,7 @@ Two artifacts at `docs/skill-optimizer//`: human review of the design reasoning; downstream steps do NOT read it. The subagent rewrites it on re-runs. -2. **`tests//spec.yaml`** — one folder per +2. **`skill-evals///spec.yaml`** — one folder per proposed functionality, with frontmatter per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): @@ -59,14 +59,14 @@ For the body template of `03-test-proposals.md` and the exact `01-functionality.md` must exist. If not, tell the user to run `skill-optimizer-investigate-functionality` first and stop here. -If `docs/skill-optimizer//tests/` doesn't exist yet, create -the empty directory. +If `skill-evals//` doesn't exist yet, create the empty +directory. ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply the maintenance-step flow. Re-runs read the current -`tests/` tree as state and extend or modify it. Collect +`skill-evals//` tree as state and extend or modify it. Collect `${OPERATOR_DIRECTIVES}` per the protocol. If a directive will be destructive (e.g., "remove the X functionality"), commit a checkpoint before dispatching per the protocol's safe-destructive- @@ -90,8 +90,9 @@ tool's `prompt` parameter (per [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s Dispatch protocol). -The subagent sees: `01-functionality.md`, the current `tests/` -tree (load-bearing state per the maintenance rule), +The subagent sees: `01-functionality.md`, the current +`skill-evals//` tree (load-bearing state per the +maintenance rule), `${OPERATOR_DIRECTIVES}`, output paths. The subagent does NOT see: the skill's source content (would @@ -110,7 +111,7 @@ regression-defensive ones. The subagent writes: - `03-test-proposals.md` (the audit report) -- One `tests//spec.yaml` per proposed +- One `skill-evals///spec.yaml` per proposed functionality. Existing spec.yaml files the user has already edited (e.g., `picked: true` set) are preserved verbatim unless a directive explicitly targets that functionality. @@ -121,7 +122,7 @@ preserved-unchanged list, newly-added list. ### (d) Confirm subagent output Verify `03-test-proposals.md` parses and each new -`tests//spec.yaml` parses with required fields. +`skill-evals///spec.yaml` parses with required fields. If any spec.yaml fails to parse, surface to the user; don't repair the subagent's output yourself. @@ -135,7 +136,7 @@ step 1 yourself. Show the user the ranked list from `03-test-proposals.md` and tell them how to indicate picks: edit `picked: true|false` in each -`tests//spec.yaml`. They can also edit +`skill-evals///spec.yaml`. They can also edit `suggested_probes` or ask for a revised proposal. Three realistic responses: diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/skill-optimizer-improve/SKILL.md index 9602fd0..970d682 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/skill-optimizer-improve/SKILL.md @@ -69,9 +69,10 @@ Three checks: agree, exit honestly." Do NOT proceed. 3. `01-functionality.md` must exist. Skill content must be - readable: `improved-skill/` if it exists (accumulated state - from prior step-9 approvals), else `vendored-skill/` (step 1 - vendored the source regardless of upstream/local). + readable: `docs/skill-optimizer//improved-skill/` if it + exists (accumulated state from prior step-9 approvals), else + `.skill-optimizer//vendored-skill/` (step 1 vendored the + source regardless of upstream/local). If `pr_submission_intent: true`, `02-submissions.md` should exist so the optimizer can shape the diff to upstream conventions from @@ -107,13 +108,15 @@ the rendered text as the Agent tool's `prompt` parameter (per Dispatch protocol). The optimizer sees: `07-analysis.md`; `01-functionality.md`; -`${SKILL_CURRENT_PATH}` — `improved-skill/` if it exists, else -`vendored-skill/`; `02-submissions.md` if PR-bound; -`${OPERATOR_DIRECTIVES}`. +`${SKILL_CURRENT_PATH}` — `docs/skill-optimizer//improved-skill/` +if it exists, else `.skill-optimizer//vendored-skill/`; +`02-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. The optimizer does NOT see: raw failed trials, `findings.txt`, -`trace.jsonl`; grader internals (`tests//grader.mjs`); test -inputs (`tests//workspace/`); its own prior +`trace.jsonl`; grader internals +(`skill-evals////grader.mjs`); test +inputs (`skill-evals////workspace/`); +its own prior `08-improvement-proposal.md` or git history; step 9's prior `09-validator-verdict.md` (would bias toward defending or pivoting away from the prior attempt). diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/skill-optimizer-investigate-functionality/SKILL.md index 1592f5f..42ffa26 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/skill-optimizer-investigate-functionality/SKILL.md @@ -53,11 +53,14 @@ proceed treating the wrapper itself as the skill to optimize. For the body template, see [`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). -The source skill is vendored to `vendored-skill/` regardless of -source type (upstream fetch or local copy) so downstream steps -have a single canonical input — they never branch on local vs -upstream, and the user's original local file is not touched by -the chain. +The source skill is vendored to +`.skill-optimizer//vendored-skill/` regardless of source +type (upstream fetch or local copy) so downstream steps have a +single canonical input — they never branch on local vs upstream, +and the user's original local file is not touched by the chain. +The vendored copy is gitignored (re-fetchable from +`skill_source` any time); only the reports and improved-skill +under `docs/skill-optimizer//` get tracked. ## Workflow @@ -65,10 +68,13 @@ the chain. Two quick setup tasks before doing any research: -**Branch check.** The chain writes to -`docs/skill-optimizer//`, `vendored-skill/`, and (on step 9 -approval) `improved-skill/` — keeping these on a feature branch -makes the run easy to discard, iterate on, or merge in one piece. +**Branch check.** The chain writes to three locations: +`docs/skill-optimizer//` (reports + the improved-skill +deliverable, tracked), `skill-evals//` (test suites, +tracked), and `.skill-optimizer//` (vendored source + raw +bench output, gitignored). Keeping the tracked locations on a +feature branch makes the run easy to discard, iterate on, or +merge in one piece. If the user is on `main`, `master`, `development`, or a similar long-lived branch, recommend creating a dedicated branch (e.g., @@ -128,22 +134,25 @@ history. ### (d) Vendor the source -Copy the skill's files into `vendored-skill/` at the -working-directory root, regardless of source type: +Copy the skill's files into `.skill-optimizer//vendored-skill/` +(create the parent dirs if missing), regardless of source type: - **Upstream** — `gh api` or equivalent fetch of the skill's directory contents -- **Local** — `cp -r` of the local skill's directory into - `vendored-skill/` +- **Local** — `cp -r` of the local skill's directory into the + vendored path -If `vendored-skill/` already exists from a prior run, reuse it -unless: (1) the source URL changed (upstream — a different repo or -skill is being investigated), or (2) the user explicitly asks to -re-vendor (e.g., they edited a local skill between runs, or the -upstream got new commits worth refetching). +If `.skill-optimizer//vendored-skill/` already exists from a +prior run, reuse it unless: (1) the source URL changed (upstream +— a different repo or skill is being investigated), or (2) the +user explicitly asks to re-vendor (e.g., they edited a local +skill between runs, or the upstream got new commits worth +refetching). The user's original local file is never modified by the chain — -`vendored-skill/` is a separate copy that downstream steps read. +the vendored copy is a separate read-only reference for +downstream steps. `.skill-optimizer/` is gitignored, so the +vendored copy doesn't add repo weight. ### (e) Determine the slug and the report path @@ -175,8 +184,9 @@ vendored path, then pass the rendered text as the Agent tool's Dispatch protocol — do not pass a short prompt that points at the template path). -The subagent sees: the vendored skill files at `vendored-skill/`, -targeted web-search results, output path, frontmatter fields, +The subagent sees: the vendored skill files at +`.skill-optimizer//vendored-skill/`, targeted web-search +results, output path, frontmatter fields, `${OPERATOR_DIRECTIVES}`. The subagent does NOT see: prior `01-functionality.md` or its git @@ -211,14 +221,15 @@ content elsewhere), surface to the user before handing off: > 3. **Cancel and provide a different source.** If the subagent's > judgment is wrong or you meant to point elsewhere. -If the user picks (1): delete `vendored-skill/`, vendor the -content at `wrapper_points_to` into it, update `${SKILL_SOURCE}` -to that URL/path, and re-invoke this skill from (d). The new -step 1 run rebuilds `01-functionality.md` from the underlying -content — `skill_source` reflects the new location and -`likely_wrapper` will typically be `false` (unless the underlying -is itself another wrapper, in which case repeat the gate). After -this re-run, downstream steps see a clean non-wrapper source. +If the user picks (1): delete +`.skill-optimizer//vendored-skill/`, vendor the content at +`wrapper_points_to` into it, update `${SKILL_SOURCE}` to that +URL/path, and re-invoke this skill from (d). The new step 1 run +rebuilds `01-functionality.md` from the underlying content — +`skill_source` reflects the new location and `likely_wrapper` +will typically be `false` (unless the underlying is itself +another wrapper, in which case repeat the gate). After this +re-run, downstream steps see a clean non-wrapper source. If (2): keep the report as-is and proceed. Downstream steps see `likely_wrapper: true` with `wrapper_points_to` documented; the diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/skill-optimizer-run-bench/SKILL.md index 6e8a410..93559cc 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/skill-optimizer-run-bench/SKILL.md @@ -1,13 +1,13 @@ --- name: skill-optimizer-run-bench -description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once `tests/suite.yml` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. +description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once `skill-evals//suite.yml` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. --- # skill-optimizer-run-bench Step 6 of the skill-optimizer chain. **Fresh-derivation step (for the summary).** Invokes the skill-optimizer CLI's `run-suite` -command against `tests/suite.yml` (generated by step 4), captures +command against `skill-evals//suite.yml` (generated by step 4), captures the raw results under a timestamped directory, and writes a small summary report that step 7 reads. No subagent dispatch — this is a thin operator-driven CLI step. @@ -16,7 +16,7 @@ thin operator-driven CLI step. Two outputs at `docs/skill-optimizer//`: -1. **`06-bench-results//`** — raw CLI output: +1. **`.skill-optimizer//bench-results//`** — raw CLI output: `suite-result.json`, per-trial `trace.jsonl`, per-trial `findings.txt`, preserved workspaces. Timestamped per invocation; old runs are NEVER overwritten. Outside the @@ -28,7 +28,7 @@ Two outputs at `docs/skill-optimizer//`: ```yaml --- - bench_results_path: 06-bench-results// + bench_results_path: .skill-optimizer//bench-results// overall_pass_rate: --- ``` @@ -42,21 +42,21 @@ Two outputs at `docs/skill-optimizer//`: ### (a) Confirm prerequisites -`tests/suite.yml` must exist (per step 4's output contract). If +`skill-evals//suite.yml` must exist (per step 4's output contract). If missing, tell the user to complete step 4 first. -`vendored-skill/` should exist (step 1 vendored the source regardless -of upstream/local). +`.skill-optimizer//vendored-skill/` should exist (step 1 +vendored the source regardless of upstream/local). ### (b) Handle iteration Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply it. Two pieces: -- **Raw bench output** at `06-bench-results//`: always +- **Raw bench output** at `.skill-optimizer//bench-results//`: always write to a fresh timestamped directory (`date -u +%Y%m%dT%H%M%SZ`). Outside the iteration protocol; never overwrite. - **Summary file** at `06-bench-summary.md`: fresh-derivation per - the protocol. Staleness check: if `tests/suite.yml` has a newer + the protocol. Staleness check: if `skill-evals//suite.yml` has a newer git mtime than `06-bench-summary.md`, the existing summary is stale. If summary is current AND no re-run directive, tell the user and exit. @@ -69,7 +69,7 @@ OUT_DIR="docs/skill-optimizer//06-bench-results/${TIMESTAMP}" mkdir -p "${OUT_DIR}" npx tsx /src/cli.ts run-suite \ - docs/skill-optimizer//tests/suite.yml \ + skill-evals//suite.yml \ --out "${OUT_DIR}" \ --trials 3 ``` @@ -126,5 +126,5 @@ Read `overall_pass_rate`. Two messages: accept a case filter, so there's no first-class partial-rebench mode. Operators with expensive suites can run `run-case` manually for changed probes and splice into the prior - `06-bench-results//` dir — outside-the-chain escape hatch, + `.skill-optimizer//bench-results//` dir — outside-the-chain escape hatch, not a supported flow. diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/skill-optimizer-shared/iteration-protocol.md index d687c87..ba0fa5b 100644 --- a/skills/skill-optimizer-shared/iteration-protocol.md +++ b/skills/skill-optimizer-shared/iteration-protocol.md @@ -48,8 +48,8 @@ extend or modify it. | Step | What accumulates | |---|---| -| 3. design-tests | `tests//spec.yaml` grows as functionalities are added/refined | -| 4. write-tests | `tests///` probes grow as the user adds coverage | +| 3. design-tests | `skill-evals///spec.yaml` grows as functionalities are added/refined | +| 4. write-tests | `skill-evals////` probes grow as the user adds coverage | ## Staleness detection (git-native) @@ -58,7 +58,7 @@ modification times against direct upstream: ```bash UPSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//01-functionality.md) -DOWNSTREAM_T=$(git log -1 --format=%ct -- docs/skill-optimizer//tests/) +DOWNSTREAM_T=$(git log -1 --format=%ct -- skill-evals//) [ "$UPSTREAM_T" -gt "$DOWNSTREAM_T" ] && echo "stale" ``` @@ -116,18 +116,22 @@ needed — git is the archive. ## What this protocol does NOT cover -- **Step 5's raw bench results** are timestamped under - `06-bench-results//`. Each run preserves naturally as its own - directory. The summary file `06-bench-summary.md` is a - fresh-derivation artifact per this protocol; the timestamped raw - output is outside. +- **Step 6's raw bench results** are timestamped under + `.skill-optimizer//bench-results//`. Each run preserves + naturally as its own directory. Gitignored — `bench-results/` + accumulates trace.jsonl and findings.txt per trial (tens of MB + per run); the durable record is the + `docs/skill-optimizer//06-bench-summary.md` aggregate, + which IS tracked and falls under this protocol. - **Auto-pilot's summary report** at `docs/skill-optimizer//autopilot-summary-.md` is timestamped per run. Each auto-pilot invocation produces a fresh summary; old ones are preserved naturally. -- **`vendored-skill/`** (the canonical input to the chain — step 1 - copies the source skill here regardless of upstream/local) is - reused across iterations of the same slug unless: (1) the - source URL changed (upstream), or (2) the user explicitly asks - to re-vendor (e.g., they edited the local skill between runs). - The source URL itself is the identity; no versioning. +- **`.skill-optimizer//vendored-skill/`** (the canonical + input to the chain — step 1 copies the source skill here + regardless of upstream/local) is reused across iterations of the + same slug unless: (1) the source URL changed (upstream), or (2) + the user explicitly asks to re-vendor (e.g., they edited the + local skill between runs). The source URL itself is the + identity; no versioning. Gitignored — re-fetchable from + `skill_source` any time. diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/skill-optimizer-shared/subagent-dispatch.md index 6667012..a81eb5b 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/skill-optimizer-shared/subagent-dispatch.md @@ -54,11 +54,12 @@ Three rules every reasoning subagent must follow: target being improved) is upstream input, not your own canonical. Both the optimizer (step 8) and the validator (step 9) read the **current state of the skill** — which is - `improved-skill/` if it exists (the accumulated state from - prior step-9 approvals), else `vendored-skill/` (step 1 vendored - the source regardless of upstream/local). Neither step modifies - the original. The "own canonical" off-limits to each subagent - is its report (`08-improvement-proposal.md` for the optimizer; + `docs/skill-optimizer//improved-skill/` if it exists (the + accumulated state from prior step-9 approvals), else + `.skill-optimizer//vendored-skill/` (step 1 vendored the + source regardless of upstream/local). Neither step modifies the + original. The "own canonical" off-limits to each subagent is + its report (`08-improvement-proposal.md` for the optimizer; `09-validator-verdict.md` for the validator), not the skill content itself. diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index 858f638..8ae1f45 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -9,16 +9,18 @@ re-runs, and the relationship between steps. | # | Skill | Kind | Input | Output | |---|---|---|---|---| -| 1 | `investigate-functionality` | fresh-derivation | source skill (URL or local) | `01-functionality.md` | -| 2 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `02-submissions.md` | -| 3 | `design-tests` | maintenance | step 1 | `03-test-proposals.md`, `tests//spec.yaml` | -| 4 | `write-tests` | maintenance | steps 1+3 | `tests///`, `tests/suite.yml` | -| 5 | `validate-tests` | fresh-derivation | step 4 + step 1 + skill | `05-tests-verdict.md` | -| 6 | `run-bench` | fresh-derivation (summary) | step 5 (must approve) + skill | `06-bench-results//`, `06-bench-summary.md` | -| 7 | `analyze` | fresh-derivation | step 6 + skill | `07-analysis.md` | -| 8 | `improve` | fresh-derivation | step 7 + skill (+ step 2 if PR-bound) | `08-improvement-proposal.md` | -| 9 | `validate` | fresh-derivation | step 8 + skill (+ step 2 if PR-bound) | `09-validator-verdict.md`, `improved-skill/` (on approve) | -| 10 | `autopilot` | chain driver | same as step 1 + flags | `autopilot-summary-.md` | +| 1 | `investigate-functionality` | fresh-derivation | source skill (URL or local) | `docs/.../01-functionality.md`, `.skill-optimizer/.../vendored-skill/` | +| 2 | `investigate-submissions` (optional) | fresh-derivation | step 1 (PR-bound) | `docs/.../02-submissions.md` | +| 3 | `design-tests` | maintenance | step 1 | `docs/.../03-test-proposals.md`, `skill-evals/...//spec.yaml` | +| 4 | `write-tests` | maintenance | steps 1+3 | `skill-evals/...///`, `skill-evals/.../suite.yml` | +| 5 | `validate-tests` | fresh-derivation | step 4 + step 1 + skill | `docs/.../05-tests-verdict.md` | +| 6 | `run-bench` | fresh-derivation (summary) | step 5 (must approve) + skill | `.skill-optimizer/.../bench-results//`, `docs/.../06-bench-summary.md` | +| 7 | `analyze` | fresh-derivation | step 6 + skill | `docs/.../07-analysis.md` | +| 8 | `improve` | fresh-derivation | step 7 + skill (+ step 2 if PR-bound) | `docs/.../08-improvement-proposal.md` | +| 9 | `validate` | fresh-derivation | step 8 + skill (+ step 2 if PR-bound) | `docs/.../09-validator-verdict.md`, `docs/.../improved-skill/` (on approve) | +| 10 | `autopilot` | chain driver | same as step 1 + flags | `docs/.../autopilot-summary-.md` | + +(Paths abbreviated; full state layout below.) The PR-or-not decision is made ONCE at step 1; subsequent steps know from `01-functionality.md`'s `pr_submission_intent` field @@ -44,9 +46,9 @@ when to re-run each step. Common triggers: | 1 | source URL changed; PR-intent changed; user wants fresh research with new directives | | 2 | upstream updated `CONTRIBUTING.md`/license/CLA; PR conventions visibly shifted; step 1 changed | | 3 | user wants different coverage; step 4/5/6/7 surfaced a coverage gap; step 1 changed | -| 4 | step 3's `tests/` tree changed; a probe's smoke check failed; step 5 flagged probes for revision; step 7 showed probes systematically too easy/hard | +| 4 | step 3's `skill-evals//` tree changed; a probe's smoke check failed; step 5 flagged probes for revision; step 7 showed probes systematically too easy/hard | | 5 | step 4 produced new or revised probes; user disagrees with prior test verdict | -| 6 | step 4's `tests/` changed (and step 5 re-approved); user wants fresh trial data (flakiness, model list changed); step 7 wants more trials | +| 6 | step 4's `skill-evals//` changed (and step 5 re-approved); user wants fresh trial data (flakiness, model list changed); step 7 wants more trials | | 7 | step 6 produced new results; user disagrees with prior analysis; step 8 was unable to address a named weakness | | 8 | step 7 produced a new analysis; step 9 returned `needs-revision`/`reject` with a distillable rationale; user wants a different approach | | 9 | step 8 produced a new proposal; user disagrees with prior verdict; step 2 was updated and prior external check is stale | @@ -75,36 +77,66 @@ auto-fire; they're surfaced to the user. ## State layout +Chain output splits across three top-level locations, each with +its own tracking discipline: + ```text -docs/skill-optimizer// - 01-functionality.md - 02-submissions.md # only if step 2 ran (PR-bound) - 03-test-proposals.md # step 3's audit report - tests/ - / - spec.yaml # step 3 writes - / # step 4 writes one per probe - spec.yaml - workspace/ - grader.mjs - smoke/{good,bad,empty}/ - checks/smoke.mjs - suite.yml # step 4 generates - 05-tests-verdict.md # step 5 writes (per-probe + aggregate) - 06-bench-results// # raw, timestamped per run - 06-bench-summary.md # single canonical - 07-analysis.md - 08-improvement-proposal.md - 09-validator-verdict.md - vendored-skill/ # frozen original (always — step 1 vendors upstream OR local) - improved-skill/ # step 9 materializes on approve - autopilot-summary-.md # step 10 writes per run +docs/skill-optimizer// # TRACKED — reports + deliverable (PR-reviewable) + 01-functionality.md # step 1 + 02-submissions.md # step 2 — only if PR-bound + 03-test-proposals.md # step 3 audit report + 05-tests-verdict.md # step 5 verdict (per-probe + aggregate) + 06-bench-summary.md # step 6 aggregate summary + 07-analysis.md # step 7 + 08-improvement-proposal.md # step 8 + 09-validator-verdict.md # step 9 + improved-skill/ # step 9 materializes on approve — the deliverable + autopilot-summary-.md # step 10 per run + +skill-evals// # TRACKED — eval suites + probes (reusable across runs) + / + spec.yaml # step 3 writes + / # step 4 writes one per probe + spec.yaml + workspace/ + grader.mjs + smoke/{good,bad,empty}/ + checks/smoke.mjs + suite.yml # step 4 generates + +.skill-optimizer// # GITIGNORED — heavy ephemeral artifacts + vendored-skill/ # step 1 vendors (upstream OR local) + bench-results// # step 6 raw output (suite-result.json, trace.jsonl, findings.txt) ``` `` is `--` for upstream skills, or -`` for local skills. Filesystem IS the state; -history is git (no `version:` fields or `archive/` directories -per [`frontmatter-discipline.md`](./frontmatter-discipline.md)). +`` for local skills. The same `` is reused +across all three locations so a slug's full state can be located +by name. + +**Why the split:** + +- **`docs/skill-optimizer//`** holds what a human reviews + during PR: the markdown reports, plus the improved-skill the + PR is about. Tracked so reviewers see history. +- **`skill-evals//`** holds reusable test machinery + (workbench probes). Tracked because tests are valuable on + their own — they run against future skill versions, get reused + across runs, get reviewed for fairness. +- **`.skill-optimizer//`** holds heavy or trivially + re-derivable artifacts. Gitignored: + - `vendored-skill/` can be re-fetched from `skill_source` any + time; tracking it doubles repo size for no review value. + - `bench-results//` accumulates `trace.jsonl` and + `findings.txt` per trial — tens of MB per run. The summary + at `docs/skill-optimizer//06-bench-summary.md` is the + durable record. + +Filesystem IS the state; history is git (no `version:` fields or +`archive/` directories per +[`frontmatter-discipline.md`](./frontmatter-discipline.md)). +`.skill-optimizer//` is the one location where state is +intentionally NOT versioned — re-runs overwrite freely. ## Shared docs diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/skill-optimizer-subagents/analyzer.md index 23eaf9e..ad32d4d 100644 --- a/skills/skill-optimizer-subagents/analyzer.md +++ b/skills/skill-optimizer-subagents/analyzer.md @@ -12,12 +12,12 @@ address it AND the anti-patterns that would NOT. - `${SUMMARY_PATH}` — `06-bench-summary.md` (entry point; failed-probe pointer list) -- `${BENCH_RESULTS_PATH}` — `06-bench-results//` with +- `${BENCH_RESULTS_PATH}` — `.skill-optimizer//bench-results//` with per-trial `trace.jsonl` and per-trial `findings.txt` -- `${TESTS_TREE_PATH}` — `tests/` tree. **Read ONLY each probe's +- `${TESTS_TREE_PATH}` — `skill-evals//` tree. **Read ONLY each probe's `spec.yaml`** — what each probe was probing at the level of INTENT. Do NOT read `workspace/` contents (raw input fixtures). -- `${SKILL_SOURCE_PATH}` — the skill's content at `vendored-skill/` +- `${SKILL_SOURCE_PATH}` — the skill's content at `.skill-optimizer//vendored-skill/` (step 1 vendored the source regardless of upstream/local) - `${OUTPUT_PATH}` — where to write `07-analysis.md` - `${OPERATOR_DIRECTIVES}` — atomic new requirements (empty unless @@ -34,7 +34,7 @@ address it AND the anti-patterns that would NOT. ## What you do NOT see -- **The test inputs themselves** (`tests//workspace/` +- **The test inputs themselves** (`skill-evals////workspace/` files). This is load-bearing — you must think about the SKILL (what it instructs the agent to do), not the SOLUTIONS (what the agent should have done in this specific input shape). @@ -58,7 +58,7 @@ Frontmatter (runtime-relevant only): --- has_structural_weakness: true | false weakness_count: -bench_results_path: 06-bench-results// +bench_results_path: .skill-optimizer//bench-results// --- ``` @@ -78,7 +78,7 @@ For each weakness, a section with **all five required parts** - **Pattern**: Across trials, systematically failed to detect . Specifically: //>. + probe-IDs from skill-evals////>. - **Hypothesized cause**: . diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md index e2d57da..ae0fa1f 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -13,8 +13,8 @@ skill (that's step 9's job after the validator approves). - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (the skill's stated responsibilities — your change must not contradict them) - `${SKILL_CURRENT_PATH}` — the **current state of the skill**: - `improved-skill/` if it exists (accumulated state from prior - step-9 approvals), else `vendored-skill/` (step 1 vendored the + `docs/skill-optimizer//improved-skill/` if it exists (accumulated state from prior + step-9 approvals), else `.skill-optimizer//vendored-skill/` (step 1 vendored the source regardless of upstream/local). You propose a new improvement on top of whatever current state you read. - `${SUBMISSIONS_PATH}` — `02-submissions.md` if PR-bound (lets @@ -37,15 +37,15 @@ skill (that's step 9's job after the validator approves). ## What you do NOT see - **Raw failed trials** (`findings.txt`, `trace.jsonl` from - `06-bench-results/`). The analyzer translated raw failures into + `.skill-optimizer//bench-results/`). The analyzer translated raw failures into general principles + anti-patterns at step 7; your job is to apply the principle. Seeing the raw failures would pull you toward pattern-matching a patch for THOSE specific trials — the textbook ducktape mode. -- Grader internals (`tests//grader.mjs` source). You +- Grader internals (`skill-evals////grader.mjs` source). You don't optimize against the grader; you optimize against the named weakness. -- Test inputs (`tests//workspace/`). Same reason — you +- Test inputs (`skill-evals////workspace/`). Same reason — you reason about the SKILL, not about specific inputs. - Your own prior `08-improvement-proposal.md` or git history. No consistency-with-prior-attempts bias. diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/skill-optimizer-subagents/research-functionality.md index 1c460d1..cc9533e 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/skill-optimizer-subagents/research-functionality.md @@ -10,7 +10,7 @@ the chain consumes. - `${SKILL_SOURCE}` — URL (`//`) or local filesystem path to the skill being investigated (recorded in the report frontmatter, but you read from `${VENDORED_PATH}`) -- `${VENDORED_PATH}` — `vendored-skill/` directory (the operator +- `${VENDORED_PATH}` — `.skill-optimizer//vendored-skill/` directory (the operator session has already copied the source skill here, whether the original was upstream or local; this is your read path) - `${PR_SUBMISSION_INTENT}` — `true` or `false`, captured at step 1 diff --git a/skills/skill-optimizer-subagents/test-designer.md b/skills/skill-optimizer-subagents/test-designer.md index 8755746..9e854ac 100644 --- a/skills/skill-optimizer-subagents/test-designer.md +++ b/skills/skill-optimizer-subagents/test-designer.md @@ -3,8 +3,8 @@ You are dispatched by `skill-optimizer-design-tests` to enumerate the skill's responsibilities (from `01-functionality.md`) and propose a ranked set of **functionalities** to test, then write -a per-functionality `tests//spec.yaml` for each -plus a one-time audit report `03-test-proposals.md`. +a per-functionality `skill-evals///spec.yaml` +for each plus a one-time audit report `03-test-proposals.md`. A "functionality" here means **one distinct responsibility the skill must fulfill**. It's what step 4 will build probes for (each @@ -14,9 +14,9 @@ angles). ## Inputs (templated by the operator session) - `${FUNCTIONALITY_PATH}` — current `01-functionality.md` -- `${TESTS_TREE_PATH}` — current `tests/` directory (may be empty +- `${TESTS_TREE_PATH}` — current `skill-evals//` directory (may be empty on first invocation; on re-runs it has existing - `tests//spec.yaml` files) + `skill-evals///spec.yaml` files) - `${PROPOSALS_PATH}` — where to write the audit report (typically `docs/skill-optimizer//03-test-proposals.md`) - `${OPERATOR_DIRECTIVES}` — bulleted list of atomic new @@ -27,7 +27,7 @@ angles). - `${FUNCTIONALITY_PATH}` — the responsibilities, classification, triggers, terminology - `${TESTS_TREE_PATH}` — every existing - `tests//spec.yaml` (load-bearing state per + `skill-evals///spec.yaml` (load-bearing state per the maintenance rule; you extend, not replace) - The operator's directives @@ -78,7 +78,7 @@ Body: one-line reason (e.g., "trivially tested by N1's probe set", "not testable in a static workbench"). -### `tests//spec.yaml` — one per proposed functionality +### `skill-evals///spec.yaml` — one per proposed functionality ```yaml name: @@ -106,7 +106,7 @@ X"). Don't overwrite the user's `picked: true` edits. 1. **Read `${FUNCTIONALITY_PATH}` first** — enumerate the responsibilities. Each becomes a candidate functionality. -2. **Read the current `tests/` tree.** Existing functionalities +2. **Read the current `skill-evals//` tree.** Existing functionalities are constraints: you don't re-propose them (unless directives say otherwise); you extend or modify the set. 3. **Group adjacent responsibilities** if they're not meaningfully diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/skill-optimizer-subagents/test-validator.md index edda94b..2816630 100644 --- a/skills/skill-optimizer-subagents/test-validator.md +++ b/skills/skill-optimizer-subagents/test-validator.md @@ -19,12 +19,12 @@ coincidentally match rather than truly distinguish. ## Inputs (templated by the operator session) - `${PROBE_NAME}` — slug of the probe under judgment -- `${PROBE_DIR}` — `tests///` (full +- `${PROBE_DIR}` — `skill-evals////` (full contents: spec.yaml, workspace/, grader.mjs, smoke/) - `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's spec.yaml (the responsibility this probe is supposed to test) - `${FUNCTIONALITY_PATH}` — `01-functionality.md` -- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (step 1 vendored the +- `${SKILL_SOURCE_PATH}` — `.skill-optimizer//vendored-skill/` (step 1 vendored the source regardless of upstream/local) — needed to judge whether the probe exercises what the skill actually instructs the agent to do diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/skill-optimizer-subagents/test-writer.md index cf50a5d..9e94cad 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/skill-optimizer-subagents/test-writer.md @@ -9,7 +9,7 @@ its own probe folder and nothing else. ## Inputs (templated by the operator session) - `${PROBE_NAME}` — slug of this probe (becomes the folder name - under `tests//`) + under `skill-evals///`) - `${FUNCTIONALITY_SPEC_PATH}` — the parent functionality's `spec.yaml` (gives you context on what responsibility this probe is testing) @@ -17,9 +17,9 @@ its own probe folder and nothing else. `spec.yaml` describing what this probe sets up + expects - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (skill's stated responsibilities and classification) -- `${SKILL_SOURCE_PATH}` — `vendored-skill/` (step 1 vendored the +- `${SKILL_SOURCE_PATH}` — `.skill-optimizer//vendored-skill/` (step 1 vendored the source regardless of upstream/local) -- `${OUTPUT_PROBE_DIR}` — `tests///` +- `${OUTPUT_PROBE_DIR}` — `skill-evals////` - `${OPERATOR_DIRECTIVES}` — case-level revision hints for this probe (empty unless this is a rebuild) @@ -70,7 +70,7 @@ ${OUTPUT_PROBE_DIR}/ **`grader.mjs` lives at the probe root, NOT under `checks/`.** The smoke runner at `checks/smoke.mjs` references it as `../grader.mjs`. This layout is load-bearing for the workbench -(`tests/suite.yml` generated by the operator points at +(`skill-evals//suite.yml` generated by the operator points at `grader.mjs` at the probe root); a grader put under `checks/` will not be found. diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/skill-optimizer-subagents/validator.md index a1e19ee..0bda4b9 100644 --- a/skills/skill-optimizer-subagents/validator.md +++ b/skills/skill-optimizer-subagents/validator.md @@ -3,7 +3,7 @@ You are dispatched by `skill-optimizer-validate` to **independently judge** an improvement proposal (from step 8) and write `09-validator-verdict.md`. On `verdict: approve`, the operator -session materializes `improved-skill/` from the proposal. On +session materializes `docs/skill-optimizer//improved-skill/` from the proposal. On `needs-revision`/`reject`, the proposal is sent back to step 8. Your independence is the chain's load-bearing property at this @@ -17,12 +17,12 @@ without your own prior verdicts. - `${PROPOSAL_PATH}` — `08-improvement-proposal.md` (the optimizer's diff + rationale + self-check) - `${SKILL_BEFORE_PATH}` — the current state of the skill (same - state the optimizer read at step 8: `improved-skill/` if it - exists, else `vendored-skill/`) + state the optimizer read at step 8: `docs/skill-optimizer//improved-skill/` if it + exists, else `.skill-optimizer//vendored-skill/`) - `${SKILL_AFTER_PATH}` — temporary materialization of the proposal applied to a copy of BEFORE (the operator session prepares this; you read it but it's NOT the canonical - `improved-skill/` yet) + `docs/skill-optimizer//improved-skill/` yet) - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the internal consistency check: does the change make sense given the skill's stated responsibilities?) @@ -48,7 +48,7 @@ without your own prior verdicts. - Raw failed trials, `findings.txt`, `trace.jsonl` — your job is to validate the PROPOSAL, not re-analyze the failures -- Test inputs (`tests//workspace/`) — same reason +- Test inputs (`skill-evals////workspace/`) — same reason - The optimizer's internal reasoning trace from step 8 (only the proposal artifact, not how the optimizer arrived at it). Reading the optimizer's reasoning would lead you to accept @@ -128,7 +128,7 @@ directive for step 8. ## Verdict definitions - **approve** — internal + external (if applicable) checks both - hold. Proposal is sound; step 9 materializes `improved-skill/` + hold. Proposal is sound; step 9 materializes `docs/skill-optimizer//improved-skill/` and the chain reaches its endpoint for this iteration. - **needs-revision** — one or more checks has a fixable issue. Examples: anti-pattern self-check is hand-waving in a way diff --git a/skills/skill-optimizer-validate-tests/SKILL.md b/skills/skill-optimizer-validate-tests/SKILL.md index e968684..700b2ba 100644 --- a/skills/skill-optimizer-validate-tests/SKILL.md +++ b/skills/skill-optimizer-validate-tests/SKILL.md @@ -1,6 +1,6 @@ --- name: skill-optimizer-validate-tests -description: Use when the user wants to validate the probes built by `skill-optimizer-write-tests` before running the bench — phrases like "validate the tests", "check the graders", "are these probes fair", "review the test suite". Triggers after `skill-optimizer-write-tests` has populated `tests///` folders and before `skill-optimizer-run-bench`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking that the probes actually probe what they claim should trigger this. +description: Use when the user wants to validate the probes built by `skill-optimizer-write-tests` before running the bench — phrases like "validate the tests", "check the graders", "are these probes fair", "review the test suite". Triggers after `skill-optimizer-write-tests` has populated `skill-evals////` folders and before `skill-optimizer-run-bench`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking that the probes actually probe what they claim should trigger this. --- # skill-optimizer-validate-tests @@ -67,8 +67,8 @@ Body template and the validator's reasoning protocol live at ### (a) Confirm prerequisites -`tests/` must exist with at least one -`tests///` folder containing `spec.yaml`, +`skill-evals//` must exist with at least one +`skill-evals////` folder containing `spec.yaml`, `workspace/`, `grader.mjs`, `smoke/`. If any picked functionality has zero built probes, surface as a step-4 problem and tell the user to run step 4 first. @@ -160,7 +160,7 @@ Don't auto-invoke step 4 or step 6. ## Edge cases -- **No probes built (tests/ empty)** — caught at (a). Tell the +- **No probes built (`skill-evals//` empty)** — caught at (a). Tell the user to run step 4 first. - **Validator's verdict contradicts the smoke check** (e.g., validator rejects a probe whose smoke check passed) — that's diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/skill-optimizer-validate/SKILL.md index 9261db2..4ac9db4 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/skill-optimizer-validate/SKILL.md @@ -10,7 +10,7 @@ the proposal from step 8, dispatches a validator subagent to independently check whether the proposed change is sound (internal consistency) and conformant (external PR conventions if PR-bound), then — on `verdict: approve` — materializes the improved skill at -`improved-skill/`. Writes `09-validator-verdict.md` regardless of +`docs/skill-optimizer//improved-skill/`. Writes `09-validator-verdict.md` regardless of verdict. **Single-shot per invocation.** No in-step revision loop. If the @@ -44,13 +44,13 @@ One or two artifacts at `docs/skill-optimizer//`: step 9 doesn't produce one. Body template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). -2. **`improved-skill/`** — improved skill content, only +2. **`docs/skill-optimizer//improved-skill/`** — improved skill content, only materialized when `verdict: approve`. Mirrors the source's directory structure with the optimizer's diff applied. **The - original is never modified**: `vendored-skill/` (the canonical + original is never modified**: `.skill-optimizer//vendored-skill/` (the canonical input regardless of upstream/local) stays frozen, and the user's original local file (if any) is untouched. Git tracks - `improved-skill/` history across iterations. + `docs/skill-optimizer//improved-skill/` history across iterations. ## Workflow @@ -61,9 +61,9 @@ Three checks: 1. `08-improvement-proposal.md` must exist with valid frontmatter. If not, tell the user to run `skill-optimizer-improve` first. -2. Current skill state must be readable: `improved-skill/` if it - exists (prior accumulated state), else `vendored-skill/`. The - validator needs this as "skill BEFORE". If `vendored-skill/` +2. Current skill state must be readable: `docs/skill-optimizer//improved-skill/` if it + exists (prior accumulated state), else `.skill-optimizer//vendored-skill/`. The + validator needs this as "skill BEFORE". If `.skill-optimizer//vendored-skill/` was re-vendored between step 8 and step 9 (e.g., the operator re-ran step 1 mid-chain), surface to the user — the BEFORE must match what the optimizer read. The fix is re-invoking step 8 @@ -95,9 +95,9 @@ session** — dispatch the subagent via the `Agent` tool. Render the prompt template at [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) inline by substituting `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` -(current state — `improved-skill/` if it exists, else source), +(current state — `docs/skill-optimizer//improved-skill/` if it exists, else source), `${SKILL_AFTER_PATH}` (proposed result — apply the diff to a temp -copy; do NOT touch the canonical `improved-skill/` yet), +copy; do NOT touch the canonical `docs/skill-optimizer//improved-skill/` yet), `${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), `${VERDICT_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per @@ -110,7 +110,7 @@ materialization); `08-improvement-proposal.md`; `${OPERATOR_DIRECTIVES}`. The validator does NOT see: raw failed trials, `findings.txt`, -`trace.jsonl`; test inputs (`tests//workspace/`); the +`trace.jsonl`; test inputs (`skill-evals////workspace/`); the optimizer's reasoning trace (only the proposal artifact); its own prior `09-validator-verdict.md` or git history; `07-analysis.md` directly. @@ -125,7 +125,7 @@ judging the new proposal on its own merits. ### (d) Handle the verdict + materialize on approve -**`verdict: approve`:** materialize `improved-skill/` by copying +**`verdict: approve`:** materialize `docs/skill-optimizer//improved-skill/` by copying the source's directory structure and applying the optimizer's diff. **Do NOT modify the source.** Commit: @@ -135,7 +135,7 @@ git commit -m "step 9: validate + apply improvement for " ``` The source skill file is NOT in the commit — git tracks -`improved-skill/` alongside the verdict report, but never touches +`docs/skill-optimizer//improved-skill/` alongside the verdict report, but never touches the user's source or the vendored reference. **`verdict: needs-revision`** or **`reject`:** do NOT materialize. @@ -145,11 +145,11 @@ the user's source or the vendored reference. Three messages by verdict: - **approve:** "Validation complete; improvement applied at - `improved-skill/`. The chain has reached its natural endpoint + `docs/skill-optimizer//improved-skill/`. The chain has reached its natural endpoint for this iteration. Three realistic next steps: review and copy locally; hand to a PR composer (auto-pilot can do this end-to-end if running step 10); or re-bench against a workbench - that points at `improved-skill/`." + that points at `docs/skill-optimizer//improved-skill/`." - **needs-revision:** "Validator says needs-revision. Distill the rationale into a directive and re-invoke step 8, then re-invoke this step. Operator owns distillation; chain skills don't diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/skill-optimizer-write-tests/SKILL.md index a2b11fe..fbbaad5 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/skill-optimizer-write-tests/SKILL.md @@ -1,27 +1,27 @@ --- name: skill-optimizer-write-tests -description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-design-tests` has produced `tests//spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. +description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-design-tests` has produced `skill-evals///spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. --- # skill-optimizer-write-tests Step 4 of the skill-optimizer chain. **Maintenance step.** Reads the picked functionalities from step 3 -(`tests//spec.yaml` files where `picked: true`), +(`skill-evals///spec.yaml` files where `picked: true`), decides the probe set per functionality, dispatches one test-writer subagent per probe (in parallel) to build the concrete workspace files + grader scripts + smoke fixtures, then generates -`tests/suite.yml` for the run-bench step. +`skill-evals//suite.yml` for the run-bench step. ## What you produce Two kinds of output under `docs/skill-optimizer//`: -1. **Probe folders** at `tests///`. For - each picked functionality, one or more probe folders containing: +1. **Probe folders** at `skill-evals////`. + For each picked functionality, one or more probe folders containing: ```text - tests/// + skill-evals//// spec.yaml # probe-level intent workspace/ # the files the agent sees in /work grader.mjs # the grading script (or .py) @@ -36,8 +36,8 @@ Two kinds of output under `docs/skill-optimizer//`: with these contents; `grader.mjs` present + smoke-check passed = built and verified. -2. **`tests/suite.yml`** — generated from the current tree. Lists - every probe under every `picked: true` functionality. +2. **`skill-evals//suite.yml`** — generated from the current + tree. Lists every probe under every `picked: true` functionality. Regenerated by step (f) on every run. Probe-level `spec.yaml` format and the workbench schema for @@ -54,8 +54,8 @@ respectively. Three prerequisites: 1. `01-functionality.md` exists. -2. `tests/` exists with at least one - `tests//spec.yaml` having `picked: true`. +2. `skill-evals//` exists with at least one + `/spec.yaml` having `picked: true`. 3. Every picked functionality's spec.yaml parses with required fields. @@ -67,7 +67,7 @@ guess. Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) and apply the maintenance-step flow: -- Walk `tests/`. For each picked functionality, determine which +- Walk `skill-evals//`. For each picked functionality, determine which probes already have built folders (`grader.mjs` exists). Existing built probes are preserved by default; don't re-dispatch unless `${OPERATOR_DIRECTIVES}` explicitly names them @@ -165,7 +165,7 @@ Each probe has its own `checks/smoke.mjs`. After test-writer dispatches return, run each new or rebuilt probe's smoke runner: ```bash -node tests///checks/smoke.mjs +node skill-evals////checks/smoke.mjs ``` Expected per probe: GOOD passes, BAD fails, EMPTY fails. If any @@ -181,15 +181,15 @@ fails, surface to the user. Two realistic responses: ### (f) Generate suite.yml + commit + hand off -Walk `tests/`. For each `tests//` where `spec.yaml` +Walk `skill-evals//`. For each `/` where `spec.yaml` has `picked: true`, list every probe folder and emit a suite.yml -entry per the workbench schema. Overwrite `tests/suite.yml`. +entry per the workbench schema. Overwrite `skill-evals//suite.yml`. -Commit `tests/` and hand off to `skill-optimizer-run-bench`. +Commit `skill-evals//` and hand off to `skill-optimizer-run-bench`. ## Edge cases - **User de-picks a functionality between step 4 runs** — its probe folders stay on disk; step (f) excludes them from the - regenerated `tests/suite.yml`. If re-picked later, probes are - already there. + regenerated `skill-evals//suite.yml`. If re-picked later, + probes are already there. From fe797566152c1213e9d7749df8b10b197b13d768 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 07:25:43 -0500 Subject: [PATCH 059/121] feat(philosophy): cross-vendor skill-design synthesis for analyzer/optimizer/validator MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds skills/skill-optimizer-shared/skill-design-philosophy.md — a shared reference doc that the analyzer (step 7), optimizer (step 8), and validator (step 9) subagents all consult when reasoning about skill quality. Previously these three subagents each carried their own implicit notion of "what makes a skill good" and "what makes an improvement principled vs ducktape"; this doc makes the vocabulary shared and grounded. Research sources behind the synthesis (documented in the doc's Sources section with URLs): - Anthropic: Agent Skills overview + best-practices + Claude Code skills guide + skill-creator + superpowers writing-skills - OpenAI: Apps SDK + GPT-4.1 prompting guide + Custom GPT instructions guide + function-calling docs - Google: Gemini API system instructions + function calling + Gemini CLI extension best practices + GEMINI.md docs - Community: SwirlAI analysis, awesome-cursorrules, awesome- agent-skills, leaked-system-prompt analyses Doc structure: - 12 core principles (vendor-convergent) with source citations - 14 anti-patterns to always flag - Principled-vs-ducktape rubric (8-row comparison table) - Per-subagent usage guidance Wired into: - analyzer.md, optimizer.md, validator.md subagent prompts — each gets a "Read first" line at the top pointing at the doc and explaining how their specific output maps to it - workflow.md shared-docs list — new entry alongside the four existing shared docs Co-Authored-By: Claude Opus 4.7 --- .../skill-design-philosophy.md | 218 ++++++++++++++++++ skills/skill-optimizer-shared/workflow.md | 4 + skills/skill-optimizer-subagents/analyzer.md | 9 + skills/skill-optimizer-subagents/optimizer.md | 10 + skills/skill-optimizer-subagents/validator.md | 9 + 5 files changed, 250 insertions(+) create mode 100644 skills/skill-optimizer-shared/skill-design-philosophy.md diff --git a/skills/skill-optimizer-shared/skill-design-philosophy.md b/skills/skill-optimizer-shared/skill-design-philosophy.md new file mode 100644 index 0000000..d52adaa --- /dev/null +++ b/skills/skill-optimizer-shared/skill-design-philosophy.md @@ -0,0 +1,218 @@ +# Skill design philosophy + +A cross-vendor synthesis of what makes an agent skill **good** and what +makes an improvement **principled** vs **ducktape**. Loaded by the +analyzer (step 7), optimizer (step 8), and validator (step 9) +subagents — these are the three places in the chain that reason +about skill quality, so they share this reference instead of each +re-deriving the criteria. + +Sources behind the synthesis: Anthropic Agent Skills docs + +skill-creator, OpenAI Apps SDK + GPT-4.1 prompting guide + Custom +GPT instructions, Google Gemini API + CLI extension best +practices, and convergent community patterns (Anthropic skills +collection, awesome-cursorrules, SwirlAI analysis, etc.). Specific +URLs in the "Sources" section at the bottom. + +## Core principles (vendor-convergent) + +1. **Progressive disclosure is the organizing principle.** Metadata + always loaded (~100 tokens), SKILL.md body on trigger (target + <500 lines), bundled references / scripts only when needed. + Heavy reference material belongs in flat one-level + `references/.md` files, NOT inlined. *Sources: + Anthropic Agent Skills overview; Microsoft `agent-skills`; + SwirlAI.* + +2. **The description IS the triggering contract.** Third person, + starts with "Use when…", names concrete triggering conditions + (symptoms, error messages, contexts). It must answer "should + Claude load this skill right now?" — not "what does this skill + do?". Workflow summaries in descriptions cause shortcuts where + the agent follows the description and skips the body. *Sources: + superpowers writing-skills CSO section; Anthropic skill-creator; + OpenAI Apps SDK ("Use this when…"); Gemini API ("Be extremely + clear and specific").* + +3. **Conciseness is public stewardship.** Every paragraph competes + with conversation history and other skills for context. Challenge + each paragraph: "Does Claude already know this? Does this + justify its cost?" Remove a sentence and ask whether Claude + would make mistakes without it. *Source: Anthropic best + practices, "Concise is key".* + +4. **Specificity matches task fragility.** High-stakes / fragile + operations (destructive edits, security boundaries) need + procedural specifics. Flexible / creative operations tolerate + high-level guidance and degrade under over-prescription. The + common framing: narrow bridge needs handrails, open field + doesn't. *Source: Anthropic best practices, "Degrees of + freedom".* + +5. **Explain WHY, not just WHAT.** Smart models reason past rote + instructions when they understand intent. "ALWAYS"/"NEVER" + without rationale is a yellow flag; reframing to explain the + underlying reason is usually more effective AND more concise. + *Sources: Anthropic skill-creator ("explain to the model why + things are important in lieu of heavy-handed musty MUSTs"); + superpowers writing-skills; Gemini API ("avoid persuasive or + flowery language").* + +6. **Procedural + declarative hybrid.** Pure procedural scripting + under-performs on planning. Pure declarative under-performs on + agentic execution. Effective skills mix: declarative reasoning + constraints ("analyze before acting") plus procedural numbered + steps where order matters. *Source: Gemini API agentic + instruction template.* + +7. **Evaluation-driven content.** Write skill content (or + improvements) only for gaps that have been measured. Build + evals first, establish a baseline (agent without the skill), + then add the minimal content that closes the gap. Anticipatory + content for hypothetical edge cases is bloat. *Sources: + Anthropic best practices ("Build evaluations first"); + superpowers writing-skills (TDD for documentation).* + +8. **One excellent example beats many mediocre ones.** Multiple + examples in the same shape risk the model copying them + verbatim. Pick one canonical example, make it complete and + runnable, explain WHY each part exists. *Sources: superpowers + writing-skills; OpenAI GPT-4.1 guide ("models can use those + quotes verbatim … vary them as necessary"); Gemini API ("Few-shot + examples must be format-consistent").* + +9. **Single-action granularity, not kitchen-sink.** Split read and + write into separate skills/tools so confirmation flows work. + Overlapping triggers between sibling skills is a primary + structural weakness — the model hesitates or picks wrong. + *Sources: OpenAI Apps SDK ("keep each tool focused on a single + read or write action"); Anthropic Agent Skills.* + +10. **Knowledge vs. instructions separation.** Reference material + (data, API docs, tables) belongs in attached files. Behavioral + rules and workflow guidance belong in the instruction body. + Mixing them dilutes both. *Source: OpenAI Custom GPT + guidelines.* + +11. **Executable scripts > generated code.** For deterministic + operations, a pre-written utility script (in `scripts/`) is + more reliable and saves tokens compared to asking the agent + to write the same code inline. *Source: Anthropic best + practices, "Provide utility scripts".* + +12. **Validation loops > one-shot execution.** For multi-step or + destructive operations, include verifiable intermediate + outputs (JSON checklists, plan files) that can be checked + before the next step. Catches errors before the cost is paid. + *Source: Anthropic best practices, "Implement feedback loops".* + +## Anti-patterns (always flag) + +- **Workflow summaries in the description field.** Causes the + agent to follow the description and skip the body. +- **Narrative bloat.** Explaining what Claude already knows + (e.g., "PDFs are documents that contain text…"). +- **MUST/NEVER without rationale** for non-critical operations. +- **Magic numbers** in scripts or skill prose, with no comment + explaining why this number. +- **Try blocks that swallow + re-throw** ("punt to Claude" — push + error reasoning onto the agent without giving it information). +- **Multi-language dilution.** Re-implementing the same example + in JS + Python + Go. One excellent example > three mediocre. +- **Inconsistent terminology** ("field"/"box"/"element" mixed) — + degrades retrieval and confuses the agent. +- **Many options without a default.** "You can use X, Y, or Z" — + pick one and provide an escape hatch. +- **Time-sensitive content** ("before August 2025") in the main + body — relegate to a collapsible "old patterns" section. +- **Tool/skill descriptions without a "Use when…" clause.** +- **Persuasive language** ("critically important", "you MUST + always", emoji), instead of clean explanation. +- **Active skill set >20** (per Gemini's bound) — at that size, + triggering accuracy degrades and overlap becomes the dominant + failure mode. +- **Monolithic SKILL.md** that should have been decomposed via + `references/` or sibling skills. +- **First-person voice** in descriptions ("I help you with X"). + +## Principled vs. ducktape — the operational test + +The chain's anti-ducktape architecture rests on this distinction. +When judging a proposed improvement: + +| Ducktape (symptomatic patch) | Principled (mechanism change) | +|---|---| +| Adds tokens to "make it more clear" | Removes tokens not pulling their weight, OR adds tokens that close a measured eval gap | +| Adds bare MUST/NEVER for one observed failure | Reframes the rule with reasoning, or restructures hierarchy | +| Inlines more content in existing structure | Decomposes into flat `references/.md` | +| Patches the specific failing input shape | Adds an eval that captures the failure mode AND a general principle | +| Inflates description to be "more discoverable" | Adds concrete triggering symptoms / exclusions | +| Generates code ad-hoc each invocation | Bundles a script in `scripts/` and tells the agent to use it | +| Adds repetitive emphasis ("CRITICAL: NEVER…") | Explains WHY the rule matters and where the cost lands | +| Adds the patch for THIS bench's failures | Names the structural weakness any reasonable bench would surface | + +The convergent test: **does the change remove tokens that aren't +pulling their weight, or add tokens that close a measured eval +gap?** If neither — it's bloat at best, ducktape at worst. + +A second test, from Anthropic's skill-creator: **"if there's some +stubborn issue, you might try branching out and using different +metaphors, or recommending different patterns of working"** — i.e., +the principled response to recurring failure is often to REFRAME, +not to add more constraints. + +## How the chain subagents use this doc + +- **Analyzer (step 7).** When naming a structural weakness, the + "What WOULD address this" principle should align with one or + more core principles above. The "What WOULD NOT address this" + anti-pattern list should call out specific ducktape moves from + the rubric. + +- **Optimizer (step 8).** The proposed change should map to a + named principle. The self-check against the anti-pattern list + must walk each ducktape move the analyzer flagged AND check + against the universal anti-patterns in this doc. + +- **Validator (step 9).** The verdict reasoning should reference + which principles the change embodies and which anti-patterns it + avoids. "Additive vs destructive" and "general vs ducktape" are + rubric-table rows; use them as concrete tests, not abstractions. + +This doc is the shared vocabulary. If a subagent's rationale +doesn't connect to a principle or anti-pattern named here, that +rationale is probably under-grounded and the operator should ask +for revision. + +## Sources + +**Anthropic:** + +- [Agent Skills overview](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview) +- [Agent Skills best practices](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/best-practices) +- [Claude Code skills guide](https://docs.claude.com/en/docs/claude-code/skills) +- [anthropics/skills (skill-creator)](https://github.com/anthropics/skills) +- superpowers writing-skills (local plugin) — TDD-for-documentation framing + +**OpenAI:** + +- [GPT-4.1 prompting guide](https://cookbook.openai.com/examples/gpt4-1_prompting_guide) +- [Apps SDK — define tools](https://developers.openai.com/apps-sdk/plan/tools) +- [Function calling](https://platform.openai.com/docs/guides/function-calling) +- [Key guidelines for Custom GPT instructions](https://help.openai.com/en/articles/9358033) +- [A practical guide to building agents (PDF)](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) + +**Google Gemini:** + +- [Prompt design strategies](https://ai.google.dev/gemini-api/docs/prompting-strategies) +- [System instructions](https://ai.google.dev/gemini-api/docs/system-instructions) +- [Function calling](https://ai.google.dev/gemini-api/docs/function-calling) +- [Gemini CLI extension best practices](https://geminicli.com/docs/extensions/best-practices/) +- [GEMINI.md context files](https://geminicli.com/docs/cli/gemini-md/) + +**Community:** + +- [SwirlAI: Agent Skills Progressive Disclosure](https://www.newsletter.swirlai.com/p/agent-skills-progressive-disclosure) +- [VoltAgent/awesome-agent-skills](https://github.com/VoltAgent/awesome-agent-skills) +- [PatrickJS/awesome-cursorrules](https://github.com/PatrickJS/awesome-cursorrules) +- [HackerNoon: The Moat is a Config File — leaked system prompt analysis](https://hackernoon.com/the-moat-is-a-config-file-analysis-of-leaked-system-prompts-from-openai-anthropic-google-and-more) diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/skill-optimizer-shared/workflow.md index 8ae1f45..13e95a4 100644 --- a/skills/skill-optimizer-shared/workflow.md +++ b/skills/skill-optimizer-shared/workflow.md @@ -153,3 +153,7 @@ need them: runtime facts vs. history rule - [`workbench.md`](./workbench.md) — workbench schema (probe layout, grader contract, smoke-check format); referenced by step 4 +- [`skill-design-philosophy.md`](./skill-design-philosophy.md) — + cross-vendor synthesis of skill-quality principles, anti-patterns, + and the principled-vs-ducktape rubric; referenced by analyzer + (step 7), optimizer (step 8), validator (step 9) diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/skill-optimizer-subagents/analyzer.md index ad32d4d..074a53b 100644 --- a/skills/skill-optimizer-subagents/analyzer.md +++ b/skills/skill-optimizer-subagents/analyzer.md @@ -8,6 +8,15 @@ step 8 (improve) refuses to fire unless this report names at least one structural weakness with the general principle that WOULD address it AND the anti-patterns that would NOT. +**Read first:** +[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +— the cross-vendor synthesis of what makes a skill good and what +makes an improvement principled vs ducktape. Your "What WOULD +address this" principle and "What WOULD NOT address this" +anti-pattern list should map to specific entries in that doc. If +your reasoning doesn't connect to a named principle or anti-pattern +there, your weakness framing is probably under-grounded. + ## Inputs (templated by the operator session) - `${SUMMARY_PATH}` — `06-bench-summary.md` (entry point; diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/skill-optimizer-subagents/optimizer.md index ae0fa1f..8c57110 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/skill-optimizer-subagents/optimizer.md @@ -6,6 +6,16 @@ principled fix for a named structural weakness from proposal is the SOLE output; you do NOT materialize the improved skill (that's step 9's job after the validator approves). +**Read first:** +[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +— the cross-vendor synthesis of what makes a skill good and what +makes an improvement principled vs ducktape. Your proposed change +should map to one or more named principles. Your self-check +against the anti-pattern list (see "Output" below) must walk both +the analyzer's named anti-patterns AND the universal ducktape +moves listed in the philosophy doc's "Principled vs ducktape" +rubric. + ## Inputs (templated by the operator session) - `${ANALYSIS_PATH}` — `07-analysis.md` (named weaknesses + what diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/skill-optimizer-subagents/validator.md index 0bda4b9..162adde 100644 --- a/skills/skill-optimizer-subagents/validator.md +++ b/skills/skill-optimizer-subagents/validator.md @@ -6,6 +6,15 @@ judge** an improvement proposal (from step 8) and write session materializes `docs/skill-optimizer//improved-skill/` from the proposal. On `needs-revision`/`reject`, the proposal is sent back to step 8. +**Read first:** +[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +— the cross-vendor synthesis of what makes a skill good and what +makes an improvement principled vs ducktape. Your verdict +reasoning should cite which principles the change embodies and +which anti-patterns it avoids. "Additive vs destructive" and +"general vs ducktape" are rows in the rubric — use them as +concrete tests, not abstractions. + Your independence is the chain's load-bearing property at this gate. The optimizer at step 8 produced the proposal with incentive to have it accepted; you must judge as if seeing the From 1b2da2809e5226e3aa49c24ea9432c65a3f79e8d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 07:36:17 -0500 Subject: [PATCH 060/121] refactor(philosophy): drop human-side provenance from agent-facing doc MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first draft included content for the curator (me), not for the analyzer/optimizer/validator subagents that load this doc each chain run: - "Sources behind the synthesis" framing paragraph at the top - Per-principle italic "*Sources: ...*" citations (12 of them) - Bottom "## Sources" URL list (Anthropic/OpenAI/Google/Community) - Attribution-style quotes ("from Anthropic's skill-creator: ...") The agent applies principles on their merit, not on authority. Citations are useful for the curator deciding whether to trust the principle, not for the subagent applying it. They also cost tokens on every dispatch. Trimmed: - Header to one sentence stating who reads this and when - Principles to lean wording, no source attributions - "Anti-pattern: active skill set >20" no longer attributes to Gemini (the bound stands on its own) - "Second test" no longer cites Anthropic; the test stands on its own merit Net: 219 → 145 lines (-34%). Same principles, same anti-patterns, same rubric, same per-subagent usage — just without the provenance baggage. Sources preserved in git history at commit fe79756 if curator review is ever needed. Co-Authored-By: Claude Opus 4.7 --- .../skill-design-philosophy.md | 139 +++++------------- 1 file changed, 33 insertions(+), 106 deletions(-) diff --git a/skills/skill-optimizer-shared/skill-design-philosophy.md b/skills/skill-optimizer-shared/skill-design-philosophy.md index d52adaa..2d5a5aa 100644 --- a/skills/skill-optimizer-shared/skill-design-philosophy.md +++ b/skills/skill-optimizer-shared/skill-design-philosophy.md @@ -1,110 +1,80 @@ # Skill design philosophy -A cross-vendor synthesis of what makes an agent skill **good** and what -makes an improvement **principled** vs **ducktape**. Loaded by the -analyzer (step 7), optimizer (step 8), and validator (step 9) -subagents — these are the three places in the chain that reason -about skill quality, so they share this reference instead of each -re-deriving the criteria. - -Sources behind the synthesis: Anthropic Agent Skills docs + -skill-creator, OpenAI Apps SDK + GPT-4.1 prompting guide + Custom -GPT instructions, Google Gemini API + CLI extension best -practices, and convergent community patterns (Anthropic skills -collection, awesome-cursorrules, SwirlAI analysis, etc.). Specific -URLs in the "Sources" section at the bottom. - -## Core principles (vendor-convergent) +Read by the analyzer (step 7), optimizer (step 8), and validator +(step 9) when reasoning about skill quality. Every per-weakness +principle, every proposed change, and every verdict rationale +should connect to a specific entry below — if it doesn't, the +reasoning is under-grounded. + +## Core principles 1. **Progressive disclosure is the organizing principle.** Metadata always loaded (~100 tokens), SKILL.md body on trigger (target <500 lines), bundled references / scripts only when needed. Heavy reference material belongs in flat one-level - `references/.md` files, NOT inlined. *Sources: - Anthropic Agent Skills overview; Microsoft `agent-skills`; - SwirlAI.* + `references/.md` files, NOT inlined. 2. **The description IS the triggering contract.** Third person, starts with "Use when…", names concrete triggering conditions (symptoms, error messages, contexts). It must answer "should Claude load this skill right now?" — not "what does this skill do?". Workflow summaries in descriptions cause shortcuts where - the agent follows the description and skips the body. *Sources: - superpowers writing-skills CSO section; Anthropic skill-creator; - OpenAI Apps SDK ("Use this when…"); Gemini API ("Be extremely - clear and specific").* + the agent follows the description and skips the body. 3. **Conciseness is public stewardship.** Every paragraph competes with conversation history and other skills for context. Challenge each paragraph: "Does Claude already know this? Does this justify its cost?" Remove a sentence and ask whether Claude - would make mistakes without it. *Source: Anthropic best - practices, "Concise is key".* + would make mistakes without it. 4. **Specificity matches task fragility.** High-stakes / fragile operations (destructive edits, security boundaries) need procedural specifics. Flexible / creative operations tolerate high-level guidance and degrade under over-prescription. The common framing: narrow bridge needs handrails, open field - doesn't. *Source: Anthropic best practices, "Degrees of - freedom".* + doesn't. 5. **Explain WHY, not just WHAT.** Smart models reason past rote instructions when they understand intent. "ALWAYS"/"NEVER" without rationale is a yellow flag; reframing to explain the underlying reason is usually more effective AND more concise. - *Sources: Anthropic skill-creator ("explain to the model why - things are important in lieu of heavy-handed musty MUSTs"); - superpowers writing-skills; Gemini API ("avoid persuasive or - flowery language").* 6. **Procedural + declarative hybrid.** Pure procedural scripting under-performs on planning. Pure declarative under-performs on agentic execution. Effective skills mix: declarative reasoning constraints ("analyze before acting") plus procedural numbered - steps where order matters. *Source: Gemini API agentic - instruction template.* + steps where order matters. 7. **Evaluation-driven content.** Write skill content (or improvements) only for gaps that have been measured. Build evals first, establish a baseline (agent without the skill), then add the minimal content that closes the gap. Anticipatory - content for hypothetical edge cases is bloat. *Sources: - Anthropic best practices ("Build evaluations first"); - superpowers writing-skills (TDD for documentation).* + content for hypothetical edge cases is bloat. 8. **One excellent example beats many mediocre ones.** Multiple examples in the same shape risk the model copying them verbatim. Pick one canonical example, make it complete and - runnable, explain WHY each part exists. *Sources: superpowers - writing-skills; OpenAI GPT-4.1 guide ("models can use those - quotes verbatim … vary them as necessary"); Gemini API ("Few-shot - examples must be format-consistent").* + runnable, explain WHY each part exists. 9. **Single-action granularity, not kitchen-sink.** Split read and write into separate skills/tools so confirmation flows work. Overlapping triggers between sibling skills is a primary structural weakness — the model hesitates or picks wrong. - *Sources: OpenAI Apps SDK ("keep each tool focused on a single - read or write action"); Anthropic Agent Skills.* 10. **Knowledge vs. instructions separation.** Reference material (data, API docs, tables) belongs in attached files. Behavioral rules and workflow guidance belong in the instruction body. - Mixing them dilutes both. *Source: OpenAI Custom GPT - guidelines.* + Mixing them dilutes both. 11. **Executable scripts > generated code.** For deterministic operations, a pre-written utility script (in `scripts/`) is more reliable and saves tokens compared to asking the agent - to write the same code inline. *Source: Anthropic best - practices, "Provide utility scripts".* + to write the same code inline. 12. **Validation loops > one-shot execution.** For multi-step or destructive operations, include verifiable intermediate outputs (JSON checklists, plan files) that can be checked before the next step. Catches errors before the cost is paid. - *Source: Anthropic best practices, "Implement feedback loops".* ## Anti-patterns (always flag) @@ -128,9 +98,8 @@ URLs in the "Sources" section at the bottom. - **Tool/skill descriptions without a "Use when…" clause.** - **Persuasive language** ("critically important", "you MUST always", emoji), instead of clean explanation. -- **Active skill set >20** (per Gemini's bound) — at that size, - triggering accuracy degrades and overlap becomes the dominant - failure mode. +- **Active skill set >20** — at that size, triggering accuracy + degrades and overlap becomes the dominant failure mode. - **Monolithic SKILL.md** that should have been decomposed via `references/` or sibling skills. - **First-person voice** in descriptions ("I help you with X"). @@ -155,64 +124,22 @@ The convergent test: **does the change remove tokens that aren't pulling their weight, or add tokens that close a measured eval gap?** If neither — it's bloat at best, ducktape at worst. -A second test, from Anthropic's skill-creator: **"if there's some -stubborn issue, you might try branching out and using different -metaphors, or recommending different patterns of working"** — i.e., -the principled response to recurring failure is often to REFRAME, -not to add more constraints. - -## How the chain subagents use this doc - -- **Analyzer (step 7).** When naming a structural weakness, the - "What WOULD address this" principle should align with one or - more core principles above. The "What WOULD NOT address this" - anti-pattern list should call out specific ducktape moves from - the rubric. - -- **Optimizer (step 8).** The proposed change should map to a - named principle. The self-check against the anti-pattern list - must walk each ducktape move the analyzer flagged AND check - against the universal anti-patterns in this doc. - -- **Validator (step 9).** The verdict reasoning should reference - which principles the change embodies and which anti-patterns it - avoids. "Additive vs destructive" and "general vs ducktape" are - rubric-table rows; use them as concrete tests, not abstractions. - -This doc is the shared vocabulary. If a subagent's rationale -doesn't connect to a principle or anti-pattern named here, that -rationale is probably under-grounded and the operator should ask -for revision. - -## Sources - -**Anthropic:** - -- [Agent Skills overview](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview) -- [Agent Skills best practices](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/best-practices) -- [Claude Code skills guide](https://docs.claude.com/en/docs/claude-code/skills) -- [anthropics/skills (skill-creator)](https://github.com/anthropics/skills) -- superpowers writing-skills (local plugin) — TDD-for-documentation framing - -**OpenAI:** - -- [GPT-4.1 prompting guide](https://cookbook.openai.com/examples/gpt4-1_prompting_guide) -- [Apps SDK — define tools](https://developers.openai.com/apps-sdk/plan/tools) -- [Function calling](https://platform.openai.com/docs/guides/function-calling) -- [Key guidelines for Custom GPT instructions](https://help.openai.com/en/articles/9358033) -- [A practical guide to building agents (PDF)](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) +A second test: **the principled response to recurring failure is +often to REFRAME, not to add more constraints.** If you find +yourself stacking MUST clauses against a stubborn issue, try a +different metaphor or restructure the section instead. -**Google Gemini:** +## How each subagent uses this doc -- [Prompt design strategies](https://ai.google.dev/gemini-api/docs/prompting-strategies) -- [System instructions](https://ai.google.dev/gemini-api/docs/system-instructions) -- [Function calling](https://ai.google.dev/gemini-api/docs/function-calling) -- [Gemini CLI extension best practices](https://geminicli.com/docs/extensions/best-practices/) -- [GEMINI.md context files](https://geminicli.com/docs/cli/gemini-md/) +- **Analyzer (step 7).** "What WOULD address this" must align with + one or more core principles above. "What WOULD NOT address this" + must call out specific ducktape moves from the rubric. -**Community:** +- **Optimizer (step 8).** The proposed change maps to one or more + named principles. The self-check walks the analyzer's named + anti-patterns AND the universal anti-pattern list here. -- [SwirlAI: Agent Skills Progressive Disclosure](https://www.newsletter.swirlai.com/p/agent-skills-progressive-disclosure) -- [VoltAgent/awesome-agent-skills](https://github.com/VoltAgent/awesome-agent-skills) -- [PatrickJS/awesome-cursorrules](https://github.com/PatrickJS/awesome-cursorrules) -- [HackerNoon: The Moat is a Config File — leaked system prompt analysis](https://hackernoon.com/the-moat-is-a-config-file-analysis-of-leaked-system-prompts-from-openai-anthropic-google-and-more) +- **Validator (step 9).** The verdict cites which principles the + change embodies and which anti-patterns it avoids. "Additive vs + destructive" and "general vs ducktape" are concrete rubric tests, + not abstractions. From 1f79681dbcd0734c37af8f67669e88093e695827 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 07:42:44 -0500 Subject: [PATCH 061/121] docs(philosophy): expand to 4-vendor survey + synthesis rationale MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The existing docs/skill-writing-philosophy.md was Claude-only (Anthropic official + skill-creator + superpowers writing-skills). Expanded to a four-vendor survey backing the lean agent-facing skills/skill-optimizer-shared/skill-design-philosophy.md. Added: - Per-vendor "findings" sections for OpenAI (Apps SDK, GPT-4.1 guide, Custom GPT instructions, function calling) and Google Gemini (API system instructions, function calling, CLI extension best practices, GEMINI.md guidance) with their distinctive principles - "Community" section covering SwirlAI, awesome-cursorrules, awesome-agent-skills, leaked-prompt analyses - "Cross-vendor consensus" + "Cross-vendor divergence" sections summarizing where the four buckets agree and where they pick different stances - "Synthesis" section with a per-principle source table showing which principles got promoted into the shipped doc and why, plus what was deliberately excluded - "Open questions" section flagging cross-vendor portability, the "explain WHY" vs literal-following tension, tool-count bound implications for chain extensions, and self-application - References organized by vendor bucket Kept (still load-bearing): - "Skill type → philosophy mapping" 5-row table - "Bootstrapping limit" section - "Description: triggers, not summaries" section - "When the optimizer is revising someone else's skill" 6-rule guidance - Decision-tree at the end This doc is the human-facing companion to the agent-facing philosophy doc — explains provenance, lets a curator audit which principles came from where, and provides the type-mapping + meta- guidance the chain's analyzer/optimizer/validator subagents don't need at dispatch time. Co-Authored-By: Claude Opus 4.7 --- docs/skill-writing-philosophy.md | 534 ++++++++++++++++++++++++------- 1 file changed, 424 insertions(+), 110 deletions(-) diff --git a/docs/skill-writing-philosophy.md b/docs/skill-writing-philosophy.md index 8f9c539..bd8d71f 100644 --- a/docs/skill-writing-philosophy.md +++ b/docs/skill-writing-philosophy.md @@ -4,53 +4,305 @@ > and before designing the optimizer subagent's prompt — choosing the > wrong philosophy for the skill type is itself a form of ducktape. +The lean agent-facing distillation lives at +[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skills/skill-optimizer-shared/skill-design-philosophy.md). +That doc is what the chain's analyzer/optimizer/validator subagents +load each run — terse, no provenance, just the rules. **This** doc +is the research backing it: per-vendor findings, where vendors +agree or diverge, and the rationale for which principles got +promoted into the shipped distillation. + ## Background -There are three named bodies of guidance on writing Claude Agent Skills. -They mostly agree, but they diverge on tone and testing rigor in ways -that matter: - -1. **Anthropic official** — the canonical [skill authoring best - practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices). -2. **skill-creator** — Anthropic's plugin for iterative skill authoring - with an eval-viewer + description-optimizer loop. -3. **superpowers:writing-skills** — third-party plugin (obra/superpowers) - that frames skill authoring as TDD applied to process documentation. - -## Consensus (load-bearing — all three sources agree) - -- **The description is the trigger mechanism.** Claude scans - descriptions to decide which skills to consult. Get this right or the - skill never fires. -- **SKILL.md stays ≤ ~500 lines; bulk goes in `references/`, scripts in - `scripts/`.** This is progressive disclosure. -- **References are one level deep from SKILL.md** — nested references - cause partial reads and lost info. -- **Evaluation-driven development beats imagined-requirements - development.** Run baseline tasks, document where the agent fails, - write the skill to address those specific failures. -- **Use Claude to iterate on Claude's skills** — one instance drafts, - another tests, observations feed back to the drafter. - -## Divergence (where you must pick) - -| Axis | Anthropic + skill-creator | superpowers:writing-skills | +Four bodies of guidance on writing agent skills were surveyed: + +1. **Anthropic** — official Agent Skills docs ([best + practices](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/best-practices)), + the [anthropics/skills](https://github.com/anthropics/skills) + repo (skill-creator plugin), and the third-party + [superpowers writing-skills](https://github.com/obra/superpowers) + that frames skill authoring as TDD applied to process docs. +2. **OpenAI** — [Apps SDK tool design](https://developers.openai.com/apps-sdk/plan/tools), + [GPT-4.1 prompting guide](https://cookbook.openai.com/examples/gpt4-1_prompting_guide), + [Custom GPT instructions guide](https://help.openai.com/en/articles/9358033), + [function calling](https://platform.openai.com/docs/guides/function-calling). +3. **Google Gemini** — [system instructions](https://ai.google.dev/gemini-api/docs/system-instructions), + [function calling](https://ai.google.dev/gemini-api/docs/function-calling), + [Gemini CLI extension best practices](https://geminicli.com/docs/extensions/best-practices/), + [GEMINI.md context files](https://geminicli.com/docs/cli/gemini-md/). +4. **Community** — [SwirlAI analysis](https://www.newsletter.swirlai.com/p/agent-skills-progressive-disclosure), + [awesome-cursorrules](https://github.com/PatrickJS/awesome-cursorrules), + [awesome-agent-skills](https://github.com/VoltAgent/awesome-agent-skills), + leaked-system-prompt analyses, Cursor `.cursorrules` ecosystem. + +Full source URLs in References at the bottom. + +## Per-vendor findings + +### Anthropic + +The Claude ecosystem has the most-developed skill philosophy. Three +named bodies of guidance, mostly agreeing, diverging on tone and +testing rigor: + +- **Anthropic official** treats skills as reference docs for an + agent: progressive disclosure, concise prose, explain WHY rather + than stack MUSTs, build evals first. +- **skill-creator** is Anthropic's plugin for iterative skill + authoring — adds an eval-viewer + description-optimizer loop. Its + most distinctive contribution: be a little "pushy" in descriptions + because Claude tends to under-trigger ("Make sure to use this + skill whenever the user mentions dashboards…"). +- **superpowers writing-skills** is third-party and frames skill + authoring as TDD applied to documentation. Its strongest empirical + finding: workflow summaries in the description cause agents to + follow the description and skip the body. Switch to triggering + conditions only. + +#### Anthropic divergence axis: discipline-enforcing skills + +| Axis | Anthropic + skill-creator | superpowers writing-skills | |---|---|---| | Tone for must-follow rules | "Explain the WHY; MUSTs/NEVERs in all caps are a yellow flag" | "Authority framing, no exceptions, close every loophole" | | What gets a skill | Workflows, techniques, references | Same, **plus** discipline-enforcing rules (TDD, verification) | | Testing rigor | ≥3 evals (Anthropic) / full with-skill vs baseline pipeline (skill-creator) | Pressure scenarios with combined pressures (time + sunk cost + authority) | | Description content | "Include both what + when" | "**Only** when, **never** what — the body becomes documentation Claude skips otherwise" | -**The divergence is not a contradiction.** writing-skills explicitly +The divergence is not a contradiction. writing-skills explicitly says: discipline-enforcing skills get authority framing + pressure tests; reference / technique skills get application tests. The other two sources mostly assume the workflow / technique case and don't address discipline rules. +### OpenAI + +OpenAI's framing is more literal and prescriptive than Anthropic's +"explain why" stance. Distinctive principles: + +- **Literal-instruction-following.** GPT-4.1 and later are + *"trained to follow instructions more closely and more literally + than its predecessors."* State behavior unambiguously; don't rely + on the model to infer intent. This *contradicts* Anthropic's + preference for terse-with-reasoning — OpenAI says be explicit + even when it feels redundant. +- **Positive instructions over prohibitions.** Custom GPT guide: + *"Prefer positive, concrete instructions ('Do X') over long lists + of prohibitions ('Don't do Y') when possible."* +- **Single-action granularity.** Apps SDK: *"keep each tool focused + on a single read or write action ('fetch_board', 'create_ticket'), + rather than a kitchen-sink endpoint."* Read and write should be + separate tools so confirmation flows work. +- **Tool description shape.** Apps SDK mandates one-or-two-sentence + descriptions that *"start with 'Use this when…' so the model + knows exactly when to pick the tool."* +- **Knowledge vs. instructions separation.** Custom GPT guide: + *"Use knowledge for reference material, not rules or behavior. + Put rules, tone, and workflow guidance in instructions."* +- **Structured prompt hierarchy.** GPT-4.1 guide recommends a + canonical order: Role → Instructions → Reasoning Steps → Output + Format → Examples → Context. For long contexts, put instructions + at both ends. +- **Three-line agentic harness.** *"Keep going until resolved", + "use tools instead of guessing", "plan before each call and + reflect after"* — produced ~20% SWE-bench gain. +- **Tool-count failure mode.** Up to ~100 tools is "in-distribution", + but the named failure mode is overlap: *"If multiple tools have + overlapping purposes or vague descriptions, models may call the + wrong one or hesitate to call any at all."* +- **Examples sit in their own section, not in tool descriptions.** + *"When provided sample phrases, models can use those quotes + verbatim and start to sound repetitive."* Instruct the model to + vary them. + +Anti-patterns OpenAI explicitly names: bribes/all-caps/tip threats +(*"not necessary"*), conflicting instructions (silently resolved in +favor of the one closer to the end), forced tool loops, conflated +read/write tools. + +### Google Gemini + +Gemini's framing emphasizes terseness, hierarchical context, and +hard numeric bounds. Distinctive principles: + +- **Terse default.** *"Gemini 3 models provide direct and efficient + answers. If you need more conversational or detailed response, + you must explicitly request it."* Don't assume verbose output — + instructions that don't ask for elaboration shouldn't expect it. +- **Critical instructions first; long-context question last.** + *"Place essential behavioral constraints and role definitions at + the beginning."* For long-context prompts, supply documents first + and put the actual instruction at the very end with a transition + phrase. *Divergence*: Anthropic emphasizes consistent positioning; + Google is more prescriptive about question-at-end for long-context. +- **Hierarchical, modular context.** Three tiers: global + (`~/.gemini/GEMINI.md`), workspace, just-in-time (file-adjacent). + Large files should be decomposed using `@file.md` imports. + Monolithic single-file context is an anti-pattern. +- **Hard 10-20 tool cap.** *"Providing too many can increase the + risk of selecting an incorrect or suboptimal tool. Aim to provide + only the relevant tools for the context or task, ideally keeping + the active set to a maximum of 10-20."* Numeric bound where + others are vague. +- **Constrain with schema, not prose.** *"Use 'enum' to list the + allowed values instead of just describing them in the + description."* Prefer `integer` over `number` when applicable. +- **Few-shot examples format-consistent.** Inconsistency across + examples actively degrades behavior. +- **Procedural + declarative hybrid endorsed.** Google's agentic + system-instruction template combines declarative reasoning steps + ("analyze constraints", "evaluate consequences") with procedural + numbered steps. *Divergence*: explicitly endorses the hybrid, + whereas Anthropic guidance has shifted more toward declarative- + only for skills. + +Anti-patterns Gemini explicitly names: *"overly persuasive or +flowery language"*, monolithic context with no decomposition, +buried critical constraints, enum-able parameters described only +in prose, granting broad permissions when narrow ones suffice. + +### Community + +Convergent patterns across multiple unrelated skill collections +(SwirlAI, awesome-agent-skills, awesome-cursorrules, Microsoft +`agent-skills`, Vercel/community packs, leaked-prompt analyses): + +- **Progressive disclosure as the core organizing principle.** + Three-tier loading: metadata (~60-100 tokens always loaded) → + SKILL.md body (loaded on trigger, <500 lines) → bundled + references/scripts (loaded only as needed). SwirlAI: *"Context + window is the agent's cognitive space. Overloading it degrades + performance, while keeping it focused lets the agent reason + sharply."* +- **Description = triggering contract.** anthropics/skills + explicitly recommends being *"pushy"* to combat under-triggering. +- **One-level-deep references.** Universally warned-against: + nested reference chains (SKILL.md → advanced.md → details.md). + Claude does partial reads with `head -100`, losing info in deep + trees. +- **Conditional activation.** Cursor `.mdc` rules converge on + `globs:` + `alwaysApply:` — *"conditional patterns outperform + blanket instructions because they reduce noise."* +- **Evals-before-prose.** Anthropic explicit: *"Create evaluations + BEFORE writing extensive documentation. This ensures your Skill + solves real problems rather than documenting imagined ones."* + +Divergent patterns (context-dependent): + +- **Strictness of constraint language.** Anthropic warns against + ALWAYS/NEVER caps, but the same doc later recommends *"using + stronger language like 'MUST filter' instead of 'always filter'"* + when Claude under-applies a rule. Resolution in practice: match + constraint strictness to task fragility (the "narrow bridge vs. + open field" framing). +- **Skill scope — narrow vs. multi-domain.** Microsoft / SwirlAI + warn against "kitchen-sink skills"; some community packs + deliberately bundle (rohitg00/awesome-claude-code-toolkit ships + 135 agents in one). Convergent compromise: split by domain + inside one skill (Pattern 2 in Anthropic docs — + `reference/finance.md`, `reference/sales.md`). + +Community-flagged anti-patterns: + +- Verbose descriptions that summarize workflow. +- Voodoo constants / bare try-blocks that re-throw. +- Overfitting to one observed failure. +- Time-sensitive content ("before August 2025") in the main body. +- Offering many options without a default. +- Inconsistent terminology ("field"/"box"/"element" mixed). + +## Cross-vendor consensus (load-bearing — agreed across all four) + +- **Progressive disclosure.** Anthropic + OpenAI (implicit) + + Gemini + community. +- **Description = triggering contract.** Anthropic + OpenAI ("Use + this when…") + Gemini ("extremely clear and specific") + + community. +- **Conciseness / context as scarce resource.** Anthropic + ("concise is key") + Gemini ("terse default") + community + ("don't dump exhaustive docs"). +- **Evaluation-driven content.** Anthropic + superpowers (TDD) + plus community. +- **Examples teach behavior, but don't overfit.** superpowers + + OpenAI ("models can use those quotes verbatim") + Gemini + ("few-shot must be format-consistent"). + +## Cross-vendor divergence (pick your stance) + +| Axis | Anthropic side | OpenAI side | Gemini side | +|---|---|---|---| +| Tone for rules | "Explain WHY; MUSTs are a yellow flag" | "Literal; be explicit even when redundant" | "Avoid flowery / persuasive language" | +| Output verbosity default | Implicit (depends on skill) | Implicit | **Terse** — must explicitly request elaboration | +| Tool/skill count bound | Implicit (works at any count) | ~100 tools in-distribution | **Hard 10-20 active cap** | +| Long-context positioning | Implicit | "Instructions at both ends" | "Critical first, question last" | +| Procedural vs declarative | Trending declarative-only | Mixed (agentic harness is procedural) | **Explicit hybrid endorsed** | +| Type constraints | Prose acceptable | Function-calling schema | **Enum/schema > prose** | + +Where vendors disagree, the v1.4 chain takes: + +- **Tone:** Anthropic — "explain WHY" by default, allow MUSTs for + fragile/discipline-enforcing rules (per superpowers). +- **Verbosity:** match the skill type; no universal stance. +- **Tool count:** Gemini's 10-20 bound applies to the chain itself + (we ship 9 + 1 autopilot). +- **Positioning:** Anthropic conventional + Gemini's "critical + first" for long-context dispatches. +- **Procedural/declarative:** Gemini's hybrid. +- **Type constraints:** N/A for prose skills, but enums-over-prose + applies to any spec.yaml fields the chain writes. + +## Synthesis: how we reached skill-design-philosophy.md + +The shipped lean doc has 12 principles, 14 anti-patterns, and an +8-row rubric. Each entry has at least two sources backing it. +Justification per principle: + +| # | Principle | Sources | +|---|---|---| +| 1 | Progressive disclosure | Anthropic + OpenAI (implicit) + Gemini + SwirlAI + Microsoft | +| 2 | Description = triggering contract | All four buckets | +| 3 | Conciseness as public stewardship | Anthropic + Gemini + community | +| 4 | Specificity matches task fragility | Anthropic explicit ("degrees of freedom"); others implicit | +| 5 | Explain WHY, not just WHAT | Anthropic + superpowers + Gemini (anti-flowery) | +| 6 | Procedural + declarative hybrid | Gemini explicit; others implicit | +| 7 | Evaluation-driven content | Anthropic + superpowers | +| 8 | One example beats many | superpowers + OpenAI + Gemini | +| 9 | Single-action granularity | OpenAI Apps SDK explicit | +| 10 | Knowledge vs instructions separation | OpenAI Custom GPT explicit | +| 11 | Executable scripts > generated code | Anthropic | +| 12 | Validation loops > one-shot | Anthropic | + +**Principles deliberately NOT included:** + +- OpenAI's "literal-instruction-following" — contradicts the + more-effective "explain WHY" framing for most skill types. We + treat it as a fragile-task fallback under principle 4, not a + universal rule. +- OpenAI's "structured prompt hierarchy" (Role → Instructions → + Reasoning Steps → …) — useful for system prompts, less so for + skill bodies. Not load-bearing for our agent's reasoning. +- Gemini's "long-context positioning" (critical first, question + last) — useful for long-context dispatches but the chain skills + are short enough that positioning doesn't matter much. +- The 10-20 tool cap was promoted into anti-pattern 12 (active + skill set >20), not into a top-level principle, because it's + a quantitative bound rather than a quality principle. + +**Anti-patterns:** each anti-pattern in the shipped doc was named +by at least two sources. The 14 include workflow-summary +descriptions (superpowers explicit), narrative bloat (Anthropic), +MUST/NEVER without rationale (Anthropic + Gemini), magic numbers +(community), persuasive language (Gemini), monolithic SKILL.md +(SwirlAI + Anthropic), etc. + +**Principled-vs-ducktape rubric:** derived from the convergent +test *"does the change pay its way in tokens or evals?"* Each +row is a specific ducktape pattern flagged by ≥2 sources, paired +with the principled alternative also flagged by ≥2 sources. + ## Skill type → philosophy mapping -This is the single most important decision when authoring or revising -a skill. Diagnose first, write/revise second. +This is the single most important decision when authoring or +revising a skill. Diagnose first, write/revise second. | Skill type | Examples | Philosophy | Testing | |---|---|---|---| @@ -65,108 +317,109 @@ are workflows with one or two embedded discipline rules (e.g., "dispatch the subagent, do NOT do the research yourself"). The right pattern: -1. Write the workflow body in Anthropic/skill-creator style — lean, - explain why, set freedom per step. -2. For each embedded discipline rule, mark it explicitly (bold, "Do - NOT" framing at the action site) AND give a why-this-matters +1. Write the workflow body in Anthropic / skill-creator style — + lean, explain why, set freedom per step. +2. For each embedded discipline rule, mark it explicitly (bold, + "Do NOT" framing at the action site) AND give a why-this-matters paragraph nearby. 3. Reserve full writing-skills treatment (rationalization tables, - red-flags lists, no-exceptions language) for the small handful of - rules where compliance under pressure is load-bearing and was - verified via pressure scenarios. + red-flags lists, no-exceptions language) for the small handful + of rules where compliance under pressure is load-bearing and + was verified via pressure scenarios. ## The bootstrapping limit The skill-optimizer chain runs an empirical loop (eval → analyze → -improve) on target skills. But that loop cannot validate **itself** — -you can't use the chain to author its own seven SKILL.md files, -because the chain doesn't exist yet when those files are being -written. This is the same shape as Thompson's "Reflections on -Trusting Trust": validating a compiler with itself is circular. +improve) on target skills. But that loop cannot validate **itself** +— you can't use the chain to author its own SKILL.md files, because +the chain doesn't exist yet when those files are being written. +Thompson's "Reflections on Trusting Trust": validating a compiler +with itself is circular. Practical consequences: -- **The seven skill-optimizer SKILL.md files are authored from - philosophy + best judgment**, not from an eval loop. The "test" - for these skills is end-to-end runs on real targets and ongoing - observation in actual use. - -- **The optimizer subagent should not be surprised by the absence of - eval data for the skill-optimizer's own skills.** If asked to - improve one of them, it should treat absence of empirical data as - a known limit, not a gap to fill speculatively. Real-world - observations from end-to-end runs are the only valid signal for - self-improvement of the chain itself. - -- **Self-application is deferred** until the chain has earned trust - on external targets. Running skill-optimizer on the skill-optimizer - is a future-state exercise; doing it pre-launch is bootstrapping in - a loop. +- **The chain SKILL.md files are authored from philosophy + best + judgment**, not from an eval loop. The "test" for these skills + is end-to-end runs on real targets and ongoing observation in + actual use. +- **The optimizer subagent should not be surprised by the absence + of eval data for the skill-optimizer's own skills.** If asked + to improve one of them, it should treat absence of empirical + data as a known limit, not a gap to fill speculatively. +- **Self-application is deferred** until the chain has earned + trust on external targets. Running skill-optimizer on + skill-optimizer is a future-state exercise; doing it pre-launch + is bootstrapping in a loop. ## Description: triggers, not summaries writing-skills' strongest empirical finding: when a description -summarizes the skill's workflow, agents follow the summary instead of -reading the full skill body. The skill body becomes documentation the -agent skips. +summarizes the skill's workflow, agents follow the summary instead +of reading the full skill body. The skill body becomes +documentation the agent skips. -The example they cite: a description saying "code review between tasks" -caused agents to do ONE review even though the skill body clearly -specified TWO. Switching the description to pure trigger language -("Use when executing implementation plans with independent tasks in -the current session") restored compliance. +The example they cite: a description saying "code review between +tasks" caused agents to do ONE review even though the skill body +clearly specified TWO. Switching the description to pure trigger +language ("Use when executing implementation plans with independent +tasks in the current session") restored compliance. -**Practical guidance for skill-optimizer descriptions:** +Practical guidance for skill-optimizer descriptions: -- Start with "Use when ..." and list the user-language symptoms that +- Start with "Use when …" and list the user-language symptoms that should trigger this skill. -- Do NOT summarize the workflow, the dispatches, or the output format - in the description. Those go in the body. -- Use concrete trigger phrases, not abstract ones: "Use when the user - asks 'what does this skill do' or 'investigate this skill'" beats - "Use when investigating skills." +- Do NOT summarize the workflow, the dispatches, or the output + format in the description. Those go in the body. +- Use concrete trigger phrases, not abstract ones: "Use when the + user asks 'what does this skill do' or 'investigate this skill'" + beats "Use when investigating skills." - Anthropic's "both what + when" guidance is fine for simple - workflow skills with no embedded discipline rules. For skills with - embedded discipline (most of the skill-optimizer chain skills), - trigger-only is safer. -- skill-creator's "be a little pushy" guidance applies: Claude tends to - under-trigger; include adjacent phrasings explicitly. + workflow skills with no embedded discipline rules. For skills + with embedded discipline (most of the skill-optimizer chain + skills), trigger-only is safer. +- skill-creator's "be a little pushy" guidance applies: Claude + tends to under-trigger; include adjacent phrasings explicitly. ## When the optimizer is revising someone else's skill -This is the meta-payoff and the load-bearing reason this doc exists. -The Phase 7 optimizer subagent revises target skills based on weaknesses -the analyzer surfaced. To avoid ducktape, the optimizer MUST: +This is the meta-payoff and the load-bearing reason this doc +exists. The Phase 7 optimizer subagent revises target skills based +on weaknesses the analyzer surfaced. To avoid ducktape, the +optimizer MUST: -1. **Diagnose the target skill's type before proposing changes.** A - reference skill and a discipline-enforcing skill need different - fixes for the same observed failure. +1. **Diagnose the target skill's type before proposing changes.** + A reference skill and a discipline-enforcing skill need + different fixes for the same observed failure. -2. **Preserve the existing style unless the analyzer flagged the style - itself as the weakness.** If the original uses explanatory prose, - the fix uses explanatory prose. Don't impose authority framing - because MUSTs feel clearer to the optimizer — that's ducktape. +2. **Preserve the existing style unless the analyzer flagged the + style itself as the weakness.** If the original uses + explanatory prose, the fix uses explanatory prose. Don't + impose authority framing because MUSTs feel clearer to the + optimizer — that's ducktape. -3. **Bias toward "Claude is smart" — pruning beats adding.** If the - skill restates what Claude already knows, removing the restatement - is often a more principled fix than adding new rules. Token weight - competes with conversation context once the skill loads. +3. **Bias toward "Claude is smart" — pruning beats adding.** If + the skill restates what Claude already knows, removing the + restatement is often a more principled fix than adding new + rules. Token weight competes with conversation context once + the skill loads. 4. **Reserve authority framing / rationalization tables / - no-exceptions language for skills where the analyzer specifically - documented a discipline failure under pressure.** Adding MUSTs to - patch a reference-skill bug is the canonical ducktape pattern. + no-exceptions language for skills where the analyzer + specifically documented a discipline failure under pressure.** + Adding MUSTs to patch a reference-skill bug is the canonical + ducktape pattern. 5. **Test the proposed change against a scenario the analyzer documented, not a new scenario invented by the optimizer.** - Optimizer-invented scenarios drift into solving imagined problems - instead of the real surfaced weakness. + Optimizer-invented scenarios drift into solving imagined + problems instead of the real surfaced weakness. -6. **Treat description changes as separate, conservative edits.** The - description determines triggering, not behavior. Changing the - description to "fix" a behavioral problem is misdirected. Change - the body for behavior; change the description only if the analyzer - flagged a trigger problem (over-triggering or under-triggering). +6. **Treat description changes as separate, conservative edits.** + The description determines triggering, not behavior. Changing + the description to "fix" a behavioral problem is misdirected. + Change the body for behavior; change the description only if + the analyzer flagged a trigger problem (over-triggering or + under-triggering). ## When to use each philosophy: a quick decision tree @@ -193,13 +446,74 @@ If the skill is mixed (workflow with embedded discipline rules): write the body in workflow style, mark the discipline rules explicitly at their action sites, and reserve the heavy writing-skills treatment (rationalization tables etc.) for the -specific rules whose compliance under pressure was actually verified. +specific rules whose compliance under pressure was actually +verified. + +## Open questions + +- **Cross-vendor portability.** The shipped philosophy doc + assumes Claude as the operator. Whether the synthesis holds for + Codex / Gemini operators (when the chain skills run on those + platforms via the cross-agent plugin metadata) is untested. + OpenAI's "literal-instruction-following" might mean Codex + operators handle our terse-with-rationale style worse than + Claude does. Future work: re-run a chain end-to-end on Codex + and Gemini to see whether the principles still apply. +- **Where the "explain WHY" preference breaks down.** Both + OpenAI's "be literal" and Anthropic's "stronger language like + MUST" suggest there's a regime where rationale-prose + under-performs. We currently route this through principle 4 + (specificity matches fragility) but haven't validated the + threshold empirically. +- **Tool-count bound for chain skills.** Gemini's 10-20 cap is + the only quantitative bound we found. The chain ships 9 + 1 + (autopilot pending) = 10. Adding more chain steps would push + toward Gemini's degradation regime. Need to consider this when + scoping future chain extensions. +- **Self-application of the chain.** Per the bootstrapping limit + above, we can't yet use the chain to improve its own SKILL.md + files. The threshold for doing so is "enough external + validation to trust the chain's verdicts" — currently + un-quantified. ## References -- [Anthropic skill authoring best practices](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices) +**Anthropic:** + +- [Agent Skills overview](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview) +- [Agent Skills best practices](https://docs.claude.com/en/docs/agents-and-tools/agent-skills/best-practices) +- [Claude Code skills guide](https://docs.claude.com/en/docs/claude-code/skills) +- [anthropics/skills (skill-creator)](https://github.com/anthropics/skills) - skill-creator plugin: `~/.claude/plugins/cache/claude-plugins-official/skill-creator/` -- superpowers:writing-skills: - `~/.claude/plugins/cache/claude-plugins-official/superpowers//skills/writing-skills/` -- Meincke et al. 2025 — persuasion principles in AI compliance (cited - by writing-skills/persuasion-principles.md) +- superpowers:writing-skills: `~/.claude/plugins/cache/claude-plugins-official/superpowers//skills/writing-skills/` +- [Equipping Agents for the Real World with Agent Skills (engineering blog)](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills) + +**OpenAI:** + +- [GPT-4.1 prompting guide](https://cookbook.openai.com/examples/gpt4-1_prompting_guide) +- [Apps SDK — define tools](https://developers.openai.com/apps-sdk/plan/tools) +- [Function calling](https://platform.openai.com/docs/guides/function-calling) +- [Key guidelines for Custom GPT instructions](https://help.openai.com/en/articles/9358033) +- [A practical guide to building agents (PDF)](https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf) +- [Agents — OpenAI Agents SDK](https://openai.github.io/openai-agents-python/agents/) + +**Google Gemini:** + +- [Prompt design strategies](https://ai.google.dev/gemini-api/docs/prompting-strategies) +- [System instructions](https://ai.google.dev/gemini-api/docs/system-instructions) +- [Function calling](https://ai.google.dev/gemini-api/docs/function-calling) +- [Gemini CLI extension best practices](https://geminicli.com/docs/extensions/best-practices/) +- [GEMINI.md context files](https://geminicli.com/docs/cli/gemini-md/) +- [Building Gemini CLI Extensions](https://geminicli.com/docs/extensions/writing-extensions/) + +**Community:** + +- [SwirlAI: Agent Skills Progressive Disclosure](https://www.newsletter.swirlai.com/p/agent-skills-progressive-disclosure) +- [VoltAgent/awesome-agent-skills](https://github.com/VoltAgent/awesome-agent-skills) +- [karanb192/awesome-claude-skills](https://github.com/karanb192/awesome-claude-skills) +- [ComposioHQ/awesome-claude-skills](https://github.com/ComposioHQ/awesome-claude-skills) +- [PatrickJS/awesome-cursorrules](https://github.com/PatrickJS/awesome-cursorrules) +- [dev.to: 5 .cursorrules patterns that make Cursor actually reliable](https://dev.to/olivia_craft/5-cursorrules-patterns-that-make-cursor-actually-reliable-41ni) +- [HackerNoon: The Moat is a Config File — leaked system prompt analysis](https://hackernoon.com/the-moat-is-a-config-file-analysis-of-leaked-system-prompts-from-openai-anthropic-google-and-more) +- Meincke et al. 2025 — persuasion principles in AI compliance + (cited by writing-skills/persuasion-principles.md) From 7f159117ce6d6501f63b8953dbba22015b89d14d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 08:05:04 -0500 Subject: [PATCH 062/121] refactor(naming): drop skill-optimizer- prefix; co-locate subagents under each step MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit User-visible names like `skill-optimizer:skill-optimizer-analyze` had redundant prefixes — the plugin namespace already says `skill-optimizer:`, so the skill name shouldn't repeat it. Matches the superpowers convention (`superpowers:brainstorming`, not `superpowers:superpowers-brainstorming`). Also moved subagent prompts from the sibling-folder `skill-optimizer-subagents/` (idiosyncratic) into per-step `agents/` subfolders. Aligns with skill-creator's `agents/.md` convention and makes each chain step self-contained. Renames: - skills/skill-optimizer-investigate-functionality/ → skills/investigate-functionality/ - skills/skill-optimizer-investigate-submissions/ → skills/investigate-submissions/ - skills/skill-optimizer-design-tests/ → skills/design-tests/ - skills/skill-optimizer-write-tests/ → skills/write-tests/ - skills/skill-optimizer-validate-tests/ → skills/validate-tests/ - skills/skill-optimizer-run-bench/ → skills/run-bench/ - skills/skill-optimizer-analyze/ → skills/analyze/ - skills/skill-optimizer-improve/ → skills/improve/ - skills/skill-optimizer-validate/ → skills/validate/ - skills/skill-optimizer-shared/ → skills/shared/ Subagent prompts moved into each dispatching step's agents/: - research-functionality.md → investigate-functionality/agents/ - research-submissions.md → investigate-submissions/agents/ - test-designer.md → design-tests/agents/ - test-writer.md → write-tests/agents/ - test-validator.md → validate-tests/agents/ - analyzer.md → analyze/agents/ - optimizer.md → improve/agents/ - validator.md → validate/agents/ (skills/skill-optimizer-subagents/ folder removed.) Updated throughout: - name: frontmatter in each SKILL.md (now just the bare slug, no namespace prefix) - H1 headers in each SKILL.md (bare slug) - All cross-references between chain skills now use the plugin-namespaced form `skill-optimizer:` (e.g., "Next, invoke `skill-optimizer:write-tests`") - All path refs to shared/ and to per-step agents/ subagent prompts updated (relative paths now `./agents/X.md` from SKILL.md and `../../shared/Y.md` from agents/X.md) - .claude-plugin/marketplace.json skills array - tests/smoke-skill-distribution.ts (chain skill paths and the Gemini @import regex) - AGENTS.md, CLAUDE.md, README.md, CONTRIBUTING.md, GEMINI.md - docs/README.codex.md, docs/README.opencode.md - .codex/INSTALL.md, .cursor/INSTALL.md, .opencode/INSTALL.md - docs/skill-writing-philosophy.md (refs to philosophy doc) 11/11 smoke tests pass; typecheck clean. Co-Authored-By: Claude Opus 4.7 --- .claude-plugin/marketplace.json | 18 ++++---- .codex/INSTALL.md | 18 ++++---- .cursor/INSTALL.md | 20 ++++----- .opencode/INSTALL.md | 2 +- AGENTS.md | 8 ++-- CLAUDE.md | 8 ++-- CONTRIBUTING.md | 12 ++--- GEMINI.md | 4 +- README.md | 40 ++++++++--------- docs/README.codex.md | 20 ++++----- docs/README.opencode.md | 4 +- docs/skill-writing-philosophy.md | 2 +- .../SKILL.md | 22 +++++----- .../agents}/analyzer.md | 4 +- .../SKILL.md | 20 ++++----- .../agents}/test-designer.md | 2 +- .../SKILL.md | 22 +++++----- .../agents}/optimizer.md | 4 +- .../SKILL.md | 22 +++++----- .../agents}/research-functionality.md | 2 +- .../SKILL.md | 22 +++++----- .../agents}/research-submissions.md | 2 +- .../SKILL.md | 10 ++--- .../frontmatter-discipline.md | 0 .../iteration-protocol.md | 0 .../skill-design-philosophy.md | 0 .../subagent-dispatch.md | 2 +- .../workbench.md | 0 .../workflow.md | 0 skills/skill-optimizer-subagents/.gitkeep | 0 .../SKILL.md | 22 +++++----- .../agents}/test-validator.md | 2 +- .../SKILL.md | 20 ++++----- .../agents}/validator.md | 4 +- .../SKILL.md | 20 ++++----- .../agents}/test-writer.md | 4 +- tests/smoke-skill-distribution.ts | 44 +++++++++---------- 37 files changed, 203 insertions(+), 203 deletions(-) rename skills/{skill-optimizer-analyze => analyze}/SKILL.md (89%) rename skills/{skill-optimizer-subagents => analyze/agents}/analyzer.md (97%) rename skills/{skill-optimizer-design-tests => design-tests}/SKILL.md (90%) rename skills/{skill-optimizer-subagents => design-tests/agents}/test-designer.md (99%) rename skills/{skill-optimizer-improve => improve}/SKILL.md (89%) rename skills/{skill-optimizer-subagents => improve/agents}/optimizer.md (97%) rename skills/{skill-optimizer-investigate-functionality => investigate-functionality}/SKILL.md (92%) rename skills/{skill-optimizer-subagents => investigate-functionality/agents}/research-functionality.md (99%) rename skills/{skill-optimizer-investigate-submissions => investigate-submissions}/SKILL.md (87%) rename skills/{skill-optimizer-subagents => investigate-submissions/agents}/research-submissions.md (99%) rename skills/{skill-optimizer-run-bench => run-bench}/SKILL.md (94%) rename skills/{skill-optimizer-shared => shared}/frontmatter-discipline.md (100%) rename skills/{skill-optimizer-shared => shared}/iteration-protocol.md (100%) rename skills/{skill-optimizer-shared => shared}/skill-design-philosophy.md (100%) rename skills/{skill-optimizer-shared => shared}/subagent-dispatch.md (99%) rename skills/{skill-optimizer-shared => shared}/workbench.md (100%) rename skills/{skill-optimizer-shared => shared}/workflow.md (100%) delete mode 100644 skills/skill-optimizer-subagents/.gitkeep rename skills/{skill-optimizer-validate-tests => validate-tests}/SKILL.md (88%) rename skills/{skill-optimizer-subagents => validate-tests/agents}/test-validator.md (99%) rename skills/{skill-optimizer-validate => validate}/SKILL.md (90%) rename skills/{skill-optimizer-subagents => validate/agents}/validator.md (97%) rename skills/{skill-optimizer-write-tests => write-tests}/SKILL.md (91%) rename skills/{skill-optimizer-subagents => write-tests/agents}/test-writer.md (98%) diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index bd669b5..8fec79d 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -14,15 +14,15 @@ "name": "Fast" }, "skills": [ - "./skills/skill-optimizer-investigate-functionality", - "./skills/skill-optimizer-investigate-submissions", - "./skills/skill-optimizer-design-tests", - "./skills/skill-optimizer-write-tests", - "./skills/skill-optimizer-validate-tests", - "./skills/skill-optimizer-run-bench", - "./skills/skill-optimizer-analyze", - "./skills/skill-optimizer-improve", - "./skills/skill-optimizer-validate" + "./skills/investigate-functionality", + "./skills/investigate-submissions", + "./skills/design-tests", + "./skills/write-tests", + "./skills/validate-tests", + "./skills/run-bench", + "./skills/analyze", + "./skills/improve", + "./skills/validate" ] } ] diff --git a/.codex/INSTALL.md b/.codex/INSTALL.md index 855155a..5cc99a2 100644 --- a/.codex/INSTALL.md +++ b/.codex/INSTALL.md @@ -26,15 +26,15 @@ Install the 9 chain skills with the open skills CLI: ```bash npx skills add fastxyz/skill-optimizer \ - --skill skill-optimizer-investigate-functionality \ - --skill skill-optimizer-investigate-submissions \ - --skill skill-optimizer-design-tests \ - --skill skill-optimizer-write-tests \ - --skill skill-optimizer-validate-tests \ - --skill skill-optimizer-run-bench \ - --skill skill-optimizer-analyze \ - --skill skill-optimizer-improve \ - --skill skill-optimizer-validate \ + --skill investigate-functionality \ + --skill investigate-submissions \ + --skill design-tests \ + --skill write-tests \ + --skill validate-tests \ + --skill run-bench \ + --skill analyze \ + --skill improve \ + --skill validate \ -a codex -y ``` diff --git a/.cursor/INSTALL.md b/.cursor/INSTALL.md index 8016f32..be3cbc0 100644 --- a/.cursor/INSTALL.md +++ b/.cursor/INSTALL.md @@ -6,15 +6,15 @@ Install the 9 chain skills into Cursor's project or global skill directory throu ```bash npx skills add fastxyz/skill-optimizer \ - --skill skill-optimizer-investigate-functionality \ - --skill skill-optimizer-investigate-submissions \ - --skill skill-optimizer-design-tests \ - --skill skill-optimizer-write-tests \ - --skill skill-optimizer-validate-tests \ - --skill skill-optimizer-run-bench \ - --skill skill-optimizer-analyze \ - --skill skill-optimizer-improve \ - --skill skill-optimizer-validate \ + --skill investigate-functionality \ + --skill investigate-submissions \ + --skill design-tests \ + --skill write-tests \ + --skill validate-tests \ + --skill run-bench \ + --skill analyze \ + --skill improve \ + --skill validate \ -a cursor -y ``` @@ -22,4 +22,4 @@ Cursor can also import remote skills from GitHub in Settings -> Rules -> Project ## Plugin metadata -This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The skills live at `skills/skill-optimizer-*/`; the plugin manifest exposes all 9 as a chain that triggers based on the user's intent (investigate, design tests, run bench, analyze, improve, validate). +This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The skills live at `skills//`; the plugin manifest exposes all 9 as a chain that triggers based on the user's intent (investigate, design tests, run bench, analyze, improve, validate). diff --git a/.opencode/INSTALL.md b/.opencode/INSTALL.md index c11e621..61e3996 100644 --- a/.opencode/INSTALL.md +++ b/.opencode/INSTALL.md @@ -8,7 +8,7 @@ Add the plugin to `opencode.json` at user or project scope: } ``` -Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can discover the 9 chain skills (`skill-optimizer-investigate-functionality`, `skill-optimizer-investigate-submissions`, `skill-optimizer-design-tests`, `skill-optimizer-write-tests`, `skill-optimizer-validate-tests`, `skill-optimizer-run-bench`, `skill-optimizer-analyze`, `skill-optimizer-improve`, `skill-optimizer-validate`). +Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can discover the 9 chain skills (`investigate-functionality`, `investigate-submissions`, `design-tests`, `write-tests`, `validate-tests`, `run-bench`, `analyze`, `improve`, `validate`). Verify by listing skills with the skill tool; each chain skill triggers on its own description (investigate, design tests, run bench, analyze, improve, validate). diff --git a/AGENTS.md b/AGENTS.md index a412443..8c2d832 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -20,9 +20,9 @@ npx tsx src/cli.ts run-suite --help - `src/cli.ts`: public CLI entrypoint - `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces - `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases -- `skills/skill-optimizer-*/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) -- `skills/skill-optimizer-shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench reference; loaded on-demand by chain skills -- `skills/skill-optimizer-subagents/`: prompt templates dispatched by chain skills via the Agent tool +- `skills//`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) +- `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench reference; loaded on-demand by chain skills +- `skills//agents/`: prompt templates dispatched by chain skills via the Agent tool - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support - `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin - `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file @@ -49,7 +49,7 @@ Keep the README installation section aligned with packaged plugin metadata: - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state. - The agent phase sees only `/work`, not `/case` or `/results`. -- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills//`; do not create divergent skill copies. - Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`. - Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies. - Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials. diff --git a/CLAUDE.md b/CLAUDE.md index 4c2d4e9..47c2ec8 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -22,9 +22,9 @@ npx tsx src/cli.ts run-suite --help - `src/cli.ts`: public CLI entrypoint - `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces - `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases -- `skills/skill-optimizer-*/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) -- `skills/skill-optimizer-shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills -- `skills/skill-optimizer-subagents/`: prompt templates dispatched by chain skills via the Agent tool +- `skills//`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) +- `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills +- `skills//agents/`: prompt templates dispatched by chain skills via the Agent tool - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support - `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin - `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file @@ -51,7 +51,7 @@ Keep the README installation section aligned with packaged plugin metadata: - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state. - The agent phase sees only `/work`, not `/case` or `/results`. -- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills//`; do not create divergent skill copies. - Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`. - Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies. - Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 505da10..e080d76 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,6 +1,6 @@ # Contributing to skill-optimizer -Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills/skill-optimizer-*/`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths. +Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills//`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths. ## Installing The Skill @@ -24,9 +24,9 @@ All three commands must pass before opening a PR when code changes are involved. - `src/cli.ts` — public CLI entry point for `run-case` and `run-suite`. - `src/workbench/` — case/suite loading, Docker runner, Pi agent wiring, graders, traces, metrics, MCP support, and trial aggregation. - `docker/workbench-runner.Dockerfile` — non-root container image for setup, agent, grade, and cleanup phases. -- `skills/skill-optimizer-*/` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). -- `skills/skill-optimizer-shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. -- `skills/skill-optimizer-subagents/` — prompt templates dispatched by chain skills via the Agent tool. +- `skills//` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). +- `skills/shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. +- `skills//agents/` — prompt templates dispatched by chain skills via the Agent tool. - `examples/workbench/` — packaged example suites. - `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`, `.agents/plugins/marketplace.json`, `gemini-extension.json`, `GEMINI.md` — cross-agent plugin and extension metadata. - `tests/` — hand-rolled smoke tests (`tsx tests/smoke-*.ts`). @@ -47,7 +47,7 @@ All three commands must pass before opening a PR when code changes are involved. - `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override. - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - The agent phase sees only `/work`, not `/case`, `/results`, graders, hidden answers, or hidden metadata. -- Keep plugin metadata pointed at every chain skill under `skills/skill-optimizer-*/`; do not create divergent skill copies. +- Keep plugin metadata pointed at every chain skill under `skills//`; do not create divergent skill copies. ## Testing guidance @@ -59,7 +59,7 @@ All three commands must pass before opening a PR when code changes are involved. ## Adding workbench capabilities -Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/skill-optimizer-shared/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior. +Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/shared/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior. ## Commit style diff --git a/GEMINI.md b/GEMINI.md index ed6d83b..7b5b5ac 100644 --- a/GEMINI.md +++ b/GEMINI.md @@ -1,5 +1,5 @@ @./AGENTS.md @./README.md @./CONTRIBUTING.md -@./skills/skill-optimizer-shared/workflow.md -@./skills/skill-optimizer-shared/workbench.md +@./skills/shared/workflow.md +@./skills/shared/workbench.md diff --git a/README.md b/README.md index 72c09a9..85153a4 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,12 @@ Docker workbench and Agent Skills for running deterministic evals against agent Use this repo in two ways: -- Install the `skill-optimizer` plugin into your agent so it can investigate a target skill, design and run an eval suite, analyze failures, and propose improvements. The plugin bundles a 9-step chain of Agent Skills under `skills/skill-optimizer-*/`. +- Install the `skill-optimizer` plugin into your agent so it can investigate a target skill, design and run an eval suite, analyze failures, and propose improvements. The plugin bundles a 9-step chain of Agent Skills under `skills//`. - Run the local CLI to execute cases and suites in Docker against OpenRouter models. ## Installation -Installation differs by agent. Every plugin manifest exposes all 9 chain skills from `skills/skill-optimizer-*/`; once installed, each skill triggers on its own description (e.g. "investigate this skill", "design tests for this skill", "run the bench", "analyze the results"). +Installation differs by agent. Every plugin manifest exposes all 9 chain skills from `skills//`; once installed, each skill triggers on its own description (e.g. "investigate this skill", "design tests for this skill", "run the bench", "analyze the results"). ### Claude Code @@ -57,15 +57,15 @@ Install the chain skills with the open skills CLI (pass each skill name explicit ```bash npx skills add fastxyz/skill-optimizer \ - --skill skill-optimizer-investigate-functionality \ - --skill skill-optimizer-investigate-submissions \ - --skill skill-optimizer-design-tests \ - --skill skill-optimizer-write-tests \ - --skill skill-optimizer-validate-tests \ - --skill skill-optimizer-run-bench \ - --skill skill-optimizer-analyze \ - --skill skill-optimizer-improve \ - --skill skill-optimizer-validate \ + --skill investigate-functionality \ + --skill investigate-submissions \ + --skill design-tests \ + --skill write-tests \ + --skill validate-tests \ + --skill run-bench \ + --skill analyze \ + --skill improve \ + --skill validate \ -a cursor -y ``` @@ -109,15 +109,15 @@ If you only want the skill files without plugin metadata, use the open skills CL ```bash npx skills add fastxyz/skill-optimizer \ - --skill skill-optimizer-investigate-functionality \ - --skill skill-optimizer-investigate-submissions \ - --skill skill-optimizer-design-tests \ - --skill skill-optimizer-write-tests \ - --skill skill-optimizer-validate-tests \ - --skill skill-optimizer-run-bench \ - --skill skill-optimizer-analyze \ - --skill skill-optimizer-improve \ - --skill skill-optimizer-validate \ + --skill investigate-functionality \ + --skill investigate-submissions \ + --skill design-tests \ + --skill write-tests \ + --skill validate-tests \ + --skill run-bench \ + --skill analyze \ + --skill improve \ + --skill validate \ -a claude-code -a opencode -a codex -a cursor -y ``` diff --git a/docs/README.codex.md b/docs/README.codex.md index 846f3ba..e64f590 100644 --- a/docs/README.codex.md +++ b/docs/README.codex.md @@ -26,16 +26,16 @@ Install only the chain skill files with the open skills CLI (pass each skill nam ```bash npx skills add fastxyz/skill-optimizer \ - --skill skill-optimizer-investigate-functionality \ - --skill skill-optimizer-investigate-submissions \ - --skill skill-optimizer-design-tests \ - --skill skill-optimizer-write-tests \ - --skill skill-optimizer-validate-tests \ - --skill skill-optimizer-run-bench \ - --skill skill-optimizer-analyze \ - --skill skill-optimizer-improve \ - --skill skill-optimizer-validate \ + --skill investigate-functionality \ + --skill investigate-submissions \ + --skill design-tests \ + --skill write-tests \ + --skill validate-tests \ + --skill run-bench \ + --skill analyze \ + --skill improve \ + --skill validate \ -a codex -y ``` -Restart Codex if the skills do not appear immediately. The chain skills live at `skills/skill-optimizer-*/SKILL.md`. +Restart Codex if the skills do not appear immediately. The chain skills live at `skills//SKILL.md`. diff --git a/docs/README.opencode.md b/docs/README.opencode.md index 083f70b..515b285 100644 --- a/docs/README.opencode.md +++ b/docs/README.opencode.md @@ -16,7 +16,7 @@ Restart OpenCode. The plugin registers this repository's `skills/` directory so ## Verify -Use OpenCode's native `skill` tool to list skills. You should see all 9 chain skills (`skill-optimizer-investigate-functionality`, `skill-optimizer-investigate-submissions`, `skill-optimizer-design-tests`, `skill-optimizer-write-tests`, `skill-optimizer-validate-tests`, `skill-optimizer-run-bench`, `skill-optimizer-analyze`, `skill-optimizer-improve`, `skill-optimizer-validate`); each triggers on its own description. +Use OpenCode's native `skill` tool to list skills. You should see all 9 chain skills (`investigate-functionality`, `investigate-submissions`, `design-tests`, `write-tests`, `validate-tests`, `run-bench`, `analyze`, `improve`, `validate`); each triggers on its own description. ## Updating @@ -32,4 +32,4 @@ OpenCode reinstalls git plugins when it starts. To pin a tag or commit, append a The plugin exposes `.opencode/plugins/skill-optimizer.js` and adds the repository `skills/` directory to `config.skills.paths`. -The chain skills live at `skills/skill-optimizer-*/SKILL.md` — 9 user-invocable Agent Skills that orchestrate investigation, test design, bench runs, analysis, and improvement of a target skill. +The chain skills live at `skills//SKILL.md` — 9 user-invocable Agent Skills that orchestrate investigation, test design, bench runs, analysis, and improvement of a target skill. diff --git a/docs/skill-writing-philosophy.md b/docs/skill-writing-philosophy.md index bd8d71f..eb1604d 100644 --- a/docs/skill-writing-philosophy.md +++ b/docs/skill-writing-philosophy.md @@ -5,7 +5,7 @@ > wrong philosophy for the skill type is itself a form of ducktape. The lean agent-facing distillation lives at -[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skills/skill-optimizer-shared/skill-design-philosophy.md). +[`skills/shared/skill-design-philosophy.md`](../skills/shared/skill-design-philosophy.md). That doc is what the chain's analyzer/optimizer/validator subagents load each run — terse, no provenance, just the rules. **This** doc is the research backing it: per-vendor findings, where vendors diff --git a/skills/skill-optimizer-analyze/SKILL.md b/skills/analyze/SKILL.md similarity index 89% rename from skills/skill-optimizer-analyze/SKILL.md rename to skills/analyze/SKILL.md index 1a6282c..e3ee200 100644 --- a/skills/skill-optimizer-analyze/SKILL.md +++ b/skills/analyze/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-analyze -description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer-run-bench` has produced a `06-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. +name: analyze +description: Use when the user wants to diagnose why a bench run produced failures — phrases like "analyze the results", "diagnose what failed", "find structural weaknesses", "why did the skill miss X". Triggers mid-way through skill-optimizer chain work, after `skill-optimizer:run-bench` has produced a `06-bench-summary.md` with at least one failed trial. Use even when the user doesn't explicitly say "analyze" — any phrasing about understanding bench failures should trigger this. --- -# skill-optimizer-analyze +# analyze Step 7 of the skill-optimizer chain. **Fresh-derivation step.** Takes the bench summary + raw trial output from step 6, dispatches an @@ -19,7 +19,7 @@ principle that WOULD address it and the anti-patterns that would NOT. A single report at `docs/skill-optimizer//07-analysis.md`. Frontmatter (runtime-relevant facts only, per -[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): +[`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -48,7 +48,7 @@ a ducktape" signal to check against. Losing the anti-pattern list breaks the gate. Full body template and reasoning protocol in -[`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md). +[`skil./agents/analyzer.md`](./agents/analyzer.md). ## Workflow @@ -56,7 +56,7 @@ Full body template and reasoning protocol in `06-bench-summary.md` must exist with `bench_results_path` and `overall_pass_rate`. If not, tell the user to run -`skill-optimizer-run-bench` first. +`skill-optimizer:run-bench` first. If `overall_pass_rate == 1.0`: there's nothing to analyze. Surface honestly — either accept that probes don't expose a weakness, or @@ -71,7 +71,7 @@ bench run) and tell the user to re-run step 6. ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from current bench data plus directives. Collect `${OPERATOR_DIRECTIVES}` per the protocol — examples: "focus on @@ -83,16 +83,16 @@ contradiction before re-dispatching. ### (c) Dispatch the analyzer subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT cluster failures or name weaknesses yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/analyzer.md`](../skill-optimizer-subagents/analyzer.md) +[`skil./agents/analyzer.md`](./agents/analyzer.md) inline by substituting `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). The subagent sees: `06-bench-summary.md` (entry point with @@ -134,7 +134,7 @@ optimizer avoid?" Two messages depending on `has_structural_weakness`: - **`true`:** "Analysis complete. `` structural weakness(es) - identified. Next, invoke `skill-optimizer-improve`." + identified. Next, invoke `skill-optimizer:improve`." - **`false`:** "No structural weakness identified — failures consistent with noise rather than a fixable defect. Step 8 will refuse to fire. Either accept the conclusion, or re-invoke step diff --git a/skills/skill-optimizer-subagents/analyzer.md b/skills/analyze/agents/analyzer.md similarity index 97% rename from skills/skill-optimizer-subagents/analyzer.md rename to skills/analyze/agents/analyzer.md index 074a53b..dee0105 100644 --- a/skills/skill-optimizer-subagents/analyzer.md +++ b/skills/analyze/agents/analyzer.md @@ -1,6 +1,6 @@ # Analyzer subagent -You are dispatched by `skill-optimizer-analyze` to read the bench +You are dispatched by `skill-optimizer:analyze` to read the bench results, cluster failures into named **structural weaknesses** of the skill (or explicitly say there are none), and write `07-analysis.md`. This report is the chain's anti-ducktape gate: @@ -9,7 +9,7 @@ one structural weakness with the general principle that WOULD address it AND the anti-patterns that would NOT. **Read first:** -[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your "What WOULD address this" principle and "What WOULD NOT address this" diff --git a/skills/skill-optimizer-design-tests/SKILL.md b/skills/design-tests/SKILL.md similarity index 90% rename from skills/skill-optimizer-design-tests/SKILL.md rename to skills/design-tests/SKILL.md index 0e579dc..dbaa27f 100644 --- a/skills/skill-optimizer-design-tests/SKILL.md +++ b/skills/design-tests/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-design-tests +name: design-tests description: Use when the user wants to design or propose test cases for a skill — phrases like "design tests for this skill", "propose test cases", "what should we test", "plan test coverage for this skill". Also triggers mid-way through skill-optimizer chain work, once a functionality report exists (and submissions research, if PR-bound) and the next thing is figuring out what to test. Use even when the user doesn't explicitly say "design" — any phrasing about figuring out what tests to build for a skill should trigger this. --- -# skill-optimizer-design-tests +# design-tests Step 3 of the skill-optimizer chain. **Maintenance step.** Takes the functionality report from step 1, dispatches a designer subagent to @@ -29,7 +29,7 @@ Two artifacts at `docs/skill-optimizer//`: 2. **`skill-evals///spec.yaml`** — one folder per proposed functionality, with frontmatter per - [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): + [`frontmatter-discipline.md`](../shared/frontmatter-discipline.md): ```yaml name: refuses-malformed-input @@ -50,21 +50,21 @@ Two artifacts at `docs/skill-optimizer//`: For the body template of `03-test-proposals.md` and the exact `spec.yaml` field list, see -[`skills/skill-optimizer-subagents/test-designer.md`](../skill-optimizer-subagents/test-designer.md). +[`skil./agents/test-designer.md`](./agents/test-designer.md). ## Workflow ### (a) Confirm prerequisites `01-functionality.md` must exist. If not, tell the user to run -`skill-optimizer-investigate-functionality` first and stop here. +`skill-optimizer:investigate-functionality` first and stop here. If `skill-evals//` doesn't exist yet, create the empty directory. ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply the maintenance-step flow. Re-runs read the current `skill-evals//` tree as state and extend or modify it. Collect `${OPERATOR_DIRECTIVES}` per the protocol. If a directive will be @@ -78,16 +78,16 @@ re-dispatching; don't try to resolve it yourself. ### (c) Dispatch the test-designer subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT enumerate responsibilities or design proposals yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/test-designer.md`](../skill-optimizer-subagents/test-designer.md) +[`skil./agents/test-designer.md`](./agents/test-designer.md) inline by substituting `${FUNCTIONALITY_PATH}`, `${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). The subagent sees: `01-functionality.md`, the current @@ -161,7 +161,7 @@ paying for probe-building in step 4 and the bench run in step 6). ### (f) Hand off -> Next, invoke `skill-optimizer-write-tests`. +> Next, invoke `skill-optimizer:write-tests`. ## Edge cases diff --git a/skills/skill-optimizer-subagents/test-designer.md b/skills/design-tests/agents/test-designer.md similarity index 99% rename from skills/skill-optimizer-subagents/test-designer.md rename to skills/design-tests/agents/test-designer.md index 9e854ac..d7b1498 100644 --- a/skills/skill-optimizer-subagents/test-designer.md +++ b/skills/design-tests/agents/test-designer.md @@ -1,6 +1,6 @@ # Test-designer subagent -You are dispatched by `skill-optimizer-design-tests` to +You are dispatched by `skill-optimizer:design-tests` to enumerate the skill's responsibilities (from `01-functionality.md`) and propose a ranked set of **functionalities** to test, then write a per-functionality `skill-evals///spec.yaml` diff --git a/skills/skill-optimizer-improve/SKILL.md b/skills/improve/SKILL.md similarity index 89% rename from skills/skill-optimizer-improve/SKILL.md rename to skills/improve/SKILL.md index 970d682..0a0d1a4 100644 --- a/skills/skill-optimizer-improve/SKILL.md +++ b/skills/improve/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-improve -description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer-analyze` has produced `07-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. +name: improve +description: Use when the user wants to improve a skill based on identified structural weaknesses — phrases like "improve this skill", "fix the structural weakness", "optimize the skill", "apply the analysis". Triggers after `skill-optimizer:analyze` has produced `07-analysis.md` with `has_structural_weakness: true`. Refuses to fire if no weakness was identified (anti-ducktape gate). Use even when the user doesn't explicitly say "improve" — any phrasing about acting on the analysis or modifying the skill should trigger this. --- -# skill-optimizer-improve +# improve Step 8 of the skill-optimizer chain. **Fresh-derivation step.** Takes the named structural weaknesses from step 7, dispatches an optimizer @@ -34,7 +34,7 @@ One artifact at `docs/skill-optimizer//`: **`08-improvement-proposal.md`** — the optimizer's proposed change plus rationale. Frontmatter (runtime-relevant facts only, per -[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): +[`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -50,7 +50,7 @@ applied, and an explicit self-check against the "What WOULD NOT address this" anti-pattern list (the optimizer states why its proposal is NOT one of the ducktape moves the analyzer flagged). Subagent prompt template at -[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) +[`skil./agents/optimizer.md`](./agents/optimizer.md) specifies the section shape. ## Workflow @@ -60,7 +60,7 @@ specifies the section shape. Three checks: 1. `07-analysis.md` must exist with valid frontmatter. If not, - tell the user to run `skill-optimizer-analyze` first. + tell the user to run `skill-optimizer:analyze` first. 2. **Anti-ducktape gate:** `has_structural_weakness: true` must be set. If `false`, REFUSE — print: "Step 7 found no structural @@ -81,7 +81,7 @@ external check still runs if it appears later. ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from `07-analysis.md` plus current skill plus directives. Collect `${OPERATOR_DIRECTIVES}` — examples: "prefer additive changes", @@ -95,16 +95,16 @@ Translate to an atomic new requirement before passing. ### (c) Dispatch the optimizer subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT propose the diff yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/optimizer.md`](../skill-optimizer-subagents/optimizer.md) +[`skil./agents/optimizer.md`](./agents/optimizer.md) inline by substituting `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), `${PROPOSAL_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). The optimizer sees: `07-analysis.md`; `01-functionality.md`; @@ -150,7 +150,7 @@ requirement explicit — don't fill it in yourself. > Improvement proposal complete at > `08-improvement-proposal.md`. Next, invoke -> `skill-optimizer-validate` to check the proposal +> `skill-optimizer:validate` to check the proposal > independently. If the validator returns `needs-revision` or > `reject`, you'll come back here with a directive distilling the > validator's concerns. diff --git a/skills/skill-optimizer-subagents/optimizer.md b/skills/improve/agents/optimizer.md similarity index 97% rename from skills/skill-optimizer-subagents/optimizer.md rename to skills/improve/agents/optimizer.md index 8c57110..553ea7c 100644 --- a/skills/skill-optimizer-subagents/optimizer.md +++ b/skills/improve/agents/optimizer.md @@ -1,13 +1,13 @@ # Optimizer subagent -You are dispatched by `skill-optimizer-improve` to draft a +You are dispatched by `skill-optimizer:improve` to draft a principled fix for a named structural weakness from `07-analysis.md` and write `08-improvement-proposal.md`. The proposal is the SOLE output; you do NOT materialize the improved skill (that's step 9's job after the validator approves). **Read first:** -[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your proposed change should map to one or more named principles. Your self-check diff --git a/skills/skill-optimizer-investigate-functionality/SKILL.md b/skills/investigate-functionality/SKILL.md similarity index 92% rename from skills/skill-optimizer-investigate-functionality/SKILL.md rename to skills/investigate-functionality/SKILL.md index 42ffa26..1791983 100644 --- a/skills/skill-optimizer-investigate-functionality/SKILL.md +++ b/skills/investigate-functionality/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-investigate-functionality +name: investigate-functionality description: Use when the user wants to understand what an existing agent skill does — phrases like "what does this skill do", "investigate this skill", "understand this skill", or when they hand you a URL or local path to a skill they want analyzed. Also triggers at the start of any skill-optimizer chain work, before test design, analysis, or improvement. Use even when the user doesn't explicitly say "investigate" — any phrasing that signals they want to understand a skill before doing anything with it should trigger this. --- -# skill-optimizer-investigate-functionality +# investigate-functionality Step 1 of the skill-optimizer chain. **Fresh-derivation step.** Takes a skill (URL or local path), runs a researcher subagent to figure out @@ -18,7 +18,7 @@ where `` is the source skill's directory name (e.g., `firecrawl-build-scrape`). Frontmatter (runtime-relevant facts only, per -[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): +[`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -52,7 +52,7 @@ to do — re-vendor and re-research the referenced content, or proceed treating the wrapper itself as the skill to optimize. For the body template, see -[`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md). +[`skil./agents/research-functionality.md`](./agents/research-functionality.md). The source skill is vendored to `.skill-optimizer//vendored-skill/` regardless of source type (upstream fetch or local copy) so downstream steps have a @@ -164,23 +164,23 @@ Report path: `docs/skill-optimizer//01-functionality.md`. ### (f) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from the current source plus directives; git captures prior state. Collect `${OPERATOR_DIRECTIVES}` per the protocol. ### (g) Dispatch the functionality-researcher subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT do the research yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/research-functionality.md`](../skill-optimizer-subagents/research-functionality.md) +[`skil./agents/research-functionality.md`](./agents/research-functionality.md) inline by substituting `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, `${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, and the vendored path, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol — do not pass a short prompt that points at the template path). @@ -240,9 +240,9 @@ If (3): exit; the user will re-invoke with a new source. Once resolved, hand off based on this report's `pr_submission_intent` field: -- **`true`** — next, invoke `skill-optimizer-investigate-submissions` +- **`true`** — next, invoke `skill-optimizer:investigate-submissions` (step 2). After that, the chain proceeds to - `skill-optimizer-design-tests` (step 3). -- **`false`** — skip step 2 and invoke `skill-optimizer-design-tests` + `skill-optimizer:design-tests` (step 3). +- **`false`** — skip step 2 and invoke `skill-optimizer:design-tests` (step 3) directly. No late "submit a PR?" prompts; the decision was recorded above. diff --git a/skills/skill-optimizer-subagents/research-functionality.md b/skills/investigate-functionality/agents/research-functionality.md similarity index 99% rename from skills/skill-optimizer-subagents/research-functionality.md rename to skills/investigate-functionality/agents/research-functionality.md index cc9533e..15f0f9c 100644 --- a/skills/skill-optimizer-subagents/research-functionality.md +++ b/skills/investigate-functionality/agents/research-functionality.md @@ -1,6 +1,6 @@ # Functionality-researcher subagent -You are dispatched by `skill-optimizer-investigate-functionality` to +You are dispatched by `skill-optimizer:investigate-functionality` to research what a skill is supposed to do and produce `01-functionality.md` — the briefing document every later step in the chain consumes. diff --git a/skills/skill-optimizer-investigate-submissions/SKILL.md b/skills/investigate-submissions/SKILL.md similarity index 87% rename from skills/skill-optimizer-investigate-submissions/SKILL.md rename to skills/investigate-submissions/SKILL.md index af771bf..f1417c7 100644 --- a/skills/skill-optimizer-investigate-submissions/SKILL.md +++ b/skills/investigate-submissions/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-investigate-submissions +name: investigate-submissions description: Use when the user wants to research a skill's upstream PR conventions — phrases like "research PR conventions for this skill", "what does the upstream repo require for contributions", "investigate submissions for X", or when prepping a PR-bound optimization run and you need to know the upstream's rules. Triggers mid-way through skill-optimizer chain work when `01-functionality.md` has `pr_submission_intent: true`. Use even when the user doesn't explicitly say "investigate submissions" — any phrasing about figuring out the upstream's contribution rules should trigger this. --- -# skill-optimizer-investigate-submissions +# investigate-submissions Step 2 of the skill-optimizer chain — OPTIONAL, runs only when the target skill is bound for upstream PR submission. **Fresh-derivation @@ -18,7 +18,7 @@ its external consistency check. A single report at `docs/skill-optimizer//02-submissions.md`. Frontmatter (runtime-relevant facts only, per -[`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): +[`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -40,7 +40,7 @@ Body covers license details, frontmatter spec extracted from existing skills, file-location conventions, prefix taxonomy, PR-shape patterns from recent merged + closed-without-merge PRs, and rejection signals. Body template lives in -[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md). +[`skil./agents/research-submissions.md`](./agents/research-submissions.md). ## Workflow @@ -50,14 +50,14 @@ Two prerequisites: 1. `01-functionality.md` must exist with valid frontmatter. If not, tell the user to run - `skill-optimizer-investigate-functionality` first. + `skill-optimizer:investigate-functionality` first. 2. That report's `pr_submission_intent` field must be `true`. If `false`, this skill should not run — tell the user step 2 is skipped for local-only optimization runs. ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical from current upstream facts plus directives. Staleness against `01-functionality.md` is determined by git-mtime comparison. @@ -69,16 +69,16 @@ known. ### (c) Dispatch the submission-researcher subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT scrape the upstream repo yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/research-submissions.md`](../skill-optimizer-subagents/research-submissions.md) +[`skil./agents/research-submissions.md`](./agents/research-submissions.md) inline by substituting `${UPSTREAM_REPO}` from `01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). The subagent sees: the upstream repo via `gh` CLI (PR list — both @@ -128,8 +128,8 @@ be merged. Report the file path, a one-line summary (license / CLA / branch target / any flagged blockers), then: -> Next, invoke `skill-optimizer-design-tests`. The validator in -> `skill-optimizer-validate` will later read this report for its +> Next, invoke `skill-optimizer:design-tests`. The validator in +> `skill-optimizer:validate` will later read this report for its > external consistency check. ## Edge cases diff --git a/skills/skill-optimizer-subagents/research-submissions.md b/skills/investigate-submissions/agents/research-submissions.md similarity index 99% rename from skills/skill-optimizer-subagents/research-submissions.md rename to skills/investigate-submissions/agents/research-submissions.md index c92ec73..b3c0bfd 100644 --- a/skills/skill-optimizer-subagents/research-submissions.md +++ b/skills/investigate-submissions/agents/research-submissions.md @@ -1,6 +1,6 @@ # Submission-researcher subagent -You are dispatched by `skill-optimizer-investigate-submissions` to +You are dispatched by `skill-optimizer:investigate-submissions` to research the upstream repo's PR conventions and produce `02-submissions.md` — the verbatim-pastable context block the validator (step 9) uses for its external consistency check. diff --git a/skills/skill-optimizer-run-bench/SKILL.md b/skills/run-bench/SKILL.md similarity index 94% rename from skills/skill-optimizer-run-bench/SKILL.md rename to skills/run-bench/SKILL.md index 93559cc..8bf6819 100644 --- a/skills/skill-optimizer-run-bench/SKILL.md +++ b/skills/run-bench/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-run-bench +name: run-bench description: Use when the user wants to run the eval suite against a skill and capture results — phrases like "run the bench", "measure the skill", "benchmark this", "run the eval suite", "execute the workbench". Also triggers mid-way through skill-optimizer chain work, once `skill-evals//suite.yml` exists from step 4 and the next thing is to measure. Use even when the user doesn't explicitly say "bench" — any phrasing about running the test suite for the skill should trigger this. --- -# skill-optimizer-run-bench +# run-bench Step 6 of the skill-optimizer chain. **Fresh-derivation step (for the summary).** Invokes the skill-optimizer CLI's `run-suite` @@ -24,7 +24,7 @@ Two outputs at `docs/skill-optimizer//`: naturally-accumulated archive. 2. **`06-bench-summary.md`** — single canonical aggregate, per - [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md): + [`frontmatter-discipline.md`](../shared/frontmatter-discipline.md): ```yaml --- @@ -49,7 +49,7 @@ vendored the source regardless of upstream/local). ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Two pieces: - **Raw bench output** at `.skill-optimizer//bench-results//`: always @@ -114,7 +114,7 @@ Don't auto-invoke step 7. Read `overall_pass_rate`. Two messages: -- `< 1.0`: invoke `skill-optimizer-analyze` to diagnose +- `< 1.0`: invoke `skill-optimizer:analyze` to diagnose failures. - `== 1.0`: surface the choice — accept that probes don't expose a weakness, or re-run step 3 with a "make probes harder" diff --git a/skills/skill-optimizer-shared/frontmatter-discipline.md b/skills/shared/frontmatter-discipline.md similarity index 100% rename from skills/skill-optimizer-shared/frontmatter-discipline.md rename to skills/shared/frontmatter-discipline.md diff --git a/skills/skill-optimizer-shared/iteration-protocol.md b/skills/shared/iteration-protocol.md similarity index 100% rename from skills/skill-optimizer-shared/iteration-protocol.md rename to skills/shared/iteration-protocol.md diff --git a/skills/skill-optimizer-shared/skill-design-philosophy.md b/skills/shared/skill-design-philosophy.md similarity index 100% rename from skills/skill-optimizer-shared/skill-design-philosophy.md rename to skills/shared/skill-design-philosophy.md diff --git a/skills/skill-optimizer-shared/subagent-dispatch.md b/skills/shared/subagent-dispatch.md similarity index 99% rename from skills/skill-optimizer-shared/subagent-dispatch.md rename to skills/shared/subagent-dispatch.md index a81eb5b..e915172 100644 --- a/skills/skill-optimizer-shared/subagent-dispatch.md +++ b/skills/shared/subagent-dispatch.md @@ -104,7 +104,7 @@ that needs no file I/O to begin work. Concretely, for every step's dispatch: 1. **Read the subagent's prompt template** at - `skills/skill-optimizer-subagents/.md`. The chain skill's + `skills//agents/.md`. The chain skill's workflow tells you which template. 2. **Substitute every `${VAR}` placeholder** with the concrete value the chain skill's workflow specifies. `${OPERATOR_DIRECTIVES}` diff --git a/skills/skill-optimizer-shared/workbench.md b/skills/shared/workbench.md similarity index 100% rename from skills/skill-optimizer-shared/workbench.md rename to skills/shared/workbench.md diff --git a/skills/skill-optimizer-shared/workflow.md b/skills/shared/workflow.md similarity index 100% rename from skills/skill-optimizer-shared/workflow.md rename to skills/shared/workflow.md diff --git a/skills/skill-optimizer-subagents/.gitkeep b/skills/skill-optimizer-subagents/.gitkeep deleted file mode 100644 index e69de29..0000000 diff --git a/skills/skill-optimizer-validate-tests/SKILL.md b/skills/validate-tests/SKILL.md similarity index 88% rename from skills/skill-optimizer-validate-tests/SKILL.md rename to skills/validate-tests/SKILL.md index 700b2ba..eecc9ad 100644 --- a/skills/skill-optimizer-validate-tests/SKILL.md +++ b/skills/validate-tests/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-validate-tests -description: Use when the user wants to validate the probes built by `skill-optimizer-write-tests` before running the bench — phrases like "validate the tests", "check the graders", "are these probes fair", "review the test suite". Triggers after `skill-optimizer-write-tests` has populated `skill-evals////` folders and before `skill-optimizer-run-bench`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking that the probes actually probe what they claim should trigger this. +name: validate-tests +description: Use when the user wants to validate the probes built by `skill-optimizer:write-tests` before running the bench — phrases like "validate the tests", "check the graders", "are these probes fair", "review the test suite". Triggers after `skill-optimizer:write-tests` has populated `skill-evals////` folders and before `skill-optimizer:run-bench`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking that the probes actually probe what they claim should trigger this. --- -# skill-optimizer-validate-tests +# validate-tests Step 5 of the skill-optimizer chain. **Fresh-derivation step.** Takes the probes built by step 4, dispatches a test-validator @@ -37,7 +37,7 @@ One artifact at `docs/skill-optimizer//`: **`05-tests-verdict.md`** — aggregate verdict across all probes plus per-probe sections. Frontmatter (runtime-relevant facts only, -per [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): +per [`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -61,7 +61,7 @@ outputs), and — for needs-revision/reject — a concrete suggestion the operator can distill into a directive for step 4. Body template and the validator's reasoning protocol live at -[`skills/skill-optimizer-subagents/test-validator.md`](../skill-optimizer-subagents/test-validator.md). +[`skil./agents/test-validator.md`](./agents/test-validator.md). ## Workflow @@ -78,7 +78,7 @@ context on what the skill is supposed to do). ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical verdict from current probe state plus directives. Collect `${OPERATOR_DIRECTIVES}` per the protocol — examples: "be stricter @@ -87,18 +87,18 @@ requires a specific JSON shape the skill never asks for". ### (c) Dispatch test-validator subagents (parallel, one per probe) -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT judge probes yourself in this session** — dispatch test-validator subagents via the `Agent` tool, in parallel (emit all dispatches in a single message). For each probe, render the prompt template at -[`skills/skill-optimizer-subagents/test-validator.md`](../skill-optimizer-subagents/test-validator.md) +[`skil./agents/test-validator.md`](./agents/test-validator.md) inline by substituting that probe's `${PROBE_NAME}`, `${PROBE_DIR}`, `${FUNCTIONALITY_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, `${VERDICT_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as that Agent dispatch's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). Each test-validator sees: its single probe's full contents @@ -148,11 +148,11 @@ reconcile yourself. Two messages depending on `all_probes_approved`: - **`true`:** "All probes approved. Next, invoke - `skill-optimizer-run-bench` to measure baseline performance." + `skill-optimizer:run-bench` to measure baseline performance." - **`false`:** "`` probes need revision, `` rejected. The validator's per-probe rationale is in `05-tests-verdict.md`. Distill the concerns into directives and re-invoke - `skill-optimizer-write-tests` for the affected probes (per the + `skill-optimizer:write-tests` for the affected probes (per the per-probe rebuild flow at step 4), then re-invoke this step. Don't proceed to bench against bad probes." diff --git a/skills/skill-optimizer-subagents/test-validator.md b/skills/validate-tests/agents/test-validator.md similarity index 99% rename from skills/skill-optimizer-subagents/test-validator.md rename to skills/validate-tests/agents/test-validator.md index 2816630..2684934 100644 --- a/skills/skill-optimizer-subagents/test-validator.md +++ b/skills/validate-tests/agents/test-validator.md @@ -1,6 +1,6 @@ # Test-validator subagent -You are dispatched by `skill-optimizer-validate-tests` to +You are dispatched by `skill-optimizer:validate-tests` to **independently judge** whether ONE probe (built by step 4) fairly tests the responsibility its parent functionality claims, and whether its grader is sound. Many test-validator subagents run in diff --git a/skills/skill-optimizer-validate/SKILL.md b/skills/validate/SKILL.md similarity index 90% rename from skills/skill-optimizer-validate/SKILL.md rename to skills/validate/SKILL.md index 4ac9db4..a4cb8ee 100644 --- a/skills/skill-optimizer-validate/SKILL.md +++ b/skills/validate/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-validate -description: Use when the user wants to validate an improvement proposal from `skill-optimizer-improve` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer-improve` has produced `08-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. +name: validate +description: Use when the user wants to validate an improvement proposal from `skill-optimizer:improve` — phrases like "validate the improvement", "check the proposal", "is the optimizer's change sound", "review the fix". Triggers after `skill-optimizer:improve` has produced `08-improvement-proposal.md`. Use even when the user doesn't explicitly say "validate" — any phrasing about checking whether the proposed change is sound and conformant should trigger this. --- -# skill-optimizer-validate +# validate Step 9 of the skill-optimizer chain. **Fresh-derivation step.** Takes the proposal from step 8, dispatches a validator subagent to @@ -24,7 +24,7 @@ One or two artifacts at `docs/skill-optimizer//`: 1. **`09-validator-verdict.md`** — the validator's independent judgment. Frontmatter (runtime-relevant facts only, per - [`frontmatter-discipline.md`](../skill-optimizer-shared/frontmatter-discipline.md)): + [`frontmatter-discipline.md`](../shared/frontmatter-discipline.md)): ```yaml --- @@ -42,7 +42,7 @@ One or two artifacts at `docs/skill-optimizer//`: rules?). External check is forward-looking — verifies the improved skill COULD be turned into a valid PR, even though step 9 doesn't produce one. Body template at - [`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md). + [`skil./agents/validator.md`](./agents/validator.md). 2. **`docs/skill-optimizer//improved-skill/`** — improved skill content, only materialized when `verdict: approve`. Mirrors the source's @@ -59,7 +59,7 @@ One or two artifacts at `docs/skill-optimizer//`: Three checks: 1. `08-improvement-proposal.md` must exist with valid frontmatter. - If not, tell the user to run `skill-optimizer-improve` + If not, tell the user to run `skill-optimizer:improve` first. 2. Current skill state must be readable: `docs/skill-optimizer//improved-skill/` if it exists (prior accumulated state), else `.skill-optimizer//vendored-skill/`. The @@ -77,7 +77,7 @@ verdict body will note the omission). ### (b) Handle iteration -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply it. Each invocation overwrites the canonical verdict. Collect `${OPERATOR_DIRECTIVES}` — examples: "be stricter on additive-vs-destructive", "the external check missed the @@ -89,11 +89,11 @@ an atomic new requirement before passing. ### (c) Dispatch the validator subagent -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT judge the proposal yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skills/skill-optimizer-subagents/validator.md`](../skill-optimizer-subagents/validator.md) +[`skil./agents/validator.md`](./agents/validator.md) inline by substituting `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (current state — `docs/skill-optimizer//improved-skill/` if it exists, else source), `${SKILL_AFTER_PATH}` (proposed result — apply the diff to a temp @@ -101,7 +101,7 @@ copy; do NOT touch the canonical `docs/skill-optimizer//improved-skill/` y `${FUNCTIONALITY_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), `${VERDICT_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). The validator sees: skill BEFORE; skill AFTER (temporary diff --git a/skills/skill-optimizer-subagents/validator.md b/skills/validate/agents/validator.md similarity index 97% rename from skills/skill-optimizer-subagents/validator.md rename to skills/validate/agents/validator.md index 162adde..8620283 100644 --- a/skills/skill-optimizer-subagents/validator.md +++ b/skills/validate/agents/validator.md @@ -1,13 +1,13 @@ # Validator subagent -You are dispatched by `skill-optimizer-validate` to **independently +You are dispatched by `skill-optimizer:validate` to **independently judge** an improvement proposal (from step 8) and write `09-validator-verdict.md`. On `verdict: approve`, the operator session materializes `docs/skill-optimizer//improved-skill/` from the proposal. On `needs-revision`/`reject`, the proposal is sent back to step 8. **Read first:** -[`skills/skill-optimizer-shared/skill-design-philosophy.md`](../skill-optimizer-shared/skill-design-philosophy.md) +[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your verdict reasoning should cite which principles the change embodies and diff --git a/skills/skill-optimizer-write-tests/SKILL.md b/skills/write-tests/SKILL.md similarity index 91% rename from skills/skill-optimizer-write-tests/SKILL.md rename to skills/write-tests/SKILL.md index fbbaad5..5b60343 100644 --- a/skills/skill-optimizer-write-tests/SKILL.md +++ b/skills/write-tests/SKILL.md @@ -1,9 +1,9 @@ --- -name: skill-optimizer-write-tests -description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer-design-tests` has produced `skill-evals///spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. +name: write-tests +description: Use when the user wants to build or implement the eval probes for a skill — phrases like "build the tests", "implement the probes", "set up the eval cases", "write test fixtures for X". Triggers after `skill-optimizer:design-tests` has produced `skill-evals///spec.yaml` files with at least one marked `picked: true`, and before any bench run. Use even when the user doesn't explicitly say "write tests" — any phrasing about turning the picked functionalities into concrete probes should trigger this. --- -# skill-optimizer-write-tests +# write-tests Step 4 of the skill-optimizer chain. **Maintenance step.** Reads the picked functionalities from step 3 @@ -42,9 +42,9 @@ Two kinds of output under `docs/skill-optimizer//`: Probe-level `spec.yaml` format and the workbench schema for `suite.yml` are defined in -[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) +[`skil./agents/test-writer.md`](./agents/test-writer.md) and -[`skills/skill-optimizer-shared/workbench.md`](../skill-optimizer-shared/workbench.md) +[`skills/shared/workbench.md`](../shared/workbench.md) respectively. ## Workflow @@ -64,7 +64,7 @@ guess. ### (b) Handle iteration (maintenance step) -Read [`iteration-protocol.md`](../skill-optimizer-shared/iteration-protocol.md) +Read [`iteration-protocol.md`](../shared/iteration-protocol.md) and apply the maintenance-step flow: - Walk `skill-evals//`. For each picked functionality, determine which @@ -113,18 +113,18 @@ Three realistic responses: ### (d) Dispatch test-writer subagents (parallel, one per probe) -Read [`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md) +Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT write workspace files or graders yourself in this session** — dispatch test-writer subagents via the `Agent` tool, in parallel (emit all dispatches in a single message). For each probe, render the prompt template at -[`skills/skill-optimizer-subagents/test-writer.md`](../skill-optimizer-subagents/test-writer.md) +[`skil./agents/test-writer.md`](./agents/test-writer.md) inline by substituting that probe's `${PROBE_NAME}`, `${FUNCTIONALITY_SPEC_PATH}`, `${PROBE_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PROBE_DIR}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as that Agent dispatch's `prompt` parameter (per -[`subagent-dispatch.md`](../skill-optimizer-shared/subagent-dispatch.md)'s +[`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). Each test-writer sees: its single probe's intent (slug + parent @@ -185,7 +185,7 @@ Walk `skill-evals//`. For each `/` where `spec.yaml` has `picked: true`, list every probe folder and emit a suite.yml entry per the workbench schema. Overwrite `skill-evals//suite.yml`. -Commit `skill-evals//` and hand off to `skill-optimizer-run-bench`. +Commit `skill-evals//` and hand off to `skill-optimizer:run-bench`. ## Edge cases diff --git a/skills/skill-optimizer-subagents/test-writer.md b/skills/write-tests/agents/test-writer.md similarity index 98% rename from skills/skill-optimizer-subagents/test-writer.md rename to skills/write-tests/agents/test-writer.md index 9e94cad..054cd97 100644 --- a/skills/skill-optimizer-subagents/test-writer.md +++ b/skills/write-tests/agents/test-writer.md @@ -1,6 +1,6 @@ # Test-writer subagent -You are dispatched by `skill-optimizer-write-tests` to build ONE +You are dispatched by `skill-optimizer:write-tests` to build ONE probe for ONE functionality — the concrete workspace files, the grader script, and the smoke-check fixtures. Many test-writer subagents run in parallel (one per probe); each is responsible for @@ -76,7 +76,7 @@ will not be found. The workbench schema (suite.yml, grader contract, smoke format) lives in -[`../skill-optimizer-shared/workbench.md`](../skill-optimizer-shared/workbench.md). +[`../../shared/workbench.md`](../../shared/workbench.md). ### 1. `${PROBE_SPEC_PATH}` — the probe's own spec.yaml diff --git a/tests/smoke-skill-distribution.ts b/tests/smoke-skill-distribution.ts index 5172d43..b3e97f0 100644 --- a/tests/smoke-skill-distribution.ts +++ b/tests/smoke-skill-distribution.ts @@ -17,15 +17,15 @@ function readText(relativePath: string): string { test('chain skills follow the portable agent skills contract', () => { const chainSkillDirs = [ - 'skill-optimizer-investigate-functionality', - 'skill-optimizer-investigate-submissions', - 'skill-optimizer-design-tests', - 'skill-optimizer-write-tests', - 'skill-optimizer-validate-tests', - 'skill-optimizer-run-bench', - 'skill-optimizer-analyze', - 'skill-optimizer-improve', - 'skill-optimizer-validate', + 'investigate-functionality', + 'investigate-submissions', + 'design-tests', + 'write-tests', + 'validate-tests', + 'run-bench', + 'analyze', + 'improve', + 'validate', ]; for (const dir of chainSkillDirs) { @@ -41,7 +41,7 @@ test('chain skills follow the portable agent skills contract', () => { }); test('shared workbench reference documents current CLI patterns', () => { - const reference = readText('skills/skill-optimizer-shared/workbench.md'); + const reference = readText('skills/shared/workbench.md'); assert.doesNotMatch(reference, /verify-suite/); assert.doesNotMatch(reference, /runWorkbenchReferenceSolutions/); @@ -54,7 +54,7 @@ test('shared workbench reference documents current CLI patterns', () => { }); test('shared workbench reference documents bin directory visibility accurately', () => { - const reference = readText('skills/skill-optimizer-shared/workbench.md'); + const reference = readText('skills/shared/workbench.md'); assert.match( reference, @@ -126,15 +126,15 @@ test('Claude plugin and marketplace metadata expose all v1.4 chain skills', () = assert.equal(marketplace.plugins[0].description, pluginDescription); assert.equal(marketplace.plugins[0].source, './'); assert.deepEqual(marketplace.plugins[0].skills, [ - './skills/skill-optimizer-investigate-functionality', - './skills/skill-optimizer-investigate-submissions', - './skills/skill-optimizer-design-tests', - './skills/skill-optimizer-write-tests', - './skills/skill-optimizer-validate-tests', - './skills/skill-optimizer-run-bench', - './skills/skill-optimizer-analyze', - './skills/skill-optimizer-improve', - './skills/skill-optimizer-validate', + './skills/investigate-functionality', + './skills/investigate-submissions', + './skills/design-tests', + './skills/write-tests', + './skills/validate-tests', + './skills/run-bench', + './skills/analyze', + './skills/improve', + './skills/validate', ]); for (const skillPath of marketplace.plugins[0].skills) { const resolved = skillPath.replace(/^\.\//, ''); @@ -203,6 +203,6 @@ test('Gemini extension metadata imports the v1.4 chain overview + workbench refe assert.match(geminiInstructions, /^@\.\/AGENTS\.md$/m); assert.match(geminiInstructions, /^@\.\/README\.md$/m); assert.match(geminiInstructions, /^@\.\/CONTRIBUTING\.md$/m); - assert.match(geminiInstructions, /^@\.\/skills\/skill-optimizer-shared\/workflow\.md$/m); - assert.match(geminiInstructions, /^@\.\/skills\/skill-optimizer-shared\/workbench\.md$/m); + assert.match(geminiInstructions, /^@\.\/skills\/shared\/workflow\.md$/m); + assert.match(geminiInstructions, /^@\.\/skills\/shared\/workbench\.md$/m); }); From 336c27c2149b3a6ab3077457db656d2e34835727 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 08:16:32 -0500 Subject: [PATCH 063/121] fix(links): repair display text mangled by unescaped sed pattern MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The big rename commit's sed had \`s|../skill-optimizer-subagents/|./agents/|g\` where \`..\` matched any 2 chars (sed BRE treats unescaped \`.\` as "any character"). On display text \`skills/skill-optimizer-subagents/X.md\`, the pattern matched "ls/skill-optimizer-subagents/" producing \`skil./agents/X.md\` instead of leaving the display text alone. Same bug hit \`s|../skill-optimizer-shared/|../../shared/|g\` on agents/*.md files: \`skills/skill-optimizer-shared/X.md\` became \`skil../../shared/X.md\`. Fix: - All 16 broken display texts in chain SKILL.md files (\`[\`skil./agents/X.md\`]...\` → \`[\`agents/X.md\`]...\`) - 2 broken display texts in analyzer.md and validator.md (\`[\`skil../../shared/X.md\`]...\` → \`[\`../../shared/X.md\`]...\`) URLs were already correct; only display text was mangled. Verified: tests + typecheck pass; full skills/ audit clean for legacy patterns, bare refs, and step-letter ordering. Co-Authored-By: Claude Opus 4.7 --- skills/analyze/SKILL.md | 4 ++-- skills/analyze/agents/analyzer.md | 2 +- skills/design-tests/SKILL.md | 4 ++-- skills/improve/SKILL.md | 4 ++-- skills/improve/agents/optimizer.md | 4 ++-- skills/investigate-functionality/SKILL.md | 4 ++-- skills/investigate-submissions/SKILL.md | 4 ++-- skills/validate-tests/SKILL.md | 4 ++-- skills/validate/SKILL.md | 4 ++-- skills/validate/agents/validator.md | 2 +- skills/write-tests/SKILL.md | 4 ++-- 11 files changed, 20 insertions(+), 20 deletions(-) diff --git a/skills/analyze/SKILL.md b/skills/analyze/SKILL.md index e3ee200..991f342 100644 --- a/skills/analyze/SKILL.md +++ b/skills/analyze/SKILL.md @@ -48,7 +48,7 @@ a ducktape" signal to check against. Losing the anti-pattern list breaks the gate. Full body template and reasoning protocol in -[`skil./agents/analyzer.md`](./agents/analyzer.md). +[`agents/analyzer.md`](./agents/analyzer.md). ## Workflow @@ -87,7 +87,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT cluster failures or name weaknesses yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/analyzer.md`](./agents/analyzer.md) +[`agents/analyzer.md`](./agents/analyzer.md) inline by substituting `${BENCH_RESULTS_PATH}`, `${SUMMARY_PATH}`, `${TESTS_TREE_PATH}`, `${SKILL_SOURCE_PATH}`, `${OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the diff --git a/skills/analyze/agents/analyzer.md b/skills/analyze/agents/analyzer.md index dee0105..715ef86 100644 --- a/skills/analyze/agents/analyzer.md +++ b/skills/analyze/agents/analyzer.md @@ -9,7 +9,7 @@ one structural weakness with the general principle that WOULD address it AND the anti-patterns that would NOT. **Read first:** -[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) +[`../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your "What WOULD address this" principle and "What WOULD NOT address this" diff --git a/skills/design-tests/SKILL.md b/skills/design-tests/SKILL.md index dbaa27f..6a1eeaa 100644 --- a/skills/design-tests/SKILL.md +++ b/skills/design-tests/SKILL.md @@ -50,7 +50,7 @@ Two artifacts at `docs/skill-optimizer//`: For the body template of `03-test-proposals.md` and the exact `spec.yaml` field list, see -[`skil./agents/test-designer.md`](./agents/test-designer.md). +[`agents/test-designer.md`](./agents/test-designer.md). ## Workflow @@ -82,7 +82,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT enumerate responsibilities or design proposals yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/test-designer.md`](./agents/test-designer.md) +[`agents/test-designer.md`](./agents/test-designer.md) inline by substituting `${FUNCTIONALITY_PATH}`, `${TESTS_TREE_PATH}`, `${PROPOSALS_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent diff --git a/skills/improve/SKILL.md b/skills/improve/SKILL.md index 0a0d1a4..731214c 100644 --- a/skills/improve/SKILL.md +++ b/skills/improve/SKILL.md @@ -50,7 +50,7 @@ applied, and an explicit self-check against the "What WOULD NOT address this" anti-pattern list (the optimizer states why its proposal is NOT one of the ducktape moves the analyzer flagged). Subagent prompt template at -[`skil./agents/optimizer.md`](./agents/optimizer.md) +[`agents/optimizer.md`](./agents/optimizer.md) specifies the section shape. ## Workflow @@ -99,7 +99,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT propose the diff yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/optimizer.md`](./agents/optimizer.md) +[`agents/optimizer.md`](./agents/optimizer.md) inline by substituting `${ANALYSIS_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_CURRENT_PATH}`, `${SUBMISSIONS_PATH}` (if PR-bound), `${PROPOSAL_OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass diff --git a/skills/improve/agents/optimizer.md b/skills/improve/agents/optimizer.md index 553ea7c..09b3d53 100644 --- a/skills/improve/agents/optimizer.md +++ b/skills/improve/agents/optimizer.md @@ -7,7 +7,7 @@ proposal is the SOLE output; you do NOT materialize the improved skill (that's step 9's job after the validator approves). **Read first:** -[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) +[`../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your proposed change should map to one or more named principles. Your self-check @@ -170,7 +170,7 @@ Examples that don't (context dumps — reject): addressing the weakness would require contradicting the skill's stated responsibilities, surface — there's a deeper framing problem. -- **All anti-patterns are pre-empted by the skill's current +- **All anti-patterns are preempted by the skill's current shape.** If the only diffs that would address the weakness are on the anti-pattern list, the gate is doing its job — surface honestly and exit. diff --git a/skills/investigate-functionality/SKILL.md b/skills/investigate-functionality/SKILL.md index 1791983..afcd2bf 100644 --- a/skills/investigate-functionality/SKILL.md +++ b/skills/investigate-functionality/SKILL.md @@ -52,7 +52,7 @@ to do — re-vendor and re-research the referenced content, or proceed treating the wrapper itself as the skill to optimize. For the body template, see -[`skil./agents/research-functionality.md`](./agents/research-functionality.md). +[`agents/research-functionality.md`](./agents/research-functionality.md). The source skill is vendored to `.skill-optimizer//vendored-skill/` regardless of source type (upstream fetch or local copy) so downstream steps have a @@ -175,7 +175,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT do the research yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/research-functionality.md`](./agents/research-functionality.md) +[`agents/research-functionality.md`](./agents/research-functionality.md) inline by substituting `${SKILL_SOURCE}`, `${OUTPUT_PATH}`, `${PR_SUBMISSION_INTENT}`, `${OPERATOR_DIRECTIVES}`, and the vendored path, then pass the rendered text as the Agent tool's diff --git a/skills/investigate-submissions/SKILL.md b/skills/investigate-submissions/SKILL.md index f1417c7..29167dc 100644 --- a/skills/investigate-submissions/SKILL.md +++ b/skills/investigate-submissions/SKILL.md @@ -40,7 +40,7 @@ Body covers license details, frontmatter spec extracted from existing skills, file-location conventions, prefix taxonomy, PR-shape patterns from recent merged + closed-without-merge PRs, and rejection signals. Body template lives in -[`skil./agents/research-submissions.md`](./agents/research-submissions.md). +[`agents/research-submissions.md`](./agents/research-submissions.md). ## Workflow @@ -73,7 +73,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT scrape the upstream repo yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/research-submissions.md`](./agents/research-submissions.md) +[`agents/research-submissions.md`](./agents/research-submissions.md) inline by substituting `${UPSTREAM_REPO}` from `01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the diff --git a/skills/validate-tests/SKILL.md b/skills/validate-tests/SKILL.md index eecc9ad..65dc558 100644 --- a/skills/validate-tests/SKILL.md +++ b/skills/validate-tests/SKILL.md @@ -61,7 +61,7 @@ outputs), and — for needs-revision/reject — a concrete suggestion the operator can distill into a directive for step 4. Body template and the validator's reasoning protocol live at -[`skil./agents/test-validator.md`](./agents/test-validator.md). +[`agents/test-validator.md`](./agents/test-validator.md). ## Workflow @@ -92,7 +92,7 @@ for the constraints. **Do NOT judge probes yourself in this session** — dispatch test-validator subagents via the `Agent` tool, in parallel (emit all dispatches in a single message). For each probe, render the prompt template at -[`skil./agents/test-validator.md`](./agents/test-validator.md) +[`agents/test-validator.md`](./agents/test-validator.md) inline by substituting that probe's `${PROBE_NAME}`, `${PROBE_DIR}`, `${FUNCTIONALITY_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, `${VERDICT_OUTPUT_PATH}`, and diff --git a/skills/validate/SKILL.md b/skills/validate/SKILL.md index a4cb8ee..af1bb54 100644 --- a/skills/validate/SKILL.md +++ b/skills/validate/SKILL.md @@ -42,7 +42,7 @@ One or two artifacts at `docs/skill-optimizer//`: rules?). External check is forward-looking — verifies the improved skill COULD be turned into a valid PR, even though step 9 doesn't produce one. Body template at - [`skil./agents/validator.md`](./agents/validator.md). + [`agents/validator.md`](./agents/validator.md). 2. **`docs/skill-optimizer//improved-skill/`** — improved skill content, only materialized when `verdict: approve`. Mirrors the source's @@ -93,7 +93,7 @@ Read [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for the constraints. **Do NOT judge the proposal yourself in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at -[`skil./agents/validator.md`](./agents/validator.md) +[`agents/validator.md`](./agents/validator.md) inline by substituting `${PROPOSAL_PATH}`, `${SKILL_BEFORE_PATH}` (current state — `docs/skill-optimizer//improved-skill/` if it exists, else source), `${SKILL_AFTER_PATH}` (proposed result — apply the diff to a temp diff --git a/skills/validate/agents/validator.md b/skills/validate/agents/validator.md index 8620283..289d135 100644 --- a/skills/validate/agents/validator.md +++ b/skills/validate/agents/validator.md @@ -7,7 +7,7 @@ session materializes `docs/skill-optimizer//improved-skill/` from the prop `needs-revision`/`reject`, the proposal is sent back to step 8. **Read first:** -[`skil../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) +[`../../shared/skill-design-philosophy.md`](../../shared/skill-design-philosophy.md) — the cross-vendor synthesis of what makes a skill good and what makes an improvement principled vs ducktape. Your verdict reasoning should cite which principles the change embodies and diff --git a/skills/write-tests/SKILL.md b/skills/write-tests/SKILL.md index 5b60343..600b57a 100644 --- a/skills/write-tests/SKILL.md +++ b/skills/write-tests/SKILL.md @@ -42,7 +42,7 @@ Two kinds of output under `docs/skill-optimizer//`: Probe-level `spec.yaml` format and the workbench schema for `suite.yml` are defined in -[`skil./agents/test-writer.md`](./agents/test-writer.md) +[`agents/test-writer.md`](./agents/test-writer.md) and [`skills/shared/workbench.md`](../shared/workbench.md) respectively. @@ -118,7 +118,7 @@ for the constraints. **Do NOT write workspace files or graders yourself in this session** — dispatch test-writer subagents via the `Agent` tool, in parallel (emit all dispatches in a single message). For each probe, render the prompt template at -[`skil./agents/test-writer.md`](./agents/test-writer.md) +[`agents/test-writer.md`](./agents/test-writer.md) inline by substituting that probe's `${PROBE_NAME}`, `${FUNCTIONALITY_SPEC_PATH}`, `${PROBE_SPEC_PATH}`, `${FUNCTIONALITY_PATH}`, `${SKILL_SOURCE_PATH}`, From 64e3970d4dd116272a6b2d684bfa57cc470e92bd Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 11:59:40 -0500 Subject: [PATCH 064/121] docs(spec): multi-agent ACP workbench design Brainstormed design for replacing the embedded pi-agent runtime with host-side ACP client driving 5 agent CLIs (Claude Code, Codex, Gemini, OpenCode, pi-acp) in per-trial sandboxed containers. Captures architecture, registry pattern, schema changes, skill deployment, auth shuttling, raw ACP trace capture, per-agent MCP config writing, migration path, and testing strategy. Co-Authored-By: Claude Opus 4.7 --- .../2026-05-25-multi-agent-acp-design.md | 490 ++++++++++++++++++ 1 file changed, 490 insertions(+) create mode 100644 docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md diff --git a/docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md b/docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md new file mode 100644 index 0000000..dd54eb2 --- /dev/null +++ b/docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md @@ -0,0 +1,490 @@ +# Multi-agent workbench via Agent Client Protocol + +**Date:** 2026-05-25 +**Status:** design (brainstormed with `superpowers:brainstorming`) +**Supersedes:** the embedded-pi-agent runtime in +`src/workbench/pi-agent.ts` and the `--agent` mode of +`src/workbench/container-runner.ts`. + +## Goal + +Run skill-optimizer's bench against five real agent CLIs instead of one +embedded pi-agent library: Claude Code, OpenAI Codex, Google Gemini, +OpenCode, and pi (rehosted via pi-acp). Adopt the Agent Client Protocol +(ACP) as the uniform runtime interface, with the host driving each agent +inside a per-trial sandboxed container. + +## Why + +skill-optimizer today calls `@mariozechner/pi-coding-agent` as an +in-process library inside the Docker container. That gives one runtime +flavor — bench results reflect how a skill behaves under pi-agent as a +proxy, not under the agents real users actually load it in. The chain's +analyzer/optimizer/validator make decisions based on this proxy data. + +ACP is Zed's standardized JSON-RPC-over-stdio protocol for driving +coding agents. The official TypeScript SDK +(`@agentclientprotocol/sdk@0.22.1`) ships a stable client. +benchflow/skillsbench has already validated the pattern in production +across the same five agents. Adopting ACP gives skill-optimizer: + +- Realistic measurement: a skill's behavior under Claude Code is what + matters when shipping a skill targeting Claude Code users. +- Apples-to-apples cross-agent comparison: same trace schema across all + five agents. +- One uniform code path: no more "pi via library + others somehow." +- Future-proof: new ACP-supporting agents become registry entries, not + runtime work. + +## Architecture + +### Host / container split + +Host owns all bench orchestration logic. The container is a pure +sandbox: it holds the agent CLI binary, the read-only auth files, the +skill under test, and `/work`. It has no bench logic of its own. + +```text +HOST + docker-runner.ts + - resolve agent from registry + - detect subscription-auth files on host + - prepare workspace + - docker run -d agent-sandbox container (per-trial fresh) + - docker exec -i -> stdio pipes + - ACP client (from @agentclientprotocol/sdk): + initialize -> session/new -> session/prompt + <- session/update events streamed live + - write raw ACP messages verbatim to trace.jsonl + - on prompt completion: docker exec graders + - write result.json + summary.json + - docker rm -f container + +CONTAINER (per-trial, fresh from cached image) + - /opt/skill-opt/agents/ (all 5 CLIs pre-installed at image build) + - /home/agent/.claude/ (auth files + skill under test, ro) + - /work (rw, agent's cwd) + - /case (ro, bundled case) + - /results (rw, graders write here) + - agent process, spawned by docker exec, speaks ACP to host +``` + +### Per-trial container lifecycle + +Each trial gets a fresh container. The Docker **image** is built once +and cached (all five agent CLIs baked in at image build time); the +**container** is per-trial, ensuring full state isolation: + +| State | Mechanism | Isolated? | +|---|---|---| +| Agent conversation memory | ACP `session/new` per trial | yes | +| Filesystem (`/work`) | Container fresh, per-trial tmp mount | yes | +| Agent local state | Container fresh | yes | +| Auth files | Bind-mounted read-only | yes | +| Skill under test | Bind-mounted read-only per trial | yes | +| MCP servers | Sidecar containers per-trial on private network | yes | +| Agent CLI binaries | Image-baked, shared across trials | shared by design | + +`install_cmd` from each agent's registry entry runs at **image build +time** — not per trial. Per-trial `docker run` is ~1s startup, not 30s +of `npm install`. + +`launch_cmd` runs at trial time via `docker exec -i`, stays alive as +the ACP server until the trial completes. + +## Agent registry + +One file: `src/workbench/agents/registry.ts`. TypeScript port of +[benchflow's +`registry.py`](../../../skillsbench/.venv/lib/python3.12/site-packages/benchflow/agents/registry.py). +Per-agent quirks (e.g. `ANTHROPIC_AUTH_TOKEN` vs `_API_KEY`, opencode's +`provider/model` ID format, codex's base-url shell expansion) are +direct ports — those are battle-tested. + +```typescript +export interface AgentConfig { + name: string; + description: string; + installCmd: string; // bash, runs at IMAGE BUILD time + launchCmd: string; // bash, runs per-trial via docker exec + requiresEnv: string[]; // API key env vars + apiProtocol: 'anthropic-messages' | 'openai-completions' + | 'openai-responses' | ''; + envMapping: Record; + skillPaths: string[]; // e.g. ["$HOME/.claude/skills"] + credentialFiles: CredentialFile[]; + homeDirs: string[]; + subscriptionAuth: SubscriptionAuth | null; + acpModelFormat: 'bare' | 'provider/model'; + supportsAcpSetModel: boolean; +} + +export const AGENTS: Record = { + 'claude-agent-acp': { ... }, + 'codex-acp': { ... }, + 'gemini': { ... }, + 'opencode': { ... }, + 'pi-acp': { ... }, +}; + +export const AGENT_ALIASES: Record = { + 'claude': 'claude-agent-acp', + 'codex': 'codex-acp', + 'pi': 'pi-acp', +}; +``` + +Resolution: `resolveAgent(name)` looks up via alias map, then registry, +throws with fuzzy-match suggestion on unknown. + +## Case + suite schema + +### case.yml + +`agent:` is a **required** field. The case loader fails loudly if +missing. No backwards-compatible default. + +```yaml +name: review-product-card +agent: claude-agent-acp # REQUIRED +model: claude-haiku-4-5-20251001 # agent-native model ID +task: | + Review /work/ProductCard.tsx ... + +graders: [...] +setup: [...] +cleanup: [...] +env: [...] +mcpServers: {...} # per-agent native config writing (see MCP section) +mcpServices: {...} +``` + +`model:` field semantics: **agent-native** model ID, not a normalized +form. Each agent expects its own ID format: + +- `claude-agent-acp` → `claude-haiku-4-5-20251001` +- `codex-acp` → `gpt-5-mini` +- `gemini` → `gemini-3.1-pro-preview` +- `opencode` → `google/gemini-3.1-pro-preview` (provider/model) +- `pi-acp` → `openrouter/anthropic/claude-haiku-4-5` (unchanged) + +### suite.yml + +Matrix shape: `runs:` is the **only** matrix dimension. No `models:` +sugar. trials × runs.length = total trials. + +```yaml +name: my-suite +runs: + - agent: claude-agent-acp + model: claude-haiku-4-5-20251001 + - agent: claude-agent-acp + model: claude-sonnet-4-6-20250929 + - agent: codex-acp + model: gpt-5-mini + - agent: gemini + model: gemini-3.1-pro-preview + +cases: + - name: extract-pdf-facts + task: ... + graders: [...] +``` + +### Trial directory naming + +`------/` (gains the `` segment). + +## Skill deployment + +The skill under test is mounted **only at the agent's native skill +path** (Option A from the brainstorming). No `/work//` mounting. + +| Agent | Mount point inside container | +|---|---| +| claude-agent-acp | `/home/agent/.claude/skills//` | +| codex-acp | `/home/agent/.agents/skills//` | +| gemini | `/home/agent/.gemini/skills//` | +| opencode | `/home/agent/.opencode/skills//` | +| pi-acp | `/home/agent/.pi/agent/skills//` | + +Mounted read-only via `-v ::ro`. + +Task prompts describe the **user's actual task** ("Review +/work/ProductCard.tsx"), not "use the skill at X to do Y." This is the +realistic invocation pattern — the agent decides whether to trigger the +skill based on frontmatter `description`. "Skill didn't trigger" then +becomes a measurable weakness class, not a harness convenience. + +Chain implication: `skills/write-tests/agents/test-writer.md` updates +to drop "use the skill at /work/X" boilerplate from generated task +prompts. + +## Authentication + +Resolution priority per agent at trial start: + +1. **Subscription auth (host CLI login files)** — if the agent's + `subscriptionAuth.detectFile` exists on host (e.g., + `~/.claude/.credentials.json`), copy it into the container's + `$HOME` read-only. Done. +2. **API key from env** — if subscription file absent, require each + var in `requiresEnv` to be set on host. Pass into container via + `-e ENV_NAME` and applied per `envMapping`. +3. **Fail loudly** — if neither present, abort the trial with a + clear message indicating how to authenticate (login command or env + var name). + +Per-agent file paths: + +| Agent | Subscription detect path | +|---|---| +| claude-agent-acp | `~/.claude/.credentials.json` | +| codex-acp | `~/.codex/auth.json` | +| gemini | `~/.gemini/oauth_creds.json` (with siblings, see note) [1] | +| opencode | (no subscription auth; provider API keys only) | +| pi-acp | (no subscription auth; provider API keys only) | + +Subscription files are mounted **read-only**. The skill-optimizer +process on host never reads file contents — just checks existence and +hands the path to Docker. + +[1] Gemini's OAuth login writes three files that must all be copied +together: `~/.gemini/oauth_creds.json`, `~/.gemini/settings.json`, +`~/.gemini/google_accounts.json`. The detect path is the first; all +three are listed in the agent's `subscriptionAuth.files`. + +## Trace pipeline + +### Raw ACP capture + +`trace.jsonl` is the **raw ACP wire format**. One JSON-RPC envelope per +line, plus a `trace_start` header with trial metadata. + +```jsonl +{"type":"trace_start","caseName":"...","agent":"claude-agent-acp","model":"...","startedAt":"..."} +{"jsonrpc":"2.0","id":1,"method":"initialize","params":{...}} +{"jsonrpc":"2.0","id":1,"result":{...}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"...","update":{"sessionUpdate":"agent_message_chunk","content":{...}}}} +{"jsonrpc":"2.0","method":"session/update","params":{...,"update":{"sessionUpdate":"tool_call","toolCallId":"...","kind":"execute","content":{...}}}} +... +``` + +No normalization layer. The `WorkbenchTraceEntry` type is deleted. + +### Parse helpers + +`src/workbench/parse-trace.ts` exposes typed views over the raw trace: + +```typescript +iterMessages(trace): IterableIterator<{role, text?, thinking?, stopReason?, usage?}> +iterToolCalls(trace): IterableIterator<{id, kind, args, result?, isError?}> +computeMetrics(trace): WorkbenchMetrics +getFinalAssistantMessage(trace): string | undefined +getFailureEvidence(trace): string[] +``` + +`computeMetrics` returns the existing `WorkbenchMetrics` shape minus +the `cost` field (dropped permanently — tokens × downstream pricing +table if anyone needs cost). + +### Agent-internal trace archive + +In addition to `trace.jsonl`, we `docker cp` the agent's home dot-dir +at trial end into `/agent-internal/`. This captures whatever +the agent's own log machinery wrote (Claude Code's session jsonl, +Codex's session log, etc.) for human debugging. Bonus archive — chain +skills do not read this. + +```text +/ +├── trace.jsonl # canonical, raw ACP wire format +├── agent-internal/ # bonus archive, agent-specific format +│ └── (claude/projects/... | codex/sessions/... | ...) +├── result.json # graded outcome +├── summary.json # derived from trace.jsonl +└── workspace/ # filesystem state (on fail or --keep) +``` + +If disk pressure shows up later, gate `agent-internal/` behind +`--keep-internal-trace` (default off). For v1 it is always on. + +## MCP handling + +Per-agent native config writing. case.yml `mcpServers:` / +`mcpServices:` declarations stay the same; the harness translates them +into each agent's native MCP config file before launch: + +| Agent | Native MCP config target | +|---|---| +| claude-agent-acp | `~/.claude.json` `mcpServers` field | +| codex-acp | `~/.codex/config.toml` `[mcp_servers.*]` blocks | +| gemini | `~/.gemini/settings.json` `mcpServers` | +| opencode | `~/.config/opencode/opencode.json` `mcp` field | +| pi-acp | TBD at implementation (see note) [2] | + +MCP sidecar containers (the `mcpServices:` declaration in case.yml) +keep the existing pattern: spun up per-trial on a private Docker +network, agents reach them over HTTP/SSE. + +Implementation file: `src/workbench/acp/mcp-config-writer.ts`. One +small writer function per agent. Each writer is mechanical (case +config → agent-native shape) and unit-testable against fixture cases. + +[2] pi-acp's native MCP config target needs verification at +implementation time. benchflow's pi-acp registry entry doesn't +document an MCP config path explicitly. Three resolution options for +the implementer to pick from after inspecting pi-acp source: +(a) reuse the existing mcporter pattern (write `mcporter.json`, +inject `MCPORTER_CONFIG` env into the launch shim), (b) write to +pi-agent's native config location if pi-acp exposes one, (c) skip +MCP for pi-acp in v1 (fail loudly if a case using pi-acp declares +`mcpServers:`) and defer. Whichever path is chosen, surface in the +plan as an explicit task with the inspection step preceding the +implementation. + +## Chain integration + +Affected chain skills: + +- **`skills/write-tests/agents/test-writer.md`** — drop "use the skill + at /work/X" boilerplate from generated task prompts. Task prompts + describe the user's task only. + +- **`skills/analyze/agents/analyzer.md`** — analyzer reads ACP-format + `trace.jsonl` instead of normalized `WorkbenchTraceEntry`. Add + pointer to new `skills/shared/acp-trace-format.md` reference. + +- **`skills/run-bench/SKILL.md`** — `06-bench-summary.md` gains + `agent:` column alongside `model:`. Drops cost. Adds tokens + + duration + tool-call counts per probe and overall (closing the + pre-existing metrics gap). + +- **`skills/shared/acp-trace-format.md`** — NEW reference doc (~50 + lines) summarizing ACP `session/update` event types relevant to + skill-behavior analysis, plus pointers at the official ACP spec. + +## Code surface + +### New files + +- `src/workbench/agents/registry.ts` — `AgentConfig` type + 5 entries + aliases +- `src/workbench/agents/install-snippets.ts` — per-agent bash for the Dockerfile +- `src/workbench/acp/client.ts` — wrapper around `@agentclientprotocol/sdk` +- `src/workbench/acp/transport.ts` — `docker exec -i` stdio bridge +- `src/workbench/acp/auth.ts` — subscription detection + mount resolution +- `src/workbench/acp/mcp-config-writer.ts` — case → per-agent MCP config +- `src/workbench/acp/skill-deploy.ts` — mount skill to agent's `skillPaths[0]` +- `src/workbench/parse-trace.ts` — ACP trace helpers +- `docker/skill-optimizer-agent.Dockerfile` — replaces `workbench-runner.Dockerfile` +- `skills/shared/acp-trace-format.md` — chain analyzer reference +- `tests/smoke-agents/.spec.ts` — one smoke probe per agent + +### Deleted + +- `src/workbench/pi-agent.ts` +- `--agent` mode of `src/workbench/container-runner.ts` +- `WorkbenchTraceEntry` type + normalization code in `trace.ts` +- System-prompt logging +- mcporter injection into pi-agent (replaced by `mcp-config-writer.ts`) + +### Kept (modified as needed) + +- `src/workbench/container-runner.ts` — `--setup` and `--grade` modes only +- `src/workbench/docker-runner.ts` — heavily updated to orchestrate ACP host-side +- `src/workbench/trials.ts` — aggregation surfaces tokens + duration +- Workspace prep, MCP sidecar networking, cleanup orchestration + +## Migration + +### Existing tracked cases + +- `examples/workbench/pdf/suite.yml`, `examples/workbench/mcp/suite.yml` + — convert `models:` to `runs:`. Add `agent: pi-acp` per run + (these use OpenRouter refs that pi-acp accepts unchanged). +- Other `examples/workbench/*/.run.log` and stale result directories — + remove (past run artifacts, not source). +- Any `skill-evals//...` directories — past run artifacts, can + be removed entirely. Future probes generated by the updated + test-writer subagent are correct from creation. + +### Existing chain-skill docs + +Updated in the chain integration section above. Mechanical edits, no +schema migration. + +## Testing strategy + +### Per-agent smoke probes (5) + +Each of the 5 agents gets one minimal probe in `tests/smoke-agents/`: + +- Task: "write 'hello' to /work/out.txt; report what you did to + /work/findings.txt" +- Grader: assert `out.txt` exists with content `hello`, assert + `findings.txt` mentions writing it +- Verifies the full lifecycle: image-time install, container launch, + ACP handshake, prompt, tool call (write), termination, trace + well-formedness + +Runs in CI on every PR. Failure of any smoke probe blocks merge. + +### Pi-acp regression probe + +Run the existing `examples/workbench/pdf/suite.yml` through pi-acp +end-to-end. Compare pass rates against a captured baseline from the +last known-good direct-pi-agent run. Catches behavioral drift from the +runtime swap. + +### Unit tests + +- Registry resolution (alias → config, fail-loud on unknown) +- Auth resolution (subscription detected → uses it; missing → falls + back to env; both missing → throws) +- `parse-trace.ts` helpers against fixture ACP traces from each agent +- MCP config translation (one fixture case per agent → assertion on + emitted config file) + +## Out of scope (v1) + +- Cost computation (dropped permanently; tokens are the measurement) +- Web-tools disabling per agent (registry fields exist, no CLI flag + yet — adds when needed) +- Non-Docker sandbox backends (Daytona, e2b, firecracker, k8s) +- Runtime agent registration (hardcoded registry is enough for now) +- Custom agents beyond the 5 + +## Risks + +1. **Per-agent quirks surface only at runtime.** Budget ~1 day per + agent for smoke-test debugging. Known unknowns from benchflow's + experience: codex base-url shell expansion, opencode + `provider/model` ID format, gemini's multi-file subscription auth, + pi-acp's launcher shim for OpenRouter routing. + +2. **ACP version skew.** Zed's SDK (`@agentclientprotocol/sdk@0.22.1`) + and each agent CLI must agree on `session/update` semantics. We + pin SDK version; if a specific agent CLI version breaks, document + the compatible range in the agent's registry entry. + +3. **Subscription auth file paths vary across agent versions** (Claude + Code CLI vs Desktop login, gemini OAuth flow changes). Stick with + benchflow's tested paths; fall back to API key if subscription + detection fails. + +4. **Docker image size grows.** Five agent CLIs ≈ 200-300MB on top of + the base. Mitigation: multi-stage Dockerfile, agent install layer + cached. Acceptable for a controlled workbench. + +5. **Chain re-validation.** Existing `06-bench-summary.md` consumers + (analyzer subagent) need updated mental model. Pi-acp regression + probe catches the major drift; chain skill prompts get updated in + the same plan. + +## Open questions + +None at design time — all design questions resolved during +brainstorming. Implementation will surface tactical questions +(specific ACP message shapes per agent, exact docker exec semantics +for long-lived processes) which the writing-plans skill should +sequence as discrete tasks. From 63d94c96e147f34add6833f2cd6bfb659a25d153 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:12:44 -0500 Subject: [PATCH 065/121] docs(plan): multi-agent ACP workbench implementation plan 26 task plan implementing the spec at docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md. Sequenced as: ACP SDK + transport + client (Tasks 1-3), agent registry + entries (Tasks 4-5), Dockerfile bake (Task 6), auth + skill-deploy + MCP config writer + parse-trace + trace-recorder (Tasks 7-11), schema changes + agent: required + runs: matrix (Tasks 12-13), metrics/trace refactor (Tasks 14-15), docker-runner integration (Task 16), container-runner cleanup + pi-agent.ts deletion (Task 17), trials aggregation + run-case/suite migration (Tasks 18-19), example migration (Task 20), smoke tests + pi-acp regression (Tasks 21-22), chain skill updates + reference doc (Tasks 23-24), docs refresh + final E2E check (Tasks 25-26). Co-Authored-By: Claude Opus 4.7 --- .../plans/2026-05-25-multi-agent-acp.md | 3558 +++++++++++++++++ 1 file changed, 3558 insertions(+) create mode 100644 docs/superpowers/plans/2026-05-25-multi-agent-acp.md diff --git a/docs/superpowers/plans/2026-05-25-multi-agent-acp.md b/docs/superpowers/plans/2026-05-25-multi-agent-acp.md new file mode 100644 index 0000000..0ae6e49 --- /dev/null +++ b/docs/superpowers/plans/2026-05-25-multi-agent-acp.md @@ -0,0 +1,3558 @@ +# Multi-agent ACP Workbench Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use +> superpowers:subagent-driven-development (recommended) or +> superpowers:executing-plans to implement this plan task-by-task. Steps +> use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Replace the embedded pi-agent runtime with a host-side ACP +client driving five agent CLIs (Claude Code, Codex, Gemini, OpenCode, +pi-acp) inside per-trial sandboxed Docker containers. + +**Architecture:** Host orchestrates everything (case load, container +lifecycle, ACP handshake, trace capture, grading). The container is a +pure sandbox with the agent CLI pre-baked at image-build time. The +official `@agentclientprotocol/sdk` provides the ACP runtime; an agent +registry holds per-agent install/launch/auth specifics, ported from +benchflow's battle-tested entries. + +**Tech Stack:** TypeScript (existing `src/workbench/`), +`@agentclientprotocol/sdk@0.22.1`, Docker (existing), `yaml` (existing +case-loader dependency), `node:test` (existing test runner). + +**Spec:** +[`docs/superpowers/specs/2026-05-25-multi-agent-acp-design.md`](../specs/2026-05-25-multi-agent-acp-design.md) + +--- + +## File structure + +### New files + +| Path | Responsibility | +|---|---| +| `src/workbench/acp/transport.ts` | Wrap `docker exec -i` stdio as ACP `Stream` | +| `src/workbench/acp/client.ts` | `ClientSideConnection` wrapper + our `Client` impl | +| `src/workbench/acp/auth.ts` | Subscription detection + Docker mount spec | +| `src/workbench/acp/skill-deploy.ts` | Host skill dir → container mount | +| `src/workbench/acp/mcp-config-writer.ts` | Case `mcpServers:` → per-agent native config | +| `src/workbench/acp/trace-recorder.ts` | Buffer raw ACP messages → `trace.jsonl` | +| `src/workbench/agents/registry.ts` | `AgentConfig` type + 5 entries + `resolveAgent` | +| `src/workbench/agents/install-snippets.ts` | Bash install snippets used by Dockerfile | +| `src/workbench/parse-trace.ts` | `iterMessages` / `iterToolCalls` / `computeMetrics` | +| `docker/skill-optimizer-agent.Dockerfile` | Bakes node + python + all 5 agent CLIs | +| `skills/shared/acp-trace-format.md` | Chain analyzer reference for ACP event types | +| `tests/acp/transport.test.ts` | Mock-subprocess unit tests | +| `tests/acp/auth.test.ts` | Resolution priority unit tests | +| `tests/acp/mcp-config-writer.test.ts` | Per-agent fixture assertions | +| `tests/agents/registry.test.ts` | Alias + resolution unit tests | +| `tests/parse-trace.test.ts` | Fixture trace assertions | +| `tests/smoke-agents/claude-agent-acp.smoke.test.ts` | E2E smoke probe | +| `tests/smoke-agents/codex-acp.smoke.test.ts` | E2E smoke probe | +| `tests/smoke-agents/gemini.smoke.test.ts` | E2E smoke probe | +| `tests/smoke-agents/opencode.smoke.test.ts` | E2E smoke probe | +| `tests/smoke-agents/pi-acp.smoke.test.ts` | E2E smoke probe | +| `tests/regression/pi-acp-pdf-suite.test.ts` | Pi-acp regression against pdf suite | +| `tests/fixtures/acp-traces/-sample.jsonl` | Captured fixture traces | + +### Modified files + +| Path | Change | +|---|---| +| `package.json` | Add `@agentclientprotocol/sdk@0.22.1` dependency | +| `src/workbench/case-loader.ts` | Require `agent:` field, fail loud | +| `src/workbench/suite-loader.ts` | `runs:` matrix replaces `models:` | +| `src/workbench/types.ts` | Add `AgentConfig` types; remove `WorkbenchTraceEntry` | +| `src/workbench/docker-runner.ts` | Heavy rewrite: ACP host-side orchestration | +| `src/workbench/container-runner.ts` | Remove `--agent` mode, keep `--setup` / `--grade` | +| `src/workbench/run-case.ts` | Adapt to `runs:` matrix; per-run trial dirs | +| `src/workbench/run-suite.ts` | Same | +| `src/workbench/trials.ts` | Aggregate tokens + duration (close metrics gap) | +| `src/workbench/metrics.ts` | Thin wrapper around `parse-trace.ts`; drop cost | +| `src/workbench/trace.ts` | Collapse to raw JSONL writer; drop normalization | +| `skills/write-tests/agents/test-writer.md` | Drop `/work/` skill-path boilerplate | +| `skills/analyze/agents/analyzer.md` | Point at `acp-trace-format.md` reference | +| `skills/run-bench/SKILL.md` | Add `agent:` column, drop cost, surface tokens + duration | +| `examples/workbench/pdf/suite.yml` | Convert `models:` → `runs:` with `agent: pi-acp` | +| `examples/workbench/mcp/suite.yml` | Convert `models:` → `runs:` with `agent: pi-acp` | +| `.gitignore` | Add `tests/fixtures/acp-traces/*.local.jsonl` if regenerating | + +### Deleted files + +| Path | Why | +|---|---| +| `src/workbench/pi-agent.ts` | Pi runs through pi-acp; no embedded library | +| `docker/workbench-runner.Dockerfile` | Replaced by `skill-optimizer-agent.Dockerfile` | +| `examples/workbench/*/.run.log` | Stale run logs, not source | +| `examples/workbench/*/.results/` | Stale result dirs, not source | + +--- + +## Tasks + +### Task 1: Add ACP SDK dependency + smoke import + +**Files:** +- Modify: `package.json` +- Create: `tests/acp/sdk-smoke.test.ts` + +- [ ] **Step 1: Add the dependency** + +```bash +npm install --save-exact @agentclientprotocol/sdk@0.22.1 +``` + +- [ ] **Step 2: Write a smoke test asserting key SDK exports are importable** + +Create `tests/acp/sdk-smoke.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; + +test('@agentclientprotocol/sdk exports ClientSideConnection', async () => { + const mod = await import('@agentclientprotocol/sdk'); + assert.equal(typeof mod.ClientSideConnection, 'function'); + assert.equal(typeof mod.ndJsonStream, 'function'); + assert.equal(typeof mod.RequestError, 'function'); +}); +``` + +- [ ] **Step 3: Run the test** + +```bash +npx tsx --test tests/acp/sdk-smoke.test.ts +``` + +Expected: PASS + +- [ ] **Step 4: Run typecheck** + +```bash +npm run typecheck +``` + +Expected: no errors + +- [ ] **Step 5: Commit** + +```bash +git add package.json package-lock.json tests/acp/sdk-smoke.test.ts +git commit -m "feat(acp): add @agentclientprotocol/sdk dependency" +``` + +--- + +### Task 2: ACP transport over `docker exec -i` + +**Files:** +- Create: `src/workbench/acp/transport.ts` +- Create: `tests/acp/transport.test.ts` + +- [ ] **Step 1: Write the failing test (mock subprocess)** + +Create `tests/acp/transport.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { spawn } from 'node:child_process'; +import { createDockerExecStream } from '../../src/workbench/acp/transport.js'; + +test('createDockerExecStream produces a Stream that round-trips ndjson', async () => { + // Use `cat` as a stand-in for the agent: it echoes stdin to stdout. + const child = spawn('cat', [], { stdio: ['pipe', 'pipe', 'pipe'] }); + const stream = createDockerExecStream(child); + + const message = { jsonrpc: '2.0', id: 1, method: 'initialize', params: {} }; + const writer = stream.outgoing.getWriter(); + await writer.write(new TextEncoder().encode(JSON.stringify(message) + '\n')); + writer.releaseLock(); + + const reader = stream.incoming.getReader(); + const { value } = await reader.read(); + const echoed = new TextDecoder().decode(value).trim(); + assert.equal(echoed, JSON.stringify(message)); + + child.kill(); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +```bash +npx tsx --test tests/acp/transport.test.ts +``` + +Expected: FAIL with "Cannot find module './transport.js'" or similar + +- [ ] **Step 3: Implement the transport** + +Create `src/workbench/acp/transport.ts`: + +```typescript +import type { ChildProcessWithoutNullStreams } from 'node:child_process'; +import type { Stream } from '@agentclientprotocol/sdk'; + +export interface DockerExecStream extends Stream { + outgoing: WritableStream; + incoming: ReadableStream; + close(): Promise; +} + +export function createDockerExecStream( + child: ChildProcessWithoutNullStreams, +): DockerExecStream { + const outgoing = new WritableStream({ + write(chunk) { + return new Promise((resolve, reject) => { + child.stdin.write(chunk, (err) => (err ? reject(err) : resolve())); + }); + }, + close() { + child.stdin.end(); + }, + }); + + const incoming = new ReadableStream({ + start(controller) { + child.stdout.on('data', (chunk: Buffer) => controller.enqueue(chunk)); + child.stdout.on('end', () => controller.close()); + child.stdout.on('error', (err) => controller.error(err)); + }, + }); + + return { + outgoing, + incoming, + async close() { + try { child.stdin.end(); } catch {} + child.kill(); + }, + }; +} +``` + +- [ ] **Step 4: Run test to verify it passes** + +```bash +npx tsx --test tests/acp/transport.test.ts +``` + +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/acp/transport.ts tests/acp/transport.test.ts +git commit -m "feat(acp): docker exec stdio bridge implementing Stream" +``` + +--- + +### Task 3: ACP client wrapper + +**Files:** +- Create: `src/workbench/acp/client.ts` +- Create: `tests/acp/client.test.ts` + +- [ ] **Step 1: Write the failing test** + +Create `tests/acp/client.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { spawn } from 'node:child_process'; +import { createWorkbenchClient } from '../../src/workbench/acp/client.js'; +import { createDockerExecStream } from '../../src/workbench/acp/transport.js'; + +test('createWorkbenchClient returns object with initialize/newSession/prompt/cancel', () => { + const child = spawn('cat', []); + const stream = createDockerExecStream(child); + const client = createWorkbenchClient({ stream, onSessionUpdate: () => {} }); + + assert.equal(typeof client.connection.initialize, 'function'); + assert.equal(typeof client.connection.newSession, 'function'); + assert.equal(typeof client.connection.prompt, 'function'); + assert.equal(typeof client.connection.cancel, 'function'); + assert.equal(typeof client.close, 'function'); + + child.kill(); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +```bash +npx tsx --test tests/acp/client.test.ts +``` + +Expected: FAIL + +- [ ] **Step 3: Implement the wrapper** + +Create `src/workbench/acp/client.ts`: + +```typescript +import { ClientSideConnection, ndJsonStream } from '@agentclientprotocol/sdk'; +import type { + Client, + SessionNotification, + RequestPermissionRequest, + RequestPermissionResponse, +} from '@agentclientprotocol/sdk'; +import type { DockerExecStream } from './transport.js'; + +export interface WorkbenchClientOptions { + stream: DockerExecStream; + onSessionUpdate: (notification: SessionNotification) => void; +} + +export interface WorkbenchClient { + connection: ClientSideConnection; + close(): Promise; +} + +export function createWorkbenchClient(opts: WorkbenchClientOptions): WorkbenchClient { + // Note: ndJsonStream wraps the raw byte streams as JSON-RPC envelopes. + const ndjson = ndJsonStream(opts.stream.outgoing, opts.stream.incoming); + + const connection = new ClientSideConnection( + (_agent): Client => ({ + sessionUpdate: async (params) => { + opts.onSessionUpdate(params); + }, + // Auto-approve all tool calls. The container is the security boundary, + // not the permission gate. Approving everything matches the headless- + // bench model and matches benchflow's pattern. + requestPermission: async ( + params: RequestPermissionRequest, + ): Promise => ({ + outcome: { outcome: 'selected', optionId: params.options[0]?.optionId ?? 'allow' }, + }), + // We don't expose filesystem capabilities to the agent over ACP; the + // agent uses its own tools to touch /work directly. + }), + ndjson, + ); + + return { + connection, + async close() { + await opts.stream.close(); + }, + }; +} +``` + +- [ ] **Step 4: Run test to verify it passes** + +```bash +npx tsx --test tests/acp/client.test.ts && npm run typecheck +``` + +Expected: PASS + no type errors + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/acp/client.ts tests/acp/client.test.ts +git commit -m "feat(acp): client wrapper with auto-approve permission handler" +``` + +--- + +### Task 4: Agent registry types + +**Files:** +- Create: `src/workbench/agents/registry.ts` +- Create: `tests/agents/registry.test.ts` + +- [ ] **Step 1: Write the failing test** + +Create `tests/agents/registry.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { resolveAgent, AGENTS, AGENT_ALIASES } from '../../src/workbench/agents/registry.js'; + +test('AGENTS contains all 5 expected entries', () => { + for (const name of ['claude-agent-acp', 'codex-acp', 'gemini', 'opencode', 'pi-acp']) { + assert.ok(AGENTS[name], `missing agent: ${name}`); + assert.equal(AGENTS[name].name, name); + } +}); + +test('resolveAgent accepts aliases', () => { + assert.equal(resolveAgent('claude').name, 'claude-agent-acp'); + assert.equal(resolveAgent('codex').name, 'codex-acp'); + assert.equal(resolveAgent('pi').name, 'pi-acp'); +}); + +test('resolveAgent accepts canonical names', () => { + assert.equal(resolveAgent('claude-agent-acp').name, 'claude-agent-acp'); +}); + +test('resolveAgent throws with suggestion on unknown', () => { + assert.throws(() => resolveAgent('claud-agent'), /Did you mean/); +}); + +test('every agent declares at least one skillPath', () => { + for (const cfg of Object.values(AGENTS)) { + assert.ok(cfg.skillPaths.length > 0, `${cfg.name} missing skillPaths`); + assert.match(cfg.skillPaths[0], /^\$HOME\//); + } +}); + +test('AGENT_ALIASES does not collide with canonical names', () => { + for (const alias of Object.keys(AGENT_ALIASES)) { + if (AGENTS[alias]) { + assert.equal(alias, AGENT_ALIASES[alias], `alias ${alias} collides`); + } + } +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +```bash +npx tsx --test tests/agents/registry.test.ts +``` + +Expected: FAIL + +- [ ] **Step 3: Implement the registry types and resolver** + +Create `src/workbench/agents/registry.ts`: + +```typescript +export interface CredentialFile { + path: string; // Target path in container; may use {home} + envSource: string; // Env var on host to read value from + template?: string; // If set, value is inserted into template at {value} + mkdir?: boolean; // Create parent dir; default true +} + +export interface HostAuthFile { + hostPath: string; // ~/.claude/.credentials.json + containerPath: string; // {home}/.claude/.credentials.json +} + +export interface SubscriptionAuth { + replacesEnv: string; // e.g. "ANTHROPIC_API_KEY" + detectFile: string; // host path to check for login + files: HostAuthFile[]; // all files to copy when sub-auth used +} + +export type ApiProtocol = + | 'anthropic-messages' + | 'openai-completions' + | 'openai-responses' + | ''; + +export type AcpModelFormat = 'bare' | 'provider/model'; + +export interface AgentConfig { + name: string; + description: string; + installCmd: string; // bash, runs at IMAGE BUILD time + launchCmd: string; // bash, runs per-trial via docker exec + requiresEnv: string[]; + apiProtocol: ApiProtocol; + envMapping: Record; // SKILL_OPT_PROVIDER_* → agent-native + skillPaths: string[]; // e.g. ["$HOME/.claude/skills"] + credentialFiles: CredentialFile[]; + homeDirs: string[]; + subscriptionAuth: SubscriptionAuth | null; + acpModelFormat: AcpModelFormat; + supportsAcpSetModel: boolean; + loginHint?: string; // shown in fail-loud message +} + +export const AGENTS: Record = { + // Entries filled in Task 5. Empty here to make this task self-contained. +}; + +export const AGENT_ALIASES: Record = { + claude: 'claude-agent-acp', + codex: 'codex-acp', + gemini: 'gemini', + pi: 'pi-acp', + openclaw: 'openclaw', +}; + +export function resolveAgent(spec: string): AgentConfig { + const canonical = AGENT_ALIASES[spec] ?? spec; + const cfg = AGENTS[canonical]; + if (cfg) return cfg; + + const known = Object.keys(AGENTS); + const close = closestMatch(canonical, known); + if (close) { + throw new Error(`Unknown agent: ${spec!r}. Did you mean: ${close!r}?`.replace(/!r/g, '')); + } + throw new Error(`Unknown agent: ${spec}. Available: ${known.join(', ')}`); +} + +function closestMatch(needle: string, haystack: string[]): string | undefined { + // Simple shared-prefix heuristic; good enough for typo suggestion. + let best: { name: string; score: number } | undefined; + for (const name of haystack) { + const score = sharedPrefix(needle, name) + sharedPrefix(needle.split('').reverse().join(''), name.split('').reverse().join('')); + if (score > 4 && (!best || score > best.score)) { + best = { name, score }; + } + } + return best?.name; +} + +function sharedPrefix(a: string, b: string): number { + let i = 0; + while (i < a.length && i < b.length && a[i] === b[i]) i++; + return i; +} +``` + +- [ ] **Step 4: Run test to verify it fails differently (now: missing entries)** + +```bash +npx tsx --test tests/agents/registry.test.ts +``` + +Expected: FAIL on "missing agent" assertions (entries empty by design at this task) + +- [ ] **Step 5: Commit the skeleton** + +```bash +git add src/workbench/agents/registry.ts tests/agents/registry.test.ts +git commit -m "feat(agents): registry types + resolver (entries pending in task 5)" +``` + +--- + +### Task 5: Populate registry with 5 agent entries + +**Files:** +- Modify: `src/workbench/agents/registry.ts` + +These are direct ports of [benchflow's registry entries](../../../skillsbench/.venv/lib/python3.12/site-packages/benchflow/agents/registry.py). +The install snippets become `BUILD-TIME` only — no `command -v ... ||` +guards needed since the Dockerfile runs them once. + +- [ ] **Step 1: Add the `claude-agent-acp` entry** + +In `src/workbench/agents/registry.ts`, replace the empty `AGENTS = {}` +with: + +```typescript +export const AGENTS: Record = { + 'claude-agent-acp': { + name: 'claude-agent-acp', + description: 'Claude Code via ACP (Anthropic CLI)', + installCmd: `npm install -g @zed-industries/claude-agent-acp@latest`, + launchCmd: `claude-agent-acp`, + requiresEnv: ['ANTHROPIC_API_KEY'], + apiProtocol: 'anthropic-messages', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'ANTHROPIC_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'ANTHROPIC_AUTH_TOKEN', + SKILL_OPT_PROVIDER_MODEL: 'ANTHROPIC_MODEL', + }, + skillPaths: ['$HOME/.claude/skills'], + credentialFiles: [], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'ANTHROPIC_API_KEY', + detectFile: '~/.claude/.credentials.json', + files: [ + { hostPath: '~/.claude/.credentials.json', containerPath: '{home}/.claude/.credentials.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'claude login', + }, +}; +``` + +- [ ] **Step 2: Add the `codex-acp` entry** + +Append to the `AGENTS` object: + +```typescript + 'codex-acp': { + name: 'codex-acp', + description: 'OpenAI Codex via ACP', + installCmd: `npm install -g @zed-industries/codex-acp@latest`, + launchCmd: `codex-acp \${OPENAI_BASE_URL:+-c openai_base_url=$OPENAI_BASE_URL}`, + requiresEnv: ['OPENAI_API_KEY'], + apiProtocol: 'openai-responses', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'OPENAI_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'OPENAI_API_KEY', + }, + skillPaths: ['$HOME/.agents/skills'], + credentialFiles: [ + { + path: '{home}/.codex/auth.json', + envSource: 'OPENAI_API_KEY', + template: '{"OPENAI_API_KEY": "{value}"}', + }, + ], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'OPENAI_API_KEY', + detectFile: '~/.codex/auth.json', + files: [ + { hostPath: '~/.codex/auth.json', containerPath: '{home}/.codex/auth.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'codex login', + }, +``` + +- [ ] **Step 3: Add the `gemini` entry** + +Append to the `AGENTS` object: + +```typescript + 'gemini': { + name: 'gemini', + description: 'Google Gemini CLI via ACP', + installCmd: `npm install -g @google/gemini-cli@latest`, + launchCmd: `gemini --acp --yolo`, + requiresEnv: ['GOOGLE_API_KEY'], + apiProtocol: '', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'GEMINI_API_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'GOOGLE_API_KEY', + }, + skillPaths: ['$HOME/.gemini/skills'], + credentialFiles: [], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'GEMINI_API_KEY', + detectFile: '~/.gemini/oauth_creds.json', + files: [ + { hostPath: '~/.gemini/oauth_creds.json', containerPath: '{home}/.gemini/oauth_creds.json' }, + { hostPath: '~/.gemini/settings.json', containerPath: '{home}/.gemini/settings.json' }, + { hostPath: '~/.gemini/google_accounts.json', containerPath: '{home}/.gemini/google_accounts.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'gemini auth login', + }, +``` + +- [ ] **Step 4: Add the `opencode` entry** + +Append to the `AGENTS` object: + +```typescript + 'opencode': { + name: 'opencode', + description: 'OpenCode via ACP — open-source coding agent', + installCmd: `npm install -g opencode-ai@latest`, + launchCmd: `opencode acp`, + requiresEnv: [], + apiProtocol: '', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'OPENAI_BASE_URL', + }, + skillPaths: ['$HOME/.opencode/skills'], + credentialFiles: [], + homeDirs: ['.opencode'], + subscriptionAuth: null, + acpModelFormat: 'provider/model', + supportsAcpSetModel: true, + loginHint: 'set OPENAI_API_KEY (or provider-specific key)', + }, +``` + +- [ ] **Step 5: Add the `pi-acp` entry** + +Append to the `AGENTS` object. The pi-acp launcher needs a wrapper +that bridges `SKILL_OPT_PROVIDER_*` env vars into pi's config; we +deploy this wrapper script in Task 6 (the Dockerfile). + +```typescript + 'pi-acp': { + name: 'pi-acp', + description: 'Pi coding agent via ACP', + installCmd: `npm install -g @mariozechner/pi-coding-agent@latest pi-acp@latest`, + launchCmd: `/opt/skill-opt/bin/pi-acp-launcher`, + requiresEnv: [], + apiProtocol: '', + envMapping: {}, + skillPaths: ['$HOME/.pi/agent/skills', '$HOME/.agents/skills'], + credentialFiles: [], + homeDirs: ['.pi'], + subscriptionAuth: null, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'set OPENROUTER_API_KEY', + }, +``` + +- [ ] **Step 6: Run the registry tests** + +```bash +npx tsx --test tests/agents/registry.test.ts +``` + +Expected: all PASS + +- [ ] **Step 7: Run typecheck** + +```bash +npm run typecheck +``` + +Expected: no errors + +- [ ] **Step 8: Commit** + +```bash +git add src/workbench/agents/registry.ts +git commit -m "feat(agents): populate 5-agent registry (claude/codex/gemini/opencode/pi-acp)" +``` + +--- + +### Task 6: Build the new Dockerfile + +**Files:** +- Create: `docker/skill-optimizer-agent.Dockerfile` +- Create: `src/workbench/agents/install-snippets.ts` +- Create: `docker/pi-acp-launcher.sh` + +- [ ] **Step 1: Create the pi-acp launcher script** + +Create `docker/pi-acp-launcher.sh`: + +```bash +#!/bin/sh +# Bridges SKILL_OPT_PROVIDER_* env vars to pi-acp's expected config. +# pi-acp reads its model/provider from a runtime config; we set +# OPENROUTER_API_KEY here if SKILL_OPT_PROVIDER_API_KEY is set. +set -e + +if [ -n "$SKILL_OPT_PROVIDER_API_KEY" ] && [ -z "$OPENROUTER_API_KEY" ]; then + export OPENROUTER_API_KEY="$SKILL_OPT_PROVIDER_API_KEY" +fi + +exec pi-acp "$@" +``` + +- [ ] **Step 2: Generate install snippets from the registry** + +Create `src/workbench/agents/install-snippets.ts`: + +```typescript +import { AGENTS } from './registry.js'; + +export function generateDockerInstallBlock(): string { + const lines: string[] = []; + for (const cfg of Object.values(AGENTS)) { + lines.push(`# Install ${cfg.name}: ${cfg.description}`); + lines.push(`RUN ${cfg.installCmd}`); + lines.push(''); + } + return lines.join('\n'); +} + +// Used at Dockerfile generation time; printed by `npm run dockerfile:print`. +if (import.meta.url === `file://${process.argv[1]}`) { + console.log(generateDockerInstallBlock()); +} +``` + +- [ ] **Step 3: Add npm script to print install block** + +Edit `package.json`, add to `scripts`: + +```json +"dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts" +``` + +- [ ] **Step 4: Write the new Dockerfile** + +Create `docker/skill-optimizer-agent.Dockerfile`: + +```dockerfile +FROM node:22-bookworm + +ENV PATH="/opt/skill-opt/bin:/app/node_modules/.bin:/work/.venv/bin:${PATH}" \ + PIP_REQUIRE_VIRTUALENV=1 + +WORKDIR /app + +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + bash \ + ca-certificates \ + coreutils \ + curl \ + file \ + findutils \ + gawk \ + git \ + grep \ + jq \ + less \ + python-is-python3 \ + python3 \ + python3-pip \ + python3-venv \ + ripgrep \ + sed \ + unzip \ + wget \ + zip \ + && rm -rf /var/lib/apt/lists/* + +# --- Agent CLI install layer (cached as one layer for build speed) --- +# Each agent's installCmd is also kept in src/workbench/agents/registry.ts +# (single source of truth). If you change one here, change it there too. +RUN npm install -g \ + @zed-industries/claude-agent-acp@latest \ + @zed-industries/codex-acp@latest \ + @google/gemini-cli@latest \ + opencode-ai@latest \ + @mariozechner/pi-coding-agent@latest \ + pi-acp@latest + +# --- pi-acp launcher wrapper --- +COPY docker/pi-acp-launcher.sh /opt/skill-opt/bin/pi-acp-launcher +RUN chmod +x /opt/skill-opt/bin/pi-acp-launcher + +# --- Workbench container-runner (setup + grade modes only) --- +COPY package.json package-lock.json tsconfig.json ./ +COPY src ./src +COPY docs ./docs + +RUN npm ci \ + && npm run build \ + && useradd -m -u 10001 agent +USER agent + +# Container-runner is now used only for --setup and --grade modes. +# Agent dispatch happens host-side via ACP; the agent CLIs above are +# invoked directly by docker exec. +ENTRYPOINT ["node", "/app/dist/workbench/container-runner.js"] +``` + +- [ ] **Step 5: Build the image** + +```bash +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . +``` + +Expected: successful build; final image ~600MB - ~1GB + +- [ ] **Step 6: Smoke-test that each agent binary is reachable** + +```bash +for agent in claude-agent-acp codex-acp gemini opencode pi-acp; do + echo "=== $agent ===" + docker run --rm --entrypoint sh skill-optimizer-agent:local -lc "command -v $agent && $agent --version 2>&1 | head -2" +done +``` + +Expected: each agent prints a version line (or a usage line that confirms the binary exists). Failure of any agent means the install snippet is wrong; fix at registry.ts AND the Dockerfile RUN line. + +- [ ] **Step 7: Delete the old Dockerfile** + +```bash +git rm docker/workbench-runner.Dockerfile +``` + +- [ ] **Step 8: Update the project's docker:build reference** + +Search and replace `workbench-runner.Dockerfile` and `skill-optimizer-workbench:local`: + +```bash +grep -rln "workbench-runner.Dockerfile\|skill-optimizer-workbench:local" src/ CLAUDE.md README.md CONTRIBUTING.md +``` + +For each file in the result, update the references to `skill-optimizer-agent.Dockerfile` and `skill-optimizer-agent:local`. + +Files known to mention these: +- `src/workbench/docker-runner.ts` (constant `DEFAULT_WORKBENCH_IMAGE`) +- `CLAUDE.md` (test guidance) +- `README.md` (install docs) +- `CONTRIBUTING.md` + +- [ ] **Step 9: Run typecheck** + +```bash +npm run typecheck +``` + +- [ ] **Step 10: Commit** + +```bash +git add docker/ package.json src/workbench/agents/install-snippets.ts \ + src/workbench/docker-runner.ts CLAUDE.md README.md CONTRIBUTING.md +git commit -m "feat(docker): new image with all 5 agents pre-baked" +``` + +--- + +### Task 7: Auth resolution module + +**Files:** +- Create: `src/workbench/acp/auth.ts` +- Create: `tests/acp/auth.test.ts` + +- [ ] **Step 1: Write the failing tests** + +Create `tests/acp/auth.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, writeFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { resolveAuth } from '../../src/workbench/acp/auth.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +test('resolveAuth uses subscription file when present', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + writeFileSync(join(home, '.credentials.json'), '{}'); + + const auth = resolveAuth(AGENTS['claude-agent-acp'], { + home, + env: {}, + }); + + assert.equal(auth.mode, 'subscription'); + assert.equal(auth.files.length, 1); + assert.equal(auth.files[0].hostPath, join(home, '.credentials.json')); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth falls back to env API key when subscription absent', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + const auth = resolveAuth(AGENTS['claude-agent-acp'], { + home, + env: { ANTHROPIC_API_KEY: 'sk-test-key' }, + }); + + assert.equal(auth.mode, 'env'); + assert.deepEqual(auth.envNames, ['ANTHROPIC_API_KEY']); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth throws when neither subscription nor env key present', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + assert.throws( + () => resolveAuth(AGENTS['claude-agent-acp'], { home, env: {} }), + /claude login|ANTHROPIC_API_KEY/, + ); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth returns env mode for agents without subscriptionAuth', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + const auth = resolveAuth(AGENTS['pi-acp'], { + home, + env: { OPENROUTER_API_KEY: 'sk-test' }, + }); + + assert.equal(auth.mode, 'env'); + + rmSync(home, { recursive: true }); +}); +``` + +- [ ] **Step 2: Run tests to verify they fail** + +```bash +npx tsx --test tests/acp/auth.test.ts +``` + +Expected: FAIL (module missing) + +- [ ] **Step 3: Implement auth resolution** + +Create `src/workbench/acp/auth.ts`: + +```typescript +import { existsSync } from 'node:fs'; +import { homedir } from 'node:os'; +import { join, normalize } from 'node:path'; +import type { AgentConfig, HostAuthFile } from '../agents/registry.js'; + +export interface AuthContext { + home?: string; // override $HOME for testing + env: Record; +} + +export interface SubscriptionAuthResult { + mode: 'subscription'; + files: Array<{ hostPath: string; containerPath: string }>; +} + +export interface EnvAuthResult { + mode: 'env'; + envNames: string[]; +} + +export type AuthResult = SubscriptionAuthResult | EnvAuthResult; + +export function resolveAuth(agent: AgentConfig, ctx: AuthContext): AuthResult { + const home = ctx.home ?? homedir(); + + if (agent.subscriptionAuth) { + const detectPath = expandHome(agent.subscriptionAuth.detectFile, home); + if (existsSync(detectPath)) { + return { + mode: 'subscription', + files: agent.subscriptionAuth.files.map((f) => ({ + hostPath: expandHome(f.hostPath, home), + containerPath: f.containerPath, // {home} placeholder, resolved at mount time + })), + }; + } + } + + const missing = agent.requiresEnv.filter((name) => !ctx.env[name]); + if (missing.length > 0 && agent.requiresEnv.length > 0) { + const hint = agent.loginHint ? ` (run \`${agent.loginHint}\` or set ${missing.join(', ')})` : ''; + throw new Error( + `Agent ${agent.name} requires auth but none is available${hint}. Missing: ${missing.join(', ')}.`, + ); + } + + return { mode: 'env', envNames: agent.requiresEnv }; +} + +function expandHome(p: string, home: string): string { + if (p.startsWith('~/')) return normalize(join(home, p.slice(2))); + if (p === '~') return home; + return p; +} + +export function resolveContainerPath(template: string, agentHome: string): string { + return template.replace('{home}', agentHome); +} +``` + +- [ ] **Step 4: Run tests to verify they pass** + +```bash +npx tsx --test tests/acp/auth.test.ts && npm run typecheck +``` + +Expected: all PASS, no type errors + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/acp/auth.ts tests/acp/auth.test.ts +git commit -m "feat(acp): subscription-first auth resolution with fail-loud env fallback" +``` + +--- + +### Task 8: Skill deployment module + +**Files:** +- Create: `src/workbench/acp/skill-deploy.ts` +- Create: `tests/acp/skill-deploy.test.ts` + +- [ ] **Step 1: Write the failing test** + +Create `tests/acp/skill-deploy.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { computeSkillMount } from '../../src/workbench/acp/skill-deploy.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +test('computeSkillMount returns native path for claude-agent-acp', () => { + const mount = computeSkillMount({ + agent: AGENTS['claude-agent-acp'], + skillSlug: 'web-design-guidelines', + hostSkillDir: '/host/path/to/skill', + agentHome: '/home/agent', + }); + + assert.equal(mount.hostPath, '/host/path/to/skill'); + assert.equal(mount.containerPath, '/home/agent/.claude/skills/web-design-guidelines'); + assert.equal(mount.readOnly, true); +}); + +test('computeSkillMount handles agents whose skill_paths use $HOME', () => { + for (const cfg of Object.values(AGENTS)) { + const mount = computeSkillMount({ + agent: cfg, + skillSlug: 'test-skill', + hostSkillDir: '/host/dir', + agentHome: '/home/agent', + }); + assert.ok(!mount.containerPath.includes('$HOME')); + assert.ok(mount.containerPath.startsWith('/home/agent/')); + assert.ok(mount.containerPath.endsWith('/test-skill')); + } +}); +``` + +- [ ] **Step 2: Run tests to verify they fail** + +```bash +npx tsx --test tests/acp/skill-deploy.test.ts +``` + +- [ ] **Step 3: Implement skill-deploy** + +Create `src/workbench/acp/skill-deploy.ts`: + +```typescript +import { join } from 'node:path'; +import type { AgentConfig } from '../agents/registry.js'; + +export interface SkillMount { + hostPath: string; + containerPath: string; + readOnly: boolean; +} + +export interface ComputeSkillMountParams { + agent: AgentConfig; + skillSlug: string; + hostSkillDir: string; // absolute path on host to the skill folder + agentHome: string; // /home/agent (or whatever the container user's $HOME is) +} + +export function computeSkillMount(params: ComputeSkillMountParams): SkillMount { + // Use the first skillPath (primary discovery location for the agent). + const skillRoot = params.agent.skillPaths[0]; + if (!skillRoot) { + throw new Error(`Agent ${params.agent.name} has no skillPaths configured`); + } + const expanded = skillRoot.replace('$HOME', params.agentHome); + return { + hostPath: params.hostSkillDir, + containerPath: join(expanded, params.skillSlug), + readOnly: true, + }; +} + +export function dockerMountFlag(mount: SkillMount): string { + const ro = mount.readOnly ? ':ro' : ':rw'; + return `-v ${shellQuote(mount.hostPath)}:${shellQuote(mount.containerPath)}${ro}`; +} + +function shellQuote(s: string): string { + return `'${s.replace(/'/g, `'\\''`)}'`; +} +``` + +- [ ] **Step 4: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/acp/skill-deploy.test.ts && npm run typecheck +``` + +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/acp/skill-deploy.ts tests/acp/skill-deploy.test.ts +git commit -m "feat(acp): skill-deploy module mounts to agent's native skill path" +``` + +--- + +### Task 9: MCP config writer + +**Files:** +- Create: `src/workbench/acp/mcp-config-writer.ts` +- Create: `tests/acp/mcp-config-writer.test.ts` +- Create: `tests/fixtures/mcp-cases/sample-stdio.yml` + +Per the spec's note [2], pi-acp's MCP target is TBD-at-implementation. +This task implements the four known agents (claude/codex/gemini/opencode); +pi-acp falls back to a "not yet supported" error. Resolution for pi-acp +will be in Task 9b. + +- [ ] **Step 1: Create the fixture case** + +Create `tests/fixtures/mcp-cases/sample-stdio.yml`: + +```yaml +name: sample-with-mcp +agent: claude-agent-acp +model: claude-haiku-4-5 +task: do nothing +graders: + - name: noop + command: 'true' +mcpServers: + calculator: + command: node + args: ['/work/mcp/calculator.mjs'] +``` + +- [ ] **Step 2: Write the failing test** + +Create `tests/acp/mcp-config-writer.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, readFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { writeMcpConfig } from '../../src/workbench/acp/mcp-config-writer.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +const sampleCase = { + mcpServers: { + calculator: { command: 'node', args: ['/work/mcp/calculator.mjs'] }, + }, +}; + +test('writeMcpConfig for claude writes to .claude.json mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['claude-agent-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.claude.json'), 'utf-8')); + assert.ok(data.mcpServers?.calculator); + assert.equal(data.mcpServers.calculator.command, 'node'); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for codex writes to .codex/config.toml', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['codex-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const toml = readFileSync(join(home, '.codex/config.toml'), 'utf-8'); + assert.match(toml, /\[mcp_servers\.calculator\]/); + assert.match(toml, /command = "node"/); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for gemini writes to .gemini/settings.json mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['gemini'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.gemini/settings.json'), 'utf-8')); + assert.ok(data.mcpServers?.calculator); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for opencode writes to .config/opencode/opencode.json mcp', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['opencode'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.config/opencode/opencode.json'), 'utf-8')); + assert.ok(data.mcp?.calculator); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for pi-acp throws "not yet supported" (pending Task 9b)', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + assert.throws( + () => writeMcpConfig({ + agent: AGENTS['pi-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }), + /pi-acp MCP support pending/, + ); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig is a no-op when caseConfig has no mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['claude-agent-acp'], + caseConfig: {}, + agentHomeOnHost: home, + }); + // No files should be written. + rmSync(home, { recursive: true }); +}); +``` + +- [ ] **Step 3: Run tests to verify they fail** + +```bash +npx tsx --test tests/acp/mcp-config-writer.test.ts +``` + +- [ ] **Step 4: Implement the writer** + +Create `src/workbench/acp/mcp-config-writer.ts`: + +```typescript +import { mkdirSync, readFileSync, writeFileSync, existsSync } from 'node:fs'; +import { dirname, join } from 'node:path'; +import type { AgentConfig } from '../agents/registry.js'; + +export interface McpServerSpec { + command?: string; + args?: string[]; + env?: Record; + url?: string; + headers?: Record; +} + +export interface WriteMcpConfigParams { + agent: AgentConfig; + caseConfig: { mcpServers?: Record }; + agentHomeOnHost: string; // path on host that will become container's $HOME +} + +export function writeMcpConfig(params: WriteMcpConfigParams): void { + const servers = params.caseConfig.mcpServers; + if (!servers || Object.keys(servers).length === 0) return; + + switch (params.agent.name) { + case 'claude-agent-acp': + return writeClaude(servers, params.agentHomeOnHost); + case 'codex-acp': + return writeCodex(servers, params.agentHomeOnHost); + case 'gemini': + return writeGemini(servers, params.agentHomeOnHost); + case 'opencode': + return writeOpencode(servers, params.agentHomeOnHost); + case 'pi-acp': + throw new Error('pi-acp MCP support pending — see plan Task 9b'); + default: + throw new Error(`MCP not implemented for agent ${params.agent.name}`); + } +} + +function writeClaude(servers: Record, home: string): void { + const path = join(home, '.claude.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcpServers = { ...(existing.mcpServers ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} + +function writeCodex(servers: Record, home: string): void { + const path = join(home, '.codex/config.toml'); + mkdirSync(dirname(path), { recursive: true }); + const blocks: string[] = []; + for (const [name, spec] of Object.entries(servers)) { + blocks.push(`[mcp_servers.${name}]`); + if (spec.command) blocks.push(`command = ${JSON.stringify(spec.command)}`); + if (spec.args) blocks.push(`args = ${JSON.stringify(spec.args)}`); + if (spec.env) { + blocks.push(`[mcp_servers.${name}.env]`); + for (const [k, v] of Object.entries(spec.env)) { + blocks.push(`${k} = ${JSON.stringify(v)}`); + } + } + blocks.push(''); + } + writeFileSync(path, blocks.join('\n')); +} + +function writeGemini(servers: Record, home: string): void { + const path = join(home, '.gemini/settings.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcpServers = { ...(existing.mcpServers ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} + +function writeOpencode(servers: Record, home: string): void { + const path = join(home, '.config/opencode/opencode.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcp = { ...(existing.mcp ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} +``` + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/acp/mcp-config-writer.test.ts && npm run typecheck +``` + +Expected: all PASS + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/acp/mcp-config-writer.ts tests/acp/mcp-config-writer.test.ts tests/fixtures/mcp-cases/ +git commit -m "feat(acp): per-agent MCP config writer (claude/codex/gemini/opencode)" +``` + +--- + +### Task 9b: Resolve pi-acp MCP target + +**Files:** +- Modify: `src/workbench/acp/mcp-config-writer.ts` +- Modify: `tests/acp/mcp-config-writer.test.ts` + +- [ ] **Step 1: Inspect pi-acp source to determine native MCP target** + +Run: + +```bash +docker run --rm --entrypoint sh skill-optimizer-agent:local -lc "find / -name 'pi-acp' -type f 2>/dev/null | head -5; npm root -g" +docker run --rm --entrypoint sh skill-optimizer-agent:local -lc "cat \$(npm root -g)/pi-acp/package.json | head -20" +docker run --rm --entrypoint sh skill-optimizer-agent:local -lc "find \$(npm root -g)/pi-acp -name '*.js' | xargs grep -l 'mcp' | head" +``` + +Decision criteria (pick one): +- **(a) pi-acp has a native MCP config path** (e.g., `~/.pi/agent/mcp.json`) → implement a `writePiAcp` function like the other agents +- **(b) pi-acp uses mcporter via env** (e.g., reads `MCPORTER_CONFIG` env at launch) → write `mcporter.json` to a known path and inject `MCPORTER_CONFIG` into the launch env +- **(c) pi-acp has no MCP support** → keep the "not yet supported" throw; document the limitation in `examples/workbench/mcp/README.md` + +- [ ] **Step 2: Document the decision in a comment block** + +Add to top of `src/workbench/acp/mcp-config-writer.ts`: + +```typescript +// pi-acp MCP resolution (from plan Task 9b investigation on 2026-XX-XX): +// Decision: +// Evidence: +// Implementation: +``` + +- [ ] **Step 3: Implement the chosen option** + +For option (a) — add to `writeMcpConfig` switch: + +```typescript +case 'pi-acp': + return writePiAcp(servers, params.agentHomeOnHost); +``` + +And add the `writePiAcp` function appropriate to the path found. + +For option (b) — write `mcporter.json` to the discovered path and return additional launch env (this requires extending the function's return type to optionally carry env additions). + +For option (c) — keep the throw, add a regression test that confirms the case-loader rejects `mcpServers:` for pi-acp at load time. + +- [ ] **Step 4: Update the pi-acp test in `mcp-config-writer.test.ts`** + +Replace the "pending" test with one that asserts the implemented behavior. Example for option (a): + +```typescript +test('writeMcpConfig for pi-acp writes to pi-agent native MCP config', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['pi-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + // Assert based on whichever path was discovered. + const path = join(home, '.pi/agent/mcp.json'); // adjust per evidence + assert.ok(existsSync(path)); + rmSync(home, { recursive: true }); +}); +``` + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/acp/mcp-config-writer.test.ts && npm run typecheck +``` + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/acp/mcp-config-writer.ts tests/acp/mcp-config-writer.test.ts +git commit -m "feat(acp): resolve pi-acp MCP target — option " +``` + +--- + +### Task 10: parse-trace.ts helpers + +**Files:** +- Create: `src/workbench/parse-trace.ts` +- Create: `tests/parse-trace.test.ts` +- Create: `tests/fixtures/acp-traces/claude-sample.jsonl` + +- [ ] **Step 1: Create the fixture trace** + +Create `tests/fixtures/acp-traces/claude-sample.jsonl`. Each line is +the raw ACP JSON-RPC envelope as captured by the trace-recorder: + +```jsonl +{"type":"trace_start","schemaVersion":2,"caseName":"sample","agent":"claude-agent-acp","model":"claude-haiku-4-5","startedAt":"2026-05-25T00:00:00Z"} +{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"0.22","clientCapabilities":{}}} +{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"0.22","agentCapabilities":{},"authMethods":[]}} +{"jsonrpc":"2.0","id":2,"method":"session/new","params":{"cwd":"/work","mcpServers":[]}} +{"jsonrpc":"2.0","id":2,"result":{"sessionId":"s-1"}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"Thinking about it..."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"agent_message_chunk","content":{"type":"text","text":"I'll write the file now."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"tool_call","toolCallId":"tc-1","title":"Write file","kind":"edit","status":"in_progress","content":[{"type":"content","content":{"type":"text","text":"out.txt"}}]}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"tool_call_update","toolCallId":"tc-1","status":"completed","content":[{"type":"content","content":{"type":"text","text":"Wrote 6 bytes"}}]}}} +{"jsonrpc":"2.0","id":3,"method":"session/prompt","params":{"sessionId":"s-1","prompt":[{"type":"text","text":"Write hello to out.txt"}]}} +{"jsonrpc":"2.0","id":3,"result":{"stopReason":"end_turn","usage":{"inputTokens":120,"outputTokens":45,"cacheReadTokens":0,"cacheCreationTokens":0}}} +``` + +- [ ] **Step 2: Write the failing test** + +Create `tests/parse-trace.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { readFileSync } from 'node:fs'; +import { + iterMessages, + iterToolCalls, + computeMetrics, + getFinalAssistantMessage, +} from '../src/workbench/parse-trace.js'; + +const fixturePath = 'tests/fixtures/acp-traces/claude-sample.jsonl'; +const trace = readFileSync(fixturePath, 'utf-8'); + +test('iterMessages yields assistant text + thinking', () => { + const msgs = [...iterMessages(trace)]; + const assistant = msgs.find((m) => m.role === 'assistant'); + assert.ok(assistant); + assert.match(assistant.text ?? '', /write the file/); + assert.match(assistant.thinking ?? '', /Thinking about it/); +}); + +test('iterToolCalls pairs tool_call and tool_call_update', () => { + const calls = [...iterToolCalls(trace)]; + assert.equal(calls.length, 1); + assert.equal(calls[0].id, 'tc-1'); + assert.equal(calls[0].kind, 'edit'); + assert.equal(calls[0].status, 'completed'); + assert.match(calls[0].resultText ?? '', /Wrote 6 bytes/); +}); + +test('computeMetrics returns tokens, duration, tool counts', () => { + const m = computeMetrics(trace); + assert.equal(m.tokens.input, 120); + assert.equal(m.tokens.output, 45); + assert.equal(m.tokens.total, 165); + assert.equal(m.toolCalls, 1); + assert.equal(m.editCalls, 1); + assert.equal(m.bashCalls, 0); + assert.equal(m.stopReason, 'end_turn'); +}); + +test('getFinalAssistantMessage returns concatenated assistant text', () => { + const final = getFinalAssistantMessage(trace); + assert.match(final ?? '', /write the file/); +}); +``` + +- [ ] **Step 3: Run tests to verify they fail** + +```bash +npx tsx --test tests/parse-trace.test.ts +``` + +- [ ] **Step 4: Implement parse-trace** + +Create `src/workbench/parse-trace.ts`: + +```typescript +export interface ParsedMessage { + role: 'user' | 'assistant'; + text?: string; + thinking?: string; + timestamp?: string; +} + +export interface ParsedToolCall { + id: string; + kind: string; // execute, read, write, edit, search, fetch, think, other + title?: string; + status: 'in_progress' | 'completed' | 'failed' | 'pending'; + argsText?: string; + resultText?: string; + isError?: boolean; +} + +export interface ComputedMetrics { + durationMs: number; + turns: number; + toolCalls: number; + bashCalls: number; + readCalls: number; + writeCalls: number; + editCalls: number; + stopReason?: string; + tokens: { input: number; output: number; cacheRead: number; cacheWrite: number; total: number }; +} + +interface TraceHeader { + type: 'trace_start'; + caseName?: string; + agent?: string; + model?: string; + startedAt?: string; + endedAt?: string; +} + +function parseLines(jsonl: string): unknown[] { + return jsonl.split(/\r?\n/).filter(Boolean).map((line) => { + try { return JSON.parse(line); } catch { return null; } + }).filter((x) => x !== null); +} + +function getHeader(rows: unknown[]): TraceHeader | undefined { + for (const row of rows) { + if (typeof row === 'object' && row !== null && (row as any).type === 'trace_start') { + return row as TraceHeader; + } + } + return undefined; +} + +function extractText(content: unknown): string | undefined { + if (!content || typeof content !== 'object') return undefined; + const c = content as any; + if (typeof c.text === 'string') return c.text; + if (c.content && typeof c.content.text === 'string') return c.content.text; + if (Array.isArray(c)) { + return c.map((part: any) => extractText(part)).filter(Boolean).join('\n') || undefined; + } + return undefined; +} + +export function* iterMessages(jsonl: string): IterableIterator { + const rows = parseLines(jsonl); + // Buffer per role; ACP streams assistant text in chunks. + let assistantText = ''; + let assistantThinking = ''; + + for (const row of rows) { + const r = row as any; + if (r.method !== 'session/update') continue; + const update = r.params?.update; + if (!update) continue; + + switch (update.sessionUpdate) { + case 'agent_message_chunk': + assistantText += extractText(update.content) ?? ''; + break; + case 'agent_thought_chunk': + assistantThinking += extractText(update.content) ?? ''; + break; + case 'user_message_chunk': + yield { role: 'user', text: extractText(update.content) }; + break; + } + } + + if (assistantText || assistantThinking) { + yield { + role: 'assistant', + text: assistantText || undefined, + thinking: assistantThinking || undefined, + }; + } +} + +export function* iterToolCalls(jsonl: string): IterableIterator { + const rows = parseLines(jsonl); + const calls = new Map(); + + for (const row of rows) { + const r = row as any; + if (r.method !== 'session/update') continue; + const update = r.params?.update; + if (!update) continue; + + if (update.sessionUpdate === 'tool_call') { + calls.set(update.toolCallId, { + id: update.toolCallId, + kind: update.kind ?? 'other', + title: update.title, + status: update.status ?? 'in_progress', + argsText: extractText(update.content), + }); + } else if (update.sessionUpdate === 'tool_call_update') { + const existing = calls.get(update.toolCallId); + if (existing) { + existing.status = update.status ?? existing.status; + if (update.content) existing.resultText = extractText(update.content); + if (update.status === 'failed') existing.isError = true; + } + } + } + + yield* calls.values(); +} + +export function computeMetrics(jsonl: string): ComputedMetrics { + const rows = parseLines(jsonl); + const header = getHeader(rows); + + const m: ComputedMetrics = { + durationMs: 0, + turns: 0, + toolCalls: 0, + bashCalls: 0, + readCalls: 0, + writeCalls: 0, + editCalls: 0, + tokens: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 }, + }; + + if (header?.startedAt) { + const start = Date.parse(header.startedAt); + const end = header.endedAt ? Date.parse(header.endedAt) : Date.now(); + if (Number.isFinite(start) && Number.isFinite(end)) { + m.durationMs = Math.max(0, end - start); + } + } + + for (const row of rows) { + const r = row as any; + if (r.method === 'session/update') { + const update = r.params?.update; + if (!update) continue; + if (update.sessionUpdate === 'agent_message_chunk') m.turns += 1; + if (update.sessionUpdate === 'tool_call') { + m.toolCalls += 1; + const kind = update.kind ?? 'other'; + if (kind === 'execute') m.bashCalls += 1; + if (kind === 'read') m.readCalls += 1; + if (kind === 'write') m.writeCalls += 1; + if (kind === 'edit') m.editCalls += 1; + } + } + // session/prompt response carries final usage + stopReason + if (r.id && r.result?.stopReason) { + m.stopReason = r.result.stopReason; + const u = r.result.usage; + if (u) { + m.tokens.input += Number(u.inputTokens ?? 0); + m.tokens.output += Number(u.outputTokens ?? 0); + m.tokens.cacheRead += Number(u.cacheReadTokens ?? 0); + m.tokens.cacheWrite += Number(u.cacheCreationTokens ?? 0); + } + } + } + + m.tokens.total = m.tokens.input + m.tokens.output + m.tokens.cacheRead + m.tokens.cacheWrite; + return m; +} + +export function getFinalAssistantMessage(jsonl: string): string | undefined { + for (const msg of iterMessages(jsonl)) { + if (msg.role === 'assistant' && msg.text) return msg.text; + } + return undefined; +} + +export function getFailureEvidence(jsonl: string): string[] { + const evidence: string[] = []; + for (const call of iterToolCalls(jsonl)) { + if (call.isError && call.resultText) { + evidence.push(`tool ${call.kind} (${call.id}) failed: ${call.resultText}`); + } + } + return evidence; +} +``` + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/parse-trace.test.ts && npm run typecheck +``` + +Expected: all PASS + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/parse-trace.ts tests/parse-trace.test.ts tests/fixtures/acp-traces/ +git commit -m "feat(parse-trace): ACP trace helpers (iterMessages/iterToolCalls/computeMetrics)" +``` + +--- + +### Task 11: Trace recorder + +**Files:** +- Create: `src/workbench/acp/trace-recorder.ts` +- Create: `tests/acp/trace-recorder.test.ts` + +- [ ] **Step 1: Write the failing test** + +Create `tests/acp/trace-recorder.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, readFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { createTraceRecorder } from '../../src/workbench/acp/trace-recorder.js'; + +test('createTraceRecorder writes header then verbatim messages', () => { + const dir = mkdtempSync(join(tmpdir(), 'trace-rec-')); + const path = join(dir, 'trace.jsonl'); + const rec = createTraceRecorder({ + tracePath: path, + header: { + caseName: 'x', + agent: 'claude-agent-acp', + model: 'claude-haiku-4-5', + startedAt: '2026-05-25T00:00:00Z', + }, + }); + rec.recordRaw({ jsonrpc: '2.0', id: 1, method: 'initialize', params: {} }); + rec.recordRaw({ jsonrpc: '2.0', id: 1, result: {} }); + rec.finalize('2026-05-25T00:00:01Z'); + + const lines = readFileSync(path, 'utf-8').trim().split('\n'); + assert.equal(lines.length, 3); + const header = JSON.parse(lines[0]); + assert.equal(header.type, 'trace_start'); + assert.equal(header.agent, 'claude-agent-acp'); + assert.equal(header.endedAt, '2026-05-25T00:00:01Z'); + const msg = JSON.parse(lines[1]); + assert.equal(msg.method, 'initialize'); + + rmSync(dir, { recursive: true }); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +```bash +npx tsx --test tests/acp/trace-recorder.test.ts +``` + +- [ ] **Step 3: Implement the recorder** + +Create `src/workbench/acp/trace-recorder.ts`: + +```typescript +import { writeFileSync, appendFileSync, mkdirSync } from 'node:fs'; +import { dirname } from 'node:path'; + +export interface TraceHeader { + caseName: string; + agent: string; + model: string; + startedAt: string; + endedAt?: string; +} + +export interface TraceRecorder { + recordRaw(message: unknown): void; + finalize(endedAt?: string): void; +} + +export function createTraceRecorder(params: { + tracePath: string; + header: TraceHeader; +}): TraceRecorder { + mkdirSync(dirname(params.tracePath), { recursive: true }); + const headerLine = JSON.stringify({ + type: 'trace_start', + schemaVersion: 2, + ...params.header, + }); + // Write header up front so partial traces are still self-describing on crash. + writeFileSync(params.tracePath, headerLine + '\n'); + + return { + recordRaw(message: unknown) { + appendFileSync(params.tracePath, JSON.stringify(message) + '\n'); + }, + finalize(endedAt?: string) { + if (!endedAt) return; + // Rewrite header with endedAt; preserve remaining lines. + const updatedHeader = JSON.stringify({ + type: 'trace_start', + schemaVersion: 2, + ...params.header, + endedAt, + }); + const existing = require('node:fs').readFileSync(params.tracePath, 'utf-8') as string; + const rest = existing.split('\n').slice(1).join('\n'); + writeFileSync(params.tracePath, updatedHeader + '\n' + rest); + }, + }; +} +``` + +- [ ] **Step 4: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/acp/trace-recorder.test.ts && npm run typecheck +``` + +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/acp/trace-recorder.ts tests/acp/trace-recorder.test.ts +git commit -m "feat(acp): trace recorder writes raw ACP messages with header+endedAt" +``` + +--- + +### Task 12: case.yml schema — require `agent:` field + +**Files:** +- Modify: `src/workbench/case-loader.ts` +- Modify: `src/workbench/types.ts` +- Create: `tests/case-loader-agent-required.test.ts` + +- [ ] **Step 1: Add `agent` + `skillUnderTest` to `WorkbenchCaseConfig` and `ResolvedWorkbenchCase`** + +Edit `src/workbench/types.ts`. In `WorkbenchCaseConfig`: + +```typescript +export interface SkillUnderTestSpec { + slug: string; // e.g., "web-design-guidelines" + hostPath: string; // absolute path on host to the skill dir containing SKILL.md +} + +export interface WorkbenchCaseConfig { + name: string; + references: string; + task: string; + graders: WorkbenchGraderConfig[]; + agent: string; // NEW — required + model: string; // semantics now agent-native + skillUnderTest?: SkillUnderTestSpec; // NEW — optional; if set, mount to agent's native skill path + // ... existing fields below unchanged +} +``` + +And in `ResolvedWorkbenchCase`: + +```typescript +export interface ResolvedWorkbenchCase { + configPath: string; + configDir: string; + name: string; + referencesDir: string; + task: string; + graders: WorkbenchGraderConfig[]; + mcpServers: WorkbenchMcpServersConfig; + mcpServices: WorkbenchMcpServicesConfig; + env: string[]; + setup: string[]; + cleanup: string[]; + agent: string; // NEW + model: string; + skillUnderTest?: SkillUnderTestSpec; // NEW + timeoutSeconds: number; +} +``` + +Document the convention: when a case is run by the chain skills, the +chain operator sets `skillUnderTest.hostPath` to either +`docs/skill-optimizer//improved-skill/` (if it exists) or +`.skill-optimizer//vendored-skill/` (else). For standalone +workbench usage (the examples in `examples/workbench/`), the field is +omitted — those cases don't need skill discovery. + +- [ ] **Step 2: Write the failing test** + +Create `tests/case-loader-agent-required.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, writeFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { loadWorkbenchCase } from '../src/workbench/case-loader.js'; + +test('loadWorkbenchCase throws when agent: is missing', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +task: do nothing +graders: + - name: noop + command: 'true' +model: openrouter/anthropic/claude-haiku-4-5 +`); + assert.throws( + () => loadWorkbenchCase(join(dir, 'case.yml')), + /agent.*required|missing.*agent/i, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchCase throws when agent: is unknown', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +agent: nonexistent-agent +model: x +task: do nothing +graders: + - name: noop + command: 'true' +`); + assert.throws( + () => loadWorkbenchCase(join(dir, 'case.yml')), + /Unknown agent/, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchCase accepts valid agent', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +agent: pi-acp +model: openrouter/anthropic/claude-haiku-4-5 +task: do nothing +graders: + - name: noop + command: 'true' +`); + // Refs dir doesn't exist; loader may complain about references. Test + // only that agent validation passes — catch and ignore non-agent errors. + try { + const c = loadWorkbenchCase(join(dir, 'case.yml')); + assert.equal(c.agent, 'pi-acp'); + } catch (e: any) { + assert.doesNotMatch(e.message, /agent/i); + } + rmSync(dir, { recursive: true }); +}); +``` + +- [ ] **Step 3: Run test to see it fails (loader doesn't validate yet)** + +```bash +npx tsx --test tests/case-loader-agent-required.test.ts +``` + +- [ ] **Step 4: Modify `case-loader.ts` to enforce the field** + +In `src/workbench/case-loader.ts`, find the parsing/validation function +and add agent validation. Add at top of the file: + +```typescript +import { resolveAgent } from './agents/registry.js'; +``` + +In the function that builds `ResolvedWorkbenchCase` from parsed YAML, +add (near the top of validation): + +```typescript +if (!parsed.agent || typeof parsed.agent !== 'string') { + throw new Error( + `Case ${configPath} is missing required field \`agent:\`. ` + + `Valid agents: claude-agent-acp, codex-acp, gemini, opencode, pi-acp.`, + ); +} +// Validate the agent exists (throws with suggestion on unknown). +const agentCfg = resolveAgent(parsed.agent); +``` + +And include `agent: agentCfg.name` in the returned object (canonical +name, in case the input was an alias). + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/case-loader-agent-required.test.ts && npm run typecheck +``` + +Expected: all PASS + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/case-loader.ts src/workbench/types.ts tests/case-loader-agent-required.test.ts +git commit -m "feat(schema): require agent: field in case.yml, fail loud on missing/unknown" +``` + +--- + +### Task 13: suite.yml schema — `runs:` matrix + +**Files:** +- Modify: `src/workbench/suite-loader.ts` +- Modify: `src/workbench/types.ts` +- Create: `tests/suite-loader-runs.test.ts` + +- [ ] **Step 1: Add `WorkbenchRunSpec` type** + +Edit `src/workbench/types.ts`, add: + +```typescript +export interface WorkbenchRunSpec { + agent: string; + model: string; +} + +// If WorkbenchSuiteConfig exists in this file, modify it: +// - REMOVE: models: string[]; +// - ADD: runs: WorkbenchRunSpec[]; +``` + +- [ ] **Step 2: Write the failing test** + +Create `tests/suite-loader-runs.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, writeFileSync, rmSync, mkdirSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { loadWorkbenchSuite } from '../src/workbench/suite-loader.js'; + +test('loadWorkbenchSuite parses runs: matrix', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +runs: + - agent: claude-agent-acp + model: claude-haiku-4-5 + - agent: pi-acp + model: openrouter/anthropic/claude-haiku-4-5 +cases: + - name: test-case + task: do nothing + graders: + - name: noop + command: 'true' +`); + const suite = loadWorkbenchSuite(join(dir, 'suite.yml')); + assert.equal(suite.runs.length, 2); + assert.equal(suite.runs[0].agent, 'claude-agent-acp'); + assert.equal(suite.runs[1].agent, 'pi-acp'); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchSuite rejects legacy `models:` shape', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +models: + - openrouter/anthropic/claude-haiku-4-5 +cases: [] +`); + assert.throws( + () => loadWorkbenchSuite(join(dir, 'suite.yml')), + /runs:.*replaced.*models:/i, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchSuite validates each run.agent', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +runs: + - agent: bogus + model: x +cases: [] +`); + assert.throws( + () => loadWorkbenchSuite(join(dir, 'suite.yml')), + /Unknown agent.*bogus/, + ); + rmSync(dir, { recursive: true }); +}); +``` + +- [ ] **Step 3: Run test to verify it fails** + +```bash +npx tsx --test tests/suite-loader-runs.test.ts +``` + +- [ ] **Step 4: Modify suite-loader.ts** + +In `src/workbench/suite-loader.ts`, add: + +```typescript +import { resolveAgent } from './agents/registry.js'; +``` + +In the loader function, after parsing the YAML: + +```typescript +if ('models' in parsed && !('runs' in parsed)) { + throw new Error( + `Suite ${configPath} uses legacy 'models:' field. ` + + `Replace with 'runs:' matrix (list of { agent, model } objects).`, + ); +} +if (!Array.isArray(parsed.runs) || parsed.runs.length === 0) { + throw new Error(`Suite ${configPath} must declare a non-empty 'runs:' list.`); +} +for (const run of parsed.runs) { + if (!run.agent || !run.model) { + throw new Error(`Each run in ${configPath} must have agent: and model: fields.`); + } + // Validate; throws on unknown agent. + run.agent = resolveAgent(run.agent).name; +} +``` + +In the return object, replace `models:` with `runs:`. + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/suite-loader-runs.test.ts && npm run typecheck +``` + +Expected: all PASS + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/suite-loader.ts src/workbench/types.ts tests/suite-loader-runs.test.ts +git commit -m "feat(schema): suite.yml uses runs: matrix; legacy models: rejected" +``` + +--- + +### Task 14: Refactor `metrics.ts` as thin wrapper over `parse-trace` + +**Files:** +- Modify: `src/workbench/metrics.ts` +- Modify: `src/workbench/types.ts` + +- [ ] **Step 1: Remove `cost` from `WorkbenchMetrics`** + +In `src/workbench/types.ts`: + +```typescript +export interface WorkbenchMetrics { + durationMs: number; + turns: number; + toolCalls: number; + toolResults: number; // KEEP for backward compat with downstream readers; equals toolCalls completed + bashCalls: number; + readCalls: number; + writeCalls: number; + editCalls: number; + stopReason?: string; + tokens: WorkbenchTokenMetrics; + // REMOVED: cost: WorkbenchCostMetrics +} +``` + +Delete the `WorkbenchCostMetrics` interface entirely. + +- [ ] **Step 2: Rewrite `metrics.ts` as thin wrapper** + +Replace `src/workbench/metrics.ts`: + +```typescript +import { readFileSync } from 'node:fs'; +import { computeMetrics as computeFromTrace, getFailureEvidence, getFinalAssistantMessage } from './parse-trace.js'; +import type { WorkbenchMetrics, WorkbenchResult, WorkbenchTrialSummaryFile } from './types.js'; + +export function buildWorkbenchMetricsFromTrace(tracePath: string): WorkbenchMetrics { + const jsonl = readFileSync(tracePath, 'utf-8'); + const m = computeFromTrace(jsonl); + return { + durationMs: m.durationMs, + turns: m.turns, + toolCalls: m.toolCalls, + toolResults: m.toolCalls, // ACP doesn't distinguish; treat each call as one result + bashCalls: m.bashCalls, + readCalls: m.readCalls, + writeCalls: m.writeCalls, + editCalls: m.editCalls, + stopReason: m.stopReason, + tokens: m.tokens, + }; +} + +export function buildTrialSummary(params: { + tracePath: string; + result: WorkbenchResult; +}): WorkbenchTrialSummaryFile { + const jsonl = readFileSync(params.tracePath, 'utf-8'); + const metrics = params.result.metrics ?? buildWorkbenchMetricsFromTrace(params.tracePath); + const failedGraders = params.result.graders?.filter((g) => !g.pass).map((g) => g.name) ?? []; + return { + finalAssistantMessage: getFinalAssistantMessage(jsonl), + failedGraders, + evidence: [...params.result.evidence, ...getFailureEvidence(jsonl)], + bashCommands: extractBashCommands(jsonl), + stopReason: metrics.stopReason, + errorMessage: undefined, + metrics, + }; +} + +function extractBashCommands(jsonl: string): string[] { + const out: string[] = []; + const rows = jsonl.split(/\r?\n/).filter(Boolean).map((l) => { + try { return JSON.parse(l); } catch { return null; } + }); + for (const row of rows) { + if ((row as any)?.method !== 'session/update') continue; + const u = (row as any).params?.update; + if (u?.sessionUpdate === 'tool_call' && u.kind === 'execute') { + const cmd = u.content?.[0]?.content?.text ?? u.argsText; + if (typeof cmd === 'string') out.push(cmd); + } + } + return out; +} +``` + +- [ ] **Step 3: Update any callers that referenced `metrics.cost`** + +```bash +grep -rn "metrics.cost\|cost:" src/workbench/ | grep -v test +``` + +For each match, delete the `cost:` line / property. + +- [ ] **Step 4: Run typecheck and existing tests** + +```bash +npm run typecheck && npm test +``` + +Expected: no type errors; tests pass (some may need adjustment if they +asserted on the cost field — fix those by removing cost assertions). + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/metrics.ts src/workbench/types.ts +git commit -m "refactor(metrics): thin wrapper over parse-trace; drop cost field" +``` + +--- + +### Task 15: Refactor `trace.ts` — drop normalization layer + +**Files:** +- Modify: `src/workbench/trace.ts` +- Modify: `src/workbench/types.ts` + +- [ ] **Step 1: Delete `WorkbenchTraceEntry` from types.ts** + +In `src/workbench/types.ts`, delete the `WorkbenchTraceEntry` union +type and the `WorkbenchTraceEvent` interface. Keep `WorkbenchTrace` but +simplify: + +```typescript +export interface WorkbenchTrace { + schemaVersion?: 2; // bumped from 1 (raw ACP format) + caseName: string; + agent: string; // NEW + model: string; + startedAt: string; + endedAt: string; + // No more entries[]; the trace.jsonl file IS the source of truth. +} +``` + +- [ ] **Step 2: Replace `trace.ts` with raw-only writer** + +Replace `src/workbench/trace.ts` with: + +```typescript +// Re-export trace-recorder as the public API. +export { createTraceRecorder } from './acp/trace-recorder.js'; +export type { TraceHeader, TraceRecorder } from './acp/trace-recorder.js'; + +// Legacy buildWorkbenchTrace removed; raw ACP capture replaces it. +``` + +- [ ] **Step 3: Update callers that imported `buildWorkbenchTrace`, `createTraceCollector`, `WorkbenchTraceEntry`** + +```bash +grep -rn "buildWorkbenchTrace\|createTraceCollector\|WorkbenchTraceEntry\|WorkbenchTraceEvent" src/workbench/ | grep -v test +``` + +The only legitimate caller after Task 16 will be docker-runner.ts. +For now, comment out those imports with `// TODO: replaced by ACP +client in Task 16` to allow typecheck to proceed. (We'll delete those +TODOs in Task 16.) + +- [ ] **Step 4: Run typecheck** + +```bash +npm run typecheck +``` + +Expected: passes after TODO comments hide unresolved references. + +- [ ] **Step 5: Commit** + +```bash +git add src/workbench/trace.ts src/workbench/types.ts +git commit -m "refactor(trace): drop WorkbenchTraceEntry; trace.jsonl is raw ACP" +``` + +--- + +### Task 16: Integrate ACP into `docker-runner.ts` + +This is the integration centerpiece. Replace the existing agent-spawn +flow inside `runDockerWorkbenchCase` with: per-trial container, +auth/skill/MCP setup, ACP client lifecycle, trace recording. + +**Files:** +- Modify: `src/workbench/docker-runner.ts` + +- [ ] **Step 1: Add the new imports** + +At the top of `src/workbench/docker-runner.ts`, add: + +```typescript +import { spawn } from 'node:child_process'; +import { resolveAgent } from './agents/registry.js'; +import { resolveAuth, resolveContainerPath } from './acp/auth.js'; +import { computeSkillMount, dockerMountFlag } from './acp/skill-deploy.js'; +import { writeMcpConfig } from './acp/mcp-config-writer.js'; +import { createTraceRecorder } from './acp/trace-recorder.js'; +import { createDockerExecStream } from './acp/transport.js'; +import { createWorkbenchClient } from './acp/client.js'; +``` + +- [ ] **Step 2: Add a helper that runs one agent trial via ACP** + +Add the following function. It runs ONE trial: starts the container, +mounts auth/skill, exec's the agent, drives the ACP handshake, captures +the trace, and runs graders. + +```typescript +async function runOneAcpTrial(params: { + resolvedCase: ResolvedWorkbenchCase; + tempDir: string; + workDir: string; + caseDir: string; + resultsDir: string; + agentName: string; + model: string; + image: string; + hostSkillDir: string; + skillSlug: string; + repoRoot: string; + timeoutSeconds: number; +}): Promise<{ pass: boolean; tracePath: string; resultPath: string }> { + const agent = resolveAgent(params.agentName); + const auth = resolveAuth(agent, { env: process.env }); + + // Stage auth files + MCP config into a host-side dir that we'll bind-mount. + const agentHomeHost = join(params.tempDir, 'agent-home'); + mkdirSync(agentHomeHost, { recursive: true }); + if (auth.mode === 'subscription') { + for (const f of auth.files) { + const targetPath = resolveContainerPath(f.containerPath, agentHomeHost); + mkdirSync(dirname(targetPath), { recursive: true }); + cpSync(f.hostPath, targetPath); + } + } + writeMcpConfig({ + agent, + caseConfig: { mcpServers: params.resolvedCase.mcpServers as any }, + agentHomeOnHost: agentHomeHost, + }); + + // Compute mounts + const skillMount = computeSkillMount({ + agent, + skillSlug: params.skillSlug, + hostSkillDir: params.hostSkillDir, + agentHome: '/home/agent', + }); + const authMounts = [`-v ${shellQuote(`${agentHomeHost}:/home/agent:rw`)}`]; + const skillMountFlag = dockerMountFlag(skillMount); + + // Env passthrough + const envFlags: string[] = []; + if (auth.mode === 'env') { + for (const name of auth.envNames) { + if (process.env[name]) envFlags.push(`-e ${name}`); + } + } + for (const name of params.resolvedCase.env) { + if (process.env[name]) envFlags.push(`-e ${name}`); + } + + // Start the container (detached, idle) + const containerName = `skill-opt-trial-${params.tempDir.split('/').pop()}`; + const runCmd = [ + 'docker run -d', + `--name ${shellQuote(containerName)}`, + '--cap-drop=ALL', '--security-opt no-new-privileges', '--pids-limit 512', + `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, + `-v ${shellQuote(`${params.caseDir}:/case:ro`)}`, + `-v ${shellQuote(`${params.resultsDir}:/results:rw`)}`, + ...authMounts, + skillMountFlag, + ...envFlags, + '--workdir /work', + '--entrypoint sleep', + shellQuote(params.image), + 'infinity', + ].join(' '); + const runResult = await runShellCommand(runCmd, { cwd: params.repoRoot }); + if (runResult.exitCode !== 0) { + throw new Error(`Failed to start container: ${runResult.stderr}`); + } + + const tracePath = join(params.resultsDir, 'trace.jsonl'); + const resultPath = join(params.resultsDir, 'result.json'); + const startedAt = new Date().toISOString(); + const recorder = createTraceRecorder({ + tracePath, + header: { + caseName: params.resolvedCase.name, + agent: agent.name, + model: params.model, + startedAt, + }, + }); + + try { + // Run setup phase (still inside the container, via the existing entrypoint) + if (params.resolvedCase.setup.length > 0) { + const setupCmd = `docker exec ${shellQuote(containerName)} sh -c ${shellQuote( + params.resolvedCase.setup.join(' && '), + )}`; + const setupResult = await runShellCommand(setupCmd, { cwd: params.repoRoot }); + if (setupResult.exitCode !== 0) { + throw new Error(`Setup failed: ${setupResult.stderr}`); + } + } + + // Exec the agent CLI inside the container with stdin/stdout piped + const child = spawn('docker', [ + 'exec', '-i', + containerName, + 'sh', '-c', agent.launchCmd, + ], { stdio: ['pipe', 'pipe', 'pipe'] }); + + const stream = createDockerExecStream(child as any); + const client = createWorkbenchClient({ + stream, + onSessionUpdate: (notification) => { + // Re-wrap as JSON-RPC notification for the trace + recorder.recordRaw({ + jsonrpc: '2.0', + method: 'session/update', + params: notification, + }); + }, + }); + + // ACP handshake + await client.connection.initialize({ + protocolVersion: '0.22', + clientCapabilities: {}, + }); + const session = await client.connection.newSession({ + cwd: '/work', + mcpServers: [], + }); + const promptPromise = client.connection.prompt({ + sessionId: session.sessionId, + prompt: [{ type: 'text', text: params.resolvedCase.task }], + }); + + // Timeout wrapper + const promptResult = await Promise.race([ + promptPromise.then((r) => ({ ok: true, result: r }) as const), + new Promise<{ ok: false }>((_, reject) => + setTimeout(() => reject(new Error(`Timeout after ${params.timeoutSeconds}s`)), params.timeoutSeconds * 1000), + ), + ]); + + if (promptResult.ok) { + // Record the final response with its usage + recorder.recordRaw({ jsonrpc: '2.0', id: 'final', result: promptResult.result }); + } + + await client.close(); + recorder.finalize(new Date().toISOString()); + + // Run graders + const gradeCmd = `docker exec ${shellQuote(containerName)} node /app/dist/workbench/container-runner.js --grade --case /case/case.yml --work /work --results /results`; + const gradeResult = await runShellCommand(gradeCmd, { cwd: params.repoRoot }); + + // Copy agent-internal traces out. Each agent writes to its own dot-dir; + // we copy any that exist. Failure to copy (dir absent) is non-fatal — the + // archive is a debugging bonus, not the canonical trace. + const internalDir = join(params.resultsDir, 'agent-internal'); + mkdirSync(internalDir, { recursive: true }); + for (const dotDir of ['.claude', '.codex', '.gemini', '.opencode', '.pi']) { + await runShellCommand( + `docker cp ${shellQuote(`${containerName}:/home/agent/${dotDir}`)} ${shellQuote(internalDir)} 2>/dev/null || true`, + { cwd: params.repoRoot }, + ); + } + + const pass = readTrialPass(resultPath) ?? false; + return { pass, tracePath, resultPath }; + } finally { + await runShellCommand(`docker rm -f ${shellQuote(containerName)}`, { cwd: params.repoRoot }); + } +} +``` + +- [ ] **Step 3: Replace the body of `runDockerWorkbenchCase` to call `runOneAcpTrial`** + +In the existing `runDockerWorkbenchCase` function, after `prepareDockerWorkbenchRun`, +replace the agent + grade phase block with: + +```typescript +// Compute the skill mount params if the case declares a skill under test. +// (Cases that don't declare skillUnderTest skip the skill mount entirely — +// the agent runs without any skill-discovery deployment.) +const skillMountParams = resolvedCase.skillUnderTest + ? { + hostSkillDir: resolvedCase.skillUnderTest.hostPath, + skillSlug: resolvedCase.skillUnderTest.slug, + } + : null; + +const trialOutcome = await runOneAcpTrial({ + resolvedCase, + tempDir: prepared.tempDir, + workDir: prepared.workDir, + caseDir: prepared.caseDir, + resultsDir: prepared.resultsDir, + agentName: resolvedCase.agent, + model: options.model ?? resolvedCase.model, + image, + skillMountParams, + repoRoot, + timeoutSeconds: resolvedCase.timeoutSeconds, +}); +``` + +Adjust `runOneAcpTrial`'s signature to accept `skillMountParams: { hostSkillDir, skillSlug } | null` +and branch on it in step 2 above: only call `computeSkillMount` / +`dockerMountFlag` when non-null. If null, no `-v` for the skill. + +- [ ] **Step 4: Delete the old MCP service / setup / agent code paths from `runDockerWorkbenchCase`** + +Remove the calls to: +- `startMcpServices` (replaced by per-agent native MCP config inside `runOneAcpTrial`) +- `waitForMcpServices` +- The `buildDockerAgentCommand` path (the old --agent flow) + +Keep `removeContainer` / `removeDockerNetwork` cleanups in the `finally` block. + +- [ ] **Step 5: Run typecheck** + +```bash +npm run typecheck +``` + +Fix any errors. Common issues will be deleted types (`WorkbenchTraceEntry`) +that are still imported somewhere. + +- [ ] **Step 6: Run existing unit tests to confirm nothing else broke** + +```bash +npm test +``` + +Expected: PASS (existing tests are unit-level and don't depend on the +end-to-end flow we just refactored). + +- [ ] **Step 7: Commit** + +```bash +git add src/workbench/docker-runner.ts +git commit -m "feat(docker-runner): host-side ACP orchestration per trial" +``` + +--- + +### Task 17: Container-runner cleanup — remove `--agent` mode + +**Files:** +- Modify: `src/workbench/container-runner.ts` + +- [ ] **Step 1: Delete the `--agent` mode** + +In `src/workbench/container-runner.ts`, delete: +- `interface AgentRunnerArgs` +- The `--agent` branch of `parseContainerRunnerArgs` +- The `runAgentMode` function +- The branch in `runContainerWorkbenchCase` that dispatches to `runAgentMode` + +Keep: +- `--setup` mode (`runSetupMode`) +- `--grade` mode (`runGradeMode`) +- `writeTraceFile` (still used by grade mode if it derives a trace) + +- [ ] **Step 2: Simplify `parseContainerRunnerArgs`** + +Replace with: + +```typescript +export type ContainerRunnerArgs = GradeRunnerArgs | SetupRunnerArgs; + +export function parseContainerRunnerArgs(args: string[]): ContainerRunnerArgs { + const workDir = getFlagValue(args, '--work'); + const casePath = getFlagValue(args, '--case'); + + if (args.includes('--setup')) { + if (!casePath || !workDir) { + throw new Error('Usage: container-runner --setup --case --work '); + } + return { mode: 'setup', casePath, workDir }; + } + + if (args.includes('--grade')) { + const resultsDir = getFlagValue(args, '--results'); + if (!casePath || !workDir || !resultsDir) { + throw new Error('Usage: container-runner --grade --case --work --results '); + } + return { mode: 'grade', casePath, workDir, resultsDir }; + } + + throw new Error('container-runner: expected --setup or --grade'); +} +``` + +- [ ] **Step 3: Simplify `runContainerWorkbenchCase`** + +```typescript +export async function runContainerWorkbenchCase(args: string[]): Promise { + const parsed = parseContainerRunnerArgs(args); + if (parsed.mode === 'setup') return runSetupMode(parsed); + return runGradeMode(parsed); +} +``` + +- [ ] **Step 4: Delete pi-agent.ts** + +```bash +git rm src/workbench/pi-agent.ts +``` + +- [ ] **Step 5: Build and run typecheck** + +```bash +npm run build && npm run typecheck +``` + +Fix any remaining imports of `pi-agent.ts` or `runAgentMode`. + +- [ ] **Step 6: Run tests** + +```bash +npm test +``` + +- [ ] **Step 7: Commit** + +```bash +git add src/workbench/container-runner.ts src/workbench/pi-agent.ts +git commit -m "refactor(container-runner): drop --agent mode and pi-agent.ts (ACP host-side now)" +``` + +--- + +### Task 18: Update `trials.ts` to aggregate tokens + duration + +**Files:** +- Modify: `src/workbench/trials.ts` +- Modify: `src/workbench/types.ts` +- Create: `tests/trials-aggregate.test.ts` + +- [ ] **Step 1: Add aggregate fields to types** + +In `src/workbench/types.ts`, extend `TrialAggregate` and `WorkbenchModelAggregateResult`: + +```typescript +export interface TrialAggregate { + totalTrials: number; + passedTrials: number; + failedTrials: number; + trialPassRate: number; + meanScore: number; + passAtK: boolean; + passHatK: boolean; + totalTokens: number; // NEW + totalDurationMs: number; // NEW +} +``` + +- [ ] **Step 2: Write the failing test** + +Create `tests/trials-aggregate.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { aggregateTrials } from '../src/workbench/trials.js'; + +test('aggregateTrials sums tokens and duration', () => { + const result = aggregateTrials([ + { trial: 1, pass: true, score: 1, tokens: 100, durationMs: 1000 }, + { trial: 2, pass: false, score: 0, tokens: 200, durationMs: 2000 }, + { trial: 3, pass: true, score: 1, tokens: 150, durationMs: 1500 }, + ]); + assert.equal(result.totalTokens, 450); + assert.equal(result.totalDurationMs, 4500); + assert.equal(result.passedTrials, 2); +}); +``` + +- [ ] **Step 3: Run test to verify it fails** + +```bash +npx tsx --test tests/trials-aggregate.test.ts +``` + +- [ ] **Step 4: Update `aggregateTrials`** + +In `src/workbench/trials.ts`, extend `TrialScoreInput`: + +```typescript +export interface TrialScoreInput { + trial: number; + pass: boolean; + score: number; + tokens?: number; + durationMs?: number; +} +``` + +Update the function body to also sum tokens + duration. Update the +return object. + +Update callers (`run-case.ts`, `run-suite.ts`) to populate +`tokens` and `durationMs` when building `TrialScoreInput` by reading +each trial's `result.json` `metrics.tokens.total` and `metrics.durationMs`. + +- [ ] **Step 5: Verify tests pass and typecheck** + +```bash +npx tsx --test tests/trials-aggregate.test.ts && npm run typecheck && npm test +``` + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/trials.ts src/workbench/types.ts src/workbench/run-case.ts src/workbench/run-suite.ts tests/trials-aggregate.test.ts +git commit -m "feat(trials): aggregate tokens and durationMs across trials" +``` + +--- + +### Task 19: Update `run-case.ts` and `run-suite.ts` for `runs:` matrix + +**Files:** +- Modify: `src/workbench/run-case.ts` +- Modify: `src/workbench/run-suite.ts` + +- [ ] **Step 1: In `run-case.ts`, change matrix dimension from `models` to per-(case, agent, model) jobs** + +Find the `runWorkbenchCaseMatrix` function. Replace the job-building block: + +```typescript +// Before: +// const jobs = params.models.flatMap((model) => Array.from({ length: trials }, (_, index) => ({ model, trial: index + 1 }))); + +// After: +const jobs = params.runs.flatMap((run) => Array.from({ length: trials }, (_, index) => ({ + agent: run.agent, + model: run.model, + trial: index + 1, +}))); +``` + +And in the trial directory naming: + +```typescript +function trialDirName(agent: string, model: string, trial: number): string { + return `${agent}--${slugModelRef(model)}--${formatTrialNumber(trial)}`; +} +``` + +Update all references to use `(agent, model, trial)` instead of +`(model, trial)`. + +- [ ] **Step 2: In `run-suite.ts`, do the same** + +Replace the suite's matrix iteration from `models × cases` to `runs × cases`. + +- [ ] **Step 3: Update CLI in cli-args.ts / cli.ts to drop `--models` flag** + +`run-case --models` flag was the old escape hatch. With `runs:` only, +remove `--models` from the run-case CLI flag list. `run-suite` already +read from `suite.yml`; no CLI change there. + +- [ ] **Step 4: Run typecheck and build** + +```bash +npm run typecheck && npm run build +``` + +Fix any compilation errors. + +- [ ] **Step 5: Run tests** + +```bash +npm test +``` + +Expected: PASS (some unit tests may need updating to use the new shape). + +- [ ] **Step 6: Commit** + +```bash +git add src/workbench/run-case.ts src/workbench/run-suite.ts src/workbench/cli-args.ts src/workbench/cli.ts +git commit -m "feat(run-*): replace models matrix with runs (agent + model) matrix" +``` + +--- + +### Task 20: Migrate example suites to new schema + +**Files:** +- Modify: `examples/workbench/pdf/suite.yml` +- Modify: `examples/workbench/mcp/suite.yml` + +- [ ] **Step 1: Migrate pdf/suite.yml** + +Replace the `models:` block with a `runs:` block, agent set to pi-acp. + +In `examples/workbench/pdf/suite.yml`, change: + +```yaml +models: + - openrouter/google/gemini-2.5-flash +``` + +To: + +```yaml +runs: + - agent: pi-acp + model: openrouter/google/gemini-2.5-flash +``` + +Add `agent: pi-acp` to each inline case definition. Iterate the four +cases in the file (`extract-pdf-facts`, `split-customer-packet`, +`build-briefing-pdf`, `no-pdf-skill-needed`) — add `agent: pi-acp` +right after each `name:` line. + +- [ ] **Step 2: Migrate mcp/suite.yml the same way** + +- [ ] **Step 3: Delete stale result + log dirs** + +```bash +find examples/workbench -type d -name ".results" -exec rm -rf {} + 2>/dev/null +find examples/workbench -name ".run.log" -delete +``` + +- [ ] **Step 4: Run a dry-load of each suite to confirm it parses** + +```bash +npx tsx -e "import('./src/workbench/suite-loader.js').then(m => { console.log(m.loadWorkbenchSuite('examples/workbench/pdf/suite.yml')); })" +npx tsx -e "import('./src/workbench/suite-loader.js').then(m => { console.log(m.loadWorkbenchSuite('examples/workbench/mcp/suite.yml')); })" +``` + +Expected: both print suite objects with `runs:` field populated. + +- [ ] **Step 5: Commit** + +```bash +git add examples/workbench/ +git commit -m "chore(examples): migrate pdf+mcp suites to runs: schema; remove stale results" +``` + +--- + +### Task 21: Per-agent smoke probes + +**Files:** +- Create: `tests/smoke-agents/_common.ts` +- Create: `tests/smoke-agents/claude-agent-acp.smoke.test.ts` +- Create: `tests/smoke-agents/codex-acp.smoke.test.ts` +- Create: `tests/smoke-agents/gemini.smoke.test.ts` +- Create: `tests/smoke-agents/opencode.smoke.test.ts` +- Create: `tests/smoke-agents/pi-acp.smoke.test.ts` + +- [ ] **Step 1: Write a shared smoke harness** + +Create `tests/smoke-agents/_common.ts`: + +```typescript +import { mkdtempSync, writeFileSync, readFileSync, existsSync, rmSync, mkdirSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { runDockerWorkbenchCase } from '../../src/workbench/docker-runner.js'; + +export async function runSmokeTrial(params: { + agent: string; + model: string; + env: string[]; +}): Promise<{ pass: boolean; outputContent?: string; tracePath: string }> { + const dir = mkdtempSync(join(tmpdir(), 'smoke-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'case.yml'), ` +name: smoke-${params.agent} +references: ./references +agent: ${params.agent} +model: ${params.model} +task: | + Write the literal string "hello" to /work/out.txt. + Then write a one-line description of what you did to /work/findings.txt. +graders: + - name: out-txt-exists-with-hello + command: 'test "$(cat /work/out.txt 2>/dev/null)" = "hello"' +env: +${params.env.map(e => ` - ${e}`).join('\n') || ' []'} +timeoutSeconds: 120 +`); + const result = await runDockerWorkbenchCase({ + casePath: join(dir, 'case.yml'), + image: 'skill-optimizer-agent:local', + keepWorkspace: true, + }); + const tracePath = result.tracePath; + let outputContent: string | undefined; + try { + outputContent = readFileSync(join(result.workspacePath!, 'out.txt'), 'utf-8'); + } catch {} + + const resultData = JSON.parse(readFileSync(result.resultPath, 'utf-8')); + rmSync(dir, { recursive: true }); + return { pass: resultData.pass, outputContent, tracePath }; +} +``` + +- [ ] **Step 2: Per-agent test files** + +For each agent, create a file that calls `runSmokeTrial` with the +right model + env. Example for claude: + +Create `tests/smoke-agents/claude-agent-acp.smoke.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { runSmokeTrial } from './_common.js'; + +const skip = !process.env.ANTHROPIC_API_KEY && !require('node:fs').existsSync( + require('node:path').join(require('node:os').homedir(), '.claude/.credentials.json'), +); + +test('claude-agent-acp smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'claude-agent-acp', + model: 'claude-haiku-4-5-20251001', + env: ['ANTHROPIC_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); +``` + +Create equivalents for the other 4 agents, adjusting model and env: +- codex: model `gpt-5-mini`, env `OPENAI_API_KEY`, sub-auth file `~/.codex/auth.json` +- gemini: model `gemini-3.1-pro-preview`, env `GOOGLE_API_KEY`, sub-auth file `~/.gemini/oauth_creds.json` +- opencode: model `google/gemini-3.1-pro-preview`, env `OPENAI_API_KEY` +- pi-acp: model `openrouter/anthropic/claude-haiku-4-5`, env `OPENROUTER_API_KEY` + +- [ ] **Step 3: Run each smoke test** + +```bash +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . +# Run each one (will skip if auth unavailable) +for agent in claude-agent-acp codex-acp gemini opencode pi-acp; do + npx tsx --test tests/smoke-agents/${agent}.smoke.test.ts +done +``` + +Expected: each tests passes or skips. Failures indicate agent-specific +runtime issues (likely registry quirks); fix per-agent and re-run. + +- [ ] **Step 4: Add smoke tests to npm test script** + +In `package.json`, ensure `npm test` includes `tests/smoke-agents/**`. +If smoke tests should be gated behind a flag (e.g., they need Docker +to be running locally), add a `npm run test:smoke` script and call it +from CI only when Docker is present. + +- [ ] **Step 5: Commit** + +```bash +git add tests/smoke-agents/ package.json +git commit -m "test(smoke): per-agent end-to-end smoke probes" +``` + +--- + +### Task 22: Pi-acp regression probe against pdf suite + +**Files:** +- Create: `tests/regression/pi-acp-pdf-suite.test.ts` +- Create: `tests/regression/baselines/pi-acp-pdf-2026-05-25.json` + +- [ ] **Step 1: Capture a baseline run of the pdf suite under pi-acp** + +```bash +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . +OPENROUTER_API_KEY=$OPENROUTER_API_KEY npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials 3 +``` + +This produces `examples/workbench/pdf/.results//suite-result.json`. + +Copy the per-case pass rates into `tests/regression/baselines/pi-acp-pdf-2026-05-25.json`: + +```json +{ + "captured": "2026-05-25", + "agent": "pi-acp", + "model": "openrouter/google/gemini-2.5-flash", + "trials": 3, + "perCasePassRate": { + "extract-pdf-facts": 1.0, + "split-customer-packet": 1.0, + "build-briefing-pdf": 1.0, + "no-pdf-skill-needed": 1.0 + } +} +``` + +(Use whatever rates the actual run produces; rates here are illustrative.) + +- [ ] **Step 2: Write the regression test** + +Create `tests/regression/pi-acp-pdf-suite.test.ts`: + +```typescript +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { readFileSync, readdirSync } from 'node:fs'; +import { join } from 'node:path'; +import { execSync } from 'node:child_process'; + +const skip = !process.env.OPENROUTER_API_KEY; + +test('pi-acp regression: pdf suite pass rates match baseline within tolerance', { skip }, () => { + const baseline = JSON.parse(readFileSync('tests/regression/baselines/pi-acp-pdf-2026-05-25.json', 'utf-8')); + + execSync(`npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials ${baseline.trials}`, { + stdio: 'inherit', + }); + + const resultsDir = 'examples/workbench/pdf/.results'; + const runs = readdirSync(resultsDir).sort(); + const latest = runs[runs.length - 1]; + const data = JSON.parse(readFileSync(join(resultsDir, latest, 'suite-result.json'), 'utf-8')); + + for (const [caseName, expectedRate] of Object.entries(baseline.perCasePassRate)) { + const actual = data.results.find((r: any) => r.caseName === caseName)?.trialPassRate; + assert.ok(actual !== undefined, `case ${caseName} not in suite result`); + // Tolerance: ±0.34 (one trial of three can flip) + assert.ok(Math.abs(actual - (expectedRate as number)) <= 0.34, + `case ${caseName}: expected ~${expectedRate}, got ${actual}`); + } +}); +``` + +- [ ] **Step 3: Run the regression test once to confirm it passes against the captured baseline** + +```bash +OPENROUTER_API_KEY=$OPENROUTER_API_KEY npx tsx --test tests/regression/pi-acp-pdf-suite.test.ts +``` + +Expected: PASS + +- [ ] **Step 4: Commit** + +```bash +git add tests/regression/ +git commit -m "test(regression): pi-acp pdf suite baseline + tolerance check" +``` + +--- + +### Task 23: Write `skills/shared/acp-trace-format.md` + +**Files:** +- Create: `skills/shared/acp-trace-format.md` + +- [ ] **Step 1: Write the reference doc** + +Create `skills/shared/acp-trace-format.md`: + +```markdown +# ACP trace format (for chain analyzer) + +`trace.jsonl` captures the raw Agent Client Protocol (ACP) messages +exchanged between the workbench (client) and the agent CLI (server) +for one trial. The first line is a `trace_start` header with trial +metadata; every subsequent line is a JSON-RPC envelope per the ACP +spec at . + +This doc summarizes the message types the chain analyzer cares about. +For the full spec, follow the link above. + +## Header (line 1) + +```json +{ + "type": "trace_start", + "schemaVersion": 2, + "caseName": "...", + "agent": "claude-agent-acp", + "model": "claude-haiku-4-5-20251001", + "startedAt": "ISO-8601", + "endedAt": "ISO-8601" +} +``` + +## Initialization (early lines) + +- `initialize` request from client; `initialize` response from agent +- `session/new` request; response carries `sessionId` + +These confirm the agent started. If they're absent, the trial failed +before reaching the prompt — bench infrastructure issue, not skill +weakness. + +## Session updates (the meat) + +All have shape: + +```json +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"...","update":{...}}} +``` + +The `update` object's `sessionUpdate` field tells you what happened: + +| `sessionUpdate` value | Meaning | Analyzer cares because | +|---|---|---| +| `agent_message_chunk` | Streaming chunk of assistant text | Final response — what the agent told the user | +| `agent_thought_chunk` | Streaming chunk of assistant reasoning | The agent's reasoning — useful for diagnosing why it did X | +| `tool_call` | Agent is calling a tool (start) | `kind` field tells you what kind: `execute`, `read`, `write`, `edit`, `search`, `fetch`, `think`, `other` | +| `tool_call_update` | Tool result/progress (end) | `status: "completed" \| "failed"`, `content` carries the tool output | +| `plan` | Agent's high-level plan | Optional sidebar; skip in most analysis | +| `user_message_chunk` | Agent echoing user input | Rare; skip | + +## Final prompt response (last numbered response) + +```json +{ + "jsonrpc": "2.0", + "id": , + "result": { + "stopReason": "end_turn" | "max_tokens" | "refusal" | "cancelled", + "usage": { + "inputTokens": 120, + "outputTokens": 45, + "cacheReadTokens": 0, + "cacheCreationTokens": 0 + } + } +} +``` + +`stopReason` is critical for diagnosing failures: +- `end_turn` — normal completion (still check `findings.txt` for correctness) +- `max_tokens` — agent ran out of context; skill may be too verbose +- `refusal` — agent declined the task; skill description may have triggered a safety pattern +- `cancelled` — workbench timed out the prompt + +## Helpers + +Don't parse the JSONL manually. Use `src/workbench/parse-trace.ts`: + +- `iterMessages(jsonl)` — assistant/user messages (text + thinking) +- `iterToolCalls(jsonl)` — paired tool_call + tool_call_update +- `computeMetrics(jsonl)` — tokens, duration, per-tool counts, stopReason +- `getFinalAssistantMessage(jsonl)` — last assistant chunk concatenated +- `getFailureEvidence(jsonl)` — failed-tool result snippets +``` + +- [ ] **Step 2: Lint** + +```bash +pnpm dlx markdownlint-cli --fix --disable MD013 MD031 MD032 MD033 MD040 MD041 MD060 -- skills/shared/acp-trace-format.md +``` + +- [ ] **Step 3: Commit** + +```bash +git add skills/shared/acp-trace-format.md +git commit -m "docs(shared): ACP trace format reference for chain analyzer" +``` + +--- + +### Task 24: Update chain skill files for new trace format + agent column + +**Files:** +- Modify: `skills/write-tests/agents/test-writer.md` +- Modify: `skills/analyze/agents/analyzer.md` +- Modify: `skills/run-bench/SKILL.md` + +- [ ] **Step 1: Update test-writer.md to drop `/work/` skill-path boilerplate** + +In `skills/write-tests/agents/test-writer.md`, do the following exact +substitutions: + +1. Grep for occurrences: + +```bash +grep -n "work/\|skill at\|SKILL\.md\|/work" skills/write-tests/agents/test-writer.md +``` + +2. For each match that instructs the test-writer to make task prompts +reference the skill explicitly (e.g., "the task prompt should say +'use the skill at /work/X/SKILL.md'"), replace with the realistic- +invocation guidance below. Add the spec.yaml convention pointer. + +3. Add a new paragraph near the section that defines the probe spec's +`task:` field: + +```markdown +**Task prompts describe the user's actual task; do not reference the +skill explicitly.** The harness mounts the skill at the agent's native +discovery path (e.g., `~/.claude/skills//SKILL.md`). The agent +decides whether to invoke it based on the skill's frontmatter +`description`. "Skill didn't trigger" is then a measurable weakness +class — don't pre-trigger it via the task prompt. + +Example task prompt (good): "Review /work/ProductCard.tsx for +compliance issues. Write findings to /work/findings.txt." + +Example task prompt (bad — pre-triggers): "Use the skill at +/work/web-design-guidelines/SKILL.md to review /work/ProductCard.tsx." +``` + +- [ ] **Step 2: Update analyzer.md to point at acp-trace-format.md** + +In `skills/analyze/agents/analyzer.md`, find any reference to +`WorkbenchTraceEntry` or the normalized trace format. Replace with: + +```markdown +**Read first:** [`../../shared/acp-trace-format.md`](../../shared/acp-trace-format.md) +— `trace.jsonl` is raw ACP wire format. Use `parse-trace.ts` helpers +(`iterMessages`, `iterToolCalls`, `computeMetrics`, +`getFinalAssistantMessage`, `getFailureEvidence`) rather than reading +JSONL by hand. The format reference doc summarizes the event types +relevant to skill-behavior analysis. +``` + +- [ ] **Step 3: Update run-bench/SKILL.md** + +In the "Write the summary" section, expand the requirements: + +```markdown +### (d) Write the summary + +Parse `${OUT_DIR}/suite-result.json` and write `06-bench-summary.md` +with the frontmatter above plus a body containing: + +- **Overall:** total trials, passed, failed, overall pass rate. Total + tokens consumed. Total duration. +- **Per agent + model:** trial count, pass rate, mean tokens per + trial, mean duration per trial for each row in `runs:` from the + suite. +- **Per probe:** trial count, pass rate, one-line note if any + trial failed. Mean tokens and duration per probe. +- **Failed-probe pointer list:** probe IDs where any trial failed, + with paths to their `trace.jsonl` and `findings.txt`. +- **Raw output:** the `bench_results_path` value. + +Do NOT include a cost column. tokens × downstream pricing is computed +offline if needed. +``` + +- [ ] **Step 4: Lint all three** + +```bash +for f in skills/write-tests/agents/test-writer.md skills/analyze/agents/analyzer.md skills/run-bench/SKILL.md; do + pnpm dlx markdownlint-cli --fix --disable MD013 MD031 MD032 MD033 MD040 MD041 MD060 -- "$f" +done +``` + +- [ ] **Step 5: Commit** + +```bash +git add skills/ +git commit -m "feat(chain): update test-writer, analyzer, run-bench for ACP trace format + agent column" +``` + +--- + +### Task 25: Documentation refresh + +**Files:** +- Modify: `CLAUDE.md` +- Modify: `CONTRIBUTING.md` +- Modify: `README.md` + +- [ ] **Step 1: Update CLAUDE.md** + +In `CLAUDE.md`: +- Update the "Key Commands" section: add `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .` +- Update the "Important Files" section: add `src/workbench/acp/`, `src/workbench/agents/`, `src/workbench/parse-trace.ts`; remove `src/workbench/pi-agent.ts` +- Update "Invariants": replace "uses models from `suite.yml`" with "uses runs (agent + model) from `suite.yml`"; add "Every case.yml must declare `agent:`" +- Update "Testing Guidance": replace `docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile .` with the new image/Dockerfile + +- [ ] **Step 2: Update CONTRIBUTING.md** + +Same set of replacements. Add a section noting the ACP architecture +(host-side client, per-trial container). + +- [ ] **Step 3: Update README.md** + +In the README, replace any "set OPENROUTER_API_KEY" prereq with a +matrix of supported agents and their auth options: + +```markdown +## Supported agents + +| Agent | Auth options | +|---|---| +| claude-agent-acp | `claude login` (subscription) or `ANTHROPIC_API_KEY` | +| codex-acp | `codex login` (subscription) or `OPENAI_API_KEY` | +| gemini | `gemini auth login` (subscription) or `GOOGLE_API_KEY` | +| opencode | `OPENAI_API_KEY` (or provider-specific key) | +| pi-acp | `OPENROUTER_API_KEY` | +``` + +- [ ] **Step 4: Lint** + +```bash +for f in CLAUDE.md CONTRIBUTING.md README.md; do + pnpm dlx markdownlint-cli --fix --disable MD013 MD031 MD032 MD033 MD040 MD041 MD060 -- "$f" +done +``` + +- [ ] **Step 5: Commit** + +```bash +git add CLAUDE.md CONTRIBUTING.md README.md +git commit -m "docs: refresh CLAUDE.md, CONTRIBUTING.md, README.md for multi-agent ACP" +``` + +--- + +### Task 26: Final integration check + +- [ ] **Step 1: Full typecheck + build + tests** + +```bash +npm run typecheck && npm run build && npm test +``` + +Expected: all PASS + +- [ ] **Step 2: Run smoke distribution test** + +```bash +npx tsx tests/smoke-skill-distribution.ts +``` + +Expected: PASS (verifies plugin metadata still references all chain +skills correctly) + +- [ ] **Step 3: Re-build the Docker image to confirm reproducibility** + +```bash +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . +``` + +- [ ] **Step 4: Run one end-to-end probe** + +```bash +OPENROUTER_API_KEY=$OPENROUTER_API_KEY npx tsx src/cli.ts run-case examples/workbench/pdf/suite.yml +``` + +Wait, run-case takes a case path. For an end-to-end check, run-suite: + +```bash +OPENROUTER_API_KEY=$OPENROUTER_API_KEY npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials 1 +``` + +Expected: completes with at least one trial reported, results in +`examples/workbench/pdf/.results//`. Inspect the result: + +```bash +ls examples/workbench/pdf/.results/$(ls -t examples/workbench/pdf/.results | head -1)/ +cat examples/workbench/pdf/.results/$(ls -t examples/workbench/pdf/.results | head -1)/suite-result.json | jq '.summary' +``` + +Confirm `summary.trialPassRate` is sensible. + +- [ ] **Step 5: Inspect a trace to confirm raw ACP format** + +```bash +cat examples/workbench/pdf/.results/$(ls -t examples/workbench/pdf/.results | head -1)/trials/*/trace.jsonl | head -5 | jq . +``` + +Expected: first line is a `trace_start` header; following lines have +`jsonrpc: "2.0"` envelopes. + +- [ ] **Step 6: Commit any final fixes if needed** + +If steps 1-5 surfaced bugs, fix them in small targeted commits. + +--- + +## Verification checklist (run after all tasks) + +- [ ] `npm run typecheck` — clean +- [ ] `npm run build` — clean +- [ ] `npm test` — all pass +- [ ] `npx tsx tests/smoke-skill-distribution.ts` — passes +- [ ] `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .` — succeeds +- [ ] Each of the 5 per-agent smoke tests passes (or skips gracefully) +- [ ] Pi-acp regression test against pdf suite passes within tolerance +- [ ] One end-to-end `run-suite` call against the pdf suite produces sensible results with raw ACP trace From 64cf9487079275a2eab0265e20e3093c46dd08f0 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:24:02 -0500 Subject: [PATCH 066/121] feat(acp): add @agentclientprotocol/sdk dependency --- package-lock.json | 11 ++++++++++- package.json | 1 + tests/acp/sdk-smoke.test.ts | 9 +++++++++ 3 files changed, 20 insertions(+), 1 deletion(-) create mode 100644 tests/acp/sdk-smoke.test.ts diff --git a/package-lock.json b/package-lock.json index d985d28..91fb901 100644 --- a/package-lock.json +++ b/package-lock.json @@ -8,8 +8,8 @@ "name": "skill-optimizer", "version": "2.0.0", "license": "MIT", - "main": ".opencode/plugins/skill-optimizer.js", "dependencies": { + "@agentclientprotocol/sdk": "0.22.1", "@mariozechner/pi-agent-core": "^0.66.1", "@mariozechner/pi-ai": "^0.66.1", "@mariozechner/pi-coding-agent": "^0.66.1", @@ -29,6 +29,15 @@ "node": ">=20" } }, + "node_modules/@agentclientprotocol/sdk": { + "version": "0.22.1", + "resolved": "https://registry.npmjs.org/@agentclientprotocol/sdk/-/sdk-0.22.1.tgz", + "integrity": "sha512-DfqXtl/8gO9NImq094MTaCXEU2vkhh6v7q/kT+9UjZxUqj8hYaya2OjLVIqn16MzNHcXEpShTR2RIauLSYeDQQ==", + "license": "Apache-2.0", + "peerDependencies": { + "zod": "^3.25.0 || ^4.0.0" + } + }, "node_modules/@anthropic-ai/sdk": { "version": "0.73.0", "resolved": "https://registry.npmjs.org/@anthropic-ai/sdk/-/sdk-0.73.0.tgz", diff --git a/package.json b/package.json index 1f1be4c..adc8fd7 100644 --- a/package.json +++ b/package.json @@ -82,6 +82,7 @@ "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts" }, "dependencies": { + "@agentclientprotocol/sdk": "0.22.1", "@mariozechner/pi-agent-core": "^0.66.1", "@mariozechner/pi-ai": "^0.66.1", "@mariozechner/pi-coding-agent": "^0.66.1", diff --git a/tests/acp/sdk-smoke.test.ts b/tests/acp/sdk-smoke.test.ts new file mode 100644 index 0000000..e48a70d --- /dev/null +++ b/tests/acp/sdk-smoke.test.ts @@ -0,0 +1,9 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; + +test('@agentclientprotocol/sdk exports ClientSideConnection', async () => { + const mod = await import('@agentclientprotocol/sdk'); + assert.equal(typeof mod.ClientSideConnection, 'function'); + assert.equal(typeof mod.ndJsonStream, 'function'); + assert.equal(typeof mod.RequestError, 'function'); +}); From c7f93bc10aa0664651295eae3f8643a5fe9dc572 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:25:52 -0500 Subject: [PATCH 067/121] chore(test): wire ACP smoke test into npm test script --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index adc8fd7..6f636f6 100644 --- a/package.json +++ b/package.json @@ -79,7 +79,7 @@ "build": "tsc && chmod +x dist/cli.js", "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 0260339b8ee2debdd9583ad54b76ae36969b4bd0 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:27:51 -0500 Subject: [PATCH 068/121] feat(acp): docker exec stdio bridge implementing Stream --- src/workbench/acp/transport.ts | 50 ++++++++++++++++++++++++++++++++++ tests/acp/transport.test.ts | 22 +++++++++++++++ 2 files changed, 72 insertions(+) create mode 100644 src/workbench/acp/transport.ts create mode 100644 tests/acp/transport.test.ts diff --git a/src/workbench/acp/transport.ts b/src/workbench/acp/transport.ts new file mode 100644 index 0000000..1e1ff3e --- /dev/null +++ b/src/workbench/acp/transport.ts @@ -0,0 +1,50 @@ +import type { ChildProcessWithoutNullStreams } from 'node:child_process'; + +/** + * Raw byte transport for ACP over docker exec stdio. + * + * Exposes `outgoing` (bytes to write to the subprocess stdin) and + * `incoming` (bytes read from subprocess stdout) so that the ACP client + * (Task 3) can wrap them with `ndJsonStream` from `@agentclientprotocol/sdk`. + */ +export interface DockerExecStream { + outgoing: WritableStream; + incoming: ReadableStream; + close(): Promise; +} + +export function createDockerExecStream( + child: ChildProcessWithoutNullStreams, +): DockerExecStream { + const outgoing = new WritableStream({ + write(chunk) { + return new Promise((resolve, reject) => { + child.stdin.write(chunk, (err) => (err ? reject(err) : resolve())); + }); + }, + close() { + child.stdin.end(); + }, + }); + + const incoming = new ReadableStream({ + start(controller) { + child.stdout.on('data', (chunk: Buffer) => controller.enqueue(chunk)); + child.stdout.on('end', () => controller.close()); + child.stdout.on('error', (err) => controller.error(err)); + }, + }); + + return { + outgoing, + incoming, + async close() { + try { + child.stdin.end(); + } catch { + // ignore + } + child.kill(); + }, + }; +} diff --git a/tests/acp/transport.test.ts b/tests/acp/transport.test.ts new file mode 100644 index 0000000..7f09b23 --- /dev/null +++ b/tests/acp/transport.test.ts @@ -0,0 +1,22 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { spawn } from 'node:child_process'; +import { createDockerExecStream } from '../../src/workbench/acp/transport.js'; + +test('createDockerExecStream produces a Stream that round-trips ndjson', async () => { + // Use `cat` as a stand-in for the agent: it echoes stdin to stdout. + const child = spawn('cat', [], { stdio: ['pipe', 'pipe', 'pipe'] }); + const stream = createDockerExecStream(child); + + const message = { jsonrpc: '2.0', id: 1, method: 'initialize', params: {} }; + const writer = stream.outgoing.getWriter(); + await writer.write(new TextEncoder().encode(JSON.stringify(message) + '\n')); + writer.releaseLock(); + + const reader = stream.incoming.getReader(); + const { value } = await reader.read(); + const echoed = new TextDecoder().decode(value).trim(); + assert.equal(echoed, JSON.stringify(message)); + + child.kill(); +}); From ad388b42eb7743801e46e2aa86fcf641b0c9dd0e Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:28:12 -0500 Subject: [PATCH 069/121] chore(test): wire ACP transport test into npm test --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 6f636f6..586561e 100644 --- a/package.json +++ b/package.json @@ -79,7 +79,7 @@ "build": "tsc && chmod +x dist/cli.js", "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 9aeef030f40cef32d865bae7b4b12a7108b565d3 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:29:42 -0500 Subject: [PATCH 070/121] fix(acp): make transport outgoing close() return a Promise --- src/workbench/acp/transport.ts | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/src/workbench/acp/transport.ts b/src/workbench/acp/transport.ts index 1e1ff3e..f99e0ff 100644 --- a/src/workbench/acp/transport.ts +++ b/src/workbench/acp/transport.ts @@ -23,7 +23,9 @@ export function createDockerExecStream( }); }, close() { - child.stdin.end(); + return new Promise((resolve, reject) => { + child.stdin.end((err?: Error | null) => (err ? reject(err) : resolve())); + }); }, }); From 7be33416e14f3ce4784071b10ad26d683736f016 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:31:50 -0500 Subject: [PATCH 071/121] feat(acp): client wrapper with auto-approve permission handler Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/acp/client.ts | 49 +++++++++++++++++++++++++++++++++++++ tests/acp/client.test.ts | 19 ++++++++++++++ 2 files changed, 68 insertions(+) create mode 100644 src/workbench/acp/client.ts create mode 100644 tests/acp/client.test.ts diff --git a/src/workbench/acp/client.ts b/src/workbench/acp/client.ts new file mode 100644 index 0000000..d807617 --- /dev/null +++ b/src/workbench/acp/client.ts @@ -0,0 +1,49 @@ +import { ClientSideConnection, ndJsonStream } from '@agentclientprotocol/sdk'; +import type { + Client, + SessionNotification, + RequestPermissionRequest, + RequestPermissionResponse, +} from '@agentclientprotocol/sdk'; +import type { DockerExecStream } from './transport.js'; + +export interface WorkbenchClientOptions { + stream: DockerExecStream; + onSessionUpdate: (notification: SessionNotification) => void; +} + +export interface WorkbenchClient { + connection: ClientSideConnection; + close(): Promise; +} + +export function createWorkbenchClient(opts: WorkbenchClientOptions): WorkbenchClient { + // Note: ndJsonStream wraps the raw byte streams as JSON-RPC envelopes. + const ndjson = ndJsonStream(opts.stream.outgoing, opts.stream.incoming); + + const connection = new ClientSideConnection( + (_agent): Client => ({ + sessionUpdate: async (params: SessionNotification) => { + opts.onSessionUpdate(params); + }, + // Auto-approve all tool calls. The container is the security boundary, + // not the permission gate. Approving everything matches the headless- + // bench model and matches benchflow's pattern. + requestPermission: async ( + params: RequestPermissionRequest, + ): Promise => ({ + outcome: { outcome: 'selected', optionId: params.options[0]?.optionId ?? 'allow' }, + }), + // We don't expose filesystem capabilities to the agent over ACP; the + // agent uses its own tools to touch /work directly. + }), + ndjson, + ); + + return { + connection, + async close() { + await opts.stream.close(); + }, + }; +} diff --git a/tests/acp/client.test.ts b/tests/acp/client.test.ts new file mode 100644 index 0000000..a4ad178 --- /dev/null +++ b/tests/acp/client.test.ts @@ -0,0 +1,19 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { spawn } from 'node:child_process'; +import { createWorkbenchClient } from '../../src/workbench/acp/client.js'; +import { createDockerExecStream } from '../../src/workbench/acp/transport.js'; + +test('createWorkbenchClient returns object with initialize/newSession/prompt/cancel', () => { + const child = spawn('cat', []); + const stream = createDockerExecStream(child); + const client = createWorkbenchClient({ stream, onSessionUpdate: () => {} }); + + assert.equal(typeof client.connection.initialize, 'function'); + assert.equal(typeof client.connection.newSession, 'function'); + assert.equal(typeof client.connection.prompt, 'function'); + assert.equal(typeof client.connection.cancel, 'function'); + assert.equal(typeof client.close, 'function'); + + child.kill(); +}); From c1c965d7613717e2f536d508fa2fe2beda3f4b3e Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:32:09 -0500 Subject: [PATCH 072/121] chore(test): wire ACP client test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 586561e..a49cd4c 100644 --- a/package.json +++ b/package.json @@ -79,7 +79,7 @@ "build": "tsc && chmod +x dist/cli.js", "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 2732c6083674fb5c91e98ec7f77c9442e4761613 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:35:14 -0500 Subject: [PATCH 073/121] feat(agents): registry types + resolver (entries pending in task 5) Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/agents/registry.ts | 87 ++++++++++++++++++++++++++++++++ tests/agents/registry.test.ts | 39 ++++++++++++++ 2 files changed, 126 insertions(+) create mode 100644 src/workbench/agents/registry.ts create mode 100644 tests/agents/registry.test.ts diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts new file mode 100644 index 0000000..a1dfaa5 --- /dev/null +++ b/src/workbench/agents/registry.ts @@ -0,0 +1,87 @@ +export interface CredentialFile { + path: string; // Target path in container; may use {home} + envSource: string; // Env var on host to read value from + template?: string; // If set, value is inserted into template at {value} + mkdir?: boolean; // Create parent dir; default true +} + +export interface HostAuthFile { + hostPath: string; // ~/.claude/.credentials.json + containerPath: string; // {home}/.claude/.credentials.json +} + +export interface SubscriptionAuth { + replacesEnv: string; // e.g. "ANTHROPIC_API_KEY" + detectFile: string; // host path to check for login + files: HostAuthFile[]; // all files to copy when sub-auth used +} + +export type ApiProtocol = + | 'anthropic-messages' + | 'openai-completions' + | 'openai-responses' + | ''; + +export type AcpModelFormat = 'bare' | 'provider/model'; + +export interface AgentConfig { + name: string; + description: string; + installCmd: string; // bash, runs at IMAGE BUILD time + launchCmd: string; // bash, runs per-trial via docker exec + requiresEnv: string[]; + apiProtocol: ApiProtocol; + envMapping: Record; // SKILL_OPT_PROVIDER_* → agent-native + skillPaths: string[]; // e.g. ["$HOME/.claude/skills"] + credentialFiles: CredentialFile[]; + homeDirs: string[]; + subscriptionAuth: SubscriptionAuth | null; + acpModelFormat: AcpModelFormat; + supportsAcpSetModel: boolean; + loginHint?: string; // shown in fail-loud message +} + +export const AGENTS: Record = { + // Entries filled in Task 5. Empty here to make this task self-contained. +}; + +export const AGENT_ALIASES: Record = { + claude: 'claude-agent-acp', + codex: 'codex-acp', + gemini: 'gemini', + pi: 'pi-acp', + openclaw: 'openclaw', +}; + +export function resolveAgent(spec: string): AgentConfig { + const canonical = AGENT_ALIASES[spec] ?? spec; + const cfg = AGENTS[canonical]; + if (cfg) return cfg; + + const known = Object.keys(AGENTS); + const close = closestMatch(canonical, known); + if (close) { + throw new Error(`Unknown agent: ${spec}. Did you mean: ${close}?`); + } + throw new Error(`Unknown agent: ${spec}. Available: ${known.join(', ')}`); +} + +function closestMatch(needle: string, haystack: string[]): string | undefined { + let best: { name: string; score: number } | undefined; + for (const name of haystack) { + const score = sharedPrefix(needle, name) + sharedPrefix( + needle.split('').reverse().join(''), + name.split('').reverse().join(''), + ); + if (score > 4 && (!best || score > best.score)) { + best = { name, score }; + } + } + return best?.name; +} + +function sharedPrefix(a: string, b: string): number { + let i = 0; + while (i < a.length && i < b.length && a[i] === b[i]) i++; + return i; +} diff --git a/tests/agents/registry.test.ts b/tests/agents/registry.test.ts new file mode 100644 index 0000000..be3ad79 --- /dev/null +++ b/tests/agents/registry.test.ts @@ -0,0 +1,39 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { resolveAgent, AGENTS, AGENT_ALIASES } from '../../src/workbench/agents/registry.js'; + +test('AGENTS contains all 5 expected entries', () => { + for (const name of ['claude-agent-acp', 'codex-acp', 'gemini', 'opencode', 'pi-acp']) { + assert.ok(AGENTS[name], `missing agent: ${name}`); + assert.equal(AGENTS[name].name, name); + } +}); + +test('resolveAgent accepts aliases', () => { + assert.equal(resolveAgent('claude').name, 'claude-agent-acp'); + assert.equal(resolveAgent('codex').name, 'codex-acp'); + assert.equal(resolveAgent('pi').name, 'pi-acp'); +}); + +test('resolveAgent accepts canonical names', () => { + assert.equal(resolveAgent('claude-agent-acp').name, 'claude-agent-acp'); +}); + +test('resolveAgent throws with suggestion on unknown', () => { + assert.throws(() => resolveAgent('claud-agent'), /Did you mean/); +}); + +test('every agent declares at least one skillPath', () => { + for (const cfg of Object.values(AGENTS)) { + assert.ok(cfg.skillPaths.length > 0, `${cfg.name} missing skillPaths`); + assert.match(cfg.skillPaths[0], /^\$HOME\//); + } +}); + +test('AGENT_ALIASES does not collide with canonical names', () => { + for (const alias of Object.keys(AGENT_ALIASES)) { + if (AGENTS[alias]) { + assert.equal(alias, AGENT_ALIASES[alias], `alias ${alias} collides`); + } + } +}); From 7d2c79db3a701b10711d552b31d889621b4b2ec9 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:36:55 -0500 Subject: [PATCH 074/121] feat(agents): populate 5-agent registry (claude/codex/gemini/opencode/pi-acp) --- src/workbench/agents/registry.ts | 119 ++++++++++++++++++++++++++++++- 1 file changed, 118 insertions(+), 1 deletion(-) diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts index a1dfaa5..b33a292 100644 --- a/src/workbench/agents/registry.ts +++ b/src/workbench/agents/registry.ts @@ -42,7 +42,124 @@ export interface AgentConfig { } export const AGENTS: Record = { - // Entries filled in Task 5. Empty here to make this task self-contained. + 'claude-agent-acp': { + name: 'claude-agent-acp', + description: 'Claude Code via ACP (Anthropic CLI)', + installCmd: `npm install -g @zed-industries/claude-agent-acp@latest`, + launchCmd: `claude-agent-acp`, + requiresEnv: ['ANTHROPIC_API_KEY'], + apiProtocol: 'anthropic-messages', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'ANTHROPIC_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'ANTHROPIC_AUTH_TOKEN', + SKILL_OPT_PROVIDER_MODEL: 'ANTHROPIC_MODEL', + }, + skillPaths: ['$HOME/.claude/skills'], + credentialFiles: [], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'ANTHROPIC_API_KEY', + detectFile: '~/.claude/.credentials.json', + files: [ + { hostPath: '~/.claude/.credentials.json', containerPath: '{home}/.claude/.credentials.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'claude login', + }, + 'codex-acp': { + name: 'codex-acp', + description: 'OpenAI Codex via ACP', + installCmd: `npm install -g @zed-industries/codex-acp@latest`, + launchCmd: `codex-acp \${OPENAI_BASE_URL:+-c openai_base_url=$OPENAI_BASE_URL}`, + requiresEnv: ['OPENAI_API_KEY'], + apiProtocol: 'openai-responses', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'OPENAI_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'OPENAI_API_KEY', + }, + skillPaths: ['$HOME/.agents/skills'], + credentialFiles: [ + { + path: '{home}/.codex/auth.json', + envSource: 'OPENAI_API_KEY', + template: '{"OPENAI_API_KEY": "{value}"}', + }, + ], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'OPENAI_API_KEY', + detectFile: '~/.codex/auth.json', + files: [ + { hostPath: '~/.codex/auth.json', containerPath: '{home}/.codex/auth.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'codex login', + }, + 'gemini': { + name: 'gemini', + description: 'Google Gemini CLI via ACP', + installCmd: `npm install -g @google/gemini-cli@latest`, + launchCmd: `gemini --acp --yolo`, + requiresEnv: ['GOOGLE_API_KEY'], + apiProtocol: '', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'GEMINI_API_BASE_URL', + SKILL_OPT_PROVIDER_API_KEY: 'GOOGLE_API_KEY', + }, + skillPaths: ['$HOME/.gemini/skills'], + credentialFiles: [], + homeDirs: [], + subscriptionAuth: { + replacesEnv: 'GEMINI_API_KEY', + detectFile: '~/.gemini/oauth_creds.json', + files: [ + { hostPath: '~/.gemini/oauth_creds.json', containerPath: '{home}/.gemini/oauth_creds.json' }, + { hostPath: '~/.gemini/settings.json', containerPath: '{home}/.gemini/settings.json' }, + { hostPath: '~/.gemini/google_accounts.json', containerPath: '{home}/.gemini/google_accounts.json' }, + ], + }, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'gemini auth login', + }, + 'opencode': { + name: 'opencode', + description: 'OpenCode via ACP — open-source coding agent', + installCmd: `npm install -g opencode-ai@latest`, + launchCmd: `opencode acp`, + requiresEnv: [], + apiProtocol: '', + envMapping: { + SKILL_OPT_PROVIDER_BASE_URL: 'OPENAI_BASE_URL', + }, + skillPaths: ['$HOME/.opencode/skills'], + credentialFiles: [], + homeDirs: ['.opencode'], + subscriptionAuth: null, + acpModelFormat: 'provider/model', + supportsAcpSetModel: true, + loginHint: 'set OPENAI_API_KEY (or provider-specific key)', + }, + 'pi-acp': { + name: 'pi-acp', + description: 'Pi coding agent via ACP', + installCmd: `npm install -g @mariozechner/pi-coding-agent@latest pi-acp@latest`, + launchCmd: `/opt/skill-opt/bin/pi-acp-launcher`, + requiresEnv: [], + apiProtocol: '', + envMapping: {}, + skillPaths: ['$HOME/.pi/agent/skills', '$HOME/.agents/skills'], + credentialFiles: [], + homeDirs: ['.pi'], + subscriptionAuth: null, + acpModelFormat: 'bare', + supportsAcpSetModel: true, + loginHint: 'set OPENROUTER_API_KEY', + }, }; export const AGENT_ALIASES: Record = { From 35f2060d4a2b56b92b75186f79b63091d7bed423 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:37:18 -0500 Subject: [PATCH 075/121] chore(test): wire registry test into npm test --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index a49cd4c..250bbad 100644 --- a/package.json +++ b/package.json @@ -79,7 +79,7 @@ "build": "tsc && chmod +x dist/cli.js", "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 13bef0b4c70421b63999f6b076861bf18e485380 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:41:09 -0500 Subject: [PATCH 076/121] feat(docker): new image with all 5 agents pre-baked Replace workbench-runner.Dockerfile with skill-optimizer-agent.Dockerfile, which installs claude-agent-acp, codex-acp, gemini-cli, opencode-ai, and pi-acp/pi-coding-agent in a single cached layer. Add pi-acp-launcher.sh to bridge SKILL_OPT_PROVIDER_API_KEY to OPENROUTER_API_KEY. Add install-snippets.ts and a dockerfile:print-installs script to surface the registry-driven install block. Update the image-name constant and all doc references from skill-optimizer-workbench:local to skill-optimizer-agent:local. Co-Authored-By: Claude Sonnet 4.6 --- CLAUDE.md | 4 +- CONTRIBUTING.md | 4 +- README.md | 2 +- docker/pi-acp-launcher.sh | 11 +++++ docker/skill-optimizer-agent.Dockerfile | 60 ++++++++++++++++++++++++ docker/workbench-runner.Dockerfile | 41 ---------------- package.json | 1 + src/workbench/agents/install-snippets.ts | 15 ++++++ src/workbench/docker-runner.ts | 2 +- 9 files changed, 93 insertions(+), 47 deletions(-) create mode 100755 docker/pi-acp-launcher.sh create mode 100644 docker/skill-optimizer-agent.Dockerfile delete mode 100644 docker/workbench-runner.Dockerfile create mode 100644 src/workbench/agents/install-snippets.ts diff --git a/CLAUDE.md b/CLAUDE.md index 47c2ec8..42a061f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -21,7 +21,7 @@ npx tsx src/cli.ts run-suite --help - `src/cli.ts`: public CLI entrypoint - `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces -- `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases +- `docker/skill-optimizer-agent.Dockerfile`: container image with all 5 agent CLIs pre-baked, used for setup and grade phases - `skills//`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) - `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills - `skills//agents/`: prompt templates dispatched by chain skills via the Agent tool @@ -60,6 +60,6 @@ Keep the README installation section aligned with packaged plugin metadata: - Run `npm run typecheck` after TypeScript changes. - Run `npm test` before finishing behavior changes. -- For Docker runner or image changes, also run `docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile .`. +- For Docker runner or image changes, also run `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .`. - For CLI/docs changes, verify `npx tsx src/cli.ts --help` if touched docs mention CLI behavior. - For plugin/package metadata changes, run `npx tsx tests/smoke-skill-distribution.ts` and verify `npm pack --dry-run --json` includes required plugin files without result/cache directories. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index e080d76..a706ece 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -23,7 +23,7 @@ All three commands must pass before opening a PR when code changes are involved. - `src/cli.ts` — public CLI entry point for `run-case` and `run-suite`. - `src/workbench/` — case/suite loading, Docker runner, Pi agent wiring, graders, traces, metrics, MCP support, and trial aggregation. -- `docker/workbench-runner.Dockerfile` — non-root container image for setup, agent, grade, and cleanup phases. +- `docker/skill-optimizer-agent.Dockerfile` — container image with all 5 agent CLIs pre-baked, used for setup and grade phases. - `skills//` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). - `skills/shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. - `skills//agents/` — prompt templates dispatched by chain skills via the Agent tool. @@ -53,7 +53,7 @@ All three commands must pass before opening a PR when code changes are involved. - Run `npm run typecheck` after TypeScript changes. - Run `npm test` before finishing behavior changes. -- For Docker runner or image changes, also run `docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile .`. +- For Docker runner or image changes, also run `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .`. - For CLI/docs changes, verify `npx tsx src/cli.ts --help` if touched docs mention CLI behavior. - For plugin/package metadata changes, run `npx tsx tests/smoke-skill-distribution.ts` and verify `npm pack --dry-run --json` includes required plugin files without result/cache directories. diff --git a/README.md b/README.md index 85153a4..3d84b69 100644 --- a/README.md +++ b/README.md @@ -196,7 +196,7 @@ npx tsx src/cli.ts --help For Docker runner or image changes: ```bash -docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile . +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . ``` Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials. diff --git a/docker/pi-acp-launcher.sh b/docker/pi-acp-launcher.sh new file mode 100755 index 0000000..f168bf0 --- /dev/null +++ b/docker/pi-acp-launcher.sh @@ -0,0 +1,11 @@ +#!/bin/sh +# Bridges SKILL_OPT_PROVIDER_* env vars to pi-acp's expected config. +# pi-acp reads its model/provider from a runtime config; we set +# OPENROUTER_API_KEY here if SKILL_OPT_PROVIDER_API_KEY is set. +set -e + +if [ -n "$SKILL_OPT_PROVIDER_API_KEY" ] && [ -z "$OPENROUTER_API_KEY" ]; then + export OPENROUTER_API_KEY="$SKILL_OPT_PROVIDER_API_KEY" +fi + +exec pi-acp "$@" diff --git a/docker/skill-optimizer-agent.Dockerfile b/docker/skill-optimizer-agent.Dockerfile new file mode 100644 index 0000000..0786ee3 --- /dev/null +++ b/docker/skill-optimizer-agent.Dockerfile @@ -0,0 +1,60 @@ +FROM node:22-bookworm + +ENV PATH="/opt/skill-opt/bin:/app/node_modules/.bin:/work/.venv/bin:${PATH}" \ + PIP_REQUIRE_VIRTUALENV=1 + +WORKDIR /app + +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + bash \ + ca-certificates \ + coreutils \ + curl \ + file \ + findutils \ + gawk \ + git \ + grep \ + jq \ + less \ + python-is-python3 \ + python3 \ + python3-pip \ + python3-venv \ + ripgrep \ + sed \ + unzip \ + wget \ + zip \ + && rm -rf /var/lib/apt/lists/* + +# --- Agent CLI install layer (cached as one layer for build speed) --- +# Each agent's installCmd is also kept in src/workbench/agents/registry.ts +# (single source of truth). If you change one here, change it there too. +RUN npm install -g \ + @zed-industries/claude-agent-acp@latest \ + @zed-industries/codex-acp@latest \ + @google/gemini-cli@latest \ + opencode-ai@latest \ + @mariozechner/pi-coding-agent@latest \ + pi-acp@latest + +# --- pi-acp launcher wrapper --- +COPY docker/pi-acp-launcher.sh /opt/skill-opt/bin/pi-acp-launcher +RUN chmod +x /opt/skill-opt/bin/pi-acp-launcher + +# --- Workbench container-runner (setup + grade modes only) --- +COPY package.json package-lock.json tsconfig.json ./ +COPY src ./src +COPY docs ./docs + +RUN npm ci \ + && npm run build \ + && useradd -m -u 10001 agent +USER agent + +# Container-runner is now used only for --setup and --grade modes. +# Agent dispatch happens host-side via ACP; the agent CLIs above are +# invoked directly by docker exec. +ENTRYPOINT ["node", "/app/dist/workbench/container-runner.js"] diff --git a/docker/workbench-runner.Dockerfile b/docker/workbench-runner.Dockerfile deleted file mode 100644 index 54087ee..0000000 --- a/docker/workbench-runner.Dockerfile +++ /dev/null @@ -1,41 +0,0 @@ -FROM node:22-bookworm - -ENV PATH="/app/node_modules/.bin:/work/.venv/bin:${PATH}" \ - PIP_REQUIRE_VIRTUALENV=1 - -WORKDIR /app - -RUN apt-get update \ - && apt-get install -y --no-install-recommends \ - bash \ - ca-certificates \ - coreutils \ - curl \ - file \ - findutils \ - gawk \ - git \ - grep \ - jq \ - less \ - python-is-python3 \ - python3 \ - python3-pip \ - python3-venv \ - ripgrep \ - sed \ - unzip \ - wget \ - zip \ - && rm -rf /var/lib/apt/lists/* - -COPY package.json package-lock.json tsconfig.json ./ -COPY src ./src -COPY docs ./docs - -RUN npm ci \ - && npm run build \ - && useradd -m -u 10001 agent -USER agent - -ENTRYPOINT ["node", "/app/dist/workbench/container-runner.js"] diff --git a/package.json b/package.json index 250bbad..356840f 100644 --- a/package.json +++ b/package.json @@ -79,6 +79,7 @@ "build": "tsc && chmod +x dist/cli.js", "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", + "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts" }, "dependencies": { diff --git a/src/workbench/agents/install-snippets.ts b/src/workbench/agents/install-snippets.ts new file mode 100644 index 0000000..26264c9 --- /dev/null +++ b/src/workbench/agents/install-snippets.ts @@ -0,0 +1,15 @@ +import { AGENTS } from './registry.js'; + +export function generateDockerInstallBlock(): string { + const lines: string[] = []; + for (const cfg of Object.values(AGENTS)) { + lines.push(`# Install ${cfg.name}: ${cfg.description}`); + lines.push(`RUN ${cfg.installCmd}`); + lines.push(''); + } + return lines.join('\n'); +} + +if (import.meta.url === `file://${process.argv[1]}`) { + console.log(generateDockerInstallBlock()); +} diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index b0f5451..b36be19 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -12,7 +12,7 @@ import type { ResolvedWorkbenchCase, WorkbenchCaseConfig } from './types.js'; import { timestampSlug } from './utils.js'; import { prepareWorkbenchDirectory } from './workspace.js'; -const DEFAULT_WORKBENCH_IMAGE = 'skill-optimizer-workbench:local'; +const DEFAULT_WORKBENCH_IMAGE = 'skill-optimizer-agent:local'; const AGENT_RESULTS_DIR = '/tmp/workbench-results'; const AGENT_PATH = '/work/bin:/app/node_modules/.bin:/work/.venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin'; From ec1520e80f5b84648592484017ee994d222c133c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:44:22 -0500 Subject: [PATCH 077/121] feat(acp): subscription-first auth resolution with fail-loud env fallback Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/acp/auth.ts | 58 ++++++++++++++++++++++++++++++ src/workbench/agents/registry.ts | 4 +-- tests/acp/auth.test.ts | 61 ++++++++++++++++++++++++++++++++ 3 files changed, 121 insertions(+), 2 deletions(-) create mode 100644 src/workbench/acp/auth.ts create mode 100644 tests/acp/auth.test.ts diff --git a/src/workbench/acp/auth.ts b/src/workbench/acp/auth.ts new file mode 100644 index 0000000..0c5b8e9 --- /dev/null +++ b/src/workbench/acp/auth.ts @@ -0,0 +1,58 @@ +import { existsSync } from 'node:fs'; +import { homedir } from 'node:os'; +import { join, normalize } from 'node:path'; +import type { AgentConfig } from '../agents/registry.js'; + +export interface AuthContext { + home?: string; // override $HOME for testing + env: Record; +} + +export interface SubscriptionAuthResult { + mode: 'subscription'; + files: Array<{ hostPath: string; containerPath: string }>; +} + +export interface EnvAuthResult { + mode: 'env'; + envNames: string[]; +} + +export type AuthResult = SubscriptionAuthResult | EnvAuthResult; + +export function resolveAuth(agent: AgentConfig, ctx: AuthContext): AuthResult { + const home = ctx.home ?? homedir(); + + if (agent.subscriptionAuth) { + const detectPath = expandHome(agent.subscriptionAuth.detectFile, home); + if (existsSync(detectPath)) { + return { + mode: 'subscription', + files: agent.subscriptionAuth.files.map((f) => ({ + hostPath: expandHome(f.hostPath, home), + containerPath: f.containerPath, // {home} placeholder, resolved at mount time + })), + }; + } + } + + const missing = agent.requiresEnv.filter((name) => !ctx.env[name]); + if (missing.length > 0 && agent.requiresEnv.length > 0) { + const hint = agent.loginHint ? ` (run \`${agent.loginHint}\` or set ${missing.join(', ')})` : ''; + throw new Error( + `Agent ${agent.name} requires auth but none is available${hint}. Missing: ${missing.join(', ')}.`, + ); + } + + return { mode: 'env', envNames: agent.requiresEnv }; +} + +function expandHome(p: string, home: string): string { + if (p.startsWith('~/')) return normalize(join(home, p.slice(2))); + if (p === '~') return home; + return p; +} + +export function resolveContainerPath(template: string, agentHome: string): string { + return template.replace('{home}', agentHome); +} diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts index b33a292..641e69a 100644 --- a/src/workbench/agents/registry.ts +++ b/src/workbench/agents/registry.ts @@ -59,9 +59,9 @@ export const AGENTS: Record = { homeDirs: [], subscriptionAuth: { replacesEnv: 'ANTHROPIC_API_KEY', - detectFile: '~/.claude/.credentials.json', + detectFile: '~/.credentials.json', files: [ - { hostPath: '~/.claude/.credentials.json', containerPath: '{home}/.claude/.credentials.json' }, + { hostPath: '~/.credentials.json', containerPath: '{home}/.credentials.json' }, ], }, acpModelFormat: 'bare', diff --git a/tests/acp/auth.test.ts b/tests/acp/auth.test.ts new file mode 100644 index 0000000..28db063 --- /dev/null +++ b/tests/acp/auth.test.ts @@ -0,0 +1,61 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, writeFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { resolveAuth } from '../../src/workbench/acp/auth.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +test('resolveAuth uses subscription file when present', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + writeFileSync(join(home, '.credentials.json'), '{}'); + + const auth = resolveAuth(AGENTS['claude-agent-acp'], { + home, + env: {}, + }); + + assert.equal(auth.mode, 'subscription'); + assert.equal(auth.files.length, 1); + assert.equal(auth.files[0].hostPath, join(home, '.credentials.json')); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth falls back to env API key when subscription absent', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + const auth = resolveAuth(AGENTS['claude-agent-acp'], { + home, + env: { ANTHROPIC_API_KEY: 'sk-test-key' }, + }); + + assert.equal(auth.mode, 'env'); + assert.deepEqual(auth.envNames, ['ANTHROPIC_API_KEY']); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth throws when neither subscription nor env key present', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + assert.throws( + () => resolveAuth(AGENTS['claude-agent-acp'], { home, env: {} }), + /claude login|ANTHROPIC_API_KEY/, + ); + + rmSync(home, { recursive: true }); +}); + +test('resolveAuth returns env mode for agents without subscriptionAuth', () => { + const home = mkdtempSync(join(tmpdir(), 'auth-test-')); + + const auth = resolveAuth(AGENTS['pi-acp'], { + home, + env: { OPENROUTER_API_KEY: 'sk-test' }, + }); + + assert.equal(auth.mode, 'env'); + + rmSync(home, { recursive: true }); +}); From 45160430a276cc1db2c92871a0f01c354485695d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:45:09 -0500 Subject: [PATCH 078/121] chore(test): wire ACP auth test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 356840f..7d6bcec 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 196c347b943b43bdbd0c94bf54c639ac1a0e7c5f Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:46:13 -0500 Subject: [PATCH 079/121] fix(acp): restore .claude subdir in registry; fix auth test to match --- src/workbench/agents/registry.ts | 4 ++-- tests/acp/auth.test.ts | 7 ++++--- 2 files changed, 6 insertions(+), 5 deletions(-) diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts index 641e69a..b33a292 100644 --- a/src/workbench/agents/registry.ts +++ b/src/workbench/agents/registry.ts @@ -59,9 +59,9 @@ export const AGENTS: Record = { homeDirs: [], subscriptionAuth: { replacesEnv: 'ANTHROPIC_API_KEY', - detectFile: '~/.credentials.json', + detectFile: '~/.claude/.credentials.json', files: [ - { hostPath: '~/.credentials.json', containerPath: '{home}/.credentials.json' }, + { hostPath: '~/.claude/.credentials.json', containerPath: '{home}/.claude/.credentials.json' }, ], }, acpModelFormat: 'bare', diff --git a/tests/acp/auth.test.ts b/tests/acp/auth.test.ts index 28db063..107755a 100644 --- a/tests/acp/auth.test.ts +++ b/tests/acp/auth.test.ts @@ -1,6 +1,6 @@ import { test } from 'node:test'; import { strict as assert } from 'node:assert'; -import { mkdtempSync, writeFileSync, rmSync } from 'node:fs'; +import { mkdtempSync, writeFileSync, mkdirSync, rmSync } from 'node:fs'; import { tmpdir } from 'node:os'; import { join } from 'node:path'; import { resolveAuth } from '../../src/workbench/acp/auth.js'; @@ -8,7 +8,8 @@ import { AGENTS } from '../../src/workbench/agents/registry.js'; test('resolveAuth uses subscription file when present', () => { const home = mkdtempSync(join(tmpdir(), 'auth-test-')); - writeFileSync(join(home, '.credentials.json'), '{}'); + mkdirSync(join(home, '.claude'), { recursive: true }); + writeFileSync(join(home, '.claude', '.credentials.json'), '{}'); const auth = resolveAuth(AGENTS['claude-agent-acp'], { home, @@ -17,7 +18,7 @@ test('resolveAuth uses subscription file when present', () => { assert.equal(auth.mode, 'subscription'); assert.equal(auth.files.length, 1); - assert.equal(auth.files[0].hostPath, join(home, '.credentials.json')); + assert.equal(auth.files[0].hostPath, join(home, '.claude', '.credentials.json')); rmSync(home, { recursive: true }); }); From b21fc5e761e43889e6e914341c6d8b24c274571f Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:47:48 -0500 Subject: [PATCH 080/121] feat(acp): skill-deploy module mounts to agent's native skill path --- src/workbench/acp/skill-deploy.ts | 37 +++++++++++++++++++++++++++++++ tests/acp/skill-deploy.test.ts | 31 ++++++++++++++++++++++++++ 2 files changed, 68 insertions(+) create mode 100644 src/workbench/acp/skill-deploy.ts create mode 100644 tests/acp/skill-deploy.test.ts diff --git a/src/workbench/acp/skill-deploy.ts b/src/workbench/acp/skill-deploy.ts new file mode 100644 index 0000000..0e7ad91 --- /dev/null +++ b/src/workbench/acp/skill-deploy.ts @@ -0,0 +1,37 @@ +import { join } from 'node:path'; +import type { AgentConfig } from '../agents/registry.js'; + +export interface SkillMount { + hostPath: string; + containerPath: string; + readOnly: boolean; +} + +export interface ComputeSkillMountParams { + agent: AgentConfig; + skillSlug: string; + hostSkillDir: string; // absolute path on host to the skill folder + agentHome: string; // /home/agent (or whatever the container user's $HOME is) +} + +export function computeSkillMount(params: ComputeSkillMountParams): SkillMount { + const skillRoot = params.agent.skillPaths[0]; + if (!skillRoot) { + throw new Error(`Agent ${params.agent.name} has no skillPaths configured`); + } + const expanded = skillRoot.replace('$HOME', params.agentHome); + return { + hostPath: params.hostSkillDir, + containerPath: join(expanded, params.skillSlug), + readOnly: true, + }; +} + +export function dockerMountFlag(mount: SkillMount): string { + const ro = mount.readOnly ? ':ro' : ':rw'; + return `-v ${shellQuote(mount.hostPath)}:${shellQuote(mount.containerPath)}${ro}`; +} + +function shellQuote(s: string): string { + return `'${s.replace(/'/g, `'\\''`)}'`; +} diff --git a/tests/acp/skill-deploy.test.ts b/tests/acp/skill-deploy.test.ts new file mode 100644 index 0000000..0bcb3b5 --- /dev/null +++ b/tests/acp/skill-deploy.test.ts @@ -0,0 +1,31 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { computeSkillMount } from '../../src/workbench/acp/skill-deploy.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +test('computeSkillMount returns native path for claude-agent-acp', () => { + const mount = computeSkillMount({ + agent: AGENTS['claude-agent-acp'], + skillSlug: 'web-design-guidelines', + hostSkillDir: '/host/path/to/skill', + agentHome: '/home/agent', + }); + + assert.equal(mount.hostPath, '/host/path/to/skill'); + assert.equal(mount.containerPath, '/home/agent/.claude/skills/web-design-guidelines'); + assert.equal(mount.readOnly, true); +}); + +test('computeSkillMount handles agents whose skill_paths use $HOME', () => { + for (const cfg of Object.values(AGENTS)) { + const mount = computeSkillMount({ + agent: cfg, + skillSlug: 'test-skill', + hostSkillDir: '/host/dir', + agentHome: '/home/agent', + }); + assert.ok(!mount.containerPath.includes('$HOME')); + assert.ok(mount.containerPath.startsWith('/home/agent/')); + assert.ok(mount.containerPath.endsWith('/test-skill')); + } +}); From 6cddae5bb7d51d079d99934d854de6b0b98c1421 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:48:15 -0500 Subject: [PATCH 081/121] chore(test): wire ACP skill-deploy test into npm test --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 7d6bcec..9487a70 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 77e1c0d836138d6241c0a896bffa02e18b5d221f Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:49:04 -0500 Subject: [PATCH 082/121] fix(docker): update stale Dockerfile path refs after Task 6 rename --- .claude/worktrees/agent-a0620e810e20af755 | 1 + .claude/worktrees/agent-a2bdbcde626027d00 | 1 + .claude/worktrees/agent-aac9dc7ea895a8a09 | 1 + .claude/worktrees/agent-aaf823bf7de0c6401 | 1 + .claude/worktrees/agent-abe33dd2c200c608a | 1 + .claude/worktrees/agent-browser-rerun | 1 + .claude/worktrees/firebase-v1.3 | 1 + .claude/worktrees/supabase-pilot-v2 | 1 + .claude/worktrees/v1.3-impl | 1 + .claude/worktrees/v1.4-rationale | 1 + src/workbench/docker-runner.ts | 2 +- tests/smoke-workbench-docker-runner.ts | 4 ++-- 12 files changed, 13 insertions(+), 3 deletions(-) create mode 160000 .claude/worktrees/agent-a0620e810e20af755 create mode 160000 .claude/worktrees/agent-a2bdbcde626027d00 create mode 160000 .claude/worktrees/agent-aac9dc7ea895a8a09 create mode 160000 .claude/worktrees/agent-aaf823bf7de0c6401 create mode 160000 .claude/worktrees/agent-abe33dd2c200c608a create mode 160000 .claude/worktrees/agent-browser-rerun create mode 160000 .claude/worktrees/firebase-v1.3 create mode 160000 .claude/worktrees/supabase-pilot-v2 create mode 160000 .claude/worktrees/v1.3-impl create mode 160000 .claude/worktrees/v1.4-rationale diff --git a/.claude/worktrees/agent-a0620e810e20af755 b/.claude/worktrees/agent-a0620e810e20af755 new file mode 160000 index 0000000..1744daf --- /dev/null +++ b/.claude/worktrees/agent-a0620e810e20af755 @@ -0,0 +1 @@ +Subproject commit 1744daf5b8f4c0c73ff08842424810422b4de977 diff --git a/.claude/worktrees/agent-a2bdbcde626027d00 b/.claude/worktrees/agent-a2bdbcde626027d00 new file mode 160000 index 0000000..41009f8 --- /dev/null +++ b/.claude/worktrees/agent-a2bdbcde626027d00 @@ -0,0 +1 @@ +Subproject commit 41009f87e932bc519c1560d0190c3363c47d2d6b diff --git a/.claude/worktrees/agent-aac9dc7ea895a8a09 b/.claude/worktrees/agent-aac9dc7ea895a8a09 new file mode 160000 index 0000000..f0883ad --- /dev/null +++ b/.claude/worktrees/agent-aac9dc7ea895a8a09 @@ -0,0 +1 @@ +Subproject commit f0883adc271cb6a73b7d1f58ab29da421faa2da0 diff --git a/.claude/worktrees/agent-aaf823bf7de0c6401 b/.claude/worktrees/agent-aaf823bf7de0c6401 new file mode 160000 index 0000000..b342220 --- /dev/null +++ b/.claude/worktrees/agent-aaf823bf7de0c6401 @@ -0,0 +1 @@ +Subproject commit b34222097d65ffbfc545eaaeff716b4b0794c9e0 diff --git a/.claude/worktrees/agent-abe33dd2c200c608a b/.claude/worktrees/agent-abe33dd2c200c608a new file mode 160000 index 0000000..f9463be --- /dev/null +++ b/.claude/worktrees/agent-abe33dd2c200c608a @@ -0,0 +1 @@ +Subproject commit f9463bec339d63d71482a50bff7c36dcac97417d diff --git a/.claude/worktrees/agent-browser-rerun b/.claude/worktrees/agent-browser-rerun new file mode 160000 index 0000000..26601a0 --- /dev/null +++ b/.claude/worktrees/agent-browser-rerun @@ -0,0 +1 @@ +Subproject commit 26601a0adee88601ff0004bb2795f77993496911 diff --git a/.claude/worktrees/firebase-v1.3 b/.claude/worktrees/firebase-v1.3 new file mode 160000 index 0000000..c519d28 --- /dev/null +++ b/.claude/worktrees/firebase-v1.3 @@ -0,0 +1 @@ +Subproject commit c519d281a48141e6dae37435d22c5539ec868cd5 diff --git a/.claude/worktrees/supabase-pilot-v2 b/.claude/worktrees/supabase-pilot-v2 new file mode 160000 index 0000000..59c3e85 --- /dev/null +++ b/.claude/worktrees/supabase-pilot-v2 @@ -0,0 +1 @@ +Subproject commit 59c3e85e3fe254832a79bc8948cf0da147e523f4 diff --git a/.claude/worktrees/v1.3-impl b/.claude/worktrees/v1.3-impl new file mode 160000 index 0000000..96dbe31 --- /dev/null +++ b/.claude/worktrees/v1.3-impl @@ -0,0 +1 @@ +Subproject commit 96dbe31569e151bc2be960994201a767275bcf6e diff --git a/.claude/worktrees/v1.4-rationale b/.claude/worktrees/v1.4-rationale new file mode 160000 index 0000000..2e4ea27 --- /dev/null +++ b/.claude/worktrees/v1.4-rationale @@ -0,0 +1 @@ +Subproject commit 2e4ea27ccf9e4a3e55a9a4c86ea113a6151768ce diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index b36be19..b72f75b 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -456,7 +456,7 @@ async function ensureDockerImage(image: string, repoRoot: string): Promise return; } - const dockerfilePath = join(repoRoot, 'docker', 'workbench-runner.Dockerfile'); + const dockerfilePath = join(repoRoot, 'docker', 'skill-optimizer-agent.Dockerfile'); if (!existsSync(dockerfilePath)) { throw new Error(`Dockerfile not found: ${dockerfilePath}`); } diff --git a/tests/smoke-workbench-docker-runner.ts b/tests/smoke-workbench-docker-runner.ts index 1285595..acdf61c 100644 --- a/tests/smoke-workbench-docker-runner.ts +++ b/tests/smoke-workbench-docker-runner.ts @@ -27,13 +27,13 @@ test('packageRootFromModuleUrl resolves repo root independently of cwd', () => { }); test('workbench image runs agents as non-root with venv-only pip installs', () => { - const dockerfile = readFileSync(join(process.cwd(), 'docker', 'workbench-runner.Dockerfile'), 'utf-8'); + const dockerfile = readFileSync(join(process.cwd(), 'docker', 'skill-optimizer-agent.Dockerfile'), 'utf-8'); assert.match(dockerfile, /useradd .* agent/); assert.match(dockerfile, /USER agent/); assert.match(dockerfile, /ENTRYPOINT \["node", "\/app\/dist\/workbench\/container-runner\.js"\]/); assert.match(dockerfile, /PIP_REQUIRE_VIRTUALENV=1/); - assert.match(dockerfile, /PATH="\/app\/node_modules\/\.bin:\/work\/\.venv\/bin:/); + assert.match(dockerfile, /PATH="\/opt\/skill-opt\/bin:\/app\/node_modules\/\.bin:\/work\/\.venv\/bin:/); assert.doesNotMatch(dockerfile, /PIP_BREAK_SYSTEM_PACKAGES/); }); From 301fd58331ae7d29c9d77b518f001a5f68188c34 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:51:43 -0500 Subject: [PATCH 083/121] feat(acp): per-agent MCP config writer (claude/codex/gemini/opencode) Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/acp/mcp-config-writer.ts | 80 +++++++++++++++++++++ tests/acp/mcp-config-writer.test.ts | 88 +++++++++++++++++++++++ tests/fixtures/mcp-cases/sample-stdio.yml | 11 +++ 3 files changed, 179 insertions(+) create mode 100644 src/workbench/acp/mcp-config-writer.ts create mode 100644 tests/acp/mcp-config-writer.test.ts create mode 100644 tests/fixtures/mcp-cases/sample-stdio.yml diff --git a/src/workbench/acp/mcp-config-writer.ts b/src/workbench/acp/mcp-config-writer.ts new file mode 100644 index 0000000..0bd7012 --- /dev/null +++ b/src/workbench/acp/mcp-config-writer.ts @@ -0,0 +1,80 @@ +import { mkdirSync, readFileSync, writeFileSync, existsSync } from 'node:fs'; +import { dirname, join } from 'node:path'; +import type { AgentConfig } from '../agents/registry.js'; + +export interface McpServerSpec { + command?: string; + args?: string[]; + env?: Record; + url?: string; + headers?: Record; +} + +export interface WriteMcpConfigParams { + agent: AgentConfig; + caseConfig: { mcpServers?: Record }; + agentHomeOnHost: string; // path on host that will become container's $HOME +} + +export function writeMcpConfig(params: WriteMcpConfigParams): void { + const servers = params.caseConfig.mcpServers; + if (!servers || Object.keys(servers).length === 0) return; + + switch (params.agent.name) { + case 'claude-agent-acp': + return writeClaude(servers, params.agentHomeOnHost); + case 'codex-acp': + return writeCodex(servers, params.agentHomeOnHost); + case 'gemini': + return writeGemini(servers, params.agentHomeOnHost); + case 'opencode': + return writeOpencode(servers, params.agentHomeOnHost); + case 'pi-acp': + throw new Error('pi-acp MCP support pending — see plan Task 9b'); + default: + throw new Error(`MCP not implemented for agent ${params.agent.name}`); + } +} + +function writeClaude(servers: Record, home: string): void { + const path = join(home, '.claude.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcpServers = { ...(existing.mcpServers ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} + +function writeCodex(servers: Record, home: string): void { + const path = join(home, '.codex/config.toml'); + mkdirSync(dirname(path), { recursive: true }); + const blocks: string[] = []; + for (const [name, spec] of Object.entries(servers)) { + blocks.push(`[mcp_servers.${name}]`); + if (spec.command) blocks.push(`command = ${JSON.stringify(spec.command)}`); + if (spec.args) blocks.push(`args = ${JSON.stringify(spec.args)}`); + if (spec.env) { + blocks.push(`[mcp_servers.${name}.env]`); + for (const [k, v] of Object.entries(spec.env)) { + blocks.push(`${k} = ${JSON.stringify(v)}`); + } + } + blocks.push(''); + } + writeFileSync(path, blocks.join('\n')); +} + +function writeGemini(servers: Record, home: string): void { + const path = join(home, '.gemini/settings.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcpServers = { ...(existing.mcpServers ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} + +function writeOpencode(servers: Record, home: string): void { + const path = join(home, '.config/opencode/opencode.json'); + mkdirSync(dirname(path), { recursive: true }); + const existing = existsSync(path) ? JSON.parse(readFileSync(path, 'utf-8')) : {}; + existing.mcp = { ...(existing.mcp ?? {}), ...servers }; + writeFileSync(path, JSON.stringify(existing, null, 2)); +} diff --git a/tests/acp/mcp-config-writer.test.ts b/tests/acp/mcp-config-writer.test.ts new file mode 100644 index 0000000..2987359 --- /dev/null +++ b/tests/acp/mcp-config-writer.test.ts @@ -0,0 +1,88 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, readFileSync, rmSync, existsSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { writeMcpConfig } from '../../src/workbench/acp/mcp-config-writer.js'; +import { AGENTS } from '../../src/workbench/agents/registry.js'; + +const sampleCase = { + mcpServers: { + calculator: { command: 'node', args: ['/work/mcp/calculator.mjs'] }, + }, +}; + +test('writeMcpConfig for claude writes to .claude.json mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['claude-agent-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.claude.json'), 'utf-8')); + assert.ok(data.mcpServers?.calculator); + assert.equal(data.mcpServers.calculator.command, 'node'); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for codex writes to .codex/config.toml', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['codex-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const toml = readFileSync(join(home, '.codex/config.toml'), 'utf-8'); + assert.match(toml, /\[mcp_servers\.calculator\]/); + assert.match(toml, /command = "node"/); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for gemini writes to .gemini/settings.json mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['gemini'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.gemini/settings.json'), 'utf-8')); + assert.ok(data.mcpServers?.calculator); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for opencode writes to .config/opencode/opencode.json mcp', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['opencode'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }); + const data = JSON.parse(readFileSync(join(home, '.config/opencode/opencode.json'), 'utf-8')); + assert.ok(data.mcp?.calculator); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig for pi-acp throws "not yet supported" (pending Task 9b)', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + assert.throws( + () => writeMcpConfig({ + agent: AGENTS['pi-acp'], + caseConfig: sampleCase, + agentHomeOnHost: home, + }), + /pi-acp MCP support pending/, + ); + rmSync(home, { recursive: true }); +}); + +test('writeMcpConfig is a no-op when caseConfig has no mcpServers', () => { + const home = mkdtempSync(join(tmpdir(), 'mcp-test-')); + writeMcpConfig({ + agent: AGENTS['claude-agent-acp'], + caseConfig: {}, + agentHomeOnHost: home, + }); + // No files should be written. + assert.equal(existsSync(join(home, '.claude.json')), false); + rmSync(home, { recursive: true }); +}); diff --git a/tests/fixtures/mcp-cases/sample-stdio.yml b/tests/fixtures/mcp-cases/sample-stdio.yml new file mode 100644 index 0000000..9704bc7 --- /dev/null +++ b/tests/fixtures/mcp-cases/sample-stdio.yml @@ -0,0 +1,11 @@ +name: sample-with-mcp +agent: claude-agent-acp +model: claude-haiku-4-5 +task: do nothing +graders: + - name: noop + command: 'true' +mcpServers: + calculator: + command: node + args: ['/work/mcp/calculator.mjs'] From 943cdc52bf7a72acbb9cf4dccd5fe11393e3f53c Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:51:55 -0500 Subject: [PATCH 084/121] chore(test): wire MCP config writer test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 9487a70..367ee71 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 2417bfbf05552ba738a266c65a29b5bcc1c3fc84 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:53:00 -0500 Subject: [PATCH 085/121] =?UTF-8?q?docs(mcp):=20pick=20option=20(c)=20for?= =?UTF-8?q?=20pi-acp=20MCP=20=E2=80=94=20defer=20to=20v1.1?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- CONTRIBUTING.md | 4 ++++ src/workbench/acp/mcp-config-writer.ts | 9 +++++++++ 2 files changed, 13 insertions(+) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a706ece..a59714b 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -57,6 +57,10 @@ All three commands must pass before opening a PR when code changes are involved. - For CLI/docs changes, verify `npx tsx src/cli.ts --help` if touched docs mention CLI behavior. - For plugin/package metadata changes, run `npx tsx tests/smoke-skill-distribution.ts` and verify `npm pack --dry-run --json` includes required plugin files without result/cache directories. +## Agent capabilities and limitations (v1) + +MCP server support is available for `claude-agent-acp`, `codex-acp`, `gemini`, and `opencode` agents. `pi-acp` MCP support is deferred to v1.1 pending documentation of pi-acp's native MCP config path. + ## Adding workbench capabilities Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/shared/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior. diff --git a/src/workbench/acp/mcp-config-writer.ts b/src/workbench/acp/mcp-config-writer.ts index 0bd7012..80fb4d6 100644 --- a/src/workbench/acp/mcp-config-writer.ts +++ b/src/workbench/acp/mcp-config-writer.ts @@ -2,6 +2,15 @@ import { mkdirSync, readFileSync, writeFileSync, existsSync } from 'node:fs'; import { dirname, join } from 'node:path'; import type { AgentConfig } from '../agents/registry.js'; +// pi-acp MCP resolution (plan Task 9b, decided 2026-05-25): +// Decision: option (c) — defer pi-acp MCP support to a follow-up. +// Reason: pi-acp's native MCP config path is not documented in the +// pi-coding-agent or pi-acp npm packages; investigating + implementing +// would block v1 on a tangent. For v1, any case declaring `mcpServers:` +// with `agent: pi-acp` gets a clear "pending" error from writeMcpConfig +// (see switch case below). Users requiring MCP should use one of the +// supported agents: claude-agent-acp, codex-acp, gemini, opencode. + export interface McpServerSpec { command?: string; args?: string[]; From 62ce6adb253aea6ae4262d06c36f15750d6e7b41 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:55:17 -0500 Subject: [PATCH 086/121] feat(parse-trace): ACP trace helpers (iterMessages/iterToolCalls/computeMetrics) Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/parse-trace.ts | 198 ++++++++++++++++++ tests/fixtures/acp-traces/claude-sample.jsonl | 11 + tests/parse-trace.test.ts | 45 ++++ 3 files changed, 254 insertions(+) create mode 100644 src/workbench/parse-trace.ts create mode 100644 tests/fixtures/acp-traces/claude-sample.jsonl create mode 100644 tests/parse-trace.test.ts diff --git a/src/workbench/parse-trace.ts b/src/workbench/parse-trace.ts new file mode 100644 index 0000000..1447641 --- /dev/null +++ b/src/workbench/parse-trace.ts @@ -0,0 +1,198 @@ +export interface ParsedMessage { + role: 'user' | 'assistant'; + text?: string; + thinking?: string; + timestamp?: string; +} + +export interface ParsedToolCall { + id: string; + kind: string; + title?: string; + status: 'in_progress' | 'completed' | 'failed' | 'pending'; + argsText?: string; + resultText?: string; + isError?: boolean; +} + +export interface ComputedMetrics { + durationMs: number; + turns: number; + toolCalls: number; + bashCalls: number; + readCalls: number; + writeCalls: number; + editCalls: number; + stopReason?: string; + tokens: { input: number; output: number; cacheRead: number; cacheWrite: number; total: number }; +} + +interface TraceHeader { + type: 'trace_start'; + caseName?: string; + agent?: string; + model?: string; + startedAt?: string; + endedAt?: string; +} + +function parseLines(jsonl: string): unknown[] { + return jsonl.split(/\r?\n/).filter(Boolean).map((line) => { + try { return JSON.parse(line); } catch { return null; } + }).filter((x) => x !== null); +} + +function getHeader(rows: unknown[]): TraceHeader | undefined { + for (const row of rows) { + if (typeof row === 'object' && row !== null && (row as any).type === 'trace_start') { + return row as TraceHeader; + } + } + return undefined; +} + +function extractText(content: unknown): string | undefined { + if (!content || typeof content !== 'object') return undefined; + const c = content as any; + if (typeof c.text === 'string') return c.text; + if (c.content && typeof c.content.text === 'string') return c.content.text; + if (Array.isArray(c)) { + return c.map((part: any) => extractText(part)).filter(Boolean).join('\n') || undefined; + } + return undefined; +} + +export function* iterMessages(jsonl: string): IterableIterator { + const rows = parseLines(jsonl); + let assistantText = ''; + let assistantThinking = ''; + + for (const row of rows) { + const r = row as any; + if (r.method !== 'session/update') continue; + const update = r.params?.update; + if (!update) continue; + + switch (update.sessionUpdate) { + case 'agent_message_chunk': + assistantText += extractText(update.content) ?? ''; + break; + case 'agent_thought_chunk': + assistantThinking += extractText(update.content) ?? ''; + break; + case 'user_message_chunk': + yield { role: 'user', text: extractText(update.content) }; + break; + } + } + + if (assistantText || assistantThinking) { + yield { + role: 'assistant', + text: assistantText || undefined, + thinking: assistantThinking || undefined, + }; + } +} + +export function* iterToolCalls(jsonl: string): IterableIterator { + const rows = parseLines(jsonl); + const calls = new Map(); + + for (const row of rows) { + const r = row as any; + if (r.method !== 'session/update') continue; + const update = r.params?.update; + if (!update) continue; + + if (update.sessionUpdate === 'tool_call') { + calls.set(update.toolCallId, { + id: update.toolCallId, + kind: update.kind ?? 'other', + title: update.title, + status: update.status ?? 'in_progress', + argsText: extractText(update.content), + }); + } else if (update.sessionUpdate === 'tool_call_update') { + const existing = calls.get(update.toolCallId); + if (existing) { + existing.status = update.status ?? existing.status; + if (update.content) existing.resultText = extractText(update.content); + if (update.status === 'failed') existing.isError = true; + } + } + } + + yield* calls.values(); +} + +export function computeMetrics(jsonl: string): ComputedMetrics { + const rows = parseLines(jsonl); + const header = getHeader(rows); + + const m: ComputedMetrics = { + durationMs: 0, + turns: 0, + toolCalls: 0, + bashCalls: 0, + readCalls: 0, + writeCalls: 0, + editCalls: 0, + tokens: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 }, + }; + + if (header?.startedAt) { + const start = Date.parse(header.startedAt); + const end = header.endedAt ? Date.parse(header.endedAt) : Date.now(); + if (Number.isFinite(start) && Number.isFinite(end)) { + m.durationMs = Math.max(0, end - start); + } + } + + for (const row of rows) { + const r = row as any; + if (r.method === 'session/update') { + const update = r.params?.update; + if (!update) continue; + if (update.sessionUpdate === 'agent_message_chunk') m.turns += 1; + if (update.sessionUpdate === 'tool_call') { + m.toolCalls += 1; + const kind = update.kind ?? 'other'; + if (kind === 'execute') m.bashCalls += 1; + if (kind === 'read') m.readCalls += 1; + if (kind === 'write') m.writeCalls += 1; + if (kind === 'edit') m.editCalls += 1; + } + } + if (r.id && r.result?.stopReason) { + m.stopReason = r.result.stopReason; + const u = r.result.usage; + if (u) { + m.tokens.input += Number(u.inputTokens ?? 0); + m.tokens.output += Number(u.outputTokens ?? 0); + m.tokens.cacheRead += Number(u.cacheReadTokens ?? 0); + m.tokens.cacheWrite += Number(u.cacheCreationTokens ?? 0); + } + } + } + + m.tokens.total = m.tokens.input + m.tokens.output + m.tokens.cacheRead + m.tokens.cacheWrite; + return m; +} + +export function getFinalAssistantMessage(jsonl: string): string | undefined { + for (const msg of iterMessages(jsonl)) { + if (msg.role === 'assistant' && msg.text) return msg.text; + } + return undefined; +} + +export function getFailureEvidence(jsonl: string): string[] { + const evidence: string[] = []; + for (const call of iterToolCalls(jsonl)) { + if (call.isError && call.resultText) { + evidence.push(`tool ${call.kind} (${call.id}) failed: ${call.resultText}`); + } + } + return evidence; +} diff --git a/tests/fixtures/acp-traces/claude-sample.jsonl b/tests/fixtures/acp-traces/claude-sample.jsonl new file mode 100644 index 0000000..e539af9 --- /dev/null +++ b/tests/fixtures/acp-traces/claude-sample.jsonl @@ -0,0 +1,11 @@ +{"type":"trace_start","schemaVersion":2,"caseName":"sample","agent":"claude-agent-acp","model":"claude-haiku-4-5","startedAt":"2026-05-25T00:00:00Z"} +{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"0.22","clientCapabilities":{}}} +{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"0.22","agentCapabilities":{},"authMethods":[]}} +{"jsonrpc":"2.0","id":2,"method":"session/new","params":{"cwd":"/work","mcpServers":[]}} +{"jsonrpc":"2.0","id":2,"result":{"sessionId":"s-1"}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"agent_thought_chunk","content":{"type":"text","text":"Thinking about it..."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"agent_message_chunk","content":{"type":"text","text":"I'll write the file now."}}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"tool_call","toolCallId":"tc-1","title":"Write file","kind":"edit","status":"in_progress","content":[{"type":"content","content":{"type":"text","text":"out.txt"}}]}}} +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"s-1","update":{"sessionUpdate":"tool_call_update","toolCallId":"tc-1","status":"completed","content":[{"type":"content","content":{"type":"text","text":"Wrote 6 bytes"}}]}}} +{"jsonrpc":"2.0","id":3,"method":"session/prompt","params":{"sessionId":"s-1","prompt":[{"type":"text","text":"Write hello to out.txt"}]}} +{"jsonrpc":"2.0","id":3,"result":{"stopReason":"end_turn","usage":{"inputTokens":120,"outputTokens":45,"cacheReadTokens":0,"cacheCreationTokens":0}}} diff --git a/tests/parse-trace.test.ts b/tests/parse-trace.test.ts new file mode 100644 index 0000000..75f4c80 --- /dev/null +++ b/tests/parse-trace.test.ts @@ -0,0 +1,45 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { readFileSync } from 'node:fs'; +import { + iterMessages, + iterToolCalls, + computeMetrics, + getFinalAssistantMessage, +} from '../src/workbench/parse-trace.js'; + +const fixturePath = 'tests/fixtures/acp-traces/claude-sample.jsonl'; +const trace = readFileSync(fixturePath, 'utf-8'); + +test('iterMessages yields assistant text + thinking', () => { + const msgs = [...iterMessages(trace)]; + const assistant = msgs.find((m) => m.role === 'assistant'); + assert.ok(assistant); + assert.match(assistant.text ?? '', /write the file/); + assert.match(assistant.thinking ?? '', /Thinking about it/); +}); + +test('iterToolCalls pairs tool_call and tool_call_update', () => { + const calls = [...iterToolCalls(trace)]; + assert.equal(calls.length, 1); + assert.equal(calls[0].id, 'tc-1'); + assert.equal(calls[0].kind, 'edit'); + assert.equal(calls[0].status, 'completed'); + assert.match(calls[0].resultText ?? '', /Wrote 6 bytes/); +}); + +test('computeMetrics returns tokens, duration, tool counts', () => { + const m = computeMetrics(trace); + assert.equal(m.tokens.input, 120); + assert.equal(m.tokens.output, 45); + assert.equal(m.tokens.total, 165); + assert.equal(m.toolCalls, 1); + assert.equal(m.editCalls, 1); + assert.equal(m.bashCalls, 0); + assert.equal(m.stopReason, 'end_turn'); +}); + +test('getFinalAssistantMessage returns concatenated assistant text', () => { + const final = getFinalAssistantMessage(trace); + assert.match(final ?? '', /write the file/); +}); From 41b62e601fb43bee278b4bfa47e4fa5036850daf Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:55:20 -0500 Subject: [PATCH 087/121] chore(test): wire parse-trace test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 367ee71..513e6b8 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From e7d36b0a73a586897df8af1784518d0c87c947df Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:56:22 -0500 Subject: [PATCH 088/121] feat(acp): trace recorder writes raw ACP messages with header+endedAt Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/acp/trace-recorder.ts | 48 +++++++++++++++++++++++++++++ tests/acp/trace-recorder.test.ts | 34 ++++++++++++++++++++ 2 files changed, 82 insertions(+) create mode 100644 src/workbench/acp/trace-recorder.ts create mode 100644 tests/acp/trace-recorder.test.ts diff --git a/src/workbench/acp/trace-recorder.ts b/src/workbench/acp/trace-recorder.ts new file mode 100644 index 0000000..367e0b0 --- /dev/null +++ b/src/workbench/acp/trace-recorder.ts @@ -0,0 +1,48 @@ +import { writeFileSync, appendFileSync, mkdirSync, readFileSync } from 'node:fs'; +import { dirname } from 'node:path'; + +export interface TraceHeader { + caseName: string; + agent: string; + model: string; + startedAt: string; + endedAt?: string; +} + +export interface TraceRecorder { + recordRaw(message: unknown): void; + finalize(endedAt?: string): void; +} + +export function createTraceRecorder(params: { + tracePath: string; + header: TraceHeader; +}): TraceRecorder { + mkdirSync(dirname(params.tracePath), { recursive: true }); + const headerLine = JSON.stringify({ + type: 'trace_start', + schemaVersion: 2, + ...params.header, + }); + // Write header up front so partial traces are still self-describing on crash. + writeFileSync(params.tracePath, headerLine + '\n'); + + return { + recordRaw(message: unknown) { + appendFileSync(params.tracePath, JSON.stringify(message) + '\n'); + }, + finalize(endedAt?: string) { + if (!endedAt) return; + // Rewrite header with endedAt; preserve remaining lines. + const updatedHeader = JSON.stringify({ + type: 'trace_start', + schemaVersion: 2, + ...params.header, + endedAt, + }); + const existing = readFileSync(params.tracePath, 'utf-8'); + const rest = existing.split('\n').slice(1).join('\n'); + writeFileSync(params.tracePath, updatedHeader + '\n' + rest); + }, + }; +} diff --git a/tests/acp/trace-recorder.test.ts b/tests/acp/trace-recorder.test.ts new file mode 100644 index 0000000..918795c --- /dev/null +++ b/tests/acp/trace-recorder.test.ts @@ -0,0 +1,34 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, readFileSync, rmSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { createTraceRecorder } from '../../src/workbench/acp/trace-recorder.js'; + +test('createTraceRecorder writes header then verbatim messages', () => { + const dir = mkdtempSync(join(tmpdir(), 'trace-rec-')); + const path = join(dir, 'trace.jsonl'); + const rec = createTraceRecorder({ + tracePath: path, + header: { + caseName: 'x', + agent: 'claude-agent-acp', + model: 'claude-haiku-4-5', + startedAt: '2026-05-25T00:00:00Z', + }, + }); + rec.recordRaw({ jsonrpc: '2.0', id: 1, method: 'initialize', params: {} }); + rec.recordRaw({ jsonrpc: '2.0', id: 1, result: {} }); + rec.finalize('2026-05-25T00:00:01Z'); + + const lines = readFileSync(path, 'utf-8').trim().split('\n'); + assert.equal(lines.length, 3); + const header = JSON.parse(lines[0]); + assert.equal(header.type, 'trace_start'); + assert.equal(header.agent, 'claude-agent-acp'); + assert.equal(header.endedAt, '2026-05-25T00:00:01Z'); + const msg = JSON.parse(lines[1]); + assert.equal(msg.method, 'initialize'); + + rmSync(dir, { recursive: true }); +}); From f420a604dee43580b70cb5c0bbaa5a8e6b74f88d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 12:56:39 -0500 Subject: [PATCH 089/121] chore(test): wire ACP trace recorder test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 513e6b8..dc490c5 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 3383084e547f0fa13048041ba40c882cb14118c7 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 13:04:04 -0500 Subject: [PATCH 090/121] feat(schema): require agent: field in case.yml, fail loud on missing/unknown Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/case-loader.ts | 11 +++++ src/workbench/docker-runner.ts | 1 + src/workbench/types.ts | 11 ++++- tests/case-loader-agent-required.test.ts | 63 ++++++++++++++++++++++++ tests/smoke-workbench-case.ts | 16 ++++++ tests/smoke-workbench-docker-runner.ts | 5 ++ tests/smoke-workbench-models.ts | 2 +- tests/smoke-workbench-suite.ts | 2 + tests/smoke-workbench-trials.ts | 1 + 9 files changed, 110 insertions(+), 2 deletions(-) create mode 100644 tests/case-loader-agent-required.test.ts diff --git a/src/workbench/case-loader.ts b/src/workbench/case-loader.ts index 5f55653..5941fbb 100644 --- a/src/workbench/case-loader.ts +++ b/src/workbench/case-loader.ts @@ -3,6 +3,7 @@ import { dirname, extname, resolve } from 'node:path'; import { parse as parseYaml } from 'yaml'; +import { resolveAgent } from './agents/registry.js'; import { ensureOpenRouterModelRef } from './models.js'; import type { ResolvedWorkbenchCase, @@ -56,6 +57,14 @@ export function resolveWorkbenchCaseConfig( throw new Error(`Workbench case ${resolvedConfigPath}: field "artifacts" is invalid; inspect outputs in the workspace or use --keep-workspace`); } + if (!parsed.agent || typeof parsed.agent !== 'string') { + throw new Error( + `Case ${resolvedConfigPath} is missing required field \`agent:\`. ` + + `Valid agents: claude-agent-acp, codex-acp, gemini, opencode, pi-acp.`, + ); + } + const agentCfg = resolveAgent(parsed.agent); + const name = requireNonEmptyString(parsed, 'name', resolvedConfigPath); const references = requireNonEmptyString(parsed, 'references', resolvedConfigPath); const task = requireNonEmptyString(parsed, 'task', resolvedConfigPath); @@ -93,7 +102,9 @@ export function resolveWorkbenchCaseConfig( env, setup, cleanup, + agent: agentCfg.name, model, + skillUnderTest: parsed.skillUnderTest as ResolvedWorkbenchCase['skillUnderTest'], timeoutSeconds, }; } diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index b72f75b..370bd80 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -254,6 +254,7 @@ function buildBundledCaseFile(params: { references: './references', task: params.source.task, graders: params.source.graders.map((grader) => ({ ...grader })), + agent: params.source.agent, model: params.modelOverride ?? params.source.model, timeoutSeconds: params.source.timeoutSeconds, }; diff --git a/src/workbench/types.ts b/src/workbench/types.ts index d80c25d..c062b33 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -3,6 +3,11 @@ export interface WorkbenchGraderConfig { command: string; } +export interface SkillUnderTestSpec { + slug: string; // e.g., "web-design-guidelines" + hostPath: string; // absolute path on host to the skill dir containing SKILL.md +} + export type WorkbenchMcpJsonValue = | string | number @@ -41,12 +46,14 @@ export interface WorkbenchCaseConfig { references: string; task: string; graders: WorkbenchGraderConfig[]; + agent: string; // required + model?: string; // semantics now agent-native + skillUnderTest?: SkillUnderTestSpec; // optional mcpServers?: WorkbenchMcpServersConfig; mcpServices?: WorkbenchMcpServicesConfig; env?: string[]; setup?: string[]; cleanup?: string[]; - model?: string; timeoutSeconds?: number; } @@ -62,7 +69,9 @@ export interface ResolvedWorkbenchCase { env: string[]; setup: string[]; cleanup: string[]; + agent: string; model: string; + skillUnderTest?: SkillUnderTestSpec; timeoutSeconds: number; } diff --git a/tests/case-loader-agent-required.test.ts b/tests/case-loader-agent-required.test.ts new file mode 100644 index 0000000..9fdd1c4 --- /dev/null +++ b/tests/case-loader-agent-required.test.ts @@ -0,0 +1,63 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { mkdtempSync, writeFileSync, rmSync, mkdirSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { loadWorkbenchCase } from '../src/workbench/case-loader.js'; + +test('loadWorkbenchCase throws when agent: is missing', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + mkdirSync(join(dir, 'refs'), { recursive: true }); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +task: do nothing +graders: + - name: noop + command: 'true' +model: openrouter/anthropic/claude-haiku-4-5 +`); + assert.throws( + () => loadWorkbenchCase(join(dir, 'case.yml')), + /agent.*required|missing.*agent/i, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchCase throws when agent: is unknown', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + mkdirSync(join(dir, 'refs'), { recursive: true }); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +agent: nonexistent-agent +model: x +task: do nothing +graders: + - name: noop + command: 'true' +`); + assert.throws( + () => loadWorkbenchCase(join(dir, 'case.yml')), + /Unknown agent/, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchCase accepts valid agent', () => { + const dir = mkdtempSync(join(tmpdir(), 'case-test-')); + mkdirSync(join(dir, 'refs'), { recursive: true }); + writeFileSync(join(dir, 'case.yml'), ` +name: test +references: ./refs +agent: pi-acp +model: openrouter/anthropic/claude-haiku-4-5 +task: do nothing +graders: + - name: noop + command: 'true' +`); + const c = loadWorkbenchCase(join(dir, 'case.yml')); + assert.equal(c.agent, 'pi-acp'); + rmSync(dir, { recursive: true }); +}); diff --git a/tests/smoke-workbench-case.ts b/tests/smoke-workbench-case.ts index 2b2d6c2..648ca42 100644 --- a/tests/smoke-workbench-case.ts +++ b/tests/smoke-workbench-case.ts @@ -25,6 +25,7 @@ test('type supports minimal fields', () => { graders: [ { name: 'merged-output', command: 'node $CASE/check.js' }, ], + agent: 'pi-acp', }; assert.equal(minimal.name, 'merge-pdfs'); @@ -37,6 +38,7 @@ test('YAML case loads and resolves relative references', () => { const casePath = writeCaseFile(root, 'case.yaml', [ 'name: merge-pdfs', 'references: ./references', + 'agent: pi-acp', 'task: Merge the PDFs in inputs/ into outputs/book.pdf.', 'graders:', ' - name: merged-output', @@ -62,6 +64,7 @@ test('JSON case loads', () => { const casePath = writeCaseFile(root, 'case.json', JSON.stringify({ name: 'merge-pdfs-json', references: './refs', + agent: 'pi-acp', task: 'Merge the PDFs.', graders: [ { name: 'merged-output', command: 'node $CASE/checks/merge-pdfs.js' }, @@ -84,6 +87,7 @@ test('YAML case loads MCP server definitions', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: mcp-docs', 'references: ./references', + 'agent: pi-acp', 'task: Use the configured MCP docs server.', 'graders:', ' - name: output', @@ -132,6 +136,7 @@ test('invalid MCP server without transport throws', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: mcp-docs', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -157,6 +162,7 @@ test('MCP service without matching MCP server throws', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: mcp-docs', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -184,6 +190,7 @@ test('invalid MCP service command reports mcpServices field', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: mcp-docs', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -212,6 +219,7 @@ test('MCP service port is rejected because mcpServers URL owns the port', () => const casePath = writeCaseFile(root, 'case.yml', [ 'name: mcp-docs', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -243,6 +251,7 @@ test('defaults are applied', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: merge-pdfs', 'references: ./references', + 'agent: pi-acp', 'task: Merge files', 'graders:', ' - name: merged-output', @@ -267,6 +276,7 @@ test('invalid case model ref is rejected while loading', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: merge-pdfs', 'references: ./references', + 'agent: pi-acp', 'task: Merge files', 'model: anthropic/claude-3-5-haiku-latest', 'graders:', @@ -288,6 +298,7 @@ test('invalid missing references throws', () => { try { const casePath = writeCaseFile(root, 'case.yaml', [ 'name: merge-pdfs', + 'agent: pi-acp', 'task: Merge files', 'graders:', ' - name: merged-output', @@ -310,6 +321,7 @@ test('invalid non-array env throws', () => { const casePath = writeCaseFile(root, 'case.json', JSON.stringify({ name: 'merge-pdfs', references: './references', + agent: 'pi-acp', task: 'Merge files', graders: [ { name: 'merged-output', command: 'node $CASE/checks/merge-pdfs.js' }, @@ -335,6 +347,7 @@ test('invalid env variable names are rejected', () => { const casePath = writeCaseFile(root, `case-${envName.length}.json`, JSON.stringify({ name: 'merge-pdfs', references: './references', + agent: 'pi-acp', task: 'Merge files', graders: [ { name: 'merged-output', command: 'node $CASE/checks/merge-pdfs.js' }, @@ -359,6 +372,7 @@ test('valid env variable names still load', () => { const casePath = writeCaseFile(root, 'case.json', JSON.stringify({ name: 'merge-pdfs', references: './references', + agent: 'pi-acp', task: 'Merge files', graders: [ { name: 'merged-output', command: 'node $CASE/checks/merge-pdfs.js' }, @@ -380,6 +394,7 @@ test('invalid missing graders throws', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: merge-pdfs', 'references: ./references', + 'agent: pi-acp', 'task: Merge files', ].join('\n')); @@ -446,6 +461,7 @@ test('invalid grader command throws', () => { const casePath = writeCaseFile(root, 'case.yml', [ 'name: merge-pdfs', 'references: ./references', + 'agent: pi-acp', 'task: Merge files', 'graders:', ' - name: merged-output', diff --git a/tests/smoke-workbench-docker-runner.ts b/tests/smoke-workbench-docker-runner.ts index acdf61c..644b876 100644 --- a/tests/smoke-workbench-docker-runner.ts +++ b/tests/smoke-workbench-docker-runner.ts @@ -48,6 +48,7 @@ test('prepareDockerWorkbenchRun writes results under case .results and keeps bun writeFileSync(join(sourceCaseDir, 'case.yml'), [ 'name: pdf-merge', 'references: ./references', + 'agent: pi-acp', 'task: Merge PDFs.', 'graders:', ' - name: merged-output', @@ -98,6 +99,7 @@ test('prepareDockerWorkbenchRun bundles case support directories', () => { writeFileSync(join(sourceCaseDir, 'case.yml'), [ 'name: support-case', 'references: ./references', + 'agent: pi-acp', 'task: Test support dirs.', 'graders:', ' - name: passes', @@ -136,6 +138,7 @@ test('prepareDockerWorkbenchRun writes isolated mcporter config for MCP servers' writeFileSync(join(sourceCaseDir, 'case.yml'), [ 'name: mcp-case', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -194,6 +197,7 @@ test('prepareDockerWorkbenchRun bundles hidden MCP service support outside work' writeFileSync(join(sourceCaseDir, 'case.yml'), [ 'name: mcp-service-case', 'references: ./references', + 'agent: pi-acp', 'task: Use MCP.', 'graders:', ' - name: output', @@ -412,6 +416,7 @@ test('prepareDockerWorkbenchRun honors --out as the results root', () => { writeFileSync(join(sourceCaseDir, 'case.yml'), [ 'name: pdf-merge', 'references: ./references', + 'agent: pi-acp', 'task: Merge PDFs.', 'graders:', ' - name: merged-output', diff --git a/tests/smoke-workbench-models.ts b/tests/smoke-workbench-models.ts index 7527321..2b08f7f 100644 --- a/tests/smoke-workbench-models.ts +++ b/tests/smoke-workbench-models.ts @@ -153,7 +153,7 @@ test('runWorkbenchCase --trials uses the case model when no model override is pr mkdirSync(outDir, { recursive: true }); writeFileSync( casePath, - 'name: model-case\nreferences: ./refs\nmodel: openrouter/openai/gpt-5.4\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', + 'name: model-case\nreferences: ./refs\nagent: pi-acp\nmodel: openrouter/openai/gpt-5.4\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8', ); diff --git a/tests/smoke-workbench-suite.ts b/tests/smoke-workbench-suite.ts index 7ea1792..57ab694 100644 --- a/tests/smoke-workbench-suite.ts +++ b/tests/smoke-workbench-suite.ts @@ -46,6 +46,7 @@ test('loadWorkbenchSuite supports inline cases with suite defaults', () => { ' - openrouter/google/gemini-2.5-flash', 'cases:', ' - name: async-parallel', + ' agent: pi-acp', ' task: Make this faster', ' graders:', ' - name: async-parallel', @@ -88,6 +89,7 @@ test('loadWorkbenchSuite applies and merges MCP defaults for inline cases', () = ' - mcp/default-server.mjs', 'cases:', ' - name: mcp-inline', + ' agent: pi-acp', ' task: Use MCP.', ' mcpServers:', ' context7:', diff --git a/tests/smoke-workbench-trials.ts b/tests/smoke-workbench-trials.ts index ea49b51..2c3cd1d 100644 --- a/tests/smoke-workbench-trials.ts +++ b/tests/smoke-workbench-trials.ts @@ -51,6 +51,7 @@ test('runWorkbenchSuite writes trial directories and case-model aggregates', asy ' - openrouter/google/gemini-2.5-flash', 'cases:', ' - name: trial-case', + ' agent: pi-acp', ' task: Test trials', ' graders:', ' - name: passes', From 15a12e494971cdea6fd88e8215fb5adc74130add Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 13:04:08 -0500 Subject: [PATCH 091/121] chore(test): wire case-loader agent-required test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index dc490c5..9535018 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From f76384005015995223ff1f612cf43633b230425e Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 13:10:16 -0500 Subject: [PATCH 092/121] feat(schema): suite.yml uses runs: matrix; legacy models: rejected Co-Authored-By: Claude Sonnet 4.6 --- src/workbench/run-suite.ts | 10 ++--- src/workbench/suite-loader.ts | 42 ++++++++++++++++---- src/workbench/types.ts | 5 +++ tests/smoke-workbench-suite.ts | 70 +++++++++++++++++++-------------- tests/smoke-workbench-trials.ts | 5 ++- tests/suite-loader-runs.test.ts | 68 ++++++++++++++++++++++++++++++++ 6 files changed, 155 insertions(+), 45 deletions(-) create mode 100644 tests/suite-loader-runs.test.ts diff --git a/src/workbench/run-suite.ts b/src/workbench/run-suite.ts index 6003523..290d8b6 100644 --- a/src/workbench/run-suite.ts +++ b/src/workbench/run-suite.ts @@ -76,11 +76,8 @@ export async function runWorkbenchSuite( deps: RunWorkbenchSuiteDeps = {}, ): Promise { const suite = loadWorkbenchSuite(params.suitePath); - const models = suite.models; + const runs = suite.runs; const trials = params.trials ?? 1; - if (models.length === 0) { - throw new Error('Workbench suite requires at least one model in suite.yml via the suite "models" field'); - } const dockerRunner = deps.runDockerWorkbenchCase ?? runDockerWorkbenchCase; const startedAt = new Date().toISOString(); @@ -92,11 +89,11 @@ export async function runWorkbenchSuite( mkdirSync(resultsDir, { recursive: true }); - const jobs = suite.cases.flatMap((suiteCase) => models.flatMap((model) => ( + const jobs = suite.cases.flatMap((suiteCase) => runs.flatMap((run) => ( Array.from({ length: trials }, (_, index) => ({ suiteCase, caseName: suiteCase.slug, - model, + model: run.model, trial: index + 1, })) ))); @@ -125,6 +122,7 @@ export async function runWorkbenchSuite( return { ...job, trialResult }; }); + const models = runs.map((run) => run.model); const results: WorkbenchCaseModelAggregateResult[] = []; for (const suiteCase of suite.cases) { for (const model of models) { diff --git a/src/workbench/suite-loader.ts b/src/workbench/suite-loader.ts index a6813b4..84dbae5 100644 --- a/src/workbench/suite-loader.ts +++ b/src/workbench/suite-loader.ts @@ -3,9 +3,9 @@ import { basename, dirname, extname, resolve } from 'node:path'; import { parse as parseYaml } from 'yaml'; +import { resolveAgent } from './agents/registry.js'; import { readMcpServers, readMcpServices, resolveWorkbenchCaseConfig } from './case-loader.js'; -import { ensureOpenRouterModelRef } from './models.js'; -import type { ResolvedWorkbenchCase, WorkbenchMcpServersConfig, WorkbenchMcpServicesConfig } from './types.js'; +import type { ResolvedWorkbenchCase, WorkbenchMcpServersConfig, WorkbenchMcpServicesConfig, WorkbenchRunSpec } from './types.js'; import { slugPathSegment } from './utils.js'; export interface ResolvedWorkbenchSuiteCase { @@ -21,7 +21,7 @@ export interface ResolvedWorkbenchSuite { appendSystemPrompt?: string; casePaths: string[]; cases: ResolvedWorkbenchSuiteCase[]; - models: string[]; + runs: WorkbenchRunSpec[]; } export function loadWorkbenchSuite(configPath: string): ResolvedWorkbenchSuite { @@ -46,11 +46,10 @@ export function loadWorkbenchSuite(configPath: string): ResolvedWorkbenchSuite { const name = requireNonEmptyString(parsed, 'name', resolvedConfigPath); const appendSystemPrompt = readOptionalString(parsed, 'appendSystemPrompt', resolvedConfigPath); const suiteDefaults = readSuiteCaseDefaults(parsed, resolvedConfigPath); + const runs = readRunsMatrix(parsed, resolvedConfigPath); const cases = readCaseEntries(parsed, resolvedConfigPath) .map((entry, index) => resolveSuiteCase(entry, index, resolvedConfigPath, configDir, suiteDefaults)); const casePaths = cases.flatMap((suiteCase) => suiteCase.path ? [suiteCase.path] : []); - const models = readStringArray(parsed, 'models', resolvedConfigPath, true) - .map((model) => ensureOpenRouterModelRef(model)); return { configPath: resolvedConfigPath, @@ -59,7 +58,7 @@ export function loadWorkbenchSuite(configPath: string): ResolvedWorkbenchSuite { appendSystemPrompt, casePaths, cases, - models, + runs, }; } @@ -194,7 +193,7 @@ function requireNonEmptyString( function readStringArray( parsed: Record, - field: 'models' | 'env' | 'setup' | 'cleanup', + field: 'env' | 'setup' | 'cleanup', configPath: string, optional = false, ): string[] { @@ -261,3 +260,32 @@ function readOptionalTimeoutSeconds(parsed: Record, configPath: } return value; } + +function readRunsMatrix(parsed: Record, configPath: string): WorkbenchRunSpec[] { + if ('models' in parsed && !('runs' in parsed)) { + throw new Error( + `Suite ${configPath}: 'runs:' has replaced 'models:'. ` + + `Replace 'models:' with a 'runs:' list of { agent, model } objects.`, + ); + } + if (!Array.isArray(parsed.runs) || parsed.runs.length === 0) { + throw new Error(`Suite ${configPath} must declare a non-empty 'runs:' list.`); + } + return parsed.runs.map((run: unknown, index: number) => { + if (!run || typeof run !== 'object' || Array.isArray(run)) { + throw new Error(`Suite ${configPath}: 'runs[${index}]' must be an object with agent: and model: fields.`); + } + const r = run as Record; + if (typeof r.agent !== 'string' || r.agent.trim() === '') { + throw new Error(`Each run in ${configPath} must have agent: and model: fields.`); + } + if (typeof r.model !== 'string' || r.model.trim() === '') { + throw new Error(`Each run in ${configPath} must have agent: and model: fields.`); + } + const resolvedAgent = resolveAgent(r.agent.trim()); + return { + agent: resolvedAgent.name, + model: r.model.trim(), + }; + }); +} diff --git a/src/workbench/types.ts b/src/workbench/types.ts index c062b33..1824fca 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -3,6 +3,11 @@ export interface WorkbenchGraderConfig { command: string; } +export interface WorkbenchRunSpec { + agent: string; + model: string; +} + export interface SkillUnderTestSpec { slug: string; // e.g., "web-design-guidelines" hostPath: string; // absolute path on host to the skill dir containing SKILL.md diff --git a/tests/smoke-workbench-suite.ts b/tests/smoke-workbench-suite.ts index 57ab694..fe7dee6 100644 --- a/tests/smoke-workbench-suite.ts +++ b/tests/smoke-workbench-suite.ts @@ -4,18 +4,19 @@ import { tmpdir } from 'node:os'; import { join } from 'node:path'; import { test } from 'node:test'; -import { loadWorkbenchSuite } from '../src/workbench/suite-loader.js'; import { runWorkbenchSuite, runWorkbenchSuiteFromCli } from '../src/workbench/run-suite.js'; +import { loadWorkbenchSuite } from '../src/workbench/suite-loader.js'; -test('loadWorkbenchSuite resolves case paths and validates models', () => { +test('loadWorkbenchSuite resolves case paths and validates runs', () => { const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-suite-load-')); try { const suitePath = join(root, 'suite.yml'); mkdirSync(join(root, 'cases', 'missing-index'), { recursive: true }); writeFileSync(suitePath, [ 'name: supabase-postgres-best-practices', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - cases/missing-index/case.yml', ].join('\n'), 'utf-8'); @@ -23,7 +24,7 @@ test('loadWorkbenchSuite resolves case paths and validates models', () => { const suite = loadWorkbenchSuite(suitePath); assert.equal(suite.name, 'supabase-postgres-best-practices'); - assert.deepEqual(suite.models, ['openrouter/google/gemini-2.5-flash']); + assert.deepEqual(suite.runs, [{ agent: 'pi-acp', model: 'openrouter/google/gemini-2.5-flash' }]); assert.deepEqual(suite.casePaths, [join(root, 'cases', 'missing-index', 'case.yml')]); } finally { rmSync(root, { recursive: true, force: true }); @@ -42,8 +43,9 @@ test('loadWorkbenchSuite supports inline cases with suite defaults', () => { 'env:', ' - OPENROUTER_API_KEY', 'timeoutSeconds: 123', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - name: async-parallel', ' agent: pi-acp', @@ -78,8 +80,9 @@ test('loadWorkbenchSuite applies and merges MCP defaults for inline cases', () = writeFileSync(suitePath, [ 'name: mcp-suite', 'references: ./references', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'mcpServers:', ' context7:', ' baseUrl: https://mcp.context7.com/mcp', @@ -126,8 +129,9 @@ test('loadWorkbenchSuite reads suite appendSystemPrompt', () => { 'name: prompted-suite', 'appendSystemPrompt: |', ' Prefer simple shell commands when possible.', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - cases/noop/case.yml', ].join('\n'), 'utf-8'); @@ -146,8 +150,9 @@ test('loadWorkbenchSuite rejects suite artifacts defaults', () => { const suitePath = join(root, 'suite.yml'); writeFileSync(suitePath, [ 'name: artifact-suite', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'artifacts:', ' - output.json', 'cases:', @@ -177,9 +182,11 @@ test('runWorkbenchSuite writes case-model matrix aggregate output', async () => writeFileSync(caseB, 'name: partial-index\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); writeFileSync(suitePath, [ 'name: supabase-postgres-best-practices', - 'models:', - ' - openrouter/google/gemini-2.5-flash', - ' - openrouter/openai/gpt-5.4', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', + ' - agent: pi-acp', + ' model: openrouter/openai/gpt-5.4', 'cases:', ' - cases/missing-index/case.yml', ' - cases/partial-index/case.yml', @@ -248,8 +255,9 @@ test('runWorkbenchSuite passes suite appendSystemPrompt to every trial', async ( 'name: prompted-suite', 'appendSystemPrompt: |', ' Prefer simple shell commands when possible.', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - cases/prompted/case.yml', ].join('\n'), 'utf-8'); @@ -295,24 +303,24 @@ test('runWorkbenchSuiteFromCli rejects model overrides because suites own models ); }); -test('runWorkbenchSuite missing models error references suite models only', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-suite-missing-models-')); +test('runWorkbenchSuite missing runs error references suite runs only', async () => { + const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-suite-missing-runs-')); try { const suitePath = join(root, 'suite.yml'); - const casePath = join(root, 'cases', 'no-models', 'case.yml'); - mkdirSync(join(root, 'cases', 'no-models'), { recursive: true }); - writeFileSync(casePath, 'name: no-models\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); + const casePath = join(root, 'cases', 'no-runs', 'case.yml'); + mkdirSync(join(root, 'cases', 'no-runs'), { recursive: true }); + writeFileSync(casePath, 'name: no-runs\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); writeFileSync(suitePath, [ - 'name: no-models-suite', + 'name: no-runs-suite', 'cases:', - ' - cases/no-models/case.yml', + ' - cases/no-runs/case.yml', ].join('\n'), 'utf-8'); await assert.rejects( () => runWorkbenchSuite({ suitePath }), (error: unknown) => { assert.ok(error instanceof Error); - assert.match(error.message, /suite\.yml|models/); + assert.match(error.message, /suite\.yml|runs/); assert.doesNotMatch(error.message, /--models/); return true; }, @@ -332,8 +340,9 @@ test('runWorkbenchSuite rejects non-integer programmatic concurrency', async () writeFileSync(casePath, 'name: invalid-concurrency\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); writeFileSync(suitePath, [ 'name: invalid-concurrency-suite', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - cases/invalid-concurrency/case.yml', ].join('\n'), 'utf-8'); @@ -358,8 +367,9 @@ test('runWorkbenchSuite honors concurrency for independent trials', async () => writeFileSync(casePath, 'name: parallel\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); writeFileSync(suitePath, [ 'name: parallel-suite', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - cases/parallel/case.yml', ].join('\n'), 'utf-8'); diff --git a/tests/smoke-workbench-trials.ts b/tests/smoke-workbench-trials.ts index 2c3cd1d..5061b7f 100644 --- a/tests/smoke-workbench-trials.ts +++ b/tests/smoke-workbench-trials.ts @@ -47,8 +47,9 @@ test('runWorkbenchSuite writes trial directories and case-model aggregates', asy writeFileSync(suitePath, [ 'name: trial-suite', 'references: ./references', - 'models:', - ' - openrouter/google/gemini-2.5-flash', + 'runs:', + ' - agent: pi-acp', + ' model: openrouter/google/gemini-2.5-flash', 'cases:', ' - name: trial-case', ' agent: pi-acp', diff --git a/tests/suite-loader-runs.test.ts b/tests/suite-loader-runs.test.ts new file mode 100644 index 0000000..495c579 --- /dev/null +++ b/tests/suite-loader-runs.test.ts @@ -0,0 +1,68 @@ +import { strict as assert } from 'node:assert'; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { test } from 'node:test'; + +import { loadWorkbenchSuite } from '../src/workbench/suite-loader.js'; + +test('loadWorkbenchSuite parses runs: matrix', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +runs: + - agent: claude-agent-acp + model: claude-haiku-4-5 + - agent: pi-acp + model: openrouter/anthropic/claude-haiku-4-5 +cases: + - name: test-case + agent: pi-acp + task: do nothing + graders: + - name: noop + command: 'true' +`); + const suite = loadWorkbenchSuite(join(dir, 'suite.yml')); + assert.equal(suite.runs.length, 2); + assert.equal(suite.runs[0].agent, 'claude-agent-acp'); + assert.equal(suite.runs[1].agent, 'pi-acp'); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchSuite rejects legacy `models:` shape', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +models: + - openrouter/anthropic/claude-haiku-4-5 +cases: [] +`); + assert.throws( + () => loadWorkbenchSuite(join(dir, 'suite.yml')), + /runs:.*replaced.*models:|legacy.*models|models.*replaced/i, + ); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchSuite validates each run.agent', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +runs: + - agent: bogus + model: x +cases: [] +`); + assert.throws( + () => loadWorkbenchSuite(join(dir, 'suite.yml')), + /Unknown agent.*bogus/, + ); + rmSync(dir, { recursive: true }); +}); From a100661f2fa86ceff9536ef2d2442a622cb8548d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 13:10:18 -0500 Subject: [PATCH 093/121] chore(test): wire suite-loader runs test into npm test Co-Authored-By: Claude Sonnet 4.6 --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index 9535018..04de91c 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From 61e8fa35a98981605d3d35cdd14e938dd64f3557 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 14:51:06 -0500 Subject: [PATCH 094/121] refactor(metrics): thin wrapper over parse-trace; drop cost field --- src/workbench/container-runner.ts | 4 +- src/workbench/metrics.ts | 156 ++++++++---------------------- src/workbench/types.ts | 9 -- tests/smoke-workbench-metrics.ts | 119 ++++++++++------------- 4 files changed, 93 insertions(+), 195 deletions(-) diff --git a/src/workbench/container-runner.ts b/src/workbench/container-runner.ts index 742183b..58f3f65 100644 --- a/src/workbench/container-runner.ts +++ b/src/workbench/container-runner.ts @@ -3,7 +3,7 @@ import { dirname, join } from 'node:path'; import { runGraderCommands } from './check-runner.js'; import { loadWorkbenchCase } from './case-loader.js'; -import { buildWorkbenchMetrics } from './metrics.js'; +import { buildWorkbenchMetricsFromTrace } from './metrics.js'; import { createWorkbenchPiSession } from './pi-agent.js'; import { runShellCommand } from './process.js'; import { buildAgentSystemPrompt } from './sandbox.js'; @@ -539,7 +539,7 @@ async function runGradeMode(parsed: GradeRunnerArgs): Promise { endedAt: new Date().toISOString(), grade: { ...grade, - metrics: buildWorkbenchMetrics(trace), + metrics: buildWorkbenchMetricsFromTrace(tracePath), }, }); diff --git a/src/workbench/metrics.ts b/src/workbench/metrics.ts index 6c37c67..e989bf6 100644 --- a/src/workbench/metrics.ts +++ b/src/workbench/metrics.ts @@ -1,135 +1,55 @@ -import type { - WorkbenchMetrics, - WorkbenchResult, - WorkbenchTrace, - WorkbenchTraceEntry, - WorkbenchTrialSummaryFile, -} from './types.js'; +import { readFileSync } from 'node:fs'; -function emptyMetrics(): WorkbenchMetrics { +import { computeMetrics as computeFromTrace, getFailureEvidence, getFinalAssistantMessage } from './parse-trace.js'; +import type { WorkbenchMetrics, WorkbenchResult, WorkbenchTrialSummaryFile } from './types.js'; + +export function buildWorkbenchMetricsFromTrace(tracePath: string): WorkbenchMetrics { + const jsonl = readFileSync(tracePath, 'utf-8'); + const m = computeFromTrace(jsonl); return { - durationMs: 0, - turns: 0, - toolCalls: 0, - toolResults: 0, - bashCalls: 0, - readCalls: 0, - writeCalls: 0, - editCalls: 0, - tokens: { - input: 0, - output: 0, - cacheRead: 0, - cacheWrite: 0, - total: 0, - }, - cost: { - input: 0, - output: 0, - cacheRead: 0, - cacheWrite: 0, - total: 0, - }, + durationMs: m.durationMs, + turns: m.turns, + toolCalls: m.toolCalls, + toolResults: m.toolCalls, // ACP doesn't distinguish; treat each call as one result + bashCalls: m.bashCalls, + readCalls: m.readCalls, + writeCalls: m.writeCalls, + editCalls: m.editCalls, + stopReason: m.stopReason, + tokens: m.tokens, }; } -export function buildWorkbenchMetrics(trace: WorkbenchTrace): WorkbenchMetrics { - const metrics = emptyMetrics(); - const started = Date.parse(trace.startedAt); - const ended = Date.parse(trace.endedAt); - metrics.durationMs = Number.isFinite(started) && Number.isFinite(ended) - ? Math.max(0, ended - started) - : 0; - - for (const entry of trace.entries) { - if (entry.type === 'message') { - metrics.turns += 1; - if (typeof entry.stopReason === 'string') { - metrics.stopReason = entry.stopReason; - } - addUsage(metrics, entry.usage); - continue; - } - - if (entry.type === 'tool_result') { - metrics.toolResults += 1; - continue; - } - - metrics.toolCalls += 1; - if (entry.name === 'bash') metrics.bashCalls += 1; - if (entry.name === 'read') metrics.readCalls += 1; - if (entry.name === 'write') metrics.writeCalls += 1; - if (entry.name === 'edit') metrics.editCalls += 1; - } - - return metrics; -} - export function buildTrialSummary(params: { - trace: WorkbenchTrace; + tracePath: string; result: WorkbenchResult; }): WorkbenchTrialSummaryFile { - const failedGraders = params.result.graders - ?.filter((grader) => !grader.pass) - .map((grader) => grader.name) ?? []; - const metrics = params.result.metrics ?? buildWorkbenchMetrics(params.trace); - const terminalMessage = [...params.trace.entries] - .reverse() - .find((entry): entry is Extract => entry.type === 'message' && entry.role === 'assistant'); - + const jsonl = readFileSync(params.tracePath, 'utf-8'); + const metrics = params.result.metrics ?? buildWorkbenchMetricsFromTrace(params.tracePath); + const failedGraders = params.result.graders?.filter((g) => !g.pass).map((g) => g.name) ?? []; return { - finalAssistantMessage: terminalMessage?.text, + finalAssistantMessage: getFinalAssistantMessage(jsonl), failedGraders, - evidence: [...params.result.evidence], - bashCommands: extractBashCommands(params.trace), - stopReason: typeof terminalMessage?.stopReason === 'string' ? terminalMessage.stopReason : undefined, - errorMessage: terminalMessage?.errorMessage, + evidence: [...params.result.evidence, ...getFailureEvidence(jsonl)], + bashCommands: extractBashCommands(jsonl), + stopReason: metrics.stopReason, + errorMessage: undefined, metrics, }; } -function extractBashCommands(trace: WorkbenchTrace): string[] { - return trace.entries.flatMap((entry) => { - if (entry.type !== 'tool_call' || entry.name !== 'bash') { - return []; - } - - const args = entry.arguments; - if (!args || typeof args !== 'object' || Array.isArray(args)) { - return []; - } - - const command = (args as Record).command; - return typeof command === 'string' ? [command] : []; +function extractBashCommands(jsonl: string): string[] { + const out: string[] = []; + const rows = jsonl.split(/\r?\n/).filter(Boolean).map((l) => { + try { return JSON.parse(l); } catch { return null; } }); -} - -function addUsage(metrics: WorkbenchMetrics, usage: unknown): void { - if (!usage || typeof usage !== 'object' || Array.isArray(usage)) { - return; - } - - const record = usage as Record; - metrics.tokens.input += readNumber(record.input); - metrics.tokens.output += readNumber(record.output); - metrics.tokens.cacheRead += readNumber(record.cacheRead); - metrics.tokens.cacheWrite += readNumber(record.cacheWrite); - metrics.tokens.total += readNumber(record.totalTokens); - - const cost = record.cost; - if (!cost || typeof cost !== 'object' || Array.isArray(cost)) { - return; + for (const row of rows) { + if ((row as any)?.method !== 'session/update') continue; + const u = (row as any).params?.update; + if (u?.sessionUpdate === 'tool_call' && u.kind === 'execute') { + const cmd = u.content?.[0]?.content?.text ?? u.argsText; + if (typeof cmd === 'string') out.push(cmd); + } } - - const costRecord = cost as Record; - metrics.cost.input += readNumber(costRecord.input); - metrics.cost.output += readNumber(costRecord.output); - metrics.cost.cacheRead += readNumber(costRecord.cacheRead); - metrics.cost.cacheWrite += readNumber(costRecord.cacheWrite); - metrics.cost.total += readNumber(costRecord.total); -} - -function readNumber(value: unknown): number { - return typeof value === 'number' && Number.isFinite(value) ? value : 0; + return out; } diff --git a/src/workbench/types.ts b/src/workbench/types.ts index 1824fca..f8cdafa 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -115,14 +115,6 @@ export interface WorkbenchTokenMetrics { total: number; } -export interface WorkbenchCostMetrics { - input: number; - output: number; - cacheRead: number; - cacheWrite: number; - total: number; -} - export interface WorkbenchMetrics { durationMs: number; turns: number; @@ -134,7 +126,6 @@ export interface WorkbenchMetrics { editCalls: number; stopReason?: string; tokens: WorkbenchTokenMetrics; - cost: WorkbenchCostMetrics; } export interface WorkbenchTrialSummaryFile { diff --git a/tests/smoke-workbench-metrics.ts b/tests/smoke-workbench-metrics.ts index 96eeae3..6075daf 100644 --- a/tests/smoke-workbench-metrics.ts +++ b/tests/smoke-workbench-metrics.ts @@ -1,75 +1,62 @@ import assert from 'node:assert/strict'; +import { mkdtempSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; import { test } from 'node:test'; -import { buildTrialSummary, buildWorkbenchMetrics } from '../src/workbench/metrics.js'; -import type { WorkbenchResult, WorkbenchTrace } from '../src/workbench/types.js'; +import { buildTrialSummary, buildWorkbenchMetricsFromTrace } from '../src/workbench/metrics.js'; +import type { WorkbenchResult } from '../src/workbench/types.js'; -test('buildWorkbenchMetrics counts tool calls and sums usage', () => { - const trace: WorkbenchTrace = { - caseName: 'metrics-case', - model: 'openrouter/test/model', - startedAt: '2026-04-27T10:00:00.000Z', - endedAt: '2026-04-27T10:00:02.500Z', - entries: [ - { type: 'message', role: 'user', text: 'task' }, - { - type: 'message', - role: 'assistant', - usage: { - input: 10, - output: 5, - cacheRead: 2, - cacheWrite: 1, - totalTokens: 18, - cost: { input: 0.1, output: 0.2, cacheRead: 0.03, cacheWrite: 0.04, total: 0.37 }, - }, - stopReason: 'toolUse', - }, - { type: 'tool_call', name: 'bash', arguments: { command: 'npm test' } }, - { type: 'tool_call', name: 'read', arguments: { path: 'file.ts' } }, - { type: 'tool_result', name: 'bash', text: 'ok' }, - { type: 'message', role: 'assistant', text: 'done', stopReason: 'stop' }, - ], - }; +const FIXTURE_JSONL = [ + JSON.stringify({ type: 'trace_start', schemaVersion: 2, caseName: 'm', agent: 'claude-agent-acp', model: 'x', startedAt: '2026-05-25T00:00:00Z', endedAt: '2026-05-25T00:00:02.500Z' }), + JSON.stringify({ jsonrpc: '2.0', method: 'session/update', params: { sessionId: 's', update: { sessionUpdate: 'agent_message_chunk', content: { type: 'text', text: 'final answer' } } } }), + JSON.stringify({ jsonrpc: '2.0', method: 'session/update', params: { sessionId: 's', update: { sessionUpdate: 'tool_call', toolCallId: 'tc1', kind: 'execute', content: [{ type: 'content', content: { type: 'text', text: 'firecrawl search "x"' } }] } } }), + JSON.stringify({ jsonrpc: '2.0', method: 'session/update', params: { sessionId: 's', update: { sessionUpdate: 'tool_call', toolCallId: 'tc2', kind: 'read' } } }), + JSON.stringify({ jsonrpc: '2.0', id: 3, result: { stopReason: 'end_turn', usage: { inputTokens: 10, outputTokens: 5, cacheReadTokens: 2, cacheCreationTokens: 1 } } }), +].join('\n'); - const metrics = buildWorkbenchMetrics(trace); - assert.equal(metrics.durationMs, 2500); - assert.equal(metrics.turns, 3); - assert.equal(metrics.toolCalls, 2); - assert.equal(metrics.toolResults, 1); - assert.equal(metrics.bashCalls, 1); - assert.equal(metrics.readCalls, 1); - assert.equal(metrics.stopReason, 'stop'); - assert.equal(metrics.tokens.total, 18); - assert.equal(metrics.cost.total, 0.37); +function withFixture(): { path: string; cleanup: () => void } { + const dir = mkdtempSync(join(tmpdir(), 'metrics-test-')); + const path = join(dir, 'trace.jsonl'); + writeFileSync(path, FIXTURE_JSONL); + return { path, cleanup: () => rmSync(dir, { recursive: true }) }; +} + +test('buildWorkbenchMetricsFromTrace counts tool calls and sums usage', () => { + const { path, cleanup } = withFixture(); + try { + const m = buildWorkbenchMetricsFromTrace(path); + assert.equal(m.durationMs, 2500); + assert.equal(m.toolCalls, 2); + assert.equal(m.bashCalls, 1); + assert.equal(m.readCalls, 1); + assert.equal(m.stopReason, 'end_turn'); + assert.equal(m.tokens.input, 10); + assert.equal(m.tokens.output, 5); + assert.equal(m.tokens.total, 18); + // Cost field must not exist on the metrics object + assert.equal((m as unknown as { cost?: unknown }).cost, undefined); + } finally { cleanup(); } }); test('buildTrialSummary extracts final text, failed graders, and bash commands', () => { - const trace: WorkbenchTrace = { - caseName: 'summary-case', - model: 'openrouter/test/model', - startedAt: '2026-04-27T10:00:00.000Z', - endedAt: '2026-04-27T10:00:01.000Z', - entries: [ - { type: 'tool_call', name: 'bash', arguments: { command: 'firecrawl search "x"' } }, - { type: 'message', role: 'assistant', text: 'final answer', stopReason: 'stop' }, - ], - }; - const result: WorkbenchResult = { - caseName: 'summary-case', - model: 'openrouter/test/model', - pass: false, - score: 0.5, - evidence: ['missing output'], - graders: [ - { name: 'uses-tool', command: 'true', pass: true, score: 1, evidence: [] }, - { name: 'saves-output', command: 'false', pass: false, score: 0, evidence: ['missing output'] }, - ], - }; - - const summary = buildTrialSummary({ trace, result }); - assert.equal(summary.finalAssistantMessage, 'final answer'); - assert.deepEqual(summary.failedGraders, ['saves-output']); - assert.deepEqual(summary.bashCommands, ['firecrawl search "x"']); - assert.equal(summary.metrics.bashCalls, 1); + const { path, cleanup } = withFixture(); + try { + const result: WorkbenchResult = { + caseName: 'summary-case', + model: 'x', + pass: false, + score: 0.5, + evidence: ['missing output'], + graders: [ + { name: 'uses-tool', command: 'true', pass: true, score: 1, evidence: [] }, + { name: 'saves-output', command: 'false', pass: false, score: 0, evidence: ['missing output'] }, + ], + }; + const summary = buildTrialSummary({ tracePath: path, result }); + assert.equal(summary.finalAssistantMessage, 'final answer'); + assert.deepEqual(summary.failedGraders, ['saves-output']); + assert.deepEqual(summary.bashCommands, ['firecrawl search "x"']); + assert.equal(summary.metrics.bashCalls, 1); + } finally { cleanup(); } }); From 139e85a6b94594f3352b0676cc88f1fa7ae5e1c1 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 14:55:17 -0500 Subject: [PATCH 095/121] refactor(metrics): delegate to iterToolCalls; single trace read in buildTrialSummary --- src/workbench/metrics.ts | 23 +++++++++-------------- tests/smoke-workbench-metrics.ts | 1 + 2 files changed, 10 insertions(+), 14 deletions(-) diff --git a/src/workbench/metrics.ts b/src/workbench/metrics.ts index e989bf6..a204b1c 100644 --- a/src/workbench/metrics.ts +++ b/src/workbench/metrics.ts @@ -1,10 +1,9 @@ import { readFileSync } from 'node:fs'; -import { computeMetrics as computeFromTrace, getFailureEvidence, getFinalAssistantMessage } from './parse-trace.js'; +import { computeMetrics as computeFromTrace, getFailureEvidence, getFinalAssistantMessage, iterToolCalls } from './parse-trace.js'; import type { WorkbenchMetrics, WorkbenchResult, WorkbenchTrialSummaryFile } from './types.js'; -export function buildWorkbenchMetricsFromTrace(tracePath: string): WorkbenchMetrics { - const jsonl = readFileSync(tracePath, 'utf-8'); +function buildMetricsFromJsonl(jsonl: string): WorkbenchMetrics { const m = computeFromTrace(jsonl); return { durationMs: m.durationMs, @@ -20,12 +19,16 @@ export function buildWorkbenchMetricsFromTrace(tracePath: string): WorkbenchMetr }; } +export function buildWorkbenchMetricsFromTrace(tracePath: string): WorkbenchMetrics { + return buildMetricsFromJsonl(readFileSync(tracePath, 'utf-8')); +} + export function buildTrialSummary(params: { tracePath: string; result: WorkbenchResult; }): WorkbenchTrialSummaryFile { const jsonl = readFileSync(params.tracePath, 'utf-8'); - const metrics = params.result.metrics ?? buildWorkbenchMetricsFromTrace(params.tracePath); + const metrics = params.result.metrics ?? buildMetricsFromJsonl(jsonl); const failedGraders = params.result.graders?.filter((g) => !g.pass).map((g) => g.name) ?? []; return { finalAssistantMessage: getFinalAssistantMessage(jsonl), @@ -40,16 +43,8 @@ export function buildTrialSummary(params: { function extractBashCommands(jsonl: string): string[] { const out: string[] = []; - const rows = jsonl.split(/\r?\n/).filter(Boolean).map((l) => { - try { return JSON.parse(l); } catch { return null; } - }); - for (const row of rows) { - if ((row as any)?.method !== 'session/update') continue; - const u = (row as any).params?.update; - if (u?.sessionUpdate === 'tool_call' && u.kind === 'execute') { - const cmd = u.content?.[0]?.content?.text ?? u.argsText; - if (typeof cmd === 'string') out.push(cmd); - } + for (const call of iterToolCalls(jsonl)) { + if (call.kind === 'execute' && call.argsText) out.push(call.argsText); } return out; } diff --git a/tests/smoke-workbench-metrics.ts b/tests/smoke-workbench-metrics.ts index 6075daf..178d8c2 100644 --- a/tests/smoke-workbench-metrics.ts +++ b/tests/smoke-workbench-metrics.ts @@ -28,6 +28,7 @@ test('buildWorkbenchMetricsFromTrace counts tool calls and sums usage', () => { const m = buildWorkbenchMetricsFromTrace(path); assert.equal(m.durationMs, 2500); assert.equal(m.toolCalls, 2); + assert.equal(m.toolResults, m.toolCalls); assert.equal(m.bashCalls, 1); assert.equal(m.readCalls, 1); assert.equal(m.stopReason, 'end_turn'); From 65cc4b4b84c51c7acd1c1adf2ee2c910f65c840d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:01:25 -0500 Subject: [PATCH 096/121] refactor(trace): drop WorkbenchTraceEntry; trace.jsonl is raw ACP --- package.json | 2 +- src/workbench/container-runner.ts | 233 +++-------------------- src/workbench/trace.ts | 296 +---------------------------- src/workbench/types.ts | 39 +--- tests/smoke-workbench-container.ts | 241 ----------------------- tests/smoke-workbench-trace.ts | 176 ----------------- 6 files changed, 36 insertions(+), 951 deletions(-) delete mode 100644 tests/smoke-workbench-container.ts delete mode 100644 tests/smoke-workbench-trace.ts diff --git a/package.json b/package.json index 04de91c..f4fdb35 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-trace.ts && tsx tests/smoke-workbench-container.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", diff --git a/src/workbench/container-runner.ts b/src/workbench/container-runner.ts index 58f3f65..2299a10 100644 --- a/src/workbench/container-runner.ts +++ b/src/workbench/container-runner.ts @@ -7,9 +7,10 @@ import { buildWorkbenchMetricsFromTrace } from './metrics.js'; import { createWorkbenchPiSession } from './pi-agent.js'; import { runShellCommand } from './process.js'; import { buildAgentSystemPrompt } from './sandbox.js'; -import { buildWorkbenchTrace, createTraceRecorder } from './trace.js'; -import type { TraceRecorder } from './trace.js'; -import type { WorkbenchGrade, WorkbenchResult, WorkbenchTrace, WorkbenchTraceEntry } from './types.js'; +// TODO(Task 17): --agent mode is removed in Task 17, along with the legacy +// trace functions below. These imports stay only to satisfy the type system +// until then. +import type { WorkbenchGrade, WorkbenchResult } from './types.js'; import { isRecord, writeJsonFile } from './utils.js'; import { buildWorkbenchEnv, prepareWorkbenchDirectory } from './workspace.js'; @@ -160,57 +161,24 @@ export async function runAgentPromptWithTimeout( } } -export function writeBestEffortTrace(params: { +// TODO(Task 17): writeBestEffortTrace is part of the --agent mode flow which +// is removed in Task 17. Stubbed to keep typecheck green during transition. +export function writeBestEffortTrace(_params: { tracePath: string; caseName?: string; model?: string; startedAt?: string; endedAt?: string; - session?: PromptSession; - recorder?: TraceRecorder; + session?: unknown; + recorder?: unknown; }): boolean { - const messages = params.session?.state?.messages; - if (!params.caseName || !params.model || !params.startedAt) { - return false; - } - - if (params.recorder && params.recorder.events.length > 0) { - writeTraceFile(params.tracePath, params.recorder.toTrace({ - caseName: params.caseName, - model: params.model, - startedAt: params.startedAt, - endedAt: params.endedAt ?? new Date().toISOString(), - messages: messages ?? [], - })); - return true; - } - - if (!messages) { - return false; - } - - writeTraceFile(params.tracePath, buildWorkbenchTrace({ - caseName: params.caseName, - model: params.model, - startedAt: params.startedAt, - endedAt: params.endedAt ?? new Date().toISOString(), - messages, - })); - return true; + return false; } -export function writeTraceFile(tracePath: string, trace: WorkbenchTrace): void { - const header = { - type: 'trace_start', - schemaVersion: trace.schemaVersion ?? 1, - caseName: trace.caseName, - model: trace.model, - startedAt: trace.startedAt, - endedAt: trace.endedAt, - }; - const lines = [header, ...trace.entries] - .map((entry) => JSON.stringify(entry)); - writeFileSync(tracePath, `${lines.join('\n')}\n`, 'utf-8'); +// TODO(Task 17): writeTraceFile is removed with --agent mode. Stub kept for +// the brief window between Task 15 and Task 17. +export function writeTraceFile(_tracePath: string, _trace: unknown): void { + // no-op } async function runCleanupCommands( @@ -291,53 +259,8 @@ function buildResult(params: { }; } -function readTraceFile(tracePath: string, fallback: Omit): WorkbenchTrace { - try { - const raw = readFileSync(tracePath, 'utf-8'); - const trimmed = raw.trim(); - if (trimmed.startsWith('{')) { - try { - const parsed = JSON.parse(trimmed) as unknown; - if (isRecord(parsed) && Array.isArray(parsed.entries)) { - return parsed as unknown as WorkbenchTrace; - } - } catch { - // Fall through to JSONL parsing. - } - } - - const rows = trimmed.length > 0 - ? trimmed.split(/\r?\n/).flatMap((line) => { - try { - const parsed = JSON.parse(line) as unknown; - return isRecord(parsed) ? [parsed] : []; - } catch { - return []; - } - }) - : []; - const header = rows.find((row) => row.type === 'trace_start'); - const entries = rows.filter(isTraceEntry) as WorkbenchTraceEntry[]; - if (header || entries.length > 0) { - return { - schemaVersion: 1, - caseName: isRecord(header) && typeof header.caseName === 'string' ? header.caseName : fallback.caseName, - model: isRecord(header) && typeof header.model === 'string' ? header.model : fallback.model, - startedAt: isRecord(header) && typeof header.startedAt === 'string' ? header.startedAt : fallback.startedAt, - endedAt: isRecord(header) && typeof header.endedAt === 'string' ? header.endedAt : fallback.endedAt, - entries, - }; - } - } catch { - // Grade results should still be written if trace persistence failed. - } - - return { ...fallback, entries: [] }; -} - -function isTraceEntry(value: Record): boolean { - return value.type === 'message' || value.type === 'tool_call' || value.type === 'tool_result'; -} +// TODO(Task 17): readTraceFile is gone — replaced by ACP trace-recorder; +// grade mode reads metrics directly via buildWorkbenchMetricsFromTrace(). function summarizeContent(content: unknown): string | undefined { if (!Array.isArray(content)) { @@ -389,112 +312,12 @@ function logAgentSystemPrompt(systemPrompt: string): void { console.log('[agent:system_prompt_end]'); } -async function runAgentMode(parsed: AgentRunnerArgs): Promise { - const resultPath = join(parsed.resultsDir, 'result.json'); - const tracePath = join(parsed.resultsDir, 'trace.jsonl'); - let session: PromptSession | undefined; - let recorder: TraceRecorder | undefined; - let startedAt: string | undefined; - const previousWork = process.env.WORK; - const previousResults = process.env.RESULTS; - const previousMcporterConfig = process.env.MCPORTER_CONFIG; - - mkdirSync(parsed.resultsDir, { recursive: true }); - process.env.WORK = parsed.workDir; - process.env.RESULTS = parsed.resultsDir; - if (parsed.mcpConfigPath) { - process.env.MCPORTER_CONFIG = parsed.mcpConfigPath; - } else { - delete process.env.MCPORTER_CONFIG; - } - - try { - try { - startedAt = new Date().toISOString(); - const created = await createWorkbenchPiSession({ - cwd: parsed.workDir, - modelRef: parsed.model, - apiKeyEnv: 'OPENROUTER_API_KEY', - appendSystemPrompt: parsed.appendSystemPrompt, - mcpConfigPath: parsed.mcpConfigPath, - }); - session = created.session as PromptSession; - const systemPrompt = typeof session.systemPrompt === 'string' - ? session.systemPrompt - : buildAgentSystemPrompt(); - logAgentSystemPrompt(systemPrompt); - recorder = createTraceRecorder(); - const unsubscribe = session.subscribe?.((event) => { - recorder?.record(event); - logAgentEvent(event); - }); - - try { - await runAgentPromptWithTimeout(session, parsed.task, parsed.timeoutSeconds); - } finally { - unsubscribe?.(); - } - - const endedAt = new Date().toISOString(); - const trace = recorder.toTrace({ - caseName: parsed.caseName, - model: parsed.model, - startedAt, - endedAt, - messages: session.state?.messages ?? [], - }); - trace.entries.unshift({ - type: 'message', - role: 'system', - text: systemPrompt, - timestamp: startedAt, - }); - - writeTraceFile(tracePath, trace); - return 0; - } catch (error) { - const endedAt = new Date().toISOString(); - try { - writeBestEffortTrace({ - tracePath, - caseName: parsed.caseName, - model: parsed.model, - startedAt, - endedAt, - session, - recorder, - }); - } catch { - // Fatal result writing is more important than partial trace persistence. - } - writeJsonFile(resultPath, { - caseName: parsed.caseName, - model: parsed.model, - endedAt, - ...buildFatalGrade(error), - error: error instanceof Error ? error.message : String(error), - }); - return 1; - } - } finally { - if (previousWork === undefined) { - delete process.env.WORK; - } else { - process.env.WORK = previousWork; - } - - if (previousResults === undefined) { - delete process.env.RESULTS; - } else { - process.env.RESULTS = previousResults; - } - - if (previousMcporterConfig === undefined) { - delete process.env.MCPORTER_CONFIG; - } else { - process.env.MCPORTER_CONFIG = previousMcporterConfig; - } - } +// TODO(Task 17): --agent mode is removed entirely in Task 17. The body +// previously here drove a Pi session, normalized events into the legacy +// WorkbenchTrace shape, and wrote out trace.jsonl. ACP-based agents replace +// all of it. Stubbed to keep typecheck green during the transition. +async function runAgentMode(_parsed: AgentRunnerArgs): Promise { + throw new Error('container-runner --agent mode is disabled; ACP agents replace it in Task 17'); } async function runSetupMode(parsed: SetupRunnerArgs): Promise { @@ -518,13 +341,9 @@ async function runGradeMode(parsed: GradeRunnerArgs): Promise { const cleanupErrorPath = join(parsed.resultsDir, 'cleanup-error.txt'); const resolved = loadWorkbenchCase(parsed.casePath); const env = buildContainerWorkbenchEnv(parsed); - const now = new Date().toISOString(); - const trace = readTraceFile(tracePath, { - caseName: resolved.name, - model: resolved.model, - startedAt: now, - endedAt: now, - }); + const startedAt = new Date().toISOString(); + // TODO(Task 17): grade mode no longer reads the trace file; metrics come + // straight from buildWorkbenchMetricsFromTrace(tracePath) below. try { const grade = await runGraderCommands(resolved.graders, { @@ -535,7 +354,7 @@ async function runGradeMode(parsed: GradeRunnerArgs): Promise { const result = buildResult({ caseName: resolved.name, model: resolved.model, - startedAt: trace.startedAt, + startedAt, endedAt: new Date().toISOString(), grade: { ...grade, diff --git a/src/workbench/trace.ts b/src/workbench/trace.ts index 3ea5916..64df84c 100644 --- a/src/workbench/trace.ts +++ b/src/workbench/trace.ts @@ -1,290 +1,6 @@ -import type { WorkbenchTrace, WorkbenchTraceEntry, WorkbenchTraceEvent } from './types.js'; -import { isRecord } from './utils.js'; - -export interface TraceRecorder { - events: WorkbenchTraceEvent[]; - record(event: unknown): void; - toTrace(params: { - caseName: string; - model: string; - startedAt: string; - endedAt: string; - messages?: unknown[]; - }): WorkbenchTrace; -} - -export function createTraceCollector(): { record(event: unknown): void; events: unknown[] } { - const events: unknown[] = []; - return { - events, - record(event: unknown) { - events.push(event); - }, - }; -} - -export function createTraceRecorder(options: { now?: () => string } = {}): TraceRecorder { - const now = options.now ?? (() => new Date().toISOString()); - const events: WorkbenchTraceEvent[] = []; - - return { - events, - record(event: unknown) { - events.push(normalizeTraceEvent(event, now())); - }, - toTrace(params) { - const eventEntries = normalizeEvents(events); - const entries = eventEntries.length > 0 - ? mergeSessionMessages(eventEntries, params.messages ?? []) - : normalizeMessages(params.messages ?? []); - return { - schemaVersion: 1, - caseName: params.caseName, - model: params.model, - startedAt: params.startedAt, - endedAt: params.endedAt, - events: [...events], - entries, - }; - }, - }; -} - -export function buildWorkbenchTrace(params: { - caseName: string; - model: string; - startedAt: string; - endedAt: string; - messages: unknown[]; -}): WorkbenchTrace { - return { - caseName: params.caseName, - model: params.model, - startedAt: params.startedAt, - endedAt: params.endedAt, - entries: normalizeMessages(params.messages), - }; -} - -function normalizeTraceEvent(event: unknown, timestamp: string): WorkbenchTraceEvent { - if (!isRecord(event) || typeof event.type !== 'string') { - return { type: 'unknown', timestamp, value: toJsonSafe(event) }; - } - - const normalized: WorkbenchTraceEvent = { type: event.type, timestamp }; - for (const [key, value] of Object.entries(event)) { - if (key === 'type') { - continue; - } - const safeValue = toJsonSafe(value); - if (safeValue !== undefined) { - normalized[key] = safeValue; - } - } - return normalized; -} - -function normalizeEvents(events: WorkbenchTraceEvent[]): WorkbenchTraceEntry[] { - const entries: WorkbenchTraceEntry[] = []; - - for (const event of events) { - if (event.type === 'message_end' && isRecord(event.message)) { - const messageEntry = normalizeMessageOnly(event.message, event.timestamp); - if (messageEntry) { - entries.push(messageEntry); - } - continue; - } - - if (event.type === 'tool_execution_start') { - entries.push({ - type: 'tool_call', - id: typeof event.toolCallId === 'string' ? event.toolCallId : undefined, - name: typeof event.toolName === 'string' ? event.toolName : 'unknown', - arguments: event.args, - timestamp: event.timestamp, - }); - continue; - } - - if (event.type === 'tool_execution_end') { - entries.push({ - type: 'tool_result', - id: typeof event.toolCallId === 'string' ? event.toolCallId : undefined, - name: typeof event.toolName === 'string' ? event.toolName : undefined, - text: extractToolEventText(event.result), - isError: typeof event.isError === 'boolean' ? event.isError : undefined, - timestamp: event.timestamp, - }); - } - } - - return entries; -} - -function mergeSessionMessages(eventEntries: WorkbenchTraceEntry[], messages: unknown[]): WorkbenchTraceEntry[] { - const sessionMessages = normalizeMessages(messages) - .filter((entry): entry is Extract => entry.type === 'message'); - const missingSessionMessages = sessionMessages.filter((message) => !eventEntries.some((entry) => sameMessageEntry(entry, message))); - return [...missingSessionMessages, ...eventEntries]; -} - -function sameMessageEntry(left: WorkbenchTraceEntry, right: Extract): boolean { - if (left.type !== 'message') { - return false; - } - return left.role === right.role - && left.text === right.text - && left.thinking === right.thinking - && left.stopReason === right.stopReason - && left.errorMessage === right.errorMessage; -} - -function normalizeMessageOnly(message: Record, timestamp: string): WorkbenchTraceEntry | undefined { - const role = typeof message.role === 'string' ? message.role : 'unknown'; - if (role === 'toolResult') { - return undefined; - } - - const content = Array.isArray(message.content) ? message.content : []; - const text = extractContentByType(content, 'text', 'text'); - const thinking = extractContentByType(content, 'thinking', 'thinking'); - const hasTerminalMetadata = typeof message.stopReason === 'string' || typeof message.errorMessage === 'string'; - if (text.length === 0 && thinking.length === 0 && role === 'assistant' && !hasTerminalMetadata) { - return undefined; - } - - return { - type: 'message', - role, - text: text.length > 0 ? text : undefined, - thinking: thinking.length > 0 ? thinking : undefined, - timestamp, - usage: message.usage, - stopReason: message.stopReason, - errorMessage: typeof message.errorMessage === 'string' ? message.errorMessage : undefined, - }; -} - -function normalizeMessages(messages: unknown[]): WorkbenchTraceEntry[] { - const entries: WorkbenchTraceEntry[] = []; - - for (const message of messages) { - if (!isRecord(message)) { - continue; - } - - const role = typeof message.role === 'string' ? message.role : 'unknown'; - const timestamp = message.timestamp; - - if (role === 'toolResult') { - entries.push({ - type: 'tool_result', - id: typeof message.toolCallId === 'string' ? message.toolCallId : undefined, - name: typeof message.toolName === 'string' ? message.toolName : undefined, - text: extractText(message.content), - isError: typeof message.isError === 'boolean' ? message.isError : undefined, - timestamp, - }); - continue; - } - - const content = Array.isArray(message.content) ? message.content : []; - const text = extractContentByType(content, 'text', 'text'); - const thinking = extractContentByType(content, 'thinking', 'thinking'); - - const hasTerminalMetadata = typeof message.stopReason === 'string' || typeof message.errorMessage === 'string'; - if (text.length > 0 || thinking.length > 0 || role !== 'assistant' || hasTerminalMetadata) { - entries.push({ - type: 'message', - role, - text: text.length > 0 ? text : undefined, - thinking: thinking.length > 0 ? thinking : undefined, - timestamp, - usage: message.usage, - stopReason: message.stopReason, - errorMessage: typeof message.errorMessage === 'string' ? message.errorMessage : undefined, - }); - } - - for (const item of content) { - if (!isRecord(item) || item.type !== 'toolCall') { - continue; - } - - entries.push({ - type: 'tool_call', - id: typeof item.id === 'string' ? item.id : undefined, - name: typeof item.name === 'string' ? item.name : 'unknown', - arguments: item.arguments, - timestamp, - }); - } - } - - return entries; -} - -function extractContentByType(content: unknown[], type: string, field: string): string { - return content - .map((item) => { - if (!isRecord(item) || item.type !== type) { - return ''; - } - const value = item[field]; - return typeof value === 'string' ? value : ''; - }) - .filter((value) => value.length > 0) - .join('\n'); -} - -function extractText(content: unknown): string | undefined { - if (typeof content === 'string') { - return content; - } - - if (!Array.isArray(content)) { - return undefined; - } - - const text = extractContentByType(content, 'text', 'text'); - return text.length > 0 ? text : undefined; -} - -function extractToolEventText(result: unknown): string | undefined { - if (!isRecord(result)) { - return undefined; - } - return extractText(result.content); -} - -function toJsonSafe(value: unknown, seen = new WeakSet(), depth = 0): unknown { - if (value === null || typeof value === 'string' || typeof value === 'number' || typeof value === 'boolean') { - return value; - } - if (value === undefined || typeof value === 'function' || typeof value === 'symbol') { - return undefined; - } - if (depth > 8) { - return '[MaxDepth]'; - } - if (Array.isArray(value)) { - return value.map((item) => toJsonSafe(item, seen, depth + 1)); - } - if (typeof value === 'object') { - if (seen.has(value)) { - return '[Circular]'; - } - seen.add(value); - const record: Record = {}; - for (const [key, item] of Object.entries(value)) { - const safeItem = toJsonSafe(item, seen, depth + 1); - if (safeItem !== undefined) { - record[key] = safeItem; - } - } - seen.delete(value); - return record; - } - return String(value); -} +// Re-export trace-recorder as the public API. +// The legacy normalization layer (buildWorkbenchTrace, createTraceCollector, +// the entries[]-shaped WorkbenchTrace) is gone — trace.jsonl is raw ACP wire +// format now, captured by createTraceRecorder from ./acp/trace-recorder.js. +export { createTraceRecorder } from './acp/trace-recorder.js'; +export type { TraceHeader, TraceRecorder } from './acp/trace-recorder.js'; diff --git a/src/workbench/types.ts b/src/workbench/types.ts index f8cdafa..c371929 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -194,45 +194,12 @@ export interface RunSuiteAggregateResultFile { results: WorkbenchCaseModelAggregateResult[]; } -export type WorkbenchTraceEntry = - | { - type: 'message'; - role: string; - text?: string; - thinking?: string; - timestamp?: unknown; - usage?: unknown; - stopReason?: unknown; - errorMessage?: string; - } - | { - type: 'tool_call'; - id?: string; - name: string; - arguments?: unknown; - timestamp?: unknown; - } - | { - type: 'tool_result'; - id?: string; - name?: string; - text?: string; - isError?: boolean; - timestamp?: unknown; - }; - -export interface WorkbenchTraceEvent { - type: string; - timestamp: string; - [key: string]: unknown; -} - export interface WorkbenchTrace { - schemaVersion?: 1; + schemaVersion?: 2; // bumped from 1 (raw ACP format) caseName: string; + agent: string; // NEW model: string; startedAt: string; endedAt: string; - events?: WorkbenchTraceEvent[]; - entries: WorkbenchTraceEntry[]; + // No more entries[]; the trace.jsonl file IS the source of truth. } diff --git a/tests/smoke-workbench-container.ts b/tests/smoke-workbench-container.ts deleted file mode 100644 index b6f9362..0000000 --- a/tests/smoke-workbench-container.ts +++ /dev/null @@ -1,241 +0,0 @@ -import assert from 'node:assert/strict'; -import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; -import { tmpdir } from 'node:os'; -import { join } from 'node:path'; -import { test } from 'node:test'; - -import { - buildAgentSystemPrompt, - buildContainerWorkbenchEnv, - parseContainerRunnerArgs, - prepareWorkbenchDirectory, - runContainerWorkbenchCase, - runAgentPromptWithTimeout, - writeBestEffortTrace, -} from '../src/workbench/container-runner.js'; -import { createTraceRecorder } from '../src/workbench/trace.js'; - -test('buildAgentSystemPrompt describes operating constraints without eval/sandbox hints', () => { - const prompt = buildAgentSystemPrompt(); - - assert.match(prompt, /Current working directory is \/work/); - assert.match(prompt, /Do not use global pip installs/); - assert.match(prompt, /python -m venv \/work\/\.venv/); - assert.match(prompt, /Write all outputs under \/work/); - assert.doesNotMatch(prompt, /sandbox/i); - assert.doesNotMatch(prompt, /skill\/reference/i); - assert.doesNotMatch(prompt, /grader/i); - assert.doesNotMatch(prompt, /expected answer/i); - assert.doesNotMatch(prompt, /suite metadata/i); - assert.doesNotMatch(prompt, /\/case/); - assert.doesNotMatch(prompt, /Task:/); -}); - -test('buildContainerWorkbenchEnv exposes CASE as the mounted case directory', () => { - const env = buildContainerWorkbenchEnv({ - casePath: '/case/case.yml', - workDir: '/work', - resultsDir: '/results', - baseEnv: {}, - }); - - assert.equal(env.CASE, '/case'); - assert.equal(env.WORK, '/work'); - assert.equal(env.RESULTS, '/results'); -}); - -test('buildContainerWorkbenchEnv prepends work and case bin to PATH when present', () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-env-')); - try { - const caseDir = join(root, 'case'); - const workDir = join(root, 'work'); - mkdirSync(join(caseDir, 'bin'), { recursive: true }); - mkdirSync(workDir, { recursive: true }); - - const env = buildContainerWorkbenchEnv({ - casePath: join(caseDir, 'case.yml'), - workDir, - resultsDir: join(root, 'results'), - baseEnv: { PATH: '/usr/bin' }, - }); - - assert.equal(env.PATH, `${join(workDir, 'bin')}:${join(caseDir, 'bin')}:/usr/bin`); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('prepareWorkbenchDirectory copies references then optional workspace seed', () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-prepare-')); - try { - const referencesDir = join(root, 'references'); - const workspaceDir = join(root, 'workspace'); - const workDir = join(root, 'work'); - mkdirSync(referencesDir, { recursive: true }); - mkdirSync(workspaceDir, { recursive: true }); - mkdirSync(workDir, { recursive: true }); - writeFileSync(join(workDir, 'stale.txt'), 'stale\n', 'utf-8'); - writeFileSync(join(referencesDir, 'SKILL.md'), '# Skill\n', 'utf-8'); - writeFileSync(join(workspaceDir, 'seed.txt'), 'seed\n', 'utf-8'); - - prepareWorkbenchDirectory({ referencesDir, workspaceDir, workDir }); - - assert.equal(existsSync(join(workDir, 'stale.txt')), false); - assert.equal(readFileSync(join(workDir, 'SKILL.md'), 'utf-8'), '# Skill\n'); - assert.equal(readFileSync(join(workDir, 'seed.txt'), 'utf-8'), 'seed\n'); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('parseContainerRunnerArgs reads optional MCP config path for agent mode', () => { - const parsed = parseContainerRunnerArgs([ - '--agent', - '--case-name', 'mcp-case', - '--model', 'openrouter/google/gemini-2.5-flash', - '--task-base64', Buffer.from('Use MCP.', 'utf-8').toString('base64'), - '--timeout-seconds', '600', - '--work', '/work', - '--results', '/tmp/workbench-results', - '--mcp-config', '/work/mcporter.json', - ]); - - assert.equal(parsed.mode, 'agent'); - assert.equal(parsed.mcpConfigPath, '/work/mcporter.json'); -}); - -test('runContainerWorkbenchCase restores global WORK/RESULTS/MCPORTER_CONFIG after agent mode failure', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-env-restore-')); - const workDir = join(root, 'work'); - const resultsDir = join(root, 'results'); - mkdirSync(workDir, { recursive: true }); - mkdirSync(resultsDir, { recursive: true }); - - const previousWork = process.env.WORK; - const previousResults = process.env.RESULTS; - const previousMcporterConfig = process.env.MCPORTER_CONFIG; - - process.env.WORK = 'existing-work'; - process.env.RESULTS = 'existing-results'; - delete process.env.MCPORTER_CONFIG; - - try { - const exitCode = await runContainerWorkbenchCase([ - '--agent', - '--case-name', 'env-restore', - '--model', 'openrouter/google/gemini-2.5-flash', - '--task-base64', Buffer.from('test', 'utf-8').toString('base64'), - '--timeout-seconds', '1', - '--work', workDir, - '--results', resultsDir, - '--mcp-config', join(root, 'missing-mcporter-config.json'), - ]); - - assert.equal(exitCode, 1); - assert.equal(process.env.WORK, 'existing-work'); - assert.equal(process.env.RESULTS, 'existing-results'); - assert.equal(process.env.MCPORTER_CONFIG, undefined); - } finally { - if (previousWork === undefined) { - delete process.env.WORK; - } else { - process.env.WORK = previousWork; - } - - if (previousResults === undefined) { - delete process.env.RESULTS; - } else { - process.env.RESULTS = previousResults; - } - - if (previousMcporterConfig === undefined) { - delete process.env.MCPORTER_CONFIG; - } else { - process.env.MCPORTER_CONFIG = previousMcporterConfig; - } - rmSync(root, { recursive: true, force: true }); - } -}); - -test('runAgentPromptWithTimeout rejects when agent exceeds timeout', async () => { - await assert.rejects( - runAgentPromptWithTimeout({ prompt: () => new Promise(() => {}) }, 'task', 0.001), - /Agent timed out after 0.001 seconds/, - ); -}); - -test('runAgentPromptWithTimeout rejects when agent ends with provider error', async () => { - await assert.rejects( - runAgentPromptWithTimeout({ - prompt: async () => undefined, - state: { - messages: [ - { role: 'assistant', content: [], stopReason: 'error', errorMessage: 'Upstream request failed' }, - ], - }, - }, 'task', 1), - /Upstream request failed/, - ); -}); - -test('writeBestEffortTrace writes trace from available session messages', () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-trace-')); - try { - const tracePath = join(root, 'trace.jsonl'); - const wrote = writeBestEffortTrace({ - tracePath, - caseName: 'partial-trace', - model: 'openrouter/test/model', - startedAt: '2026-04-27T10:11:12.000Z', - endedAt: '2026-04-27T10:11:13.000Z', - session: { - state: { - messages: [ - { role: 'user', content: [{ type: 'text', text: 'hello' }] }, - ], - }, - }, - }); - - assert.equal(wrote, true); - const lines = readFileSync(tracePath, 'utf-8').trim().split('\n').map((line) => JSON.parse(line) as { type: string; caseName?: string }); - assert.equal(lines[0]?.caseName, 'partial-trace'); - assert.equal(lines.filter((line) => line.type === 'message').length, 1); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('writeBestEffortTrace prefers recorded Pi events when available', () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-event-trace-')); - try { - const tracePath = join(root, 'trace.jsonl'); - const recorder = createTraceRecorder({ now: () => '2026-04-27T10:11:12.500Z' }); - recorder.record({ - type: 'tool_execution_start', - toolCallId: 'call-1', - toolName: 'bash', - args: { command: 'npm test' }, - }); - - const wrote = writeBestEffortTrace({ - tracePath, - caseName: 'event-trace', - model: 'openrouter/test/model', - startedAt: '2026-04-27T10:11:12.000Z', - endedAt: '2026-04-27T10:11:13.000Z', - recorder, - session: { state: { messages: [] } }, - }); - - assert.equal(wrote, true); - const lines = readFileSync(tracePath, 'utf-8').trim().split('\n').map((line) => JSON.parse(line) as { - type: string; - arguments?: { command?: string }; - }); - assert.equal(lines[1]?.type, 'tool_call'); - assert.equal(lines[1]?.arguments?.command, 'npm test'); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); diff --git a/tests/smoke-workbench-trace.ts b/tests/smoke-workbench-trace.ts deleted file mode 100644 index 7c5781b..0000000 --- a/tests/smoke-workbench-trace.ts +++ /dev/null @@ -1,176 +0,0 @@ -import { buildWorkbenchTrace, createTraceCollector, createTraceRecorder } from '../src/workbench/trace.js'; - -let passed = 0; -let failed = 0; - -async function test(name: string, fn: () => Promise | void) { - try { - await fn(); - passed++; - console.log(` ✓ ${name}`); - } catch (error: any) { - failed++; - console.log(` ✗ ${name}`); - console.log(` ${error.message}`); - } -} - -function assert(condition: boolean, message: string) { - if (!condition) throw new Error(`Assertion failed: ${message}`); -} - -function assertEqual(actual: T, expected: T, message: string) { - if (actual !== expected) { - throw new Error(`${message}: expected ${JSON.stringify(expected)}, got ${JSON.stringify(actual)}`); - } -} - -console.log('\n=== Workbench Trace Smoke Tests ===\n'); - -await test('buildWorkbenchTrace stores a deduped interaction timeline', () => { - const trace = buildWorkbenchTrace({ - caseName: 'case-1', - model: 'openrouter/google/gemini-2.5-flash', - startedAt: '2026-01-01T00:00:00.000Z', - endedAt: '2026-01-01T00:00:02.000Z', - messages: [ - { role: 'user', content: [{ type: 'text', text: 'Do the task' }] }, - { - role: 'assistant', - content: [ - { type: 'thinking', thinking: 'Need to inspect files' }, - { type: 'text', text: 'I will read the skill.' }, - { type: 'toolCall', id: 'call-1', name: 'read', arguments: { path: '/work/SKILL.md' } }, - ], - usage: { totalTokens: 10 }, - }, - { - role: 'toolResult', - toolCallId: 'call-1', - toolName: 'read', - content: [{ type: 'text', text: '# Skill' }], - isError: false, - }, - ], - }); - - assertEqual(trace.caseName, 'case-1', 'trace should preserve caseName'); - assertEqual(trace.entries.length, 4, 'trace should normalize messages into entries'); - assertEqual(trace.entries[0].type, 'message', 'first entry should be user message'); - assertEqual(trace.entries[1].type, 'message', 'second entry should be assistant message'); - assertEqual(trace.entries[2].type, 'tool_call', 'third entry should be tool call'); - assertEqual(trace.entries[3].type, 'tool_result', 'fourth entry should be tool result'); - assert(!('events' in trace), 'trace should not include raw streaming events'); - assert(!('messages' in trace), 'trace should not duplicate raw messages'); -}); - -await test('buildWorkbenchTrace preserves assistant provider error messages', () => { - const trace = buildWorkbenchTrace({ - caseName: 'case-error', - model: 'openrouter/google/gemini-2.5-flash', - startedAt: '2026-01-01T00:00:00.000Z', - endedAt: '2026-01-01T00:00:02.000Z', - messages: [ - { - role: 'assistant', - content: [], - stopReason: 'error', - errorMessage: 'Provider returned 500', - }, - ], - }); - - const entry = trace.entries[0] as { type: string; errorMessage?: string; stopReason?: unknown }; - assertEqual(entry.type, 'message', 'entry should be a message'); - assertEqual(entry.stopReason, 'error', 'entry should preserve stop reason'); - assertEqual(entry.errorMessage, 'Provider returned 500', 'entry should preserve provider error message'); -}); - -await test('createTraceCollector records arbitrary events in order', () => { - const collector = createTraceCollector(); - collector.record({ step: 1 }); - collector.record('tool-call'); - collector.record(42); - - assertEqual(collector.events.length, 3, 'collector should record all events'); - assertEqual((collector.events[0] as { step?: number }).step, 1, 'collector should preserve object payload'); - assertEqual(collector.events[1], 'tool-call', 'collector should preserve string payload'); - assertEqual(collector.events[2], 42, 'collector should preserve numeric payload'); -}); - -await test('createTraceRecorder captures Pi session events and normalized entries', () => { - const recorder = createTraceRecorder({ now: () => '2026-01-01T00:00:01.000Z' }); - - recorder.record({ - type: 'message_end', - message: { - role: 'assistant', - content: [{ type: 'text', text: 'I will run the command.' }], - stopReason: 'toolUse', - }, - }); - recorder.record({ - type: 'tool_execution_start', - toolCallId: 'call-1', - toolName: 'bash', - args: { command: 'firecrawl search browser --scrape' }, - }); - recorder.record({ - type: 'tool_execution_end', - toolCallId: 'call-1', - toolName: 'bash', - result: { content: [{ type: 'text', text: 'ok' }] }, - isError: false, - }); - - const trace = recorder.toTrace({ - caseName: 'case-events', - model: 'openrouter/test/model', - startedAt: '2026-01-01T00:00:00.000Z', - endedAt: '2026-01-01T00:00:02.000Z', - }); - - assertEqual(trace.events?.length, 3, 'trace should preserve raw-ish Pi events'); - assertEqual(trace.events?.[0]?.timestamp, '2026-01-01T00:00:01.000Z', 'trace events should have capture timestamps'); - assertEqual(trace.entries.length, 3, 'trace should derive normalized entries from events'); - assertEqual(trace.entries[0].type, 'message', 'first entry should be assistant message'); - assertEqual(trace.entries[1].type, 'tool_call', 'second entry should be tool call'); - assertEqual(trace.entries[2].type, 'tool_result', 'third entry should be tool result'); - assertEqual( - ((trace.entries[1] as { arguments?: { command?: string } }).arguments)?.command, - 'firecrawl search browser --scrape', - 'tool call entry should preserve bash command', - ); -}); - -await test('createTraceRecorder preserves session messages when events are partial', () => { - const recorder = createTraceRecorder({ now: () => '2026-01-01T00:00:01.000Z' }); - - recorder.record({ - type: 'tool_execution_start', - toolCallId: 'call-1', - toolName: 'bash', - args: { command: 'node parse-pdf.mjs' }, - }); - - const trace = recorder.toTrace({ - caseName: 'partial-events', - model: 'openrouter/test/model', - startedAt: '2026-01-01T00:00:00.000Z', - endedAt: '2026-01-01T00:00:02.000Z', - messages: [ - { role: 'user', content: [{ type: 'text', text: 'Extract the PDF facts.' }] }, - { role: 'assistant', content: [{ type: 'text', text: 'I will parse the PDF.' }] }, - ], - }); - - assertEqual(trace.entries.length, 3, 'trace should keep session messages plus partial event entries'); - assertEqual(trace.entries[0].type, 'message', 'first entry should be a session message'); - assertEqual((trace.entries[0] as { role?: string }).role, 'user', 'first session message should be user'); - assertEqual(trace.entries[1].type, 'message', 'second entry should be a session message'); - assertEqual((trace.entries[1] as { role?: string }).role, 'assistant', 'second session message should be assistant'); - assertEqual(trace.entries[2].type, 'tool_call', 'partial tool event should still be included'); -}); - -console.log(`\n${passed} passed, ${failed} failed\n`); -process.exit(failed > 0 ? 1 : 0); From 0053d2d4ab27745ac6025695444a7a41b4bdc7fb Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:08:44 -0500 Subject: [PATCH 097/121] feat(docker-runner): add runOneAcpTrial helper for host-side ACP orchestration --- src/workbench/docker-runner.ts | 204 +++++++++++++++++++++++++++++++++ 1 file changed, 204 insertions(+) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index 370bd80..ffedf23 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -1,3 +1,4 @@ +import { spawn } from 'node:child_process'; import { chmodSync, cpSync, existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; import { tmpdir } from 'node:os'; import { dirname, join, resolve } from 'node:path'; @@ -5,6 +6,13 @@ import { fileURLToPath } from 'node:url'; import { stringify as stringifyYaml } from 'yaml'; +import { resolveAuth, resolveContainerPath } from './acp/auth.js'; +import { createWorkbenchClient } from './acp/client.js'; +import { writeMcpConfig } from './acp/mcp-config-writer.js'; +import { computeSkillMount, dockerMountFlag } from './acp/skill-deploy.js'; +import { createTraceRecorder } from './acp/trace-recorder.js'; +import { createDockerExecStream } from './acp/transport.js'; +import { resolveAgent } from './agents/registry.js'; import { loadWorkbenchCase } from './case-loader.js'; import { MCPORTER_CONFIG_CONTAINER_PATH, writeWorkbenchMcpConfig } from './mcp/index.js'; import { runShellCommand } from './process.js'; @@ -631,3 +639,199 @@ export async function runDockerWorkbenchCase( prepared.cleanup(); } } + +export interface RunOneAcpTrialOptions { + resolvedCase: ResolvedWorkbenchCase; + tempDir: string; + workDir: string; + caseDir: string; + resultsDir: string; + agentName: string; + model: string; + image: string; + skillMountParams: { hostSkillDir: string; skillSlug: string } | null; + repoRoot: string; + timeoutSeconds: number; +} + +export interface RunOneAcpTrialResult { + pass: boolean; + tracePath: string; + resultPath: string; +} + +export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise { + const agent = resolveAgent(params.agentName); + const auth = resolveAuth(agent, { env: process.env }); + + // Stage host-side $HOME for the agent: auth files + MCP config land here. + const agentHomeHost = join(params.tempDir, 'agent-home'); + mkdirSync(agentHomeHost, { recursive: true }); + + if (auth.mode === 'subscription') { + for (const f of auth.files) { + const targetPath = resolveContainerPath(f.containerPath, agentHomeHost); + mkdirSync(dirname(targetPath), { recursive: true }); + cpSync(f.hostPath, targetPath); + } + } + + writeMcpConfig({ + agent, + caseConfig: { + mcpServers: params.resolvedCase.mcpServers as Parameters[0]['caseConfig']['mcpServers'], + }, + agentHomeOnHost: agentHomeHost, + }); + + // Skill mount (optional) + let skillMountFlag = ''; + if (params.skillMountParams) { + const mount = computeSkillMount({ + agent, + skillSlug: params.skillMountParams.skillSlug, + hostSkillDir: params.skillMountParams.hostSkillDir, + agentHome: '/home/agent', + }); + skillMountFlag = dockerMountFlag(mount); + } + + // Env passthrough + const envFlags: string[] = []; + if (auth.mode === 'env') { + for (const name of auth.envNames) { + if (process.env[name] !== undefined) envFlags.push(`-e ${name}`); + } + } + for (const name of params.resolvedCase.env) { + if (process.env[name] !== undefined) envFlags.push(`-e ${name}`); + } + + // Start the detached idle container + const containerName = `skill-opt-trial-${params.tempDir.split('/').pop() ?? 'run'}`; + const runCmd = [ + 'docker run -d', + `--name ${shellQuote(containerName)}`, + ...dockerSandboxFlags(), + `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, + `-v ${shellQuote(`${params.caseDir}:/case:ro`)}`, + `-v ${shellQuote(`${params.resultsDir}:/results:rw`)}`, + `-v ${shellQuote(`${agentHomeHost}:/home/agent:rw`)}`, + skillMountFlag, + ...envFlags, + '--workdir /work', + '--entrypoint sleep', + shellQuote(params.image), + 'infinity', + ].filter(Boolean).join(' '); + const runResult = await runShellCommand(runCmd, { cwd: params.repoRoot }); + if (runResult.exitCode !== 0) { + throw new Error(`Failed to start trial container: ${runResult.stderr.trim() || runResult.stdout.trim()}`); + } + + const tracePath = join(params.resultsDir, 'trace.jsonl'); + const resultPath = join(params.resultsDir, 'result.json'); + const startedAt = new Date().toISOString(); + const recorder = createTraceRecorder({ + tracePath, + header: { + caseName: params.resolvedCase.name, + agent: agent.name, + model: params.model, + startedAt, + }, + }); + + try { + // Case setup, if any + if (params.resolvedCase.setup.length > 0) { + const setupScript = params.resolvedCase.setup.join(' && '); + const setupCmd = `docker exec ${shellQuote(containerName)} sh -c ${shellQuote(setupScript)}`; + const setupRun = await runShellCommand(setupCmd, { cwd: params.repoRoot }); + if (setupRun.exitCode !== 0) { + throw new Error(`Case setup failed: ${setupRun.stderr.trim() || setupRun.stdout.trim()}`); + } + } + + // Spawn the agent CLI in the container, talking ACP over stdio + const child = spawn('docker', [ + 'exec', '-i', + containerName, + 'sh', '-c', agent.launchCmd, + ], { stdio: ['pipe', 'pipe', 'pipe'] }); + + const stream = createDockerExecStream(child as Parameters[0]); + const client = createWorkbenchClient({ + stream, + onSessionUpdate: (notification) => { + recorder.recordRaw({ + jsonrpc: '2.0', + method: 'session/update', + params: notification, + }); + }, + }); + + // ACP handshake + await client.connection.initialize({ + protocolVersion: '0.22' as unknown as Parameters[0]['protocolVersion'], + clientCapabilities: {}, + }); + const session = await client.connection.newSession({ + cwd: '/work', + mcpServers: [], + }); + const promptPromise = client.connection.prompt({ + sessionId: session.sessionId, + prompt: [{ type: 'text', text: params.resolvedCase.task }], + }); + + let promptOutcome: { ok: true; result: unknown } | { ok: false; error: Error }; + let timeoutHandle: NodeJS.Timeout | undefined; + try { + const timeoutPromise = new Promise((_resolve, reject) => { + timeoutHandle = setTimeout( + () => reject(new Error(`Prompt timeout after ${params.timeoutSeconds}s`)), + params.timeoutSeconds * 1000, + ); + }); + const result = await Promise.race([promptPromise, timeoutPromise]); + promptOutcome = { ok: true, result }; + } catch (error) { + promptOutcome = { ok: false, error: error instanceof Error ? error : new Error(String(error)) }; + } finally { + if (timeoutHandle) clearTimeout(timeoutHandle); + } + + if (promptOutcome.ok) { + recorder.recordRaw({ jsonrpc: '2.0', id: 'final', result: promptOutcome.result }); + } else { + recorder.recordRaw({ jsonrpc: '2.0', id: 'final', error: { message: promptOutcome.error.message } }); + } + + await client.close(); + recorder.finalize(new Date().toISOString()); + + // Run graders inside the same container + const gradeCmd = `docker exec ${shellQuote(containerName)} node /app/dist/workbench/container-runner.js --grade --case /case/case.yml --work /work --results /results`; + const gradeRun = await runShellCommand(gradeCmd, { cwd: params.repoRoot }); + if (gradeRun.exitCode !== 0 && !existsSync(resultPath)) { + throw new Error(`Grade run failed: ${gradeRun.stderr.trim() || gradeRun.stdout.trim()}`); + } + + // Copy agent-internal logs out (non-fatal) + const internalDir = join(params.resultsDir, 'agent-internal'); + mkdirSync(internalDir, { recursive: true }); + for (const dotDir of ['.claude', '.codex', '.gemini', '.opencode', '.pi']) { + await runShellCommand( + `docker cp ${shellQuote(`${containerName}:/home/agent/${dotDir}`)} ${shellQuote(internalDir)} 2>/dev/null || true`, + { cwd: params.repoRoot }, + ); + } + + const pass = readTrialPass(resultPath) ?? false; + return { pass, tracePath, resultPath }; + } finally { + await runShellCommand(`docker rm -f ${shellQuote(containerName)}`, { cwd: params.repoRoot }); + } +} From 4c6d8c7f645a2d01193b750086a6a92b497ca1e1 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:10:33 -0500 Subject: [PATCH 098/121] fix(acp): use SDK PROTOCOL_VERSION constant (integer) for initialize --- src/workbench/docker-runner.ts | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index ffedf23..1d71e2d 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -4,6 +4,7 @@ import { tmpdir } from 'node:os'; import { dirname, join, resolve } from 'node:path'; import { fileURLToPath } from 'node:url'; +import { PROTOCOL_VERSION } from '@agentclientprotocol/sdk'; import { stringify as stringifyYaml } from 'yaml'; import { resolveAuth, resolveContainerPath } from './acp/auth.js'; @@ -774,7 +775,7 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise[0]['protocolVersion'], + protocolVersion: PROTOCOL_VERSION, clientCapabilities: {}, }); const session = await client.connection.newSession({ From 64f99a78655caf88fac6d210993c57c4b8dd8666 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:15:36 -0500 Subject: [PATCH 099/121] feat(docker-runner): host-side ACP orchestration per trial --- src/workbench/docker-runner.ts | 366 ++----------------------- tests/smoke-workbench-docker-runner.ts | 186 ------------- 2 files changed, 27 insertions(+), 525 deletions(-) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index 1d71e2d..0f637ba 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -15,15 +15,13 @@ import { createTraceRecorder } from './acp/trace-recorder.js'; import { createDockerExecStream } from './acp/transport.js'; import { resolveAgent } from './agents/registry.js'; import { loadWorkbenchCase } from './case-loader.js'; -import { MCPORTER_CONFIG_CONTAINER_PATH, writeWorkbenchMcpConfig } from './mcp/index.js'; +import { writeWorkbenchMcpConfig } from './mcp/index.js'; import { runShellCommand } from './process.js'; import type { ResolvedWorkbenchCase, WorkbenchCaseConfig } from './types.js'; import { timestampSlug } from './utils.js'; import { prepareWorkbenchDirectory } from './workspace.js'; const DEFAULT_WORKBENCH_IMAGE = 'skill-optimizer-agent:local'; -const AGENT_RESULTS_DIR = '/tmp/workbench-results'; -const AGENT_PATH = '/work/bin:/app/node_modules/.bin:/work/.venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin'; export function packageRootFromModuleUrl(moduleUrl: string): string { return dirname(dirname(dirname(fileURLToPath(moduleUrl)))); @@ -84,176 +82,6 @@ function dockerCacheEnvFlags(): string[] { ]; } -export function buildDockerAgentCommand(params: { - image: string; - containerName: string; - workDir: string; - caseName: string; - model: string; - task: string; - appendSystemPrompt?: string; - mcpConfigPath?: string; - networkName?: string; - timeoutSeconds: number; - envNames: string[]; -}): string { - const envArgs = params.envNames.map((name) => `-e ${name}`).join(' '); - const mcpEnvArg = params.mcpConfigPath ? `-e MCPORTER_CONFIG=${params.mcpConfigPath}` : ''; - const networkArg = params.networkName ? `--network ${shellQuote(params.networkName)}` : ''; - const taskBase64 = Buffer.from(params.task, 'utf-8').toString('base64'); - const appendSystemPromptBase64 = params.appendSystemPrompt - ? Buffer.from(params.appendSystemPrompt, 'utf-8').toString('base64') - : undefined; - return [ - 'docker run', - `--name ${shellQuote(params.containerName)}`, - ...dockerSandboxFlags(), - '--workdir /work', - `-e PATH=${AGENT_PATH}`, - ...dockerCacheEnvFlags(), - networkArg, - `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, - envArgs, - mcpEnvArg, - shellQuote(params.image), - '--agent', - '--work /work', - `--results ${AGENT_RESULTS_DIR}`, - `--case-name ${shellQuote(params.caseName)}`, - `--model ${shellQuote(params.model)}`, - `--timeout-seconds ${params.timeoutSeconds}`, - `--task-base64 ${shellQuote(taskBase64)}`, - params.mcpConfigPath ? `--mcp-config ${shellQuote(params.mcpConfigPath)}` : '', - appendSystemPromptBase64 - ? `--append-system-prompt-base64 ${shellQuote(appendSystemPromptBase64)}` - : '', - ].filter(Boolean).join(' '); -} - -export function buildDockerMcpServiceCommand(params: { - image: string; - containerName: string; - networkName: string; - alias: string; - mcpDir: string; - command: string; - args: string[]; -}): string { - const serviceCommand = [params.command, ...params.args].map(shellQuote).join(' '); - return [ - 'docker run -d', - `--name ${shellQuote(params.containerName)}`, - ...dockerSandboxFlags(), - `--network ${shellQuote(params.networkName)}`, - `--network-alias ${shellQuote(params.alias)}`, - '--workdir /mcp', - `-v ${shellQuote(`${params.mcpDir}:/mcp:ro`)}`, - '--entrypoint /bin/sh', - shellQuote(params.image), - '-lc', - shellQuote(serviceCommand), - ].filter(Boolean).join(' '); -} - -export function buildDockerMcpServiceProbeCommand(params: { - image: string; - networkName: string; - workDir: string; - serverName: string; -}): string { - const probeCommand = [ - '/app/node_modules/.bin/mcporter', - '--config /work/mcporter.json', - '--root /work', - 'list', - shellQuote(params.serverName), - '--schema', - ].join(' '); - return [ - 'docker run --rm', - ...dockerSandboxFlags(), - `--network ${shellQuote(params.networkName)}`, - '--workdir /work', - `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, - '--entrypoint /bin/sh', - shellQuote(params.image), - '-lc', - shellQuote(probeCommand), - ].filter(Boolean).join(' '); -} - -export function buildDockerSetupCommand(params: { - image: string; - caseDir: string; - workDir: string; - envNames: string[]; -}): string { - const envArgs = params.envNames.map((name) => `-e ${name}`).join(' '); - return [ - 'docker run --rm', - ...dockerSandboxFlags(), - '--workdir /work', - ...dockerCacheEnvFlags(), - `-v ${shellQuote(`${params.caseDir}:/case:ro`)}`, - `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, - envArgs, - shellQuote(params.image), - '--setup', - '--case /case/case.yml', - '--work /work', - ].filter(Boolean).join(' '); -} - -function agentContainerName(tempDir: string): string { - return `skill-optimizer-agent-${tempDir.split('/').pop() ?? 'run'}`; -} - -async function copyAgentResults(containerName: string, resultsDir: string, repoRoot: string): Promise { - const copy = await runShellCommand( - `docker cp ${shellQuote(`${containerName}:${AGENT_RESULTS_DIR}/.`)} ${shellQuote(resultsDir)}`, - { cwd: repoRoot }, - ); - - if (copy.exitCode !== 0) { - throw new Error([ - 'Failed to copy agent results from Docker container', - copy.stdout.trim(), - copy.stderr.trim(), - ].filter(Boolean).join('\n\n')); - } - - await runShellCommand(`chmod -R a+rw ${shellQuote(resultsDir)}`, { cwd: repoRoot }); -} - -async function removeContainer(containerName: string, repoRoot: string): Promise { - await runShellCommand(`docker rm -f ${shellQuote(containerName)}`, { cwd: repoRoot }); -} - -export function buildDockerGradeCommand(params: { - image: string; - caseDir: string; - workDir: string; - resultsDir: string; - envNames: string[]; -}): string { - const envArgs = params.envNames.map((name) => `-e ${name}`).join(' '); - return [ - 'docker run --rm', - ...dockerSandboxFlags(), - '--workdir /work', - ...dockerCacheEnvFlags(), - `-v ${shellQuote(`${params.caseDir}:/case:ro`)}`, - `-v ${shellQuote(`${params.workDir}:/work:rw`)}`, - `-v ${shellQuote(`${params.resultsDir}:/results:rw`)}`, - envArgs, - shellQuote(params.image), - '--grade', - '--case /case/case.yml', - '--work /work', - '--results /results', - ].filter(Boolean).join(' '); -} - function buildBundledCaseFile(params: { source: ReturnType; modelOverride?: string; @@ -311,82 +139,6 @@ function copyCaseSupportDirs(sourceCaseDir: string, bundledCaseDir: string): voi } } -function mcpNetworkName(tempDir: string): string { - return `skill-optimizer-mcp-${tempDir.split('/').pop() ?? 'run'}`; -} - -async function createDockerNetwork(networkName: string, repoRoot: string): Promise { - const create = await runShellCommand(`docker network create ${shellQuote(networkName)}`, { cwd: repoRoot }); - if (create.exitCode !== 0) { - throw new Error(['Failed to create MCP Docker network', create.stdout.trim(), create.stderr.trim()].filter(Boolean).join('\n\n')); - } -} - -async function removeDockerNetwork(networkName: string | undefined, repoRoot: string): Promise { - if (!networkName) return; - await runShellCommand(`docker network rm ${shellQuote(networkName)}`, { cwd: repoRoot }); -} - -export async function startMcpServices(params: { - image: string; - networkName: string; - caseDir: string; - tempDir: string; - services: ResolvedWorkbenchCase['mcpServices']; - repoRoot: string; - startedContainers?: string[]; - runCommand?: typeof runShellCommand; -}): Promise { - const containerNames = params.startedContainers ?? []; - const runCommand = params.runCommand ?? runShellCommand; - for (const [name, service] of Object.entries(params.services)) { - const containerName = `${mcpNetworkName(params.tempDir)}-${name}`; - console.log(`Starting MCP service ${name}...`); - const command = buildDockerMcpServiceCommand({ - image: params.image, - containerName, - networkName: params.networkName, - alias: name, - mcpDir: join(params.caseDir, 'mcp'), - command: service.command, - args: service.args, - }); - const run = await runCommand(command, { cwd: params.repoRoot }); - if (run.exitCode !== 0) { - throw new Error([`Failed to start MCP service ${name}`, run.stdout.trim(), run.stderr.trim()].filter(Boolean).join('\n\n')); - } - containerNames.push(containerName); - } - return containerNames; -} - -async function waitForMcpServices(params: { - image: string; - networkName: string; - workDir: string; - services: ResolvedWorkbenchCase['mcpServices']; - repoRoot: string; -}): Promise { - for (const name of Object.keys(params.services)) { - console.log(`Waiting for MCP service ${name}...`); - const command = buildDockerMcpServiceProbeCommand({ - image: params.image, - networkName: params.networkName, - workDir: params.workDir, - serverName: name, - }); - const probe = await runShellCommand(command, { cwd: params.repoRoot, timeoutSeconds: 30 }); - if (probe.exitCode !== 0) { - throw new Error([ - `MCP service ${name} did not become ready`, - probe.stdout.trim(), - probe.stderr.trim(), - ].filter(Boolean).join('\n\n')); - } - console.log(`MCP service ${name} ready.`); - } -} - function copyAgentSupportDirs(sourceCaseDir: string, workDir: string): void { copyCaseSupportDir(sourceCaseDir, workDir, 'bin'); } @@ -535,108 +287,44 @@ export async function runDockerWorkbenchCase( const image = options.image ?? DEFAULT_WORKBENCH_IMAGE; const resolvedCase = resolveDockerWorkbenchCase(options); const prepared = prepareDockerWorkbenchRun({ ...options, case: resolvedCase }); - const containerName = agentContainerName(prepared.tempDir); - const networkName = Object.keys(resolvedCase.mcpServices).length > 0 ? mcpNetworkName(prepared.tempDir) : undefined; - let mcpServiceContainers: string[] = []; try { await ensureDockerImage(image, repoRoot); - const envNames = resolvedCase.env - .filter((name) => process.env[name] !== undefined) - .map((name) => name); - - if (resolvedCase.setup.length > 0) { - const setupCommand = buildDockerSetupCommand({ - image, - caseDir: prepared.caseDir, - workDir: prepared.workDir, - envNames, - }); - const setupRun = await runShellCommand(setupCommand, { cwd: repoRoot }); - if (setupRun.exitCode !== 0) { - writeFatalResult({ - resultPath: prepared.resultPath, - caseName: resolvedCase.name, - model: options.model ?? resolvedCase.model, - evidence: [ - 'setup failed', - setupRun.stdout.trim(), - setupRun.stderr.trim(), - ].filter(Boolean), - }); - return copyWorkspaceIfRequested(prepared, true); - } - } + const skillMountParams = resolvedCase.skillUnderTest + ? { + hostSkillDir: resolvedCase.skillUnderTest.hostPath, + skillSlug: resolvedCase.skillUnderTest.slug, + } + : null; - if (networkName) { - await createDockerNetwork(networkName, repoRoot); - mcpServiceContainers = await startMcpServices({ - image, - networkName, - caseDir: prepared.caseDir, - tempDir: prepared.tempDir, - services: resolvedCase.mcpServices, - repoRoot, - startedContainers: mcpServiceContainers, - }); - await waitForMcpServices({ - image, - networkName, - workDir: prepared.workDir, - services: resolvedCase.mcpServices, - repoRoot, - }); - } - - const agentCommand = buildDockerAgentCommand({ - image, - containerName, + await runOneAcpTrial({ + resolvedCase, + tempDir: prepared.tempDir, workDir: prepared.workDir, - caseName: resolvedCase.name, + caseDir: prepared.caseDir, + resultsDir: prepared.resultsDir, + agentName: resolvedCase.agent, model: options.model ?? resolvedCase.model, - task: resolvedCase.task, - appendSystemPrompt: options.appendSystemPrompt, - mcpConfigPath: prepared.mcpConfigPath ? MCPORTER_CONFIG_CONTAINER_PATH : undefined, - networkName, + image, + skillMountParams, + repoRoot, timeoutSeconds: resolvedCase.timeoutSeconds, - envNames, }); - const agentRun = await runShellCommand(agentCommand, { cwd: repoRoot }); - await copyAgentResults(containerName, prepared.resultsDir, repoRoot); - - if (agentRun.exitCode !== 0) { - if (!existsSync(prepared.resultPath)) { - throw new Error([ - 'Docker agent run failed', - agentRun.stdout.trim(), - agentRun.stderr.trim(), - ].filter(Boolean).join('\n\n')); - } - } else { - const gradeCommand = buildDockerGradeCommand({ - image, - caseDir: prepared.caseDir, - workDir: prepared.workDir, - resultsDir: prepared.resultsDir, - envNames, - }); - const gradeRun = await runShellCommand(gradeCommand, { cwd: repoRoot }); - - if (gradeRun.exitCode !== 0 && !existsSync(prepared.resultPath)) { - throw new Error([ - 'Docker grade run failed', - gradeRun.stdout.trim(), - gradeRun.stderr.trim(), - ].filter(Boolean).join('\n\n')); - } - } return copyWorkspaceIfRequested(prepared, options.keepWorkspace); + } catch (error) { + // If runOneAcpTrial threw before any result.json was written, persist a fatal record. + if (!existsSync(prepared.resultPath)) { + writeFatalResult({ + resultPath: prepared.resultPath, + caseName: resolvedCase.name, + model: options.model ?? resolvedCase.model, + evidence: [error instanceof Error ? error.message : String(error)], + }); + } + return copyWorkspaceIfRequested(prepared, true); } finally { - await removeContainer(containerName, repoRoot); - await Promise.all(mcpServiceContainers.map((name) => removeContainer(name, repoRoot))); - await removeDockerNetwork(networkName, repoRoot); prepared.cleanup(); } } diff --git a/tests/smoke-workbench-docker-runner.ts b/tests/smoke-workbench-docker-runner.ts index 644b876..c32313b 100644 --- a/tests/smoke-workbench-docker-runner.ts +++ b/tests/smoke-workbench-docker-runner.ts @@ -5,14 +5,8 @@ import { join } from 'node:path'; import { test } from 'node:test'; import { - buildDockerAgentCommand, - buildDockerGradeCommand, - buildDockerMcpServiceCommand, - buildDockerMcpServiceProbeCommand, - buildDockerSetupCommand, packageRootFromModuleUrl, prepareDockerWorkbenchRun, - startMcpServices, } from '../src/workbench/docker-runner.js'; test('packageRootFromModuleUrl resolves repo root independently of cwd', () => { @@ -227,186 +221,6 @@ test('prepareDockerWorkbenchRun bundles hidden MCP service support outside work' } }); -test('startMcpServices records services started before a later service fails', async () => { - const startedContainers: string[] = []; - - await assert.rejects( - startMcpServices({ - image: 'skill-optimizer-workbench:local', - networkName: 'skill-optimizer-mcp-test', - caseDir: '/tmp/case', - tempDir: '/tmp/skill-optimizer-workbench-test', - services: { - ok: { command: 'node', args: ['server.mjs'] }, - fail: { command: 'node', args: ['server.mjs'] }, - }, - repoRoot: '/tmp/repo', - startedContainers, - runCommand: async (command) => ({ - exitCode: command.includes('-fail') ? 7 : 0, - stdout: '', - stderr: command.includes('-fail') ? 'boom' : '', - }), - }), - /Failed to start MCP service fail/, - ); - - assert.deepEqual(startedContainers, ['skill-optimizer-mcp-skill-optimizer-workbench-test-ok']); -}); - -test('setup docker command mounts case and work before agent phase', () => { - const command = buildDockerSetupCommand({ - image: 'skill-optimizer-workbench:local', - caseDir: '/tmp/case', - workDir: '/tmp/work', - envNames: [], - }); - - assert.match(command, /--setup/); - assert.match(command, /-v '\/tmp\/case:\/case:ro'/); - assert.match(command, /-v '\/tmp\/work:\/work:rw'/); - assert.doesNotMatch(command, /\/results/); - assert.doesNotMatch(command, /docker\.sock/); -}); - -test('agent docker command mounts only work and uses sandbox hardening flags', () => { - const command = buildDockerAgentCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-agent-test', - workDir: '/tmp/work', - caseName: 'extract-pdf-facts', - model: 'openrouter/google/gemini-2.5-flash', - task: 'Read the PDF and write answer.json.', - timeoutSeconds: 600, - envNames: ['OPENROUTER_API_KEY'], - }); - - assert.match(command, /--agent/); - assert.match(command, /--name 'skill-optimizer-agent-test'/); - assert.match(command, /-v '\/tmp\/work:\/work:rw'/); - assert.match(command, /--workdir \/work/); - assert.match(command, /-e PATH=\/work\/bin:\/app\/node_modules\/\.bin:\/work\/\.venv\/bin:\/usr\/local\/sbin:\/usr\/local\/bin:\/usr\/sbin:\/usr\/bin:\/sbin:\/bin/); - assert.match(command, /--cap-drop=ALL/); - assert.match(command, /--security-opt no-new-privileges/); - assert.match(command, /-e OPENROUTER_API_KEY/); - assert.doesNotMatch(command, /\/case/); - assert.doesNotMatch(command, /\/results/); - assert.doesNotMatch(command, /docker\.sock/); -}); - -test('agent docker command passes optional appended system prompt', () => { - const command = buildDockerAgentCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-agent-test', - workDir: '/tmp/work', - caseName: 'prompted-case', - model: 'openrouter/google/gemini-2.5-flash', - task: 'Write output.txt.', - timeoutSeconds: 600, - envNames: [], - appendSystemPrompt: 'Prefer simple shell commands when possible.', - }); - - assert.match(command, /--append-system-prompt-base64/); - assert.match(command, new RegExp(Buffer.from('Prefer simple shell commands when possible.', 'utf-8').toString('base64'))); -}); - -test('agent docker command passes optional MCP config path', () => { - const command = buildDockerAgentCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-agent-test', - workDir: '/tmp/work', - caseName: 'mcp-case', - model: 'openrouter/google/gemini-2.5-flash', - task: 'Use MCP.', - timeoutSeconds: 600, - envNames: [], - mcpConfigPath: '/work/mcporter.json', - }); - - assert.match(command, /-e MCPORTER_CONFIG=\/work\/mcporter\.json/); - assert.match(command, /--mcp-config '\/work\/mcporter\.json'/); -}); - -test('agent docker command joins optional MCP network', () => { - const command = buildDockerAgentCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-agent-test', - workDir: '/tmp/work', - caseName: 'mcp-case', - model: 'openrouter/google/gemini-2.5-flash', - task: 'Use MCP.', - timeoutSeconds: 600, - envNames: [], - networkName: 'skill-optimizer-mcp-test', - }); - - assert.match(command, /--network 'skill-optimizer-mcp-test'/); -}); - -test('MCP service docker command mounts hidden service files outside agent work', () => { - const command = buildDockerMcpServiceCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-mcp-test-calculator', - networkName: 'skill-optimizer-mcp-test', - alias: 'calculator', - mcpDir: '/tmp/case/mcp', - command: 'node', - args: ['server.mjs'], - }); - - assert.match(command, /-v '\/tmp\/case\/mcp:\/mcp:ro'/); - assert.match(command, /--workdir \/mcp/); - assert.match(command, /--network-alias 'calculator'/); - assert.doesNotMatch(command, /\/work/); -}); - -test('MCP service docker command does not forward case env vars', () => { - const command = buildDockerMcpServiceCommand({ - image: 'skill-optimizer-workbench:local', - containerName: 'skill-optimizer-mcp-test-calculator', - networkName: 'skill-optimizer-mcp-test', - alias: 'calculator', - mcpDir: '/tmp/case/mcp', - command: 'node', - args: ['server.mjs'], - }); - - assert.doesNotMatch(command, /-e OPENROUTER_API_KEY/); -}); - -test('MCP service probe command verifies service through mcporter on private network', () => { - const command = buildDockerMcpServiceProbeCommand({ - image: 'skill-optimizer-workbench:local', - networkName: 'skill-optimizer-mcp-test', - workDir: '/tmp/work', - serverName: 'calculator', - }); - - assert.match(command, /--network 'skill-optimizer-mcp-test'/); - assert.match(command, /-v '\/tmp\/work:\/work:rw'/); - assert.match(command, /mcporter --config \/work\/mcporter\.json --root \/work list/); - assert.match(command, /calculator/); - assert.match(command, /--schema/); -}); - -test('grade docker command mounts case after agent phase', () => { - const command = buildDockerGradeCommand({ - image: 'skill-optimizer-workbench:local', - caseDir: '/tmp/case', - workDir: '/tmp/work', - resultsDir: '/tmp/results', - envNames: [], - }); - - assert.match(command, /--grade/); - assert.match(command, /-v '\/tmp\/case:\/case:ro'/); - assert.match(command, /-v '\/tmp\/work:\/work:rw'/); - assert.match(command, /-v '\/tmp\/results:\/results:rw'/); - assert.match(command, /--cap-drop=ALL/); - assert.match(command, /--security-opt no-new-privileges/); -}); - test('prepareDockerWorkbenchRun honors --out as the results root', () => { const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-out-')); try { From 7b59fb96d60f346ec5574435172c1b3e77222e14 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:18:51 -0500 Subject: [PATCH 100/121] fix(docker-runner): restore cache env flags on per-trial container --- src/workbench/docker-runner.ts | 1 + 1 file changed, 1 insertion(+) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index 0f637ba..eabef6f 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -402,6 +402,7 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise Date: Mon, 25 May 2026 15:21:41 -0500 Subject: [PATCH 101/121] refactor(container-runner): drop --agent mode and pi-agent.ts (ACP host-side now) --- package.json | 2 +- src/workbench/container-runner.ts | 197 ++---------------------------- src/workbench/index.ts | 1 - src/workbench/pi-agent.ts | 156 ----------------------- src/workbench/sandbox.ts | 13 -- tests/smoke-workbench-pi-agent.ts | 159 ------------------------ 6 files changed, 13 insertions(+), 515 deletions(-) delete mode 100644 src/workbench/pi-agent.ts delete mode 100644 src/workbench/sandbox.ts delete mode 100644 tests/smoke-workbench-pi-agent.ts diff --git a/package.json b/package.json index f4fdb35..f31dfda 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-pi-agent.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", diff --git a/src/workbench/container-runner.ts b/src/workbench/container-runner.ts index 2299a10..cef8032 100644 --- a/src/workbench/container-runner.ts +++ b/src/workbench/container-runner.ts @@ -1,43 +1,15 @@ -import { existsSync, mkdirSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { writeFileSync } from 'node:fs'; import { dirname, join } from 'node:path'; import { runGraderCommands } from './check-runner.js'; import { loadWorkbenchCase } from './case-loader.js'; import { buildWorkbenchMetricsFromTrace } from './metrics.js'; -import { createWorkbenchPiSession } from './pi-agent.js'; import { runShellCommand } from './process.js'; -import { buildAgentSystemPrompt } from './sandbox.js'; -// TODO(Task 17): --agent mode is removed in Task 17, along with the legacy -// trace functions below. These imports stay only to satisfy the type system -// until then. import type { WorkbenchGrade, WorkbenchResult } from './types.js'; -import { isRecord, writeJsonFile } from './utils.js'; +import { writeJsonFile } from './utils.js'; import { buildWorkbenchEnv, prepareWorkbenchDirectory } from './workspace.js'; export { prepareWorkbenchDirectory } from './workspace.js'; -export { buildAgentSystemPrompt } from './sandbox.js'; - -interface PromptSession { - prompt(prompt: string): Promise; - systemPrompt?: string; - subscribe?: (listener: (event: unknown) => void) => () => void; - dispose?: () => void; - state?: { - messages?: unknown[]; - }; -} - -interface AgentRunnerArgs { - mode: 'agent'; - caseName: string; - model: string; - task: string; - appendSystemPrompt?: string; - timeoutSeconds: number; - workDir: string; - resultsDir: string; - mcpConfigPath?: string; -} interface GradeRunnerArgs { mode: 'grade'; @@ -52,7 +24,7 @@ interface SetupRunnerArgs { workDir: string; } -export type ContainerRunnerArgs = AgentRunnerArgs | GradeRunnerArgs | SetupRunnerArgs; +export type ContainerRunnerArgs = GradeRunnerArgs | SetupRunnerArgs; export function buildContainerWorkbenchEnv(params: { casePath: string; @@ -70,34 +42,8 @@ export function buildContainerWorkbenchEnv(params: { export function parseContainerRunnerArgs(args: string[]): ContainerRunnerArgs { const workDir = getFlagValue(args, '--work'); - const resultsDir = getFlagValue(args, '--results'); - - if (args.includes('--agent')) { - const caseName = getFlagValue(args, '--case-name'); - const model = getFlagValue(args, '--model'); - const taskBase64 = getFlagValue(args, '--task-base64'); - const appendSystemPromptBase64 = getFlagValue(args, '--append-system-prompt-base64'); - const mcpConfigPath = getFlagValue(args, '--mcp-config'); - const timeoutSeconds = Number(getFlagValue(args, '--timeout-seconds')); - if (!caseName || !model || !taskBase64 || !Number.isFinite(timeoutSeconds) || timeoutSeconds <= 0 || !workDir || !resultsDir) { - throw new Error('Usage: container-runner --agent --case-name --model --task-base64 --timeout-seconds --work --results '); - } - return { - mode: 'agent', - caseName, - model, - task: Buffer.from(taskBase64, 'base64').toString('utf-8'), - appendSystemPrompt: appendSystemPromptBase64 - ? Buffer.from(appendSystemPromptBase64, 'base64').toString('utf-8') - : undefined, - timeoutSeconds, - workDir, - resultsDir, - mcpConfigPath, - }; - } - const casePath = getFlagValue(args, '--case'); + if (args.includes('--setup')) { if (!casePath || !workDir) { throw new Error('Usage: container-runner --setup --case --work '); @@ -105,11 +51,15 @@ export function parseContainerRunnerArgs(args: string[]): ContainerRunnerArgs { return { mode: 'setup', casePath, workDir }; } - if (!args.includes('--grade') || !casePath || !workDir || !resultsDir) { - throw new Error('Usage: container-runner --agent ... or --setup --case --work or --grade --case --work --results '); + if (args.includes('--grade')) { + const resultsDir = getFlagValue(args, '--results'); + if (!casePath || !workDir || !resultsDir) { + throw new Error('Usage: container-runner --grade --case --work --results '); + } + return { mode: 'grade', casePath, workDir, resultsDir }; } - return { mode: 'grade', casePath, workDir, resultsDir }; + throw new Error('container-runner: expected --setup or --grade'); } function getFlagValue(args: string[], flag: string): string | undefined { @@ -126,61 +76,6 @@ function getFlagValue(args: string[], flag: string): string | undefined { return value; } -export async function runAgentPromptWithTimeout( - session: PromptSession, - prompt: string, - timeoutSeconds: number, -): Promise { - let timeout: NodeJS.Timeout | undefined; - try { - await Promise.race([ - session.prompt(prompt), - new Promise((_, reject) => { - timeout = setTimeout(() => { - reject(new Error(`Agent timed out after ${timeoutSeconds} seconds`)); - }, timeoutSeconds * 1000); - }), - ]); - } finally { - if (timeout) { - clearTimeout(timeout); - } - } - - const messages = session.state?.messages ?? []; - const lastMessage = messages[messages.length - 1]; - if (!isRecord(lastMessage) || lastMessage.role !== 'assistant') { - return; - } - - if (lastMessage.stopReason === 'error' || lastMessage.stopReason === 'aborted') { - const errorMessage = typeof lastMessage.errorMessage === 'string' - ? lastMessage.errorMessage - : `Agent request ${lastMessage.stopReason}`; - throw new Error(errorMessage); - } -} - -// TODO(Task 17): writeBestEffortTrace is part of the --agent mode flow which -// is removed in Task 17. Stubbed to keep typecheck green during transition. -export function writeBestEffortTrace(_params: { - tracePath: string; - caseName?: string; - model?: string; - startedAt?: string; - endedAt?: string; - session?: unknown; - recorder?: unknown; -}): boolean { - return false; -} - -// TODO(Task 17): writeTraceFile is removed with --agent mode. Stub kept for -// the brief window between Task 15 and Task 17. -export function writeTraceFile(_tracePath: string, _trace: unknown): void { - // no-op -} - async function runCleanupCommands( commands: string[], opts: { cwd: string; env: NodeJS.ProcessEnv }, @@ -259,67 +154,6 @@ function buildResult(params: { }; } -// TODO(Task 17): readTraceFile is gone — replaced by ACP trace-recorder; -// grade mode reads metrics directly via buildWorkbenchMetricsFromTrace(). - -function summarizeContent(content: unknown): string | undefined { - if (!Array.isArray(content)) { - return undefined; - } - - const text = content - .flatMap((item) => isRecord(item) && typeof item.text === 'string' ? [item.text] : []) - .join('\n') - .replace(/\s+/g, ' ') - .trim(); - return text.length > 160 ? `${text.slice(0, 157)}...` : text || undefined; -} - -function logAgentEvent(event: unknown): void { - if (!isRecord(event) || typeof event.type !== 'string') { - return; - } - - if (event.type === 'message_end' && isRecord(event.message)) { - const role = typeof event.message.role === 'string' ? event.message.role : 'unknown'; - const text = summarizeContent(event.message.content); - console.log(`[agent:${event.type}] ${role}${text ? `: ${text}` : ''}`); - return; - } - - if (event.type === 'tool_execution_start') { - const name = typeof event.toolName === 'string' ? event.toolName : 'unknown'; - const args = event.args === undefined ? '' : ` ${JSON.stringify(event.args)}`; - console.log(`[agent:${event.type}] ${name}${args}`); - return; - } - - if (event.type === 'tool_execution_end') { - const name = typeof event.toolName === 'string' ? event.toolName : 'unknown'; - const status = event.isError === true ? 'error' : 'ok'; - console.log(`[agent:${event.type}] ${name} ${status}`); - return; - } - - if (event.type === 'turn_start' || event.type === 'turn_end' || event.type === 'agent_start' || event.type === 'agent_end') { - console.log(`[agent:${event.type}]`); - } -} - -function logAgentSystemPrompt(systemPrompt: string): void { - console.log('[agent:system_prompt_start]'); - console.log(systemPrompt); - console.log('[agent:system_prompt_end]'); -} - -// TODO(Task 17): --agent mode is removed entirely in Task 17. The body -// previously here drove a Pi session, normalized events into the legacy -// WorkbenchTrace shape, and wrote out trace.jsonl. ACP-based agents replace -// all of it. Stubbed to keep typecheck green during the transition. -async function runAgentMode(_parsed: AgentRunnerArgs): Promise { - throw new Error('container-runner --agent mode is disabled; ACP agents replace it in Task 17'); -} - async function runSetupMode(parsed: SetupRunnerArgs): Promise { const resolved = loadWorkbenchCase(parsed.casePath); const env = buildContainerWorkbenchEnv({ @@ -342,8 +176,6 @@ async function runGradeMode(parsed: GradeRunnerArgs): Promise { const resolved = loadWorkbenchCase(parsed.casePath); const env = buildContainerWorkbenchEnv(parsed); const startedAt = new Date().toISOString(); - // TODO(Task 17): grade mode no longer reads the trace file; metrics come - // straight from buildWorkbenchMetricsFromTrace(tracePath) below. try { const grade = await runGraderCommands(resolved.graders, { @@ -379,12 +211,7 @@ async function runGradeMode(parsed: GradeRunnerArgs): Promise { export async function runContainerWorkbenchCase(args: string[]): Promise { const parsed = parseContainerRunnerArgs(args); - if (parsed.mode === 'agent') { - return runAgentMode(parsed); - } - if (parsed.mode === 'setup') { - return runSetupMode(parsed); - } + if (parsed.mode === 'setup') return runSetupMode(parsed); return runGradeMode(parsed); } diff --git a/src/workbench/index.ts b/src/workbench/index.ts index 6376574..8dc809e 100644 --- a/src/workbench/index.ts +++ b/src/workbench/index.ts @@ -5,7 +5,6 @@ export * from './check-runner.js'; export * from './trace.js'; export * from './models.js'; export * from './mcp/index.js'; -export * from './pi-agent.js'; export * from './docker-runner.js'; export * from './run-case.js'; export * from './suite-loader.js'; diff --git a/src/workbench/pi-agent.ts b/src/workbench/pi-agent.ts deleted file mode 100644 index 60c99bb..0000000 --- a/src/workbench/pi-agent.ts +++ /dev/null @@ -1,156 +0,0 @@ -import { - createAgentSession, - createBashTool, - createEditTool, - createFindTool, - createGrepTool, - createLsTool, - createReadTool, - createWriteTool, - AuthStorage, - DefaultResourceLoader, - ModelRegistry, - SessionManager, - type ResourceLoader, -} from '@mariozechner/pi-coding-agent'; -import type { AgentTool } from '@mariozechner/pi-agent-core'; -import { getModel, type Api, type Model } from '@mariozechner/pi-ai'; -import { resolve } from 'node:path'; - -import { buildAgentSystemPrompt } from './sandbox.js'; - -export function stripSensitiveEnv(env: NodeJS.ProcessEnv): NodeJS.ProcessEnv { - return { ...env }; -} - -export function createWorkbenchPiTools(cwd: string): AgentTool[] { - return [ - createReadTool(cwd), - createBashTool(cwd, { - spawnHook: (context) => ({ - ...context, - env: stripSensitiveEnv(context.env), - }), - }), - createEditTool(cwd), - createWriteTool(cwd), - createGrepTool(cwd), - createFindTool(cwd), - createLsTool(cwd), - ]; -} - -export async function createWorkbenchPiResourceLoader(params: { - cwd: string; - appendSystemPrompt?: string; - mcpConfigPath?: string; -}): Promise { - const cwd = resolve(params.cwd); - const appendSystemPrompt = [buildAgentSystemPrompt(), buildMcpSystemPrompt(params.mcpConfigPath), params.appendSystemPrompt] - .filter((value): value is string => typeof value === 'string' && value.trim().length > 0) - .join('\n\n'); - const loader = new DefaultResourceLoader({ - cwd, - noExtensions: true, - noSkills: true, - additionalSkillPaths: [cwd], - appendSystemPrompt, - }); - - await loader.reload(); - return loader; -} - -export async function createWorkbenchPiSession(params: { - cwd: string; - modelRef: string; - apiKeyEnv?: string; - appendSystemPrompt?: string; - mcpConfigPath?: string; - thinkingLevel?: 'off' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh'; -}) { - const { provider, model } = parseModelRef(params.modelRef); - if (provider !== 'openrouter') { - throw new Error(`Workbench only supports OpenRouter model refs, got: ${params.modelRef}`); - } - - const authStorage = AuthStorage.create(); - const apiKeyEnv = params.apiKeyEnv ?? 'OPENROUTER_API_KEY'; - const apiKey = process.env[apiKeyEnv]; - if (apiKey) { - authStorage.setRuntimeApiKey('openrouter' as never, apiKey); - } - - const modelRegistry = ModelRegistry.create(authStorage); - const resolvedModel = modelRegistry.find(provider, model) - ?? getModel(provider as never, model) - ?? synthesizeOpenRouterModel(provider, model); - if (!resolvedModel) { - throw new Error(`Could not resolve Pi model ${provider}/${model}`); - } - - const auth = await modelRegistry.getApiKeyAndHeaders(resolvedModel); - if (!auth.ok) { - throw new Error(auth.error); - } - - const resourceLoader = await createWorkbenchPiResourceLoader({ - cwd: params.cwd, - appendSystemPrompt: params.appendSystemPrompt, - mcpConfigPath: params.mcpConfigPath, - }); - - return createAgentSession({ - cwd: params.cwd, - model: resolvedModel, - thinkingLevel: params.thinkingLevel ?? 'medium', - authStorage, - modelRegistry, - resourceLoader, - tools: createWorkbenchPiTools(params.cwd), - sessionManager: SessionManager.inMemory(), - }); -} - -function buildMcpSystemPrompt(mcpConfigPath: string | undefined): string | undefined { - if (!mcpConfigPath) { - return undefined; - } - - return [ - 'Additional command:', - '- `mcp` is available on PATH for configured MCP servers.', - '- Run `mcp list --schema` to inspect available tools when needed.', - '- Run `mcp call key=value` to call a tool from bash.', - ].join('\n'); -} - -function parseModelRef(modelRef: string): { provider: string; model: string } { - const slash = modelRef.indexOf('/'); - if (slash <= 0 || slash === modelRef.length - 1) { - throw new Error(`Invalid model ref: ${modelRef}`); - } - return { - provider: modelRef.slice(0, slash), - model: modelRef.slice(slash + 1), - }; -} - -function synthesizeOpenRouterModel(provider: string, modelName: string): Model | undefined { - if (provider !== 'openrouter') { - return undefined; - } - - return { - id: modelName, - name: modelName, - api: 'openai-completions' as const, - provider: 'openrouter' as const, - baseUrl: 'https://openrouter.ai/api/v1', - reasoning: false, - input: ['text'], - cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, - contextWindow: 128000, - maxTokens: 16384, - }; -} diff --git a/src/workbench/sandbox.ts b/src/workbench/sandbox.ts deleted file mode 100644 index 8cb0864..0000000 --- a/src/workbench/sandbox.ts +++ /dev/null @@ -1,13 +0,0 @@ -export function buildAgentSystemPrompt(): string { - return [ - 'Operating environment:', - '- Current working directory is /work.', - '- Write all outputs under /work.', - '- The Docker socket is not mounted.', - '- Internet access is available for task dependencies unless the network is unavailable.', - '- Node.js, npm, Python, pip, and venv are installed.', - '- Do not use global pip installs.', - '- If you need Python packages, run: python -m venv /work/.venv && /work/.venv/bin/pip install .', - '- Run Python scripts with /work/.venv/bin/python when using installed packages.', - ].join('\n'); -} diff --git a/tests/smoke-workbench-pi-agent.ts b/tests/smoke-workbench-pi-agent.ts deleted file mode 100644 index f2c01ee..0000000 --- a/tests/smoke-workbench-pi-agent.ts +++ /dev/null @@ -1,159 +0,0 @@ -import assert from 'node:assert/strict'; -import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs'; -import { tmpdir } from 'node:os'; -import { join } from 'node:path'; -import { test } from 'node:test'; - -import { createWorkbenchPiResourceLoader, createWorkbenchPiSession, createWorkbenchPiTools, stripSensitiveEnv } from '../src/workbench/pi-agent.js'; - -function toolText(result: unknown): string { - const content = (result as { content?: Array<{ text?: string }> }).content ?? []; - return content.map((item) => item.text ?? '').join(''); -} - -test('createWorkbenchPiTools enables coding plus repo-scale search tools', () => { - const tools = createWorkbenchPiTools('/work'); - const names = tools.map((tool) => tool.name).sort(); - - assert.deepEqual(names, ['bash', 'edit', 'find', 'grep', 'ls', 'read', 'write']); -}); - -test('stripSensitiveEnv preserves all case-allowed credentials for tool subprocesses', () => { - const env = stripSensitiveEnv({ - OPENROUTER_API_KEY: 'secret', - OPENAI_API_KEY: 'secret', - GOOGLE_WORKSPACE_CLI_TOKEN: 'gws-token', - GOOGLE_WORKSPACE_CLI_CLIENT_SECRET: 'gws-secret', - GOOGLE_WORKSPACE_CLI_CREDENTIALS_FILE: '/work/gws-credentials.json', - MODEL_AUTH_FILE: '/run/secrets/model-auth.json', - WHATSAPP_ACCESS_TOKEN: 'secret', - DASHBOARD_TOKEN_SECRET: 'secret', - PATH: '/usr/bin', - WORK: '/work', - }); - - assert.equal(env.OPENROUTER_API_KEY, 'secret'); - assert.equal(env.OPENAI_API_KEY, 'secret'); - assert.equal(env.GOOGLE_WORKSPACE_CLI_TOKEN, 'gws-token'); - assert.equal(env.GOOGLE_WORKSPACE_CLI_CLIENT_SECRET, 'gws-secret'); - assert.equal(env.GOOGLE_WORKSPACE_CLI_CREDENTIALS_FILE, '/work/gws-credentials.json'); - assert.equal(env.MODEL_AUTH_FILE, '/run/secrets/model-auth.json'); - assert.equal(env.WHATSAPP_ACCESS_TOKEN, 'secret'); - assert.equal(env.DASHBOARD_TOKEN_SECRET, 'secret'); - assert.equal(env.PATH, '/usr/bin'); - assert.equal(env.WORK, '/work'); -}); - -test('createWorkbenchPiResourceLoader discovers a root SKILL.md from references', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-skill-')); - try { - writeFileSync(root + '/SKILL.md', [ - '---', - 'name: pdf', - 'description: PDF merge instructions', - '---', - '', - '# PDF Skill', - ].join('\n'), 'utf-8'); - mkdirSync(join(root, 'inputs')); - - const loader = await createWorkbenchPiResourceLoader({ cwd: root }); - const loaded = loader.getSkills().skills.map((skill) => skill.name); - - assert.ok(loaded.includes('pdf')); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('createWorkbenchPiResourceLoader appends suite prompt after workbench prompt', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-append-prompt-')); - try { - const loader = await createWorkbenchPiResourceLoader({ - cwd: root, - appendSystemPrompt: 'Prefer simple shell commands when possible.', - }); - - const appended = loader.getAppendSystemPrompt().join('\n\n'); - assert.match(appended, /Operating environment:/); - assert.match(appended, /Prefer simple shell commands when possible\./); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('createWorkbenchPiResourceLoader documents MCP command when configured', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-mcp-prompt-')); - try { - const loader = await createWorkbenchPiResourceLoader({ - cwd: root, - mcpConfigPath: '/work/mcporter.json', - }); - - const appended = loader.getAppendSystemPrompt().join('\n\n'); - assert.match(appended, /`mcp` is available on PATH/); - assert.match(appended, /Run `mcp list --schema`/); - assert.doesNotMatch(appended, /calculator\.add/); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); - -test('createWorkbenchPiTools passes process env through bash subprocesses', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-tool-env-')); - const previousSecret = process.env.WORKBENCH_AGENT_SECRET; - try { - process.env.WORKBENCH_AGENT_SECRET = 'agent-secret'; - const bashTool = createWorkbenchPiTools(root).find((tool) => tool.name === 'bash'); - assert.ok(bashTool); - - const result = await bashTool.execute( - 'call-1', - { command: 'printf "%s" "$WORKBENCH_AGENT_SECRET"', timeout: 5 }, - new AbortController().signal, - ); - - assert.equal(toolText(result), 'agent-secret'); - } finally { - if (previousSecret === undefined) { - delete process.env.WORKBENCH_AGENT_SECRET; - } else { - process.env.WORKBENCH_AGENT_SECRET = previousSecret; - } - rmSync(root, { recursive: true, force: true }); - } -}); - -test('createWorkbenchPiSession leaves runtime API key env available after session creation', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-session-env-')); - const previousApiKey = process.env.OPENROUTER_API_KEY; - try { - process.env.OPENROUTER_API_KEY = 'test-openrouter-key'; - const created = await createWorkbenchPiSession({ - cwd: root, - modelRef: 'openrouter/google/gemini-2.5-flash', - }); - - assert.equal(process.env.OPENROUTER_API_KEY, 'test-openrouter-key'); - (created.session as { dispose?: () => void }).dispose?.(); - } finally { - if (previousApiKey === undefined) { - delete process.env.OPENROUTER_API_KEY; - } else { - process.env.OPENROUTER_API_KEY = previousApiKey; - } - rmSync(root, { recursive: true, force: true }); - } -}); - -test('createWorkbenchPiSession rejects non-OpenRouter model refs', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-openrouter-only-')); - try { - await assert.rejects( - () => createWorkbenchPiSession({ cwd: root, modelRef: 'direct/model' }), - /only supports OpenRouter/, - ); - } finally { - rmSync(root, { recursive: true, force: true }); - } -}); From 439bb6633b4c62ff1cbeedb9dfae1ebdf02193dc Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:25:10 -0500 Subject: [PATCH 102/121] feat(trials): aggregate tokens and durationMs across trials Extend TrialAggregate / WorkbenchModelAggregateResult with totalTokens and totalDurationMs, sourced from per-trial result.metrics.tokens.total and result.metrics.durationMs. Missing values fall back to 0 so legacy result.json files remain backward-compatible. Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/run-case.ts | 4 ++++ src/workbench/run-suite.ts | 4 ++++ src/workbench/trials.ts | 8 ++++++++ src/workbench/types.ts | 4 ++++ tests/trials-aggregate.test.ts | 30 ++++++++++++++++++++++++++++++ 5 files changed, 50 insertions(+) create mode 100644 tests/trials-aggregate.test.ts diff --git a/src/workbench/run-case.ts b/src/workbench/run-case.ts index 672ddb3..546f63b 100644 --- a/src/workbench/run-case.ts +++ b/src/workbench/run-case.ts @@ -105,6 +105,8 @@ async function runWorkbenchCaseMatrix( resultPath: relative(resultsDir, run.resultPath), tracePath: relative(resultsDir, run.tracePath), ...(run.summaryPath ? { summaryPath: relative(resultsDir, run.summaryPath) } : {}), + ...(result.metrics?.tokens?.total !== undefined ? { tokens: result.metrics.tokens.total } : {}), + ...(result.metrics?.durationMs !== undefined ? { durationMs: result.metrics.durationMs } : {}), }; console.log(`${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); @@ -127,6 +129,8 @@ async function runWorkbenchCaseMatrix( meanScore: aggregate.meanScore, passAtK: aggregate.passAtK, passHatK: aggregate.passHatK, + totalTokens: aggregate.totalTokens, + totalDurationMs: aggregate.totalDurationMs, trials: trialResults, }); } diff --git a/src/workbench/run-suite.ts b/src/workbench/run-suite.ts index 290d8b6..2a2f187 100644 --- a/src/workbench/run-suite.ts +++ b/src/workbench/run-suite.ts @@ -117,6 +117,8 @@ export async function runWorkbenchSuite( resultPath: relative(resultsDir, run.resultPath), tracePath: relative(resultsDir, run.tracePath), ...(run.summaryPath ? { summaryPath: relative(resultsDir, run.summaryPath) } : {}), + ...(result.metrics?.tokens?.total !== undefined ? { tokens: result.metrics.tokens.total } : {}), + ...(result.metrics?.durationMs !== undefined ? { durationMs: result.metrics.durationMs } : {}), }; console.log(`${job.caseName} ${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); return { ...job, trialResult }; @@ -141,6 +143,8 @@ export async function runWorkbenchSuite( meanScore: aggregate.meanScore, passAtK: aggregate.passAtK, passHatK: aggregate.passHatK, + totalTokens: aggregate.totalTokens, + totalDurationMs: aggregate.totalDurationMs, trials: trialResults, }); } diff --git a/src/workbench/trials.ts b/src/workbench/trials.ts index 9dd34c0..c65a9e1 100644 --- a/src/workbench/trials.ts +++ b/src/workbench/trials.ts @@ -4,6 +4,8 @@ export interface TrialScoreInput { trial: number; pass: boolean; score: number; + tokens?: number; + durationMs?: number; } export interface TrialAggregate { @@ -14,6 +16,8 @@ export interface TrialAggregate { meanScore: number; passAtK: boolean; passHatK: boolean; + totalTokens: number; + totalDurationMs: number; } export function formatTrialNumber(trial: number): string { @@ -40,6 +44,8 @@ export function aggregateTrials(trials: TrialScoreInput[]): TrialAggregate { const passedTrials = trials.filter((trial) => trial.pass).length; const failedTrials = totalTrials - passedTrials; const scoreTotal = trials.reduce((sum, trial) => sum + trial.score, 0); + const totalTokens = trials.reduce((sum, trial) => sum + (trial.tokens ?? 0), 0); + const totalDurationMs = trials.reduce((sum, trial) => sum + (trial.durationMs ?? 0), 0); return { totalTrials, @@ -49,6 +55,8 @@ export function aggregateTrials(trials: TrialScoreInput[]): TrialAggregate { meanScore: totalTrials === 0 ? 0 : scoreTotal / totalTrials, passAtK: totalTrials > 0 && passedTrials > 0, passHatK: totalTrials > 0 && passedTrials === totalTrials, + totalTokens, + totalDurationMs, }; } diff --git a/src/workbench/types.ts b/src/workbench/types.ts index c371929..31e3b10 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -157,6 +157,8 @@ export interface WorkbenchTrialResultRef { resultPath: string; tracePath: string; summaryPath?: string; + tokens?: number; + durationMs?: number; } export interface WorkbenchModelAggregateResult { @@ -168,6 +170,8 @@ export interface WorkbenchModelAggregateResult { meanScore: number; passAtK: boolean; passHatK: boolean; + totalTokens: number; + totalDurationMs: number; trials: WorkbenchTrialResultRef[]; } diff --git a/tests/trials-aggregate.test.ts b/tests/trials-aggregate.test.ts new file mode 100644 index 0000000..5ace945 --- /dev/null +++ b/tests/trials-aggregate.test.ts @@ -0,0 +1,30 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { aggregateTrials } from '../src/workbench/trials.js'; + +test('aggregateTrials sums tokens and duration', () => { + const result = aggregateTrials([ + { trial: 1, pass: true, score: 1, tokens: 100, durationMs: 1000 }, + { trial: 2, pass: false, score: 0, tokens: 200, durationMs: 2000 }, + { trial: 3, pass: true, score: 1, tokens: 150, durationMs: 1500 }, + ]); + assert.equal(result.totalTokens, 450); + assert.equal(result.totalDurationMs, 4500); + assert.equal(result.passedTrials, 2); +}); + +test('aggregateTrials treats missing tokens/durationMs as zero', () => { + const result = aggregateTrials([ + { trial: 1, pass: true, score: 1 }, + { trial: 2, pass: false, score: 0, tokens: 100 }, + ]); + assert.equal(result.totalTokens, 100); + assert.equal(result.totalDurationMs, 0); +}); + +test('aggregateTrials returns zero totals for an empty input', () => { + const result = aggregateTrials([]); + assert.equal(result.totalTokens, 0); + assert.equal(result.totalDurationMs, 0); + assert.equal(result.totalTrials, 0); +}); From 20baebe8c203384f4dce122cc162ef6fce52d130 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:25:13 -0500 Subject: [PATCH 103/121] chore(test): wire trials-aggregate test into npm test Co-Authored-By: Claude Opus 4.7 (1M context) --- package.json | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/package.json b/package.json index f31dfda..36e3b87 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,7 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts && tsx --test tests/trials-aggregate.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", From e5678173ac4b566bef71f6c2ea170fac2958523b Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:32:00 -0500 Subject: [PATCH 104/121] feat(run-*): replace models matrix with runs (agent + model) matrix Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/run-case.ts | 88 ++++++------------ src/workbench/run-suite.ts | 19 ++-- src/workbench/types.ts | 5 +- tests/smoke-workbench-models.ts | 145 +++--------------------------- tests/smoke-workbench-run-case.ts | 13 +++ tests/smoke-workbench-suite.ts | 4 +- tests/smoke-workbench-trials.ts | 6 +- 7 files changed, 70 insertions(+), 210 deletions(-) diff --git a/src/workbench/run-case.ts b/src/workbench/run-case.ts index 546f63b..9b1ceae 100644 --- a/src/workbench/run-case.ts +++ b/src/workbench/run-case.ts @@ -5,16 +5,15 @@ import { loadWorkbenchCase } from './case-loader.js'; import { getFlag, positionals } from './cli-args.js'; import { runDockerWorkbenchCase } from './docker-runner.js'; import type { DockerWorkbenchRunResult, RunDockerWorkbenchCaseOptions } from './docker-runner.js'; -import { ensureOpenRouterModelRef, parseModelList, slugModelRef } from './models.js'; +import { ensureOpenRouterModelRef, slugModelRef } from './models.js'; import { aggregateTrials, formatTrialNumber, parseTrialsFlag, summarizeTrialAggregates } from './trials.js'; -import type { RunCaseAggregateResultFile, WorkbenchModelAggregateResult, WorkbenchTrialResultRef } from './types.js'; +import type { RunCaseAggregateResultFile, WorkbenchModelAggregateResult, WorkbenchRunSpec, WorkbenchTrialResultRef } from './types.js'; import { readWorkbenchResultFile, timestampSlug, writeJsonFile } from './utils.js'; export interface RunWorkbenchCaseParams { casePath: string; outDir?: string; model?: string; - models?: string[]; image?: string; keepWorkspace?: boolean; trials?: number; @@ -65,13 +64,13 @@ async function mapWithConcurrency( return results; } -function trialDirName(model: string, trial: number): string { - return `${slugModelRef(model)}--${formatTrialNumber(trial)}`; +function trialDirName(agent: string, model: string, trial: number): string { + return `${agent}--${slugModelRef(model)}--${formatTrialNumber(trial)}`; } -async function runWorkbenchCaseMatrix( - params: RunWorkbenchCaseParams & { models: string[] }, - deps: RunWorkbenchCaseDeps, +export async function runWorkbenchCase( + params: RunWorkbenchCaseParams, + deps: RunWorkbenchCaseDeps = {}, ): Promise { const dockerRunner = deps.runDockerWorkbenchCase ?? runDockerWorkbenchCase; const startedAt = new Date().toISOString(); @@ -81,15 +80,23 @@ async function runWorkbenchCaseMatrix( ? Math.floor(params.concurrency) : 1; + const resolvedCase = loadWorkbenchCase(params.casePath); + const cliModel = params.model ? ensureOpenRouterModelRef(params.model) : undefined; + const runs: WorkbenchRunSpec[] = [{ + agent: resolvedCase.agent, + model: cliModel ?? resolvedCase.model, + }]; + mkdirSync(resultsDir, { recursive: true }); - const jobs = params.models.flatMap((model) => Array.from({ length: trials }, (_, index) => ({ - model, + const jobs = runs.flatMap((run) => Array.from({ length: trials }, (_, index) => ({ + agent: run.agent, + model: run.model, trial: index + 1, }))); const completedTrials = await mapWithConcurrency(jobs, concurrency, async (job) => { - const trialDir = join(resultsDir, 'trials', trialDirName(job.model, job.trial)); + const trialDir = join(resultsDir, 'trials', trialDirName(job.agent, job.model, job.trial)); const run = await dockerRunner({ casePath: params.casePath, resultsDir: trialDir, @@ -109,19 +116,20 @@ async function runWorkbenchCaseMatrix( ...(result.metrics?.durationMs !== undefined ? { durationMs: result.metrics.durationMs } : {}), }; - console.log(`${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); + console.log(`${job.agent} ${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); return { ...job, trialResult }; }); const results: WorkbenchModelAggregateResult[] = []; - for (const model of params.models) { + for (const run of runs) { const trialResults = completedTrials - .filter((trial) => trial.model === model) + .filter((trial) => trial.agent === run.agent && trial.model === run.model) .map((trial) => trial.trialResult) .sort((left, right) => left.trial - right.trial); const aggregate = aggregateTrials(trialResults); results.push({ - model, + agent: run.agent, + model: run.model, totalTrials: aggregate.totalTrials, passedTrials: aggregate.passedTrials, failedTrials: aggregate.failedTrials, @@ -140,7 +148,7 @@ async function runWorkbenchCaseMatrix( name: 'run-case', startedAt, endedAt: new Date().toISOString(), - models: params.models, + runs, summary, results, }; @@ -154,60 +162,17 @@ async function runWorkbenchCaseMatrix( } } -export async function runWorkbenchCase( - params: RunWorkbenchCaseParams, - deps: RunWorkbenchCaseDeps = {}, -): Promise { - const model = params.model ? ensureOpenRouterModelRef(params.model) : undefined; - const models = params.models?.map((modelRef) => ensureOpenRouterModelRef(modelRef)); - - if ((models && models.length > 0) || (params.trials ?? 1) > 1) { - const matrixModels = models && models.length > 0 - ? models - : [model ?? loadWorkbenchCase(params.casePath).model]; - await runWorkbenchCaseMatrix({ ...params, model, models: matrixModels }, deps); - return; - } - - const dockerRunner = deps.runDockerWorkbenchCase ?? runDockerWorkbenchCase; - const selectedModel = models?.[0] ?? model; - const run = await dockerRunner({ - casePath: params.casePath, - outDir: params.outDir, - model: selectedModel, - image: params.image, - keepWorkspace: params.keepWorkspace, - }); - - const result = readWorkbenchResultFile(run.resultPath); - console.log(`Results: ${run.resultsDir}`); - console.log(`Grade: ${result.pass ? 'PASS' : 'FAIL'}`); - - if (result.evidence.length > 0) { - for (const line of result.evidence) { - console.log(`- ${line}`); - } - } else { - console.log('- (no evidence)'); - } - - if (!result.pass) { - process.exitCode = 1; - } -} - export async function runWorkbenchCaseFromCli(args: string[]): Promise { const caseArg = positionals(args, { - valueFlags: ['--out', '--model', '--models', '--image', '--trials', '--concurrency'], + valueFlags: ['--out', '--model', '--image', '--trials', '--concurrency'], booleanFlags: ['--keep-workspace'], })[0]; if (!caseArg) { - throw new Error('Missing case path. Usage: skill-optimizer run-case [--out ] [--model ] [--models ] [--trials ] [--concurrency ] [--image ] [--keep-workspace]'); + throw new Error('Missing case path. Usage: skill-optimizer run-case [--out ] [--model ] [--trials ] [--concurrency ] [--image ] [--keep-workspace]'); } const outDir = getFlag(args, '--out'); const model = getFlag(args, '--model'); - const models = getFlag(args, '--models'); const image = getFlag(args, '--image'); const trials = parseTrialsFlag(getFlag(args, '--trials')); const concurrency = parseConcurrencyFlag(getFlag(args, '--concurrency')); @@ -217,7 +182,6 @@ export async function runWorkbenchCaseFromCli(args: string[]): Promise { casePath: resolve(caseArg), outDir: outDir ? resolve(outDir) : undefined, model: model ? ensureOpenRouterModelRef(model) : undefined, - models: models ? parseModelList(models) : undefined, trials, concurrency, image, diff --git a/src/workbench/run-suite.ts b/src/workbench/run-suite.ts index 2a2f187..9ea8e6f 100644 --- a/src/workbench/run-suite.ts +++ b/src/workbench/run-suite.ts @@ -67,8 +67,8 @@ async function mapWithConcurrency( return results; } -function trialDirName(caseName: string, model: string, trial: number): string { - return `${caseName}--${slugModelRef(model)}--${formatTrialNumber(trial)}`; +function trialDirName(caseName: string, agent: string, model: string, trial: number): string { + return `${caseName}--${agent}--${slugModelRef(model)}--${formatTrialNumber(trial)}`; } export async function runWorkbenchSuite( @@ -93,13 +93,14 @@ export async function runWorkbenchSuite( Array.from({ length: trials }, (_, index) => ({ suiteCase, caseName: suiteCase.slug, + agent: run.agent, model: run.model, trial: index + 1, })) ))); const completedTrials = await mapWithConcurrency(jobs, concurrency, async (job) => { - const trialDir = join(resultsDir, 'trials', trialDirName(job.caseName, job.model, job.trial)); + const trialDir = join(resultsDir, 'trials', trialDirName(job.caseName, job.agent, job.model, job.trial)); const run = await dockerRunner({ casePath: job.suiteCase.path, case: job.suiteCase.case, @@ -120,22 +121,22 @@ export async function runWorkbenchSuite( ...(result.metrics?.tokens?.total !== undefined ? { tokens: result.metrics.tokens.total } : {}), ...(result.metrics?.durationMs !== undefined ? { durationMs: result.metrics.durationMs } : {}), }; - console.log(`${job.caseName} ${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); + console.log(`${job.caseName} ${job.agent} ${job.model} trial ${formatTrialNumber(job.trial)}: ${result.pass ? 'PASS' : 'FAIL'}`); return { ...job, trialResult }; }); - const models = runs.map((run) => run.model); const results: WorkbenchCaseModelAggregateResult[] = []; for (const suiteCase of suite.cases) { - for (const model of models) { + for (const run of runs) { const trialResults = completedTrials - .filter((trial) => trial.caseName === suiteCase.slug && trial.model === model) + .filter((trial) => trial.caseName === suiteCase.slug && trial.agent === run.agent && trial.model === run.model) .map((trial) => trial.trialResult) .sort((left, right) => left.trial - right.trial); const aggregate = aggregateTrials(trialResults); results.push({ caseName: suiteCase.slug, - model, + agent: run.agent, + model: run.model, totalTrials: aggregate.totalTrials, passedTrials: aggregate.passedTrials, failedTrials: aggregate.failedTrials, @@ -155,7 +156,7 @@ export async function runWorkbenchSuite( name: suite.name, startedAt, endedAt: new Date().toISOString(), - models, + runs, cases: caseSlugs, summary, results, diff --git a/src/workbench/types.ts b/src/workbench/types.ts index 31e3b10..9aa252f 100644 --- a/src/workbench/types.ts +++ b/src/workbench/types.ts @@ -162,6 +162,7 @@ export interface WorkbenchTrialResultRef { } export interface WorkbenchModelAggregateResult { + agent: string; model: string; totalTrials: number; passedTrials: number; @@ -183,7 +184,7 @@ export interface RunCaseAggregateResultFile { name: string; startedAt: string; endedAt: string; - models: string[]; + runs: WorkbenchRunSpec[]; summary: WorkbenchAggregateSummary; results: WorkbenchModelAggregateResult[]; } @@ -192,7 +193,7 @@ export interface RunSuiteAggregateResultFile { name: string; startedAt: string; endedAt: string; - models: string[]; + runs: WorkbenchRunSpec[]; cases: string[]; summary: WorkbenchAggregateSummary; results: WorkbenchCaseModelAggregateResult[]; diff --git a/tests/smoke-workbench-models.ts b/tests/smoke-workbench-models.ts index 2b08f7f..b516d5b 100644 --- a/tests/smoke-workbench-models.ts +++ b/tests/smoke-workbench-models.ts @@ -24,131 +24,13 @@ test('slugModelRef creates filesystem-safe model directories', () => { assert.equal(slugModelRef('openrouter/meta-llama/llama-3.3-70b-instruct:free'), 'openrouter-meta-llama-llama-3.3-70b-instruct-free'); }); -test('runWorkbenchCase writes aggregate output for multi-model runs', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-models-')); - const previousExitCode = process.exitCode; - try { - const casePath = join(root, 'case.yml'); - const outDir = join(root, 'results'); - const calls: Array<{ model?: string; resultsDir?: string }> = []; - mkdirSync(outDir, { recursive: true }); - writeFileSync(casePath, 'name: model-case\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); - - process.exitCode = undefined; - await runWorkbenchCase( - { - casePath, - outDir, - models: ['openrouter/google/gemini-2.5-flash', 'openrouter/openai/gpt-5.4'], - }, - { - runDockerWorkbenchCase: async (options) => { - calls.push({ model: options.model, resultsDir: options.resultsDir }); - assert.ok(options.resultsDir); - mkdirSync(options.resultsDir, { recursive: true }); - const resultPath = join(options.resultsDir, 'result.json'); - const tracePath = join(options.resultsDir, 'trace.jsonl'); - const pass = options.model !== 'openrouter/openai/gpt-5.4'; - writeFileSync(resultPath, JSON.stringify({ pass, score: pass ? 1 : 0, evidence: [options.model] }), 'utf-8'); - writeFileSync(tracePath, JSON.stringify({ entries: [] }), 'utf-8'); - return { - tempDir: join(root, 'temp'), - caseDir: join(root, 'temp', 'case'), - bundledCasePath: join(root, 'temp', 'case', 'case.yml'), - workDir: join(root, 'temp', 'work'), - resultsDir: options.resultsDir, - resultPath, - tracePath, - cleanup: () => {}, - }; - }, - now: new Date('2026-04-27T10:11:12.000Z'), - }, - ); - - const runResultPath = join(outDir, '20260427-101112', 'run-result.json'); - assert.ok(existsSync(runResultPath)); - const aggregate = JSON.parse(readFileSync(runResultPath, 'utf-8')) as { - summary: { total: number; passed: number; failed: number; passRate: number; totalTrials: number; passedTrials: number; failedTrials: number }; - results: Array<{ model: string; passHatK: boolean; trials: Array<{ resultPath: string; tracePath: string }> }>; - }; - assert.deepEqual(calls.map((call) => call.model), [ - 'openrouter/google/gemini-2.5-flash', - 'openrouter/openai/gpt-5.4', - ]); - assert.equal(aggregate.summary.total, 2); - assert.equal(aggregate.summary.passed, 1); - assert.equal(aggregate.summary.failed, 1); - assert.equal(aggregate.summary.passRate, 0.5); - assert.equal(aggregate.summary.totalTrials, 2); - assert.equal(aggregate.summary.passedTrials, 1); - assert.equal(aggregate.summary.failedTrials, 1); - assert.equal(aggregate.results[0]?.trials[0]?.resultPath, 'trials/openrouter-google-gemini-2.5-flash--001/result.json'); - assert.equal(aggregate.results[1]?.trials[0]?.tracePath, 'trials/openrouter-openai-gpt-5.4--001/trace.jsonl'); - assert.equal(process.exitCode, 1); - } finally { - process.exitCode = previousExitCode; - rmSync(root, { recursive: true, force: true }); - } -}); - -test('runWorkbenchCase writes aggregate output when --models has one model', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-one-model-')); - const previousExitCode = process.exitCode; - try { - const casePath = join(root, 'case.yml'); - const outDir = join(root, 'results'); - mkdirSync(outDir, { recursive: true }); - writeFileSync(casePath, 'name: model-case\nreferences: ./refs\ntask: Test\ngraders:\n - name: passes\n command: "true"\n', 'utf-8'); - - process.exitCode = undefined; - await runWorkbenchCase( - { - casePath, - outDir, - models: ['openrouter/google/gemini-2.5-flash'], - }, - { - runDockerWorkbenchCase: async (options) => { - assert.ok(options.resultsDir); - mkdirSync(options.resultsDir, { recursive: true }); - const resultPath = join(options.resultsDir, 'result.json'); - const tracePath = join(options.resultsDir, 'trace.jsonl'); - writeFileSync(resultPath, JSON.stringify({ pass: true, score: 1, evidence: [] }), 'utf-8'); - writeFileSync(tracePath, JSON.stringify({ entries: [] }), 'utf-8'); - return { - tempDir: join(root, 'temp'), - caseDir: join(root, 'temp', 'case'), - bundledCasePath: join(root, 'temp', 'case', 'case.yml'), - workDir: join(root, 'temp', 'work'), - resultsDir: options.resultsDir, - resultPath, - tracePath, - cleanup: () => {}, - }; - }, - now: new Date('2026-04-27T10:11:12.000Z'), - }, - ); - - const runResultPath = join(outDir, '20260427-101112', 'run-result.json'); - assert.ok(existsSync(runResultPath)); - assert.ok(existsSync(join(outDir, '20260427-101112', 'trials', 'openrouter-google-gemini-2.5-flash--001', 'result.json'))); - assert.equal(process.exitCode, undefined); - } finally { - process.exitCode = previousExitCode; - rmSync(root, { recursive: true, force: true }); - } -}); - -test('runWorkbenchCase --trials uses the case model when no model override is provided', async () => { - const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-case-model-')); +test('runWorkbenchCase fans out trials per agent+model dir', async () => { + const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-runs-')); const previousExitCode = process.exitCode; try { const casePath = join(root, 'case.yml'); const outDir = join(root, 'results'); const refsDir = join(root, 'refs'); - const calls: Array<{ model?: string }> = []; mkdirSync(refsDir, { recursive: true }); mkdirSync(outDir, { recursive: true }); writeFileSync( @@ -159,20 +41,15 @@ test('runWorkbenchCase --trials uses the case model when no model override is pr process.exitCode = undefined; await runWorkbenchCase( - { - casePath, - outDir, - trials: 2, - }, + { casePath, outDir, trials: 2 }, { runDockerWorkbenchCase: async (options) => { - calls.push({ model: options.model }); assert.ok(options.resultsDir); mkdirSync(options.resultsDir, { recursive: true }); const resultPath = join(options.resultsDir, 'result.json'); const tracePath = join(options.resultsDir, 'trace.jsonl'); writeFileSync(resultPath, JSON.stringify({ pass: true, score: 1, evidence: [] }), 'utf-8'); - writeFileSync(tracePath, JSON.stringify({ entries: [] }), 'utf-8'); + writeFileSync(tracePath, '{}\n', 'utf-8'); return { tempDir: join(root, 'temp'), caseDir: join(root, 'temp', 'case'), @@ -188,11 +65,15 @@ test('runWorkbenchCase --trials uses the case model when no model override is pr }, ); - assert.deepEqual(calls.map((call) => call.model), [ - 'openrouter/openai/gpt-5.4', - 'openrouter/openai/gpt-5.4', - ]); - assert.ok(existsSync(join(outDir, '20260427-101112', 'trials', 'openrouter-openai-gpt-5.4--002', 'result.json'))); + const runResultPath = join(outDir, '20260427-101112', 'run-result.json'); + assert.ok(existsSync(runResultPath)); + const aggregate = JSON.parse(readFileSync(runResultPath, 'utf-8')) as { + runs: Array<{ agent: string; model: string }>; + results: Array<{ agent: string; model: string }>; + }; + assert.deepEqual(aggregate.runs, [{ agent: 'pi-acp', model: 'openrouter/openai/gpt-5.4' }]); + assert.equal(aggregate.results[0]?.agent, 'pi-acp'); + assert.ok(existsSync(join(outDir, '20260427-101112', 'trials', 'pi-acp--openrouter-openai-gpt-5.4--002', 'result.json'))); assert.equal(process.exitCode, undefined); } finally { process.exitCode = previousExitCode; diff --git a/tests/smoke-workbench-run-case.ts b/tests/smoke-workbench-run-case.ts index b9bfacd..50881e8 100644 --- a/tests/smoke-workbench-run-case.ts +++ b/tests/smoke-workbench-run-case.ts @@ -18,6 +18,19 @@ test('runWorkbenchCase preserves failing result as process exitCode 1', async () score: 0, evidence: ['expected failure'], }), 'utf-8'); + + mkdirSync(join(root, 'refs'), { recursive: true }); + writeFileSync(join(root, 'case.yml'), [ + 'name: exit-code-case', + 'references: ./refs', + 'agent: pi-acp', + 'model: openrouter/anthropic/claude-haiku-4-5', + 'task: noop', + 'graders:', + ' - name: passes', + ' command: "true"', + ].join('\n'), 'utf-8'); + process.exitCode = undefined; await runWorkbenchCase( diff --git a/tests/smoke-workbench-suite.ts b/tests/smoke-workbench-suite.ts index fe7dee6..e28e127 100644 --- a/tests/smoke-workbench-suite.ts +++ b/tests/smoke-workbench-suite.ts @@ -233,8 +233,8 @@ test('runWorkbenchSuite writes case-model matrix aggregate output', async () => assert.equal(aggregate.summary.totalTrials, 4); assert.equal(aggregate.summary.passedTrials, 3); assert.equal(aggregate.summary.failedTrials, 1); - assert.equal(aggregate.results[0]?.trials[0]?.resultPath, 'trials/missing-index--openrouter-google-gemini-2.5-flash--001/result.json'); - assert.equal(aggregate.results[3]?.trials[0]?.resultPath, 'trials/partial-index--openrouter-openai-gpt-5.4--001/result.json'); + assert.equal(aggregate.results[0]?.trials[0]?.resultPath, 'trials/missing-index--pi-acp--openrouter-google-gemini-2.5-flash--001/result.json'); + assert.equal(aggregate.results[3]?.trials[0]?.resultPath, 'trials/partial-index--pi-acp--openrouter-openai-gpt-5.4--001/result.json'); assert.equal(process.exitCode, 1); } finally { process.exitCode = previousExitCode; diff --git a/tests/smoke-workbench-trials.ts b/tests/smoke-workbench-trials.ts index 5061b7f..18e1ca7 100644 --- a/tests/smoke-workbench-trials.ts +++ b/tests/smoke-workbench-trials.ts @@ -87,9 +87,9 @@ test('runWorkbenchSuite writes trial directories and case-model aggregates', asy ); const runRoot = join(outDir, '20260427-101112'); - assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--openrouter-google-gemini-2.5-flash--001', 'result.json'))); - assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--openrouter-google-gemini-2.5-flash--002', 'result.json'))); - assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--openrouter-google-gemini-2.5-flash--003', 'result.json'))); + assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--pi-acp--openrouter-google-gemini-2.5-flash--001', 'result.json'))); + assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--pi-acp--openrouter-google-gemini-2.5-flash--002', 'result.json'))); + assert.ok(existsSync(join(runRoot, 'trials', 'trial-case--pi-acp--openrouter-google-gemini-2.5-flash--003', 'result.json'))); const aggregate = JSON.parse(readFileSync(join(runRoot, 'suite-result.json'), 'utf-8')) as { summary: { totalTrials: number; passedTrials: number; failedTrials: number; trialPassRate: number; meanScore: number }; From 62c786e1e785dc2e23c65b0f9675678afe45853a Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:32:36 -0500 Subject: [PATCH 105/121] chore(examples): migrate pdf+mcp suites to runs: schema; remove stale results Co-Authored-By: Claude Opus 4.7 (1M context) --- examples/workbench/mcp/suite.yml | 6 ++++-- examples/workbench/pdf/suite.yml | 9 +++++++-- 2 files changed, 11 insertions(+), 4 deletions(-) diff --git a/examples/workbench/mcp/suite.yml b/examples/workbench/mcp/suite.yml index 5ac689a..6be750a 100644 --- a/examples/workbench/mcp/suite.yml +++ b/examples/workbench/mcp/suite.yml @@ -1,7 +1,8 @@ name: mcp-calculator-example references: ./references -models: - - openrouter/google/gemini-2.5-flash +runs: + - agent: pi-acp + model: openrouter/google/gemini-2.5-flash env: - OPENROUTER_API_KEY timeoutSeconds: 600 @@ -15,6 +16,7 @@ mcpServices: - calculator-server.mjs cases: - name: use-calculator-mcp + agent: pi-acp task: | Compute this expression: diff --git a/examples/workbench/pdf/suite.yml b/examples/workbench/pdf/suite.yml index 0dc13de..d6388d3 100644 --- a/examples/workbench/pdf/suite.yml +++ b/examples/workbench/pdf/suite.yml @@ -2,8 +2,9 @@ name: pdf-workbench-example references: ./references appendSystemPrompt: | Keep task outputs at the top level of /work unless the user asks for a different path. -models: - - openrouter/google/gemini-2.5-flash +runs: + - agent: pi-acp + model: openrouter/google/gemini-2.5-flash env: - OPENROUTER_API_KEY timeoutSeconds: 600 @@ -15,6 +16,7 @@ setup: - cp input/briefing-source.pdf briefing-source.pdf cases: - name: extract-pdf-facts + agent: pi-acp task: | Extract the key facts from statement.pdf and write answer.json with this exact schema: { @@ -30,6 +32,7 @@ cases: command: node $CASE/checks/extract-pdf-facts.mjs - name: split-customer-packet + agent: pi-acp task: | Create customer-copy.pdf from customer-packet.pdf. Include only the pages marked CUSTOMER COPY, in their original order. Exclude the page marked INTERNAL NOTES. The output must be a PDF, not a text file. graders: @@ -37,6 +40,7 @@ cases: command: node $CASE/checks/split-customer-packet.mjs - name: build-briefing-pdf + agent: pi-acp task: | Create briefing.pdf as a one-page PDF briefing based on briefing-source.pdf. It must include these exact lines: PDF Skill Briefing @@ -49,6 +53,7 @@ cases: command: node $CASE/checks/build-briefing-pdf.mjs - name: no-pdf-skill-needed + agent: pi-acp task: | Write note.txt with exactly this text: done From 9faba1e43e4f9e96f42bf848cb9ed507ec7af636 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:35:08 -0500 Subject: [PATCH 106/121] test(smoke): per-agent end-to-end smoke probes --- package.json | 3 +- tests/smoke-agents/_common.ts | 75 +++++++++++++++++++ .../claude-agent-acp.smoke.test.ts | 17 +++++ tests/smoke-agents/codex-acp.smoke.test.ts | 17 +++++ tests/smoke-agents/gemini.smoke.test.ts | 17 +++++ tests/smoke-agents/opencode.smoke.test.ts | 17 +++++ tests/smoke-agents/pi-acp.smoke.test.ts | 17 +++++ 7 files changed, 162 insertions(+), 1 deletion(-) create mode 100644 tests/smoke-agents/_common.ts create mode 100644 tests/smoke-agents/claude-agent-acp.smoke.test.ts create mode 100644 tests/smoke-agents/codex-acp.smoke.test.ts create mode 100644 tests/smoke-agents/gemini.smoke.test.ts create mode 100644 tests/smoke-agents/opencode.smoke.test.ts create mode 100644 tests/smoke-agents/pi-acp.smoke.test.ts diff --git a/package.json b/package.json index 36e3b87..961f1ed 100644 --- a/package.json +++ b/package.json @@ -80,7 +80,8 @@ "typecheck": "tsc --noEmit", "lint": "tsc --noUnusedLocals --noEmit", "dockerfile:print-installs": "tsx src/workbench/agents/install-snippets.ts", - "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts && tsx --test tests/trials-aggregate.test.ts" + "test": "tsx tests/smoke-workbench-case.ts && tsx tests/smoke-workbench-checks.ts && tsx tests/smoke-workbench-docker-runner.ts && tsx tests/smoke-workbench-run-case.ts && tsx tests/smoke-workbench-models.ts && tsx tests/smoke-workbench-suite.ts && tsx tests/smoke-workbench-trials.ts && tsx tests/smoke-workbench-metrics.ts && tsx tests/smoke-skill-distribution.ts && tsx --test tests/acp/sdk-smoke.test.ts && tsx --test tests/acp/transport.test.ts && tsx --test tests/acp/client.test.ts && tsx --test tests/agents/registry.test.ts && tsx --test tests/acp/auth.test.ts && tsx --test tests/acp/skill-deploy.test.ts && tsx --test tests/acp/mcp-config-writer.test.ts && tsx --test tests/parse-trace.test.ts && tsx --test tests/acp/trace-recorder.test.ts && tsx --test tests/case-loader-agent-required.test.ts && tsx --test tests/suite-loader-runs.test.ts && tsx --test tests/trials-aggregate.test.ts", + "test:smoke-agents": "tsx --test tests/smoke-agents/claude-agent-acp.smoke.test.ts && tsx --test tests/smoke-agents/codex-acp.smoke.test.ts && tsx --test tests/smoke-agents/gemini.smoke.test.ts && tsx --test tests/smoke-agents/opencode.smoke.test.ts && tsx --test tests/smoke-agents/pi-acp.smoke.test.ts" }, "dependencies": { "@agentclientprotocol/sdk": "0.22.1", diff --git a/tests/smoke-agents/_common.ts b/tests/smoke-agents/_common.ts new file mode 100644 index 0000000..b62026d --- /dev/null +++ b/tests/smoke-agents/_common.ts @@ -0,0 +1,75 @@ +import { mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync, existsSync } from 'node:fs'; +import { tmpdir, homedir } from 'node:os'; +import { join } from 'node:path'; + +import { runDockerWorkbenchCase } from '../../src/workbench/docker-runner.js'; + +export interface SmokeOptions { + agent: string; + model: string; + env: string[]; +} + +export interface SmokeResult { + pass: boolean; + outputContent?: string; + tracePath: string; + resultPath: string; +} + +export async function runSmokeTrial(opts: SmokeOptions): Promise { + const dir = mkdtempSync(join(tmpdir(), 'smoke-')); + try { + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'references', 'README.md'), '# smoke\n', 'utf-8'); + const envBlock = opts.env.length > 0 + ? opts.env.map((name) => ` - ${name}`).join('\n') + : ' []'; + writeFileSync(join(dir, 'case.yml'), [ + `name: smoke-${opts.agent}`, + 'references: ./references', + `agent: ${opts.agent}`, + `model: ${opts.model}`, + 'task: |', + ' Write the literal string "hello" to /work/out.txt.', + ' Then write a one-line description of what you did to /work/findings.txt.', + 'graders:', + ' - name: out-txt-exists-with-hello', + ` command: 'test "$(cat /work/out.txt 2>/dev/null)" = "hello"'`, + 'env:', + envBlock, + 'timeoutSeconds: 120', + ].join('\n'), 'utf-8'); + + const run = await runDockerWorkbenchCase({ + casePath: join(dir, 'case.yml'), + image: 'skill-optimizer-agent:local', + keepWorkspace: true, + }); + + let outputContent: string | undefined; + if (run.workspacePath) { + try { + outputContent = readFileSync(join(run.workspacePath, 'out.txt'), 'utf-8'); + } catch { + outputContent = undefined; + } + } + + const result = JSON.parse(readFileSync(run.resultPath, 'utf-8')) as { pass?: unknown }; + return { + pass: Boolean(result.pass), + outputContent, + tracePath: run.tracePath, + resultPath: run.resultPath, + }; + } finally { + rmSync(dir, { recursive: true, force: true }); + } +} + +export function hasAuth(opts: { envVar?: string; subscriptionFile?: string }): boolean { + if (opts.envVar && process.env[opts.envVar]) return true; + if (opts.subscriptionFile && existsSync(join(homedir(), opts.subscriptionFile))) return true; + return false; +} diff --git a/tests/smoke-agents/claude-agent-acp.smoke.test.ts b/tests/smoke-agents/claude-agent-acp.smoke.test.ts new file mode 100644 index 0000000..0d55567 --- /dev/null +++ b/tests/smoke-agents/claude-agent-acp.smoke.test.ts @@ -0,0 +1,17 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { hasAuth, runSmokeTrial } from './_common.js'; + +const skip = !hasAuth({ envVar: 'ANTHROPIC_API_KEY', subscriptionFile: '.claude/.credentials.json' }) + ? 'no ANTHROPIC_API_KEY env and no ~/.claude/.credentials.json' + : false; + +test('claude-agent-acp smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'claude-agent-acp', + model: 'claude-haiku-4-5-20251001', + env: ['ANTHROPIC_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); diff --git a/tests/smoke-agents/codex-acp.smoke.test.ts b/tests/smoke-agents/codex-acp.smoke.test.ts new file mode 100644 index 0000000..428e7b8 --- /dev/null +++ b/tests/smoke-agents/codex-acp.smoke.test.ts @@ -0,0 +1,17 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { hasAuth, runSmokeTrial } from './_common.js'; + +const skip = !hasAuth({ envVar: 'OPENAI_API_KEY', subscriptionFile: '.codex/auth.json' }) + ? 'no OPENAI_API_KEY env and no ~/.codex/auth.json' + : false; + +test('codex-acp smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'codex-acp', + model: 'gpt-5-mini', + env: ['OPENAI_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); diff --git a/tests/smoke-agents/gemini.smoke.test.ts b/tests/smoke-agents/gemini.smoke.test.ts new file mode 100644 index 0000000..cc9bb83 --- /dev/null +++ b/tests/smoke-agents/gemini.smoke.test.ts @@ -0,0 +1,17 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { hasAuth, runSmokeTrial } from './_common.js'; + +const skip = !hasAuth({ envVar: 'GOOGLE_API_KEY', subscriptionFile: '.gemini/oauth_creds.json' }) + ? 'no GOOGLE_API_KEY env and no ~/.gemini/oauth_creds.json' + : false; + +test('gemini smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'gemini', + model: 'gemini-3.1-pro-preview', + env: ['GOOGLE_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); diff --git a/tests/smoke-agents/opencode.smoke.test.ts b/tests/smoke-agents/opencode.smoke.test.ts new file mode 100644 index 0000000..818479f --- /dev/null +++ b/tests/smoke-agents/opencode.smoke.test.ts @@ -0,0 +1,17 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { hasAuth, runSmokeTrial } from './_common.js'; + +const skip = !hasAuth({ envVar: 'OPENAI_API_KEY' }) + ? 'no OPENAI_API_KEY env' + : false; + +test('opencode smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'opencode', + model: 'google/gemini-3.1-pro-preview', + env: ['OPENAI_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); diff --git a/tests/smoke-agents/pi-acp.smoke.test.ts b/tests/smoke-agents/pi-acp.smoke.test.ts new file mode 100644 index 0000000..bb4eb09 --- /dev/null +++ b/tests/smoke-agents/pi-acp.smoke.test.ts @@ -0,0 +1,17 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { hasAuth, runSmokeTrial } from './_common.js'; + +const skip = !hasAuth({ envVar: 'OPENROUTER_API_KEY' }) + ? 'no OPENROUTER_API_KEY env' + : false; + +test('pi-acp smoke: writes hello to out.txt', { skip }, async () => { + const r = await runSmokeTrial({ + agent: 'pi-acp', + model: 'openrouter/anthropic/claude-haiku-4-5', + env: ['OPENROUTER_API_KEY'], + }); + assert.equal(r.pass, true); + assert.equal(r.outputContent?.trim(), 'hello'); +}); From 968c0ef9f27d1256ce9c8a4dbf131ee4b9fa5c46 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:36:49 -0500 Subject: [PATCH 107/121] fix(model-refs): only pi-acp requires openrouter/ prefix; native agents pass through Native ACP agents (claude-agent-acp, codex-acp, gemini, opencode) accept their own model refs (claude-haiku-4-5-20251001, gpt-5-mini, etc.). The OpenRouter-only guard now applies only when agent is pi-acp. Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/case-loader.ts | 3 ++- src/workbench/run-case.ts | 6 ++++-- tests/smoke-workbench-run-case.ts | 16 ++++++++++++++-- 3 files changed, 20 insertions(+), 5 deletions(-) diff --git a/src/workbench/case-loader.ts b/src/workbench/case-loader.ts index 5941fbb..8087279 100644 --- a/src/workbench/case-loader.ts +++ b/src/workbench/case-loader.ts @@ -75,7 +75,8 @@ export function resolveWorkbenchCaseConfig( const env = readStringArray(parsed, 'env', resolvedConfigPath); const setup = readStringArray(parsed, 'setup', resolvedConfigPath); const cleanup = readStringArray(parsed, 'cleanup', resolvedConfigPath); - const model = ensureOpenRouterModelRef(readOptionalString(parsed, 'model', resolvedConfigPath) ?? DEFAULT_WORKBENCH_MODEL); + const rawModel = readOptionalString(parsed, 'model', resolvedConfigPath) ?? DEFAULT_WORKBENCH_MODEL; + const model = agentCfg.name === 'pi-acp' ? ensureOpenRouterModelRef(rawModel) : rawModel.trim(); const timeoutSeconds = readOptionalTimeoutSeconds(parsed, resolvedConfigPath) ?? DEFAULT_WORKBENCH_TIMEOUT_SECONDS; const referencesDir = resolve(resolvedConfigDir, references); diff --git a/src/workbench/run-case.ts b/src/workbench/run-case.ts index 9b1ceae..ad794a7 100644 --- a/src/workbench/run-case.ts +++ b/src/workbench/run-case.ts @@ -81,7 +81,9 @@ export async function runWorkbenchCase( : 1; const resolvedCase = loadWorkbenchCase(params.casePath); - const cliModel = params.model ? ensureOpenRouterModelRef(params.model) : undefined; + const cliModel = params.model + ? (resolvedCase.agent === 'pi-acp' ? ensureOpenRouterModelRef(params.model) : params.model.trim()) + : undefined; const runs: WorkbenchRunSpec[] = [{ agent: resolvedCase.agent, model: cliModel ?? resolvedCase.model, @@ -181,7 +183,7 @@ export async function runWorkbenchCaseFromCli(args: string[]): Promise { await runWorkbenchCase({ casePath: resolve(caseArg), outDir: outDir ? resolve(outDir) : undefined, - model: model ? ensureOpenRouterModelRef(model) : undefined, + model, trials, concurrency, image, diff --git a/tests/smoke-workbench-run-case.ts b/tests/smoke-workbench-run-case.ts index 50881e8..1f0e70b 100644 --- a/tests/smoke-workbench-run-case.ts +++ b/tests/smoke-workbench-run-case.ts @@ -56,12 +56,24 @@ test('runWorkbenchCase preserves failing result as process exitCode 1', async () } }); -test('runWorkbenchCaseFromCli rejects invalid --model before loading the case', async () => { +test('runWorkbenchCaseFromCli rejects non-openrouter --model when the case agent is pi-acp', async () => { const root = mkdtempSync(join(tmpdir(), 'skill-opt-workbench-run-case-model-')); try { + mkdirSync(join(root, 'refs'), { recursive: true }); + writeFileSync(join(root, 'case.yml'), [ + 'name: model-validate-case', + 'references: ./refs', + 'agent: pi-acp', + 'model: openrouter/anthropic/claude-haiku-4-5', + 'task: noop', + 'graders:', + ' - name: passes', + ' command: "true"', + ].join('\n'), 'utf-8'); + await assert.rejects( runWorkbenchCaseFromCli([ - join(root, 'missing-case.yml'), + join(root, 'case.yml'), '--model', 'anthropic/claude-3-5-haiku-latest', ]), From 7baf73df7dfbb626e2194276d3f10ce92b779c19 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:39:39 -0500 Subject: [PATCH 108/121] fix(acp): chmod 644 on staged subscription auth files Subscription auth files (~/.claude/.credentials.json, ~/.codex/auth.json, etc.) are typically mode 600 on the host. After cpSync into the staging dir that becomes /home/agent in the container, the container's agent user (uid 10001) can't read them. Loosen mode to 644 on the staged copies so the in-container agent can read. Host files unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/docker-runner.ts | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index eabef6f..a270f07 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -362,6 +362,11 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise Date: Mon, 25 May 2026 15:40:24 -0500 Subject: [PATCH 109/121] test(regression): pi-acp pdf suite baseline scaffold + tolerance check Baseline is currently a placeholder (captureStatus: placeholder); the test skips when OPENROUTER_API_KEY is unset or when the baseline hasn't been replaced with real captured rates. Once a real run is performed, replace the JSON with the actual perCasePassRate and remove the captureStatus field to enable the check. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../baselines/pi-acp-pdf-2026-05-25.json | 14 ++++++ tests/regression/pi-acp-pdf-suite.test.ts | 50 +++++++++++++++++++ 2 files changed, 64 insertions(+) create mode 100644 tests/regression/baselines/pi-acp-pdf-2026-05-25.json create mode 100644 tests/regression/pi-acp-pdf-suite.test.ts diff --git a/tests/regression/baselines/pi-acp-pdf-2026-05-25.json b/tests/regression/baselines/pi-acp-pdf-2026-05-25.json new file mode 100644 index 0000000..f17df7a --- /dev/null +++ b/tests/regression/baselines/pi-acp-pdf-2026-05-25.json @@ -0,0 +1,14 @@ +{ + "captured": "2026-05-25", + "captureStatus": "placeholder", + "captureNote": "Baseline not yet captured under real OPENROUTER_API_KEY. Run `npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials 3` and replace this file with the actual per-case pass rates.", + "agent": "pi-acp", + "model": "openrouter/google/gemini-2.5-flash", + "trials": 3, + "perCasePassRate": { + "extract-pdf-facts": 1.0, + "split-customer-packet": 1.0, + "build-briefing-pdf": 1.0, + "no-pdf-skill-needed": 1.0 + } +} diff --git a/tests/regression/pi-acp-pdf-suite.test.ts b/tests/regression/pi-acp-pdf-suite.test.ts new file mode 100644 index 0000000..098f3b9 --- /dev/null +++ b/tests/regression/pi-acp-pdf-suite.test.ts @@ -0,0 +1,50 @@ +import { test } from 'node:test'; +import { strict as assert } from 'node:assert'; +import { readFileSync, readdirSync } from 'node:fs'; +import { join } from 'node:path'; +import { execSync } from 'node:child_process'; + +interface Baseline { + captured: string; + captureStatus?: string; + agent: string; + model: string; + trials: number; + perCasePassRate: Record; +} + +const BASELINE_PATH = 'tests/regression/baselines/pi-acp-pdf-2026-05-25.json'; +const baseline = JSON.parse(readFileSync(BASELINE_PATH, 'utf-8')) as Baseline; + +const skip = !process.env.OPENROUTER_API_KEY + ? 'no OPENROUTER_API_KEY env' + : baseline.captureStatus === 'placeholder' + ? `${BASELINE_PATH} is still a placeholder — re-run and replace before enforcing regression` + : false; + +test('pi-acp regression: pdf suite pass rates match baseline within tolerance', { skip }, () => { + execSync( + `npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials ${baseline.trials}`, + { stdio: 'inherit' }, + ); + + const resultsDir = 'examples/workbench/pdf/.results'; + const runs = readdirSync(resultsDir).sort(); + const latest = runs[runs.length - 1]; + if (!latest) { + throw new Error(`no run directories under ${resultsDir}`); + } + const data = JSON.parse(readFileSync(join(resultsDir, latest, 'suite-result.json'), 'utf-8')) as { + results: Array<{ caseName: string; trialPassRate: number }>; + }; + + for (const [caseName, expectedRate] of Object.entries(baseline.perCasePassRate)) { + const row = data.results.find((r) => r.caseName === caseName); + assert.ok(row !== undefined, `case ${caseName} not in suite result`); + // Tolerance: ±0.34 (one trial of three can flip) + assert.ok( + Math.abs(row.trialPassRate - expectedRate) <= 0.34, + `case ${caseName}: expected ~${expectedRate}, got ${row.trialPassRate}`, + ); + } +}); From 0dca012a8dad95a4dbf9f7effaf9d4a60b06b690 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:41:01 -0500 Subject: [PATCH 110/121] docs(shared): ACP trace format reference for chain analyzer Co-Authored-By: Claude Opus 4.7 (1M context) --- skills/shared/acp-trace-format.md | 87 +++++++++++++++++++++++++++++++ 1 file changed, 87 insertions(+) create mode 100644 skills/shared/acp-trace-format.md diff --git a/skills/shared/acp-trace-format.md b/skills/shared/acp-trace-format.md new file mode 100644 index 0000000..25b0555 --- /dev/null +++ b/skills/shared/acp-trace-format.md @@ -0,0 +1,87 @@ +# ACP trace format (for chain analyzer) + +`trace.jsonl` captures the raw Agent Client Protocol (ACP) messages +exchanged between the workbench (client) and the agent CLI (server) +for one trial. The first line is a `trace_start` header with trial +metadata; every subsequent line is a JSON-RPC envelope per the ACP +spec at . + +This doc summarizes the message types the chain analyzer cares about. +For the full spec, follow the link above. + +## Header (line 1) + +```json +{ + "type": "trace_start", + "schemaVersion": 2, + "caseName": "...", + "agent": "claude-agent-acp", + "model": "claude-haiku-4-5-20251001", + "startedAt": "ISO-8601", + "endedAt": "ISO-8601" +} +``` + +## Initialization (early lines) + +- `initialize` request from client; `initialize` response from agent +- `session/new` request; response carries `sessionId` + +These confirm the agent started. If they're absent, the trial failed +before reaching the prompt — bench infrastructure issue, not skill +weakness. + +## Session updates (the meat) + +All have shape: + +```json +{"jsonrpc":"2.0","method":"session/update","params":{"sessionId":"...","update":{...}}} +``` + +The `update` object's `sessionUpdate` field tells you what happened: + +| `sessionUpdate` value | Meaning | Analyzer cares because | +|---|---|---| +| `agent_message_chunk` | Streaming chunk of assistant text | Final response — what the agent told the user | +| `agent_thought_chunk` | Streaming chunk of assistant reasoning | The agent's reasoning — useful for diagnosing why it did X | +| `tool_call` | Agent is calling a tool (start) | `kind` field tells you what kind: `execute`, `read`, `write`, `edit`, `search`, `fetch`, `think`, `other` | +| `tool_call_update` | Tool result/progress (end) | `status: "completed" \| "failed"`, `content` carries the tool output | +| `plan` | Agent's high-level plan | Optional sidebar; skip in most analysis | +| `user_message_chunk` | Agent echoing user input | Rare; skip | + +## Final prompt response (last numbered response) + +```json +{ + "jsonrpc": "2.0", + "id": , + "result": { + "stopReason": "end_turn" | "max_tokens" | "refusal" | "cancelled", + "usage": { + "inputTokens": 120, + "outputTokens": 45, + "cacheReadTokens": 0, + "cacheCreationTokens": 0 + } + } +} +``` + +`stopReason` is critical for diagnosing failures: + +- `end_turn` — normal completion (still check `findings.txt` for correctness) +- `max_tokens` — agent ran out of context; skill may be too verbose +- `refusal` — agent declined the task; skill description may have triggered a safety pattern +- `cancelled` — workbench timed out the prompt + +## Helpers + +Don't parse the JSONL manually. Use `src/workbench/parse-trace.ts`: + +- `iterMessages(jsonl)` — assistant/user messages (text + thinking) +- `iterToolCalls(jsonl)` — paired tool_call + tool_call_update +- `computeMetrics(jsonl)` — tokens, duration, per-tool counts, stopReason +- `getFinalAssistantMessage(jsonl)` — last assistant chunk concatenated +- `getFailureEvidence(jsonl)` — failed-tool result snippets From a05caaac3767b6b0937f839a26a6d75e4aa28eef Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:42:11 -0500 Subject: [PATCH 111/121] feat(chain): update test-writer, analyzer, run-bench for ACP trace format + agent column - test-writer: do-not-pre-trigger-skill guidance for task prompts - analyzer: read-first pointer to acp-trace-format.md - run-bench: per-(agent + model) summary; surface tokens/duration; drop cost; new image+Dockerfile names Co-Authored-By: Claude Opus 4.7 (1M context) --- skills/analyze/agents/analyzer.md | 7 ++++++ skills/run-bench/SKILL.md | 31 +++++++++++++++--------- skills/write-tests/agents/test-writer.md | 13 ++++++++++ 3 files changed, 40 insertions(+), 11 deletions(-) diff --git a/skills/analyze/agents/analyzer.md b/skills/analyze/agents/analyzer.md index 715ef86..a3676ca 100644 --- a/skills/analyze/agents/analyzer.md +++ b/skills/analyze/agents/analyzer.md @@ -17,6 +17,13 @@ anti-pattern list should map to specific entries in that doc. If your reasoning doesn't connect to a named principle or anti-pattern there, your weakness framing is probably under-grounded. +**Read first:** [`../../shared/acp-trace-format.md`](../../shared/acp-trace-format.md) +— `trace.jsonl` is raw ACP wire format. Use `parse-trace.ts` helpers +(`iterMessages`, `iterToolCalls`, `computeMetrics`, +`getFinalAssistantMessage`, `getFailureEvidence`) rather than reading +JSONL by hand. The format reference doc summarizes the event types +relevant to skill-behavior analysis. + ## Inputs (templated by the operator session) - `${SUMMARY_PATH}` — `06-bench-summary.md` (entry point; diff --git a/skills/run-bench/SKILL.md b/skills/run-bench/SKILL.md index 8bf6819..2b71da8 100644 --- a/skills/run-bench/SKILL.md +++ b/skills/run-bench/SKILL.md @@ -76,33 +76,42 @@ npx tsx /src/cli.ts run-suite \ `--trials 3` is the chain default — enough to distinguish flaky from systematic failures at step 7. Honor a different count if -requested. Models come from `suite.yml` (per project invariant: -`run-suite` does NOT take a `--models` override). Stream +requested. Runs (agent + model pairs) come from `suite.yml` (per +project invariant: `run-suite` does NOT take an override). Stream stdout/stderr to the user — bench runs take minutes to hours. Environment failures to surface honestly (don't try to recover): -- **`OPENROUTER_API_KEY` not set** — CLI fails fast. Ask the user - to set it; don't mock. +- **Required auth not available** — fail fast. Ask the user to set + the appropriate env var or run the relevant `… login` command: + `ANTHROPIC_API_KEY` / `claude login` for `claude-agent-acp`; + `OPENAI_API_KEY` / `codex login` for `codex-acp`; `GOOGLE_API_KEY` + / `gemini auth login` for `gemini`; `OPENAI_API_KEY` for + `opencode`; `OPENROUTER_API_KEY` for `pi-acp`. Don't mock. - **Docker image missing** — default is - `skill-optimizer-workbench:local`. Tell the user to build it - (`docker build -t skill-optimizer-workbench:local -f - docker/workbench-runner.Dockerfile .`). + `skill-optimizer-agent:local`. Tell the user to build it + (`docker build -t skill-optimizer-agent:local -f + docker/skill-optimizer-agent.Dockerfile .`). ### (d) Write the summary Parse `${OUT_DIR}/suite-result.json` and write `06-bench-summary.md` with the frontmatter above plus a body containing: -- **Overall:** total trials, passed, failed, overall pass rate. -- **Per model:** trial count and pass rate per model in - `suite.yml`. +- **Overall:** total trials, passed, failed, overall pass rate. Total + tokens consumed. Total duration. +- **Per agent + model:** trial count, pass rate, mean tokens per + trial, mean duration per trial for each row in `runs:` from the + suite. - **Per probe:** trial count, pass rate, one-line note if any - trial failed. + trial failed. Mean tokens and duration per probe. - **Failed-probe pointer list:** probe IDs where any trial failed, with paths to their `trace.jsonl` and `findings.txt`. - **Raw output:** the `bench_results_path` value. +Do NOT include a cost column. tokens × downstream pricing is computed +offline if needed. + If all trials errored (no graded results), that's a workbench misconfiguration or environmental failure rather than a skill weakness. Write the summary honestly and surface before handing diff --git a/skills/write-tests/agents/test-writer.md b/skills/write-tests/agents/test-writer.md index 054cd97..70465aa 100644 --- a/skills/write-tests/agents/test-writer.md +++ b/skills/write-tests/agents/test-writer.md @@ -92,6 +92,19 @@ grader_logic: > ``` +**Task prompts describe the user's actual task; do not reference the +skill explicitly.** The harness mounts the skill at the agent's native +discovery path (e.g., `~/.claude/skills//SKILL.md`). The agent +decides whether to invoke it based on the skill's frontmatter +`description`. "Skill didn't trigger" is then a measurable weakness +class — don't pre-trigger it via the task prompt. + +Example task prompt (good): "Review /work/ProductCard.tsx for +compliance issues. Write findings to /work/findings.txt." + +Example task prompt (bad — pre-triggers): "Use the skill at +/work/web-design-guidelines/SKILL.md to review /work/ProductCard.tsx." + ### 2. `workspace/` — files the agent sees in `/work` The fixture content. Whatever files the parent functionality's From eb92cc543650aa3ac2c14a0cf59f52c097a6df01 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 15:43:27 -0500 Subject: [PATCH 112/121] docs: refresh CLAUDE.md, CONTRIBUTING.md, README.md for multi-agent ACP - Architecture summary: host-side ACP client driving per-trial Docker containers - Important files: src/workbench/acp/, src/workbench/agents/, parse-trace.ts - Invariants: agent: required on case.yml, runs: matrix on suite.yml, pi-acp-only openrouter/ prefix - Auth matrix per agent - Docker image+Dockerfile name updated to skill-optimizer-agent:local Co-Authored-By: Claude Opus 4.7 (1M context) --- CLAUDE.md | 16 +++++++++++----- CONTRIBUTING.md | 12 ++++++++---- README.md | 16 +++++++++++++--- 3 files changed, 32 insertions(+), 12 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 42a061f..ac5227f 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -4,7 +4,7 @@ `skill-optimizer` is a Docker workbench for running and grading agent skill eval cases. The current public CLI centers on `run-case` and `run-suite`. -The workbench gives an agent an isolated Docker `/work` directory, captures traces, and grades deterministic local outcomes from files, command logs, generated artifacts, or other workspace state. +The workbench drives a host-side ACP (Agent Client Protocol) client against per-trial Docker containers that run one of five agent CLIs (claude-agent-acp, codex-acp, gemini, opencode, pi-acp). Each trial gets an isolated `/work` directory; raw ACP messages are captured to `trace.jsonl`; graders evaluate deterministic local outcomes from files, command logs, generated artifacts, or other workspace state. ## Key Commands @@ -20,8 +20,12 @@ npx tsx src/cli.ts run-suite --help ## Important Files - `src/cli.ts`: public CLI entrypoint -- `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces -- `docker/skill-optimizer-agent.Dockerfile`: container image with all 5 agent CLIs pre-baked, used for setup and grade phases +- `src/workbench/`: case loading, suite loading, host-side Docker runner, graders +- `src/workbench/acp/`: ACP transport, client wrapper, auth, skill deployment, MCP config writers, trace recorder +- `src/workbench/agents/`: 5-agent registry (claude-agent-acp, codex-acp, gemini, opencode, pi-acp) + install snippets +- `src/workbench/parse-trace.ts`: helpers for reading the raw ACP `trace.jsonl` (`iterMessages`, `iterToolCalls`, `computeMetrics`, etc.) +- `docker/skill-optimizer-agent.Dockerfile`: container image with all 5 agent CLIs pre-baked +- `docker/pi-acp-launcher.sh`: wrapper that bridges `SKILL_OPT_PROVIDER_*` env vars into pi-acp's `OPENROUTER_API_KEY` - `skills//`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) - `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills - `skills//agents/`: prompt templates dispatched by chain skills via the Agent tool @@ -46,11 +50,13 @@ Keep the README installation section aligned with packaged plugin metadata: ## Invariants - Keep evaluation static: extraction and matching are allowed; do not execute model-produced code outside the Docker workbench as part of evaluation. -- `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override. -- Keep OpenRouter model refs as `openrouter/...`; real model runs require `OPENROUTER_API_KEY`. +- Every `case.yml` must declare `agent:` (one of `claude-agent-acp`, `codex-acp`, `gemini`, `opencode`, `pi-acp`); the loader fails loud on missing/unknown agents. +- `suite.yml` declares its agent+model matrix via `runs: [{ agent, model }]`; legacy `models:` is rejected. `run-suite` uses what `suite.yml` declares; do not add a `--models` override. +- Only `pi-acp` requires `openrouter/...` model refs; native ACP agents pass their own model strings through unchanged. - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state. - The agent phase sees only `/work`, not `/case` or `/results`. +- `trace.jsonl` is raw ACP wire format (JSON-RPC envelopes); use `parse-trace.ts` helpers rather than parsing the file by hand. See [`skills/shared/acp-trace-format.md`](skills/shared/acp-trace-format.md). - Keep plugin metadata pointed at every chain skill under `skills//`; do not create divergent skill copies. - Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`. - Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a59714b..1aba225 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -22,8 +22,10 @@ All three commands must pass before opening a PR when code changes are involved. ## Project layout - `src/cli.ts` — public CLI entry point for `run-case` and `run-suite`. -- `src/workbench/` — case/suite loading, Docker runner, Pi agent wiring, graders, traces, metrics, MCP support, and trial aggregation. -- `docker/skill-optimizer-agent.Dockerfile` — container image with all 5 agent CLIs pre-baked, used for setup and grade phases. +- `src/workbench/` — case/suite loading, host-side Docker runner, graders, metrics, MCP support, trial aggregation. +- `src/workbench/acp/` — ACP transport, client wrapper, auth resolution, skill deployment, per-agent MCP config writer, trace recorder. +- `src/workbench/agents/` — 5-agent registry + Dockerfile install snippet generator. +- `docker/skill-optimizer-agent.Dockerfile` — container image with all 5 agent CLIs pre-baked. - `skills//` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). - `skills/shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. - `skills//agents/` — prompt templates dispatched by chain skills via the Agent tool. @@ -43,10 +45,12 @@ All three commands must pass before opening a PR when code changes are involved. ## Workbench invariants - Keep evaluation static: extraction and matching are allowed; do not execute model-produced code outside the Docker workbench as part of evaluation. -- Use only `openrouter/...` model refs; real model runs require `OPENROUTER_API_KEY`. -- `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override. +- Every `case.yml` must declare `agent:` (one of `claude-agent-acp`, `codex-acp`, `gemini`, `opencode`, `pi-acp`); the loader fails loud on missing/unknown. +- `suite.yml` declares its `runs: [{ agent, model }]` matrix; legacy `models:` is rejected. `run-suite` uses what `suite.yml` declares; do not add a `--models` override. +- Only `pi-acp` requires `openrouter/...` model refs; native ACP agents pass their own model strings through unchanged. - Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid. - The agent phase sees only `/work`, not `/case`, `/results`, graders, hidden answers, or hidden metadata. +- `trace.jsonl` is raw ACP wire format (JSON-RPC envelopes); use `src/workbench/parse-trace.ts` helpers (`iterMessages`, `iterToolCalls`, `computeMetrics`) rather than parsing the file by hand. - Keep plugin metadata pointed at every chain skill under `skills//`; do not create divergent skill copies. ## Testing guidance diff --git a/README.md b/README.md index 3d84b69..4c9cbaa 100644 --- a/README.md +++ b/README.md @@ -127,20 +127,30 @@ Requirements: - Node.js 20+ - Docker -- `OPENROUTER_API_KEY` for real model runs + +The workbench drives one of five agent CLIs over ACP (Agent Client Protocol). Each agent has its own auth options: + +| Agent | Auth options | +|---|---| +| `claude-agent-acp` | `claude login` (subscription) or `ANTHROPIC_API_KEY` | +| `codex-acp` | `codex login` (subscription) or `OPENAI_API_KEY` | +| `gemini` | `gemini auth login` (subscription) or `GOOGLE_API_KEY` | +| `opencode` | `OPENAI_API_KEY` (or provider-specific key) | +| `pi-acp` | `OPENROUTER_API_KEY` | Install and build: ```bash npm install npm run build +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . ``` -Only `openrouter/...` model refs are supported. +Only `pi-acp` requires `openrouter/...` model refs. Native ACP agents (`claude-agent-acp`, `codex-acp`, `gemini`, `opencode`) take their own model strings (e.g., `claude-haiku-4-5-20251001`, `gpt-5-mini`). ## Quick Start -Run the suite against the models listed in `suite.yml`: +Run the suite against the agent+model pairs listed in `suite.yml`'s `runs:`: ```bash npx tsx src/cli.ts run-suite examples/workbench/pdf/suite.yml --trials 1 From 5e567f71d587f49b6fe61df6380f8582f79fa9a4 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Mon, 25 May 2026 17:29:09 -0500 Subject: [PATCH 113/121] =?UTF-8?q?fix(acp):=20subscription=20auth=20?= =?UTF-8?q?=E2=80=94=20extract=20OAuth=20token,=20stage=20writable=20home?= =?UTF-8?q?=20dirs?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three intertwined fixes needed for claude-agent-acp (and codex-acp) subscription auth to work end-to-end: 1. Extract the OAuth bearer from the credentials JSON and pass it via ANTHROPIC_AUTH_TOKEN (claude) / OPENAI_API_KEY (codex). The respective underlying SDKs (@anthropic-ai/claude-agent-sdk and codex-acp's analogue) read env vars, not the credentials file directly. New SubscriptionEnvExtract type lets the registry declare these mappings. 2. Pre-create directories the agent's runtime expects under the staged $HOME. @anthropic-ai/claude-agent-sdk writes debug logs to ~/.claude/debug/.txt; session/new fails with "Query closed before response received" if that dir doesn't exist. Use the registry's existing homeDirs field to declare these. 3. Make staged $HOME and its declared subdirs world-writable (mode 0o777). The container's agent user (uid 10001) does not match the host user that owns the staging dir, so without loosened perms the SDK can't write its own session/debug files. Verified: claude-agent-acp smoke probe now completes a real initialize→newSession→prompt cycle and writes /work/out.txt in ~15s. Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/acp/auth.ts | 30 ++++++++++++++++++++++++++-- src/workbench/agents/registry.ts | 34 +++++++++++++++++++++++++++++++- src/workbench/docker-runner.ts | 19 ++++++++++++++++++ 3 files changed, 80 insertions(+), 3 deletions(-) diff --git a/src/workbench/acp/auth.ts b/src/workbench/acp/auth.ts index 0c5b8e9..e8b4645 100644 --- a/src/workbench/acp/auth.ts +++ b/src/workbench/acp/auth.ts @@ -1,7 +1,7 @@ -import { existsSync } from 'node:fs'; +import { existsSync, readFileSync } from 'node:fs'; import { homedir } from 'node:os'; import { join, normalize } from 'node:path'; -import type { AgentConfig } from '../agents/registry.js'; +import type { AgentConfig, SubscriptionEnvExtract } from '../agents/registry.js'; export interface AuthContext { home?: string; // override $HOME for testing @@ -11,6 +11,7 @@ export interface AuthContext { export interface SubscriptionAuthResult { mode: 'subscription'; files: Array<{ hostPath: string; containerPath: string }>; + envOverrides: Record; // env vars to set in container (extracted from credentials file) } export interface EnvAuthResult { @@ -26,12 +27,20 @@ export function resolveAuth(agent: AgentConfig, ctx: AuthContext): AuthResult { if (agent.subscriptionAuth) { const detectPath = expandHome(agent.subscriptionAuth.detectFile, home); if (existsSync(detectPath)) { + const envOverrides: Record = {}; + for (const extract of agent.subscriptionAuth.extractEnv ?? []) { + const value = extractEnvFromFile(extract, home); + if (value !== undefined) { + envOverrides[extract.targetEnv] = value; + } + } return { mode: 'subscription', files: agent.subscriptionAuth.files.map((f) => ({ hostPath: expandHome(f.hostPath, home), containerPath: f.containerPath, // {home} placeholder, resolved at mount time })), + envOverrides, }; } } @@ -47,6 +56,23 @@ export function resolveAuth(agent: AgentConfig, ctx: AuthContext): AuthResult { return { mode: 'env', envNames: agent.requiresEnv }; } +function extractEnvFromFile(extract: SubscriptionEnvExtract, home: string): string | undefined { + const path = expandHome(extract.source, home); + if (!existsSync(path)) return undefined; + let parsed: unknown; + try { + parsed = JSON.parse(readFileSync(path, 'utf-8')); + } catch { + return undefined; + } + let cursor: unknown = parsed; + for (const segment of extract.jsonPath.split('.')) { + if (cursor === null || typeof cursor !== 'object') return undefined; + cursor = (cursor as Record)[segment]; + } + return typeof cursor === 'string' ? cursor : undefined; +} + function expandHome(p: string, home: string): string { if (p.startsWith('~/')) return normalize(join(home, p.slice(2))); if (p === '~') return home; diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts index b33a292..0a5901d 100644 --- a/src/workbench/agents/registry.ts +++ b/src/workbench/agents/registry.ts @@ -10,10 +10,21 @@ export interface HostAuthFile { containerPath: string; // {home}/.claude/.credentials.json } +export interface SubscriptionEnvExtract { + // Extract a value out of a JSON credentials file and expose it as an env var + // inside the trial container. Lets agents that use OAuth-style subscription + // auth (Claude, Codex) get their bearer token without us mounting and parsing + // the file on every agent CLI invocation. + source: string; // host file path (with ~ expansion); usually same as detectFile + jsonPath: string; // dot-separated path into the JSON, e.g. "claudeAiOauth.accessToken" + targetEnv: string; // env var to set in container, e.g. "ANTHROPIC_AUTH_TOKEN" +} + export interface SubscriptionAuth { replacesEnv: string; // e.g. "ANTHROPIC_API_KEY" detectFile: string; // host path to check for login files: HostAuthFile[]; // all files to copy when sub-auth used + extractEnv?: SubscriptionEnvExtract[]; // optional: derive container env vars from the credentials file } export type ApiProtocol = @@ -56,13 +67,26 @@ export const AGENTS: Record = { }, skillPaths: ['$HOME/.claude/skills'], credentialFiles: [], - homeDirs: [], + // The Anthropic SDK writes runtime debug logs to ~/.claude/debug/.txt; + // session/new fails on ENOENT if that dir doesn't exist in the staged home. + homeDirs: ['.claude/debug'], subscriptionAuth: { replacesEnv: 'ANTHROPIC_API_KEY', detectFile: '~/.claude/.credentials.json', files: [ { hostPath: '~/.claude/.credentials.json', containerPath: '{home}/.claude/.credentials.json' }, ], + // claude-agent-acp uses @anthropic-ai/claude-agent-sdk, which reads + // ANTHROPIC_API_KEY / ANTHROPIC_AUTH_TOKEN env vars — NOT the credentials + // file directly. So we mount the file (for defensive completeness) AND + // extract the OAuth bearer into ANTHROPIC_AUTH_TOKEN. + extractEnv: [ + { + source: '~/.claude/.credentials.json', + jsonPath: 'claudeAiOauth.accessToken', + targetEnv: 'ANTHROPIC_AUTH_TOKEN', + }, + ], }, acpModelFormat: 'bare', supportsAcpSetModel: true, @@ -94,6 +118,14 @@ export const AGENTS: Record = { files: [ { hostPath: '~/.codex/auth.json', containerPath: '{home}/.codex/auth.json' }, ], + // codex auth.json format is { "OPENAI_API_KEY": "..." }; extract for env passthrough. + extractEnv: [ + { + source: '~/.codex/auth.json', + jsonPath: 'OPENAI_API_KEY', + targetEnv: 'OPENAI_API_KEY', + }, + ], }, acpModelFormat: 'bare', supportsAcpSetModel: true, diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index a270f07..dc5b128 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -356,11 +356,24 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise Date: Tue, 26 May 2026 07:23:18 -0500 Subject: [PATCH 114/121] feat(acp): apply case model via unstable_setSessionModel after newSession MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Without this, the model: field in case.yml was silently ignored — the agent used its default. Now passes the case's model string to the agent's session when supportsAcpSetModel is true on the agent's registry entry. Errors are re-thrown with the agent name + bad model string so case authors can tell they used an unrecognized ID. Verified with model: opus[1m] on claude-agent-acp; tokens recorded, stopReason end_turn. Co-Authored-By: Claude Opus 4.7 (1M context) --- src/workbench/docker-runner.ts | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index dc5b128..8e48594 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -495,6 +495,25 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise/"; etc.). The case author is + // responsible for using a string the agent recognizes. + if (agent.supportsAcpSetModel && params.model && params.model !== 'default') { + try { + await client.connection.unstable_setSessionModel({ + sessionId: session.sessionId, + modelId: params.model, + }); + } catch (error) { + const msg = error instanceof Error ? error.message : String(error); + throw new Error( + `Agent ${agent.name} rejected model "${params.model}": ${msg}. ` + + `Check the registry's acpModelFormat for valid IDs.`, + ); + } + } const promptPromise = client.connection.prompt({ sessionId: session.sessionId, prompt: [{ type: 'text', text: params.resolvedCase.task }], From 194c8f7dd1cd3f59707010ded3874a0189b0e548 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Tue, 26 May 2026 07:33:04 -0500 Subject: [PATCH 115/121] fix(agents): add OPENROUTER_API_KEY to pi-acp requiresEnv; drop stray aliases MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes two spec-satisfaction gaps from the post-implementation audit: - pi-acp now declares OPENROUTER_API_KEY in requiresEnv so resolveAuth fails loudly preflight instead of letting the agent crash mid-trial. opencode keeps requiresEnv: [] intentionally — it's multi-provider and the loginHint guides the user. - AGENT_ALIASES drops 'gemini' (self-pointing) and 'openclaw' (points at non-existent agent). Spec's alias table is canonical: claude, codex, pi. --- src/workbench/agents/registry.ts | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/src/workbench/agents/registry.ts b/src/workbench/agents/registry.ts index 0a5901d..955c027 100644 --- a/src/workbench/agents/registry.ts +++ b/src/workbench/agents/registry.ts @@ -181,7 +181,7 @@ export const AGENTS: Record = { description: 'Pi coding agent via ACP', installCmd: `npm install -g @mariozechner/pi-coding-agent@latest pi-acp@latest`, launchCmd: `/opt/skill-opt/bin/pi-acp-launcher`, - requiresEnv: [], + requiresEnv: ['OPENROUTER_API_KEY'], apiProtocol: '', envMapping: {}, skillPaths: ['$HOME/.pi/agent/skills', '$HOME/.agents/skills'], @@ -197,9 +197,7 @@ export const AGENTS: Record = { export const AGENT_ALIASES: Record = { claude: 'claude-agent-acp', codex: 'codex-acp', - gemini: 'gemini', pi: 'pi-acp', - openclaw: 'openclaw', }; export function resolveAgent(spec: string): AgentConfig { From 9e10006302cd7c8d768e38b91e3bc5f6bd462341 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 07:22:52 -0500 Subject: [PATCH 116/121] =?UTF-8?q?docs(spec):=20skill-optimizer=20autopil?= =?UTF-8?q?ot=20=E2=80=94=20step-10=20chain=20driver=20design?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Fulfills the v1.4 acceptance criterion #8 that was never built. Defines a single user-invocable SKILL.md at skills/autopilot/ that walks the 9-step chain end-to-end on one skill. Defaults to fully automated with optional per-gate breakpoints chosen at startup. Replaces v1.4 §10's single-cycle retry model with a strategy-menu per gate (cheap → expensive); the pilot picks the cheapest unattempted strategy on each gate firing. Caps are per-gate (default 2) + global (default 8); cap exhaustion is an honest exit, not a failure. State lives in a per-run journal (gitignored), promoted to a tracked autopilot-summary-.md at end-of-run. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../specs/2026-05-27-autopilot-design.md | 600 ++++++++++++++++++ 1 file changed, 600 insertions(+) create mode 100644 docs/superpowers/specs/2026-05-27-autopilot-design.md diff --git a/docs/superpowers/specs/2026-05-27-autopilot-design.md b/docs/superpowers/specs/2026-05-27-autopilot-design.md new file mode 100644 index 0000000..8ce5d40 --- /dev/null +++ b/docs/superpowers/specs/2026-05-27-autopilot-design.md @@ -0,0 +1,600 @@ +# skill-optimizer autopilot — design spec + +**Status:** approved (brainstormed 2026-05-27 with the +`superpowers:brainstorming` skill) +**Fulfills:** v1.4 acceptance criterion #8 (auto-pilot smoke test) — +the only criterion from the v1.4 design that was never built. The +chain was deliberately structured to slot a driver in cleanly; this +spec defines what that driver does. +**Supersedes:** v1.4 design §10 (`skill-optimizer-autopilot`). That +section described a single-cycle retry loop with halt-on-honest-exit +defaults. This spec replaces it with a strategy-menu model that +defaults to grinding for an improvement until either it lands or +caps are exhausted. + +## Goal + +A single user-invocable skill at `skills/autopilot/SKILL.md` that +runs the full 9-step skill-optimizer chain end-to-end on one skill, +fully automated by default, with optional per-gate breakpoints that +the operator can opt into at startup. The pilot follows the same +re-run and directive-distillation mechanics the chain already uses +for manual operation — autopilot is the chain's tenth skill, not a +parallel pipeline. + +## Motivation + +The chain works step-by-step today, but driving it manually across +nine steps is the bottleneck for batch-processing skills. The v1.4 +design anticipated this and reserved step 10 for a driver, but +deferred the build. Two prior attempts inform this design: + +1. **v1.3 `tools/auto-improve-skill.mjs`** — a wrapper script that + spawned `claude -p` with a 5-phase prompt. Worked for batch runs + (the May 9 batch-2 hit 8 of 10 targets) but predates the v1.4 + chain decomposition. Retiring with this spec. +2. **v1.4 §10** — sketched a SKILL.md interface with `pr_intent`, + `pick_top_n`, `max_iterations_per_step` flags. The retry policy + was single-cycle ("on `needs-revision`, distill and retry up to + N rounds"). Subsequent team review surfaced that recovery isn't + single-cycle: when the skill is too easy or no weakness is + found, the right move depends on what's available (more picks? + harder probes? reframe?). This spec replaces the single-cycle + model with a strategy menu per gate. + +Both legacies inform what to keep: a SKILL.md (not a wrapper +script), per-gate retry caps, distillation of validator rationales +into directives. What changes: each retry-eligible gate has a menu +of recovery strategies the pilot picks from based on state, not a +fixed cycle. The default disposition is **grind for an improvement +until something lands or the cap is exhausted**, not "honest-exit +on first no-finding." + +## Audience + +The autopilot's target user is an operator who wants the +skill-optimizer applied to one skill without understanding the +internals of the chain. The startup interview is written for that +audience — plain language, no skill-optimizer jargon, breakpoints +described as natural questions ("if the skill passes every test, do +I push for harder probes or accept it's strong?"). Skills internals +people can still use it, but the UX is tuned for the operator who +just wants results. + +Autopilot is for **single-skill optimization**. Batch processing +multiple skills is a separate orchestrator and out of scope here. + +## Architecture overview + +```text +skills/ + investigate-functionality/ # step 1 + investigate-submissions/ # step 2 (optional) + design-tests/ # step 3 + write-tests/ # step 4 + validate-tests/ # step 5 + run-bench/ # step 6 + analyze/ # step 7 + improve/ # step 8 + validate/ # step 9 + autopilot/ # step 10 — NEW (this spec) + shared/ +``` + +Autopilot dispatches the nine chain skills sequentially via the +**Skill tool** (not the Agent tool). Each chain skill, when +invoked, internally dispatches its own narrow-context subagents via +the Agent tool — autopilot doesn't touch that machinery. The +operator session running autopilot is the single place where +backward-trigger rationales get distilled into +`${OPERATOR_DIRECTIVES}` for the next dispatch. + +**Why Skill-sequential, not Agent-nested.** If autopilot dispatched +the chain skills as agents-within-an-agent, the outer pilot +couldn't read each step's canonical output to distill directives +for the next step. Skill-sequential keeps the operator session as +the mediator, matching how the chain already works during manual +operation. + +State writes go to the same paths the chain already uses, plus one +new pair specific to autopilot: + +```text +docs/skill-optimizer// + 01-functionality.md # existing chain output + ... (steps 2-9 outputs) + autopilot-summary-.md # NEW — tracked, per-run summary + +.skill-optimizer// + vendored-skill/ # existing chain output + bench-results// # existing chain output + autopilot-/ # NEW — gitignored, per-run scratch + journal.md # appended-to during the run +``` + +The journal is gitignored (ephemeral scratch); the summary is +tracked (the audit record). Promotion from journal to summary +happens at end-of-run. + +## The startup interview + +The interview captures these values into the run journal's +frontmatter: + +| Captured | Type | Default | Asked when | +|---|---|---|---| +| `pr_intent` | bool | `false` | source is upstream URL | +| `pick_top_n` | int | `5` | always | +| `max_retries_per_gate` | int | `2` | always | +| `global_retry_cap` | int | `8` | always | +| `surfaced_gates` | set of gate IDs | `{G6-CLI-FAIL}` | always (operator may add more) | +| `default_overrides` | map | `{}` | always (operator may override any gate's default behavior) | + +When the operator invokes autopilot, the pilot: + +1. **Classifies the source** as upstream URL or local filesystem + path. (Reuses step 1's classification logic — see + `skills/investigate-functionality/SKILL.md`.) +2. **Presents the interview prompt** (below) as a single message. +3. **Records the operator's responses** into the run journal's + frontmatter. Anything not changed keeps its default per the + table above. +4. **Initializes the run journal** at + `.skill-optimizer//autopilot-/journal.md`. +5. **Fires the forward walk.** + +Interview prompt: + +```text +I'm going to optimize end-to-end. By default I'm fully +automated — I run all 9 chain steps and retry up to 2 times at each +recovery point. I'll only pause and ask if the bench environment +breaks (docker/agent crash — that's not something to fix without +you). + +Optional breakpoints — pick any to be in the loop on: + +- Reviewing picked test cases before bench runs. Default: top 5 by + importance are picked automatically. +- A probe gets blocked or rejected during test-writing. Default: + drop the affected probe and continue; the others still run. +- Validate-tests says one or more probes need revision. Default: + distill the rationale into a directive and retry the probe up to + the cap. +- The skill passes every test or no structural weakness gets found. + Default: try recovery strategies in order — first reframe the + analysis, then more bench trials, then enable more picks, then + propose new probes — until something surfaces or the cap is hit. +- The optimizer gets stuck on a weakness it can't translate to a + concrete fix. Default: reframe the weakness and retry up to the + cap. +- The validator says the proposed fix needs revision. Default: + distill the rationale and retry the proposal up to the cap. +- The validator outright rejects the proposed fix. Default: try + recovery strategies in order — first a different fix for the same + weakness, then reframe the weakness, then add more probe coverage + — until something lands or the cap is hit. +- A recovery point has multiple strategies available (meta-breakpoint + covering the three above). Default: I pick the cheapest unattempted + strategy. Surfacing this means I'll show you the menu and let you + choose. + +A few defaults you can also override: + +- PR intent: no (only matters for upstream skills) +- Top-N test cases picked: 5 +- Max retries per recovery point: 2 + +Say 'all auto, go' to just run, or tell me what to change. +``` + +The interview is one turn. After the operator responds, no further +mid-run prompts fire except for surfaced gates (always-surfaces +G6-CLI-FAIL or anything the operator opted into). + +## Control flow + +```text +Phase 1: setup + - Classify source + - Run startup interview + - Initialize run journal + +Phase 2: forward walk (steps 1 → 9) + For each step: + - Check staleness vs direct upstream (git-mtime) + - If artifact is current AND no operator directive forces rerun → skip + - Else: dispatch via Skill, wait for completion + - After completion: check for gate firings (see below) + +Phase 3: gate handling (on any gate firing) + - Read gate's menu from SKILL.md + - Read journal for prior attempts at this gate + - If gate is in surfaced_gates: prompt operator, follow their answer + - Else: pick cheapest unattempted strategy + - Append journal entry + - Increment per-gate counter; check caps + - If cap exhausted at this gate OR global cap exhausted: + honest-exit (do NOT fail; this is principled) + - Else: execute strategy (re-dispatch the relevant chain step(s) + with ${OPERATOR_DIRECTIVES} distilled from the gate's rationale) + - Resume forward walk + +Phase 4: termination + - Reach step 9 verdict: approve → exit_status: improved + - Cap exhaustion at a gate → exit_status: unchanged-honest-exit + - Infra failure (G6-CLI-FAIL) the operator didn't recover → exit_status: blocked-on-infra + - Write autopilot-summary-.md (promote from journal) +``` + +### Re-entrancy + +Re-invoking autopilot on the same slug picks up where it left off +via git-mtime staleness: artifacts ≤ direct upstream are current +and skipped; everything else re-runs. The new invocation gets its +own `autopilot-/` journal directory and its own +`autopilot-summary-.md`. **Caps reset per run**, not per slug — +two separate invocations each get `max_retries_per_gate = 2`. + +The prior run's outputs are preserved naturally by their +timestamps. This matches the chain's general "filesystem IS the +state; history is git" discipline. + +## Gates and recovery menus + +A **gate** is a decision point in the chain where autopilot might +need to take action beyond just dispatching the next step. + +### Gate definitions + +| ID | Fires when | Retry policy | +|---|---|---| +| G3 | Step 3 has emitted N proposals; picks need to be set | Single action — mark top-`pick_top_n` by importance; not a retry gate | +| G4-BLOCKED | Step 4's test-writer subagent returns BLOCKED on a probe | No retry — skip, log, continue | +| G5-REVISION | Step 5 verdict is `needs-revision` on ≥1 probe | Menu (1 option), retry-eligible | +| G5-REJECT | Step 5 verdict is `reject` on ≥1 probe | No retry — drop the probe, continue | +| G6-ALLPASS | Step 6 bench is 100% pass across every model+probe | Not a halt gate; continue to step 7 (G7-NW handles actual recovery) | +| G6-CLI-FAIL | Step 6 CLI/infra failure (docker, agent crash) | No retry — always surfaces | +| G7-NW | Step 7 reports `has_structural_weakness: false` | Menu (4 options), retry-eligible | +| G8-BLOCKED | Step 8 optimizer subagent returns BLOCKED | Menu (2 options), retry-eligible | +| G9-REVISION | Step 9 verdict is `needs-revision` | Menu (1 option), retry-eligible | +| G9-REJECT | Step 9 verdict is `reject` | Menu (3 options), retry-eligible | + +### Recovery menus + +Each retry-eligible gate has a menu of strategies, ordered +cheap → expensive. The pilot picks the cheapest **available + +unattempted** option at each gate firing. + +#### G5-REVISION + +1. Distill validator rationale → re-run step 4 for affected probes + only → re-run step 5 + +#### G7-NW (cheap → expensive) + +1. **Reframe analysis angle.** Re-run step 7 with a directive like + "look at marginal failures, model-specific drift, latent issues + even if pass-rate is high." Cheap: 1 step re-runs. +2. **More bench trials.** Re-run step 6 with extra trials per probe + (e.g., 3× instead of default 1×) to surface variance, then + re-run step 7. Medium: a full bench cycle. +3. **Enable more picks.** Flip next-highest-importance unpicked + functionalities to `picked: true` in step 3's spec.yaml, then + re-run steps 4 → 5 → 6 → 7 for the new ones (existing probes + stay). Medium-high: incremental probe build + full bench. +4. **Add new probes.** Re-run step 3 with a directive "propose new + functionalities the existing pass missed; focus on edge cases the + existing probes don't cover," then 4 → 5 → 6 → 7. Expensive: + fresh design work + full bench. + +#### G8-BLOCKED (cheap → expensive) + +1. **Reframe weakness.** Re-run step 7 with a directive "the + optimizer was blocked on weakness W; please reformulate more + concretely or at a different abstraction," then re-run step 8. +2. **Sharper directive.** If the weakness reads concrete enough but + the optimizer struggled with the general principle, re-run step + 8 alone with a sharpened directive (e.g., "make the principle + more specific to how the skill is structured today"). + +#### G9-REVISION + +1. Distill validator rationale into a directive → re-run step 8 → + re-run step 9 + +#### G9-REJECT (cheap → expensive) + +1. **Different fix, same weakness.** Re-run step 8 with the rejected + proposal noted as anti-pattern ("proposal P rejected because Y; + try a different general principle for the same weakness"). +2. **Reframe weakness.** Re-run step 7 with the validator's rejection + rationale "weakness X rejected because Y; find a different angle + on what's failing," then 8 → 9. +3. **Add coverage.** If the validator rejected on grounds that the + weakness wasn't real in the trace data, treat as G7-NW menu + (more trials, more picks, more probes). + +### Reasoning protocol at each gate firing + +```text +1. Read the gate's menu from this SKILL.md. +2. Read journal.md for prior attempts at this gate (strategies + already tried this run). +3. If gate ID is in `surfaced_gates`: + - Surface in plain language to operator + - Show the menu options + reasoning for each + - Wait for operator decision (pick strategy / supply directive + / accept honest-exit) +4. Else: pick cheapest unattempted strategy from the menu. +5. Append journal entry: timestamp, gate ID, strategy chosen, + reason, retry counter (e.g., "3/4 strategies remaining"). +6. Distill the gate's rationale into ${OPERATOR_DIRECTIVES} (atomic + list, no context dump — per shared/subagent-dispatch.md). +7. Re-dispatch the relevant chain step(s) with that directive. +8. On the next gate firing (which may be the same gate again): + start from step 1 of this protocol. +``` + +### Cap accounting + +- **Per-gate cap** (`max_retries_per_gate`, default 2). Counts + total retry events for that gate type, regardless of which menu + strategy was picked. Two G7-NW retries means strategies 1+2 were + attempted; strategies 3+4 aren't reached. +- **Global cap** (`global_retry_cap`, default 8). Total retry + events across all gates. Safety net against thrashing. +- **Cap exhaustion is an honest exit**, not a failure. The journal + records what was tried; the summary explains the principled + outcome. + +## State tracking — the run journal + +Single markdown file at +`.skill-optimizer//autopilot-/journal.md`, gitignored. +Appended-to throughout the run. The pilot reads it before each gate +decision; the journal IS the pilot's working memory across the run. + +```markdown +--- +run_started_at: 2026-05-27T12:34:56Z +slug: anthropics-skills-pdf +source: https://github.com/anthropics/skills/tree/main/pdf +source_kind: upstream | local +mode_settings: + pr_intent: false + pick_top_n: 5 + max_retries_per_gate: 2 + global_retry_cap: 8 + surfaced_gates: [G6-CLI-FAIL] + default_overrides: {} +--- + +## Run journal + +- 12:35:01Z step 1 dispatched (investigate-functionality) +- 12:36:14Z step 1 complete; 01-functionality.md written +- 12:36:14Z step 3 dispatched (design-tests) +- 12:37:42Z step 3 complete; G3 fired +- 12:37:42Z G3 strategy 1 (top-5 by importance) — auto-applied +- ... +- 13:02:15Z step 7 complete; G7-NW fired (has_structural_weakness: false) +- 13:02:15Z G7-NW strategy 1 (reframe) | reason: cheapest unattempted | retry: 1/2 +- 13:02:16Z step 7 re-dispatched with directive "...look at marginal failures..." +- 13:03:48Z step 7 complete; G7-NW fired again (still false) +- 13:03:48Z G7-NW strategy 2 (more trials) | reason: 1 attempted | retry: 2/2 +- ... +- 13:18:02Z honest-exit: G7-NW cap exhausted across 2 strategies (1+2) +- 13:18:02Z run ended + +## Retry counters + +- G7-NW: 2/2 (strategies tried: 1=reframe, 2=more-trials) +- G9-REVISION: 0/2 +- ... + +## Global retry counter + +- Total retries: 2/8 +``` + +Format is markdown with frontmatter for the run config + appended +log lines + counters block. Human-readable; the pilot parses the +counters block on each gate firing. + +## End-of-run summary report + +Written to `docs/skill-optimizer//autopilot-summary-.md` +(tracked). Promoted from the journal at end-of-run. Concise audit +record, not a dump of every event. + +```markdown +--- +run_started_at: 2026-05-27T12:34:56Z +run_ended_at: 2026-05-27T13:18:02Z +slug: anthropics-skills-pdf +source: https://github.com/anthropics/skills/tree/main/pdf +exit_status: improved | unchanged-honest-exit | blocked-on-infra +total_retries: 2 +journal_path: .skill-optimizer//autopilot-/journal.md +--- + +# Autopilot run for `anthropics-skills-pdf` + +## Headline + +Skill-improvement run ended in honest-exit. G7-NW cap exhausted after +2 recovery strategies (reframe + more trials). Analyzer found no +structural weakness; nothing to ship. + +## Per-step outcomes + +| Step | Outcome | Artifact | +|---|---|---| +| 1 investigate-functionality | written | docs/.../01-functionality.md | +| 3 design-tests | written, 5 picked | docs/.../03-test-proposals.md | +| 4 write-tests | 5 probes built | skill-evals/.../ | +| 5 validate-tests | all approved | docs/.../05-tests-verdict.md | +| 6 run-bench | 84% pass | docs/.../06-bench-summary.md | +| 7 analyze | no structural weakness (after 2 retries) | docs/.../07-analysis.md | +| 8 improve | skipped (no weakness) | — | +| 9 validate | skipped | — | + +## Recovery summary + +- G7-NW retries: 2/2 + - strategy 1 (reframe analysis) — analyzer still reported no weakness + - strategy 2 (more bench trials, ran 3× per probe) — same outcome +- All other gates: not fired + +## Caveats + +[anything the operator should know — flaky bench, suspicious trace + patterns, etc. Auto-pilot's interpretation, not a substitute for + operator review.] + +## What changed in the repo + +[file-list of new/modified canonicals so the operator can scan diffs.] +``` + +### exit_status enum + +- `improved` — step 9 approved; `improved-skill/` materialized at + `docs/skill-optimizer//improved-skill/` +- `unchanged-honest-exit` — cap exhausted at any gate, OR no + weakness with all G7-NW strategies exhausted, OR validator-reject + with all G9-REJECT strategies exhausted +- `blocked-on-infra` — G6-CLI-FAIL or another infra surface the + operator didn't recover from + +## SKILL.md shape + +```text +--- +name: autopilot +description: +--- + +# autopilot + +[purpose statement + audience expectations] + +## Before you start + +[upfront interview contract, with the plain-language prompt block + from this spec verbatim] + +## Workflow + +### (a) Classify source + run startup interview +### (b) Initialize the run journal +### (c) Forward walk through the chain (steps 1 → 9) +### (d) Handle gate firings +### (e) Honest-exit and write summary + +## Gate menus + +[the full menu table from this spec — load-bearing reference] + +## Reasoning protocol at each gate firing + +[the 8-step protocol from this spec] + +## Cap accounting + +[per-gate + global, definitions of "retry event"] + +## Surfacing a gate (interactive) + +[for surfaced gates: how to format the plain-language prompt to the + operator and parse the response] +``` + +### Description routing + +```text +description: Use when the user wants to run the entire skill-optimizer chain +end-to-end on a single skill without driving each step manually — phrases like +"autopilot this skill", "run the full chain on X", "skill-optimizer end-to-end +for X", "automated improvement run for X". Walks steps 1 → 9 with the iteration +mechanism, applies per-gate recovery strategies up to a configurable cap, and +honest-exits when no improvement is doable. Use even when the user doesn't +explicitly say "autopilot" — any phrasing about running the full chain +end-to-end on one skill should trigger this. +``` + +## Acceptance criteria + +1. `skills/autopilot/SKILL.md` exists with the structure above. The + plain-language interview block is verbatim from this spec. +2. Plugin metadata lists the new skill across all five provider + manifests (`.claude-plugin/`, `.codex-plugin/`, + `.cursor-plugin/`, `.opencode/`, `gemini-extension.json`). +3. `skills/shared/workflow.md` row 10 description matches the + as-built behavior (cap names, exit-status enum, journal + location). +4. `skills/shared/subagent-dispatch.md` step 2↔3 classification + swap is fixed (existing bug surfaced during this brainstorm). +5. **End-to-end smoke test** on a local skill (e.g., one of this + repo's own chain skills, or a tiny purpose-built skill): + autopilot walks 1 → 9, writes the summary, exits cleanly. + Acceptance: a real `autopilot-summary-.md` exists, journal + is coherent, retries (if any) match the per-gate and global + caps. +6. **Surfaced-gate smoke test**: invoke autopilot with one gate + explicitly surfaced (e.g., G7-NW); verify the operator-prompt + fires and the run resumes after operator response. +7. **Cap exhaustion smoke test**: rig a probe deliberately too-easy + so G7-NW fires; verify autopilot tries strategies 1 and 2, then + honest-exits with `exit_status: unchanged-honest-exit`. + +## Out of scope (deferred) + +- **Batch processing.** Autopilot is single-skill. A batch + orchestrator that loops autopilot over a list of slugs is a + separate workstream. +- **Cost tracking.** Autopilot inherits the chain's cost + characteristics (subagents are free per Claude Code plan; + OpenRouter usage tracked by the run-bench CLI). No new budget + logic at the autopilot layer. +- **Cross-run learning.** Each autopilot run is independent. The + `default_overrides` recorded in a run's journal don't propagate + to subsequent runs. Possible future feature. +- **Surfacing partial bench results.** During a long bench run, + autopilot doesn't surface intermediate trial results to the + operator. The bench is treated atomically (start → finish → + evaluate G6-* gates). +- **Multi-skill dependencies.** If a skill being optimized depends + on another skill (chain of plugins), autopilot doesn't follow + those dependencies. One slug per run. + +## Open questions + +1. **Should `default_overrides` be persisted to a per-user config?** + Right now they're one-shot per invocation. A user who always + wants `pick_top_n: 10` re-supplies that every run. A persistent + per-user config (e.g., `~/.skill-optimizer/autopilot.yml`) could + eliminate the repetition but adds another piece of state. Defer + until real usage shows the friction. +2. **How should autopilot interact with the v1.3 + `tools/auto-improve-skill.mjs` wrapper?** The v1.3 wrapper still + exists in the repo on some branches. Once autopilot lands and is + validated, the v1.3 wrapper can be removed in a separate + cleanup PR. No coexistence needed during rollout — they're for + different chain generations. +3. **Strategy menu evolution.** The menus in this spec reflect the + patterns visible today. As operators use autopilot across more + skills, new strategies may emerge (e.g., "lower a probe's + acceptance threshold by N% before re-bench" as a cheap G6-ALLPASS + move). The SKILL.md is the source of truth for the menus; new + strategies get added by edit, no architectural change. + +## Side cleanup flagged by this brainstorm + +`skills/shared/subagent-dispatch.md:37` says fresh-derivation steps +are `(1, 3, 5, 6-summary, 7, 8, 9)` and `:66` says maintenance steps +are `(2, 4)`. Reading each chain skill's own iteration section +confirms: step 2 is fresh-derivation, step 3 is maintenance. The +two doc lines have the step numbers swapped. Fix during the +autopilot implementation, not a blocker. From c897c9407f0ba477b4b9cc90ee30b23b998c61d7 Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 07:43:25 -0500 Subject: [PATCH 117/121] feat(autopilot): step-10 chain driver SKILL + agent templates MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds `skills/autopilot/SKILL.md` as the 10th skill — the chain driver that walks steps 1→9 end-to-end on a single skill. Fully automated by default; operator picks optional per-gate breakpoints at startup. Honest-exits when caps exhaust (per-gate=2, global=8) or step 9 approves. Long static content lives in agents/ so the SKILL.md stays operational: interview-prompt.md (startup UX), recovery-menus.md (strategy menus consulted at each gate firing), surfacing-prompt.md (operator-facing gate prompt format), summary-template.md (end-of-run audit report shape). Plugin metadata + smoke tests updated. README/CONTRIBUTING mention the new driver. Fixed a pre-existing step 2↔3 classification swap in shared/subagent-dispatch.md that this work surfaced. Fulfills v1.4 acceptance criterion #8 (auto-pilot smoke test) — the last criterion from the v1.4 design that was never built. Co-Authored-By: Claude Opus 4.7 (1M context) --- .claude-plugin/marketplace.json | 3 +- CONTRIBUTING.md | 4 +- README.md | 6 +- skills/autopilot/SKILL.md | 285 ++++++++++++++++++++ skills/autopilot/agents/interview-prompt.md | 83 ++++++ skills/autopilot/agents/recovery-menus.md | 65 +++++ skills/autopilot/agents/summary-template.md | 98 +++++++ skills/autopilot/agents/surfacing-prompt.md | 73 +++++ skills/shared/subagent-dispatch.md | 4 +- skills/shared/workflow.md | 4 +- tests/smoke-skill-distribution.ts | 2 + 11 files changed, 619 insertions(+), 8 deletions(-) create mode 100644 skills/autopilot/SKILL.md create mode 100644 skills/autopilot/agents/interview-prompt.md create mode 100644 skills/autopilot/agents/recovery-menus.md create mode 100644 skills/autopilot/agents/summary-template.md create mode 100644 skills/autopilot/agents/surfacing-prompt.md diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index 8fec79d..e5da08f 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -22,7 +22,8 @@ "./skills/run-bench", "./skills/analyze", "./skills/improve", - "./skills/validate" + "./skills/validate", + "./skills/autopilot" ] } ] diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 1aba225..94a5e26 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,6 +1,6 @@ # Contributing to skill-optimizer -Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills//`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths. +Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills//`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill, plus an `autopilot` driver that walks the chain end-to-end. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths. ## Installing The Skill @@ -26,7 +26,7 @@ All three commands must pass before opening a PR when code changes are involved. - `src/workbench/acp/` — ACP transport, client wrapper, auth resolution, skill deployment, per-agent MCP config writer, trace recorder. - `src/workbench/agents/` — 5-agent registry + Dockerfile install snippet generator. - `docker/skill-optimizer-agent.Dockerfile` — container image with all 5 agent CLIs pre-baked. -- `skills//` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate). +- `skills//` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) plus the `autopilot` driver. - `skills/shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills. - `skills//agents/` — prompt templates dispatched by chain skills via the Agent tool. - `examples/workbench/` — packaged example suites. diff --git a/README.md b/README.md index 4c9cbaa..6fce926 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,12 @@ Docker workbench and Agent Skills for running deterministic evals against agent Use this repo in two ways: -- Install the `skill-optimizer` plugin into your agent so it can investigate a target skill, design and run an eval suite, analyze failures, and propose improvements. The plugin bundles a 9-step chain of Agent Skills under `skills//`. +- Install the `skill-optimizer` plugin into your agent so it can investigate a target skill, design and run an eval suite, analyze failures, and propose improvements. The plugin bundles a 9-step chain of Agent Skills under `skills//`, plus an `autopilot` driver that runs the chain end-to-end. - Run the local CLI to execute cases and suites in Docker against OpenRouter models. ## Installation -Installation differs by agent. Every plugin manifest exposes all 9 chain skills from `skills//`; once installed, each skill triggers on its own description (e.g. "investigate this skill", "design tests for this skill", "run the bench", "analyze the results"). +Installation differs by agent. Every plugin manifest exposes all 9 chain skills from `skills//` plus the `autopilot` driver; once installed, each skill triggers on its own description (e.g. "investigate this skill", "design tests for this skill", "run the bench", "analyze the results", "autopilot this skill"). ### Claude Code @@ -66,6 +66,7 @@ npx skills add fastxyz/skill-optimizer \ --skill analyze \ --skill improve \ --skill validate \ + --skill autopilot \ -a cursor -y ``` @@ -118,6 +119,7 @@ npx skills add fastxyz/skill-optimizer \ --skill analyze \ --skill improve \ --skill validate \ + --skill autopilot \ -a claude-code -a opencode -a codex -a cursor -y ``` diff --git a/skills/autopilot/SKILL.md b/skills/autopilot/SKILL.md new file mode 100644 index 0000000..19473c6 --- /dev/null +++ b/skills/autopilot/SKILL.md @@ -0,0 +1,285 @@ +--- +name: autopilot +description: Use when the user wants to run the entire skill-optimizer chain end-to-end on a single skill without driving each step manually — phrases like "autopilot this skill", "run the full chain on X", "skill-optimizer end-to-end for X", "automated improvement run for X". Walks steps 1 through 9 with the chain's iteration mechanism, applies per-gate recovery strategies up to a configurable cap, and honest-exits when no improvement is doable. Use even when the user doesn't explicitly say "autopilot" — any phrasing about running the full chain end-to-end on one skill should trigger this. +--- + +# autopilot + +Step 10 of the skill-optimizer chain. **Chain driver, not a step.** +Takes a skill (URL or local path), walks the 9-step chain +sequentially via the `Skill` tool, handles each step's gate firings +with a strategy-menu, and writes a per-run summary at +`docs/skill-optimizer//autopilot-summary-.md`. + +Fully automated by default. Operator picks optional breakpoints at +startup; everything else auto-recovers up to a per-gate cap. The +only always-on breakpoint is bench infrastructure failure (docker +or agent crash) — autopilot doesn't recover from that. + +## What you produce + +Two outputs per run: + +1. **`.skill-optimizer//autopilot-/journal.md`** — the + per-run journal. Appended to throughout the run; this IS your + working memory across the chain. Gitignored. +2. **`docs/skill-optimizer//autopilot-summary-.md`** — + the per-run audit summary. Promoted from the journal at + end-of-run. Tracked. + +Multiple runs accumulate naturally with their own ``. Caps +reset per run; the new invocation gets fresh +`max_retries_per_gate`. + +## Workflow + +### (a) Confirm prerequisites and classify source + +The user must have provided a skill source (URL or local path). If +they invoked autopilot without a source, ask which skill to +optimize. + +Classify the source the same way step 1 does: + +- **Upstream:** URL or `//`. Triggers the + PR-intent question in the interview below. +- **Local:** filesystem path to a SKILL.md (or a directory + containing one). Skips the PR question. + +If the user is on `main`, `master`, `development`, or a similar +long-lived branch, recommend creating a dedicated branch (e.g., +`git checkout -b eval/`) before continuing. Same branch +discipline as step 1. + +Create a TodoWrite list with all 10 chain steps so progress is +visible across the long run. Mark step 10 (this skill) as +`in_progress`; subsequent chain skills update their own entry as +they fire. + +### (b) Run the startup interview + +Load +[`agents/interview-prompt.md`](./agents/interview-prompt.md), +substitute `${SKILL_NAME}`, and present it to the user as a single +message. Wait for the response. + +Capture the response into the run journal frontmatter (see (c) for +the shape). Defaults if the user said "all auto, go": + +| Field | Default | Source | +|---|---|---| +| `pr_intent` | `false` | upstream source only | +| `pick_top_n` | `5` | always | +| `max_retries_per_gate` | `2` | always | +| `global_retry_cap` | `8` | always | +| `surfaced_gates` | `{G6-CLI-FAIL}` | operator may add more | +| `default_overrides` | `{}` | operator may override any gate default | + +The response-parsing rules (including the user-phrasing → gate-ID +mapping) live in `agents/interview-prompt.md`. + +### (c) Initialize the run journal + +```bash +TS=$(date -u +%Y%m%dT%H%M%SZ) +mkdir -p ".skill-optimizer//autopilot-${TS}" +``` + +Write `journal.md` with YAML frontmatter holding the `mode_settings` +fields from (b), plus three empty H2 sections to be appended to as +the run proceeds: `## Run journal` (event log), `## Retry counters` +(per-gate `N/max` lines), `## Global retry counter`. The journal is +your working memory across the run — append every dispatch, every +gate firing, and every retry decision. Read it before each gate +decision to know what's been tried. + +### (d) Forward walk through the chain (steps 1 → 9) + +For each step in order: + +1. **Staleness check.** Read the canonical artifact's git mtime + against direct upstream (per + [`iteration-protocol.md`](../shared/iteration-protocol.md)). If + current and no operator directive forces a re-run, skip and log + `step N skipped (current)` in the journal. +2. **Dispatch.** Call the chain skill via the `Skill` tool: + + ```text + Skill skill-optimizer:investigate-functionality + Skill skill-optimizer:investigate-submissions # only if pr_intent + Skill skill-optimizer:design-tests + Skill skill-optimizer:write-tests + Skill skill-optimizer:validate-tests + Skill skill-optimizer:run-bench + Skill skill-optimizer:analyze + Skill skill-optimizer:improve + Skill skill-optimizer:validate + ``` + + Each chain skill does its own work (subagent dispatch, file + writes); you wait for it to complete. +3. **Read the canonical.** After the chain skill returns, read its + canonical output (frontmatter is the load-bearing signal — see + gate triggers below). +4. **Check for gate firings.** Match the canonical against the gate + triggers in [Gate menus](#gate-menus). If any gate fires, go to + workflow step (e). Otherwise, continue to the next chain step. + +Step 2 (`investigate-submissions`) runs only if +`pr_intent: true` (set in the interview or carried from a prior +step 1 canonical). Skip otherwise. + +### (e) Handle gate firings + +Follow the [Reasoning protocol at each gate firing](#reasoning-protocol-at-each-gate-firing). + +The protocol resolves to one of: + +- **Execute a recovery strategy.** Re-dispatch one or more chain + skills with a distilled `${OPERATOR_DIRECTIVES}` derived from the + gate's rationale. Then resume the forward walk from the earliest + re-dispatched step. +- **Skip and continue.** For `No retry` gates (G4-BLOCKED, + G5-REJECT) — drop the affected probe(s), log, and resume the + forward walk where you left off. +- **Honest-exit.** Cap exhausted (per-gate or global) — go to + workflow step (f). +- **Surface.** Gate is in `surfaced_gates` — prompt the operator + per [Surfacing a gate](#surfacing-a-gate), then follow their + answer. + +Every gate firing gets one journal entry. Every retry decision +increments the gate's counter and the global counter. + +### (f) Honest-exit and write the summary + +Set `exit_status` per the run's final state: + +- **`improved`** — step 9 verdict was `approve`; + `docs/skill-optimizer//improved-skill/` materialized. +- **`unchanged-honest-exit`** — cap exhausted at any gate, OR + step 7 reported no structural weakness with all G7-NW strategies + exhausted, OR step 9 rejected with all G9-REJECT strategies + exhausted. +- **`blocked-on-infra`** — G6-CLI-FAIL surfaced and the operator + did not recover the environment. + +Promote the journal into +`docs/skill-optimizer//autopilot-summary-.md` using +[`agents/summary-template.md`](./agents/summary-template.md). The +summary is an aggregate; trim narrative, keep retry trail + exit +reason + repo changes. + +Update the chain TodoWrite list to mark step 10 `completed` and +print the path to the summary report. + +## Gate menus + +A **gate** is a decision point in the chain where autopilot might +need to take action beyond just dispatching the next step. Gates +are detected from the chain skill's canonical frontmatter and body. + +### Gate definitions + +| ID | Fires when (frontmatter / body signal) | Retry policy | +|---|---|---| +| G3 | Step 3's `03-test-proposals.md` lists N functionalities; picks need to be set | Single action — mark top-`pick_top_n` by importance; not retry-eligible | +| G4-BLOCKED | Step 4 test-writer subagent returned `status: blocked` on a probe | No retry — skip, log, continue | +| G5-REVISION | `05-tests-verdict.md` has `needs-revision` on ≥1 probe | Menu (1 option) | +| G5-REJECT | `05-tests-verdict.md` has `reject` on ≥1 probe | No retry — drop those probes, continue | +| G6-ALLPASS | `06-bench-summary.md` shows `overall_pass_rate: 1.0` | Not a halt gate; continue to step 7 (G7-NW handles real recovery) | +| G6-CLI-FAIL | Step 6 CLI exited non-zero or no `suite-result.json` | No retry — always surfaces (infra) | +| G7-NW | `07-analysis.md` has `has_structural_weakness: false` | Menu (4 options) | +| G8-BLOCKED | Step 8 optimizer subagent returned `status: blocked` | Menu (2 options) | +| G9-REVISION | `09-validator-verdict.md` has `verdict: needs-revision` | Menu (1 option) | +| G9-REJECT | `09-validator-verdict.md` has `verdict: reject` | Menu (3 options) | + +### Recovery menus + +Each retry-eligible gate has a menu of strategies, ordered +cheap → expensive. Full menu reference (loaded at each gate +firing): [`agents/recovery-menus.md`](./agents/recovery-menus.md). + +Read it before deciding, and pick the cheapest **available + +unattempted** option each time the gate fires this run. + +## Reasoning protocol at each gate firing + +When a gate fires: + +1. Read the gate's menu in the section above. +2. Read the journal's retry counters and the appended event log + to see which strategies have already been attempted this run + for this gate. +3. If the gate ID is in `surfaced_gates` (from interview): surface + per [Surfacing a gate](#surfacing-a-gate). Wait for the + operator's response and follow it. +4. Else: pick the cheapest unattempted strategy from the menu. +5. Append a journal entry. Format: + + ```text + strategy () | + reason: | + retry: / + ``` + +6. Distill the gate's rationale into `${OPERATOR_DIRECTIVES}` — an + atomic list of new requirements for the chain skill being + re-dispatched. Never a context dump. See + [`subagent-dispatch.md`](../shared/subagent-dispatch.md) for + the directive contract. +7. Re-dispatch the relevant chain step(s) via the `Skill` tool + with that directive in scope. +8. After the re-dispatch returns, resume the forward walk. If the + same gate fires again, repeat from step 1 of this protocol. + +## Cap accounting + +- **Per-gate cap** (`max_retries_per_gate`, default 2). Counts + total retry events for that gate type, regardless of which menu + strategy was picked. Two G7-NW retries means strategies 1 and 2 + have been attempted; strategies 3 and 4 are unreachable this + run. +- **Global cap** (`global_retry_cap`, default 8). Counts retry + events across all gates. Safety net for thrashing. +- A retry event is one re-dispatch of a chain step driven by a + gate's recovery menu. Mechanical re-dispatches (e.g., step 4 + re-running per-probe inside step 5's revision loop) count once + at the gate that triggered them, not per chain skill. +- **Cap exhaustion is honest-exit**, not failure. The journal + records what was tried; the summary explains the principled + outcome. + +## Surfacing a gate + +When a gate ID is in `surfaced_gates` (or the meta-breakpoint for +multi-strategy menus is enabled), pause the run and present the +template at +[`agents/surfacing-prompt.md`](./agents/surfacing-prompt.md) with +the gate's context substituted in. Wait for the operator's +response; the parsing rules (strategy number, custom directive, +"honest exit") live in that template. + +## Edge cases + +- **Re-entrancy on the same slug.** Re-invoking autopilot on the + same slug picks up where it left off via the staleness check + in (d). The new invocation gets its own + `autopilot-/journal.md` and its own summary. Caps reset. + Prior run's outputs are preserved naturally by timestamp. +- **Step 1 wrapper detection.** If step 1 flags + `likely_wrapper: true`, the chain skill itself surfaces the + decision to the operator. Autopilot does not auto-resolve + wrapper redirection — treat this as a structural surface, not a + gate. The operator's response determines whether to re-vendor + and continue. +- **No suite.yml yet on a fresh slug.** First-time invocation has + no canonicals; staleness check treats everything as stale and + dispatches step 1 normally. Same for downstream steps. +- **G6-CLI-FAIL during a retry.** Infra failure during a + recovery step is still G6-CLI-FAIL. Surface, do not retry, + `exit_status: blocked-on-infra` if operator does not recover. +- **All-pass after an enable-more-picks attempt.** If G7-NW + strategy 3 succeeds (picks are added, bench runs, weakness is + found), the forward walk resumes from the new step 7. No + special handling. diff --git a/skills/autopilot/agents/interview-prompt.md b/skills/autopilot/agents/interview-prompt.md new file mode 100644 index 0000000..715fab8 --- /dev/null +++ b/skills/autopilot/agents/interview-prompt.md @@ -0,0 +1,83 @@ +# Startup interview prompt + +Operator-facing text presented by `skill-optimizer:autopilot` at +workflow step (b). Render with the inputs below substituted, then +send as a single message to the user. Wait for the user's response +before continuing. + +## Inputs (templated by the operator session) + +- `${SKILL_NAME}` — display name of the skill being optimized + (the slug or the user's original phrasing) + +## Prompt body + +```text +I'm going to optimize ${SKILL_NAME} end-to-end. By default I'm +fully automated — I run all 9 chain steps and retry up to 2 times +at each recovery point. I'll only pause and ask if the bench +environment breaks (docker/agent crash — that's not something to +fix without you). + +Optional breakpoints — pick any to be in the loop on: + +- Reviewing picked test cases before bench runs. Default: top 5 by + importance are picked automatically. +- A probe gets blocked or rejected during test-writing. Default: + drop the affected probe and continue; the others still run. +- Validate-tests says one or more probes need revision. Default: + distill the rationale into a directive and retry the probe up to + the cap. +- The skill passes every test or no structural weakness gets found. + Default: try recovery strategies in order — first reframe the + analysis, then more bench trials, then enable more picks, then + propose new probes — until something surfaces or the cap is hit. +- The optimizer gets stuck on a weakness it can't translate to a + concrete fix. Default: reframe the weakness and retry up to the + cap. +- The validator says the proposed fix needs revision. Default: + distill the rationale and retry the proposal up to the cap. +- The validator outright rejects the proposed fix. Default: try + recovery strategies in order — first a different fix for the same + weakness, then reframe the weakness, then add more probe coverage + — until something lands or the cap is hit. +- A recovery point has multiple strategies available (meta-breakpoint + covering the three above). Default: I pick the cheapest unattempted + strategy. Surfacing this means I'll show you the menu and let you + choose. + +A few defaults you can also override: + +- PR intent: no (only matters for upstream skills) +- Top-N test cases picked: 5 +- Max retries per recovery point: 2 + +Say 'all auto, go' to just run, or tell me what to change. +``` + +## Parsing the response + +The operator's response may be: + +- **"all auto, go"** (or equivalent) — proceed with all defaults. +- **A list of breakpoints to surface** — add the listed gate IDs + to `surfaced_gates`. The mapping from the user-facing phrasing + to gate IDs: + + | User phrasing | Gate ID(s) | + |---|---| + | reviewing picked test cases | G3 | + | probe blocked / rejected during test-writing | G4-BLOCKED, G5-REJECT | + | validate-tests probe needs-revision | G5-REVISION | + | skill passes every test / no structural weakness | G6-ALLPASS, G7-NW | + | optimizer stuck | G8-BLOCKED | + | validator needs-revision | G9-REVISION | + | validator rejects | G9-REJECT | + | meta-breakpoint (multi-strategy menus) | mark G7-NW, G8-BLOCKED, G9-REJECT as surfaced | + +- **Override values** for `pr_intent`, `pick_top_n`, or + `max_retries_per_gate` — capture into the matching field. +- **Per-gate default overrides** — capture into + `default_overrides` as `{: }`. +- **Ambiguous response** — ask one clarifying question; do not + guess. diff --git a/skills/autopilot/agents/recovery-menus.md b/skills/autopilot/agents/recovery-menus.md new file mode 100644 index 0000000..debcb84 --- /dev/null +++ b/skills/autopilot/agents/recovery-menus.md @@ -0,0 +1,65 @@ +# Recovery menus + +Reference loaded by `skill-optimizer:autopilot` at each gate firing +(per the Reasoning protocol in SKILL.md). For each retry-eligible +gate, lists strategies ordered cheap → expensive. The pilot picks +the cheapest **available + unattempted** option, consulting the +journal's retry log for what's been tried this run. + +## G5-REVISION + +1. Distill the validator's per-probe rationale into directives. + Re-run step 4 for the affected probes only. Re-run step 5. + +## G7-NW (cheap → expensive) + +1. **Reframe analysis angle.** Re-run step 7 with a directive + pointing at marginal failures, model-specific drift, latent + issues even when pass-rate is high. +2. **More bench trials.** Re-run step 6 with an elevated trial + count (e.g., 3× the normal trials per probe) to surface + variance, then re-run step 7. +3. **Enable more picks.** Flip the next-highest-importance + `picked: false` functionalities in step 3's spec to + `picked: true`. Re-run steps 4 → 5 → 6 → 7 for the new probes + only (existing probes stay). +4. **Add new probes.** Re-run step 3 with a directive to propose + new functionalities the existing pass missed (focus on edge + cases the existing probes don't cover). Then 4 → 5 → 6 → 7. + +## G8-BLOCKED (cheap → expensive) + +1. **Reframe weakness.** Re-run step 7 with a directive that the + optimizer was blocked on weakness W; ask for a reformulation + that is more concrete or at a different abstraction. Then + re-run step 8. +2. **Sharper directive to optimizer.** If the weakness reads + concrete enough but the optimizer struggled with the general + principle, re-run step 8 alone with a sharpened directive + tying the principle to the skill's structure. + +## G9-REVISION + +1. Distill the validator's rationale into a directive. Re-run + step 8. Re-run step 9. + +## G9-REJECT (cheap → expensive) + +1. **Different fix, same weakness.** Re-run step 8 with the + rejected proposal noted as anti-pattern (proposal P rejected + because Y; try a different general principle for the same + weakness). +2. **Reframe weakness.** Re-run step 7 with the validator's + rejection rationale (weakness X rejected because Y; find a + different angle on what's failing). Then 8 → 9. +3. **Add coverage.** If the validator rejected on grounds that + the weakness wasn't real in the trace data, treat as G7-NW + menu (more trials, more picks, more probes). + +## Gates that don't appear here + +- **G3** (picks) — single action, not retry-eligible. +- **G4-BLOCKED, G5-REJECT** — no retry: skip/drop, log, continue. +- **G6-ALLPASS** — not a halt gate; continue to step 7 (G7-NW + handles the actual recovery). +- **G6-CLI-FAIL** — no retry: always surfaces; infra needs human. diff --git a/skills/autopilot/agents/summary-template.md b/skills/autopilot/agents/summary-template.md new file mode 100644 index 0000000..765f3b9 --- /dev/null +++ b/skills/autopilot/agents/summary-template.md @@ -0,0 +1,98 @@ +# End-of-run summary template + +Operator-facing audit record written by `skill-optimizer:autopilot` +at workflow step (f), promoted from the run journal. Tracked at +`docs/skill-optimizer//autopilot-summary-.md`. + +## Inputs (templated by the operator session) + +- `${RUN_STARTED_AT}` / `${RUN_ENDED_AT}` — ISO 8601 UTC +- `${SLUG}` — chain slug +- `${SOURCE}` — URL or local path the operator provided +- `${EXIT_STATUS}` — one of `improved`, `unchanged-honest-exit`, + `blocked-on-infra` +- `${TOTAL_RETRIES}` — integer total across all gates +- `${JOURNAL_PATH}` — gitignored journal path +- `${HEADLINE}` — one paragraph: what happened end-to-end and why + it ended this way (operator writes from journal) +- `${PER_STEP_ROWS}` — table rows for the per-step outcomes +- `${RECOVERY_SUMMARY}` — bullet list of fired gates with retry + trail; bullet list of gates not fired +- `${CAVEATS}` — operator's interpretation of anything off (flaky + bench, suspicious trace patterns); not a substitute for operator + review +- `${REPO_CHANGES}` — file list of new/modified canonicals + +## Template body + +```markdown +--- +run_started_at: ${RUN_STARTED_AT} +run_ended_at: ${RUN_ENDED_AT} +slug: ${SLUG} +source: ${SOURCE} +exit_status: ${EXIT_STATUS} +total_retries: ${TOTAL_RETRIES} +journal_path: ${JOURNAL_PATH} +--- + +# Autopilot run for `${SLUG}` + +## Headline + +${HEADLINE} + +## Per-step outcomes + +| Step | Outcome | Artifact | +|---|---|---| +${PER_STEP_ROWS} + +## Recovery summary + +${RECOVERY_SUMMARY} + +## Caveats + +${CAVEATS} + +## What changed in the repo + +${REPO_CHANGES} +``` + +## Per-step outcome row format + +```markdown +| | | | +``` + +Examples: + +- `| 1 investigate-functionality | written | docs/.../01-functionality.md |` +- `| 7 analyze | written (re-run 2x) | docs/.../07-analysis.md |` +- `| 8 improve | skipped (no weakness) | — |` + +## Recovery summary format + +For each gate that fired: + +```markdown +- retries: / + - strategy () — + - strategy () — +``` + +Then one line listing gates that did not fire: + +```markdown +- Gates not fired: +``` + +## Discipline + +- The summary is an aggregate, not a journal dump. Trim narrative; + keep retry trail + exit reason + repo changes. +- `exit_status` enum is the contract — do not invent new values. +- Do NOT embed trace excerpts. Pointers to canonical files are + enough; the operator opens those for detail. diff --git a/skills/autopilot/agents/surfacing-prompt.md b/skills/autopilot/agents/surfacing-prompt.md new file mode 100644 index 0000000..643da52 --- /dev/null +++ b/skills/autopilot/agents/surfacing-prompt.md @@ -0,0 +1,73 @@ +# Gate-surfacing prompt + +Operator-facing text presented by `skill-optimizer:autopilot` when +a gate in `surfaced_gates` fires (or when the meta-breakpoint for +multi-strategy menus is enabled). Render with the inputs below +substituted, then send as a single message to the user. Wait for +the user's response before continuing. + +## Inputs (templated by the operator session) + +- `${GATE_ID}` — e.g. `G7-NW` +- `${STEP_NUMBER}` — chain step that just produced the canonical + triggering the gate +- `${WHAT_HAPPENED}` — plain-language description of the canonical + finding (one or two sentences; no jargon if avoidable) +- `${MENU_OPTIONS}` — numbered list of strategies from this gate's + recovery menu, with one-sentence descriptions and an attempted + marker (yes/no) per option +- `${CURRENT_RETRY}` / `${MAX_RETRIES}` — current and max per-gate + retry counter +- `${GLOBAL_CURRENT}` / `${GLOBAL_CAP}` — current and max global + retry counter + +## Prompt body + +```text +Gate ${GATE_ID} just fired at step ${STEP_NUMBER}. + +What happened: +${WHAT_HAPPENED} + +Recovery options (cheap → expensive): +${MENU_OPTIONS} + +You can: +- Pick a strategy number to try it +- Supply your own directive ("look at the X cluster", "treat Y as a + marginal failure") +- Say "honest exit" to stop here + +Currently ${CURRENT_RETRY}/${MAX_RETRIES} retries used on this +gate; global ${GLOBAL_CURRENT}/${GLOBAL_CAP}. +``` + +## Menu option format + +Each line of `${MENU_OPTIONS}`: + +```text +. — [attempted: yes/no] +``` + +Example: + +```text +1. Reframe analysis angle — re-run step 7 with a directive pointing at + marginal failures and model-specific drift [attempted: yes] +2. More bench trials — re-run step 6 with 3x trials, then step 7 + [attempted: no] +``` + +## Parsing the response + +- **Strategy number** — execute that option (as if autopilot had + picked it itself). +- **Custom directive** — treat as a one-off operator directive; + pick the most-fitting menu option (based on which strategy the + directive aligns with) and prepend the operator's text to the + distilled `${OPERATOR_DIRECTIVES}` for that dispatch. +- **"Honest exit"** — go straight to workflow step (f), + `exit_status: unchanged-honest-exit`. +- **Ambiguous response** — ask one clarifying question; do not + guess which strategy to pick. diff --git a/skills/shared/subagent-dispatch.md b/skills/shared/subagent-dispatch.md index e915172..0cce4ea 100644 --- a/skills/shared/subagent-dispatch.md +++ b/skills/shared/subagent-dispatch.md @@ -34,7 +34,7 @@ Three rules every reasoning subagent must follow: filesystem is the source of truth; do not walk `git log` looking for prior versions of upstream files. -2. **For fresh-derivation steps (1, 3, 5, 6-summary, 7, 8, 9): do NOT +2. **For fresh-derivation steps (1, 2, 5, 6-summary, 7, 8, 9): do NOT read your own canonical file, and do NOT read git history of it.** Each invocation derives a new artifact from upstream + the operator's directives, without direct access to prior @@ -63,7 +63,7 @@ Three rules every reasoning subagent must follow: `09-validator-verdict.md` for the validator), not the skill content itself. -3. **For maintenance steps (2, 4): DO read your own canonical +3. **For maintenance steps (3, 4): DO read your own canonical tree** (when it exists). Your job on a re-run is to extend or modify the current state per `${OPERATOR_DIRECTIVES}`, preserving entries the user has invested in unless a directive diff --git a/skills/shared/workflow.md b/skills/shared/workflow.md index 13e95a4..58539f0 100644 --- a/skills/shared/workflow.md +++ b/skills/shared/workflow.md @@ -18,7 +18,7 @@ re-runs, and the relationship between steps. | 7 | `analyze` | fresh-derivation | step 6 + skill | `docs/.../07-analysis.md` | | 8 | `improve` | fresh-derivation | step 7 + skill (+ step 2 if PR-bound) | `docs/.../08-improvement-proposal.md` | | 9 | `validate` | fresh-derivation | step 8 + skill (+ step 2 if PR-bound) | `docs/.../09-validator-verdict.md`, `docs/.../improved-skill/` (on approve) | -| 10 | `autopilot` | chain driver | same as step 1 + flags | `docs/.../autopilot-summary-.md` | +| 10 | `autopilot` | chain driver | source skill + startup-interview flags | `.skill-optimizer/.../autopilot-/journal.md`, `docs/.../autopilot-summary-.md` | (Paths abbreviated; full state layout below.) @@ -107,6 +107,8 @@ skill-evals// # TRACKED — eval suites + probes (reusable a .skill-optimizer// # GITIGNORED — heavy ephemeral artifacts vendored-skill/ # step 1 vendors (upstream OR local) bench-results// # step 6 raw output (suite-result.json, trace.jsonl, findings.txt) + autopilot-/ # step 10 per-run scratch + journal.md # appended-to during the run ``` `` is `--` for upstream skills, or diff --git a/tests/smoke-skill-distribution.ts b/tests/smoke-skill-distribution.ts index b3e97f0..659b574 100644 --- a/tests/smoke-skill-distribution.ts +++ b/tests/smoke-skill-distribution.ts @@ -26,6 +26,7 @@ test('chain skills follow the portable agent skills contract', () => { 'analyze', 'improve', 'validate', + 'autopilot', ]; for (const dir of chainSkillDirs) { @@ -135,6 +136,7 @@ test('Claude plugin and marketplace metadata expose all v1.4 chain skills', () = './skills/analyze', './skills/improve', './skills/validate', + './skills/autopilot', ]); for (const skillPath of marketplace.plugins[0].skills) { const resolved = skillPath.replace(/^\.\//, ''); From f41710f24ddc6f9026d6a4f58613950835dbdf0d Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 08:15:47 -0500 Subject: [PATCH 118/121] feat: workbench reference doc refresh + tmp cleanup + wrapper/branch UX MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit skills/shared/workbench.md is significantly refreshed to match current truth: drops stale --model/--models flags, replaces models:[...] with runs:[{agent,model}], adds agent: as required on cases, fixes the default image name. Adds a new Per-Agent Model Strings section listing the model-naming convention each agent CLI accepts plus the env vars they require — closes the trace-friction where Claude wrote 'claude-opus-4-7-1m' and the CLI rejected it (valid: opus[1m]). src/workbench/docker-runner.ts: two changes. - copyCaseSupportDirs() now enumerates the case dir dynamically with an exclusion set (references/, node_modules/, hidden dirs), instead of a hardcoded ['checks','fixtures','bin','workspace','mcp'] list. Case authors can add seeds/, data/, helpers/ etc. without editing the runner. - Pre-removal docker exec --user 0 chmod -R 0777 /home/agent /work before docker rm -f, so the host's rmSync(tempDir) doesn't trip EACCES on uid-10001-owned subdirs (.cache, .venv). Best-effort; ignored if the container is already gone. skills/investigate-functionality/SKILL.md: two UX changes. - Wrapper-redirect now keeps the wrapper intact at vendored-skill/ and writes the underlying content to a sibling vendored-underlying/. Tests target vendored-skill/ (the ship target) regardless of re-vendor outcome. Surfaces clearer 3-option prompt. - Branch check is now ALWAYS asked, with current branch surfaced and eval/ recommended as option 1. Replaces the prior 'only ask if on main/master/development' heuristic. skills/shared/workflow.md: state layout adds vendored-underlying/ and notes that test target is vendored-skill/. Addresses 4 of 5 friction points from a manual trial run trace: workbench model-string confusion, wrapper-handoff prose, branching timing ambiguity, tmp /tmp permission noise. Co-Authored-By: Claude Opus 4.7 (1M context) --- skills/investigate-functionality/SKILL.md | 87 ++++++++++++++++------- skills/shared/workbench.md | 73 ++++++++++++++----- skills/shared/workflow.md | 10 ++- src/workbench/docker-runner.ts | 24 ++++++- 4 files changed, 147 insertions(+), 47 deletions(-) diff --git a/skills/investigate-functionality/SKILL.md b/skills/investigate-functionality/SKILL.md index afcd2bf..16446da 100644 --- a/skills/investigate-functionality/SKILL.md +++ b/skills/investigate-functionality/SKILL.md @@ -76,12 +76,31 @@ bench output, gitignored). Keeping the tracked locations on a feature branch makes the run easy to discard, iterate on, or merge in one piece. -If the user is on `main`, `master`, `development`, or a similar -long-lived branch, recommend creating a dedicated branch (e.g., -`git checkout -b eval/`) before continuing. If they decline, -proceed but note that chain outputs will accumulate on the current -branch. Skip this check entirely if they're already on a -purpose-named feature branch. +**Always ask the user explicitly** — don't infer from branch +name. After running `git rev-parse --abbrev-ref HEAD`, surface +this prompt verbatim: + +> I'm about to start an optimization run for ``. The chain +> will create tracked files under `docs/skill-optimizer//` +> and `skill-evals//`. **I recommend creating a dedicated +> branch** (e.g., `git checkout -b eval/`) so the run is +> easy to discard, iterate on, or merge in one piece. +> +> You're currently on ``. Options: +> +> 1. **Create `eval/` and switch to it now.** (Recommended.) +> 2. **Use the current branch.** Chain outputs will accumulate +> on `` — that's fine if it's already a +> purpose-named feature branch. +> 3. **Cancel and let me sort out branches first.** + +Default to recommending option 1 unless the current branch is +already a clearly purpose-named feature branch matching this +slug (e.g., `eval/`, `feat/`). If the user picks +option 1, run `git checkout -b eval/` for them before +proceeding. If they pick option 2, proceed on the current +branch. If option 3, exit so they can create the branch they +want and re-invoke. **Chain todos.** Create a TodoWrite list with the chain's 9 (or 10, if autopilot will run) steps so progress is visible across a @@ -212,28 +231,46 @@ content elsewhere), surface to the user before handing off: > reference the actual content at ``. Three > realistic responses: > -> 1. **Re-vendor the referenced content and re-research.** The -> chain will then target the actual content for testing, -> analysis, and optimization. (Common — usually what you -> want.) -> 2. **Proceed treating this wrapper as the skill.** Useful if -> the wrapper itself is what you want to improve (rare). +> 1. **Re-vendor the referenced content alongside the wrapper.** +> Investigation re-runs against the underlying content so the +> `01-functionality.md` reflects the actual rules; tests will +> still target the wrapper (the wrapper is what ships, so +> that's what should be exercised). The optimizer may modify +> files in either folder — both stay available. (Common — +> usually what you want.) +> 2. **Proceed treating this wrapper as the skill.** Don't fetch +> the underlying content; everything downstream uses the +> wrapper as-is. Useful if you only want to improve the wrapper +> itself. > 3. **Cancel and provide a different source.** If the subagent's > judgment is wrong or you meant to point elsewhere. -If the user picks (1): delete -`.skill-optimizer//vendored-skill/`, vendor the content at -`wrapper_points_to` into it, update `${SKILL_SOURCE}` to that -URL/path, and re-invoke this skill from (d). The new step 1 run -rebuilds `01-functionality.md` from the underlying content — -`skill_source` reflects the new location and `likely_wrapper` -will typically be `false` (unless the underlying is itself -another wrapper, in which case repeat the gate). After this -re-run, downstream steps see a clean non-wrapper source. - -If (2): keep the report as-is and proceed. Downstream steps see -`likely_wrapper: true` with `wrapper_points_to` documented; the -PR target is the wrapper itself (the user's deliberate choice). +If the user picks (1): + +1. **Keep the wrapper at `.skill-optimizer//vendored-skill/` + intact.** That's the ship target — tests install this folder; + from the user's perspective it IS the skill. +2. Vendor the content at `wrapper_points_to` into a sibling + directory: `.skill-optimizer//vendored-underlying/`. +3. Re-invoke this skill from (d) with the **underlying** content + as the research target. `${SKILL_SOURCE}` in this re-run points + at `wrapper_points_to` so the researcher builds an accurate + `01-functionality.md`, but `vendored-skill/` is not touched. +4. After the re-run, the report frontmatter still records the + user's original `skill_source` (the wrapper URL/path) and + `wrapper_points_to`. Downstream steps know: test against + `vendored-skill/` (the wrapper); the optimizer/validator may + read `vendored-underlying/` for context and may patch files in + either folder depending on where the actual rule content lives. +5. If the underlying content is itself a wrapper, the re-run will + flag `likely_wrapper: true` again and the gate repeats. Recurse + into a deeper `vendored-underlying/` chain only if you want to + walk it. + +If (2): keep the report as-is and proceed. No `vendored-underlying/` +is created. Downstream steps see `likely_wrapper: true` with +`wrapper_points_to` documented; the PR target is the wrapper +itself (the user's deliberate choice). If (3): exit; the user will re-invoke with a new source. diff --git a/skills/shared/workbench.md b/skills/shared/workbench.md index e0c5a0f..0e42745 100644 --- a/skills/shared/workbench.md +++ b/skills/shared/workbench.md @@ -18,8 +18,7 @@ Avoid evals that require running model-produced arbitrary production code outsid ```bash npx tsx src/cli.ts run-case -npx tsx src/cli.ts run-case --model openrouter/google/gemini-2.5-flash -npx tsx src/cli.ts run-case --models openrouter/google/gemini-2.5-flash,openrouter/openai/gpt-5.4 --trials 3 --concurrency 2 +npx tsx src/cli.ts run-case --trials 3 --concurrency 2 npx tsx src/cli.ts run-suite --trials 3 --concurrency 2 ``` @@ -28,16 +27,20 @@ Options: | Command | Option | Meaning | |---------|--------|---------| | `run-case` | `--out ` | Results root, default `/.results` | -| `run-case` | `--model ` | Single OpenRouter model override | -| `run-case` | `--models ` | Comma-separated OpenRouter model refs | -| `run-case` | `--trials ` | Independent trials per model | +| `run-case` | `--trials ` | Independent trials per case | | `run-suite` | `--out ` | Results root, default `/.results` | -| `run-suite` | `--trials ` | Independent trials per case/model | +| `run-suite` | `--trials ` | Independent trials per case × run | | both | `--concurrency ` | Maximum concurrent trial containers | -| both | `--image ` | Docker image, default `skill-optimizer-workbench:local` | +| both | `--image ` | Docker image, default `skill-optimizer-agent:local` | | both | `--keep-workspace` | Preserve successful workspaces too; failures are always preserved | -Only `openrouter/...` model refs are accepted. `run-suite` uses the `models:` array in the suite file. +There is **no `--model` / `--models` override**. The case file's `model:` (and the suite's `runs: [{ agent, model }]` matrix) is the only source of truth — this is intentional, so a suite's matrix is reproducible from its file. + +The image must contain the agent CLIs. Build it once with: + +```bash +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . +``` ## Case Schema @@ -45,6 +48,8 @@ Case files may be `.yml`, `.yaml`, or `.json`. ```yaml name: extract-pdf-facts +agent: pi-acp +model: openrouter/google/gemini-2.5-flash references: ./references task: | Read statement.pdf and write answer.json with the account, quarter, approval code, and risk flags. @@ -64,7 +69,6 @@ mcpServices: command: node args: - calculator-server.mjs -model: openrouter/google/gemini-2.5-flash timeoutSeconds: 600 ``` @@ -73,35 +77,40 @@ Required fields: | Field | Type | Meaning | |-------|------|---------| | `name` | string | Human-readable case name; suite inline cases slug this for result dirs | +| `agent` | string | One of `claude-agent-acp`, `codex-acp`, `gemini`, `opencode`, `pi-acp`. Fails loud on missing/unknown | | `references` | string | Directory copied into `/work` before the agent starts | | `task` | string | User-like task sent to the agent | -| `graders` | array | Non-empty list of `{ name, command }` grader commands | +| `graders` | array | Non-empty list of `{ name, command }` grader commands. **Legacy `check:` and `artifacts:` are rejected.** | Optional fields: | Field | Type | Meaning | |-------|------|---------| +| `model` | string | Per-agent model identifier (see [Per-Agent Model Strings](#per-agent-model-strings)). Defaults to `openrouter/google/gemini-2.5-flash`. Only `pi-acp` enforces the `openrouter/...` prefix; native ACP agents (`claude-agent-acp`, `codex-acp`, `gemini`, `opencode`) pass the model string through to the agent CLI unchanged | | `setup` | string[] | Commands run in `/work` before the agent phase | | `cleanup` | string[] | Commands run after grading | | `env` | string[] | Host environment variable names forwarded into setup, agent, grading, and cleanup containers | | `mcpServers` | object | MCP servers exposed through the agent `mcp` tool | | `mcpServices` | object | Hidden local MCP services started as separate Docker containers | -| `model` | string | Default model for `run-case`; defaults to `openrouter/google/gemini-2.5-flash` | | `timeoutSeconds` | number | Agent timeout; defaults to `600` | All relative paths resolve from the case file directory. ## Suite Schema -Suites may contain inline case objects or paths to external case files. +Suites declare their agent + model matrix explicitly via `runs:`. Inline cases inherit the matrix; external case files declare their own `agent:`. ```yaml name: pdf-workbench-example references: ./references -models: - - openrouter/google/gemini-2.5-flash +runs: + - agent: pi-acp + model: openrouter/google/gemini-2.5-flash + - agent: claude-agent-acp + model: opus[1m] env: - OPENROUTER_API_KEY + - ANTHROPIC_API_KEY timeoutSeconds: 600 setup: - node $CASE/checks/_pdf.mjs write-inputs input @@ -109,6 +118,7 @@ appendSystemPrompt: | Keep task outputs at the top level of /work unless the user asks otherwise. cases: - name: extract-pdf-facts + agent: pi-acp task: | Read statement.pdf and write answer.json with the account, quarter, approval code, and risk flags. graders: @@ -122,8 +132,8 @@ Suite fields: | Field | Required | Meaning | |-------|----------|---------| | `name` | yes | Suite name in aggregate output | -| `models` | yes | OpenRouter model refs for the case/model matrix | -| `cases` | yes | Inline case objects or paths to case files | +| `runs` | yes | Non-empty list of `{ agent, model }` rows. Each row is one (agent, model) combination the suite executes. **Legacy `models:` is rejected.** | +| `cases` | yes | Inline case objects or paths to case files. Inline cases must still declare `agent:` (an inline case's `agent:` field is what its trial uses; the suite's `runs:` matrix is what's iterated) | | `references` | no | Default references dir for inline cases; defaults to `./references` | | `env` | no | Default env allowlist for inline cases | | `setup` | no | Default setup commands for inline cases | @@ -137,6 +147,30 @@ Inline case fields override suite defaults. External case files are loaded from Environment variables listed in `env` are forwarded unchanged. This intentionally supports live integration evals such as authenticated CLI calls, but it also means the agent can read or print those values through shell tools. Use dedicated test accounts, least-privilege credentials, and cleanup routines for live systems. Treat `trace.jsonl`, `result.json`, grader evidence, stdout/stderr, and preserved `workspace/` directories as potentially sensitive if an agent or grader prints or writes secret values. +## Per-Agent Model Strings + +Each agent CLI has its own model-naming convention. The workbench passes the case's `model:` string through to the agent's CLI unchanged (except for `pi-acp`, which enforces an `openrouter/...` prefix). If the model string isn't valid for that agent, the agent CLI fails the trial — the workbench will NOT catch this at suite-load time. **Use the table below when writing `suite.yml` to avoid runtime rejections.** + +| Agent | Format | Valid examples | Notes | +|-------|--------|----------------|-------| +| `claude-agent-acp` | Claude CLI enum | `opus`, `opus[1m]`, `sonnet`, `sonnet[1m]`, `haiku`, `default` | Bracket suffix `[1m]` selects the 1M-context window. Full model IDs like `claude-opus-4-7` are **not** accepted by the CLI. Set via `unstable_setSessionModel` after `newSession` | +| `codex-acp` | OpenAI model IDs | `gpt-5`, `gpt-5-mini`, `gpt-4o`, `o3-mini` | Whatever Codex's `--model` flag accepts | +| `gemini` | Gemini model IDs | `gemini-2.5-pro`, `gemini-2.5-flash`, `gemini-2.0-flash` | Whatever `gemini --model` accepts | +| `opencode` | `provider/model` form | `anthropic/claude-sonnet-4`, `openai/gpt-5`, `google/gemini-2.5-pro` | Must include the provider prefix | +| `pi-acp` | `openrouter/...` only | `openrouter/google/gemini-2.5-flash`, `openrouter/anthropic/claude-sonnet-4`, `openrouter/openai/gpt-5` | Enforced by the case loader; non-`openrouter/` prefixes are rejected at load time | + +**Default if `model:` is omitted:** `openrouter/google/gemini-2.5-flash`. This default only makes sense for `pi-acp`; specify `model:` explicitly when using any other agent. + +**Required env per agent** (declare these in `env:` so they propagate into the trial container): + +| Agent | Env var | Alternative | +|-------|---------|-------------| +| `claude-agent-acp` | `ANTHROPIC_API_KEY` | `~/.claude/.credentials.json` (subscription) | +| `codex-acp` | `OPENAI_API_KEY` | `~/.codex/auth.json` (subscription) | +| `gemini` | `GOOGLE_API_KEY` | `~/.gemini/oauth_creds.json` (subscription) | +| `opencode` | `OPENAI_API_KEY` (or provider-specific key) | — | +| `pi-acp` | `OPENROUTER_API_KEY` | — | + ## MCP Servers `mcpServers` uses mcporter-compatible server entries. During each Docker trial, the workbench writes `/work/mcporter.json` with `imports: []` and exposes an `mcp` command on `PATH`. @@ -507,7 +541,7 @@ Files to inspect: | File | Purpose | |------|---------| | `examples/workbench/README.md` | Top-level example command walkthrough | -| `examples/workbench/pdf/suite.yml` | Inline suite using models, setup, graders, and append prompt | +| `examples/workbench/pdf/suite.yml` | Inline suite using `runs:`, setup, graders, and append prompt | | `examples/workbench/pdf/references/pdf-skill/SKILL.md` | Skill under test copied into `/work` | | `examples/workbench/pdf/checks/*.mjs` | Deterministic graders and setup helpers | | `examples/workbench/pdf/README.md` | Demo walkthrough | @@ -526,8 +560,9 @@ npx tsx src/cli.ts --help node dist/cli.js --help ``` -For runner/Docker changes, rebuild the image: +For runner/Docker changes, rebuild the image (this is also what +`run-case` / `run-suite` expect as the default image): ```bash -docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile . +docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile . ``` diff --git a/skills/shared/workflow.md b/skills/shared/workflow.md index 58539f0..89958c3 100644 --- a/skills/shared/workflow.md +++ b/skills/shared/workflow.md @@ -105,12 +105,20 @@ skill-evals// # TRACKED — eval suites + probes (reusable a suite.yml # step 4 generates .skill-optimizer// # GITIGNORED — heavy ephemeral artifacts - vendored-skill/ # step 1 vendors (upstream OR local) + vendored-skill/ # step 1 vendors (upstream OR local) — the wrapper / ship target + vendored-underlying/ # only if step 1's wrapper-redirect gate fired and user chose re-vendor bench-results// # step 6 raw output (suite-result.json, trace.jsonl, findings.txt) autopilot-/ # step 10 per-run scratch journal.md # appended-to during the run ``` +When `vendored-underlying/` exists, **the testing target is still +`vendored-skill/`** (that's what ships when the user installs the +skill). The underlying directory carries the actual rules for +analysis/improvement reference; the optimizer/validator may patch +files in either folder depending on where the live rule content +sits. + `` is `--` for upstream skills, or `` for local skills. The same `` is reused across all three locations so a slug's full state can be located diff --git a/src/workbench/docker-runner.ts b/src/workbench/docker-runner.ts index 8e48594..564a4c6 100644 --- a/src/workbench/docker-runner.ts +++ b/src/workbench/docker-runner.ts @@ -133,9 +133,19 @@ function copyCaseSupportDir(sourceCaseDir: string, bundledCaseDir: string, name: cpSync(sourceDir, destinationDir, { recursive: true }); } +// `references/` is handled separately via copyDirectoryContents; skipping it here avoids +// double-copy. Hidden dirs (e.g., `.results`, `.skill-optimizer`) are bench output, not +// case content. `node_modules/` is excluded for cost reasons; nothing in a case should +// depend on it at grade time. +const EXCLUDED_CASE_SUBDIRS = new Set(['references', 'node_modules']); + function copyCaseSupportDirs(sourceCaseDir: string, bundledCaseDir: string): void { - for (const name of ['checks', 'fixtures', 'bin', 'workspace', 'mcp']) { - copyCaseSupportDir(sourceCaseDir, bundledCaseDir, name); + if (!existsSync(sourceCaseDir)) return; + for (const entry of readdirSync(sourceCaseDir, { withFileTypes: true })) { + if (!entry.isDirectory()) continue; + if (entry.name.startsWith('.')) continue; + if (EXCLUDED_CASE_SUBDIRS.has(entry.name)) continue; + copyCaseSupportDir(sourceCaseDir, bundledCaseDir, entry.name); } } @@ -565,6 +575,16 @@ export async function runOneAcpTrial(params: RunOneAcpTrialOptions): Promise/dev/null || true`, + { cwd: params.repoRoot }, + ); await runShellCommand(`docker rm -f ${shellQuote(containerName)}`, { cwd: params.repoRoot }); } } From a37b75fa4ea8fc78ec0381cf15c7275e751e0cda Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 08:46:02 -0500 Subject: [PATCH 119/121] fix(autopilot): auto-handle branch + add CWD discipline; remove stale worktrees MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Trial run on web-interface-guidelines surfaced two issues that this commit addresses, plus the long-standing worktree-gitlink clutter. Branch handling. The step 1 branching prompt was "always ask", which interrupts an autopilot full-auto run after the operator already said "all auto". Autopilot now handles branching itself in workflow (a): reads current branch, creates eval/ if not already on a purpose-named branch, logs to the journal, and passes 'branching: already-resolved' in OPERATOR_DIRECTIVES so step 1 skips its prompt. Step 1 still asks when invoked manually (directive absent). CWD discipline. The orchestrator ran 'cd .skill-optimizer/.../vendored-skill && curl ...' to vendor source files. The Bash tool keeps CWD between calls, so the cd persisted; later 'mkdir -p docs/skill-optimizer//' resolved under vendored-skill/ silently, while the dispatched subagent (with absolute output path) wrote the canonical 01-functionality.md to the correct main-checkout location. The orchestrator then "couldn't find" the file because its relative-path verification looked in the wrong subtree. Added a Working-directory discipline section to autopilot SKILL.md and a CWD-safe vendor-step note to investigate-functionality: prefer absolute paths or (subshell) scoping; never bare cd that persists. Worktree cleanup. .claude/worktrees/ held 10 stale gitlinks (mode 160000) from old auto-pilot pilot runs in May. All worktrees removed via 'git worktree remove --force'; gitlinks dropped from the tree. Branches themselves stay (eval/auto-pilot/shadcn-ui, eval/auto-pilot/firecrawl-build-scrape, etc.) — if you want them gone, 'git branch -D' separately. Co-Authored-By: Claude Opus 4.7 (1M context) --- .claude/worktrees/agent-a0620e810e20af755 | 1 - .claude/worktrees/agent-a2bdbcde626027d00 | 1 - .claude/worktrees/agent-aac9dc7ea895a8a09 | 1 - .claude/worktrees/agent-aaf823bf7de0c6401 | 1 - .claude/worktrees/agent-abe33dd2c200c608a | 1 - .claude/worktrees/agent-browser-rerun | 1 - .claude/worktrees/firebase-v1.3 | 1 - .claude/worktrees/supabase-pilot-v2 | 1 - .claude/worktrees/v1.3-impl | 1 - .claude/worktrees/v1.4-rationale | 1 - skills/autopilot/SKILL.md | 52 +++++++++++++++++++++-- skills/investigate-functionality/SKILL.md | 25 +++++++++-- 12 files changed, 70 insertions(+), 17 deletions(-) delete mode 160000 .claude/worktrees/agent-a0620e810e20af755 delete mode 160000 .claude/worktrees/agent-a2bdbcde626027d00 delete mode 160000 .claude/worktrees/agent-aac9dc7ea895a8a09 delete mode 160000 .claude/worktrees/agent-aaf823bf7de0c6401 delete mode 160000 .claude/worktrees/agent-abe33dd2c200c608a delete mode 160000 .claude/worktrees/agent-browser-rerun delete mode 160000 .claude/worktrees/firebase-v1.3 delete mode 160000 .claude/worktrees/supabase-pilot-v2 delete mode 160000 .claude/worktrees/v1.3-impl delete mode 160000 .claude/worktrees/v1.4-rationale diff --git a/.claude/worktrees/agent-a0620e810e20af755 b/.claude/worktrees/agent-a0620e810e20af755 deleted file mode 160000 index 1744daf..0000000 --- a/.claude/worktrees/agent-a0620e810e20af755 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 1744daf5b8f4c0c73ff08842424810422b4de977 diff --git a/.claude/worktrees/agent-a2bdbcde626027d00 b/.claude/worktrees/agent-a2bdbcde626027d00 deleted file mode 160000 index 41009f8..0000000 --- a/.claude/worktrees/agent-a2bdbcde626027d00 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 41009f87e932bc519c1560d0190c3363c47d2d6b diff --git a/.claude/worktrees/agent-aac9dc7ea895a8a09 b/.claude/worktrees/agent-aac9dc7ea895a8a09 deleted file mode 160000 index f0883ad..0000000 --- a/.claude/worktrees/agent-aac9dc7ea895a8a09 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit f0883adc271cb6a73b7d1f58ab29da421faa2da0 diff --git a/.claude/worktrees/agent-aaf823bf7de0c6401 b/.claude/worktrees/agent-aaf823bf7de0c6401 deleted file mode 160000 index b342220..0000000 --- a/.claude/worktrees/agent-aaf823bf7de0c6401 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit b34222097d65ffbfc545eaaeff716b4b0794c9e0 diff --git a/.claude/worktrees/agent-abe33dd2c200c608a b/.claude/worktrees/agent-abe33dd2c200c608a deleted file mode 160000 index f9463be..0000000 --- a/.claude/worktrees/agent-abe33dd2c200c608a +++ /dev/null @@ -1 +0,0 @@ -Subproject commit f9463bec339d63d71482a50bff7c36dcac97417d diff --git a/.claude/worktrees/agent-browser-rerun b/.claude/worktrees/agent-browser-rerun deleted file mode 160000 index 26601a0..0000000 --- a/.claude/worktrees/agent-browser-rerun +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 26601a0adee88601ff0004bb2795f77993496911 diff --git a/.claude/worktrees/firebase-v1.3 b/.claude/worktrees/firebase-v1.3 deleted file mode 160000 index c519d28..0000000 --- a/.claude/worktrees/firebase-v1.3 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit c519d281a48141e6dae37435d22c5539ec868cd5 diff --git a/.claude/worktrees/supabase-pilot-v2 b/.claude/worktrees/supabase-pilot-v2 deleted file mode 160000 index 59c3e85..0000000 --- a/.claude/worktrees/supabase-pilot-v2 +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 59c3e85e3fe254832a79bc8948cf0da147e523f4 diff --git a/.claude/worktrees/v1.3-impl b/.claude/worktrees/v1.3-impl deleted file mode 160000 index 96dbe31..0000000 --- a/.claude/worktrees/v1.3-impl +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 96dbe31569e151bc2be960994201a767275bcf6e diff --git a/.claude/worktrees/v1.4-rationale b/.claude/worktrees/v1.4-rationale deleted file mode 160000 index 2e4ea27..0000000 --- a/.claude/worktrees/v1.4-rationale +++ /dev/null @@ -1 +0,0 @@ -Subproject commit 2e4ea27ccf9e4a3e55a9a4c86ea113a6151768ce diff --git a/skills/autopilot/SKILL.md b/skills/autopilot/SKILL.md index 19473c6..4dda490 100644 --- a/skills/autopilot/SKILL.md +++ b/skills/autopilot/SKILL.md @@ -46,10 +46,29 @@ Classify the source the same way step 1 does: - **Local:** filesystem path to a SKILL.md (or a directory containing one). Skips the PR question. -If the user is on `main`, `master`, `development`, or a similar -long-lived branch, recommend creating a dedicated branch (e.g., -`git checkout -b eval/`) before continuing. Same branch -discipline as step 1. +**Auto-handle the branch in full-auto mode.** Step 1 has its own +explicit branching prompt for manual invocation, but in autopilot +the operator already said "all auto" — interrupting with a branch +question contradicts that. Autopilot handles it up-front: + +1. Run `git rev-parse --abbrev-ref HEAD` to read the current branch. +2. If the current branch already matches `eval/` or `feat/`, + reuse it. +3. Otherwise, create `eval/` and switch to it: + `git checkout -b eval/`. +4. Log the decision to the journal: either + `auto-created branch eval/` or + `using existing branch `. +5. When dispatching step 1 (in workflow (d)), include in + `${OPERATOR_DIRECTIVES}` the line + `branching: already-resolved ()` so step 1 skips + its branching prompt. + +If the operator picked the interactive surface for this (an +exception — autopilot's startup interview doesn't surface branching +as a breakpoint by default), prompt instead per +[`agents/surfacing-prompt.md`](./agents/surfacing-prompt.md) and +follow their answer. Create a TodoWrite list with all 10 chain steps so progress is visible across the long run. Mark step 10 (this skill) as @@ -260,6 +279,31 @@ the gate's context substituted in. Wait for the operator's response; the parsing rules (strategy number, custom directive, "honest exit") live in that template. +## Working-directory discipline + +A long chain run with many `Bash` calls is sensitive to CWD drift. +The `Bash` tool preserves CWD between calls — a single bare `cd` +into a subdirectory persists for the rest of the run, and later +relative paths (e.g., `mkdir -p docs/skill-optimizer//...`) +silently resolve in the wrong place. + +**Two rules:** + +1. **Never use bare `cd subdir && cmd`** that would persist. Either: + - Use absolute paths: `curl -o /home/.../vendored-skill/SKILL.md ...` + - Or scope with a subshell: `(cd vendored-skill && curl -o SKILL.md ...)` +2. **Verify with absolute paths** when checking that a downstream + step's output landed: `ls -la /home/.../docs/skill-optimizer//` + rather than `ls -la docs/...`. The orchestrator and the + dispatched subagents may have different effective CWDs; absolute + paths remove that ambiguity. + +If a step's canonical "appears missing" after a successful +dispatch, the first thing to check is the orchestrator's current +working directory — `pwd` it, then re-look at the canonical with +an absolute path. The subagent was likely correct; the relative +path was resolving to a sibling tree. + ## Edge cases - **Re-entrancy on the same slug.** Re-invoking autopilot on the diff --git a/skills/investigate-functionality/SKILL.md b/skills/investigate-functionality/SKILL.md index 16446da..c2cd3b6 100644 --- a/skills/investigate-functionality/SKILL.md +++ b/skills/investigate-functionality/SKILL.md @@ -76,9 +76,15 @@ bench output, gitignored). Keeping the tracked locations on a feature branch makes the run easy to discard, iterate on, or merge in one piece. -**Always ask the user explicitly** — don't infer from branch -name. After running `git rev-parse --abbrev-ref HEAD`, surface -this prompt verbatim: +**Honor an autopilot directive first.** If +`${OPERATOR_DIRECTIVES}` contains a line like +`branching: already-resolved ()`, the autopilot driver has +already handled branch creation for this run. Log "branch +confirmed by autopilot: " and proceed without prompting. + +**Otherwise, always ask the user explicitly** — don't infer from +branch name. After running `git rev-parse --abbrev-ref HEAD`, +surface this prompt verbatim: > I'm about to start an optimization run for ``. The chain > will create tracked files under `docs/skill-optimizer//` @@ -161,6 +167,19 @@ Copy the skill's files into `.skill-optimizer//vendored-skill/` - **Local** — `cp -r` of the local skill's directory into the vendored path +**Do not `cd` into the vendored directory.** The `Bash` tool +preserves CWD between calls — a bare `cd vendored-skill && curl ...` +will leak CWD into the rest of the chain, causing later relative +paths (e.g., `mkdir -p docs/skill-optimizer//`) to silently +resolve under `vendored-skill/`. Use one of: + +- **Absolute paths:** + `curl -sSfL -o //.skill-optimizer//vendored-skill/SKILL.md ...` +- **Subshell scoping (parentheses):** + `(cd .skill-optimizer//vendored-skill && curl -sSfL -o SKILL.md ...)` +- **Tool's `-o`/`-O` flags or `cp` directly:** + `cp source.md .skill-optimizer//vendored-skill/SKILL.md` + If `.skill-optimizer//vendored-skill/` already exists from a prior run, reuse it unless: (1) the source URL changed (upstream — a different repo or skill is being investigated), or (2) the From 69b414c7438c35ed5c6ada3e423f84a245a881fc Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 10:39:28 -0500 Subject: [PATCH 120/121] feat(autopilot): empirical-verification re-bench + suite skillUnderTest inheritance MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A validator-approved improvement is now backed by an actual bench comparison before autopilot claims `improved`. Without this, the exit_status `improved` only meant "validator believed the proposal" — no measurement of whether the proposal moved the needle. Autopilot workflow gains phase (e.5) — fires after step 9 approve: 1. Snapshot baseline pass-rate from 06-bench-summary.md. 2. Patch suite.yml with a top-level skillUnderTest pointing at improved-skill/; back the original up to the journal dir. 3. Re-dispatch step 6 against the improved skill. 4. Restore suite.yml; the re-bench raw output lives at bench-results/-after-improvement/. 5. Per-case comparison: count pass→fail (regressions), fail→pass (new passes), and the net delta. New gate G10-NO-EMPIRICAL-GAIN fires when the re-bench shows pass-rate ≤ baseline or any regression. Recovery menu (3 strategies, cheap→expensive): different fix, reframe weakness, increase test coverage. Cap exhaustion → exit_status: unchanged-empirical. New exit_status values: improved-with-regression (net gain but ≥1 case regressed — surface clearly for review), unchanged-empirical (re-bench didn't demonstrate the validator's claim). Suite-loader change (the only code touched): suite-level skillUnderTest now inherits to inline cases (with per-case override), mirroring how references/env/setup/etc. already work. Lets autopilot patch one top-level field instead of N cases. Test covers inheritance + override + malformed-field rejection. workbench.md Suite Schema doc gains a row for skillUnderTest with its purpose called out. Co-Authored-By: Claude Opus 4.7 (1M context) --- skills/autopilot/SKILL.md | 78 +++++++++++++++++++-- skills/autopilot/agents/recovery-menus.md | 26 +++++++ skills/autopilot/agents/summary-template.md | 36 +++++++++- skills/shared/workbench.md | 1 + src/workbench/suite-loader.ts | 26 ++++++- tests/suite-loader-runs.test.ts | 59 ++++++++++++++++ 6 files changed, 216 insertions(+), 10 deletions(-) diff --git a/skills/autopilot/SKILL.md b/skills/autopilot/SKILL.md index 4dda490..61662ad 100644 --- a/skills/autopilot/SKILL.md +++ b/skills/autopilot/SKILL.md @@ -170,16 +170,78 @@ The protocol resolves to one of: Every gate firing gets one journal entry. Every retry decision increments the gate's counter and the global counter. +### (e.5) Empirical verification re-bench + +Fires only when step 9 returns `verdict: approve` AND +`docs/skill-optimizer//improved-skill/` exists. This phase +empirically demonstrates that the change actually helps — pass-rate +goes up, with no regressions. **Without this verification an +"improved" claim is unbacked.** + +1. **Snapshot baseline.** Read `06-bench-summary.md` frontmatter + for `overall_pass_rate` and `bench_results_path`; capture both + into the journal as `baseline_pass_rate` and + `baseline_bench_results_path`. The raw per-case results live at + `/suite-result.json` — preserve + that path; it's what we compare per-case against. +2. **Temporarily patch suite.yml.** Back up + `skill-evals//suite.yml` to + `.skill-optimizer//autopilot-/suite.yml.backup`. Then + add a top-level `skillUnderTest` field pointing at the improved + skill (the suite loader propagates this to all inline cases — + see `skills/shared/workbench.md` Suite Schema): + + ```yaml + skillUnderTest: + slug: + hostPath: //docs/skill-optimizer//improved-skill + ``` + +3. **Re-dispatch step 6** via `Skill skill-optimizer:run-bench`. + Output writes to a fresh + `.skill-optimizer//bench-results/-after-improvement/` + directory; `06-bench-summary.md` overwrites with the + after-improvement state (this is fine — it's now the latest + bench for the improved skill). +4. **Restore suite.yml** from the backup. The skill-evals tree + stays clean; future re-runs of the chain start from the original + suite. +5. **Compare.** Read the new `06-bench-summary.md` for + `overall_pass_rate` (after) and the new `bench_results_path`. + Per-case comparison: parse both `suite-result.json` files; for + each `(probe × run)` pair, note `baseline=pass/fail` vs + `after=pass/fail`. A `pass → fail` flip is a **regression**. +6. **Classify.** Decide which exit_status applies in (f) based on: + - `after_pass_rate > baseline_pass_rate` AND zero regressions → + `improved` + - `after_pass_rate > baseline_pass_rate` AND ≥1 regression → + `improved-with-regression` (surface clearly; the optimizer + bought one case at the cost of another) + - `after_pass_rate ≤ baseline_pass_rate` (or `==` with no + winning case flip) → **G10-NO-EMPIRICAL-GAIN fires** (see + gate menu) + +Record all four numbers into the journal: +`baseline_pass_rate`, `after_pass_rate`, +`pass_to_fail_regressions: []`, +`fail_to_pass_flips: []`. + ### (f) Honest-exit and write the summary Set `exit_status` per the run's final state: -- **`improved`** — step 9 verdict was `approve`; - `docs/skill-optimizer//improved-skill/` materialized. -- **`unchanged-honest-exit`** — cap exhausted at any gate, OR - step 7 reported no structural weakness with all G7-NW strategies - exhausted, OR step 9 rejected with all G9-REJECT strategies - exhausted. +- **`improved`** — step 9 approve AND the empirical re-bench + showed strictly increased pass-rate with zero regressions. +- **`improved-with-regression`** — pass-rate up but ≥1 case + regressed. Distinct from clean `improved` because it's worth a + human eye before merging. +- **`unchanged-empirical`** — step 9 approved but the empirical + re-bench showed pass-rate ≤ baseline (or zero net gain), and + all G10 recovery strategies were exhausted. Validator believed + the proposal helps; the bench says it doesn't. +- **`unchanged-honest-exit`** — cap exhausted at some earlier + gate (G7-NW, G8-BLOCKED, G9-REVISION, G9-REJECT) before reaching + step 9 approval. The chain never produced an improved-skill. - **`blocked-on-infra`** — G6-CLI-FAIL surfaced and the operator did not recover the environment. @@ -187,7 +249,8 @@ Promote the journal into `docs/skill-optimizer//autopilot-summary-.md` using [`agents/summary-template.md`](./agents/summary-template.md). The summary is an aggregate; trim narrative, keep retry trail + exit -reason + repo changes. +reason + repo changes + the empirical-verification numbers +(baseline → after pass-rate, regressions list, flips list). Update the chain TodoWrite list to mark step 10 `completed` and print the path to the summary report. @@ -212,6 +275,7 @@ are detected from the chain skill's canonical frontmatter and body. | G8-BLOCKED | Step 8 optimizer subagent returned `status: blocked` | Menu (2 options) | | G9-REVISION | `09-validator-verdict.md` has `verdict: needs-revision` | Menu (1 option) | | G9-REJECT | `09-validator-verdict.md` has `verdict: reject` | Menu (3 options) | +| G10-NO-EMPIRICAL-GAIN | (e.5) re-bench `overall_pass_rate` ≤ baseline OR ≥1 case regressed | Menu (3 options) | ### Recovery menus diff --git a/skills/autopilot/agents/recovery-menus.md b/skills/autopilot/agents/recovery-menus.md index debcb84..8b3d042 100644 --- a/skills/autopilot/agents/recovery-menus.md +++ b/skills/autopilot/agents/recovery-menus.md @@ -56,6 +56,32 @@ journal's retry log for what's been tried this run. the weakness wasn't real in the trace data, treat as G7-NW menu (more trials, more picks, more probes). +## G10-NO-EMPIRICAL-GAIN (cheap → expensive) + +Validator approved (step 9) but the empirical re-bench (phase +(e.5)) showed pass-rate ≤ baseline or introduced a regression. +The proposal sounded right but didn't move the needle. + +1. **Try a different fix, same weakness.** Re-run step 8 with the + rejected proposal noted as "approved by validator but added zero + bench gain — likely too abstract or addressed a non-load-bearing + detail." Then 9 → re-bench. Cheap (one optimization cycle). +2. **Reframe weakness.** Re-run step 7 with directive "prior + weakness X passed validator but didn't improve the bench; + reconsider whether the named pattern is the real failure mode." + Then 8 → 9 → re-bench. Medium. +3. **Increase test coverage.** If the after-rebench is identical + to baseline (e.g., both at 100%, no fail cases for the + improvement to flip), the probes can't *show* an improvement + that exists. Treat as G7-NW menu (more trials, enable more + picks, add new probes). Then continue from step 6 forward. + Expensive but the right move when the probes themselves are + the bottleneck. + +After cap exhaustion: exit_status `unchanged-empirical`. The +validator approved but the empirical evidence doesn't support +shipping the proposal as an improvement. + ## Gates that don't appear here - **G3** (picks) — single action, not retry-eligible. diff --git a/skills/autopilot/agents/summary-template.md b/skills/autopilot/agents/summary-template.md index 765f3b9..1ca0cbb 100644 --- a/skills/autopilot/agents/summary-template.md +++ b/skills/autopilot/agents/summary-template.md @@ -9,8 +9,10 @@ at workflow step (f), promoted from the run journal. Tracked at - `${RUN_STARTED_AT}` / `${RUN_ENDED_AT}` — ISO 8601 UTC - `${SLUG}` — chain slug - `${SOURCE}` — URL or local path the operator provided -- `${EXIT_STATUS}` — one of `improved`, `unchanged-honest-exit`, - `blocked-on-infra` +- `${EXIT_STATUS}` — one of `improved`, `improved-with-regression`, + `unchanged-empirical`, `unchanged-honest-exit`, `blocked-on-infra` +- `${EMPIRICAL_BLOCK}` — empirical-verification block (see format + below); omit if the run never reached step 9 approve - `${TOTAL_RETRIES}` — integer total across all gates - `${JOURNAL_PATH}` — gitignored journal path - `${HEADLINE}` — one paragraph: what happened end-to-end and why @@ -52,6 +54,10 @@ ${PER_STEP_ROWS} ${RECOVERY_SUMMARY} +## Empirical verification + +${EMPIRICAL_BLOCK} + ## Caveats ${CAVEATS} @@ -61,6 +67,32 @@ ${CAVEATS} ${REPO_CHANGES} ``` +## Empirical verification block format + +Always include this section when the run reached step 9 approve +(even if the empirical re-bench failed). Omit only if the chain +never produced an improved-skill (e.g., honest-exit at G7-NW with +no weakness found). + +```markdown +- **Baseline pass rate:** / (%) + — `.skill-optimizer//bench-results//` +- **After-improvement pass rate:** / (%) + — `.skill-optimizer//bench-results/-after-improvement/` +- **Net gain:** + cases passed (or 0, or negative) +- **Regressions** (pass → fail after improvement): +- **New passes** (fail → pass after improvement): +- **Verdict:** +``` + +When `exit_status: unchanged-empirical`, the verdict line should +spell out why this is honest rather than failed — e.g., "validator +approved but bench can't demonstrate the improvement (probes too +easy, weakness too subtle, or the fix is structural without +behavioral impact in the current test set)". + ## Per-step outcome row format ```markdown diff --git a/skills/shared/workbench.md b/skills/shared/workbench.md index 0e42745..78c415c 100644 --- a/skills/shared/workbench.md +++ b/skills/shared/workbench.md @@ -142,6 +142,7 @@ Suite fields: | `mcpServices` | no | Default hidden MCP service containers for inline cases, merged by service name | | `timeoutSeconds` | no | Default agent timeout for inline cases | | `appendSystemPrompt` | no | Extra suite-wide system prompt appended after the workbench prompt | +| `skillUnderTest` | no | `{ slug, hostPath }` — overrides which skill directory is mounted into the trial container. Inline cases inherit this; an inline case may override it with its own `skillUnderTest`. Used by autopilot's empirical re-bench to point the same suite at `improved-skill/` instead of the original vendored source | Inline case fields override suite defaults. External case files are loaded from their own file directory and do not inherit suite defaults. diff --git a/src/workbench/suite-loader.ts b/src/workbench/suite-loader.ts index 84dbae5..6c60740 100644 --- a/src/workbench/suite-loader.ts +++ b/src/workbench/suite-loader.ts @@ -5,7 +5,7 @@ import { parse as parseYaml } from 'yaml'; import { resolveAgent } from './agents/registry.js'; import { readMcpServers, readMcpServices, resolveWorkbenchCaseConfig } from './case-loader.js'; -import type { ResolvedWorkbenchCase, WorkbenchMcpServersConfig, WorkbenchMcpServicesConfig, WorkbenchRunSpec } from './types.js'; +import type { ResolvedWorkbenchCase, SkillUnderTestSpec, WorkbenchMcpServersConfig, WorkbenchMcpServicesConfig, WorkbenchRunSpec } from './types.js'; import { slugPathSegment } from './utils.js'; export interface ResolvedWorkbenchSuiteCase { @@ -70,6 +70,7 @@ interface SuiteCaseDefaults { mcpServers: WorkbenchMcpServersConfig; mcpServices: WorkbenchMcpServicesConfig; timeoutSeconds?: number; + skillUnderTest?: SkillUnderTestSpec; } function readSuiteCaseDefaults(parsed: Record, configPath: string): SuiteCaseDefaults { @@ -85,9 +86,31 @@ function readSuiteCaseDefaults(parsed: Record, configPath: stri mcpServers: readMcpServers(parsed, configPath), mcpServices: readMcpServices(parsed, configPath), timeoutSeconds: readOptionalTimeoutSeconds(parsed, configPath), + skillUnderTest: readOptionalSkillUnderTest(parsed, configPath), }; } +function readOptionalSkillUnderTest( + parsed: Record, + configPath: string, +): SkillUnderTestSpec | undefined { + const value = parsed.skillUnderTest; + if (value === undefined) return undefined; + if (!value || typeof value !== 'object' || Array.isArray(value)) { + throw new Error(`Workbench suite ${configPath}: field "skillUnderTest" must be an object with slug and hostPath`); + } + const obj = value as Record; + const slug = obj.slug; + const hostPath = obj.hostPath; + if (typeof slug !== 'string' || !slug.trim()) { + throw new Error(`Workbench suite ${configPath}: field "skillUnderTest.slug" must be a non-empty string`); + } + if (typeof hostPath !== 'string' || !hostPath.trim()) { + throw new Error(`Workbench suite ${configPath}: field "skillUnderTest.hostPath" must be a non-empty string`); + } + return { slug: slug.trim(), hostPath: hostPath.trim() }; +} + function resolveSuiteCase( entry: string | Record, index: number, @@ -133,6 +156,7 @@ function applySuiteDefaults( ...(Object.keys(mcpServers).length > 0 ? { mcpServers } : {}), ...(Object.keys(mcpServices).length > 0 ? { mcpServices } : {}), ...(defaults.timeoutSeconds !== undefined ? { timeoutSeconds: defaults.timeoutSeconds } : {}), + ...(defaults.skillUnderTest !== undefined ? { skillUnderTest: defaults.skillUnderTest } : {}), ...entry, ...(Object.keys(mcpServers).length > 0 ? { mcpServers } : {}), ...(Object.keys(mcpServices).length > 0 ? { mcpServices } : {}), diff --git a/tests/suite-loader-runs.test.ts b/tests/suite-loader-runs.test.ts index 495c579..5b98cd0 100644 --- a/tests/suite-loader-runs.test.ts +++ b/tests/suite-loader-runs.test.ts @@ -66,3 +66,62 @@ cases: [] ); rmSync(dir, { recursive: true }); }); + +test('loadWorkbenchSuite propagates suite-level skillUnderTest to inline cases', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +skillUnderTest: + slug: my-skill + hostPath: /tmp/my-skill-source +runs: + - agent: pi-acp + model: openrouter/anthropic/claude-haiku-4-5 +cases: + - name: case-a + agent: pi-acp + task: noop + graders: + - name: noop + command: 'true' + - name: case-b + agent: pi-acp + task: noop + graders: + - name: noop + command: 'true' + skillUnderTest: + slug: my-skill + hostPath: /tmp/override +`); + const suite = loadWorkbenchSuite(join(dir, 'suite.yml')); + assert.equal(suite.cases.length, 2); + // case-a inherits the suite-level skillUnderTest + assert.equal(suite.cases[0].case?.skillUnderTest?.slug, 'my-skill'); + assert.equal(suite.cases[0].case?.skillUnderTest?.hostPath, '/tmp/my-skill-source'); + // case-b overrides the suite-level value + assert.equal(suite.cases[1].case?.skillUnderTest?.hostPath, '/tmp/override'); + rmSync(dir, { recursive: true }); +}); + +test('loadWorkbenchSuite rejects malformed skillUnderTest', () => { + const dir = mkdtempSync(join(tmpdir(), 'suite-test-')); + mkdirSync(join(dir, 'references'), { recursive: true }); + writeFileSync(join(dir, 'suite.yml'), ` +name: my-suite +references: ./references +skillUnderTest: + slug: '' +runs: + - agent: pi-acp + model: openrouter/anthropic/claude-haiku-4-5 +cases: [] +`); + assert.throws( + () => loadWorkbenchSuite(join(dir, 'suite.yml')), + /skillUnderTest\.slug.*non-empty/, + ); + rmSync(dir, { recursive: true }); +}); From c453b1ff2818a8a6c86807388ba9f5d50be629ce Mon Sep 17 00:00:00 2001 From: Yuqing Zhai Date: Wed, 27 May 2026 20:43:53 -0500 Subject: [PATCH 121/121] =?UTF-8?q?feat(chain):=20unify=20wrapper=20handli?= =?UTF-8?q?ng=20=E2=80=94=20single=20vendored-skill,=20optimization=5Ftarg?= =?UTF-8?q?et=20picks=20the=20file?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Replaces the wrapper-detection branch (vendored-skill/ + vendored-underlying/ + 3-option re-vendor prompt) with a single-directory model that's both simpler and more honest about test reproducibility. What changes structurally - vendored-skill/ is now a single flat dir holding the source SKILL.md PLUS any files the source content fetches at runtime (the researcher vendors all referenced URLs as part of step 1). This makes the test bench reproducible without live network — whether the skill is a wrapper or not, the workbench mount is self-contained. - New frontmatter on 01-functionality.md: optimization_target (relative path within vendored-skill/, the file the chain modifies + submits upstream) + optimization_target_source (upstream URL of that file, used by step 2 to decide which repo to research for PR conventions — may differ from skill_source when a wrapper redirects). - optimization_target_candidates lists the viable files the researcher identified; downstream skills ignore it, but step 1's picker uses it. Picker UX - Replaces the 3-option re-vendor prompt with a "which file to optimize" picker. Surfaces only when multiple candidates exist AND no operator directive pre-resolves the choice. - Autopilot full-auto passes 'optimization-target: auto-pick-recommended' in OPERATOR_DIRECTIVES so step 1 honors the researcher's recommendation without surfacing. Same pattern as 'branching: already-resolved'. Per-skill updates - investigate-functionality: vendor-all-references rule + picker; new frontmatter fields documented. likely_wrapper kept as diagnostic only. - research-functionality (subagent): new section "Vendoring referenced content" + "Identifying the optimization target". Researcher fetches external URLs into vendored-skill/, lists candidates, recommends one. - investigate-submissions: reads optimization_target_source (not skill_source) to derive UPSTREAM_REPO for the gh CLI subagent. - improve / optimizer: reads ${OPTIMIZATION_TARGET}; diff scope is that single file. - validate / validator: BEFORE/AFTER comparison at ${OPTIMIZATION_TARGET}; other files in the trees should be identical (if they differ, the optimizer exceeded scope — validator flags). - autopilot: passes 'optimization-target: auto-pick-recommended' alongside the existing 'branching: already-resolved' in step 1's directives. - shared/workflow.md: dropped vendored-underlying/ from state layout; documented the new optimization_target / optimization_target_source contract. The wrapper concept doesn't disappear — likely_wrapper + wrapper_points_to are still recorded as diagnostic info, just no longer load-bearing for behavior. optimization_target is the field that drives the chain. Co-Authored-By: Claude Opus 4.7 (1M context) --- skills/autopilot/SKILL.md | 16 +- skills/improve/SKILL.md | 28 +++- skills/improve/agents/optimizer.md | 7 + skills/investigate-functionality/SKILL.md | 149 ++++++++++-------- .../agents/research-functionality.md | 116 +++++++++++--- skills/investigate-submissions/SKILL.md | 13 +- skills/shared/workflow.md | 20 ++- skills/validate/SKILL.md | 27 ++-- skills/validate/agents/validator.md | 6 + 9 files changed, 267 insertions(+), 115 deletions(-) diff --git a/skills/autopilot/SKILL.md b/skills/autopilot/SKILL.md index 61662ad..5f90490 100644 --- a/skills/autopilot/SKILL.md +++ b/skills/autopilot/SKILL.md @@ -64,9 +64,19 @@ question contradicts that. Autopilot handles it up-front: `branching: already-resolved ()` so step 1 skips its branching prompt. -If the operator picked the interactive surface for this (an -exception — autopilot's startup interview doesn't surface branching -as a breakpoint by default), prompt instead per +**Auto-resolve the optimization-target picker.** Step 1's +researcher subagent may identify multiple candidate optimization +targets (typically the source SKILL.md plus any underlying file +that the source fetches at runtime, for wrapper skills). In +manual invocation, step 1 surfaces a picker; in autopilot +full-auto mode, that's another contradiction. When dispatching +step 1, also include in `${OPERATOR_DIRECTIVES}` the line +`optimization-target: auto-pick-recommended` so step 1 honors +the researcher's recommendation without surfacing. + +If the operator picked the interactive surface for branching or +optimization-target (exceptions — autopilot's startup interview +doesn't surface either by default), prompt instead per [`agents/surfacing-prompt.md`](./agents/surfacing-prompt.md) and follow their answer. diff --git a/skills/improve/SKILL.md b/skills/improve/SKILL.md index 731214c..87cf9cb 100644 --- a/skills/improve/SKILL.md +++ b/skills/improve/SKILL.md @@ -68,11 +68,16 @@ Three checks: If you disagree, re-invoke step 7 with a directive; if you agree, exit honestly." Do NOT proceed. -3. `01-functionality.md` must exist. Skill content must be - readable: `docs/skill-optimizer//improved-skill/` if it - exists (accumulated state from prior step-9 approvals), else - `.skill-optimizer//vendored-skill/` (step 1 vendored the - source regardless of upstream/local). +3. `01-functionality.md` must exist with both + `optimization_target` and `optimization_target_source` + frontmatter fields set (step 1 records these — see step 1's + "What you produce" section). Skill content must be readable at + `docs/skill-optimizer//improved-skill/` if it exists + (accumulated state from prior step-9 approvals), else + `.skill-optimizer//vendored-skill/`. The + **`optimization_target` field identifies the specific file in + that directory that the optimizer modifies** — everything else + in the dir is context only. If `pr_submission_intent: true`, `02-submissions.md` should exist so the optimizer can shape the diff to upstream conventions from @@ -107,10 +112,19 @@ the rendered text as the Agent tool's `prompt` parameter (per [`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). -The optimizer sees: `07-analysis.md`; `01-functionality.md`; +The optimizer sees: `07-analysis.md`; `01-functionality.md` +(including the `optimization_target` and +`optimization_target_source` frontmatter fields); `${SKILL_CURRENT_PATH}` — `docs/skill-optimizer//improved-skill/` if it exists, else `.skill-optimizer//vendored-skill/`; -`02-submissions.md` if PR-bound; `${OPERATOR_DIRECTIVES}`. +`${OPTIMIZATION_TARGET}` — the relative path within +`${SKILL_CURRENT_PATH}` of the file to modify, copied from +`01-functionality.md`'s frontmatter; `02-submissions.md` if +PR-bound; `${OPERATOR_DIRECTIVES}`. + +The optimizer modifies **only the file at `${OPTIMIZATION_TARGET}`**. +The proposal's diff scope is that single file. Other files in +`${SKILL_CURRENT_PATH}` are reference context, not edit targets. The optimizer does NOT see: raw failed trials, `findings.txt`, `trace.jsonl`; grader internals diff --git a/skills/improve/agents/optimizer.md b/skills/improve/agents/optimizer.md index 09b3d53..352864c 100644 --- a/skills/improve/agents/optimizer.md +++ b/skills/improve/agents/optimizer.md @@ -27,6 +27,13 @@ rubric. step-9 approvals), else `.skill-optimizer//vendored-skill/` (step 1 vendored the source regardless of upstream/local). You propose a new improvement on top of whatever current state you read. +- `${OPTIMIZATION_TARGET}` — the **relative path within + `${SKILL_CURRENT_PATH}` of the file you modify**. From step 1's + `01-functionality.md` frontmatter. Default is `SKILL.md`; for + wrapper skills it points at the underlying rule file (e.g., + `command.md`). Your proposed diff is scoped to this single file — + other files in `${SKILL_CURRENT_PATH}` are context, not edit + targets. - `${SUBMISSIONS_PATH}` — `02-submissions.md` if PR-bound (lets you shape the diff to match upstream conventions from the start, reducing step-9 round-trips) diff --git a/skills/investigate-functionality/SKILL.md b/skills/investigate-functionality/SKILL.md index c2cd3b6..863f822 100644 --- a/skills/investigate-functionality/SKILL.md +++ b/skills/investigate-functionality/SKILL.md @@ -25,8 +25,15 @@ Frontmatter (runtime-relevant facts only, per skill_source: pr_submission_intent: true | false classification: +optimization_target: +optimization_target_source: +optimization_target_candidates: + - path: SKILL.md + source: + - path: command.md + source: likely_wrapper: true | false -wrapper_points_to: +wrapper_points_to: --- ``` @@ -36,20 +43,28 @@ cleanly, write a short descriptive label of your own (`dataset-extraction`, `deployment-runbook`, etc.) rather than falling back to `other`. -**`likely_wrapper`** — the subagent's judgment that the source the -user pointed at appears to be a thin wrapper file (e.g., a -SKILL.md that mostly references content elsewhere, common in -multi-agent plugins where one canonical agent-agnostic content -file is wrapped by several agent-flavored SKILL.md files). When -`true`, **`wrapper_points_to`** records where the subagent thinks -the actual content lives. Downstream chain steps (step 2 submission -research, step 8 optimization, step 9 validation) consult these fields -when deciding what to target. - -If the subagent flags `likely_wrapper: true`, the operator surfaces -to the user (workflow step (h) below) and the user decides what -to do — re-vendor and re-research the referenced content, or -proceed treating the wrapper itself as the skill to optimize. +**`optimization_target`** — the file in `vendored-skill/` that the +chain optimizes and submits upstream. For non-wrapper skills this is +`SKILL.md`. For wrapper skills (where the source file fetches its +real rules from another URL) this is the underlying file +(e.g. `command.md`). Identified by the researcher subagent; surfaced +to the operator via a picker prompt at step (h) only when more than +one viable candidate exists. + +**`optimization_target_source`** — the upstream URL of the +optimization target. Step 2 (investigate-submissions) reads this +to determine which repo to research for PR conventions — may differ +from `skill_source` when a wrapper redirects to another repo. + +**`optimization_target_candidates`** — the candidate list the +researcher identified. Used by step (h)'s picker; downstream steps +ignore it (the resolved `optimization_target` is what matters). + +**`likely_wrapper`** + **`wrapper_points_to`** — diagnostic. Kept +for human review and debugging; they no longer drive behavior +(`optimization_target` is the load-bearing field). The researcher +still records them because downstream skill optimizers may surface +them in body sections. For the body template, see [`agents/research-functionality.md`](./agents/research-functionality.md). @@ -167,6 +182,15 @@ Copy the skill's files into `.skill-optimizer//vendored-skill/` - **Local** — `cp -r` of the local skill's directory into the vendored path +**`vendored-skill/` is a single flat dir holding everything the +chain needs.** The researcher subagent (step (g)) is responsible +for identifying any external URLs the source content references at +runtime (e.g. a wrapper that WebFetches `command.md` from another +repo) and fetching them into `vendored-skill/` too — so the test +environment is fully self-contained and reproducible without live +network. You don't need to do that here; you just vendor the source +itself at this step. + **Do not `cd` into the vendored directory.** The `Bash` tool preserves CWD between calls — a bare `cd vendored-skill && curl ...` will leak CWD into the rest of the chain, causing later relative @@ -237,61 +261,52 @@ user has already expressed. A subagent walled off from that context produces a fresh derivation from the source itself, not a rationalization of expectations. -### (h) Confirm + handle wrapper detection + hand off +### (h) Confirm + handle optimization-target picker + hand off Verify the report file exists and frontmatter parses. -**If the subagent flagged `likely_wrapper: true`** in the -frontmatter (it judged the source to be a thin wrapper pointing at -content elsewhere), surface to the user before handing off: - -> The researcher subagent thinks the source you pointed at -> (``) is likely a wrapper file — it appears to -> reference the actual content at ``. Three -> realistic responses: -> -> 1. **Re-vendor the referenced content alongside the wrapper.** -> Investigation re-runs against the underlying content so the -> `01-functionality.md` reflects the actual rules; tests will -> still target the wrapper (the wrapper is what ships, so -> that's what should be exercised). The optimizer may modify -> files in either folder — both stay available. (Common — -> usually what you want.) -> 2. **Proceed treating this wrapper as the skill.** Don't fetch -> the underlying content; everything downstream uses the -> wrapper as-is. Useful if you only want to improve the wrapper -> itself. -> 3. **Cancel and provide a different source.** If the subagent's -> judgment is wrong or you meant to point elsewhere. - -If the user picks (1): - -1. **Keep the wrapper at `.skill-optimizer//vendored-skill/` - intact.** That's the ship target — tests install this folder; - from the user's perspective it IS the skill. -2. Vendor the content at `wrapper_points_to` into a sibling - directory: `.skill-optimizer//vendored-underlying/`. -3. Re-invoke this skill from (d) with the **underlying** content - as the research target. `${SKILL_SOURCE}` in this re-run points - at `wrapper_points_to` so the researcher builds an accurate - `01-functionality.md`, but `vendored-skill/` is not touched. -4. After the re-run, the report frontmatter still records the - user's original `skill_source` (the wrapper URL/path) and - `wrapper_points_to`. Downstream steps know: test against - `vendored-skill/` (the wrapper); the optimizer/validator may - read `vendored-underlying/` for context and may patch files in - either folder depending on where the actual rule content lives. -5. If the underlying content is itself a wrapper, the re-run will - flag `likely_wrapper: true` again and the gate repeats. Recurse - into a deeper `vendored-underlying/` chain only if you want to - walk it. - -If (2): keep the report as-is and proceed. No `vendored-underlying/` -is created. Downstream steps see `likely_wrapper: true` with -`wrapper_points_to` documented; the PR target is the wrapper -itself (the user's deliberate choice). - -If (3): exit; the user will re-invoke with a new source. +**Pick the optimization target.** The researcher subagent has +recorded `optimization_target_candidates` (one entry per file it +identified as a viable target — typically the source `SKILL.md` +itself, plus any external file the source fetches at runtime that +it also vendored locally). It has also recorded its **recommended** +choice as `optimization_target` in the frontmatter. + +Resolve the choice: + +- **If `${OPERATOR_DIRECTIVES}` contains an `optimization-target:` + line** (autopilot passes this in full-auto mode): honor it + without prompting. Values: `auto-pick-recommended` (use the + researcher's recommendation as-is) or a specific relative path + matching one of the candidates. + +- **If `optimization_target_candidates` has exactly one entry** + (non-wrapper, non-ambiguous case): no prompt; the single + candidate is already recorded as `optimization_target`. Proceed. + +- **Otherwise (multiple candidates, no directive)**: surface this + prompt to the user: + + > The source you provided needs a clear optimization target — + > the file we'll modify and submit upstream as the PR. + > Multiple candidates exist in `vendored-skill/`: + > + > 1. **``** *(recommended)* — ``. + > Upstream source: ``. PR would target + > ``. + > 2. **``** — ``. + > Upstream source: ``. PR would target + > ``. + > 3. ... (one entry per remaining candidate) + > N. **Other** — name a relative path inside `vendored-skill/` + > not in the list. + > + > Say "default" or "go" to accept option 1. + + If the user picks a different candidate (or "other"), update the + `optimization_target` and `optimization_target_source` frontmatter + fields in `01-functionality.md` to reflect their choice. Use the + `Edit` tool — don't rewrite the whole file. Once resolved, hand off based on this report's `pr_submission_intent` field: diff --git a/skills/investigate-functionality/agents/research-functionality.md b/skills/investigate-functionality/agents/research-functionality.md index 15f0f9c..781987e 100644 --- a/skills/investigate-functionality/agents/research-functionality.md +++ b/skills/investigate-functionality/agents/research-functionality.md @@ -28,6 +28,18 @@ the chain consumes. skill is about - The operator's directives +## What you write (your job, in order) + +1. **Fetch all externally-referenced files** into `${VENDORED_PATH}` + so the test bench is reproducible without live network. See + "Vendoring referenced content" below. +2. **Identify optimization-target candidates** — record each viable + target file (the source SKILL.md + any underlying files you + vendored) into `optimization_target_candidates`. Recommend one. +3. **Write the briefing report** at `${OUTPUT_PATH}` with the + frontmatter shape below and the body covering classification, + responsibilities, dependencies, etc. + ## What you do NOT see - Prior `01-functionality.md` drafts or git history of the file @@ -50,8 +62,17 @@ Frontmatter (runtime-relevant only, no version-tracking metadata): skill_source: pr_submission_intent: classification: -likely_wrapper: -wrapper_points_to: +optimization_target: +optimization_target_source: +optimization_target_candidates: + - path: + source: + rationale: + - path: + source: + rationale: +likely_wrapper: +wrapper_points_to: --- ``` @@ -102,6 +123,61 @@ Body (markdown) covering: who decides whether to re-vendor the referenced content and re-research, treat the wrapper as the skill, or cancel. +## Vendoring referenced content (load-bearing) + +The test environment must be **fully self-contained** — the agent +running in a trial container should not need live network to +exercise the skill. As part of your research: + +1. Scan the skill files at `${VENDORED_PATH}` for any external URL + the source content references at runtime (e.g. a wrapper SKILL.md + that says "WebFetch this URL for the actual rules", or a `Read` + reference to a remote file). +2. For each such reference, **fetch the file into `${VENDORED_PATH}` + with its basename** (use absolute paths or subshell `(cd && curl)` + per the CWD discipline rules). Don't `cd` into `${VENDORED_PATH}` + in a way that persists across your tool calls. +3. After fetching, every URL referenced by the source content should + now have a local copy under `${VENDORED_PATH}`. Downstream chain + steps mount this directory as the test bench; an agent running + against it should be able to satisfy the skill's references + without external network. + +This applies whether the source is a wrapper or not. Many skills +reference `references/*.md` files that may be missing from the +vendored copy; pull anything cited via URL. + +## Identifying the optimization target + +The **optimization target** is the file the chain will modify +(step 8) and submit upstream (when PR-bound). For each viable +candidate file at `${VENDORED_PATH}`, record an entry in +`optimization_target_candidates`: + +```yaml +- path: + source: + rationale: +``` + +Common patterns: + +- **Non-wrapper skill**: one candidate, the source `SKILL.md`. Set + `optimization_target: SKILL.md` and `likely_wrapper: false`. +- **Wrapper skill** (the source file fetches its real rules from + another URL): two candidates — the wrapper SKILL.md itself, AND + the underlying file. Set `optimization_target` to the **underlying** + (the file with the actual rules — that's where meaningful + optimization happens) and `likely_wrapper: true`. +- **Ambiguous skill** (e.g., multi-file plugin where multiple files + carry rule content): record all viable candidates; recommend the + one most likely to be the rule corpus. The operator will surface + the picker to the user. + +The `optimization_target_source` field records the upstream URL of +your recommended target. This is what step 2 (investigate-submissions) +will use to decide which repo to research for PR conventions. + ## Reasoning protocol 1. **Read the skill files first.** Don't research the technology @@ -124,9 +200,9 @@ Body (markdown) covering: the secondary aspect in the body. Don't invent classification subtypes. -## Wrapper detection +## Wrapper detection (diagnostic) -While reading the source, judge whether what you're looking at is +While reading the source, judge whether the source file is **likely a thin wrapper file** rather than the actual skill content. Common patterns: @@ -144,18 +220,18 @@ content. Common patterns: If you judge this to be likely a wrapper, set `likely_wrapper: true` in the frontmatter and record where the actual content -appears to live in `wrapper_points_to`. Add a body section 9 -("Wrapper observation") explaining the signals and your -confidence. - -This is your judgment — there's no mechanical signal that's -definitive. If you're not sure, err toward `likely_wrapper: -false` and proceed; the operator will catch obvious wrappers in -review of your output. The operator surfaces a `true` finding to -the user, who decides whether to re-vendor the referenced -content, treat the wrapper as the skill, or cancel and provide a -different source. **Do NOT try to follow the pointer yourself or -re-vendor** — that's the operator's job after the user gate. +appears to live in `wrapper_points_to`. The wrapper finding is +**diagnostic only** — it doesn't drive behavior; the +`optimization_target_candidates` list (and the operator's picker) +is what determines what gets optimized. You DO follow the pointer +and vendor the underlying file into `${VENDORED_PATH}` per the +"Vendoring referenced content" section above, and record both the +wrapper and the underlying as candidates. + +If you're not sure whether it's a wrapper, err toward +`likely_wrapper: false`. If there's only one viable optimization +target (the source SKILL.md itself), the operator will pick it +automatically with no surface to the user. ## Return summary @@ -164,9 +240,11 @@ After writing `${OUTPUT_PATH}`, return a brief summary: - Classification - 3-5 key responsibilities (the ones step 3 will design probes for) - Any PR submission notes captured -- **Wrapper finding** (if any): `likely_wrapper: true` + one-line - rationale + `wrapper_points_to` value. Operator needs this to - surface the user gate. +- **Optimization target candidates** — list each `path` and a + one-line rationale; mark your recommended pick. Operator uses + this to decide whether to surface a picker to the user. +- **Wrapper finding** (only if `likely_wrapper: true`) — one-line + rationale + `wrapper_points_to` value. Keep it under 200 words — the operator session reads this for the handoff message; full detail is in the report. diff --git a/skills/investigate-submissions/SKILL.md b/skills/investigate-submissions/SKILL.md index 29167dc..1a7a437 100644 --- a/skills/investigate-submissions/SKILL.md +++ b/skills/investigate-submissions/SKILL.md @@ -75,12 +75,21 @@ in this session** — dispatch the subagent via the `Agent` tool. Render the prompt template at [`agents/research-submissions.md`](./agents/research-submissions.md) inline by substituting `${UPSTREAM_REPO}` from -`01-functionality.md`'s `skill_source` field, `${OUTPUT_PATH}`, -and `${OPERATOR_DIRECTIVES}`, then pass the rendered text as the +`01-functionality.md`'s `optimization_target_source` field (NOT +`skill_source` — the optimization target may live in a different +upstream repo than the source URL the user pointed at, e.g. when +a wrapper SKILL.md in repo A fetches its rules from `command.md` +in repo B; the PR goes to repo B), `${OUTPUT_PATH}`, and +`${OPERATOR_DIRECTIVES}`, then pass the rendered text as the Agent tool's `prompt` parameter (per [`subagent-dispatch.md`](../shared/subagent-dispatch.md)'s Dispatch protocol). +Derive `${UPSTREAM_REPO}` from `optimization_target_source` by +extracting `/` from the URL. If the URL form is +unfamiliar (e.g., GitLab / Bitbucket), surface to the user — see +Edge cases below. + The subagent sees: the upstream repo via `gh` CLI (PR list — both merged and closed-without-merge for shape patterns and rejection signals, repo-file API, `CONTRIBUTING.md`, license file, existing diff --git a/skills/shared/workflow.md b/skills/shared/workflow.md index 89958c3..bf09476 100644 --- a/skills/shared/workflow.md +++ b/skills/shared/workflow.md @@ -105,19 +105,23 @@ skill-evals// # TRACKED — eval suites + probes (reusable a suite.yml # step 4 generates .skill-optimizer// # GITIGNORED — heavy ephemeral artifacts - vendored-skill/ # step 1 vendors (upstream OR local) — the wrapper / ship target - vendored-underlying/ # only if step 1's wrapper-redirect gate fired and user chose re-vendor + vendored-skill/ # step 1 vendors source + all externally-referenced files (single flat dir) bench-results// # step 6 raw output (suite-result.json, trace.jsonl, findings.txt) autopilot-/ # step 10 per-run scratch journal.md # appended-to during the run ``` -When `vendored-underlying/` exists, **the testing target is still -`vendored-skill/`** (that's what ships when the user installs the -skill). The underlying directory carries the actual rules for -analysis/improvement reference; the optimizer/validator may patch -files in either folder depending on where the live rule content -sits. +`vendored-skill/` is a **single self-contained dir**: step 1's +researcher vendors the source plus any external URLs the source +references at runtime, so the test bench reproduces the install +behavior without live network. The specific file the chain +**optimizes and submits upstream** is named explicitly by +`01-functionality.md`'s `optimization_target` frontmatter field — +either the source `SKILL.md` (non-wrapper) or an underlying file +the source points at (wrapper). Step 2 reads +`optimization_target_source` (the upstream URL of that file) to +decide which repo to research for PR conventions; this may differ +from `skill_source`. `` is `--` for upstream skills, or `` for local skills. The same `` is reused diff --git a/skills/validate/SKILL.md b/skills/validate/SKILL.md index af1bb54..30f69c7 100644 --- a/skills/validate/SKILL.md +++ b/skills/validate/SKILL.md @@ -45,12 +45,17 @@ One or two artifacts at `docs/skill-optimizer//`: [`agents/validator.md`](./agents/validator.md). 2. **`docs/skill-optimizer//improved-skill/`** — improved skill content, only - materialized when `verdict: approve`. Mirrors the source's - directory structure with the optimizer's diff applied. **The - original is never modified**: `.skill-optimizer//vendored-skill/` (the canonical - input regardless of upstream/local) stays frozen, and the - user's original local file (if any) is untouched. Git tracks - `docs/skill-optimizer//improved-skill/` history across iterations. + materialized when `verdict: approve`. Mirrors the + `vendored-skill/` directory structure exactly, with the + optimizer's diff applied **only to the + `optimization_target` file** (from + `01-functionality.md`'s frontmatter). All other files in the + tree are copied verbatim from `vendored-skill/` (or the prior + `improved-skill/` if accumulated). **The original is never + modified**: `.skill-optimizer//vendored-skill/` stays + frozen, and the user's original local file (if any) is + untouched. Git tracks `docs/skill-optimizer//improved-skill/` + history across iterations. ## Workflow @@ -68,7 +73,9 @@ Three checks: re-ran step 1 mid-chain), surface to the user — the BEFORE must match what the optimizer read. The fix is re-invoking step 8 against the new state. -3. `01-functionality.md` must exist. +3. `01-functionality.md` must exist with `optimization_target` + set. The validator compares BEFORE and AFTER **at that + specific file path** — not the whole skill tree. If `pr_submission_intent: true`, `02-submissions.md` should exist; if missing, the external check can't run. Ask whether to @@ -126,8 +133,10 @@ judging the new proposal on its own merits. ### (d) Handle the verdict + materialize on approve **`verdict: approve`:** materialize `docs/skill-optimizer//improved-skill/` by copying -the source's directory structure and applying the optimizer's -diff. **Do NOT modify the source.** Commit: +the source's directory structure (from `.skill-optimizer//vendored-skill/` +or the existing `improved-skill/`) and applying the optimizer's +diff **only to the `optimization_target` file**. All other files +are copied verbatim. **Do NOT modify the source.** Commit: ```bash git add docs/skill-optimizer// diff --git a/skills/validate/agents/validator.md b/skills/validate/agents/validator.md index 289d135..75af250 100644 --- a/skills/validate/agents/validator.md +++ b/skills/validate/agents/validator.md @@ -32,6 +32,12 @@ without your own prior verdicts. proposal applied to a copy of BEFORE (the operator session prepares this; you read it but it's NOT the canonical `docs/skill-optimizer//improved-skill/` yet) +- `${OPTIMIZATION_TARGET}` — relative path within both + `${SKILL_BEFORE_PATH}` and `${SKILL_AFTER_PATH}` of the file + whose diff you're judging. From `01-functionality.md`'s + frontmatter. Other files in the trees should be identical + between BEFORE and AFTER — if they differ, the optimizer + exceeded its scope; flag that. - `${FUNCTIONALITY_PATH}` — `01-functionality.md` (for the internal consistency check: does the change make sense given the skill's stated responsibilities?)