Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
121 commits
Select commit Hold shift + click to select a range
bbb6158
docs: spec for skill-optimizer v1.4 — 7-skill decomposition
Zhaiyuqing2003 May 19, 2026
a01d0cc
docs(v1.4-spec): PR-or-not decision moves to step 1 (was: end of step 7)
Zhaiyuqing2003 May 19, 2026
d517348
docs: implementation plan for skill-optimizer v1.4
Zhaiyuqing2003 May 19, 2026
29f72a2
docs(v1.4-plan): mark Phases B/C/D interactive, assign skill-creator …
Zhaiyuqing2003 May 19, 2026
a7dce70
feat(v1.4): create 7 skill directory shells with frontmatter
Zhaiyuqing2003 May 19, 2026
7c0541a
chore(v1.4): create subagents/ and references/ dirs (populated in Pha…
Zhaiyuqing2003 May 19, 2026
2774984
feat(v1.4): seed references/recipes.md from v1.3 lessons.md
Zhaiyuqing2003 May 19, 2026
ebb9eab
docs(v1.4): clean up skills/ layout — drop references/, scope subagen…
Zhaiyuqing2003 May 19, 2026
47124f4
feat(skill-optimizer-investigate-functionality): SKILL.md body — clas…
Zhaiyuqing2003 May 19, 2026
0009066
docs(v1.4): add iteration patterns, 8th auto-pilot skill, subagent fo…
Zhaiyuqing2003 May 19, 2026
b349056
feat(skill-optimizer-investigate-functionality): SKILL.md aligned wit…
Zhaiyuqing2003 May 19, 2026
c82776f
feat(v1.4-shared): factor iteration mechanics into shared protocol doc
Zhaiyuqing2003 May 19, 2026
d18ac79
docs(v1.4-shared): tighten iteration-protocol against authoring philo…
Zhaiyuqing2003 May 19, 2026
552e7fa
docs(skill-optimizer-investigate-functionality): disambiguate step nu…
Zhaiyuqing2003 May 19, 2026
2fd60b1
feat(skill-optimizer-investigate-test-case): SKILL.md + spec: source …
Zhaiyuqing2003 May 19, 2026
ab3506a
feat(skill-optimizer-investigate-submissions): SKILL.md body
Zhaiyuqing2003 May 19, 2026
fb1e269
feat(v1.4-shared): add "Re-run authorization" section to iteration pr…
Zhaiyuqing2003 May 20, 2026
97b67a4
fix(skill-optimizer-investigate-submissions): three corrections per r…
Zhaiyuqing2003 May 20, 2026
0df252f
feat(v1.4): introduce maintenance-step pattern for B2 (and future B4)
Zhaiyuqing2003 May 20, 2026
eb6c53e
feat(skill-optimizer-write-tests): SKILL.md body
Zhaiyuqing2003 May 20, 2026
3fc554b
fix(skill-optimizer-write-tests): correct the smoke-check section — n…
Zhaiyuqing2003 May 20, 2026
a7cac46
feat(skill-optimizer-run-bench): SKILL.md body
Zhaiyuqing2003 May 20, 2026
0b2d1c2
feat(v1.4-chain): case-rename convention + defer partial re-bench
Zhaiyuqing2003 May 20, 2026
59fe82e
feat(v1.4-iteration): git-native + filesystem-as-state redesign
Zhaiyuqing2003 May 20, 2026
6e65c18
feat(v1.4-chain): rewrite B1-B5 SKILL.md for git-native model
Zhaiyuqing2003 May 20, 2026
a31d720
docs(iteration-protocol): clarify directive channel + step 7 carve-out
Zhaiyuqing2003 May 21, 2026
56867f8
docs(iteration-protocol): prescriptive frontmatter discipline rule
Zhaiyuqing2003 May 21, 2026
0dbb998
feat(skill-optimizer-analyze-result + improve-skill): SKILL.md bodies
Zhaiyuqing2003 May 21, 2026
f68339d
fix(skill-optimizer-analyze-result): move body template to subagent p…
Zhaiyuqing2003 May 21, 2026
d3d91b1
feat(v1.4-chain): B7 scope reduction — drop PR draft + preserve original
Zhaiyuqing2003 May 21, 2026
1a50a9f
fix(skill-optimizer-investigate-test-case): three review fixes
Zhaiyuqing2003 May 21, 2026
0528d8c
fix(skill-optimizer-investigate-submissions): drop over-specific PR c…
Zhaiyuqing2003 May 21, 2026
08a40c0
feat(v1.4-chain): split B7 into improve-skill + validate-improvement
Zhaiyuqing2003 May 21, 2026
e737fc1
refactor(v1.4-chain): verbosity sweep + split shared docs (-930 lines)
Zhaiyuqing2003 May 21, 2026
cfc79f2
refactor(v1.4-chain): closeness sweep — inline edges, drop iteration …
Zhaiyuqing2003 May 21, 2026
169abca
refactor(docs): move chain-specific docs out of top-level docs/
Zhaiyuqing2003 May 21, 2026
165b053
refactor: nuke legacy skills/skill-optimizer/ + co-locate workbench.md
Zhaiyuqing2003 May 21, 2026
6617c70
refactor(v1.4-chain): trim 3 skill names + drop step-range disambigua…
Zhaiyuqing2003 May 21, 2026
096237a
feat(v1.4-chain): add validate-tests step (B5) + renumber downstream
Zhaiyuqing2003 May 21, 2026
f862528
feat(v1.4-subagents): write all 8 subagent prompt templates
Zhaiyuqing2003 May 21, 2026
422c831
refactor(v1.4-chain): vendor-always — single canonical input regardle…
Zhaiyuqing2003 May 21, 2026
1a19d54
feat(b3-submissions): detect entry-file pattern + map linked consumers
Zhaiyuqing2003 May 21, 2026
f9d2737
feat(b1-functionality): detect entry-file pattern in vendor step
Zhaiyuqing2003 May 21, 2026
11c8c0c
fix(subagent-prompts): drop invisible internal-step references
Zhaiyuqing2003 May 22, 2026
f8a7f7b
refactor(wrapper-detection): subagent judgment, not operator mechanic…
Zhaiyuqing2003 May 22, 2026
efde54f
docs(v1.4): swap chain steps 2/3 — submissions before design-tests
Zhaiyuqing2003 May 22, 2026
6866d1e
docs(v1.4): replace internal B-number shorthand with step numbers
Zhaiyuqing2003 May 22, 2026
2c813e8
refactor(wrapper-flow): resolve wrapper-vs-target at step 1, not step 2
Zhaiyuqing2003 May 22, 2026
fcde748
fix(tests): relink smoke distribution test to v1.4 chain skills
Zhaiyuqing2003 May 22, 2026
640524b
fix(metadata): relink plugin manifests + docs to v1.4 chain skills
Zhaiyuqing2003 May 22, 2026
5d652b0
docs(dispatch): require inline-rendered subagent prompts (Pattern A)
Zhaiyuqing2003 May 25, 2026
a51d9b0
fix(workbench): unblock Linux Docker bind-mount permissions
Zhaiyuqing2003 May 7, 2026
8e9e39b
fix(docker): drop stale COPY of removed scripts/ directory
Zhaiyuqing2003 May 25, 2026
d16df9c
fix(test-writer): pin probe layout + smoke.mjs import path explicitly
Zhaiyuqing2003 May 25, 2026
46ccdf1
fix(write-tests): present minimal+full probe-set options at step 4 gate
Zhaiyuqing2003 May 25, 2026
c5dd466
feat(step-1): recommend a dedicated branch before chain starts
Zhaiyuqing2003 May 25, 2026
786c5fe
feat(step-1): create chain-progress TodoWrite at start
Zhaiyuqing2003 May 25, 2026
d5f3e2b
refactor(layout): split chain state across three locations
Zhaiyuqing2003 May 25, 2026
fe79756
feat(philosophy): cross-vendor skill-design synthesis for analyzer/op…
Zhaiyuqing2003 May 25, 2026
1b2da28
refactor(philosophy): drop human-side provenance from agent-facing doc
Zhaiyuqing2003 May 25, 2026
1f79681
docs(philosophy): expand to 4-vendor survey + synthesis rationale
Zhaiyuqing2003 May 25, 2026
7f15911
refactor(naming): drop skill-optimizer- prefix; co-locate subagents u…
Zhaiyuqing2003 May 25, 2026
336c27c
fix(links): repair display text mangled by unescaped sed pattern
Zhaiyuqing2003 May 25, 2026
64e3970
docs(spec): multi-agent ACP workbench design
Zhaiyuqing2003 May 25, 2026
63d94c9
docs(plan): multi-agent ACP workbench implementation plan
Zhaiyuqing2003 May 25, 2026
64cf948
feat(acp): add @agentclientprotocol/sdk dependency
Zhaiyuqing2003 May 25, 2026
c7f93bc
chore(test): wire ACP smoke test into npm test script
Zhaiyuqing2003 May 25, 2026
0260339
feat(acp): docker exec stdio bridge implementing Stream
Zhaiyuqing2003 May 25, 2026
ad388b4
chore(test): wire ACP transport test into npm test
Zhaiyuqing2003 May 25, 2026
9aeef03
fix(acp): make transport outgoing close() return a Promise
Zhaiyuqing2003 May 25, 2026
7be3341
feat(acp): client wrapper with auto-approve permission handler
Zhaiyuqing2003 May 25, 2026
c1c965d
chore(test): wire ACP client test into npm test
Zhaiyuqing2003 May 25, 2026
2732c60
feat(agents): registry types + resolver (entries pending in task 5)
Zhaiyuqing2003 May 25, 2026
7d2c79d
feat(agents): populate 5-agent registry (claude/codex/gemini/opencode…
Zhaiyuqing2003 May 25, 2026
35f2060
chore(test): wire registry test into npm test
Zhaiyuqing2003 May 25, 2026
13bef0b
feat(docker): new image with all 5 agents pre-baked
Zhaiyuqing2003 May 25, 2026
ec1520e
feat(acp): subscription-first auth resolution with fail-loud env fall…
Zhaiyuqing2003 May 25, 2026
4516043
chore(test): wire ACP auth test into npm test
Zhaiyuqing2003 May 25, 2026
196c347
fix(acp): restore .claude subdir in registry; fix auth test to match
Zhaiyuqing2003 May 25, 2026
b21fc5e
feat(acp): skill-deploy module mounts to agent's native skill path
Zhaiyuqing2003 May 25, 2026
6cddae5
chore(test): wire ACP skill-deploy test into npm test
Zhaiyuqing2003 May 25, 2026
77e1c0d
fix(docker): update stale Dockerfile path refs after Task 6 rename
Zhaiyuqing2003 May 25, 2026
301fd58
feat(acp): per-agent MCP config writer (claude/codex/gemini/opencode)
Zhaiyuqing2003 May 25, 2026
943cdc5
chore(test): wire MCP config writer test into npm test
Zhaiyuqing2003 May 25, 2026
2417bfb
docs(mcp): pick option (c) for pi-acp MCP — defer to v1.1
Zhaiyuqing2003 May 25, 2026
62ce6ad
feat(parse-trace): ACP trace helpers (iterMessages/iterToolCalls/comp…
Zhaiyuqing2003 May 25, 2026
41b62e6
chore(test): wire parse-trace test into npm test
Zhaiyuqing2003 May 25, 2026
e7d36b0
feat(acp): trace recorder writes raw ACP messages with header+endedAt
Zhaiyuqing2003 May 25, 2026
f420a60
chore(test): wire ACP trace recorder test into npm test
Zhaiyuqing2003 May 25, 2026
3383084
feat(schema): require agent: field in case.yml, fail loud on missing/…
Zhaiyuqing2003 May 25, 2026
15a12e4
chore(test): wire case-loader agent-required test into npm test
Zhaiyuqing2003 May 25, 2026
f763840
feat(schema): suite.yml uses runs: matrix; legacy models: rejected
Zhaiyuqing2003 May 25, 2026
a100661
chore(test): wire suite-loader runs test into npm test
Zhaiyuqing2003 May 25, 2026
61e8fa3
refactor(metrics): thin wrapper over parse-trace; drop cost field
Zhaiyuqing2003 May 25, 2026
139e85a
refactor(metrics): delegate to iterToolCalls; single trace read in bu…
Zhaiyuqing2003 May 25, 2026
65cc4b4
refactor(trace): drop WorkbenchTraceEntry; trace.jsonl is raw ACP
Zhaiyuqing2003 May 25, 2026
0053d2d
feat(docker-runner): add runOneAcpTrial helper for host-side ACP orch…
Zhaiyuqing2003 May 25, 2026
4c6d8c7
fix(acp): use SDK PROTOCOL_VERSION constant (integer) for initialize
Zhaiyuqing2003 May 25, 2026
64f99a7
feat(docker-runner): host-side ACP orchestration per trial
Zhaiyuqing2003 May 25, 2026
7b59fb9
fix(docker-runner): restore cache env flags on per-trial container
Zhaiyuqing2003 May 25, 2026
c55ef48
refactor(container-runner): drop --agent mode and pi-agent.ts (ACP ho…
Zhaiyuqing2003 May 25, 2026
439bb66
feat(trials): aggregate tokens and durationMs across trials
Zhaiyuqing2003 May 25, 2026
20baebe
chore(test): wire trials-aggregate test into npm test
Zhaiyuqing2003 May 25, 2026
e567817
feat(run-*): replace models matrix with runs (agent + model) matrix
Zhaiyuqing2003 May 25, 2026
62c786e
chore(examples): migrate pdf+mcp suites to runs: schema; remove stale…
Zhaiyuqing2003 May 25, 2026
9faba1e
test(smoke): per-agent end-to-end smoke probes
Zhaiyuqing2003 May 25, 2026
968c0ef
fix(model-refs): only pi-acp requires openrouter/ prefix; native agen…
Zhaiyuqing2003 May 25, 2026
7baf73d
fix(acp): chmod 644 on staged subscription auth files
Zhaiyuqing2003 May 25, 2026
e68d03b
test(regression): pi-acp pdf suite baseline scaffold + tolerance check
Zhaiyuqing2003 May 25, 2026
0dca012
docs(shared): ACP trace format reference for chain analyzer
Zhaiyuqing2003 May 25, 2026
a05caaa
feat(chain): update test-writer, analyzer, run-bench for ACP trace fo…
Zhaiyuqing2003 May 25, 2026
eb92cc5
docs: refresh CLAUDE.md, CONTRIBUTING.md, README.md for multi-agent ACP
Zhaiyuqing2003 May 25, 2026
5e567f7
fix(acp): subscription auth — extract OAuth token, stage writable hom…
Zhaiyuqing2003 May 25, 2026
f0fb4aa
feat(acp): apply case model via unstable_setSessionModel after newSes…
Zhaiyuqing2003 May 26, 2026
194c8f7
fix(agents): add OPENROUTER_API_KEY to pi-acp requiresEnv; drop stray…
Zhaiyuqing2003 May 26, 2026
9e10006
docs(spec): skill-optimizer autopilot — step-10 chain driver design
Zhaiyuqing2003 May 27, 2026
c897c94
feat(autopilot): step-10 chain driver SKILL + agent templates
Zhaiyuqing2003 May 27, 2026
f41710f
feat: workbench reference doc refresh + tmp cleanup + wrapper/branch UX
Zhaiyuqing2003 May 27, 2026
a37b75f
fix(autopilot): auto-handle branch + add CWD discipline; remove stale…
Zhaiyuqing2003 May 27, 2026
69b414c
feat(autopilot): empirical-verification re-bench + suite skillUnderTe…
Zhaiyuqing2003 May 27, 2026
c453b1f
feat(chain): unify wrapper handling — single vendored-skill, optimiza…
Zhaiyuqing2003 May 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,16 @@
"name": "Fast"
},
"skills": [
"./skills/skill-optimizer"
"./skills/investigate-functionality",
"./skills/investigate-submissions",
"./skills/design-tests",
"./skills/write-tests",
"./skills/validate-tests",
"./skills/run-bench",
"./skills/analyze",
"./skills/improve",
"./skills/validate",
"./skills/autopilot"
]
}
]
Expand Down
16 changes: 13 additions & 3 deletions .codex/INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,10 +22,20 @@ codex plugin marketplace add fastxyz/skill-optimizer --ref main

## Skill-Only Install

Install the canonical skill with the open skills CLI:
Install the 9 chain skills with the open skills CLI:

```bash
npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a codex -y
npx skills add fastxyz/skill-optimizer \
--skill investigate-functionality \
--skill investigate-submissions \
--skill design-tests \
--skill write-tests \
--skill validate-tests \
--skill run-bench \
--skill analyze \
--skill improve \
--skill validate \
-a codex -y
```

Restart Codex if the skill does not appear immediately.
Restart Codex if the skills do not appear immediately.
16 changes: 13 additions & 3 deletions .cursor/INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,14 +2,24 @@

## Skill install

Install the skill into Cursor's project or global skill directory through the open skills CLI:
Install the 9 chain skills into Cursor's project or global skill directory through the open skills CLI:

```bash
npx skills add fastxyz/skill-optimizer --skill skill-optimizer -a cursor -y
npx skills add fastxyz/skill-optimizer \
--skill investigate-functionality \
--skill investigate-submissions \
--skill design-tests \
--skill write-tests \
--skill validate-tests \
--skill run-bench \
--skill analyze \
--skill improve \
--skill validate \
-a cursor -y
```

Cursor can also import remote skills from GitHub in Settings -> Rules -> Project Rules -> Add Rule -> Remote Rule (Github).

## Plugin metadata

This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The canonical skill remains `skills/skill-optimizer/SKILL.md`.
This repository includes `.cursor-plugin/plugin.json` for Cursor-compatible plugin metadata. The skills live at `skills/<chain-skill>/`; the plugin manifest exposes all 9 as a chain that triggers based on the user's intent (investigate, design tests, run bench, analyze, improve, validate).
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,7 @@ temp/
docs/superpowers/
docs/plans/
docs/specs/
.superpowers/

# Skill-optimizer generated artifacts
.skill-optimizer/
Expand Down
4 changes: 2 additions & 2 deletions .opencode/INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,9 +8,9 @@ Add the plugin to `opencode.json` at user or project scope:
}
```

Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can load `skill-optimizer`.
Restart OpenCode. The plugin registers the repository `skills/` directory so the native `skill` tool can discover the 9 chain skills (`investigate-functionality`, `investigate-submissions`, `design-tests`, `write-tests`, `validate-tests`, `run-bench`, `analyze`, `improve`, `validate`).

Verify with the skill tool by listing skills or loading `skill-optimizer`.
Verify by listing skills with the skill tool; each chain skill triggers on its own description (investigate, design tests, run bench, analyze, improve, validate).

To pin a version, append a tag or commit ref:

Expand Down
6 changes: 4 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,9 @@ npx tsx src/cli.ts run-suite --help
- `src/cli.ts`: public CLI entrypoint
- `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces
- `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases
- `skills/skill-optimizer/SKILL.md`: canonical distributable Agent Skill
- `skills/<chain-skill>/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate)
- `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench reference; loaded on-demand by chain skills
- `skills/<chain-skill>/agents/`: prompt templates dispatched by chain skills via the Agent tool
- `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support
- `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin
- `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file
Expand All @@ -47,7 +49,7 @@ Keep the README installation section aligned with packaged plugin metadata:
- Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid.
- Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state.
- The agent phase sees only `/work`, not `/case` or `/results`.
- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies.
- Keep plugin metadata pointed at every chain skill under `skills/<chain-skill>/`; do not create divergent skill copies.
- Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`.
- Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies.
- Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials.
Expand Down
25 changes: 16 additions & 9 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

`skill-optimizer` is a Docker workbench for running and grading agent skill eval cases. The current public CLI centers on `run-case` and `run-suite`.

The workbench gives an agent an isolated Docker `/work` directory, captures traces, and grades deterministic local outcomes from files, command logs, generated artifacts, or other workspace state.
The workbench drives a host-side ACP (Agent Client Protocol) client against per-trial Docker containers that run one of five agent CLIs (claude-agent-acp, codex-acp, gemini, opencode, pi-acp). Each trial gets an isolated `/work` directory; raw ACP messages are captured to `trace.jsonl`; graders evaluate deterministic local outcomes from files, command logs, generated artifacts, or other workspace state.

## Key Commands

Expand All @@ -20,10 +20,15 @@ npx tsx src/cli.ts run-suite --help
## Important Files

- `src/cli.ts`: public CLI entrypoint
- `src/workbench/`: workbench case loading, suite loading, Docker runner, Pi agent, graders, and traces
- `docker/workbench-runner.Dockerfile`: generic non-root container image for setup, agent, grade, and cleanup phases
- `skills/skill-optimizer/SKILL.md`: canonical distributable Agent Skill
- `skills/skill-optimizer/references/workbench.md`: detailed workbench schema and usage reference
- `src/workbench/`: case loading, suite loading, host-side Docker runner, graders
- `src/workbench/acp/`: ACP transport, client wrapper, auth, skill deployment, MCP config writers, trace recorder
- `src/workbench/agents/`: 5-agent registry (claude-agent-acp, codex-acp, gemini, opencode, pi-acp) + install snippets
- `src/workbench/parse-trace.ts`: helpers for reading the raw ACP `trace.jsonl` (`iterMessages`, `iterToolCalls`, `computeMetrics`, etc.)
- `docker/skill-optimizer-agent.Dockerfile`: container image with all 5 agent CLIs pre-baked
- `docker/pi-acp-launcher.sh`: wrapper that bridges `SKILL_OPT_PROVIDER_*` env vars into pi-acp's `OPENROUTER_API_KEY`
- `skills/<chain-skill>/`: the v1.4 chain of 9 user-invocable skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate)
- `skills/shared/`: workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, and the workbench schema reference; loaded on-demand by chain skills
- `skills/<chain-skill>/agents/`: prompt templates dispatched by chain skills via the Agent tool
- `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`: cross-agent plugin manifests and install support
- `.agents/plugins/marketplace.json`: Codex repo marketplace entry for the root plugin
- `gemini-extension.json`, `GEMINI.md`: Gemini extension metadata and context file
Expand All @@ -45,12 +50,14 @@ Keep the README installation section aligned with packaged plugin metadata:
## Invariants

- Keep evaluation static: extraction and matching are allowed; do not execute model-produced code outside the Docker workbench as part of evaluation.
- `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override.
- Keep OpenRouter model refs as `openrouter/...`; real model runs require `OPENROUTER_API_KEY`.
- Every `case.yml` must declare `agent:` (one of `claude-agent-acp`, `codex-acp`, `gemini`, `opencode`, `pi-acp`); the loader fails loud on missing/unknown agents.
- `suite.yml` declares its agent+model matrix via `runs: [{ agent, model }]`; legacy `models:` is rejected. `run-suite` uses what `suite.yml` declares; do not add a `--models` override.
- Only `pi-acp` requires `openrouter/...` model refs; native ACP agents pass their own model strings through unchanged.
- Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid.
- Graders are the acceptance contract; evaluate outputs from `/work`, generated artifacts, `answer.json`, `trace.jsonl`, and result state.
- The agent phase sees only `/work`, not `/case` or `/results`.
- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies.
- `trace.jsonl` is raw ACP wire format (JSON-RPC envelopes); use `parse-trace.ts` helpers rather than parsing the file by hand. See [`skills/shared/acp-trace-format.md`](skills/shared/acp-trace-format.md).
- Keep plugin metadata pointed at every chain skill under `skills/<chain-skill>/`; do not create divergent skill copies.
- Codex plugin metadata lives in `.codex-plugin/plugin.json`; the repo marketplace lives in `.agents/plugins/marketplace.json` and points at `./`.
- Provider install docs should link to the same canonical skill/plugin metadata, not separate skill copies.
- Do not commit `.skill-eval/`, `.results/`, `.env`, or credentials.
Expand All @@ -59,6 +66,6 @@ Keep the README installation section aligned with packaged plugin metadata:

- Run `npm run typecheck` after TypeScript changes.
- Run `npm test` before finishing behavior changes.
- For Docker runner or image changes, also run `docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile .`.
- For Docker runner or image changes, also run `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .`.
- For CLI/docs changes, verify `npx tsx src/cli.ts --help` if touched docs mention CLI behavior.
- For plugin/package metadata changes, run `npx tsx tests/smoke-skill-distribution.ts` and verify `npm pack --dry-run --json` includes required plugin files without result/cache directories.
29 changes: 19 additions & 10 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Contributing to skill-optimizer

Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills. Changes should preserve deterministic grading, isolated agent workspaces, and the canonical `skills/skill-optimizer/SKILL.md` distribution path.
Thanks for contributing! This project is a small, opinionated Docker workbench for evaluating agent skills, plus a 9-step chain of Agent Skills (`skills/<chain-skill>/`) that orchestrates investigation, test design, bench runs, analysis, and improvement of a target skill, plus an `autopilot` driver that walks the chain end-to-end. Changes should preserve deterministic grading, isolated agent workspaces, and the chain skills' distribution paths.

## Installing The Skill

Expand All @@ -22,10 +22,13 @@ All three commands must pass before opening a PR when code changes are involved.
## Project layout

- `src/cli.ts` — public CLI entry point for `run-case` and `run-suite`.
- `src/workbench/` — case/suite loading, Docker runner, Pi agent wiring, graders, traces, metrics, MCP support, and trial aggregation.
- `docker/workbench-runner.Dockerfile` — non-root container image for setup, agent, grade, and cleanup phases.
- `skills/skill-optimizer/SKILL.md` — canonical distributable Agent Skill.
- `skills/skill-optimizer/references/workbench.md` — detailed workbench schema and authoring reference.
- `src/workbench/` — case/suite loading, host-side Docker runner, graders, metrics, MCP support, trial aggregation.
- `src/workbench/acp/` — ACP transport, client wrapper, auth resolution, skill deployment, per-agent MCP config writer, trace recorder.
- `src/workbench/agents/` — 5-agent registry + Dockerfile install snippet generator.
- `docker/skill-optimizer-agent.Dockerfile` — container image with all 5 agent CLIs pre-baked.
- `skills/<chain-skill>/` — the 9-step chain of user-invocable Agent Skills (investigate-functionality, investigate-submissions, design-tests, write-tests, validate-tests, run-bench, analyze, improve, validate) plus the `autopilot` driver.
- `skills/shared/` — workflow overview, iteration protocol, subagent-dispatch rules, frontmatter discipline, workbench schema reference; loaded on-demand by chain skills.
- `skills/<chain-skill>/agents/` — prompt templates dispatched by chain skills via the Agent tool.
- `examples/workbench/` — packaged example suites.
- `.claude-plugin/`, `.codex-plugin/`, `.cursor-plugin/`, `.opencode/`, `.agents/plugins/marketplace.json`, `gemini-extension.json`, `GEMINI.md` — cross-agent plugin and extension metadata.
- `tests/` — hand-rolled smoke tests (`tsx tests/smoke-*.ts`).
Expand All @@ -42,23 +45,29 @@ All three commands must pass before opening a PR when code changes are involved.
## Workbench invariants

- Keep evaluation static: extraction and matching are allowed; do not execute model-produced code outside the Docker workbench as part of evaluation.
- Use only `openrouter/...` model refs; real model runs require `OPENROUTER_API_KEY`.
- `run-suite` uses models from `suite.yml`; do not add a `run-suite --models` override.
- Every `case.yml` must declare `agent:` (one of `claude-agent-acp`, `codex-acp`, `gemini`, `opencode`, `pi-acp`); the loader fails loud on missing/unknown.
- `suite.yml` declares its `runs: [{ agent, model }]` matrix; legacy `models:` is rejected. `run-suite` uses what `suite.yml` declares; do not add a `--models` override.
- Only `pi-acp` requires `openrouter/...` model refs; native ACP agents pass their own model strings through unchanged.
- Cases use `graders: [{ name, command }]`; legacy `check:` and `artifacts:` are invalid.
- The agent phase sees only `/work`, not `/case`, `/results`, graders, hidden answers, or hidden metadata.
- Keep plugin metadata pointed at the canonical `skills/skill-optimizer/SKILL.md`; do not create divergent skill copies.
- `trace.jsonl` is raw ACP wire format (JSON-RPC envelopes); use `src/workbench/parse-trace.ts` helpers (`iterMessages`, `iterToolCalls`, `computeMetrics`) rather than parsing the file by hand.
- Keep plugin metadata pointed at every chain skill under `skills/<chain-skill>/`; do not create divergent skill copies.

## Testing guidance

- Run `npm run typecheck` after TypeScript changes.
- Run `npm test` before finishing behavior changes.
- For Docker runner or image changes, also run `docker build -t skill-optimizer-workbench:local -f docker/workbench-runner.Dockerfile .`.
- For Docker runner or image changes, also run `docker build -t skill-optimizer-agent:local -f docker/skill-optimizer-agent.Dockerfile .`.
- For CLI/docs changes, verify `npx tsx src/cli.ts --help` if touched docs mention CLI behavior.
- For plugin/package metadata changes, run `npx tsx tests/smoke-skill-distribution.ts` and verify `npm pack --dry-run --json` includes required plugin files without result/cache directories.

## Agent capabilities and limitations (v1)

MCP server support is available for `claude-agent-acp`, `codex-acp`, `gemini`, and `opencode` agents. `pi-acp` MCP support is deferred to v1.1 pending documentation of pi-acp's native MCP config path.

## Adding workbench capabilities

Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/skill-optimizer/references/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior.
Keep new capabilities small and deterministic. Add validation in the relevant loader, tests in `tests/smoke-workbench-*.ts`, and docs in `skills/shared/workbench.md` or `docs/workbench.md` when users need to author new YAML fields or understand new runtime behavior.

## Commit style

Expand Down
4 changes: 2 additions & 2 deletions GEMINI.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
@./AGENTS.md
@./README.md
@./CONTRIBUTING.md
@./skills/skill-optimizer/SKILL.md
@./skills/skill-optimizer/references/workbench.md
@./skills/shared/workflow.md
@./skills/shared/workbench.md
Loading
Loading