Repository navigation
v1.1 - #21
v1.1#21
Conversation
…API formats (#12) * fix(benchmark): strip provider prefix from model ID for direct API formats When `benchmark.format` is `anthropic` or `openai` with a direct `baseUrl`, the model ID (e.g. `anthropic/claude-sonnet-4-6`) includes a provider prefix that OpenRouter expects but direct APIs reject. The Anthropic API returns: 404 {"error":{"message":"model: anthropic/claude-sonnet-4-6"}} Root cause: `createLLMClient` in `src/benchmark/llm/index.ts` passes `model.id` verbatim to the API request body. The config validator requires the prefix (`anthropic/...`) but the direct API expects the bare name (`claude-sonnet-4-6`). Fix: added `stripProviderPrefix()` that removes everything up to the first `/` when format is `anthropic` or `openai`. Applied in all three call paths (chat, chatWithTools, chatAgentLoop). The `pi` format is unaffected. Tests: 4 new cases in `smoke-llm.ts`: - anthropic: strips prefix in chat - anthropic: strips prefix in chatWithTools - anthropic: no-op when prefix absent - openai: strips prefix in chat * fix: use prefix-aware stripping instead of indexOf in stripProviderPrefix indexOf('/') only strips the first slash segment, so openrouter/anthropic/ model IDs would become anthropic/claude-... (still wrong for direct Anthropic API). validate.ts accepts openrouter/ prefixes as valid for any format, making this reachable by documented usage. New approach: only strip the direct-provider prefixes (anthropic/, openai/). openrouter/ IDs are left intact — they signal format:'pi' usage, so encountering one in a direct-API context is a misconfiguration that should fail visibly at the API level rather than being silently misrouted. Also adds a test confirming openrouter/-prefixed IDs are passed through unchanged. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: fix env var restore and add chatAgentLoop prefix-stripping coverage - Restore OPENAI_API_KEY to its original value (or delete it if absent) in a finally block so CI environments don't lose a pre-existing key - Add chatAgentLoop test to cover the third call path for prefix stripping, matching the chat and chatWithTools cases Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: add missing models field to openaiConfig test fixture LLMConfig.models is required; the inline config in the OpenAI prefix-stripping test was missing it, causing a structural type error in editors. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Add current OpenRouter model presets * fix: align optimize model ID to dot notation used by benchmark presets commonOptimize.model was left as 'openrouter/anthropic/claude-sonnet-4-6' (dashes) while the rest of this PR uses 'openrouter/anthropic/claude-sonnet-4.6' (dots). A mismatched model ID would cause `skill-optimizer optimize` to call a different or invalid OpenRouter ref than the benchmark uses. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: use hyphen notation in model IDs to match codebase convention validate.ts flags digit.digit patterns in model IDs as bad format (warning: "OpenRouter expects hyphens") and fix.ts auto-corrects them. All new model IDs in this PR used dot notation (e.g. claude-sonnet-4.6) which would trigger warnings on every init-generated config. Convert all version segments in model ID strings to hyphens to match the existing convention. Display names (name/label fields) keep dots for readability (e.g. 'Claude Sonnet 4.6'). Also reverts the previous incorrect fix to commonOptimize.model — the dash format was already correct. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs(claude): document model ID hyphen convention and display name distinction validate.ts warns on digit.digit patterns in model IDs and fix.ts auto-corrects them. This tripped up PR #10 where new presets used dot notation. Document the rule clearly so it's not re-discovered the hard way. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: enforce hyphen convention for model ID version segments Adds a check to the MODEL_PRESETS block in smoke-init.ts that fails if any preset value contains a digit.digit pattern (e.g. claude-sonnet-4.6). validate.ts already flags these at runtime; this makes CI catch them at commit time so the convention is enforced on every PR without manual review. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * ci: rename workflow to 'Build & Test', add development and staging branches Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * ci: restore job ID to build-test to preserve required check names Renaming the job broke branch protection rules that require build-test (20) and build-test (22). Keep the job ID as-is; the display name still shows 'Build & Test (node X)' in the GitHub UI. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: dmn <damian.ovidiu27@gmail.com>
* Add Codex auth support
* Fix Codex auth CI smoke test
* Fix codex auth in optimize task generation
* Fix Codex MCP schema handling
* fix(auth): prioritize tokens.access_token over static OPENAI_API_KEY in readCodexApiKey
The browser-login JWT (tokens.access_token) now takes highest priority,
followed by tokens.OPENAI_API_KEY, then the top-level OPENAI_API_KEY.
Previously, a stale static key could shadow a valid browser-login token,
defeating the purpose of Codex auth. The expiry check on access_token is
preserved and still applied.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): resolve OpenAI credential per-call to avoid stale Codex token
Move resolvedOpenAICredential and useCodexBridgeForOpenAI out of the
createLLMClient closure and into each method body, so long-running
benchmarks pick up refreshed Codex tokens instead of a stale snapshot.
Also pass apiKeyOverride (pre-resolved key) to Pi bridge calls instead
of authMode+apiKeyEnv, preventing a second auth.json read on every call.
* test(llm): cover chatWithTools+chatAgentLoop codex bridge; narrow resolveDirectApiKey return type
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(doctor): probe each model individually instead of bailing when first model is non-openrouter
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tests): restore HOME env var with delete-vs-assign to avoid "undefined" string pollution
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use benchmark.authMode as field for codex auth warning instead of benchmark.apiKeyEnv
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(optimizer): reference PiAuthMode instead of duplicating literal union
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: clarify Codex auth paths and format/apiKeyEnv descriptions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: simplify code clarity in Codex auth changes
Drop the redundant useCodexBridgeForOpenAI variable in createLLMClient.
Since resolveOpenAICredential() returns undefined unless format is
'openai', checking openAICredential?.source === 'codex' is sufficient
on its own.
Also regenerate docs/reference/config-schema.md to reflect the
schema.ts description updates from the PR.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(review): add format guard comment, accurate doctor skip message, and auto-mode priority test
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(auth): skip apiKeyEnv lookup when authMode=codex; avoid writing OPENROUTER_API_KEY in openai-format scaffold
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use integer index instead of model.id in error field paths
Replace `for...of` loops over `benchmark.models` with `entries()` variants
so error field paths use the numeric index (e.g. `benchmark.models[0].weight`)
instead of the model ID string, which produced malformed paths like
`benchmark.models[openai/gpt-5.4].weight`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): check all providers in mixed-model configs and validate optimize auth
Fix 2: when benchmark.format is 'pi', models can carry different provider
prefixes (openrouter/*, openai/*, anthropic/*). The old code derived
benchmarkProvider from models[0] and validated only that one credential.
Now all unique provider prefixes are collected and each is checked
separately, so a mixed config like ["openai/gpt-4", "openrouter/…"]
warns when either key is missing.
Fix 3: the optimize section's authMode/apiKeyEnv was never validated for
key presence — only codex provider-compatibility was checked. Now the
optimize credential is resolved and a warning is pushed if no key is
found, parallel to the benchmark check.
Both fixes share a warnMissingApiKey helper that formats the hint and
field path consistently for both benchmark and optimize sections.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(scaffold): drop redundant benchmark.apiKeyEnv from scaffold defaults
benchmark.apiKeyEnv: 'OPENROUTER_API_KEY' was written into every
scaffolded config by both scaffold.ts (commonBenchmark) and
init.ts (commonBenchmark + commonOptimize). The field is redundant:
resolveApiKey already falls back to OPENROUTER_API_KEY for the
openrouter provider via defaultApiKeyEnvForProvider.
With authMode:'auto' now in the PR, the explicit field becomes a
footgun: when a user edits their config to use openai/* models and
forgets to remove apiKeyEnv, resolveApiCredential reads OPENROUTER_API_KEY
from the env (present and non-empty) and returns it as the OpenAI key —
skipping the Codex fallback and sending a wrong key to the OpenAI API.
Fix 4 already removed it from scaffold.ts commonOptimize. This commit
removes it from the three remaining sites.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): restore authMode:codex in bridge calls to preserve Pi routing signal
The Codex bridge was passing apiKeyOverride to chatPi, which caused
resolveApiCredential to short-circuit with source:'override'. The routing
decision in resolvePiModel checks credential.source === 'codex' to switch
to the 'openai-codex' provider (synthesizeOpenAICodexModel, Codex endpoint
at chatgpt.com/backend-api/codex). With source:'override' that branch was
never taken — the browser JWT was being sent to the standard OpenAI API
endpoint instead, causing a 401.
Restore the original approach: pass authMode:'codex' so Pi re-reads
~/.codex/auth.json, returns source:'codex', and routes to openai-codex.
The second disk read is load-bearing for the provider routing decision.
Update the three smoke tests that were asserting the old (incorrect)
apiKeyOverride behavior.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix dot-notation model IDs in README; extend CLAUDE.md ID convention
README used openai/gpt-5.4 (dot notation) in the Codex auth example —
two occurrences. Corrected to openai/gpt-5-4 to match the project's
model ID convention (hyphens in version segments).
CLAUDE.md model ID convention section only documented the openrouter/
format. Extended it to also cover the direct-provider formats
(openai/<model>, anthropic/<model>) that this PR introduces, with the
same hyphen rule and an openai/gpt-5-4 example.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate,models): restore dot-version warning for anthropic/; add OpenAI OPENROUTER_API_KEY guard
Finding 1: validate.ts narrowed the model-id-bad-format check to
openrouter/* only, which stopped flagging anthropic/claude-sonnet-4.6 —
a real error since the Anthropic API expects hyphens. The exemption should
only apply to openai/ IDs, where OpenAI's own API uses dots in some model
slugs (e.g. gpt-4.5). Change the guard to !startsWith('openai/') and
update the message to be provider-agnostic.
Finding 2: resolvePiModel had an Anthropic-specific guard that throws when
apiKeyEnv is OPENROUTER_API_KEY, preventing silent 401s from stale
scaffolded configs. No equivalent existed for the openai provider.
Add a parallel guard so direct openai/* calls with an OpenRouter key
fail fast with a clear error and a migration hint.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(cli,auth,validate): three preflight and credential correctness gaps
cli.ts: run preflight checked only models[0] provider for pi format,
so a mixed-provider config (e.g. openrouter + openai models) would pass
startup and then silently record per-model failures for whichever
provider's key was missing. Now all unique providers are checked, matching
what validate.ts already does across the model list.
auth.ts: readCodexApiKey returned {} immediately on an expired JWT,
making the tokens.OPENAI_API_KEY and top-level OPENAI_API_KEY fallbacks
unreachable when access_token was present-but-expired. The comment's
intent was correct (don't shadow valid JWT with stale static key) but the
inverse was wrong (expired JWT was shadowing valid static key). Now falls
through to the static key branches on JWT expiry.
validate.ts: optimize auth check ran even when optimize.enabled === false,
causing spurious auth warnings for benchmark-only setups. Added the same
enabled !== false guard that neighboring optimize checks already use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add SKILL design spec for skill-optimizer Spec for a multi-file SKILL/ folder that guides AI agents through setting up, benchmarking, and optimizing SDK/CLI/MCP documentation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add implementation plan for SKILL folder 7-task plan covering branch creation, all 5 markdown files with complete content, verification, and PR creation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(skill): add SKILL.md entry point with context detection and routing * feat(skill): add setup reference — prerequisites, init wizard, discovery * feat(skill): add benchmark reference — run, interpret, diagnose, compare * feat(skill): add optimize reference — loop mechanics, safety, troubleshooting * feat(skill): add config reference — schema, models, scope, errors * fix(skill): address Copilot review — config path, safety boundary, minImprovement, gitignore dev docs * docs(skill): update routing concept; add direct-provider auth to prerequisites; fix dot-notation * docs(skill): expand compare section in benchmark reference --------- Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#20) * docs: fix repo URL in CONTRIBUTING.md (bucurdavid → fastxyz) Fixes #18 * feat(prompt): add prompt surface type for skill/template benchmarking Discovers phases, instructions, output formats, and decision points from markdown skill files. Evaluates model responses with content-based criteria (required sections, format patterns, keywords, structural elements) instead of tool-call matching. New files: - src/project/discover-prompt.ts — markdown parser for capabilities - src/benchmark/prompt-evaluator.ts — content-based scorer (weighted) - src/discovery/prompt.ts — file-based discoverer + types - tests/smoke-discovery-prompt.ts — 10 tests - tests/smoke-prompt-evaluator.ts — 15 tests - tests/fixtures/sample-skill.md — 3-phase sample Modified: types.ts, validate.ts, snapshot.ts, runner.ts, evaluator.ts, extractors/index.ts, prompts.ts, cli.ts, init/, CHANGELOG.md, README.md * fix(prompt): guard snapshot surface, fix format inconsistency, remove dead interface, fix stale comment * fix(discovery): anchor frontmatter closing delimiter to line boundary Replace indexOf('---', 3) with a regex that matches --- only at the start of a line, preventing early termination when a YAML value contains ---. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(discover-prompt): clear preamble on first heading, strip frontmatter before section split * fix(discover-prompt): fallback to whole-content instruction extraction for heading-less files Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(discover-prompt): document interface difference between production and standalone paths Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(prompt): wire evaluatePromptResponse into runner, expose section text, fix generate schema Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(prompt): unique temp dir + cleanup guard in test, enforce empty actions in grounding, bypass maxTasks constraint for prompt Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: DRY deduplication across LLM format handlers and Pi task helpers
Extract shared retry logic (isAbortError, isRetryableError, sleep), tool-result
truncation, and agent-loop accumulator helpers into src/benchmark/llm/shared.ts,
eliminating identical code that existed in both openai-format.ts and
anthropic-format.ts.
Extract the Pi completeSimple call sequence (timeout setup, error check, text
block extraction) into src/tasks/pi-simple-complete.ts, used by both
default-pi-generator.ts and default-pi-critic.ts to replace ~40 lines of
near-identical implementation in each file.
Net: -129 lines of duplication, no behavior changes, all tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): consolidate duplicate type definitions
- DiscoveredActionArg (discovery/types) was identical to ActionArgSchema
(actions/types); replaced with a re-export alias to eliminate the copy
- BenchmarkSurface now aliases ActionSurface instead of re-declaring the
same 'sdk' | 'cli' | 'mcp' union literal
- Removed dead SurfaceActionArg alias from benchmark/types (never imported)
- Removed dead SurfaceSnapshotArg alias from project/types and project/index
(exported but never consumed)
- Replaced all inline import('...').PiAuthMode dynamic-import type refs
with proper named import type statements in pi-format, models,
coding-orchestrator, benchmark-agent, default-pi-critic, default-pi-generator
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove unused code found by knip
- Delete src/runtime/pi/benchmark-agent.ts (createReadOnlyBenchmarkModel was dead — exported but never imported)
- Delete src/discovery/index.ts and src/optimizer/mutation/types.ts (barrel re-exports with no consumers)
- Remove requireApiKeyFromEnv from auth.ts (never called internally or imported from outside)
- Remove 5 unused re-export lines from benchmark/extractors/index.ts (extractCodeBlock, extractShellBlock etc. — tests import directly from sub-files)
- Unexport 11 internal-only types/interfaces (MAX_TOOL_RESULT_CHARS, buildConfigFromAnswers, LLMClient, ToolNameAliasCodec, PromptSurface/Options, DoctorOptions, DetectedProject, BenchmarkAdapterRunOptions/Result, FeedbackPackage, RawExtraction, PiSimpleCompleteOptions/Input)
- Remove @mariozechner/pi-agent-core from package.json dependencies (transitive dep, never directly imported)
- Add smoke-actions.ts to the npm test script (17 passing tests were orphaned from the suite)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(deps): break circular dependency between actions/discover and benchmark/config
Move loadCliCommands and loadMcpTools from benchmark/config.ts into a new
actions/loaders.ts file, resolving the cycle:
actions/discover → benchmark/config → project → project/snapshot → actions/discover
benchmark/config.ts re-exports both functions for backward compatibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): replace weak types with precise types throughout
- Replace `Model<any>` with `Model<Api>` and specific API literal types
(`Model<'openai-completions'>`, `Model<'openai-codex-responses'>`) in
models.ts and pi-format.ts
- Replace `(err as any).status` with typed intersection cast
`Error & { status: number }` in anthropic-format.ts and openai-format.ts
- Remove redundant `as unknown[]` cast on `session.state.messages` and
import `AgentMessage` from `@mariozechner/pi-agent-core` in pi-coding.ts;
update `extractLatestAssistantText` and `extractToolActivity` parameter
types accordingly
- Remove unnecessary double casts `(c as unknown as { method: string }).method`
in failure-details.ts and passing-failing-diff.ts — `ExtractedCall` is
`ActionAttempt` which already has `method: string` and `args`
- Remove `as any` from `Type.Unsafe(...)` call in pi-format.ts — the argument
type is compatible with `UnsafeOptions` (SchemaOptions index signature)
- Fix `catch (error: any)` to `catch (error)` in two snapshot files,
using `instanceof Error` guard to access `.message`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: remove unnecessary optional chaining on non-optional reportConfig
reportConfig is always defined (report.config is non-optional per
BenchmarkReport). The ?. on reportConfig before .outputDir was
defensive against a value that TypeScript knows can never be null.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove deprecated and legacy code paths
- Remove toLegacyOptimizeManifest (never called anywhere in src or tests)
- Remove deprecated SdkSurfaceConfig fields: classes, functions, functionReturns, methods
- Remove dead config.sdk?.methods fallback in runner.ts (no config ever sets methods)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove AI slop comments and stubs
Removed or rewrote ~75 comments across 15 files:
- Dropped "// NEW:", "// Transitional alias", numbered step comments (1-14
in runner.ts), and other in-motion migration language
- Removed redundant comments that restated the immediately-following code
- Replaced multi-line comment pairs with single consolidated explanations
where the context was non-obvious and worth preserving
- Kept all comments explaining invariants, non-obvious behavior, or
security rationale (e.g. git injection safety, OpenAI dot-format exemption)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove backwards-compatibility shims — SemVer allows breaking changes
- Remove LEGACY_PROJECT_CONFIG_NAME and skill-benchmark.json detection in load.ts
- Remove legacy-config-name warning from validate.ts (old filename alongside new one)
- Remove E_LEGACY_CONFIG error definition from errors.ts and its docs entry
- Remove DiscoveredActionArg deprecated alias from discovery/types.ts; update all four
discovery modules (cli, mcp, sdk, optique) to import ActionArgSchema directly
- Remove CodeModeConfig, McpModeConfig, ExpectedTool, ToolMatch backward-compat
type aliases from benchmark/types.ts and benchmark/index.ts
- Remove toolMatches dual-field from TaskResult; make actionMatches required
- Remove unnecessaryCalls/hallucinatedCalls alias fields from TaskResult.metrics;
consolidate to unnecessaryActions/hallucinatedActions throughout
- Remove method? alias from ExpectedAction; make name required
- Remove expected_tools? alias from TaskDefinition; make expected_actions required
- Simplify getExpectedActionName and getExpectedActions helper functions
- Remove deprecated extractFromCode function from code-analyzer.ts
- Update all internal consumers (optimizer, evaluator, tasks, benchmark) to use
canonical field names without the ?? fallback patterns
- Update all test files to use canonical field names and types
- Rewrite smoke-code.ts Group 2 tests for extractAllFromCode (the modern API)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address Copilot review issues from simplify PR
- Restore @mariozechner/pi-agent-core as explicit dependency (still
imported as a type in src/optimizer/mutation/pi-coding.ts; relying on
transitive availability is fragile)
- Fix piSimpleComplete: throw unconditionally when stopReason === 'error'
(previously silent fallthrough when errorMessage was falsy) and throw
with content-type diagnostic when the response has no text blocks
- Remove redundant empty-text guard in default-pi-generator.ts now that
piSimpleComplete handles that case itself
- Replace unsafe `as` casts for verify/expected_fetches in config.ts with
Array.isArray guards before assignment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address secondary reviewer issues
- loadReport: detect pre-1.x report format (toolMatches field) and throw
a clear error instead of letting callers hit silent undefined-access later
- loadProjectConfig: when skill-optimizer.json is missing, check for
skill-benchmark.json and surface a targeted rename hint instead of the
generic "config not found" error
- CHANGELOG: add BREAKING CHANGES section documenting all removed public
API exports and renamed TaskResult fields
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename minOverallPassDelta, fix mock repo configs, remove ghost repo names
- Part F: rename `minOverallPassDelta` → `minImprovement` in optimizer
types, loop resolution, adapters, and tests to unify with the public
ProjectOptimizeConfig field name
- Part G: fix dot-notation model IDs (claude-sonnet-4.6 → 4-6,
gpt-5.4 → 5-4) in all three mock-repo skill-optimizer.json files;
delete stale legacy benchmark.config.json and optimize.config.json
from mcp-tracker-demo (superseded by skill-optimizer.json)
- Part H: remove ghost names sdk-demo, cli-demo, mcp-demo from
MOCK_REPO_NAMES and update materialize-mock-repo.ts usage string to
only list the three templates that actually exist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove expected_tools/method input shims and benchmark.tasks deprecated field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: reapply Copilot comment fixes that were lost before agent commits
- loadProjectConfig: only check for skill-benchmark.json when configPath
was not explicitly supplied; use dirname() for clarity
- default-pi-critic: catch no-text-blocks error and return '[]' (graceful
degradation) while re-throwing real provider errors
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename matchTools→matchActions, remove pre-1.x report guard, clean straddling test comments
- Renamed exported matchTools → matchActions in evaluator.ts; updated
internal vars (toolMatches→actionMatches, unnecessaryCalls→unnecessaryActions,
hallucinatedCalls→hallucinatedActions) for consistency
- Stripped pre-1.x schema guard from loadReport in compare.ts — old format
check on 'toolMatches' field was pure backwards-compat dead code
- Removed "Accept either old or new error" branching from two smoke tests;
both now only assert the current validation path
- Updated CLAUDE.md: manifest fallback is a supported feature, not a
transitional shim
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove skill-benchmark.json legacy filename check from loadProjectConfig
The check for the old config filename was a migration hint for users
upgrading from a pre-1.x version. Current config is skill-optimizer.json;
drop the dead code.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename misleading 'legacy' test names in smoke-actions
Test names said 'legacy' and 'legacy shape' but were testing current
SurfaceSnapshot ↔ ActionCatalog conversion behaviour. Rename variables
and assertions to remove the legacy framing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove/reword remaining legacy and old language
- README.md: remove skill-benchmark.json migration hint (the detection
code was removed in the previous commit)
- cli.ts: rename example path results/old/ → results/baseline/ (clearer)
- snapshot.ts: rephrase error from "old format" to "not supported"
- smoke-actions.ts: align test name and assertion strings with the
updated error message; drop "legacy" framing from assertion strings
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: final cleanup — remove unused exports, trivial comments, stale doc entries
Unused exports:
- Make extractPythonSdk/extractRustSdk/extractTypeScriptSdk unexported;
they are private implementation details assigned to adapter objects in
the same file, not part of the public API
Trivial JSDoc removal:
- shared.ts: remove one-liner comments that restate function names
- reporter.ts: remove pad/center JSDoc; remove redundant generateMarkdown doc
- anthropic-format.ts: remove comment restating toAnthropicTool's name
- openai-format.ts: remove HTTP-method comments that add no information
Docs:
- README.md: fix benchmark.models description to cover all three providers
(openrouter, anthropic, openai) — not OpenRouter only
- CLAUDE.md: fix layer count (5, not 4); fix dot-notation validator
description (skips openai/, not just openrouter/)
- SKILL/references/setup.md: remove troubleshooting row for
skill-benchmark.json detection that was removed last commit
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct inaccurate CHANGELOG claim and wrong GitHub URL in CONTRIBUTING
- CHANGELOG: loadReport does not validate old field names (the validation
guard was intentionally removed); update wording to match actual behavior
- CONTRIBUTING: fix clone URL from bucurdavid fork to fastxyz canonical repo
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: repair three regressions in the generate-tasks → run workflow
Issue A (freeze.ts): surface.snapshot.json was written as a plain
{surface,actions} SurfaceSnapshot, which loadSurfaceSnapshotFile now
rejects. Switch to writeActionSnapshotFile + fromSurfaceSnapshot so
the file is written in the versioned {version,catalog} format.
Issue B (shared.ts): truncateToolResult sliced to MAX_TOOL_RESULT_CHARS
then appended the 16-char suffix, producing 50,016-char strings.
Subtract TRUNCATED_SUFFIX.length from the slice index so the final
string is exactly MAX_TOOL_RESULT_CHARS.
Issue C (types/resolve/adapters): benchmark.tasks was removed from
ProjectBenchmarkConfig during the simplify cleanup. freeze.ts still
writes tasks: tasksPath into benchmark.generated.json, but the field
was silently dropped on load and adapters.ts hard-coded
tasks: '__generated__', causing the runner to throw "run generate-tasks
first" even after generate-tasks had succeeded. Re-add tasks?: string
to both benchmark config interfaces, thread it through resolve.ts, and
use project.benchmark.tasks ?? '__generated__' in adapters.ts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* typos
* license typo
* fix(validation): restore missing-tasks guard, validate verify/expected_fetches elements, document tasks.json migration
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add prompt surface to init comment, E_INVALID_SURFACE fix text, and config-schema reference
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
This PR rolls forward the codebase to the “v1.1” shape by introducing a new prompt surface (content-based benchmarking), completing the migration from legacy tools/calls terminology to actions, and expanding LLM credential handling to support Codex auth alongside env-based keys.
Changes:
- Add prompt surface discovery + evaluation flow (task generation/grounding + runner support).
- Rename/migrate evaluation/task/report structures from
expected_tools/toolMatches/hallucinatedCalls→expected_actions/actionMatches/hallucinatedActions. - Add authMode (
env|codex|auto) and shared Pi runtime helpers (including tool-name aliasing for provider constraints).
Reviewed changes
Copilot reviewed 118 out of 119 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/smoke-snapshot-prompt.ts | New smoke test for prompt surface snapshot loading |
| tests/smoke-sdk-rust.ts | Update expected action shape/naming for Rust SDK smoke |
| tests/smoke-sdk-python.ts | Update expected action shape/naming for Python SDK smoke |
| tests/smoke-optimize.ts | Update optimize report/task fields to action terminology + minImprovement |
| tests/smoke-mcp.ts | Update evaluator matching tests from tools→actions |
| tests/smoke-init.ts | Update preset count + enforce hyphenated model IDs in preset values |
| tests/smoke-generation.ts | Ensure generated benchmark config preserves authMode and omits optimize-only settings; expected_actions rename |
| tests/smoke-feedback.ts | Update TaskResult construction to new action fields |
| tests/smoke-errors.ts | Remove legacy config filename smoke path |
| tests/smoke-e2e.ts | Rename optimizer policy field to minImprovement |
| tests/smoke-discovery-mcp.ts | Tighten missing-source error expectation |
| tests/smoke-coverage.ts | Remove legacy expected_tools duplication in generated tasks |
| tests/smoke-actions.ts | Add nested MCP schema roundtrip tests; update snapshot compatibility expectations |
| tests/fixtures/sample-skill.md | Add sample markdown skill fixture for prompt surface |
| src/tasks/types.ts | Make expected_actions required on GeneratedTask |
| src/tasks/pi-simple-complete.ts | New shared Pi completeSimple helper with timeout + response normalization |
| src/tasks/index.ts | Skip coverage enforcement loop for prompt surface task generation |
| src/tasks/ground.ts | Enforce prompt tasks have empty expected_actions; remove legacy expected_tools fallback |
| src/tasks/generate.ts | Prompt-surface task prompt; remove expected_tools/method aliases; simplify normalization |
| src/tasks/freeze.ts | Freeze action snapshot artifact instead of plain surface snapshot; preserve benchmark authMode; omit optimize |
| src/tasks/default-pi-generator.ts | Use piSimpleComplete + plumb authMode |
| src/tasks/default-pi-critic.ts | Use piSimpleComplete, treat “no text blocks” as empty recs, plumb authMode |
| src/tasks/coverage.ts | Coverage now uses expected_actions only |
| src/runtime/pi/models.ts | Add Codex routing and provider/key mismatch guards; plumb authMode |
| src/runtime/pi/index.ts | Export new auth helpers and types |
| src/runtime/pi/coding-orchestrator.ts | Plumb authMode into Pi model resolution |
| src/runtime/pi/benchmark-agent.ts | Remove unused legacy benchmark-agent helper |
| src/runtime/pi/auth.ts | Add authMode + Codex credential resolution (~/.codex/auth.json) + requireConfiguredApiKey |
| src/project/validate.ts | Add prompt surface validation; auth-mode/provider checks; updated API key diagnostics; remove legacy config checks |
| src/project/types.ts | Add authMode to benchmark/optimize configs; remove SurfaceSnapshotArg |
| src/project/snapshot.ts | Prompt surface snapshot building; disallow legacy plain snapshot format; preserve nested MCP arg schemas |
| src/project/schema.ts | Add prompt surface + authMode fields to schema docs |
| src/project/resolve.ts | Default authMode; resolve authMode into resolved config |
| src/project/load.ts | Remove legacy config filename auto-detection |
| src/project/index.ts | Remove legacy exports; remove toLegacyOptimizeManifest export |
| src/project/fix.ts | Remove deprecated-tasks-field fixer logic |
| src/project/adapters.ts | Add authMode to benchmark/optimize manifests; rename minOverallPassDelta→minImprovement; remove legacy optimize manifest adapter |
| src/optimizer/types.ts | Rename minOverallPassDelta→minImprovement; add mutation authMode |
| src/optimizer/mutation/types.ts | Remove redundant re-export file |
| src/optimizer/mutation/pi-coding.ts | Plumb authMode to orchestrator; tighten typing of agent messages |
| src/optimizer/mock-repos.ts | Reduce supported mock repo templates list |
| src/optimizer/materialize-mock-repo.ts | Update usage/examples for new mock repo templates list |
| src/optimizer/main.ts | Pass mutation authMode into generation deps |
| src/optimizer/loop.ts | Rename minOverallPassDelta usage to minImprovement; remove legacy toolMatches fallbacks |
| src/optimizer/feedback/passing-failing-diff.ts | Stop casting extractedCalls; use method directly |
| src/optimizer/feedback/mutation-context.ts | Make FeedbackPackage internal (non-exported) |
| src/optimizer/feedback/failure-details.ts | Remove casts and legacy fallbacks; use actionMatches + hallucinatedActions |
| src/optimizer/failure-analysis.ts | Remove legacy fallbacks; use actionMatches + hallucinatedActions |
| src/optimizer/benchmark-adapter.ts | Make run option/result interfaces internal |
| src/init/wizard.ts | Add prompt surface option; update model presets to hyphenated version segments |
| src/init/scaffold.ts | Update known models list; make config builder internal; remove hard-coded apiKeyEnv defaults; add prompt scaffolding |
| src/init/detect-project.ts | Make DetectedProject internal |
| src/init/answers.ts | Add prompt to surface union; update default models; validate prompt as allowed |
| src/import/extractors/ts-commander.ts | Remove redundant comments |
| src/import/extractors/rs-clap.ts | Remove redundant comments; minor cleanup |
| src/import/extractors/py-click.ts | Remove redundant comment |
| src/errors.ts | Add prompt to invalid-surface fix; remove legacy config error; update issue URL |
| src/doctor/index.ts | Make DoctorOptions internal |
| src/doctor/checks.ts | Use requireConfiguredApiKey; skip reachability for non-openrouter models |
| src/discovery/types.ts | Reuse ActionArgSchema for args |
| src/discovery/sdk.ts | Switch to ActionArgSchema typing for discovered args |
| src/discovery/prompt.ts | New prompt surface discovery (phases/capabilities) |
| src/discovery/optique.ts | Switch to ActionArgSchema typing for args |
| src/discovery/mcp.ts | Preserve full JSON schema in discovered MCP args |
| src/discovery/index.ts | Remove discovery barrel exports |
| src/discovery/cli.ts | Switch to ActionArgSchema typing for args |
| src/cli.ts | Add prompt surface to usage; credential checks via requireConfiguredApiKey; various messaging updates |
| src/benchmark/types.ts | Add prompt surface; remove legacy aliases; make ExpectedAction.name required; rename TaskResult fields |
| src/benchmark/scoring.ts | Minor comment update |
| src/benchmark/runner.ts | Add prompt surface runner path + prompt evaluator integration |
| src/benchmark/prompts.ts | Add prompt surface system/task prompt behavior |
| src/benchmark/llm/tool-name-aliases.ts | New tool-name alias codec to satisfy provider tool-name constraints |
| src/benchmark/llm/shared.ts | New shared retry/sleep/truncation/usage helpers for LLM formats |
| src/benchmark/llm/pi-format.ts | Add tool-name aliasing + authMode plumb + stricter PI error handling |
| src/benchmark/llm/openai-format.ts | Add tool-name aliasing; shared retry/truncation helpers; usage normalization |
| src/benchmark/llm/index.ts | Support authMode/codex bridging; strip provider prefixes for direct APIs; centralize credential resolution |
| src/benchmark/llm/anthropic-format.ts | Use shared retry/truncation/usage helpers |
| src/benchmark/config.ts | Make expected_actions required; remove expected_tools/method normalization; move loaders to actions/loaders |
| src/benchmark/compare.ts | Remove redundant comments |
| src/benchmark/index.ts | Update exported types to new names (ExpectedAction/ActionMatch) |
| src/benchmark/extractors/sdk/typescript.ts | Make extractor function internal |
| src/benchmark/extractors/sdk/rust.ts | Make extractor function internal |
| src/benchmark/extractors/sdk/python.ts | Make extractor function internal |
| src/benchmark/extractors/index.ts | Add prompt surface no-op extractor; remove re-export bundle |
| src/benchmark/extractors/code-analyzer.ts | Make RawExtraction internal; remove deprecated API |
| src/benchmark/evaluator.ts | Rename matchTools→matchActions; remove legacy toolMatches/hallucinatedCalls plumbing; add prompt to surface union |
| src/benchmark/init.ts | Add prompt surface scaffolding; update default model set; remove hard-coded apiKeyEnv defaults |
| src/benchmark/reporter.ts | Remove redundant doc comments |
| src/actions/types.ts | Add prompt to ActionSurface; add optional nested schema to ActionArgSchema |
| src/actions/snapshot.ts | Validate + normalize nested arg schemas; allow prompt surface in snapshot artifacts |
| src/actions/loaders.ts | New shared loaders for MCP tools and CLI commands |
| src/actions/index.ts | Export new loaders |
| src/actions/discover.ts | Use new loaders; preserve MCP arg schema field |
| package.json | Expand test suite to include prompt + snapshot tests; append smoke-actions |
| mock-repos/sdk-counter-demo/skill-optimizer.json | Hyphenate model IDs in demo config |
| mock-repos/mcp-tracker-demo/skill-optimizer.json | Hyphenate model IDs in demo config; remove legacy configs |
| mock-repos/mcp-tracker-demo/optimize.config.json | Remove legacy optimize config |
| mock-repos/mcp-tracker-demo/benchmark.config.json | Remove legacy benchmark config |
| mock-repos/cli-taskfile-demo/skill-optimizer.json | Hyphenate model IDs in demo config |
| docs/reference/errors.md | Update surface list; remove legacy-config entry; update issue URL |
| docs/reference/config-schema.md | Add prompt + authMode fields; update apiKeyEnv description |
| SKILL/references/setup.md | New setup/init reference doc |
| SKILL/references/optimize.md | New optimization loop reference doc |
| SKILL/references/config.md | New configuration reference doc |
| SKILL/references/benchmark.md | New benchmark interpretation reference doc |
| SKILL/SKILL.md | New repo skill doc content |
| README.md | Update requirements/auth guidance; add prompt surface overview; update examples |
| LICENSE | Update copyright holder |
| CONTRIBUTING.md | Update clone URL |
| CLAUDE.md | Update architecture layer count + model ID conventions section |
| CHANGELOG.md | Add breaking-change notes + prompt surface + auth-related notes |
| .gitignore | Ignore additional internal docs dirs |
| .github/workflows/ci.yml | Rename workflow; expand branches; add matrix job name |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| printSummary(report); | ||
|
|
||
| // Print coverage | ||
| printCoverage(report.coverage); |
There was a problem hiding this comment.
printCoverage() assumes coverage.length > 0 (it uses Math.max(...coverage.map(...)) and then padEnd(maxMethodLen)), but the new prompt surface currently produces an empty knownMethods set in the runner, so report.coverage can be []. That will cause padEnd(-Infinity) to throw at runtime. Guard this call (skip coverage output when empty / prompt surface), or update printCoverage() to handle an empty array safely.
| @@ -0,0 +1,25 @@ | |||
| import { test } from 'node:test'; | |||
| import assert from 'node:assert/strict'; | |||
| import { writeFileSync, mkdirSync, mkdtempSync, rmSync } from 'node:fs'; | |||
There was a problem hiding this comment.
Unused import: mkdirSync is imported from node:fs but never used in this test. Removing it will keep the smoke test lint/tsc-clean.
| format: z.enum(['pi', 'openai', 'anthropic']).optional().describe('LLM transport format: "pi" routes through OpenRouter/Pi (use openrouter/* or openai/* model refs); "openai" calls the OpenAI API directly (supports Codex auth); "anthropic" calls the Anthropic API directly'), | ||
| authMode: z.enum(['env', 'codex', 'auto']).optional().describe('How to resolve credentials: env var, ~/.codex/auth.json browser-login tokens, or env-then-codex fallback'), | ||
| apiKeyEnv: z.string().optional().describe('Env var name for the API key (default: OPENROUTER_API_KEY for format:pi, OPENAI_API_KEY for format:openai, ANTHROPIC_API_KEY for format:anthropic)'), |
There was a problem hiding this comment.
The benchmark.apiKeyEnv schema description says the default is based on benchmark.format, but in code credential resolution ultimately depends on the model/provider being used (e.g. pi format can still use openai/* model refs that default to OPENAI_API_KEY). Consider updating this description to reflect provider/model-prefix based defaults to avoid misleading config docs.
| throw new Error( | ||
| `Snapshot file format is not supported — delete .skill-optimizer/ and re-run the benchmark to regenerate.`, | ||
| ); |
There was a problem hiding this comment.
When rejecting the legacy plain SurfaceSnapshot JSON shape, the thrown error does not include the snapshotPath. Including the path in this message would make it much easier to diagnose which snapshot file is being rejected (especially when multiple generated dirs exist).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: strip provider prefix from model ID for direct Anthropic/OpenAI API formats (#12)
* fix(benchmark): strip provider prefix from model ID for direct API formats
When `benchmark.format` is `anthropic` or `openai` with a direct
`baseUrl`, the model ID (e.g. `anthropic/claude-sonnet-4-6`) includes a
provider prefix that OpenRouter expects but direct APIs reject.
The Anthropic API returns:
404 {"error":{"message":"model: anthropic/claude-sonnet-4-6"}}
Root cause: `createLLMClient` in `src/benchmark/llm/index.ts` passes
`model.id` verbatim to the API request body. The config validator
requires the prefix (`anthropic/...`) but the direct API expects the
bare name (`claude-sonnet-4-6`).
Fix: added `stripProviderPrefix()` that removes everything up to the
first `/` when format is `anthropic` or `openai`. Applied in all three
call paths (chat, chatWithTools, chatAgentLoop). The `pi` format is
unaffected.
Tests: 4 new cases in `smoke-llm.ts`:
- anthropic: strips prefix in chat
- anthropic: strips prefix in chatWithTools
- anthropic: no-op when prefix absent
- openai: strips prefix in chat
* fix: use prefix-aware stripping instead of indexOf in stripProviderPrefix
indexOf('/') only strips the first slash segment, so openrouter/anthropic/
model IDs would become anthropic/claude-... (still wrong for direct Anthropic API).
validate.ts accepts openrouter/ prefixes as valid for any format, making this
reachable by documented usage.
New approach: only strip the direct-provider prefixes (anthropic/, openai/).
openrouter/ IDs are left intact — they signal format:'pi' usage, so encountering
one in a direct-API context is a misconfiguration that should fail visibly at the
API level rather than being silently misrouted.
Also adds a test confirming openrouter/-prefixed IDs are passed through unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: fix env var restore and add chatAgentLoop prefix-stripping coverage
- Restore OPENAI_API_KEY to its original value (or delete it if absent)
in a finally block so CI environments don't lose a pre-existing key
- Add chatAgentLoop test to cover the third call path for prefix stripping,
matching the chat and chatWithTools cases
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: add missing models field to openaiConfig test fixture
LLMConfig.models is required; the inline config in the OpenAI prefix-stripping
test was missing it, causing a structural type error in editors.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Add current OpenRouter model presets (#10)
* Add current OpenRouter model presets
* fix: align optimize model ID to dot notation used by benchmark presets
commonOptimize.model was left as 'openrouter/anthropic/claude-sonnet-4-6'
(dashes) while the rest of this PR uses 'openrouter/anthropic/claude-sonnet-4.6'
(dots). A mismatched model ID would cause `skill-optimizer optimize` to call
a different or invalid OpenRouter ref than the benchmark uses.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: use hyphen notation in model IDs to match codebase convention
validate.ts flags digit.digit patterns in model IDs as bad format
(warning: "OpenRouter expects hyphens") and fix.ts auto-corrects them.
All new model IDs in this PR used dot notation (e.g. claude-sonnet-4.6)
which would trigger warnings on every init-generated config.
Convert all version segments in model ID strings to hyphens to match
the existing convention. Display names (name/label fields) keep dots
for readability (e.g. 'Claude Sonnet 4.6').
Also reverts the previous incorrect fix to commonOptimize.model —
the dash format was already correct.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(claude): document model ID hyphen convention and display name distinction
validate.ts warns on digit.digit patterns in model IDs and fix.ts
auto-corrects them. This tripped up PR #10 where new presets used dot
notation. Document the rule clearly so it's not re-discovered the hard way.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: enforce hyphen convention for model ID version segments
Adds a check to the MODEL_PRESETS block in smoke-init.ts that fails if any
preset value contains a digit.digit pattern (e.g. claude-sonnet-4.6).
validate.ts already flags these at runtime; this makes CI catch them at
commit time so the convention is enforced on every PR without manual review.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* ci: rename workflow to 'Build & Test', add development and staging branches
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* ci: restore job ID to build-test to preserve required check names
Renaming the job broke branch protection rules that require build-test (20)
and build-test (22). Keep the job ID as-is; the display name still shows
'Build & Test (node X)' in the GitHub UI.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: dmn <damian.ovidiu27@gmail.com>
* Add Codex auth support (#11)
* Add Codex auth support
* Fix Codex auth CI smoke test
* Fix codex auth in optimize task generation
* Fix Codex MCP schema handling
* fix(auth): prioritize tokens.access_token over static OPENAI_API_KEY in readCodexApiKey
The browser-login JWT (tokens.access_token) now takes highest priority,
followed by tokens.OPENAI_API_KEY, then the top-level OPENAI_API_KEY.
Previously, a stale static key could shadow a valid browser-login token,
defeating the purpose of Codex auth. The expiry check on access_token is
preserved and still applied.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): resolve OpenAI credential per-call to avoid stale Codex token
Move resolvedOpenAICredential and useCodexBridgeForOpenAI out of the
createLLMClient closure and into each method body, so long-running
benchmarks pick up refreshed Codex tokens instead of a stale snapshot.
Also pass apiKeyOverride (pre-resolved key) to Pi bridge calls instead
of authMode+apiKeyEnv, preventing a second auth.json read on every call.
* test(llm): cover chatWithTools+chatAgentLoop codex bridge; narrow resolveDirectApiKey return type
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(doctor): probe each model individually instead of bailing when first model is non-openrouter
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tests): restore HOME env var with delete-vs-assign to avoid "undefined" string pollution
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use benchmark.authMode as field for codex auth warning instead of benchmark.apiKeyEnv
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(optimizer): reference PiAuthMode instead of duplicating literal union
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: clarify Codex auth paths and format/apiKeyEnv descriptions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: simplify code clarity in Codex auth changes
Drop the redundant useCodexBridgeForOpenAI variable in createLLMClient.
Since resolveOpenAICredential() returns undefined unless format is
'openai', checking openAICredential?.source === 'codex' is sufficient
on its own.
Also regenerate docs/reference/config-schema.md to reflect the
schema.ts description updates from the PR.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(review): add format guard comment, accurate doctor skip message, and auto-mode priority test
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(auth): skip apiKeyEnv lookup when authMode=codex; avoid writing OPENROUTER_API_KEY in openai-format scaffold
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use integer index instead of model.id in error field paths
Replace `for...of` loops over `benchmark.models` with `entries()` variants
so error field paths use the numeric index (e.g. `benchmark.models[0].weight`)
instead of the model ID string, which produced malformed paths like
`benchmark.models[openai/gpt-5.4].weight`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): check all providers in mixed-model configs and validate optimize auth
Fix 2: when benchmark.format is 'pi', models can carry different provider
prefixes (openrouter/*, openai/*, anthropic/*). The old code derived
benchmarkProvider from models[0] and validated only that one credential.
Now all unique provider prefixes are collected and each is checked
separately, so a mixed config like ["openai/gpt-4", "openrouter/…"]
warns when either key is missing.
Fix 3: the optimize section's authMode/apiKeyEnv was never validated for
key presence — only codex provider-compatibility was checked. Now the
optimize credential is resolved and a warning is pushed if no key is
found, parallel to the benchmark check.
Both fixes share a warnMissingApiKey helper that formats the hint and
field path consistently for both benchmark and optimize sections.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(scaffold): drop redundant benchmark.apiKeyEnv from scaffold defaults
benchmark.apiKeyEnv: 'OPENROUTER_API_KEY' was written into every
scaffolded config by both scaffold.ts (commonBenchmark) and
init.ts (commonBenchmark + commonOptimize). The field is redundant:
resolveApiKey already falls back to OPENROUTER_API_KEY for the
openrouter provider via defaultApiKeyEnvForProvider.
With authMode:'auto' now in the PR, the explicit field becomes a
footgun: when a user edits their config to use openai/* models and
forgets to remove apiKeyEnv, resolveApiCredential reads OPENROUTER_API_KEY
from the env (present and non-empty) and returns it as the OpenAI key —
skipping the Codex fallback and sending a wrong key to the OpenAI API.
Fix 4 already removed it from scaffold.ts commonOptimize. This commit
removes it from the three remaining sites.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): restore authMode:codex in bridge calls to preserve Pi routing signal
The Codex bridge was passing apiKeyOverride to chatPi, which caused
resolveApiCredential to short-circuit with source:'override'. The routing
decision in resolvePiModel checks credential.source === 'codex' to switch
to the 'openai-codex' provider (synthesizeOpenAICodexModel, Codex endpoint
at chatgpt.com/backend-api/codex). With source:'override' that branch was
never taken — the browser JWT was being sent to the standard OpenAI API
endpoint instead, causing a 401.
Restore the original approach: pass authMode:'codex' so Pi re-reads
~/.codex/auth.json, returns source:'codex', and routes to openai-codex.
The second disk read is load-bearing for the provider routing decision.
Update the three smoke tests that were asserting the old (incorrect)
apiKeyOverride behavior.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix dot-notation model IDs in README; extend CLAUDE.md ID convention
README used openai/gpt-5.4 (dot notation) in the Codex auth example —
two occurrences. Corrected to openai/gpt-5-4 to match the project's
model ID convention (hyphens in version segments).
CLAUDE.md model ID convention section only documented the openrouter/
format. Extended it to also cover the direct-provider formats
(openai/<model>, anthropic/<model>) that this PR introduces, with the
same hyphen rule and an openai/gpt-5-4 example.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate,models): restore dot-version warning for anthropic/; add OpenAI OPENROUTER_API_KEY guard
Finding 1: validate.ts narrowed the model-id-bad-format check to
openrouter/* only, which stopped flagging anthropic/claude-sonnet-4.6 —
a real error since the Anthropic API expects hyphens. The exemption should
only apply to openai/ IDs, where OpenAI's own API uses dots in some model
slugs (e.g. gpt-4.5). Change the guard to !startsWith('openai/') and
update the message to be provider-agnostic.
Finding 2: resolvePiModel had an Anthropic-specific guard that throws when
apiKeyEnv is OPENROUTER_API_KEY, preventing silent 401s from stale
scaffolded configs. No equivalent existed for the openai provider.
Add a parallel guard so direct openai/* calls with an OpenRouter key
fail fast with a clear error and a migration hint.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(cli,auth,validate): three preflight and credential correctness gaps
cli.ts: run preflight checked only models[0] provider for pi format,
so a mixed-provider config (e.g. openrouter + openai models) would pass
startup and then silently record per-model failures for whichever
provider's key was missing. Now all unique providers are checked, matching
what validate.ts already does across the model list.
auth.ts: readCodexApiKey returned {} immediately on an expired JWT,
making the tokens.OPENAI_API_KEY and top-level OPENAI_API_KEY fallbacks
unreachable when access_token was present-but-expired. The comment's
intent was correct (don't shadow valid JWT with stale static key) but the
inverse was wrong (expired JWT was shadowing valid static key). Now falls
through to the static key branches on JWT expiry.
validate.ts: optimize auth check ran even when optimize.enabled === false,
causing spurious auth warnings for benchmark-only setups. Added the same
enabled !== false guard that neighboring optimize checks already use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: add SKILL folder for AI agent guidance (#14)
* docs: add SKILL design spec for skill-optimizer
Spec for a multi-file SKILL/ folder that guides AI agents through
setting up, benchmarking, and optimizing SDK/CLI/MCP documentation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: add implementation plan for SKILL folder
7-task plan covering branch creation, all 5 markdown files with
complete content, verification, and PR creation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(skill): add SKILL.md entry point with context detection and routing
* feat(skill): add setup reference — prerequisites, init wizard, discovery
* feat(skill): add benchmark reference — run, interpret, diagnose, compare
* feat(skill): add optimize reference — loop mechanics, safety, troubleshooting
* feat(skill): add config reference — schema, models, scope, errors
* fix(skill): address Copilot review — config path, safety boundary, minImprovement, gitignore dev docs
* docs(skill): update routing concept; add direct-provider auth to prerequisites; fix dot-notation
* docs(skill): expand compare section in benchmark reference
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(prompt): add prompt surface type for skill/template benchmarking (#20)
* docs: fix repo URL in CONTRIBUTING.md (bucurdavid → fastxyz)
Fixes #18
* feat(prompt): add prompt surface type for skill/template benchmarking
Discovers phases, instructions, output formats, and decision points from
markdown skill files. Evaluates model responses with content-based criteria
(required sections, format patterns, keywords, structural elements) instead
of tool-call matching.
New files:
- src/project/discover-prompt.ts — markdown parser for capabilities
- src/benchmark/prompt-evaluator.ts — content-based scorer (weighted)
- src/discovery/prompt.ts — file-based discoverer + types
- tests/smoke-discovery-prompt.ts — 10 tests
- tests/smoke-prompt-evaluator.ts — 15 tests
- tests/fixtures/sample-skill.md — 3-phase sample
Modified: types.ts, validate.ts, snapshot.ts, runner.ts, evaluator.ts,
extractors/index.ts, prompts.ts, cli.ts, init/, CHANGELOG.md, README.md
* fix(prompt): guard snapshot surface, fix format inconsistency, remove dead interface, fix stale comment
* fix(discovery): anchor frontmatter closing delimiter to line boundary
Replace indexOf('---', 3) with a regex that matches --- only at the start
of a line, preventing early termination when a YAML value contains ---.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(discover-prompt): clear preamble on first heading, strip frontmatter before section split
* fix(discover-prompt): fallback to whole-content instruction extraction for heading-less files
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(discover-prompt): document interface difference between production and standalone paths
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(prompt): wire evaluatePromptResponse into runner, expose section text, fix generate schema
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(prompt): unique temp dir + cleanup guard in test, enforce empty actions in grounding, bypass maxTasks constraint for prompt
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Refactoring & Removal of legacy code (#15)
* refactor: DRY deduplication across LLM format handlers and Pi task helpers
Extract shared retry logic (isAbortError, isRetryableError, sleep), tool-result
truncation, and agent-loop accumulator helpers into src/benchmark/llm/shared.ts,
eliminating identical code that existed in both openai-format.ts and
anthropic-format.ts.
Extract the Pi completeSimple call sequence (timeout setup, error check, text
block extraction) into src/tasks/pi-simple-complete.ts, used by both
default-pi-generator.ts and default-pi-critic.ts to replace ~40 lines of
near-identical implementation in each file.
Net: -129 lines of duplication, no behavior changes, all tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): consolidate duplicate type definitions
- DiscoveredActionArg (discovery/types) was identical to ActionArgSchema
(actions/types); replaced with a re-export alias to eliminate the copy
- BenchmarkSurface now aliases ActionSurface instead of re-declaring the
same 'sdk' | 'cli' | 'mcp' union literal
- Removed dead SurfaceActionArg alias from benchmark/types (never imported)
- Removed dead SurfaceSnapshotArg alias from project/types and project/index
(exported but never consumed)
- Replaced all inline import('...').PiAuthMode dynamic-import type refs
with proper named import type statements in pi-format, models,
coding-orchestrator, benchmark-agent, default-pi-critic, default-pi-generator
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove unused code found by knip
- Delete src/runtime/pi/benchmark-agent.ts (createReadOnlyBenchmarkModel was dead — exported but never imported)
- Delete src/discovery/index.ts and src/optimizer/mutation/types.ts (barrel re-exports with no consumers)
- Remove requireApiKeyFromEnv from auth.ts (never called internally or imported from outside)
- Remove 5 unused re-export lines from benchmark/extractors/index.ts (extractCodeBlock, extractShellBlock etc. — tests import directly from sub-files)
- Unexport 11 internal-only types/interfaces (MAX_TOOL_RESULT_CHARS, buildConfigFromAnswers, LLMClient, ToolNameAliasCodec, PromptSurface/Options, DoctorOptions, DetectedProject, BenchmarkAdapterRunOptions/Result, FeedbackPackage, RawExtraction, PiSimpleCompleteOptions/Input)
- Remove @mariozechner/pi-agent-core from package.json dependencies (transitive dep, never directly imported)
- Add smoke-actions.ts to the npm test script (17 passing tests were orphaned from the suite)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(deps): break circular dependency between actions/discover and benchmark/config
Move loadCliCommands and loadMcpTools from benchmark/config.ts into a new
actions/loaders.ts file, resolving the cycle:
actions/discover → benchmark/config → project → project/snapshot → actions/discover
benchmark/config.ts re-exports both functions for backward compatibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): replace weak types with precise types throughout
- Replace `Model<any>` with `Model<Api>` and specific API literal types
(`Model<'openai-completions'>`, `Model<'openai-codex-responses'>`) in
models.ts and pi-format.ts
- Replace `(err as any).status` with typed intersection cast
`Error & { status: number }` in anthropic-format.ts and openai-format.ts
- Remove redundant `as unknown[]` cast on `session.state.messages` and
import `AgentMessage` from `@mariozechner/pi-agent-core` in pi-coding.ts;
update `extractLatestAssistantText` and `extractToolActivity` parameter
types accordingly
- Remove unnecessary double casts `(c as unknown as { method: string }).method`
in failure-details.ts and passing-failing-diff.ts — `ExtractedCall` is
`ActionAttempt` which already has `method: string` and `args`
- Remove `as any` from `Type.Unsafe(...)` call in pi-format.ts — the argument
type is compatible with `UnsafeOptions` (SchemaOptions index signature)
- Fix `catch (error: any)` to `catch (error)` in two snapshot files,
using `instanceof Error` guard to access `.message`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: remove unnecessary optional chaining on non-optional reportConfig
reportConfig is always defined (report.config is non-optional per
BenchmarkReport). The ?. on reportConfig before .outputDir was
defensive against a value that TypeScript knows can never be null.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove deprecated and legacy code paths
- Remove toLegacyOptimizeManifest (never called anywhere in src or tests)
- Remove deprecated SdkSurfaceConfig fields: classes, functions, functionReturns, methods
- Remove dead config.sdk?.methods fallback in runner.ts (no config ever sets methods)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove AI slop comments and stubs
Removed or rewrote ~75 comments across 15 files:
- Dropped "// NEW:", "// Transitional alias", numbered step comments (1-14
in runner.ts), and other in-motion migration language
- Removed redundant comments that restated the immediately-following code
- Replaced multi-line comment pairs with single consolidated explanations
where the context was non-obvious and worth preserving
- Kept all comments explaining invariants, non-obvious behavior, or
security rationale (e.g. git injection safety, OpenAI dot-format exemption)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove backwards-compatibility shims — SemVer allows breaking changes
- Remove LEGACY_PROJECT_CONFIG_NAME and skill-benchmark.json detection in load.ts
- Remove legacy-config-name warning from validate.ts (old filename alongside new one)
- Remove E_LEGACY_CONFIG error definition from errors.ts and its docs entry
- Remove DiscoveredActionArg deprecated alias from discovery/types.ts; update all four
discovery modules (cli, mcp, sdk, optique) to import ActionArgSchema directly
- Remove CodeModeConfig, McpModeConfig, ExpectedTool, ToolMatch backward-compat
type aliases from benchmark/types.ts and benchmark/index.ts
- Remove toolMatches dual-field from TaskResult; make actionMatches required
- Remove unnecessaryCalls/hallucinatedCalls alias fields from TaskResult.metrics;
consolidate to unnecessaryActions/hallucinatedActions throughout
- Remove method? alias from ExpectedAction; make name required
- Remove expected_tools? alias from TaskDefinition; make expected_actions required
- Simplify getExpectedActionName and getExpectedActions helper functions
- Remove deprecated extractFromCode function from code-analyzer.ts
- Update all internal consumers (optimizer, evaluator, tasks, benchmark) to use
canonical field names without the ?? fallback patterns
- Update all test files to use canonical field names and types
- Rewrite smoke-code.ts Group 2 tests for extractAllFromCode (the modern API)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address Copilot review issues from simplify PR
- Restore @mariozechner/pi-agent-core as explicit dependency (still
imported as a type in src/optimizer/mutation/pi-coding.ts; relying on
transitive availability is fragile)
- Fix piSimpleComplete: throw unconditionally when stopReason === 'error'
(previously silent fallthrough when errorMessage was falsy) and throw
with content-type diagnostic when the response has no text blocks
- Remove redundant empty-text guard in default-pi-generator.ts now that
piSimpleComplete handles that case itself
- Replace unsafe `as` casts for verify/expected_fetches in config.ts with
Array.isArray guards before assignment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address secondary reviewer issues
- loadReport: detect pre-1.x report format (toolMatches field) and throw
a clear error instead of letting callers hit silent undefined-access later
- loadProjectConfig: when skill-optimizer.json is missing, check for
skill-benchmark.json and surface a targeted rename hint instead of the
generic "config not found" error
- CHANGELOG: add BREAKING CHANGES section documenting all removed public
API exports and renamed TaskResult fields
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename minOverallPassDelta, fix mock repo configs, remove ghost repo names
- Part F: rename `minOverallPassDelta` → `minImprovement` in optimizer
types, loop resolution, adapters, and tests to unify with the public
ProjectOptimizeConfig field name
- Part G: fix dot-notation model IDs (claude-sonnet-4.6 → 4-6,
gpt-5.4 → 5-4) in all three mock-repo skill-optimizer.json files;
delete stale legacy benchmark.config.json and optimize.config.json
from mcp-tracker-demo (superseded by skill-optimizer.json)
- Part H: remove ghost names sdk-demo, cli-demo, mcp-demo from
MOCK_REPO_NAMES and update materialize-mock-repo.ts usage string to
only list the three templates that actually exist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove expected_tools/method input shims and benchmark.tasks deprecated field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: reapply Copilot comment fixes that were lost before agent commits
- loadProjectConfig: only check for skill-benchmark.json when configPath
was not explicitly supplied; use dirname() for clarity
- default-pi-critic: catch no-text-blocks error and return '[]' (graceful
degradation) while re-throwing real provider errors
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename matchTools→matchActions, remove pre-1.x report guard, clean straddling test comments
- Renamed exported matchTools → matchActions in evaluator.ts; updated
internal vars (toolMatches→actionMatches, unnecessaryCalls→unnecessaryActions,
hallucinatedCalls→hallucinatedActions) for consistency
- Stripped pre-1.x schema guard from loadReport in compare.ts — old format
check on 'toolMatches' field was pure backwards-compat dead code
- Removed "Accept either old or new error" branching from two smoke tests;
both now only assert the current validation path
- Updated CLAUDE.md: manifest fallback is a supported feature, not a
transitional shim
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove skill-benchmark.json legacy filename check from loadProjectConfig
The check for the old config filename was a migration hint for users
upgrading from a pre-1.x version. Current config is skill-optimizer.json;
drop the dead code.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename misleading 'legacy' test names in smoke-actions
Test names said 'legacy' and 'legacy shape' but were testing current
SurfaceSnapshot ↔ ActionCatalog conversion behaviour. Rename variables
and assertions to remove the legacy framing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove/reword remaining legacy and old language
- README.md: remove skill-benchmark.json migration hint (the detection
code was removed in the previous commit)
- cli.ts: rename example path results/old/ → results/baseline/ (clearer)
- snapshot.ts: rephrase error from "old format" to "not supported"
- smoke-actions.ts: align test name and assertion strings with the
updated error message; drop "legacy" framing from assertion strings
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: final cleanup — remove unused exports, trivial comments, stale doc entries
Unused exports:
- Make extractPythonSdk/extractRustSdk/extractTypeScriptSdk unexported;
they are private implementation details assigned to adapter objects in
the same file, not part of the public API
Trivial JSDoc removal:
- shared.ts: remove one-liner comments that restate function names
- reporter.ts: remove pad/center JSDoc; remove redundant generateMarkdown doc
- anthropic-format.ts: remove comment restating toAnthropicTool's name
- openai-format.ts: remove HTTP-method comments that add no information
Docs:
- README.md: fix benchmark.models description to cover all three providers
(openrouter, anthropic, openai) — not OpenRouter only
- CLAUDE.md: fix layer count (5, not 4); fix dot-notation validator
description (skips openai/, not just openrouter/)
- SKILL/references/setup.md: remove troubleshooting row for
skill-benchmark.json detection that was removed last commit
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct inaccurate CHANGELOG claim and wrong GitHub URL in CONTRIBUTING
- CHANGELOG: loadReport does not validate old field names (the validation
guard was intentionally removed); update wording to match actual behavior
- CONTRIBUTING: fix clone URL from bucurdavid fork to fastxyz canonical repo
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: repair three regressions in the generate-tasks → run workflow
Issue A (freeze.ts): surface.snapshot.json was written as a plain
{surface,actions} SurfaceSnapshot, which loadSurfaceSnapshotFile now
rejects. Switch to writeActionSnapshotFile + fromSurfaceSnapshot so
the file is written in the versioned {version,catalog} format.
Issue B (shared.ts): truncateToolResult sliced to MAX_TOOL_RESULT_CHARS
then appended the 16-char suffix, producing 50,016-char strings.
Subtract TRUNCATED_SUFFIX.length from the slice index so the final
string is exactly MAX_TOOL_RESULT_CHARS.
Issue C (types/resolve/adapters): benchmark.tasks was removed from
ProjectBenchmarkConfig during the simplify cleanup. freeze.ts still
writes tasks: tasksPath into benchmark.generated.json, but the field
was silently dropped on load and adapters.ts hard-coded
tasks: '__generated__', causing the runner to throw "run generate-tasks
first" even after generate-tasks had succeeded. Re-add tasks?: string
to both benchmark config interfaces, thread it through resolve.ts, and
use project.benchmark.tasks ?? '__generated__' in adapters.ts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* typos
* license typo
* fix(validation): restore missing-tasks guard, validate verify/expected_fetches elements, document tasks.json migration
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add prompt surface to init comment, E_INVALID_SURFACE fix text, and config-schema reference
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(release): bump version to 1.1.0, update CHANGELOG and docs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* v1.1 (#21)
* fix: strip provider prefix from model ID for direct Anthropic/OpenAI API formats (#12)
* fix(benchmark): strip provider prefix from model ID for direct API formats
When `benchmark.format` is `anthropic` or `openai` with a direct
`baseUrl`, the model ID (e.g. `anthropic/claude-sonnet-4-6`) includes a
provider prefix that OpenRouter expects but direct APIs reject.
The Anthropic API returns:
404 {"error":{"message":"model: anthropic/claude-sonnet-4-6"}}
Root cause: `createLLMClient` in `src/benchmark/llm/index.ts` passes
`model.id` verbatim to the API request body. The config validator
requires the prefix (`anthropic/...`) but the direct API expects the
bare name (`claude-sonnet-4-6`).
Fix: added `stripProviderPrefix()` that removes everything up to the
first `/` when format is `anthropic` or `openai`. Applied in all three
call paths (chat, chatWithTools, chatAgentLoop). The `pi` format is
unaffected.
Tests: 4 new cases in `smoke-llm.ts`:
- anthropic: strips prefix in chat
- anthropic: strips prefix in chatWithTools
- anthropic: no-op when prefix absent
- openai: strips prefix in chat
* fix: use prefix-aware stripping instead of indexOf in stripProviderPrefix
indexOf('/') only strips the first slash segment, so openrouter/anthropic/
model IDs would become anthropic/claude-... (still wrong for direct Anthropic API).
validate.ts accepts openrouter/ prefixes as valid for any format, making this
reachable by documented usage.
New approach: only strip the direct-provider prefixes (anthropic/, openai/).
openrouter/ IDs are left intact — they signal format:'pi' usage, so encountering
one in a direct-API context is a misconfiguration that should fail visibly at the
API level rather than being silently misrouted.
Also adds a test confirming openrouter/-prefixed IDs are passed through unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: fix env var restore and add chatAgentLoop prefix-stripping coverage
- Restore OPENAI_API_KEY to its original value (or delete it if absent)
in a finally block so CI environments don't lose a pre-existing key
- Add chatAgentLoop test to cover the third call path for prefix stripping,
matching the chat and chatWithTools cases
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: add missing models field to openaiConfig test fixture
LLMConfig.models is required; the inline config in the OpenAI prefix-stripping
test was missing it, causing a structural type error in editors.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Add current OpenRouter model presets (#10)
* Add current OpenRouter model presets
* fix: align optimize model ID to dot notation used by benchmark presets
commonOptimize.model was left as 'openrouter/anthropic/claude-sonnet-4-6'
(dashes) while the rest of this PR uses 'openrouter/anthropic/claude-sonnet-4.6'
(dots). A mismatched model ID would cause `skill-optimizer optimize` to call
a different or invalid OpenRouter ref than the benchmark uses.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: use hyphen notation in model IDs to match codebase convention
validate.ts flags digit.digit patterns in model IDs as bad format
(warning: "OpenRouter expects hyphens") and fix.ts auto-corrects them.
All new model IDs in this PR used dot notation (e.g. claude-sonnet-4.6)
which would trigger warnings on every init-generated config.
Convert all version segments in model ID strings to hyphens to match
the existing convention. Display names (name/label fields) keep dots
for readability (e.g. 'Claude Sonnet 4.6').
Also reverts the previous incorrect fix to commonOptimize.model —
the dash format was already correct.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(claude): document model ID hyphen convention and display name distinction
validate.ts warns on digit.digit patterns in model IDs and fix.ts
auto-corrects them. This tripped up PR #10 where new presets used dot
notation. Document the rule clearly so it's not re-discovered the hard way.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test: enforce hyphen convention for model ID version segments
Adds a check to the MODEL_PRESETS block in smoke-init.ts that fails if any
preset value contains a digit.digit pattern (e.g. claude-sonnet-4.6).
validate.ts already flags these at runtime; this makes CI catch them at
commit time so the convention is enforced on every PR without manual review.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* ci: rename workflow to 'Build & Test', add development and staging branches
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* ci: restore job ID to build-test to preserve required check names
Renaming the job broke branch protection rules that require build-test (20)
and build-test (22). Keep the job ID as-is; the display name still shows
'Build & Test (node X)' in the GitHub UI.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: dmn <damian.ovidiu27@gmail.com>
* Add Codex auth support (#11)
* Add Codex auth support
* Fix Codex auth CI smoke test
* Fix codex auth in optimize task generation
* Fix Codex MCP schema handling
* fix(auth): prioritize tokens.access_token over static OPENAI_API_KEY in readCodexApiKey
The browser-login JWT (tokens.access_token) now takes highest priority,
followed by tokens.OPENAI_API_KEY, then the top-level OPENAI_API_KEY.
Previously, a stale static key could shadow a valid browser-login token,
defeating the purpose of Codex auth. The expiry check on access_token is
preserved and still applied.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): resolve OpenAI credential per-call to avoid stale Codex token
Move resolvedOpenAICredential and useCodexBridgeForOpenAI out of the
createLLMClient closure and into each method body, so long-running
benchmarks pick up refreshed Codex tokens instead of a stale snapshot.
Also pass apiKeyOverride (pre-resolved key) to Pi bridge calls instead
of authMode+apiKeyEnv, preventing a second auth.json read on every call.
* test(llm): cover chatWithTools+chatAgentLoop codex bridge; narrow resolveDirectApiKey return type
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(doctor): probe each model individually instead of bailing when first model is non-openrouter
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tests): restore HOME env var with delete-vs-assign to avoid "undefined" string pollution
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use benchmark.authMode as field for codex auth warning instead of benchmark.apiKeyEnv
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(optimizer): reference PiAuthMode instead of duplicating literal union
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: clarify Codex auth paths and format/apiKeyEnv descriptions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor: simplify code clarity in Codex auth changes
Drop the redundant useCodexBridgeForOpenAI variable in createLLMClient.
Since resolveOpenAICredential() returns undefined unless format is
'openai', checking openAICredential?.source === 'codex' is sufficient
on its own.
Also regenerate docs/reference/config-schema.md to reflect the
schema.ts description updates from the PR.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(review): add format guard comment, accurate doctor skip message, and auto-mode priority test
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(auth): skip apiKeyEnv lookup when authMode=codex; avoid writing OPENROUTER_API_KEY in openai-format scaffold
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): use integer index instead of model.id in error field paths
Replace `for...of` loops over `benchmark.models` with `entries()` variants
so error field paths use the numeric index (e.g. `benchmark.models[0].weight`)
instead of the model ID string, which produced malformed paths like
`benchmark.models[openai/gpt-5.4].weight`.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate): check all providers in mixed-model configs and validate optimize auth
Fix 2: when benchmark.format is 'pi', models can carry different provider
prefixes (openrouter/*, openai/*, anthropic/*). The old code derived
benchmarkProvider from models[0] and validated only that one credential.
Now all unique provider prefixes are collected and each is checked
separately, so a mixed config like ["openai/gpt-4", "openrouter/…"]
warns when either key is missing.
Fix 3: the optimize section's authMode/apiKeyEnv was never validated for
key presence — only codex provider-compatibility was checked. Now the
optimize credential is resolved and a warning is pushed if no key is
found, parallel to the benchmark check.
Both fixes share a warnMissingApiKey helper that formats the hint and
field path consistently for both benchmark and optimize sections.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(scaffold): drop redundant benchmark.apiKeyEnv from scaffold defaults
benchmark.apiKeyEnv: 'OPENROUTER_API_KEY' was written into every
scaffolded config by both scaffold.ts (commonBenchmark) and
init.ts (commonBenchmark + commonOptimize). The field is redundant:
resolveApiKey already falls back to OPENROUTER_API_KEY for the
openrouter provider via defaultApiKeyEnvForProvider.
With authMode:'auto' now in the PR, the explicit field becomes a
footgun: when a user edits their config to use openai/* models and
forgets to remove apiKeyEnv, resolveApiCredential reads OPENROUTER_API_KEY
from the env (present and non-empty) and returns it as the OpenAI key —
skipping the Codex fallback and sending a wrong key to the OpenAI API.
Fix 4 already removed it from scaffold.ts commonOptimize. This commit
removes it from the three remaining sites.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(llm): restore authMode:codex in bridge calls to preserve Pi routing signal
The Codex bridge was passing apiKeyOverride to chatPi, which caused
resolveApiCredential to short-circuit with source:'override'. The routing
decision in resolvePiModel checks credential.source === 'codex' to switch
to the 'openai-codex' provider (synthesizeOpenAICodexModel, Codex endpoint
at chatgpt.com/backend-api/codex). With source:'override' that branch was
never taken — the browser JWT was being sent to the standard OpenAI API
endpoint instead, causing a 401.
Restore the original approach: pass authMode:'codex' so Pi re-reads
~/.codex/auth.json, returns source:'codex', and routes to openai-codex.
The second disk read is load-bearing for the provider routing decision.
Update the three smoke tests that were asserting the old (incorrect)
apiKeyOverride behavior.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix dot-notation model IDs in README; extend CLAUDE.md ID convention
README used openai/gpt-5.4 (dot notation) in the Codex auth example —
two occurrences. Corrected to openai/gpt-5-4 to match the project's
model ID convention (hyphens in version segments).
CLAUDE.md model ID convention section only documented the openrouter/
format. Extended it to also cover the direct-provider formats
(openai/<model>, anthropic/<model>) that this PR introduces, with the
same hyphen rule and an openai/gpt-5-4 example.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(validate,models): restore dot-version warning for anthropic/; add OpenAI OPENROUTER_API_KEY guard
Finding 1: validate.ts narrowed the model-id-bad-format check to
openrouter/* only, which stopped flagging anthropic/claude-sonnet-4.6 —
a real error since the Anthropic API expects hyphens. The exemption should
only apply to openai/ IDs, where OpenAI's own API uses dots in some model
slugs (e.g. gpt-4.5). Change the guard to !startsWith('openai/') and
update the message to be provider-agnostic.
Finding 2: resolvePiModel had an Anthropic-specific guard that throws when
apiKeyEnv is OPENROUTER_API_KEY, preventing silent 401s from stale
scaffolded configs. No equivalent existed for the openai provider.
Add a parallel guard so direct openai/* calls with an OpenRouter key
fail fast with a clear error and a migration hint.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(cli,auth,validate): three preflight and credential correctness gaps
cli.ts: run preflight checked only models[0] provider for pi format,
so a mixed-provider config (e.g. openrouter + openai models) would pass
startup and then silently record per-model failures for whichever
provider's key was missing. Now all unique providers are checked, matching
what validate.ts already does across the model list.
auth.ts: readCodexApiKey returned {} immediately on an expired JWT,
making the tokens.OPENAI_API_KEY and top-level OPENAI_API_KEY fallbacks
unreachable when access_token was present-but-expired. The comment's
intent was correct (don't shadow valid JWT with stale static key) but the
inverse was wrong (expired JWT was shadowing valid static key). Now falls
through to the static key branches on JWT expiry.
validate.ts: optimize auth check ran even when optimize.enabled === false,
causing spurious auth warnings for benchmark-only setups. Added the same
enabled !== false guard that neighboring optimize checks already use.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: add SKILL folder for AI agent guidance (#14)
* docs: add SKILL design spec for skill-optimizer
Spec for a multi-file SKILL/ folder that guides AI agents through
setting up, benchmarking, and optimizing SDK/CLI/MCP documentation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* docs: add implementation plan for SKILL folder
7-task plan covering branch creation, all 5 markdown files with
complete content, verification, and PR creation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(skill): add SKILL.md entry point with context detection and routing
* feat(skill): add setup reference — prerequisites, init wizard, discovery
* feat(skill): add benchmark reference — run, interpret, diagnose, compare
* feat(skill): add optimize reference — loop mechanics, safety, troubleshooting
* feat(skill): add config reference — schema, models, scope, errors
* fix(skill): address Copilot review — config path, safety boundary, minImprovement, gitignore dev docs
* docs(skill): update routing concept; add direct-provider auth to prerequisites; fix dot-notation
* docs(skill): expand compare section in benchmark reference
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(prompt): add prompt surface type for skill/template benchmarking (#20)
* docs: fix repo URL in CONTRIBUTING.md (bucurdavid → fastxyz)
Fixes #18
* feat(prompt): add prompt surface type for skill/template benchmarking
Discovers phases, instructions, output formats, and decision points from
markdown skill files. Evaluates model responses with content-based criteria
(required sections, format patterns, keywords, structural elements) instead
of tool-call matching.
New files:
- src/project/discover-prompt.ts — markdown parser for capabilities
- src/benchmark/prompt-evaluator.ts — content-based scorer (weighted)
- src/discovery/prompt.ts — file-based discoverer + types
- tests/smoke-discovery-prompt.ts — 10 tests
- tests/smoke-prompt-evaluator.ts — 15 tests
- tests/fixtures/sample-skill.md — 3-phase sample
Modified: types.ts, validate.ts, snapshot.ts, runner.ts, evaluator.ts,
extractors/index.ts, prompts.ts, cli.ts, init/, CHANGELOG.md, README.md
* fix(prompt): guard snapshot surface, fix format inconsistency, remove dead interface, fix stale comment
* fix(discovery): anchor frontmatter closing delimiter to line boundary
Replace indexOf('---', 3) with a regex that matches --- only at the start
of a line, preventing early termination when a YAML value contains ---.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(discover-prompt): clear preamble on first heading, strip frontmatter before section split
* fix(discover-prompt): fallback to whole-content instruction extraction for heading-less files
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(discover-prompt): document interface difference between production and standalone paths
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(prompt): wire evaluatePromptResponse into runner, expose section text, fix generate schema
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(prompt): unique temp dir + cleanup guard in test, enforce empty actions in grounding, bypass maxTasks constraint for prompt
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* Refactoring & Removal of legacy code (#15)
* refactor: DRY deduplication across LLM format handlers and Pi task helpers
Extract shared retry logic (isAbortError, isRetryableError, sleep), tool-result
truncation, and agent-loop accumulator helpers into src/benchmark/llm/shared.ts,
eliminating identical code that existed in both openai-format.ts and
anthropic-format.ts.
Extract the Pi completeSimple call sequence (timeout setup, error check, text
block extraction) into src/tasks/pi-simple-complete.ts, used by both
default-pi-generator.ts and default-pi-critic.ts to replace ~40 lines of
near-identical implementation in each file.
Net: -129 lines of duplication, no behavior changes, all tests pass.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): consolidate duplicate type definitions
- DiscoveredActionArg (discovery/types) was identical to ActionArgSchema
(actions/types); replaced with a re-export alias to eliminate the copy
- BenchmarkSurface now aliases ActionSurface instead of re-declaring the
same 'sdk' | 'cli' | 'mcp' union literal
- Removed dead SurfaceActionArg alias from benchmark/types (never imported)
- Removed dead SurfaceSnapshotArg alias from project/types and project/index
(exported but never consumed)
- Replaced all inline import('...').PiAuthMode dynamic-import type refs
with proper named import type statements in pi-format, models,
coding-orchestrator, benchmark-agent, default-pi-critic, default-pi-generator
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove unused code found by knip
- Delete src/runtime/pi/benchmark-agent.ts (createReadOnlyBenchmarkModel was dead — exported but never imported)
- Delete src/discovery/index.ts and src/optimizer/mutation/types.ts (barrel re-exports with no consumers)
- Remove requireApiKeyFromEnv from auth.ts (never called internally or imported from outside)
- Remove 5 unused re-export lines from benchmark/extractors/index.ts (extractCodeBlock, extractShellBlock etc. — tests import directly from sub-files)
- Unexport 11 internal-only types/interfaces (MAX_TOOL_RESULT_CHARS, buildConfigFromAnswers, LLMClient, ToolNameAliasCodec, PromptSurface/Options, DoctorOptions, DetectedProject, BenchmarkAdapterRunOptions/Result, FeedbackPackage, RawExtraction, PiSimpleCompleteOptions/Input)
- Remove @mariozechner/pi-agent-core from package.json dependencies (transitive dep, never directly imported)
- Add smoke-actions.ts to the npm test script (17 passing tests were orphaned from the suite)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(deps): break circular dependency between actions/discover and benchmark/config
Move loadCliCommands and loadMcpTools from benchmark/config.ts into a new
actions/loaders.ts file, resolving the cycle:
actions/discover → benchmark/config → project → project/snapshot → actions/discover
benchmark/config.ts re-exports both functions for backward compatibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* refactor(types): replace weak types with precise types throughout
- Replace `Model<any>` with `Model<Api>` and specific API literal types
(`Model<'openai-completions'>`, `Model<'openai-codex-responses'>`) in
models.ts and pi-format.ts
- Replace `(err as any).status` with typed intersection cast
`Error & { status: number }` in anthropic-format.ts and openai-format.ts
- Remove redundant `as unknown[]` cast on `session.state.messages` and
import `AgentMessage` from `@mariozechner/pi-agent-core` in pi-coding.ts;
update `extractLatestAssistantText` and `extractToolActivity` parameter
types accordingly
- Remove unnecessary double casts `(c as unknown as { method: string }).method`
in failure-details.ts and passing-failing-diff.ts — `ExtractedCall` is
`ActionAttempt` which already has `method: string` and `args`
- Remove `as any` from `Type.Unsafe(...)` call in pi-format.ts — the argument
type is compatible with `UnsafeOptions` (SchemaOptions index signature)
- Fix `catch (error: any)` to `catch (error)` in two snapshot files,
using `instanceof Error` guard to access `.message`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: remove unnecessary optional chaining on non-optional reportConfig
reportConfig is always defined (report.config is non-optional per
BenchmarkReport). The ?. on reportConfig before .outputDir was
defensive against a value that TypeScript knows can never be null.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove deprecated and legacy code paths
- Remove toLegacyOptimizeManifest (never called anywhere in src or tests)
- Remove deprecated SdkSurfaceConfig fields: classes, functions, functionReturns, methods
- Remove dead config.sdk?.methods fallback in runner.ts (no config ever sets methods)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove AI slop comments and stubs
Removed or rewrote ~75 comments across 15 files:
- Dropped "// NEW:", "// Transitional alias", numbered step comments (1-14
in runner.ts), and other in-motion migration language
- Removed redundant comments that restated the immediately-following code
- Replaced multi-line comment pairs with single consolidated explanations
where the context was non-obvious and worth preserving
- Kept all comments explaining invariants, non-obvious behavior, or
security rationale (e.g. git injection safety, OpenAI dot-format exemption)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove backwards-compatibility shims — SemVer allows breaking changes
- Remove LEGACY_PROJECT_CONFIG_NAME and skill-benchmark.json detection in load.ts
- Remove legacy-config-name warning from validate.ts (old filename alongside new one)
- Remove E_LEGACY_CONFIG error definition from errors.ts and its docs entry
- Remove DiscoveredActionArg deprecated alias from discovery/types.ts; update all four
discovery modules (cli, mcp, sdk, optique) to import ActionArgSchema directly
- Remove CodeModeConfig, McpModeConfig, ExpectedTool, ToolMatch backward-compat
type aliases from benchmark/types.ts and benchmark/index.ts
- Remove toolMatches dual-field from TaskResult; make actionMatches required
- Remove unnecessaryCalls/hallucinatedCalls alias fields from TaskResult.metrics;
consolidate to unnecessaryActions/hallucinatedActions throughout
- Remove method? alias from ExpectedAction; make name required
- Remove expected_tools? alias from TaskDefinition; make expected_actions required
- Simplify getExpectedActionName and getExpectedActions helper functions
- Remove deprecated extractFromCode function from code-analyzer.ts
- Update all internal consumers (optimizer, evaluator, tasks, benchmark) to use
canonical field names without the ?? fallback patterns
- Update all test files to use canonical field names and types
- Rewrite smoke-code.ts Group 2 tests for extractAllFromCode (the modern API)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address Copilot review issues from simplify PR
- Restore @mariozechner/pi-agent-core as explicit dependency (still
imported as a type in src/optimizer/mutation/pi-coding.ts; relying on
transitive availability is fragile)
- Fix piSimpleComplete: throw unconditionally when stopReason === 'error'
(previously silent fallthrough when errorMessage was falsy) and throw
with content-type diagnostic when the response has no text blocks
- Remove redundant empty-text guard in default-pi-generator.ts now that
piSimpleComplete handles that case itself
- Replace unsafe `as` casts for verify/expected_fetches in config.ts with
Array.isArray guards before assignment
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address secondary reviewer issues
- loadReport: detect pre-1.x report format (toolMatches field) and throw
a clear error instead of letting callers hit silent undefined-access later
- loadProjectConfig: when skill-optimizer.json is missing, check for
skill-benchmark.json and surface a targeted rename hint instead of the
generic "config not found" error
- CHANGELOG: add BREAKING CHANGES section documenting all removed public
API exports and renamed TaskResult fields
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename minOverallPassDelta, fix mock repo configs, remove ghost repo names
- Part F: rename `minOverallPassDelta` → `minImprovement` in optimizer
types, loop resolution, adapters, and tests to unify with the public
ProjectOptimizeConfig field name
- Part G: fix dot-notation model IDs (claude-sonnet-4.6 → 4-6,
gpt-5.4 → 5-4) in all three mock-repo skill-optimizer.json files;
delete stale legacy benchmark.config.json and optimize.config.json
from mcp-tracker-demo (superseded by skill-optimizer.json)
- Part H: remove ghost names sdk-demo, cli-demo, mcp-demo from
MOCK_REPO_NAMES and update materialize-mock-repo.ts usage string to
only list the three templates that actually exist
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove expected_tools/method input shims and benchmark.tasks deprecated field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: reapply Copilot comment fixes that were lost before agent commits
- loadProjectConfig: only check for skill-benchmark.json when configPath
was not explicitly supplied; use dirname() for clarity
- default-pi-critic: catch no-text-blocks error and return '[]' (graceful
degradation) while re-throwing real provider errors
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename matchTools→matchActions, remove pre-1.x report guard, clean straddling test comments
- Renamed exported matchTools → matchActions in evaluator.ts; updated
internal vars (toolMatches→actionMatches, unnecessaryCalls→unnecessaryActions,
hallucinatedCalls→hallucinatedActions) for consistency
- Stripped pre-1.x schema guard from loadReport in compare.ts — old format
check on 'toolMatches' field was pure backwards-compat dead code
- Removed "Accept either old or new error" branching from two smoke tests;
both now only assert the current validation path
- Updated CLAUDE.md: manifest fallback is a supported feature, not a
transitional shim
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove skill-benchmark.json legacy filename check from loadProjectConfig
The check for the old config filename was a migration hint for users
upgrading from a pre-1.x version. Current config is skill-optimizer.json;
drop the dead code.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: rename misleading 'legacy' test names in smoke-actions
Test names said 'legacy' and 'legacy shape' but were testing current
SurfaceSnapshot ↔ ActionCatalog conversion behaviour. Rename variables
and assertions to remove the legacy framing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: remove/reword remaining legacy and old language
- README.md: remove skill-benchmark.json migration hint (the detection
code was removed in the previous commit)
- cli.ts: rename example path results/old/ → results/baseline/ (clearer)
- snapshot.ts: rephrase error from "old format" to "not supported"
- smoke-actions.ts: align test name and assertion strings with the
updated error message; drop "legacy" framing from assertion strings
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: final cleanup — remove unused exports, trivial comments, stale doc entries
Unused exports:
- Make extractPythonSdk/extractRustSdk/extractTypeScriptSdk unexported;
they are private implementation details assigned to adapter objects in
the same file, not part of the public API
Trivial JSDoc removal:
- shared.ts: remove one-liner comments that restate function names
- reporter.ts: remove pad/center JSDoc; remove redundant generateMarkdown doc
- anthropic-format.ts: remove comment restating toAnthropicTool's name
- openai-format.ts: remove HTTP-method comments that add no information
Docs:
- README.md: fix benchmark.models description to cover all three providers
(openrouter, anthropic, openai) — not OpenRouter only
- CLAUDE.md: fix layer count (5, not 4); fix dot-notation validator
description (skips openai/, not just openrouter/)
- SKILL/references/setup.md: remove troubleshooting row for
skill-benchmark.json detection that was removed last commit
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: correct inaccurate CHANGELOG claim and wrong GitHub URL in CONTRIBUTING
- CHANGELOG: loadReport does not validate old field names (the validation
guard was intentionally removed); update wording to match actual behavior
- CONTRIBUTING: fix clone URL from bucurdavid fork to fastxyz canonical repo
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: repair three regressions in the generate-tasks → run workflow
Issue A (freeze.ts): surface.snapshot.json was written as a plain
{surface,actions} SurfaceSnapshot, which loadSurfaceSnapshotFile now
rejects. Switch to writeActionSnapshotFile + fromSurfaceSnapshot so
the file is written in the versioned {version,catalog} format.
Issue B (shared.ts): truncateToolResult sliced to MAX_TOOL_RESULT_CHARS
then appended the 16-char suffix, producing 50,016-char strings.
Subtract TRUNCATED_SUFFIX.length from the slice index so the final
string is exactly MAX_TOOL_RESULT_CHARS.
Issue C (types/resolve/adapters): benchmark.tasks was removed from
ProjectBenchmarkConfig during the simplify cleanup. freeze.ts still
writes tasks: tasksPath into benchmark.generated.json, but the field
was silently dropped on load and adapters.ts hard-coded
tasks: '__generated__', causing the runner to throw "run generate-tasks
first" even after generate-tasks had succeeded. Re-add tasks?: string
to both benchmark config interfaces, thread it through resolve.ts, and
use project.benchmark.tasks ?? '__generated__' in adapters.ts.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* typos
* license typo
* fix(validation): restore missing-tasks guard, validate verify/expected_fetches elements, document tasks.json migration
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: add prompt surface to init comment, E_INVALID_SURFACE fix text, and config-schema reference
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(release): bump version to 1.1.0, update CHANGELOG and docs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Matías Selser <mati.selser1995@gmail.com>
Co-authored-by: OpenClaw Agent (basd) <basd@openclaw.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Chris <25006584+seichris@users.noreply.github.com>
Co-authored-by: Matías Selser <mati.selser@hotmail.com>
* cleanup
* fix(tests): update smoke-code assertions to expect openrouter/ dots preserved
Update three test assertions (two in smoke-code.ts, one in smoke-init.ts)
that were checking the old behavior where openrouter/ model ID dots were
rewritten to hyphens. Now that fix.ts and validate.ts both exempt openrouter/
IDs from dot-rewriting, the tests reflect the correct invariant: openrouter/
slugs are passed verbatim to the OpenRouter API and must not be modified.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tests): use correct openrouter/ claude slug with dot in smoke-llm
* test(model-ids): add regression test for openrouter/ dot preservation
Adds smoke-model-ids.ts verifying that the validate+fix pipeline preserves
dots in openrouter/ model IDs, rewrites dots in ant…
No description provided.