模块状态(step B 收口): level-up 当前 = 冻结的执行环 + 活的机器源模块。 执行环(autopilot/strategy/evaluator/apply/dev-loop/self-review/runner/metric/repair-adapter/redline)已冻结废弃,与 looper 执行环重复,牌面保留 looper;勿新增对它们的依赖。 活的机器源(auto-research/util/outpact-adapter/machine-intake)独立,不依赖执行环,可被外部复用。 详见
docs/STATUS.md。
level-up is an agent-facing autoresearch runtime for L3 local autopilot work.
It is not a giant skill. It is a thin experiment loop with optional capability slots. An agent may use or skip each slot, but every skip must be explicit.
L3 local autopilot can:
- inspect a target repository;
- turn a user goal into a goal contract;
- generate experiment candidates;
- create an isolated git worktree;
- implement one experiment per round;
- run validation and evaluators;
- self-review the result;
- keep improved experiments as commits;
- discard failed or no-op experiments;
- stop at max rounds, blockers, or no-improvement thresholds.
L3 local autopilot cannot:
- merge;
- deploy;
- modify global config;
- touch secrets, billing, ads, production data, or unrelated repositories;
- silently skip hard gates.
interview -> goal contract -> strategy -> worktree experiment -> validation -> evaluator -> review -> keep/discard/crash -> ledger -> PR/MR packet -> Chinese run report
The core runtime owns the loop, state, ledger, and safety boundaries. Slots add domain-specific capability:
interview: lightweight front gate for only high-impact decisions; replaces defaultgrill-me.ideation: divergent experiment generation.strategy: choose the next untried candidate, adapt after failed rounds, generate repair candidates with safe apply plans, and record why.metric: scoring for performance, UI, tests, code health, or custom goals.evaluator: turn apply, validation, review, and worktree delta into keep/discard evidence for the next strategy step.repair-adapter: turn validation/review failure evidence into targeted repair proposals and bounded safe apply plans.review: self-review or review-hub style independent review.recovery: nsr-lite milestones, next action, and resume state.policy: hard gates, forbidden actions, and human approval boundaries.runner: current session now; future opencode/MMS/external model process adapters.apply: structured worktree mutation via command, patch, or file-write manifests.notify: Feishu/GitHub/GitLab notification adapters after PR or MR creation.redline: optional PR/MR merge-readiness evidence adapter through siblingredline-guard.redline-final: final pre-merge gate wrapper; onlymergeablepasses, and it never approves, merges, deploys, or force-pushes.cleanup: remove clean merged experiment worktree folders after PR/MR merge.
Preferred agent-facing entry:
用 level-up 升级 /path/to/project,目标是优化首页加载速度和事件逻辑。
The agent should inspect the target repo, ask only blocking questions, create the run contract, use isolated worktrees, run validation, keep/discard experiments, open a PR/MR when there is a useful change, notify Feishu when configured, and leave a Chinese report for the user. The user should not need to manually drive the CLI.
Initialize a run for a local project:
npm run level-up -- init --target /path/to/project \
--goal "Make the homepage faster without changing product behavior" \
--metric "Improve mobile LCP while keeping build, tests, and SSR safe-access green"Run the practical L3 loop:
npm run level-up -- run --run /path/to/project/.level-up/runs/<run-id> --execute --pr-packlevel-up run ensures scan, ideas, work-pack, baseline validation, an isolated worktree, experiment/final validation, deterministic self-review, ledger recording, and optional PR evidence. If the round makes no change or fails validation/review, it records discard instead of pretending the attempt worked. Adaptive rounds can turn validation/review failures into focused repair candidates; synthetic repair candidates use their own targeted repair proposal and safe apply plan instead of repeating the failed input. A narrow validation repair can execute a safe command for git diff --check whitespace failures.
Run multiple experiments under a budget instead of a single round:
# Keep experimenting until a 5-minute wall-clock budget runs out,
# or 3 consecutive rounds fail to improve (whichever comes first).
npm run level-up -- run --run /path/to/project/.level-up/runs/<run-id> \
--execute --budget 5m --max-no-improvement 3 --pr-packStop conditions default from the goal contract's stopConditions (maxRounds, maxWallClockMs, maxMinutesPerRound, maxNoImprovementRounds) and are overridden per run by --rounds, --budget, and --max-no-improvement. The summary records stopReason (rounds-exhausted, budget-exhausted, no-improvement, round-timeout, blocked), budgetMs, elapsedMs, and noImprovementRounds. To let a numeric metric decide keep/discard, write metric-baseline.json at the run root and metric.json in each experiments/round-NNN/; each round is scored against the best kept value so far (the incumbent), and the runtime advances metric-incumbent.json after every keep. Absent those files, the binary gates decide. See docs/experiment-loop.md.
Generate the same loop with a user-readable Chinese report:
npm run level-up -- run --run /path/to/project/.level-up/runs/<run-id> --execute --pr-pack --reportRun a round with a structured apply adapter:
npm run level-up -- run --run /path/to/project/.level-up/runs/<run-id> \
--apply-patch /tmp/experiment.patch \
--execute --pr-pack --reportlevel-up also supports --apply-write-file <path> --apply-content <text> for small generated files and keeps --apply-command <cmd> for narrow local commands. Unsafe command patterns are blocked before validation.
Generate or refresh a report for an existing run:
npm run level-up -- report --run /path/to/project/.level-up/runs/<run-id> \
--link "https://github.com/org/repo/pull/123" \
--notify-status "Feishu 已通知"The report is written to REPORT.zh.md inside the run root and summarizes what happened, why, experiment results, metric evidence, validation, PR/MR links, Feishu status, and next step.
Generate a runner packet for the current session or a future model process:
npm run level-up -- runner-pack --run /path/to/project/.level-up/runs/<run-id> \
--runner current-session \
--runner-profile codex-session \
--skills level-up,interview \
--mcp github,browserThe current recommended mode is hybrid: the Codex/MMS session acts as the model runner, while level-up records runtime state, validation, self-review, ledger, and PR evidence. Future opencode-profile and mms-runner adapters should consume the same packet.
Notify Feishu after a PR or MR is created.
Clean up merged worktree folders after a PR or MR is merged:
npm run level-up -- cleanup-worktrees --repo /path/to/repo --base-ref origin/main --execute
npm run level-up -- post-merge --repo /path/to/repo --base-ref origin/main \
--run /path/to/project/.level-up/runs/<run-id> --execute --delete-branches \
--prune-branches --branch-prefix codex/The cleanup command skips the current worktree, protected branches, dirty worktrees, and worktrees whose HEAD is not already merged into the base ref. Without --execute, it only reports what would be removed. Add --delete-branches only when the local merged branch reference should be removed after the worktree folder is removed. post-merge wraps the same safety checks and writes POST_MERGE_CLEANUP.zh.md plus post-merge-cleanup.json when --run or --output-dir is provided. Branch pruning is off by default; use --prune-branches --branch-prefix codex/ only for merged agent-owned local branches.
Run the final redline-guard pre-merge gate after a PR/MR exists:
npm run level-up -- redline-final --run /path/to/project/.level-up/runs/<run-id> \
--url "https://github.com/org/repo/pull/123" \
--validate --notifyYou can also attach it while refreshing the Chinese report:
npm run level-up -- report --run /path/to/project/.level-up/runs/<run-id> \
--link "https://github.com/org/repo/pull/123" \
--redlineredline is evidence-only; redline-final is the final pre-merge gate. Both first look for a configured --redline-bin or LEVEL_UP_REDLINE_BIN, then a sibling ../redline-guard/src/cli.mjs, then redline-guard on PATH. By default level-up passes --evidence <run-root> so the adapter can inspect local run artifacts; pass --evidence false when evidence should be omitted. For final pre-merge use, prefer redline-final; only mergeable passes finalGateStatus. needs-review, blocked, unknown, missing URL, or adapter failure must stop before merge. --notify does not request PR/MR comments; comments are never posted unless --comment is explicitly passed.
npm run level-up -- notify \
--channel feishu \
--repo level-up \
--branch "codex/example -> main" \
--title "perf: 优化首页首屏加载和事件逻辑" \
--link "https://github.com/CtriXin/level-up/pull/5" \
--status "check/build/self-review 通过" \
--effect "首页主 JS gzip 下降"The webhook must come from runtime environment such as FEISHU_WEBHOOK_URL; do not commit webhook URLs.
Step-by-step commands remain available for debugging or manual control:
npm run level-up -- scan --run /path/to/project/.level-up/runs/<run-id>
npm run level-up -- ideas --run /path/to/project/.level-up/runs/<run-id>
npm run level-up -- work-pack --run /path/to/project/.level-up/runs/<run-id>
npm run level-up -- runner-pack --run /path/to/project/.level-up/runs/<run-id> --runner current-session
npm run level-up -- worktree --run /path/to/project/.level-up/runs/<run-id>
npm run level-up -- dev-loop --run /path/to/project/.level-up/runs/<run-id> --phase baseline
npm run level-up -- dev-loop --run /path/to/project/.level-up/runs/<run-id> --phase final --execute
npm run level-up -- record --run /path/to/project/.level-up/runs/<run-id> --status keep --score 84.2 --description "Lazy-load non-critical hero media"
npm run level-up -- pr-pack --run /path/to/project/.level-up/runs/<run-id> --visual
npm run level-up -- report --run /path/to/project/.level-up/runs/<run-id>The CLI is only the fallback renderer. The durable contract is the files under .level-up/runs/<run-id>/.
level-up can bind a run to a canonical state-core task when Mommy hands off a task_id.
STATE_CORE_DIR=/Users/xin/auto-skills/CtriXin-repo/state-core \
npm run level-up -- init --target /path/to/project --task-id <task-id>The adapter resolves state-core from STATE_CORE_DIR, falling back to a sibling ../state-core, and calls python3 <state-core>/src/cli.py. Node never imports Python code directly.
Binding behavior:
- init reads
task-state.jsonthroughcli.py read, usesintent.goal/intent.rawas the level-up goal, and setsrunner=level-up; - keep records report
passto the related slot (verifyfor medium/small,executorfor large); - discard/crash records report
fail, producing a state-core blocker; - finalize writes
ledger_refto the.level-up/runs/<run-id>/root and advances the canonical phase toverifying; doneremains state-core's decision throughcli.py advance --phase done; if the gate blocks, level-up reports the unmet slots instead of declaring completion.
Canonical truth lives in state-core. .level-up/runs/* is runtime state and evidence for the loop, not a replacement for task-state.json.
Use interview as the default intake slot: inspect local evidence first, ask 1-3 questions only when answers change objective, metric, guardrails, irreversible scope, or human gates, and accept defaults / 你定 / 先做.
Keep grill-me as an explicit deep stress-test mode for strategy or design branches that are too risky to infer.
docs/ Product and runtime design
schemas/ Renderer-neutral JSON schemas
skills/level-up/ Agent entry skill
src/ Minimal dependency-free local runtime
tests/ Node test runner coverage