English Β· δΈζ Β· EspaΓ±ol Β· PortuguΓͺs Β· ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯
Second-model AI approval for DeepSeek Harness β the Codex approvals_reviewer=auto_review / Claude Code auto mode pattern, built as a pure Cordis plugin.
When an agent's action crosses the sandbox boundary, a read-only reviewer subagent decides allow/deny β with a reason β so humans approve nothing while nothing unsafe slips through.
Zero human operations. The request goes to the AI reviewer, the verdict is allow/deny + reason + risk level, and every decision is reconstructable from the session log: approval/asked β autoReview/verdict (or autoReview/rejection for hard disables) β approval/decided.
One real evidence run (real server, real API key, two real model rounds): the AI reviewer allows an escalated workspace write (risk low, 5.2 s), then denies a recursive out-of-workspace delete (risk high, 8.9 s) β the deny reason is fed back to the model, visible in the transcript.
Pattern-based auto-approvers decide before dispatch, with no evidence. dsh-auto-review gives the decision to a reviewer subagent that reads the actual workspace (through its read-only tool face), the already-streamed tool-call arguments (sensitive values redacted), the request reason, and your risk rules β then returns a structured verdict. A deny verdict feeds its reason back to the calling model, so the agent learns why instead of retrying blindly.
| π Official seam | An answerer on the approval/request waterfall. Requests it does not own are delegated via next() β the human approval flow is never short-circuited. |
| π§ Second-model verdict | One-shot fork subagent with a read-only tool allow-list (read/glob/grep) and a structured verdict schema { decision, reason, riskLevel }. |
| π‘οΈ Fail closed | Reviewer crash, timeout, or schema mismatch never opens the gate: fallbackPolicy applies, default rejected. |
| π§© Config-driven routing | Per-tool policies (ai/human/never) + regex risk rules, all changeable from cordis.yml. |
| π¬ Deny reasons reach the model | The reviewer's reason is injected into the denied tool result (callId-linked), so the agent adapts. Fail-closed fallback rejections and never-policy hard disables inject auditable failure texts too ([auto-review] / [auto-review-fallback] / [auto-review-never] markers). |
| π Full audit trail | autoReview/verdict + autoReview/rejection session events (reviewer identity, verdict, reason, risk, duration) + an invariant companion enforcing model-visible βΊ logged. |
| π No recursion | Reviewer asks are recognized by identity and delegated; maxDepth + the tool allow-list keep the reviewer non-delegating. |
| π§― Rejection circuit breaker | 3 consecutive denials (or 6 within the last 10 verdicts) in one turn trip the breaker: later requests delegate, reject, or abort the turn β no endless denial loops. |
| ποΈ Risk-level policy | An allow verdict whose risk exceeds riskPolicy.maxAutoAllow never settles the request: it delegates to a human or denies. |
| β One-shot human override | /auto-review approve [n] authorizes ONE retry of a recent denial; the next same-tool review carries that authorization as reviewer context (the reviewer still decides). |
| π Reviewer context | Optional compact transcript (recent messages and tool results, bounded) + a Codex-style Markdown reviewerPolicyText ruling policy. |
| β¨οΈ Session command | `/auto-review on |
| π₯οΈ Web review panel | A session-header panel (Web GUI) shows the switch (with on/off buttons), both per-turn budgets, cumulative statistics, the circuit trip, recent verdicts, and one-shot approve buttons β driven by the autoReview session projection. |
approval/request waterfall (answerer chain)
β
βββββββββββββββββββββββββ΄βββββββββββββββββββββββ
β dsh-auto-review answerer β
β Β· session enabled? Β· policy = ai? β no ββ next() βββΆ human answerer (UI)
β Β· risk rules β toolsPolicy β default β
βββββββββββββββββββββββββ¬βββββββββββββββββββββββ
β yes
βΌ
βββββββββββββββββββββββββββββββββββββ
β reviewer subagent (fork, one-shot)β
β Β· toolFilter: read/glob/grep β
β Β· outputSchema: {decision, β
β reason, riskLevel} β
β Β· timeout + req.signal abort β
βββββββββββββββββ¬ββββββββββββββββββββ
β verdict / failure (fail-closed fallback)
βΌ
allow β allowed-once deny β rejected + reason injected into the
denied tool result (callId-linked)
β never β rejected + [auto-review-never] feedback
β (hard disable, no reviewer runs)
βΌ
audit: approval/asked β autoReview/verdict | autoReview/rejection
β approval/decided (session events, log-only, invariant-checked)
Composition order. The answerer runs at its registration position in the waterfall: if a human UI answerer is composed BEFORE the auto-review row, humans answer first and the reviewer only sees what is delegated downstream. Verify with dsh --profile <name> --dump-config and place the auto-review row before your human answerer rows when you want ai-policy tools routed to the reviewer first.
Three install channels; the plugin is a bundle ("dsh": { "bundle": { "patch": "./cordis.patch.yml" } }).
# 1. npm (published artifact, no build step)
dsh plugin --profile web add dsh-auto-review
dsh --profile web # restart
# 2. npm tarball (built artifacts, offline install)
pnpm pack # β dsh-auto-review-<version>.tgz
dsh plugin --profile web add ./dsh-auto-review-<version>.tgz
dsh --profile web # restart
# 3. git source (pin the commit; self-contained `prepare` builds it)
# pnpm β₯ 10 blocks lifecycle builds: add the printed allowBuilds key
# to the profile's pnpm-workspace.yaml first.
dsh plugin --profile web add "github:PerryLink/dsh-auto-review#<commit>"
# 4. local link (development)
dsh plugin --profile web add link:/path/to/dsh-auto-reviewVerify:
dsh --profile web --dump-config | grep -A4 'id: auto-review'Out of the box the shipped patch AI-reviews bash and write; every other tool (including edit β in-place modification) delegates to the human chain. Add edit: ai explicitly if you accept in-place edits without a human in the loop.
All tunables are Schemastery Config fields (changeable from cordis.yml). An id-targeted override replaces the whole config row β restate every key you need.
| Key | Default | Meaning |
|---|---|---|
enableByDefault |
true |
Sessions start with auto-review enabled; /auto-review on|off writes a durable override that beats this |
toolsPolicy.default |
human |
Policy for unlisted tools (delegate to the human answerer) |
toolsPolicy.overrides |
{} |
Per-tool policy: ai (reviewer decides), human (force human), never (deterministic reject) |
riskRules |
[] |
{pattern, policy, field?} matched (first match wins) before the tool table; field selects reason (default), toolName, or arguments (the redacted presented call arguments) |
reviewerProvider |
fork |
Subagent provider for the reviewer (in-process fork backend) |
reviewerModel |
(inherit) | Reviewer model id; unset inherits the session agent's route |
reviewerTimeoutMs |
60000 |
Verdict deadline; on expiry the fallback policy applies |
reviewerTools |
[read, glob, grep] |
The reviewer child's tool allow-list (must be non-empty) β everything else is invisible there |
fallbackPolicy |
rejected |
Reviewer failure: rejected (fail closed), delegate (continue the chain), allow-once (grant β see Security). Renamed from allow-readonly in 0.2.0; the old spelling fails loudly |
maxReviewsPerTurn |
10 |
Real AI-verdict budget per open turn; beyond it, requests delegate to humans |
maxFailuresPerTurn |
10 |
Reviewer-failure budget per open turn (timeout/unavailable/schema, not cancellations); beyond it, requests delegate instead of paying another full timeout. Defaults to maxReviewsPerTurn |
reasonMaxChars |
2000 |
Cap for reviewer reasons, the request reason, and the redacted argument preview |
reviewerGuidance |
(none) | Optional advisory guidance appended to the reviewer prompt |
reviewerPolicyText |
(none) | Markdown ruling policy injected into the reviewer prompt (Codex-style; template at fixtures/config/policy-template.md) |
denyGuidance |
(anti-circumvention text) | Guidance appended to every injected deny reason |
contextBudget |
{turns: 0, maxChars: 4000} |
Compact transcript budget for the reviewer prompt; turns: 0 disables |
riskPolicy |
{maxAutoAllow: high, onHighRisk: delegate} |
allow verdicts above maxAutoAllow delegate (delegate) or deny (deny) |
circuitBreaker |
{consecutiveDenies: 3, windowDenies: 6, windowSize: 10, action: delegate} |
Rejection circuit breaker; trips on 3 consecutive denies or 6 of the last 10 verdicts in a turn; action: delegate / reject / abort-turn |
overrideTtlMs |
300000 |
How long a /auto-review approve override stays usable |
language |
en |
UI language of the /auto-review command output (en | zh) |
Example (annotated full form: fixtures/config/config-full.yaml):
- insert:
- id: auto-review
name: dsh-auto-review
config:
toolsPolicy:
overrides: { bash: ai, write: ai }
riskRules:
- pattern: '(?i)(rm\s+(-[a-z]+\s+)*/|git\s+push\s+--force)'
policy: never
- pattern: 'write'
policy: never
field: toolName
reviewerTimeoutMs: 30000
fallbackPolicy: delegate
riskPolicy: { maxAutoAllow: medium, onHighRisk: delegate }
circuitBreaker: { consecutiveDenies: 3, windowDenies: 6, windowSize: 10, action: delegate }/auto-review on|off|status|approve [n]
on/off append the durable autoReview/state override (the fold survives restart/resume β replay IS the state) and inject a switch notice the model sees (logged as a user/message event). status reports the effective state, both per-turn budgets (AI verdicts and reviewer failures), a tripped circuit breaker when one is active, and the session's cumulative statistics (allows/denies/fallbacks/never rejects, mean duration, recent verdicts). approve [n] records a single-use autoReview/override for the n-th most recent denial (1 = most recent): the next same-tool review within overrideTtlMs carries the authorization as reviewer context β the reviewer still decides, and the override is consumed by that review regardless of its outcome.
In the Web GUI (web profile), the package contributes a session-header action (AI Review) that opens a panel with the session's auto-review state: the switch with on/off buttons (they execute /auto-review on|off), both per-turn budgets, cumulative statistics (including hard-disable rejections), the circuit trip, the recent verdicts, and one-shot approve buttons for recent denials (they execute /auto-review approve [n]).
How it is wired:
- The host registers an
autoReviewsession projection (folded from the log-onlyautoReview/*events) and serves it through the session-projection channel. - The browser half is a client module (auto-discovered from the
dsh.clientdeclaration) registered on theconversation.session.header.actionsseat. - No extra patch rows are needed: the panel loads whenever the plugin is installed in a profile whose web build provides the session-projection capability (the web profile does). Without that capability the panel reports itself unavailable; the answerer is unaffected.
The panel reads only whole projection values β it never receives the raw session event stream.
- The reviewer runs in a read-only tool face (
toolFilterallow-list). It cannot write, edit, run bash, fetch the network, or delegate (maxDepth= its own depth). Its session log is persisted and auditable. - Sensitive arguments are redacted (key-name matching:
token,password,api_key,Authorization, credentials, private keys β¦) before entering the reviewer prompt; the plugin never executes the reviewed arguments. Redaction is key-based, not content-based β do not AI-review tools whose argument values you cannot afford to show a model. - Fail closed by default. Every abnormal path (provider missing, capability gaps, start rejection, timeout, non-
completedstop reason, missing/malformed verdict, audit-correlation failure) resolves throughfallbackPolicy, defaultrejectedβ and the rejection feeds an auditable reason back to the model instead of the generic "user rejected" text.allow-oncegrants unconditionally β it exists only for unattended deployments whose admin accepts that risk. - Hard disables explain themselves. A
nevertool or risk rule rejects deterministically AND records a log-onlyautoReview/rejectionevent with the matched rule/table entry, then injects a[auto-review-never]marker text into the denied tool result β the model learns the action is hard-disabled instead of retrying it (invariant-checked: marker βΊ event). - Rejection circuit breaker. A run of denials in one turn trips the breaker (
consecutiveDenies/windowDeniesinsidewindowSize), recorded as a log-onlyautoReview/circuitevent; later requests follow itsaction(delegate/reject/abort-turn).abort-turninjects a model-visible warning and cancels the agent. - Reviewer context is presented transcript.
contextBudgetfeeds already-presented session content (messages, tool results) to the reviewer. With the default same-route reviewer model that content stays inside one provider; configurereviewerModelto a different provider only if you accept presenting that transcript to it. neveris one-way at this layer. Anevertool or risk rule rejects before the human chain sees the request β a lockdown knob, not a default.- The reviewer is a model. Its verdicts are advisory policy, not a security kernel. Prefer
human/neverrules for irreversible operations.
- The reviewer needs a working LLM route (inherited from the session agent by default); without one every review falls back per
fallbackPolicyβ never a silent grant. reviewerToolsnames must exist as global tools in the profile; an unknown name fails the reviewer child loudly at the earliest point and falls back.- Risk rules match the request
reason, thetoolName, or the redacted callargumentsper theirfield; other conditions belong intoolsPolicy.overrides. - The
/auto-review approveoverride authorizes the next same-tool review, not the exact historical call; a different action on the same tool consumes it. - The verdict events are log-only; the dedicated Web review panel reads the folded
autoReviewprojection (the raw event stream never reaches browser plugins). autoReview/stateandautoReview/verdictare appended with the envelope'signorable: truemarker, so any harness build loads the log β readers that do not know the out-of-repo types simply skip those records instead of refusing the session. (rc.6 hosts accept and ignore the marker, keeping the exact pre-marker behavior; sessions written by pre-0.1.1 versions can be repaired withscripts/repair-session-logs.mjsfromdsh-permission-rules.)- The git channel needs the single
allowBuildskey thedshCLI prints fordsh-auto-reviewitself. The repo ships its ownpnpm-workspace.yamlwithallowBuilds: { esbuild: true }so the isolated prepare environment does not fail on esbuild's (harmless platform-binary validation) postinstall;typescript+tsdownare regulardependenciesso that environment always has the build tools. - The optional invariant companion (
dsh-auto-review/invariant) needs theinvariantsservice (agent-spine compositions such as headless/ACP); the plain web profile does not provide it, so the row ships commented out in the bundle patch.
Recommended when you publish: dsh Β· dsh-plugin Β· deepseek-harness Β· deepseek Β· cordis Β· ai-safety Β· approval Β· sandbox Β· subagent Β· llm
- Andy8647/dsh-auto-approval β two-state allow/deny classifier on the
tools/pre-executewaterfall with file-log audit.dsh-auto-reviewdeliberately differs: official answerer chain, always delegates what it does not own, read-only second model with a structured verdict, deny reasons fed back to the model, session-log audit. - ACP automation bridge β one-shot machine decisions for its own ACP-owned agents.
dsh-auto-reviewis session- and tool-policy-scoped for the interactive harness; it never infers durable grants.
pnpm install # node ^22.19 || >=24
pnpm run typecheck # tsc, src + tests
pnpm test # vitest: 135 tests, 8 suites
pnpm run build # tsc declarations + tsdown bundles (lib/, incl. the client bundle)
pnpm run verify:self-contained
pnpm pack # publish artifactRepository layout (plugin-template structure): src/index.ts (plugin contract) Β· src/config.ts (Schemastery schema + resolution) Β· src/runtime.ts (answerer, command, deny-reason injection) Β· src/review.ts (reviewer orchestration, prompt, sanitization) Β· src/events.ts (session-event vocabulary + folds) Β· src/projection.ts + src/projection-types.ts (the autoReview session projection) Β· src/invariant.ts (invariant companion) Β· src/client/ (browser half: review panel, locales, styles) Β· test/ Β· fixtures/.
Thanks to everyone who has contributed to dsh-auto-review:
- PerryLink β author and maintainer: approval answerer, reviewer subagent, risk policy and circuit breaker, session-projection review panel, invariant companion, docs, CI/CD and releases.
Want to help? Check the issue templates, the security policy, and AGENTS.md for repo conventions β PRs are welcome in English or Chinese.