Agent Release Gate turns a completed agent-benchmark report into a deterministic,
machine-readable go or no_go decision. It is designed for release automation
that needs an explicit answer after evaluation: does this agent version satisfy
this release policy?
The v0 adapter reads ClawProBench reports. The gate does not fetch, install, import, or execute benchmark code.
A benchmark score is evidence, not a release policy. Agent Release Gate keeps those concerns separate:
- a benchmark produces a report;
- an adapter validates and normalizes that report;
- a versioned TOML policy defines acceptable capability, reliability, coverage, safety, and execution integrity;
- the evaluator applies every rule in a stable order;
- the CLI writes a decision artifact and returns an automation-friendly exit code.
benchmark report ──> adapter ──> normalized evidence ──┐
├──> evaluator ──> decision JSON
policy TOML ───────────────────────────────────────────┘ + exit code
pinned benchmark checkout ──> provenance validation ────────────────────┘
The result records the report hash, policy hash, benchmark source commit,
observed metrics, and every blocking rule. Except for evaluated_at, equivalent
inputs produce equivalent substantive output.
Clone Agent Release Gate into the directory that will contain both repositories:
cd /path/to/Projects
git clone \
--no-recurse-submodules \
-c core.hooksPath=/dev/null \
https://github.com/bsha6/agent-release-gate.git agent-release-gateThe project and audited benchmark checkout must be direct siblings:
Projects/
├── agent-release-gate/
└── ClawProBench/
Clone ClawProBench without running hooks or checking out its vendored agent implementations:
cd /path/to/Projects
git clone \
--filter=blob:none \
--no-checkout \
--no-recurse-submodules \
-c core.hooksPath=/dev/null \
https://github.com/suyoumo/ClawProBench.git ClawProBench
cd ClawProBench
git sparse-checkout init --cone
git sparse-checkout set \
config custom_checks datasets fixtures frameworks harness \
mock_tools scenarios scripts tests
git -c core.hooksPath=/dev/null checkout --detach \
c4b8395854fe0752eef435b44f140366efd44d8eThe default integration manifest requires that exact origin and commit, a clean
worktree, and no checked-out ironclaw/ or nanoclaw/ directories.
Agent Release Gate is not published to PyPI. From a trusted source checkout:
cd /path/to/Projects/agent-release-gate
python3.14 -m venv .venv
.venv/bin/python -m pip install .Requirements are Python 3.14+, Git, and a POSIX-style operating system with descriptor-relative filesystem operations. The installed package has no third-party runtime dependencies. Run commands from the source checkout when using the included default policy and integration manifest.
.venv/bin/agent-release-gate doctorA valid checkout produces JSON and exits 0:
{
"integration": {
"adapter": "clawprobench",
"commit": "c4b8395854fe0752eef435b44f140366efd44d8e",
"name": "ClawProBench",
"repository_url": "https://github.com/suyoumo/ClawProBench.git"
},
"schema_version": 1,
"valid": true
}doctor uses read-only Git commands. It never runs upstream code or changes the
benchmark checkout.
mkdir -p decisions
.venv/bin/agent-release-gate evaluate \
--adapter clawprobench \
--report tests/fixtures/clawprobench_go.json \
--policy policies/default.toml \
--integration integrations/clawprobench.lock.json \
--output decisions/release-decision.jsonAn abridged passing decision looks like this:
{
"adapter": "clawprobench",
"decision": "go",
"blockers": [],
"benchmark": {
"name": "ClawProBench",
"source_version": "c4b8395854fe0752eef435b44f140366efd44d8e",
"subject": "agent-go"
},
"observed": {
"capability_score": 0.8,
"coverage_ratio": 1.0,
"execution_failures": 0,
"safety_passed": true,
"strict_pass_rate": 0.8
},
"policy": {
"name": "default",
"min_capability_score": 0.7,
"min_coverage_ratio": 1.0,
"min_strict_pass_rate": 0.7
},
"schema_version": 1
}The committed fixtures are synthetic examples, not copied benchmark results or user data.
Consumers should use both the exit code and decision artifact:
0: valid evaluation with decisiongo;1: valid evaluation with decisionno_go;2: invalid arguments, unproven integration, malformed evidence, invalid policy, unknown adapter, or output failure.
Treat exit 2 as a pipeline error, not as an ordinary failed release gate:
gate_status=0
agent-release-gate evaluate \
--adapter clawprobench \
--report "$REPORT_PATH" \
--policy policies/default.toml \
--integration integrations/clawprobench.lock.json \
--output release-decision.json || gate_status=$?
case "$gate_status" in
0) echo "release approved" ;;
1) echo "release blocked"; exit 1 ;;
2) echo "release evaluation invalid" >&2; exit 2 ;;
esacThe CLI evaluates every policy rule, so a no_go decision reports all observed
blockers rather than stopping at the first failure. If evaluation fails before a
valid decision is produced, an existing output file is preserved.
The ClawProBench adapter accepts an existing JSON report and validates required field types, numeric ranges, scenario counts, per-trial safety state, and execution status. Unknown additive fields are ignored. The exact report bytes are hashed with SHA-256 and recorded in the decision.
The report is evidence supplied to the gate. v0 does not prove that it was produced honestly or by the pinned benchmark source.
The default [gate] policy requires:
- capability score of at least
0.70; - strict pass rate of at least
0.70; - complete requested-scenario coverage;
- every observed trial to pass its safety gate;
- zero execution failures.
A custom TOML policy can change thresholds without changing the adapter or evaluation logic. The exact policy bytes are hashed into the decision.
The ClawProBench lock file binds an adapter to its expected repository URL, full Git commit, sibling checkout path, and prohibited checked-out paths. Evaluation fails before reading the report if the requested adapter and manifest do not match.
The output is sorted, indented JSON containing:
goorno_goand ordered blockers;- normalized observed metrics;
- benchmark identity, report hash, and report timestamp;
- policy values and policy hash;
- validated repository URL and source commit;
- schema version and UTC evaluation time.
Absolute local checkout paths are never serialized. Output must be separate from the report, policy, manifest, and benchmark checkout.
Inputs and the validated benchmark checkout are pinned by file descriptor. Output is written atomically through a held directory descriptor. The CLI defends against path replacement, parent-symlink swaps, case variants, hard-link aliases, mixed-checkout provenance, hidden untracked files, and Git index flags that can conceal modifications.
Filesystem identity and ancestry are checked when descriptors are acquired and immediately before output commit. A hostile concurrent process that can write both the output and benchmark directory trees is outside the v0 threat model; run evaluations where untrusted processes cannot rename those directories.
See architecture, dependency boundaries, and the security policy for the detailed model.
Benchmark-specific parsing stops at the adapter boundary. The domain model, policy evaluator, decision schema, and exit-code contract do not depend on ClawProBench.
To add a benchmark:
- implement the
BenchmarkAdapterprotocol; - normalize its report into
BenchmarkEvidence; - register one stable lowercase adapter name;
- add a pinned integration manifest;
- cover passing, blocking, malformed, and additive report cases with synthetic fixtures.
See Adding a Benchmark Adapter for the full contract.
No dependency installation is required for repository tests:
PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=src \
python3.14 -m unittest discover -s tests -vCI runs the full suite and installed-package smoke tests on macOS and Ubuntu. It also builds the wheel and a source archive with neutral ownership metadata.
Version 0.1.0 is an early, source-distributed release supporting the documented
ClawProBench report shape and integration. Outside pull requests and feature
requests are not being solicited for v0.
Agent Release Gate is licensed under the Apache License 2.0. Report vulnerabilities through the private process in SECURITY.md.