Problem
A benchmark with results from one lab is a claim, not a benchmark. AOBench currently
publishes numbers the maintainer produced. Independent results, from independent
hardware, are what make it credible.
A model that scores badly is a more useful submission than one that scores well. The
interesting finding in this project is that capable systems answer respectably and comply
with access policy almost not at all — every independent confirmation or refutation of
that is worth more than a good headline number.
Desired result
One reproducible leaderboard entry for a model you already have access to.
Good candidates:
Please do not buy credits for this. If you have no access to a model, D1 (authoring a
task) is the contribution to make instead.
Where to look
| Path |
Why |
docs/guides/evaluating-your-own-agent.md |
The bring-your-own-agent path |
src/aobench/cli/run_cmd.py, src/aobench/cli/leaderboard_cmd.py |
The commands you will use |
docs/leaderboard.md |
Where results live |
Implementation hints
uv run aobench run all --adapter openai:<model> --split dev
uv run aobench report json data/runs/<run_id>
uv run aobench clear run data/runs/<run_id> --output data/runs/<run_id>/clear_report.json
All three were re-run end to end on main on 2026-09-11 and do what this issue says
they do. Two things that will save you a confused half-hour, both found by actually running
them rather than by reading the code:
- Pass
--output explicitly. aobench clear run defaults --output to the bare name
clear_report.json, so without it the scorecard lands in whatever directory you ran the
command from — your repo root — rather than beside the run it describes. Run two models
and the second silently overwrites the first. (main now gitignores that path so it
cannot sneak into your PR, but putting it in the run directory is what you actually want.)
- Assurance is
N/A if you dry-run with direct_qa, and that is correct. The free,
no-API-key adapter is a zero-tool baseline, and the A axis is engagement-aware: it
averages governance only over runs that actually invoked a tool. A model that never
engages has no defined Assurance, so the column reads N/A. You have not done anything
wrong — you will get a real A as soon as you point it at a model that uses the tools.
Worth knowing before you conclude your setup is broken, because this issue also says a
submission without the assurance axis is not usable.
Report the full CLEAR scorecard, not just the headline. The governance/assurance axis
is the one this benchmark exists to measure — a submission that omits it is not usable.
Acceptance criteria
Tests
None — this is a data contribution. It is checked by re-running the command you state.
Difficulty
A few hours, most of it waiting on the run.
What you get for it
Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:
- Your name lands on a public page with a stable URL. The contributor
wall and
AUTHORS.md describe what each
person actually did, in specific terms, not as a row of avatars. If that is useful to you
as third-party evidence — a CV, a profile, a funding application — please use it. That is
what it is for.
- Release notes name contributors for the version their change shipped in.
- Co-authorship is on the table for this kind of work. The
recognition policy
says it in its own words: "Substantial corpus or methodological contributions may warrant
co-authorship on a paper that depends on them. If you believe that applies to your work,
say so — the awkwardness of asking should not decide who gets credit." It says may, and
it means may — it is a conversation, not a promise attached to a single merged file. But
corpus work is exactly the category the policy was written for, and asking is explicitly
welcome.
- Contributions stay listed even if you later step away.
If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.
Getting started
Use the
leaderboard submission form, or
comment here with what you plan to run. If your numbers look embarrassing for the model,
submit them anyway — that is the finding.
Problem
A benchmark with results from one lab is a claim, not a benchmark. AOBench currently
publishes numbers the maintainer produced. Independent results, from independent
hardware, are what make it credible.
A model that scores badly is a more useful submission than one that scores well. The
interesting finding in this project is that capable systems answer respectably and comply
with access policy almost not at all — every independent confirmation or refutation of
that is worth more than a good headline number.
Desired result
One reproducible leaderboard entry for a model you already have access to.
Good candidates:
OLLAMA_BASE_URL, seedocs/guides/evaluating-your-own-agent.mdPlease do not buy credits for this. If you have no access to a model, D1 (authoring a
task) is the contribution to make instead.
Where to look
docs/guides/evaluating-your-own-agent.mdsrc/aobench/cli/run_cmd.py,src/aobench/cli/leaderboard_cmd.pydocs/leaderboard.mdImplementation hints
All three were re-run end to end on
mainon 2026-09-11 and do what this issue saysthey do. Two things that will save you a confused half-hour, both found by actually running
them rather than by reading the code:
--outputexplicitly.aobench clear rundefaults--outputto the bare nameclear_report.json, so without it the scorecard lands in whatever directory you ran thecommand from — your repo root — rather than beside the run it describes. Run two models
and the second silently overwrites the first. (
mainnow gitignores that path so itcannot sneak into your PR, but putting it in the run directory is what you actually want.)
N/Aif you dry-run withdirect_qa, and that is correct. The free,no-API-key adapter is a zero-tool baseline, and the A axis is engagement-aware: it
averages governance only over runs that actually invoked a tool. A model that never
engages has no defined Assurance, so the column reads
N/A. You have not done anythingwrong — you will get a real A as soon as you point it at a model that uses the tools.
Worth knowing before you conclude your setup is broken, because this issue also says a
submission without the assurance axis is not usable.
Report the full CLEAR scorecard, not just the headline. The governance/assurance axis
is the one this benchmark exists to measure — a submission that omits it is not usable.
Acceptance criteria
--split devonly (thetestsplit is held out)Tests
None — this is a data contribution. It is checked by re-running the command you state.
Difficulty
A few hours, most of it waiting on the run.
What you get for it
Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:
wall and
AUTHORS.mddescribe what eachperson actually did, in specific terms, not as a row of avatars. If that is useful to you
as third-party evidence — a CV, a profile, a funding application — please use it. That is
what it is for.
recognition policy
says it in its own words: "Substantial corpus or methodological contributions may warrant
co-authorship on a paper that depends on them. If you believe that applies to your work,
say so — the awkwardness of asking should not decide who gets credit." It says may, and
it means may — it is a conversation, not a promise attached to a single merged file. But
corpus work is exactly the category the policy was written for, and asking is explicitly
welcome.
If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.
Getting started
Use the
leaderboard submission form, or
comment here with what you plan to run. If your numbers look embarrassing for the model,
submit them anyway — that is the finding.