Problem
The documentation contains 136 fenced shell commands across 24 pages, and nothing checks
that any of them do what the surrounding prose says they do.
This is not hypothetical. Two instances found so far:
- The leaderboard page told submitters to run
aobench report json <run> > my_result.json
to produce a submission file. That command writes <run>/run_summary.json and prints a
short human summary to stdout — so the documented step captured prose, not JSON. Someone
followed it and attached the wrong thing, through no fault of their own
(#60).
- The pinned "Start here" issue recommended five first contributions, four of which had
already shipped (#20, corrected
2026-09-10).
Both were written from memory rather than executed. A command that no longer works is
worse than a missing one: the reader assumes they broke it, and the ones who quietly give
up are invisible.
Desired result
Every command block in the docs has been run from a clean checkout, and the page says
what actually happens — including the output, when the page claims an output.
Scope this one page at a time
Please do not try to take all 24 pages in one PR. One page, or one small group of
related pages, per PR. A PR that fixes three commands on one page and says "I ran the
other nine and they are correct" is a complete, mergeable contribution.
Highest value first, because these are the pages a newcomer hits:
| Page |
Why it matters |
docs/getting-started/installation.md |
If this is wrong, nothing else gets read |
docs/getting-started/quickstart.md |
The first commands anyone runs |
docs/getting-started/first-10-minutes.md |
The unbranched newcomer path |
docs/reference/commands.md |
The CLI reference — must match --help exactly |
docs/guides/evaluating-your-own-agent.md |
The adapter path; needs an API key for parts |
docs/guides/programmatic-access.md |
Python API rather than CLI |
docs/tutorials/serving-the-benchmark.md |
The server surface |
docs/leaderboard.md is already done and is the model: every command on it has been
executed against a real run directory. Use it as the standard for what "verified" means.
How to do it
git clone https://github.com/MSKazemi/aobench && cd aobench
uv sync --all-extras
Then, for each command on the page:
- Run it. Actually run it, in that order, from that state.
- If the page shows output, compare it to what you got. Paths, counts and IDs drift.
- If it fails, or succeeds while doing something other than what the prose says, fix the
page — or say so in the PR if the right fix is a code change rather than a docs change.
What to report even if you change nothing: which commands you ran, and which you could
not. That list is the deliverable as much as the diff is.
You will not be able to run everything, and that is expected:
- Anything using
--adapter openai:* needs an OPENAI_API_KEY. Skip those and say so.
--adapter anthropic:* cannot be verified by the maintainer either, so flag rather than
guess.
- The
direct_qa adapter needs no API key at all and runs offline against frozen
snapshots, so the majority of the getting-started path is fully checkable on a laptop.
- If you are on Windows, please say so in the PR — a Windows contributor lost a 67-task run
to a cp1252 encoding crash (#60), and a
second reported he could not run make check at all. Both were right, both are fixed,
and a third pair of Windows eyes is genuinely valuable.
Acceptance criteria
Tests
Optional and welcome, not required: a test that extracts the fenced commands from a page
and asserts the offline ones exit zero would stop this from rotting again. tests/ has no
docs-execution test today. If you would rather do that than the manual pass, say so — it is
arguably the better contribution, and it is a bigger one.
Difficulty
A few hours per page. No internals knowledge needed — being new to the project is an
advantage here, because the places you get stuck are the finding.
Getting started
Comment with the page you are taking so two people do not check the same one.
@aawhan0 — you have first refusal on any page here,
including all of getting-started if you want it. You volunteered for exactly this work on
#32 ("go through the workflow from a
clean checkout and make sure all commands and expected outputs are actually verified")
and I failed to answer you. No obligation at all — but the offer is real and it is yours
first.
Problem
The documentation contains 136 fenced shell commands across 24 pages, and nothing checks
that any of them do what the surrounding prose says they do.
This is not hypothetical. Two instances found so far:
aobench report json <run> > my_result.jsonto produce a submission file. That command writes
<run>/run_summary.jsonand prints ashort human summary to stdout — so the documented step captured prose, not JSON. Someone
followed it and attached the wrong thing, through no fault of their own
(#60).
already shipped (#20, corrected
2026-09-10).
Both were written from memory rather than executed. A command that no longer works is
worse than a missing one: the reader assumes they broke it, and the ones who quietly give
up are invisible.
Desired result
Every command block in the docs has been run from a clean checkout, and the page says
what actually happens — including the output, when the page claims an output.
Scope this one page at a time
Please do not try to take all 24 pages in one PR. One page, or one small group of
related pages, per PR. A PR that fixes three commands on one page and says "I ran the
other nine and they are correct" is a complete, mergeable contribution.
Highest value first, because these are the pages a newcomer hits:
docs/getting-started/installation.mddocs/getting-started/quickstart.mddocs/getting-started/first-10-minutes.mddocs/reference/commands.md--helpexactlydocs/guides/evaluating-your-own-agent.mddocs/guides/programmatic-access.mddocs/tutorials/serving-the-benchmark.mddocs/leaderboard.mdis already done and is the model: every command on it has beenexecuted against a real run directory. Use it as the standard for what "verified" means.
How to do it
Then, for each command on the page:
page — or say so in the PR if the right fix is a code change rather than a docs change.
What to report even if you change nothing: which commands you ran, and which you could
not. That list is the deliverable as much as the diff is.
You will not be able to run everything, and that is expected:
--adapter openai:*needs anOPENAI_API_KEY. Skip those and say so.--adapter anthropic:*cannot be verified by the maintainer either, so flag rather thanguess.
direct_qaadapter needs no API key at all and runs offline against frozensnapshots, so the majority of the getting-started path is fully checkable on a laptop.
to a cp1252 encoding crash (#60), and a
second reported he could not run
make checkat all. Both were right, both are fixed,and a third pair of Windows eyes is genuinely valuable.
Acceptance criteria
as not executable and why
Tests
Optional and welcome, not required: a test that extracts the fenced commands from a page
and asserts the offline ones exit zero would stop this from rotting again.
tests/has nodocs-execution test today. If you would rather do that than the manual pass, say so — it is
arguably the better contribution, and it is a bigger one.
Difficulty
A few hours per page. No internals knowledge needed — being new to the project is an
advantage here, because the places you get stuck are the finding.
Getting started
Comment with the page you are taking so two people do not check the same one.