Skip to content

Verify every command in the docs actually does what the page says — from a clean checkout #62

Description

@MSKazemi

Problem

The documentation contains 136 fenced shell commands across 24 pages, and nothing checks
that any of them do what the surrounding prose says they do.

This is not hypothetical. Two instances found so far:

  • The leaderboard page told submitters to run aobench report json <run> > my_result.json
    to produce a submission file. That command writes <run>/run_summary.json and prints a
    short human summary to stdout — so the documented step captured prose, not JSON. Someone
    followed it and attached the wrong thing, through no fault of their own
    (#60).
  • The pinned "Start here" issue recommended five first contributions, four of which had
    already shipped (#20, corrected
    2026-09-10).

Both were written from memory rather than executed. A command that no longer works is
worse than a missing one: the reader assumes they broke it, and the ones who quietly give
up are invisible.

Desired result

Every command block in the docs has been run from a clean checkout, and the page says
what actually happens — including the output, when the page claims an output.

Scope this one page at a time

Please do not try to take all 24 pages in one PR. One page, or one small group of
related pages, per PR. A PR that fixes three commands on one page and says "I ran the
other nine and they are correct" is a complete, mergeable contribution.

Highest value first, because these are the pages a newcomer hits:

Page Why it matters
docs/getting-started/installation.md If this is wrong, nothing else gets read
docs/getting-started/quickstart.md The first commands anyone runs
docs/getting-started/first-10-minutes.md The unbranched newcomer path
docs/reference/commands.md The CLI reference — must match --help exactly
docs/guides/evaluating-your-own-agent.md The adapter path; needs an API key for parts
docs/guides/programmatic-access.md Python API rather than CLI
docs/tutorials/serving-the-benchmark.md The server surface

docs/leaderboard.md is already done and is the model: every command on it has been
executed against a real run directory. Use it as the standard for what "verified" means.

How to do it

git clone https://github.com/MSKazemi/aobench && cd aobench
uv sync --all-extras

Then, for each command on the page:

  1. Run it. Actually run it, in that order, from that state.
  2. If the page shows output, compare it to what you got. Paths, counts and IDs drift.
  3. If it fails, or succeeds while doing something other than what the prose says, fix the
    page — or say so in the PR if the right fix is a code change rather than a docs change.

What to report even if you change nothing: which commands you ran, and which you could
not. That list is the deliverable as much as the diff is.

You will not be able to run everything, and that is expected:

  • Anything using --adapter openai:* needs an OPENAI_API_KEY. Skip those and say so.
  • --adapter anthropic:* cannot be verified by the maintainer either, so flag rather than
    guess.
  • The direct_qa adapter needs no API key at all and runs offline against frozen
    snapshots, so the majority of the getting-started path is fully checkable on a laptop.
  • If you are on Windows, please say so in the PR — a Windows contributor lost a 67-task run
    to a cp1252 encoding crash (#60), and a
    second reported he could not run make check at all. Both were right, both are fixed,
    and a third pair of Windows eyes is genuinely valuable.

Acceptance criteria

  • Every command block on the page(s) you took has been executed, or explicitly listed
    as not executable and why
  • Prose and shown output match observed behaviour
  • Any command that needs a key, a cluster or a network call is marked as such on the page
  • The PR body lists what you ran and what you skipped

Tests

Optional and welcome, not required: a test that extracts the fenced commands from a page
and asserts the offline ones exit zero would stop this from rotting again. tests/ has no
docs-execution test today. If you would rather do that than the manual pass, say so — it is
arguably the better contribution, and it is a bigger one.

Difficulty

A few hours per page. No internals knowledge needed — being new to the project is an
advantage here, because the places you get stuck are the finding.

Getting started

Comment with the page you are taking so two people do not check the same one.

@aawhan0 — you have first refusal on any page here,
including all of getting-started if you want it. You volunteered for exactly this work on
#32 ("go through the workflow from a
clean checkout and make sure all commands and expected outputs are actually verified")
and I failed to answer you. No obligation at all — but the offer is real and it is yours
first.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions