This issue stays open until every cell has more than one task. It is a standing call
across all 50 cells, not a single task — GitHub closed it automatically when the first one
landed (PR #78 by
@userfypp, which filled DOCS_USR), and it has been
reopened. 31 cells are still thin. If you are looking for the most useful thing you
can do here, it is still this.
Problem
Corpus breadth is the main thing limiting what AOBench can measure, and coverage is
lopsided. MON_SYS and ENERGY_FAC have 6 tasks each, while 31 of the 50
QCAT × role cells have exactly one task. A cell with a single task cannot distinguish
"the agent understands this category" from "the agent got lucky on this item".
Cells with exactly one task today:
AIOPS_DES AIOPS_RES AIOPS_USR ARCH_FAC ARCH_RES ARCH_SYS ARCH_USR
DATA_DES DATA_FAC DATA_RES DATA_SYS DATA_USR DOCS_DES DOCS_FAC
DOCS_RES DOCS_SYS ENERGY_DES ENERGY_RES FAC_DES FAC_FAC FAC_RES
FAC_SYS FAC_USR JOB_DES JOB_RES MON_DES PERF_DES PERF_FAC
PERF_RES SEC_DES SEC_RES
Regenerate the list yourself at any time:
ls benchmark/tasks/specs/*.json | xargs -n1 basename \
| sed -E 's/^M100_//; s/_[0-9]+\.json$//' | sort | uniq -c | sort -n
# The `s/^M100_//` matters: an M100 task spec is named M100_<QCAT>_<ROLE>_NNN.json,
# so without stripping that prefix the count shows spurious "M100_ENERGY_SYS" style
# rows instead of folding those tasks into the real cell they belong to.
Desired result
One new task spec in a cell that currently has one task, reusing an existing environment
bundle.
You do not need cluster access, an API key, or a GPU to do this. That is the entire
point of the snapshot design, and it is why this is the most valuable contribution
anyone can make to AOBench.
Where to look
| Path |
Why |
docs/guides/adding-a-task.md |
Read this first. The complete path from idea to merged, including the reviewer checklist. 159 lines, all of them worth it |
src/aobench/schemas/task.py |
The TaskSpec model — authoritative, more so than the guide |
benchmark/tasks/specs/ |
Where your new <QCAT>_<ROLE>_<NNN>.json goes; the existing task in your chosen cell is the model to copy |
docs/reference/environment-catalog.md |
Pick an environment that already supports your task |
Implementation hints
- Pick a cell from the list above. Open its one existing task and read it end to end.
- Reuse that task's environment bundle. Authoring a new environment is a separate,
larger contribution — see docs/guides/adding-an-environment.md.
- Your gold answer must be derivable from evidence that actually exists in the
snapshot. If answering requires knowledge the agent has no way to obtain, the task
measures memorisation and it will not merge.
- The task must be role-sensitive: a competent operator in a different role should
answer differently, or need different tool access. If not, it is a knowledge question
and belongs in DOCS at best.
- Set
benchmark_split to dev. The test split is held out and is not open for
contribution.
Acceptance criteria
Tests
No new test file is normally needed — the corpus validators cover it:
uv run aobench validate benchmark
uv run python -m pytest tests/unit/ -k "task or spec or validate"
Difficulty
Roughly a day for your first one, most of it spent reading the environment snapshot
rather than writing JSON. Much faster after that — which is why this is the lane we most
want people in.
What you get for it
Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:
- Your name lands on a public page with a stable URL. The contributor
wall and
AUTHORS.md describe what each
person actually did, in specific terms, not as a row of avatars. If that is useful to you
as third-party evidence — a CV, a profile, a funding application — please use it. That is
what it is for.
- Release notes name contributors for the version their change shipped in.
- Co-authorship is on the table for this kind of work. The
recognition policy
says it in its own words: "Substantial corpus or methodological contributions may warrant
co-authorship on a paper that depends on them. If you believe that applies to your work,
say so — the awkwardness of asking should not decide who gets credit." It says may, and
it means may — it is a conversation, not a promise attached to a single merged file. But
corpus work is exactly the category the policy was written for, and asking is explicitly
welcome.
- Contributions stay listed even if you later step away.
If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.
Getting started
Comment here with the cell you want to take. I will confirm it is unclaimed and
sanity-check the idea before you write anything — that costs you ten minutes and can
save you a day. Questions are expected, not a nuisance.
Problem
Corpus breadth is the main thing limiting what AOBench can measure, and coverage is
lopsided.
MON_SYSandENERGY_FAChave 6 tasks each, while 31 of the 50QCAT × role cells have exactly one task. A cell with a single task cannot distinguish
"the agent understands this category" from "the agent got lucky on this item".
Cells with exactly one task today:
Regenerate the list yourself at any time:
Desired result
One new task spec in a cell that currently has one task, reusing an existing environment
bundle.
You do not need cluster access, an API key, or a GPU to do this. That is the entire
point of the snapshot design, and it is why this is the most valuable contribution
anyone can make to AOBench.
Where to look
docs/guides/adding-a-task.mdsrc/aobench/schemas/task.pyTaskSpecmodel — authoritative, more so than the guidebenchmark/tasks/specs/<QCAT>_<ROLE>_<NNN>.jsongoes; the existing task in your chosen cell is the model to copydocs/reference/environment-catalog.mdImplementation hints
larger contribution — see
docs/guides/adding-an-environment.md.snapshot. If answering requires knowledge the agent has no way to obtain, the task
measures memorisation and it will not merge.
answer differently, or need different tool access. If not, it is a knowledge question
and belongs in
DOCSat best.benchmark_splittodev. Thetestsplit is held out and is not open forcontribution.
Acceptance criteria
uv run aobench validate benchmarkpasses with the new task includedAOBENCH_SKIP_FIDELITY=1)environment's
rbac_policy.yamluv run aobench run task --task <NEW_ID> --env <ENV> --adapter direct_qacompletesadding-a-task.mdis satisfiedTests
No new test file is normally needed — the corpus validators cover it:
uv run aobench validate benchmark uv run python -m pytest tests/unit/ -k "task or spec or validate"Difficulty
Roughly a day for your first one, most of it spent reading the environment snapshot
rather than writing JSON. Much faster after that — which is why this is the lane we most
want people in.
What you get for it
Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:
wall and
AUTHORS.mddescribe what eachperson actually did, in specific terms, not as a row of avatars. If that is useful to you
as third-party evidence — a CV, a profile, a funding application — please use it. That is
what it is for.
recognition policy
says it in its own words: "Substantial corpus or methodological contributions may warrant
co-authorship on a paper that depends on them. If you believe that applies to your work,
say so — the awkwardness of asking should not decide who gets credit." It says may, and
it means may — it is a conversation, not a promise attached to a single merged file. But
corpus work is exactly the category the policy was written for, and asking is explicitly
welcome.
If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.
Getting started
Comment here with the cell you want to take. I will confirm it is unclaimed and
sanity-check the idea before you write anything — that costs you ten minutes and can
save you a day. Questions are expected, not a nuisance.