You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
π Welcome. This issue is the map. Everything on the roadmap that is not yet built has a tracking issue, grouped so you can find something that fits the time you have.
You do not need HPC access, a cluster, or an API key to contribute. The whole benchmark runs against frozen snapshots on a laptop, and the direct_qa adapter needs no model provider.
The most valuable thing you can do β no code required
AOBench is a benchmark. Its two highest-value contributions are data, not code, and
neither needs cluster access, an API key, or any knowledge of the internals:
Write a second task for a thin QCAT Γ role cell (32 of 50 have only one)Β #26 β write a task. 31 of the 50 QCAT Γ role cells currently have exactly one task,
which means the benchmark cannot tell "understands the category" from "got lucky".
One JSON file against an existing environment snapshot. This is the single most useful
contribution to this project.
Run a model and submit a leaderboard entry β negative results welcomeΒ #27 β run a model and submit its numbers. Independent results from independent
hardware are what make a benchmark credible rather than a claim. Negative results are
more useful than good ones β a model that complies with access policy badly is the
finding, not an embarrassment.
Suggested first contributions
Every issue named here was checked against main on 2026-09-11 and is genuinely open.
No code required
Record a 30-second demo GIF for the READMEΒ #9 β record a 30-second demo GIF for the README. No Python at all. The README asks
people to picture what a run looks like and then does not show them.
Write a second task for a thin QCAT Γ role cell (32 of 50 have only one)Β #26 β write a second task for a thin cell. One JSON file against an existing
environment snapshot. Still the single most useful contribution to this project, and aobench new task --thinnest now picks the emptiest cell and writes a valid spec for you.
DOCS_USR is filled β by @userfypp, the project's first
corpus contribution (#78); 31 cells are not.
Run a model and submit a leaderboard entry β negative results welcomeΒ #27 β run a model and submit its numbers. Listed twice on this page on purpose. It
is the one contribution nobody else can make for you: independent numbers from
independent hardware. Negative results are more useful than good ones.
examples/ is outside the encoding gate β and has 2 live Windows-breaking callsΒ #69 β a quality gate with a blind spot aimed at newcomers. scripts/check_text_encoding.py keeps encoding-less text I/O out of the tree, and its own
docstring argues that a contributor who cannot run the tests is as blocked as one whose
run dies β but it scans src/aobench, tests and scripts and notexamples/, which
is the first thing a new user opens. Add one tuple entry and it goes red on two real
calls in the CI-gate example. Two-word fix; the point is closing the hole.
Verify every command in the docs actually does what the page says β from a clean checkoutΒ #62 β check that the documented commands actually work. 136 fenced shell commands
across 24 pages, and nothing verifies any of them; two have already been caught doing
something other than what the page said. One page per PR, no internals knowledge needed,
and being new to the project is an advantage β the places you get stuck are the finding.
(Three issues stood in this section this morning and all three are now closed, because
contributors finished them within nine hours of each other on 2026-09-11: #61 β
the tools/mypy --strict slice, the last package-sized piece of #7 β by @Akimbo92i (PR #64); #65 β the patch.dict test-isolation bug that #61 surfaced on its way through β by @motodriver (PR #67);
and #66 β the MockSlurmTool._load_json return type that three of the seven job_details.json snapshots disagreed with β by @QIU-Guanzong (PR #68).
The tree is now at a single mypy --strict error and it is a third-party import-untyped. Each of those three was found by the previous one β which is the honest
reason the small-issue shelf keeps emptying.)
About a day
Google Gemini adapterΒ #34 / AWS Bedrock adapterΒ #35 β a Gemini or AWS Bedrock adapter. Nobody can benchmark a model AOBench
cannot call, so every missing provider is a result that never gets submitted.
Triage the 96 CodeQL findings β path injection on the server surfaces firstΒ #23 β triage the CodeQL findings. 80 open alerts, of which 25 are py/path-injection on the server surfaces and the rest is largely noise. The value is
in the triage, not the patching; a dismissed alert with a one-line reason is a real
deliverable here.
Claimed right now β do not duplicate these
Kept current here so an issue someone has spoken for is visibly taken, rather than taken
on trust that I am reading my notifications.
Claiming an issue protects it. If you comment to say you are taking something, it is
yours and nothing will be merged over you. That promise exists because it was broken:
someone claimed #32 at 07:29 one morning, a parallel PR doing the same work was opened
four and a half hours later, and the claim was never even acknowledged. Details and
apology are on that issue.
If you claim something and I go quiet for more than a couple of days, ping it β that is my
failure, not impatience on your part.
If you have claimed something and changed your mind, just say so. Nobody minds, and it
frees it for the next person.
Housekeeping note β 2026-08-10
Three issues that were labelled good first issue (#1, #4, #5) described work that was already on main when they were opened. They have been closed with an explanation and
replaced with a verified set (#26β#39). If you were looking at one of them, sorry β each
closing comment names its live replacement.
Every issue under the label is now checked against main. If you find another that is
already done, say so on the issue and I will close it and credit you for the catch.
Housekeeping note β 2026-09-10
This map had drifted into exactly the failure it apologises for above: four of the five
issues it recommended as first contributions had already shipped β #29, #31 and #32
merged from external contributors, #37 likewise. Anyone arriving here since late August
clicked four suggestions and found closed issues. That is a bad first impression and it
was mine to prevent, not yours to work around. The list above is rebuilt from issues
verified open today, and the thin-cell count is corrected from 26 to 32 of 50, which
is what the corpus actually shows.
If you find another stale issue, say so on it β I will close it and credit you for the
catch, and that catch goes on the contributor
wall the same as a patch does.
Housekeeping note β 2026-09-11
This list was rebuilt yesterday and went stale again the same day: #65 and #66 were
both recommended here as first contributions this morning, and both were merged and closed
by the afternoon. Fixed above, with the finished work credited rather than deleted.
Then, four minutes after that rebuild, the shelf was refilled β #69 and #70, both good first issue / effort: small, both free, both listed above. Both were found the
honest way, by running this project's own quality gates and reading what they printed, and
both were proved against main before the issue was typed: #69's gate genuinely reports OK
today while examples/04_ci_gate.py has two live encoding-less read_text() calls, and #70's root cause was reproduced and reverted to confirm the warning comes back. The
diagnosis in each is tested rather than guessed, so neither sends you chasing a premise that
was never checked.
That churn is worth being straight about, because it is the honest shape of this project
right now: the shelf of small, self-contained issues empties faster than one maintainer
refills it, and it refills mostly because finishing one issue surfaces the next β #61
surfaced #65 and #66; #70 is the same class of defect as #65, found by looking for more of
it. If the small list looks short on the day you arrive, check again, and check the good first issue label
directly rather than trusting this page alone.
The two most valuable contributions here have never been the small ones, though, and
neither needs cluster access, an API key, or any knowledge of the internals:
If small and mechanical is what you have time for, #62 is the deepest seam left: 136
documented commands across 24 pages, one page per PR, and being new here is an advantage,
because the places you get stuck are the finding.
Housekeeping note β 2026-09-12
Seven issues were opened in the last day and every one of them came out of doing the work
rather than from a planning session: #69 and #70 fell out of finishing #61; #71 and #72 out of re-running the gates; #73 out of testing aobench review task while it was two hours old; and #74 and #75 out of sweeping every spec in the
corpus with it. #73 is already closed β @mgalore fixed it
in PR #77 within hours of it being filed.
Of the six still open, none is claimed and four are effort: small (#69, #70, #71, #74).
(Corrected 2026-09-12: this note first said six issues, that #73 was filed by a
contributor rather than fixed by one, and that four were unclaimed with three small. The
counts are now checked against the tracker rather than written from memory.)
That is the honest shape of this project: the backlog refills from contact with the code,
so the list above goes stale within hours of a merge rather than within weeks. If something
here is already closed, say so on it β that catch is credited on the contributor wall exactly like a
patch.
Housekeeping note β 2026-09-19
Three issues suggested above as open work are now closed, all three by @motodriver: #70 (the leaked-coroutine warning, PR #80), #75 (the scaffold-scoring bug, PR #82), and #71 (the test-count fact
gate, PR #83). PERF_RES_002
(PR #81 by @yangziao56) also merged. #74 remains open and is now
claimed β see the table below.
If you click through from this page and find something already closed, say so on it; that
catch is credited exactly like a patch, as noted above.
Housekeeping note β 2026-09-12 (evening)
The first corpus contribution landed.@userfypp filled the DOCS_USR cell (#78) β the first task AOBench
has ever received from outside the project. Thin cells: 32 β 31.
Two things changed as a result, both worth knowing before you pick something:
Corpus work is much cheaper now than when this page was written.aobench new task --thinnest writes a valid spec for the emptiest cell, and make review runs the same
checklist CI runs on your PR. Neither existed a week ago.
How it works
Comment on the issue to say you are taking it. No need to wait for permission on anything labelled good first issue.
Read CONTRIBUTING.md β clone to green tests is three commands.
Open a PR. Prefer under ~300 changed lines.
You will get a first response within 3 working days, even if it is only "seen, I will look properly on Friday". If a PR of yours goes quiet for more than a week, ping it β that is our failure, not rudeness on your part.
Questions are welcome and expected. Ask here, in the issue itself, or in Discussions.
π Welcome. This issue is the map. Everything on the roadmap that is not yet built has a tracking issue, grouped so you can find something that fits the time you have.
You do not need HPC access, a cluster, or an API key to contribute. The whole benchmark runs against frozen snapshots on a laptop, and the
direct_qaadapter needs no model provider.Pick by how much time you have
effort: smallΒ·good first issueeffort: mediumeffort: largePick by what you like working on
area: cliarea: apiarea: mcpΒ·area: a2aarea: scorersarea: benchmarkarea: docsarea: reportingarea: infraΒ·type: tech-debtPick by milestone
blocked: upstream. Discussion welcome, PRs not yet.The most valuable thing you can do β no code required
AOBench is a benchmark. Its two highest-value contributions are data, not code, and
neither needs cluster access, an API key, or any knowledge of the internals:
which means the benchmark cannot tell "understands the category" from "got lucky".
One JSON file against an existing environment snapshot. This is the single most useful
contribution to this project.
hardware are what make a benchmark credible rather than a claim. Negative results are
more useful than good ones β a model that complies with access policy badly is the
finding, not an embarrassment.
Suggested first contributions
Every issue named here was checked against
mainon 2026-09-11 and is genuinely open.No code required
people to picture what a run looks like and then does not show them.
summary emits
mean_dimension_scores, so the per-dimension breakdown the form asks foris one
jqaway. Negative results are more useful than good ones. The firstcommunity submission ([Result] Claude Sonnet BenchmarkingΒ #60) found five bugs in this project on its way to a score, and
its numbers are now the first community row on the
leaderboard.
A few hours
environment snapshot. Still the single most useful contribution to this project, and
aobench new task --thinnestnow picks the emptiest cell and writes a valid spec for you.DOCS_USR is filled β by @userfypp, the project's first
corpus contribution (#78); 31 cells are not.
is the one contribution nobody else can make for you: independent numbers from
independent hardware. Negative results are more useful than good ones.
scripts/check_text_encoding.pykeeps encoding-less text I/O out of the tree, and its owndocstring argues that a contributor who cannot run the tests is as blocked as one whose
run dies β but it scans
src/aobench,testsandscriptsand notexamples/, whichis the first thing a new user opens. Add one tuple entry and it goes red on two real
calls in the CI-gate example. Two-word fix; the point is closing the hole.
(PR #80).
Ten tasks grant tool families that do not exist (
topology,inventory,filesystem,and
bmsin the RBAC policy files), silently dropped byToolRegistryβ three of theten hand the agent no tools at all.
(PR #83).
(PR #82).
test_report_json_flag_emits_clean_json_on_stdout.Flakiness is worth fixing for its own sake: a test that fails at random teaches everyone
to re-run instead of read.
across 24 pages, and nothing verifies any of them; two have already been caught doing
something other than what the page said. One page per PR, no internals knowledge needed,
and being new to the project is an advantage β the places you get stuck are the finding.
(Three issues stood in this section this morning and all three are now closed, because
contributors finished them within nine hours of each other on 2026-09-11: #61 β
the
tools/mypy --strictslice, the last package-sized piece of #7 β by@Akimbo92i (PR #64);
#65 β the
patch.dicttest-isolation bug that #61 surfaced on its way through β by@motodriver (PR #67);
and #66 β the
MockSlurmTool._load_jsonreturn type that three of the sevenjob_details.jsonsnapshots disagreed with β by@QIU-Guanzong (PR #68).
The tree is now at a single
mypy --stricterror and it is a third-partyimport-untyped. Each of those three was found by the previous one β which is the honestreason the small-issue shelf keeps emptying.)
About a day
cannot call, so every missing provider is a result that never gets submitted.
py/path-injectionon the server surfaces and the rest is largely noise. The value isin the triage, not the patching; a dismissed alert with a one-line reason is a real
deliverable here.
Claimed right now β do not duplicate these
Kept current here so an issue someone has spoken for is visibly taken, rather than taken
on trust that I am reading my notifications.
docs/getting-started/quickstart.md(other pages still free)Everything else on this page is free.
Claiming an issue protects it. If you comment to say you are taking something, it is
yours and nothing will be merged over you. That promise exists because it was broken:
someone claimed #32 at 07:29 one morning, a parallel PR doing the same work was opened
four and a half hours later, and the claim was never even acknowledged. Details and
apology are on that issue.
If you claim something and I go quiet for more than a couple of days, ping it β that is my
failure, not impatience on your part.
If you have claimed something and changed your mind, just say so. Nobody minds, and it
frees it for the next person.
Housekeeping note β 2026-08-10
Three issues that were labelled
good first issue(#1, #4, #5) described work that wasalready on
mainwhen they were opened. They have been closed with an explanation andreplaced with a verified set (#26β#39). If you were looking at one of them, sorry β each
closing comment names its live replacement.
Every issue under the label is now checked against
main. If you find another that isalready done, say so on the issue and I will close it and credit you for the catch.
Housekeeping note β 2026-09-10
This map had drifted into exactly the failure it apologises for above: four of the five
issues it recommended as first contributions had already shipped β #29, #31 and #32
merged from external contributors, #37 likewise. Anyone arriving here since late August
clicked four suggestions and found closed issues. That is a bad first impression and it
was mine to prevent, not yours to work around. The list above is rebuilt from issues
verified open today, and the thin-cell count is corrected from 26 to 32 of 50, which
is what the corpus actually shows.
If you find another stale issue, say so on it β I will close it and credit you for the
catch, and that catch goes on the contributor
wall the same as a patch does.
Housekeeping note β 2026-09-11
This list was rebuilt yesterday and went stale again the same day: #65 and #66 were
both recommended here as first contributions this morning, and both were merged and closed
by the afternoon. Fixed above, with the finished work credited rather than deleted.
Then, four minutes after that rebuild, the shelf was refilled β #69 and #70, both
good first issue/effort: small, both free, both listed above. Both were found thehonest way, by running this project's own quality gates and reading what they printed, and
both were proved against
mainbefore the issue was typed: #69's gate genuinely reports OKtoday while
examples/04_ci_gate.pyhas two live encoding-lessread_text()calls, and#70's root cause was reproduced and reverted to confirm the warning comes back. The
diagnosis in each is tested rather than guessed, so neither sends you chasing a premise that
was never checked.
That churn is worth being straight about, because it is the honest shape of this project
right now: the shelf of small, self-contained issues empties faster than one maintainer
refills it, and it refills mostly because finishing one issue surfaces the next β #61
surfaced #65 and #66; #70 is the same class of defect as #65, found by looking for more of
it. If the small list looks short on the day you arrive, check again, and check the
good first issuelabeldirectly rather than trusting this page alone.
The two most valuable contributions here have never been the small ones, though, and
neither needs cluster access, an API key, or any knowledge of the internals:
those cells the benchmark cannot distinguish "understands the category" from "got lucky".
One JSON file fixes one cell.
own author is a claim, not a benchmark.
If small and mechanical is what you have time for, #62 is the deepest seam left: 136
documented commands across 24 pages, one page per PR, and being new here is an advantage,
because the places you get stuck are the finding.
Housekeeping note β 2026-09-12
Seven issues were opened in the last day and every one of them came out of doing the work
rather than from a planning session: #69 and #70 fell out of finishing #61;
#71 and #72 out of re-running the gates; #73 out of testing
aobench review taskwhile it was two hours old; and #74 and #75 out of sweeping every spec in thecorpus with it. #73 is already closed β @mgalore fixed it
in PR #77 within hours of it being filed.
Of the six still open, none is claimed and four are
effort: small(#69, #70, #71,#74).
(Corrected 2026-09-12: this note first said six issues, that #73 was filed by a
contributor rather than fixed by one, and that four were unclaimed with three small. The
counts are now checked against the tracker rather than written from memory.)
That is the honest shape of this project: the backlog refills from contact with the code,
so the list above goes stale within hours of a merge rather than within weeks. If something
here is already closed, say so on it β that catch is credited on the
contributor wall exactly like a
patch.
Housekeeping note β 2026-09-19
Three issues suggested above as open work are now closed, all three by
@motodriver: #70 (the leaked-coroutine warning,
PR #80), #75 (the scaffold-scoring bug,
PR #82), and #71 (the test-count fact
gate, PR #83).
PERF_RES_002(PR #81 by
@yangziao56) also merged. #74 remains open and is now
claimed β see the table below.
If you click through from this page and find something already closed, say so on it; that
catch is credited exactly like a patch, as noted above.
Housekeeping note β 2026-09-12 (evening)
The first corpus contribution landed. @userfypp filled the
DOCS_USRcell (#78) β the first task AOBenchhas ever received from outside the project. Thin cells: 32 β 31.
Two things changed as a result, both worth knowing before you pick something:
Closes #26, which was correctof the author and wrong for the backlog β Write a second task for a thin QCAT Γ role cell (32 of 50 have only one)Β #26 is a standing call across all 50 cells. If
you see it closed again after a task merges, that is the same mechanic, not a decision.
aobench new task --thinnestwrites a valid spec for the emptiest cell, andmake reviewruns the samechecklist CI runs on your PR. Neither existed a week ago.
How it works
good first issue.You will get a first response within 3 working days, even if it is only "seen, I will look properly on Friday". If a PR of yours goes quiet for more than a week, ping it β that is our failure, not rudeness on your part.
Questions are welcome and expected. Ask here, in the issue itself, or in Discussions.