Skip to content

Write a second task for a thin QCAT × role cell (32 of 50 have only one) #26

Description

@MSKazemi

This issue stays open until every cell has more than one task. It is a standing call
across all 50 cells, not a single task — GitHub closed it automatically when the first one
landed (PR #78 by
@userfypp, which filled DOCS_USR), and it has been
reopened. 31 cells are still thin. If you are looking for the most useful thing you
can do here, it is still this.

Problem

Corpus breadth is the main thing limiting what AOBench can measure, and coverage is
lopsided. MON_SYS and ENERGY_FAC have 6 tasks each, while 31 of the 50
QCAT × role cells have exactly one task
. A cell with a single task cannot distinguish
"the agent understands this category" from "the agent got lucky on this item".

Cells with exactly one task today:

AIOPS_DES  AIOPS_RES  AIOPS_USR  ARCH_FAC  ARCH_RES  ARCH_SYS  ARCH_USR
DATA_DES   DATA_FAC   DATA_RES   DATA_SYS  DATA_USR  DOCS_DES  DOCS_FAC
DOCS_RES   DOCS_SYS   ENERGY_DES ENERGY_RES FAC_DES   FAC_FAC   FAC_RES
FAC_SYS    FAC_USR    JOB_DES    JOB_RES    MON_DES   PERF_DES  PERF_FAC
PERF_RES   SEC_DES    SEC_RES

Regenerate the list yourself at any time:

ls benchmark/tasks/specs/*.json | xargs -n1 basename \
  | sed -E 's/^M100_//; s/_[0-9]+\.json$//' | sort | uniq -c | sort -n

# The `s/^M100_//` matters: an M100 task spec is named M100_<QCAT>_<ROLE>_NNN.json,
# so without stripping that prefix the count shows spurious "M100_ENERGY_SYS" style
# rows instead of folding those tasks into the real cell they belong to.

Desired result

One new task spec in a cell that currently has one task, reusing an existing environment
bundle.

You do not need cluster access, an API key, or a GPU to do this. That is the entire
point of the snapshot design, and it is why this is the most valuable contribution
anyone can make to AOBench.

Where to look

Path Why
docs/guides/adding-a-task.md Read this first. The complete path from idea to merged, including the reviewer checklist. 159 lines, all of them worth it
src/aobench/schemas/task.py The TaskSpec model — authoritative, more so than the guide
benchmark/tasks/specs/ Where your new <QCAT>_<ROLE>_<NNN>.json goes; the existing task in your chosen cell is the model to copy
docs/reference/environment-catalog.md Pick an environment that already supports your task

Implementation hints

  1. Pick a cell from the list above. Open its one existing task and read it end to end.
  2. Reuse that task's environment bundle. Authoring a new environment is a separate,
    larger contribution — see docs/guides/adding-an-environment.md.
  3. Your gold answer must be derivable from evidence that actually exists in the
    snapshot
    . If answering requires knowledge the agent has no way to obtain, the task
    measures memorisation and it will not merge.
  4. The task must be role-sensitive: a competent operator in a different role should
    answer differently, or need different tool access. If not, it is a knowledge question
    and belongs in DOCS at best.
  5. Set benchmark_split to dev. The test split is held out and is not open for
    contribution.

Acceptance criteria

  • uv run aobench validate benchmark passes with the new task included
  • The F1–F7 fidelity gate passes (run without AOBENCH_SKIP_FIDELITY=1)
  • Every tool call in the gold trajectory is permitted for the declared role under the
    environment's rbac_policy.yaml
  • Every claim in the gold answer traces to a file in the environment snapshot
  • uv run aobench run task --task <NEW_ID> --env <ENV> --adapter direct_qa completes
  • The reviewer checklist at the end of adding-a-task.md is satisfied

Tests

No new test file is normally needed — the corpus validators cover it:

uv run aobench validate benchmark
uv run python -m pytest tests/unit/ -k "task or spec or validate"

Difficulty

Roughly a day for your first one, most of it spent reading the environment snapshot
rather than writing JSON. Much faster after that — which is why this is the lane we most
want people in.

What you get for it

Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:

  • Your name lands on a public page with a stable URL. The contributor
    wall
    and
    AUTHORS.md describe what each
    person actually did, in specific terms, not as a row of avatars. If that is useful to you
    as third-party evidence — a CV, a profile, a funding application — please use it. That is
    what it is for.
  • Release notes name contributors for the version their change shipped in.
  • Co-authorship is on the table for this kind of work. The
    recognition policy
    says it in its own words: "Substantial corpus or methodological contributions may warrant
    co-authorship on a paper that depends on them. If you believe that applies to your work,
    say so — the awkwardness of asking should not decide who gets credit."
    It says may, and
    it means may — it is a conversation, not a promise attached to a single merged file. But
    corpus work is exactly the category the policy was written for, and asking is explicitly
    welcome.
  • Contributions stay listed even if you later step away.

If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.

Getting started

Comment here with the cell you want to take. I will confirm it is unclaimed and
sanity-check the idea before you write anything — that costs you ten minutes and can
save you a day. Questions are expected, not a nuisance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions