Skip to content

Run a model and submit a leaderboard entry — negative results welcome #27

Description

@MSKazemi

Problem

A benchmark with results from one lab is a claim, not a benchmark. AOBench currently
publishes numbers the maintainer produced. Independent results, from independent
hardware, are what make it credible.

A model that scores badly is a more useful submission than one that scores well. The
interesting finding in this project is that capable systems answer respectably and comply
with access policy almost not at all — every independent confirmation or refutation of
that is worth more than a good headline number.

Desired result

One reproducible leaderboard entry for a model you already have access to.

Good candidates:

Please do not buy credits for this. If you have no access to a model, D1 (authoring a
task) is the contribution to make instead.

Where to look

Path Why
docs/guides/evaluating-your-own-agent.md The bring-your-own-agent path
src/aobench/cli/run_cmd.py, src/aobench/cli/leaderboard_cmd.py The commands you will use
docs/leaderboard.md Where results live

Implementation hints

uv run aobench run all --adapter openai:<model> --split dev
uv run aobench report json data/runs/<run_id>
uv run aobench clear run data/runs/<run_id> --output data/runs/<run_id>/clear_report.json

All three were re-run end to end on main on 2026-09-11 and do what this issue says
they do. Two things that will save you a confused half-hour, both found by actually running
them rather than by reading the code:

  • Pass --output explicitly. aobench clear run defaults --output to the bare name
    clear_report.json, so without it the scorecard lands in whatever directory you ran the
    command from — your repo root — rather than beside the run it describes. Run two models
    and the second silently overwrites the first. (main now gitignores that path so it
    cannot sneak into your PR, but putting it in the run directory is what you actually want.)
  • Assurance is N/A if you dry-run with direct_qa, and that is correct. The free,
    no-API-key adapter is a zero-tool baseline, and the A axis is engagement-aware: it
    averages governance only over runs that actually invoked a tool. A model that never
    engages has no defined Assurance, so the column reads N/A. You have not done anything
    wrong
    — you will get a real A as soon as you point it at a model that uses the tools.
    Worth knowing before you conclude your setup is broken, because this issue also says a
    submission without the assurance axis is not usable.

Report the full CLEAR scorecard, not just the headline. The governance/assurance axis
is the one this benchmark exists to measure — a submission that omits it is not usable.

Acceptance criteria

  • --split dev only (the test split is held out)
  • Exact model identifier, adapter string, AOBench version/tag, and date recorded
  • Hardware or endpoint stated well enough that someone else could repeat it
  • Full CLEAR scorecard included, hard-fails reported separately
  • The run is reproducible from the stated command

Tests

None — this is a data contribution. It is checked by re-running the command you state.

Difficulty

A few hours, most of it waiting on the run.

What you get for it

Stated plainly, because this project asks for corpus work more than for code and the credit
for it is easy to leave implicit:

  • Your name lands on a public page with a stable URL. The contributor
    wall
    and
    AUTHORS.md describe what each
    person actually did, in specific terms, not as a row of avatars. If that is useful to you
    as third-party evidence — a CV, a profile, a funding application — please use it. That is
    what it is for.
  • Release notes name contributors for the version their change shipped in.
  • Co-authorship is on the table for this kind of work. The
    recognition policy
    says it in its own words: "Substantial corpus or methodological contributions may warrant
    co-authorship on a paper that depends on them. If you believe that applies to your work,
    say so — the awkwardness of asking should not decide who gets credit."
    It says may, and
    it means may — it is a conversation, not a promise attached to a single merged file. But
    corpus work is exactly the category the policy was written for, and asking is explicitly
    welcome.
  • Contributions stay listed even if you later step away.

If you would rather not be listed at all, say so in the PR and you will not be. Opting out
is honoured immediately and without being asked why.

Getting started

Use the
leaderboard submission form, or
comment here with what you plan to run. If your numbers look embarrassing for the model,
submit them anyway — that is the finding.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: reportingReports, leaderboard, scorecardseffort: smallA few hours — good place to startgood first issueGood for newcomershelp wantedExtra attention is needed

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions