Repository navigation
Show and tell — post your AOBench numbers (negative results especially welcome) #41
Replies: 1 comment
|
The first community result is in, and it is worth telling you what happened to it — @hari760 ran The submission found four defects in this project, every one of them real:
Four bugs, from someone running the thing rather than reading it, on a machine I do not What this means if you were thinking about posting numbers
Post here, or open a result issue. And if you know someone working on agent evaluation, HPC operations, or LLM benchmarking |
Uh oh!
There was an error while loading. Please reload this page.
If you have run AOBench against anything — a frontier model, an open-weight model on your own hardware, your own agent through a custom adapter, an MCP server — post the numbers here.
Negative results are the point. The finding this benchmark exists to surface is that capable systems answer HPC operations questions respectably and comply with access policy almost not at all. Every independent confirmation or refutation of that is worth more than a good headline score. A model that scores badly is a contribution, not an embarrassment — and it is certainly not an embarrassment for you.
What makes a post useful:
--split dev(thetestsplit is held out)Open-weight models work today through Ollama via
OLLAMA_BASE_URL— see evaluating your own agent.If you would like your run to land on the leaderboard, #27 is the tracking issue and there is a submission form under New issue.
Please do not buy credits to participate. If you have no model access, writing a task (#26) is the contribution to make instead, and it needs nothing but a laptop.
All reactions