You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Roadmap for v0.5 and beyond — what should we build next?
#40
AOBench has four milestones and about thirty open issues, but the ordering is currently one person's opinion. If you operate a cluster, evaluate agents, or are thinking about giving an AI assistant access to real infrastructure, your opinion is better than mine on what matters next.
The question: of the things below, which two would change whether you actually use AOBench?
Real facility grounding. Six environments are grounded in Marconi100 ExaData today. Extending that — or grounding against a second real facility — is the biggest credibility lever and needs someone with data access.
Second question, more important than the first: what would you need AOBench to measure that it currently does not? The scoring dimensions were chosen from HPC operations experience, not from a survey, and I would like to know where that shows.
No wrong answers, and "I looked at this and bounced off because X" is the single most useful reply in the thread.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
AOBench has four milestones and about thirty open issues, but the ordering is currently one person's opinion. If you operate a cluster, evaluate agents, or are thinking about giving an AI assistant access to real infrastructure, your opinion is better than mine on what matters next.
The question: of the things below, which two would change whether you actually use AOBench?
Second question, more important than the first: what would you need AOBench to measure that it currently does not? The scoring dimensions were chosen from HPC operations experience, not from a survey, and I would like to know where that shows.
No wrong answers, and "I looked at this and bounced off because X" is the single most useful reply in the thread.
All reactions