Problem
AOBench cannot call models hosted on AWS Bedrock. For anyone evaluating agents inside an
AWS environment — which includes a lot of the HPC and enterprise operations audience this
benchmark is aimed at — that is the blocker between them and a submitted result.
Desired result
aobench run task --task JOB_USR_001 --env env_01 --adapter bedrock:<model-id>
Read this first
If the LiteLLM adapter (#33) lands, this may reduce to a documentation page plus a regression
test. Please check #33 before starting, and comment here so we do not do the work
twice.
Where to look
| Path |
Why |
src/aobench/adapters/openai_adapter.py |
The shape to mirror |
src/aobench/adapters/base.py |
The contract: BaseAdapter.run(context) -> Trace |
src/aobench/adapters/__init__.py |
Register it |
pyproject.toml |
Optional extra, following openai/anthropic |
Implementation hints
- SDK:
boto3 (bedrock-runtime). The Converse API is the one with a uniform tool-use
shape across Bedrock-hosted models — prefer it over per-model invoke payloads.
- Credentials come from the AWS credential chain, not an env var you can check for.
So the "not configured" error path is genuinely different from the other adapters and
needs its own message — say which credential sources were tried.
- Model IDs are region-qualified. The adapter string has to accommodate a region;
decide how (bedrock:<region>/<model-id>?) and say so in the PR.
- Populate token counts — the efficiency scorer and CLEAR cost axis depend on them.
Acceptance criteria
Tests
tests/unit/ with a mocked boto3 client (or botocore stubber), following the existing
OpenAI adapter tests. No network calls, and no AWS credentials required to run the suite.
Difficulty
Roughly a day.
Getting started
Comment before starting — the region-in-adapter-string question is worth settling first.
Problem
AOBench cannot call models hosted on AWS Bedrock. For anyone evaluating agents inside an
AWS environment — which includes a lot of the HPC and enterprise operations audience this
benchmark is aimed at — that is the blocker between them and a submitted result.
Desired result
Read this first
If the LiteLLM adapter (#33) lands, this may reduce to a documentation page plus a regression
test. Please check #33 before starting, and comment here so we do not do the work
twice.
Where to look
src/aobench/adapters/openai_adapter.pysrc/aobench/adapters/base.pyBaseAdapter.run(context) -> Tracesrc/aobench/adapters/__init__.pypyproject.tomlopenai/anthropicImplementation hints
boto3(bedrock-runtime). The Converse API is the one with a uniform tool-useshape across Bedrock-hosted models — prefer it over per-model invoke payloads.
So the "not configured" error path is genuinely different from the other adapters and
needs its own message — say which credential sources were tried.
decide how (
bedrock:<region>/<model-id>?) and say so in the PR.Acceptance criteria
aobench run task --adapter bedrock:<model-id> ...produces a scoreableTraceopenaiadapterboto3and missing/invalid AWS credentials produce distinct, actionable messagesaobench list adaptersincludes itdocs/guides/evaluating-your-own-agent.mdupdated, including how regions are handledTests
tests/unit/with a mocked boto3 client (orbotocorestubber), following the existingOpenAI adapter tests. No network calls, and no AWS credentials required to run the suite.
Difficulty
Roughly a day.
Getting started
Comment before starting — the region-in-adapter-string question is worth settling first.