Skip to content

AWS Bedrock adapter #35

Description

@MSKazemi

Problem

AOBench cannot call models hosted on AWS Bedrock. For anyone evaluating agents inside an
AWS environment — which includes a lot of the HPC and enterprise operations audience this
benchmark is aimed at — that is the blocker between them and a submitted result.

Desired result

aobench run task --task JOB_USR_001 --env env_01 --adapter bedrock:<model-id>

Read this first

If the LiteLLM adapter (#33) lands, this may reduce to a documentation page plus a regression
test.
Please check #33 before starting, and comment here so we do not do the work
twice.

Where to look

Path Why
src/aobench/adapters/openai_adapter.py The shape to mirror
src/aobench/adapters/base.py The contract: BaseAdapter.run(context) -> Trace
src/aobench/adapters/__init__.py Register it
pyproject.toml Optional extra, following openai/anthropic

Implementation hints

  • SDK: boto3 (bedrock-runtime). The Converse API is the one with a uniform tool-use
    shape across Bedrock-hosted models — prefer it over per-model invoke payloads.
  • Credentials come from the AWS credential chain, not an env var you can check for.
    So the "not configured" error path is genuinely different from the other adapters and
    needs its own message — say which credential sources were tried.
  • Model IDs are region-qualified. The adapter string has to accommodate a region;
    decide how (bedrock:<region>/<model-id>?) and say so in the PR.
  • Populate token counts — the efficiency scorer and CLEAR cost axis depend on them.

Acceptance criteria

  • aobench run task --adapter bedrock:<model-id> ... produces a scoreable Trace
  • Tool calls and results recorded in the same shape as the openai adapter
  • Token usage populated
  • Missing boto3 and missing/invalid AWS credentials produce distinct, actionable messages
  • aobench list adapters includes it
  • docs/guides/evaluating-your-own-agent.md updated, including how regions are handled

Tests

tests/unit/ with a mocked boto3 client (or botocore stubber), following the existing
OpenAI adapter tests. No network calls, and no AWS credentials required to run the suite.

Difficulty

Roughly a day.

Getting started

Comment before starting — the region-in-adapter-string question is worth settling first.

Activity

  1. added
    enhancementNew feature or request
    help wantedExtra attention is needed
    area: adaptersModel/agent adapters (OpenAI, LiteLLM, Gemini, Bedrock, MCP)
    on Aug 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: adaptersModel/agent adapters (OpenAI, LiteLLM, Gemini, Bedrock, MCP)effort: mediumRoughly a dayenhancementNew feature or requesthelp wantedExtra attention is needed

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions