Skip to content

Add long-horizon Arazzo goal evaluation - #124

Merged
SonAIengine merged 2 commits into
mainfrom
codex/arazzo-long-horizon-e2e-0.42
Aug 16, 2026
Merged

Add long-horizon Arazzo goal evaluation#124
SonAIengine merged 2 commits into
mainfrom
codex/arazzo-long-horizon-e2e-0.42

Conversation

@SonAIengine

Copy link
Copy Markdown
Owner

Summary

  • add a public outcome evaluator and replayable long-horizon baseline
  • benchmark paired 1,000-tool catalogs with and without Arazzo across 3, 10, and 30-step workflows
  • record target, required-tool, ordering, binding, goal-state, latency, and token-budget metrics
  • document the methodology and expose focused Make targets

Results

  • Arazzo: target/required-tool/order/binding/goal metrics all 1.00
  • OpenAPI-only control: order/binding/goal metrics 0.00
  • average admitted schema tokens: 260.67; maximum: 301

This benchmark is deterministic (model=none) and does not claim LLM planning quality.

Validation

  • poetry run ruff check .
  • poetry run ruff format --check .
  • poetry run pytest tests/ -q (1235 passed, 47 skipped)
  • frozen Arazzo long-horizon benchmark result passes its gates

@SonAIengine
SonAIengine marked this pull request as ready for review August 16, 2026 07:46
@SonAIengine
SonAIengine merged commit 86874d0 into main Aug 16, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant