Benchmark for coding agents on large multi-repo enterprise codebases: 112 tasks across 10 task types covering cross-repo dependency tracing, incident investigation, and non-patch artifacts. WIP; results not yet validated.
-
Updated
Aug 16, 2026 - HTML