docs: add OpenCode + Qwen3.6-27B benchmark report - #28
Conversation
|
Thanks for the submission. First-touch impression: this is a documentation-only agent result report, not a leaderboard row. Checks: no GitHub status checks are currently reported on the PR, though the author says Result/source: Harbor job is listed as public at https://hub.harborframework.com/jobs/80a2855e-6706-4b12-8759-f32cbe3d0e0a with OpenCode 1.18.15 + Qwen3.6-27B via Groq, reporting 6/14 passed, 42.86% mean reward, and 2 execution errors. Main risk: readers may compare a k=1/n=1 pilot with leaderboard methodology, so the non-leaderboard framing should stay clear. Maintainer verification points:
|
|
Thanks for the review. I independently rechecked each verification point using unauthenticated HTTP requests:
GitHub created CI run https://github.com/devrev/enterprise-bench/actions/runs/31281779123 with conclusion |
Summary
Adds a reproducible report for a complete OpenCode + Qwen3.6-27B pilot across all 14 Enterprise-Bench L1-L2 tasks.
k=1,n=1This is submitted as an honest full-suite community-challenge run report rather than a leaderboard row. A single attempt per task cannot produce the required pass@2 through pass@10 metrics, so none have been fabricated.
Type of contribution
Validation
make validatepassesAll 14 benchmark tasks were run. The public Harbor job and trajectories are accessible without authentication.
Commands run
Core benchmark invocation, with credentials omitted:
harbor run \ -p tasks \ -a opencode \ -m 'groq/qwen/qwen3.6-27b' \ --mcp-config mcp.json \ -k 1 -n 1 --yes \ --jobs-dir jobs/opencode-groq-qwen36-terra-full-v4Notes for reviewers
eng-l1-cended withNonZeroAgentExitCodeErrorafter an invalid tool call.sales-l2-aended with a providerApiRateLimitError.k=1pilot is not claimed to be equivalent to the leaderboard's 140-trial methodology.