Skip to content

Latest commit

 

History

History
35 lines (23 loc) · 3.55 KB

File metadata and controls

35 lines (23 loc) · 3.55 KB

amber-opencode

Public periodic AMBER benchmark results of models served by OpenCode Go (opencode.ai/zen/go). Cases stay private; results are public. 中文说明:README.md

What this is

  • A 'lane' is one vendor's shop/API for a model name; a 'case' is one task, a 'run' is one sitting (a multi-variant case has several runs).

  • One results/YYYY-Www.md per issue: same cases, same harness (the program that runs the exam and scores it), full library per model; same model name across vendors side by side.

  • Each issue pins: library size and hashes, per-case defect-hunt score and pass/fail, terminal states (how the run process exited), token usage (when the lane reports it) and latency, environment fingerprint, and a qualitative verdict written under evidence discipline.

  • Cases, oracles, transcripts (full answer logs)s and intermediates are never published.

  • Sister repos: amber-deepseek (official DeepSeek lane), amber-commandcode (CommandCode lane), amber-gpt, amber-crof, amber-ollama, amber-devin, amber-workbuddy (WorkBuddy ACP lane), amber-doubao, amber-goldenpotato, amber-kimi, amber-stepfun. This repo's comparison axis is same-name cross-vendor duels — the same model name on OpenCode Go / CommandCode / the official DeepSeek API can be a different endpoint, and every cross-repo citation carries an explicit date and band declaration.

Publication red lines

  1. Publish only: scores and aggregates, token usage (when reported), speed, qualitative verdicts.
  2. Never publish: case content, oracles/graders, transcripts, candidate workspaces, anything that could reconstruct a case.
  3. Every issue pins: model ID, effort band (the thinking-effort setting), date (UTC), harness version, per-case bundle hash — verifiable against the public hash index in amber.
  4. Case numbering is private: public matrices use stable aliases (A-xxxxxxxx, hash-derived) plus bundle hashes only.
  5. Tone: community measurement, not vendor attacks.

A methodological premise

Same model name, same provider, two runs can still score differently — inference parameters, load, and server-side versions drift. Relay/aggregator lanes add an upstream routing layer: the same name may not be the same endpoint. Every conclusion here is dated and banded, and we re-test periodically. A single day's number is a snapshot, not a law.

Results index

Issue Content Headline
2026-W37 deepseek-flash (V4.1 GA) full-library debut (23 cases, GA day) See the issue for scores and verdicts; all three same-name lanes verified genuine v4.1; strong build/ops, with the no-tools phantom-tool-call disease on review/vision papers
2026-W38 correction notice W38 full-library review: 0 cells reversed · 3 held here 3 W37 ocgo-column cells held; the 16/23 headline may move up

Disclaimer

Not affiliated with or sponsored by OpenCode or DeepSeek. Scores are dated, band-specific snapshots, not purchasing advice.