Skip to content

Latest commit

 

History

History
181 lines (117 loc) · 40 KB

File metadata and controls

181 lines (117 loc) · 40 KB

AMBER — Sealed-in-Amber Historical Replay Evaluation

Seal the scene in amber. Retake the exam of that moment.

Library 24 cases / 27 papers (headline scoring now on the /24 basis; dead or paused lanes are frozen at ∅) · Spec v0.2.2 (draft) · Hash index v2026-09 · Result repos × 12 · Rule scores always public, cases never

⚠️ Correction 2026-09-18: published "swe-2-low @ Devin" results were actually swe-2-high — that model id does not exist; the server silently routed to its default band. The medium 15 < high 16 < max 18 ladder is unaffected. ⚠️ Correction 2026-W38: first wave of reversals and holds from the full-library review; per-repo notices in the 12 result repos. The 4 observation points in this repo's docs screen are annotated only — not score cells.

中文说明:README.md. The normative document is AMBER-Core-Specification.md.

What this is

A formal evaluation method. Take a real, auditable historical incident and rebuild the scene exactly as it stood before the answer was known. The candidate gets only the information available at that moment; all later evidence and scoring material is physically sealed. It diagnoses, decides, acts, abstains, refuses, or escalates. The rubric is frozen before anyone sees the output. Everything is archived, and everything can be audited.

AMBER in three minutes (for first-time visitors)

① Current leader — swe-2-max @ Devin, 19'/24 (since 2026-09-21 the convergence case A-3f2a9cdd counts toward the headline: swe-2-max ✓ fast 37s; previously 18/23 on the 23-case basis; band curve medium 15 < high 16 < max 18). The same model name on a different endpoint can be a different brain, so we record scores per endpoint × name:

Current top five 2026-W39

② Completion profile: same score, different shape — ten-axis completion matrix (nine axes until 09-21, when the Convergence axis joined), seven lanes side by side; the former five-way 17/23 tie broke up in the convergence make-ups: k3 / glm-5.3-flash / v4.1-flash @ Ollama / hy4 now sit at 18/24, while the CommandCode lane was frozen at ∅ at the owner's call — same score, five different shapes (pin-level completion; negative defect-hunt scores still count). The new k3 row at a glance: an empty dot on UI (the incomplete-deliverable case), attribution below the top pair, vision tied for best with glm-5.3-flash. Added 09-16, the doubao 16/23 row: full marks on all four construction axes (coding/delivery/ops/requirements), level with the leaders — but empty dots on UI and vision, review at a third — the most lopsided shape on the field. Its two verification axes (attribution/defense) were harness-walled in the main sweep and landed real scores in a 3600s makeup (0.47 / 0.50 — low but real; amber-doubao W38 addendum). Negative defect-hunt scores are floored at 0. Added 09-17, the gp27b 14/23 row (Qwen3.8-27B @ goldenpotato community self-hosted): an even more lopsided shape than doubao — construction group at the leaders' level (coding 0.83 / delivery full / ops 0.97 / requirements full) while review/vision/UI sit at absolute zero and attribution/defense at the field's bottom band; aggressive NVFP4 quantization cost nothing on the hands-on faces and everything on the judgment faces (amber-goldenpotato W38):

Completion matrix, five-way 17/23 tie + doubao 16/23 + gp27b 14/23

The ten axes, in plain language — each cell is the lane's completion (0–1) across that axis's cases; pin-scored cases fold in by pins:

  • Coding · cook from the recipe: implement the spec correctly (mean of 6 cases)
  • Delivery · done ≠ handed in: no artifact means 0, however good the plan (1 case)
  • Defense · night-shift guard: plug every hole in the validator without turning away legit input (mean of 2 cases)
  • Attribution · a doctor matching symptoms to causes: pin each defect to the right root cause (1 case, 15 pins)
  • Review · be the inspector: find real defects in someone's deliverable — misses and false alarms both cost, and the score can go negative (2 cases, defect-hunt score)
  • Ops · follow the runbook: backups, cutovers, reconciliation — no skipped steps (mean of 6 cases)
  • Requirements · the client asked for A, not B — ship A (1 case)
  • UI · build the page to the mock, pin-level acceptance (1 case, 12 pins)
  • Vision · spot defects in real screenshots: overlaps, cropping, missing legends — did it actually see them (1 case, defect-hunt score)
  • Convergence · real finish or busywork loops: did the work land, how fast, and did it spin in place farming temp files (1 case; joined 2026-09-21, every examined lane passed so far; blank cell = not yet examined, not a zero)

Axes grow with the library: each new case family can add a column — extend this list the same way.

Who tops each axis? Recomputed across 40 published lanes: Nine-axis podium 2026-09-18 (written before the Convergence axis joined — covers the first nine) — only defense/attribution/review/vision actually rank the field, the four champions belong to four different vendors, and the overall board leader wins none of them; on Convergence every examined lane is tied at full marks, pending more cases.

③ What it is — a private case library plus public scores: the cases never go public, the scores and hashes always do.

What is AMBER

④ How a result is produced — it works like an exam: papers sealed at authoring, sat in a clean room, audited paper by paper, published redacted, checkable by anyone.

How a result is produced

⑤ A counterintuitive finding, with limits — on the families measured so far, high is the sweet spot and top bands backfire (three of the four families plotted). A pattern, not a law: swe-2's band curve is monotone to the top (medium 15 < high 16 < max 18, amber-devin W37), and on some models the band barely moves the score — pick your band by cost and speed, not score; the WorkBuddy lane, Kimi's K2.8, and the stepfun plan lane all have only one band scored so far, no curve to plot (the stepfun endpoint's usage carries no reasoning_tokens, so high is declared but the burn is unverifiable).

effort curves

⑥ Score vs thinking budget — same band (high), same library, and the output-token bill spans 17× (147K vs 2.5M) for scores within a case of each other (token totals from each lane's published issue). The axis is tokens, not dollars: billing is mixed (subscription lanes have no marginal price; the devin and workbuddy lanes report no usage at all, so swe-2 and the wb pair sit this one out). Within a family, more tokens bought no score (luna flat; astra's top bands lose four cases); across families the shape varies — which is why ⑤ is a pattern, not a law.

score vs thinking budget W37

⑦ High TPS only holds on easy problems — same library, per-problem wall time (log axis): a vendor's high TPS is decode speed measured on easy problems, and hard problems mean more thinking, slower effective decoding, and ballooning wall time. Dot = one problem, bar = median; problem IDs are anonymized (the mapping stays private; per-point data in wallclock-2026-w37.csv). On the families measured so far, swe-2 gets slower with each higher band yet solves more (median 82 s → 280 s), while v4.1-flash backfires at the top band (68 s median, two fewer solves). Shapes vary by family — a pattern on these families, not a law.

High TPS only holds on easy problems

Chart sources (.puml for PlantUML, .vega-lite.json / .vg.json for Vega) sit next to the PNGs in docs/images/ — edit a source, re-render, done.

Why

Public benchmarks are frozen, public corpora: training contamination is rampant and unauditable, scores keep inflating, and static Q&A can't measure what real work demands — multi-step diagnosis, abstention, refusal, escalation under uncertainty. AMBER replays provenance-controlled real events whose leak status can be checked. It runs alongside public benchmarks; it doesn't replace them.

Two caveats that won't go away

  • Runtime sealing ≠ training-data purity: we seal runtime evidence; we can't prove the model never saw the future in training.
  • A historical outcome is evidence, not the one correct answer.

Status

Draft v0.2.2. The Core spec is stable; profiles/ and the case-building/runner tooling are not yet published, so you can't execute a compliant run from this repo alone yet. Roadmap and milestones: PLAN.md.

schemas/ has started landing: the Case Manifest JSON Schema, its validator, and the SHRE→AMBER identifier mapping Core §9.3 requires (see below).

Manifest schema and validator

The Case Manifest is the control-plane record: private-channel material, never candidate-visible (Distribution §1, Core §2). Its field set is taken entirely from the protocol text: protocols/distribution.md §3 (provenance, cutoff_utc, the resolved cutoff commit, the time-to-topology mapping rule and its evidence class, the preregistered cutoff rule and its script-output hash, spec_sha256, the sha256 of both artifacts, the declared candidate_input_bundle, the available-information manifest, the eligibility determination and its evidence class, producer identity and signing-key identifier, the leak-check procedure / last run date / result, retirement state, and the sha256 of the protocol document), §3.1 (construction parameters: git version, bundle format version, hash algorithm, bundle size), §3.2, §5.1, §5.2. No field is invented.

  • schemas/manifest.schema.json — the manifest JSON Schema (2020-12); the top level and every field block are closed (additionalProperties: false), so a field the protocol does not enumerate is an error.
  • tools/validate_manifest.py — the validator. It checks the manifest against the schema and additionally rejects any redacted manifest summary carrying a field outside the closed set of Distribution §5.1.
  • schemas/shre-amber-mapping.md — the shre↔amber identifier mapping required by Core §9.3; anything unmappable from public material is marked TBD with the missing artefact named.
  • schemas/examples/ — regression fixtures: one valid manifest, three malformed manifests (missing required fields / wrong types and enum values / fields outside the enumerated set) and one redacted summary carrying an excluded field.

Usage (YAML and JSON are both accepted):

python3 tools/validate_manifest.py schemas/examples/manifest.valid.yaml        # passes, exit code 0
python3 tools/validate_manifest.py schemas/examples/manifest.wrong-types.yaml  # fails, exit code 1 + the offending fields on stderr
python3 tools/validate_manifest.py schemas/examples/summary.invalid-fieldset.yaml   # redacted summary with an excluded field

Exit codes: 0 valid; 1 invalid (schema, closed-set, or cross-field violation); 2 unreadable/unparseable, undetectable kind, or schema load failure. JSON input is dependency-free; YAML uses PyYAML when it is installed and otherwise falls back to a bundled conservative parser that refuses structures it cannot read safely (anchors, aliases, folded scalars) instead of guessing. Install PyYAML with python3 -m pip install pyyaml or uv run --with pyyaml python3 tools/validate_manifest.py <file>.

Contents

  • AMBER-Core-Specification.md — the normative spec: purpose, definitions, mechanisms, 8 invariants, 8 boundaries, epistemic limits, naming review, adoption rules
  • protocols/distribution.md — cross-host case distribution protocol (v0.3): public/private channel split, fixed-form git bundles, detached signature manifests, public index, sealing probes, leak-window adjudication, run records, comparability and verification matrices
  • protocols/stability.md — stability protocol draft (v0.1): same-arm repeats, recovery-after-fail, side-effect counts, separate infrastructure accounting, decision-driven sample sizes
  • docs/instability-memo-2026-09-14.md — $0 historical-drift memo: 900 published matrix cells across 31 canonical arms (republished columns flagged), showing why a score snapshot is not a stability certificate
  • docs/stage0-flip-analysis-2026-09-14.md — stage-0 flip inventory: 17 drift/recovery events (matrix-derived + prose-flagged) with the stage-1 screening candidate list
  • docs/stage1-v001-screen-20260915.md — first designed-stability dataset: A-ea80d793 × glm-5.3-flash@ollama n=20 same-arm repeats, 8/20 pass (~40%, verdict discipline stable 20/20) — boundary-margin cases must report score distributions, not binary flips
  • docs/stage1-reqdrift-screen-20260915.md — A-0676097b × luna-high n=5 case attempts: 3/5 pass; the r3 regression reproduces twice in one sitting on the same two variants (names private) — an escalation-classification boundary case, incl. one hallucinated-delivery no deliverable
  • docs/stage1-sfail-screen-20260915.md — six-case stable-fail set × glm-5.3-flash@ollama n=5 screen: five cases keep the label with zero recoveries (A-87c472cb/A-d511f9e8 partials glued), A-d9b79b46 goes 3/4 pass — label torn (stage-2 n=20 complete: 6/20 ≈30% combined, Wilson [14.5%, 51.9%] — the n=5 screen's 75% over-read; only the binary verdict is trustworthy at n=5); A-a317e74b needs a 3600 s cap to finish — protocol gains per-case timeout table and abort-on-billing
  • docs/nine-axis-top3-2026-09-18.en.md — nine-axis podium (40 lanes recomputed, written before the Convergence axis joined): only defense/attribution/review/vision discriminate, and the four champions belong to four different vendors; the other five axes are saturated at the top
  • hash-index/v2026-09.md — the public hash index: alias + bundle/oracle dual hashes for every case in the current library (24 cases); every results-repo matrix is checked against it
  • schemas/manifest.schema.json · tools/validate_manifest.py · schemas/shre-amber-mapping.md — the Case Manifest JSON Schema, its validator (including the closed-set Distribution §5.1 redacted-summary check), and the SHRE→AMBER identifier mapping
  • PLAN.md — status, milestones (case tooling → reference runner → scoring/adjudication → statistics → public index), open design questions
  • CONTRIBUTING.md — contribution rules: this repo never accepts case content, document versioning and revision policy, what byte-hash-locking the Core spec means

Weekly results (sister repos)

Scores and cases are published separately: results are public, cases never are. These repos run the full library weekly (aliases + bundle hashes verifiable against the public hash index):

Current top five (source snapshot as of 2026-09-24; 24-case set¹; convergence case A-3f2a9cdd counts toward the headline; case = one scored task; ∅ = convergence not sat / frozen lane; ' = contested (held for safety refusal) or invalid (infrastructure-related (test harness or scoring environment) cases: held, void or awaiting re-scoring); neither counts as a win or a loss. Every lane with NA carries an apostrophe, including frozen display rows; a hold does not settle the cause):

# model @ endpoint pass source
1 swe-2-max @ Devin 19'/24 amber-devin W37
2 glm-5.3-flash @ Ollama Cloud 18/24 amber-ollama W37 (also 18/24: deepseek-v4.1-flash @ Ollama Cloud, amber-ollama W37 Addendum 09-11; k3 @ Kimi official coding endpoint, amber-kimi W38; hy4-preview-f @ WorkBuddy (passed on retake after the 429 cooldown, amber-workbuddy W37). Convergence not sat, frozen at 17/23∅: deepseek-v4.1-flash @ CommandCode (frozen at the owner's call, amber-commandcode W37))
3 deepseek-flash @ DeepSeek official 17/24 amber-deepseek W37 (also 17/24: deepseek-flash @ OpenCode Go, amber-opencode W37; swe-2-high @ Devin, amber-devin W37; doubao-seed-evolving @ Volcengine Ark Agent Plan (17'/24, 3 NA), amber-doubao W38; claude-opus-5-5 @ Anthropic subscription lane (17'/24² — full-library debut on the model's release day, amber-claude W39); gpt-6-sol-900k @ OpenAI Codex (17/24 clean, zero held cases — full-library debut the day after the 6-series launch, amber-gpt W39); mimo-v2.6-pro @ CommandCode (17'/24³ — 1 case invalid and held pending re-judge, amber-commandcode W39). Frozen 16/23∅: qwen3.8-27b @ CrofAI (CrofAI lanes retired, amber-crof W37))
4 gpt-5.6-luna-900k @ OpenAI Codex (high) 16'/24³ amber-gpt W37³ (also 16/24: gpt-6-astra-900k @ OpenAI Codex (correction 2026-09-21: its W36 UI-build cells were fallback-delivered, 16/23→15/23, +1 convergence; full-library re-verification in amber-gpt W38 Addendum 5; a fresh same-day full-library run on 2026-09-23 reproduces 16/24 exactly, double-run pair with zero flips — see amber-gpt W39); deepseek-v4-flash:0731 @ Ollama Cloud, amber-ollama W37; swe-2-medium @ Devin, amber-devin W37; deepseek-v4.1-flash @ WorkBuddy, amber-workbuddy W37; step-5-preview @ stepfun plan endpoint (16'/24¹, 4 held), amber-stepfun W38¹. Frozen/paused 15'/23∅: gpt-5.6-sol-900k @ OpenAI Codex (3 NA; paused at the owner's call). Other frozen 15/23∅ lanes: deepseek-v4-flash-0731 @ CrofAI (CrofAI lanes retired); swe-2-low @ Devin (correction 2026-09-18: swe-2-low does not exist — that run was a swe-2-high re-run, so no separate convergence sitting))
5 kimi-for-coding (K2.8 Preview) @ Kimi official coding endpoint 15/24 amber-kimi W38 (also 15/24: gpt-6-luna-900k @ OpenAI Codex (full-library debut the day after the 6-series launch, amber-gpt W39); space-bunny-alpha @ CommandCode (stealth trial model, free-window full-library debut 09-24, amber-commandcode W39); swe-1-7-medium @ Devin, amber-devin W37; Qwen3.8-27B @ goldenpotato community self-hosted endpoint, amber-goldenpotato W38. Frozen 14/23∅: glm-5.3-flash @ CrofAI (CrofAI lanes retired); deepseek-v4.1-flash-exp (preview) @ DeepSeek official (id pulled from the official endpoint, frozen at the owner's call, amber-deepseek W37))

Not yet on the board: Fable and friends — this week's token budget didn't stretch to their exam fees; they sit the library as soon as the budget lands, and scores publish with the next issue. (kimi-k3 sat the library on 2026-09-15 and entered the board at the #2 tie — see the table above.)

¹ Since 2026-09-08 the board runs on the full 23-case library: CrofAI/Ollama/astra sat makeup runs of the 2 ops cases added 09-07 (12/12 papers wire- and hash-verified). astra's 16/23 = W36 -900k 14 cases + a W37 makeup on bare gpt-6-astra (the -900k variant was revoked server-side; the mixed lineage is noted in the issue). Added 2026-09-10: a three-lane duel on DeepSeek V4.1-Flash's GA day — the CommandCode and OpenCode Go relays plus the official DeepSeek API, same day, same band, full library (26/26 papers wire-verified each; CommandCode 17/23 joins the #2 tie, OpenCode Go and the official lane 16/23 join #3 — result repos in the table). Added 2026-09-11: deepseek-v4.1-flash @ Ollama Cloud debuts 17/23, joining the #2 tie (same-day glm-5.3-flash re-run 16/23, inside the known drift band; amber-ollama W37 Addendum); swe-2-max @ Devin takes the board at 18/23 (16/21 on the public 21-case subset) — band curve monotone to the top (medium 15 < high 16 < max 18), an OPS 6/6 sweep (lane-only — gpt's luna and ollama's g53f swept the ops face earlier), at ~5 h wall clock, roughly 4× the previous leader's (25/25 sessioned rows wire-verified, hashes 26/26 against the public index); swe-2-high 16/23 joins the #3 tie. The devin lane's harness does not forward effort — the true band is the UID suffix — and the lane reports no token usage. Out / not in: glm-5.3-flash @ CrofAI 14/23 (former #3 tie), devin swe-1-7-medium 14/23, deepseek-v4.1-flash-exp (preview, official lane) 14/23, gpt-5.6-sol-900k 15/23 (corrected on the 2026-09-16 re-test: the ui-build "zero delivery" was a client-side watchdog kill — small-prompt tiers vs 100–170s of silent high-effort reasoning — with a makeup 12/12 perfect; zero capability drift vs the prior run, see amber-gpt W38; upstream hermes-agent#112909), gpt-5.6-luna-900k 15/23, deepseek-v4-flash-0731 @ CrofAI 15/23, deepseek-v4-flash:0731 @ Ollama Cloud 15/23 (former #3 tie), swe-2-medium 15/23 (best-ever vision 4.0), swe-2-low 15/23 (correction 2026-09-18: swe-2-low is a nonexistent id — the server silently kept its default band, so this was actually a second swe-2-high run; vs the 2026-09-10 first run's 16/23 it is same-model variance, not a band difference; the signature contrast with medium still holds: keeps the two heavy-judgment cases A-442d4aab 7/7 and A-a317e74b 7/15, loses ops archaeology; 2026-09-12 Addendum, 26/26 wire-verified, 23/23 hashes against the public index), deepseek-v4.1-flash @ WorkBuddy 15/23 (see below). Added 2026-09-12: gpt-5.6-luna-900k's third high run lands 15/23 again (fail-set drifts ±2 across days; board composition unchanged; sol paused at the owner's call, 9/26 unscored — the lane was completed 09-16 as one fresh full-library sweep, 15/23, zero drift). Added 2026-09-13: first WorkBuddy ACP-channel entry, two models — hy4-preview-f 17/23 joins the #2 tie (OPS 6/6 sweep, A-d9b79b46 12/12 perfect; its A-a317e74b main run hit the 1800s cap and the 3600s-cap makeup scored 14/15, tying the case's second-best published score — the only pass remains crof q38's 15/15); deepseek-v4.1-flash 15/23 misses the board but lands A-be92627f 9/9 — the first-ever pass on that case across all published entries — second verify-face pass overall (previous best 8/9; the face's first break was crof q38's A-a317e74b 15/15). Audit: 82/82 usage rows on the pinned lane; both V001 papers via a direct-acp bypass (marked runner in the manifest); 23/23 hashes against the public index. Lane caveats: no token usage reported, effort pinned via the ACP set_config_option side channel, native tool surface is bypassPermissions. hy3 halted at the owner's call, unscored. Added 2026-09-15: k3 @ Kimi official coding endpoint debuts 17/23 (15/21 on the public 21-case subset), joining the #2 tie — coding face 5/6 with two perfect scores (hard discriminator A-442d4aab 7/7, A-569dbe0d 10/10), OPS 6/6 sweep, and the riding-line vision case A-ea80d793 passed at 3.0; verification face 0/3, A-d9b79b46 failed on an incomplete deliverable, A-cdc3d11a -2. 26/26 sessioned rows verified (k3, kimi-coding) with zero stand-ins, 23/23 hashes against the public index; 26 papers ~2.8 h wall-clock sum (amber-kimi W38). Added 2026-09-16: doubao-seed-evolving @ Volcengine Ark Agent Plan debuts 16/23, joining the #3 tie — the channel is officially version-synced with Doubao-Seed-2.1-pro-0915 (the plan catalog carries no version-pinned ID; verified against official plan docs). A sharply split profile: build/text/ops/req-drift sub-scores 92/94 (97.9%) with a perfect 7/7 on the hard discriminator A-442d4aab, while review/vision/ui-build net −2 and all three verify cases hit the harness wall with zero deliveries (a 3600s-cap makeup lands in the result repo's addendum); reasoning tokens = 65% of output and mean wall 557s/paper, ~2× the GPT lane at the same band; 23/23 sessioned rows wire-verified on (doubao-seed-evolving, …/api/plan/v3) with zero stand-ins, 0/26 hash mismatches (amber-doubao W38). Added 2026-09-17: kimi-for-coding (K2.8 Preview — the ID was silently re-brained on 09-11) @ Kimi official coding endpoint debuts 14/23 (12/21 on the public subset), off the then-top-three board — coding, text, and req-drift faces match k3 paper-for-paper at 27% faster wall clock, and the defense-verify case A-be92627f 7/9 beats k3's 4/9; but the adversarial-review case A-cdc3d11a scores -17 (k3: -2; hallucination-flood class, the worst published band on that case), and the review face is judged unusable. The 3-case gap to k3 sits entirely inside the known drift / riding-line band. 26/26 papers wire-verified (clean-room profile, zero fallback stand-ins), 26/26 hashes against the public index (amber-kimi W38 Addendum 09-17). Added 2026-09-17: gpt-5.6-luna-900k's first max-band run lands 16/23 — the +1 case is ui-build A-d9b79b46's first-ever delivery at 12/12 (luna becomes the fifth published lane to max that case), but it is confounded with the same-day 09-16 harness watchdog fix (any band re-run would now deliver), so net of the confound it is 15/23, tied with high; the headline stays 15/23 (the cell-level correction is owed a same-band re-run). Vision case A-ea80d793 scores 5.0, a new published best for that case (previous record: swe-2-medium 4.0); attribution case A-a317e74b slides 14/15→7/15, the top-band backfire again (amber-gpt W38 Addendum 4). Changed 2026-09-17: the board display widens from top three to top five — #4 (15/23, seven tied lanes) and #5 (14/23, five tied lanes) appear on the table and chart for the first time. Added 2026-09-17: Qwen3.8-27B @ goldenpotato community self-hosted endpoint debuts 14/23, joining the #5 tie — a hobbyist lane on 3× V100 32GB running NVIDIA's official NVFP4 weights via a heavily patched vLLM TP3 (FP8 KV cache); build face 5/6 with a perfect 7/7 on the hard discriminator A-442d4aab, zero passes across the judgment faces (review/verify/vision/ui-build); the effort knob is proven inert (reasoning_tokens = 0 across all 57 usage records, so the row is the endpoint's default band); the endpoint was a time-limited stress test (~one day per its announcement), making this row a one-shot, non-reproducible snapshot; 26/26 papers wire-verified on (Qwen3.8-27B, custom, 27b.goldenpotato.cn) with zero stand-ins, 0/26 hash mismatches (amber-goldenpotato W38). Added 2026-09-20: step-5-preview @ stepfun plan endpoint debuts 15/23 (13/21 on the public 21-case subset), joining the #4 tie — StepFun's new flagship (600B/27B MoE, advertised 1M context + vision) on its official plan endpoint (OpenAI-compatible surface). Build side is top-tier: ops 6/6 clean, req-drift all four variants green, coding 5/6 including the library's only hard discriminator A-442d4aab at a perfect 7/7, delivery perfect; the judgement side trails: the attribution axis beats the reference anchor k3 (0.800 vs 0.667 — the same case 12/15 vs 10/15), but verify 0/3, review net −2, vision −2, ui-build void. Three disclosures travel with the row: (1) the endpoint's usage carries no reasoning_tokens (all 72 usage records 0/missing), so high is a request label only and the row is the endpoint's default band; (2) the vision face hung on the endpoint at launch (zero-byte hang; the same endpoint answered correctly for another model on the same image), so that case was first not run (not a failure) and, after the endpoint recovered, re-took the case for a real d2 = −2; (3) the three verify cases converged only under extended wall-clock caps (7200→4326s / 10800→5353s / 3600s), making the 26-paper wall clock 5.9 h, a heavy cost. 26/26 bundle hashes match the public index; 72/72 usage rows pin to (step-5-preview, stepfun) with zero substitutes (amber-stepfun W39). Added 2026-09-21: the Convergence axis's first case A-3f2a9cdd joins the library (24 cases); later the same day the owner approved folding it into the headline — of the 25 board lanes, 17 live lanes sat the make-up and all passed (fastmedium, zero loop signals), hy4 @ WorkBuddy passed on retake at 11:18 UTC after five 429s (fast 51s) and joins 18/24; d41f @ CommandCode was frozen at ∅ at the owner's call (17/23 sealed), and 6 dead/paused lanes are frozen at ∅ (CrofAI ×3, the delisted ds-v4.1-exp, paused sol, the phantom swe-2-low), lifting the completion matrix from nine axes to ten. Added 2026-09-23: claude-opus-5-5 @ Anthropic subscription lane (OAuth; subscription-quota lane, no per-token price sheet, no USD cost reported) debuts 17'/24², joining the #3 tie — a full-library first sitting on the model's own release day (09-22), scoring window ~50 minutes for 27 papers; build 5/6 including a perfect 7/7 on the hard discriminator A-442d4aab, ui-build A-d9b79b46 a perfect 12/12, req-drift 4/4 all green, vision A-ea80d793 4/5 hits with zero false positives; the short boards are stated plainly: verify 0/3 (including a true zero-delivery loss on A-be92627f), review case A-cdc3d11a judged the right verdict but flooded 18 of 22 findings as false positives for -17, attribution pin A-a317e74b 7/15. Wire audit: 30/30 exam sessions pinned to (claude-opus-5-5, Anthropic subscription lane), zero stand-ins, zero throttling, all closed cleanly (truncation forensics: 0 wronged papers); 3 first-sitting attempts across 2 cases were voided for case-level infrastructure faults and re-sat after repair, never scored; bundle hashes are published per paper in the issue against the hash-index (amber-claude W39). Added 2026-09-23 (II): full-library debuts the day after OpenAI's 6-series launch — gpt-6-sol-900k 17/24 joins the #3 tie (24 cases, zero held; attribution axis 0.93 near the leaders' band, vision 1.0 below the passing line is the weakness), gpt-6-luna-900k 15/24 joins #5 (the 5.6-era family pattern "luna ≥ sol" flips in the 6 series), and gpt-6-astra-900k's same-day double-run reproduces the board's 16/24 exactly (identical pass/fail sets — reproducibility evidence); the brain-verification gate is now hard-wired into the harness (declared brain ≠ seated brain → refuse to run); 54 ledger rows pin gpt-6-sol/luna-900k with zero stand-ins, bundle hashes against the public index (amber-gpt W39). Also added 2026-09-23: mimo-v2.6-pro @ CommandCode 17'/24³ joins the #3 tie — OPS 6/6 sweep, text face all green, tied best on A-47eea242; one case (A-d511f9e8) invalid on case-level infrastructure, held pending re-judge; its lane-mates grok-4.7 hit the quota wall unfinished and mimo-v2.6-flash is still running (amber-commandcode W39). Added 2026-09-24: space-bunny-alpha @ CommandCode 15/24 joins the #5 tie — a stealth trial model (no upstream vendor named) listed by the platform that day, full-library debut inside its free window, 27/27 papers complete with zero held and zero invalid: an ops 6/6 sweep, a perfect 12/12 on the ui-build case A-d9b79b46 (the lane's only pass this week), A-be92627f 9/9 (the lane's previous best there was no-deliverable/invalid), text face 2/3; the losses concentrate on the judgment faces plus the shared build pit (A-87c472cb 6/8 — every lane falls in it), and req-drift 3/4 (one variant self-synced a signed-decision gap that should have gone to the owner). Three papers first died on bench-side timer/grader faults with deliverables intact; offline re-judging cleared two and confirmed one true fail (re-scored, never rewritten; append-only). 27/27 manifest rows pin (stealth/space-bunny-alpha, commandcode) with zero stand-ins, bundle hashes 24/24 against the public index (amber-commandcode W39).

¹ Addendum: week labels mean the baseline test week. step-5-preview: full test 2026-09-20 (W38), publication and convergence make-up 2026-09-21 (W39); 16'/24, 4 held cells excluded from wins/losses. Luna's W37 base and W39 make-up are explained in ³. Earlier holds are not cleared.

² Claude correction: A-1fd3683a (6a980035b42f) is invalid / NA due to test-harness infrastructure: original answer not saved, zero-traffic gate misclassification. Refusal stays a diagnostic note only; the owner has withdrawn the safety-boundary deployment advice. 17'/24 = 17 wins · 6 losses · 1 void case. Details.

³ ' = contested (held for safety refusal) or invalid (infrastructure-related (test harness or scoring environment) cases: held, void or awaiting re-scoring); neither counts as a win or a loss. Every lane with NA carries an apostrophe, including frozen display rows; a hold does not settle the cause. MiMo (mimo-v2.6-pro), 17'/24: A-d511f9e8 remains invalid / NA, held for re-scoring due to scoring-environment infrastructure, not a test-harness failure. 5.6-luna high: baseline test 2026-09-07 (W37), high make-up 2026-09-21 (W39); 16'/24 = 16 wins · 4 losses · 4 held. Three exonerated_infra rulings are first published here (capability scores withheld); A-ea80d793 remains held, so Vision is NA, not a valid loss. See its correction and ledger citation. Other lanes keep their own named states.

5.6-luna high profile correction — completion is a scaled score from 0 to 1, not a pass count; NA is unscored, not zero. The original seven-model profile does not include luna. This table clarifies luna separately and does not replace another model.

Coding Delivery Ops Requirements UI Vision Defense Attribution Review Convergence
0.958 1.000 1.000 1.000 NA NA 0.292 0.667 NA 1.000

How to read a results matrix

Every issue is a single results/YYYY-Www.md built around a matrix. Five things to know:

  • Alias (A-xxxxxxxx) — the case's public handle. Internal case numbers never appear, so scores can't be reverse-engineered into case content.
  • bundle_sha — the content hash of the case bundle. Match it against the hash index: identical means the library hasn't changed.
  • ✓ / ✗ (case level) — the pass line is "all required checks green": 8/9 still fails — a defense that leaks one pin leaks.
  • Defect-hunt score (review/vision cases) — hits − false positives − flattery − verdict penalty. Positive is hard; negative is common. (Published result-repo write-ups and charts label it d2; same ruler.)
  • Comparability trio — compare only at the same effort band and same library version, and read the date; the same model name on another endpoint may be another brain. One day's number is a snapshot, not a law.

FAQ

If cases are private, why trust the scores? Trust comes from the chain, not from showing you the paper: clean-room profiles (no fallback chain), per-paper wire audits (every call reconciled; a substitute call voids the paper), a closing gate (zero pollution or nothing ships), alias + hash publishing (you can verify the library is unchanged, case by case), and a per-issue harness pin (current runs: Hermes, version + upstream commit pinned in every issue). Verifiable process is what makes up for secret cases.

Why keep cases private at all? Public corpora get eaten by training data; scores inflate and become unauditable. That's the chronic disease of public benchmarks. Private cases make leak status checkable.

Can I compare two models' scores directly? Only at the same band and same library version, ideally the same day. Cross-week comparisons must carry explicit date and band declarations (pinned in every issue); cross-repo citations likewise.

Can I reproduce or join? Not from the public repos alone yet (spec v0.2.2 draft; schemas/, profiles/, and tooling unpublished — see PLAN.md). Follow the result repos for scores; methodology questions are welcome as issues.

License

Apache-2.0 — see LICENSE.