Skip to content

Re-verify RLM note vs arXiv:2512.24601v3 — verified clean, date refreshed - #1

Open
OpenCnid wants to merge 1 commit into
mainfrom
reverify/2026-07-20
Open

Re-verify RLM note vs arXiv:2512.24601v3 — verified clean, date refreshed#1
OpenCnid wants to merge 1 commit into
mainfrom
reverify/2026-07-20

Conversation

@OpenCnid

Copy link
Copy Markdown
Owner

Re-verification: density-chain.md vs pinned source

Source: arXiv:2512.24601v3 — Recursive Language Models (Alex L. Zhang, Tim Kraska, Omar Khattab; MIT CSAIL). Pin resolves at source (arXiv API + abs page); v3 confirmed as the latest version (v1 2025-12-31, v2 2026-01-28, v3 2026-05-11). Studied fresh this session from the arXiv v3 PDF and the arXiv HTML rendering; the PDF was treated as session material and is not committed.

Result: verified clean against arXiv:2512.24601v3 on 2026-07-20; refreshed verification date only.

Discrepancy ledger

None. Every quantitative claim, benchmark name, coefficient, and locator matches the source.

Checked and confirmed against the source:

Claim Note Source Locator
Headline medians (GPT-5) 26% / 130% / 13% 26% / 130% / 13% Abstract
Input scale >1 order of magnitude "more than an order of magnitude" Abstract
RLM-Qwen3-8B gain median 28% (Abstract) / 28.3% (§1) "median of 28%" / "median of 28.3%" Abstract; §1
OOLONG-Pairs (GPT-5) 0.1 → d1 58.0 → d2 65.5 → d3 76.0 0.1 / 58.0 / 65.5 / 76.0 Table 1
OOLONG-Pairs (Qwen3-Coder d1) 23.1 23.1 Table 1; Obs. 1
BrowseComp-Plus (GPT-5) 0.0 → d1 91.3 @ $0.99±$1.22; d2 92.0 0.0 / 91.3 ($0.99±$1.22) / 92.0 Table 1; Obs. 1
Extrapolated ingestion $1.50–2.75 vs $0.99; ≥29% $1.50–$2.75; $0.99; "over 29%" Obs. 1
OOLONG gain (d1) +28.4% GPT-5 / +33.3% Qwen3-Coder 28.4% / 33.3% Obs. 1
CodeQA GPT-5 base 24.0; d2 66.0; Qwen3-Coder best d0 66.0 24.0 / 66.0 / 66.0 Table 1; Obs. 2
Claude Code (+offloading) CodeQA 62.0; BC+ 84.0; OOLONG 48.0; OP 6.5 62.0 / 84.0 / 48.0 / 6.5 Table 1
LONGCOT-MINI (GPT-5.2) 38.7 → 50.6 → 65.6 (+69.5%); LOGIC/CHESS 99.0 38.7 / 50.6 / 65.6; +69.5%; 99.0/99.0 Table 2; Obs. 5
Training 1,000 rejection-FT trajectories (LongBenchPro); median +28.3%; >3× faster matches §3.2; §1; Obs. 6
Length generalization RLVR MRCRv2 64k/2-needle → 1M/8-needle matches Obs. 6; Fig. 3
Cost profile median RLM cheaper; average dearer (outliers) matches Obs. 4
Syntax errors RLM(Qwen3-Coder) ≫ RLM(GPT-5); propagate at depth matches §5; Fig. 4
Code github.com/alexzhang13/rlm matches Abstract
License CC BY 4.0 CC BY 4.0 arXiv record

Also confirmed: authors/affiliations (all MIT CSAIL), venue, three design choices (symbolic handle / Final variable / symbolic recursion), Ω(|P|)/Ω(|P|²) framing, task-complexity taxonomy (S-NIAH / BrowseComp-Plus 1K docs / OOLONG / OOLONG-Pairs / CodeQA), model roots and depths 0–3, and the §7 limitations (guardrails, exploding sub-call costs, blocking implementation).

Note on arXiv metadata

The arXiv API <summary> field currently reads "up to two orders of magnitude" and "28.3% on average", which differs from the v3 PDF/HTML ("more than an order of magnitude" and "median of 28%"). The note pins and follows the v3 PDF, so the note is correct as written — flagging only so a future reviewer isn't misled by the metadata abstract.

Changes

  • verified_against_source: 2026-07-18 → 2026-07-20 (frontmatter + provenance block + index.json)
  • index.json updated: 2026-07-18 → 2026-07-20
  • No changes to tiers, Key results, entity ledger, or Our take — the one-way rule had nothing to correct this round.

🤖 Generated with Claude Code

Re-verified the density-chain.md note against its pinned source,
arXiv:2512.24601v3 (Recursive Language Models, Zhang/Kraska/Khattab,
MIT CSAIL), fetched fresh from arXiv. Checked every T1-T5 claim, all
14 Key results rows, Table 1, Table 2, Observations 1-6, and every
locator against the source PDF and HTML rendering.

No substantive discrepancies found. Only the verification date is
refreshed (2026-07-18 -> 2026-07-20) in the frontmatter, the
provenance block, and index.json.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant