Skip to content

Show Oh's full 500-question LongMemEval-S study on benchmarks - #153

Merged
0thernet merged 3 commits into
mainfrom
claude/benchmarks-longmemeval-500
Sep 27, 2026
Merged

0thernet merged 3 commits into
mainfrom
claude/benchmarks-longmemeval-500

Conversation

@0thernet

Copy link
Copy Markdown
Member

Summary

  • Vendors docs/evaluations/oh/memory-longmemeval-s-500-v1.json byte for byte from hraness/oh 21c500c (PR #195) and registers it in sources.json.
  • Adds a typed LongMemEval-S study in site/wordcell/oh-evidence.ts. Every displayed figure is derived from the vendored JSON, with validation.
  • /benchmarks: the Oh section opens with the full-500 chart (Oh semantic 88.87% vs BM25 86.13%, mean of three runs). The pre-registered 2-of-3 measure is +2.8 points with a 95% interval of 0.0 to 5.6, and the page says it does not rule out a tie. The section adds a by-type table, a limits list quoting Oh's own language, and then LoCoMo. The 60-question pilot moves under matched comparisons as a smaller, earlier development pilot and remains the only matched Oh-vs-Supermemory run. The hero, description, and reproduce links are updated.
  • The in-sample lab pipeline figure (93.07%) appears once, on /benchmarks only. It is explicitly framed as neither Oh's nor Wordcell's score and as not in the Oh package.
  • Sweep: /compare/supermemory, README, CHANGELOG, docs/evidence.md, the launch blog post (via evidence placeholders), and llms.txt.
  • Makes no claim that Oh beats Supermemory or other frameworks. These are Oh retrieval benchmarks, not measurements of Wordcell search.

Test plan

  • cd site && bun run check
  • Root bun run check: 1927 pass, 1 fail. The failure is a 5 s timeout in src/evaluation-evidence.test.ts, a file this branch does not touch; it passes in isolation (13/0). Host load average was about 50.

🤖 Generated with Claude Code

Vendor Oh's memory-longmemeval-s-500-v1.json (hraness/oh 21c500c) byte for
byte, parse it into a typed BenchmarkStudy, and lead the Oh section of
/benchmarks with it: Oh semantic 88.87% vs BM25 86.13% (mean of three runs),
with the pre-registered 2-of-3 measure (+2.8 points, 95% interval 0.0 to 5.6)
stated as not ruling out a tie. The in-sample lab pipeline figure is framed
as neither Oh's nor Wordcell's score. The 60-question pilot stays as the only
matched Oh-vs-Supermemory run. README, changelog, evidence doc, compare page,
blog post, and llms.txt updated to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@0thernet
0thernet enabled auto-merge (squash) September 27, 2026 02:06
@vercel

vercel Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
wordcell Ready Ready Preview Sep 27, 2026 3:29am UTC

Request Review

Move the benchmarks changelog entry to a new Unreleased section above 0.23.0,
name LongMemEval-S in the Oh section heading, and regenerate generated files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@0thernet
0thernet merged commit 7f5200a into main Sep 27, 2026
9 checks passed
@0thernet
0thernet deleted the claude/benchmarks-longmemeval-500 branch September 27, 2026 03:31

This branch was successfully deployed

1 active deployment
Preview — e59341b1 Deployed Sep 27, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant