Skip to content

docs(site): compare LHTB LoopX results with Plain and native Goal - #4708

Merged
huangruiteng merged 1 commit into
mainfrom
codex/lhtb-baseline-narrative-20260919
Sep 18, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/lhtb-baseline-narrative-20260919

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

The LHTB brief previously centered its result strip, cases, and task filters on new-versus-legacy Heartbeat. It now answers the reader’s main comparison: what the study’s LoopX 1.0.3 Heartbeat achieves relative to Plain and native Goal.

The bilingual opening, results, execution explanation, cases, and next experiments follow those two baselines. Mean deltas, relative gains, and strict win/tie/loss counts are derived from the public 46-task rows. Cases render all three scores from that same source. Historical arms remain available in a disclosure and an optional matrix view. The application-scenarios articles and study reading guide use the same framing.

The result is deliberately bounded: LoopX has higher mean reward than both baselines, matches Plain at 7/46 strict solves, and exceeds native Goal’s 4/46. Per-task losses and recorded-versus-estimated costs remain explicit. The raw experiment data, scoring, and runner are unchanged.

Validation: TypeScript/Vite build; complete frontstage share-bundle smoke and bilingual blog smoke; final production export; independent aggregate and case-score checks; EN/ZH production-browser checks of both filters (19 and 21 tasks), query/filter intersections, empty search, historical columns and row maxima, disclosure, article-to-category navigation, and 390px layouts. The old production bundle fails the new three-arm/two-baseline oracle; the current bundle passes. Public/private boundary scan and diff hygiene pass. The owner requested refinement and self-merge; the session’s no-CI-wait preference is retained.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng merged commit 4d73654 into main Sep 18, 2026
18 checks passed
@huangruiteng
huangruiteng deleted the codex/lhtb-baseline-narrative-20260919 branch September 18, 2026 16:37

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

问题是真实的读者体验问题。LHTB 简报是这项五臂研究对外的入口(/benchmarks/lhtb/),但改动前它的结果条、案例和筛选器都围绕"新版 Heartbeat vs 旧版 Heartbeat"(+0.0154、17/19/10)展开。而读者真正要判断的是:这项研究里的 LoopX 1.0.3 Heartbeat 相对 Plain 和 原生 Goal 到底多做了什么——这两条正是多数人日常会用的基线。把主要对比放在次要维度上,等于让读者自己去算差值。

我把这一页当作证据页来验证,而不是当作文案改动:所有对外发布的数字都必须能从公开的 46 题数据里复现。这一点成立(见下)。

改动思路

方向对,而且形状是最小的:对比量在组件内直接从 data.json 推导(不新增构建步骤、不新增数据文件),案例卡片改成从同一份数据读三个分数,历史两臂用 disclosure + 矩阵开关保留而不是删除,study README 用同一套口径记录这张对比表。

几处判断我认同:

  • 不删历史臂。 五臂研究的证据完整性没有被牺牲,只是降级为"历史参照",页面的 historyTitle/historyBody 与 README 的措辞都说明了这一点。
  • 口径写在页面上。 comparisonNote 明确写了"差值由 46 题原始 Reward 计算后再四舍五入、相对增幅 = 均分差 ÷ 基线均分、胜负按原始分数严格比较、平局不代表统计等价";limits 保留了替代 trial、时限不一致、seed 不可重复、成本口径不同(LoopX 记录 $551.73 vs Plain/Goal 估算 $212.97 / $87.58)这些边界。这正是这项研究最容易被误读的地方,它没有被弱化。
  • 删掉了旧的 spread 筛选器。 这是有代价的(原来能一眼筛出五臂差异大的题),但筛选标签现在直接写明"vs Plain |Δ| ≥ 0.05 / vs Goal |Δ| ≥ 0.05",与谓词一致,读者不会误解筛的是什么。scoresBody 也写清了"默认展示三者、蓝色标出可见列最高分、历史两臂可按需展开",所以默认视图的变化是被披露的。

具体改动

6 个文件、+301/-244,data.json 未被改动:

  • apps/presentation/site/src/LhtbBrief.tsx(+78/-62):新增 comparisons(逐题 new_heartbeat - baseline → 均分差、相对增幅、严格胜负平)、summaryTable(同一渲染器同时服务主表与历史表,/46 改为 study.tasks.length,表头补 scope="col")、caseCards(每个案例列出三者分数)、showHistory 开关;TableMode 由 "all" | "spread" | "heartbeat" 改为 "all" | Baseline;标题/document.title 改为两种基线的说法。
  • lhtb-copy.json(+144/-124):EN/ZH 双语重写(title、deck、resultBody、insightBody、insights、limits、filters、mechanismRows 由 5 行收敛到 3 行、新增 comparisonTitle/pairCounts/pairSolves/historyTitle 等),案例文案去掉硬编码的 before → after。
  • lhtb-brief.css(+32/-28):.lhtb-delta-strip → .lhtb-comparisons(2 列,窄屏 1 列)、.lhtb-arm-flow 由 5 列改 3 列(与 3 行 mechanismRows 对齐)、新增 .lhtb-case-scores / .lhtb-history / .lhtb-history-toggle。
  • 两个 application-scenarios 静态页(各 +11/-14):LHTB 区块改为两基线表述,表格由 5 行变 3 行,Boundary 段落补上成本差异。
  • benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md(+25/-2):新增 "Reading the baseline comparisons" 表与口径说明,并把原来"New-vs-Legacy 不是单变量因果"的边界改写为覆盖两基线的版本。

我在这个 exact head 上自己跑的验证(专用 worktree):

  • 数字复现:从 data.json 重算 —— Plain 均分差 0.073072(显示 +0.0731)、相对 17.33%(+17.3%)、胜/平/负 17/13/16、通过 7 vs 7;原生 Goal 0.047283(+0.0473)、10.56%(+10.6%)、23/13/10、通过 7 vs 4;各臂 mean_reward 与 pass_095 也与表格一致(0.4218 / 0.4475 / 0.4948、7 / 4 / 7)。简报、两个静态页、README 三处全部一致。
  • 构建:npm run build(tsc --noEmit && vite build)通过。
  • 已提交的守卫:frontstage-share-bundle-smoke: ok(含 /loopx/ base 下的静态导出与全部相对链接校验)、blog-bilingual-index-smoke: ok (3 paired articles)。注:本地跑 share-bundle smoke 需要 site/dashboard 的 node_modules 与 3.10+ 的 python3,我用 checkout 的 3.13 解释器做了 PATH shim;这个 ref 上 CI 的 presentation/deploy 是 skipping,所以这一步 CI 并不覆盖。
  • CI:gh pr checks 4708 在该 head 上 25 项全 pass,无失败/等待。
  • 案例数字:8 个案例的三臂分数与文案一致(Tabular 0.000/0.205/0.743、Sokoban 0.271/0.594/0.890、PoC 0.000/0.892/0.892、DuckDB 0.000/0.770/0.000 等)。

对主干的风险

改动面只在展示层,data.json、评分、runner、CLI、权限边界一律未动,回退成本很低。两条非阻塞记录项:

P3:三处发布同一组数字,但只有一处是从 data.json 推导的,另外两处没有守卫。 简报在组件里算,而两个 application-scenarios 静态页(0.0731/17.3%/17-13-16、0.0473/10.6%/23-13-10、0.4218/0.4475/0.4948、7/4)和新增的 README 表(README.md:35-50)都是手写复述。我把它们逐项重算过,现在完全一致——所以这不是当前错误,而是缺一道守卫:blog-bilingual-index-smoke.mjs 只校验双语配对、canonical/hreflang、section id 一致与标题存在;frontstage-share-bundle-smoke.mjs 只校验 bundle 结构、链接完整与脚本/样式接线。将来任何一次数据修订(改一个 reward、补一条替代 trial、改 pass_095)都会让简报自动更新、而两个静态页与 README 继续发布旧数字,形成两个公开页面互相矛盾且无人告警。最小修复:加一个小 oracle,从 data.json 重算这两个基线对比并断言静态页/README 的表格与差值一致(或直接生成)。这样正文里"same source"的说法才是持久的,而不只是编辑期的。

P3:案例卡片对"复制文案里的 task 名一定存在于数据里"用了非空断言,复制写错会直接崩页。 caseCards 用 study.tasks.find((item) => item.task === task)! 取行再读 row[arm](LhtbBrief.tsx:98-100)。task 名来自双语 copy 文件(gainCases/lossCases 各 4 条),而没有任何校验:tsc --noEmit 只看到 string[][],也没有测试或 smoke 把名字对回 data.json。今天 8 个名字全都能解析(我逐个核对过),所以这是潜在健壮性问题而非现存缺陷;但失效方式是公开且难看的——任一门语言里一个 id 改名或拼错,就会在渲染时抛错,这条 SPA 路由上整页简报都会挂掉,而不是只显示一张坏卡片。最小修复:把非空断言换成显式查找(跳过并报告未知 task,或带 task 名抛出),并在已有 smoke 里加两行断言,检查双语 gainCases/lossCases 的 task id 都存在于 data.json。

另有两点观察(不作为 finding):一是这次改动动了公开页面的首屏(LHTB 简报的 hero 与标题),而 PR 里没有保留视觉预览产物;作者本人提出并自合并了这次精修,所以批准是存在的,但如果留一张截图或托管预览,"首屏评审门"事后会更可审计。二是 lhtb-copy.json 的 benchmarkSetupNote 仍写 "five-arm study / 五臂研究",在 3 臂主视图下依然准确(历史两臂仍在页面上),我没有把它算作陈旧文案。

我的整体评价

这是一次正确的重新定位:把公开简报的主要对比换成读者真正要判断的那一个,并且没有靠删证据来简化——历史两臂保留在 disclosure 与矩阵开关里,边界与成本口径写得更清楚而不是更含糊,案例分数从"编辑手写的 before → after"改成"从研究数据读三臂分数",这本身就是去重复的方向。我把全部对外数字从 data.json 重算了一遍,简报、两个静态页与 README 三处现在完全一致;站点构建与两个已提交的 smoke 在本 head 上通过,CI 25 项全绿。

两条 P3(静态复述缺守卫、案例查找缺兜底)都不阻塞这次 post-merge audit:它们既不改变当前发布内容,也不影响任何运行时、评分或权限行为;但都很便宜且值得补,第一条尤其重要——它是"同一个结论在多个公开面上被手写多份"的典型入口,而这次新增的 README 表又多了第三份。

English verdict: APPROVE (exact head 9d42520)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant