fix(site): correct native Goal token accounting - #4714
huangruiteng merged 1 commit into
Conversation
Signed-off-by: shangzh0 <2586756592@qq.com>
huangruiteng
left a comment
There was a problem hiding this comment.
审阅对象:PR #4714(已合并),exact head fd90b2528e9f07331cf4f5b99e784527a9f9bf19(作者 shangzh0),2026-09-18T17:53:11Z 合并。本文是该 exact head 的 post-merge audit。
动机
已发布的研究与文章把 Native Goal 的用量和成本写低了:benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json 记为 104,495,564 tokens / $87.58,而同一张表里其他臂的等效费率都在 $0.85–$1.06 每百万 token 区间——$87.58 对应约 $0.30/M,是全表唯一的离群值。这个数字还出现在站点的 LHTB 简报文案和 application-scenarios 文章的范围段里,用来和 LoopX 记录的 $551.73 做对比,因此低估基线成本会把 LoopX 的支出显得比实际更失衡,也削弱了该研究"各臂计费口径不同"的既有边界声明。
改动思路
没有引入新机制,只做一次同一 diff 内、跨三处表面的数据更正:把研究数据文件作为唯一来源,同步更新站点文案(lhtb-copy.json,中英各一处)与静态文章(application-scenarios/index.html,中英各一处)。之所以必须同 diff 改完三处,是因为只改数据文件会让简报和文章继续引用 $87.58,与它们链接的表格自相矛盾;只改一种语言则会让两个版本互相打架。研究的结论、奖励矩阵与通过数都不动,成本仍按 README 的既有口径保持为"从留存 token 遥测估算、仅作描述"。
具体改动
4 个文件、+6/-6,全部是单行数值替换:benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json:28 的 tokens 与 estimated_cost_usd;apps/presentation/site/src/lhtb-copy.json 第 130 行(英)与第 285 行(中)的 limitations 条目;apps/presentation/site/public/blog/application-scenarios/index.html:101(英)与其中文版第 98 行的范围段落。
我在 origin/main(6f0f9b06)上做了全树检索:87.58、104495564、104,495,564 已无任何残留,344.25、293812465 恰好出现在上述六处。费率一致性上,修正后 Native Goal 为 $1.1717/M,与 Plain $0.8461/M、SSH-Goal $1.0562/M、Legacy Heartbeat $0.8593/M 同处一个区间,从"全表最低的离群值"回到可比范围;中英两份文案与两版文章的数字完全一致。需要说明的证据边界:产生 293,812,465 这条 token 遥测不在仓库内,因此 token 数本身无法在仓库内重算,本轮只验证了传播完整性、跨表面一致性与费率合理性。
对主干的风险
纯数据与展示文本,无运行时、权限或持久状态影响,回滚即回退四个文件。唯一的结构性风险是这类数字目前靠人工在三处镜像:将来若只改其中一处,公开简报、文章与研究表格会再次互相矛盾。仓库当前的检查(benchmark/tests/test_publication_scope.py 覆盖聚合语义)不会发现这种镜像漂移,本轮只能靠全树检索来兜底。此外,一个"内部一致但错误"的数字能通过所有现有检查;本次之所以接受该更正,是因为它把该臂从明显离群拉回同类区间,而不是因为有任何检查证明了它。
我的整体评价
结论 APPROVE。这是一次小而完整的更正:来源唯一、三处表面同 diff 同步、中英一致,且修正后的费率与同表其他臂可比,同时保留了研究原有的"计费口径不同、成本仅作描述"的边界。合并 head 上没有留下任何评审记录,本文补上该 exact head 的审计记录。残留风险只有一条且已如实写明——token 遥测在仓库外,无法在此独立重算;若后续拿到原始用量,应以同一 diff 再核对三处。
English verdict: APPROVE - Audit of merged exact head fd90b25 of PR #4714. The change corrects the published Native Goal accounting from 104,495,564 tokens / $87.58 to 293,812,465 tokens / $344.25 across the study data file, the site brief copy and both language editions of the application-scenarios article. I verified on origin/main that no stale occurrence of 87.58 or 104495564 survives anywhere in the tree and that the corrected values appear in exactly the six intended places, that both locales agree, and that the corrected effective rate (~$1.17 per million tokens) sits inside the range of the other arms ($0.85-$1.06) whereas the previous figure implied ~$0.30/M, far below every peer. The remaining evidence boundary is disclosed in the review: the token telemetry that produced the corrected count lives outside the repository, so the count itself could not be recomputed here, and the study already documents cost as an estimate and treats it as descriptive. No blocking finding. Note that this head carried no review record at merge time; this exact-head audit supplies it.
Summary
Accounting
The previous aggregate retained only the final app-server phase output for multi-phase trials. The corrected total sums the maximum cumulative usage for each unique archived session: 290,319,642 input tokens, including 255,677,676 cached input tokens and 33,550,772 cache-write tokens, plus 3,492,823 output tokens. The estimate uses the published GPT-5.6 Sol short-context Standard rates: $4/M uncached input, $0.40/M cached input, $5/M cache writes, and $20/M output.
Rewards, solve counts, other arms, and mechanism claims are unchanged.
Validation