From fd90b2528e9f07331cf4f5b99e784527a9f9bf19 Mon Sep 17 00:00:00 2001 From: shangzh0 <2586756592@qq.com> Date: Sat, 19 Sep 2026 01:49:58 +0800 Subject: [PATCH] fix(site): correct native Goal token accounting Signed-off-by: shangzh0 <2586756592@qq.com> --- .../site/public/blog/application-scenarios/index.html | 2 +- .../site/public/blog/zh/application-scenarios/index.html | 2 +- apps/presentation/site/src/lhtb-copy.json | 4 ++-- benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json | 4 ++-- 4 files changed, 6 insertions(+), 6 deletions(-) diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index 9e4040a26d..16df33e25c 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -98,7 +98,7 @@

LHTB: what does LoopX accomplish beyond Plain and native Goal?

Observation: against Plain, LoopX gains 0.0731 mean reward (17.3%), with 17 wins, 13 ties and 16 losses, and the same solve count. Against native Goal, it gains 0.0473 (10.6%), with 23 wins, 13 ties and 10 losses, and 7 solves versus 4. Counts compare published unrounded scores strictly; higher mean reward does not imply better results on every task.

Insight: Plain advances within a session, native Goal retains the objective, and LoopX also organizes subsequent work through registry and Todos. Tabular and Sokoban improve over both baselines; the PoC task ties native Goal; DuckDB scores below it. The questions are whether durable work state reliably adds useful progress, and whether acceptance and rollback preserve that progress as a valid result.

-

Boundary: LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $87.58. Accounting also differs; equal-budget efficiency gains are not established.

+

Boundary: LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $344.25. Accounting also differs; equal-budget efficiency gains are not established.

diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index 6bf17e6dd8..a96466afdf 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -95,7 +95,7 @@

LHTB:比 Plain 和原生 Goal 多做成了什么?

观察:相对 Plain,LoopX 均分增加 0.0731(17.3%),逐题为 17 胜、13 平、16 负,通过数相同;相对原生 Goal,均分增加 0.0473(10.6%),为 23 胜、13 平、10 负,通过数从 4 到 7。胜负按公开原始分数严格比较;均分提升不代表每题都更强。

Insight:Plain 在会话内推进,原生 Goal 保留目标,LoopX 进一步用 Registry 与 Todo 组织后续工作。Tabular 与 Sokoban 相对两种基线均有收益;PoC 与原生 Goal 持平;DuckDB 却低于原生 Goal。值得验证的是,持久工作状态能否稳定增加有效进展,以及验收和回滚能否把进展变成保得住的结果。

-

范围:这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $87.58;计费口径也不一致,尚不能宣称同预算效率优势。

+

范围:这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $344.25;计费口径也不一致,尚不能宣称同预算效率优势。

diff --git a/apps/presentation/site/src/lhtb-copy.json b/apps/presentation/site/src/lhtb-copy.json index d7a087abf9..9a2e5862d6 100644 --- a/apps/presentation/site/src/lhtb-copy.json +++ b/apps/presentation/site/src/lhtb-copy.json @@ -127,7 +127,7 @@ "Model and reasoning effort match, but transport, session lifetime, and scheduling differ between LoopX and the baselines. No single mechanism is isolated.", "Effective aggregates include replacement trials, some with longer LoopX limits. This is not a strictly matched-budget comparison.", "Hidden evaluation grades the final result; continuation does not automatically reveal an unknown correct solution. Recovering work and solving the task remain distinct.", - "Accounting differs: LoopX records $551.73, while Plain and native Goal estimate $212.97 and $87.58. A higher mean does not establish better efficiency." + "Accounting differs: LoopX records $551.73, while Plain and native Goal estimate $212.97 and $344.25. A higher mean does not establish better efficiency." ], "nextTitle": "Three questions worth testing", "nextBody": "At equal budgets, which tasks does LoopX finish beyond native Goal? Do its larger gains over Plain survive repeated runs? Can independent acceptance and rollback reduce losses like DuckDB?", @@ -282,7 +282,7 @@ "模型与推理强度一致,但 LoopX 与基线同时存在传输、会话和调度策略差异,不能把收益归因于某一个机制。", "有效聚合包含替代 trial,部分 LoopX 替代运行采用更长时限;当前结果不是严格匹配预算的对照。", "隐藏评测负责最终评分,续跑不能自动获得未知的正确解;恢复能力与任务本身能否解决是两件事。", - "成本记录口径不同;LoopX 记录 $551.73,Plain 与原生 Goal 分别估算 $212.97、$87.58。均分更高不等于效率更高。" + "成本记录口径不同;LoopX 记录 $551.73,Plain 与原生 Goal 分别估算 $212.97、$344.25。均分更高不等于效率更高。" ], "nextTitle": "最值得追问的三个问题", "nextBody": "相同预算下,LoopX 还能比原生 Goal 多完成哪些任务?对 Plain 的较大增益能否跨重复运行保留?独立验收和回滚能否减少 DuckDB 这类退步?", diff --git a/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json b/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json index c77d70f324..6e370b11a1 100644 --- a/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json +++ b/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json @@ -25,8 +25,8 @@ "label": "Native Goal", "mean_reward": 0.4475494760979694, "pass_095": 4, - "tokens": 104495564, - "estimated_cost_usd": 87.58, + "tokens": 293812465, + "estimated_cost_usd": 344.25, "runtime": "Codex app-server; native persistent Goal" }, "ssh_goal": {