Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -87,22 +87,19 @@ <h2>Benchmarks: start with LHTB, then three supporting signals</h2>
<p>Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.</p>

<div class="study-signal" id="benchmark-lhtb">
<h3>LHTB: recover progress and preserve working results</h3>
<p class="study-setting">GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell</p>
<p><strong>What LHTB tests:</strong> Long-Horizon Terminal-Bench places an agent in a stateful container and asks it to sustain work across hundreds of dependent terminal actions. Its official introduction uses nine categories: software and reverse engineering; scientific computing and simulation; earth, climate, and energy; multimodal and imaging analysis; research reproduction and ML; systems, performance, and security; interactive games; APEX professional workflows; and logic and constraint puzzles.</p>
<p>Tasks include migrating an old framework, recovering data from scientific figures, reproducing paper experiments, handling investment-banking or legal matters, playing 2048 turn by turn, and searching for puzzle solutions. <a href="../../benchmarks/lhtb/?lang=en#categories">Nine categories and representative tasks →</a></p>
<p><strong>How it scores:</strong> hidden verifiers check final artifacts or replayable outcomes and award continuous reward from 0 to 1. This study counts reward ≥ 0.95 as solved, keeping partial progress distinct from full acceptance.</p>
<div class="table-scroll"><table class="benchmark-table"><caption>LHTB five-arm results: mean reward and strict solves</caption><thead><tr><th scope="col">Execution mechanism</th><th scope="col">Mean reward</th><th scope="col">Solved / 46</th></tr></thead><tbody>
<h3>LHTB: what does LoopX accomplish beyond Plain and native Goal?</h3>
<p class="study-setting">GPT-5.6 Sol / max · 46 matched tasks · one effective run per task-arm cell</p>
<p><strong>What LHTB tests:</strong> Long-Horizon Terminal-Bench asks an agent to complete hundreds of dependent terminal actions in a stateful container. Its 46 tasks cover nine categories: software and reverse engineering; scientific computing; earth, climate and energy; multimodal work; research reproduction; systems, performance and security; games; APEX professional workflows; and logic puzzles. Examples include framework migration, paper reproduction, investment-banking deliverables and playing 2048. <a href="../../benchmarks/lhtb/?lang=en#categories">Categories and representative tasks →</a></p>
<p><strong>How it scores:</strong> hidden verifiers grade final artifacts or replayable outcomes on a 0–1 reward scale. This study counts ≥ 0.95 as solved, reporting average progress and full acceptance separately.</p>
<div class="table-scroll"><table class="benchmark-table"><caption>LHTB: LoopX Heartbeat and two baselines</caption><thead><tr><th scope="col">Execution</th><th scope="col">Mean reward</th><th scope="col">Solved / 46</th></tr></thead><tbody>
<tr><th scope="row">Plain</th><td>0.4218</td><td>7</td></tr>
<tr><th scope="row">Native Goal</th><td>0.4475</td><td>4</td></tr>
<tr><th scope="row">LoopX SSH-Goal</th><td>0.4678</td><td>6</td></tr>
<tr><th scope="row">Legacy Heartbeat</th><td>0.4794</td><td>7</td></tr>
<tr><th scope="row">New Heartbeat</th><td>0.4948</td><td>7</td></tr>
<tr><th scope="row">LoopX Heartbeat (1.0.3)</th><td>0.4948</td><td>7</td></tr>
</tbody></table></div>
<p><strong>Observation:</strong> New Heartbeat has the highest mean reward, 0.0154 above Legacy Heartbeat. Its strict solve count matches Plain and Legacy at 7/46. Higher average progress has not translated into more fully solved tasks.</p>
<p><strong>Insight:</strong> each new Heartbeat wake starts a fresh executor, recovering progress from the workspace, registry, and Todos. The brief reports a consulting task resumed at S27 after interruption, alongside a DuckDB regression where later optimization broke correctness. The next questions are whether externalized state reliably enables recovery and whether checkpoints and rollback preserve working results.</p>
<p class="signal-boundary"><strong>Boundary:</strong> transport, session lifetime, LoopX version, and replanning change together. Effective runs include replacement trials, some with longer time limits. There are no repeated seeds, and cost telemetry differs, so this does not isolate a mechanism or establish an equal-budget efficiency advantage.</p>
<p class="study-link"><a href="../../benchmarks/lhtb/?lang=en">Full study: five mechanisms, gains and losses, and the 46-task matrix →</a></p>
<p><strong>Observation:</strong> against Plain, LoopX gains 0.0731 mean reward (17.3%), with 17 wins, 13 ties and 16 losses, and the same solve count. Against native Goal, it gains 0.0473 (10.6%), with 23 wins, 13 ties and 10 losses, and 7 solves versus 4. Counts compare published unrounded scores strictly; higher mean reward does not imply better results on every task.</p>
<p><strong>Insight:</strong> Plain advances within a session, native Goal retains the objective, and LoopX also organizes subsequent work through registry and Todos. Tabular and Sokoban improve over both baselines; the PoC task ties native Goal; DuckDB scores below it. The questions are whether durable work state reliably adds useful progress, and whether acceptance and rollback preserve that progress as a valid result.</p>
<p class="signal-boundary"><strong>Boundary:</strong> LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $87.58. Accounting also differs; equal-budget efficiency gains are not established.</p>
<p class="study-link"><a href="../../benchmarks/lhtb/?lang=en">Full study: two baselines, gains and losses, 46 tasks, and historical context →</a></p>
</div>

<div class="study-signal" id="benchmark-marathon">
Expand Down Expand Up @@ -201,7 +198,7 @@ <h2>Evolution: add one verifiable capability at a time</h2>
<h2>Public sources and further reading</h2>
<p class="small-note">The technical content is derived only from public repository material. Source and data links are pinned to the revisions read for this article; historical experiments retain their own versions. This page introduces no new experimental results.</p>
<ol class="source-list">
<li><a href="../../benchmarks/lhtb/?lang=en">LHTB: five long-horizon execution mechanisms</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json">public aggregates for 46 tasks</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md">setup and evidence boundary</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json">mechanisms and cases reported in the brief</a>; <a href="https://zli12321.github.io/LHTB/index.html">official benchmark</a>. See the study for contributor attribution.</li>
<li><a href="../../benchmarks/lhtb/?lang=en">LHTB: compared with Plain and native Goal</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json">public aggregates for 46 tasks</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md">setup and evidence boundary</a>; <a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json">mechanisms and cases reported in the brief</a>; <a href="https://zli12321.github.io/LHTB/index.html">official benchmark</a>. See the study for contributor attribution.</li>
<li><a href="../../benchmarks/swe-marathon/">SWE-Marathon: continuous self-verification</a>; <a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/README.md">setup, positive and negative cases, and limitations</a>; <a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/data.json">public aggregate data</a>. See the study for contributors and case provenance.</li>
<li><a href="../../benchmarks/deepswe-sol/">DeepSWE × Sol: from continued execution to valid delivery</a> (Chinese). The standalone brief covers historical results over 113 tasks, mechanism diagrams, and pinned primary sources; research archive contribution: <a href="https://github.com/huangruiteng/loopx/pull/4502">@gwh6669999, #4502</a>.</li>
<li><a href="../../benchmarks/deepswe/behavior-discovery/">DeepSWE × V4 Flash max: from hints to behavior</a>; <a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/README.md">disclosure boundary</a>; <a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/index.html">charts, cases, and metrics at the pinned revision</a>.</li>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -84,22 +84,19 @@ <h2>Benchmark:从 LHTB 看长程工作,再看三个补充信号</h2>
<p>先看跨领域长程任务中的恢复与回退,再用三项软件工程研究补充续跑、交付和验证行为。四项研究的任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。</p>

<div class="study-signal" id="benchmark-lhtb">
<h3>LHTB:工作能恢复,已有成果也要保住</h3>
<p class="study-setting">GPT-5.6 Sol / max · 46 个匹配任务、五种执行机制 · 每个任务 × 实验臂一条有效运行</p>
<p><strong>LHTB 考什么:</strong>Long-Horizon Terminal-Bench 把 Agent 放进有状态的容器环境,要求它跨数百个相互依赖的终端动作完成工作。官方介绍页分为九类:软件与逆向工程、科学计算与仿真、地球/气候/能源、多模态与图像分析、研究复现与机器学习、系统/性能/安全、交互式游戏、APEX 专业工作流、逻辑与约束谜题。</p>
<p>具体会遇到旧版框架迁移、从科学图还原数据、复现论文实验、投行或法律事项,以及逐步玩 2048、搜索谜题解。<a href="../../../benchmarks/lhtb/?lang=zh#categories">九类任务与代表例子 →</a></p>
<p><strong>怎样评分:</strong>隐藏验证器检查最终产物或可重放结果,给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。因此,部分进展与完整验收可以分开观察。</p>
<div class="table-scroll"><table class="benchmark-table"><caption>LHTB 五臂结果:平均 Reward 与严格通过数</caption><thead><tr><th scope="col">执行机制</th><th scope="col">平均 Reward</th><th scope="col">通过 / 46</th></tr></thead><tbody>
<h3>LHTB:比 Plain 和原生 Goal 多做成了什么?</h3>
<p class="study-setting">GPT-5.6 Sol / max · 46 个匹配任务 · 每题每臂一条有效运行</p>
<p><strong>LHTB 考什么:</strong>Long-Horizon Terminal-Bench 要求 Agent 在有状态的容器中,跨数百个有依赖的终端动作完成工作。46 题覆盖九类:软件与逆向工程、科学计算、地球/气候/能源、多模态、研究复现、系统/性能/安全、游戏、APEX 专业工作流、逻辑谜题。比如迁移旧框架、复现论文、完成投行交付或逐步玩 2048。<a href="../../../benchmarks/lhtb/?lang=zh#categories">类别与代表任务 →</a></p>
<p><strong>怎样评分:</strong>隐藏验证器检查最终产物或可重放结果,给出 0–1 的 Reward;本研究以 ≥ 0.95 计为通过。平均进展与完整验收分别报告。</p>
<div class="table-scroll"><table class="benchmark-table"><caption>LHTB:LoopX Heartbeat 与两种基线</caption><thead><tr><th scope="col">执行方式</th><th scope="col">平均 Reward</th><th scope="col">通过 / 46</th></tr></thead><tbody>
<tr><th scope="row">Plain</th><td>0.4218</td><td>7</td></tr>
<tr><th scope="row">原生 Goal</th><td>0.4475</td><td>4</td></tr>
<tr><th scope="row">LoopX SSH-Goal</th><td>0.4678</td><td>6</td></tr>
<tr><th scope="row">旧版 Heartbeat</th><td>0.4794</td><td>7</td></tr>
<tr><th scope="row">新版 Heartbeat</th><td>0.4948</td><td>7</td></tr>
<tr><th scope="row">LoopX Heartbeat(1.0.3)</th><td>0.4948</td><td>7</td></tr>
</tbody></table></div>
<p><strong>观察:</strong>新版 Heartbeat 平均 Reward 最高,比旧版高 0.0154;严格通过数与 Plain、旧版相同,均为 7/46。平均进展增加,尚未转化为更多完全通过的任务。</p>
<p><strong>Insight:</strong>新版每次唤醒启动新的执行器,依靠工作区、Registry 与 Todo 恢复进度。简报记录了咨询任务中断后从 S27 接续的案例,也记录了 DuckDB 后续优化破坏正确性的退步。值得继续验证的是:外部化状态能否稳定支持恢复,以及检查点和回滚能否保住已有成果。</p>
<p class="signal-boundary"><strong>范围:</strong>新旧同时改变了传输、会话生命周期、LoopX 版本和 replan 策略;有效运行包含替代 trial,部分替代运行时限更长。没有重复 seed,成本记录口径也不同,不能据此归因单个机制或宣称同预算效率优势。</p>
<p class="study-link"><a href="../../../benchmarks/lhtb/?lang=zh">研究全文:五种执行机制、正反案例与 46 题矩阵 →</a></p>
<p><strong>观察:</strong>相对 Plain,LoopX 均分增加 0.0731(17.3%),逐题为 17 胜、13 平、16 负,通过数相同;相对原生 Goal,均分增加 0.0473(10.6%),为 23 胜、13 平、10 负,通过数从 4 到 7。胜负按公开原始分数严格比较;均分提升不代表每题都更强。</p>
<p><strong>Insight:</strong>Plain 在会话内推进,原生 Goal 保留目标,LoopX 进一步用 Registry 与 Todo 组织后续工作。Tabular 与 Sokoban 相对两种基线均有收益;PoC 与原生 Goal 持平;DuckDB 却低于原生 Goal。值得验证的是,持久工作状态能否稳定增加有效进展,以及验收和回滚能否把进展变成保得住的结果。</p>
<p class="signal-boundary"><strong>范围:</strong>这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $87.58;计费口径也不一致,尚不能宣称同预算效率优势。</p>
<p class="study-link"><a href="../../../benchmarks/lhtb/?lang=zh">研究全文:两种基线、正反案例、46 题矩阵与历史参照 →</a></p>
</div>

<div class="study-signal" id="benchmark-marathon">
Expand Down Expand Up @@ -198,7 +195,7 @@ <h2>演进:每一步增加一种可验证的能力</h2>
<h2>公开依据与延伸阅读</h2>
<p class="small-note">技术内容仅据公开仓库材料整理。源码与数据引用固定到本次读取修订;历史实验使用其自身版本。本页没有新增实验结果。</p>
<ol class="source-list">
<li><a href="../../../benchmarks/lhtb/?lang=zh">LHTB:五种长程执行机制</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json">46 题公开聚合数据</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md">实验设置与范围</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json">简报中的机制与正反案例</a>;<a href="https://zli12321.github.io/LHTB/index.html">LHTB 官方说明</a>。研究贡献者见原文。</li>
<li><a href="../../../benchmarks/lhtb/?lang=zh">LHTB:与 Plain、原生 Goal 的对比</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json">46 题公开聚合数据</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md">实验设置与范围</a>;<a href="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json">简报中的机制与正反案例</a>;<a href="https://zli12321.github.io/LHTB/index.html">LHTB 官方说明</a>。研究贡献者见原文。</li>
<li><a href="../../../benchmarks/swe-marathon/?lang=zh">SWE-Marathon:持续自我验证</a>;<a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/README.md">设置、正反个案与局限</a>;<a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/data.json">公开聚合数据</a>。贡献者与案例来源见研究原文。</li>
<li><a href="../../../benchmarks/deepswe-sol/">DeepSWE × Sol:从继续执行到有效交付</a>。独立研究简报,含 113 任务历史结果、机制图与固定修订的一手来源;研究归档贡献:<a href="https://github.com/huangruiteng/loopx/pull/4502">@gwh6669999,#4502</a>。</li>
<li><a href="../../../benchmarks/deepswe/behavior-discovery/">DeepSWE × V4 Flash max:从提示到行为</a>;<a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/README.md">披露范围</a>;<a href="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/index.html">固定修订的图表、案例与指标</a>。</li>
Expand Down
Loading
Loading