From 68b480e38ee6b615d1ac182aa6f54087cbf7c723 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Fri, 18 Sep 2026 23:52:54 +0800 Subject: [PATCH 1/2] docs(blog): lead application benchmarks with LHTB Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- .../blog/application-scenarios/index.html | 26 ++++++++++++++++--- .../blog/zh/application-scenarios/index.html | 26 ++++++++++++++++--- examples/export-frontstage-share-bundle.mjs | 2 +- examples/frontstage-share-bundle-smoke.mjs | 1 + 4 files changed, 46 insertions(+), 9 deletions(-) diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index b7cd356482..22468d7d4f 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -83,8 +83,25 @@
The studies below observe continuation, delivery, and validation behavior under different tasks, model settings, and measurement rules. Each needs to be read on its own terms. The current evidence does not establish a universal LoopX gain.
+Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.
+ +GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell
+Long-Horizon Terminal-Bench spans nine categories, including software, games, science, and professional workflows. Final artifacts receive a continuous reward from 0 to 1; this study counts reward ≥ 0.95 as solved. Partial progress and full acceptance can therefore be examined separately.
+| Execution mechanism | Mean reward | Solved / 46 |
|---|---|---|
| Plain | 0.4218 | 7 |
| Native Goal | 0.4475 | 4 |
| LoopX SSH-Goal | 0.4678 | 6 |
| Legacy Heartbeat | 0.4794 | 7 |
| New Heartbeat | 0.4948 | 7 |
Observation: New Heartbeat has the highest mean reward, 0.0154 above Legacy Heartbeat. Its strict solve count matches Plain and Legacy at 7/46. Higher average progress has not translated into more fully solved tasks.
+Insight: each new Heartbeat wake starts a fresh executor, recovering progress from the workspace, registry, and Todos. The brief reports a consulting task resumed at S27 after interruption, alongside a DuckDB regression where later optimization broke correctness. The next questions are whether externalized state reliably enables recovery and whether checkpoints and rollback preserve working results.
+Boundary: transport, session lifetime, LoopX version, and replanning change together. Effective runs include replacement trials, some with longer time limits. There are no repeated seeds, and cost telemetry differs, so this does not isolate a mechanism or establish an equal-budget efficiency advantage.
+Full study: five mechanisms, gains and losses, and the 46-task matrix →
+Full study: requirement coverage, behavior cases, and the long-duration slice →
The current signal is that useful continuation, reliable delivery, and counterexample-driven repair deserve further study.
-The studies offer local positive observations while exposing cost, failure, and attribution problems. Matched budgets and repeated experiments are needed to determine which gains reproduce reliably.
+Recover progress, preserve working results, and check whether continuation closes acceptance gaps.
+The four studies offer local positive observations while exposing regressions, cost, and attribution problems. Matched budgets and repeated experiments should examine recovery, rollback, delivery, and counterexample-driven repair separately.
Withdrawn SSH Goal and Codex CLI scores and conclusions from SWE-Marathon and DeepSWE × Sol are not used in these comparisons.
The technical content is derived only from public repository material. Source and data links are pinned to the revisions read for this article; historical experiments retain their own versions. This page introduces no new experimental results.
三项研究分别观察续跑、交付和验证行为。任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。
+先看跨领域长程任务中的恢复与回退,再用三项软件工程研究补充续跑、交付和验证行为。四项研究的任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。
+ +GPT-5.6 Sol / max · 46 个匹配任务、五种执行机制 · 每个任务 × 实验臂一条有效运行
+Long-Horizon Terminal-Bench 覆盖软件、游戏、科学与专业工作流等九类任务,按最终产物给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。它把长程工作中的部分进展与完整验收分开呈现。
+| 执行机制 | 平均 Reward | 通过 / 46 |
|---|---|---|
| Plain | 0.4218 | 7 |
| 原生 Goal | 0.4475 | 4 |
| LoopX SSH-Goal | 0.4678 | 6 |
| 旧版 Heartbeat | 0.4794 | 7 |
| 新版 Heartbeat | 0.4948 | 7 |
观察:新版 Heartbeat 平均 Reward 最高,比旧版高 0.0154;严格通过数与 Plain、旧版相同,均为 7/46。平均进展增加,尚未转化为更多完全通过的任务。
+Insight:新版每次唤醒启动新的执行器,依靠工作区、Registry 与 Todo 恢复进度。简报记录了咨询任务中断后从 S27 接续的案例,也记录了 DuckDB 后续优化破坏正确性的退步。值得继续验证的是:外部化状态能否稳定支持恢复,以及检查点和回滚能否保住已有成果。
+范围:新旧同时改变了传输、会话生命周期、LoopX 版本和 replan 策略;有效运行包含替代 trial,部分替代运行时限更长。没有重复 seed,成本记录口径也不同,不能据此归因单个机制或宣称同预算效率优势。
+ +当前的信号是:有效续跑、可靠交付和反例驱动的修复值得继续研究。
-三项研究提供了局部正向观察,也暴露了成本、失败与归因问题。下一步需要在匹配预算和重复实验中,判断哪些收益能够稳定复现。
+让进度可恢复,让已有成果可保留,再检查续跑是否补上了验收缺口。
+四项研究提供了局部正向观察,也暴露了回退、成本和归因问题。下一步需要在匹配预算和重复实验中,分别验证恢复、回滚、交付与反例驱动修复的收益。
SWE-Marathon 与 DeepSWE × Sol 已撤回的 SSH Goal、Codex CLI 成绩和结论均不用于这里的比较。
技术内容仅据公开仓库材料整理。源码与数据引用固定到本次读取修订;历史实验使用其自身版本。本页没有新增实验结果。
GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell
-Long-Horizon Terminal-Bench spans nine categories, including software, games, science, and professional workflows. Final artifacts receive a continuous reward from 0 to 1; this study counts reward ≥ 0.95 as solved. Partial progress and full acceptance can therefore be examined separately.
+What LHTB tests: Long-Horizon Terminal-Bench places an agent in a stateful container and asks it to sustain work across hundreds of dependent terminal actions. Its official introduction uses nine categories: software and reverse engineering; scientific computing and simulation; earth, climate, and energy; multimodal and imaging analysis; research reproduction and ML; systems, performance, and security; interactive games; APEX professional workflows; and logic and constraint puzzles.
+Tasks include migrating an old framework, recovering data from scientific figures, reproducing paper experiments, handling investment-banking or legal matters, playing 2048 turn by turn, and searching for puzzle solutions. Nine categories and representative tasks →
+How it scores: hidden verifiers check final artifacts or replayable outcomes and award continuous reward from 0 to 1. This study counts reward ≥ 0.95 as solved, keeping partial progress distinct from full acceptance.
| Execution mechanism | Mean reward | Solved / 46 |
|---|---|---|
| Plain | 0.4218 | 7 |
| Native Goal | 0.4475 | 4 |
| 执行机制 | 平均 Reward | 通过 / 46 |
|---|---|---|
| Plain | 0.4218 | 7 |
| 原生 Goal | 0.4475 | 4 |
| {label} | )}|
|---|---|
| {category} | {examples} |
{c.taskExample}
+{c.benchmarkSetupNote}
+