From 68b480e38ee6b615d1ac182aa6f54087cbf7c723 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Fri, 18 Sep 2026 23:52:54 +0800 Subject: [PATCH 1/2] docs(blog): lead application benchmarks with LHTB Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- .../blog/application-scenarios/index.html | 26 ++++++++++++++++--- .../blog/zh/application-scenarios/index.html | 26 ++++++++++++++++--- examples/export-frontstage-share-bundle.mjs | 2 +- examples/frontstage-share-bundle-smoke.mjs | 1 + 4 files changed, 46 insertions(+), 9 deletions(-) diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index b7cd356482..22468d7d4f 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -83,8 +83,25 @@

01 / Finish a complex task

-

Benchmarks: three studies, three signals worth investigating

-

The studies below observe continuation, delivery, and validation behavior under different tasks, model settings, and measurement rules. Each needs to be read on its own terms. The current evidence does not establish a universal LoopX gain.

+

Benchmarks: start with LHTB, then three supporting signals

+

Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.

+ +
+

LHTB: recover progress and preserve working results

+

GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell

+

Long-Horizon Terminal-Bench spans nine categories, including software, games, science, and professional workflows. Final artifacts receive a continuous reward from 0 to 1; this study counts reward ≥ 0.95 as solved. Partial progress and full acceptance can therefore be examined separately.

+
+ + + + + +
LHTB five-arm results: mean reward and strict solves
Execution mechanismMean rewardSolved / 46
Plain0.42187
Native Goal0.44754
LoopX SSH-Goal0.46786
Legacy Heartbeat0.47947
New Heartbeat0.49487
+

Observation: New Heartbeat has the highest mean reward, 0.0154 above Legacy Heartbeat. Its strict solve count matches Plain and Legacy at 7/46. Higher average progress has not translated into more fully solved tasks.

+

Insight: each new Heartbeat wake starts a fresh executor, recovering progress from the workspace, registry, and Todos. The brief reports a consulting task resumed at S27 after interruption, alongside a DuckDB regression where later optimization broke correctness. The next questions are whether externalized state reliably enables recovery and whether checkpoints and rollback preserve working results.

+

Boundary: transport, session lifetime, LoopX version, and replanning change together. Effective runs include replacement trials, some with longer time limits. There are no repeated seeds, and cost telemetry differs, so this does not isolate a mechanism or establish an equal-budget efficiency advantage.

+ +

SWE-Marathon: continuation must close real gaps

@@ -113,8 +130,8 @@

DeepSWE × V4 Flash max: let counterexamples change the implementation

-

The current signal is that useful continuation, reliable delivery, and counterexample-driven repair deserve further study.

-

The studies offer local positive observations while exposing cost, failure, and attribution problems. Matched budgets and repeated experiments are needed to determine which gains reproduce reliably.

+

Recover progress, preserve working results, and check whether continuation closes acceptance gaps.

+

The four studies offer local positive observations while exposing regressions, cost, and attribution problems. Matched budgets and repeated experiments should examine recovery, rollback, delivery, and counterexample-driven repair separately.

Withdrawn SSH Goal and Codex CLI scores and conclusions from SWE-Marathon and DeepSWE × Sol are not used in these comparisons.

@@ -182,6 +199,7 @@

Evolution: add one verifiable capability at a time

Public sources and further reading

The technical content is derived only from public repository material. Source and data links are pinned to the revisions read for this article; historical experiments retain their own versions. This page introduces no new experimental results.

    +
  1. LHTB: five long-horizon execution mechanisms; public aggregates for 46 tasks; setup and evidence boundary; mechanisms and cases reported in the brief; official benchmark. See the study for contributor attribution.
  2. SWE-Marathon: continuous self-verification; setup, positive and negative cases, and limitations; public aggregate data. See the study for contributors and case provenance.
  3. DeepSWE × Sol: from continued execution to valid delivery (Chinese). The standalone brief covers historical results over 113 tasks, mechanism diagrams, and pinned primary sources; research archive contribution: @gwh6669999, #4502.
  4. DeepSWE × V4 Flash max: from hints to behavior; disclosure boundary; charts, cases, and metrics at the pinned revision.
  5. diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index 307bfa74fa..d8f1305f84 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -80,8 +80,25 @@

    01 / 把复杂任务做完

    -

    Benchmark:三项研究,三个值得追查的信号

    -

    三项研究分别观察续跑、交付和验证行为。任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。

    +

    Benchmark:从 LHTB 看长程工作,再看三个补充信号

    +

    先看跨领域长程任务中的恢复与回退,再用三项软件工程研究补充续跑、交付和验证行为。四项研究的任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。

    + +
    +

    LHTB:工作能恢复,已有成果也要保住

    +

    GPT-5.6 Sol / max · 46 个匹配任务、五种执行机制 · 每个任务 × 实验臂一条有效运行

    +

    Long-Horizon Terminal-Bench 覆盖软件、游戏、科学与专业工作流等九类任务,按最终产物给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。它把长程工作中的部分进展与完整验收分开呈现。

    +
    + + + + + +
    LHTB 五臂结果:平均 Reward 与严格通过数
    执行机制平均 Reward通过 / 46
    Plain0.42187
    原生 Goal0.44754
    LoopX SSH-Goal0.46786
    旧版 Heartbeat0.47947
    新版 Heartbeat0.49487
    +

    观察:新版 Heartbeat 平均 Reward 最高,比旧版高 0.0154;严格通过数与 Plain、旧版相同,均为 7/46。平均进展增加,尚未转化为更多完全通过的任务。

    +

    Insight:新版每次唤醒启动新的执行器,依靠工作区、Registry 与 Todo 恢复进度。简报记录了咨询任务中断后从 S27 接续的案例,也记录了 DuckDB 后续优化破坏正确性的退步。值得继续验证的是:外部化状态能否稳定支持恢复,以及检查点和回滚能否保住已有成果。

    +

    范围:新旧同时改变了传输、会话生命周期、LoopX 版本和 replan 策略;有效运行包含替代 trial,部分替代运行时限更长。没有重复 seed,成本记录口径也不同,不能据此归因单个机制或宣称同预算效率优势。

    + +

    SWE-Marathon:续跑要补上真实缺口

    @@ -110,8 +127,8 @@

    DeepSWE × V4 Flash max:让反例改变实现

    -

    当前的信号是:有效续跑、可靠交付和反例驱动的修复值得继续研究。

    -

    三项研究提供了局部正向观察,也暴露了成本、失败与归因问题。下一步需要在匹配预算和重复实验中,判断哪些收益能够稳定复现。

    +

    让进度可恢复,让已有成果可保留,再检查续跑是否补上了验收缺口。

    +

    四项研究提供了局部正向观察,也暴露了回退、成本和归因问题。下一步需要在匹配预算和重复实验中,分别验证恢复、回滚、交付与反例驱动修复的收益。

    SWE-Marathon 与 DeepSWE × Sol 已撤回的 SSH Goal、Codex CLI 成绩和结论均不用于这里的比较。

    @@ -179,6 +196,7 @@

    演进:每一步增加一种可验证的能力

    公开依据与延伸阅读

    技术内容仅据公开仓库材料整理。源码与数据引用固定到本次读取修订;历史实验使用其自身版本。本页没有新增实验结果。

      +
    1. LHTB:五种长程执行机制;46 题公开聚合数据;实验设置与范围;简报中的机制与正反案例;LHTB 官方说明。研究贡献者见原文。
    2. SWE-Marathon:持续自我验证;设置、正反个案与局限;公开聚合数据。贡献者与案例来源见研究原文。
    3. DeepSWE × Sol:从继续执行到有效交付。独立研究简报,含 113 任务历史结果、机制图与固定修订的一手来源;研究归档贡献:@gwh6669999,#4502。
    4. DeepSWE × V4 Flash max:从提示到行为;披露范围;固定修订的图表、案例与指标。
    5. diff --git a/examples/export-frontstage-share-bundle.mjs b/examples/export-frontstage-share-bundle.mjs index 359717abb7..e01d694da9 100644 --- a/examples/export-frontstage-share-bundle.mjs +++ b/examples/export-frontstage-share-bundle.mjs @@ -136,7 +136,7 @@ async function copyHomepage(siteDir, base) { async function copyPublicSiteRoutes(siteDir) { const homepage = resolve(siteDir, "index.html"); - for (const route of ["benchmarks/swe-marathon"]) { + for (const route of ["benchmarks/swe-marathon", "benchmarks/lhtb"]) { const routeDir = resolve(siteDir, route); await mkdir(routeDir, { recursive: true }); await copyFile(homepage, resolve(routeDir, "index.html")); diff --git a/examples/frontstage-share-bundle-smoke.mjs b/examples/frontstage-share-bundle-smoke.mjs index 2d97dfea62..5d3a0c642f 100644 --- a/examples/frontstage-share-bundle-smoke.mjs +++ b/examples/frontstage-share-bundle-smoke.mjs @@ -118,6 +118,7 @@ const siteDir = resolve(outDir, "site"); assertExists(resolve(siteDir, "index.html")); assertExists(resolve(siteDir, "frontstage/index.html")); assertExists(resolve(siteDir, "benchmarks/swe-marathon/index.html")); +assertExists(resolve(siteDir, "benchmarks/lhtb/index.html")); assertExists(resolve(siteDir, "benchmarks/deepswe/behavior-discovery/index.html")); // Static research articles must remain readable and navigable in the shipped // bundle without falling back to the homepage SPA. From 3fa0a79821a15f78ac84828a5e4533ffe9b994f9 Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Sat, 19 Sep 2026 00:00:19 +0800 Subject: [PATCH 2/2] docs(site): explain LHTB categories and artifact grading Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- .../blog/application-scenarios/index.html | 4 ++- .../blog/zh/application-scenarios/index.html | 4 ++- apps/presentation/site/src/LhtbBrief.tsx | 15 ++++++++ apps/presentation/site/src/lhtb-brief.css | 15 ++++++++ apps/presentation/site/src/lhtb-copy.json | 36 +++++++++++++++++++ 5 files changed, 72 insertions(+), 2 deletions(-) diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index 22468d7d4f..2e747f711a 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -89,7 +89,9 @@

      Benchmarks: start with LHTB, then three supporting signals

      LHTB: recover progress and preserve working results

      GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell

      -

      Long-Horizon Terminal-Bench spans nine categories, including software, games, science, and professional workflows. Final artifacts receive a continuous reward from 0 to 1; this study counts reward ≥ 0.95 as solved. Partial progress and full acceptance can therefore be examined separately.

      +

      What LHTB tests: Long-Horizon Terminal-Bench places an agent in a stateful container and asks it to sustain work across hundreds of dependent terminal actions. Its official introduction uses nine categories: software and reverse engineering; scientific computing and simulation; earth, climate, and energy; multimodal and imaging analysis; research reproduction and ML; systems, performance, and security; interactive games; APEX professional workflows; and logic and constraint puzzles.

      +

      Tasks include migrating an old framework, recovering data from scientific figures, reproducing paper experiments, handling investment-banking or legal matters, playing 2048 turn by turn, and searching for puzzle solutions. Nine categories and representative tasks →

      +

      How it scores: hidden verifiers check final artifacts or replayable outcomes and award continuous reward from 0 to 1. This study counts reward ≥ 0.95 as solved, keeping partial progress distinct from full acceptance.

      diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index d8f1305f84..c5ff1eeec9 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -86,7 +86,9 @@

      Benchmark:从 LHTB 看长程工作,再看三个补充信号

      LHTB:工作能恢复,已有成果也要保住

      GPT-5.6 Sol / max · 46 个匹配任务、五种执行机制 · 每个任务 × 实验臂一条有效运行

      -

      Long-Horizon Terminal-Bench 覆盖软件、游戏、科学与专业工作流等九类任务,按最终产物给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。它把长程工作中的部分进展与完整验收分开呈现。

      +

      LHTB 考什么:Long-Horizon Terminal-Bench 把 Agent 放进有状态的容器环境,要求它跨数百个相互依赖的终端动作完成工作。官方介绍页分为九类:软件与逆向工程、科学计算与仿真、地球/气候/能源、多模态与图像分析、研究复现与机器学习、系统/性能/安全、交互式游戏、APEX 专业工作流、逻辑与约束谜题。

      +

      具体会遇到旧版框架迁移、从科学图还原数据、复现论文实验、投行或法律事项,以及逐步玩 2048、搜索谜题解。九类任务与代表例子 →

      +

      怎样评分:隐藏验证器检查最终产物或可重放结果,给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。因此,部分进展与完整验收可以分开观察。

      LHTB five-arm results: mean reward and strict solves
      Execution mechanismMean rewardSolved / 46
      Plain0.42187
      Native Goal0.44754
      diff --git a/apps/presentation/site/src/LhtbBrief.tsx b/apps/presentation/site/src/LhtbBrief.tsx index 3c97b08f81..acbefb04bf 100644 --- a/apps/presentation/site/src/LhtbBrief.tsx +++ b/apps/presentation/site/src/LhtbBrief.tsx @@ -174,6 +174,21 @@ export function LhtbBrief() {
      {value}{label}{note}
      ))} +
      +

      {c.categoriesTitle}

      +

      {c.categoriesBody} {c.categoriesSource} · {c.categoriesExamplesSource}

      +
      +
      LHTB 五臂结果:平均 Reward 与严格通过数
      执行机制平均 Reward通过 / 46
      Plain0.42187
      原生 Goal0.44754
      + + {c.categoryColumns.map(label => )} + {c.categories.map(([category, examples]) => ( + + ))} +
      {c.categoriesTitle}
      {label}
      {category}{examples}
      +
      +

      {c.taskExample}

      +

      {c.benchmarkSetupNote}

      +
      diff --git a/apps/presentation/site/src/lhtb-brief.css b/apps/presentation/site/src/lhtb-brief.css index 9c418f066a..fbbac63318 100644 --- a/apps/presentation/site/src/lhtb-brief.css +++ b/apps/presentation/site/src/lhtb-brief.css @@ -113,6 +113,21 @@ border-left: 1px solid var(--bm-border); } +.lhtb-taxonomy { + grid-column: 1 / -1; + min-width: 0; + scroll-margin-top: 72px; +} + +.lhtb-taxonomy h3 { margin: 0 0 12px; font-size: 24px; } +.lhtb-taxonomy > p { color: var(--bm-body); font-size: 15px; line-height: 1.7; } +.lhtb-category-table { margin-block: 20px; } +.lhtb-category-table table { min-width: 0; table-layout: fixed; } +.lhtb-category-table caption { padding: 14px 18px; text-align: left; color: var(--bm-muted); font-size: 12px; } +.lhtb-category-table th:first-child { width: 35%; } +.lhtb-category-table tbody th { white-space: normal; } +.lhtb-category-table th, .lhtb-category-table td { overflow-wrap: anywhere; } + .lhtb-arm-flow article { display: grid; grid-template-rows: auto 1fr; diff --git a/apps/presentation/site/src/lhtb-copy.json b/apps/presentation/site/src/lhtb-copy.json index ecae53b7fd..c1ee5f5dfe 100644 --- a/apps/presentation/site/src/lhtb-copy.json +++ b/apps/presentation/site/src/lhtb-copy.json @@ -51,6 +51,24 @@ ["0–1", "continuous reward", "partial credit exposes progress below full solve"], ["≥ .95", "solved threshold", "the strict outcome used in this study"] ], + "categoriesTitle": "What work do the nine categories cover?", + "categoriesBody": "Categories follow the official introduction page; the examples show the tools and deliverables involved.", + "categoriesSource": "Official task taxonomy →", + "categoriesExamplesSource": "Task examples →", + "categoryColumns": ["Category", "Representative work"], + "categories": [ + ["Software & reverse engineering", "Migrate an old LangChain application; debug a RISC-V core."], + ["Scientific computing & simulation", "Accelerate an N-body simulator; restore an airfoil simulation to regression tolerances."], + ["Earth, climate & energy", "Audit groundwater simulations and power-system optimization."], + ["Multimodal & imaging analysis", "Reconstruct data from scientific figures; audit medical images."], + ["Research reproduction & ML", "Reproduce the UNISON simulation experiment or a Foldseek paper result."], + ["Systems, performance & security", "Work on DuckDB optimization or construct an exploit proof of concept."], + ["Interactive games", "Play 2048 or Super Mario through sustained interaction."], + ["APEX professional workflows", "Complete investment-banking deliverables or a legal matter."], + ["Logic & constraint puzzles", "Search for a checkmate or solve a constrained state-space puzzle."] + ], + "taskExample": "For example, UNISON paper reproduction requires a runnable simulation pipeline and reports. Hidden checks examine metric fidelity, partitioning and scheduling, repeatability, and behavior under held-out seeds. Partial credit records which requirements the artifacts satisfy.", + "benchmarkSetupNote": "The official model comparison used Terminus-2 and a 90-minute budget. This page reports a separate LoopX five-arm study with its disclosed runtimes and replacement trials. Upstream tasks and verifiers continue to evolve; historical scores retain their original evaluation settings.", "mechanismEyebrow": "02 / Five execution mechanisms", "mechanismTitle": "The model is shared. Continuation ownership and durable state are not.", "mechanismBody": "All five arms use GPT-5.6 Sol at max reasoning with web search disabled. The experiment varies who decides that work should continue, how the next unit of work is represented, and whether a new executor can recover the frontier without old conversation instructions.", @@ -178,6 +196,24 @@ ["0–1", "连续 Reward", "Partial credit 能显示未完全通过时的进度"], ["≥ .95", "通过阈值", "本研究采用的严格完成标准"] ], + "categoriesTitle": "九类任务,具体在做什么?", + "categoriesBody": "类别按官方介绍页划分;代表工作展示每类涉及的工具与交付物。", + "categoriesSource": "官方任务分类 →", + "categoriesExamplesSource": "任务示例 →", + "categoryColumns": ["类别", "代表工作"], + "categories": [ + ["软件与逆向工程", "迁移旧版 LangChain 应用;调试 RISC-V 处理器。"], + ["科学计算与仿真", "加速 N-body 模拟器;让翼型仿真重新满足回归容差。"], + ["地球、气候与能源", "地下水仿真审计、电力系统优化。"], + ["多模态与图像分析", "从科学图还原数据;医学影像审计。"], + ["研究复现与机器学习", "复现 UNISON 仿真实验或 Foldseek 论文结果。"], + ["系统、性能与安全", "DuckDB 优化、漏洞利用概念验证。"], + ["交互式游戏", "逐步玩 2048 或 Super Mario,持续根据环境反馈行动。"], + ["APEX 专业工作流", "完成投资银行交付物或法律事项。"], + ["逻辑与约束谜题", "搜索将死解,或求解受约束的状态空间谜题。"] + ], + "taskExample": "例如 UNISON 论文复现,交付物是能运行的仿真流水线和报告。隐藏检查会核对指标、分区与调度、重复运行的一致性,以及留出 seed 下的行为;部分得分反映产物已经满足了哪些要求。", + "benchmarkSetupNote": "官方模型比较使用 Terminus-2 和每题 90 分钟预算;本页是另行开展的 LoopX 五臂研究,执行设置与替代运行以本页披露为准。上游任务与验证器仍在演进,历史成绩保留其原始评测设置。", "mechanismEyebrow": "02 / 五种执行机制", "mechanismTitle": "模型相同;续跑权和持久状态的归属不同。", "mechanismBody": "五臂均使用 GPT-5.6 Sol、max 推理,并关闭 web search。实验改变的是谁决定继续工作、下一段工作如何表示,以及新执行器能否不依赖旧对话指令恢复任务前沿。",