diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index 2e747f711a..9e4040a26d 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -87,22 +87,19 @@

Benchmarks: start with LHTB, then three supporting signals

Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.

-

LHTB: recover progress and preserve working results

-

GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell

-

What LHTB tests: Long-Horizon Terminal-Bench places an agent in a stateful container and asks it to sustain work across hundreds of dependent terminal actions. Its official introduction uses nine categories: software and reverse engineering; scientific computing and simulation; earth, climate, and energy; multimodal and imaging analysis; research reproduction and ML; systems, performance, and security; interactive games; APEX professional workflows; and logic and constraint puzzles.

-

Tasks include migrating an old framework, recovering data from scientific figures, reproducing paper experiments, handling investment-banking or legal matters, playing 2048 turn by turn, and searching for puzzle solutions. Nine categories and representative tasks →

-

How it scores: hidden verifiers check final artifacts or replayable outcomes and award continuous reward from 0 to 1. This study counts reward ≥ 0.95 as solved, keeping partial progress distinct from full acceptance.

-
+

LHTB: what does LoopX accomplish beyond Plain and native Goal?

+

GPT-5.6 Sol / max · 46 matched tasks · one effective run per task-arm cell

+

What LHTB tests: Long-Horizon Terminal-Bench asks an agent to complete hundreds of dependent terminal actions in a stateful container. Its 46 tasks cover nine categories: software and reverse engineering; scientific computing; earth, climate and energy; multimodal work; research reproduction; systems, performance and security; games; APEX professional workflows; and logic puzzles. Examples include framework migration, paper reproduction, investment-banking deliverables and playing 2048. Categories and representative tasks →

+

How it scores: hidden verifiers grade final artifacts or replayable outcomes on a 0–1 reward scale. This study counts ≥ 0.95 as solved, reporting average progress and full acceptance separately.

+
LHTB five-arm results: mean reward and strict solves
Execution mechanismMean rewardSolved / 46
- - - +
LHTB: LoopX Heartbeat and two baselines
ExecutionMean rewardSolved / 46
Plain0.42187
Native Goal0.44754
LoopX SSH-Goal0.46786
Legacy Heartbeat0.47947
New Heartbeat0.49487
LoopX Heartbeat (1.0.3)0.49487
-

Observation: New Heartbeat has the highest mean reward, 0.0154 above Legacy Heartbeat. Its strict solve count matches Plain and Legacy at 7/46. Higher average progress has not translated into more fully solved tasks.

-

Insight: each new Heartbeat wake starts a fresh executor, recovering progress from the workspace, registry, and Todos. The brief reports a consulting task resumed at S27 after interruption, alongside a DuckDB regression where later optimization broke correctness. The next questions are whether externalized state reliably enables recovery and whether checkpoints and rollback preserve working results.

-

Boundary: transport, session lifetime, LoopX version, and replanning change together. Effective runs include replacement trials, some with longer time limits. There are no repeated seeds, and cost telemetry differs, so this does not isolate a mechanism or establish an equal-budget efficiency advantage.

- +

Observation: against Plain, LoopX gains 0.0731 mean reward (17.3%), with 17 wins, 13 ties and 16 losses, and the same solve count. Against native Goal, it gains 0.0473 (10.6%), with 23 wins, 13 ties and 10 losses, and 7 solves versus 4. Counts compare published unrounded scores strictly; higher mean reward does not imply better results on every task.

+

Insight: Plain advances within a session, native Goal retains the objective, and LoopX also organizes subsequent work through registry and Todos. Tabular and Sokoban improve over both baselines; the PoC task ties native Goal; DuckDB scores below it. The questions are whether durable work state reliably adds useful progress, and whether acceptance and rollback preserve that progress as a valid result.

+

Boundary: LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $87.58. Accounting also differs; equal-budget efficiency gains are not established.

+
@@ -201,7 +198,7 @@

Evolution: add one verifiable capability at a time

Public sources and further reading

The technical content is derived only from public repository material. Source and data links are pinned to the revisions read for this article; historical experiments retain their own versions. This page introduces no new experimental results.

    -
  1. LHTB: five long-horizon execution mechanisms; public aggregates for 46 tasks; setup and evidence boundary; mechanisms and cases reported in the brief; official benchmark. See the study for contributor attribution.
  2. +
  3. LHTB: compared with Plain and native Goal; public aggregates for 46 tasks; setup and evidence boundary; mechanisms and cases reported in the brief; official benchmark. See the study for contributor attribution.
  4. SWE-Marathon: continuous self-verification; setup, positive and negative cases, and limitations; public aggregate data. See the study for contributors and case provenance.
  5. DeepSWE × Sol: from continued execution to valid delivery (Chinese). The standalone brief covers historical results over 113 tasks, mechanism diagrams, and pinned primary sources; research archive contribution: @gwh6669999, #4502.
  6. DeepSWE × V4 Flash max: from hints to behavior; disclosure boundary; charts, cases, and metrics at the pinned revision.
  7. diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index c5ff1eeec9..6bf17e6dd8 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -84,22 +84,19 @@

    Benchmark:从 LHTB 看长程工作,再看三个补充信号

    先看跨领域长程任务中的恢复与回退,再用三项软件工程研究补充续跑、交付和验证行为。四项研究的任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。

    -

    LHTB:工作能恢复,已有成果也要保住

    -

    GPT-5.6 Sol / max · 46 个匹配任务、五种执行机制 · 每个任务 × 实验臂一条有效运行

    -

    LHTB 考什么:Long-Horizon Terminal-Bench 把 Agent 放进有状态的容器环境,要求它跨数百个相互依赖的终端动作完成工作。官方介绍页分为九类:软件与逆向工程、科学计算与仿真、地球/气候/能源、多模态与图像分析、研究复现与机器学习、系统/性能/安全、交互式游戏、APEX 专业工作流、逻辑与约束谜题。

    -

    具体会遇到旧版框架迁移、从科学图还原数据、复现论文实验、投行或法律事项,以及逐步玩 2048、搜索谜题解。九类任务与代表例子 →

    -

    怎样评分:隐藏验证器检查最终产物或可重放结果,给出 0–1 的连续 Reward;本研究以 Reward ≥ 0.95 计为通过。因此,部分进展与完整验收可以分开观察。

    -
    +

    LHTB:比 Plain 和原生 Goal 多做成了什么?

    +

    GPT-5.6 Sol / max · 46 个匹配任务 · 每题每臂一条有效运行

    +

    LHTB 考什么:Long-Horizon Terminal-Bench 要求 Agent 在有状态的容器中,跨数百个有依赖的终端动作完成工作。46 题覆盖九类:软件与逆向工程、科学计算、地球/气候/能源、多模态、研究复现、系统/性能/安全、游戏、APEX 专业工作流、逻辑谜题。比如迁移旧框架、复现论文、完成投行交付或逐步玩 2048。类别与代表任务 →

    +

    怎样评分:隐藏验证器检查最终产物或可重放结果,给出 0–1 的 Reward;本研究以 ≥ 0.95 计为通过。平均进展与完整验收分别报告。

    +
    LHTB 五臂结果:平均 Reward 与严格通过数
    执行机制平均 Reward通过 / 46
    - - - +
    LHTB:LoopX Heartbeat 与两种基线
    执行方式平均 Reward通过 / 46
    Plain0.42187
    原生 Goal0.44754
    LoopX SSH-Goal0.46786
    旧版 Heartbeat0.47947
    新版 Heartbeat0.49487
    LoopX Heartbeat(1.0.3)0.49487
    -

    观察:新版 Heartbeat 平均 Reward 最高,比旧版高 0.0154;严格通过数与 Plain、旧版相同,均为 7/46。平均进展增加,尚未转化为更多完全通过的任务。

    -

    Insight:新版每次唤醒启动新的执行器,依靠工作区、Registry 与 Todo 恢复进度。简报记录了咨询任务中断后从 S27 接续的案例,也记录了 DuckDB 后续优化破坏正确性的退步。值得继续验证的是:外部化状态能否稳定支持恢复,以及检查点和回滚能否保住已有成果。

    -

    范围:新旧同时改变了传输、会话生命周期、LoopX 版本和 replan 策略;有效运行包含替代 trial,部分替代运行时限更长。没有重复 seed,成本记录口径也不同,不能据此归因单个机制或宣称同预算效率优势。

    - +

    观察:相对 Plain,LoopX 均分增加 0.0731(17.3%),逐题为 17 胜、13 平、16 负,通过数相同;相对原生 Goal,均分增加 0.0473(10.6%),为 23 胜、13 平、10 负,通过数从 4 到 7。胜负按公开原始分数严格比较;均分提升不代表每题都更强。

    +

    Insight:Plain 在会话内推进,原生 Goal 保留目标,LoopX 进一步用 Registry 与 Todo 组织后续工作。Tabular 与 Sokoban 相对两种基线均有收益;PoC 与原生 Goal 持平;DuckDB 却低于原生 Goal。值得验证的是,持久工作状态能否稳定增加有效进展,以及验收和回滚能否把进展变成保得住的结果。

    +

    范围:这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $87.58;计费口径也不一致,尚不能宣称同预算效率优势。

    +
    @@ -198,7 +195,7 @@

    演进:每一步增加一种可验证的能力

    公开依据与延伸阅读

    技术内容仅据公开仓库材料整理。源码与数据引用固定到本次读取修订;历史实验使用其自身版本。本页没有新增实验结果。

      -
    1. LHTB:五种长程执行机制;46 题公开聚合数据;实验设置与范围;简报中的机制与正反案例;LHTB 官方说明。研究贡献者见原文。
    2. +
    3. LHTB:与 Plain、原生 Goal 的对比;46 题公开聚合数据;实验设置与范围;简报中的机制与正反案例;LHTB 官方说明。研究贡献者见原文。
    4. SWE-Marathon:持续自我验证;设置、正反个案与局限;公开聚合数据。贡献者与案例来源见研究原文。
    5. DeepSWE × Sol:从继续执行到有效交付。独立研究简报,含 113 任务历史结果、机制图与固定修订的一手来源;研究归档贡献:@gwh6669999,#4502。
    6. DeepSWE × V4 Flash max:从提示到行为;披露范围;固定修订的图表、案例与指标。
    7. diff --git a/apps/presentation/site/src/LhtbBrief.tsx b/apps/presentation/site/src/LhtbBrief.tsx index acbefb04bf..a2c7210df1 100644 --- a/apps/presentation/site/src/LhtbBrief.tsx +++ b/apps/presentation/site/src/LhtbBrief.tsx @@ -13,18 +13,24 @@ import { usePublicPageNavigation } from "./usePublicPageNavigation"; import study from "../../../../benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json"; import copy from "./lhtb-copy.json"; -type Language = "en" | "zh"; type ArmKey = keyof typeof study.arms; -type TaskRow = (typeof study.tasks)[number]; -type TableMode = "all" | "spread" | "heartbeat"; +type Baseline = "plain" | "native_goal"; +type TableMode = "all" | Baseline; -const armOrder: ArmKey[] = [ - "plain", - "native_goal", - "ssh_goal", - "legacy_heartbeat", - "new_heartbeat", -]; +const primaryArms: ArmKey[] = ["plain", "native_goal", "new_heartbeat"]; +const historicalArms: ArmKey[] = ["ssh_goal", "legacy_heartbeat"]; +const baselines: Baseline[] = ["plain", "native_goal"]; +const comparisons = baselines.map((baseline) => { + const deltas = study.tasks.map((row) => row.new_heartbeat - row[baseline]); + const meanDelta = deltas.reduce((sum, delta) => sum + delta, 0) / deltas.length; + const baselineMean = study.tasks.reduce((sum, row) => sum + row[baseline], 0) / deltas.length; + return { + baseline, meanDelta, relativeGain: meanDelta / baselineMean, + wins: deltas.filter((delta) => delta > 0).length, + ties: deltas.filter((delta) => delta === 0).length, + losses: deltas.filter((delta) => delta < 0).length, + }; +}); const contributorLinks = [ { label: "@shangzh0", href: "https://github.com/shangzh0" }, @@ -40,39 +46,68 @@ function formatMillions(value: number) { } function formatReward(value: number) { - return value.toFixed(3); -} - -function rewardSpread(row: TaskRow) { - const values = armOrder.map((arm) => row[arm]); - return Math.max(...values) - Math.min(...values); + return value.toFixed(4); } export function LhtbBrief() { const [language, setLanguage] = usePublicPageNavigation(); const [query, setQuery] = useState(""); const [tableMode, setTableMode] = useState("all"); + const [showHistory, setShowHistory] = useState(false); const c = copy[language]; + const visibleArms = showHistory ? [...primaryArms, ...historicalArms] : primaryArms; const basePath = import.meta.env.BASE_URL; useEffect(() => { document.title = language === "zh" - ? "LoopX × LHTB:五种长程执行机制" - : "LoopX × LHTB: five long-horizon execution mechanisms"; + ? "LoopX × LHTB:与 Plain、原生 Goal 的对比" + : "LoopX × LHTB: compared with Plain and native Goal"; }, [language]); const visibleTasks = useMemo(() => { const normalized = query.trim().toLowerCase(); return study.tasks.filter((row) => { if (normalized && !row.task.toLowerCase().includes(normalized)) return false; - if (tableMode === "spread") return rewardSpread(row) >= 0.2; - if (tableMode === "heartbeat") { - return Math.abs(row.new_heartbeat - row.legacy_heartbeat) >= 0.05; - } + if (tableMode !== "all") return Math.abs(row.new_heartbeat - row[tableMode]) >= 0.05; return true; }); }, [query, tableMode]); + const summaryTable = (arms: ArmKey[]) => ( +
      + + {c.summaryColumns.map((label) => )} + {arms.map((arm) => { + const row = study.arms[arm]; + const tokens = "tokens" in row + ? formatMillions(row.tokens) + : `${formatMillions(row.input_tokens)} in / ${formatMillions(row.output_tokens)} out`; + const cost = "estimated_cost_usd" in row ? row.estimated_cost_usd : row.recorded_cost_usd; + return ( + + + + + + ); + })} +
      {label}
      {c.armLabels[arm]}{c.armKinds[arm]}{row.mean_reward.toFixed(4)}{row.pass_095}/{study.tasks.length}{tokens}${cost.toFixed(2)}
      +
      + ); + + const caseCards = (cases: string[][]) => cases.map(([task, note]) => { + const row = study.tasks.find((item) => item.task === task)!; + return ( +
      +

      {task}

      +
      {primaryArms.map((arm) => ( +
      {c.armLabels[arm]}
      {formatReward(row[arm])}
      + ))}
      +

      {note}

      +
      + ); + }); + return (
      @@ -123,37 +158,21 @@ export function LhtbBrief() {

      {c.resultTitle}

      {c.resultBody}

      -
      - - {c.summaryColumns.map((label) => )} - - {armOrder.map((arm) => { - const row = study.arms[arm]; - const isNew = arm === "new_heartbeat"; - const tokens = "tokens" in row - ? formatMillions(row.tokens) - : `${formatMillions(row.input_tokens)} in / ${formatMillions(row.output_tokens)} out`; - const cost = "estimated_cost_usd" in row ? row.estimated_cost_usd : row.recorded_cost_usd; - return ( - - - - - - - - ); - })} - -
      {label}
      {c.armLabels[arm]}{c.armKinds[arm]}{row.mean_reward.toFixed(4)}{row.pass_095}/46{tokens}${cost.toFixed(2)}
      -
      -
      -
      +{study.heartbeat_comparison.mean_delta.toFixed(4)}{c.deltaMean}
      -
      {study.heartbeat_comparison.wins}{c.deltaWins}
      -
      {study.heartbeat_comparison.ties}{c.deltaTies}
      -
      {study.heartbeat_comparison.losses}{c.deltaLosses}
      + {summaryTable(primaryArms)} +
      + {comparisons.map((row) => ( +
      +

      {c.comparisonTitle.replace("{baseline}", c.armLabels[row.baseline])}

      + +{row.meanDelta.toFixed(4)}{c.deltaMean} +

      +{(row.relativeGain * 100).toFixed(1)}% {c.relativeGain}

      +

      {c.pairCounts.replace("{wins}", String(row.wins)).replace("{ties}", String(row.ties)).replace("{losses}", String(row.losses))}

      +

      {c.pairSolves.replace("{current}", `${study.arms.new_heartbeat.pass_095}/46`).replace("{baseline}", c.armLabels[row.baseline]).replace("{base}", `${study.arms[row.baseline].pass_095}/46`)}

      +
      + ))}
      +

      {c.comparisonNote}

      {c.readingNoteLabel}{c.readingNote}

      +
      {c.historyTitle}

      {c.historyBody}

      {summaryTable(historicalArms)}
      @@ -221,15 +240,11 @@ export function LhtbBrief() {

      {c.gainTitle}

      - {c.gainCases.map(([task, before, after, note]) => ( -

      {task}

      {before} → {after}

      {note}

      - ))} + {caseCards(c.gainCases)}

      {c.lossTitle}

      - {c.lossCases.map(([task, before, after, note]) => ( -

      {task}

      {before} → {after}

      {note}

      - ))} + {caseCards(c.lossCases)}
      @@ -243,23 +258,24 @@ export function LhtbBrief() {
      - {(["all", "spread", "heartbeat"] as TableMode[]).map((mode) => ( - ))}
      +
      - {armOrder.map((arm) => )} + {visibleArms.map((arm) => )} {visibleTasks.map((row) => { - const best = Math.max(...armOrder.map((arm) => row[arm])); + const best = Math.max(...visibleArms.map((arm) => row[arm])); return ( - {armOrder.map((arm) => ( + {visibleArms.map((arm) => ( ))} diff --git a/apps/presentation/site/src/lhtb-brief.css b/apps/presentation/site/src/lhtb-brief.css index fbbac63318..f3bf8e90d4 100644 --- a/apps/presentation/site/src/lhtb-brief.css +++ b/apps/presentation/site/src/lhtb-brief.css @@ -41,35 +41,43 @@ box-shadow: inset 3px 0 0 var(--bm-blue); } -.lhtb-delta-strip { +.lhtb-comparisons { grid-column: 1 / -1; display: grid; - grid-template-columns: repeat(4, 1fr); + grid-template-columns: repeat(2, minmax(0, 1fr)); border-top: 1px solid var(--bm-border); border-left: 1px solid var(--bm-border); } -.lhtb-delta-strip > div { - display: flex; - flex-direction: column; - min-height: 94px; - padding: 16px 18px; +.lhtb-comparisons article { + padding: 24px; border-right: 1px solid var(--bm-border); border-bottom: 1px solid var(--bm-border); background: var(--bm-surface); } -.lhtb-delta-strip strong { +.lhtb-comparisons h3 { margin: 0 0 12px; font-size: 18px; } +.lhtb-comparisons strong { color: var(--bm-blue); font-family: "Geist Mono", "SFMono-Regular", Consolas, monospace; - font-size: 26px; - font-weight: 560; -} - -.lhtb-delta-strip span { - margin-top: auto; - color: var(--bm-muted); - font-size: 11px; + font-size: 32px; + font-weight: 600; +} +.lhtb-comparisons span { margin-left: 12px; color: var(--bm-muted); font-size: 12px; } +.lhtb-comparisons p { margin: 8px 0 0; color: var(--bm-body); font-size: 14px; } +.lhtb-history { grid-column: 1 / -1; min-width: 0; } +.lhtb-history summary { cursor: pointer; padding-block: 16px; font-weight: 600; } +.lhtb-history > p { color: var(--bm-body); font-size: 14px; line-height: 1.7; } +.lhtb-history-toggle { + justify-self: start; + padding: 12px 16px; + border: 1px solid var(--bm-border); + border-radius: 6px; + background: var(--bm-surface); + color: var(--bm-body); + font: inherit; + font-size: 12px; + cursor: pointer; } .lhtb-official-links { @@ -108,7 +116,7 @@ .lhtb-arm-flow { grid-column: 1 / -1; display: grid; - grid-template-columns: repeat(5, 1fr); + grid-template-columns: repeat(3, minmax(0, 1fr)); border-top: 1px solid var(--bm-border); border-left: 1px solid var(--bm-border); } @@ -219,15 +227,11 @@ letter-spacing: 0; } -.lhtb-cases strong { - color: var(--bm-blue); - font-family: "Geist Mono", "SFMono-Regular", Consolas, monospace; - font-size: 16px; -} - -.lhtb-cases > div:last-child strong { - color: var(--bm-amber); -} +.lhtb-case-scores { margin: 0; } +.lhtb-case-scores > div { display: flex; justify-content: space-between; gap: 12px; padding-block: 3px; } +.lhtb-case-scores dt { color: var(--bm-muted); font-size: 11px; } +.lhtb-case-scores dd { margin: 0; font-family: "Geist Mono", monospace; font-size: 13px; } +.lhtb-case-scores > div:last-child { color: var(--bm-blue); font-weight: 600; } .lhtb-cases p { margin: 10px 0 0; @@ -425,8 +429,8 @@ font-size: 40px; } - .lhtb-delta-strip { - grid-template-columns: 1fr 1fr; + .lhtb-comparisons { + grid-template-columns: 1fr; } .lhtb-arm-flow { diff --git a/apps/presentation/site/src/lhtb-copy.json b/apps/presentation/site/src/lhtb-copy.json index c1ee5f5dfe..3009fa9fcb 100644 --- a/apps/presentation/site/src/lhtb-copy.json +++ b/apps/presentation/site/src/lhtb-copy.json @@ -1,46 +1,43 @@ { "en": { - "meta": "LHTB / FIVE-ARM EXPLORATORY STUDY", + "meta": "LHTB / LOOPX AND TWO BASELINES", "back": "LoopX home", "source": "Study source", - "title": "Five ways to keep an agent working across a long horizon.", - "deck": "A 46-task comparison of plain Codex, native Goal, LoopX SSH-Goal, legacy Heartbeat, and a new fresh-exec Heartbeat. The result is a mechanism study: how durable state, bounded Todos, and replanning change continuation behavior.", + "title": "What does LoopX accomplish beyond Plain and native Goal?", + "deck": "Across 46 tasks with the same model, LoopX Heartbeat has a higher mean reward than Plain and native Goal. It matches Plain on strict solves and exceeds native Goal. Task-level gains, losses, and execution settings put this signal in context.", "evidenceTag": "GPT-5.6 Sol · max reasoning · 46 tasks · one effective trial per cell", "contributorsLabel": "Contributors", "readingLabel": "Reading path", "readingPath": [ - ["01", "Result", "#result"], - ["02", "Mechanisms", "#mechanisms"], - ["03", "Insights", "#insights"], - ["04", "46 tasks", "#scores"] + ["01", "Comparison", "#result"], + ["02", "Execution", "#mechanisms"], + ["03", "Gains and losses", "#insights"], + ["04", "Task evidence", "#scores"] ], "resultEyebrow": "00 / Core result", - "resultTitle": "The newest Heartbeat has the highest mean reward, with gains concentrated in recoverable, externally visible work.", - "resultBody": "Mean reward rises from 0.4218 for Plain to 0.4948 for New Heartbeat. Against Legacy Heartbeat, the new runtime wins 17 tasks, ties 19, and loses 10. The +0.0154 mean delta is useful directional evidence, not a causal estimate.", - "summaryColumns": ["Arm", "Mean reward", "Solved ≥ 0.95", "Recorded tokens", "USD"], + "resultTitle": "Higher mean reward than both baselines; no more strict solves than Plain.", + "resultBody": "LoopX averages 0.4948, Plain 0.4218, and native Goal 0.4475. At reward ≥ 0.95, they solve 7, 7, and 4 tasks respectively. Mean reward captures partial progress; full acceptance remains a separate outcome.", + "summaryColumns": ["Execution", "Mean reward", "Solved ≥ 0.95", "Token usage", "USD"], "armLabels": { "plain": "Plain", "native_goal": "Native Goal", "ssh_goal": "LoopX SSH-Goal", "legacy_heartbeat": "Legacy Heartbeat", - "new_heartbeat": "New Heartbeat" + "new_heartbeat": "LoopX Heartbeat" }, "armKinds": { "plain": "baseline · app-server", "native_goal": "baseline · app-server", - "ssh_goal": "LoopX · app-server", + "ssh_goal": "historical · app-server", "legacy_heartbeat": "LoopX 0.5.3 · app-server", "new_heartbeat": "LoopX 1.0.3 · generic_cli" }, - "deltaMean": "mean reward vs legacy Heartbeat", - "deltaWins": "task wins", - "deltaTies": "ties", - "deltaLosses": "losses", + "deltaMean": "mean reward increase", "readingNoteLabel": "How to read this", - "readingNote": "The first four arms are historical Codex app-server runs. New Heartbeat uses generic_cli with a fresh codex exec per wake, so the comparison changes several runtime dimensions together. Historical token costs are estimates; New Heartbeat cost is recorded telemetry.", + "readingNote": "LoopX here means the study’s 1.0.3 Heartbeat with fresh exec, not a new evaluation of the latest release. Both baselines use app-server; LoopX uses generic_cli. Effective runs include replacement trials, some with longer limits. Historical USD is estimated; LoopX USD is recorded. This is not an equal-budget efficiency ranking.", "benchmarkEyebrow": "01 / The benchmark", "benchmarkTitle": "LHTB grades durable artifacts after hundreds of dependent terminal actions.", - "benchmarkBody": "Long-Horizon Terminal-Bench contains 46 containerized tasks across nine categories. Hidden rebuild-from-artifact verifiers grade the final workspace on a continuous [0,1] scale; self-reported completion does not count. Reward ≥ 0.95 is the solved threshold.", + "benchmarkBody": "Long-Horizon Terminal-Bench contains 46 containerized tasks across nine categories. Hidden verifiers grade artifacts or replayable outcomes on a continuous [0,1] scale; self-reported completion does not count. This study uses reward ≥ 0.95 as its solved threshold.", "officialLinks": [ ["Official benchmark", "Task taxonomy, leaderboard, methodology, and harness notes.", "https://zli12321.github.io/LHTB/index.html"], ["LHTB paper", "Long-Horizon Terminal-Bench, arXiv:2607.08964.", "https://arxiv.org/abs/2607.08964"] @@ -69,123 +66,133 @@ ], "taskExample": "For example, UNISON paper reproduction requires a runnable simulation pipeline and reports. Hidden checks examine metric fidelity, partitioning and scheduling, repeatability, and behavior under held-out seeds. Partial credit records which requirements the artifacts satisfy.", "benchmarkSetupNote": "The official model comparison used Terminus-2 and a 90-minute budget. This page reports a separate LoopX five-arm study with its disclosed runtimes and replacement trials. Upstream tasks and verifiers continue to evolve; historical scores retain their original evaluation settings.", - "mechanismEyebrow": "02 / Five execution mechanisms", - "mechanismTitle": "The model is shared. Continuation ownership and durable state are not.", - "mechanismBody": "All five arms use GPT-5.6 Sol at max reasoning with web search disabled. The experiment varies who decides that work should continue, how the next unit of work is represented, and whether a new executor can recover the frontier without old conversation instructions.", + "mechanismEyebrow": "02 / Three execution settings", + "mechanismTitle": "How does the goal persist, and who organizes the next step?", + "mechanismBody": "All three can use the same model and tools for long tasks. Plain advances within its session; native Goal persists the objective; LoopX also maintains work items and progress state to govern subsequent wakes. These are configuration differences, not independently established causes of the score gains.", "mechanismRows": [ - ["plain", "Plain", "Model session owns stopping", "A normal app-server turn plans, edits, and returns. There is no persistent Goal or LoopX registry."], - ["native_goal", "Native Goal", "Codex Goal owns continuity", "A persistent “Finish the task” Goal survives turns, without LoopX Todo, quota, or registry control."], - ["ssh_goal", "LoopX SSH-Goal", "LoopX dispatcher owns the next slice", "Registry, quota, claimed Todo, bounded work, and writeback make the next action explicit across phases."], - ["legacy_heartbeat", "Legacy Heartbeat", "Periodic LoopX obligation in app-server", "LoopX 0.5.3 periodically re-enters a long-lived app-server setup and asks it to advance one bounded slice."], - ["new_heartbeat", "New Heartbeat", "Durable control, replaceable executor", "LoopX 1.0.3 starts a fresh codex exec on every wake. Workspace and registry persist; old conversation instructions do not. Three completed advancement Todos open a replan review."] + ["plain", "Plain", "Planning and execution within a session", "A normal app-server turn plans, uses tools, and returns. It can run for a long time, but has no native persistent Goal or LoopX control plane."], + ["native_goal", "Native Goal", "An objective that persists across turns", "Codex maintains a “Finish the task” Goal for continued work. This arm does not use LoopX registry, Todos, or scheduling."], + ["new_heartbeat", "LoopX Heartbeat", "Work state persists across executors", "Workspace, registry, and Todos persist across fresh codex exec wakes. Three qualifying advancement Todos trigger a review that can retain or revise the plan."] ], "insightEyebrow": "03 / LoopX insight", - "insightTitle": "The advantage is not a longer chat. It is a recoverable task frontier.", - "insightBody": "The strongest evidence appears when progress can be externalized into files and Todos: the current stage, passed checks, remaining modules, and next falsifiable action. A fresh executor can then continue from project evidence instead of reconstructing intent from prose history.", + "insightTitle": "Average progress increased. Task-level evidence explains where.", + "insightBody": "Against Plain, LoopX scores higher on 17 tasks and lower on 16: larger gains outweigh losses, rather than most tasks improving. Against native Goal, it scores higher on 23 and lower on 10. Each case below shows all three scores so that the choice of baseline stays visible.", "insights": [ - ["State survives the session", "Workspace artifacts plus LoopX registry and Todo state preserve where work stopped. Executor sessions become replaceable rather than the sole memory owner."], - ["Todos make continuation inspectable", "A claimed bounded slice separates current work, completed checks, and successor work. This lowers restart ambiguity on staged repairs and multi-artifact workflows."], - ["Replan can reopen false endings", "After three qualifying Todo completions, New Heartbeat opens a review obligation. Negative evidence can redirect the frontier instead of accepting a stale “done.”"], - ["Fresh execution limits instruction drift", "Each wake gets the current task body and project state, not an accumulated conversation. This creates a cleaner boundary between durable control state and model-session state."] + ["More partial progress; full acceptance remains hard", "LoopX’s 0.4948 exceeds both baselines, yet it matches Plain at 7/46 solves. Large gains on Tabular and Sokoban still fall below 0.95. Advancing further and finishing remain distinct."], + ["Native Goal is a substantive control", "On the PoC task, Goal and LoopX both reach 0.892 while Plain scores zero. On DuckDB, Goal reaches 0.770 and LoopX zero. A LoopX-specific claim must account for what native continuation already achieves."], + ["Recovery is a mechanism clue, not attribution", "The original brief reports a consulting task resuming at S27 after interruption. Workspace and Todos provide a place to recover progress; matched-budget ablations are needed to separate recovery, fresh execution, and replanning."], + ["Continuation must preserve results and justify cost", "The reported DuckDB case suggests later optimization can break correctness. Checkpointing, acceptance, and rollback merit separate tests. LoopX’s recorded cost exceeds both baseline estimates; an efficiency gain is not established."] ], - "gainTitle": "Where the new Heartbeat gained", + "gainTitle": "Gain cases: keep both baselines in view", "gainCases": [ - ["tabular-data-feature-covshift", "0.292", "0.743", "A compact artifact-and-check workflow gave successor work a concrete target; the largest observed gain."], - ["rush-hour-campaign", "0.481", "0.772", "Board state and unfinished stages were recoverable, so subsequent work continued the search instead of restarting it."], - ["apex-management-consulting-matter", "0.433", "0.513", "After interruption, a fresh executor recovered at stage S27 and completed the final four stages."], - ["unknown-config-semantics", "0.000", "0.231", "Externalized artifacts gave a new wake enough evidence to make partial progress where the legacy run produced none."] + ["tabular-data-feature-covshift", "A substantial gain over both baselines, but 0.743 still falls short of strict acceptance."], + ["sokoban", "Higher than both baselines; the final 0.890 still does not count as a strict solve."], + ["poc-exploit-craft", "A large gain over Plain, but a tie with native Goal. This result is not unique to LoopX."], + ["apex-management-consulting-matter", "Higher than both baselines. The original brief separately reports recovery at S27; that association does not establish causality."] ], - "lossTitle": "Where more continuation was not enough", + "lossTitle": "Counterexamples: the task and baseline change the conclusion", "lossCases": [ - ["duckdb-optimizer-closure", "0.771", "0.000", "A later optimization replaced a correctness-safe version with a nondeterministic one. LoopX needs best-so-far checkpointing and final rollback."], - ["satellite-flood-change-detection-audit", "0.285", "0.040", "Additional work did not recover the hidden semantic target; continuation without diagnostic evidence can extend the wrong hypothesis."], - ["sokoban", "0.981", "0.890", "The task remained strong, but the independent run found a weaker solution. Single-trial sampling variation remains material."], - ["apex-openroad-ibex-signoff", "0.000", "0.000", "One physical-design baseline consumed most of the budget. The limiting factor was tool runtime, not frontier recovery."] + ["duckdb-optimizer-closure", "Native Goal retains a substantial score; LoopX and Plain score zero. The original brief reports a correctness regression after optimization, motivating best-version retention and final rollback."], + ["snake-obstacle-campaign", "Higher than native Goal, lower than Plain. Choosing a different baseline changes the apparent result."], + ["great-expectations-audit", "Lower than both baselines. Adding persistent scheduling does not guarantee a better final artifact."], + ["apex-openroad-ibex-signoff", "All three score zero. The control plane did not solve this task; continued execution alone is not evidence of useful progress."] ], "scoresEyebrow": "04 / Complete task matrix", - "scoresTitle": "All 46 tasks, on one denominator.", - "scoresBody": "Scores are LHTB continuous rewards. The highest observed value in each row is highlighted. Filters expose large cross-arm spreads and material New-vs-Legacy Heartbeat changes.", + "scoresTitle": "Inspect LoopX’s gains and losses against each baseline.", + "scoresBody": "The default view shows three arms; blue marks the highest visible reward in each row. Filters select |LoopX − baseline| ≥ 0.05. Values display four decimals; comparisons use full precision. Historical arms can be expanded.", "searchPlaceholder": "Search task", "filterLabel": "Task filters", - "filters": {"all": "All 46", "spread": "Spread ≥ 0.20", "heartbeat": "Heartbeat |Δ| ≥ 0.05"}, + "filters": { + "all": "All 46", + "plain": "vs Plain |Δ| ≥ 0.05", + "native_goal": "vs Goal |Δ| ≥ 0.05" + }, "taskColumn": "Task", "visibleCount": "Showing {count} tasks", - "programEyebrow": "05 / Research-program mapping", - "programTitle": "This study feeds the LoopX long-horizon benchmark program.", - "programBody": "The five-arm study predates the shared runtime matrix introduced later in LoopX. It should be read as evidence that motivated the program, not as a completed run of every current execution-mode × task-entry cell.", + "programEyebrow": "05 / Next experiments", + "programTitle": "Test whether the gains beyond the baselines reproduce.", + "programBody": "This study provides outcome differences and mechanism clues. Align Plain, native Goal, and LoopX on environment and budget first, then isolate recovery, plan review, and rollback instead of changing several variables together.", "programSteps": [ - ["Native metric", "Keep LHTB reward and the ≥0.95 solve threshold; do not collapse it into a private LoopX score."], - ["Typed treatment", "Record continuation owner, session mode, task entry, replan policy, model, effort, and budget as separate dimensions."], - ["Mechanism evidence", "Use trajectories to test stall, recovery, repeated work, and correction after contradictory evidence."], - ["Integrity gate", "Treat verifier isolation and benchmark fairness as independent qualification axes, not optional footnotes."] + ["Fix the comparison", "Use the same model, task version, time limit, and accounting; repeat every task and report every attempt."], + ["Separate mechanisms", "Vary fresh/resumed sessions, Todo planning, and review cadence independently to identify what changes outcomes."], + ["Measure gains and losses", "Report mean reward, strict solves, per-task regressions, cost, and evidence of recovery and repeated work together."], + ["Protect final delivery", "Test independent acceptance and best-version retention so later optimization cannot erase correct results."] ], "programAction": "Read the Long-Horizon Harness RFC", "programUrl": "https://github.com/huangruiteng/loopx/blob/main/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md", "limitsEyebrow": "06 / Evidence boundary", - "limitsTitle": "Promising mechanism evidence, not a universal capability claim.", + "limitsTitle": "A reason to investigate further, not proof of reliable gains at equal budgets.", "limits": [ - "One effective trial per task-arm cell; there are no repeated seeds or confidence intervals.", - "New Heartbeat changes transport, session lifetime, LoopX version, and replan policy together; the comparison is not a single-variable ablation.", - "The effective aggregates include designated replacement trials. Some New Heartbeat replacements used longer budgets; all replacements are disclosed in the underlying study artifacts.", - "LoopX receives no hidden verifier diagnostics. More continuation cannot identify an invisible semantic mismatch by itself.", - "Tokens and USD use different telemetry conventions across historical and new runs; treat cost as descriptive, not a normalized efficiency ranking." + "One effective run per task-arm cell, with no repeated seeds or confidence intervals. Win/loss counts are not significance tests.", + "Model and reasoning effort match, but transport, session lifetime, and scheduling differ between LoopX and the baselines. No single mechanism is isolated.", + "Effective aggregates include replacement trials, some with longer LoopX limits. This is not a strictly matched-budget comparison.", + "Hidden evaluation grades the final result; continuation does not automatically reveal an unknown correct solution. Recovering work and solving the task remain distinct.", + "Accounting differs: LoopX records $551.73, while Plain and native Goal estimate $212.97 and $87.58. A higher mean does not establish better efficiency." ], - "nextTitle": "The next decisive experiment", - "nextBody": "Run the shared LoopX runtime matrix with repeated seeds and one variable changed at a time: task entry, fresh vs resumed executor context, replan cadence, and best-so-far rollback. Pre-register reward, solve rate, cost, stall recovery, and repeated-work metrics.", + "nextTitle": "Three questions worth testing", + "nextBody": "At equal budgets, which tasks does LoopX finish beyond native Goal? Do its larger gains over Plain survive repeated runs? Can independent acceptance and rollback reduce losses like DuckDB?", "sourcesEyebrow": "07 / Sources", "sourcesTitle": "Public links and reproducible aggregate evidence.", "sourceItems": [ + ["46-task public aggregate", "Full scores, runtime settings, and cost provenance.", "https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json"], + ["Original mechanism case report", "Reported S27 recovery and DuckDB regression; observational case evidence.", "https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json"], ["LoopX LHTB runner", "Current generic_cli + fresh codex exec Heartbeat implementation and fairness contract.", "https://github.com/huangruiteng/loopx/tree/main/benchmark/LHTB"], ["LHTB official site", "Benchmark definition, leaderboard, and methodology.", "https://zli12321.github.io/LHTB/index.html"], ["LHTB paper", "Long-Horizon Terminal-Bench, arXiv:2607.08964.", "https://arxiv.org/abs/2607.08964"], ["LoopX benchmark RFC", "The broader long-horizon harness benchmark and research program.", "https://github.com/huangruiteng/loopx/blob/main/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md"] ], - "attestation": "Public aggregate only: no credentials, local paths, gateway addresses, hidden verifier data, or raw trajectories are published here.", + "attestation": "Scores come from the public 46-task aggregate; comparisons are computed from task rewards. Mechanism cases refer to the original research brief, not new experimental findings.", "footer": "An exploratory LoopX research brief built from 46-task aggregate evidence.", - "backTop": "Back to top" + "backTop": "Back to top", + "comparisonTitle": "Relative to {baseline}", + "relativeGain": "relative mean gain", + "pairCounts": "{wins} wins · {ties} ties · {losses} losses", + "pairSolves": "Solved: LoopX {current}, {baseline} {base}", + "comparisonNote": "Deltas are computed from all 46 unrounded task rewards, then rounded for display. Relative gain = mean delta / baseline mean. Wins, ties and losses use strict comparisons of published rewards; equal values do not establish statistical equivalence.", + "historyTitle": "Historical context: SSH-Goal and legacy Heartbeat", + "historyBody": "The other two arms remain available for tracing implementation history. The primary comparison throughout this page is this study’s LoopX Heartbeat against Plain and native Goal.", + "showHistory": "Show historical arms", + "hideHistory": "Hide historical arms" }, "zh": { - "meta": "LHTB / 五臂探索性研究", + "meta": "LHTB / LOOPX 与两种基线", "back": "返回 LoopX", "source": "研究代码", - "title": "五种让 Agent 跨越长程任务持续工作的方式。", - "deck": "对比 Plain、原生 Goal、LoopX SSH-Goal、旧版 Heartbeat 与 fresh-exec 新版 Heartbeat 在 46 个任务上的表现。重点不是只看榜单,而是研究持久状态、Todo 与 replan 如何改变续跑行为。", + "title": "长程任务里,LoopX 比 Plain 和原生 Goal 多做成了什么?", + "deck": "在同一模型的 46 个任务上,LoopX Heartbeat 的平均 Reward 高于 Plain 和原生 Goal;完整通过数与 Plain 相同,高于原生 Goal。下面从逐题得失、工作机制和失败案例看清这组信号。", "evidenceTag": "GPT-5.6 Sol · max 推理 · 46 任务 · 每个 cell 1 条有效轨迹", "contributorsLabel": "研究贡献者", "readingLabel": "阅读路径", "readingPath": [ - ["01", "结果", "#result"], - ["02", "五臂机制", "#mechanisms"], - ["03", "LoopX Insight", "#insights"], - ["04", "46 题", "#scores"] + ["01", "结果对比", "#result"], + ["02", "工作机制", "#mechanisms"], + ["03", "得失与启示", "#insights"], + ["04", "逐题核对", "#scores"] ], "resultEyebrow": "00 / 核心结果", - "resultTitle": "新版 Heartbeat 均分最高,收益集中在进度可外部化、可恢复的任务。", - "resultBody": "Plain 均分为 0.4218,新版 Heartbeat 为 0.4948。与旧版 Heartbeat 相比,新版在 46 题中 17 胜、19 平、10 负,均分增加 0.0154。这是有价值的方向性证据,不是因果效应估计。", - "summaryColumns": ["实验臂", "平均 Reward", "通过 ≥ 0.95", "记录 Tokens", "USD"], + "resultTitle": "均分高于两种基线,完整通过数尚未超过 Plain。", + "resultBody": "LoopX 为 0.4948,Plain 为 0.4218,原生 Goal 为 0.4475。以 Reward ≥ 0.95 计通过,三者分别为 7、7、4 题。均分反映部分进展,完整验收还需单独看。", + "summaryColumns": ["执行方式", "平均 Reward", "通过 ≥ 0.95", "Token 用量", "USD"], "armLabels": { "plain": "Plain", "native_goal": "原生 Goal", "ssh_goal": "LoopX SSH-Goal", "legacy_heartbeat": "旧版 Heartbeat", - "new_heartbeat": "新版 Heartbeat" + "new_heartbeat": "LoopX Heartbeat" }, "armKinds": { "plain": "基线 · app-server", "native_goal": "基线 · app-server", - "ssh_goal": "LoopX · app-server", + "ssh_goal": "历史参照 · app-server", "legacy_heartbeat": "LoopX 0.5.3 · app-server", "new_heartbeat": "LoopX 1.0.3 · generic_cli" }, - "deltaMean": "相对旧 Heartbeat 均分", - "deltaWins": "任务胜出", - "deltaTies": "持平", - "deltaLosses": "下降", + "deltaMean": "平均 Reward 增量", "readingNoteLabel": "阅读口径", - "readingNote": "前四臂是历史 Codex app-server 实验;新版 Heartbeat 使用 generic_cli,每次 wake 启动全新的 codex exec。因此新旧比较同时改变了多个 runtime 维度。历史成本为估算值,新版 Heartbeat 成本来自运行记录。", + "readingNote": "这里的 LoopX 指实验中的 1.0.3 Heartbeat(fresh exec),不是对最新版本的重新测试。两种基线使用 app-server,LoopX 使用 generic_cli;有效运行包含替代 trial,部分时限更长。历史 USD 为估算,LoopX USD 来自运行记录,不能作同预算效率排名。", "benchmarkEyebrow": "01 / Benchmark", "benchmarkTitle": "LHTB 在数百个依赖终端动作之后,按最终产物评分。", - "benchmarkBody": "Long-Horizon Terminal-Bench 包含 46 个容器化任务,覆盖九大类别。隐藏的 rebuild-from-artifact verifier 在 [0,1] 连续区间评价最终工作区;模型自报完成不计分。本研究采用 Reward ≥ 0.95 作为通过阈值。", + "benchmarkBody": "Long-Horizon Terminal-Bench 包含 46 个容器化任务,覆盖九大类别。隐藏验证器在 [0,1] 连续区间评价产物或可重放结果;模型自报完成不计分。本研究采用 Reward ≥ 0.95 作为通过阈值。", "officialLinks": [ ["LHTB 官方网站", "任务分类、排行榜、方法和 harness 说明。", "https://zli12321.github.io/LHTB/index.html"], ["LHTB 论文", "Long-Horizon Terminal-Bench,arXiv:2607.08964。", "https://arxiv.org/abs/2607.08964"] @@ -214,79 +221,92 @@ ], "taskExample": "例如 UNISON 论文复现,交付物是能运行的仿真流水线和报告。隐藏检查会核对指标、分区与调度、重复运行的一致性,以及留出 seed 下的行为;部分得分反映产物已经满足了哪些要求。", "benchmarkSetupNote": "官方模型比较使用 Terminus-2 和每题 90 分钟预算;本页是另行开展的 LoopX 五臂研究,执行设置与替代运行以本页披露为准。上游任务与验证器仍在演进,历史成绩保留其原始评测设置。", - "mechanismEyebrow": "02 / 五种执行机制", - "mechanismTitle": "模型相同;续跑权和持久状态的归属不同。", - "mechanismBody": "五臂均使用 GPT-5.6 Sol、max 推理,并关闭 web search。实验改变的是谁决定继续工作、下一段工作如何表示,以及新执行器能否不依赖旧对话指令恢复任务前沿。", + "mechanismEyebrow": "02 / 三种工作方式", + "mechanismTitle": "目标如何持续,下一步由谁组织?", + "mechanismBody": "三者都能使用同一模型和工具完成长任务。区别在于:Plain 由会话自行推进;原生 Goal 持续保留目标;LoopX 另外维护工作项与推进状态,并按状态唤醒执行器。以下是配置差异,分数尚不能证明其中哪个机制贡献了增益。", "mechanismRows": [ - ["plain", "Plain", "由模型会话决定停止", "普通 app-server turn 自己规划、修改并返回;没有持久 Goal,也没有 LoopX registry。"], - ["native_goal", "原生 Goal", "由 Codex Goal 维持连续性", "持久的“Finish the task”目标可跨 turn 保留,但不引入 LoopX Todo、quota 或 registry 控制。"], - ["ssh_goal", "LoopX SSH-Goal", "由 LoopX dispatcher 决定下一工作切片", "Registry、quota、已领取 Todo、bounded work 与 writeback 让跨 phase 的下一动作保持显式。"], - ["legacy_heartbeat", "旧版 Heartbeat", "app-server 中的周期 LoopX 义务", "LoopX 0.5.3 周期性重入长生命周期 app-server,并要求推进一个有边界的工作切片。"], - ["new_heartbeat", "新版 Heartbeat", "持久控制面,可替换执行器", "LoopX 1.0.3 每次 wake 启动全新的 codex exec。工作区与 registry 持久化,旧对话指令不复用;每完成三个有效推进 Todo 打开一次 replan review。"] + ["plain", "Plain", "会话内规划与执行", "普通 app-server turn 自己规划、调用工具并返回;可以长时间工作,但没有原生持久 Goal 或 LoopX 控制面。"], + ["native_goal", "原生 Goal", "跨 turn 保留目标", "Codex 维持“Finish the task”目标,支持后续推进;本实验未接入 LoopX 的 Registry、Todo 与调度。"], + ["new_heartbeat", "LoopX Heartbeat", "工作状态跨执行器保留", "工作区、Registry 与 Todo 持久化,每次唤醒启动新的 codex exec;三个有效推进 Todo 后触发复核,可以维持或调整计划。"] ], "insightEyebrow": "03 / LoopX Insight", - "insightTitle": "优势不是更长的聊天,而是可恢复的任务前沿。", - "insightBody": "当进度能写入文件和 Todo 时,证据最清楚:当前阶段、已通过检查、剩余模块以及下一条可证伪动作都留在环境中。新的执行器无需从长对话叙述重建意图,就能从项目证据继续。", + "insightTitle": "均分增加了;优势出现在哪里,还需要逐题解释。", + "insightBody": "对 Plain,17 题更高、16 题更低,均分提升来自收益幅度大于损失,而非多数任务都更强。对原生 Goal,23 题更高、10 题更低。下列案例同时列出三者分数,避免把选择基线造成的差异误读为普遍收益。", "insights": [ - ["状态跨会话存续", "工作区产物加 LoopX registry/Todo 保留停止位置,执行器会话不再是唯一的记忆所有者。"], - ["Todo 让续跑可检查", "已领取的 bounded slice 分开当前工作、已完成检查和 successor 工作,降低分阶段修复和多产物任务的恢复歧义。"], - ["Replan 能重新打开假终点", "三个有效 Todo 完成后,新版 Heartbeat 打开 review 义务;负面证据可以重定向任务前沿,而不是接受陈旧的“完成”。"], - ["Fresh exec 降低指令漂移", "每次 wake 只取得当前 task body 与项目状态,不继承累积对话,把持久控制状态和模型会话状态分开。"] + ["部分进展增加,完整验收仍是难点", "LoopX 的 0.4948 高于两种基线,但与 Plain 同为 7/46。Tabular、Sokoban 等收益明显的任务仍未达到 0.95,说明更多进展与做完之间还有距离。"], + ["原生 Goal 已经是有力的对照", "PoC 任务中,Goal 与 LoopX 都是 0.892,Plain 为 0;DuckDB 则是 Goal 0.770、LoopX 0。需要解释 LoopX 在原生续跑之上多带来了什么,也保留它做得更差的任务。"], + ["恢复是机制线索,尚不是增益归因", "研究简报记录咨询任务中断后从 S27 接续。工作区与 Todo 提供了恢复位置的载体;要证明它提升了分数,还需在相同预算下隔离状态恢复、fresh exec 与 replan。"], + ["续跑要保住结果,也要算成本", "DuckDB 的报告案例提示后续优化可能破坏正确性。检查点、验收与回滚值得单独验证;本轮 LoopX 记录成本高于两种基线的估算成本,尚未证明更经济。"] ], - "gainTitle": "新版 Heartbeat 提升样例", + "gainTitle": "收益案例:看三者,而非只看一个增量", "gainCases": [ - ["tabular-data-feature-covshift", "0.292", "0.743", "紧凑的产物—检查闭环给 successor 工作提供了明确目标,是本轮最大提升。"], - ["rush-hour-campaign", "0.481", "0.772", "棋盘状态和未完成阶段可以恢复,后续工作继续搜索,而不是重新开始。"], - ["apex-management-consulting-matter", "0.433", "0.513", "首轮中断后,fresh executor 从 S27 恢复并完成最后四个 stage。"], - ["unknown-config-semantics", "0.000", "0.231", "外部化产物为下一次 wake 提供了足够证据,使旧版零分任务取得部分进展。"] + ["tabular-data-feature-covshift", "相对两种基线均有明显增益,但 0.743 仍未达到严格通过阈值。"], + ["sokoban", "高于两种基线;最终 0.890,更多进展仍未转化为完整通过。"], + ["poc-exploit-craft", "相对 Plain 提升明显,却与原生 Goal 持平。这项收益不能直接归为 LoopX 特有。"], + ["apex-management-consulting-matter", "分数高于两种基线;原简报另记录了从 S27 恢复的轨迹案例。两者相关,尚不足以证明因果。"] ], - "lossTitle": "仅靠继续工作仍解决不了的情况", + "lossTitle": "反例:续跑收益随任务和基线变化", "lossCases": [ - ["duckdb-optimizer-closure", "0.771", "0.000", "后续优化把正确安全版本替换成非确定版本;LoopX 仍缺 best-so-far checkpoint 和最终回滚。"], - ["satellite-flood-change-detection-audit", "0.285", "0.040", "额外工作没有命中隐藏语义目标;缺少诊断证据时,续跑也可能延长错误假设。"], - ["sokoban", "0.981", "0.890", "新版仍然较高,但独立采样找到的解更弱;单 trial 的采样方差不可忽略。"], - ["apex-openroad-ibex-signoff", "0.000", "0.000", "一次物理设计 baseline 消耗了大部分时限,瓶颈是外部工具时长,不是任务前沿恢复。"] + ["duckdb-optimizer-closure", "原生 Goal 保住了较高得分,LoopX 与 Plain 为零。原简报报告后续优化破坏正确性,值得验证最佳版本保留与最终回滚。"], + ["snake-obstacle-campaign", "高于原生 Goal,却低于 Plain。换一个基线,结论就会变化。"], + ["great-expectations-audit", "同时低于两种基线,说明引入持续调度并不保证更好的最终产物。"], + ["apex-openroad-ibex-signoff", "三者均为零。控制面未解决这道任务,不能将继续执行本身当成有效进展。"] ], "scoresEyebrow": "04 / 46 题完整矩阵", - "scoresTitle": "46 个任务使用同一分母。", - "scoresBody": "表格展示 LHTB 连续 Reward,并高亮每题观测到的最高值。筛选器可查看臂间差异较大的任务,以及新旧 Heartbeat 变化明显的任务。", + "scoresTitle": "逐题看 LoopX 相对两种基线的得失。", + "scoresBody": "默认展示三者,蓝色标出可见列中的最高 Reward。筛选条件是 LoopX 与指定基线的绝对差 ≥ 0.05;显示四位小数,比较使用原始精度。历史两臂可按需展开。", "searchPlaceholder": "搜索任务", "filterLabel": "任务筛选", - "filters": {"all": "全部 46 题", "spread": "臂间差异 ≥ 0.20", "heartbeat": "Heartbeat |Δ| ≥ 0.05"}, + "filters": { + "all": "全部 46 题", + "plain": "对 Plain |Δ| ≥ 0.05", + "native_goal": "对 Goal |Δ| ≥ 0.05" + }, "taskColumn": "任务", "visibleCount": "当前展示 {count} 个任务", - "programEyebrow": "05 / 映射到研究计划", - "programTitle": "这项实验为 LoopX 长程 benchmark 研究计划提供输入。", - "programBody": "五臂实验早于 LoopX 后续统一的 shared runtime matrix。它应被理解为推动研究计划形成的证据,而不是已经跑完所有新版 execution-mode × task-entry cell。", + "programEyebrow": "05 / 下一步研究", + "programTitle": "下一步:验证比基线多做成的部分能否稳定复现。", + "programBody": "本轮提供了结果差异与机制线索。后续先用统一环境和预算对齐 Plain、原生 Goal 与 LoopX,再分别检查状态恢复、计划复核和回滚;避免一次更换多个变量。", "programSteps": [ - ["保留原生 Metric", "继续报告 LHTB Reward 与 ≥0.95 通过率,不把它折叠成私有 LoopX 分数。"], - ["Treatment 类型化", "把续跑所有者、会话模式、任务入口、replan 策略、模型、推理强度和时限分开记录。"], - ["机制证据", "用轨迹研究 stall、恢复、重复工作,以及矛盾证据出现后的纠偏质量。"], - ["独立诚信门槛", "Verifier 隔离和 benchmark 公平性是独立资格轴,不能只放在脚注里。"] + ["先固定对照", "同模型、任务版本、时限与计费口径,对每题重复运行并报告所有尝试。"], + ["再拆机制", "分别控制 fresh/resume、Todo 规划、replan 节奏,检验哪个环节改变结果。"], + ["同时看收益和损失", "共同报告均分、严格通过数、逐题退步、成本,以及中断恢复与重复工作的证据。"], + ["验证最终交付", "检查独立验收与最佳版本保留;不能让后续优化冲掉已经正确的成果。"] ], "programAction": "阅读 Long-Horizon Harness RFC", "programUrl": "https://github.com/huangruiteng/loopx/blob/main/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md", "limitsEyebrow": "06 / 证据边界", - "limitsTitle": "这是有希望的机制证据,不是普适能力结论。", + "limitsTitle": "这组结果支持继续研究,尚不能证明同预算下稳定领先。", "limits": [ - "每个 task-arm cell 只有一条有效轨迹,没有多 seed 重复或置信区间。", - "新版 Heartbeat 同时改变 transport、会话生命周期、LoopX 版本和 replan 策略,不是单变量消融。", - "有效聚合包含指定的替代 trial;部分新版 Heartbeat 替代运行使用更长时限,底层研究产物中保留了披露。", - "LoopX 不会获得隐藏 verifier 的诊断信息;单纯增加续跑无法自己识别不可见的语义偏差。", - "历史臂与新版臂的 token、USD telemetry 口径不同;成本只作描述,不作为归一化效率排行。" + "每题每臂仅一条有效运行,没有重复 seed 或置信区间;逐题胜负也不是显著性检验。", + "模型与推理强度一致,但 LoopX 与基线同时存在传输、会话和调度策略差异,不能把收益归因于某一个机制。", + "有效聚合包含替代 trial,部分 LoopX 替代运行采用更长时限;当前结果不是严格匹配预算的对照。", + "隐藏评测负责最终评分,续跑不能自动获得未知的正确解;恢复能力与任务本身能否解决是两件事。", + "成本记录口径不同;LoopX 记录 $551.73,Plain 与原生 Goal 分别估算 $212.97、$87.58。均分更高不等于效率更高。" ], - "nextTitle": "下一项决定性实验", - "nextBody": "按统一 LoopX runtime matrix 做多 seed 对照,每次只改变一个变量:task entry、fresh/resume executor context、replan cadence,以及 best-so-far rollback。预先注册 Reward、通过率、成本、stall recovery 和重复工作指标。", + "nextTitle": "最值得追问的三个问题", + "nextBody": "相同预算下,LoopX 还能比原生 Goal 多完成哪些任务?对 Plain 的较大增益能否跨重复运行保留?独立验收和回滚能否减少 DuckDB 这类退步?", "sourcesEyebrow": "07 / 来源", "sourcesTitle": "公开链接与可复现的聚合证据。", "sourceItems": [ + ["46 题公开聚合", "全部分数、实验设置与成本来源。", "https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json"], + ["原始机制案例报告", "S27 恢复与 DuckDB 回退的报告来源,属于观察性案例。", "https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json"], ["LoopX LHTB runner", "当前 generic_cli + fresh codex exec Heartbeat 实现与公平性契约。", "https://github.com/huangruiteng/loopx/tree/main/benchmark/LHTB"], ["LHTB 官方网站", "Benchmark 定义、排行榜与方法。", "https://zli12321.github.io/LHTB/index.html"], ["LHTB 论文", "Long-Horizon Terminal-Bench,arXiv:2607.08964。", "https://arxiv.org/abs/2607.08964"], ["LoopX benchmark RFC", "更完整的长程 harness benchmark 与研究计划。", "https://github.com/huangruiteng/loopx/blob/main/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md"] ], - "attestation": "这里只发布公开聚合数据,不包含凭据、本地路径、网关地址、隐藏 verifier 数据或原始轨迹。", + "attestation": "分数来自公开 46 题聚合;比较指标由逐题分数计算。机制案例沿用原研究简报的报告,不新增实验结论。", "footer": "基于 46 题聚合证据的 LoopX 探索性研究简报。", - "backTop": "回到顶部" + "backTop": "回到顶部", + "comparisonTitle": "相对 {baseline}", + "relativeGain": "相对均分", + "pairCounts": "{wins} 胜 · {ties} 平 · {losses} 负", + "pairSolves": "通过数:LoopX {current},{baseline} {base}", + "comparisonNote": "差值由 46 题原始 Reward 计算后四舍五入;相对增幅 = 均分差 ÷ 基线均分。胜 / 平 / 负按公开分数逐题严格比较,平局表示数值相等,不代表统计等价。", + "historyTitle": "历史参照:SSH-Goal 与旧版 Heartbeat", + "historyBody": "另外两臂保留在这里,便于回溯实现演进。全文的主比较是本研究的 LoopX Heartbeat 与 Plain、原生 Goal。", + "showHistory": "显示历史两臂", + "hideHistory": "收起历史两臂" } } diff --git a/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md b/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md index 5e53af6b78..7f0e2ad31d 100644 --- a/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md +++ b/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md @@ -20,8 +20,10 @@ gateway addresses, credentials, and verifier artifacts. - Each task-arm cell contributes one effective trial. This is not a repeated-seed estimate. -- The New-vs-Legacy Heartbeat comparison changes multiple runtime dimensions; - it is mechanism evidence, not a single-variable causal ablation. +- The primary reading compares the 1.0.3 Heartbeat arm with Plain and Native + Goal. Runtime and budget differences prevent a single-variable causal or + equal-budget efficiency interpretation. SSH-Goal and Legacy Heartbeat are + retained as historical context. - The effective aggregates include designated replacement trials. Some New Heartbeat replacements used longer budgets. - Historical cost is estimated from retained token telemetry. New Heartbeat @@ -29,3 +31,24 @@ gateway addresses, credentials, and verifier artifacts. - LHTB reward and the `>= 0.95` solved threshold remain benchmark-native. The runnable current Heartbeat implementation lives in `benchmark/LHTB/`. + +## Reading the baseline comparisons + +The brief derives these comparisons from the 46 `tasks` rows in `data.json`, +without modifying the experiment data: + +| LoopX 1.0.3 Heartbeat versus | Mean reward delta | Relative mean gain | Wins / ties / losses | Strict solves (LoopX / baseline) | +| --- | ---: | ---: | --- | --- | +| Plain | +0.0731 | +17.3% | 17 / 13 / 16 | 7 / 7 | +| Native Goal | +0.0473 | +10.6% | 23 / 13 / 10 | 7 / 4 | + +Mean delta is the mean of per-task differences; relative gain divides that +unrounded delta by the baseline mean. Wins, ties and losses compare published +unrounded rewards strictly; equality is not statistical equivalence. Display +rounding is applied only afterwards. The historical `heartbeat_comparison` +summary is retained in the source archive but is not used for these comparisons. + +The task matrix defaults to Plain, Native Goal and LoopX Heartbeat; readers can +expand both historical arms. Case scores come from the same task rows rather +than a separate editorial copy. The recorded scores support outcome comparisons; +reported recovery or regression cases are mechanism clues, not causal estimates.
      {c.taskColumn}{c.armLabels[arm]}
      {c.taskColumn}{c.armLabels[arm]}
      {row.task}{formatReward(row[arm])}