From f26910438023cc1cb049e3a831086669d7601b1c Mon Sep 17 00:00:00 2001 From: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Date: Fri, 18 Sep 2026 01:35:46 +0800 Subject: [PATCH] feat(site): publish application scenarios and DeepSWE Sol research brief Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> --- apps/presentation/site/README.md | 15 ++ .../public/benchmarks/deepswe-sol/index.html | 137 ++++++++++++ .../public/benchmarks/deepswe-sol/study.css | 148 +++++++++++++ .../blog/zh/application-scenarios/index.html | 195 ++++++++++++++++++ .../zh/application-scenarios/presentation.js | 28 +++ .../site/public/blog/zh/index.html | 4 + apps/presentation/site/src/App.tsx | 7 +- apps/presentation/site/vite.config.ts | 4 +- .../dashboard-frontstage-browser-smoke.mjs | 4 +- ...board-frontstage-design-baseline-smoke.mjs | 2 +- examples/export-frontstage-share-bundle.mjs | 4 + examples/frontstage-share-bundle-smoke.mjs | 27 ++- 12 files changed, 565 insertions(+), 10 deletions(-) create mode 100644 apps/presentation/site/public/benchmarks/deepswe-sol/index.html create mode 100644 apps/presentation/site/public/benchmarks/deepswe-sol/study.css create mode 100644 apps/presentation/site/public/blog/zh/application-scenarios/index.html create mode 100644 apps/presentation/site/public/blog/zh/application-scenarios/presentation.js diff --git a/apps/presentation/site/README.md b/apps/presentation/site/README.md index a480765c7..b4a27ec8a 100644 --- a/apps/presentation/site/README.md +++ b/apps/presentation/site/README.md @@ -20,6 +20,21 @@ navigation work without JavaScript. Vite copies these pages into both the local build and the existing Pages export; no separate hosting or content service is needed. Relative navigation supports both root and repository base paths. +The Chinese DeepSWE × Sol research brief is a static page at +`public/benchmarks/deepswe-sol/`, linked from the homepage research collection. +It uses the shared editorial tokens and ships through the same public-directory +copy as the Blog. Its historical results and mechanism explanations cite the +immutable v1 archive; editing the brief must not rewrite that archive or restore +withdrawn scores. The article works without JavaScript and has stable section +anchors for other articles to cite. + +The Chinese application-scenarios article at `public/blog/zh/application-scenarios/` +links the three research briefs with comparable summaries. Its full text and +initial PR-state example are static HTML; `presentation.js` progressively adds +presentation typography, section navigation, and the synthetic state selector. +These controls stay hidden when JavaScript is unavailable. No real repository +state or write APIs are involved. + Edit the paired HTML editions together, including their index summaries and metadata. Preserve matching section anchors and public source attribution. Keep source-document exports, private references, and unreviewed media outside diff --git a/apps/presentation/site/public/benchmarks/deepswe-sol/index.html b/apps/presentation/site/public/benchmarks/deepswe-sol/index.html new file mode 100644 index 000000000..08552af5b --- /dev/null +++ b/apps/presentation/site/public/benchmarks/deepswe-sol/index.html @@ -0,0 +1,137 @@ + + + + + + + DeepSWE × Sol:从继续执行到有效交付 · LoopX + + + + + + + + +
+
+ LoopX / RESEARCH + +
+
+
+
+
+

DeepSWE · GPT-5.6 Sol · v1 研究解读

+

从继续执行,
走到有效交付。

+

113 个软件工程任务,把长程 Agent 的三个问题放到一起:工作怎样继续,补丁是否交付,结果是否正确。

+

研究归档贡献 @gwh6669999 ↗
2026 年 9 月 16 日合入 · 实验修订 2cef51d

+ 看运行器怎样做出下一步判断 +
+
+

113 TASKS / HISTORICAL BEST-VALID

+

三种配置,三组完成结果

+
裸 Codex54 / 113
47.8%
+
原生 Goal60 / 113
53.1%
+
LoopX + Heartbeat70 / 113
61.9%
+ +
每任务取最佳有效结果,非单次通过率。不同配置的时间窗口不一致,不能解释为同预算下的机制增益。
+
+
+ + + +
+

01 / RESULTS

先把数字放回它的统计口径。

Heartbeat 相比原生 Goal 多完成 10 个任务,完成比例相差约 8.8 个百分点。这是值得继续研究的观察;汇总表本身还不能说明,差异来自额外计算量、恢复策略,还是交付修复。

+
+ + + + + + + +
来源:v1 README。统一分母 113,聚合方式为 per-task best valid。
配置完成任务完成比例PartialF2PP2P
裸 Codex plain54 / 11347.8%0.92060.7450.997
原生 Goal goal60 / 11353.1%0.96200.8680.997
LoopX + Heartbeat heartbeat70 / 11361.9%0.97390.8860.993
+

Partial、F2P 与 P2P 沿用原表字段;F2P 关注失败测试转为通过,P2P 关注原有通过测试是否保持通过。公开归档未提供逐题归约数据与完整计分脚本,本页不重新定义权重,也不将这些比例当作独立任务数。

+
+

已能看到什么

原生 Goal 相比单次执行已经出现差异;Heartbeat 的历史完成数进一步上升。三组 P2P 均接近 1,但 Heartbeat 略低,不能只用完成数概括全部质量变化。

+

还不能推导什么

没有逐任务的成败配对,无法知道哪些任务转正或转负;没有完整尝试次数和成本,无法估计成功率稳定性、单位成本收益或单项机制的因果效应。

+
+
+ +
+

02 / EXECUTION CONFIGURATIONS

比较的核心:谁负责让工作继续?

这五种配置同时改变了运行接口、状态组织和继续方式。它们提供不同的实验切口,不能看成只开关一个功能的消融实验。

+
+
PLAIN

裸 Codex

使用 codex app-server,不启用 Goal 或 LoopX,按单次执行方式工作。

历史结果见上表
+
GOAL

Codex 原生 Goal

仍使用 app-server;原生 Goal 保持活跃时,由 app-server 继续推进。

历史结果见上表
+ +
CODEX-CLI

LoopX 控制面驱动 CLI

由外部调用 loopx turn run-once --host codex-cli,组织多段执行。

成绩与结论已撤回
+
SSH-GOAL

LoopX + 原生 Goal

使用 app-server,在同一会话与 Goal 中处理 blocked 状态并重新启动 turn。

成绩与结论已撤回
+
+

配置描述来自归档 README 与运行器源码,只说明历史实验怎样组织执行;不证明当前 LoopX 全部能力已经得到验证。

+
+ +
+

03 / CONTINUATION

Heartbeat 每轮都要重新检查交付。

归档里的 supervisor 并非只重复一句“继续”。它取回任务提示,恢复已有会话,检查代码交付和工作项状态,再决定结束或进入下一轮。

+
    +
  1. 01

    取回工作

    heartbeat-prompt --thin
    取出当前任务,附加补丁交付要求。

  2. +
  3. 02

    执行或恢复

    exec → exec resume
    从事件中保留会话标识,限制单段运行。

  4. +
  5. 03

    核对结果

    normalize_delivery
    整理补丁,再读取 P0 工作项状态。

  6. +
  7. 04

    收口或续跑

    交付与状态同时满足才收口;缺失交付时补修复工作项。

  8. +
+
+
有效补丁 + 收口条件满足

supervisor 返回执行完成;功能正确性仍交给独立验证器。

+
状态已收口,但补丁无效

新增交付修复 Todo,在剩余预算内继续。

+
预算或唤醒次数耗尽

返回失败,不把已运行很久当作完成依据。

+
+
源码细节:这里的“收口条件”具体检查什么?

历史函数 _goal_terminal 实际检查 Todo 列表:至少存在一个 P0 项,且所有 P0 项处于 done 或 deferred。它没有直接证明整个业务目标完成,也不验证补丁功能。

默认最多唤醒 8 次,轮间间隔 5 秒,单段超时 7,200 秒、总窗口 14,400 秒;命令行可覆盖这些值。单段子进程超时后 supervisor 仍可继续判断。这些是源码中的默认配置,不是历史每个任务的实际唤醒数或耗时。

阅读 Heartbeat supervisor 源码 ↗

+

继续执行需要判断依据,停止执行也需要。

+
+ +
+

04 / DELIVERY & VERIFICATION

代码写好了,验证器却可能拿到空补丁。

归档记录了一个具体的交付边界:Agent 可以在关联 worktree 中开发,但 DeepSWE 从主工作目录收集 git diff <base> HEAD。代码存在、提交存在,仍可能没有进入最终收集的产物。

+
+
工作发生处Agent worktree

修改与提交

交付整理唯一补丁与哈希核对

可应用 · 非空 · 内容一致

验收入口主目录的已提交补丁

交给独立验证器

+
机制示意,依据 workspace_delivery.py;不代表每次历史运行都发生了 worktree 恢复。
+
+
+

交付整理器检查什么

枚举工作树,收集相对任务基线的补丁,并按补丁哈希去重。没有改动时返回 empty;存在多个不同补丁时返回 ambiguous。只有唯一补丁才继续检查可应用性、恢复到主目录,并核对恢复前后的哈希。

+

独立验证器检查什么

交付整理器确认“补丁被交到了正确位置”。研究口径还要求无运行异常、任务校验和匹配、独立验证结果存在且一致。功能是否正确由验证器判断,不能用流程终态或非空补丁代替。

+
+
为什么这属于 Harness,而不能都算成模型能力?

历史整理器会提交未提交的改动,并在隔离实验工作区恢复补丁。这是运行器介入交付的行为。成功提交不一定完全由 Agent 自主完成;缺失交付也不一定意味着模型不会解题。比较时应把生成代码、运行器恢复和最终验证分开记录。

这份实现会对实验主目录执行重置与补丁恢复,依赖专用、隔离的任务环境。本文解释它的交付机制,不把原脚本当作日常仓库操作指南。

阅读补丁交付整理器源码 ↗

+

流程完成、产物交付、功能正确,是三个要分别验证的事实。

+
+ +
+

05 / SCOPE & NEXT EXPERIMENT

把观察变成结论,还缺哪些证据?

这份贡献提供了可读的运行器和历史汇总。要进一步判断哪种机制有效,需要把计算预算、重复尝试和交付处理重新放到可对照的实验里。

+
+ + + + + +
维度已公开的依据对结论的影响
模型与推理强度两个启动脚本默认 openai/gpt-5.6-sol / xhigh。逐次运行配置未公开,默认值不等于每次调用的证明。
时间与资源remaining59 脚本给 LoopX 的 Goal 超时为 14,400 秒,plain / goal 为 3,600 秒;LoopX 另设 3.0 的 agent timeout multiplier。不是等预算对比;超时上限也不是实际耗时。不能推导成本效率。
聚合与尝试113 任务,逐任务选最佳有效结果;没有完整逐题尝试清单。不是 pass@1,无法重建成败配对、方差或置信区间。
版本与复现实验基于 2cef51d。原始 14 个文件归档;依赖外部任务定义、支持模块与运行环境。不是独立可运行的发布包。文件身份、语法与导入检查不等于实验复现。
已知执行缺陷归档说明保留了 59 任务启动器与 54 任务准入不匹配、profile 初始化竞争、启动和重试处理等问题。不能把准入回执视为每次运行均有效的保证;需要在修复后的独立实验中验证。
+
    +
  1. 01

    冻结总预算与尝试口径

    匹配任务、模型、运行环境、总时间与 token 预算;保留每次尝试及基础设施失败,先定义重试如何计入。

  2. +
  3. 02

    拆开继续策略和交付恢复

    分别比较原生 Goal、外部续跑和补丁恢复;检查新完成的任务在哪个环节改变,避免多项机制同时变化。

  4. +
  5. 03

    同时报告质量、成本和人工介入

    独立验证最终交付物,展示逐题转正与转负、回归、总成本和人的处理次数,再决定哪些机制值得保留。

  6. +
+
+ +
+

06 / SOURCES

从研究汇总读到执行代码。

本文只使用公开归档与源码,引用固定到读取修订。没有运行新实验,没有新增或恢复已撤回的成绩。

+
    +
  1. 研究贡献与合入说明@gwh6669999 · #4502 · 保留原始 v1 文件及历史结果。
  2. +
  3. v1 README113 任务结果、五种配置与有效交付口径。
  4. +
  5. 归档状态与修订说明撤回范围、版本身份、外部依赖与已知缺陷。
  6. +
  7. remaining59 启动脚本Sol / xhigh 默认值,plain / goal 与 LoopX 的时间窗口差异。
  8. +
  9. 54 任务重跑启动脚本三种 LoopX 配置的准入与执行组织。
  10. +
  11. Heartbeat supervisor提示读取、会话恢复、状态检查、修复后继和停止条件。
  12. +
  13. Workspace delivery工作树枚举、补丁去重、恢复与交付哈希验证。
  14. +
+ +
+
+ + + diff --git a/apps/presentation/site/public/benchmarks/deepswe-sol/study.css b/apps/presentation/site/public/benchmarks/deepswe-sol/study.css new file mode 100644 index 000000000..c19521101 --- /dev/null +++ b/apps/presentation/site/public/benchmarks/deepswe-sol/study.css @@ -0,0 +1,148 @@ +/* Research layout; shared typography and tokens come from blog/blog.css. */ +.research-page { font-family: "Geist", "Inter", "Helvetica Neue", "PingFang SC", "Microsoft YaHei", Arial, sans-serif; line-height: 1.75; } +.research-page .shell { width: min(1160px, calc(100% - 64px)); margin-inline: auto; } +.research-page h1, .research-page h2, .research-page h3 { color: var(--ink); font-weight: 600; } +.research-page h2 { font-size: 30px; line-height: 1.4; letter-spacing: -.035em; margin: 12px 0 20px; } +.research-page h3 { font-size: 19px; line-height: 1.5; margin: 0 0 12px; } +.research-page p { color: var(--body); margin: 0 0 20px; } +.research-page code { font: .86em/1.7 var(--font-mono); overflow-wrap: anywhere; } +.research-page a { text-underline-offset: 4px; } +.research-page .caption { font-size: 13px; line-height: 1.7; } +.research-page .eyebrow { font-size: 11px; letter-spacing: .07em; margin-bottom: 16px; } +.research-header { position: sticky; top: 0; z-index: 5; background: var(--canvas); border-bottom: 1px solid var(--border); } +.research-nav { display: flex; align-items: center; justify-content: space-between; gap: 24px; min-height: 64px; } +.research-nav .brand { flex-shrink: 0; } +.nav-edition { font: 11px/1.5 var(--font-mono); color: var(--body); letter-spacing: .06em; margin-left: 8px; } +.research-nav nav { display: flex; gap: 28px; overflow-x: auto; white-space: nowrap; } +.research-nav nav a { font-size: 13px; min-height: 44px; display: inline-flex; align-items: center; } +.research-page .hero { display: grid; grid-template-columns: 1.12fr 1fr; gap: 64px; align-items: center; padding: 72px 0 48px; } +.research-page h1 { font-size: clamp(36px, 4vw, 52px); line-height: 1.25; margin: 20px 0 24px; letter-spacing: -.045em; } +.research-page .deck { font-size: 18px; line-height: 1.85; max-width: 510px; } +.research-page .contributor { margin: 28px 0 16px; font-size: 13px; line-height: 1.9; } +.contributor a { color: #005cc5; } +.text-link { display: inline-flex; align-items: center; gap: 16px; min-height: 44px; font-size: 14px; font-weight: 500; } +.result-figure { padding: 28px; background: var(--surface); border: 1px solid var(--border); border-radius: 12px; margin: 0; } +.result-figure h2 { font-size: 21px; margin-bottom: 28px; } +.result-figure figcaption { border-top: 1px solid var(--border); padding-top: 16px; font-size: 12px; line-height: 1.75; } +.bar-row { margin: 18px 0 0; position: relative; padding-bottom: 2px; } +.bar-label { display: flex; align-items: baseline; justify-content: space-between; gap: 16px; font-size: 14px; } +.bar-label strong { font-size: 26px; font-weight: 600; font-variant-numeric: tabular-nums; line-height: 1.2; } +.bar-label small { font-size: 13px; font-weight: 400; color: var(--body); } +.bar-track { height: 8px; background: var(--soft); margin: 10px 0 0; border-radius: 2px; } +.bar-track i { display: block; height: 100%; border-radius: inherit; background: #858585; } +.rate { display: block; font: 12px/1.5 var(--font-mono); margin-top: 4px; color: var(--body); } +.heartbeat .bar-label, .heartbeat .rate { color: #005cc5; } +.heartbeat .bar-track i { background: var(--blue); } +.bar-axis { display: flex; justify-content: space-between; font: 10px/1.5 var(--font-mono); color: var(--body); margin-top: 8px; } +.scope-note { padding: 20px 24px; border-left: 2px solid var(--ink); background: var(--soft); font-size: 13px; color: var(--body); } +.scope-note strong { color: var(--ink); } +.scope-note a { color: #005cc5; white-space: nowrap; } +.study-section { padding: 72px 0; border-bottom: 1px solid var(--border); scroll-margin-top: 92px; } +.section-lead { max-width: 820px; } +.section-lead > p:not(.eyebrow) { font-size: 17px; line-height: 1.85; } +.research-page .table-scroll { margin: 28px 0 16px; } +.table-scroll:focus-visible { outline: 2px solid var(--blue); outline-offset: 4px; } +.score-table { font-variant-numeric: tabular-nums; } +.score-table td { white-space: nowrap; } +.score-table th code { display: block; margin-top: 4px; font-weight: 400; color: var(--body); } +.score-table tbody th { background: transparent; } +.score-table .highlight { background: #edf5ff; } +.score-table .highlight td, .score-table .highlight th { color: #005cc5; } +.reading-pair { display: grid; grid-template-columns: 1fr 1fr; gap: 48px; margin: 32px 0; } +.reading-pair > div { border-top: 1px solid var(--border); padding-top: 24px; } +.reading-pair p { font-size: 15px; margin-bottom: 0; } +.arm-list { margin-top: 32px; } +.arm-list article { display: grid; grid-template-columns: 124px 1fr 150px; gap: 24px; align-items: start; padding: 24px 0; border-top: 1px solid var(--border); } +.arm-id { font: 12px/1.7 var(--font-mono); margin-top: 3px; } +.arm-list h3 { font-size: 17px; margin-bottom: 6px; } +.arm-list p { font-size: 14px; margin: 0; } +.arm-status { text-align: right; font-size: 12px; color: var(--body); } +.arm-featured .arm-id, .arm-featured .arm-status { color: #005cc5; } +.withdrawn { color: #8b5c00; } +.execution-flow { list-style: none; padding: 0; display: grid; grid-template-columns: repeat(4, 1fr); margin: 36px 0; border: 1px solid var(--border); border-radius: 12px; background: var(--surface); } +.execution-flow li { padding: 24px; border-right: 1px solid var(--border); min-width: 0; } +.execution-flow li:last-child { border: 0; } +.execution-flow li > span { display: block; margin-bottom: 20px; font: 12px/1.6 var(--font-mono); color: #005cc5; } +.execution-flow h3 { font-size: 17px; } +.execution-flow p { font-size: 13px; margin: 0; } +.decision-rows > div { display: grid; grid-template-columns: 290px 1fr; gap: 32px; padding: 16px 0; border-bottom: 1px solid var(--border); } +.decision-rows b { font-size: 14px; font-weight: 500; } +.decision-rows p { font-size: 14px; margin: 0; } +.research-page details { border: 1px solid var(--border); border-radius: 12px; padding: 12px 24px; margin-top: 28px; } +.research-page summary { min-height: 44px; padding: 10px 0; cursor: pointer; font-size: 14px; font-weight: 500; } +.research-page details p { font-size: 14px; margin: 12px 0 16px; } +.research-page details a { color: #005cc5; } +.research-page .takeaway { font-size: 23px; line-height: 1.6; color: var(--ink); margin: 36px 0 0; font-weight: 500; } +.delivery-figure { padding: 28px; border: 1px solid var(--border); border-radius: 12px; background: var(--surface); } +.delivery-path { display: grid; grid-template-columns: 1fr 40px 1.2fr 40px 1.2fr; gap: 12px; align-items: center; } +.delivery-path > div { min-width: 0; } +.delivery-path span:not(.path-arrow) { display: block; font-size: 12px; color: var(--body); margin-bottom: 10px; } +.delivery-path b { font-size: 16px; font-weight: 500; } +.delivery-path p { font-size: 13px; margin: 4px 0 0; } +.path-arrow { color: #005cc5; text-align: center; } +.limits-table { min-width: 640px; } +.limits-table th:first-child { width: 150px; } +.limits-table th:nth-child(2) { width: 48%; } +.limits-table tbody th { background: transparent; } +.next-steps { list-style: none; padding: 0; margin: 36px 0 0; } +.next-steps li { display: grid; grid-template-columns: 60px 1fr; gap: 24px; padding: 24px 0; border-bottom: 1px solid var(--border); } +.next-steps li > span { font: 16px/1.7 var(--font-mono); color: var(--body); } +.next-steps h3 { margin-bottom: 6px; } +.next-steps p { font-size: 15px; max-width: 800px; margin-bottom: 0; } +.research-page .source-list { list-style: decimal-leading-zero; padding-left: 32px; margin-top: 32px; } +.source-list li { padding: 14px 0 14px 12px; border-bottom: 1px solid var(--border); } +.source-list li::marker { font: 12px var(--font-mono); color: var(--body); } +.source-list a { display: block; } +.source-list b { font-size: 15px; font-weight: 500; } +.source-list span { display: block; font-size: 13px; color: var(--body); margin-top: 4px; } +.related { margin-top: 48px; } +.related > a { display: block; padding: 16px 0; color: #005cc5; font-size: 15px; } +.related > a span { display: block; color: var(--body); font-size: 13px; } +.related .caption { margin-top: 16px; } +.research-footer { display: flex; justify-content: space-between; align-items: center; gap: 24px; min-height: 100px; color: var(--body); font-size: 12px; } +.research-footer a { min-height: 44px; display: inline-flex; align-items: center; } +@media (max-width: 1000px) { + .research-page .hero { gap: 32px; } + .research-nav nav { gap: 20px; } + .nav-edition { display: none; } + .execution-flow { grid-template-columns: 1fr 1fr; } + .execution-flow li:nth-child(2) { border-right: 0; } + .execution-flow li:nth-child(-n+2) { border-bottom: 1px solid var(--border); } + .arm-list article { grid-template-columns: 110px 1fr; gap: 16px; } + .arm-status { grid-column: 2; text-align: left; } +} +@media (max-width: 760px) { + .research-page .shell { width: calc(100% - 40px); } + .research-page .hero { grid-template-columns: 1fr; gap: 32px; padding-top: 36px; } + .research-page h1 { font-size: 40px; } + .research-page h2 { font-size: 26px; } + .research-page .deck { font-size: 17px; } + .research-nav { flex-wrap: wrap; gap: 0; padding-top: 8px; } + .research-nav .brand { min-height: 40px; font-size: 16px; } + .research-nav nav { width: 100%; gap: 24px; } + .result-figure { padding: 24px; } + .result-figure h2 { font-size: 21px; } + .scope-note { padding: 16px 20px; } + .study-section { padding: 48px 0; scroll-margin-top: 116px; } + .section-lead > p:not(.eyebrow) { font-size: 16px; } + .reading-pair { grid-template-columns: 1fr; gap: 24px; } + .arm-list article { grid-template-columns: 1fr; gap: 8px; } + .arm-status { grid-column: 1; } + .execution-flow li { padding: 20px 16px; } + .decision-rows > div { grid-template-columns: 1fr; gap: 8px; } + .delivery-path { grid-template-columns: 1fr; gap: 16px; } + .path-arrow { transform: rotate(90deg); text-align: center; } + .delivery-path > div { text-align: center; } + .delivery-figure { padding: 24px; } + .next-steps li { grid-template-columns: 32px 1fr; gap: 16px; } + .research-footer { align-items: flex-start; flex-direction: column; gap: 4px; padding-block: 24px; } +} +@media print { + .research-header, .skip-link, .text-link, .research-footer > a { display: none; } + .research-page .shell { width: 100%; } + .research-page .hero { padding-top: 0; } + .study-section { padding-block: 32px; } + .table-scroll { overflow: visible; } + .limits-table { min-width: 0; } + .result-figure, .delivery-figure, .execution-flow { break-inside: avoid; } +} diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html new file mode 100644 index 000000000..86dc536ac --- /dev/null +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -0,0 +1,195 @@ + + + + + + + LoopX 用在哪里:复杂任务、开放探索与持续交付 + + + + + + + + + + + + + +
+
+

场景与实践 · 长程 Agent

+

LoopX 用在哪里?
复杂任务、开放探索与持续交付

+

有些工作需要把一件事做完,有些需要找到下一条值得走的路,还有些需要长期对结果负责。它们需要不同的领域能力,也需要跨会话保留下来的目标、证据与约束。

+
Ruiteng Huang三个场景 · 一套长程控制面
+
+
+ +
+
+

先看工作,再看抽象

+
+
01 / COMPLETION

验收明确的复杂任务

重构、迁移、协议实现。终点比较明确,执行路线仍有不确定性。

保留:验收缺口、版本、修复证据
关注:完成率与成本
+
02 / DISCOVERY

开放探索类任务

调研、算法实验、系统优化。下一条路线由新证据不断改变。

保留:假设、反证、候选方向
关注:有效发现与验证
+
03 / DELIVERY

持续交付的数字员工

接收 issue、修复、跟进评审,处理新反馈,并对结果持续负责。

保留:职责、外部状态、等待条件
关注:交付质量与人的投入
+
+

这三类可以嵌套。一个长期维护仓库的 Agent,可能先探索性能问题,再完成一项验收明确的修复,最后持续跟进 PR。前两类主要描述工作如何求解,第三类还增加了持续的职责与到达的新任务。

+

真正延续下来的,应当是工作及其判断依据。

+
+ +
+

01 / 把复杂任务做完

+

“实现一个符合规范的解码器”听起来很确定。但可见样例通过之后,仍可能遗漏边界输入、内存安全和兼容性;重开一个会话,又可能丢掉此前确认的失败条件。

+
+
定义验收规范、基线、预算
实现并验证产物绑定具体版本
找出缺口测试之外还缺什么
继续或停止补缺口,或明确收口
+
+

LoopX 的作用,是让目标和验收缺口跨轮次保持连续,让新证据进入下一步决策。模型负责实现与判断,领域验证器负责检验结果;控制面记录哪些结果被接受、哪些工作仍未完成。

+

适合它的任务,通常会经历多轮验证、等待或交接。短小、一次会话就能验收的工作,现有 Agent 可能已经足够,新增控制面需要证明自身开销值得。

+
+ +
+

Benchmark:三项研究,三个值得追查的信号

+

三项研究分别观察续跑、交付和验证行为。任务、模型设置与统计口径不同,应分别理解;现有结果还不足以概括 LoopX 的普遍增益。

+ +
+

SWE-Marathon:续跑要补上真实缺口

+

GPT-5.6 Sol / high · 15 个匹配任务 · 3 个保留模式,每格运行 1 次 · 超时系数 0.3

+

观察:原生 Goal 完成 4/15,Heartbeat 完成 5/15;两组总成本从 $533 增至 $830,约多 56%。

+

Insight:zstd 个案中,续跑补充了可见测试之外的验收;excel 个案多次续跑仍未完成。值得追查的是下一轮补了什么缺口,以及这份收益是否值得新增成本。

+

范围:每格只有一次运行,多项机制同时变化。已有个案收益,也有无效续跑;尚不能确认稳定增益或成本优势。

+ +
+ +
+

DeepSWE × Sol:完成差异要拆开看

+

脚本默认 GPT-5.6 Sol / xhigh · 113 个任务 · 历史 best-valid 汇总

+

观察:裸 Codex、原生 Goal、Heartbeat 分别完成 54、60、70 题。Heartbeat 比 Goal 多 10 题,但默认时间窗口也更长。

+

Insight:源码把“继续执行”与“有效交付”分开处理:worktree 中有代码,收集器仍可能拿到空补丁。应分别检查恢复、补丁交付与独立验收,再追查它们对完成差异的贡献。

+

范围:每任务取最佳有效结果,非单次通过率;预算不同,尝试次数和成本未完整披露。合入未复验成绩,不能视为同预算净增益。

+ +
+ +
+

DeepSWE × V4 Flash max:让反例改变实现

+

DeepSeek V4 Flash / max + Codex · 冻结 113 题 · 按相同 hint 条件比较 Goal 与 LoopX

+

观察:LoopX 的按题平均 feature 覆盖高约 2 个百分点。两组均有 hint 的事后长时切片中,完成数从 12/29 增至 14/29,累计耗时少 16.4%。

+

Insight:精选案例里的差别在于,失败反例是否被保留,外部契约能否推翻实现假设。验证的价值要体现在修复与复验上,而不只是增加检查次数。

+

范围:覆盖不等于成功率;长时切片为事后分组,耗时包含成功与失败运行。局部发现不能外推为总体提升或等质量加速。

+ +
+ +

当前的信号是:有效续跑、可靠交付和反例驱动的修复值得继续研究。

+

三项研究提供了局部正向观察,也暴露了成本、失败与归因问题。下一步需要在匹配预算和重复实验中,判断哪些收益能够稳定复现。

+

SWE-Marathon 与 DeepSWE × Sol 已撤回的 SSH Goal、Codex CLI 成绩和结论均不用于这里的比较。

+
+ +
+

02 / 在开放探索中找路

+

“找到一个更好的算法方案”没有预先写好的完整任务链。真正的进展可能是验证一个假设,也可能是排除一条看似有希望的路线。只保留成功结论,下一轮就容易重试已被否定的想法。

+
+
提出问题约束与待验证假设
尝试多条路线隔离实验,控制预算
记录正反证据支持、反驳、引出新问题
更新探索方向继续、合并、淘汰或暂停
+
+

Explore 已有可选的证据图和有界分支规划:节点表示问题与发现,关系表达 supports、refutes、leads_to。规划建议需要经过正常执行边界,图本身不启动 Worker,也不授予花费权限。

+

Auto Research 把这一思路组织成研究工作:选题、提出假设、执行实验、独立评价、形成报告。现有协议与命令提供了基础,公开 showcase 文档仍包含蓝图和待验证项;实际效果需要在具体任务中证明。

+ +

RSI 放在这里,但收紧含义

+

如果改进对象变成 Agent 自己的工具、策略或 Harness,就进入自改进研究。能修改自身代码,只是能够提出候选版本;持续变强还需要独立评价、跨任务泛化、回归检查和回滚。

+

把历史经验用于下一次决策、自动生成实验、修改 Harness,是不同层次。这里可以讨论通往递归自改进(RSI)的实验路径,不能把“自动迭代”直接说成已经实现开放式自我提升。

+
+ +
+

03 / 持续承担交付责任

+

一个 PR / issue fix 数字员工,工作的起点是“这件事是否值得修”,交付后还要面对 CI、评审意见、分支变化和合并结果。任务不断到来,外部事实也不断变化。

+

生成补丁之后,责任还没有结束。

+

下面用一个构造场景说明:修复已提交,CI 和评审仍在进行。点击外部状态,看下一步如何改变。

+ +
+
外部事实
GitHub 返回当前修订的 checks pending。
+
领域状态
记录 PR、修订与检查状态;相同观察无需制造新进展。
+
能力提出下一步
保留监控和恢复条件,等待结果。
+
控制面检查
按授权、预算和执行资格准入;等待范围之外的工作仍可推进。
+
对人的意义
没有重要变化时保持安静,避免把轮询变成催促。
+
+ +

交互图为设计说明,不连接真实仓库,也不执行任何 GitHub 操作。

+

“数字员工”在这里意味着有持续职责、边界、记忆与反馈闭环。是否值得使用,要看被接受的修复、重开与回归、处理时延、每次交付成本,以及需要人反复盯守多少次。PR 数量和在线时长只是过程信号。

+
+ +
+

领域越丰富,分工越要清楚

+
+
Kernel
控制权与通用生命周期
哪个目标与工作项有效,谁可执行,哪些授权和预算适用,怎样交接、等待与提交状态。
+
Domain State
领域连续性
这个 issue 是否可修,PR 对应哪个修订,CI 和评审到了哪一步,哪些结论仍然有效。
+
Capability
结果契约与判断
把领域事实翻译成可验证的下一步:修复、等待、交给人判断或收口,并定义这个场景怎样验收。
+
Provider / Runtime
外部读写与实际执行
调用代码仓库、测试和工具,执行已授权动作并回读结果。GitHub 仍拥有代码、CI 和评审的外部事实。
+
+

以 checks failing 为例:领域能力可以提出一个修 CI 的后继任务;它不能因“需要修”就自行得到写仓库或合并权限。换成实验任务,CI 状态变成指标和留出集结论,通用的认领、授权、恢复规则仍应保持一致。

+

这也是 Capability 的价值:把一个场景中反复出现的判断与结果契约沉淀下来。它并不要求把所有业务阶段都写进 Kernel。

+
+ +
+

演进:每一步增加一种可验证的能力

+
    +
  1. 单次执行 → 可恢复的长程任务换会话、进程中断后,仍能恢复目标、证据与下一步。先证明一件事能可靠做完。
  2. +
  3. 通用任务 → 垂域交付与探索用 issue-fix 和研究任务承接真实工作,保留领域状态,定义有价值的结果。
  4. +
  5. 单个 Agent → 小团队协作先让 2–3 个 Worker 完成两轮依赖交付:接收产物、独立验收、吸收纠偏,并把结果送回。
  6. +
  7. 本地协作 → 跨 Host 的持续工作验证共享状态、权限撤回、旧执行者隔离、网络故障和共同预算;注册成功不等于协作成立。
  8. +
  9. 可持续运行 → 可衡量地改进用结果反馈改善选择、记忆与策略;用对照实验、留出验证和回滚证明收益。规模扩张另行验收。
  10. +
+

以上是建设顺序与验证目标,不是完成清单。现有基础、局部实现、蓝图和端到端资格应分别理解。

+

我希望交给 Agent 的,逐渐从一条指令变成一份可以持续履行的工作约定。

+

这份约定要让人看得懂、改得动,让 Agent 接得住,也让失败之后仍有证据可查。

+
+ +
+

公开依据与延伸阅读

+

技术内容仅据公开仓库材料整理。源码与数据引用固定到本次读取修订;历史实验使用其自身版本。本页没有新增实验结果。

+
    +
  1. SWE-Marathon:持续自我验证;设置、正反个案与局限;公开聚合数据。贡献者与案例来源见研究原文。
  2. +
  3. DeepSWE × Sol:从继续执行到有效交付。独立研究简报,含 113 任务历史结果、机制图与固定修订的一手来源;研究归档贡献:@gwh6669999,#4502。
  4. +
  5. DeepSWE × V4 Flash max:从提示到行为;披露范围;固定修订的图表、案例与指标。
  6. +
  7. Explore:证据图、规划与权限边界。
  8. +
  9. Auto Research:公开蓝图与命令路径。
  10. +
  11. PR / Issue Fix:State Kernel 与领域状态如何协同。
  12. +
  13. 整体路线图:产品目标与独立验收里程碑。
  14. +
  15. 从一次性 Agent 到长程控制面;LoopX:长程 Agent 的原生 Kanban。
  16. +
+
+
+
+
+ + + + diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/presentation.js b/apps/presentation/site/public/blog/zh/application-scenarios/presentation.js new file mode 100644 index 000000000..34dde2bbd --- /dev/null +++ b/apps/presentation/site/public/blog/zh/application-scenarios/presentation.js @@ -0,0 +1,28 @@ +const presentation = document.querySelector('#presentation'); +presentation.hidden = false; +document.querySelector('.stage-buttons').hidden = false; +presentation.addEventListener('click', () => { + const enabled = document.body.classList.toggle('presenting'); + presentation.setAttribute('aria-pressed', String(enabled)); + presentation.textContent = enabled ? '阅读模式' : '演讲模式'; +}); +const states = { + pending: ['GitHub 返回当前修订的 checks pending。','记录 PR、修订与检查状态;相同观察无需制造新进展。','保留监控和恢复条件,等待结果。','按授权、预算和执行资格准入;等待范围之外的工作仍可推进。','没有重要变化时保持安静,避免把轮询变成催促。'], + failed: ['当前修订的 CI 出现失败。','区分代码回归、环境故障和未知原因,保留证据。','提出一个有范围、有验收方式的修复后继。','重新检查执行者、写入范围与预算;获得资格后才执行。','把失败转成可接手的工作,而不只发一条报错消息。'], + changed: ['PR 从修订 A 更新为修订 B。','原评审和测试绑定 A,不能直接证明 B 合格。','重新确认变更与验证范围,需要时重新评审。','保留版本关联和当前权限;旧回执不能冒充新结果。','让“通过”始终对应清楚的交付物。'], + merged: ['GitHub 确认 PR 已合并。','记录终态和结果;合并不自动等于整个目标已完成。','检查目标验收:补集成验证、接续任务,或给出无后继理由。','结算当前监控与后继;发布等其他动作仍需相应授权。','每项工作有明确去向,完成后不凭惯性继续找活。'] +}; +document.querySelectorAll('[data-state]').forEach(button => button.addEventListener('click', () => { + document.querySelectorAll('[data-state]').forEach(b => b.setAttribute('aria-pressed',String(b === button))); + ['fact','domain','proposal','kernel','meaning'].forEach((id,i) => document.getElementById(id).textContent = states[button.dataset.state][i]); +})); +document.addEventListener('keydown', event => { + if (!document.body.classList.contains('presenting') || event.altKey || event.ctrlKey || event.metaKey) return; + if (event.key === 'Escape') { presentation.click(); return; } + if (event.target.closest('input,textarea,select,[contenteditable=true]')) return; + if (!['ArrowRight','ArrowLeft'].includes(event.key)) return; + const sections = [...document.querySelectorAll('article > section')]; + let index = sections.findIndex(s => s.getBoundingClientRect().bottom > 100); + index = Math.max(0, Math.min(sections.length - 1, index + (event.key === 'ArrowRight' ? 1 : -1))); + event.preventDefault(); sections[index].scrollIntoView({behavior: 'instant'}); history.replaceState(null,'','#'+sections[index].id); +}); diff --git a/apps/presentation/site/public/blog/zh/index.html b/apps/presentation/site/public/blog/zh/index.html index 757a4c43d..46b8b6dfd 100644 --- a/apps/presentation/site/public/blog/zh/index.html +++ b/apps/presentation/site/public/blog/zh/index.html @@ -32,6 +32,10 @@

LoopX · 博客

长程工作的工程实践。

关于长程 Agent 的设计、工程与实践。

+
+

场景 · 实践与研究

2026 年 9 月 18 日

+

LoopX 用在哪里:复杂任务、开放探索与持续交付

沿三个工作场景理解目标、证据与交付责任,并列解读 SWE-Marathon、DeepSWE × Sol 与 V4 Flash max 的研究信号。

阅读全文
+

Agent-native Kanban · 长程协作

2026 年 9 月 15 日

LoopX:长程 Agent 的原生 Kanban

用四张图理解 Goal、Todo、认领、租约、门禁、证据与共享权威状态,再看它们如何支持长程协作。

阅读全文
diff --git a/apps/presentation/site/src/App.tsx b/apps/presentation/site/src/App.tsx index 0ddd8b213..cde5a94e9 100644 --- a/apps/presentation/site/src/App.tsx +++ b/apps/presentation/site/src/App.tsx @@ -139,8 +139,9 @@ const content = { body: "Try the product guide, read the research, and browse the full case collection.", cards: [ ["Product demo", "Personal Workspace", "A shared view of Goals, Tasks, Chat, and outputs. Watch the demo and start your own local workspace.", "Watch demo & read guide"], - ["Research & evaluation", "SWE-Marathon", "Compare five execution modes across 15 matched tasks, with results, costs, and study limitations.", "Read the study"], + ["Research & evaluation", "SWE-Marathon", "Compare three retained execution modes across 15 matched tasks, with results, costs, and study limitations.", "Read the study"], ["Research · Chinese", "DeepSWE behavior discoveries", "Explore how domain hints affect implementation and validation in individual cases.", "Read the behavior analysis"], + ["Research · Chinese", "DeepSWE × Sol", "Read the historical 113-task results, continuation mechanisms, and patch-delivery checks.", "Read the research brief"], ["Cases", "All LoopX showcases", "Browse public cases, interactive walkthroughs, and their evidence boundaries.", "Browse all cases"], ], developer: "Projection developer tools", @@ -252,8 +253,9 @@ const content = { body: "从产品演示到评测结果,找到你想深入了解的内容。", cards: [ ["产品演示", "Personal Workspace", "在一个工作区查看目标、任务、对话与产出。观看演示,开始使用本地工作区。", "观看演示与使用指南"], - ["研究与评测", "SWE-Marathon", "查看五种执行模式在 15 个匹配任务上的结果、成本与研究局限。", "阅读研究简报"], + ["研究与评测", "SWE-Marathon", "查看三种保留模式在 15 个匹配任务上的结果、成本与研究局限。", "阅读研究简报"], ["研究与评测", "DeepSWE 行为分析", "从具体案例观察领域提示如何影响实现选择与验证行为。", "阅读行为分析"], + ["研究与评测", "DeepSWE × Sol", "解读 113 任务的历史结果、持续执行机制,以及补丁如何进入最终验收。", "阅读研究简报"], ["案例", "完整案例目录", "浏览公开案例、交互式讲解及其证据边界。", "浏览全部案例"], ], developer: "开发者投影工具", @@ -1010,6 +1012,7 @@ export function App() { "docs/guides/personal-workspace-user-guide/", `benchmarks/swe-marathon/${language === "zh" ? "?lang=zh" : ""}`, "benchmarks/deepswe/behavior-discovery/", + "benchmarks/deepswe-sol/", `docs/showcases/index${language === "en" ? ".en" : ""}.html`, ]; return ( diff --git a/apps/presentation/site/vite.config.ts b/apps/presentation/site/vite.config.ts index a7970a634..3cd15a841 100644 --- a/apps/presentation/site/vite.config.ts +++ b/apps/presentation/site/vite.config.ts @@ -6,12 +6,12 @@ import { defineConfig } from "vite"; export default defineConfig({ base: process.env.LOOPX_SITE_BASE ?? "/", plugins: [react(), { - name: "blog-directory-index", + name: "public-directory-index", configureServer(server) { // Match static-host directory URLs before Vite's SPA fallback. server.middlewares.use((request, _response, next) => { const url = new URL(request.url ?? "/", "http://localhost"); - if (url.pathname.startsWith("/blog/") && url.pathname.endsWith("/")) { + if ((url.pathname.startsWith("/blog/") || url.pathname.startsWith("/benchmarks/")) && url.pathname.endsWith("/")) { const index = new URL(`./public${url.pathname}index.html`, import.meta.url); if (existsSync(fileURLToPath(index))) request.url = `${url.pathname}index.html${url.search}`; } diff --git a/examples/dashboard-frontstage-browser-smoke.mjs b/examples/dashboard-frontstage-browser-smoke.mjs index e56aeb4e6..9a2cf674a 100644 --- a/examples/dashboard-frontstage-browser-smoke.mjs +++ b/examples/dashboard-frontstage-browser-smoke.mjs @@ -235,7 +235,7 @@ try { await page.locator("#explore").waitFor(); assert.equal(await page.locator("html").getAttribute("lang"), lang === "zh" ? "zh-CN" : "en"); assert.equal(await page.locator('a[href*="deprecated"], a[href*="frontstage/"]').count(), 0); - const expected = ["docs/guides/personal-workspace-user-guide/", `benchmarks/swe-marathon/${lang === "zh" ? "?lang=zh" : ""}`, "benchmarks/deepswe/behavior-discovery/", `docs/showcases/index${lang === "en" ? ".en" : ""}.html`]; + const expected = ["docs/guides/personal-workspace-user-guide/", `benchmarks/swe-marathon/${lang === "zh" ? "?lang=zh" : ""}`, "benchmarks/deepswe/behavior-discovery/", "benchmarks/deepswe-sol/", `docs/showcases/index${lang === "en" ? ".en" : ""}.html`]; assert.deepEqual(await page.locator("#explore .resource-card").evaluateAll((links) => links.map((a) => a.getAttribute("href"))), expected.map((path) => `/loopx/${path}`)); if (width === 390) { await page.getByRole("button", { name: "Open navigation" }).click(); @@ -260,7 +260,7 @@ try { if (process.env.LOOPX_PUBLIC_SITE_DIR) { // Check the assembled publication, including the separately built books. const checked = new Set(); - for (const path of ["", "?lang=zh", "benchmarks/swe-marathon/", "benchmarks/deepswe/behavior-discovery/", "docs/showcases/index.html", "docs/showcases/index.en.html", "docs/guides/personal-workspace-user-guide/", "docs/book/", "docs/book/en/", "blog/", "blog/zh/"]) { + for (const path of ["", "?lang=zh", "benchmarks/swe-marathon/", "benchmarks/deepswe/behavior-discovery/", "benchmarks/deepswe-sol/", "docs/showcases/index.html", "docs/showcases/index.en.html", "docs/guides/personal-workspace-user-guide/", "docs/book/", "docs/book/en/", "blog/", "blog/zh/"]) { await page.goto(`${publicOrigin}/loopx/${path}`); await page.locator("h1").first().waitFor(); const links = await page.locator("a[href]").evaluateAll((links) => links.map((link) => link.href)); diff --git a/examples/dashboard-frontstage-design-baseline-smoke.mjs b/examples/dashboard-frontstage-design-baseline-smoke.mjs index 2c35a1f3c..03a2a1eda 100644 --- a/examples/dashboard-frontstage-design-baseline-smoke.mjs +++ b/examples/dashboard-frontstage-design-baseline-smoke.mjs @@ -6,7 +6,7 @@ const read = (path) => readFileSync(fileURLToPath(new URL(`../${path}`, import.m const home = read("apps/presentation/site/src/App.tsx"); const styles = read("apps/presentation/site/src/styles.css"); assert.doesNotMatch(home, /href=.[^\n]*(?:deprecated|frontstage\/)/, "homepage must not promote retired surfaces"); -for (const destination of ["docs/guides/personal-workspace-user-guide/", "benchmarks/swe-marathon/", "benchmarks/deepswe/behavior-discovery/", "docs/showcases/index", "developers/projections/"]) { +for (const destination of ["docs/guides/personal-workspace-user-guide/", "benchmarks/swe-marathon/", "benchmarks/deepswe/behavior-discovery/", "benchmarks/deepswe-sol/", "docs/showcases/index", "developers/projections/"]) { assert.ok(home.includes(destination), `missing public destination: ${destination}`); } assert.ok(styles.includes("prefers-reduced-motion")); diff --git a/examples/export-frontstage-share-bundle.mjs b/examples/export-frontstage-share-bundle.mjs index fe18b98a9..359717abb 100644 --- a/examples/export-frontstage-share-bundle.mjs +++ b/examples/export-frontstage-share-bundle.mjs @@ -17,6 +17,7 @@ const showcaseCatalogPath = "docs/showcases/showcase-catalog.json"; const projectionFixturePath = "examples/goal-channel-frontstage-fixture.py"; const installerScriptPath = "scripts/install-from-github.sh"; const deepSweBehaviorArticlePath = "benchmark/deepswe/behavior-discovery/index.html"; +const deepSweSolArticlePath = "apps/presentation/site/public/benchmarks/deepswe-sol/index.html"; const homepageEvidenceAssets = [ "docs/assets/long-running-loop-openviking-trajectory.png", "docs/assets/long-running-loop-ml-experiment-trajectory.png", @@ -322,6 +323,7 @@ ${previewBlock} - Homepage source: \`apps/presentation/site\`. - SWE-Marathon research brief: \`${sweMarathonBriefUrl}\`, built from the pinned public-safe aggregate and case-insight projection under \`benchmark/swe-marathon/\`. - DeepSWE behavior discoveries: \`${deepSweBehaviorArticleUrl}\`, copied byte-for-byte from the reviewed standalone article at \`${deepSweBehaviorArticlePath}\`. +- DeepSWE × Sol research brief: \`${base}benchmarks/deepswe-sol/\`, a static historical-study interpretation from \`${deepSweSolArticlePath}\`. - Homepage evidence assets: ${homepageEvidenceAssets.map((path) => `\`${path}\``).join(", ")}. - Personal Workspace demo and guide: docs/guides/personal-workspace-user-guide/. - Legacy Frontstage URLs redirect to the case directory without loading a dashboard or forwarding status parameters. @@ -344,6 +346,7 @@ async function writeManifest(outDir, base, interactivePages) { homepage_entry: "site/index.html", swe_marathon_brief_entry: "site/benchmarks/swe-marathon/index.html", deepswe_behavior_article_entry: "site/benchmarks/deepswe/behavior-discovery/index.html", + deepswe_sol_article_entry: "site/benchmarks/deepswe-sol/index.html", installer_entry: "site/install.sh", frontstage_entry: "site/frontstage/index.html", frontstage_redirect: "docs/showcases/index.en.html", @@ -352,6 +355,7 @@ async function writeManifest(outDir, base, interactivePages) { public_homepage: "apps/presentation/site", swe_marathon_brief: "benchmark/swe-marathon", deepswe_behavior_article: deepSweBehaviorArticlePath, + deepswe_sol_article: deepSweSolArticlePath, installer_script: installerScriptPath, homepage_evidence_assets: homepageEvidenceAssets, primary_public_story: showcaseCatalogPath, diff --git a/examples/frontstage-share-bundle-smoke.mjs b/examples/frontstage-share-bundle-smoke.mjs index 6c8bd5b89..58dd6a04f 100644 --- a/examples/frontstage-share-bundle-smoke.mjs +++ b/examples/frontstage-share-bundle-smoke.mjs @@ -118,21 +118,42 @@ assertExists(resolve(siteDir, "index.html")); assertExists(resolve(siteDir, "frontstage/index.html")); assertExists(resolve(siteDir, "benchmarks/swe-marathon/index.html")); assertExists(resolve(siteDir, "benchmarks/deepswe/behavior-discovery/index.html")); +// Static research articles must remain readable and navigable in the shipped +// bundle without falling back to the homepage SPA. +for (const route of ["benchmarks/deepswe-sol/"]) { + const pagePath = resolve(siteDir, route, "index.html"); + const html = await readFile(pagePath, "utf8"); + if (!/') || html.includes("`) || !html.includes("

") || html.includes("`) || !html.includes("

") || (!interactive && html.includes("]*>[\s\S]*?<\/script>/g)].map((match) => match[0]); + if (scripts.length !== 1 || scripts[0] !== '') { + throw new Error("Application article must keep enhancement in its local deferred script"); + } + assertExists(resolve(dirname(pagePath), "presentation.js")); + } // A single-language article must not advertise a nonexistent translation. - const alternates = article === "agent-facing-kanban/" ? [] : ["en", "zh-CN", "x-default"]; + const alternates = ["agent-facing-kanban/", "application-scenarios/"].includes(article) ? [] : ["en", "zh-CN", "x-default"]; for (const hreflang of alternates) { if (!html.includes(`hreflang="${hreflang}"`)) throw new Error(`Missing Blog language alternate: ${hreflang}`); }