Skip to content

Commit 4bf0048

Browse files
committed
fix(site): complete LHTB Pages publication
Signed-off-by: shangzh0 <2586756592@qq.com>
1 parent 45abe1d commit 4bf0048

8 files changed

Lines changed: 64 additions & 5 deletions

File tree

‎README.md‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -229,12 +229,16 @@ creator dogfooding, reproducible demos, and explicit evidence-strength labels.
229229
- **[SWE-Marathon](https://huangruiteng.github.io/loopx/benchmarks/swe-marathon/):**
230230
Five execution modes on 15 matched tasks compare self-verification, scores,
231231
and cost. More self-verification did not consistently yield higher scores.
232+
- **[LHTB × LoopX](https://huangruiteng.github.io/loopx/benchmarks/lhtb/):**
233+
Five execution mechanisms on 46 long-horizon terminal tasks compare durable
234+
state, bounded Todos, replanning, and fresh executor sessions.
232235
- **[DeepSWE behavior analysis](https://huangruiteng.github.io/loopx/benchmarks/deepswe/behavior-discovery/)** (Chinese):
233236
Selected cases examine how domain hints relate to requirement retention and
234237
verification choices, offering mechanism hypotheses for further testing.
235238

236-
SWE-Marathon has one trial per task and mode; DeepSWE uses selected cases and
237-
post-hoc analysis. Neither establishes a general performance gain.
239+
SWE-Marathon and LHTB have one effective trial per task and mode; DeepSWE uses
240+
selected cases and post-hoc analysis. None establishes a general performance
241+
gain.
238242

239243
More inspectable surfaces:
240244

‎README.zh-CN.md‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -200,11 +200,14 @@ creator dogfooding、reproducible demo 和证据强度标签见
200200
- **[SWE-Marathon](https://huangruiteng.github.io/loopx/benchmarks/swe-marathon/?lang=zh)**:
201201
在相同的 15 个任务上对照 5 种执行模式,比较自验证行为、得分与成本。
202202
更多自验证并未稳定转化为更高得分。
203+
- **[LHTB × LoopX](https://huangruiteng.github.io/loopx/benchmarks/lhtb/?lang=zh)**:
204+
在 46 个长程终端任务上对比 5 种执行机制,研究持久状态、Todo、replan
205+
与 fresh executor session。
203206
- **[DeepSWE 行为分析](https://huangruiteng.github.io/loopx/benchmarks/deepswe/behavior-discovery/)**:
204207
通过精选案例观察领域提示、需求保留与验证选择之间的关系,提出有待复验的机制假设。
205208

206-
SWE-Marathon 每个任务、每种模式仅运行一次;DeepSWE 包含精选案例与事后分析。
207-
目前两者均不足以证明普遍的性能提升。
209+
SWE-Marathon 与 LHTB 每个任务、每种模式仅保留一条有效轨迹;DeepSWE 包含精选案例与
210+
事后分析。目前三者均不足以证明普遍的性能提升。
208211

209212
更多可检查入口:
210213

‎apps/presentation/site/src/App.tsx‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -140,6 +140,7 @@ const content = {
140140
cards: [
141141
["Product demo", "Personal Workspace", "A shared view of Goals, Tasks, Chat, and outputs. Watch the demo and start your own local workspace.", "Watch demo & read guide"],
142142
["Research & evaluation", "SWE-Marathon", "Compare three retained execution modes across 15 matched tasks, with results, costs, and study limitations.", "Read the study"],
143+
["Research & evaluation", "LHTB × LoopX", "Compare five execution mechanisms across 46 long-horizon terminal tasks and inspect where recoverable state helps.", "Read the study"],
143144
["Research · Chinese", "DeepSWE behavior discoveries", "Explore how domain hints affect implementation and validation in individual cases.", "Read the behavior analysis"],
144145
["Research · Chinese", "DeepSWE × Sol", "Read the historical 113-task results, continuation mechanisms, and patch-delivery checks.", "Read the research brief"],
145146
["Cases", "All LoopX showcases", "Browse public cases, interactive walkthroughs, and their evidence boundaries.", "Browse all cases"],
@@ -254,6 +255,7 @@ const content = {
254255
cards: [
255256
["产品演示", "Personal Workspace", "在一个工作区查看目标、任务、对话与产出。观看演示,开始使用本地工作区。", "观看演示与使用指南"],
256257
["研究与评测", "SWE-Marathon", "查看三种保留模式在 15 个匹配任务上的结果、成本与研究局限。", "阅读研究简报"],
258+
["研究与评测", "LHTB × LoopX", "对比五种执行机制在 46 个长程终端任务上的结果,观察可恢复状态在什么场景有效。", "阅读研究简报"],
257259
["研究与评测", "DeepSWE 行为分析", "从具体案例观察领域提示如何影响实现选择与验证行为。", "阅读行为分析"],
258260
["研究与评测", "DeepSWE × Sol", "解读 113 任务的历史结果、持续执行机制,以及补丁如何进入最终验收。", "阅读研究简报"],
259261
["案例", "完整案例目录", "浏览公开案例、交互式讲解及其证据边界。", "浏览全部案例"],
@@ -1011,6 +1013,7 @@ export function App() {
10111013
const paths = [
10121014
"docs/guides/personal-workspace-user-guide/",
10131015
`benchmarks/swe-marathon/${language === "zh" ? "?lang=zh" : ""}`,
1016+
`benchmarks/lhtb/${language === "zh" ? "?lang=zh" : ""}`,
10141017
"benchmarks/deepswe/behavior-discovery/",
10151018
"benchmarks/deepswe-sol/",
10161019
`docs/showcases/index${language === "en" ? ".en" : ""}.html`,

‎benchmark/README.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -48,6 +48,9 @@ by itself establish a C2 uplift claim.
4848

4949
- [`swe-marathon/README.md`](swe-marathon/README.md) links the published
5050
[SWE-Marathon research brief](https://huangruiteng.github.io/loopx/benchmarks/swe-marathon/).
51+
- [`LHTB/studies/five-arm-gpt56sol-max/README.md`](LHTB/studies/five-arm-gpt56sol-max/README.md)
52+
documents the public-safe aggregate behind the bilingual
53+
[LHTB research brief](https://huangruiteng.github.io/loopx/benchmarks/lhtb/).
5154
- [`deepswe/behavior-discovery/README.md`](deepswe/behavior-discovery/README.md)
5255
links the standalone
5356
[DeepSWE behavior-discovery article](https://huangruiteng.github.io/loopx/benchmarks/deepswe/behavior-discovery/).

‎benchmark/tests/test_publication_scope.py‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -80,3 +80,31 @@ def test_published_data_and_bilingual_tables_share_scope(exporters):
8080
for localized in copy.values():
8181
assert {row[0] for row in localized["armRows"]} == allowed
8282
assert set(localized["executiveReads"]) == allowed
83+
84+
85+
def test_lhtb_published_data_and_bilingual_copy_share_scope():
86+
study = STUDY.parent / "LHTB" / "studies" / "five-arm-gpt56sol-max"
87+
data = json.loads((study / "data.json").read_text())
88+
arms = {
89+
"plain",
90+
"native_goal",
91+
"ssh_goal",
92+
"legacy_heartbeat",
93+
"new_heartbeat",
94+
}
95+
assert set(data["arms"]) == arms
96+
assert len(data["tasks"]) == len({row["task"] for row in data["tasks"]}) == 46
97+
assert all(set(row) == arms | {"task"} for row in data["tasks"])
98+
99+
for arm, summary in data["arms"].items():
100+
rewards = [row[arm] for row in data["tasks"]]
101+
assert summary["mean_reward"] == pytest.approx(sum(rewards) / 46, rel=0, abs=1e-12)
102+
assert summary["pass_095"] == sum(reward >= 0.95 for reward in rewards)
103+
104+
site = STUDY.parents[1] / "apps/presentation/site/src"
105+
localized_copy = json.loads((site / "lhtb-copy.json").read_text())
106+
assert set(localized_copy) == {"en", "zh"}
107+
for localized in localized_copy.values():
108+
assert set(localized["armLabels"]) == arms
109+
assert set(localized["armKinds"]) == arms
110+
assert {row[0] for row in localized["mechanismRows"]} == arms

‎examples/dashboard-frontstage-design-baseline-smoke.mjs‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ const read = (path) => readFileSync(fileURLToPath(new URL(`../${path}`, import.m
66
const home = read("apps/presentation/site/src/App.tsx");
77
const styles = read("apps/presentation/site/src/styles.css");
88
assert.doesNotMatch(home, /href=.[^\n]*(?:deprecated|frontstage\/)/, "homepage must not promote retired surfaces");
9-
for (const destination of ["docs/guides/personal-workspace-user-guide/", "benchmarks/swe-marathon/", "benchmarks/deepswe/behavior-discovery/", "benchmarks/deepswe-sol/", "docs/showcases/index", "developers/projections/"]) {
9+
for (const destination of ["docs/guides/personal-workspace-user-guide/", "benchmarks/swe-marathon/", "benchmarks/lhtb/", "benchmarks/deepswe/behavior-discovery/", "benchmarks/deepswe-sol/", "docs/showcases/index", "developers/projections/"]) {
1010
assert.ok(home.includes(destination), `missing public destination: ${destination}`);
1111
}
1212
assert.ok(styles.includes("prefers-reduced-motion"));

‎examples/export-frontstage-share-bundle.mjs‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -268,6 +268,7 @@ async function writeShareReadme(outDir, base, interactivePages) {
268268
const homepageUrl = base;
269269
const frontstageUrl = `${base}frontstage/`;
270270
const sweMarathonBriefUrl = `${base}benchmarks/swe-marathon/`;
271+
const lhtbBriefUrl = `${base}benchmarks/lhtb/`;
271272
const deepSweBehaviorArticleUrl = `${base}benchmarks/deepswe/behavior-discovery/`;
272273
const previewBlock = base === "/"
273274
? `## Try It Locally
@@ -283,6 +284,7 @@ Then open the homepage or showcase:
283284
http://127.0.0.1:8080${homepageUrl}
284285
http://127.0.0.1:8080${frontstageUrl}
285286
http://127.0.0.1:8080${sweMarathonBriefUrl}
287+
http://127.0.0.1:8080${lhtbBriefUrl}
286288
http://127.0.0.1:8080${deepSweBehaviorArticleUrl}
287289
\`\`\`
288290
`
@@ -298,6 +300,7 @@ Hosted entries:
298300
${homepageUrl}
299301
${frontstageUrl}
300302
${sweMarathonBriefUrl}
303+
${lhtbBriefUrl}
301304
${deepSweBehaviorArticleUrl}
302305
\`\`\`
303306
`;
@@ -322,6 +325,7 @@ ${previewBlock}
322325
catalog-declared interactive case pages.
323326
- Homepage source: \`apps/presentation/site\`.
324327
- SWE-Marathon research brief: \`${sweMarathonBriefUrl}\`, built from the pinned public-safe aggregate and case-insight projection under \`benchmark/swe-marathon/\`.
328+
- LHTB research brief: \`${lhtbBriefUrl}\`, built from the public-safe five-arm aggregate under \`benchmark/LHTB/studies/five-arm-gpt56sol-max/\`.
325329
- DeepSWE behavior discoveries: \`${deepSweBehaviorArticleUrl}\`, copied byte-for-byte from the reviewed standalone article at \`${deepSweBehaviorArticlePath}\`.
326330
- DeepSWE × Sol research brief: \`${base}benchmarks/deepswe-sol/\`, a static historical-study interpretation from \`${deepSweSolArticlePath}\`.
327331
- Homepage evidence assets: ${homepageEvidenceAssets.map((path) => `\`${path}\``).join(", ")}.
@@ -345,6 +349,7 @@ async function writeManifest(outDir, base, interactivePages) {
345349
status_fixture: `site/${statusFileName}`,
346350
homepage_entry: "site/index.html",
347351
swe_marathon_brief_entry: "site/benchmarks/swe-marathon/index.html",
352+
lhtb_brief_entry: "site/benchmarks/lhtb/index.html",
348353
deepswe_behavior_article_entry: "site/benchmarks/deepswe/behavior-discovery/index.html",
349354
deepswe_sol_article_entry: "site/benchmarks/deepswe-sol/index.html",
350355
installer_entry: "site/install.sh",
@@ -354,6 +359,7 @@ async function writeManifest(outDir, base, interactivePages) {
354359
content_sources: {
355360
public_homepage: "apps/presentation/site",
356361
swe_marathon_brief: "benchmark/swe-marathon",
362+
lhtb_brief: "benchmark/LHTB/studies/five-arm-gpt56sol-max",
357363
deepswe_behavior_article: deepSweBehaviorArticlePath,
358364
deepswe_sol_article: deepSweSolArticlePath,
359365
installer_script: installerScriptPath,
@@ -474,6 +480,7 @@ async function main() {
474480
site_dir: siteDir,
475481
homepage_url: args.base,
476482
swe_marathon_brief_url: `${args.base}benchmarks/swe-marathon/`,
483+
lhtb_brief_url: `${args.base}benchmarks/lhtb/`,
477484
deepswe_behavior_article_url: `${args.base}benchmarks/deepswe/behavior-discovery/`,
478485
frontstage_url: `${args.base}frontstage/`,
479486
status_fixture: `site/${statusFileName}`,

‎examples/frontstage-share-bundle-smoke.mjs‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -234,6 +234,10 @@ const benchmarkHtml = await readFile(resolve(siteDir, "benchmarks/swe-marathon/i
234234
if (benchmarkHtml !== homepageHtml) {
235235
throw new Error("SWE-Marathon static route must reuse the compiled public-site entry");
236236
}
237+
const lhtbHtml = await readFile(resolve(siteDir, "benchmarks/lhtb/index.html"), "utf8");
238+
if (lhtbHtml !== homepageHtml) {
239+
throw new Error("LHTB static route must reuse the compiled public-site entry");
240+
}
237241
const deepSweBehaviorHtml = await readFile(
238242
resolve(siteDir, "benchmarks/deepswe/behavior-discovery/index.html"),
239243
"utf8",
@@ -367,6 +371,7 @@ if (manifest.base !== "/loopx/") {
367371
if (
368372
manifest.homepage_entry !== "site/index.html" ||
369373
manifest.swe_marathon_brief_entry !== "site/benchmarks/swe-marathon/index.html" ||
374+
manifest.lhtb_brief_entry !== "site/benchmarks/lhtb/index.html" ||
370375
manifest.deepswe_behavior_article_entry !== "site/benchmarks/deepswe/behavior-discovery/index.html" ||
371376
manifest.frontstage_entry !== "site/frontstage/index.html" ||
372377
manifest.installer_entry !== "site/install.sh"
@@ -379,6 +384,9 @@ if (manifest.content_sources?.public_homepage !== "apps/presentation/site") {
379384
if (manifest.content_sources?.swe_marathon_brief !== "benchmark/swe-marathon") {
380385
throw new Error(`manifest benchmark brief source mismatch: ${JSON.stringify(manifest.content_sources)}`);
381386
}
387+
if (manifest.content_sources?.lhtb_brief !== "benchmark/LHTB/studies/five-arm-gpt56sol-max") {
388+
throw new Error(`manifest LHTB brief source mismatch: ${JSON.stringify(manifest.content_sources)}`);
389+
}
382390
if (
383391
manifest.content_sources?.deepswe_behavior_article !==
384392
"benchmark/deepswe/behavior-discovery/index.html"
@@ -458,6 +466,9 @@ if (!readmeText.includes("frontstage/")) {
458466
if (!readmeText.includes("benchmarks/swe-marathon/")) {
459467
throw new Error("share bundle README must publish the SWE-Marathon research brief entry");
460468
}
469+
if (!readmeText.includes("benchmarks/lhtb/")) {
470+
throw new Error("share bundle README must publish the LHTB research brief entry");
471+
}
461472
if (!readmeText.includes("benchmarks/deepswe/behavior-discovery/")) {
462473
throw new Error("share bundle README must publish the DeepSWE behavior article entry");
463474
}

0 commit comments

Comments
 (0)