diff --git a/apps/presentation/site/public/blog/application-scenarios/index.html b/apps/presentation/site/public/blog/application-scenarios/index.html index 16df33e25c..db182422d9 100644 --- a/apps/presentation/site/public/blog/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/application-scenarios/index.html @@ -97,7 +97,9 @@

LHTB: what does LoopX accomplish beyond Plain and native Goal?

LoopX Heartbeat (1.0.3)0.49487

Observation: against Plain, LoopX gains 0.0731 mean reward (17.3%), with 17 wins, 13 ties and 16 losses, and the same solve count. Against native Goal, it gains 0.0473 (10.6%), with 23 wins, 13 ties and 10 losses, and 7 solves versus 4. Counts compare published unrounded scores strictly; higher mean reward does not imply better results on every task.

-

Insight: Plain advances within a session, native Goal retains the objective, and LoopX also organizes subsequent work through registry and Todos. Tabular and Sokoban improve over both baselines; the PoC task ties native Goal; DuckDB scores below it. The questions are whether durable work state reliably adds useful progress, and whether acceptance and rollback preserve that progress as a valid result.

+

Where gains cluster: in post-hoc groups based on task demands, research/modeling (4 tasks) gains +0.2802 / +0.2179 mean reward over Plain / Goal; logic puzzles (4 tasks) gain +0.2739 / +0.1374. Research stays positive without Tabular. Sokoban retains solved levels, Rush Hour requires a route ledger and replay checks, and Tabular requires experiments and validation—promising conditions to test.

+

Where the advantage is unclear: multimodal analysis (6 tasks) yields −0.0034 / +0.0058; science and simulation (7 tasks) yield +0.0161 / +0.0099. Games are mixed: 2048 improves, but Snake trails Plain. APEX law trails Goal, and DuckDB scores 0 versus about 0.770. Continuation cannot substitute for perception, domain judgment or correctness checks.

+

Insight: prioritize tasks where the next step can be checked and useful progress retained. The median paired delta across all tasks is 0 against Plain and about 0.0003 against Goal: mean gains are concentrated. Groups contain only 4–7 tasks and are not an official task mapping. Prompt structure suggests explanations, but does not establish the mechanism behind a trial’s score. Nine groups: counts, wins/ties/losses, sensitivity and task instructions →

Boundary: LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $344.25. Accounting also differs; equal-budget efficiency gains are not established.

diff --git a/apps/presentation/site/public/blog/zh/application-scenarios/index.html b/apps/presentation/site/public/blog/zh/application-scenarios/index.html index a96466afdf..00cc828646 100644 --- a/apps/presentation/site/public/blog/zh/application-scenarios/index.html +++ b/apps/presentation/site/public/blog/zh/application-scenarios/index.html @@ -94,7 +94,9 @@

LHTB:比 Plain 和原生 Goal 多做成了什么?

LoopX Heartbeat(1.0.3)0.49487

观察:相对 Plain,LoopX 均分增加 0.0731(17.3%),逐题为 17 胜、13 平、16 负,通过数相同;相对原生 Goal,均分增加 0.0473(10.6%),为 23 胜、13 平、10 负,通过数从 4 到 7。胜负按公开原始分数严格比较;均分提升不代表每题都更强。

-

Insight:Plain 在会话内推进,原生 Goal 保留目标,LoopX 进一步用 Registry 与 Todo 组织后续工作。Tabular 与 Sokoban 相对两种基线均有收益;PoC 与原生 Goal 持平;DuckDB 却低于原生 Goal。值得验证的是,持久工作状态能否稳定增加有效进展,以及验收和回滚能否把进展变成保得住的结果。

+

哪些题更受益:按任务要求做事后分组,研究复现/建模(4 题)相对 Plain / Goal 的平均差为 +0.2802 / +0.2179;逻辑谜题(4 题)为 +0.2739 / +0.1374。去掉 Tabular,研究组仍正向。Sokoban 能保留已解关卡,Rush Hour 要记录并复查长路线,Tabular 要反复实验与验证——这是值得检验的适用条件。

+

哪里没有清晰优势:多模态(6 题)为 −0.0034 / +0.0058,科学计算与仿真(7 题)为 +0.0161 / +0.0099。游戏组不能一概而论:2048 提升,Snake 却低于 Plain;APEX 法律事项低于 Goal,DuckDB 更是 0 对约 0.770。持续推进不能替代感知、领域判断与正确性验收。

+

Insight:优先关注“能验证下一步,也能保留有效进展”的任务。全体逐题差值中位数,对 Plain 为 0、对 Goal 约 0.0003,均分收益集中。分组每类仅 4–7 题,且不是官方逐题分类;题目结构只提供解释假设,尚不能确认实际收益来自哪种机制。九组的样本量、胜平负、敏感性与题目原文 →

范围:这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $344.25;计费口径也不一致,尚不能宣称同预算效率优势。

diff --git a/apps/presentation/site/src/LhtbBrief.tsx b/apps/presentation/site/src/LhtbBrief.tsx index a2c7210df1..372abc8f28 100644 --- a/apps/presentation/site/src/LhtbBrief.tsx +++ b/apps/presentation/site/src/LhtbBrief.tsx @@ -11,25 +11,39 @@ import { import { useEffect, useMemo, useState } from "react"; import { usePublicPageNavigation } from "./usePublicPageNavigation"; import study from "../../../../benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json"; +import taskGroups from "../../../../benchmark/LHTB/studies/five-arm-gpt56sol-max/task-groups.json"; import copy from "./lhtb-copy.json"; type ArmKey = keyof typeof study.arms; type Baseline = "plain" | "native_goal"; type TableMode = "all" | Baseline; +type GroupKey = keyof typeof taskGroups.groups; +const groupKeys = Object.keys(taskGroups.groups) as GroupKey[]; +const taskGroup = new Map(groupKeys.flatMap((key) => taskGroups.groups[key].map((task) => [task, key] as const))); + +function promptUrl(task: string) { + const aliases: Record = taskGroups.prompt_aliases; + return `${taskGroups.prompt_base_url}${aliases[task] ?? task}/instruction.md`; +} const primaryArms: ArmKey[] = ["plain", "native_goal", "new_heartbeat"]; const historicalArms: ArmKey[] = ["ssh_goal", "legacy_heartbeat"]; const baselines: Baseline[] = ["plain", "native_goal"]; -const comparisons = baselines.map((baseline) => { - const deltas = study.tasks.map((row) => row.new_heartbeat - row[baseline]); +function compareTasks(tasks: typeof study.tasks, baseline: Baseline) { + const deltas = tasks.map((row) => row.new_heartbeat - row[baseline]); const meanDelta = deltas.reduce((sum, delta) => sum + delta, 0) / deltas.length; - const baselineMean = study.tasks.reduce((sum, row) => sum + row[baseline], 0) / deltas.length; + const baselineMean = tasks.reduce((sum, row) => sum + row[baseline], 0) / deltas.length; return { baseline, meanDelta, relativeGain: meanDelta / baselineMean, wins: deltas.filter((delta) => delta > 0).length, ties: deltas.filter((delta) => delta === 0).length, losses: deltas.filter((delta) => delta < 0).length, }; +} +const comparisons = baselines.map((baseline) => compareTasks(study.tasks, baseline)); +const groupComparisons = groupKeys.map((key) => { + const tasks = study.tasks.filter((row) => taskGroup.get(row.task) === key); + return { key, count: tasks.length, comparisons: baselines.map((baseline) => compareTasks(tasks, baseline)) }; }); const contributorLinks = [ @@ -53,6 +67,7 @@ export function LhtbBrief() { const [language, setLanguage] = usePublicPageNavigation(); const [query, setQuery] = useState(""); const [tableMode, setTableMode] = useState("all"); + const [group, setGroup] = useState("all"); const [showHistory, setShowHistory] = useState(false); const c = copy[language]; const visibleArms = showHistory ? [...primaryArms, ...historicalArms] : primaryArms; @@ -67,11 +82,14 @@ export function LhtbBrief() { const visibleTasks = useMemo(() => { const normalized = query.trim().toLowerCase(); return study.tasks.filter((row) => { + if (group !== "all" && taskGroup.get(row.task) !== group) return false; if (normalized && !row.task.toLowerCase().includes(normalized)) return false; if (tableMode !== "all") return Math.abs(row.new_heartbeat - row[tableMode]) >= 0.05; return true; }); - }, [query, tableMode]); + }, [query, tableMode, group]); + + const resetFilters = () => { setQuery(""); setTableMode("all"); setGroup("all"); }; const summaryTable = (arms: ArmKey[]) => (
@@ -104,6 +122,7 @@ export function LhtbBrief() {
{c.armLabels[arm]}
{formatReward(row[arm])}
))}

{note}

+ {c.promptLink} ); }); @@ -226,27 +245,54 @@ export function LhtbBrief() {
-
-
+
+

{c.insightEyebrow}

{c.insightTitle}

{c.insightBody}

+
+

{c.groupMethod}

+
+ + + {c.groupColumns.map((label) => )} + {groupComparisons.map((row) => ( + + + + {row.comparisons.map((comparison) => ( + + ))} + + + ))} +
{c.groupCountLabel} · Δ Reward
{label}
{ resetFilters(); setGroup(row.key); }}>{c.groupLabels[row.key]}{row.count} + {comparison.meanDelta > 0 ? "+" : ""}{comparison.meanDelta.toFixed(4)} + {comparison.wins} / {comparison.ties} / {comparison.losses} + {c.groupNotes[row.key]}
+
+
{c.sensitivityTitle}

{c.sensitivityBody}

+

{c.promptBoundary}

+
{c.insights.map(([title, body], index) => (
0{index + 1}

{title}

{body}

))}
-
-
-

{c.gainTitle}

- {caseCards(c.gainCases)} +
+ {c.caseDetails} +
+
+

{c.gainTitle}

+ {caseCards(c.gainCases)} +
+
+

{c.lossTitle}

+ {caseCards(c.lossCases)} +
-
-

{c.lossTitle}

- {caseCards(c.lossCases)} -
-
+
@@ -256,7 +302,7 @@ export function LhtbBrief() {

{c.scoresBody}

- +
{(["all", ...baselines] as TableMode[]).map((mode) => (
+
+ + + +
@@ -274,17 +328,18 @@ export function LhtbBrief() { const best = Math.max(...visibleArms.map((arm) => row[arm])); return ( - + {visibleArms.map((arm) => ( ))} ); })} + {visibleTasks.length === 0 && }
{row.task}{row.task}{c.groupLabels[taskGroup.get(row.task)!]}{formatReward(row[arm])}
{c.emptyTasks}
-

{c.visibleCount.replace("{count}", String(visibleTasks.length))}

+

{c.visibleCount.replace("{count}", String(visibleTasks.length))}

diff --git a/apps/presentation/site/src/lhtb-brief.css b/apps/presentation/site/src/lhtb-brief.css index f3bf8e90d4..5b984f05e2 100644 --- a/apps/presentation/site/src/lhtb-brief.css +++ b/apps/presentation/site/src/lhtb-brief.css @@ -184,6 +184,29 @@ grid-column: 1 / -1; } +.lhtb-group-analysis { grid-column: 1 / -1; min-width: 0; } +.lhtb-group-analysis > p { color: var(--bm-body); font-size: 14px; line-height: 1.7; } +.lhtb-group-table table { min-width: 820px; table-layout: fixed; } +.lhtb-group-table caption { padding: 12px; text-align: left; color: var(--bm-body); font-size: 12px; } +.lhtb-group-table th:first-child { width: 23%; } +.lhtb-group-table th:nth-child(2) { width: 7%; } +.lhtb-group-table th:nth-child(3), .lhtb-group-table th:nth-child(4) { width: 17%; } +.lhtb-group-table tbody th { white-space: normal; } +.lhtb-group-table td { font-size: 13px; } +.lhtb-group-table strong, .lhtb-group-table small { display: block; font-family: "Geist Mono", monospace; } +.lhtb-group-table small { margin-top: 4px; color: var(--bm-body); font-size: 11px; } +.lhtb-group-table a { display: inline-flex; align-items: center; min-height: 44px; } +.lhtb-group-tools { grid-column: 1 / -1; display: flex; flex-wrap: wrap; align-items: center; gap: 12px; font-size: 13px; } +.lhtb-group-tools select, .lhtb-group-tools button { + min-height: 44px; max-width: 100%; padding: 8px 12px; border: 1px solid var(--bm-border); + border-radius: 6px; color: var(--bm-ink); background: var(--bm-surface); font: inherit; cursor: pointer; +} +.lhtb-page a:focus-visible, .lhtb-page button:focus-visible, .lhtb-page select:focus-visible, .lhtb-page input:focus-visible, .lhtb-page summary:focus-visible { + outline: 2px solid var(--bm-blue); outline-offset: 3px; +} +.lhtb-case-details summary { line-height: 1.6; } +.lhtb-cases article > a { display: inline-flex; align-items: center; gap: 6px; min-height: 44px; font-size: 12px; } + .lhtb-cases { grid-column: 1 / -1; display: grid; diff --git a/apps/presentation/site/src/lhtb-copy.json b/apps/presentation/site/src/lhtb-copy.json index 9a2e5862d6..71483eaf96 100644 --- a/apps/presentation/site/src/lhtb-copy.json +++ b/apps/presentation/site/src/lhtb-copy.json @@ -11,7 +11,7 @@ "readingPath": [ ["01", "Comparison", "#result"], ["02", "Execution", "#mechanisms"], - ["03", "Gains and losses", "#insights"], + ["03", "Task-type signals", "#task-types"], ["04", "Task evidence", "#scores"] ], "resultEyebrow": "00 / Core result", @@ -75,27 +75,27 @@ ["new_heartbeat", "LoopX Heartbeat", "Work state persists across executors", "Workspace, registry, and Todos persist across fresh codex exec wakes. Three qualifying advancement Todos trigger a review that can retain or revise the plan."] ], "insightEyebrow": "03 / LoopX insight", - "insightTitle": "Average progress increased. Task-level evidence explains where.", - "insightBody": "Against Plain, LoopX scores higher on 17 tasks and lower on 16: larger gains outweigh losses, rather than most tasks improving. Against native Goal, it scores higher on 23 and lower on 10. Each case below shows all three scores so that the choice of baseline stays visible.", + "insightTitle": "Which task types benefit, and where is the signal weak?", + "insightBody": "Research and puzzles have higher group means against both baselines. Multimodal analysis and specialist simulations show little separation; other groups need task-level inspection. Across all 46 tasks, the median paired delta is 0 against Plain and about 0.0003 against Goal: mean gains are concentrated, not universal.", "insights": [ - ["More partial progress; full acceptance remains hard", "LoopX’s 0.4948 exceeds both baselines, yet it matches Plain at 7/46 solves. Large gains on Tabular and Sokoban still fall below 0.95. Advancing further and finishing remain distinct."], - ["Native Goal is a substantive control", "On the PoC task, Goal and LoopX both reach 0.892 while Plain scores zero. On DuckDB, Goal reaches 0.770 and LoopX zero. A LoopX-specific claim must account for what native continuation already achieves."], - ["Recovery is a mechanism clue, not attribution", "The original brief reports a consulting task resuming at S27 after interruption. Workspace and Todos provide a place to recover progress; matched-budget ablations are needed to separate recovery, fresh execution, and replanning."], - ["Continuation must preserve results and justify cost", "The reported DuckDB case suggests later optimization can break correctness. Checkpointing, acceptance, and rollback merit separate tests. LoopX’s recorded cost exceeds both baseline estimates; an efficiency gain is not established."] + ["Retained progress and repeated checks are promising conditions", "Sokoban retains solved levels; Tabular requires experiments and an auditable model; Rush Hour asks for a route ledger and replay checks. These structures suit durable work state. Scores alone cannot attribute gains to Todos, fresh exec, or replanning."], + ["Long workflows need feedback about correctness", "APEX consulting beats both baselines, but law trails Goal. The law task requires late corrections and exact cross-document references; validate checks only the schema. Advancing the workflow does not ensure correct content. Persistent state still needs reliable acceptance."], + ["Continuation has not shown a general capability lift", "Multimodal tasks require segmentation, geometry/OCR and generalization; simulation tasks require correct solver use and physical constraints. Small differences do not establish that continuation resolves those bottlenecks. Perfect ties on LangChain and RISC-V also differ from the all-zero OpenRoad floor."], + ["Test both progress and preservation", "2048 retains peak score and improves, yet Snake also retains peaks and still trails Plain. DuckDB requires every query to be correct; optimization errors can zero the reward. Matched-budget repeats should isolate feedback quality, retained progress and best-version rollback."] ], - "gainTitle": "Gain cases: keep both baselines in view", + "gainTitle": "Gains and possible explanations", "gainCases": [ - ["tabular-data-feature-covshift", "A substantial gain over both baselines, but 0.743 still falls short of strict acceptance."], - ["sokoban", "Higher than both baselines; the final 0.890 still does not count as a strict solve."], - ["poc-exploit-craft", "A large gain over Plain, but a tie with native Goal. This result is not unique to LoopX."], - ["apex-management-consulting-matter", "Higher than both baselines. The original brief separately reports recovery at S27; that association does not establish causality."] + ["tabular-data-feature-covshift", "Sparse inverse recovery under covariate shift: recover a 200-dimensional signal from 50 observations, using experiments, validation history and an auditable model. Gains against both, but 0.743 is still below the solve threshold."], + ["sokoban", "155 sequential puzzles; reset the current puzzle while retaining solved levels. Ahead of both baselines, still below 0.95."], + ["rush-hour-campaign", "Four long spatial-route puzzles with a ledger and replay checks. Programmatic search solvers are explicitly forbidden; the gain is not evidence of automated search-code generation."], + ["unison-paper-reproduction", "Implement parallel simulation, partitioning and adaptive scheduling, with multiple artifacts and held-out seeds. Ahead of both baselines, but only 0.5."] ], - "lossTitle": "Counterexamples: the task and baseline change the conclusion", + "lossTitle": "Counterexamples and limits", "lossCases": [ - ["duckdb-optimizer-closure", "Native Goal retains a substantial score; LoopX and Plain score zero. The original brief reports a correctness regression after optimization, motivating best-version retention and final rollback."], - ["snake-obstacle-campaign", "Higher than native Goal, lower than Plain. Choosing a different baseline changes the apparent result."], - ["great-expectations-audit", "Lower than both baselines. Adding persistent scheduling does not guarantee a better final artifact."], - ["apex-openroad-ibex-signoff", "All three score zero. The control plane did not solve this task; continued execution alone is not evidence of useful progress."] + ["apex-law433-matter", "70 stages with cross-document references and late corrections. Validate checks schema, not content correctness. Above Plain, below Goal: length alone does not guarantee gains."], + ["snake-obstacle-campaign", "Reset and peak-score retention are available, yet LoopX trails Plain while beating Goal. Retaining progress is a candidate condition, not a sufficient one."], + ["duckdb-optimizer-closure", "All 22 TPC-H queries must be correct; a wrong query can zero reward. LoopX and Plain score 0, Goal about 0.770. Terminal scores alone cannot locate the failure."], + ["scientific-figure-data-reconstruction", "Reconstruct curves and tables from images, including axis calibration, line separation and numeric accuracy. Scores are close; no clear control-plane benefit is visible."] ], "scoresEyebrow": "04 / Complete task matrix", "scoresTitle": "Inspect LoopX’s gains and losses against each baseline.", @@ -103,7 +103,7 @@ "searchPlaceholder": "Search task", "filterLabel": "Task filters", "filters": { - "all": "All 46", + "all": "Any delta", "plain": "vs Plain |Δ| ≥ 0.05", "native_goal": "vs Goal |Δ| ≥ 0.05" }, @@ -152,7 +152,47 @@ "historyTitle": "Historical context: SSH-Goal and legacy Heartbeat", "historyBody": "The other two arms remain available for tracing implementation history. The primary comparison throughout this page is this study’s LoopX Heartbeat against Plain and native Goal.", "showHistory": "Show historical arms", - "hideHistory": "Hide historical arms" + "hideHistory": "Hide historical arms", + "groupMethod": "These are analyst-defined, post-hoc groups based on task demands, not an official task mapping. Each of the 46 tasks appears once; groups have only 4–7 tasks. Δ is mean LoopX reward minus the baseline. Wins/ties/losses use unrounded scores and do not imply statistical significance. Select a group to inspect every member.", + "groupColumns": [ + "Analysis group", + "Tasks", + "Δ vs Plain", + "Δ vs native Goal", + "Current signal" + ], + "groupLabels": { + "research": "Research & modeling", + "puzzles": "Logic & spatial puzzles", + "professional": "APEX professional work", + "multimodal": "Multimodal analysis", + "science": "Science & simulation", + "earth": "Earth, climate & energy", + "software": "Software repair & toolchains", + "games": "Games & strategy", + "systems": "Systems, performance & security" + }, + "groupNotes": { + "research": "All 4 beat Plain; ALP and Foldseek tie Goal.", + "puzzles": "3 wins, 1 loss on each side; Sudoku trails Goal, Chess trails Plain.", + "professional": "Positive means; law trails Goal, IB244 ties Plain.", + "multimodal": "All 6 have absolute task deltas < 0.05 against both baselines.", + "science": "Similar means; median paired deltas are 0 against both.", + "earth": "No mean advantage over Plain; slightly ahead of Goal.", + "software": "No wins over Plain: 4 ties, 2 losses; some gains are Goal-only.", + "games": "2048 improves; Snake, Mario and Generals all trail Plain.", + "systems": "PoC drives the gain over Plain; DuckDB drives the loss against Goal." + }, + "groupCountLabel": "W / T / L", + "groupFilterLabel": "Filter by analysis group", + "allGroups": "All groups", + "resetFilters": "Reset filters", + "emptyTasks": "No matching tasks. Reset filters to try again.", + "promptLink": "Task instructions", + "promptBoundary": "Task structure is read from pinned upstream revision d78f5eb (accessed 2026-09-19). The study specifies only a July 2026 snapshot, without per-task source digests. Byte identity with evaluated prompts is unverified; these instructions explain task demands, not what happened in a trial.", + "sensitivityTitle": "Does one task drive the group mean?", + "sensitivityBody": "Without Tabular, research still gains +0.1259 / +0.1111 over Plain / Goal. Removing any one puzzle leaves positive means against both. Systems is different: removing PoC changes its Plain delta from +0.1732 to −0.0065; removing DuckDB changes its Goal delta from −0.1450 to +0.0111. These exclusions are sensitivity checks, not an alternative leaderboard or confidence intervals.", + "caseDetails": "Expand task structure and counterexamples (scores from the same data)" }, "zh": { "meta": "LHTB / LOOPX 与两种基线", @@ -166,7 +206,7 @@ "readingPath": [ ["01", "结果对比", "#result"], ["02", "工作机制", "#mechanisms"], - ["03", "得失与启示", "#insights"], + ["03", "题型信号", "#task-types"], ["04", "逐题核对", "#scores"] ], "resultEyebrow": "00 / 核心结果", @@ -230,27 +270,27 @@ ["new_heartbeat", "LoopX Heartbeat", "工作状态跨执行器保留", "工作区、Registry 与 Todo 持久化,每次唤醒启动新的 codex exec;三个有效推进 Todo 后触发复核,可以维持或调整计划。"] ], "insightEyebrow": "03 / LoopX Insight", - "insightTitle": "均分增加了;优势出现在哪里,还需要逐题解释。", - "insightBody": "对 Plain,17 题更高、16 题更低,均分提升来自收益幅度大于损失,而非多数任务都更强。对原生 Goal,23 题更高、10 题更低。下列案例同时列出三者分数,避免把选择基线造成的差异误读为普遍收益。", + "insightTitle": "哪些任务更受益,哪些还看不出优势?", + "insightBody": "研究复现与逻辑谜题的组均值领先两种基线;多模态和专业仿真差异小,其余组需逐题看。46 题的配对差值中位数,对 Plain 为 0、对 Goal 约 0.0003:均分收益集中,并非普遍提升。", "insights": [ - ["部分进展增加,完整验收仍是难点", "LoopX 的 0.4948 高于两种基线,但与 Plain 同为 7/46。Tabular、Sokoban 等收益明显的任务仍未达到 0.95,说明更多进展与做完之间还有距离。"], - ["原生 Goal 已经是有力的对照", "PoC 任务中,Goal 与 LoopX 都是 0.892,Plain 为 0;DuckDB 则是 Goal 0.770、LoopX 0。需要解释 LoopX 在原生续跑之上多带来了什么,也保留它做得更差的任务。"], - ["恢复是机制线索,尚不是增益归因", "研究简报记录咨询任务中断后从 S27 接续。工作区与 Todo 提供了恢复位置的载体;要证明它提升了分数,还需在相同预算下隔离状态恢复、fresh exec 与 replan。"], - ["续跑要保住结果,也要算成本", "DuckDB 的报告案例提示后续优化可能破坏正确性。检查点、验收与回滚值得单独验证;本轮 LoopX 记录成本高于两种基线的估算成本,尚未证明更经济。"] + ["可保留进度、反复验证:值得优先验证的适用条件", "Sokoban 解出的关卡可以保留;Tabular 要反复实验并提交可审计模型;Rush Hour 要维护路线记录并复查。这些结构适合让工作状态跨轮延续。但仅凭分数,不能证明收益来自 Todo、fresh exec 或重规划中的哪一项。"], + ["长流程本身不够,反馈必须接近正确性", "APEX 咨询高于两侧基线,法律事项却低于 Goal。法律题有晚到修订、跨文档精确引用,validate 只检查格式。流程继续向前,并不等于内容得到纠正;状态管理还需要可靠验收。"], + ["续跑没有显示出普遍的能力补足", "多模态题涉及图像分割、几何/OCR 和跨输入泛化;仿真题还要求正确使用求解器与物理约束。当前差异小,尚无证据说明增加续跑能解决这些瓶颈。也要区分 LangChain、RISC-V 的满分平局与 OpenRoad 的零分平局。"], + ["下一步要验证“能前进,也能保住成果”", "2048 保留峰值并有收益,但同样保留峰值的 Snake 仍低于 Plain。DuckDB 则要求所有查询正确,优化失误可能让分数归零。应在相同预算、重复运行中,单独检验反馈质量、进度保留与最佳版本回滚。"] ], - "gainTitle": "收益案例:看三者,而非只看一个增量", + "gainTitle": "收益与可能解释", "gainCases": [ - ["tabular-data-feature-covshift", "相对两种基线均有明显增益,但 0.743 仍未达到严格通过阈值。"], - ["sokoban", "高于两种基线;最终 0.890,更多进展仍未转化为完整通过。"], - ["poc-exploit-craft", "相对 Plain 提升明显,却与原生 Goal 持平。这项收益不能直接归为 LoopX 特有。"], - ["apex-management-consulting-matter", "分数高于两种基线;原简报另记录了从 S27 恢复的轨迹案例。两者相关,尚不足以证明因果。"] + ["tabular-data-feature-covshift", "稀疏逆问题与分布偏移:50 个观测恢复 200 维信号,需要实验、验证历史与可审计模型。两侧均有增益,但 0.743 仍未通过。"], + ["sokoban", "155 个顺序关卡;可重置当前题而保留已解关卡。两侧均领先,仍未达到 0.95。"], + ["rush-hour-campaign", "4 道长程空间路线题,要求记账与路线复查,明确禁止程序化搜索求解。收益不能解释为自动写搜索算法。"], + ["unison-paper-reproduction", "实现并行仿真、分区与自适应调度,交付多个产物并检查隐藏 seed。两侧均领先,最终仅 0.5。"] ], - "lossTitle": "反例:续跑收益随任务和基线变化", + "lossTitle": "反例与能力边界", "lossCases": [ - ["duckdb-optimizer-closure", "原生 Goal 保住了较高得分,LoopX 与 Plain 为零。原简报报告后续优化破坏正确性,值得验证最佳版本保留与最终回滚。"], - ["snake-obstacle-campaign", "高于原生 Goal,却低于 Plain。换一个基线,结论就会变化。"], - ["great-expectations-audit", "同时低于两种基线,说明引入持续调度并不保证更好的最终产物。"], - ["apex-openroad-ibex-signoff", "三者均为零。控制面未解决这道任务,不能将继续执行本身当成有效进展。"] + ["apex-law433-matter", "70 个阶段含跨文档引用与晚到修订;validate 只查 schema,不查内容正确性。LoopX 高于 Plain、低于 Goal,长流程不保证收益。"], + ["snake-obstacle-campaign", "可重置且保留峰值,但 LoopX 仍低于 Plain、高于 Goal。保留进度只是候选条件,并不充分。"], + ["duckdb-optimizer-closure", "22 个 TPC-H 查询要求全正确,错误可使 Reward 归零。LoopX 与 Plain 为 0、Goal 约 0.770;仅凭终值不能定位失败环节。"], + ["scientific-figure-data-reconstruction", "从图像重建曲线与表格,要处理坐标标定、线条分离与数值误差。三者接近,尚看不到控制面带来的明确收益。"] ], "scoresEyebrow": "04 / 46 题完整矩阵", "scoresTitle": "逐题看 LoopX 相对两种基线的得失。", @@ -258,7 +298,7 @@ "searchPlaceholder": "搜索任务", "filterLabel": "任务筛选", "filters": { - "all": "全部 46 题", + "all": "不限差值", "plain": "对 Plain |Δ| ≥ 0.05", "native_goal": "对 Goal |Δ| ≥ 0.05" }, @@ -307,6 +347,46 @@ "historyTitle": "历史参照:SSH-Goal 与旧版 Heartbeat", "historyBody": "另外两臂保留在这里,便于回溯实现演进。全文的主比较是本研究的 LoopX Heartbeat 与 Plain、原生 Goal。", "showHistory": "显示历史两臂", - "hideHistory": "收起历史两臂" + "hideHistory": "收起历史两臂", + "groupMethod": "以下是按任务要求做的事后分析分组,非官方逐题分类。46 题各归一组;每组仅 4–7 题。Δ 是 LoopX 减基线的平均 Reward,胜/平/负按未四舍五入的分数严格比较,不代表统计显著性。点击组名查看完整成员。", + "groupColumns": [ + "分析分组", + "题数", + "Δ 对 Plain", + "Δ 对原生 Goal", + "当前信号" + ], + "groupLabels": { + "research": "研究复现与建模", + "puzzles": "逻辑与空间谜题", + "professional": "APEX 专业工作流", + "multimodal": "多模态分析", + "science": "科学计算与仿真", + "earth": "地球、气候与能源", + "software": "软件修复与工具链", + "games": "游戏与策略", + "systems": "系统、性能与安全" + }, + "groupNotes": { + "research": "4 题均高于 Plain;ALP、Foldseek 与 Goal 持平。", + "puzzles": "两侧均 3 胜 1 负;Sudoku 低于 Goal,Chess 低于 Plain。", + "professional": "组均值正向;法律事项低于 Goal,IB244 与 Plain 持平。", + "multimodal": "6 题对两种基线的单题绝对差均 < 0.05。", + "science": "均值接近;两侧配对差值中位数均为 0。", + "earth": "相对 Plain 没有组均值优势,对 Goal 略高。", + "software": "对 Plain 无胜出题:4 平 2 负;部分收益只对 Goal 成立。", + "games": "2048 提升;Snake、Mario、Generals 均低于 Plain。", + "systems": "对 Plain 的组收益由 PoC 拉动;对 Goal 的损失由 DuckDB 主导。" + }, + "groupCountLabel": "胜 / 平 / 负", + "groupFilterLabel": "按分析分组查看", + "allGroups": "全部分组", + "resetFilters": "重置筛选", + "emptyTasks": "没有匹配任务。可重置筛选后重试。", + "promptLink": "题目原文", + "promptBoundary": "题目结构参考上游固定版本 d78f5eb(2026-09-19 读取)。实验仅注明 July 2026 snapshot,未公开逐题源文件摘要,不能确认这里的题目与当时运行版本逐字一致;原文用于理解任务要求,不能据此还原运行轨迹。", + "sensitivityTitle": "组均值是否被一道题左右?", + "sensitivityBody": "去掉 Tabular,研究组仍对 Plain / Goal 为 +0.1259 / +0.1111;谜题组任去一道题,两侧均值仍正向。系统组则不同:去掉 PoC,对 Plain 从 +0.1732 变为 −0.0065;去掉 DuckDB,对 Goal 从 −0.1450 变为 +0.0111。删题只是敏感性检查,不是另一份榜单,也不是置信区间。", + "caseDetails": "展开题目结构与正反案例(分数来自同一份数据)" } } diff --git a/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md b/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md index 7f0e2ad31d..ac1559acf6 100644 --- a/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md +++ b/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md @@ -52,3 +52,110 @@ The task matrix defaults to Plain, Native Goal and LoopX Heartbeat; readers can expand both historical arms. Case scores come from the same task rows rather than a separate editorial copy. The recorded scores support outcome comparisons; reported recovery or regression cases are mechanism clues, not causal estimates. + +## Exploratory task-type analysis + +`task-groups.json` supplies an exhaustive, analyst-defined **post-hoc** partition +of these 46 tasks. It changes no rewards, arms, scoring or trial selection. +The brief computes group deltas and win/tie/loss counts from `data.json` using +the same reducer as its overall comparison. Selecting a group reveals its +complete membership in the task matrix; search and baseline filters intersect. + +The nine headings resemble the upstream introduction, but upstream website, +README and task metadata do not supply one consistent task-to-category map. +This is **not an official category leaderboard**. Membership follows the main +deliverable: paper-method pipelines and inverse modeling; constraint puzzles; +staged professional matters; perception/reconstruction; physical simulation; +earth/energy tools; software/toolchain repair; games; or systems optimization, +performance and security. These boundaries overlap (e.g. a research task also +requires software); they describe tasks, not measured causes of success. + +### Complete membership and outcomes + +Δ = mean(task LoopX 1.0.3 Heartbeat reward − task baseline reward). +Every task has equal weight within its group. Counts are strict W/T/L on +unrounded values; rounded zero or a tiny win does not mean statistical equality +or meaningful superiority. These small, selected groups are descriptive. + +| Group | n | Δ Plain; W/T/L | Δ Goal; W/T/L | +| --- | ---: | --- | --- | +| Research & modeling | 4 | +0.2802; 4/0/0 | +0.2179; 2/2/0 | +| Logic & spatial puzzles | 4 | +0.2739; 3/0/1 | +0.1374; 3/0/1 | +| APEX professional work | 4 | +0.1022; 3/1/0 | +0.0407; 3/0/1 | +| Multimodal analysis | 6 | -0.0034; 1/3/2 | +0.0058; 3/3/0 | +| Science & simulation | 7 | +0.0161; 2/2/3 | +0.0099; 3/2/2 | +| Earth, climate & energy | 6 | -0.0108; 2/2/2 | +0.0382; 3/2/1 | +| Software repair & toolchains | 6 | -0.0188; 0/4/2 | +0.0991; 2/3/1 | +| Games & strategy | 4 | -0.0112; 1/0/3 | +0.0971; 3/0/1 | +| Systems, performance & security | 5 | +0.1732; 1/1/3 | -0.1450; 1/1/3 | + +Each link opens the pinned public task instructions used for structural reading: + +- **Research & modeling:** [alp-paper-reproduction](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/alp-paper-reproduction/instruction.md), [foldseek-paper-reproduction](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/foldseek-paper-reproduction/instruction.md), [tabular-data-feature-covshift](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/tabular-data-feature-covshift/instruction.md), [unison-paper-reproduction](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/unison-paper-reproduction/instruction.md). +- **Logic & spatial puzzles:** [chess-mate](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/chess-mate/instruction.md), [rush-hour-campaign](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/rush_hour_campaign/instruction.md), [sokoban](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/sokoban/instruction.md), [sudoku-recovery](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/sudoku-recovery/instruction.md). +- **APEX professional work:** [apex-ib244-matter](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/apex-ib244-matter/instruction.md), [apex-investment-banking-matter](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/apex-investment-banking-matter/instruction.md), [apex-law433-matter](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/apex-law433-matter/instruction.md), [apex-management-consulting-matter](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/apex-management-consulting-matter/instruction.md). +- **Multimodal analysis:** [audio-visual-event-alignment](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/audio-visual-event-alignment/instruction.md), [dicom-radiology-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/dicom-radiology-audit/instruction.md), [document-table-layout-reconstruction](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/document-table-layout-reconstruction/instruction.md), [microscopy-cell-count-qc-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/microscopy-cell-count-qc-audit/instruction.md), [satellite-flood-change-detection-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/satellite-flood-change-detection-audit/instruction.md), [scientific-figure-data-reconstruction](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/scientific-figure-data-reconstruction/instruction.md). +- **Science & simulation:** [epidemic-inverse-control-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/epidemic-inverse-control-audit/instruction.md), [materials-phase-diagram-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/materials-phase-diagram-audit/instruction.md), [nbody-accel-iterative](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/nbody-accel-iterative/instruction.md), [opensees-seismic-structural-regression-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/opensees-seismic-structural-regression-audit/instruction.md), [robotics-slam-benchmark-repair](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/robotics-slam-benchmark-repair/instruction.md), [spice-ephemeris-regression](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/spice-ephemeris-regression/instruction.md), [su2-airfoil-regression](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/su2-airfoil-regression/instruction.md). +- **Earth, climate & energy:** [climate-netcdf-extreme-event-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/climate-netcdf-extreme-event-audit/instruction.md), [epa-swmm-stormwater-regression-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/epa-swmm-stormwater-regression-audit/instruction.md), [gdal-proj-raster-regression](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/gdal-proj-raster-regression/instruction.md), [matpower-opf-regression](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/matpower-opf-regression/instruction.md), [modflow6-groundwater-regression-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/modflow6-groundwater-regression-audit/instruction.md), [nrel-pysam-hybrid-renewables-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/nrel-pysam-hybrid-renewables-audit/instruction.md). +- **Software repair & toolchains:** [apex-openroad-ibex-signoff](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/apex-openroad-ibex-signoff/instruction.md), [commit0-multilib-tdd](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/commit0-multilib-tdd/instruction.md), [great-expectations-audit](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/great-expectations-audit/instruction.md), [langchain-version-migration](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/langchain-version-migration/instruction.md), [riscv-core-debug](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/riscv-core-debug/instruction.md), [unknown-config-semantics](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/unknown-config-semantics/instruction.md). +- **Games & strategy:** [2048](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/2048/instruction.md), [generals-bot-arena](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/generals-bot-arena/instruction.md), [snake-obstacle-campaign](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/snake_maze_campaign/instruction.md), [super-mario](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/super-mario/instruction.md). +- **Systems, performance & security:** [duckdb-optimizer-closure](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/duckdb-optimizer-closure/instruction.md), [grammar-fuzz-coverage-hunt](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/grammar-fuzz-coverage-hunt/instruction.md), [poc-exploit-craft](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/poc-exploit-craft/instruction.md), [spot-scheduler-traces](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/spot-scheduler-traces/instruction.md), [vector-db-iterative-build](https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/vector-db-iterative-build/instruction.md). + +### Concentration and sensitivity + +- Across all 46 tasks, median paired delta is **0** against Plain and + **0.000263888889** against Goal. Higher means do not imply a typical task gains. +- Research: without Tabular, mean deltas remain **+0.1259 / +0.1111** against + Plain / Goal. ALP and Foldseek tie Goal; not all four beat both baselines. +- Puzzles: removing any one task leaves positive mean deltas on both sides + (ranges **+0.1588…+0.3708** and **+0.0843…+0.2069**). Sudoku trails Goal; + Chess trails Plain. This is still only four tasks. +- Systems: without PoC, the Plain delta changes from **+0.1732 to −0.0065**; + without DuckDB, the Goal delta changes from **−0.1450 to +0.0111**. + The group label hides two very different baseline-dependent cases. +- All six multimodal tasks have absolute paired deltas below 0.05 against + both baselines. Science has zero median paired deltas on both sides. +- Software has no wins over Plain (4 ties, 2 losses). Perfect ties on + LangChain and RISC-V indicate no headroom; all-zero OpenRoad is a failure + floor. Neither establishes equal underlying capability. + +Leave-one-out values remove exactly one task from that group's arithmetic +mean. They are sensitivity checks, not confidence intervals, revised study +results or permission to discard inconvenient tasks. + +### What reading the tasks adds + +The public instructions suggest testable conditions, not trajectory evidence: + +- Sokoban retains solved levels when resetting the current puzzle; 2048 + retains peak scores. Tabular calls for iterative experiments and an audited + model. Rush Hour requires manual spatial reasoning, route bookkeeping and + replay checks, and **forbids programmatic search solvers**. These are plausible + settings for durable progress, but Snake also retains peaks and trails Plain. +- APEX consulting has 33 dependent stages; law has 70, with exact document + references and late corrections. Both expose a schema-only `validate` + command. Law trails Goal: advancing stages is not evidence of semantic + correctness. Long workflows alone do not predict an advantage. +- DuckDB requires all 22 TPC-H queries to remain correct. This illustrates a + correctness veto and a reason to test best-version retention/rollback; + final scores alone do not identify the trial's failure mechanism. +- Multimodal and simulation tasks require working perception/numerical + pipelines and generalization, not just persistent task state. Small score + differences do not show that extra continuation supplies those capabilities. + +Instructions are pinned to upstream commit +`d78f5eb52ad754c5ee9154741af73130a85a65b8`, accessed 2026-09-19. Study IDs +`rush-hour-campaign` and `snake-obstacle-campaign` map to upstream folders +`rush_hour_campaign` and `snake_maze_campaign`. The extra current upstream task +`genetic-convergence-testing` is not in the 46-task study and is excluded. +The study identifies a **July 2026 snapshot without per-task source digests**; +byte identity between these public prompts and the evaluated versions is +unverified. Only representative instruction bodies were closely read; links +provide all members for inspection. No hidden tests or solutions were used. + +There is one effective trial per cell, replacement trials and unequal runtime +budgets. Post-hoc grouping adds selection and taxonomy uncertainty. None of +these observations establishes statistical significance, a causal mechanism, +an equal-budget efficiency gain or generalization to another model/provider. +The next discriminating study would fix budgets, repeat matched task-arm runs, +and isolate progress retention, correctness feedback and rollback with traces. diff --git a/benchmark/LHTB/studies/five-arm-gpt56sol-max/task-groups.json b/benchmark/LHTB/studies/five-arm-gpt56sol-max/task-groups.json new file mode 100644 index 0000000000..459b8b80e9 --- /dev/null +++ b/benchmark/LHTB/studies/five-arm-gpt56sol-max/task-groups.json @@ -0,0 +1,76 @@ +{ + "schema_version": 1, + "classification": "analyst-defined post-hoc task-demand groups; not an official task mapping", + "prompt_revision": "d78f5eb52ad754c5ee9154741af73130a85a65b8", + "prompt_base_url": "https://github.com/zli12321/LHTB/blob/d78f5eb52ad754c5ee9154741af73130a85a65b8/tasks/", + "prompt_aliases": { + "rush-hour-campaign": "rush_hour_campaign", + "snake-obstacle-campaign": "snake_maze_campaign" + }, + "groups": { + "research": [ + "alp-paper-reproduction", + "foldseek-paper-reproduction", + "tabular-data-feature-covshift", + "unison-paper-reproduction" + ], + "puzzles": [ + "chess-mate", + "rush-hour-campaign", + "sokoban", + "sudoku-recovery" + ], + "professional": [ + "apex-ib244-matter", + "apex-investment-banking-matter", + "apex-law433-matter", + "apex-management-consulting-matter" + ], + "multimodal": [ + "audio-visual-event-alignment", + "dicom-radiology-audit", + "document-table-layout-reconstruction", + "microscopy-cell-count-qc-audit", + "satellite-flood-change-detection-audit", + "scientific-figure-data-reconstruction" + ], + "science": [ + "epidemic-inverse-control-audit", + "materials-phase-diagram-audit", + "nbody-accel-iterative", + "opensees-seismic-structural-regression-audit", + "robotics-slam-benchmark-repair", + "spice-ephemeris-regression", + "su2-airfoil-regression" + ], + "earth": [ + "climate-netcdf-extreme-event-audit", + "epa-swmm-stormwater-regression-audit", + "gdal-proj-raster-regression", + "matpower-opf-regression", + "modflow6-groundwater-regression-audit", + "nrel-pysam-hybrid-renewables-audit" + ], + "software": [ + "apex-openroad-ibex-signoff", + "commit0-multilib-tdd", + "great-expectations-audit", + "langchain-version-migration", + "riscv-core-debug", + "unknown-config-semantics" + ], + "games": [ + "2048", + "generals-bot-arena", + "snake-obstacle-campaign", + "super-mario" + ], + "systems": [ + "duckdb-optimizer-closure", + "grammar-fuzz-coverage-hunt", + "poc-exploit-craft", + "spot-scheduler-traces", + "vector-db-iterative-build" + ] + } +} diff --git a/benchmark/tests/test_publication_scope.py b/benchmark/tests/test_publication_scope.py index af96fbb560..df48cf5c21 100644 --- a/benchmark/tests/test_publication_scope.py +++ b/benchmark/tests/test_publication_scope.py @@ -114,3 +114,26 @@ def test_lhtb_published_data_and_bilingual_copy_share_scope(): assert "7, 7, and 4 tasks" in localized_copy["en"]["resultBody"] assert localized_copy["zh"]["resultTitle"] == "均分高于 Plain 和原生 Goal。" assert "7、7、4 题" in localized_copy["zh"]["resultBody"] + + +def test_lhtb_task_groups_partition_the_published_study_and_have_prompt_links(): + study = STUDY.parent / "LHTB" / "studies" / "five-arm-gpt56sol-max" + data = json.loads((study / "data.json").read_text()) + analysis = json.loads((study / "task-groups.json").read_text()) + groups = analysis["groups"] + members = [task for tasks in groups.values() for task in tasks] + assert len(members) == len(set(members)) == len(data["tasks"]) + assert set(members) == {row["task"] for row in data["tasks"]} + assert all(groups.values()) + assert set(analysis["prompt_aliases"]) <= set(members) + revision = analysis["prompt_revision"] + assert len(revision) == 40 and all(c in "0123456789abcdef" for c in revision) + assert analysis["prompt_base_url"] == f"https://github.com/zli12321/LHTB/blob/{revision}/tasks/" + assert all("/" not in folder and folder not in {".", ".."} + for folder in analysis["prompt_aliases"].values()) + + site = STUDY.parents[1] / "apps/presentation/site/src" + for localized in json.loads((site / "lhtb-copy.json").read_text()).values(): + assert set(localized["groupLabels"]) == set(localized["groupNotes"]) == set(groups) + for task, _ in localized["gainCases"] + localized["lossCases"]: + assert task in members