Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -97,7 +97,9 @@ <h3>LHTB: what does LoopX accomplish beyond Plain and native Goal?</h3>
<tr><th scope="row">LoopX Heartbeat (1.0.3)</th><td>0.4948</td><td>7</td></tr>
</tbody></table></div>
<p><strong>Observation:</strong> against Plain, LoopX gains 0.0731 mean reward (17.3%), with 17 wins, 13 ties and 16 losses, and the same solve count. Against native Goal, it gains 0.0473 (10.6%), with 23 wins, 13 ties and 10 losses, and 7 solves versus 4. Counts compare published unrounded scores strictly; higher mean reward does not imply better results on every task.</p>
<p><strong>Insight:</strong> Plain advances within a session, native Goal retains the objective, and LoopX also organizes subsequent work through registry and Todos. Tabular and Sokoban improve over both baselines; the PoC task ties native Goal; DuckDB scores below it. The questions are whether durable work state reliably adds useful progress, and whether acceptance and rollback preserve that progress as a valid result.</p>
<p><strong>Where gains cluster:</strong> in post-hoc groups based on task demands, research/modeling (4 tasks) gains +0.2802 / +0.2179 mean reward over Plain / Goal; logic puzzles (4 tasks) gain +0.2739 / +0.1374. Research stays positive without Tabular. Sokoban retains solved levels, Rush Hour requires a route ledger and replay checks, and Tabular requires experiments and validation—promising conditions to test.</p>
<p><strong>Where the advantage is unclear:</strong> multimodal analysis (6 tasks) yields −0.0034 / +0.0058; science and simulation (7 tasks) yield +0.0161 / +0.0099. Games are mixed: 2048 improves, but Snake trails Plain. APEX law trails Goal, and DuckDB scores 0 versus about 0.770. Continuation cannot substitute for perception, domain judgment or correctness checks.</p>
<p><strong>Insight:</strong> prioritize tasks where the next step can be checked and useful progress retained. The median paired delta across all tasks is 0 against Plain and about 0.0003 against Goal: mean gains are concentrated. Groups contain only 4–7 tasks and are not an official task mapping. Prompt structure suggests explanations, but does not establish the mechanism behind a trial’s score. <a href="../../benchmarks/lhtb/?lang=en#task-types">Nine groups: counts, wins/ties/losses, sensitivity and task instructions →</a></p>
<p class="signal-boundary"><strong>Boundary:</strong> LoopX here is the study’s 1.0.3 fresh-exec Heartbeat. Execution settings differ, replacement trials include longer limits, and there are no repeated seeds. LoopX records $551.73, above the Plain / Goal estimates of $212.97 / $344.25. Accounting also differs; equal-budget efficiency gains are not established.</p>
<p class="study-link"><a href="../../benchmarks/lhtb/?lang=en">Full study: two baselines, gains and losses, 46 tasks, and historical context →</a></p>
</div>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -94,7 +94,9 @@ <h3>LHTB:比 Plain 和原生 Goal 多做成了什么?</h3>
<tr><th scope="row">LoopX Heartbeat(1.0.3)</th><td>0.4948</td><td>7</td></tr>
</tbody></table></div>
<p><strong>观察:</strong>相对 Plain,LoopX 均分增加 0.0731(17.3%),逐题为 17 胜、13 平、16 负,通过数相同;相对原生 Goal,均分增加 0.0473(10.6%),为 23 胜、13 平、10 负,通过数从 4 到 7。胜负按公开原始分数严格比较;均分提升不代表每题都更强。</p>
<p><strong>Insight:</strong>Plain 在会话内推进,原生 Goal 保留目标,LoopX 进一步用 Registry 与 Todo 组织后续工作。Tabular 与 Sokoban 相对两种基线均有收益;PoC 与原生 Goal 持平;DuckDB 却低于原生 Goal。值得验证的是,持久工作状态能否稳定增加有效进展,以及验收和回滚能否把进展变成保得住的结果。</p>
<p><strong>哪些题更受益:</strong>按任务要求做事后分组,研究复现/建模(4 题)相对 Plain / Goal 的平均差为 +0.2802 / +0.2179;逻辑谜题(4 题)为 +0.2739 / +0.1374。去掉 Tabular,研究组仍正向。Sokoban 能保留已解关卡,Rush Hour 要记录并复查长路线,Tabular 要反复实验与验证——这是值得检验的适用条件。</p>
<p><strong>哪里没有清晰优势:</strong>多模态(6 题)为 −0.0034 / +0.0058,科学计算与仿真(7 题)为 +0.0161 / +0.0099。游戏组不能一概而论:2048 提升,Snake 却低于 Plain;APEX 法律事项低于 Goal,DuckDB 更是 0 对约 0.770。持续推进不能替代感知、领域判断与正确性验收。</p>
<p><strong>Insight:</strong>优先关注“能验证下一步,也能保留有效进展”的任务。全体逐题差值中位数,对 Plain 为 0、对 Goal 约 0.0003,均分收益集中。分组每类仅 4–7 题,且不是官方逐题分类;题目结构只提供解释假设,尚不能确认实际收益来自哪种机制。<a href="../../../benchmarks/lhtb/?lang=zh#task-types">九组的样本量、胜平负、敏感性与题目原文 →</a></p>
<p class="signal-boundary"><strong>范围:</strong>这里的 LoopX 是实验中的 1.0.3 fresh-exec Heartbeat。运行设置不同,且包含更长时限的替代 trial,没有重复 seed。LoopX 记录成本 $551.73,高于 Plain / Goal 的估算 $212.97 / $344.25;计费口径也不一致,尚不能宣称同预算效率优势。</p>
<p class="study-link"><a href="../../../benchmarks/lhtb/?lang=zh">研究全文:两种基线、正反案例、46 题矩阵与历史参照 →</a></p>
</div>
Expand Down
91 changes: 73 additions & 18 deletions apps/presentation/site/src/LhtbBrief.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -11,25 +11,39 @@
import { useEffect, useMemo, useState } from "react";
import { usePublicPageNavigation } from "./usePublicPageNavigation";
import study from "../../../../benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json";
import taskGroups from "../../../../benchmark/LHTB/studies/five-arm-gpt56sol-max/task-groups.json";
import copy from "./lhtb-copy.json";

type ArmKey = keyof typeof study.arms;
type Baseline = "plain" | "native_goal";
type TableMode = "all" | Baseline;
type GroupKey = keyof typeof taskGroups.groups;
const groupKeys = Object.keys(taskGroups.groups) as GroupKey[];
const taskGroup = new Map(groupKeys.flatMap((key) => taskGroups.groups[key].map((task) => [task, key] as const)));

function promptUrl(task: string) {
const aliases: Record<string, string> = taskGroups.prompt_aliases;
return `${taskGroups.prompt_base_url}${aliases[task] ?? task}/instruction.md`;
}

const primaryArms: ArmKey[] = ["plain", "native_goal", "new_heartbeat"];
const historicalArms: ArmKey[] = ["ssh_goal", "legacy_heartbeat"];
const baselines: Baseline[] = ["plain", "native_goal"];
const comparisons = baselines.map((baseline) => {
const deltas = study.tasks.map((row) => row.new_heartbeat - row[baseline]);
function compareTasks(tasks: typeof study.tasks, baseline: Baseline) {
const deltas = tasks.map((row) => row.new_heartbeat - row[baseline]);
const meanDelta = deltas.reduce((sum, delta) => sum + delta, 0) / deltas.length;
const baselineMean = study.tasks.reduce((sum, row) => sum + row[baseline], 0) / deltas.length;
const baselineMean = tasks.reduce((sum, row) => sum + row[baseline], 0) / deltas.length;
return {
baseline, meanDelta, relativeGain: meanDelta / baselineMean,
wins: deltas.filter((delta) => delta > 0).length,
ties: deltas.filter((delta) => delta === 0).length,
losses: deltas.filter((delta) => delta < 0).length,
};
}
const comparisons = baselines.map((baseline) => compareTasks(study.tasks, baseline));
const groupComparisons = groupKeys.map((key) => {
const tasks = study.tasks.filter((row) => taskGroup.get(row.task) === key);
return { key, count: tasks.length, comparisons: baselines.map((baseline) => compareTasks(tasks, baseline)) };
});

const contributorLinks = [
Expand All @@ -53,6 +67,7 @@
const [language, setLanguage] = usePublicPageNavigation();
const [query, setQuery] = useState("");
const [tableMode, setTableMode] = useState<TableMode>("all");
const [group, setGroup] = useState<GroupKey | "all">("all");
const [showHistory, setShowHistory] = useState(false);
const c = copy[language];
const visibleArms = showHistory ? [...primaryArms, ...historicalArms] : primaryArms;
Expand All @@ -67,11 +82,14 @@
const visibleTasks = useMemo(() => {
const normalized = query.trim().toLowerCase();
return study.tasks.filter((row) => {
if (group !== "all" && taskGroup.get(row.task) !== group) return false;
if (normalized && !row.task.toLowerCase().includes(normalized)) return false;
if (tableMode !== "all") return Math.abs(row.new_heartbeat - row[tableMode]) >= 0.05;
return true;
});
}, [query, tableMode]);
}, [query, tableMode, group]);

const resetFilters = () => { setQuery(""); setTableMode("all"); setGroup("all"); };

const summaryTable = (arms: ArmKey[]) => (
<div className="bm-table-wrap lhtb-summary-table">
Expand Down Expand Up @@ -104,6 +122,7 @@
<div key={arm}><dt>{c.armLabels[arm]}</dt><dd>{formatReward(row[arm])}</dd></div>
))}</dl>
<p>{note}</p>
<a href={promptUrl(task)} target="_blank" rel="noreferrer">{c.promptLink} <ExternalLink size={11} /></a>
</article>
);
});
Expand Down Expand Up @@ -226,27 +245,54 @@
</div>
</section>

<section className="bm-section bm-shell lhtb-insight-section" id="insights">
<div className="bm-section-lead bm-section-lead-wide">
<section className="bm-section bm-shell lhtb-insight-section" id="task-types">
<div className="bm-section-lead bm-section-lead-wide" id="insights">
<p className="bm-kicker">{c.insightEyebrow}</p>
<h2>{c.insightTitle}</h2>
<p>{c.insightBody}</p>
</div>
<div className="lhtb-group-analysis">
<p>{c.groupMethod}</p>
<div className="bm-table-wrap lhtb-group-table">
<table>
<caption>{c.groupCountLabel} · Δ Reward</caption>
<thead><tr>{c.groupColumns.map((label) => <th scope="col" key={label}>{label}</th>)}</tr></thead>
<tbody>{groupComparisons.map((row) => (
<tr key={row.key} data-group={row.key}>
<th scope="row"><a href="#scores" onClick={() => { resetFilters(); setGroup(row.key); }}>{c.groupLabels[row.key]}</a></th>
<td>{row.count}</td>
{row.comparisons.map((comparison) => (
<td key={comparison.baseline}>
<strong>{comparison.meanDelta > 0 ? "+" : ""}{comparison.meanDelta.toFixed(4)}</strong>
<small>{comparison.wins} / {comparison.ties} / {comparison.losses}</small>
</td>
))}
<td>{c.groupNotes[row.key]}</td>
</tr>
))}</tbody>
</table>
</div>
<details className="lhtb-history"><summary>{c.sensitivityTitle}</summary><p>{c.sensitivityBody}</p></details>
<p className="bm-runner-note">{c.promptBoundary}</p>
</div>
<div className="bm-insight-grid lhtb-insight-grid">
{c.insights.map(([title, body], index) => (
<article key={title}><span>0{index + 1}</span><h3>{title}</h3><p>{body}</p></article>
))}
</div>
<div className="lhtb-cases">
<div>
<p className="bm-kicker">{c.gainTitle}</p>
{caseCards(c.gainCases)}
<details className="lhtb-history lhtb-case-details">
<summary>{c.caseDetails}</summary>
<div className="lhtb-cases">
<div>
<p className="bm-kicker">{c.gainTitle}</p>
{caseCards(c.gainCases)}
</div>
<div>
<p className="bm-kicker lhtb-caution">{c.lossTitle}</p>
{caseCards(c.lossCases)}
</div>
</div>
<div>
<p className="bm-kicker lhtb-caution">{c.lossTitle}</p>
{caseCards(c.lossCases)}
</div>
</div>
</details>
</section>

<section className="bm-section bm-shell" id="scores">
Expand All @@ -256,7 +302,7 @@
<p>{c.scoresBody}</p>
</div>
<div className="lhtb-table-tools">
<label><Search size={15} /><input value={query} onChange={(event) => setQuery(event.target.value)} placeholder={c.searchPlaceholder} /></label>
<label><Search size={15} /><input aria-label={c.searchPlaceholder} value={query} onChange={(event) => setQuery(event.target.value)} placeholder={c.searchPlaceholder} /></label>
<div className="lhtb-segments" aria-label={c.filterLabel}>
{(["all", ...baselines] as TableMode[]).map((mode) => (
<button aria-pressed={tableMode === mode} className={tableMode === mode ? "is-active" : undefined} onClick={() => setTableMode(mode)} type="button" key={mode}>
Expand All @@ -265,6 +311,14 @@
))}
</div>
</div>
<div className="lhtb-group-tools">
<label htmlFor="lhtb-group">{c.groupFilterLabel}</label>
<select id="lhtb-group" value={group} onChange={(event) => setGroup(event.target.value as GroupKey | "all")}>
<option value="all">{c.allGroups}</option>
{groupKeys.map((key) => <option value={key} key={key}>{c.groupLabels[key]}</option>)}
</select>
<button type="button" onClick={resetFilters}>{c.resetFilters}</button>
</div>
<button className="lhtb-history-toggle" type="button" aria-pressed={showHistory} onClick={() => setShowHistory(!showHistory)}>{showHistory ? c.hideHistory : c.showHistory}</button>
<div className="bm-table-wrap lhtb-task-table">
<table>
Expand All @@ -274,17 +328,18 @@
const best = Math.max(...visibleArms.map((arm) => row[arm]));
return (
<tr key={row.task}>
<th scope="row"><code>{row.task}</code></th>
<th scope="row"><a href={promptUrl(row.task)} target="_blank" rel="noreferrer"><code>{row.task}</code></a><span>{c.groupLabels[taskGroup.get(row.task)!]}</span></th>
{visibleArms.map((arm) => (
<td className={row[arm] === best ? "is-best" : undefined} key={arm}>{formatReward(row[arm])}</td>
))}
</tr>
);
})}
{visibleTasks.length === 0 && <tr><td colSpan={visibleArms.length + 1}>{c.emptyTasks}</td></tr>}
</tbody>
</table>
</div>
<p className="lhtb-visible-count">{c.visibleCount.replace("{count}", String(visibleTasks.length))}</p>
<p className="lhtb-visible-count" role="status">{c.visibleCount.replace("{count}", String(visibleTasks.length))}</p>

Check warning on line 342 in apps/presentation/site/src/LhtbBrief.tsx

View check run for this annotation

SonarQubeCloud / SonarCloud Code Analysis

Use <output> instead of the "status" role to ensure accessibility across all devices.

See more on https://sonarcloud.io/project/issues?id=huangruiteng_loopx&issues=AaC1-Q--I3lxymM_-KF4&open=AaC1-Q--I3lxymM_-KF4&pullRequest=4717
</section>

<section className="bm-section bm-shell lhtb-program" id="program">
Expand Down
23 changes: 23 additions & 0 deletions apps/presentation/site/src/lhtb-brief.css
Original file line number Diff line number Diff line change
Expand Up @@ -184,6 +184,29 @@
grid-column: 1 / -1;
}

.lhtb-group-analysis { grid-column: 1 / -1; min-width: 0; }
.lhtb-group-analysis > p { color: var(--bm-body); font-size: 14px; line-height: 1.7; }
.lhtb-group-table table { min-width: 820px; table-layout: fixed; }
.lhtb-group-table caption { padding: 12px; text-align: left; color: var(--bm-body); font-size: 12px; }
.lhtb-group-table th:first-child { width: 23%; }
.lhtb-group-table th:nth-child(2) { width: 7%; }
.lhtb-group-table th:nth-child(3), .lhtb-group-table th:nth-child(4) { width: 17%; }
.lhtb-group-table tbody th { white-space: normal; }
.lhtb-group-table td { font-size: 13px; }
.lhtb-group-table strong, .lhtb-group-table small { display: block; font-family: "Geist Mono", monospace; }
.lhtb-group-table small { margin-top: 4px; color: var(--bm-body); font-size: 11px; }
.lhtb-group-table a { display: inline-flex; align-items: center; min-height: 44px; }
.lhtb-group-tools { grid-column: 1 / -1; display: flex; flex-wrap: wrap; align-items: center; gap: 12px; font-size: 13px; }
.lhtb-group-tools select, .lhtb-group-tools button {
min-height: 44px; max-width: 100%; padding: 8px 12px; border: 1px solid var(--bm-border);
border-radius: 6px; color: var(--bm-ink); background: var(--bm-surface); font: inherit; cursor: pointer;
}
.lhtb-page a:focus-visible, .lhtb-page button:focus-visible, .lhtb-page select:focus-visible, .lhtb-page input:focus-visible, .lhtb-page summary:focus-visible {
outline: 2px solid var(--bm-blue); outline-offset: 3px;
}
.lhtb-case-details summary { line-height: 1.6; }
.lhtb-cases article > a { display: inline-flex; align-items: center; gap: 6px; min-height: 44px; font-size: 12px; }

.lhtb-cases {
grid-column: 1 / -1;
display: grid;
Expand Down
Loading
Loading