docs(lhtb): show task-type gains and counterexamples against both baselines - #4717
Conversation
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed head: 93461e75441d9a5bcdfe540ad951c83d977f4aef · base: 42620170d
无阻塞性发现。以下结论覆盖整个 PR,不把探索性分组当作显著性或因果结果。
动机
原页面有总均分和个案,缺少可检查的题型分布,容易把少数大幅收益概括成普遍优势。本次完成九组、46 题的完整分析,并把精简结论接入中英文应用场景博客。公开数据和运行方式均未改变。
改动思路
沿用现有 LHTB 页面、数据源、矩阵和本地化机制。task-groups.json 只维护明确标为事后分析的归类与题目来源,均分差和胜平负从原始分数派生。官方介绍、README 和 task metadata 没有一致的逐题分类,因此不冒充官方类别榜单。
正向路径是点击“研究复现与建模”后清除其它筛选,显示全部四道成员及原文。反向路径是将 Tabular 查询与多模态组组合:应显示空结果并能重置,不能残留误导性的行。两条路径均在真实浏览器验证。相比另建分析系统,这里扩展既有组件更合适;仅补案例文字则无法核查完整成员。
具体改动
关键代码讲解
LhtbBrief.compareTasks:将原总表的配对差计算复用于分组;严格按未舍入分数计胜平负,只在显示时保留四位小数。visibleTasks:新增明确的 GroupKey 筛选,与原查询和 ≥0.05 差值筛选取交集;重置恢复 46 题,历史两臂保持可展开。promptUrl:使用固定上游 revision,并处理 Rush Hour、Snake 的两个目录别名。矩阵和案例共用此来源逻辑。test_lhtb_task_groups_partition_the_published_study_and_have_prompt_links:验证所有任务恰好归类一次、双语键完整、案例属于实验范围及 source pin 格式;临时重复成员的反例能触发失败。
其它修改包括复用设计 token 的表格/控件样式、双语分析与折叠案例、完整分组方法和敏感性说明,以及两篇博客的对应摘要。保留旧 #insights 锚点。原 data.json 无差异;未提交生成物、原始轨迹、题目全文或本地状态。
对主干的风险
最大风险是解释过度。研究组去掉 Tabular 后仍对两侧正向;谜题组留一法方向不变。但系统组对 Plain 的收益由 PoC 拉动,对 Goal 的损失由 DuckDB 主导;Snake、法律事项的基线反转也被保留。多模态均分接近,不能据此定位能力瓶颈的实际成因。
页面明确写出每组仅 4–7 题、每格一次有效运行、时限不一致,以及当前 pinned prompt 未证实与实验的 July snapshot 逐字相同。任务要求只能形成解释假设。若要声称稳定优势或机制归因,仍需同预算重复试验和轨迹分析。
验证覆盖:TypeScript/Vite 构建;publication-scope pytest 5 passed;design baseline、双语 blog、静态 export 三项 smoke;Python 独立计算与浏览器中九组/18 对照逐项一致;实际 base/head 页面 46×3 分数、Sudoku 查询、Goal 差值筛选与历史分数一致;桌面/手机、中英文、空结果/重置、分组、原文别名、案例展开和键盘焦点通过。静态发布验证同时覆盖 / 与 /loopx/ 路径。
初次 export 因独立 worktree 缺依赖及系统 Python 版本失败,补齐 npm 依赖并使用项目 uv 环境后通过;一次自动点击被 sticky header 遮挡,居中目标后点击通过。没有未解决的本地必需检查或人工 hold。上述集合构成本次 presentation/publication 变更的风险对应 premerge 验证;没有更改 runtime、quota、权限或评分,无需启动 benchmark。CI 按已解析的 wait_for_ci=false 不查询、不等待。
我的整体评价
APPROVE。复用现有 reducer 和矩阵,归类数据与评分数据分离,完整成员和反例让解释可核验。修改对当前问题范围适当,未引入新的运行期权威或通用控制面术语;无 default-off 能力或自动加载指令变化。用户已明确授权本次自合并并豁免首屏确认;合并仍须对同一 head 通过即时 readiness 检查。
English verdict: APPROVE — 93461e75441d9a5bcdfe540ad951c83d977f4aef. Exhaustive exploratory grouping with both baselines, counterexamples and source-version limits; original rewards are unchanged. Build, five publication tests, design/blog/export smokes, independent numeric checks and real base/head browser parity passed. No blocking finding; causal and equal-budget generalization remain unproven. CI was not consulted under the resolved policy.
LHTB’s overall mean hides where LoopX gains over Plain and native Goal. Add an exploratory task-type analysis to the bilingual research brief and application-scenarios blog: all 46 tasks in nine explicit post-hoc groups, paired mean deltas and W/T/L, source-linked drilldown, and concrete counterexamples.
Research/modeling and puzzles show positive group means against both baselines; multimodal/scientific tasks show little separation. The copy discloses single-trial and unequal-budget limits, baseline-sensitive outliers, leave-one-out sensitivity, and the unverified identity between current pinned prompts and the evaluated snapshot. Published rewards, scoring and execution are unchanged.
Validation: site TypeScript/Vite build; publication-scope pytest (5 passed); design and bilingual-blog smokes; static export smoke; independent Python-to-browser checks of all nine memberships and 18 comparisons; desktop/mobile EN/ZH filters, empty/reset states, history, prompt aliases and case expansion. Owner authorized self-merge and waived the first-screen preview gate; CI waiting is disabled for this review, with local checks and exact-head review still required.