Archive DeepSWE v1 with reviewed placement and stable verification - #4502
huangruiteng merged 3 commits into
Conversation
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
原始 DeepSWE GPT xhigh v1 的实验代码与历史结果此前只存在于被取代的 PR 里(#4498 已关闭,本 PR 独立替代它),而 deprecate/benchmark-legacy/ 是"源代码考古"目录,没有字节身份契约。结果是:任何一次搬运都会被两类正常操作改写——premerge 门禁的 py_compile 会在归档目录里生成 __pycache__,core.autocrlf=true 的检出会把行尾变成 CRLF。两者都会让"原始 v1"变成无法复现的历史文本,后续再想核对当年的执行逻辑或结果表,就只能凭记忆重建。这个 PR 要做的是给这段历史一个可验证的家:目录冻结、身份可独立复核、并且明确声明它不参与当前 benchmark 执行。
改动思路
作者没有继续往 deprecate/benchmark-legacy/ 里塞,也没有把规则散落到三处,而是把放置规则收敛到唯一权威 benchmark/README.md#archive-placement:退役实现与独立研究包仍归 deprecate/benchmark-legacy/,而"明确标识、不可变、且保持惰性"的实验快照可以在 benchmark/ 下按版本目录存放,并附带四个准入条件(公开安全、伴随文档写清 provenance/依赖/已知缺陷、固定哈希加只读校验、不得被产品导入或自动执行)。AGENTS.md 相应从"历史版本一律放 deprecate/"改成指向该锚点,benchmark-toolkit README 与 docs/development/documentation-layout.md 也改成同一锚点,全文只有一处权威。验证侧则是一个只读检查器加一组回归测试:检查器不再依赖运行环境,而是重建 git 树对象(mode + 名字 + blob sha1)与固定 EXPECTED_TREE 比对,跳过字节码、编译全部 Python、再对两个 launcher 跑 bash -n;行尾问题交给 .gitattributes 的 benchmark/deepswe-gptxhigh-v1/** text eol=lf。
具体改动
benchmark/deepswe-gptxhigh-v1/:14 个原始文件按树1bc5d2b3…冻结,内容与可执行位不变;benchmark/deepswe-gptxhigh-versions.md把"SSH Goal / Codex CLI 数据已撤回待复验"的提示放在最前面,并写明 provenance(贡献者本地 commit 在上游不可达)、运行前置条件与遗留缺陷。benchmark/check_deepswe_v1.py:新增只读检查器,输出archive_tree、original_snapshot_preserved、benchmark_executed: false、standalone_runnable: false,明确不调用模型、不跑任务。benchmark/tests/test_deepswe_v1_archive.py:9 个用例覆盖"编译后身份不变""7 类篡改仍然拒绝""autocrlf 检出保持字节"。- 治理文本:
AGENTS.md、benchmark/README.md(新增## Archive placement)、loopx/capabilities/benchmark_toolkit/README.md、docs/development/documentation-layout.md四处改成同一权威锚点;.gitattributes新增 LF 规则。
关键代码讲解
benchmark/check_deepswe_v1.py:15 EXPECTED_TREE:唯一的身份常量。我在 head 上独立跑了git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1,得到的就是1bc5d2b3b74761a97d34ba3f3612e977fd610340,与常量一致——这条声明不依赖作者的自证。benchmark/check_deepswe_v1.py:18 check_archive:把目录重建成 git 风格树对象后比对。跳过__pycache__/与裸*.pyc(第 24–33 行)是让"编译后仍可校验"成立的关键分支;随后对每个.py执行compile()、对_RUNNER/_BOOTSTRAP两个内嵌常量单独编译、对.sh跑bash -n。我用独立探针验证了它的失败面:改字节、删源文件、加源文件、源文件被裸.pyc顶替,四种都在树比对处拒绝。benchmark/tests/test_deepswe_v1_archive.py:66 test_autocrlf_checkout_preserves_archive_bytes:这个用例带正控——先断言control.txt真的被转成 CRLF(第 85 行),再断言归档文件逐字节不变且树哈希仍等于EXPECTED_TREE。正控很重要,否则"行尾没变"可能只是因为测试环境根本没开 autocrlf。.gitattributes:5:benchmark/deepswe-gptxhigh-v1/** text eol=lf,把归档目录钉在 LF 上,普通文本仍可被转换。- 反例对照:父提交
5f18c63ad在"先编译再检查"场景下报unexpected archive entry: __pycache__,而 head 通过;也就是说最后一个提交修的是真实故障,不是把结论写进文档。
对主干的风险
范围与风险面本身很窄:22 个文件里没有产品运行路径、quota、scheduler、authority 或持久化状态的改动,归档也不会被任何产品模块 import。仓库原生校验我都跑了:python examples/docs-governance-smoke.py ok,9 个新用例通过,loopx canary premerge --from-git-diff 的 direct checks、10 个 catalog canary 与 public boundary 扫描零失败;canary 最终状态是 manual_review_required,因为它把 benchmark 敏感路径列成人工 hold——这是流程约束,不是缺陷,合并决定仍归 owner。公开边界方面,我对 14 个归档文件和说明文档做了本地路径、凭据、内网主机与原始日志特征扫描,只命中 --token-budget 这类参数名,没有私密材料。
两条非阻塞 P3:
- 裸
*.pyc被整体豁免(check_deepswe_v1.py:24-33)。我的探针里"额外加一个裸.pyc"仍然返回original_snapshot_preserved: true。文档写了"排除字节码",但没写后果:在 legacy 布局下,与源文件同名的.pyc(头部 mtime/size 匹配时)会遮蔽同名.py,也就是身份检查通过、但按说明去组装工作区的人可能加载到被换掉的模块。建议把豁免收窄到 premerge 真正会生成的__pycache__/*.pyc,或在文档里补一句裸.pyc的遮蔽风险。 - 冻结快照新增了常驻 CI 用例。
pyproject.toml只排除deprecate/,所以benchmark/tests/会被默认收集;而仓库策略说实验专属验证属于研究工作区,放置规则第 3 条也只要求"存在只读校验"。如果这是有意的 durable invariant guard,建议在归档说明里写明这个取舍,否则把它留成按需运行的检查命令更贴合策略。
English verdict: APPROVE — exact head 71cc156925a6a8fae26f2db0cea1129db8964b1a of #4502. The archive is byte-frozen (I reproduced tree 1bc5d2b3b74761a97d34ba3f3612e977fd610340 with git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1, independent of the author's checker), the placement rule has a single canonical owner that four other surfaces now point at, and the final commit fixes a real failure mode: the parent commit 5f18c63ad fails on a pre-merge-compiled copy with unexpected archive entry: __pycache__ while the head passes. Independent probes confirm the identity check fails closed on byte mutation, source removal, source addition and source replacement by a bare .pyc; the nine new tests pass; examples/docs-governance-smoke.py passes; loopx canary premerge --from-git-diff reports zero check failures across direct checks, ten catalog canaries and the public-boundary scan, with a single benchmark_sensitive manual hold that keeps the merge decision with the owner; and a local scan found no credentials, private configuration, raw task text or trajectories in the frozen files. Two non-blocking P3s: a deliberately tolerated bare .pyc can shadow a same-named archived module while the identity check still reports success (narrow the exclusion to __pycache__/*.pyc, or document the consequence), and the archive test is collected by default CI even though the policy routes experiment-specific validation to the research workspace. No product runtime, quota, scheduler, authority or persistence behaviour changes, and no benchmark was executed.
huangruiteng
left a comment
There was a problem hiding this comment.
动机
原始 DeepSWE GPT xhigh v1 的实验代码与历史结果此前只存在于被取代的 PR 里(#4498 已关闭,本 PR 独立替代它),而 deprecate/benchmark-legacy/ 是"源代码考古"目录,没有字节身份契约。结果是:任何一次搬运都会被两类正常操作改写——premerge 门禁的 py_compile 会在归档目录里生成 __pycache__,core.autocrlf=true 的检出会把行尾变成 CRLF。两者都会让"原始 v1"变成无法复现的历史文本,后续再想核对当年的执行逻辑或结果表,就只能凭记忆重建。这个 PR 要做的是给这段历史一个可验证的家:目录冻结、身份可独立复核、并且明确声明它不参与当前 benchmark 执行。
改动思路
作者没有继续往 deprecate/benchmark-legacy/ 里塞,也没有把规则散落到三处,而是把放置规则收敛到唯一权威 benchmark/README.md#archive-placement:退役实现与独立研究包仍归 deprecate/benchmark-legacy/,而"明确标识、不可变、且保持惰性"的实验快照可以在 benchmark/ 下按版本目录存放,并附带四个准入条件(公开安全、伴随文档写清 provenance/依赖/已知缺陷、固定哈希加只读校验、不得被产品导入或自动执行)。AGENTS.md 相应从"历史版本一律放 deprecate/"改成指向该锚点,benchmark-toolkit README 与 docs/development/documentation-layout.md 也改成同一锚点,全文只有一处权威。验证侧则是一个只读检查器加一组回归测试:检查器不再依赖运行环境,而是重建 git 树对象(mode + 名字 + blob sha1)与固定 EXPECTED_TREE 比对,跳过字节码、编译全部 Python、再对两个 launcher 跑 bash -n;行尾问题交给 .gitattributes 的 benchmark/deepswe-gptxhigh-v1/** text eol=lf。
具体改动
benchmark/deepswe-gptxhigh-v1/:14 个原始文件按树1bc5d2b3…冻结,内容与可执行位不变;benchmark/deepswe-gptxhigh-versions.md把"SSH Goal / Codex CLI 数据已撤回待复验"的提示放在最前面,并写明 provenance(贡献者本地 commit 在上游不可达)、运行前置条件与遗留缺陷。benchmark/check_deepswe_v1.py:新增只读检查器,输出archive_tree、original_snapshot_preserved、benchmark_executed: false、standalone_runnable: false,明确不调用模型、不跑任务。benchmark/tests/test_deepswe_v1_archive.py:9 个用例覆盖"编译后身份不变""7 类篡改仍然拒绝""autocrlf 检出保持字节"。- 治理文本:
AGENTS.md、benchmark/README.md(新增## Archive placement)、loopx/capabilities/benchmark_toolkit/README.md、docs/development/documentation-layout.md四处改成同一权威锚点;.gitattributes新增 LF 规则。
关键代码讲解
benchmark/check_deepswe_v1.py:15 EXPECTED_TREE:唯一的身份常量。我在 head 上独立跑了git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1,得到的就是1bc5d2b3b74761a97d34ba3f3612e977fd610340,与常量一致——这条声明不依赖作者的自证。benchmark/check_deepswe_v1.py:18 check_archive:把目录重建成 git 风格树对象后比对。跳过__pycache__/与裸*.pyc(第 24–33 行)是让"编译后仍可校验"成立的关键分支;随后对每个.py执行compile()、对_RUNNER/_BOOTSTRAP两个内嵌常量单独编译、对.sh跑bash -n。我用独立探针验证了它的失败面:改字节、删源文件、加源文件、源文件被裸.pyc顶替,四种都在树比对处拒绝。benchmark/tests/test_deepswe_v1_archive.py:66 test_autocrlf_checkout_preserves_archive_bytes:这个用例带正控——先断言control.txt真的被转成 CRLF(第 85 行),再断言归档文件逐字节不变且树哈希仍等于EXPECTED_TREE。正控很重要,否则"行尾没变"可能只是因为测试环境根本没开 autocrlf。.gitattributes:5:benchmark/deepswe-gptxhigh-v1/** text eol=lf,把归档目录钉在 LF 上,普通文本仍可被转换。- 反例对照:父提交
5f18c63ad在"先编译再检查"场景下报unexpected archive entry: __pycache__,而 head 通过;也就是说最后一个提交修的是真实故障,不是把结论写进文档。
对主干的风险
范围与风险面本身很窄:22 个文件里没有产品运行路径、quota、scheduler、authority 或持久化状态的改动,归档也不会被任何产品模块 import。仓库原生校验我都跑了:python examples/docs-governance-smoke.py ok,9 个新用例通过,loopx canary premerge --from-git-diff 的 direct checks、10 个 catalog canary 与 public boundary 扫描零失败;canary 最终状态是 manual_review_required,因为它把 benchmark 敏感路径列成人工 hold——这是流程约束,不是缺陷,合并决定仍归 owner。公开边界方面,我对 14 个归档文件和说明文档做了本地路径、凭据、内网主机与原始日志特征扫描,只命中 --token-budget 这类参数名,没有私密材料。
两条非阻塞 P3:
- 裸
*.pyc被整体豁免(check_deepswe_v1.py:24-33)。我的探针里"额外加一个裸.pyc"仍然返回original_snapshot_preserved: true。文档写了"排除字节码",但没写后果:在 legacy 布局下,与源文件同名的.pyc(头部 mtime/size 匹配时)会遮蔽同名.py,也就是身份检查通过、但按说明去组装工作区的人可能加载到被换掉的模块。建议把豁免收窄到 premerge 真正会生成的__pycache__/*.pyc,或在文档里补一句裸.pyc的遮蔽风险。 - 冻结快照新增了常驻 CI 用例。
pyproject.toml只排除deprecate/,所以benchmark/tests/会被默认收集;而仓库策略说实验专属验证属于研究工作区,放置规则第 3 条也只要求"存在只读校验"。如果这是有意的 durable invariant guard,建议在归档说明里写明这个取舍,否则把它留成按需运行的检查命令更贴合策略。
English verdict: APPROVE — exact head 71cc156925a6a8fae26f2db0cea1129db8964b1a of #4502. The archive is byte-frozen (I reproduced tree 1bc5d2b3b74761a97d34ba3f3612e977fd610340 with git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1, independent of the author's checker), the placement rule has a single canonical owner that four other surfaces now point at, and the final commit fixes a real failure mode: the parent commit 5f18c63ad fails on a pre-merge-compiled copy with unexpected archive entry: __pycache__ while the head passes. Independent probes confirm the identity check fails closed on byte mutation, source removal, source addition and source replacement by a bare .pyc; the nine new tests pass; examples/docs-governance-smoke.py passes; loopx canary premerge --from-git-diff reports zero check failures across direct checks, ten catalog canaries and the public-boundary scan, with a single benchmark_sensitive manual hold that keeps the merge decision with the owner; and a local scan found no credentials, private configuration, raw task text or trajectories in the frozen files. Two non-blocking P3s: a deliberately tolerated bare .pyc can shadow a same-named archived module while the identity check still reports success (narrow the exclusion to __pycache__/*.pyc, or document the consequence), and the archive test is collected by default CI even though the policy routes experiment-specific validation to the research workspace. No product runtime, quota, scheduler, authority or persistence behaviour changes, and no benchmark was executed.
我的整体评价
整体 APPROVE。归档本身的可验证性我是独立复现的:git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1 得到与固定常量一致的树 id,篡改探针(改字节、删源文件、加源文件、源文件被裸 .pyc 顶替)全部 fail closed,父提交在"先编译再检查"下失败而 head 通过,说明最后一个提交修的是真实故障而不是补文档。放置规则的治理改动是这次最需要盯的部分:它把仓库自有策略从"历史版本一律进 deprecate/"改成"不可变快照可按四条件留在 benchmark/",属于行为/政策变更,但披露面完整——PR 描述、AGENTS.md、benchmark/README.md#archive-placement、benchmark-toolkit README 与文档布局表指向同一锚点,docs-governance smoke 仍通过,没有留下第二份互相矛盾的规则。canary 的 benchmark_sensitive 人工 hold 把合并决定留在 owner 手里,这是流程而非缺陷。两条 P3(裸 .pyc 豁免带来的遮蔽风险、冻结快照用例进入默认 CI 收集)都不阻塞,但建议在文档里补上取舍说明。
huangruiteng
left a comment
There was a problem hiding this comment.
审查对象:4502@71cc156925a6a8fae26f2db0cea1129db8964b1a(已合并,合并后审计)。merge base f4ed58de9e44ff58e1a3e46731c520309e486e2f,22 个文件 +4866/-9,3 个 commit。
动机
DeepSWE GPT xhigh v1 的原始 harness 和 113 题历史结果原本只存在于贡献者本地工作区;仓库既没有"不可变快照"的落位规则,也没有任何办法核对"这就是当时那份字节"。旧规则只有一句"legacy runner 与 dated packet 归到 deprecate/benchmark-legacy/",而那个目录明确写着仅供 source archaeology、禁止 import 或扩展——它无法表达"这批字节就是那张历史结果表的记录"。结果是:历史数字要么无法被引用,要么只能靠一个上游不存在的 commit 号自证。
改动思路
把"落位"和"身份"分开:benchmark/README.md#archive-placement 成为唯一的落位规则(两类归档:退休实现/dated packet 去 deprecate/benchmark-legacy/;命名且不可变的实验快照可在四个条件满足时留在 benchmark/),再由一个只读 checker 用 Git tree hash 固化身份,并用一个随 CI 收集的测试模块守住它;配套的 deepswe-gptxhigh-versions.md 负责版本、provenance 边界、运行依赖与已知缺陷,且把撤回提示放在最前面。
具体改动
- 新增
benchmark/deepswe-gptxhigh-v1/(14 个文件,约 4600 行,属冻结载荷)与.gitattributes的benchmark/deepswe-gptxhigh-v1/** text eol=lf。 - 新增
benchmark/check_deepswe_v1.py(84 行,只读):逐文件按mode + name + blob组成 Git tree 记录并与EXPECTED_TREE比对;对.py做compile()、对两个内嵌_RUNNER/_BOOTSTRAP常量单独编译、对两个.sh跑bash -n;排除编译产物,拒绝额外源码、目录、符号链接。 - 新增
benchmark/tests/test_deepswe_v1_archive.py(9 个测试):编译后身份不变、7 类变异 fail-closed、真实core.autocrlf=true检出模拟(同时断言普通文件确实变成 CRLF)。 - 新增
benchmark/deepswe-gptxhigh-versions.md(121 行,撤回提示置首);benchmark/README.md新增## Archive placement并把规则 3 改写;AGENTS.md、docs/development/documentation-layout.md、loopx/capabilities/benchmark_toolkit/README.md改为指向该锚点。
我做的独立核对:
- 身份:
git rev-parse 71cc15692:benchmark/deepswe-gptxhigh-v1=1bc5d2b3b74761a97d34ba3f3612e977fd610340,与文档与 checker 的EXPECTED_TREE完全一致(tree 对象同时证明 5 个 100755、9 个 100644);python3 benchmark/check_deepswe_v1.py→ok:true,11 个 python、2 个内嵌、2 个 shell 检查,benchmark_executed:false。 - 变异:追加 1 字节、改执行位、加
extra.py、加符号链接、__pycache__放非字节码 → 全部按预期拒绝;__pycache__仅放.pyc与顶层散落.pyc→ 通过(有意设计,见保留意见)。 - 仓库检查:
pytest benchmark/tests/test_deepswe_v1_archive.py→ 9 passed;examples/docs-governance-smoke.py→ ok;git diff --check干净;git check-attr确认为text:set/eol:lf;pyproject的norecursedirs只排除deprecate,CI 又是全量收集分片,因此该守卫测试确实会在 CI 跑;3 个 commit 均带 DCO。 - 边界与惰性:归档目录仅被 README、版本说明、checker 三处引用,无任何 import;归档文件内无绝对路径、无内部域名、无凭证,只有
127.0.0.1占位符与公共镜像地址。
对主干的风险
该快照不参与产品运行时、权限、配额或调度,失败模式是"字节被悄悄改动而无人发现",而这一点已被 CI 中的 tree hash 断言挡住,.gitattributes 也消掉了最常见的意外改动来源(CRLF 归一)。三点保留意见(均 P3、非阻塞):(1) 归档中 loopx_turn_runner.py 与现役 benchmark/swe-marathon/runtime/turn/loopx_turn_runner.py 字节相同,另外 3 个同名文件已分叉,但没有任何文档说明"现役实现是 v1 的后继",读者会看到同一份文件的两处副本而不知道谁当前有效;(2) checker 忽略任意位置的 .pyc 而非仅 __pycache__,所以归档文件集合并没有被枚举,被 -f 强加入的 .pyc 不会改变身份;(3) pyproject.toml 的 norecursedirs 只排除 deprecate,冻结目录位于 pytest 默认收集树内——当前没有 test_*.py 或 conftest.py 所以无事,但后续快照若含此类文件会被收集甚至 import,与"惰性"条件冲突。另外,快照本身不构成任何实验结论:来源 commit 在上游不可达、未重跑、未复验,SWE Marathon 已撤回的 SSH Goal/Codex CLI 行仍是历史文本——这些边界文档都写明了,我认同这种"只对字节负责"的克制。
我的整体评价
这是一次把"归档"从口头约定变成可验证契约的改动:身份用 Git tree hash 钉死、篡改与夹带 fail-closed、编译产物与 LF 检出这两个最容易造成假阳性的来源被显式处理并测到,落位规则写在唯一 owner 文档里、其他三处只做指针,撤回提示置于最前且未对历史结果做任何新主张。我独立复算了 tree hash、跑了 checker 与 9 个测试、并自己构造了 7 组变异,结论与文档一致。剩下的三点都是后继可读性与枚举严格性的小修,不影响本次交付成立。作为已合并精确 head 的合并后审计,证据支持通过。
English verdict: APPROVE (exact head 71cc156)
Archive the original DeepSWE GPT xhigh v1 harness and historical 113-task results, with the maintainer-requested placement policy and environment-stable identity verification. This new PR replaces #4498 and carries its completed fixes on current
main.本 PR 独立替代 #4498,包含完整归档及维护者要求的修复。原始 v1 的 14 个文件、执行逻辑和结果数字不变,不包含正在运行的新 benchmark 或
v1-revised。Changes:
benchmark/deepswe-gptxhigh-v1/exactly at tree1bc5d2b3b74761a97d34ba3f3612e977fd610340.benchmark/README.md#archive-placementthe canonical placement rule, with AGENTS and related documentation referring to it..gitattributes, including undercore.autocrlf=true.Validation on the new base: 9 focused regression tests pass (compile-before-check, bytecode handling, mutation rejection, and actual Git autocrlf checkout simulation); the archive checker, docs governance smoke, and whitespace checks pass. All commits have DCO sign-offs. The branch was rebuilt on current main without conflicts. No native Windows run or full premerge run is claimed.
This is an inert historical archive, not a standalone runnable benchmark release. External dependencies and original execution defects remain documented. No model call, benchmark rerun, new score, or revalidation of historical results is claimed.