Add an Appendix E incident retrodiction harness - #4621
Conversation
The RFC's Appendix E collects 15 lessons from real escapes, but nothing answers the question they were written for: does the guard as it exists today reject the defect form each lesson describes? Recorded lessons without a reproducing probe decay into folklore, and the appendix's own E10 lesson says that a corpus exercised only locally is not a test. `loopx/semantics/incident_v0.json` pairs each lesson with a probe, the rule it expects, and an `expectation` authored from the lesson text. `scripts/ semantic_incident_retrodiction.py` drives the real pre-merge guard for each one, plus five negative controls: legal variations the guard must accept, which is what stops a fail-always check from scoring as coverage. A retrodiction asks whether today's guard rejects the defect form, never whether a historical revision passed; the guard did not exist then, so that answer would be no for all 15 and would measure nothing. Measured at 3937f6b: would_have_caught=12/15, negative_control_false_positives=0/5 E02 uncovered no source outside loopx/ is scanned (declared root) E12 uncovered dropping 83 sources under loopx/extensions/ was unnoticed by every reach-consuming check; upper-bound budgets cannot see a narrowed reach, so the declared scope is unverified E14 uncovered 11 open decisions, no check enforces the inputs they name E12 and E14 are unstated gaps rather than declared boundaries. Recording them is the point: the ledger is a measurement, and the two gaps are the candidate work, not a silent pass. Probes fault-inject an input and then call the unmodified check, which requires the guard's real globals dict. `runpy.run_path` returns a copy, so an injection there is a silent no-op that reports "uncovered" without measuring anything; the module is imported instead, and every injection re-reads its own effect before its verdict is trusted. The registry-file probe additionally reads the committed registry back through the patched path before mutating. The negative controls pin the other direction, including that the equality anchors accept the committed registry and that the identifier-counting ratchet does not count a longer identifier containing the field token. tests/architecture/test_semantic_incident_retrodiction.py makes the ledger a committed check and freezes the baseline, so a lesson that was caught stays caught. Signed-off-by: song <liusongstep@gmail.com>
补充:全局定位与本 PR 的必要性本节是为审阅者补的上下文。PR body 是变更本身,这里是它在一整盘计划里的位置,以及为什么它必须先做。 一、全局定位唯一目标源是 RFC( 目标态一句话:把语义分歧的发现点,从生产环境左移到提交时。 指标结构(四索引 + 一层 + 四字段):
字段(不是索引):覆盖率、误报率、每次改动时延、最老未决决策年龄。 N1 批次(四件事,互不依赖,各自可独立回滚):
当前可核对快照(base
二、本 PR 的必要性按必要性六维逐条(生产者存在 / 判别性 / 多值可达 / 不可推导 / 消费者存在 / 边界角色):
它不买什么(明确排除):
结论: 三、验证与状态
四、附带发现(与代码无关但阻塞流程)
|
…trodiction Signed-off-by: song <22676124+songoow@users.noreply.github.com>
exact-head 复核(
|
…trodiction Signed-off-by: song <22676124+songoow@users.noreply.github.com>
本 PR 在 #4447 计划中的位置issue #4447 现在有一节统一协调(中英双语),把这 13 个在开 PR 作为一个计划列出:各自修什么、为何必要、以及实测出的合并顺序。 冲突实测:对全部 78 对做了试合并,9 对冲突,分四簇,每一处都是文本相邻,没有一处是语义分歧。
建议顺序(代价从低到高):#4628 → #4625、#4626 → #4627 → #4619、#4621 → #4630 → #4614 → #4631 → #4629 → #4617 → #4606 → #4608。四个棘轮 PR 放最后,因为每落地一个,下一个的数字就从估算变成确定值。 全部 13 个 PR 现已同步到 |
…trodiction Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: fb9a5fad58f2256057ee044a1ed030d96c38d065 (codex/n1b-incident-retrodiction).
动机
RFC 的 Appendix E 收了 15 条来自真实逃逸的教训,但它自己列的验收面(probe)从来没有落地:仓库里没有任何东西回答「今天的 guard 还认不认这条缺陷形态」。教训只以散文存在,就会在 guard 漂移时变成 folklore——而且这正是 Appendix E 的 E10 教训本身说的:只在本机跑过的语料不算测试。
改动后这个问题有了提交物:loopx/semantics/incident_v0.json 把每条教训、期望规则、probe 和 expectation 配起来,scripts/semantic_incident_retrodiction.py 驱动真实 guard,tests/architecture/test_semantic_incident_retrodiction.py 把它变成 CI 内的检查并冻结基线。这是一个完整的、可独立复核的切片:它回答了「现状是否还拦住」这个问题,并明确停下、不去顺手改 guard 本身(那是另一个该单独 review 的变更)。
改动思路
入口只有一个:脚本的 main();测试用 sys.executable 把同一个脚本当子进程跑,所以这份测量真的被 CI 执行,不是无人调用的脚手架。权威输入是 ledger 加 guard 模块本身——Guard.__init__ 用 importlib.util.spec_from_file_location 把 examples/semantic-vocabulary-drift-smoke.py 导入成模块,拿到它真实的 globals,因此 probe 才能做故障注入;docstring 里解释得很对:runpy.run_path 返回的是副本,在那里注入是静默 no-op,会报「uncovered」却什么都没测。判定边界在 INCIDENT_PROBES / CONTROL_PROBES 里每个小函数:注入一个输入,然后调用未改动的 check 函数,看它是否拒绝。副作用为空——registry 只改内存深拷贝,唯一的临时文件走 tempfile.NamedTemporaryFile 并在 finally 里删掉。
复用面我确认过:guard 的谓词(check_coverage_floor、check_owned_vocabularies、check_scope_declarations、check_relations、evaluate_inventory_budget_findings、collect_literal_uses、count_identifier_modules)全部是调用原模块,没有第二份实现;loopx/semantics/ 下已经住着 vocabulary_v0.json / inventory_v0.json,ledger 放这里、且 load_ledger 要求 vocabulary_v0.json 就在旁边,符合既有形状。五条 negative control(guard 必须接受的合法变体)是这份测量最关键的设计:没有它们,一个永远报错的检查也能刷满覆盖率。
具体改动
3 个新文件、+789/-0,没有任何既有文件被改:ledger 166 行、driver 566 行、测试 57 行。
我在这个 head 上独立跑了 python scripts/semantic_incident_retrodiction.py,结果与作者声明一致:appendix_e_lessons=15、would_have_caught=12/15、negative_control_false_positives=0/5、caught_baseline=12/15,未被拦住的是 E02 / E12 / E14。pytest tests/architecture/test_semantic_incident_retrodiction.py -q → 2 passed;ruff check 两个新文件 → All checks passed。ledger 每条 defect 文本与 RFC Appendix E 的条目逐条对得上,measured_at: "3937f6b99" 是真实存在的合并提交。
关键代码讲解
Guard.__init__/Guard.with_registry_file(脚本 82 起):前者以模块方式导入 guard,保证注入落在真实 globals 上;后者在改 registry 内容前,先把已提交的 registry 写进同一个临时路径并读回比对,注入不是活的就报错,而不是默默判定。这是整份 harness 的可信度基础,做对了。probe_registry_weakness_rejected/probe_scan_reach_disclosed(242 起):故障注入的代表。前者同时验证「floor 被调低」「owner 不是 module::Symbol」「扫描根被收窄」三条都被拒;后者把load_sources换成收窄树(丢掉loopx/extensions/下 83 个源文件),再跑三个吃 reach 的检查——全部没察觉,于是 E12 如实记为 uncovered,而不是假装覆盖。main()(500 起):先对账「Appendix E 的条目数 == ledger 的 incidents 数」(行数变了就失败,不靠人记),再校验每个 probe 名可解析,然后统计 caught / control violations 并逐条打印;--report补打印 defect/rule 文本,--json给机器读。tests/architecture/test_semantic_incident_retrodiction.py:41:用sys.executable跑子进程并钉住打印出来的计数,符合仓库「子进程沿用选定解释器」的规则。loopx/semantics/incident_v0.json:load_ledger用精确 key-set 校验(LEDGER_KEYS/INCIDENT_KEYS/CONTROL_KEYS)加VERDICTS枚举,expectation只允许caught/uncovered,非法状态很难表达。
对主干的风险
最强回归不是崩溃,而是「verdict 与它声称测量的东西脱钩」:读者把 uncovered(或将来的 caught)当成 guard 的行为,而 probe 其实没问 guard。我在这条思路上做了两个反例。
一条非阻塞 P2(F1):E02 的 probe 结构上是常量,不是观测。 probe_fixture_copy_visible(188-192)用 guard.sources 里不以 loopx/ 开头的路径来判定,而 Guard.__init__ 调的是 load_sources(REPO_ROOT),默认 root 就是 loopx,路径又是 relative_to(repo_root) 生成的。我在这个 head 上量到 1206 个源文件、其中 0 个在 loopx/ 之外;即使把 guard.ns["LITERAL_SCAN_ROOTS"] 改成 ["loopx","tests","examples"],probe 仍然返回 False。也就是说 sources_outside=0 是构造出来的,E02 记成 uncovered 从未问过 guard 任何问题。建议让这个 probe 真的去问 guard:要么把 guard 声明的扫描根与已知 fixture 副本位置(tests/**、examples/**)对比并报出未覆盖的边界,要么往被扫描的根里注入一份合成 fixture 副本,看 check_literal_vocabularies 是否报出 fork,然后按观测重新写 expectation。
一条非阻塞 P2(F2):冻结基线把 PR 自己承诺的下一步变成红灯。 测试断言的是相等(tests/architecture/test_semantic_incident_retrodiction.py:53:(caught,total) == (baseline.would_have_caught, baseline.total)),失败信息是「a lesson that was caught must stay caught」;而 harness 自己只要求 caught >= baseline(脚本 553),PR 正文也把 E12/E14 称作 candidate work。我在这个 head 上复现了这个冲突:新增一个未跟踪的 loopx/canary/*.py,内容只是一个 open_decision 标识符,测量立刻变成 13/15——脚本打印 would_have_caught=13/15 加一行非致命的 expectation-mismatch、退出码 0,而提交的测试以 assert (13, 15) == (12, 15) 失败,报的还是「倒退」文案。也就是说按 PR 宣称的方向去补覆盖会先撞 CI,而真正需要跟进的 mismatch 只被打印、没有义务。建议测试改成 caught >= baseline 且 total == len(incidents),并让 expectation 变化时必须伴随 ledger 更新(或让 harness 对 mismatch 直接失败),这样规则只有一处表达。
一条 P3(F3,非阻塞):E14 是子串判定,且 ledger 的 meaning 声明了它没有的性质。 probe_dangling_decision 用 "open_decision" in 拼接 loopx/canary/*.py 的文本判定,所以任何只是提到这个名字的模块都能把它刷成 caught(上面那个临时文件就产生了 check_enforcing_named_inputs=True、E14 caught),而没有任何检查真的读了 open decisions 的输入。另外 ledger 的 meaning 写着 expectation 只从教训文本撰写、绝不来自 probe 输出,但 Appendix E 里没有任何一句能推出 E02/E12/E14 会被漏掉——这三个 uncovered 只能是读测量结果写的。建议 E14 改成匹配一个可解析的具名输入检查,并把 meaning 改写成「作者依据教训文本 + 首次测量记录」或把这三条明确标为 declared gap。
语义与 CI 对齐
这份变更复用既有语义面而不是造新的:没有新增产品词汇,唯一的两个新名字是 ledger 的 schema 版本和它自己的文件名,且都落在 loopx/semantics/ 既有目录里;guard 依旧是唯一的判定权威,ledger 只记录观测。语义上需要跟进的正是 F3 里那句 meaning 声明与文件实际内容的偏差,以及被冻结的 expectation。
我的整体评价
结论 APPROVE。这是一个有真实价值的增量:它把 Appendix E 从散文变成可重跑的测量,用真实 guard 而不是复刻实现,用五条 negative control 挡住「永远报错也算覆盖」,并把「先证明注入是活的再信 verdict」这条做对了——这些正是同类 harness 最容易糊弄的地方。我独立复现了作者声明的全部数字(12/15、0/5、E02/E12/E14 uncovered),核对了两条测试与 ruff,也确认 diff 只新增三个文件、没有碰任何既有路径。回退成本是一个 commit。
同时我把这份测量的可信度边界说清楚:15 条 verdict 里有 2 条(E02、E14)不是对 guard 的观测,冻结基线又把 PR 自己列的 candidate work 变成红灯,meaning 的独立撰写声明与实际内容不符。这三条都是新文件内部的可维护性问题,非阻塞,不改变本次结论;但在这个 harness 被用来判断「哪条教训还被拦住」之前,值得先修掉——否则它会以很高的话术质量给出偏乐观或偏悲观的读数。
English verdict: APPROVE - exact head fb9a5fa; I reproduced the claimed measurement independently (would_have_caught=12/15, negative_control_false_positives=0/5, E02/E12/E14 uncovered, 2 architecture tests passed, ruff clean, three new files and no existing file modified), and the harness genuinely drives the real pre-merge guard through its own globals with five negative controls. Three non-blocking findings, all inside the new files: the E02 probe's verdict is structurally constant because it filters an already root-limited source list instead of observing the guard; the committed test asserts equality against the frozen baseline, so the E12/E14 improvement the PR advertises fails CI with a regression message while the harness only prints the mismatch and exits 0; and E14 decides with a substring over loopx/canary/*.py while the ledger's meaning claims an expectation authorship the three uncovered rows do not have.
What
The RFC's Appendix E collects 15 lessons from real escapes, but nothing answers the question they were written for: does the guard as it exists today reject the defect form each lesson describes? Recorded lessons without a reproducing probe decay into folklore, and the appendix's own E10 lesson says a corpus exercised only locally is not a test.
This adds the missing measurement:
loopx/semantics/incident_v0.json— pairs each lesson with a probe, the rule it expects, and anexpectationauthored from the lesson text.scripts/semantic_incident_retrodiction.py— drives the real pre-merge guard for each lesson, plus five negative controls (legal variations the guard must accept, which is what stops a fail-always check from scoring as coverage).tests/architecture/test_semantic_incident_retrodiction.py— makes the ledger a committed check and freezes the baseline.A retrodiction asks whether today's guard rejects the defect form, never whether a historical revision passed. The guard did not exist then, so that answer would be "no" for all 15 and would measure nothing. The ledger's
meaningfield states this.Measured at
3937f6b99E12 and E14 are unstated gaps, not declared boundaries. Recording them is the point of the harness: the ledger is a measurement, and those two are candidate work rather than a silent pass.
Probe honesty
Probes fault-inject an input and then call the unmodified check, which requires the guard's real globals dict.
runpy.run_pathreturns a copy, so an injection there is a silent no-op that would report "uncovered" without measuring anything — the module is imported instead, and every injection re-reads its own effect before its verdict is trusted. The registry-file probe reads the committed registry back through the patched path before mutating.Validation
python examples/semantic-vocabulary-drift-smoke.py --report— rc 0python -m pytest -q tests/architecture— 290 passed in 80sruff checkon both new files — cleanNotes