Skip to content

Add an Appendix E incident retrodiction harness - #4621

Merged
huangruiteng merged 4 commits into
loopx-project:mainfrom
songoow:codex/n1b-incident-retrodiction
Sep 17, 2026
Merged

huangruiteng merged 4 commits into
loopx-project:mainfrom
songoow:codex/n1b-incident-retrodiction

Conversation

@songoow

@songoow songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

What

The RFC's Appendix E collects 15 lessons from real escapes, but nothing answers the question they were written for: does the guard as it exists today reject the defect form each lesson describes? Recorded lessons without a reproducing probe decay into folklore, and the appendix's own E10 lesson says a corpus exercised only locally is not a test.

This adds the missing measurement:

  • loopx/semantics/incident_v0.json — pairs each lesson with a probe, the rule it expects, and an expectation authored from the lesson text.
  • scripts/semantic_incident_retrodiction.py — drives the real pre-merge guard for each lesson, plus five negative controls (legal variations the guard must accept, which is what stops a fail-always check from scoring as coverage).
  • tests/architecture/test_semantic_incident_retrodiction.py — makes the ledger a committed check and freezes the baseline.

A retrodiction asks whether today's guard rejects the defect form, never whether a historical revision passed. The guard did not exist then, so that answer would be "no" for all 15 and would measure nothing. The ledger's meaning field states this.

Measured at 3937f6b99

would_have_caught=12/15, negative_control_false_positives=0/5
E02 uncovered  no source outside loopx/ is scanned (declared root)
E12 uncovered  dropping 83 sources under loopx/extensions/ was unnoticed by every
               reach-consuming check; upper-bound budgets cannot see a narrowed
               reach, so the declared scope is unverified
E14 uncovered  11 open decisions, no check enforces the inputs they name

E12 and E14 are unstated gaps, not declared boundaries. Recording them is the point of the harness: the ledger is a measurement, and those two are candidate work rather than a silent pass.

Probe honesty

Probes fault-inject an input and then call the unmodified check, which requires the guard's real globals dict. runpy.run_path returns a copy, so an injection there is a silent no-op that would report "uncovered" without measuring anything — the module is imported instead, and every injection re-reads its own effect before its verdict is trusted. The registry-file probe reads the committed registry back through the patched path before mutating.

Validation

  • python examples/semantic-vocabulary-drift-smoke.py --report — rc 0
  • python -m pytest -q tests/architecture — 290 passed in 80s
  • ruff check on both new files — clean

Notes

  • Report-only: this is evidence, not a new gate on product code. The committed test guards the harness's own ratchet.
  • No existing file is modified; the diff is three new files.

The RFC's Appendix E collects 15 lessons from real escapes, but nothing
answers the question they were written for: does the guard as it exists today
reject the defect form each lesson describes? Recorded lessons without a
reproducing probe decay into folklore, and the appendix's own E10 lesson says
that a corpus exercised only locally is not a test.

`loopx/semantics/incident_v0.json` pairs each lesson with a probe, the rule it
expects, and an `expectation` authored from the lesson text. `scripts/
semantic_incident_retrodiction.py` drives the real pre-merge guard for each
one, plus five negative controls: legal variations the guard must accept, which
is what stops a fail-always check from scoring as coverage.

A retrodiction asks whether today's guard rejects the defect form, never whether
a historical revision passed; the guard did not exist then, so that answer
would be no for all 15 and would measure nothing.

Measured at 3937f6b:

  would_have_caught=12/15, negative_control_false_positives=0/5
  E02 uncovered  no source outside loopx/ is scanned (declared root)
  E12 uncovered  dropping 83 sources under loopx/extensions/ was unnoticed by
                 every reach-consuming check; upper-bound budgets cannot see a
                 narrowed reach, so the declared scope is unverified
  E14 uncovered  11 open decisions, no check enforces the inputs they name

E12 and E14 are unstated gaps rather than declared boundaries. Recording them
is the point: the ledger is a measurement, and the two gaps are the candidate
work, not a silent pass.

Probes fault-inject an input and then call the unmodified check, which requires
the guard's real globals dict. `runpy.run_path` returns a copy, so an injection
there is a silent no-op that reports "uncovered" without measuring anything;
the module is imported instead, and every injection re-reads its own effect
before its verdict is trusted. The registry-file probe additionally reads the
committed registry back through the patched path before mutating.

The negative controls pin the other direction, including that the equality
anchors accept the committed registry and that the identifier-counting ratchet
does not count a longer identifier containing the field token.

tests/architecture/test_semantic_incident_retrodiction.py makes the ledger a
committed check and freezes the baseline, so a lesson that was caught stays
caught.

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

补充:全局定位与本 PR 的必要性

本节是为审阅者补的上下文。PR body 是变更本身,这里是它在一整盘计划里的位置,以及为什么它必须先做。

一、全局定位

唯一目标源是 RFC(docs/architecture/rfcs/semantic-vocabulary-convergence-v0.md / .zh-CN.md)。不新建"总体方案"文档——否则"该做什么"会同时存在 RFC §11、新方案文档、LoopX todo 三个权威,而这正是本工作流要消灭的失效模式。人的叙述放 .local/reviews/(已 gitignore)。

目标态一句话:把语义分歧的发现点,从生产环境左移到提交时。

指标结构(四索引 + 一层 + 四字段):

层 内容 今天
S 自洽 声明 = 强制 = 测量 ≈0%:声明覆盖 enforced_by 0/6;3/3 虚假强制声明被接受;盲区披露率 0%
I 自我迭代 G1 变异检出 × G2 分类覆盖 × E1 回测命中(乘法,任一为 0 则环路为 0) ≈0:G1 未进 CI、G2 0/38、E1 未测
退役轴 R 有退出条件的检查占比 × C 理由(C1 误报率 / C2 每次改动时延 / C3 挡下缺陷的成本) R ≈1/4;C 未知
U 用户可见影响 未披露静默替换、operator-visible parity、逃逸到用户的缺陷数 不可计算,必须显式声明 unknown
D 债务存量 单独一层,永不与工具层平均 见下

字段(不是索引):覆盖率、误报率、每次改动时延、最老未决决策年龄。

N1 批次(四件事,互不依赖,各自可独立回滚):

  • N1-a 存量收尾:3 个 Track A 分支补 commit body 边界说明;冲突预算继续降(conflicting_values_semantic 2→0 已做)
  • N1-b → 本 PR:E1 回测
  • N1-c 观测层披露 + AST 缓存(report-only,永不升门禁)
  • N1-d RFC 规范性修订(en/zh 同步、Last normative revision 提升、Appendix B 一条决策)——唯一需要 owner 批准的步骤,不阻塞 a/b/c

当前可核对快照(base 3937f6b99):

  • RFC 附录 E = 15 条真实事故教训
  • 合并候选组 38 = 18 个 registry 已声明双 owner(工具假阳性,merge_candidate_groups 从不查 registry 自己的 owner 对)+ 2 个未注册跨运行时对 + 18 个纯 Python 组
  • producer 扫描实际触达 432/1203 = 35%
  • 违反"单一 owner"的提交:0/1420(6月)→ 0/1818(7月)→ 0/1127(8月)→ 8/1036(9月)
  • 4 个预算落后于实际(约 9 单位静默回退空间)
  • smoke 18.5s,其中 6,186 次 ast.parse / 1,033 个不同源文件(83% 重复)
  • 引入率 8/1036(9月)vs 0/4365(前 3 个月)→ 无统计功效,不得当证据使用

二、本 PR 的必要性

按必要性六维逐条(生产者存在 / 判别性 / 多值可达 / 不可推导 / 消费者存在 / 边界角色):

  1. 消费者存在。E1 的输出被 M2/M3 排期、S 指标 S2(声称可信度)、I 指标 G1 直接消费。没有它,"这个守卫到底强制了什么"只能靠读代码猜。
  2. 不可推导。12/15 无法由代码阅读得出。可以读到"存在这个检查",但"这个检查是否真的拒绝该缺陷形式"只有故障注入能回答。本次实测就踩到了反例:runpy.run_path 返回的是 globals 副本,注入静默失效,探针会假报"未覆盖"——不跑就不知道。
  3. 判别性。12/15 不是全绿也不是全红。三条未覆盖各有实测原因:
    • E02 未覆盖:loopx/ 之外 0 个源文件被扫(已声明的边界)
    • E12 未覆盖:丢掉 loopx/extensions/ 下 83 个源文件后,所有读 reach 的检查无一察觉(上界预算看不见被收窄的 reach)→ 未声明盲区
    • E14 未覆盖:11 条未决决策,无任何检查强制其输入 → 未声明盲区
  4. 把声明变成可证伪。今天的 S 指标 ≈ 0%。E1 让"盲区披露率"从 0% 变成 2/15 显式登记,并给出证据。这是 S 从 0 起步唯一一块真实基座——它测的不是"我们写了检查",而是"检查真的拦得住"。
  5. 它本身就是守卫的退化陷阱。附录 E 的 E10 教训正是"只在本地跑过的变异语料不算测试"。本 PR 把语料变成 CI 里的 committed check 并冻结 12/15 基线:任何检查被删除都会让计数掉到 12 以下,PR 直接变红。这是 I 指标 G1 第一次真正进 CI。
  6. 成本有上界。9.5s(一次 inventory 构建),完全落在 tests/architecture 的 80s 内。零运行时改动、零新门禁、三份新文件、整文件可回滚。

它不买什么(明确排除):

  • 不改变任何运行时行为,用户可见影响 U 目前 = 0
  • 不提升检出能力——它测量检出能力,不增加
  • 不减少债务存量 D
  • 不修 E12 / E14 两个盲区,只是把它们登记为可追踪项
  • 不能作为"项目更有效了"的证据。它是测量工具,不是效果证据

结论:load_bearing,但承重的是"决策"而不是"运行时"。 它把 15 条教训从"文档里的说法"变成"12 条机器强制 / 3 条具名缺口",这是后续排期与 S/I 指标的地基。按同一判据,它既不是 duplicate_candidate,也不只是 derivable_only。

三、验证与状态

  • 本地:semantic-vocabulary-drift-smoke.py --report rc 0;pytest -q tests/architecture 290 passed / 80s;ruff check 两个新文件干净
  • CI:9 passed / 0 failed / 8 pending
  • reviewDecision=REVIEW_REQUIRED——控制平面改动,不自行合并,等 owner 批准

四、附带发现(与代码无关但阻塞流程)

GH_TOKEN / GITHUB_TOKEN 里的 github_pat_* 已失效,会使 gh 与 loopx 的 PR 轮询全线 401。绕过方式:env -u GH_TOKEN -u GITHUB_TOKEN gh ...(走 keychain 认证)。建议纳入本地 P0 清理。

…trodiction

Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

exact-head 复核(0287d87d0)— 同步 main

之前 BEHIND main 8 个提交,检查零失败。干净合并,无冲突,内容未变。合并提交带 DCO 签名。

…trodiction

Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

本 PR 在 #4447 计划中的位置

issue #4447 现在有一节统一协调(中英双语),把这 13 个在开 PR 作为一个计划列出:各自修什么、为何必要、以及实测出的合并顺序。

冲突实测:对全部 78 对做了试合并,9 对冲突,分四簇,每一处都是文本相邻,没有一处是语义分歧。

冲突簇 涉及 PR 后合并者的解法
棘轮锚点 #4606、#4608、#4617、#4629 四者全部合并后的终态已实测:same_runtime_forks 18、same_runtime_fork_definitions 41、same_runtime_forks_semantic 11、conflicting_values 16、conflicting_definitions 55、multi_value_twins 13。registry 与 BUDGET_ANCHOR 两处都要带同一组值(该检查是相等而非 <=)
漂移测试文件 #4626、#4629、#4631 三者都在文件末尾追加,全部保留即可
RFC 双镜像 #4614、#4627、#4631 附录 B 决策行,按日期先后保留两条
清单与生成器 #4614、#4630 加法式:owner 对过滤与改名不变性证据同时保留

建议顺序(代价从低到高):#4628 → #4625、#4626 → #4627 → #4619、#4621 → #4630 → #4614 → #4631 → #4629 → #4617 → #4606 → #4608。四个棘轮 PR 放最后,因为每落地一个,下一个的数字就从估算变成确定值。

全部 13 个 PR 现已同步到 main、零失败检查。

…trodiction

Signed-off-by: song <22676124+songoow@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: fb9a5fad58f2256057ee044a1ed030d96c38d065 (codex/n1b-incident-retrodiction).

动机

RFC 的 Appendix E 收了 15 条来自真实逃逸的教训,但它自己列的验收面(probe)从来没有落地:仓库里没有任何东西回答「今天的 guard 还认不认这条缺陷形态」。教训只以散文存在,就会在 guard 漂移时变成 folklore——而且这正是 Appendix E 的 E10 教训本身说的:只在本机跑过的语料不算测试。

改动后这个问题有了提交物:loopx/semantics/incident_v0.json 把每条教训、期望规则、probe 和 expectation 配起来,scripts/semantic_incident_retrodiction.py 驱动真实 guard,tests/architecture/test_semantic_incident_retrodiction.py 把它变成 CI 内的检查并冻结基线。这是一个完整的、可独立复核的切片:它回答了「现状是否还拦住」这个问题,并明确停下、不去顺手改 guard 本身(那是另一个该单独 review 的变更)。

改动思路

入口只有一个:脚本的 main();测试用 sys.executable 把同一个脚本当子进程跑,所以这份测量真的被 CI 执行,不是无人调用的脚手架。权威输入是 ledger 加 guard 模块本身——Guard.__init__ 用 importlib.util.spec_from_file_location 把 examples/semantic-vocabulary-drift-smoke.py 导入成模块,拿到它真实的 globals,因此 probe 才能做故障注入;docstring 里解释得很对:runpy.run_path 返回的是副本,在那里注入是静默 no-op,会报「uncovered」却什么都没测。判定边界在 INCIDENT_PROBES / CONTROL_PROBES 里每个小函数:注入一个输入,然后调用未改动的 check 函数,看它是否拒绝。副作用为空——registry 只改内存深拷贝,唯一的临时文件走 tempfile.NamedTemporaryFile 并在 finally 里删掉。

复用面我确认过:guard 的谓词(check_coverage_floor、check_owned_vocabularies、check_scope_declarations、check_relations、evaluate_inventory_budget_findings、collect_literal_uses、count_identifier_modules)全部是调用原模块,没有第二份实现;loopx/semantics/ 下已经住着 vocabulary_v0.json / inventory_v0.json,ledger 放这里、且 load_ledger 要求 vocabulary_v0.json 就在旁边,符合既有形状。五条 negative control(guard 必须接受的合法变体)是这份测量最关键的设计:没有它们,一个永远报错的检查也能刷满覆盖率。

具体改动

3 个新文件、+789/-0,没有任何既有文件被改:ledger 166 行、driver 566 行、测试 57 行。

我在这个 head 上独立跑了 python scripts/semantic_incident_retrodiction.py,结果与作者声明一致:appendix_e_lessons=15、would_have_caught=12/15、negative_control_false_positives=0/5、caught_baseline=12/15,未被拦住的是 E02 / E12 / E14。pytest tests/architecture/test_semantic_incident_retrodiction.py -q → 2 passed;ruff check 两个新文件 → All checks passed。ledger 每条 defect 文本与 RFC Appendix E 的条目逐条对得上,measured_at: "3937f6b99" 是真实存在的合并提交。

关键代码讲解

  • Guard.__init__ / Guard.with_registry_file(脚本 82 起):前者以模块方式导入 guard,保证注入落在真实 globals 上;后者在改 registry 内容前,先把已提交的 registry 写进同一个临时路径并读回比对,注入不是活的就报错,而不是默默判定。这是整份 harness 的可信度基础,做对了。
  • probe_registry_weakness_rejected / probe_scan_reach_disclosed(242 起):故障注入的代表。前者同时验证「floor 被调低」「owner 不是 module::Symbol」「扫描根被收窄」三条都被拒;后者把 load_sources 换成收窄树(丢掉 loopx/extensions/ 下 83 个源文件),再跑三个吃 reach 的检查——全部没察觉,于是 E12 如实记为 uncovered,而不是假装覆盖。
  • main()(500 起):先对账「Appendix E 的条目数 == ledger 的 incidents 数」(行数变了就失败,不靠人记),再校验每个 probe 名可解析,然后统计 caught / control violations 并逐条打印;--report 补打印 defect/rule 文本,--json 给机器读。
  • tests/architecture/test_semantic_incident_retrodiction.py:41:用 sys.executable 跑子进程并钉住打印出来的计数,符合仓库「子进程沿用选定解释器」的规则。
  • loopx/semantics/incident_v0.json:load_ledger 用精确 key-set 校验(LEDGER_KEYS / INCIDENT_KEYS / CONTROL_KEYS)加 VERDICTS 枚举,expectation 只允许 caught/uncovered,非法状态很难表达。

对主干的风险

最强回归不是崩溃,而是「verdict 与它声称测量的东西脱钩」:读者把 uncovered(或将来的 caught)当成 guard 的行为,而 probe 其实没问 guard。我在这条思路上做了两个反例。

一条非阻塞 P2(F1):E02 的 probe 结构上是常量,不是观测。 probe_fixture_copy_visible(188-192)用 guard.sources 里不以 loopx/ 开头的路径来判定,而 Guard.__init__ 调的是 load_sources(REPO_ROOT),默认 root 就是 loopx,路径又是 relative_to(repo_root) 生成的。我在这个 head 上量到 1206 个源文件、其中 0 个在 loopx/ 之外;即使把 guard.ns["LITERAL_SCAN_ROOTS"] 改成 ["loopx","tests","examples"],probe 仍然返回 False。也就是说 sources_outside=0 是构造出来的,E02 记成 uncovered 从未问过 guard 任何问题。建议让这个 probe 真的去问 guard:要么把 guard 声明的扫描根与已知 fixture 副本位置(tests/**、examples/**)对比并报出未覆盖的边界,要么往被扫描的根里注入一份合成 fixture 副本,看 check_literal_vocabularies 是否报出 fork,然后按观测重新写 expectation。

一条非阻塞 P2(F2):冻结基线把 PR 自己承诺的下一步变成红灯。 测试断言的是相等(tests/architecture/test_semantic_incident_retrodiction.py:53:(caught,total) == (baseline.would_have_caught, baseline.total)),失败信息是「a lesson that was caught must stay caught」;而 harness 自己只要求 caught >= baseline(脚本 553),PR 正文也把 E12/E14 称作 candidate work。我在这个 head 上复现了这个冲突:新增一个未跟踪的 loopx/canary/*.py,内容只是一个 open_decision 标识符,测量立刻变成 13/15——脚本打印 would_have_caught=13/15 加一行非致命的 expectation-mismatch、退出码 0,而提交的测试以 assert (13, 15) == (12, 15) 失败,报的还是「倒退」文案。也就是说按 PR 宣称的方向去补覆盖会先撞 CI,而真正需要跟进的 mismatch 只被打印、没有义务。建议测试改成 caught >= baseline 且 total == len(incidents),并让 expectation 变化时必须伴随 ledger 更新(或让 harness 对 mismatch 直接失败),这样规则只有一处表达。

一条 P3(F3,非阻塞):E14 是子串判定,且 ledger 的 meaning 声明了它没有的性质。 probe_dangling_decision 用 "open_decision" in 拼接 loopx/canary/*.py 的文本判定,所以任何只是提到这个名字的模块都能把它刷成 caught(上面那个临时文件就产生了 check_enforcing_named_inputs=True、E14 caught),而没有任何检查真的读了 open decisions 的输入。另外 ledger 的 meaning 写着 expectation 只从教训文本撰写、绝不来自 probe 输出,但 Appendix E 里没有任何一句能推出 E02/E12/E14 会被漏掉——这三个 uncovered 只能是读测量结果写的。建议 E14 改成匹配一个可解析的具名输入检查,并把 meaning 改写成「作者依据教训文本 + 首次测量记录」或把这三条明确标为 declared gap。

语义与 CI 对齐

这份变更复用既有语义面而不是造新的:没有新增产品词汇,唯一的两个新名字是 ledger 的 schema 版本和它自己的文件名,且都落在 loopx/semantics/ 既有目录里;guard 依旧是唯一的判定权威,ledger 只记录观测。语义上需要跟进的正是 F3 里那句 meaning 声明与文件实际内容的偏差,以及被冻结的 expectation。

我的整体评价

结论 APPROVE。这是一个有真实价值的增量:它把 Appendix E 从散文变成可重跑的测量,用真实 guard 而不是复刻实现,用五条 negative control 挡住「永远报错也算覆盖」,并把「先证明注入是活的再信 verdict」这条做对了——这些正是同类 harness 最容易糊弄的地方。我独立复现了作者声明的全部数字(12/15、0/5、E02/E12/E14 uncovered),核对了两条测试与 ruff,也确认 diff 只新增三个文件、没有碰任何既有路径。回退成本是一个 commit。

同时我把这份测量的可信度边界说清楚:15 条 verdict 里有 2 条(E02、E14)不是对 guard 的观测,冻结基线又把 PR 自己列的 candidate work 变成红灯,meaning 的独立撰写声明与实际内容不符。这三条都是新文件内部的可维护性问题,非阻塞,不改变本次结论;但在这个 harness 被用来判断「哪条教训还被拦住」之前,值得先修掉——否则它会以很高的话术质量给出偏乐观或偏悲观的读数。

English verdict: APPROVE - exact head fb9a5fa; I reproduced the claimed measurement independently (would_have_caught=12/15, negative_control_false_positives=0/5, E02/E12/E14 uncovered, 2 architecture tests passed, ruff clean, three new files and no existing file modified), and the harness genuinely drives the real pre-merge guard through its own globals with five negative controls. Three non-blocking findings, all inside the new files: the E02 probe's verdict is structurally constant because it filters an already root-limited source list instead of observing the guard; the committed test asserts equality against the frozen baseline, so the E12/E14 improvement the PR advertises fails CI with a regression message while the harness only prints the mismatch and exits 0; and E14 decides with a substring over loopx/canary/*.py while the ledger's meaning claims an expectation authorship the three uncovered rows do not have.

@huangruiteng
huangruiteng merged commit 1924c6f into loopx-project:main Sep 17, 2026
23 checks passed
@songoow
songoow deleted the codex/n1b-incident-retrodiction branch September 28, 2026 03:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants