fix(semantics): two measurements that contradicted themselves - #4680
huangruiteng merged 6 commits into
Conversation
Collision identity compared `tuple(item["values"])`, which is the order the source happens to list elements in. Swapping three lines inside a `set` literal, membership unchanged, turned a twin into a fork and failed the gate on three budgets at once -- while `divergent_value_sets`, computed from the same inventory, correctly reported no divergence. One scan produced two contradictory answers and the wrong one held the gate. Identity is now membership when every definition of a name is a `set` or `frozenset`, and source order otherwise. The narrowness is the point: `LIFECYCLE_PRIORITY` is a `tuple` defined in two modules whose order *is* the priority, so normalizing every carrier would have replaced a false positive with a false negative. A name carried by mixed containers also stays order-sensitive, which makes the rule a pure relaxation -- it can only merge definitions the old rule split, so no untouched tree starts failing. Enums and `Literal` aliases carry no container and are left order-sensitive; their ordering semantics are not established here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The `projections[*]` selector returned `len(registry["projections"])` for both `verified` and `registered`, while `check_projections` imported exactly one hardcoded projection. A projection whose owner module and function do not exist anywhere in the tree, declared as 2/2, passed the whole smoke and printed `F5:2/2` under the evidence bound `executable_owner_function`. The declaration was counting as its own proof. `verified` now counts only the projections named in `EXECUTED_PROJECTIONS`, which lives in code for the same reason as `COVERAGE_ANCHOR`, and each one's registry `owner` must equal the function the check imports. A registered projection with no executed check raises `registered` without raising `verified`, the way F1 reports 6 of 26, so the gap is reported rather than blocked or hidden. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Both were reproduced against d8e7af1 before being written, and neither fix relaxes a budget, floor or anchor. Appendix A states what each measurement claimed, what it actually did, and what the correction still does not establish -- F5 walks one projection, and the ordering semantics of enums and `Literal` aliases remain open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这个 PR 修的是语义门禁报出证据不支持的东西的两个位置,两处都在 d8e7af141 上复现过,且都不以放松预算/下限/锚点为代价:
- F5 把"声明"当成了自己的证据:
FORMAL_DOMAIN_SELECTORS["projections[*]"]对verified和registered都返回注册表长度,而check_projections实际只 import 并执行了一个写死的投影。于是在 main 上加一个"owner 模块和函数在树里根本不存在"的投影、把域声明成 2/2,整套 smoke 会通过并打印F5:2/2——没有任何东西定位到那个 owner,更不用说执行它。 - 把
set的重排当成语义分叉:碰撞身份用tuple(item["values"]),也就是源码顺序。把一个set字面量里三行的顺序换一下(成员集合完全不变),门禁会一次挂掉三个预算(multi_value_forks2→3、multi_value_forks_semantic1→2、multi_value_fork_definitions6→9),而同一份 inventory 算出的divergent_value_sets却正确报告无分叉——一次扫描给出两个互相矛盾的答案,而且错的那个握着门禁。
改动思路
- F5:
verified只数EXECUTED_PROJECTIONS(放在代码里,理由与既有COVERAGE_ANCHOR相同:只改数据的编辑不得扩大不变式的声明),并要求每个投影的 registryowner与被 import 的函数一致。注入后要么报claims 2 verified members ... the m0 check walks 1,要么老实声明并得到F5:1/2——差距可见,和 F1 报 6/26 是同一种诚实。 - 身份:仅当某名字的每一个定义都是
set/frozenset时按成员判等,否则保持顺序敏感。这个"窄"是关键:LIFECYCLE_PRIORITY是tuple,它的顺序就是优先级语义,一视同仁地归一化会把假阳性换成假阴性(把真实分叉藏起来)。把规则做成纯粹的放松——只会合并旧规则切开的东西,永远不会切开旧规则合并的东西——因此不会让任何未改动的树开始失败。
具体改动
关键代码讲解
loopx/semantics/inventory.py::_collision_keys+UNORDERED_CONTAINERS:身份判定改为"全为无序容器时按sorted(set(values)),否则按tuple(values)";混合容器与无容器(enum、Literal、TSas const数组)保持顺序敏感,方向偏保守(宁可报出来也不隐藏)。examples/semantic-vocabulary-drift-smoke.py::EXECUTED_PROJECTIONS与check_projections的 owner 一致性检查:把"声明"和"实际执行"绑在一起;owner 串指向别的函数时报crediting a function it does not run。- 新增 9 条回归(
tests/architecture/test_semantic_inventory.py:400-470、test_semantic_vocabulary_drift.py:733-780):两个缺陷各有余量方向的反向测试——成员变化仍是分叉、有序载体重排仍是分叉、混合容器仍顺序敏感、无容器载体仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个。
对主干的风险
无阻塞发现。两点记录:
- P3:F5 仍然只走一个投影。 这次是把数字变诚实,没有增加覆盖——PR 正文与两份 RFC 都明确写了这一点,我按同一口径记录,避免被读成"覆盖提升"。
- P3:本 head 的第一次 CI 是"被取消"而不是"失败"。
Python Testsrun 35296728937 的kernel-static-checks在跑cli-output-budget-regression-smoke.py时收到##[error]The operation was canceled.(跑约 15 分钟后),于是聚合任务checks/pytest/merge-gate因KERNEL_RESULT: cancelled报红。这不是代码失败,但会让这个 head 的 rollup 变红;我已对该 head 重跑该工作流(attempt 2),结果以重跑为准。
独立验证(exact head 5a6ba847):pytest tests/architecture → 375 passed in 62.84s(与 PR 声称一致);semantic-vocabulary-drift-smoke.py 打印 F5:1/1、projections=1/1;docs-governance-smoke ok。另一个工作流(DCO / Dependency Review / Release Artifacts / Frontstage Pages)在该 head 上全部 success。
我的整体评价
两处都是"门禁说了证据不支持的话",而且第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住——都属于会误导维护者的方向性错误。修法窄、可逆、有余量方向的反向测试,"纯粹放松"这个论证我按代码核过(all(... in UNORDERED_CONTAINERS) 才走成员判定,混合与无容器都退回顺序)。可以接受;合并前只需看重跑后的 CI rollup。
English verdict: APPROVE - head 5a6ba84 makes the F5 verified count name only projections the check actually executes (with an owner must-match-the-imported-function rule) and keys collision identity on membership for set/frozenset carriers while keeping source order for tuples, mixed and container-less carriers, so the change is a pure relaxation; verified by 375 passing architecture tests, the drift smoke printing F5:1/1 and projections=1/1, nine new regression tests covering both defects and their overshoot directions, and docs-governance ok, with the caveat that this head's first CI attempt was cancelled at the runner level and has been re-run.
…dence Both mirrors conflicted on the Appendix A append cluster: this branch's entry for the two corrected measurements against main's entry for the cross_runtime value notes (loopx-project#4662). Disjoint additions, so both are kept, ours first. The only line this branch removes from main's version of each mirror is the F5 sentence it deliberately rewrites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…dence Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这个 head(5d426d83)相对我上一轮评过的 5a6ba847 没有任何 PR 自有内容的变化:差异全部来自 origin/main 的合并。我批准的是精确 head,所以 head 一动就重新取证——但这一轮要确认的只是"两处修正还在不在、数还对不对",不是重做设计评审。
改动本身修的仍是 main 上的两个"门禁报出证据不支持的东西":FORMAL_DOMAIN_SELECTORS["projections[*]"] 对 verified 和 registered 都返回注册表长度,而 check_projections 实际只 import 并执行一个写死的投影;碰撞身份用源码顺序(tuple(item["values"])),于是把 set 字面量的三行换个顺序就会一次挂掉三个预算,而同一份 inventory 算出的 divergent_value_sets 却正确报告无分叉。
改动思路
(与上一轮相同)verified 只数代码里声明的 EXECUTED_PROJECTIONS,并要求每个投影的 registry owner 与被 import 的函数一致——只改数据的编辑不得扩大不变式的声明;身份判定仅在某个名字的每一个定义都是 set/frozenset 时按成员判等,其余保持顺序敏感,因此规则是纯粹的放松,只会合并旧规则切开的东西。
具体改动
在 5d426d8 上的复核
先证明"内容没变":
git diff --stat 5a6ba847..5d426d83 -- loopx/semantics/inventory.py \
examples/semantic-vocabulary-drift-smoke.py \
tests/architecture/test_semantic_inventory.py \
tests/architecture/test_semantic_vocabulary_drift.py → 空
git log --oneline --no-merges 5a6ba847..5d426d83 → 只有 main 侧提交
再证明"数还成立"(实跑,非引用):
/Users/bytedance/goal-harness/.venv/bin/python -m pytest tests/architecture -q
→ 419 passed in 61.30s
/Users/bytedance/goal-harness/.venv/bin/python examples/semantic-vocabulary-drift-smoke.py
→ ok
F5:1/1, projections=1/1
multi_value_forks=2/2, multi_value_forks_semantic=1/1,
multi_value_fork_definitions=6/6, formal_domain_bounds: projections=1/1
(上一轮在 5a6ba847 上是 375 passed;差额来自 main 新增的测试,不是本 PR 的。tests/architecture 整目录包含那 9 条新回归。)
CI 也比上一轮干净:这条 head 上 kernel-static-checks pass(12m41s),dashboard-acceptance、node-minimum/forward-compatibility、dependency-review、Sign-off 均 pass;我读 rollup 时 test-shard 还在跑,所以"聚合 checks 变绿"是合并前唯一还要看的读回。上一轮那次的 15 分钟被取消是上一条 head 的容量症状,不是这份 diff 的属性。
对主干的风险
无阻塞发现。两点记录:
- P3:F5 仍然只走一个投影。 这次是把数字变诚实,不是提升覆盖;PR 正文与两份 RFC 镜像都写明了同一口径,我按同样措辞记录,避免被读成"覆盖提升"。
- P3:head 移动只发生在 main 合并上。 这个 PR 自有文件在两版 head 之间零差异,所以"合并让语义悄悄变化"这条风险在本轮不成立;但两版都是 docs 之外无动作,仍建议合并前确认
test-shard收尾。
我的整体评价
两处都是"门禁说了证据不支持的话",第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住,都属于会误导维护者的方向性错误。修法窄、可逆、每个修正都有余量方向的反向测试(成员变化仍分叉、有序载体重排仍分叉、混合容器仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个)。head 移动时我按路径范围取证而不是重读整棵树;结论沿用,证据换新。可以接受。
English verdict: APPROVE - head 5d426d8 differs from the head I previously reviewed (5a6ba84) only by the merge of origin/main, with the path-scoped diff over this PR's own files empty, so the same two corrections still stand and I re-verified them by running rather than quoting: tests/architecture gives 419 passed in 61.30s (375 at the earlier head, the difference being tests main added, and this run includes the nine new regression tests covering both defects and their overshoot directions), the drift smoke exits ok printing F5:1/1, projections=1/1, multi_value_forks=2/2 and multi_value_fork_definitions=6/6, and CI at this head shows kernel-static-checks passing in 12m41s (the earlier 15-minute cancellation belonged to the previous head) with dashboard-acceptance, node compatibility, dependency-review and Sign-off green and test-shards still pending; the two non-blocking P3 notes remain that F5 still walks a single projection so the corrected number is honesty rather than coverage, and that the aggregated checks are the last readback before merge.
…dence Appendix A append cluster in both mirrors again; both sides kept, ours first. The only line removed from main's version of each mirror is the F5 sentence this branch deliberately rewrites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这个 head(3b7d69a8)已经是我第三次评它:5a6ba847 → 5d426d83 → 3b7d69a8,两次移动都只是 origin/main 的合并,PR 自有代码与测试文件在两次移动中都没有变化(变化的两份 RFC 镜像是 main 合并带进来的内容)。我批准的是精确 head,所以 head 一动就重新取证——但这一轮要确认的只是"两处修正还在不在、数还对不对",不是重做设计评审。
改动本身修的仍是 main 上的两个"门禁报出证据不支持的东西":FORMAL_DOMAIN_SELECTORS["projections[*]"] 对 verified 和 registered 都返回注册表长度,而 check_projections 实际只 import 并执行一个写死的投影;碰撞身份用源码顺序(tuple(item["values"])),于是把 set 字面量的三行换个顺序就会一次挂掉三个预算,而同一份 inventory 算出的 divergent_value_sets 却正确报告无分叉。
改动思路
(与上一轮相同)verified 只数代码里声明的 EXECUTED_PROJECTIONS,并要求每个投影的 registry owner 与被 import 的函数一致——只改数据的编辑不得扩大不变式的声明;身份判定仅在某个名字的每一个定义都是 set/frozenset 时按成员判等,其余保持顺序敏感,因此规则是纯粹的放松,只会合并旧规则切开的东西。
具体改动
在 5d426d8 上的复核
先证明"内容没变":
git diff --stat 5a6ba847..3b7d69a8 -- loopx/semantics/inventory.py \
examples/semantic-vocabulary-drift-smoke.py \
tests/architecture/test_semantic_inventory.py \
tests/architecture/test_semantic_vocabulary_drift.py → 空
git log --oneline --no-merges 5a6ba847..3b7d69a8 → 只有 main 侧提交
再证明"数还成立"(实跑,非引用):
/Users/bytedance/goal-harness/.venv/bin/python -m pytest tests/architecture -q
→ 517 passed in 94.03s
/Users/bytedance/goal-harness/.venv/bin/python examples/semantic-vocabulary-drift-smoke.py
→ ok
F5:1/1, projections=1/1
multi_value_forks=2/2, multi_value_forks_semantic=1/1,
multi_value_fork_definitions=6/6, formal_domain_bounds: projections=1/1
(上一轮在 5a6ba847 上是 375 passed、5d426d83 上是 419 passed;差额来自 main 新增的测试,不是本 PR 的。tests/architecture 整目录包含那 9 条新回归。)
CI 读回是针对上一条 head 5d426d83 的:kernel-static-checks pass(12m41s),dashboard-acceptance、node-minimum/forward-compatibility、dependency-review、Sign-off 均 pass;读 rollup 时 test-shard 仍在跑。再往前那次的 15 分钟被取消是容量症状,不是这份 diff 的属性。本 head(3b7d69a8)的 rollup 我没有重读,所以"聚合 checks 变绿"仍是合并前最后一项读回。
对主干的风险
无阻塞发现。两点记录:
- P3:F5 仍然只走一个投影。 这次是把数字变诚实,不是提升覆盖;PR 正文与两份 RFC 镜像都写明了同一口径,我按同样措辞记录,避免被读成"覆盖提升"。
- P3:head 移动只发生在 main 合并上。 这个 PR 自有文件在两版 head 之间零差异,所以"合并让语义悄悄变化"这条风险在本轮不成立;但两版都是 docs 之外无动作,仍建议合并前确认
test-shard收尾。
我的整体评价
两处都是"门禁说了证据不支持的话",第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住,都属于会误导维护者的方向性错误。修法窄、可逆、每个修正都有余量方向的反向测试(成员变化仍分叉、有序载体重排仍分叉、混合容器仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个)。head 移动时我按路径范围取证而不是重读整棵树;结论沿用,证据换新。可以接受。
English verdict: APPROVE - head 3b7d69a has now moved twice (5a6ba84 to 5d426d8 to 3b7d69a), both times only by merging origin/main, and the path-scoped diff over this PR's own code and test files is empty across both moves, so the same two corrections still stand and I re-verified them by running rather than quoting: tests/architecture gives 517 passed in 94.03s (419 at the previous moved head, 375 before that, the difference being tests main added, and this run includes the nine new regression tests covering both defects and their overshoot directions), the drift smoke exits ok printing F5:1/1, projections=1/1, multi_value_forks=2/2 and multi_value_fork_definitions=6/6, and the CI rollup I read was the previous head's (kernel-static-checks pass in 12m41s with dashboard-acceptance, node compatibility, dependency-review and Sign-off green and test-shards pending) - I did not re-read 3b7d69a's rollup, so that plus the two standing P3 notes (F5 still walks a single projection, so the corrected number is honesty rather than coverage) are the remaining readbacks before merge.
Three append clusters, all disjoint, both sides kept: the two RFC mirrors, and `test_semantic_vocabulary_drift.py`, where main's side is the F5 projection regression from loopx-project#4680 and this branch's is the B3 budget group. No top-level function is defined twice after the merge, and nothing was removed from main's copy of the test file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Two defects on
mainwhere the semantic gate reports something the evidence does not support. Both were reproduced againstd8e7af141before anything was written, and neither fix relaxes a budget, floor or anchor.1. A declaration counted as its own evidence (F5)
FORMAL_DOMAIN_SELECTORS["projections[*]"]returnedlen(registry["projections"])for bothverifiedandregistered, whilecheck_projectionsimported exactly one hardcoded projection name.Reproduction on
main: add a projection whose owner module and function do not exist anywhere in the tree, declare the domain as 2/2, and the full smoke passes, printingunder the evidence bound
executable_owner_function. Nothing located that owner, let alone executed it.verifiednow counts only the projections named inEXECUTED_PROJECTIONS, which lives in code for the same reason asCOVERAGE_ANCHOR— a data-only edit must not be able to widen what an invariant claims. Each one's registryownermust equal the function the check imports.After: the same injection either fails with
claims 2 verified members of projections[*]; the m0 check walks 1, or is declared honestly and reportsF5:1/2— the gap visible, the way F1 reports 6 of 26. Tampering with the executed projection's owner string fails withcrediting a function it does not run.2. Reordering a
setcounted as a semantic forkCollision identity used
tuple(item["values"])— source order. Swapping three lines inside asetliteral, membership unchanged, failed the gate:while
divergent_value_sets, computed from the same inventory, correctly reported no divergence. One scan, two contradictory answers, and the wrong one held the gate.Identity is now membership when every definition of a name is a
set/frozenset, and source order otherwise.Why the normalization is deliberately narrow
LIFECYCLE_PRIORITYis atupledefined in two modules whose order is the priority, andRAW_MATERIAL_KEY_HINTSis another multi-defined tuple. Normalizing every carrier — as the obvious one-line version does — would have replaced a false positive with a false negative and hidden a real divergence.A name carried by mixed containers also stays order-sensitive. That makes the rule a pure relaxation: it can only merge definitions the old rule split, never split a pair it merged, so no untouched tree starts failing. (The first version of this patch keyed per carrier instead of per name group and grew the fork count 2 → 4; the tests below pin that shut.)
What this does not establish
Literalaliases carry no container and are left order-sensitive — their ordering semantics are not settled here.Verification
All from the checkout on the merged tree:
semantic-vocabulary-drift-smoke.pyF5:1/1,projections=1/1pytest tests/architecture/docs-governance-smoke.pycanary premerge --from-git-diffNine new regression tests cover both defects and, just as important, the ways the fixes could overshoot: membership change is still a fork, reordering an ordered carrier is still a fork, mixed containers stay order-sensitive, and a container-less carrier stays order-sensitive.
Both RFC mirrors record what each measurement claimed, what it actually did, and what the correction still does not establish.
Refs #4447
🤖 Generated with Claude Code