Skip to content

fix(semantics): two measurements that contradicted themselves - #4680

Merged
huangruiteng merged 6 commits into
loopx-project:mainfrom
songoow:codex/correct-fork-identity-and-projection-evidence
Sep 18, 2026
Merged

huangruiteng merged 6 commits into
loopx-project:mainfrom
songoow:codex/correct-fork-identity-and-projection-evidence

Conversation

@songoow

@songoow songoow commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator

Two defects on main where the semantic gate reports something the evidence does not support. Both were reproduced against d8e7af141 before anything was written, and neither fix relaxes a budget, floor or anchor.

1. A declaration counted as its own evidence (F5)

FORMAL_DOMAIN_SELECTORS["projections[*]"] returned len(registry["projections"]) for both verified and registered, while check_projections imported exactly one hardcoded projection name.

Reproduction on main: add a projection whose owner module and function do not exist anywhere in the tree, declare the domain as 2/2, and the full smoke passes, printing

formal_domain=F1:6/26,F2:6/26,F3:0/26,F4:4/4,F5:2/2,F6:0/0 (verified/registered)

under the evidence bound executable_owner_function. Nothing located that owner, let alone executed it.

verified now counts only the projections named in EXECUTED_PROJECTIONS, which lives in code for the same reason as COVERAGE_ANCHOR — a data-only edit must not be able to widen what an invariant claims. Each one's registry owner must equal the function the check imports.

After: the same injection either fails with claims 2 verified members of projections[*]; the m0 check walks 1, or is declared honestly and reports F5:1/2 — the gap visible, the way F1 reports 6 of 26. Tampering with the executed projection's owner string fails with crediting a function it does not run.

2. Reordering a set counted as a semantic fork

Collision identity used tuple(item["values"]) — source order. Swapping three lines inside a set literal, membership unchanged, failed the gate:

inventory multi_value_forks grew to 3; budget is 2 (unreviewed);
multi_value_forks_semantic grew to 2; budget is 1; multi_value_fork_definitions grew to 9; budget is 6

while divergent_value_sets, computed from the same inventory, correctly reported no divergence. One scan, two contradictory answers, and the wrong one held the gate.

Identity is now membership when every definition of a name is a set/frozenset, and source order otherwise.

Why the normalization is deliberately narrow

LIFECYCLE_PRIORITY is a tuple defined in two modules whose order is the priority, and RAW_MATERIAL_KEY_HINTS is another multi-defined tuple. Normalizing every carrier — as the obvious one-line version does — would have replaced a false positive with a false negative and hidden a real divergence.

A name carried by mixed containers also stays order-sensitive. That makes the rule a pure relaxation: it can only merge definitions the old rule split, never split a pair it merged, so no untouched tree starts failing. (The first version of this patch keyed per carrier instead of per name group and grew the fork count 2 → 4; the tests below pin that shut.)

What this does not establish

  • F5 still walks one projection. This makes the reported number honest; it does not add coverage.
  • Enums and Literal aliases carry no container and are left order-sensitive — their ordering semantics are not settled here.

Verification

All from the checkout on the merged tree:

Gate Result
semantic-vocabulary-drift-smoke.py ok — F5:1/1, projections=1/1
pytest tests/architecture/ 375 passed
docs-governance-smoke.py ok
canary premerge --from-git-diff ok — direct/catalog/risk/boundary all 0 failures

Nine new regression tests cover both defects and, just as important, the ways the fixes could overshoot: membership change is still a fork, reordering an ordered carrier is still a fork, mixed containers stay order-sensitive, and a container-less carrier stays order-sensitive.

Both RFC mirrors record what each measurement claimed, what it actually did, and what the correction still does not establish.

Refs #4447

🤖 Generated with Claude Code

songoow and others added 3 commits September 17, 2026 21:38
Collision identity compared `tuple(item["values"])`, which is the order the
source happens to list elements in. Swapping three lines inside a `set` literal,
membership unchanged, turned a twin into a fork and failed the gate on three
budgets at once -- while `divergent_value_sets`, computed from the same
inventory, correctly reported no divergence. One scan produced two contradictory
answers and the wrong one held the gate.

Identity is now membership when every definition of a name is a `set` or
`frozenset`, and source order otherwise. The narrowness is the point:
`LIFECYCLE_PRIORITY` is a `tuple` defined in two modules whose order *is* the
priority, so normalizing every carrier would have replaced a false positive with
a false negative. A name carried by mixed containers also stays order-sensitive,
which makes the rule a pure relaxation -- it can only merge definitions the old
rule split, so no untouched tree starts failing.

Enums and `Literal` aliases carry no container and are left order-sensitive;
their ordering semantics are not established here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The `projections[*]` selector returned `len(registry["projections"])` for both
`verified` and `registered`, while `check_projections` imported exactly one
hardcoded projection. A projection whose owner module and function do not exist
anywhere in the tree, declared as 2/2, passed the whole smoke and printed
`F5:2/2` under the evidence bound `executable_owner_function`. The declaration
was counting as its own proof.

`verified` now counts only the projections named in `EXECUTED_PROJECTIONS`,
which lives in code for the same reason as `COVERAGE_ANCHOR`, and each one's
registry `owner` must equal the function the check imports. A registered
projection with no executed check raises `registered` without raising
`verified`, the way F1 reports 6 of 26, so the gap is reported rather than
blocked or hidden.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Both were reproduced against d8e7af1 before being written, and neither fix
relaxes a budget, floor or anchor. Appendix A states what each measurement
claimed, what it actually did, and what the correction still does not
establish -- F5 walks one projection, and the ordering semantics of enums and
`Literal` aliases remain open.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
huangruiteng previously approved these changes Sep 18, 2026

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 PR 修的是语义门禁报出证据不支持的东西的两个位置,两处都在 d8e7af141 上复现过,且都不以放松预算/下限/锚点为代价:

  1. F5 把"声明"当成了自己的证据:FORMAL_DOMAIN_SELECTORS["projections[*]"] 对 verified 和 registered 都返回注册表长度,而 check_projections 实际只 import 并执行了一个写死的投影。于是在 main 上加一个"owner 模块和函数在树里根本不存在"的投影、把域声明成 2/2,整套 smoke 会通过并打印 F5:2/2——没有任何东西定位到那个 owner,更不用说执行它。
  2. 把 set 的重排当成语义分叉:碰撞身份用 tuple(item["values"]),也就是源码顺序。把一个 set 字面量里三行的顺序换一下(成员集合完全不变),门禁会一次挂掉三个预算(multi_value_forks 2→3、multi_value_forks_semantic 1→2、multi_value_fork_definitions 6→9),而同一份 inventory 算出的 divergent_value_sets 却正确报告无分叉——一次扫描给出两个互相矛盾的答案,而且错的那个握着门禁。

改动思路

  • F5:verified 只数 EXECUTED_PROJECTIONS(放在代码里,理由与既有 COVERAGE_ANCHOR 相同:只改数据的编辑不得扩大不变式的声明),并要求每个投影的 registry owner 与被 import 的函数一致。注入后要么报 claims 2 verified members ... the m0 check walks 1,要么老实声明并得到 F5:1/2——差距可见,和 F1 报 6/26 是同一种诚实。
  • 身份:仅当某名字的每一个定义都是 set/frozenset 时按成员判等,否则保持顺序敏感。这个"窄"是关键:LIFECYCLE_PRIORITY 是 tuple,它的顺序就是优先级语义,一视同仁地归一化会把假阳性换成假阴性(把真实分叉藏起来)。把规则做成纯粹的放松——只会合并旧规则切开的东西,永远不会切开旧规则合并的东西——因此不会让任何未改动的树开始失败。

具体改动

关键代码讲解

  • loopx/semantics/inventory.py::_collision_keys + UNORDERED_CONTAINERS:身份判定改为"全为无序容器时按 sorted(set(values)),否则按 tuple(values)";混合容器与无容器(enum、Literal、TS as const 数组)保持顺序敏感,方向偏保守(宁可报出来也不隐藏)。
  • examples/semantic-vocabulary-drift-smoke.py::EXECUTED_PROJECTIONS 与 check_projections 的 owner 一致性检查:把"声明"和"实际执行"绑在一起;owner 串指向别的函数时报 crediting a function it does not run。
  • 新增 9 条回归(tests/architecture/test_semantic_inventory.py:400-470、test_semantic_vocabulary_drift.py:733-780):两个缺陷各有余量方向的反向测试——成员变化仍是分叉、有序载体重排仍是分叉、混合容器仍顺序敏感、无容器载体仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个。

对主干的风险

无阻塞发现。两点记录:

  1. P3:F5 仍然只走一个投影。 这次是把数字变诚实,没有增加覆盖——PR 正文与两份 RFC 都明确写了这一点,我按同一口径记录,避免被读成"覆盖提升"。
  2. P3:本 head 的第一次 CI 是"被取消"而不是"失败"。 Python Tests run 35296728937 的 kernel-static-checks 在跑 cli-output-budget-regression-smoke.py 时收到 ##[error]The operation was canceled.(跑约 15 分钟后),于是聚合任务 checks/pytest/merge-gate 因 KERNEL_RESULT: cancelled 报红。这不是代码失败,但会让这个 head 的 rollup 变红;我已对该 head 重跑该工作流(attempt 2),结果以重跑为准。

独立验证(exact head 5a6ba847):pytest tests/architecture → 375 passed in 62.84s(与 PR 声称一致);semantic-vocabulary-drift-smoke.py 打印 F5:1/1、projections=1/1;docs-governance-smoke ok。另一个工作流(DCO / Dependency Review / Release Artifacts / Frontstage Pages)在该 head 上全部 success。

我的整体评价

两处都是"门禁说了证据不支持的话",而且第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住——都属于会误导维护者的方向性错误。修法窄、可逆、有余量方向的反向测试,"纯粹放松"这个论证我按代码核过(all(... in UNORDERED_CONTAINERS) 才走成员判定,混合与无容器都退回顺序)。可以接受;合并前只需看重跑后的 CI rollup。

English verdict: APPROVE - head 5a6ba84 makes the F5 verified count name only projections the check actually executes (with an owner must-match-the-imported-function rule) and keys collision identity on membership for set/frozenset carriers while keeping source order for tuples, mixed and container-less carriers, so the change is a pure relaxation; verified by 375 passing architecture tests, the drift smoke printing F5:1/1 and projections=1/1, nine new regression tests covering both defects and their overshoot directions, and docs-governance ok, with the caveat that this head's first CI attempt was cancelled at the runner level and has been re-run.

…dence

Both mirrors conflicted on the Appendix A append cluster: this branch's entry
for the two corrected measurements against main's entry for the cross_runtime
value notes (loopx-project#4662). Disjoint additions, so both are kept, ours first. The only
line this branch removes from main's version of each mirror is the F5 sentence
it deliberately rewrites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…dence

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
huangruiteng previously approved these changes Sep 18, 2026

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 head(5d426d83)相对我上一轮评过的 5a6ba847 没有任何 PR 自有内容的变化:差异全部来自 origin/main 的合并。我批准的是精确 head,所以 head 一动就重新取证——但这一轮要确认的只是"两处修正还在不在、数还对不对",不是重做设计评审。

改动本身修的仍是 main 上的两个"门禁报出证据不支持的东西":FORMAL_DOMAIN_SELECTORS["projections[*]"] 对 verified 和 registered 都返回注册表长度,而 check_projections 实际只 import 并执行一个写死的投影;碰撞身份用源码顺序(tuple(item["values"])),于是把 set 字面量的三行换个顺序就会一次挂掉三个预算,而同一份 inventory 算出的 divergent_value_sets 却正确报告无分叉。

改动思路

(与上一轮相同)verified 只数代码里声明的 EXECUTED_PROJECTIONS,并要求每个投影的 registry owner 与被 import 的函数一致——只改数据的编辑不得扩大不变式的声明;身份判定仅在某个名字的每一个定义都是 set/frozenset 时按成员判等,其余保持顺序敏感,因此规则是纯粹的放松,只会合并旧规则切开的东西。

具体改动

在 5d426d8 上的复核

先证明"内容没变":

git diff --stat 5a6ba847..5d426d83 -- loopx/semantics/inventory.py \
  examples/semantic-vocabulary-drift-smoke.py \
  tests/architecture/test_semantic_inventory.py \
  tests/architecture/test_semantic_vocabulary_drift.py        → 空
git log --oneline --no-merges 5a6ba847..5d426d83               → 只有 main 侧提交

再证明"数还成立"(实跑,非引用):

/Users/bytedance/goal-harness/.venv/bin/python -m pytest tests/architecture -q
                                                        → 419 passed in 61.30s
/Users/bytedance/goal-harness/.venv/bin/python examples/semantic-vocabulary-drift-smoke.py
                                                        → ok
  F5:1/1, projections=1/1
  multi_value_forks=2/2, multi_value_forks_semantic=1/1,
  multi_value_fork_definitions=6/6, formal_domain_bounds: projections=1/1

(上一轮在 5a6ba847 上是 375 passed;差额来自 main 新增的测试,不是本 PR 的。tests/architecture 整目录包含那 9 条新回归。)

CI 也比上一轮干净:这条 head 上 kernel-static-checks pass(12m41s),dashboard-acceptance、node-minimum/forward-compatibility、dependency-review、Sign-off 均 pass;我读 rollup 时 test-shard 还在跑,所以"聚合 checks 变绿"是合并前唯一还要看的读回。上一轮那次的 15 分钟被取消是上一条 head 的容量症状,不是这份 diff 的属性。

对主干的风险

无阻塞发现。两点记录:

  1. P3:F5 仍然只走一个投影。 这次是把数字变诚实,不是提升覆盖;PR 正文与两份 RFC 镜像都写明了同一口径,我按同样措辞记录,避免被读成"覆盖提升"。
  2. P3:head 移动只发生在 main 合并上。 这个 PR 自有文件在两版 head 之间零差异,所以"合并让语义悄悄变化"这条风险在本轮不成立;但两版都是 docs 之外无动作,仍建议合并前确认 test-shard 收尾。

我的整体评价

两处都是"门禁说了证据不支持的话",第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住,都属于会误导维护者的方向性错误。修法窄、可逆、每个修正都有余量方向的反向测试(成员变化仍分叉、有序载体重排仍分叉、混合容器仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个)。head 移动时我按路径范围取证而不是重读整棵树;结论沿用,证据换新。可以接受。

English verdict: APPROVE - head 5d426d8 differs from the head I previously reviewed (5a6ba84) only by the merge of origin/main, with the path-scoped diff over this PR's own files empty, so the same two corrections still stand and I re-verified them by running rather than quoting: tests/architecture gives 419 passed in 61.30s (375 at the earlier head, the difference being tests main added, and this run includes the nine new regression tests covering both defects and their overshoot directions), the drift smoke exits ok printing F5:1/1, projections=1/1, multi_value_forks=2/2 and multi_value_fork_definitions=6/6, and CI at this head shows kernel-static-checks passing in 12m41s (the earlier 15-minute cancellation belonged to the previous head) with dashboard-acceptance, node compatibility, dependency-review and Sign-off green and test-shards still pending; the two non-blocking P3 notes remain that F5 still walks a single projection so the corrected number is honesty rather than coverage, and that the aggregated checks are the last readback before merge.

…dence

Appendix A append cluster in both mirrors again; both sides kept, ours first.
The only line removed from main's version of each mirror is the F5 sentence this
branch deliberately rewrites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这个 head(3b7d69a8)已经是我第三次评它:5a6ba847 → 5d426d83 → 3b7d69a8,两次移动都只是 origin/main 的合并,PR 自有代码与测试文件在两次移动中都没有变化(变化的两份 RFC 镜像是 main 合并带进来的内容)。我批准的是精确 head,所以 head 一动就重新取证——但这一轮要确认的只是"两处修正还在不在、数还对不对",不是重做设计评审。

改动本身修的仍是 main 上的两个"门禁报出证据不支持的东西":FORMAL_DOMAIN_SELECTORS["projections[*]"] 对 verified 和 registered 都返回注册表长度,而 check_projections 实际只 import 并执行一个写死的投影;碰撞身份用源码顺序(tuple(item["values"])),于是把 set 字面量的三行换个顺序就会一次挂掉三个预算,而同一份 inventory 算出的 divergent_value_sets 却正确报告无分叉。

改动思路

(与上一轮相同)verified 只数代码里声明的 EXECUTED_PROJECTIONS,并要求每个投影的 registry owner 与被 import 的函数一致——只改数据的编辑不得扩大不变式的声明;身份判定仅在某个名字的每一个定义都是 set/frozenset 时按成员判等,其余保持顺序敏感,因此规则是纯粹的放松,只会合并旧规则切开的东西。

具体改动

在 5d426d8 上的复核

先证明"内容没变":

git diff --stat 5a6ba847..3b7d69a8 -- loopx/semantics/inventory.py \
  examples/semantic-vocabulary-drift-smoke.py \
  tests/architecture/test_semantic_inventory.py \
  tests/architecture/test_semantic_vocabulary_drift.py        → 空
git log --oneline --no-merges 5a6ba847..3b7d69a8               → 只有 main 侧提交

再证明"数还成立"(实跑,非引用):

/Users/bytedance/goal-harness/.venv/bin/python -m pytest tests/architecture -q
                                                        → 517 passed in 94.03s
/Users/bytedance/goal-harness/.venv/bin/python examples/semantic-vocabulary-drift-smoke.py
                                                        → ok
  F5:1/1, projections=1/1
  multi_value_forks=2/2, multi_value_forks_semantic=1/1,
  multi_value_fork_definitions=6/6, formal_domain_bounds: projections=1/1

(上一轮在 5a6ba847 上是 375 passed、5d426d83 上是 419 passed;差额来自 main 新增的测试,不是本 PR 的。tests/architecture 整目录包含那 9 条新回归。)

CI 读回是针对上一条 head 5d426d83 的:kernel-static-checks pass(12m41s),dashboard-acceptance、node-minimum/forward-compatibility、dependency-review、Sign-off 均 pass;读 rollup 时 test-shard 仍在跑。再往前那次的 15 分钟被取消是容量症状,不是这份 diff 的属性。本 head(3b7d69a8)的 rollup 我没有重读,所以"聚合 checks 变绿"仍是合并前最后一项读回。

对主干的风险

无阻塞发现。两点记录:

  1. P3:F5 仍然只走一个投影。 这次是把数字变诚实,不是提升覆盖;PR 正文与两份 RFC 镜像都写明了同一口径,我按同样措辞记录,避免被读成"覆盖提升"。
  2. P3:head 移动只发生在 main 合并上。 这个 PR 自有文件在两版 head 之间零差异,所以"合并让语义悄悄变化"这条风险在本轮不成立;但两版都是 docs 之外无动作,仍建议合并前确认 test-shard 收尾。

我的整体评价

两处都是"门禁说了证据不支持的话",第一处会让未执行的投影自证 verified,第二处让正确的摘要被错误的身份判定压住,都属于会误导维护者的方向性错误。修法窄、可逆、每个修正都有余量方向的反向测试(成员变化仍分叉、有序载体重排仍分叉、混合容器仍顺序敏感、未执行投影不得自证 verified、owner 必须是被 import 的那个)。head 移动时我按路径范围取证而不是重读整棵树;结论沿用,证据换新。可以接受。

English verdict: APPROVE - head 3b7d69a has now moved twice (5a6ba84 to 5d426d8 to 3b7d69a), both times only by merging origin/main, and the path-scoped diff over this PR's own code and test files is empty across both moves, so the same two corrections still stand and I re-verified them by running rather than quoting: tests/architecture gives 517 passed in 94.03s (419 at the previous moved head, 375 before that, the difference being tests main added, and this run includes the nine new regression tests covering both defects and their overshoot directions), the drift smoke exits ok printing F5:1/1, projections=1/1, multi_value_forks=2/2 and multi_value_fork_definitions=6/6, and the CI rollup I read was the previous head's (kernel-static-checks pass in 12m41s with dashboard-acceptance, node compatibility, dependency-review and Sign-off green and test-shards pending) - I did not re-read 3b7d69a's rollup, so that plus the two standing P3 notes (F5 still walks a single projection, so the corrected number is honesty rather than coverage) are the remaining readbacks before merge.

@huangruiteng
huangruiteng merged commit 71cf164 into loopx-project:main Sep 18, 2026
22 of 23 checks passed
songoow added a commit to songoow/loopx that referenced this pull request Sep 18, 2026
Three append clusters, all disjoint, both sides kept: the two RFC mirrors, and
`test_semantic_vocabulary_drift.py`, where main's side is the F5 projection
regression from loopx-project#4680 and this branch's is the B3 budget group. No top-level
function is defined twice after the merge, and nothing was removed from main's
copy of the test file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow
songoow deleted the codex/correct-fork-identity-and-projection-evidence branch September 28, 2026 03:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants