feat(semantics): print the merge candidates the registry does not explain - #4630
Conversation
…lain `merge_candidate_groups` finds distinct names carrying one identical multi-value set, but it never consulted the registry and `--report` never printed it, so the signal was both noisy and invisible. 18 of its 38 groups were one registered vocabulary's own two owner symbols, where Python spells the concept `EffectiveAction` and TypeScript spells it `EFFECTIVE_ACTIONS`. The registry entry is already the decision that those are one vocabulary, so each of those groups asks a question that has an answer, and together they buried the 20 groups that nobody has ruled on. Pass the registry to drop the groups whose names are exactly one vocabulary's owner symbols, mark each surviving group cross-runtime or Python-only, and print the survivors from `--report` with names, values and modules in a stable order. Calling `merge_candidate_groups` without a registry still returns the unfiltered list for a raw audit. A group stays advisory either way; nothing here merges anything. This lowers no debt count and classifies nothing. Read 38 -> 20 as noise removed from an advisory review list, not as debt repaired: the 18 dropped pairs are a cross-runtime naming convention the registry already settled, the 20 that remain are unchanged and still unclassified, and the RFC's rule holds that classification work is not required to lower any count and that renaming never counts as a fix. No budget, anchor, coverage floor or registry value moves in this diff. Measured on this branch: 38 raw groups, 18 explained, 20 to review, of which 2 span both runtimes and are therefore registry gaps a human still has to decide on: MATERIAL_DELIVERY_OUTCOMES / VISION_OUTCOME_CHECKPOINT_MATERIAL_OUTCOMES, and MONITOR_METADATA_FIELDS / TODO_MONITOR_METADATA_FIELDS. The other 18 are Python-only. tests/architecture: 288 -> 291 passed; the drift smoke report is byte-identical to its pre-change output. Refs loopx-project#4447. Signed-off-by: song <liusongstep@gmail.com>
The first version of this column reported the raw `merge_candidate_groups` output, 38, and drew the conclusion that merge candidates are "the one surface that grew". Measured against the registry, that reading is wrong in both the number and the direction. Of the 38 groups, exactly 18 are the Python and TypeScript owner symbols of one registered vocabulary -- `AGENT_SCOPE_FRONTIER_ACTIONS` with `AgentScopeFrontierAction`, `EFFECTIVE_ACTIONS` with `EffectiveAction`, and 16 more. A registered cross-runtime vocabulary is *required* to have both ends; grouping them as merge candidates is the grouping function not consulting the registry it sits next to. The real backlog is 20, which is smaller than the historical 32, not larger. Reporting 38 would have put a scanner artifact into the plan as debt growth -- the exact failure mode this RFC exists to prevent, committed in the column that claims to measure it. loopx-project#4630 fixes the grouping function itself; this states the split so the column is right whether or not that lands first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…-visibility Signed-off-by: song <22676124+songoow@users.noreply.github.com>
本 PR 在 #4447 计划中的位置issue #4447 现在有一节统一协调(中英双语),把这 13 个在开 PR 作为一个计划列出:各自修什么、为何必要、以及实测出的合并顺序。 冲突实测:对全部 78 对做了试合并,9 对冲突,分四簇,每一处都是文本相邻,没有一处是语义分歧。
建议顺序(代价从低到高):#4628 → #4625、#4626 → #4627 → #4619、#4621 → #4630 → #4614 → #4631 → #4629 → #4617 → #4606 → #4608。四个棘轮 PR 放最后,因为每落地一个,下一个的数字就从估算变成确定值。 全部 13 个 PR 现已同步到 |
…-visibility loopx-project#4614 merged while this branch was open. Both append a section to the same --report block: loopx-project#4614 prints the name-keyed fork advisory, this branch prints the merge candidates. Resolved additively, fork section first, because the two answer different questions -- a fork is one name whose modules disagree, a candidate group is two or more names carrying identical values, so neither report can substitute for the other. Verified after the merge: --report prints both sections, and the candidate header reads '20 to review, 18 explained by a registered vocabulary's own owner symbols, 38 raw groups'. Signed-off-by: song <22676124+songoow@users.noreply.github.com>
exact-head 复核(
|
…-visibility Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: 6ed92d8c1d93e53b592bea6fe839b542c43f22db (codex/merge-candidate-visibility).
动机
#4447 里合并候选本来是「供人评审的建议列表」,但分组完全不看注册表声明的 owner,于是同一个已注册词表的两端——Python 的 EffectiveAction 与 TypeScript 的 EFFECTIVE_ACTIONS——被当成待合并的重复项。基线实测:38 组里 18 组是这种配对,真正需要决策的行被埋在噪音里;更糟的是没有任何命令打印这些组,所以清单既看不见也排不出序。
改动后 38 → 20 组可评审(2 组跨运行时注册表缺口 + 18 组纯 Python 待分类),并且 --report 真的把它们打出来。这是可独立复核的完整切片:只做「让未定论的候选可见」,分类与合并都留给后续。
改动思路
入口是 scripts/generate_semantic_inventory.py --report → print_merge_candidates() → loopx/semantics/inventory.py::merge_candidate_groups(inventory, registry)。
判定归属设计得对:注册表是「两种拼法同一个概念」的裁定者,所以「哪些名字集合已被解释」由新助手 registered_owner_symbol_sets(registry) 从 owners 派生(owner.split("::")[-1],且只保留有多个不同符号的集合——单 owner 或两端同名的词表什么也解释不了),而不是在报告脚本里再写一份规则。过滤条件也保守:只有 frozenset(names) 恰好等于某个已注册词表的 owner 符号集才丢弃,所以 3 个以上名字的组永远不会被误删;未传注册表时行为与之前完全一致(旧调用点 scripts/semantic_incident_retrodiction.py 的对照组正是靠这一点)。
具体改动
5 个文件、+218/-11(inventory.py ~46 行、报告脚本 ~39 行、测试 118 行、RFC 中英各一段)。
我独立复算了核心数字,与作者声明一致:raw 38 / 过滤后 20 / 丢弃 18,且 18 个被丢弃的名字集合逐一对应某个已注册词表的两个 owner 符号;存活的 2 个跨运行时组是 MATERIAL_DELIVERY_OUTCOMES ↔ VISION_OUTCOME_CHECKPOINT_MATERIAL_OUTCOMES 与 MONITOR_METADATA_FIELDS ↔ TODO_MONITOR_METADATA_FIELDS,我把这四个符号都拿去注册表里查过,没有任何词表把它们声明为 owner——所以「注册表缺口」这个说法成立。pytest tests/architecture/test_semantic_inventory.py -q → 18 passed(含 3 个新用例);drift smoke → ok,计数与 merge-base 逐项一致。
关键代码讲解
registered_owner_symbol_sets(inventory.py:406):从owners抽符号名集合,len(symbols) > 1才返回;docstring 把「单 owner 或同名两端解释不了任何东西」写清楚了,避免把 python-only 词表也当成可解释来源。merge_candidate_groups(419):新增可选registry,len(names) < 2 or frozenset(names) in explained才跳过;同时给每组加cross_runtime(模块同时含.py与.ts)。文档里明确「过滤只让未定论的组可读,既不退休任何东西也不给任何留存组下定论」,与 RFC 「advisory, never auto-merged」一致。print_merge_candidates(报告脚本 33):先打印20 to review, 18 explained …, 38 raw groups这类计数行,再按(not cross_runtime, names)排序输出(注册表缺口排最前),每组带 names/values/modules。tests/.../test_semantic_inventory.py的merge_candidate_repofixture:合成一个仓库,里面有「已注册且两端拼法不同的一对」「未注册的纯 Python 一对」「未注册的跨运行时一对」「已注册但只有一个 owner 的词表」,四个行为各自钉住。
对主干的风险
最强回归不是崩溃,而是过滤掉仍有待评审的组。我在真实仓库上反向验证过:所有被丢弃的集合都能映射到一个真实词表的 owner 对,两个存活跨运行时组则确实无人注册;配合「恰好相等才丢」的条件,3 名以上的组不会受影响。另一条被现有护栏兜住的风险是:若某词表的 owner 模块值与注册表声明不一致,这个配对会在这里被当作「已解释」丢掉——但那种漂移由 smoke 的 owned-vocabulary 检查单独报错,不会因此静默。
一条 P3(F1,非阻塞):新块不吃 --top,而它上面两块(consumer ranking、divergent value sets)都按 [: args.top] 截断,flag 的说明还是「rows to print with --report」(默认 25)。今天 20 行看不出来,但候选空间是 38,而这块存在的意义就是让不断累积的分类清单可读;一旦超过 25,报告会在两个有截断的块旁边悄悄打印一段无界输出。建议要么同样按 args.top 截断(截断时打印 showing N of M),要么在 --top 说明里写明候选列表刻意全量。
我的整体评价
结论 APPROVE。这是一个方向正确的可见性增量:把注册表已有的裁定接入派生报告,而不是新增权威;过滤条件精确到「恰好等于某个词表的 owner 符号集」,未传注册表时行为不变,还顺手把两个真实的跨运行时注册表缺口顶到列表最前。我独立复算了 38/20/18 与两个存活缺口,跑了 inventory 测试与 drift smoke,并确认报告真的打印这些行。回退成本是一个 commit。
唯一 P3 是 --top 在该块上不生效,属显示一致性问题,不影响本次判断。
English verdict: APPROVE - exact head 6ed92d8; the registry-aware filter is precise (a group is dropped only when its name set is exactly one registered vocabulary's owner-symbol set, so larger groups and the no-registry call path are untouched), I re-derived 38 raw / 20 to review / 18 explained independently, confirmed the two surviving cross-runtime pairs are owned by no vocabulary, ran the 18 inventory tests and the unchanged drift smoke, and the report prints the filtered list cross-runtime first. One non-blocking P3: the new merge-candidate block ignores --top while the two blocks above it honour it.
No conflict; the branch was only behind. loopx-project#4628 memoized python_facts in the same module this branch extends, and the two changes compose. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
Head moved: merged No conflict — the branch was only behind. Worth noting for the re-read: #4628 memoized Revalidated: Per the exact-head rule in |
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: 6f11b951fb4cd5158e35d2dcf1a97446aeaca233 (codex/merge-candidate-visibility, re-review after the earlier head was superseded by a main merge).
动机
未变:--report 里的 merge candidate 是按「值集相同、名字不同」分组的,但完全不看注册表声明的 owner——于是 EffectiveAction 与 EFFECTIVE_ACTIONS 这种同一个已注册词表的两种拼写也被当成待合并候选。38 组里 18 组属于这一类,真正需要人裁决的组被埋在噪音里,而且没有任何输出说明哪些组注册表已经拍过板。这个动机在本次 head 上仍然成立。
改动思路
把「注册表已经裁定过」这件事接进派生报告,而不是新增权威:过滤规则放进已经负责分组的 merge_candidate_groups(inventory, registry)(registry 可选,传 None 保留原始审计视图),报告行同时打印「待评审 / 已解释 / 原始」三个数,并把跨运行时组排到最前——跨运行时且无人注册,本身就是注册表缺口。
具体改动
5 个文件、+218/-11(相对当前 main;本 head 相对我上次 review 只多了一次 main 合并,内容增量相同):
loopx/semantics/inventory.py:新增registered_owner_symbol_sets(只收集「有多个不同 owner 符号」的词表),merge_candidate_groups增加可选registry参数,命中「名字集合恰好等于某个词表的 owner 符号集」时跳过,并给存活组加cross_runtime标记。scripts/generate_semantic_inventory.py:新增print_merge_candidates,--report调用它;打印计数行并把跨运行时组排前。docs/.../semantic-vocabulary-convergence-v0.md及其中文镜像:记录新的计数与含义。tests/architecture/test_semantic_inventory.py:+118 行,钉住过滤精度、cross_runtime、无 registry 的旧行为与打印内容。
我复核的关键点(都在这个 head 上自己跑过):
python scripts/generate_semantic_inventory.py --report→ exit 0,打印merge candidates (advisory, never auto-merged): 20 to review, 18 explained by a registered vocabulary's own owner symbols, 38 raw groups,并把两个跨运行时组排在最前。- 独立复算过滤精度:直接调用库函数得到 raw 38 / filtered 20,被丢掉的 18 组全部能映射到某个词表的 owner 符号集(
unexplained_drops = 0);两个存活的跨运行时组(MATERIAL_DELIVERY_OUTCOMES/VISION_OUTCOME_CHECKPOINT_MATERIAL_OUTCOMES、MONITOR_METADATA_FIELDS/TODO_MONITOR_METADATA_FIELDS)都不属于任何词表,是真实缺口。 pytest -q tests/architecture/test_semantic_inventory.py→ 18 passed;examples/semantic-vocabulary-drift-smoke.py→ exit 0。
遗留问题(非阻塞,P3)
新块不吃 --top:它上面两块(consumer ranking、divergent value sets)都按 [: args.top] 截断,--top 的说明还是「rows to print with --report」(默认 25)。今天 20 行看不出差异,但这块存在的意义正是让不断累积的分类清单可读,而候选空间是 38。建议同样按 args.top 截断(截断时打印 showing N of M),或在 --top 帮助里写明候选列表刻意全量。
对主干的风险
最强回归不是崩溃,而是丢掉了仍需评审的组。我用真实仓库反向验证过:只有「名字集合恰好等于某个词表 owner 符号集」的组会被丢,unexplained_drops = 0,所以 3 名以上的组和未注册的组都不受影响;registry=None 的旧路径原样返回 38 组。另一条已有护栏兜住的风险是「词表声明本身漂移导致误判为已解释」——那种漂移由 drift smoke 的 owned-vocabulary 检查单独报错,不会静默。除此之外:纯读路径,无写、无状态、无 quota,回退成本一个 commit。本次 head 只多了 main 合并,我重新对当前 main 比对了内容增量并重跑了上述证据,没有继承上一轮结论。
我的整体评价
结论 APPROVE。方向正确的可见性增量:让派生报告反映注册表已有的裁定,而不是再造一个权威;过滤条件窄到「恰好相等」,无 registry 的行为不变,还顺手把两个真实的跨运行时注册表缺口顶到最前。我独立复算 38/20/18 与丢弃集合、跑通 18 个测试与 drift smoke,确认报告真的打印这些行。唯一 P3 是 --top 在该块上不生效,属显示一致性问题。
English verdict: APPROVE - exact head 6f11b95 (re-review after a main merge; content delta versus current main unchanged at 5 files, +218/-11). The registry-aware filter is precise: I re-derived 38 raw / 20 to review / 18 explained independently, found zero unexplained drops, and confirmed both surviving cross-runtime pairs are owned by no vocabulary while the report prints them first. 18 inventory tests pass and the drift smoke exits 0. One non-blocking P3: the new merge-candidate block ignores --top while the two blocks above it honour it.
What
The merge-candidate report (
scripts/generate_semantic_inventory.py --report) now prints the candidate groups the registry does not explain, and the candidate grouping itself filters out registered owner pairs.Before:
merge_candidate_groups()grouped by "identical value set + different name" without consulting the registry's declared owners, so 18 of 38 groups were the Python and TypeScript ends of the same registered vocabulary (e.g.LoopXTurnRoute↔TURN_ROUTES) — tool noise, not debt. The CLI printed none of it, so the 38 groups were invisible and unsortable.After: 38 → 20 groups (2 cross-runtime registry gaps + 18 pure-Python candidates needing classification), and
--reportprints them so the remaining work is visible and sortable.The two cross-runtime gaps this surfaces
These two pairs have identical value sets across runtimes but no vocabulary owner declaration in the registry:
MATERIAL_DELIVERY_OUTCOMES(outcome_continuity.py) ↔VISION_OUTCOME_CHECKPOINT_MATERIAL_OUTCOMES(delivery_outcome.ts)MONITOR_METADATA_FIELDS(todos/contract.py) ↔TODO_MONITOR_METADATA_FIELDS(todos/monitor_metadata.ts)They should either be registered as
cross_runtimevocabularies or annotated with why both ends are intentionally held.Classification stays out of scope
This branch makes candidates visible, not classified. The 18 pure-Python groups still need the four-label classification (same-concept-same-scope / same-concept-different-scope / unrelated collision / shared shape) before any merge — renaming to make groups disappear scores nothing (I14).
Known interactions
inventory.pycandidate logic and the report generator); resolution is additive — keep both the owner-pair filter and the rename-invariance evidence.codex/lock-inventory-ratchets,codex/invariant-domain,codex/semantic-parse-memoization(perf(semantics): parse each Python source once per inventory run #4628).Validation
python3.11 examples/semantic-vocabulary-drift-smoke.py— green, output unchanged (the smoke does not consume the report path)python3.11 -m pytest -q tests/architecture/test_semantic_inventory.py— 18 passed (3 new tests cover the owner-pair filter and the report output)