docs(semantics): name what measures each target surface, not a snapshot of it - #4627
Conversation
exact-head 复核(
|
| 表面 | 基线 | 2026-09-17 实测 | 变化 |
|---|---|---|---|
effective_action |
33 值、无 owner | 32 值、一个枚举 owner、5 个已退休 | 仍是单槽位、仍是 merge_candidate——剩的是 M1 拆槽,不是数量 |
| Turn 词表 | 3 / 28 / 21 | 3 / 28 / 21 | 未变;投影与 29 条规则已生成校验 |
| 语义分叉 | 18 | 18 | 未变;两个在开 PR 合并后为 11 |
| 语义冲突 | 2 | 0 | 已达目标 |
| 多值分叉 | 4(1 个误分类) | 原始 4 / 语义 3 | SOURCE_SURFACES 已声明为 bounded_context |
| 多值孪生 | 19 | 19 | 未变 |
| 旧字段 | 124 py / 10 ts | 109 py / 10 ts | 按标识符计数,每个字段都低于基线 |
| 合并候选 | 32 | 38 | 唯一增长的表面 |
| py/ts 孪生 | 43 | 43 | 未变 |
这一列暴露的三件事
- M2 先于 M1 落地——与路线图里的依赖箭头相反。契约生成不需要等拆槽,所以它自己先走了。
- 语义分叉的目标由两个在开 PR 达成,而不是由某个里程碑——
Reached by写的是"基线窄 PR",这不会告诉读者这项工作只差一次评审。 - 合并候选实测 38,不是 32。按历史数字排序会少算六组。这正是 issue 自己那句"remeasure before ordering work"要防的错,现在它有地方落脚了。
对主干的风险
纯文档:git diff --stat 两个文件都在 docs/architecture/rfcs/,无注册表、代码、预算或锚点改动。实测列不是承诺:每个值都带产生它的 PR 或日期,评审者可以核而不是信。中英两表同十行、同数字,单元格是翻译而非各自撰写。
验证(当前 exact head)
python3 examples/docs-governance-smoke.py # ok
两份表均解析为 12 行、每行 5 格
edf4e28 to
806a74e
Compare
exact-head 复核(
|
| 表面 | 基线 | 实测 |
|---|---|---|
effective_action |
33 值、无 owner | 32 值、一个 owner、5 个已退休;仍是单槽位、仍是 merge_candidate |
| 语义冲突 | 2 | 0 |
| 语义分叉 | 18 | 18(两个在开 PR 后 → 11) |
| 旧字段 | 124 py / 10 ts | 109 py / 10 ts |
| 合并候选 | 32 | 原始 38,真实 20 |
| py/ts 孪生 | 43 | 43 |
验证(当前 exact head)
python3 examples/docs-governance-smoke.py # ok
两份表均解析为 12 行、每行 5 格
污染检查:三处命中数均为 0
本 PR 在 #4447 计划中的位置issue #4447 现在有一节统一协调(中英双语),把这 13 个在开 PR 作为一个计划列出:各自修什么、为何必要、以及实测出的合并顺序。 冲突实测:对全部 78 对做了试合并,9 对冲突,分四簇,每一处都是文本相邻,没有一处是语义分歧。
建议顺序(代价从低到高):#4628 → #4625、#4626 → #4627 → #4619、#4621 → #4630 → #4614 → #4631 → #4629 → #4617 → #4606 → #4608。四个棘轮 PR 放最后,因为每落地一个,下一个的数字就从估算变成确定值。 全部 13 个 PR 现已同步到 |
…ot of it Section 11's target table gives a frozen baseline and an end state, and nothing says how a reader finds the current value. The first version of this PR filled that gap with a dated `Measured 2026-09-17` column. That was the wrong fix, for two reasons this rework removes. It was in the wrong place. The document map says Section 11 is the normative delivery plan and that "dated progress entries do not amend normative sections", and issue loopx-project#4447 owns delivery status. A dated column inside the normative table, carrying a note claiming it is not normative, argues with the maintenance contract instead of following it. It was also redundant work with a failure mode. Eight of the ten surfaces are already printed by the drift smoke on every run -- `same_runtime_forks_semantic`, `multi_value_twins`, one `<field>.py` / `<field>.ts` pair per legacy field, `independently_maintained`. A hand-written column transcribes output a command already produces, goes stale on the next merge, and invites exactly the error the first version shipped: a merge-candidate count copied from a raw grouping without the registry filter the surrounding numbers carry. So the column now names the command and field that measure each surface, the way Section 9 names a test for each claim, and carries no values at all. A reader who wants the current state runs the command. Two surfaces honestly say no counter exists: the `effective_action` slot split is a Q6 decision, and merge candidates are only reachable through `merge_candidate_groups()` because no command prints them yet. Where a capability is not on `main`, the cell says which PR adds it rather than describing it as present. Dated values stay in issue loopx-project#4447, which the RFC already designates as the owner of delivery status. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
806a74e to
1b67f3a
Compare
exact-head 复核(
|
| 带日期的数值列 | 现在的「由什么度量」列 | |
|---|---|---|
| 与维护契约的关系 | 把带日期的进度放进规范章节,再附声明说它不规范 | 与第 9 节同惯例,点明检查手段 |
| 下次合并后 | 过期 | 不变 |
| 与 issue #4447 | 重复它拥有的交付状态 | 互补:这里是「真值在哪被度量」,那里是「当天的值」 |
| 誊抄出错的可能 | 就是这么错的 | 无数可誊 |
对主干的风险
纯文档,两个文件都在 docs/architecture/rfcs/。无数值、无日期——这正是重点。所引用的每个 smoke 字段都已在真实 smoke 输出中核对存在,所引用的每条命令都已核对在 main 上存在。
python3 examples/docs-governance-smoke.py # ok
…ogress-column Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: 73a07bf551c696312756eac55e739e4502d8732c (codex/rfc-measured-progress-column).
动机
#4447 的文档切片:RFC 第 11 节是规范性交付计划,它的目标表原本只写「表面 / 基线 / 目标 / 由谁达成」,没有说这些表面在哪里被度量。结果就是十个基线里有八个其实每次跑 smoke 都会打印,但文档从没说过;读者只能「相信」而不能「核对」计划,而且誊抄的数字会在下一次合并后过期。
改动后表格增加 Measured by 列,点名今天打印每个表面的命令与字段(或注册表键),不带任何数值、不带日期,并明确两处「今天没有计数器」。这正是作者上一版(带 Measured 2026-09-17 数值列)被否后的修正方向:文档地图规定带日期的进度条目不改写规范章节,而带日期的数值归 issue #4447 管;把「真值在哪里被度量」写进规范表,才既不过期也不越权。
改动思路
入口是 docs/architecture/rfcs/semantic-vocabulary-convergence-v0.md 第 11 节表与其 zh-CN 镜像。权威来源是真实存在的产生者:examples/semantic-vocabulary-drift-smoke.py 打印的计数(multi_value_twins、same_runtime_forks_semantic、conflicting_values_semantic、independently_maintained、每个旧字段一对 <field>.py/<field>.ts)、注册表 vocabularies.effective_action.values、loopx/semantics/inventory.py::merge_candidate_groups(),以及 scripts/generate_semantic_inventory.py --report。
判定边界是「列只记来源,不记数值」:数值归 #4447,读者按需跑命令。两处诚实标注「无计数器」(effective_action 拆槽是 Q6 决策而非数字;合并候选只有函数、今天没有命令打印),以及「尚未在 main 的能力要点名是哪个 PR 加」的写法,都是这份 PR 的核心价值。
具体改动
2 个文件、+45/-24,全部在 docs/architecture/rfcs/ 下:英中两版目标表各加一列(原三列内容逐字保留),并在表后各补两段说明(「Measured by 只给来源不给数值」「合并候选要读可评审数而非原始数」)。
我核到的可核对结果:python examples/docs-governance-smoke.py → ok;被引用的 smoke 字段在真实输出里全部存在(multi_value_twins=13/13、same_runtime_forks_semantic=11/11、conflicting_values_semantic=0/0、independently_maintained=43/43、以及 execution_obligation/heartbeat_recommendation/work_lane_contract/external_evidence_observation/goal_boundary/protocol_action_packet 六对 .py/.ts);英中两表列数与行数一一对应。
关键内容讲解
merge_candidate_groups()那一格:作者写「今天没有任何命令打印它,#4630 增加该 CLI 行」,我确认属实——loopx/cli_commands/*里没有任何 merge-candidate 打印,rg全仓只有文档、测试与inventory.py自身引用它。- 「无计数器」的两格也站得住:
effective_action的槽位拆分确实是 Q6 决策(没有可数的量),所以指向relations.shared_field_names;合并候选今天只能通过函数取。
对主干的风险
这份表本身是「声明值从哪来」的规范性文字,而 docs-governance-smoke 只查结构与链接、不查这类陈述的真假,所以每一格都要靠人对着树核。我逐格核对后,有一格是过期的,另有一格的指引今天没有落地。
一条 P2(F1,非阻塞):Multi-value forks 那格把已合并的 #4614 写成了待办,并且与本文件第 9 节自相矛盾。 新格写「Only the count is printed today; #4614 adds divergent_value_sets to name the surviving forks」(中文:「#4614 增加 divergent_value_sets 以按名字列出存活的分叉」)。但 #4614 已于 2026-09-17T09:28:53Z 合并(其提交 ddc240c4f/a0826c93e 都在本 PR 的 merge-base 897e9aedb 上),divergent_value_sets 就在本 head 的 loopx/semantics/inventory.py:435,并由 scripts/generate_semantic_inventory.py --report 打印(value_sets name definition_modules 段);更直接的是同一份 RFC 第 9 节第 758 行本来就写着「divergent_value_sets(inventory), printed by --report, …」,用的是现在时。本 head 在 10:19Z 合并了 main(晚于 #4614 落地),所以这句话在这个 exact head 上已经过期;PR 正文里「我硬着头皮核过:第一版引用的 --report 在今天的 main 上都不存在」这一条对 #4614 而言不成立。最小修复:把这句改成现在时并点名真实产生者(scripts/generate_semantic_inventory.py --report 打印按名字的 divergent_value_sets),中文同步;#4630 那半句保留,它是对的。
一条 P3(F2,非阻塞):「读可评审数而非原始数」这条指引今天没有产生者。 该格点名的是 merge_candidate_groups(),但在本 head 上它的签名是 merge_candidate_groups(inventory)(inventory.py:406-432),完全不看注册表,会把「注册的跨运行时词表按其 Python/TypeScript 两个符号」也配成一组——也就是单元格自己说的那类度量伪影。真正加上注册表过滤的是 #4630(其 diff 新增 registered_owner_symbol_sets 与可选 registry 参数,正好丢掉这些配对)。所以读者今天按格子的指引去「读可评审数」是读不到的。建议在该格说明 #4630 同时提供注册表过滤与 CLI 行(这样可评审数才有产生者),或在 #4630 落地前先不加这句。
我的整体评价
结论 APPROVE。方向是对的且比上一版明显更好:用「点名产生者」代替「誊抄数值」,让第 11 节从不可核对变成可核对,两处诚实承认没有计数器,并把未落地的能力点名为 PR 而不是写成既成事实。我逐格把十个表面回到今天的树核对过,smoke 字段全部存在、docs-governance-smoke 通过、英中两版一致。
两条都是文档准确性问题、非阻塞:一条是 #4614 那半句已过期且与本文件第 9 节冲突(P2),一条是「可评审数」的过滤其实随 #4630 才到(P3)。考虑到这张表存在的意义就是「让计划可被核对而不是被相信」,这两处值得在合并前顺手改掉——它们恰好落在同一张表里。
English verdict: APPROVE - exact head 73a07bf; a docs-only change that adds a "Measured by" column to the RFC's normative Section 11 target table (both languages), naming the real producer per surface and carrying no values, which is the correct repair of the earlier dated-column version. I verified every cited smoke field in real smoke output, confirmed no command prints merge_candidate_groups today (#4630 open), ran docs-governance-smoke (ok), and checked the two tables line up. Two non-blocking findings: the multi-value-fork cell describes the already-merged #4614 as pending even though divergent_value_sets is on the merge-base and printed by scripts/generate_semantic_inventory.py --report, contradicting Section 9 of the same document; and the instruction to read the reviewable merge-candidate count cites merge_candidate_groups(), which is unfiltered on this head because the registry filter arrives with #4630.
Six tracker PRs landed while this branch was open. Three touched files it also edits, in the way loopx-project#4447's merge-order note predicted: - loopx-project#4626 and loopx-project#4625 append to the end of `test_semantic_vocabulary_drift.py`; both blocks are kept, theirs first. - loopx-project#4627 replaced the Section 11 target table with a *Measured by* column and a rule that the table carries no dated values, since those belong to the tracker. This branch's row had added dated numbers, so the resolution takes loopx-project#4627's table and puts the migration surface in *Measured by* as the `--report` line that prints it. The dated table stays in Appendix A. - loopx-project#4628 memoized `python_facts`. The retirement scan needs the tree rather than the facts, so `parse_python` is factored out for one error path and left uncached: caching the trees held about two million AST nodes for the rest of the run and measured 0.7s worse overall, while slowing `check_inventory` from 2.4s to 5.4s -- the pass loopx-project#4628 had just made cheaper. Remeasured on the integrated tree: every role count is unchanged, and `dynamic_mapping_key_sites` moved 1704 to 1712 with the new code. Both mirrors carry the new number. `loopx/semantics/field_use.py` also had to stop spelling the six field names in its own docstrings. The scan reads tracked sources under `loopx/`, this module is one of them, and committing it pushed `heartbeat_recommendation` to 18 of a budget of 17 -- the check catching its own module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Refs #4447 — Section 11 says where each surface is measured, so the plan can be checked instead of trusted.
Why the first version was wrong
The first version of this PR added a dated
Measured 2026-09-17column. Two things were wrong with it, and both are instructive.Wrong place. The document map says Section 11 is the normative delivery plan and that "dated progress entries do not amend normative sections", and issue #4447 owns delivery status. A dated column inside the normative table — carrying a note asserting it is not normative — argues with the maintenance contract instead of following it.
Redundant, with a failure mode. Eight of the ten surfaces are already printed by the drift smoke on every run:
same_runtime_forks_semantic,conflicting_values_semantic,multi_value_twins, one<field>.py/<field>.tspair per legacy field,independently_maintained. A hand-written column transcribes output a command already produces, goes stale on the next merge, and invites precisely the error the first version shipped — a merge-candidate count copied from a raw grouping without the registry filter the surrounding numbers carry.What this version does
The column names the command and field that measure each surface, the way Section 9 names a test for each claim, and carries no values at all.
semantic-vocabulary-drift-smoke.py:same_runtime_forks_semanticconflicting_values_semanticmulti_value_twins<field>.py/<field>.tspair per fieldindependently_maintainedeffective_actionvaluesvocabularies.effective_action.valuesTwo surfaces honestly say no counter exists: the
effective_actionslot split is a Q6 decision rather than a number, and merge candidates are reachable only throughmerge_candidate_groups()because no command prints them yet.Where a capability is not on
main, the cell says which PR adds it rather than describing it as present — #4614 for naming surviving forks, #4630 for the merge-candidate CLI line. I checked this the hard way: my first draft cited--reportfor both, and neither exists onmaintoday.Boundary
docs/architecture/rfcs/.Validation
Every cited smoke field verified present in real smoke output; every cited command verified to exist on
main.🤖 Generated with Claude Code