Skip to content

docs(semantics): say which condition produces each effective_action value - #4626

Merged
huangruiteng merged 3 commits into
loopx-project:mainfrom
songoow:codex/effective-action-value-notes
Sep 17, 2026
Merged

huangruiteng merged 3 commits into
loopx-project:mainfrom
songoow:codex/effective-action-value-notes

Conversation

@songoow

@songoow songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator

Refs #4447 — Track A / documentation half of the registry contract. Companion to #4625, independent of it.

Why

effective_action is the most overloaded registered slot: 32 values under one field name, carrying a decision verdict, a frontier verdict and a replay phase at once. It is the vocabulary M1 is due to split, and the one where a bare token tells a reader the least.

Five of its 32 values carried a note, and all five recorded why the literal scan had missed them — not what the value means. So the registry said quota_skip was legal without saying it is the ladder's final fallback, and listed five near-identical *_projection_repair names with nothing to tell them apart.

What

The remaining 27 values are documented against the code that decides them.

The quota_effective_action ladder now reads in order — normal_run → outcome_floor_recovery → agent_workspace_repair → self-repair → capability_bridge_repair → the blocked states → quota_skip, which is named as the final fallback rather than a generic skip.

The five self-repair spend actions are separated by trigger, not by name:

Value Trigger
control_plane_health_repair health_blocker
control_plane_projection_repair waiting_without_owner_projection
state_projection_gap_repair state_projection_gap
boundary_projection_repair required_write_scope_missing_from_goal_boundary
todo_decision_scope_projection_repair user_gate_scope_projection_drift

Routing consequences are stated where they exist. governed_capability_intent is the only value that reaches capability_action_required, and only when the intent projection matches the envelope goal and agent and names a command. autonomous_replan_required is a REPLAN_ACTION. Every *_repair name routes to repair_required. terminal_no_followup is what the controller reads as its terminal_action partition.

The four values a name alone misleads on are spelled out: monitor_due and monitor_quiet_skip are the two arms of one branch; coordinate_task_bundle replaces normal_run rather than adding to it; scoped_user_gate_fallback replaces quota_skip or monitor_quiet_skip; unsettled_host_turn_recovery drops the selected Todo and refuses all three delivery modes.

The ratchet test is extended to effective_action and lease_action, so a new value in either fails the PR path until the same diff says what produces it. Mutation-checked with a whitespace-only note.

Reading the diff

The five pre-existing notes are preserved verbatim and reordered to follow the registered value order, which is why they appear as moved lines. Verified programmatically: every prior note string is unchanged.

Boundary

  • No behaviour, budget or inventory count changes. value_notes is registry documentation; the drift smoke already validates that it names only registered values.
  • These notes describe today's producers. When M1 splits the slot, the notes move with their values rather than being rewritten — which is the point of recording the producing condition rather than a paraphrase of the name.
  • This PR touches effective_action only; docs(semantics): say what each Turn kernel vocabulary value means #4625 covers the four canonical Turn kernel vocabularies. Neither depends on the other, and they edit different regions of both files.

Validation

python3 examples/semantic-vocabulary-drift-smoke.py                  # ok
python3 -m pytest tests/architecture/test_semantic_vocabulary_drift.py \
                  tests/architecture/test_semantic_inventory.py -q   # 75 passed

Every smoke counter is byte-identical to the baseline run, including unresolved_producer_sites=41 and conflicting_values_semantic=0/0.

Environment note: without node_modules the TypeScript production parser raises and 16 tests in that file fail on a clean origin/main tree as well. With the Node dependencies present, the tree is green.

🤖 Generated with Claude Code

…alue

`effective_action` is the most overloaded registered slot: 32 values under one
field name, carrying a decision verdict, a frontier verdict and a replay phase
at once. It is the vocabulary M1 is due to split, and the one where a bare
token tells a reader the least. Five of its 32 values carried a note, and all
five recorded why the literal scan had missed them rather than what the value
means.

Document the remaining 27 against the code that decides them:

- The `quota_effective_action` ladder is now readable in order: `normal_run`,
  `outcome_floor_recovery`, `agent_workspace_repair`, self-repair,
  `capability_bridge_repair`, then the blocked states, with `quota_skip` named
  as the final fallback rather than as a generic skip.
- The five self-repair spend actions are separated by their trigger
  (`health_blocker`, `waiting_without_owner_projection`, `state_projection_gap`,
  `required_write_scope_missing_from_goal_boundary`,
  `user_gate_scope_projection_drift`) instead of by five similar names.
- Routing consequences are stated where they exist: `governed_capability_intent`
  is the only value that reaches `capability_action_required`, and only with a
  matching intent projection; `autonomous_replan_required` is a REPLAN_ACTION;
  every `*_repair` name routes to `repair_required`; `terminal_no_followup` is
  what the controller reads as its `terminal_action` partition.
- The four values a name alone misleads on are spelled out: `monitor_due` and
  `monitor_quiet_skip` are the two arms of one branch, `coordinate_task_bundle`
  replaces `normal_run` rather than adding to it, `scoped_user_gate_fallback`
  replaces `quota_skip` or `monitor_quiet_skip`, and
  `unsettled_host_turn_recovery` drops the selected Todo and refuses all three
  delivery modes.

The five pre-existing notes are preserved verbatim and reordered to follow the
registered value order, which is why the diff shows them moving.

Extend the note-coverage ratchet to `effective_action` and `lease_action`, so
a new value in either fails the PR path until the same diff says what produces
it. Mutation-checked with a whitespace-only note.

No behaviour, budget or inventory count changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

exact-head 复核(c98f3f9d6)

为什么挑这个词表

effective_action 是注册表里最重载的一个槽:一个字段名下 32 个取值,同时承载决策判定、frontier 判定和 replay 阶段。它是 M1 计划要拆开的那个槽,也是"光看 token 最读不出含义"的那个。

现存 5 条注记记录的是字面量扫描为何漏掉它们,不是它们的含义。于是注册表说 quota_skip 合法,却不说它是整条阶梯的最后兜底;列出五个长得几乎一样的 *_projection_repair,却没有任何东西能把它们区分开。

27 条注记的取材

  • quota_effective_action 阶梯按顺序可读:normal_run → outcome_floor_recovery → agent_workspace_repair → self-repair → capability_bridge_repair → 各 blocked 状态 → quota_skip(点名为最后兜底)。
  • 五个 self-repair 按 trigger 区分,而不是按名字:health_blocker、waiting_without_owner_projection、state_projection_gap、required_write_scope_missing_from_goal_boundary、user_gate_scope_projection_drift。
  • 有路由后果的都写出后果:governed_capability_intent 是唯一能到达 capability_action_required 的取值,且需 intent 投影与信封 goal/agent 匹配并带 command;autonomous_replan_required 属 REPLAN_ACTIONS;所有 *_repair 走 repair_required;terminal_no_followup 正是控制器读作 terminal_action 分区的那个值。
  • 四个"名字会误导"的单独点出:monitor_due 与 monitor_quiet_skip 是同一分支的两臂;coordinate_task_bundle 是替换 normal_run 而非追加;scoped_user_gate_fallback 替换 quota_skip/monitor_quiet_skip;unsettled_host_turn_recovery 会丢弃已选 Todo 并同时拒绝三种交付模式。

怎么读这个 diff

原有 5 条注记逐字保留,只是重排到与注册取值同序,所以它们表现为移动行。已用程序核对:每条旧注记字符串未变。

对主干的风险

零行为变更;预算与清单计数与基线逐项一致。棘轮测试扩到 effective_action 与 lease_action,两者新增取值时不写注记即在 PR 路径失败。

值得注意的一点:这些注记描述的是今天的生产者。M1 拆槽时,注记随取值一起走,而不需要重写——这正是记录"产生条件"而非"名字的同义改写"的意义。

验证(当前 exact head)

semantic-vocabulary-drift-smoke: ok        # 全部计数器与基线逐字节一致
pytest tests/architecture/test_semantic_vocabulary_drift.py \
       tests/architecture/test_semantic_inventory.py -q  ->  75 passed
变异验证:把 monitor_due 注记改成纯空白 -> 测试失败并点名该取值

与 #4625 相互独立:两者改的是同一文件的不同区域,任一先合并都不阻塞另一个。

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

详细评审(exact head c98f3f9d6)

动机

effective_action 是注册表里最"超载"的一个槽位:一个字段名下面同时装着决策裁决、frontier 裁决和 replay 阶段三类含义,共 32 个值,也正是 M1 计划要拆分的那一个。作者给出的基线是:32 个值里只有 5 个带 note,而这 5 条的正文记的都是"字面扫描当初为什么漏了它",不是"这个值是什么意思"。于是注册表告诉读者 quota_skip 合法,却不说它是阶梯的最后兜底;五个名字几乎一样的 *_repair 也无法从注册表里区分。这个 gap 是真的,而且正好落在"裸 token 信息量最低"的那个槽位上。

改动思路

与 #4625 同一 Track A 的另一半(两者相互独立):把剩下 27 个值的含义写在值旁边,对着产生它们的代码写;既有的 5 条 provenance note 逐字保留、只是挪进同一个块内以保持值序,因此 diff 看起来"移动"了它们。再把 note 覆盖率棘轮扩到 effective_action 与 lease_action,让以后新增值必须在同一个 diff 里说明"什么条件产生它"。方向对:含义与注册表同源,读者不必再回读整条阶梯。

具体改动

  • loopx/semantics/vocabulary_v0.json(+41/-7):quota_effective_action 阶梯按顺序可读——normal_run → outcome_floor_recovery → agent_workspace_repair → 自修复 → capability_bridge_repair → 各 blocked 态 → quota_skip(明确写成最后兜底而非泛泛的 skip);五个 stall 自修复支出动作按触发条件而非名字区分(health_blocker、waiting_without_owner_projection、state_projection_gap、required_write_scope_missing_from_goal_boundary、user_gate_scope_projection_drift);路由后果写在有的地方(governed_capability_intent 是唯一能走到 capability_action_required 的值且需要 intent 匹配、autonomous_replan_required 属于 REPLAN_ACTIONS、*_repair 一律走 repair_required、terminal_no_followup 就是控制器的 terminal_action 分区);四个"名字会误导"的值被点明(monitor_due/monitor_quiet_skip 是同一分支的两臂、coordinate_task_bundle 是替换 normal_run、scoped_user_gate_fallback 替换 quota_skip/monitor_quiet_skip/缺省动作、unsettled_host_turn_recovery 丢弃所选 Todo 并拒绝三种投递模式)。原有 5 条 note 逐字保留。
  • tests/architecture/test_semantic_vocabulary_drift.py(+20):新增 test_remaining_kernel_values_each_carry_a_note,参数化为 effective_action 与 lease_action,空串与纯空白都算缺失;用文件里既有的 runpy 加载器读注册表。

验证(都在 c98f3f9d6 上):pytest tests/architecture/test_semantic_vocabulary_drift.py -q → 60 passed(main 同环境 58 passed,差值正是新增 2 条);pytest tests/architecture/test_semantic_production.py tests/capabilities/test_pr_review_contract.py -q → 65 passed;examples/semantic-vocabulary-drift-smoke.py → exit 0,conflicting_values=16/18、kernel_producer_coverage_pending 为空,说明 note 没有挪动语义冲突预算或清单计数。

对主干的风险

我逐个把 note 与产生它的代码对齐,没有发现不符:阶梯顺序与 decision_summary.quota_effective_action(388–411 行)逐条一致,包括 agent_workspace_repair 优先于自修复、capability_bridge_repair 在两个修复之后、quota_skip 是最终 return;五个 stall 动作的 effective_action/trigger 在 stall_repair.py 里成对出现(health_blocker → CONTROL_PLANE_HEALTH_REPAIR,waiting_without_owner_projection → CONTROL_PLANE_PROJECTION_REPAIR,其余三个触发名与模块常量一致);automation_prompt_upgrade_required 确实在 autonomous_replan_decision_allowed 里显式 withhold;unsettled_host_turn_recovery 确实 pop 掉 selected_todo/todo_id/action_portfolio 并把 normal/recovery/self-repair 三种投递全部置 False;scoped_user_gate_fallback 确实只在原动作是 quota_skip/monitor_quiet_skip/缺省时替换;peer_coordination_blocked 在 scheduler/arbitration.py 映射为 PEER_COORDINATION_STOP(而不是 monitor wait),monitor_quiet_skip 映射为 MONITOR_WAIT,heartbeat_settled_skip 映射为 QUIET_WAIT,agent_monitor_only 映射为 AGENT_MONITOR_ONLY_WAIT——与四条 note 的说法一致。

风险面很小:没有任何生产行为、值集、schema 或预算变化,回滚只需撤掉文本与两条测试参数。唯二限制仍是:(1) 棘轮保证"有 note",不保证"note 正确"(正确性靠人读代码,我按上面方式代读了一遍);(2) M1 对这个槽位本身的拆分不在本 PR 内。

同样提醒一个环境性现象,避免误判:在没有 node_modules 的干净 worktree 里,tests/architecture/test_semantic_vocabulary_drift.py 会报 16 条失败(TypeScript production parser 不可用),origin/main 上同样 16 条,主检出(带依赖)58/58 全绿;本 head 补上依赖后 60/60 全绿。这不是本 PR 引入的回归。

我的整体评价

APPROVE。 目标清晰、范围克制、保留了既有的 provenance note 而不是覆盖掉它们,并把棘轮扩到了这两个词表;我用独立复算确认结论(逐值比对产生代码;自己把 quota_skip 的 note 改成空白后,恰好只有 effective_action 参数失败,恢复后 2/2 通过)。剩下的(M1 拆分)是明确写在范围外的后续工作。

关键代码讲解

  • loopx/control_plane/quota/decision_summary.py::quota_effective_action:阶梯就是本 PR 的主线——第一次 return 决定一切,所以 note 写成"阶梯第几位 + 什么条件"比"名字释义"有用得多;quota_skip 是最后一个 return,note 才敢说它是兜底而非跳过。
  • loopx/control_plane/quota/stall_repair.py:五个 *_repair 值由不同 trigger 产生(health_blocker 走 293–300 行、waiting_without_owner_projection 走 336 行附近),note 用 trigger 而不是用名字把它们分开,这正是评审这类值最需要的信息。
  • loopx/control_plane/quota/unsettled_host_turn.py(274–300 行):unsettled_host_turn_recovery 会 pop 掉 selected_todo/todo_id/action_portfolio,并把三种投递开关全部置 False——note 把这三件事都写出来了,读者不必再从 payload 变更里推断。
  • tests/architecture/test_semantic_vocabulary_drift.py::test_remaining_kernel_values_each_carry_a_note:与 #4625 的棘轮同形,参数改为 effective_action/lease_action;lease_action 的 4 个值目前都是 compatibility-only 说明,因此这条断言同时守住了"兼容词表也要说清处置"。

English verdict: APPROVE - reviewed exact head c98f3f9. It documents the 27 remaining effective_action values against the code that selects them, keeps the five pre-existing provenance notes verbatim, and extends the note-coverage ratchet to effective_action and lease_action. I verified every claim against its producer: the ladder order matches decision_summary.quota_effective_action lines 388-411 with quota_skip as the terminal return; the five stall-repair actions pair each action with its trigger in stall_repair.py (health_blocker, waiting_without_owner_projection, state_projection_gap, required_write_scope_missing_from_goal_boundary, user_gate_scope_projection_drift); automation_prompt_upgrade_required really is withheld inside autonomous_replan_decision_allowed; unsettled_host_turn_recovery drops selected_todo/todo_id/action_portfolio and sets normal, recovery and self-repair delivery to false; scoped_user_gate_fallback substitutes only for quota_skip, monitor_quiet_skip or a missing action; and the scheduler arbitration turns peer_coordination_blocked into a stop, monitor_quiet_skip into a monitor wait, heartbeat_settled_skip into a quiet wait and agent_monitor_only into an agent-monitor wait. Validation at this head: drift file 60 passed (main 58 in the same environment, the delta being the two new cases), semantic_production plus pr_review_contract 65 passed, drift smoke exit 0 with unchanged conflicting_values and empty kernel_producer_coverage_pending. My own whitespace mutation on the quota_skip note fails exactly the effective_action parameter and passes 2/2 after restore. Note for other reviewers: the drift file reports 16 failures in clean worktrees lacking node_modules, identically on origin/main, so those are environmental and not introduced here.

…n-value-notes

Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

本 PR 在 #4447 计划中的位置

issue #4447 现在有一节统一协调(中英双语),把这 13 个在开 PR 作为一个计划列出:各自修什么、为何必要、以及实测出的合并顺序。

冲突实测:对全部 78 对做了试合并,9 对冲突,分四簇,每一处都是文本相邻,没有一处是语义分歧。

冲突簇 涉及 PR 后合并者的解法
棘轮锚点 #4606、#4608、#4617、#4629 四者全部合并后的终态已实测:same_runtime_forks 18、same_runtime_fork_definitions 41、same_runtime_forks_semantic 11、conflicting_values 16、conflicting_definitions 55、multi_value_twins 13。registry 与 BUDGET_ANCHOR 两处都要带同一组值(该检查是相等而非 <=)
漂移测试文件 #4626、#4629、#4631 三者都在文件末尾追加,全部保留即可
RFC 双镜像 #4614、#4627、#4631 附录 B 决策行,按日期先后保留两条
清单与生成器 #4614、#4630 加法式:owner 对过滤与改名不变性证据同时保留

建议顺序(代价从低到高):#4628 → #4625、#4626 → #4627 → #4619、#4621 → #4630 → #4614 → #4631 → #4629 → #4617 → #4606 → #4608。四个棘轮 PR 放最后,因为每落地一个,下一个的数字就从估算变成确定值。

全部 13 个 PR 现已同步到 main、零失败检查。

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复确认评审(exact head c0e52df3d)— 分支已 rebase,内容未变

动机

与首次评审相同:effective_action 是注册表里最超载的槽位(32 个值同时承载决策、frontier 与 replay 阶段三类语义),而 32 个值里只有 5 个带 note、且都只记"当初字面扫描为什么漏了它"。我在 c98f3f9d6 上已给出完整评审与结论(APPROVE)。这次 head 变化的唯一原因是作者合并了最新 origin/main。

改动思路 与 具体改动

本 head 的两个文件与 c98f3f9d6 逐字节相同(loopx/semantics/vocabulary_v0.json 与 tests/architecture/test_semantic_vocabulary_drift.py 的 md5 一致),diff 规模同样是 2 文件 +54/-7。也就是说:27 个 effective_action 值的含义说明、五个 stall 自修复动作按触发条件的区分、以及扩到 effective_action/lease_action 的覆盖率棘轮都没有变化;变化只来自被合并的主干提交。

对主干的风险

无新增风险。本 head 重跑:tests/architecture/test_semantic_vocabulary_drift.py、tests/architecture/test_semantic_production.py、tests/capabilities/test_pr_review_contract.py → 125 passed(drift 60 条含新增 2 条,另两份 65 条),与首次评审在旧 head 上的 60/65 一致;examples/semantic-vocabulary-drift-smoke.py 仍 exit 0(conflicting_values、kernel_producer_coverage_pending 未变)。首次评审中逐个把 note 与产生它的代码对齐(阶梯顺序对 decision_summary.quota_effective_action、五个 stall 触发对 stall_repair.py、automation_prompt_upgrade_required 对 autonomous_replan_decision_allowed、unsettled_host_turn_recovery 的三个投递开关、peer_coordination_blocked 与 monitor_quiet_skip 的调度处置)以及我自己做的 whitespace 变异验证,在内容不变的前提下继续成立。

我的整体评价

APPROVE。 纯 rebase 的复确认:改动内容与此前被批准的 head 逐字节一致,验证结果一致,结论不变。合并仍由维护者决定;本 head 的精确评审记录即此卡片。

English verdict: APPROVE - reviewed exact head c0e52df. Re-confirmation after a pure rebase: both PR files are byte-identical to c98f3f9 (md5 match) with the same 2-file +54/-7 diff, and the branch only merged origin/main. Re-run at this head: tests/architecture/test_semantic_vocabulary_drift.py plus test_semantic_production.py and test_pr_review_contract.py = 125 passed (60 including the 2 new cases, plus 65), and the drift smoke exits 0 with unchanged conflicting_values and empty kernel_producer_coverage_pending. The producer-by-producer comparison against decision_summary.quota_effective_action, stall_repair.py, autonomous_replan_decision_allowed, unsettled_host_turn.py and the scheduler arbitration, plus my whitespace mutation on the quota_skip note, remain valid because the content did not change.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复确认评审 — exact head c0e52df3de682d1a682f13fccb1bc29bb898e920

动机

与首次评审相同:effective_action 是注册表里最超载的槽位(32 个值同时承载决策、frontier 与 replay 阶段语义),而原先 32 个值里只有 5 个带 note、且都只记"当初字面扫描为什么漏了它"。我在 c98f3f9d6 上已给出完整评审与结论(APPROVE)。这次 head 变化的唯一原因是作者合并了最新 origin/main。

改动思路

不改内容,只做 rebase:本 head 的两个文件与 c98f3f9d6 逐字节相同(loopx/semantics/vocabulary_v0.json 与 tests/architecture/test_semantic_vocabulary_drift.py 的 md5 一致),diff 规模同样是 2 文件 +54/-7。27 个 effective_action 值的含义说明、五个 stall 自修复动作按触发条件的区分、以及扩到 effective_action/lease_action 的覆盖率棘轮都没有变化;变化只来自被合并的主干提交。

具体改动

  • loopx/semantics/vocabulary_v0.json:quota_effective_action 阶梯按顺序可读(normal_run → outcome_floor_recovery → agent_workspace_repair → 自修复 → capability_bridge_repair → blocked 态 → quota_skip 兜底);五个 stall 修复动作按触发条件区分(health_blocker、waiting_without_owner_projection、state_projection_gap、required_write_scope_missing_from_goal_boundary、user_gate_scope_projection_drift);四条"名字会误导"的值被点明;原有 5 条 provenance note 逐字保留。
  • tests/architecture/test_semantic_vocabulary_drift.py:test_remaining_kernel_values_each_carry_a_note,参数为 effective_action 与 lease_action。

本 head 上重跑:tests/architecture/test_semantic_vocabulary_drift.py、tests/architecture/test_semantic_production.py、tests/capabilities/test_pr_review_contract.py → 125 passed(drift 60 条含新增 2 条,另两份 65 条);examples/semantic-vocabulary-drift-smoke.py → exit 0,conflicting_values、kernel_producer_coverage_pending 未变。与首次评审在旧 head 上得到的 60/65 一致。

对主干的风险

无新增风险。首次评审中我逐个把 note 与产生它的代码对齐(阶梯顺序对 decision_summary.quota_effective_action、五个 stall 触发对 stall_repair.py、automation_prompt_upgrade_required 对 autonomous_replan_decision_allowed、unsettled_host_turn_recovery 的三个投递开关、peer_coordination_blocked/monitor_quiet_skip 的调度处置),并做过 whitespace 变异验证(抹白 quota_skip 的 note 恰好只让 effective_action 参数失败);这些在内容逐字节不变的前提下继续成立。合并仍由维护者决定。

我的整体评价

APPROVE。 纯 rebase 的复确认:内容与此前被我批准的 head 逐字节一致,验证结果一致,结论不变。本 head 的精确评审记录即此卡片。

English verdict: APPROVE - reviewed exact head c0e52df. Re-confirmation after a pure rebase: both PR files are byte-identical to c98f3f9 (md5 match) with the same 2-file +54/-7 diff, and only origin/main was merged in. Re-run at this head: tests/architecture/test_semantic_vocabulary_drift.py plus tests/architecture/test_semantic_production.py and tests/capabilities/test_pr_review_contract.py = 125 passed (60 including the two new cases, plus 65), and the drift smoke exits 0 with unchanged conflicting_values and empty kernel_producer_coverage_pending. The producer-by-producer comparison against decision_summary.quota_effective_action, stall_repair.py, autonomous_replan_decision_allowed, unsettled_host_turn.py and the scheduler arbitration, plus my whitespace mutation on the quota_skip note, remain valid because the content did not change.

…n-value-notes

Signed-off-by: song <22676124+songoow@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: d1ea66a5541f75490800ce319f70b1e31dbd0d3b (codex/effective-action-value-notes).

动机

#4447 Track A 的文档半边,与 #4625 是同一个改动理由的两个切片:effective_action 是注册表里最过载的槽位——一个字段名下 32 个值,同时装着决策判决、frontier 判决和 replay 阶段,而且正是 M1 要拆的那个词表。基线 main 只文档化了 5/32,而且那五条说明写的都是「当初字面扫描为什么漏了它」,不是值是什么意思;读者面对五个几乎同名的 *_projection_repair 无从区分,quota_skip 合法却没人说它是阶梯的最后兜底。

改动后 32 个值都对着产生它们的代码写清了条件:阶梯顺序、五个 self-repair 动作各自的 trigger、以及存在时的路由后果(哪个值是唯一进 capability_action_required 的、哪些名字结尾 _repair 必然路由 repair_required、terminal_no_followup 就是控制器读的 terminal_action 分区)。这是可独立复核的完整切片:把「今天谁产生这个值」记下来,把 M1 拆分留到它自己的变更里。

改动思路

入口是 loopx/semantics/vocabulary_v0.json(经 load_registry),权威来源是产生值的代码:decision_summary.py::quota_effective_action 的阶梯、_task_orchestration_effective_action、should_run_packet.py 各分支、stall_repair.py、projection_repair.py、decision_scope.py、unsettled_host_turn.py、heartbeat_receipt.py、scheduler/arbitration.py、user_gate.py::apply_scoped_user_gate_fallback_projection。

只扩展既有的可选键 value_notes(smoke 本来就校验它只能命名已注册的值),没有新 key、新 owner、新词表;把说明留在值旁边,正是为了 M1 拆分时「说明跟着值走」而不是重写。副作用为零,唯一消费者是 smoke 的键校验与新增测试。generate_semantic_bindings.py 的 render_glossary/render_binding 都不读 value_notes,所以本 PR 不需要重新生成任何产物。

具体改动

2 个文件、+54/-7:注册表 34 行(27 条新说明 + 被移动的 5 条)+ 测试 19 行。那 7 行删除是把五条既有说明按注册值顺序上移后的位移,以及 JSON 键序调整。

我独立核对:effective_action 5/32 → 32/32,lease_action 仍是 4/4,没有值被增删;并且用程序比对了 main 与 head 的 value_notes 映射,五条既有说明字符串逐字未变(作者声明成立)。smoke 在 head 与 merge-base(origin/main 897e9ae)上输出完全一致:coverage 26/26、same_runtime_forks=18/18、conflicting_values_semantic=0/0、unresolved_producer_sites=41。pytest tests/architecture/test_semantic_vocabulary_drift.py tests/architecture/test_semantic_inventory.py -q → 78 passed(正文在其较早 head 报 75)。

关键代码讲解

  • effective_action.value_notes(注册表 388 行起):阶梯说明与 decision_summary.py:377-411 的分支顺序逐条对得上——normal_run → outcome_floor_recovery → agent_workspace_repair → self-repair(stall_self_repair.effective_action 或 control_plane_repair)→ capability_bridge_repair → operator_gate_notify/blocked_health/throttled_skip/blocked_wait → quota_skip 作为最后兜底;agent_workspace_repair 里「outranks self-repair」「名字结尾 _repair 所以路由 repair_required」两句话分别对应阶梯位置与 _typed_route 的 endswith("_repair")。
  • 五个 self-repair 名字的 trigger:control_plane_health_repair ← health_blocker(stall_repair.py:299-301)、control_plane_projection_repair ← waiting_without_owner_projection(stall_repair.py:336-338)、state_projection_gap_repair ← state_projection_gap(project_asset.py:292 / execution_obligation.py:119-126)、boundary_projection_repair ← required_write_scope_missing_from_goal_boundary(projection_repair.py:206)、todo_decision_scope_projection_repair ← user_gate_scope_projection_drift(decision_scope.py:237)。五条我都读到对应常量。
  • 路由后果:governed_capability_intent 是唯一进 capability_action_required 的值(_typed_route 只在该 effective_action 下返回它,且 intent 投影需匹配 goal/agent 并带 command);autonomous_replan_required 确在 REPLAN_ACTIONS;terminal_no_followup 由「frontier 终态 + 空 frontier」产生(decision_summary.py:302-321),控制器把它当 terminal_action 分区(loop_controller.py:452)。
  • 调度侧措辞我逐条核过 scheduler/arbitration.py:78-95:terminal_no_followup → TERMINAL_STOP、peer_coordination_blocked → PEER_COORDINATION_STOP(所以说明写「stops rather than holding it in a monitor wait」是准确的)、agent_monitor_only → AGENT_MONITOR_ONLY_WAIT、monitor_quiet_skip → MONITOR_WAIT、heartbeat_settled_skip → QUIET_WAIT。
  • scoped_user_gate_fallback 与 unsettled_host_turn_recovery 两条最容易写错的也对得上:前者在 user_gate.py:149-193 只在 not replan_decision_allowed 且原值是 quota_skip/monitor_quiet_skip/None 时替换,义务是 one_non_gated_fallback_segment_after_user_gate_notice;后者在 unsettled_host_turn.py:275-295 先 pop 掉 selected_todo/todo_id/action_portfolio,再设 should_run=True 且 normal/recovery/self-repair 全部 False。
  • 新测试 test_remaining_kernel_values_each_carry_a_note:2 个用例通过;我用它自己的谓词在内存副本里删一条、把另一条改成纯空白,报出 ['agent_monitor_only','agent_workspace_repair'],与正文的 mutation check 一致。

对主干的风险

最强回归不是崩溃,而是「一条读起来权威的说明其实是错的/不全」,而 ratchet 只要求字符串非空。我把有实质断言的那些说明都回到产生侧核对(阶梯、五个 trigger、能力路由、调度处置、scoped user-gate、unsettled host turn),没有发现与代码不符的条目;这一点我按「读产生者」而不是「跑通测试」来验证,因为测试只能证明说明存在。

一条 P3(F2,非阻塞):kernel 覆盖规则现在有两份手写清单——本 PR 的 ['effective_action','lease_action'] 与 #4625 的四个 Turn 词表,合起来正好是注册表 tier == 'kernel' 的六个(我实测 kernel 共 6 个)。规则本可以只由注册表的分类表达;将来真出现第七个 kernel 词表,注册表会加、两个测试都不会覆盖——而这正是该 ratchet 想防的漂移。建议把两个 parametrize 合成一个按 tier == 'kernel' 派生的列表,或让它们各自断言等于该分类。

语义与 CI 对齐

本 PR 复用既有语义面:只扩展注册表既有的 value_notes 键,未新增词表/owner/键,唯一收紧的规则是「这两个 kernel 词表的值必须有说明」,且由 CI 内的架构测试执行。生成产物不受影响(value_notes 不进入 binding/glossary 渲染)。M1 拆分时说明应随值迁移,本 PR 的写法(记录产生条件而非名字改写)正是为那一刻准备的。

我的整体评价

结论 APPROVE。这是把注册表里最过载槽位从「合法但无解」补成「说明产生条件」的一次文档增量:32/32 有说明、五条既有说明逐字保留、无行为/计数/产物变化,并且把 ratchet 扩到这两个 kernel 词表,使新增值必须带说明。我按产生侧代码逐条核对了有实质断言的说明,并独立复现了文档化计数、说明保留、smoke 计数面(与 merge-base 逐项一致)与 78 个架构测试。回退成本是一个 commit。

唯一 P3 是与 #4625 重复的 kernel 覆盖清单,属维护性问题,不影响本次判断。

English verdict: APPROVE - exact head d1ea66a; I verified the registry delta directly (effective_action 5/32 -> 32/32, lease_action 4/4, no value added or removed, the five pre-existing notes byte-identical), traced the substantive note claims back to their producers (the quota_effective_action ladder, the five stall/projection repair triggers, capability routing, the terminal/frontier/scheduler dispositions, scoped user-gate fallback and unsettled host turn recovery), reproduced the test's mutation check, and confirmed the drift smoke counters match the merge-base (origin/main 897e9ae) exactly with 78 architecture tests passing. One non-blocking P3: the kernel note-coverage rule now lives in two hand-written lists (this PR plus #4625) that together equal the registry's tier == 'kernel' set, so it should be derived from that classification instead.

@huangruiteng
huangruiteng merged commit 9ccbf80 into loopx-project:main Sep 17, 2026
22 checks passed
songoow added a commit to songoow/loopx that referenced this pull request Sep 17, 2026
Six tracker PRs landed while this branch was open. Three touched files it
also edits, in the way loopx-project#4447's merge-order note predicted:

- loopx-project#4626 and loopx-project#4625 append to the end of `test_semantic_vocabulary_drift.py`;
  both blocks are kept, theirs first.
- loopx-project#4627 replaced the Section 11 target table with a *Measured by* column and
  a rule that the table carries no dated values, since those belong to the
  tracker. This branch's row had added dated numbers, so the resolution takes
  loopx-project#4627's table and puts the migration surface in *Measured by* as the
  `--report` line that prints it. The dated table stays in Appendix A.
- loopx-project#4628 memoized `python_facts`. The retirement scan needs the tree rather
  than the facts, so `parse_python` is factored out for one error path and
  left uncached: caching the trees held about two million AST nodes for the
  rest of the run and measured 0.7s worse overall, while slowing
  `check_inventory` from 2.4s to 5.4s -- the pass loopx-project#4628 had just made cheaper.

Remeasured on the integrated tree: every role count is unchanged, and
`dynamic_mapping_key_sites` moved 1704 to 1712 with the new code. Both
mirrors carry the new number.

`loopx/semantics/field_use.py` also had to stop spelling the six field names
in its own docstrings. The scan reads tracked sources under `loopx/`, this
module is one of them, and committing it pushed `heartbeat_recommendation`
to 18 of a budget of 17 -- the check catching its own module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
songoow added a commit to songoow/loopx that referenced this pull request Sep 17, 2026
Only conflict is the end of test_semantic_vocabulary_drift.py, where loopx-project#4625
and loopx-project#4626 append their value-notes coverage blocks and this branch appends
the invariant-domain fixtures. loopx-project#4447's merge-order table calls this cluster
out: keep every block, there is no overlap. Both are kept, theirs first.

Revalidated on the integrated tree: drift smoke ok, docs governance ok,
99 drift tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
songoow added a commit to songoow/loopx that referenced this pull request Sep 17, 2026
Same append-at-end cluster as loopx-project#4631: loopx-project#4625 and loopx-project#4626 landed their blocks at
the end of test_semantic_vocabulary_drift.py while this branch appends the
ratchet-lock fixtures. Both blocks kept, theirs first.

The locked values still hold on the integrated tree: conflicting_values
16/16 and conflicting_definitions 55/55, so no anchor moves in this merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
huangruiteng pushed a commit that referenced this pull request Sep 18, 2026
The kernel tier gained per-value meaning in #4625 and #4626; the 20
cross_runtime vocabularies did not, so 81 of 149 registered values were
bare tokens whose meaning a reader had to recover from the generated
rule table. Per-value coverage moves from 68/149 to 149/149.

A note says which condition produces the value: what has to be true at
runtime for the code to choose it. Not a restatement of the identifier,
and not only the disposition that follows -- that was the failure mode of
the three sets of notes M0 started with.

Two values could not be established and say so rather than guess.
settlement_failure_kind.cancelled is declared in both owners and admitted
by the decoders but selected by no branch under loopx/, exercised only by
tests that fabricate it, with no compatibility_only declaration marking it
reserved. todo_decision_scope_kind.other is an accepted member with no
producer and no fallback -- a kind outside the set is rejected, not
coerced to it -- and no documented rule for when an author picks it. Both
name the missing evidence.

The notes also state a boundary the registry previously left implicit:
several cross_runtime values are author-declared and only
membership-validated, never selected by a branch. That is the whole of
goal_amendment_class, todo_decision_scope_kind,
todo_decision_scope_granularity, and delivery_outcome.primary_goal_outcome.
Their notes name who declares the value, the criterion, where that
criterion is normative, and that no code branch selects it.

The ratchet is a new file rather than an addition to the tail of
test_semantic_vocabulary_drift.py, where the kernel ratchet lives and
where open branches already collide. It derives its population from the
registry, so a new cross_runtime vocabulary is covered without editing the
test, and it fails a missing note, an empty note, a note carrying no words
beyond its own identifier, and an unresolved marker that does not name its
missing evidence. The unresolved count is pinned at 2.

The tracking issue called this remainder 117 values; that count predates
#4626 and included effective_action (32) and lease_action (4), both
kernel-tier and already documented. The outstanding work was 81.

Refs #4447

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow
songoow deleted the codex/effective-action-value-notes branch September 28, 2026 03:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants