fix(quota): settle an autonomous replan spend from its own terminal frontier - #4536
Conversation
…rontier An autonomous replan whose accepted semantic outcome is a coverage-backed `no_followup` derives `terminal_no_followup` from its durable writeback. The strict terminal guard then rejected the quota spend of that exact settlement, which stranded the settlement and forced an unnecessary control-plane self-repair turn after the benchmark work was already complete (loopx-project#4501). The terminal guard is correct for new and unrelated work. What it must not do is reject the remaining step of the typed settlement that produced the terminal state. `build_quota_slot_preview_for_decision()` now admits a delivery-completion spend from `terminal_no_followup` only when the request is bound to an autonomous-replan settlement that already owns a committed durable-writeback receipt and does not yet own a committed spend receipt. Every other terminal spend - unbound, mismatched, or already accounted - keeps failing closed. Post-spend terminal closeout needs no new mutation step for the replan path: `receiptBoundReplayPhase` already reports an autonomous-replan settlement as `settled` only once writeback and spend receipts both exist, so ordering the spend before closeout is what makes the existing rule observable. Adds tests/control_plane/test_replan_terminal_spend_ordering.py covering the coverage-backed replan round trip and the unbound terminal rejection. Signed-off-by: YZJF,YCDG,DJLY,ZZZB <emmmmyh.inori@gmail.com> Signed-off-by: yilin-succeed <204474593+yilin-succeed@users.noreply.github.com>
d3f1c90 to
30ca03d
Compare
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这条修的是一个很具体的"半结账"窗口(即 #4501 那个排序缺口):autonomous-replan 的 settlement 在 refresh-state 里写下 writeback 之后,Goal 就变成了 terminal 前沿;但同一个 settlement 还没记录的 spend 步骤,会被严格的 terminal 守卫拒绝——于是 writeback 已提交、spend 被卡住,账本与 goal 状态互相矛盾,重试还会撞同一堵墙。我今天在本 lane 里就真实碰到过同一形态(guard 要求带 --replan-obligation-id、spend 报"matching accountable refresh-state receipt is missing"),所以这条修复的对象是活的。
改动思路
保持 terminal 守卫对"其它一切情形"依旧严格,只为本 settlement 自己开放一步:把原来内联的 spend 状态集合提升为命名常量 DELIVERY_COMPLETION_SPEND_STATES,再增加 _admits_delivery_completion_spend_state,它仅在三个条件同时成立时额外接纳 terminal_no_followup——settlement identity 的 binding_kind 是 AUTONOMOUS_REPLAN、该 settlement 的 writeback 已提交、且 spend 尚未提交。为了让这个判断可用,preview 现在还会输出 writeback_committed / spend_committed 两个派生字段。
具体改动
loopx/control_plane/quota/slot_accounting.py(+50/-2):新增_receipt_committed(:120)与_admits_delivery_completion_spend_state(:513,含DELIVERY_COMPLETION_SPEND_STATES、TERMINAL_NO_FOLLOWUP_DECISION_STATE两个常量),preview 增加两个 receipt 派生字段,原内联状态判断(:810)改为调用该谓词。tests/control_plane/test_replan_terminal_spend_ordering.py(新增 250 行,2 个用例):正向用真实 CLI 走 should-run → refresh-state → spend-slot 的排序窗口;反向test_unbound_terminal_spend_stays_rejected保证例外不被放大。
关键代码讲解
loopx/control_plane/quota/slot_accounting.py:120—_receipt_committed:把"这一步是否已有 committed 收据"从 preview 里暴露出来,且result is None or result.failure is not None一律视为未提交。这个字段是后面那个严格合取的前提,也让 preview 的读者能直接看到结账进度。loopx/control_plane/quota/slot_accounting.py:513—_admits_delivery_completion_spend_state:先看普通状态集合,只有在state == terminal_no_followup时才继续检查identity.binding_kind is SettlementBindingKind.AUTONOMOUS_REPLAN、writeback_committed is True、spend_committed is not True三者全部成立。这是本 PR 的核心:守卫的严格性没有降低,只是承认"产生这个 terminal 前沿的那笔 settlement 自己还剩最后一步没记"。loopx/control_plane/quota/slot_accounting.py:810— 调用点:delivery_completion_spend的内联集合判断被替换为谓词调用,diff 只有几行,说明修改面被压在决策点上,没有扩散到别的结账路径。tests/control_plane/test_replan_terminal_spend_ordering.py:163—test_coverage_backed_no_followup_replan_does_not_strand_its_quota_spend:这个测试值得单独说——它不是 mock 掉守卫,而是用合成 project/runtime/registry 真跑 CLI:先断言 next_cli_actions 里 refresh 与 spend 都带--replan-obligation-id且不带--todo-id,再断言 refresh 的 receipts 恰为[validation, durable_writeback],然后在这个"Goal 已 terminal 但 spend 未记录"的窗口里继续花掉这笔 spend。配合反向用例,正/负两侧都被钉住。
对主干的风险
没有阻塞项。 这个例外紧挨着"不再花配额"的终态守卫,所以我把它的收紧条件逐条核对过:绑定类型必须是 autonomous replan、writeback 必须已提交、spend 必须尚未提交——三者缺一即回落到原先的拒绝;test_unbound_terminal_spend_stays_rejected 正是这条的反向控制。另外我核对了它引用的修复模式:skills/loopx-self-repair/references/repair-patterns.md:51 的 terminal_settlement_ordering_gap 条目写的正是"保持 Goal 终态守卫严格,把最终 no_followup 当作条件化的 post-spend 终态收尾",本 PR 的实现与该模式一致,而不是自创语义。terminal_no_followup 这个状态值也与其它产出方(协议文档、opencode/pi goal-bridge)一致,不是孤字面量。
残余风险(已写入 result):我验证的是合成项目/运行时上的真实 CLI 排序,没有在真实 Goal 上跑一次收尾;例外本身很窄,未来若新增终态类型,守卫会继续保持拒绝直到被显式接纳(保守方向)。按本 lane 配置不拉取 CI。
我的整体评价
APPROVE。修得准:把"终态守卫"和"这笔 settlement 自己的最后一步"分开对待,而不是为了结账方便去放宽守卫——这正好落在仓库里那条已文档化的修复模式上。我特别认可两件事:一是 preview 先暴露 writeback_committed/spend_committed 让判断基于收据而不是猜测;二是回归测试用真 CLI 复现了排序窗口并配了反向控制,这类"半结账"缺陷最容易被单元测试漏掉。剩下的只是没在真实 Goal 上验证过这一层,属于证据缺口而非风险。
English verdict: APPROVE — exact head 30ca03d32bba1da678a681c8c95b0737e64851eb of #4536. The change keeps the terminal no-more-spend guard strict and admits exactly one extra case: a terminal_no_followup preview is allowed to settle its own spend only when the settlement identity is AUTONOMOUS_REPLAN, that settlement's durable writeback is committed, and its spend is not yet recorded (loopx/control_plane/quota/slot_accounting.py:513), with the preview now exposing writeback_committed/spend_committed (:120) and the previous inline state set promoted to DELIVERY_COMPLETION_SPEND_STATES. Validation at this head: tests/control_plane/test_replan_terminal_spend_ordering.py passes 2 tests (18.9s), the positive one driving should-run, refresh-state and spend-slot through the real CLI over the ordering window and the negative one keeping an unbound terminal spend rejected. The implementation matches the documented terminal_settlement_ordering_gap repair pattern in skills/loopx-self-repair/references/repair-patterns.md:51. No blocking findings; the residual gap is that I did not exercise a real Goal closeout.
Summary
An autonomous replan whose accepted semantic outcome is a coverage-backed
no_followupderivesterminal_no_followupfrom its durable writeback. The strict terminal guard then rejected the quota spend of that exact settlement, so the settlement was stranded and the control plane spent an extra self-repair turn after the benchmark work was already complete.This keeps the terminal guard strict and only stops it from rejecting the remaining step of the typed settlement that produced the terminal state.
build_quota_slot_preview_for_decision()admits a delivery-completion spend fromterminal_no_followuponly when the request is bound to an autonomous-replan settlement that already owns a committed durable-writeback receipt and does not yet own a committed spend receipt._resolve_preview_settlement()now surfaces committed writeback and spend receipt presence, so the admission rule is receipt-gated rather than state-only.No new terminal mutation step is added for the replan path:
receiptBoundReplayPhasealready reports an autonomous-replan settlement assettledonly once writeback and spend receipts both exist, so ordering the spend before closeout is what makes that existing rule observable instead of stranding it.Issue Or Task
Addresses #4501.
Validation
tests/control_plane/test_replan_terminal_spend_ordering.py:no_followupreplan writeback no longer strands its quota spend (receiptsvalidation -> durable_writeback -> quota_spend, one effect id, replay adds no spend);python3 -m pytest -q tests/control_plane/test_replan_terminal_spend_ordering.py— 2 passed (both fail on currentmain).python3 -m pytest -q tests/control_plane/test_quota_slot_accounting.py tests/control_plane/test_goal_terminal_no_followup.py tests/control_plane/test_replan_terminal_spend_ordering.py— 51 passed.python3 -m pytest -q tests/control_plane/test_quota_settlement_cli.py tests/control_plane/test_quota_settlement.py tests/control_plane/test_terminal_settlement_tool_behavior.py tests/control_plane/test_todo_next_action_settlement.py— 135 passed, 2 failed. Both failures aretest_prior_host_closeout_survives_hidden_todo_lifecycle[0-sqlite]/[6-sqlite]and are environmental: the local Node is 22.22.2 while the SQLite authority runtime requires the qualified 22.22.3. They fail identically on unmodifiedmain.tests/canary/test_maintainability_ratchet.pyreports the same two pre-existing unreviewed findings before and after this change (loopx/extensions/lark/goal_topic_runtime.py,should_run_prepare._prepare_quota_should_run_item);loopx/control_plane/quota/slot_accounting.pyis not flagged.ruff checkon the changed files reports only pre-existing findings on untouched import lines.Signed-off-by: yilin-succeed 204474593+yilin-succeed@users.noreply.github.com