Skip to content

fix(quota): settle an autonomous replan spend from its own terminal frontier - #4536

Merged
huangruiteng merged 1 commit into
loopx-project:mainfrom
yilin-succeed:codex/fix-replan-terminal-spend-ordering
Sep 16, 2026
Merged

huangruiteng merged 1 commit into
loopx-project:mainfrom
yilin-succeed:codex/fix-replan-terminal-spend-ordering

Conversation

@yilin-succeed

Copy link
Copy Markdown
Contributor

Summary

An autonomous replan whose accepted semantic outcome is a coverage-backed no_followup derives terminal_no_followup from its durable writeback. The strict terminal guard then rejected the quota spend of that exact settlement, so the settlement was stranded and the control plane spent an extra self-repair turn after the benchmark work was already complete.

This keeps the terminal guard strict and only stops it from rejecting the remaining step of the typed settlement that produced the terminal state.

  • build_quota_slot_preview_for_decision() admits a delivery-completion spend from terminal_no_followup only when the request is bound to an autonomous-replan settlement that already owns a committed durable-writeback receipt and does not yet own a committed spend receipt.
  • _resolve_preview_settlement() now surfaces committed writeback and spend receipt presence, so the admission rule is receipt-gated rather than state-only.
  • Unbound, mismatched, and already-accounted terminal spends keep failing closed.

No new terminal mutation step is added for the replan path: receiptBoundReplayPhase already reports an autonomous-replan settlement as settled only once writeback and spend receipts both exist, so ordering the spend before closeout is what makes that existing rule observable instead of stranding it.

Issue Or Task

Addresses #4501.

Validation

  • New tests/control_plane/test_replan_terminal_spend_ordering.py:
    • a coverage-backed no_followup replan writeback no longer strands its quota spend (receipts validation -> durable_writeback -> quota_spend, one effect id, replay adds no spend);
    • an unbound spend on the same terminal Goal is still rejected.
      python3 -m pytest -q tests/control_plane/test_replan_terminal_spend_ordering.py — 2 passed (both fail on current main).
  • Focused suites: python3 -m pytest -q tests/control_plane/test_quota_slot_accounting.py tests/control_plane/test_goal_terminal_no_followup.py tests/control_plane/test_replan_terminal_spend_ordering.py — 51 passed.
  • python3 -m pytest -q tests/control_plane/test_quota_settlement_cli.py tests/control_plane/test_quota_settlement.py tests/control_plane/test_terminal_settlement_tool_behavior.py tests/control_plane/test_todo_next_action_settlement.py — 135 passed, 2 failed. Both failures are test_prior_host_closeout_survives_hidden_todo_lifecycle[0-sqlite] / [6-sqlite] and are environmental: the local Node is 22.22.2 while the SQLite authority runtime requires the qualified 22.22.3. They fail identically on unmodified main.
  • tests/canary/test_maintainability_ratchet.py reports the same two pre-existing unreviewed findings before and after this change (loopx/extensions/lark/goal_topic_runtime.py, should_run_prepare._prepare_quota_should_run_item); loopx/control_plane/quota/slot_accounting.py is not flagged.
  • ruff check on the changed files reports only pre-existing findings on untouched import lines.

Signed-off-by: yilin-succeed 204474593+yilin-succeed@users.noreply.github.com

…rontier

An autonomous replan whose accepted semantic outcome is a coverage-backed
`no_followup` derives `terminal_no_followup` from its durable writeback. The
strict terminal guard then rejected the quota spend of that exact settlement,
which stranded the settlement and forced an unnecessary control-plane
self-repair turn after the benchmark work was already complete (loopx-project#4501).

The terminal guard is correct for new and unrelated work. What it must not do
is reject the remaining step of the typed settlement that produced the
terminal state. `build_quota_slot_preview_for_decision()` now admits a
delivery-completion spend from `terminal_no_followup` only when the request is
bound to an autonomous-replan settlement that already owns a committed
durable-writeback receipt and does not yet own a committed spend receipt.
Every other terminal spend - unbound, mismatched, or already accounted -
keeps failing closed.

Post-spend terminal closeout needs no new mutation step for the replan path:
`receiptBoundReplayPhase` already reports an autonomous-replan settlement as
`settled` only once writeback and spend receipts both exist, so ordering the
spend before closeout is what makes the existing rule observable.

Adds tests/control_plane/test_replan_terminal_spend_ordering.py covering the
coverage-backed replan round trip and the unbound terminal rejection.

Signed-off-by: YZJF,YCDG,DJLY,ZZZB <emmmmyh.inori@gmail.com>
Signed-off-by: yilin-succeed <204474593+yilin-succeed@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

这条修的是一个很具体的"半结账"窗口(即 #4501 那个排序缺口):autonomous-replan 的 settlement 在 refresh-state 里写下 writeback 之后,Goal 就变成了 terminal 前沿;但同一个 settlement 还没记录的 spend 步骤,会被严格的 terminal 守卫拒绝——于是 writeback 已提交、spend 被卡住,账本与 goal 状态互相矛盾,重试还会撞同一堵墙。我今天在本 lane 里就真实碰到过同一形态(guard 要求带 --replan-obligation-id、spend 报"matching accountable refresh-state receipt is missing"),所以这条修复的对象是活的。

改动思路

保持 terminal 守卫对"其它一切情形"依旧严格,只为本 settlement 自己开放一步:把原来内联的 spend 状态集合提升为命名常量 DELIVERY_COMPLETION_SPEND_STATES,再增加 _admits_delivery_completion_spend_state,它仅在三个条件同时成立时额外接纳 terminal_no_followup——settlement identity 的 binding_kind 是 AUTONOMOUS_REPLAN、该 settlement 的 writeback 已提交、且 spend 尚未提交。为了让这个判断可用,preview 现在还会输出 writeback_committed / spend_committed 两个派生字段。

具体改动

  • loopx/control_plane/quota/slot_accounting.py(+50/-2):新增 _receipt_committed(:120)与 _admits_delivery_completion_spend_state(:513,含 DELIVERY_COMPLETION_SPEND_STATES、TERMINAL_NO_FOLLOWUP_DECISION_STATE 两个常量),preview 增加两个 receipt 派生字段,原内联状态判断(:810)改为调用该谓词。
  • tests/control_plane/test_replan_terminal_spend_ordering.py(新增 250 行,2 个用例):正向用真实 CLI 走 should-run → refresh-state → spend-slot 的排序窗口;反向 test_unbound_terminal_spend_stays_rejected 保证例外不被放大。

关键代码讲解

  1. loopx/control_plane/quota/slot_accounting.py:120 — _receipt_committed:把"这一步是否已有 committed 收据"从 preview 里暴露出来,且 result is None or result.failure is not None 一律视为未提交。这个字段是后面那个严格合取的前提,也让 preview 的读者能直接看到结账进度。
  2. loopx/control_plane/quota/slot_accounting.py:513 — _admits_delivery_completion_spend_state:先看普通状态集合,只有在 state == terminal_no_followup 时才继续检查 identity.binding_kind is SettlementBindingKind.AUTONOMOUS_REPLAN、writeback_committed is True、spend_committed is not True 三者全部成立。这是本 PR 的核心:守卫的严格性没有降低,只是承认"产生这个 terminal 前沿的那笔 settlement 自己还剩最后一步没记"。
  3. loopx/control_plane/quota/slot_accounting.py:810 — 调用点:delivery_completion_spend 的内联集合判断被替换为谓词调用,diff 只有几行,说明修改面被压在决策点上,没有扩散到别的结账路径。
  4. tests/control_plane/test_replan_terminal_spend_ordering.py:163 — test_coverage_backed_no_followup_replan_does_not_strand_its_quota_spend:这个测试值得单独说——它不是 mock 掉守卫,而是用合成 project/runtime/registry 真跑 CLI:先断言 next_cli_actions 里 refresh 与 spend 都带 --replan-obligation-id 且不带 --todo-id,再断言 refresh 的 receipts 恰为 [validation, durable_writeback],然后在这个"Goal 已 terminal 但 spend 未记录"的窗口里继续花掉这笔 spend。配合反向用例,正/负两侧都被钉住。

对主干的风险

没有阻塞项。 这个例外紧挨着"不再花配额"的终态守卫,所以我把它的收紧条件逐条核对过:绑定类型必须是 autonomous replan、writeback 必须已提交、spend 必须尚未提交——三者缺一即回落到原先的拒绝;test_unbound_terminal_spend_stays_rejected 正是这条的反向控制。另外我核对了它引用的修复模式:skills/loopx-self-repair/references/repair-patterns.md:51 的 terminal_settlement_ordering_gap 条目写的正是"保持 Goal 终态守卫严格,把最终 no_followup 当作条件化的 post-spend 终态收尾",本 PR 的实现与该模式一致,而不是自创语义。terminal_no_followup 这个状态值也与其它产出方(协议文档、opencode/pi goal-bridge)一致,不是孤字面量。

残余风险(已写入 result):我验证的是合成项目/运行时上的真实 CLI 排序,没有在真实 Goal 上跑一次收尾;例外本身很窄,未来若新增终态类型,守卫会继续保持拒绝直到被显式接纳(保守方向)。按本 lane 配置不拉取 CI。

我的整体评价

APPROVE。修得准:把"终态守卫"和"这笔 settlement 自己的最后一步"分开对待,而不是为了结账方便去放宽守卫——这正好落在仓库里那条已文档化的修复模式上。我特别认可两件事:一是 preview 先暴露 writeback_committed/spend_committed 让判断基于收据而不是猜测;二是回归测试用真 CLI 复现了排序窗口并配了反向控制,这类"半结账"缺陷最容易被单元测试漏掉。剩下的只是没在真实 Goal 上验证过这一层,属于证据缺口而非风险。

English verdict: APPROVE — exact head 30ca03d32bba1da678a681c8c95b0737e64851eb of #4536. The change keeps the terminal no-more-spend guard strict and admits exactly one extra case: a terminal_no_followup preview is allowed to settle its own spend only when the settlement identity is AUTONOMOUS_REPLAN, that settlement's durable writeback is committed, and its spend is not yet recorded (loopx/control_plane/quota/slot_accounting.py:513), with the preview now exposing writeback_committed/spend_committed (:120) and the previous inline state set promoted to DELIVERY_COMPLETION_SPEND_STATES. Validation at this head: tests/control_plane/test_replan_terminal_spend_ordering.py passes 2 tests (18.9s), the positive one driving should-run, refresh-state and spend-slot through the real CLI over the ordering window and the negative one keeping an unbound terminal spend rejected. The implementation matches the documented terminal_settlement_ordering_gap repair pattern in skills/loopx-self-repair/references/repair-patterns.md:51. No blocking findings; the residual gap is that I did not exercise a real Goal closeout.

@huangruiteng
huangruiteng merged commit d3e3901 into loopx-project:main Sep 16, 2026
18 of 26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants