fix(post-writeback): preserve composition failure identity with retryable receipt - #3978
Conversation
…able receipt A failed post-writeback composition projection no longer collapses its durable failure to the generic composition/unknown identity. The failure dispatch now reports each registered hook's concrete public-safe hook_id/capability_id, records one retryable receipt bound to the successful primary writeback in an append-only per-goal journal (goals/<goal_id>/post_writeback_hooks/composition-retry-receipts.jsonl), and settles that receipt after a later replay composes the projection cleanly. The primary writeback is never rolled back or repeated, and no job queue, retry scheduler, or second writeback transaction is added. Closes loopx-project#3917 Signed-off-by: now-ing <now-ing@users.noreply.github.com>
…semantics Add the two counterfactuals a strict reviewer asks for next: a broken composition journal must degrade to identity-only dispatch output (no durable_receipt_ref, no journal file, concrete hook identity kept), and a composed projection must settle the composition receipt even when a hook-level producer fails, because the receipt tracks projection composition while hook failures keep their own per-hook trail. Document that division on the settle call. Signed-off-by: now-ing <now-ing@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Request changes
评审 exact head:774df713a17f0e1ab8ab54d95bd8ed4aabebf002。结论:Request changes;不是合并许可。
动机
primary writeback 已成功、但 post-writeback projection composition 失败时,旧实现只返回 composition / unknown / source_projection_failed,既丢失实际 hook/capability 身份,也没有可持久查询的“仍需重试”状态。这样后续 Agent 无法判断是同一已提交写入的暂时性 projection 失败,还是另一个失败;如果直接重做 primary mutation,又有重复副作用风险。
本 PR 的目标与 issue #3917 一致:保留具体 hook identity,为已成功的 primary writeback 写一条 retryable receipt,composition 恢复后将其 supersede 为 settled,并保证 primary effect 不回滚、不重复。这个可靠性问题与发布稳定性直接相关,值得修;最小解也确实需要 durable state,而不只是改错误文案。
改动思路
正向恢复路径是:dispatch_committed_cli_post_writeback_hooks 先解析 runtime root 和 projection;projection builder 或 dispatch transport 抛错时,按 Goal/event/Agent/Todo/turn/effect/state-version 生成稳定 pwcr_… identity,将每个注册 hook 的公开 hook_id / capability_id 写入每 Goal JSONL journal,并在返回的 failure 中附 durable_receipt_ref。相同 primary writeback 重放成功后,既有 hook sidecar 保证 provider 不重复执行;composition journal 再追加 settled row,pending readback 不再返回该 receipt。
负向路径也做了隔离:journal 写失败时仍返回具体 hook identity,只去掉无法兑现的 receipt ref;没有 pending row 时 settle 不创建文件;settled 状态终态,不会被后续 retryable append 回退。primary writeback 的真值与外部写权限不归新 journal 所有。
但“同一个 receipt”当前只绑定 primary writeback identity,没有绑定失败时的 hook set。恢复时用当前 hooks 直接结算旧 receipt;另外 readback 只扫描 JSONL 最后 512 行。因此配置变化或高 churn 时,durable pending truth 会被错误结算或直接不可见。
具体改动
loopx/control_plane/post_writeback_composition_retry.py(新增 288 行):定义 receipt schema/error codes、稳定 id/ref、per-goal journal path、append/settle 与 pending suffix readback。loopx/cli_commands/post_writeback.py(+165/-19):把 runtime-root、source projection 与 dispatch 分成三个失败阶段;composition 失败时记录 receipt 并返回真实 hook identities;成功返回后安静 settle。tests/control_plane/test_post_writeback_composition_retry.py(新增 490 行):覆盖 identity preservation、transient recovery、primary idempotency、append/settle terminal 语义、malformed rows、journal unavailable degradation 和 hook-level failure 分层。
关键代码讲解
composition_retry_receipt_id用 primary writeback identity + state version 生成幂等键。这避免同一次命令重放创建不同逻辑 receipt,但由于没有纳入原始 hooks digest,无法证明后来成功的是原先失败的同一 composition。_recorded_composition_failure把 hook identity 与 durable ref 同时投影到返回值,journal 不可用时 fail-soft;这是本 PR 最有价值的用户可诊断改进。settle_composition_retry_receipt查到同 receipt id 就追加 settled row,却不比较当前 row 的hooks和待写 settled row 的hooks。pending_composition_retry_receipts对最后 512 行做 newest-per-id 折叠;它不是“最多返回 512 个 pending”,而是先截断历史,因此可把仍未 settled 的旧 receipt 从视图中移除。
三文件 +943/-19 中,生产 +453/-19、测试 +490。按 issue 的 durable-receipt边界看整体是单一主题,未引入 scheduler 或第二次 primary transaction;不过 pending_composition_retry_receipts 和 durable_receipt_ref 目前没有任何生产 readback/resolver caller,只有测试引用,这使“后续 turn 可辨认并恢复”的 active callsite 仍未闭环。
对主干的风险
[P1] 当前 hook 集合可以错误结算原始失败 receipt
位置:loopx/control_plane/post_writeback_composition_retry.py:216-252。receipt id 不包含 hooks;settle_composition_retry_receipt 找到同 id 后,不检查原 retryable row 的 hook identities 是否与当前 composition 一致,就写入 settled。
独立反例:先用 hook-a/cap-a 记录 projection failure,再以相同 primary identity、但只注册 hook-b/cap-b 调用 settle。当前实现返回 changed=true 并把 receipt 标成 settled;hook-a 从未成功获得 projection,它的意图却永久从 pending truth 中消失。最小修复是把规范化 hook-set digest 纳入 receipt identity,或在 settle 前精确比较原 row 的 hooks,并让配置变更走明确的 supersede/cancel 规则。加入 hook removed/replaced/reordered 的恢复测试。
[P1] 512 行 suffix 会遗忘仍未结算的 durable receipt
位置:同文件 _iter_composition_retry_rows 与 pending_composition_retry_receipts。实现先读取最后 512 行,之后才按 receipt id 折叠状态。独立反例写入一条旧 retryable receipt,再写 512 条不同 receipt;旧 receipt 仍在 append-only journal 中且从未 settled,但 pending API 已不返回它。
这破坏了 issue 要求的“durable indication of whether retry is still needed”。最小修复是维护有界的 materialized pending index/compacted snapshot,或扫描/压缩到能证明每个未终结 id 的最新状态;限制应约束返回数量和存储治理,不能静默把未终结状态当不存在。加入超过边界、settled/unsolved 交错和 restart readback 测试。
[P1] retry receipt 尚无生产 readback 或 ref resolver,后续 turn 无法通过已发布控制面消费
全仓调用扫描中,pending_composition_retry_receipts 只出现在新测试;post-writeback-composition:pwcr_… ref 也只有生成器与测试,没有 status/quota/CLI resolver。当前只有触发失败的同一命令响应能看到 ref,进程结束或下一 turn 不能通过普通 LoopX read model 判断应该 replay 哪个 primary identity。
最小修复不要求新增 scheduler:把 bounded pending projection 接到现有 status/quota/repair read model,提供可解析 receipt identity 与明确 replay action;或者删掉无 caller 的 reader/ref 并把 PR 的承诺收窄为“即时错误身份 + 原命令人工重放”。需要一个真实 CLI/turn restart 测试证明 later turn 能发现并清除 pending receipt。
默认路径方面,新逻辑只在已有 post-writeback hooks 的 committed mutation 上工作;无 hooks、不发生 composition failure 时没有新增权限或外部副作用。receipt 的字段为控制面 identity,不包含 projection payload/provider output,public/private 边界合理。retryable/settled 与 error codes 是显式 schema 值,没有 substring classification;命名也未扩大 Agent authority。主要风险是持久状态真值与恢复生命周期不完整,而不是权限提升。
我的整体评价
这是正确问题、正确大方向:保留具体 hook identity、把 composition failure 与 hook-level failure 分开、用稳定 primary identity 防重复写,都是正向可靠性修复。现有 9 个 focused tests 全部通过,git diff --check 通过;远端 Python、Windows、DCO、dependency、build 与 Sonar checks 也成功。
但两个独立 durable-state 反例均失败,而且新增 pending reader/ref 没有生产消费者。当前实现可能把未执行的原 hook 错标 settled,也可能在 512 行后遗忘未结算状态,因此尚不能兑现“后续 turn 可可靠恢复”的核心承诺。修复 hook-set identity、pending retention 和最小 readback 闭环后再复审;本次不建议作为 v1.0 抢合并项,也未执行合并。
English verdict: REQUEST_CHANGES at 774df71. Concrete hook identity and idempotent composition receipts are the right reliability direction, and 9 focused tests plus remote CI pass. However, a changed hook set can settle an older failure without running the original hook, the 512-row suffix drops unresolved receipts, and the pending reader/ref has no production readback consumer. Two independent durable-state probes fail; no merge performed.
…rface pending retries Three durability gaps from the review are closed: - The receipt identity now includes a normalized, order-insensitive digest of the registered hook set, so a composition that later succeeds with a different hook set settles a different receipt (or none at all). A replaced, removed, or reordered hook can no longer mark the original failure resolved; the cross-hook-set counterexample is pinned as a test. - Pending readback scans every journal row before folding to the newest row per receipt id, so no unresolved receipt can fall out of the view past a row bound. Storage governance moved to append-time compaction under the lock, which folds repeated observations of one receipt while keeping every distinct unresolved receipt durable. - The pending view gained a production readback consumer: status attaches a bounded pending-composition-retry projection (only when something is pending, so healthy status output is unchanged) with the replay action, and a later-turn test discovers a pending receipt through that projection, replays the committed mutation, and observes it cleared. Signed-off-by: now-ing <now-ing@users.noreply.github.com>
|
Thanks for the detailed probes — all three P1s are closed at P1-1 (a changed hook set can settle the original failure's receipt). The receipt identity now includes a normalized, order-insensitive digest of the registered hook set, so the receipt a later composition settles is computed from its own hooks. Your probe — hook-a/cap-a records the failure, then the same primary identity with only hook-b registered — now settles nothing: the settle computes a different receipt id, finds no current row, and returns unchanged, while the hook-a receipt stays pending. Reordered registrations hash identically (deliberately: reordering is not a semantic change); removed/replaced hooks are different identities and are covered by dedicated regressions. P1-2 (the bounded suffix forgets unresolved receipts). P1-3 (no production readback consumer). The pending view is now consumed by Verification: 12 tests in the focused suite (the three probes plus the prior nine), 30 across the maintainability ratchet and CLI output budget gates, 74 across the status server fast-path and cli_commands suites, Ruff check and format clean, and a red/green probe on the hook-set digest (neutralizing it turns the cross-hook-set and identity tests red, restoring turns them green). |
…sion The hook-set digest keyed only (hook_id, capability_id), so replacing a hook adapter with a new policy version (the manifest-discovered adapter shape) computed the same receipt identity and could settle a failure the replacement never actually resolved. The digest and the receipt's hook rows now carry policy_version, and a version-replacement regression pins that only the original version's composition settles its own receipt. Signed-off-by: now-ing <now-ing@users.noreply.github.com>
…cement Signed-off-by: now-ing <now-ing@users.noreply.github.com>
|
Quick CI note on |
huangruiteng
left a comment
There was a problem hiding this comment.
动机
REQUEST_CHANGES,精确 head c7efbbeb25754c8e53192921b70a9086cfaa03b4。
[P1] journal compaction 原地 truncate 会永久丢失已有 pending receipts。 post_writeback_composition_retry.py 的 _compact_composition_retry_journal() 在超过 512 行后先 handle.truncate(),再逐行重写。进程退出、磁盘写入故障或这两步间的异常会破坏唯一 journal;写锁只串行化 writer,不提供 crash atomicity,reader 也不拿该锁。独立真实文件实验从 512 个 pending receipt 出发,通过生产 dispatch_committed_cli_post_writeback_hooks 的 projection-failure 分支触发第 513 次 append;仅在真实 truncate 系统操作完成后注入 OSError。最终 journal 为 0 字节,独立 pending readback 从 512 变为 0,dispatch 仍返回普通 failure 而没有持久 receipt ref。旧的未处理故障因此全部消失。最小修复:在同目录临时文件完成并验证 replacement 后原子替换,保留失败前原文件;复用现有原子写/锁边界,补 compaction 中断及 concurrent reader 回归。不要通过吞异常或再次截断来“恢复”。
改动思路
#3917 要保留失败的 hook/capability identity,并能在主写入成功后找回 composition 故障。PR 把 runtime resolution、projection 和 dispatch 分成三个 failure boundary,按 primary identity 与 hook-set/policy-version digest 记录 retryable receipt,成功 composition 后 settle;hook 自身失败仍由原 sidecar 跟踪。这个区分合理,也修复了先前不同 hook-set 误 settle 与 suffix read 丢旧 pending 的问题。
正向:projection 首次失败 → journal pending → 同 identity 的 projection 恢复 → 既有 TS hook lifecycle 去重 provider → composition settled。负向:journal 不可用 → 保留具体 hook identity 且不伪造 durable ref;但新增 compaction 会把原本可恢复的旧记录一并破坏,违反自身“所有 unresolved 都保留”的核心约束。
具体改动
四文件 +1372/-19:post_writeback.py +169/-19,status +10,新 Python journal 433 行,测试 760 行。无 UI、CLI option 或数据库 provider 改动。
_composition_failure_dispatch()返回每个注册 hook 的具体身份及可选 durable ref,区分 projection failure 与 dispatch exception;无 journal 时不声称持久化成功。composition_retry_receipt_id()纳入 goal/event/agent/todo/turn/effect/state_version 与 order-insensitive hook-set/policy-version digest;避免升级或移除 hook 后误将旧失败标为 resolved。append_composition_retry_receipt()/settle_composition_retry_receipt()在文件锁内折叠状态,settled 不回退;pending 全量读取解决旧 suffix 遗漏,却新增了危险的原地压缩。既有post_writeback_hook_transaction.ts已使用 atomicWriteJson,需优先复用 owning boundary 而非另造不具备同等耐久性的 journal lifecycle。handle_status_command()在正常 status 组装后追加 bounded pending/replay guidance;健康无 pending 时不加字段。该 read model 让遗留失败可见,但 guidance 不能充当实际 replay qualification。- 13 个新增 tests 覆盖 identity、retry/settle、changed hook-set、policy version、超过阈值与 discovery;没有覆盖压缩中断。
test_composition_replay_keeps_primary_writeback_idempotent只初始化一个从未被调用路径修改的primary_effects列表,再重入 adapter;它不能证明重新执行真正的 Todo/refresh CLI 不重复主效果。
对主干的风险
独立运行新增 13 tests 与既有 post-writeback/turn-start 45 tests,共 58 passed;Ruff 通过。真实 temporary registry/journal + production dispatcher 的 fault injection 稳定复现 512→0 pending,未访问任何活跃 Goal。未运行完整实际主写入 CLI 的 response-loss/replay、进程级并发或 crash suite,这些仍是未验证项,不能用 adapter fixture 代替。
主写入不应被 optional hook 回滚的边界仍在;然而 Python 现在拥有独立 retry/settle/compaction 状态机,而 established hook effect/receipt owner 在 TS。既有 owner 的 reuse/separation 依据需要补齐。没有新增 actor 权限,也不应把 status 的 replay 文案解释成执行授权。生命周期状态用固定枚举值而非 prose 猜测是正向的,但 durability 必须覆盖错误路径。
我的整体评价
当前完整机制仍不能兑现“失败不会再丢”的承诺。Future-facing pass 应收敛为既有 typed post-writeback owner 下的小而完整的 composition-failure seam,保留 hook-set 隔离;不要因为上一轮反馈而继续增长第二套状态机。首先修复原子 compaction,并用真实 primary CLI replay 证明恢复路径。范围/value 的判断也需从原 issue 重新评估,绿灯和新增字段本身不能成为批准理由。未合并。
English verdict: REQUEST_CHANGES at c7efbbeb25754c8e53192921b70a9086cfaa03b4. In-place journal compaction truncates the only durable copy before replacement is ready. A real-file, production-dispatch fault injection reduced 512 pending receipts to zero after an error immediately following truncate. 58 focused/adjacent tests and Ruff pass but do not cover this failure. Use atomic replacement under the existing owning boundary, and qualify actual primary-CLI replay rather than a constant-list adapter assertion.
In-place journal compaction truncated the only durable copy before the folded replacement existed, so a crash or I/O failure between truncate and rewrite lost every pending receipt (owner probe: 512 pending -> 0 after an OSError injected right after the real truncate). Materialize the folded journal in a sibling temporary file, fsync and verify it, then swap it in with one atomic replace: the failure propagates, the previous journal survives untouched, and lock-free readers only ever observe a complete old or complete folded file. Reuse the existing NamedTemporaryFile + os.replace atomic-write boundary under the append lock; add regressions for interrupted compaction and concurrent readers (red on the in-place truncation, green here). Signed-off-by: now-ing <now-ing@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
REQUEST_CHANGES,审查 exact head 731ff3574aec77b7d9d0437fe293f8381c364a28。
[P2] agent-scoped status 会投影其他 peer 的 retry 工作。 loopx/cli_commands/status.py:238-244 在既有 agent-lane compaction 之后,调用只接受 runtime_root / goal_id 的 collect_pending_composition_retry_projection;新 reader 不按 identity.agent_id 过滤。真实 status --goal-id … --agent-id A 仍返回只属于 B 的 pending receipt,附带“重放原 CLI mutation”的通用行动指导。现有 run-history / agent-management projection 按 A 过滤,新队列却越过这个边界;A 没有 pending 时仍看到 B 的待重放项,且 20 条截断可能让 A 的实际工作被其他 peer 占满。
最小修复:将可选 agent selector 传到 reader,在计数和 max_items 截断前过滤;无 agent selector 的全局视图保留全体。加真实 CLI A/B/global 三种读回,覆盖 B 排在 A 前且超过展示上限的场景。这里是工作范围和指导错投,不声称 CLI 已绕过底层写权限。
#3917 的原问题是 primary writeback 成功、projection composition 失败后身份和重试状态不可追踪。durable receipt 有实际价值,但后续消费者必须保留与已有控制面相同的 scope。
改动思路
runtime-root / projection / dispatch 三阶段失败隔离;receipt 绑定 primary identity、state_version、规范化 hook-set / policy-version。成功 composition 后 settled,hook 自身失败仍留给原 TS sidecar。新 status readback 让后续 turn 可发现 pending,不引入 scheduler 或自动执行权限。
最新提交用同目录临时文件写入、flush/fsync、验证后 os.replace,修复上一版原地 truncate 的丢失窗口。锁使用既有独立 lock 文件,替换 journal 不会替换锁本身。原始 hook-set 与 policy-version 隔离、全量折叠 pending 的修复也保留。
具体改动
四文件 +1586/-19:CLI bridge +169/-19、status +10、新 journal helper 463 行、测试 944 行。完整生产和测试均已阅读;最新 delta 为原子压缩与三项故障/读者测试,不只检查最后一个 finding。
关键代码讲解
_recorded_composition_failure:保存公开 lifecycle identity,不包含 projection/provider payload;journal 错误时保留 hook/capability 身份而不伪造 durable ref。composition_retry_receipt_id:primary identity + state_version + order-insensitive hook-set/policy-version digest,防止另一配置误结算原失败。_compact_composition_retry_journal:先 materialize replacement,再原子替换;故障测试证明原 pending 保留,live reader 测试覆盖重复压缩。它不是有界后缀裁剪,未终结 ID 不被遗忘。settle_composition_retry_receipt/ CLI bridge:相同 receipt settled 为终态;没有 pending 时不创建文件。实际 Todo complete 调用方保留原 updated_at,使同一 settlement replay 可命中 receipt。collect_pending_composition_retry_projection/handle_status_command:真正接入了生产 read model,但只按 Goal 选择,缺少既有 agent-lane selector,形成上述反例。
对主干的风险
61 项新增/邻接 Python 测试通过,Ruff、diff check 通过;远端可见检查通过,发布任务跳过。新增 compaction 故障、替换前后观察和 live reader 回归通过。未跑全仓库、真实进程终止/断电或 Windows 本地 fault injection;atomic rename 不能被描述成已验证所有崩溃模型。
另外独立执行真实 Todo complete CLI handler → 实际 Todo 文件 / settlement → TS provider/sidecar:第一次仅注入 projection builder 故障,primary 成功并记录 pending;隔一秒以相同身份重放,Todo 文件不变、updated_at 不变,projection 恢复、provider 调用一次、pending 清空。这补上原测试仅检查常量 primary_effects 列表的证据缺口。
随后在同一隔离 runtime 为 B 写一条合成 pending,独立 status CLI 指定 A,结果仍包含 B 的 receipt 和 replay_action。合成 Goal 的 status 整体返回健康告警退出码 1、无异常 error;新 pending 字段已正常渲染,错投不是 exception fallback。既有 agent-lane history/management 筛选与新 reader 的缺失参数构成直接代码证据。所有探针仅使用隔离合成状态,没有操作活动 Goal。
我的整体评价
上一轮原子压缩 blocker 已修,真实 primary replay 也得到正向验证。当前剩余阻塞应收敛为读模型 scope,不必再建一套任务调度。Python composition journal 记录 TS 调用前的失败,与 TS per-hook 执行 receipt 有不同失败边界;不要把它扩大成第二个 primary-effect owner。未来向整理建议压缩重复 journal/序列化知识、将长测试按持久化与 CLI 消费分组,但不以行数本身判定失败。
全机制与 durable recovery 的问题相称;新增行动型状态必须保持 agent 隔离,修复后再审 exact head。另请更新 PR body 的旧行数/测试数/“512 行读取”描述。未批准、未合并。
English verdict: REQUEST_CHANGES at 731ff3574aec77b7d9d0437fe293f8381c364a28. Atomic compaction fixes the previous loss window, and an independent real Todo-complete replay preserves primary state and clears the receipt. However, agent-scoped status appends unfiltered peer retry receipts and replay guidance. Filter by agent before counting/truncation; cover A/B/global CLI readback. 61 focused/adjacent Python tests, Ruff, diff checks and visible CI pass. No merge performed.
|
Thanks for the crash-atomicity probe — fixed at
Your experiment is pinned as tests: 512 pending receipts, a production dispatch driving the 513th append, and an injected Verification: focused composition-retry suite 16 passed (including the new atomicity regressions), maintainability ratchet 12, CLI output budget 21 plus the differential smoke against |
The pending composition-retry projection ignored the status command's agent selector, so one agent's --agent-id status answer surfaced other peers' pending retry receipts with generic replay guidance, and peer rows could fill the display cap before the agent's own work. The collector now filters by the receipt's identity agent before counting and truncating; a selector with no pending work stays absent from the output, and the global view keeps every agent's rows. Signed-off-by: now-ing <now-ing@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
动机
APPROVE,审查 exact head 47f65fe695653c0d5b0d86d36b0d3dc25ee8635e,全量基线 bf217e1e01bec79f357c9ecbd580cf2dfa73db8b。此前 head 731ff3574aec77b7d9d0437fe293f8381c364a28 的 agent-scoped status blocker 已修复,本轮没有发现新的阻塞问题。
#3917 的问题是 primary writeback 已提交、后续 projection composition 失败时,具体 hook 身份与后续恢复线索丢失。这个 PR 让失败可持久发现,不把 primary success 改成失败,也不另建 scheduler。单纯返回错误文本不能解决下一进程/下一轮发现问题,因此一个窄的 composition receipt 有实际价值。
改动思路
CLI bridge 隔离 runtime-root 解析、projection composition 与 hook dispatch 的失败;Python journal 只记录 lifecycle identity,真正 hook 执行和幂等 sidecar 仍由既有控制面负责。成功 composition 结算自己的 receipt,producer 的失败由原 hook trail 负责,不混淆两个失败阶段。
比较了 base/head 的 bridge、既有 TS hook sidecar、Todo complete 调用方、status 的 agent-lane compaction 与新 reader。composition 在 TS hook 调用前失败,不能假装已有 per-hook receipt 能记录未到达的执行;分离有理由,但不应扩展成第二套 primary-effect authority。现有独立文件锁被复用。
具体改动
全量四文件 +1666/-19:post-writeback bridge +169/-19、status +10、新 journal helper 472 行、测试 1015 行。不是只审最后一个过滤条件。
_recorded_composition_failure:持久记录 hook/capability 身份和 failure code;journal 也失败时保留可读身份、不伪造 durable ref。无 projection/provider 原始内容。composition_retry_receipt_id:绑定 Goal/event/Agent/Todo/Turn/effect/state_version,以及无序 hook-set/policy-version;另一配置不能误结算原失败。_compact_composition_retry_journal:全量折叠后同目录临时文件、flush/fsync、校验、原子替换;不截断未处理 ID,读者观察完整旧版或新版。settle_composition_retry_receipt:settled 是同 ID 终态;没有 pending 不创建 journal。真实 Todo complete 调用链的同身份 replay 保留 primary updated_at。collect_pending_composition_retry_projection/handle_status_command:现在传递 Agent selector,先过滤再统计和截断;global view 保留所有 Agent,总数不再错误等于展示条数。
正路径:composition 失败保存 pending,后续同身份 replay composition 成功结算,后续 status 不再显示。上一轮已用实际 Todo complete handler、文件/settlement 与 TS sidecar 做过隔离故障注入,确认 primary 文件和 updated_at 不变、provider 只调用一次;当前这些生产路径没有变化。
本轮独立运行真实 status CLI + 合成 registry/journal:25 个 B receipt 故意排在 A 的唯一 receipt 前面。A 读回总数/展示 1/1 且只有 A;B 为 25/20 且只有 B;无待办的 C 没有该字段;global 为 26/20。不是仅调用新 helper。各 CLI 因合成 Goal 健康告警返回 1,但没有 exception/error,目标字段正常渲染。旧 head 的真实 A 读取 B 反例与本轮读回构成红/绿证据。
对主干的风险
本轮 composition retry + post-writeback capability hooks 合计 58 tests passed;Ruff、diff check 通过。远端可见检查成功,发布任务按配置跳过。压缩故障、并发读者、hook-set/policy-version、journal 不可用及失败后恢复均有覆盖。
没有把测试里的常量 primary_effects 列表当作真实副作用证明;真实 primary replay 证据来自上轮独立探针。未跑全仓库、真实断电/进程终止或本地 Windows 故障注入,不能把 rename/fsync 描述成所有崩溃模型均已验证。journal 仍随独立 receipt 数增长,512 是折叠阈值而非全局存储上限。
无新 opt-in/actor 生命周期,无新执行权限;replay_action 是恢复指导,不会强制调度或授权其他 peer 写入。状态使用固定 schema/status 与身份匹配,没有 prose/substring 分类规则。健康且无 pending 时保持原 status 形状。
我的整体评价
阻塞已收敛并修好:身份隔离、未处理记录保留、原子压缩、Agent scope 都有对应证据。整个机制与 durable failure recovery 的损失风险相称,没有仅因修复旧 finding 就忽略全量成本。
未来向整理:建议把超过千行的测试按 journal 持久化与 CLI 消费分组,更新 PR body 中旧的测试数和“512 行读取”描述;不需要因此引入通用重试框架。批准此 exact head,未合并。
English verdict: APPROVE at 47f65fe695653c0d5b0d86d36b0d3dc25ee8635e. Agent-scoped status now filters before counting and truncation. Independent real CLI readbacks with 25 earlier peer receipts prove A=1/1, B=25/20, empty C, and global=26/20. 58 focused/adjacent tests, Ruff, diff checks and visible CI pass. Earlier real primary-replay evidence remains applicable to unchanged paths. Power-loss and local Windows fault testing remain unverified; no merge performed.
|
Fixed at
The regression covers exactly the shape you described: six peer receipts ahead of one own receipt — the agent-scoped view returns |
Summary
A failed post-writeback composition projection no longer collapses its durable failure to the generic
composition/unknownidentity: the failure dispatch now reports each registered hook's concrete public-safehook_id/capability_idwith adurable_receipt_ref.One retryable receipt is appended to an append-only per-goal journal (
goals/<goal_id>/post_writeback_hooks/composition-retry-receipts.jsonl), deterministically bound (pwcr_<sha256[:16]>) to the successful primary writeback identity (goal / event kind / agent / todo / turn / effect / state version), and is superseded by a settled row once a later replay composes the projection cleanly. Settled is terminal; settling without a pending receipt is a no-op that never creates the file; reads are bounded (512 rows).The failure path is also split into three precise stages (runtime-root resolution / projection builder / dispatch) so
dispatch_failedis no longer mislabeled assource_projection_failed.Issue Or Task
Closes #3917
Validation
python3 -m py_compile loopx/*.pyloopx check --scan-root .tests/control_plane/test_post_writeback_composition_retry.py→ 7 passed, one per acceptance check plus journal unit tests:test_failed_projection_preserves_actionable_hook_identity(acceptance 1)test_replay_after_transient_recovery_projects_once_and_settles_receipt(acceptance 2)test_composition_replay_keeps_primary_writeback_idempotent(acceptance 3)pytest tests/control_plane/ -q→ 2477 passed, 8 skipped (includes M6 ratchet and absolute output-budget gates)pytest tests/cli_commands/ tests/architecture/ -q→ 82 passedupstream/main→ ok (no surface growth)Type of Change
LoopX Area
loopx/command surface)apps/presentation/dashboard)Technical Direction
Boundary Checklist
.loopx/, credentials, private traces, raw sessions, or local machine paths committedprimary_writeback_preserved=True/external_writes_performed=False)Signed-off-bytrailer (git commit -s)