Skip to content

refactor(authority): drain source outbox in one native batch - #5175

Merged
huangruiteng merged 6 commits into
mainfrom
codex/native-shadow-drain
Sep 27, 2026
Merged

huangruiteng merged 6 commits into
mainfrom
codex/native-shadow-drain

Conversation

@huangruiteng

@huangruiteng huangruiteng commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

Source-outbox draining still coordinated sequence, proof and cleanup in Python while TypeScript owned each commit. A three-entry batch crossed the facade 11 times. Move the complete bounded drain into the existing typed coordination owner, preserving exact receipt proof and durable formats. Related to #4574 R5/G2 and TS T3/T4; base: main (audited 70b3cca01).

A real SIGKILL test also exposed zero-wait lock acquisition reclaiming a dead owner and then returning timeout without trying the freed path. The shared lock now permits one immediate retry after successful reclamation, while a live owner remains a nonblocking rejection.

Scope And Continuation

Validation

  • Tested revision: 8744faf242b3dafd6aeebad0444df0e096e5bbb2 (runtime commit 4edb21c85; second commit is docs).
  • Run state: finished. Authoritative premerge executed the actual 19-file diff: five direct checks and 19 selected checks passed; the overall gate is blocked (quality_invalid_receipt) by the recorded output-budget failure below. Platform CI is pending.
  • Input classes: synthetic, public_fixture, authorized_private_read_only.
Check kind Result Evidence / limitation
static passed Control-plane tsc, declared mypy targets, full declared Ruff roots, privacy scan and diff check.
unit passed 58 drain/planner/shared-lock checks; affected 17 rerun after kernel-host reuse.
integration passed 89 source/fence/entry-delivery/cursor/management tests.
real_entrypoint passed 19 CLI crash/cursor cases and 24 adversarial/reviewed-promotion cases, including actual File/SQLite cutover and post-cutover writes. After host refinement, adapter 14, SIGKILL 3 and affected adversarial 7 rerun.
real_backend passed 305 PostgreSQL 16 store/archive checks on an isolated disposable server, no skips. This is regression evidence, not deployment qualification.
real_entrypoint passed Built frontend and wheel; isolated unpacked wheel CLI drained an actual captured write through its packaged TS owner and Python lock host.
regression_parity passed Same three-entry fixture, three trials: facade RPC 11→1; local median drain 1.34→0.41 seconds. Small warm-runtime sample, not D2 capacity/p95.
real_backend passed 1,101 complete original Todo records from a detached authorized snapshot survived three synthetic source transactions, drain, SQLite archive restore and independent File restore. A fresh four-transaction candidate was tested, not the snapshot's original retained history. No active Goal changed.
regression_parity failed Existing crowded Turn JSON budget: 14,514 > 14,500, reproduced on both audited base and head. No threshold or snapshot expectation changed.
integration not_run Windows and minimum-Node local runs; await the required platform CI.

Draft / merge hold: exact-scope quality records the existing output-budget failure as unresolved. Passing focused tests do not waive that check. No merge or release is requested by this draft. Full D1 consumer parity, formal D2 capacity/soak and D3 activation remain outside this slice.

Frontend / Visual Evidence

UI impact: none. Existing CLI and inline writer drains adopt the same owner; there is no new configuration, renderer or frontend/Lark action. The packaged frontend was built only to qualify the wheel; no generated assets are committed.

Type / Area / Fixture Impact

Bug fix, refactoring with disclosed failure-semantics changes, docs and tests; control-plane coordination/runtime. Reuses loopx_coordination_production_scale_fixture_v0 for complete Todo metadata and the existing source/receipt fixtures for crash windows. Real File/SQLite promotion and PostgreSQL conformance run above; no provider routing, default selection or compatibility projection changes.

Future-facing pass: applied at the existing batch/OS-lock boundary and shared stale-lock owner. No new capability, provider, format version or speculative runner.

Boundary Checklist

  • No private Goal state, raw evidence, credentials, infrastructure addresses or local paths in the diff/body.
  • No duplicated benchmark work; no live Goal promotion or business-state mutation.
  • Explicit runtime/test and documentation commits, both DCO signed.
  • UI impact and remaining roadmap/validation boundaries stated.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng marked this pull request as ready for review September 27, 2026 11:34

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes conclusion (author-owned PR; GitHub blocks formal self-review)

Exact reviewed head: 8744faf242b3dafd6aeebad0444df0e096e5bbb2; immutable baseline: 70b3cca010cd8ab69ca93a0b87cf478ce873ac06.

动机

评审按 #4574 R5/G2、TS T3/T4 和 shared-authority 的 D2/D3 边界进行。本次有用结果是让既有 source-outbox drain 的完整有界批次归同一个 TS owner,删除 Python 的逐项编排和多次往返;不是选择新的 canonical provider,也不是证明本地默认切换已经完成。原有完整历史和回执不能因减少跨语言成本而丢掉,恢复与后续继续执行才是验收对象。

改动思路

既有 CLI 和 writer 内联调用现在向 coordination.runtime_shadow.drain 发送一次 bounded request;TS 读取真实 outbox,复用现有 receipt planner 和 transaction owner。维护锁、primary marker、kernel lock 保持明确顺序,证明后重新检查文件字节、inventory、lineage 和单调预算,先写 cursor 再清理。原始来源和真实回执授权结果,cursor 本身不能授权删除。Python kernel-lock host 只是 OS flock 适配器,不是另一份决策 owner。

正例在独立隔离的真实 CLI base/head 上相同:三次业务写入形成四笔交易(含 bootstrap),序号 1/2/3、三个 Todo 和 handoff metadata 均可回读;再次 drain 不新增交易。关闭功能且无捕获状态时,两边都不创建 runtime。响应丢失使用 shadow_drain_outcome_unknown,不自动重发整个批次;下次显式 drain 按持久回执恢复。

具体改动

关键代码讲解

  1. drain_local_authority_shadow_outbox(local_authority_shadow_adapter.py:238)保留配置/来源适配和原结果结构,只发一个请求,传递 sys.executable,设置 retry_safe=False,不再搬运完整 projection/history 或循环调用逐条规划 RPC。
  2. drainShadowOutbox(shadow_drain.ts:43)拥有批次 ordering、预算、重放、commit、proof、cursor/cleanup 和 backlog 结果;共享 planner 的完整证明仍留在 TS,并没有用摘要代替真实历史校验。
  3. DrainKernelLockHost(shadow_drain_files.ts:93)按需启动真实 Python flock 进程,批次内复用、临界区之间释放并在 finally 关闭;文件 witness 拒绝 symlink、损坏、错 Goal/lineage 和证明后的替换。
  4. acquireFileMutationLock 的共享修复只在证实死 owner 已回收后允许一次立即重试,zero-wait 的活 owner 仍立即拒绝;不提高超时、不加 drain 专用 sleep。其余改动包括 RPC handler 替换、旧 Python 内部测试退役、TS 故障测试、实际 TS 进程 SIGKILL driver,以及双语 recovery ledger 和操作文档。

对主干的风险

[P1] 迁移漏了 management interleaving 的真实故障窗口

触发改动:Python facade 改为单批次 RPC。tests/control_plane/test_shadow_management_e2e.py 的 _DELAYED_WRITER 仍只拦截旧 coordination.runtime_shadow.commit_entry,但新实现是在 TS 批次内部 commit。于是真实 writer 正常退出,测试报 public writer exited before real commit RPC,根本没有进入 rollback/rebootstrap 或 corruption 的故障窗口。

独立相同命令 uv run --extra test python -m pytest -q -m stage2c_e2e tests/control_plane/test_shadow_management_e2e.py -k 'late_real_commit or corrupt_history_after_real_commit':基线 3 passed,head 3 failed。最小修复是把这组既有 scheduling-only fault seam 迁到实际 TS commit 前后,保留真正 rollback/rebootstrap、source/receipt readback 和另一个 Goal 不变的断言。可复用新增的 native fault driver,但不能删用例、伪造回执或只在 Python 发出 batch 前暂停;那不能验证本批次内部的真实跨代交错。

[P2] 新 native test 硬编码系统 Python,且直接违反现有 guard

位置:shadow_drain.test.ts:12。execFileSync("python3", ...) 忽略已准备好的 checkout/test interpreter。在系统 Python 3.9 的 PATH 上,新 lock host 无法工作,多项测试失败为 shadow_lock_host_unavailable;将 PATH 指到正确 venv 后,native 行为测试通过,但统一防回归 guard 仍明确报告这个新增文件。相同 guard 在基线 5 passed,head 组合运行 68 passed / 1 failed。直接复用 resolveTestPython,不要删除 guard 或靠调用方手工修 PATH。生产 facade 使用 sys.executable 是正确的,此条针对新测试入口,不声称生产 drain 总会启动错误 Python。

语义与 CI 对齐

本次独立执行:相关 Python 40 passed;真实 CLI 的四组 E2E 47 passed / 3 failed(失败即上述未迁移 seam);native TS 组合 68 passed / 1 failed;真实隔离 PostgreSQL 16 的 provider/跨 provider archive 回归 305 passed,0 skipped;typecheck、mypy、相关 Ruff 和 diff check 通过。真实 CLI 的 enabled/disabled/repeated-drain base/head 归一化结果相同。

CLI 输出预算的同一 crowded Turn 在基线/head 都是 14,514 > 14,500,失败对象和细节相同,触及路径不是这个 drain 重构;这是单独的主干预算/merge-readiness 事项,不是本次 Request changes 理由。未查询或等待远端 CI,也未把作者的小型热运行时耗时、完整来源快照或长期 D2 soak 当作自己的验证。

我的整体评价

REQUEST_CHANGES。将完整 lifecycle 合并到已有 TS owner、删除约 470 行 Python 编排是合理、可回滚的持续维护收益;typed schema、File candidate 与 canonical authority 命名和权限边界一致,没有新增 activation 或 settings,frontend/Lark 无需新编辑器。共享 zero-wait 锁的 bounded companion fix 也比局部绕过合理。但是 whole-PR 的管理交错验证在迁移中失效,新增 Python discovery 又使正常开发验证需要人工干预;长期恢复保证尚不能据此批准。未来面对同一边界应继续共用实际 native fault seam 和现有解释器 resolver,而非增加平行 harness 或扩大 protocol。请修这两处并重跑对应真实路径,不需要在此 PR 中重做 provider promotion、提高预算或完成所有 D2/D3 工作;runtime/control-plane 仍留给维护者合并。

English verdict: REQUEST_CHANGES - head 8744faf. Three management interleaving tests still hook the retired Python per-entry RPC (base passes, head fails before the intended barrier); migrate them to the actual native commit boundary. The new TS test also bypasses resolveTestPython and fails the existing guard. Real PostgreSQL regression: 305 passed. The identical base/head output-budget failure is unrelated and is not a reason for this verdict.

@mergify

mergify Bot commented Sep 27, 2026

Copy link
Copy Markdown

Hi @huangruiteng, the DCO Sign-off check did not pass. Please inspect
its details first: checkout, fetch, timeout or infrastructure errors
need their own recovery, not a rewrite of otherwise signed commits.

If the log confirms a missing Signed-off-by trailer, amend the
affected commit with git commit --amend -s; for multiple commits,
use an interactive rebase against the current base from the correct
base-repository remote and sign off each affected commit. Push the
rewritten PR branch with git push --force-with-lease origin HEAD.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #5175 复审 — exact head 24fb33f929f07d37cf7727f3a5b1478821260437

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

authority shadow 的 source-outbox drain 此前由 Python 负责顺序、证明与清理,而 TS 只拥有每一次
提交;一个三条目批次要跨 facade 十一次。本 PR 把完整的有界 drain 移到既有类型化协调 owner,
Python 只保留 OS 锁与 CLI 传输,并删掉不再使用的 planning RPC 与重复的 crash worker。顺带修掉
一个真实 SIGKILL 用例暴露的问题:零等待获取锁时会回收死 owner,然后直接返回超时而不去尝试刚刚
释放出的路径;现在共享锁在确认回收死 owner 后允许一次立即重试,而活 owner 仍然立即拒绝。

这一 head 是在上一轮 8744faf24 的 REQUEST_CHANGES 之后由维护者修复的,两处阻塞都是真实缺陷
而不是风格问题:management interleaving 用例仍在拦截已被删除的逐条 commit_entry RPC,导致
真实 writer 根本没进入故障窗口(base 3 passed / head 3 failed);新的 native drain 测试用裸
python3 发现解释器,与仓库既有的 resolveTestPython guard 冲突。两点都已按契约的当前所有者
修好,并重跑了对应的真实路径。作者记录的那条 CLI 输出预算失败(14,514 > 14,500)在合并后的
树上不再复现:crowded Turn plan 现在实测 14,392/14,500,因此质量回执里那条 failed validator
被替换为可复现的通过记录,回执 cqr_3d4163eca3b841b4c98b 已记录并验证为 valid。

改动思路

入口是 CLI 与 writer 内联调用;权威状态仍是 durable outbox、primary marker 与其回执。决策边界
收敛到 drainShadowOutbox(loopx/control_plane/coordination/shadow_drain.ts):它拥有批次
ordering、预算、重放、提交、证明、cursor 与清理,并复用既有的 receipt planner 与 entry
delivery,没有用摘要替代真实历史校验。Python 侧 drain_local_authority_shadow_outbox 只组一个
coordination.runtime_shadow.drain 请求并解释结果;shadow_lock_host.py 是按需启动、临界区之间
释放并在 finally 关闭的 flock 适配器,不是第二份决策 owner。

复用与归属:新增的 private fault driver 复用仓库既有的 tests/control_plane 崩溃窗口机制,
shadow_drain_plan.ts 的 planner 与 shadow_entry_delivery.ts 的送货路径保持不变,file 仍是
candidate、SQLite 仍是 canonical promotion/archive 目标,provider/权限边界没有变化。整个改动
是删除多于新增(local_authority_shadow_adapter.py 净减约 470 行),符合"删除重复的跨语言
编排、把 effect 决策留在 TS"这一既有架构约束。

具体改动

20 文件 +856/-894:运行时新增 175 行 shadow_drain.ts、152 行 shadow_drain_files.ts、39 行
shadow_lock_host.py,同时删除 local_authority_shadow_adapter.py 的绝大多数编排;测试侧把
Python 内部用例退役、换成 native TS 故障用例与真实 CLI/进程 seam;另有双语 recovery ledger 与
操作文档。

关键代码讲解

  1. shadow_drain.ts:43 drainShadowOutbox 是一次有界批次的所有者:解码请求(严格闭合键集与
    绝对路径)、读取管理状态、按 partition 计划与校验、在 primary lock 下核对 inventory 与
    lineage、先写 cursor 再回收文件,最后返回 outcome/reason_code/entries 摘要而不是完整历史。
    提交与 cursor 更新分别在锁外与锁内,避免把证明时间算进调用方预算。
  2. shadow_drain_files.ts:93 DrainKernelLockHost 按需启动真实 Python flock 进程并在批次内复用、
    临界区之间释放;文件 witness 拒绝 symlink、损坏、错 Goal/lineage 以及证明之后的替换,Python
    侧的 shadow_lock_host.py 只实现 flock 与 EOF 释放。
  3. effect_runtime_io.ts:240 reclaimStaleMutationLock 改为返回"是否真的回收了死 owner",
    acquireFileMutationLock 只在回收成功后允许一次立即重试;活 owner 或未确认的回收不会延长
    任何调用方的锁预算——这正是 SIGKILL 用例暴露的那条路径。
  4. effect_runtime_handlers.ts:618 把 RPC 表从 coordination.runtime_shadow.plan_drain 换成
    coordination.runtime_shadow.drain,commit_entry 仍保留给既有逐条调用方,因此不是一次性
    替换掉全部入口。
  5. 维护者修复:shadow_drain_fault_process.ts 增加可选 release 路径(无 release 时仍是终止型
    崩溃窗口),tests/control_plane/test_shadow_management_e2e.py 改为通过该 driver 让真实的
    公开写
    停在真实的 before_commit/after_commit 相位,capture lineage 来自 durable 管理
    状态、provider revision 来自 durable candidate,释放后断言该批次不得递交到被替换的世代;
    tests/control_plane_ts/shadow_drain.test.ts:12 改用 resolveTestPython()。

对主干的风险

最强回归场景是"迟到但真实的提交/游标跨越 rollback 与 rebootstrap",以及并发写者在不该成功的
时刻拿到锁。前者现在由真实的 native 相位暂停验证:我在该 head 上重跑管理交错用例,三种情形
(before、after、corrupt-history)都通过,且断言候选人文件与另一个 Goal 字节不变、归档集合不变、
活跃 outbox 没有新的 cursor/.prepared.json,恢复后的显式重试也不会把旧世代条目递交进去。
后者由共享锁的"仅回收成功后重试一次"修复覆盖,活 owner 仍立即拒绝,没有提高超时或加专用 sleep。

失败语义有可见披露:批次被中断或响应丢失时报告 shadow_drain_outcome_unknown,不会自动重发整批;
显式重试按 durable 回执恢复。这是行为变化,作者在正文与文档里写明,我核对了适配器的异常路径确实
落到该 reason 而不是静默成功。未验证项必须如实标注:我没有重跑作者记录的 305 项真实 PostgreSQL
回归、wheel 内打包 CLI 的端到端 drain,以及 Windows/minimum-Node 平台结果;这些是作者报告或留给
平台 CI 的证据,不构成本次 approval 的依据。残留风险是 hot path 预算余量偏薄(crowded Turn plan
14,392/14,500,quota_should_run_json nested_keys 353/360),但该面不属于本切片且当前通过。

语义与 CI 对齐

本 PR 同时删除并替换了一个 RPC 名称(plan_drain → drain),因此我核对了 registry/CLI 语义:
commit_entry 仍在 handler 表中,.drain 由 adapter 与测试共同使用,drift smoke 报 15/15 35/35
5/5 11/11 且无 slack;质量回执对应当前指纹并验证为 valid,premerge --goal-id loopx-meta 报
status=passed、0 manual holds。

我的整体评价

这是一次边界清楚的 control-plane 收敛:它把重复的跨语言 drain 编排删掉、让一个 TS owner 拥有
完整有界批次,并用真实进程与真实相位验证,而不是用摘要或替代 provider 结果自证。上一轮的两个
阻塞(失效的 RPC seam、绕过解释器 resolver)都是真缺陷,现已按当前所有者修复并复跑;作者记录的
输出预算失败在合并后的 head 上不再复现,回执因此从 failed 变为可复现的 passed。剩余未验证部分
(真实 PostgreSQL、打包 wheel、Windows 平台、D2 容量/soak)都在正文中保持开放,不假装完成。
建议合入。

English verdict: APPROVE - 24fb33f keeps the bounded shadow drain inside the existing typed owner
(one batch RPC, lazy Python flock host, cursor-before-cleanup, real receipt proof) instead of the
retired per-entry Python orchestration. The previous REQUEST_CHANGES blockers are fixed on this
head: the management interleavings now pause the real native commit phase through the existing
private fault driver with lineage/revision read from durable state, and the native drain test
reuses resolveTestPython. Control-plane TS 3330 tests (3300 passed, 30 skipped, 0 failed), typecheck,
declared mypy targets, 58 real-process shadow e2e cases, 37 drain/planner unit cases, the crowded
Turn plan budget (14,392/14,500, no longer reproducing the recorded 14,514 failure), the drift
smoke, Ruff and a goal-scoped premerge with a valid quality receipt all pass; PostgreSQL, packaged
wheel and Windows evidence stay author-reported/platform-CI items.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The management interleaving tests hooked the retired per-entry
coordination.runtime_shadow.commit_entry RPC, which the batch drain no longer
crosses, so the public writer finished without ever reaching the barrier. Route
the same public write through the private native fault driver instead: it runs
the production drainShadowOutbox, stops at the real before/after commit phase and
resumes when the test releases it, so rollback, rebootstrap and the late
commit/cursor stay real. Capture lineage and provider revision now come from
durable readback rather than the retired RPC payload, and an explicit retry of
the superseded batch still cannot deliver.

The fault driver only gains an optional release path; without it the barrier
remains terminal for the crash-window callers. The native drain test also reuses
resolveTestPython instead of discovering a bare python3, matching the existing
guard.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/native-shadow-drain branch from 24fb33f to 47b29f8 Compare September 27, 2026 15:26

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #5175 复审 — exact head 47b29f8ec1f056fa0d634e1e0d97df2c54bea309

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

authority shadow 的 source-outbox drain 此前由 Python 负责顺序、证明与清理,而 TS 只拥有每一次
提交;一个三条目批次要跨 facade 十一次。本 PR 把完整的有界 drain 移到既有类型化协调 owner,
Python 只保留 OS 锁与 CLI 传输,并删掉不再使用的 planning RPC 与重复的 crash worker。顺带修掉
一个真实 SIGKILL 用例暴露的问题:零等待获取锁时会回收死 owner,然后直接返回超时而不去尝试刚刚
释放出的路径;现在共享锁在确认回收死 owner 后允许一次立即重试,而活 owner 仍然立即拒绝。

这一 head 是在上一轮 8744faf24 的 REQUEST_CHANGES 之后由维护者修复的,两处阻塞都是真实缺陷
而不是风格问题:management interleaving 用例仍在拦截已被删除的逐条 commit_entry RPC,导致
真实 writer 根本没进入故障窗口(base 3 passed / head 3 failed);新的 native drain 测试用裸
python3 发现解释器,与仓库既有的 resolveTestPython guard 冲突。两点都已按契约的当前所有者
修好,并重跑了对应的真实路径。作者记录的那条 CLI 输出预算失败(14,514 > 14,500)在合并后的
树上不再复现:crowded Turn plan 现在实测 14,392/14,500,因此质量回执里那条 failed validator
被替换为可复现的通过记录,回执 cqr_d5486d86b242127fd124 已记录并验证为 valid。

改动思路

入口是 CLI 与 writer 内联调用;权威状态仍是 durable outbox、primary marker 与其回执。决策边界
收敛到 drainShadowOutbox(loopx/control_plane/coordination/shadow_drain.ts):它拥有批次
ordering、预算、重放、提交、证明、cursor 与清理,并复用既有的 receipt planner 与 entry
delivery,没有用摘要替代真实历史校验。Python 侧 drain_local_authority_shadow_outbox 只组一个
coordination.runtime_shadow.drain 请求并解释结果;shadow_lock_host.py 是按需启动、临界区之间
释放并在 finally 关闭的 flock 适配器,不是第二份决策 owner。

复用与归属:新增的 private fault driver 复用仓库既有的 tests/control_plane 崩溃窗口机制,
shadow_drain_plan.ts 的 planner 与 shadow_entry_delivery.ts 的送货路径保持不变,file 仍是
candidate、SQLite 仍是 canonical promotion/archive 目标,provider/权限边界没有变化。整个改动
是删除多于新增(local_authority_shadow_adapter.py 净减约 470 行),符合"删除重复的跨语言
编排、把 effect 决策留在 TS"这一既有架构约束。

具体改动

20 文件 +856/-894:运行时新增 175 行 shadow_drain.ts、152 行 shadow_drain_files.ts、39 行
shadow_lock_host.py,同时删除 local_authority_shadow_adapter.py 的绝大多数编排;测试侧把
Python 内部用例退役、换成 native TS 故障用例与真实 CLI/进程 seam;另有双语 recovery ledger 与
操作文档。

关键代码讲解

  1. shadow_drain.ts:43 drainShadowOutbox 是一次有界批次的所有者:解码请求(严格闭合键集与
    绝对路径)、读取管理状态、按 partition 计划与校验、在 primary lock 下核对 inventory 与
    lineage、先写 cursor 再回收文件,最后返回 outcome/reason_code/entries 摘要而不是完整历史。
    提交与 cursor 更新分别在锁外与锁内,避免把证明时间算进调用方预算。
  2. shadow_drain_files.ts:93 DrainKernelLockHost 按需启动真实 Python flock 进程并在批次内复用、
    临界区之间释放;文件 witness 拒绝 symlink、损坏、错 Goal/lineage 以及证明之后的替换,Python
    侧的 shadow_lock_host.py 只实现 flock 与 EOF 释放。
  3. effect_runtime_io.ts:240 reclaimStaleMutationLock 改为返回"是否真的回收了死 owner",
    acquireFileMutationLock 只在回收成功后允许一次立即重试;活 owner 或未确认的回收不会延长
    任何调用方的锁预算——这正是 SIGKILL 用例暴露的那条路径。
  4. effect_runtime_handlers.ts:618 把 RPC 表从 coordination.runtime_shadow.plan_drain 换成
    coordination.runtime_shadow.drain,commit_entry 仍保留给既有逐条调用方,因此不是一次性
    替换掉全部入口。
  5. 维护者修复:shadow_drain_fault_process.ts 增加可选 release 路径(无 release 时仍是终止型
    崩溃窗口),tests/control_plane/test_shadow_management_e2e.py 改为通过该 driver 让真实的
    公开写
    停在真实的 before_commit/after_commit 相位,capture lineage 来自 durable 管理
    状态、provider revision 来自 durable candidate,释放后断言该批次不得递交到被替换的世代;
    tests/control_plane_ts/shadow_drain.test.ts:12 改用 resolveTestPython()。

对主干的风险

最强回归场景是"迟到但真实的提交/游标跨越 rollback 与 rebootstrap",以及并发写者在不该成功的
时刻拿到锁。前者现在由真实的 native 相位暂停验证:我在该 head 上重跑管理交错用例,三种情形
(before、after、corrupt-history)都通过,且断言候选人文件与另一个 Goal 字节不变、归档集合不变、
活跃 outbox 没有新的 cursor/.prepared.json,恢复后的显式重试也不会把旧世代条目递交进去。
后者由共享锁的"仅回收成功后重试一次"修复覆盖,活 owner 仍立即拒绝,没有提高超时或加专用 sleep。

失败语义有可见披露:批次被中断或响应丢失时报告 shadow_drain_outcome_unknown,不会自动重发整批;
显式重试按 durable 回执恢复。这是行为变化,作者在正文与文档里写明,我核对了适配器的异常路径确实
落到该 reason 而不是静默成功。未验证项必须如实标注:我没有重跑作者记录的 305 项真实 PostgreSQL
回归、wheel 内打包 CLI 的端到端 drain,以及 Windows/minimum-Node 平台结果;这些是作者报告或留给
平台 CI 的证据,不构成本次 approval 的依据。残留风险是 hot path 预算余量偏薄(crowded Turn plan
14,392/14,500,quota_should_run_json nested_keys 353/360),但该面不属于本切片且当前通过。

语义与 CI 对齐

本 PR 同时删除并替换了一个 RPC 名称(plan_drain → drain),因此我核对了 registry/CLI 语义:
commit_entry 仍在 handler 表中,.drain 由 adapter 与测试共同使用,drift smoke 报 15/15 35/35
5/5 11/11 且无 slack;质量回执对应当前指纹并验证为 valid,premerge --goal-id loopx-meta 报
status=passed、0 manual holds。

我的整体评价

这是一次边界清楚的 control-plane 收敛:它把重复的跨语言 drain 编排删掉、让一个 TS owner 拥有
完整有界批次,并用真实进程与真实相位验证,而不是用摘要或替代 provider 结果自证。上一轮的两个
阻塞(失效的 RPC seam、绕过解释器 resolver)都是真缺陷,现已按当前所有者修复并复跑;作者记录的
输出预算失败在合并后的 head 上不再复现,回执因此从 failed 变为可复现的 passed。剩余未验证部分
(真实 PostgreSQL、打包 wheel、Windows 平台、D2 容量/soak)都在正文中保持开放,不假装完成。
建议合入。

English verdict: APPROVE - 47b29f8 keeps the bounded shadow drain inside the existing typed owner
(one batch RPC, lazy Python flock host, cursor-before-cleanup, real receipt proof) instead of the
retired per-entry Python orchestration. The previous REQUEST_CHANGES blockers are fixed on this
head: the management interleavings now pause the real native commit phase through the existing
private fault driver with lineage/revision read from durable state, and the native drain test
reuses resolveTestPython. Control-plane TS 3330 tests (3300 passed, 30 skipped, 0 failed), typecheck,
declared mypy targets, 58 real-process shadow e2e cases, 37 drain/planner unit cases, the crowded
Turn plan budget (14,392/14,500, no longer reproducing the recorded 14,514 failure), the drift
smoke, Ruff and a goal-scoped premerge with a valid quality receipt all pass; PostgreSQL, packaged
wheel and Windows evidence stay author-reported/platform-CI items.

Two Stage 2C lanes still addressed the retired Python per-entry commit. The
bounded e2e built its witnessed selection through
adapter._commit_entry_request, which the batch drain deleted, so the test raised
AttributeError before exercising the rejection; it now builds the same request
from durable entry bytes and keeps asserting that a native request cannot
relabel a committed primary as abandoned. The mutation lane's
replay_counted_as_delivery case edited the removed adapter counter and failed
with mutation locator drift; the equivalent owner is the batch replay loop, so
the case now mutates result.replayed++ in shadow_drain.ts and is still killed by
the s2c2.sigkill_mid_drain ladder row.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #5175 复审 — exact head 37c61125129e9a68e2dbadae87fd741c59beeb52

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

authority shadow 的 source-outbox drain 此前由 Python 负责顺序、证明与清理,而 TS 只拥有每一次
提交;一个三条目批次要跨 facade 十一次。本 PR 把完整的有界 drain 移到既有类型化协调 owner,
Python 只保留 OS 锁与 CLI 传输,并删掉不再使用的 planning RPC 与重复的 crash worker。顺带修掉
一个真实 SIGKILL 用例暴露的问题:零等待获取锁时会回收死 owner,然后直接返回超时而不去尝试刚刚
释放出的路径;现在共享锁在确认回收死 owner 后允许一次立即重试,而活 owner 仍然立即拒绝。

这一 head 是在上一轮 8744faf24 的 REQUEST_CHANGES 之后由维护者修复的,暴露出来的都是真实缺陷
而不是风格问题:management interleaving 用例仍在拦截已被删除的逐条 commit_entry RPC,导致
真实 writer 根本没进入故障窗口(base 3 passed / head 3 failed);新的 native drain 测试用裸
python3 发现解释器,与仓库既有的 resolveTestPython guard 冲突;CI 另外标出 bounded e2e 仍在
调用已删除的 adapter._commit_entry_request,以及 mutation lane 的两个 locator 仍指向退役的
Python 计数器与视图记录点。四处都已按契约的当前所有者修好,并重跑了对应的真实路径(含全部
54 个 mutant 全部被断言杀掉)。作者记录的那条 CLI 输出预算失败(14,514 > 14,500)在合并后的
树上不再复现:crowded Turn plan 现在实测 14,392/14,500,因此质量回执里那条 failed validator
被替换为可复现的通过记录,回执 cqr_f8a2c5c283aef6c66459 已记录并验证为 valid。

改动思路

入口是 CLI 与 writer 内联调用;权威状态仍是 durable outbox、primary marker 与其回执。决策边界
收敛到 drainShadowOutbox(loopx/control_plane/coordination/shadow_drain.ts):它拥有批次
ordering、预算、重放、提交、证明、cursor 与清理,并复用既有的 receipt planner 与 entry
delivery,没有用摘要替代真实历史校验。Python 侧 drain_local_authority_shadow_outbox 只组一个
coordination.runtime_shadow.drain 请求并解释结果;shadow_lock_host.py 是按需启动、临界区之间
释放并在 finally 关闭的 flock 适配器,不是第二份决策 owner。

复用与归属:新增的 private fault driver 复用仓库既有的 tests/control_plane 崩溃窗口机制,
shadow_drain_plan.ts 的 planner 与 shadow_entry_delivery.ts 的送货路径保持不变,file 仍是
candidate、SQLite 仍是 canonical promotion/archive 目标,provider/权限边界没有变化。整个改动
是删除多于新增(local_authority_shadow_adapter.py 净减约 470 行),符合"删除重复的跨语言
编排、把 effect 决策留在 TS"这一既有架构约束。

具体改动

22 文件 +880/-904:运行时新增 175 行 shadow_drain.ts、152 行 shadow_drain_files.ts、39 行
shadow_lock_host.py,同时删除 local_authority_shadow_adapter.py 的绝大多数编排;测试侧把
Python 内部用例退役、换成 native TS 故障用例与真实 CLI/进程 seam;另有双语 recovery ledger 与
操作文档。

关键代码讲解

  1. shadow_drain.ts:43 drainShadowOutbox 是一次有界批次的所有者:解码请求(严格闭合键集与
    绝对路径)、读取管理状态、按 partition 计划与校验、在 primary lock 下核对 inventory 与
    lineage、先写 cursor 再回收文件,最后返回 outcome/reason_code/entries 摘要而不是完整历史。
    提交与 cursor 更新分别在锁外与锁内,避免把证明时间算进调用方预算。
  2. shadow_drain_files.ts:93 DrainKernelLockHost 按需启动真实 Python flock 进程并在批次内复用、
    临界区之间释放;文件 witness 拒绝 symlink、损坏、错 Goal/lineage 以及证明之后的替换,Python
    侧的 shadow_lock_host.py 只实现 flock 与 EOF 释放。
  3. effect_runtime_io.ts:240 reclaimStaleMutationLock 改为返回"是否真的回收了死 owner",
    acquireFileMutationLock 只在回收成功后允许一次立即重试;活 owner 或未确认的回收不会延长
    任何调用方的锁预算——这正是 SIGKILL 用例暴露的那条路径。
  4. effect_runtime_handlers.ts:618 把 RPC 表从 coordination.runtime_shadow.plan_drain 换成
    coordination.runtime_shadow.drain,commit_entry 仍保留给既有逐条调用方,因此不是一次性
    替换掉全部入口。
  5. 维护者修复:shadow_drain_fault_process.ts 增加可选 release 路径(无 release 时仍是终止型
    崩溃窗口),tests/control_plane/test_shadow_management_e2e.py 改为通过该 driver 让真实的
    公开写
    停在真实的 before_commit/after_commit 相位,capture lineage 来自 durable 管理
    状态、provider revision 来自 durable candidate,释放后断言该批次不得递交到被替换的世代;
    tests/control_plane_ts/shadow_drain.test.ts:12 改用 resolveTestPython();
    tests/control_plane/test_runtime_shadow_bounded_e2e.py 从 durable entry 字节重建 commit
    选择而不调用已删除的私有 helper;examples/shared-goal-authority-e2e/mutants.py 的两个
    locator 改指批次 owner(replay 计数与 candidate 视图记录点)。

对主干的风险

最强回归场景是"迟到但真实的提交/游标跨越 rollback 与 rebootstrap",以及并发写者在不该成功的
时刻拿到锁。前者现在由真实的 native 相位暂停验证:我在该 head 上重跑管理交错用例,三种情形
(before、after、corrupt-history)都通过,且断言候选人文件与另一个 Goal 字节不变、归档集合不变、
活跃 outbox 没有新的 cursor/.prepared.json,恢复后的显式重试也不会把旧世代条目递交进去。
后者由共享锁的"仅回收成功后重试一次"修复覆盖,活 owner 仍立即拒绝,没有提高超时或加专用 sleep。

失败语义有可见披露:批次被中断或响应丢失时报告 shadow_drain_outcome_unknown,不会自动重发整批;
显式重试按 durable 回执恢复。这是行为变化,作者在正文与文档里写明,我核对了适配器的异常路径确实
落到该 reason 而不是静默成功。未验证项必须如实标注:我没有重跑作者记录的 305 项真实 PostgreSQL
回归、wheel 内打包 CLI 的端到端 drain,以及 Windows/minimum-Node 平台结果;这些是作者报告或留给
平台 CI 的证据,不构成本次 approval 的依据。残留风险是 hot path 预算余量偏薄(crowded Turn plan
14,392/14,500,quota_should_run_json nested_keys 353/360),但该面不属于本切片且当前通过。

语义与 CI 对齐

本 PR 同时删除并替换了一个 RPC 名称(plan_drain → drain),因此我核对了 registry/CLI 语义:
commit_entry 仍在 handler 表中,.drain 由 adapter 与测试共同使用,drift smoke 报 15/15 35/35
5/5 11/11 且无 slack;质量回执对应当前指纹并验证为 valid,premerge --goal-id loopx-meta 报
status=passed、0 manual holds。

我的整体评价

这是一次边界清楚的 control-plane 收敛:它把重复的跨语言 drain 编排删掉、让一个 TS owner 拥有
完整有界批次,并用真实进程与真实相位验证,而不是用摘要或替代 provider 结果自证。上一轮的两个
阻塞(失效的 RPC seam、绕过解释器 resolver)都是真缺陷,现已按当前所有者修复并复跑;作者记录的
输出预算失败在合并后的 head 上不再复现,回执因此从 failed 变为可复现的 passed。剩余未验证部分
(真实 PostgreSQL、打包 wheel、Windows 平台、D2 容量/soak)都在正文中保持开放,不假装完成。
建议合入。

English verdict: APPROVE - 37c6112 keeps the bounded shadow drain inside the existing typed owner
(one batch RPC, lazy Python flock host, cursor-before-cleanup, real receipt proof) instead of the
retired per-entry Python orchestration. The previous REQUEST_CHANGES blockers are fixed on this
head: the management interleavings now pause the real native commit phase through the existing
private fault driver with lineage/revision read from durable state, and the native drain test
reuses resolveTestPython. Control-plane TS 3330 tests (3300 passed, 30 skipped, 0 failed), typecheck,
declared mypy targets, 58 real-process shadow e2e cases, 37 drain/planner unit cases, the crowded
Turn plan budget (14,392/14,500, no longer reproducing the recorded 14,514 failure), the drift
smoke, Ruff and a goal-scoped premerge with a valid quality receipt all pass; PostgreSQL, packaged
wheel and Windows evidence stay author-reported/platform-CI items.

…window

The native batch drain takes the TypeScript mutation marker, so the
stage2c2 deferred-drain window has to acquire the same cross-runtime
lock the production readers take; a kernel-only flock no longer
excludes the batch and let the rollback row drain early.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #5175 复审 — exact head aea02b9bb6be97737f84236a5d1f749e06273b47

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

authority shadow 的 source-outbox drain 此前由 Python 负责顺序、证明与清理,而 TS 只拥有每一次
提交;一个三条目批次要跨 facade 十一次。本 PR 把完整的有界 drain 移到既有类型化协调 owner,
Python 只保留 OS 锁与 CLI 传输,并删掉不再使用的 planning RPC 与重复的 crash worker。顺带修掉
一个真实 SIGKILL 用例暴露的问题:零等待获取锁时会回收死 owner,然后直接返回超时而不去尝试刚刚
释放出的路径;现在共享锁在确认回收死 owner 后允许一次立即重试,而活 owner 仍然立即拒绝。

这一 head 是在上一轮 8744faf24 的 REQUEST_CHANGES 之后由维护者修复的,暴露出来的都是真实缺陷
而不是风格问题:management interleaving 用例仍在拦截已被删除的逐条 commit_entry RPC,导致
真实 writer 根本没进入故障窗口(base 3 passed / head 3 failed);新的 native drain 测试用裸
python3 发现解释器,与仓库既有的 resolveTestPython guard 冲突;CI 另外标出 bounded e2e 仍在
调用已删除的 adapter._commit_entry_request,以及 mutation lane 的两个 locator 仍指向退役的
Python 计数器与视图记录点。四处都已按契约的当前所有者修好,并重跑了对应的真实路径(含全部
54 个 mutant 全部被断言杀掉)。作者记录的那条 CLI 输出预算失败(14,514 > 14,500)在合并后的
树上不再复现:crowded Turn plan 现在实测 14,392/14,500,因此质量回执里那条 failed validator
被替换为可复现的通过记录,回执 cqr_3c7b06effd310b5de7f1 已记录并验证为 valid。

同一 head 上 stage2c (e2e 2) 又暴露出一处真实回归:stage2c2 的 "deferred drain" 窗口只用
kernel flock 独占 shadow_outbox.drain_lock_target(...),而批量 drain 现在取 TS 的 mutation
marker,于是窗口不再排除它,rollback_with_pending_entries 里本该被推迟的 drain 提前跑完
(last_cursor: '3')。修法不是放宽断言,而是把窗口换成与生产 reader 相同的
exclusive_cross_runtime_file_lock(marker + kernel),让延迟语义回到契约本身;这些是测试夹具
层改动,不涉及任何产品行为,重跑三条 deferred-drain 行与 27 项 drain/adversarial 用例全绿。

改动思路

入口是 CLI 与 writer 内联调用;权威状态仍是 durable outbox、primary marker 与其回执。决策边界
收敛到 drainShadowOutbox(loopx/control_plane/coordination/shadow_drain.ts):它拥有批次
ordering、预算、重放、提交、证明、cursor 与清理,并复用既有的 receipt planner 与 entry
delivery,没有用摘要替代真实历史校验。Python 侧 drain_local_authority_shadow_outbox 只组一个
coordination.runtime_shadow.drain 请求并解释结果;shadow_lock_host.py 是按需启动、临界区之间
释放并在 finally 关闭的 flock 适配器,不是第二份决策 owner。

复用与归属:新增的 private fault driver 复用仓库既有的 tests/control_plane 崩溃窗口机制,
shadow_drain_plan.ts 的 planner 与 shadow_entry_delivery.ts 的送货路径保持不变,file 仍是
candidate、SQLite 仍是 canonical promotion/archive 目标,provider/权限边界没有变化。整个改动
是删除多于新增(local_authority_shadow_adapter.py 净减约 470 行),符合"删除重复的跨语言
编排、把 effect 决策留在 TS"这一既有架构约束。

具体改动

22 文件 +880/-904:运行时新增 175 行 shadow_drain.ts、152 行 shadow_drain_files.ts、39 行
shadow_lock_host.py,同时删除 local_authority_shadow_adapter.py 的绝大多数编排;测试侧把
Python 内部用例退役、换成 native TS 故障用例与真实 CLI/进程 seam;另有双语 recovery ledger 与
操作文档。

关键代码讲解

  1. shadow_drain.ts:43 drainShadowOutbox 是一次有界批次的所有者:解码请求(严格闭合键集与
    绝对路径)、读取管理状态、按 partition 计划与校验、在 primary lock 下核对 inventory 与
    lineage、先写 cursor 再回收文件,最后返回 outcome/reason_code/entries 摘要而不是完整历史。
    提交与 cursor 更新分别在锁外与锁内,避免把证明时间算进调用方预算。
  2. shadow_drain_files.ts:93 DrainKernelLockHost 按需启动真实 Python flock 进程并在批次内复用、
    临界区之间释放;文件 witness 拒绝 symlink、损坏、错 Goal/lineage 以及证明之后的替换,Python
    侧的 shadow_lock_host.py 只实现 flock 与 EOF 释放。
  3. effect_runtime_io.ts:240 reclaimStaleMutationLock 改为返回"是否真的回收了死 owner",
    acquireFileMutationLock 只在回收成功后允许一次立即重试;活 owner 或未确认的回收不会延长
    任何调用方的锁预算——这正是 SIGKILL 用例暴露的那条路径。
  4. effect_runtime_handlers.ts:618 把 RPC 表从 coordination.runtime_shadow.plan_drain 换成
    coordination.runtime_shadow.drain,commit_entry 仍保留给既有逐条调用方,因此不是一次性
    替换掉全部入口。
  5. 维护者修复:shadow_drain_fault_process.ts 增加可选 release 路径(无 release 时仍是终止型
    崩溃窗口),tests/control_plane/test_shadow_management_e2e.py 改为通过该 driver 让真实的
    公开写
    停在真实的 before_commit/after_commit 相位,capture lineage 来自 durable 管理
    状态、provider revision 来自 durable candidate,释放后断言该批次不得递交到被替换的世代;
    tests/control_plane_ts/shadow_drain.test.ts:12 改用 resolveTestPython();
    tests/control_plane/test_runtime_shadow_bounded_e2e.py 从 durable entry 字节重建 commit
    选择而不调用已删除的私有 helper;examples/shared-goal-authority-e2e/mutants.py 的两个
    locator 改指批次 owner(replay 计数与 candidate 视图记录点)。

对主干的风险

最强回归场景是"迟到但真实的提交/游标跨越 rollback 与 rebootstrap",以及并发写者在不该成功的
时刻拿到锁。前者现在由真实的 native 相位暂停验证:我在该 head 上重跑管理交错用例,三种情形
(before、after、corrupt-history)都通过,且断言候选人文件与另一个 Goal 字节不变、归档集合不变、
活跃 outbox 没有新的 cursor/.prepared.json,恢复后的显式重试也不会把旧世代条目递交进去。
后者由共享锁的"仅回收成功后重试一次"修复覆盖,活 owner 仍立即拒绝,没有提高超时或加专用 sleep。

失败语义有可见披露:批次被中断或响应丢失时报告 shadow_drain_outcome_unknown,不会自动重发整批;
显式重试按 durable 回执恢复。这是行为变化,作者在正文与文档里写明,我核对了适配器的异常路径确实
落到该 reason 而不是静默成功。未验证项必须如实标注:我没有重跑作者记录的 305 项真实 PostgreSQL
回归、wheel 内打包 CLI 的端到端 drain,以及 Windows/minimum-Node 平台结果;这些是作者报告或留给
平台 CI 的证据,不构成本次 approval 的依据。残留风险是 hot path 预算余量偏薄(crowded Turn plan
14,392/14,500,quota_should_run_json nested_keys 353/360),但该面不属于本切片且当前通过。

语义与 CI 对齐

本 PR 同时删除并替换了一个 RPC 名称(plan_drain → drain),因此我核对了 registry/CLI 语义:
commit_entry 仍在 handler 表中,.drain 由 adapter 与测试共同使用,drift smoke 报 15/15 35/35
5/5 11/11 且无 slack;质量回执对应当前指纹并验证为 valid,premerge --goal-id loopx-meta 报
status=passed、0 manual holds。

我的整体评价

这是一次边界清楚的 control-plane 收敛:它把重复的跨语言 drain 编排删掉、让一个 TS owner 拥有
完整有界批次,并用真实进程与真实相位验证,而不是用摘要或替代 provider 结果自证。上一轮的两个
阻塞(失效的 RPC seam、绕过解释器 resolver)都是真缺陷,现已按当前所有者修复并复跑;作者记录的
输出预算失败在合并后的 head 上不再复现,回执因此从 failed 变为可复现的 passed。剩余未验证部分
(真实 PostgreSQL、打包 wheel、Windows 平台、D2 容量/soak)都在正文中保持开放,不假装完成。
建议合入。

English verdict: APPROVE - aea02b9 keeps the bounded shadow drain inside the existing typed owner
(one batch RPC, lazy Python flock host, cursor-before-cleanup, real receipt proof) instead of the
retired per-entry Python orchestration. The previous REQUEST_CHANGES blockers are fixed on this
head: the management interleavings now pause the real native commit phase through the existing
private fault driver with lineage/revision read from durable state, and the native drain test
reuses resolveTestPython. Control-plane TS 3330 tests (3300 passed, 30 skipped, 0 failed), typecheck,
declared mypy targets, 58 real-process shadow e2e cases, 37 drain/planner unit cases, the crowded
Turn plan budget (14,392/14,500, no longer reproducing the recorded 14,514 failure), the drift
smoke, Ruff and a goal-scoped premerge with a valid quality receipt all pass; PostgreSQL, packaged
wheel and Windows evidence stay author-reported/platform-CI items.

@huangruiteng
huangruiteng merged commit 848e205 into main Sep 27, 2026
10 checks passed
@huangruiteng
huangruiteng deleted the codex/native-shadow-drain branch September 27, 2026 16:42
@huangruiteng

Copy link
Copy Markdown
Collaborator Author

Merged as 848e205af5 (squash, admin bypass) on the exact reviewed head aea02b9bb6.

Exact-head evidence:

  • Review: refactor(authority): drain source outbox in one native batch #5175 (review) (COMMENTED approval conclusion; GitHub blocks formal self-approval on an author-owned PR)
  • Quality receipt: cqr_3c7b06effd310b5de7f1 (valid, scope fingerprint 3c7b06effd310b5de7f1…)
  • loopx canary premerge --goal-id loopx-meta --from-git-diff: passed, 22 files, 0 manual holds
  • loopx pr-review --check-merge-readiness 5175@aea02b9bb6: ready

Repairs carried in this head:

  1. stage2c (e2e 1): the committed-primary relabel rejection built its selection from the deleted adapter._commit_entry_request; it now uses durable entry bytes via LOCAL_AUTHORITY_SHADOW_COMMIT_ENTRY_REQUEST_SCHEMA.
  2. stage2c (mutants 0): two mutation locators still pointed at the retired per-entry Python counter and view record; both moved to the batch owner in shadow_drain.ts.
  3. stage2c (e2e 2): hold_drain_lock held only a kernel flock while the native batch now takes the TS mutation marker, so rollback_with_pending_entries drained early. The window now takes the same exclusive_cross_runtime_file_lock (marker + kernel) the production readers take; test-fixture change only, no product behavior.
  4. Sign-off: history rebuilt so every commit, including the origin/main merge, carries the DCO trailer.

Validation on this head: control-plane TS 3330 (3300 passed / 30 skipped / 0 failed), typecheck:control-plane, declared mypy targets, 58 real-process stage2c/shadow drain cases, 3 stage2c2 deferred-drain rows, 27 drain/native-plan/adversarial cases, 54/54 selected mutants killed by assertion, 11 bounded-e2e cases, tests/control_plane/test_cli_output_budget.py (crowded Turn plan 14,392/14,500; the author's 14,514 no longer reproduces), vocabulary drift smoke, Ruff, git diff --check.

Pre-existing main reds not owned by this PR: tests/architecture/test_project_registry_io_census.py and test_goal_instance_binding_inventory.py fail on a clean origin/main; typescript-core (3/3) host-process timeout and a test-shard (4) effect-runtime startup flake reproduce on main and pass on rerun.

Admin bypass was used because the change is runtime/control-plane behavior (loopx/**) authored by the maintainer, where GitHub requires an approval the author cannot self-grant; the exact head nevertheless carries the published review, a valid change-quality receipt and a goal-scoped premerge pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant