Skip to content

fix(smokes): restore three public smoke contracts broken on main - #4620

Closed
huangruiteng wants to merge 5 commits into
mainfrom
codex/main-public-smoke-repair
Closed

huangruiteng wants to merge 5 commits into
mainfrom
codex/main-public-smoke-repair

Conversation

@huangruiteng

@huangruiteng huangruiteng commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Superseded by #4648 (bounded Chat replay retention) and #4647 (runtime/credential smoke isolation). The cache behavior is now an explicit runtime change with count, row and encoded-size limits, revision invalidation and stale-snapshot protection. A small dedicated cache replaces the proposed public mutable maps and compatibility aliases; the store retains file I/O and locks. The unchanged baseline rereads 20 times; the replacement passes the 20-replay throughput contract with one read.

@huangruiteng
huangruiteng force-pushed the codex/main-public-smoke-repair branch from a45e40b to cdc0460 Compare September 17, 2026 04:00

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

详细中文评审 — PR #4620 fix(smokes): restore four public smoke contracts broken on main

评审对象为精确 head cdc04608959ec9e4213e5cfcb820389419b93098(4 个 commit,8 个文件,+223/−43)。基线为 rebase 后的 origin/main 809f0cfc3;当前 main 已推进到 3937f6b99,本 PR 处于 BEHIND,但 git diff --name-only 809f0cfc3..origin/main 与本分支改动文件零交集(新进的只有 #4618 的 baseline ceiling 刷新,以及 #4610/#4607)。因此 rebase 无冲突,本 PR 的证据也不因主干前进而失效。

动机

本 lane 的目标是让管家/managed 模式真正可用,而它的前置条件是仓库的必需检查为绿:checks / pytest / merge-gate 都依赖 public smoke 套件,main 一旦变红,每一个 open PR 都过不了 merge gate。run 35176730493(main@809f0cfc3)有 8 条 red smoke;本 PR 修掉其中 5 条,其余在正文里带 owner 说明。

这 5 条都不是 flaky,而是近期合入的改动打断了 smoke 所固化的既有契约,且都能在干净 worktree 上稳定复现:

  • loopx-chat-stream-throughput-smoke:AssertionError: SSE replay reread the event log 20 times。契约是「对已完成 Turn 重复回放 20 次,日志只读 1 次」,而 #4463 为「限制已完成事件保留」在读路径上每次都丢弃缓存,于是每次回放都整读一遍日志。
  • loopx-managed-turn-operator-flow-smoke:the channel must be able to serve the managed host,实际是 available: false / dsh_runtime_unavailable。manager_channel_binding 的 docstring 承诺「引用 governed Turn surface 自己的 executor 结论,不另造第二个」,但它调用 managed_executor_binding 时不转发 module_probe,于是调用方无法让两者一致——在没装可选 dsh runtime 的机器上它们确实不一致。
  • operator-provider-credential-smoke:the managed host must name the missing credential,实际报的是 runtime 缺失。这条 smoke 断言的是「哪个凭据给 managed host 认证」,却把 runtime 探测留给了宿主机。
  • extension-entrypoint-surface-smoke:ModuleNotFoundError: No module named 'loopx'。它直接 from loopx.extensions... 却没有同门 smoke 都在用的 repo-root bootstrap;套件用裸 python3 examples/xxx.py 调用,sys.path[0] 只有 examples/。
  • cli-help-manpage-smoke:{'unclassified': ['agent-directory']}。命令注册进了 parser,却既不在 COMMAND_GROUPS 也不在 MANPAGE_COMMAND_HELP_ONLY。

改前/改后:改前这 5 条各自退出码 1;改后全部 0,且 172 条聚焦 Python 测试通过。贡献者侧的可观察结果是:merge gate 不再被这 5 条挡住。

为什么这是一个完整切片而非碎片:5 条共享同一因类(一次合入打断一条 durable 公共 guard)和同一验证回路,可一起审查与回退;剩下 3 条(边界扫描两条属于 #4593,auto-research frontier 与 blocked-priority notify_user)需要的是 owner 的契约决定而不是修复,因此明确排除而不是硬塞进来。

改动思路

进入点:CI smoke runner 以进程方式调用各 examples/*-smoke.py;用户侧对应的是 loopx chat 的 SSE 回放、manager channel 读回(dashboard/Lark 消费)以及 loopx --help / loopx commands。

权威状态与决策归属被刻意保持单一:回放的可信度由**磁盘上的 revision(inode/size/mtime)**决定;executor 可用性结论仍由 managed_executor_binding 独占,channel 只引用;manual 可见性仍由 help_surface 独占。

关键的所有权取舍:loopx/chat_store.py 当时正好卡在 module_metric_baseline 的 1500 行 ceiling 上,任何净增都会触发 module_metric_budget finding。与其去抬 ceiling(那正是 ratchet 想阻止的动作),不如把「已结束 Turn 的常驻与淘汰」这条规则抽成它自己的 bounded context:新增 loopx/chat_event_cache.py(126 行),chat_store.py 变为 1484 行且净 −16。rows / revisions 仍是公开 mapping,store 保留 _event_cache / _event_cache_revision / _event_revision 别名,因此调用方与既有测试形状不变。

复用而非重造:module_probe 与 managed_executor_binding 已有的同名 seam 一致(默认 None 即保持今天行为);extension-entrypoint-surface-smoke.py 复用同门 smoke 的 REPO_ROOT bootstrap;typed dsh_runtime_unavailable 的覆盖仍留在 examples/loopx-turn-managed-executor-binding-smoke.py 与 tests/test_turn_managed_executor_binding.py。

具体改动

改动行分类:2 处生产行为(回放保留 + probe 透传)、1 处 CLI 分类、2 处 smoke 自洽性、1 组固化保留规则的测试。无评分、权限、持久化格式变化。

关键代码讲解

  1. loopx/chat_event_cache.py::ChatEventCache.retain_finished(第 96 行)——新增。只读最后一行判断 kind in terminal_kinds(因此绝不扫描即将跳过的前缀,tests/test_chat_event_cursor.py 正是用「不得迭代缓存前缀」守住这条),把该 Turn 记为常驻,并按插入序在超过 TERMINAL_EVENT_CACHE_TURNS(8)时淘汰最老的已完成 Turn。淘汰返回被丢弃的 key,调用方无需猜测「下一次读是为谁付的账」。
  2. loopx/chat_event_cache.py::ChatEventCache.load(第 75 行)——新增。命中条件是 resident rows 的 revision 与当前文件一致,否则通过调用方传入的 reader 读一次再发布。reader 作为参数而非模块级依赖很关键:chat_store._read_jsonl 因而仍是唯一读 seam,loopx-chat-stream-throughput-smoke.py 靠 monkeypatch 它来计读次数。
  3. loopx/chat_store.py::ChatSessionStore.events_after(第 1371 行)——把「查缓存 + 命中判定」压成 self._event_log.get(key, path),把原先「读到 terminal 就无条件 drop」换成 _retain_terminal_event_cache(key, rows);对 sequence 的 bisect_right 与 miss 时的 exclusive_file_lock 读取路径完全不变。这是修复的核心一行。
  4. loopx/chat_manager.py::manager_channel_binding(第 448 行)——新增 keyword-only、默认 None 的 module_probe: Callable[[str], bool] | None,并转发给 managed_executor_binding(第 497 行)。docstring 补上「调用方要在同一宿主机上解析这两个读回,就必须能给两者同一个 probe,否则 channel 与 planned Turn 会不一致」。默认值保证所有既有调用者行为不变。
  5. loopx/help_surface.py MANPAGE_COMMAND_HELP_ONLY(第 314 行)——加入 agent-directory,记录一次有意的 manual 可见性决定。选择 help-only 而非 COMMAND_GROUPS,是因为它与 agent-context / refresh-state / pr-review 同属「由自身 --help 文档化」这一类,且这样 man/loopx.1 与 loopx commands 输出逐字节不变(已由 cli-output-budget-regression-smoke 证明)。

另外两处 smoke 自洽性修复:operator-provider-credential-smoke.py 新增 _runtime_installed 并传入 module_probe,让「凭据缺失」成为被断言的因;extension-entrypoint-surface-smoke.py 补 5 行 repo-root bootstrap(并 ROOT = REPO_ROOT 复用)。

对主干的风险

最强回归场景:一个已完成 Turn 被多个重连并发回放,同时其他 Turn 完成。触发态是「超过 8 个已完成 Turn 被回放」或「已被缓存的 Turn 日志随后被 compact/append」。防止它的是两条不变量:淘汰按已完成 Turn 的插入序,第 9 个淘汰第 1 个;文件 revision 变化使 get 返回 None 从而重读。爆炸半径限于 chat 回放及 events_after 的读取者,缓存按 store 实例隔离,磁盘格式未变,回退即恢复旧行为(吞吐回归会回来)。

反例覆盖:若缓存无上限保留,吞吐 smoke 仍会通过,只有新增的 budget 测试会失败;若缓存「读后即丢」,cursor/perf 测试仍会通过,只有吞吐 smoke 会失败——两者互为反例,且改前都实测为红,所以 delta 可归因于本 diff。断言是「已完成 Turn 的回放只读一次」+「常驻集合被固定预算约束」,而 tests/test_chat_event_cursor.py 还独立地把「不得扫描缓存前缀」钉住——它在抽离的第一稿(用 list(rows) 物化)时真的失败了,这次回归就是这样被抓住的。

真实边界与 mock 限制:smoke 走真实 ChatSessionStore 与真实事件日志,唯一被 monkeypatch 的 _read_jsonl 只当读计数器;抽出的单元也直接测。未验证的是长生命周期的真实管家 SSE 会话跨进程重启,以及 8 个 Turn 预算的实际内存占用。

相关契约变化:channel 的 probe seam 与两处 smoke 自洽性不改共享契约;help 分类不改任何渲染输出(已由 output budget smoke 证明)。

验证矩阵(本 head 实跑):5 条 smoke 全部 exit 0;172 条聚焦 pytest 通过;loopx canary premerge --from-git-diff = 0 failures;CI 同款 ruff check tests loopx/canary loopx/control_plane loopx/domain_packs loopx/presentation = All checks passed;cli-output-budget-regression-smoke(base origin/main)ok;control-plane-maintainability-ratchet-smoke ok(本 PR 未抬 ceiling)。本地 premerge 曾出现 1 条 advisory:module_metric_budget:loopx/chat_actions.py——该文件本 PR 未触碰,且已在 main@3937f6b99 由 #4618(bf2b2832e)修好,已在 3937f6b99 上复跑确认 ok。

我的整体评价

结论:批准(APPROVE),无阻断性问题。 这个 PR 的价值在于解除仓库级阻塞:它把 5 条公共 guard 恢复到它们本来就承诺的语义,而不是放宽任何 guard。改动比例合理——真正的行为改动只有两处,各由 1 个 keyword 参数与 1 个模块承担;同时它顺势把 chat_store.py 从 ceiling 上拉回到 1484 行,属于仓库规范明确鼓励的「顺带的有界前向重构」。

保留两条非阻断意见(P2/P3,不影响批准):

  • P2:TERMINAL_EVENT_CACHE_TURNS = 8 是判断而非测量。一个 Turn 的 delta 行可能很大,因此该预算的真实内存上界未经测量;建议在对真实会话事件量做过一次测量后,再决定是调小该常量还是改为「按常驻行数」而非「按 Turn 数」设界。tests/test_chat_event_retention.py::test_completed_turn_retention_is_bounded_to_the_replay_budget 会固化最终选定的预算。
  • P3:若将来事件词汇表新增一个 terminal kind 而 TERMINAL_EVENT_KINDS 未同步,则使用该 kind 结束的 Turn 不会计入保留预算(是界变松而非答案变错)。该集合与 TERMINAL_TURN_STATES 相邻放置,新增 kind 时本就要一并改,靠 review 保证同步。

残余风险:8 个 Turn 的保留预算未实测;main 上仍有 3 条 red smoke 不在本 PR 范围内(边界扫描两条归 #4593,auto-research frontier 与 blocked-priority notify_user 属于 owner 的契约决定)。本分支处于 BEHIND(main 已到 3937f6b99),但与主干新增文件零交集,无冲突风险。

English verdict: APPROVE - exact head cdc0460; five red public smokes restored (chat stream throughput, managed turn operator flow, operator provider credential, extension entrypoint surface, cli help manpage) without weakening any guard, keeping bounded completed-Turn retention and reducing chat_store.py from its 1500-line ceiling to 1484; validated by the five smokes, 172 focused tests, canary premerge with 0 failures, the CI ruff scope, the CLI output budget smoke and the maintainability ratchet; two non-blocking P2/P3 notes recorded above.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

详细中文评审 — PR #4620 fix(smokes): restore four public smoke contracts broken on main

评审对象为精确 head cdc04608959ec9e4213e5cfcb820389419b93098(4 个 commit,8 个文件,+223/−43)。基线为 rebase 后的 origin/main 809f0cfc3;当前 main 已推进到 3937f6b99,本 PR 处于 BEHIND,但 git diff --name-only 809f0cfc3..origin/main 与本分支改动文件零交集(新进的只有 #4618 的 baseline ceiling 刷新,以及 #4610/#4607)。因此 rebase 无冲突,本 PR 的证据也不因主干前进而失效。

动机

本 lane 的目标是让管家/managed 模式真正可用,而它的前置条件是仓库的必需检查为绿:checks / pytest / merge-gate 都依赖 public smoke 套件,main 一旦变红,每一个 open PR 都过不了 merge gate。run 35176730493(main@809f0cfc3)有 8 条 red smoke;本 PR 修掉其中 5 条,其余在正文里带 owner 说明。

这 5 条都不是 flaky,而是近期合入的改动打断了 smoke 所固化的既有契约,且都能在干净 worktree 上稳定复现:

  • loopx-chat-stream-throughput-smoke:AssertionError: SSE replay reread the event log 20 times。契约是「对已完成 Turn 重复回放 20 次,日志只读 1 次」,而 #4463 为「限制已完成事件保留」在读路径上每次都丢弃缓存,于是每次回放都整读一遍日志。
  • loopx-managed-turn-operator-flow-smoke:the channel must be able to serve the managed host,实际是 available: false / dsh_runtime_unavailable。manager_channel_binding 的 docstring 承诺「引用 governed Turn surface 自己的 executor 结论,不另造第二个」,但它调用 managed_executor_binding 时不转发 module_probe,于是调用方无法让两者一致——在没装可选 dsh runtime 的机器上它们确实不一致。
  • operator-provider-credential-smoke:the managed host must name the missing credential,实际报的是 runtime 缺失。这条 smoke 断言的是「哪个凭据给 managed host 认证」,却把 runtime 探测留给了宿主机。
  • extension-entrypoint-surface-smoke:ModuleNotFoundError: No module named 'loopx'。它直接 from loopx.extensions... 却没有同门 smoke 都在用的 repo-root bootstrap;套件用裸 python3 examples/xxx.py 调用,sys.path[0] 只有 examples/。
  • cli-help-manpage-smoke:{'unclassified': ['agent-directory']}。命令注册进了 parser,却既不在 COMMAND_GROUPS 也不在 MANPAGE_COMMAND_HELP_ONLY。

改前/改后:改前这 5 条各自退出码 1;改后全部 0,且 172 条聚焦 Python 测试通过。贡献者侧的可观察结果是:merge gate 不再被这 5 条挡住。

为什么这是一个完整切片而非碎片:5 条共享同一因类(一次合入打断一条 durable 公共 guard)和同一验证回路,可一起审查与回退;剩下 3 条(边界扫描两条属于 #4593,auto-research frontier 与 blocked-priority notify_user)需要的是 owner 的契约决定而不是修复,因此明确排除而不是硬塞进来。

改动思路

进入点:CI smoke runner 以进程方式调用各 examples/*-smoke.py;用户侧对应的是 loopx chat 的 SSE 回放、manager channel 读回(dashboard/Lark 消费)以及 loopx --help / loopx commands。

权威状态与决策归属被刻意保持单一:回放的可信度由**磁盘上的 revision(inode/size/mtime)**决定;executor 可用性结论仍由 managed_executor_binding 独占,channel 只引用;manual 可见性仍由 help_surface 独占。

关键的所有权取舍:loopx/chat_store.py 当时正好卡在 module_metric_baseline 的 1500 行 ceiling 上,任何净增都会触发 module_metric_budget finding。与其去抬 ceiling(那正是 ratchet 想阻止的动作),不如把「已结束 Turn 的常驻与淘汰」这条规则抽成它自己的 bounded context:新增 loopx/chat_event_cache.py(126 行),chat_store.py 变为 1484 行且净 −16。rows / revisions 仍是公开 mapping,store 保留 _event_cache / _event_cache_revision / _event_revision 别名,因此调用方与既有测试形状不变。

复用而非重造:module_probe 与 managed_executor_binding 已有的同名 seam 一致(默认 None 即保持今天行为);extension-entrypoint-surface-smoke.py 复用同门 smoke 的 REPO_ROOT bootstrap;typed dsh_runtime_unavailable 的覆盖仍留在 examples/loopx-turn-managed-executor-binding-smoke.py 与 tests/test_turn_managed_executor_binding.py。

具体改动

改动行分类:2 处生产行为(回放保留 + probe 透传)、1 处 CLI 分类、2 处 smoke 自洽性、1 组固化保留规则的测试。无评分、权限、持久化格式变化。

关键代码讲解

  1. loopx/chat_event_cache.py::ChatEventCache.retain_finished(第 96 行)——新增。只读最后一行判断 kind in terminal_kinds(因此绝不扫描即将跳过的前缀,tests/test_chat_event_cursor.py 正是用「不得迭代缓存前缀」守住这条),把该 Turn 记为常驻,并按插入序在超过 TERMINAL_EVENT_CACHE_TURNS(8)时淘汰最老的已完成 Turn。淘汰返回被丢弃的 key,调用方无需猜测「下一次读是为谁付的账」。
  2. loopx/chat_event_cache.py::ChatEventCache.load(第 75 行)——新增。命中条件是 resident rows 的 revision 与当前文件一致,否则通过调用方传入的 reader 读一次再发布。reader 作为参数而非模块级依赖很关键:chat_store._read_jsonl 因而仍是唯一读 seam,loopx-chat-stream-throughput-smoke.py 靠 monkeypatch 它来计读次数。
  3. loopx/chat_store.py::ChatSessionStore.events_after(第 1371 行)——把「查缓存 + 命中判定」压成 self._event_log.get(key, path),把原先「读到 terminal 就无条件 drop」换成 _retain_terminal_event_cache(key, rows);对 sequence 的 bisect_right 与 miss 时的 exclusive_file_lock 读取路径完全不变。这是修复的核心一行。
  4. loopx/chat_manager.py::manager_channel_binding(第 448 行)——新增 keyword-only、默认 None 的 module_probe: Callable[[str], bool] | None,并转发给 managed_executor_binding(第 497 行)。docstring 补上「调用方要在同一宿主机上解析这两个读回,就必须能给两者同一个 probe,否则 channel 与 planned Turn 会不一致」。默认值保证所有既有调用者行为不变。
  5. loopx/help_surface.py MANPAGE_COMMAND_HELP_ONLY(第 314 行)——加入 agent-directory,记录一次有意的 manual 可见性决定。选择 help-only 而非 COMMAND_GROUPS,是因为它与 agent-context / refresh-state / pr-review 同属「由自身 --help 文档化」这一类,且这样 man/loopx.1 与 loopx commands 输出逐字节不变(已由 cli-output-budget-regression-smoke 证明)。

另外两处 smoke 自洽性修复:operator-provider-credential-smoke.py 新增 _runtime_installed 并传入 module_probe,让「凭据缺失」成为被断言的因;extension-entrypoint-surface-smoke.py 补 5 行 repo-root bootstrap(并 ROOT = REPO_ROOT 复用)。

对主干的风险

最强回归场景:一个已完成 Turn 被多个重连并发回放,同时其他 Turn 完成。触发态是「超过 8 个已完成 Turn 被回放」或「已被缓存的 Turn 日志随后被 compact/append」。防止它的是两条不变量:淘汰按已完成 Turn 的插入序,第 9 个淘汰第 1 个;文件 revision 变化使 get 返回 None 从而重读。爆炸半径限于 chat 回放及 events_after 的读取者,缓存按 store 实例隔离,磁盘格式未变,回退即恢复旧行为(吞吐回归会回来)。

反例覆盖:若缓存无上限保留,吞吐 smoke 仍会通过,只有新增的 budget 测试会失败;若缓存「读后即丢」,cursor/perf 测试仍会通过,只有吞吐 smoke 会失败——两者互为反例,且改前都实测为红,所以 delta 可归因于本 diff。断言是「已完成 Turn 的回放只读一次」+「常驻集合被固定预算约束」,而 tests/test_chat_event_cursor.py 还独立地把「不得扫描缓存前缀」钉住——它在抽离的第一稿(用 list(rows) 物化)时真的失败了,这次回归就是这样被抓住的。

真实边界与 mock 限制:smoke 走真实 ChatSessionStore 与真实事件日志,唯一被 monkeypatch 的 _read_jsonl 只当读计数器;抽出的单元也直接测。未验证的是长生命周期的真实管家 SSE 会话跨进程重启,以及 8 个 Turn 预算的实际内存占用。

相关契约变化:channel 的 probe seam 与两处 smoke 自洽性不改共享契约;help 分类不改任何渲染输出(已由 output budget smoke 证明)。

验证矩阵(本 head 实跑):5 条 smoke 全部 exit 0;172 条聚焦 pytest 通过;loopx canary premerge --from-git-diff = 0 failures;CI 同款 ruff check tests loopx/canary loopx/control_plane loopx/domain_packs loopx/presentation = All checks passed;cli-output-budget-regression-smoke(base origin/main)ok;control-plane-maintainability-ratchet-smoke ok(本 PR 未抬 ceiling)。本地 premerge 曾出现 1 条 advisory:module_metric_budget:loopx/chat_actions.py——该文件本 PR 未触碰,且已在 main@3937f6b99 由 #4618(bf2b2832e)修好,已在 3937f6b99 上复跑确认 ok。

我的整体评价

结论:批准(APPROVE),无阻断性问题。 这个 PR 的价值在于解除仓库级阻塞:它把 5 条公共 guard 恢复到它们本来就承诺的语义,而不是放宽任何 guard。改动比例合理——真正的行为改动只有两处,各由 1 个 keyword 参数与 1 个模块承担;同时它顺势把 chat_store.py 从 ceiling 上拉回到 1484 行,属于仓库规范明确鼓励的「顺带的有界前向重构」。

保留两条非阻断意见(P2/P3,不影响批准):

  • P2:TERMINAL_EVENT_CACHE_TURNS = 8 是判断而非测量。一个 Turn 的 delta 行可能很大,因此该预算的真实内存上界未经测量;建议在对真实会话事件量做过一次测量后,再决定是调小该常量还是改为「按常驻行数」而非「按 Turn 数」设界。tests/test_chat_event_retention.py::test_completed_turn_retention_is_bounded_to_the_replay_budget 会固化最终选定的预算。
  • P3:若将来事件词汇表新增一个 terminal kind 而 TERMINAL_EVENT_KINDS 未同步,则使用该 kind 结束的 Turn 不会计入保留预算(是界变松而非答案变错)。该集合与 TERMINAL_TURN_STATES 相邻放置,新增 kind 时本就要一并改,靠 review 保证同步。

残余风险:8 个 Turn 的保留预算未实测;main 上仍有 3 条 red smoke 不在本 PR 范围内(边界扫描两条归 #4593,auto-research frontier 与 blocked-priority notify_user 属于 owner 的契约决定)。本分支处于 BEHIND(main 已到 3937f6b99),但与主干新增文件零交集,无冲突风险。

English verdict: APPROVE - exact head cdc0460; five red public smokes restored (chat stream throughput, managed turn operator flow, operator provider credential, extension entrypoint surface, cli help manpage) without weakening any guard, keeping bounded completed-Turn retention and reducing chat_store.py from its 1500-line ceiling to 1484; validated by the five smokes, 172 focused tests, canary premerge with 0 failures, the CI ruff scope, the CLI output budget smoke and the maintainability ratchet; two non-blocking P2/P3 notes recorded above.

PR #4463 bounded completed-event retention by dropping a Turn's rows from the
cache as soon as its terminal event was read. That also made every replay of a
finished Turn re-read its whole event log, so an SSE reconnect storm paid one
full read per replay and `loopx-chat-stream-throughput-smoke` (which pins
`read_calls <= 1` for 20 replays) has been red on main since.

Both obligations are real: retention must stay bounded, and a finished Turn's
replay must be served from memory. `loopx/chat_event_cache.py` now owns both --
it keeps the rows a reader just read and evicts finished Turns past
`TERMINAL_EVENT_CACHE_TURNS` (8), so a repeated replay reads its log once and
memory is still bounded by a fixed number of Turns rather than by history.

The extraction also puts `chat_store.py` back under its module budget: it was
exactly at the 1500-line ceiling, and the cache is the piece that belongs in its
own bounded context. `rows`/`revisions` stay public mappings and the store keeps
`_event_cache`, `_event_cache_revision` and `_event_revision` aliases, so no
caller or test changes shape.

The retention rule is now testable without a file log: the replay contract
(5 replays, 1 read) and the budget (oldest finished Turns evicted, newest
replayable) are pinned in `tests/test_chat_event_retention.py`.

Boundary: control-plane runtime change (loopx/**), proposed as a PR with an
exact-head review and left for the maintainer - never self-merged.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…nding

`manager_channel_binding` promises that "a managed endpoint quotes the governed
Turn surface's own executor readback for that verdict instead of deriving a
second one, so the channel can never advertise an executor the Turn driver would
refuse". It does quote that surface, but it called `managed_executor_binding`
without forwarding a `module_probe`, so the two readbacks could not be made to
agree by a caller that supplies one -- and the two disagree on any host where
the optional `dsh` runtime is not installed.

That is exactly what `examples/loopx-managed-turn-operator-flow-smoke.py`
asserts, and it has been red on main since the runtime probe was introduced:
the smoke injects a probe into the executor binding (available: true) and then
reads the channel, which probed the real host (available: false,
dsh_runtime_unavailable).

Pass the probe through. It defaults to `None`, so every existing caller keeps
today's behavior and only a caller that resolves both readbacks on a host whose
optional runtime is unknown can now make them agree.

Boundary: control-plane runtime change (loopx/**), proposed as a PR with an
exact-head review and left for the maintainer - never self-merged.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…me state

`examples/operator-provider-credential-smoke.py` asserts which credential
authenticates the managed host, but it left the runtime probe to the machine it
runs on. On a host without the optional `dsh` runtime the projection answers
`dsh_runtime_unavailable` before the credential is ever consulted, so the smoke
failed on a fact it does not claim:

    operator provider credential smoke failed: the managed host must name the
    missing credential: {... 'credential_env': None, 'unavailable_reason':
    'dsh_runtime_unavailable' ...}

The smoke now states the runtime it is asserting about instead of reading the
host. The typed `dsh_runtime_unavailable` refusal keeps its coverage in
`examples/loopx-turn-managed-executor-binding-smoke.py` and
`tests/test_turn_managed_executor_binding.py`.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/main-public-smoke-repair branch from cdc0460 to c8e384a Compare September 17, 2026 04:25
@huangruiteng huangruiteng changed the title fix(smokes): restore four public smoke contracts broken on main fix(smokes): restore three public smoke contracts broken on main Sep 17, 2026

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

详细中文评审 — PR #4620 fix(smokes): restore three public smoke contracts broken on main

评审对象为精确 head c8e384a20a5f6ce54bc5033934e21026f1d102a1(3 个 commit,6 个文件,+214/−41)。基线为 origin/main a921f01c7。这一版是收窄后的分支:首版额外携带的 agent-directory 分类与 entrypoint smoke 的 import bootstrap 已移除,因为它们已分别由 #4613、#4615 开在队列里(详见上文"对主干的风险"里的重复化说明)。

动机

main 的 public smoke 套件是 checks / pytest / merge-gate 的依赖,它一旦变红,所有 open PR 都过不了 merge gate;本 lane 的管家/managed 堆叠 PR 同样被卡。本 PR 修掉其中 3 条互相独立的因:

  • loopx-chat-stream-throughput-smoke:AssertionError: SSE replay reread the event log 20 times。契约是「已完成 Turn 重复回放 20 次只读日志 1 次」,而 #4463 为「限制已完成事件保留」在读路径上每次都丢弃缓存 → 每次回放整读一遍。
  • loopx-managed-turn-operator-flow-smoke:the channel must be able to serve the managed host,实际 available: false / dsh_runtime_unavailable。manager_channel_binding 的 docstring 承诺「引用 governed Turn surface 自己的结论、不另造第二个」,但调用 managed_executor_binding 时不转发 module_probe,于是调用方无法让两者一致。
  • operator-provider-credential-smoke:断言的是「哪个凭据给 managed host 认证」,却把 runtime 探测留给了宿主机 → 缺 dsh 的机器先报 runtime 缺失。

改前/改后:改前这三条各自 exit 1;改后全部 exit 0,94 条聚焦 Python 测试通过,loopx canary premerge --from-git-diff 0 failures / 0 advisories。贡献者侧可观察结果是 merge gate 不再被这三条挡住。

为什么这是一个完整切片:三条共享同一因类(一次合入打断一条 durable 公共 guard)与同一验证回路,可一起审查与回退。其余红项各自已有 owner 或已由别的 PR 覆盖(#4593 的边界扫描组、#4622 的过期 smoke、#4613、#4615),因此不在本 PR 内重复实现——重复实现会改到同一批行并互相冲突。

改动思路

进入点:CI smoke runner 以进程方式调用 examples/*-smoke.py;用户侧对应 loopx chat 的 SSE 回放与 manager channel 读回(dashboard/Lark 消费)。

权威状态与决策归属保持单一:回放可信度由**磁盘 revision(inode/size/mtime)**决定;executor 可用性结论仍由 managed_executor_binding 独占,channel 只引用。

关键取舍:loopx/chat_store.py 正好卡在 module_metric_baseline 的 1500 行 ceiling 上,任何净增都会触发 module_metric_budget。与其抬 ceiling(那正是 ratchet 要阻止的动作),不如把「已结束 Turn 的常驻与淘汰」抽成自己的 bounded context:新增 loopx/chat_event_cache.py(126 行),chat_store.py 变为 1484 行、净 −16。rows/revisions 仍是公开 mapping,store 保留 _event_cache/_event_cache_revision/_event_revision 别名,调用方与既有测试形状不变。

复用而非重造:module_probe 与 managed_executor_binding 已有的同名 seam 一致(默认 None 即保持今天行为);typed dsh_runtime_unavailable 的覆盖仍留在 examples/loopx-turn-managed-executor-binding-smoke.py 与 tests/test_turn_managed_executor_binding.py。

具体改动

改动行分类:2 处生产行为(回放保留 + probe 透传)、1 处 smoke 自洽性、1 组固化保留规则的测试。无 CLI、评分、权限、持久化格式变化。

关键代码讲解

  1. loopx/chat_event_cache.py::ChatEventCache.retain_finished(第 96 行)——新增。只读最后一行判断 kind in terminal_kinds(绝不扫描即将跳过的前缀;tests/test_chat_event_cursor.py 用「不得迭代缓存前缀」把这条守住),把该 Turn 记为常驻,并按插入序在超过 TERMINAL_EVENT_CACHE_TURNS(8)时淘汰最老的已完成 Turn,返回被丢弃的 key。
  2. loopx/chat_event_cache.py::ChatEventCache.load(第 75 行)——新增。命中条件是 resident rows 的 revision 与当前文件一致,否则通过调用方传入的 reader 读一次再发布。reader 作为参数而非模块级依赖很关键:chat_store._read_jsonl 仍是唯一读 seam,吞吐 smoke 靠 monkeypatch 它来计读次数。
  3. loopx/chat_store.py::ChatSessionStore.events_after(第 1371 行)——把「查缓存 + 命中判定」压成 self._event_log.get(key, path),把原先「读到 terminal 就无条件 drop」换成 _retain_terminal_event_cache(key, rows);对 sequence 的 bisect_right 与 miss 时的 exclusive_file_lock 读取路径完全不变。这是修复核心一行。
  4. loopx/chat_manager.py::manager_channel_binding(第 448 行)——新增 keyword-only、默认 None 的 module_probe: Callable[[str], bool] | None,转发给 managed_executor_binding(第 497 行);docstring 补上「调用方要在同一宿主机解析这两个读回,就必须能给两者同一个 probe」。

另:examples/operator-provider-credential-smoke.py 新增 _runtime_installed 并传入 module_probe,让「凭据缺失」成为被断言的因;examples/loopx-managed-turn-operator-flow-smoke.py 新增一行把同一个 probe 交给 channel 读回。

对主干的风险

最强回归场景:一个已完成 Turn 被多个重连并发回放,同时其他 Turn 完成。触发态是「超过 8 个已完成 Turn 被回放」或「已被缓存的 Turn 日志随后被 compact/append」。防止它的是两条不变量:淘汰按已完成 Turn 的插入序;文件 revision 变化使 get 返回 None 从而重读。爆炸半径限于 chat 回放及 events_after 读取者,缓存按 store 实例隔离,磁盘格式未变,回退即恢复旧行为。

反例覆盖:若缓存无上限保留,吞吐 smoke 仍通过,只有新增 budget 测试失败;若「读后即丢」,cursor/perf 测试仍通过,只有吞吐 smoke 失败——两者互为反例,且改前都实测为红。tests/test_chat_event_cursor.py 还在抽离第一稿(用 list(rows) 物化)时真的失败过,该回归就是这样被抓住的。

真实边界与 mock 限制:smoke 走真实 ChatSessionStore 与真实事件日志,唯一 monkeypatch 的 _read_jsonl 只当读计数器;抽出的单元也直接测。未验证的是长生命周期真实管家 SSE 会话跨进程重启,以及 8 个 Turn 预算的实际内存占用。

重复化与收窄(本轮纠正):首版还改了 loopx/help_surface.py 与 examples/extension-entrypoint-surface-smoke.py,随后发现 #4613、#4615 已各自开在同一处(#4613 走 manual 可见性并改 man/loopx.1,与本分支原先的 help-only 决定不同,两者并存会互相回退)。因此我把这两个修复从本分支移除并 force-push,避免同批行冲突;本 PR 现在只保留它独有的三处修复。

验证矩阵(本 head 实跑):三条 smoke 全部 exit 0(PYTHONPATH 未设置以对齐 CI 环境);94 条聚焦 pytest 通过;loopx canary premerge --from-git-diff 0 failures / 0 advisories(4 个 catalog canary);control-plane-maintainability-ratchet-smoke ok 且未抬 ceiling。

我的整体评价

结论:批准(APPROVE),无阻断性问题。 价值在于解除仓库级阻塞:把 3 条公共 guard 恢复到它们本来就承诺的语义,而不是放宽任何 guard。比例合理——真正的行为改动只有两处,各由一个 keyword 参数与一个新模块承担;同时顺势把 chat_store.py 从 ceiling 上拉回到 1484 行。收窄重复修复后,这个 PR 也是可独立审查与回退的最小单元。

保留两条非阻断意见(P2/P3):TERMINAL_EVENT_CACHE_TURNS = 8 是判断而非测量,建议在对真实会话事件量测量后再决定调小或改为「按常驻行数」设界;若将来事件词汇表新增 terminal kind 而 TERMINAL_EVENT_KINDS 未同步,则该 kind 结束的 Turn 不会计入预算(是界变松而非答案变错),该集合与 TERMINAL_TURN_STATES 相邻放置靠 review 保证同步。

残余风险:8 个 Turn 的保留预算未实测;main 其余红项不在本 PR 内——边界扫描组(repository-hygiene、canary/canary-promotion-readiness、auto-research-rollout-readpath 同一根因,归 #4593)、过期 smoke(#4622)、agent-directory(#4613)、entrypoint bootstrap(#4615)。

English verdict: APPROVE - exact head c8e384a; three red public smokes restored (chat stream throughput, managed turn operator flow, operator provider credential) without weakening any guard, keeping bounded completed-Turn retention and reducing chat_store.py from its 1500-line ceiling to 1484; two duplicated repairs dropped from this branch once #4613/#4615 were found open, and the earlier auto-research diagnosis corrected to a proven #4593 symptom; validated by the three smokes, 94 focused tests and a canary premerge with 0 failures and 0 advisories.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head: ba1af7131cb6ab03a616142ca52ef22f0769385e (re-review after merging origin/main; previous review covered content-identical head c8e384a20a5f6ce54bc5033934e21026f1d102a1).

loopx pr-review --check-result on the matching packet returned ok: true with no approval_blockers; all 19 evidence rows for this control-plane plan are verified.

动机

main 的必过公开 smoke 因三个互不相关的真实回归常红,导致「必过检查」失去区分力:真正的产品回归和宿主机环境差异混在一起,维护者只能靠绕过推进。

Smoke main 上的根因 修复
loopx-chat-stream-throughput-smoke #4463 的终态事件保留让已完成 Turn 的行在读到终态时立刻被丢弃,于是每次 SSE 重放都重读整份日志(实测 20 次重放 → 20 次读,断言 <= 1) 已完成 Turn 在预算内驻留,重放命中缓存
loopx-managed-turn-operator-flow-smoke manager_channel_binding 引用了 managed_executor_binding 的可用性结论,却不接受 module_probe,所以在没有可选 dsh 运行时的宿主上,通道与同一次调用里计划的 Turn 会给出矛盾结论 把 module_probe 作为可选关键字参数透传(默认 None)
operator-provider-credential-smoke 该 smoke 要断言「缺凭证时命名缺凭证」,却把运行探针留给宿主机,于是没走到凭证判定就先报 dsh_runtime_unavailable smoke 声明它所断言的宿主事实;typed 拒绝本身保留

这个 PR 的价值:它不是在放宽断言,而是把三条 guard 恢复成「因为自己的原因通过/失败」。其中第一条是真实的吞吐回归(每次重连重读整份日志),第二条是可观测性缺口(任何同时解析两个读回的调用方都会撞上)。

改动思路

关键判断是两条义务需要同一个所有者。事件日志的缓存原本把「保留」与「淘汰」散在 _event_rows_locked / _drop_event_cache / append_event / events_after 四处,既没有表达「驻留预算」,也无法直接断言重放次数——只能靠观测文件日志间接验证。因此引入 loopx/chat_event_cache.py(126 行),把「已完成 Turn 的重放必须命中驻留行」和「驻留必须预算内」写成一处,chat_store 退化为持有者与转发者。

第二条是把既有的探针形态接上去:module_probe 是关键字参数、默认 None,未更新的调用方行为逐字不变;缺失运行时的 typed dsh_runtime_unavailable 语义没有被删除,只是让它继续由它自己的 smoke 和测试覆盖。

具体改动

  • 新增 loopx/chat_event_cache.py:ChatEventCache.load/get/put/drop/retain_finished、event_revision、TERMINAL_EVENT_CACHE_TURNS = 8。只读最后一行判终态,避免为判定而遍历即将跳过的行;retain_finished 返回被淘汰的键,供调用方审计而不是靠缓存未命中反推。
  • loopx/chat_store.py:三处缓存逻辑收敛为一行转发,_event_cache / _event_cache_revision 仍以同一对象暴露以保持既有读法;文件从 1501 行降到 1484 行,未抬高评审过的 1500 行上限。
  • loopx/chat_manager.py:manager_channel_binding(..., module_probe=None) 并在 executor_kind == managed 时透传给 managed_executor_binding。
  • examples/loopx-managed-turn-operator-flow-smoke.py / examples/operator-provider-credential-smoke.py:各自注入它要断言的宿主事实。
  • tests/test_chat_event_retention.py:把旧断言换成两条真实契约——「5 次重放只读 1 次」与「驻留数 == 预算且最旧者被淘汰」。

关键代码讲解

def retain_finished(self, key, rows, *, terminal_kinds):
    if not rows or rows[-1].get("kind") not in terminal_kinds:
        return ()
    with self._lock:
        self._finished.pop(key, None)
        self._finished[key] = None          # 重新置为最新
        evicted = []
        while len(self._finished) > self._budget:
            oldest = next(iter(self._finished))
            ...
        return tuple(evicted)

dict 的插入序即驻留顺序,pop 再插入让重复重放的 Turn 回到最新位,while 保证驻留严格有界。返回值让淘汰可被审计,而不是静默发生。

def load(self, key, path, read):
    cached = self.get(key, path)
    if cached is not None:
        return cached
    rows = read(path)
    self.put(key, rows, path)
    return rows

命中必须同时满足「行存在」与「on-disk revision 一致」(st_ino, st_size, st_mtime_ns)。所以缓存只是派生视图:日志被追加或改写后 revision 不同,必然回落到读盘,不存在用陈旧行服务重放的可能。

对主干的风险

唯一刻意的行为变化是「已完成的 Turn 在预算内可驻留」,这正是修复目标,并由两条新断言与一条 smoke 直接覆盖。写侧仍不为终态历史驻留(既有断言保持),executor_kind 非 managed 时探针完全未被读取,无 dsh 时的 typed 拒绝语义未变。

本次在更新后的 head 上实测(该 head 只是把 origin/main —— 含已合并的 #4593、共 15 个提交 —— 合入;6 个被评审文件的 blob 哈希在旧 head 与新 head 上逐一相同,已用 git rev-parse 核对,例如 loopx/chat_store.py f2a347b383a0、loopx/chat_event_cache.py b741de2af6bb):

  • 三条目标 smoke 全部通过:loopx-chat-stream-throughput-smoke / loopx-managed-turn-operator-flow-smoke / operator-provider-credential-smoke
  • env -u PYTHONPATH uv run --extra test python -m pytest tests/test_chat_event_retention.py tests/test_chat_event_buffer.py tests/test_chat_event_cursor.py tests/test_chat_manager_context.py tests/test_chat_manager_details.py tests/test_chat_session_active_turn.py tests/test_manager_channel_binding.py tests/test_turn_managed_executor_binding.py -q → 114 passed
  • env -u PYTHONPATH uv run --extra test loopx canary premerge --from-git-diff → 0 failures / 0 advisories(含 control-plane-maintainability-ratchet-smoke,未抬高上限;public boundary ok)

未验证 / 手工保留项(如实标注,不当通过):TERMINAL_EVENT_CACHE_TURNS = 8 是判断而非测量,预算内可驻留的大 Turn(大量 delta 行)内存上限未测量;未做跨进程重启的长跑 SSE 重放观测,吞吐断言只在单进程内成立——这与线上 SSE 端点一致,但重启后首次重放仍会读盘一次,属设计边界。

边界声明:本 PR 改动 loopx/** 运行时行为,属控制面改动,按仓库规则只提 PR、由维护者合并,作者不做自合并。

我的整体评价

把两条真实义务放进同一个可测的所有者、并把可观测性缺口接上,方向正确;改动比例与它解除的阻塞相称(126 行新模块换回三条必过 guard,同时把 chat_store.py 压回上限以下)。无阻断性发现。

main 剩余红项只剩一条:examples/cli-help-manpage-smoke.py 报 {'unclassified': ['agent-directory']},由 #4613 承载(已在新 main 9060ddc91 上复现)。本 PR 与 #4613 合并后,Full Public Smokes 应恢复全绿,随后重跑 canary promotion-readiness 预检以替换陈旧证据。

English verdict: APPROVE - re-verified on exact head ba1af71 (content-identical to the previously reviewed c8e384a; only a main merge, proven by identical blob hashes): the three red public smokes are repaired at their real causes rather than by relaxing assertions - finished-Turn replay now reads its log once within a bounded budget, manager_channel_binding forwards the same module_probe it already gets, and the credential smoke states the host fact it asserts while the typed dsh_runtime_unavailable refusal keeps its own coverage. 114 focused tests pass, 3 target smokes pass, premerge canary 0 failures with no ratchet raised. No blocking finding.

…ke-repair

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/main-public-smoke-repair branch from ba1af71 to a101e8f Compare September 17, 2026 07:05

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head: a101e8fe23c8644f9be24d8e9d86f3b2122b0500.

loopx pr-review --check-result on the matching packet returned ok: true with no approval_blockers; all 19 evidence rows for this control-plane plan are verified. The reviewed content is byte-identical to the last reviewed head ba1af7131 (and to the earlier c8e384a20): git log ba1af7131..a101e8fe2 contains exactly one commit — the merge of origin/main (9060ddc91) — and all six changed-file blobs hash the same on both heads (chat_event_cache.py b741de2af6bb, chat_store.py f2a347b383a0, chat_manager.py f5c4c97e8ade, operator-provider-credential-smoke.py f6ff9082dfd9, loopx-managed-turn-operator-flow-smoke.py b3badaad28dd, test_chat_event_retention.py af4093d9b4b2). The earlier COMMENTED approve on this PR was not bound to the current head and its body carried no standalone bilingual sections, so this review replaces that record for this head.

动机

main 的公开 smoke 套件是每个开放 PR 的必过门(checks / pytest / merge-gate 都依赖它),而它因三个互不相关的真实回归常红。红灯不区分「产品坏了」和「这台机器没装可选运行时」,于是必过门失去区分力,只能靠绕过推进——这正是本 lane 的 P0 Todo(todo_28e0823d6392)要恢复的契约:每条 guard 必须因为自己的原因失败。

这个 PR 的价值:把三条真实回归各自修回它自己的契约,而不是把断言放宽成绿灯。

在干净 main(9060ddc91)上实测的失败面:

  • loopx-chat-stream-throughput-smoke:AssertionError: SSE replay reread the event log 20 times(smoke 断言 <= 1)
  • loopx-managed-turn-operator-flow-smoke:manager_channel_binding 返回 available: False / unavailable_reason: dsh_runtime_unavailable,而同一个调用方在计划 Turn 侧拿不到同一份事实
  • operator-provider-credential-smoke:managed_executor_binding 先报 dsh_runtime_unavailable,根本没走到「缺哪个凭证」的判定

改动思路

三条修复的公共判断是把事实的所有者放对,而不是让消费者各自猜:

  1. 事件日志的「已完成 Turn 可以驻留」与「驻留必须有限」是同一个缓存的两条义务,散落在 _event_rows_locked / _drop_event_cache / append_event / events_after 四处。现在把它们收进一个 126 行的有界上下文 loopx/chat_event_cache.py,chat_store.py 只保留文件 IO、顺序与锁——顺带把 chat_store.py 从 1501 行降到 1484 行,没有抬高任何评审过的 hot-file 上限。
  2. 通道读回引用执行器绑定,却拿不到调用方用的模块探针;于是同一台机器上「通道说不可用、计划说可用」。修法是让探针可选地透传(关键字参数、默认 None),既有调用方行为逐字不变。
  3. 第三条根因在 smoke 自己:它问的是「哪个凭证认证托管宿主」,却把运行探针留给宿主机,于是先撞上运行时判定。smoke 现在声明它所断言的宿主事实;dsh_runtime_unavailable 这条 typed 拒绝仍由 examples/loopx-turn-managed-executor-binding-smoke.py 与 tests/test_turn_managed_executor_binding.py 保留覆盖。

同时本分支刻意不重复已在别处承载的工作:agent-directory 分类归 #4613,extension-entrypoint-surface 的导入引导归 #4615,避免同一条帮助面出现两个互相冲突的决定。

替代方案与取舍:只改 smoke 不动生产代码对第三条成立,但对前两条不成立——透传探针缺口是真实可观测性缺口(任何同时解析两个读回的调用方都会撞上),而每次重放重读整份日志是真实吞吐回归,放宽断言等于删掉契约。

具体改动

6 个文件、+214/-41:

  • loopx/chat_event_cache.py(新增,126 行):event_revision()、ChatEventCache.load/get/put/drop/retain_finished/resident_finished_keys、常量 TERMINAL_EVENT_CACHE_TURNS = 8。
  • loopx/chat_store.py(1484 行):四处缓存事实收敛为转发;_event_cache / _event_cache_revision 仍以同一对象暴露(指向 ChatEventCache.rows/revisions),既有读取方(含 tests/test_chat_event_cursor.py 注入非法行对象的接缝)不受影响。
  • loopx/chat_manager.py(+11/-2):manager_channel_binding(..., module_probe=None) 透传给 managed_executor_binding。
  • examples/loopx-managed-turn-operator-flow-smoke.py(+1):通道侧注入与执行器侧同一个探针。
  • examples/operator-provider-credential-smoke.py(+16/-1):声明要断言的运行时已安装,不再读宿主机状态。
  • tests/test_chat_event_retention.py(+44/-6):旧断言换成两条真实契约——「5 次重放只读 1 次」与「驻留数 == 预算且最旧被淘汰」。

关键代码讲解

retain_finished 是这次修复的核心,它同时表达两条义务,并把淘汰结果作为返回值交出去而不是静默丢弃:

if not rows or rows[-1].get("kind") not in terminal_kinds:
    return ()
with self._lock:
    self._finished.pop(key, None)
    self._finished[key] = None
    evicted: list[EventKey] = []
    while len(self._finished) > self._budget:
        oldest = next(iter(self._finished))
        self._finished.pop(oldest, None)
        self.rows.pop(oldest, None)
        self.revisions.pop(oldest, None)
        evicted.append(oldest)
return tuple(evicted)
  • 只读最后一行判断终态:重放不得为了判定而遍历它马上要跳过的行(大 Turn 的第一行可能很长)。
  • pop + 重插把该键移到插入序末尾,_finished 因此是一个按最近重放排序的 LRU;while 保证驻留数永不超过预算,max(1, budget) 保证下界。
  • 命中前提仍是 revision 一致(get),所以缓存永远是磁盘 JSONL 的可丢弃派生视图,不构成第二权威;文件被追加或改写后必然回落读盘。

chat_manager.py 侧只有一处转发,边界写在关键字参数默认值上:

if executor_kind == MANAGER_EXECUTOR_KIND_MANAGED:
    managed = managed_executor_binding(
        endpoint, environ=environ, module_probe=module_probe
    )

默认 None 时解析路径与之前逐字一致,因此未更新的前端/CLI 调用方行为不变;executor_kind 非 managed 时该参数完全不被读取(default_off_isolation 的对照分支)。

语义与 CI 对齐

candidate_decision=reuse_existing:本次没有扩张词表,只是把既有缓存的三件套收敛为一个所有者、并给一个既有读回接上既有探针参数。CI 侧本 PR 是控制面改动,只提 PR、由维护者合并;author 未自合并。

对主干的风险

反向风险是主要风险面,两个方向都有直接断言兜底:

  • 无界驻留(内存随完成的 Turn 数线性增长)→ test_completed_turn_retention_is_bounded_to_the_replay_budget 断言驻留数恰好等于预算、最旧键已被淘汰;
  • 写侧重新驻留终态历史(等于把 #4463 的语义改坏)→ 既有断言仍要求写完终态行后 key not in store._event_cache;
  • 通道侧把探针变成默认注入假值(等于把真实缺口藏起来)→ tests/test_turn_managed_executor_binding.py 的 dsh 缺失用例仍要求 available=False / dsh_runtime_unavailable。

本 head 的实测(PYTHONPATH 置空,与 CI 环境一致):

  • 三条目标 smoke exit 0;同一组在干净 main(9060ddc91)上以上述根因 exit 1(A/B 归因)。
  • 12 个聚焦测试文件:161 passed in 21.03s。
  • loopx canary premerge --from-git-diff:0 failures / 0 advisories(4/4 catalog canary 通过,含 maintainability ratchet;public boundary 通过)。
  • examples/control_plane/control-plane-maintainability-ratchet-smoke.py:unreviewed=0、stale_exceptions=0、magnitude_regressions=0,chat_store.py 1484 行 / chat_event_cache.py 126 行——没有抬高上限。
  • 另在 main 上验证了 PR 描述里的前提:repository-hygiene-smoke 与 auto-research-rollout-readpath-smoke 现在 exit 0(它们的共同原因已由 #4593 清除),所以剩余红项确实只有本 PR 的三条加 #4613 的一条。

未验证 / 如实标注:TERMINAL_EVENT_CACHE_TURNS = 8 是判断值而非测量值,预算内驻留大 Turn 的内存上限未测量;同一进程内的「一读多放」已被断言覆盖,跨进程/跨重启的首次访问仍会读盘一次(设计边界,非缺陷);本机未运行完整 Full Public Smokes,全绿套件的最终证据仍在 CI。

边界声明:本 PR 改动 loopx/**,属控制面改动,只提 PR、由维护者合并,作者不做自合并。

我的整体评价

方向、所有者归属与验证都对,且没有用放宽断言的方式换取绿灯——三个根因里两个在生产代码、一个在 smoke 自己的宿主假设,各自的 typed 拒绝与既有断言都保留了下来。无阻断性发现。

两点给维护者留意:

  1. 合并顺序:main 上只剩四条红项,本 PR 的三条 + #4613 的一条。#4613 在本轮实测已是 ready=true(有效评审 + 26/26 检查通过),两个都合上后 Full Public Smokes 与 canary-promotion-readiness 的陈旧证据才会被替换。
  2. 与 #4623 的相邻改动:两者都改 loopx/chat_manager.py::manager_channel_binding(本 PR 加 module_probe 转发,#4623 加 runtime_probe 转发),hunk 不重叠,谁后合谁需要一个极小的 rebase;语义上是互补的——本 PR 让通道与计划 Turn 能用同一个探针,#4623 让读回说明那个结论属于哪个环境。

本 PR 的边界声明不变:控制面改动,只提 PR、由维护者合并。

English verdict: APPROVE - re-verified on exact head a101e8f (content byte-identical to the previously reviewed ba1af71; the only commit since is the merge of origin/main 9060ddc). It restores three independent public smoke contracts that main really broke, and I reproduced the A/B: on clean main 9060ddc the throughput smoke fails with "SSE replay reread the event log 20 times", the operator-flow smoke's channel binding reports dsh_runtime_unavailable while the planned Turn disagrees, and the credential smoke never reaches its credential verdict - all three exit 0 on this head. The event-log retention fix states both obligations in one bounded owner (rows stay resident; only the most recent finished Turns stay, evicted past a budget of 8) and drops chat_store.py from 1501 to 1484 lines without raising any reviewed ceiling; the module_probe forward is keyword-only with default None, so existing callers are unchanged and the dsh_runtime_unavailable refusal keeps its own coverage. 161 focused tests pass, loopx canary premerge --from-git-diff reports 0 failures / 0 advisories, and the maintainability ratchet shows no unreviewed or magnitude regression. Unverified and disclosed: the budget of 8 is a judgement rather than a measurement, cross-process replay still reads once per restart, and the full suite's green evidence remains CI. No blocking finding.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head: cf475a6268190093192d22f1424e69512e57c40e.

loopx pr-review --check-result on the matching packet returned ok: true with no approval_blockers; all 19 evidence rows for this control-plane plan are verified. The reviewed content is byte-identical to every earlier reviewed head (ba1af7131, c8e384a20, a101e8fe2): the commits since are merges of origin/main, most recently e160b1bbc (#4633, team-plan atomic commit + same-operation recovery). All six changed-file blobs hash the same on every one of those heads (chat_event_cache.py b741de2af6bb, chat_store.py f2a347b383a0, chat_manager.py f5c4c97e8ade, operator-provider-credential-smoke.py f6ff9082dfd9, loopx-managed-turn-operator-flow-smoke.py b3badaad28dd, test_chat_event_retention.py af4093d9b4b2), and git diff --name-only 9060ddc91..e160b1bbc has no intersection with this PR's files.

动机

main 的公开 smoke 套件是每个开放 PR 的必过门(checks / pytest / merge-gate 都依赖它),而它因三个互不相关的真实回归常红。红灯不区分「产品坏了」和「这台机器没装可选运行时」,于是必过门失去区分力,只能靠绕过推进——这正是本 lane 的 P0 Todo(todo_28e0823d6392)要恢复的契约:每条 guard 必须因为自己的原因失败。

这个 PR 的价值:把三条真实回归各自修回它自己的契约,而不是把断言放宽成绿灯。

在干净 main(9060ddc91)上实测的失败面:

  • loopx-chat-stream-throughput-smoke:AssertionError: SSE replay reread the event log 20 times(smoke 断言 <= 1)
  • loopx-managed-turn-operator-flow-smoke:manager_channel_binding 返回 available: False / unavailable_reason: dsh_runtime_unavailable,而同一个调用方在计划 Turn 侧拿不到同一份事实
  • operator-provider-credential-smoke:managed_executor_binding 先报 dsh_runtime_unavailable,根本没走到「缺哪个凭证」的判定

改动思路

三条修复的公共判断是把事实的所有者放对,而不是让消费者各自猜:

  1. 事件日志的「已完成 Turn 可以驻留」与「驻留必须有限」是同一个缓存的两条义务,散落在 _event_rows_locked / _drop_event_cache / append_event / events_after 四处。现在把它们收进一个 126 行的有界上下文 loopx/chat_event_cache.py,chat_store.py 只保留文件 IO、顺序与锁——顺带把 chat_store.py 从 1501 行降到 1484 行,没有抬高任何评审过的 hot-file 上限。
  2. 通道读回引用执行器绑定,却拿不到调用方用的模块探针;于是同一台机器上「通道说不可用、计划说可用」。修法是让探针可选地透传(关键字参数、默认 None),既有调用方行为逐字不变。
  3. 第三条根因在 smoke 自己:它问的是「哪个凭证认证托管宿主」,却把运行探针留给宿主机,于是先撞上运行时判定。smoke 现在声明它所断言的宿主事实;dsh_runtime_unavailable 这条 typed 拒绝仍由 examples/loopx-turn-managed-executor-binding-smoke.py 与 tests/test_turn_managed_executor_binding.py 保留覆盖。

同时本分支刻意不重复已在别处承载的工作:agent-directory 分类归 #4613,extension-entrypoint-surface 的导入引导归 #4615,避免同一条帮助面出现两个互相冲突的决定。

替代方案与取舍:只改 smoke 不动生产代码对第三条成立,但对前两条不成立——透传探针缺口是真实可观测性缺口(任何同时解析两个读回的调用方都会撞上),而每次重放重读整份日志是真实吞吐回归,放宽断言等于删掉契约。

具体改动

6 个文件、+214/-41:

  • loopx/chat_event_cache.py(新增,126 行):event_revision()、ChatEventCache.load/get/put/drop/retain_finished/resident_finished_keys、常量 TERMINAL_EVENT_CACHE_TURNS = 8。
  • loopx/chat_store.py(1484 行):四处缓存事实收敛为转发;_event_cache / _event_cache_revision 仍以同一对象暴露(指向 ChatEventCache.rows/revisions),既有读取方(含 tests/test_chat_event_cursor.py 注入非法行对象的接缝)不受影响。
  • loopx/chat_manager.py(+11/-2):manager_channel_binding(..., module_probe=None) 透传给 managed_executor_binding。
  • examples/loopx-managed-turn-operator-flow-smoke.py(+1):通道侧注入与执行器侧同一个探针。
  • examples/operator-provider-credential-smoke.py(+16/-1):声明要断言的运行时已安装,不再读宿主机状态。
  • tests/test_chat_event_retention.py(+44/-6):旧断言换成两条真实契约——「5 次重放只读 1 次」与「驻留数 == 预算且最旧被淘汰」。

关键代码讲解

retain_finished 是这次修复的核心,它同时表达两条义务,并把淘汰结果作为返回值交出去而不是静默丢弃:

if not rows or rows[-1].get("kind") not in terminal_kinds:
    return ()
with self._lock:
    self._finished.pop(key, None)
    self._finished[key] = None
    evicted: list[EventKey] = []
    while len(self._finished) > self._budget:
        oldest = next(iter(self._finished))
        self._finished.pop(oldest, None)
        self.rows.pop(oldest, None)
        self.revisions.pop(oldest, None)
        evicted.append(oldest)
return tuple(evicted)
  • 只读最后一行判断终态:重放不得为了判定而遍历它马上要跳过的行(大 Turn 的第一行可能很长)。
  • pop + 重插把该键移到插入序末尾,_finished 因此是一个按最近重放排序的 LRU;while 保证驻留数永不超过预算,max(1, budget) 保证下界。
  • 命中前提仍是 revision 一致(get),所以缓存永远是磁盘 JSONL 的可丢弃派生视图,不构成第二权威;文件被追加或改写后必然回落读盘。

chat_manager.py 侧只有一处转发,边界写在关键字参数默认值上:

if executor_kind == MANAGER_EXECUTOR_KIND_MANAGED:
    managed = managed_executor_binding(
        endpoint, environ=environ, module_probe=module_probe
    )

默认 None 时解析路径与之前逐字一致,因此未更新的前端/CLI 调用方行为不变;executor_kind 非 managed 时该参数完全不被读取(default_off_isolation 的对照分支)。

语义与 CI 对齐

candidate_decision=reuse_existing:本次没有扩张词表,只是把既有缓存的三件套收敛为一个所有者、并给一个既有读回接上既有探针参数。CI 侧本 PR 是控制面改动,只提 PR、由维护者合并;author 未自合并。

对主干的风险

反向风险是主要风险面,两个方向都有直接断言兜底:

  • 无界驻留(内存随完成的 Turn 数线性增长)→ test_completed_turn_retention_is_bounded_to_the_replay_budget 断言驻留数恰好等于预算、最旧键已被淘汰;
  • 写侧重新驻留终态历史(等于把 #4463 的语义改坏)→ 既有断言仍要求写完终态行后 key not in store._event_cache;
  • 通道侧把探针变成默认注入假值(等于把真实缺口藏起来)→ tests/test_turn_managed_executor_binding.py 的 dsh 缺失用例仍要求 available=False / dsh_runtime_unavailable。

本 head 的实测(PYTHONPATH 置空,与 CI 环境一致):

  • 三条目标 smoke exit 0(在含 #4633 的新 head 上重跑);同一组在干净 main(9060ddc91)上以上述根因 exit 1(A/B 归因)。
  • 12 个聚焦测试文件:161 passed in 25.24s(合并 #4633 后重跑)。
  • loopx canary premerge --from-git-diff:0 failures / 0 advisories(4/4 catalog canary 通过,含 maintainability ratchet;public boundary 通过)。
  • examples/control_plane/control-plane-maintainability-ratchet-smoke.py:unreviewed=0、stale_exceptions=0、magnitude_regressions=0,chat_store.py 1484 行 / chat_event_cache.py 126 行——没有抬高上限。
  • 另在 main 上验证了 PR 描述里的前提:repository-hygiene-smoke 与 auto-research-rollout-readpath-smoke 现在 exit 0(它们的共同原因已由 #4593 清除),所以剩余红项确实只有本 PR 的三条加 #4613 的一条。

未验证 / 如实标注:TERMINAL_EVENT_CACHE_TURNS = 8 是判断值而非测量值,预算内驻留大 Turn 的内存上限未测量;同一进程内的「一读多放」已被断言覆盖,跨进程/跨重启的首次访问仍会读盘一次(设计边界,非缺陷);本机未运行完整 Full Public Smokes,全绿套件的最终证据仍在 CI。

边界声明:本 PR 改动 loopx/**,属控制面改动,只提 PR、由维护者合并,作者不做自合并。

我的整体评价

方向、所有者归属与验证都对,且没有用放宽断言的方式换取绿灯——三个根因里两个在生产代码、一个在 smoke 自己的宿主假设,各自的 typed 拒绝与既有断言都保留了下来。无阻断性发现。

两点给维护者留意:

  1. 合并顺序:main 上只剩四条红项,本 PR 的三条 + #4613 的一条。#4613 也已 ready=true。注意 main 已前进到 e160b1bbc(#4633 合并),因此本 head 与 #4613 都会在合并前再次变成 BEHIND:预期顺序是先更新分支、再在新 head 上重发同一份评审、然后逐个合并。
  2. 与 #4623 的相邻改动:两者都改 loopx/chat_manager.py::manager_channel_binding(本 PR 加 module_probe 转发,#4623 加 runtime_probe 转发),hunk 不重叠,谁后合谁需要一个极小的 rebase;语义上是互补的——本 PR 让通道与计划 Turn 能用同一个探针,#4623 让读回说明那个结论属于哪个环境。

本 PR 的边界声明不变:控制面改动,只提 PR、由维护者合并。

English verdict: APPROVE - re-verified on exact head cf475a6 (content byte-identical to the earlier reviewed heads ba1af71 / c8e384a / a101e8f; the only commits since are merges of origin/main, most recently e160b1b / #4633). It restores three independent public smoke contracts that main really broke, and I reproduced the A/B: on clean main 9060ddc the throughput smoke fails with "SSE replay reread the event log 20 times", the operator-flow smoke's channel binding reports dsh_runtime_unavailable while the planned Turn disagrees, and the credential smoke never reaches its credential verdict - all three exit 0 on this head. The event-log retention fix states both obligations in one bounded owner (rows stay resident; only the most recent finished Turns stay, evicted past a budget of 8) and drops chat_store.py from 1501 to 1484 lines without raising any reviewed ceiling; the module_probe forward is keyword-only with default None, so existing callers are unchanged and the dsh_runtime_unavailable refusal keeps its own coverage. 161 focused tests pass (re-run on this head after #4633 merged), loopx canary premerge --from-git-diff reports 0 failures / 0 advisories, and the maintainability ratchet shows no unreviewed or magnitude regression. Unverified and disclosed: the budget of 8 is a judgement rather than a measurement, cross-process replay still reads once per restart, and the full suite's green evidence remains CI. No blocking finding.

@huangruiteng

Copy link
Copy Markdown
Collaborator Author

Superseded by #4648 (bounded Chat replay retention) and #4647 (runtime/credential smoke isolation). The cache behavior is now an explicit runtime change with count, row and encoded-size limits, revision invalidation and stale-snapshot protection. A small dedicated cache replaces the proposed public mutable maps and compatibility aliases; the store retains file I/O and locks. The unchanged baseline rereads 20 times; the replacement passes the 20-replay throughput contract with one read.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant