Skip to content

fix(update): activate managed services when only extensions are blocked - #4458

Merged
huangruiteng merged 1 commit into
mainfrom
codex/restart-managed-services-after-archive-activation
Sep 15, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/restart-managed-services-after-archive-activation

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

Why

loopx update apply replaced the installed release snapshot but restarted the
LaunchAgent-managed status/chat services only inside its fully-clean branch.
On this machine a real update apply produced:

  • install rc 0, core doctor rc 0, release snapshot promoted;
  • extension doctor --all-enabled --execute rc 1 because two pre-existing
    enabled providers (ark-flow-lane-agent-runtime, ark-flow-lane-delivery)
    were provider_unavailable/probe_nonzero_exit;
  • therefore restarted_services was never attempted, and the next action told
    the operator to review_or_rollback a runtime that was already installed and
    healthy.

The operator-visible effect: the local steward chat kept serving the previous
release's packaged frontend (index-DscLlBrK.js, no steward execution chip)
until the services were restarted by hand. Activation silently depended on the
health of unrelated extension providers.

What changed

  • The restart decision now follows the runtime install plus its core doctor
    readback. Blocked enabled extension providers no longer prevent activation.
  • When the runtime is installed but an enabled provider is blocked, the outcome
    reports repair_blocked_extensions (non-mutating) instead of
    review_or_rollback, and no longer recommends rolling back a healthy runtime.
    ok stays false so the blocked provider remains a visible defect.
  • restart_managed_loopx_services() and the new activation decision moved to
    loopx/runtime_activation.py, a bounded owner for "how an installed runtime
    starts serving". loopx/self_update.py was at 1480 lines against a 1500-line
    module budget and now stays inside it (1485) instead of needing an exception.
  • Executed loopx/semantics/inventory_v0.json regeneration for the new module.

Both execution drivers (archive snapshot and pip/pipx) now share one rule; the
pip path keeps requiring its host-material steps (workflow-skills,
slash-commands) before activation.

Validation

  • loopx canary premerge --from-git-diff --timeout-seconds 300:
    selected=14 executed=14 failures=0, public boundary passed. (The default
    120s per-check budget times out examples/install-local-smoke.py, which needs
    ~157s on this host; it passes standalone.)
  • pytest tests/test_self_update_runtime_activation.py tests/test_archive_installer_commit_response.py tests/test_doctor_install_freshness.py tests/canary/test_maintainability_ratchet.py -q → 63 passed, 1 skipped.
  • examples/loopx-update-smoke.py, examples/control_plane/control-plane-maintainability-ratchet-smoke.py,
    examples/semantic-vocabulary-drift-smoke.py, examples/docs-governance-smoke.py,
    examples/dev-book-publication-smoke.py → ok.
  • New regression tests cover both directions: extensions blocked → services
    restarted + repair action; runtime unhealthy → no restart + rollback action.
  • Live local readback after the upgrade: http://127.0.0.1:8767/api/chat/capabilities
    now reports runtime_identity.release_id=20260915T144512Z,
    source_revision=9719dc0d4, and the steward channel binding
    codex · individual · gpt-6-astra.

Notes

  • Steward chat channel default stays codex; managed/worker execution stays the
    credential-resolved Turn host. This PR does not change either default.
  • Docs: one paragraph each in the English and Chinese safe-upgrade runbooks so
    the activation semantics are stated where operators read them.

`loopx update apply` restarted the LaunchAgent-managed status/chat services
only in its fully-clean branch. A blocked enabled extension provider made the
run report `review_or_rollback` even though the release snapshot was installed
and core doctor readback passed, so status/chat kept serving the previous
release until an operator restarted them by hand.

Decide the restart from the runtime install plus its core doctor readback,
report blocked providers as a repair action instead of a rollback, and move
the LaunchAgent restart into a bounded runtime_activation owner so the update
boundary keeps describing what was installed rather than how it starts
serving.

The new archive-path cases carry the same native-Windows skip as the existing
archive update test, because `execute_update_plan` fails closed before the
installer on that platform.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/restart-managed-services-after-archive-activation branch from c451e52 to 6cf3cdc Compare September 15, 2026 15:23

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

精确 head:6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b(取代 c451e52e8b36185c13460c50114bcfb1efc663d4,差异只是测试的平台守卫)。评审按 pull_request_review_execution_contract_v2(policy_revision=3)执行,pr-review --check-result 对该 head 返回 approval_consistent=true、无 blocker。

动机

loopx update apply 会替换已安装的 release snapshot,但 LaunchAgent 托管的 status/chat 服务仍然跑着启动时的旧代码。旧实现只在“所有步骤全绿”的分支里重启服务,而 extension doctor --all-enabled --execute 只要有一个已启用的 provider 被阻塞就会返回非零,于是更新直接落到 review_or_rollback 分支:既没有重启,又建议操作者回滚一个其实已经装好、核心 doctor 也读通的运行时。

这是本机真实发生的路径:install rc 0、核心 doctor rc 0、release snapshot 已提升,但因为两个既有的 provider 阻塞(ark-flow-lane-agent-runtime、ark-flow-lane-delivery,provider_unavailable/probe_nonzero_exit),服务的重启被跳过。操作者看到的是“机器报告了新 release,管家聊天却继续吐旧版本打包前端”,必须手动 launchctl kickstart 才能对齐;而这些 provider 的健康与“能不能把服务切到新 release”本来没有因果关系,属于把两件事绑在一起。

影响面是本机操作者加上所有 chat/dashboard 消费者;成本是静默的版本错配(报告的身份 ≠ 实际服务的身份),恢复方式只有手动重启或走回滚。更小的修法试过两条:只放宽 ok 计算,重启仍然留在全绿分支里,观察到的结果不变;安装返回 0 就无条件重启,则会把服务切到一个核心 doctor 读不通的 release 上。以“安装 + 核心 doctor 读回”作为切换门槛,是能覆盖该问题的 B/A 折中。

非目标:不改管家 chat 通道默认(仍是 codex)、不改由凭据解析的 managed Turn host、不修复那两个 provider、不动 installer 下载与 release 提升规则。

改动思路

入口是 loopx update apply → loopx.self_update.execute_update_plan(归档快照驱动)与 _execute_python_distribution_update(pip/pipx 驱动)。权威输入是 install_lifecycle.execution_driver、plan.backup.current_release_root,以及 install / 核心 doctor / extension doctor 三个返回码。改动把原来一个布尔拆成两个派生事实:runtime_ready = install rc == 0 and doctor rc == 0 决定“能不能切换”,ok = runtime_ready and 已启用扩展复检通过 决定“这次更新干不干净”。

生效副作用(launchctl kickstart -k)与判定逻辑被抽到 loopx/runtime_activation.py:restart_managed_loopx_services() 从 loopx/self_update.py 原样搬过去,新增 restart_services_for_runtime_activation() 作为唯一规则所有者,两个驱动都调它。原先那份模块内定义被删除而不是包装保留,仓内其他引用只有测试模块,已同 head 更新。抽出的直接原因是:修复本身会把 loopx/self_update.py 从基线 1480 行推到 1543 行,超过 1500 行模块预算;抽出后是 1485 行,等于把债务消掉而不是登记例外。

具体改动

六个文件 +212/-52。分类:生产代码 loopx/self_update.py(净 +5 行)与新增 loopx/runtime_activation.py(74 行,主体是搬移);测试 tests/test_self_update_runtime_activation.py(+84/-6);公开文档两篇 docs/book/en/chapters/appendix-reference.md 与 docs/book/chapters/appendix-reference.md 各一段;生成物 loopx/semantics/inventory_v0.json(source_files 1175 → 1176,由 scripts/generate_semantic_inventory.py 重新生成)。

关键代码讲解

  1. loopx/runtime_activation.py:52 restart_services_for_runtime_activation:唯一判定点。仅当 changes_applied and runtime_ready 时才调用重启,返回 {restarted_services, restart_status};否则返回空列表加 restart_status=skipped_runtime_not_activated。扩展 provider 健康被明确排除在这个判定之外。
  2. loopx/self_update.py:1030 execute_update_plan:runtime_ready 由 install/doctor 返回码构成,ok 再叠加扩展复检;installed 但扩展被阻塞 这条路径现在返回非阻塞的 next_action.kind=repair_blocked_extensions(mutating=false,命令是操作者自己的 loopx extension doctor --all-enabled --execute),并保留 ok=false 让阻塞可见。install 或核心 doctor 失败时仍走原 review_or_rollback 与回滚命令。
  3. loopx/self_update.py:861 _execute_python_distribution_update:同一规则;差异是 runtime_steps 里仍要求 workflow-skills、slash-commands 主机材料步骤通过,避免 PyPI 更新把服务切到技能/命令未安装的环境。
  4. loopx/runtime_activation.py:17 restart_managed_loopx_services:原样搬移的 LaunchAgent 选择与重启逻辑(仅匹配 loopx/goal-harness 且以 .status/.chat 结尾的 plist,逐条 30s 超时,best-effort)。公开语义未变。

对主干的风险

最强回归场景是“install rc 0 但核心 doctor 读回失败”——新门槛若判错,会把服务切到坏 release。防线是 runtime_ready 同时要求两者为 0 返回码,且负向用例 test_archive_apply_skips_managed_service_restart_when_the_runtime_is_unhealthy 断言此时不重启、并保留 review_or_rollback 与回滚命令。爆炸半径仅限本机用户级服务(com.loopx.chat、com.loopx.status 及同类 plist),不触碰 release 提升、quota、todo 与 registry 状态;可观测性上 restarted_services/restart_status 在所有分支都会写出并渲染进 markdown 计划。

真正值得记录的是上一 head 的反例:Windows CI 会原生跑这个测试模块,而 execute_update_plan 在 os.name == 'nt' 下会先返回 unsupported_platform,导致两个新的归档路径用例在那里失败。当前 head 给这两个用例补了与该文件既有归档用例一致的 native-Windows skip,windows-powershell 已通过,同时 test_windows_execute_update_fails_closed_without_launching_bash 继续钉住 Windows 的 fail-closed 行为。校验覆盖:本地 pytest 63 passed/1 skipped、loopx canary premerge --from-git-diff --timeout-seconds 300 14/14 通过(默认 120s 会在本机把 install-local-smoke 判超时,单独跑约 157s 通过)、loopx-update-smoke、maintainability ratchet(unreviewed 0)、semantic inventory drift、两篇文档 smoke,以及该 head 全部 24 项非 skip 的 CI 检查。

scope_fit:真实生产调用点存在并在本机端到端跑过;change_proportionality:机制成本是一个 74 行、以搬移为主、零新增持久化状态与 CLI 选项的模块,对照的是一个每次更新都可能静默复现的版本错配。

我的整体评价

基线/head 对比:全绿路径不变;install ok + doctor ok + 扩展阻塞 从“ok=false、不重启、建议回滚”变成“ok=false、记录重启、给出修复动作”;install/doctor 失败路径在两个版本都不重启且保留回滚。等价性风险集中在新增的 next_action.kind=repair_blocked_extensions 消费方,未知 kind 在现渲染里按文本处理,属可接受。

结论:APPROVE。没有阻塞项,按 P2 留下的残余风险是 Windows 激活分支刻意不可达因而未被真实覆盖,以及 launchctl 语义靠未改动的选择用例加一次真实本机 update apply 验证。若下一步要更硬的证据,值得补一条 CI 级集成:在模拟主机上让 extension doctor 返回非零,断言 restarted_services 被记录。

English verdict: APPROVE — exact head 6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b. The update driver now gates the managed-service restart on the runtime install plus core doctor readback, so a blocked enabled extension provider no longer leaves status/chat silently serving the previous release, and the outcome reports a non-mutating repair_blocked_extensions step instead of a rollback of a healthy runtime. Evidence: targeted and negative unit cases for both drivers, loopx canary premerge --from-git-diff 14/14 on the changed surfaces, loopx-update-smoke, maintainability ratchet and semantic inventory drift green, and all 24 non-skipped CI checks passing on this head (the superseded head failed the native Windows job, which the exact-head platform guards fix). Key finding: none blocking; residual risk is the intentionally unreachable Windows activation branch and launchctl semantics verified by an unchanged selection test plus a real local upgrade rather than a CI integration test.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

动机

loopx update apply 只在"全部干净"的分支里重启 LaunchAgent 托管的 status/chat,而"全部干净"包含了 enabled 扩展健康。现场结果是:release snapshot 已安装、核心 doctor 读回通过,但两个既有扩展 provider(ark-flow-lane-agent-runtime、ark-flow-lane-delivery)处于 provider_unavailable,于是这一次升级既没有重启服务,还把 next action 报成 review_or_rollback——而 runtime 其实已经装好且健康。操作者可见的后果是本地 steward chat 继续用上一个 release 的打包前端(index-DscLlBrK.js,没有 steward execution chip),直到手动重启。也就是说"安装成功却静默没有开始服务",并且被一条误导性的回滚建议掩盖。更小的修法存在且本 PR 走的就是它:只把"已安装的 runtime 如何开始服务"这层独立出来,不去重写安装/读回/扩展健康的语义。

改动思路

  • 拆开两个就绪概念:runtime_ready = install + 核心 doctor(pip/pipx 路径额外要求 workflow_skills/slash_commands 这两个 host-material 步骤),扩展健康单独成 extension_ready;重启跟随前者,不跟随可选 provider 的健康。
  • 新的 bounded owner:loopx/runtime_activation.py 承载 restart_managed_loopx_services()(从 self_update.py 原样搬移)与新的 restart_services_for_runtime_activation(changes_applied, runtime_ready),返回 {restarted_services, restart_status}。这样 update 边界继续只描述"装了什么","怎么开始服务"归新模块;同时 self_update.py 从 1480 行落回模块预算内(现 1485 < 1500),不需要例外。
  • 结果语义:ok 仍然是 False(被阻塞的扩展依旧是可见缺陷),但 next_action.kind 从 review_or_rollback 换成非变更的 repair_blocked_extensions,不再建议回滚一个已经健康的 runtime;两条执行驱动(archive snapshot 与 pip/pipx)共用同一条规则。
  • 失败关闭:只有 changes_applied and runtime_ready 才重启,否则 restart_status: skipped_runtime_not_activated,绝不"先把服务打起来再说"。
  • 正例路径:apply → install + doctor 通过 → restart_services_for_runtime_activation → execution.restarted_services/restart_status 写入 → 扩展被阻塞时给出 repair 动作。

具体改动

关键代码讲解

  1. loopx/runtime_activation.py:52-74(新增):restart_services_for_runtime_activation 是唯一的激活判定入口,changes_applied and runtime_ready 为真才调用重启,否则返回 skipped_runtime_not_activated;两个返回值都由 execution.update(...) 写回,调用方不需要再判断。
  2. loopx/runtime_activation.py:17-49:restart_managed_loopx_services() 为原样搬移的 macOS 行为(只匹配 stem 含 loopx/goal-harness 且以 .status/.chat 结尾的 plist,逐 label launchctl kickstart -k,非 darwin 直接返回空列表),语义未变。
  3. loopx/self_update.py:1126-1160(archive 路径):runtime_ready = install_returncode == 0 and doctor_returncode == 0;ok 仍要求 extension doctor 为 0;重启交给新 helper;新增 if not updated["ok"] and runtime_ready: 提前返回 repair_blocked_extensions,把原先的 review_or_rollback 留给真正的运行时失败。
  4. loopx/self_update.py:957-1009(pip/pipx 路径):runtime_steps = (install, workflow_skills, slash_commands, doctor) 与 extension_ready 分离,ok = runtime_ready and extension_ready,并新增 elif runtime_ready: 分支给出同一个 repair 动作与命令。
  5. tests/test_self_update_runtime_activation.py:386-456:新增两条 archive 路径回归(扩展被阻塞 → restarted_services 有值 + repair_blocked_extensions + 文案不含 rollback;安装失败 → 不重启 + skipped_runtime_not_activated + review_or_rollback),并把所有 patch 目标改到 loopx.runtime_activation.*。
  6. loopx/semantics/inventory_v0.json:source_files 1175 → 1176,对应新增模块。
  7. docs/book/{,en/}chapters/appendix-reference.md:双语 runbook 各加一段,说明"运行时安装 + 核心 doctor 通过后就会重启托管服务,被阻塞的扩展只报告待修复"。

实测(head 6cf3cdcb9):pytest tests/test_self_update_runtime_activation.py tests/test_archive_installer_commit_response.py tests/test_doctor_install_freshness.py tests/canary/test_maintainability_ratchet.py -q → 63 passed, 1 skipped;examples/loopx-update-smoke.py → ok;scripts/generate_semantic_inventory.py --check → up to date(head 与 head 合入当前 main 9719dc0d 的 merge ref a2fa2ad36 都是干净);examples/docs-governance-smoke.py → ok。另外我用未提交的探针独立复现了 pip 路径的 repair 分支:5 次 subprocess 全部执行、install+doctor 通过而 extension doctor 返回 1 时得到 ok=False、changes_applied=True、restarted_services=['com.loopx.status']、next_action.kind='repair_blocked_extensions'(即行为正确,只是没有用例守着)。

对主干的风险

  1. P2(非阻塞)机器可见的"失败"信号没有跟着改。 loopx update 的进程退出码是 0 if payload["ok"] else 1(loopx/cli_commands/support_control.py:718),而 ok 在本 PR 之后仍然为 False(这是刻意的);同时 loopx/control_plane/heartbeat/installed_prompt_update.py:157-161,169,210 用 updated["ok"] 推导 runtime_ready,于是"已安装且已在服务"的这次升级 upgrade_complete 仍为 False。本 PR 修掉的是面向人的误导(repair 而不是 rollback),但面向脚本/自动化的信号依旧读作失败——这正是本 PR 动机里的那一类混淆。二选一即可:在 runbook 段落里写清"非零退出 + next_action.kind=repair_blocked_extensions 表示 runtime 已安装并在服务",或让进程级信号跟随 changes_applied and runtime_ready、把被阻塞的扩展单独作为一个字段/动作。我没有按阻塞处理,因为这两处语义在 PR 之前同样如此、本 PR 没有引入回归,但请在本轮明确一种口径。
  2. P3 新分支缺回归保护。 repair_blocked_extensions 与 restart_status 只在 archive 路径被断言(tests/test_self_update_runtime_activation.py:419,421,453),pip 路径新增的 elif runtime_ready 没有用例;行为我已独立验证通过,缺的是防止未来改动的测试。
  3. P3 "best-effort" 并非真的 best-effort。 restart_managed_loopx_services() 的 subprocess.run(..., timeout=30) 没有捕获 subprocess.TimeoutExpired,launchctl 缺失时也不捕获 FileNotFoundError;本 PR 让这段代码在更多路径上执行,一旦某个 label 卡住,整个 update 会抛异常而不是报告部分重启。建议逐 label 捕获并把 restart_status 扩成 partial/failed。
  4. P3 归属说明。 loopx/runtime_activation.py 作为"已安装 runtime 如何开始服务"的单一 owner 我认为合适(模块预算 + 单一变更理由),建议在 docstring 里一句话说明它与 loopx/control_plane/runtime/* 的边界,避免以后被读成第二个 runtime 概念。

默认行为与披露:steward 通道默认仍是 codex、托管/worker 执行仍走凭据解析出的 Turn host,这一点 PR 描述已声明且我在 diff 中未发现相反改动;本次默认行为变化(扩展被阻塞时也会重启服务)已在两份 runbook 镜像披露。无凭证、私有路径、raw 证据或本地绝对路径进入仓库。

我的整体评价

APPROVE。这个修复对准的是真实且已被现场观测到的缺陷(服务没跟着新 release 重启、更新却报"回滚"),拆法也克制:只把"如何让已安装的 runtime 开始服务"独立成一个 bounded owner,安装/读回/扩展健康的语义没有被重写,两条驱动共享同一条规则,失败方向是关闭的(没装好绝不重启),并且 self_update.py 回到了模块预算内。正反两个方向的回归测试都在,清单计数按新模块更新过,head 与"合入当前 main"的 merge ref 都干净。剩下三点都不阻塞合并:把进程级退出码/upgrade_complete 的失败信号口径写清(或改)、补 pip 路径的 repair 用例、让 restart 真正 best-effort(吞掉 launchctl 超时/缺失并报告部分重启)。本轮不做合并动作。

English verdict: APPROVE at 6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b. The change fixes an observed defect rather than a hypothetical one: update apply restarted the LaunchAgent-managed status/chat services only in its fully-clean branch, so a blocked enabled extension provider left status/chat serving the previous release while the reported next action told the operator to roll back an already installed, healthy runtime. The fix separates runtime_ready (install plus core doctor, plus host-material steps on the pip/pipx driver) from extension_ready, restarts on the former via a new bounded owner loopx/runtime_activation.py, keeps ok false so the blocked provider stays visible, and replaces the rollback advice with a non-mutating repair_blocked_extensions action shared by both drivers. Verified at the exact head: 63 passed / 1 skipped across the runtime-activation, archive-installer, doctor-freshness and maintainability-ratchet modules, loopx-update-smoke ok, semantic inventory up to date both at the head and on the merge ref a2fa2ad36 (head into current main 9719dc0d), docs-governance smoke ok, and I independently reproduced the pip-driver repair branch (services restarted, repair_blocked_extensions, changes_applied=True, ok=False) since it has no committed test. Three non-blocking notes: the process-level signal still reads as failure (loopx update exits 1 and upgrade_complete stays false for installed_prompt_update.py because both key on ok) - state the exit-code contract in the runbook or let the process signal follow the runtime state; add a pip-driver regression case; and make the launchctl restart genuinely best-effort by catching its timeout/missing-binary paths instead of letting them propagate. No merge performed in this round.

@huangruiteng
huangruiteng merged commit c62b24f into main Sep 15, 2026
28 checks passed
@huangruiteng
huangruiteng deleted the codex/restart-managed-services-after-archive-activation branch September 15, 2026 15:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant