fix(update): activate managed services when only extensions are blocked - #4458
Conversation
`loopx update apply` restarted the LaunchAgent-managed status/chat services only in its fully-clean branch. A blocked enabled extension provider made the run report `review_or_rollback` even though the release snapshot was installed and core doctor readback passed, so status/chat kept serving the previous release until an operator restarted them by hand. Decide the restart from the runtime install plus its core doctor readback, report blocked providers as a repair action instead of a rollback, and move the LaunchAgent restart into a bounded runtime_activation owner so the update boundary keeps describing what was installed rather than how it starts serving. The new archive-path cases carry the same native-Windows skip as the existing archive update test, because `execute_update_plan` fails closed before the installer on that platform. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
c451e52 to
6cf3cdc
Compare
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
精确 head:6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b(取代 c451e52e8b36185c13460c50114bcfb1efc663d4,差异只是测试的平台守卫)。评审按 pull_request_review_execution_contract_v2(policy_revision=3)执行,pr-review --check-result 对该 head 返回 approval_consistent=true、无 blocker。
动机
loopx update apply 会替换已安装的 release snapshot,但 LaunchAgent 托管的 status/chat 服务仍然跑着启动时的旧代码。旧实现只在“所有步骤全绿”的分支里重启服务,而 extension doctor --all-enabled --execute 只要有一个已启用的 provider 被阻塞就会返回非零,于是更新直接落到 review_or_rollback 分支:既没有重启,又建议操作者回滚一个其实已经装好、核心 doctor 也读通的运行时。
这是本机真实发生的路径:install rc 0、核心 doctor rc 0、release snapshot 已提升,但因为两个既有的 provider 阻塞(ark-flow-lane-agent-runtime、ark-flow-lane-delivery,provider_unavailable/probe_nonzero_exit),服务的重启被跳过。操作者看到的是“机器报告了新 release,管家聊天却继续吐旧版本打包前端”,必须手动 launchctl kickstart 才能对齐;而这些 provider 的健康与“能不能把服务切到新 release”本来没有因果关系,属于把两件事绑在一起。
影响面是本机操作者加上所有 chat/dashboard 消费者;成本是静默的版本错配(报告的身份 ≠ 实际服务的身份),恢复方式只有手动重启或走回滚。更小的修法试过两条:只放宽 ok 计算,重启仍然留在全绿分支里,观察到的结果不变;安装返回 0 就无条件重启,则会把服务切到一个核心 doctor 读不通的 release 上。以“安装 + 核心 doctor 读回”作为切换门槛,是能覆盖该问题的 B/A 折中。
非目标:不改管家 chat 通道默认(仍是 codex)、不改由凭据解析的 managed Turn host、不修复那两个 provider、不动 installer 下载与 release 提升规则。
改动思路
入口是 loopx update apply → loopx.self_update.execute_update_plan(归档快照驱动)与 _execute_python_distribution_update(pip/pipx 驱动)。权威输入是 install_lifecycle.execution_driver、plan.backup.current_release_root,以及 install / 核心 doctor / extension doctor 三个返回码。改动把原来一个布尔拆成两个派生事实:runtime_ready = install rc == 0 and doctor rc == 0 决定“能不能切换”,ok = runtime_ready and 已启用扩展复检通过 决定“这次更新干不干净”。
生效副作用(launchctl kickstart -k)与判定逻辑被抽到 loopx/runtime_activation.py:restart_managed_loopx_services() 从 loopx/self_update.py 原样搬过去,新增 restart_services_for_runtime_activation() 作为唯一规则所有者,两个驱动都调它。原先那份模块内定义被删除而不是包装保留,仓内其他引用只有测试模块,已同 head 更新。抽出的直接原因是:修复本身会把 loopx/self_update.py 从基线 1480 行推到 1543 行,超过 1500 行模块预算;抽出后是 1485 行,等于把债务消掉而不是登记例外。
具体改动
六个文件 +212/-52。分类:生产代码 loopx/self_update.py(净 +5 行)与新增 loopx/runtime_activation.py(74 行,主体是搬移);测试 tests/test_self_update_runtime_activation.py(+84/-6);公开文档两篇 docs/book/en/chapters/appendix-reference.md 与 docs/book/chapters/appendix-reference.md 各一段;生成物 loopx/semantics/inventory_v0.json(source_files 1175 → 1176,由 scripts/generate_semantic_inventory.py 重新生成)。
关键代码讲解
loopx/runtime_activation.py:52restart_services_for_runtime_activation:唯一判定点。仅当changes_applied and runtime_ready时才调用重启,返回{restarted_services, restart_status};否则返回空列表加restart_status=skipped_runtime_not_activated。扩展 provider 健康被明确排除在这个判定之外。loopx/self_update.py:1030execute_update_plan:runtime_ready由 install/doctor 返回码构成,ok再叠加扩展复检;installed 但扩展被阻塞这条路径现在返回非阻塞的next_action.kind=repair_blocked_extensions(mutating=false,命令是操作者自己的loopx extension doctor --all-enabled --execute),并保留ok=false让阻塞可见。install 或核心 doctor 失败时仍走原review_or_rollback与回滚命令。loopx/self_update.py:861_execute_python_distribution_update:同一规则;差异是runtime_steps里仍要求workflow-skills、slash-commands主机材料步骤通过,避免 PyPI 更新把服务切到技能/命令未安装的环境。loopx/runtime_activation.py:17restart_managed_loopx_services:原样搬移的 LaunchAgent 选择与重启逻辑(仅匹配 loopx/goal-harness 且以.status/.chat结尾的 plist,逐条 30s 超时,best-effort)。公开语义未变。
对主干的风险
最强回归场景是“install rc 0 但核心 doctor 读回失败”——新门槛若判错,会把服务切到坏 release。防线是 runtime_ready 同时要求两者为 0 返回码,且负向用例 test_archive_apply_skips_managed_service_restart_when_the_runtime_is_unhealthy 断言此时不重启、并保留 review_or_rollback 与回滚命令。爆炸半径仅限本机用户级服务(com.loopx.chat、com.loopx.status 及同类 plist),不触碰 release 提升、quota、todo 与 registry 状态;可观测性上 restarted_services/restart_status 在所有分支都会写出并渲染进 markdown 计划。
真正值得记录的是上一 head 的反例:Windows CI 会原生跑这个测试模块,而 execute_update_plan 在 os.name == 'nt' 下会先返回 unsupported_platform,导致两个新的归档路径用例在那里失败。当前 head 给这两个用例补了与该文件既有归档用例一致的 native-Windows skip,windows-powershell 已通过,同时 test_windows_execute_update_fails_closed_without_launching_bash 继续钉住 Windows 的 fail-closed 行为。校验覆盖:本地 pytest 63 passed/1 skipped、loopx canary premerge --from-git-diff --timeout-seconds 300 14/14 通过(默认 120s 会在本机把 install-local-smoke 判超时,单独跑约 157s 通过)、loopx-update-smoke、maintainability ratchet(unreviewed 0)、semantic inventory drift、两篇文档 smoke,以及该 head 全部 24 项非 skip 的 CI 检查。
scope_fit:真实生产调用点存在并在本机端到端跑过;change_proportionality:机制成本是一个 74 行、以搬移为主、零新增持久化状态与 CLI 选项的模块,对照的是一个每次更新都可能静默复现的版本错配。
我的整体评价
基线/head 对比:全绿路径不变;install ok + doctor ok + 扩展阻塞 从“ok=false、不重启、建议回滚”变成“ok=false、记录重启、给出修复动作”;install/doctor 失败路径在两个版本都不重启且保留回滚。等价性风险集中在新增的 next_action.kind=repair_blocked_extensions 消费方,未知 kind 在现渲染里按文本处理,属可接受。
结论:APPROVE。没有阻塞项,按 P2 留下的残余风险是 Windows 激活分支刻意不可达因而未被真实覆盖,以及 launchctl 语义靠未改动的选择用例加一次真实本机 update apply 验证。若下一步要更硬的证据,值得补一条 CI 级集成:在模拟主机上让 extension doctor 返回非零,断言 restarted_services 被记录。
English verdict: APPROVE — exact head 6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b. The update driver now gates the managed-service restart on the runtime install plus core doctor readback, so a blocked enabled extension provider no longer leaves status/chat silently serving the previous release, and the outcome reports a non-mutating repair_blocked_extensions step instead of a rollback of a healthy runtime. Evidence: targeted and negative unit cases for both drivers, loopx canary premerge --from-git-diff 14/14 on the changed surfaces, loopx-update-smoke, maintainability ratchet and semantic inventory drift green, and all 24 non-skipped CI checks passing on this head (the superseded head failed the native Windows job, which the exact-head platform guards fix). Key finding: none blocking; residual risk is the intentionally unreachable Windows activation branch and launchctl semantics verified by an unchanged selection test plus a real local upgrade rather than a CI integration test.
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
动机
loopx update apply 只在"全部干净"的分支里重启 LaunchAgent 托管的 status/chat,而"全部干净"包含了 enabled 扩展健康。现场结果是:release snapshot 已安装、核心 doctor 读回通过,但两个既有扩展 provider(ark-flow-lane-agent-runtime、ark-flow-lane-delivery)处于 provider_unavailable,于是这一次升级既没有重启服务,还把 next action 报成 review_or_rollback——而 runtime 其实已经装好且健康。操作者可见的后果是本地 steward chat 继续用上一个 release 的打包前端(index-DscLlBrK.js,没有 steward execution chip),直到手动重启。也就是说"安装成功却静默没有开始服务",并且被一条误导性的回滚建议掩盖。更小的修法存在且本 PR 走的就是它:只把"已安装的 runtime 如何开始服务"这层独立出来,不去重写安装/读回/扩展健康的语义。
改动思路
- 拆开两个就绪概念:
runtime_ready = install + 核心 doctor(pip/pipx 路径额外要求workflow_skills/slash_commands这两个 host-material 步骤),扩展健康单独成extension_ready;重启跟随前者,不跟随可选 provider 的健康。 - 新的 bounded owner:
loopx/runtime_activation.py承载restart_managed_loopx_services()(从self_update.py原样搬移)与新的restart_services_for_runtime_activation(changes_applied, runtime_ready),返回{restarted_services, restart_status}。这样 update 边界继续只描述"装了什么","怎么开始服务"归新模块;同时self_update.py从 1480 行落回模块预算内(现 1485 < 1500),不需要例外。 - 结果语义:
ok仍然是 False(被阻塞的扩展依旧是可见缺陷),但next_action.kind从review_or_rollback换成非变更的repair_blocked_extensions,不再建议回滚一个已经健康的 runtime;两条执行驱动(archive snapshot 与 pip/pipx)共用同一条规则。 - 失败关闭:只有
changes_applied and runtime_ready才重启,否则restart_status: skipped_runtime_not_activated,绝不"先把服务打起来再说"。 - 正例路径:apply → install + doctor 通过 →
restart_services_for_runtime_activation→execution.restarted_services/restart_status写入 → 扩展被阻塞时给出 repair 动作。
具体改动
关键代码讲解
loopx/runtime_activation.py:52-74(新增):restart_services_for_runtime_activation是唯一的激活判定入口,changes_applied and runtime_ready为真才调用重启,否则返回skipped_runtime_not_activated;两个返回值都由execution.update(...)写回,调用方不需要再判断。loopx/runtime_activation.py:17-49:restart_managed_loopx_services()为原样搬移的 macOS 行为(只匹配 stem 含loopx/goal-harness且以.status/.chat结尾的 plist,逐 labellaunchctl kickstart -k,非 darwin 直接返回空列表),语义未变。loopx/self_update.py:1126-1160(archive 路径):runtime_ready = install_returncode == 0 and doctor_returncode == 0;ok仍要求 extension doctor 为 0;重启交给新 helper;新增if not updated["ok"] and runtime_ready:提前返回repair_blocked_extensions,把原先的review_or_rollback留给真正的运行时失败。loopx/self_update.py:957-1009(pip/pipx 路径):runtime_steps = (install, workflow_skills, slash_commands, doctor)与extension_ready分离,ok = runtime_ready and extension_ready,并新增elif runtime_ready:分支给出同一个 repair 动作与命令。tests/test_self_update_runtime_activation.py:386-456:新增两条 archive 路径回归(扩展被阻塞 →restarted_services有值 +repair_blocked_extensions+ 文案不含 rollback;安装失败 → 不重启 +skipped_runtime_not_activated+review_or_rollback),并把所有 patch 目标改到loopx.runtime_activation.*。loopx/semantics/inventory_v0.json:source_files1175 → 1176,对应新增模块。docs/book/{,en/}chapters/appendix-reference.md:双语 runbook 各加一段,说明"运行时安装 + 核心 doctor 通过后就会重启托管服务,被阻塞的扩展只报告待修复"。
实测(head 6cf3cdcb9):pytest tests/test_self_update_runtime_activation.py tests/test_archive_installer_commit_response.py tests/test_doctor_install_freshness.py tests/canary/test_maintainability_ratchet.py -q → 63 passed, 1 skipped;examples/loopx-update-smoke.py → ok;scripts/generate_semantic_inventory.py --check → up to date(head 与 head 合入当前 main 9719dc0d 的 merge ref a2fa2ad36 都是干净);examples/docs-governance-smoke.py → ok。另外我用未提交的探针独立复现了 pip 路径的 repair 分支:5 次 subprocess 全部执行、install+doctor 通过而 extension doctor 返回 1 时得到 ok=False、changes_applied=True、restarted_services=['com.loopx.status']、next_action.kind='repair_blocked_extensions'(即行为正确,只是没有用例守着)。
对主干的风险
- P2(非阻塞)机器可见的"失败"信号没有跟着改。
loopx update的进程退出码是0 if payload["ok"] else 1(loopx/cli_commands/support_control.py:718),而ok在本 PR 之后仍然为 False(这是刻意的);同时loopx/control_plane/heartbeat/installed_prompt_update.py:157-161,169,210用updated["ok"]推导runtime_ready,于是"已安装且已在服务"的这次升级upgrade_complete仍为 False。本 PR 修掉的是面向人的误导(repair 而不是 rollback),但面向脚本/自动化的信号依旧读作失败——这正是本 PR 动机里的那一类混淆。二选一即可:在 runbook 段落里写清"非零退出 +next_action.kind=repair_blocked_extensions表示 runtime 已安装并在服务",或让进程级信号跟随changes_applied and runtime_ready、把被阻塞的扩展单独作为一个字段/动作。我没有按阻塞处理,因为这两处语义在 PR 之前同样如此、本 PR 没有引入回归,但请在本轮明确一种口径。 - P3 新分支缺回归保护。
repair_blocked_extensions与restart_status只在 archive 路径被断言(tests/test_self_update_runtime_activation.py:419,421,453),pip 路径新增的elif runtime_ready没有用例;行为我已独立验证通过,缺的是防止未来改动的测试。 - P3 "best-effort" 并非真的 best-effort。
restart_managed_loopx_services()的subprocess.run(..., timeout=30)没有捕获subprocess.TimeoutExpired,launchctl缺失时也不捕获FileNotFoundError;本 PR 让这段代码在更多路径上执行,一旦某个 label 卡住,整个 update 会抛异常而不是报告部分重启。建议逐 label 捕获并把restart_status扩成 partial/failed。 - P3 归属说明。
loopx/runtime_activation.py作为"已安装 runtime 如何开始服务"的单一 owner 我认为合适(模块预算 + 单一变更理由),建议在 docstring 里一句话说明它与loopx/control_plane/runtime/*的边界,避免以后被读成第二个 runtime 概念。
默认行为与披露:steward 通道默认仍是 codex、托管/worker 执行仍走凭据解析出的 Turn host,这一点 PR 描述已声明且我在 diff 中未发现相反改动;本次默认行为变化(扩展被阻塞时也会重启服务)已在两份 runbook 镜像披露。无凭证、私有路径、raw 证据或本地绝对路径进入仓库。
我的整体评价
APPROVE。这个修复对准的是真实且已被现场观测到的缺陷(服务没跟着新 release 重启、更新却报"回滚"),拆法也克制:只把"如何让已安装的 runtime 开始服务"独立成一个 bounded owner,安装/读回/扩展健康的语义没有被重写,两条驱动共享同一条规则,失败方向是关闭的(没装好绝不重启),并且 self_update.py 回到了模块预算内。正反两个方向的回归测试都在,清单计数按新模块更新过,head 与"合入当前 main"的 merge ref 都干净。剩下三点都不阻塞合并:把进程级退出码/upgrade_complete 的失败信号口径写清(或改)、补 pip 路径的 repair 用例、让 restart 真正 best-effort(吞掉 launchctl 超时/缺失并报告部分重启)。本轮不做合并动作。
English verdict: APPROVE at 6cf3cdcb9948f0b50906b6291c72c8cc694c3d5b. The change fixes an observed defect rather than a hypothetical one: update apply restarted the LaunchAgent-managed status/chat services only in its fully-clean branch, so a blocked enabled extension provider left status/chat serving the previous release while the reported next action told the operator to roll back an already installed, healthy runtime. The fix separates runtime_ready (install plus core doctor, plus host-material steps on the pip/pipx driver) from extension_ready, restarts on the former via a new bounded owner loopx/runtime_activation.py, keeps ok false so the blocked provider stays visible, and replaces the rollback advice with a non-mutating repair_blocked_extensions action shared by both drivers. Verified at the exact head: 63 passed / 1 skipped across the runtime-activation, archive-installer, doctor-freshness and maintainability-ratchet modules, loopx-update-smoke ok, semantic inventory up to date both at the head and on the merge ref a2fa2ad36 (head into current main 9719dc0d), docs-governance smoke ok, and I independently reproduced the pip-driver repair branch (services restarted, repair_blocked_extensions, changes_applied=True, ok=False) since it has no committed test. Three non-blocking notes: the process-level signal still reads as failure (loopx update exits 1 and upgrade_complete stays false for installed_prompt_update.py because both key on ok) - state the exit-code contract in the runbook or let the process signal follow the runtime state; add a pip-driver regression case; and make the launchctl restart genuinely best-effort by catching its timeout/missing-binary paths instead of letting them propagate. No merge performed in this round.
Why
loopx update applyreplaced the installed release snapshot but restarted theLaunchAgent-managed
status/chatservices only inside its fully-clean branch.On this machine a real
update applyproduced:installrc 0, coredoctorrc 0, release snapshot promoted;extension doctor --all-enabled --executerc 1 because two pre-existingenabled providers (
ark-flow-lane-agent-runtime,ark-flow-lane-delivery)were
provider_unavailable/probe_nonzero_exit;restarted_serviceswas never attempted, and the next action toldthe operator to
review_or_rollbacka runtime that was already installed andhealthy.
The operator-visible effect: the local steward chat kept serving the previous
release's packaged frontend (
index-DscLlBrK.js, no steward execution chip)until the services were restarted by hand. Activation silently depended on the
health of unrelated extension providers.
What changed
doctorreadback. Blocked enabled extension providers no longer prevent activation.
reports
repair_blocked_extensions(non-mutating) instead ofreview_or_rollback, and no longer recommends rolling back a healthy runtime.okstaysfalseso the blocked provider remains a visible defect.restart_managed_loopx_services()and the new activation decision moved toloopx/runtime_activation.py, a bounded owner for "how an installed runtimestarts serving".
loopx/self_update.pywas at 1480 lines against a 1500-linemodule budget and now stays inside it (1485) instead of needing an exception.
loopx/semantics/inventory_v0.jsonregeneration for the new module.Both execution drivers (archive snapshot and pip/pipx) now share one rule; the
pip path keeps requiring its host-material steps (
workflow-skills,slash-commands) before activation.Validation
loopx canary premerge --from-git-diff --timeout-seconds 300:selected=14 executed=14 failures=0, public boundary passed. (The default120s per-check budget times out
examples/install-local-smoke.py, which needs~157s on this host; it passes standalone.)
pytest tests/test_self_update_runtime_activation.py tests/test_archive_installer_commit_response.py tests/test_doctor_install_freshness.py tests/canary/test_maintainability_ratchet.py -q→ 63 passed, 1 skipped.examples/loopx-update-smoke.py,examples/control_plane/control-plane-maintainability-ratchet-smoke.py,examples/semantic-vocabulary-drift-smoke.py,examples/docs-governance-smoke.py,examples/dev-book-publication-smoke.py→ ok.restarted + repair action; runtime unhealthy → no restart + rollback action.
http://127.0.0.1:8767/api/chat/capabilitiesnow reports
runtime_identity.release_id=20260915T144512Z,source_revision=9719dc0d4, and the steward channel bindingcodex · individual · gpt-6-astra.Notes
codex; managed/worker execution stays thecredential-resolved Turn host. This PR does not change either default.
the activation semantics are stated where operators read them.