Skip to content

fix(steward): let the machine setting decide the manager channel's executor - #4510

Merged
huangruiteng merged 1 commit into
mainfrom
codex/steward-connection-executor-observation
Sep 16, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/steward-connection-executor-observation

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

Motivation

A manager (steward) connection stored the executor endpoint as a decision when it was created. The Lark route, the authorized-connection resolution and the Turn that answers on the channel all read that stored value instead of the machine.

The consequence on a real machine: the operator's steward_executor machine configuration selected the managed host (dsh), /api/chat/capabilities reported executor_endpoint=dsh, source=machine_configuration, available=true — and the steward group still answered on codex, failing with a host gate because that endpoint belonged to an exhausted personal subscription. The connection write path had the same shape (the stored value outranked the machine), and no surface could correct it: the Dashboard never sends executor_endpoint_id, and the edit API re-read the stored value.

Change

  • The machine configuration is the one owner of the manager channel's executor.
  • New loopx/chat_manager.py::manager_connection_executor_endpoint(runtime_root) resolves the machine document, then the service environment, then the shipped default — the same precedence the capability readback already publishes.
  • decide_manager_event, the Lark event decision and answer_lark_goal_topic re-resolve through it, so a record written under an earlier default cannot decide a later Turn. authorized_manager_goal_ids compares the bound Session against the same resolution.
  • A connection write records the machine's current resolution as an observation (executor_endpoint_id plus the new executor_endpoint_source) and refuses a request that asks for a different endpoint, naming the owner, instead of writing a value the route would have to ignore.
  • When a machine changes its selection, the channel's already-bound Session runs on the earlier endpoint. That Turn is now refused with the typed manager_channel_executor_rebind_required receipt, and the reply names the single repairing action (re-apply the connection) instead of the generic "manager failed" label.

Scope and delivery

Backend only. The operator-facing control for this setting already exists and is unchanged: the Dashboard machine-configuration editor renders the steward_executor namespace fields (machine-configuration-settings.tsx + capability-localization.ts, field executor_endpoint), with preview, apply and readback. Because the connection record no longer decides, that editor is now sufficient to move the steward channel; no packaged-frontend asset changes, and therefore no npm run build:chat parity rebuild.

Validation

  • pytest tests/extensions/test_lark_goal_topic_connections.py tests/extensions/test_lark_goal_topic_runtime.py tests/test_manager_channel_binding.py tests/test_manager_context_handoff.py tests/test_chat_lark_api_contract.py → 182 passed, 1 failed.
  • The single failure, test_every_production_steward_caller_passes_the_machine_defaults, is pre-existing: it asserts loopx/chat_server.py still holds a steward resolver call, and that call was extracted earlier. It reproduces identically on a clean origin/main checkout (verified by stashing the whole diff).
  • Four new cases: route precedence over a stale connection record, the write record and its refusal of an override, authorized-session matching against the machine endpoint, and the typed rebind reply through manager_failure_reply.
  • Smokes: examples/loopx-steward-channel-binding-smoke.py, examples/loopx-steward-managed-chat-smoke.py, examples/loopx-managed-turn-operator-flow-smoke.py all pass.
  • loopx canary premerge --from-git-diff --git-diff-base origin/main --tier standard → 8 checks executed, 0 failures; public boundary scan clean.

Live evidence

On the affected machine the same Telegram-free repair path (re-applying the manager connection via the edit API with the machine's endpoint restated) moved the channel Session from codex to dsh, and a real Turn then completed in 25s on deepseek-v4-flash@high. This PR removes the possibility of that drift returning.

Risk

Behaviour change, disclosed: a manager connection can no longer pin a different endpoint than the machine configuration, and such a request now fails with a typed error instead of silently writing a value the route ignores. Every affected connection would today be answering on an endpoint the readback does not report. Unconfigured machines keep resolving exactly as before (shipped default codex).

…ecutor

A manager connection stored the executor endpoint as a decision when it was
created, and the Lark route, the authorized-connection resolution and the
answering Turn all read that record instead of the machine. A machine that
later selected another steward executor therefore kept answering on the
endpoint that was the default on the day of the connection: its capability
readback reported the managed host while the channel still ran, and failed,
on the interactive CLI endpoint. The connection write path had the same shape
and no surface to change it, so the value could not be corrected at all.

The machine configuration is now the one owner of that choice. The connection
record keeps the resolved endpoint as an observation with its source; every
read path re-resolves through the new
`manager_connection_executor_endpoint` owner, and a connection write records
the machine's current resolution while refusing a request that tries to
override it. A Session left behind by a machine that changed its selection is
refused with the typed `manager_channel_executor_rebind_required` receipt, and
the reply names the one action that repairs it instead of the generic manager
failure.

Verified: the changed Lark, manager-channel, handoff and Lark-API suites pass
(182 passed), including four new cases covering route precedence, the write
record and refusal, authorized-session matching and the typed rebind reply; the
steward channel-binding, steward managed-chat and managed-turn operator-flow
smokes pass. The pre-existing failure of
`test_every_production_steward_caller_passes_the_machine_defaults` reproduces
unchanged on origin/main.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head: bdca4e33cd9873b8301184d38c14c6941d109dfb (12 files, +278/-25).

动机

管家通道的执行器有两个所有者:machine configuration 的 steward_executor 命名空间,以及创建连接时写进 binding.routing.executor_endpoint_id 的那份值。读取路径只认后者,于是"本机选了 dsh"这件事对实际应答没有任何约束力。在本机上的真实表现正是如此:/api/chat/capabilities 报 executor_endpoint=dsh、source=machine_configuration、available=true,而管家群仍走 codex,并因为该端点属于已耗尽额度的个人订阅而返回 host_gate 失败。写入路径同形,且没有任何界面能改它——Dashboard 从不发送 executor_endpoint_id,编辑接口又会回读旧值,所以这个错误值无法被修正。

改动思路

删掉第二个决定,而不是加一层对账。machine configuration 成为唯一所有者;连接记录保留解析结果,但把它降级为观测值(连同新的 executor_endpoint_source)。所有读取路径统一走同一条解析,写入路径记录当前解析、并拒绝试图覆盖它的请求(点名所有者)。本机确实改了选择时,通道上已绑定的 Session 仍跑在旧端点——这一种情况现在给出类型化的 manager_channel_executor_rebind_required,回复直接写出唯一能修复它的动作,而不是笼统的"管家处理失败"。

具体改动

  • loopx/chat_manager.py:新增 manager_connection_executor_endpoint(runtime_root),复用既有的 load_effective_steward_executor_defaults 与 selected_manager_executor_endpoint,保持"本机文档 → LOOPX_MANAGER_ENDPOINT → 出货默认"这一既有优先级。
  • loopx/extensions/lark/manager_routing.py:decide_manager_event 的路由端点改为该解析结果并附 executor_endpoint_source;authorized_manager_goal_ids 用同一解析与绑定 Session 比对。
  • loopx/extensions/lark/goal_topic_edit.py:resolve_conversation_policy 对 manager 返回机器解析与来源,并拒绝与其不一致的显式请求。
  • loopx/extensions/lark/goal_topic_connections.py / goal_topic_batch.py / chat_lark_api.py:写入路径透传 runtime_root,记录 executor_endpoint_id + executor_endpoint_source;manager 连接的端点一律由机器解析决定。
  • loopx/extensions/lark/goal_topic_runtime.py:answer_lark_goal_topic 对 manager 不再读路由上的旧值;仅"同一受众、执行器不同"的失配给出类型化回执,其它失配(受众/Goal/已关闭)维持原有 RuntimeError,避免误报。
  • loopx/extensions/lark/manager_context.py:为新 code 增加用户可见文案。
  • docs/architecture/rfcs/harness-selection-dsh-pi-v0.md / .zh-CN.md:写入该不变量(连接记录是观测值;改选择后需重新应用一次连接)。
  • 测试:新增 4 例——陈旧记录下路由仍走机器选择、写入记录并拒绝覆盖、授权会话必须与机器端点一致、陈旧会话得到类型化重绑回复。

对主干的风险

已披露的行为变更有两处,且都是本 PR 的目的本身:连接记录不再能压过 machine configuration;试图用连接级 executor_endpoint_id 覆盖机器选择的请求现在报错(invalid_lark_connection,点名所有者),而不是静默写入一个路由会忽略的值。未配置该命名空间的机器解析与行为完全不变(codex / product_default,由既有 shipped-default 测试与 steward channel-binding smoke 守住)。worker(goal)连接不受影响,仍用 agent_id。前端无需改动:可编辑该设置的面已经存在(machine-configuration-settings.tsx 渲染 steward_executor 命名空间字段,capability-localization.ts 中 executor_endpoint 即"管家执行器"),此前只是不生效;因此没有打包资产变动,也不需要 npm run build:chat 对齐重建。

我的整体评价

无阻断性问题,建议合并。这是一次删除重复决定的重构,而不是新增抽象:改动集中在唯一的渠道所有者与其读取路径,量级与问题相称。验证取自被改动的 Lark、manager-channel、handoff、Lark-API 套件(182 passed),以及三个 steward/managed smoke 与 canary premerge --tier standard(8/8 检查、公开边界扫描干净);唯一的失败 test_every_production_steward_caller_passes_the_machine_defaults 在干净 origin/main 上同样失败(已通过 stash 整个 diff 复现),属于既有的过期不变量测试,不在本 PR 范围内。线上证据:在受影响机器上把管家连接重新应用后,通道 Session 由 codex 迁到 dsh,真实 Turn 在 25s 内以 deepseek-v4-flash@high 完成。

残留风险与最强缺口:本 PR 之后,改了机器执行器的机器仍需手工重新应用一次连接才能让已绑定的 Session 跟上(现阶段靠类型化回执提示,而非自动重绑——因为读取路径不能写);另外 test_every_production_steward_caller_passes_the_machine_defaults 仍红在 main,应单独修掉,否则会持续掩盖真正的调用点回归。

English verdict: APPROVE - exact head bdca4e33c. The steward channel's executor now has one owner (the machine configuration); connection records store an observation, all read paths re-resolve through manager_connection_executor_endpoint, a write that tries to override the machine choice fails with the owner named, and a Session stranded on the previous endpoint is refused with the typed manager_channel_executor_rebind_required receipt instead of the generic manager failure. Validated by 182 passing tests in the changed suites plus four new cases, three steward/managed smokes, a clean canary premerge --tier standard run, and a live Turn on the repaired machine; the single suite failure is pre-existing on origin/main.

@huangruiteng
huangruiteng merged commit 0346a31 into main Sep 16, 2026
20 of 22 checks passed
@huangruiteng
huangruiteng deleted the codex/steward-connection-executor-observation branch September 16, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant