feat(chat): select the steward channel executor explicitly - #4446
Conversation
The steward channel now selects its executor and its model, and a configured operator credential re-points neither. The previous rule resolved the channel onto the managed host when DEEPSEEK_API_KEY was present while the endpoint stayed codex, because dsh has no interactive Chat transport -- so the executor and the model disagreed, and the swapped model was handed to the Codex adapter. - manager_channel_binding resolves the endpoint from explicit configuration only; `codex` is the shipped default and LOOPX_MANAGER_ENDPOINT re-points it. - The steward model follows the selected executor, so the shipped CLI endpoint keeps the vendor default and manager_model_config no longer reads the credential. - The managed Turn host fails closed as the typed managed_host_chat_transport_unsupported host-tool gate instead of an unknown endpoint ValueError, in the Chat service and Lark routing. - The three hardcoded endpoint fallbacks now resolve through the one function that owns the rule, and Chat capabilities carry the binding for frontend readback. Replaces the steward half of the stacked chain (#4417). Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…n gates The RFC still described the managed host as resolved from the operator credential. It now records what the repository ships and what the managed stack will ship: the Turn host is selected, never inferred (shipped default `dsh`, `LOOPX_TURN_HOST` re-points it, an explicit `--host` wins), the steward channel selects its own executor (shipped default `codex`, because it is the only interactive Chat transport today), and a configured credential only authenticates the selection instead of changing it. - The role table gains the steward channel executor row and states each selection source, including the `individual` versus `managed` executor kind. - The steward-channel section replaces the credential-resolved binding rule and records the promotion gate for the managed host (an interactive Chat transport), so the credential is explicitly not the promotion signal. - The milestone table records the steward channel's M1 contract as selected -- endpoint, model, source and executor kind -- and states the session-identity gap that still blocks a channel answer from proving which session served it. - Evidence and dsh-pin rows are re-pointed at the replacement PRs (#4443, #4446) and the dsh pin PR #4420. Docs only: no runtime is promoted, no default behavior changes here, and the English and Chinese mirrors are updated together. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
详细中文评审
审查对象:91c440ca021dd2f7378f29bb539db7f305555c4f(feat(chat): select the steward channel executor explicitly,base main,12 文件 +710/-45)。执行契约 policy_revision=3;结果 JSON 已通过 pr-review --check-result。
动机
旧规则让凭据决定选择:只要 DEEPSEEK_API_KEY 存在,管家通道就把模型换成 deepseek-flash(来源标记 operator_credential_default)。但端点仍然是 codex,因为 dsh 没有交互式 Chat 传输。于是这个模型名被交给 Codex adapter —— 执行器和模型互相矛盾,而运营方无法从配置里判断这是「我选的」还是「环境碰巧给的」。本 PR 把两侧都改成显式选择:端点默认 codex、LOOPX_MANAGER_ENDPOINT 可改指、模型跟随所选端点(默认 gpt-6-astra),凭据只作为事实被回报,并为「托管宿主无法承载会话」给出 typed gate。
改动思路
拆分两个 owner:选择归 loopx/chat_manager.py(纯函数 selected_manager_executor_endpoint),认证事实归已被 Turn 侧引入的 loopx/control_plane/operator_credential.py;能否承载会话归 loopx/chat_agent.py 的既有 managed-host 判据。三个原先硬编码 "codex" 的兜底(Lark chat API、goal topic runtime、open_manager_session)改为调用同一解析函数,消除三处重复判据。可用性只在能证明不可用时为 False,否则为 None(不做无证据的 True 断言),与 Turn 侧 managed_executor_binding 的语义保持一致。
具体改动
selected_manager_executor_endpoint/manager_executor_endpoint_default(loopx/chat_manager.py:134):仅从LOOPX_MANAGER_ENDPOINT或产品默认取值;空白输入不覆盖默认值;凭据完全不参与。manager_channel_binding(loopx/chat_manager.py:152):输出executor_endpoint、executor_endpoint_source、executor_kind(individual/managed,与 Turn 侧同词表)、model、model_source、credential_env_var、operator_credential_configured、available、unavailable_reason。凭据只以变量名出现。agent_endpoint_error(loopx/chat_agent.py:50):managed Turn 宿主抛CodexChatAgentError(error_code=managed_host_chat_transport_unsupported)并带host_tool_gate(next_action 指向改用可承载端点或loopx turn);其他未知 id 保留原有无类型ValueError。chat_runtime.py:527改用它,chat_lark_api.py捕获后返回 400 + gate。manager_model_config(loopx/chat_manager.py:244):参数化为可注入environ,默认值不再读取凭据;LOOPX_MANAGER_MODEL/LOOPX_MANAGER_REASONING_EFFORT优先级不变。manager_runtime_capability_projection(machine_profile.py):新增可选channel_binding透传,由调用方(chat_server.py:1272)提供,投影层不重新推导规则。- 新增
tests/test_manager_channel_binding.py(12 用例)与examples/loopx-steward-channel-binding-smoke.py;文档在docs/reference/protocols/manager-evidence-and-continuity-v0.md增加Steward channel host selection。
对主干的风险
- 默认行为变化已披露:有凭据时模型不再变成
deepseek-flash。若有人把「配了 key 就换模型」当作依赖,需改用LOOPX_MANAGER_MODEL显式表达。无凭据路径逐字段与main等价(tests/test_chat_manager_context.py直接断言)。 - capabilities 新增字段是 additive,前端 schema 早已
.optional();无channel_binding参数时投影与main相同(已测)。 - 与同系列 #4443 共享新文件
loopx/control_plane/operator_credential.py:先合并的一方落地后,另一方会得到一次 add/add 冲突(内容相同,取任一即可),合并队列需要一次 rebase。 - 未验证维度:真实 CLI/Codex 会话建立、真实飞书往返均未覆盖,因此本评审不声称线上可用性已验收;托管宿主的管家传输仍不存在,显式选择
dsh之后仍是一条被拒绝的路径。
我的整体评价
APPROVE。这是对一处默认路径语义错误的正确修复,change_proportionality=proportionate、repository_reuse=separation_justified、observable_semantics=intentional_change_validated、authority_semantics=aligned。相对被它替代的旧版本(凭据解析 + provider→模型映射 + dsh_chat_transport_unsupported 这套并行词表)范围是缩小的:删掉了死代码路径,只保留唯一可达的 typed 拒绝。非阻塞建议两条:把 executor_kind 词表抽到一个双方都能导入的叶子模块(不要放进 turn_driver 包——实测会让 Chat 服务多加载 18 个模块、约 0.25s);把真实管家往返验收放进后续切片而不是本 PR。
验证:tests/test_manager_channel_binding.py + test_chat_manager_context.py 24 passed;pytest tests -k "chat or manager or lark" 783 passed;examples/loopx-steward-channel-binding-smoke.py、loopx-chat-server-smoke.py、docs-governance-smoke.py 全通过;loopx canary premerge --from-git-diff gate passed;远端 CI 20 pass / 4 skipping / 1 pending,无失败。
English verdict: APPROVE at 91c440ca021dd2f7378f29bb539db7f305555c4f. The steward channel now selects its executor instead of resolving it from a credential: shipped endpoint codex, LOOPX_MANAGER_ENDPOINT re-points it, the model follows the selected endpoint, and a configured operator credential is reported as a fact (operator_credential_configured, env var name only) rather than silently swapping the model to deepseek-flash while the endpoint stayed on the CLI. A host without an interactive Chat transport now fails as the typed managed_host_chat_transport_unsupported host-tool gate instead of an untyped unknown Agent endpoint error. Key non-blocking note: loopx/control_plane/operator_credential.py is added by both this PR and #4443, so the second one to merge needs a one-line add/add rebase. Validation: 24 focused tests, 783 chat/manager/lark tests, the steward channel binding smoke, the chat-server smoke, the docs-governance smoke, and loopx canary premerge --from-git-diff all pass; remote CI has no failures.
feat(chat): select the steward channel executor explicitly
Summary
The steward channel now selects its executor, and a configured operator
credential no longer re-points it. This replaces the previous steward rule,
where
DEEPSEEK_API_KEYbeing present resolved the channel onto the managedhost and simultaneously swapped the model to the operator model -- while the
endpoint stayed
codexbecause dsh has no interactive Chat transport. Theresult was an executor and a model that disagreed with each other, and a model
that was handed to the Codex adapter.
Discovering a provider key is not a decision to change how the steward runs.
The operator's decision is now explicit, and the credential only
authenticates the endpoint that decision selected.
What changed
codex(interactive CLI transport)LOOPX_MANAGER_ENDPOINTexecutor_endpoint_id, oropen_manager_session(executor_endpoint_id=...)gpt-6-astra, overridable only byLOOPX_MANAGER_MODELhigh, overridable only byLOOPX_MANAGER_REASONING_EFFORTmanager_channel_bindingresolves the endpoint from explicit configurationonly (
selected_manager_executor_endpoint). A credential value or presence isnot an input to that choice, which is now covered by a test that asserts the
endpoint and the model are identical with and without a credential.
keeps the vendor default.
manager_model_configno longer consults thecredential, and the previous
deepseek-flashdefault (which is not even thedsh host's model id) is gone.
"codex"fallbacks inchat_lark_api.pyandextensions/lark/goal_topic_runtime.pynow resolve through the one functionthat owns the rule, and
open_manager_sessionresolves its endpoint only whenthe caller did not name one.
The managed host fails closed, with a typed gate
The managed Turn host (
dsh) runs one bounded work segment per request and hasno interactive Chat transport, so it cannot hold a steward session yet.
Selecting it now produces a typed outcome instead of an "unknown Agent endpoint"
ValueError:manager_channel_bindingavailable: false,unavailable_reason: managed_host_chat_transport_unsupportedCodexChatAgentError(managed_host_chat_transport_unsupported) carrying ahost_tool_gatewhose next action namesloopx turnPromoting
dshto the steward default is gated on that transport, not on acredential appearing. The steward still drives managed work on
dsh: thosebounded Turns are the managed execution unit, and they are a different surface
from the channel the steward answers on.
Frontend readback
The Chat capabilities payload now carries
manager.channel_binding: resolvedexecutor endpoint and its source, executor kind in the same vocabulary as the
governed Turn surface (
individual/managed), resolved model and its source,whether an operator credential is configured, and
available/unavailable_reasonwhen LoopX can prove the selected endpointcannot serve this channel.
availableisnullwhen the projection makes noclaim. Credential facts are the env var name only; values never appear.
Behavior change
Affected lane: the steward channel.
codexand themodel stays
gpt-6-astra(previously: model silently becamedeepseek-flashwhile the executor stayed
codex).LOOPX_MANAGER_ENDPOINT,LOOPX_MANAGER_MODEL, orLOOPX_MANAGER_REASONING_EFFORTstill wins, and did before.Disclosed in
docs/reference/protocols/manager-evidence-and-continuity-v0.mdunder a new Steward channel host selection section.
Validation
tests/test_manager_channel_binding.py,tests/test_chat_manager_context.py(24 tests)examples/loopx-steward-channel-binding-smoke.pyexamples/loopx-chat-server-smoke.py,examples/docs-governance-smoke.pyloopx canary premerge --from-git-diffEvery credential in the tests and smokes is a fixture string whose only role is
to prove it does not select anything. No provider calls are made.
Boundaries
chat_managerimports credential facts fromcontrol_plane/operator_credential.py(already introduced on the Turn side)and reuses the Chat runtime's own managed-host set and error code rather than
restating them.
Delivery status: readback only, frontend lands separately
This slice ships the backend rule, the typed refusal, and the capabilities
readback. It deliberately does not ship a frontend consumer: no dashboard
schema, no manager-header chip, and no packaged
loopx/web/chatassets. Thesteward execution chip is an owner-approved surface, and it is refreshed onto
this corrected
channel_bindingshape (plus the rebuilt packaged bundle) in thesibling frontend/docs slice instead of being merged here. Delivery is therefore
partial on the frontend entry point by design, and the chat entry point
(
/chat/capabilities) is complete and covered above.Supersedes
This main-based slice replaces
#4417(credential-resolved steward binding),the steward half of the stacked chain. The managed-execution half is replaced by
the sibling
#4443. Both slices are main-based, so they can be merged in eitherorder without layer-by-layer refreshes.
#4419(steward execution chip in the manager header) is not replaced: it isthe frontend consumer of this readback and is refreshed onto the corrected
channel_bindingshape in the docs/frontend slice rather than merged as-is,because it currently renders the credential-selected binding this PR removes.