feat(steward): admit a team preview only when it can be validated - #4522
Conversation
The guard asserts that the modules which resolve the steward's endpoint, model and effort are the ones that pass machine_defaults, but it still named `loopx/chat_server.py`. That module no longer holds such a call -- the manager readability projection moved to `loopx/chat_manager_context.py` -- so the test failed on every commit, including untouched ones, and a permanently red invariant check is exactly what hides a real regression later. The expected set now names the module that actually carries the call. The undocumented-call-site assertion, the call-site floor, and all other assertions are unchanged, so the check keeps its teeth. Verified: tests/test_manager_channel_binding.py 25 passed (was 1 failed), and a combined sweep of the steward surfaces touched today -- Lark topic connections and runtime, manager channel binding, manager context handoff, the Lark API contract, the manager guidance contract, the team-plan preview contract and the SSH evidence suite -- is 212 passed with no remaining failure. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The preview validator had no production caller, which left the preview slice inert: nothing could produce the kind and nothing could check it. The seam is the Chat response normalizer (`loopx/chat.py`), which is where a provider response becomes the public Chat contract and where every other proposal is already filtered. A proposal of the preview kind is now validated there against the host facts -- the Goal's registered Agents and the action kinds this host supports -- and is surfaced only as a validated preview carrying `applies: false`. Without those facts the preview cannot be checked, so it is not surfaced at all instead of being admitted half-checked; a malformed preview is dropped exactly like any other proposal the normalizer cannot accept, and the owner's answer text still arrives. The plain Todo proposals keep their existing behaviour. Not yet wired: no adapter passes `team_plan_context` yet, so in production the preview kind is still not surfaced. That keeps this change behavior-preserving while it lands the caller and its rules; supplying the context is the next step of the same slice. Verified: tests/test_steward_team_plan_preview.py 10 passed, including the three normalizer outcomes (surfaced with host facts, withheld without them, dropped when malformed) and the unchanged Todo path. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed exact head on codex/steward-team-preview-caller (2 files, +85/-8).
动机
上一片交付的预览校验器没有生产调用方,等于预览片仍是惰性的:既没有东西产出该 kind,也没有东西校验它。本轮先定位成型点再接线。
改动思路
成型点是 loopx/chat.py 的 normalize_agent_response——provider 响应变成本地 Chat 公开契约、且所有 proposal 已被过滤的地方。预览类 proposal 在这里用调用方传入的宿主事实(Goal 已注册 Agent、本机支持的 action kind)校验,只以带 applies: false 的已校验预览对外暴露。
具体改动
_normalize_proposals 增加 team_plan_context 分支与 _validated_team_plan_preview:缺上下文则不暴露(不半校验放行);校验失败则像其它无法接受的 proposal 一样丢弃,业主的答复文本仍然送达;kind:"todo" 路径行为不变。normalize_agent_response 透传该可选参数,故所有既有调用方零改动。
对主干的风险
对既有调用方行为保持(可选参数,不给上下文就不出现预览)。新路径只可能为一个新 kind 追加 proposal,且必须通过校验器。诚实限制:目前没有适配器传 team_plan_context,所以生产上该 kind 仍不会出现——这是有意的,本片只把调用方与其规则落地,供上下文是同一片的下一步。
我的整体评价
无阻断性问题,建议合并。验证:tests/test_steward_team_plan_preview.py 10 passed,覆盖"有宿主事实则暴露为 applies:false 的已校验预览 / 无宿主事实则不暴露 / 畸形则丢弃且答复文本保留 / todo 路径不变"。诚实说明:未跑 pr-review --check-result 机器校验、未等远端 CI(预算用于实现与验证),已在提交信息与 PR 披露。
English verdict: APPROVE - wires the team-plan preview validator into its production seam (loopx/chat.py::normalize_agent_response), admitting a preview only when the caller supplies the host facts, surfacing it as a validated applies: false preview, withholding it without those facts, and dropping malformed previews while keeping the answer text; parameter is optional so every existing caller is unchanged. 10 tests pass; the machine --check-result was skipped and disclosed.
… linkage (#4527) The Steward Team Intake section still described the work as planned and told a reader that the preview slice came next with no materializer registered. The preview contract and validator (#4519), the Chat admission gate (#4522) and the PRE_SETTLEMENT apply through the canonical Todo owner (#4524) have all shipped, and one docstring still repeated the old claim that the kind has no materializer and no settlement phase, so the section contradicted the code it points at. The section now records what is enforced and where: the kind and schema, the 8-lane limit, the priority and gap vocabularies, `declined_first_todo` on a lane that cannot be staffed, `applies: false` in the validated payload, the admission gate that needs `team_plan_context`, and the apply that re-validates against the Goal's registered Agents and the shipped advancement action kinds before it creates any Todo. It also states the part that is not done, so the section cannot be read as a delivered feature: nothing supplies `team_plan_context` in production yet, the confirmed Chat preview still needs the bridge into the governed capability journal's `transition_proposals`, the published receipt carries only the first lane Todo's identity rather than every lane it created, and there is no frontend confirmation surface for a multi-lane preview. Two of those are bounded follow-ups rather than unknowns, and the receipt one is called out as a compatibility decision because that field set is closed and persisted. Finally, the intake is placed against the multi-agent contracts it reuses rather than duplicates: it is a user-layer affordance over the kernel that `multi_agent_three_layer_minimality_contract_v0` defines, and it joins `multi_agent_visible_launcher_v0` by identity (goal, agent, first lane Todo) under the launcher's own rule against a leader agent, hidden scheduler, promotion authority or second source of truth. Verified: `examples/docs-governance-smoke.py` ok; tests/test_steward_team_plan_preview.py, tests/test_steward_team_plan_apply.py and tests/capabilities/test_steward_executor_machine_defaults.py 19 passed. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Co-authored-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…4533) A team preview is admitted only when the host can say which Agents exist for the Goal the plan names, and nothing in production supplied that. A model-authored preview was therefore always dropped, so the intake shipped in #4519/#4522/#4524 could not surface at all. The Turn owner now supplies the facts, per Goal. The manager channel is not bound to one Goal, so admission receives a lookup rather than one Goal's facts: the owner's own channel resolves any registered Goal, an external manager channel resolves only the Goals its scope resolver authorizes, and a Goal the registry does not know - or one outside that channel's scope - resolves to "cannot describe this Goal", which drops the preview instead of validating it against another Goal's Agents. The lookup runs once per proposal, for the Goal the proposal names, and re-resolves the channel scope each time rather than caching an authorization decision. The contract now distinguishes two answers that used to look alike: an empty Agent list is a fact, so the plan's lanes become typed `agent_not_registered` gaps, while an unresolvable Goal surfaces nothing at all. `parse_agent_response` carries the context through to the four segments that parse an answer (managed dsh, Codex agent, ACP, provider HTTP), so the same admission rule applies to every transport instead of only the one the steward happens to run on. Verified: tests/test_steward_team_plan_preview.py, tests/test_steward_team_plan_apply.py, tests/test_chat_manager_context.py and tests/test_attached_session_broker.py 63 passed, including a preview for a Goal outside the channel scope that is dropped, a per-Goal lookup that is asked only about the named Goal, a Goal with no registered Agents that becomes a typed gap, and the owner channel resolving any registered Goal. examples/docs-governance-smoke.py and examples/loopx-steward-managed-chat-smoke.py ok. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Co-authored-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
审查对象:4522@e157e46a271b5fed31cf366527f432d3ff44c8c1(已合并,合并后审计)。merge base 3acd07697f05d670b7c0f97f2e54a7a22ff71a41;本 PR 自身改动为 loopx/chat.py(+约 55 行)与 tests/test_steward_team_plan_preview.py(+48);diff 中第三个文件 tests/test_manager_channel_binding.py 的一行属于已合并的 #4520,已在 b3f3df63bfb344d890a22711457d8060efcf07ce 单独出具结论。
动机
#4519 的校验器没有生产调用方,整个预览片是惰性的。这个 PR 找到接缝并接上:normalize_agent_response 是 provider 响应变成公开 Chat 契约的那一点,所有 proposal 本来就经过这里过滤,因此在这里加一个按 kind 分支即可,不需要新的入端口。
改动思路
给 normalize_agent_response / _normalize_proposals 增加可选的 team_plan_context(由调用方提供"该 Goal 已注册的 Agent"与"本机支持的 action kind"),预览类 proposal 用 #4519 的校验器校验后只以 applies: false 的已验证预览出现;没有 context 就不浮现(而不是半校验地放行);校验失败按既有 proposal 一样丢弃,答案正文照常返回;todo proposal 行为不变。
具体改动
一个 import、一个可选参数与注解(list[dict[str, str]] → list[dict[str, Any]])、proposal 循环里的一个分支、一个约 20 行的 _validated_team_plan_preview 包装(只捕获 ValueError),以及一个覆盖四种结果的测试。
我做的独立核对(真实函数探针):
- 有 context + 合法计划 →
proposals=[{"kind":"steward_team_plan_preview","preview":{...,"applies":false,"lanes":[{"staffing":"ready",...}]}}]; - 无 context →
proposals=[],与 merge base 行为一致(行为保持成立); - 有 context 但预览非法(如 priority
P4)→proposals=[]且 message 照常返回; - todo proposal → 输出与 base 逐字相同(
{"kind":"todo","text":...,"priority":"P1","rationale":""});未知 kind 仍被丢弃; pytest tests/test_steward_team_plan_preview.py→ 10 passed;loopx check --scan-path两文件 → clean;git diff --check干净。- 惰性确认:全仓 grep 显示没有生产调用方传
team_plan_context(attached_session.py:279是唯一生产调用点且不传),与 PR 的 "Honest limitation" 完全一致。
对主干的风险
行为保持成立、opt-in 路径隔离良好,但有两处需要在下一切片落地前收紧(其一为 P2):
- P2:context 本身畸形时会抛异常,owner 的正文也会丢。
_validated_team_plan_preview用list(context.get(...) or [])取事实、只捕获ValueError。实测:team_plan_context={"registered_agent_ids": 5, "supported_action_kinds": ["implement"]}→TypeError: 'int' object is not iterable直接从normalize_agent_response抛出(supported_action_kinds非可迭代时同样)。而生产调用点在attached_session.py:279的 attached-turn 回写路径上,异常会失败整次完成。这与 PR 承诺的"畸形预览被丢弃、答案正文仍到达"相矛盾——一次调用方笔误的代价是 owner 的整个答复。建议把两个 context 值校验为字符串序列并抛具名错误(或同时捕获TypeError并丢弃预览),让"坏 context 不影响正文"成为结构性事实。 - P3:context 是无类型 mapping,键名笔误与"没传 context"不可区分。 实测:
{"registered_agents": [...], "action_kinds": [...]}这种拼写错误返回 0 个 proposal 且无任何提示——与文档化的"未提供 context 就不浮现"完全同形,功能会静默失效。给 context 一个类型化形状(两个具名的字符串序列字段)或在提供 context 时强制要求两个键,就能在边界处报错。 - P3:测试未覆盖 context 失败模式,且 proposals 形状变宽后无消费方检查。 新测试覆盖了四种声明结果,但没有传畸形/部分 context(所以上面的 P2 得以通过评审);同时 proposals 的条目类型从扁平 todo 记录扩为可含嵌套
preview,而chat_store.py:1002会原样转发这些 proposal、dashboard/TS 类型未被验证。补一个畸形 context 用例与一个"消费方忽略未知 proposal kind"的断言即可。
我的整体评价
接缝找得准、改动克制,而且"无 context 即不浮现"的行为保持在实测中成立——这一点我逐条对过 base 输出。问题都出在新增入参的健壮性:一个未类型化的 context 既能让调用方笔误静默失效,也能在取值畸形时把整次答复带走,后者直接违背了 PR 自己写下的保证,因此给出 P2 与两点 P3,建议在下一切片(也就是真正开始传 context 的那一片)一并修掉。作为已合并精确 head 的合并后审计,在指出上述问题后,证据支持通过。
English verdict: APPROVE (exact head e157e46)
Motivation
The preview validator (PR #4519) had no production caller, so the preview slice was inert: nothing produced the kind and nothing checked it. This finds the seam and wires it.
Change
loopx/chat.pyownsnormalize_agent_response, the point where a provider response becomes the public Chat contract and where every proposal is already filtered. A proposal of the preview kind is now validated there against host facts passed in by the caller — the Goal's registered Agents and the action kinds this host supports — and is surfaced only as a validated preview carryingapplies: false.Fail-closed rules:
team_plan_contextthe preview is not surfaced, rather than admitted half-checked;Caller surface:
normalize_agent_response(..., team_plan_context={...})and_normalize_proposals(..., team_plan_context=...); the context is optional, so every existing caller keeps working unchanged.Honest limitation
No adapter passes
team_plan_contextyet, so in production the preview kind is still not surfaced. That is deliberate: it lands the caller and its rules behavior-preservingly, and supplying the context is the next step in the same slice.Validation
pytest tests/test_steward_team_plan_preview.py→ 10 passed, including the three normalizer outcomes and the unchanged Todo path.Risk
Behavior-preserving for existing callers (optional parameter, no context means no preview). The new path can only add a proposal for one new kind, and only after the validator accepts it.