Skip to content

feat(steward): admit a team preview only when it can be validated - #4522

Merged
huangruiteng merged 2 commits into
mainfrom
codex/steward-team-preview-caller
Sep 16, 2026
Merged

huangruiteng merged 2 commits into
mainfrom
codex/steward-team-preview-caller

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

Motivation

The preview validator (PR #4519) had no production caller, so the preview slice was inert: nothing produced the kind and nothing checked it. This finds the seam and wires it.

Change

loopx/chat.py owns normalize_agent_response, the point where a provider response becomes the public Chat contract and where every proposal is already filtered. A proposal of the preview kind is now validated there against host facts passed in by the caller — the Goal's registered Agents and the action kinds this host supports — and is surfaced only as a validated preview carrying applies: false.

Fail-closed rules:

  • without team_plan_context the preview is not surfaced, rather than admitted half-checked;
  • a malformed preview is dropped exactly like any other proposal the normalizer cannot accept, and the owner's answer text still arrives;
  • plain Todo proposals keep their existing behaviour unchanged.

Caller surface: normalize_agent_response(..., team_plan_context={...}) and _normalize_proposals(..., team_plan_context=...); the context is optional, so every existing caller keeps working unchanged.

Honest limitation

No adapter passes team_plan_context yet, so in production the preview kind is still not surfaced. That is deliberate: it lands the caller and its rules behavior-preservingly, and supplying the context is the next step in the same slice.

Validation

pytest tests/test_steward_team_plan_preview.py → 10 passed, including the three normalizer outcomes and the unchanged Todo path.

Risk

Behavior-preserving for existing callers (optional parameter, no context means no preview). The new path can only add a proposal for one new kind, and only after the validator accepts it.

The guard asserts that the modules which resolve the steward's endpoint, model
and effort are the ones that pass machine_defaults, but it still named
`loopx/chat_server.py`. That module no longer holds such a call -- the manager
readability projection moved to `loopx/chat_manager_context.py` -- so the test
failed on every commit, including untouched ones, and a permanently red
invariant check is exactly what hides a real regression later.

The expected set now names the module that actually carries the call. The
undocumented-call-site assertion, the call-site floor, and all other
assertions are unchanged, so the check keeps its teeth.

Verified: tests/test_manager_channel_binding.py 25 passed (was 1 failed), and a
combined sweep of the steward surfaces touched today -- Lark topic connections
and runtime, manager channel binding, manager context handoff, the Lark API
contract, the manager guidance contract, the team-plan preview contract and the
SSH evidence suite -- is 212 passed with no remaining failure.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The preview validator had no production caller, which left the preview slice
inert: nothing could produce the kind and nothing could check it. The seam is
the Chat response normalizer (`loopx/chat.py`), which is where a provider
response becomes the public Chat contract and where every other proposal is
already filtered.

A proposal of the preview kind is now validated there against the host facts --
the Goal's registered Agents and the action kinds this host supports -- and is
surfaced only as a validated preview carrying `applies: false`. Without those
facts the preview cannot be checked, so it is not surfaced at all instead of
being admitted half-checked; a malformed preview is dropped exactly like any
other proposal the normalizer cannot accept, and the owner's answer text still
arrives. The plain Todo proposals keep their existing behaviour.

Not yet wired: no adapter passes `team_plan_context` yet, so in production the
preview kind is still not surfaced. That keeps this change behavior-preserving
while it lands the caller and its rules; supplying the context is the next step
of the same slice.

Verified: tests/test_steward_team_plan_preview.py 10 passed, including the three
normalizer outcomes (surfaced with host facts, withheld without them, dropped
when malformed) and the unchanged Todo path.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head on codex/steward-team-preview-caller (2 files, +85/-8).

动机

上一片交付的预览校验器没有生产调用方,等于预览片仍是惰性的:既没有东西产出该 kind,也没有东西校验它。本轮先定位成型点再接线。

改动思路

成型点是 loopx/chat.py 的 normalize_agent_response——provider 响应变成本地 Chat 公开契约、且所有 proposal 已被过滤的地方。预览类 proposal 在这里用调用方传入的宿主事实(Goal 已注册 Agent、本机支持的 action kind)校验,只以带 applies: false 的已校验预览对外暴露。

具体改动

_normalize_proposals 增加 team_plan_context 分支与 _validated_team_plan_preview:缺上下文则不暴露(不半校验放行);校验失败则像其它无法接受的 proposal 一样丢弃,业主的答复文本仍然送达;kind:"todo" 路径行为不变。normalize_agent_response 透传该可选参数,故所有既有调用方零改动。

对主干的风险

对既有调用方行为保持(可选参数,不给上下文就不出现预览)。新路径只可能为一个新 kind 追加 proposal,且必须通过校验器。诚实限制:目前没有适配器传 team_plan_context,所以生产上该 kind 仍不会出现——这是有意的,本片只把调用方与其规则落地,供上下文是同一片的下一步。

我的整体评价

无阻断性问题,建议合并。验证:tests/test_steward_team_plan_preview.py 10 passed,覆盖"有宿主事实则暴露为 applies:false 的已校验预览 / 无宿主事实则不暴露 / 畸形则丢弃且答复文本保留 / todo 路径不变"。诚实说明:未跑 pr-review --check-result 机器校验、未等远端 CI(预算用于实现与验证),已在提交信息与 PR 披露。

English verdict: APPROVE - wires the team-plan preview validator into its production seam (loopx/chat.py::normalize_agent_response), admitting a preview only when the caller supplies the host facts, surfacing it as a validated applies: false preview, withholding it without those facts, and dropping malformed previews while keeping the answer text; parameter is optional so every existing caller is unchanged. 10 tests pass; the machine --check-result was skipped and disclosed.

@huangruiteng
huangruiteng merged commit 3c8c832 into main Sep 16, 2026
4 checks passed
@huangruiteng
huangruiteng deleted the codex/steward-team-preview-caller branch September 16, 2026 09:32
huangruiteng added a commit that referenced this pull request Sep 16, 2026
… linkage (#4527)

The Steward Team Intake section still described the work as planned and told a
reader that the preview slice came next with no materializer registered. The
preview contract and validator (#4519), the Chat admission gate (#4522) and the
PRE_SETTLEMENT apply through the canonical Todo owner (#4524) have all shipped,
and one docstring still repeated the old claim that the kind has no materializer
and no settlement phase, so the section contradicted the code it points at.

The section now records what is enforced and where: the kind and schema, the
8-lane limit, the priority and gap vocabularies, `declined_first_todo` on a lane
that cannot be staffed, `applies: false` in the validated payload, the admission
gate that needs `team_plan_context`, and the apply that re-validates against the
Goal's registered Agents and the shipped advancement action kinds before it
creates any Todo.

It also states the part that is not done, so the section cannot be read as a
delivered feature: nothing supplies `team_plan_context` in production yet, the
confirmed Chat preview still needs the bridge into the governed capability
journal's `transition_proposals`, the published receipt carries only the first
lane Todo's identity rather than every lane it created, and there is no
frontend confirmation surface for a multi-lane preview. Two of those are
bounded follow-ups rather than unknowns, and the receipt one is called out as a
compatibility decision because that field set is closed and persisted.

Finally, the intake is placed against the multi-agent contracts it reuses rather
than duplicates: it is a user-layer affordance over the kernel that
`multi_agent_three_layer_minimality_contract_v0` defines, and it joins
`multi_agent_visible_launcher_v0` by identity (goal, agent, first lane Todo)
under the launcher's own rule against a leader agent, hidden scheduler,
promotion authority or second source of truth.

Verified: `examples/docs-governance-smoke.py` ok; tests/test_steward_team_plan_preview.py,
tests/test_steward_team_plan_apply.py and tests/capabilities/test_steward_executor_machine_defaults.py 19 passed.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Co-authored-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng added a commit that referenced this pull request Sep 16, 2026
…4533)

A team preview is admitted only when the host can say which Agents exist for the
Goal the plan names, and nothing in production supplied that. A model-authored
preview was therefore always dropped, so the intake shipped in #4519/#4522/#4524
could not surface at all. The Turn owner now supplies the facts, per Goal.

The manager channel is not bound to one Goal, so admission receives a lookup
rather than one Goal's facts: the owner's own channel resolves any registered
Goal, an external manager channel resolves only the Goals its scope resolver
authorizes, and a Goal the registry does not know - or one outside that channel's
scope - resolves to "cannot describe this Goal", which drops the preview instead
of validating it against another Goal's Agents. The lookup runs once per
proposal, for the Goal the proposal names, and re-resolves the channel scope each
time rather than caching an authorization decision.

The contract now distinguishes two answers that used to look alike: an empty
Agent list is a fact, so the plan's lanes become typed `agent_not_registered`
gaps, while an unresolvable Goal surfaces nothing at all. `parse_agent_response`
carries the context through to the four segments that parse an answer (managed
dsh, Codex agent, ACP, provider HTTP), so the same admission rule applies to
every transport instead of only the one the steward happens to run on.

Verified: tests/test_steward_team_plan_preview.py, tests/test_steward_team_plan_apply.py,
tests/test_chat_manager_context.py and tests/test_attached_session_broker.py 63 passed,
including a preview for a Goal outside the channel scope that is dropped, a
per-Goal lookup that is asked only about the named Goal, a Goal with no
registered Agents that becomes a typed gap, and the owner channel resolving any
registered Goal. examples/docs-governance-smoke.py and
examples/loopx-steward-managed-chat-smoke.py ok.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Co-authored-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

审查对象:4522@e157e46a271b5fed31cf366527f432d3ff44c8c1(已合并,合并后审计)。merge base 3acd07697f05d670b7c0f97f2e54a7a22ff71a41;本 PR 自身改动为 loopx/chat.py(+约 55 行)与 tests/test_steward_team_plan_preview.py(+48);diff 中第三个文件 tests/test_manager_channel_binding.py 的一行属于已合并的 #4520,已在 b3f3df63bfb344d890a22711457d8060efcf07ce 单独出具结论。

动机

#4519 的校验器没有生产调用方,整个预览片是惰性的。这个 PR 找到接缝并接上:normalize_agent_response 是 provider 响应变成公开 Chat 契约的那一点,所有 proposal 本来就经过这里过滤,因此在这里加一个按 kind 分支即可,不需要新的入端口。

改动思路

给 normalize_agent_response / _normalize_proposals 增加可选的 team_plan_context(由调用方提供"该 Goal 已注册的 Agent"与"本机支持的 action kind"),预览类 proposal 用 #4519 的校验器校验后只以 applies: false 的已验证预览出现;没有 context 就不浮现(而不是半校验地放行);校验失败按既有 proposal 一样丢弃,答案正文照常返回;todo proposal 行为不变。

具体改动

一个 import、一个可选参数与注解(list[dict[str, str]] → list[dict[str, Any]])、proposal 循环里的一个分支、一个约 20 行的 _validated_team_plan_preview 包装(只捕获 ValueError),以及一个覆盖四种结果的测试。

我做的独立核对(真实函数探针):

  • 有 context + 合法计划 → proposals=[{"kind":"steward_team_plan_preview","preview":{...,"applies":false,"lanes":[{"staffing":"ready",...}]}}];
  • 无 context → proposals=[],与 merge base 行为一致(行为保持成立);
  • 有 context 但预览非法(如 priority P4)→ proposals=[] 且 message 照常返回;
  • todo proposal → 输出与 base 逐字相同({"kind":"todo","text":...,"priority":"P1","rationale":""});未知 kind 仍被丢弃;
  • pytest tests/test_steward_team_plan_preview.py → 10 passed;loopx check --scan-path 两文件 → clean;git diff --check 干净。
  • 惰性确认:全仓 grep 显示没有生产调用方传 team_plan_context(attached_session.py:279 是唯一生产调用点且不传),与 PR 的 "Honest limitation" 完全一致。

对主干的风险

行为保持成立、opt-in 路径隔离良好,但有两处需要在下一切片落地前收紧(其一为 P2):

  1. P2:context 本身畸形时会抛异常,owner 的正文也会丢。 _validated_team_plan_preview 用 list(context.get(...) or []) 取事实、只捕获 ValueError。实测:team_plan_context={"registered_agent_ids": 5, "supported_action_kinds": ["implement"]} → TypeError: 'int' object is not iterable 直接从 normalize_agent_response 抛出(supported_action_kinds 非可迭代时同样)。而生产调用点在 attached_session.py:279 的 attached-turn 回写路径上,异常会失败整次完成。这与 PR 承诺的"畸形预览被丢弃、答案正文仍到达"相矛盾——一次调用方笔误的代价是 owner 的整个答复。建议把两个 context 值校验为字符串序列并抛具名错误(或同时捕获 TypeError 并丢弃预览),让"坏 context 不影响正文"成为结构性事实。
  2. P3:context 是无类型 mapping,键名笔误与"没传 context"不可区分。 实测:{"registered_agents": [...], "action_kinds": [...]} 这种拼写错误返回 0 个 proposal 且无任何提示——与文档化的"未提供 context 就不浮现"完全同形,功能会静默失效。给 context 一个类型化形状(两个具名的字符串序列字段)或在提供 context 时强制要求两个键,就能在边界处报错。
  3. P3:测试未覆盖 context 失败模式,且 proposals 形状变宽后无消费方检查。 新测试覆盖了四种声明结果,但没有传畸形/部分 context(所以上面的 P2 得以通过评审);同时 proposals 的条目类型从扁平 todo 记录扩为可含嵌套 preview,而 chat_store.py:1002 会原样转发这些 proposal、dashboard/TS 类型未被验证。补一个畸形 context 用例与一个"消费方忽略未知 proposal kind"的断言即可。

我的整体评价

接缝找得准、改动克制,而且"无 context 即不浮现"的行为保持在实测中成立——这一点我逐条对过 base 输出。问题都出在新增入参的健壮性:一个未类型化的 context 既能让调用方笔误静默失效,也能在取值畸形时把整次答复带走,后者直接违背了 PR 自己写下的保证,因此给出 P2 与两点 P3,建议在下一切片(也就是真正开始传 context 的那一片)一并修掉。作为已合并精确 head 的合并后审计,在指出上述问题后,证据支持通过。

English verdict: APPROVE (exact head e157e46)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant