Skip to content

feat(steward): report lane readiness as verified rungs, not as "ready" - #4604

Closed
huangruiteng wants to merge 6 commits into
mainfrom
codex/steward-lane-readiness-ladder
Closed

huangruiteng wants to merge 6 commits into
mainfrom
codex/steward-lane-readiness-ladder

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

Roadmap card R2 / audit finding F6: a team plan lane's ready only ever meant "this Goal registers the Agent" and "this host ships the action kind". The card showed it beside the Agent, the first bounded Todo and the acceptance signal, so an owner could read it as work an executor will pick up. Nothing in the preview said which conditions were actually machine-detected, which is the whole of F6: the weaker fact read as the stronger one.

What a staffed lane now carries

readiness:
  schema_version: steward_lane_readiness_ladder_v0
  launchable: false
  verified_rungs: [registered, action_kind_supported]
  unverified_rungs:
    - {rung: addressable, reason_code: host_has_no_agent_presence_provider}
    - {rung: bound,       reason_code: host_has_no_runtime_binding_readback}
    - {rung: launchable,  reason_code: host_has_no_launch_probe}
    - {rung: executing,   reason_code: lane_not_materialized_by_a_preview}

The unverified rungs are precisely the ones this host has no provider for: the agent directory has no presence provider, and there is no runtime-binding readback, launch probe, or execution reading. Naming them is the honest half of the contract — a lane is not launchable today, and the preview says so rather than leaving absence of evidence to be read as evidence.

Compatibility

staffing: "ready" stays the plan's own vocabulary and is unchanged, so no existing consumer breaks. Both the frontend reducer and the TS review plan branch on staffing === "gap", and that is untouched. A gap lane carries no readiness object, because a lane the host could not staff has no rungs to report.

Owner-visible effect

The card renders the ladder beside the lane:

agent-backend · P1 · implement · Implement the bounded intake ·
acceptance: the bounded Todo is created through the canonical owner ·
尚未就绪 —— 已验证:registered, action_kind_supported;未验证:addressable, bound, launchable, executing

Validation

uv run --extra test python -m pytest tests/test_steward_team_plan_preview.py -q   # 20 passed
uv run --extra test python -m pytest tests/ -q -k "steward or team_plan"          # 58 passed
LOOPX_PERSONAL_WORKSPACE_SCENARIO=team-plan node examples/personal-workspace-browser-smoke.mjs
# ok — asserts the rendered readiness line names both halves
node examples/personal-workspace-browser-smoke.mjs    # all six scenarios ok
node src/features/personal-workspace/workspace-theme.test.mjs   # contract ok

The packaged chat bundle is rebuilt and is a build fixed point. smoke:team-plan-proposal cannot run on main today (TS2688: Cannot find type definition file for 'node'); that failure reproduces on a clean origin/main worktree, so the browser scenario is the end-to-end proof for this change. The check added to that smoke file is kept for when it runs again.

Boundary

Public-safe: synthetic fixtures only, no private state, credentials, local paths or raw evidence. No change to staffing decisions, apply/settlement semantics, or which lanes a plan admits — only to what the preview states about them.

Roadmap card R2 / audit F6: `ready` validated only that the Goal registers the
Agent and that this host ships the action kind, but the plan card showed it next
to an Agent, a first Todo and an acceptance signal, which an owner can read as
work an executor will pick up. Nothing in the preview said which conditions were
actually machine-detected, so the weaker fact read as the stronger one.

A staffed lane now carries the host's reading of its own check:

    readiness:
      schema_version: steward_lane_readiness_ladder_v0
      launchable: false
      verified_rungs: [registered, action_kind_supported]
      unverified_rungs:
        - {rung: addressable, reason_code: host_has_no_agent_presence_provider}
        - {rung: bound,       reason_code: host_has_no_runtime_binding_readback}
        - {rung: launchable,  reason_code: host_has_no_launch_probe}
        - {rung: executing,   reason_code: lane_not_materialized_by_a_preview}

The unverified rungs are the ones this host has no provider for: the agent
directory has no presence provider, and there is no runtime-binding readback,
launch probe, or execution reading. Naming them is the honest half of the
contract — a lane is not launchable today, and the preview says so instead of
leaving the absence of evidence to be read as evidence. `staffing: "ready"`
stays the plan's own vocabulary and is unchanged, so no consumer breaks; a gap
lane carries no ladder because it has no rungs to report.

The owner-facing card renders the ladder, so the sentence a reader sees is
"尚未就绪 —— 已验证: registered, action_kind_supported; 未验证: addressable,
bound, launchable, executing" beside the lane it describes.

Validated: `pytest tests/test_steward_team_plan_preview.py` 20 passed and
`pytest -k "steward or team_plan"` 58 passed; the `team-plan` browser scenario
asserts the rendered readiness line end to end, including that it names both
halves; the full development workspace suite and `workspace-theme` contract pass.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
… ladder

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/steward-lane-readiness-ladder branch from c4e65f6 to 8d2e421 Compare September 16, 2026 21:56
huangruiteng added a commit that referenced this pull request Sep 16, 2026
…ng it (#4609)

The model-provider category assertion waited for one credential panel to
become visible and then took a one-shot count. A category switch is a state
transition, not a settled fact, so that count can still observe the pane
mid-mount: the required dashboard-acceptance job failed on that line for #4593
and #4604 while the same tree passed on main's own Python Tests run, and a
one-shot count cannot say whether the page hosted zero or two panels.

The assertion now waits for exactly one panel, and when the count never settles
it names every matching node with its owning section and visibility, so the next
CI failure is attributable without a machine that reproduces it. Because two
matching panels also make a strict-mode locator wait throw, the count settles
before the visibility check instead of after it.

Assertion strength is unchanged: exactly one visible credential panel is still
required inside the model-provider category, and the readback labels are still
asserted below.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Co-authored-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

计划卡上一条 lane 标着 staffing: "ready" 时,owner 很容易读成“有执行者会接手”。但 ready 只说“本 Goal 注册了这个 Agent、本 host 提供这个 action kind”,并不等于 lane 能寻址、能绑定运行时、能启动或正在执行。本车道在 steward 旅程里把这一点记成了缺口:没有 per-lane readiness 阶梯,读者无法区分“已核对的两级”和“本 host 根本没有探针的两级”。

改动思路

不动 staffing 的既有词汇,只增补一个 host 视角的 ladder:明确列出已验证的两级(registered、action_kind_supported)与未提供的四级及其原因码(addressable / bound / launchable / executing),并给出显式的 launchable: False,使弱主张不可能被读成强主张。

具体改动

  • loopx/control_plane/work_items/governed_transition_proposal.py:新增 ladder schema 常量、已验证/未提供两级集合与 _staffed_lane_readiness(),并在 staffed lane 的预览里挂上 readiness 字段;gap lane 不带该字段。
  • tests/test_steward_team_plan_preview.py:新增“staffed lane 报告两级已验证 + 四级未提供(含原因码)”与“gap lane 不带 ladder”两条命名测试。
  • loopx/web/chat 随包 bundle 重建,使打包前端展示同一份读数。

对主干的风险

新增字段只增不改:staffing、first_todo 等既有键保持原义,已有断言不变。风险点在于“未提供”的原因码是否会被未来读成永久事实——它们描述的是当前 host 能力,代码里以常量列出,替换成真实探针时应同时改这里与测试;测试对四个原因码逐项断言,避免悄悄变成空列表。另外这条改动影响打包前端展示,按本仓库的打包契约用重建后的 bundle 一并提交。

我的整体评价

把“ready 被读成 running”这一具体误解用类型化阶梯堵住,并对未验证项给出可替换的原因码,方向与最小性都对。建议在该 head 的必过检查转绿后合并;合并后 steward 旅程里对应的 readiness 缺口可从 gap 转为 proven(仍需 smoke 断言同步)。

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Head: 4093037

English verdict: APPROVE

…adiness-ladder

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

# Conflicts:
#	loopx/web/chat/asset-retention.json
#	loopx/web/chat/assets/index-CFOC0-T9.js
#	loopx/web/chat/assets/index-CXvjZarX.js
#	loopx/web/chat/assets/index-DMfpskuZ.js
#	loopx/web/chat/index.html

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head: 06d7d9d5ce7ad6b9d4f807a37963e848ac675cb1 (re-review after merging origin/main and rebuilding the packaged chat bundle).

loopx pr-review --check-result on the matching packet returned ok: true with no approval_blockers; all 19 evidence rows for this control-plane plan are verified.

动机

计划卡上的 staffing: "ready" 只证明两件事:这个 Goal 注册了该 Agent,以及本机提供该 action kind。但卡片把 ready 摆在 Agent 与首个 Todo 旁边,owner 很容易读成「有执行者会接手」。而真正决定「能不能跑」的四档——addressable、bound、launchable、executing——本机今天一个都验证不了(没有 agent presence provider、没有 runtime binding readback、没有 launch probe,且预览本身不会 materialize lane)。

这个 PR 的价值:它不把弱事实删除或改名,而是把宿主真正验证过的档位与未提供的档位并排说清楚,让卡片不再暗示一个不存在的执行者。对 owner 来说,这是「我确认的是入队,不是开工」的可见区分。

改动思路

关键在于谁拥有这个事实。launchable 这类读数只有控制面知道,前端无从得知——如果让前端判断或硬编码提示,就会出现第二个权威;如果把 staffing 从 ready 改成别的值,就会破坏一个既有字段的语义,用一个值掩盖两种事实。所以选择:控制面在预览里产出 typed 阶梯,前端只渲染。

launchable 的合取规则被写死在常量与注释里:只有当所有前置档位(含 launchable 自身)都被验证时才可为真,当前无一档被验证,因此恒为 false。四个未提供档位各自带 reason_code,而不是一个笼统的布尔。

具体改动

  • loopx/control_plane/work_items/governed_transition_proposal.py(+37):新增 steward_lane_readiness_ladder_v0、STEWARD_LANE_READINESS_VERIFIED_RUNGS = ("registered", "action_kind_supported")、四个 (档位, reason_code) 未提供项与 _staffed_lane_readiness();在已配齐 lane 上输出 readiness。缺口 lane 不带该字段。
  • apps/presentation/dashboard/src/features/personal-workspace/team-plan-preview.ts(+24):teamPlanLaneReadiness() 渲染该句;readiness.launchable === true 或两半皆空时返回空串(前端绝不自行补默认文案)。
  • i18n.tsx(+2):中英双语模板,语义一致。
  • tests/test_steward_team_plan_preview.py(+37)、apps/presentation/dashboard/smoke/team-plan-proposal-smoke.ts(+18)、examples/personal-workspace-browser/team-plan.mjs(+17):断言阶梯内容与「缺口 lane 无阶梯」。
  • loopx/web/chat/*:打包产物在最新 main 上重跑构建生成(见下)。

关键代码讲解

STEWARD_LANE_READINESS_VERIFIED_RUNGS = ("registered", "action_kind_supported")
STEWARD_LANE_READINESS_UNPROVIDED_RUNGS = (
    ("addressable", "host_has_no_agent_presence_provider"),
    ("bound", "host_has_no_runtime_binding_readback"),
    ("launchable", "host_has_no_launch_probe"),
    ("executing", "lane_not_materialized_by_a_preview"),
)

四个 reason_code 都是「本机没有这个探针」的直白陈述,而不是把未知写成已知。前端只在有内容时加话:

if (readiness.launchable === true) return "";
...
if (!verified.length && !unverified.length) return "";

这一早返回让「就绪时不显示」「字段缺失时不猜值」由同一处保证。

对主干的风险

唯一的刻意变化是新增 readiness 子结构与卡片追加一句话。staffing 取值域未变,缺口 lane 行为未变(并有专门断言锁定),前端在字段缺失时行为与改动前逐字一致。新增字段可被旧消费者忽略,不需要迁移。

本次的关键工作是解决长期冲突:该分支与 main 的冲突全部来自打包 frontend 产物(loopx/web/chat/{index.html,asset-retention.json,assets/*} 的 rename/rename 于 hashed 入口文件)。正确解法不是手工合并压缩过的 bundle,而是把该目录重置到 main 后重新构建。我按该流程处理:git rm -r -f --cached loopx/web/chat → git checkout origin/main -- loopx/web/chat → 移走残留未跟踪产物 → npm run build:chat → git add -A。重建后相对 main 的差异只有 asset-retention.json 6 行、index.html 2 行与入口 bundle 34 行,与该分支 +44 行前端源码改动相称。

在更新后的 head 上实测:

  • env -u PYTHONPATH uv run --extra test python -m pytest tests/test_steward_team_plan_preview.py -q → 20 passed
  • cd apps/presentation/dashboard && npm run smoke:team-plan-proposal → ok
  • env -u PYTHONPATH uv run --extra test python examples/dashboard-pwa-bundle-smoke.py → ok(bundle 与 asset-retention 一致)
  • env -u PYTHONPATH uv run --extra test loopx canary premerge --from-git-diff → 0 failures / 0 advisories(catalog canaries 4/4、risk smokes 8/8、public boundary 1/1)

未验证 / 环境说明(如实标注,不当通过):未做真实浏览器像素级视觉审查,也未经 Lark 侧同一卡片的展示复核(载荷相同、展示层不同)。本机 semantic-vocabulary-drift-smoke 与前端 smoke 需要本地未跟踪的 npm 依赖(根 node_modules、@types/node),缺依赖时会以解析器/TS2688 报错;同一 smoke 在未补依赖的干净 main worktree 上以完全相同的方式失败,补齐后复跑通过,证明与本 PR 无关。

边界声明:本 PR 改动 loopx/** 与 apps/**,按仓库规则属控制面改动,只提 PR、由维护者合并,作者不做自合并。

我的整体评价

把「ready ≠ launchable」这件事交回它的所有者(控制面预览)去如实产出,并让前端只做渲染,方向和边界都对;改动比例与它消除的产品级失真相称。无阻断性发现。

残余风险是四个档位的探针尚未实现,因此 launchable 会长期为 false、卡片持续显示「尚未就绪」——这是有意的保守表述;若长期不接入探针,读者可能对提示脱敏,应作为后续切片跟进。

English verdict: APPROVE - re-verified on exact head 06d7d9d: a staffed lane no longer presents "ready" as if an executor would pick the work up. The host emits a typed readiness ladder naming the two rungs it actually verified and the four it has no provider for (each with a reason_code), and the card renders that instead of guessing; launchable is guarded by an explicit conjunction so a registered Agent cannot be read as launchable. The long-standing conflict was only in packaged frontend output and was resolved by rebuilding the chat bundle on the new main rather than hand-merging a minified bundle. 20 preview tests pass, the dashboard team-plan smoke and packaged-bundle smoke pass, and premerge canary reports 0 failures. No blocking finding.

@huangruiteng

Copy link
Copy Markdown
Collaborator Author

“已分配不等于正在执行”的判断正确,已在 #4633 的结果表达和 RFC 边界中保留。目前新增的 readiness ladder 主要是常量缺省值,尚未接入接收方采纳、真实运行绑定或执行探针;把这些值包装成新协议会增加未获验证的状态层。后续应由 R1–R4 的实际依赖交接、采纳、独立验收和结果回传提供事实。

本 PR 关闭并保留分支,避免与 #4633 重复推进。替代 PR 尚待维护者审阅合并。

Superseded by #4633: retained the useful contract, consolidated implementation and validation, and kept assignment separate from collaboration/execution authority.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant