Skip to content

docs(steward): read the confirm result as proven, and check the case against the run - #4716

Merged
huangruiteng merged 2 commits into
mainfrom
codex/steward-case-confirm-slice-20260919
Sep 18, 2026
Merged

huangruiteng merged 2 commits into
mainfrom
codex/steward-case-confirm-slice-20260919

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

  • Goal/source and gap: the steward qualification case is the owner-facing reading of examples/personal-workspace-browser/steward-journey.mjs, and it had drifted. It still listed gap 1 as "a confirmed plan does not distinguish committed / partial / all-gap / stale / rejected per lane", with the recorded evidence "the only outcome sentence is the generic applied notice", and beat 4 as "Proven, but see gap 1".
  • Observable before → after: the scenario has recorded that confirm result as present for some time — its printed gap list starts at 4-readiness and its report shows .personal-team-plan-result naming agent-backend + Implement the bounded intake beside agent-reviewer + Independently wait… and 待安排 · 尚未加入此目标. After this PR both locales say the confirm result is proven, renumber the remaining gaps to 1–4, and the scenario fails if either table drifts again.
  • Issue/task and intended base: main (4262017). No tracking issue.

Scope And Continuation

  • Completed scope and remaining work: the case text now says what the surface proves and, just as important, what it does not: the result states that the assignment is recorded and points at the Goal for progress, so beat 5 ("checks who can actually work") stays gap 1 instead of being absorbed into a proven confirm. Gaps 2–4 (correction, recovery, return) are unchanged.
  • Why the guard lives in the scenario: this table drifted within a day, because another lane delivered the confirm result while the doc kept the old conclusion. Nothing in review fires when a surface changes, but the run knows which gaps it recorded, so the scenario now reads both case files and requires the gap section of each to list exactly one row per recorded gap — and to have dropped the retired confirm-result wording.
  • Slice boundary / successor: complete within this scope. Advances against the remaining gaps stay where they are: docs/product/use-cases/steward names an owner surface per gap, and the readiness gap is tracked by the lane's own P0 todo rather than by this PR.

Validation

  • Tested revision: 4e684ba
  • Run state: finished
  • Input classes: synthetic (the fixture substitutes the agent turn; no live Goal, Agent, credential or local path is read)
Check kind Result Public-safe evidence / limitation
real_entrypoint passed LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey node examples/personal-workspace-browser-smoke.mjs — development ok, beats=4 gaps=4-readiness,5-correction,6-recovery,7-return; the packaged run (LOOPX_PERSONAL_WORKSPACE_PACKAGED=1) is ok against the shipped loopx/web/chat bundle with the same gap list
regression_parity passed Negative control: adding one extra row to the English gap table makes the run fail with docs/product/use-cases/steward/README.md lists 5 gaps but this run recorded 4 (4-readiness, 5-correction, 6-recovery, 7-return); removing it returns the run to green. The guard has teeth on both the count and the retired wording
static passed docs-governance-smoke ok; docs-asset-integrity-smoke: ok (6 assets verified, all PNG hashes distinct); loopx check --scan-path over the three changed files reports errors=0 and public boundary scan clean: 3 files
  • Coverage and gaps: the scenario drives the real packaged frontend over the fixture and now also reads the two case files, so the doc/run pairing is checked on every run of this scenario. The case's own untested paths are unchanged and stay listed: Lark audience and remote/cloud host are not exercised.

Frontend / Visual Evidence

  • UI impact: none (the scenario, and both case mirrors, change; no shipped frontend file is touched)
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: synthetic

Type of Change

  • Test update
  • Documentation update

LoopX Area

  • Public docs or presentation surface (README, protocols, dashboard)

Technical Direction

  • Direction / acceptance reference, when applicable: the steward qualification case (docs/product/use-cases/steward), the frontend-first end-to-end journey for the local steward lane.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A
  • Semantic dimensions changed, or reviewed no-impact rationale: N/A
  • Provider conformance arms run: N/A
  • Read-only legacy/file/PostgreSQL three-arm rehearsal: N/A

Boundary Checklist

  • Neither the diff nor this PR body/comments/attachments disclose private state, credentials, raw traces or verifier output, internal links, or local machine paths
  • I did not duplicate maintainer-owned benchmark work
  • I kept the change scoped to the linked issue/task
  • I completed the visual evidence section for UI changes, or marked UI impact none
  • Every commit includes a DCO Signed-off-by trailer

中文摘要

  • 漂移:用例文档仍把「确认后不区分 committed / partial / all-gap / stale / rejected」记为缺口 1,证据是「只有一条通用的『已应用』提示」,而场景早已把这条记成 present——它打印的缺口清单从 4-readiness 开始,报告里 .personal-team-plan-result 已经逐条点名已分配的 lane 与未派工项及原因。
  • 变化:中英两份用例都改为「确认结果已证明」,并把剩余缺口重编号为 1–4;同时写清它不主张的部分——结果只说「分配已记录,执行进度请查看目标」,所以第 5 拍(谁能真正干活)仍然是缺口 1。
  • 防复发:这张表在一天内就漂移过一次,因此由场景自己检查——它读取两份用例文档,要求缺口小节的行数正好等于本次运行记录的缺口数,且不得保留已退役的确认结果措辞。
  • 验证:开发态与打包态(走已发布 /chat/ bundle)场景均通过,gap 列表一致;负向对照(英文表多一行)会以「lists 5 gaps but this run recorded 4 …」失败;docs 两个 smoke 通过;三个改动文件 loopx check --scan-path errors=0 且边界扫描干净。

…against the run

The product case still recorded gap 1 as "a confirmed plan does not
distinguish committed / partial / all-gap / stale / rejected per lane" with
the evidence "the only outcome sentence is the generic applied notice". That
stopped being true when the team-plan result became a per-lane record: the
scenario now finds `.personal-team-plan-result` naming the assigned lane and
its Todo beside the unstaffed lane and its reason, and its printed gap list
starts at `4-readiness`.

So the case now says what the surface actually proves, including the limit it
does not cross: the result says the assignment is recorded and points at the
Goal for progress, so beat 5 (who can actually work) stays gap 1 instead of
being absorbed into a proven confirm. The remaining gaps renumber to 1-4 in
both locales.

The table drifted within a day of the previous flip, so the scenario now
checks it: the gap section of each locale must list exactly one row per gap
the run recorded, and must not keep the retired confirm-result wording. The
run is the only thing that knows which gaps are open, and review is not a
mechanism that fires when a surface changes.

Validation: steward-journey in development and packaged modes (beats=4,
gaps=4-readiness,5-correction,6-recovery,7-return); the negative control (one
extra row in the English table) fails with "lists 5 gaps but this run recorded
4 (4-readiness, 5-correction, 6-recovery, 7-return)"; docs-governance-smoke ok;
docs-asset-integrity-smoke ok; loopx check --scan-path over the three changed
files reports errors=0 and a clean boundary scan.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

用例文档是这条端到端旅程给 owner 看的读法,而它漂移了:beat 4 仍写着「已证明,但见缺口 1」,缺口表第一条仍是「确认后不区分 committed / partial / all-gap / stale / rejected」,证据是「只有一条通用的『已应用』提示」。事实上场景早就把这条记成 present:它打印的缺口清单从 4-readiness 开始,报告里 .personal-team-plan-result 会逐条点名已分配的 lane 与其 Todo,旁边是未派工的 lane 及其原因(待安排 · 尚未加入此目标)。文档比产品落后了一天,而这一天里没有任何机制会自己发现这件事。

改动思路

先说实话,再钉住它。说部分:中英两份用例都改为「确认结果已证明」,并把剩余缺口重编号为 1–4;同时写清它不主张的部分——结果只说明「分配已记录,执行进度请查看目标」,所以第 5 拍(谁能真正干活)仍然是缺口 1,不会被并进一个"已证明的确认"里冒充执行真相。

钉部分:这张表在一天内就漂移过一次,因为另一个车道交付了确认结果而文档保留了旧结论;评审不会在某个面变化时自动触发,但这次运行知道它记录了哪些缺口。所以场景现在自己读两份用例文档,要求缺口小节的行数正好等于本次运行记录的缺口数,并要求已退役的确认结果措辞不再出现。

具体改动

  • examples/personal-workspace-browser/steward-journey.mjs:新增用例文档一致性检查——按 ## Recorded Gaps And Owners / ## 已记录缺口与归属 定位缺口小节,比对行数与本次 gaps 中 status === "gap" 的数量,并断言两份文档都不再包含旧结论措辞。
  • docs/product/use-cases/steward/README.md、README.zh-CN.md:beat 4 改为「已证明」并写明结果逐条点名已分配与未派工项及其原因;缺口表删除原缺口 1、其余重编号 1–4;正文交叉引用(含「缺口 5 关闭前」)同步;新增一段说明这次翻转与其边界。
  • 没有改产品代码:apps/presentation/dashboard 与 loopx/web/chat 在本 PR 里逐字节未动。

对主干的风险

需要说清这次断言的边界:它证明的是文档与本次运行记录的缺口集合一致,不是"缺口一定描述得准"——真正的准确性仍由场景对每个面的探针负责(这也是为什么探针证据留在报告里)。另一个风险是过度收紧:如果有人把缺口拆成两行来写,检查会红。这是有意的取舍——缺口表的行就是 owner 的阅读单位,拆行应当是一次显式决定,而不是无声的编辑。

负向对照做过了:给英文缺口表临时加一行,运行以 docs/product/use-cases/steward/README.md lists 5 gaps but this run recorded 4 (4-readiness, 5-correction, 6-recovery, 7-return) 失败;删掉这行后回到绿色。开发态与打包态(走已发布的 /chat/ bundle)都跑到 beats=4,gap 列表一致为 4-readiness,5-correction,6-recovery,7-return。文档侧 docs-governance-smoke ok、docs-asset-integrity-smoke ok,三个改动文件的 loopx check --scan-path 为 errors=0 且边界扫描干净。

边界声明:本 PR 只动 docs/** 与 examples/**,不含 loopx/**、apps/** 或 packages/**,也未改权限、状态契约或 CLI 契约,因此落在仓库允许自合并的小范围改动里。

我的整体评价

把"产品已经做到、用例还说是缺口"的漂移收口,同时明确不把"分配已记录"读成"执行已证实",方向是对的;更重要的是这次的修复不只改文本,而是让运行本身承担一致性检查,避免同一张表第三次静默漂移。建议在该 head 的必过检查全绿的前提下合并。

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Head: 4e684ba

English verdict: APPROVE

…nfirm-slice-20260919

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

审阅对象:PR #4716(开放中,未合并),exact head 4e684ba6bf0abcd67ba394b9e4feab5ee468ad6d(作者 huangruiteng)。本文是该 exact head 的评审记录。

动机

管家产品用例(docs/product/use-cases/steward/README.md 与中文版)是这个旅程"到底证明了什么"的 owner 面向读法,而它漂了:确认结果早就能逐条说明分配了哪些工作、分给谁、哪些没派出去以及原因,用例表却仍把"确认后不区分 committed / partial / all-gap / stale / rejected"记为缺口 1,证据写着"只有一条通用的『已应用』提示"。漂移还是静默的——没有任何东西比对用例表与场景运行记录,所以一个 lane 可以把面往前推,用例照旧说那是缺口。

改动思路

两件事一起做,缺一不可:把确认结果从"缺口"改成"已证明"并重新编号(缺口 1→4 依次前移,第 5 拍仍保留为缺口 1,因为它只证明分配、不证明执行),同时让运行本身成为权威——场景在跑完七拍后读两份用例文件,断言其中的缺口行数等于本次运行记录的 gap 拍数,且已退役的那句话不再出现。只改文案,下次还会漂;只加断言,表格会继续描述一个已被证明存在的缺口。

具体改动

3 个文件、+55/-20。两份用例文档:第 4 拍改为"已证明"并写明结果会点名已分配与未派工的工作及原因、且不主张执行进度;表格由 5 条缺口改为 4 条并加一段说明为什么旧的确认结果不再是缺口。examples/personal-workspace-browser/steward-journey.mjs:新增 24 行,在 finally 之前读取两份用例、定位"已记录缺口与归属/Recorded Gaps And Owners"小节、按 | N | 统计行数与本次 gaps.filter(status==='gap') 的拍数比较,并断言退役句子不存在(repoRoot 已由 fixture.mjs 导出,导入已补)。

验证:LOOPX_PERSONAL_WORKSPACE_PACKAGED=1 LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey node examples/personal-workspace-browser-smoke.mjs → beats=4 gaps=4-readiness,5-correction,6-recovery,7-return + personal-workspace-browser-smoke (packaged): ok,即第 3 拍确认结果现在是 present,文档 4 条正好对上。我用两次变异确认断言真的会咬:往英文表加第 5 行 → 失败信息为 ... README.md lists 5 gaps but this run recorded 4 (...);把退役句子塞回小节 → 失败信息为 ... still keeps the confirm-result gap this run proves present(两处均已还原)。另外离线复现了同一段提取逻辑(两份文件各 4 行、无退役句子),examples/docs-governance-smoke.py → ok,PR 检查 7 pass / 2 pending / 2 skipped、无失败。

顺带核对文档所述与已发布面一致:确认结果文案确实来自已发布的 i18n(proposal.teamPlan.appliedPartially "已分配 {created} 项,{gaps} 项待安排"、laneUnstaffed、gapReason.*),而"只记录分配、执行进度请看目标"对应 proposal.teamPlan.assignedHint,所以"仍不主张执行"这句是有出处的,不是新加的主张。

对主干的风险

纯文档 + 示例场景断言,无运行时、权限、配额或持久状态影响,回滚即回退三个文件。唯一的结构性弱点是新断言只比数量、不比身份:小节里的 | N | 行数与本次 gap 拍数相等即可,因此把 readiness 与 correction 两行对调(数量不变)仍会通过,而一次合法的表格重构(例如把某个拍的两行合并、数量变化)会误报失败。这是"文档读法"的脆弱性,不是产品风险;失败信息本身写清了文件、行数与拍号,修复路径明确。若要收紧,改成把每一行映射到 record 里的拍 id,而不是比总数。

证据边界:本机只跑了 packaged 模式(未跑开发服务器模式),且该旅程按设计用合成 fixture 替换 agent turn——它驱动的是真实打包工作区,但断言对象是文档而非产品路径本身。

我的整体评价

结论 APPROVE。这是一次小而完整的"让文档跟上已交付面"的收尾:它没有把缺口删掉了事,而是先说明为什么那条不再是缺口(分配已点名、执行仍未主张),再把这件事交给运行来守——正因为它已经漂过一次,把权威放在场景运行而不是评审记忆里是对的。断言会咬(两次变异都按预期失败),双语表格保持一致,文档治理 smoke 通过,且新增的 24 行沿用既有场景与报告,没有引入第二个 runner 或新的夹具。合并仍归维护者;若要继续加厚,建议把"行数相等"升级为"每行对应哪一拍"的映射断言。

English verdict: APPROVE - Review of open PR #4716 at exact head 4e684ba (author-owned; recorded as a COMMENTED approval because GitHub blocks formal self-approval). The steward product case had drifted: it still listed the confirm result as gap 1 with the evidence "the only outcome sentence is the generic applied notice", even though the confirm surface now names each assigned lane and each unstaffed item with its reason and explicitly points at the Goal for execution progress. The PR retires that gap, renumbers both language editions to four gaps, explains why beat 5 remains open, and adds a run-time check so the scenario run is the authority: after driving the seven beats it reads both case files, requires the gap-row count to equal the run's recorded gap beats, and rejects the retired sentence. I ran the packaged journey at the exact head (beats=4 gaps=4-readiness,5-correction,6-recovery,7-return, smoke ok), confirmed both language files carry four rows with no stale sentence, verified the documented claims against the shipped i18n keys (appliedPartially, laneUnstaffed, gapReason.*, and assignedHint for "Assignment recorded. See the Goal for execution progress."), ran docs-governance-smoke (ok), and proved the new assertions bite with two mutations (an extra gap row and the re-inserted retired sentence each fail the run with the intended message). Residual weakness, non-blocking: the check compares gap-row counts rather than mapping each row to its recorded beat, so a same-count reordering would pass and a legitimate restructure could fail. Evidence limit: only the packaged variant was exercised, and the journey uses a synthetic agent turn by design. No blocking finding. Merges remain with the maintainer.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

动机

用例文档是这条端到端旅程给 owner 看的读法,而它漂移了:beat 4 仍写着「已证明,但见缺口 1」,缺口表第一条仍是「确认后不区分 committed / partial / all-gap / stale / rejected」,证据是「只有一条通用的『已应用』提示」。事实上场景早就把这条记成 present:它打印的缺口清单从 4-readiness 开始,报告里 .personal-team-plan-result 会逐条点名已分配的 lane 与其 Todo,旁边是未派工的 lane 及其原因(待安排 · 尚未加入此目标)。文档比产品落后了一天,而这一天里没有任何机制会自己发现这件事。

改动思路

先说实话,再钉住它。说部分:中英两份用例都改为「确认结果已证明」,并把剩余缺口重编号为 1–4;同时写清它不主张的部分——结果只说明「分配已记录,执行进度请查看目标」,所以第 5 拍(谁能真正干活)仍然是缺口 1,不会被并进一个"已证明的确认"里冒充执行真相。

钉部分:这张表在一天内就漂移过一次,因为另一个车道交付了确认结果而文档保留了旧结论;评审不会在某个面变化时自动触发,但这次运行知道它记录了哪些缺口。所以场景现在自己读两份用例文档,要求缺口小节的行数正好等于本次运行记录的缺口数,并要求已退役的确认结果措辞不再出现。

具体改动

  • examples/personal-workspace-browser/steward-journey.mjs:新增用例文档一致性检查——按 ## Recorded Gaps And Owners / ## 已记录缺口与归属 定位缺口小节,比对行数与本次 gaps 中 status === "gap" 的数量,并断言两份文档都不再包含旧结论措辞。
  • docs/product/use-cases/steward/README.md、README.zh-CN.md:beat 4 改为「已证明」并写明结果逐条点名已分配与未派工项及其原因;缺口表删除原缺口 1、其余重编号 1–4;正文交叉引用(含「缺口 5 关闭前」)同步;新增一段说明这次翻转与其边界。
  • 没有改产品代码:apps/presentation/dashboard 与 loopx/web/chat 在本 PR 里逐字节未动。

对主干的风险

需要说清这次断言的边界:它证明的是文档与本次运行记录的缺口集合一致,不是"缺口一定描述得准"——真正的准确性仍由场景对每个面的探针负责(这也是为什么探针证据留在报告里)。另一个风险是过度收紧:如果有人把缺口拆成两行来写,检查会红。这是有意的取舍——缺口表的行就是 owner 的阅读单位,拆行应当是一次显式决定,而不是无声的编辑。

负向对照做过了:给英文缺口表临时加一行,运行以 docs/product/use-cases/steward/README.md lists 5 gaps but this run recorded 4 (4-readiness, 5-correction, 6-recovery, 7-return) 失败;删掉这行后回到绿色。开发态与打包态(走已发布的 /chat/ bundle)都跑到 beats=4,gap 列表一致为 4-readiness,5-correction,6-recovery,7-return。文档侧 docs-governance-smoke ok、docs-asset-integrity-smoke ok,三个改动文件的 loopx check --scan-path 为 errors=0 且边界扫描干净。

边界声明:本 PR 只动 docs/** 与 examples/**,不含 loopx/**、apps/** 或 packages/**,也未改权限、状态契约或 CLI 契约,因此落在仓库允许自合并的小范围改动里。

我的整体评价

把"产品已经做到、用例还说是缺口"的漂移收口,同时明确不把"分配已记录"读成"执行已证实",方向是对的;更重要的是这次的修复不只改文本,而是让运行本身承担一致性检查,避免同一张表第三次静默漂移。建议在该 head 的必过检查全绿的前提下合并。

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Head: bd45c6a

English verdict: APPROVE

@huangruiteng
huangruiteng merged commit 96364a3 into main Sep 18, 2026
25 checks passed
@huangruiteng
huangruiteng deleted the codex/steward-case-confirm-slice-20260919 branch September 18, 2026 19:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant