Skip to content

Commit e614af4

Browse files
committed
Merge remote-tracking branch 'origin/main' into feat/track-a-single-source-todo-task-class
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
2 parents e385e0e + 6b3264f commit e614af4

8 files changed

Lines changed: 577 additions & 19 deletions

File tree

‎docs/product/use-cases/README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,3 +7,4 @@ They do not grant domain providers core control-plane authority.
77
- [Issue and PR work](issue-pr/README.md)
88
- [Cross-runtime implementation and review](cross-runtime/README.md)
99
- [Office operations](office-operations/README.md)
10+
- [Steward: an owner sentence becomes a confirmed team](steward/README.md)
Lines changed: 90 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,90 @@
1+
# Steward: An Owner Sentence Becomes A Confirmed Team
2+
3+
Status: qualification case. It records what the local steward journey proves on
4+
the workspace today, which beats are still unproven, and how to reproduce both.
5+
It is product guidance, not a new capability, contract or scheduler.
6+
7+
The case is grounded in one deterministic browser scenario
8+
(`examples/personal-workspace-browser/steward-journey.mjs`) that runs on synthetic
9+
data. The fixture substitutes the agent turn; everything the case describes is a
10+
fact the workspace surfaces render, not a claim about a live Goal.
11+
12+
## When This Case Applies
13+
14+
- an owner wants work to start from one sentence instead of a filled-in form;
15+
- the work needs more than one Agent or more than one lane, so staffing is
16+
itself part of the answer;
17+
- the owner wants to keep confirming, correcting and reading results in one
18+
place instead of relaying between Agent conversations.
19+
20+
## The Journey
21+
22+
| Beat | What the owner does | What the workspace shows | State |
23+
| --- | --- | --- | --- |
24+
| 1 | Looks at the first screen | Goal board lanes (needs you / running / observing / scheduled), each Goal card naming its Agent and its next sentence | Proven |
25+
| 2 | Asks the steward in the Goal conversation | The ask becomes an accepted Turn and the admitted team plan card lands in the same conversation | Proven |
26+
| 3 | Reads the card | Per lane: the Agent, the first bounded Todo with priority and action kind, the acceptance signal, and an explicitly unstaffed lane that keeps the work it did not staff; the quota envelope and stop condition; a statement that confirming is what creates the lanes | Proven |
27+
| 4 | Confirms | Exactly one apply and one durable write; the card reports that LoopX state will refresh | Proven, but see gap 2 |
28+
| 5 | Checks who can actually work | — | Gap 3 |
29+
| 6 | Corrects or pauses one lane | — | Gap 4 |
30+
| 7 | Waits for a lane to fail and asks who fixes it / judges completion | — | Gaps 5, 6 |
31+
32+
Beats 5–7 are recorded by the scenario as typed gaps with the probe that looked
33+
for them. They are not "not implemented here" hand-waving: the scenario names
34+
the selectors and phrases it searched for and what it found instead.
35+
36+
## Patterns
37+
38+
1. **Ask for an outcome, not an org chart.** One sentence with the outcome and
39+
the constraint produces a plan card; naming Agents before the outcome turns
40+
coordination into the owner's job.
41+
2. **Read four facts before confirming.** Agent, first bounded Todo, acceptance
42+
signal and staffing gap. A card that cannot show a gap is not yet reviewable.
43+
3. **Treat the gap lane as information, not failure.** An unstaffed lane keeps
44+
the work it could not staff and names the reason, so the owner can decide to
45+
drop it, staff it, or accept partial delivery.
46+
4. **Confirmation is a durable write.** Confirming sends exactly one apply and
47+
performs one durable write; the surface must not claim a lane exists before
48+
that write, and must say what the write produced afterwards.
49+
5. **Judge delivery by the returned result, not by the conversation.** A reply
50+
or a message is not a completed lane. Until gap 6 closes, treat the
51+
conversation as the request channel and the Goal's own state as the truth.
52+
6. **Correct in the conversation the work came from.** Steering an active run is
53+
supported today; correcting a confirmed lane commitment is not yet, so avoid
54+
confirming a plan whose lanes may need to be withdrawn.
55+
56+
## Reproduce
57+
58+
```sh
59+
# development surfaces
60+
LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey \
61+
node examples/personal-workspace-browser-smoke.mjs
62+
63+
# packaged workspace bundle
64+
LOOPX_PERSONAL_WORKSPACE_PACKAGED=1 \
65+
LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey \
66+
node examples/personal-workspace-browser-smoke.mjs
67+
```
68+
69+
The run writes `steward-journey-report.json` (beats, gaps, probe evidence) and
70+
per-beat screenshots under `output/playwright/personal-workspace/`, which is
71+
gitignored. No live Goal, Agent, credential or local path is read or captured.
72+
73+
## Recorded Gaps And Owners
74+
75+
| # | Gap | Evidence the scenario recorded | Owner surface |
76+
| --- | --- | --- | --- |
77+
| 1 | The steward's bounded prompt set (`找下一步` / `看阻塞` / `查证据`) is defined in the client model but not reachable from the conversation | probe: no steward-prompt element, no prompt phrases before the owner types | workspace composer |
78+
| 2 | A confirmed plan does not distinguish committed / partial / all-gap / stale / rejected per lane | probe: the only outcome sentence is the generic applied notice | steward plan commit (roadmap R1 remainder) |
79+
| 3 | No per-lane readiness ladder (registered → bound → launchable → executing) | probe: no lane-readiness element or phrase | steward readiness (roadmap R2 / audit F6) |
80+
| 4 | No lane-level correction (pause or supersede a confirmed commitment) | probe: no lane-correction element; only run steering exists | shared alignment (roadmap R4) |
81+
| 5 | A failed lane does not name its blocker owner and next step | probe: no lane-blocker element or phrase | recovery/continuation (roadmap R3) |
82+
| 6 | Completion is not judged by the lane's returned result | probe: no lane-return element or phrase | return delivery (roadmap R3) |
83+
84+
## What This Case Does Not Claim
85+
86+
- It does not qualify a live steward conversation: the fixture substitutes the
87+
agent turn, so the model/runtime behind the intake stays untested here.
88+
- It does not qualify Lark audiences or any cloud/remote worker.
89+
- It does not turn a passing smoke into product acceptance for a Goal whose
90+
plan was confirmed with real consequences.
Lines changed: 66 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,66 @@
1+
# 管家:一句话变成一个被确认的团队
2+
3+
状态:qualification 案例。它记录本机管家旅程今天在工作台面上证明了什么、哪几拍还没有证明,以及如何复现两者。它是产品侧使用说明,不是新能力、新契约或新调度器。
4+
5+
本文案的事实来源是一个确定性的浏览器场景(`examples/personal-workspace-browser/steward-journey.mjs`),全部跑在合成数据上。fixture 顶替了 agent turn;本文描述的每一句都是工作台面渲染出来的事实,不是对某个真实 Goal 的断言。
6+
7+
## 什么情况下适用
8+
9+
- 老板想用一句话启动工作,而不是填表;
10+
- 这项工作需要不止一个 Agent 或不止一条 lane,所以“有没有人干”本身就是答案的一部分;
11+
- 老板想在同一个地方确认、纠偏、看结果,而不是在多个 Agent 会话之间来回转发。
12+
13+
## 旅程七拍
14+
15+
| 拍 | 老板做什么 | 界面显示什么 | 状态 |
16+
| --- | --- | --- | --- |
17+
| 1 | 看首屏 | Goal 看板四条 lane(需要你/执行中/观察中/已安排),每张 Goal 卡给出 Agent 与下一步那句话 | 已证明 |
18+
| 2 | 在 Goal 对话里向管家提要求 | 这句话成为一个被接受的 Turn,被准入的团队计划卡落在同一个对话里 | 已证明 |
19+
| 3 | 阅读计划卡 | 每条 lane 的 Agent、第一刀 Todo(含优先级与 action kind)、验收信号,以及一条明确“未配齐”并保留未派工工作的 lane;配额包络与停止条件;以及“确认才会建 lane”的说明 | 已证明 |
20+
| 4 | 确认 | 恰好一次 apply、一次 durable write;卡片提示 LoopX 状态将刷新 | 已证明,但见缺口 2 |
21+
| 5 | 检查到底谁能干活 | — | 缺口 3 |
22+
| 6 | 暂停或撤销某条 lane | — | 缺口 4 |
23+
| 7 | 等某条 lane 失败,问谁负责修 / 用什么判定完成 | — | 缺口 5、6 |
24+
25+
第 5–7 拍由场景以 typed gap 记录,并带上“探针找过什么”的证据:不是“这里先不做”的一句话,而是列出了查找的选择器和文本、以及实际找到什么。
26+
27+
## 推荐姿势
28+
29+
1. **说要结果,不要点将。** 一句话给出结果与约束,让计划卡来回答“谁来做”;先点名 Agent 会把协调变成老板的活。
30+
2. **确认前先看四件事**:Agent、第一刀 bounded Todo、验收信号、缺人情况。看不到缺口的卡片还不具备可评审性。
31+
3. **缺人 lane 是信息,不是失败。** 未配齐的 lane 会保留它没能派出去的工作并给出原因,老板可以选择砍掉、补人、或接受部分交付。
32+
4. **确认是一次落地的 durable 写入。** 确认只发一次 apply、只做一次 durable write;卡片不能在写入前声称 lane 已存在,写入后必须说明产生了什么。
33+
5. **用回传结果判定交付,不要用对话判定。** 一条回复不等于一条 lane 完成。在缺口 6 关闭前,把对话当请求通道,把 Goal 自身状态当事实。
34+
6. **在产生工作的那个对话里纠偏。** 对运行中的 Turn 纠偏今天已支持;对已确认 lane 承诺的纠偏还没有,所以不要确认一张可能需要撤回 lane 的计划。
35+
36+
## 如何复现
37+
38+
```sh
39+
# 开发态台面
40+
LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey \
41+
node examples/personal-workspace-browser-smoke.mjs
42+
43+
# 打包态工作台
44+
LOOPX_PERSONAL_WORKSPACE_PACKAGED=1 \
45+
LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey \
46+
node examples/personal-workspace-browser-smoke.mjs
47+
```
48+
49+
运行会在 `output/playwright/personal-workspace/`(已 gitignore)下写出 `steward-journey-report.json`(拍子、缺口、探针证据)与每拍截图。不读取、不截取任何真实 Goal、Agent、凭证或本地路径。
50+
51+
## 已记录缺口与归属
52+
53+
| # | 缺口 | 场景记录的证据 | 归属面 |
54+
| --- | --- | --- | --- |
55+
| 1 | 管家快捷提示(找下一步 / 看阻塞 / 查证据)只定义在客户端模型里,对话里点不到 | 探针:老板输入前既无 steward-prompt 元素,也无提示文本 | 工作台输入区 |
56+
| 2 | 确认后不区分 committed / partial / all-gap / stale / rejected | 探针:只有一条通用的“已应用”提示 | 管家计划落地(roadmap R1 剩余项) |
57+
| 3 | 没有 per-lane readiness 阶梯(registered → bound → launchable → executing) | 探针:无 lane-readiness 元素或文本 | 管家 readiness(roadmap R2 / 审计 F6) |
58+
| 4 | 没有 lane 级纠偏(暂停或撤销已确认承诺) | 探针:无 lane-correction 元素;只有运行中 Turn 的纠偏 | shared alignment(roadmap R4) |
59+
| 5 | lane 失败后不说明阻塞归属与下一步 | 探针:无 lane-blocker 元素或文本 | 恢复与继续(roadmap R3) |
60+
| 6 | 不用 lane 的回传结果判定完成 | 探针:无 lane-return 元素或文本 | 交付回收(roadmap R3) |
61+
62+
## 本文案不主张什么
63+
64+
- 它不资格化真实管家对话:fixture 顶替了 agent turn,因此接入背后的模型/运行时在这里仍未测试;
65+
- 它不资格化飞书受众,也不资格化任何云端/远端 worker;
66+
- 它不把一条通过的 smoke 当成“某个真实确认过的 Goal 已被产品验收”。

‎examples/personal-workspace-browser-smoke.mjs‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -23,10 +23,11 @@ import {
2323
} from "./personal-workspace-browser/fixture.mjs";
2424
import { navigationSortingScenario } from "./personal-workspace-browser/navigation-sorting.mjs";
2525
import { progressiveLoadingScenario } from "./personal-workspace-browser/progressive-loading.mjs";
26+
import { stewardJourneyScenario } from "./personal-workspace-browser/steward-journey.mjs";
2627
import { teamPlanScenario } from "./personal-workspace-browser/team-plan.mjs";
2728
import { typedActionsScenario } from "./personal-workspace-browser/typed-actions.mjs";
2829

29-
const scenarioCatalog = [navigationSortingScenario, chatRecoveryScenario, typedActionsScenario, teamPlanScenario, executionChipScenario, progressiveLoadingScenario];
30+
const scenarioCatalog = [navigationSortingScenario, chatRecoveryScenario, typedActionsScenario, teamPlanScenario, stewardJourneyScenario, executionChipScenario, progressiveLoadingScenario];
3031
const requestedScenario = process.env.LOOPX_PERSONAL_WORKSPACE_SCENARIO;
3132
const scenarios = requestedScenario
3233
? scenarioCatalog.filter((scenario) => scenario.id === requestedScenario)

‎examples/personal-workspace-browser/execution-chip.mjs‎

Lines changed: 65 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -94,6 +94,64 @@ async function chipText(page) {
9494
return (await page.locator(".personal-execution-chip").innerText()).replace(/\s+/g, " ").trim();
9595
}
9696

97+
const pickerSelector = "div.personal-agent-select button.personal-select-trigger";
98+
99+
async function readPickerLabel(page) {
100+
return (await page.locator(pickerSelector).innerText()).replace(/\s+/g, " ").trim();
101+
}
102+
103+
// What the page resolved, phrased so a failed CI run is diagnosable without a
104+
// local reproduction: the declared endpoint, the adapters the page actually
105+
// received and the label it rendered are the three facts that decide the label.
106+
async function pickerResolution(page) {
107+
const capabilities = await page.evaluate(async () => {
108+
try {
109+
const response = await fetch("/api/chat/capabilities");
110+
const body = await response.json();
111+
return {
112+
status: response.status,
113+
declaredEndpoint: body.manager?.channel_binding?.executor_endpoint ?? null,
114+
adapters: (body.adapters ?? []).map(
115+
(adapter) => `${adapter.agent_id}:${adapter.available ? "available" : "unavailable"}`,
116+
),
117+
};
118+
} catch (error) {
119+
return { status: "unavailable", declaredEndpoint: null, adapters: [], error: String(error) };
120+
}
121+
});
122+
return [
123+
`label=${await readPickerLabel(page)}`,
124+
`declared=${capabilities.declaredEndpoint ?? "<none>"}`,
125+
`adapters=[${capabilities.adapters.join(", ")}]`,
126+
`capabilities=${capabilities.status}${capabilities.error ? ` (${capabilities.error})` : ""}`,
127+
].join(" ");
128+
}
129+
130+
// The picker resolves from the same capabilities response that carries the
131+
// channel binding, and the execution chip only renders once that binding
132+
// lands. A single read therefore asserts on a state the page never promised was
133+
// settled, and it can still hold the pre-fetch Codex fallback. Wait for the
134+
// declared value instead, and name what the page really resolved if the wait
135+
// runs out.
136+
async function waitForPickerLabel(page, settled, timeoutMs = 15_000) {
137+
const deadline = Date.now() + timeoutMs;
138+
let label = await readPickerLabel(page);
139+
while (!settled(label) && Date.now() < deadline) {
140+
await page.waitForTimeout(100);
141+
label = await readPickerLabel(page);
142+
}
143+
if (settled(label)) {
144+
return label;
145+
}
146+
const resolution = await pickerResolution(page);
147+
await page.screenshot({
148+
animations: "disabled",
149+
fullPage: false,
150+
path: resolve(outputDir, "execution-chip-picker-unresolved.png"),
151+
});
152+
throw new Error(`Chat runtime picker never settled: ${resolution}`);
153+
}
154+
97155
async function assertHairlineRow(page) {
98156
const headerBox = await page.locator(".personal-channel-header").boundingBox();
99157
const chipBox = await page.locator(".personal-execution-chip").boundingBox();
@@ -214,12 +272,10 @@ export const executionChipScenario = {
214272
collectCoverage,
215273
});
216274
try {
217-
const pickerLabel = (await stewardPicker.page
218-
.locator("div.personal-agent-select button.personal-select-trigger")
219-
.innerText()).replace(/\s+/g, " ").trim();
220-
if (!pickerLabel.includes("DeepSeek Harness (managed)")) {
221-
throw new Error(`Chat runtime picker ignored the declared steward executor: ${pickerLabel}`);
222-
}
275+
const pickerLabel = await waitForPickerLabel(
276+
stewardPicker.page,
277+
(label) => label.includes("DeepSeek Harness (managed)"),
278+
);
223279
if (pickerLabel.includes("Codex")) {
224280
throw new Error(`Chat runtime picker advertised a discovered CLI as the steward: ${pickerLabel}`);
225281
}
@@ -242,12 +298,9 @@ export const executionChipScenario = {
242298
collectCoverage,
243299
});
244300
try {
245-
const pickerLabel = (await undeclaredSteward.page
246-
.locator("div.personal-agent-select button.personal-select-trigger")
247-
.innerText()).replace(/\s+/g, " ").trim();
248-
if (pickerLabel !== "Chat Codex") {
249-
throw new Error(`An undeclared steward executor no longer used the shipped default: ${pickerLabel}`);
250-
}
301+
// The shipped default is the settled value a machine with no declared
302+
// steward executor must keep.
303+
await waitForPickerLabel(undeclaredSteward.page, (label) => label === "Chat Codex");
251304
} finally {
252305
coverageEntries.push(...await undeclaredSteward.close());
253306
}

0 commit comments

Comments
 (0)