diff --git a/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.md b/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.md index 7becd26656..ac6d6b9113 100644 --- a/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.md +++ b/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.md @@ -221,7 +221,9 @@ Do not put the entire handoff inside Goal Vision's bounded summary or expand eve ### 5.5 Responsibility discovery and receiver-owned planning -Discover all registered active Goals and their agents within authorized host scope. Stopped Goals are excluded by default but can be requested. Discoverability, read access, context delivery and execution permission are four separate facts. “Best effort” means the manager investigates role, current work, repository and availability; it does not mean broadcasting private context or guessing an identity. +**Broad discovery is the owner default.** The private steward discovers the owner's registered Goals and Agents across the local host and already-authorized connected sources, without requiring per-recipient enrollment just to find them. Search and pagination must reach the full permitted inventory; a small prompt or first-page limit is not a smaller discovery scope. Keep stopped Goals available through history/search, while excluding them from default work assignment. Offline, unbound, capacity-limited and unknown-presence Agents remain discoverable with their actual state. + +Discoverability, evidence read access, context-delivery grants and execution readiness are separate facts. A narrow delivery allowlist must not become the owner discovery catalog or justify “no Agent exists.” A discovered relevant owner without a delivery grant is a specific delegation gap, with an existing configuration/recovery route; missing runtime evidence is an unknown-readiness gap. Reuse the broad authorized inventory for investigation, then check the requested effect separately. Shared/external audiences receive only their authorized metadata and evidence; owner-private coverage does not implicitly become group-visible coverage. Do not broadcast private context, discover arbitrary unconnected hosts or expand execution authority through discovery. These are target defaults; the current recipient catalog alone does not implement them. The existing manager context-delivery catalog now excludes Goals marked stopped in its selected registry, and delivery rechecks that activation guard before creating or replaying an inbox request. A malformed activation state cannot grant a recipient. This is one admission boundary only: a lagging global mirror still needs source-authority reconciliation, and registration does not establish a live session, executable capacity, suitable model or a returned result. Explicit inspection of a stopped Goal remains separate from delivering it new work. @@ -506,7 +508,13 @@ The shared [conversation work surface](intelligent-review-presentation-surfaces- Routing uses §5.5 rather than a static Agent name list. For a product-design request addressed to the steward, first inspect authorized current Goal, registration, claimed work and fresh session reachability; then rank eligible receivers by responsibility and context, with model/profile fit and actual capacity as separate constraints. Explain the selected recipient or the exact gap. The receiver must acknowledge and assess the full corrected intent, then either work or defer with an owner and condition. The original conversation receives the assessment and final evidenced result through §5.6; the catalog, a stored inbox request and a spinner are three distinct incomplete states. This must pass with a real active worker plus stopped, registered-only, stale and model-mismatched decoys before advertising automatic delegation. -Delivery order is: (1) qualify the reusable report, activity and stop/steer surface in Chat, frontend and Lark; (2) add context-affine steward selection through the existing collaboration owner; (3) prove one original-conversation request through real receiver assessment, work and final return. These are independently reviewable slices; the steward product claim waits for the third. Characterize and retire duplicated answer-shape prose and message/Turn correlation rules where parity is proven. +Delivery now prioritizes one complete supported intent→receiver→work→result journey, with the shared report, activity and stop/steer behavior needed by that journey. Do not make routing wait for presentation polish across every channel. Frontend and Lark still need separate real acceptance before an equivalence claim. Characterize and retire duplicated answer-shape prose and message/Turn correlation rules where parity is proven. + +The [golden-query pack](../../product/use-cases/steward/golden-queries.md) supplies short user requests and independent outcome/attention oracles. GQ01/GQ02 qualify creation and existing-Agent connection; GQ03/GQ04 qualify responsible dispatch; GQ05/GQ11–GQ13 qualify two-cycle small-team coordination, with GQ07–GQ09 continuity; GQ06/GQ10/GQ14–GQ15 extend materials, attention and replanning. GQ16/GQ17 remain later cross-host/scale qualification. These are scenario tests over existing A1–A24, not new Core protocol states. + +When routing fails, distinguish unread/unavailable sources, incomplete or stale directory coverage, no relevant registered owner, unauthorized scope, missing binding, unknown runtime readiness, capacity wait and receiver rejection. Probe/refresh permitted sources and repair an eligible binding through its owner before asking the user to locate an Agent. Registration grants neither reachability nor authority. If no existing receiver qualifies, an already-authorized creation path is valid; otherwise retain the request and ask only for the concrete missing decision. Never select an irrelevant sole candidate or silently replace an explicitly requested model. A well-written recommendation with no requested dispatch is still an undelivered task. + +Ordinary follow-ups retain the responsible owner; direct small reads need no team. One binding retains one execution driver. Existing request/outbox and continuous-monitor owners handle event-driven continuation and delayed return; do not compensate with faster polling. Retain useful constraints through scoped context, distinguishing preference, current fact and action grant; optional memory providers remain optional. GQ03's CI example requires causal evidence and all other review blockers, not a blanket approve policy. ## 6. Alternatives and disposition of #4306 diff --git a/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.zh-CN.md b/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.zh-CN.md index 8fa9d3215c..ca83923250 100644 --- a/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.zh-CN.md +++ b/docs/architecture/rfcs/capable-manager-semantic-handoff-v0.zh-CN.md @@ -206,7 +206,9 @@ LoopX 不是只有任务队列。交接应让接收方结合权威状态和持 ### 5.5 职责发现与接收方重规划 -在已授权主机范围发现所有注册的活跃 Goal 和 Agent。默认排除停止的 Goal,用户明确询问时可读。能发现、能读、能投递上下文、能执行,是四种不同事实。best effort 指管家认真检查职责、当前工作、仓库和可用性;不是群发私人背景或猜身份。 +**主人管家默认广泛发现。** 私有管家应发现本机以及已授权接入来源中,主人已登记的 Goal 和 Agent;仅为了找到负责人,不要求逐个加入委派名单。搜索与分页可达完整的允许目录,prompt 或首屏条数上限不能缩小发现范围。停止的 Goal 保留历史/搜索入口,默认不参与工作指派;离线、未绑定、容量不足和在线状态未知的 Agent 仍可发现,并标明实际状态。 + +能发现、能读证据、能委派、当前能执行,是不同事实。窄委派名单不能充当主人的发现目录,更不能据此回答“没有 Agent”。发现相关负责人但未获委派授权时,明确说明委派缺口及既有配置/恢复入口;缺运行证据则说明就绪状态未知。先复用广泛的已授权目录调查,再按请求效果判断准入。共享/外部受众只能看到该受众已授权的元数据与证据;主人私有可见范围不会自动变成群聊可见范围。不群发私人背景、不探测任意未接入主机,不通过发现扩张执行权限。这是目标默认行为,当前接收者名单本身尚未实现。 优先用户指定接收方;否则基于当前状态选合适责任人。缺少便捷 routing profile 不应让一个已知、已授权的 worker 变得不存在。不能把固定请求目录当唯一职责模型。多个接收者适合时,先选一个评估负责人并说明依据;仅在歧义实质影响权限或结果时询问。没有合适 worker,则自己做允许的工作,或报告真实能力缺口;不偷偷新建用户任务或唤醒已停止 Goal。 @@ -471,7 +473,13 @@ Stage A 替换当前 Todo note,不提供不可变历史版本或私有 memory 路由沿用 §5.5,而非固定 Agent 名单。管家收到产品设计请求时,先查看获授权的当前 Goal、注册、已认领工作及新鲜的 session 可达性;再按职责与上下文选择合格接收者,单独约束模型/profile 适配和实际容量。说明选择理由或真实缺口。接收者要确认、评估包含历史修正的完整意图,随后执行或明确延期条件和负责人。通过 §5.6 将评估与有证据的最终结果送回原对话;候选列表、已存入收件箱、正在处理是三种不同的未完成状态。对真实活跃 worker 与已停止、仅注册、过期、模型不适配的干扰候选都验收后,才能宣传自动委派。 -交付顺序:(1)在 Chat、前端和飞书验收可复用的报告、活动与停止/纠偏界面;(2)通过现有 collaboration owner 做好管家的上下文亲和选择;(3)让同一原始对话经过真实接收方评估、执行与最终回传。各阶段可独立审阅,管家产品主张要等第三阶段。等价行为刻画通过后,消除重复的回答格式说明和消息/Turn 归属规则。 +交付改为优先跑通一条已支持的意图→接收者→执行→结果旅程,同批带上它需要的共用报告、活动与停止/纠偏能力。路由不再等待所有渠道的展示打磨;宣称前端/飞书等价前,仍分别完成真实验收。等价行为刻画通过后,消除重复的回答格式说明和消息/Turn 归属规则。 + +[Golden-query 集](../../product/use-cases/steward/golden-queries.md)提供简短用户请求及独立的结果/注意力验收。GQ01/GQ02 验收创建与接入已有 Agent;GQ03/GQ04 验收找负责人派单;GQ05/GQ11–GQ13 验收两轮小团队协调,GQ07–GQ09 验收连续性;GQ06/GQ10/GQ14–GQ15 扩展材料、注意力和重排。GQ16/GQ17 保留为后续跨主机/规模验收。这些是既有 A1–A24 上的场景,不新增 Core 协议状态。 + +路由失败要区分:来源未读/不可用、目录不完整/过期、没有职责匹配的注册人、未获授权、缺少绑定、runtime 可执行性未知、容量等待、接收者拒绝。先在许可范围刷新来源、探测和通过原 owner 修复合格绑定,再要求用户定位 Agent;注册本身既不授权,也不证明可达。没有现成接收者时可沿已有授权的创建路径处理;否则保留请求,只询问确实缺失的决定。不误投给唯一但不相关的候选,也不暗中替换用户指定模型。只给出好看的建议却没有执行用户要求的委派,仍是未交付。 + +普通追问保留原负责人,简单查询直接处理;同一绑定只保留一个执行驱动。事件续接和延迟回传复用既有 request/outbox 与 continuous-monitor owner,不靠提高轮询频率补偿。通过有作用域的上下文复用约束,分别保存偏好、当前事实和动作授权;可选记忆 provider 不变成前置条件。GQ03 的 CI 例子要求查明因果和其他 review blocker,不是无条件 approve 政策。 ## 6. 备选与 #4306 裁决 diff --git a/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.md b/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.md index 72344e9aac..66d994da32 100644 --- a/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.md +++ b/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.md @@ -727,6 +727,24 @@ same identity, interruption, replay and audience-isolation cases at their own display densities. This shared contract reuses Chat/session, artifact, and presentation owners; it creates no second conversation store or scheduler. +**Attention-oriented return.** The original conversation distinguishes requested +results, routine progress and decisions needing the owner. Requested results +return normally; unchanged progress folds into a digest; material decisions show +the concrete object, recommendation, evidence and consequence of inaction. +Keep deeper evidence and the responsible Agent's conversation directly reachable. +A correction made there returns its relevant decision/work change to the steward +through the same request lineage, without copying the entire private dialogue. +Source coverage and unresolved work remain visible. An empty directory or a +saved answer with failed delivery cannot be presented as a successful conclusion. + +Creation and first submission are part of this shared surface: preserve the +message and its pending/failed state across navigation or reload, expose a safe +retry/readback, and distinguish Goal created, Agent connected and work started. +An optimistic disappearing composer is not an accepted request. The +[golden-query pack](../../product/use-cases/steward/golden-queries.md) checks these +entry states alongside the full conversation, in packaged frontend and each +claimed channel; it does not specialize activity presentation to reports. + Botmux is an interaction reference: its [live card](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs-site/docs/zh/cards.md) keeps final text ahead of collapsible recorded activity, its [session model](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs-site/docs/zh/session-model.md) distinguishes talk and operation rights, and its [Codex steering study](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs/design/2026-05-28-codex-type-ahead-steer-design.md) @@ -1065,6 +1083,14 @@ Measure both attention cost and outcome quality: - model-advice override, hallucination, over-escalation, and dangerous suppression rates. +The [golden-query evaluation](../../product/use-cases/steward/golden-queries.md) +operationalizes these measures with paired baseline/candidate tasks and separate +per-surface results. Count avoidable finding/context/repetition/chasing/relay work; +report legitimate authorization, goal changes and voluntary learning separately. +Include failed/abandoned attempts and unknown cost/coverage. Silence, a short +answer or fewer messages alone cannot improve the score. All live case outcomes +remain unqualified until their independent evidence exists. + Reducing clicks while lowering accepted outcome quality is a regression, not a success. diff --git a/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.zh-CN.md b/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.zh-CN.md index aad9063874..26227c77a7 100644 --- a/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.zh-CN.md +++ b/docs/architecture/rfcs/intelligent-review-presentation-surfaces-v0.zh-CN.md @@ -568,6 +568,10 @@ Adaptive policy 必须 inspectable、resettable,其输出携带 reason codes 停止和修正绑定当前 session 与 Turn,旧控件不能影响后来的 Turn。停止读回区分真实中断、已经结束、不支持或拒绝。忙时修正要么作为本轮原生 steering 接收,要么带可恢复入口回执明确排入后续回合;完成竞态与丢失 ACK 不能使它无声消失。停止对话 Turn 不隐式影响被委派 worker 的 Todo/lease 或待回传义务。前端与飞书以各自展示密度验收同一身份、中断、重放及受众隔离用例。此共用合同复用 Chat/session、工件和展示 owner,不新增第二套对话存储或调度器。 +**以注意力为中心回传。** 原对话区分用户请求的成果、日常进展、需要主人决定的事项:成果正常返回,未变化的进展合并进摘要,重要决定给出具体对象、建议、证据和不处理的影响。保留深入证据及负责人对话的直接入口;在负责人处形成的纠偏,通过同一请求关系将相关决定/工作变化带回管家,不复制整段私人对话。来源覆盖和未完成工作保持可见;空目录或已保存但投递失败的答案都不算成功交付。 + +创建与首次发送同样属于共用界面:切换页面或刷新后保留消息及 pending/failed 状态,提供安全重试/读回,区分 Goal 已创建、Agent 已接入、工作已开始。乐观清空输入框不等于请求已接收。[Golden-query 集](../../product/use-cases/steward/golden-queries.md)在打包前端及每个对外承诺的渠道检查这些入口和完整对话;中间过程展示不为报告特化。 + Botmux 是交互参考:[实时卡片](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs-site/docs/zh/cards.md)优先保留最终答复、折叠真实过程;[会话模型](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs-site/docs/zh/session-model.md)区分对话权与操作权;[Codex 纠偏研究](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs/design/2026-05-28-codex-type-ahead-steer-design.md)记录忙时消息会合并或分开回复。停止能力因后端而异。这些公开来源指导竞态和展示验收,不能替代 LoopX adapter 的真实资格,也不要求把 Botmux 装到已有飞书 token 上。 ## 9. 覆盖长程工作的完整生命周期 @@ -824,6 +828,8 @@ Attention cost 与 outcome quality 必须同时度量: - Agent throughput、acceptance quality 与 safety outcomes; - model-advice override、hallucination、over-escalation 与 dangerous suppression rate。 +[Golden-query evaluation](../../product/use-cases/steward/golden-queries.md)用成对基线/候选任务及逐入口结果落实这些指标。计量可避免的找人、补背景、重复、催办、搬运;必要授权、主动改目标和自愿学习分开报告。保留失败/放弃样本及未知成本/覆盖,不能靠沉默、短回答或少消息刷高分。真实案例在取得独立证据前均未验收。 + 减少点击但降低 accepted outcome quality 是回归,不是成功。 ## 15. Failure 与 Fallback Rules diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md index fe2c4b8550..70ce963899 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md @@ -348,6 +348,45 @@ Chat, session, collaboration and presentation owners as R2/R3 integration work. Do not call a candidate catalog, spinner or queued inbox receipt a completed worker handoff. This checkpoint does not lower R1–R3 or G1 gates. +### Attention-cost acceptance: create, connect, collaborate, understand + +The [public-safe golden-query pack](../../product/use-cases/steward/golden-queries.md) +is the next product evaluation target, not a statement of current support. +Users should express a short outcome or correction without finding Agent IDs, +repeating known constraints, chasing work or transporting results themselves. +Keep valuable learning, deliberate choices and necessary approvals separate +from avoidable coordination. Preserve direct conversations with long-term owners. + +P0 starts at the entry: a creation request cannot disappear after submission; +Goal creation, reuse of an existing Agent and authorized new-worker setup have +observable outcomes. Then qualify R3/M2/M3 routing **alongside** R2 adoption: +responsible receiver, actual work, exact artifact use, independent acceptance, +and synthesized return to the initiating conversation. A registration is not a +live binding; a missing catalog row is not proof that no suitable owner exists. +Use scoped discovery and permitted recovery before requesting manual IDs. +Private-owner discovery is broad by default across registered local resources +and authorized connected sources; a delivery allowlist, missing live binding or +bounded first page must not hide an otherwise visible responsible Agent. +Keep discovery, audience evidence access, delegation and execution readiness +separate, as specified by [manager §5.5](capable-manager-semantic-handoff-v0.md#55-responsibility-discovery-and-receiver-owned-planning). + +The P0 pilots are responsibility routing and real 2–3-worker coordination: +parallel joins, peer help, disagreement, independent review and engineering-to- +research adoption across two cycles. Entry checks, correction and recovery +accompany them. P1 adds materials, decision summaries, dependency replanning, +qualified mixed-model allocation and retained constraints; +P2 broadens scale and cinematic presentation after real outcomes are legible. +Reports, truthful activity and scoped controls support this same P0 journey; +finishing every presentation surface is not a prerequisite for attempting it. +Necessary R1/TS transaction repairs retain their owners, without making the +journey wait for the entire migration or optional memory infrastructure. + +The pack freezes paired baseline/candidate tasks, outcome and attention measures, +negative cases and per-surface evidence. All live case results start unqualified. +Keep existing R/G/M/A identifiers and canonical Todos; do not create a parallel +roadmap, scheduler or achievement ledger. Release claims require observed +results, not this plan or merged prerequisite PRs. + ### Collaboration and Handoff Between LoopX Agents Participants are long-running LoopX Agents with their own goals, commitments, frontiers and execution bindings, not merely temporary subtasks inside the steward process. Manager→worker and worker→worker share one collaboration contract. Workers can request help, provide results, challenge dependencies and propose replanning without asking the steward to relay every message. The steward owns overall progress and synthesis, not a serial transit point for every message or commit. @@ -370,7 +409,7 @@ R2's dependency must use real requests/artifact handoff between LoopX Agents. Th | --- | --- | --- | --- | | R1 | P0: confirmed commitments survive materialization; failure/retry is recoverable | Current main and F1–F4 regressions | Other TS transactions and provider promotion | | R2 | P0: one steward drives 2–3 bound managed workers through continued work | R1 and real selected runtime/profile qualification | Full collaboration migration and PostgreSQL | -| R3 | P1: semantic handoff and automatic return survive restart without owner polling | Existing inbox/outbox; new generic producers require transaction migration | Early answer/transport recovery can start with R1/R2 | +| R3 | P0: eligible owner → actual work → original-request return on supported paths; P1: general migration and broader recovery | Existing inbox/outbox and qualified runtime; new generic producers require transaction migration | Complete a real journey with R1/R2 before universal transport/storage migration | | R4 | P1: shared intent/work basis and governed amendment close the loop | Alignment Stage 1/2 and affected TS transactions | R1–R3 that preserve shared intent | | R5 | P1: durable local authority and long-horizon storage qualification | Affected T0–T3 transactions and D1/D2/D3 | Preparation alongside R1–R4; no full TS rewrite prerequisite | | R6 | P2: local and cloud workers share authority and recover execution | R2/R3, selected shared profile, authenticated service and applicable R4 contracts | PostgreSQL service engineering can start earlier without promotion | @@ -475,15 +514,13 @@ This qualifies a local execution-facts readback, not the full R2 ladder: assignm receipt integration, provisioning, remote probes, two-cycle continuation and Lark qualification remain open under their existing owners. For managed non-Chat work, [PR #4978](https://github.com/loopx-project/loopx/pull/4978) -proposes an accepted, exact-version Goal result readback: canonical Todo +added an accepted, exact-version Goal result readback and is merged: canonical Todo completion binds local report bytes, the CLI verifies them, and packaged Goal Files reads a Goal-scoped loopback projection. The [Live Team Workspace RFC](live-team-workspace-v0.md) -defines this producer-to-reader boundary. Focused local File/SQLite and -desktop/mobile checks pass on the proposed head; maintainer review and CI -remain open. Return to the original requester conversation, mixed-team -continuation and stop/recovery remain separate unqualified outcomes. +defines this producer-to-reader boundary. This merged report readback does not +qualify the full intent-routing, requester-adoption or recovery journey. -The proposed managed-result readback binds a current accepted Todo report to +The managed-result readback binds a current accepted Todo report to canonical completion and its exact digest. Goal Files can open it, and a manager conversation can show one report only when its confirmed team-plan receipt names that Todo. Multiple matching reports require selection in the diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md index 581fdf06e0..5ab591416c 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md @@ -315,13 +315,23 @@ R2 的一条依赖必须通过真实 LoopX Agent 间的请求/产物交接完成 下一项可复用的对话检查点以管家为首个接入者,聚焦原始对话:问题得到相称、可读的结果,真实工作事件可查;停止和纠偏作用于正确的 Turn;实质请求到达上下文最合适的活跃 Agent;其评估和有证据的结论返回同一个前端或飞书受众。[展示 RFC §8.8](intelligent-review-presentation-surfaces-v0.zh-CN.md#88-可复用的对话工作界面)拥有共用交互;[管家 RFC §5.14](capable-manager-semantic-handoff-v0.zh-CN.md#514-管家接入可复用的对话工作界面)拥有管家接入与 A21–A24 旅程。沿既有 Chat、session、collaboration、presentation owner 作为 R2/R3 集成推进。候选列表、转圈提示、已排队的收件箱回执均不等于完成 worker 交接;该检查点不降低 R1–R3 或 G1 门槛。 +### 注意力成本验收:创建、接入、协作、看懂 + +[公开安全的 golden-query 集](../../product/use-cases/steward/golden-queries.md)是下一阶段产品验收目标,不是当前支持声明。用户只表达简短目标和纠偏,无需自己找 Agent ID、重述已知约束、催办和搬运结果。保留有价值的学习、主动判断、必要批准,以及与长期负责人的直接讨论;这些不算应该消除的注意力成本。 + +P0 从入口开始:创建请求发出后不能消失;创建 Goal、接入已有 Agent、按授权建立新执行者,都要有可观察结果。随后把 R3/M2/M3 意图分发与 R2 实际采用并列验收:找到合适负责人、真实执行、使用精确版本产物、独立验收、综合结果返回原对话。注册不等于有效执行绑定;目录缺一行不等于没有合适负责人。应先做范围内的发现和可恢复操作,再向用户索要确实缺失的信息。主人私有管家默认广泛发现本机已登记资源与已授权接入来源;委派名单窄、缺活跃绑定或首屏条数受限,不能隐藏原本可见的负责人。发现、受众证据读取、委派和执行就绪分别判断,具体规则归[管家 §5.5](capable-manager-semantic-handoff-v0.zh-CN.md#55-职责发现与接收方重规划)。 + +P0 首批是负责人路由和真实 2–3-worker 协调:两轮并行汇合、同伴求助、分歧处理、独立复核、工程交付给研究采用,伴随入口检查、纠偏、中断和恢复。P1 做材料分发、决策摘要、依赖重排、合格混合模型分配和约束复用;P2 在真实结果可理解后扩展规模和宣传表现力。可读结果、真实活动和 scoped controls 服务于同一 P0 旅程,不再要求先打磨完所有展示渠道才能尝试路由。必要的 R1/TS 事务修复保留 owner,但整条旅程不等待全面迁移或新的可选记忆设施。 + +验收集在运行前冻结成对基线/候选任务、结果与注意力指标、反例和逐入口证据;所有真实运行初始都未验收。沿用 R/G/M/A 编号和 canonical Todo,不另建路线图、调度器或成绩账本。发布主张依据实际结果,不能以规划或前置 PR 合并代替。 + ## 6. 核心交付路径:R1–R7 执行卡 | 卡 | 优先级 / 可验收结果 | 硬前置 | 可同时推进但无需等待 | | --- | --- | --- | --- | | R1 | P0:确认内容与工作落盘一致,失败/重试可恢复 | 当前 main + F1–F4 回归 | TS 其余事务、provider 晋升 | | R2 | P0:一个管家驱动 2–3 个已绑定 managed worker 完成连续工作 | R1;选定 runtime/profile 的真实资格 | 通用 collaboration 全迁移、PostgreSQL | -| R3 | P1:语义交接与自动回报,重启后不用人追问 | 现有 inbox/outbox;通用 producer 依赖事务迁移 | 早期正文/传输恢复可与 R1/R2 同期 | +| R3 | P0:在已支持路径跑通合格负责人→真实执行→原请求回传;P1:通用迁移和更广恢复 | 现有 inbox/outbox 与合格 runtime;通用 producer 依赖事务迁移 | 与 R1/R2 先交付真实旅程,不等待所有传输和存储迁移 | | R4 | P1:共享意图/工作基线与受治理修订形成闭环 | alignment Stage 1/2、相关 TS 事务 | 不阻塞不改变共享意图的 R1–R3 | | R5 | P1:本地 durable authority + 长程存储资格 | T0–T3 受影响事务、D1/D2/D3 | 与 R1–R4 同期准备,不把全部 TS 重写设为前置 | | R6 | P2:本地与云端使用同一权威、可恢复执行 | R2/R3;所选 shared profile、认证服务与相关 R4 合同 | PostgreSQL service 工程可早做;不得提前晋升 | @@ -400,13 +410,12 @@ executor/profile 检查启动条件。任务准入、当前 pinned 验收绑定 这只验收本地执行事实的读回;计划分配回执 整合、通用创建、远端探针、两轮持续协作及 Lark 等价仍由原 owner 继续推进。 对于托管非 Chat 工作,[PR #4978](https://github.com/loopx-project/loopx/pull/4978) -提议已验收、精确版本的 Goal 成果回读:canonical Todo 完成时绑定本地报告字节, +已合入已验收、精确版本的 Goal 成果回读:canonical Todo 完成时绑定本地报告字节, CLI 重新核验,打包 Goal「文件」读取按 Goal 限定的本机回环投影。 [团队实时工作区 RFC](live-team-workspace-v0.zh-CN.md) 定义产出到读取的边界。 -提议版本已通过本地 File/SQLite 和桌面 / 手机的针对性检查;维护者评审与 CI -仍待完成。向原请求方对话回送、混合团队持续协作及整队停止 / 恢复仍需另行验收。 +该合入证明报告读回边界,不证明完整意图路由、请求方采用或恢复旅程。 -待合入的托管结果读回将当前已验收的 Todo 报告绑定到 canonical 完成事实和精确摘要。 +托管结果读回将当前已验收的 Todo 报告绑定到 canonical 完成事实和精确摘要。 Goal 成果页可打开正文;原管家对话仅在已确认团队计划的回执明确包含该 Todo, 且恰好有一份匹配报告时显示。多份报告留在 Goal 内供选择;失败、过期或跨 Goal 读回会撤下正文。这证明报告返回,不证明请求方采用或最终综合答案。要验收一键 diff --git a/docs/product/use-cases/steward/README.md b/docs/product/use-cases/steward/README.md index 18a2e92d4b..e5f167dcfd 100644 --- a/docs/product/use-cases/steward/README.md +++ b/docs/product/use-cases/steward/README.md @@ -99,3 +99,7 @@ regression fails with the missing labels named, so it cannot pass silently. - It does not qualify Lark audiences or any cloud/remote worker. - It does not turn a passing smoke into product acceptance for a Goal whose plan was confirmed with real consequences. + +## Live intent-to-result evaluation + +The [golden-query pack](golden-queries.md) covers creation, existing-Agent connection, responsibility routing, receiver adoption, correction/recovery and attention. It defines proposed live acceptance beyond the synthetic journey above; no live pass is inferred from this browser fixture. diff --git a/docs/product/use-cases/steward/README.zh-CN.md b/docs/product/use-cases/steward/README.zh-CN.md index 45b936b575..80fdb7c10c 100644 --- a/docs/product/use-cases/steward/README.zh-CN.md +++ b/docs/product/use-cases/steward/README.zh-CN.md @@ -71,3 +71,7 @@ LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey \ - 它不资格化真实管家对话:fixture 顶替了 agent turn,因此接入背后的模型/运行时在这里仍未测试; - 它不资格化飞书受众,也不资格化任何云端/远端 worker; - 它不把一条通过的 smoke 当成“某个真实确认过的 Goal 已被产品验收”。 + +## 真实意图到结果验收 + +[Golden-query 集](golden-queries.md)覆盖创建、接入已有 Agent、负责人路由、接收方采用、纠偏恢复与注意力。它定义以上合成旅程之外的真实验收目标,不能从浏览器 fixture 推断真实运行通过。 diff --git a/docs/product/use-cases/steward/golden-queries.md b/docs/product/use-cases/steward/golden-queries.md new file mode 100644 index 0000000000..df0011290f --- /dev/null +++ b/docs/product/use-cases/steward/golden-queries.md @@ -0,0 +1,289 @@ +# Steward golden queries: less coordination, accepted results + +Status: proposed evaluation specification, **not a passing score or a list of +shipped capabilities**. These are public-safe adaptations of real user intent +patterns, not published chat transcripts. Setup and failure injections are +constructed fixtures. No private conversations, organization projects, account +identifiers or live operating state belong in this pack. + +The product objective is simple: say what you need, let LoopX find or establish +the right execution path, change direction when needed, and receive a useful +result. Creation, connection, collaboration and presentation are all part of +that journey. A beautiful answer cannot compensate for a lost request. + +This pack qualifies the [overall roadmap R1–R3](../../../architecture/rfcs/loopx-overall-roadmap-v0.md), +[manager M1–M3 / A1–A24](../../../architecture/rfcs/capable-manager-semantic-handoff-v0.md) +and [shared conversation surface](../../../architecture/rfcs/intelligent-review-presentation-surfaces-v0.md#88-reusable-conversation-work-surface). +It adds no capability, runtime schema, routing keyword list or scheduler. +Canonical Todos own implementation work; this document owns case intent and +evaluation rules. Existing [workspace qualification](README.md) uses synthetic +turns and does not establish these live outcomes. + +## Request style / 请求写法 + +Only the request and ordinary attachments reach the user-facing conversation. +The user does not need to recite Agent IDs, routing policy, receipt vocabulary, +budget rules or acceptance machinery. The evaluator prepares that context below. +Equivalent natural phrasings must work; exact wording is not the oracle. + +## Capability layers and priorities / 能力分层与优先级 + +Layers describe increasing product responsibility, not new roadmap milestones. +The first two layers are both P0: a reliable single-owner path is the entry gate; +**small-team coordination is the next mandatory product outcome**, not optional +polish deferred to hundred-Agent scale. Run only the current bounded cohort; +listing a later query does not launch that experiment or authorize cloud spend. + +| Layer | Product promise / 产品承诺 | Exit and next gate | Existing roadmap | +| --- | --- | --- | --- | +| P0 · Reliable entry and responsible execution / 可靠入口与负责人闭环 | Create or connect, find the right owner, execute, correct/stop and return | Qualify the entry/owner boundaries needed by the selected journey on one installed host; GQ01–04/GQ07–09 retain messages, constraints, authority and results | R1 and supported R2/R3; G0→G1 entry | +| P0 · Real small-team coordination / 真正的小团队协调 | Parallel branches, peer help, disagreement, independent review, dependency adoption and synthesis | GQ05/GQ11–13; 2–3 real workers complete two cycles with correction and recovery; no manual relay or fabricated consensus | R2/R3, G1, M2/M3 | +| P1 · Continuous portfolio coordination / 跨目标持续管理 | Material distribution, decision summaries, dynamic replanning and qualified mixed-model allocation | GQ06/GQ10/GQ14–15; reuse small-team contracts across goals, keep grants/budgets/constraints scoped and user attention bounded | R3/R4, G2; no full storage rewrite prerequisite | +| P2 · Distributed and scaled work / 跨主机与规模化 | Local/cloud cooperation, larger useful parallel cohorts and explainable aggregate progress | GQ16 before GQ17; real two-host failure/recovery, then separately frozen 10→30→100 throughput/cost/attention evidence | R6/G3 then R7/G4 | + +Readable results, evidence, truthful activity and controls belong to every layer. +Visual polish follows these facts; it does not create a fifth coordination engine. +Independent first-use and release qualification remain G5. A configured team, +multiple processes or several isolated answers do not qualify coordination. +Broad owner discovery is a P0 default: all registered local Goals/Agents and +already-authorized connected sources remain searchable without per-recipient +delivery enrollment. Keep prompt pages bounded, disclose coverage and allow +drill-down. Offline/unbound/unknown states stay visible; stopped work is available +through history/search and is not assigned new work. Discovery, evidence access, +delegation and execution readiness are evaluated separately. An external group's +metadata/evidence scope remains its own; private-owner access is not group access. +These layers are delivery priorities, not a rule that every channel and variant +must finish before work on the next layer can start. GQ17 repeats the ordinary +parallel-work intent at larger fixture sizes; basic parallel work is already P0 +in GQ11, and the steward should not overstaff a small task. + +### P0: reliable entry and responsible execution + +| Case | 中文请求 | English equivalent | Accepted outcome | Priority / owner | +| --- | --- | --- | --- | --- | +| GQ01 Create and start | 建个目标:研究微软近三年的现金流,给我份报告。 | Start a goal to research Microsoft's cash flow over the last three years. Bring me a report. | The request remains visible; a durable Goal and an eligible execution path exist, actual work starts, and a sourced report returns | P0 · R1/R2; M1/M3 | +| GQ02 Connect existing work | 用我已经在跑的那个 Codex,接着做。 | Continue with the Codex I already have running. | Resolve the established referent, retain its work/context and qualified binding; no duplicate Goal, worker or execution driver | P0 · R2/R3; A4/A13/A18/A20 | +| GQ03 Find the responsible owner | 让负责 PR review 的 Agent 改一下:别把基线 CI 失败算到这次 PR 上。 | Ask the PR review owner to stop treating baseline CI failures as regressions in the PR. | Find a qualified owner, investigate causal evidence, deliver a reviewed change or evidenced no-change decision, and return it here | P0 · R3; A1/A4/A24 | +| GQ04 Investigate and act | LoopX 首页该不该突出个人 Agent?合适就提个 PR。 | Should LoopX's homepage emphasize personal agents? Open a PR if it makes sense. | Investigate current product and code, retain owner presentation gates, return a concrete preview/PR or supported decision against changing it | P0 · R3; A1/A3/A21/A24 | +| GQ07 Keep working | 把刚才那个问题修好。 | Fix the issue we were just discussing. | Reuse the current task and constraints; repair, validate and return without a sequence of user “continue” prompts | P0 · R1–R3; A5/A7/A8/A13 | +| GQ08 Correct or stop | 先只看微软,亚马逊下次再说。 / 先停下。 | Focus on Microsoft for now; leave Amazon for later. / Stop for now. | Resolve the affected work; the actual owner adopts the correction or the supported stop takes effect, with precise feedback | P0 · R2/R3; A5/A7/A23 | +| GQ09 Resume and return | 接着昨天的做,有要我决定的再说。 | Pick up where we left off yesterday. Ask if you need a decision. | Resume the right unfinished work without duplicate effects; surface necessary decisions and return requested final results despite quiet routine progress | P0 · R3; A8/A13/A17–A20 | + +### P0: real small-team coordination + +| Case | 中文请求 | English equivalent | Accepted outcome | Priority / owner | +| --- | --- | --- | --- | --- | +| GQ05 Deliver for another Agent to use | 把微软现金流数据接进研究工具,跑份分析给我。 | Connect Microsoft's cash-flow data to the research tool and bring me an analysis. | Engineering delivers a versioned input; research actually consumes it, challenges errors, obtains independent acceptance and returns a usable analysis | P0 · R2/R3; A5–A8/A13 | +| GQ11 Parallel team and synthesis | 组个小队,看看微软的 AI 投入能不能赚回来。 | Put together a small team to investigate whether Microsoft can earn a return on its AI investment. | Parallel financial and technical research share a question and evidence basis, exchange dependencies and produce one independently checked synthesis with bounded uncertainty | P0 · R2/G1; A6/A8/A13 | +| GQ12 Independent peer review | 这个结论找个人独立复核一下。 | Have someone independently check this conclusion. | An eligible different reviewer reads the actual version, checks sources and challenges or accepts it; author revises when needed; requester receives the assessed result | P0 · R2/R3; A6/A8 | +| GQ13 Resolve disagreement | 你们结论不一样,把分歧查清楚再给我。 | Your conclusions differ. Check the disagreement before coming back to me. | Locate conflicting evidence or assumptions, assign bounded checks, exchange/revise artifacts, preserve justified remaining dissent and return a reasoned conclusion | P0 · R2/R3; A5/A6/A8 | + +### P1: continuous portfolio coordination + +| Case | 中文请求 | English equivalent | Accepted outcome | Priority / owner | +| --- | --- | --- | --- | --- | +| GQ06 Distribute a material | 看看这篇文章,有用的记下来,值得改的推进。 | Read this article. Save what's useful and follow through on worthwhile changes. | Deduplicate, preserve source and uncertainty, update authorized notes, route an actionable delta to its owner and return its disposition | P1 · R3; A3/A5/A6/A8 | +| GQ10 Protect attention | 这周我只能管两件事,先做什么? | I can focus on only two things this week. What comes first? | Recommend at most two concrete priorities with tradeoffs, evidence and consequences; show gaps and leave Agent-owned work in the background | P1 · R2/R3; presentation §14.4 | +| GQ14 Replan around a dependency | 数据没齐也别都等着,能做的先做。 | Keep moving on what you can while the data is incomplete. | Continue independent branches, resolve the missing dependency with its owner, preserve budget/claims and block only joins that genuinely need that input | P1 · R2/R3/R4 | +| GQ15 Allocate an explicit mixed team | 用两个 Luna、一个 DSH 来做,预算按之前的。 | Use two Luna workers and one DSH, within the budget we agreed. | Configure the requested qualified profiles, allocate actual complementary work, enforce the existing budget and join usable results; no silent substitution or duplicate driver | P1 · R2/R3; G1 before broader profiles | + +### P2: distributed and scaled work + +| Case | 中文请求 | English equivalent | Accepted outcome | Priority / owner | +| --- | --- | --- | --- | --- | +| GQ16 Coordinate across hosts | 本机安排,云上跑,最后一起给我。 | Plan locally, run in the cloud, and bring the results together. | Authorized real hosts resolve versioned dependencies, recover network loss and return once with cost and coverage; stale executors cannot commit | P2 · R6/G3 | +| GQ17 Scale useful parallel work | 把这个项目拆开,能并行的并行,别重复做。 | Split up this project and parallelize where it helps, without duplicate work. | Scale only the useful independent work within capacity, cost and attention budgets; qualify fairness, failure isolation and complete joins at 10/30/100 cohorts | P2 · R7/G4 | + +GQ03 does not require automatic approval of a red build: prove the failure is +unrelated and check other blockers under the existing review contract. Unknown +cause remains unknown. GQ04 permits a reasoned “do not change”; the desired +answer is not predetermined. GQ06 does not authorize public posting. GQ08 is two +separate variants: correcting scope and stopping work must not be conflated. + +## Freeze a reproducible setup before running + +Use disposable Goals, isolated workspaces and a test conversation. Do not modify +an active user's registry, grants or leases to create failures. Public examples +must contain synthetic identities, safe evidence references and no credentials. + +For each baseline/candidate pair, record: + +- source/package commit, installed version, surface (packaged frontend, Goal + Chat, direct Agent Chat or Lark), host/runtime versions and effective profile; +- the initial Goal/Todo/work/artifact versions, known responsible owners, + relevant prior conversation, policy/grants, deadlines and budget; +- source URLs plus retrieved version/date and content digest; material bodies + are linked or stored only when redistribution is permitted; +- independent acceptance facts, permitted receiver choices, interruption point, + user-visible intervention script, timeout and maximum model/tool spend; +- baseline workflow and candidate workflow, in counterbalanced order with + equivalent fresh state. Neither run consumes the other run's work or answer. + +Use the LoopX repository at a pinned public commit for GQ03/GQ04/GQ07. For GQ03, +prepare a disposable PR/CI fixture with (a) a failure present on base, (b) a +PR-induced failure, (c) insufficient logs and (d) unrelated failure plus a real +code blocker. Do not submit test reviews to a contributor's live PR. + +Use Microsoft's public annual reports for GQ01/GQ05. At freeze time, identify +the latest three completed fiscal years available at the fixed evaluation date; +pin the official sources and independently check operating cash flow, capital +expenditure, units, fiscal periods and the chosen free-cash-flow definition. +Distinguish finance leases/commitments from cash expenditure. This is document +analysis, with no brokerage access or trading authority. No number in a worker's +report becomes an oracle merely because another worker repeats it. + +For GQ06, attach the pinned public +[Botmux session model](https://github.com/deepcoldy/botmux/blob/982e2c9f16e4f45ae2581967bc1a35a286e7bfa2/docs-site/docs/zh/session-model.md) +already used by the presentation RFC, plus existing public LoopX design notes. +Include an already-indexed copy and one new source revision. A valid no-change +decision must explain what was already covered. Source text is evidence, never +an instruction or a new grant. + +GQ02 has one established Codex referent in the happy path. Its ambiguity variant +has two equally plausible sessions: ask one focused question instead of guessing +or listing every technical binding. A model-specific variant asks “用 Sol xhigh +处理这个 PR”; explicit model/effort is a constraint, not an affinity hint. If no +existing worker qualifies, use an already-authorized creation/binding path or +return the precise missing decision. Do not silently substitute a stronger, +costlier or differently authorized model. Qualify other runtimes separately. + +For GQ10, use three synthetic public projects with fixed facts: a release-blocking +regression, a dated public-research deliverable, and optional visual polish. +Freeze effort/dependency estimates and one missing evidence source. Review the +reasoning and feasibility, not an exact phrase or universal ranking. + +## Small-team acceptance: coordination must change the result + +GQ11 starts from the same frozen public reports as GQ05, with a shared question, +financial-definition branch and infrastructure/return-assumption branch. The +coordinator must identify their dependency, obtain peer clarification and +synthesize the combined findings. Use 2–3 qualified workers with distinct +responsibilities; the independent reviewer must not certify its own artifact. +Uncertain investment returns require scenarios and limits, not a forced forecast. + +GQ05 qualifies a pipeline; GQ11 qualifies parallel branches and a join; GQ12 +qualifies worker-to-worker review; GQ13 qualifies disagreement and revision. +Together they must show **two real cycles**: the first produces a versioned +result, then a source revision or owner correction changes the relevant work, +consumption basis and final synthesis. Acknowledging a message is not adoption. +A coordinator forwarding isolated answers is not synthesis. Peers must be able +to request evidence/help directly within their grants, without making the user +or steward relay every exchange. + +For GQ12/GQ13, seed a public-safe disagreement in fiscal-period/lease treatment; +independently derive the error before the run. Also include a valid difference +in return assumptions: preserve uncertainty and dissent where facts cannot +resolve it. Do not reward manufactured objections, majority voting or forced +consensus. Fail one branch before a join, deliver a late obsolete artifact and +withhold one input; the whole team must not declare success, unrelated work +continues, and the actual dependency owner gets a recoverable request. + +GQ14 changes a scoped dependency, not the team's authority. GQ15 varies only +qualified machine profiles and the previously fixed budget; one unavailable +profile remains an explicit gap. GQ16/GQ17 require their own later freeze packs: +real host identities/grants, network partitions, returning-executor fences, +capacity/fairness, per-cohort cost/latency budgets and a problem with enough +independent work to justify the cohort. These two are roadmap targets, not +ready-to-run scale fixtures or a reason to spawn idle workers now. + +## Required variants and independent observations + +| Boundary | Fixture or intervention | What the observer must establish | +| --- | --- | --- | +| Message/Goal creation | Submit, delayed response, failed create, reload after response loss | Source message remains pending/failed/recoverable until acknowledged; one canonical creation; retry does not create a second Goal; “sent” is not “started” | +| Broad default discovery | Relevant registered owner is outside the delivery allowlist or first page; another is offline/unbound; one authorized connected source is temporarily unreadable | Owner discovery still finds the permitted relevant identity and explains the specific delegation/readiness gap; pagination reaches it, source failure stays visible, and routine investigation does not require the user to supply IDs. Shared audiences cannot see out-of-scope private identities/content; no grants or worker wake are inferred | +| Discovery versus eligibility | Correct registered owner, sole irrelevant candidate, stopped Goal, stale binding, missing presence probe, unavailable remote source | Incomplete discovery is not “no Agent exists”; relevance is not eligibility; an irrelevant sole candidate is not selected; unknown reachability is not fabricated readiness | +| Recoverable route gap | Stale authorized catalog or absent binding with a supported repair/probe path | Refresh scoped sources and attempt permitted repair before asking the owner for IDs; retain the request and say precisely which stage needs help; do not broaden grants | +| Actual execution | Receiver inbox ACK without a worker starting | The result remains queued/waiting with owner and recovery condition, never “being handled” or completed without evidence | +| Adoption | Engineering artifact v1 has a seeded unit/period error; revised v2 fixes it | Research rejects v1, reads and uses v2 in a recalculated result; independent checker validates it; coordinator incorporates that accepted result before final return | +| Steering | Deliver GQ08 while a tool is pending and again at completion boundary | Receiver uses the latest intent; superseded result stays labeled; no silent loss or new unrelated task; uncertainty follows existing ingress/Turn semantics | +| Stop scope | Stop current conversation; separately request stopping delegated work | Actual native interruption/unsupported state is visible; conversation stop does not imply team stop; delegated stop requires its own authority and readback; unrelated work continues | +| Recovery and duplicate delivery | Lose reply ACK, restart supported host, replay same ingress/callback; later deliberately repeat a similar request | Recover saved result/request first; one effect and one logical answer per source; identical text is not a global deduplication key; a genuinely new request still works | +| Long and simple answers | Long comparative result; “刚才那条 PR 合了吗?” with an unambiguous prior reference | Readable leading answer and resolvable evidence, full safe Markdown/report accessible on each supported surface; simple read stays direct and need not create a team | +| Retained constraints | Prior “do not publish” and cost preference; explicit scoped override later | Retrieve applicable constraints without user repetition; separate durable preference, fresh fact and action authorization; no stale preference overrides the current instruction | +| Attention and depth | Unchanged blocker, new deadline, requested result, voluntary deep discussion | No repeated unchanged alert; timely material decision with object/recommendation/evidence/inaction consequence; requested final result still returns; discussion is not penalized | +| Understanding work | User opens owner conversation and returns; evidence unavailable on one host | Same work/result lineage, meaningful decisions flow back; unsupported coverage is named; user can tell who owns work, what changed and what needs a decision | + +The discovery diagnosis must separate source coverage, registration, authorized +scope, binding, runtime/profile, capacity, receiver assessment, execution and +return. Existing typed facts and receipts own those states. A text reason such +as “no suitable agent” is insufficient evidence; this table is not a proposal +for a second universal state machine. Observe both actual permission denial and +an empty/incomplete projection. Do not “repair” absence by granting all Agents +access or hard-coding the current developer's Agent name. + +## Scoring: outcome and attention together + +An evaluator who did not implement the candidate checks each task's sources, +result and state readback. A model judge may assist narrative scoring; it cannot +certify authorization, version use, message delivery or effect deduplication. +Capture the input, receiver assessment, work/result references, independent +acceptance and original-route answer. Redact the evidence before publication. + +Record **pass / fail / blocked / not run** for each case, variant and surface. +Only pass is success. Missing permissions/runtime, budget exhaustion and timeout +are blocked or failed with the cause retained, not omitted from the cohort. +Development fixtures, packaged browser checks and real host runs have separate +columns; none substitutes for another. Cross-host and Lark claims require their +own runs. All live golden-query cells start **not run** in this specification. + +Measure avoidable human coordination by category: finding the owner, supplying +already-available context, repeating retained constraints, manually assigning, +chasing progress, transporting results and synthesizing what the task requested. +Count an input once in the total and retain secondary category tags. Review the +cause rather than classifying words: “continue” can be a new instruction, and a +clarification can be necessary. Human goal changes, required authorization, +voluntary learning and careful final judgment are reported separately and are +not waste to eliminate. + +Report task-level accepted quality, avoidable interventions and active attention +minutes, plus completion latency, model/tool cost, duplicate effects, missed +decisions and false alerts. Unknown cost/time is unknown, not zero. Show raw +per-attempt totals and the accepted-outcome denominator; include failed and +abandoned attempts in the comparison so silence/failure cannot look efficient. + +Initial exit targets, frozen before candidate execution: + +1. No unauthorized effects, wrong-target stop, false completion/adoption, + duplicate effect or silently lost request/result in any required variant. +2. The responsibility-routing and small-team pilot happy paths pass on the + supported installed host and packaged frontend; no manual Agent-ID handoff, copy/paste relay or reminder is needed. + Run the corresponding Lark cases before claiming Lark equivalence. +3. Candidate accepted-outcome quality is no worse than the paired baseline; + aggregate avoidable coordination decreases by at least **30%**, with + attention time not increased. This is a proposed target, **not measured + improvement**. If baseline coordination is zero, use non-regression instead + of a percentage; report the case counts, no population-level claim. +4. Latency and spend remain within the predeclared per-case envelope. Do not + hide high-cost model substitution behind fewer user messages. + +After pilot repair, qualify all P0 small-team cases before the P1 families. +Use independently authored held-out paraphrases and the required variants. Freeze expected semantics before viewing +candidate answers. Do not tune production routing to these example strings. Larger +first-use cohorts and release acceptance retain their existing gates. + +## Delivery order and reuse + +| Batch | Useful exit | Reused owner / next dependency | +| --- | --- | --- | +| P0 entry and route | GQ01/GQ02 request durability plus GQ03/GQ04 eligible responsibility, actual work and same-conversation result | Existing creation/Chat services, directory, host binding and collaboration/outbox; ship the complete supported path before general migration | +| P0 use and continuity | GQ05/GQ11–GQ13 + GQ07–GQ09: dependency adoption, parallel join, peer review and resolved disagreement across two cycles; correct once, interrupt once, resume and return | R2 small-team and R3/M2/M3, artifact versions, existing driver/monitor and return recovery | +| P1 material and attention | GQ06/GQ10/GQ14–GQ15: materials, attention, dependency replan, explicit mixed profiles and retained constraints | Existing material lifecycle, scoped context and presentation; no new memory installation prerequisite | +| P2 breadth and launch | GQ16 then GQ17: real host and scale qualification; public-safe showcase/film only claims the separately proven cohort | Existing R6/R7 and release/first-use gates; visual motion explains actual transitions | + +The first focused pilots are **GQ03/GQ04 responsibility routing and +GQ05/GQ11–GQ13 real small-team coordination**; **GQ06 material distribution** +follows as the first P1 transfer test. GQ01/GQ02 are prerequisite entry +checks, not deferred onboarding work. GQ07–GQ09 are perturbations of those same +journeys, not extra schedulers. Shared answers/activity/controls are companion +acceptance on the journey; polish across every channel is not a prerequisite +for attempting the first real route. + +Reuse the [collaboration delivery example](../../../../examples/collaboration-delivery/README.md) +for the real receiver/correction/independent-oracle pattern and the +[managed research team](../../../../examples/managed-research-team/README.md) +for isolated runtime qualification. Neither currently proves this entire pack. +Attach case-level evidence to the existing implementation PR and canonical Todo; +only extract a durable regression into product tests after reproducing it. +No speculative benchmark runner, public transcript corpus or additional polling +automation is required to begin.