diff --git a/docs/architecture/rfcs/desktop-execution-frontends-v0.md b/docs/architecture/rfcs/desktop-execution-frontends-v0.md index cf982fc6bd..b2b662095a 100644 --- a/docs/architecture/rfcs/desktop-execution-frontends-v0.md +++ b/docs/architecture/rfcs/desktop-execution-frontends-v0.md @@ -5,7 +5,7 @@ session and an end-to-end LoopX-managed desktop runtime - Initial attached runtime: Codex App / app-server - Initial managed runtimes: Pi and DeepSeek Harness (`dsh`) -- Default managed provider profile: Volcengine Ark Agent Plan +- Managed provider selection: explicit guided configuration and user intent; Agent allocation within authorized eligible profiles. Ark Agent Plan remains an optional distribution preset. ## Summary @@ -16,9 +16,10 @@ LoopX Desktop should support two explicit execution frontend modes: process, conversation, interruption, resume, and execution-loop ownership. 2. **Managed Agent Runtime.** LoopX Desktop launches and supervises Pi or DeepSeek Harness, selects an explicit provider profile, and advances work - through bounded `loopx_turn_v0` transactions. The default distribution - profile uses Volcengine Ark Agent Plan, while the runtime and provider - contracts remain replaceable. + through bounded `loopx_turn_v0` transactions. A distribution may offer + Volcengine Ark Agent Plan as a named preset. The operator's configuration and + explicit intent bound autonomous Agent allocation; no provider is a universal + product default. Runtime and provider contracts remain replaceable. Both modes present the same LoopX Goal, Todo, gate, quota, evidence, and status truth. They do not share process ownership. The frontend must never infer a @@ -44,6 +45,46 @@ eligible, invokes the selected runtime adapter, validates its result, and commits accepted state. `loopx_turn_v0` remains one transaction rather than a second recurring scheduler. +## Guided configuration, user intent and autonomous allocation + +This proposed selection rule refines the provider-default language; it does not +change shipped runtime defaults or qualify new adapters. Reuse the existing +machine/Goal capability editor, provider store, session binding and dispatch +admission. It adds no capability id, provider implementation or routing service. + +- **Configuration makes choices usable.** Show detected installation and login, + supported tools, model/account, cost or unknown cost, host availability and + permission scope. Discovery does not grant access. A named preset is an + offered choice; it cannot override an existing configured selection. +- **Explicit user intent binds the choice.** A fixed model/account, local-only + requirement, budget or member restriction applies at its declared scope. + A preference is not a hard lock unless the user made it one. Do not turn a + single user's Codex preference, or one distribution's Ark preset, into a + global provider rule. Unresolved conflicting instructions require clarification. +- **The Agent allocates inside that boundary.** Choose eligible models/runtimes, + reuse or request workers and redistribute future work by task fit, tool access, + cost and observed availability. A flexible authorized pool does not need a + fresh confirmation for each assignment. Selection is semantic Agent judgment; + typed owners enforce eligibility, budget, authority and session fences. + +Project the effective runtime/model/profile, assignment reason and readiness in +settings and the team detail. Recheck admission at dispatch. An unavailable pinned +route blocks with a repair action; an unavailable member of a flexible pool may +be replaced by another eligible route with visible readback. No substitution may +expand data exposure, credentials, cost authority or supported tools. Active +sessions retain their binding and context ownership; an authorized autonomous +reassignment uses the existing explicit rebinding/continuation contract, not a +silent session migration. New paid resources or scope changes retain their +existing decision boundary. + +Qualify through packaged setup, Chat/steward and independent CLI: pinned choice +wins over a preset; flexible allocation succeeds without repeated approval; +unavailable pinned choice stays blocked; exhausted budget or an unauthorized +fallback creates no execution; restart preserves the effective configuration. +Optional Lark must project the same choice and its own audience restrictions. +Until these cases pass, this is the allocation design, not a runtime guarantee. +See the [near-term launch route](loopx-overall-roadmap-v0.md#near-term-local-agent-product-and-launch). + ## Problem LoopX already has the pieces of two different products: @@ -536,8 +577,8 @@ The managed desktop path is end to end: 1. select or create a LoopX Goal and working Agent binding; 2. select Pi or `dsh` as the runtime; -3. select a managed provider profile, with Ark Agent Plan as the default - distribution profile; +3. guide provider configuration and user constraints, then let the Agent select + within the authorized eligible profiles; offer Ark Agent Plan as a named preset; 4. validate runtime installation, provider authentication, and advertised capabilities; 5. launch one runtime and create one opaque resumable session; @@ -596,9 +637,10 @@ for reconciliation, validation, and resume. ### Provider profile contract -Runtime choice and provider choice are orthogonal. Ark Agent Plan is the -default managed product profile, not a special case embedded throughout the -LoopX kernel. +Runtime choice and provider choice are orthogonal. Ark Agent Plan is a named +managed-provider preset. Guided configuration, scoped user intent and Agent +allocation select an eligible profile; no provider rule is embedded throughout +the LoopX kernel. A provider profile must expose or resolve: @@ -907,7 +949,7 @@ running. 2. one Desktop runtime supervisor with start, interrupt, close, reconcile, and resume; 3. `dsh` as the first reference runtime by reusing its accepted Turn adapter; -4. Ark Agent Plan as the default configured provider profile; +4. Ark Agent Plan as one explicitly selected provider preset; 5. one resumable conversation and one-at-a-time bounded Turn execution; and 6. joined runtime, Turn, and LoopX status in Desktop. diff --git a/docs/architecture/rfcs/desktop-execution-frontends-v0.zh-CN.md b/docs/architecture/rfcs/desktop-execution-frontends-v0.zh-CN.md index 3a7df7bbe2..d492f2670e 100644 --- a/docs/architecture/rfcs/desktop-execution-frontends-v0.zh-CN.md +++ b/docs/architecture/rfcs/desktop-execution-frontends-v0.zh-CN.md @@ -4,7 +4,7 @@ - 决策边界:同时支持挂接到外部拥有的 Agent 会话,以及端到端由 LoopX 托管的桌面运行时 - 初始挂接运行时:Codex App / app-server - 初始托管运行时:Pi 与 DeepSeek Harness(`dsh`) -- 默认托管 provider 配置:火山方舟 Agent Plan +- 托管 provider 选择:显式配置引导与用户意图约束,Agent 在已授权可用 profile 内自主分配;火山方舟 Agent Plan 保留为可选发行预设。 ## 摘要 @@ -15,8 +15,9 @@ LoopX Desktop 应支持两种显式的执行前端模式: 中断、恢复和执行循环的所有权。 2. **托管 Agent 运行时(Managed Agent Runtime)**。LoopX Desktop 启动并 监督 Pi 或 DeepSeek Harness,选择显式的 provider 配置,并通过有界的 - `loopx_turn_v0` 事务推进工作。默认发行配置使用火山方舟 Agent Plan, - 而运行时与 provider 契约保持可替换。 + `loopx_turn_v0` 事务推进工作。发行版可提供命名的火山方舟 Agent Plan 预设, + 操作者配置与明确意图约束 Agent 的自主分配;不把任何 provider 设为统一产品 + 默认。运行时与 provider 契约保持可替换。 两种模式呈现相同的 LoopX Goal、Todo、gate、quota、evidence 和状态事实。 它们不共享进程所有权。前端绝不能从聊天散文推断模式切换,也不得静默启动 @@ -36,6 +37,34 @@ manager Agent,也不把连接器硬编码到 Codex。 下一个有界 Turn 是否有资格执行,调用所选运行时适配器,验证其结果,并提交 被接受的状态。`loopx_turn_v0` 始终是一个事务,而不是第二个常驻调度器。 +## 显式配置引导、用户意图与自主分配 + +这里细化 provider 默认策略,属于提案,不改变已发布运行时默认值,也不认证新 +适配器。复用现有 machine/Goal capability editor、provider store、会话绑定与 +调度准入,不新增 capability id、provider 实现或选路服务。 + +- **配置让选择可用。** 展示已探测安装与登录、支持工具、模型/账户、费用或费用 + 未知、宿主在线条件和权限范围。发现不等于授权。命名预设是可选方案,不能覆盖 + 已有明确配置。 +- **用户明确意图约束选择。** 固定模型/账户、本地限定、预算和成员限制,按用户 + 声明的作用域生效。偏好不自动变成硬锁。不能把某位用户的 Codex 偏好,或某个 + 发行版的 Ark 预设,推广成全产品规则。未解决的明确约束冲突需要澄清。 +- **Agent 在范围内自主分配。** 按任务适配性、工具权限、成本和观测到的可用性, + 选择可用模型/运行时、复用或请求 worker、重新分配后续工作。已授权的灵活资源池 + 不要求逐次确认。选择由 Agent 作语义判断;typed owner 执行资格、预算、权限和 + 会话 fence。 + +在设置与团队详情投影有效 runtime/model/profile、分配理由和就绪状态,派发时 +重新验准入。锁定路径不可用就阻塞并给修复入口;灵活池中的成员不可用,可以选择 +另一条合格路径并明确回读。不能借替代扩大数据外发、凭证、费用权限或支持工具。 +活跃会话保留绑定与上下文所有权;已授权自主改派走现有显式重新绑定/continuation +契约,不静默迁移会话。新增付费资源和范围变化保留原有决策边界。 + +通过打包引导、Chat/管家和独立 CLI 验收:固定选择优先于预设;灵活分配不反复 +审批;锁定路径失效时保持阻塞;预算耗尽或未授权回退不产生执行;重启保持有效 +配置。可选 Lark 投影同一选择及自己的受众约束。用例通过前,这是分配设计而非 +运行保证。见[近期发布路线](loopx-overall-roadmap-v0.zh-CN.md#近期本地-agent-产品与发布路线)。 + ## 问题 LoopX 已经拥有两个不同产品的零件: @@ -446,7 +475,7 @@ managed 面板或 supervisor。 1. 选择或创建 LoopX Goal 和工作 Agent 绑定; 2. 选择 Pi 或 `dsh` 作为运行时; -3. 选择托管 provider 配置,默认发行配置为 Ark Agent Plan; +3. 引导配置 provider 与用户约束,由 Agent 在已授权可用 profile 内选择;提供命名的 Ark Agent Plan 预设; 4. 验证运行时安装、provider 认证和已宣称能力; 5. 启动一个运行时并创建一个不透明可恢复会话; 6. 向同一会话发送用户输入; @@ -497,8 +526,9 @@ Pi 和 `dsh` 实现同一个窄托管运行时契约,而不假装其内部循 ### Provider 配置契约 -运行时选择与 provider 选择正交。Ark Agent Plan 是默认托管产品配置,而不是 -散落在 LoopX 内核各处的特例。 +运行时选择与 provider 选择正交。Ark Agent Plan 是命名的托管 provider 预设。 +显式配置引导、作用域内的用户意图与 Agent 自主分配共同选择合格 profile, +不把 provider 规则散落在 LoopX 内核各处。 一个 provider 配置必须暴露或解析: @@ -770,7 +800,7 @@ executor 精确版本、完成情况、延迟、动作数、人工介入、禁 2. 一个 Desktop 运行时监督器,支持 start、interrupt、close、reconcile 和 resume; 3. `dsh` 作为第一个参考运行时,复用其已被接受的 Turn 适配器; -4. Ark Agent Plan 作为默认配置的 provider 配置; +4. Ark Agent Plan 作为一个显式选择的 provider 预设; 5. 一个可恢复对话和一次一个的有界 Turn 执行;以及 6. 在 Desktop 中联合展示运行时、Turn 和 LoopX 状态。 diff --git a/docs/architecture/rfcs/live-team-workspace-v0.md b/docs/architecture/rfcs/live-team-workspace-v0.md index da8276e9a9..ea485dffd6 100644 --- a/docs/architecture/rfcs/live-team-workspace-v0.md +++ b/docs/architecture/rfcs/live-team-workspace-v0.md @@ -309,12 +309,17 @@ Pausing the coordinator does not stop dispatched workers: name those workers and the remaining execution scope. If worker cancellation is unsupported, state it explicitly; a whole-team stop claim requires worker-owner termination readback. -The current producer gap is substantive: `consume_return` records consumption, -not version-bound requester adoption; research-specific adoption checks in an -example are not a generic production projection. Extend the existing owning -contract and its real caller where necessary, rather than inventing frontend -completion from prose. Do not claim L1 complete until this gap is closed for the -selected episode. Motion is retained only when it clarifies these transitions; +Implementation checkpoint (2026-09-20): [#4762](https://github.com/loopx-project/loopx/pull/4762) +is open and proposes version-bound response/revision/use inputs, explicit +requester adoption into an accepted downstream artifact, evidence navigation, +contextual feedback and scoped coordinator pause. File/SQLite, CLI/MCP and +packaged-browser checks cover that local slice; the real GPT attempt stopped at +MCP tool approval before an accepted artifact. This is neither shipped acceptance +nor L1 qualification. Review and reuse that producer rather than rebuilding it. +`consume_return` alone still means consumption, not version-bound adoption. +Qualify the substantive independent objection, correction and useful synthesis +through the selected real-model episode before claiming L1. Motion is retained +only when it clarifies these transitions; remove effects that obscure absent execution, absent acceptance or source loss. Broader semantic zoom, extra actors and renderer experiments follow this exit. @@ -354,6 +359,8 @@ sensitive event bodies in diagnostics. ## 11. Delivery order and relationship to aggressive R2 progress +The [near-term local-agent launch](loopx-overall-roadmap-v0.md#near-term-local-agent-product-and-launch) uses L1 as its real-run demonstration and L2/G1 for sustained-team claims. Marketing preparation can run alongside implementation, but cannot advance these exits. + | Slice | Complete useful result | Entry / exit | Owner and rollback | | --- | --- | --- | --- | | L0 Design and traceability | Interactive synthetic study, source audit and executable acceptance plan | Design review; V1/V3 concept checks; no live claim | Presentation; discard prototype, retain decisions | @@ -403,7 +410,8 @@ motion, correction replay, evidence drill-down and degraded-state presentation. It launches no Agents and reads no private research data. It is not shipped in the product or used as live runtime evidence. -**Current checkpoint:** L0 design proposal; L1–L3 unimplemented here. V1/V3 may be +**Current checkpoint:** L0 design is merged; Section 9 records the open L1 +implementation candidate. Full L1 and L2–L3 remain unqualified. V1/V3 may be explored with the synthetic study; V2/V4/V5/V6/V7 remain unqualified until their required production or measured evidence exists. No G1/G3/G4 promotion follows. The exact delivery PR carries validation and review; roadmap pointers retain diff --git a/docs/architecture/rfcs/live-team-workspace-v0.zh-CN.md b/docs/architecture/rfcs/live-team-workspace-v0.zh-CN.md index 6ff5ca50e9..57f30ebf97 100644 --- a/docs/architecture/rfcs/live-team-workspace-v0.zh-CN.md +++ b/docs/architecture/rfcs/live-team-workspace-v0.zh-CN.md @@ -252,10 +252,13 @@ session、提高历史 receipt 的语义强度或重置 provider。默认概览 什么。暂停协调员不会停止已派发成员:须显示这些成员及仍在执行的范围。成员 不支持取消时明确说明;宣称整队停止需要成员执行 owner 的终止回读。 -当前生产者缺口有实质意义:`consume_return` 只记录消费,不是绑定版本的请求方 -采用;示例中的投研专用采用检查,也不是通用生产投影。必要时扩展既有 owning -contract 及其真实调用方,不能在前端从文字制造完成状态。所选过程补齐这些 -证据前,不宣称 L1 完成。动态只有让上述变化更清楚才保留;掩盖未执行、未验收 +实施检查点(2026-09-20):[#4762](https://github.com/loopx-project/loopx/pull/4762) +仍开放,提出版本绑定的响应/修订/使用输入、请求方对已验收下游产物的显式采用、 +证据导航、上下文反馈和范围限定的协调者暂停。File/SQLite、CLI/MCP 与打包浏览器 +检查覆盖本地切片;真实 GPT 尝试在产物被验收前被 MCP 工具审批拦截。这不代表 +已发布验收或 L1 资格。review 并复用该 producer,避免重做;`consume_return` 单独 +仍只代表消费。宣称 L1 前须以真实模型过程验收实质的独立异议、纠偏与有用综合。 +动态只有让上述变化更清楚才保留;掩盖未执行、未验收 或来源失联的效果应移除。更广的语义缩放、成员扩张和 renderer 探索在此后推进。 ### 资格化矩阵 @@ -288,6 +291,8 @@ V5 候选目标在实验前定稿:前台动态目标 60fps / 优雅降级底 ## 11. 交付顺序与激进推进 R2 的关系 +[近期本地 Agent 发布路线](loopx-overall-roadmap-v0.zh-CN.md#近期本地-agent-产品与发布路线)以 L1 作为真实运行演示,以 L2/G1 支撑持续团队承诺。宣传准备可与实现同期推进,但不替代这些验收。 + | 切片 | 完整有用结果 | 进入 / 退出 | Owner 与回滚 | | --- | --- | --- | --- | | L0 设计与可追溯性 | 可交互合成研究、源审计、可执行验收计划 | 设计评审;V1/V3 概念检查;不声明 live | Presentation;丢弃原型,保留决策 | @@ -325,6 +330,6 @@ SSE。第 6 节一手资料支撑设计选项,不构成产品验收。一个 空间成员、单次交接动态、纠偏回放、证据深入与降级状态;不启动 Agent,不读取 私人投研数据,未作为产品发布,也不充当真实 runtime 证据。 -**当前检查点:** L0 设计提案;本文未实现 L1–L3。合成研究可探索 V1/V3; +**当前检查点:** L0 设计已合并;第 9 节记录开放中的 L1 实现候选。完整 L1 及 L2–L3 仍未验收。合成研究可探索 V1/V3; V2/V4/V5/V6/V7 在获得其要求的生产或测量证据前都未资格化。不晋升 G1/G3/G4。 准确交付 PR 承载验证与 review;总路线只投影此边界,不再追加运行任务账本。 diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md index 7c119a1fd4..83d285ec6d 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md @@ -85,6 +85,133 @@ First resource ordering: complete R1 commitments/recovery, then qualify G1. One Pause/downscope when a design duplicates authority or widens defaults, real-backend qualification is missing, rollback is not independently possible, cost/attention grows uncontrollably with scale, benchmark integrity fails, or users cannot explain work and blockers. Preserve facts/receipts, stop affected new admission and repair rules or reduce the cohort. Do not weaken acceptance, remove failed samples, refresh test expectations or rename a stage into completion. +### Near-term local-agent product and launch + +**Planning revision: 2026-09-20; proposal, not a release qualification.** +Make the native route approachable as **an open-source personal Agent that uses +an existing local Agent to carry work through to a checkable result**. Start +with developers and research-heavy users who already have a supported Agent +login. One persistent steward is the entry; specialists appear when useful. +This is a bounded S1/S4/S5/S12/S13 delivery of G0/G1 and early G5, not a third +runtime, another roadmap, or a prerequisite for the observer-first route. + +“Make your local Agent work like Grok Bot or Muse” is a useful comparison hook. +The primary promise should be “Give your Agent a goal. Come back to the work.” +Use the comparison with the precise supported scope, not “full open-source +replacement”, model equivalence, free inference, guaranteed unattended success, +or an implication of affiliation. Local orchestration does not mean offline +inference or that providers never receive task data. Background work requires +an available execution host; closing a browser is not shutting down a laptop. + +#### Research basis and product decisions + +Sources inspected on 2026-09-20; vendor descriptions are **stated**, not tested. + +| Direct source | Relevant observation | LoopX decision / remaining limit | +| --- | --- | --- | +| [Grok Bot product](https://x.ai/bot) and [launch](https://x.ai/news/introducing-grok-bot) | Vendor describes persistent cloud teammates, shared tools, handoff and returned work | Adopt conversational delegation and visible returned artifacts; local availability cannot inherit the vendor's cloud 24/7 promise | +| [Meta Muse announcement](https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/) and [design rationale](https://introducing.muse.ai/) | Vendor describes a persistent main chat, Goals, meaningful proactive updates, rich artifacts and explicit action controls | Make one steward and first useful outcome primary; use deterministic controls for consequential actions and progressive disclosure for tools | +| [Original X comparison](https://x.com/Michaelzsguo/status/2101057424115253645) | One user's account favors Muse's integrated experience over configuring components | Treat as a productization hypothesis, not comparative performance evidence; test first use independently | + +The X post was inspected through Ego Lite. Official pages were read; no signed-in +product run or complete demo video was tested. The Muse design page timed out in +the browser; its article was available through web retrieval. Do not infer actual +reliability, privacy equivalence or conversion rates from this investigation. + +#### One first-use journey and one flagship result + +1. Install the released package using the existing installation owner; open + `loopx dashboard`. Diagnose Python/Node, the host login, selected model/profile + and tool readiness in place. A package install alone is not a ready Agent. +2. Guide explicit configuration of eligible runtimes, model accounts, tools, + budgets and authority. Honor the user's explicit intent; the Agent chooses and + allocates eligible routes within that scope. A pinned steward/Chat profile is + binding, while an authorized flexible pool permits autonomous selection. + Codex and Ark are configurations, not universal preferences. Surface the + effective choice; never copy credentials, migrate old sessions or silently + fall back outside authorization. Missing readiness offers a concrete repair. +3. Give the steward a real task without first constructing a team or learning + Goal/Todo/Turn vocabulary. Preview relevant scope, outputs and resource limits; + routine already-authorized reversible work proceeds without repeated cards. +4. Deliver a sourced research brief from public documents. Add a contradictory + or newer source, show an independent objection, the exact revision and check, + then the requester's adopted synthesis. The first single-Agent result must + remain useful even when team execution is unavailable. +5. Return the artifact and the next meaningful decision to the original + conversation. Inspect evidence, change direction, pause or stop the named + execution scope. Reopen the browser and recover context; stale/disconnected + work remains explicit. Proactive updates require a material change. + +Use this same episode for the product, a reproducible public sample and the +launch film. Domain-specific research can build on it without making financial +recommendations or transaction authority part of the generic kernel. Engineering +issue-to-reviewed-patch is the next already-owned use case, not a second blocker +for the first research preview. Lark is an optional later transport using the +same session, audience and result owners; first local onboarding must not require +Lark configuration. CLI readback remains part of every claimed local journey. + +#### Delivery order, owners and release claims + +Estimates are elapsed working time from a staffed start, assuming one dedicated +engineer with Agent assistance, shared design/QA support, daily maintainer review, +one macOS + Codex configuration, available model usage and authorized test tools. +They are planning ranges, not measured velocity, SLAs or a promise of all-host +support. A missing dependency moves the date, not the acceptance bar. + +| Stage / cumulative target | Complete work and existing owner | Exit / permitted claim | +| --- | --- | --- | +| Narrative and storyboard · 2–3 working days | S1/S5/S13: one audience, one task, first-screen design, 60–90s script, bilingual copy and public evidence plan | Reviewable polished draft; label concepts and replays. No runnable-product claim | +| Recorded product preview · 5–7 working days | R2/R3 + S5: review/integrate #4762, resolve supported host-tool authorization, run one real L1 correction; refine the actual result and evidence view | Packaged frontend plus CLI show exact-version objection/revision/acceptance/adoption, intervention and scoped stop. Record an actual run; no synthetic activity presented as live | +| Focused public alpha · 2–3 weeks | S1/S4/S12: release-package onboarding, intent-respecting steward/Chat selection, readiness repair, single-Agent fallback, restart/reconnect and useful result return; S5 supplies focused team view | Freeze five independent clean-install attempts before testing; at least four reach the first useful result without maintainer shell intervention, all failures retained. G1's two real cycles must pass before advertising continued team collaboration | +| Repeatable beta / full launch kit · 4–6 weeks | S4/S5/S10/S12/S13: upgrade/rollback, revoke/unavailable/quota recovery, pilot-driven UX, measured cost and support; qualify optional Lark separately | At least three independent users repeat a task on another day; publish denominator, interruptions and limitations. Each advertised platform/transport has its own release-artifact evidence | + +Suggested alpha usability targets, to freeze before recruitment: setup within +10 minutes after documented prerequisites and authentication; first useful +artifact within 15 minutes for the bounded sample task at a declared budget. +These are proposed targets, not current performance facts. Report both complete +onboarding time (including prerequisites/login) and the post-readiness interval; +authentication or environment failures stay in the funnel. A small pilot is +feedback, not statistical reliability or product-market fit. + +The critical path is host/tool readiness → real L1 episode → packaged first-use +repair → independent reproduction. #4758 is merged design; #4762 is open at this +revision and does not establish a shipped or real-model-qualified journey. +Its GPT trial ended at MCP approval with no accepted artifact. Resolve the +supported authorization path with its owner; never turn off approval globally +to obtain a recording. Existing implementation successors remain authoritative; +inspect canonical Todos and related PRs before assigning a new slice. + +The first launch does not wait for hundred-Agent scale, a new state provider, +full 3D, all hosts, mobile-native shells, a connector marketplace, payment/booking +or teach-by-demonstration automation. Those retain their existing owners and +independent demand/qualification. Do not divert R1 reliability work to them. + +#### Presentation and launch assets + +The signature moment is **the conclusion changing because a peer found better +evidence**. A precise command surface holds conversation, result and decisions; +stable spatial stations and event-local motion make the exchange legible. Show +before/after claims, cited sources and the accepted adopted version. Keep one +focal event, an accessible list, reduced motion and honest quiet/failed states. +No ambient animation or Agent count substitutes for work. Follow the existing +[design contract](../../development/design.md); actual public first screens +still require their concrete preview approval. + +Prepare one consistent kit: a 60–90s real-run film with disclosed time cuts, +a 15s excerpt, three annotated product screenshots, an accessible replay marked +historical, bilingual launch article and short social draft, one install/try +path, a supported-version/cost/privacy/availability FAQ, and a public-safe sample +with expected output, correction input and recovery steps. Draft these alongside +implementation, then replace placeholders with qualified evidence. Present the +user's request, returned result and correction first; configuration details +belong in setup help. Publication is separate from preparing the kit. + +Each release sentence must trace to the tested package/version, entry point and +user-visible result. “Open source” describes LoopX under its existing license; +models, paid usage, third-party connectors and hosted infrastructure keep their +own terms. No license change, public post, live service provisioning or permission +expansion follows from this roadmap revision. + ### Cross-Domain Measurement and Validation | Dimension | Shared definition | Minimum evidence | diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md index 77a8cae363..4a2a7d2a60 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md @@ -85,6 +85,102 @@ S 工作流表示长期责任,G 里程碑表示一次可验收的产品组合 暂停/降级条件:新方案复制权威或扩大默认权限、真实 backend 资格缺失、改动无法独立回滚、成本或人工干预随规模失控、benchmark integrity 失败、用户无法解释当前工作与阻塞。保留事实/回执,停止相关新 admission,修复规则或缩小 cohort;不能靠放松验收、删失败样本、刷测试预期或换名晋级。 +### 近期本地 Agent 产品与发布路线 + +**规划修订:2026-09-20;属于提案,不代表版本验收通过。** +把 native 路线收敛成 **用已有本地 Agent 持续办事、带回可检查成果的开源个人 Agent**。 +首批用户是已拥有受支持 Agent 登录态的开发者和重度研究用户。先面对一位长期管家, +需要时再展开专家团队。这是 S1/S4/S5/S12/S13 对 G0/G1 与早期 G5 的有界交付, +不新增运行时、平行路线图,也不阻塞 observer-first 路线。 + +“让你的 local Agent 像 Grok Bot / Muse 一样工作”可以作为理解入口。 +主要承诺应是“给 Agent 一个目标,回来验收成果”。比较必须说明支持范围,不宣称 +完整开源替代、模型能力等价、免费推理、保证无人值守成功或官方关联。本地编排不 +等于离线推理,也不代表模型服务商不接收任务数据。后台执行依赖宿主在线;关闭 +浏览器与关闭电脑是两件事。 + +#### 调研依据与产品取舍 + +以下来源于 2026-09-20 检索;厂商能力属于 **声明,未经本次实测**。 + +| 原始来源 | 有关发现 | LoopX 取舍与限制 | +| --- | --- | --- | +| [Grok Bot 产品页](https://x.ai/bot)与[发布说明](https://x.ai/news/introducing-grok-bot) | 厂商描述常驻云端队友、共享工具、工作交接与成果返回 | 借鉴对话委派和可见成果;本地宿主不能继承云端 24/7 承诺 | +| [Meta Muse 发布](https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/)与[设计说明](https://introducing.muse.ai/) | 厂商描述长期主对话、Goals、有价值才主动通知、丰富产物和明确操作控件 | 以一位管家和首次有用成果为中心;重要动作采用确定性控件,工具细节渐进展开 | +| [X 原始体验比较](https://x.com/Michaelzsguo/status/2101057424115253645) | 一位用户更认可 Muse 的成品体验,而非配置组件 | 作为产品化假设,不作为性能优劣证据;另做独立首次使用验证 | + +X 原帖通过 Ego Lite 阅读;官方文章已读取,未登录实测产品或完整观看演示视频。 +Muse 设计页在浏览器超时,其文章通过网页检索读取。本次调研不能证明实际可靠性、 +隐私保护等价或转化率。 + +#### 一条首次使用路径,一个标志性成果 + +1. 沿现有安装 owner 安装发行包,打开 `loopx dashboard`。在原处诊断 Python/Node、 + 宿主登录、所选模型/profile 和工具就绪;安装成功不能直接显示 Agent 已就绪。 +2. 显式引导配置可用运行时、模型账户、工具、预算与权限。以用户明确意图为约束, + Agent 在范围内自主选择和分配。锁定管家/Chat profile 时必须遵守;授权灵活资源池 + 时可自主选路。Codex 与 Ark 都是配置方案,不是统一偏好。展示有效选择,不复制 + 密钥、不迁移旧会话、不静默越权回退。缺少条件时给出具体修复入口。 +3. 直接交给管家真实任务,不先建团队,不要求理解 Goal/Todo/Turn。预览有关范围、 + 交付物与资源上限;已经授权、可逆的日常步骤不反复弹确认卡。 +4. 从公开资料交付带来源的研究简报。加入矛盾或更新资料,展现独立异议、精确修订、 + 验收和请求方采用后的综合结论。团队不可用时,单 Agent 的首个成果仍应有用。 +5. 把产物和下一个值得决策的问题带回原对话。可以查证据、改方向、暂停或停止明确的 + 执行范围;重开浏览器能恢复上下文,过期/失联状态明确。主动通知需要实质变化。 + +同一案例同时服务产品、可复现公开样例和宣传片。领域研究可建立在它之上,但金融 +建议与交易权限不进入通用内核。工程 issue→审阅后的补丁是下一个已有 owner 的 +场景,不成为首次研究预览的第二个阻塞项。Lark 后续作为可选 transport,复用相同 +会话、受众和结果 owner;本地首次使用不要求配置 Lark。每条宣称可用的本地路径 +均保留独立 CLI 回读。 + +#### 交付顺序、owner 与可宣传边界 + +下列是人员到位后的累计工作时间估计,假设一位专职工程师配合 Agent 辅助、共享 +设计/QA 支持、维护者每日可 review、先验收一种 macOS + Codex 配置,且模型额度 +和测试工具授权可用。不是实测速度、SLA 或全宿主承诺。依赖没到位就顺延日期, +不能降低验收标准。 + +| 阶段 / 累计目标 | 完整工作与既有 owner | 出口 / 可以宣称什么 | +| --- | --- | --- | +| 定位与分镜 · 2–3 个工作日 | S1/S5/S13:一个人群、一个任务、首屏设计、60–90 秒脚本、双语文案、公开证据计划 | 可审阅的精美草稿;概念与回放明确标注,不宣称产品已可用 | +| 产品实录预览 · 5–7 个工作日 | R2/R3 + S5:review/集成 #4762,解决受支持宿主工具授权,跑通一次真实 L1 纠偏,打磨真实结果与证据视图 | 打包前端 + CLI 展示精确版本的异议/修订/验收/采用、干预和范围准确的停止;实录运行,不把合成活动当 live | +| 聚焦公开 alpha · 2–3 周 | S1/S4/S12:发行包引导、遵守用户意图的管家/Chat 选路、就绪修复、单 Agent 降级、重启/重连和有用结果返回;S5 提供聚焦团队视图 | 测试前冻结五次独立干净安装,至少四次无需维护者 shell 救场拿到首个有用成果,保留全部失败。宣传持续团队协作前必须通过 G1 两轮真实协作 | +| 可重复 beta / 完整发布包 · 4–6 周 | S4/S5/S10/S12/S13:升级/回滚、撤权/失联/额度恢复、试用反馈驱动 UX、实测成本与支持负担;可选 Lark 单独验收 | 至少三位独立用户隔天再次完成任务,公开样本分母、干预和限制;每个宣传平台/transport 都有发行包证据 | + +建议 alpha 易用性目标,在招募前冻结:满足文档前置条件并登录后,10 分钟内配置 +完成;明确预算的有界样例 15 分钟内拿到首个有用产物。这是拟定目标,不是当前 +性能数据。同时记录包含环境准备/登录的完整引导耗时与就绪后耗时,认证或环境 +失败仍计入漏斗。小规模试用用于发现问题,不能据此宣称统计可靠性或 PMF。 + +关键路径是:宿主/工具就绪 → 一次真实 L1 → 打包首次使用修复 → 独立复现。 +本修订时 #4758 已合并但属于设计,#4762 仍开放,不能证明已发布或真实模型验收。 +其中 GPT 尝试被 MCP 审批拦截,没有被验收的产物;通过所属 owner 解决受支持授权 +路径,不能为了录屏全局关闭审批。既有实施 successor 保持权威;分配新工作前核对 +canonical Todos 和相关 PR。 + +首次发布不等待百 Agent、新状态 provider、完整 3D、全宿主、原生移动壳、连接器 +市场、支付/订票或示教自动化。这些保留既有 owner 与独立需求/验收,不挤占 R1 +可靠性工作。 + +#### 视觉与宣传资产 + +标志性瞬间是 **同伴找到更好的证据,结论因此改变**。精确指挥台承载对话、结果和 +决策;稳定空间位置与局部事件动画解释交接。展示前后结论、引用来源和被验收采用 +的版本。一次突出一件事,保留可访问列表、减少动态以及诚实的安静/失败状态。 +不能用持续动画或 Agent 数量替代工作。遵循既有[设计契约](../../development/design.md), +实际公开首屏依然先提供具体预览并取得批准。 + +同步准备统一发布包:60–90 秒真实运行影片(披露时间剪辑)、15 秒短版、三张有 +注释的产品截图、标明历史的可访问回放、双语发布文章与社交短稿、一条安装/试用 +路径、支持版本/成本/隐私/在线条件 FAQ,以及含预期结果、纠偏输入和恢复步骤的 +公开安全样例。随实现准备草稿,再用合格证据替换占位;先讲用户请求、成果和纠偏, +配置细节放进帮助。准备发布包不等于已授权对外发帖。 + +每句发布承诺关联被测包/版本、入口与用户可见结果。“开源”按 LoopX 既有 license +解释;模型、付费额度、第三方连接器与托管设施各有条件。本修订不改变 license, +不发布社交内容、不开通在线服务,也不扩大权限。 + ### 跨领域指标与验证矩阵 | 维度 | 统一口径 | 最低证据 |