Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions docs/architecture/rfcs/loopx-overall-roadmap-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -424,6 +424,16 @@ sessions, generic Agent creation, dynamic governed work derivation, complete
inbox/queue/steer, authenticated remote authority and packaged frontend/Lark
companion work remain R2/R3/R4/R6 boundaries. Existing Goals are not promoted.

The disposable example now also accepts `--team-size LUNA DSH ARK`: independent
Luna max Turns consume the initial filing, DSH consumes accepted analysis and
the correction, and Ark consumes accepted corrected evidence. Preparation
reuses machine credentials and creates only Goal-owned bindings; the DSH lead
uses the same delegation service and must adopt every configured result. This
removes hand-edited roster setup for the synthetic qualification route. It does
not close the release frontier: actual public-source research, a visual launch
control, original-conversation result return and member stop/recovery must be
qualified together before advertising the one-action research showcase.

An existing shell-capable coordinator uses `delegation list/operations/start/read/wait/resume` without replacing its session. Requester-scoped `operations` recovers durable work after context loss, independently rechecks accepted results and preserves unavailable branches and pagination; enabled MCP and newly tool-equipped Goal Chat use the same read model. Existing native threads retain their tool schema on resume. It starts no work and does not infer overall readiness from a display list. The example's `prepare` still only provisions isolated operator bindings. A Codex binding can now pass an exact model/reasoning effort into an independent resumable Turn Session and expose the same profile through preflight and planning; this remains separate from a native temporary child profile inside the parent execution. Next, feed actual execution/acceptance facts into existing R2 readiness, extend registration/runtime configuration for approved identity provisioning, and qualify original-request return/lead continuation. Unattended wake, full cross-host inbox/queue/steer and Lark parity remain separate requirements; an exact profile parameter does not promote G1/G3.

The local Goal conversation now exposes that inventory on demand, with per-binding
Expand Down
7 changes: 7 additions & 0 deletions docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -369,6 +369,13 @@ Todo 完成入口分别执行当前 pinned 检查,accepted 返回读 canonical
完成。长期 attached 会话、通用 Agent 创建、动态受治理工作派生、完整 inbox/queue/steer、
认证远端权威与 packaged frontend/Lark 配套仍归 R2/R3/R4/R6;不晋升已有 Goal。

隔离示例现在支持 `--team-size LUNA DSH ARK`:独立 Luna max Turn 分析原始财报,
DSH 采用已验收分析并处理修订,Ark 采用已验收的修订证据。准备阶段复用机器凭证,
仅生成 Goal 范围的执行绑定;DSH 协调员沿用同一委派服务,最终必须采用每位成员的
产物。这消除了合成验收场景中手工修改成员配置的步骤,尚未关闭 release 缺口:
真实公开资料投研、可视化启动、原对话返回和成员停止/恢复,需要一起验证后才能
宣传一键投研 showcase。

有 shell 能力的原 coordinator 可通过 `delegation list/operations/start/read/wait/resume` 调用已有执行 owner,无需替换会话。`operations` 从自身持久记录找回上下文丢失前的工作,重新核验 accepted,保留不可用分支与分页;已启用的 MCP 和新挂载工具的 Goal Chat 共用该读模型;已有原生线程恢复时保留原工具 schema。读取不启动工作,也不把展示列表当整体 readiness。合成示例 `prepare` 仍只准备隔离绑定。Codex binding 现可将精确 model/reasoning effort 送入独立且可续接的 Turn Session,并由同一 preflight/规划投影读回;这与父执行内部的原生临时 child profile 分开。下一步先将实际执行/验收事实接入已有 R2 readiness,再沿现有注册及 runtime 配置扩展经授权的身份创建,并验证原请求返回与主力继续推进。无人值守唤醒、完整跨宿主 inbox/queue/steer 和 Lark 等价仍分别验收,不因 profile 参数接通而晋升 G1/G3。

本地 Goal 对话现可按需读取该持久目录,并按绑定调用实际 Turn dry-run 与所选
Expand Down
76 changes: 68 additions & 8 deletions examples/managed-research-team/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,9 @@ uv sync --extra test --extra deepseek-harness
uv pip install --python .venv/bin/python -e packages/loopx-ark-turn
```

Set `ARK_API_KEY`, `ARK_MODEL_ID`, `ARK_ENVIRONMENT_ID` and `DEEPSEEK_API_KEY`
in the environment. The Ark Environment must already belong to the operator;
Set `ARK_API_KEY`, `ARK_MODEL_ID` and `ARK_ENVIRONMENT_ID` in the process
environment. DSH reuses the machine operator credential; `DEEPSEEK_API_KEY`
is an explicit environment override. The Ark Environment must already belong to the operator;
this launcher never creates or deletes it. The local model defaults to
`deepseek-v4-flash@high`; select another profile with `--dsh-model`.

Expand Down Expand Up @@ -59,6 +60,65 @@ uv run --no-sync --extra test python -m loopx.cli \
--format json goal-acceptance verify --goal-id synthetic-managed-research --execute
```

## Configure a Luna / DSH / Ark team

Use `--team-size LUNA DSH ARK` to select positive counts for all three member
runtimes. Counts exclude the coordinator. For example, this single command
prepares the isolated Goal and starts a DSH coordinator with one member per
runtime:

```bash
uv run --no-sync --extra test python examples/managed-research-team/research_team.py \
run "$DEMO_ROOT" --team-size 1 1 1 \
--model "$ARK_MODEL_ID" --environment-id "$ARK_ENVIRONMENT_ID"
```

Codex members use independent `gpt-5.6-luna@max` Turns and the machine's existing
Codex login. Install `codex` on `PATH` first. DSH uses the machine operator
credential or an explicit process environment override; no credential is
copied into the Goal. Ark uses the selected model and existing Environment.
This does not advertise a different DSH model version than the one the provider
actually serves. A generated binding is configuration, not a successful login
or model-availability check. `run` makes real, billable provider calls.

Luna members analyze the initial filing. DSH members consume an accepted Luna
artifact and analyze the correction; Ark members check an accepted DSH
artifact. Predecessors are assigned round-robin within this sample graph. The
coordinator chooses execution order and questions through the existing
collaboration tools. The report must adopt **every** configured member,
including an extra Luna member without a downstream consumer. Missing canonical
completion or a changed artifact prevents report acceptance. Repeating the
same analysis with more models does not create independent source families.

`--team-size 2 1 1` creates four member tasks. The same option works with
`prepare` and `prepare-chat`, which make **no model calls**. It cannot be combined
with `--topology`; omitting it preserves the existing local-led four-member
example. Zero/negative counts are rejected before creating a directory. More
members increase spend and can exhaust the isolated Goal's quota or lead's
20-minute deadline; this option is not a capacity or cost guarantee.

Read `project/team.json`, `delegation-config.json` and canonical Todo state to
inspect the generated roster. The `validate-report` and acceptance commands
above recheck every configured dependency. The Chat route below consumes the
same generated configuration; it still requires explicit owner settings and
enabling, rather than silently starting from preparation.

This remains a **synthetic acceptance example**, not a financial-research
showcase or proof of autonomous source discovery. An interrupted run is not
successful: use the existing delegation inventory/read/wait/resume operations
for the original executions. Do not rerun `run` against the same directory or
start replacements while their status is unknown. Pausing a Chat coordinator
does not stop already running members; retain their receipts and use the
individual provider's execution controls. Removing bindings prevents new
admission but does not cancel accepted work.

中文:`--team-size 1 1 1` 表示一名 Luna max、一名 DSH、一名 Ark 成员,协调员
另计。`run` 准备隔离团队并发起真实模型执行;`prepare`/`prepare-chat` 只准备。
可改为 `2 1 1` 等正整数;每名成员都必须通过独立验收并被最终报告采用,不能用
人数或注册成功代替协作证据。凭证复用机器配置,Goal 只持有分工和授权。此处是
明确标注的合成财报验收场景,真实投研、可视化一键启动、团队级停止和宣传影片
仍需分别验证。暂停协调员不会取消成员;运行中断时先恢复原执行,勿重复拉起。

## Goal Chat coordinator

To use the local Goal conversation as the lead, prepare a new disposable team
Expand All @@ -82,7 +142,7 @@ advances business phases. Pause while a member works, refresh the page, then
continue to observe that original member's accepted result. Queue a correction
for the next native turn or explicitly select inbox/steer.

This route returns the report in the conversation; the four member tasks have
This route returns the report in the conversation; all configured member tasks have
independent canonical acceptance. It does **not** write `lead/report.json` or
complete `todo_lead-report`. The `validate-report` command above applies to the
DSH/Ark lead route, which has an explicitly bound report-writing tool. Both
Expand All @@ -92,7 +152,7 @@ command above; do not infer report acceptance from a native completion label.
中文:用 `prepare-chat` 准备隔离团队,再启动上述本地 Chat。进入该 Goal 的
对话,原地选择 `lead`、已生成的执行配置和协调员额度,开启后让模型组织协作。
可在成员执行时暂停、刷新、恢复,检查成员结果仍回到原对话。此入口把综合报告
返回对话;四个成员任务分别验收,报告 Todo 和整体 Goal 保留给所有者处理。
返回对话;各个成员任务分别验收,报告 Todo 和整体 Goal 保留给所有者处理。
详见 [Goal 对话运行模式](../../docs/reference/goal-chat-continuation.md)。

## Collaboration path
Expand Down Expand Up @@ -121,7 +181,7 @@ its own analysis while members run. The nested cloud analyst still requests
its local reviewer through the same service. No business phase argument is
introduced.

After reading all four canonical completions and exact artifact hashes, the
After reading all configured canonical completions and exact artifact hashes, the
lead writes `lead/report.json` with the fields described by `scenario.py` and
the acceptance table below. Run `validate-report`, then complete the report
through ordinary `todo complete --todo-id todo_lead-report --agent-id lead
Expand Down Expand Up @@ -162,16 +222,16 @@ an additional route. It does not substitute for the primary local-led path.
| Repost of issuer material | Same source | Old figures retained | One current-period source family; corrected repost is stale |

`bootstrap.ts` creates only a fresh disposable canonical runtime and invokes
the production owner configuration API once. It binds four member criteria and
the production owner configuration API once. It binds each configured member criterion and
one report criterion. Task instructions, the roster and verifier files are
pinned; a member cannot change its own acceptance. Existing Goals are never
promoted or rewritten by this bootstrap.

Core delegation asks the TS acceptance owner for the exact task's criteria and
runs those checks as its Turn validator. It then uses ordinary `todo complete`,
which re-executes validation and atomically commits through the same TS owner.
The report separately checks all four canonical completions, current binding
guards, adopted hashes and financial conclusions. All five Todos may be done
The report separately checks all configured canonical completions, current binding
guards, adopted hashes and financial conclusions. All configured Todos may be done
while the overall Goal remains active for its owner.

A member's own peer conclusion is preserved. The delegation result independently
Expand Down
5 changes: 5 additions & 0 deletions examples/managed-research-team/execution.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,11 @@ def host_arguments(root: Path, actor: str, revision: str, *, host: str, attempt:
settings = json.loads((root / "settings.json").read_text())
coordinator = actor == "lead"
workspace = root / "lead" if coordinator else root / actor / revision
if host == "codex":
# Delegation injects its existing native MCP tools. Model/effort are an
# independent Turn binding; no user profile or credential is copied.
return ["--host", "codex-cli", "--codex-model", "gpt-5.6-luna",
"--codex-reasoning-effort", "max", "--codex-sandbox", "workspace-write"]
if host == "dsh":
args = ["--host", "dsh", "--dsh-model", settings["dsh_model"], "--dsh-reasoning-effort", "high",
"--dsh-home", str(root / "homes" / (actor + "-" + revision + "-" + str(attempt)))]
Expand Down
Loading
Loading