Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/architecture/rfcs/loopx-overall-roadmap-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -411,7 +411,7 @@ sessions, generic Agent creation, dynamic governed work derivation, complete
inbox/queue/steer, authenticated remote authority and packaged frontend/Lark
companion work remain R2/R3/R4/R6 boundaries. Existing Goals are not promoted.

An existing shell-capable coordinator uses `delegation list/operations/start/read/wait/resume` without replacing its session. Requester-scoped `operations` recovers durable work after context loss, independently rechecks accepted results and preserves unavailable branches and pagination; enabled MCP and newly tool-equipped Goal Chat use the same read model. Existing native threads retain their tool schema on resume. It starts no work and does not infer overall readiness from a display list. The example's `prepare` still only provisions isolated operator bindings. Next, feed actual execution/acceptance facts into existing R2 readiness, then extend existing registration/runtime configuration for approved identity/profile provisioning and qualify original-request return/lead continuation. Unattended wake, full cross-host inbox/queue/steer and Lark parity remain separate requirements; no G1/G3 promotion follows from this recovery entrypoint.
An existing shell-capable coordinator uses `delegation list/operations/start/read/wait/resume` without replacing its session. Requester-scoped `operations` recovers durable work after context loss, independently rechecks accepted results and preserves unavailable branches and pagination; enabled MCP and newly tool-equipped Goal Chat use the same read model. Existing native threads retain their tool schema on resume. It starts no work and does not infer overall readiness from a display list. The example's `prepare` still only provisions isolated operator bindings. A Codex binding can now pass an exact model/reasoning effort into an independent resumable Turn Session and expose the same profile through preflight and planning; this remains separate from a native temporary child profile inside the parent execution. Next, feed actual execution/acceptance facts into existing R2 readiness, extend registration/runtime configuration for approved identity provisioning, and qualify original-request return/lead continuation. Unattended wake, full cross-host inbox/queue/steer and Lark parity remain separate requirements; an exact profile parameter does not promote G1/G3.

The local Goal conversation now exposes that inventory on demand, with per-binding
preflight through the actual Turn dry-run and selected executor/profile. Task
Expand Down
2 changes: 1 addition & 1 deletion docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -360,7 +360,7 @@ Todo 完成入口分别执行当前 pinned 检查,accepted 返回读 canonical
完成。长期 attached 会话、通用 Agent 创建、动态受治理工作派生、完整 inbox/queue/steer、
认证远端权威与 packaged frontend/Lark 配套仍归 R2/R3/R4/R6;不晋升已有 Goal。

有 shell 能力的原 coordinator 可通过 `delegation list/operations/start/read/wait/resume` 调用已有执行 owner,无需替换会话。`operations` 从自身持久记录找回上下文丢失前的工作,重新核验 accepted,保留不可用分支与分页;已启用的 MCP 和新挂载工具的 Goal Chat 共用该读模型;已有原生线程恢复时保留原工具 schema。读取不启动工作,也不把展示列表当整体 readiness。合成示例 `prepare` 仍只准备隔离绑定。下一步先将实际执行/验收事实接入已有 R2 readiness,再沿现有注册及 runtime 配置扩展经授权的身份/profile 创建,并验证原请求返回与主力继续推进。无人值守唤醒、完整跨宿主 inbox/queue/steer 和 Lark 等价仍分别验收,不因新增恢复入口晋升 G1/G3。
有 shell 能力的原 coordinator 可通过 `delegation list/operations/start/read/wait/resume` 调用已有执行 owner,无需替换会话。`operations` 从自身持久记录找回上下文丢失前的工作,重新核验 accepted,保留不可用分支与分页;已启用的 MCP 和新挂载工具的 Goal Chat 共用该读模型;已有原生线程恢复时保留原工具 schema。读取不启动工作,也不把展示列表当整体 readiness。合成示例 `prepare` 仍只准备隔离绑定。Codex binding 现可将精确 model/reasoning effort 送入独立且可续接的 Turn Session,并由同一 preflight/规划投影读回;这与父执行内部的原生临时 child profile 分开。下一步先将实际执行/验收事实接入已有 R2 readiness,再沿现有注册及 runtime 配置扩展经授权的身份创建,并验证原请求返回与主力继续推进。无人值守唤醒、完整跨宿主 inbox/queue/steer 和 Lark 等价仍分别验收,不因 profile 参数接通而晋升 G1/G3。

本地 Goal 对话现可按需读取该持久目录,并按绑定调用实际 Turn dry-run 与所选
executor/profile 检查启动条件。任务准入、当前 pinned 验收绑定、运行时可用性分开
Expand Down
38 changes: 38 additions & 0 deletions docs/reference/local-delegation.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,44 @@ select `generic-cli`, `fresh`, and the optional adapter's `--config` invocation.
Profiles, executables, workspace isolation and credential custody remain the
operator's responsibility. No model tool accepts those values.

A Codex binding launches an independent, resumable Codex Agent Session through
the same governed Turn path. Pin both fields when the worker must use an exact
profile:

```json
{
"id": "strong-independent-review",
"agent_id": "managed-reviewer",
"todo_id": "todo_review",
"requesters": ["lead"],
"workspace": "/absolute/reviewer-worktree",
"host_args": [
"--host", "codex-cli",
"--codex-model", "gpt-5.6-sol",
"--codex-reasoning-effort", "xhigh",
"--codex-sandbox", "workspace-write"
],
"timeout_seconds": 300,
"output_refs": ["output.json"]
}
```

This is not a native `multi_subagent` child. Native children remain temporary
workers inside one parent execution and use the Goal's child model preference.
The binding above has its own Agent identity, Todo, workspace, Codex Session and
durable delegation operation. `delegation inspect` and the planning projection
show `gpt-5.6-sol@xhigh`; start and resume pass both fields to that same Session.
The requester still needs the exact binding grant, and a profile is not an
acceptance or result-return receipt.

中文:Codex binding 通过同一条受治理 Turn 链启动独立、可续接的 Codex Agent
Session。需要精确执行配置时同时固定 `--codex-model` 与
`--codex-reasoning-effort`。它不是 `multi_subagent` 的原生临时 child:后者仍在
单个父执行内部使用 Goal 的 child model 偏好;前者拥有独立 Agent 身份、Todo、
workspace、Codex Session 与持久 delegation operation。`delegation inspect` 和
规划投影会读回例如 `gpt-5.6-sol@xhigh`,start/resume 也把同一配置送入原 Session。
这不扩大 requester grant,也不把 profile 冒充验收或结果返回回执。

## Use an existing Agent conversation through its shell

An attached Codex or other shell-capable Agent can use the same execution
Expand Down
29 changes: 25 additions & 4 deletions loopx/cli_commands/turn.py
Original file line number Diff line number Diff line change
Expand Up @@ -195,10 +195,30 @@ def handle_turn_command(
# have to name the same credential.
environ=operator_environ,
dsh_runner_configured=bool(getattr(args, "dsh_runner", None)),
provider=getattr(args, "dsh_provider", None),
model=getattr(args, "dsh_model", None),
reasoning_effort=getattr(args, "dsh_reasoning_effort", None),
max_tokens=getattr(args, "dsh_max_tokens", None),
provider=(
getattr(args, "dsh_provider", None)
if args.host == "dsh"
else None
),
model=(
getattr(args, "dsh_model", None)
if args.host == "dsh"
else getattr(args, "codex_model", None)
if args.host == "codex-cli"
else None
),
reasoning_effort=(
getattr(args, "dsh_reasoning_effort", None)
if args.host == "dsh"
else getattr(args, "codex_reasoning_effort", None)
if args.host == "codex-cli"
else None
),
max_tokens=(
getattr(args, "dsh_max_tokens", None)
if args.host == "dsh"
else None
),
)
if (
args.turn_command == "run-once"
Expand Down Expand Up @@ -974,6 +994,7 @@ def run_built_in_host(
codex_bin=args.codex_bin,
sandbox=args.codex_sandbox,
model=args.codex_model,
reasoning_effort=args.codex_reasoning_effort,
timeout_seconds=max(1.0, args.timeout_seconds - 5.0),
)

Expand Down
10 changes: 10 additions & 0 deletions loopx/cli_commands/turn_registration.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
MANAGED_TURN_HOST,
resolve_default_turn_host,
)
from ..control_plane.turn_driver.execution_profile import REASONING_EFFORTS
from ..paths import default_public_scan_root

# Explicit host choices stay per-command: planning may name any host the Turn
Expand Down Expand Up @@ -208,6 +209,15 @@ def register_turn_commands(
help="Codex CLI executable used by the built-in codex-cli host.",
)
run_once.add_argument("--codex-model")
run_once.add_argument(
"--codex-reasoning-effort",
choices=list(REASONING_EFFORTS),
help=(
"Reasoning effort for the independent Codex CLI Turn. This is an "
"operator-bound independent Agent profile, not the native child-agent "
"model preference."
),
)
run_once.add_argument(
"--codex-sandbox",
choices=["read-only", "workspace-write", "danger-full-access"],
Expand Down
13 changes: 13 additions & 0 deletions loopx/control_plane/turn_driver/codex_cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
HOST_RESULT_TEXT_LIMITS,
LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION,
)
from .execution_profile import require_supported_reasoning_effort
from .host_failure import BuiltInHostError
from .transaction import LOOPX_TURN_RESULT_SCHEMA_VERSION, TRANSACTION_PHASES

Expand Down Expand Up @@ -665,6 +666,7 @@ def _codex_command(
output_path: Path,
sandbox: str,
model: str | None,
reasoning_effort: str | None,
session_id: str | None,
) -> list[str]:
if session_id:
Expand Down Expand Up @@ -700,6 +702,13 @@ def _codex_command(
]
if model:
command.extend(["--model", model])
if reasoning_effort:
command.extend(
[
"-c",
f"model_reasoning_effort={json.dumps(reasoning_effort)}",
]
)
if session_id:
command.append(session_id)
command.append("-")
Expand All @@ -714,12 +723,15 @@ def run_codex_cli_host(
codex_bin: str = "codex",
sandbox: str = "read-only",
model: str | None = None,
reasoning_effort: str | None = None,
timeout_seconds: float = 115.0,
) -> dict[str, Any]:
if request.get("schema_version") != LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION:
raise ValueError("unsupported LoopX Turn host request schema")
if sandbox not in CODEX_CLI_SANDBOXES:
raise ValueError(f"Codex CLI sandbox must be one of {CODEX_CLI_SANDBOXES}")
if reasoning_effort is not None:
reasoning_effort = require_supported_reasoning_effort(reasoning_effort)
resolved = shutil.which(codex_bin) if os.path.sep not in codex_bin else codex_bin
if not resolved or not Path(resolved).exists():
raise ValueError("Codex CLI executable is unavailable")
Expand Down Expand Up @@ -760,6 +772,7 @@ def run_codex_cli_host(
output_path=output_path,
sandbox=sandbox,
model=model,
reasoning_effort=reasoning_effort,
session_id=session_id,
)
proc = subprocess.Popen(
Expand Down
38 changes: 30 additions & 8 deletions loopx/control_plane/turn_driver/host_binding.py
Original file line number Diff line number Diff line change
Expand Up @@ -195,10 +195,11 @@ def managed_executor_binding(
records that this projection does not probe that executor kind, so it makes
no claim rather than an unproven ``True``.

``execution_profile`` is the managed profile this executor would run -- and,
with it, the provider claims to authenticate. It is ``None`` for every
non-managed executor because neither the profile nor the credential belongs
to an individual or generic host.
``execution_profile`` is the explicit profile this executor would run.
Managed DSH resolves its complete provider profile. An individual Codex
executor projects only operator-pinned model/effort arguments; absent fields
remain ``host-default`` rather than being guessed. Generic hosts have no
shared profile contract and keep this field ``None``.

``runtime_probe`` states what the ``dsh_runtime_unavailable`` verdict is a
claim about, so a reader does not take a process-level answer for a
Expand Down Expand Up @@ -251,6 +252,11 @@ def managed_executor_binding(
"available": runtime_available,
},
}
individual_profile: str | None = None
if host == INDIVIDUAL_TURN_HOST and (model or reasoning_effort):
individual_profile = (
f"{model or 'host-default'}@{reasoning_effort or 'host-default'}"
)
return {
"schema_version": MANAGED_EXECUTOR_BINDING_SCHEMA_VERSION,
"executor": host,
Expand All @@ -261,7 +267,7 @@ def managed_executor_binding(
),
"credential_env": None,
"endpoint_env": None,
"execution_profile": None,
"execution_profile": individual_profile,
"operator_credential_bound": False,
"available": None,
"unavailable_reason": None,
Expand Down Expand Up @@ -313,9 +319,25 @@ def managed_executor_binding_from_host_args(
host,
environ=environ,
module_probe=module_probe,
provider=turn_host_arg_option(host_args, "--dsh-provider"),
model=turn_host_arg_option(host_args, "--dsh-model"),
reasoning_effort=turn_host_arg_option(host_args, "--dsh-reasoning-effort"),
provider=(
turn_host_arg_option(host_args, "--dsh-provider")
if host == MANAGED_HOST
else None
),
model=(
turn_host_arg_option(host_args, "--dsh-model")
if host == MANAGED_HOST
else turn_host_arg_option(host_args, "--codex-model")
if host == INDIVIDUAL_TURN_HOST
else None
),
reasoning_effort=(
turn_host_arg_option(host_args, "--dsh-reasoning-effort")
if host == MANAGED_HOST
else turn_host_arg_option(host_args, "--codex-reasoning-effort")
if host == INDIVIDUAL_TURN_HOST
else None
),
)


Expand Down
27 changes: 27 additions & 0 deletions tests/test_delegation_preflight.py
Original file line number Diff line number Diff line change
Expand Up @@ -168,6 +168,33 @@ def test_selected_dsh_profile_is_not_replaced_by_the_default(service):
assert not (root / "host-started").exists()


def test_selected_codex_managed_agent_profile_is_projected_exactly(service):
root, runner = service
config = json.loads(runner.config.read_text())
config["bindings"][0]["host_args"] = [
"--host",
"codex-cli",
"--codex-model",
"gpt-5.6-sol",
"--codex-reasoning-effort",
"xhigh",
]
runner.config.write_text(json.dumps(config))

status, result = cli(runner, "inspect", "--binding-id", "analysis")

assert status == 0, result
assert result["executor"] == {
"host": "codex-cli",
"available": None,
"reason": None,
"profile": "gpt-5.6-sol@xhigh",
}
assert result["state"] == "runtime_unverified"
assert not any(result["effects"].values())
assert not (root / "host-started").exists()


def test_preflight_does_not_call_an_invalidated_acceptance_ready(service):
root, runner = service
from loopx.agent_registry import load_goal_from_registry
Expand Down
34 changes: 34 additions & 0 deletions tests/test_loopx_turn_codex_cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -368,6 +368,8 @@ def test_codex_cli_host_starts_then_resumes_opaque_session(
project=project,
codex_bin=str(executable),
sandbox=sandbox,
model="gpt-5.6-sol",
reasoning_effort="xhigh",
timeout_seconds=5,
)
with pytest.raises(RuntimeError, match="binding changed after planning"):
Expand All @@ -388,6 +390,8 @@ def test_codex_cli_host_starts_then_resumes_opaque_session(
project=project,
codex_bin=str(executable),
sandbox=sandbox,
model="gpt-5.6-sol",
reasoning_effort="xhigh",
timeout_seconds=5,
)

Expand All @@ -399,6 +403,14 @@ def test_codex_cli_host_starts_then_resumes_opaque_session(
assert "resume" not in argv_rows[0]
assert "resume" in argv_rows[1]
assert "session-fixture-0001" in argv_rows[1]
for argv in argv_rows:
assert argv[argv.index("--model") + 1] == "gpt-5.6-sol"
config_values = [
argv[index + 1]
for index, value in enumerate(argv)
if value == "-c"
]
assert 'model_reasoning_effort="xhigh"' in config_values
resume_argv = argv_rows[1]
assert resume_argv[resume_argv.index("-c") + 1] == (
f'sandbox_mode="{sandbox}"'
Expand Down Expand Up @@ -435,6 +447,28 @@ def test_codex_cli_host_starts_then_resumes_opaque_session(
assert "private_material" not in persisted


def test_codex_cli_host_rejects_unknown_reasoning_effort_before_launch(
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
executable, log_path = _fake_codex(tmp_path)
monkeypatch.setenv("FAKE_CODEX_LOG", str(log_path))
project = tmp_path / "project"
project.mkdir()

with pytest.raises(ValueError, match="unsupported reasoning effort"):
run_codex_cli_host(
_request(),
runtime_root=tmp_path / "runtime",
project=project,
codex_bin=str(executable),
reasoning_effort="turbo",
timeout_seconds=5,
)

assert not log_path.exists()


def test_codex_cli_host_fresh_iteration_ignores_stored_session(
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
Expand Down
17 changes: 17 additions & 0 deletions tests/test_turn_default_host_binding.py
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,23 @@ def test_explicit_host_flag_wins_over_the_default(monkeypatch):
assert args.host == "generic-cli"


def test_codex_managed_agent_profile_is_explicit_cli_configuration():
args = build_parser().parse_args(
[
*_turn_argv("run-once"),
"--host",
"codex-cli",
"--codex-model",
"gpt-5.6-sol",
"--codex-reasoning-effort",
"xhigh",
]
)

assert args.codex_model == "gpt-5.6-sol"
assert args.codex_reasoning_effort == "xhigh"


@pytest.mark.parametrize("command", ["plan", "run-once"])
def test_default_execution_mode_follows_the_selected_host(command, monkeypatch):
for name in ("DEEPSEEK_API_KEY", TURN_HOST_ENV_VAR):
Expand Down
Loading
Loading