diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md index a01b7459c2..023aff1773 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md @@ -411,7 +411,7 @@ sessions, generic Agent creation, dynamic governed work derivation, complete inbox/queue/steer, authenticated remote authority and packaged frontend/Lark companion work remain R2/R3/R4/R6 boundaries. Existing Goals are not promoted. -An existing shell-capable coordinator uses `delegation list/operations/start/read/wait/resume` without replacing its session. Requester-scoped `operations` recovers durable work after context loss, independently rechecks accepted results and preserves unavailable branches and pagination; enabled MCP and newly tool-equipped Goal Chat use the same read model. Existing native threads retain their tool schema on resume. It starts no work and does not infer overall readiness from a display list. The example's `prepare` still only provisions isolated operator bindings. Next, feed actual execution/acceptance facts into existing R2 readiness, then extend existing registration/runtime configuration for approved identity/profile provisioning and qualify original-request return/lead continuation. Unattended wake, full cross-host inbox/queue/steer and Lark parity remain separate requirements; no G1/G3 promotion follows from this recovery entrypoint. +An existing shell-capable coordinator uses `delegation list/operations/start/read/wait/resume` without replacing its session. Requester-scoped `operations` recovers durable work after context loss, independently rechecks accepted results and preserves unavailable branches and pagination; enabled MCP and newly tool-equipped Goal Chat use the same read model. Existing native threads retain their tool schema on resume. It starts no work and does not infer overall readiness from a display list. The example's `prepare` still only provisions isolated operator bindings. A Codex binding can now pass an exact model/reasoning effort into an independent resumable Turn Session and expose the same profile through preflight and planning; this remains separate from a native temporary child profile inside the parent execution. Next, feed actual execution/acceptance facts into existing R2 readiness, extend registration/runtime configuration for approved identity provisioning, and qualify original-request return/lead continuation. Unattended wake, full cross-host inbox/queue/steer and Lark parity remain separate requirements; an exact profile parameter does not promote G1/G3. The local Goal conversation now exposes that inventory on demand, with per-binding preflight through the actual Turn dry-run and selected executor/profile. Task diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md index a07ffac2b3..6c46533375 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md @@ -360,7 +360,7 @@ Todo 完成入口分别执行当前 pinned 检查,accepted 返回读 canonical 完成。长期 attached 会话、通用 Agent 创建、动态受治理工作派生、完整 inbox/queue/steer、 认证远端权威与 packaged frontend/Lark 配套仍归 R2/R3/R4/R6;不晋升已有 Goal。 -有 shell 能力的原 coordinator 可通过 `delegation list/operations/start/read/wait/resume` 调用已有执行 owner,无需替换会话。`operations` 从自身持久记录找回上下文丢失前的工作,重新核验 accepted,保留不可用分支与分页;已启用的 MCP 和新挂载工具的 Goal Chat 共用该读模型;已有原生线程恢复时保留原工具 schema。读取不启动工作,也不把展示列表当整体 readiness。合成示例 `prepare` 仍只准备隔离绑定。下一步先将实际执行/验收事实接入已有 R2 readiness,再沿现有注册及 runtime 配置扩展经授权的身份/profile 创建,并验证原请求返回与主力继续推进。无人值守唤醒、完整跨宿主 inbox/queue/steer 和 Lark 等价仍分别验收,不因新增恢复入口晋升 G1/G3。 +有 shell 能力的原 coordinator 可通过 `delegation list/operations/start/read/wait/resume` 调用已有执行 owner,无需替换会话。`operations` 从自身持久记录找回上下文丢失前的工作,重新核验 accepted,保留不可用分支与分页;已启用的 MCP 和新挂载工具的 Goal Chat 共用该读模型;已有原生线程恢复时保留原工具 schema。读取不启动工作,也不把展示列表当整体 readiness。合成示例 `prepare` 仍只准备隔离绑定。Codex binding 现可将精确 model/reasoning effort 送入独立且可续接的 Turn Session,并由同一 preflight/规划投影读回;这与父执行内部的原生临时 child profile 分开。下一步先将实际执行/验收事实接入已有 R2 readiness,再沿现有注册及 runtime 配置扩展经授权的身份创建,并验证原请求返回与主力继续推进。无人值守唤醒、完整跨宿主 inbox/queue/steer 和 Lark 等价仍分别验收,不因 profile 参数接通而晋升 G1/G3。 本地 Goal 对话现可按需读取该持久目录,并按绑定调用实际 Turn dry-run 与所选 executor/profile 检查启动条件。任务准入、当前 pinned 验收绑定、运行时可用性分开 diff --git a/docs/reference/local-delegation.md b/docs/reference/local-delegation.md index a8d160bc94..456354ffa3 100644 --- a/docs/reference/local-delegation.md +++ b/docs/reference/local-delegation.md @@ -36,6 +36,44 @@ select `generic-cli`, `fresh`, and the optional adapter's `--config` invocation. Profiles, executables, workspace isolation and credential custody remain the operator's responsibility. No model tool accepts those values. +A Codex binding launches an independent, resumable Codex Agent Session through +the same governed Turn path. Pin both fields when the worker must use an exact +profile: + +```json +{ + "id": "strong-independent-review", + "agent_id": "managed-reviewer", + "todo_id": "todo_review", + "requesters": ["lead"], + "workspace": "/absolute/reviewer-worktree", + "host_args": [ + "--host", "codex-cli", + "--codex-model", "gpt-5.6-sol", + "--codex-reasoning-effort", "xhigh", + "--codex-sandbox", "workspace-write" + ], + "timeout_seconds": 300, + "output_refs": ["output.json"] +} +``` + +This is not a native `multi_subagent` child. Native children remain temporary +workers inside one parent execution and use the Goal's child model preference. +The binding above has its own Agent identity, Todo, workspace, Codex Session and +durable delegation operation. `delegation inspect` and the planning projection +show `gpt-5.6-sol@xhigh`; start and resume pass both fields to that same Session. +The requester still needs the exact binding grant, and a profile is not an +acceptance or result-return receipt. + +中文:Codex binding 通过同一条受治理 Turn 链启动独立、可续接的 Codex Agent +Session。需要精确执行配置时同时固定 `--codex-model` 与 +`--codex-reasoning-effort`。它不是 `multi_subagent` 的原生临时 child:后者仍在 +单个父执行内部使用 Goal 的 child model 偏好;前者拥有独立 Agent 身份、Todo、 +workspace、Codex Session 与持久 delegation operation。`delegation inspect` 和 +规划投影会读回例如 `gpt-5.6-sol@xhigh`,start/resume 也把同一配置送入原 Session。 +这不扩大 requester grant,也不把 profile 冒充验收或结果返回回执。 + ## Use an existing Agent conversation through its shell An attached Codex or other shell-capable Agent can use the same execution diff --git a/loopx/cli_commands/turn.py b/loopx/cli_commands/turn.py index d2d96451f6..c3e8802e26 100644 --- a/loopx/cli_commands/turn.py +++ b/loopx/cli_commands/turn.py @@ -195,10 +195,30 @@ def handle_turn_command( # have to name the same credential. environ=operator_environ, dsh_runner_configured=bool(getattr(args, "dsh_runner", None)), - provider=getattr(args, "dsh_provider", None), - model=getattr(args, "dsh_model", None), - reasoning_effort=getattr(args, "dsh_reasoning_effort", None), - max_tokens=getattr(args, "dsh_max_tokens", None), + provider=( + getattr(args, "dsh_provider", None) + if args.host == "dsh" + else None + ), + model=( + getattr(args, "dsh_model", None) + if args.host == "dsh" + else getattr(args, "codex_model", None) + if args.host == "codex-cli" + else None + ), + reasoning_effort=( + getattr(args, "dsh_reasoning_effort", None) + if args.host == "dsh" + else getattr(args, "codex_reasoning_effort", None) + if args.host == "codex-cli" + else None + ), + max_tokens=( + getattr(args, "dsh_max_tokens", None) + if args.host == "dsh" + else None + ), ) if ( args.turn_command == "run-once" @@ -974,6 +994,7 @@ def run_built_in_host( codex_bin=args.codex_bin, sandbox=args.codex_sandbox, model=args.codex_model, + reasoning_effort=args.codex_reasoning_effort, timeout_seconds=max(1.0, args.timeout_seconds - 5.0), ) diff --git a/loopx/cli_commands/turn_registration.py b/loopx/cli_commands/turn_registration.py index db6bb8d0a7..b5a541d976 100644 --- a/loopx/cli_commands/turn_registration.py +++ b/loopx/cli_commands/turn_registration.py @@ -9,6 +9,7 @@ MANAGED_TURN_HOST, resolve_default_turn_host, ) +from ..control_plane.turn_driver.execution_profile import REASONING_EFFORTS from ..paths import default_public_scan_root # Explicit host choices stay per-command: planning may name any host the Turn @@ -208,6 +209,15 @@ def register_turn_commands( help="Codex CLI executable used by the built-in codex-cli host.", ) run_once.add_argument("--codex-model") + run_once.add_argument( + "--codex-reasoning-effort", + choices=list(REASONING_EFFORTS), + help=( + "Reasoning effort for the independent Codex CLI Turn. This is an " + "operator-bound independent Agent profile, not the native child-agent " + "model preference." + ), + ) run_once.add_argument( "--codex-sandbox", choices=["read-only", "workspace-write", "danger-full-access"], diff --git a/loopx/control_plane/turn_driver/codex_cli.py b/loopx/control_plane/turn_driver/codex_cli.py index cc19372af0..60802997e9 100644 --- a/loopx/control_plane/turn_driver/codex_cli.py +++ b/loopx/control_plane/turn_driver/codex_cli.py @@ -25,6 +25,7 @@ HOST_RESULT_TEXT_LIMITS, LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION, ) +from .execution_profile import require_supported_reasoning_effort from .host_failure import BuiltInHostError from .transaction import LOOPX_TURN_RESULT_SCHEMA_VERSION, TRANSACTION_PHASES @@ -665,6 +666,7 @@ def _codex_command( output_path: Path, sandbox: str, model: str | None, + reasoning_effort: str | None, session_id: str | None, ) -> list[str]: if session_id: @@ -700,6 +702,13 @@ def _codex_command( ] if model: command.extend(["--model", model]) + if reasoning_effort: + command.extend( + [ + "-c", + f"model_reasoning_effort={json.dumps(reasoning_effort)}", + ] + ) if session_id: command.append(session_id) command.append("-") @@ -714,12 +723,15 @@ def run_codex_cli_host( codex_bin: str = "codex", sandbox: str = "read-only", model: str | None = None, + reasoning_effort: str | None = None, timeout_seconds: float = 115.0, ) -> dict[str, Any]: if request.get("schema_version") != LOOPX_TURN_HOST_REQUEST_SCHEMA_VERSION: raise ValueError("unsupported LoopX Turn host request schema") if sandbox not in CODEX_CLI_SANDBOXES: raise ValueError(f"Codex CLI sandbox must be one of {CODEX_CLI_SANDBOXES}") + if reasoning_effort is not None: + reasoning_effort = require_supported_reasoning_effort(reasoning_effort) resolved = shutil.which(codex_bin) if os.path.sep not in codex_bin else codex_bin if not resolved or not Path(resolved).exists(): raise ValueError("Codex CLI executable is unavailable") @@ -760,6 +772,7 @@ def run_codex_cli_host( output_path=output_path, sandbox=sandbox, model=model, + reasoning_effort=reasoning_effort, session_id=session_id, ) proc = subprocess.Popen( diff --git a/loopx/control_plane/turn_driver/host_binding.py b/loopx/control_plane/turn_driver/host_binding.py index d67bdd604e..43ac3ce965 100644 --- a/loopx/control_plane/turn_driver/host_binding.py +++ b/loopx/control_plane/turn_driver/host_binding.py @@ -195,10 +195,11 @@ def managed_executor_binding( records that this projection does not probe that executor kind, so it makes no claim rather than an unproven ``True``. - ``execution_profile`` is the managed profile this executor would run -- and, - with it, the provider claims to authenticate. It is ``None`` for every - non-managed executor because neither the profile nor the credential belongs - to an individual or generic host. + ``execution_profile`` is the explicit profile this executor would run. + Managed DSH resolves its complete provider profile. An individual Codex + executor projects only operator-pinned model/effort arguments; absent fields + remain ``host-default`` rather than being guessed. Generic hosts have no + shared profile contract and keep this field ``None``. ``runtime_probe`` states what the ``dsh_runtime_unavailable`` verdict is a claim about, so a reader does not take a process-level answer for a @@ -251,6 +252,11 @@ def managed_executor_binding( "available": runtime_available, }, } + individual_profile: str | None = None + if host == INDIVIDUAL_TURN_HOST and (model or reasoning_effort): + individual_profile = ( + f"{model or 'host-default'}@{reasoning_effort or 'host-default'}" + ) return { "schema_version": MANAGED_EXECUTOR_BINDING_SCHEMA_VERSION, "executor": host, @@ -261,7 +267,7 @@ def managed_executor_binding( ), "credential_env": None, "endpoint_env": None, - "execution_profile": None, + "execution_profile": individual_profile, "operator_credential_bound": False, "available": None, "unavailable_reason": None, @@ -313,9 +319,25 @@ def managed_executor_binding_from_host_args( host, environ=environ, module_probe=module_probe, - provider=turn_host_arg_option(host_args, "--dsh-provider"), - model=turn_host_arg_option(host_args, "--dsh-model"), - reasoning_effort=turn_host_arg_option(host_args, "--dsh-reasoning-effort"), + provider=( + turn_host_arg_option(host_args, "--dsh-provider") + if host == MANAGED_HOST + else None + ), + model=( + turn_host_arg_option(host_args, "--dsh-model") + if host == MANAGED_HOST + else turn_host_arg_option(host_args, "--codex-model") + if host == INDIVIDUAL_TURN_HOST + else None + ), + reasoning_effort=( + turn_host_arg_option(host_args, "--dsh-reasoning-effort") + if host == MANAGED_HOST + else turn_host_arg_option(host_args, "--codex-reasoning-effort") + if host == INDIVIDUAL_TURN_HOST + else None + ), ) diff --git a/tests/test_delegation_preflight.py b/tests/test_delegation_preflight.py index a136dcea57..6d77fdde35 100644 --- a/tests/test_delegation_preflight.py +++ b/tests/test_delegation_preflight.py @@ -168,6 +168,33 @@ def test_selected_dsh_profile_is_not_replaced_by_the_default(service): assert not (root / "host-started").exists() +def test_selected_codex_managed_agent_profile_is_projected_exactly(service): + root, runner = service + config = json.loads(runner.config.read_text()) + config["bindings"][0]["host_args"] = [ + "--host", + "codex-cli", + "--codex-model", + "gpt-5.6-sol", + "--codex-reasoning-effort", + "xhigh", + ] + runner.config.write_text(json.dumps(config)) + + status, result = cli(runner, "inspect", "--binding-id", "analysis") + + assert status == 0, result + assert result["executor"] == { + "host": "codex-cli", + "available": None, + "reason": None, + "profile": "gpt-5.6-sol@xhigh", + } + assert result["state"] == "runtime_unverified" + assert not any(result["effects"].values()) + assert not (root / "host-started").exists() + + def test_preflight_does_not_call_an_invalidated_acceptance_ready(service): root, runner = service from loopx.agent_registry import load_goal_from_registry diff --git a/tests/test_loopx_turn_codex_cli.py b/tests/test_loopx_turn_codex_cli.py index 5fed62efc0..2da8292868 100644 --- a/tests/test_loopx_turn_codex_cli.py +++ b/tests/test_loopx_turn_codex_cli.py @@ -368,6 +368,8 @@ def test_codex_cli_host_starts_then_resumes_opaque_session( project=project, codex_bin=str(executable), sandbox=sandbox, + model="gpt-5.6-sol", + reasoning_effort="xhigh", timeout_seconds=5, ) with pytest.raises(RuntimeError, match="binding changed after planning"): @@ -388,6 +390,8 @@ def test_codex_cli_host_starts_then_resumes_opaque_session( project=project, codex_bin=str(executable), sandbox=sandbox, + model="gpt-5.6-sol", + reasoning_effort="xhigh", timeout_seconds=5, ) @@ -399,6 +403,14 @@ def test_codex_cli_host_starts_then_resumes_opaque_session( assert "resume" not in argv_rows[0] assert "resume" in argv_rows[1] assert "session-fixture-0001" in argv_rows[1] + for argv in argv_rows: + assert argv[argv.index("--model") + 1] == "gpt-5.6-sol" + config_values = [ + argv[index + 1] + for index, value in enumerate(argv) + if value == "-c" + ] + assert 'model_reasoning_effort="xhigh"' in config_values resume_argv = argv_rows[1] assert resume_argv[resume_argv.index("-c") + 1] == ( f'sandbox_mode="{sandbox}"' @@ -435,6 +447,28 @@ def test_codex_cli_host_starts_then_resumes_opaque_session( assert "private_material" not in persisted +def test_codex_cli_host_rejects_unknown_reasoning_effort_before_launch( + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + executable, log_path = _fake_codex(tmp_path) + monkeypatch.setenv("FAKE_CODEX_LOG", str(log_path)) + project = tmp_path / "project" + project.mkdir() + + with pytest.raises(ValueError, match="unsupported reasoning effort"): + run_codex_cli_host( + _request(), + runtime_root=tmp_path / "runtime", + project=project, + codex_bin=str(executable), + reasoning_effort="turbo", + timeout_seconds=5, + ) + + assert not log_path.exists() + + def test_codex_cli_host_fresh_iteration_ignores_stored_session( tmp_path: Path, monkeypatch: pytest.MonkeyPatch, diff --git a/tests/test_turn_default_host_binding.py b/tests/test_turn_default_host_binding.py index 945e2410c9..e3893c73a5 100644 --- a/tests/test_turn_default_host_binding.py +++ b/tests/test_turn_default_host_binding.py @@ -128,6 +128,23 @@ def test_explicit_host_flag_wins_over_the_default(monkeypatch): assert args.host == "generic-cli" +def test_codex_managed_agent_profile_is_explicit_cli_configuration(): + args = build_parser().parse_args( + [ + *_turn_argv("run-once"), + "--host", + "codex-cli", + "--codex-model", + "gpt-5.6-sol", + "--codex-reasoning-effort", + "xhigh", + ] + ) + + assert args.codex_model == "gpt-5.6-sol" + assert args.codex_reasoning_effort == "xhigh" + + @pytest.mark.parametrize("command", ["plan", "run-once"]) def test_default_execution_mode_follows_the_selected_host(command, monkeypatch): for name in ("DEEPSEEK_API_KEY", TURN_HOST_ENV_VAR): diff --git a/tests/test_turn_managed_executor_binding.py b/tests/test_turn_managed_executor_binding.py index cf3d2a602b..7ab7261dc6 100644 --- a/tests/test_turn_managed_executor_binding.py +++ b/tests/test_turn_managed_executor_binding.py @@ -87,6 +87,27 @@ def test_trusted_host_args_project_the_same_last_explicit_profile_without_raw_ar assert turn_host_arg_option(["--dsh-model", "fixture"], "--host") is None +def test_trusted_codex_binding_projects_its_independent_agent_profile(): + binding = managed_executor_binding_from_host_args( + [ + "--host", + "codex-cli", + "--codex-model", + "gpt-5.6-sol", + "--codex-reasoning-effort", + "xhigh", + ], + environ={"DEEPSEEK_API_KEY": "unrelated-managed-credential"}, + ) + + assert binding["executor"] == "codex-cli" + assert binding["executor_kind"] == EXECUTOR_KIND_INDIVIDUAL + assert binding["execution_profile"] == "gpt-5.6-sol@xhigh" + assert binding["available"] is None + assert binding["operator_credential_bound"] is False + assert "unrelated-managed-credential" not in json.dumps(binding) + + def test_managed_executor_reports_the_operator_credential_and_endpoint(): binding = managed_executor_binding( "dsh", @@ -401,7 +422,7 @@ def test_a_deviating_provider_is_named_in_the_profile_line(): assert binding["execution_profile"] == "fixture-provider/deepseek-v4-flash@high" -def test_non_managed_hosts_carry_no_execution_profile(): +def test_unconfigured_non_managed_hosts_carry_no_execution_profile(): for host in ("codex-cli", "generic-cli"): binding = managed_executor_binding( host, environ={"DEEPSEEK_API_KEY": "sk-operator"}, module_probe=_RUNTIME