diff --git a/benchmark/deepswe-gptxhigh-v1/README.md b/benchmark/deepswe-gptxhigh-v1/README.md new file mode 100644 index 0000000000..7956579f80 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/README.md @@ -0,0 +1,94 @@ +# DeepSWE Five-Arm Benchmark Harness (v1) + +How we evaluate five agent configurations ("arms") on the DeepSWE task set +(113 SWE tasks). Runner/methodology code only — no API/gateway config, no trajectories. +Run against LoopX revision `2cef51d` (the evaluated revision, not the PR base). +This versioned snapshot lives at `benchmark/deepswe-gptxhigh-v1/`. + +## Reproduction prerequisites + +This is an archived source and methodology snapshot, not a standalone runnable +bundle. The launch scripts retain the original experiment workspace layout. +Running them requires the external Pier environment, `run.sh`, task manifests +(`remaining59.txt`, `goal30_subset.py`, `hard24_subset.py`, and +`remaining4_subset.py`), a configured Python environment, and separately supplied +model gateway configuration. Adapt the environment paths and set `MR_LOOPX_ROOT` +to a checkout of the evaluated revision before running. These prerequisites and +raw verifier artifacts are not included here; the reported results below are +preserved from the v1 summary and have not been independently reproduced by this +publication change. + +Set `MR_PYTHON` to the interpreter containing the Pier dependencies (defaults to +`python3`). Both launchers require `MR_MODELONLY_COMPOSE` to name an existing +external Docker Compose overlay defining the model-only network. Its gateway +must be reachable from the task containers; loopback placeholders are not a +portable network configuration. Supply the overlay and gateway configuration +for your environment before launching. + +The remaining-59 launcher accepts `MR_TASK_LIST` (defaults to `remaining59.txt`) +and requires 59 unique task directory names, one per line. The 54-task launcher +loads its external frozen subset modules and validates any explicit task subset. +Each launcher passes a snapshot of its actual selected task ids to preflight; +the admission receipt hashes those tasks' `upstream/tasks//task.toml` files. +The evaluated LoopX revision is fixed and cannot be overridden by environment. +All LoopX admissions finish before any arm starts; launcher failure is nonzero +if admission or any arm fails. + +Publication fixes add the missing plain runner and harden launch/admission and +profile initialization. Unsupported legacy `MR_CODEX_ARM=loopx` and Claude +adapters are excluded from this five-arm package. These fixes do not constitute +a rerun or revalidation of the historical results. + +Retry decisions use structured provider status/code fields; unknown prose-only +errors now fail instead of being classified by substring. A terminal Todo with +invalid delivery stops the Codex CLI runner with failure. Offline regression +coverage can be run with +`python3 -m pytest -q benchmark/deepswe-gptxhigh-v1/tests/` from the repository root. + +## The five arms +| Arm | Transport | Goal / LoopX | Continuation | +|---|---|---|---| +| `plain` | codex app-server | no Goal, no LoopX | single pass | +| `goal` | codex app-server | native Codex Goal | app-server continues while Goal active | +| `heartbeat` | `codex exec` (fresh, then `resume`) | LoopX Goal/Todo | recurring supervisor wakes | +| `codex-cli` | `codex exec` (CLI) | LoopX control plane | external `loopx turn run-once --host codex-cli`, multi-segment | +| `ssh-goal` | codex app-server | LoopX + native Goal (official full path) | same thread/Goal; LoopX clears blocked + restarts turn | + +- Arm dispatch: `pier_cn.py` (`MR_CODEX_ARM=plain|goal|loopx-native|loopx-native-codex-cli|loopx-native-heartbeat`) +- Arm classes: `goal_codex.py` (`PlainAppServerCodex`, `GoalCodex`) and + `loopx_native_codex.py` (the three LoopX variants) +- Runners: `loopx_wen_native_runner.py` (ssh-goal), `loopx_codex_cli_runner.py` (codex-cli), + `loopx_heartbeat_supervisor.py` (heartbeat) +- Delivery gate: `workspace_delivery.py` (recover agent work from linked git worktrees into + `/app` so the collected patch is non-empty) +- Admission: `preflight_loopx_rerun.py` (pins LoopX revision, delivery self-test) + +## Validity (strict) +`exception_info == null`, independent `verifier/reward.json` present & consistent, task +checksum matches, and a **non-empty committed patch** exists. Goal/Todo state is lifecycle +evidence only; the independent verifier is the sole correctness authority. `partial > 0` +alone does NOT count as valid delivery. + +## Results — v1 (113 tasks, per-task best valid) + +Historical reported summary only. The current +[SWE Marathon publication](../swe-marathon/README.md) withdraws SSH Goal and +Codex CLI data and conclusions pending revalidation. Their rows are retained +below as part of this v1 archive, not as currently validated results or ranking +claims. + +| Rank | Arm | Solved | Solve rate | Partial | F2P | P2P | +|---|---|---:|---:|---:|---:|---:| +| 1 | heartbeat | 70/113 | 61.9% | 0.9739 | 0.886 | 0.993 | +| 2 | codex-cli | 66/113 | 58.4% | 0.9546 | 0.884 | 0.996 | +| 3 | goal | 60/113 | 53.1% | 0.9620 | 0.868 | 0.997 | +| 4 | ssh-goal | 58/113 | 51.3% | 0.9654 | 0.883 | 0.997 | +| 5 | plain | 54/113 | 47.8% | 0.9206 | 0.745 | 0.997 | + +P2P (regression) ≈ 1.0 for all arms; spread is driven by F2P and solve rate. +The table records the original v1 comparison; conclusions involving the +withdrawn arms require revalidation. + +> Model & gateway endpoints are configured via `MR_*` env vars (not included). +> Internal hosts/paths replaced with placeholders (`127.0.0.1`, ``, ``). +> v2 (latest LoopX main) evaluation is in progress and will be published separately. diff --git a/benchmark/deepswe-gptxhigh-v1/codex_nosandbox_wrapper.py b/benchmark/deepswe-gptxhigh-v1/codex_nosandbox_wrapper.py new file mode 100755 index 0000000000..dc044d9104 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/codex_nosandbox_wrapper.py @@ -0,0 +1,131 @@ +#!/usr/bin/env python3 +"""A `codex` stand-in that drops LoopX's sandbox flags before running the real one. + +LoopX's codex-cli host is the path that works: driven through it, a Turn loop +ran four times on one task, committed a 22 KB patch and scored f2p 31/35. Its +one problem is the sandbox — it always passes `--sandbox ` (or +`-c sandbox_mode=...` when resuming), only permits read-only and +workspace-write, and both need bubblewrap, which needs unprivileged user +namespaces these containers do not have: + + bwrap: No permissions to create a new namespace + +Switching to `--host generic-cli` avoided that but bought a worse problem: the +generic host carries its own scheduler contract, and eleven of sixteen turns +died at "LoopX Turn route is not host executable" before any model work, with +no route recorded to explain why. + +So keep the working host and fix the flag instead. This sits earlier on PATH +than the real codex, strips the sandbox arguments, and substitutes the same +`--dangerously-bypass-approvals-and-sandbox` the other two arms already use — +which is also what keeps the three arms identical in permissions. LoopX's +contracts are untouched: it still believes it is driving codex-cli, because it +is. + +Set MR_REAL_CODEX to the real binary; defaults to /usr/local/bin/codex. +MR_LOOPX_CODEX_LOG names the log file; defaults to /tmp/loopx-goal/codex-wrapper.log. +""" + +from __future__ import annotations + +import os +import shutil +import subprocess +import sys +from pathlib import Path + +# Resolve the real binary rather than assuming /usr/local/bin/codex. Codex is +# installed into the image through nvm, so it lives under the Node version's +# bin directory and the hardcoded path does not exist — which made the wrapper +# die before it ever reached Codex, and LoopX report the indistinguishable +# `codex_cli_exit_nonzero`. The wrapper is invoked by absolute path through +# --codex-bin and is not itself on PATH, so a PATH lookup finds the real one. +REAL = ( + os.environ.get("MR_REAL_CODEX") + or shutil.which("codex") + or "/usr/local/bin/codex" +) +BYPASS = "--dangerously-bypass-approvals-and-sandbox" +REASONING_EFFORT = os.environ.get("MR_CODEX_REASONING_EFFORT", "").strip() +LOG = Path( + os.environ.get("MR_LOOPX_CODEX_LOG", "/tmp/loopx-goal/codex-wrapper.log") +) + + +def rewrite(argv: list[str]) -> list[str]: + out: list[str] = [] + skip_next = False + for i, arg in enumerate(argv): + if skip_next: + skip_next = False + continue + # `--sandbox ` — new-session form. + if arg == "--sandbox": + skip_next = True + continue + if arg.startswith("--sandbox="): + continue + # `-c sandbox_mode="..."` — resume form. The value is a separate argv + # item after -c, so both have to go, and only when it is that key: -c + # carries every other config override too. + if arg == "-c" and i + 1 < len(argv) and argv[i + 1].startswith("sandbox_mode="): + skip_next = True + continue + out.append(arg) + + # Insert the bypass right after the subcommand so it lands before `--`, + # which codex treats as the end of flags. + if out and out[0] == "exec": + out.insert(1, BYPASS) + if REASONING_EFFORT and not any( + value.startswith("model_reasoning_effort=") for value in out + ): + out[2:2] = ["-c", f"model_reasoning_effort={REASONING_EFFORT}"] + else: + out.insert(0, BYPASS) + return out + + +def _log(text: str) -> None: + try: + LOG.parent.mkdir(parents=True, exist_ok=True) + with LOG.open("a", encoding="utf-8") as handle: + handle.write(text.rstrip("\n") + "\n") + except OSError: + pass + + +def main() -> int: + argv = rewrite(sys.argv[1:]) + # Run the real codex as a child rather than execv'ing it, so its stderr can + # be recorded. LoopX reports a failed Turn as `codex_cli_exit_nonzero` and + # keeps neither the exit code's cause nor any output, and the container is + # gone by the time anyone looks — so an execv here means the only evidence + # of why Codex refused is destroyed at the moment it is produced. + # + # stdout stays inherited and untouched: LoopX parses Codex's `--json` + # stream off it, so anything written there would corrupt the Turn. + _log(f"--- argv in : {sys.argv[1:]}") + _log(f"--- argv out: {argv}") + _log(f"--- real : {REAL} (exists={os.path.exists(REAL)})") + try: + completed = subprocess.run( # noqa: S603 + [REAL, *argv], stderr=subprocess.PIPE, check=False + ) + except OSError as exc: + # Without this the wrapper's own failure to start Codex is reported by + # LoopX as `codex_cli_exit_nonzero`, which reads as "the model refused" + # rather than "the binary is not there". + _log(f"--- launch failed: {type(exc).__name__}: {exc}") + sys.stderr.write(f"codex wrapper could not launch {REAL}: {exc}\n") + return 127 + stderr = completed.stderr.decode("utf-8", "replace") if completed.stderr else "" + _log(f"--- exit {completed.returncode}") + if stderr: + _log(stderr) + sys.stderr.write(stderr) + return completed.returncode + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/goal_codex.py b/benchmark/deepswe-gptxhigh-v1/goal_codex.py new file mode 100644 index 0000000000..a8f7cabf46 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/goal_codex.py @@ -0,0 +1,451 @@ +"""Codex driven through its native Goal API, as a Pier agent. + +Pier's stock Codex agent runs `codex exec` once. That is enough to *create* a +Goal — with `features.goals` on, the model will call `create_goal` when asked — +but not to exercise one: the process exits after the first turn, so the +automatic continuation loop that is the entire point of Goal mode never runs. +Measured that way, Goal mode looks like a no-op, and the experiment would +conclude the wrong thing for a purely mechanical reason. + +So this subclass keeps everything Pier does to stand Codex up — npm install +through the CN mirror, CODEX_HOME, auth.json, config.toml, skills, MCP, session +capture — and swaps only the final invocation for the app-server transaction: + + initialize(experimentalApi=true) -> thread/start -> thread/goal/set(active) + -> turn/start -> observe continuation turns while the Goal stays active + +That transaction is not reimplemented here. `native_codex_goal.py` from LoopX +already owns it, is stdlib-only, and is the same code path the LoopX arm will +use later — sharing it is what keeps the two arms differing in LoopX alone +rather than in how each one talks to Codex. It is copied into the container at +run time rather than baked into the image so that the two arms cannot drift. + +The swap is done by intercepting `exec_as_agent` rather than by reimplementing +`run()`. Pier's `run()` is one long method whose setup (auth resolution, +ownership fixes, config blocks) would have to be duplicated and kept in step +with upstream; intercepting the one command that matters leaves that setup +untouched. If Pier ever changes how it invokes Codex, the marker below stops +matching and this fails loudly instead of silently reverting to plain +`codex exec` — which would look like a successful Goal run with no Goal in it. + +Usage: + + MR_AGENT=goal_codex:GoalCodex MR_MODEL=openai/gpt-5.5 ./run.sh --all -i + +Environment: + + MR_GOAL_PREFLIGHT=1 prove Goal attachment and stop before any model + turn — costs nothing, use it first on a new box + MR_GOAL_TIMEOUT_SEC ceiling for the continuation loop (default 3600, + under the 5400 s task budget so the loop stops + itself instead of being killed mid-turn) + MR_GOAL_TOKEN_BUDGET optional Goal token budget + MR_NATIVE_GOAL_MODULE path to LoopX's native_codex_goal.py on the host +""" + +from __future__ import annotations + +import os +import shlex +import tempfile +from pathlib import Path + +from pier.agents.installed.codex import Codex +from pier.models.trial.paths import EnvironmentPaths + +# Pier builds exactly one command containing this; see +# pier/agents/installed/codex.py, Codex.run(). +_CODEX_EXEC_MARKER = "codex exec " + +_REMOTE_DIR = "/tmp/loopx-goal" +_DEFAULT_LOOPX_ROOT = str( + Path(__file__).resolve().parents[1] / "wen" / "loopx" +) +_DEFAULT_MODULE_RELATIVE = ( + "loopx/capabilities/benchmark_toolkit/native_codex_goal.py" +) + +# The Goal objective is fixed rather than derived from the task text. A Goal is +# meant to state the durable intent that survives across continuation turns, +# while the task file already carries the specifics; restating the task as the +# objective gave the model two copies of the same thing and nothing to hold on +# to between turns. DeepSWE grades a committed patch, so committing belongs in +# the objective — a run that solves the task and never commits scores zero. +_OBJECTIVE = ( + "Complete the software engineering task described in the task file. " + "Work in the repository, keep existing behaviour intact, verify the change " + "against the repository's own tests, and commit the finished work to a new " + "branch off main. The goal is complete only once the change is committed." +) + +# A Goal only continues while it is still active, so an objective that one turn +# can satisfy never exercises the continuation loop — and the objective above +# says outright that committing completes it. Across 53 runs the continuation +# count was zero every time, which makes the measured "Goal API has no effect" +# a statement about an objective that never needed the API, not about the API. +# +# This variant withholds completion until work that cannot plausibly finish in +# one turn is done: pass, then re-derive from the tests, then hunt regressions, +# then edge cases. Whether that actually keeps the Goal active is the thing +# being tested — if the continuation count is still zero, single-turn +# termination is Codex's behaviour here rather than an artefact of the wording. +_OBJECTIVE_STAGED = ( + "Complete the software engineering task described in the task file, in " + "stages, and do not consider the goal complete until every stage is done.\n" + "Stage 1: make the target behaviour work and commit it.\n" + "Stage 2: re-read the task description and check your implementation " + "against every requirement it states, including ones you did not address " + "in stage 1. Fix what is missing and commit.\n" + "Stage 3: look for behaviour you may have broken elsewhere in the " + "repository, run the wider test suite, and fix any regression you find.\n" + "Stage 4: consider edge cases the tests may not cover — empty inputs, " + "concurrent use, error paths — and handle the ones the task implies.\n" + "The goal is complete only after stage 4." +) + + +def _objective() -> str: + return _OBJECTIVE_STAGED if os.environ.get("MR_GOAL_OBJECTIVE") == "staged" else _OBJECTIVE + + +# All three arms keep Pier's own agent name. Overriding name() per arm looked +# tidy but fed straight into AgentInstallSpec.fingerprint(), whose first input is +# agent_name — so each arm produced a different PIER_AGENT_INSTALL_FINGERPRINT, +# invalidated the Docker layer cache, and rebuilt `nvm install 22` plus the npm +# install of Codex for every task in every arm: 162 builds where 54 would do. +# It also removed the only fallback for a network outage, since a cached layer +# needs no proxy. The arm is selected by pier_cn.py rebinding AgentName.CODEX, +# which needs no distinct name. + +_WEB_SEARCH_OFF = 'printf "\\nweb_search = \\"disabled\\"\\n" >> "$CODEX_HOME/config.toml"' + + +class PlainCodex(Codex): + """The control arm: same everything, no Goal attached. + + Exists so that Goal vs no-Goal differs in the Goal API and nothing else. + Two things have to be carried over from GoalCodex or the comparison measures + the wrong difference: + + * ``web_search`` off. Stock Codex leaves it at its default, and it is a + hosted tool the container's egress allowlist cannot block, so one arm + could look answers up. + * the objective text. A Goal cannot exist without an objective, so the Goal + arm is necessarily prompted with those three sentences. Withholding them + here would fold "the effect of that wording" into the measured difference. + Appending them leaves the API as the only variable. + + What remains different is intrinsic: `codex exec` runs one turn and exits, + while the app-server keeps serving continuations while the Goal stays + active. That *is* the treatment. + """ + + async def run(self, instruction, environment, context): # type: ignore[override] + return await super().run( + f"{instruction}\n\n{_objective()}", environment, context + ) + + async def exec_as_agent(self, environment, command: str = "", env=None, **kwargs): # type: ignore[override] + if _CODEX_EXEC_MARKER in command: + await super().exec_as_agent(environment, command=_WEB_SEARCH_OFF, env=env) + return await super().exec_as_agent( + environment, command=command, env=env, **kwargs + ) +class GoalCodex(Codex): + """Codex with its native Goal loop actually running.""" + + def __init__(self, *args, **kwargs): + super().__init__(*args, **kwargs) + self._goal_instruction: str | None = None + self._goal_swapped = False + + async def run(self, instruction, environment, context): # type: ignore[override] + self._goal_instruction = instruction + self._goal_swapped = False + try: + return await super().run(instruction, environment, context) + finally: + if not self._goal_swapped: + raise RuntimeError( + "GoalCodex never intercepted a `codex exec` command — Pier's " + "Codex.run() no longer matches the expected shape, and this " + "run would have been plain Codex with no Goal attached." + ) + + async def exec_as_agent(self, environment, command: str = "", env=None, **kwargs): # type: ignore[override] + if _CODEX_EXEC_MARKER not in command: + return await super().exec_as_agent( + environment, command=command, env=env, **kwargs + ) + + self._goal_swapped = True + instruction = self._goal_instruction or "" + model = self._command_model_name or (self.model_name or "").split("/")[-1] + + loopx_root = Path( + os.environ.get("MR_LOOPX_ROOT", _DEFAULT_LOOPX_ROOT) + ).expanduser() + module_src = Path( + os.environ.get( + "MR_NATIVE_GOAL_MODULE", + str(loopx_root / _DEFAULT_MODULE_RELATIVE), + ) + ) + if not module_src.is_file(): + raise FileNotFoundError( + f"native_codex_goal.py not found at {module_src}; set " + "MR_NATIVE_GOAL_MODULE to LoopX's copy" + ) + runner_source = getattr(self, "_goal_runner_source", None) + + await super().exec_as_agent( + environment, command=f"mkdir -p {shlex.quote(_REMOTE_DIR)}", env=env + ) + + # Upload rather than heredoc: the instruction is arbitrary user text and + # routinely contains quotes, backticks and $ — shell-quoting it into a + # container command is how the objective silently loses characters. + with tempfile.TemporaryDirectory() as tmp: + objective_path = Path(tmp) / "objective.txt" + task_path = Path(tmp) / "task.txt" + objective_path.write_text(_objective(), encoding="utf-8") + task_path.write_text(instruction, encoding="utf-8") + + uploads = [ + (module_src, f"{_REMOTE_DIR}/native_codex_goal.py"), + (objective_path, f"{_REMOTE_DIR}/objective.txt"), + (task_path, f"{_REMOTE_DIR}/task.txt"), + ] + if runner_source is not None: + uploads.append((Path(runner_source), f"{_REMOTE_DIR}/run_goal.py")) + for local, remote in uploads: + await environment.upload_file(str(local), remote) + + # upload_file writes as root; the agent user has to be able to read them. + if environment.default_user is not None: + await self.exec_as_root( + environment, + command=f"chown -R {environment.default_user} {shlex.quote(_REMOTE_DIR)}", + ) + + # Codex ships a `web_search` tool. It is a hosted tool — the provider + # runs the search, so it does not traverse the container's network and + # the egress allowlist that blocks apt, pip and the open internet does + # not block it. Left at its default, one arm of this experiment could + # look things up while the 113-task mini-swe-agent baseline could not, + # and no amount of network isolation would show it. Disabling it in + # config.toml covers both `codex exec` and `codex app-server`; the key + # and its values come from the binary's own override message ("live", + # "cached", "disabled"). The no-Goal arm needs the same line for the + # comparison to hold. + await super().exec_as_agent( + environment, + command=_WEB_SEARCH_OFF, + env=env, + ) + + if runner_source is None: + runner = _RUNNER.format(remote_dir=_REMOTE_DIR) + await super().exec_as_agent( + environment, + command=f"cat > {shlex.quote(_REMOTE_DIR)}/run_goal.py <<'LOOPX_EOF'\n{runner}\nLOOPX_EOF", + env=env, + ) + + timeout = os.environ.get("MR_GOAL_TIMEOUT_SEC", "3600") + budget = os.environ.get("MR_GOAL_TOKEN_BUDGET", "") + effort = os.environ.get("MR_REASONING_EFFORT", "").strip() + preflight = os.environ.get("MR_GOAL_PREFLIGHT", "") not in ("", "0") + + args = [ + "python3", + f"{_REMOTE_DIR}/run_goal.py", + "--cwd", + "__PWD__", + "--objective-file", + # The LoopX arm renders its objective with LoopX's own CLI and + # points here; a subclass cannot substitute it by rewriting the + # command, because this method calls `super().exec_as_agent` rather + # than `self.`, so an override never sees the built command at all. + os.environ.get("MR_GOAL_OBJECTIVE_FILE", f"{_REMOTE_DIR}/objective.txt"), + "--task-file", + f"{_REMOTE_DIR}/task.txt", + "--codex-bin", + "codex", + "--model", + model, + "--goal-timeout-seconds", + timeout, + # NativeGoalConfig defaults to sandbox="workspace-write", which on + # Linux is enforced with landlock/seccomp and cannot be set up + # inside these task containers. Codex then declines to run + # commands: the first attempt produced 431 assistant-message deltas, + # zero command_execution events, zero file_change events and an + # empty model.patch, which the verifier scored 0/24 — a harness + # failure that reads exactly like the model failing the task. + # `codex exec` avoids it with --dangerously-bypass-approvals-and- + # sandbox; this is the app-server equivalent, so both arms execute + # under the same permissions. + "--sandbox", + os.environ.get("MR_GOAL_SANDBOX", "danger-full-access"), + ] + if effort: + args += ["--effort", effort] + if budget: + args += ["--token-budget", budget] + # Empty on this arm. The LoopX arm sets it to the skill ids its + # installed profile materialized, which arms `skills/list` as a + # precondition of thread creation: without it a run in which Codex never + # discovered LoopX still finishes and still scores, and is then filed as + # a LoopX result. One such run has already happened. + skills = os.environ.get("MR_GOAL_REQUIRED_SKILL_IDS", "").strip() + if skills: + args += ["--required-skill-ids", skills] + runner_env = getattr(self, "_goal_runner_env", {}) + if runner_env: + args += [ + "--loopx-cli", runner_env["LOOPX_CLI"], + "--registry", runner_env["LOOPX_REGISTRY"], + "--runtime-root", runner_env["LOOPX_RUNTIME_ROOT"], + "--goal-id", runner_env["LOOPX_GOAL_ID"], + "--agent-id", runner_env["LOOPX_AGENT_ID"], + ] + if runner_env.get("LOOPX_MODE"): + args += ["--mode", runner_env["LOOPX_MODE"]] + if runner_env.get("LOOPX_RECOVER_BLOCKED") == "1": + args.append("--recover-blocked") + if runner_env.get("LOOPX_MODE") in {"codex-cli", "heartbeat"}: + args += [ + "--segment-timeout-seconds", + os.environ.get("MR_HEARTBEAT_SEGMENT_TIMEOUT_SEC", "7200"), + ] + if runner_env.get("LOOPX_MODE") == "ssh-goal": + args += [ + "--turn-idle-timeout-seconds", + os.environ.get("MR_LOOPX_TURN_IDLE_TIMEOUT_SEC", "7200"), + ] + if preflight: + args.append("--preflight-only") + + # No task declares a working directory, so `codex exec` would have run + # in whatever WORKDIR the image sets — different per repository. Resolve + # it in the container instead of guessing. Both spellings are replaced + # because shlex.join only quotes a token that needs it, and this one + # (letters and underscores) comes back bare: matching only the quoted + # form silently leaves the placeholder in the command, which surfaces as + # FileNotFoundError('__PWD__') from inside the runner. + rendered = shlex.join(args) + rendered = rendered.replace("'__PWD__'", '"$(pwd)"').replace( + "__PWD__", '"$(pwd)"' + ) + + output = (EnvironmentPaths.agent_dir / "codex-goal.json").as_posix() + # The LoopX arm points CODEX_HOME at its installed profile so app-server + # discovers the LoopX skills; this arm leaves it as Pier set it up. + codex_home = os.environ.get("MR_GOAL_CODEX_HOME", "").strip() + prefix = f"export CODEX_HOME={shlex.quote(codex_home)}; " if codex_home else "" + for key, value in runner_env.items(): + prefix += f"export {key}={shlex.quote(str(value))}; " + # The profile's `loopx` launcher was built on the host and records the + # host's own Python path, which does not exist in the container. This + # was fixed for the one-shot bootstrap script by exporting the variable + # in that command's own shell -- but bootstrap and this long-lived + # app-server process are separate `docker exec` invocations, each with + # its own shell, so the export did not carry over. Any `loopx` command + # Codex itself runs *during* a turn -- todo claim, quota should-run, + # heartbeat-prompt -- inherits app-server's process environment, not + # bootstrap's, and failed with the same "configured Python executable + # not found" every time until this export is repeated here. A session + # transcript showed exactly that: two failed `loopx` calls mid-turn. + if os.environ.get("MR_GOAL_ARM_LOOPX_PYTHON", "") not in ("", "0"): + prefix += 'export LOOPX_PYTHON="$(command -v python3)"; ' + return await super().exec_as_agent( + environment, + command=( + "if [ -s ~/.nvm/nvm.sh ]; then . ~/.nvm/nvm.sh; fi; " + f"{prefix}{rendered} 2>&1 dict[str, Any]: + for line in reversed(text.splitlines()): + try: + value = json.loads(line) + except ValueError: + continue + if isinstance(value, dict): + return value + try: + value = json.loads(text) + except ValueError: + return {} + return value if isinstance(value, dict) else {} + + +def _goal_terminal(todo_payload: dict[str, Any]) -> bool: + todos = todo_payload.get("todos") + if not isinstance(todos, list): + return False + statuses = [ + item.get("status") + for item in todos + if isinstance(item, dict) and item.get("priority") == "P0" + ] + return bool(statuses) and all(status in {"done", "deferred"} for status in statuses) + + +def _todo_state(args: argparse.Namespace, project: Path) -> dict[str, Any]: + completed = subprocess.run( + [ + args.loopx_cli, + "--registry", + args.registry, + "--runtime-root", + args.runtime_root, + "--format", + "json", + "todo", + "list", + "--goal-id", + args.goal_id, + ], + cwd=project, + capture_output=True, + text=True, + timeout=180, + check=False, + ) + payload = _json_payload(completed.stdout) + return { + "returncode": completed.returncode, + "terminal": completed.returncode == 0 and _goal_terminal(payload), + "payload": payload, + "stdout_tail": completed.stdout[-800:], + "stderr_tail": completed.stderr[-800:], + } + + +def _claim_primary_p0( + args: argparse.Namespace, + project: Path, +) -> dict[str, Any]: + """Keep the benchmark P0 selected across validated-progress segments.""" + state = _todo_state(args, project) + if state["returncode"] != 0: + return { + "ok": False, + "reason": "todo_list_failed", + "returncode": state["returncode"], + "stdout_tail": state["stdout_tail"], + "stderr_tail": state["stderr_tail"], + } + if state["terminal"]: + return {"ok": True, "terminal": True, "claimed": False} + + todos = state["payload"].get("todos") + open_p0 = [ + item + for item in todos if isinstance(item, dict) + and item.get("priority") == "P0" + and item.get("role") == "agent" + and item.get("status") == "open" + ] if isinstance(todos, list) else [] + if len(open_p0) != 1: + return { + "ok": False, + "reason": f"expected_one_open_agent_p0_got_{len(open_p0)}", + } + + todo = open_p0[0] + todo_id = str(todo.get("todo_id") or "") + if not todo_id: + return {"ok": False, "reason": "open_agent_p0_missing_todo_id"} + if todo.get("claimed_by") == args.agent_id: + return { + "ok": True, + "terminal": False, + "claimed": False, + "todo_id": todo_id, + "claimed_by": args.agent_id, + } + if todo.get("claimed_by") not in {None, ""}: + return { + "ok": False, + "reason": "benchmark_p0_claimed_by_other_agent", + "todo_id": todo_id, + "claimed_by": todo.get("claimed_by"), + } + + completed = subprocess.run( + [ + args.loopx_cli, + "--registry", + args.registry, + "--runtime-root", + args.runtime_root, + "--format", + "json", + "todo", + "claim", + "--goal-id", + args.goal_id, + "--todo-id", + todo_id, + "--claimed-by", + args.agent_id, + "--agent-id", + args.agent_id, + ], + cwd=project, + capture_output=True, + text=True, + timeout=180, + check=False, + ) + payload = _json_payload(completed.stdout) + ok = completed.returncode == 0 and payload.get("ok") is not False + return { + "ok": ok, + "terminal": False, + "claimed": ok, + "todo_id": todo_id, + "claimed_by": args.agent_id if ok else None, + "returncode": completed.returncode, + "stdout_tail": completed.stdout[-800:], + "stderr_tail": completed.stderr[-800:], + } + + +def _retry_delay( + payload: dict[str, Any], + args: argparse.Namespace, + failure_streak: int, +) -> float: + if not failure_streak: + return args.segment_interval_seconds + delay = args.retry_backoff_base_seconds * (2 ** (failure_streak - 1)) + host_failure = payload.get("host_failure") + if isinstance(host_failure, dict) and host_failure.get("retryable") is True: + retry = host_failure.get("retry") + if isinstance(retry, dict): + recommended = retry.get("backoff_seconds") + if isinstance(recommended, (int, float)) and recommended > 0: + delay = max(delay, float(recommended)) + return min(args.retry_backoff_cap_seconds, delay) + + +def run(args: argparse.Namespace) -> tuple[dict[str, Any], int]: + project = Path(args.cwd).resolve() + base_sha = head_sha(project) + deadline = time.monotonic() + args.goal_timeout_seconds + delivery_path = Path("/logs/agent/delivery_receipt.json") + receipts: list[dict[str, Any]] = [] + failure_streak = 0 + wrapper = Path(__file__).with_name("codex_nosandbox_wrapper.py") + delivery_helper = Path(__file__).with_name("workspace_delivery.py") + + if args.preflight_only: + payload = { + "schema_version": "deepswe_codex_cli_runner_v1", + "host_surface": "codex_cli_turn_run_once", + "preflight": wrapper.is_file() and delivery_helper.is_file(), + } + return payload, 0 if payload["preflight"] else 2 + + for segment in range(1, args.max_segments + 1): + remaining = deadline - time.monotonic() + if remaining <= 0: + break + claim = _claim_primary_p0(args, project) + if claim.get("terminal"): + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + if delivery["treatment_valid"]: + result = { + "schema_version": "deepswe_codex_cli_runner_v1", + "execution_mode": "loopx_turn_run_once", + "host_surface": "codex_cli", + "model": args.model, + "effort": args.effort, + "turn_status": "completed", + "segments": receipts, + "segment_count": len(receipts), + "delivery": delivery, + "treatment_valid": True, + "task_correctness_authority": "independent_verifier", + } + return result, 0 + receipts.append({ + "segment": segment, + "status": "terminal_todo_without_valid_delivery", + "delivery": delivery, + }) + break + if not claim.get("ok"): + receipts.append({ + "segment": segment, + "returncode": claim.get("returncode", 1), + "status": "claim_failed", + "result_kind": None, + "validation_status": None, + "p0_claim": claim, + }) + break + segment_timeout = max(60.0, min(args.segment_timeout_seconds, remaining)) + validation = [ + "python3", + str(delivery_helper), + "--project", + str(project), + "--base-sha", + base_sha, + "--receipt", + str(delivery_path), + "--require-valid", + ] + command = [ + args.loopx_cli, + "--registry", + args.registry, + "--runtime-root", + args.runtime_root, + "--format", + "json", + "turn", + "run-once", + "--goal-id", + args.goal_id, + "--agent-id", + args.agent_id, + "--turn-instance-id", + f"deepswe-codex-cli-{segment}", + "--host", + "codex-cli", + "--execution-mode", + "isolated-headless", + "--scheduler-owner", + "agent_cli_loop", + "--project", + str(project), + "--codex-bin", + str(wrapper), + "--codex-sandbox", + "workspace-write", + "--codex-model", + args.model or "", + "--validation-command-json", + json.dumps(validation), + "--validation-timeout-seconds", + "180", + "--timeout-seconds", + str(segment_timeout), + "--no-global-sync", + "--execute", + ] + env = os.environ.copy() + env["MR_REAL_CODEX"] = args.codex_bin + env["MR_LOOPX_PROJECT"] = str(project) + env["MR_CODEX_REASONING_EFFORT"] = args.effort or "" + completed = subprocess.run( + command, + cwd=project, + env=env, + capture_output=True, + text=True, + timeout=segment_timeout + 240, + check=False, + ) + payload = _json_payload(completed.stdout) + receipts.append( + { + "segment": segment, + "returncode": completed.returncode, + "status": payload.get("status"), + "result_kind": payload.get("result_kind"), + "validation_status": payload.get("validation_status"), + "p0_claim": claim, + "stdout_tail": completed.stdout[-1200:], + "stderr_tail": completed.stderr[-1200:], + } + ) + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + todo_state = _todo_state(args, project) + receipts[-1]["goal_terminal"] = todo_state["terminal"] + receipts[-1]["todo_state_returncode"] = todo_state["returncode"] + if delivery["treatment_valid"] and ( + completed.returncode == 0 + and ( + payload.get("status") == "committed" + or payload.get("validation_status") == "passed" + ) + and todo_state["terminal"] + ): + result = { + "schema_version": "deepswe_codex_cli_runner_v1", + "execution_mode": "loopx_turn_run_once", + "host_surface": "codex_cli", + "model": args.model, + "effort": args.effort, + "turn_status": "completed", + "segments": receipts, + "segment_count": len(receipts), + "delivery": delivery, + "treatment_valid": True, + "task_correctness_authority": "independent_verifier", + } + return result, 0 + if completed.returncode != 0 and not payload: + break + failure_streak = failure_streak + 1 if completed.returncode else 0 + if segment < args.max_segments: + delay = _retry_delay(payload, args, failure_streak) + time.sleep(min(delay, max(0.0, deadline - time.monotonic()))) + + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + result = { + "schema_version": "deepswe_codex_cli_runner_v1", + "execution_mode": "loopx_turn_run_once", + "host_surface": "codex_cli", + "model": args.model, + "effort": args.effort, + "turn_status": "failed", + "segments": receipts, + "segment_count": len(receipts), + "delivery": delivery, + "treatment_valid": False, + "task_correctness_authority": "independent_verifier", + "error": "codex_cli_did_not_reach_validated_delivery", + } + return result, 12 + + +def parser() -> argparse.ArgumentParser: + result = argparse.ArgumentParser() + result.add_argument("--cwd", required=True) + result.add_argument("--objective-file", required=True) + result.add_argument("--task-file", required=True) + result.add_argument("--codex-bin", default="codex") + result.add_argument("--model") + result.add_argument("--effort") + result.add_argument("--token-budget", type=int) + result.add_argument("--response-timeout-seconds", type=float, default=180) + result.add_argument("--goal-timeout-seconds", type=float, default=5400) + result.add_argument("--segment-timeout-seconds", type=float, default=7200) + result.add_argument("--max-segments", type=int, default=256) + result.add_argument("--segment-interval-seconds", type=float, default=5) + result.add_argument("--retry-backoff-base-seconds", type=float, default=10) + result.add_argument("--retry-backoff-cap-seconds", type=float, default=60) + result.add_argument("--sandbox", default="danger-full-access") + result.add_argument("--required-skill-ids", default="") + result.add_argument("--preflight-only", action="store_true") + result.add_argument("--recover-blocked", action="store_true") + result.add_argument("--max-unblocks", type=int, default=8) + result.add_argument("--loopx-cli", required=True) + result.add_argument("--registry", required=True) + result.add_argument("--runtime-root", required=True) + result.add_argument("--goal-id", required=True) + result.add_argument("--agent-id", required=True) + result.add_argument("--mode") + return result + + +def main() -> int: + payload, code = run(parser().parse_args()) + print(json.dumps(payload, indent=2, sort_keys=True)) + return code + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/loopx_heartbeat_supervisor.py b/benchmark/deepswe-gptxhigh-v1/loopx_heartbeat_supervisor.py new file mode 100755 index 0000000000..9317447d81 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/loopx_heartbeat_supervisor.py @@ -0,0 +1,403 @@ +#!/usr/bin/env python3 +"""Runner-owned recurring heartbeat host for the DeepSWE LoopX arm.""" + +from __future__ import annotations + +import argparse +import json +import os +import signal +import subprocess +import tempfile +import time +from pathlib import Path +from typing import Any + +from workspace_delivery import head_sha, normalize_delivery, write_receipt + + +def _decode_tail(value: bytes | str | None) -> str: + if value is None: + return "" + if isinstance(value, bytes): + value = value.decode("utf-8", "replace") + return value[-800:] + + +def _run_codex_segment( + command: list[str], + *, + cwd: Path, + env: dict[str, str], + trace: Path, + timeout: float, +) -> dict[str, Any]: + """Run one heartbeat wake without letting its timeout kill the supervisor.""" + timed_out = False + with trace.open("wb") as output: + process = subprocess.Popen( + command, + cwd=cwd, + env=env, + stdout=output, + stderr=subprocess.PIPE, + start_new_session=True, + ) + try: + _, stderr = process.communicate(timeout=timeout) + returncode = process.returncode + stderr_tail = _decode_tail(stderr) + except subprocess.TimeoutExpired as exc: + returncode = 124 + stderr_tail = _decode_tail(exc.stderr) + timed_out = True + try: + os.killpg(process.pid, signal.SIGTERM) + except ProcessLookupError: + pass + try: + process.wait(timeout=5) + except subprocess.TimeoutExpired: + pass + # The CLI launcher can exit before its native child. Kill the whole + # isolated group even when the direct child has already returned. + try: + os.killpg(process.pid, signal.SIGKILL) + except ProcessLookupError: + pass + _, final_stderr = process.communicate() + if final_stderr: + stderr_tail = _decode_tail(final_stderr) + return { + "returncode": returncode, + "stderr_tail": stderr_tail, + "timed_out": timed_out, + } + + +def _invoke_json(command: list[str], *, cwd: Path, timeout: float = 180) -> dict[str, Any]: + completed = subprocess.run( + command, cwd=cwd, capture_output=True, text=True, timeout=timeout, check=False + ) + try: + payload = json.loads(completed.stdout) + except ValueError: + payload = {} + return { + "returncode": completed.returncode, + "payload": payload if isinstance(payload, dict) else {}, + "stdout_tail": completed.stdout[-800:], + "stderr_tail": completed.stderr[-800:], + } + + +def _loopx(args: argparse.Namespace, *command: str, cwd: Path) -> dict[str, Any]: + return _invoke_json( + [ + args.loopx_cli, + "--registry", + args.registry, + "--runtime-root", + args.runtime_root, + "--format", + "json", + *command, + ], + cwd=cwd, + ) + + +def _goal_terminal(todo_payload: dict[str, Any]) -> bool: + todos = todo_payload.get("todos") + if not isinstance(todos, list): + return False + statuses = [ + item.get("status") + for item in todos + if isinstance(item, dict) and item.get("priority") == "P0" + ] + return bool(statuses) and all(status in {"done", "deferred"} for status in statuses) + + +def _trace_thread_id(trace: Path) -> str | None: + try: + with trace.open(encoding="utf-8", errors="replace") as handle: + for line in handle: + try: + event = json.loads(line) + except ValueError: + continue + thread_id = event.get("thread_id") + if event.get("type") == "thread.started" and isinstance(thread_id, str): + return thread_id + except OSError: + return None + return None + + +def _release_canonical_duplicate( + project: Path, base_sha: str, delivery: dict[str, Any] +) -> bool: + """Keep linked continuation work without leaving two changed checkouts.""" + if not delivery.get("treatment_valid") or not delivery.get( + "recovered_from_linked_worktree" + ): + return False + source = Path(str(delivery.get("source_worktree") or "")).resolve() + if source == project: + return False + completed = subprocess.run( + ["git", "-C", str(project), "reset", "--hard", base_sha], + capture_output=True, + check=False, + ) + if completed.returncode: + raise RuntimeError( + "could not release canonical continuation duplicate: " + + _decode_tail(completed.stderr) + ) + return True + + +def _repair_todo(args: argparse.Namespace, wake: int, cwd: Path) -> dict[str, Any]: + return _loopx( + args, + "todo", + "add", + "--goal-id", + args.goal_id, + "--role", + "agent", + "--todo-id", + f"deepswe-delivery-repair-{wake}", + "--text", + "Recover or implement the requested task in /app, run public tests, and leave a non-empty committed patch. Do not complete the Goal before this delivery exists.", + "--task-class", + "advancement_task", + "--status", + "open", + "--execute", + cwd=cwd, + ) + + +def run(args: argparse.Namespace) -> tuple[dict[str, Any], int]: + project = Path(args.cwd).resolve() + base_sha = head_sha(project) + deadline = time.monotonic() + args.goal_timeout_seconds + task_text = Path(args.task_file).read_text(encoding="utf-8").strip() + logs_dir = Path(args.logs_dir) + delivery_path = logs_dir / "delivery_receipt.json" + wakes: list[dict[str, Any]] = [] + resume_session_id: str | None = None + continuation_worktree: str | None = None + + if args.preflight_only: + prompt = _loopx( + args, + "heartbeat-prompt", + "--thin", + "--goal-id", + args.goal_id, + "--agent-id", + args.agent_id, + "--available-capability", + "shell", + "--available-capability", + "filesystem_write", + "--runtime-profile", + "outer_controller", + cwd=project, + ) + payload = { + "schema_version": "deepswe_heartbeat_supervisor_v1", + "host_surface": "recurring_outer_controller_heartbeat", + "preflight": prompt["returncode"] == 0 and bool(prompt["payload"].get("task_body")), + } + return payload, 0 if payload["preflight"] else 2 + + for wake in range(1, args.max_wakes + 1): + remaining = deadline - time.monotonic() + if remaining <= 0: + break + prompt_result = _loopx( + args, + "heartbeat-prompt", + "--thin", + "--goal-id", + args.goal_id, + "--agent-id", + args.agent_id, + "--available-capability", + "shell", + "--available-capability", + "filesystem_write", + "--runtime-profile", + "outer_controller", + cwd=project, + ) + task_body = str(prompt_result["payload"].get("task_body") or "") + if prompt_result["returncode"] != 0 or not task_body: + wakes.append({"wake": wake, "prompt": prompt_result, "error": "heartbeat_prompt_failed"}) + break + prompt = ( + task_body + + "\n\nBenchmark delivery fence: work on the task below in /app (a linked " + "worktree is allowed but will be recovered), run the repository's public tests, " + "and do not mark the Goal/Todo complete until a non-empty committed patch exists.\n\n" + + task_text + ) + if continuation_worktree: + prompt += ( + "\n\nHeartbeat continuation fence: continue the existing implementation in " + f"{continuation_worktree}; do not create another worktree or edit the " + "canonical checkout directly. Finish validation and the LoopX lifecycle " + "settlement before starting unrelated work." + ) + segment_timeout = max(60.0, min(args.segment_timeout_seconds, remaining)) + with tempfile.TemporaryDirectory(prefix="deepswe-heartbeat-") as tmp: + last = Path(tmp) / "last.txt" + trace = logs_dir / f"heartbeat-wake-{wake}.jsonl" + trace.parent.mkdir(parents=True, exist_ok=True) + common = [ + "--dangerously-bypass-approvals-and-sandbox", + "--skip-git-repo-check", + "--model", + args.model or "", + "-c", + f"model_reasoning_effort={args.effort}", + "--output-last-message", + str(last), + "--json", + ] + if resume_session_id: + command = [ + args.codex_bin, + "exec", + "resume", + *common, + resume_session_id, + prompt, + ] + else: + command = [ + args.codex_bin, + "exec", + *common, + "-C", + str(project), + "--", + prompt, + ] + segment = _run_codex_segment( + command, + cwd=project, + env=os.environ.copy(), + trace=trace, + timeout=segment_timeout, + ) + observed_session_id = _trace_thread_id(trace) + if observed_session_id: + resume_session_id = observed_session_id + last_text = last.read_text(encoding="utf-8", errors="replace") if last.exists() else "" + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + todos = _loopx(args, "todo", "list", "--goal-id", args.goal_id, cwd=project) + terminal = _goal_terminal(todos["payload"]) + wakes.append( + { + "wake": wake, + "host_process": "codex_exec", + "returncode": segment["returncode"], + "timed_out": segment["timed_out"], + "resumed": wake > 1 and resume_session_id is not None, + "session_id": resume_session_id, + "prompt_sha256": __import__("hashlib").sha256(prompt.encode()).hexdigest(), + "last_message_tail": last_text[-800:], + "stderr_tail": segment["stderr_tail"], + "delivery_status": delivery["status"], + "goal_terminal": terminal, + } + ) + if delivery["treatment_valid"] and terminal: + result = { + "schema_version": "deepswe_heartbeat_supervisor_v1", + "execution_mode": "recurring_heartbeat", + "host_surface": "outer_controller_heartbeat", + "model": args.model, + "effort": args.effort, + "continuation_owner": "benchmark_supervisor", + "turn_status": "completed", + "wake_count": len(wakes), + "wakes": wakes, + "delivery": delivery, + "treatment_valid": True, + "task_correctness_authority": "independent_verifier", + } + return result, 0 + if terminal and not delivery["treatment_valid"]: + _repair_todo(args, wake, project) + canonical_released = _release_canonical_duplicate(project, base_sha, delivery) + wakes[-1]["canonical_released_for_continuation"] = canonical_released + if canonical_released: + continuation_worktree = str(delivery["source_worktree"]) + if wake < args.max_wakes: + time.sleep(args.heartbeat_interval_seconds) + + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + result = { + "schema_version": "deepswe_heartbeat_supervisor_v1", + "execution_mode": "recurring_heartbeat", + "host_surface": "outer_controller_heartbeat", + "model": args.model, + "effort": args.effort, + "continuation_owner": "benchmark_supervisor", + "turn_status": "failed", + "wake_count": len(wakes), + "wakes": wakes, + "delivery": delivery, + "treatment_valid": False, + "task_correctness_authority": "independent_verifier", + "error": "heartbeat_did_not_reach_terminal_validated_delivery", + } + return result, 12 + + +def parser() -> argparse.ArgumentParser: + result = argparse.ArgumentParser() + result.add_argument("--cwd", required=True) + result.add_argument("--logs-dir", default="/logs/agent") + result.add_argument("--objective-file", required=True) + result.add_argument("--task-file", required=True) + result.add_argument("--codex-bin", default="codex") + result.add_argument("--model") + result.add_argument("--effort") + result.add_argument("--token-budget", type=int) + result.add_argument("--response-timeout-seconds", type=float, default=180) + result.add_argument("--goal-timeout-seconds", type=float, default=14400) + result.add_argument("--segment-timeout-seconds", type=float, default=7200) + result.add_argument("--max-wakes", type=int, default=8) + result.add_argument("--heartbeat-interval-seconds", type=float, default=5) + result.add_argument("--sandbox", default="danger-full-access") + result.add_argument("--required-skill-ids", default="") + result.add_argument("--preflight-only", action="store_true") + result.add_argument("--recover-blocked", action="store_true") + result.add_argument("--max-unblocks", type=int, default=8) + result.add_argument("--loopx-cli", required=True) + result.add_argument("--registry", required=True) + result.add_argument("--runtime-root", required=True) + result.add_argument("--goal-id", required=True) + result.add_argument("--agent-id", required=True) + result.add_argument("--mode") + return result + + +def main() -> int: + payload, code = run(parser().parse_args()) + print(json.dumps(payload, indent=2, sort_keys=True)) + return code + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/loopx_native_codex.py b/benchmark/deepswe-gptxhigh-v1/loopx_native_codex.py new file mode 100644 index 0000000000..cad1efa4f7 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/loopx_native_codex.py @@ -0,0 +1,475 @@ +#!/usr/bin/env python3 +"""The LoopX arm, driven through LoopX's own product path. + +The first LoopX arm measured the wrong thing. It called `loopx turn run-once` +from an outer loop of my own, four times, against a goal document I hand-wrote +with four stages in it. `benchmark/deepswe/README.md` rules that out in as +many words -- the treatment must not be "an outer polling loop labeled as Goal +mode" -- and it makes the distinction a formal gate: "a clean run can still be +uncountable when the treatment did not execute the preregistered LoopX path". +Under LoopX's own standard those 54 results are uncountable. They describe my +wrapper, not this product. + +The real path, all of it LoopX's: + + install_native_codex_profile scripts/install-local.sh builds a release + snapshot and installs the seven LoopX skills + into a profile-owned CODEX_HOME + loopx bootstrap LoopX writes .loopx/registry.json and + .codex/goals//ACTIVE_GOAL_STATE.md, + with its own execution profile + loopx configure-goal registers the peer identity + render_native_codex_goal_prompt the installed CLI renders the real Goal + body -- the thing I used to hand-write + run_native_goal_process_until_terminal + codex app-server owns continuation while + the Goal stays active; nothing here counts + turns + +`required_skill_ids` makes `skills/list` a precondition of thread creation, so a +run in which Codex never discovered the LoopX skills fails before model work +instead of quietly scoring as a LoopX result. + +Everything Pier does to stand Codex up is inherited from GoalCodex, which also +already owns the app-server transaction (it copies LoopX's own +`native_codex_goal.py` into the container). Only three things differ from the +Goal arm: where CODEX_HOME points, who wrote the objective, and whether the +skills gate is armed. That is the intended contrast -- the same host, the same +transaction, LoopX present or absent. + +The profile is built on the host and shipped, rather than installed in the +container: `install-local.sh` is offline (no pip, npm, curl, wget, git clone or +apt in 872 lines, so the sandbox is not the obstacle), but it verifies that its +source tree is a clean checkout, and the container receives a tarball of the +`loopx` package with no `.git` to verify. Building it once on the host keeps +`source_clean` a real claim. Both sides use the same absolute path so the +release snapshot's symlinks stay valid. + +Usage: + + MR_AGENT=loopx_native_codex:LoopxNativeCodex MR_MODEL=openai/gpt-5.5 \ + ./run.sh --all -i + +Environment: + + MR_LOOPX_PREFLIGHT=1 prove skills discovery and Goal attachment, then + stop before any model turn -- costs nothing + MR_LOOPX_PROFILE_ROOT where the profile lives on host and in container + (default /tmp/loopx-profile; must match on both) + MR_LOOPX_ROOT LoopX checkout to install from + MR_GOAL_TIMEOUT_SEC ceiling for the continuation loop +""" + +from __future__ import annotations + +import fcntl +import json +import os +import shlex +import subprocess +import sys +import tempfile +from pathlib import Path + +from goal_codex import GoalCodex, _CODEX_EXEC_MARKER, _REMOTE_DIR + +_PROFILE_ROOT = os.environ.get( + "MR_LOOPX_PROFILE_ROOT", "/tmp/loopx-profile" +) +_LOOPX_ROOT = os.path.expanduser( + os.environ.get( + "MR_LOOPX_ROOT", + str(Path(__file__).resolve().parents[1] / "wen" / "loopx"), + ) +) +_GOAL_ID = "deepswe-task" +_AGENT_ID = "deepswe-codex" +_PROJECT = "/app" +_RUNTIME_PROFILES = { + "ssh-goal": "codex_app_ssh_goal", + "codex-cli": "codex_cli", + "heartbeat": "outer_controller", +} +# Deliberately not GoalCodex's objective.txt: that file is written after this +# runs, so sharing the name would let the Goal arm's hand-written objective +# overwrite the one LoopX rendered. +_OBJECTIVE_FILE = f"{_REMOTE_DIR}/loopx_objective.txt" +_GOAL_DOC_FILE = f"{_PROFILE_ROOT}/goal-doc.md" +_WEN_COMPAT = os.environ.get("MR_LOOPX_WEN_COMPAT", "1") not in ("", "0") +_SOURCE_DIR = Path(__file__).resolve().parent +_RUNNER_SOURCES = { + "ssh-goal": _SOURCE_DIR / "loopx_wen_native_runner.py", + "codex-cli": _SOURCE_DIR / "loopx_codex_cli_runner.py", + "heartbeat": _SOURCE_DIR / "loopx_heartbeat_supervisor.py", +} +_SUPPORT_SOURCES = ( + _SOURCE_DIR / "workspace_delivery.py", + _SOURCE_DIR / "codex_nosandbox_wrapper.py", +) + + +def build_host_profile(loopx_root: str = _LOOPX_ROOT, + profile_root: str = _PROFILE_ROOT) -> dict: + """Install the formal release snapshot once, on the host. + + Reuses an existing profile: the installer refuses a non-empty target on + purpose, because mixing installation revisions would invalidate the + treatment, and re-installing per task would repeat that work 54 times. + """ + sys.path.insert(0, loopx_root) + from loopx.capabilities.benchmark_toolkit.native_codex_profile import ( + compact_native_codex_profile_receipt, + inspect_native_codex_profile, + install_native_codex_profile, + ) + + target = Path(profile_root).resolve() + target.parent.mkdir(parents=True, exist_ok=True) + # The sibling lock survives installation and serializes independent workers. + with target.with_name(target.name + ".lock").open("a") as lock: + fcntl.flock(lock, fcntl.LOCK_EX) + if target.exists() and any(target.iterdir()): + profile = inspect_native_codex_profile(target, source_root=loopx_root) + else: + profile = install_native_codex_profile(loopx_root, target) + return compact_native_codex_profile_receipt(profile) + + +class LoopxNativeCodex(GoalCodex): + """Codex with LoopX installed, driven by LoopX's rendered Goal.""" + + loopx_mode = "ssh-goal" + + async def exec_as_agent(self, environment, command, env=None, **kwargs): # noqa: ANN001 + # Intercept the same marker GoalCodex does, not the command it builds. + # GoalCodex reaches the launch through `super().exec_as_agent`, so an + # override keyed on `run_goal.py` is never called: the first version of + # this arm silently shipped no profile, armed no skills gate, and still + # produced a receipt that looked like a LoopX run. Everything below is + # therefore handed over through the environment, which GoalCodex reads + # while building its own command. + if _CODEX_EXEC_MARKER not in str(command): + return await super().exec_as_agent(environment, command, env=env, **kwargs) + + profile_receipt = build_host_profile() + mode = self.loopx_mode + runtime_profile = _RUNTIME_PROFILES[mode] + # GoalCodex owns the common Pier setup and command interception, but the + # three treatments must not share one execution surface. ssh-goal uses + # native app-server Goal attachment, codex-cli uses LoopX's built-in + # ``turn run-once`` Codex host, and heartbeat uses a recurring external + # supervisor which launches one fresh host wake per cadence tick. + self._goal_runner_source = _RUNNER_SOURCES[mode] + self._goal_runner_env = { + "LOOPX_CLI": f"{_PROFILE_ROOT}/bin/loopx", + "LOOPX_REGISTRY": f"{_PROJECT}/.loopx/registry.json", + "LOOPX_RUNTIME_ROOT": f"{_PROJECT}/.loopx/runtime", + "LOOPX_GOAL_ID": _GOAL_ID, + "LOOPX_AGENT_ID": _AGENT_ID, + "LOOPX_RECOVER_BLOCKED": "1", + "LOOPX_MODE": mode, + } + treatment_env = dict(env or {}) + treatment_env["LOOPX_WEN_COMPAT"] = "1" if _WEN_COMPAT else "0" + parent = str(Path(_PROFILE_ROOT).parent) + name = Path(_PROFILE_ROOT).name + + await super().exec_as_agent( + environment, command=f"mkdir -p {shlex.quote(_REMOTE_DIR)}", env=treatment_env + ) + + # Ship the installed profile to the identical absolute path: the release + # snapshot's `bin/loopx` is a symlink into releases/, so a different path + # would leave a dangling CLI and no Goal body could be rendered at all. + with tempfile.TemporaryDirectory() as tmp: + tarball = Path(tmp) / "profile.tar.gz" + subprocess.run( + ["tar", "czf", str(tarball), "-C", parent, name], check=True + ) + receipt_path = Path(tmp) / "profile_receipt.json" + receipt_path.write_text( + json.dumps(profile_receipt, indent=2), encoding="utf-8" + ) + bootstrap_path = Path(tmp) / "loopx_product_bootstrap.py" + bootstrap_path.write_text(_BOOTSTRAP, encoding="utf-8") + for local, remote in ( + (tarball, f"{_REMOTE_DIR}/profile.tar.gz"), + (receipt_path, f"{_REMOTE_DIR}/profile_receipt.json"), + (bootstrap_path, f"{_REMOTE_DIR}/loopx_product_bootstrap.py"), + (self._goal_runner_source, f"{_REMOTE_DIR}/{self._goal_runner_source.name}"), + *( + (source, f"{_REMOTE_DIR}/{source.name}") + for source in _SUPPORT_SOURCES + ), + ): + await environment.upload_file(str(local), remote) + + await super().exec_as_agent( + environment, + command=( + f"mkdir -p {shlex.quote(parent)} && " + f"tar xzf {shlex.quote(_REMOTE_DIR)}/profile.tar.gz " + f"-C {shlex.quote(parent)} && " + f"test -x {shlex.quote(_PROFILE_ROOT)}/bin/loopx && " + f"chmod +x {shlex.quote(_REMOTE_DIR)}/codex_nosandbox_wrapper.py" + ), + env=treatment_env, + ) + + # Bootstrap renders the product-owned objective before GoalCodex reaches + # its normal upload phase. Seed the task file here so a fresh container + # does not fail before the later upload (which also sends objective.txt + # and the native module). + with tempfile.TemporaryDirectory() as tmp: + task_path = Path(tmp) / "task.txt" + task_path.write_text(self._goal_instruction or "", encoding="utf-8") + await environment.upload_file(str(task_path), f"{_REMOTE_DIR}/task.txt") + + # LoopX writes its own registry, goal state and Goal body. Nothing in + # this arm authors goal content: the hand-written four-stage document + # the previous arm used is exactly what made its results describe a + # prompt of mine rather than this product. + await super().exec_as_agent( + environment, + command=( + # The profile is installed on the host, so its launcher records + # the host's Python path -- a uv-managed interpreter that does + # not exist in the task image, which made every CLI call exit 2 + # with "configured Python executable not found". LOOPX_PYTHON + # redirects the launcher at the container's own interpreter + # without reinstalling, so the release snapshot stays the one + # whose cleanliness was proven on the host. + f"cd {shlex.quote(_PROJECT)} && " + "export LOOPX_PYTHON=\"$(command -v python3)\" && " + "python3 -c 'import sys; assert sys.version_info >= (3, 11), sys.version' && " + f"python3 {shlex.quote(_REMOTE_DIR)}/loopx_product_bootstrap.py " + f"--profile-root {shlex.quote(_PROFILE_ROOT)} " + f"--project {shlex.quote(_PROJECT)} " + f"--goal-id {_GOAL_ID} --agent-id {_AGENT_ID} " + f"--runtime-profile {runtime_profile} " + f"--goal-doc-file {shlex.quote(_GOAL_DOC_FILE)} " + f"--task-file {shlex.quote(f'{_REMOTE_DIR}/task.txt')} " + f"--objective-out {shlex.quote(_OBJECTIVE_FILE)} " + f"--receipt-out {shlex.quote(_REMOTE_DIR)}/loopx_product.json" + ), + env=treatment_env, + ) + + os.environ["MR_GOAL_OBJECTIVE_FILE"] = _OBJECTIVE_FILE + os.environ["MR_GOAL_CODEX_HOME"] = f"{_PROFILE_ROOT}/codex-home" + os.environ["MR_GOAL_REQUIRED_SKILL_IDS"] = ",".join( + profile_receipt["required_skill_ids"] + ) + # Arm GoalCodex's LOOPX_PYTHON export for the actual app-server process, + # not just the earlier bootstrap script -- Codex's own mid-turn `loopx` + # calls run inside app-server's environment and hit the same failure + # bootstrap did until this is set here too. + os.environ["MR_GOAL_ARM_LOOPX_PYTHON"] = "1" + if os.environ.get("MR_LOOPX_PREFLIGHT", "") not in ("", "0"): + os.environ["MR_GOAL_PREFLIGHT"] = "1" + + # Give the profile's CODEX_HOME the credentials and provider settings + # Pier wrote into its own. Pointing app-server at the formally + # installed CODEX_HOME is what makes `skills/list` find LoopX, but that + # directory ships only `skills/` -- no auth.json, no config.toml, so no + # API key and no gateway base_url. A run configured that way discovers + # every skill, starts, and then waits for a model it has no address + # for: one smoke hung for 85 minutes and the gateway's call count never + # moved. The preflight cannot catch it, because Goal attachment stops + # before the first model turn and needs no credentials. + # + # Only top-level files are copied. `cp -a` of the whole directory + # would merge Pier's own skills over the formal install, and which + # skills app-server discovers is the one thing this arm must not + # improvise. + profile_home = f"{_PROFILE_ROOT}/codex-home" + await super().exec_as_agent( + environment, + command=( + 'for f in "$CODEX_HOME"/*; do ' + f'[ -f "$f" ] && cp -f "$f" {shlex.quote(profile_home)}/; ' + "done; " + # GoalCodex disables web_search by appending to *its* CODEX_HOME + # after this runs, so the copy above would leave the profile + # without it and give this arm a hosted search tool the other + # two do not have -- an advantage no network isolation would + # reveal, since the provider runs the search. + f'grep -q "^web_search" {shlex.quote(profile_home)}/config.toml ' + f'|| printf "\\nweb_search = \\"disabled\\"\\n" ' + f'>> {shlex.quote(profile_home)}/config.toml; ' + f'test -s {shlex.quote(profile_home)}/config.toml ' + '|| { echo "no config.toml reached the LoopX profile" >&2; exit 1; }' + ), + env=env, + ) + + # Capture LoopX's own trace before the container is torn down. Pier's + # own teardown copies `$CODEX_HOME/sessions` into the job directory, + # but that is Pier's CODEX_HOME -- this arm points app-server at the + # profile's instead, so codex writes its session transcript to + # `{profile_home}/sessions` and Pier's copy finds nothing. The first + # smoke run's job directory logged "No Codex session directory found" + # for exactly this reason: the transcript existed, just one directory + # over from where anyone looked for it. + # + # This has to be a second, separate call rather than the launch command + # rewritten to append a capture step. `command` at this point still + # reads "codex exec ..." -- GoalCodex.exec_as_agent below only checks + # for that marker's presence and then throws the string away, building + # its own `run_goal.py` invocation from internal state. Appending shell + # onto a string GoalCodex never looks at silently does nothing, which is + # exactly the bug that made the CODEX_HOME swap above necessary: this + # arm keeps stumbling on places where a value must go through GoalCodex + # rather than through the string it happens to be holding. + capture_dir = "/logs/agent/loopx_trace" + try: + return await super().exec_as_agent( + environment, command=command, env=env, **kwargs + ) + finally: + await super().exec_as_agent( + environment, + command=( + f"mkdir -p {shlex.quote(capture_dir)}; " + f'if [ -d {shlex.quote(profile_home)}/sessions ]; then ' + f'cp -R {shlex.quote(profile_home)}/sessions ' + f'{shlex.quote(capture_dir)}/sessions; fi; ' + f'find /app/.codex/goals -name ACTIVE_GOAL_STATE.md ' + f'-exec cp {{}} {shlex.quote(capture_dir)}/ACTIVE_GOAL_STATE.md \\; ' + f'2>/dev/null; ' + # Same LOOPX_PYTHON fix as the app-server launch above: this + # CLI call is yet another separate shell, and without its own + # export it fails with the same "configured Python + # executable not found" that showed up twice in mid-turn + # calls before the launch-side fix existed. + 'export LOOPX_PYTHON="$(command -v python3)"; ' + f'{shlex.quote(_PROFILE_ROOT)}/bin/loopx ' + f'--registry {shlex.quote(_PROJECT)}/.loopx/registry.json ' + f'--runtime-root {shlex.quote(_PROJECT)}/.loopx/runtime ' + f'--format json todo list --goal-id {_GOAL_ID} ' + f'> {shlex.quote(capture_dir)}/todo_list.json 2>&1; ' + "true" + ), + env=env, + ) + + +class LoopxCodexCliCodex(LoopxNativeCodex): + """LoopX governed Turns executed by the real Codex CLI host.""" + + loopx_mode = "codex-cli" + + +class LoopxHeartbeatCodex(LoopxNativeCodex): + """LoopX generic CLI heartbeat executed by a recurring supervisor.""" + + loopx_mode = "heartbeat" + + +# Runs inside the container: bootstrap, register the peer, render the Goal body. +_BOOTSTRAP = '''\ +import argparse, json, os, subprocess, sys +from pathlib import Path + +p = argparse.ArgumentParser() +p.add_argument("--profile-root", required=True) +p.add_argument("--project", required=True) +p.add_argument("--goal-id", required=True) +p.add_argument("--agent-id", required=True) +p.add_argument("--goal-doc-file", required=True) +p.add_argument("--task-file", required=True) +p.add_argument("--runtime-profile", required=True) +p.add_argument("--objective-out", required=True) +p.add_argument("--receipt-out", required=True) +a = p.parse_args() + +cli = f"{a.profile_root}/bin/loopx" +registry = f"{a.project}/.loopx/registry.json" +runtime = f"{a.project}/.loopx/runtime" +base = [cli, "--registry", registry, "--runtime-root", runtime, "--format", "json"] +steps = {} + + +def run(name, args): + out = subprocess.run(base + args, capture_output=True, text=True) + try: + payload = json.loads(out.stdout) + except Exception: + payload = {} + # Keep returncode and both streams for every step. Pier reports a failed + # agent command as "Command failed" and discards its output, so a step that + # dies here is otherwise invisible: the first run of this arm failed exactly + # once and left nothing but that phrase in the log. + steps[name] = { + "returncode": out.returncode, + "ok": payload.get("ok"), + "error": payload.get("error"), + "stdout_tail": out.stdout[-400:] if not payload else None, + "stderr_tail": out.stderr[-400:], + } + if out.returncode or payload.get("ok") is not True: + detail = payload.get("error") or out.stderr[-240:] or out.stdout[-240:] + raise RuntimeError(f"{name}_failed:{detail}") + return payload + + +goal_doc = Path(a.goal_doc_file) +goal_doc.write_text( + Path(a.task_file).read_text(encoding="utf-8"), + encoding="utf-8", +) +task_text = goal_doc.read_text(encoding="utf-8") +bootstrap_args = [ + "bootstrap", "--project", a.project, "--goal-id", a.goal_id, + "--objective", "Complete the software engineering task described in the task file and commit the finished work.", + "--goal-doc", str(goal_doc), + "--adapter-kind", "read_only_project_map_v0", + "--adapter-status", "connected-read-only", +] +if os.environ.get("LOOPX_WEN_COMPAT", "1") not in ("", "0"): + bootstrap_args += [ + # Benchmark admission supplies one frozen task. Project-map onboarding + # todos otherwise outrank it and measure repository housekeeping rather + # than the requested software-engineering task. + "--no-onboarding-scan", + "--begin-autonomous-advance", + "--codex-app-heartbeat", + "yes" if a.runtime_profile == "codex_app_ssh_goal" else "no", + "--write-scope", a.project, + ] +else: + bootstrap_args += ["--no-onboarding-scan", "--codex-app-heartbeat", "ask"] +run("bootstrap", bootstrap_args) +run("configure_goal", ["configure-goal", "--goal-id", a.goal_id, + "--registered-agent", a.agent_id, "--execute"]) +# The gate selects work from the todo frontier. Keeping the task only in the +# turn input lets the model spend the whole budget on onboarding todos while +# the target repository remains untouched; wen's native driver explicitly +# inserts this P0 advancement todo for the same reason. +run("add_task_todo", ["todo", "add", "--goal-id", a.goal_id, + "--role", "agent", "--todo-id", "deepswe-task", + "--text", "[P0] " + task_text, "--task-class", "advancement_task", + "--action-kind", "implement", + "--required-capability", "filesystem_write", + "--status", "open", "--execute"]) +if os.environ.get("LOOPX_WEN_COMPAT", "1") not in ("", "0"): + run("clear_waiting_on", ["configure-goal", "--goal-id", a.goal_id, + "--clear-waiting-on", + "--agent-work-mode", f"{a.agent_id}=active", + "--write-scope", a.project, "--execute"]) +prompt = run("heartbeat_prompt", + ["heartbeat-prompt", "--thin", "--goal-id", a.goal_id, + "--agent-id", a.agent_id, "--available-capability", "shell", + "--available-capability", "filesystem_write", + "--runtime-profile", a.runtime_profile, "--cli-bin", cli]) + +body = prompt.get("task_body") +if body: + Path(a.objective_out).write_text(body, encoding="utf-8") + steps["goal_body_chars"] = len(body) +else: + steps["fatal"] = "heartbeat-prompt returned no task_body" +Path(a.receipt_out).write_text(json.dumps(steps, indent=2), encoding="utf-8") +print(json.dumps(steps, indent=2)) +raise SystemExit(0 if body else 1) +''' diff --git a/benchmark/deepswe-gptxhigh-v1/loopx_wen_native_runner.py b/benchmark/deepswe-gptxhigh-v1/loopx_wen_native_runner.py new file mode 100644 index 0000000000..84e00070c0 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/loopx_wen_native_runner.py @@ -0,0 +1,457 @@ +#!/usr/bin/env python3 +"""Run the ssh-goal treatment on the native Codex app-server Goal surface. + +The LoopX toolkit owns the JSON-RPC transaction and event reducer. This +runner only adds the benchmark-host action that wen performs when a Goal is +blocked: clear the waiting gate and start the next turn on the existing thread. +""" + +from __future__ import annotations + +import argparse +import json +import os +import subprocess +import time +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Mapping + +from native_codex_goal import ( + NativeGoalConfig, + NativeGoalProtocolError, + StdioNativeGoalTransport, + compact_native_goal_receipt, + probe_native_goal_process, + observe_native_goal_event, + refresh_native_goal_status, + start_native_goal_turn, +) +from workspace_delivery import head_sha, normalize_delivery, write_receipt + + +def _restart_turn_same_thread(transport, config, turn, instruction=None): + result = transport.request( + "turn/start", + { + "threadId": turn.thread_id, + "input": [{"type": "text", "text": instruction or config.task_instruction}], + "cwd": config.cwd, + "approvalPolicy": config.approval_policy, + **({"model": config.model} if config.model else {}), + **({"effort": config.effort} if config.effort else {}), + **( + {"sandboxPolicy": dict(config.sandbox_policy)} + if config.sandbox_policy is not None + else {} + ), + }, + ) + turn.methods.append("turn/start") + nested = result.get("turn") + nested = nested if isinstance(nested, dict) else {} + turn_id = str(nested.get("id") or result.get("turnId") or "") + if not turn_id: + raise NativeGoalProtocolError("turn_start_id_missing") + turn.turn_id = turn_id + turn.response_turn_id = turn_id + turn.turn_status = str(nested.get("status") or "accepted") + return turn + + +def _reactivate_goal(transport, config, turn) -> None: + payload = { + "threadId": turn.thread_id, + "objective": config.objective, + "status": "active", + } + if config.token_budget is not None: + payload["tokenBudget"] = config.token_budget + transport.request("thread/goal/set", payload) + turn.methods.append("thread/goal/set") + turn.post_goal_status = "active" + + +def _clear_blocked( + *, + loopx_cli: str, + registry: str, + runtime_root: str, + goal_id: str, + agent_id: str, + cwd: str, +) -> None: + command = [ + loopx_cli, + "--registry", + registry, + "--runtime-root", + runtime_root, + "--format", + "json", + "configure-goal", + "--goal-id", + goal_id, + "--clear-waiting-on", + "--agent-work-mode", + f"{agent_id}=active", + "--write-scope", + cwd, + "--execute", + ] + result = subprocess.run(command, capture_output=True, text=True, timeout=180) + if result.returncode: + detail = (result.stderr or result.stdout)[-240:] + raise NativeGoalProtocolError(f"blocked_recovery_failed:{detail}") + + +@dataclass(frozen=True) +class ModelFailure: + message: str + retryable: bool + + +def classify_model_error(error: Any) -> ModelFailure: + # Unknown/unstructured failures fail closed; prose is not a retry contract. + if not isinstance(error, Mapping): + return ModelFailure(str(error), False) + message = str(error.get("message") or error.get("type") or error) + code = error.get("code") or error.get("type") + status = error.get("status") or error.get("status_code") or error.get("httpStatusCode") + info = error.get("codexErrorInfo") + if isinstance(info, Mapping): + code = next(iter(info), code) + details = info.get(code) + if isinstance(details, Mapping): + status = details.get("httpStatusCode", status) + elif isinstance(info, str): + code = info + code = str(code or "") + if code in {"insufficient_quota", "deployment_disabled", "invalid_api_key", "authentication_error", "permission_denied"}: + return ModelFailure(message, False) + try: + status = int(status) + except (TypeError, ValueError): + status = None + if status is not None: + return ModelFailure(message, status in {408, 429, 500, 502, 503, 504}) + retryable = code in { + "rate_limit_exceeded", "overloaded", "server_error", "request_timeout", + "temporarily_unavailable", "httpConnectionFailed", "responseStreamConnectionFailed", + "responseStreamDisconnected", "responseTooManyFailedAttempts", + } + return ModelFailure(message, retryable) + + +def _terminal_error(event: Mapping[str, Any]) -> ModelFailure | None: + method = str(event.get("method") or "") + params = event.get("params") if isinstance(event.get("params"), Mapping) else {} + event_type = str(event.get("type") or "") + payload = event.get("payload") if isinstance(event.get("payload"), Mapping) else {} + payload_type = str(payload.get("type") or "") + terminal = method == "turn/completed" or ( + event_type == "event_msg" and payload_type in {"task_complete", "task_completed", "turn_completed"} + ) + if not terminal: + return None + turn = params.get("turn") if isinstance(params.get("turn"), Mapping) else {} + for container in (payload, turn, params): + error = container.get("error") if isinstance(container, Mapping) else None + if error is None: + continue + return classify_model_error(error) + return None + + +def _wait_turn( + transport, + turn, + *, + deadline: float, + completed_before: int, + idle_timeout_seconds: float, + interrupt_grace_seconds: float, +) -> ModelFailure | None: + error: ModelFailure | None = None + last_event_at = time.monotonic() + interrupt_deadline: float | None = None + while turn.turn_completed_count <= completed_before: + now = time.monotonic() + remaining = deadline - now + if remaining <= 0: + raise NativeGoalProtocolError("goal_timeout_before_terminal") + event = transport.next_event(timeout_sec=min(0.25, remaining)) + if event is not None: + last_event_at = time.monotonic() + error = _terminal_error(event) or error + observe_native_goal_event(turn, event) + continue + now = time.monotonic() + if interrupt_deadline is not None: + if now >= interrupt_deadline: + raise NativeGoalProtocolError("turn_interrupt_timeout") + continue + if now - last_event_at >= idle_timeout_seconds: + transport.request( + "turn/interrupt", + {"threadId": turn.thread_id, "turnId": turn.turn_id}, + ) + error = ModelFailure(f"turn idle timeout after {idle_timeout_seconds:g} seconds", True) + interrupt_deadline = min(deadline, now + interrupt_grace_seconds) + return error + + + +def _benchmark_todo_terminal(args: argparse.Namespace, project: Path) -> bool: + result = subprocess.run( + [ + args.loopx_cli, + "--registry", + args.registry, + "--runtime-root", + args.runtime_root, + "--format", + "json", + "todo", + "list", + "--goal-id", + args.goal_id, + ], + cwd=project, + capture_output=True, + text=True, + timeout=180, + check=False, + ) + if result.returncode: + return False + try: + payload = json.loads(result.stdout) + except ValueError: + return False + todos = payload.get("todos") if isinstance(payload, dict) else None + if not isinstance(todos, list): + return False + statuses = [ + item.get("status") + for item in todos + if isinstance(item, dict) and item.get("priority") == "P0" + ] + return bool(statuses) and all(status in {"done", "deferred"} for status in statuses) + + +def run_goal(args: argparse.Namespace): + project = Path(args.cwd).resolve() + base_sha = head_sha(project) + delivery_path = Path("/logs/agent/delivery_receipt.json") + config = NativeGoalConfig( + cwd=args.cwd, + objective=Path(args.objective_file).read_text(encoding="utf-8").strip(), + task_instruction=Path(args.task_file).read_text(encoding="utf-8").strip(), + model=args.model, + effort=args.effort, + token_budget=args.token_budget, + sandbox=args.sandbox, + required_skill_ids=tuple( + item for item in (part.strip() for part in args.required_skill_ids.split(",")) if item + ), + ) + process_command = [ + args.codex_bin, + "app-server", + "--listen", + "stdio://", + "--enable", + "goals", + "--enable", + "unified_exec", + ] + process_env = os.environ.copy() + if args.preflight_only: + return probe_native_goal_process( + config, + codex_bin=args.codex_bin, + process_command=process_command, + process_env=process_env, + process_cwd=args.cwd, + response_timeout_sec=args.response_timeout_seconds, + ), 0 + + with StdioNativeGoalTransport.spawn( + process_command, + cwd=args.cwd, + env=process_env, + response_timeout_sec=args.response_timeout_seconds, + ) as transport: + turn = start_native_goal_turn(transport, config) + deadline = time.monotonic() + args.goal_timeout_seconds + completed_before = turn.turn_completed_count + unblocks = 0 + delivery_retries = 0 + transient_retries = 0 + delivery = None + benchmark_todo_terminal = False + while True: + remaining = deadline - time.monotonic() + if remaining <= 0: + raise NativeGoalProtocolError("goal_timeout_before_terminal") + turn_error = _wait_turn( + transport, + turn, + deadline=deadline, + completed_before=completed_before, + idle_timeout_seconds=args.turn_idle_timeout_seconds, + interrupt_grace_seconds=args.interrupt_grace_seconds, + ) + completed_before = turn.turn_completed_count + if turn_error is not None: + if not turn_error.retryable: + raise NativeGoalProtocolError(f"non_transient_model_error:{turn_error.message}") + if transient_retries >= args.max_transient_retries: + raise NativeGoalProtocolError( + f"transient_model_retries_exhausted:{turn_error.message}" + ) + transient_retries += 1 + delay = min( + args.retry_backoff_cap_seconds, + args.retry_backoff_base_seconds * (2 ** (transient_retries - 1)), + max(0.0, deadline - time.monotonic()), + ) + if delay <= 0: + raise NativeGoalProtocolError("goal_timeout_during_transient_backoff") + time.sleep(delay) + _reactivate_goal(transport, config, turn) + turn = _restart_turn_same_thread( + transport, + config, + turn, + "The previous model turn ended in a transient provider error. " + "Resume the active P0 task from the current Goal and workspace state; " + "do not repeat completed investigation.", + ) + completed_before = turn.turn_completed_count + continue + status = refresh_native_goal_status(transport, turn) + if status == "active": + continue + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + benchmark_todo_terminal = _benchmark_todo_terminal(args, project) + if delivery["treatment_valid"] and benchmark_todo_terminal: + break + + can_recover_blocked = ( + status == "blocked" and args.recover_blocked and unblocks < args.max_unblocks + ) + can_recover_delivery = ( + (not delivery["treatment_valid"] or not benchmark_todo_terminal) + and delivery_retries < args.max_delivery_retries + ) + if not can_recover_blocked and not can_recover_delivery: + raise NativeGoalProtocolError( + f"goal_terminal_without_valid_delivery:status={status}:" + f"delivery={delivery.get('status')}" + ) + if can_recover_blocked: + unblocks += 1 + _clear_blocked( + loopx_cli=args.loopx_cli, + registry=args.registry, + runtime_root=args.runtime_root, + goal_id=args.goal_id, + agent_id=args.agent_id, + cwd=args.cwd, + ) + if can_recover_delivery: + delivery_retries += 1 + _reactivate_goal(transport, config, turn) + repair = ( + "Runner delivery check failed. Continue the active P0 task in /app, or recover your linked " + "worktree, and do not complete the Goal until git diff from the task base " + "is non-empty, committed, and the repository's public tests pass." + ) + turn = _restart_turn_same_thread(transport, config, turn, repair) + completed_before = turn.turn_completed_count + if delivery is None: + delivery = normalize_delivery(project, base_sha) + write_receipt(delivery_path, delivery) + if not delivery["treatment_valid"]: + raise NativeGoalProtocolError("native_goal_finished_without_valid_delivery") + return ( + turn, + unblocks, + delivery_retries, + transient_retries, + benchmark_todo_terminal, + delivery, + ) + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--cwd", required=True) + parser.add_argument("--objective-file", required=True) + parser.add_argument("--task-file", required=True) + parser.add_argument("--codex-bin", default="codex") + parser.add_argument("--model") + parser.add_argument("--effort") + parser.add_argument("--token-budget", type=int) + parser.add_argument("--response-timeout-seconds", type=float, default=30) + parser.add_argument("--goal-timeout-seconds", type=float, default=3600) + parser.add_argument("--sandbox", default="danger-full-access") + parser.add_argument("--required-skill-ids", default="") + parser.add_argument("--preflight-only", action="store_true") + parser.add_argument("--recover-blocked", action="store_true") + parser.add_argument("--max-unblocks", type=int, default=8) + parser.add_argument("--max-delivery-retries", type=int, default=3) + parser.add_argument("--max-transient-retries", type=int, default=12) + parser.add_argument("--retry-backoff-base-seconds", type=float, default=5) + parser.add_argument("--retry-backoff-cap-seconds", type=float, default=60) + parser.add_argument("--turn-idle-timeout-seconds", type=float, default=300) + parser.add_argument("--interrupt-grace-seconds", type=float, default=60) + parser.add_argument("--loopx-cli", default="loopx") + parser.add_argument("--registry", required=True) + parser.add_argument("--runtime-root", required=True) + parser.add_argument("--goal-id", required=True) + parser.add_argument("--agent-id", required=True) + parser.add_argument("--mode", choices=("ssh-goal",), default="ssh-goal") + args = parser.parse_args() + + result = run_goal(args) + if args.preflight_only: + turn, unblocks = result + delivery_retries = 0 + transient_retries = 0 + benchmark_todo_terminal = False + delivery = None + else: + ( + turn, + unblocks, + delivery_retries, + transient_retries, + benchmark_todo_terminal, + delivery, + ) = result + receipt = compact_native_goal_receipt(turn) + receipt["execution_mode"] = "goal_attachment_preflight" if args.preflight_only else "goal_until_terminal" + receipt["loopx_mode"] = args.mode + receipt["continuation_owner"] = "codex" + receipt["loopx_unblock_count"] = unblocks + receipt["loopx_blocked_recovery"] = bool(args.recover_blocked) + receipt["delivery_retry_count"] = delivery_retries + receipt["transient_retry_count"] = transient_retries + receipt["turn_idle_timeout_seconds"] = args.turn_idle_timeout_seconds + receipt["benchmark_todo_terminal"] = benchmark_todo_terminal + receipt["delivery"] = delivery + receipt["treatment_valid"] = bool(delivery and delivery.get("treatment_valid")) + receipt["host_surface"] = "native_goal_appserver" + receipt["model"] = args.model + receipt["effort"] = args.effort + receipt["task_correctness_authority"] = "independent_verifier" + print(json.dumps(receipt, indent=2, sort_keys=True)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/pier_cn.py b/benchmark/deepswe-gptxhigh-v1/pier_cn.py new file mode 100644 index 0000000000..8ff37abac1 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/pier_cn.py @@ -0,0 +1,360 @@ +#!/usr/bin/env python +"""Pier, with agent-image builds pointed at domestic mirrors. + +Pier bakes the agent harness into every task image. Each of DeepSWE's 113 tasks +ships its own base image, so this layer is rebuilt 113 times and never reused +between tasks — whatever it costs, it costs 113 times, per harness. From here +the upstream install path is slow or fatal at three points: + + * ``apt-get update`` hits deb.debian.org, which does not answer at all from + this host; a build sat there for 25 min before it was killed. + * ``curl -LsSf https://astral.sh/uv/... | sh`` (mini-swe-agent) pulls a 17 MB + binary from GitHub only to *downgrade* the uv 0.9.18 these images already + carry in /root/.local/bin. + * ``uv tool install`` resolves against pypi.org, unreachable here without the + proxy and timing out even with it; ``npm install -g`` against + registry.npmjs.org (claude-code, codex, opencode) works but at 0.8 MB/s. + +So: keep the image's own uv when it has one, and resolve everything from +mirrors.aliyun.com / registry.npmmirror.com, which answer directly in ~0.2-0.5 s +at 2.8-5.6 MB/s and need no proxy. Measured end to end on one task image for +mini-swe-agent: 366 s and then a failure, against 8.8 s. + +The proxy still applies to the build (run.sh points DOCKER_CONFIG at +mr_common/docker-config); this only removes the steps that depend on it being +both up and fast. The mirror hosts are in that file's ``noProxy`` on purpose — +routed through the proxy they come back 502, and they need no proxy anyway. + +run.sh calls this instead of .venv/bin/pier. A plain `pier` is untouched and +still behaves exactly as upstream ships it. +""" + +from __future__ import annotations + +import importlib +import os +import re +import sys + +PYPI_INDEX = "https://mirrors.aliyun.com/pypi/simple/" +NPM_REGISTRY = "https://registry.npmmirror.com" +APT_MIRROR_HOST = "mirrors.aliyun.com" + +# UV_DEFAULT_INDEX is what uv >= 0.6 reads; UV_INDEX_URL keeps older uv (and +# pier's pinned 0.7.13, should an image ever need it fetched) on the mirror too. +PYPI_ENV = { + "UV_DEFAULT_INDEX": PYPI_INDEX, + "UV_INDEX_URL": PYPI_INDEX, + "PIP_INDEX_URL": PYPI_INDEX, + "UV_HTTP_TIMEOUT": "120", +} + +# npm reads either spelling depending on version; setting both is free. +NPM_ENV = { + "NPM_CONFIG_REGISTRY": NPM_REGISTRY, + "npm_config_registry": NPM_REGISTRY, +} + +# Matched verbatim against pier's install script, so an upstream change to the +# pinned uv version makes the patch fall through loudly rather than silently +# stop applying. +UV_INSTALL_LINE = "curl -LsSf https://astral.sh/uv/0.7.13/install.sh | sh" + +UV_REUSE_BLOCK = f"""if ! command -v uv >/dev/null 2>&1 && [ ! -x "$HOME/.local/bin/uv" ]; then + {UV_INSTALL_LINE} +fi +mkdir -p "$HOME/.local/bin" +[ -f "$HOME/.local/bin/env" ] || : > "$HOME/.local/bin/env" """ + +# Runs before the package manager. Images that are not Debian match no files. +APT_MIRROR_PREFIX = ( + "for f in /etc/apt/sources.list /etc/apt/sources.list.d/*.sources" + " /etc/apt/sources.list.d/*.list; do" + ' [ -f "$f" ] || continue;' + f' sed -i "s|deb.debian.org|{APT_MIRROR_HOST}|g;' + f' s|security.debian.org|{APT_MIRROR_HOST}|g" "$f";' + " done 2>/dev/null; true; " +) + +# Every harness Pier can install that we might sweep with. Listed explicitly +# rather than walked from __subclasses__: the subclasses only exist once their +# module is imported, and an agent that quietly missed the patch would show up +# as a 25-minute hang rather than an error. +HARNESSES = ( + ("pier.agents.installed.mini_swe_agent", "MiniSweAgent"), + ("pier.agents.installed.claude_code", "ClaudeCode"), + ("pier.agents.installed.codex", "Codex"), + ("pier.agents.installed.opencode", "OpenCode"), + ("pier.agents.installed.cursor_cli", "CursorCli"), + ("pier.agents.installed.gemini_cli", "GeminiCli"), +) + + +def rewrite_step(step) -> list[str]: + """Point one install step at the mirrors. Returns what it changed.""" + changed = [] + run = step.run + + if ("apt-get" in run or "apk add" in run) and APT_MIRROR_PREFIX not in run: + run = APT_MIRROR_PREFIX + run + changed.append("apt") + + if UV_INSTALL_LINE in run: + run = run.replace(UV_INSTALL_LINE, UV_REUSE_BLOCK) + changed.append("uv-reuse") + if "uv tool install" in run or "uv pip install" in run or "pip install" in run: + step.env = {**(step.env or {}), **PYPI_ENV} + changed.append("pypi") + + if "npm install" in run or "npm i " in run: + step.env = {**(step.env or {}), **NPM_ENV} + # Registry downloads occasionally reset the TLS connection on this host; + # make image builds retry instead of turning a transient reset into an + # invalid benchmark trial. + run = run.replace( + "npm install -g @openai/codex@latest", + "npm install --fetch-retries=5 --fetch-retry-mintimeout=5000 " + "--fetch-retry-maxtimeout=60000 -g @openai/codex@latest", + ) + # Debian task images already ship node/npm. Avoid the upstream nvm + # bootstrap from GitHub, which is unreachable when the host proxy is + # down; install the harness directly from the configured npm mirror. + if "nvm-sh/nvm" in run: + run = re.sub( + r"else\s+curl -o- https://raw\.githubusercontent\.com/nvm-sh/nvm/.*?; fi && codex --version", + "else npm install --fetch-retries=5 --fetch-retry-mintimeout=5000 " + "--fetch-retry-maxtimeout=60000 -g @openai/codex@latest; " + "fi && codex --version", + run, + flags=re.DOTALL, + ) + changed.append("nvm-bypass") + changed.append("npm") + + step.run = run + return changed + + +def patch_harnesses() -> None: + """Wrap install_spec on every known harness class.""" + patched_any = False + for module_path, class_name in HARNESSES: + try: + cls = getattr(importlib.import_module(module_path), class_name) + except (ImportError, AttributeError) as exc: + print(f"pier_cn: skipping {class_name} ({exc})", file=sys.stderr) + continue + + original = cls.install_spec + + def install_spec(self, _original=original, _name=class_name): + spec = _original(self) + changed = [] + for step in spec.steps: + changed += rewrite_step(step) + if not changed: + # Pier changed its install script out from under the patch: say + # so rather than let a sweep crawl or die one task at a time. + raise RuntimeError( + f"pier_cn: no install step of {_name} matched; refusing an unpatched harness" + ) + return spec + + cls.install_spec = install_spec + patched_any = True + + if not patched_any: + raise SystemExit("pier_cn: could not patch any harness — refusing to run") + + +def patch_goal_mode() -> None: + """Point the ``codex`` agent name at one of the experiment's two arms. + + Pier's CLI validates ``--agent`` against the ``AgentName`` enum, so the + factory's import-path loader ('module:Class') cannot be reached from the + command line and a new name cannot simply be added. Rebinding the existing + name suits the experiment better anyway: both arms then run byte-for-byte + identical commands and differ only in this one environment variable, which + is what makes the comparison controlled. + + MR_CODEX_ARM=goal Goal attached over the app-server (treatment) + MR_CODEX_ARM=plain same objective text and web_search setting, but + app-server transport with no Goal (control) + MR_CODEX_ARM=loopx-native + LoopX through its own product path: formal release + install, LoopX-rendered Goal body, and continuation + owned by app-server. `loopx` predates it and drove + `turn run-once` from an outer loop with a + hand-written goal document, which LoopX's own + benchmark method rules out; its results describe + that wrapper rather than this product. + MR_CODEX_ARM=loopx-native-deepseek + Byte-identical to loopx-native; only the arm name + differs, so it lands in its own + deepswe--loopx-native-deepseek job + directory instead of overwriting the gpt-5.5 + native run already recorded there. Combine with + MR_MODEL to actually route to a different model. + MR_CODEX_ARM=loopx-native-deepseek-flash + Same as loopx-native-deepseek; a separate alias so + a deepseek-v4-flash sweep lands in its own job + directory instead of mixing into + deepswe--loopx-native-deepseek, which + already holds the deepseek-v4-pro results. + MR_CODEX_ARM=loopx-native-codex-cli + Official `codex_cli` profile through LoopX + `turn run-once`, launching real `codex exec` turns. + MR_CODEX_ARM=loopx-native-heartbeat + Official `generic_cli` heartbeat profile; recurring + wakes are owned by the benchmark supervisor. + + MR_GOAL_MODE=1 is still honoured as a synonym for the goal arm. Unset, an + ordinary sweep is untouched. + """ + arm = os.environ.get("MR_CODEX_ARM", "").strip().lower() + if not arm and os.environ.get("MR_GOAL_MODE", "") not in ("", "0"): + arm = "goal" + if not arm: + return + if arm not in ("goal", "plain", "loopx-native", "loopx-native-deepseek", + "loopx-native-deepseek-flash", "loopx-native-codex-cli", + "loopx-native-heartbeat"): + raise SystemExit( + "pier_cn: MR_CODEX_ARM must be 'goal', 'plain', 'loopx-native', " + "'loopx-native-deepseek', 'loopx-native-deepseek-flash', " + "'loopx-native-codex-cli' or 'loopx-native-heartbeat', " + f"got {arm!r}" + ) + + from pier.agents.factory import AgentFactory + from pier.models.agent.name import AgentName + + import goal_codex + + if arm in ("loopx-native", "loopx-native-deepseek", "loopx-native-deepseek-flash", + "loopx-native-codex-cli", "loopx-native-heartbeat"): + os.environ["MR_LOOPX_MODE"] = { + "loopx-native": "ssh-goal", + "loopx-native-deepseek": "ssh-goal", + "loopx-native-deepseek-flash": "ssh-goal", + "loopx-native-codex-cli": "codex-cli", + "loopx-native-heartbeat": "heartbeat", + }[arm] + import loopx_native_codex + + agent = { + "loopx-native": loopx_native_codex.LoopxNativeCodex, + "loopx-native-deepseek": loopx_native_codex.LoopxNativeCodex, + "loopx-native-deepseek-flash": loopx_native_codex.LoopxNativeCodex, + "loopx-native-codex-cli": loopx_native_codex.LoopxCodexCliCodex, + "loopx-native-heartbeat": loopx_native_codex.LoopxHeartbeatCodex, + }[arm] + else: + agent = { + "goal": goal_codex.GoalCodex, + "plain": goal_codex.PlainAppServerCodex, + }[arm] + AgentFactory._AGENT_MAP[AgentName.CODEX] = agent + print( + f"pier_cn: MR_CODEX_ARM={arm} — 'codex' now runs as {agent.__qualname__}", + file=sys.stderr, + ) + + +def patch_claude_arm() -> None: + """Reject unsupported Claude arms before starting this Codex study.""" + if os.environ.get("MR_CLAUDE_ARM", "").strip(): + raise SystemExit("pier_cn: Claude arms are not part of this five-arm snapshot") + + +def patch_modelonly_network() -> None: + """Attach Pier's task container to the internal model-only network. + + The compose overlay is appended after Pier's generated files, so the task + can reach the host gateway at 127.0.0.1 without relying on the egress + proxy or the container's loopback address. + """ + if os.environ.get("MR_MODELONLY_NET", "0") not in ("1", "true", "yes"): + return + from pathlib import Path + from pier.environments.docker.docker import DockerEnvironment + + local_agent_base = os.environ.get("MR_LOCAL_AGENT_BASE_IMAGE", "").strip() + if local_agent_base and not getattr( + DockerEnvironment._prepare_agent_build_context, "_deepswe_local_base", False + ): + original_prepare = DockerEnvironment._prepare_agent_build_context + + def prepare_agent_build_context(self): + # BuildKit performs a registry metadata HEAD even when the remote + # base image was already loaded. Use the explicit local alias only + # while Pier writes the generated agent Dockerfile; the task config + # is restored before Compose starts, so scoring metadata remains + # unchanged. + original_image = self.task_env_config.docker_image + self.task_env_config.docker_image = local_agent_base + try: + return original_prepare(self) + finally: + self.task_env_config.docker_image = original_image + + prepare_agent_build_context._deepswe_local_base = True + DockerEnvironment._prepare_agent_build_context = prepare_agent_build_context + print( + f"pier_cn: local agent base override={local_agent_base}", + file=sys.stderr, + ) + + overlay_path = os.environ.get("MR_MODELONLY_COMPOSE", "").strip() + if not overlay_path: + raise SystemExit("pier_cn: set MR_MODELONLY_COMPOSE to an external model-only compose overlay") + overlay = Path(overlay_path).expanduser().resolve() + if not overlay.is_file(): + raise SystemExit(f"pier_cn: missing model-only compose overlay: {overlay}") + original = DockerEnvironment._docker_compose_paths + if getattr(original, "_deepswe_modelonly", False): + return + + def paths(self): + # Verifier environments are deliberately no-network. Adding a + # `networks` key there conflicts with Compose's `network_mode: none`. + # Pier uses `no-network` for both agent and verifier task declarations; + # the verifier's environment directory is the reliable distinction. + if Path(self.environment_dir).name == "tests": + return list(original.fget(self)) + result = list(original.fget(self)) + result.append(overlay) + return result + + paths._deepswe_modelonly = True + DockerEnvironment._docker_compose_paths = property(paths) + original_agent_env = DockerEnvironment.agent_process_env + + def agent_process_env(self, env): + result = dict(original_agent_env(self, env) or {}) + for key in ("HTTP_PROXY", "HTTPS_PROXY", "http_proxy", "https_proxy"): + result[key] = "" + no_proxy = result.get("NO_PROXY", "") + result["NO_PROXY"] = ",".join( + item for item in (no_proxy, "localhost", "127.0.0.1", "127.0.0.1") if item + ) + result["no_proxy"] = result["NO_PROXY"] + return result + + agent_process_env._deepswe_modelonly = True + DockerEnvironment.agent_process_env = agent_process_env + print(f"pier_cn: model-only network overlay={overlay}", file=sys.stderr) + + +def main() -> int: + patch_harnesses() + patch_goal_mode() + patch_claude_arm() + patch_modelonly_network() + from pier.cli.main import app + + return app() + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/plain_appserver_runner.py b/benchmark/deepswe-gptxhigh-v1/plain_appserver_runner.py new file mode 100644 index 0000000000..84865d6d86 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/plain_appserver_runner.py @@ -0,0 +1,158 @@ +#!/usr/bin/env python3 +"""Run one Codex app-server turn without attaching a Goal. + +This is the wen-aligned plain control: the transport and permissions match the +Goal/LoopX arms, but no ``thread/goal/set`` call is made and no LoopX command is +invoked. It accepts the same CLI shape as the native Goal runner because +GoalCodex owns the common Pier upload/setup path. +""" + +from __future__ import annotations + +import argparse +import json +import os +from pathlib import Path + +from native_codex_goal import ( + NativeGoalConfig, + NativeGoalProtocolError, + NativeGoalTurn, + StdioNativeGoalTransport, + _nested, + compact_native_goal_receipt, + wait_native_goal_turn, +) + + +def start_turn_without_goal( + transport: StdioNativeGoalTransport, config: NativeGoalConfig, *, execute: bool = True +) -> NativeGoalTurn: + transport.request( + "initialize", + { + "clientInfo": { + "name": "deepswe_plain_appserver", + "title": "DeepSWE plain (app-server, no Goal)", + "version": "0.1.0", + }, + "capabilities": {"experimentalApi": True}, + }, + ) + transport.notify("initialized", {}) + thread_result = transport.request( + "thread/start", + { + "cwd": config.cwd, + "sandbox": config.sandbox, + "approvalPolicy": config.approval_policy, + **({"model": config.model} if config.model else {}), + }, + ) + thread = _nested(thread_result, "thread") + thread_id = str(thread.get("id") or thread_result.get("threadId") or "") + if not thread_id: + raise NativeGoalProtocolError("thread_start_id_missing") + + turn = NativeGoalTurn( + thread_id=thread_id, + turn_id="", + response_turn_id="", + goal_status="none", + objective_sha256="", + objective_chars=0, + task_instruction_sha256="", + task_instruction_chars=len(config.task_instruction), + token_budget_present=False, + methods=["initialize", "initialized", "thread/start"], + ) + if not execute: + return turn + turn_result = transport.request( + "turn/start", + { + "threadId": thread_id, + "input": [{"type": "text", "text": config.task_instruction}], + "cwd": config.cwd, + "approvalPolicy": config.approval_policy, + **({"model": config.model} if config.model else {}), + **({"effort": config.effort} if config.effort else {}), + }, + ) + turn.methods.append("turn/start") + response_turn = _nested(turn_result, "turn") + turn_id = str(response_turn.get("id") or turn_result.get("turnId") or "") + if not turn_id: + raise NativeGoalProtocolError("turn_start_id_missing") + turn.turn_id = turn_id + turn.response_turn_id = turn_id + turn.turn_status = str(response_turn.get("status") or "accepted") + return turn + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--cwd", required=True) + parser.add_argument("--objective-file", required=True) + parser.add_argument("--task-file", required=True) + parser.add_argument("--codex-bin", default="codex") + parser.add_argument("--model") + parser.add_argument("--effort") + parser.add_argument("--token-budget", type=int) + parser.add_argument("--response-timeout-seconds", type=float, default=180) + parser.add_argument("--goal-timeout-seconds", type=float, default=3600) + parser.add_argument("--sandbox", default="danger-full-access") + parser.add_argument("--required-skill-ids", default="") + parser.add_argument("--preflight-only", action="store_true") + args = parser.parse_args() + + config = NativeGoalConfig( + cwd=args.cwd, + objective="No Goal is attached in the plain control arm.", + task_instruction=Path(args.task_file).read_text(encoding="utf-8").strip(), + model=args.model, + effort=args.effort, + token_budget=None, + approval_policy="never", + sandbox=args.sandbox, + ) + command = [ + args.codex_bin, + "app-server", + "--listen", + "stdio://", + "--enable", + "goals", + "--enable", + "unified_exec", + ] + transport = StdioNativeGoalTransport.spawn( + command, + cwd=args.cwd, + env=os.environ.copy(), + response_timeout_sec=args.response_timeout_seconds, + stderr=None, + ) + turn = None + try: + turn = start_turn_without_goal(transport, config, execute=not args.preflight_only) + try: + if not args.preflight_only: + wait_native_goal_turn( + transport, turn, timeout_sec=args.goal_timeout_seconds + ) + except NativeGoalProtocolError as exc: + if str(exc) != "goal_turn_timeout": + raise + receipt = compact_native_goal_receipt(turn) + receipt["execution_mode"] = ( + "plain_appserver_preflight" if args.preflight_only else "plain_appserver_single_turn" + ) + print(json.dumps(receipt, indent=2, sort_keys=True)) + return 0 if args.preflight_only or turn.turn_status == "completed" else 1 + finally: + transport.close() + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py b/benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py new file mode 100644 index 0000000000..93f067d858 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py @@ -0,0 +1,304 @@ +#!/usr/bin/env python3 +"""Fail-closed admission check run independently before every LoopX arm.""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import re +import socket +import subprocess +import sys +import tempfile +import tomllib +from datetime import datetime, timezone +from pathlib import Path + +from workspace_delivery import head_sha, normalize_delivery + + +EXPECTED_MODEL = "openai/gpt-5.6-sol" +EXPECTED_EFFORT = "xhigh" +EXPECTED_AGENT_TIMEOUT_MULTIPLIER = 3.0 +EXPECTED_GOAL_TIMEOUT_SECONDS = 14400.0 +EXPECTED_HEARTBEAT_SEGMENT_TIMEOUT_SECONDS = 7200.0 +EXPECTED_TURN_IDLE_TIMEOUT_SECONDS = 7200.0 +EXPECTED_LOOPX_REVISION = "2cef51d08b2a0103f4ba026bf47fd70dc8acee30" +RUNNERS = { + "ssh-goal": { + "file": "loopx_wen_native_runner.py", + "surface": "native_goal_appserver", + "required": ('"app-server"', '"turn/interrupt"', '"task_correctness_authority"', '"independent_verifier"'), + "forbidden": (), + }, + "codex-cli": { + "file": "loopx_codex_cli_runner.py", + "surface": "codex_cli", + "required": ('"run-once"', '"--host"', '"codex-cli"', '"task_correctness_authority"', '"independent_verifier"'), + "forbidden": ('"app-server"',), + }, + "heartbeat": { + "file": "loopx_heartbeat_supervisor.py", + "surface": "outer_controller_heartbeat", + "required": ('"heartbeat-prompt"', '"outer_controller"', '"codex_exec"', '"continuation_owner": "benchmark_supervisor"', '"task_correctness_authority": "independent_verifier"'), + "forbidden": ('"app-server"',), + }, +} + + +def digest(path: Path) -> str: + return hashlib.sha256(path.read_bytes()).hexdigest() + + +def _git(*args: str, cwd: Path) -> subprocess.CompletedProcess[str]: + return subprocess.run( + ("git", *args), cwd=cwd, capture_output=True, text=True, check=False + ) + + +def delivery_self_test() -> tuple[bool, str]: + """Exercise the exact ignored-control-state failure seen in smoke.""" + with tempfile.TemporaryDirectory(prefix="deepswe-delivery-preflight-") as tmp: + project = Path(tmp) / "project" + project.mkdir() + commands = ( + ("init", "-q"), + ("config", "user.name", "DeepSWE preflight"), + ("config", "user.email", "preflight@deepswe.invalid"), + ) + for command in commands: + result = _git(*command, cwd=project) + if result.returncode: + return False, result.stderr[-300:] + (project / "source.txt").write_text("base\n", encoding="utf-8") + if _git("add", "source.txt", cwd=project).returncode: + return False, "could_not_stage_base" + if _git("commit", "-qm", "base", cwd=project).returncode: + return False, "could_not_commit_base" + base = head_sha(project) + for name in (".loopx", ".codex"): + (project / name).mkdir() + (project / name / "state.json").write_text("{}\n", encoding="utf-8") + (project / "delivered.txt").write_text("ok\n", encoding="utf-8") + receipt = normalize_delivery(project, base) + tracked = _git("ls-files", cwd=project).stdout.splitlines() + valid = ( + receipt.get("treatment_valid") is True + and receipt.get("patch_applies") is True + and "delivered.txt" in tracked + and not any(path.startswith((".loopx/", ".codex/")) for path in tracked) + ) + return valid, str(receipt.get("reason") or "unknown") + + +def port_is_free(port: int) -> bool: + for host in ("127.0.0.1", "127.0.0.1"): + with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as probe: + probe.settimeout(0.1) + if probe.connect_ex((host, port)) == 0: + return False + return True + + +def load_task_list(path: Path, expected_count: int) -> tuple[str, ...]: + tasks = tuple(path.read_text(encoding="utf-8").splitlines()) + if expected_count <= 0 or len(tasks) != expected_count: + raise ValueError(f"expected {expected_count} task ids, got {len(tasks)}") + if len(set(tasks)) != len(tasks): + raise ValueError("duplicate task ids") + if any(re.fullmatch(r"[A-Za-z0-9][A-Za-z0-9_.-]*", task) is None for task in tasks): + raise ValueError("invalid task id: expected a single directory name per line") + return tasks + + +def task_manifest(task_root: Path, tasks: tuple[str, ...]) -> tuple[str, float, list[str]]: + entries = [] + timeouts = [] + missing = [] + for task in tasks: + path = task_root / task / "task.toml" + if not path.is_file(): + missing.append(task) + continue + try: + raw = path.read_bytes() + data = tomllib.loads(raw.decode("utf-8")) + base = str(data.get("metadata", {}).get("base_commit_hash") or "") + timeout = float(data.get("agent", {}).get("timeout_sec") or 0) + except (OSError, ValueError, TypeError, AttributeError): + missing.append(task) + continue + if re.fullmatch(r"[0-9a-fA-F]{7,40}", base) is None or timeout <= 0: + missing.append(task) + continue + timeouts.append(timeout) + entries.append(f"{task}\t{base}\t{hashlib.sha256(raw).hexdigest()}") + payload = "\n".join(entries).encode("utf-8") + return hashlib.sha256(payload).hexdigest(), min(timeouts, default=0), missing + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--arm", choices=tuple(RUNNERS), required=True) + parser.add_argument("--loopx-root", type=Path, required=True) + parser.add_argument("--port", type=int, required=True) + parser.add_argument("--model", required=True) + parser.add_argument("--effort", required=True) + parser.add_argument("--goal-timeout", type=float, required=True) + parser.add_argument("--heartbeat-segment-timeout", type=float, required=True) + parser.add_argument("--turn-idle-timeout", type=float, required=True) + parser.add_argument("--agent-timeout-multiplier", type=float, required=True) + parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--task-list", type=Path, required=True) + parser.add_argument("--expected-task-count", type=int, required=True) + args = parser.parse_args() + + root = Path(__file__).resolve().parent + errors: list[str] = [] + contract = RUNNERS[args.arm] + runner_name = str(contract["file"]) + surface = str(contract["surface"]) + runner = root / runner_name + delivery = root / "workspace_delivery.py" + adapter = root / "loopx_native_codex.py" + wrapper = root / "codex_nosandbox_wrapper.py" + harness = root / "goal_codex.py" + task_root = root / "upstream" / "tasks" + try: + tasks = load_task_list(args.task_list, args.expected_task_count) + except (OSError, ValueError) as exc: + parser.error(str(exc)) + manifest_sha256, minimum_agent_timeout, malformed_tasks = task_manifest(task_root, tasks) + if args.port in {4141, 4250}: + errors.append("prohibited_gateway_port") + if not port_is_free(args.port): + errors.append("gateway_port_occupied") + if args.model != EXPECTED_MODEL: + errors.append("model_not_fair_control") + if args.effort != EXPECTED_EFFORT: + errors.append("effort_not_fair_control") + if args.agent_timeout_multiplier != EXPECTED_AGENT_TIMEOUT_MULTIPLIER: + errors.append("agent_timeout_multiplier_not_fair_control") + if args.goal_timeout != EXPECTED_GOAL_TIMEOUT_SECONDS: + errors.append("goal_timeout_not_pinned") + if args.heartbeat_segment_timeout != EXPECTED_HEARTBEAT_SEGMENT_TIMEOUT_SECONDS: + errors.append("heartbeat_segment_timeout_not_pinned") + if args.turn_idle_timeout != EXPECTED_TURN_IDLE_TIMEOUT_SECONDS: + errors.append("turn_idle_timeout_not_pinned") + if args.heartbeat_segment_timeout >= args.goal_timeout: + errors.append("heartbeat_segment_not_below_goal_timeout") + if minimum_agent_timeout <= 0 or args.goal_timeout >= minimum_agent_timeout * args.agent_timeout_multiplier: + errors.append("inner_timeout_not_below_agent_timeout") + source = runner.read_text(encoding="utf-8") if runner.is_file() else "" + if not runner.is_file(): + errors.append("runner_missing") + if any(marker not in source for marker in contract["required"]): + errors.append("runner_contract_marker_missing") + if any(marker in source for marker in contract["forbidden"]): + errors.append("runner_uses_wrong_surface") + if not delivery.is_file(): + errors.append("delivery_gate_missing") + adapter_source = adapter.read_text(encoding="utf-8") if adapter.is_file() else "" + if not adapter.is_file() or '"--accept-onboarding-agent-todos"' in adapter_source: + errors.append("benchmark_adapter_allows_onboarding_todos") + if '"--no-onboarding-scan"' not in adapter_source or '"--text", "[P0] " + task_text' not in adapter_source: + errors.append("benchmark_task_admission_contract_missing") + wrapper_source = wrapper.read_text(encoding="utf-8") if wrapper.is_file() else "" + if args.arm == "codex-cli" and ( + not wrapper.is_file() or "MR_CODEX_REASONING_EFFORT" not in wrapper_source + ): + errors.append("codex_cli_effort_injection_missing") + if args.arm == "heartbeat" and "model_reasoning_effort=" not in source: + errors.append("heartbeat_effort_injection_missing") + harness_source = harness.read_text(encoding="utf-8") if harness.is_file() else "" + if not harness.is_file() or 'args += ["--effort", effort]' not in harness_source: + errors.append("runner_effort_forwarding_missing") + if args.arm in {"codex-cli", "heartbeat"} and ( + 'runner_env.get("LOOPX_MODE") in {"codex-cli", "heartbeat"}' not in harness_source + or '"--segment-timeout-seconds"' not in harness_source + or 'MR_HEARTBEAT_SEGMENT_TIMEOUT_SEC' not in harness_source + ): + errors.append("runner_segment_timeout_forwarding_missing") + if args.arm == "ssh-goal" and ( + 'runner_env.get("LOOPX_MODE") == "ssh-goal"' not in harness_source + or '"--turn-idle-timeout-seconds"' not in harness_source + or "MR_LOOPX_TURN_IDLE_TIMEOUT_SEC" not in harness_source + ): + errors.append("runner_turn_idle_timeout_forwarding_missing") + compile_check = subprocess.run( + [sys.executable, "-m", "py_compile", str(runner), str(delivery)], + capture_output=True, + text=True, + check=False, + ) + if compile_check.returncode: + errors.append("runner_or_delivery_not_compilable") + delivery_probe_ok, delivery_probe_detail = delivery_self_test() + if not delivery_probe_ok: + errors.append("delivery_self_test_failed") + if malformed_tasks: + errors.append("missing_or_malformed_task_definitions") + git = subprocess.run( + ["git", "-C", str(args.loopx_root), "status", "--porcelain"], + capture_output=True, + text=True, + check=False, + ) + if git.returncode or git.stdout.strip(): + errors.append("loopx_source_not_clean") + revision = subprocess.run( + ["git", "-C", str(args.loopx_root), "rev-parse", "HEAD"], + capture_output=True, + text=True, + check=False, + ).stdout.strip() + if len(revision) != 40: + errors.append("loopx_revision_missing") + elif revision != EXPECTED_LOOPX_REVISION: + errors.append("loopx_revision_not_pinned") + + receipt = { + "schema_version": "deepswe_loopx_arm_admission_v1", + "generated_at": datetime.now(timezone.utc).isoformat(), + "arm": args.arm, + "host_surface": surface, + "runner": runner_name, + "runner_sha256": digest(runner) if runner.is_file() else None, + "delivery_gate_sha256": digest(delivery) if delivery.is_file() else None, + "benchmark_adapter_sha256": digest(adapter) if adapter.is_file() else None, + "codex_wrapper_sha256": digest(wrapper) if wrapper.is_file() else None, + "goal_harness_sha256": digest(harness) if harness.is_file() else None, + "loopx_revision": revision or None, + "loopx_source_clean": "loopx_source_not_clean" not in errors, + "model": args.model, + "effort": args.effort, + "goal_timeout_seconds": args.goal_timeout, + "heartbeat_segment_timeout_seconds": args.heartbeat_segment_timeout, + "turn_idle_timeout_seconds": args.turn_idle_timeout, + "agent_timeout_multiplier": args.agent_timeout_multiplier, + "minimum_native_agent_timeout_seconds": minimum_agent_timeout, + "task_count": len(tasks), + "task_ids": tasks, + "task_manifest_sha256": manifest_sha256, + "missing_or_malformed_tasks": malformed_tasks, + "gateway_port": args.port, + "gateway_port_free": "gateway_port_occupied" not in errors, + "network_policy": "model_only_dedicated_gateway", + "sandbox_policy": "danger_full_access_equivalent", + "web_search": "disabled", + "delivery_self_test": delivery_probe_ok, + "delivery_self_test_detail": delivery_probe_detail, + "lifecycle_authority": "loopx_goal_or_todo", + "task_correctness_authority": "independent_verifier", + "admitted": not errors, + "errors": errors, + } + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(receipt, indent=2, sort_keys=True) + "\n", encoding="utf-8") + print(json.dumps(receipt, sort_keys=True)) + return 0 if receipt["admitted"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/benchmark/deepswe-gptxhigh-v1/run_five_arms_remaining59_20260910.sh b/benchmark/deepswe-gptxhigh-v1/run_five_arms_remaining59_20260910.sh new file mode 100755 index 0000000000..5bcbfc19e6 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/run_five_arms_remaining59_20260910.sh @@ -0,0 +1,101 @@ +#!/usr/bin/env bash +# Full 5-arm launch over the remaining 59 tasks (113 total - frozen 54). +# 3 LoopX arms mirror run_loopx_rerun_54_20260908.sh; plain/goal mirror +# run_goal_plain_infra_complete.sh. Task set: remaining59.txt. +set -euo pipefail +DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +cd "$DIR" +PY="${MR_PYTHON:-python3}" +: "${MR_MODELONLY_COMPOSE:?set MR_MODELONLY_COMPOSE to the model-only compose overlay}" +[[ -f "$MR_MODELONLY_COMPOSE" ]] || { echo "missing model-only compose overlay" >&2; exit 2; } +MODEL="${MR_MODEL:-openai/gpt-5.6-sol}"; EFFORT="${MR_EFFORT:-xhigh}" +LOOPX_ROOT="${MR_LOOPX_ROOT:-$DIR/../loopx-official-latest}" +SHA="$(git -C "$LOOPX_ROOT" rev-parse --verify HEAD | cut -c1-12)" +OUT_ROOT="$DIR/jobs/loopx-five-arms-remaining59-20260910" +LOG_ROOT="$DIR/logs/loopx-five-arms-remaining59-20260910" +mkdir -p "$OUT_ROOT" "$LOG_ROOT" +GOAL_TIMEOUT=14400; HB_SEG=7200; TURN_IDLE=7200; MULT=3.0; LOOPX_CC=2 +PG_TIMEOUT=3600; PG_CC=5 +TASK_TEXT="$("$PY" - "${MR_TASK_LIST:-$DIR/remaining59.txt}" <<'PY' +import sys +from pathlib import Path +from preflight_loopx_rerun import load_task_list +print("\n".join(load_task_list(Path(sys.argv[1]), 59))) +PY +)" +mapfile -t TASKS <<< "$TASK_TEXT" +TASK_LIST="$(mktemp)" +printf '%s\n' "${TASKS[@]}" > "$TASK_LIST" +PIDS=() +cleanup() { + local pid + for pid in "${PIDS[@]}"; do kill "$pid" 2>/dev/null || true; done + rm -f "$TASK_LIST" +} +trap cleanup EXIT +trap 'exit 130' INT TERM +INC=(); for t in "${TASKS[@]}"; do INC+=(-i "$t"); done +echo "$(date -Is) launching 5 arms x ${#TASKS[@]} tasks -> $OUT_ROOT" + +preflight_loopx() { + local m="$1" port="$2" + echo "$(date -Is) preflight $m (port $port)" + "$PY" "$DIR/preflight_loopx_rerun.py" --arm "$m" --loopx-root "$LOOPX_ROOT" \ + --port "$port" --model "$MODEL" --effort "$EFFORT" --goal-timeout "$GOAL_TIMEOUT" \ + --heartbeat-segment-timeout "$HB_SEG" --turn-idle-timeout "$TURN_IDLE" \ + --agent-timeout-multiplier "$MULT" --output "$LOG_ROOT/admission-$m.json" \ + --task-list "$TASK_LIST" --expected-task-count 59 \ + >"$LOG_ROOT/preflight-$m.log" 2>&1 || { echo "$(date -Is) PREFLIGHT FAILED $m"; return 1; } +} +launch_loopx() { # arm_mode codex_arm port + local m="$1" arm="$2" port="$3" jobs="$OUT_ROOT/$1"; mkdir -p "$jobs" + echo "$(date -Is) start $m" + ( cd "$DIR" + PIER_CUSTOM_NETWORKS=1 MR_MODELONLY_NET=1 MR_MODELONLY_HOST=127.0.0.1 MR_MODELONLY_PORT="$port" \ + MR_API_BASE="http://127.0.0.1:$port/v1" MR_AGENT=codex MR_MODEL="$MODEL" MR_EFFORT="$EFFORT" MR_REASONING_EFFORT="$EFFORT" \ + MR_CODEX_ARM="$arm" MR_LOOPX_MODE="$m" MR_LOOPX_ROOT="$LOOPX_ROOT" \ + MR_LOOPX_PROFILE_ROOT="/tmp/loopx-profile-fair54-$SHA-$m" MR_LOOPX_PREFLIGHT=0 MR_LOOPX_WEN_COMPAT=1 \ + MR_GOAL_TIMEOUT_SEC="$GOAL_TIMEOUT" MR_UPSTREAM_TIMEOUT="$GOAL_TIMEOUT" \ + MR_HEARTBEAT_SEGMENT_TIMEOUT_SEC="$HB_SEG" MR_LOOPX_TURN_IDLE_TIMEOUT_SEC="$TURN_IDLE" \ + MR_RUN_LABEL="five-arms-rem59-$m" MR_GATEWAY_PORT="$port" MR_JOBS_DIR="$jobs" \ + MR_LOG_PATH="$LOG_ROOT/gateway_calls_$m.jsonl" MR_GATEWAY_OUT="$LOG_ROOT/gateway_$m.out" \ + MR_GATEWAY_MAX_RETRIES=12 MR_GATEWAY_RETRY_BASE_SEC=5 MR_GATEWAY_RETRY_CAP_SEC=60 \ + exec ./run.sh --all "${INC[@]}" -k 1 -n "$LOOPX_CC" --agent-timeout-multiplier "$MULT" + ) >"$LOG_ROOT/$m.log" 2>&1 & + PIDS+=("$!") + echo "$(date -Is) $m launched pid=${PIDS[-1]}" +} +launch_pg() { # mode port + local m="$1" port="$2" jobs="$OUT_ROOT/$1"; mkdir -p "$jobs" + echo "$(date -Is) start $m (gateway $port)" + ( cd "$DIR" + PIER_CUSTOM_NETWORKS=1 MR_MODELONLY_NET=1 MR_MODELONLY_HOST=127.0.0.1 MR_MODELONLY_PORT="$port" \ + MR_API_BASE="http://127.0.0.1:$port/v1" MR_AGENT=codex MR_MODEL="$MODEL" MR_EFFORT="$EFFORT" MR_REASONING_EFFORT="$EFFORT" \ + MR_CODEX_ARM="$m" MR_LOOPX_MODE=x MR_LOOPX_ROOT="$LOOPX_ROOT" MR_LOOPX_PREFLIGHT=0 MR_LOOPX_WEN_COMPAT=1 \ + MR_GOAL_TIMEOUT_SEC="$PG_TIMEOUT" MR_UPSTREAM_TIMEOUT="$PG_TIMEOUT" \ + MR_RUN_LABEL="five-arms-rem59-$m" MR_GATEWAY_PORT="$port" MR_JOBS_DIR="$jobs" \ + MR_LOG_PATH="$LOG_ROOT/gateway_calls_$m.jsonl" MR_GATEWAY_OUT="$LOG_ROOT/gateway_$m.out" \ + exec ./run.sh --all "${INC[@]}" -k 1 -n "$PG_CC" + ) >"$LOG_ROOT/$m.log" 2>&1 & + PIDS+=("$!") + echo "$(date -Is) $m launched pid=${PIDS[-1]}" +} +preflight_loopx ssh-goal 4411 +preflight_loopx codex-cli 4412 +preflight_loopx heartbeat 4413 +launch_loopx ssh-goal loopx-native 4411 +launch_loopx codex-cli loopx-native-codex-cli 4412 +launch_loopx heartbeat loopx-native-heartbeat 4413 +launch_pg goal 4393 +launch_pg plain 4394 +echo "$(date -Is) all 5 arms launched; logs -> $LOG_ROOT" +status=0 +for pid in "${PIDS[@]}"; do + if wait "$pid"; then :; else status=1; fi +done +PIDS=() +if (( status )); then + echo "$(date -Is) one or more arms failed" >&2 + exit "$status" +fi +echo "$(date -Is) all arms finished" diff --git a/benchmark/deepswe-gptxhigh-v1/run_loopx_rerun_54_20260908.sh b/benchmark/deepswe-gptxhigh-v1/run_loopx_rerun_54_20260908.sh new file mode 100755 index 0000000000..f2af172e26 --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/run_loopx_rerun_54_20260908.sh @@ -0,0 +1,194 @@ +#!/usr/bin/env bash +# Fair rerun of the three LoopX treatments after host/delivery fixes. +set -euo pipefail +trap 'exit 130' INT TERM + +DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +cd "$DIR" +PY="${MR_PYTHON:-python3}" +: "${MR_MODELONLY_COMPOSE:?set MR_MODELONLY_COMPOSE to the model-only compose overlay}" +[[ -f "$MR_MODELONLY_COMPOSE" ]] || { echo "missing model-only compose overlay" >&2; exit 2; } +MODE="${1:-all}" +shift $(( $# > 0 ? 1 : 0 )) + +case "$MODE" in + all|ssh-goal|codex-cli|heartbeat) ;; + *) echo "usage: $0 [all|ssh-goal|codex-cli|heartbeat] [task ...]" >&2; exit 2 ;; +esac + +MODEL="${MR_MODEL:-openai/gpt-5.6-sol}" +EFFORT="${MR_EFFORT:-xhigh}" +CONCURRENT="${MR_CONCURRENT:-2}" +GOAL_TIMEOUT="${MR_GOAL_TIMEOUT_SEC:-14400}" +HEARTBEAT_SEGMENT_TIMEOUT="${MR_HEARTBEAT_SEGMENT_TIMEOUT_SEC:-7200}" +TURN_IDLE_TIMEOUT="${MR_LOOPX_TURN_IDLE_TIMEOUT_SEC:-7200}" +AGENT_TIMEOUT_MULTIPLIER="${MR_AGENT_TIMEOUT_MULTIPLIER:-3.0}" +PORT_BASE="${MR_GATEWAY_BASE_PORT:-4411}" +OUT_ROOT="${MR_LOOPX_RERUN_ROOT:-$DIR/jobs/loopx-three-arms-54-fair-rerun-20260908}" +LOG_ROOT="${MR_LOOPX_RERUN_LOG_ROOT:-$DIR/logs/loopx-three-arms-54-fair-rerun-20260908}" +LOOPX_ROOT="${MR_LOOPX_ROOT:-$DIR/../loopx-official-latest}" +LOOPX_REVISION="$(git -C "$LOOPX_ROOT" rev-parse --verify HEAD)" +LOOPX_REVISION_SHORT="${LOOPX_REVISION:0:12}" + +ALL_TASK_TEXT="$( + cd "$DIR" + "$PY" - <<'PY' +from goal30_subset import SUBSET +from hard24_subset import HARD_SUBSET +from remaining4_subset import REMAINING_SUBSET + +tasks = list(dict.fromkeys([*SUBSET, *HARD_SUBSET, *REMAINING_SUBSET])) +if len(tasks) != 54: + raise SystemExit(f"expected 54 tasks, got {len(tasks)}") +print("\n".join(tasks)) +PY +)" +mapfile -t ALL_TASKS <<< "$ALL_TASK_TEXT" +(( ${#ALL_TASKS[@]} == 54 )) || { echo "expected 54 tasks" >&2; exit 2; } + +if (( $# )); then + TASKS=("$@") + for requested in "${TASKS[@]}"; do + found=0 + for allowed in "${ALL_TASKS[@]}"; do + [[ "$requested" == "$allowed" ]] && found=1 && break + done + (( found == 1 )) || { echo "task is outside the frozen 54: $requested" >&2; exit 2; } + done +else + TASKS=("${ALL_TASKS[@]}") +fi + +TASK_LIST="$(mktemp)" +trap 'rm -f "$TASK_LIST"' EXIT +printf '%s\n' "${TASKS[@]}" > "$TASK_LIST" +"$PY" - "$TASK_LIST" "${#TASKS[@]}" <<'PY' +import sys +from pathlib import Path +from preflight_loopx_rerun import load_task_list +load_task_list(Path(sys.argv[1]), int(sys.argv[2])) +PY + +port_for() { + case "$1" in + ssh-goal) echo "$PORT_BASE" ;; + codex-cli) echo "$((PORT_BASE + 1))" ;; + heartbeat) echo "$((PORT_BASE + 2))" ;; + esac +} + +arm_for() { + case "$1" in + ssh-goal) echo loopx-native ;; + codex-cli) echo loopx-native-codex-cli ;; + heartbeat) echo loopx-native-heartbeat ;; + esac +} + +preflight_arm() { + local arm_mode="$1" + local port jobs + port="$(port_for "$arm_mode")" + jobs="$OUT_ROOT/$arm_mode" + [[ "$port" != 4141 && "$port" != 4250 ]] || { + echo "prohibited DeepSWE gateway port: $port" >&2 + return 2 + } + if ss -ltnH "sport = :$port" | grep -q .; then + echo "gateway port already occupied: $port" >&2 + return 1 + fi + if pgrep -af "[p]ier_cn.py.*--jobs-dir $jobs" >/dev/null; then + echo "$arm_mode already running: $jobs" >&2 + return 1 + fi + + mkdir -p "$jobs" "$LOG_ROOT" + "$PY" "$DIR/preflight_loopx_rerun.py" \ + --arm "$arm_mode" --loopx-root "$LOOPX_ROOT" --port "$port" \ + --model "$MODEL" --effort "$EFFORT" --goal-timeout "$GOAL_TIMEOUT" \ + --heartbeat-segment-timeout "$HEARTBEAT_SEGMENT_TIMEOUT" \ + --turn-idle-timeout "$TURN_IDLE_TIMEOUT" \ + --agent-timeout-multiplier "$AGENT_TIMEOUT_MULTIPLIER" \ + --task-list "$TASK_LIST" --expected-task-count "${#TASKS[@]}" \ + --output "$LOG_ROOT/admission-$arm_mode.json" +} + +run_arm() { + local arm_mode="$1" + local port arm jobs label + port="$(port_for "$arm_mode")" + arm="$(arm_for "$arm_mode")" + jobs="$OUT_ROOT/$arm_mode" + label="loopx-fair54-20260908-$arm_mode" + local -a include=() + local task + for task in "${TASKS[@]}"; do include+=(-i "$task"); done + echo "$(date -Is) start $arm_mode tasks=${#TASKS[@]} concurrency=$CONCURRENT port=$port" + ( + cd "$DIR" + PIER_CUSTOM_NETWORKS=1 \ + MR_MODELONLY_NET=1 MR_MODELONLY_HOST=127.0.0.1 MR_MODELONLY_PORT="$port" \ + MR_API_BASE="http://127.0.0.1:$port/v1" \ + MR_AGENT=codex MR_MODEL="$MODEL" MR_EFFORT="$EFFORT" MR_REASONING_EFFORT="$EFFORT" \ + MR_CODEX_ARM="$arm" MR_LOOPX_MODE="$arm_mode" MR_LOOPX_ROOT="$LOOPX_ROOT" \ + MR_LOOPX_PROFILE_ROOT="/tmp/loopx-profile-fair54-$LOOPX_REVISION_SHORT-$arm_mode" \ + MR_LOOPX_PREFLIGHT=0 MR_LOOPX_WEN_COMPAT=1 \ + MR_GOAL_TIMEOUT_SEC="$GOAL_TIMEOUT" MR_UPSTREAM_TIMEOUT="$GOAL_TIMEOUT" \ + MR_HEARTBEAT_SEGMENT_TIMEOUT_SEC="$HEARTBEAT_SEGMENT_TIMEOUT" \ + MR_LOOPX_TURN_IDLE_TIMEOUT_SEC="$TURN_IDLE_TIMEOUT" \ + MR_RUN_LABEL="$label" MR_GATEWAY_PORT="$port" MR_JOBS_DIR="$jobs" \ + MR_LOG_PATH="$LOG_ROOT/gateway_calls_$arm_mode.jsonl" \ + MR_GATEWAY_OUT="$LOG_ROOT/gateway_$arm_mode.out" \ + MR_GATEWAY_MAX_RETRIES=12 MR_GATEWAY_RETRY_BASE_SEC=5 MR_GATEWAY_RETRY_CAP_SEC=60 \ + ./run.sh --all "${include[@]}" -k 1 -n "$CONCURRENT" \ + --agent-timeout-multiplier "$AGENT_TIMEOUT_MULTIPLIER" + ) >"$LOG_ROOT/$arm_mode.log" 2>&1 + echo "$(date -Is) finish $arm_mode" +} + +if [[ "$MODE" == all ]]; then + for arm_mode in ssh-goal codex-cli heartbeat; do + preflight_arm "$arm_mode" + done + "$PY" - "$LOG_ROOT" <<'PY' +import json +import sys +from pathlib import Path + +root = Path(sys.argv[1]) +receipts = [json.loads((root / f"admission-{arm}.json").read_text()) for arm in ("ssh-goal", "codex-cli", "heartbeat")] +fixed = ( + "model", + "effort", + "agent_timeout_multiplier", + "goal_timeout_seconds", + "heartbeat_segment_timeout_seconds", + "turn_idle_timeout_seconds", + "task_count", + "task_manifest_sha256", + "loopx_revision", + "network_policy", + "sandbox_policy", + "web_search", + "task_correctness_authority", +) +for key in fixed: + values = {json.dumps(receipt.get(key), sort_keys=True) for receipt in receipts} + if len(values) != 1: + raise SystemExit(f"cross-arm admission mismatch for {key}: {values}") +surfaces = {receipt["host_surface"] for receipt in receipts} +if len(surfaces) != 3: + raise SystemExit(f"host surfaces are not distinct: {surfaces}") +if not all(receipt.get("admitted") for receipt in receipts): + raise SystemExit("at least one arm was not admitted") +print("cross-arm admission: matched controls, three distinct execution surfaces") +PY + for arm_mode in ssh-goal codex-cli heartbeat; do + run_arm "$arm_mode" + done + exit 0 +fi + +preflight_arm "$MODE" +run_arm "$MODE" diff --git a/benchmark/deepswe-gptxhigh-v1/tests/test_execution_contracts.py b/benchmark/deepswe-gptxhigh-v1/tests/test_execution_contracts.py new file mode 100644 index 0000000000..2844f92cbf --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/tests/test_execution_contracts.py @@ -0,0 +1,329 @@ +"""Offline regression tests for experiment admission and launcher failures.""" + +import importlib.util +import json +import os +from pathlib import Path +import shutil +import subprocess +import sys +import threading +import time +from concurrent.futures import ThreadPoolExecutor +from types import ModuleType, SimpleNamespace + +import pytest + + +ROOT = Path(__file__).resolve().parents[1] + + +def load_module(name, path): + spec = importlib.util.spec_from_file_location(name, path) + module = importlib.util.module_from_spec(spec) + sys.modules[name] = module + spec.loader.exec_module(module) + return module + + +@pytest.fixture +def admission(monkeypatch): + monkeypatch.syspath_prepend(str(ROOT)) + return load_module("study_admission", ROOT / "preflight_loopx_rerun.py") + + +@pytest.mark.parametrize("contents,count", [("", 59), ("a\na\n", 2), ("../a\n", 1), ("a\n\n", 2), ("a\n", 59)]) +def test_reject_invalid_task_lists(admission, tmp_path, contents, count): + manifest = tmp_path / "tasks.txt" + manifest.write_text(contents) + with pytest.raises(ValueError): + admission.load_task_list(manifest, count) + + +def test_manifest_covers_actual_selected_tasks(admission, tmp_path): + tasks = tuple(f"task-{i}" for i in range(59)) + for task in tasks: + directory = tmp_path / task + directory.mkdir() + (directory / "task.toml").write_text( + '[metadata]\nbase_commit_hash = "abcdef1"\n[agent]\ntimeout_sec = 6000\n' + ) + digest, timeout, missing = admission.task_manifest(tmp_path, tasks) + assert timeout == 6000 and not missing + assert digest != admission.task_manifest(tmp_path, tasks[:54])[0] + (tmp_path / tasks[-1] / "task.toml").write_text("invalid toml = [") + assert admission.task_manifest(tmp_path, tasks)[2] == [tasks[-1]] + + +def test_revision_cannot_be_overridden(monkeypatch, admission): + monkeypatch.setenv("MR_EXPECTED_LOOPX_REVISION", "f" * 40) + reloaded = load_module("study_pinned_admission", ROOT / "preflight_loopx_rerun.py") + assert reloaded.EXPECTED_LOOPX_REVISION == "2cef51d08b2a0103f4ba026bf47fd70dc8acee30" + + +@pytest.fixture +def launcher(tmp_path): + for name in ( + "run_five_arms_remaining59_20260910.sh", "run_loopx_rerun_54_20260908.sh", + "preflight_loopx_rerun.py", "workspace_delivery.py", + ): + shutil.copy2(ROOT / name, tmp_path / name) + (tmp_path / "remaining59.txt").write_text("".join(f"task-{i}\n" for i in range(59))) + (tmp_path / "overlay.yaml").write_text("services: {}\n") + subprocess.run(["git", "init", "-q", str(tmp_path / "repo")], check=True) + subprocess.run( + ["git", "-C", str(tmp_path / "repo"), "-c", "user.name=Test", "-c", + "user.email=test@example.invalid", "commit", "--allow-empty", "-qm", "base"], + check=True, + ) + python = tmp_path / "python" + python.write_text( + f"#!{sys.executable}\n" + "import json, os, pathlib, sys\n" + "if sys.argv[1].endswith('preflight_loopx_rerun.py'):\n" + " args = sys.argv[2:]\n" + " arm = args[args.index('--arm') + 1]\n" + " tasks = pathlib.Path(args[args.index('--task-list') + 1]).read_text().splitlines()\n" + " assert len(tasks) == int(args[args.index('--expected-task-count') + 1])\n" + " with open(os.environ['ADMISSION_LOG'], 'a') as log:\n" + " log.write(json.dumps({'arm': arm, 'tasks': tasks}) + '\\n')\n" + " sys.exit(7 if os.environ.get('FAIL_ADMISSION') == arm else 0)\n" + "os.execv(sys.executable, [sys.executable, *sys.argv[1:]])\n" + ) + python.chmod(0o755) + runner = tmp_path / "run.sh" + runner.write_text( + '#!/usr/bin/env bash\n' + 'printf "%s %s %s\\n" "$MR_CODEX_ARM" "$MR_GATEWAY_PORT" "$MR_API_BASE" >> "$LAUNCH_LOG"\n' + '[[ "$MR_CODEX_ARM" != "${FAIL_ARM:-}" ]]\n' + ) + runner.chmod(0o755) + env = { + **os.environ, "MR_PYTHON": str(python), "MR_LOOPX_ROOT": str(tmp_path / "repo"), + "MR_MODELONLY_COMPOSE": str(tmp_path / "overlay.yaml"), + "LAUNCH_LOG": str(tmp_path / "launched"), "ADMISSION_LOG": str(tmp_path / "admitted"), + } + for key in ("MR_TASK_LIST", "FAIL_ARM", "FAIL_ADMISSION"): + env.pop(key, None) + return tmp_path, env + + +def run_launcher(launcher, name="run_five_arms_remaining59_20260910.sh", *args): + directory, env = launcher + return subprocess.run(["bash", str(directory / name), *args], env=env, capture_output=True, text=True, timeout=20) + + +@pytest.mark.parametrize("contents", [None, "", "task-0\n" * 59]) +def test_no_launch_with_missing_empty_or_duplicate_tasks(launcher, contents): + directory, _ = launcher + tasks = directory / "remaining59.txt" + tasks.unlink() if contents is None else tasks.write_text(contents) + assert run_launcher(launcher).returncode != 0 + assert not (directory / "launched").exists() + + +def test_admission_failure_stops_all_arms(launcher): + directory, env = launcher + env["FAIL_ADMISSION"] = "codex-cli" + assert run_launcher(launcher).returncode != 0 + assert not (directory / "launched").exists() + + +@pytest.mark.parametrize("failed_arm", ["", "plain", "loopx-native-heartbeat"]) +def test_waits_for_each_arm_and_uses_selected_tasks_and_ports(launcher, failed_arm): + directory, env = launcher + env["FAIL_ARM"] = failed_arm + result = run_launcher(launcher) + assert (result.returncode == 0) == (not failed_arm), result.stderr + receipts = [json.loads(line) for line in (directory / "admitted").read_text().splitlines()] + assert len(receipts) == 3 + assert all(row["tasks"] == [f"task-{i}" for i in range(59)] for row in receipts) + launches = (directory / "launched").read_text().splitlines() + assert len(launches) == 5 + assert "plain 4394 http://127.0.0.1:4394/v1" in launches + assert "goal 4393 http://127.0.0.1:4393/v1" in launches + assert ("all arms finished" in result.stdout) == (not failed_arm) + + +def test_missing_54_subset_modules_fail_before_launch(launcher): + directory, _ = launcher + result = run_launcher(launcher, "run_loopx_rerun_54_20260908.sh") + assert result.returncode != 0 + assert not (directory / "launched").exists() + + +@pytest.mark.parametrize("selected", [("task-4", "task-8"), ("task-4", "task-4"), ("outside",)]) +def test_54_launcher_admits_exact_selection(launcher, selected): + directory, env = launcher + (directory / "goal30_subset.py").write_text(f"SUBSET = {[f'task-{i}' for i in range(30)]!r}\n") + (directory / "hard24_subset.py").write_text(f"HARD_SUBSET = {[f'task-{i}' for i in range(30, 54)]!r}\n") + (directory / "remaining4_subset.py").write_text("REMAINING_SUBSET = []\n") + for command, code in (("ss", 0), ("pgrep", 1)): + executable = directory / command + executable.write_text(f"#!/bin/sh\nexit {code}\n") + executable.chmod(0o755) + env["PATH"] = str(directory) + os.pathsep + env["PATH"] + result = run_launcher(launcher, "run_loopx_rerun_54_20260908.sh", "heartbeat", *selected) + if selected == ("task-4", "task-8"): + assert result.returncode == 0, result.stderr + receipts = [json.loads(line) for line in (directory / "admitted").read_text().splitlines()] + assert receipts == [{"arm": "heartbeat", "tasks": list(selected)}] + else: + assert result.returncode != 0 + assert not (directory / "launched").exists() + + +def test_profile_initialization_is_serialized(monkeypatch, tmp_path): + goal = ModuleType("goal_codex") + goal.GoalCodex = object + goal._CODEX_EXEC_MARKER = "codex exec " + goal._REMOTE_DIR = "/tmp/test-goal" + monkeypatch.setitem(sys.modules, "goal_codex", goal) + api = ModuleType("loopx.capabilities.benchmark_toolkit.native_codex_profile") + installs = [] + + def install(source, target): + installs.append(target) + target.mkdir() + (target / "partial").touch() + time.sleep(0.05) + (target / "ready").touch() + return {"ready": True} + + def inspect(target, **kwargs): + assert (target / "ready").is_file(), "observed partially installed profile" + return {"ready": True} + + api.install_native_codex_profile = install + api.inspect_native_codex_profile = inspect + api.compact_native_codex_profile_receipt = lambda value: value + monkeypatch.setitem(sys.modules, api.__name__, api) + native = load_module("study_native_profile", ROOT / "loopx_native_codex.py") + barrier = threading.Barrier(2) + + def build(): + barrier.wait() + return native.build_host_profile(str(tmp_path), str(tmp_path / "profile")) + + with ThreadPoolExecutor(max_workers=2) as pool: + results = list(pool.map(lambda _: build(), range(2))) + assert results == [{"ready": True}, {"ready": True}] + assert len(installs) == 1 + + +@pytest.mark.parametrize("execute", [False, True]) +def test_plain_transport_never_attaches_goal(monkeypatch, execute): + native = load_module("native_codex_goal", ROOT.parents[1] / "loopx/capabilities/benchmark_toolkit/native_codex_goal.py") + monkeypatch.setitem(sys.modules, "native_codex_goal", native) + plain = load_module("study_plain", ROOT / "plain_appserver_runner.py") + + class Transport: + def __init__(self): + self.calls = [] + + def request(self, method, payload): + self.calls.append((method, payload)) + return {"thread": {"id": "thread"}, "turn": {"id": "turn"}} + + def notify(self, method, payload): + self.calls.append((method, payload)) + + transport = Transport() + config = native.NativeGoalConfig(cwd="/tmp", objective="unused", task_instruction="task", effort="xhigh") + plain.start_turn_without_goal(transport, config, execute=execute) + methods = [method for method, _ in transport.calls] + assert methods == ["initialize", "initialized", "thread/start"] + (["turn/start"] if execute else []) + if execute: + assert transport.calls[-1][1]["effort"] == "xhigh" + + +@pytest.mark.parametrize("key,value", [("MR_CLAUDE_ARM", "loopx"), ("MR_CODEX_ARM", "loopx")]) +def test_unsupported_legacy_arms_fail_before_importing_pier(monkeypatch, key, value): + module = load_module("study_pier", ROOT / "pier_cn.py") + monkeypatch.setenv(key, value) + with pytest.raises(SystemExit): + (module.patch_claude_arm if key == "MR_CLAUDE_ARM" else module.patch_goal_mode)() + + +def test_terminal_todo_without_delivery_stops_immediately(monkeypatch): + monkeypatch.syspath_prepend(str(ROOT)) + runner = load_module("study_cli_runner", ROOT / "loopx_codex_cli_runner.py") + claims = [] + monkeypatch.setattr(runner, "head_sha", lambda _: "base") + monkeypatch.setattr(runner, "normalize_delivery", lambda *_: {"treatment_valid": False}) + monkeypatch.setattr(runner, "write_receipt", lambda *_: None) + + def claim(*args): + claims.append(True) + return {"terminal": True, "ok": True} + + monkeypatch.setattr(runner, "_claim_primary_p0", claim) + result, code = runner.run(SimpleNamespace( + cwd="/tmp", goal_timeout_seconds=60, preflight_only=False, max_segments=5, + model="test", effort="xhigh", + )) + assert code != 0 and result["treatment_valid"] is False + assert result["segments"][0]["status"] == "terminal_todo_without_valid_delivery" + assert len(claims) == 1 + + +@pytest.mark.parametrize("error,retryable", [ + ({"status": 429, "message": "busy"}, True), + ({"status": 503, "message": "temporarily unavailable"}, True), + ({"status": 401, "message": "server error timeout"}, False), + ({"code": "insufficient_quota", "status": 429}, False), + ({"codexErrorInfo": {"httpConnectionFailed": {"httpStatusCode": 403}}}, False), + ({"code": "request_timeout"}, True), + ("unknown timeout text", False), +]) +def test_retry_classification_uses_structured_error(monkeypatch, error, retryable): + monkeypatch.syspath_prepend(str(ROOT)) + native = load_module("native_codex_goal", ROOT.parents[1] / "loopx/capabilities/benchmark_toolkit/native_codex_goal.py") + monkeypatch.setitem(sys.modules, "native_codex_goal", native) + runner = load_module("study_native_runner", ROOT / "loopx_wen_native_runner.py") + event = {"method": "turn/completed", "params": {"turn": {"error": error}}} + assert runner._terminal_error(event).retryable is retryable + + +def test_unmatched_install_step_fails_closed(monkeypatch): + pier = load_module("study_mirrors", ROOT / "pier_cn.py") + + class Agent: + def install_spec(self): + return SimpleNamespace(steps=[SimpleNamespace(run="unknown installer", env={})]) + + monkeypatch.setattr(pier, "HARNESSES", (("fake", "Agent"),)) + monkeypatch.setattr(pier.importlib, "import_module", lambda _: SimpleNamespace(Agent=Agent)) + pier.patch_harnesses() + with pytest.raises(RuntimeError, match="refusing an unpatched harness"): + Agent().install_spec() + + +def test_external_network_overlay_is_required_and_used(monkeypatch, tmp_path): + pier = load_module("study_network", ROOT / "pier_cn.py") + + class Docker: + environment_dir = tmp_path / "environment" + _docker_compose_paths = property(lambda self: []) + + def agent_process_env(self, env): + return env + + module = ModuleType("pier.environments.docker.docker") + module.DockerEnvironment = Docker + monkeypatch.setitem(sys.modules, module.__name__, module) + monkeypatch.setenv("MR_MODELONLY_NET", "1") + monkeypatch.delenv("MR_LOCAL_AGENT_BASE_IMAGE", raising=False) + monkeypatch.delenv("MR_MODELONLY_COMPOSE", raising=False) + with pytest.raises(SystemExit, match="MR_MODELONLY_COMPOSE"): + pier.patch_modelonly_network() + overlay = tmp_path / "overlay.yaml" + monkeypatch.setenv("MR_MODELONLY_COMPOSE", str(overlay)) + with pytest.raises(SystemExit, match="missing model-only"): + pier.patch_modelonly_network() + overlay.write_text("services: {}\n") + pier.patch_modelonly_network() + assert Docker()._docker_compose_paths == [overlay] + verifier = Docker() + verifier.environment_dir = tmp_path / "tests" + assert verifier._docker_compose_paths == [] diff --git a/benchmark/deepswe-gptxhigh-v1/workspace_delivery.py b/benchmark/deepswe-gptxhigh-v1/workspace_delivery.py new file mode 100755 index 0000000000..d6f539b1db --- /dev/null +++ b/benchmark/deepswe-gptxhigh-v1/workspace_delivery.py @@ -0,0 +1,279 @@ +#!/usr/bin/env python3 +"""Normalize an agent's work into the canonical DeepSWE checkout. + +DeepSWE grades ``git diff HEAD`` from ``/app``. LoopX may legitimately +place an agent in a linked worktree, so a commit can exist while the collector +still sees an empty patch. This module makes that delivery boundary explicit: +it finds the one changed worktree, commits any remaining agent changes, copies +the resulting patch into the canonical checkout, and proves that the patch +applies to the task base before the verifier is allowed to run. + +LoopX control state is runner-owned and is never included in the submission. +""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import os +import subprocess +import tempfile +from dataclasses import dataclass +from pathlib import Path +from typing import Any + + +CONTROL_PATHS = (".loopx", ".codex", ".worktrees") + + +class DeliveryError(RuntimeError): + pass + + +@dataclass(frozen=True) +class Candidate: + path: Path + head: str + patch: bytes + + @property + def patch_sha256(self) -> str: + return hashlib.sha256(self.patch).hexdigest() + + +def _run( + argv: list[str], + *, + cwd: Path | None = None, + input_bytes: bytes | None = None, + check: bool = True, +) -> subprocess.CompletedProcess[bytes]: + result = subprocess.run( + argv, + cwd=cwd, + input=input_bytes, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + check=False, + ) + if check and result.returncode: + detail = result.stderr.decode("utf-8", "replace")[-600:] + raise DeliveryError(f"command failed ({result.returncode}): {argv!r}: {detail}") + return result + + +def head_sha(project: Path) -> str: + return _run(["git", "-C", str(project), "rev-parse", "HEAD"]).stdout.decode().strip() + + +def _exclude_control_state(project: Path) -> None: + git_dir = _run( + ["git", "-C", str(project), "rev-parse", "--git-common-dir"] + ).stdout.decode().strip() + common = Path(git_dir) + if not common.is_absolute(): + common = (project / common).resolve() + exclude = common / "info" / "exclude" + exclude.parent.mkdir(parents=True, exist_ok=True) + existing = exclude.read_text(encoding="utf-8") if exclude.exists() else "" + additions = [f"/{name}/" for name in CONTROL_PATHS if f"/{name}/" not in existing] + if additions: + with exclude.open("a", encoding="utf-8") as handle: + handle.write("\n# DeepSWE runner control state\n") + handle.write("\n".join(additions) + "\n") + + +def worktrees(project: Path) -> list[Path]: + output = _run( + ["git", "-C", str(project), "worktree", "list", "--porcelain"] + ).stdout.decode("utf-8", "replace") + paths = [] + for line in output.splitlines(): + if line.startswith("worktree "): + paths.append(Path(line.removeprefix("worktree ")).resolve()) + canonical = project.resolve() + return sorted(set(paths), key=lambda path: (path != canonical, str(path))) + + +def _commit_pending(path: Path) -> None: + _exclude_control_state(path) + # The runner state is ignored through the repository's shared exclude file. + # Passing ignored paths as explicit negative pathspecs makes Git reject the + # entire add operation in some task repositories. + _run(["git", "-C", str(path), "add", "-A", "--", "."]) + staged = _run( + ["git", "-C", str(path), "diff", "--cached", "--quiet"], check=False + ) + if staged.returncode not in (0, 1): + raise DeliveryError(f"could not inspect staged changes in {path}") + if staged.returncode == 1: + env = os.environ.copy() + env.setdefault("GIT_AUTHOR_NAME", "DeepSWE delivery runner") + env.setdefault("GIT_AUTHOR_EMAIL", "runner@deepswe.invalid") + env.setdefault("GIT_COMMITTER_NAME", env["GIT_AUTHOR_NAME"]) + env.setdefault("GIT_COMMITTER_EMAIL", env["GIT_AUTHOR_EMAIL"]) + result = subprocess.run( + ["git", "-C", str(path), "commit", "-m", "chore: deliver agent work"], + env=env, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + check=False, + ) + if result.returncode: + raise DeliveryError( + "could not commit pending agent changes: " + + result.stderr.decode("utf-8", "replace")[-600:] + ) + + +def _candidate(path: Path, base_sha: str) -> Candidate | None: + _commit_pending(path) + head = head_sha(path) + patch = _run( + [ + "git", + "-C", + str(path), + "diff", + "--binary", + base_sha, + "HEAD", + "--", + ".", + *(f":(exclude){name}" for name in CONTROL_PATHS), + ] + ).stdout + return Candidate(path=path, head=head, patch=patch) if patch.strip() else None + + +def _prove_applies(project: Path, base_sha: str, patch: bytes) -> None: + with tempfile.TemporaryDirectory(prefix="deepswe-delivery-check-") as tmp: + checkout = Path(tmp) / "base" + _run(["git", "clone", "--shared", "--no-checkout", str(project), str(checkout)]) + _run(["git", "-C", str(checkout), "checkout", "--detach", base_sha]) + _run(["git", "-C", str(checkout), "apply", "--check", "--binary", "-"], input_bytes=patch) + + +def _install_patch(project: Path, base_sha: str, patch: bytes) -> str: + _run(["git", "-C", str(project), "reset", "--hard", base_sha]) + _run(["git", "-C", str(project), "apply", "--index", "--binary", "-"], input_bytes=patch) + env = os.environ.copy() + env.setdefault("GIT_AUTHOR_NAME", "DeepSWE delivery runner") + env.setdefault("GIT_AUTHOR_EMAIL", "runner@deepswe.invalid") + env.setdefault("GIT_COMMITTER_NAME", env["GIT_AUTHOR_NAME"]) + env.setdefault("GIT_COMMITTER_EMAIL", env["GIT_AUTHOR_EMAIL"]) + result = subprocess.run( + ["git", "-C", str(project), "commit", "-m", "chore: recover agent worktree delivery"], + env=env, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + check=False, + ) + if result.returncode: + raise DeliveryError( + "could not install recovered patch: " + + result.stderr.decode("utf-8", "replace")[-600:] + ) + return head_sha(project) + + +def normalize_delivery(project: Path, base_sha: str) -> dict[str, Any]: + project = project.resolve() + receipt: dict[str, Any] = { + "schema_version": "deepswe_delivery_receipt_v1", + "project": str(project), + "base_sha": base_sha, + "status": "error", + "treatment_valid": False, + } + try: + candidates = [ + item + for path in worktrees(project) + if (item := _candidate(path, base_sha)) is not None + ] + by_patch: dict[str, list[Candidate]] = {} + for item in candidates: + by_patch.setdefault(item.patch_sha256, []).append(item) + receipt["changed_worktree_count"] = len(candidates) + receipt["distinct_patch_count"] = len(by_patch) + receipt["changed_worktrees"] = [str(item.path) for item in candidates] + if not candidates: + receipt["status"] = "empty" + receipt["reason"] = "no_agent_delta_from_task_base" + return receipt + if len(by_patch) != 1: + receipt["status"] = "ambiguous" + receipt["reason"] = "multiple_distinct_agent_patches" + return receipt + + selected = next(iter(by_patch.values()))[0] + _prove_applies(project, base_sha, selected.patch) + canonical = project.resolve() + recovered = selected.path.resolve() != canonical + if recovered: + final_head = _install_patch(project, base_sha, selected.patch) + else: + final_head = selected.head + final_patch = _run( + [ + "git", + "-C", + str(project), + "diff", + "--binary", + base_sha, + "HEAD", + "--", + ".", + *(f":(exclude){name}" for name in CONTROL_PATHS), + ] + ).stdout + if not final_patch.strip(): + raise DeliveryError("canonical patch became empty after normalization") + if hashlib.sha256(final_patch).hexdigest() != selected.patch_sha256: + raise DeliveryError("canonical patch digest differs after normalization") + _prove_applies(project, base_sha, final_patch) + + receipt.update( + { + "status": "valid", + "reason": "canonical_patch_verified", + "treatment_valid": True, + "source_worktree": str(selected.path), + "recovered_from_linked_worktree": recovered, + "final_head": final_head, + "patch_bytes": len(final_patch), + "patch_sha256": selected.patch_sha256, + "patch_applies": True, + } + ) + except Exception as exc: + receipt["status"] = "error" + receipt["reason"] = f"{type(exc).__name__}: {exc}" + return receipt + + +def write_receipt(path: Path, receipt: dict[str, Any]) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(receipt, indent=2, sort_keys=True) + "\n", encoding="utf-8") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--project", type=Path, required=True) + parser.add_argument("--base-sha", required=True) + parser.add_argument("--receipt", type=Path) + parser.add_argument("--require-valid", action="store_true") + args = parser.parse_args() + + receipt = normalize_delivery(args.project, args.base_sha) + if args.receipt: + write_receipt(args.receipt, receipt) + print(json.dumps(receipt, sort_keys=True)) + return 0 if receipt["treatment_valid"] or not args.require_valid else 12 + + +if __name__ == "__main__": + raise SystemExit(main())