Skip to content

feat(routing): route Codex "Approve for me" reviews and add codex_approve_all - #872

Open
elyasmnvidian wants to merge 2 commits into
grclark/mac-deamonfrom
emehtabuddin/switch-1630-codex-approve-for-me-switchyard
Open

elyasmnvidian wants to merge 2 commits into
grclark/mac-deamonfrom
emehtabuddin/switch-1630-codex-approve-for-me-switchyard

Conversation

@elyasmnvidian

@elyasmnvidian elyasmnvidian commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

What

When "Approve for me" is on and Codex runs through Switchyard, Codex declines every action that needs approval. Codex asks a reviewer model to approve each of these actions. Unless Codex is logged in with an OpenAI API key, it sends the review request to the model ID codex-auto-review, even when it has no OpenAI login at all. Before this PR, neither examples/run_codex.sh nor the config that scripts/macos/install.sh writes had a route with that ID. So the server returned 404 model_not_found, and Codex treated the failed review as a denial.

This PR adds a codex-auto-review route to both configs and adds a route type, codex_approve_all. The route's config decides what answers the review requests:

  • Codex's own reviewer. A passthrough route sends the review request unchanged to codex-auto-review on the ChatGPT backend. The chatgpt_backend client sets forward_auth = true, so Switchyard forwards the caller's login. Both configs use this setup.
  • Another model. Point the route's target at any other target. Codex reads the reviewer's final message as a JSON verdict, so the model must follow a JSON output schema.
  • Approve everything. The new route type codex_approve_all answers every review request with Codex's allow verdict and calls no model. Codex then runs every action that needs approval, including commands that ask to run outside the sandbox.
[llm_clients.chatgpt_backend]
format = "openai_responses"
base_url = "https://chatgpt.com/backend-api/codex"
forward_auth = true

[targets.reviewer]
id = "codex-auto-review"
llm_client = "chatgpt_backend"

[routes.codex_auto_review]
id = "codex-auto-review"
type = "passthrough"        # or: type = "codex_approve_all", without `target`
target = "reviewer"

A config without a codex-auto-review route behaves as before: the server returns 404, and Codex declines the action. codex_approve_all accepts only the keys that every route takes (id, type, context_window, tool_calling, reasoning, vision), so the server refuses to start if target is left in:

unknown field `target`, there are no fields

codex_approve_all reuses the reply builder from noop, which this PR moves to fixed_reply in crates/libsy/src/algorithms/util.rs. The reply is streamed or buffered to match the request, and the route keeps no state.

Menu bar prices. The menu bar hides a period's dollar figures when any model seen in that period has no price. In the installer's setup, the routing log records each review under codex-auto-review, so the installer now prices that model ID at the baseline_model rates. The dollars saved stay the same, and the percentage drops a little (details below). The installer never overwrites an existing composite.toml or menubar.toml, so an existing install needs the route and the price added by hand.

Why

Codex picks the reviewer model ID from its login type, not from Switchyard. codex-auto-review is a hidden model in Codex's built-in catalog. Switchyard's GET /v1/models returns "models": [] on purpose, so Codex keeps that catalog and sends the ID even though Switchyard does not list it. With an OpenAI API-key login, Codex uses gpt-5.6-luna instead. Neither config adds a route for that ID. The server README explains how to add one, and warns that the route then receives every Codex request for gpt-5.6-luna, not only review requests.

The alternative is a Codex setting. Codex can move reviews to another model only through a full replacement model catalog (model_catalog_json), in which the session model's entry sets auto_review_model_override. I tested this: with that override and no reviewer route in Switchyard, the review request went to the switchyard route, and the command ran. But the bundled catalog is about 420 KB and changes with Codex releases, so a Switchyard route is the simpler fix.

Notes for reviewers

This PR is stacked on #863 because it changes the config files that the macOS installer writes. Start with crates/libsy/src/algorithms/codex_approve_all.rs and the CodexApproveAll variant in crates/switchyard-runner/src/algorithm.rs.

The new route also changes GET /v1/models. The server sorts route IDs, and codex-auto-review sorts before switchyard, so default_model and first_id change from switchyard to codex-auto-review, and the startup banner's example curl uses codex-auto-review. I found no client in this repo or in Codex that reads default_model, and this PR does not change how the server picks it.

All live runs below used this branch rebased onto the current #863. Every Codex run used codex exec --approve-for-me (Codex 0.152.0), a temporary CODEX_HOME with no OpenAI login, and a prompt that makes Codex ask for network access to run curl https://example.com. No run sent traffic to OpenAI or ChatGPT. The main switchyard route used an OpenAI-compatible LiteLLM gateway (claude-haiku-4-5). The gateway's own model IDs are replaced below with public model IDs.

Reviewer route Server log for the review request Codex result
None (before) status=404 requested_model="codex-auto-review" command declined
codex_approve_all codex_approve_all approved an action without a review, status=200 command ran and printed 200
passthrough to claude-opus-4-8 (gateway) status=200 requested_model="codex-auto-review" selected_model="claude-opus-4-8" command ran and printed 200
passthrough to gpt-5.6-luna (gateway) status=200 requested_model="codex-auto-review" selected_model="gpt-5.6-luna" command ran and printed 200
None, with the model_catalog_json override the review request went to switchyard, status=200; no codex-auto-review request command ran and printed 200

The Opus target needed omit_body_fields = ["reasoning_effort"]; without it, the gateway returned 400. For the installer's setup, a local stub replaced the ChatGPT backend and recorded what Switchyard forwarded. The details below have both, plus the full server log and Codex events from the run without a reviewer route. I did not test how the real ChatGPT backend answers (that needs ChatGPT traffic) or an API-key login.

codex_approve_all_route_returns_an_allow_verdict sends a request built like Codex's review request to /v1/responses, streamed and buffered, and checks that the reply text parses as {"outcome": "allow"}. The config has no targets, so the test fails if the route tries to call a model. It fails on the old code because the config does not load.

Server log and Codex events without a reviewer route
WARN switchyard_server::request: LLM request failed wire_format=openai_responses status=404 requested_model="codex-auto-review" selected_model="" streaming=true error="No route registered for model codex-auto-review"

Codex events, trimmed from codex exec --json:

{"type":"command_execution","command":"/bin/zsh -lc \"curl -sS -o /dev/null -w '%{http_code}' https://example.com\"","status":"declined"}
{"type":"agent_message","text":"I encountered a system-level issue with the automatic approval review. The escalation request was rejected due to an infrastructure error rather than a policy decision. ..."}
Installer setup: what Switchyard forwards to the ChatGPT backend

I replayed a captured Codex review request through the installer config, with a local stub in place of the ChatGPT backend. The stub recorded:

  • model codex-auto-review, with stream: true and store: false, and a body identical to the one sent
  • the tools inside input[0] as an additional_tools item, no instructions, and reasoning.effort: "low"
  • the headers x-openai-subagent: guardian and x-openai-internal-codex-responses-lite: true
  • the caller's authorization and chatgpt-account-id headers
Opus reviewer: 400 without omit_body_fields

On the gateway I tested, the Opus reviewer first returned 400. Codex sends reasoning.effort: "low" with every review request, and Switchyard passes that setting to the gateway as reasoning_effort. The gateway's 400 error says "thinking.type.enabled" is not supported for this model. Adding omit_body_fields = ["reasoning_effort"] to that target fixed it: the same request without reasoning_effort returns 200.

Menu bar prices and totals with the installer's menubar.toml

The menu bar hides a period's dollar figures (Today, This week) when any model seen in that period has no price. Without a codex-auto-review price, the first review in the installer's setup would hide that period's Saved row. The installer prices codex-auto-review at the baseline_model rates (gpt-5.6-sol), because with a ChatGPT login Codex sends review requests there even without Switchyard. Each review then adds the same amount to the actual cost and to the baseline. The dollars saved stay the same, and the percentage drops a little because the baseline grows. A lower price would make reviews look like savings that routing did not produce.

The docs now say which model ID needs a price in menubar.toml: the target's id when another model reviews, or codex-auto-review (logged with zero tokens) for codex_approve_all.

A throwaway test loaded the installer's menubar.toml, wrote routing-log lines, and printed the menu rows:

one gpt-5.6-luna request:                         Saved $0.12 (80% of $0.14), 1 request
same request + one 10k-token codex-auto-review:   Saved $0.12 (73% of $0.16), 2 requests
same request + one codex_approve_all review:      Saved $0.12 (80% of $0.14), 2 requests
same, with a menubar.toml from before this PR:    no Saved rows
same request + one claude-opus-4-8 review:        no Saved rows (no price for that model)
codex_approve_all replay on /v1/responses

Config: only a codex-auto-review route with type = "codex_approve_all" and an empty [targets] table. The request body is a captured Codex review request, sent once with stream: true and once with stream: false, with the headers x-openai-subagent: guardian and x-openai-internal-codex-responses-lite: true.

streamed: HTTP 200
event: response.created event: response.output_item.added event: response.content_part.added event: response.output_text.delta event: response.output_text.done event: response.content_part.done event: response.output_item.done event: response.completed
text: {"outcome":"allow","rationale":"Switchyard's codex_approve_all route approves every action without a review."}

buffered: HTTP 200
text: {"outcome":"allow","rationale":"Switchyard's codex_approve_all route approves every action without a review."}

server log:
codex_approve_all approved an action without a review route=codex-auto-review
LLM request handled wire_format=openai_responses status=200 requested_model="codex-auto-review" selected_model="codex-auto-review" streaming=true
codex_approve_all approved an action without a review route=codex-auto-review
LLM request handled wire_format=openai_responses status=200 requested_model="codex-auto-review" selected_model="codex-auto-review" streaming=false

routing log (each line): "route_id":"codex-auto-review","algorithm":"codex_approve_all","model":"codex-auto-review", all token counts 0

@elyasmnvidian
elyasmnvidian requested a review from a team as a code owner September 29, 2026 18:08
Signed-off-by: Greg Clark <grclark@nvidia.com>

chore: cleanup

Signed-off-by: Greg Clark <grclark@nvidia.com>

chore: cleanup

Signed-off-by: Greg Clark <grclark@nvidia.com>
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-872/

Built to branch gh-pages at 2026-10-01 20:08 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

…rove_all

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@elyasmnvidian
elyasmnvidian force-pushed the emehtabuddin/switch-1630-codex-approve-for-me-switchyard branch from ee44fbc to 3e9a724 Compare October 1, 2026 20:07
@messiaen
messiaen force-pushed the grclark/mac-deamon branch 3 times, most recently from c871c35 to 042211d Compare October 2, 2026 00:02

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants