Skip to content

feat(server): add forward_auth_from to send an OpenAI login to anthropic_messages clients - #873

Draft
elyasmnvidian wants to merge 3 commits into
mainfrom
emehtabuddin/switch-1631-forward-auth-across-formats
Draft

elyasmnvidian wants to merge 3 commits into
mainfrom
emehtabuddin/switch-1631-forward-auth-across-formats

Conversation

@elyasmnvidian

@elyasmnvidian elyasmnvidian commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

A forward_auth = true route with an OpenAI-format judge cannot send its Claude requests to /v1/messages. If the judge uses an openai_responses client and the Claude tiers use an anthropic_messages client, the server refuses to load the config. For the route in the live run below, renamed here to claude_route, the error is:

route claude_route cannot forward both Anthropic and OpenAI caller credentials

This PR adds an opt-in client key, forward_auth_from = "openai". With it set on the anthropic_messages client, the route loads and forwards the caller's one bearer token to both endpoints.

Why a Claude target needs /v1/messages

Some OpenAI-compatible gateways, such as a LiteLLM proxy that serves Claude and GPT models with one API key, accept that key as Authorization: Bearer on /v1/chat/completions, /v1/responses, and /v1/messages. A coding agent such as Oh My Pi calls Switchyard on /v1/responses with that key, and Switchyard forwards it. On the gateway we tested, a Claude target must use /v1/messages to get two things:

  • Thinking at the requested effort. Direct requests to /v1/chat/completions for Claude Opus 5.5 or Sonnet 5 return 400 whenever they set any thinking field (reasoning_effort, thinking, or output_config), with "thinking.type.enabled" is not supported for this model. On /v1/messages, Switchyard's Anthropic encoder turns reasoning.effort into thinking: {type: "adaptive"} plus output_config.effort, and the gateway accepts it.
  • Prompt caching. On that gateway, /v1/responses never returns cache hits for Claude. /v1/messages does, even without cache_control.

So a composite route with a GPT judge on /v1/responses and Claude tiers on /v1/messages must forward one caller key to both formats. The config check rejects that route even though the gateway accepts the key on both endpoints.

Why the check exists

forward_auth sends the caller's credential header and other application headers upstream. Claude Code sends an Anthropic login, and Codex sends an OpenAI login. If one route forwarded to both formats, it could send one provider's login to the other provider. So switchyard-runner gives each forwarding route one caller credential family, based on its clients' formats, and the server returns 400 when a caller uses the other API family. That default stays. The runner just has no way to learn that one gateway accepts the same key on every endpoint; forward_auth_from tells it.

What changed

[llm_clients.<name>] has a new optional key, forward_auth_from. It names the caller API whose credential a forwarding client accepts: openai or anthropic. When it is unset, format decides, as before.

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[llm_clients.gateway_messages]
format = "anthropic_messages"
base_url = "https://gateway.example.com/v1"
forward_auth = true
forward_auth_from = "openai"
  • The runner takes the route's caller credential family from forward_auth_from when it is set. Both the config check and the server's caller-format check use that value. The check covers every target a route can call, including classifier, judge, advisor, and subagent targets. A route whose forwarding clients all accept OpenAI callers serves /v1/chat/completions and /v1/responses callers.
  • No backend code changed; backend.rs and client.rs have doc-comment changes only. The Anthropic backend already forwards the caller's authorization header unchanged, forwards x-api-key only if the caller sent one, and always adds its own anthropic-version. It never builds x-api-key from the bearer token.
  • extra_headers still cannot set authorization, x-api-key, anthropic-version, or anthropic-beta on a forwarding Anthropic client, so a header stored in the config cannot replace the caller's token.

Costs

  • OpenAI application headers reach the Anthropic upstream. With the opt-in, the caller's chatgpt-account-id and x-openai-fedramp go to the anthropic_messages upstream as ordinary application headers. The client does not mark them sensitive, but it does redact their values from upstream error bodies. extra_headers on that client may also set chatgpt-account-id; the upstream then receives two values, the caller's first and the configured one second. An OpenAI forwarding client rejects that extra_headers entry.
  • Nothing checks that both clients use the same gateway. The docs recommend it, but the server does not enforce it. With the opt-in, every OpenAI login a caller sends, including a ChatGPT OAuth token, goes to the anthropic_messages client's base_url.
  • Every route that uses the opted-in client refuses Messages callers. forward_auth_from applies to the client, not to one route. Every route that uses the client returns the existing 400 to /v1/messages and /v1/messages/count_tokens callers. That includes a passthrough route whose only target uses the client.
  • A Claude Code caller needs a second client. To serve Claude Code from the same gateway, add a second anthropic_messages client without forward_auth_from and give Claude Code a route that uses it.

GET /v1/models still lists anthropic-messages in supported_inbound_formats for these routes, as it already does on main for OpenAI forwarding routes that refuse /v1/messages.

What stays the same

  • Without forward_auth_from, a route that mixes forwarding formats fails with the same error as before.
  • forward_auth_from = "openai" on an OpenAI-format client and forward_auth_from = "anthropic" on an anthropic_messages client match the default, so the server accepts them and they change nothing.
  • The NeMo Relay plugin still rejects every route that uses forward_auth = true.

Rejected inputs

The config check rejects these inputs:

  • forward_auth_from without forward_auth = true: llm client <name> sets forward_auth_from without forward_auth = true
  • An unknown value: unknown variant `azure`, expected `anthropic` or `openai`
  • forward_auth_from = "anthropic" on an openai_chat or openai_responses client: llm client <name> cannot set forward_auth_from = "anthropic"; only anthropic_messages clients can forward Anthropic caller credentials. OpenAI-format upstreams read the key only from authorization. Anthropic callers often send x-api-key, and an OpenAI backend does not turn that header into a bearer token.

Alternative

Accepting forward_auth = "openai" would avoid a second key. That would change forward_auth from a boolean to a boolean-or-string field. A separate key leaves the type of forward_auth unchanged, and existing configs parse without changes.

Docs

docs/reference/toml_schema.md documents the key with the example above, the headers the Anthropic client forwards, and the costs listed here. The forwarding paragraphs in crates/switchyard-server/README.md, crates/libsy-llm-client/README.md, and docs/getting_started.md now say that forwarding clients must accept the same caller credential, and they point to the new key. crates/switchyard-nemo-relay-plugin/README.md and docs/integrations/nemo_relay.md no longer say that standalone forwarding always needs the caller and target to use the same credential family.

Testing

  • forward_auth_from_sends_one_openai_login_to_both_formats in crates/switchyard-server/tests/server.rs builds an llm_classifier route through build_switchyard_router. The judge uses an openai_responses client, and the tiers use an anthropic_messages client with forward_auth_from = "openai". Both clients point at one local stub server. The test sends Authorization: Bearer gateway-key and chatgpt-account-id: account-1 through /v1/chat/completions and then /v1/responses. It checks that the stub's /v1/responses and /v1/messages calls both received the bearer token unchanged and chatgpt-account-id, that neither received x-api-key, and that only the /v1/messages call received anthropic-version: 2023-06-01. A /v1/messages caller gets the existing 400, and the stub receives no request for it. If apply_forwarded_auth also sends x-api-key, this test fails.
  • forward_auth_from_opts_a_route_into_mixed_formats in crates/switchyard-runner/src/config.rs loads the opted-in route and checks that its caller credential family is OpenAI. It then checks the exact errors for four configs: a mixed route without the opt-in, the opt-in without forward_auth, an unknown value, and "anthropic" on a Responses client.
cargo test -p switchyard-server --test server forward_auth_from
cargo test -p switchyard-runner --lib forward_auth_from

When switchyard-runner, switchyard-server, and switchyard-llm-client run together, the existing test sse::tests::stream_client_error_warn_redacts_upstream_body sometimes fails. cargo test -p switchyard-server --lib also fails intermittently on main, and this PR does not touch crates/switchyard-server/src.

Live run

I ran switchyard-server from this branch on a local port against the gateway described above, with --routing-log-file. GPT-5.6 Terra judges on /v1/responses, and Claude Opus 5.5 and Sonnet 5 serve on /v1/messages. The composite route matches the one used with Oh My Pi. The two passthrough routes send every request to Opus or Sonnet through the same opted-in client. The config stores no key; every caller sent Authorization: Bearer <gateway key>.

The snippet below shows the config's shape. The base URL is a placeholder, the gateway's own model IDs are replaced with public model IDs, and the client, target, route, and file names are renamed. The command output, tables, error messages, and routing-log records in this section use the same replacements.

schema_version = 1

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[llm_clients.gateway_messages]
format = "anthropic_messages"
base_url = "https://gateway.example.com"
forward_auth = true
forward_auth_from = "openai"

[targets.capable]
id = "claude-opus-5-5"
llm_client = "gateway_messages"

[targets.efficient]
id = "claude-sonnet-5"
llm_client = "gateway_messages"

[targets.judge]
id = "gpt-5.6-terra"
llm_client = "gateway_responses"

[routes.claude_route]
id = "claude-route"
type = "composite"

[routes.claude_route.classifier]
target = "judge"
base_threshold = 0.5
classify_trigger = "user_turn"
message_hash_fallback = true

[routes.claude_route.stage]
capable_target = "capable"
efficient_target = "efficient"
confidence_threshold = 0.5

[routes.opus_route]
id = "opus-route"
type = "passthrough"
target = "capable"

[routes.sonnet_route]
id = "sonnet-route"
type = "passthrough"
target = "efficient"

The first run (requests 1–4) used this config without the two passthrough routes. The second run (requests A–H and the 400s) also set max_retries = 0 on both clients.

--dry-run accepts the config. Without the forward_auth_from line, it prints the existing error. The text appears twice because of existing error formatting: the outer invalid server config {path}: {error} message already includes the inner message, and printing the error chain repeats it. This PR does not change that output.

$ switchyard-server --config gateway.toml --dry-run
server OK: claude-route, opus-route, sonnet-route
$ switchyard-server --config gateway-no-optin.toml --dry-run
invalid server config gateway-no-optin.toml: route claude_route cannot forward both Anthropic and OpenAI caller credentials: route claude_route cannot forward both Anthropic and OpenAI caller credentials

Both caller APIs, streaming and buffered

Requests A–H cover /v1/responses and /v1/chat/completions, streaming and buffered. On claude-route, the judge chose Sonnet both times (confidence=0.0), so Opus was reached through opus-route. Every request returned 200, and the server log has a matching LLM request handled ... status=200 line with the same wire format, streaming flag, and selected model.

# Caller endpoint Mode Effort Route Served model Result
A /v1/responses stream high claude-route Claude Sonnet 5 reasoning item with encrypted_content, 29 reasoning tokens, answer 9. Events: created, two output_item.added (reasoning, message), 7 output_text.delta, completed
B /v1/chat/completions buffered high claude-route Claude Sonnet 5 1081, finish_reason stop, 0 reasoning tokens
C /v1/responses buffered high opus-route Claude Opus 5.5 reasoning item with encrypted_content, 77 reasoning tokens, answer 23
D /v1/chat/completions stream high opus-route Claude Opus 5.5 Canberra, finish_reason stop, one data: [DONE]
E /v1/responses stream high opus-route Claude Opus 5.5 completed, No. 91 = 7 × 13, ...
F /v1/chat/completions buffered — opus-route Claude Opus 5.5 Bonjour, finish_reason stop
G /v1/responses buffered — sonnet-route Claude Sonnet 5 **¡Hola!**
H /v1/chat/completions stream — sonnet-route Claude Sonnet 5 **Hallo!**, one data: [DONE], 17 reasoning tokens

The routing log has 10 records: a judge call and an answer call for each of the two claude-route requests, four opus-route calls, and two sonnet-route calls.

Messages callers get 400

/v1/messages and /v1/messages/count_tokens callers get the existing 400 without a gateway call, on the composite route and on a passthrough route. The route names in these messages are the renamed routes:

POST /v1/messages (claude-route)
400 {"type":"error","error":{"type":"invalid_request_error","message":"route claude-route forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses"}}

POST /v1/messages/count_tokens (claude-route)
400, same message

POST /v1/messages, streaming (opus-route)
400 route opus-route forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses

Thinking and prompt caching

I sent four requests to /v1/responses on claude-route with reasoning: {"effort": "high"}. The judge picked the efficient tier each time, so Sonnet 5 answered every request through /v1/messages:

# Prompt Status Served model Result
1 Bat-and-ball question, 57 input tokens 200 Claude Sonnet 5 Answer only, 0 reasoning tokens
2 Three-remainder puzzle, 70 input tokens 200 Claude Sonnet 5 reasoning item with encrypted_content (the signed thinking block), 281 reasoning tokens, answer 213
3 Inventory log, 17,101 input tokens 200 Claude Sonnet 5 17,099 tokens written to cache
4 Same request as #3 200 Claude Sonnet 5 17,099 tokens read from cache

With adaptive thinking, the model decides whether to think. It skipped thinking on the easy question and thought on the puzzle.

Routing log records for requests 3 and 4. Each request made one judge call (GPT-5.6 Terra) and one answer call (Claude Sonnet 5):

{"ts":"2026-09-30T06:17:13.496Z","route_id":"claude-route","algorithm":"composite","origin":null,"task":null,"trial_id":null,"session_id":null,"model":"gpt-5.6-terra","tier":"classifier","prompt_tokens":13150,"cached_tokens":0,"cache_creation_tokens":13147,"completion_tokens":188,"reasoning_tokens":130,"total_tokens":13338}
{"ts":"2026-09-30T06:17:15.220Z","route_id":"claude-route","algorithm":"composite","origin":null,"task":null,"trial_id":null,"session_id":null,"model":"claude-sonnet-5","tier":"","prompt_tokens":17101,"cached_tokens":0,"cache_creation_tokens":17099,"completion_tokens":25,"reasoning_tokens":0,"total_tokens":17126}
{"ts":"2026-09-30T06:17:28.278Z","route_id":"claude-route","algorithm":"composite","origin":null,"task":null,"trial_id":null,"session_id":null,"model":"gpt-5.6-terra","tier":"classifier","prompt_tokens":13150,"cached_tokens":13147,"cache_creation_tokens":0,"completion_tokens":152,"reasoning_tokens":94,"total_tokens":13302}
{"ts":"2026-09-30T06:17:31.160Z","route_id":"claude-route","algorithm":"composite","origin":null,"task":null,"trial_id":null,"session_id":null,"model":"claude-sonnet-5","tier":"","prompt_tokens":17101,"cached_tokens":17099,"cache_creation_tokens":0,"completion_tokens":25,"reasoning_tokens":0,"total_tokens":17126}

Related PR

#874 documents the current workarounds in the pi and Oh My Pi guides, including the forwarded-key limit that this PR removes. Whichever PR merges second should update the "Forwarded keys" paragraphs in those guides to mention forward_auth_from.

…pic_messages clients

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
…ps application headers

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@github-actions

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-873/

Built to branch gh-pages at 2026-09-30 06:45 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant