feat(server): add forward_auth_from to send an OpenAI login to anthropic_messages clients - #873
Draft
elyasmnvidian wants to merge 3 commits into
Draft
elyasmnvidian wants to merge 3 commits into
elyasmnvidian wants to merge 3 commits into
Conversation
…pic_messages clients Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
…ps application headers Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A
forward_auth = trueroute with an OpenAI-format judge cannot send its Claude requests to/v1/messages. If the judge uses anopenai_responsesclient and the Claude tiers use ananthropic_messagesclient, the server refuses to load the config. For the route in the live run below, renamed here toclaude_route, the error is:This PR adds an opt-in client key,
forward_auth_from = "openai". With it set on theanthropic_messagesclient, the route loads and forwards the caller's one bearer token to both endpoints.Why a Claude target needs
/v1/messagesSome OpenAI-compatible gateways, such as a LiteLLM proxy that serves Claude and GPT models with one API key, accept that key as
Authorization: Beareron/v1/chat/completions,/v1/responses, and/v1/messages. A coding agent such as Oh My Pi calls Switchyard on/v1/responseswith that key, and Switchyard forwards it. On the gateway we tested, a Claude target must use/v1/messagesto get two things:/v1/chat/completionsfor Claude Opus 5.5 or Sonnet 5 return 400 whenever they set any thinking field (reasoning_effort,thinking, oroutput_config), with"thinking.type.enabled" is not supported for this model. On/v1/messages, Switchyard's Anthropic encoder turnsreasoning.effortintothinking: {type: "adaptive"}plusoutput_config.effort, and the gateway accepts it./v1/responsesnever returns cache hits for Claude./v1/messagesdoes, even withoutcache_control.So a
compositeroute with a GPT judge on/v1/responsesand Claude tiers on/v1/messagesmust forward one caller key to both formats. The config check rejects that route even though the gateway accepts the key on both endpoints.Why the check exists
forward_authsends the caller's credential header and other application headers upstream. Claude Code sends an Anthropic login, and Codex sends an OpenAI login. If one route forwarded to both formats, it could send one provider's login to the other provider. Soswitchyard-runnergives each forwarding route one caller credential family, based on its clients' formats, and the server returns 400 when a caller uses the other API family. That default stays. The runner just has no way to learn that one gateway accepts the same key on every endpoint;forward_auth_fromtells it.What changed
[llm_clients.<name>]has a new optional key,forward_auth_from. It names the caller API whose credential a forwarding client accepts:openaioranthropic. When it is unset,formatdecides, as before.forward_auth_fromwhen it is set. Both the config check and the server's caller-format check use that value. The check covers every target a route can call, including classifier, judge, advisor, and subagent targets. A route whose forwarding clients all accept OpenAI callers serves/v1/chat/completionsand/v1/responsescallers.backend.rsandclient.rshave doc-comment changes only. The Anthropic backend already forwards the caller'sauthorizationheader unchanged, forwardsx-api-keyonly if the caller sent one, and always adds its ownanthropic-version. It never buildsx-api-keyfrom the bearer token.extra_headersstill cannot setauthorization,x-api-key,anthropic-version, oranthropic-betaon a forwarding Anthropic client, so a header stored in the config cannot replace the caller's token.Costs
chatgpt-account-idandx-openai-fedrampgo to theanthropic_messagesupstream as ordinary application headers. The client does not mark them sensitive, but it does redact their values from upstream error bodies.extra_headerson that client may also setchatgpt-account-id; the upstream then receives two values, the caller's first and the configured one second. An OpenAI forwarding client rejects thatextra_headersentry.anthropic_messagesclient'sbase_url.forward_auth_fromapplies to the client, not to one route. Every route that uses the client returns the existing 400 to/v1/messagesand/v1/messages/count_tokenscallers. That includes a passthrough route whose only target uses the client.anthropic_messagesclient withoutforward_auth_fromand give Claude Code a route that uses it.GET /v1/modelsstill listsanthropic-messagesinsupported_inbound_formatsfor these routes, as it already does onmainfor OpenAI forwarding routes that refuse/v1/messages.What stays the same
forward_auth_from, a route that mixes forwarding formats fails with the same error as before.forward_auth_from = "openai"on an OpenAI-format client andforward_auth_from = "anthropic"on ananthropic_messagesclient match the default, so the server accepts them and they change nothing.forward_auth = true.Rejected inputs
The config check rejects these inputs:
forward_auth_fromwithoutforward_auth = true:llm client <name> sets forward_auth_from without forward_auth = trueunknown variant `azure`, expected `anthropic` or `openai`forward_auth_from = "anthropic"on anopenai_chatoropenai_responsesclient:llm client <name> cannot set forward_auth_from = "anthropic"; only anthropic_messages clients can forward Anthropic caller credentials. OpenAI-format upstreams read the key only fromauthorization. Anthropic callers often sendx-api-key, and an OpenAI backend does not turn that header into a bearer token.Alternative
Accepting
forward_auth = "openai"would avoid a second key. That would changeforward_authfrom a boolean to a boolean-or-string field. A separate key leaves the type offorward_authunchanged, and existing configs parse without changes.Docs
docs/reference/toml_schema.mddocuments the key with the example above, the headers the Anthropic client forwards, and the costs listed here. The forwarding paragraphs incrates/switchyard-server/README.md,crates/libsy-llm-client/README.md, anddocs/getting_started.mdnow say that forwarding clients must accept the same caller credential, and they point to the new key.crates/switchyard-nemo-relay-plugin/README.mdanddocs/integrations/nemo_relay.mdno longer say that standalone forwarding always needs the caller and target to use the same credential family.Testing
forward_auth_from_sends_one_openai_login_to_both_formatsincrates/switchyard-server/tests/server.rsbuilds anllm_classifierroute throughbuild_switchyard_router. The judge uses anopenai_responsesclient, and the tiers use ananthropic_messagesclient withforward_auth_from = "openai". Both clients point at one local stub server. The test sendsAuthorization: Bearer gateway-keyandchatgpt-account-id: account-1through/v1/chat/completionsand then/v1/responses. It checks that the stub's/v1/responsesand/v1/messagescalls both received the bearer token unchanged andchatgpt-account-id, that neither receivedx-api-key, and that only the/v1/messagescall receivedanthropic-version: 2023-06-01. A/v1/messagescaller gets the existing 400, and the stub receives no request for it. Ifapply_forwarded_authalso sendsx-api-key, this test fails.forward_auth_from_opts_a_route_into_mixed_formatsincrates/switchyard-runner/src/config.rsloads the opted-in route and checks that its caller credential family is OpenAI. It then checks the exact errors for four configs: a mixed route without the opt-in, the opt-in withoutforward_auth, an unknown value, and"anthropic"on a Responses client.When
switchyard-runner,switchyard-server, andswitchyard-llm-clientrun together, the existing testsse::tests::stream_client_error_warn_redacts_upstream_bodysometimes fails.cargo test -p switchyard-server --libalso fails intermittently onmain, and this PR does not touchcrates/switchyard-server/src.Live run
I ran
switchyard-serverfrom this branch on a local port against the gateway described above, with--routing-log-file. GPT-5.6 Terra judges on/v1/responses, and Claude Opus 5.5 and Sonnet 5 serve on/v1/messages. Thecompositeroute matches the one used with Oh My Pi. The two passthrough routes send every request to Opus or Sonnet through the same opted-in client. The config stores no key; every caller sentAuthorization: Bearer <gateway key>.The snippet below shows the config's shape. The base URL is a placeholder, the gateway's own model IDs are replaced with public model IDs, and the client, target, route, and file names are renamed. The command output, tables, error messages, and routing-log records in this section use the same replacements.
The first run (requests 1–4) used this config without the two passthrough routes. The second run (requests A–H and the 400s) also set
max_retries = 0on both clients.--dry-runaccepts the config. Without theforward_auth_fromline, it prints the existing error. The text appears twice because of existing error formatting: the outerinvalid server config {path}: {error}message already includes the inner message, and printing the error chain repeats it. This PR does not change that output.Both caller APIs, streaming and buffered
Requests A–H cover
/v1/responsesand/v1/chat/completions, streaming and buffered. Onclaude-route, the judge chose Sonnet both times (confidence=0.0), so Opus was reached throughopus-route. Every request returned 200, and the server log has a matchingLLM request handled ... status=200line with the same wire format, streaming flag, and selected model./v1/responsesclaude-routereasoningitem withencrypted_content, 29 reasoning tokens, answer9. Events:created, twooutput_item.added(reasoning, message), 7output_text.delta,completed/v1/chat/completionsclaude-route1081,finish_reasonstop, 0 reasoning tokens/v1/responsesopus-routereasoningitem withencrypted_content, 77 reasoning tokens, answer23/v1/chat/completionsopus-routeCanberra,finish_reasonstop, onedata: [DONE]/v1/responsesopus-routecompleted,No. 91 = 7 × 13, .../v1/chat/completionsopus-routeBonjour,finish_reasonstop/v1/responsessonnet-route**¡Hola!**/v1/chat/completionssonnet-route**Hallo!**, onedata: [DONE], 17 reasoning tokensThe routing log has 10 records: a judge call and an answer call for each of the two
claude-routerequests, fouropus-routecalls, and twosonnet-routecalls.Messages callers get 400
/v1/messagesand/v1/messages/count_tokenscallers get the existing 400 without a gateway call, on the composite route and on a passthrough route. The route names in these messages are the renamed routes:Thinking and prompt caching
I sent four requests to
/v1/responsesonclaude-routewithreasoning: {"effort": "high"}. The judge picked the efficient tier each time, so Sonnet 5 answered every request through/v1/messages:reasoningitem withencrypted_content(the signed thinking block), 281 reasoning tokens, answer213With adaptive thinking, the model decides whether to think. It skipped thinking on the easy question and thought on the puzzle.
Routing log records for requests 3 and 4. Each request made one judge call (GPT-5.6 Terra) and one answer call (Claude Sonnet 5):
Related PR
#874 documents the current workarounds in the pi and Oh My Pi guides, including the forwarded-key limit that this PR removes. Whichever PR merges second should update the "Forwarded keys" paragraphs in those guides to mention
forward_auth_from.