Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion crates/libsy-llm-client/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -241,7 +241,10 @@ fn build_multi_format_client(
normalize `authorization`, `chatgpt-account-id`, and `x-openai-fedramp`.
Anthropic backends forward `authorization` or `x-api-key`; they also keep
`oauth-*` values from `anthropic-beta` and remove other caller-supplied beta values.
All backends reachable through a forwarding route must use the same provider.
All backends reachable through a forwarding route must use one credential family
(OpenAI or Anthropic) unless they all use the same scheme, host, and port. Such a
route serves Chat Completions and Responses callers and forwards their bearer
token to every backend.
Headers owned by other providers are preserved as application headers.
- Per-backend custom headers go in `HttpBackendConfig::extra_headers`. Set credentials with
`api_key`. OpenAI backends reject `Authorization`; Anthropic backends reject `x-api-key`
Expand Down
3 changes: 2 additions & 1 deletion crates/libsy-llm-client/src/backend.rs
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,8 @@ pub struct HttpBackendConfig {
pub api_key: Option<String>,
/// Whether this backend forwards the caller's provider credential and application headers.
///
/// All backends reachable through a forwarding route must use the same provider.
/// All backends reachable through a forwarding route must use one credential family
/// (OpenAI or Anthropic) unless they all use the same scheme, host, and port.
pub forward_auth: bool,
/// Custom headers added to every outbound call to this backend.
///
Expand Down
5 changes: 3 additions & 2 deletions crates/libsy-llm-client/src/client.rs
Original file line number Diff line number Diff line change
Expand Up @@ -604,8 +604,9 @@ impl TranslatingLlmClient {
/// `http_headers` are carried through as the request's
/// [`Metadata::http_headers`]. Backends with `forward_auth` disabled forward only
/// allowed metadata headers; `forward_auth` backends forward all application
/// headers. All backends reachable through a forwarding route must use the same
/// provider. Transport headers are always rebuilt. Pass `None` to forward nothing.
/// headers. All backends reachable through a forwarding route must use one credential
/// family unless they all use the same scheme, host, and port. Transport headers are
/// always rebuilt. Pass `None` to forward nothing.
pub async fn call_rewrite_model_raw(
&self,
raw_http_request: Value,
Expand Down
4 changes: 3 additions & 1 deletion crates/switchyard-nemo-relay-plugin/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -216,7 +216,9 @@ and configure `api_key_env` on each authenticated client. The two options cannot
be enabled together. If each caller must use its own provider credential, use
standalone `switchyard-server`. Standalone forwarding requires the caller and
target to use the same credential family: OpenAI-compatible (Chat Completions
and Responses) or Anthropic (Messages).
and Responses) or Anthropic (Messages). A route may mix both families when all
of its forwarding clients use the same scheme, host, and port; such a route
serves Chat Completions and Responses callers.

Support for provider-specific fields depends on the source and target formats.
Test any fields that your application relies on before deploying a translated
Expand Down
109 changes: 104 additions & 5 deletions crates/switchyard-runner/src/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -363,6 +363,8 @@ impl DeploymentConfig {
let mut by_model = HashMap::new();
let mut targets_by_model: HashMap<&str, (&str, &TargetConfig)> = HashMap::new();
let mut caller_auth = None;
let mut mixes_families = false;
let mut forwarding_origins = BTreeSet::new();
for name in route.callable_target_names() {
let target = self.targets.get(name).ok_or_else(|| {
RunnerError::configuration(format!("route references unknown target {name}"))
Expand All @@ -386,16 +388,26 @@ impl DeploymentConfig {
})?;
if client_config.forward_auth {
let target_auth = client_config.format.caller_auth_kind();
if caller_auth.is_some_and(|kind| kind != target_auth) {
return Err(RunnerError::configuration(format!(
"route {route_name} cannot forward both Anthropic and OpenAI caller credentials"
)));
}
mixes_families |= caller_auth.is_some_and(|kind| kind != target_auth);
caller_auth = Some(target_auth);
forwarding_origins.insert(client_config.base_url.0.origin().ascii_serialization());
}
let client: Arc<dyn RoutedLlmClient> = client.clone();
by_model.insert(target.id.clone(), client);
}
// A route may mix credential families only when all of its forwarding clients use the
// same scheme, host, and port, so the caller's credential reaches only that host. Such a
// route serves Chat Completions and Responses callers because Anthropic clients forward
// the caller's `authorization` header unchanged.
if mixes_families {
if forwarding_origins.len() > 1 {
let origins = Vec::from_iter(forwarding_origins).join(", ");
return Err(RunnerError::configuration(format!(
"route {route_name} cannot forward both Anthropic and OpenAI caller credentials to different hosts ({origins}); point all of its forwarding clients at one host"
)));
}
caller_auth = Some(CallerAuthKind::OpenAi);
}
let completion_targets = route
.routing_target_names()
.into_iter()
Expand Down Expand Up @@ -1897,6 +1909,93 @@ confidence_threshold = 0.5
}
}

// Builds a classifier route whose judge uses a Responses client and whose tiers use a
// Messages client, plus a passthrough route that uses only the Messages client.
fn mixed_forwarding_config(messages_url: &str) -> String {
format!(
r#"
schema_version = 1

[llm_clients.responses]
format = "openai_responses"
base_url = "https://gateway.example.test/v1"
forward_auth = true

[llm_clients.messages]
format = "anthropic_messages"
base_url = "{messages_url}"
forward_auth = true

[targets.judge]
id = "judge/model"
llm_client = "responses"

[targets.capable]
id = "capable/model"
llm_client = "messages"

[targets.efficient]
id = "efficient/model"
llm_client = "messages"

[routes.hub]
id = "switchyard/hub"
type = "llm_classifier"
classifier_target = "judge"
strong_target = "capable"
weak_target = "efficient"
base_threshold = 0.5

[routes.claude]
id = "switchyard/claude"
type = "passthrough"
target = "capable"
"#
)
}

#[test]
fn forwarding_route_mixes_credential_families_only_on_one_host() -> RunnerResult<()> {
// The URL path does not count, so both base URLs name one host.
for messages_url in [
"https://gateway.example.test",
"https://gateway.example.test/v1",
] {
let runner = runner_from_toml(&mixed_forwarding_config(messages_url))?;
let caller_auth = |id| runner.route(id).and_then(Route::caller_auth);
assert_eq!(caller_auth("switchyard/hub"), Some(CallerAuthKind::OpenAi));
assert_eq!(
caller_auth("switchyard/claude"),
Some(CallerAuthKind::Anthropic)
);
}

// A different host, port, or scheme is a different origin.
for (messages_url, origins) in [
(
"https://api.anthropic.test",
"https://api.anthropic.test, https://gateway.example.test",
),
(
"https://gateway.example.test:8443",
"https://gateway.example.test, https://gateway.example.test:8443",
),
(
"http://gateway.example.test",
"http://gateway.example.test, https://gateway.example.test",
),
] {
let error = error_message(&mixed_forwarding_config(messages_url));
assert!(
error.contains(&format!(
"route hub cannot forward both Anthropic and OpenAI caller credentials to different hosts ({origins}); point all of its forwarding clients at one host"
)),
"{error}"
);
}
Ok(())
}

const ADVISOR_CONFIG: &str = r#"
schema_version = 1

Expand Down
12 changes: 8 additions & 4 deletions crates/switchyard-server/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,10 +91,14 @@ A client can set `forward_auth = true` instead of `api_key_env` to send the
caller's credential to the configured upstream. OpenAI clients forward
`authorization`, `chatgpt-account-id`, and `x-openai-fedramp`. Anthropic clients
forward `authorization` or `x-api-key`. Enable this only when every forwarding
client's `base_url` should receive the caller's login. All backends reachable
through the route must use the same provider. Other application headers are
preserved and may contain provider-specific credentials. A forwarding route
must be called through the matching provider API.
client's `base_url` should receive the caller's login. Other application headers
are preserved and may contain provider-specific credentials. A route's
forwarding clients must use one credential family: all OpenAI formats or all
`anthropic_messages`. A route may mix the two families only when all of its
forwarding clients use the same scheme, host, and port. Such a route serves
Chat Completions and Responses callers and forwards the caller's bearer token to
every client.
The server returns 400 to a caller whose API the route does not serve.
Target-level `extra_body` values are shallow-merged into the upstream request when
the request does not already contain that key.
Target-level `system_prompt` values are prepended when that target serves a completion.
Expand Down
170 changes: 170 additions & 0 deletions crates/switchyard-server/tests/server.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3280,6 +3280,176 @@ target = "openai"
Ok(())
}

/// Serves `/v1/responses` and `/v1/messages` for a stub gateway on one host. For each call, it
/// records every value of the `authorization`, `x-api-key`, `chatgpt-account-id`, and
/// `anthropic-version` headers, so a header sent twice shows up as two values. Every
/// `/v1/responses` call returns a judge verdict that picks the efficient tier.
async fn upstream_gateway_records_auth(
State(calls): State<Arc<Mutex<Vec<Value>>>>,
uri: Uri,
headers: HeaderMap,
Json(body): Json<Value>,
) -> HttpResponse {
let header = |name: &str| {
headers
.get_all(name)
.iter()
.filter_map(|value| value.to_str().ok())
.collect::<Vec<_>>()
};
calls.lock().await.push(json!({
"path": uri.path(),
"authorization": header("authorization"),
"x_api_key": header("x-api-key"),
"chatgpt_account_id": header("chatgpt-account-id"),
"anthropic_version": header("anthropic-version"),
}));
let model = body["model"].as_str().unwrap_or_default();
if uri.path() == "/v1/responses" {
let verdict = json!({
"crux": "bounded task", "primary_rule": "SUP-1",
"capability_boundary": "supported", "p_solve": 0.9,
});
return Json(responses_body("resp_judge", model, &verdict.to_string())).into_response();
}
Json(json!({
"id": "msg_gateway", "type": "message", "role": "assistant", "model": model,
"content": [{"type": "text", "text": "ok"}],
"stop_reason": "end_turn", "stop_sequence": null,
"usage": {"input_tokens": 1, "output_tokens": 1}
}))
.into_response()
}

#[tokio::test]
async fn route_on_one_host_forwards_the_bearer_token_to_responses_and_messages() -> TestResult {
let calls = Arc::new(Mutex::new(Vec::new()));
let gateway = Router::new()
.route("/v1/responses", post(upstream_gateway_records_auth))
.route("/v1/messages", post(upstream_gateway_records_auth))
.with_state(Arc::clone(&calls));
let listener = TcpListener::bind("127.0.0.1:0").await?;
let base_url = format!("http://{}/v1", listener.local_addr()?);
tokio::spawn(async move { axum::serve(listener, gateway).await });
let state = load_test_config(&format!(
r#"
schema_version = 1

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "{base_url}"
forward_auth = true
max_retries = 0

[llm_clients.gateway_messages]
format = "anthropic_messages"
base_url = "{base_url}"
forward_auth = true
max_retries = 0

[targets]
judge = {{ id = "model/judge", llm_client = "gateway_responses" }}
capable = {{ id = "model/capable", llm_client = "gateway_messages" }}
efficient = {{ id = "model/efficient", llm_client = "gateway_messages" }}

[routes.agent]
id = "switchyard/agent"
type = "composite"
classifier = {{ target = "judge", base_threshold = 0.5, classify_trigger = "user_turn" }}
stage = {{ capable_target = "capable", efficient_target = "efficient", confidence_threshold = 0.5 }}

[routes.claude]
id = "switchyard/claude"
type = "passthrough"
target = "efficient"
"#
))?;
let app = build_switchyard_router(state);
let bearer = [
("authorization", "Bearer gateway-key"),
("chatgpt-account-id", "account-1"),
];

for (path, body) in [
(
"/v1/chat/completions",
json!({"model": "switchyard/agent", "messages": [{"role": "user", "content": "hello"}]}),
),
(
"/v1/responses",
json!({"model": "switchyard/agent", "input": "hi there"}),
),
] {
let response = send_with_headers(&app, "POST", path, Some(body), &bearer).await?;
assert_eq!(
response.status,
StatusCode::OK,
"{path}: {}",
response.text()?
);
assert_eq!(
response.headers["x-model-router-selected-model"],
"model/efficient"
);
}
// Both endpoints receive exactly one copy of the caller's bearer token and no `x-api-key`.
// Only the Messages call gets `anthropic-version`.
let judge = json!({
"path": "/v1/responses", "authorization": ["Bearer gateway-key"], "x_api_key": [],
"chatgpt_account_id": ["account-1"], "anthropic_version": []
});
let answer = json!({
"path": "/v1/messages", "authorization": ["Bearer gateway-key"], "x_api_key": [],
"chatgpt_account_id": ["account-1"], "anthropic_version": ["2023-06-01"]
});
assert_eq!(
*calls.lock().await,
[judge.clone(), answer.clone(), judge, answer]
);

// The mixed route serves OpenAI callers only, so a Messages caller gets 400 before any call.
let messages_body = |model: &str| {
json!({
"model": model,
"max_tokens": 16,
"messages": [{"role": "user", "content": "hello"}]
})
};
let wrong_api = send_with_headers(
&app,
"POST",
"/v1/messages",
Some(messages_body("switchyard/agent")),
&bearer,
)
.await?;
assert_eq!(wrong_api.status, StatusCode::BAD_REQUEST);
assert_eq!(
wrong_api.json()?["error"]["message"],
"route switchyard/agent forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses"
);
assert_eq!(calls.lock().await.len(), 4);

// A route that uses only the Messages client still serves Messages callers.
let claude = send_with_headers(
&app,
"POST",
"/v1/messages",
Some(messages_body("switchyard/claude")),
&[("authorization", "Bearer gateway-key")],
)
.await?;
assert_eq!(claude.status, StatusCode::OK, "{}", claude.text()?);
assert_eq!(
calls.lock().await[4..],
[json!({
"path": "/v1/messages", "authorization": ["Bearer gateway-key"], "x_api_key": [],
"chatgpt_account_id": [], "anthropic_version": ["2023-06-01"]
})]
);
Ok(())
}

#[tokio::test]
async fn routes_dispatch_and_discovery_endpoints_are_stable() -> TestResult {
let (upstream, app) = test_app(&[
Expand Down
12 changes: 8 additions & 4 deletions docs/getting_started.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,10 +109,14 @@ A client can set `forward_auth = true` instead of `api_key_env` to send each
caller's credential to that upstream. OpenAI clients forward `authorization`,
`chatgpt-account-id`, and `x-openai-fedramp`. Anthropic clients forward
`authorization` or `x-api-key`. Enable this only for an upstream that should
receive the caller's login. All backends reachable through the route, including
efficient and capable targets, must use the same provider. Other application
headers are preserved, so they may contain provider-specific credentials. The
server rejects a forwarding route called through the other provider's API.
receive the caller's login. Other application headers are preserved, so they
may contain provider-specific credentials. The forwarding clients in a route,
including efficient and capable targets, must use one credential family: all
OpenAI formats or all `anthropic_messages`. A route may mix the two families
only when all of its forwarding clients use the same scheme, host, and port.
Such a route serves Chat Completions and Responses callers and forwards the
caller's bearer token to every client. The server returns 400 to a caller whose
API the route does not serve.

### Run the server

Expand Down
Loading
Loading