Skip to content

[bug] stage_router no longer falls back to the capable target when the connection to the efficient upstream times out (0.3.0; 0.2.0 fell back) #853

Description

@mill101

Summary

With 0.3.0, if a stage_router route's efficient target is down (connection refused), the request is retried for ~91 s and then fails with 504 upstream_timeout. The capable target is never tried. With 0.2.0 and the same config, the request falls back to the capable target (reason="unavailable") and succeeds.

This matters for setups where the efficient tier is a local model that may be stopped: in 0.3.0 the route becomes unusable instead of degrading to the capable tier.

Versions

  • switchyard-server 0.3.0 (cargo install --locked switchyard-server --version 0.3.0) — fails
  • switchyard-server 0.2.0 (git 9b6efb94) — falls back
  • Ubuntu 24.04.4 LTS on WSL2
  • x86_64; inbound: Chat Completions (non-streaming); upstreams: OpenAI-compatible (openai_chat)

Minimal config

schema_version = 1

[llm_clients.up]
format = "openai_chat"
base_url = "http://127.0.0.1:18100/v1"   # any OpenAI-compatible server that answers

[llm_clients.down]
format = "openai_chat"
base_url = "http://127.0.0.1:18199/v1"   # nothing listens here -> connection refused

[targets.good]
id = "mock-good"
llm_client = "up"

[targets.dead]
id = "mock-dead"
llm_client = "down"

[routes.fb_dead]
id = "fb-dead"
type = "stage_router"
efficient_target = "dead"
capable_target = "good"
picker = "efficient_first"
confidence_threshold = 0.5
recent_turn_window = 3

Steps

RUST_LOG=info switchyard-server --config min.toml --host 127.0.0.1 --port 4005
curl -s -w '\nHTTP %{http_code} %{time_total}s\n' -H 'content-type: application/json' \
  -d '{"model":"fb-dead","stream":false,"max_tokens":8,"messages":[{"role":"user","content":"ping"}]}' \
  http://127.0.0.1:4005/v1/chat/completions

Observed

0.3.0 0.2.0
HTTP 504 after 91.37 s 200 after 91.37 s
Body {"error":{"message":"error sending request","type":"upstream_error","code":"upstream_timeout"}} answer from mock-good
Log LLM request failed … status=504 requested_model="fb-dead" selected_model="" model call failed; trying next candidate from=mock-dead to=mock-good reason="unavailable"

Other cases behaved identically in both versions (429/500 → retried 3× then fell back; 403 → fell back; 400/404 → no fallback; streaming fallback on 429 works).

What seems to happen (from reading 0.3.0 source)

  1. switchyard-llm-client client.rs: AttemptFailure::is_retryable treats LlmClientError::Transport as retryable, so a refused connection is retried until the response deadline.
  2. When the deadline expires, deadline_error() returns LlmClientError::Timeout ("response did not finish within … ms, retries included").
  3. run.rs fallback_reason() maps Transport to RoutingFallbackReason::Unavailable, but has no arm for Timeout, so the route stops instead of trying the next candidate.

So a persistent transport failure never reaches fallback_reason() as Transport; it arrives as Timeout and falls through to None.

Is this intended?

This looks related to #702 (merged 2026-09-16), whose description says:

After retries, the HTTP driver stops routing on client errors; a timeout never tries another candidate.

Compatibility: HTTP client errors stop routing even when timeout_ms is unset or an advisor has fail_open = true.

So stopping on a timeout is clearly intended. What I'm unsure about is whether an upstream that refuses the connection was meant to fall under this too. Before 0.3.0 that case fell back as Unavailable, and fallback_reason() still has a Transport → Unavailable arm, but with the new retry loop a persistent transport failure seems to always end as Timeout, so that arm is effectively never reached for a down upstream.

For deployments whose efficient tier is a local model that may be stopped, this changes stage_router from "degrade to the capable tier" to "fail after ~91 s". If this is intended, it would help to document it (release notes / routing docs), since the 0.3.0 notes mention fallback correctness but not this change. If it is not intended, some options:

  • keep the last Transport error (→ Unavailable) when the deadline expires after transport-only failures, or
  • stop retrying a candidate on connection-refused and hand off to the next candidate.

Separately, ~91 s before giving up on a refused connection seems long for a local upstream; is there a per-client setting to shorten it?

Happy to test a patch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions