Skip to content

proposal(client): let embedded LLM clients classify routing failures #865

Description

@afourniernv

Problem

An embedded RoutedLlmClient may already know why a model call failed and whether another Switchyard candidate can safely serve it. Today Switchyard infers fallback behavior from transport- and HTTP-shaped LlmClientError variants. A host that owns provider pools, circuit state, or policy must either disguise its failure as an HTTP condition or lose the correct fallback behavior.

Proposed direction

Agree on the smallest provider-neutral contract that lets a client return:

  • a bounded, stable failure kind for telemetry; and
  • an explicit routing disposition such as Stop or NextTarget.

Existing built-in errors should keep their current behavior. Focused coverage should prove that NextTarget tries the next candidate and Stop terminates routing.

Non-goals

  • adopting a large provider failure taxonomy;
  • provider-exhaustion partitions or summaries;
  • implementing circuit-breaker state;
  • changing retry policy;
  • defining new server error bodies;
  • judge fail-open behavior; or
  • reclassifying connection timeouts.

LlmClientError is already #[non_exhaustive], so a new variant would be an additive Rust API change for external callers.

Related work

ConductorOne has a working version of the broader contract in ConductorOne/Switchyard#1: typed failure contract and routing/status integration. Its full taxonomy should be treated as prior art rather than adopted wholesale. @pquerna

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions