Skip to content

[feature] Make Responses continuation state shareable across replicas #832

Description

@cpakkamisaac-sae

Problem

Responses continuation state currently lives inside each route's ClientRouter in a
process-local Mutex<StateOwners>. Each server replica constructs its own router and
therefore has an independent state map.

A continuation works only when it reaches the replica that observed the original
response. A restart has the same failure mode.

Verified reproduction

Reproduced on current main (64def564) using two independent
switchyard-server processes and one deterministic loopback upstream:

  1. Both replicas use the same route and provider configuration.
  2. A seed request sent to replica A is routed to provider A.
  3. Provider A returns resp_A_*.
  4. A follow-up containing that ID in previous_response_id asks the classifier
    to choose provider B if routing runs again.
  5. Sending the follow-up to replica A returns HTTP 200 and remains pinned to
    provider A.
  6. Sending it to replica B misses the process-local map, runs routing again,
    selects provider B, and returns HTTP 400 state_not_found.

The same-replica control passed 3/3 times. The cross-replica failure reproduced
3/3 times.

Why it happens

ClientRouter::stored_state_owner looks up previous_response_id or
conversation in the router's local StateOwners map. A miss is treated as if
there were no known continuation, so run executes the routing algorithm again.

Replica B cannot distinguish these cases:

  • a genuinely new request;
  • an unknown continuation;
  • a continuation whose record expired; or
  • a continuation created through another replica.

Selecting another provider after such a miss is unsafe because provider response
IDs are opaque and provider-owned.

This is related to, but not solved by, #802/#803. That work handles materialized
conversation history within one router. It does not share provider ownership or
canonical history across replicas.

Proposed ownership boundary

libsy-llm-client owns

  • detecting Responses continuation selectors;
  • constructing scoped keys;
  • the versioned state-record format;
  • provider ownership and canonical-history semantics;
  • bypassing routing when state is found;
  • conflict and concurrent-update rules;
  • mapping store outcomes to request or stream behavior; and
  • bounded telemetry that does not expose IDs or message content.

A host-provided store owns

  • durable and atomic persistence;
  • availability and connection management;
  • retention, tombstones, and capacity;
  • encryption, access control, and deletion; and
  • implementing the required conditional or transactional operations.

Out of scope for the first increment

  • selecting Redis, PostgreSQL, or another production backend;
  • deployment TOML and credentials;
  • operating a shared-store service;
  • changing the wire API; and
  • adding a server adapter before the storage contract is accepted.

Existing constructors should continue to use the current bounded in-memory
behavior.

Interface sketch

This is conceptual rather than a proposed final Rust signature:

#[async_trait]
pub trait ResponseStateStore: Send + Sync {
    async fn load(
        &self,
        key: &ResponseStateKey,
    ) -> Result<StateRead, StateStoreError>;

    async fn commit(
        &self,
        write: ResponseStateWrite,
    ) -> Result<(), StateStoreError>;
}

pub enum StateRead {
    Found { revision: u64, record: OpaqueStateRecord },
    Missing,
    Expired,
}

ResponseStateWrite represents one atomic operation containing:

  • the response-ID alias;
  • an optional conversation alias;
  • a versioned opaque record;
  • an optional expected conversation revision; and
  • retention metadata.

The library should own the record encoding. Store implementations should not need
to construct or understand canonical message-history internals.

A ClientRouter builder or constructor extension would accept the store and its
stable namespace. Existing constructors would retain the in-memory default.

Proposed runtime behavior

  • Requests without previous_response_id or conversation perform no state
    lookup.
  • A found record bypasses routing and selects its recorded target.
  • Canonical history is loaded only when a cross-format continuation needs it.
    Native Responses calls should not perform a redundant history lookup.
  • Response and conversation aliases for one completed turn are committed
    atomically.
  • A lookup outage must stop before routing instead of selecting another provider.
  • A buffered response should not report successful durable state if its commit
    failed.
  • For streaming, the terminal completion should not be exposed until the required
    commit succeeds. A failure after headers were sent must become an explicit
    terminal stream error.

Decisions needed before implementation

Key scope and target identity

A key likely needs:

  • a deployment or configuration namespace shared by replicas;
  • the route identity;
  • a caller or tenant scope;
  • selector kind (response or conversation); and
  • the opaque selector ID.

The namespace must remain stable across equivalent replicas but rotate when target
mappings change incompatibly. Storing only ModelId is not sufficient if a later
configuration maps the same model ID to another provider.

Caller isolation

The current proposal in #803 keys conversation history only by the caller-supplied
conversation ID. I reproduced two different session_id values sharing history
when they reused the same conversation ID.

session_id can prevent accidental collisions, but it is client-controlled and is
not an authentication boundary. We need to decide whether shared deployments
require a trusted host-provided tenant/principal scope and what happens when one is
unavailable.

The same policy should be considered for response IDs, even though they are
normally harder to guess.

Missing versus expired state

A never-seen conversation may represent a new conversation. An expired
conversation represents lost continuation state. Those cases should not have the
same result.

A shared adapter may need tombstones so it can return Expired separately from
Missing. For previous_response_id on a route spanning providers, an
unclassified miss should fail closed rather than run routing again.

Store outages and capacity

Lookup failures should return an explicit service error before routing. Capacity
or commit failures should not be logged and silently ignored, because that returns
an ID which the next replica cannot safely resolve.

The exact HTTP and stream error codes remain to be agreed.

Concurrent conversation turns

Atomic alias insertion alone is not enough when two continuations of the same
conversation run concurrently. The contract should decide between:

  • revision-based compare-and-swap with a conflict response;
  • explicitly supported branching; or
  • completion-order last-write-wins.

This decision should align with the latest-history behavior proposed in #803.

Retention and privacy

The production adapter must define:

  • active-record TTL;
  • tombstone TTL;
  • whether reads refresh retention;
  • maximum records and history size;
  • deletion behavior; and
  • encryption and access controls for canonical history containing user content.

First implementation increment

The first PR should only establish and test the agreed boundary:

  • keep the current in-memory behavior as the default;
  • allow two independently constructed routers to share one test store;
  • cover provider-owned affinity and cross-format canonical history;
  • cover buffered and completed streaming responses;
  • test missing, expired, unavailable, conflict, capacity, and stale-revision
    outcomes;
  • test caller isolation using the policy agreed for fix(llm-client): store canonical history under the conversation id #803;
  • prove response and conversation aliases are committed atomically;
  • prove selector-free requests do not incur a lookup;
  • prove native Responses output does not incur a redundant history lookup; and
  • keep production adapters and server TOML out of scope.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions