You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Responses continuation state currently lives inside each route's ClientRouter in a
process-local Mutex<StateOwners>. Each server replica constructs its own router and
therefore has an independent state map.
A continuation works only when it reaches the replica that observed the original
response. A restart has the same failure mode.
Verified reproduction
Reproduced on current main (64def564) using two independent switchyard-server processes and one deterministic loopback upstream:
Both replicas use the same route and provider configuration.
A seed request sent to replica A is routed to provider A.
Provider A returns resp_A_*.
A follow-up containing that ID in previous_response_id asks the classifier
to choose provider B if routing runs again.
Sending the follow-up to replica A returns HTTP 200 and remains pinned to
provider A.
Sending it to replica B misses the process-local map, runs routing again,
selects provider B, and returns HTTP 400 state_not_found.
The same-replica control passed 3/3 times. The cross-replica failure reproduced
3/3 times.
Why it happens
ClientRouter::stored_state_owner looks up previous_response_id or conversation in the router's local StateOwners map. A miss is treated as if
there were no known continuation, so run executes the routing algorithm again.
Replica B cannot distinguish these cases:
a genuinely new request;
an unknown continuation;
a continuation whose record expired; or
a continuation created through another replica.
Selecting another provider after such a miss is unsafe because provider response
IDs are opaque and provider-owned.
This is related to, but not solved by, #802/#803. That work handles materialized
conversation history within one router. It does not share provider ownership or
canonical history across replicas.
Proposed ownership boundary
libsy-llm-client owns
detecting Responses continuation selectors;
constructing scoped keys;
the versioned state-record format;
provider ownership and canonical-history semantics;
bypassing routing when state is found;
conflict and concurrent-update rules;
mapping store outcomes to request or stream behavior; and
bounded telemetry that does not expose IDs or message content.
A host-provided store owns
durable and atomic persistence;
availability and connection management;
retention, tombstones, and capacity;
encryption, access control, and deletion; and
implementing the required conditional or transactional operations.
Out of scope for the first increment
selecting Redis, PostgreSQL, or another production backend;
deployment TOML and credentials;
operating a shared-store service;
changing the wire API; and
adding a server adapter before the storage contract is accepted.
Existing constructors should continue to use the current bounded in-memory
behavior.
Interface sketch
This is conceptual rather than a proposed final Rust signature:
ResponseStateWrite represents one atomic operation containing:
the response-ID alias;
an optional conversation alias;
a versioned opaque record;
an optional expected conversation revision; and
retention metadata.
The library should own the record encoding. Store implementations should not need
to construct or understand canonical message-history internals.
A ClientRouter builder or constructor extension would accept the store and its
stable namespace. Existing constructors would retain the in-memory default.
Proposed runtime behavior
Requests without previous_response_id or conversation perform no state
lookup.
A found record bypasses routing and selects its recorded target.
Canonical history is loaded only when a cross-format continuation needs it.
Native Responses calls should not perform a redundant history lookup.
Response and conversation aliases for one completed turn are committed
atomically.
A lookup outage must stop before routing instead of selecting another provider.
A buffered response should not report successful durable state if its commit
failed.
For streaming, the terminal completion should not be exposed until the required
commit succeeds. A failure after headers were sent must become an explicit
terminal stream error.
Decisions needed before implementation
Key scope and target identity
A key likely needs:
a deployment or configuration namespace shared by replicas;
the route identity;
a caller or tenant scope;
selector kind (response or conversation); and
the opaque selector ID.
The namespace must remain stable across equivalent replicas but rotate when target
mappings change incompatibly. Storing only ModelId is not sufficient if a later
configuration maps the same model ID to another provider.
Caller isolation
The current proposal in #803 keys conversation history only by the caller-supplied
conversation ID. I reproduced two different session_id values sharing history
when they reused the same conversation ID.
session_id can prevent accidental collisions, but it is client-controlled and is
not an authentication boundary. We need to decide whether shared deployments
require a trusted host-provided tenant/principal scope and what happens when one is
unavailable.
The same policy should be considered for response IDs, even though they are
normally harder to guess.
Missing versus expired state
A never-seen conversation may represent a new conversation. An expired
conversation represents lost continuation state. Those cases should not have the
same result.
A shared adapter may need tombstones so it can return Expired separately from Missing. For previous_response_id on a route spanning providers, an
unclassified miss should fail closed rather than run routing again.
Store outages and capacity
Lookup failures should return an explicit service error before routing. Capacity
or commit failures should not be logged and silently ignored, because that returns
an ID which the next replica cannot safely resolve.
The exact HTTP and stream error codes remain to be agreed.
Concurrent conversation turns
Atomic alias insertion alone is not enough when two continuations of the same
conversation run concurrently. The contract should decide between:
revision-based compare-and-swap with a conflict response;
explicitly supported branching; or
completion-order last-write-wins.
This decision should align with the latest-history behavior proposed in #803.
Retention and privacy
The production adapter must define:
active-record TTL;
tombstone TTL;
whether reads refresh retention;
maximum records and history size;
deletion behavior; and
encryption and access controls for canonical history containing user content.
First implementation increment
The first PR should only establish and test the agreed boundary:
keep the current in-memory behavior as the default;
allow two independently constructed routers to share one test store;
cover provider-owned affinity and cross-format canonical history;
cover buffered and completed streaming responses;
test missing, expired, unavailable, conflict, capacity, and stale-revision
outcomes;
Problem
Responses continuation state currently lives inside each route's
ClientRouterin aprocess-local
Mutex<StateOwners>. Each server replica constructs its own router andtherefore has an independent state map.
A continuation works only when it reaches the replica that observed the original
response. A restart has the same failure mode.
Verified reproduction
Reproduced on current main (
64def564) using two independentswitchyard-serverprocesses and one deterministic loopback upstream:resp_A_*.previous_response_idasks the classifierto choose provider B if routing runs again.
provider A.
selects provider B, and returns HTTP 400
state_not_found.The same-replica control passed 3/3 times. The cross-replica failure reproduced
3/3 times.
Why it happens
ClientRouter::stored_state_ownerlooks upprevious_response_idorconversationin the router's localStateOwnersmap. A miss is treated as ifthere were no known continuation, so
runexecutes the routing algorithm again.Replica B cannot distinguish these cases:
Selecting another provider after such a miss is unsafe because provider response
IDs are opaque and provider-owned.
This is related to, but not solved by, #802/#803. That work handles materialized
conversation history within one router. It does not share provider ownership or
canonical history across replicas.
Proposed ownership boundary
libsy-llm-clientownsA host-provided store owns
Out of scope for the first increment
Existing constructors should continue to use the current bounded in-memory
behavior.
Interface sketch
This is conceptual rather than a proposed final Rust signature:
ResponseStateWriterepresents one atomic operation containing:The library should own the record encoding. Store implementations should not need
to construct or understand canonical message-history internals.
A
ClientRouterbuilder or constructor extension would accept the store and itsstable namespace. Existing constructors would retain the in-memory default.
Proposed runtime behavior
previous_response_idorconversationperform no statelookup.
Native Responses calls should not perform a redundant history lookup.
atomically.
failed.
commit succeeds. A failure after headers were sent must become an explicit
terminal stream error.
Decisions needed before implementation
Key scope and target identity
A key likely needs:
responseorconversation); andThe namespace must remain stable across equivalent replicas but rotate when target
mappings change incompatibly. Storing only
ModelIdis not sufficient if a laterconfiguration maps the same model ID to another provider.
Caller isolation
The current proposal in #803 keys conversation history only by the caller-supplied
conversation ID. I reproduced two different
session_idvalues sharing historywhen they reused the same conversation ID.
session_idcan prevent accidental collisions, but it is client-controlled and isnot an authentication boundary. We need to decide whether shared deployments
require a trusted host-provided tenant/principal scope and what happens when one is
unavailable.
The same policy should be considered for response IDs, even though they are
normally harder to guess.
Missing versus expired state
A never-seen conversation may represent a new conversation. An expired
conversation represents lost continuation state. Those cases should not have the
same result.
A shared adapter may need tombstones so it can return
Expiredseparately fromMissing. Forprevious_response_idon a route spanning providers, anunclassified miss should fail closed rather than run routing again.
Store outages and capacity
Lookup failures should return an explicit service error before routing. Capacity
or commit failures should not be logged and silently ignored, because that returns
an ID which the next replica cannot safely resolve.
The exact HTTP and stream error codes remain to be agreed.
Concurrent conversation turns
Atomic alias insertion alone is not enough when two continuations of the same
conversation run concurrently. The contract should decide between:
This decision should align with the latest-history behavior proposed in #803.
Retention and privacy
The production adapter must define:
First implementation increment
The first PR should only establish and test the agreed boundary:
outcomes;