Problem
classify_trigger = "user_turn" (#487) re-decides the tier on every new user message, but the task the judge is shown is still anchored to the conversation's first user message. task_messages sends "the opening task and the latest user follow-up", and trim_messages sends "the opening task and the last N messages after it". Both treat the newest user message as a follow-up to the opening task.
In an interactive coding session that assumption stops holding after the first job. A session that opens with "add caching to the fetcher" and, two hundred messages later, asks "now write the migration for the new table" is judged with the caching request as its task and the migration request as a follow-up. Every re-decision user_turn makes after the first is anchored to a task the user finished long ago, and recent_turn_window only widens the tail after that stale anchor. recent_turn_window = 0 makes it worse: the judge sees only the opening task.
The client cannot fix this from outside. Coding agents send the whole transcript on every request, so there is no way to make the current request the "first user message" without dropping history the answer model needs. Related: #252 asks to separate the user's question from injected context; #495 covers the session-id requirement of user_turn.
Proposed solution
A route-level key on capability and custom llm_classifier routes that says which user message is the task:
[routes.auto]
type = "llm_classifier"
classify_trigger = "user_turn"
task_anchor = "latest_user_turn" # default: "opening_task" (today's behavior)
recent_turn_window = 4
task_anchor |
without recent_turn_window |
with recent_turn_window = N |
opening_task (default) |
opening task + latest user follow-up when they differ |
opening task + the last N messages after it |
latest_user_turn |
the latest ordinary user message alone |
the N messages before the latest user message, then that message last |
"Ordinary user message" uses the existing rule: a user-role message with at least one block that is not a tool call, tool result, or reasoning, so a tool-result-only user message never becomes the task. In the windowed form the window keeps tool call and result pairs whole, the same way the trailing window does today, and the trailing routing instruction is still appended after the task. Client system and developer instructions are handled as before.
The default is unchanged. Escalation mode keeps its own anchors and rejects the key like the other capability-only settings; stage and composite routes and the Python bindings keep the default.
Alternatives considered
Scope notes
- Surface: routing algorithm (
switchyard-libsy llm_classifier, capability and custom modes) and the deployment TOML parsed by switchyard-runner.
- Adds one optional TOML key and one public enum field with a default on
TaskClassifierConfig and CustomClassifierConfig; no PyO3 or Python API change.
- Backward compatible: the default reproduces current behavior exactly.
Additional context
The motivating setup is a fine-tuned judge in front of an interactive coding agent, where the session-level trigger (new_session) was too coarse (one judgement per multi-job session) and every_request too fine (tier flips on tool turns). user_turn is the right trigger for that harness once the judge sees the current request as the task. I have a patch with tests and docs ready and will open a PR referencing this issue.
Problem
classify_trigger = "user_turn"(#487) re-decides the tier on every new user message, but the task the judge is shown is still anchored to the conversation's first user message.task_messagessends "the opening task and the latest user follow-up", andtrim_messagessends "the opening task and the last N messages after it". Both treat the newest user message as a follow-up to the opening task.In an interactive coding session that assumption stops holding after the first job. A session that opens with "add caching to the fetcher" and, two hundred messages later, asks "now write the migration for the new table" is judged with the caching request as its task and the migration request as a follow-up. Every re-decision
user_turnmakes after the first is anchored to a task the user finished long ago, andrecent_turn_windowonly widens the tail after that stale anchor.recent_turn_window = 0makes it worse: the judge sees only the opening task.The client cannot fix this from outside. Coding agents send the whole transcript on every request, so there is no way to make the current request the "first user message" without dropping history the answer model needs. Related: #252 asks to separate the user's question from injected context; #495 covers the session-id requirement of
user_turn.Proposed solution
A route-level key on capability and custom
llm_classifierroutes that says which user message is the task:task_anchorrecent_turn_windowrecent_turn_window = Nopening_task(default)latest_user_turn"Ordinary user message" uses the existing rule: a
user-role message with at least one block that is not a tool call, tool result, or reasoning, so a tool-result-only user message never becomes the task. In the windowed form the window keeps tool call and result pairs whole, the same way the trailing window does today, and the trailing routing instruction is still appended after the task. Client system and developer instructions are handled as before.The default is unchanged. Escalation mode keeps its own anchors and rejects the key like the other capability-only settings; stage and composite routes and the Python bindings keep the default.
Alternatives considered
classify_trigger = "user_turn": changes the judge input for every existinguser_turnroute, and some of them want the opening task as context. An explicit key keeps that choice with the operator.<history>and<query>sections as proposed in [Feature] Separate the user's question from the context Claude Code injects around it #252: complementary rather than competing. This issue only decides which message is the task; [Feature] Separate the user's question from the context Claude Code injects around it #252 decides how the surrounding context is labelled.Scope notes
switchyard-libsyllm_classifier, capability and custom modes) and the deployment TOML parsed byswitchyard-runner.TaskClassifierConfigandCustomClassifierConfig; no PyO3 or Python API change.Additional context
The motivating setup is a fine-tuned judge in front of an interactive coding agent, where the session-level trigger (
new_session) was too coarse (one judgement per multi-job session) andevery_requesttoo fine (tier flips on tool turns).user_turnis the right trigger for that harness once the judge sees the current request as the task. I have a patch with tests and docs ready and will open a PR referencing this issue.