From 7f1b303b8500f6ed583388c24385054ae65c64c6 Mon Sep 17 00:00:00 2001 From: Ryan Lempka Date: Mon, 28 Sep 2026 11:29:38 -0500 Subject: [PATCH] docs: clarify stats and composite configuration Signed-off-by: Ryan Lempka --- docs/reference/toml_schema.md | 8 +++++++- docs/routing_algorithms/escalation_router_routing.md | 5 +++-- 2 files changed, 10 insertions(+), 3 deletions(-) diff --git a/docs/reference/toml_schema.md b/docs/reference/toml_schema.md index 7fc96b709..aef6319cb 100644 --- a/docs/reference/toml_schema.md +++ b/docs/reference/toml_schema.md @@ -340,8 +340,13 @@ configuration. Today a classifier sets the tier a stage router falls open to whe |---|:---:|---|---| | `classifier.target` | Yes | — | Target the tier judge is called through. Not a routing destination. | | `classifier.base_threshold` | Yes | — | `p_solve` floor that still routes to the efficient tier. In `[0, 1]`. | +| `classifier.threshold_step` | No | `0.0` | Finite, non-negative amount added once for uncertain or unmatched verdicts and twice for unsupported verdicts. `base_threshold + 2 * threshold_step` must be at most `1`. | | `classifier.classify_trigger` | Yes | — | `user_turn` re-picks the tier whenever the user speaks, `new_session` picks once and holds it. `every_request` is rejected here: a judge call per tool step is the cost this route exists to avoid. | -| `classifier.message_hash_fallback` | No | `false` | Retains the tier by hashing the first user message, for clients that send no session ID. Unlike the `llm_classifier` route, this works on either trigger. Conversations opening with the same text share a tier. | +| `classifier.message_hash_fallback` | No | `false` | Retains the tier by hashing the first user message, for clients that send no session ID. Works with either supported trigger. Conversations opening with the same text share a tier. | +| `classifier.recent_turn_window` | No | unset | When unset, the judge sees the opening task and latest user follow-up, when present. When set, it also sees trailing turns. | +| `classifier.prompt` | No | packaged prompt | Replaces the classifier judge prompt. The verdict schema is unchanged. | +| `classifier.response_format_type` | No | `json_schema` | Structured-output mode for the judge. Use `json_object` when the provider does not support JSON Schema. Switchyard adds the schema to the prompt and validates the verdict locally. | +| `classifier.max_output_tokens` | No | `4096` | Maximum completion tokens for the judge verdict. Must be at least `1`. | | `stage.capable_target` | Yes | — | Capable tier. | | `stage.efficient_target` | Yes | — | Efficient tier. | | `stage.confidence_threshold` | Yes | — | Corroboration a decisive signal needs. In `[0, 1]`. | @@ -351,6 +356,7 @@ configuration. Today a classifier sets the tier a stage router falls open to whe | `stage.tool_semantics.mutate` | No | `[]` | Additional exact ASCII case-insensitive tool names that count as mutation. | | `stage.tool_semantics.plan` | No | `[]` | Additional exact ASCII case-insensitive tool names that count as planning. | | `stage.tool_semantics.new` | No | `[]` | Additional exact ASCII case-insensitive tool names that count as neutral forward activity. | +| `stage.handoff_notes` | No | unset | Optional guidance appended to forwarded requests. Uses the same fields and behavior as [stage-router handoff notes](../routing_algorithms/stage_router_routing.md#optional-handoff-notes). | | `subagents` | No | unset | Nested policy used only for delegated sub-agent work. | The tier is retained per session. A deployment that sends no session ID needs diff --git a/docs/routing_algorithms/escalation_router_routing.md b/docs/routing_algorithms/escalation_router_routing.md index 3fecddab6..734c01e04 100644 --- a/docs/routing_algorithms/escalation_router_routing.md +++ b/docs/routing_algorithms/escalation_router_routing.md @@ -153,9 +153,10 @@ Read the standard stats endpoint: curl -s http://localhost:4000/v1/stats ``` -The snapshot reports per-model calls, tokens, latency, and cost for the strong +The snapshot reports per-model calls, token usage, and latency for the strong and weak tiers. Judge calls are recorded in the classifier stats bucket, so their -token cost and latency remain visible as routing overhead. +token usage and latency remain visible as routing overhead. Dollar costs are not +included. Calculate them separately using the recorded usage and model pricing. When the server runs with a routing log, successful judge calls also appear in per-session routing stats under the judge's model id, tagged with the