Skip to content

[feature] make the unmatched verdict threshold step configurable (unmatched_steps) #845

Description

@himorishige

Problem

In capability mode, an unmatched verdict (no capability rule in the Card applied to the request) uses the same one-step threshold as uncertain. The judge still emits a p_solve, but there is no rule behind it, so the number carries less evidence than an uncertain verdict, which does name a rule. Operators who want unmatched requests to prove more before they go to the efficient tier have no setting for it short of raising base_threshold, which also cuts the supported-tier sends they want to keep.

We hit this running a fine-tuned judge in front of a coding agent with base_threshold = 0.75 and threshold_step = 0.1. A set of ten design-discussion requests (no code, no tools) came back unmatched from every judge variant we tried. Three of them carried p_solve = 0.85, exactly the one-step threshold, so they went to the efficient tier, where a blind review found technical errors in two of the answers. Retraining did not move this distribution, and adding a catch-all rule to the Card would change what the rules mean.

Proposed solution

A route-level key on capability-mode llm_classifier routes:

[routes.auto]
type = "llm_classifier"
base_threshold = 0.75
threshold_step = 0.1
unmatched_steps = 2   # 0, 1 or 2; default 1 (today's behavior)
verdict threshold
supported base_threshold
uncertain base_threshold + threshold_step
unmatched base_threshold + unmatched_steps * threshold_step
unsupported base_threshold + 2 * threshold_step

Values above 2 are rejected, so the existing base_threshold + 2 * threshold_step <= 1 check still bounds every threshold. Setting the key on an escalation or custom mode route is rejected like the other capability-only keys. Stage, composite, and the Python bindings keep the default.

With unmatched_steps = 2 in our setup, all ten design requests routed to the capable tier and an 87-request regression path stayed at 85/87 across three runs.

Alternatives considered

  • A hard policy that pins primary_rule = none to the capable tier: larger change, and it also pins requests the judge is clearly confident about (p_solve >= 0.95).
  • Raising base_threshold: moves the supported boundary too and reduces efficient-tier sends we want.
  • Adding an "unmatched" rule to the Card: changes the meaning of the rules and of unmatched itself.

Scope notes

  • Surface: routing algorithm (switchyard-libsy llm_classifier) and the deployment TOML parsed by switchyard-runner.
  • Adds one optional TOML key and one public u8 field on TaskClassifierConfig with a default; no PyO3 or Python API change.
  • Backward compatible: the default reproduces current behavior exactly.

Additional context

I have a patch with tests and docs ready and will open a PR referencing this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions