Skip to content

feat(runner): configure Jev decision judges - #916

Merged
ayushag-nv merged 1 commit into
mainfrom
nachiketb/switch-1688-jev-config
Oct 6, 2026
Merged

ayushag-nv merged 1 commit into
mainfrom
nachiketb/switch-1688-jev-config

Conversation

@nachiketb-nvidia

@nachiketb-nvidia nachiketb-nvidia commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Closes SWITCH-1688.

The capability classifier can call decision models, but deployments cannot configure them. This wires Jev into the shared runner so server requests can use a decision judge, then call the selected LLM.

How

  • Add decision_clients for System One endpoints, environment-based credentials and deadlines, plus separate decision_targets.
  • Select relative-advantage judgment with routes.<name>.decision. Resolve candidate labels to LLM target IDs; pass evidence and optional instructions unchanged.
  • Register the decision client with the existing ClientRouter. Keep existing LLM settings, caller authentication, target prompts and capable fallback behavior.
  • Reject decision targets as answer destinations and reject incompatible judge settings at startup. Reuse typed URLs and credential redaction.
  • Include a configuration example. Both normal requests and /v1/decision use the same wiring; existing classifier stats and usage logging apply to normal requests.

Validation

  • All 49 existing runner configuration tests pass (cargo test -p switchyard-runner --locked config::).
  • All 8 targeted server tests pass (--test server classifier_ and --test server decision_).
  • Local HTTP smoke checks with the real server and mock providers: 21 configuration cases and 14 request scenarios. Verified both routing outcomes, cutoff equality, invalid answers, provider errors, deadlines, fail_open = false, decision-only routing, custom instructions, target prompts, credential separation and redaction. Also checked classifier latency, outcomes and token usage in stats and the routing log.
  • cargo fmt --all --check and workspace Clippy pass. No tests added or modified.

Live validation

Passed 5 live HTTP checks on 33d215449 using Jev (jev-latest) and the LiteLLM gateway configured in secrets/secrets.json:

  • Buffered answers from GPT-OSS 120B and 20B. Cutoffs 0.0 and 1.0 exercised both routing branches.
  • Streaming at cutoff 0.4 selected 20B and completed with text, usage and [DONE].
  • /v1/decision selected each tier with exactly one upstream call and no final LLM call.

All returned HTTP 200 with fail_open = false. Counters confirmed five Jev calls and three LLM calls. Server stats and the routing log recorded Jev latency, successful outcomes and token usage. Credentials stayed out of logs. These are integration checks, not routing-quality calibration.

Based directly on latest main; independent of #915.

Summary by CodeRabbit

  • New Features
    • Capability-based classification routes can now use a decision judge with configurable candidate labels, evidence, instructions, and cutoff.
    • Deployments can configure decision clients and targets for these routes. Without a decision judge, capability routes continue to use the existing LLM judge.
  • Documentation
    • Added configuration guidance, including how decision-judge errors and deadlines are handled and how to configure failure behavior.

Signed-off-by: nachiketb <nachiketb@nvidia.com>
@nachiketb-nvidia
nachiketb-nvidia requested a review from a team as a code owner October 6, 2026 00:28
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 314a08ec-c6b6-4e5f-8e6d-50e0221030d7
📥 Commits

Reviewing files that changed from the base of the PR and between db0a6bc and 33d2154.

📒 Files selected for processing (4)
  • crates/switchyard-runner/src/algorithm.rs
  • crates/switchyard-runner/src/config.rs
  • crates/switchyard-runner/src/lib.rs
  • crates/switchyard-server/CONFIGURATION.md

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


Walkthrough

Capability routes can now use a decision judge with configured candidates, evidence, and a cutoff. Deployment configuration adds decision clients and targets, validates their route use, and connects them to the router.

Changes

Decision Judge Routing

Layer / File(s) Summary
Capability decision judge
crates/switchyard-runner/src/algorithm.rs, crates/switchyard-runner/src/lib.rs
Capability route configuration accepts either a decision judge or the existing LLM judge. Decision settings are validated, and candidate target names are resolved to model IDs during algorithm construction.
Decision client configuration
crates/switchyard-runner/src/config.rs
Deployment configuration adds decision clients and targets. It constructs System One clients and shares API-key validation with LLM clients.
Target validation and routing
crates/switchyard-runner/src/config.rs, crates/switchyard-server/CONFIGURATION.md
Build validates target references and route target kinds, then passes decision clients to the router. The configuration guide documents the settings, route behavior, and failure handling.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Merge Risk: ⚪ Minimal · up to 33d21

The investigated configuration cases do not disrupt decision routing. No merge-blocking issue remains after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 15.38% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 13 functions across 3 files. (1 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the addition of Jev decision-judge configuration to the runner.
Full details: Docstring Coverage

Explanation

Docstring coverage is 15.38% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 13 functions across 3 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

I’m a rabbit with a config to read,
A decision judge joins the speed.
Candidate names find model IDs,
The router links clients as it bids.
With evidence set and cutoff in view,
I hop through the settings, neat and true.

Comment @coderabbitai help to get the list of available commands.

@ayushag-nv
ayushag-nv merged commit fdf6454 into main Oct 6, 2026
17 checks passed
@ayushag-nv
ayushag-nv deleted the nachiketb/switch-1688-jev-config branch October 6, 2026 00:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants