Skip to content

feat(menubar): change a route's algorithm and models from the menu bar - #875

Draft
elyasmnvidian wants to merge 3 commits into
grclark/mac-deamonfrom
emehtabuddin/switch-1633-menubar-route-picker
Draft

elyasmnvidian wants to merge 3 commits into
grclark/mac-deamonfrom
emehtabuddin/switch-1633-menubar-route-picker

Conversation

@elyasmnvidian

@elyasmnvidian elyasmnvidian commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #863. This PR targets grclark/mac-deamon, not main, so the diff shows only the menu bar changes.

What

Moving a route to other models meant editing the server config by hand. This PR adds a Change routing… item to the menu bar app. It opens a window where you pick a route, its algorithm, and a model for each role the algorithm needs. For composite, the roles are Judge, Capable, and Efficient. Apply checks the new config with switchyard-server --dry-run, saves it with a backup, and restarts the server.

Each model box lists the models from the LLM client's GET /models endpoint. Type to filter the list, or type any model ID. The app caches each model list in model-lists.json. After that, it fetches the list again only when you click Refresh models.

Why

Today the menu bar app can open the server config and restart the server, but it cannot change the config. To move a route to other models, you look up exact model IDs in the provider's model list, edit the TOML by hand, run switchyard-server --config <file> --dry-run, and restart the LaunchAgent with launchctl kickstart -k. Some mistakes show up only as a --dry-run error. For example, two targets in one route cannot use the same model ID on different LLM clients.

Example: a composite route uses a GPT judge and Claude for Capable and Efficient. To move it to GPT models, pick the route in the window, type sol in the Capable box, and pick gpt-5.6-sol. Do the same for Efficient, then click Apply.

Not in this PR: the feature request asked Apply to add prices for the new models when the prices are known. The app has no price source besides menubar.toml, so the result names each chosen model that has no price there, instead of guessing one. Savings stay hidden until you add the price and restart the menu bar app.

Notes for reviewers

Start with server_config.rs. ServerConfig::edit and choose_target decide which targets to keep, change, or copy. Then read server.rs::save_checked (check, backup, rename) and models.rs::load (cache and API keys). picker.rs is the AppKit window. It only collects choices, and it runs model lists, Keychain calls, and Apply on worker threads. The README describes the same behavior for users.

Existing menu items, the --print mode, and the server's routing do not change. restart and command moved from tray.rs to server.rs, so the window can use them too.

How it works

  • Target choice (server_config.rs). The edit goes through toml_edit, so comments and untouched tables stay as they were. For each role, the app keeps the route's target when it already names the chosen model. Otherwise it uses another target with that model and client, if the route would get the same system_prompt, reasoning_effort, extra_body, and omit_body_fields from it, or if the last three differ, because the server rejects two targets for one model on one client with different values for them. Otherwise the app changes the route's own target in place, or copies it when another route or an earlier role in the same Apply uses it.
  • Notes instead of fixes. --dry-run does not catch a setting that the new model rejects. So when a changed or copied target moves to another model family (for example from gpt-… to claude-…) or to a client with another format, the result lists the omit_body_fields, reasoning_effort, and extra_body that the target kept, and the app leaves them as they are. The result also lists the settings that an algorithm switch removed, and says when a role now shares a target with another route.
  • Check, save, and restart (server.rs). The app checks a temporary file next to the real file with switchyard-server --dry-run. It saves only when the check passes and the file did not change during the check. It then writes a backup that never replaces an older one, renames the temporary file over the real file, runs launchctl kickstart -k, and waits up to 10 seconds for /health.
  • API keys (models.rs). The app never writes a key to a file, and model-lists.json holds only model IDs and fetch times. To list a client's models, the app uses the key you just typed, then the client's api_key_env variable, then the login Keychain item for the client's base_url. Save key saves a key only when the models endpoint did not reject it. I chose the Keychain over reading your login shell's environment, because that runs your whole shell profile from the app and hangs if the profile waits for input. curl gets the key on stdin and runs with -q, so a verbose line in ~/.curlrc cannot print the key in the window.

Costs and limits

  • The lock file gains toml_edit 0.25. The app now uses security-framework, which was already in the lock file. The crate gains a test-only dependency on switchyard-runner, so tests can parse edited configs with the server's own parser.
  • A cached model list can go stale. The note shows the list's age, and Refresh models fetches it again.
  • A role's note shows three lines. Its tooltip shows the whole text.
  • A LaunchAgent does not load your shell profile. So for the check, a client with api_key_env needs that variable in the menu bar app's LaunchAgent, which puts the key in that plist file as plain text. The window names a missing variable under each role that needs it. A forward_auth client avoids this, because the server sends each caller's own key upstream.
  • Applying to a custom-mode llm_classifier route, or to a route whose type is not in the Algorithm list, replaces its settings with the chosen algorithm.
  • The family check compares the first word of each model ID's last path segment (gpt, claude, and so on). A provider that names models differently can get a note it does not need, or miss one.

Evidence

All runs used a debug build of this branch under temporary LaunchAgents, with a config in /tmp, clients on an OpenAI-compatible LiteLLM gateway that use api_key_env, and a local /v1/models stub for the API key cases. The gateway appears as https://gateway.example.com/v1, and its model IDs are replaced with public IDs. Automation could not click the status item on macOS 26, so a temporary code change, not in this PR, called Picker::show() at launch, which is the same call the menu item makes.

Apply on the final code

The test config had 14 routes, including a composite route (judge gpt-5.6-terra, Capable gpt-5.6-sol, Efficient gpt-5.6-luna, all on an openai_chat client) and a claude-opus-5-5 target with omit_body_fields = ["reasoning_effort"].

  1. GPT pair to Claude pair. I set Capable to claude-opus-5-5 and Efficient to claude-sonnet-5 and clicked Apply. The server's PID changed, and the result area showed:

    Saved /tmp/…/composite.toml. The old file is at /private/tmp/…/composite.toml.switchyard-backup.20261001105544.
    Restarted com.nvidia.switchyard.<test-label>-server, and the server answers http://127.0.0.1:18755/health.
    [targets.efficient] now names claude-sonnet-5 and has no omit_body_fields, reasoning_effort, or extra_body. Check whether the new model needs one of them.
    The switchyard-server --dry-run check passed for 14 routes.
    menubar.toml has no price for claude-opus-5-5. Savings stay hidden until you add one and restart the menu bar app.
    menubar.toml has no price for claude-sonnet-5. Savings stay hidden until you add one and restart the menu bar app.
    

    In the config diff, the id of [targets.efficient] changed to claude-sonnet-5, and capable_target changed to the existing claude-opus-5-5 target. A chat request to the route was answered by claude-sonnet-5 ("pong") after a judge call to gpt-5.6-terra.

  2. Back to the GPT pair. The result had the same kind of note for [targets.efficient] and said the check passed for 14 routes. cmp found the config byte-for-byte equal to the original, with the same permissions, and a chat request was answered by gpt-5.6-luna after a judge call.

  3. A failed check. With a client whose api_key_env variable the app does not have, Apply showed the server's error, then "The app's environment has no , which Apply needs: add it to the menu bar app's LaunchAgent." The config's checksum, the backup count, and the server's PID stayed the same, and no temporary file was left.

  4. Missing switchyard-server. The code gives "Not saved. Could not run …/switchyard-server: No such file or directory (os error 2). Apply needs switchyard-server next to the menu bar app. Run make install-macos from the Switchyard repository to install it." Only an earlier build was run without switchyard-server, and that Apply changed nothing.

Window behavior on the final code: a long role list, keyboard, Save key, a slow Refresh, and accessibility labels
  • A random route with 10 models: the window was 986 points tall on a screen whose visible area is about that height. Models 1–7 showed, the rest scrolled, and Refresh models, Close, and Apply stayed on the screen. Tab from Model 7 to Model 10 scrolled Model 10 into view.
  • Escape and Cmd-W each closed the window. With a model box's dropdown open, Return picked the highlighted model and did not run Apply.
  • Save key with a wrong key: the stub got one request with the wrong key, the result said "Did not save the key for http://127.0.0.1:…/v1, because the models endpoint rejected it.", and security find-generic-password found no item. With the right key, the app listed 5 models and saved the key. I deleted that item afterwards.
  • Refresh models with one URL that takes 12 seconds: 4 seconds later, the gateway's roles already showed their 249-model list, while the slow role still said "Refreshing…".
  • The accessibility tree names each control, for example AXPopUpButton (Capable LLM client), AXComboBox (Capable model), and AXTextField (API key for http://127.0.0.1:…/v1).
Model lists on an earlier build: Save key under launchd, the dropdown, Refresh, a failed refresh, restart, and the real gateway

An earlier build had a bug. When a refresh failed and model-lists.json did not have the list, the role note stayed at "Refreshing…" until the app restarted. The final code numbers each load, ignores a result from an older load, and never lets an older list replace a newer one in model-lists.json.

That run used clients stub (Responses) and stub_chat (Chat Completions) at a local /v1/models stub that counts requests and answers only one test key, plus a client for the gateway.

  1. No key. Each role showed the no-key note, the key field named the stub's base_url, and the stub counted 0 requests.
  2. Save key under launchd. Save key created the Keychain item, the stub counted 1 request for both clients, and every role showed the 10-model list. model-lists.json held only model IDs and the fetch time.
  3. The dropdown. With the Capable box cleared, the dropdown listed all 10 models, and clicking one put it in the box. After I typed SOL, the note said "2 of 10 models match."
  4. Refresh models. After the stub's list changed to 8 models, Refresh sent 1 more request, and every role and model-lists.json showed the new list.
  5. A failed refresh. With the stub stopped, each role kept its list and added the curl error.
  6. Restart. After launchctl kickstart -k, the window showed the cached lists, the stub still counted 2 requests, and model-lists.json was unchanged.
  7. The real gateway (2 requests). Save key listed 249 models, typing sonnet 5 gave "5 of 249 models match.", and Refresh fetched the list again.

Save key's order changed after that run: the final code lists the models first and saves only a key that the endpoint did not reject (see above). After the run I booted out the test LaunchAgent and deleted the Keychain items I had created.

Tests

cargo test -p switchyard-menubar runs 50 tests on macOS, and all pass. The new tests call the production functions with a local GET /models stub and a shell script in place of switchyard-server. They cover the model-list cache and API keys, the models-URL rules, the target rules and their notes, every algorithm's output against the server's own parser, and saving. Each new rule's test fails when that rule is removed.

Signed-off-by: Greg Clark <grclark@nvidia.com>

chore: cleanup

Signed-off-by: Greg Clark <grclark@nvidia.com>

chore: cleanup

Signed-off-by: Greg Clark <grclark@nvidia.com>
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-875/

Built to branch gh-pages at 2026-10-01 20:08 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@elyasmnvidian
elyasmnvidian force-pushed the emehtabuddin/switch-1633-menubar-route-picker branch from 2ad6add to a0c6783 Compare September 30, 2026 17:33
… check keys before saving them

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@messiaen
messiaen force-pushed the grclark/mac-deamon branch 3 times, most recently from c871c35 to 042211d Compare October 2, 2026 00:02

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants