BrainRoute picks the right model for each request. It classifies the prompt, applies admission policy, scores every eligible local and cloud model, executes the winner with a fallback chain, records what happened, and feeds that back into the next decision.
Local by default. Cloud only when a local model genuinely cannot do the job — and it tells you when that happens.
It ships a browser UI, a CLI, and an OpenAI-compatible gateway, with no third-party Python dependencies.
Requires Python 3.11+.
python3 -m unittest discover -s tests -t . # hermetic: no providers, no network
python3 -m brainroute.cli eval # routing evals
python3 -m brainroute.cli route "Design a distributed ingestion pipeline" --explain
python3 -m brainroute.server # then open http://127.0.0.1:8765Nothing above calls a model provider. -t . matters on the test command: it lets the tests package redirect state to a temporary directory instead of writing telemetry into your real data/.
To actually execute a request, add --execute:
ollama serve && ollama pull qwen2.5:7b
python3 -m brainroute.cli ask "Write a short project tagline" --model local-qwen2-5 --executeA posture is a named stance on where data may go, so the common decision is one setting rather than four correlated flags.
| Posture | Cloud | Per-request opt-in | Use when |
|---|---|---|---|
open |
Competes freely | n/a | Capability and cost matter more than data locality |
local_first (default) |
Available, must earn it | n/a | Local by default, without losing hard tasks |
local_only |
Rejected by policy | Allowed | Nothing leaves the machine unless you say so per request |
air_gapped |
Rejected, discovery disabled | Refused | No outbound traffic except your own local runtime |
python3 -m brainroute.cli posture # show postures and the active one
python3 -m brainroute.cli posture local_only # switch
python3 -m brainroute.cli policy # resolved policy; * marks flags you pinnedTwo rules hold under every posture, including open:
- A prompt classified as sensitive never goes to an external provider.
allow_cloud_for_privateis off in every preset. - A per-request cloud opt-in lifts the blanket cloud block only. It never extends to sensitive prompts — opting in to cloud for a request is a different decision from consenting to send confidential data off the machine.
Under local_only a caller can escalate a single request explicitly, and the decision is flagged:
python3 -m brainroute.cli route "Design a distributed ingestion pipeline" --allow-cloudair_gapped refuses that opt-in outright. A posture a single API call can switch off is not a guarantee.
A posture supplies every policy flag you have not set yourself. Pin a flag explicitly and it survives posture switches — the UI marks pinned flags, brainroute policy marks them with *, and brainroute posture --reset-overrides hands control back.
This works because only your explicit values are written to data/settings.json:
{
"policy": { "posture": "local_only" },
"routing": { "capability_floor": { "high": 0.7 } }
}A local model that is cheap, fast, and private is still the wrong choice for work it cannot do. The capability floor enforces that. When the floor is what pushed a request off local hardware, the decision carries a warning rather than escalating silently:
Recommended: premium-reasoner (0.8214) [cloud]
! privacy_escalation: local-qwen2-5 would have won (0.9319 vs 0.8214) but fell
below the capability floor for this task, so it routes to premium-reasoner.
Lower routing.capability_floor.high to keep it local.
Every part of that is tunable: lower routing.capability_floor.high to keep hard work local at the cost of answer quality, raise policy.local_bonus to win more close races locally, or set routing.warn_on_privacy_escalation to false once you are happy with the trade.
The most effective fix is a bigger local model — see below.
prompt
-> classify task type, complexity, privacy, urgency, context size
-> policy hard constraints: context window, budget, privacy, circuit breakers
-> score task-modulated weights x (declared scores blended with telemetry)
-> execute selected model, then the ranked fallback chain
-> record latency, tokens, cost, success -> feeds the next decision
A keyword-and-phrase heuristic runs always and costs microseconds. Optionally a local model does structured classification against a JSON schema; results are cached, every field is validated against the heuristic, and any failure degrades to the heuristic rather than failing the request. Privacy and urgency signals the heuristic detected are never downgraded by the model — under-detecting either is the costlier error.
Policy decides what is eligible, not what is preferred. A model is rejected if it cannot hold the prompt, would breach a budget, would leak a sensitive request, or its provider circuit is open. Every rejection is reported with a reason. The active posture supplies the cloud-access rules.
If every model is rejected for reliability reasons, the unhealthy ones are re-admitted — a degraded answer beats no answer.
Five factors — quality, speed, cost, privacy, task fit — each 0-1.
- Weights are normalised, so profiles can use intuitive numbers that need not sum to 1, and scores stay comparable across profiles.
- Weights are then tilted by the task. A high-complexity request shifts weight toward quality and task fit and away from cost and speed; an urgent one toward speed; a sensitive one toward privacy. Weights are renormalised after tilting.
- Declared scores are blended with telemetry. Speed moves from the configured prior toward measured p95 latency as samples accumulate. Quality is corrected by thumbs up/down feedback once enough votes exist.
- Capability floors stop a cheap, fast, private model from winning a task it cannot do, via a penalty proportional to the shortfall. Floors and penalty are configurable under
routing. - The local bonus (
policy.local_bonus, default 0.15) applies whenprefer_localis on. Deliberately sized to win close races without overpowering the capability floor.
--explain prints the full arithmetic:
$ brainroute route "Design a complex production architecture for a routing layer" --explain
Recommended: premium-reasoner (0.8214) [cloud]
Fallbacks: gpt-5.4-mini -> local-qwen2-5
Task: reasoning / high (confidence 0.74, via heuristic)
Margin: +0.0869
Weights: quality=0.46 speed=0.11 cost=0.09 privacy=0.11 task_fit=0.22
Score breakdown:
premium-reasoner final=0.8214 base=0.7614
quality 0.950 x 0.465 = +0.4417
speed 0.550 x 0.111 = +0.0609
cost 0.345 x 0.092 = +0.0319
privacy 0.350 x 0.111 = +0.0387
task_fit 0.850 x 0.221 = +0.1882
capability_headroom +0.0600 comfortable headroom for a high-complexity task
local-qwen2-5 final=0.7069 base=0.7819
quality 0.650 x 0.465 = +0.3022
...
capability_floor -0.2250 below the quality floor (0.65 < 0.80) for a high-complexity task
policy_bonus +0.1500 policy prefers local providers
The selected model is tried first, then each model in the ranked fallback chain. Transient failures (429, 5xx, connection errors, timeouts) are retried with jittered exponential backoff; permanent ones (401, 400) are not. Streaming is never retried once bytes have reached the client — a silent retry would duplicate output.
Every attempt writes a run record: latency, token counts, estimated cost, success. Those records drive:
- Speed and quality blending — measured behaviour gradually overrides declared priors.
- Circuit breakers — after N consecutive failures a model is removed from routing, then half-opened after a cooldown to probe recovery.
- Budget enforcement — month-to-date spend is real, not estimated from a config file.
Optional epsilon-greedy exploration occasionally promotes a near-tied runner-up so the router keeps learning about models it would otherwise never pick. Off by default, so routing stays deterministic and evals reproducible.
brainroute reliability shows observed success rates, p95 latency, and circuit state.
Three provider adapters cover everything: ollama for local, openai for the Responses API, and openai_compatible for anything else.
Every model your local runtime reports becomes its own routing candidate with a distinct id, so qwen2.5:7b, qwen2.5-coder:14b, and llama3.3:70b compete with each other, not just with cloud. Capability is derived from parameter count, and models under 30B are marked weak for legal and medical work. Discovered models arrive disabled until you enable them.
This is the most effective way to keep work local: a 32B scores 0.824 and a 70B 0.85, both above the default 0.80 capability floor. Install one and hard tasks stop escalating.
id params quality speed ctx strengths
ollama-qwen2-5-7b 7.6 0.613 0.745 32k coding,general,private,summarization
ollama-qwen2-5-coder-14b 14.8 0.709 0.678 32k coding,general,private,summarization
ollama-qwen2-5-32b 32.8 0.824 0.598 32k coding,general,private,summarization
ollama-llama3-3-70b 70.6 0.850 0.520 128k general,private,summarization
Context windows are detected, not guessed. The tags endpoint does not report context length, so discovery asks the runtime for each model's real window (in parallel, best-effort, cached), falling back to a table of known families. This matters because an under-reported window escalates long prompts to cloud that a local model could have handled. The UI marks any window that is still an estimate.
Requests ask for the window they were promised. Ollama defaults num_ctx to about 4k regardless of what a model supports, and silently truncates beyond it — so correct detection alone would trade a visible escalation for an invisible quality loss. The adapter sizes num_ctx to what each request needs, in buckets, capped at the model's window:
| Prompt | Eligible models | num_ctx requested |
|---|---|---|
| 2k | 32B, 70B, cloud | default (no option sent) |
| 9k | 32B, 70B, cloud | 16384 |
| 20k | 32B, 70B, cloud | 32768 |
| 60k | 70B, cloud | 65536 |
Memory therefore scales with the prompt rather than the model's ceiling — a 5k prompt on a 128k model does not allocate a 128k KV cache. Set catalog.detect_local_context to false to skip the probe and rely on the family table alone.
The openai_compatible provider reaches anything speaking /v1/chat/completions — self-hosted vLLM or llama.cpp, LM Studio, a managed inference gateway, a corporate proxy, or another BrainRoute. Adding one is a config change, not new code:
- id: my-frontier-model
provider: openai_compatible
provider_label: my-gateway
model: vendor/some-model
endpoint: https://gateway.internal/v1/chat/completions
api_key_env: MY_GATEWAY_KEY
local: false
enabled: true
quality_score: 0.93
speed_score: 0.70
privacy_score: 0.35
input_price_per_mtok: 3.00
output_price_per_mtok: 15.00
max_context: 200000Use requires_api_key: false for endpoints with no credential, headers for static headers, and extra_body for endpoint-specific sampling knobs. Streaming, retries, fallback, health checks, and cost tracking work the same as for any other provider.
Set max_context — without it the context-window admission check cannot protect that model from an oversized prompt.
The seeded premium-reasoner is a placeholder of exactly this shape. It routes as a high-capability cloud option but cannot execute until you give it an endpoint, at which point it fails fast with a clear message rather than pretending to work.
None are configured by default. Any endpoint serving an OpenAI-style /v1/models list can be added under catalog.remote_sources; its models arrive as disabled candidates and nothing routes to them until you enable one.
{
"catalog": {
"remote_sources": [
{
"id": "my-gateway",
"enabled": true,
"models_url": "https://gateway.internal/v1/models",
"chat_url": "https://gateway.internal/v1/chat/completions",
"api_key_env": "MY_GATEWAY_KEY",
"limit": 25
}
]
}
}Set limit — a large catalog will otherwise import hundreds of models, inflating /api/config and the model picker. Discovered models take pricing and context length from the source when reported; where prices are reported, quality_score is inferred from price, a crude proxy worth reviewing before enabling anything for serious work.
Discovery is outbound traffic even though it carries no prompt data, so air_gapped blocks it as a hard constraint — even if a source is pinned enabled.
| File | Purpose |
|---|---|
| config/models.yaml | Model registry: capabilities, priors, prices, context windows |
| config/router.weights.yaml | Routing profiles |
| config/evals.yaml | Routing evals |
data/settings.json |
Only the settings you explicitly changed |
Profiles make trade-offs explicit:
balanced— quality, speed, cost, and privacy all consideredcheap— prefers low-cost modelsurgent— prefers low-latency modelsprivate— heavily prefers local models
cost_score is derived from prices, not hand-tuned. Set input_price_per_mtok and output_price_per_mtok (USD per 1M tokens); the router computes a cost score on a logarithmic curve, estimates per-request spend before committing, and enforces per-request and monthly budgets against real recorded spend. Prices for well-known models come from a built-in table; discovered models carry their own when the source reports it.
A model with no price signal scores mid-market (0.5) rather than free, so an unpriced model never wins the cheap profile by accident.
| Variable | Purpose |
|---|---|
OPENAI_API_KEY |
OpenAI credential |
per-model api_key_env |
Credential for any openai_compatible endpoint |
BRAINROUTE_API_KEY |
Require Authorization: Bearer ... on the gateway and state-changing admin endpoints |
BRAINROUTE_DATA_DIR |
Move SQLite state and settings off the checkout |
BRAINROUTE_CONFIG_DIR |
Move YAML config off the checkout |
At http://127.0.0.1:8765: enter a prompt, choose a profile, inspect the ranked decision and why each model scored as it did, execute or stream into a persisted conversation, compare two candidates side by side, discover and enable local models, switch privacy posture, tune capability floors and budgets, watch provider health and circuit state, run evals, and record good/bad feedback that feeds back into routing.
| Method | Path | Purpose |
|---|---|---|
| GET | /healthz |
Liveness |
| GET | /api/config |
Profiles, models, postures, settings |
| GET | /api/health |
Provider reachability (checked in parallel) |
| GET | /api/reliability |
Observed per-model behaviour and circuit state |
| GET | /api/catalog |
Discovered models |
| GET | /api/dashboard |
Aggregate run metrics |
| GET | /api/evals |
Run routing evals |
| GET | /api/telemetry, /api/sessions |
Event and conversation history |
| GET / POST | /api/postures |
Read or switch the privacy posture |
| POST | /api/route |
Route without executing |
| POST | /api/ask |
Route and optionally execute |
| POST | /api/chat/stream |
Server-sent event stream |
| POST | /api/compare |
Run two candidates side by side |
| POST | /api/feedback |
Record a good/bad rating |
| POST | /api/models, /api/settings |
Enable a model, change settings |
| POST | /api/maintenance/prune |
Trim telemetry to its retention cap |
| GET | /v1/models |
OpenAI-compatible model list |
| POST | /v1/chat/completions |
OpenAI-compatible completion, streaming supported |
/api/route and /api/ask accept either prompt or a full messages array, plus "allow_cloud": true to opt in to cloud for that request where the posture permits. Decisions return local, policy.posture, and a warnings array.
Point any OpenAI client at it:
curl http://127.0.0.1:8765/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"brainroute-auto","messages":[{"role":"user","content":"Write a release note."}],"stream":true}'Use "model": "brainroute-auto" to let the router choose, or name a model to pin it. Add "brainroute_profile": "cheap" to select a profile, or "brainroute_allow_cloud": true for a per-request cloud opt-in. Non-streaming responses include the full routing decision under a brainroute key.
config/evals.yaml holds routing evals that never call a provider. Beyond asserting which model wins, a case can pin a posture, assert the classification itself, require a specific warning, or require a minimum score margin — which catches the common failure where the router picks the right model for the wrong reason.
- name: local_first_escalates_hard_work_and_warns
prompt: Design a complex production architecture for a multi-provider AI routing layer.
posture: local_first
profile: balanced
expected_model: premium-reasoner
expected_warning: privacy_escalationpython3 -m brainroute.cli evalRun it after any change to weights, model scores, prices, or postures.
- Binding. Run behind a host firewall or reverse proxy when binding beyond loopback. The server warns if you do so without an API key set.
- Auth.
BRAINROUTE_API_KEYprotects the gateway and every state-changing admin endpoint — including posture changes, the most security-relevant write the API exposes — compared in constant time. Read-only endpoints stay open for the local UI; setgateway.protect_admin_apitofalseto opt out. - Rate limiting. Per-client fixed window over the API surface, configurable via
gateway.rate_limit_per_minute. Static assets are exempt. - State. SQLite in WAL mode with indexes and per-thread connections, falling back to a rollback journal on filesystems that cannot support WAL.
brainroute prunecaps table growth; the JSONL development log rotates at 16 MiB. - Caching. Provider discovery, assembled config, and classification are TTL-cached, so a routing decision does not depend on live network calls. Settings writes invalidate the caches immediately.
- Failure handling. Every HTTP handler runs inside an error boundary that returns JSON rather than dropping the connection or leaking a traceback.
- Credentials belong in the process environment, never in config files.