Auto-arbitrate between your Anthropic Claude Pro subscription tokens and OpenRouter fallback. When your Anthropic rate limit is hit, automatically route to OpenRouter (same Claude model). When the window resets, seamlessly return to Anthropic.
- 🔄 Auto-arbitration: Anthropic → OpenRouter → Anthropic as windows reset
- 📊 Proactive switching: Monitors token budget, switches before hitting hard 429
- 🔌 Drop-in OpenAI-compatible: Change only
base_urlin your code - 💾 Persistent state: Survives restarts without losing fallback context
- 🔍 Streaming support: Full SSE streaming for both providers
- 📈 Health endpoint: Monitor current provider state and token budget
pip install -e .cp .env.example .env
# Edit .env with your API keys
cat > .env << 'EOF'
ANTHROPIC_API_KEY=sk-ant-...
OPENROUTER_API_KEY=sk-or-...
EOFOptionally edit config.yaml to tune thresholds:
anthropic_model: claude-sonnet-4-6
openrouter_model: anthropic/claude-sonnet-4-6
listen_host: 127.0.0.1
listen_port: 8080
proactive_threshold: 0.05 # switch at 5% tokens remaining
recovery_hysteresis_seconds: 30 # wait 30s after reset before probingauto-arbiterThe proxy listens on http://127.0.0.1:8080.
import openai
# Only line that changes from normal OpenAI setup:
client = openai.OpenAI(
base_url="http://127.0.0.1:8080/v1",
api_key="ignored" # proxy doesn't validate this
)
response = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[
{"role": "user", "content": "Hello, Claude!"}
]
)
print(response.choices[0].message.content)stream = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[...],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [{"role": "user", "content": "What is 2+2?"}]
}'POST /v1/chat/completions— OpenAI-compatible chat endpointGET /v1/models— List available modelsGET /health— Current provider state and token budget
ANTHROPIC_ACTIVE
↓ (tokens < 5%)
ANTHROPIC_PROACTIVE
↓ (429/529 or tokens exhausted)
OPENROUTER_FALLBACK
↓ (window resets + 30s hysteresis)
RECOVERING (probe request)
↓ (success)
ANTHROPIC_ACTIVE (back to normal)
Rate limit state is saved to ~/.auto_arbiter_state.json. On restart:
- If the saved reset time hasn't passed yet, the proxy resumes in fallback mode
- If the reset time has passed, the proxy starts fresh on Anthropic
This prevents hammering Anthropic immediately after restart if the window isn't actually open.
- FastAPI proxy server on localhost:8080
- Anthropic: native SDK with raw response header extraction
- OpenRouter: httpx client (OpenAI-compatible)
- State machine: ANTHROPIC_ACTIVE → PROACTIVE → FALLBACK → RECOVERING → ANTHROPIC_ACTIVE
- Recovery watcher: background task probes Anthropic every 10s
ANTHROPIC_API_KEY— RequiredOPENROUTER_API_KEY— Required- Loaded from
.envfile in working directory
Stuck in fallback?
curl http://127.0.0.1:8080/healthCheck the tokens_reset time. If it's in the past, the recovery watcher should probe soon.
Wrong model?
Edit config.yaml and restart:
anthropic_model: claude-opus-4-6
openrouter_model: anthropic/claude-opus-4-6Debug logs?
Edit config.yaml:
debug_log: trueRun tests (when tests are added):
pytest tests/MIT