OpenAI-compatible proxy for the tmiland-lab/opencode self-hosting fork.
One endpoint, our quotas (per-key requests/minute + daily budgets in local
sqlite), routed to free-tier provider keys + local Ollama. No third party
can throttle or reshape what we expose.
mkdir -p ~/.config/llm-gateway ~/.local/share/llm-gateway
cp config.example.json ~/.config/llm-gateway/config.json
python3 -c "import secrets; print(secrets.token_hex(32))" # -> client key
# put the key into config.json clients + chmod 600, then:
printf 'GROQ_API_KEY=\nCEREBRAS_API_KEY=\nOPENROUTER_API_KEY=\nGEMINI_API_KEY=\n' > ~/.config/llm-gateway/env
chmod 600 ~/.config/llm-gateway/config.json ~/.config/llm-gateway/env
python3 ~/llm-gateway/gateway.pysystemd (user): see llm-gateway.service — systemctl --user enable --now llm-gateway.
The gateway serves its own UI — no build step, no JS deps:
http://localhost:4143/ui— routes (ready/missing key), models, today's per-client usage, last 20 requests (auto-refresh 30s).
cp config.docker.json config/config.json # then put your client key in it
docker compose up --build -d # UI at http://localhost:4143/uiconfig/ is gitignored (keys stay local). The container reaches host
Ollama via host.docker.internal. Data (sqlite) lives in the gwdata
volume.
Then /models → gateway models, or "model": "gateway/groq/llama-3.3-70b-versatile".
- Groq: https://console.groq.com/keys
- Cerebras: https://cloud.cerebras.ai/ → API keys
- OpenRouter: https://openrouter.ai/keys (
:freemodels cost nothing) - Google AI Studio: https://aistudio.google.com/apikey
Paste into ~/.config/llm-gateway/env, restart the service. Routes
without a key answer 402 with the exact missing var — nothing breaks.
rpm/daily_requestsper client key;429with the spent budget.GET /v1/usage(gateway key): today's per-client counters.- sqlite:
~/.local/share/llm-gateway/usage.db, tablehits.