A lightweight, high-performance, rate-limiting, and header-spoofing reverse proxy designed for LLM APIs.
If your LLM gateway enforces strict rate limits (e.g., 429 Too Many Requests) or restricts API access to specific developer clients, this proxy lets you queue incoming requests in memory and inject required spoofed headers (such as specific User-Agent or custom authentication metadata) before forwarding them upstream.
- Token Bucket Rate Limiting with Reservation: Smooths out incoming request spikes. If you hit your rate limit, the proxy buffers and queues the request in memory rather than returning a
429error. - Client Disconnection Support: Automatically detects if a client aborts or times out while waiting in the queue. It cancels the queue slot and refunds the reserved token so it isn't wasted on upstream calls.
- Dynamic Header Injection: Inject any HTTP headers (such as
User-Agentor custom API keys/metadata) dynamically using simple environment variables. - OAuth Token Management: Read OAuth tokens from an
auth.jsonfile created by an external CLI tool. The proxy handles token refresh, expiry detection, and hot-reload via SIGHUP — no restart required. - Zero Dependencies: Written in standard Go, compiling down to a single self-contained binary.
This proxy acts as a centralized middleware layer between your downline services and upstream LLM providers, unlocking several key architectural patterns:
- Centralized API Key Management & Decoupling: Instead of distributing and rotating secret upstream provider keys across all downline applications, configure downline apps to use virtual/internal keys and let the proxy swap them centrally at the edge with
API_KEY_REPLACE. - Dynamic Routing & Model Version Upgrades: Avoid deploying configuration changes to multiple client apps when upgrading models. By configuring
MODEL_REPLACE(e.g. mappinggpt-3.5-turbotogpt-4o-mini), all downline requests are centrally and transparently mapped to the new model (updating request bodies and URL paths). - Client Identification & Header Spoofing: Some LLM gateways require specific HTTP headers (like a specific
User-Agent). The proxy lets you spoof these credentials centrally to bypass access restrictions. - OAuth Token Gateway: If you use CLI tools that authenticate via OAuth device flow (e.g., tools that store tokens in
auth.json), the proxy can read those tokens, refresh them before expiry, and inject them as upstream API keys. Downline services never need to know about OAuth — they just send a virtual key, and the proxy swaps it for the live OAuth token. SendSIGHUPto hot-reload a refreshed token without restarting. - Resiliency Against Hard Rate Limits (429s): The token-bucket rate limiter intercepts client requests and buffers/queues them in memory when limits are reached, gradually releasing them to fit upstream quotas instead of failing downstream calls with
429 Too Many Requests. - Central Audits & Cost Analysis: With all transaction details, response statuses, and raw bodies saved to a local SQLite database, you can centrally audit all LLM traffic, debug payloads, and compute usage costs.
Configure the proxy at runtime using the following environment variables:
| Variable | Description | Default |
|---|---|---|
PROXY_TARGET_URL |
Required. The upstream LLM API base URL to proxy requests to (e.g., https://api.openai.com or any custom gateway). |
None |
PROXY_PORT |
The port the proxy server will listen on. | 8318 |
RATE_LIMIT_RPM |
Requests per minute to allow. | 20 |
RATE_LIMIT_BURST |
The maximum burst capacity (tokens) allowed before queuing kicks in. | 5 |
HEADER_<NAME> |
Injects an HTTP header named NAME with the specified value. Single underscores are replaced with hyphens (e.g., HEADER_User_Agent maps to User-Agent). |
None |
INJECT_HEADERS_JSON |
A JSON-formatted string representing a key-value map of headers to inject (useful for complex headers). | None |
API_KEY_REPLACE |
Maps client API keys to upstream API keys. Replaces keys in standard headers (Authorization, api-key, x-api-key) and query parameters (key, api_key, api-key). Supports comma-separated format (e.g. client-key-1:upstream-key-1) or JSON format. |
None |
OAUTH_AUTH_PATH |
Path to an auth.json file created by an external CLI tool. Enables OAuth token management mode. When set, the proxy reads the token, injects it as the upstream API key, and optionally refreshes it before expiry. |
None |
OAUTH_TOKEN_URL |
OAuth token refresh endpoint. Required when OAUTH_AUTH_PATH is set. The proxy POSTs here with grant_type=refresh_token when the token is near expiry. |
None |
OAUTH_CLIENT_ID |
OAuth client ID sent in refresh requests. | None |
OAUTH_PROXY_TARGET_URL |
Overrides PROXY_TARGET_URL when OAuth mode is enabled. Useful when the OAuth provider also serves as the proxy target. |
None |
OAUTH_REFRESH_INTERVAL |
Background refresh interval in minutes. 0 = disabled (token only refreshed on startup if expired). |
0 |
OAUTH_EAGER_REFRESH_SECONDS |
How many seconds before expiry to trigger a refresh. | 300 |
OAUTH_FIELD_ACCESS |
JSON key for the access token in auth.json. |
access |
OAUTH_FIELD_REFRESH |
JSON key for the refresh token in auth.json. |
refresh |
OAUTH_FIELD_EXPIRES |
JSON key for the expiry timestamp in auth.json. |
expires |
Build and run the proxy locally:
# Set configuration env variables and run
export PROXY_TARGET_URL="https://your-upstream-api.com"
export RATE_LIMIT_RPM=20
export RATE_LIMIT_BURST=5
export HEADER_User_Agent="my-custom-client/1.0"
go run main.goBuild a minimal, multi-stage Docker container:
# Build the container
docker build -t gatepass .
# Run the container
docker run -d \
-p 8318:8318 \
-e PROXY_TARGET_URL="https://your-upstream-api.com" \
-e RATE_LIMIT_RPM=20 \
-e RATE_LIMIT_BURST=5 \
-e HEADER_User_Agent="my-custom-client/1.0" \
--name gatepass \
gatepassSome API gateways restrict access to official developer command-line tools by matching on specific headers (like User-Agent or version keys). You can bypass these restrictions by running this proxy to inject the expected headers.
export PROXY_TARGET_URL="https://api.upstream-service.com"
export PROXY_PORT="8318"
export RATE_LIMIT_RPM=20
export RATE_LIMIT_BURST=5
# Inject official client spoofing headers
export HEADER_User_Agent="official-cli-client/1.0.0"
export HEADER_X_Client_Version="1.0.0"
go run main.goPoint your client tool's base URL to the local proxy:
export UPSTREAM_BASE_URL="http://localhost:8318"
export UPSTREAM_API_KEY="your-api-key"
cli-tool-runIf you use a CLI tool that authenticates via OAuth (e.g., device code flow) and stores tokens in auth.json, the proxy can manage those tokens transparently.
# Use the tool's built-in login (one-time)
some-cli-tool account login
# This creates ~/.local/share/some-tool/auth.jsonexport OAUTH_AUTH_PATH="$HOME/.local/share/some-tool/auth.json"
export OAUTH_TOKEN_URL="https://example.com/auth/device/token"
export OAUTH_CLIENT_ID="my-cli-client"
export OAUTH_PROXY_TARGET_URL="https://example.com/zen/v1"
# Optional: enable background refresh every 30 minutes
export OAUTH_REFRESH_INTERVAL=30
./gatepassPoint downstream services at the proxy. They don't need to know about OAuth:
export ANTHROPIC_BASE_URL="http://localhost:8318"
export DEFAULT_MODEL="anthropic/claude-sonnet-4"If the CLI tool refreshes its token externally, reload without restarting:
kill -HUP $(pgrep gatepass){
"access": "tok_abc123",
"refresh": "ref_xyz789",
"expires": "2026-12-31T23:59:59Z"
}If your tool uses different field names, configure them:
export OAUTH_FIELD_ACCESS=token
export OAUTH_FIELD_REFRESH=refresh_token
export OAUTH_FIELD_EXPIRES=expiryTo run the compiled proxy binary as a background systemd service:
Create /etc/systemd/system/gatepass.service:
[Unit]
Description=Gatepass LLM Proxy
After=network.target
[Service]
ExecStart=/usr/local/bin/gatepass
Restart=always
RestartSec=3
Environment=PROXY_TARGET_URL=https://agentrouter.org
Environment=PROXY_PORT=8318
#Environment=RATE_LIMIT_RPM=20
#Environment=RATE_LIMIT_BURST=5
Environment=HEADER_Originator=codex_cli_rs
Environment="HEADER_User_Agent=codex_cli_rs/0.138.0 (Mac OS 26.0.1; arm64) Apple_Terminal/464"
Environment=HEADER_Version=0.138.0
# Environment=API_KEY_REPLACE="client-key:upstream-key"
[Install]
WantedBy=default.targetEnable and start the service:
sudo systemctl daemon-reload
sudo systemctl enable --now gatepassIf you are using Podman, you can manage the container lifecycle through systemd using a Quadlet file.
Create /etc/containers/systemd/gatepass.container (or ~/.config/containers/systemd/gatepass.container for rootless setup):
[Unit]
Description=Gatepass LLM Proxy Container
After=network.target
[Container]
Image=ghcr.io/deadrat-in/gatepass:latest
PublishPort=8318:8318
# For system-wide (runs as root, Podman will auto-create the directory):
Volume=/srv/gatepass/data:/data:Z
# For rootless (runs as user, use this instead so Podman can auto-create it in home):
# Volume=%h/.local/share/gatepass/data:/data:Z
Environment=PROXY_TARGET_URL=https://agentrouter.org
Environment=PROXY_PORT=8318
#Environment=RATE_LIMIT_RPM=20
#Environment=RATE_LIMIT_BURST=5
Environment=HEADER_Originator=codex_cli_rs
Environment="HEADER_User_Agent=codex_cli_rs/0.138.0 (Mac OS 26.0.1; arm64) Apple_Terminal/464"
Environment=HEADER_Version=0.138.0
# Environment=API_KEY_REPLACE="client-key:upstream-key"
[Service]
Restart=always
[Install]
WantedBy=default.targetReload systemd to generate the service unit and start it:
# For system-wide:
sudo systemctl daemon-reload
sudo systemctl enable --now gatepass
# For rootless (run without sudo):
systemctl --user daemon-reload
systemctl --user enable --now gatepassThis project is open-source and available under the GNU Affero General Public License v3.0 (AGPLv3).