Unified model-agnostic local AI gateway with automatic failover, load balancing, and token accounting in pure PHP.
eidcloud, ai-gateway, llm-proxy, ollama, load-balancer, failover, php8
eidcloud-ai-gateway is an ultra-lightweight, zero-dependency, model-agnostic reverse proxy and routing gateway built in pure modern PHP (8.2+). It unifies local inference daemons (Ollama, llama.cpp) and commercial AI clouds (OpenAI, Anthropic Claude, Google Gemini) under a standard OpenAI-compatible API interface (/v1/chat/completions).
flowchart TD
Client["Client / Agent Application<br/>(OpenAI SDK, cURL, Python, etc.)"] -->|"POST /v1/chat/completions"| Gateway["🌐 EidCloud AI Gateway<br/>(PHP 8.2+ Zero Dependencies)"]
Gateway --> Router{"Provider Router & Load Balancer"}
Gateway --> Ledger[("💰 Token Ledger &<br/>Budget Accounting")]
Gateway --> Monitor["🩺 Health Check Engine"]
Router -->|"Primary Target"| Ollama["🦙 Ollama Local<br/>(qwen2.5 / llama3.2)"]
Router -->|"Failover Target 1"| LlamaCpp["⚡ llama.cpp Server<br/>(localhost:8080)"]
Router -->|"Failover Target 2"| OpenAI["☁️ OpenAI<br/>(gpt-4o / gpt-4o-mini)"]
Router -->|"Fallback Target 3"| Anthropic["☁️ Anthropic<br/>(claude-3-5-haiku)"]
Router -->|"Fallback Target 4"| Gemini["☁️ Google Gemini<br/>(gemini-1.5-flash)"]
Ollama -.->|"Fail (Timeout / 5xx)"| OpenAI
OpenAI -->|"Normalized OpenAI JSON"| Gateway
Gateway -->|"Unified Response"| Client
- Zero External Dependencies: 100% pure PHP 8.2+ without Composer bloat.
- Model-Agnostic OpenAI Compatibility: Route any client to Ollama, llama.cpp, Claude, or Gemini without changing your code.
- Automatic Failover: Transparently switch to cloud providers if your local Ollama or llama.cpp instance is offline or out of memory.
- Flexible Load Balancing: Supports sequential failover, round-robin, and weighted probability distribution.
- Model Alias Abstraction: Map generic model requests (e.g.,
coder,fast,general) to ordered provider cascades. - Token Accounting & Hard Budgeting: In-memory and file-persisted token usage tracking with automated spend caps.
- Streaming Support: SSE chunk streaming simulation compatible with standard client chat UI streaming.
- CLI & Diagnostic Suite: Built-in CLI commands for server hosting, real-time health checks, and JSON programmatic diagnostics.
Clone the repository into your environment:
git clone https://github.com/shadialhasan/eidcloud-ai-gateway.git
cd eidcloud-ai-gatewayNo external packages required. Just ensure PHP 8.2+ is installed:
php -vThe gateway includes an executable CLI tool:
php bin/eidcloud-gateway serve --port=8000# Formatted human terminal diagnostics
php bin/eidcloud-gateway check-health
# Programmatic JSON output
php bin/eidcloud-gateway check-health --jsonphp bin/eidcloud-gateway accounting
php bin/eidcloud-gateway accounting --jsonphp bin/eidcloud-gateway routes| Method | Endpoint | Description |
|---|---|---|
POST |
/v1/chat/completions |
Standard OpenAI-compatible completion with streaming & failover |
GET |
/v1/models |
List available gateway models and alias mappings |
GET |
/v1/accounting |
Real-time token usage, per-provider stats, and spend totals |
GET |
/health |
Live provider health checks with millisecond latency |
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "coder",
"messages": [
{"role": "user", "content": "Write a clean binary search in PHP."}
],
"temperature": 0.2
}'Run the zero-dependency test suite:
php tests/run_tests.phpAll 27 assertion tests verify routing, failover cascades, weighted and round-robin load distribution, token budgeting, and streaming chunk mechanics.
Eng. MHD. Shadi AL-Hasan
- Role: Executive CTO & Enterprise Solutions Architect
- Email: mhd.shadi.alhasan@gmail.com
- Phone / WhatsApp: +963934005922
- Location: Damascus, Syria
- GitHub: shadialhasan
This project is licensed under the MIT License - see the LICENSE file for details.
Copyright (c) 2026 MHD. Shadi AL-Hasan. All rights reserved.