Learning project — built to explore LLMOps in practice: multi-agent orchestration, RAG pipelines, evaluation, and production deployment on GKE. The mountaineering domain is the vehicle; the LLMOps architecture is the point.
Mountain planning involves a lot of moving parts — weather, route conditions, gear, risk. CairnOps pulls them together into one API call.
You describe a climb in plain text. The system fetches real weather data, checks the route (GPX/topo where available), assembles a gear list, scores the risk, and returns a structured plan.
Example:
"Technical climb in Aladağlar, July, 3 days, above 3 200 m."
Output: 7-day forecast, route overview, equipment list, risk score (LOW / MODERATE / HIGH), day-by-day schedule. If risk is HIGH, the tone and recommendations change accordingly.
- Developers studying multi-agent LLM systems, RAG, and production MLOps — this is the primary audience. See
docs/roadmap.mdfor the full LLMOps breakdown. - Mountaineers and alpinists who want a quick planning baseline
- Outdoor guides building on top of the API
A request hits POST /api/v1/plan. The SupervisorAgent (LangGraph StateGraph) passes it through six domain agents in sequence:
- InputParser — extracts location, activity type, elevation, terrain (LLM @ temp 0.0)
- RouteAgent — uses GPX/topo data if available; otherwise LLM suggestion
- KnowledgeRetriever — queries Qdrant for relevant context (UIAA manuals, route guides)
- WeatherAgent — calls Open-Meteo for 7-day forecast; falls back to estimates if API down
- EquipmentAgent — rule-based base list + LLM enrichment using retrieved context
- SafetyAgent — deterministic risk score (0–9, five factors) + LLM explanation
Finally, PlanWriter generates either a standard plan (LOW/MODERATE) or a flagged high-risk plan (HIGH).
POST /api/v1/plan
│
▼
SupervisorAgent (LangGraph StateGraph)
│
├─► InputParser — text → structured fields
├─► RouteAgent — GPX/topo → LLM commentary
├─► KnowledgeRetriever — RAG query → context
├─► WeatherAgent — Open-Meteo → 7-day forecast
├─► EquipmentAgent — rules + RAG + LLM → gear list
├─► SafetyAgent — score + explanation
│
├─ risk HIGH ──► PlanWriter (hard tone, mandatory bail-out)
└─ otherwise ──► PlanWriter (standard tone)
│
▼
PlanResponse (JSON)
| Layer | Technology |
|---|---|
| API | FastAPI + Uvicorn (dev) / Gunicorn (prod) |
| Agent orchestration | LangGraph StateGraph with conditional routing |
| LLM Proxy | LiteLLM (abstracts away all the provider quirks) |
| LLM (models) | Ollama llama3.2 (local) / Groq (prod) — both routed through LiteLLM |
| Vector store | Qdrant — local Docker or Qdrant Cloud |
| Embeddings | nomic-embed-text via Ollama |
| Weather | Open-Meteo (free, no key) + Nominatim geocoding |
| Geospatial | GPX parsing, elevation lookup via OpenTopoData |
| Caching / checkpoints | Redis + langgraph-checkpoint-redis |
| Database | PostgreSQL (async SQLAlchemy) + Alembic migrations |
| Auth | JWT + bcrypt |
| Observability | Langfuse (LLM traces), MLflow (eval tracking), Prometheus (metrics), structlog |
| Evaluation | DeepEval (plan quality) + Ragas (retrieval quality) |
| Config | Pydantic Settings — fails fast on missing keys |
| Code quality | Ruff, Mypy (strict), pre-commit, pytest ≥70% coverage |
| Infra | Docker Compose (local), GKE + Terraform (prod) |
Requirements: Python 3.11+, Ollama, Docker.
# 1. Clone
git clone <repo-url>
cd CairnOps
# 2. Configure
cp .env.example .env
# Set JWT_SECRET_KEY — generate with:
python -c "import secrets; print(secrets.token_hex(32))"
# 3. Start backing services
docker-compose up -d qdrant redis postgres mlflow
# 4. Install dependencies
uv sync --all-extras
# 5. Pull models
ollama pull llama3.2
ollama pull nomic-embed-text
# 6. Run migrations
uv run alembic upgrade head
# 7. Ingest knowledge base
python -m cairnops.rag.ingester
# 8. Start API
cairnops-apiServer at http://localhost:8000 — Swagger at /docs.
Make a plan request:
curl -X POST http://localhost:8000/api/v1/plan \
-H "Content-Type: application/json" \
-d '{
"user_input": "Technical climbing in Aladağlar, July, 3 days.",
"elevation_m": 3200,
"terrain_type": "rocky ridge"
}'Run tests:
pytest # unit + API (no network, mocked LLM)
pytest -m integration # hits Open-Meteo, full pipeline
pytest -m rag # requires Qdrant + nomic-embed-text
pytest -m evaluation # full DeepEval + Ragas suite| Doc | What's in it |
|---|---|
docs/architecture.md |
Agent graph, state schema, LangGraph design |
docs/agents.md |
Per-agent breakdown — inputs, outputs, fallbacks |
docs/rag-pipeline.md |
Ingestion, retrieval, embedding, evaluation |
docs/evaluation.md |
DeepEval + Ragas setup, benchmark cases, CI integration |
docs/observability.md |
Langfuse, MLflow, Prometheus, structlog |
docs/api-reference.md |
Endpoint reference, request/response schemas |
docs/getting-started.md |
Full local setup walkthrough |
docs/development.md |
Dev tools, testing, code quality |
docs/terraform.md |
GCP infrastructure — two-stage Terraform |
docs/kubernetes.md |
GKE deployment, manifests, health probes |
Active development — Phase 5 complete (production deployment on GKE). Phase 6 covers prompt tuning, RAG expansion, and UI improvements.
Hilal Alpak — Apache-2.0