Skip to content

Architecture Overview

Mohsen Seyedkazemi Ardebili edited this page Aug 7, 2026 · 1 revision

Architecture Overview

How a plain-English question becomes a grounded answer — and, when needed, a human-approved action.

Current version is v4/. This page is the conceptual map; the authoritative, code-level reference is in v4/docs/.

The flow

You: "why is the api-server pod crashlooping?"
        │
        ▼
  Coordinator  ──fans out to──►  specialized agents
        │                          • pod / workload state (kubectl)
        │                          • metrics (Prometheus / PromQL)
        │                          • logs (Loki / LogQL)
        │                          • events
        ▼
  Correlate the evidence  ──►  a root-cause answer
        │
        ▼
  If you ask for a CHANGE (scale/restart/delete/patch):
        propose → server-side dry-run diff → ⛔ HUMAN APPROVAL (RBAC) → execute → verify

Read-only questions return immediately. Anything that mutates the cluster stops and waits for an explicit human approve, subject to role-based access control (admin / operator / readonly).

The pieces

  • Coordinator — routes a request to the agents that can answer it, then synthesizes.
  • Evidence agents — each owns one source of truth (cluster state, metrics, logs, events) and returns compact, structured results — never raw dumps.
  • Detectors / sensorium — declarative signals and failure playbooks (detect → investigate → remediate) that can fire without spending LLM tokens.
  • Memory — episodes plus a temporal knowledge graph, so the system can answer "what changed before this incident?"
  • The safety gate (HITL) — the invariant that makes acting safe. See The Safety Model.

Packaging

v4/ ships as a uv monorepo with three packages:

  • kubeintellect-server — the agent runtime and API
  • kube-q — the kq CLI you talk to
  • ki-protocol — the shared protocol/types between them

All generations run against one shared infrastructure stack (a Kind cluster, a Prometheus + Grafana + Loki observability stack, and a Langfuse instance) managed from the repo-root Makefile.

Design principle: earn every layer

plain LLM call
  → single agent + tools        (add only when the LLM alone fails)
    → multi-agent orchestration (add only for real specialization / parallelism / security separation)

If a capability can be a tool on an existing agent, it doesn't become a new agent. This keeps the system debuggable and cheap. Contributions are held to this bar — see CONTRIBUTING.md.

Go deeper

Clone this wiki locally