diff --git a/README.md b/README.md index 66cf717..76041f0 100644 --- a/README.md +++ b/README.md @@ -8,6 +8,7 @@

Drop a container. Your stack is monitored.
+ When something breaks, fixed rules catch it, the alert is pushed to you, and your AI agent investigates through the built-in MCP server.

Docker, Kubernetes, uptime, TLS, cron jobs, live logs, image updates, CVEs: auto-discovered, alerting on every one of them,
from a single Go binary that idles under 30 MB of RAM. No PromQL, no exporters, no dashboards to build.

@@ -20,7 +21,7 @@

- Quick Start  •  Why maintenant  •  Features  •  Documentation  •  Editions  •  Pricing + Quick Start  •  Why maintenant  •  Agents  •  Features  •  Documentation  •  Editions  •  Pricing

--- @@ -113,6 +114,7 @@ Against the tools usually stacked up next to it: | | maintenant | Uptime Kuma | Portainer | Dozzle | | ---------------------------- |:--------------:|:-----------:|:----------:|:----------:| | Container auto-discovery | **Yes** | No | Yes | Yes | +| MCP server, agent-ready | **Built in** | Third-party | Separate | No | | Live container logs | **Yes** | No | Yes | Yes | | HTTP/TCP endpoint checks | **Yes** | Yes | No | No | | Cron/heartbeat monitoring | **Yes** | Yes | No | No | @@ -211,6 +213,46 @@ Against the tools usually stacked up next to it: --- +## Agents + +maintenant watches, your AI agent investigates. The loop has three steps: + +1. **Detect.** Fixed rules decide that something is down: consecutive failures, thresholds, missed deadlines, restart loops. No model is involved, so an alert is never invented and nothing is spent while everything is fine. +2. **Push.** The alert leaves as a [webhook](https://docs.maintenant.dev/features/alerts/#webhook-payload): `alert.fired` when it starts, `alert.resolved` when it recovers, with the host it belongs to (`agent_id`) and a link back to it. +3. **Investigate.** Your agent receives the alert and queries maintenant's built-in MCP server: the container's logs, the endpoint's check history, CPU and memory, the other active alerts on the same host. You get a diagnosis, not just a red light. + +```mermaid +flowchart LR + S[containers, endpoints,
certificates, cron jobs, hosts] --> M[maintenant
rules fire an alert] + M -- webhook --> A[your agent
Claude Code, OpenCode, n8n] + A -- MCP: logs, history,
active alerts --> M +``` + +**With Claude Code.** A [small receiver](examples/agent-webhook/receiver.py) (Python, standard library only) runs on the host next to maintenant. Each fired alert starts a headless Claude Code session that can only call maintenant's read-only MCP tools, and the diagnosis lands in `reports/.md`: + +```bash +export RECEIVER_TOKEN=$(openssl rand -hex 32) +python3 examples/agent-webhook/receiver.py # listens on 127.0.0.1:9099 +``` + +```bash +# what the receiver runs for each alert +claude -p "maintenant just fired this alert: {...} Investigate it read-only." \ + --setting-sources project --mcp-config mcp.json --strict-mcp-config --tools "" \ + --allowedTools mcp__maintenant__list_alerts mcp__maintenant__get_container_logs \ + mcp__maintenant__get_endpoint_history mcp__maintenant__get_resources +``` + +`mcp.json` reaches maintenant over stdio with `docker exec -i maintenant /app/maintenant --mcp-stdio`, so no MCP port is opened. Point a webhook channel at the receiver with an `Authorization: Bearer ` header and route alerts to it with a trigger. The [AI agents guide](https://docs.maintenant.dev/guides/ai-agents/) walks through the setup. + +**With OpenCode**, the same receiver calls `opencode run` instead of `claude -p`, with maintenant declared as a local MCP server in `opencode.json`. **With n8n**, a Webhook trigger feeds an AI Agent node whose MCP Client tool points at your instance's `/mcp` endpoint. + +### [MCP server](https://docs.maintenant.dev/features/mcp/) + +Built-in [Model Context Protocol](https://modelcontextprotocol.io/) server with 51 tools. Ask your AI assistant what is burning, read a container's logs, check the alert queue, acknowledge an alert, open an incident. stdio and Streamable HTTP transports, OAuth2 with a client id and secret for remote clients (Claude web, mobile and Desktop). + +--- + ## Features Every section links to its full documentation. @@ -286,10 +328,6 @@ Channels: Discord and webhooks (Community), email and Telegram (Personal), Slack Real-time status page with severity aggregation across every monitor, live over SSE. **Personal** adds incident timelines, **Pro** adds email subscribers (double opt-in, through your own SMTP server), maintenance windows and branding. -### [MCP server](https://docs.maintenant.dev/features/mcp/) - -Built-in [Model Context Protocol](https://modelcontextprotocol.io/) server with 51 tools. Ask your AI assistant what is burning, read a container's logs, check the alert queue, acknowledge an alert, open an incident. stdio and Streamable HTTP transports, OAuth2 with a client id and secret for remote clients (Claude web, mobile and Desktop). - --- ## Configuration diff --git a/docs/api/reference.md b/docs/api/reference.md index 8a9b88c..13a42b1 100644 --- a/docs/api/reference.md +++ b/docs/api/reference.md @@ -434,7 +434,7 @@ All routes are open in every edition. - `GET /api/v1/alerts/active` answers `{ "critical": [...], "warning": [...], "info": [...] }`. Acknowledged and resolved alerts are left out. - `GET /api/v1/alerts/{id}` answers `404 NOT_FOUND` for an unknown alert. - A daily job purges the alerts that are not active (resolved or silenced) once they are older than 90 days, counted from their resolution, or from their creation when they never resolved. An alert that is still active is never purged, however old. -- An alert has `id`, `source`, `alert_type`, `severity` (`critical`, `warning` or `info`), `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `details` (a string holding JSON), `fired_at`, `resolved_at`, `resolved_by_id`, `acknowledged_at`, `acknowledged_by`, `escalated_at` (set when an escalation policy notifies a level for the alert) and `created_at`. +- An alert has `id`, `source`, `alert_type`, `severity` (`critical`, `warning` or `info`), `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `agent_id` (the host the alert belongs to: the agent that runs the container, endpoint, heartbeat or certificate, the disconnected agent itself, or `00000000-0000-0000-0000-000000000000` for the server's own runtime and for alerts with no host, such as the infrastructure security score), `details` (a string holding JSON), `fired_at`, `resolved_at`, `resolved_by_id`, `acknowledged_at`, `acknowledged_by`, `escalated_at` (set when an escalation policy notifies a level for the alert) and `created_at`. - `POST .../acknowledge` takes `{ "acknowledged_by": "..." }` (required). It answers the alert, emits `alert.acknowledged` and stops its running escalation. The dashboard, the MCP server and the security posture acknowledgments all go through this one path, so they behave the same. An acknowledged alert whose severity later rises does not start a new escalation. Errors: `400 INVALID_BODY`, `400 INVALID_REQUEST`, `404 NOT_FOUND` (unknown alert) and `409 CONFLICT` (the alert is not active or is already acknowledged). --- @@ -464,6 +464,7 @@ A channel has `id`, `name`, `type`, `url`, `headers` (a string holding a JSON ob - `POST` body: `name` (required), `url` (required), `type`, `headers`, `secret` (required for Telegram), `config` and `enabled` (default `true`). A URL must be HTTPS and must not resolve to a private or internal address, unless `MAINTENANT_ALLOW_PRIVATE_WEBHOOKS` is set. Answers `201` with the channel. - `PUT` takes the same fields, all optional. A secret cannot be cleared. A request that only sets `enabled` to `false` is always accepted, even after the edition dropped. +- A `webhook` channel receives `{ "event": "alert.fired" | "alert.resolved", "alert": {...}, "timestamp": "..." }`. `alert` holds the fields of the [`alert.fired` SSE event](#sse-event-stream), plus `url`, the link to the alert in the UI (`/alerts/history?alert=`). `url` is absent when the base URL has no host a recipient could open, as with the default built from a listen address such as `:8080` or `0.0.0.0:8080`. - `test` answers `404 NOT_FOUND` for an unknown channel. Otherwise it answers `200`: `{ "status": "delivered", "response_code": n }` or `{ "status": "failed", "error": "..." }`. - Creating, updating or testing a channel of a type the edition does not open answers `403 EDITION_REQUIRED` (`feature` is the capability: `slack`, `teams`, `smtp` or `telegram`). After a downgrade such a channel is `suspended`: it stops delivering and `GET /api/v1/edition` lists it under `suspended_channels`. - Errors: `400 INVALID_BODY`, `400 VALIDATION_ERROR`, `404 NOT_FOUND` and `409 DUPLICATE_NAME` on create. @@ -877,13 +878,13 @@ The SSE `event:` field is the event type, and `data:` holds the JSON payload onl | `certificate.created` | A certificate monitor is created or auto-detected | `monitor_id`, `hostname`, `port`, `source`, plus `server_name` and `agent_id` when they apply | | `certificate.check_completed` | A check finishes, including failed ones | `monitor_id`, `hostname`, `status`, `checked_at`, plus `subject_cn`, `issuer_cn`, `not_after`, `days_remaining`, `chain_valid`, `hostname_match` when known | | `certificate.status_changed` | A monitor changes status | `monitor_id`, `hostname`, `previous_status`, `new_status`, `days_remaining`, `timestamp` | -| `certificate.alert` | Expiry threshold, invalid chain, hostname mismatch, revoked OCSP response or expiry | `monitor_id`, `hostname`, `port`, `alert_type`, `severity`, `timestamp` and details | -| `certificate.recovery` | A certificate alert clears: the certificate was renewed, or the chain, hostname or OCSP problem is gone | `monitor_id`, `hostname`, `port`, `previous_alert_type` (`expiring`, `expired`, `chain_invalid`, `hostname_mismatch` or `ocsp_revoked`), `days_remaining`, `timestamp`, plus `new_not_after` and `server_name` when they apply | +| `certificate.alert` | Expiry threshold, invalid chain, hostname mismatch, revoked OCSP response or expiry | `monitor_id`, `hostname`, `port`, `alert_type`, `severity`, `timestamp`, `agent_id` and details | +| `certificate.recovery` | A certificate alert clears: the certificate was renewed, or the chain, hostname or OCSP problem is gone | `monitor_id`, `hostname`, `port`, `previous_alert_type` (`expiring`, `expired`, `chain_invalid`, `hostname_mismatch` or `ocsp_revoked`), `days_remaining`, `timestamp`, `agent_id`, plus `new_not_after` and `server_name` when they apply | | `certificate.deleted` | A certificate monitor is deleted | `monitor_id`, `hostname` | | `resource.snapshot` | A sample is stored (live samples only) | `container_id`, `cpu_percent`, `mem_used`, `mem_limit`, `mem_percent`, `net_rx_bytes`, `net_tx_bytes`, `block_read_bytes`, `block_write_bytes`, `timestamp`, `agent_id` | -| `resource.alert` | CPU or memory stays over its threshold (one event per metric) | `container_id`, `container_name`, `alert_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp` | -| `resource.recovery` | CPU or memory returns to normal (one event per metric) | `container_id`, `container_name`, `recovered_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp` | -| `alert.fired` | An alert is raised, or its severity escalates | the alert: `id`, `source`, `alert_type`, `severity`, `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `details` (an object), `fired_at`, `created_at` | +| `resource.alert` | CPU or memory stays over its threshold (one event per metric) | `container_id`, `container_name`, `alert_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp`, `agent_id` | +| `resource.recovery` | CPU or memory returns to normal (one event per metric) | `container_id`, `container_name`, `recovered_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp`, `agent_id` | +| `alert.fired` | An alert is raised, or its severity escalates | the alert: `id`, `source`, `alert_type`, `severity`, `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `agent_id`, `details` (an object), `fired_at`, `created_at` | | `alert.silenced` | An alert is raised while a silence rule or a maintenance window matches | same as `alert.fired` | | `alert.resolved` | An alert resolves | same as `alert.fired`, with `resolved_at` | | `alert.acknowledged` | An alert is acknowledged, from the REST route, the MCP server or a security posture acknowledgment | the alert as stored: `details` is a string holding JSON | @@ -898,8 +899,8 @@ The SSE `event:` field is the event type, and `data:` holds the JSON payload onl | `storage.availability_changed` | The database becomes unreachable, or answers again | `engine`, `connected` | | `update.scan_started` | A scan starts | `scan_id`, `started_at` | | `update.scan_completed` | A scan ends | `scan_id`, `updates_found`, `errors` | -| `update.detected` | A scan finds an update (on every scan, for each container) | `container_id`, `container_uid`, `container_name`, `image`, `current_tag`, `latest_tag`, `update_type`, `risk_score`, `alert_on`, plus `update_command`, `rollback_command` | -| `update.resolved` | An update is no longer pending | `container_id`, `container_uid`, `container_name` | +| `update.detected` | A scan finds an update (on every scan, for each container) | `container_id`, `container_uid`, `container_name`, `image`, `current_tag`, `latest_tag`, `update_type`, `risk_score`, `alert_on`, `agent_id`, plus `update_command`, `rollback_command` | +| `update.resolved` | An update is no longer pending | `container_id`, `container_uid`, `container_name`, `agent_id` (empty when the scan no longer sees the container) | | `security.insights_changed` | A container's insights change | `container_id`, `container_name`, `highest_severity`, `count`, `change` | | `security.insights_resolved` | All insights of a container are gone | `container_id`, `container_name` | | `security.posture_changed` | The posture score moved by 5 points or more, or changed colour (needs `MAINTENANT_SECURITY_SCORE_THRESHOLD`; evaluated at every scoring) | `score`, `previous_score`, `color` | diff --git a/docs/features/alerts.md b/docs/features/alerts.md index df1b869..a532873 100644 --- a/docs/features/alerts.md +++ b/docs/features/alerts.md @@ -158,13 +158,38 @@ POST /api/v1/channels } ``` -maintenant POSTs this JSON body, with `event` set to `alert.fired` or `alert.resolved` (`test` for a test notification): +#### Webhook payload + +maintenant sends one `POST` per notification, with `Content-Type: application/json` and the channel's own headers. Any 2xx answer counts as delivered; anything else, or no answer within 10 seconds, is retried as described in [Delivery and Retries](#delivery-and-retries). + +| Field | Type | Description | +|-------|------|-------------| +| `event` | string | `alert.fired`, `alert.resolved` or `test` | +| `timestamp` | string | When the notification was sent, RFC 3339 UTC | +| `alert.id` | string | UUID of the alert. The resolved notification carries the id of the alert it resolves. | +| `alert.agent_id` | string | UUID of the host the alert belongs to. The server's own runtime is `00000000-0000-0000-0000-000000000000`; remote hosts are listed by `GET /api/v1/agents`. | +| `alert.url` | string | Link to the alert in the UI, `/alerts/history?alert=`. Set `MAINTENANT_BASE_URL` to the address your team opens: without it, the link is built from the listen address. Absent when that address names no reachable host (`0.0.0.0`, `::` or an empty host), and on `test` notifications. | +| `alert.source` | string | `container`, `endpoint`, `heartbeat`, `certificate`, `resource`, `update`, `security`, `agent`, `host`, `swarm` or `kubernetes` (see [Alert Sources](#alert-sources)) | +| `alert.alert_type` | string | Source-specific type, such as `consecutive_failure` or `restart_loop` | +| `alert.severity` | string | `critical`, `warning` or `info` | +| `alert.status` | string | `active` on `alert.fired`, `resolved` on `alert.resolved` | +| `alert.message` | string | Human-readable description; on `alert.resolved`, the recovery message | +| `alert.entity_type`, `alert.entity_id`, `alert.entity_name` | string | The monitored object: its type, UUID and display name | +| `alert.fired_at` | string | When the condition was detected, RFC 3339 UTC | +| `alert.created_at` | string | When the alert was stored | +| `alert.details` | object | Source-specific values (target, failure count, threshold, last error...). Absent when the source has none. | +| `alert.resolved_at`, `alert.resolved_by_id` | string | On `alert.resolved` only: the recovery time and the UUID of the recovery record | +| `alert.acknowledged_at`, `alert.acknowledged_by`, `alert.escalated_at` | string | Present once the alert has been acknowledged or escalated | + +An alert that fires: ```json { "event": "alert.fired", "alert": { "id": "0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8", + "agent_id": "0198a0f4-1c2d-7e3f-8a4b-5c6d7e8f9a0b", + "url": "https://now.example.com/alerts/history?alert=0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8", "source": "endpoint", "alert_type": "consecutive_failure", "severity": "critical", @@ -186,7 +211,41 @@ maintenant POSTs this JSON body, with `event` set to `alert.fired` or `alert.res } ``` -A resolved notification carries the same alert with `status` set to `resolved`, a `resolved_at` time, the `resolved_by_id` of the recovery record, the `info` severity and the recovery message. The payload is maintenant's own: Slack and Teams expect theirs, which is why they are native channels. +The same alert when it recovers. The id, the host, the entity, `fired_at` and `details` are those of the original alert; the severity and the message come from the recovery: + +```json +{ + "event": "alert.resolved", + "alert": { + "id": "0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8", + "agent_id": "0198a0f4-1c2d-7e3f-8a4b-5c6d7e8f9a0b", + "url": "https://now.example.com/alerts/history?alert=0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8", + "source": "endpoint", + "alert_type": "consecutive_failure", + "severity": "info", + "status": "resolved", + "message": "Endpoint https://api.example.com/health recovered after 2 consecutive successes", + "entity_type": "endpoint", + "entity_id": "0198b1c2-8b4f-7a11-8d22-3e5f6071b8c9", + "entity_name": "api", + "fired_at": "2026-03-01T02:00:00Z", + "created_at": "2026-03-01T02:00:00Z", + "resolved_at": "2026-03-01T02:04:30Z", + "resolved_by_id": "0198b1c6-9d01-7c22-b733-4f6a7182c9da", + "details": { + "target": "https://api.example.com/health", + "failures": 3, + "threshold": 3, + "last_error": "connection refused" + } + }, + "timestamp": "2026-03-01T02:04:30Z" +} +``` + +A few cases send `alert.fired` again for the same `id`: a [severity raise](#life-of-an-alert) (the message starts with "Severity raised from X to Y") and each level of an [escalation policy](alert-escalation.md) that targets the channel. A receiver that must act once per alert keys on `id`. The **Test** button sends `event: "test"` with an `alert` whose `id`, `agent_id` and `entity_id` are empty and whose `source` and `alert_type` are `test`. + +The payload is maintenant's own: Slack and Teams expect theirs, which is why they are native channels. To hand alerts to an AI agent, see [AI Agents](../guides/ai-agents.md). ### Email (SMTP) :material-star-four-points:{ title="Personal" } diff --git a/docs/guides/ai-agents.md b/docs/guides/ai-agents.md new file mode 100644 index 0000000..7efdb92 --- /dev/null +++ b/docs/guides/ai-agents.md @@ -0,0 +1,136 @@ +# AI Agents + +maintenant decides that something is down, an AI agent finds out why. This guide connects the two: every fired alert is pushed to a small receiver, which starts a headless [Claude Code](https://docs.claude.com/en/docs/claude-code/overview) session that investigates through maintenant's [MCP server](../features/mcp.md) and writes a diagnosis. + +```mermaid +flowchart LR + M[maintenant
rules fire an alert] -- webhook --> R[receiver.py] + R -- claude -p --> A[Claude Code] + A -- MCP over stdio:
logs, history, active alerts --> M + A --> D[reports/alert-id.md] +``` + +1. **Detect.** Alerts come from fixed rules (consecutive failures, thresholds, deadlines). No model decides whether something is down, and nothing runs while everything is fine. +2. **Push.** The alert leaves as a [webhook](../features/alerts.md#webhook-payload) carrying the host it belongs to (`agent_id`) and a link back to it (`url`). +3. **Investigate.** The agent reads what it needs through the MCP tools, read-only. + +--- + +## Requirements + +- maintenant running in a Docker container, or installed natively. `mcp.json` calls the container `maintenant`: if `docker ps` shows another name (Compose names it `-maintenant-1` unless `container_name` is set), put yours in its place. +- On the same host: Python 3 and [Claude Code](https://docs.claude.com/en/docs/claude-code/setup), logged in once interactively by the user that will run the receiver. +- A webhook channel, available in every edition. + +The receiver and its MCP configuration live in [`examples/agent-webhook/`](https://github.com/kolapsis/maintenant/tree/main/examples/agent-webhook) in the repository. + +--- + +## 1. Start the receiver + +```bash +git clone --depth 1 https://github.com/kolapsis/maintenant.git +cd maintenant/examples/agent-webhook +export RECEIVER_TOKEN=$(openssl rand -hex 32) +python3 receiver.py +``` + +| Variable | Default | Description | +|----------|---------|-------------| +| `RECEIVER_TOKEN` | (required) | Shared secret. Requests without `Authorization: Bearer ` get `401`. | +| `RECEIVER_LISTEN` | `127.0.0.1:9099` | Address and port to listen on | +| `RECEIVER_MCP_CONFIG` | `mcp.json` next to the script | MCP configuration passed to Claude Code | +| `RECEIVER_REPORTS` | `reports/` next to the script | Where diagnoses are written, one file per alert | +| `RECEIVER_TIMEOUT` | `600` | Seconds an investigation may run | + +The receiver answers `202` at once and queues the alert: investigations run one at a time, so a burst of alerts never starts a burst of sessions. `alert.resolved` and `test` notifications are logged and not investigated. + +For each fired alert it runs: + +```bash +claude -p "" \ + --setting-sources project \ + --mcp-config mcp.json --strict-mcp-config \ + --tools "" \ + --allowedTools mcp__maintenant__list_alerts mcp__maintenant__get_container_logs ... +``` + +`--setting-sources project` keeps the user's own settings and hooks out of the session (it runs in `reports/`, which has no project settings), `--tools ""` removes every built-in tool (no shell, no file access), `--strict-mcp-config` ignores any other MCP server configured for that user, and `--allowedTools` lists the read-only maintenant tools the session may call: `list_alerts`, `list_agents`, `list_containers`, `get_container`, `get_container_logs`, `get_resources`, `get_top_consumers`, `list_endpoints`, `get_endpoint_history`, `list_heartbeats`, `list_certificates`, `get_updates`, `get_security_insights`, `get_health`. Add `acknowledge_alert` to `READ_TOOLS` in `receiver.py` if you want the agent to acknowledge what it has diagnosed. + +`mcp.json` reaches maintenant over stdio, inside its own container, so no MCP port is opened: + +```json +{ + "mcpServers": { + "maintenant": { + "command": "docker", + "args": ["exec", "-i", "maintenant", "/app/maintenant", "--mcp-stdio"] + } + } +} +``` + +On a native install, call the binary directly: `"command": "maintenant", "args": ["--mcp-stdio"]`, with `MAINTENANT_DB` in `env` when the database is not at its default path. See [MCP Server](../features/mcp.md#authentication) for what the stdio transport can reach. + +Run the receiver under your service manager (systemd, a supervisor, a container with the Docker CLI and Claude Code) once it works by hand. + +--- + +## 2. Let maintenant reach it + +Webhook channels refuse plain `http` and private addresses, at save time and at every connection (see [the SSRF guard](../features/alerts.md#webhook-destinations-and-the-ssrf-guard)). Two ways through: + +- **Behind your reverse proxy** (recommended). Publish the receiver at a public HTTPS name, for example `https://hooks.example.com/maintenant`, proxied to `127.0.0.1:9099`. Keep `RECEIVER_LISTEN` on loopback: only the proxy talks to it, and the bearer token protects it from the rest of the internet. +- **On a private host**, set `MAINTENANT_ALLOW_PRIVATE_WEBHOOKS=true` on maintenant and point the channel at the Docker bridge gateway, for example `http://172.17.0.1:9099/`, with `RECEIVER_LISTEN=172.17.0.1:9099`. This lifts the guard for every channel and event webhook, not only this one. + +--- + +## 3. Create the channel and route alerts to it + +In **Alerts → Channels**, add a **Webhook** channel with the receiver's URL and the header `Authorization: Bearer `, then press **Test**: the receiver logs a `test` line. Through the API: + +```bash +POST /api/v1/channels +{ + "name": "claude-code", + "type": "webhook", + "url": "https://hooks.example.com/maintenant", + "headers": "{\"Authorization\": \"Bearer \"}" +} +``` + +Channels are silent until a [trigger](../features/alerts.md#alert-triggers) routes alerts to them. Start narrow, critical alerts only, so the agent works on what matters: + +```bash +POST /api/v1/alert-triggers +{ + "name": "Critical alerts to Claude Code", + "filter_severities": "critical", + "enabled": true, + "notify_on_resolve": false, + "channel_ids": [""] +} +``` + +--- + +## 4. Read the diagnosis + +When the next critical alert fires, the receiver logs it and, once the session ends, writes `reports/.md`: what is failing, the most likely cause, the log lines it rests on, and the fix to try first. The alert's `url` takes you to it in maintenant. + +A severity raise and each escalation level send `alert.fired` again for the same alert id, so the same alert can be investigated more than once. + +--- + +## Other agents + +- **OpenCode.** Replace the `claude` command in `receiver.py` with `opencode run ""`, and declare maintenant as a local MCP server in `opencode.json` with the same `docker exec -i maintenant /app/maintenant --mcp-stdio` command. +- **n8n.** A **Webhook** trigger node receives the payload and feeds an **AI Agent** node whose **MCP Client** tool points at your instance's Streamable HTTP endpoint, `https:///mcp`, authenticated with the OAuth2 client described in [MCP Server](../features/mcp.md#streamable-http-oauth2). + +--- + +## Related + +- [Webhook payload](../features/alerts.md#webhook-payload): every field, fired and resolved examples +- [MCP Server](../features/mcp.md): the 51 tools, transports and authentication +- [Alert Engine](../features/alerts.md): sources, triggers, delivery and retries diff --git a/examples/agent-webhook/mcp.json b/examples/agent-webhook/mcp.json new file mode 100644 index 0000000..5398b87 --- /dev/null +++ b/examples/agent-webhook/mcp.json @@ -0,0 +1,8 @@ +{ + "mcpServers": { + "maintenant": { + "command": "docker", + "args": ["exec", "-i", "maintenant", "/app/maintenant", "--mcp-stdio"] + } + } +} diff --git a/examples/agent-webhook/receiver.py b/examples/agent-webhook/receiver.py new file mode 100644 index 0000000..d7af553 --- /dev/null +++ b/examples/agent-webhook/receiver.py @@ -0,0 +1,109 @@ +#!/usr/bin/env python3 +"""Receive maintenant alert webhooks and hand each fired alert to Claude Code.""" + +import hmac +import json +import os +import pathlib +import queue +import subprocess +import sys +import threading +from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer + +HERE = pathlib.Path(__file__).resolve().parent +LISTEN = os.environ.get("RECEIVER_LISTEN", "127.0.0.1:9099") +TOKEN = os.environ.get("RECEIVER_TOKEN", "") +MCP_CONFIG = os.environ.get("RECEIVER_MCP_CONFIG", str(HERE / "mcp.json")) +REPORTS = pathlib.Path(os.environ.get("RECEIVER_REPORTS", HERE / "reports")) +TIMEOUT = int(os.environ.get("RECEIVER_TIMEOUT", "600")) + +READ_TOOLS = [ + "list_alerts", + "list_agents", + "list_containers", + "get_container", + "get_container_logs", + "get_resources", + "get_top_consumers", + "list_endpoints", + "get_endpoint_history", + "list_heartbeats", + "list_certificates", + "get_updates", + "get_security_insights", + "get_health", +] + +PROMPT = """maintenant just fired this alert: + +{alert} + +Investigate it with the maintenant MCP tools, read-only. Start from the entity +in the alert, read its logs and recent history, check the other active alerts +on the same agent, and look for what changed shortly before fired_at. + +Answer in under 15 lines: what is failing, the most likely cause, the evidence +you saw (quote log lines), and the fix you would try first. +""" + +jobs = queue.Queue() + + +def investigate(alert): + REPORTS.mkdir(parents=True, exist_ok=True) + cmd = [ + "claude", "-p", PROMPT.format(alert=json.dumps(alert, indent=2)), + "--setting-sources", "project", + "--mcp-config", MCP_CONFIG, + "--strict-mcp-config", + "--tools", "", + "--allowedTools", *[f"mcp__maintenant__{t}" for t in READ_TOOLS], + ] + try: + out = subprocess.run(cmd, cwd=REPORTS, capture_output=True, text=True, timeout=TIMEOUT) + report = out.stdout if out.returncode == 0 else f"claude exited {out.returncode}\n{out.stderr}" + except subprocess.TimeoutExpired: + report = f"investigation timed out after {TIMEOUT}s" + path = REPORTS / f"{alert['id']}.md" + path.write_text(f"# {alert.get('message', alert['id'])}\n\n{report}\n") + print(f"report written to {path}", flush=True) + + +def worker(): + while True: + alert = jobs.get() + try: + investigate(alert) + except Exception as exc: + print(f"investigation of {alert.get('id')} failed: {exc}", file=sys.stderr, flush=True) + jobs.task_done() + + +class Handler(BaseHTTPRequestHandler): + def do_POST(self): + if TOKEN and not hmac.compare_digest(self.headers.get("Authorization", ""), f"Bearer {TOKEN}"): + self.send_response(401) + self.end_headers() + return + try: + body = json.loads(self.rfile.read(int(self.headers.get("Content-Length", 0)))) + except ValueError: + self.send_response(400) + self.end_headers() + return + self.send_response(202) + self.end_headers() + event, alert = body.get("event"), body.get("alert") or {} + print(f"{event} {alert.get('severity', '')} {alert.get('message', '')}", flush=True) + if event == "alert.fired": + jobs.put(alert) + + +if __name__ == "__main__": + if not TOKEN: + sys.exit("set RECEIVER_TOKEN, and send it from the channel as 'Authorization: Bearer '") + threading.Thread(target=worker, daemon=True).start() + host, port = LISTEN.rsplit(":", 1) + print(f"listening on {LISTEN}", flush=True) + ThreadingHTTPServer((host, int(port)), Handler).serve_forever() diff --git a/frontend/src/components/AlertList.vue b/frontend/src/components/AlertList.vue index f70bbb2..fd19532 100644 --- a/frontend/src/components/AlertList.vue +++ b/frontend/src/components/AlertList.vue @@ -4,18 +4,71 @@ --> + + diff --git a/frontend/src/components/__tests__/AlertList.spec.ts b/frontend/src/components/__tests__/AlertList.spec.ts index 9c4371b..bae1e81 100644 --- a/frontend/src/components/__tests__/AlertList.spec.ts +++ b/frontend/src/components/__tests__/AlertList.spec.ts @@ -8,6 +8,18 @@ import AlertList from '@/components/AlertList.vue' import { useAlertsStore } from '@/stores/alerts' import { detailSlideOverKey } from '@/composables/useDetailSlideOver' import type { Alert } from '@/services/alertApi' +import { LOCAL_AGENT } from '@/services/apiFetch' + +const sse = vi.hoisted(() => new Map void>()) +vi.mock('@/services/sseBus', () => ({ + sseBus: { + on: (type: string, fn: (e: MessageEvent) => void) => sse.set(type, fn), + off: (type: string) => sse.delete(type), + connect: () => {}, + disconnect: () => {}, + connected: false, + }, +})) function alertAt(id: string, firedAt: string): Alert { return { @@ -20,14 +32,132 @@ function alertAt(id: string, firedAt: string): Alert { entity_type: 'container', entity_id: id, entity_name: id, + agent_id: LOCAL_AGENT, fired_at: firedAt, created_at: firedAt, } } +function jsonResponse(body: unknown, status = 200): Response { + return new Response(JSON.stringify(body), { status, headers: { 'content-type': 'application/json' } }) +} + +function mountList(props: { linkedId?: string } = {}) { + return mount(AlertList, { + props, + attachTo: document.body, + global: { + provide: { [detailSlideOverKey as symbol]: { openDetail: vi.fn() } }, + stubs: { AcknowledgeButton: true }, + }, + }) +} + describe('AlertList', () => { afterEach(() => { vi.unstubAllGlobals() + document.body.innerHTML = '' + }) + + it('highlights the linked alert when it is in the loaded page', async () => { + const fetchMock = vi.fn().mockImplementation(async () => jsonResponse(alertAt('a1', '2026-09-30T10:00:00Z'))) + vi.stubGlobal('fetch', fetchMock) + setActivePinia(createPinia()) + const store = useAlertsStore() + store.alerts = [alertAt('a2', '2026-09-30T11:00:00Z'), alertAt('a1', '2026-09-30T10:00:00Z')] + + const wrapper = mountList({ linkedId: 'a1' }) + await flushPromises() + + const rows = wrapper.findAll('tr[data-alert-id]') + expect(rows.map((r) => r.attributes('data-alert-id'))).toEqual(['a2', 'a1']) + expect(rows[1]!.classes()).toContain('alert-linked') + expect(rows[0]!.classes()).not.toContain('alert-linked') + }) + + it('fetches the linked alert by id and shows it above a page that does not hold it', async () => { + const fetchMock = vi.fn().mockImplementation(async () => jsonResponse(alertAt('old', '2026-08-01T10:00:00Z'))) + vi.stubGlobal('fetch', fetchMock) + setActivePinia(createPinia()) + const store = useAlertsStore() + store.alerts = [alertAt('a2', '2026-09-30T11:00:00Z')] + + const wrapper = mountList({ linkedId: 'old' }) + await flushPromises() + + expect(String(fetchMock.mock.calls[0]![0])).toMatch(/\/alerts\/old$/) + const rows = wrapper.findAll('tr[data-alert-id]') + expect(rows.map((r) => r.attributes('data-alert-id'))).toEqual(['old', 'a2']) + expect(rows[0]!.classes()).toContain('alert-linked') + }) + + it('applies live updates to a linked alert fetched outside the loaded page', async () => { + const old = alertAt('old', '2026-08-01T10:00:00Z') + vi.stubGlobal('fetch', vi.fn().mockImplementation(async () => jsonResponse(old))) + setActivePinia(createPinia()) + const store = useAlertsStore() + store.alerts = [alertAt('a2', '2026-09-30T11:00:00Z')] + store.connectSSE() + + const wrapper = mountList({ linkedId: 'old' }) + await flushPromises() + + sse.get('alert.acknowledged')!({ data: JSON.stringify({ ...old, acknowledged_at: '2026-10-01T09:00:00Z', acknowledged_by: 'ops' }) } as MessageEvent) + expect(store.alerts.find((a) => a.id === 'old')!.acknowledged_at).toBe('2026-10-01T09:00:00Z') + + sse.get('alert.resolved')!({ data: JSON.stringify({ ...old, status: 'resolved', resolved_at: '2026-10-01T09:05:00Z' }) } as MessageEvent) + await flushPromises() + const row = wrapper.find('tr[data-alert-id="old"]') + expect(row.classes()).toContain('alert-linked') + expect(row.text()).toContain('resolved') + store.disconnectSSE() + }) + + it('updates a linked alert outside the loaded page when it is acknowledged from its row', async () => { + const old = alertAt('old', '2026-08-01T10:00:00Z') + const acked = { ...old, acknowledged_at: '2026-10-01T09:00:00Z', acknowledged_by: 'maintenant-ui' } + const fetchMock = vi.fn().mockImplementation(async (url: string) => jsonResponse(String(url).endsWith('/acknowledge') ? acked : old)) + vi.stubGlobal('fetch', fetchMock) + setActivePinia(createPinia()) + const store = useAlertsStore() + store.alerts = [alertAt('a2', '2026-09-30T11:00:00Z')] + + mountList({ linkedId: 'old' }) + await flushPromises() + await store.acknowledgeAlert('old') + + expect(store.alerts.find((a) => a.id === 'old')!.acknowledged_at).toBe('2026-10-01T09:00:00Z') + }) + + it('keeps the linked alert when the first page arrives after it, and drops it on a filter change', async () => { + const old = alertAt('old', '2026-08-01T10:00:00Z') + const page = { alerts: [alertAt('a2', '2026-09-30T11:00:00Z')], has_more: false } + vi.stubGlobal('fetch', vi.fn().mockImplementation(async (url: string) => jsonResponse(String(url).includes('/alerts/old') ? old : page))) + setActivePinia(createPinia()) + const store = useAlertsStore() + + const wrapper = mountList({ linkedId: 'old' }) + await flushPromises() + await store.fetchAlerts() + expect(store.alerts.map((a) => a.id)).toEqual(['old', 'a2']) + + await wrapper.findAllComponents({ name: 'SelectInput' })[0]!.vm.$emit('update:modelValue', 'endpoint') + await flushPromises() + expect(store.alerts.map((a) => a.id)).toEqual(['a2']) + }) + + it('says so when the linked alert no longer exists', async () => { + const fetchMock = vi.fn().mockImplementation(async () => + jsonResponse({ error: { code: 'NOT_FOUND', message: 'alert not found' } }, 404), + ) + vi.stubGlobal('fetch', fetchMock) + setActivePinia(createPinia()) + + const wrapper = mountList({ linkedId: 'gone' }) + await flushPromises() + + expect(wrapper.text()).toContain('Linked alert not found') + expect(wrapper.text()).toContain('This alert no longer exists') }) it('pages from the last alert listed, not only from its second', async () => { diff --git a/frontend/src/pages/AlertsPage.vue b/frontend/src/pages/AlertsPage.vue index f69760a..952f7cd 100644 --- a/frontend/src/pages/AlertsPage.vue +++ b/frontend/src/pages/AlertsPage.vue @@ -23,8 +23,14 @@ const router = useRouter() const store = useAlertsStore() const triggersStore = useTriggersStore() +const linkedAlertId = computed(() => { + const id = route.query.alert + return typeof id === 'string' && id !== '' ? id : undefined +}) + const activeTab = computed({ get: () => { + if (linkedAlertId.value) return 'history' const t = route.params.tab as string if (t === 'channels') return 'triggers' // legacy redirect if (t === 'triggers' || t === 'silence' || t === 'history') return t @@ -77,7 +83,7 @@ onUnmounted(() => { - +