Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 43 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@

<p align="center">
<strong>Drop a container. Your stack is monitored.</strong><br>
<em>When something breaks, fixed rules catch it, the alert is pushed to you, and your AI agent investigates through the built-in MCP server.</em><br><br>
Docker, Kubernetes, uptime, TLS, cron jobs, live logs, image updates, CVEs: auto-discovered, alerting on every one of them,<br>
from a single Go binary that idles under 30 MB of RAM. No PromQL, no exporters, no dashboards to build.
</p>
Expand All @@ -20,7 +21,7 @@
</p>

<p align="center">
<a href="#quick-start">Quick Start</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#why-maintenant">Why maintenant</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#features">Features</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="https://docs.maintenant.dev/">Documentation</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#editions">Editions</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="https://maintenant.dev/pricing/">Pricing</a>
<a href="#quick-start">Quick Start</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#why-maintenant">Why maintenant</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#agents">Agents</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#features">Features</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="https://docs.maintenant.dev/">Documentation</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="#editions">Editions</a>&nbsp;&nbsp;&bull;&nbsp;&nbsp;<a href="https://maintenant.dev/pricing/">Pricing</a>
</p>

---
Expand Down Expand Up @@ -113,6 +114,7 @@ Against the tools usually stacked up next to it:
| | maintenant | Uptime Kuma | Portainer | Dozzle |
| ---------------------------- |:--------------:|:-----------:|:----------:|:----------:|
| Container auto-discovery | **Yes** | No | Yes | Yes |
| MCP server, agent-ready | **Built in** | Third-party | Separate | No |
| Live container logs | **Yes** | No | Yes | Yes |
| HTTP/TCP endpoint checks | **Yes** | Yes | No | No |
| Cron/heartbeat monitoring | **Yes** | Yes | No | No |
Expand Down Expand Up @@ -211,6 +213,46 @@ Against the tools usually stacked up next to it:

---

## Agents

maintenant watches, your AI agent investigates. The loop has three steps:

1. **Detect.** Fixed rules decide that something is down: consecutive failures, thresholds, missed deadlines, restart loops. No model is involved, so an alert is never invented and nothing is spent while everything is fine.
2. **Push.** The alert leaves as a [webhook](https://docs.maintenant.dev/features/alerts/#webhook-payload): `alert.fired` when it starts, `alert.resolved` when it recovers, with the host it belongs to (`agent_id`) and a link back to it.
3. **Investigate.** Your agent receives the alert and queries maintenant's built-in MCP server: the container's logs, the endpoint's check history, CPU and memory, the other active alerts on the same host. You get a diagnosis, not just a red light.

```mermaid
flowchart LR
S[containers, endpoints,<br>certificates, cron jobs, hosts] --> M[maintenant<br>rules fire an alert]
M -- webhook --> A[your agent<br>Claude Code, OpenCode, n8n]
A -- MCP: logs, history,<br>active alerts --> M
```

**With Claude Code.** A [small receiver](examples/agent-webhook/receiver.py) (Python, standard library only) runs on the host next to maintenant. Each fired alert starts a headless Claude Code session that can only call maintenant's read-only MCP tools, and the diagnosis lands in `reports/<alert-id>.md`:

```bash
export RECEIVER_TOKEN=$(openssl rand -hex 32)
python3 examples/agent-webhook/receiver.py # listens on 127.0.0.1:9099
```

```bash
# what the receiver runs for each alert
claude -p "maintenant just fired this alert: {...} Investigate it read-only." \
--setting-sources project --mcp-config mcp.json --strict-mcp-config --tools "" \
--allowedTools mcp__maintenant__list_alerts mcp__maintenant__get_container_logs \
mcp__maintenant__get_endpoint_history mcp__maintenant__get_resources
```

`mcp.json` reaches maintenant over stdio with `docker exec -i maintenant /app/maintenant --mcp-stdio`, so no MCP port is opened. Point a webhook channel at the receiver with an `Authorization: Bearer <token>` header and route alerts to it with a trigger. The [AI agents guide](https://docs.maintenant.dev/guides/ai-agents/) walks through the setup.

**With OpenCode**, the same receiver calls `opencode run` instead of `claude -p`, with maintenant declared as a local MCP server in `opencode.json`. **With n8n**, a Webhook trigger feeds an AI Agent node whose MCP Client tool points at your instance's `/mcp` endpoint.

### [MCP server](https://docs.maintenant.dev/features/mcp/)

Built-in [Model Context Protocol](https://modelcontextprotocol.io/) server with 51 tools. Ask your AI assistant what is burning, read a container's logs, check the alert queue, acknowledge an alert, open an incident. stdio and Streamable HTTP transports, OAuth2 with a client id and secret for remote clients (Claude web, mobile and Desktop).

---

## Features

Every section links to its full documentation.
Expand Down Expand Up @@ -286,10 +328,6 @@ Channels: Discord and webhooks (Community), email and Telegram (Personal), Slack

Real-time status page with severity aggregation across every monitor, live over SSE. **Personal** adds incident timelines, **Pro** adds email subscribers (double opt-in, through your own SMTP server), maintenance windows and branding.

### [MCP server](https://docs.maintenant.dev/features/mcp/)

Built-in [Model Context Protocol](https://modelcontextprotocol.io/) server with 51 tools. Ask your AI assistant what is burning, read a container's logs, check the alert queue, acknowledge an alert, open an incident. stdio and Streamable HTTP transports, OAuth2 with a client id and secret for remote clients (Claude web, mobile and Desktop).

---

## Configuration
Expand Down
17 changes: 9 additions & 8 deletions docs/api/reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -434,7 +434,7 @@ All routes are open in every edition.
- `GET /api/v1/alerts/active` answers `{ "critical": [...], "warning": [...], "info": [...] }`. Acknowledged and resolved alerts are left out.
- `GET /api/v1/alerts/{id}` answers `404 NOT_FOUND` for an unknown alert.
- A daily job purges the alerts that are not active (resolved or silenced) once they are older than 90 days, counted from their resolution, or from their creation when they never resolved. An alert that is still active is never purged, however old.
- An alert has `id`, `source`, `alert_type`, `severity` (`critical`, `warning` or `info`), `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `details` (a string holding JSON), `fired_at`, `resolved_at`, `resolved_by_id`, `acknowledged_at`, `acknowledged_by`, `escalated_at` (set when an escalation policy notifies a level for the alert) and `created_at`.
- An alert has `id`, `source`, `alert_type`, `severity` (`critical`, `warning` or `info`), `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `agent_id` (the host the alert belongs to: the agent that runs the container, endpoint, heartbeat or certificate, the disconnected agent itself, or `00000000-0000-0000-0000-000000000000` for the server's own runtime and for alerts with no host, such as the infrastructure security score), `details` (a string holding JSON), `fired_at`, `resolved_at`, `resolved_by_id`, `acknowledged_at`, `acknowledged_by`, `escalated_at` (set when an escalation policy notifies a level for the alert) and `created_at`.
- `POST .../acknowledge` takes `{ "acknowledged_by": "..." }` (required). It answers the alert, emits `alert.acknowledged` and stops its running escalation. The dashboard, the MCP server and the security posture acknowledgments all go through this one path, so they behave the same. An acknowledged alert whose severity later rises does not start a new escalation. Errors: `400 INVALID_BODY`, `400 INVALID_REQUEST`, `404 NOT_FOUND` (unknown alert) and `409 CONFLICT` (the alert is not active or is already acknowledged).

---
Expand Down Expand Up @@ -464,6 +464,7 @@ A channel has `id`, `name`, `type`, `url`, `headers` (a string holding a JSON ob

- `POST` body: `name` (required), `url` (required), `type`, `headers`, `secret` (required for Telegram), `config` and `enabled` (default `true`). A URL must be HTTPS and must not resolve to a private or internal address, unless `MAINTENANT_ALLOW_PRIVATE_WEBHOOKS` is set. Answers `201` with the channel.
- `PUT` takes the same fields, all optional. A secret cannot be cleared. A request that only sets `enabled` to `false` is always accepted, even after the edition dropped.
- A `webhook` channel receives `{ "event": "alert.fired" | "alert.resolved", "alert": {...}, "timestamp": "..." }`. `alert` holds the fields of the [`alert.fired` SSE event](#sse-event-stream), plus `url`, the link to the alert in the UI (`<MAINTENANT_BASE_URL>/alerts/history?alert=<id>`). `url` is absent when the base URL has no host a recipient could open, as with the default built from a listen address such as `:8080` or `0.0.0.0:8080`.
- `test` answers `404 NOT_FOUND` for an unknown channel. Otherwise it answers `200`: `{ "status": "delivered", "response_code": n }` or `{ "status": "failed", "error": "..." }`.
- Creating, updating or testing a channel of a type the edition does not open answers `403 EDITION_REQUIRED` (`feature` is the capability: `slack`, `teams`, `smtp` or `telegram`). After a downgrade such a channel is `suspended`: it stops delivering and `GET /api/v1/edition` lists it under `suspended_channels`.
- Errors: `400 INVALID_BODY`, `400 VALIDATION_ERROR`, `404 NOT_FOUND` and `409 DUPLICATE_NAME` on create.
Expand Down Expand Up @@ -877,13 +878,13 @@ The SSE `event:` field is the event type, and `data:` holds the JSON payload onl
| `certificate.created` | A certificate monitor is created or auto-detected | `monitor_id`, `hostname`, `port`, `source`, plus `server_name` and `agent_id` when they apply |
| `certificate.check_completed` | A check finishes, including failed ones | `monitor_id`, `hostname`, `status`, `checked_at`, plus `subject_cn`, `issuer_cn`, `not_after`, `days_remaining`, `chain_valid`, `hostname_match` when known |
| `certificate.status_changed` | A monitor changes status | `monitor_id`, `hostname`, `previous_status`, `new_status`, `days_remaining`, `timestamp` |
| `certificate.alert` | Expiry threshold, invalid chain, hostname mismatch, revoked OCSP response or expiry | `monitor_id`, `hostname`, `port`, `alert_type`, `severity`, `timestamp` and details |
| `certificate.recovery` | A certificate alert clears: the certificate was renewed, or the chain, hostname or OCSP problem is gone | `monitor_id`, `hostname`, `port`, `previous_alert_type` (`expiring`, `expired`, `chain_invalid`, `hostname_mismatch` or `ocsp_revoked`), `days_remaining`, `timestamp`, plus `new_not_after` and `server_name` when they apply |
| `certificate.alert` | Expiry threshold, invalid chain, hostname mismatch, revoked OCSP response or expiry | `monitor_id`, `hostname`, `port`, `alert_type`, `severity`, `timestamp`, `agent_id` and details |
| `certificate.recovery` | A certificate alert clears: the certificate was renewed, or the chain, hostname or OCSP problem is gone | `monitor_id`, `hostname`, `port`, `previous_alert_type` (`expiring`, `expired`, `chain_invalid`, `hostname_mismatch` or `ocsp_revoked`), `days_remaining`, `timestamp`, `agent_id`, plus `new_not_after` and `server_name` when they apply |
| `certificate.deleted` | A certificate monitor is deleted | `monitor_id`, `hostname` |
| `resource.snapshot` | A sample is stored (live samples only) | `container_id`, `cpu_percent`, `mem_used`, `mem_limit`, `mem_percent`, `net_rx_bytes`, `net_tx_bytes`, `block_read_bytes`, `block_write_bytes`, `timestamp`, `agent_id` |
| `resource.alert` | CPU or memory stays over its threshold (one event per metric) | `container_id`, `container_name`, `alert_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp` |
| `resource.recovery` | CPU or memory returns to normal (one event per metric) | `container_id`, `container_name`, `recovered_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp` |
| `alert.fired` | An alert is raised, or its severity escalates | the alert: `id`, `source`, `alert_type`, `severity`, `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `details` (an object), `fired_at`, `created_at` |
| `resource.alert` | CPU or memory stays over its threshold (one event per metric) | `container_id`, `container_name`, `alert_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp`, `agent_id` |
| `resource.recovery` | CPU or memory returns to normal (one event per metric) | `container_id`, `container_name`, `recovered_type` (`cpu` or `memory`), `current_value`, `threshold`, `timestamp`, `agent_id` |
| `alert.fired` | An alert is raised, or its severity escalates | the alert: `id`, `source`, `alert_type`, `severity`, `status`, `message`, `entity_type`, `entity_id`, `entity_name`, `agent_id`, `details` (an object), `fired_at`, `created_at` |
| `alert.silenced` | An alert is raised while a silence rule or a maintenance window matches | same as `alert.fired` |
| `alert.resolved` | An alert resolves | same as `alert.fired`, with `resolved_at` |
| `alert.acknowledged` | An alert is acknowledged, from the REST route, the MCP server or a security posture acknowledgment | the alert as stored: `details` is a string holding JSON |
Expand All @@ -898,8 +899,8 @@ The SSE `event:` field is the event type, and `data:` holds the JSON payload onl
| `storage.availability_changed` | The database becomes unreachable, or answers again | `engine`, `connected` |
| `update.scan_started` | A scan starts | `scan_id`, `started_at` |
| `update.scan_completed` | A scan ends | `scan_id`, `updates_found`, `errors` |
| `update.detected` | A scan finds an update (on every scan, for each container) | `container_id`, `container_uid`, `container_name`, `image`, `current_tag`, `latest_tag`, `update_type`, `risk_score`, `alert_on`, plus `update_command`, `rollback_command` |
| `update.resolved` | An update is no longer pending | `container_id`, `container_uid`, `container_name` |
| `update.detected` | A scan finds an update (on every scan, for each container) | `container_id`, `container_uid`, `container_name`, `image`, `current_tag`, `latest_tag`, `update_type`, `risk_score`, `alert_on`, `agent_id`, plus `update_command`, `rollback_command` |
| `update.resolved` | An update is no longer pending | `container_id`, `container_uid`, `container_name`, `agent_id` (empty when the scan no longer sees the container) |
| `security.insights_changed` | A container's insights change | `container_id`, `container_name`, `highest_severity`, `count`, `change` |
| `security.insights_resolved` | All insights of a container are gone | `container_id`, `container_name` |
| `security.posture_changed` | The posture score moved by 5 points or more, or changed colour (needs `MAINTENANT_SECURITY_SCORE_THRESHOLD`; evaluated at every scoring) | `score`, `previous_score`, `color` |
Expand Down
63 changes: 61 additions & 2 deletions docs/features/alerts.md
Original file line number Diff line number Diff line change
Expand Up @@ -158,13 +158,38 @@ POST /api/v1/channels
}
```

maintenant POSTs this JSON body, with `event` set to `alert.fired` or `alert.resolved` (`test` for a test notification):
#### Webhook payload

maintenant sends one `POST` per notification, with `Content-Type: application/json` and the channel's own headers. Any 2xx answer counts as delivered; anything else, or no answer within 10 seconds, is retried as described in [Delivery and Retries](#delivery-and-retries).

| Field | Type | Description |
|-------|------|-------------|
| `event` | string | `alert.fired`, `alert.resolved` or `test` |
| `timestamp` | string | When the notification was sent, RFC 3339 UTC |
| `alert.id` | string | UUID of the alert. The resolved notification carries the id of the alert it resolves. |
| `alert.agent_id` | string | UUID of the host the alert belongs to. The server's own runtime is `00000000-0000-0000-0000-000000000000`; remote hosts are listed by `GET /api/v1/agents`. |
| `alert.url` | string | Link to the alert in the UI, `<MAINTENANT_BASE_URL>/alerts/history?alert=<id>`. Set `MAINTENANT_BASE_URL` to the address your team opens: without it, the link is built from the listen address. Absent when that address names no reachable host (`0.0.0.0`, `::` or an empty host), and on `test` notifications. |
| `alert.source` | string | `container`, `endpoint`, `heartbeat`, `certificate`, `resource`, `update`, `security`, `agent`, `host`, `swarm` or `kubernetes` (see [Alert Sources](#alert-sources)) |
| `alert.alert_type` | string | Source-specific type, such as `consecutive_failure` or `restart_loop` |
| `alert.severity` | string | `critical`, `warning` or `info` |
| `alert.status` | string | `active` on `alert.fired`, `resolved` on `alert.resolved` |
| `alert.message` | string | Human-readable description; on `alert.resolved`, the recovery message |
| `alert.entity_type`, `alert.entity_id`, `alert.entity_name` | string | The monitored object: its type, UUID and display name |
| `alert.fired_at` | string | When the condition was detected, RFC 3339 UTC |
| `alert.created_at` | string | When the alert was stored |
| `alert.details` | object | Source-specific values (target, failure count, threshold, last error...). Absent when the source has none. |
| `alert.resolved_at`, `alert.resolved_by_id` | string | On `alert.resolved` only: the recovery time and the UUID of the recovery record |
| `alert.acknowledged_at`, `alert.acknowledged_by`, `alert.escalated_at` | string | Present once the alert has been acknowledged or escalated |

An alert that fires:

```json
{
"event": "alert.fired",
"alert": {
"id": "0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8",
"agent_id": "0198a0f4-1c2d-7e3f-8a4b-5c6d7e8f9a0b",
"url": "https://now.example.com/alerts/history?alert=0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8",
"source": "endpoint",
"alert_type": "consecutive_failure",
"severity": "critical",
Expand All @@ -186,7 +211,41 @@ maintenant POSTs this JSON body, with `event` set to `alert.fired` or `alert.res
}
```

A resolved notification carries the same alert with `status` set to `resolved`, a `resolved_at` time, the `resolved_by_id` of the recovery record, the `info` severity and the recovery message. The payload is maintenant's own: Slack and Teams expect theirs, which is why they are native channels.
The same alert when it recovers. The id, the host, the entity, `fired_at` and `details` are those of the original alert; the severity and the message come from the recovery:

```json
{
"event": "alert.resolved",
"alert": {
"id": "0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8",
"agent_id": "0198a0f4-1c2d-7e3f-8a4b-5c6d7e8f9a0b",
"url": "https://now.example.com/alerts/history?alert=0198b1c2-7a3e-7f00-9c11-2d4e5f60a7b8",
"source": "endpoint",
"alert_type": "consecutive_failure",
"severity": "info",
"status": "resolved",
"message": "Endpoint https://api.example.com/health recovered after 2 consecutive successes",
"entity_type": "endpoint",
"entity_id": "0198b1c2-8b4f-7a11-8d22-3e5f6071b8c9",
"entity_name": "api",
"fired_at": "2026-03-01T02:00:00Z",
"created_at": "2026-03-01T02:00:00Z",
"resolved_at": "2026-03-01T02:04:30Z",
"resolved_by_id": "0198b1c6-9d01-7c22-b733-4f6a7182c9da",
"details": {
"target": "https://api.example.com/health",
"failures": 3,
"threshold": 3,
"last_error": "connection refused"
}
},
"timestamp": "2026-03-01T02:04:30Z"
}
```

A few cases send `alert.fired` again for the same `id`: a [severity raise](#life-of-an-alert) (the message starts with "Severity raised from X to Y") and each level of an [escalation policy](alert-escalation.md) that targets the channel. A receiver that must act once per alert keys on `id`. The **Test** button sends `event: "test"` with an `alert` whose `id`, `agent_id` and `entity_id` are empty and whose `source` and `alert_type` are `test`.

The payload is maintenant's own: Slack and Teams expect theirs, which is why they are native channels. To hand alerts to an AI agent, see [AI Agents](../guides/ai-agents.md).

### Email (SMTP) :material-star-four-points:{ title="Personal" }

Expand Down
Loading