Skip to content

fix: harden agents, alerting, status page and deployment; align the docs with the code - #129

Merged
btouchard merged 54 commits into
mainfrom
039-audit-fixes
Oct 1, 2026
Merged

btouchard merged 54 commits into
mainfrom
039-audit-fixes

Conversation

@btouchard

Copy link
Copy Markdown
Contributor

An audit of the documentation against the code surfaced behaviour that drifted from what the product promises. This PR fixes it and brings every page back in line.

Fixes

  • Agents: the gRPC listener starts again in Personal and Pro. Agents re-enroll after revocation, keep spool replays out of live state, honour per-endpoint intervals, start without a container runtime and report runtime changes.
  • Deployment: hardened Kubernetes pods start, manifests carry their namespace and complete RBAC, the entrypoint prepares bind-mounted data directories, the generated agent manifest is a single Deployment. install.sh restarts on update, keeps the env file and installs offline.
  • Alerting: cumulative escalation delays, one acknowledgment path, ordered deliveries. Every certificate, resource and Swarm alert resolves, Kubernetes alerts are raised, retention never drops an active alert.
  • Status page: subscriber emails go through MAINTENANT_SMTP_*, hidden components never reach a public surface, incident and maintenance inputs are validated.
  • Runtimes: Docker and Kubernetes loss is detected and the UI degrades accordingly. Swarm service labels, cluster counts and crash loops are read cluster-wide.
  • Updates: official images are scanned, floating tags are tracked against the running digest, rollbacks target the previous image, maintenant.update.* labels apply to agents and Kubernetes.
  • Security: cross-origin writes to /api are refused, MAINTENANT_CA_CERT covers every outbound TLS client, public SVG assets are parsed in full, /mcp works behind a local reverse proxy.

Behaviour changes

  • Migrations 35 to 39: daily uptime rollups, heartbeat pauses, unused columns dropped.
  • Removed: filter_tags, the maintenant.alert.channels label, GET/PUT /api/v1/status/smtp, and the alert_entity_routing capability, which no code ever implemented.
  • Cross-origin POST/PUT/PATCH/DELETE on /api return 403 CROSS_ORIGIN_REFUSED unless the origin is listed in MAINTENANT_CORS_ORIGINS.
  • An invalid MAINTENANT_MAX_BODY_SIZE or MAINTENANT_SECURITY_SCORE_THRESHOLD, or half a gRPC TLS key pair, now stops startup.

Docs

SECURITY.md, the README and every page under docs/ were checked against the code: the API reference lists the registered routes, configuration covers every variable, and the reverse proxy examples were tested.

…custom data paths

- Replace the binary through a temporary file and a rename, then restart
  the service when it already runs (start it otherwise).
- Merge the env file key by key: only the keys passed as flags change,
  every other line (comments, DOCKER_HOST, names with digits) is kept.
- Take the flag lists from the binary: every FlagTypeBool flag works bare
  or as --flag=true|false, --flag=value works everywhere, unknown flags
  are refused. The help lists every flag; a Go test keeps both in step.
- Add --binary and --sha256sums for an install with no network access.
- Follow --data-dir and --db: create those directories for the service
  user and render the unit with them as working and writable paths.
- Show the configured listen address in the summary, not the script's
  own environment.
- Exit 20 on a failed download and 30 on a filesystem error.
- Keep the telemetry directory next to the database instead of /data/shm,
  which the native unit cannot write.
- Check Docker Hub official images on latest; skip only images the local
  Docker runtime reports as never pulled (no RepoDigests, read from the
  image list at scan time).
- Rollback commands return to the image the container ran before the
  update: running repo digest, else digest baseline, else fixed version
  tag. Compose retags the unchanged tag or pins the previous image,
  Kubernetes sets the image explicitly. previous_digest is persisted.
- The Compose update command asks to write the new tag in the compose
  file when the tag changes.
- Apply maintenant.update.track, ignore_major, digest_only and alert_on.
- Image exclusions match the image name without tag or digest.
- Tag filters keep the current tag as the digest comparison reference.
…refusals

Unsafe cross-origin browser requests on /api/ are refused with a 403
CROSS_ORIGIN_REFUSED through http.CrossOriginProtection, trusting the
origins listed in MAINTENANT_CORS_ORIGINS. /ping, /status, /mcp, /oauth
and /.well-known keep accepting cross-origin calls.

CORS preflights now allow PATCH, every SSE response (MCP transport
included) carries X-Accel-Buffering: no, a PostgreSQL server that declines
TLS is reported apart from an unreachable one, and a MCP client secret
shorter than 32 characters logs a warning at startup.
… as shipped

- Entrypoint: exec the command directly when not root (runAsUser, --user)
  instead of failing in setpriv; detect agent mode from MAINTENANT_MODE and
  the real data dir; own the database directory, never its content nor /.
- Image: default MAINTENANT_DB to /data/maintenant.db, on the data volume.
- Config: read MAINTENANT_NODE_NAME from the environment.
- Kubernetes runtime: Connect retries until the API server answers, and the
  local topology reconcile only runs once connected, which removes the nil
  pointer panic on an unreachable cluster.
- RBAC: one read-only rule list derived from the runtime's API calls, used by
  the agent manifest and checked against the server manifests (adds nodes,
  batch/jobs, metrics nodes; drops replicasets).
- Agent manifest: a single-replica Deployment with a PVC for identity and
  spool, hardened, instead of a DaemonSet on hostPath.
- deploy/kubernetes: pin the maintenant namespace on namespaced objects.
- Standalone install tab: hand out the install.maintenant.dev command.
On a loopback listener the MCP SDK refuses any request whose Host header
is not loopback, so a local nginx forwarding the public Host got 403
"invalid Host header" on every /mcp call. When MCP OAuth is configured,
the bearer token already authenticates each request, so the SDK's
localhost protection is turned off in that case only. The unauthenticated
mode keeps it.
The subscriber service and the notifier were built with a nil mailer, so no
confirmation or notification email ever left. Both now use the SMTP client of
the email channel, built from MAINTENANT_SMTP_*. Mails are plain text, and the
SMTP exchange honours its context with a 30 s ceiling.

Manual incidents (API and MCP) now notify confirmed subscribers on creation,
on each update and on resolution, through the same announce path as automatic
incidents. A resolving update sends one resolution mail, not an update too.

The in-memory status page SMTP settings are removed (GET/PUT
/api/v1/status/smtp, form, types). POST /api/v1/status/smtp/test takes a
recipient and uses the environment configuration.

/status/api exposes subscriptions_enabled and the public page shows the
subscription form only then. POST /status/subscribe answers JSON errors
(subscriptions_unavailable, confirmation_failed, ...) and its 429 matches the
other limiters, with Retry-After.
- POST /api/v1/channels without "enabled" creates an enabled channel, as MCP does.
- Escalation levels fire at run start + delay (cumulative), shifted only by
  a maintenance pause; the ack notification reaches every channel the run
  notified, once; ack and exhaustion messages are in English.
- Trigger scope-filter refusals use EDITION_REQUIRED with feature and
  required_edition.
- Remove trigger filter_tags, escalation policy tags and the
  maintenant.alert.channels label/annotation, none of which had any effect;
  migration 36 drops their columns on both engines.
- MCP list_agents checks the multihost capability; a conformance test now
  holds every edition-gated tool to its description.
… trust MAINTENANT_CA_CERT everywhere

Replayed spool events now feed history only: resource samples skip the
threshold pipeline, container state and health changes write the timeline
from the recorded state at their time without touching the current row or
emitting, and a replayed container absent from the inventory is kept archived.

An agent whose identity the server revokes or no longer knows now enrolls a
fresh identity once with the configured token, or exits non-zero with a
message naming MAINTENANT_ENROLLMENT_TOKEN. The stream handler also ends when
the receive side closes, so a refusal is seen without waiting for a send.

The agent probes each labelled endpoint at its own interval and timeout.
MAINTENANT_CA_CERT now applies to the agent gRPC client, webhooks and
channels, SMTP STARTTLS, license, OSV, changelog, EOL and registry clients.
Half a gRPC TLS keypair stops startup, Community logs why no agent listener
starts, and the commercial set wires the multi-host extension again.
…r stale kubeconfigs

- Each event stream builds and starts its own informers after registering
  them; before, the factory was started empty at connect time and no
  Kubernetes event ever reached the server.
- The stream ends when the API server misses three probes in a row (15 s
  apart), so the supervisor goes degraded and reconnects as it does for
  Docker. Any API answer, a refusal included, counts as reachable.
- Container events log a shortened external ID without slicing past the end
  of short Kubernetes IDs such as "ns/pod", which now reach this code.
- Auto-detection from a kubeconfig falls back to Docker when that cluster is
  unreachable; in-cluster and MAINTENANT_RUNTIME=kubernetes still wait.
- Helm chart 1.3.0, appVersion 1.8.0.
Subscribing an address already on file failed with a 500 on its unique
constraint. The store now upserts a pending subscription: a new or still
unconfirmed address gets a fresh confirmation token and a new 24 h window,
a confirmed one is left as is. The confirmation email goes out in the
background, so POST /status/subscribe gives the same code, body and delay
whether the address is new, pending or confirmed, and whether the mail
server answers or not. The confirmation_failed code is gone.

buildMIME encodes the Subject header as an RFC 2047 encoded word, so
accented incident titles and alert messages reach mail clients intact,
for the status page and the email channel alike.
…d Kubernetes service insights

Daily uptime is now rolled up into endpoint_, heartbeat_ and
container_uptime_daily (migration 35, both engines) before the raw rows
are purged, and kept 365 days; the current day is still computed live
with portable SQL. Heartbeat days count completion pings only, exit code
0 as success. Transition purge keeps each container's latest transition.

A container marked maintenant.ignore raises no alert and declares no
endpoint or certificate, locally, through agents and from Kubernetes
annotations; label-derived fields are refreshed on reconcile.

Swarm task containers read their service labels (deploy.labels) under
their own and are grouped by stack; the unused task-to-container mapping
is removed.

LoadBalancer and NodePort Services exposing a local workload raise
security insights; missing_network_policy is removed.
…gle errors

- Triggers: an absent "enabled" means enabled on creation and keeps the
  stored value on update, in REST and MCP alike.
- Triggers: the advanced-filters gate applies to a scope filter the request
  changes, so a downgraded instance can still switch a scoped trigger on or
  off; a failed toggle or delete is shown in the trigger list.
- Escalation editor: loads, shows and sends back the policy's scopes instead
  of wiping them on save.
- AcknowledgeButton labels are in English, like the rest of the interface.
The clientset and the metrics client were plain fields, rewritten by every
connection while logs, topology, stats and the event stream read them from
other goroutines. They now live together in an atomic.Pointer, set once per
connection and read through a single accessor; before the first connection
the accessor returns an error instead of a nil client.
Fix the Caddy and nginx examples (public and MCP routes were behind auth), add /assets and MAINTENANT_TRUSTED_PROXIES, document CSRF protection, exact CSP and CORS, rate limiters, MCP OAuth, agent gRPC, demo mode, SSRF guard and secrets at rest, and cover release binary verification in SECURITY.md.
… fix false Pro comments

The API reference now lists every registered route with its real edition gate, parameters, bodies and error codes, the full error format and the events the server actually emits. The architecture page covers the modes, the agent protocol, the packages, the flows and the retention as they are. The Pro labels on routes, swarm handlers and events that no edition check backs are removed.
…pdate labels to agents and Kubernetes

A floating tag was compared with the digest seen at the previous scan, so an
update was raised once and resolved on the next scan without any action. The
scan now compares the tag's multi-platform digest (or one of its platform
manifests) with the digest the container actually runs: Docker repo digests,
the new repo_digests field agents send with their inventory, and the imageID
of Kubernetes pods. The baseline only stands in when the runtime cannot tell,
and is kept as the multi-platform digest so a rollback never pins the
server's platform.

maintenant.update.* labels of agent containers are kept in memory from their
events, and Kubernetes workloads read their maintenant.update.* annotations.

The Kubernetes update command for a republished tag now sets the image to
tag@digest, since setting an unchanged reference rolls nothing out, and names
the pod container instead of the workload.
Document every environment variable the code reads, grouped by area, with
exact defaults, the license offline behavior, the exact telemetry payload
and the PostgreSQL and SQLite storage roles.

Rewrite the native install page for the offline install, the env file
merge, custom paths, the docker group, version pinning and exit codes,
and the container page for image tags, the entrypoint and the hardened
Kubernetes pod.

Fix the flag help texts, .env.example and the installer README.
… audited code

Container states, restart loop, ignore label and compose one-offs; real endpoint defaults, status versus alert thresholds, manual endpoints and quotas; heartbeat deadlines, statuses and routes; certificate severities, hostname_mismatch and Personal OCSP; resource alert body and sampling; host OS support states and Updates page; multi-host editions, spool replay, re-enrollment, CA bundle, Kubernetes Deployment agent and sentinel agent_id; agent setup installer, RBAC and variables; PostgreSQL local hosts, TLS refusal, MAINTENANT_DB and Helm values.
Troubleshooting now quotes the real log messages and covers the current pitfalls: loopback bind, busy port, degraded runtime, Kubernetes fallback, edition refusals, cross-origin refusals, trusted proxies, MCP startup, revoked agents, PostgreSQL TLS and long migrations. The cloud guides describe the data bind mount, the multi-host setup (Personal edition, gRPC listener and certificates), the reverse proxy settings, the socket warning and backups without sqlite3 in the image, and fix the ordering and per-provider inconsistencies. The cloud-init header lists all five providers.
…, fill container Swarm fields

- The Kubernetes alert checker now runs in the local cluster's reconcile
  loop. replica_health, crash_loop and node_condition reach the alert
  engine with an entity id, fire once, escalate, and resolve when the
  condition clears, the object is ignored or gone, including alerts left
  by a previous run. Each transition broadcasts kubernetes.workload_changed,
  pod_changed or node_changed. Jobs raise no replica alert, a node's
  conditions share one alert, and a node whose Ready condition is Unknown
  counts as NotReady.
- Pods created by a Deployment's ReplicaSet point to the Deployment.
- A workload or pod annotated maintenant.ignore, and a Swarm service
  labelled maintenant.ignore, raise no alert; an alert raised before the
  label is resolved.
- Swarm service alerts carry the service id, so two services no longer
  share one alert; alerts of services the manager no longer lists are
  resolved.
- Enabling Swarm while the server runs builds, wires and starts the same
  services as at boot, so its events and alerts go out.
- Swarm task containers read their service, node and slot from the labels
  Docker sets on them, locally and through agents, and the container
  detail returns them. swarm_service_mode and swarm_desired_replicas, which
  no container label carries, are dropped (migration 37, both engines).
…race-free activation

- Swarm node alerts carry the node ID, so two nodes no longer share one
  alert; alerts of nodes the manager no longer lists are resolved, as for
  removed services.
- A rolling update that completes resolves the service's update_rollback
  and update_stalled alerts; ignoring the service resolves them too.
- Replica, crash-loop and rolling-update alerts still active from a
  previous run are taken over at startup, and when Swarm is enabled later,
  so they resolve once their condition clears.
- The Swarm manager's services are built together and published through
  one atomic pointer, the cluster through another: enabling or leaving
  Swarm while events flow is race-free, and the API reads the update
  tracker and crash-loop detector in place at request time. The task
  tracker and the handler's replica checker, which nothing read, are gone.
quorum_degraded is raised once when fewer managers than the quorum are
ready, instead of on every node reconcile, and resolved when the quorum
is reached again. An alert still active from a previous run is taken
over at startup, and when Swarm is enabled later, like the other Swarm
alerts.
…agent container exits, weigh heartbeat uptime by time

- certificates: expired, chain_invalid, hostname_mismatch and ocsp_revoked
  now recover once a scan no longer shows them, not only expiring
- certificates: endpoint probes and agent scans store a result once per
  monitor interval or when the outcome changes, instead of every probe
- resources: CPU and memory alerts fire and resolve independently; leaving
  both_alert no longer sends an unmatched both_threshold recovery
- resources: live agent samples feed the current snapshot and the live top
- containers: an OOM kill (137 with OOMKilled) is a crash, docker stop stays
  completed; agents carry the exit code and a new oom_killed proto field in
  their inventory and events, and an EXITED report without a code is a stop
- embedded agent dials grpc:// when the listener runs in h2c mode
- heartbeats: daily uptime replays pings against the deadline, so missed
  deadlines lower it and days without pings get a value
- /ping records the client address resolved like the rate limit
Status page:
- validate component overrides, incident severities and statuses against
  the model, on REST and MCP, with an explicit 400
- list upcoming maintenance soonest first; the admin list shows running,
  then upcoming, then completed windows
- deleting a running maintenance window ends it like its scheduled end
  (incident resolved, components released, visitors told)
- a maintenance window no longer destroys a manual override: it is kept
  (migration 38) and restored when the last window holding the component
  ends; overlapping windows keep the component under maintenance
- the Atom feed carries ongoing incidents as well as recent ones
- SVG assets are read whole and refused when they carry scripts, event
  handlers, javascript: links, foreignObject or embedded HTML; an SVG
  without an XML declaration is recognised

Escalation:
- a policy holds at most 5 levels, enforced by the service for REST and
  MCP and reported in the plan limits the editor now reads
- disabling a policy stops its runs with stopped_by_policy_disabled
- drop the never-written skipped_maintenance delivery status from the
  contract and the schema (migration 38, SQLite rebuild after conversion)

Security posture:
- every infrastructure scoring evaluates the score threshold, and a
  background check scores every 5 minutes when a threshold is set
- the posture page help names the weighted categories of the scorer

Transport and config:
- SMTP fails when an announced STARTTLS fails, and uses implicit TLS on
  port 465
- SSE streams send a keep-alive comment every 25 s
- an invalid MAINTENANT_MAX_BODY_SIZE stops startup
The API returns affected components as {id, name}, but the front typed them
as {component_id, name}, so editing a window sent undefined component ids.
Align both component ref types and the editor with the API.

The admin API now answers 400 when component_ids holds an empty, unknown or
repeated id, on incident and maintenance create and update, instead of a 500
that left a half-written window.
Swarm is open in every edition, with real alert severities, delays and resolutions, deploy.labels and the agent path. The Kubernetes guide now carries the actual RBAC, namespace defaults, annotations, alerts, runtime behaviour and Helm values. The labels reference lists every label read, with real defaults and update label semantics.
…quotas

- Broadcast endpoint.* events to SSE and webhooks; notify the status page on status changes only
- Report MAINTENANT_STATUS_URL as status_url in GET /api/v1/edition
- Record webhook deliveries: last status, failure count, auto-disable after 10 failures, reset and re-enable on success; store webhook timestamps as epoch seconds
- Send the webhook test with the headers and signature of a real delivery, and record it like one
- Read the endpoint, heartbeat and certificate caps from the running edition at each creation
- Compute container uptime over every window the edition's history cap opens
- Answer a duplicate exclusion with the stored row and 200
- Start the escalation retention loop once, without an error at shutdown
- Emit status page events on the admin SSE bus as well
- Drop the unused endpoint orchestration_group filter and the untracked update count; fill agent_id on check results; standard error body on the logs stream 503; snake_case restart alert event
- Updates: keep the image reference when a pulled tag moves, kubectl commands for bare pods, purge the digest baselines of gone containers, keep scanning and warn once when the image list is refused
- Retry runtime connection errors with backoff and log their cause
- Correct comments that tied Personal or open capabilities to Pro
create_incident and create_maintenance wrote the incident or the window,
then failed on an unknown component id and left it without its components.
The MCP server now receives the status component store and runs the same
check as the REST API (status.CheckComponentIDs) before any write, with the
same message.
…component changes

- Deliver an alert's notifications to one channel, and a webhook subscription's events, on a single notifier worker, so a recovery can no longer overtake the alert it resolves while that alert's delivery waits or retries
- Emit status.component_created, status.component_updated and status.component_deleted (component id only) followed by status.global_changed on both SSE buses when an administrator changes a component
- Refresh the status admin components and the edition quotas on those events
- Match status.component_changed on component_id, as the server sends it, and apply its name, status and monitor breakdown in place
- Apply status.global_changed in place instead of refetching, which wiped the breakdown the previous event had just brought
- Refetch on component creation, update and deletion, alongside the incident and maintenance events
…cluster counts

- replica_unhealthy: service events feed the 5-minute ReplicaHealthChecker
  instead of raising at once, so a scale-up no longer leaves an alert open
- crash_loop: count the tasks Swarm marks failed on every node, read from the
  task list, instead of every local die; shutdowns and SIGTERM stops do not count
- SwarmCluster: fill manager/worker counts and creation date from docker info,
  refreshed on every recheck
- detection is armed for any Docker runtime and runs as soon as it connects
- quorum_degraded: raised when the node list fails because the swarm has no leader
- emit swarm.node_updated when a node joins or changes role, host, engine or address
- maintenant.update.registry now queries that registry as a mirror of the image
… Swarm services

Container health alerts and their recovery had no entity name, so every
notification was titled "Alert: ". The health event now carries the
container name, and a test walks every alert.Event built in internal/app
for an EntityName.

The stale update cleanup treated a container whose registry could not be
reached as scanned: its pending update was deleted and its alert resolved,
then the next scan raised both again. A container whose scan failed now
keeps its update and its alert until a scan reaches its registry.

Scanner.Scan returns only the containers with an update, so the CVE pass
meant for every container ran on those alone and left the others "not
evaluated" in the posture. The enrichment now also gets the running image
of every other container the scan lists.

A Swarm task got the standalone docker stop, rm and run commands, which
Swarm undoes by rescheduling the task. Its update, rollback and fix
commands now run docker service update --image on the task's service,
with the Kubernetes reference rule: the tag, or the tag and its new index
digest for a republished tag.
… guides with the audited code

Endpoint events, webhook delivery tracking, status_url, SSE keep-alive and new events, SMTP implicit TLS, strict MAINTENANT_MAX_BODY_SIZE, escalation limits, status page validation, Swarm task commands, runtime reconnection, and the help texts and category of the license flag.
Covers Swarm alerts and detection, certificate and resource alert recovery, status page events and validation, escalation limits, endpoint events, update commands for Swarm tasks, SMTP TLS, MCP parameters and the three-attempt notification retries.
…fig gaps

- agent: an unreachable Docker or Kubernetes no longer blocks startup. The
  agent enrolls, streams host metrics and OS identity, and retries the
  runtime in the background with backoff (1s to 1m), starting container
  monitoring as soon as it answers.
- swarm: a crash-looping service that becomes maintenant.ignore resolves its
  crash_loop alert at once, like replica_unhealthy and update alerts.
- config: MAINTENANT_SECURITY_SCORE_THRESHOLD and --securityScoreThreshold
  share one rule: a whole number from 0 to 100, 0 or unset disabling the
  alert; anything else stops the startup.
- help: drop the em dash from the --help header.
- deploy: remove MAINTENANT_TELEMETRY_DATADIR and MAINTENANT_TZ from
  compose.test.yml (read by nothing), replace the MAINTENANT_WEBHOOK_URL
  example in the Helm values with MAINTENANT_BASE_URL, and name Personal
  next to Pro for the license key.
The writer's stop called WaitGroup.Wait while submitters could still
Add from zero. The race detector reports that misuse, and it failed
TestStart_FailsOnBusyPort under load: the early return cancelled Start's
context while the resource rollup was writing. The same Wait also
deadlocked a write queued with a context the stop does not cancel.
Submitters now select on a stopped channel instead, and the vacuum test
no longer starts a writer that openTestDB already runs.

Start also binds its HTTP port before any service starts, so a taken
port fails the boot before it has side effects.
…the public routes

- Send component changes of a hidden component to the dashboard bus only;
  created, updated and deleted events reach the public page only when the
  component is or was visible there.
- Return the per-monitor breakdown in /status/api, the same fields as the
  status.component_changed event, so a reload keeps it.
- Parse the Content-Type of POST /status/subscribe as a media type and answer
  415 unsupported_media_type to anything but JSON or a urlencoded form.
- Build the Atom feed links from MAINTENANT_STATUS_URL, or MAINTENANT_BASE_URL
  followed by /status, instead of the request host.
- Refuse a maintenance window update whose end precedes its start.
- Answer the lower case internal code on a failed component creation, like the
  rest of the status admin routes.
- Drop the unused DeleteIncidentsOlderThan: incidents are kept until deleted.
The runtime an agent reports was frozen at enrollment: an agent enrolled
before Docker, Swarm or Kubernetes answered, or a Docker host that later
joined or left a Swarm, kept showing a wrong runtime.

- proto: add RuntimeMsg (AgentEvent field 19). An older server ignores it
  and keeps the enrollment value; an older agent never sends it.
- agent: report the runtime once it answers, keep it as current state in the
  spool so every new stream restates it, and recheck Swarm membership every
  minute on a Docker host. Joining or leaving a Swarm restarts collection
  under the new runtime; leaving sends an empty Swarm topology so the
  server drops the cluster it held for the agent.
- server: store the reported runtime in agents.detected_runtime and
  broadcast agent.updated when it changes, so /api/v1/agents, MCP and the
  UI follow it.
…Swarm leaves

- The Docker event stream closes once the daemon stops answering after a
  cut, so the supervisor marks the runtime disconnected, broadcasts
  runtime.availability_changed and reconnects. A cut while the daemon
  answers resumes right after the last event (the resume point was far in
  the future). On the agent, a closed stream ends the runtime collection,
  which restarts through the same loop as a Swarm switch once the runtime
  answers again: inventory resent, one new subscription.
- The header banner and the containers store read the runtime store, which
  follows /runtime/status and runtime.availability_changed; the
  runtime.status listener waited for an event never sent.
- Swarm detection logs only role changes. Deactivating Swarm stops the
  manager loops; reactivation starts a new manager.
- swarm.status is sent when the Swarm state changes (activation,
  deactivation, node counts), no longer at boot before any client listens.
- GET /containers/{id}/logs answers 503 RUNTIME_UNAVAILABLE while the local
  runtime is disconnected, like logs/stream.
- Drop recent_events from /swarm/dashboard: it had no source.
- Help: --runtime lists swarm for agents; the spool budgets say what 0
  does and that both must be 0 to turn the spool off.
The role took the first 16 hex characters of a UUIDv7, which are its
millisecond timestamp and sub-millisecond sequence with no random bit.
Test processes sharing one server could draw the same name and fail
CREATE ROLE. The whole UUID carries its 62 random bits, like the test
database names already do.
…ts quiet

- REST, MCP and the security posture acknowledge through Engine.Acknowledge:
  stored once, broadcast as alert.acknowledged, escalation stopped. An unknown
  alert id now answers 404 instead of 500.
- A severity raise on an acknowledged alert keeps the acknowledgment, never
  starts an escalation again and is notified once, as "Severity raised from X
  to Y". An escalation run that outlives an acknowledgment stops at its next
  cycle.
- escalated_at is stamped each time an escalation level goes out.
- Retention deletes resolved and silenced alerts 90 days after their
  resolution (or creation when they have none), never an active one.
- GET /api/v1/alerts pages with before and before_id, so alerts fired in the
  same second are neither lost nor repeated between pages.
- MCP list_alerts leaves acknowledged alerts out, as GET /alerts/active does,
  and returns at most 100 recent alerts. Fix the escalation scope kinds and the
  misplaced doc comments.
- Drop the 25s notifier backoff that no retry ever reached.
- An auto_incident component that is hidden no longer opens or updates a
  public incident; an incident opened while it was visible still resolves.
- Incident and maintenance component references carry the component's
  visibility. /status/api, the public status.incident_created and
  status.maintenance_* events and the subscriber emails name only the visible
  components; the dashboard bus keeps the full list.
- Drop the resolved-incident query GetPageData ran on every /status/api call
  although nothing read it.
…tbeats and endpoints

Updates
- Retention no longer deletes pending image updates behind their alert; scans
  still named by a pending update are kept so its scan_id never turns NULL.
- The enrichment log reports the running edition instead of "(Pro)".
- The risk score reads network exposure (security insights) and the last
  day's restarts; criticality and dependents, which had no source, are gone.
- Posture scoring reads certificates and updates once per pass, indexed by
  container, instead of once per container.

Certificates
- Chain validation no longer checks the hostname, so a name mismatch no
  longer also reports an invalid chain.
- Correct the comment on the crossed expiry threshold.

Resources
- Network totals are real rates in bytes per second, computed from two
  samples; unavailable (-1) or reset counters no longer count.
- Top consumers over 24h and more include the hour or day in progress.
- The cap of 20 containers lives in the service, so MCP gets it too.

Heartbeats
- Pauses are recorded (migration 39) and left out of the daily uptime; a
  resume or a ping restarts the deadline. Pausing a paused heartbeat is 400.
- Deleting a heartbeat removes it and its history; the active column goes.
- MCP pause/resume detect a missing heartbeat with the real sentinel error.

Endpoints
- Creation and update reject a TCP target that is not host:port and an HTTP
  target without an http(s) scheme and host.
- A label interval below the 5s API minimum is raised to it, with a warning.
…atus page, updates, resources, heartbeats and endpoints
…ration

The daily uptime reads, the top consumers ranking and every cleanup of the
retention pass each called time.Now() on their own. A test that built its
fixtures from an earlier reading saw its days shift by one when UTC midnight
passed in between (and its ranking lose the hour in progress at any hour
boundary), which made TestContainerDailyUptime fail on PostgreSQL at 23:58.

These operations now take their reference time from the caller: the HTTP
handler and the resource service pass time.Now(), the retention pass reads
the clock once and hands it to every rollup and cutoff, and the resource
rollup derives its hourly and daily buckets from a single reading. The tests
pass the same reference to their fixtures and to the code, and assert on
complete days rather than on a day in progress that is still empty at
00:00:00.
…nted

The Pro capability was advertised but no router was ever wired into the
engine. Trigger scopes already route an entity's alerts to chosen channels.
Comment thread internal/store/uptime_daily.go Fixed
…uptime window, patch base packages

The MCP authorize endpoint now redirects to a URL rebuilt from the
configured allowlist or a canonical loopback host and numeric port,
never to the raw redirect_uri. The daily uptime window is bounded with
an explicit guard, and the runtime image upgrades its Alpine packages
so OpenSSL fixes land at build time.
@btouchard
btouchard merged commit e97f381 into main Oct 1, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants