fix: harden agents, alerting, status page and deployment; align the docs with the code - #129
Merged
Merged
Conversation
…custom data paths - Replace the binary through a temporary file and a rename, then restart the service when it already runs (start it otherwise). - Merge the env file key by key: only the keys passed as flags change, every other line (comments, DOCKER_HOST, names with digits) is kept. - Take the flag lists from the binary: every FlagTypeBool flag works bare or as --flag=true|false, --flag=value works everywhere, unknown flags are refused. The help lists every flag; a Go test keeps both in step. - Add --binary and --sha256sums for an install with no network access. - Follow --data-dir and --db: create those directories for the service user and render the unit with them as working and writable paths. - Show the configured listen address in the summary, not the script's own environment. - Exit 20 on a failed download and 30 on a filesystem error. - Keep the telemetry directory next to the database instead of /data/shm, which the native unit cannot write.
- Check Docker Hub official images on latest; skip only images the local Docker runtime reports as never pulled (no RepoDigests, read from the image list at scan time). - Rollback commands return to the image the container ran before the update: running repo digest, else digest baseline, else fixed version tag. Compose retags the unchanged tag or pins the previous image, Kubernetes sets the image explicitly. previous_digest is persisted. - The Compose update command asks to write the new tag in the compose file when the tag changes. - Apply maintenant.update.track, ignore_major, digest_only and alert_on. - Image exclusions match the image name without tag or digest. - Tag filters keep the current tag as the digest comparison reference.
…refusals Unsafe cross-origin browser requests on /api/ are refused with a 403 CROSS_ORIGIN_REFUSED through http.CrossOriginProtection, trusting the origins listed in MAINTENANT_CORS_ORIGINS. /ping, /status, /mcp, /oauth and /.well-known keep accepting cross-origin calls. CORS preflights now allow PATCH, every SSE response (MCP transport included) carries X-Accel-Buffering: no, a PostgreSQL server that declines TLS is reported apart from an unreachable one, and a MCP client secret shorter than 32 characters logs a warning at startup.
… as shipped - Entrypoint: exec the command directly when not root (runAsUser, --user) instead of failing in setpriv; detect agent mode from MAINTENANT_MODE and the real data dir; own the database directory, never its content nor /. - Image: default MAINTENANT_DB to /data/maintenant.db, on the data volume. - Config: read MAINTENANT_NODE_NAME from the environment. - Kubernetes runtime: Connect retries until the API server answers, and the local topology reconcile only runs once connected, which removes the nil pointer panic on an unreachable cluster. - RBAC: one read-only rule list derived from the runtime's API calls, used by the agent manifest and checked against the server manifests (adds nodes, batch/jobs, metrics nodes; drops replicasets). - Agent manifest: a single-replica Deployment with a PVC for identity and spool, hardened, instead of a DaemonSet on hostPath. - deploy/kubernetes: pin the maintenant namespace on namespaced objects. - Standalone install tab: hand out the install.maintenant.dev command.
On a loopback listener the MCP SDK refuses any request whose Host header is not loopback, so a local nginx forwarding the public Host got 403 "invalid Host header" on every /mcp call. When MCP OAuth is configured, the bearer token already authenticates each request, so the SDK's localhost protection is turned off in that case only. The unauthenticated mode keeps it.
The subscriber service and the notifier were built with a nil mailer, so no confirmation or notification email ever left. Both now use the SMTP client of the email channel, built from MAINTENANT_SMTP_*. Mails are plain text, and the SMTP exchange honours its context with a 30 s ceiling. Manual incidents (API and MCP) now notify confirmed subscribers on creation, on each update and on resolution, through the same announce path as automatic incidents. A resolving update sends one resolution mail, not an update too. The in-memory status page SMTP settings are removed (GET/PUT /api/v1/status/smtp, form, types). POST /api/v1/status/smtp/test takes a recipient and uses the environment configuration. /status/api exposes subscriptions_enabled and the public page shows the subscription form only then. POST /status/subscribe answers JSON errors (subscriptions_unavailable, confirmation_failed, ...) and its 429 matches the other limiters, with Retry-After.
- POST /api/v1/channels without "enabled" creates an enabled channel, as MCP does. - Escalation levels fire at run start + delay (cumulative), shifted only by a maintenance pause; the ack notification reaches every channel the run notified, once; ack and exhaustion messages are in English. - Trigger scope-filter refusals use EDITION_REQUIRED with feature and required_edition. - Remove trigger filter_tags, escalation policy tags and the maintenant.alert.channels label/annotation, none of which had any effect; migration 36 drops their columns on both engines. - MCP list_agents checks the multihost capability; a conformance test now holds every edition-gated tool to its description.
… trust MAINTENANT_CA_CERT everywhere Replayed spool events now feed history only: resource samples skip the threshold pipeline, container state and health changes write the timeline from the recorded state at their time without touching the current row or emitting, and a replayed container absent from the inventory is kept archived. An agent whose identity the server revokes or no longer knows now enrolls a fresh identity once with the configured token, or exits non-zero with a message naming MAINTENANT_ENROLLMENT_TOKEN. The stream handler also ends when the receive side closes, so a refusal is seen without waiting for a send. The agent probes each labelled endpoint at its own interval and timeout. MAINTENANT_CA_CERT now applies to the agent gRPC client, webhooks and channels, SMTP STARTTLS, license, OSV, changelog, EOL and registry clients. Half a gRPC TLS keypair stops startup, Community logs why no agent listener starts, and the commercial set wires the multi-host extension again.
…r stale kubeconfigs - Each event stream builds and starts its own informers after registering them; before, the factory was started empty at connect time and no Kubernetes event ever reached the server. - The stream ends when the API server misses three probes in a row (15 s apart), so the supervisor goes degraded and reconnects as it does for Docker. Any API answer, a refusal included, counts as reachable. - Container events log a shortened external ID without slicing past the end of short Kubernetes IDs such as "ns/pod", which now reach this code. - Auto-detection from a kubeconfig falls back to Docker when that cluster is unreachable; in-cluster and MAINTENANT_RUNTIME=kubernetes still wait. - Helm chart 1.3.0, appVersion 1.8.0.
Subscribing an address already on file failed with a 500 on its unique constraint. The store now upserts a pending subscription: a new or still unconfirmed address gets a fresh confirmation token and a new 24 h window, a confirmed one is left as is. The confirmation email goes out in the background, so POST /status/subscribe gives the same code, body and delay whether the address is new, pending or confirmed, and whether the mail server answers or not. The confirmation_failed code is gone. buildMIME encodes the Subject header as an RFC 2047 encoded word, so accented incident titles and alert messages reach mail clients intact, for the status page and the email channel alike.
…d Kubernetes service insights Daily uptime is now rolled up into endpoint_, heartbeat_ and container_uptime_daily (migration 35, both engines) before the raw rows are purged, and kept 365 days; the current day is still computed live with portable SQL. Heartbeat days count completion pings only, exit code 0 as success. Transition purge keeps each container's latest transition. A container marked maintenant.ignore raises no alert and declares no endpoint or certificate, locally, through agents and from Kubernetes annotations; label-derived fields are refreshed on reconcile. Swarm task containers read their service labels (deploy.labels) under their own and are grouped by stack; the unused task-to-container mapping is removed. LoadBalancer and NodePort Services exposing a local workload raise security insights; missing_network_policy is removed.
…gle errors - Triggers: an absent "enabled" means enabled on creation and keeps the stored value on update, in REST and MCP alike. - Triggers: the advanced-filters gate applies to a scope filter the request changes, so a downgraded instance can still switch a scoped trigger on or off; a failed toggle or delete is shown in the trigger list. - Escalation editor: loads, shows and sends back the policy's scopes instead of wiping them on save. - AcknowledgeButton labels are in English, like the rest of the interface.
The clientset and the metrics client were plain fields, rewritten by every connection while logs, topology, stats and the event stream read them from other goroutines. They now live together in an atomic.Pointer, set once per connection and read through a single accessor; before the first connection the accessor returns an error instead of a nil client.
Fix the Caddy and nginx examples (public and MCP routes were behind auth), add /assets and MAINTENANT_TRUSTED_PROXIES, document CSRF protection, exact CSP and CORS, rate limiters, MCP OAuth, agent gRPC, demo mode, SSRF guard and secrets at rest, and cover release binary verification in SECURITY.md.
… fix false Pro comments The API reference now lists every registered route with its real edition gate, parameters, bodies and error codes, the full error format and the events the server actually emits. The architecture page covers the modes, the agent protocol, the packages, the flows and the retention as they are. The Pro labels on routes, swarm handlers and events that no edition check backs are removed.
…pdate labels to agents and Kubernetes A floating tag was compared with the digest seen at the previous scan, so an update was raised once and resolved on the next scan without any action. The scan now compares the tag's multi-platform digest (or one of its platform manifests) with the digest the container actually runs: Docker repo digests, the new repo_digests field agents send with their inventory, and the imageID of Kubernetes pods. The baseline only stands in when the runtime cannot tell, and is kept as the multi-platform digest so a rollback never pins the server's platform. maintenant.update.* labels of agent containers are kept in memory from their events, and Kubernetes workloads read their maintenant.update.* annotations. The Kubernetes update command for a republished tag now sets the image to tag@digest, since setting an unchanged reference rolls nothing out, and names the pod container instead of the workload.
Document every environment variable the code reads, grouped by area, with exact defaults, the license offline behavior, the exact telemetry payload and the PostgreSQL and SQLite storage roles. Rewrite the native install page for the offline install, the env file merge, custom paths, the docker group, version pinning and exit codes, and the container page for image tags, the entrypoint and the hardened Kubernetes pod. Fix the flag help texts, .env.example and the installer README.
… audited code Container states, restart loop, ignore label and compose one-offs; real endpoint defaults, status versus alert thresholds, manual endpoints and quotas; heartbeat deadlines, statuses and routes; certificate severities, hostname_mismatch and Personal OCSP; resource alert body and sampling; host OS support states and Updates page; multi-host editions, spool replay, re-enrollment, CA bundle, Kubernetes Deployment agent and sentinel agent_id; agent setup installer, RBAC and variables; PostgreSQL local hosts, TLS refusal, MAINTENANT_DB and Helm values.
Troubleshooting now quotes the real log messages and covers the current pitfalls: loopback bind, busy port, degraded runtime, Kubernetes fallback, edition refusals, cross-origin refusals, trusted proxies, MCP startup, revoked agents, PostgreSQL TLS and long migrations. The cloud guides describe the data bind mount, the multi-host setup (Personal edition, gRPC listener and certificates), the reverse proxy settings, the socket warning and backups without sqlite3 in the image, and fix the ordering and per-provider inconsistencies. The cloud-init header lists all five providers.
…, fill container Swarm fields - The Kubernetes alert checker now runs in the local cluster's reconcile loop. replica_health, crash_loop and node_condition reach the alert engine with an entity id, fire once, escalate, and resolve when the condition clears, the object is ignored or gone, including alerts left by a previous run. Each transition broadcasts kubernetes.workload_changed, pod_changed or node_changed. Jobs raise no replica alert, a node's conditions share one alert, and a node whose Ready condition is Unknown counts as NotReady. - Pods created by a Deployment's ReplicaSet point to the Deployment. - A workload or pod annotated maintenant.ignore, and a Swarm service labelled maintenant.ignore, raise no alert; an alert raised before the label is resolved. - Swarm service alerts carry the service id, so two services no longer share one alert; alerts of services the manager no longer lists are resolved. - Enabling Swarm while the server runs builds, wires and starts the same services as at boot, so its events and alerts go out. - Swarm task containers read their service, node and slot from the labels Docker sets on them, locally and through agents, and the container detail returns them. swarm_service_mode and swarm_desired_replicas, which no container label carries, are dropped (migration 37, both engines).
…race-free activation - Swarm node alerts carry the node ID, so two nodes no longer share one alert; alerts of nodes the manager no longer lists are resolved, as for removed services. - A rolling update that completes resolves the service's update_rollback and update_stalled alerts; ignoring the service resolves them too. - Replica, crash-loop and rolling-update alerts still active from a previous run are taken over at startup, and when Swarm is enabled later, so they resolve once their condition clears. - The Swarm manager's services are built together and published through one atomic pointer, the cluster through another: enabling or leaving Swarm while events flow is race-free, and the API reads the update tracker and crash-loop detector in place at request time. The task tracker and the handler's replica checker, which nothing read, are gone.
quorum_degraded is raised once when fewer managers than the quorum are ready, instead of on every node reconcile, and resolved when the quorum is reached again. An alert still active from a previous run is taken over at startup, and when Swarm is enabled later, like the other Swarm alerts.
…agent container exits, weigh heartbeat uptime by time - certificates: expired, chain_invalid, hostname_mismatch and ocsp_revoked now recover once a scan no longer shows them, not only expiring - certificates: endpoint probes and agent scans store a result once per monitor interval or when the outcome changes, instead of every probe - resources: CPU and memory alerts fire and resolve independently; leaving both_alert no longer sends an unmatched both_threshold recovery - resources: live agent samples feed the current snapshot and the live top - containers: an OOM kill (137 with OOMKilled) is a crash, docker stop stays completed; agents carry the exit code and a new oom_killed proto field in their inventory and events, and an EXITED report without a code is a stop - embedded agent dials grpc:// when the listener runs in h2c mode - heartbeats: daily uptime replays pings against the deadline, so missed deadlines lower it and days without pings get a value - /ping records the client address resolved like the rate limit
Status page: - validate component overrides, incident severities and statuses against the model, on REST and MCP, with an explicit 400 - list upcoming maintenance soonest first; the admin list shows running, then upcoming, then completed windows - deleting a running maintenance window ends it like its scheduled end (incident resolved, components released, visitors told) - a maintenance window no longer destroys a manual override: it is kept (migration 38) and restored when the last window holding the component ends; overlapping windows keep the component under maintenance - the Atom feed carries ongoing incidents as well as recent ones - SVG assets are read whole and refused when they carry scripts, event handlers, javascript: links, foreignObject or embedded HTML; an SVG without an XML declaration is recognised Escalation: - a policy holds at most 5 levels, enforced by the service for REST and MCP and reported in the plan limits the editor now reads - disabling a policy stops its runs with stopped_by_policy_disabled - drop the never-written skipped_maintenance delivery status from the contract and the schema (migration 38, SQLite rebuild after conversion) Security posture: - every infrastructure scoring evaluates the score threshold, and a background check scores every 5 minutes when a threshold is set - the posture page help names the weighted categories of the scorer Transport and config: - SMTP fails when an announced STARTTLS fails, and uses implicit TLS on port 465 - SSE streams send a keep-alive comment every 25 s - an invalid MAINTENANT_MAX_BODY_SIZE stops startup
The API returns affected components as {id, name}, but the front typed them
as {component_id, name}, so editing a window sent undefined component ids.
Align both component ref types and the editor with the API.
The admin API now answers 400 when component_ids holds an empty, unknown or
repeated id, on incident and maintenance create and update, instead of a 500
that left a half-written window.
Swarm is open in every edition, with real alert severities, delays and resolutions, deploy.labels and the agent path. The Kubernetes guide now carries the actual RBAC, namespace defaults, annotations, alerts, runtime behaviour and Helm values. The labels reference lists every label read, with real defaults and update label semantics.
…quotas - Broadcast endpoint.* events to SSE and webhooks; notify the status page on status changes only - Report MAINTENANT_STATUS_URL as status_url in GET /api/v1/edition - Record webhook deliveries: last status, failure count, auto-disable after 10 failures, reset and re-enable on success; store webhook timestamps as epoch seconds - Send the webhook test with the headers and signature of a real delivery, and record it like one - Read the endpoint, heartbeat and certificate caps from the running edition at each creation - Compute container uptime over every window the edition's history cap opens - Answer a duplicate exclusion with the stored row and 200 - Start the escalation retention loop once, without an error at shutdown - Emit status page events on the admin SSE bus as well - Drop the unused endpoint orchestration_group filter and the untracked update count; fill agent_id on check results; standard error body on the logs stream 503; snake_case restart alert event - Updates: keep the image reference when a pulled tag moves, kubectl commands for bare pods, purge the digest baselines of gone containers, keep scanning and warn once when the image list is refused - Retry runtime connection errors with backoff and log their cause - Correct comments that tied Personal or open capabilities to Pro
create_incident and create_maintenance wrote the incident or the window, then failed on an unknown component id and left it without its components. The MCP server now receives the status component store and runs the same check as the REST API (status.CheckComponentIDs) before any write, with the same message.
…component changes - Deliver an alert's notifications to one channel, and a webhook subscription's events, on a single notifier worker, so a recovery can no longer overtake the alert it resolves while that alert's delivery waits or retries - Emit status.component_created, status.component_updated and status.component_deleted (component id only) followed by status.global_changed on both SSE buses when an administrator changes a component - Refresh the status admin components and the edition quotas on those events
- Match status.component_changed on component_id, as the server sends it, and apply its name, status and monitor breakdown in place - Apply status.global_changed in place instead of refetching, which wiped the breakdown the previous event had just brought - Refetch on component creation, update and deletion, alongside the incident and maintenance events
…cluster counts - replica_unhealthy: service events feed the 5-minute ReplicaHealthChecker instead of raising at once, so a scale-up no longer leaves an alert open - crash_loop: count the tasks Swarm marks failed on every node, read from the task list, instead of every local die; shutdowns and SIGTERM stops do not count - SwarmCluster: fill manager/worker counts and creation date from docker info, refreshed on every recheck - detection is armed for any Docker runtime and runs as soon as it connects - quorum_degraded: raised when the node list fails because the swarm has no leader - emit swarm.node_updated when a node joins or changes role, host, engine or address - maintenant.update.registry now queries that registry as a mirror of the image
… Swarm services Container health alerts and their recovery had no entity name, so every notification was titled "Alert: ". The health event now carries the container name, and a test walks every alert.Event built in internal/app for an EntityName. The stale update cleanup treated a container whose registry could not be reached as scanned: its pending update was deleted and its alert resolved, then the next scan raised both again. A container whose scan failed now keeps its update and its alert until a scan reaches its registry. Scanner.Scan returns only the containers with an update, so the CVE pass meant for every container ran on those alone and left the others "not evaluated" in the posture. The enrichment now also gets the running image of every other container the scan lists. A Swarm task got the standalone docker stop, rm and run commands, which Swarm undoes by rescheduling the task. Its update, rollback and fix commands now run docker service update --image on the task's service, with the Kubernetes reference rule: the tag, or the tag and its new index digest for a republished tag.
… guides with the audited code Endpoint events, webhook delivery tracking, status_url, SSE keep-alive and new events, SMTP implicit TLS, strict MAINTENANT_MAX_BODY_SIZE, escalation limits, status page validation, Swarm task commands, runtime reconnection, and the help texts and category of the license flag.
Covers Swarm alerts and detection, certificate and resource alert recovery, status page events and validation, escalation limits, endpoint events, update commands for Swarm tasks, SMTP TLS, MCP parameters and the three-attempt notification retries.
…fig gaps - agent: an unreachable Docker or Kubernetes no longer blocks startup. The agent enrolls, streams host metrics and OS identity, and retries the runtime in the background with backoff (1s to 1m), starting container monitoring as soon as it answers. - swarm: a crash-looping service that becomes maintenant.ignore resolves its crash_loop alert at once, like replica_unhealthy and update alerts. - config: MAINTENANT_SECURITY_SCORE_THRESHOLD and --securityScoreThreshold share one rule: a whole number from 0 to 100, 0 or unset disabling the alert; anything else stops the startup. - help: drop the em dash from the --help header. - deploy: remove MAINTENANT_TELEMETRY_DATADIR and MAINTENANT_TZ from compose.test.yml (read by nothing), replace the MAINTENANT_WEBHOOK_URL example in the Helm values with MAINTENANT_BASE_URL, and name Personal next to Pro for the license key.
The writer's stop called WaitGroup.Wait while submitters could still Add from zero. The race detector reports that misuse, and it failed TestStart_FailsOnBusyPort under load: the early return cancelled Start's context while the resource rollup was writing. The same Wait also deadlocked a write queued with a context the stop does not cancel. Submitters now select on a stopped channel instead, and the vacuum test no longer starts a writer that openTestDB already runs. Start also binds its HTTP port before any service starts, so a taken port fails the boot before it has side effects.
…the public routes - Send component changes of a hidden component to the dashboard bus only; created, updated and deleted events reach the public page only when the component is or was visible there. - Return the per-monitor breakdown in /status/api, the same fields as the status.component_changed event, so a reload keeps it. - Parse the Content-Type of POST /status/subscribe as a media type and answer 415 unsupported_media_type to anything but JSON or a urlencoded form. - Build the Atom feed links from MAINTENANT_STATUS_URL, or MAINTENANT_BASE_URL followed by /status, instead of the request host. - Refuse a maintenance window update whose end precedes its start. - Answer the lower case internal code on a failed component creation, like the rest of the status admin routes. - Drop the unused DeleteIncidentsOlderThan: incidents are kept until deleted.
The runtime an agent reports was frozen at enrollment: an agent enrolled before Docker, Swarm or Kubernetes answered, or a Docker host that later joined or left a Swarm, kept showing a wrong runtime. - proto: add RuntimeMsg (AgentEvent field 19). An older server ignores it and keeps the enrollment value; an older agent never sends it. - agent: report the runtime once it answers, keep it as current state in the spool so every new stream restates it, and recheck Swarm membership every minute on a Docker host. Joining or leaving a Swarm restarts collection under the new runtime; leaving sends an empty Swarm topology so the server drops the cluster it held for the agent. - server: store the reported runtime in agents.detected_runtime and broadcast agent.updated when it changes, so /api/v1/agents, MCP and the UI follow it.
…Swarm leaves
- The Docker event stream closes once the daemon stops answering after a
cut, so the supervisor marks the runtime disconnected, broadcasts
runtime.availability_changed and reconnects. A cut while the daemon
answers resumes right after the last event (the resume point was far in
the future). On the agent, a closed stream ends the runtime collection,
which restarts through the same loop as a Swarm switch once the runtime
answers again: inventory resent, one new subscription.
- The header banner and the containers store read the runtime store, which
follows /runtime/status and runtime.availability_changed; the
runtime.status listener waited for an event never sent.
- Swarm detection logs only role changes. Deactivating Swarm stops the
manager loops; reactivation starts a new manager.
- swarm.status is sent when the Swarm state changes (activation,
deactivation, node counts), no longer at boot before any client listens.
- GET /containers/{id}/logs answers 503 RUNTIME_UNAVAILABLE while the local
runtime is disconnected, like logs/stream.
- Drop recent_events from /swarm/dashboard: it had no source.
- Help: --runtime lists swarm for agents; the spool budgets say what 0
does and that both must be 0 to turn the spool off.
The role took the first 16 hex characters of a UUIDv7, which are its millisecond timestamp and sub-millisecond sequence with no random bit. Test processes sharing one server could draw the same name and fail CREATE ROLE. The whole UUID carries its 62 random bits, like the test database names already do.
…ts quiet - REST, MCP and the security posture acknowledge through Engine.Acknowledge: stored once, broadcast as alert.acknowledged, escalation stopped. An unknown alert id now answers 404 instead of 500. - A severity raise on an acknowledged alert keeps the acknowledgment, never starts an escalation again and is notified once, as "Severity raised from X to Y". An escalation run that outlives an acknowledgment stops at its next cycle. - escalated_at is stamped each time an escalation level goes out. - Retention deletes resolved and silenced alerts 90 days after their resolution (or creation when they have none), never an active one. - GET /api/v1/alerts pages with before and before_id, so alerts fired in the same second are neither lost nor repeated between pages. - MCP list_alerts leaves acknowledged alerts out, as GET /alerts/active does, and returns at most 100 recent alerts. Fix the escalation scope kinds and the misplaced doc comments. - Drop the 25s notifier backoff that no retry ever reached.
- An auto_incident component that is hidden no longer opens or updates a public incident; an incident opened while it was visible still resolves. - Incident and maintenance component references carry the component's visibility. /status/api, the public status.incident_created and status.maintenance_* events and the subscriber emails name only the visible components; the dashboard bus keeps the full list. - Drop the resolved-incident query GetPageData ran on every /status/api call although nothing read it.
…tbeats and endpoints Updates - Retention no longer deletes pending image updates behind their alert; scans still named by a pending update are kept so its scan_id never turns NULL. - The enrichment log reports the running edition instead of "(Pro)". - The risk score reads network exposure (security insights) and the last day's restarts; criticality and dependents, which had no source, are gone. - Posture scoring reads certificates and updates once per pass, indexed by container, instead of once per container. Certificates - Chain validation no longer checks the hostname, so a name mismatch no longer also reports an invalid chain. - Correct the comment on the crossed expiry threshold. Resources - Network totals are real rates in bytes per second, computed from two samples; unavailable (-1) or reset counters no longer count. - Top consumers over 24h and more include the hour or day in progress. - The cap of 20 containers lives in the service, so MCP gets it too. Heartbeats - Pauses are recorded (migration 39) and left out of the daily uptime; a resume or a ping restarts the deadline. Pausing a paused heartbeat is 400. - Deleting a heartbeat removes it and its history; the active column goes. - MCP pause/resume detect a missing heartbeat with the real sentinel error. Endpoints - Creation and update reject a TCP target that is not host:port and an HTTP target without an http(s) scheme and host. - A label interval below the 5s API minimum is raised to it, with a warning.
…atus page, updates, resources, heartbeats and endpoints
…ration The daily uptime reads, the top consumers ranking and every cleanup of the retention pass each called time.Now() on their own. A test that built its fixtures from an earlier reading saw its days shift by one when UTC midnight passed in between (and its ranking lose the hour in progress at any hour boundary), which made TestContainerDailyUptime fail on PostgreSQL at 23:58. These operations now take their reference time from the caller: the HTTP handler and the resource service pass time.Now(), the retention pass reads the clock once and hands it to every rollup and cutoff, and the resource rollup derives its hourly and daily buckets from a single reading. The tests pass the same reference to their fixtures and to the code, and assert on complete days rather than on a day in progress that is still empty at 00:00:00.
…nted The Pro capability was advertised but no router was ever wired into the engine. Trigger scopes already route an entity's alerts to chosen channels.
…uptime window, patch base packages The MCP authorize endpoint now redirects to a URL rebuilt from the configured allowlist or a canonical loopback host and numeric port, never to the raw redirect_uri. The daily uptime window is bounded with an explicit guard, and the runtime image upgrades its Alpine packages so OpenSSL fixes land at build time.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An audit of the documentation against the code surfaced behaviour that drifted from what the product promises. This PR fixes it and brings every page back in line.
Fixes
install.shrestarts on update, keeps the env file and installs offline.MAINTENANT_SMTP_*, hidden components never reach a public surface, incident and maintenance inputs are validated.maintenant.update.*labels apply to agents and Kubernetes./apiare refused,MAINTENANT_CA_CERTcovers every outbound TLS client, public SVG assets are parsed in full,/mcpworks behind a local reverse proxy.Behaviour changes
filter_tags, themaintenant.alert.channelslabel,GET/PUT /api/v1/status/smtp, and thealert_entity_routingcapability, which no code ever implemented.POST/PUT/PATCH/DELETEon/apireturn 403CROSS_ORIGIN_REFUSEDunless the origin is listed inMAINTENANT_CORS_ORIGINS.MAINTENANT_MAX_BODY_SIZEorMAINTENANT_SECURITY_SCORE_THRESHOLD, or half a gRPC TLS key pair, now stops startup.Docs
SECURITY.md, the README and every page underdocs/were checked against the code: the API reference lists the registered routes, configuration covers every variable, and the reverse proxy examples were tested.