Add a CloudNativePG workspace: fleet triage, protection evidence, declarations and composed CNPG detail - #1921
nadaverell wants to merge 17 commits into
Conversation
… composed Cluster summary - GET /api/cnpg/workspace: every CNPG kind plus owner-validated instance Pods, authorized per kind (namespaced kinds fall back per namespace; ClusterImageCatalog needs a cluster-scope list), with coverage states, CNPG issues from the Issues engine and the no-schedule audit finding withheld where coverage is missing. - k8s-ui: buildCNPGFleet derives per-cluster facts (instances, replication, protection as separate schedule/destination/last-backup/WAL/recovery-window/ restore-validation facts, declarations) without inventing values: replication lag is unknown without runtime data, restore validation is never green, and unreadable kinds read "No access" rather than none. - ResourcesSidebar gains optional category workspaces (destinations above a collapsible Resource kinds block, grouped by API group). - WorkloadView gains an optional composed summary: Overview shows it and the kind's renderer moves to a Spec & status tab. - /cnpg: Overview fleet with Needs attention / All, problem-category chips, search, namespace chip and a URL-backed drawer (?drawer=kind:group:ns:name).
…ens and composed summaries Screens (/cnpg/protection, /declarations, /pooling, /operator): - Protection keeps schedule, destination, last successful backup (with its source), WAL archiving, recovery window and restore validation as separate facts; ObjectStore upload health is labelled inferred with its evidence. - Declarations groups Databases, Publications, Subscriptions and managed roles by cluster, separating applied, not applied and pending, with controller errors and GitOps source. - Pooling lists Poolers with readiness; connection pressure reads "Not measured" until PgBouncer metrics are read. - Operator shows operator/plugin workloads and versions, image catalogs and their visible users, and operator configuration (Secret contents never read). Composed Overview summaries (Spec & status keeps the existing renderers) for Backup, ScheduledBackup, ObjectStore, Database, Publication, Subscription, Pooler and image catalogs; drawer trail back-link on workspace screens. GET /api/cnpg/operator discovers operator and plugin Deployments and their config references, gated per namespace on list deployments/services and get configmaps; it does not follow the namespace view filter. Workspace hardening from review: - issues are read flat, so Pod evidence only reaches callers with Pod access; - instance Pods must be controller-owned by the visible Cluster's UID; - partial coverage names denied namespaces only when the caller supplied the candidates, and always reports the allowed ones; - unobserved managed roles are pending, unreadable schedules or Poolers are unknown rather than absent, restore evidence requires a ready restored cluster resolved by exact reference, and an unreadable Cluster list is not reported as "no clusters".
…ctivity
- Every CNPG kind now has a CNPG-framed full detail at /cnpg/<plural>/<ns|_>/<name>,
reached from row Open, drawer expand and a redirect from /workload/...: the
workspace sidebar stays highlighted with the object nested under its home, a
return link names the page the drilldown came from (history state, kept
across tab changes), and a crumb names the object's place.
- Detail URLs pin the kube context (ctx=, added on first view). After a context
switch the page says the object is not in the new context and offers
Switch back / Go to Overview instead of loading a same-named object.
- A Cluster adds Protection (its recovery evidence), Activity (replacing
Timeline) and merged instance Logs; WorkloadView gains additive extraTabs and
the logs viewer an initialPods selection.
- GET /api/cnpg/clusters/{ns}/{name}/logs (+/stream): logs from controller-owned
instance Pods, gated on get clusters, list pods and get pods/log, with CNPG's
JSON lines parsed into level/logger/message.
- GET /api/cnpg/clusters/{ns}/{name}/activity: timeline events for the Cluster,
its instance Pods and CNPG objects attributed to it, per-kind gated. The
timeline now records the owning cluster (cnpg.io/cluster label or
spec.cluster.name) at ingestion so deleted Backups and declarations stay
attributed; earlier history is reported as incomplete.
- The drawer trail back link moves to the drawer shell so it survives hops to
Pods and Secrets.
Review fixes on the workspace screens and summaries: empty states follow each
kind's coverage, Declarations separates not applied from pending, GitOps
provenance reads "not recorded" instead of "applied directly", ObjectStore no
longer asserts recoverability, a Backup's destination inferred from the
Cluster's current config is labelled, schedule runs require the owner UID,
partial catalog coverage keeps the known users, and CNPG's six-field cron is
shown verbatim.
Docs: docs/cnpg.md (destinations, navigation, certainty table, access).
… redirects - Activity rows from Kubernetes Events now also require list events in the namespace; they carry their subject's kind, so the subject check alone let Event messages through. - The /workload redirect to a CNPG detail keeps ctx, pod and every other parameter, and maps a Cluster's timeline tab to Activity. - The drawer trail back link keeps the page's return state. - Activity describes its earliest linked child event as a recorded boundary, not a completeness guarantee.
PR Summary by QodoAdd a read-only CloudNativePG workspace and composed object details
AI Description
Diagram
High-Level Assessment
Files changed (60)
|
Code Review by Qodo
1.
|
…severity - Replication and ObjectStore-sourced facts read "No access" when Pods or ObjectStores are unreadable, instead of asserting absence. - Restore evidence only counts barman-cloud plugin recovery sources. - Inferred ObjectStore health is "Accepting uploads" only when every user cluster reports archiving; the failed-backup window uses completion time. - Log level detection reads CloudNativePG's nested PostgreSQL severity, and live stream lines keep their instance role label. - A restarted instance log stream resumes from the last delivered line instead of replaying its tail; activity Pod rows must match the live Cluster's UID; activity and stream failures are logged. - Workspace components use the shared Tooltip rather than native titles.
Conflicts: - CLAUDE.md: main condensed the endpoint list; the CloudNativePG endpoints move to docs/cnpg.md#api behind a one-line pointer. - useLogBuffer.ts: main moved level detection to utils/log-level.ts; the CloudNativePG record.error_severity rule moves with it into selectLevelField.
CloudNativePG's record.error_severity is one of PostgreSQL's severity names; another logger's numeric or custom record field no longer overrides the line's own level.
…nreadable Backups - PostgreSQL writes every DEBUGn level as DEBUG in error_severity. - Restore validation says the Backup a recovery names couldn't be read, instead of None recorded, when Backups are denied in the namespace.
Recovery replays archived WAL past the latest base backup, and no status reports the newest recoverable point, so the window shows its earliest point and reaches the newest archived WAL. It reads not advancing only while WAL archiving fails; a failed base backup no longer implies it.
Recovery replays archived WAL past the last base backup, so a failed backup makes recovery slower, not shorter; only archiving stops it.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 47f1b1b. Configure here.
The failing-backups banner names servers whose cluster also stopped archiving and drops the claim that recovery reaches the newest WAL, and the row keeps its archiving-stopped note while backups fail.
With failures and no success the store holds nothing to restore from, and retention isn't shown to move the earliest point.
The plugin keeps a successful backup when it fails to update the ObjectStore status, so an absent success time leaves recoverability unestablished rather than proving there is nothing to restore.
Status can miss a success whose update failed, so a failure is stated as recorded after the last recorded success, and an entry without timestamps reads as a recovery point not reported.
It is internal planning; docs/cnpg.md stops pointing at it.

Summary
CloudNativePG support in Radar today is ten separate CRD lists and renderers. Answering "which Postgres cluster needs attention, why, and what do I look at next" means stitching Clusters, Backups, ObjectStores, declarations and Pods together by hand. This adds a CloudNativePG workspace inside Resources that does that stitching. It covers fleet triage, recovery evidence, declared vs reconciled state, pooling and the operator. Every CNPG object gets one composed detail, whichever way you reach it.
The workspace is read-only. Every value it shows is something the cluster actually reports, labelled with its source. When something isn't reported it says so, never zero, "none" or green. The full contract is in
docs/cnpg.md.What changed
Workspace screens (
/cnpg/*, inside Resources)The Resources sidebar's CloudNativePG group gains a Workspace block above the exact kinds. Kinds sit under a collapsible "Resource kinds" block, grouped by API group.
Overview is the cluster fleet, defaulting to Needs attention. It shows instance pills, replication, protection, declarations, PG version and the top problem, with problem-category chips, search and a namespace chip. Each row inspects in the drawer and has explicit Logs and Open → actions.
Protection keeps these as separate facts per cluster, each with its source:
It also lists failed backups from the last 7 days, destinations (ObjectStore upload health, labelled inferred with its evidence) and schedules. Restore validation is never green: Kubernetes records no restore tests.
Declarations lists Databases, Publications, Subscriptions and managed roles by cluster. It separates applied, not applied and pending, and shows controller errors and the GitOps source ("not recorded" when there isn't one).
Pooling lists Poolers and the clusters they front. Connection pressure reads "Not measured".
Operator shows operator and plugin Deployments with versions and readiness, image catalogs and their visible users, and operator configuration. Secret contents are never read.
One composed detail per object
WorkloadViewgains an optional composed summary. For CNPG kinds, Overview is the summary (facts grouped by task, e.g. Backup → Outcome / Relationships) and the existing renderer moves to a Spec & status tab, so nothing repeats./cnpg/<plural>/<ns|_>/<name>. You reach it from Open, from drawer expand, or by redirect from/workload/.... It keeps the workspace sidebar with the object nested under its home destination.record.error_severityis read by the sharedutils/log-level.ts, so every log viewer levels CNPG lines the same way).?pod=preselects an instance.docs/cnpg.md#api;CLAUDE.mdpoints there.Navigation
?drawer=kind:group:ns:name. Links inside it build a trail with a "← previous object" link in the drawer shell.ctx=<kube context>, added on first view. After a context switch the page says " is not in ", with Switch back and Go to Overview, instead of opening a same-named object from another cluster.Backend
GET /api/cnpg/workspacereturns every CNPG kind plus instance Pods, authorized per kind, with coverage states:full,partial(including the allowed namespaces),denied,syncing,errorornotInstalled.ClusterImageCatalogneeds a cluster-scopelist.cnpgNoDeclarativeBackupaudit finding. Both are withheld where coverage is missing.GET /api/cnpg/operatordiscovers operator and plugin workloads and their config references.get configmaps; Secrets are never read.GET /api/cnpg/clusters/{ns}/{name}/logs(plus/logs/stream) returns merged logs from owner-validated instance Pods. It is gated onget clusters,list podsandget pods/log.GET /api/cnpg/clusters/{ns}/{name}/activityreturns timeline events. Each kind is gated onlist <kind>; Event rows additionally needlist events.ExtractLabelsnow retainscnpg.io/cluster(orspec.cluster.name) on CNPG objects and Pods, so deleted Backups and declarations stay attributed. No storage schema change.Shared package (
@skyhook-io/k8s-ui), all additiveResourcesSidebar: optionalcategoryWorkspaces, with a pass-throughsidebarCategoryWorkspacesonResourcesView.WorkloadView: optionalrenderSummaryandextraTabs, plus a'spec'value inWorkloadTabType.WorkloadLogsViewer: optionalinitialPods.components/cnpg:buildCNPGFleet, the relation helpers and the summaries.Consumers that pass none of these props see no change. That includes Radar Hub, which only needs to opt in if it wants the workspace.
Worth reviewing closely
internal/server/cnpg_workspace.go: coverage, issue withholding, and not naming namespaces.docs/cnpg.md, andpackages/k8s-ui/src/components/cnpg/workspace.ts+relations.ts, which implement it.?drawer=two-way sync inweb/src/components/cnpg/CNPGView.tsx.App.tsx, which preservesctxon/cnpg/routes.Not included (needs a product/RBAC decision first)
pods/proxy) and/or Prometheus. Replication reads "lag unknown" until then.These ship in #1922, stacked on this PR.
Testing
main; CI is green on the head.make tscpasses.go test ./internal/server/ ./internal/timeline/ ./internal/k8s/andpkg/timelinepass, including new handler tests for:make test: one unrelatedcmd/desktopenv test (TestEnrichEnvPrecedenceAndDiagnostics) failed under the full parallel run and passes on its own; nothing incmd/changed.make cnpg-demoand walked these paths:Note
Medium Risk
New read-only APIs touch per-kind RBAC, merged pod logs, and timeline attribution; mistakes could leak withheld data or cross-namespace hints, though behavior is heavily tested and documented.
Overview
Adds a read-only CloudNativePG workspace under Resources (
/cnpg/*) that composes fleet triage, protection evidence, declarations, pooling, and operator views from one workspace payload instead of separate CRD lists.Backend: New
GET /api/cnpg/workspace(per-kind coverage, issues, audit),GET /api/cnpg/operator, cluster logs (snapshot + SSE with instance validation), and activity (timeline withcnpg.io/clusterattribution). Routes are registered inserver.go; pod log entries gain optional parsedlevel/logger/message.Frontend (
@skyhook-io/k8s-ui): Newcomponents/cnpg(fleet derivation, relations, per-kind Overview summaries) and optional hooks onWorkloadView/ sidebar for composed CNPG detail (Overview + Spec & status, Cluster Protection/Activity/Logs).Docs:
docs/cnpg.md(certainty contract, navigation, API); README/CLAUDE/integrations updated for expanded CNPG kinds and workspace entry point.Reviewed by Cursor Bugbot for commit a027865. Bugbot is set up for automated code reviews on this repo. Configure here.