Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
0c80dce
Add CloudNativePG workspace: fleet Overview, sidebar destinations and…
nadaverell Sep 28, 2026
91241a9
Add CloudNativePG Protection, Declarations, Pooling and Operator scre…
nadaverell Sep 28, 2026
b698092
Add the CloudNativePG full detail, merged instance logs and cluster A…
nadaverell Sep 28, 2026
a0cab4a
Gate CNPG activity Event rows on list events and keep context through…
nadaverell Sep 28, 2026
3867b1c
Fix CNPGView hook dependencies and simplify the summary test's text e…
nadaverell Sep 29, 2026
d4e5889
Address PR review: coverage-aware facts, log stream resume, CNPG log …
nadaverell Sep 29, 2026
382bfea
Word an unreadable ObjectStore's last-backup fact by its coverage state
nadaverell Sep 29, 2026
57e8b47
Merge main into feature/cnpg-workspace
nadaverell Oct 2, 2026
ed03116
Read a nested record's level only when it is a PostgreSQL severity
nadaverell Oct 2, 2026
3eb0685
Read PostgreSQL's DEBUG severity, and don't claim no restore behind u…
nadaverell Oct 3, 2026
e732198
End the recovery window at WAL archiving, not the last base backup
nadaverell Oct 3, 2026
47f1b1b
Don't read a failed base backup as recovery that stopped advancing
nadaverell Oct 3, 2026
9fec7a4
Say archiving stopped too when an ObjectStore's backups are failing
nadaverell Oct 3, 2026
545f74e
Claim recovery from an ObjectStore only after a base backup succeeded
nadaverell Oct 3, 2026
c1572dc
Read a missing ObjectStore success as not recorded, not as no backup
nadaverell Oct 3, 2026
a6d98ea
Word ObjectStore backup outcomes as recorded, not as complete history
nadaverell Oct 3, 2026
a027865
Keep the CNPG workspace planning doc out of the repository
nadaverell Oct 3, 2026
135d8fd
Merge remote-tracking branch 'origin/main' into feature/cnpg-workspace
nadaverell Oct 4, 2026
e860514
CloudNativePG: Logs tab for Clusters without related Pods; one store-…
nadaverell Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ Not everything is in this file. The following files contain critical details tha
| Adding or modifying **HTTP endpoints** | `internal/server/server.go` — all routes are defined here — **plus** the handler's doc comments (why the route is gated the way it is lives there; copy the gate of the closest sibling only after reading it) and the integration's section in [docs/integrations.md](docs/integrations.md) |
| Adding or modifying **CLI flags** | `cmd/explorer/main.go` — flag definitions and defaults |
| Adding a **new CRD integration** (renderer, topology, discovery) | [docs/INTEGRATION_GUIDE.md](docs/INTEGRATION_GUIDE.md) — full checklist with collision gotchas |
| Working on the **CloudNativePG workspace** (`/cnpg`) | [docs/cnpg.md](docs/cnpg.md) — destinations, navigation (drawer trail, return label, `ctx` guard) and the certainty table: which source each fact comes from and what it reads when unknown. Data from `/api/cnpg/workspace` (per-kind coverage); derivations in `packages/k8s-ui/src/components/cnpg/workspace.ts` + `relations.ts`; screens in `web/src/components/cnpg/` |
| Working on **local per-cluster integration settings** (Metrics, Argo CD, Cost in `~/.radar/clusters.json`) | [docs/configuration.md](docs/configuration.md#local-integration-connections) — store `internal/config/profiles.go`, resolve/update `internal/connections`, activation `internal/connectionruntime`, routes `GET/PUT /api/integrations/connections`. In local mode the older `PUT /api/integrations/{prometheus,argocd,cost}` return 409 |
| Working on **GitOps** (Argo CD / Flux detail pages, operations, Terminating lifecycle, drift, per-resource health, remote destinations) | [docs/gitops.md](docs/gitops.md) — detail-page tabs, operation semantics, the Terminating severity ramp, nested navigation, single-cluster scope. Engine in `pkg/gitops/`, handlers `internal/server/gitops_handlers.go` |
| Working on an **integration's reverse-lookup or actions** (Velero, CloudNativePG, Kyverno, Argo Rollouts, …) | That integration's section in [docs/integrations.md](docs/integrations.md) + the doc comments in `internal/server/<name>_handlers.go` — both carry the per-integration gating and scope rules this file only summarizes |
Expand Down Expand Up @@ -152,7 +153,7 @@ After `make <name>-demo`, run `kubectl config use-context kind-radar-<name>-demo
- RBAC reverse-lookup: `/api/rbac/subject/{kind}/{namespace}/{name}` (ServiceAccount, plus `usedByPods`) and `/api/rbac/subject/{kind}/{name}` (User/Group) — direct + group-inherited bindings and flattened effective rules; `/api/rbac/role/{kind}/{namespace}/{name}` (`_` for a ClusterRole's namespace) — the bindings that reference it; `/api/rbac/namespace/{namespace}` — backs the Namespace RBAC section (group-only ClusterRoleBindings deliberately excluded); `/api/rbac/whoami` — `SelfSubjectRulesReview` pass-through. All gate on `list rolebindings` AND `list clusterrolebindings`: **403 when either is denied, never a silent partial view**
- Policy (Kyverno): `/api/policy/resource/{kind}/{ns}/{name}` (one resource's findings), `/api/policy/policies/{policy}` (every resource one policy recorded an outcome for). Report families are authorized **per subject scope** (`policyreports` cluster-wide ≠ `clusterpolicyreports`); findings from an unreadable family are dropped from lists AND counts, with the withheld count reported; `counts` describe the cluster while subject lists are capped and view-filtered. `/api/policy/policies/{policy}/queued` reads Kyverno's `UpdateRequest`s cluster-wide, gated on `list updaterequests`
- Velero: `/api/velero/backupstoragelocations/{ns}/{name}/backups` (what a location holds; gated on `list backups`); `POST /api/velero/{backups|restores}/{ns}/{name}/messages` (a run's warnings/errors via a `DownloadRequest` — impersonated; needs a running Velero controller and object storage reachable from Radar, and reports which one failed)
- CloudNativePG: `/api/cnpg/imagecatalogs/{ns}/{name}/clusters`, `/api/cnpg/clusterimagecatalogs/{name}/clusters` (Clusters pinned to a catalog). Cluster-scoped catalogs are referenceable from any namespace, so that route reads cluster-wide gated on `list clusters` — a view-filtered answer would report "nothing uses this" before an edit
- CloudNativePG: `/api/cnpg/...` — the workspace, operator, catalog reverse-lookups, Cluster logs and activity. Routes, gates and coverage states are listed in [docs/cnpg.md](docs/cnpg.md#api); each is gated on the caller's own access, and a partial answer says what it withheld

## Key Patterns

Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -548,7 +548,7 @@ Upgrade impact also gets list-only access to CSIStorageCapacities, FlowSchemas,
| **Strimzi** | [KafkaConnector failure evidence](docs/integrations.md#strimzi-kafka-connectors) (connector/task status) |
| **Velero** | Backup, Restore, Schedule, BackupStorageLocation, VolumeSnapshotLocation |
| **External Secrets** | ExternalSecret, ClusterExternalSecret, SecretStore, ClusterSecretStore |
| **CloudNativePG** | Cluster, Backup, ScheduledBackup, Pooler |
| **CloudNativePG** | Cluster, Backup, ScheduledBackup, Pooler, Database, Publication, Subscription, ImageCatalog, ClusterImageCatalog, ObjectStore — plus a [workspace](docs/cnpg.md) for fleet, protection and declaration triage |
| **Crossplane** | Managed Resources (any provider), Composite Resources, Claims, Provider, ProviderConfig, Function, Configuration, Composition, CompositionRevision, XRD |
| **Kyverno** | Policy, ClusterPolicy, PolicyReport, ClusterPolicyReport |
| **Sealed Secrets** | SealedSecret |
Expand Down
76 changes: 76 additions & 0 deletions docs/cnpg.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
# CloudNativePG workspace

A task-shaped view over [CloudNativePG](https://cloudnative-pg.io/) (CNPG): which PostgreSQL cluster needs attention, why, and what to inspect next — without assembling the story from ten separate CRD lists. The per-kind renderers, issue detection and audit check it builds on are described in [integrations.md](integrations.md#cloudnativepg).

The workspace is read-only. It never writes to a cluster.

## Where it lives

CNPG stays inside **Resources**; there is no new global navigation item. When the `postgresql.cnpg.io` CRDs are discovered, the Resources sidebar's CloudNativePG group gains a **Workspace** block above its exact kinds:

| Destination | Route | Job | Detail home for |
|---|---|---|---|
| Overview | `/cnpg` | The fleet: every Cluster with instances, replication, protection, declarations and its top problem. Defaults to **Needs attention**. | Cluster |
| Protection | `/cnpg/protection` | Recovery evidence per cluster, failed backups (7 days), destinations, schedules. | Backup, ScheduledBackup, ObjectStore |
| Declarations | `/cnpg/declarations` | Databases, Publications, Subscriptions and managed roles by cluster; declared vs reconciled. | Database, Publication, Subscription |
| Pooling | `/cnpg/pooling` | Poolers and the clusters they front. | Pooler |
| Operator | `/cnpg/operator` | Operator and plugin workloads, image catalogs, operator configuration. | ImageCatalog, ClusterImageCatalog |

Destination badges count **affected clusters**, not findings, and follow the namespace filter (the sidebar says so). The exact kinds stay under a collapsible **Resource kinds** block, grouped by API group; on workspace screens it starts collapsed.

Every CNPG kind's full detail is `/cnpg/<plural>/<namespace|_>/<name>` — reached from a row's **Open**, from the drawer's expand control, and by redirect from the generic `/workload/...` URL. The page keeps the workspace sidebar (its destination highlighted, the object nested under it) and uses Radar's detail view underneath: **Overview** is a composed summary, **Spec & status** is the kind's existing renderer, then YAML and the rest. A Cluster adds **Protection** (its recovery evidence), **Activity** (in place of Timeline) and merged instance **Logs**.

## Navigation

- **One drawer.** Rows inspect in the app's single drawer; `?drawer=kind:group:namespace:name` backs it, so refresh, share and Back restore it. Links inside the drawer append to that chain and show "← <previous object>" at the top of the drawer.
- **Return vs location.** A full detail shows "← <previous page>" only when it was reached by a drilldown (the label travels in history state); sidebar and global-nav hops are location changes and carry no return label. The crumb (`CloudNativePG / Protection / name`) always names the object's place, so a fresh tab has a parent without a fabricated previous task.
- **Context.** Detail URLs carry `ctx=<kube context>` (added on first view when absent). After a context switch the page says "<name> is not in <context>" with **Switch back** and **Go to …** — Radar never opens a same-named object from another cluster.
- **Namespace filter** narrows collections and counts. An explicitly opened object stays open, with a note when it is outside the filter.

## The certainty contract

Every value is something the cluster reports, labelled with where it came from. When the cluster does not report something the UI says so; it never shows zero, "none" or green in its place.

| Fact | Source | When it is not known |
|---|---|---|
| Instances, primary | `status.readyInstances`, `status.currentPrimary`, instance Pods (controller-owned by the Cluster's UID) | `–` |
| Replication | Pod readiness only | Always "lag unknown": readiness does not show whether a replica is streaming. Lag needs runtime data Radar does not read yet. |
| Schedule | ScheduledBackups targeting the Cluster (`spec.suspend` → suspended) | "No access to ScheduledBackups" when unreadable in that namespace |
| Destination | barman-cloud plugin `barmanObjectName`, in-tree `barmanObjectStore`, or volume snapshots | "No destination configured" |
| Last successful backup | Newest of: completed Backup CRs (7-day window plus the newest per cluster), ObjectStore `serverRecoveryWindow[...].lastSuccessfulBackupTime`, in-tree `status.lastSuccessfulBackup` (ignored for plugin clusters, where CNPG no longer sets it) — the winning source is shown | "None observed", or "No access to Backups" |
| WAL archiving | `ContinuousArchiving` condition | "Not reported" |
| Recovery window | Earliest point from ObjectStore `status.serverRecoveryWindow` for the cluster's server name. The latest point follows WAL archiving, not the last base backup, and no status reports it; it reads "not advancing" only while `ContinuousArchiving` is False | "Not reported" |
| Restore validation | A Cluster in the same namespace bootstrapped (`bootstrap.recovery`) from this cluster's store/server or one of its Backups, **with a ready instance** | "None recorded" (unknown tone) — Kubernetes records no restore tests, so this is never green. A matching cluster without a ready instance reads "Recovery declared in …". |
| ObjectStore upload health | **Inferred** from its user clusters' WAL archiving and recovery windows (ObjectStore has no status of its own) | "Unknown" |
| Declarations | `status.applied` (true / false / absent = pending); managed roles from `status.managedRolesStatus` (`reconciled`, `cannotReconcile`; anything else pending) | Pending, never failed |
| GitOps source | Argo CD / Flux labels and the Argo tracking annotation | "GitOps source not recorded" |
| Pooler pressure | — | "Not measured": needs PgBouncer metrics |
| ScheduledBackup cron | Shown verbatim | CNPG's cron is six-field (seconds first) and is never translated |

Problems come from Radar's Issues engine (the same detections as `/issues`) plus the audit's `cnpgNoDeclarativeBackup`, worded "No declarative backup schedule" because that is all it proves. A cluster **needs attention** when it has an issue of warning or worse on itself, an instance Pod, or an object that references it.

## Access

All data comes from `GET /api/cnpg/workspace`, authorized **per kind**: namespaced kinds use a cluster-wide `list` or fall back per namespace; `ClusterImageCatalog` needs a cluster-scope `list`. Each kind reports coverage (`full`, `partial` with the namespaces read, `denied`, `syncing`, `error`, `notInstalled`). Issues and audit findings are withheld where the underlying kind is not covered — Pod evidence only reaches callers who can list Pods. Denied namespaces are named only when the caller supplied the namespace list. A partial or denied kind makes the screen show a coverage notice, and its facts read "No access" rather than none.

`GET /api/cnpg/operator` reads operator and plugin Deployments (label `app.kubernetes.io/name=cloudnative-pg`, plugin Services labelled `cnpg.io/pluginName`) and the operator's config references. It ignores the namespace view filter (the operator lives in its own namespace), returns ConfigMap data only with `get configmaps`, and never reads Secrets.

Cluster logs (`/api/cnpg/clusters/{ns}/{name}/logs`) need `get pods/log`; Activity (`.../activity`) drops events for kinds the caller cannot list. Deleted child objects stay attributed to their Cluster because Radar records the owning cluster on timeline events at ingestion; history recorded before that is marked incomplete.

## Not in this version

Runtime data (replication lag, sessions, locks, WAL and slots via the instance manager or Prometheus), Pooler pressure, and operations (Backup now, Switchover, Restart, Hibernate, Restore).

## API

Every route is gated on the caller's own access, and a partial answer names what it withheld rather than shrinking silently.

- Workspace: `/api/cnpg/workspace` returns every CNPG kind (plus owner-validated instance Pods) with per-kind `coverage` (`full|partial|denied|notInstalled|syncing|error`, `partial` naming only in-scope denied namespaces), CNPG issues from the Issues engine and `cnpgNoDeclarativeBackup` audit findings, each withheld where the caller lacks coverage. Namespaced kinds follow the view filter and the capacity per-namespace `list` fallback; `ClusterImageCatalog` needs a cluster-scope `list`. Handler `internal/server/cnpg_workspace.go`
- Operator: `/api/cnpg/operator` returns the operator Deployments (`app.kubernetes.io/name=cloudnative-pg`) and plugin Deployments (served by Services labelled `cnpg.io/pluginName`) with image-tag version and readiness (`null` when unreported), plus the operator's ConfigMap/Secret/monitoring-queries references from its args and env. ConfigMap data only with `get configmaps`; the Secret is name-only, never read. Deployments and Services carry the workspace `coverage` states; deliberately ignores the namespace view filter (the operator lives in its own namespace). Handler `internal/server/cnpg_operator.go`
- Catalog reverse-lookup: `/api/cnpg/imagecatalogs/{ns}/{name}/clusters` and `/api/cnpg/clusterimagecatalogs/{name}/clusters` return the Clusters pinned to an image catalog, with the major each asks for and the image it actually resolved. Cluster-scoped catalogs are referenceable from any namespace, so the cluster-scoped route reads cluster-wide gated on `list clusters` — a view-filtered answer would report "nothing uses this" before an edit.
- Cluster logs: `/api/cnpg/clusters/{ns}/{name}/logs` (bounded snapshot) and `/logs/stream` (SSE, re-resolves instances every 5s) merge every instance Pod — label `cnpg.io/cluster` AND controller ownerRef to the Cluster's UID, never the label alone. Gated on `get clusters` + `list pods` + `get pods/log` before the Cluster lookup (404 after). `container` defaults to `postgres`, `tailLines` 200, `sinceTime` is converted to seconds and trimmed, `pod` must be a validated instance (400). Entries keep raw `content` and add `level`/`logger`/`message` parsed from the instance manager's JSON (`record.error_severity` wins over `level`).
- Cluster activity: `/api/cnpg/clusters/{ns}/{name}/activity?since=&limit=` reads the timeline store for the Cluster, its instance Pods (by owner) and CNPG children attributed by the retained `cnpg.io/cluster` label — `pkg/timeline.ExtractLabels` records it from the label or `spec.cluster.name` on CNPG-group objects and Pods, so deleted children stay attributed. K8s Event rows join by subject UID. Rows of a kind the caller can't `list` in the namespace are dropped; `oldest` is the namespace's retention floor and `attributionSince` the earliest labelled row — history before it cannot attribute deleted children.

## Testing

`make cnpg-demo` (read `scripts/cnpg-demo/README.md` first) produces WAL archiving failure, failed and unrecognised-phase Backups, failing declarations, a Pooler, both catalog kinds and an ObjectStore with a failing server — every state the workspace distinguishes, except successful restores.
2 changes: 2 additions & 0 deletions docs/integrations.md
Original file line number Diff line number Diff line change
Expand Up @@ -851,6 +851,8 @@ The source contract is Strimzi's [KafkaConnector status schema](https://strimzi.

[CloudNativePG](https://cloudnative-pg.io/) (CNPG) is the Kubernetes operator for PostgreSQL, covering the full lifecycle from bootstrapping to monitoring, with high availability, automated failover, and backup management.

Beyond the per-kind views below, the CloudNativePG **workspace** (`/cnpg`) composes them into fleet, protection, declaration, pooling and operator screens — see [cnpg.md](cnpg.md).

### What Radar Shows

**Cluster Detail View:**
Expand Down
Loading
Loading