CloudNativePG: task-shaped cluster pages, operations, live diagnosis and repair paths - #1922
nadaverell wants to merge 245 commits into
Conversation
…e guard
Operations (parity with the Freelens and Headlamp CNPG integrations): back up
now, switchover, restart cluster or one instance, reload, fence and lift,
hibernate and resume, schedule suspend/resume/run-now, and restore to a new
cluster through the existing create flow in strict create mode.
- GET .../capabilities answers per action and per instance from SubjectAccess
Reviews and the cluster's state; buttons are disabled with the reason.
- POST .../actions/{action} writes as the user and binds the confirmation to
the facts the dialog showed (cluster and target Pod UIDs, primaries, fencing,
hibernation, context): any change returns 409 instead of acting on a
different situation. Status transitions mirror kubectl-cnpg; Pod deletes use
UID preconditions; a Backup create that times out is resolved by reading the
same name back.
- ActionConfirmDialog (k8s-ui) leads with the effect, keeps the literal API
writes in an expandable section, and asks for the object's name before
disruptive actions.
Runtime: GET .../runtime reads each instance's /pg/status and exporter metrics
through the impersonated pods/proxy (fixed GET paths only, no redirects,
per-Pod/per-endpoint TLS, byte caps, a short identity-keyed memo). The Cluster
page gains a Runtime tab (replication topology with lag, sessions by state,
transactions, storage and WAL, slots, trends sampled while open) and Pooling
shows live PgBouncer pressure. Denied, unreachable and missing measurements
are reported as such, never as zero.
GitOps write guard, shared by every write dialog: POST /api/gitops/write-
evidence reads the target and its owner as the caller and reports, per field,
last-applied presence, owning field managers and ignore rules plus the owner's
sync policy; evaluateGitOpsWriteGuard classifies it and GitOpsWriteWarning
asks for acknowledgment when a sync may overwrite the change. SetImageDialog
and the Diagnose apply dialog now use it (ResourceActionsBar's
managedImageSources prop becomes imageOwnership).
…ard, and honest runtime totals - ScheduledBackup run now binds the schedule's generation; changed settings return 409 - Helm drift exemptions no longer read as safe: the next upgrade still overwrites - Transaction rates use the scrape time, deduplicating memoized samples - Pooler totals are lower bounds unless every pod and pool reported them - An unknown write outcome locks confirm instead of inviting a duplicate - Refused-for-changed-facts actions re-read capabilities - Restore dialog reads the Cluster from useResource correctly
PR Summary by QodoAdd CloudNativePG actions, live runtime, and a shared GitOps write guard
AI Description
Diagram
High-Level Assessment
Files changed (44)
|
Code Review by Qodo
1.
|
…s, scheme in runtime memo key
… lock chain and Prometheus
…erged logs, and call an idle pooler idle
…scheduled runs that produced no backup
GET /api/cnpg/clusters/{ns}/{name}/storage reports each instance's claims by
role (owner- and instance-validated), requested vs capacity, resize state,
StorageClass expansion, kubelet volume usage from Prometheus and the WAL facts
(size, segments, archive backlog, slot retention) through the Runtime tab's
memoized pods/proxy reads, each source with its own coverage. GET
/api/cnpg/disk returns the fullest measured volume per visible cluster, one
claim list and one usage batch per namespace.
The Issues engine raises CNPGScheduledRunNoBackup on a Cluster when an active
ScheduledBackup fired after its newest successful backup and no run since then
succeeded or is still running, using the operator's six-field cron parser.
…e and Pooler pause/resume endpoints Sessions run fixed SQL (pg_stat_activity + pg_blocking_pids) with psql over the caller's pods/exec; cancel/terminate revalidate pod UID, pid and backend_start in the same statement and pass the user's values only as psql variables. Destroy instance mirrors kubectl cnpg destroy (PVCs detached or deleted first, then the Pod, then the instance's Jobs), binds the reviewed Pod and PVC UIDs and refuses the primary. Poolers gain capabilities with Deployment/Service readiness, pause/resume bound to the reviewed paused value, and per-Pod observed PgBouncer state via SHOW STATE. Pooler runtime reports the pool mode PgBouncer uses.
…tion and DatabaseRole kind
GET /api/cnpg/clusters/{ns}/{name}/ha reads FailoverQuorum (with R + W > N),
PodDisruptionBudgets, primary and operator Leases, per-instance node, zone,
QoS and image drift, cluster Jobs, -rw endpoints and certificate renewal
ownership, each sub-read authorized on its own. The runtime status adds
pendingRestartForDecrease, pg_rewind, instance-manager version and a role
detail derived from the instance's own report, plus the postmaster start time.
Certificates close to expiry become issues. setMaintenance/unsetMaintenance
mirror kubectl cnpg maintenance. DatabaseRole joins the workspace kinds.
…nd disk use in the fleet The Runtime tab's Storage & WAL section becomes the storage decision: each instance's claims with used/capacity from kubelet volume stats (or why it is unknown), resize state, whether the StorageClass allows expansion, and the WAL on disk, archive backlog and slot retention side by side. Expansion names the spec field and opens the apply flow with only that size, behind the GitOps write guard. Volumes stay visible without pods/proxy. The fleet gains a Disk column and the Cluster overview a Storage fact from /api/cnpg/disk; a measured volume at 80% or more puts the cluster in Needs attention. The renderer's configured sizes read as requested sizes and the Runtime database sizes as logical sizes.
…ce and Pooler detail to the CNPG workspace psql opens the dock terminal in an instance's postgres container (?shell=psql) from instance rows and the cluster menu, labelled by the instance's current role. The Runtime Sessions section adds the blocker-to-victim tree with connection headroom and per-instance CPU/memory, and cancel/terminate dialogs bound to pod, pid and backend_start. Standby rows gain Destroy with the delete / keep-volumes choice. The Pooler summary shows Deployment readiness, per-pool live pressure, limits with verified PgBouncer defaults, requested vs observed pause and the Service path; the Pooling table reads readiness from the Deployment instead of the scheduled count. Pause/resume is a header action.
…CNPG demo, and document storage Schedule fire times were computed in Radar's local zone; the operator's clock is UTC. The runtime demo's Prometheus now scrapes each kubelet through the apiserver node proxy (kubelet_volume_stats_* only); kind's hostPath volumes publish no stats, which the README states and the UI shows as no usage metrics. Claims kept after an instance was destroyed are named as such, a standby's archive fact says the primary archives, and a volume sharing a larger filesystem says the measurement is that filesystem's.
…ow readiness source once Pooler readiness detail names its source instead of repeating the count, and the Pooler runtime's reported pool mode is pinned by a test.
# Conflicts: # web/src/components/cnpg/CNPGClusterRuntime.tsx
…he header - Overview keeps State and Protection open; HA and instances, and Certificates, fold to one summary line each and open themselves when something in them is out of line. - Connect is a header button opening the connection facts in a dialog (?connect=ns/name), which the restore Next steps link opens too. - The Replication chip reads no calmer than a sustained-lag finding, and the fleet says Pods ready rather than ready. - WAL archiving shows since when it has worked, when that changed within the last day and well after creation. - The pooler summary breaks PgBouncer pressure down per Pod. - Switching the namespace picker to All refetches what is on screen once the server holds the new pick, so no request reads the old one.
…alog - The HA summary names the facts it could not read instead of reading calm; the Certificates summary says nearest reported expiry, counts unreadable ones and opens itself for them. - Connect's URL value names the surface, so a drawer showing the page's own Cluster does not open a second dialog, and opening or closing it keeps the page's return label. - Switching to All namespaces cancels requests still in flight before refetching, so a first fetch that read the old pick is replaced.
A link following from the dialog and the dialog closing would both write the URL in the same tick and undo each other, leaving the dialog over the drawer it opened. The param now only asks for the dialog: the button opens it and drops the param. The dialog is a step wider so endpoint descriptions stay on one line.
A full-screen drawer can report the same surface as the page beneath it, so both would open a dialog for the same request.
…ton answer its request A Cluster with no instance Pods now opens HA with "no instance Pods" rather than "0/0 instances ready", and "images match" needs every instance to say so. The Connect request is held by the button that answered it only while that button is mounted, so a remount before the param is dropped still opens the dialog. Pins five-field schedules to the parser's next runs.
…workspace (#1969) ## Summary The CloudNativePG workspace (#1921, #1922) carried a generic layer under CNPG names: the facts and problem list on every summary, per-kind coverage, the reviewed-action contract, grants, Prometheus series attribution, and screen layout borrowed from Capacity's internals. This PR moves that layer into shared, integration-neutral modules, so the next workspace-style integration (Velero is the likely one) imports it instead of copying CNPG. Nothing here is released yet, so the three CNPG wire shapes that every future workspace would copy are fixed now: - grants are structured objects, not sentences; - "not cached by Radar" is told apart from "no access"; - each object's GitOps manager comes from the server. It also adds the version-skew gate the CNPG endpoints were missing. Only mechanisms whose interface is already evident in today's code are shared: each has two or more consumers, or is a contract every workspace must follow. The rest is listed below as not shared yet. ## What changed ### Shared UI (`@skyhook-io/k8s-ui`) - **`components/facts/`**, for any surface that shows observed values (single-kind renderers included): - `Fact`, with `FactGrid`/`FactRow`/`FactValue`/`FactSource` - `CertaintyGlyph`, moved from Capacity; `CapacityCertainty` stays as the same type - `ManagedByText` - **`components/problems/`**: `WorkspaceProblem<Category>` and `ProblemCallout`/`ProblemList`/`ProblemMeta`, which take the workspace's `rootKind` instead of a hard-coded `'Cluster'`, plus `OpenIssueContext`. - **`ui/FoldSection`**: `SectionHeading`, `FoldSection`, `FoldSummary`. - **Elsewhere in k8s-ui**: - `ui/RefLink` now uses the existing `ResourceRef`. - `toneTextClass` and `worseTone` live in `ui/status-tone`. - `utils/grant` adds `Grant`, `formatGrant` and `grantParts`. - `issueReasonTitle` gives a reason one title on both the Issues page and the workspace. - **Kept CNPG-specific:** categories, ordering, wording and builders. CNPG binds `WorkspaceProblem` to its own categories. - **Badges:** CNPG badges now use `healthToSeverity` like the rest of the app. - **Buttons:** secondary buttons use a new `.btn-secondary` class (documented in DESIGN.md) instead of ten hand-rolled class strings. - **Sidebar:** `SidebarCategoryDestination.countLowerBound` renders a count over partly read data as `≥N`, and its zero as unknown rather than none. ### App (`web/`) - **`components/workspace/`** holds the screen primitives Capacity and CNPG share: `ScreenBody`, `ScreenEmptyState`, `Notice`, `Segments`, `FilterChips`, `SectionTable` and its table classes, `RefreshFailedNotice`, `GrantText`, and the text helpers. CNPG no longer imports from `capacity/shared.tsx`. - **`api/actions.ts`** is the client half of the action contract: - the `ActionCapability` and `ActionRequest` types - `actionErrorCode`, which an integration widens with its own codes - `actionOutcomeLocked` and `actionCompleted` - `capabilityReason`, now the one "why is this disabled" wording - **Utilities:** `utils/drawer-trail.ts` holds the `?drawer=` codec, and `utils/page-links.ts` the back label and the subject-filtered Issues link. - **Kind table:** CNPG's detail kinds derive kind and group from the workspace kind table. ### Server - **`internal/server/actions.go`**: the reviewed-action contract. - `ActionCapability`, `ActionRequest`, `decodeActionRequest` and `decodeActionParams` - refusals: `changedAction`, `blockedAction`, `partialAction` and `writeActionError` - the permission decision: `grantPermission` and `capabilityVerdict`, with a real SelfSubjectAccessReview in local mode - `mergePatchAtVersion` - **Grants:** `Grant` lives in `internal/auth` and carries `In(ns)` and `String()`. - **Kind access:** `kind_access.go` covers per-kind access and coverage (`readWorkspaceKind`, `typedKindScope`, `KindCoverage`); `cache_scope.go` has `namespacesWithinCache`. - **Other shared helpers:** - `read_source.go`: one `ReadSource` shape for HA, recovery and storage - `fanout.go`: the bounded per-namespace fan-out - `internal/prometheus/series_scope.go`: `SeriesIsolation`, already used by the PVC usage charts - **Kept CNPG-specific:** facts, guards, runners and the extra refusal codes. - **Exported manager detection:** `pkg/topology` now exports `ManagedByFromMeta`. This is additive. ### Wire changes (CNPG endpoints, unreleased) - **`grant`** is `{verb, group?, resource, subresource?, namespace?}` (no namespace means cluster-wide) on every capability, read source, chart and report item. The wording is unchanged and comes from `Grant.String()` / `formatGrant`. - **Kind coverage** gains: - a state `uncached`, for a scope Radar's cache does not cover at all; - `uncachedNamespaces`, named under the same rule as `deniedNamespaces`. Before, a namespace the caller can read but Radar does not cache was reported as denied, and the UI said "No access". It now says "Radar does not cache Pods in db". - **`/api/cnpg/workspace`** gains `managedBy`, keyed `Kind/ns/name`. - "Declared in" now uses the server's manager detection and links to the Argo CD application or Flux object. - Before, the client parsed labels itself. That misnamed Argo CD applications living outside Argo CD's namespace (`ns_app`) and ignored the Flux namespace label. ### Version skew New `FeatureCapabilities` flags: `cnpgWorkspace` (every `/api/cnpg/*` route except the two image-catalog lookups that predate it) and `gitopsWriteEvidence`. Each has a `radarFeatures` entry. Every CNPG query, mutation, log stream, report download and the write-evidence query goes through `useRadarFeature`. Mutations are never retried. On a Radar without the workspace: - the sidebar offers no workspace destinations; - `/workload` links and drawer Expand stay on the standard views; - a Cluster's Logs tab shows its Pods' logs; - a workspace screen opened directly says it needs a newer Radar. ### Vocabulary On screen these areas are named by their subject. The block above a sidebar category's kinds reads **Views**, and the copy says "CloudNativePG views" and "CloudNativePG data", never "workspace", because Radar Hub uses "workspace" for the customer's account. "Workspace" remains the internal term for an integration with several kinds and its own screens (the integration guide, `SidebarCategoryWorkspace`, `/api/cnpg/workspace`). ### Docs - **DESIGN.md:** the rules for unknown, partial and denied values are written once. `capacity.md` and `cnpg.md` link to them and keep their per-value tables. - **`docs/INTEGRATION_GUIDE.md`:** a new "Workspace integrations" chapter covering when a workspace is warranted, the pieces to import, and what is not shared yet. CLAUDE.md points to it. ## Not shared yet These have one consumer, and their interface should come from the second: - the workspace registry / App wiring - the operation tracker - the generic action runner - the fixed-path `pods/proxy` reader - the merged log stream - the report bundle - operator diagnosis - a TTL-memo helper Rollouts' capabilities answer "allowed" in local mode without asking the apiserver. That is a real bug, but Rollouts returns bare booleans and hides denied actions, so it needs its own PR. ## Public surface - **k8s-ui:** exports that exist on main are unchanged. Renamed exports were all added in #1921/#1922. `CapacityCertainty` is kept. - **`pkg/`:** one additive export (`topology.ManagedByFromMeta`). - **`pkg/capacityapi` v1alpha1:** unchanged. - **radar-app:** surface unchanged. ## Testing - **Go:** `go build ./...`; `go test ./...` in both modules. The `cmd/desktop` env test is flaky and passes alone. `gofmt -l` is clean. - **Frontend:** - type checks pass for web and k8s-ui; - vitest: k8s-ui 236 files / 4,369 tests, web 182 files / 2,028 tests; - eslint: 0 errors, and no new warnings. - **Live on the kind CloudNativePG demo, full access and a view-only identity:** - every CNPG page rendered with structured grants and no "[object Object]" or "undefined"; - a namespaced Argo CD tracking ID on a demo Cluster showed "Declared in: Argo CD application argocd/payments", linked; - with the capability flags removed from `/api/capabilities` (an older Radar behind Radar Hub), the sidebar workspace, redirects, composed summaries and actions fell back to the standard views, and `/cnpg` showed the upgrade note. Stacked on #1922 (which is stacked on #1921). Merge order: #1921, #1922, then this PR, before a release ships `/api/cnpg/*`. <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Touches CNPG API shapes, RBAC/coverage semantics, and cluster write paths via the shared action layer; behavior is intended to be equivalent with clearer grant and cache messaging. > > **Overview** > Extracts the generic workspace layer that CloudNativePG had under CNPG-specific names into shared modules, so future integrations (e.g. Capacity-style screens) import one contract instead of copying. > > **Server:** Adds `internal/auth.Grant` (structured RBAC on the wire), `internal/server/actions.go` (reviewed `ActionRequest`, capabilities, 409/partial errors, `mergePatchAtVersion`), `cache_scope.go` / `kind_access` patterns, and `prometheus/series_scope.go` (shared Prometheus isolation). CNPG handlers now use these; history/PVC denial fields expose `*Grant` instead of strings. `FeatureCapabilities` adds `cnpgWorkspace` and `gitopsWriteEvidence`. CNPG workspace coverage gains `uncached` / `uncachedNamespaces` (distinct from denied) and `managedBy` from server-side GitOps detection. > > **UI:** New k8s-ui `facts/`, `problems/`, fold sections, `formatGrant`, and app `components/workspace/` plus `api/actions.ts`. CNPG is rewired to structured grants and shared action types; secondary buttons use `.btn-secondary`. > > **Docs:** DESIGN.md documents unknown/partial/denied values once; INTEGRATION_GUIDE adds a workspace-integrations chapter; CLAUDE.md points builders at it. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit 5af6ed1. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->
…ture/cnpg-actions-runtime # Conflicts: # web/src/components/cnpg/CNPGClusterLogs.tsx
A Backup list that failed or is not watched was treated as no backups, so a cluster that still had backups could get "no successful backup since a scheduled run". The detector now evaluates a cluster only when its Backups were read, and, for a cluster backing up to an ObjectStore, that store's list too.
# Conflicts: # packages/k8s-ui/src/components/cnpg/index.ts
…operator and plugin Pod restarts Fleet metrics: a standby whose WAL receiver is down reads replication lag 0, so the lag reading in state ok now carries standbys, receiving and receiverDown from cnpg_pg_replication_is_wal_receiver_up (receiverUnknown when no standby reports it), and each cluster gets a slots reading listing inactive physical replication slots with the reporting instance's role and retained WAL. Both are one query per namespace under the same get pods gate, selector and isolation as lag. Operator: each operator and plugin component lists its Deployment's Pods (ready, restarts, current container start, last termination) with podCoverage gated on list pods; the diagnosis Pods carry lastTermination too.
…surface The Cluster page is organized by task: Overview, Replication, Storage, Performance, Backups, Activity, Logs, Configuration, YAML. Replication joins each instance's Pod readiness with its PostgreSQL role, timeline, streaming state and the HA slot the primary keeps for it, plus the HA configuration; Performance holds Sessions, Database health and History (chart groups); Backups shows the cluster's recovery evidence as facts with every backup run and each schedule's last outcome; Configuration adds connections and certificates above the declared settings. The fleet row, the header chips (now on every tab), the Overview and the switchover dialog read one assessment. A standby that receives nothing is a problem from Prometheus (WAL receiver down) or the instance managers (no pg_stat_replication row), never "lag 0 s"; an inactive slot on the primary holding 1 GiB or more is a problem naming the standby it serves. Switchover candidates show what the live read sees wrong with them. The fleet views are named Clusters and Backups; the fleet table shows Replication, Storage and Backups beside the cluster's readiness. The Operator screen leads with its current state and shows restart history. Escape that closes a menu no longer also leaves the page.
…hat they disprove A standby's WAL receiver must be down in every Prometheus sample for 5 minutes before the fleet raises it, so a restarting standby is shown but not alarmed. A live read replaces the fleet's answer for what it covers: a standby the primary streams to again, or a slot no longer inactive, clears the earlier problem, and slots are judged from the primary's slot list even when its replication rows are unusable. "None receiving" needs every expected standby accounted for. Switchover never preselects a standby with concerns. The Operator's current state only claims what it read and lists what it could not. Backup runs show the newest 10 with a way to see all, and a schedule's last run never borrows an older Backup's outcome. The Overview's links name the tab they open, and the Configuration tab leads with connections instead of repeating the Overview's problems.
The tab strip drops its icons and tightens spacing only while every tab does not fit, measured against the space it has, so a page with many tabs (a CloudNativePG Cluster's nine) stays on one line at 1280 px and pages that fit keep their look. Header badges no longer break mid-word. The Operator's current state names webhook configurations it could not read, and a Cluster's Backups tab leaves out the Cluster column.
…ups tab A live read clears a fleet slot warning only when it shows the slot active, measured below the threshold, or gone from a complete list; an unmeasured slot or a capped list leaves it standing. Sustained receiver loss needs the instance to have been a standby in every sample, so a former primary just after a switchover never qualifies. The Operator confirms its webhooks only when every configuration was read. A Cluster's Backups tab offers restore to a new cluster, and Storage names the standby each HA slot is kept for.
…ilures A Cluster's Backups tab leads with a WAL archiving repair panel while archiving fails: the operator's reason, the last archived and failed WAL, where the destination is declared, the credential Secrets it references (named, never read), where the archiver logs, then whether archiving resumed and a base backup completed after it. Storage explains WAL an inactive slot holds, names the standby, and how it is freed; a class that cannot expand points to restoring into a larger cluster. A controller phase that is not healthy links to the Operator view, which lists clusters stuck on a plugin beside that plugin's readiness and restart history. The fleet's lag says how many standbys it covers.
…he archiver's logs, and waits for a backup started after resume The repair panel reads the ContinuousArchiving message instead of claiming there is none, opens logs on the archiving container (plugin-barman-cloud when the plugin is the WAL archiver), links to the Operator when the Cluster's phase says a plugin blocks reconciliation, shows a workload identity and the endpoint CA together, and counts only a completed Backup that started after archiving resumed. The fleet's lag coverage counts the standbys whose lag was read, and expects every instance of a replica cluster.
…ent that just restarted
…r trusts only a backup a restore can use A blocked phase no longer claims the operator stopped and needs manual intervention: it retries every one, the plugin phases clear once the plugin loads and answers, and the banner and Overview quote status.phaseReason. Unrecoverable keeps upstream's own wording. The archiving checklist counts a Backup only when a restore can start from it with this cluster's WAL archive (same archiver, or a volume snapshot), and reads resumption from the ContinuousArchiving condition when the instance manager's stats restarted with the primary. A standby that is not streaming can still replay from the archive; copy now says the slot does not advance without streaming instead of claiming the standby cannot catch up.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit c5038c0. Configure here.
…at failed to archive A Backup that started after archiving resumed can still begin earlier in the WAL when it was taken from a lagging standby (CloudNativePG's default target). The checklist now compares the Backup's own beginWal with the last failed WAL, says when a backup begins before it, and stays unverified when the failed WAL cannot be read.
…t diagnosed as beginning before the failure
…stance manager recorded The ContinuousArchiving condition turning True also happens when archiving is first set up, so on its own it no longer opens the resumed panel.

Stacked on #1921 (CloudNativePG workspace). Merge that first; this PR's base is
feature/cnpg-workspace.Summary
This PR turns the CloudNativePG workspace from a read-only view into something on-call can run the common incidents from, largely without dropping to
kubectl cnpgor Grafana. It covers the recurring CNPG incidents:For each one, Radar shows the evidence, offers the intervention, and follows it until it is observed complete. It never shows an unknown value as zero or healthy, and it clears a warning only on evidence that disproves it.
Actions cover everything the Freelens and Headlamp CNPG plugins offer, except guided creation forms (deferred). From
kubectl cnpgthey cover backup, promote, restart, reload, fencing, hibernate, maintenance, destroy and psql. They do not coverpgbench,fio,certificate,publication/subscription create,report operatororinstall generate. Every write is made with the caller's identity, is bound to the facts they reviewed (409 if those changed), and passes a shared GitOps write guard.What changed
Information architecture
Fleet views (Resources sidebar → CloudNativePG → Views): Clusters (the fleet, defaulting to Needs attention), Backups (recovery evidence, runs, schedules, destinations), Declarations, Pooling, Operator. Badges count affected clusters, not findings.
A Cluster's page is organized by task, each tab the one home of its facts, the others linking to it:
Serving · Replication · Storage · Backups chips sit under the title on every tab and open the tab that explains them. The chips, the Overview, every tab and the fleet row read one assessment (
useCNPGClusterAssessment): the fleet's evidence (cached objects, issues, Prometheus) enriched with what the instance managers report live, so no surface reads calmer than another.WorkloadView(k8s-ui) gains additivetabOrder,specTab(relabel and lead content for the spec tab) andsubheaderprops;DetailShelldrops tab icons only when the tabs would otherwise overflow.Repair paths
plugin-barman-cloudwhen the plugin is the WAL archiver, otherwisepostgres); a link to the Operator when the Cluster's phase says a plugin blocks reconciliation. Then a checklist: archiving resumed (after a failure the primary's instance manager recorded), and a completed Backup that a restore can start from with this cluster's WAL archive (same archiver, or a volume snapshot) and whose ownbeginWalcomes after the last WAL that failed to archive — a backup from a lagging standby can start after the resume yet begin before the failure.pg_stat_replication. Switchover never preselects a candidate with concerns, and shows each candidate's.status.phaseReason; only "unrecoverable" keeps upstream's "needs manual intervention". The Overview links a plugin phase to the Operator, whose Current state names those clusters beside the plugin Deployment's readiness and restart history.Actions and operations (
internal/server/cnpg_actions*.go,cnpg_destroy.go,cnpg_pooler_actions.go,web/src/components/cnpg/actions/)Cluster.IsReplica().SHOW STATE.backend_startinside the same statement.ActionConfirmDialog(k8s-ui, additive): effect first, literal API writes in an expandable section, typed confirmation for disruptive actions; an unknown or partial outcome locks confirm instead of inviting a duplicate.-rwand the old primary streams again). Missing telemetry reads "unobservable", never "stalled".Problem provenance and the Issues link
/api/issues/resource/...) requiresgeton the subject's kind, withholds grouped issues and members the caller can't read, and with?coverage=1reports whether the kind is watched and how many issues were withheld./api/issueslist, the MCPissuestool, and evidence inside grouped issues still filter by namespace only. The link is hidden under Radar Hub until Hub forwards the subject params.Restore (
internal/server/cnpg_recovery.go,web/src/components/cnpg/recovery/)create clustersin the target namespace and names the missing grant before the review step.Live reads, fleet evidence and history (
cnpg_runtime.go,cnpg_sessions.go,cnpg_history.go,internal/prometheus/cnpg_history.go)/pg/statusand the exporters through the caller'spods/proxy— fixed GET paths, no redirects, bounded, memoized per identity.pg_stat_replicationrow for reads "not connected to the primary"; its own backlog is shown only when it is on the primary's timeline.pods/exec; connection headroom.Storage (
cnpg_storage.go,CNPGStorage.tsx)HA, certificates, declarations (
cnpg_cluster_ha.go,internal/issues/source_cnpg_*.go)Operator (
cnpg_operator*.go,/api/cnpg/operator/status)Failwebhook with no ready endpoint, restart history per component (with the last termination), what it confirmed and what it could not read. Then leader, watched namespaces, webhooks and reconcile errors from the operator's metrics.Parity with Freelens and Headlamp (excluding guided creation forms)
-rw/-ro/-r/ Pooler hosts with honest labels, database, owner and the app Secret name, never read); schedules in plain language with next runs from the operator's own cron parser; logical replication Subscription → Publication → publisher slot and whether it survives failover; replica clone progress; per-database health; checkpoints; in-page trends without Prometheus (rates per exporter run and per primary, so a switchover is a gap, never a spike).Shared pieces
POST /api/gitops/write-evidence+evaluateGitOpsWriteGuard/GitOpsWriteWarningin k8s-ui; evidence classes last-applied, GitOps field manager, ignore rules, owner sync policy; used by the CNPG actions, Set image and the investigation apply dialog.k8score.CanIsends subresources in their own field, so webhook authorizers such as GKE IAM answerpods/execandpods/proxycorrectly.initialContainerpreselects a container.ManagedImageSource/SetImageDialog.managedSources/ResourceActionsBar.managedImageSources→imageOwnership. Radar Hub does not use them.make cnpg-demo-runtimeadds a live namespace with real plugin backups and a restore, a pooler, pgbench load, a blocked lock chain, and a Prometheus that also scrapes kubelet volume stats.All k8s-ui changes are additive props or new exports. Radar Hub is unaffected until it opts in.
Testing
go build ./...,go test ./...,make tsc, k8s-uitsc, vitest (k8s-ui 4,399, web 2,053) and web lint (0 errors). CI's workflow runs only on PRs intomain, so this PR shows Bugbot only until it is retargeted.view+ CNPG read: nopods/proxy, no exec, no writes): the fleet; each Cluster tab at 1920 and 1280; switchover candidates with concerns; a standby stuck on an old timeline with its inactive slot holding ~5 GiB (Storage relief → Replication); a WAL-failing cluster's repair panel, its link landing on theplugin-barman-cloudlogs that show the unreachable endpoint; the Operator naming plugin-blocked clusters. The same pass under real degradation — operator and plugin crash-looping on startup-probe timeouts, primaries flapping — read as stale, unassessed or unknown with the reason, never as healthy. Restricted mode names each missing grant and shows nothing as zero.Notes / follow-ups
Note
High Risk
Adds many impersonated write and pods/proxy/exec paths for PostgreSQL operations (switchover, destroy, session terminate, restore) plus GitOps write-evidence; mis-gating or fact-binding bugs could allow unsafe changes or misleading health signals.
Overview
CloudNativePG grows from a read-only fleet workspace into task-shaped cluster pages (Replication, Storage, Performance, Backups, etc.) backed by many new
/api/cnpg/*routes: live instance/pooler runtime via callerpods/proxy, storage and HA facts, Prometheus history with strict cluster-identity scoping, sessions/blocking and destroy flows viapods/exec, capabilities + POST actions bound to reviewed facts (409 on drift), restore/recovery, report bundles, and an operation tracker on the UI. Fleet evidence and problems are tightened so unknown/partial data is never shown as zero or healthy, with shared facts/problems/workspace components and documented certainty rules in DESIGN.md.Cross-cutting additions include
Grant/ReadSourcefor permission wording,POST /api/gitops/write-evidenceand client GitOps write guard wiring,cnpgWorkspace(and related) feature flags, Issues for certificate expiry and missed scheduled backups, DatabaseRole support, timeline Pod health on change, Prometheus PVC usage batching and clearer unreachable discovery errors, Radar Cloud Lease read for viewers plus expanded CNPG integration-read grants, andmake cnpg-demo-runtimefor live backup/restore/load/lock fixtures. Docs and INTEGRATION_GUIDE add a full workspace integrations checklist.Reviewed by Cursor Bugbot for commit 349ccfb. Bugbot is set up for automated code reviews on this repo. Configure here.