Skip to content

CloudNativePG: task-shaped cluster pages, operations, live diagnosis and repair paths - #1922

Open
nadaverell wants to merge 245 commits into
feature/cnpg-workspacefrom
feature/cnpg-actions-runtime
Open

nadaverell wants to merge 245 commits into
feature/cnpg-workspacefrom
feature/cnpg-actions-runtime

Conversation

@nadaverell

@nadaverell nadaverell commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #1921 (CloudNativePG workspace). Merge that first; this PR's base is feature/cnpg-workspace.

Summary

This PR turns the CloudNativePG workspace from a read-only view into something on-call can run the common incidents from, largely without dropping to kubectl cnpg or Grafana. It covers the recurring CNPG incidents:

  • WAL archiving failures that fill the disk
  • full volumes and WAL held by an inactive replication slot
  • replicas that won't rejoin
  • stuck switchovers
  • restore/PITR
  • overdue backups
  • blocked queries
  • pooler saturation
  • certificate expiry
  • an operator, or a CNPG-I plugin, that has stopped reconciling

For each one, Radar shows the evidence, offers the intervention, and follows it until it is observed complete. It never shows an unknown value as zero or healthy, and it clears a warning only on evidence that disproves it.

Actions cover everything the Freelens and Headlamp CNPG plugins offer, except guided creation forms (deferred). From kubectl cnpg they cover backup, promote, restart, reload, fencing, hibernate, maintenance, destroy and psql. They do not cover pgbench, fio, certificate, publication/subscription create, report operator or install generate. Every write is made with the caller's identity, is bound to the facts they reviewed (409 if those changed), and passes a shared GitOps write guard.

What changed

Information architecture

Fleet views (Resources sidebar → CloudNativePG → Views): Clusters (the fleet, defaulting to Needs attention), Backups (recovery evidence, runs, schedules, destinations), Declarations, Pooling, Operator. Badges count affected clusters, not findings.

A Cluster's page is organized by task, each tab the one home of its facts, the others linking to it:

Tab Job
Overview What needs attention and where to act: problems, then one line per health dimension linking to its tab, instances, controller phase
Replication Diagnose instances and assess a primary change: per instance Pod readiness, node and QoS beside role, timeline and streaming state; per standby its backlog and the HA slot the primary keeps for it; other slots; HA configuration. Instance actions live here
Storage What occupies space, why, and resize
Performance Sessions (headroom, blocking tree, cancel/terminate) · Database health · History (every chart, filterable by group; Replication and Storage link to their group)
Backups Restore, this cluster's recovery evidence, runs, schedules, destination, restore validation
Activity, Logs Events and changes; merged instance logs
Configuration Connections and certificates above the declared spec

Serving · Replication · Storage · Backups chips sit under the title on every tab and open the tab that explains them. The chips, the Overview, every tab and the fleet row read one assessment (useCNPGClusterAssessment): the fleet's evidence (cached objects, issues, Prometheus) enriched with what the instance managers report live, so no surface reads calmer than another.

WorkloadView (k8s-ui) gains additive tabOrder, specTab (relabel and lead content for the spec tab) and subheader props; DetailShell drops tab icons only when the tabs would otherwise overflow.

Repair paths

  • WAL archiving failing (Backups tab, while failing and for a day after it resumes): the operator's ContinuousArchiving message; the last archived and failed WAL from the primary's instance manager; the destination and where it is declared; the credential Secrets and endpoint CA it references (named, never read) or the workload identity it uses; a link that opens Logs on the archiving container (plugin-barman-cloud when the plugin is the WAL archiver, otherwise postgres); a link to the Operator when the Cluster's phase says a plugin blocks reconciliation. Then a checklist: archiving resumed (after a failure the primary's instance manager recorded), and a completed Backup that a restore can start from with this cluster's WAL archive (same archiver, or a volume snapshot) and whose own beginWal comes after the last WAL that failed to archive — a backup from a lagging standby can start after the resume yet begin before the failure.
  • WAL held by an inactive slot (Storage tab, ≥ 1 GiB on the primary): the standby the HA slot is kept for, whether it receives WAL, and the two ways the WAL is freed (get the standby streaming, or destroy it so the operator recreates it and drops its slot), linking to that standby on Replication. For a StorageClass that cannot expand, resize points to restoring into a larger cluster.
  • A standby not receiving WAL: raised in the fleet only when its receiver was down in every sample for 5 minutes while it was a standby throughout; on the Cluster page from the primary's own pg_stat_replication. Switchover never preselects a candidate with concerns, and shows each candidate's.
  • Blocked reconciliation: a blocked phase says the operator retries it — the plugin phases clear once the plugin loads and answers — and quotes the operator's status.phaseReason; only "unrecoverable" keeps upstream's "needs manual intervention". The Overview links a plugin phase to the Operator, whose Current state names those clusters beside the plugin Deployment's readiness and restart history.

Actions and operations (internal/server/cnpg_actions*.go, cnpg_destroy.go, cnpg_pooler_actions.go, web/src/components/cnpg/actions/)

  • Back up now binds the cluster's backup target when the Backup inherits it, names that target in the dialog, and warns when WAL archiving is failing. Replica-cluster detection follows the operator's Cluster.IsReplica().
  • Cluster: back up now; switchover/promote; restart the cluster or one instance; reload; fence/unfence; hibernate/rehydrate; node maintenance set/unset; destroy instance (standbys only, fenced first because CNPG never promotes a fenced instance; keep or delete volumes; every grant checked before the first write; the fence lifted last); psql on any instance; a redacted report bundle.
  • ScheduledBackup: suspend, resume, run now, edit schedule (validated server-side with the operator's own cron parser).
  • Pooler: pause/resume, with the observed PgBouncer state read via SHOW STATE.
  • Backends: stop query / terminate from the blocking view; the server re-checks pid and backend_start inside the same statement.
  • Capabilities endpoints return the facts a confirmation binds to plus a per-action verdict naming the missing grant. With auth off, Radar asks the kubeconfig identity (SelfSubjectAccessReview) instead of assuming access. Actions that go through the operator's webhook are refused, with the reason, while the operator is not reconciling.
  • ActionConfirmDialog (k8s-ui, additive): effect first, literal API writes in an expandable section, typed confirmation for disruptive actions; an unknown or partial outcome locks confirm instead of inviting a duplicate.
  • Operation tracker: requested → observed → progressing → completed | failed | stalled | superseded | unobservable, with per-operation completion criteria (a switchover completes only when the new primary serves -rw and the old primary streams again). Missing telemetry reads "unobservable", never "stalled".

Problem provenance and the Issues link

  • Every CNPG problem names where its evidence comes from ("Reported by CNPG", "Backup status", "Pod readiness probe", "Radar check", "Measured by Prometheus"). "See in Issues →" opens Issues filtered to that object using the uncapped per-resource lookup.
  • That lookup (/api/issues/resource/...) requires get on the subject's kind, withholds grouped issues and members the caller can't read, and with ?coverage=1 reports whether the kind is watched and how many issues were withheld.
  • Known gap, fixed in a follow-up PR: the main /api/issues list, the MCP issues tool, and evidence inside grouped issues still filter by namespace only. The link is hidden under Radar Hub until Hub forwards the subject params.

Restore (internal/server/cnpg_recovery.go, web/src/components/cnpg/recovery/)

  • Restore starts from a Cluster (its Backups tab), a Backup or an ObjectStore. Every entry point checks create clusters in the target namespace and names the missing grant before the review step.
  • The point-in-time picker shows the recovery evidence and warns but never blocks. A preflight lists what is copied. The new cluster is followed through its recovery Pods to healthy.
  • A restore-validation note (an annotation) makes "Restore validation" read "recorded by … at …"; it is never green. A links-only Next steps list covers Connect, validation and setting up backups.

Live reads, fleet evidence and history (cnpg_runtime.go, cnpg_sessions.go, cnpg_history.go, internal/prometheus/cnpg_history.go)

  • Live reads: /pg/status and the exporters through the caller's pods/proxy — fixed GET paths, no redirects, bounded, memoized per identity.
  • Replication: role detail, pending restart, replay backlog in bytes. A standby the primary has no pg_stat_replication row for reads "not connected to the primary"; its own backlog is shown only when it is on the primary's timeline.
  • Fleet evidence from Prometheus: replay lag (sustained ≥ 30 s for 10 min is a warning, ≥ 5 min critical), WAL receiver state (a lag of 0 never stands in for streaming), inactive physical slots and the WAL each retains, volume growth. Lag says how many standbys it covers ("lag 0 s (1 of 2 standbys reporting)") and is never healthy when partial. Live reads on the Cluster page replace a fleet warning only with evidence that disproves it. When series could only be matched to the cluster by namespace and Pod name, the finding says so.
  • Sessions: aggregates for everyone; a blocker → victims tree for callers with pods/exec; connection headroom.
  • History: Prometheus (15 m–24 h) or samples taken while the page is open, labelled as such; gaps hatched with their reason; selecting an interval opens Logs and Activity for exactly that window.

Storage (cnpg_storage.go, CNPGStorage.tsx)

  • Per-instance data, WAL and tablespace volumes: requested vs capacity, resize state. Used % from kubelet volume stats; without them, the WAL an instance reports is a lower bound ("≥ 3.3 GiB used by WAL alone"). What holds WAL (archive queue, slot retention per standby, WAL size) side by side, never summed. Guided resize through the GitOps guard. ≥ 80 % / ≥ 90 % measured usage raises a problem.

HA, certificates, declarations (cnpg_cluster_ha.go, internal/issues/source_cnpg_*.go)

  • FailoverQuorum as recorded configuration (R + W > N), PDBs, zones and nodes, the primary-election and operator Leases, Jobs and image drift — on the Replication tab. Certificate expiry with who renews it — on Configuration and as an issue.
  • A schedule-aware "the last scheduled run produced no successful backup" issue, judged only from Backup and ObjectStore lists that were read.
  • DatabaseRole (CNPG 1.30) across the workspace and Declarations.

Operator (cnpg_operator*.go, /api/cnpg/operator/status)

  • Current state first: clusters blocked on a plugin, components not ready or scaled to 0, a leader Lease not renewed, a Fail webhook with no ready endpoint, restart history per component (with the last termination), what it confirmed and what it could not read. Then leader, watched namespaces, webhooks and reconcile errors from the operator's metrics.
  • A per-namespace "is the operator reconciling" verdict drives a "status may be stale" banner on the fleet and cluster pages.

Parity with Freelens and Headlamp (excluding guided creation forms)

  • Fleet ordered by urgency; a Connect section (-rw / -ro / -r / Pooler hosts with honest labels, database, owner and the app Secret name, never read); schedules in plain language with next runs from the operator's own cron parser; logical replication Subscription → Publication → publisher slot and whether it survives failover; replica clone progress; per-database health; checkpoints; in-page trends without Prometheus (rates per exporter run and per primary, so a switchover is a gap, never a spike).

Shared pieces

  • GitOps write guard: POST /api/gitops/write-evidence + evaluateGitOpsWriteGuard / GitOpsWriteWarning in k8s-ui; evidence classes last-applied, GitOps field manager, ignore rules, owner sync policy; used by the CNPG actions, Set image and the investigation apply dialog.
  • Radar Cloud RBAC: FailoverQuorum added to the reviewed CNPG integration-read grant; Leases added to the cluster-read add-on (only viewers gain). Secrets deliberately not added.
  • Timeline: Pod rows are tagged by the state the change produced (Radar-wide).
  • Permission checks (Radar-wide): k8score.CanI sends subresources in their own field, so webhook authorizers such as GKE IAM answer pods/exec and pods/proxy correctly.
  • Shared Prometheus: PVC usage and CNPG history pin queries only to a proven cluster identity; an ambiguous identity reads as unavailable, never as a merged value.
  • Logs viewer (k8s-ui, additive): initialContainer preselects a container.
  • Renamed: ManagedImageSource / SetImageDialog.managedSources / ResourceActionsBar.managedImageSources → imageOwnership. Radar Hub does not use them.
  • Demo: make cnpg-demo-runtime adds a live namespace with real plugin backups and a restore, a pooler, pgbench load, a blocked lock chain, and a Prometheus that also scrapes kubelet volume stats.

All k8s-ui changes are additive props or new exports. Radar Hub is unaffected until it opts in.

Testing

  • go build ./..., go test ./..., make tsc, k8s-ui tsc, vitest (k8s-ui 4,399, web 2,053) and web lint (0 errors). CI's workflow runs only on PRs into main, so this PR shows Bugbot only until it is retargeted.
  • Tests cover every action's write shape and fact binding, permission reviews, runtime/storage/history gating, the operator verdict and current state, the operation tracker, sustained lag and receiver loss (a standby throughout the window), lag coverage (including replica clusters), slot evidence and when a live read may clear it, the archiving repair's resume boundary and which Backups count, blocked-phase wording, and the guard evaluator.
  • Exercised live on the kind demo, full access and a restricted viewer (view + CNPG read: no pods/proxy, no exec, no writes): the fleet; each Cluster tab at 1920 and 1280; switchover candidates with concerns; a standby stuck on an old timeline with its inactive slot holding ~5 GiB (Storage relief → Replication); a WAL-failing cluster's repair panel, its link landing on the plugin-barman-cloud logs that show the unreachable endpoint; the Operator naming plugin-blocked clusters. The same pass under real degradation — operator and plugin crash-looping on startup-probe timeouts, primaries flapping — read as stale, unassessed or unknown with the reason, never as healthy. Restricted mode names each missing grant and shows nothing as zero.

Notes / follow-ups

  • Not included (each a product decision): checks run inside Radar to prove a restore usable, a guided configuration-change journey (restart vs reload, availability impact), a connection test through the selected route, and verifying the certificate actually served after renewal.
  • Guided creation forms (Freelens has them) are not included; creation goes through Radar's YAML create flow.
  • On kind's local-path volumes kubelet publishes no volume stats, so the demo shows "no usage metrics". The measured paths are unit-tested.

Note

High Risk
Adds many impersonated write and pods/proxy/exec paths for PostgreSQL operations (switchover, destroy, session terminate, restore) plus GitOps write-evidence; mis-gating or fact-binding bugs could allow unsafe changes or misleading health signals.

Overview
CloudNativePG grows from a read-only fleet workspace into task-shaped cluster pages (Replication, Storage, Performance, Backups, etc.) backed by many new /api/cnpg/* routes: live instance/pooler runtime via caller pods/proxy, storage and HA facts, Prometheus history with strict cluster-identity scoping, sessions/blocking and destroy flows via pods/exec, capabilities + POST actions bound to reviewed facts (409 on drift), restore/recovery, report bundles, and an operation tracker on the UI. Fleet evidence and problems are tightened so unknown/partial data is never shown as zero or healthy, with shared facts/problems/workspace components and documented certainty rules in DESIGN.md.

Cross-cutting additions include Grant / ReadSource for permission wording, POST /api/gitops/write-evidence and client GitOps write guard wiring, cnpgWorkspace (and related) feature flags, Issues for certificate expiry and missed scheduled backups, DatabaseRole support, timeline Pod health on change, Prometheus PVC usage batching and clearer unreachable discovery errors, Radar Cloud Lease read for viewers plus expanded CNPG integration-read grants, and make cnpg-demo-runtime for live backup/restore/load/lock fixtures. Docs and INTEGRATION_GUIDE add a full workspace integrations checklist.

Reviewed by Cursor Bugbot for commit 349ccfb. Bugbot is set up for automated code reviews on this repo. Configure here.

…e guard

Operations (parity with the Freelens and Headlamp CNPG integrations): back up
now, switchover, restart cluster or one instance, reload, fence and lift,
hibernate and resume, schedule suspend/resume/run-now, and restore to a new
cluster through the existing create flow in strict create mode.
- GET .../capabilities answers per action and per instance from SubjectAccess
  Reviews and the cluster's state; buttons are disabled with the reason.
- POST .../actions/{action} writes as the user and binds the confirmation to
  the facts the dialog showed (cluster and target Pod UIDs, primaries, fencing,
  hibernation, context): any change returns 409 instead of acting on a
  different situation. Status transitions mirror kubectl-cnpg; Pod deletes use
  UID preconditions; a Backup create that times out is resolved by reading the
  same name back.
- ActionConfirmDialog (k8s-ui) leads with the effect, keeps the literal API
  writes in an expandable section, and asks for the object's name before
  disruptive actions.

Runtime: GET .../runtime reads each instance's /pg/status and exporter metrics
through the impersonated pods/proxy (fixed GET paths only, no redirects,
per-Pod/per-endpoint TLS, byte caps, a short identity-keyed memo). The Cluster
page gains a Runtime tab (replication topology with lag, sessions by state,
transactions, storage and WAL, slots, trends sampled while open) and Pooling
shows live PgBouncer pressure. Denied, unreachable and missing measurements
are reported as such, never as zero.

GitOps write guard, shared by every write dialog: POST /api/gitops/write-
evidence reads the target and its owner as the caller and reports, per field,
last-applied presence, owning field managers and ignore rules plus the owner's
sync policy; evaluateGitOpsWriteGuard classifies it and GitOpsWriteWarning
asks for acknowledgment when a sync may overwrite the change. SetImageDialog
and the Diagnose apply dialog now use it (ResourceActionsBar's
managedImageSources prop becomes imageOwnership).
…ard, and honest runtime totals

- ScheduledBackup run now binds the schedule's generation; changed settings return 409
- Helm drift exemptions no longer read as safe: the next upgrade still overwrites
- Transaction rates use the scrape time, deduplicating memoized samples
- Pooler totals are lower bounds unless every pod and pool reported them
- An unknown write outcome locks confirm instead of inviting a duplicate
- Refused-for-changed-facts actions re-read capabilities
- Restore dialog reads the Cluster from useResource correctly
@nadaverell
nadaverell requested a review from hisco as a code owner September 29, 2026 09:39
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Add CloudNativePG actions, live runtime, and a shared GitOps write guard

✨ Enhancement 🐞 Bug fix 🧪 Tests 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Add CloudNativePG operations that refuse writes when reviewed cluster facts have changed.
• Show live PostgreSQL and PgBouncer state through the caller’s Kubernetes permissions.
• Share field-aware GitOps revert warnings across CloudNativePG, Set image, and investigation apply.
Diagram

graph TD
  UI["Web UI"] --> API["CNPG API"] --> Actions["Action handlers"] --> K8s["Kubernetes API"]
  API --> Runtime["Runtime readers"] --> Pods["CNPG Pod endpoints"]
  UI --> Guard["Shared write guard"] --> Evidence["Evidence endpoint"] --> K8s
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Read runtime through pods/exec
  • ➕ Could support richer SQL queries, including per-session detail.
  • ➖ Requires a more invasive grant and could expose query contents; loses the fixed-path proxy boundary.
2. Keep write warnings within each dialog
  • ➕ Smaller immediate change to existing Set image and investigation flows.
  • ➖ Duplicates policy interpretation and risks inconsistent revert warnings.

Recommendation: Keep the shared evidence-backed guard and bounded pods/proxy approach: they reuse caller authorization without requiring exec or duplicating GitOps policy logic. Review proxy safety, fact binding, and fail-closed confirmation closely; successful Backup creation and switchover were not verified live.

Files changed (44) +9994 / -427

Enhancement (27) +7086 / -19
cnpg_actions.goImplement guarded CNPG cluster and schedule actions +2045/-0

Implement guarded CNPG cluster and schedule actions

• Adds capabilities and impersonated handlers for backups, switchover, restarts, reload, fencing, hibernation, and schedules. Rechecks reviewed facts and permissions, uses optimistic locking or Pod UID preconditions, and reports stable conflict and uncertain-outcome codes.

internal/server/cnpg_actions.go

cnpg_runtime.goRead bounded CNPG and PgBouncer runtime data +1472/-0

Read bounded CNPG and PgBouncer runtime data

• Adds impersonated pods/proxy readers for fixed instance-manager and exporter paths. Validates Pod ownership, bounds and memoizes reads, and preserves denial and missing-data states.

internal/server/cnpg_runtime.go

gitops_write_evidence.goDerive field-level GitOps write evidence +787/-0

Derive field-level GitOps write evidence

• Checks requested paths against last-applied data, managed fields, ignore rules, and Argo, Flux, or Helm policies as the caller. Returns derived evidence rather than raw ownership metadata.

internal/server/gitops_write_evidence.go

server.goRegister CNPG and write-evidence routes +7/-0

Register CNPG and write-evidence routes

• Wires runtime, capability, action, and GitOps evidence endpoints into the API router.

internal/server/server.go

ActionConfirmDialog.tsxAdd reusable effect-first confirmation +227/-0

Add reusable effect-first confirmation

• Introduces context, warnings, expandable API writes, optional typed confirmation, guard gating, inline errors, and unknown-outcome locking.

packages/k8s-ui/src/components/shared/ActionConfirmDialog.tsx

CreateResourceDialog.tsxAllow strict create as the initial mode +5/-2

Allow strict create as the initial mode

• Adds an initial mode prop so prefilled restore manifests need not default to applying over existing objects.

packages/k8s-ui/src/components/shared/CreateResourceDialog.tsx

GitOpsWriteWarning.tsxRender shared GitOps write-risk warnings +131/-0

Render shared GitOps write-risk warnings

• Displays revert reasons, owner navigation, pending and informational states, and acknowledgement controls when needed.

packages/k8s-ui/src/components/shared/GitOpsWriteWarning.tsx

index.tsExport new shared confirmation components +3/-1

Export new shared confirmation components

• Exports action confirmation, GitOps warning, and the Set image ownership contract.

packages/k8s-ui/src/components/shared/index.ts

drawer-components.tsxAllow alert banners to control spacing +4/-2

Allow alert banners to control spacing

• Adds an optional class name while preserving default spacing so warnings fit within dialogs.

packages/k8s-ui/src/components/ui/drawer-components.tsx

WorkloadView.tsxExpose domain actions in workload headers +5/-0

Expose domain actions in workload headers

• Adds an optional header-action renderer for drawer and expanded views.

packages/k8s-ui/src/components/workload/WorkloadView.tsx

gitops-write-guard.tsClassify per-write GitOps revert risk +445/-0

Classify per-write GitOps revert risk

• Adds a pure evaluator yielding none, info, may-revert, or will-revert assessments. Treats missing evidence conservatively and exposes a common acknowledgement gate.

packages/k8s-ui/src/utils/gitops-write-guard.ts

index.tsExport write-guard utilities +1/-0

Export write-guard utilities

• Exposes the evaluator and types through the k8s-ui utilities entry point.

packages/k8s-ui/src/utils/index.ts

client.tsAdd write-evidence API client +19/-0

Add write-evidence API client

• Posts resource, owner, and written paths to the server and types its evidence response.

web/src/api/client.ts

cnpg.tsType and fetch CNPG actions and runtime +273/-2

Type and fetch CNPG actions and runtime

• Adds capability and runtime contracts, polling hooks, action mutations, error-code extraction, and cache invalidation.

web/src/api/cnpg.ts

CNPGClusterRuntime.tsxShow live PostgreSQL runtime +472/-0

Show live PostgreSQL runtime

• Adds replication, session, transaction, storage, WAL, slot, and page-local trend views with source availability and instance actions.

web/src/components/cnpg/CNPGClusterRuntime.tsx

CNPGDetailPage.tsxAdd a Runtime tab to cluster details +16/-2

Add a Runtime tab to cluster details

• Mounts live runtime on Cluster pages and connects Pod-log navigation to the logs tab.

web/src/components/cnpg/CNPGDetailPage.tsx

CNPGPooling.tsxDisplay measured PgBouncer pressure +50/-7

Display measured PgBouncer pressure

• Replaces the placeholder with exporter-backed waiting clients, active servers, maximum wait, and reporting coverage; marks incomplete totals as lower bounds.

web/src/components/cnpg/CNPGPooling.tsx

CNPGSummaryHost.tsxUse live replication in the cluster summary +27/-1

Use live replication in the cluster summary

• Shows primary-reported streaming and replay lag when available, retaining the Kubernetes-only fallback otherwise.

web/src/components/cnpg/CNPGSummaryHost.tsx

CNPGClusterActions.tsxProvide cluster operations and confirmations +525/-0

Provide cluster operations and confirmations

• Adds capability-gated backup, switchover, restart, reload, fencing, and hibernation flows with effects, write plans, GitOps checks, and error handling.

web/src/components/cnpg/actions/CNPGClusterActions.tsx

CNPGInstanceActions.tsxAdd per-instance operations +48/-0

Add per-instance operations

• Exposes eligible promote, restart, fence, and lift-fence controls using per-instance capability reasons.

web/src/components/cnpg/actions/CNPGInstanceActions.tsx

CNPGRestoreDialog.tsxReview restore sources before creation +86/-0

Review restore sources before creation

• Selects a Backup or object-store source and optional point-in-time target, then opens the standard create flow with a recovery manifest.

web/src/components/cnpg/actions/CNPGRestoreDialog.tsx

CNPGScheduleActions.tsxAdd ScheduledBackup controls +99/-0

Add ScheduledBackup controls

• Provides capability-gated run-now, suspend, and resume dialogs with catch-up information and GitOps checks.

web/src/components/cnpg/actions/CNPGScheduleActions.tsx

actionModel.tsModel standby choice and recovery manifests +95/-0

Model standby choice and recovery manifests

• Supplies backup naming, method descriptions, standby ranking, source selection, and recovery manifest construction.

web/src/components/cnpg/actions/actionModel.ts

renderCNPGHeaderActions.tsxRoute CNPG header actions by kind +12/-0

Route CNPG header actions by kind

• Renders cluster or ScheduledBackup controls only for CloudNativePG resources.

web/src/components/cnpg/actions/renderCNPGHeaderActions.tsx

useCNPGWriteGuard.tsxAdapt shared guard to CNPG actions +58/-0

Adapt shared guard to CNPG actions

• Maps CNPG write effects to the common warning and confirmation contract.

web/src/components/cnpg/actions/useCNPGWriteGuard.tsx

InvestigationView.tsxResolve write risk for proposed fixes +35/-2

Resolve write risk for proposed fixes

• Requests a conservative guard for unspecified spec changes when Apply opens and passes owner navigation to the dialog.

web/src/components/diagnose/InvestigationView.tsx

useGitOpsWriteGuard.tsFetch and evaluate GitOps evidence on demand +139/-0

Fetch and evaluate GitOps evidence on demand

• Resolves ownership, requests field evidence for managed targets, and evaluates intended writes against the shared guard.

web/src/hooks/useGitOpsWriteGuard.ts

Refactor (5) +395 / -391
ResourceActionsBar.tsxPass evaluated ownership to Set image +4/-4

Pass evaluated ownership to Set image

• Replaces bespoke managed-image sources with the shared guard assessment passed to the image dialog.

packages/k8s-ui/src/components/shared/ResourceActionsBar.tsx

SetImageDialog.tsxReplace bespoke image warning with shared guard +43/-82

Replace bespoke image warning with shared guard

• Defines image-field write paths and uses common warning and confirmation logic instead of dialog-specific GitOps copy.

packages/k8s-ui/src/components/shared/SetImageDialog.tsx

ApplyDialog.tsxUse shared warning for investigation apply +19/-26

Use shared warning for investigation apply

• Replaces the managed-by warning with a guard assessment, owner link, and common confirmation gate.

web/src/components/diagnose/ApplyDialog.tsx

WorkloadView.tsxWire CNPG actions and reuse ownership resolution +87/-279

Wire CNPG actions and reuse ownership resolution

• Connects CNPG header actions, extracts ownership handling into a shared hook, and supplies the evaluated guard to Set image.

web/src/components/workload/WorkloadView.tsx

useResolvedGitOpsOwner.tsExtract reusable owner resolution +242/-0

Extract reusable owner resolution

• Moves direct and inherited GitOps or Helm ownership lookup, Argo namespace resolution, owner fetching, and source descriptions out of WorkloadView.

web/src/hooks/useResolvedGitOpsOwner.ts

Tests (9) +2428 / -12
cnpg_actions_test.goTest CNPG action guards and write shapes +898/-0

Test CNPG action guards and write shapes

• Uses fake Kubernetes clients to verify capabilities, conflicts, fencing and plugin cases, and action writes and errors.

internal/server/cnpg_actions_test.go

cnpg_runtime_test.goTest runtime parsing and proxy boundaries +693/-0

Test runtime parsing and proxy boundaries

• Exercises status and exporter fixtures, pooler metrics, access failures, Pod validation, and request limits.

internal/server/cnpg_runtime_test.go

gitops_write_evidence_test.goTest write-evidence classification +339/-0

Test write-evidence classification

• Covers field paths, last-applied and managed-field evidence, owner policies, and ignore rules.

internal/server/gitops_write_evidence_test.go

GitOpsWriteWarning.test.tsxTest shared warning presentation +70/-0

Test shared warning presentation

• Verifies revert copy, owner links, acknowledgement, pending checks, informational writes, and unmanaged rendering.

packages/k8s-ui/src/components/shared/GitOpsWriteWarning.test.tsx

SetImageDialog.test.tsTest guard gating for image changes +23/-6

Test guard gating for image changes

• Updates submit tests for managed acknowledgement, unresolved ownership, and existing draft and loading gates.

packages/k8s-ui/src/components/shared/SetImageDialog.test.ts

gitops-write-guard.test.tsTest GitOps risk decisions +295/-0

Test GitOps risk decisions

• Covers status and child writes, evidence, sync policies, exemptions, uncertainty, and acknowledgement gates.

packages/k8s-ui/src/utils/gitops-write-guard.test.ts

actionModel.test.tsTest action selection and restore manifests +56/-0

Test action selection and restore manifests

• Checks backup naming, standby preference, recovery sources, and manifests that do not inherit source archiving settings.

web/src/components/cnpg/actions/actionModel.test.ts

ApplyDialog.test.tsxTest investigation apply ownership gates +52/-5

Test investigation apply ownership gates

• Checks warning and acknowledgement behavior, unmanaged confirmation, and blocking during ownership resolution.

web/src/components/diagnose/ApplyDialog.test.tsx

WorkloadView.test.tsFollow extracted owner-lookup helper +2/-1

Follow extracted owner-lookup helper

• Updates ownership tests to import the helper from its reusable hook module.

web/src/components/workload/WorkloadView.test.ts

Documentation (3) +85 / -5
CLAUDE.mdDocument new CNPG and GitOps endpoints +4/-1

Document new CNPG and GitOps endpoints

• Lists runtime, action, and write-evidence routes, authorization boundaries, and the expanded CNPG demo mode.

CLAUDE.md

cnpg.mdDescribe live runtime, operations, and write warnings +34/-4

Describe live runtime, operations, and write warnings

• Documents runtime availability and permissions, action write shapes, fact-bound confirmation, restore, and the GitOps guard.

docs/cnpg.md

CNPG_WORKSPACE.mdRecord phase-two CNPG design +47/-0

Record phase-two CNPG design

• Adds the actions, runtime, fixtures, and shared-guard plan, including revisions on fact binding, fencing, and proxy safety.

docs/plans/CNPG_WORKSPACE.md

@qodo-free-for-open-source-projects

qodo-free-for-open-source-projects Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (2) 📘 Rule violations (3) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Backups can use an unreviewed target ✓ Resolved
Description
cnpgClusterActionRunners binds backup confirmation to hibernation but not to the cluster's backup
target. When the operator selects “Cluster default” and that target changes before the POST, the
Backup is created without an explicit target and inherits the new value.
Code

internal/server/cnpg_actions.go[1244]

+	"backup":          {binds: []string{"hibernation"}, run: cnpgRunBackup},
Evidence
The server supplies a backup target as a reviewed fact, but neither decodes nor compares it for
backup actions; an omitted explicit target leaves the created Backup to use the current cluster
default.

internal/server/cnpg_actions.go[263-274]
internal/server/cnpg_actions.go[1243-1245]
internal/server/cnpg_actions.go[1488-1493]
web/src/components/cnpg/actions/CNPGClusterActions.tsx[243-249]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A backup using the cluster default can inherit a target that changed after confirmation.
## Fix Focus Areas
- internal/server/cnpg_actions.go[263-274]
- internal/server/cnpg_actions.go[1243-1245]
- internal/server/cnpg_actions.go[1265-1290]
## Recommended Fix
Decode `backupTarget` from the reviewed facts and include it in the backup runner's fact comparison. Return `changed` when the current value differs.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Action errors bypass the shared format ✓ Resolved
Description
writeCNPGActionError writes typed action errors with w.WriteHeader and json.NewEncoder instead
of s.writeError. When an action is refused or its reviewed facts have changed, these responses use
a separate error-writing path that also carries the dialog's action code and current facts.
Code

internal/server/cnpg_actions.go[R1999-2000]

+		w.WriteHeader(ae.Status)
+		if encErr := json.NewEncoder(w).Encode(body); encErr != nil {
Evidence
The new action-error path directly writes an error status and JSON body rather than using the
required helper.

Rule 3036617: Use standard JSON error helper in HTTP handlers
internal/server/cnpg_actions.go[1987-2003]


3. Action failures lack the required log format ✓ Resolved
Description
writeCNPGActionError logs fallback errors with a quoted action and -> %d rather than the
required [module] Failed to ... %s/%s: %v structure. Unexpected errors retain the default 500
status and reach this log before s.writeError sends the response.
Code

internal/server/cnpg_actions.go[2026]

+	log.Printf("[cnpg] %q %s/%s -> %d: %v", action, sanitizeForLog(namespace), sanitizeForLog(name), status, err)
Evidence
The fallback status is 500, but its preceding log has a different literal prefix and placeholder
structure.

Rule 3036628: Log 500 errors with standardized module/action format before writing the response
internal/server/cnpg_actions.go[2005-2034]


4. Write-evidence failures use a different log format 📘 Rule violation ◔ Observability
Description
handleGitOpsWriteEvidence logs target-read failures as `Failed to read %s %s/%s for write
evidence: %v`, with an extra kind placeholder and words between the resource name and error. An
unclassified Kubernetes Get error then reaches the 500 response on that branch.
Code

internal/server/gitops_write_evidence.go[161]

+			log.Printf("[gitops] Failed to read %s %s/%s for write evidence: %v", sanitizeForLog(req.Kind), sanitizeForLog(req.Namespace), sanitizeForLog(req.Name), err)
Evidence
The log immediately precedes a 500 response but does not use the required %s/%s: %v ending.

Rule 3036628: Log 500 errors with standardized module/action format before writing the response
internal/server/gitops_write_evidence.go[153-164]


View medium (9)
5. Invalid action input receives status 422 📘 Rule violation ≡ Correctness
Description
writeCNPGActionError maps apierrors.IsInvalid(err) to http.StatusUnprocessableEntity rather
than 400. When Kubernetes rejects an action write as invalid, that response follows a different
client-input status mapping from the one required for these handlers.
Code

internal/server/cnpg_actions.go[R2023-2024]

+	case apierrors.IsInvalid(err):
+		status = http.StatusUnprocessableEntity
Evidence
The added invalid-error branch explicitly selects 422 instead of the checklist's 400 for invalid
client input.

Rule 3036624: Use canonical HTTP status codes for backend error responses
internal/server/cnpg_actions.go[2021-2025]


6. Cluster actions omit mutation messages ✓ Resolved
Description
useCNPGAction creates a mutation without meta.errorMessage or meta.successMessage. Both
cluster and scheduled-backup actions use this hook, so neither supplies the message metadata
expected by the centralized mutation feedback path.
Code

web/src/api/cnpg.ts[R241-243]

+  return useMutation<CNPGActionResult, Error, { action: string; request: CNPGActionRequest; successMessage: string }>({
+    mutationFn: ({ action, request }) =>
+      fetchJSON<CNPGActionResult>(`${cnpgPath(kind, namespace, name)}/actions/${action}`, {
Evidence
The added mutation options contain mutationFn, onSuccess, and onError, but no meta object.

Rule 3036648: React Query mutations must define toast metadata and avoid per-mutation toast handlers
web/src/api/cnpg.ts[239-257]


7. Replication bars use raw background colors ✓ Resolved
Description
ReplicationView selects bg-emerald-500, bg-amber-500, and bg-red-500 for its lag bar instead
of theme background tokens. Each replica row uses one of these classes according to its lag tone,
and no design-spec exception is documented above the element.
Code

web/src/components/cnpg/CNPGClusterRuntime.tsx[231]

+                    <div className={clsx('h-full', tone === 'healthy' ? 'bg-emerald-500' : tone === 'degraded' ? 'bg-amber-500' : tone === 'unhealthy' ? 'bg-red-500' : 'bg-transparent')} style={{ width: `${pct}%` }} />
Evidence
The newly added bar uses three non-theme background color utilities.

Rule 3036653: Use theme background tokens instead of hardcoded utility color classes
web/src/components/cnpg/CNPGClusterRuntime.tsx[229-233]


8. Destructive buttons use raw red backgrounds 📘 Rule violation ⚙ Maintainability
Description
ActionConfirmDialog assigns bg-red-600 and hover:bg-red-700 to its disruptive confirm button
instead of theme background tokens. Every action marked disruptive takes that branch, with no
documented design-spec exception above the button.
Code

packages/k8s-ui/src/components/shared/ActionConfirmDialog.tsx[213]

+            disruptive ? 'bg-red-600 text-white hover:bg-red-700' : 'btn-brand',
Evidence
The added button styling contains hardcoded background color utilities on the disruptive branch.

Rule 3036653: Use theme background tokens instead of hardcoded utility color classes
packages/k8s-ui/src/components/shared/ActionConfirmDialog.tsx[207-215]


9. An action comment repeats its control flow ⊘ Outdated
Description
The comment above renderCNPGHeaderActions says it returns operations for supported kinds and
null for everything else. The immediately following kind checks and returns express the same
behavior, so the comment adds no rationale or constraint for a later change.
Code

web/src/components/cnpg/actions/renderCNPGHeaderActions.tsx[6]

+/** Operations for the CNPG kinds that have them; null for everything else. */
Evidence
The added comment only describes the branches visible directly beneath it.

Rule 3036542: Avoid explanatory comments that restate obvious code behavior
web/src/components/cnpg/actions/renderCNPGHeaderActions.tsx[6-11]


10. Replication reports the wrong sent position 🐞 Bug ≡ Correctness
Description
cnpgPgStatus.ReplicationInfo.SentLsn decodes receivedLsn instead of sentLsn. When a status
response supplies the sent position, the runtime response cannot report that value and instead
displays the received position under its sent-position label.
Code

internal/server/cnpg_runtime.go[926]

+		SentLsn         string          `json:"receivedLsn"`
Evidence
The decoder's JSON tag names a different position, and the decoded value is subsequently exposed as
sentLsn.

internal/server/cnpg_runtime.go[923-928]
internal/server/cnpg_runtime.go[658-678]
internal/server/cnpg_runtime_test.go[35-40]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The replication status decoder reads the received position into its sent-position field.
## Fix Focus Areas
- internal/server/cnpg_runtime.go[923-928]
- internal/server/cnpg_runtime_test.go[35-40]
## Recommended Fix
Correct the JSON tag for `SentLsn` and test a response with distinct sent and received positions.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


11. Runtime can show stale protocol results ✓ Resolved
Description
cnpgMemoized omits the proxy target's HTTP scheme from its cache key. If TLS configuration changes
while the Pod, port and path remain the same, subsequent reads can reuse an HTTP result for an HTTPS
target, or vice versa, until the memo expires.
Code

internal/server/cnpg_runtime.go[722]

+	key := fmt.Sprintf("%s\x00%s/%s\x00%s\x00%d%s", identity, target.namespace, target.pod, target.podUID, target.port, target.path)
Evidence
The proxy target selects a scheme for the request, while memo lookup identifies targets without that
scheme; metrics entries can remain cached for 25 seconds.

internal/server/cnpg_runtime.go[72-78]
internal/server/cnpg_runtime.go[721-728]
internal/server/cnpg_runtime.go[760-765]
internal/server/cnpg_runtime.go[810-817]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Runtime reads under different proxy schemes share one memo entry.
## Fix Focus Areas
- internal/server/cnpg_runtime.go[721-728]
## Recommended Fix
Add `target.scheme` to the memo key and test that HTTP and HTTPS targets cannot reuse each other's results.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


12. Unknown instances appear as replicas 🐞 Bug ≡ Correctness
Description
CNPGClusterRuntime classifies every instance whose role is not primary as a replica. The backend
also emits unknown, so an instance with an indeterminate role enters the replica cards even though
its role was not established.
Code

web/src/components/cnpg/CNPGClusterRuntime.tsx[140]

+  const replicas = data.instances.filter((i) => i.role !== 'primary')
Evidence
The backend has an unknown role, and the frontend's non-primary filter passes it to the
replica-card rendering path.

internal/server/cnpg_runtime.go[625-631]
web/src/components/cnpg/CNPGClusterRuntime.tsx[139-152]
web/src/components/cnpg/CNPGClusterRuntime.tsx[211-243]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Instances of unknown role are included in the runtime replica display.
## Fix Focus Areas
- web/src/components/cnpg/CNPGClusterRuntime.tsx[139-152]
- web/src/components/cnpg/CNPGClusterRuntime.tsx[211-243]
## Recommended Fix
Select replicas by an explicitly known replica role and display unknown-role instances separately without presenting them as confirmed standbys.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


13. Named restores retain a prior time ✓ Resolved
Description
CNPGRestoreDialog passes targetTime to buildRestoreManifest even after the user switches to a
named Backup source, which disables the time input without clearing it. If the user entered a
point-in-time value under an object-store source first, the resulting named-Backup manifest still
includes recoveryTarget.targetTime despite the disabled control saying that a named Backup
restores to its endpoint.
Code

web/src/components/cnpg/actions/CNPGRestoreDialog.tsx[45]

+        const m = buildRestoreManifest(cluster, source, newName, targetTime ? new Date(targetTime).toISOString() : undefined)
Evidence
The source switch only disables the input; the confirmation still forwards its retained state, and
the manifest builder writes that state for every source kind.

web/src/components/cnpg/actions/CNPGRestoreDialog.tsx[43-45]
web/src/components/cnpg/actions/CNPGRestoreDialog.tsx[69-80]
web/src/components/cnpg/actions/actionModel.ts[70-75]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A time entered for one restore source remains in the manifest after switching to a named Backup.
## Fix Focus Areas
- web/src/components/cnpg/actions/CNPGRestoreDialog.tsx[43-45]
- web/src/components/cnpg/actions/actionModel.ts[70-75]
## Recommended Fix
Only pass or emit `targetTime` for restore sources that support point-in-time selection. Test switching from an object store to a named Backup after entering a time.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Informational

14. Revert warnings appear for fields Argo ignores ✓ Resolved
Description
gitOpsIgnoredPath checks matching ignoreDifferences entries for jsonPointers and
managedFieldsManagers but never reads jqPathExpressions. When a jq expression covers a declared
write path, the resulting evidence does not mark it as effectively ignored, so the guard can
classify the write as reverting and request acknowledgement even with
RespectIgnoreDifferences=true.
Code

internal/server/gitops_write_evidence.go[R450-453]

+			covered := false
+			for _, p := range stringSlice(entry["jsonPointers"]) {
+				if pointerCovers(p, tokens) {
+					covered = true
Evidence
The Argo ignore-rule loop inspects only jsonPointers and managedFieldsManagers, leaving jq-based
coverage out of the ignored-path evidence. The frontend uses that evidence to distinguish
effectively ignored declared fields from fields subject to revert, which explains the false
classification.

internal/server/gitops_write_evidence.go[450-467]
packages/k8s-ui/src/utils/gitops-write-guard.ts[279-301]
internal/server/gitops_write_evidence.go[441-468]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Argo `ignoreDifferences` entries using `jqPathExpressions` are absent from write-path evidence and can produce false revert warnings.
## Fix Focus Areas
- internal/server/gitops_write_evidence.go[441-468]
- packages/k8s-ui/src/utils/gitops-write-guard.ts[279-301]
## Recommended Fix
Read `entry["jqPathExpressions"]` when establishing ignore coverage. Evaluate the expressions against written paths where coverage can be determined safely; otherwise, emit evidence that a jq rule exists but was not evaluated and classify coverage as unknown or may-revert rather than definitely unignored or will-revert. Add a classification test for a declared write path covered by a jq rule.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


15. A 'null' fencing value is treated as unfenced ✓ Resolved
Description
parseCNPGFenced flags the annotation as malformed only when json.Unmarshal fails, and null
decodes cleanly into a nil slice, so Malformed stays false. As a result the malformed-value block
on fence and unfence is skipped: unfence reports that no instance is fenced, and fence overwrites
the invalid annotation.
Code

internal/server/cnpg_actions.go[R325-329]

+	var names []string
+	if err := json.Unmarshal([]byte(raw), &names); err != nil {
+		out.Malformed = true
+		return out
+	}
Evidence
Unmarshalling null into a []string succeeds and leaves the slice nil. Only the Malformed flag
blocks fence and unfence.

internal/server/cnpg_actions.go[320-329]
internal/server/cnpg_actions.go[603-642]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
A fencedInstances annotation with the value `null` is accepted as valid and unfenced.
## Fix Focus Areas
- internal/server/cnpg_actions.go[320-329]
## Recommended Fix
After unmarshalling, set out.Malformed = true when raw is non-empty and names == nil, or require the trimmed raw value to start with '['.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Tip of the day
💡 Did you know, you can describe a rule in plain language on the Rules page and Qodo drafts it for you

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread internal/server/cnpg_actions.go Outdated
Comment thread internal/server/cnpg_actions.go Outdated
Comment thread internal/server/gitops_write_evidence.go
Comment thread internal/server/cnpg_actions.go Outdated
Comment thread web/src/api/cnpg.ts Outdated
Comment thread internal/server/cnpg_runtime.go Outdated
Comment thread web/src/components/cnpg/CNPGClusterRuntime.tsx Outdated
Comment thread web/src/components/cnpg/actions/CNPGRestoreDialog.tsx Outdated
Comment thread internal/server/gitops_write_evidence.go
Comment thread internal/server/cnpg_actions.go

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/server/cnpg_actions.go

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread web/src/components/cnpg/CNPGClusterRuntime.tsx Outdated
Comment thread web/src/components/cnpg/CNPGClusterRuntime.tsx Outdated

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread web/src/components/cnpg/CNPGPooling.tsx
…scheduled runs that produced no backup

GET /api/cnpg/clusters/{ns}/{name}/storage reports each instance's claims by
role (owner- and instance-validated), requested vs capacity, resize state,
StorageClass expansion, kubelet volume usage from Prometheus and the WAL facts
(size, segments, archive backlog, slot retention) through the Runtime tab's
memoized pods/proxy reads, each source with its own coverage. GET
/api/cnpg/disk returns the fullest measured volume per visible cluster, one
claim list and one usage batch per namespace.

The Issues engine raises CNPGScheduledRunNoBackup on a Cluster when an active
ScheduledBackup fired after its newest successful backup and no run since then
succeeded or is still running, using the operator's six-field cron parser.
…e and Pooler pause/resume endpoints

Sessions run fixed SQL (pg_stat_activity + pg_blocking_pids) with psql over the
caller's pods/exec; cancel/terminate revalidate pod UID, pid and backend_start in
the same statement and pass the user's values only as psql variables. Destroy
instance mirrors kubectl cnpg destroy (PVCs detached or deleted first, then the
Pod, then the instance's Jobs), binds the reviewed Pod and PVC UIDs and refuses
the primary. Poolers gain capabilities with Deployment/Service readiness,
pause/resume bound to the reviewed paused value, and per-Pod observed PgBouncer
state via SHOW STATE. Pooler runtime reports the pool mode PgBouncer uses.
…tion and DatabaseRole kind

GET /api/cnpg/clusters/{ns}/{name}/ha reads FailoverQuorum (with R + W > N),
PodDisruptionBudgets, primary and operator Leases, per-instance node, zone,
QoS and image drift, cluster Jobs, -rw endpoints and certificate renewal
ownership, each sub-read authorized on its own. The runtime status adds
pendingRestartForDecrease, pg_rewind, instance-manager version and a role
detail derived from the instance's own report, plus the postmaster start time.
Certificates close to expiry become issues. setMaintenance/unsetMaintenance
mirror kubectl cnpg maintenance. DatabaseRole joins the workspace kinds.
…nd disk use in the fleet

The Runtime tab's Storage & WAL section becomes the storage decision: each
instance's claims with used/capacity from kubelet volume stats (or why it is
unknown), resize state, whether the StorageClass allows expansion, and the WAL
on disk, archive backlog and slot retention side by side. Expansion names the
spec field and opens the apply flow with only that size, behind the GitOps
write guard. Volumes stay visible without pods/proxy.

The fleet gains a Disk column and the Cluster overview a Storage fact from
/api/cnpg/disk; a measured volume at 80% or more puts the cluster in Needs
attention. The renderer's configured sizes read as requested sizes and the
Runtime database sizes as logical sizes.
…ce and Pooler detail to the CNPG workspace

psql opens the dock terminal in an instance's postgres container (?shell=psql)
from instance rows and the cluster menu, labelled by the instance's current
role. The Runtime Sessions section adds the blocker-to-victim tree with
connection headroom and per-instance CPU/memory, and cancel/terminate dialogs
bound to pod, pid and backend_start. Standby rows gain Destroy with the
delete / keep-volumes choice. The Pooler summary shows Deployment readiness,
per-pool live pressure, limits with verified PgBouncer defaults, requested vs
observed pause and the Service path; the Pooling table reads readiness from the
Deployment instead of the scheduled count. Pause/resume is a header action.
…CNPG demo, and document storage

Schedule fire times were computed in Radar's local zone; the operator's clock
is UTC. The runtime demo's Prometheus now scrapes each kubelet through the
apiserver node proxy (kubelet_volume_stats_* only); kind's hostPath volumes
publish no stats, which the README states and the UI shows as no usage
metrics. Claims kept after an instance was destroyed are named as such, a
standby's archive fact says the primary archives, and a volume sharing a larger
filesystem says the measurement is that filesystem's.
…ow readiness source once

Pooler readiness detail names its source instead of repeating the count, and
the Pooler runtime's reported pool mode is pinned by a test.
# Conflicts:
#	web/src/components/cnpg/CNPGClusterRuntime.tsx

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/issues/cnpg_schedule_describe.go
…he header

- Overview keeps State and Protection open; HA and instances, and
  Certificates, fold to one summary line each and open themselves when
  something in them is out of line.
- Connect is a header button opening the connection facts in a dialog
  (?connect=ns/name), which the restore Next steps link opens too.
- The Replication chip reads no calmer than a sustained-lag finding,
  and the fleet says Pods ready rather than ready.
- WAL archiving shows since when it has worked, when that changed within
  the last day and well after creation.
- The pooler summary breaks PgBouncer pressure down per Pod.
- Switching the namespace picker to All refetches what is on screen once
  the server holds the new pick, so no request reads the old one.
…alog

- The HA summary names the facts it could not read instead of reading
  calm; the Certificates summary says nearest reported expiry, counts
  unreadable ones and opens itself for them.
- Connect's URL value names the surface, so a drawer showing the page's
  own Cluster does not open a second dialog, and opening or closing it
  keeps the page's return label.
- Switching to All namespaces cancels requests still in flight before
  refetching, so a first fetch that read the old pick is replaced.
A link following from the dialog and the dialog closing would both write
the URL in the same tick and undo each other, leaving the dialog over
the drawer it opened. The param now only asks for the dialog: the button
opens it and drops the param. The dialog is a step wider so endpoint
descriptions stay on one line.
A full-screen drawer can report the same surface as the page beneath it,
so both would open a dialog for the same request.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread packages/k8s-ui/src/components/cnpg/ha.ts Outdated
Comment thread web/src/components/cnpg/actions/CNPGConnectButton.tsx Outdated
…ton answer its request

A Cluster with no instance Pods now opens HA with "no instance Pods" rather
than "0/0 instances ready", and "images match" needs every instance to say so.
The Connect request is held by the button that answered it only while that
button is mounted, so a remount before the param is dropped still opens the
dialog. Pins five-field schedules to the parser's next runs.
nadaverell and others added 2 commits October 4, 2026 11:52
…workspace (#1969)

## Summary

The CloudNativePG workspace (#1921, #1922) carried a generic layer under
CNPG names: the facts and problem list on every summary, per-kind
coverage, the reviewed-action contract, grants, Prometheus series
attribution, and screen layout borrowed from Capacity's internals. This
PR moves that layer into shared, integration-neutral modules, so the
next workspace-style integration (Velero is the likely one) imports it
instead of copying CNPG.

Nothing here is released yet, so the three CNPG wire shapes that every
future workspace would copy are fixed now:
- grants are structured objects, not sentences;
- "not cached by Radar" is told apart from "no access";
- each object's GitOps manager comes from the server.

It also adds the version-skew gate the CNPG endpoints were missing.

Only mechanisms whose interface is already evident in today's code are
shared: each has two or more consumers, or is a contract every workspace
must follow. The rest is listed below as not shared yet.

## What changed

### Shared UI (`@skyhook-io/k8s-ui`)
- **`components/facts/`**, for any surface that shows observed values
(single-kind renderers included):
  - `Fact`, with `FactGrid`/`FactRow`/`FactValue`/`FactSource`
- `CertaintyGlyph`, moved from Capacity; `CapacityCertainty` stays as
the same type
  - `ManagedByText`
- **`components/problems/`**: `WorkspaceProblem<Category>` and
`ProblemCallout`/`ProblemList`/`ProblemMeta`, which take the workspace's
`rootKind` instead of a hard-coded `'Cluster'`, plus `OpenIssueContext`.
- **`ui/FoldSection`**: `SectionHeading`, `FoldSection`, `FoldSummary`.
- **Elsewhere in k8s-ui**:
  - `ui/RefLink` now uses the existing `ResourceRef`.
  - `toneTextClass` and `worseTone` live in `ui/status-tone`.
  - `utils/grant` adds `Grant`, `formatGrant` and `grantParts`.
- `issueReasonTitle` gives a reason one title on both the Issues page
and the workspace.
- **Kept CNPG-specific:** categories, ordering, wording and builders.
CNPG binds `WorkspaceProblem` to its own categories.
- **Badges:** CNPG badges now use `healthToSeverity` like the rest of
the app.
- **Buttons:** secondary buttons use a new `.btn-secondary` class
(documented in DESIGN.md) instead of ten hand-rolled class strings.
- **Sidebar:** `SidebarCategoryDestination.countLowerBound` renders a
count over partly read data as `≥N`, and its zero as unknown rather than
none.

### App (`web/`)
- **`components/workspace/`** holds the screen primitives Capacity and
CNPG share: `ScreenBody`, `ScreenEmptyState`, `Notice`, `Segments`,
`FilterChips`, `SectionTable` and its table classes,
`RefreshFailedNotice`, `GrantText`, and the text helpers. CNPG no longer
imports from `capacity/shared.tsx`.
- **`api/actions.ts`** is the client half of the action contract:
  - the `ActionCapability` and `ActionRequest` types
  - `actionErrorCode`, which an integration widens with its own codes
  - `actionOutcomeLocked` and `actionCompleted`
  - `capabilityReason`, now the one "why is this disabled" wording
- **Utilities:** `utils/drawer-trail.ts` holds the `?drawer=` codec, and
`utils/page-links.ts` the back label and the subject-filtered Issues
link.
- **Kind table:** CNPG's detail kinds derive kind and group from the
workspace kind table.

### Server
- **`internal/server/actions.go`**: the reviewed-action contract.
- `ActionCapability`, `ActionRequest`, `decodeActionRequest` and
`decodeActionParams`
- refusals: `changedAction`, `blockedAction`, `partialAction` and
`writeActionError`
- the permission decision: `grantPermission` and `capabilityVerdict`,
with a real SelfSubjectAccessReview in local mode
  - `mergePatchAtVersion`
- **Grants:** `Grant` lives in `internal/auth` and carries `In(ns)` and
`String()`.
- **Kind access:** `kind_access.go` covers per-kind access and coverage
(`readWorkspaceKind`, `typedKindScope`, `KindCoverage`);
`cache_scope.go` has `namespacesWithinCache`.
- **Other shared helpers:**
- `read_source.go`: one `ReadSource` shape for HA, recovery and storage
  - `fanout.go`: the bounded per-namespace fan-out
- `internal/prometheus/series_scope.go`: `SeriesIsolation`, already used
by the PVC usage charts
- **Kept CNPG-specific:** facts, guards, runners and the extra refusal
codes.
- **Exported manager detection:** `pkg/topology` now exports
`ManagedByFromMeta`. This is additive.

### Wire changes (CNPG endpoints, unreleased)
- **`grant`** is `{verb, group?, resource, subresource?, namespace?}`
(no namespace means cluster-wide) on every capability, read source,
chart and report item. The wording is unchanged and comes from
`Grant.String()` / `formatGrant`.
- **Kind coverage** gains:
  - a state `uncached`, for a scope Radar's cache does not cover at all;
- `uncachedNamespaces`, named under the same rule as `deniedNamespaces`.

Before, a namespace the caller can read but Radar does not cache was
reported as denied, and the UI said "No access". It now says "Radar does
not cache Pods in db".
- **`/api/cnpg/workspace`** gains `managedBy`, keyed `Kind/ns/name`.
- "Declared in" now uses the server's manager detection and links to the
Argo CD application or Flux object.
- Before, the client parsed labels itself. That misnamed Argo CD
applications living outside Argo CD's namespace (`ns_app`) and ignored
the Flux namespace label.

### Version skew
New `FeatureCapabilities` flags: `cnpgWorkspace` (every `/api/cnpg/*`
route except the two image-catalog lookups that predate it) and
`gitopsWriteEvidence`. Each has a `radarFeatures` entry.

Every CNPG query, mutation, log stream, report download and the
write-evidence query goes through `useRadarFeature`. Mutations are never
retried. On a Radar without the workspace:
- the sidebar offers no workspace destinations;
- `/workload` links and drawer Expand stay on the standard views;
- a Cluster's Logs tab shows its Pods' logs;
- a workspace screen opened directly says it needs a newer Radar.

### Vocabulary
On screen these areas are named by their subject. The block above a
sidebar category's kinds reads **Views**, and the copy says
"CloudNativePG views" and "CloudNativePG data", never "workspace",
because Radar Hub uses "workspace" for the customer's account.
"Workspace" remains the internal term for an integration with several
kinds and its own screens (the integration guide,
`SidebarCategoryWorkspace`, `/api/cnpg/workspace`).

### Docs
- **DESIGN.md:** the rules for unknown, partial and denied values are
written once. `capacity.md` and `cnpg.md` link to them and keep their
per-value tables.
- **`docs/INTEGRATION_GUIDE.md`:** a new "Workspace integrations"
chapter covering when a workspace is warranted, the pieces to import,
and what is not shared yet. CLAUDE.md points to it.

## Not shared yet
These have one consumer, and their interface should come from the
second:
- the workspace registry / App wiring
- the operation tracker
- the generic action runner
- the fixed-path `pods/proxy` reader
- the merged log stream
- the report bundle
- operator diagnosis
- a TTL-memo helper

Rollouts' capabilities answer "allowed" in local mode without asking the
apiserver. That is a real bug, but Rollouts returns bare booleans and
hides denied actions, so it needs its own PR.

## Public surface
- **k8s-ui:** exports that exist on main are unchanged. Renamed exports
were all added in #1921/#1922. `CapacityCertainty` is kept.
- **`pkg/`:** one additive export (`topology.ManagedByFromMeta`).
- **`pkg/capacityapi` v1alpha1:** unchanged.
- **radar-app:** surface unchanged.

## Testing
- **Go:** `go build ./...`; `go test ./...` in both modules. The
`cmd/desktop` env test is flaky and passes alone. `gofmt -l` is clean.
- **Frontend:**
  - type checks pass for web and k8s-ui;
  - vitest: k8s-ui 236 files / 4,369 tests, web 182 files / 2,028 tests;
  - eslint: 0 errors, and no new warnings.
- **Live on the kind CloudNativePG demo, full access and a view-only
identity:**
- every CNPG page rendered with structured grants and no "[object
Object]" or "undefined";
- a namespaced Argo CD tracking ID on a demo Cluster showed "Declared
in: Argo CD application argocd/payments", linked;
- with the capability flags removed from `/api/capabilities` (an older
Radar behind Radar Hub), the sidebar workspace, redirects, composed
summaries and actions fell back to the standard views, and `/cnpg`
showed the upgrade note.

Stacked on #1922 (which is stacked on #1921). Merge order: #1921, #1922,
then this PR, before a release ships `/api/cnpg/*`.

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Medium Risk**
> Touches CNPG API shapes, RBAC/coverage semantics, and cluster write
paths via the shared action layer; behavior is intended to be equivalent
with clearer grant and cache messaging.
> 
> **Overview**
> Extracts the generic workspace layer that CloudNativePG had under
CNPG-specific names into shared modules, so future integrations (e.g.
Capacity-style screens) import one contract instead of copying.
> 
> **Server:** Adds `internal/auth.Grant` (structured RBAC on the wire),
`internal/server/actions.go` (reviewed `ActionRequest`, capabilities,
409/partial errors, `mergePatchAtVersion`), `cache_scope.go` /
`kind_access` patterns, and `prometheus/series_scope.go` (shared
Prometheus isolation). CNPG handlers now use these; history/PVC denial
fields expose `*Grant` instead of strings. `FeatureCapabilities` adds
`cnpgWorkspace` and `gitopsWriteEvidence`. CNPG workspace coverage gains
`uncached` / `uncachedNamespaces` (distinct from denied) and `managedBy`
from server-side GitOps detection.
> 
> **UI:** New k8s-ui `facts/`, `problems/`, fold sections,
`formatGrant`, and app `components/workspace/` plus `api/actions.ts`.
CNPG is rewired to structured grants and shared action types; secondary
buttons use `.btn-secondary`.
> 
> **Docs:** DESIGN.md documents unknown/partial/denied values once;
INTEGRATION_GUIDE adds a workspace-integrations chapter; CLAUDE.md
points builders at it.
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
5af6ed1. Bugbot is set up for automated
code reviews on this repo. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->
…ture/cnpg-actions-runtime

# Conflicts:
#	web/src/components/cnpg/CNPGClusterLogs.tsx

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/issues/source_cnpg_schedule.go
A Backup list that failed or is not watched was treated as no backups, so
a cluster that still had backups could get "no successful backup since a
scheduled run". The detector now evaluates a cluster only when its
Backups were read, and, for a cluster backing up to an ObjectStore, that
store's list too.
# Conflicts:
#	packages/k8s-ui/src/components/cnpg/index.ts
…operator and plugin Pod restarts

Fleet metrics: a standby whose WAL receiver is down reads replication lag 0,
so the lag reading in state ok now carries standbys, receiving and
receiverDown from cnpg_pg_replication_is_wal_receiver_up (receiverUnknown
when no standby reports it), and each cluster gets a slots reading listing
inactive physical replication slots with the reporting instance's role and
retained WAL. Both are one query per namespace under the same get pods gate,
selector and isolation as lag.

Operator: each operator and plugin component lists its Deployment's Pods
(ready, restarts, current container start, last termination) with
podCoverage gated on list pods; the diagnosis Pods carry lastTermination too.
…surface

The Cluster page is organized by task: Overview, Replication, Storage,
Performance, Backups, Activity, Logs, Configuration, YAML. Replication joins
each instance's Pod readiness with its PostgreSQL role, timeline, streaming
state and the HA slot the primary keeps for it, plus the HA configuration;
Performance holds Sessions, Database health and History (chart groups);
Backups shows the cluster's recovery evidence as facts with every backup run
and each schedule's last outcome; Configuration adds connections and
certificates above the declared settings.

The fleet row, the header chips (now on every tab), the Overview and the
switchover dialog read one assessment. A standby that receives nothing is a
problem from Prometheus (WAL receiver down) or the instance managers (no
pg_stat_replication row), never "lag 0 s"; an inactive slot on the primary
holding 1 GiB or more is a problem naming the standby it serves. Switchover
candidates show what the live read sees wrong with them. The fleet views are
named Clusters and Backups; the fleet table shows Replication, Storage and
Backups beside the cluster's readiness. The Operator screen leads with its
current state and shows restart history. Escape that closes a menu no longer
also leaves the page.
…hat they disprove

A standby's WAL receiver must be down in every Prometheus sample for 5
minutes before the fleet raises it, so a restarting standby is shown but not
alarmed. A live read replaces the fleet's answer for what it covers: a standby
the primary streams to again, or a slot no longer inactive, clears the
earlier problem, and slots are judged from the primary's slot list even when
its replication rows are unusable. "None receiving" needs every expected
standby accounted for. Switchover never preselects a standby with concerns.
The Operator's current state only claims what it read and lists what it
could not. Backup runs show the newest 10 with a way to see all, and a
schedule's last run never borrows an older Backup's outcome. The Overview's
links name the tab they open, and the Configuration tab leads with
connections instead of repeating the Overview's problems.
The tab strip drops its icons and tightens spacing only while every tab does
not fit, measured against the space it has, so a page with many tabs (a
CloudNativePG Cluster's nine) stays on one line at 1280 px and pages that fit
keep their look. Header badges no longer break mid-word. The Operator's
current state names webhook configurations it could not read, and a
Cluster's Backups tab leaves out the Cluster column.
…ups tab

A live read clears a fleet slot warning only when it shows the slot active,
measured below the threshold, or gone from a complete list; an unmeasured
slot or a capped list leaves it standing. Sustained receiver loss needs the
instance to have been a standby in every sample, so a former primary just
after a switchover never qualifies. The Operator confirms its webhooks only
when every configuration was read. A Cluster's Backups tab offers restore to
a new cluster, and Storage names the standby each HA slot is kept for.
…ilures

A Cluster's Backups tab leads with a WAL archiving repair panel while
archiving fails: the operator's reason, the last archived and failed WAL,
where the destination is declared, the credential Secrets it references
(named, never read), where the archiver logs, then whether archiving resumed
and a base backup completed after it. Storage explains WAL an inactive slot
holds, names the standby, and how it is freed; a class that cannot expand
points to restoring into a larger cluster. A controller phase that is not
healthy links to the Operator view, which lists clusters stuck on a plugin
beside that plugin's readiness and restart history. The fleet's lag says how
many standbys it covers.
…he archiver's logs, and waits for a backup started after resume

The repair panel reads the ContinuousArchiving message instead of claiming
there is none, opens logs on the archiving container (plugin-barman-cloud
when the plugin is the WAL archiver), links to the Operator when the
Cluster's phase says a plugin blocks reconciliation, shows a workload
identity and the endpoint CA together, and counts only a completed Backup
that started after archiving resumed.

The fleet's lag coverage counts the standbys whose lag was read, and
expects every instance of a replica cluster.
…r trusts only a backup a restore can use

A blocked phase no longer claims the operator stopped and needs manual
intervention: it retries every one, the plugin phases clear once the plugin
loads and answers, and the banner and Overview quote status.phaseReason.
Unrecoverable keeps upstream's own wording.

The archiving checklist counts a Backup only when a restore can start from
it with this cluster's WAL archive (same archiver, or a volume snapshot),
and reads resumption from the ContinuousArchiving condition when the
instance manager's stats restarted with the primary.

A standby that is not streaming can still replay from the archive; copy
now says the slot does not advance without streaming instead of claiming
the standby cannot catch up.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit c5038c0. Configure here.

Comment thread web/src/components/cnpg/CNPGArchivingRepair.tsx Outdated
…at failed to archive

A Backup that started after archiving resumed can still begin earlier in
the WAL when it was taken from a lagging standby (CloudNativePG's default
target). The checklist now compares the Backup's own beginWal with the
last failed WAL, says when a backup begins before it, and stays unverified
when the failed WAL cannot be read.
@nadaverell nadaverell changed the title CloudNativePG operations, live diagnosis, history and a reusable GitOps write guard CloudNativePG: task-shaped cluster pages, operations, live diagnosis and repair paths Oct 4, 2026
…stance manager recorded

The ContinuousArchiving condition turning True also happens when archiving
is first set up, so on its own it no longer opens the resumed panel.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant