` variants.
@@ -172,6 +173,17 @@ Use CSS classes from `components.css` for status cells in table rows:
| `.border-r-subtle` | Right border |
| `.border-t-subtle` | Top border |
+### Unknown, partial and denied values
+Radar shows only what the cluster reports, and says where it came from. These rules apply to every surface that reads several sources at once (workspace integrations such as Capacity and CloudNativePG, multi-source detail pages):
+
+- **Unavailable ≠ zero.** A value Radar could not read (no access, not installed, not cached, the request failed) renders as unread, naming why — never as `0`, "none" or a healthy colour. A missing grant is named exactly (`GrantText`, `formatGrant`).
+- **Partial ≠ exact.** A count or total over data read only in part is a lower bound: `≥N` (`CertaintyGlyph`, `SidebarCategoryDestination.countLowerBound`), and a zero over partial data is unknown, not none.
+- **Unread is listed, not left out.** A summary line names what it could not read ("not read: Pods, zones"); a combined tone is never calmer than a part that was not read (`worseTone` ranks `unknown` above `healthy`).
+- **Recorded ≠ observed.** A value copied from a status field, an annotation or a declaration says so; a value matched to its subject by name rather than by identity says that too.
+- **Facts keep their rows.** `FactRow` renders the unread text in place; never hide a row because its value is missing (unlike `Property`, which hides empty values).
+
+The shared pieces live in `packages/k8s-ui/src/components/facts/` (facts, certainty, GitOps manager), `components/problems/` (problems with their sources), `components/ui/FoldSection.tsx` (section headings and folded sections) and `web/src/components/workspace/` (workspace screen layout and tables); each integration's own doc lists which source each value comes from and how it reads when unknown.
+
## 5. Layout Principles
### Spacing
diff --git a/Makefile b/Makefile
index eac7209c34..cf3625ee02 100644
--- a/Makefile
+++ b/Makefile
@@ -278,6 +278,16 @@ cnpg-demo-status:
cnpg-demo-live:
./scripts/cnpg-demo.sh live
+# Runtime fixtures on the EXISTING CNPG demo cluster (operator thawed if frozen):
+# MinIO + real plugin WAL archiving and backups, a restored cluster, a Pooler,
+# pgbench load, a blocked lock chain and a Prometheus Radar auto-discovers.
+# `./scripts/cnpg-demo.sh runtime-lag on|off` induces replica lag.
+cnpg-demo-runtime:
+ ./scripts/cnpg-demo.sh runtime
+
+cnpg-demo-runtime-down:
+ ./scripts/cnpg-demo.sh runtime-down
+
# Grafana Beyla on kind: eBPF loaded, a minimal Prometheus scraping it, and two
# conversations to observe. Which labels Beyla exports depends on configuration —
# dst_port and transport are off by default, direction is on and doubles every
@@ -523,6 +533,7 @@ help:
@echo " make rollouts-demo - Argo Rollouts progression fixtures"
@echo " make cnpg-demo - Frozen CNPG rendering fixtures"
@echo " make cnpg-demo-live - CNPG fixtures with the operator running"
+ @echo " make cnpg-demo-runtime - CNPG runtime: backups/restore, load, locks, Prometheus"
@echo " make velero-demo - Velero fixtures, all 13 backup phases at once"
@echo " make velero-demo-live - Velero with real object storage; states produced by the controller"
@echo " make beyla-demo - Grafana Beyla eBPF traffic fixtures"
diff --git a/README.md b/README.md
index e196a4cf18..09e6d0bb35 100644
--- a/README.md
+++ b/README.md
@@ -362,7 +362,7 @@ View TLS certificate details and expiry dates across all namespaces — catch ex
### GitOps
-Monitor, diagnose, and manage FluxCD and ArgoCD resources from a dedicated GitOps workspace.
+Monitor, diagnose, and manage FluxCD and ArgoCD resources in one place.
@@ -548,7 +548,7 @@ Upgrade impact also gets list-only access to CSIStorageCapacities, FlowSchemas,
| **Strimzi** | [KafkaConnector failure evidence](docs/integrations.md#strimzi-kafka-connectors) (connector/task status) |
| **Velero** | Backup, Restore, Schedule, BackupStorageLocation, VolumeSnapshotLocation |
| **External Secrets** | ExternalSecret, ClusterExternalSecret, SecretStore, ClusterSecretStore |
-| **CloudNativePG** | Cluster, Backup, ScheduledBackup, Pooler, Database, Publication, Subscription, ImageCatalog, ClusterImageCatalog, ObjectStore — plus a [workspace](docs/cnpg.md) for fleet, protection and declaration triage |
+| **CloudNativePG** | Cluster, Backup, ScheduledBackup, Pooler, Database, Publication, Subscription, ImageCatalog, ClusterImageCatalog, ObjectStore — plus [dedicated views](docs/cnpg.md) for fleet, protection and declaration triage |
| **Crossplane** | Managed Resources (any provider), Composite Resources, Claims, Provider, ProviderConfig, Function, Configuration, Composition, CompositionRevision, XRD |
| **Kyverno** | Policy, ClusterPolicy, PolicyReport, ClusterPolicyReport |
| **Sealed Secrets** | SealedSecret |
diff --git a/deploy/helm/radar/files/integration-read-baseline.yaml b/deploy/helm/radar/files/integration-read-baseline.yaml
index 1bd404b76e..6de59583c6 100644
--- a/deploy/helm/radar/files/integration-read-baseline.yaml
+++ b/deploy/helm/radar/files/integration-read-baseline.yaml
@@ -613,7 +613,7 @@ entries:
reason: "Includes spec.externalClusters[].connectionParameters; inline connection settings (potentially passwords) are intentionally visible, like Helm values."
source: "https://cloudnative-pg.io/docs/1.28/logical_replication/"
- group: "postgresql.cnpg.io"
- resources: ["backups","databases","imagecatalogs","poolers","publications","scheduledbackups","subscriptions"]
+ resources: ["backups","databaseroles","databases","failoverquorums","imagecatalogs","poolers","publications","scheduledbackups","subscriptions"]
scope: Namespaced
collection: "cloudnativePg"
decision: grant
diff --git a/deploy/helm/radar/templates/cloud-rbac-cluster-read.yaml b/deploy/helm/radar/templates/cloud-rbac-cluster-read.yaml
index 21b7a928e6..be7d8ed16a 100644
--- a/deploy/helm/radar/templates/cloud-rbac-cluster-read.yaml
+++ b/deploy/helm/radar/templates/cloud-rbac-cluster-read.yaml
@@ -12,6 +12,11 @@ bindings to the radar:* groups — deliberately NOT via `aggregate-to-view` labe
which would widen the shared built-in `view`/`edit` roles cluster-wide for every
other subject bound to them (mirrors the radar-helm add-on pattern).
+Leases are namespaced but read here like infrastructure: they carry only a
+holder identity and renew times (controller leader election, node heartbeats,
+CloudNativePG primary election). `edit` and `admin` already include them, so
+only the viewer tier gains anything.
+
No Secrets and no RBAC objects here — those stay gated behind the tier roles
(view excludes Secrets; RBAC visibility is rbac.viewRBAC) so this is a pure
infrastructure-read widening.
@@ -71,6 +76,9 @@ rules:
- apiGroups: ["scheduling.k8s.io"]
resources: ["priorityclasses"]
verbs: ["get", "list", "watch"]
+ - apiGroups: ["coordination.k8s.io"]
+ resources: ["leases"]
+ verbs: ["get", "list", "watch"]
- apiGroups: ["certificates.k8s.io"]
resources: ["clustertrustbundles"]
verbs: ["get", "list", "watch"]
diff --git a/deploy/helm/radar/tests/cloud_rbac_cluster_read_test.yaml b/deploy/helm/radar/tests/cloud_rbac_cluster_read_test.yaml
index f3d9decbe6..e558efe99b 100644
--- a/deploy/helm/radar/tests/cloud_rbac_cluster_read_test.yaml
+++ b/deploy/helm/radar/tests/cloud_rbac_cluster_read_test.yaml
@@ -43,6 +43,12 @@ tests:
apiGroups: ["scheduling.k8s.io"]
resources: ["priorityclasses"]
verbs: ["get", "list", "watch"]
+ - contains:
+ path: rules
+ content:
+ apiGroups: ["coordination.k8s.io"]
+ resources: ["leases"]
+ verbs: ["get", "list", "watch"]
- contains:
path: rules
content:
diff --git a/docs/CNPG_ARCHITECTURE_PLAN.md b/docs/CNPG_ARCHITECTURE_PLAN.md
new file mode 100644
index 0000000000..8fc8729d19
--- /dev/null
+++ b/docs/CNPG_ARCHITECTURE_PLAN.md
@@ -0,0 +1,154 @@
+# CNPG service and workspace host boundaries
+
+Status: implemented and validated. Independent subscription cross-review remains blocked by its weekly quota. The original plan follows the implementation record below.
+
+## Implementation record
+
+- `internal/cnpg` owns workspace/catalog assembly, runtime and memoization, storage, HA, recovery, SQL inspection, sessions, operator diagnosis, history, capabilities, action execution, report composition, log interpretation/stream lifecycle, and activity attribution. Operations accept contexts, parsed inputs and explicit dependencies. Runtime parsing, proxy transport, memoization and response types are separate modules.
+- `internal/server/cnpg_service.go` binds the caller's authorization, cache, discovery, typed/dynamic/proxy/exec clients and Prometheus client to one cluster snapshot. Read captures share a lock and refuse a context transition; superseded adapters withhold observations and permissions. Write client bundles recheck the reviewed context before execution. HTTP handlers retain decoding, transport framing, serialization and audit logging.
+- `internal/integration` contains the existing shared coverage/read/action contracts, conditional patching, cache-scope helpers and bounded fan-out. `internal/podlogs` contains the existing bounded log collector and Pod/container projections used by CNPG, workloads and JobSet. `internal/imageutil` provides the existing explicit image-tag parser to both application and CNPG observations.
+- `web/src/integrations` composes app-private resource hosts and workspace routes. Generic detail/drawer/list hosts consume it. Exact API-group identity, exclusive-slot ownership, additive slots and fixed hook ordering are enforced. CNPG and Capacity retain their distinct navigation and scope policies. Workspace screen imports are separate from route metadata to avoid cycles.
+- Service tests moved beside their implementation; HTTP contract tests remain in server. An AST boundary test rejects server/router imports, browser HTTP ports and global Kubernetes/Prometheus resolution in the service. Public Go/npm exports and endpoint contracts are unchanged.
+
+The uncertain sidebar data-provider and batched SSE invalidation seams remain documented in the integration guide, as allowed by the approved plan. They need another concrete consumer before choosing a shared shape.
+
+Validation passed: `make tsc`, shared UI type checks, the full app suite (2,295 tests), the shared UI suite (4,572 tests; one existing skip), `make test`, shared Go module tests and `make build` with frontend embedding. Live checks compared the rebuilt binary with the pre-extraction binary on the CNPG demo: 20 API read/refusal contracts, report ZIP contents/coverage, bounded logs/activity and 16 concurrent reads. An isolated browser exercised the workspace, Cluster detail, kind-list takeover and drawer query, and Capacity's cluster-wide posture without Karpenter; no page errors occurred. Visual screenshots were skipped because this extraction preserves rendered behavior. An existing desktop shell test timed out under load and passed both its isolated recheck and the final full suite without changes to its timeout. Independent plan and code review attempts were blocked by the subscription's weekly quota and produced no findings; this is an outstanding review gap.
+
+Extract CNPG orchestration into a request-independent service, then give application resource views a small, typed way to compose integration-owned extensions. Deliver these as several compiling changes. Each step must preserve the current endpoints, permission decisions, evidence coverage, navigation and action binding.
+
+## Evidence and confidence
+
+The source locations below describe the pre-extraction architecture at `8800398ae`; they are historical evidence for the approved plan.
+
+| Certain observation | Evidence |
+|---|---|
+| Pure CNPG interpretation is already separate from HTTP and permissions. | `pkg/cnpg/backup.go:1`, `pkg/cnpg/instances.go:1` |
+| Runtime, storage and operator have typed builders, but those builders still receive HTTP requests and use the server. | `internal/server/cnpg_runtime.go:473`, `internal/server/cnpg_storage.go:203`, `internal/server/cnpg_operator.go:108` |
+| The common cached read gate still depends on server permission and cache helpers. | `internal/server/cnpg_reads.go:21` |
+| Actions already have a useful context/client-based execution seam; capabilities still depend on HTTP/server methods. | `internal/server/cnpg_actions.go:296`, `internal/server/cnpg_actions.go:760`, `internal/server/cnpg_actions.go:1284` |
+| Shared read/action contracts and conditional patching are defined in the server package. Importing them from a new service would create a dependency cycle. | `internal/server/read_source.go:8`, `internal/server/kind_access.go:39`, `internal/server/actions.go:53`, `internal/server/actions.go:228` |
+| Runtime memoization separates context, caller identity and Pod UID, with bounded detached reads. | `internal/server/cnpg_runtime.go:805`, `internal/server/cnpg_runtime.go:815` |
+| Reports still retain both a server and an HTTP request while composing their reads. | `internal/server/cnpg_report.go:207` |
+| A service with a caller authorization adapter already exists elsewhere in Radar. Its permission policy is specific to upgrade scans. | `internal/upgrade/evidence.go:26`, `internal/server/upgrade_readiness_handler.go:101` |
+| Generic views import the CNPG adapter for navigation, detail slots and logs. | `web/src/App.tsx:2431`, `web/src/components/workload/WorkloadView.tsx:286`, `web/src/components/workload/WorkloadView.tsx:1209`, `web/src/components/workload/WorkloadView.tsx:1510`, `web/src/components/resources/ResourceDetailDrawer.tsx:36` |
+| Other integrations already extend resource detail through renderer wrappers. | `web/src/components/workload/WorkloadView.tsx:181`, `web/src/components/resources/renderers/KarpenterNodePoolRenderer.tsx:1`, `web/src/components/resources/renderers/JobAdmissionRenderers.tsx:1` |
+| Workspace routes and context-switch handling have integration-specific branches in App. | `web/src/App.tsx:288`, `web/src/App.tsx:1254` |
+| CNPG's resource-list replacement controls both fetching and rendering. Its sidebar also has an integration-specific data hook. | `web/src/components/resources/ResourcesView.tsx:286`, `web/src/components/resources/ResourcesView.tsx:392`, `web/src/components/resources/ResourcesView.tsx:114` |
+
+**Confident recommendations:** the backend dependency direction and the existing resource-detail extension slots are clear enough to implement. Workspace route ownership also has two concrete consumers, CNPG and Capacity. Keep those consumers' existing route parsers and scope policies.
+
+**Defer until there is another concrete consumer:** category-workspace sidebar aggregation, operation tracking, operator diagnosis, and a general-purpose proxy/report framework. Document the places a second integration must revisit. No user decision is required to draft the plan; a new route takeover or public plugin API would be a separate decision.
+
+**Needs-input:** none for the proposed direction. The user authorized implementation; it preserves current product behavior and public interfaces.
+
+## Backend ownership
+
+```mermaid
+flowchart TD
+ Server[HTTP handlers and caller adapters] --> Service[internal/cnpg service]
+ Server --> Existing[Existing authorization, cache and metrics helpers]
+ Service --> Contracts[internal/integration contracts]
+ Service --> Domain[pkg/cnpg pure rules]
+ Server --> Contracts
+ Contracts --> Grant[internal/auth Grant]
+```
+
+`internal/cnpg` owns runtime, storage, HA, recovery, history, operator assessment, capabilities, action execution and report composition. It receives `context.Context`, resource identities, parsed options and caller-scoped dependencies. It does not receive browser HTTP requests or response writers, import `internal/server` or chi, or reach for the process-wide Kubernetes singleton. Pod/API HTTP transport can still use `net/http` internally.
+
+`internal/server` owns routes, bounded decoding, response serialization, browser streaming, HTTP error mapping, audit logging and adapting the authenticated caller to existing permission/cache/client helpers. The service's operations use authorized read ports so authorization is also applied when a report calls another service method; endpoint-only gates are insufficient.
+
+`pkg/cnpg` remains the shared pure domain package used by Audit, Issues and the service. Cached-object finding production remains in Issues. Start the orchestration service in `internal/`: its current callers are Radar handlers and report composition. Keep its ports portable, and move a service surface to the public Go module only when a supported consumer actually needs that surface.
+
+The extraction reduces server coupling; it will not remove CNPG's genuine complexity. Runtime, recovery and destructive actions keep separate modules and source-specific result types. Prefer several small operation interfaces over a shared workspace engine with interchangeable permission or certainty policies.
+
+### Dependency seams
+
+Create narrow dependencies alongside the operation that consumes them. Start with runtime's existing Cluster read, instance-Pod inventory and fixed-path proxy read; add storage and operator dependencies as those operations move. Avoid one interface containing every Kubernetes operation.
+
+| Dependency | Responsibility |
+|---|---|
+| Authorized observations | Return the object/list and its coverage or read-source outcome. Preserve unread, incomplete, missing and empty as distinct results. Delegate permission and cache-scope decisions to existing helpers. |
+| Proxy reader | Caller-scoped client, current fixed path/port allowlist, bounded response, scheme handling and error classification. No arbitrary proxy request API. |
+| Exec function | Reuse the current context/namespace/Pod/container/argv/stdin seam and output cap. Preserve fixed SQL and psql-variable binding. |
+| Metrics reader | Adapt existing Prometheus queries and identity-isolation results. Carry the same ambiguity and partial-coverage metadata. |
+| Action clients | Caller-impersonated dynamic/typed clients and exec, using the existing execution seam. Write preflights must read fresh facts directly, not through the observation cache. |
+| Clock and memo identity | Explicit time and captured cluster/caller identity. Preserve existing key components, TTLs, timeout behavior and result timestamps. |
+
+Read adapters may capture the server and request internally; their service-facing methods do not expose either. Capture the selected cluster/cache/client identity when preparing an operation rather than resolving a different global context halfway through it. Long-lived service state holds memoization, not a previous caller's clients or grants.
+
+Each operation receives only its dependencies: runtime observations and proxy reads for runtime, additional volume/metrics reads for storage, and fresh write clients for actions. CNPG request/response models move with their slice into `internal/cnpg`, retaining JSON tags and optional fields. The shared leaf package contains existing cross-integration contracts, not CNPG runtime models or orchestration.
+
+Authorizers preserve allowed, denied and non-authoritative permission checks. Preserve the distinction between ordinary cached permission checks and subresource checks. Required grants must be checked before returning a memoized answer. Keep per-kind namespace coverage, including nil versus empty namespace scope and the rules about disclosing denied namespaces. Retain the existing action behavior: an unknown permission check can leave a capability offered, while the impersonated apiserver write is authoritative (`internal/server/actions.go:335`).
+
+### Implementation sequence
+
+1. **Remove the dependency-cycle blockers.** Add `internal/integration/{reads,actions,patch}.go` for the existing `ReadSource`, `KindCoverage`, action envelope/capability, pure capability calculation and version-bound patch helper. Reuse `internal/auth.Grant`. Keep HTTP body decoding, connection handling, authorization adapters and serialization in server. Give service read failures explicit categories such as disconnected, syncing, denied and missing; keep the existing action refusal fields and test their current status/message/code mapping. Do not change wire output as part of the extraction. Temporary internal type aliases can keep intermediate commits compiling and are removed as callers migrate.
+
+2. **Extract runtime as the first complete slice.** Add `internal/cnpg/{ports,types,runtime,runtime_decode,proxy,memo}.go` and `internal/server/cnpg_service.go`. Move runtime parsing, target assessment, bounded proxy reads and memoization from `cnpg_runtime.go`, plus the transport classification it needs from `cnpg_errors.go`. Convert Cluster and Pooler runtime handlers to thin adapters. Reports call the same service operation. Prove this slice works with fake ports and no HTTP test request or live singleton before extending the design.
+
+3. **Move related read orchestration.** Extract storage and HA next, reusing runtime observations. Then move recovery/restore validation, inspection/sessions, history/fleet metrics and operator/status/diagnosis. Source files are the corresponding `internal/server/cnpg_{storage,cluster_ha,recovery,inspect,sessions,history,operator,operator_status,operator_diagnosis}.go`. Move catalog reverse-lookup orchestration from `cnpg_handlers.go`, schedule preview from `cnpg_schedule.go`, and fleet disk assessment from `cnpg_storage.go`. Keep source-specific error/coverage vocabularies. Keep Prometheus query engines in their existing package. Assembly from `cnpg_workspace.go` receives already-authorized per-kind observations and canonical Issues results; it does not become another finding producer.
+
+4. **Move capabilities and actions together.** Extract `cnpg_actions.go`, `cnpg_actions_maintenance.go`, `cnpg_destroy.go` and `cnpg_pooler_actions.go` into service action modules. Preserve the existing per-action binding table and current client injection, rather than introducing a new workflow engine. Capability calculation and submit-time guards use the same facts and blockers. Keep context binding, fresh UID/resourceVersion checks, Pod/backend/PVC identity checks, exact write shapes, preflight order, no automatic retries, partial results and ambiguous outcomes.
+
+5. **Finish reports, logs and activity.** Replace `cnpgReportBuilder`'s server/request fields with operation-scoped reads and the extracted service. Put archive construction, redaction, source selection and CNPG log interpretation in the service; HTTP headers and stream framing stay in server. CNPG logs already use shared workload-log entries/sources from `workload_logs.go`; move the smallest shared protocol/helper subset to an app-level log package if the extraction needs it, after checking existing helpers. Do not copy the merged-log engine or move unrelated workload behavior into CNPG. Preserve interval/restart cursors, cancellation, byte bounds and report error classification.
+
+Each slice retains handler integration tests. Move pure parsers/guards and service-level tests beside their implementation; do not export private helpers merely to keep tests in server. Remove migrated wrappers and aliases before declaring the extraction complete.
+
+## Frontend generalization
+
+### Resource detail and kind-list extension points
+
+Add an app-owned composition root under `web/src/integrations/` with `resourceHosts.tsx` and a small internal contract. It imports integration adapters; generic resource views import this composition root. Keep it private to radar-app and reuse existing k8s-ui `RendererOverrides`, summary/header-action slots and extra-tab types.
+
+The first entries are CNPG and the existing renderer contributions, including Karpenter and batch/Ray. This changes where their existing behavior is composed. It does not introduce automatic Capacity redirects or new tabs.
+
+| Existing host responsibility | Planned ownership |
+|---|---|
+| Drawer expansion and generic-detail redirect | Resolve the owning adapter by exact resource identity; delegate path/query construction to it. |
+| Summary, header actions, extra tabs and Diagnose decoration | Typed contributions using existing detail slots. |
+| Integration log view | Optional contribution alongside the existing built-in log strategies, preserving their selection order. |
+| Drawer navigation/trail | Optional adapter-owned component. |
+| Renderer wrappers | One stable composed `RendererOverrides` object in the app composition root. Existing dispatch still handles curated kinds, collisions and version/spec-shape predicates. |
+| Resources kind-list replacement | Separate optional kind-list contribution keyed by API group and resource plural, with a pure `table / wait / view` decision used for both query enabling and rendering. |
+
+Resolve identity through existing group/kind/navigation helpers. A selection uses resource plurals while fetched objects use Kinds; normalize at the boundary. Core group `""` and an unknown group remain distinct. A colliding Kind without a resolved group cannot select an integration-specific destination.
+
+An object may have contributions from several integrations: a core Job can have Kueue admission and batch execution. Additive actions/tabs compose in a declared order. Summary, renderer override, logs, detail destination and kind-list takeover each have an explicit owner; incompatible contributions are rejected in validation/tests rather than resolved by registration order. Existing wrapper components may compose related contributions internally.
+
+Use stable React components for data-bearing contributions. Never call a selected adapter's hook conditionally inside a generic host. Capability loading and support decisions use the existing feature-gating machinery; adapters retain their current supported/unknown/unsupported policies. Registration does not bypass the guards on fetch hooks.
+
+Migrate `App.tsx` expansion, `WorkloadView.tsx` redirect/slots/logs/Diagnose, `ResourceDetailDrawer.tsx` navigation and `ResourcesView.tsx` kind-list ownership. Those host operations should no longer import CNPG implementations directly. The sidebar data hook is a documented exception until its provider composition has another consumer. Renderer dispatch and standard workload actions remain shared through their current interfaces.
+
+### Workspace route ownership
+
+CNPG and Capacity already have separate route parsers (`web/src/components/cnpg/routes.ts` and `web/src/components/capacity/shared.tsx`, consumed by `web/src/components/capacity/CapacityView.tsx:9`). Add a small app-owned workspace table for those two consumers: route-prefix ownership, label/breadcrumb contribution, sidebar ownership, screen component and optional context-switch search policy. Move the corresponding App branches behind that table while retaining each adapter's parser and view logic.
+
+Keep CNPG's resource-category placement and namespace filter, and Capacity's top-level placement and cluster-wide scope. Preserve the current CNPG `ctx` pin on a context switch; do not apply that policy to every workspace. Keep connection/feature gates, navigation customization, route state, basename handling and the radar-app embedding API unchanged. Routes and screens stay explicitly registered in the application; no public plugin loader or universal workspace data model is needed.
+
+Implement the frontend in four compiling steps: move the existing stable renderer-override composition first; migrate detail/navigation/log contributions next; move kind-list ownership together with its query guard; then extract workspace route metadata into a separate module. Keep workspace screen imports separate from the resource-host root so screens can use the app's WorkloadView without creating a composition-root cycle. No integration adapter imports App.
+
+### Document extension points whose shape is still uncertain
+
+Document these in the integration guide now. Add a source comment only when implementing a seam with a non-obvious invariant; avoid repeated TODOs explaining obvious imports.
+
+| Place to revisit | Trigger and constraint |
+|---|---|
+| `ResourcesView.tsx:114` and `sidebarCategoryWorkspaces` | A second category workspace needs aggregate counts/coverage. Reuse the existing sidebar prop; mount data hooks in integration-owned components and preserve per-kind unread scopes. Capacity's global workspace is not evidence for a common namespace-scoped sidebar provider. |
+| `App.tsx:1117` and `App.tsx:1162` | Another integration needs event-driven workspace-query invalidation. Extract change-to-query-key contributions into the existing batching mechanism; preserve its timing and context lifetime, rather than adding another SSE listener. |
+| CNPG operation tracker, proxy protocol, report options and operator diagnosis | A second integration needs the same semantics. Extract the demonstrated common part; resource-specific evidence, permission policies and outcomes remain with that integration. |
+
+## Validation and completion criteria
+
+Backend tests must exercise exact permissions/subresources, caller/context/Pod-UID memo separation, stale or partial cache coverage, disconnected/syncing states, proxy/exec bounds and cancellation, report redaction, and action fact binding/partial writes. Preserve existing HTTP response shapes, including null/empty/omitted fields, statuses, refusal codes and source timestamps. Use focused golden fixtures for meaningful wire contracts, not snapshots of every implementation detail.
+
+Frontend tests must cover CNPG/CAPI/core collisions, unknown group, JobSet version dispatch, simultaneous batch/Kueue contributions, all feature-support states, query suppression for kind-list takeover, tab/query/context preservation, drawer history, default log fallback and duplicate exclusive-slot ownership. Exercise OSS and embedded basename navigation using current public interfaces.
+
+For each code slice run the affected tests and required QA; complete the refactor with `make tsc`, shared-package type checks and frontend suites, `make test`, the separate `pkg` Go suite where affected, and `make build`. Consider a focused real CNPG smoke pass for runtime/report/permission behavior and navigation once the demo is available. Visual capture is useful only for any slice that changes rendered behavior.
+
+The final architecture should allow service tests without constructing an HTTP request or server, retain one implementation of each shared CNPG rule and finding, and allow an integration to extend resource views by adding its adapter/registration rather than editing several generic views. Context/caller adaptation and wire serialization remain centralized in server; package dependency direction is enforced by an import-boundary check.
+
+## Original plan review
+
+The draft has been checked against the current call sites for dependency cycles, permission/memo leakage, fresh-write preflights, competing detail contributions, React hook ordering and workspace scope differences. Source anchors and the integration-guide link were verified, and `git diff --check` passed. Visual testing is skipped because this change contains only documentation.
+
+Independent plan review was attempted on the Claude subscription on 2026-10-06 and blocked by its weekly usage limit; no independent findings were produced. The user subsequently authorized implementation. The implementation record above tracks its current state.
diff --git a/docs/INTEGRATION_GUIDE.md b/docs/INTEGRATION_GUIDE.md
index ad993ddb76..47a337844f 100644
--- a/docs/INTEGRATION_GUIDE.md
+++ b/docs/INTEGRATION_GUIDE.md
@@ -126,7 +126,164 @@ shared renderers.
- `packages/k8s-ui/src/components/topology/`: `TopologyFilterSidebar.tsx`,
`K8sResourceNode.tsx`, `layout.ts`, `topology.css`.
-## 3. Before submitting
+## 3. Workspace integrations
+
+A workspace is a set of screens for several related kinds (CloudNativePG at
+`/cnpg`, Karpenter at `/capacity`). Build one only when the operator question
+spans objects: fleet health, a chain such as Backup → ObjectStore → Cluster, or
+actions that need facts from more than one object. With fewer kinds, or no
+cross-object question, a renderer plus Issues plus the detail slots is enough.
+The batch kinds, for example, use only the detail seams.
+
+The pieces below are shared and should be imported, not copied. Read
+[DESIGN.md](../DESIGN.md#unknown-partial-and-denied-values) for the rules on
+unknown, partial and denied values first: every piece here exists to keep them.
+
+- [ ] **Placement.** Under Resources as a category workspace
+ (`ResourcesSidebar` `categoryWorkspaces`, keyed by the category name from
+ `api-resources.ts`), or a top-level page when the subject is cluster-wide.
+ Say whether the namespace filter applies, and why.
+- [ ] **Routes and navigation.** Use `/x`, `/x/` and
+ `/x///?ctx=`.
+ - The in-drawer trail is `?drawer=`, encoded by `web/src/utils/drawer-trail.ts`.
+ - Back labels come from `currentPageLabel` and subject-filtered Issues links
+ from `issuesPathForSubject` (both in `web/src/utils/page-links.ts`).
+ - Pin `ctx` on a detail page. After a context switch, say the object is not in
+ this context; never open a same-named object from another cluster.
+- [ ] **One aggregate endpoint.**
+ - List each kind with `readWorkspaceKind` (dynamic kinds) or
+ `typedKindScope` (typed kinds such as Pods), both in
+ `internal/server/kind_access.go`. Their answer's `Coverage()` is a
+ `internal/integration.KindCoverage`: `full|partial|denied|notInstalled|syncing|uncached|error`,
+ with denied and uncached namespaces named only when the caller supplied the
+ namespace list.
+ - Radar's own cache scope comes from `integration.NamespacesWithinCache`
+ (`internal/integration/cache_scope.go`).
+ - A per-object read with several sources reports a `ReadSource`
+ `{state, grant, reason}` per source (`internal/integration/reads.go`), with the missing
+ permission as a `Grant` (`internal/auth/grant.go`), never a sentence.
+ - Fan out over namespaces with `integration.FanOut` (`internal/integration/fanout.go`), behind a cap.
+ - Prometheus series matched to an object by name rather than identity carry a
+ `SeriesIsolation` (`internal/prometheus/series_scope.go`).
+ - A GitOps or Helm manager comes from `topology.ManagedByFromMeta`, never from
+ labels read on the client.
+- [ ] **Version skew.** Give the workspace's endpoints one `FeatureCapabilities`
+ flag and a `radarFeatures.ts` entry with `flagShippedWithEndpoint: true`.
+ - Gate every hook with `useRadarFeature`, including mutations, streams and
+ downloads. Never add `retry` to a mutation.
+ - When the Radar is too old, hide the sidebar destinations and fall back to
+ the standard detail views.
+- [ ] **Guided multi-resource setup.** Keep integration-specific target models and workflow state in the app. CNPG reuses its target fields between Create and Restore, its schedule inputs between Edit and Create, and the shared strict-create/YAML review for both. Put preflight and reviewed conditional writes in the integration service, with thin HTTP adapters and a new feature flag when the endpoint/verb is new. Each independent write needs its own review and observed follow-through; saving declarations is not end-to-end success. Do not grow a universal provisioning registry around one integration. A future second concrete setup flow should extract only the seam it actually shares.
+- [ ] **Domain rules.** Put interpretation reused by checks, findings and actions in a pure integration package (CNPG uses `pkg/cnpg`), independent of HTTP and caller permissions. Keep matching frontend derivations together; share fixture cases for rules represented in both languages. Put request-independent orchestration in an integration service (`internal/cnpg` is the current example), with caller-scoped observations/clients supplied by server adapters. Share typed service reads between endpoints and reports rather than invoking handlers through a response recorder. Keep HTTP requests, routing and process-wide client resolution outside the service.
+- [ ] **Findings.** The Issues engine owns findings from cached Kubernetes objects. Cross-resource findings retain their inventory operations in `Issue.RequiredReads`; hosts supply `CanReadEvidence` for both composition and cached related-issue projections. A live proxy or Prometheus measurement stays in the workspace as a `WorkspaceProblem` with `source: 'measurement'`, `measuredBy` and, when matched by name only, `unverifiedMatch`. Add no new severity ladder, and title reasons the Issues page already titles with `issueReasonTitle`. Carry semantic reasons/states through presentation; never branch on generated IDs or display text.
+- [ ] **Screens.**
+ - k8s-ui `components/facts`: `Fact`, `FactGrid`, `FactRow`, `FactValue`,
+ `FactSource`, `CertaintyGlyph` and `ManagedByText`. These are for any
+ surface that shows observed values, single-kind renderers included.
+ - k8s-ui `components/problems`: `WorkspaceProblem`, and
+ `ProblemCallout`/`ProblemList`/`ProblemMeta` with the workspace's
+ `rootKind`, plus `OpenIssueContext`.
+ - Also from k8s-ui: `SectionHeading`, `FoldSection` and `FoldSummary`
+ (`ui/FoldSection`), `ui/RefLink`, `toneTextClass`/`worseTone` in
+ `ui/status-tone`, and `formatGrant`.
+ - App: `web/src/components/workspace`:
+ - layout: `ScreenBody`, `ScreenEmptyState`, `Notice`
+ - controls: `Segments`, `FilterChips`
+ - tables: `SectionTable` and its table classes
+ - text: `RefreshFailedNotice` and `GrantText`
+ - Buttons are `.btn-brand` and `.btn-secondary`.
+- [ ] **Detail page.** Through `WorkloadView`:
+ - `renderSummary`: a composed Overview; the resource's renderer moves to
+ "Spec & status".
+ - `extraTabs`
+ - `renderHeaderActions`
+ - Keep an integration's composition in its own host adapter (`web/src/components/cnpg/host.tsx`) and register it in `web/src/integrations/resourceHosts.tsx`, so generic views use one interface for routing, summary, header actions and logs. Fetching stays in the app; shared presentation receives facts and callbacks.
+- [ ] **Actions.**
+ - Shared contracts/guards (`internal/integration/actions.go`); caller adaptation and HTTP decoding (`internal/server/actions.go`):
+ - A capabilities endpoint answers each action as an `ActionCapability`
+ `{allowed, reason, reasonCode?, permission, grant}`, built with `grantPermission` and
+ `integration.CapabilityVerdict`.
+ - The POST body is an `ActionRequest` `{reviewedContext, uid, facts, params}`
+ read with `decodeActionRequest`.
+ - Bind the facts the user reviewed. Refuse with 409 `changed`
+ (`integration.ChangedAction`) or `context_changed`, and use `integration.PartialAction` when a
+ multi-step write stops part-way.
+ - Writes are impersonated, version-bound (`integration.MergePatchAtVersion`) and never
+ retried.
+ - Client:
+ - `web/src/api/actions.ts`: `actionErrorCode`, widened with the integration's
+ own codes; `actionOutcomeLocked`, `actionCompleted` and `capabilityReason`.
+ - `ActionConfirmDialog` and the GitOps write guard (`useGitOpsWriteGuard`).
+ Multi-step setup can supply `onBack`/`backLabel` without giving Cancel a
+ second meaning. For YAML creation, `CreateResourceDialog.onBack(yaml)`
+ returns the edited draft to the parent; the parent retains unrepresented
+ fields and owns any explicit replacement. `onCreated(result, submittedYaml)`
+ reports the actual write so follow-through never assumes the earlier form
+ still describes edited YAML. Keep domain forms and step state in the
+ integration host; these presentation contracts contain no integration logic.
+ - An accepted POST is not a completed action: follow the outcome in status.
+- [ ] **Docs and fixtures.** Add a `docs/.md` listing which source each value
+ comes from and how it reads when unknown, a `scripts/-demo.sh` with its
+ README, and a CLAUDE.md row.
+
+### Host ownership and future extension points
+
+`web/src/integrations/resourceHosts.tsx` is the app-owned composition root for
+integration adapters. Generic resource views consume its typed routing, detail,
+log, renderer and kind-list operations. CNPG's adapter remains in
+`web/src/components/cnpg/host.tsx`; existing Karpenter, batch/admission, Ray and
+other renderer wrappers register in the same root. This interface is private to
+the app; it does not change the public radar-app embedding API.
+
+Add an adapter and registration when extending these existing slots. Reuse
+`RendererOverrides` and the detail props from k8s-ui. Resources are keyed by
+exact API group and plural; an unresolved group cannot select a colliding CRD.
+Declare exclusive ownership of summary, destination and logs. Competing
+renderer, kind-list and exclusive-slot owners are rejected. Header actions and
+extra tabs compose in order, with duplicate tab IDs rejected. Data-bearing
+contributions are mounted components whose hooks remain integration-owned.
+
+Workspace metadata lives in `workspaceRoutes.ts`, while `workspaceScreens.tsx`
+registers screens separately so they can use `WorkloadView` without a circular
+import. Register metadata and a screen for another workspace; retain its parser,
+placement, namespace policy and context-switch policy in its own adapter. The
+[architecture plan](CNPG_ARCHITECTURE_PLAN.md) records the extraction rationale.
+
+When adding another integration, check these specific places before extending
+generic hosts:
+
+| Extension | Current location | Rule to preserve |
+|---|---|---|
+| Detail summary, header actions, Diagnose, logs and renderer wrappers | `web/src/integrations/resourceHosts.tsx` and integration host adapters | Reuse the existing detail slots and `RendererOverrides`. Mount integration-owned components for their hooks. Match exact group and Kind; keep version/spec-shape dispatch where it already lives. |
+| Drawer expansion and navigation trail | `resourceHosts.tsx` destination/drawer contributions | Keep context, tab/query state, history and API-group collisions. |
+| A kind list replaced by a workspace view | `resourceHosts.tsx` kind-list contributions, consumed by `ResourcesView.tsx` | One ownership decision must control both fetching and rendering, including capability loading and unsupported Radar fallback. |
+| Workspace routes, labels and context-switch policy | `web/src/integrations/workspaceRoutes.ts` and `workspaceScreens.tsx`; integration route helpers | CNPG and Capacity share registration metadata. Their route parsing, placement and namespace scope remain integration-owned. |
+| Category-workspace counts and coverage | `ResourcesView.tsx`'s `useCNPGSidebarWorkspace` and `sidebarCategoryWorkspaces` | Generalize the data-provider composition when another category workspace needs it. Reuse the current sidebar props and retain partial/denied coverage. |
+| Event-driven workspace invalidation | `App.tsx`'s CNPG pending flag and batched invalidation | Generalize change-to-query-key contributions when another workspace needs them; reuse the existing batches and connection lifetime. |
+
+A resource can have contributions from multiple integrations, such as batch
+execution and Kueue admission on a Job. Compose additive slots explicitly; give
+summary, renderer, logs, destination and kind-list takeover an explicit owner.
+Registration order must not silently pick between conflicting owners. This is
+an internal app interface, not a new public radar-app plugin API.
+
+The sidebar-provider and event-invalidation shapes above are deliberately left
+open until another matching consumer appears. The listed locations are the
+places to revisit, rather than another set of integration-specific branches to
+copy into generic views.
+
+These other pieces are not shared yet, because they have one integration
+consumer and the second should shape them:
+- the operation tracker
+- the fixed-path `pods/proxy` reader
+- the merged log stream
+- the report bundle
+- operator diagnosis
+
+Read CloudNativePG's versions (`web/src/components/cnpg/`,
+`internal/cnpg/`), and extract the demonstrated common part when a second integration needs it. The merged log engine already lives in `internal/server/workload_logs.go`; reuse its protocol instead of duplicating it.
+
+## 4. Before submitting
- [ ] Verify status against the controller's documented API: desired versus
observed, unknown versus false, intentional pause/stop versus failure. Consider
diff --git a/docs/STRUCTURE.md b/docs/STRUCTURE.md
index d13583f487..9995089a7d 100644
--- a/docs/STRUCTURE.md
+++ b/docs/STRUCTURE.md
@@ -10,6 +10,7 @@ radar/
├── internal/
│ ├── app/ # Application lifecycle management
│ ├── audit/ # Radar-specific audit runner (cache → pkg/audit bridge)
+│ ├── cnpg/ # CNPG service: caller-scoped reads, actions, reports and log/activity interpretation
│ ├── config/ # Configuration management
│ ├── errorlog/ # Error logging utilities
│ ├── helm/ # Helm client integration
@@ -17,6 +18,8 @@ radar/
│ │ ├── handlers.go # HTTP handlers for Helm operations
│ │ ├── hook_evidence.go # Live Job/Pod/Event/log evidence for failed hooks
│ │ └── types.go # Helm release types
+│ ├── imageutil/ # Explicit image-tag parsing shared by app and integration observations
+│ ├── integration/ # Shared read coverage, action contracts, scope helpers and bounded fan-out
│ ├── images/ # Container image analysis
│ │ ├── auth.go # Registry authentication (pull secrets, ECR, GCR, ACR)
│ │ ├── handlers.go # HTTP handlers for image inspection
@@ -49,6 +52,7 @@ radar/
│ │ ├── tools_workloads.go # Workload-specific MCP tools
│ │ └── resources.go # MCP resource definitions
│ ├── opencost/ # OpenCost integration (cost analysis)
+│ ├── podlogs/ # Bounded log snapshots and Pod/container projections shared by workload integrations
│ ├── prometheus/ # Prometheus client integration
│ ├── server/
│ │ ├── server.go # chi router, main REST endpoints (SOURCE OF TRUTH for routes)
@@ -76,6 +80,7 @@ radar/
├── pkg/
│ ├── ai/context/ # AI context minification for LLM-friendly output
│ ├── audit/ # Shared cluster audit check engine (reusable by skyhook-connector)
+│ ├── cnpg/ # Pure CNPG declarations, identity, fencing and cron rules
│ ├── gitops/
│ │ ├── insights/ # Per-app diagnosis pipeline: issues + drift diff + recent events + plan + history
│ │ └── tree/ # GitOps resource tree builder for ArgoCD/FluxCD detail graphs
@@ -97,6 +102,9 @@ radar/
│ │ ├── shared/ # ResourceRendererDispatch, ResourceActionsBar, EditableYamlView
│ │ ├── gitops/ # Argo/Flux badges + actions + tree graph + insights views
│ │ ├── workload/ # WorkloadView
+│ │ ├── facts/ # Observed values with their sources: facts, certainty glyph, GitOps manager
+│ │ ├── problems/ # Problems with where their evidence came from (callout, list, origin)
+│ │ ├── cnpg/ # CloudNativePG workspace model + composed summaries
│ │ ├── timeline/ # Timeline shared components
│ │ ├── logs/ # Log viewer core
│ │ └── ui/ # Shared primitives, package-owned Monaco/YAML runtime, Problems + review
@@ -115,6 +123,8 @@ radar/
│ │ │ ├── portforward/ # Port forward manager
│ │ │ ├── resource/ # Single resource detail page
│ │ │ ├── resources/ # Resource list panels (thin wrappers over @skyhook-io/k8s-ui)
+│ │ │ ├── workspace/ # Workspace screen layout, tables and notices (Capacity, CloudNativePG)
+│ │ │ ├── cnpg/ # CloudNativePG workspace screens, actions, runtime
│ │ │ ├── audit/ # Cluster audit detail view
│ │ │ ├── cost/ # Cost tracking and visualization
│ │ │ ├── settings/ # Settings dialog
@@ -127,6 +137,7 @@ radar/
│ │ ├── context/ # React contexts (connection, theme, context-switch)
│ │ ├── contexts/ # React contexts (capabilities)
│ │ ├── hooks/ # Custom React hooks
+│ │ ├── integrations/ # App-private resource hosts and workspace route/screen composition
│ │ └── utils/ # Topology and utility helpers
│ └── package.json
├── deploy/ # Docker, Helm, Krew configs
diff --git a/docs/authentication.md b/docs/authentication.md
index 4f86c4da3a..73a31f2241 100644
--- a/docs/authentication.md
+++ b/docs/authentication.md
@@ -231,7 +231,7 @@ Under cloud-mode (`RADAR_CLOUD_MODE=true`, set automatically by the chart when `
- Forces `--auth-mode=proxy` with pinned `X-Forwarded-User` / `X-Forwarded-Groups` headers. Radar accepts those headers only on requests marked in-process by its authenticated Cloud tunnel; the ordinary pod TCP listener cannot assert Cloud identity.
- Ships three default ClusterRoleBindings mapping Cloud's `radar:owner` / `radar:member` / `radar:viewer` groups to the standard K8s `admin` / `edit` / `view` ClusterRoles. Configurable via `cloud.defaultRbac.*` in `values.yaml`.
-- Adds a **cluster-read add-on** (`cloud.defaultRbac.clusterScopedRead.{viewer,member,owner}`, each default on) granting `get/list/watch` on infrastructure the built-in `view`/`edit`/`admin` roles exclude — Nodes, PersistentVolumes, StorageClasses, IngressClasses, PriorityClasses, RuntimeClasses, ClusterTrustBundles, CRDs, and admission webhook configurations — plus list-only Upgrade impact evidence for CSIDrivers, CSIStorageCapacities, legacy PodSecurityPolicies, and API flow-control configuration. When the matching collection is enabled, it also grants APIServices, PrometheusRules, Karpenter kinds, and node metrics. It does not grant Secrets, RBAC objects, API-server metrics, or kubelet proxy access. It's an independent axis per tier: set a tier `false` to make it namespaced-only (e.g. `clusterScopedRead.viewer: false`). Owner node cordon/drain is a separate cluster-scoped *write*, off by default (`cloud.defaultRbac.nodeOps`).
+- Adds a **cluster-read add-on** (`cloud.defaultRbac.clusterScopedRead.{viewer,member,owner}`, each default on) granting `get/list/watch` on infrastructure the built-in `view`/`edit`/`admin` roles exclude — Nodes, PersistentVolumes, StorageClasses, IngressClasses, PriorityClasses, RuntimeClasses, Leases (namespaced; holder and renew times only — only viewers gain, since `edit`/`admin` include them), ClusterTrustBundles, CRDs, and admission webhook configurations — plus list-only Upgrade impact evidence for CSIDrivers, CSIStorageCapacities, legacy PodSecurityPolicies, and API flow-control configuration. When the matching collection is enabled, it also grants APIServices, PrometheusRules, Karpenter kinds, and node metrics. It does not grant Secrets, RBAC objects, API-server metrics, or kubelet proxy access. It's an independent axis per tier: set a tier `false` to make it namespaced-only (e.g. `clusterScopedRead.viewer: false`). Owner node cordon/drain is a separate cluster-scoped *write*, off by default (`cloud.defaultRbac.nodeOps`).
- Creates the default read grant for `radar:system`, Radar Cloud's alerts worker and timeline puller: a chart-owned role aggregated to match `view`, the cluster-read and integration-read add-ons, and cluster-wide Secret read. Helm release alerts and Secret changes in the timeline need Secret read; Helm stores releases as Secrets. No writes. Controlled by `cloud.systemRbac` (default `true`).
- Creates the default read grant for alert-triggered AI investigations, recorded in Kubernetes audit logs as `radar:ai:`: `radar:ai:reader` receives permissions based on `view` plus the cluster-read and integration-read add-ons. Stock `view` excludes Secrets, but includes pod logs and ConfigMaps; roles labelled `aggregate-to-view` widen this grant too. Controlled by `cloud.aiRbac` (default `true`). Requests a person starts or follows up on use that person's identity and groups, including follow-ups on alert-triggered investigations. MCP clients also use the user's permissions.
- Treats `cloud.aiRbac` and `cloud.systemRbac` as controls for the chart's default grants, independent of `cloud.defaultRbac`. False or absent (including after a plain `--reuse-values` upgrade from an older chart) creates no default grant and leaves customer-created bindings untouched. These settings do not disable background jobs or automatic analysis; they control their default Kubernetes access.
diff --git a/docs/capacity.md b/docs/capacity.md
index c2fb9e6802..ba3610a256 100644
--- a/docs/capacity.md
+++ b/docs/capacity.md
@@ -64,7 +64,7 @@ Operators rarely start at the nav. Capacity meets them where they are:
## Reading the numbers
-Capacity's core contract is **per-value certainty**. Every quantity carries one of:
+Capacity applies Radar's rules for unknown and partial values ([DESIGN.md](../DESIGN.md#unknown-partial-and-denied-values)) per quantity: its core contract is **per-value certainty**. Every quantity carries one of:
| Glyph | Meaning |
|-------|---------|
diff --git a/docs/cnpg.md b/docs/cnpg.md
index c2d43867f2..8e56d73843 100644
--- a/docs/cnpg.md
+++ b/docs/cnpg.md
@@ -1,76 +1,322 @@
-# CloudNativePG workspace
+# CloudNativePG
-A task-shaped view over [CloudNativePG](https://cloudnative-pg.io/) (CNPG): which PostgreSQL cluster needs attention, why, and what to inspect next — without assembling the story from ten separate CRD lists. The per-kind renderers, issue detection and audit check it builds on are described in [integrations.md](integrations.md#cloudnativepg).
+A task-shaped view over [CloudNativePG](https://cloudnative-pg.io/) (CNPG): which PostgreSQL cluster needs attention, why, what the instances are doing right now, and what to do about it — without assembling the story from ten separate CRD lists, `kubectl cnpg` and a Grafana dashboard. The per-kind renderers, issue detection and audit check it builds on are described in [integrations.md](integrations.md#cloudnativepg).
-The workspace is read-only. It never writes to a cluster.
+Reading never writes. Writes happen only through the [Actions](#actions) (backup, switchover, restart, fencing, hibernation, maintenance, backup schedule edit, pooler pause, destroy instance, cancel/terminate a backend, restore, a restore-validation note), each made with the caller's own identity after a confirmation bound to the facts they reviewed, behind the [GitOps write guard](#gitops-write-guard).
+
+## Code boundaries
+
+- `pkg/cnpg` owns pure interpretation of backup declarations, fencing, cron schedules and instance ownership. It has no Radar `internal/` dependencies. Audit, Issues and action guards use those rules; the TypeScript backup model in `packages/k8s-ui/src/utils/cnpg-backup.ts` runs against the same fixture cases in `pkg/cnpg/testdata/backup-declarations.json`.
+- `internal/cnpg` owns request-independent read orchestration, capabilities, actions, reports and CNPG log/activity interpretation. Its caller-scoped ports supply authorized observations, Kubernetes clients and metrics reads. `internal/server/cnpg_service.go` adapts the authenticated caller; handlers retain routing, bounded decoding, serialization and browser stream framing. Reports call service operations directly and apply typed redaction projections. They do not invoke handlers or parse their JSON output.
+- `internal/integration` owns the common read-source/coverage and action contracts, cache-scope helpers, bounded fan-out and version-bound patching. It has no server dependency. Prometheus query engines stay in `internal/prometheus`; CNPG reads bind one explicit metrics client instead of resolving global state during composition.
+- `internal/issues/source_cnpg*.go` owns findings derived from cached Kubernetes objects, including schedule destination blockers and Pod/status contradictions. Cross-resource findings retain the exact inventory grants in `Issue.RequiredReads`; composition, list/search counts and cached/grouped related-issue projections authorize them before serving REST, MCP or Diagnose results. Without an evidence authorizer these findings are withheld; only an internal memo may explicitly retain them for a later per-user projection. Unread or incomplete Pod inventories never establish a contradiction. Workspace-only live measurements remain in the frontend assessment model.
+- `packages/k8s-ui/src/components/cnpg` owns presentation and pure fleet derivations. Problems carry semantic `reason` values and WAL facts carry `state`; consumers do not inspect generated IDs or display sentences to decide behavior.
+- `web/src/api/cnpg*.ts` owns fetching and wire types. `web/src/components/cnpg/runtimeAssessment.ts` owns pure runtime assessment; `host.tsx` owns CNPG's summary, actions, logs, navigation and kind-list composition through the generic detail slots. The app-private `web/src/integrations/resourceHosts.tsx` composes it with existing renderer adapters; generic resource views consume that composition root. Exclusive slots have declared owners, while additive actions/tabs compose. `workspaceRoutes.ts` and `workspaceScreens.tsx` register CNPG and Capacity separately from their resource hosts, retaining integration-owned route parsers and scope policies.
+
+A backup destination, a WAL archive and a destination for a particular ScheduledBackup method are different facts. The shared declaration model records backup methods, enabled plugins and their WAL archiver designation; snapshots and backup-only plugins do not establish continuous archiving. Action capabilities expose `reasonCode` for behavior and `reason` for display, so a translated or reworded refusal cannot change its navigation.
## Where it lives
-CNPG stays inside **Resources**; there is no new global navigation item. When the `postgresql.cnpg.io` CRDs are discovered, the Resources sidebar's CloudNativePG group gains a **Workspace** block above its exact kinds:
+CNPG stays inside **Resources**; there is no new global navigation item. When the `postgresql.cnpg.io` CRDs are discovered, the Resources sidebar's CloudNativePG group gains a **Views** block above its exact kinds:
| Destination | Route | Job | Detail home for |
|---|---|---|---|
-| Overview | `/cnpg` | The fleet: every Cluster with instances, replication, protection, declarations and its top problem. Defaults to **Needs attention**. | Cluster |
-| Protection | `/cnpg/protection` | Recovery evidence per cluster, failed backups (7 days), destinations, schedules. | Backup, ScheduledBackup, ObjectStore |
-| Declarations | `/cnpg/declarations` | Databases, Publications, Subscriptions and managed roles by cluster; declared vs reconciled. | Database, Publication, Subscription |
+| Clusters | `/cnpg` | The fleet: every Cluster with readiness, replication, storage, backups and its top problem, ordered by urgency; **Needs attention** narrows it. | Cluster |
+| Backups | `/cnpg/protection` | Recovery evidence per cluster, backup runs (7 days; failed first when any failed), schedules with their last run's outcome, destinations. | Backup, ScheduledBackup, ObjectStore |
+| Declarations | `/cnpg/declarations` | Databases, DatabaseRoles (1.30+), Publications, Subscriptions and managed roles by cluster; declared vs reconciled. | Database, DatabaseRole, Publication, Subscription |
| Pooling | `/cnpg/pooling` | Poolers and the clusters they front. | Pooler |
-| Operator | `/cnpg/operator` | Operator and plugin workloads, image catalogs, operator configuration. | ImageCatalog, ClusterImageCatalog |
+| Operator | `/cnpg/operator` | Operator and plugin workloads, leader, watched namespaces, webhooks and reconcile errors, image catalogs, operator configuration. | ImageCatalog, ClusterImageCatalog |
+
+Destination badges count **affected clusters**, not findings, and follow the namespace filter (the sidebar says so). The exact kinds stay under a collapsible **Resource kinds** block, grouped by API group; on the CloudNativePG views it starts collapsed.
+
+**The Cluster kind's list is the Clusters view.** The generic route for the kind, `/resources/clusters?apiGroup=postgresql.cnpg.io` — where the sidebar's Cluster kind, a pinned favorite, search and every link to a CNPG Cluster from Issues, Home, Checks or an investigation land — renders the same Clusters view (one component over the same data and parameters) in place of the generic table, inside the Resources page: its drawer (`?resource=ns/name`), history and Back behave as for any kind, and expanding opens the Cluster's page. `/cnpg` stays the workspace's own home with its `?drawer=` trail; the sidebar highlights whichever was opened. A Radar without these views, or one whose capabilities could not be read, keeps the generic table; while capabilities load for the first time the page waits instead of showing the table first. Cluster API's `clusters` kind (`cluster.x-k8s.io`) is untouched. The table's column picker, column sorting, compare and bulk selection are not offered for this kind; **Create** is, on both routes, guiding name, namespace, instance count and per-instance data volume size, with optional StorageClass and image override, before opening the exact manifest. Three instances seed the choice; storage size requires an explicit value. An omitted image leaves selection to the operator. A single selected namespace seeds the choice; All or multiple namespaces require an explicit target. Both entries use strict Create, so an existing Cluster cannot be updated. Namespace suggestions are advisory: an unread list leaves manual entry available and the API server checks access. A context switch requires restarting the flow.
+
+**What "every Cluster" means.** The view lists every CNPG Cluster the caller may list, in the namespaces Radar can see for them: the namespace filter, the caller's namespace visibility (derived from their Pod/Deployment access, as everywhere in Radar) and Radar's own watch scope. When Radar's identity may not list Clusters cluster-wide, its cache watches them namespace by namespace: what it holds is listed, the rest reads as not read (named when the caller named the namespaces), and the count reads "≥N" — never an empty list or a total that looks exact.
-Destination badges count **affected clusters**, not findings, and follow the namespace filter (the sidebar says so). The exact kinds stay under a collapsible **Resource kinds** block, grouped by API group; on workspace screens it starts collapsed.
+Every CNPG kind's full detail is `/cnpg///` — reached from a row's **Open**, from the drawer's expand control, and by redirect from the generic `/workload/...` URL. The page keeps the sidebar with the CloudNativePG views (its destination highlighted, the object nested under it) and uses Radar's detail view underneath: **Overview** is a composed summary, **Spec & status** is the kind's existing renderer, then YAML and the rest.
-Every CNPG kind's full detail is `/cnpg///` — reached from a row's **Open**, from the drawer's expand control, and by redirect from the generic `/workload/...` URL. The page keeps the workspace sidebar (its destination highlighted, the object nested under it) and uses Radar's detail view underneath: **Overview** is a composed summary, **Spec & status** is the kind's existing renderer, then YAML and the rest. A Cluster adds **Protection** (its recovery evidence), **Activity** (in place of Timeline) and merged instance **Logs**.
+A Cluster's page is organized by task, each tab the one home of its facts (others link to it):
+
+| Tab | Job | Holds |
+|---|---|---|
+| Overview | What needs attention, and where to act | Standing notices and problems, then **At a glance** and **About** cards. At a glance holds one line per health dimension linking to its tab, instances, the neutral literal controller phase and live Standby cloning progress; About holds PostgreSQL version, declarations, poolers and GitOps source. A folded **Operator conditions** card preserves every `status.conditions` entry (type, status, reason, message, last transition). The drawer keeps flat summary sections |
+| Replication | Diagnose instances and assess a primary change | One card per instance: Pod readiness, node and QoS (Kubernetes) beside role, timeline and streaming state (PostgreSQL); per standby the backlog and the HA slot the primary keeps for it (as `kubectl cnpg status` shows it); other replication slots; the Serving verdict and read-write Service’s ready endpoints (with Reachability), the HA configuration (failure domains, pending restart, image, quorum, disruption budgets, Leases, Jobs). Instance actions live here |
+| Storage | What occupies space, why, and resize | WAL an inactive slot holds, with the standby it is kept for and how it is freed (get it streaming, or destroy it so the operator recreates it and drops its HA slot); volumes, measured usage or qualified lower bounds, what holds WAL (archive queue, slot retention per standby), resize; a class that cannot expand keeps existing claims unchanged, with restoration suggested only when recovery evidence is recorded |
+| Performance | Connections, blocking and how the databases are doing | **Sessions** (headroom, blocking tree, cancel/terminate) · **Database health** (cluster throughput, then the picked instance's database health and checkpoints) · **History** (every chart, filterable by group: replication, sessions, throughput, storage). Replication and Storage link to their chart group |
+| Backups | This cluster's recovery evidence and backups | **Restore to a new cluster**; while WAL archiving fails (and for a day after it resumes) a repair panel: the operator's ContinuousArchiving message, the last archived and failed WAL, where the destination is declared, the credential Secrets and endpoint CA it references (named, never read) or the workload identity it uses, the archiver's logs opened on its container (`plugin-barman-cloud` when the plugin is `isWALArchiver`, otherwise `postgres`), a link to the Operator when the Cluster's phase says a plugin blocks reconciliation, then a checklist — archiving resumed, and a completed Backup that started after the resume (the later of the last failure and the condition turning True), that a restore can start from with this cluster's WAL archive (written by the same archiver, or a volume snapshot), and whose own `status.beginWal` comes after the last WAL that failed to archive — a backup taken from a lagging standby can start after the resume yet begin before the failure. Without the failed WAL (only the primary's instance manager reports it, and its stats restart with the primary) the step stays unverified, never green. Resumption needs the instance manager's record of the failure (the condition turning True also happens when archiving is first set up); its stats restart with the primary, after which the panel has nothing to show; recovery evidence as facts, backup runs (all outcomes, newest 10 first), schedules with their last run, destination, restore validation for a restored cluster |
+| Activity, Logs | What happened | Events and changes; merged instance logs |
+| Configuration | What is declared, and what the instances report | A purpose-built Cluster page body: folded **Connect** (Services, poolers, application database; Service **Reachability** links), **PostgreSQL parameters** (Cluster declarations remain visible while loading or when exec is denied or fails; instance values name who answered, what applies a change and pending restarts), then populated declaration cards for instances/image, placement/resources, storage, bootstrap/source, backup configuration, roles/services/monitoring. Each fact names its spec path; Storage and Backups link to their evidence tabs. **Certificates** appears once, with a reason when its HA read is unavailable; folded labels and annotations finish the tab. Conditions and status problems live on Overview; events live on Activity. Only the expanded CNPG Cluster page replaces `spec`; the drawer's Spec & status and other kinds keep their renderers |
+| YAML | The object | — |
+
+Four health dimensions follow the reader on every tab. **Serving** sits on the title line beside the controller phase and opens Replication; **Replication**, **Storage** and **Backups** mark their own tabs: a dot when something needs a look, a hollow ring when it could not be assessed, nothing when it is fine, with the verdict in the mark's tooltip and in words on the Overview's At a glance. Replication always shows its Serving verdict, including healthy. They, the Overview, every tab and the fleet row read one assessment (`useCNPGClusterAssessment`): the fleet row (cached objects, issues, Prometheus) enriched with what the instance managers report, so no surface reads calmer than another.
+
+Configuration's parameters table separates the current declaration from the last instance sample. When no instance answered, every declared row remains and one shared unavailable explanation replaces repeated observation cells. A new parameter or one beyond the sample cap reads **Not sampled**; loading, denied and failed reads say **Not read** with the reason; **Not reported** means an instance answered without that parameter. A declaration added since the sample keeps the summary unknown, and an empty sample never implies there are no instance Pods. Failed refreshes retain observations and Pooler facts with a notice naming their age. `status.image` is the operator’s target image, not evidence of a running image; per-instance image observations belong to Replication.
## Navigation
- **One drawer.** Rows inspect in the app's single drawer; `?drawer=kind:group:namespace:name` backs it, so refresh, share and Back restore it. Links inside the drawer append to that chain and show "← " at the top of the drawer.
-- **Return vs location.** A full detail shows "← " only when it was reached by a drilldown (the label travels in history state); sidebar and global-nav hops are location changes and carry no return label. The crumb (`CloudNativePG / Protection / name`) always names the object's place, so a fresh tab has a parent without a fabricated previous task.
+- **Location, not history.** A full detail's title line starts with where the object lives — the workspace and the view that is its home, `CloudNativePG / Clusters /`, `CloudNativePG / Backups /` (the view is the link) — then the name. There is no "previous page" link: the crumb is the only control, the same in a fresh tab as after a drilldown, and the browser's Back returns to the previous page. Esc goes back to it after a drilldown (history state records one) and to the home view otherwise. Kind, phase and namespace follow the name on the same line — the kind is left out for the kind the view is named after (Cluster in Clusters, Backup in Backups, Pooler in Pooling), since the crumb already says it, and kept for the other kinds under the same view (ScheduledBackup, ObjectStore) — and wrap below it only when there is no room; the header actions drop below the title, as a group, before the name is squeezed. The fleet views' titles carry the same `CloudNativePG /` prefix.
+- **Operator banner.** While the operator is not reconciling (or its webhook rejects writes), a Cluster's page opens with the warning above the title line — before any status it shows — with **Open Operator →** on the warning's own title line.
- **Context.** Detail URLs carry `ctx=` (added on first view when absent). After a context switch the page says " is not in " with **Switch back** and **Go to …** — Radar never opens a same-named object from another cluster.
- **Namespace filter** narrows collections and counts. An explicitly opened object stays open, with a note when it is outside the filter.
+On Cluster/resource detail below 1300 CSS pixels, the workspace resource sidebar moves into the **Resources** dialog so the task tabs and evidence have more room. Fleet discovery keeps the sidebar visible above 1100 CSS pixels; narrower fleet views use the same Resources dialog. Desktop detail keeps it inline. The dialog retains the same workspace/kind navigation, Escape, focus restoration and context-pinned destinations. The fleet's Backups screen keeps concise archive/schedule, last-success and recovery-boundary columns; **Details** opens the full shared recovery facts with provenance, destination, schedule links and restore validation.
+
## The certainty contract
-Every value is something the cluster reports, labelled with where it came from. When the cluster does not report something the UI says so; it never shows zero, "none" or green in its place.
+Every value is something the cluster reports, labelled with where it came from. When the cluster does not report something the UI says so; it never shows zero, "none" or green in its place. These are Radar's shared rules for unknown and partial values ([DESIGN.md](../DESIGN.md#unknown-partial-and-denied-values)); the table below is where each CloudNativePG fact comes from.
| Fact | Source | When it is not known |
|---|---|---|
-| Instances, primary | `status.readyInstances`, `status.currentPrimary`, instance Pods (controller-owned by the Cluster's UID) | `–` |
-| Replication | Pod readiness only | Always "lag unknown": readiness does not show whether a replica is streaming. Lag needs runtime data Radar does not read yet. |
-| Schedule | ScheduledBackups targeting the Cluster (`spec.suspend` → suspended) | "No access to ScheduledBackups" when unreadable in that namespace |
-| Destination | barman-cloud plugin `barmanObjectName`, in-tree `barmanObjectStore`, or volume snapshots | "No destination configured" |
-| Last successful backup | Newest of: completed Backup CRs (7-day window plus the newest per cluster), ObjectStore `serverRecoveryWindow[...].lastSuccessfulBackupTime`, in-tree `status.lastSuccessfulBackup` (ignored for plugin clusters, where CNPG no longer sets it) — the winning source is shown | "None observed", or "No access to Backups" |
-| WAL archiving | `ContinuousArchiving` condition | "Not reported" |
-| Recovery window | Earliest point from ObjectStore `status.serverRecoveryWindow` for the cluster's server name. The latest point follows WAL archiving, not the last base backup, and no status reports it; it reads "not advancing" only while `ContinuousArchiving` is False | "Not reported" |
-| Restore validation | A Cluster in the same namespace bootstrapped (`bootstrap.recovery`) from this cluster's store/server or one of its Backups, **with a ready instance** | "None recorded" (unknown tone) — Kubernetes records no restore tests, so this is never green. A matching cluster without a ready instance reads "Recovery declared in …". |
+| Instances, primary | `status.readyInstances`, `status.currentPrimary`, instance Pods (controller-owned by the Cluster's UID) | Missing operator readiness reads “Not reported by the operator”; independently read Pod readiness names its observed count. An unread Pod inventory never implies zero |
+| Standby cloning | The primary's `/pg/status` `pgStatBasebackupsInfo` (`pg_stat_progress_basebackup`): phase, streamed of total bytes. CloudNativePG reads it only for application names ending in `-join` — a new instance cloning the primary, never a `Backup` — so it is shown in At a glance. A total PostgreSQL has not estimated yet reads "total not estimated yet", never 0 % | Omitted when runtime is unavailable; "None running" only when the primary's report was read and lists none |
+| Replication | The primary's `pg_stat_replication` via the instance manager; in the fleet, Prometheus `cnpg_pg_replication_lag` **and** `cnpg_pg_replication_is_wal_receiver_up` per standby; otherwise Pod readiness only | "Lag unknown" when neither is readable: readiness does not show whether a replica is streaming. A replay lag of 0 never stands in for streaming — a standby that receives nothing also reads 0 — so without receiver state it reads "streaming unverified" |
+| Schedule | ScheduledBackups targeting the Cluster (`spec.suspend` → suspended). Enabled is neutral; a method without a destination reads blocked in the Backups view and the drawer summary (amber). An enabled schedule with a verified destination blocker is a warning problem for its Cluster. The generic ScheduledBackup list status comes from that object alone; the drawer names this distinction beside a verified blocker. Only a completed Backup supplies successful run evidence | "No access to ScheduledBackups" when unreadable in that namespace |
+| Destination | Cluster barman-cloud plugin `spec.plugins[].parameters.barmanObjectName` and `serverName` (default: Cluster name), in-tree `barmanObjectStore`, or volume snapshots | "No destination configured" |
+| Last successful backup | Newest of: completed Backup CRs (7-day window plus the newest per cluster), ObjectStore `serverRecoveryWindow[...].lastSuccessfulBackupTime`, in-tree `status.lastSuccessfulBackup` (ignored for plugin clusters, where CNPG no longer sets it) — the winning source is shown. Backup runs must match the live Cluster UID recorded in `status.pluginMetadata.clusterUID` or its Cluster owner when present, and must not predate its creation; store/status success times before creation do not count | "None observed", or "No access to Backups" |
+| WAL archiving | Cluster spec's archive destination or enabled `isWALArchiver` plugin, then the operator's `ContinuousArchiving` condition with its transition time. A True condition reads "CNPG reports archiving", with its source inline. Declared archiver failures remain failing even with an invalid destination | No destination reads "Not archived: no destination configured": WAL is not archived to recovery storage, so point-in-time recovery is unavailable. Only a recorded True condition adds that CNPG accepts each WAL file without keeping it when no destination is configured. Fleet rows keep the verdict, “No point-in-time recovery” and Operator report; the explanation of CNPG’s success condition appears once for the table. The raw condition (type, status, message and relative transition age, with the exact time on hover) is folded under Operator report. Third-party archiver destinations are not assessed. With an archiver but no condition: "Not reported" |
+| Recovery window | Earliest point from ObjectStore `status.serverRecoveryWindow` for the cluster's server name. The latest point follows WAL archiving, not the last base backup, and no status reports it; it reads "not advancing" only while `ContinuousArchiving` is False | "None: no backup destination" with no destination; otherwise "Not reported" |
+| Restore validation | A person's note on a compatible restored Cluster (`radar.skyhook.io/restore-validation`, naming this Cluster's UID) reads "Validation recorded" with who and when. Recovery through a current-incarnation Backup reads "Restored into …"; matching the currently configured store/server archive alone reads "Archive restored into …", **with a ready instance**. Restores created before the current source, recovery targets before its creation, recorded predecessor source UIDs and pinned predecessor Backup IDs are excluded. A note copied from a different target UID is ignored | "None recorded" (unknown tone) — Kubernetes records no restore tests, so this is never green, and a recorded note is neutral, not passed. A matching cluster without a ready instance reads "Recovery declared in …" or "Archive recovery declared in …". Archive association does not establish which Cluster incarnation produced the restored data. |
| ObjectStore upload health | **Inferred** from its user clusters' WAL archiving and recovery windows (ObjectStore has no status of its own) | "Unknown" |
-| Declarations | `status.applied` (true / false / absent = pending); managed roles from `status.managedRolesStatus` (`reconciled`, `cannotReconcile`; anything else pending) | Pending, never failed |
-| GitOps source | Argo CD / Flux labels and the Argo tracking annotation | "GitOps source not recorded" |
-| Pooler pressure | — | "Not measured": needs PgBouncer metrics |
-| ScheduledBackup cron | Shown verbatim | CNPG's cron is six-field (seconds first) and is never translated |
+| Declarations | `status.applied` (true / false / absent = pending); managed roles from `status.managedRolesStatus` (`reconciled`, `cannotReconcile`; anything else pending). A DatabaseRole whose name also appears in the Cluster's `spec.managed.roles` is overridden — the Cluster spec wins and the operator reports it not applied; its summary says so, or "unknown" when the Cluster is not visible | Pending, never failed. Filter counts are exact only when all declaration kinds and Cluster specs were read in the selected scope; incomplete positive counts read ≥N and zero reads Unknown. A Cluster filter uses its namespace’s coverage |
+| Logical replication path | Per Subscription object: the subscriber's `spec.externalClusters[externalClusterName].connectionParameters.host`, resolved to a visible Cluster only through its `-rw`/`-ro`/`-r` Service or a Pooler Service (any other host stays external, never guessed); the publication by `publicationName` + `publicationDBName` (else the external cluster's `dbname`), linked to a Publication object when one declares it ("unknown: no access to Publications in " when that namespace's Publications are not readable, never "none declares it"). Slot name: `parameters.slot_name`, else the subscription's name, as PostgreSQL names it; observed on the publisher primary's `/pg/status` `replicationSlotsInfo` (type, active, WAL status) with retained WAL from the exporter. Survives failover: declared only — the publisher's `replicationSlots.highAvailability.synchronizeLogicalDecoding` together with `highAvailability.enabled` (default true; the operator enables synchronization only with both), then on PostgreSQL 17 the subscription's `failover` parameter (only failover slots are synchronized); before 17 it depends on `pg_failover_slots`, which Radar cannot see | Slot: "not read" / "no access (needs get pods/proxy on the publisher)", and "not found" only from a readable report. Failover: "unknown" for an external publisher, a missing PostgreSQL major, or PostgreSQL < 17. Apply errors and lag beyond the Subscription's `status.message`: not reported (the default exporter has no `pg_stat_subscription` query) |
+| GitOps source | The manager Radar's server detects from Argo CD / Flux labels and the Argo tracking annotation (`managedBy` in `/api/cnpg/workspace`) | "GitOps source not recorded" |
+| Pooler pressure | Each pooler Pod's PgBouncer exporter (`:9127/metrics`): per pool clients active/waiting, servers active/idle/used, max wait, the pool mode PgBouncer reports, and how many Pods reported | "Not measured" when no Pod could be read; partial (a lower bound) when only some reported |
+| Pooler readiness | The Pooler's Deployment (controlled by the Pooler's UID): ready of desired replicas. The Pooler's own `status.instances` counts scheduled Pods only | "Unknown" when the Deployment cannot be read |
+| Pooler paused | Requested: `spec.pgbouncer.paused`. Observed: each PgBouncer's `SHOW STATE` over the caller's `pods/exec` | Observed reads "Not observable: needs create pods/exec"; the exporter does not publish it |
+| Pooler limits | `spec.pgbouncer.parameters`; unset ones read "default" plus PgBouncer's own default where it was read from PgBouncer (`SHOW CONFIG`, 1.24): `default_pool_size` 20, `max_client_conn` 100, the per-database/user maxima and reserve/min pools 0. CloudNativePG writes only the parameters the Pooler sets | — |
+| Blocking sessions | `pg_stat_activity` + `pg_blocking_pids()` on one instance, via psql over the caller's `pods/exec` | "Blocking detail needs create pods/exec"; Sessions owns the independent aggregate-access explanation, so Blocking never repeats it |
+| Per-database health | The selected instance's exporter, one row per database: rollback ratio (`pg_stat_database` `xact_rollback` / (`xact_commit` + `xact_rollback`)), temporary files and bytes (`temp_files`, `temp_bytes`), transaction and multixact ID age (`pg_database`). Counters are cumulative since the last statistics reset; the shared-objects row (no datname) is left out | `—` for a value a family did not report, with the missing family named; a database with no transactions has no ratio, not 0 % |
+| Multixact ID age | Exporter `cnpg_pg_database_mxid_age` (default query `pg_database`, `mxid_age(datminmxid)`) per database, on the primary, beside transaction ID age | "unknown: the exporter did not report …" when a custom monitoring configuration dropped the family |
+| Extensions with updates | Exporter `cnpg_pg_extensions_update_available` (default query `pg_extensions`, every database): installed ≠ default version, with both versions | same; "every installed extension is at its default version" only when the family was reported |
+| Connect | Addresses from the Cluster spec: `-rw` / `-ro` / `-r` (minus `spec.managed.services.disabledDefaultServices`, plus `additional` Services), each Pooler fronting it (Service named after the Pooler), as `..svc` on 5432 unless a `serviceTemplate` sets the port. Database, owner and credentials Secret are resolved as CloudNativePG does — `bootstrap.recovery`, then `pg_basebackup`, then `initdb`; database defaults to `app`, owner to the database's name, Secret to `-app` ("name by convention"). The Secret is linked, **never read**; the connection string and psql templates carry a `` placeholder. Availability is separate: HA’s ready `-rw` EndpointSlices and instance Pod roles/readiness for `-ro` (standbys) and `-r` (any instance, including the primary). This is not a PostgreSQL login check. Local Radar pins port-forward to its kubeconfig context; embedded/in-cluster hosts say the command uses the current kubectl context | Unread availability reads “Not checked” with its reason or missing grant; Pooler/additional Service availability is not read. A monolithic import with no database reads "Unknown"; without `list poolers`, "No access to Poolers" |
+| Checkpoints | Exporter `pg_stat_checkpointer` (PostgreSQL 17+: `checkpoints_timed`/`_req`, `restartpoints_timed`/`_req`/`_done`, `buffers_written`) or `pg_stat_bgwriter` before 17 (`checkpoints_timed`/`_req`, `buffers_checkpoint`; a standby counts its restartpoints as checkpoints), per instance under Performance › Database health, whose instance picker (primary by default) also drives Sessions. Cumulative since the last stats reset | "Unknown: the exporter reported neither" |
+| Instance CPU / memory | metrics-server (`metrics.k8s.io`) usage of the postgres container, against its limit | "not measured: the metrics API is not available" |
+| ScheduledBackup cron | The Protection page and the Cluster's Schedule fact show the plain-language reading (`scheduleReadings` in `/api/cnpg/workspace`) in the operator’s clock, with the cron on hover or beneath; the schedule's own page shows it verbatim, with the reading and the next 3 runs from the server (`facts.preview` on the schedule's capabilities; `GET /api/cnpg/scheduledbackups/{ns}/{name}/schedule-preview?schedule=` for a draft). Parsed as the operator parses it (robfig/cron v1 `Parse`: six fields, seconds first, day of week optional, @-descriptors, no `TZ=` prefix) and counted as it counts: from `status.lastCheckTime` when set — a time that has passed since then means the operator runs one backup as soon as it sees the schedule, said so — otherwise from now; in the readable declared operator TZ, or explicitly as a UTC estimate when the clock is not established | Without the capabilities read, verbatim only; a schedule no time satisfies reads "the operator would never run it" |
-Problems come from Radar's Issues engine (the same detections as `/issues`) plus the audit's `cnpgNoDeclarativeBackup`, worded "No declarative backup schedule" because that is all it proves. A cluster **needs attention** when it has an issue of warning or worse on itself, an instance Pod, or an object that references it.
+For barman-cloud, every backup takes its destination from the **Cluster's** plugin entry: `spec.plugins[].parameters.barmanObjectName` selects the ObjectStore and `serverName` selects the archive server (default: the Cluster name). The plugin ignores `Backup.spec.pluginConfiguration.parameters` and `ScheduledBackup.spec.pluginConfiguration.parameters`, including those two keys. Backup-to-ObjectStore links and restore sources are inferred from the current Cluster configuration, which may have changed since the run; a Backup's parameters do not establish a historical destination. The standalone Backup and ScheduledBackup renderers have no Cluster to resolve, so they explain where the destination comes from and identify ignored parameters without linking them as destinations. Radar does not model third-party plugin destinations.
+
+A schedule's barman-cloud parameters cannot satisfy a missing Cluster destination: the enabled Cluster plugin must name its ObjectStore, or the schedule is blocked. Schedule declarations do not add recovery destinations. Unreadable ObjectStores and unsupported plugin methods keep recovery evidence unknown with the reason; only an assessed source with no recorded evidence reads as none.
+
+Backups from a previous same-name Cluster remain browsable, but do not supply current protection, current Cluster problems, inferred plugin destinations or restore-validation attribution. Schedule declarations still target by name; a reported run before the current Cluster was created is not a missed run for it.
+
+Backup schedules: besides "No backup has run since this schedule was due" (the operator's own `nextScheduleTime` passed), the Issues engine raises **"ScheduledBackup () has had no successful backup since its run"** on the Cluster, dated by the run (its first_seen; the workspace titles it "No successful backup since a scheduled run" and shows the run's age) when an active schedule fired (six-field cron, seconds first, parsed as the operator does) after the cluster's newest successful backup (Backup objects, the ObjectStore's `lastSuccessfulBackupTime`, in-tree `status.lastSuccessfulBackup`), allowing the last successful backup's duration plus 10 minutes, and no Backup started since is still running. Suspended schedules raise nothing. The workspace shows it only to callers who can list ScheduledBackups in the namespace.
+
+Problems come from Radar's Issues engine (the same detections as `/issues`) plus the audit's `cnpgNoDeclarativeBackup`, worded "No declarative backup schedule" because that is all it proves, plus measurements:
+
+- **" is not receiving WAL from the primary"** (warning; critical when no expected standby receives) — in the fleet from Prometheus (`cnpg_pg_replication_is_wal_receiver_up = 0` while in recovery), on the Cluster page from the instance managers (no row in the primary's `pg_stat_replication`, with what the standby reports: replay paused, no WAL receiver, another timeline). Both raise the same problem id, so the page's answer replaces the fleet's. A fenced standby, one running `pg_rewind`, and a replica cluster's designated primary are not raised.
+- **"Inactive slot holds N of WAL on for "** (warning) — an inactive physical slot on the **primary** retaining at least 1 GiB (`CNPG_SLOT_RETENTION_WARNING_BYTES`). CloudNativePG copies HA slots to standbys, where nothing streams from them, so a standby's copy never raises. The Storage chip reads "WAL held by an inactive slot" even while volume usage is unmeasured.
+- **“Backup schedule cannot run: no backup destination”** (warning) — an enabled ScheduledBackup’s method has a verified destination blocker in its readable target Cluster spec. It contributes to Needs attention, Backups counts and tab marks. Suspended or unread schedules raise nothing; a Cluster with no schedule and no destination stays posture.
+- Disk usage at or past 80 % / 90 %, and sustained replication lag (below).
+
+The fleet's lag covers only the standbys whose lag was read (a standby that has replayed nothing yet reports none): with fewer than `spec.instances − 1` — every instance, for a replica cluster — it reads "lag X (N of M standbys reporting)" and never healthy. A standby's WAL receiver must be down in every sample over 5 minutes, while the instance was a standby in every sample, before the fleet raises it — a restarting standby, or a former primary just after a switchover, is shown but not raised. A live read on the Cluster page replaces the fleet's answer only for what it disproves: a standby the primary streams to again, or a slot active, measured below the threshold, or absent from a complete list. A cluster **needs attention** when it has an issue of warning or worse on itself, an instance Pod, one of its own Job Pods, or an object that references it. A Job Pod (initdb, join, restore, clone, import, major upgrade) counts only where the caller may list Jobs, and only when the Pod's controller is that Job (name and UID) and the Job's controller is the Cluster (name and UID). Its problem names what the Job builds ("New standby pg-2: Can't be scheduled") and comes before the Cluster's own Ready condition, which sums the other problems up; a standby that cannot join is a warning, since the primary keeps serving. The Cluster page’s labelled Investigate action follows its assessment: any problem (including posture), or a degraded/unhealthy dimension, gives it prominence; this does not change the fleet’s Needs attention filter. The fleet lists clusters by urgency in every filter: worst problem first (critical, warning, posture, none), then the number of warning-or-worse problems, then all problems, then namespace and name.
## Access
-All data comes from `GET /api/cnpg/workspace`, authorized **per kind**: namespaced kinds use a cluster-wide `list` or fall back per namespace; `ClusterImageCatalog` needs a cluster-scope `list`. Each kind reports coverage (`full`, `partial` with the namespaces read, `denied`, `syncing`, `error`, `notInstalled`). Issues and audit findings are withheld where the underlying kind is not covered — Pod evidence only reaches callers who can list Pods. Denied namespaces are named only when the caller supplied the namespace list. A partial or denied kind makes the screen show a coverage notice, and its facts read "No access" rather than none.
+All data comes from `GET /api/cnpg/workspace`, authorized **per kind**: namespaced kinds use a cluster-wide `list` or fall back per namespace; `ClusterImageCatalog` needs a cluster-scope `list`. Each kind reports coverage (`full`, `partial` with the namespaces read, `denied`, `uncached`, `syncing`, `error`, `notInstalled`). A namespace the caller may read but Radar's cache does not hold reads "Radar does not cache …", never "no access". Issues and audit findings are withheld where the underlying kind is not covered — Pod evidence only reaches callers who can list Pods. Denied namespaces are named only when the caller supplied the namespace list. A partial or denied kind makes the screen show a coverage notice, and its facts read "No access" rather than none.
`GET /api/cnpg/operator` reads operator and plugin Deployments (label `app.kubernetes.io/name=cloudnative-pg`, plugin Services labelled `cnpg.io/pluginName`) and the operator's config references. It ignores the namespace view filter (the operator lives in its own namespace), returns ConfigMap data only with `get configmaps`, and never reads Secrets.
Cluster logs (`/api/cnpg/clusters/{ns}/{name}/logs`) need `get pods/log`; Activity (`.../activity`) drops events for kinds the caller cannot list. Deleted child objects stay attributed to their Cluster because Radar records the owning cluster on timeline events at ingestion; history recorded before that is marked incomplete.
-## Not in this version
+## Live instance data
+
+`GET /api/cnpg/clusters/{ns}/{name}/runtime` and `GET /api/cnpg/poolers/{ns}/{name}/runtime` read live data **through the caller's `pods/proxy`** (it feeds the Replication, Storage and Performance tabs): each instance manager's `/pg/status` (`:8000`) and the Postgres exporter (`:9187`), or each PgBouncer exporter (`:9127`). The apiserver strips the caller's credentials and `Impersonate-*` headers before forwarding, so a Pod never sees who asked. Only fixed GET paths are requested (some instance-manager paths mutate on GET), redirects are refused, and TLS is never downgraded after a certificate error. Requests are bounded: 4 in flight, 5 s deadlines, 1 MiB per status and 4 MiB per metrics body, memoized per identity and Pod UID for 5 s (status) or 25 s (metrics).
+
+Each source reports its own state (`ok`, `partial`, `denied`, `unreachable`, `error`). A status report the instance manager answered without finishing its reads — it masks errors while PostgreSQL may be unavailable (`mightBeUnavailableMaskedError`) and reads nothing while `pg_rewind` runs — is `partial` with `incomplete: true` and the masked error: a list it did not fill (replication, slots, base backups) and the archiver are `null`, never empty, and pending restart and a standby's role detail are not established from it. **Denied is never shown as zero**: without `get pods/proxy` each tab keeps what Kubernetes says (Replication lists the instance Pods with readiness and node, and the HA configuration) and names the missing grant for the live part. The tabs show replication (per-standby write, flush and replay lag), aggregated sessions and lock waits, transaction rates, storage and WAL, replication slots, and history — from Prometheus when Radar has it (see History), otherwise sampled while the tab is open — with gaps shown where a sample is missing. These reads carry no per-session query text; the blocking view (below) does, for callers who can exec into the instance. When a refresh fails while an earlier answer is still on screen (Replication, Performance, Storage, Blocking, the Cluster and Pooler summaries), `CNPGRefreshFailedNotice` (`web/src/components/cnpg/shared.tsx`) says so beside it — "Last refresh failed: · showing data from " — so cached values are never read as current.
+
+## Storage
+
+`GET /api/cnpg/clusters/{ns}/{name}/storage` backs the **Storage** tab. `GET /api/cnpg/disk` fills the fleet's **Disk** column and the Cluster overview's **Storage** fact with each cluster's fullest measured volume. Each value has its own source and its own "unknown":
+
+| Fact | Source | When it is not known |
+|---|---|---|
+| The instance's volumes | Claims labelled `cnpg.io/cluster` **and** owned by the Cluster's UID, naming an instance (`cnpg.io/instanceName`), by `cnpg.io/pvcRole` | "No access (needs list persistentvolumeclaims)". A labelled claim the Cluster does not own is listed as not counted |
+| Requested / capacity | The claim's `spec.resources.requests.storage` / `status.capacity` | `—` |
+| Resize in progress | Claim conditions (`Resizing`, `FileSystemResizePending`, resize errors), `status.allocatedResourceStatuses`, request larger than capacity, the Cluster's `status.resizingPVC` | — |
+| Healthy / dangling / unusable | The Cluster's `status.healthyPVC`, `danglingPVC`, `unusablePVC`, `initializingPVC` | Not shown |
+| Expansion allowed | The StorageClass's `allowVolumeExpansion` (unset = not allowed) | "expansion unknown" when the class cannot be read |
+| Used space | Prometheus `kubelet_volume_stats_used_bytes` / `_capacity_bytes` per claim, behind the same grant as the PVC chart | "Used space unknown" with the reason (no series, no Prometheus, no access, claim names reporting under more than one cluster identity). Never 0, never a value merged across clusters. When kubelet's filesystem is much larger than the claim (local volumes), the view says the figure is the shared filesystem's |
+| WAL on disk | Exporter `cnpg_collector_pg_wal{value="size"\|"count"}` | "unavailable" with the pods/proxy state |
+| Waiting to archive | Instance manager `readyWalFiles`, with the last archived/failed times | No destination reads “Not archived: no destination configured”; the manager’s raw record is folded and distinguished from an unread record |
+| Held by replication slots | Slot existence and activity from the instance manager; retained bytes from exporter `cnpg_pg_replication_slots_pg_wal_lsn_diff` | An inactive slot without bytes says "retained WAL not reported"; "No slots" requires a read, empty inventory. Reported bytes remain a lower bound if inventory or some byte readings are missing |
+
+The three WAL measures overlap and are shown side by side, never added up. Disk findings (≥ 80 % warning, ≥ 90 % critical) come only from a measurement, are computed where they are read, and reach the fleet's Needs attention only for callers who received the measurement (`list persistentvolumeclaims` + the Prometheus PVC grant); they are not Issues-engine issues.
+
+**Expanding.** The view names the field each size is declared in — `spec.storage.size`, `spec.walStorage.size`, or `spec.tablespaces[name=].storage.size` (or the `pvcTemplate` request when that is how it was declared) — and whether the class allows expansion. **Edit size…** shows the GitOps write guard for that field, then opens Radar's apply flow with a manifest carrying only the new size (tablespaces carry the whole list, which the CRD does not merge by key). When the class forbids expansion, the action explicitly edits only the declared size; existing claims remain unchanged. The operator can resize existing bound claims only when the class allows it; CloudNativePG does not shrink volumes. Existing Pending claims keep their original request and named class too: editing the declaration affects future claims, and a fresh never-started Cluster can use new settings. Expansion guidance uses the existing restore assessment plus recorded completed Backup, successful-backup status or recovery-window evidence: available, checked-none and unread are distinct. A declared destination alone does not justify restore advice; an unread source gets no setup or restore prescription.
+
+Volumes and their usage do not depend on `pods/proxy`: without it the Storage tab still shows them under its access notice.
+
+The replication view measures each standby's catch-up as **replay backlog in bytes**: the primary's current LSN minus the standby's replay LSN, broken into not sent / not written / not flushed / not replayed. PostgreSQL's `write_lag`, `flush_lag` and `replay_lag` are shown as what they are — acknowledgement delay for recent WAL, empty when idle and caught up — never as catch-up time. The same rule holds elsewhere: the Replication chip and fact call `replay_lag` a "replay delay", and the switchover dialog warns only from the backlog in bytes ("has 16 MiB of WAL still to replay"), showing the delay beside it. The fleet's lag and the sustained-lag problem are Prometheus `cnpg_pg_replication_lag`, the time since the standby last replayed, so they keep the word "lag". The Replication chip takes the sustained-lag problem into account and never reads calmer than it ("· sustained lag", the problem as its tooltip). Each instance also carries its own report from `/pg/status`: a role detail derived only from that report (`primary`, `streaming` standby, `fileBased` = no WAL receiver, `replayPaused`, `pgRewind`), `pendingRestart` / `pendingRestartForDecrease`, timeline and instance-manager version. Pending restart is runtime-derived: it shows on the cluster page and the instance cards for callers who can read runtime, and never enters the cached-object Issues engine.
+
+## HA and instances
+
+HA readiness uses `declaredInstances` from `spec.instances`; `expectedInstances` combines `status.instanceNames` with instance names on verified Cluster-owned pending or active join/initdb Jobs. A missing Pod is an observation, not proof that it never existed. "Waiting to join" additionally requires a pending or active owned join Job. Job phase `active` describes `status.active`; a verified Job-controlled unschedulable Pod instead reports `pending` with the scheduler's reason. Reading these Pods requires `list pods`, in addition to the Jobs' own `list jobs` gate. Image agreement requires observed instance Pods and describes their declared images.
+
+The Cluster page's neutral Phase badge is the operator's `status.phase` (shortened with the existing phase vocabulary, exact text on hover), separate from Radar's Serving verdict. Runtime captions say "checked" when no measurement was returned. Runtime `slotsTruncated` and Storage `slotInventoryTruncated` identify slot-list caps specifically; a partial read of another family does not make a complete slot inventory partial. Storage's `slotInventory` is null when unread and an empty array only when read empty; exporter retention stays in `slots` and is also joined by slot name into the inventory.
+
+`GET /api/cnpg/clusters/{ns}/{name}/ha` backs the Replication tab's instance cards and "High availability" section, the Configuration tab's "Certificates", the Serving status and tab marks, the switchover dialog and the operation tracker. Each section folds to one summary line (`cnpgHASummary`, `cnpgCertificatesSummary` in `ha.ts`) and opens itself when something in it is out of line: an instance not ready, a shared Node or single zone, a pending restart, image drift, a quorum that does not hold, missing disruption budgets, an expired or foreign lease, a failed Job, or a certificate near or past expiry. A calm summary names only what is known, never "fine" for what was not read. Connect is a header button that opens the same facts in a dialog (`?connect=/`). Reading the Cluster (`get clusters`) never implies the rest: each part is authorized on its own and reports `ok | denied (with the grant) | notFound | notInstalled | unavailable | error`.
+
+| Fact | Source | Gate | When it is not known |
+|---|---|---|---|
+| Instances: node, QoS, running image vs desired (`status.image`, else `spec.imageName`), Pod and postgres-container start | instance Pods (controller-owned) | `list pods` | "Pods not readable" |
+| Zone | `topology.kubernetes.io/zone` on each Node | cluster-scoped `get nodes` | zones unknown (never "single zone") |
+| Failover quorum | `FailoverQuorum` of the Cluster's name (1.27+): recorded sync configuration, not the operator's verdict. N = potentially synchronous standbys, W = `standbyNumber`, R = those with a ready Pod (not the recorded primary); shows whether R + W > N | `get failoverquorums` | not installed / not found (quorum failover off) / a reset object = "no configuration recorded: a failover would wait" |
+| Disruption budgets | PDBs owned by the Cluster: expected / healthy / allowed, stale when not observed | `list poddisruptionbudgets` | `enablePDB: false` is the declared state, not a fault |
+| Primary Lease | `Lease` of the Cluster's name (1.30+), holder, renew, expired | `get leases` | "creates one from 1.30" — absence on older versions is not a fault |
+| Operator leader | `db9c8771.cnpg.io` Lease in the operator Deployment's namespace | `list deployments` + `get leases` | operator namespace unknown |
+| Cluster Jobs | Jobs owned by the Cluster (`cnpg.io/jobRole`: initdb, join, major-upgrade, snapshot-recovery…), phase | `list jobs` | "No access to Jobs" |
+| Read-write endpoints | ready EndpointSlice Pods of `-rw` | `list endpointslices` | the Serving chip falls back to primary Pod readiness and says so |
+| Certificates | `status.certificates.expirations` (Go `Time.String()`, parsed server-side; unparseable = unknown) with renewal owner: operator-generated (CNPG renews) or named in `spec.certificates` (you renew). For user-provided Secrets, their **metadata only** names a cert-manager `Certificate` | `get secrets` (metadata client, never data) | "issuer unknown" |
+| Maintenance | `spec.nodeMaintenanceWindow` (`reusePVC` defaults to true) | — | — |
+
+The four dimensions — **Serving** (title line) · **Replication · Storage · Backups** (tab marks) — each come from their own source (primary Pod readiness + `-rw` endpoints; the primary's `pg_stat_replication`, or a standby another source saw receiving nothing; volume usage and inactive-slot retention; WAL archiving, destination and last backup) and read **unassessed** when it is unavailable. The controller phase stays labelled "reported by CNPG".
+
+Certificate expiry is also an Issues-engine finding (`CNPGCertificateExpiring`, or `CNPGCertificateExpired` once past, one per Secret with the same fingerprint): a certificate its owner renews is a warning under 30 days and critical under 7; an operator-managed one only once renewal is overdue (under a day — CNPG renews at 7 days by default, so earlier would light every cluster for a third of each 90-day lifetime); expired is critical.
+
+## History
+
+Prometheus history is bounded to the current Cluster's creation, plus the query lookback window (at least five minutes), aligned up to the sampling step. This excludes predecessor data under reused Pod/PVC names, including rate-window samples. A newly created Cluster waits for that window before offering history; in-page sampling is still available.
+
+`GET /api/cnpg/clusters/{ns}/{name}/history?range=15m|1h|6h|24h` backs **Performance › History** (`?charts=replication|sessions|throughput|storage` filters by group; Replication and Storage link there). `GET /api/cnpg/fleet-metrics` fills the fleet's **Replication** column with measured standby lag and adds volume growth under **Disk**. Lag becomes a **Needs attention** problem only when it is sustained: every replay-lag sample Prometheus recorded for the worst standby over the last 10 minutes was at least 30 s (warning) or 5 min (critical) (`min by (pod) (min_over_time(…))` over raw samples, so a low sample from any scrape job counts), and that standby was already reporting by the window's start (the series answers at `offset 10m`), so one that appeared a minute ago never qualifies. It reads " ≥ 55 s behind in every sample for 10 min" (whole seconds below 100, minutes above): a lower bound over the samples Prometheus recorded, saying nothing about scrapes it missed (the fleet cell shortens it to ": all samples ≥ 55 s behind (10 min)"); the queried families are on hover over its "Measured by Prometheus" source. Scrape gaps are not filled in, and the wording claims only recorded samples. A spike is shown and coloured, not raised. Volume growth is never raised.
+
+**Sources.** Only server-built PromQL runs; nothing from the request reaches a query except the range name. Steps keep every chart at 60–144 points (15 s, 30 s, 3 min, 10 min). Charts: replay lag per standby (`cnpg_pg_replication_lag` while `cnpg_pg_replication_in_recovery = 1`, so a primary's constant 0 is not drawn), client sessions by state (`cnpg_backends_total` without `streaming_replica` and the metrics exporter, plus an "all states" total that is 0 only while the exporter reports Postgres up), sessions waiting on locks, transactions per second, WAL archived/failed per minute, WAL on disk (`cnpg_collector_pg_wal{value="size"}`), volume used % (kubelet volume stats for the Cluster's owned claims, the Storage selection), database size, temporary-file writes, deadlocks and checkpoints (`pg_stat_checkpointer` on PostgreSQL 17+, `pg_stat_bgwriter` before). `max by (pod, …)` collapses duplicate scrapes before any sum. Each chart names its source and its sample coverage (evaluation steps with at least one sample).
+
+**Isolation.** Instances are selected by `namespace` and `pod=~"^-[0-9]+$"`, so instances deleted since stay in range. The `cluster` label is deliberately not used: shared Prometheus setups use it for the Kubernetes cluster while CNPG scrape configs often relabel it to the CNPG Cluster name, so its meaning cannot be read from the series. Cluster-identity labels are added when the operator configured them or when kube-state-metrics Pod UIDs prove them (the same proof the workload charts use). Without that proof, the identities are counted over the **whole requested range** (`count_over_time`, so a cluster that stopped reporting minutes ago but still fills the chart is seen): history is refused when one selected Pod name appears under more than one identity (`ambiguous`), and shown with the isolation marked unverified otherwise. Nothing is pinned without proof: labels merely seen on the series are not identity — the CNPG exporter labels its own series `cluster=`, which would drop families that lack it and make two database clusters in one namespace look like two Kubernetes clusters. Only configured or UID-proven labels become matchers. Proven labels that select no CNPG exporter series refuse history too (`scopeMismatch`). kubelet volume stats are scoped separately by the same rules over their own series, because they come from a different scrape: the volume chart, the fleet's growth (over its 6 h window), the Storage tab's used space and `/api/cnpg/disk` (instant). An ambiguous or mismatched claim scope makes those `ambiguous` / `scopeMismatch` with the reason — never a value merged across clusters. The generic PVC usage gauge (`/api/prometheus/pvc/...`) applies the same scope and answers `ambiguous_scope` / `scope_mismatch`.
+
+**What is unknown.** Every chart has its own state: `noSeries` ("not scraped", never zero), `empty` (scraped, nothing to plot — no standby, no client sessions), `denied` (the grant), `error`, `notRead`. Steps without a sample are hatched: a leading run is "before this series existed or beyond Prometheus retention", any other a missed scrape or an instance that was down. A per-instance chart names each current instance it has no line for (replay lag: each current standby), so a stopped standby does not vanish from the chart about it; the sampled fallback does the same for standbys the primary does not list as connected. `cnpg_pg_replication_lag` is the standby's replay delay for recent WAL; it reads 0 when idle and caught up and is not a catch-up time (the LSN backlog under Replication is).
+
+**Access.** `get clusters` (403 otherwise). CNPG series need the Prometheus pod-metrics grant (`get pods` in the namespace, as for the Pod charts); the volume chart needs `list persistentvolumeclaims` plus the PVC chart grant. Without Prometheus the response is `source: "none"` with the reason, and History falls back to samples this page takes while open, labelled "since this page opened": replay lag per standby; client sessions and lock waits per instance; sessions by state (active, idle, idle in transaction, other) for the instance picked in Performance; transactions per second, cache hit ratio, WAL archived/failed per minute, WAL on disk, database sizes (the five largest by default, any or all selectable; the exporter reports the 200 largest), checkpoints per minute (timed vs requested), deadlocks per minute and temporary-file bytes per second from the primary. Counters become rates between the exporter's query runs: `cnpg_last_update_timestamp` (from 1.30 CNPG caches query results for `monitoring.metricsQueriesTTL`, 30 s by default), readings of one run counted once. Without that metric, rates use the time Radar fetched each reading and the charts say "approximate". A rate is taken only between two readings of the same primary Pod (name and UID): a change of primary, a counter that went down, or an interval with no block reads is a gap, never a value. Fleet volume growth also requires a currently present volume series under Prometheus staleness/lookback rules; an older six-hour derivative alone is not current evidence. The fleet reads one lag query and one growth query per namespace (≤ 64 namespaces); without a reading the column says "lag unknown (no metrics)" or "(no access)", never a number.
+
+**Intervals.** Drag across a chart, or click one step, to select an interval. The chip opens Logs with `since`/`until` — the logs endpoint takes `sinceTime` + `untilTime` and then reads each instance **from the start of the interval** (up to 64 KiB each, streaming off), because the pod log API has no upper bound and a tail would return the lines nearest now — and Activity with `since` + `until`. An interval that ends before a container's current run started is read from its previous run (`previous=true`), and one that spans the restart reads both; the kubelet keeps only those two runs, so the notice counts runs older than that which covered the interval. A byte-cap notice appears only when the cap cut into the interval.
-Runtime data (replication lag, sessions, locks, WAL and slots via the instance manager or Prometheus), Pooler pressure, and operations (Backup now, Switchover, Restart, Hibernate, Restore).
+## Actions
+
+Capabilities (`GET /api/cnpg/clusters/{ns}/{name}/capabilities`, `/api/cnpg/scheduledbackups/{ns}/{name}/capabilities`) return the facts a confirmation is bound to, each action's verdict (allowed, or a reason naming the missing grant or the blocking state), per-instance actions, and the effects of a restart or hibernation. A Cluster's `restore` verdict is `create clusters` (postgresql.cnpg.io) in its namespace, where the restored Cluster is created; it gates the More menu's Restore item. Actions are POSTs to `.../actions/{action}`, made with the caller's impersonated client:
+
+| Action | Write | Grant |
+|---|---|---|
+| Back up now | create `Backup` with an explicit method (plugin only when the plugin reports backup capability; an unreported capability is labelled) | `create backups` |
+| Switchover / promote | status patch: `targetPrimary`, `targetPrimaryTimestamp`, phase and Ready condition, as `kubectl cnpg promote` does. Each candidate shows what the live read says is wrong with it (not connected to the primary, replay paused, another timeline); such a standby is never the default pick, and choosing it warns that it has not received WAL from the current primary and how far behind it is cannot be measured | `patch clusters/status` |
+| Restart (rolling) | annotation `kubectl.kubernetes.io/restartedAt`. Review separates each planned effect from current readiness. Automatic primary update is described from the planned effect, respecting fenced and single-instance exceptions; raw update settings are folded | `patch clusters` |
+| Restart one instance | standby: delete the Pod with a UID precondition; primary: status phase write, as `kubectl cnpg restart` does | `delete pods` / `patch clusters/status` |
+| Reload configuration | annotation `cnpg.io/reloadedAt` | `patch clusters` |
+| Fence / lift fence | annotation `cnpg.io/fencedInstances`, a JSON array (`["*"]` = all). The dialog requires an explicit instance choice, with no default (including instance-row entry points); “All instances — stops service” follows individual instances. Primary/all retain typed confirmation. Malformed JSON blocks the action; lifting one instance while `*` applies is refused; lifting is refused while an offline volume-snapshot Backup of the Cluster runs, because the operator fenced the instance for it and lifts that fence itself | `patch clusters` |
+| Hibernate / rehydrate | annotation `cnpg.io/hibernation` `on` / `off`. Kept PVCs report capacity separately from requested size; absent capacity says “capacity not reported” | `patch clusters` |
+| Node maintenance set / lift (advanced) | merge patch `spec.nodeMaintenanceWindow {inProgress, reusePVC}` as `kubectl cnpg maintenance set/unset`; binds the reviewed maintenance facts; a standing banner offers the lift | `patch clusters` |
+| ScheduledBackup suspend / resume | `spec.suspend` | `patch scheduledbackups` |
+| ScheduledBackup edit schedule | merge patch `spec.schedule` with `resourceVersion`, bound to the reviewed schedule (409 `changed` otherwise); the new value is validated server-side with the operator's parser (400 `invalid_schedule`). The dialog previews the reading and next runs as you type and warns when saving triggers an immediate run | `patch scheduledbackups` |
+| ScheduledBackup run now | create `Backup` copying method, plugin configuration, online settings and target; capability and execution require the target Cluster's destination for that method | `create backups` |
+
+| Cancel / terminate a backend | from the blocking view only: `pg_cancel_backend` / `pg_terminate_backend` run by psql in the instance, for the pid only while its `backend_start` still matches, client backends only | `create pods/exec` |
+| Destroy instance (standby) | `kubectl cnpg destroy` parity: the instance's PVCs detached (`--keep-pvc`: Cluster owner removed, `cnpg.io/pvcStatus: detached`) or deleted with UID preconditions, then the Pod, then Jobs labelled `cnpg.io/instanceName`. The operator creates a replacement instance under a new name. The primary is refused (switch over first), and the instance **must be fenced** (named in `cnpg.io/fencedInstances`, or `["*"]`): the operator never promotes a fenced instance, so no failover can make it primary between Radar's last check and a delete — re-reading alone could only narrow that window, since CloudNativePG offers no lock and upstream `kubectl cnpg destroy` does not check at all. The dialog offers "Fence first". Before every PVC, Pod and Job step the Cluster (current/target primary, switchover/failover phase, the fence) and the Pod's `cnpg.io/instanceRole` label are re-read and any change stops the sequence. Last, the destroyed name is removed from a list-form fence (merge patch with `resourceVersion`; `["*"]` is left alone); a failure there is `partial` | `delete pods`, `list`/`delete` (or `update`) `persistentvolumeclaims`, `list`/`delete jobs` |
+| Pooler pause / resume | `spec.pgbouncer.paused` (the operator runs PgBouncer `PAUSE` / `RESUME`) | `patch poolers` |
+
+**Back up now parameters.** The dialog sends the chosen plugin's name and no plugin parameters. The server returns 400 when barman-cloud `pluginParameters` contains `barmanObjectName` or `serverName`, explaining that “the barman-cloud plugin takes its destination from the Cluster.” Third-party plugin parameters are still forwarded. A manual schedule run still copies its declared settings, including parameters for third-party plugins; copied barman-cloud parameters remain ignored and cannot unblock a missing Cluster destination.
+
+## Guided backup setup and timing
+
+A Cluster's **Backups → Set up backups** joins three tasks: attach existing archive storage, establish a matching schedule, then verify uploads and a new successful base backup. Restore's backup next step opens the same guide. Each write has an independent review; a saved declaration is not a successful upload or a tested restore. Existing matching schedules remain directly reachable, including their Suspend/Resume controls.
+
+The guide attaches an **existing same-namespace Barman ObjectStore** by patching `spec.plugins`. It preserves unrelated plugins and parameters, enables the Barman WAL archiver, and uses an explicit server name (initially Cluster name plus a UID suffix). It does not author stores/credentials, install the plugin, change retention, or migrate an existing archive destination. Existing in-tree, snapshot and other WAL-archiver configurations require a separately reviewed YAML migration. ObjectStore `spec.configuration.serverName` must be empty; the Cluster's plugin parameter supplies the identity.
+
+The server reads the Cluster and selected ObjectStore as the caller, compares the recovery source and visible archive users by destination/endpoint/server identity, and performs a server dry-run. It rejects a known collision, including another ObjectStore pointing at the same archive. Unread inventory is stated, and malformed archive configuration blocks review; remote contents, endpoint aliases and credential validity are not established. Confirmation explicitly acknowledges that the identity is reserved. UID/configuration digests bind the reviewed Cluster and ObjectStore; the writer re-reads them, refuses changed facts, and patches the fresh Cluster resourceVersion without retry. These reads do not lock the ObjectStore against a subsequent concurrent change.
+
+Schedule creation uses strict create with an explicit name, Cluster reference, Barman method/plugin and `immediate: false` unless selected. A narrow method repair changes only `spec.method` and `spec.pluginConfiguration.name`, retaining timing, target, online mode and plugin parameters. It requires the Cluster's enabled Barman WAL archiver and readable ObjectStore, binds the Schedule and Cluster identities/configuration, refuses deleting resources, and does not migrate a snapshot or another plugin's schedule.
+
+Attachment needs `get clusters`, `get objectstores`, and `patch clusters`; archive-user inventory uses `list clusters` (cluster-wide, with a namespace fallback when denied). Recovery references can additionally require `get backups` or the referenced ObjectStore. Schedule repair needs `get scheduledbackups`, `get clusters`, `get objectstores`, and `patch scheduledbackups`; creation needs `create scheduledbackups`. The API server's exact denial is retained when a prerequisite read fails. The guide's final backup uses the existing capability-gated **Back up now** action and operation observer.
+
+Daily, weekly and monthly inputs are shared by schedule creation and editing. Advanced cron retains other valid expressions rather than silently projecting them onto common controls. Days 29–31 explicitly skip months without that day. The server uses CNPG's six-field seconds-first parser. Existing schedule previews count from `status.lastCheckTime`; new schedule previews count from now. Times serialize as UTC. One readable watching operator Deployment with a literal valid `TZ` supplies the **declared** clock. Partial/ambiguous inventory, indirect TZ, `envFrom`, or no TZ declaration produces a labelled UTC estimate, never a verified running-clock claim. A due-run warning uses that same certainty. Fleet descriptions say “operator clock” without claiming a zone. A missing-success finding dates itself from the operator’s reported `status.lastScheduleTime`; cron alone cannot establish that a run happened.
+
+## Local operations and handoff
+
+The Cluster header follows accepted operations using caller-readable live facts. The fleet exposes their **last checked state in this tab** without mounting an observer for every row; opening the Cluster resumes following unfinished operations. Records match context/namespace/name and the Cluster UID where it was recorded, so a recreated Cluster does not inherit a known-UID operation. A schedule run whose requester lacked the Cluster UID says that identity still needs verification.
+
+Replacement is concluded from a successful identity read since the operation started. Unfinished predecessor records remain observed while the first identity check after mounting is in flight, including records hidden by the current page's identity filter. A known-UID operation cannot finish before identity is verified; if identity remains unavailable after the first reads settle and the 15-minute follow window has elapsed, it ends unobservable with that gap stated.
+
+**Copy handoff** includes the subject, known UID, requested action/time, last check, observed state and verification checklist, plus a context-pinned live Cluster link that respects an embedded host's basename. It excludes the private baseline. Recipients see current Kubernetes facts; they do not import the sender's operation record. Session storage remains local, bounded to 50 records, with finished records pruned after 30 minutes during updates; there is no durable or shared operation history.
+
+## Restore
+
+Restore starts from the Cluster's More menu, a Backup's **Restore from this backup** or an ObjectStore's **Restore a cluster from this store**. The dialog is review, not a gate:
+
+- **Recover to** the latest archived WAL, a point in time (entered in UTC or the browser's zone; the manifest always carries UTC), or — for a plugin Backup — the end of that backup (`recoveryTarget.backupID` + `targetImmediate`; CloudNativePG restores a `bootstrap.recovery.backup` reference only for in-tree and snapshot backups, so plugin Backups go through the ObjectStore with the ID pinned).
+- **What the source holds**, each with its source and in UTC and local time: first recoverability point and last successful backup (ObjectStore `serverRecoveryWindow`, completed Backups, in-tree status), WAL archiving (`ContinuousArchiving`) and the last archived / last failed WAL (the primary's instance manager, needs `pods/proxy`). For an ObjectStore, WAL evidence requires a live Cluster's enabled Barman plugin to archive WAL into that store and server identity; a backup-only plugin can supply source settings without establishing WAL coverage. A target before the first point, after the newest evidence, or in the future is **warned about, never blocked**; missing evidence is listed as a gap.
+- **Copied from the source** (instances, image or `imageCatalogRef`, storage, WAL storage, tablespaces, PostgreSQL parameters, resources) is listed for review and editable in the manifest. The manifest never carries `plugins` or `backup`; the dialog warns that the new cluster has no archiving until configured and must not reuse the source's `serverName`.
+
+Restore first groups the source, recovery point and archive evidence; the next step summarizes that choice and guides the new Cluster name, instances, data storage and optional StorageClass. Image/catalog selection stays inherited for physical recovery. A source without readable image and storage facts requires explicit inputs instead of a guessed image or tiny volume. The source’s complex PVC template, WAL storage, tablespaces, parameters and resource requests remain in the manifest.
+
+The manifest opens in the standard create dialog (server dry-run, strict create, with the Apply/Create choice and Force hidden). Back from the review returns to YAML; **Back to restore setup** retains the current YAML draft and setup choices. Reviewing unchanged setup reopens that draft, including advanced edits. Common target-field changes update only those YAML paths and preserve comments and other configuration. Changing the recovery source/time or an unsupported resource shape requires explicitly choosing **Replace manifest**; canceling retains both the edits and the setup. A context switch requires restarting the restore. The dialog itself asks `GET /api/cnpg/restore/capability?namespace=` — `create clusters` (postgresql.cnpg.io) in the namespace the new Cluster goes to, refused while the operator's webhook rejects writes — so it behaves the same however it was opened: without the grant it names it and Review stays disabled. The Cluster's More menu also disables the item from its capabilities. If the check itself fails, the dialog says so and leaves the server dry-run at review to decide. After the create, a toast links to the actual created Cluster. The submitted manifest determines whether it is a restore: changing Advanced YAML to initdb does not start a restore tracker or claim backup recovery. For a recovery manifest, the header's operation tracker follows it (`restore`, completed at a healthy phase with every instance ready, failed on a failed recovery Job or a failing phase), and the Cluster's Overview shows **Restore in progress**: phase, the recovery Job's Pod with its init containers and a link to its logs, and Warning events, from `GET /api/cnpg/clusters/{ns}/{name}/recovery`. Once done it collapses to one "Restored from …" line with a **Next steps** checklist of links (no writes): point applications at it (opens the header's Connect dialog; never marked done, since Radar cannot tell), record what you checked (opens the validation note; done once one is recorded), and set up backups and WAL archiving (the Backups tab with its setup guide; done only once the spec archives WAL — the Barman plugin as WAL archiver with its ObjectStore named, or `barmanObjectStore` — and partly done with volume snapshots or a non-archiving ObjectStore alone, which give no point-in-time recovery).
+
+**Restore validation.** Once a restored Cluster is up, its Backups tab's Restore validation section first shows **what Radar reads in it**: in recovery or not, where recovery stopped (the current timeline's history file, newest switch first, beside the declared `recoveryTarget` — never matched against it, since a later failover or switchover writes the same history), the databases with their sizes, the roles, and in the bootstrap database the tables and the planner's row estimates. These describe what is there; the bootstrap database and owner exist either way, since CloudNativePG creates them. Then someone records what they checked (row counts, newest transaction vs target, a smoke query), optionally the recovery target, and the server writes it as the `radar.skyhook.io/restore-validation` annotation with their user name, the time and the source and restored Clusters' UIDs it read itself (`POST .../restore-validation`, `patch clusters`, GitOps write guard). It reads "Recorded", never "passed".
+
+## Report bundle
+
+**Download report…** (Cluster More menu) returns a zip shaped like `kubectl cnpg report cluster`: the Cluster, its owned Pods, Jobs and PVCs, events about them, its Backups, ScheduledBackups, Poolers and ObjectStore, operator and plugin versions, and the live instance and storage snapshots Radar shows. `report.json` lists every item with what was read, skipped or denied (and the grant), and the Secrets the cluster references by name. Secret values are never read. Pod logs are opt-in; query text inside PostgreSQL log records is removed unless also opted in: the `query`, `internal_query` and `context` fields, everything after `statement: `, everything after the keyword of a message that (after an optional `duration: … ms`) starts with `execute`, `parse` or `bind` in any case — statement and portal names are client-chosen strings that may contain spaces and colons — and auto_explain's `plan:`, and bind values after `parameters:` — in JSON records and plain-text lines alike. The match is deliberately loose; a false positive only removes more. The bundle is capped at 32 MiB.
+
+Pod ownership includes the current Job UID; matching labels or a reused Job name alone do not establish ownership. Backups use the same Cluster-incarnation matching policy as recovery. Events require the exact UID of the Cluster or a retained object; same-name Events with missing or different UIDs are excluded and counted in the contents note.
+
+Inline credential patterns are redacted from environment values, container commands and arguments (including init and ephemeral containers), and Event messages. Serialized spec copies in `cnpg.io/podSpec` and last-applied annotations are removed. Secret selector names remain available for diagnosis. Redaction operates on report copies.
+
+## Operator diagnosis
+
+The Operator screen leads with **Current state** (`cnpgOperatorState` in `operatorStatus.ts`): clusters whose phase says the operator is stuck on a CNPG-I plugin (unknown plugin, plugin error), each beside that plugin's Deployment readiness and restart history; components not ready or scaled to 0, a leader Lease not renewed, a `Fail` webhook with no ready endpoint, what it confirmed only from evidence it read and what it could not read, then each component's restart history — restarts since the Pod was created, with when the last one ended and why (`lastTermination`). Restart totals are cumulative, so only a restart that ended within the last hour is worded as current trouble; none is raised as a problem. The operator and plugin rows show the same history. It then adds, per operator Deployment: the leader-election Lease (`db9c8771.cnpg.io`: holder Pod, last renewal, leader changes, "not renewed" when older than the lease duration; "off" without `--leader-elect`), the namespaces it watches (`WATCH_NAMESPACE` on the container — the value that sets the operator's cache — with Clusters outside them named), its admission webhooks (both configurations present, failure policy, CA bundle set, and ready endpoints behind the webhook Service — none ready with `Fail` means every CNPG write is rejected), `controller_runtime_reconcile_errors_total` / `_total` per controller from its metrics port through `pods/proxy` (cumulative since the container last started; only the leader reconciles), recent events on its Deployment, ReplicaSets and Pods, and a link to its logs. Each fact carries its own access state.
+
+**Operator verdict on other pages.** `GET /api/cnpg/operator/status?namespaces=a,b` is the cheap form, memoized 10 s per identity: per namespace, whether the operator Deployment that watches it (`WATCH_NAMESPACE`) is leading — a ready Pod and a leader lease renewed within its duration — and whether a `Fail` webhook has no ready endpoint. States `reconciling | notReconciling | unknown | notWatched`; an unreadable lease or webhook is `unknown` with the grant, never "reconciling". An operator whose `WATCH_NAMESPACE` comes from a ConfigMap, Secret or field reference Radar does not resolve may or may not watch the namespace: it never makes the verdict `reconciling` (nor `notReconciling` while it could be leading); the verdict is `unknown` naming the reference, unless another operator confirmed to watch the namespace is leading. While it is `notReconciling` the fleet and every cluster page show a banner (status below may be stale) linking to Operator. Independently of the operator, when the instance Pods' Ready condition shows fewer ready than `status.readyInstances`, the fleet's Ready column and the Overview show the Pods' count with the status count beside it and raise an availability problem; when `status.currentPrimary` is not the Pod labelled primary, both are named (Overview, HA list, problem) instead of one winning silently.
+
+Capabilities carry the same verdict as `operator`. While the webhook rejects writes, actions that write through it — a new Backup, patches of the Cluster (restart, reload, fence, hibernate, maintenance) or of a ScheduledBackup — are refused with the reason; status-subresource patches (switchover, restarting the primary), Pod deletes and exec are not webhook-bound and stay allowed. Every dialog warns while the operator is observed not leading, and notes it when the state is unknown.
+
+Every request carries the facts the user reviewed (kube context, UIDs, current and target primary, fencing and hibernation values). The server re-reads them and returns **409** if anything changed; disruptive actions are never retried automatically. Errors carry a `code` (`context_changed`, `changed`, `blocked`, `all_fenced`, `operator_webhook_unavailable`, `outcome_unknown`, `partial`). `partial` means a multi-step action (destroy instance) stopped after some mutations took effect; the body's `completed` lists them (`deleted PVC x`, `detached PVC x`, `deleted Pod x`, `deleted Job x`) and the dialog keeps confirm locked, as for `outcome_unknown`. An outcome that is unknown after a timeout is resolved by re-reading the same Backup name, never by creating another.
+
+Cancel and terminate bind the instance Pod UID, the pid and the backend's `backend_start`; destroy binds the Pod UID and the name and UID of every PVC reviewed; pause/resume binds the reviewed `paused`. Each returns `target` identifying what it acted on, so an operation tracker can follow the outcome.
+
+The dialog (`ActionConfirmDialog` in k8s-ui) leads with the effect, keeps the literal API writes in an expandable section, and requires typing the name for disruptive actions (switchover, primary fence, lifting a fence, cluster restart, hibernate, cold backups).
+
+### Operation tracker
+
+A successful POST means the write was accepted, not that it happened. The dialog hands off a tracked operation (`trackCNPGOperation` in `web/src/components/cnpg/operations/`) bound to the kube context, the Cluster UID, the target (name + UID) and a baseline; the Cluster header shows it until it finishes. States: `requested → observed → progressing → completed | failed | stalled | superseded | unobservable`. A step Radar cannot see keeps the operation **unobservable**, never completed. Each source (capabilities facts, `/api/cnpg/workspace`, HA, runtime) carries when it was last fetched successfully and whether its latest fetch failed; a source not fetched successfully since the operation started, or whose refresh failed, is withheld from the observer, so a cached answer from before the action never certifies it; **stalled** needs telemetry showing no movement for 10 minutes. A newer conflicting operation, or the Cluster being recreated, supersedes it. Completion per kind: switchover — the target is `currentPrimary`, the old primary streams again, the `-rw` endpoints point at the target and the phase is healthy (unreadable endpoints leave that step unverified, so the switchover ends unobservable, never completed); restart — per instance, a recreated Pod, a restarted postgres container or a later `cnpg_pg_postmaster_start_time`; reload — no completion signal, said so; fence — the fence recorded on the Cluster; the operation then ends unobservable, because CloudNativePG reports no shutdown signal (`/pg/status` never carries `isFenced`, and a masked status error also arrives without WAL positions), so "PostgreSQL stopped" is never claimed; lift — Pod readiness and streaming; hibernate / rehydrate — the hibernation condition and ready instances; backup and schedule run — the Backup's phase; maintenance — the spec. Operations persist per browser session (`sessionStorage`, in memory when unavailable). New kinds register an observer with `registerCNPGOperationObserver(kind, fn)`.
+
+## Inside PostgreSQL
+
+**psql.** Every instance card on the Replication tab has **psql**, and the cluster menu has **Open psql on the primary**. It opens the dock terminal on the instance Pod, container `postgres`, running `psql` (`?shell=psql`, one argv element) over the caller's own `pods/exec` — the local socket as the `postgres` superuser, as `kubectl cnpg psql` does. The tab says whether the instance was the primary (read-write) or a standby (read-only while it stays one) when it opened. Disabled, with the grant named, without `create pods/exec`.
+
+**Blocking view.** Performance › Sessions keeps the exporter's aggregates and adds, for exec-capable callers, `GET /api/cnpg/clusters/{ns}/{name}/sessions`: a fixed query (never built from input) run by psql in the instance chosen by the Sessions section's one instance picker (the primary by default) with a 5 s statement timeout, excluding its own backend. It lists only backends in a blocking relation as blocker → victim trees (a victim of two blockers appears under both; a lock cycle is marked), with wait event, transaction/query/connection age, user, database, application, client address and query text cut at 200 characters. Query text is shown because the same grant lets the caller open psql and read it. The section shows one connections figure — "N in use" over "of M usable (max_connections X, R reserved)" from this read, or the exporter's count against `max_connections` with the reserve named as unread when exec is unavailable. Beside it: each instance's CPU and memory from metrics-server, so lock waits and resource pressure are told apart. Cancel is offered first (it does nothing for an idle-in-transaction holder, which the dialog says); terminate names the transaction it rolls back and requires typing the pid.
+
+## GitOps write guard
+
+Any write to an object a GitOps tool manages can be undone by the next sync. `POST /api/gitops/write-evidence` reports, for the paths an action writes, whether each is declared in `last-applied-configuration`, owned by a GitOps field manager (Argo CD, Flux, Helm), or covered by an ignore rule (Argo CD only with `RespectIgnoreDifferences`), plus the owner's sync policy (Argo CD self-heal and automated sync, Flux suspend and interval, HelmRelease drift detection). `evaluateGitOpsWriteGuard` in k8s-ui turns that into a warning, and `GitOpsWriteWarning` renders it with an acknowledgement when a revert is likely. The copy never promises a change will stick. Status writes are described as fields ordinary GitOps sync does not manage.
+
+The guard is shared: the CNPG actions, Set image and the investigation apply dialog use it (`useGitOpsWriteGuard` in `web/`). Other GitOps-aware writes (YAML force-apply, Helm and GitOps rollbacks) still carry their own warnings.
+
+Prometheus history starts after this Cluster’s creation and the query lookback, with a visible pending/clamp note. Fleet gauges wait through the default five-minute lookback; sustained lag/receiver checks also wait through their offset selector’s lookback. Six-hour volume growth waits for a full window after both the Cluster and its owned PVCs were created. These metrics do not carry the Cluster UID in their series, so reused names must not inherit the predecessor’s samples.
## API
-Every route is gated on the caller's own access, and a partial answer names what it withheld rather than shrinking silently.
+Every route is gated on the caller's own access, and a partial answer names what it withheld rather than shrinking silently. A missing grant is sent as an object, `{verb, group?, resource, subresource?, namespace?}` (no namespace means cluster-wide), and the UI words it with `formatGrant`. The core workspace routes are advertised by `cnpgWorkspace`; instance replacement uses `cnpgInstanceReplacement`, and the guided protection routes and verbs below use `cnpgProtectionSetup`; the UI gates every call on it, so a Radar that predates the workspace shows the standard views and an upgrade note instead of failed requests.
-- Workspace: `/api/cnpg/workspace` returns every CNPG kind (plus owner-validated instance Pods) with per-kind `coverage` (`full|partial|denied|notInstalled|syncing|error`, `partial` naming only in-scope denied namespaces), CNPG issues from the Issues engine and `cnpgNoDeclarativeBackup` audit findings, each withheld where the caller lacks coverage. Namespaced kinds follow the view filter and the capacity per-namespace `list` fallback; `ClusterImageCatalog` needs a cluster-scope `list`. Handler `internal/server/cnpg_workspace.go`
-- Operator: `/api/cnpg/operator` returns the operator Deployments (`app.kubernetes.io/name=cloudnative-pg`) and plugin Deployments (served by Services labelled `cnpg.io/pluginName`) with image-tag version and readiness (`null` when unreported), plus the operator's ConfigMap/Secret/monitoring-queries references from its args and env. ConfigMap data only with `get configmaps`; the Secret is name-only, never read. Deployments and Services carry the workspace `coverage` states; deliberately ignores the namespace view filter (the operator lives in its own namespace). Handler `internal/server/cnpg_operator.go`
+- Workspace: `/api/cnpg/workspace` returns every CNPG kind (plus owner-validated instance Pods, and in `jobPods` the Pods of the Jobs a Cluster controls, with `jobCoverage` saying where the caller's Jobs were read) with per-kind `coverage` (`full|partial|denied|notInstalled|syncing|uncached|error`; `partial` names in-scope namespaces left unread as `deniedNamespaces` (no access) or `uncachedNamespaces` (Radar's cache does not hold them), and only when the caller supplied the namespace list; `uncached` is a scope Radar's cache does not cover at all), CNPG issues from the Issues engine and `cnpgNoDeclarativeBackup` audit findings, each withheld where the caller lacks coverage, plus `scheduleReadings` (each readable ScheduledBackup's schedule worded as the operator parses it, keyed `ns/name`) and `managedBy` (each object's Argo CD, Flux or Helm manager, keyed `Kind/ns/name`). Namespaced kinds follow the view filter and the capacity per-namespace `list` fallback; `ClusterImageCatalog` needs a cluster-scope `list`. Handler `internal/server/cnpg_workspace.go`
+- Operator status: `/api/cnpg/operator/status?namespaces=` returns, per namespace, whether the watching operator is leading and whether a fail-closed webhook has no ready endpoint (`reconciling|notReconciling|unknown|notWatched`); capabilities embed the same verdict and refuse webhook-bound writes while it rejects them. Handler `internal/server/cnpg_operator_status.go`
+- Operator: `/api/cnpg/operator` returns the operator Deployments (`app.kubernetes.io/name=cloudnative-pg`) and plugin Deployments (served by Services labelled `cnpg.io/pluginName`) with image-tag version and readiness (`null` when unreported), plus the operator's ConfigMap/Secret/monitoring-queries references from its args and env. ConfigMap data only with `get configmaps`; the Secret is name-only, never read. Deployments and Services carry the workspace `coverage` states; deliberately ignores the namespace view filter (the operator lives in its own namespace). Each component with a Deployment adds `pods` (`{name, ready, restarts, startedAt, lastTermination?: {container, reason, exitCode, finishedAt}}` — restarts summed over containers, `startedAt` the component container's current start, `lastTermination` the `lastState.terminated` of the container that ended most recently; the diagnosis Pods carry the same `lastTermination`) and `podCoverage`, read with the caller's `list pods` in that namespace via the Deployment's selector; denied reads `podCoverage` with the grant and `pods: null`, never a 403. Handler `internal/server/cnpg_operator.go`. Additive `diagnosis[]` per operator Deployment (`cnpg_operator_diagnosis.go`), every read as the caller with its own coverage: leader Lease `db9c8771.cnpg.io` (holder Pod, renew, transitions, `stale` via `cnpgHALeaseFrom`; `disabled` without `--leader-elect`), watched namespaces from the container's `WATCH_NAMESPACE` env only (the ConfigMap value is read after the cache is built), the `cnpg-{mutating,validating}-webhook-configuration` objects (failurePolicy, caBundle set) + their Service's EndpointSlice readiness, `controller_runtime_reconcile_{errors_,}total` per controller from the operator's `metrics` port through `pods/proxy` (25s memo; counters since the container's last start), and recent events for its Deployment/ReplicaSets/Pods
- Catalog reverse-lookup: `/api/cnpg/imagecatalogs/{ns}/{name}/clusters` and `/api/cnpg/clusterimagecatalogs/{name}/clusters` return the Clusters pinned to an image catalog, with the major each asks for and the image it actually resolved. Cluster-scoped catalogs are referenceable from any namespace, so the cluster-scoped route reads cluster-wide gated on `list clusters` — a view-filtered answer would report "nothing uses this" before an edit.
-- Cluster logs: `/api/cnpg/clusters/{ns}/{name}/logs` (bounded snapshot) and `/logs/stream` (SSE, re-resolves instances every 5s) merge every instance Pod — label `cnpg.io/cluster` AND controller ownerRef to the Cluster's UID, never the label alone. Gated on `get clusters` + `list pods` + `get pods/log` before the Cluster lookup (404 after). `container` defaults to `postgres`, `tailLines` 200, `sinceTime` is converted to seconds and trimmed, `pod` must be a validated instance (400). Entries keep raw `content` and add `level`/`logger`/`message` parsed from the instance manager's JSON (`record.error_severity` wins over `level`).
-- Cluster activity: `/api/cnpg/clusters/{ns}/{name}/activity?since=&limit=` reads the timeline store for the Cluster, its instance Pods (by owner) and CNPG children attributed by the retained `cnpg.io/cluster` label — `pkg/timeline.ExtractLabels` records it from the label or `spec.cluster.name` on CNPG-group objects and Pods, so deleted children stay attributed. K8s Event rows join by subject UID. Rows of a kind the caller can't `list` in the namespace are dropped; `oldest` is the namespace's retention floor and `attributionSince` the earliest labelled row — history before it cannot attribute deleted children.
+- Runtime: `/api/cnpg/clusters/{ns}/{name}/runtime` (each instance's `/pg/status` on :8000 + exporter `/metrics` on :9187) and `/api/cnpg/poolers/{ns}/{name}/runtime` (PgBouncer exporter on :9127), read through the caller's impersonated `pods/proxy`. Fixed GET paths only, redirects refused (the :8000 server has state-changing endpoints — the fixed path is the security boundary); Pods validated by controller UID (pooler: Pooler→Deployment→ReplicaSet→Pod). Gated like logs minus `pods/log`; a `pods/proxy` denial is per-source `state:"denied"` in a 200, not a 403. Per-source `ok|partial|denied|unreachable|error` with `scheme`/`capturedAt`; absent measurements are omitted and listed in `missing`; memoized per (identity, context, Pod UID, endpoint) for 5s status / 25s metrics. Engine in `internal/server/cnpg_runtime.go`
+- Restore: `GET /api/cnpg/clusters/{ns}/{name}/recovery` is a restored Cluster's progress — phase, `bootstrap.recovery` source/target, the Pods of Cluster-owned Jobs (e.g. `-1-full-recovery`, with init containers) and instance Pods (controller-UID validated), those Jobs, and Warning events about them; each of pods/jobs/events carries `coverage` (`ok|denied(grant)|error`), plus the parsed `radar.skyhook.io/restore-validation` note. `POST .../restore-validation` `{reviewedContext, uid, params: {checked, targetTime?, source?}}` writes that annotation with an impersonated merge patch bound to the resourceVersion read; the server records `recordedBy` from the authenticated user and the source/target UIDs it read itself (never from the client), refuses a Cluster without `bootstrap.recovery` (409 `blocked`). The workspace's Restore validation fact reads it as "Validation recorded", never healthy. The restore itself is the create flow (`CreateResourceDialog`, strict create) fed by `web/src/components/cnpg/recovery/`, gated by `GET /api/cnpg/restore/capability?namespace=` (`{allowed, reason, permission, grant}` for `create clusters` in that namespace, refused while the operator's webhook rejects writes; the same verdict is the Cluster capabilities' `restore`); plugin Backups restore through the ObjectStore with `recoveryTarget.backupID` (+ `targetImmediate` for "end of this backup"). Handler `internal/server/cnpg_recovery.go`
+Named plugin-backup restores use the recorded backup ID and PostgreSQL major. The Barman plugin records the source Cluster UID but does not record the ObjectStore/server destination in the Backup. Radar rejects a recorded predecessor UID (or a backup predating the current Cluster), and asks the user to confirm that the current archive contains that backup. A source catalog uses the backup's major; a changed or unknown image major requires an explicit matching image. Physical recovery does not perform a major upgrade.
+
+- Report bundle: `GET /api/cnpg/clusters/{ns}/{name}/report[?logs=true&tailLines=N&queryText=true]` streams a zip mirroring `kubectl cnpg report cluster`: Cluster, owned Pods/Jobs/PVCs, events about them, Backups/ScheduledBackups/Poolers of the cluster, its ObjectStore(s), and the operator/runtime/storage endpoints' own JSON (captured by calling those handlers as the caller; operator ConfigMap values reduced to keys), plus `report.json` listing every item read or skipped with its grant/reason and the Secrets referenced (names only). Only the Cluster read can fail the request (403/404); every other read is SAR-gated, impersonated and skip-and-record. Inline connection passwords, initialization SQL, env values with secret-like names and credentials embedded in other env values are redacted; logs (`get pods/log`, all containers + previous runs, ≤2 MiB each) are opt-in and, without `queryText`, have `record.query`/`internal_query`/`context`, and SQL after `statement: ` / `execute …: ` / `parse …: ` / `bind …: ` / `plan:` / `parameters:` (JSON and plain lines) replaced; 32 MiB total cap, 40 s deadline. Handler `internal/server/cnpg_report.go`
+- Storage: `/api/cnpg/clusters/{ns}/{name}/storage` reports each instance's claims by `cnpg.io/pvcRole` (PG_DATA/PG_WAL/PG_TABLESPACE; kept only when the claim names an instance AND is owned by the Cluster's UID, others listed as `excluded`), requested vs capacity, resize state (claim conditions, `allocatedResourceStatuses`, the Cluster's `*PVC` status lists), StorageClass `allowVolumeExpansion` (absent when unreadable), used bytes from Prometheus `kubelet_volume_stats_*` via `prometheus.QueryPVCUsage` behind `prometheusAuthGate` (`get persistentvolumeclaims`), the spec field each size is declared in, ≥80%/≥90% disk findings, and per-instance WAL facts — size and segments, `readyWalFiles`, slot retention — side by side, never summed, from the same memoized pods/proxy reads as the runtime route. Gated on `get clusters`; claims need `list persistentvolumeclaims`, WAL needs `list pods` + `get pods/proxy`, each reported as its own coverage. `/api/cnpg/disk` is the fleet form: the fullest measured volume per visible Cluster, one claim list and one usage batch per namespace (≤64 namespaces). Handler `internal/server/cnpg_storage.go`
+- HA facts: `/api/cnpg/clusters/{ns}/{name}/ha` returns per-instance node/zone/QoS/running-vs-desired image/start times, the `FailoverQuorum` (recorded sync configuration with N, W, R and whether R + W > N — not a verdict), Cluster-owned PDBs, the primary Lease (1.30+) and the operator's leader Lease, Cluster-owned Jobs, ready `-rw` endpoints, certificate expiries with renewal owner (operator vs `spec.certificates`; cert-manager link from Secret **metadata only**) and `spec.nodeMaintenanceWindow`. Gated on `get clusters`; every sub-read is SAR-gated on its own (`get nodes` cluster-scoped, `get failoverquorums`, `list poddisruptionbudgets`, `get leases`, `list jobs`, `list endpointslices`, `get secrets`) and reports `ok|denied(grant)|notFound|notInstalled|unavailable|error`, so reading a Cluster never implies the rest. Handler `internal/server/cnpg_cluster_ha.go`
+- History: `/api/cnpg/clusters/{ns}/{name}/history?range=15m|1h|6h|24h` runs server-built Prometheus range queries only (no caller PromQL) over the CNPG exporter and kubelet volume stats, selecting instances by `namespace` + `pod=~"^-[0-9]+$"` (never the ambiguous `cluster` label) plus cluster-identity matchers when configured or proven by kube-state-metrics Pod UIDs; unproven identities that appear more than once ⇒ `state:"ambiguous"`. Per chart `ok|empty|noSeries|denied|ambiguous|scopeMismatch|error|notRead` with source, `steps`/`covered`, and `max by (pod)` before sums; ≤144 points, ≤12 series, 15 s memo. Gated on `get clusters` (403); CNPG series on `get pods`, the volume chart on `list` + `get persistentvolumeclaims`. No Prometheus ⇒ `source:"none"`. `/api/cnpg/fleet-metrics` is the fleet form: largest standby replay lag (`ok|noStandby|noSeries|denied|…`) and fastest volume growth per visible Cluster, one lag and one growth query per namespace (≤64). Under the same `get pods` gate, selector and isolation, a lag reading in state `ok` adds `standbys` (instances with `cnpg_pg_replication_in_recovery = 1`), `receiving` and `receiverDown` (pods) from `cnpg_pg_replication_is_wal_receiver_up` — never inferred from lag, which reads 0 for a standby that receives nothing — or `receiverUnknown` with `receiverReason` when no standby reports it, and each Cluster gets `slots` (`ok|noSeries|denied|ambiguous|scopeMismatch|error|notRead`; `inactive` is every physical slot with `cnpg_pg_replication_slots_active = 0` as each instance reports it, `{slot, pod, role?, bytes}` with role from that instance's recovery state and bytes from `cnpg_pg_replication_slots_pg_wal_lsn_diff`, most retained first, ≤10 plus `omitted`; standbys report CloudNativePG's synchronized HA slots, which are always inactive there), one receiver and one slots query per namespace. Activity accepts `until` alongside `since`. Engine in `internal/prometheus/cnpg_history.go`, handlers `internal/server/cnpg_history.go`
+- Cluster logs: `/api/cnpg/clusters/{ns}/{name}/logs` (bounded snapshot) and `/logs/stream` (SSE, re-resolves instances every 5s) merge every instance Pod — label `cnpg.io/cluster` AND controller ownerRef to the Cluster's UID, never the label alone. Gated on `get clusters` + `list pods` + `get pods/log` before the Cluster lookup (404 after). `container` defaults to `postgres`; the UI’s All containers sends `container=all`, reading main, init/sidecar and ephemeral containers with each line labelled by its container; a Waiting container with a retained terminated run contributes previous-run logs (labelled “previous run”), read once per retained run while streaming waits, while containers that never started remain excluded; `tailLines` 200, `sinceTime` is converted to seconds and trimmed, `untilTime` (needs `sinceTime`) trims the other end and reads each instance from the interval's start instead of its tail, `pod` must be a validated instance (400). Entries keep raw `content` and add `level`/`logger`/`message` parsed from the instance manager's JSON (`record.error_severity` wins over `level`).
+- Cluster activity: `/api/cnpg/clusters/{ns}/{name}/activity?since=&limit=` reads the timeline store for the Cluster, its instance Pods (by owner), live Pods of verified Cluster-controlled Jobs (Pod → Job → Cluster by controller name and UID, requiring the caller’s `list jobs`), and CNPG children attributed by the retained `cnpg.io/cluster` label — `pkg/timeline.ExtractLabels` records it from the label or `spec.cluster.name` on CNPG-group objects and Pods, so deleted children stay attributed. K8s Event rows join by subject UID. Rows of a kind the caller can't `list` in the namespace are dropped; `oldest` is the namespace's retention floor and `attributionSince` the earliest labelled CNPG child row, excluding instance Pods — history before it cannot attribute deleted children. Deleted Job Pods cannot be included through the live verified ownership chain.
+- Actions: `GET /api/cnpg/{clusters|scheduledbackups}/{ns}/{name}/capabilities` returns `{uid, resourceVersion, context, facts, actions}` — per action `{allowed, reason, permission: allowed|denied|unknown, grant}` from per-action state guards (no blanket "no ready instance") plus SAR (`create backups`, `patch clusters`, `patch clusters/status` separately, `delete pods`, `get pods`); Clusters add `instanceActions`, `restartPlan` and `hibernateEffects`. `POST .../actions/{action}` (cluster: backup, switchover, restart, restartInstance, reload, fence, unfence, hibernate, rehydrate, setMaintenance, unsetMaintenance, cancelBackend, terminateBackend, destroyInstance, configureArchiving; schedule: suspend, resume, run, repairMethod, setSchedule — `params.schedule`, binds `facts.schedule`, validated with the operator's cron parser, 400 `invalid_schedule`) is impersonated and **binds the reviewed facts**: body `{reviewedContext, uid, facts, params}`; context ≠ active ⇒ 409 `context_changed`, uid/bound-fact/target-Pod-UID mismatch or apiserver conflict ⇒ 409 `changed` with `current` facts, never retried. Writes mirror `kubectl cnpg` (status merge patches carry the Ready condition and `metadata.resourceVersion`; standby restart deletes the Pod with a UID precondition; `cnpg.io/fencedInstances` is a JSON-array string, lifting one instance under `["*"]` ⇒ 409 `all_fenced` unless `convertFromAll` + `remaining`). Backup confirmation binds the Cluster spec generation, so destination/plugin changes require review again while status-only updates do not. A manual ScheduledBackup run also binds its source Cluster UID and generation; follow-through uses that UID to reject a same-name replacement. A timed-out Backup create is resolved by re-reading the same name, never a second create. A multi-step action that stops after some writes took effect ⇒ `partial` with `completed: [...]`; the UI locks confirm for `partial` and `outcome_unknown` (`cnpgActionOutcomeLocked`). Handler `internal/server/cnpg_actions.go`. `GET /api/cnpg/scheduledbackups/{ns}/{name}/schedule-preview?schedule=` reads a draft schedule the way the operator would (plain-language reading, next 3 runs counted from `status.lastCheckTime`, `runsImmediately`); the capabilities `facts.preview` carries the same for the current one (`internal/server/cnpg_schedule.go`)
+- Sessions (blocking view): `GET /api/cnpg/clusters/{ns}/{name}/sessions?pod=` runs fixed SQL (`pg_stat_activity` + `pg_blocking_pids`, own backend excluded, `statement_timeout` 5s) with `psql` in the instance's `postgres` container over the caller's impersonated `pods/exec` (default: the primary; `pod` must be a validated instance). Returns only backends in a blocking relation (capped 200, `truncated`/`involvedTotal`), query text cut at 200 chars, `maxConnections`/`superuserReservedConnections`/`clientBackends` for headroom, and each instance's postgres-container requests/limits. Gated like runtime plus `create pods/exec`; without it a 200 with `state:"denied"`. Exec bounded: 4 concurrent, 10s, 1 MiB stdout. Cluster actions `cancelBackend` / `terminateBackend` (params `{pod, podUID, pid, backendStart}`) signal only when pid AND backend_start still match in the same statement (user values travel as psql `-v` variables, never in the SQL text) ⇒ else 409 `changed`. `destroyInstance` (params `{pod, podUID, keepPVC, pvcs:[{name,uid}], jobs:[{name,uid}]}`) mirrors `kubectl cnpg destroy`: PVCs detached (keep) or deleted with UID preconditions first, then the Pod, then reviewed Jobs labelled `cnpg.io/instanceName` and owned by the current Cluster UID, with UID and resourceVersion preconditions; refuses the primary and any instance not fenced (a fenced instance is never promoted), re-checking current/target primary, phase, the fence and the Pod's role label before every step, then lifts the destroyed name from a list-form fence (`["*"]` untouched); binds `fencedInstances`; reviewed PVC and Job sets must match before any write ⇒ else 409. Kept PVCs record the detaching Cluster UID; ownerless detached PVCs without that provenance are excluded. Keep refuses PVCs with additional owners, since they could still be garbage-collected. A Job delete conflict permits one bounded retry after re-listing verifies the same UID and ownership; replacement or ownership loss stops the sequence. `GET .../instances/{pod}/destroy-plan` returns the PVCs/Jobs and `delete`/`keep` verdicts. Cluster capabilities add `psql` / `destroyInstance` (and per instance `psql` / `destroy`) and `restore` (`create clusters` in the Cluster's namespace); psql itself is the dock terminal with `?shell=psql`. Handlers `internal/server/cnpg_sessions.go`, `cnpg_destroy.go`
+- Sessions without a primary instance Pod return 200 with `state:"unavailable"` and the prerequisite "Available once the primary is running"; an exec denial takes precedence. When the runtime inventory has no instance, Performance owns one shared prerequisite above Sessions and Blocking, linking to the same Cluster’s Overview and preserving its context and return state. Blocking independently names an exec denial. Logs without an instance source show the verified initdb problem, when one is known, with the same Overview destination. CNPG opts into disabling source-dependent log controls until a source exists; other shared viewer consumers can continue streaming while waiting for Pods.
+- Inspection (read-only, same exec gate and bounds as Sessions; a 200 with `state:"denied"` without `create pods/exec`): `GET /api/cnpg/clusters/{ns}/{name}/restore-checks` (400 unless `bootstrap.recovery`) runs fixed SQL on the primary — `pg_is_in_recovery()`, the timeline (`pg_control_checkpoint()`), its history file (`pg_read_file(... , missing_ok)`, 16 KiB), databases with sizes and non-system roles (each capped 50) — then, in the bootstrap database (`bootstrap.recovery.database`, default `app`, connected to only when it is a plain name and in the list read), user tables, `reltuples` estimates (a table without one — `-1`, or `0` before PostgreSQL 14 — is counted apart, never as zero rows) and the five largest. `GET .../parameters` reads, on every instance, the names in `spec.postgresql.parameters` (capped 200; names passed as a psql `-v` variable; a name that is not a parameter name is skipped): `current_setting`, `source`, `context`, `pending_restart`. That read sets only `search_path` in its session, so `statement_timeout` reads the server's value; a parameter the connection itself sets (`application_name`, and that pinned `search_path`) has no value. Every diagnostic SQL runs as the `postgres` superuser, so each starts by pinning `search_path` to `pg_catalog` (an application schema's function with an exact-match signature would otherwise be chosen over a built-in), and the ones that only read run with `default_transaction_read_only`. Handler `internal/server/cnpg_inspect.go`
+- Guided protection (`cnpgProtectionSetup`): `POST /api/cnpg/clusters/{ns}/{name}/protection/preview` accepts `{reviewedContext, objectStore, serverName}` and dry-runs attachment as the caller; `POST .../actions/configureArchiving` uses the existing action envelope with reviewed configuration digests and `{objectStore, serverName, acknowledgeArchive}`. `GET /api/cnpg/scheduledbackups/{ns}/{name}/method-preview` dry-runs the narrow repair, and `POST .../actions/repairMethod` binds its reviewed Schedule/Cluster facts. `GET /api/cnpg/clusters/{ns}/{name}/schedule-preview?schedule=` reads the target Cluster and previews a new schedule from now. Handlers are thin adapters in `internal/server/cnpg_protection.go`; orchestration and conditional writes live in `internal/cnpg/protection.go`, `schedule_repair.go` and `schedule.go`. Preview success does not verify storage uploads.
+- Pooler actions: `GET /api/cnpg/poolers/{ns}/{name}/capabilities` (facts: desired `paused`, pool mode, `spec.pgbouncer.parameters`, the Pooler-controlled Deployment's readiness and Service; actions `pause`/`resume` (`patch poolers`) and `observeState` (`create pods/exec`)), `POST .../actions/{pause|resume}` (merge patch `spec.pgbouncer.paused` with resourceVersion, binds reviewed `facts.paused` ⇒ 409 `changed`), `GET .../pgbouncer-state` (each pooler Pod's `SHOW STATE` over exec — the only observed pause; the exporter does not publish it). Handler `internal/server/cnpg_pooler_actions.go`
## Testing
`make cnpg-demo` (read `scripts/cnpg-demo/README.md` first) produces WAL archiving failure, failed and unrecognised-phase Backups, failing declarations, a Pooler, both catalog kinds and an ObjectStore with a failing server — every state the workspace distinguishes, except successful restores.
+
+The shared CNPG components use plain facts and neutral literal controller phases on every surface. Tab verdicts use `alwaysShow` only where the verdict belongs even when healthy. Expanded HA repeats its readiness summary and lists expected instances with no observed Pod as not running. With no instance, Performance shares one primary prerequisite and Overview link above Sessions and Blocking. Blocking still independently names an exec denial; Sessions does not blame unavailable proxy aggregates on exec. Resource usage follows the blocking verdict under its own heading.
diff --git a/docs/gitops.md b/docs/gitops.md
index acfac9ef79..26fde4fdbb 100644
--- a/docs/gitops.md
+++ b/docs/gitops.md
@@ -1,6 +1,6 @@
# GitOps (Argo CD & Flux)
-Radar's GitOps workspace gives Argo CD and Flux first-class treatment. Instead of treating Applications and Kustomizations as generic CRDs, you get a typed fleet view, a per-app detail page that diagnoses *why* something is misbehaving, and the controls you'd otherwise reach for `argocd` / `flux` CLI to run.
+Radar's GitOps view gives Argo CD and Flux first-class treatment. Instead of treating Applications and Kustomizations as generic CRDs, you get a typed fleet view, a per-app detail page that diagnoses *why* something is misbehaving, and the controls you'd otherwise reach for `argocd` / `flux` CLI to run.
The hard part of GitOps tooling isn't sync — it's diagnosis. Radar surfaces drift, recent events, controller-failure attribution, and lifecycle state inline so you don't have to context-switch between `kubectl get`, `argocd app diff`, controller logs, and a YAML viewer to understand a stuck reconcile.
diff --git a/docs/helm.md b/docs/helm.md
index 01df2f0051..63cb4a803b 100644
--- a/docs/helm.md
+++ b/docs/helm.md
@@ -25,7 +25,7 @@ The drawer includes:
- Resources: live status for resources rendered by the current release.
- Hooks: hook events, path, weight, status, run times, delete policies, output-log policies, and diagnostics for failed/running hooks.
-Compare opens a full-page workspace instead of rendering inside the drawer. The drawer links to Compare from history rows and operation banners when Radar can identify a useful revision pair.
+Compare opens as a full page instead of rendering inside the drawer. The drawer links to Compare from history rows and operation banners when Radar can identify a useful revision pair.
## Operation Insight
diff --git a/docs/integrations.md b/docs/integrations.md
index 5c5419a9a2..3ee4e5bc06 100644
--- a/docs/integrations.md
+++ b/docs/integrations.md
@@ -851,7 +851,7 @@ The source contract is Strimzi's [KafkaConnector status schema](https://strimzi.
[CloudNativePG](https://cloudnative-pg.io/) (CNPG) is the Kubernetes operator for PostgreSQL, covering the full lifecycle from bootstrapping to monitoring, with high availability, automated failover, and backup management.
-Beyond the per-kind views below, the CloudNativePG **workspace** (`/cnpg`) composes them into fleet, protection, declaration, pooling and operator screens — see [cnpg.md](cnpg.md).
+Beyond the per-kind views below, the CloudNativePG **workspace** (`/cnpg`) composes them into fleet, protection, declaration, pooling and operator screens. Cluster creation and restore share target inputs; Backups guides attachment to existing ObjectStores and matching schedule creation/repair, followed by observed upload/backup evidence. It does not provision storage providers or credentials — see [cnpg.md](cnpg.md).
### What Radar Shows
@@ -886,13 +886,13 @@ Beyond the per-kind views below, the CloudNativePG **workspace** (`/cnpg`) compo
- PgBouncer parameters
- Degraded state detection (AlertBanner when not all instances are scheduled)
-Note `Pooler.status.instances` counts pods *trying to be scheduled*, not ready pods — a Pooler whose PgBouncer pods are all Pending still reports the full count. Radar therefore labels the healthy state **Scheduled** rather than Ready; actual readiness lives on the Deployment CNPG generates for the Pooler (same name, same namespace).
+Note `Pooler.status.instances` counts pods *trying to be scheduled*, not ready pods — a Pooler whose PgBouncer pods are all Pending still reports the full count. Radar labels the declared `spec.instances` count neutrally as **N instances requested**; actual readiness lives on the Deployment CNPG generates for the Pooler (same name, same namespace).
**Resource Browser:** Smart columns show status, instance counts (with degraded highlighting), primary instance, image tag, storage size, cluster reference, and schedule expressions.
### Phase classification
-Cluster phases are full English sentences (`Cluster is unrecoverable and needs manual intervention`), not enum tokens, and are matched on equality. They are bucketed as healthy / transient / failing / terminal / attention; terminal phases outrank instance counts, so an unrecoverable cluster whose pods happen to still be Ready is still rendered red. An unrecognized phase from a newer CNPG minor surfaces verbatim as unknown rather than being guessed at.
+Cluster phases are full English sentences (`Cluster is unrecoverable and needs manual intervention`), not enum tokens, and are matched on equality. They are bucketed as healthy / transient / failing / terminal / attention; terminal phases outrank instance counts, so an unrecoverable cluster whose pods happen to still be Ready is still rendered red. "Terminal" means reconciliation is blocked, not that the operator gave up: it retries every one of them, the plugin phases clear on their own once the plugin loads and answers, and the banner quotes the operator's `status.phaseReason` (`cnpgBlockedPhaseExplanation`). An unrecognized phase from a newer CNPG minor surfaces verbatim as unknown rather than being guessed at.
Backup phases are lowercase tokens. `walArchivingFailing` is treated as a cluster-level signal, not an ordinary backup failure — archiving is broken upstream of that Backup, so the whole recovery window is affected.
@@ -906,7 +906,7 @@ The `ObjectStore` itself is rendered: destination and credential provider (never
Two states carry the weight. A failure newer than the last success means the window has stopped advancing while its oldest point still ages out under retention — shrinking from both ends, so it is called out rather than left to be inferred from two timestamps. An ObjectStore with an empty `serverRecoveryWindow` is reported as holding nothing restorable rather than as healthy: on the plugin path the Cluster publishes no recovery point of its own, so a green badge here would be the only claim on screen and it would be wrong.
-`Backup` and `ScheduledBackup` with `spec.method: plugin` name the plugin and link to the ObjectStore they write into, and suppress the in-tree `destinationPath` / `serverName` rows, which are never populated on that path.
+`Backup` and `ScheduledBackup` with `spec.method: plugin` name the plugin and suppress the in-tree `destinationPath` / `serverName` rows, which are never populated on that path. For barman-cloud, their destination comes from the Cluster’s plugin entry; their own plugin parameters are ignored. The standalone renderers explain this without presenting a parameter as a destination link. Workspace destination links are inferred from the current Cluster configuration. Radar does not model third-party plugin destinations.
### Declarative objects: Database, Publication, Subscription
@@ -944,6 +944,7 @@ Deliberately narrow: the absence of a ScheduledBackup does not prove a cluster i
| Database | `postgresql.cnpg.io/v1` | — | Yes | — |
| Publication | `postgresql.cnpg.io/v1` | — | Yes | — |
| Subscription | `postgresql.cnpg.io/v1` | — | Yes | — |
+| DatabaseRole | `postgresql.cnpg.io/v1` | — | Yes | — |
| ImageCatalog | `postgresql.cnpg.io/v1` | — | Yes | — |
| ClusterImageCatalog | `postgresql.cnpg.io/v1` | — | Yes | — |
diff --git a/go.mod b/go.mod
index 2ded7e0512..40bccdad70 100644
--- a/go.mod
+++ b/go.mod
@@ -23,6 +23,7 @@ require (
github.com/prometheus/client_golang v1.24.1
github.com/prometheus/client_model v0.6.3
github.com/prometheus/common v0.72.0
+ github.com/robfig/cron v1.2.0
github.com/skyhook-io/radar/pkg v0.0.0
github.com/wailsapp/wails/v2 v2.16.0
golang.org/x/net v0.59.0
diff --git a/go.sum b/go.sum
index 31eca8a7ba..adc0e06a99 100644
--- a/go.sum
+++ b/go.sum
@@ -332,6 +332,8 @@ github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qq
github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc=
github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ=
github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88=
+github.com/robfig/cron v1.2.0 h1:ZjScXvvxeQ63Dbyxy76Fj3AT3Ut0aKsyd2/tl3DTMuQ=
+github.com/robfig/cron v1.2.0/go.mod h1:JGuDeoQd7Z6yL4zQhZ3OPEVHB7fL6Ka6skscFHfmt2k=
github.com/rogpeppe/go-internal v1.14.1 h1:UQB4HGPB6osV0SQTLymcB4TgvyWu6ZyliaW0tI/otEQ=
github.com/rogpeppe/go-internal v1.14.1/go.mod h1:MaRKkUm5W0goXpeCfT7UZI6fk/L7L7so1lCWt35ZSgc=
github.com/rubenv/sql-migrate v1.8.1 h1:EPNwCvjAowHI3TnZ+4fQu3a915OpnQoPAjTXCGOy2U0=
diff --git a/internal/auth/auth.go b/internal/auth/auth.go
index a15e2c4929..f660abaa9d 100644
--- a/internal/auth/auth.go
+++ b/internal/auth/auth.go
@@ -38,6 +38,7 @@ var (
DiscoverNamespaces = pkgauth.DiscoverNamespaces
SubjectCanI = pkgauth.SubjectCanI
SubjectCanISubresource = pkgauth.SubjectCanISubresource
+ SubjectCanINamed = pkgauth.SubjectCanINamed
FilterNamespacesForUser = pkgauth.FilterNamespacesForUser
NewSessionID = pkgauth.NewSessionID
CreateSessionCookie = pkgauth.CreateSessionCookie
diff --git a/internal/auth/grant.go b/internal/auth/grant.go
new file mode 100644
index 0000000000..99bdd146e9
--- /dev/null
+++ b/internal/auth/grant.go
@@ -0,0 +1,43 @@
+package auth
+
+// Grant is one Kubernetes RBAC permission a read or an action needs, as the
+// caller would have to be given it. An empty Namespace means cluster-wide.
+type Grant struct {
+ Verb string `json:"verb"`
+ Group string `json:"group,omitempty"`
+ Resource string `json:"resource"`
+ Subresource string `json:"subresource,omitempty"`
+ Namespace string `json:"namespace,omitempty"`
+}
+
+// In returns g bound to namespace ("" for cluster-wide).
+func (g Grant) In(namespace string) Grant {
+ g.Namespace = namespace
+ return g
+}
+
+// Ref returns a copy of g for an optional field.
+func (g Grant) Ref() *Grant {
+ return &g
+}
+
+// String words the grant for a person: "patch clusters/status
+// (postgresql.cnpg.io) in namespace pg", or for a cluster-wide grant "get
+// pods/proxy cluster-wide", where the group is written resource.group so the
+// text reads without nested parentheses when quoted inside "(needs …)".
+func (g Grant) String() string {
+ res := g.Resource
+ if g.Subresource != "" {
+ res += "/" + g.Subresource
+ }
+ if g.Namespace == "" {
+ if g.Group != "" {
+ res += "." + g.Group
+ }
+ return g.Verb + " " + res + " cluster-wide"
+ }
+ if g.Group != "" {
+ res += " (" + g.Group + ")"
+ }
+ return g.Verb + " " + res + " in namespace " + g.Namespace
+}
diff --git a/internal/auth/grant_test.go b/internal/auth/grant_test.go
new file mode 100644
index 0000000000..a704f4ce70
--- /dev/null
+++ b/internal/auth/grant_test.go
@@ -0,0 +1,46 @@
+package auth
+
+import (
+ "encoding/json"
+ "testing"
+)
+
+func TestGrantString(t *testing.T) {
+ for _, c := range []struct {
+ grant Grant
+ want string
+ }{
+ {Grant{Verb: "patch", Group: "postgresql.cnpg.io", Resource: "clusters", Subresource: "status", Namespace: "pg"}, "patch clusters/status (postgresql.cnpg.io) in namespace pg"},
+ {Grant{Verb: "create", Group: "postgresql.cnpg.io", Resource: "backups", Namespace: "db"}, "create backups (postgresql.cnpg.io) in namespace db"},
+ {Grant{Verb: "get", Resource: "pods", Subresource: "proxy", Namespace: "pgrt"}, "get pods/proxy in namespace pgrt"},
+ {Grant{Verb: "get", Resource: "pods", Subresource: "proxy"}, "get pods/proxy cluster-wide"},
+ {Grant{Verb: "get", Resource: "nodes"}, "get nodes cluster-wide"},
+ {Grant{Verb: "get", Group: "admissionregistration.k8s.io", Resource: "mutatingwebhookconfigurations"}, "get mutatingwebhookconfigurations.admissionregistration.k8s.io cluster-wide"},
+ } {
+ if got := c.grant.String(); got != c.want {
+ t.Errorf("%+v.String() = %q, want %q", c.grant, got, c.want)
+ }
+ }
+}
+
+func TestGrantInBindsWithoutMutating(t *testing.T) {
+ template := Grant{Verb: "list", Resource: "pods"}
+ bound := template.In("pg")
+ if template.Namespace != "" || bound.Namespace != "pg" || bound.In("").Namespace != "" {
+ t.Errorf("template=%+v bound=%+v", template, bound)
+ }
+}
+
+func TestGrantJSON(t *testing.T) {
+ b, err := json.Marshal(Grant{Verb: "get", Resource: "pods", Subresource: "proxy", Namespace: "pg"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got := string(b); got != `{"verb":"get","resource":"pods","subresource":"proxy","namespace":"pg"}` {
+ t.Errorf("namespaced = %s", got)
+ }
+ b, _ = json.Marshal(Grant{Verb: "patch", Group: "postgresql.cnpg.io", Resource: "clusters"})
+ if got := string(b); got != `{"verb":"patch","group":"postgresql.cnpg.io","resource":"clusters"}` {
+ t.Errorf("cluster-wide = %s", got)
+ }
+}
diff --git a/internal/cnpg/actions.go b/internal/cnpg/actions.go
new file mode 100644
index 0000000000..89a8182f82
--- /dev/null
+++ b/internal/cnpg/actions.go
@@ -0,0 +1,1920 @@
+package cnpg
+
+import (
+ "bytes"
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "net"
+ "net/http"
+ "regexp"
+ "sort"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/api/meta"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/apimachinery/pkg/util/validation"
+ "k8s.io/client-go/dynamic"
+ "k8s.io/client-go/kubernetes"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+var (
+ ClusterGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "clusters"}
+ cnpgBackupGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "backups"}
+ ScheduleGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "scheduledbackups"}
+ cnpgPoolerGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "poolers"}
+ cnpgDatabaseGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "databases"}
+ cnpgPublGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "publications"}
+ cnpgSubscrGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "subscriptions"}
+ cnpgScheduleRunR = regexp.MustCompile(`-\d{14}$`)
+)
+
+const (
+ cnpgRestartAnnotation = "kubectl.kubernetes.io/restartedAt"
+ cnpgReloadAnnotation = "cnpg.io/reloadedAt"
+ cnpgHibernateAnnotation = "cnpg.io/hibernation"
+ cnpgFencedAnnotation = "cnpg.io/fencedInstances"
+ cnpgRequestedFromAnno = "radar.skyhook.io/requested-from"
+ cnpgAllInstances = "*"
+
+ cnpgPhaseHealthy = "Cluster in healthy state"
+ cnpgPhaseSwitchover = "Switchover in progress"
+ cnpgPhaseFailover = "Failing over"
+ cnpgPhaseWaitingForUser = "Waiting for user action"
+ cnpgPhaseInplaceRestart = "Primary instance is being restarted in-place"
+ cnpgInplaceReason = "Requested by the user"
+
+ // The compact UTC stamp CloudNativePG itself uses for scheduled runs.
+ compactStamp = "20060102150405"
+
+ // Refusal codes beyond the shared action contract's.
+ cnpgCodeAllFenced = "all_fenced"
+ CodeWebhook = "operator_webhook_unavailable"
+ cnpgCodeAmbiguous = "outcome_unknown"
+ cnpgCodeInvalidSchedule = "invalid_schedule"
+)
+
+// CNPGFencedFacts is the parsed cnpg.io/fencedInstances annotation. Raw is the
+// annotation exactly as stored ("" when unset) and is what a confirmation binds.
+type CNPGFencedFacts = cnpg.FencedInstances
+
+// CNPGInstanceFact is one instance as the Cluster names it, joined to its Pod
+// when the caller can read it. PodUID is empty when the Pod is missing or
+// unreadable; PodReadable distinguishes the two.
+type CNPGInstanceFact struct {
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+ Role string `json:"role"`
+ Ready bool `json:"ready"`
+ Healthy bool `json:"healthy"`
+ Fenced bool `json:"fenced"`
+ PodReadable bool `json:"podReadable"`
+ PodExists bool `json:"podExists"`
+}
+
+// CNPGBackupMethodFact is one backup method the Cluster declares. Capability
+// is "backup" when the method can take a backup, "unknown" for a plugin that
+// has not reported its capabilities, "none" for a plugin that reports it
+// cannot.
+type CNPGBackupMethodFact struct {
+ Method string `json:"method"`
+ PluginName string `json:"pluginName,omitempty"`
+ Capability string `json:"capability"`
+ Reason string `json:"reason,omitempty"`
+ Deprecated bool `json:"deprecated,omitempty"`
+}
+
+// CNPGClusterFacts is the state the confirm dialog shows and the POST binds.
+type CNPGClusterFacts struct {
+ Generation int64 `json:"generation"`
+ CurrentPrimary string `json:"currentPrimary"`
+ TargetPrimary string `json:"targetPrimary"`
+ Phase string `json:"phase"`
+ PhaseReason string `json:"phaseReason,omitempty"`
+ Hibernation string `json:"hibernation"`
+ Hibernated bool `json:"hibernated"`
+ FencedInstances CNPGFencedFacts `json:"fencedInstances"`
+ Instances []CNPGInstanceFact `json:"instances"`
+ BackupMethods []CNPGBackupMethodFact `json:"backupMethods"`
+ BackupTarget string `json:"backupTarget,omitempty"`
+ // ArchivingFailing: the ContinuousArchiving condition is False. An
+ // in-tree Barman backup then ends in phase walArchivingFailing.
+ ArchivingFailing bool `json:"archivingFailing"`
+ IsReplicaCluster bool `json:"isReplicaCluster"`
+ Terminating bool `json:"terminating"`
+ Maintenance CNPGMaintenanceFacts `json:"maintenance"`
+ // ColdSnapshotBackup names a running offline volume-snapshot Backup of
+ // the cluster: the operator fenced its instance and lifts the fence itself.
+ ColdSnapshotBackup string `json:"coldSnapshotBackup,omitempty"`
+}
+
+type CNPGClusterActions struct {
+ Backup integration.ActionCapability `json:"backup"`
+ Switchover integration.ActionCapability `json:"switchover"`
+ Restart integration.ActionCapability `json:"restart"`
+ RestartInstance integration.ActionCapability `json:"restartInstance"`
+ Reload integration.ActionCapability `json:"reload"`
+ Fence integration.ActionCapability `json:"fence"`
+ Unfence integration.ActionCapability `json:"unfence"`
+ Hibernate integration.ActionCapability `json:"hibernate"`
+ Rehydrate integration.ActionCapability `json:"rehydrate"`
+ // Psql is opening psql on the primary; DestroyInstance is the first
+ // instance that may be destroyed (per-instance verdicts are authoritative).
+ Psql integration.ActionCapability `json:"psql"`
+ DestroyInstance integration.ActionCapability `json:"destroyInstance"`
+ // Restore is creating a new Cluster in this namespace that bootstraps
+ // from this one's backups; the source is only read.
+ Restore integration.ActionCapability `json:"restore"`
+ CNPGMaintenanceActions
+}
+
+// CNPGInstanceActions is the per-row verdict for actions that name one
+// instance: a primary restarts in place (status write) where a standby's Pod
+// is deleted, so the two carry different guards and grants.
+type CNPGInstanceActions struct {
+ Restart integration.ActionCapability `json:"restart"`
+ SwitchoverTarget integration.ActionCapability `json:"switchoverTarget"`
+ Fence integration.ActionCapability `json:"fence"`
+ Unfence integration.ActionCapability `json:"unfence"`
+ Psql integration.ActionCapability `json:"psql"`
+ Destroy integration.ActionCapability `json:"destroy"`
+}
+
+// CNPGRestartStep is one instance's fate in a cluster restart, in the order
+// the operator works through them: standbys first, the primary last.
+type CNPGRestartStep struct {
+ Instance string `json:"instance"`
+ Role string `json:"role"`
+ // Effect: recreate | skipped_fenced | switchover | restart | wait_for_user | restart_only_instance
+ Effect string `json:"effect"`
+}
+
+type CNPGRestartPlan struct {
+ PrimaryUpdateStrategy string `json:"primaryUpdateStrategy"`
+ PrimaryUpdateMethod string `json:"primaryUpdateMethod"`
+ Steps []CNPGRestartStep `json:"steps"`
+}
+
+// CNPGEffectList names the objects a hibernation touches. Available is false
+// when they could not be read; an unreadable list is not an empty one.
+type CNPGEffectList struct {
+ Available bool `json:"available"`
+ Reason string `json:"reason,omitempty"`
+ Names []string `json:"names"`
+}
+
+type CNPGVolumeFact struct {
+ Name string `json:"name"`
+ Instance string `json:"instance,omitempty"`
+ Role string `json:"role,omitempty"`
+ Capacity string `json:"capacity,omitempty"`
+ Requested string `json:"requested,omitempty"`
+}
+
+type CNPGVolumeEffect struct {
+ Available bool `json:"available"`
+ Reason string `json:"reason,omitempty"`
+ Items []CNPGVolumeFact `json:"items"`
+}
+
+type CNPGHibernateEffects struct {
+ Poolers CNPGEffectList `json:"poolers"`
+ UnsuspendedScheduledBackups CNPGEffectList `json:"unsuspendedScheduledBackups"`
+ Databases CNPGEffectList `json:"databases"`
+ Publications CNPGEffectList `json:"publications"`
+ Subscriptions CNPGEffectList `json:"subscriptions"`
+ Volumes CNPGVolumeEffect `json:"volumes"`
+}
+
+// CNPGClusterCapabilitiesResponse is GET /api/cnpg/clusters/{ns}/{name}/capabilities.
+type CNPGClusterCapabilitiesResponse struct {
+ UID string `json:"uid"`
+ ResourceVersion string `json:"resourceVersion"`
+ Context string `json:"context"`
+ Facts CNPGClusterFacts `json:"facts"`
+ Actions CNPGClusterActions `json:"actions"`
+ InstanceActions map[string]CNPGInstanceActions `json:"instanceActions"`
+ RestartPlan CNPGRestartPlan `json:"restartPlan"`
+ HibernateEffects CNPGHibernateEffects `json:"hibernateEffects"`
+ // Operator is whether the operator watching this namespace is acting on
+ // it. Writes to the Cluster or new Backups are refused above when its
+ // webhook rejects them; otherwise the dialogs warn.
+ Operator CNPGOperatorVerdict `json:"operator"`
+}
+
+type CNPGScheduleFacts struct {
+ // Generation binds "run" to the settings the user reviewed: every spec
+ // change bumps it.
+ Generation int64 `json:"generation"`
+ Cluster string `json:"cluster"`
+ ClusterUID string `json:"clusterUID"`
+ ClusterGeneration int64 `json:"clusterGeneration"`
+ Suspended bool `json:"suspended"`
+ NextScheduleTime string `json:"nextScheduleTime,omitempty"`
+ Method string `json:"method,omitempty"`
+ PluginName string `json:"pluginName,omitempty"`
+ Target string `json:"target,omitempty"`
+ // ClusterState: ok | missing | hibernated | unreadable
+ ClusterState string `json:"clusterState"`
+ BackupBlockedReason string `json:"backupBlockedReason,omitempty"`
+ Terminating bool `json:"terminating"`
+ // CatchUp is set when resuming would create one backup right away: the
+ // next run the operator computed is already in the past.
+ CatchUp bool `json:"catchUp"`
+ // Schedule is spec.schedule verbatim; setSchedule binds it.
+ Schedule string `json:"schedule"`
+ // Preview explains Schedule and when the operator runs it next.
+ Preview CNPGSchedulePreview `json:"preview"`
+}
+
+type CNPGScheduleActions struct {
+ Suspend integration.ActionCapability `json:"suspend"`
+ Resume integration.ActionCapability `json:"resume"`
+ Run integration.ActionCapability `json:"run"`
+ SetSchedule integration.ActionCapability `json:"setSchedule"`
+}
+
+// CNPGScheduleCapabilitiesResponse is GET /api/cnpg/scheduledbackups/{ns}/{name}/capabilities.
+type CNPGScheduleCapabilitiesResponse struct {
+ UID string `json:"uid"`
+ ResourceVersion string `json:"resourceVersion"`
+ Context string `json:"context"`
+ Facts CNPGScheduleFacts `json:"facts"`
+ Actions CNPGScheduleActions `json:"actions"`
+ Operator CNPGOperatorVerdict `json:"operator"`
+}
+
+// cnpgReviewedFacts is the subset of facts a confirmation can bind. Decoded
+// leniently: a client may echo the whole facts object it was served.
+type cnpgReviewedFacts struct {
+ CurrentPrimary *string `json:"currentPrimary"`
+ TargetPrimary *string `json:"targetPrimary"`
+ Hibernation *string `json:"hibernation"`
+ BackupTarget *string `json:"backupTarget"`
+ FencedInstances *struct {
+ Raw *string `json:"raw"`
+ } `json:"fencedInstances"`
+ Suspended *bool `json:"suspended"`
+ Schedule *string `json:"schedule"`
+ Generation *int64 `json:"generation"`
+ ClusterUID *string `json:"clusterUID"`
+ ClusterGeneration *int64 `json:"clusterGeneration"`
+ Maintenance *CNPGMaintenanceFacts `json:"maintenance"`
+}
+
+// CNPGActionResult is a successful action's answer. Requested, not completed:
+// the operator acts on the write asynchronously, and the UI observes the
+// outcome in status.
+type CNPGActionResult struct {
+ Action string `json:"action"`
+ Message string `json:"message"`
+ // Backup is the name of the Backup created by backup / run.
+ Backup string `json:"backup,omitempty"`
+ // ResolvedAfterTimeout: the create timed out and the Backup was found by
+ // re-reading the same name. Never a second create.
+ ResolvedAfterTimeout bool `json:"resolvedAfterTimeout,omitempty"`
+ CatchUp bool `json:"catchUp,omitempty"`
+ // Target identifies what the action acted on, so a caller can follow the
+ // outcome (the backend signalled, the instance destroyed, the Pooler paused).
+ Target *CNPGActionTarget `json:"target,omitempty"`
+}
+
+type ActionClients struct {
+ Dynamic dynamic.Interface
+ Typed kubernetes.Interface
+ Now func() time.Time
+ // Exec runs a command in a Pod as the caller; nil when unavailable.
+ Exec ExecFunc
+}
+
+func (c ActionClients) clock() time.Time {
+ if c.Now != nil {
+ return c.Now().UTC()
+ }
+ return time.Now().UTC()
+}
+
+func parseCNPGFenced(raw string) CNPGFencedFacts { return cnpg.ParseFencedInstances(raw) }
+
+func cnpgIsReplicaCluster(cluster *unstructured.Unstructured) bool {
+ replica, ok, _ := unstructured.NestedMap(cluster.Object, "spec", "replica")
+ if !ok || replica == nil {
+ return false
+ }
+ if enabled, ok := replica["enabled"].(bool); ok {
+ return enabled
+ }
+ self, _ := replica["self"].(string)
+ if strings.TrimSpace(self) == "" {
+ self = cluster.GetName()
+ }
+ // The operator's Cluster.IsReplica(): without `enabled`, the cluster is a
+ // replica unless it names itself as the primary (an unset primary included).
+ primary, _ := replica["primary"].(string)
+ return strings.TrimSpace(primary) != strings.TrimSpace(self)
+}
+
+func cnpgBackupMethods(cluster *unstructured.Unstructured) []CNPGBackupMethodFact {
+ type pluginStatus struct {
+ reported bool
+ backup bool
+ }
+ statuses := map[string]pluginStatus{}
+ if list, ok, _ := unstructured.NestedSlice(cluster.Object, "status", "pluginStatus"); ok {
+ for _, item := range list {
+ m, _ := item.(map[string]any)
+ name, _ := m["name"].(string)
+ raw, present := m["backupCapabilities"]
+ if !present {
+ continue
+ }
+ caps, _ := raw.([]any)
+ statuses[name] = pluginStatus{reported: true, backup: len(caps) > 0}
+ }
+ }
+ var plugins []cnpg.BackupPlugin
+ declaration := cnpg.ParseBackupDeclaration(cluster)
+ for _, plugin := range declaration.Plugins {
+ if plugin.Enabled {
+ plugins = append(plugins, plugin)
+ }
+ }
+ sort.SliceStable(plugins, func(i, j int) bool { return plugins[i].WALArchiver && !plugins[j].WALArchiver })
+
+ out := []CNPGBackupMethodFact{}
+ for _, p := range plugins {
+ fact := CNPGBackupMethodFact{Method: "plugin", PluginName: p.Name}
+ st := statuses[p.Name]
+ switch {
+ case !st.reported:
+ fact.Capability = "unknown"
+ fact.Reason = "The plugin does not report whether it can take backups (older plugin versions omit this); the operator rejects the Backup if it cannot"
+ case st.backup:
+ fact.Capability = "backup"
+ default:
+ fact.Capability = "none"
+ fact.Reason = "The plugin reports no backup capability"
+ }
+ out = append(out, fact)
+ }
+ if declaration.SnapshotsConfigured {
+ out = append(out, CNPGBackupMethodFact{Method: "volumeSnapshot", Capability: "backup"})
+ }
+ if declaration.InTreeConfigured {
+ out = append(out, CNPGBackupMethodFact{Method: "barmanObjectStore", Capability: "backup", Deprecated: true})
+ }
+ for i := range out {
+ if reason := cnpgBackupDestinationGuard(cluster, out[i].Method, out[i].PluginName); reason != "" {
+ out[i].Capability, out[i].Reason = "none", reason
+ }
+ }
+ return out
+}
+
+func cnpgBackupDestinationGuard(cluster *unstructured.Unstructured, method, pluginName string) string {
+ blocker := cnpg.ParseBackupDeclaration(cluster).DestinationBlocker(method, pluginName)
+ if blocker == nil {
+ return ""
+ }
+ if blocker.Plugin != "" {
+ return "No " + blocker.Method + " destination on " + cluster.GetName() + ". Use method plugin with pluginConfiguration.name " + blocker.Plugin + ", or configure this method"
+ }
+ if blocker.Code == "method_destination_missing" {
+ return "No " + blocker.Method + " destination on " + cluster.GetName() + ". Configure this method or choose the cluster's configured backup method"
+ }
+ return "Configure a backup destination on " + cluster.GetName() + " first"
+}
+
+func cnpgActionPodReady(p *corev1.Pod) bool {
+ for _, c := range p.Status.Conditions {
+ if c.Type == corev1.PodReady {
+ return c.Status == corev1.ConditionTrue
+ }
+ }
+ return false
+}
+
+// cnpgClusterFactsOf builds the facts. typed may be nil, in which case the
+// instances carry no Pod join (PodReadable false).
+func cnpgClusterFactsOf(ctx context.Context, typed kubernetes.Interface, cluster *unstructured.Unstructured) (CNPGClusterFacts, map[string]*corev1.Pod) {
+ str := func(fields ...string) string {
+ v, _, _ := unstructured.NestedString(cluster.Object, fields...)
+ return v
+ }
+ anno := cluster.GetAnnotations()
+ facts := CNPGClusterFacts{
+ Generation: cluster.GetGeneration(),
+ CurrentPrimary: str("status", "currentPrimary"),
+ TargetPrimary: str("status", "targetPrimary"),
+ Phase: str("status", "phase"),
+ PhaseReason: str("status", "phaseReason"),
+ Hibernation: anno[cnpgHibernateAnnotation],
+ FencedInstances: parseCNPGFenced(anno[cnpgFencedAnnotation]),
+ BackupMethods: cnpgBackupMethods(cluster),
+ BackupTarget: str("spec", "backup", "target"),
+ IsReplicaCluster: cnpgIsReplicaCluster(cluster),
+ Terminating: !cluster.GetDeletionTimestamp().IsZero(),
+ Instances: []CNPGInstanceFact{},
+ Maintenance: cnpgMaintenanceFactsOf(cluster),
+ }
+ facts.Hibernated = facts.Hibernation == "on"
+ conds, _, _ := unstructured.NestedSlice(cluster.Object, "status", "conditions")
+ for _, raw := range conds {
+ if c, ok := raw.(map[string]any); ok && c["type"] == "ContinuousArchiving" {
+ facts.ArchivingFailing = c["status"] == "False"
+ }
+ }
+
+ names, _, _ := unstructured.NestedStringSlice(cluster.Object, "status", "instanceNames")
+ if len(names) == 0 && facts.CurrentPrimary != "" {
+ names = []string{facts.CurrentPrimary}
+ }
+ healthy := map[string]bool{}
+ if list, ok, _ := unstructured.NestedStringSlice(cluster.Object, "status", "instancesStatus", "healthy"); ok {
+ for _, n := range list {
+ healthy[n] = true
+ }
+ }
+ uids := map[string]types.UID{cluster.GetNamespace() + "/" + cluster.GetName(): cluster.GetUID()}
+ pods := map[string]*corev1.Pod{}
+ for _, n := range names {
+ inst := CNPGInstanceFact{Pod: n, Role: "standby", Healthy: healthy[n], Fenced: facts.FencedInstances.Fences(n)}
+ if n == facts.CurrentPrimary {
+ inst.Role = "primary"
+ }
+ if typed != nil {
+ pod, err := typed.CoreV1().Pods(cluster.GetNamespace()).Get(ctx, n, metav1.GetOptions{})
+ switch {
+ case err == nil:
+ inst.PodReadable = true
+ // A Pod of this name that the Cluster does not control is not
+ // this instance; treat it like no Pod at all.
+ if isCNPGInstancePod(pod, uids) {
+ inst.PodExists = true
+ inst.PodUID = string(pod.UID)
+ inst.Ready = cnpgActionPodReady(pod)
+ pods[n] = pod
+ }
+ case apierrors.IsNotFound(err):
+ inst.PodReadable = true
+ }
+ }
+ facts.Instances = append(facts.Instances, inst)
+ }
+ return facts, pods
+}
+
+func (f CNPGClusterFacts) instance(name string) (CNPGInstanceFact, bool) {
+ for _, i := range f.Instances {
+ if i.Pod == name {
+ return i, true
+ }
+ }
+ return CNPGInstanceFact{}, false
+}
+
+func (f CNPGClusterFacts) switchoverInFlight() string {
+ if f.TargetPrimary != "" && f.CurrentPrimary != "" && f.TargetPrimary != f.CurrentPrimary {
+ return f.TargetPrimary
+ }
+ return ""
+}
+
+func cnpgGuardCommon(f CNPGClusterFacts) string {
+ if f.Terminating {
+ return "The cluster is being deleted"
+ }
+ return ""
+}
+
+func cnpgGuardBackup(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: the operator fails a backup requested now"
+ }
+ for _, m := range f.BackupMethods {
+ if m.Capability != "none" {
+ return ""
+ }
+ }
+ if len(f.BackupMethods) == 0 {
+ return "Configure a backup destination first: the cluster declares no backup plugin or snapshot configuration"
+ }
+ for _, m := range f.BackupMethods {
+ if m.Reason != "" {
+ return m.Reason
+ }
+ }
+ return "No declared backup method can take a backup"
+}
+
+func cnpgGuardSwitchover(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: there is no primary to move"
+ }
+ if f.IsReplicaCluster {
+ return "This cluster follows another one: a switchover here would only move the designated primary"
+ }
+ if t := f.switchoverInFlight(); t != "" {
+ return "A switchover or failover is already in flight, to " + t
+ }
+ if f.Phase == cnpgPhaseSwitchover || f.Phase == cnpgPhaseFailover {
+ return "A switchover or failover is already in flight"
+ }
+ if f.CurrentPrimary == "" {
+ return "The cluster reports no primary yet"
+ }
+ for _, i := range f.Instances {
+ if cnpgGuardSwitchoverTarget(f, i) == "" {
+ return ""
+ }
+ }
+ return "No standby can be promoted right now"
+}
+
+func cnpgGuardSwitchoverTarget(f CNPGClusterFacts, i CNPGInstanceFact) string {
+ switch {
+ case i.Pod == f.CurrentPrimary:
+ return "It is the primary"
+ case !i.PodReadable:
+ return "Its Pod cannot be read"
+ case !i.PodExists:
+ return "Its Pod does not exist"
+ case i.Fenced:
+ return "It is fenced"
+ case !i.Ready:
+ return "Its Pod is not ready"
+ }
+ return ""
+}
+
+func cnpgGuardRestart(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: there is no instance to restart"
+ }
+ if t := f.switchoverInFlight(); t != "" {
+ return "A switchover or failover is in flight, to " + t
+ }
+ return ""
+}
+
+func cnpgGuardRestartInstance(f CNPGClusterFacts, i CNPGInstanceFact) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: there is no instance to restart"
+ }
+ if i.Fenced {
+ return "It is fenced: lift the fence first"
+ }
+ if i.Pod == f.CurrentPrimary {
+ if t := f.switchoverInFlight(); t != "" {
+ return "A switchover or failover is in flight, to " + t
+ }
+ if f.Phase != cnpgPhaseHealthy && f.Phase != cnpgPhaseWaitingForUser {
+ return fmt.Sprintf("The primary restarts in place only on a healthy cluster or one waiting for user action; the phase is %q", f.Phase)
+ }
+ return ""
+ }
+ if i.PodReadable && !i.PodExists {
+ return "Its Pod does not exist"
+ }
+ return ""
+}
+
+func cnpgGuardReload(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: there is no instance to reload"
+ }
+ return ""
+}
+
+func cnpgGuardFence(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.FencedInstances.Malformed {
+ return "The cnpg.io/fencedInstances annotation does not parse as a JSON array of names; fix it before fencing"
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated: there is no PostgreSQL to stop"
+ }
+ if f.FencedInstances.All {
+ return "Every instance is fenced already"
+ }
+ return ""
+}
+
+func cnpgGuardUnfence(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.FencedInstances.Malformed {
+ return "The cnpg.io/fencedInstances annotation does not parse as a JSON array of names; fix it before lifting a fence"
+ }
+ if len(f.FencedInstances.Instances) == 0 {
+ return "No instance is fenced"
+ }
+ if f.ColdSnapshotBackup != "" {
+ return fmt.Sprintf("Backup %s is a cold snapshot: the operator fenced the instance for it and lifts the fence itself once the snapshot is taken. Lifting it now would start PostgreSQL under the snapshot", f.ColdSnapshotBackup)
+ }
+ return ""
+}
+
+// cnpgRunningColdSnapshot names a running offline (cold) volume-snapshot
+// Backup of the cluster, or "" when there is none. A failed list also gives
+// "": lifting a fence must stay possible during an incident, and the operator
+// still owns its own fence.
+func cnpgRunningColdSnapshot(ctx context.Context, dyn dynamic.Interface, cluster *unstructured.Unstructured) string {
+ if dyn == nil {
+ return ""
+ }
+ list, err := dyn.Resource(cnpgBackupGVR).Namespace(cluster.GetNamespace()).List(ctx, metav1.ListOptions{})
+ if err != nil {
+ return ""
+ }
+ // An unset online follows the Cluster's default, which is online.
+ clusterOnline, found, _ := unstructured.NestedBool(cluster.Object, "spec", "backup", "volumeSnapshot", "online")
+ if !found {
+ clusterOnline = true
+ }
+ for i := range list.Items {
+ b := &list.Items[i]
+ if name, _, _ := unstructured.NestedString(b.Object, "spec", "cluster", "name"); name != cluster.GetName() {
+ continue
+ }
+ method, _, _ := unstructured.NestedString(b.Object, "status", "method")
+ if method == "" {
+ method, _, _ = unstructured.NestedString(b.Object, "spec", "method")
+ }
+ if method != "volumeSnapshot" {
+ continue
+ }
+ online, found, _ := unstructured.NestedBool(b.Object, "status", "online")
+ if !found {
+ if online, found, _ = unstructured.NestedBool(b.Object, "spec", "online"); !found {
+ online = clusterOnline
+ }
+ }
+ if online {
+ continue
+ }
+ switch phase, _, _ := unstructured.NestedString(b.Object, "status", "phase"); phase {
+ case "", "completed", "failed", "walArchivingFailing":
+ continue
+ }
+ return b.GetName()
+ }
+ return ""
+}
+
+func cnpgGuardHibernate(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Hibernated {
+ return "The cluster is hibernated already"
+ }
+ return ""
+}
+
+func cnpgGuardRehydrate(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if !f.Hibernated {
+ return "The cluster is not hibernated"
+ }
+ return ""
+}
+
+var (
+ cnpgGrantCreateBackups = auth.Grant{Verb: "create", Group: Group, Resource: "backups"}
+ cnpgGrantCreateClusters = auth.Grant{Verb: "create", Group: Group, Resource: "clusters"}
+ cnpgGrantPatchStatus = auth.Grant{Verb: "patch", Group: Group, Resource: "clusters", Subresource: "status"}
+ GrantPatchClusters = auth.Grant{Verb: "patch", Group: Group, Resource: "clusters"}
+ cnpgGrantDeletePods = auth.Grant{Verb: "delete", Resource: "pods"}
+ grantGetPods = auth.Grant{Verb: "get", Resource: "pods"}
+ cnpgGrantPatchSchedules = auth.Grant{Verb: "patch", Group: Group, Resource: "scheduledbackups"}
+ cnpgClusterActionGrants = map[string]auth.Grant{"backup": cnpgGrantCreateBackups, "switchover": cnpgGrantPatchStatus, "restart": GrantPatchClusters, "reload": GrantPatchClusters, "fence": GrantPatchClusters, "unfence": GrantPatchClusters, "hibernate": GrantPatchClusters, "rehydrate": GrantPatchClusters}
+ scheduleActionGrants = map[string]auth.Grant{"suspend": cnpgGrantPatchSchedules, "resume": cnpgGrantPatchSchedules, "run": cnpgGrantCreateBackups, "setSchedule": cnpgGrantPatchSchedules, "repairMethod": cnpgGrantPatchSchedules}
+ clusterActionsOrdered = []string{"backup", "switchover", "restart", "restartInstance", "reload", "fence", "unfence", "hibernate", "rehydrate", "cancelBackend", "terminateBackend", "destroyInstance"}
+)
+
+func (s *Reader) ClusterCapabilities(callerCtx context.Context, c ActionClients, contextName, namespace, name string) (*CNPGClusterCapabilitiesResponse, error) {
+ ctx := callerCtx
+ cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ facts, _ := cnpgClusterFactsOf(ctx, c.Typed, cluster)
+ facts.ColdSnapshotBackup = cnpgRunningColdSnapshot(ctx, c.Dynamic, cluster)
+
+ perm := map[auth.Grant]string{}
+ permOf := func(g auth.Grant) string {
+ if v, ok := perm[g]; ok {
+ return v
+ }
+ v := s.Access.Permission(callerCtx, g.In(namespace))
+ perm[g] = v
+ return v
+ }
+ one := func(guard string, gs ...auth.Grant) integration.ActionCapability {
+ ps := make([]string, len(gs))
+ bound := make([]auth.Grant, len(gs))
+ for i, g := range gs {
+ ps[i] = permOf(g)
+ bound[i] = g.In(namespace)
+ }
+ return integration.CapabilityVerdict(guard, ps, bound)
+ }
+
+ instanceActions := map[string]CNPGInstanceActions{}
+ // The cluster-level restartInstance verdict is the first allowed row, or
+ // the first row's refusal when none is.
+ restartInstance := one(cnpgIfNoGuard(cnpgGuardCommon(facts), "The cluster has no instance"), cnpgGrantDeletePods)
+ restartAdopted := false
+ destroyInstance := one(cnpgIfNoGuard(cnpgGuardCommon(facts), "The cluster has no standby"), cnpgDestroyGrants(false)...)
+ destroyAdopted := false
+ for _, inst := range facts.Instances {
+ grant := cnpgGrantDeletePods
+ if inst.Pod == facts.CurrentPrimary {
+ grant = cnpgGrantPatchStatus
+ }
+ restart := one(cnpgGuardRestartInstance(facts, inst), grant)
+ if !restartInstance.Allowed && (restart.Allowed || !restartAdopted) {
+ restartInstance = restart
+ restartAdopted = true
+ }
+ switchGuard := cnpgGuardSwitchover(facts)
+ if tg := cnpgGuardSwitchoverTarget(facts, inst); tg != "" {
+ switchGuard = tg
+ }
+ fenceGuard := cnpgGuardFence(facts)
+ if fenceGuard == "" && inst.Fenced {
+ fenceGuard = "It is fenced already"
+ }
+ unfenceGuard := cnpgGuardUnfence(facts)
+ if unfenceGuard == "" {
+ switch {
+ case facts.FencedInstances.All:
+ unfenceGuard = `The whole cluster is fenced with ["*"]: lift every fence, or convert it to an explicit list first`
+ case !inst.Fenced:
+ unfenceGuard = "It is not fenced"
+ }
+ }
+ destroyCode, destroyGuard := cnpgDestroyInstanceBlocker(facts, inst)
+ destroy := one(destroyGuard, cnpgDestroyGrants(false)...)
+ if destroy.Permission != integration.PermissionDenied && destroyGuard != "" {
+ destroy.ReasonCode = destroyCode
+ }
+ if !destroyInstance.Allowed && (destroy.Allowed || !destroyAdopted) {
+ destroyInstance = destroy
+ destroyAdopted = true
+ }
+ instanceActions[inst.Pod] = CNPGInstanceActions{
+ Restart: restart,
+ SwitchoverTarget: one(switchGuard, cnpgGrantPatchStatus, grantGetPods),
+ Fence: one(fenceGuard, GrantPatchClusters),
+ Unfence: one(unfenceGuard, GrantPatchClusters),
+ Psql: one(cnpgGuardPsql(facts, inst), grantCreateExec),
+ Destroy: destroy,
+ }
+ }
+ psql := one(cnpgIfNoGuard(cnpgGuardCommon(facts), "No primary is reported"), grantCreateExec)
+ if primary, ok := facts.instance(facts.CurrentPrimary); ok {
+ psql = instanceActions[primary.Pod].Psql
+ }
+
+ resp := &CNPGClusterCapabilitiesResponse{
+ UID: string(cluster.GetUID()),
+ ResourceVersion: cluster.GetResourceVersion(),
+ Context: contextName,
+ Facts: facts,
+ Actions: CNPGClusterActions{
+ Backup: one(cnpgGuardBackup(facts), cnpgGrantCreateBackups),
+ Switchover: one(cnpgGuardSwitchover(facts), cnpgGrantPatchStatus, grantGetPods),
+ Restart: one(cnpgGuardRestart(facts), GrantPatchClusters),
+ RestartInstance: restartInstance,
+ Reload: one(cnpgGuardReload(facts), GrantPatchClusters),
+ Fence: one(cnpgGuardFence(facts), GrantPatchClusters),
+ Unfence: one(cnpgGuardUnfence(facts), GrantPatchClusters),
+ Hibernate: one(cnpgGuardHibernate(facts), GrantPatchClusters),
+ Rehydrate: one(cnpgGuardRehydrate(facts), GrantPatchClusters),
+ Psql: psql,
+ DestroyInstance: destroyInstance,
+ Restore: s.RestoreCapability(callerCtx, namespace),
+ CNPGMaintenanceActions: CNPGMaintenanceActions{
+ SetMaintenance: one(cnpgGuardSetMaintenance(facts), GrantPatchClusters),
+ UnsetMaintenance: one(cnpgGuardUnsetMaintenance(facts), GrantPatchClusters),
+ },
+ },
+ InstanceActions: instanceActions,
+ RestartPlan: cnpgRestartPlanOf(cluster, facts),
+ HibernateEffects: cnpgHibernateEffectsOf(ctx, c, cluster),
+ Operator: s.operatorVerdictFor(callerCtx, namespace),
+ }
+ cnpgApplyOperatorGuard(resp)
+ return resp, nil
+}
+
+// cnpgApplyOperatorGuard refuses the actions whose write the CNPG admission
+// webhook sees — a new Backup, or a patch of the Cluster itself — while that
+// webhook rejects writes. Status patches and Pod deletes bypass it.
+func cnpgApplyOperatorGuard(resp *CNPGClusterCapabilitiesResponse) {
+ v, a := resp.Operator, &resp.Actions
+ for _, c := range []*integration.ActionCapability{&a.Backup, &a.Restart, &a.Reload, &a.Fence, &a.Unfence, &a.Hibernate, &a.Rehydrate, &a.SetMaintenance, &a.UnsetMaintenance} {
+ *c = cnpgOperatorWebhookGuard(v, *c)
+ }
+ for pod, ia := range resp.InstanceActions {
+ ia.Fence = cnpgOperatorWebhookGuard(v, ia.Fence)
+ ia.Unfence = cnpgOperatorWebhookGuard(v, ia.Unfence)
+ resp.InstanceActions[pod] = ia
+ }
+}
+
+// ifNoGuard returns guard when set, fallback otherwise.
+func cnpgIfNoGuard(guard, fallback string) string {
+ if guard != "" {
+ return guard
+ }
+ return fallback
+}
+
+func cnpgRestartPlanOf(cluster *unstructured.Unstructured, f CNPGClusterFacts) CNPGRestartPlan {
+ strategy, _, _ := unstructured.NestedString(cluster.Object, "spec", "primaryUpdateStrategy")
+ method, _, _ := unstructured.NestedString(cluster.Object, "spec", "primaryUpdateMethod")
+ if strategy == "" {
+ strategy = "unsupervised"
+ }
+ if method == "" {
+ method = "restart"
+ }
+ plan := CNPGRestartPlan{PrimaryUpdateStrategy: strategy, PrimaryUpdateMethod: method, Steps: []CNPGRestartStep{}}
+ standbys := 0
+ for _, i := range f.Instances {
+ if i.Pod == f.CurrentPrimary {
+ continue
+ }
+ standbys++
+ effect := "recreate"
+ if i.Fenced {
+ effect = "skipped_fenced"
+ }
+ plan.Steps = append(plan.Steps, CNPGRestartStep{Instance: i.Pod, Role: "standby", Effect: effect})
+ }
+ if f.CurrentPrimary == "" {
+ return plan
+ }
+ step := CNPGRestartStep{Instance: f.CurrentPrimary, Role: "primary"}
+ switch {
+ case f.FencedInstances.Fences(f.CurrentPrimary):
+ step.Effect = "skipped_fenced"
+ case standbys == 0:
+ step.Effect = "restart_only_instance"
+ case strategy == "supervised":
+ step.Effect = "wait_for_user"
+ case method == "switchover":
+ step.Effect = "switchover"
+ default:
+ step.Effect = "restart"
+ }
+ plan.Steps = append(plan.Steps, step)
+ return plan
+}
+
+func cnpgEffectListOf(ctx context.Context, dyn dynamic.Interface, gvr schema.GroupVersionResource, namespace, cluster string, keep func(*unstructured.Unstructured) bool) CNPGEffectList {
+ out := CNPGEffectList{Names: []string{}}
+ list, err := dyn.Resource(gvr).Namespace(namespace).List(ctx, metav1.ListOptions{})
+ switch {
+ case apierrors.IsNotFound(err):
+ // The kind is not served (an older operator): none can exist.
+ out.Available = true
+ return out
+ case apierrors.IsForbidden(err):
+ out.Reason = "You are not allowed to list " + gvr.Resource
+ return out
+ case err != nil:
+ out.Reason = "Could not list " + gvr.Resource + ": " + err.Error()
+ return out
+ }
+ out.Available = true
+ for i := range list.Items {
+ item := &list.Items[i]
+ ref, _, _ := unstructured.NestedString(item.Object, "spec", "cluster", "name")
+ if ref != cluster || (keep != nil && !keep(item)) {
+ continue
+ }
+ out.Names = append(out.Names, item.GetName())
+ }
+ sort.Strings(out.Names)
+ return out
+}
+
+func cnpgHibernateEffectsOf(ctx context.Context, c ActionClients, cluster *unstructured.Unstructured) CNPGHibernateEffects {
+ ns, name := cluster.GetNamespace(), cluster.GetName()
+ eff := CNPGHibernateEffects{
+ Poolers: cnpgEffectListOf(ctx, c.Dynamic, cnpgPoolerGVR, ns, name, nil),
+ UnsuspendedScheduledBackups: cnpgEffectListOf(ctx, c.Dynamic, ScheduleGVR, ns, name, func(u *unstructured.Unstructured) bool {
+ suspended, _, _ := unstructured.NestedBool(u.Object, "spec", "suspend")
+ return !suspended
+ }),
+ Databases: cnpgEffectListOf(ctx, c.Dynamic, cnpgDatabaseGVR, ns, name, nil),
+ Publications: cnpgEffectListOf(ctx, c.Dynamic, cnpgPublGVR, ns, name, nil),
+ Subscriptions: cnpgEffectListOf(ctx, c.Dynamic, cnpgSubscrGVR, ns, name, nil),
+ Volumes: CNPGVolumeEffect{Items: []CNPGVolumeFact{}},
+ }
+ if c.Typed == nil {
+ eff.Volumes.Reason = "Volumes could not be read"
+ return eff
+ }
+ pvcs, err := c.Typed.CoreV1().PersistentVolumeClaims(ns).List(ctx, metav1.ListOptions{LabelSelector: clusterLabel + "=" + name})
+ switch {
+ case apierrors.IsForbidden(err):
+ eff.Volumes.Reason = "You are not allowed to list persistentvolumeclaims"
+ return eff
+ case err != nil:
+ eff.Volumes.Reason = "Could not list persistentvolumeclaims: " + err.Error()
+ return eff
+ }
+ eff.Volumes.Available = true
+ for _, pvc := range pvcs.Items {
+ // A label is something any object can carry; an owner naming another
+ // Cluster UID disqualifies the claim.
+ foreign := false
+ for _, ref := range pvc.OwnerReferences {
+ if ref.Kind == "Cluster" && ref.UID != cluster.GetUID() {
+ foreign = true
+ }
+ }
+ if foreign {
+ continue
+ }
+ v := CNPGVolumeFact{
+ Name: pvc.Name,
+ Instance: pvc.Labels[instanceNameLabel],
+ Role: pvc.Labels[cnpgPVCRoleLabel],
+ }
+ if q, ok := pvc.Status.Capacity[corev1.ResourceStorage]; ok {
+ v.Capacity = q.String()
+ }
+ if q, ok := pvc.Spec.Resources.Requests[corev1.ResourceStorage]; ok {
+ v.Requested = q.String()
+ }
+ eff.Volumes.Items = append(eff.Volumes.Items, v)
+ }
+ sort.Slice(eff.Volumes.Items, func(i, j int) bool { return eff.Volumes.Items[i].Name < eff.Volumes.Items[j].Name })
+ return eff
+}
+
+func cnpgScheduleFactsOf(ctx context.Context, c ActionClients, sched *unstructured.Unstructured) CNPGScheduleFacts {
+ str := func(fields ...string) string {
+ v, _, _ := unstructured.NestedString(sched.Object, fields...)
+ return v
+ }
+ suspended, _, _ := unstructured.NestedBool(sched.Object, "spec", "suspend")
+ f := CNPGScheduleFacts{
+ Generation: sched.GetGeneration(),
+ Cluster: str("spec", "cluster", "name"),
+ Suspended: suspended,
+ NextScheduleTime: str("status", "nextScheduleTime"),
+ Method: str("spec", "method"),
+ PluginName: str("spec", "pluginConfiguration", "name"),
+ Target: str("spec", "target"),
+ Terminating: !sched.GetDeletionTimestamp().IsZero(),
+ Schedule: str("spec", "schedule"),
+ }
+ if t, err := time.Parse(time.RFC3339, f.NextScheduleTime); err == nil && !t.After(c.clock()) {
+ f.CatchUp = true
+ }
+ f.Preview = schedulePreview(f.Schedule, scheduleLastCheck(sched), suspended, c.clock())
+ switch cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(sched.GetNamespace()).Get(ctx, f.Cluster, metav1.GetOptions{}); {
+ case f.Cluster == "":
+ f.ClusterState = "missing"
+ case apierrors.IsNotFound(err):
+ f.ClusterState = "missing"
+ case err != nil:
+ f.ClusterState = "unreadable"
+ case cluster.GetAnnotations()[cnpgHibernateAnnotation] == "on":
+ f.ClusterState = "hibernated"
+ default:
+ f.ClusterState = "ok"
+ f.ClusterUID = string(cluster.GetUID())
+ f.ClusterGeneration = cluster.GetGeneration()
+ f.BackupBlockedReason = cnpgBackupDestinationGuard(cluster, f.Method, f.PluginName)
+ }
+ return f
+}
+
+func cnpgScheduleRunBlocker(f CNPGScheduleFacts) (string, string) {
+ switch {
+ case f.Terminating:
+ return "terminating", "The schedule is being deleted"
+ case f.Cluster == "":
+ return "cluster_missing", "The schedule names no cluster"
+ case f.ClusterState == "missing":
+ return "cluster_missing", "The cluster of the schedule does not exist in this namespace"
+ case f.ClusterState == "hibernated":
+ return "hibernated", "The cluster is hibernated: the operator fails a backup requested now"
+ case f.ClusterState == "unreadable":
+ return "cluster_unreadable", "The cluster's backup destination could not be read"
+ case f.BackupBlockedReason != "":
+ return "backup_destination", f.BackupBlockedReason
+ default:
+ return "", ""
+ }
+}
+
+func cnpgGuardScheduleRun(f CNPGScheduleFacts) string {
+ _, message := cnpgScheduleRunBlocker(f)
+ return message
+}
+
+func (s *Reader) ScheduleCapabilities(ctx context.Context, c ActionClients, contextName, namespace, name string) (*CNPGScheduleCapabilitiesResponse, error) {
+ sched, err := c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ f := cnpgScheduleFactsOf(ctx, c, sched)
+ f.Preview = s.previewOnOperatorClock(ctx, namespace, f.Schedule, scheduleLastCheck(sched), f.Suspended, c.clock())
+ one := func(guard string, g auth.Grant) integration.ActionCapability {
+ g = g.In(namespace)
+ return integration.CapabilityVerdict(guard, []string{s.Access.Permission(ctx, g)}, []auth.Grant{g})
+ }
+ suspendGuard, resumeGuard := "", ""
+ if f.Terminating {
+ suspendGuard, resumeGuard = "The schedule is being deleted", "The schedule is being deleted"
+ } else if f.Suspended {
+ suspendGuard = "The schedule is suspended already"
+ } else {
+ resumeGuard = "The schedule is not suspended"
+ }
+ operator := s.operatorVerdictFor(ctx, namespace)
+ code, message := cnpgScheduleRunBlocker(f)
+ run := one(message, cnpgGrantCreateBackups)
+ if run.Permission != integration.PermissionDenied {
+ run.ReasonCode = code
+ }
+ run = cnpgOperatorWebhookGuard(operator, run)
+ return &CNPGScheduleCapabilitiesResponse{
+ UID: string(sched.GetUID()),
+ ResourceVersion: sched.GetResourceVersion(),
+ Context: contextName,
+ Facts: f,
+ Actions: CNPGScheduleActions{
+ Suspend: cnpgOperatorWebhookGuard(operator, one(suspendGuard, cnpgGrantPatchSchedules)),
+ Resume: cnpgOperatorWebhookGuard(operator, one(resumeGuard, cnpgGrantPatchSchedules)),
+ Run: run,
+ SetSchedule: cnpgOperatorWebhookGuard(operator, one(map[bool]string{true: "The schedule is being deleted"}[f.Terminating], cnpgGrantPatchSchedules)),
+ },
+ Operator: operator,
+ }, nil
+}
+
+type cnpgClusterRun struct {
+ c ActionClients
+ cluster *unstructured.Unstructured
+ facts CNPGClusterFacts
+ reviewed cnpgReviewedFacts
+ pods map[string]*corev1.Pod
+ params json.RawMessage
+}
+
+type cnpgClusterRunner struct {
+ // binds lists the reviewed facts this action requires and compares.
+ binds []string
+ // needsPods: the action reads instance Pods to validate its target.
+ needsPods bool
+ run func(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error)
+}
+
+var clusterActionRunners = map[string]cnpgClusterRunner{
+ "backup": {binds: []string{"hibernation", "generation"}, run: cnpgRunBackup},
+ "switchover": {binds: []string{"currentPrimary", "targetPrimary", "fencedInstances"}, needsPods: true, run: cnpgRunSwitchover},
+ "restart": {binds: []string{"currentPrimary", "targetPrimary", "hibernation", "fencedInstances"}, run: cnpgRunRestart},
+ "restartInstance": {binds: []string{"currentPrimary", "targetPrimary", "fencedInstances"}, needsPods: true, run: cnpgRunRestartInstance},
+ "reload": {binds: []string{"hibernation"}, run: cnpgRunReload},
+ "fence": {binds: []string{"currentPrimary", "hibernation", "fencedInstances"}, run: cnpgRunFence},
+ "unfence": {binds: []string{"currentPrimary", "fencedInstances"}, run: cnpgRunUnfence},
+ "hibernate": {binds: []string{"hibernation"}, run: cnpgRunHibernation("on")},
+ "rehydrate": {binds: []string{"hibernation"}, run: cnpgRunHibernation("off")},
+ "cancelBackend": {needsPods: true, run: cnpgRunSignalBackend(cnpgSignalCancel)},
+ "terminateBackend": {needsPods: true, run: cnpgRunSignalBackend(cnpgSignalTerminate)},
+ "destroyInstance": {binds: []string{"currentPrimary", "targetPrimary", "fencedInstances"}, needsPods: true, run: cnpgRunDestroyInstance},
+}
+
+func decodeCNPGReviewedFacts(raw json.RawMessage) (cnpgReviewedFacts, error) {
+ var f cnpgReviewedFacts
+ if len(bytes.TrimSpace(raw)) == 0 || string(bytes.TrimSpace(raw)) == "null" {
+ return f, nil
+ }
+ if err := json.Unmarshal(raw, &f); err != nil {
+ return f, fmt.Errorf("facts: %w", err)
+ }
+ return f, nil
+}
+
+// cnpgFactsDiffer compares the reviewed facts an action binds with the facts
+// now. A bound fact missing from the request is a malformed request.
+func cnpgFactsDiffer(binds []string, reviewed cnpgReviewedFacts, now CNPGClusterFacts) (changed []string, missing string) {
+ for _, b := range binds {
+ var got *string
+ var want string
+ switch b {
+ case "generation":
+ if reviewed.Generation == nil {
+ return nil, b
+ }
+ if *reviewed.Generation != now.Generation {
+ changed = append(changed, b)
+ }
+ continue
+ case "currentPrimary":
+ got, want = reviewed.CurrentPrimary, now.CurrentPrimary
+ case "targetPrimary":
+ got, want = reviewed.TargetPrimary, now.TargetPrimary
+ case "hibernation":
+ got, want = reviewed.Hibernation, now.Hibernation
+ case "maintenance":
+ if reviewed.Maintenance == nil {
+ return nil, b
+ }
+ if *reviewed.Maintenance != now.Maintenance {
+ changed = append(changed, b)
+ }
+ continue
+ case "fencedInstances":
+ b = "fencedInstances.raw"
+ if reviewed.FencedInstances != nil {
+ got = reviewed.FencedInstances.Raw
+ }
+ want = now.FencedInstances.Raw
+ }
+ if got == nil {
+ return nil, b
+ }
+ if *got != want {
+ changed = append(changed, b)
+ }
+ }
+ return changed, ""
+}
+
+func RunCNPGClusterAction(ctx context.Context, c ActionClients, namespace, name, action string, req integration.ActionRequest) (*CNPGActionResult, error) {
+ if action == "configureArchiving" {
+ return runConfigureArchiving(ctx, c, namespace, name, req)
+ }
+ runner, ok := clusterActionRunners[action]
+ if !ok {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "unknown action %q", action)
+ }
+ reviewed, err := decodeCNPGReviewedFacts(req.Facts)
+ if err != nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "%v", err)
+ }
+ cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ typed := c.Typed
+ if !runner.needsPods {
+ typed = nil
+ }
+ facts, pods := cnpgClusterFactsOf(ctx, typed, cluster)
+ facts.ColdSnapshotBackup = cnpgRunningColdSnapshot(ctx, c.Dynamic, cluster)
+ if string(cluster.GetUID()) != req.UID {
+ return nil, integration.ChangedAction(facts, "Cluster %s/%s was deleted and recreated since you reviewed it", namespace, name)
+ }
+ changed, missing := cnpgFactsDiffer(runner.binds, reviewed, facts)
+ if missing != "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.%s is required for %s: the confirmation must bind what the dialog showed", missing, action)
+ }
+ if len(changed) > 0 {
+ return nil, integration.ChangedAction(facts, "Cluster %s/%s changed since you confirmed (%s); review the action again", namespace, name, strings.Join(changed, ", "))
+ }
+ x := &cnpgClusterRun{c: c, cluster: cluster, facts: facts, reviewed: reviewed, pods: pods, params: req.Params}
+ res, err := runner.run(ctx, x)
+ if err != nil {
+ if apierrors.IsConflict(err) {
+ // The object moved between our read and the write. Never retried:
+ // report what it looks like now so the user can review again.
+ current := facts
+ if fresh, gerr := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{}); gerr == nil {
+ current, _ = cnpgClusterFactsOf(ctx, typed, fresh)
+ }
+ return nil, integration.ChangedAction(current, "Cluster %s/%s changed while the request was being sent; review the action again", namespace, name)
+ }
+ return nil, err
+ }
+ res.Action = action
+ return res, nil
+}
+
+// cnpgConditionsWithReady returns status.conditions with the Ready condition
+// set from phase, exactly as the operator's SetClusterReadyCondition does. The
+// merge patch replaces the whole list, as upstream's MergeFrom does.
+func cnpgConditionsWithReady(cluster *unstructured.Unstructured, phase string, now time.Time) ([]any, error) {
+ var conds []metav1.Condition
+ if raw, ok, _ := unstructured.NestedSlice(cluster.Object, "status", "conditions"); ok {
+ data, err := json.Marshal(raw)
+ if err != nil {
+ return nil, err
+ }
+ if err := json.Unmarshal(data, &conds); err != nil {
+ return nil, fmt.Errorf("status.conditions: %w", err)
+ }
+ }
+ ready := metav1.Condition{Type: "Ready", Status: metav1.ConditionFalse, Reason: "ClusterIsNotReady", Message: "Cluster Is Not Ready", LastTransitionTime: metav1.NewTime(now)}
+ if phase == cnpgPhaseHealthy {
+ ready = metav1.Condition{Type: "Ready", Status: metav1.ConditionTrue, Reason: "ClusterIsReady", Message: "Cluster is Ready", LastTransitionTime: metav1.NewTime(now)}
+ }
+ if conds == nil {
+ conds = []metav1.Condition{}
+ }
+ meta.SetStatusCondition(&conds, ready)
+ data, err := json.Marshal(conds)
+ if err != nil {
+ return nil, err
+ }
+ var out []any
+ if err := json.Unmarshal(data, &out); err != nil {
+ return nil, err
+ }
+ return out, nil
+}
+
+func cnpgStatusPatch(ctx context.Context, x *cnpgClusterRun, status map[string]any) error {
+ phase, _ := status["phase"].(string)
+ conds, err := cnpgConditionsWithReady(x.cluster, phase, x.c.clock())
+ if err != nil {
+ return err
+ }
+ status["conditions"] = conds
+ return integration.MergePatchAtVersion(ctx, x.c.Dynamic, ClusterGVR, x.cluster, map[string]any{"status": status}, "status")
+}
+
+func cnpgAnnotationPatch(ctx context.Context, x *cnpgClusterRun, key string, value any) error {
+ return integration.MergePatchAtVersion(ctx, x.c.Dynamic, ClusterGVR, x.cluster, map[string]any{
+ "metadata": map[string]any{"annotations": map[string]any{key: value}},
+ })
+}
+
+type cnpgBackupParams struct {
+ Method string `json:"method"`
+ PluginName string `json:"pluginName,omitempty"`
+ PluginParameters map[string]string `json:"pluginParameters,omitempty"`
+ Target string `json:"target,omitempty"`
+ Online *bool `json:"online,omitempty"`
+ Name string `json:"name,omitempty"`
+}
+
+func cnpgRunBackup(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgBackupParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.Method == "plugin" && p.PluginName == "barman-cloud.cloudnative-pg.io" {
+ for _, key := range []string{"barmanObjectName", "serverName"} {
+ if _, ok := p.PluginParameters[key]; ok {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "the barman-cloud plugin takes its destination from the Cluster; pluginParameters.%s is ignored by the plugin", key)
+ }
+ }
+ }
+ if r := cnpgGuardBackup(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ var chosen *CNPGBackupMethodFact
+ for i, m := range x.facts.BackupMethods {
+ if m.Method == p.Method && (m.Method != "plugin" || m.PluginName == p.PluginName) {
+ chosen = &x.facts.BackupMethods[i]
+ break
+ }
+ }
+ if chosen == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "method %q%s is not a backup method this cluster declares", p.Method, cnpgPluginSuffix(p.PluginName))
+ }
+ if r := cnpgBackupDestinationGuard(x.cluster, p.Method, p.PluginName); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if chosen.Capability == "none" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "plugin %s reports no backup capability", chosen.PluginName)
+ }
+ if len(p.PluginParameters) > 0 && chosen.Method != "plugin" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "pluginParameters apply only to the plugin method")
+ }
+ if p.Target != "" && p.Target != "primary" && p.Target != "prefer-standby" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "target must be primary or prefer-standby (or omitted to inherit the cluster's)")
+ }
+ // A Backup without its own target inherits the cluster's, so the one the
+ // confirmation showed is bound. The facts omit it when unset.
+ if p.Target == "" {
+ reviewedTarget := ""
+ if x.reviewed.BackupTarget != nil {
+ reviewedTarget = *x.reviewed.BackupTarget
+ }
+ if reviewedTarget != x.facts.BackupTarget {
+ return nil, integration.ChangedAction(x.facts, "Cluster %s/%s's backup target changed since you confirmed (backupTarget); review the action again", x.cluster.GetNamespace(), x.cluster.GetName())
+ }
+ }
+ clusterName := x.cluster.GetName()
+ name := p.Name
+ if name == "" {
+ name = clusterName + "-" + x.c.clock().Format(compactStamp)
+ }
+ if err := cnpgValidateBackupName(ctx, x.c.Dynamic, x.cluster.GetNamespace(), name); err != nil {
+ return nil, err
+ }
+ spec := map[string]any{
+ "cluster": map[string]any{"name": clusterName},
+ "method": chosen.Method,
+ }
+ if chosen.Method == "plugin" {
+ pc := map[string]any{"name": chosen.PluginName}
+ if len(p.PluginParameters) > 0 {
+ params := map[string]any{}
+ for k, v := range p.PluginParameters {
+ params[k] = v
+ }
+ pc["parameters"] = params
+ }
+ spec["pluginConfiguration"] = pc
+ }
+ if p.Target != "" {
+ spec["target"] = p.Target
+ }
+ if p.Online != nil {
+ spec["online"] = *p.Online
+ }
+ backup := cnpgBackupObject(x.cluster.GetNamespace(), name, clusterName, nil, spec)
+ return cnpgCreateBackup(ctx, x.c.Dynamic, backup, clusterName)
+}
+
+func cnpgPluginSuffix(plugin string) string {
+ if plugin == "" {
+ return ""
+ }
+ return " (" + plugin + ")"
+}
+
+func cnpgBackupObject(namespace, name, cluster string, annotations map[string]any, spec map[string]any) *unstructured.Unstructured {
+ md := map[string]any{
+ "name": name,
+ "namespace": namespace,
+ "labels": map[string]any{clusterLabel: cluster},
+ }
+ if len(annotations) > 0 {
+ md["annotations"] = annotations
+ }
+ return &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": Group + "/v1",
+ "kind": "Backup",
+ "metadata": md,
+ "spec": spec,
+ }}
+}
+
+// cnpgValidateBackupName refuses a name that is not a DNS subdomain, that has
+// the form of a scheduled run (the operator would then skip that run), or that
+// already exists.
+func cnpgValidateBackupName(ctx context.Context, dyn dynamic.Interface, namespace, name string) error {
+ if errs := validation.IsDNS1123Subdomain(name); len(errs) > 0 {
+ return integration.RefuseAction(http.StatusBadRequest, "", "backup name %q is invalid: %s", name, strings.Join(errs, "; "))
+ }
+ if cnpgScheduleRunR.MatchString(name) {
+ schedule := name[:len(name)-15]
+ _, err := dyn.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, schedule, metav1.GetOptions{})
+ switch {
+ case err == nil:
+ return integration.RefuseAction(http.StatusBadRequest, "", "backup name %q has the form of a run of ScheduledBackup %s: the operator would skip that run", name, schedule)
+ case apierrors.IsNotFound(err):
+ default:
+ return fmt.Errorf("cannot check the name against ScheduledBackup %s: %w", schedule, err)
+ }
+ }
+ _, err := dyn.Resource(cnpgBackupGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ switch {
+ case err == nil:
+ return integration.RefuseAction(http.StatusConflict, "", "a Backup named %s already exists", name)
+ case apierrors.IsNotFound(err):
+ return nil
+ default:
+ return err
+ }
+}
+
+func cnpgAmbiguousOutcome(err error) bool {
+ if errors.Is(err, context.DeadlineExceeded) || apierrors.IsTimeout(err) || apierrors.IsServerTimeout(err) {
+ return true
+ }
+ var ne net.Error
+ return errors.As(err, &ne) && ne.Timeout()
+}
+
+// cnpgCreateBackup creates the Backup once. A create whose answer was lost is
+// resolved by reading the same name back, never by creating another.
+func cnpgCreateBackup(ctx context.Context, dyn dynamic.Interface, backup *unstructured.Unstructured, cluster string) (*CNPGActionResult, error) {
+ ns, name := backup.GetNamespace(), backup.GetName()
+ _, err := dyn.Resource(cnpgBackupGVR).Namespace(ns).Create(ctx, backup, metav1.CreateOptions{})
+ if err == nil {
+ return &CNPGActionResult{Backup: name, Message: "Backup " + name + " created"}, nil
+ }
+ if !cnpgAmbiguousOutcome(err) {
+ return nil, err
+ }
+ readCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), 10*time.Second)
+ defer cancel()
+ got, gerr := dyn.Resource(cnpgBackupGVR).Namespace(ns).Get(readCtx, name, metav1.GetOptions{})
+ if gerr == nil {
+ ref, _, _ := unstructured.NestedString(got.Object, "spec", "cluster", "name")
+ if ref == cluster {
+ return &CNPGActionResult{Backup: name, ResolvedAfterTimeout: true, Message: "Backup " + name + " created (confirmed by reading it back after the request timed out)"}, nil
+ }
+ }
+ return nil, integration.RefuseAction(http.StatusGatewayTimeout, cnpgCodeAmbiguous,
+ "The create of Backup %s timed out and reading it back did not find it; it may still appear. Check the cluster's backups before trying again: %v", name, err)
+}
+
+type cnpgSwitchoverParams struct {
+ Target string `json:"target"`
+ TargetPodUID string `json:"targetPodUID"`
+}
+
+func cnpgRunSwitchover(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgSwitchoverParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.Target == "" || p.TargetPodUID == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.target and params.targetPodUID are required")
+ }
+ if r := cnpgGuardSwitchover(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ inst, ok := x.facts.instance(p.Target)
+ if !ok {
+ return nil, integration.ChangedAction(x.facts, "%s is not an instance of this cluster", p.Target)
+ }
+ if inst.PodExists && inst.PodUID != p.TargetPodUID {
+ return nil, integration.ChangedAction(x.facts, "Pod %s was recreated since you reviewed it", p.Target)
+ }
+ if r := cnpgGuardSwitchoverTarget(x.facts, inst); r != "" {
+ return nil, integration.BlockedAction(p.Target + " cannot be promoted: " + r)
+ }
+ err := cnpgStatusPatch(ctx, x, map[string]any{
+ "targetPrimary": p.Target,
+ // RFC3339 with microseconds, as the operator's pgTime.GetCurrentTimestamp.
+ "targetPrimaryTimestamp": x.c.clock().Format(metav1.RFC3339Micro),
+ "phase": cnpgPhaseSwitchover,
+ "phaseReason": "Switching over to " + p.Target,
+ })
+ if err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Switchover to " + p.Target + " requested"}, nil
+}
+
+func cnpgRunRestart(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ if err := integration.DecodeActionParams(x.params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ if r := cnpgGuardRestart(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if err := cnpgAnnotationPatch(ctx, x, cnpgRestartAnnotation, x.c.clock().Format(time.RFC3339)); err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Rolling restart requested"}, nil
+}
+
+func cnpgRunReload(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ if err := integration.DecodeActionParams(x.params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ if r := cnpgGuardReload(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if err := cnpgAnnotationPatch(ctx, x, cnpgReloadAnnotation, x.c.clock().Format(metav1.RFC3339Micro)); err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Configuration reload requested"}, nil
+}
+
+type cnpgRestartInstanceParams struct {
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+}
+
+func cnpgRunRestartInstance(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgRestartInstanceParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.Pod == "" || p.PodUID == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.pod and params.podUID are required")
+ }
+ inst, ok := x.facts.instance(p.Pod)
+ if !ok {
+ return nil, integration.ChangedAction(x.facts, "%s is not an instance of this cluster", p.Pod)
+ }
+ if r := cnpgGuardRestartInstance(x.facts, inst); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if !inst.PodReadable {
+ return nil, integration.RefuseAction(http.StatusForbidden, "", "Pod %s cannot be read, so it cannot be verified as this cluster's instance", p.Pod)
+ }
+ if !inst.PodExists || inst.PodUID != p.PodUID {
+ return nil, integration.ChangedAction(x.facts, "Pod %s was recreated since you reviewed it", p.Pod)
+ }
+ if inst.Pod == x.facts.CurrentPrimary {
+ err := cnpgStatusPatch(ctx, x, map[string]any{
+ "phase": cnpgPhaseInplaceRestart,
+ "phaseReason": cnpgInplaceReason,
+ })
+ if err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "In-place restart of primary " + p.Pod + " requested"}, nil
+ }
+ uid := types.UID(p.PodUID)
+ err := x.c.Typed.CoreV1().Pods(x.cluster.GetNamespace()).Delete(ctx, p.Pod, metav1.DeleteOptions{
+ Preconditions: &metav1.Preconditions{UID: &uid},
+ })
+ if err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Pod " + p.Pod + " deleted; the operator recreates it on its volumes"}, nil
+}
+
+type cnpgFenceParams struct {
+ Instances json.RawMessage `json:"instances"`
+ ConvertFromAll bool `json:"convertFromAll,omitempty"`
+ Remaining []string `json:"remaining,omitempty"`
+}
+
+// cnpgFenceTargets reads params.instances: "*" or a list of names.
+func cnpgFenceTargets(raw json.RawMessage) (all bool, names []string, err error) {
+ var s string
+ if json.Unmarshal(raw, &s) == nil {
+ if s != cnpgAllInstances {
+ return false, nil, integration.RefuseAction(http.StatusBadRequest, "", `params.instances must be "*" or a list of instance names`)
+ }
+ return true, nil, nil
+ }
+ if err := json.Unmarshal(raw, &names); err != nil || len(names) == 0 {
+ return false, nil, integration.RefuseAction(http.StatusBadRequest, "", `params.instances must be "*" or a non-empty list of instance names`)
+ }
+ for _, n := range names {
+ if n == cnpgAllInstances {
+ return true, nil, nil
+ }
+ }
+ return false, names, nil
+}
+
+// cnpgFencedValue serializes a fenced set; nil removes the annotation.
+func cnpgFencedValue(all bool, names []string) (any, error) {
+ if all {
+ return `["*"]`, nil
+ }
+ if len(names) == 0 {
+ return nil, nil
+ }
+ sorted := append([]string(nil), names...)
+ sort.Strings(sorted)
+ b, err := json.Marshal(sorted)
+ return string(b), err
+}
+
+func cnpgRunFence(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgFenceParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.ConvertFromAll || len(p.Remaining) > 0 {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "convertFromAll and remaining apply only to unfence")
+ }
+ all, names, err := cnpgFenceTargets(p.Instances)
+ if err != nil {
+ return nil, err
+ }
+ if r := cnpgGuardFence(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ next := map[string]bool{}
+ for _, n := range x.facts.FencedInstances.Instances {
+ next[n] = true
+ }
+ for _, n := range names {
+ if _, ok := x.facts.instance(n); !ok {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "%s is not an instance of this cluster", n)
+ }
+ if next[n] {
+ return nil, integration.BlockedAction(n + " is fenced already")
+ }
+ next[n] = true
+ }
+ value, err := cnpgFencedValue(all, cnpgSortedKeys(next))
+ if err != nil {
+ return nil, err
+ }
+ if err := cnpgAnnotationPatch(ctx, x, cnpgFencedAnnotation, value); err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Fencing requested"}, nil
+}
+
+func cnpgRunUnfence(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgFenceParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ all, names, err := cnpgFenceTargets(p.Instances)
+ if err != nil {
+ return nil, err
+ }
+ if r := cnpgGuardUnfence(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ var value any
+ switch {
+ case all:
+ value = nil
+ case x.facts.FencedInstances.All:
+ if !p.ConvertFromAll {
+ return nil, integration.RefuseAction(http.StatusConflict, cnpgCodeAllFenced,
+ `Every instance is fenced with ["*"]: lifting one instance means rewriting the fence as the explicit list of the others. Lift every fence, or confirm the conversion with convertFromAll and the remaining list`)
+ }
+ // The reviewer must have seen the full explicit set that stays fenced.
+ want := map[string]bool{}
+ for _, i := range x.facts.Instances {
+ want[i.Pod] = true
+ }
+ for _, n := range names {
+ if _, ok := x.facts.instance(n); !ok {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "%s is not an instance of this cluster", n)
+ }
+ delete(want, n)
+ }
+ if !cnpgSameStringSet(p.Remaining, cnpgSortedKeys(want)) {
+ return nil, integration.ChangedAction(x.facts, "The instances that stay fenced are %s, not %s; review the conversion again",
+ strings.Join(cnpgSortedKeys(want), ", "), strings.Join(p.Remaining, ", "))
+ }
+ if value, err = cnpgFencedValue(false, cnpgSortedKeys(want)); err != nil {
+ return nil, err
+ }
+ default:
+ if p.ConvertFromAll {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", `convertFromAll applies only while ["*"] fences every instance`)
+ }
+ next := map[string]bool{}
+ for _, n := range x.facts.FencedInstances.Instances {
+ next[n] = true
+ }
+ for _, n := range names {
+ if !next[n] {
+ return nil, integration.BlockedAction(n + " is not fenced")
+ }
+ delete(next, n)
+ }
+ if value, err = cnpgFencedValue(false, cnpgSortedKeys(next)); err != nil {
+ return nil, err
+ }
+ }
+ if err := cnpgAnnotationPatch(ctx, x, cnpgFencedAnnotation, value); err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: "Lifting the fence requested"}, nil
+}
+
+func cnpgSortedKeys(m map[string]bool) []string {
+ out := make([]string, 0, len(m))
+ for k, v := range m {
+ if v {
+ out = append(out, k)
+ }
+ }
+ sort.Strings(out)
+ return out
+}
+
+func cnpgSameStringSet(a, b []string) bool {
+ if len(a) != len(b) {
+ return false
+ }
+ seen := map[string]int{}
+ for _, v := range a {
+ seen[v]++
+ }
+ for _, v := range b {
+ seen[v]--
+ }
+ for _, n := range seen {
+ if n != 0 {
+ return false
+ }
+ }
+ return true
+}
+
+func cnpgRunHibernation(value string) func(context.Context, *cnpgClusterRun) (*CNPGActionResult, error) {
+ return func(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ if err := integration.DecodeActionParams(x.params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ guard, msg := cnpgGuardHibernate, "Hibernation requested"
+ if value == "off" {
+ guard, msg = cnpgGuardRehydrate, "Rehydration requested"
+ }
+ if r := guard(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if err := cnpgAnnotationPatch(ctx, x, cnpgHibernateAnnotation, value); err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: msg}, nil
+ }
+}
+
+func RunCNPGScheduleAction(ctx context.Context, c ActionClients, namespace, name, action string, req integration.ActionRequest) (*CNPGActionResult, error) {
+ if action == "repairMethod" {
+ return runRepairScheduleMethod(ctx, c, namespace, name, req)
+ }
+ reviewed, err := decodeCNPGReviewedFacts(req.Facts)
+ if err != nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "%v", err)
+ }
+ var params struct {
+ Schedule *string `json:"schedule"`
+ }
+ if action == "setSchedule" {
+ if err := integration.DecodeActionParams(req.Params, ¶ms); err != nil {
+ return nil, err
+ }
+ if params.Schedule == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.schedule is required for setSchedule")
+ }
+ if p := schedulePreview(*params.Schedule, nil, false, c.clock()); !p.Valid {
+ return nil, integration.RefuseAction(http.StatusBadRequest, cnpgCodeInvalidSchedule, "invalid schedule %q: %s", *params.Schedule, p.Error)
+ }
+ } else if err := integration.DecodeActionParams(req.Params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ sched, err := c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ facts := cnpgScheduleFactsOf(ctx, c, sched)
+ if string(sched.GetUID()) != req.UID {
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s was deleted and recreated since you reviewed it", namespace, name)
+ }
+ switch action {
+ case "setSchedule":
+ if reviewed.Schedule == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.schedule is required for setSchedule: the confirmation must bind what the dialog showed")
+ }
+ if *reviewed.Schedule != facts.Schedule {
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s schedule changed since you reviewed it; review the action again", namespace, name)
+ }
+ case "run":
+ if reviewed.Generation == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.generation is required for run")
+ }
+ if *reviewed.Generation != facts.Generation {
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s settings changed since you confirmed; review the action again", namespace, name)
+ }
+ default:
+ if reviewed.Suspended == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.suspended is required for %s", action)
+ }
+ if *reviewed.Suspended != facts.Suspended {
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s changed since you confirmed (suspended); review the action again", namespace, name)
+ }
+ }
+ if facts.Terminating {
+ return nil, integration.BlockedAction("The schedule is being deleted")
+ }
+
+ switch action {
+ case "setSchedule":
+ if *params.Schedule == facts.Schedule {
+ return nil, integration.BlockedAction("The schedule is already " + facts.Schedule)
+ }
+ err := integration.MergePatchAtVersion(ctx, c.Dynamic, ScheduleGVR, sched, map[string]any{"spec": map[string]any{"schedule": *params.Schedule}})
+ if apierrors.IsConflict(err) {
+ if fresh, gerr := c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{}); gerr == nil {
+ facts = cnpgScheduleFactsOf(ctx, c, fresh)
+ }
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s changed while the request was being sent; review the action again", namespace, name)
+ }
+ if err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Action: action, Message: "Schedule set to " + *params.Schedule}, nil
+ case "suspend", "resume":
+ want := action == "suspend"
+ if facts.Suspended == want {
+ return nil, integration.BlockedAction(map[bool]string{true: "The schedule is suspended already", false: "The schedule is not suspended"}[want])
+ }
+ err := integration.MergePatchAtVersion(ctx, c.Dynamic, ScheduleGVR, sched, map[string]any{"spec": map[string]any{"suspend": want}})
+ if apierrors.IsConflict(err) {
+ if fresh, gerr := c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{}); gerr == nil {
+ facts = cnpgScheduleFactsOf(ctx, c, fresh)
+ }
+ return nil, integration.ChangedAction(facts, "ScheduledBackup %s/%s changed while the request was being sent; review the action again", namespace, name)
+ }
+ if err != nil {
+ return nil, err
+ }
+ res := &CNPGActionResult{Action: action, Message: "Schedule suspended"}
+ if !want {
+ res.Message = "Schedule resumed"
+ res.CatchUp = facts.CatchUp
+ }
+ return res, nil
+ }
+
+ if r := cnpgGuardScheduleRun(facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if reviewed.ClusterUID == nil || reviewed.ClusterGeneration == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.clusterUID and facts.clusterGeneration are required for run")
+ }
+ if *reviewed.ClusterUID != facts.ClusterUID || *reviewed.ClusterGeneration != facts.ClusterGeneration {
+ return nil, integration.ChangedAction(facts, "The Cluster of ScheduledBackup %s/%s changed since you reviewed it; review the action again", namespace, name)
+ }
+ backupName := name + "-manual-" + c.clock().Format(compactStamp)
+ if errs := validation.IsDNS1123Subdomain(backupName); len(errs) > 0 {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "backup name %q is invalid: %s", backupName, strings.Join(errs, "; "))
+ }
+ backup := cnpgBackupObject(namespace, backupName, facts.Cluster, map[string]any{cnpgRequestedFromAnno: name}, cnpgBackupSpecFromSchedule(sched))
+ res, err := cnpgCreateBackup(ctx, c.Dynamic, backup, facts.Cluster)
+ if err != nil {
+ return nil, err
+ }
+ res.Action = action
+ return res, nil
+}
+
+// cnpgBackupSpecFromSchedule copies the complete backup settings of the
+// schedule. Fields the schedule leaves unset stay unset, so the Backup
+// inherits the cluster's defaults exactly as a scheduled run would.
+func cnpgBackupSpecFromSchedule(sched *unstructured.Unstructured) map[string]any {
+ spec := map[string]any{}
+ src, _, _ := unstructured.NestedMap(sched.Object, "spec")
+ for _, key := range []string{"cluster", "method", "pluginConfiguration", "online", "onlineConfiguration", "target"} {
+ if v, ok := src[key]; ok && v != nil {
+ spec[key] = v
+ }
+ }
+ return spec
+}
+
+func GrantFor(action string) (auth.Grant, bool) {
+ if action == "configureArchiving" {
+ return GrantPatchClusters, true
+ }
+ if g, ok := cnpgClusterActionGrants[action]; ok {
+ return g, true
+ }
+ if g, ok := scheduleActionGrants[action]; ok {
+ return g, true
+ }
+ if g, ok := cnpgExtraActionGrants[action]; ok {
+ return g, true
+ }
+ return auth.Grant{}, false
+}
+
+func IsClusterAction(action string) bool {
+ _, ok := clusterActionRunners[action]
+ return ok || action == "configureArchiving"
+}
+
+func ClusterActions() []string {
+ return append(append([]string(nil), clusterActionsOrdered...), "configureArchiving")
+}
+
+func IsScheduleAction(action string) bool { _, ok := scheduleActionGrants[action]; return ok }
diff --git a/internal/cnpg/actions_maintenance.go b/internal/cnpg/actions_maintenance.go
new file mode 100644
index 0000000000..f5a7a00fea
--- /dev/null
+++ b/internal/cnpg/actions_maintenance.go
@@ -0,0 +1,105 @@
+package cnpg
+
+import (
+ "context"
+ "net/http"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+// CNPGMaintenanceFacts is spec.nodeMaintenanceWindow as the confirmation binds
+// it. ReusePVC is the effective value: the CRD defaults it to true.
+type CNPGMaintenanceFacts struct {
+ Declared bool `json:"declared"`
+ InProgress bool `json:"inProgress"`
+ ReusePVC bool `json:"reusePVC"`
+}
+
+type CNPGMaintenanceActions struct {
+ SetMaintenance integration.ActionCapability `json:"setMaintenance"`
+ UnsetMaintenance integration.ActionCapability `json:"unsetMaintenance"`
+}
+
+func init() {
+ for _, a := range []struct {
+ action string
+ inProgress bool
+ }{{"setMaintenance", true}, {"unsetMaintenance", false}} {
+ clusterActionRunners[a.action] = cnpgClusterRunner{binds: []string{"maintenance"}, run: cnpgRunMaintenance(a.inProgress)}
+ cnpgClusterActionGrants[a.action] = GrantPatchClusters
+ clusterActionsOrdered = append(clusterActionsOrdered, a.action)
+ }
+}
+
+func cnpgMaintenanceFactsOf(cluster *unstructured.Unstructured) CNPGMaintenanceFacts {
+ win, declared, _ := unstructured.NestedMap(cluster.Object, "spec", "nodeMaintenanceWindow")
+ m := CNPGMaintenanceFacts{Declared: declared, ReusePVC: true}
+ if !declared {
+ return m
+ }
+ if v, ok := win["inProgress"].(bool); ok {
+ m.InProgress = v
+ }
+ if v, ok := win["reusePVC"].(bool); ok {
+ m.ReusePVC = v
+ }
+ return m
+}
+
+func cnpgGuardSetMaintenance(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if f.Maintenance.InProgress {
+ return "Maintenance is already in progress"
+ }
+ return ""
+}
+
+func cnpgGuardUnsetMaintenance(f CNPGClusterFacts) string {
+ if r := cnpgGuardCommon(f); r != "" {
+ return r
+ }
+ if !f.Maintenance.InProgress {
+ return "Maintenance is not in progress"
+ }
+ return ""
+}
+
+// cnpgMaintenanceParams.ReusePVC is required for both directions: `kubectl
+// cnpg maintenance unset` writes its --reusePVC flag (default false) too, and
+// the dialog sends the value it showed rather than a default nobody reviewed.
+type cnpgMaintenanceParams struct {
+ ReusePVC *bool `json:"reusePVC"`
+}
+
+func cnpgRunMaintenance(inProgress bool) func(context.Context, *cnpgClusterRun) (*CNPGActionResult, error) {
+ return func(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgMaintenanceParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.ReusePVC == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.reusePVC is required")
+ }
+ guard, msg := cnpgGuardSetMaintenance, "Maintenance mode set"
+ if !inProgress {
+ guard, msg = cnpgGuardUnsetMaintenance, "Maintenance mode lifted"
+ }
+ if r := guard(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ err := integration.MergePatchAtVersion(ctx, x.c.Dynamic, ClusterGVR, x.cluster, map[string]any{
+ "spec": map[string]any{"nodeMaintenanceWindow": map[string]any{
+ "inProgress": inProgress,
+ "reusePVC": *p.ReusePVC,
+ }},
+ })
+ if err != nil {
+ return nil, err
+ }
+ return &CNPGActionResult{Message: msg}, nil
+ }
+}
diff --git a/internal/cnpg/actions_maintenance_test.go b/internal/cnpg/actions_maintenance_test.go
new file mode 100644
index 0000000000..374150b601
--- /dev/null
+++ b/internal/cnpg/actions_maintenance_test.go
@@ -0,0 +1,85 @@
+package cnpg
+
+import (
+ "context"
+ "net/http"
+ "testing"
+
+ "k8s.io/apimachinery/pkg/runtime"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func cnpgMaintenanceFactsMap(declared, inProgress, reusePVC bool) map[string]any {
+ f := cnpgActionFacts()
+ f["maintenance"] = map[string]any{"declared": declared, "inProgress": inProgress, "reusePVC": reusePVC}
+ return f
+}
+
+// Set and unset write exactly the two fields `kubectl cnpg maintenance`
+// writes, locked on the resourceVersion, with the reusePVC the dialog showed.
+func TestCNPGActionMaintenanceWritesLikeKubectlCNPG(t *testing.T) {
+ inMaintenance := cnpgActionCluster(func(o map[string]any) {
+ o["spec"].(map[string]any)["nodeMaintenanceWindow"] = map[string]any{"inProgress": true, "reusePVC": false}
+ })
+ for _, tc := range []struct {
+ action string
+ cluster runtime.Object
+ facts map[string]any
+ reusePVC bool
+ inProgress bool
+ }{
+ {"setMaintenance", cnpgActionCluster(nil), cnpgMaintenanceFactsMap(false, false, true), true, true},
+ {"unsetMaintenance", inMaintenance, cnpgMaintenanceFactsMap(true, true, false), false, false},
+ } {
+ t.Run(tc.action, func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{tc.cluster})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", tc.action, cnpgActionReq(t, tc.facts, map[string]any{"reusePVC": tc.reusePVC}))
+ if err != nil {
+ t.Fatalf("%s: %v", tc.action, err)
+ }
+ if len(env.patches) != 1 || env.patches[0].GetSubresource() != "" {
+ t.Fatalf("patches = %v, want one spec patch", env.patches)
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ if md, _ := body["metadata"].(map[string]any); md["resourceVersion"] != "42" || md["annotations"] != nil {
+ t.Errorf("metadata = %v, want only the resourceVersion lock", md)
+ }
+ win := body["spec"].(map[string]any)["nodeMaintenanceWindow"].(map[string]any)
+ if len(win) != 2 || win["inProgress"] != tc.inProgress || win["reusePVC"] != tc.reusePVC {
+ t.Errorf("nodeMaintenanceWindow = %v", win)
+ }
+ })
+ }
+}
+
+func TestCNPGActionMaintenanceGuardsAndBinding(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ ctx := context.Background()
+
+ // Nothing is in progress: unset is refused.
+ _, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "unsetMaintenance", cnpgActionReq(t, cnpgMaintenanceFactsMap(false, false, true), map[string]any{"reusePVC": true}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Errorf("unset while not in progress = %v, want blocked", err)
+ }
+ // The reviewed values differ from the cluster's: 409 changed, nothing written.
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "setMaintenance", cnpgActionReq(t, cnpgMaintenanceFactsMap(true, false, false), map[string]any{"reusePVC": true}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeChanged {
+ t.Errorf("stale maintenance facts = %v, want 409 changed", err)
+ }
+ // Missing binding or missing reusePVC is a malformed request.
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "setMaintenance", cnpgActionReq(t, cnpgActionFacts(), map[string]any{"reusePVC": true}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Errorf("unbound maintenance = %v, want 400", err)
+ }
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "setMaintenance", cnpgActionReq(t, cnpgMaintenanceFactsMap(false, false, true), nil))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Errorf("missing reusePVC = %v, want 400", err)
+ }
+ if len(env.patches) != 0 {
+ t.Errorf("refused requests wrote: %v", env.patches)
+ }
+ if g, ok := GrantFor("setMaintenance"); !ok || g != GrantPatchClusters {
+ t.Errorf("setMaintenance grant = %v", g)
+ }
+}
diff --git a/internal/cnpg/actions_test.go b/internal/cnpg/actions_test.go
new file mode 100644
index 0000000000..cf79104a57
--- /dev/null
+++ b/internal/cnpg/actions_test.go
@@ -0,0 +1,1287 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "net/http"
+ "strings"
+ "testing"
+ "time"
+
+ authv1 "k8s.io/api/authorization/v1"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/api/resource"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ dynamicfake "k8s.io/client-go/dynamic/fake"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+ k8stesting "k8s.io/client-go/testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+var cnpgActionTestNow = time.Date(2026, 9, 29, 10, 11, 12, 345678000, time.UTC)
+
+const cnpgActionTestUID = "cluster-uid-1"
+
+func cnpgActionCluster(mut func(obj map[string]any)) *unstructured.Unstructured {
+ obj := map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1",
+ "kind": "Cluster",
+ "metadata": map[string]any{
+ "name": "pg", "namespace": "db", "uid": cnpgActionTestUID, "resourceVersion": "42",
+ "annotations": map[string]any{},
+ },
+ "spec": map[string]any{
+ "instances": int64(3),
+ "plugins": []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io", "isWALArchiver": true, "parameters": map[string]any{"barmanObjectName": "store"}}},
+ "backup": map[string]any{"volumeSnapshot": map[string]any{"className": "csi"}},
+ },
+ "status": map[string]any{
+ "currentPrimary": "pg-1",
+ "targetPrimary": "pg-1",
+ "phase": cnpgPhaseHealthy,
+ "instanceNames": []any{"pg-1", "pg-2", "pg-3"},
+ "instancesStatus": map[string]any{
+ "healthy": []any{"pg-1", "pg-2", "pg-3"},
+ },
+ "pluginStatus": []any{map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "backupCapabilities": []any{"TYPE_BACKUP"},
+ }},
+ "conditions": []any{
+ map[string]any{"type": "ContinuousArchiving", "status": "True", "reason": "ContinuousArchivingSuccess", "message": "", "lastTransitionTime": "2026-09-01T00:00:00Z"},
+ map[string]any{"type": "Ready", "status": "True", "reason": "ClusterIsReady", "message": "Cluster is Ready", "lastTransitionTime": "2026-09-01T00:00:00Z"},
+ },
+ },
+ }
+ if mut != nil {
+ mut(obj)
+ }
+ return &unstructured.Unstructured{Object: obj}
+}
+
+func cnpgActionPod(name, uid string, ready bool) *corev1.Pod {
+ status := corev1.ConditionFalse
+ if ready {
+ status = corev1.ConditionTrue
+ }
+ controller := true
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{
+ Name: name, Namespace: "db", UID: types.UID(uid),
+ Labels: map[string]string{"cnpg.io/cluster": "pg", "cnpg.io/podRole": "instance"},
+ OwnerReferences: []metav1.OwnerReference{{
+ APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg", UID: cnpgActionTestUID, Controller: &controller,
+ }},
+ },
+ Status: corev1.PodStatus{Conditions: []corev1.PodCondition{{Type: corev1.PodReady, Status: status}}},
+ }
+}
+
+type cnpgActionEnv struct {
+ dyn *dynamicfake.FakeDynamicClient
+ typed *k8sfake.Clientset
+ patches []k8stesting.PatchAction
+ creates []*unstructured.Unstructured
+ deletes []k8stesting.DeleteAction
+ // denied lists "verb resource[/subresource]" the caller's access review refuses.
+ denied map[string]bool
+}
+
+func (e *cnpgActionEnv) clients() ActionClients {
+ return ActionClients{Dynamic: e.dyn, Typed: e.typed, Now: func() time.Time { return cnpgActionTestNow }}
+}
+
+func newCNPGActionEnv(t *testing.T, objs []runtime.Object, pods ...runtime.Object) *cnpgActionEnv {
+ t.Helper()
+ listKinds := map[schema.GroupVersionResource]string{
+ ClusterGVR: "ClusterList", cnpgBackupGVR: "BackupList", ScheduleGVR: "ScheduledBackupList",
+ cnpgPoolerGVR: "PoolerList", cnpgDatabaseGVR: "DatabaseList", cnpgPublGVR: "PublicationList", cnpgSubscrGVR: "SubscriptionList",
+ }
+ env := &cnpgActionEnv{
+ dyn: dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), listKinds, objs...),
+ typed: k8sfake.NewSimpleClientset(pods...),
+ }
+ env.dyn.PrependReactor("patch", "*", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ env.patches = append(env.patches, a.(k8stesting.PatchAction))
+ return true, &unstructured.Unstructured{Object: map[string]any{}}, nil
+ })
+ env.dyn.PrependReactor("create", "backups", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ env.creates = append(env.creates, a.(k8stesting.CreateAction).GetObject().(*unstructured.Unstructured))
+ return false, nil, nil
+ })
+ env.typed.PrependReactor("create", "selfsubjectaccessreviews", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ review := a.(k8stesting.CreateAction).GetObject().(*authv1.SelfSubjectAccessReview)
+ attrs := review.Spec.ResourceAttributes
+ res := attrs.Resource
+ if attrs.Subresource != "" {
+ res += "/" + attrs.Subresource
+ }
+ review.Status.Allowed = !env.denied[attrs.Verb+" "+res]
+ return true, review, nil
+ })
+ env.typed.PrependReactor("delete", "pods", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ env.deletes = append(env.deletes, a.(k8stesting.DeleteAction))
+ return true, nil, nil
+ })
+ return env
+}
+
+func cnpgActionReq(t *testing.T, facts map[string]any, params any) integration.ActionRequest {
+ t.Helper()
+ req := integration.ActionRequest{ReviewedContext: "kind-test", UID: cnpgActionTestUID}
+ if facts != nil {
+ b, _ := json.Marshal(facts)
+ req.Facts = b
+ }
+ if params != nil {
+ b, _ := json.Marshal(params)
+ req.Params = b
+ }
+ return req
+}
+
+// The facts a dialog would echo back for the default fixture.
+func cnpgActionFacts() map[string]any {
+ return map[string]any{
+ "generation": int64(0),
+ "currentPrimary": "pg-1",
+ "targetPrimary": "pg-1",
+ "hibernation": "",
+ "fencedInstances": map[string]any{"raw": "", "all": false, "instances": []any{}},
+ "instances": []any{map[string]any{"pod": "pg-1"}},
+ }
+}
+
+func cnpgActionPatchBody(t *testing.T, p k8stesting.PatchAction) map[string]any {
+ t.Helper()
+ if p.GetPatchType() != types.MergePatchType {
+ t.Fatalf("patch type = %s, want merge", p.GetPatchType())
+ }
+ var body map[string]any
+ if err := json.Unmarshal(p.GetPatch(), &body); err != nil {
+ t.Fatalf("patch body: %v", err)
+ }
+ return body
+}
+
+func cnpgActionAnnotations(t *testing.T, body map[string]any) map[string]any {
+ t.Helper()
+ md, _ := body["metadata"].(map[string]any)
+ if md["resourceVersion"] != "42" {
+ t.Errorf("patch is not locked on resourceVersion 42: %v", md)
+ }
+ a, _ := md["annotations"].(map[string]any)
+ return a
+}
+
+func cnpgActionStatus(t *testing.T, err error) (*integration.ActionError, bool) {
+ t.Helper()
+ var ae *integration.ActionError
+ ok := errors.As(err, &ae)
+ return ae, ok
+}
+
+func TestCNPGActionBackupCreatesExactBody(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ online := false
+ res, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{
+ "method": "plugin", "pluginName": "barman-cloud.cloudnative-pg.io",
+ "target": "primary", "online": online,
+ }))
+ if err != nil {
+ t.Fatalf("backup: %v", err)
+ }
+ if res.Backup != "pg-20260929101112" {
+ t.Errorf("backup name = %q, want pg-20260929101112", res.Backup)
+ }
+ if len(env.creates) != 1 {
+ t.Fatalf("creates = %d, want 1", len(env.creates))
+ }
+ got := env.creates[0].Object
+ want := map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1",
+ "kind": "Backup",
+ "metadata": map[string]any{
+ "name": "pg-20260929101112", "namespace": "db",
+ "labels": map[string]any{"cnpg.io/cluster": "pg"},
+ },
+ "spec": map[string]any{
+ "cluster": map[string]any{"name": "pg"},
+ "method": "plugin",
+ "pluginConfiguration": map[string]any{"name": "barman-cloud.cloudnative-pg.io"},
+ "target": "primary",
+ "online": false,
+ },
+ }
+ gb, _ := json.Marshal(got)
+ wb, _ := json.Marshal(want)
+ if string(gb) != string(wb) {
+ t.Errorf("backup body\n got %s\nwant %s", gb, wb)
+ }
+}
+
+func TestCNPGActionBackupRejectsBarmanDestinationParameters(t *testing.T) {
+ for _, key := range []string{"barmanObjectName", "serverName"} {
+ for _, value := range []string{"override", ""} {
+ t.Run(key+"="+value, func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{
+ "method": "plugin", "pluginName": "barman-cloud.cloudnative-pg.io", "pluginParameters": map[string]string{key: value},
+ }))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusBadRequest || !strings.Contains(err.Error(), "the barman-cloud plugin takes its destination from the Cluster") || !strings.Contains(err.Error(), "pluginParameters."+key) || len(env.creates) != 0 {
+ t.Fatalf("backup: %v, creates=%d, want 400 explaining the ignored parameter and no create", err, len(env.creates))
+ }
+ })
+ }
+ }
+}
+
+func TestCNPGActionBackupForwardsThirdPartyParameters(t *testing.T) {
+ const plugin = "other.example.com"
+ cluster := cnpgActionCluster(func(o map[string]any) {
+ o["spec"].(map[string]any)["plugins"] = []any{map[string]any{"name": plugin}}
+ o["status"].(map[string]any)["pluginStatus"] = []any{map[string]any{"name": plugin, "backupCapabilities": []any{"TYPE_BACKUP"}}}
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{cluster})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{
+ "method": "plugin", "pluginName": plugin, "pluginParameters": map[string]string{"barmanObjectName": "third-party-store", "serverName": "third-party-server", "custom": "value"},
+ }))
+ if err != nil || len(env.creates) != 1 {
+ t.Fatalf("backup: %v, creates=%d", err, len(env.creates))
+ }
+ got, _, err := unstructured.NestedMap(env.creates[0].Object, "spec", "pluginConfiguration")
+ body, _ := json.Marshal(got)
+ if err != nil || string(body) != `{"name":"other.example.com","parameters":{"barmanObjectName":"third-party-store","custom":"value","serverName":"third-party-server"}}` {
+ t.Fatalf("pluginConfiguration = %s, err=%v", body, err)
+ }
+}
+
+// An omitted target must stay omitted so the Backup inherits the cluster's.
+func TestCNPGActionBackupOmittedTargetInherits(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot", "name": "manual-one"})); err != nil {
+ t.Fatalf("backup: %v", err)
+ }
+ spec := env.creates[0].Object["spec"].(map[string]any)
+ if _, ok := spec["target"]; ok {
+ t.Errorf("target was invented: %v", spec)
+ }
+ if _, ok := spec["pluginConfiguration"]; ok {
+ t.Errorf("volumeSnapshot backup carries a pluginConfiguration: %v", spec)
+ }
+}
+
+// The confirmation showed the cluster's backup target; a Backup that inherits
+// it must not run against a target changed since.
+func TestCNPGActionBackupBindsInheritedTarget(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(func(obj map[string]any) {
+ obj["spec"].(map[string]any)["backup"].(map[string]any)["target"] = "primary"
+ })})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot", "name": "manual-one"}))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed", err)
+ }
+ if len(env.creates) != 0 {
+ t.Error("a Backup was created against an unreviewed target")
+ }
+ facts := cnpgActionFacts()
+ facts["backupTarget"] = "primary"
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, facts, map[string]any{"method": "volumeSnapshot", "name": "manual-one"})); err != nil {
+ t.Fatalf("backup with the reviewed target: %v", err)
+ }
+ // An explicit target doesn't read the cluster's, so its change doesn't matter.
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot", "name": "manual-two", "target": "prefer-standby"})); err != nil {
+ t.Fatalf("backup with an explicit target: %v", err)
+ }
+}
+
+func TestCNPGActionBackupRejectsScheduleRunName(t *testing.T) {
+ sched := &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "ScheduledBackup",
+ "metadata": map[string]any{"name": "nightly", "namespace": "db"},
+ "spec": map[string]any{"cluster": map[string]any{"name": "pg"}},
+ }}
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), sched})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot", "name": "nightly-20260929000000"}))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusBadRequest {
+ t.Fatalf("err = %v, want 400", err)
+ }
+ if len(env.creates) != 0 {
+ t.Error("a Backup was created under a scheduled run's name")
+ }
+}
+
+func TestCNPGActionBackupRejectsUnknownMethodAndBadName(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ for _, params := range []map[string]any{
+ {"method": "barmanObjectStore"},
+ {"method": "plugin", "pluginName": "other"},
+ {"method": "volumeSnapshot", "name": "Not_A_Name"},
+ } {
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), params))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Errorf("%v: err = %v, want 400", params, err)
+ }
+ }
+}
+
+// A create whose answer was lost is resolved by reading the same name back.
+func TestCNPGActionBackupTimeoutResolvesByReadingBack(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ attempts := 0
+ env.dyn.PrependReactor("create", "backups", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ attempts++
+ obj := a.(k8stesting.CreateAction).GetObject()
+ if err := env.dyn.Tracker().Create(cnpgBackupGVR, obj, "db"); err != nil {
+ t.Fatalf("tracker: %v", err)
+ }
+ return true, nil, apierrors.NewTimeoutError("request timed out", 1)
+ })
+ res, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot"}))
+ if err != nil {
+ t.Fatalf("backup: %v", err)
+ }
+ if !res.ResolvedAfterTimeout || attempts != 1 {
+ t.Errorf("resolved=%v attempts=%d, want resolved after exactly one create", res.ResolvedAfterTimeout, attempts)
+ }
+}
+
+func TestCNPGActionSwitchoverPatchesStatusLikeKubectlCNPG(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)},
+ cnpgActionPod("pg-1", "u1", true), cnpgActionPod("pg-2", "u2", true), cnpgActionPod("pg-3", "u3", true))
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ if err != nil {
+ t.Fatalf("switchover: %v", err)
+ }
+ if len(env.patches) != 1 {
+ t.Fatalf("patches = %d, want 1", len(env.patches))
+ }
+ p := env.patches[0]
+ if p.GetSubresource() != "status" {
+ t.Errorf("subresource = %q, want status", p.GetSubresource())
+ }
+ body := cnpgActionPatchBody(t, p)
+ if body["metadata"].(map[string]any)["resourceVersion"] != "42" {
+ t.Error("switchover is not locked on resourceVersion")
+ }
+ status := body["status"].(map[string]any)
+ if status["targetPrimary"] != "pg-2" || status["phase"] != "Switchover in progress" || status["phaseReason"] != "Switching over to pg-2" {
+ t.Errorf("status = %v", status)
+ }
+ if status["targetPrimaryTimestamp"] != "2026-09-29T10:11:12.345678Z" {
+ t.Errorf("targetPrimaryTimestamp = %v, want RFC3339Micro", status["targetPrimaryTimestamp"])
+ }
+ conds := status["conditions"].([]any)
+ if len(conds) != 2 {
+ t.Fatalf("conditions = %v, want the full list (merge patch replaces it)", conds)
+ }
+ ready := conds[1].(map[string]any)
+ if ready["type"] != "Ready" || ready["status"] != "False" || ready["reason"] != "ClusterIsNotReady" || ready["message"] != "Cluster Is Not Ready" {
+ t.Errorf("Ready = %v", ready)
+ }
+ if ready["lastTransitionTime"] != "2026-09-29T10:11:12Z" {
+ t.Errorf("Ready lastTransitionTime = %v, want the switch time", ready["lastTransitionTime"])
+ }
+ if conds[0].(map[string]any)["lastTransitionTime"] != "2026-09-01T00:00:00Z" {
+ t.Error("an unrelated condition was rewritten")
+ }
+}
+
+func TestCNPGActionSwitchoverTargetChecks(t *testing.T) {
+ fencedPg2 := cnpgActionCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: `["pg-2"]`}
+ })
+ facts := cnpgActionFacts()
+ facts["fencedInstances"] = map[string]any{"raw": `["pg-2"]`}
+ t.Run("fenced target", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{fencedPg2}, cnpgActionPod("pg-2", "u2", true))
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, facts, map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeBlocked {
+ t.Fatalf("err = %v, want 409 blocked", err)
+ }
+ })
+ t.Run("recreated pod", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-2", "u2-new", true))
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed", err)
+ }
+ })
+ t.Run("not ready", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-2", "u2", false))
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Fatalf("err = %v, want 409 blocked", err)
+ }
+ })
+ t.Run("pod of another cluster", func(t *testing.T) {
+ foreign := cnpgActionPod("pg-2", "u2", true)
+ foreign.OwnerReferences[0].UID = "other-uid"
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, foreign)
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ if _, ok := cnpgActionStatus(t, err); !ok || len(env.patches) != 0 {
+ t.Fatalf("err = %v patches=%d, want a refusal and no write", err, len(env.patches))
+ }
+ })
+}
+
+func TestCNPGActionFactMismatchReturnsCurrentFacts(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-2", "u2", true))
+ facts := cnpgActionFacts()
+ facts["currentPrimary"] = "pg-3"
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "switchover",
+ cnpgActionReq(t, facts, map[string]any{"target": "pg-2", "targetPodUID": "u2"}))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed", err)
+ }
+ if cur, ok := ae.Current.(CNPGClusterFacts); !ok || cur.CurrentPrimary != "pg-1" {
+ t.Errorf("current = %#v, want the facts as they are now", ae.Current)
+ }
+ if len(env.patches) != 0 {
+ t.Error("a write went out after a fact mismatch")
+ }
+}
+
+func TestCNPGActionUIDMismatchAndMissingFacts(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ req := cnpgActionReq(t, cnpgActionFacts(), nil)
+ req.UID = "old-uid"
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restart", req)
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed for a recreated Cluster", err)
+ }
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restart", cnpgActionReq(t, map[string]any{}, nil))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Fatalf("err = %v, want 400 when the bound facts are missing", err)
+ }
+}
+
+func TestCNPGActionReviewedContext(t *testing.T) {
+ if err := integration.CheckReviewedContext("kind-a", "kind-a"); err != nil {
+ t.Errorf("matching context refused: %v", err)
+ }
+ ae, ok := cnpgActionStatus(t, integration.CheckReviewedContext("kind-a", "prod"))
+ if !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeContextChanged {
+ t.Errorf("mismatch = %v, want 409 context_changed", ae)
+ }
+}
+
+func TestCNPGActionAnnotationWrites(t *testing.T) {
+ hibernated := cnpgActionCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgHibernateAnnotation: "on"}
+ })
+ hibFacts := cnpgActionFacts()
+ hibFacts["hibernation"] = "on"
+ for _, tc := range []struct {
+ action string
+ cluster *unstructured.Unstructured
+ facts map[string]any
+ key string
+ want any
+ }{
+ {"restart", cnpgActionCluster(nil), cnpgActionFacts(), cnpgRestartAnnotation, "2026-09-29T10:11:12Z"},
+ {"reload", cnpgActionCluster(nil), cnpgActionFacts(), cnpgReloadAnnotation, "2026-09-29T10:11:12.345678Z"},
+ {"hibernate", cnpgActionCluster(nil), cnpgActionFacts(), cnpgHibernateAnnotation, "on"},
+ {"rehydrate", hibernated, hibFacts, cnpgHibernateAnnotation, "off"},
+ } {
+ t.Run(tc.action, func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{tc.cluster})
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", tc.action, cnpgActionReq(t, tc.facts, nil)); err != nil {
+ t.Fatalf("%s: %v", tc.action, err)
+ }
+ if len(env.patches) != 1 || env.patches[0].GetSubresource() != "" {
+ t.Fatalf("patches = %v, want one main-resource patch", env.patches)
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ annos := cnpgActionAnnotations(t, body)
+ if len(annos) != 1 || annos[tc.key] != tc.want {
+ t.Errorf("annotations = %v, want only %s=%v", annos, tc.key, tc.want)
+ }
+ if _, ok := body["spec"]; ok {
+ t.Error("an annotation action wrote spec")
+ }
+ })
+ }
+}
+
+// Hibernation blocks the disruptive actions but never rehydration.
+func TestCNPGActionHibernatedGuards(t *testing.T) {
+ f := CNPGClusterFacts{Hibernated: true, Hibernation: "on", CurrentPrimary: "pg-1", BackupMethods: []CNPGBackupMethodFact{{Method: "volumeSnapshot", Capability: "backup"}}}
+ for name, reason := range map[string]string{
+ "backup": cnpgGuardBackup(f), "switchover": cnpgGuardSwitchover(f), "restart": cnpgGuardRestart(f), "fence": cnpgGuardFence(f),
+ } {
+ if reason == "" {
+ t.Errorf("%s allowed on a hibernated cluster", name)
+ }
+ }
+ if r := cnpgGuardRehydrate(f); r != "" {
+ t.Errorf("rehydrate blocked: %s", r)
+ }
+ f.Terminating = true
+ if cnpgGuardRehydrate(f) == "" || cnpgGuardUnfence(CNPGClusterFacts{Terminating: true}) == "" {
+ t.Error("a terminating cluster still offers actions")
+ }
+ replica := CNPGClusterFacts{IsReplicaCluster: true, CurrentPrimary: "pg-1"}
+ if cnpgGuardSwitchover(replica) == "" {
+ t.Error("switchover offered on a replica cluster")
+ }
+ inFlight := CNPGClusterFacts{CurrentPrimary: "pg-1", TargetPrimary: "pg-2"}
+ if cnpgGuardSwitchover(inFlight) == "" || cnpgGuardRestart(inFlight) == "" {
+ t.Error("switchover/restart offered while a switchover is in flight")
+ }
+}
+
+func TestCNPGActionRestartInstance(t *testing.T) {
+ t.Run("standby delete carries the UID precondition", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-2", "u2", false))
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restartInstance",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"pod": "pg-2", "podUID": "u2"})); err != nil {
+ t.Fatalf("restartInstance: %v", err)
+ }
+ if len(env.deletes) != 1 {
+ t.Fatalf("deletes = %d, want 1", len(env.deletes))
+ }
+ d := env.deletes[0]
+ opts := d.GetDeleteOptions()
+ if d.GetName() != "pg-2" || opts.Preconditions == nil || opts.Preconditions.UID == nil || *opts.Preconditions.UID != "u2" {
+ t.Errorf("delete %s with options %+v, want UID precondition u2", d.GetName(), opts)
+ }
+ })
+ t.Run("primary restarts in place through status", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-1", "u1", true))
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restartInstance",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"pod": "pg-1", "podUID": "u1"})); err != nil {
+ t.Fatalf("restartInstance: %v", err)
+ }
+ if len(env.deletes) != 0 || len(env.patches) != 1 || env.patches[0].GetSubresource() != "status" {
+ t.Fatalf("deletes=%d patches=%v, want one status patch", len(env.deletes), env.patches)
+ }
+ status := cnpgActionPatchBody(t, env.patches[0])["status"].(map[string]any)
+ if status["phase"] != "Primary instance is being restarted in-place" || status["phaseReason"] != "Requested by the user" {
+ t.Errorf("status = %v", status)
+ }
+ if _, ok := status["targetPrimary"]; ok {
+ t.Error("in-place restart touched targetPrimary")
+ }
+ })
+ t.Run("primary allowed while waiting for user action", func(t *testing.T) {
+ waiting := cnpgActionCluster(func(o map[string]any) { o["status"].(map[string]any)["phase"] = cnpgPhaseWaitingForUser })
+ env := newCNPGActionEnv(t, []runtime.Object{waiting}, cnpgActionPod("pg-1", "u1", true))
+ if _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restartInstance",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"pod": "pg-1", "podUID": "u1"})); err != nil {
+ t.Fatalf("restartInstance: %v", err)
+ }
+ })
+ t.Run("primary refused mid-upgrade", func(t *testing.T) {
+ upgrading := cnpgActionCluster(func(o map[string]any) { o["status"].(map[string]any)["phase"] = "Upgrading cluster" })
+ env := newCNPGActionEnv(t, []runtime.Object{upgrading}, cnpgActionPod("pg-1", "u1", true))
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "restartInstance",
+ cnpgActionReq(t, cnpgActionFacts(), map[string]any{"pod": "pg-1", "podUID": "u1"}))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Fatalf("err = %v, want 409 blocked", err)
+ }
+ })
+}
+
+func TestCNPGActionFencing(t *testing.T) {
+ withFence := func(raw string) (*unstructured.Unstructured, map[string]any) {
+ c := cnpgActionCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: raw}
+ })
+ f := cnpgActionFacts()
+ f["fencedInstances"] = map[string]any{"raw": raw}
+ return c, f
+ }
+ run := func(t *testing.T, raw, action string, params map[string]any) (*cnpgActionEnv, error) {
+ t.Helper()
+ c, f := withFence(raw)
+ env := newCNPGActionEnv(t, []runtime.Object{c})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", action, cnpgActionReq(t, f, params))
+ return env, err
+ }
+ fencedValue := func(t *testing.T, env *cnpgActionEnv) any {
+ t.Helper()
+ if len(env.patches) != 1 {
+ t.Fatalf("patches = %d, want 1", len(env.patches))
+ }
+ annos := cnpgActionAnnotations(t, cnpgActionPatchBody(t, env.patches[0]))
+ v, ok := annos[cnpgFencedAnnotation]
+ if !ok {
+ t.Fatal("patch does not set the fencing annotation")
+ }
+ return v
+ }
+
+ t.Run("fence adds to the JSON string", func(t *testing.T) {
+ env, err := run(t, `["pg-3"]`, "fence", map[string]any{"instances": []string{"pg-2"}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if v := fencedValue(t, env); v != `["pg-2","pg-3"]` {
+ t.Errorf("annotation = %v", v)
+ }
+ })
+ t.Run("fence all", func(t *testing.T) {
+ env, err := run(t, "", "fence", map[string]any{"instances": "*"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if v := fencedValue(t, env); v != `["*"]` {
+ t.Errorf("annotation = %v", v)
+ }
+ })
+ t.Run("lifting the last fence removes the annotation", func(t *testing.T) {
+ env, err := run(t, `["pg-2"]`, "unfence", map[string]any{"instances": []string{"pg-2"}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if v := fencedValue(t, env); v != nil {
+ t.Errorf("annotation = %v, want removed", v)
+ }
+ })
+ t.Run("lifting one instance under wildcard is refused", func(t *testing.T) {
+ env, err := run(t, `["*"]`, "unfence", map[string]any{"instances": []string{"pg-2"}})
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusConflict || ae.Code != cnpgCodeAllFenced {
+ t.Fatalf("err = %v, want 409 all_fenced", err)
+ }
+ if len(env.patches) != 0 {
+ t.Error("a write went out")
+ }
+ })
+ t.Run("explicit conversion from wildcard", func(t *testing.T) {
+ env, err := run(t, `["*"]`, "unfence", map[string]any{"instances": []string{"pg-2"}, "convertFromAll": true, "remaining": []string{"pg-3", "pg-1"}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if v := fencedValue(t, env); v != `["pg-1","pg-3"]` {
+ t.Errorf("annotation = %v", v)
+ }
+ })
+ t.Run("conversion with a stale remaining list", func(t *testing.T) {
+ _, err := run(t, `["*"]`, "unfence", map[string]any{"instances": []string{"pg-2"}, "convertFromAll": true, "remaining": []string{"pg-1"}})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed", err)
+ }
+ })
+ t.Run("malformed annotation blocks writes", func(t *testing.T) {
+ for _, action := range []string{"fence", "unfence"} {
+ _, err := run(t, `pg-2`, action, map[string]any{"instances": "*"})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Errorf("%s: err = %v, want 409 blocked", action, err)
+ }
+ }
+ })
+}
+
+func cnpgSnapshotBackup(name, phase string, online any) *unstructured.Unstructured {
+ spec := map[string]any{"cluster": map[string]any{"name": "pg"}, "method": "volumeSnapshot"}
+ if online != nil {
+ spec["online"] = online
+ }
+ return &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1",
+ "kind": "Backup",
+ "metadata": map[string]any{"name": name, "namespace": "db"},
+ "spec": spec,
+ "status": map[string]any{"phase": phase},
+ }}
+}
+
+// The operator fences the instance for a cold snapshot and lifts the fence
+// itself; lifting it first would start PostgreSQL under the snapshot.
+func TestCNPGActionUnfenceWaitsForColdSnapshot(t *testing.T) {
+ fenced := func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: `["pg-2"]`}
+ }
+ facts := func() map[string]any {
+ f := cnpgActionFacts()
+ f["fencedInstances"] = map[string]any{"raw": `["pg-2"]`}
+ return f
+ }
+ cases := []struct {
+ name string
+ backups []runtime.Object
+ cluster func(o map[string]any)
+ blocked bool
+ }{
+ {"running cold snapshot", []runtime.Object{cnpgSnapshotBackup("cold", "running", false)}, nil, true},
+ {"cold by the cluster's default", []runtime.Object{cnpgSnapshotBackup("cold", "started", nil)}, func(o map[string]any) {
+ o["spec"].(map[string]any)["backup"] = map[string]any{"volumeSnapshot": map[string]any{"className": "csi", "online": false}}
+ }, true},
+ {"finished cold snapshot", []runtime.Object{cnpgSnapshotBackup("cold", "completed", false)}, nil, false},
+ {"running online snapshot", []runtime.Object{cnpgSnapshotBackup("hot", "running", true)}, nil, false},
+ {"online by default", []runtime.Object{cnpgSnapshotBackup("hot", "running", nil)}, nil, false},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ c := cnpgActionCluster(func(o map[string]any) {
+ fenced(o)
+ if tc.cluster != nil {
+ tc.cluster(o)
+ }
+ })
+ env := newCNPGActionEnv(t, append([]runtime.Object{c}, tc.backups...), cnpgActionPod("pg-1", "u1", true))
+ resp, err := newTestReader(nil).ClusterCapabilities(context.Background(), env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got := !resp.Actions.Unfence.Allowed; got != tc.blocked {
+ t.Errorf("unfence blocked = %v (%q), want %v", got, resp.Actions.Unfence.Reason, tc.blocked)
+ }
+ if got := !resp.InstanceActions["pg-2"].Unfence.Allowed; got != tc.blocked {
+ t.Errorf("pg-2 unfence blocked = %v, want %v", got, tc.blocked)
+ }
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "unfence", cnpgActionReq(t, facts(), map[string]any{"instances": []string{"pg-2"}}))
+ ae, isBlocked := cnpgActionStatus(t, err)
+ isBlocked = isBlocked && ae.Code == integration.ActionCodeBlocked
+ if isBlocked != tc.blocked {
+ t.Errorf("run: err = %v, want blocked %v", err, tc.blocked)
+ }
+ if tc.blocked && len(env.patches) != 0 {
+ t.Error("a write went out")
+ }
+ })
+ }
+}
+
+func TestCNPGActionMalformedFencingBlocksCapability(t *testing.T) {
+ c := cnpgActionCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: "{not json"}
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{c}, cnpgActionPod("pg-1", "u1", true))
+ srv := newTestReader(nil)
+ ctx := context.Background()
+ resp, err := srv.ClusterCapabilities(ctx, env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if !resp.Facts.FencedInstances.Malformed {
+ t.Error("malformed annotation not reported")
+ }
+ if resp.Actions.Fence.Allowed || resp.Actions.Unfence.Allowed {
+ t.Errorf("fence=%+v unfence=%+v, want both blocked", resp.Actions.Fence, resp.Actions.Unfence)
+ }
+ if !resp.Actions.Restart.Allowed || !resp.Actions.Hibernate.Allowed {
+ t.Error("a malformed fence blocked unrelated actions")
+ }
+}
+
+func TestCNPGActionCapabilitiesPermissionDenied(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)},
+ cnpgActionPod("pg-1", "u1", true), cnpgActionPod("pg-2", "u2", true))
+ srv := newTestReader(nil)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"db"}}
+ perms.SetCanI("patch", Group, "clusters/status", "db", false)
+ perms.SetCanI("get", "", "pods", "db", true)
+ perms.SetCanI("patch", Group, "clusters", "db", true)
+ perms.SetCanI("create", Group, "backups", "db", false)
+ perms.SetCanI("delete", "", "pods", "db", true)
+ srv = newTestReader(perms)
+ ctx := context.Background()
+
+ resp, err := srv.ClusterCapabilities(ctx, env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ sw := resp.Actions.Switchover
+ if sw.Allowed || sw.Permission != integration.PermissionDenied || sw.Grant == nil || *sw.Grant != cnpgGrantPatchStatus.In("db") || !strings.Contains(sw.Reason, "patch clusters/status") {
+ t.Errorf("switchover = %+v, want denied naming patch clusters/status", sw)
+ }
+ if b := resp.Actions.Backup; b.Allowed || b.Grant == nil || *b.Grant != cnpgGrantCreateBackups.In("db") {
+ t.Errorf("backup = %+v, want denied naming create backups", b)
+ }
+ if !resp.Actions.Restart.Allowed || resp.Actions.Restart.Permission != integration.PermissionAllowed {
+ t.Errorf("restart = %+v, want allowed", resp.Actions.Restart)
+ }
+ // The primary's in-place restart needs the status grant; a standby's needs delete pods.
+ if resp.InstanceActions["pg-1"].Restart.Allowed || !resp.InstanceActions["pg-2"].Restart.Allowed {
+ t.Errorf("instance restarts = %+v", resp.InstanceActions)
+ }
+ if resp.UID != cnpgActionTestUID || resp.ResourceVersion != "42" || resp.Context != "kind-test" {
+ t.Errorf("identity = %s/%s/%s", resp.UID, resp.ResourceVersion, resp.Context)
+ }
+}
+
+func TestCNPGActionCapabilitiesPlansAndEffects(t *testing.T) {
+ c := cnpgActionCluster(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ spec["primaryUpdateStrategy"] = "unsupervised"
+ spec["primaryUpdateMethod"] = "switchover"
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: `["pg-3"]`}
+ })
+ obj := func(kind, name string, spec map[string]any) *unstructured.Unstructured {
+ return &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": kind,
+ "metadata": map[string]any{"name": name, "namespace": "db"}, "spec": spec,
+ }}
+ }
+ cl := map[string]any{"name": "pg"}
+ env := newCNPGActionEnv(t, []runtime.Object{c,
+ obj("Pooler", "pg-rw", map[string]any{"cluster": cl}),
+ obj("Pooler", "other-rw", map[string]any{"cluster": map[string]any{"name": "other"}}),
+ obj("ScheduledBackup", "nightly", map[string]any{"cluster": cl}),
+ obj("ScheduledBackup", "paused", map[string]any{"cluster": cl, "suspend": true}),
+ obj("Database", "app", map[string]any{"cluster": cl}),
+ })
+ resp, err := newTestReader(nil).ClusterCapabilities(context.Background(), env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ effects := []string{}
+ for _, s := range resp.RestartPlan.Steps {
+ effects = append(effects, s.Instance+":"+s.Effect)
+ }
+ if got := strings.Join(effects, ","); got != "pg-2:recreate,pg-3:skipped_fenced,pg-1:switchover" {
+ t.Errorf("restart plan = %s", got)
+ }
+ he := resp.HibernateEffects
+ if strings.Join(he.Poolers.Names, ",") != "pg-rw" || strings.Join(he.UnsuspendedScheduledBackups.Names, ",") != "nightly" || len(he.Databases.Names) != 1 {
+ t.Errorf("hibernate effects = %+v", he)
+ }
+ if !he.Volumes.Available {
+ t.Errorf("volumes = %+v", he.Volumes)
+ }
+ methods := []string{}
+ for _, m := range resp.Facts.BackupMethods {
+ methods = append(methods, m.Method+"/"+m.Capability)
+ }
+ if strings.Join(methods, ",") != "plugin/backup,volumeSnapshot/backup" {
+ t.Errorf("backup methods = %v", methods)
+ }
+}
+
+func TestCNPGActionBackupMethodCapability(t *testing.T) {
+ c := cnpgActionCluster(func(o map[string]any) {
+ delete(o["spec"].(map[string]any), "backup")
+ o["status"].(map[string]any)["pluginStatus"] = []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io", "backupCapabilities": []any{}}}
+ })
+ m := cnpgBackupMethods(c)
+ if len(m) != 1 || m[0].Capability != "none" {
+ t.Fatalf("methods = %+v, want the plugin marked none", m)
+ }
+ if cnpgGuardBackup(CNPGClusterFacts{BackupMethods: m}) == "" {
+ t.Error("backup offered with no backup-capable method")
+ }
+ c = cnpgActionCluster(func(o map[string]any) { delete(o["status"].(map[string]any), "pluginStatus") })
+ if m := cnpgBackupMethods(c); m[0].Capability != "unknown" {
+ t.Errorf("unreported plugin = %+v, want unknown", m[0])
+ }
+ // barman-cloud 0.14 reports status without a backupCapabilities field at all.
+ c = cnpgActionCluster(func(o map[string]any) {
+ o["status"].(map[string]any)["pluginStatus"] = []any{map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "version": "0.14.0",
+ "capabilities": []any{"TYPE_RECONCILER_HOOKS", "TYPE_LIFECYCLE_SERVICE"},
+ }}
+ })
+ if m := cnpgBackupMethods(c); m[0].Capability != "unknown" {
+ t.Errorf("plugin without a backupCapabilities field = %+v, want unknown", m[0])
+ }
+}
+
+func cnpgActionSchedule(mut func(o map[string]any)) *unstructured.Unstructured {
+ o := map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "ScheduledBackup",
+ "metadata": map[string]any{"name": "nightly", "namespace": "db", "uid": "sched-uid", "resourceVersion": "7", "generation": int64(3)},
+ "spec": map[string]any{
+ "cluster": map[string]any{"name": "pg"},
+ "schedule": "0 0 0 * * *",
+ "method": "plugin",
+ "pluginConfiguration": map[string]any{"name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "store"}},
+ "online": false,
+ "onlineConfiguration": map[string]any{"immediateCheckpoint": true},
+ "backupOwnerReference": "self",
+ },
+ "status": map[string]any{"nextScheduleTime": "2026-09-28T00:00:00Z"},
+ }
+ if mut != nil {
+ mut(o)
+ }
+ return &unstructured.Unstructured{Object: o}
+}
+
+func TestCNPGScheduleRunDestinationGuard(t *testing.T) {
+ for _, tc := range []struct {
+ name, method string
+ configure bool
+ override bool
+ }{
+ {"default without destination", "", false, false},
+ {"in-tree without destination", "barmanObjectStore", false, false},
+ {"in-tree with destination", "barmanObjectStore", true, false},
+ {"snapshot without configuration", "volumeSnapshot", false, false},
+ {"snapshot configured", "volumeSnapshot", true, false},
+ {"plugin without destination", "plugin", false, false},
+ {"plugin configured", "plugin", true, false},
+ {"plugin schedule parameter without Cluster destination", "plugin", false, true},
+ {"plugin conflicting schedule parameter with Cluster destination", "plugin", true, true},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ cluster := cnpgActionCluster(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ delete(spec, "backup")
+ delete(spec, "plugins")
+ if tc.method == "plugin" {
+ p := map[string]any{"name": "barman-cloud.cloudnative-pg.io"}
+ if tc.configure {
+ p["parameters"] = map[string]any{"barmanObjectName": "store"}
+ }
+ spec["plugins"] = []any{p}
+ } else if tc.configure {
+ if tc.method == "volumeSnapshot" {
+ spec["backup"] = map[string]any{"volumeSnapshot": map[string]any{}}
+ } else {
+ spec["backup"] = map[string]any{"barmanObjectStore": map[string]any{"destinationPath": "s3://backups"}}
+ }
+ }
+ })
+ schedule := cnpgActionSchedule(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ spec["method"] = tc.method
+ delete(spec, "pluginConfiguration")
+ if tc.method == "plugin" {
+ spec["pluginConfiguration"] = map[string]any{"name": "barman-cloud.cloudnative-pg.io"}
+ if tc.override {
+ spec["pluginConfiguration"].(map[string]any)["parameters"] = map[string]any{"barmanObjectName": "schedule-store", "serverName": "schedule-server"}
+ }
+ }
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, schedule})
+ caps, err := newTestReader(nil).ScheduleCapabilities(context.Background(), env.clients(), "kind-test", "db", "nightly")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if caps.Actions.Run.Allowed != tc.configure {
+ t.Fatalf("run capability = %+v", caps.Actions.Run)
+ }
+ if !tc.configure && caps.Actions.Run.Reason != "Configure a backup destination on pg first" {
+ t.Fatalf("reason = %s", caps.Actions.Run.Reason)
+ }
+ _, err = RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", integration.ActionRequest{UID: "sched-uid", Facts: json.RawMessage(`{"generation":3,"clusterUID":"cluster-uid-1","clusterGeneration":0}`)})
+ if tc.configure {
+ if err != nil || len(env.creates) != 1 {
+ t.Fatalf("configured run: %v, creates=%d", err, len(env.creates))
+ }
+ } else {
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Code != "blocked" || len(env.creates) != 0 {
+ t.Fatalf("unconfigured run: %v, creates=%d", err, len(env.creates))
+ }
+ }
+ })
+ }
+}
+
+func TestCNPGScheduleRunMethodMismatch(t *testing.T) {
+ cluster := cnpgActionCluster(nil)
+ schedule := cnpgActionSchedule(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ delete(spec, "method")
+ delete(spec, "pluginConfiguration")
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, schedule})
+ caps, err := newTestReader(nil).ScheduleCapabilities(context.Background(), env.clients(), "kind-test", "db", "nightly")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if caps.Actions.Run.Allowed || caps.Actions.Run.ReasonCode != "backup_destination" || !strings.Contains(caps.Actions.Run.Reason, "No barmanObjectStore destination on pg") || !strings.Contains(caps.Actions.Run.Reason, "Use method plugin") {
+ t.Fatalf("method mismatch should identify the existing plugin destination: %+v", caps.Actions.Run)
+ }
+ _, err = RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", integration.ActionRequest{UID: "sched-uid", Facts: json.RawMessage(`{"generation":3,"clusterUID":"cluster-uid-1","clusterGeneration":0}`)})
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Code != "blocked" || len(env.creates) != 0 {
+ t.Fatalf("mismatched method must not create a Backup: %v, creates=%d", err, len(env.creates))
+ }
+}
+
+func TestCNPGBackupDestinationGuard(t *testing.T) {
+ cluster := cnpgActionCluster(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ delete(spec, "plugins")
+ spec["backup"] = map[string]any{"barmanObjectStore": map[string]any{}}
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{cluster})
+ caps, err := newTestReader(nil).ClusterCapabilities(context.Background(), env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if caps.Actions.Backup.Allowed || caps.Actions.Backup.Reason != "Configure a backup destination on pg first" {
+ t.Fatalf("backup capability = %+v", caps.Actions.Backup)
+ }
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "barmanObjectStore"}))
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Code != "blocked" || len(env.creates) != 0 {
+ t.Fatalf("backup: %v, creates=%d", err, len(env.creates))
+ }
+}
+
+func TestCNPGBackupDestinationGuardRequiresClusterPluginDestination(t *testing.T) {
+ cluster := cnpgActionCluster(func(o map[string]any) {
+ o["spec"].(map[string]any)["plugins"] = []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io"}}
+ })
+ if reason := cnpgBackupDestinationGuard(cluster, "plugin", "barman-cloud.cloudnative-pg.io"); reason == "" {
+ t.Fatal("missing Cluster plugin destination was allowed")
+ }
+ env := newCNPGActionEnv(t, []runtime.Object{cluster})
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{
+ "method": "plugin", "pluginName": "barman-cloud.cloudnative-pg.io", "pluginParameters": map[string]string{"barmanObjectName": "request-store"},
+ }))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Status != http.StatusBadRequest || !strings.Contains(err.Error(), "the barman-cloud plugin takes its destination from the Cluster") || len(env.creates) != 0 {
+ t.Fatalf("request override: %v, creates=%d, want 400 and no create", err, len(env.creates))
+ }
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{
+ "method": "plugin", "pluginName": "barman-cloud.cloudnative-pg.io",
+ }))
+ ae, ok = cnpgActionStatus(t, err)
+ if !ok || ae.Code != "blocked" || len(env.creates) != 0 {
+ t.Fatalf("missing Cluster destination: %v, creates=%d, want blocked and no create", err, len(env.creates))
+ }
+}
+
+func TestCNPGActionScheduleRunCopiesSettings(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ req := integration.ActionRequest{ReviewedContext: "kind-test", UID: "sched-uid", Facts: json.RawMessage(`{"generation":3,"clusterUID":"cluster-uid-1","clusterGeneration":0}`)}
+ res, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", req)
+ if err != nil {
+ t.Fatalf("run: %v", err)
+ }
+ if res.Backup != "nightly-manual-20260929101112" || len(env.creates) != 1 {
+ t.Fatalf("backup = %q creates = %d", res.Backup, len(env.creates))
+ }
+ b := env.creates[0]
+ if b.GetLabels()["cnpg.io/cluster"] != "pg" || b.GetAnnotations()[cnpgRequestedFromAnno] != "nightly" {
+ t.Errorf("metadata = %v / %v", b.GetLabels(), b.GetAnnotations())
+ }
+ if len(b.GetOwnerReferences()) != 0 {
+ t.Error("the manual run is owned by the schedule")
+ }
+ spec := b.Object["spec"].(map[string]any)
+ got, _ := json.Marshal(spec)
+ want := `{"cluster":{"name":"pg"},"method":"plugin","online":false,"onlineConfiguration":{"immediateCheckpoint":true},"pluginConfiguration":{"name":"barman-cloud.cloudnative-pg.io","parameters":{"barmanObjectName":"store"}}}`
+ if string(got) != want {
+ t.Errorf("spec\n got %s\nwant %s", got, want)
+ }
+}
+
+func TestCNPGActionScheduleRunBindsReviewedSettings(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ req := integration.ActionRequest{ReviewedContext: "kind-test", UID: "sched-uid", Facts: json.RawMessage(`{"generation":2}`)}
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", req)
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Status != http.StatusConflict {
+ t.Fatalf("stale generation: err = %v, want 409", err)
+ }
+ if len(env.creates) != 0 {
+ t.Error("a Backup was created from settings the user did not review")
+ }
+ req.Facts = nil
+ if _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", req); !errors.As(err, &ae) || ae.Status != http.StatusBadRequest {
+ t.Fatalf("missing generation: err = %v, want 400", err)
+ }
+}
+
+func TestCNPGActionScheduleSuspendResume(t *testing.T) {
+ t.Run("resume reports catch-up", func(t *testing.T) {
+ suspended := cnpgActionSchedule(func(o map[string]any) { o["spec"].(map[string]any)["suspend"] = true })
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), suspended})
+ res, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "resume",
+ integration.ActionRequest{ReviewedContext: "kind-test", UID: "sched-uid", Facts: json.RawMessage(`{"suspended":true}`)})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if !res.CatchUp {
+ t.Error("catchUp not reported for a past nextScheduleTime")
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ if body["spec"].(map[string]any)["suspend"] != false || body["metadata"].(map[string]any)["resourceVersion"] != "7" {
+ t.Errorf("patch = %v", body)
+ }
+ })
+ t.Run("suspend binds the reviewed state", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "suspend",
+ integration.ActionRequest{ReviewedContext: "kind-test", UID: "sched-uid", Facts: json.RawMessage(`{"suspended":true}`)})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Fatalf("err = %v, want 409 changed", err)
+ }
+ })
+}
+
+func TestCNPGActionApiserverConflictIsNotRetried(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)})
+ calls := 0
+ env.dyn.PrependReactor("patch", "clusters", func(k8stesting.Action) (bool, runtime.Object, error) {
+ calls++
+ return true, nil, apierrors.NewConflict(schema.GroupResource{Group: Group, Resource: "clusters"}, "pg", errors.New("the object has been modified"))
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "hibernate", cnpgActionReq(t, cnpgActionFacts(), nil))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged || ae.Current == nil {
+ t.Fatalf("err = %v, want 409 changed with current facts", err)
+ }
+ if calls != 1 {
+ t.Errorf("patch attempts = %d, want 1", calls)
+ }
+}
+
+func TestCNPGActionCapabilitiesRestoreNeedsCreateClusters(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)},
+ cnpgActionPod("pg-1", "u1", true), cnpgActionPod("pg-2", "u2", true))
+ for _, allowed := range []bool{false, true} {
+ srv := newTestReader(nil)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"db"}}
+ perms.SetCanI("create", Group, "clusters", "db", allowed)
+ srv = newTestReader(perms)
+ ctx := context.Background()
+ resp, err := srv.ClusterCapabilities(ctx, env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ got := resp.Actions.Restore
+ if allowed {
+ if !got.Allowed || got.Permission != integration.PermissionAllowed {
+ t.Errorf("restore with create clusters = %+v, want allowed", got)
+ }
+ continue
+ }
+ if got.Allowed || got.Permission != integration.PermissionDenied || !strings.Contains(got.Reason, "create clusters (postgresql.cnpg.io) in namespace db") {
+ t.Errorf("restore without create clusters = %+v, want denied naming the grant", got)
+ }
+ }
+}
+
+func TestParseCNPGFencedRejectsNonLists(t *testing.T) {
+ for raw, malformed := range map[string]bool{"": false, `[]`: false, `["pg-1"]`: false, `null`: true, ` null `: true, `"pg-1"`: true, `{}`: true} {
+ if got := parseCNPGFenced(raw).Malformed; got != malformed {
+ t.Errorf("parseCNPGFenced(%q).Malformed = %v, want %v", raw, got, malformed)
+ }
+ }
+}
+
+func TestCNPGClusterFactsArchivingFailing(t *testing.T) {
+ failing := cnpgActionCluster(func(obj map[string]any) {
+ obj["status"].(map[string]any)["conditions"] = []any{map[string]any{"type": "ContinuousArchiving", "status": "False"}}
+ })
+ if f, _ := cnpgClusterFactsOf(context.Background(), nil, failing); !f.ArchivingFailing {
+ t.Error("ContinuousArchiving=False must read as archiving failing")
+ }
+ if f, _ := cnpgClusterFactsOf(context.Background(), nil, cnpgActionCluster(nil)); f.ArchivingFailing {
+ t.Error("ContinuousArchiving=True must not read as archiving failing")
+ }
+}
+
+func TestCNPGIsReplicaClusterMatchesOperator(t *testing.T) {
+ for _, c := range []struct {
+ replica map[string]any
+ want bool
+ }{
+ {nil, false},
+ {map[string]any{"enabled": true, "source": "east"}, true},
+ {map[string]any{"enabled": false, "primary": "east"}, false},
+ {map[string]any{"primary": "east", "source": "east"}, true},
+ {map[string]any{"primary": "pg", "source": "east"}, false},
+ {map[string]any{"self": "west", "primary": "west"}, false},
+ {map[string]any{"source": "east"}, true},
+ } {
+ cluster := cnpgActionCluster(func(obj map[string]any) {
+ if c.replica != nil {
+ obj["spec"].(map[string]any)["replica"] = c.replica
+ }
+ })
+ if got := cnpgIsReplicaCluster(cluster); got != c.want {
+ t.Errorf("replica %v: got %v, want %v", c.replica, got, c.want)
+ }
+ }
+}
+
+func TestCNPGHibernateCapacityIsNotRequestedSize(t *testing.T) {
+ cluster := cnpgActionCluster(nil)
+ claim := func(name, capacity, requested string) *corev1.PersistentVolumeClaim {
+ pvc := &corev1.PersistentVolumeClaim{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: "db", Labels: map[string]string{clusterLabel: "pg", instanceNameLabel: name}}}
+ if requested != "" {
+ pvc.Spec.Resources.Requests = corev1.ResourceList{corev1.ResourceStorage: resource.MustParse(requested)}
+ }
+ if capacity != "" {
+ pvc.Status.Capacity = corev1.ResourceList{corev1.ResourceStorage: resource.MustParse(capacity)}
+ }
+ return pvc
+ }
+ env := newCNPGActionEnv(t, []runtime.Object{cluster}, claim("pg-1", "1Gi", "2Gi"), claim("pg-2", "", "2Gi"))
+ volumes := cnpgHibernateEffectsOf(context.Background(), env.clients(), cluster).Volumes
+ if !volumes.Available || len(volumes.Items) != 2 {
+ t.Fatalf("volumes = %+v", volumes)
+ }
+ if volumes.Items[0].Capacity != "1Gi" || volumes.Items[0].Requested != "2Gi" {
+ t.Fatalf("reported = %+v", volumes.Items[0])
+ }
+ if volumes.Items[1].Capacity != "" || volumes.Items[1].Requested != "2Gi" {
+ t.Fatalf("unreported = %+v", volumes.Items[1])
+ }
+}
+
+func TestCNPGBackupBindsClusterSpecGeneration(t *testing.T) {
+ for _, schedule := range []bool{false, true} {
+ for _, change := range []string{"plugin destination", "snapshot settings", "status only", "replacement"} {
+ t.Run(fmt.Sprintf("schedule=%v/%s", schedule, change), func(t *testing.T) {
+ cluster := cnpgActionCluster(nil)
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, cnpgActionSchedule(nil)})
+ if change == "replacement" {
+ cluster.SetUID("replacement")
+ } else if change == "plugin destination" {
+ plugins, _, _ := unstructured.NestedSlice(cluster.Object, "spec", "plugins")
+ plugins[0].(map[string]any)["parameters"].(map[string]any)["barmanObjectName"] = "other-store"
+ unstructured.SetNestedSlice(cluster.Object, plugins, "spec", "plugins")
+ } else if change == "snapshot settings" {
+ unstructured.SetNestedField(cluster.Object, "other-class", "spec", "backup", "volumeSnapshot", "className")
+ } else {
+ unstructured.SetNestedField(cluster.Object, "reconciling", "status", "phaseReason")
+ }
+ if change == "plugin destination" || change == "snapshot settings" {
+ cluster.SetGeneration(1)
+ }
+ if _, err := env.dyn.Resource(ClusterGVR).Namespace("db").Update(context.Background(), cluster, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ var err error
+ if schedule {
+ _, err = RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "run", integration.ActionRequest{UID: "sched-uid", Facts: json.RawMessage(`{"generation":3,"clusterUID":"cluster-uid-1","clusterGeneration":0}`)})
+ } else {
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "backup", cnpgActionReq(t, cnpgActionFacts(), map[string]any{"method": "volumeSnapshot", "target": "primary"}))
+ }
+ if change == "status only" {
+ if err != nil || len(env.creates) != 1 {
+ t.Fatalf("status-only change blocked backup: %v creates=%d", err, len(env.creates))
+ }
+ return
+ }
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Code != integration.ActionCodeChanged || len(env.creates) != 0 {
+ t.Fatalf("changed Cluster: %v creates=%d", err, len(env.creates))
+ }
+ })
+ }
+ }
+}
diff --git a/internal/cnpg/activity.go b/internal/cnpg/activity.go
new file mode 100644
index 0000000000..ce812a35e4
--- /dev/null
+++ b/internal/cnpg/activity.go
@@ -0,0 +1,227 @@
+package cnpg
+
+import (
+ "context"
+ "sort"
+ "strings"
+ "time"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/pkg/resourceid"
+ pkgtimeline "github.com/skyhook-io/radar/pkg/timeline"
+)
+
+// The scan bounds attribution across retained rows, before the requested window.
+// Outgrowing it reports truncation rather than silently dropping older attribution.
+const cnpgActivityScanLimit = 10000
+
+type ActivityTimeline interface {
+ Query(context.Context, pkgtimeline.QueryOptions) ([]pkgtimeline.TimelineEvent, error)
+}
+
+type ActivityOptions struct {
+ Since time.Time
+ Until time.Time
+ Limit int
+}
+
+type ActivityReadError struct {
+ Operation string
+ Err error
+}
+
+func (e *ActivityReadError) Error() string { return e.Err.Error() }
+func (e *ActivityReadError) Unwrap() error { return e.Err }
+
+// CNPGClusterActivityResponse is GET /api/cnpg/clusters/{namespace}/{name}/activity.
+// Oldest is the earliest row the store still holds for the namespace — the
+// floor below which absence means "not retained", not "didn't happen".
+// AttributionSince is the earliest visible row that carries this Cluster's
+// retained cnpg.io/cluster attribution; before it, deleted children cannot be
+// attributed. Both are null when nothing is held.
+type CNPGClusterActivityResponse struct {
+ Events []pkgtimeline.TimelineEvent `json:"events"`
+ Oldest *time.Time `json:"oldest"`
+ AttributionSince *time.Time `json:"attributionSince"`
+ Truncated bool `json:"truncated"`
+}
+
+type cnpgActivityKind struct {
+ group, resource string
+}
+
+// cnpgActivityKinds are the kinds whose rows can belong to one Cluster: the
+// Cluster itself, its instance Pods, and every namespaced CNPG kind.
+var cnpgActivityKinds = func() map[string]cnpgActivityKind {
+ out := map[string]cnpgActivityKind{"/Pod": {group: "", resource: "pods"}}
+ for _, k := range workspaceKinds {
+ if !k.ClusterScoped {
+ out[k.Group+"/"+k.Kind] = cnpgActivityKind{group: k.Group, resource: k.Resource}
+ }
+ }
+ return out
+}()
+
+func cnpgActivityKindNames() []string {
+ seen := map[string]bool{}
+ var out []string
+ for key := range cnpgActivityKinds {
+ _, kind, _ := strings.Cut(key, "/")
+ if !seen[kind] {
+ seen[kind] = true
+ out = append(out, kind)
+ }
+ }
+ sort.Strings(out)
+ return out
+}
+
+// cnpgRowAttribution decides whether a timeline row is about the named
+// Cluster. Rows about the Cluster match by identity; instance Pods by their
+// controller owner; CNPG children by the retained cnpg.io/cluster label, which
+// survives their deletion. liveUID is the UID of the Cluster that exists now
+// under this name, or "" when none does; when set, only Pods it controlled
+// count, so a previous same-named Cluster's instances don't merge into a
+// recreated one's history.
+func cnpgRowAttribution(e *pkgtimeline.TimelineEvent, name, liveUID string) (matched, labelled bool) {
+ group := resourceid.GroupFromAPIVersion(e.APIVersion)
+ if _, ok := cnpgActivityKinds[group+"/"+e.Kind]; !ok {
+ return false, false
+ }
+ labelled = e.Labels[pkgtimeline.CNPGClusterLabel] == name
+ switch {
+ case e.Kind == "Cluster" && group == Group:
+ return e.Name == name, false
+ case e.Kind == "Pod" && group == "":
+ o := e.Owner
+ owned := o != nil && o.Kind == "Cluster" && o.Name == name && resourceid.GroupFromAPIVersion(o.APIVersion) == Group &&
+ (liveUID == "" || o.UID == liveUID)
+ return owned, false
+ default:
+ return labelled, labelled
+ }
+}
+
+func (s *Reader) ClusterActivity(ctx context.Context, store ActivityTimeline, clusterContext, namespace, name string, options ActivityOptions) (*CNPGClusterActivityResponse, error) {
+ rows, err := store.Query(ctx, pkgtimeline.QueryOptions{
+ Namespaces: []string{namespace},
+ Kinds: cnpgActivityKindNames(),
+ APIGroups: []string{"", Group, barmanGroup},
+ ClusterContext: clusterContext,
+ IncludeManaged: true,
+ IncludeK8sEvents: true,
+ Limit: cnpgActivityScanLimit,
+ })
+ if err != nil {
+ return nil, &ActivityReadError{Operation: "query activity", Err: err}
+ }
+ var liveUID string
+ jobPods := map[string]WorkspacePod{}
+ if cache := s.Observations.Cache; cache != nil {
+ clusters, err := s.Observations.DynamicList(ctx, cache, "Cluster", Group, namespace)
+ live, _ := SelectCluster(clusters, err, namespace, name)
+ if live != nil {
+ liveUID = string(live.GetUID())
+ _, _, jobs, _, _ := s.workspaceReadPods(ctx, cache, []string{namespace}, clusterUIDs([]*unstructured.Unstructured{live}))
+ for _, raw := range jobs {
+ pod := raw.(WorkspacePod)
+ jobPods[string(pod.Metadata.UID)] = pod
+ }
+ }
+ }
+ scanCapped := len(rows) >= cnpgActivityScanLimit
+
+ // Attribution is carried by the subject's own rows; K8s Event rows about
+ // a subject whose enrichment was already gone carry only its UID.
+ attributedUIDs := map[string]bool{}
+ matched := make([]bool, len(rows))
+ labelled := make([]bool, len(rows))
+ for i := range rows {
+ matched[i], labelled[i] = cnpgRowAttribution(&rows[i], name, liveUID)
+ if !matched[i] && rows[i].Kind == "Pod" && resourceid.GroupFromAPIVersion(rows[i].APIVersion) == "" {
+ if pod, ok := jobPods[rows[i].UID]; ok && rows[i].Owner != nil {
+ o := rows[i].Owner
+ matched[i] = controlledBy(pod.Metadata.OwnerReferences, resourceid.GroupFromAPIVersion(o.APIVersion), o.Kind, o.Name, types.UID(o.UID))
+ }
+ }
+ if matched[i] && rows[i].UID != "" {
+ attributedUIDs[rows[i].UID] = true
+ }
+ }
+
+ allowed := map[string]bool{}
+ var eventsAllowed *bool
+ canList := func(e *pkgtimeline.TimelineEvent) bool {
+ if e.Source == pkgtimeline.SourceK8sEvent {
+ if eventsAllowed == nil {
+ ok := s.Access.CanRead(ctx, "", "events", namespace, "list")
+ eventsAllowed = &ok
+ }
+ if !*eventsAllowed {
+ return false
+ }
+ }
+ key := resourceid.GroupFromAPIVersion(e.APIVersion) + "/" + e.Kind
+ ok, seen := allowed[key]
+ if !seen {
+ target, known := cnpgActivityKinds[key]
+ ok = known && s.Access.CanRead(ctx, target.group, target.resource, namespace, "list")
+ allowed[key] = ok
+ }
+ return ok
+ }
+
+ resp := CNPGClusterActivityResponse{Events: []pkgtimeline.TimelineEvent{}}
+ seenIDs := map[string]bool{}
+ var windowed []pkgtimeline.TimelineEvent
+ for i := range rows {
+ e := &rows[i]
+ if !matched[i] && (e.UID == "" || !attributedUIDs[e.UID]) {
+ continue
+ }
+ if !canList(e) || seenIDs[e.ID] {
+ continue
+ }
+ seenIDs[e.ID] = true
+ if labelled[i] && (resp.AttributionSince == nil || e.Timestamp.Before(*resp.AttributionSince)) {
+ t := e.Timestamp.UTC()
+ resp.AttributionSince = &t
+ }
+ if e.Timestamp.Before(options.Since) || (!options.Until.IsZero() && e.Timestamp.After(options.Until)) {
+ continue
+ }
+ windowed = append(windowed, *e)
+ }
+ sort.SliceStable(windowed, func(i, j int) bool {
+ if !windowed[i].Timestamp.Equal(windowed[j].Timestamp) {
+ return windowed[i].Timestamp.After(windowed[j].Timestamp)
+ }
+ return windowed[i].ID < windowed[j].ID
+ })
+ resp.Truncated = scanCapped || len(windowed) > options.Limit
+ if len(windowed) > options.Limit {
+ windowed = windowed[:options.Limit]
+ }
+ if windowed != nil {
+ resp.Events = windowed
+ }
+
+ oldest, err := store.Query(ctx, pkgtimeline.QueryOptions{
+ Namespaces: []string{namespace},
+ ClusterContext: clusterContext,
+ IncludeManaged: true,
+ IncludeK8sEvents: true,
+ SequenceOrder: pkgtimeline.SequenceOrderAscending,
+ Limit: 1,
+ })
+ if err != nil {
+ return nil, &ActivityReadError{Operation: "query retention floor", Err: err}
+ }
+ if len(oldest) > 0 {
+ t := oldest[0].Timestamp.UTC()
+ resp.Oldest = &t
+ }
+ return &resp, nil
+}
diff --git a/internal/cnpg/activity_test.go b/internal/cnpg/activity_test.go
new file mode 100644
index 0000000000..519b96e297
--- /dev/null
+++ b/internal/cnpg/activity_test.go
@@ -0,0 +1,286 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "reflect"
+ "slices"
+ "testing"
+ "time"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ pkgtimeline "github.com/skyhook-io/radar/pkg/timeline"
+)
+
+type activityTimeline struct {
+ rows, oldest []pkgtimeline.TimelineEvent
+ rowsErr, oldestErr error
+ queries []pkgtimeline.QueryOptions
+ contexts []context.Context
+}
+
+func (s *activityTimeline) Query(ctx context.Context, options pkgtimeline.QueryOptions) ([]pkgtimeline.TimelineEvent, error) {
+ s.queries = append(s.queries, options)
+ s.contexts = append(s.contexts, ctx)
+ if options.SequenceOrder == pkgtimeline.SequenceOrderAscending {
+ return s.oldest, s.oldestErr
+ }
+ return s.rows, s.rowsErr
+}
+
+func activityReader() *Reader {
+ return &Reader{Access: Access{CanRead: func(context.Context, string, string, string, string) bool { return true }}}
+}
+
+func activityEvent(id, version, kind, name, uid string, at time.Time) pkgtimeline.TimelineEvent {
+ return pkgtimeline.TimelineEvent{ID: id, APIVersion: version, Kind: kind, Name: name, Namespace: "db", UID: uid, Timestamp: at, Source: pkgtimeline.SourceInformer}
+}
+
+func activityEventIDs(response *CNPGClusterActivityResponse) []string {
+ ids := make([]string, 0, len(response.Events))
+ for _, event := range response.Events {
+ ids = append(ids, event.ID)
+ }
+ return ids
+}
+
+func TestClusterActivityRetainedAttributionAndQueryScope(t *testing.T) {
+ now := time.Date(2026, 10, 6, 12, 0, 0, 0, time.UTC)
+ labelled := func(event pkgtimeline.TimelineEvent) pkgtimeline.TimelineEvent {
+ event.Labels = map[string]string{pkgtimeline.CNPGClusterLabel: "pg"}
+ return event
+ }
+ pod := activityEvent("pod", "v1", "Pod", "pg-1", "pod-uid", now.Add(-time.Hour))
+ pod.Owner = &pkgtimeline.OwnerInfo{APIVersion: Group + "/v1", Kind: "Cluster", Name: "pg", UID: "cluster-uid"}
+ podEvent := activityEvent("pod-event", "v1", "Pod", "pg-1", "pod-uid", now.Add(-30*time.Minute))
+ podEvent.Source = pkgtimeline.SourceK8sEvent
+ backupEvent := activityEvent("backup-event", Group+"/v1", "Backup", "deleted", "backup-uid", now.Add(-10*time.Minute))
+ backupEvent.Source = pkgtimeline.SourceK8sEvent
+ oldest := now.Add(-72 * time.Hour)
+ oldLabel := now.Add(-48 * time.Hour)
+ store := &activityTimeline{
+ rows: []pkgtimeline.TimelineEvent{
+ backupEvent, podEvent, pod,
+ labelled(activityEvent("backup-delete", Group+"/v1", "Backup", "deleted", "backup-uid", now.Add(-40*time.Minute))),
+ activityEvent("cluster", Group+"/v1", "Cluster", "pg", "cluster-uid", now.Add(-2*time.Hour)),
+ labelled(activityEvent("old-pooler", Group+"/v1", "Pooler", "pg-pooler", "pooler-uid", oldLabel)),
+ labelled(activityEvent("velero", "velero.io/v1", "Backup", "pg", "velero-uid", now)),
+ activityEvent("capi", "cluster.x-k8s.io/v1beta1", "Cluster", "pg", "capi-uid", now),
+ labelled(activityEvent("impostor", "v1", "Pod", "impostor", "impostor-uid", now)),
+ labelled(activityEvent("cluster-catalog", Group+"/v1", "ClusterImageCatalog", "catalog", "catalog-uid", now)),
+ },
+ oldest: []pkgtimeline.TimelineEvent{activityEvent("unrelated-floor", "v1", "Secret", "unrelated", "secret-uid", oldest)},
+ }
+ ctx, cancel := context.WithCancel(context.Background())
+ defer cancel()
+ response, err := activityReader().ClusterActivity(ctx, store, "captured-context", "db", "pg", ActivityOptions{Since: now.Add(-24 * time.Hour), Limit: 200})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got, want := activityEventIDs(response), []string{"backup-event", "pod-event", "backup-delete", "pod", "cluster"}; !reflect.DeepEqual(got, want) {
+ t.Fatalf("events = %v, want %v", got, want)
+ }
+ if response.AttributionSince == nil || !response.AttributionSince.Equal(oldLabel) || response.Oldest == nil || !response.Oldest.Equal(oldest) || response.Truncated {
+ t.Fatalf("retention/attribution = %+v", response)
+ }
+ if len(store.queries) != 2 {
+ t.Fatalf("queries = %+v", store.queries)
+ }
+ for i, query := range store.queries {
+ if store.contexts[i] != ctx || query.ClusterContext != "captured-context" || !reflect.DeepEqual(query.Namespaces, []string{"db"}) || !query.IncludeManaged || !query.IncludeK8sEvents {
+ t.Fatalf("query escaped captured scope: %+v", query)
+ }
+ if !query.Since.IsZero() || !query.Until.IsZero() {
+ t.Fatalf("window applied before retained attribution: %+v", query)
+ }
+ }
+ scan, floor := store.queries[0], store.queries[1]
+ if scan.Limit != cnpgActivityScanLimit || !reflect.DeepEqual(scan.APIGroups, []string{"", Group, barmanGroup}) || !slices.Contains(scan.Kinds, "Pod") || slices.Contains(scan.Kinds, "ClusterImageCatalog") {
+ t.Fatalf("scan = %+v", scan)
+ }
+ if floor.Limit != 1 || floor.SequenceOrder != pkgtimeline.SequenceOrderAscending || len(floor.Kinds) != 0 || len(floor.APIGroups) != 0 {
+ t.Fatalf("floor narrowed by activity kinds: %+v", floor)
+ }
+}
+
+func TestClusterActivityPermissionFiltersApplyToUIDJoinedEvents(t *testing.T) {
+ now := time.Date(2026, 10, 6, 12, 0, 0, 0, time.UTC)
+ backup := activityEvent("backup", Group+"/v1", "Backup", "deleted", "backup-uid", now)
+ backup.Labels = map[string]string{pkgtimeline.CNPGClusterLabel: "pg"}
+ event := activityEvent("event", Group+"/v1", "Backup", "deleted", "backup-uid", now)
+ event.Source = pkgtimeline.SourceK8sEvent
+ for _, test := range []struct {
+ name string
+ backups, events bool
+ want []string
+ }{
+ {"all", true, true, []string{"backup", "event"}},
+ {"no backups", false, true, []string{}},
+ {"no events", true, false, []string{"backup"}},
+ } {
+ t.Run(test.name, func(t *testing.T) {
+ calls := map[string]int{}
+ reader := activityReader()
+ reader.Access.CanRead = func(_ context.Context, group, resource, namespace, verb string) bool {
+ if namespace != "db" || verb != "list" {
+ t.Fatalf("wrong grant: %s %s/%s in %s", verb, group, resource, namespace)
+ }
+ calls[resource]++
+ if resource == "events" {
+ return test.events
+ }
+ return resource == "backups" && group == Group && test.backups
+ }
+ response, err := reader.ClusterActivity(context.Background(), &activityTimeline{rows: []pkgtimeline.TimelineEvent{event, backup}}, "ctx", "db", "pg", ActivityOptions{Since: now.Add(-time.Hour), Limit: 20})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got := activityEventIDs(response); !reflect.DeepEqual(got, test.want) {
+ t.Fatalf("events = %v, want %v", got, test.want)
+ }
+ if calls["backups"] != 1 || calls["events"] != 1 {
+ t.Fatalf("grants were not memoized by kind: %v", calls)
+ }
+ if !test.backups && response.AttributionSince != nil {
+ t.Fatalf("denied attribution leaked: %+v", response)
+ }
+ })
+ }
+}
+
+func TestClusterActivityWindowSortDedupAndLimit(t *testing.T) {
+ now := time.Date(2026, 10, 6, 12, 0, 0, 0, time.UTC)
+ row := func(id string, at time.Time) pkgtimeline.TimelineEvent {
+ return activityEvent(id, Group+"/v1", "Cluster", "pg", "cluster-uid", at)
+ }
+ store := &activityTimeline{rows: []pkgtimeline.TimelineEvent{
+ row("later", now.Add(time.Nanosecond)), row("z", now), row("a", now), row("a", now.Add(-time.Minute)), row("start", now.Add(-time.Hour)), row("before", now.Add(-time.Hour-time.Nanosecond)),
+ }}
+ for _, test := range []struct {
+ limit int
+ want []string
+ truncated bool
+ }{{20, []string{"a", "z", "start"}, false}, {2, []string{"a", "z"}, true}} {
+ response, err := activityReader().ClusterActivity(context.Background(), store, "ctx", "db", "pg", ActivityOptions{Since: now.Add(-time.Hour), Until: now, Limit: test.limit})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got := activityEventIDs(response); !reflect.DeepEqual(got, test.want) || response.Truncated != test.truncated {
+ t.Fatalf("events = %v truncated=%v", got, response.Truncated)
+ }
+ }
+}
+
+func TestClusterActivityScanCapAndEmptyWireShape(t *testing.T) {
+ for _, count := range []int{0, cnpgActivityScanLimit} {
+ store := &activityTimeline{rows: make([]pkgtimeline.TimelineEvent, count)}
+ response, err := activityReader().ClusterActivity(context.Background(), store, "ctx", "db", "pg", ActivityOptions{Limit: 200})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if response.Truncated != (count == cnpgActivityScanLimit) {
+ t.Fatalf("scan %d: truncated=%v", count, response.Truncated)
+ }
+ body, err := json.Marshal(response)
+ if err != nil {
+ t.Fatal(err)
+ }
+ want := `{"events":[],"oldest":null,"attributionSince":null,"truncated":false}`
+ if count == 0 && string(body) != want {
+ t.Fatalf("empty shape = %s", body)
+ }
+ }
+}
+
+func TestClusterActivityLiveIncarnationAndOptionalEnrichmentFailure(t *testing.T) {
+ now := time.Date(2026, 10, 6, 12, 0, 0, 0, time.UTC)
+ row := func(id, ownerUID string) pkgtimeline.TimelineEvent {
+ event := activityEvent(id, "v1", "Pod", "pg-1", id, now)
+ event.Owner = &pkgtimeline.OwnerInfo{APIVersion: Group + "/v1", Kind: "Cluster", Name: "pg", UID: ownerUID}
+ return event
+ }
+ live := &unstructured.Unstructured{Object: map[string]any{"apiVersion": Group + "/v1", "kind": "Cluster", "metadata": map[string]any{"namespace": "db", "name": "pg", "uid": "new"}}}
+ for _, failed := range []bool{false, true} {
+ reader := activityReader()
+ cache := &k8s.ResourceCache{}
+ reader.Observations.Cache = cache
+ reader.Observations.DynamicList = func(_ context.Context, received *k8s.ResourceCache, kind, group, namespace string) ([]*unstructured.Unstructured, error) {
+ if received != cache || kind != "Cluster" || group != Group || namespace != "db" {
+ t.Fatalf("live read escaped selected cache/identity")
+ }
+ if failed {
+ return nil, errors.New("optional cache enrichment failed")
+ }
+ return []*unstructured.Unstructured{live}, nil
+ }
+ reader.Observations.TypedScope = func(context.Context, *k8s.ResourceCache, []string, string, string) (integration.KindAccess, []string) {
+ return integration.KindAccess{State: integration.KindCoverageDenied}, nil
+ }
+ response, err := reader.ClusterActivity(context.Background(), &activityTimeline{rows: []pkgtimeline.TimelineEvent{row("old", "previous"), row("new", "new")}}, "ctx", "db", "pg", ActivityOptions{Since: now.Add(-time.Hour), Limit: 200})
+ if err != nil {
+ t.Fatal(err)
+ }
+ want := []string{"new"}
+ if failed {
+ want = []string{"new", "old"}
+ }
+ if got := activityEventIDs(response); !reflect.DeepEqual(got, want) {
+ t.Fatalf("enrichment failed=%v: events=%v want=%v", failed, got, want)
+ }
+ }
+}
+
+func TestClusterActivityQueryFailuresPreserveCauseAndResponseMessage(t *testing.T) {
+ cause := errors.New("timeline unavailable")
+ for _, floor := range []bool{false, true} {
+ store := &activityTimeline{}
+ if floor {
+ store.oldestErr = cause
+ } else {
+ store.rowsErr = cause
+ }
+ response, err := activityReader().ClusterActivity(context.Background(), store, "ctx", "db", "pg", ActivityOptions{Limit: 200})
+ var readErr *ActivityReadError
+ if response != nil || !errors.Is(err, cause) || !errors.As(err, &readErr) || err.Error() != cause.Error() {
+ t.Fatalf("error contract changed: response=%v err=%v", response, err)
+ }
+ want := "query activity"
+ if floor {
+ want = "query retention floor"
+ }
+ if readErr.Operation != want {
+ t.Fatalf("operation = %q, want %q", readErr.Operation, want)
+ }
+ }
+}
+
+func TestCNPGRowAttributionUsesExactOwnerAndKindIdentity(t *testing.T) {
+ labels := map[string]string{pkgtimeline.CNPGClusterLabel: "pg"}
+ owner := &pkgtimeline.OwnerInfo{APIVersion: Group + "/v1", Kind: "Cluster", Name: "pg", UID: "current"}
+ for _, test := range []struct {
+ name, version, kind string
+ owner *pkgtimeline.OwnerInfo
+ live types.UID
+ matched, labelled bool
+ }{
+ {"owned instance", "v1", "Pod", owner, "current", true, false},
+ {"previous instance", "v1", "Pod", owner, "recreated", false, false},
+ {"uncontrolled labelled Pod", "v1", "Pod", nil, "", false, false},
+ {"deleted Backup", Group + "/v1", "Backup", nil, "", true, true},
+ {"Barman ObjectStore", barmanGroup + "/v1", "ObjectStore", nil, "", true, true},
+ {"foreign Backup", "velero.io/v1", "Backup", nil, "", false, false},
+ } {
+ t.Run(test.name, func(t *testing.T) {
+ event := &pkgtimeline.TimelineEvent{APIVersion: test.version, Kind: test.kind, Owner: test.owner, Labels: labels}
+ matched, labelled := cnpgRowAttribution(event, "pg", string(test.live))
+ if matched != test.matched || labelled != test.labelled {
+ t.Fatalf("matched=%v labelled=%v", matched, labelled)
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/architecture_test.go b/internal/cnpg/architecture_test.go
new file mode 100644
index 0000000000..33c71b6612
--- /dev/null
+++ b/internal/cnpg/architecture_test.go
@@ -0,0 +1,98 @@
+package cnpg
+
+import (
+ "go/ast"
+ "go/parser"
+ "go/token"
+ "os"
+ "strconv"
+ "strings"
+ "testing"
+)
+
+func TestServiceDependencyBoundary(t *testing.T) {
+ entries, err := os.ReadDir(".")
+ if err != nil {
+ t.Fatal(err)
+ }
+ for _, entry := range entries {
+ if !strings.HasSuffix(entry.Name(), ".go") || strings.HasSuffix(entry.Name(), "_test.go") {
+ continue
+ }
+ fset := token.NewFileSet()
+ file, err := parser.ParseFile(fset, entry.Name(), nil, 0)
+ if err != nil {
+ t.Fatal(err)
+ }
+ imports := make(map[string]string)
+ for _, spec := range file.Imports {
+ path, err := strconv.Unquote(spec.Path.Value)
+ if err != nil {
+ t.Fatal(err)
+ }
+ name := path[strings.LastIndex(path, "/")+1:]
+ if spec.Name != nil {
+ name = spec.Name.Name
+ }
+ imports[name] = path
+ if path == "github.com/skyhook-io/radar/internal/server" || strings.HasPrefix(path, "github.com/go-chi/chi") {
+ t.Errorf("%s: service imports browser routing package %s", fset.Position(spec.Pos()), path)
+ }
+ if name == "." && path == "github.com/skyhook-io/radar/internal/k8s" {
+ t.Errorf("%s: a dot import hides singleton access", fset.Position(spec.Pos()))
+ }
+ }
+ checkSignature := func(signature ast.Node) {
+ ast.Inspect(signature, func(node ast.Node) bool {
+ selector, ok := node.(*ast.SelectorExpr)
+ if !ok {
+ return true
+ }
+ owner, ok := selector.X.(*ast.Ident)
+ if ok && imports[owner.Name] == "net/http" && (selector.Sel.Name == "Request" || selector.Sel.Name == "ResponseWriter") {
+ t.Errorf("%s: service operation or port exposes browser HTTP type %s", fset.Position(selector.Pos()), selector.Sel.Name)
+ }
+ return true
+ })
+ }
+ for _, declaration := range file.Decls {
+ switch declaration := declaration.(type) {
+ case *ast.FuncDecl:
+ if declaration.Name.IsExported() {
+ checkSignature(declaration.Type)
+ }
+ case *ast.GenDecl:
+ for _, spec := range declaration.Specs {
+ if typ, ok := spec.(*ast.TypeSpec); ok && typ.Name.IsExported() {
+ checkSignature(typ.Type)
+ }
+ }
+ }
+ }
+ ast.Inspect(file, func(node ast.Node) bool {
+ call, ok := node.(*ast.CallExpr)
+ if !ok {
+ return true
+ }
+ selector, ok := call.Fun.(*ast.SelectorExpr)
+ if !ok {
+ return true
+ }
+ owner, ok := selector.X.(*ast.Ident)
+ if !ok {
+ return true
+ }
+ if imports[owner.Name] == "github.com/skyhook-io/radar/internal/prometheus" && selector.Sel.Name == "GetClient" {
+ t.Errorf("%s: service resolves global metrics client; use caller-scoped dependencies", fset.Position(call.Pos()))
+ }
+ if imports[owner.Name] != "github.com/skyhook-io/radar/internal/k8s" {
+ return true
+ }
+ name := selector.Sel.Name
+ if (strings.HasPrefix(name, "Get") && name != "GetContainersForPod") || name == "IsConnected" || name == "SnapshotCaches" || name == "ClientFromContext" || name == "DynamicClientFromContext" || name == "ConfigFromContext" {
+ t.Errorf("%s: service resolves global cluster state through %s; use caller-scoped dependencies", fset.Position(call.Pos()), name)
+ }
+ return true
+ })
+ }
+}
diff --git a/internal/cnpg/catalog.go b/internal/cnpg/catalog.go
new file mode 100644
index 0000000000..e26549fe43
--- /dev/null
+++ b/internal/cnpg/catalog.go
@@ -0,0 +1,115 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "log"
+ "net/http"
+ "sort"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// CNPGCatalogUser is one Cluster pinned to an image catalog.
+type CNPGCatalogUser struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ // Major is the PostgreSQL major the Cluster asks the catalog for. A
+ // reference without one lands here as 0, which the screen must not describe
+ // as "asks for PostgreSQL 0".
+ Major int `json:"major,omitempty"`
+ // Image the Cluster resolved and is running now. A cluster pinned to a
+ // catalog carries no spec.imageName, so this is the only place the running
+ // image appears.
+ Image string `json:"image,omitempty"`
+}
+
+// CNPGCatalogUsersResponse lists the Clusters referencing one image catalog.
+type CNPGCatalogUsersResponse struct {
+ Clusters []CNPGCatalogUser `json:"clusters"`
+}
+
+// catalogRefMatches reports whether a Cluster's imageCatalogRef names this
+// catalog.
+//
+// CloudNativePG defaults an omitted `kind` to the namespaced ImageCatalog, so a
+// reference without one must not be counted against a ClusterImageCatalog of the
+// same name — the two are different objects and may both exist.
+func catalogRefMatches(ref map[string]interface{}, name, wantKind string) bool {
+ if refName, _ := ref["name"].(string); refName != name {
+ return false
+ }
+ refKind, _ := ref["kind"].(string)
+ if refKind == "" {
+ refKind = "ImageCatalog"
+ }
+ return refKind == wantKind
+}
+
+// catalogRefMajor reads the PostgreSQL major from a catalog reference.
+//
+// Read tolerantly: the same field arrives as int64 or float64 depending on how
+// the object entered the dynamic cache, and NestedInt64 alone misses the float64
+// shape — which would drop a real major to zero, a state the screen treats as
+// "the reference carries no major at all".
+func catalogRefMajor(ref map[string]interface{}) int {
+ switch v := ref["major"].(type) {
+ case int64:
+ return int(v)
+ case float64:
+ return int(v)
+ case int:
+ return v
+ }
+ return 0
+}
+
+func (s *Reader) CatalogUsers(ctx context.Context, cache *k8s.ResourceCache, namespace, name, wantKind string) (*CNPGCatalogUsersResponse, error) {
+ // A namespaced catalog can only be referenced from its own namespace; a
+ // cluster-scoped one from anywhere.
+ items, err := s.Observations.DynamicList(ctx, cache, "Cluster", Group, namespace)
+ switch {
+ case err == nil:
+ case errors.Is(err, k8s.ErrUnknownDynamicKind):
+ // No CloudNativePG on this cluster, so nothing can be pinned to a catalog.
+ return &CNPGCatalogUsersResponse{Clusters: []CNPGCatalogUser{}}, nil
+ case errors.Is(err, integration.ErrDynamicNotSynced):
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "clusters are still loading"}
+ default:
+ // "No cluster uses this catalog" is the sentence someone reads before
+ // editing it. Never say it because the lookup failed.
+ log.Printf("[cnpg] Failed to list Clusters for catalog %s/%s: %v", k8s.SanitizeForLog(namespace), k8s.SanitizeForLog(name), err)
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "could not read CloudNativePG clusters"}
+ }
+
+ resp := CNPGCatalogUsersResponse{Clusters: []CNPGCatalogUser{}}
+ for _, u := range items {
+ if u == nil {
+ continue
+ }
+ ref, found, _ := unstructured.NestedMap(u.Object, "spec", "imageCatalogRef")
+ if !found {
+ continue
+ }
+ if !catalogRefMatches(ref, name, wantKind) {
+ continue
+ }
+ user := CNPGCatalogUser{Namespace: u.GetNamespace(), Name: u.GetName()}
+ user.Major = catalogRefMajor(ref)
+ if img, _, _ := unstructured.NestedString(u.Object, "status", "image"); img != "" {
+ user.Image = img
+ }
+ resp.Clusters = append(resp.Clusters, user)
+ }
+ sort.Slice(resp.Clusters, func(i, j int) bool {
+ if resp.Clusters[i].Namespace != resp.Clusters[j].Namespace {
+ return resp.Clusters[i].Namespace < resp.Clusters[j].Namespace
+ }
+ return resp.Clusters[i].Name < resp.Clusters[j].Name
+ })
+
+ return &resp, nil
+}
diff --git a/internal/cnpg/cluster_ha.go b/internal/cnpg/cluster_ha.go
new file mode 100644
index 0000000000..e32d1eee18
--- /dev/null
+++ b/internal/cnpg/cluster_ha.go
@@ -0,0 +1,775 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "log"
+ "slices"
+ "sort"
+ "strings"
+ "time"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ discoveryv1 "k8s.io/api/discovery/v1"
+ policyv1 "k8s.io/api/policy/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/client-go/dynamic"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/metadata"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+const (
+ cnpgHAStateOK = "ok"
+ cnpgHAStateDenied = "denied"
+ cnpgHAStateNotFound = "notFound"
+ cnpgHAStateNotInstalled = "notInstalled"
+ cnpgHAStateUnavailable = "unavailable"
+ cnpgHAStateError = "error"
+
+ jobRoleLabel = "cnpg.io/jobRole"
+ cnpgZoneLabel = "topology.kubernetes.io/zone"
+ cnpgOperatorLeaseName = "db9c8771.cnpg.io"
+ cnpgFailoverQuorumAnno = "alpha.cnpg.io/failoverQuorum"
+ cnpgHAReadTimeout = 5 * time.Second
+
+ certManagerCertificateAnno = "cert-manager.io/certificate-name"
+ certManagerIssuerAnno = "cert-manager.io/issuer-name"
+ certManagerIssuerKindAnno = "cert-manager.io/issuer-kind"
+)
+
+var (
+ cnpgFailoverQuorumGVR = schema.GroupVersionResource{Group: Group, Version: "v1", Resource: "failoverquorums"}
+ cnpgSecretsGVR = schema.GroupVersionResource{Version: "v1", Resource: "secrets"}
+)
+
+type CNPGClusterHAResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ SampledAt string `json:"sampledAt"`
+ DesiredImage string `json:"desiredImage,omitempty"`
+ DeclaredInstances *int64 `json:"declaredInstances,omitempty"`
+ ExpectedInstances []string `json:"expectedInstances"`
+ Instances []CNPGHAInstance `json:"instances"`
+ Pods integration.ReadSource `json:"pods"`
+ Nodes integration.ReadSource `json:"nodes"`
+ Quorum CNPGHAQuorum `json:"quorum"`
+ PDBs CNPGHAPDBs `json:"pdbs"`
+ PrimaryLease CNPGHALease `json:"primaryLease"`
+ OperatorLease CNPGHALease `json:"operatorLease"`
+ Jobs CNPGHAJobs `json:"jobs"`
+ RWEndpoints CNPGHAEndpoints `json:"rwEndpoints"`
+ Certificates []CNPGHACertificate `json:"certificates"`
+ Maintenance CNPGMaintenanceFacts `json:"maintenance"`
+}
+
+// CNPGHAInstance is one instance Pod. Zone is empty when the Node is not
+// readable or carries no zone label (Nodes says which). PostgresStartedAt is
+// the postgres container's current start: a Pod restart moves it, an in-place
+// PostgreSQL restart does not.
+type CNPGHAInstance struct {
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+ Role string `json:"role"`
+ Ready bool `json:"ready"`
+ Node string `json:"node,omitempty"`
+ Zone string `json:"zone,omitempty"`
+ QOSClass string `json:"qosClass,omitempty"`
+ Image string `json:"image,omitempty"`
+ ImageMatches *bool `json:"imageMatches,omitempty"`
+ PodCreatedAt string `json:"podCreatedAt,omitempty"`
+ PostgresStartedAt string `json:"postgresStartedAt,omitempty"`
+ RestartCount int32 `json:"restartCount"`
+}
+
+// CNPGHAQuorum is the recorded synchronous-replication configuration, not a
+// verdict. N is the potentially synchronous standbys, W the standbys a commit
+// waits for, R the promotable ones (in the cluster and ready now). Holds is
+// R + W > N, and is absent when any term is unknown.
+type CNPGHAQuorum struct {
+ Enabled bool `json:"enabled"`
+ EnabledBy string `json:"enabledBy,omitempty"`
+ Method string `json:"method,omitempty"`
+ Number *int64 `json:"number,omitempty"`
+ DataDurability string `json:"dataDurability,omitempty"`
+ Object integration.ReadSource `json:"object"`
+ Status *CNPGHAQuorumStatus `json:"status,omitempty"`
+ N *int `json:"n,omitempty"`
+ W *int `json:"w,omitempty"`
+ R *int `json:"r,omitempty"`
+ Promotable []string `json:"promotable,omitempty"`
+ Holds *bool `json:"holds,omitempty"`
+}
+
+type CNPGHAQuorumStatus struct {
+ Method string `json:"method,omitempty"`
+ StandbyNames []string `json:"standbyNames"`
+ StandbyNumber int `json:"standbyNumber"`
+ Primary string `json:"primary,omitempty"`
+}
+
+// CNPGHAPDBs are the budgets the operator owns for this Cluster. Enabled is
+// spec.enablePDB (default true); an absent budget under enablePDB false is
+// the declared state, not a fault.
+type CNPGHAPDBs struct {
+ integration.ReadSource
+ Enabled bool `json:"enabled"`
+ Items []CNPGHAPDB `json:"items"`
+}
+
+type CNPGHAPDB struct {
+ Name string `json:"name"`
+ Role string `json:"role"`
+ MinAvailable string `json:"minAvailable,omitempty"`
+ MaxUnavailable string `json:"maxUnavailable,omitempty"`
+ ExpectedPods int32 `json:"expectedPods"`
+ CurrentHealthy int32 `json:"currentHealthy"`
+ DesiredHealthy int32 `json:"desiredHealthy"`
+ DisruptionsAllowed int32 `json:"disruptionsAllowed"`
+ // Observed is false when the disruption controller has not caught up with
+ // the budget's current generation; its numbers are then stale.
+ Observed bool `json:"observed"`
+}
+
+// CNPGHALease is one Lease. Expired compares renewTime + duration with now.
+type CNPGHALease struct {
+ integration.ReadSource
+ Namespace string `json:"namespace,omitempty"`
+ Name string `json:"name,omitempty"`
+ Holder string `json:"holder,omitempty"`
+ RenewTime string `json:"renewTime,omitempty"`
+ DurationSeconds *int32 `json:"durationSeconds,omitempty"`
+ Expired *bool `json:"expired,omitempty"`
+ ControlledByCluster *bool `json:"controlledByCluster,omitempty"`
+}
+
+type CNPGHAJobs struct {
+ integration.ReadSource
+ Items []CNPGHAJob `json:"items"`
+}
+
+// CNPGHAJob Phase: active | succeeded | failed | pending.
+type CNPGHAJob struct {
+ Name string `json:"name"`
+ Role string `json:"role,omitempty"`
+ Instance string `json:"instance,omitempty"`
+ Phase string `json:"phase"`
+ Reason string `json:"reason,omitempty"`
+ StartTime string `json:"startTime,omitempty"`
+ CompletionTime string `json:"completionTime,omitempty"`
+}
+
+// CNPGHAEndpoints are the ready endpoints behind the -rw Service, named by Pod.
+type CNPGHAEndpoints struct {
+ integration.ReadSource
+ Service string `json:"service"`
+ Pods []string `json:"pods"`
+}
+
+// CNPGHACertificate is one entry of status.certificates.expirations. ExpiresAt
+// is RFC3339 when the operator's value parsed; Raw is always what it wrote.
+// Renewal: operator (generated by CloudNativePG) | user (named in
+// spec.certificates). CertManager is set only when the Secret's metadata was
+// read and carries cert-manager's annotations; Metadata says whether it was.
+type CNPGHACertificate struct {
+ Secret string `json:"secret"`
+ Purposes []string `json:"purposes,omitempty"`
+ Raw string `json:"raw"`
+ ExpiresAt string `json:"expiresAt,omitempty"`
+ Renewal string `json:"renewal"`
+ Metadata *integration.ReadSource `json:"metadata,omitempty"`
+ CertManager *CNPGCertManagerRef `json:"certManager,omitempty"`
+}
+
+type CNPGCertManagerRef struct {
+ Certificate string `json:"certificate"`
+ Issuer string `json:"issuer,omitempty"`
+ IssuerKind string `json:"issuerKind,omitempty"`
+}
+
+type HAClients struct {
+ Typed kubernetes.Interface
+ Dynamic dynamic.Interface
+ Metadata metadata.Interface
+}
+
+func (s *Reader) ClusterHA(callerCtx context.Context, c HAClients, cache *k8s.ResourceCache, cluster *unstructured.Unstructured, now time.Time) CNPGClusterHAResponse {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ ctx, cancel := context.WithTimeout(callerCtx, cnpgHAReadTimeout)
+ defer cancel()
+
+ resp := CNPGClusterHAResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ SampledAt: now.Format(time.RFC3339),
+ DesiredImage: cnpgDesiredImage(cluster),
+ Instances: []CNPGHAInstance{},
+ Certificates: []CNPGHACertificate{},
+ Maintenance: cnpgMaintenanceFactsOf(cluster),
+ }
+
+ if n, found, _ := unstructured.NestedInt64(cluster.Object, "spec", "instances"); found {
+ resp.DeclaredInstances = &n
+ }
+ resp.ExpectedInstances = append([]string{}, cnpgStatusStrings(cluster, "instanceNames")...)
+
+ var pods []*corev1.Pod
+ if s.Access.CanRead(callerCtx, "", "pods", namespace, "list") {
+ var err error
+ pods, err = clusterInstancePods(cache, cluster)
+ if err != nil {
+ resp.Pods = integration.ReadSource{State: cnpgHAStateUnavailable, Reason: err.Error()}
+ } else {
+ resp.Pods = integration.ReadSource{State: cnpgHAStateOK}
+ }
+ } else {
+ resp.Pods = cnpgHADenied(auth.Grant{Verb: "list", Resource: "pods"}, namespace)
+ }
+ resp.Nodes = s.hANodesSource(callerCtx, cache)
+ resp.Instances = cnpgHAInstances(cache, pods, resp.DesiredImage, resp.Nodes.State == cnpgHAStateOK)
+
+ resp.Quorum = s.hAQuorum(ctx, callerCtx, c, cluster, pods, resp.Pods.State == cnpgHAStateOK)
+ resp.PDBs = s.hAPDBs(callerCtx, cache, cluster)
+ resp.PrimaryLease = s.hAPrimaryLease(ctx, callerCtx, c, cluster, now)
+ resp.OperatorLease = s.hAOperatorLease(ctx, callerCtx, c, cache, now)
+ resp.Jobs = s.hAJobs(callerCtx, cache, cluster)
+ for _, job := range resp.Jobs.Items {
+ if job.Instance != "" && (job.Role == "join" || job.Role == "initdb") && (job.Phase == "active" || job.Phase == "pending") {
+ resp.ExpectedInstances = append(resp.ExpectedInstances, job.Instance)
+ }
+ }
+ slices.Sort(resp.ExpectedInstances)
+ resp.ExpectedInstances = slices.Compact(resp.ExpectedInstances)
+ resp.RWEndpoints = s.hARWEndpoints(ctx, callerCtx, c, cluster)
+ resp.Certificates = s.hACertificates(ctx, callerCtx, c, cluster)
+ return resp
+}
+
+func cnpgHADenied(g auth.Grant, namespace string) integration.ReadSource {
+ g = g.In(namespace)
+ return integration.ReadSource{State: cnpgHAStateDenied, Grant: g.Ref(), Reason: "You are not allowed to " + g.String()}
+}
+
+func cnpgHAClusterDenied(g auth.Grant) integration.ReadSource {
+ return integration.ReadSource{State: cnpgHAStateDenied, Grant: g.Ref(), Reason: "You are not allowed to " + g.String()}
+}
+
+// cnpgHAReadError classifies an impersonated read's failure. NotFound is left
+// to the caller: its meaning differs per object.
+func cnpgHAReadError(err error, g auth.Grant, namespace string) integration.ReadSource {
+ switch {
+ case apierrors.IsForbidden(err):
+ return cnpgHADenied(g, namespace)
+ case apierrors.IsNotFound(err):
+ return integration.ReadSource{State: cnpgHAStateNotFound}
+ case errors.Is(err, context.DeadlineExceeded) || apierrors.IsTimeout(err):
+ return integration.ReadSource{State: cnpgHAStateUnavailable, Reason: "no answer within " + cnpgHAReadTimeout.String()}
+ default:
+ if plain, ok := cnpgTransportSentence(err, 0, cnpgHAReadTimeout); ok {
+ log.Printf("[cnpg] Failed to read %s: %v", g.In(namespace).String(), err)
+ return integration.ReadSource{State: cnpgHAStateUnavailable, Reason: plain}
+ }
+ return integration.ReadSource{State: cnpgHAStateError, Reason: truncateCNPGRuntimeError(err.Error())}
+ }
+}
+
+func cnpgDesiredImage(cluster *unstructured.Unstructured) string {
+ if img, _, _ := unstructured.NestedString(cluster.Object, "status", "image"); img != "" {
+ return img
+ }
+ img, _, _ := unstructured.NestedString(cluster.Object, "spec", "imageName")
+ return img
+}
+
+// cnpgUncachedReason says what the reader cannot see when Radar's cache holds
+// no copy of a kind the caller may read: Radar lists with its own credentials
+// at connect time, which can be narrower than the caller's. Empty when the
+// kind is cached and synced. uncached: Radar holds no informer for the kind.
+// outOfScope: it watches the kind only in other namespaces, either because its
+// credentials could list it only there or because the namespace was beyond
+// the set it probed.
+func cnpgUncachedReason(fact, kind, namespace string, uncached, outOfScope, ready bool) string {
+ switch {
+ case uncached:
+ return fmt.Sprintf("%s unknown: Radar's own credentials could not list %s when it connected", fact, kind)
+ case outOfScope:
+ return fmt.Sprintf("%s unknown: Radar watches %s only in the namespaces it chose when it connected, and %s is not one of them", fact, kind, namespace)
+ case !ready:
+ return fmt.Sprintf("%s unknown: Radar is still loading %s", fact, kind)
+ }
+ return ""
+}
+
+func (s *Reader) hANodesSource(ctx context.Context, cache *k8s.ResourceCache) integration.ReadSource {
+ if !s.Access.CanRead(ctx, "", "nodes", "", "get") {
+ return cnpgHAClusterDenied(auth.Grant{Verb: "get", Resource: "nodes"})
+ }
+ if reason := cnpgUncachedReason("Zones", "Nodes", "", cache.Nodes() == nil, false, cache.IsKindReady("nodes")); reason != "" {
+ return integration.ReadSource{State: cnpgHAStateUnavailable, Reason: reason}
+ }
+ return integration.ReadSource{State: cnpgHAStateOK}
+}
+
+func cnpgHAInstances(cache *k8s.ResourceCache, pods []*corev1.Pod, desiredImage string, nodesReadable bool) []CNPGHAInstance {
+ out := make([]CNPGHAInstance, 0, len(pods))
+ for _, p := range pods {
+ inst := CNPGHAInstance{
+ Pod: p.Name,
+ PodUID: string(p.UID),
+ Role: runtimeRole(p),
+ Ready: cnpgActionPodReady(p),
+ Node: p.Spec.NodeName,
+ QOSClass: string(p.Status.QOSClass),
+ PodCreatedAt: p.CreationTimestamp.UTC().Format(time.RFC3339),
+ }
+ for _, c := range p.Spec.Containers {
+ if c.Name == defaultLogContainer {
+ inst.Image = c.Image
+ }
+ }
+ for _, cs := range p.Status.ContainerStatuses {
+ if cs.Name != defaultLogContainer {
+ continue
+ }
+ inst.RestartCount = cs.RestartCount
+ if cs.State.Running != nil {
+ inst.PostgresStartedAt = cs.State.Running.StartedAt.UTC().Format(time.RFC3339)
+ }
+ }
+ if desiredImage != "" && inst.Image != "" {
+ match := inst.Image == desiredImage
+ inst.ImageMatches = &match
+ }
+ if nodesReadable && inst.Node != "" {
+ if n, err := cache.Nodes().Get(inst.Node); err == nil {
+ inst.Zone = n.Labels[cnpgZoneLabel]
+ }
+ }
+ out = append(out, inst)
+ }
+ return out
+}
+
+// cnpgHAQuorum reads the recorded synchronous configuration. The object exists
+// only while quorum failover is enabled; an enabled cluster without one has a
+// configuration the primary has not recorded yet, and failover waits.
+func (s *Reader) hAQuorum(ctx context.Context, callerCtx context.Context, c HAClients, cluster *unstructured.Unstructured, pods []*corev1.Pod, podsKnown bool) CNPGHAQuorum {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ q := CNPGHAQuorum{}
+ sync, hasSync, _ := unstructured.NestedMap(cluster.Object, "spec", "postgresql", "synchronous")
+ if hasSync {
+ q.Method, _ = sync["method"].(string)
+ q.DataDurability, _ = sync["dataDurability"].(string)
+ if n, ok, _ := unstructured.NestedInt64(sync, "number"); ok {
+ q.Number = &n
+ }
+ if v, ok := sync["failoverQuorum"].(bool); ok && v {
+ q.Enabled, q.EnabledBy = true, "spec"
+ }
+ }
+ switch strings.ToLower(cluster.GetAnnotations()[cnpgFailoverQuorumAnno]) {
+ case "true":
+ if hasSync {
+ q.Enabled, q.EnabledBy = true, "annotation"
+ }
+ case "false":
+ q.Enabled, q.EnabledBy = false, "annotation"
+ }
+
+ if disc := s.Observations.Discovery; disc != nil {
+ if _, ok := disc.GetGVRWithGroup("FailoverQuorum", Group); !ok {
+ q.Object = integration.ReadSource{State: cnpgHAStateNotInstalled, Reason: "this CloudNativePG version has no FailoverQuorum resource"}
+ return q
+ }
+ }
+ grant := auth.Grant{Verb: "get", Group: Group, Resource: "failoverquorums"}
+ if !s.Access.CanRead(callerCtx, Group, "failoverquorums", namespace, "get") {
+ q.Object = cnpgHADenied(grant, namespace)
+ return q
+ }
+ obj, err := c.Dynamic.Resource(cnpgFailoverQuorumGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ q.Object = cnpgHAReadError(err, grant, namespace)
+ return q
+ }
+ if !controlledBy(obj.GetOwnerReferences(), Group, "Cluster", name, cluster.GetUID()) {
+ q.Object = integration.ReadSource{State: cnpgHAStateError, Reason: "a FailoverQuorum of this name exists but this Cluster does not own it"}
+ return q
+ }
+ q.Object = integration.ReadSource{State: cnpgHAStateOK}
+ st := &CNPGHAQuorumStatus{StandbyNames: []string{}}
+ st.Method, _, _ = unstructured.NestedString(obj.Object, "status", "method")
+ st.Primary, _, _ = unstructured.NestedString(obj.Object, "status", "primary")
+ if names, ok, _ := unstructured.NestedStringSlice(obj.Object, "status", "standbyNames"); ok {
+ st.StandbyNames = names
+ }
+ if n, ok, _ := unstructured.NestedInt64(obj.Object, "status", "standbyNumber"); ok {
+ st.StandbyNumber = int(n)
+ }
+ q.Status = st
+ cnpgQuorumArithmetic(&q, pods, podsKnown)
+ return q
+}
+
+// cnpgQuorumArithmetic fills N, W, R and whether R + W > N. A reset object
+// (no standby names) records no configuration: nothing is computed. R counts
+// the standby names that are instances of this Cluster with a ready Pod, other
+// than the recorded primary — Pod readiness is Radar's view of "able to report
+// its state", which the operator checks directly.
+func cnpgQuorumArithmetic(q *CNPGHAQuorum, pods []*corev1.Pod, podsKnown bool) {
+ st := q.Status
+ if st == nil || len(st.StandbyNames) == 0 {
+ return
+ }
+ n, w := len(st.StandbyNames), st.StandbyNumber
+ q.N, q.W = &n, &w
+ if !podsKnown {
+ return
+ }
+ ready := map[string]bool{}
+ for _, p := range pods {
+ if cnpgActionPodReady(p) {
+ ready[p.Name] = true
+ }
+ }
+ promotable := []string{}
+ for _, s := range st.StandbyNames {
+ if s != st.Primary && ready[s] {
+ promotable = append(promotable, s)
+ }
+ }
+ rr := len(promotable)
+ holds := rr+w > n
+ q.R, q.Promotable, q.Holds = &rr, promotable, &holds
+}
+
+func (s *Reader) hAPDBs(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured) CNPGHAPDBs {
+ namespace := cluster.GetNamespace()
+ out := CNPGHAPDBs{Enabled: true, Items: []CNPGHAPDB{}}
+ if v, ok, _ := unstructured.NestedBool(cluster.Object, "spec", "enablePDB"); ok {
+ out.Enabled = v
+ }
+ grant := auth.Grant{Verb: "list", Group: "policy", Resource: "poddisruptionbudgets"}
+ if !s.Access.CanRead(ctx, "policy", "poddisruptionbudgets", namespace, "list") {
+ out.ReadSource = cnpgHADenied(grant, namespace)
+ return out
+ }
+ lister := cache.PodDisruptionBudgets()
+ within := integration.NamespacesWithinCache(cache, "poddisruptionbudgets", []string{namespace})
+ if reason := cnpgUncachedReason("Disruption budgets", "PodDisruptionBudgets", namespace, lister == nil, within.Unavailable, cache.IsKindReady("poddisruptionbudgets")); reason != "" {
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateUnavailable, Reason: reason}
+ return out
+ }
+ list, err := lister.PodDisruptionBudgets(namespace).List(labels.SelectorFromSet(labels.Set{clusterLabel: cluster.GetName()}))
+ if err != nil {
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateError, Reason: err.Error()}
+ return out
+ }
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateOK}
+ for _, p := range list {
+ if !controlledBy(p.OwnerReferences, Group, "Cluster", cluster.GetName(), cluster.GetUID()) {
+ continue
+ }
+ out.Items = append(out.Items, cnpgHAPDBOf(p, cluster.GetName()))
+ }
+ sort.Slice(out.Items, func(i, j int) bool { return out.Items[i].Name < out.Items[j].Name })
+ return out
+}
+
+func cnpgHAPDBOf(p *policyv1.PodDisruptionBudget, cluster string) CNPGHAPDB {
+ item := CNPGHAPDB{
+ Name: p.Name,
+ Role: "other",
+ ExpectedPods: p.Status.ExpectedPods,
+ CurrentHealthy: p.Status.CurrentHealthy,
+ DesiredHealthy: p.Status.DesiredHealthy,
+ DisruptionsAllowed: p.Status.DisruptionsAllowed,
+ Observed: p.Status.ObservedGeneration >= p.Generation,
+ }
+ switch p.Name {
+ case cluster + "-primary":
+ item.Role = "primary"
+ case cluster:
+ item.Role = "replicas"
+ }
+ if p.Spec.MinAvailable != nil {
+ item.MinAvailable = p.Spec.MinAvailable.String()
+ }
+ if p.Spec.MaxUnavailable != nil {
+ item.MaxUnavailable = p.Spec.MaxUnavailable.String()
+ }
+ return item
+}
+
+func cnpgHALeaseFrom(l *leaseView, now time.Time) CNPGHALease {
+ out := CNPGHALease{ReadSource: integration.ReadSource{State: cnpgHAStateOK}, Namespace: l.namespace, Name: l.name, Holder: l.holder, DurationSeconds: l.duration}
+ if l.renew != nil {
+ out.RenewTime = l.renew.UTC().Format(time.RFC3339)
+ if l.duration != nil {
+ expired := now.After(l.renew.Add(time.Duration(*l.duration) * time.Second))
+ out.Expired = &expired
+ }
+ }
+ return out
+}
+
+type leaseView struct {
+ namespace, name, holder string
+ renew *time.Time
+ duration *int32
+}
+
+func (s *Reader) hAPrimaryLease(ctx context.Context, callerCtx context.Context, c HAClients, cluster *unstructured.Unstructured, now time.Time) CNPGHALease {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ grant := auth.Grant{Verb: "get", Group: "coordination.k8s.io", Resource: "leases"}
+ if !s.Access.CanRead(callerCtx, "coordination.k8s.io", "leases", namespace, "get") {
+ return CNPGHALease{ReadSource: cnpgHADenied(grant, namespace), Namespace: namespace, Name: name}
+ }
+ l, err := c.Typed.CoordinationV1().Leases(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ src := cnpgHAReadError(err, grant, namespace)
+ if src.State == cnpgHAStateNotFound {
+ src.Reason = "No primary Lease: CloudNativePG creates one from 1.30"
+ }
+ return CNPGHALease{ReadSource: src, Namespace: namespace, Name: name}
+ }
+ v := &leaseView{namespace: namespace, name: name, duration: l.Spec.LeaseDurationSeconds}
+ if l.Spec.HolderIdentity != nil {
+ v.holder = *l.Spec.HolderIdentity
+ }
+ if l.Spec.RenewTime != nil {
+ t := l.Spec.RenewTime.Time
+ v.renew = &t
+ }
+ out := cnpgHALeaseFrom(v, now)
+ controlled := controlledBy(l.OwnerReferences, Group, "Cluster", name, cluster.GetUID())
+ out.ControlledByCluster = &controlled
+ return out
+}
+
+// cnpgHAOperatorLease finds the operator's leader-election Lease in the
+// namespace of a visible operator Deployment. Which operator replica leads is
+// a different fact from which instance is primary.
+func (s *Reader) hAOperatorLease(ctx context.Context, callerCtx context.Context, c HAClients, cache *k8s.ResourceCache, now time.Time) CNPGHALease {
+ acc, deployments := s.operatorDeployments(callerCtx, cache, s.Observations.OperatorScope(callerCtx))
+ var namespace string
+ for _, d := range deployments {
+ if d.Labels[cnpgOperatorNameLabel] == cnpgOperatorNameValue {
+ namespace = d.Namespace
+ break
+ }
+ }
+ if namespace == "" {
+ reason := "The operator Deployment is not visible to you, so its namespace is unknown"
+ switch acc.State {
+ case integration.KindCoverageFull:
+ reason = "No operator Deployment labelled " + cnpgOperatorNameLabel + "=" + cnpgOperatorNameValue + " was found"
+ case integration.KindCoverageUncached:
+ reason = "Radar does not watch Deployments in the namespaces you can read, so the operator's namespace is unknown"
+ }
+ return CNPGHALease{ReadSource: integration.ReadSource{State: cnpgHAStateUnavailable, Reason: reason}, Name: cnpgOperatorLeaseName}
+ }
+ grant := auth.Grant{Verb: "get", Group: "coordination.k8s.io", Resource: "leases"}
+ if !s.Access.CanRead(callerCtx, "coordination.k8s.io", "leases", namespace, "get") {
+ return CNPGHALease{ReadSource: cnpgHADenied(grant, namespace), Namespace: namespace, Name: cnpgOperatorLeaseName}
+ }
+ l, err := c.Typed.CoordinationV1().Leases(namespace).Get(ctx, cnpgOperatorLeaseName, metav1.GetOptions{})
+ if err != nil {
+ src := cnpgHAReadError(err, grant, namespace)
+ if src.State == cnpgHAStateNotFound {
+ src.Reason = "No leader-election Lease: the operator may run with leader election off"
+ }
+ return CNPGHALease{ReadSource: src, Namespace: namespace, Name: cnpgOperatorLeaseName}
+ }
+ v := &leaseView{namespace: namespace, name: cnpgOperatorLeaseName, duration: l.Spec.LeaseDurationSeconds}
+ if l.Spec.HolderIdentity != nil {
+ v.holder = *l.Spec.HolderIdentity
+ }
+ if l.Spec.RenewTime != nil {
+ t := l.Spec.RenewTime.Time
+ v.renew = &t
+ }
+ return cnpgHALeaseFrom(v, now)
+}
+
+func (s *Reader) hAJobs(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured) CNPGHAJobs {
+ namespace := cluster.GetNamespace()
+ out := CNPGHAJobs{Items: []CNPGHAJob{}}
+ grant := auth.Grant{Verb: "list", Group: "batch", Resource: "jobs"}
+ if !s.Access.CanRead(ctx, "batch", "jobs", namespace, "list") {
+ out.ReadSource = cnpgHADenied(grant, namespace)
+ return out
+ }
+ lister := cache.Jobs()
+ within := integration.NamespacesWithinCache(cache, "jobs", []string{namespace})
+ if reason := cnpgUncachedReason("Instance Jobs", "Jobs", namespace, lister == nil, within.Unavailable, cache.IsKindReady("jobs")); reason != "" {
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateUnavailable, Reason: reason}
+ return out
+ }
+ list, err := lister.Jobs(namespace).List(labels.SelectorFromSet(labels.Set{clusterLabel: cluster.GetName()}))
+ if err != nil {
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateError, Reason: err.Error()}
+ return out
+ }
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateOK}
+ var jobPods []*corev1.Pod
+ if s.Access.CanRead(ctx, "", "pods", namespace, "list") && cache.Pods() != nil && cache.IsKindReady("pods") && !integration.NamespacesWithinCache(cache, "pods", []string{namespace}).Unavailable {
+ jobPods, _ = cache.Pods().Pods(namespace).List(labels.Everything())
+ }
+ for _, j := range list {
+ if !controlledBy(j.OwnerReferences, Group, "Cluster", cluster.GetName(), cluster.GetUID()) {
+ continue
+ }
+ out.Items = append(out.Items, cnpgHAJobOf(j, jobPods...))
+ }
+ sort.Slice(out.Items, func(i, j int) bool {
+ if out.Items[i].StartTime != out.Items[j].StartTime {
+ return out.Items[i].StartTime > out.Items[j].StartTime
+ }
+ return out.Items[i].Name < out.Items[j].Name
+ })
+ return out
+}
+
+func cnpgHAJobOf(j *batchv1.Job, pods ...*corev1.Pod) CNPGHAJob {
+ item := CNPGHAJob{Name: j.Name, Role: j.Labels[jobRoleLabel], Instance: j.Labels["cnpg.io/instanceName"], Phase: "pending"}
+ if j.Status.StartTime != nil {
+ item.StartTime = j.Status.StartTime.UTC().Format(time.RFC3339)
+ }
+ if j.Status.CompletionTime != nil {
+ item.CompletionTime = j.Status.CompletionTime.UTC().Format(time.RFC3339)
+ }
+ for _, c := range j.Status.Conditions {
+ if c.Status != corev1.ConditionTrue {
+ continue
+ }
+ switch c.Type {
+ case batchv1.JobComplete:
+ item.Phase = "succeeded"
+ return item
+ case batchv1.JobFailed:
+ item.Phase, item.Reason = "failed", strings.TrimSpace(c.Reason+": "+c.Message)
+ item.Reason = strings.TrimPrefix(strings.TrimSuffix(item.Reason, ":"), ": ")
+ return item
+ }
+ }
+ if j.Status.Active > 0 {
+ item.Phase = "active"
+ var schedulerReason string
+ for _, pod := range pods {
+ if pod.DeletionTimestamp != nil || !controlledBy(pod.OwnerReferences, "batch", "Job", j.Name, j.UID) {
+ continue
+ }
+ if pod.Status.Phase == corev1.PodRunning {
+ return item
+ }
+ if pod.Status.Phase != corev1.PodPending {
+ continue
+ }
+ for _, condition := range pod.Status.Conditions {
+ if condition.Type == corev1.PodScheduled && condition.Status == corev1.ConditionFalse && condition.Reason == corev1.PodReasonUnschedulable {
+ schedulerReason = "Pod cannot be scheduled: " + condition.Reason + ": " + condition.Message
+ }
+ }
+ }
+ if schedulerReason != "" {
+ item.Phase, item.Reason = "pending", schedulerReason
+ }
+ }
+ return item
+}
+
+func (s *Reader) hARWEndpoints(ctx context.Context, callerCtx context.Context, c HAClients, cluster *unstructured.Unstructured) CNPGHAEndpoints {
+ namespace := cluster.GetNamespace()
+ svc := cluster.GetName() + "-rw"
+ out := CNPGHAEndpoints{Service: svc, Pods: []string{}}
+ grant := auth.Grant{Verb: "list", Group: "discovery.k8s.io", Resource: "endpointslices"}
+ if !s.Access.CanRead(callerCtx, "discovery.k8s.io", "endpointslices", namespace, "list") {
+ out.ReadSource = cnpgHADenied(grant, namespace)
+ return out
+ }
+ list, err := c.Typed.DiscoveryV1().EndpointSlices(namespace).List(ctx, metav1.ListOptions{LabelSelector: discoveryv1.LabelServiceName + "=" + svc})
+ if err != nil {
+ out.ReadSource = cnpgHAReadError(err, grant, namespace)
+ return out
+ }
+ out.ReadSource = integration.ReadSource{State: cnpgHAStateOK}
+ out.Pods = cnpgReadyEndpointPods(list.Items)
+ return out
+}
+
+func cnpgReadyEndpointPods(slices []discoveryv1.EndpointSlice) []string {
+ seen := map[string]bool{}
+ pods := []string{}
+ for _, sl := range slices {
+ for _, ep := range sl.Endpoints {
+ if ep.Conditions.Ready != nil && !*ep.Conditions.Ready {
+ continue
+ }
+ if ep.TargetRef == nil || ep.TargetRef.Kind != "Pod" || seen[ep.TargetRef.Name] {
+ continue
+ }
+ seen[ep.TargetRef.Name] = true
+ pods = append(pods, ep.TargetRef.Name)
+ }
+ }
+ sort.Strings(pods)
+ return pods
+}
+
+// cnpgHACertificates lists every certificate the operator reports an expiry
+// for, with who renews it. A user-provided Secret's metadata is read (never
+// its data) to name a cert-manager Certificate; without get secrets that link
+// is unknown, not absent.
+func (s *Reader) hACertificates(ctx context.Context, callerCtx context.Context, c HAClients, cluster *unstructured.Unstructured) []CNPGHACertificate {
+ namespace := cluster.GetNamespace()
+ exp, _, _ := unstructured.NestedStringMap(cluster.Object, "status", "certificates", "expirations")
+ purposes := map[string][]string{}
+ for _, f := range []string{"serverCASecret", "serverTLSSecret", "replicationTLSSecret", "clientCASecret"} {
+ if n, _, _ := unstructured.NestedString(cluster.Object, "status", "certificates", f); n != "" {
+ purposes[n] = append(purposes[n], f)
+ }
+ }
+ user := issues.CNPGUserCertificateSecrets(cluster)
+ canGetSecrets := s.Access.CanRead(callerCtx, "", "secrets", namespace, "get")
+ grant := auth.Grant{Verb: "get", Resource: "secrets"}
+
+ out := make([]CNPGHACertificate, 0, len(exp))
+ for secret, raw := range exp {
+ cert := CNPGHACertificate{Secret: secret, Raw: raw, Purposes: purposes[secret], Renewal: "operator"}
+ if t, ok := issues.ParseCNPGCertExpiry(raw); ok {
+ cert.ExpiresAt = t.UTC().Format(time.RFC3339)
+ }
+ if _, ok := user[secret]; ok {
+ cert.Renewal = "user"
+ src := integration.ReadSource{State: cnpgHAStateOK}
+ switch {
+ case !canGetSecrets:
+ src = cnpgHADenied(grant, namespace)
+ default:
+ m, err := c.Metadata.Resource(cnpgSecretsGVR).Namespace(namespace).Get(ctx, secret, metav1.GetOptions{})
+ if err != nil {
+ src = cnpgHAReadError(err, grant, namespace)
+ } else if name := m.GetAnnotations()[certManagerCertificateAnno]; name != "" {
+ cert.CertManager = &CNPGCertManagerRef{
+ Certificate: name,
+ Issuer: m.GetAnnotations()[certManagerIssuerAnno],
+ IssuerKind: m.GetAnnotations()[certManagerIssuerKindAnno],
+ }
+ }
+ }
+ cert.Metadata = &src
+ }
+ out = append(out, cert)
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Secret < out[j].Secret })
+ return out
+}
diff --git a/internal/cnpg/cluster_ha_test.go b/internal/cnpg/cluster_ha_test.go
new file mode 100644
index 0000000000..dcf5498d3b
--- /dev/null
+++ b/internal/cnpg/cluster_ha_test.go
@@ -0,0 +1,82 @@
+package cnpg
+
+import (
+ "strings"
+ "testing"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+)
+
+func TestCNPGQuorumArithmetic(t *testing.T) {
+ ready := func(name string, ok bool) *corev1.Pod {
+ st := corev1.ConditionFalse
+ if ok {
+ st = corev1.ConditionTrue
+ }
+ return &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: name}, Status: corev1.PodStatus{Conditions: []corev1.PodCondition{{Type: corev1.PodReady, Status: st}}}}
+ }
+ pods := []*corev1.Pod{ready("a", true), ready("b", true), ready("c", false)}
+
+ q := CNPGHAQuorum{Status: &CNPGHAQuorumStatus{StandbyNames: []string{"b", "c"}, StandbyNumber: 2, Primary: "a"}}
+ cnpgQuorumArithmetic(&q, pods, true)
+ if *q.R != 1 || !*q.Holds {
+ t.Errorf("R=1 W=2 N=2 should hold: %+v", q)
+ }
+
+ reset := CNPGHAQuorum{Status: &CNPGHAQuorumStatus{StandbyNames: []string{}}}
+ cnpgQuorumArithmetic(&reset, pods, true)
+ if reset.N != nil || reset.Holds != nil {
+ t.Errorf("a reset object records no configuration; nothing may be computed: %+v", reset)
+ }
+
+ unknownPods := CNPGHAQuorum{Status: &CNPGHAQuorumStatus{StandbyNames: []string{"b"}, StandbyNumber: 1}}
+ cnpgQuorumArithmetic(&unknownPods, nil, false)
+ if unknownPods.N == nil || unknownPods.R != nil || unknownPods.Holds != nil {
+ t.Errorf("without Pods R and the verdict are unknown: %+v", unknownPods)
+ }
+}
+
+func TestCNPGUncachedReasonSaysWhatIsUnknownAndWhy(t *testing.T) {
+ if got := cnpgUncachedReason("Zones", "Nodes", "", true, false, false); got != "Zones unknown: Radar's own credentials could not list Nodes when it connected" {
+ t.Errorf("uncached = %q", got)
+ }
+ if got := cnpgUncachedReason("Instance Jobs", "Jobs", "pg", false, true, true); got != "Instance Jobs unknown: Radar watches Jobs only in the namespaces it chose when it connected, and pg is not one of them" {
+ t.Errorf("out of scope = %q", got)
+ }
+ if got := cnpgUncachedReason("Instance Jobs", "Jobs", "pg", false, false, false); got != "Instance Jobs unknown: Radar is still loading Jobs" {
+ t.Errorf("syncing = %q", got)
+ }
+ if got := cnpgUncachedReason("Zones", "Nodes", "", false, false, true); got != "" {
+ t.Errorf("cached = %q", got)
+ }
+}
+
+func TestCNPGHAJobSchedulerEvidence(t *testing.T) {
+ job := cnpgJob("db", "analytics-1-initdb", "job-uid", clusterRef("analytics", "cluster-uid"))
+ job.Status.Active = 1
+ pod := cnpgJobPod("db", "analytics-1-initdb-abc", "analytics", jobRef(job.Name, string(job.UID)))
+ got := cnpgHAJobOf(job, pod)
+ if got.Phase != "pending" || !strings.Contains(got.Reason, "Pod cannot be scheduled: Unschedulable: 0/2 nodes") {
+ t.Fatalf("job = %+v", got)
+ }
+ if got := cnpgHAJobOf(job); got.Phase != "active" {
+ t.Fatalf("without Pod access: %+v", got)
+ }
+ pod.OwnerReferences[0].UID = "previous-job"
+ if got := cnpgHAJobOf(job, pod); got.Phase != "active" {
+ t.Fatalf("stale Job Pod adopted: %+v", got)
+ }
+ pod.OwnerReferences[0] = jobRef(job.Name, string(job.UID))
+ pod.Status.Phase = corev1.PodRunning
+ if got := cnpgHAJobOf(job, pod); got.Phase != "active" {
+ t.Fatalf("active Job: %+v", got)
+ }
+ pending := pod.DeepCopy()
+ pending.Status.Phase = corev1.PodPending
+ for _, pods := range [][]*corev1.Pod{{pending, pod}, {pod, pending}} {
+ if got := cnpgHAJobOf(job, pods...); got.Phase != "active" {
+ t.Fatalf("Job with a running Pod: %+v", got)
+ }
+ }
+}
diff --git a/internal/cnpg/destroy.go b/internal/cnpg/destroy.go
new file mode 100644
index 0000000000..87bc78db83
--- /dev/null
+++ b/internal/cnpg/destroy.go
@@ -0,0 +1,571 @@
+package cnpg
+
+import (
+ "context"
+ "fmt"
+ "log"
+ "net/http"
+ "slices"
+ "sort"
+ "strings"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/selection"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+const (
+ cnpgPVCStatusAnnotation = "cnpg.io/pvcStatus"
+ cnpgPVCStatusDetached = "detached"
+ cnpgDetachedClusterUID = "radar.skyhook.io/detached-cluster-uid"
+)
+
+var (
+ cnpgGrantDeletePVCs = auth.Grant{Verb: "delete", Resource: "persistentvolumeclaims"}
+ cnpgGrantUpdatePVCs = auth.Grant{Verb: "update", Resource: "persistentvolumeclaims"}
+ cnpgGrantListPVCs = auth.Grant{Verb: "list", Resource: "persistentvolumeclaims"}
+ grantGetPVCs = auth.Grant{Verb: "get", Resource: "persistentvolumeclaims"}
+ cnpgGrantListJobs = auth.Grant{Verb: "list", Group: "batch", Resource: "jobs"}
+ cnpgGrantDeleteJobs = auth.Grant{Verb: "delete", Group: "batch", Resource: "jobs"}
+ cnpgGrantPatchPool = auth.Grant{Verb: "patch", Group: Group, Resource: "poolers"}
+
+ cnpgExtraActionGrants = map[string]auth.Grant{
+ "cancelBackend": grantCreateExec,
+ "terminateBackend": grantCreateExec,
+ "destroyInstance": cnpgGrantDeletePods,
+ "poolerPause": cnpgGrantPatchPool,
+ "poolerResume": cnpgGrantPatchPool,
+ }
+)
+
+// cnpgDestroyGrants: with keepPVC the PVCs are updated (detached), otherwise
+// deleted; the fence on the destroyed name is lifted last (patch clusters).
+func cnpgDestroyGrants(keepPVC bool) []auth.Grant {
+ pvc := cnpgGrantDeletePVCs
+ if keepPVC {
+ pvc = cnpgGrantUpdatePVCs
+ }
+ return []auth.Grant{cnpgGrantDeletePods, cnpgGrantListPVCs, pvc, cnpgGrantListJobs, cnpgGrantDeleteJobs, GrantPatchClusters}
+}
+
+// cnpgDestroyPreflight asks the apiserver, as the caller, for every grant the
+// sequence needs before its first write, so a missing one refuses the action
+// instead of stopping it halfway.
+func cnpgDestroyPreflight(ctx context.Context, x *cnpgClusterRun, keepPVC bool) error {
+ namespace := x.cluster.GetNamespace()
+ for _, g := range cnpgDestroyGrants(keepPVC) {
+ resource := g.Resource
+ if g.Subresource != "" {
+ resource += "/" + g.Subresource
+ }
+ allowed, apiErr := k8score.CanI(ctx, x.c.Typed, namespace, g.Group, resource, g.Verb)
+ switch {
+ case apiErr:
+ return integration.RefuseAction(http.StatusServiceUnavailable, "", "Could not confirm you may %s; nothing was changed", g.In(namespace).String())
+ case !allowed:
+ return integration.RefuseAction(http.StatusForbidden, "", "Destroying needs %s; nothing was changed", g.In(namespace).String())
+ }
+ }
+ return nil
+}
+
+func cnpgDestroyInstanceBlocker(f CNPGClusterFacts, i CNPGInstanceFact) (string, string) {
+ if r := cnpgGuardCommon(f); r != "" {
+ return "state_blocked", r
+ }
+ switch {
+ case f.Hibernated:
+ return "hibernated", "The cluster is hibernated"
+ case i.Pod == f.CurrentPrimary:
+ return "primary", "It is the primary: destroying it forces an unplanned failover. Switch over to a standby first"
+ case i.Pod == f.TargetPrimary:
+ return "switchover_target", "It is the target of a switchover"
+ case f.switchoverInFlight() != "", f.Phase == cnpgPhaseSwitchover, f.Phase == cnpgPhaseFailover:
+ return "switchover_in_progress", "A switchover or failover is in progress"
+ case !f.FencedInstances.Fences(i.Pod):
+ return "fence_required", "Fence " + i.Pod + " first: a fenced instance cannot be promoted while it is destroyed"
+ }
+ return "", ""
+}
+
+func cnpgGuardDestroyInstance(f CNPGClusterFacts, i CNPGInstanceFact) string {
+ _, reason := cnpgDestroyInstanceBlocker(f, i)
+ return reason
+}
+
+// cnpgPodLabelledPrimary reads the role the instance manager writes on its
+// own Pod, which can lead status.currentPrimary during a promotion.
+func cnpgPodLabelledPrimary(pod *corev1.Pod) bool { return cnpg.InstanceRole(pod) == "primary" }
+
+type CNPGReviewedObject struct {
+ Name string `json:"name"`
+ UID string `json:"uid"`
+}
+
+// CNPGDestroyPVC is one volume the destroy acts on.
+type CNPGDestroyPVC struct {
+ Name string `json:"name"`
+ UID string `json:"uid"`
+ Role string `json:"role,omitempty"`
+ Tablespace string `json:"tablespace,omitempty"`
+ Owned bool `json:"owned"`
+ Detached bool `json:"detached"`
+ Capacity string `json:"capacity,omitempty"`
+ StorageClass string `json:"storageClass,omitempty"`
+}
+
+// CNPGDestroyPlan is GET /api/cnpg/clusters/{ns}/{name}/instances/{pod}/destroy-plan.
+type CNPGDestroyPlan struct {
+ UID string `json:"uid"`
+ Context string `json:"context"`
+ Facts CNPGClusterFacts `json:"facts"`
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+ Role string `json:"role"`
+ PVCsReadable bool `json:"pvcsReadable"`
+ PVCReason string `json:"pvcReason,omitempty"`
+ PVCs []CNPGDestroyPVC `json:"pvcs"`
+ JobsReadable bool `json:"jobsReadable"`
+ Jobs []CNPGReviewedObject `json:"jobs"`
+ Actions struct {
+ Delete integration.ActionCapability `json:"delete"`
+ Keep integration.ActionCapability `json:"keep"`
+ } `json:"actions"`
+}
+
+// cnpgOwnedByCluster matches upstream's IsOwnedByCluster, tightened to this
+// Cluster's UID so a same-named predecessor's leftovers are not taken.
+func cnpgOwnedByCluster(refs []metav1.OwnerReference, cluster string, uid types.UID) bool {
+ for _, ref := range refs {
+ if ref.Kind != "Cluster" || ref.Name != cluster || ref.UID != uid {
+ continue
+ }
+ if gv, err := schema.ParseGroupVersion(ref.APIVersion); err == nil && gv.Group == Group {
+ return true
+ }
+ }
+ return false
+}
+
+// cnpgInstancePVCs lists the instance's PVCs as upstream's GetInstancePVCs
+// does (instance label plus a PVC role) and keeps those upstream would act
+// on: owned by the Cluster, or detached by an earlier keep-pvc destroy.
+func cnpgInstancePVCs(ctx context.Context, typed kubernetes.Interface, namespace, cluster string, clusterUID types.UID, instance string) ([]corev1.PersistentVolumeClaim, error) {
+ role, err := labels.NewRequirement(cnpgPVCRoleLabel, selection.In, []string{"PG_DATA", "PG_WAL", "PG_TABLESPACE"})
+ if err != nil {
+ return nil, err
+ }
+ inst, err := labels.NewRequirement(instanceNameLabel, selection.Equals, []string{instance})
+ if err != nil {
+ return nil, err
+ }
+ list, err := typed.CoreV1().PersistentVolumeClaims(namespace).List(ctx, metav1.ListOptions{LabelSelector: labels.NewSelector().Add(*inst, *role).String()})
+ if err != nil {
+ return nil, err
+ }
+ out := []corev1.PersistentVolumeClaim{}
+ for _, p := range list.Items {
+ owned := cnpgOwnedByCluster(p.OwnerReferences, cluster, clusterUID)
+ detached := p.Annotations[cnpgPVCStatusAnnotation] == cnpgPVCStatusDetached && p.Annotations[cnpgDetachedClusterUID] == string(clusterUID) && clusterUID != "" && p.Labels[instanceNameLabel] == instance
+ if owned || detached {
+ out = append(out, p)
+ }
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Name < out[j].Name })
+ return out, nil
+}
+
+func cnpgDestroyPVCOf(p corev1.PersistentVolumeClaim, cluster string, clusterUID types.UID) CNPGDestroyPVC {
+ out := CNPGDestroyPVC{
+ Name: p.Name, UID: string(p.UID), Role: p.Labels[cnpgPVCRoleLabel], Tablespace: p.Labels[cnpgTablespaceNameLabel],
+ Owned: cnpgOwnedByCluster(p.OwnerReferences, cluster, clusterUID),
+ Detached: p.Annotations[cnpgPVCStatusAnnotation] == cnpgPVCStatusDetached,
+ }
+ if q, ok := p.Status.Capacity[corev1.ResourceStorage]; ok {
+ out.Capacity = q.String()
+ }
+ if p.Spec.StorageClassName != nil {
+ out.StorageClass = *p.Spec.StorageClassName
+ }
+ return out
+}
+
+func cnpgInstanceJobs(ctx context.Context, typed kubernetes.Interface, namespace, cluster string, clusterUID types.UID, instance string) ([]batchv1.Job, error) {
+ list, err := typed.BatchV1().Jobs(namespace).List(ctx, metav1.ListOptions{LabelSelector: labels.Set{instanceNameLabel: instance}.String()})
+ if err != nil {
+ return nil, err
+ }
+ jobs := make([]batchv1.Job, 0, len(list.Items))
+ for _, j := range list.Items {
+ if cnpgOwnedByCluster(j.OwnerReferences, cluster, clusterUID) {
+ jobs = append(jobs, j)
+ }
+ }
+ sort.Slice(jobs, func(i, j int) bool { return jobs[i].Name < jobs[j].Name })
+ return jobs, nil
+}
+
+func (s *Reader) DestroyPlan(callerCtx context.Context, c ActionClients, contextName, namespace, name, pod string) (*CNPGDestroyPlan, error) {
+ ctx := callerCtx
+ cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ facts, _ := cnpgClusterFactsOf(ctx, c.Typed, cluster)
+ inst, ok := facts.instance(pod)
+ if !ok {
+ return nil, integration.RefuseAction(http.StatusNotFound, "", "%s is not an instance of Cluster %s/%s", pod, namespace, name)
+ }
+ plan := &CNPGDestroyPlan{
+ UID: string(cluster.GetUID()), Context: contextName, Facts: facts,
+ Pod: pod, PodUID: inst.PodUID, Role: inst.Role, PVCs: []CNPGDestroyPVC{}, Jobs: []CNPGReviewedObject{},
+ }
+ pvcs, err := cnpgInstancePVCs(ctx, c.Typed, namespace, name, cluster.GetUID(), pod)
+ if err == nil {
+ plan.PVCsReadable = true
+ for _, p := range pvcs {
+ plan.PVCs = append(plan.PVCs, cnpgDestroyPVCOf(p, name, cluster.GetUID()))
+ }
+ } else {
+ plan.PVCReason = cnpgReadReason(err)
+ }
+ if jobs, err := cnpgInstanceJobs(ctx, c.Typed, namespace, name, cluster.GetUID(), pod); err == nil {
+ plan.JobsReadable = true
+ for _, j := range jobs {
+ plan.Jobs = append(plan.Jobs, CNPGReviewedObject{Name: j.Name, UID: string(j.UID)})
+ }
+ }
+ reasonCode, guard := cnpgDestroyInstanceBlocker(facts, inst)
+ if guard == "" && !plan.JobsReadable {
+ guard = "The instance's Jobs cannot be read, so they cannot be reviewed"
+ reasonCode = "jobs_unreadable"
+ }
+ if guard == "" && !plan.PVCsReadable {
+ guard = "The instance's PVCs cannot be read (" + plan.PVCReason + "), so they cannot be reviewed"
+ reasonCode = "pvcs_unreadable"
+ }
+ cap := func(keep bool) integration.ActionCapability {
+ gs := cnpgDestroyGrants(keep)
+ ps := make([]string, len(gs))
+ for i, g := range gs {
+ gs[i] = g.In(namespace)
+ ps[i] = s.Access.Permission(callerCtx, gs[i])
+ }
+ actionGuard, actionCode := guard, reasonCode
+ if keep && actionGuard == "" {
+ actionGuard = cnpgKeepPVCBlocker(pvcs, name, cluster.GetUID())
+ if actionGuard != "" {
+ actionCode = "owners_remain"
+ }
+ }
+ capability := integration.CapabilityVerdict(actionGuard, ps, gs)
+ if capability.Permission != integration.PermissionDenied {
+ capability.ReasonCode = actionCode
+ }
+ return capability
+ }
+ plan.Actions.Delete = cap(false)
+ plan.Actions.Keep = cap(true)
+ return plan, nil
+}
+
+func cnpgKeepPVCBlocker(pvcs []corev1.PersistentVolumeClaim, cluster string, uid types.UID) string {
+ for _, pvc := range pvcs {
+ var owners []string
+ for _, ref := range pvc.OwnerReferences {
+ if !cnpgOwnedByCluster([]metav1.OwnerReference{ref}, cluster, uid) {
+ owners = append(owners, ref.Kind+" "+ref.Name+" ("+ref.APIVersion+", UID "+string(ref.UID)+")")
+ }
+ }
+ if len(owners) > 0 {
+ return "Keeping PVC " + pvc.Name + " would leave other owners able to garbage-collect it: " + strings.Join(owners, ", ") + ". Review its owner references first"
+ }
+ }
+ return ""
+}
+
+func cnpgReadReason(err error) string {
+ if apierrors.IsForbidden(err) {
+ return "no access"
+ }
+ return cnpgPlainReadError(err.Error())
+}
+
+type cnpgDestroyParams struct {
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+ KeepPVC bool `json:"keepPVC"`
+ PVCs *[]CNPGReviewedObject `json:"pvcs"`
+ Jobs *[]CNPGReviewedObject `json:"jobs"`
+}
+
+func cnpgRunDestroyInstance(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgDestroyParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.Pod == "" || p.PVCs == nil || p.Jobs == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.pod, params.podUID, params.pvcs and params.jobs are required: the confirmation must bind the Pod, volumes and Jobs reviewed")
+ }
+ inst, ok := x.facts.instance(p.Pod)
+ if !ok {
+ return nil, integration.ChangedAction(x.facts, "%s is not an instance of this cluster", p.Pod)
+ }
+ if r := cnpgGuardDestroyInstance(x.facts, inst); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ if !inst.PodReadable {
+ return nil, integration.RefuseAction(http.StatusForbidden, "", "Pod %s cannot be read, so it cannot be verified as this cluster's instance", p.Pod)
+ }
+ if inst.PodUID != p.PodUID {
+ return nil, integration.ChangedAction(x.facts, "Pod %s changed since you reviewed it", p.Pod)
+ }
+ if err := cnpgDestroyPreflight(ctx, x, p.KeepPVC); err != nil {
+ return nil, err
+ }
+ namespace, cluster, clusterUID := x.cluster.GetNamespace(), x.cluster.GetName(), x.cluster.GetUID()
+ pvcs, err := cnpgInstancePVCs(ctx, x.c.Typed, namespace, cluster, clusterUID, p.Pod)
+ if err != nil {
+ return nil, err
+ }
+ now := make([]CNPGReviewedObject, 0, len(pvcs))
+ for _, v := range pvcs {
+ now = append(now, CNPGReviewedObject{Name: v.Name, UID: string(v.UID)})
+ }
+ if !cnpgSameReviewedObjects(now, *p.PVCs) {
+ return nil, integration.ChangedAction(x.facts, "The volumes of %s changed since you reviewed them (now: %s); review the action again", p.Pod, cnpgPVCList(now))
+ }
+
+ if p.KeepPVC {
+ if reason := cnpgKeepPVCBlocker(pvcs, cluster, clusterUID); reason != "" {
+ return nil, integration.BlockedAction(reason)
+ }
+ }
+ jobs, err := cnpgInstanceJobs(ctx, x.c.Typed, namespace, cluster, clusterUID, p.Pod)
+ if err != nil {
+ return nil, err
+ }
+ reviewedJobs := make([]CNPGReviewedObject, 0, len(jobs))
+ jobNames := make([]string, 0, len(jobs))
+ for _, j := range jobs {
+ reviewedJobs = append(reviewedJobs, CNPGReviewedObject{Name: j.Name, UID: string(j.UID)})
+ jobNames = append(jobNames, j.Name)
+ }
+ if !cnpgSameReviewedObjects(reviewedJobs, *p.Jobs) {
+ return nil, integration.ChangedAction(x.facts, "The Jobs of %s changed since you reviewed them; review the action again", p.Pod)
+ }
+
+ // CloudNativePG offers no lock against a failover, so the instance is
+ // re-verified as a standby before every destructive step: the window in
+ // which a promotion can land is one request, not the whole sequence.
+ var completed []string
+ podDeleted := false
+ stop := func(err error) error {
+ if len(completed) == 0 {
+ return err
+ }
+ return integration.PartialAction(completed, err)
+ }
+ recheck := func() error {
+ fresh, err := x.c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, cluster, metav1.GetOptions{})
+ if err != nil {
+ return fmt.Errorf("re-reading Cluster %s/%s: %w", namespace, cluster, err)
+ }
+ facts, _ := cnpgClusterFactsOf(ctx, nil, fresh)
+ if fresh.GetUID() != clusterUID {
+ return integration.ChangedAction(facts, "Cluster %s/%s was deleted and recreated; nothing further was done", namespace, cluster)
+ }
+ if r := cnpgGuardDestroyInstance(facts, CNPGInstanceFact{Pod: p.Pod}); r != "" {
+ return integration.ChangedAction(facts, "%s can no longer be destroyed: %s", p.Pod, r)
+ }
+ pod, err := x.c.Typed.CoreV1().Pods(namespace).Get(ctx, p.Pod, metav1.GetOptions{})
+ switch {
+ case apierrors.IsNotFound(err):
+ return nil
+ case err != nil:
+ return fmt.Errorf("re-reading Pod %s: %w", p.Pod, err)
+ case podDeleted:
+ return nil
+ case string(pod.UID) != p.PodUID:
+ return integration.ChangedAction(facts, "Pod %s was replaced since you reviewed it", p.Pod)
+ case cnpgPodLabelledPrimary(pod):
+ return integration.ChangedAction(facts, "%s is now labelled primary: it can no longer be destroyed", p.Pod)
+ }
+ return nil
+ }
+
+ if p.KeepPVC {
+ for i := range pvcs {
+ pvc := &pvcs[i]
+ if !cnpgOwnedByCluster(pvc.OwnerReferences, cluster, clusterUID) {
+ continue
+ }
+ if err := recheck(); err != nil {
+ return nil, stop(err)
+ }
+ refs := pvc.OwnerReferences[:0:0]
+ for _, ref := range pvc.OwnerReferences {
+ if !cnpgOwnedByCluster([]metav1.OwnerReference{ref}, cluster, clusterUID) {
+ refs = append(refs, ref)
+ }
+ }
+ pvc.OwnerReferences = refs
+ if pvc.Annotations == nil {
+ pvc.Annotations = map[string]string{}
+ }
+ if pvc.Labels == nil {
+ pvc.Labels = map[string]string{}
+ }
+ pvc.Annotations[cnpgPVCStatusAnnotation] = cnpgPVCStatusDetached
+ pvc.Annotations[cnpgDetachedClusterUID] = string(clusterUID)
+ pvc.Labels[instanceNameLabel] = p.Pod
+ if _, err := x.c.Typed.CoreV1().PersistentVolumeClaims(namespace).Update(ctx, pvc, metav1.UpdateOptions{}); err != nil {
+ return nil, stop(fmt.Errorf("detaching PVC %s: %w", pvc.Name, err))
+ }
+ completed = append(completed, "detached PVC "+pvc.Name)
+ }
+ } else {
+ for _, pvc := range pvcs {
+ if err := recheck(); err != nil {
+ return nil, stop(err)
+ }
+ uid := pvc.UID
+ err := x.c.Typed.CoreV1().PersistentVolumeClaims(namespace).Delete(ctx, pvc.Name, metav1.DeleteOptions{Preconditions: &metav1.Preconditions{UID: &uid}})
+ if err != nil && !apierrors.IsNotFound(err) {
+ return nil, stop(fmt.Errorf("deleting PVC %s: %w", pvc.Name, err))
+ }
+ completed = append(completed, "deleted PVC "+pvc.Name)
+ }
+ }
+
+ if inst.PodExists {
+ if err := recheck(); err != nil {
+ return nil, stop(err)
+ }
+ uid := types.UID(p.PodUID)
+ err := x.c.Typed.CoreV1().Pods(namespace).Delete(ctx, p.Pod, metav1.DeleteOptions{Preconditions: &metav1.Preconditions{UID: &uid}})
+ if err != nil && !apierrors.IsNotFound(err) {
+ return nil, stop(fmt.Errorf("deleting Pod %s: %w", p.Pod, err))
+ }
+ completed = append(completed, "deleted Pod "+p.Pod)
+ podDeleted = true
+ }
+
+ background := metav1.DeletePropagationBackground
+ for _, reviewed := range jobs {
+ job := reviewed
+ for attempt := 0; attempt < 2; attempt++ {
+ if err := recheck(); err != nil {
+ return nil, stop(err)
+ }
+ uid, rv := job.UID, job.ResourceVersion
+ err := x.c.Typed.BatchV1().Jobs(namespace).Delete(ctx, job.Name, metav1.DeleteOptions{PropagationPolicy: &background, Preconditions: &metav1.Preconditions{UID: &uid, ResourceVersion: &rv}})
+ if err == nil || apierrors.IsNotFound(err) {
+ break
+ }
+ if attempt == 0 && apierrors.IsConflict(err) {
+ // Status updates change the version without changing the reviewed
+ // identity. Re-list with the existing grant before one bounded retry.
+ fresh, readErr := cnpgInstanceJobs(ctx, x.c.Typed, namespace, cluster, clusterUID, p.Pod)
+ if readErr != nil {
+ return nil, stop(fmt.Errorf("re-reading Jobs: %w", readErr))
+ }
+ index := slices.IndexFunc(fresh, func(j batchv1.Job) bool { return j.Name == reviewed.Name && j.UID == reviewed.UID })
+ if index < 0 {
+ return nil, stop(integration.ChangedAction(x.facts, "Job %s was replaced or is no longer owned by this instance", reviewed.Name))
+ }
+ job = fresh[index]
+ continue
+ }
+ return nil, stop(fmt.Errorf("deleting Job %s: %w", job.Name, err))
+ }
+ completed = append(completed, "deleted Job "+reviewed.Name)
+ }
+
+ lifted, err := cnpgLiftDestroyedFence(ctx, x, p.Pod)
+ if err != nil {
+ return nil, stop(fmt.Errorf("lifting the fence on %s (lift it with Unfence): %w", p.Pod, err))
+ }
+ if lifted {
+ completed = append(completed, "lifted the fence on "+p.Pod)
+ }
+ log.Printf("[cnpg] destroyed instance %s of %s/%s (keepPVC=%v, pvcs=%d, jobs=%d)", k8s.SanitizeForLog(p.Pod), k8s.SanitizeForLog(namespace), k8s.SanitizeForLog(cluster), p.KeepPVC, len(pvcs), len(jobs))
+
+ keep := p.KeepPVC
+ msg := fmt.Sprintf("Instance %s destroyed; the operator creates a replacement instance", p.Pod)
+ if keep {
+ msg = fmt.Sprintf("Instance %s destroyed and its volumes kept, detached; the operator creates a replacement instance", p.Pod)
+ }
+ if !lifted {
+ msg += `. The cluster-wide fence ["*"] stays in place`
+ }
+ return &CNPGActionResult{
+ Message: msg,
+ Target: &CNPGActionTarget{Pod: p.Pod, PodUID: p.PodUID, KeepPVC: &keep, PVCs: now, Jobs: jobNames},
+ }, nil
+}
+
+func cnpgSameReviewedObjects(a, b []CNPGReviewedObject) bool {
+ key := func(l []CNPGReviewedObject) []string {
+ out := make([]string, len(l))
+ for i, v := range l {
+ out[i] = v.Name + "\x00" + v.UID
+ }
+ return out
+ }
+ return cnpgSameStringSet(key(a), key(b))
+}
+
+func cnpgPVCList(l []CNPGReviewedObject) string {
+ if len(l) == 0 {
+ return "none"
+ }
+ names := make([]string, len(l))
+ for i, v := range l {
+ names[i] = v.Name
+ }
+ return strings.Join(names, ", ")
+}
+
+// cnpgLiftDestroyedFence removes the destroyed name from a list-form fence,
+// against the Cluster as it is now. A ["*"] fence is the user's, not this
+// action's, and stays.
+func cnpgLiftDestroyedFence(ctx context.Context, x *cnpgClusterRun, pod string) (bool, error) {
+ fresh, err := x.c.Dynamic.Resource(ClusterGVR).Namespace(x.cluster.GetNamespace()).Get(ctx, x.cluster.GetName(), metav1.GetOptions{})
+ if err != nil {
+ return false, err
+ }
+ if fresh.GetUID() != x.cluster.GetUID() {
+ return false, fmt.Errorf("Cluster %s/%s was deleted and recreated; its fences were left alone", x.cluster.GetNamespace(), x.cluster.GetName())
+ }
+ f := parseCNPGFenced(fresh.GetAnnotations()[cnpgFencedAnnotation])
+ if f.All || f.Malformed || !slices.Contains(f.Instances, pod) {
+ return false, nil
+ }
+ rest := make([]string, 0, len(f.Instances))
+ for _, n := range f.Instances {
+ if n != pod {
+ rest = append(rest, n)
+ }
+ }
+ value, err := cnpgFencedValue(false, rest)
+ if err != nil {
+ return false, err
+ }
+ err = integration.MergePatchAtVersion(ctx, x.c.Dynamic, ClusterGVR, fresh, map[string]any{
+ "metadata": map[string]any{"annotations": map[string]any{cnpgFencedAnnotation: value}},
+ })
+ return err == nil, err
+}
diff --git a/internal/cnpg/errors.go b/internal/cnpg/errors.go
new file mode 100644
index 0000000000..1c914209b9
--- /dev/null
+++ b/internal/cnpg/errors.go
@@ -0,0 +1,230 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "regexp"
+ "strings"
+ "time"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+)
+
+// Transport failures between Radar and the Kubernetes API, as client-go words
+// them. None of these can come back from the Pod through the proxy.
+var cnpgAPITransportHints = []string{
+ "i/o timeout",
+ "eof",
+ "connection reset",
+ "connection refused",
+ "tls handshake timeout",
+ "no such host",
+ "broken pipe",
+ "use of closed network connection",
+ "http2: client connection lost",
+ "timeout awaiting response headers",
+ "client.timeout exceeded",
+}
+
+// cnpgTransportSentence says in plain words why a read through the Kubernetes
+// API failed in transit. The raw error carries the API server's address and
+// proxy URL, which mean nothing to the reader; callers log it instead. ok is
+// false for any other error, whose own text the caller keeps. port is the Pod
+// port that was being read (0 when none), timeout the caller's deadline.
+func cnpgTransportSentence(err error, port int, timeout time.Duration) (string, bool) {
+ if err == nil {
+ return "", false
+ }
+ lower := strings.ToLower(err.Error())
+ on := ""
+ if port > 0 {
+ on = fmt.Sprintf(" on port %d", port)
+ }
+ var status apierrors.APIStatus
+ // The API server answered and relayed a failure to reach the Pod.
+ relayed := (errors.As(err, &status) && status.Status().Code != 0) || strings.Contains(lower, "error trying to reach service")
+ switch {
+ case errors.Is(err, context.DeadlineExceeded) || strings.Contains(lower, "context deadline exceeded"):
+ return fmt.Sprintf("the Pod did not answer%s within %s", on, cnpgSeconds(timeout)), true
+ case errors.Is(err, context.Canceled) || strings.Contains(lower, "context canceled"):
+ return "the read was cancelled before the Kubernetes API answered", true
+ case strings.Contains(lower, "container not found"):
+ return "the container is not running", true
+ case relayed && strings.Contains(lower, "connection refused"):
+ if port > 0 {
+ return fmt.Sprintf("nothing is listening on port %d in the Pod", port), true
+ }
+ return "the Pod refused the connection", true
+ case relayed && (strings.Contains(lower, "no route to host") || strings.Contains(lower, "host is unreachable") || strings.Contains(lower, "network is unreachable")):
+ return "the Kubernetes API cannot reach the Pod's address", true
+ case relayed && strings.Contains(lower, "timeout"):
+ return "the Pod did not answer" + on, true
+ case relayed:
+ return "", false
+ }
+ for _, hint := range cnpgAPITransportHints {
+ if strings.Contains(lower, hint) {
+ return "the Kubernetes API did not answer", true
+ }
+ }
+ return "", false
+}
+
+func cnpgSeconds(d time.Duration) string {
+ return fmt.Sprintf("%g s", d.Seconds())
+}
+
+// cnpgPostgresSentence says in plain words why the instance manager or psql
+// could not talk to PostgreSQL, or ok false when the text is not one it
+// recognizes. A missing or refusing local socket means the server is down.
+func cnpgPostgresSentence(raw string) (string, bool) {
+ lower := strings.ToLower(raw)
+ connecting := strings.Contains(lower, ".s.pgsql.") || strings.Contains(lower, "dial unix") || strings.Contains(lower, "failed to connect")
+ down := strings.Contains(lower, "no such file or directory") || strings.Contains(lower, "connection refused")
+ switch {
+ case strings.Contains(lower, "the database system is starting up"):
+ return "PostgreSQL is starting up on this instance", true
+ case strings.Contains(lower, "the database system is shutting down"):
+ return "PostgreSQL is shutting down on this instance", true
+ case connecting && down:
+ return "PostgreSQL is not running on this instance", true
+ }
+ return "", false
+}
+
+// PostgreSQL conditions Radar names from a relayed SQLSTATE: the ones an
+// operator can act on. Other codes add nothing.
+var cnpgSQLStateNames = map[string]string{
+ "53000": "insufficient_resources",
+ "53100": "disk_full",
+ "53200": "out_of_memory",
+ "53300": "too_many_connections",
+ "53400": "configuration_limit_exceeded",
+ "57014": "query_canceled",
+ "57P01": "admin_shutdown",
+ "57P02": "crash_shutdown",
+ "57P03": "cannot_connect_now",
+ "57P04": "database_dropped",
+ "08000": "connection_exception",
+ "08001": "sqlclient_unable_to_establish_sqlconnection",
+ "08003": "connection_does_not_exist",
+ "08004": "sqlserver_rejected_establishment_of_sqlconnection",
+ "08006": "connection_failure",
+ "28000": "invalid_authorization_specification",
+ "28P01": "invalid_password",
+ "3D000": "invalid_catalog_name",
+ "42501": "insufficient_privilege",
+ "25006": "read_only_sql_transaction",
+ "40001": "serialization_failure",
+ "40P01": "deadlock_detected",
+ "55000": "object_not_in_prerequisite_state",
+ "55006": "object_in_use",
+ "55P03": "lock_not_available",
+ "XX000": "internal_error",
+ "XX001": "data_corrupted",
+ "XX002": "index_corrupted",
+}
+
+// PostgreSQL messages Radar recognizes, each replaced by a fixed phrase.
+// Matching text is never echoed: it can sit next to user names, hosts and
+// connection strings.
+var cnpgPostgresPhrases = []struct{ match, phrase string }{
+ {"out of shared memory", "out of shared memory"},
+ {"out of memory", "out of memory"},
+ {"too many clients already", "too many clients already"},
+ {"remaining connection slots are reserved", "remaining connection slots are reserved"},
+ {"the database system is starting up", "the database system is starting up"},
+ {"the database system is shutting down", "the database system is shutting down"},
+ {"the database system is in recovery mode", "the database system is in recovery mode"},
+ {"no space left on device", "no space left on device"},
+ {"could not extend file", "could not extend a data file"},
+ {"password authentication failed", "password authentication failed"},
+}
+
+var cnpgSQLStateRe = regexp.MustCompile(`SQLSTATE[ :]*([0-9A-Z]{5})\b`)
+
+// cnpgPostgresDetail names what PostgreSQL reported inside a relayed error,
+// from an allowlist only: a fixed phrase for a recognized message and the
+// condition name of a recognized SQLSTATE. It never returns the error's own
+// text, so nothing it carries (credentials, hosts, user names) reaches the
+// UI. "" when neither is recognized.
+func cnpgPostgresDetail(raw string) string {
+ lower := strings.ToLower(raw)
+ phrase := ""
+ for _, p := range cnpgPostgresPhrases {
+ if strings.Contains(lower, p.match) {
+ phrase = p.phrase
+ break
+ }
+ }
+ condition := ""
+ if m := cnpgSQLStateRe.FindStringSubmatch(raw); m != nil {
+ if name, ok := cnpgSQLStateNames[m[1]]; ok {
+ condition = name + " (SQLSTATE " + m[1] + ")"
+ }
+ }
+ switch {
+ case phrase != "" && condition != "":
+ return phrase + ", " + condition
+ case phrase != "":
+ return phrase
+ }
+ return condition
+}
+
+// Prometheus client_golang's error page, which the exporters serve on a
+// failed scrape; nothing in front of the Pod writes it.
+const cnpgPromHTTPErrorPrefix = "An error has occurred while serving metrics"
+
+// cnpgRelayedPodSentence says in plain words what a 5xx relayed through
+// pods/proxy means. A relayed body proves nothing about where it came from:
+// a gateway or the apiserver's own timeout page carries the same shape. So
+// the instance manager or exporter is named only when the body says so —
+// a PostgreSQL failure Radar recognizes on the status port, or client_golang's
+// error page on a metrics port. Anything else is worded without an origin.
+// The body can name sockets and connection strings; callers log it instead.
+// ok is false for errors that are not a relayed 5xx.
+func cnpgRelayedPodSentence(err error, port int) (string, bool) {
+ body, code, ok := cnpgRelayedPodBody(err)
+ if !ok || code < 500 {
+ return "", false
+ }
+ if strings.Contains(strings.ToLower(body), "error trying to reach service") {
+ return "", false
+ }
+ if port == cnpgStatusPort {
+ if plain, ok := cnpgPostgresSentence(body); ok {
+ return plain, true
+ }
+ if detail := cnpgPostgresDetail(body); detail != "" {
+ return "the instance manager could not read PostgreSQL's status: " + detail, true
+ }
+ } else if strings.HasPrefix(strings.TrimSpace(body), cnpgPromHTTPErrorPrefix) {
+ return fmt.Sprintf("the metrics endpoint on port %d answered with an error (HTTP %d)", port, code), true
+ }
+ return fmt.Sprintf("the read through the Kubernetes API failed with HTTP %d (from a gateway or the Pod; Radar can't tell which)", code), true
+}
+
+// cnpgRelayedPodBody is the body of an answer the apiserver relayed from the
+// Pod through pods/proxy, with its HTTP code. client-go marks a response that
+// is not an apiserver Status (the Pod's own answer) with an
+// UnexpectedServerResponse cause carrying the body; errors the apiserver
+// produced itself (Timeout, ServiceUnavailable, InternalError) carry none.
+func cnpgRelayedPodBody(err error) (string, int, bool) {
+ var status apierrors.APIStatus
+ if err == nil || !errors.As(err, &status) {
+ return "", 0, false
+ }
+ st := status.Status()
+ if st.Details == nil {
+ return "", 0, false
+ }
+ for _, c := range st.Details.Causes {
+ if c.Type == metav1.CauseTypeUnexpectedServerResponse {
+ return c.Message, int(st.Code), true
+ }
+ }
+ return "", 0, false
+}
diff --git a/internal/cnpg/errors_test.go b/internal/cnpg/errors_test.go
new file mode 100644
index 0000000000..d65383951d
--- /dev/null
+++ b/internal/cnpg/errors_test.go
@@ -0,0 +1,207 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "strings"
+ "testing"
+ "time"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+)
+
+func TestCNPGTransportSentence(t *testing.T) {
+ relayed := apierrors.NewServiceUnavailable("error trying to reach service: dial tcp 10.0.0.5:9187: connect: connection refused")
+ cases := []struct {
+ err error
+ port int
+ want string
+ }{
+ {context.DeadlineExceeded, 8000, "the Pod did not answer on port 8000 within 5 s"},
+ {errors.New(`Get "https://127.0.0.1:55484/api/v1/namespaces/pg/pods/http:pg-1:9127/proxy/metrics": EOF`), 9127, "the Kubernetes API did not answer"},
+ {errors.New("read tcp 127.0.0.1:56435->127.0.0.1:55484: i/o timeout"), 0, "the Kubernetes API did not answer"},
+ {relayed, 9187, "nothing is listening on port 9187 in the Pod"},
+ {fmt.Errorf("error trying to reach service: dial tcp 10.0.0.5:8000: connect: no route to host"), 8000, "the Kubernetes API cannot reach the Pod's address"},
+ }
+ for _, c := range cases {
+ got, ok := cnpgTransportSentence(c.err, c.port, 5*time.Second)
+ if !ok || got != c.want {
+ t.Errorf("%v: %q (%v), want %q", c.err, got, ok, c.want)
+ }
+ if strings.Contains(got, "127.0.0.1") {
+ t.Errorf("address leaked: %q", got)
+ }
+ }
+ for _, err := range []error{
+ errors.New("psql: error: FATAL: the database system is shutting down"),
+ apierrors.NewForbidden(schema.GroupResource{Resource: "pods"}, "pg-1", errors.New("denied")),
+ } {
+ if got, ok := cnpgTransportSentence(err, 0, 5*time.Second); ok {
+ t.Errorf("%v is not a transport failure, got %q", err, got)
+ }
+ }
+}
+
+func TestCNPGRelayedPodSentence(t *testing.T) {
+ socketGone := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-wal-failing-1:8000",
+ "failed to connect to `user=postgres database=postgres`: /controller/run/.s.PGSQL.5432 (/controller/run): dial error: dial unix /controller/run/.s.PGSQL.5432: connect: no such file or directory", 0, true)
+ other := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", "pq: out of shared memory", 0, true)
+ exporter := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "http:pg-1:9187", "An error has occurred while serving metrics:\n\ncollector failed", 0, true)
+ unknownExporter := apierrors.NewGenericServerResponse(503, "get", schema.GroupResource{Resource: "pods"}, "http:pg-1:9187", "collector failed", 0, true)
+ gateway := apierrors.NewGenericServerResponse(504, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", "504 Gateway Time-out 504 Gateway Time-out nginx ", 0, true)
+ untypedStatus := apierrors.NewGenericServerResponse(504, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", `{"kind":"Status","status":"Failure","message":"Timeout: request did not complete within the allotted timeout","code":504}`, 0, true)
+ opaque := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", "internal server error", 0, true)
+ neutral := func(code int) string {
+ return fmt.Sprintf("the read through the Kubernetes API failed with HTTP %d (from a gateway or the Pod; Radar can't tell which)", code)
+ }
+ cases := []struct {
+ err error
+ port int
+ want string
+ }{
+ {socketGone, cnpgStatusPort, "PostgreSQL is not running on this instance"},
+ {other, cnpgStatusPort, "the instance manager could not read PostgreSQL's status: out of shared memory"},
+ {exporter, cnpgMetricsPort, "the metrics endpoint on port 9187 answered with an error (HTTP 500)"},
+ {unknownExporter, cnpgMetricsPort, neutral(503)},
+ {gateway, cnpgStatusPort, neutral(504)},
+ {untypedStatus, cnpgStatusPort, neutral(504)},
+ {opaque, cnpgStatusPort, neutral(500)},
+ }
+ for _, c := range cases {
+ got, ok := cnpgRelayedPodSentence(c.err, c.port)
+ if !ok || got != c.want {
+ t.Errorf("%v: %q (%v), want %q", c.err, got, ok, c.want)
+ }
+ if strings.Contains(got, "controller/run") || strings.Contains(got, "pg-") {
+ t.Errorf("raw text leaked: %q", got)
+ }
+ }
+ for _, err := range []error{
+ apierrors.NewServiceUnavailable("error trying to reach service: dial tcp 10.0.0.5:8000: connect: no route to host"),
+ apierrors.NewBadRequest("bad"),
+ errors.New("plain"),
+ } {
+ if got, ok := cnpgRelayedPodSentence(err, cnpgStatusPort); ok {
+ t.Errorf("%v is not the Pod's own error answer, got %q", err, got)
+ }
+ }
+}
+
+func TestClassifyCNPGProxyFailureReplacesRelayedInstanceManagerError(t *testing.T) {
+ err := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000",
+ "failed to connect to `user=postgres database=postgres`: dial unix /controller/run/.s.PGSQL.5432: connect: no such file or directory", 0, true)
+ out := classifyCNPGProxyFailure(context.Background(), err, cnpgProxyOutcome{}, proxyTarget{namespace: "pg", pod: "pg-1", port: cnpgStatusPort, path: cnpgStatusPath})
+ if out.err != "PostgreSQL is not running on this instance" {
+ t.Errorf("err = %q", out.err)
+ }
+}
+
+func TestCNPGPostgresSentence(t *testing.T) {
+ for raw, want := range map[string]string{
+ `psql: error: connection to server on socket "/controller/run/.s.PGSQL.5432" failed: No such file or directory`: "PostgreSQL is not running on this instance",
+ "FATAL: the database system is starting up": "PostgreSQL is starting up on this instance",
+ } {
+ if got, ok := cnpgPostgresSentence(raw); !ok || got != want {
+ t.Errorf("%q: %q, want %q", raw, got, want)
+ }
+ }
+ if got, ok := cnpgPostgresSentence("ERROR: permission denied for table x"); ok {
+ t.Errorf("unrecognized text mapped to %q", got)
+ }
+}
+
+func TestCNPGPostgresRefusalIsNotATransportFailure(t *testing.T) {
+ relayed := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000",
+ "failed to connect to `user=postgres database=postgres`: /controller/run/.s.PGSQL.5432 (/controller/run): dial error: dial unix /controller/run/.s.PGSQL.5432: connect: connection refused", 0, true)
+ out := classifyCNPGProxyFailure(context.Background(), relayed, cnpgProxyOutcome{}, proxyTarget{namespace: "pg", pod: "pg-1", port: cnpgStatusPort, path: cnpgStatusPath})
+ if out.err != "PostgreSQL is not running on this instance" {
+ t.Errorf("relayed socket refusal = %q", out.err)
+ }
+ psql := errors.New(`command terminated with exit code 2: psql: error: connection to server on socket "/controller/run/.s.PGSQL.5432" failed: Connection refused`)
+ if got := cnpgExecSourceState(psql); got.Error != "PostgreSQL is not running on this instance" || got.State != runtimeStateError {
+ t.Errorf("psql socket refusal = %+v", got)
+ }
+}
+
+func TestCNPGPostgresDetailKeepsTheDiagnosisWithoutConnectionFacts(t *testing.T) {
+ shm := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", "while reading status: pq: out of shared memory", 0, true)
+ if got, _ := cnpgRelayedPodSentence(shm, cnpgStatusPort); got != "the instance manager could not read PostgreSQL's status: out of shared memory" {
+ t.Errorf("shared memory = %q", got)
+ }
+ slots := "query failed for user=postgres database=app host=10.0.0.5:5432: FATAL: remaining connection slots are reserved (SQLSTATE 53300)"
+ if got := cnpgPostgresDetail(slots); got != "remaining connection slots are reserved, too_many_connections (SQLSTATE 53300)" {
+ t.Errorf("slots = %q", got)
+ }
+ if got := cnpgPostgresDetail("ERROR: something odd (SQLSTATE 57P03)"); got != "cannot_connect_now (SQLSTATE 57P03)" {
+ t.Errorf("code only = %q", got)
+ }
+ for _, unknown := range []string{"the instance manager crashed", "ERROR: relation \"secret_table\" does not exist (SQLSTATE 42P01)"} {
+ if d := cnpgPostgresDetail(unknown); d != "" {
+ t.Errorf("%q: unrecognized text must add nothing, got %q", unknown, d)
+ }
+ }
+}
+
+func TestCNPGPostgresDetailNeverEchoesTheError(t *testing.T) {
+ leaks := []string{
+ "pq: invalid connection string: postgresql://alice:secret@pg.private.example/db",
+ "pq: out of shared memory while connecting with password='canary secret-tail' host=pg.private.example",
+ "FATAL: password authentication failed for user \"alice\" (host = pg.private.example, password = secret tail)",
+ "ERROR: could not connect to server pg.private.example at 10.1.2.3 as alice, secret tail (SQLSTATE 08001)",
+ "FATAL: too many clients already from [2001:db8::1]:5432 user alice password secret-tail",
+ }
+ for _, raw := range leaks {
+ err := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", raw, 0, true)
+ got, _ := cnpgRelayedPodSentence(err, cnpgStatusPort)
+ for _, leak := range []string{"alice", "secret", "pg.private", "10.1.2.3", "2001:db8", "tail", "canary"} {
+ if strings.Contains(got, leak) {
+ t.Errorf("%q leaked %q: %q", raw, leak, got)
+ }
+ }
+ }
+}
+
+func TestCNPGRelayedPodSentenceIgnoresAPIServerErrors(t *testing.T) {
+ for _, err := range []error{
+ apierrors.NewTimeoutError("request did not complete within requested timeout", 0),
+ apierrors.NewServiceUnavailable("the server is currently unable to handle the request"),
+ apierrors.NewInternalError(errors.New("etcd leader changed")),
+ apierrors.NewGenericServerResponse(504, "get", schema.GroupResource{Resource: "pods"}, "https:pg-1:8000", "", 0, false),
+ } {
+ if got, ok := cnpgRelayedPodSentence(err, cnpgStatusPort); ok {
+ t.Errorf("%v is the apiserver's own error, not the Pod's; got %q", err, got)
+ }
+ }
+ timeout := apierrors.NewTimeoutError("request did not complete within requested timeout", 0)
+ out := classifyCNPGProxyFailure(context.Background(), timeout, cnpgProxyOutcome{}, proxyTarget{namespace: "pg", pod: "pg-1", port: cnpgStatusPort, path: cnpgStatusPath})
+ if strings.Contains(out.err, "instance manager") {
+ t.Errorf("an apiserver timeout blamed the instance manager: %q", out.err)
+ }
+}
+
+func TestCNPGNoPrometheusReasonIsOneSentence(t *testing.T) {
+ for msg, want := range map[string]string{
+ "": "Radar is not connected to Prometheus",
+ "Radar found 2 services that may be Prometheus but may not port-forward to them (needs create pods/portforward)": "Radar found 2 services that may be Prometheus but may not port-forward to them (needs create pods/portforward)",
+ "context deadline exceeded": "Radar is not connected to Prometheus: context deadline exceeded",
+ "No working Prometheus endpoint found.\nCandidate cert-manager/cert-manager did not respond after port-forward": "No working Prometheus endpoint found.\nCandidate cert-manager/cert-manager did not respond after port-forward",
+ } {
+ if got := cnpgNoPrometheusReason(msg); got != want {
+ t.Errorf("%q: %q, want %q", msg, got, want)
+ }
+ }
+}
+
+func TestCNPGGrantClusterStringHasNoNestedParentheses(t *testing.T) {
+ g := auth.Grant{Verb: "get", Group: "admissionregistration.k8s.io", Resource: "mutatingwebhookconfigurations"}
+ if got := g.String(); got != "get mutatingwebhookconfigurations.admissionregistration.k8s.io cluster-wide" {
+ t.Errorf("ClusterString = %q", got)
+ }
+ if got := (auth.Grant{Verb: "get", Resource: "nodes"}).String(); got != "get nodes cluster-wide" {
+ t.Errorf("core = %q", got)
+ }
+}
diff --git a/internal/cnpg/handlers_test.go b/internal/cnpg/handlers_test.go
new file mode 100644
index 0000000000..900aaa7202
--- /dev/null
+++ b/internal/cnpg/handlers_test.go
@@ -0,0 +1,118 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "net/http"
+ "testing"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// A ClusterImageCatalog is cluster-scoped and referenceable from any namespace,
+// which is why this lookup is an endpoint rather than a client-side list: asking
+// the generic resources endpoint without a namespace inherits the caller's
+// namespace view filter, and "no cluster uses this catalog" is exactly the
+// sentence someone reads before editing it.
+
+// CloudNativePG defaults an omitted `kind` to the namespaced ImageCatalog. A
+// namespaced and a cluster-scoped catalog may share a name, so a reference that
+// omits the kind must not be counted against the cluster-scoped one.
+func TestCatalogRefMatches_DefaultsToTheNamespacedKind(t *testing.T) {
+ for _, c := range []struct {
+ name string
+ ref map[string]interface{}
+ catalog string
+ wantKind string
+ want bool
+ }{
+ {"omitted kind counts as ImageCatalog",
+ map[string]interface{}{"name": "pg17"}, "pg17", "ImageCatalog", true},
+ {"omitted kind is NOT a ClusterImageCatalog",
+ map[string]interface{}{"name": "pg17"}, "pg17", "ClusterImageCatalog", false},
+ {"explicit cluster-scoped matches its own kind",
+ map[string]interface{}{"name": "pg17", "kind": "ClusterImageCatalog"}, "pg17", "ClusterImageCatalog", true},
+ {"explicit cluster-scoped does not match the namespaced kind",
+ map[string]interface{}{"name": "pg17", "kind": "ClusterImageCatalog"}, "pg17", "ImageCatalog", false},
+ {"another catalog entirely",
+ map[string]interface{}{"name": "pg16"}, "pg17", "ImageCatalog", false},
+ {"no name at all",
+ map[string]interface{}{}, "pg17", "ImageCatalog", false},
+ } {
+ t.Run(c.name, func(t *testing.T) {
+ if got := catalogRefMatches(c.ref, c.catalog, c.wantKind); got != c.want {
+ t.Errorf("catalogRefMatches(%v, %q, %q) = %v, want %v",
+ c.ref, c.catalog, c.wantKind, got, c.want)
+ }
+ })
+ }
+}
+
+// The dynamic cache hands the same field back as int64 or float64 depending on
+// how the object entered it. Missing the float64 shape would drop a real major
+// to zero — which the screen reads as "the reference carries no major", a
+// different and wrong statement.
+func TestCatalogRefMajor_ReadsEitherNumberShape(t *testing.T) {
+ for _, c := range []struct {
+ name string
+ ref map[string]interface{}
+ want int
+ }{
+ {"int64 from a typed decode", map[string]interface{}{"major": int64(17)}, 17},
+ {"float64 from a JSON decode", map[string]interface{}{"major": float64(17)}, 17},
+ {"plain int", map[string]interface{}{"major": 17}, 17},
+ {"absent", map[string]interface{}{}, 0},
+ {"a string is not a major", map[string]interface{}{"major": "17"}, 0},
+ } {
+ t.Run(c.name, func(t *testing.T) {
+ if got := catalogRefMajor(c.ref); got != c.want {
+ t.Errorf("catalogRefMajor(%v) = %d, want %d", c.ref, got, c.want)
+ }
+ })
+ }
+}
+
+func TestCNPGCatalogUsersSeparatesAbsentCRDFromFailedRead(t *testing.T) {
+ type marker struct{}
+ ctx := context.WithValue(context.Background(), marker{}, "caller")
+ for _, tc := range []struct {
+ name string
+ readError error
+ status int
+ message string
+ }{
+ {name: "absent CRD", readError: k8s.ErrUnknownDynamicKind},
+ {name: "sync pending", readError: integration.ErrDynamicNotSynced, status: http.StatusServiceUnavailable, message: "clusters are still loading"},
+ {name: "read failed", readError: errors.New("read unavailable"), status: http.StatusServiceUnavailable, message: "could not read CloudNativePG clusters"},
+ {name: "successful empty inventory"},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ reader := newTestReader(nil)
+ calls := 0
+ reader.Observations.DynamicList = func(got context.Context, cache *k8s.ResourceCache, kind, group, namespace string) ([]*unstructured.Unstructured, error) {
+ calls++
+ if got.Value(marker{}) != "caller" || kind != "Cluster" || group != Group || namespace != "db" {
+ t.Fatal("catalog read lost caller or scope")
+ }
+ return nil, tc.readError
+ }
+ got, err := reader.CatalogUsers(ctx, nil, "db", "pg17", "ImageCatalog")
+ if calls != 1 {
+ t.Fatal("catalog did not read its bound inventory")
+ }
+ if tc.status == 0 {
+ if err != nil || got == nil || got.Clusters == nil || len(got.Clusters) != 0 {
+ t.Fatalf("absence=%+v %v", got, err)
+ }
+ return
+ }
+ var failure *ReadFailure
+ if got != nil || !errors.As(err, &failure) || failure.Status != tc.status || failure.Message != tc.message {
+ t.Fatalf("failure=%+v %v", got, err)
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/history.go b/internal/cnpg/history.go
new file mode 100644
index 0000000000..52a4a48aab
--- /dev/null
+++ b/internal/cnpg/history.go
@@ -0,0 +1,603 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "sort"
+ "strings"
+ "sync"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/selection"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+const (
+ historySourcePrometheus = "prometheus"
+ historySourceNone = "none"
+
+ historyStateOK = "ok"
+ cnpgHistoryStateAmbiguous = "ambiguous"
+ cnpgHistoryStateScopeMismatch = "scopeMismatch"
+ historyStateError = "error"
+ cnpgHistoryStateDenied = "denied"
+
+ fleetLagSource = "Prometheus cnpg_pg_replication_lag, standbys only (cnpg_pg_replication_in_recovery = 1)"
+ fleetReceiverSource = "Prometheus cnpg_pg_replication_is_wal_receiver_up, standbys only (cnpg_pg_replication_in_recovery = 1)"
+ fleetSlotsSource = "Prometheus cnpg_pg_replication_slots_active = 0 for physical slots, retained WAL from cnpg_pg_replication_slots_pg_wal_lsn_diff on each reporting instance"
+ fleetGrowthSource = "Prometheus deriv(kubelet_volume_stats_used_bytes) over 6h"
+ cnpgFleetGrowthWindow = 6 * time.Hour
+ cnpgHistoryAnchorCap = 100
+)
+
+var cnpgHistoryMemoTTL = 15 * time.Second
+
+// CNPGClusterHistoryResponse is GET /api/cnpg/clusters/{ns}/{name}/history.
+// Source "none" means Radar has no Prometheus; the UI then keeps its own
+// in-browser samples. State describes the query as a whole; each chart
+// carries its own state when the whole succeeded.
+type CNPGClusterHistoryResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ Source string `json:"source"`
+ State string `json:"state,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Range string `json:"range"`
+ Start string `json:"start,omitempty"`
+ End string `json:"end,omitempty"`
+ StepSeconds int `json:"stepSeconds,omitempty"`
+ Selector string `json:"selector,omitempty"`
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+ // PVCIsolation is the volume chart's own: claims are matched apart from Pods.
+ PVCIsolation *prometheuspkg.SeriesIsolation `json:"pvcIsolation,omitempty"`
+ SampledAt string `json:"sampledAt"`
+ Charts []prometheuspkg.CNPGHistoryChart `json:"charts"`
+}
+
+type cnpgHistoryMemoEntry struct {
+ expires time.Time
+ charts []prometheuspkg.CNPGHistoryChart
+}
+
+var (
+ cnpgHistoryMemoMu sync.Mutex
+ cnpgHistoryMemo = map[string]cnpgHistoryMemoEntry{}
+)
+
+func historyMemoGet(key string, now time.Time) ([]prometheuspkg.CNPGHistoryChart, bool) {
+ cnpgHistoryMemoMu.Lock()
+ defer cnpgHistoryMemoMu.Unlock()
+ e, ok := cnpgHistoryMemo[key]
+ if !ok || now.After(e.expires) {
+ return nil, false
+ }
+ return e.charts, true
+}
+
+func historyMemoPut(key string, now time.Time, charts []prometheuspkg.CNPGHistoryChart) {
+ cnpgHistoryMemoMu.Lock()
+ defer cnpgHistoryMemoMu.Unlock()
+ for k, e := range cnpgHistoryMemo {
+ if now.After(e.expires) {
+ delete(cnpgHistoryMemo, k)
+ }
+ }
+ cnpgHistoryMemo[key] = cnpgHistoryMemoEntry{expires: now.Add(cnpgHistoryMemoTTL), charts: charts}
+}
+
+// cnpgPrometheusUnavailable returns why Radar has no Prometheus to read, or
+// "" when it does.
+// Discovery's own failures are complete diagnoses, so they are not prefixed again.
+func cnpgNoPrometheusReason(msg string) string {
+ switch {
+ case msg == "":
+ return "Radar is not connected to Prometheus"
+ case strings.HasPrefix(msg, "Radar "), strings.HasPrefix(msg, "No working Prometheus endpoint found"):
+ return truncateCNPGRuntimeError(msg)
+ }
+ return "Radar is not connected to Prometheus: " + truncateCNPGRuntimeError(msg)
+}
+
+func historyScopeFailure(err error) (string, string) {
+ switch {
+ case errors.Is(err, prometheuspkg.ErrScopeAmbiguous):
+ return cnpgHistoryStateAmbiguous, "This Prometheus holds series for these Pod names under more than one cluster identity, so history could mix clusters. An operator can configure the cluster identity labels Radar should require."
+ case errors.Is(err, prometheuspkg.ErrScopeMismatch):
+ return cnpgHistoryStateScopeMismatch, "The cluster identity labels proven for this cluster do not appear on the CNPG exporter series, so Radar cannot tell this cluster's history from another's."
+ }
+ return historyStateError, "Prometheus query failed: " + truncateCNPGRuntimeError(err.Error())
+}
+
+// cnpgUsageScopeFailure is cnpgHistoryScopeFailure for kubelet volume stats.
+func usageScopeFailure(err error) (string, string) {
+ switch {
+ case errors.Is(err, prometheuspkg.ErrScopeAmbiguous):
+ return cnpgHistoryStateAmbiguous, "This Prometheus holds volume stats for these claim names under more than one cluster identity, so a value could be another cluster's. An operator can configure the cluster identity labels Radar should require."
+ case errors.Is(err, prometheuspkg.ErrScopeMismatch):
+ return cnpgHistoryStateScopeMismatch, "The cluster identity labels proven for this cluster do not appear on these claims' volume stats, so Radar cannot tell this cluster's volumes from another's."
+ }
+ return historyStateError, "Prometheus query failed: " + truncateCNPGRuntimeError(err.Error())
+}
+
+// cnpgHistoryAnchors are current instance Pods, name and UID, that prove
+// which cluster-identity labels are this cluster's.
+func historyAnchors(cache *k8s.ResourceCache, cluster *unstructured.Unstructured) []prom.WorkloadPodIdentity {
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ return nil
+ }
+ return cnpgPodIdentities(pods)
+}
+
+func cnpgPodIdentities(pods []*corev1.Pod) []prom.WorkloadPodIdentity {
+ out := make([]prom.WorkloadPodIdentity, 0, min(len(pods), cnpgHistoryAnchorCap))
+ for _, p := range pods {
+ if len(out) == cnpgHistoryAnchorCap {
+ break
+ }
+ out = append(out, prom.WorkloadPodIdentity{Name: p.Name, UID: string(p.UID)})
+ }
+ return out
+}
+
+// CNPGFleetMetricsResponse is GET /api/cnpg/fleet-metrics: per visible
+// Cluster, its largest current standby replay lag with its standbys' WAL
+// receivers, its inactive physical replication slots and the growth of its
+// fastest-growing volume, all from Prometheus.
+type CNPGFleetMetricsResponse struct {
+ SampledAt string `json:"sampledAt"`
+ Source string `json:"source"`
+ Reason string `json:"reason,omitempty"`
+ LagSource string `json:"lagSource"`
+ ReceiverSource string `json:"receiverSource"`
+ SlotsSource string `json:"slotsSource"`
+ GrowthSource string `json:"growthSource"`
+ Clusters []CNPGClusterFleetMetrics `json:"clusters"`
+}
+
+type CNPGClusterFleetMetrics struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ Lag CNPGFleetLag `json:"lag"`
+ Slots CNPGFleetSlots `json:"slots"`
+ Growth CNPGFleetGrowth `json:"growth"`
+}
+
+// CNPGFleetLag State: ok (Seconds is the largest standby lag), noStandby
+// (scraped, but no instance is in recovery), noSeries, denied, ambiguous,
+// scopeMismatch, error or notRead.
+type CNPGFleetLag struct {
+ State string `json:"state"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Seconds *float64 `json:"seconds,omitempty"`
+ Pod string `json:"pod,omitempty"`
+ // LagStandbys counts the standbys whose replay lag was read; Seconds
+ // covers only these.
+ LagStandbys *int `json:"lagStandbys,omitempty"`
+ // SustainedSeconds is the worst standby's lowest recorded lag over
+ // SustainedWindow, for a standby already reporting when it began.
+ SustainedSeconds *float64 `json:"sustainedSeconds,omitempty"`
+ SustainedPod string `json:"sustainedPod,omitempty"`
+ SustainedWindow string `json:"sustainedWindow,omitempty"`
+ // Set only with State ok. Standbys counts instances reporting
+ // cnpg_pg_replication_in_recovery = 1; Receiving those whose WAL receiver
+ // is up, ReceiverDown those whose is not. A standby whose receiver is down
+ // can still read lag 0, so receiving is never inferred from lag: when no
+ // standby reports the receiver, ReceiverUnknown is set instead, with
+ // ReceiverReason.
+ Standbys *int `json:"standbys,omitempty"`
+ Receiving *int `json:"receiving,omitempty"`
+ ReceiverDown []string `json:"receiverDown,omitempty"`
+ ReceiverUnknown bool `json:"receiverUnknown,omitempty"`
+ ReceiverReason string `json:"receiverReason,omitempty"`
+ // ReceiverDownSustained lists the standbys whose receiver was down in every
+ // sample over ReceiverDownWindow; only these raise a problem.
+ ReceiverDownSustained []string `json:"receiverDownSustained,omitempty"`
+ ReceiverDownWindow string `json:"receiverDownWindow,omitempty"`
+ // Isolation says how the series were tied to this cluster.
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+}
+
+// CNPGFleetSlots State: ok (Inactive lists the inactive physical slots,
+// empty when none is), noSeries, denied, ambiguous, scopeMismatch, error or
+// notRead. Inactive is null unless State is ok.
+type CNPGFleetSlots struct {
+ State string `json:"state"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Inactive []CNPGFleetSlot `json:"inactive"`
+ Omitted int `json:"omitted,omitempty"`
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+}
+
+// CNPGFleetSlot is one inactive physical slot as one instance reports it.
+// Role is that instance's (primary or standby), absent when it reports no
+// recovery state; CloudNativePG copies HA slots to standbys, where they are
+// never active. Bytes is the WAL the slot retains there, null when not
+// reported.
+type CNPGFleetSlot struct {
+ Slot string `json:"slot"`
+ Pod string `json:"pod"`
+ Role string `json:"role,omitempty"`
+ Bytes *float64 `json:"bytes"`
+}
+
+// CNPGFleetGrowth State: ok (BytesPerHour of the fastest-growing claim),
+// noSeries, denied, unavailable, error or notRead.
+type CNPGFleetGrowth struct {
+ State string `json:"state"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ BytesPerHour *float64 `json:"bytesPerHour,omitempty"`
+ Claim string `json:"claim,omitempty"`
+ Instance string `json:"instance,omitempty"`
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+}
+
+func (s *Reader) namespaceFleetMetrics(callerCtx context.Context, cache *k8s.ResourceCache, namespace string, clusters []*unstructured.Unstructured) []CNPGClusterFleetMetrics {
+ ctx := callerCtx
+ names := make([]string, len(clusters))
+ var anchors []prom.WorkloadPodIdentity
+ for i, c := range clusters {
+ names[i] = c.GetName()
+ if len(anchors) < cnpgHistoryAnchorCap {
+ anchors = append(anchors, historyAnchors(cache, c)...)
+ }
+ }
+ lags := make([]CNPGFleetLag, len(clusters))
+ slots := make([]CNPGFleetSlots, len(clusters))
+ growths := make([]CNPGFleetGrowth, len(clusters))
+ fillLag := func(l CNPGFleetLag) {
+ for i := range lags {
+ lags[i] = l
+ }
+ }
+ fillSlots := func(sl CNPGFleetSlots) {
+ for i := range slots {
+ slots[i] = sl
+ }
+ }
+ fillGrowth := func(g CNPGFleetGrowth) {
+ for i := range growths {
+ growths[i] = g
+ }
+ }
+
+ podsAllowed := s.Access.MetricsRead(callerCtx, "", "pods", namespace, "get")
+ claimsByCluster, growthCov := s.fleetClaims(callerCtx, cache, namespace, clusters)
+
+ matchers, scopeErr := "", error(nil)
+ var lagIso prometheuspkg.SeriesIsolation
+ if podsAllowed {
+ matchers, lagIso, scopeErr = s.Metrics.CNPGScope(ctx, namespace, prometheuspkg.CNPGInstancesSelector(namespace, names), anchors, prometheuspkg.CNPGSustainedLagWindow)
+ }
+
+ if !podsAllowed {
+ fillLag(CNPGFleetLag{State: cnpgHistoryStateDenied, Grant: grantGetPods.In(namespace).Ref()})
+ } else if scopeErr != nil {
+ state, reason := historyScopeFailure(scopeErr)
+ fillLag(CNPGFleetLag{State: state, Reason: reason})
+ } else if res, err := s.Metrics.FleetLag(ctx, namespace, names, matchers); err != nil {
+ fillLag(CNPGFleetLag{State: historyStateError, Reason: "Prometheus query failed: " + truncateCNPGRuntimeError(err.Error())})
+ } else {
+ for i, c := range clusters {
+ switch reading, ok := res.Lag[c.GetName()]; {
+ case ok:
+ v, n := reading.Seconds, reading.Reporting
+ lags[i] = CNPGFleetLag{State: historyStateOK, Seconds: &v, Pod: reading.Pod, LagStandbys: &n, Isolation: &lagIso}
+ if sus, ok := res.Sustained[c.GetName()]; ok {
+ sv := sus.Seconds
+ lags[i].SustainedSeconds, lags[i].SustainedPod, lags[i].SustainedWindow = &sv, sus.Pod, prometheuspkg.CNPGSustainedLagWindow.String()
+ }
+ cnpgFleetReceivers(&lags[i], res, c.GetName())
+ case res.Scraped[c.GetName()]:
+ lags[i] = CNPGFleetLag{State: "noStandby", Reason: "no instance reports being a standby"}
+ default:
+ lags[i] = CNPGFleetLag{State: cnpgUsageStateNoSeries, Reason: "Prometheus has no CNPG exporter series for this cluster's instances"}
+ }
+ }
+ }
+
+ if !podsAllowed {
+ fillSlots(CNPGFleetSlots{State: cnpgHistoryStateDenied, Grant: grantGetPods.In(namespace).Ref()})
+ } else if scopeErr != nil {
+ state, reason := historyScopeFailure(scopeErr)
+ fillSlots(CNPGFleetSlots{State: state, Reason: reason})
+ } else if res, err := s.Metrics.FleetSlots(ctx, namespace, names, matchers); err != nil {
+ fillSlots(CNPGFleetSlots{State: historyStateError, Reason: "Prometheus query failed: " + truncateCNPGRuntimeError(err.Error())})
+ } else {
+ for i, c := range clusters {
+ slots[i] = cnpgFleetSlotsOf(res, c.GetName(), &lagIso)
+ }
+ }
+
+ if growthCov.State != "" {
+ fillGrowth(growthCov)
+ } else {
+ var all []string
+ for _, cs := range claimsByCluster {
+ all = append(all, claimNames(cs)...)
+ }
+ pvcMatchers, pvcIso, err := s.Metrics.PVCScope(ctx, namespace, all, anchors, cnpgFleetGrowthWindow)
+ var byClaim map[string]float64
+ if err != nil {
+ state, reason := usageScopeFailure(err)
+ fillGrowth(CNPGFleetGrowth{State: state, Reason: reason})
+ } else if byClaim, err = s.Metrics.DiskGrowth(ctx, namespace, all, cnpgFleetGrowthWindow, pvcMatchers); err != nil {
+ fillGrowth(CNPGFleetGrowth{State: historyStateError, Reason: "Prometheus query failed: " + truncateCNPGRuntimeError(err.Error())})
+ }
+ for i, c := range clusters {
+ if err != nil {
+ break
+ }
+ owned := claimsByCluster[c.GetName()]
+ if len(owned) == 0 {
+ growths[i] = CNPGFleetGrowth{State: usageStateNotRead, Reason: "no claims owned by this cluster"}
+ continue
+ }
+ g := CNPGFleetGrowth{State: cnpgUsageStateNoSeries, Reason: "Prometheus has no kubelet volume stats for this cluster's claims over the last 6h"}
+ for _, pvc := range owned {
+ v, ok := byClaim[pvc.Name]
+ if !ok {
+ continue
+ }
+ if g.BytesPerHour == nil || v > *g.BytesPerHour {
+ val := v
+ g = CNPGFleetGrowth{State: historyStateOK, BytesPerHour: &val, Claim: pvc.Name, Instance: pvc.Labels[instanceNameLabel], Isolation: &pvcIso}
+ }
+ }
+ growths[i] = g
+ }
+ }
+ return cnpgFleetMetricsRows(namespace, clusters, lags, slots, growths, claimsByCluster)
+}
+
+// cnpgFleetReceivers adds a standby-reporting Cluster's WAL receiver evidence
+// to its lag reading.
+func cnpgFleetReceivers(l *CNPGFleetLag, res prometheuspkg.CNPGFleetLag, cluster string) {
+ rec, ok := res.Receivers[cluster]
+ switch {
+ case res.ReceiversError != "":
+ l.ReceiverUnknown, l.ReceiverReason = true, res.ReceiversError
+ case !ok:
+ l.ReceiverUnknown, l.ReceiverReason = true, "no instance reported being a standby when the WAL receivers were read"
+ case rec.Unknown:
+ standbys := rec.Standbys
+ l.Standbys = &standbys
+ l.ReceiverUnknown, l.ReceiverReason = true, "the exporter does not report cnpg_pg_replication_is_wal_receiver_up for this cluster's standbys (custom monitoring queries)"
+ default:
+ standbys, receiving := rec.Standbys, rec.Receiving
+ l.Standbys, l.Receiving, l.ReceiverDown = &standbys, &receiving, rec.Down
+ if res.ReceiversDownSustained != nil {
+ l.ReceiverDownSustained, l.ReceiverDownWindow = res.ReceiversDownSustained[cluster], prometheuspkg.CNPGReceiverDownWindow.String()
+ }
+ }
+}
+
+func cnpgFleetSlotsOf(res prometheuspkg.CNPGFleetSlots, cluster string, iso *prometheuspkg.SeriesIsolation) CNPGFleetSlots {
+ cs, ok := res.Clusters[cluster]
+ switch {
+ case ok:
+ out := CNPGFleetSlots{State: historyStateOK, Inactive: make([]CNPGFleetSlot, 0, len(cs.Inactive)), Omitted: cs.Omitted, Isolation: iso}
+ for _, s := range cs.Inactive {
+ out.Inactive = append(out.Inactive, CNPGFleetSlot{Slot: s.Slot, Pod: s.Pod, Role: s.Role, Bytes: s.Bytes})
+ }
+ return out
+ case res.Scraped[cluster]:
+ return CNPGFleetSlots{State: cnpgUsageStateNoSeries, Reason: "no instance of this cluster reports a replication slot: it has none, or its monitoring queries leave pg_replication_slots out"}
+ default:
+ return CNPGFleetSlots{State: cnpgUsageStateNoSeries, Reason: "Prometheus has no CNPG exporter series for this cluster's instances"}
+ }
+}
+
+// cnpgFleetClaims lists each Cluster's owned claims under the caller's grants.
+// A non-empty returned State says why growth is not read for the namespace.
+func (s *Reader) fleetClaims(ctx context.Context, cache *k8s.ResourceCache, namespace string, clusters []*unstructured.Unstructured) (map[string][]*corev1.PersistentVolumeClaim, CNPGFleetGrowth) {
+ if !s.Access.CanRead(ctx, "", "persistentvolumeclaims", namespace, "list") {
+ return nil, CNPGFleetGrowth{State: storageStateDenied, Grant: cnpgGrantListPVCs.In(namespace).Ref()}
+ }
+ if !s.Access.MetricsRead(ctx, "", "persistentvolumeclaims", namespace, "get") {
+ return nil, CNPGFleetGrowth{State: storageStateDenied, Grant: grantGetPVCs.In(namespace).Ref()}
+ }
+ req, err := labels.NewRequirement(clusterLabel, selection.Exists, nil)
+ if err != nil {
+ return nil, CNPGFleetGrowth{State: cnpgStorageStateError, Reason: err.Error()}
+ }
+ candidates, reason := cnpgCachedClaims(cache, namespace, labels.NewSelector().Add(*req))
+ if reason != "" {
+ return nil, CNPGFleetGrowth{State: cnpgStorageStateUnavailable, Reason: reason}
+ }
+ out := map[string][]*corev1.PersistentVolumeClaim{}
+ for _, c := range clusters {
+ var mine []*corev1.PersistentVolumeClaim
+ for _, pvc := range candidates {
+ if pvc.Labels[clusterLabel] == c.GetName() {
+ mine = append(mine, pvc)
+ }
+ }
+ owned, _ := cnpgOwnedClaims(mine, c)
+ if len(owned) > 0 {
+ out[c.GetName()] = owned
+ }
+ }
+ return out, CNPGFleetGrowth{}
+}
+
+func cnpgFleetMetricsRows(namespace string, clusters []*unstructured.Unstructured, lags []CNPGFleetLag, slots []CNPGFleetSlots, growths []CNPGFleetGrowth, claimsByCluster map[string][]*corev1.PersistentVolumeClaim) []CNPGClusterFleetMetrics {
+ out := make([]CNPGClusterFleetMetrics, len(clusters))
+ for i, c := range clusters {
+ createdAt := c.GetCreationTimestamp().Time
+ if !createdAt.IsZero() {
+ age := time.Since(createdAt)
+ if age < prometheuspkg.CNPGMetricLookback {
+ if lags[i].State == historyStateOK || lags[i].State == "noStandby" {
+ lags[i] = CNPGFleetLag{State: usageStateNotRead, Reason: "Waiting for exporter samples after this Cluster was created"}
+ }
+ if slots[i].State == historyStateOK {
+ slots[i] = CNPGFleetSlots{State: usageStateNotRead, Reason: "Waiting for exporter samples after this Cluster was created"}
+ }
+ }
+ if age < prometheuspkg.CNPGSustainedLagWindow+prometheuspkg.CNPGMetricLookback {
+ lags[i].SustainedSeconds, lags[i].SustainedPod, lags[i].SustainedWindow = nil, "", ""
+ }
+ if age < prometheuspkg.CNPGReceiverDownWindow+prometheuspkg.CNPGMetricLookback {
+ lags[i].ReceiverDownSustained, lags[i].ReceiverDownWindow = nil, ""
+ }
+ if growths[i].State == historyStateOK && age < cnpgFleetGrowthWindow {
+ growths[i] = CNPGFleetGrowth{State: usageStateNotRead, Reason: "Waiting for a full volume-growth window after this Cluster was created"}
+ }
+ }
+ if growths[i].State == historyStateOK {
+ for _, pvc := range claimsByCluster[c.GetName()] {
+ if !pvc.CreationTimestamp.IsZero() && time.Since(pvc.CreationTimestamp.Time) < cnpgFleetGrowthWindow {
+ growths[i] = CNPGFleetGrowth{State: usageStateNotRead, Reason: "Waiting for a full volume-growth window after this Cluster's PVCs were created"}
+ break
+ }
+ }
+ }
+ out[i] = CNPGClusterFleetMetrics{Namespace: namespace, Name: c.GetName(), Lag: lags[i], Slots: slots[i], Growth: growths[i]}
+ }
+ return out
+}
+
+func (s *Reader) ClusterHistory(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured, rng prometheuspkg.CNPGHistoryRange) CNPGClusterHistoryResponse {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ now := time.Now()
+ resp := CNPGClusterHistoryResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ Range: rng.Name,
+ SampledAt: now.UTC().Format(time.RFC3339),
+ Charts: []prometheuspkg.CNPGHistoryChart{},
+ }
+ if reason := s.prometheusUnavailable(ctx); reason != "" {
+ resp.Source, resp.Reason = historySourceNone, reason
+ return resp
+ }
+ resp.Source = historySourcePrometheus
+ start, end := rng.Bounds(now, cluster.GetCreationTimestamp().Time)
+ resp.Start, resp.End = start.UTC().Format(time.RFC3339), end.UTC().Format(time.RFC3339)
+ if start.After(end) {
+ resp.State, resp.Reason = "pending", "This Cluster is too new for isolated history. Waiting for samples after its creation and the query lookback window."
+ return resp
+ }
+ if start.After(end.Add(-rng.Duration)) {
+ resp.Reason = "History is limited to this Cluster incarnation; the first query lookback window after creation is omitted to exclude samples from reused Pod names."
+ }
+ resp.StepSeconds = int(rng.Step / time.Second)
+ selector := prometheuspkg.CNPGInstanceSelector(namespace, name)
+ resp.Selector = selector
+
+ req := prometheuspkg.CNPGHistoryRequest{Namespace: namespace, Cluster: name, Range: rng, End: end, CreatedAt: cluster.GetCreationTimestamp().Time}
+ if !s.Access.MetricsRead(ctx, "", "pods", namespace, "get") {
+ req.PodsDenied = grantGetPods.In(namespace).Ref()
+ }
+ claims, _, claimCov := s.clusterClaims(ctx, cache, cluster)
+ switch {
+ case claimCov.State == storageStateDenied:
+ req.PVCDenied = claimCov.Grant
+ case claimCov.State != storageStateOK:
+ req.PVCReason = claimCov.Reason
+ case !s.Access.MetricsRead(ctx, "", "persistentvolumeclaims", namespace, "get"):
+ req.PVCDenied = grantGetPVCs.In(namespace).Ref()
+ default:
+ req.Claims = claimNames(claims)
+ }
+
+ anchors := historyAnchors(cache, cluster)
+ if req.PodsDenied == nil {
+ matchers, iso, err := s.Metrics.CNPGScope(ctx, namespace, selector, anchors, rng.Duration)
+ if err != nil {
+ resp.State, resp.Reason = historyScopeFailure(err)
+ return resp
+ }
+ req.Matchers, resp.Isolation = matchers, &iso
+ }
+ if len(req.Claims) > 0 {
+ matchers, iso, err := s.Metrics.PVCScope(ctx, namespace, req.Claims, anchors, rng.Duration)
+ if err != nil {
+ req.PVCScopeState, req.PVCReason = usageScopeFailure(err)
+ } else {
+ req.PVCMatchers, resp.PVCIsolation = matchers, &iso
+ }
+ }
+
+ key := strings.Join([]string{namespace, name, string(cluster.GetUID()), rng.Name, end.Format(time.RFC3339), req.Matchers, req.PVCMatchers, req.PVCScopeState, integration.GrantText(req.PodsDenied), integration.GrantText(req.PVCDenied), req.PVCReason, strings.Join(req.Claims, ",")}, "\x00")
+ charts, hit := historyMemoGet(key, now)
+ if !hit {
+ var err error
+ charts, err = s.Metrics.History(ctx, req)
+ if err != nil {
+ resp.State, resp.Reason = historyStateError, err.Error()
+ return resp
+ }
+ historyMemoPut(key, now, charts)
+ }
+ resp.State, resp.Charts = historyStateOK, charts
+ return resp
+
+}
+
+func (s *Reader) FleetMetrics(ctx context.Context, cache *k8s.ResourceCache, namespaces []string) CNPGFleetMetricsResponse {
+ resp := CNPGFleetMetricsResponse{
+ SampledAt: time.Now().UTC().Format(time.RFC3339), Source: historySourcePrometheus,
+ LagSource: fleetLagSource, ReceiverSource: fleetReceiverSource, SlotsSource: fleetSlotsSource,
+ GrowthSource: fleetGrowthSource, Clusters: []CNPGClusterFleetMetrics{},
+ }
+ if reason := s.prometheusUnavailable(ctx); reason != "" {
+ resp.Source, resp.Reason = historySourceNone, reason
+ return resp
+ }
+
+ var clusterKind integration.WorkspaceKind
+ for _, k := range workspaceKinds {
+ if k.Key == workspaceClusterKey {
+ clusterKind = k
+ }
+ }
+ acc, clusters := s.workspaceReadKind(ctx, cache, clusterKind, namespaces)
+ if acc.State != integration.KindCoverageFull && acc.State != integration.KindCoveragePartial {
+ return resp
+ }
+ byNamespace := map[string][]*unstructured.Unstructured{}
+ for _, c := range clusters {
+ byNamespace[c.GetNamespace()] = append(byNamespace[c.GetNamespace()], c)
+ }
+ nsList := make([]string, 0, len(byNamespace))
+ for ns := range byNamespace {
+ nsList = append(nsList, ns)
+ }
+ sort.Strings(nsList)
+
+ read := min(len(nsList), fleetDiskMaxNamespaces)
+ results := integration.FanOut(ctx, read, fleetDiskConcurrency, func(i int) []CNPGClusterFleetMetrics {
+ return s.namespaceFleetMetrics(ctx, cache, nsList[i], byNamespace[nsList[i]])
+ })
+ for _, rs := range results {
+ resp.Clusters = append(resp.Clusters, rs...)
+ }
+ reason := fmt.Sprintf("read for at most %d namespaces at a time; narrow the namespace filter", fleetDiskMaxNamespaces)
+ for _, ns := range nsList[read:] {
+ for _, c := range byNamespace[ns] {
+ resp.Clusters = append(resp.Clusters, CNPGClusterFleetMetrics{Namespace: ns, Name: c.GetName(),
+ Lag: CNPGFleetLag{State: usageStateNotRead, Reason: reason}, Slots: CNPGFleetSlots{State: usageStateNotRead, Reason: reason},
+ Growth: CNPGFleetGrowth{State: usageStateNotRead, Reason: reason}})
+ }
+ }
+ sort.Slice(resp.Clusters, func(i, j int) bool {
+ if resp.Clusters[i].Namespace != resp.Clusters[j].Namespace {
+ return resp.Clusters[i].Namespace < resp.Clusters[j].Namespace
+ }
+ return resp.Clusters[i].Name < resp.Clusters[j].Name
+ })
+ return resp
+
+}
diff --git a/internal/cnpg/history_incarnation_test.go b/internal/cnpg/history_incarnation_test.go
new file mode 100644
index 0000000000..c4521d2641
--- /dev/null
+++ b/internal/cnpg/history_incarnation_test.go
@@ -0,0 +1,57 @@
+package cnpg
+
+import (
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+)
+
+func TestFleetMetricsExcludePredecessorLookbackWindows(t *testing.T) {
+ lag := 30.0
+ growth := 1024.0
+ cases := []struct {
+ name string
+ age, pvcAge time.Duration
+ lagState string
+ sustained bool
+ growthState string
+ }{
+ {"new cluster", time.Minute, 0, usageStateNotRead, false, usageStateNotRead},
+ {"current gauges but incomplete lag window", 7 * time.Minute, 0, historyStateOK, false, usageStateNotRead},
+ {"offset selector still reaches predecessor lookback", 12 * time.Minute, 0, historyStateOK, false, usageStateNotRead},
+ {"current lag window but incomplete growth", time.Hour, 0, historyStateOK, true, usageStateNotRead},
+ {"all current", 24 * time.Hour, 24 * time.Hour, historyStateOK, true, historyStateOK},
+ {"recreated PVC", 24 * time.Hour, time.Hour, historyStateOK, true, usageStateNotRead},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ c := &unstructured.Unstructured{Object: map[string]any{"metadata": map[string]any{"name": "main"}}}
+ c.SetCreationTimestamp(metav1.NewTime(time.Now().Add(-tc.age)))
+ claims := map[string][]*corev1.PersistentVolumeClaim{}
+ if tc.pvcAge > 0 {
+ claims["main"] = []*corev1.PersistentVolumeClaim{{ObjectMeta: metav1.ObjectMeta{CreationTimestamp: metav1.NewTime(time.Now().Add(-tc.pvcAge))}}}
+ }
+ rows := cnpgFleetMetricsRows("pg", []*unstructured.Unstructured{c}, []CNPGFleetLag{{State: historyStateOK, Seconds: &lag, SustainedSeconds: &lag, ReceiverDownSustained: []string{"main-2"}}}, []CNPGFleetSlots{{State: historyStateOK}}, []CNPGFleetGrowth{{State: historyStateOK, BytesPerHour: &growth}}, claims)
+ row := rows[0]
+ if row.Lag.State != tc.lagState || (row.Lag.SustainedSeconds != nil) != tc.sustained || row.Growth.State != tc.growthState {
+ t.Fatalf("metrics = %+v", row)
+ }
+ if tc.age < 5*time.Minute && (row.Slots.State != usageStateNotRead || len(row.Lag.ReceiverDownSustained) != 0) {
+ t.Fatalf("predecessor gauges survived: %+v", row)
+ }
+ })
+ }
+}
+
+func TestFleetMetricsRetainDeniedAndFailedReads(t *testing.T) {
+ c := &unstructured.Unstructured{Object: map[string]any{"metadata": map[string]any{"name": "main"}}}
+ c.SetCreationTimestamp(metav1.NewTime(time.Now()))
+ grant := grantGetPods.In("pg").Ref()
+ row := cnpgFleetMetricsRows("pg", []*unstructured.Unstructured{c}, []CNPGFleetLag{{State: cnpgHistoryStateDenied, Grant: grant}}, []CNPGFleetSlots{{State: historyStateError, Reason: "query failed"}}, []CNPGFleetGrowth{{State: historyStateError, Reason: "query failed"}}, nil)[0]
+ if row.Lag.State != cnpgHistoryStateDenied || row.Lag.Grant != grant || row.Slots.Reason != "query failed" || row.Growth.Reason != "query failed" {
+ t.Fatalf("coverage was replaced by waiting: %+v", row)
+ }
+}
diff --git a/internal/cnpg/history_scope_test.go b/internal/cnpg/history_scope_test.go
new file mode 100644
index 0000000000..dcce870eef
--- /dev/null
+++ b/internal/cnpg/history_scope_test.go
@@ -0,0 +1,62 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "strings"
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+ "github.com/skyhook-io/radar/pkg/k8score"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+func TestClusterHistoryPreservesPVCIsolationFailures(t *testing.T) {
+ for _, tc := range []struct {
+ name, state, reason string
+ err error
+ }{
+ {"ambiguous", cnpgHistoryStateAmbiguous, "more than one cluster identity", prometheuspkg.ErrScopeAmbiguous},
+ {"mismatch", cnpgHistoryStateScopeMismatch, "do not appear", prometheuspkg.ErrScopeMismatch},
+ {"query error", historyStateError, "Prometheus query failed", errors.New("connection refused")},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ cluster := cnpgActionCluster(nil)
+ pod := cnpgActionPod("pg-1", "pod-uid", true)
+ pvc := &corev1.PersistentVolumeClaim{ObjectMeta: metav1.ObjectMeta{Name: "pg-1", Namespace: "db", Labels: map[string]string{clusterLabel: "pg", instanceNameLabel: "pg-1"}, OwnerReferences: pod.OwnerReferences}}
+ core, err := k8score.NewResourceCache(k8score.CacheConfig{Client: k8sfake.NewSimpleClientset(pod, pvc), ResourceTypes: map[string]bool{"pods": true, "persistentvolumeclaims": true}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer core.Stop()
+ reader := newTestReader(nil)
+ reader.Access.MetricsRead = func(context.Context, string, string, string, string) bool { return true }
+ reader.Metrics.Connection = func(context.Context) (bool, error) { return true, nil }
+ reader.Metrics.CNPGScope = func(context.Context, string, string, []prom.WorkloadPodIdentity, time.Duration) (string, prometheuspkg.SeriesIsolation, error) {
+ return "", prometheuspkg.SeriesIsolation{}, nil
+ }
+ reader.Metrics.PVCScope = func(context.Context, string, []string, []prom.WorkloadPodIdentity, time.Duration) (string, prometheuspkg.SeriesIsolation, error) {
+ return "", prometheuspkg.SeriesIsolation{}, tc.err
+ }
+ called := false
+ reader.Metrics.History = func(_ context.Context, req prometheuspkg.CNPGHistoryRequest) ([]prometheuspkg.CNPGHistoryChart, error) {
+ called = true
+ if req.PVCScopeState != tc.state || !strings.Contains(req.PVCReason, tc.reason) {
+ t.Fatalf("history request lost failure: %+v", req)
+ }
+ return nil, nil
+ }
+ rng, _ := prometheuspkg.ParseCNPGHistoryRange("1h")
+ reader.ClusterHistory(context.Background(), &k8s.ResourceCache{ResourceCache: core}, cluster, rng)
+ if !called {
+ t.Fatal("history reader did not receive the PVC scope result")
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/inspect.go b/internal/cnpg/inspect.go
new file mode 100644
index 0000000000..d30e3551ff
--- /dev/null
+++ b/internal/cnpg/inspect.go
@@ -0,0 +1,528 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "fmt"
+ "log"
+ "net/http"
+ "regexp"
+ "sort"
+ "strconv"
+ "strings"
+ "sync"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+const (
+ cnpgRestoreListCap = 50
+ cnpgHistoryReadBytes = 16 << 10
+ cnpgParametersCap = 200
+ cnpgDefaultAppDatabase = "app"
+ cnpgSessionSourceSession = "session"
+ cnpgSessionSourceClient = "client"
+)
+
+// A database name goes to psql's -d, which reads a string containing '=' or a
+// URI prefix as a connection string; only plain names are used.
+var cnpgPlainDatabaseName = regexp.MustCompile(`^[A-Za-z_][A-Za-z0-9_$-]{0,62}$`)
+
+// PostgreSQL parameter names, including extension ones ("pg_stat_statements.max").
+var cnpgParameterName = regexp.MustCompile(`^[A-Za-z_][A-Za-z0-9_.]{0,62}$`)
+
+// CNPGRestoreChecksResponse is GET /api/cnpg/clusters/{ns}/{name}/restore-checks:
+// what Radar read in the primary of a Cluster bootstrapped from a backup.
+// These are observations, not a verdict on the restore: where recovery
+// stopped is shown beside the declared target, never matched against it, and
+// a database or role existing does not prove what it holds (CloudNativePG
+// creates the bootstrap database and owner when they are missing).
+type CNPGRestoreChecksResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ Pod string `json:"pod"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGExecPermission `json:"permission"`
+ // Target is spec.bootstrap.recovery.recoveryTarget as declared.
+ Target map[string]any `json:"target,omitempty"`
+ Database string `json:"database"`
+ CNPGRuntimeSource
+ *CNPGRestoreFacts
+ // Contents describes Database; absent when it could not be read, with
+ // ContentsSource saying why.
+ Contents *CNPGDatabaseContents `json:"contents,omitempty"`
+ ContentsSource *CNPGRuntimeSource `json:"contentsSource,omitempty"`
+}
+
+type CNPGRestoreFacts struct {
+ InRecovery bool `json:"inRecovery"`
+ Timeline int64 `json:"timeline"`
+ // History is the current timeline's history file, oldest switch first.
+ // HistoryMissing is set when there is none (timeline 1 never switched).
+ History []CNPGTimelineSwitch `json:"history"`
+ HistoryMissing bool `json:"historyMissing,omitempty"`
+ DatabaseCount int `json:"databaseCount"`
+ Databases []CNPGDatabaseSize `json:"databases"`
+ RoleCount int `json:"roleCount"`
+ Roles []CNPGRoleFact `json:"roles"`
+}
+
+// CNPGTimelineSwitch is one history line: timeline From ended at SwitchLSN and
+// To began. Reason is PostgreSQL's own text — for a recovery that stopped at a
+// target, "before|after ", "at
+// restore point …" or "before|after transaction …"; a promotion without a
+// target (end of WAL, failover, switchover) reads "no recovery target specified".
+type CNPGTimelineSwitch struct {
+ From int64 `json:"from"`
+ To int64 `json:"to"`
+ SwitchLSN string `json:"switchLsn"`
+ Reason string `json:"reason"`
+}
+
+type CNPGDatabaseSize struct {
+ Name string `json:"name"`
+ Bytes int64 `json:"bytes"`
+}
+
+type CNPGRoleFact struct {
+ Name string `json:"name"`
+ CanLogin bool `json:"canLogin"`
+}
+
+// CNPGDatabaseContents: tables outside the system schemas. Row counts are the
+// planner's estimates (pg_class.reltuples), which can be stale. A table
+// without one is counted in NoEstimate, not as zero rows: from PostgreSQL 14
+// that is -1 (never vacuumed or analyzed); before 14 it is 0, which an empty
+// analyzed table also reads, so those count as unknown too.
+type CNPGDatabaseContents struct {
+ Tables int `json:"tables"`
+ EstimatedRows int64 `json:"estimatedRows"`
+ NoEstimate int `json:"noEstimate"`
+ Largest []CNPGTableSummary `json:"largest"`
+}
+
+type CNPGTableSummary struct {
+ Name string `json:"name"`
+ EstimatedRows *int64 `json:"estimatedRows,omitempty"`
+ Bytes int64 `json:"bytes"`
+}
+
+var cnpgRestoreFactsSQL = cnpgSQLPrelude + cnpgSQLReadOnly + `SET application_name = '` + cnpgDiagnosticsApp + `';
+WITH tl AS (SELECT timeline_id AS id FROM pg_control_checkpoint())
+SELECT json_build_object(
+ 'inRecovery', pg_is_in_recovery(),
+ 'timeline', (SELECT id FROM tl),
+ 'history', (SELECT pg_read_file('pg_wal/' || lpad(upper(to_hex(id)), 8, '0') || '.history', 0, ` + strconv.Itoa(cnpgHistoryReadBytes) + `, true) FROM tl),
+ 'databaseCount', (SELECT count(*) FROM pg_database WHERE datallowconn AND NOT datistemplate),
+ 'databases', coalesce((
+ SELECT json_agg(d) FROM (
+ SELECT datname AS name, pg_database_size(oid) AS bytes
+ FROM pg_database WHERE datallowconn AND NOT datistemplate
+ ORDER BY datname LIMIT ` + strconv.Itoa(cnpgRestoreListCap) + `
+ ) d
+ ), '[]'::json),
+ 'roleCount', (SELECT count(*) FROM pg_roles WHERE rolname !~ '^pg_'),
+ 'roles', coalesce((
+ SELECT json_agg(r) FROM (
+ SELECT rolname AS name, rolcanlogin AS "canLogin"
+ FROM pg_roles WHERE rolname !~ '^pg_'
+ ORDER BY rolname LIMIT ` + strconv.Itoa(cnpgRestoreListCap) + `
+ ) r
+ ), '[]'::json)
+);
+`
+
+var cnpgDatabaseContentsSQL = cnpgSQLPrelude + cnpgSQLReadOnly + `SET application_name = '` + cnpgDiagnosticsApp + `';
+WITH t AS (
+ SELECT c.oid, n.nspname, c.relname, c.relkind, c.relispartition, c.reltuples,
+ CASE WHEN current_setting('server_version_num')::int >= 140000 THEN c.reltuples < 0 ELSE c.reltuples <= 0 END AS unknown
+ FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
+ WHERE c.relkind IN ('r', 'p')
+ AND n.nspname NOT IN ('pg_catalog', 'information_schema')
+ AND n.nspname !~ '^pg_toast'
+)
+SELECT json_build_object(
+ 'tables', (SELECT count(*) FROM t WHERE NOT relispartition),
+ 'estimatedRows', (SELECT coalesce(sum(reltuples), 0)::bigint FROM t WHERE relkind = 'r' AND NOT unknown),
+ 'noEstimate', (SELECT count(*) FROM t WHERE relkind = 'r' AND unknown),
+ 'largest', coalesce((
+ SELECT json_agg(x) FROM (
+ SELECT nspname || '.' || relname AS name,
+ CASE WHEN NOT unknown THEN reltuples::bigint END AS "estimatedRows",
+ pg_total_relation_size(oid) AS bytes
+ FROM t WHERE relkind = 'r'
+ ORDER BY pg_total_relation_size(oid) DESC, nspname, relname
+ LIMIT 5
+ ) x
+ ), '[]'::json)
+);
+`
+
+func cnpgPsqlArgvFor(database string) []string {
+ return []string{"psql", "-XAtq", "-v", "ON_ERROR_STOP=1", "-d", database, "-f", "-"}
+}
+
+type cnpgRestoreFactsRaw struct {
+ InRecovery bool `json:"inRecovery"`
+ Timeline int64 `json:"timeline"`
+ History *string `json:"history"`
+ DatabaseCount int `json:"databaseCount"`
+ Databases []CNPGDatabaseSize `json:"databases"`
+ RoleCount int `json:"roleCount"`
+ Roles []CNPGRoleFact `json:"roles"`
+}
+
+func parseCNPGRestoreFacts(out []byte) (*CNPGRestoreFacts, error) {
+ var raw cnpgRestoreFactsRaw
+ if err := json.Unmarshal([]byte(strings.TrimSpace(string(out))), &raw); err != nil {
+ return nil, fmt.Errorf("unexpected psql output: %w", err)
+ }
+ f := &CNPGRestoreFacts{
+ InRecovery: raw.InRecovery,
+ Timeline: raw.Timeline,
+ DatabaseCount: raw.DatabaseCount,
+ Databases: raw.Databases,
+ RoleCount: raw.RoleCount,
+ Roles: raw.Roles,
+ History: []CNPGTimelineSwitch{},
+ }
+ if raw.History == nil {
+ f.HistoryMissing = true
+ } else {
+ f.History = parseCNPGTimelineHistory(*raw.History, raw.Timeline)
+ }
+ if f.Databases == nil {
+ f.Databases = []CNPGDatabaseSize{}
+ }
+ if f.Roles == nil {
+ f.Roles = []CNPGRoleFact{}
+ }
+ return f, nil
+}
+
+// parseCNPGTimelineHistory reads a timeline history file: one line per switch,
+// "\t\t", oldest first; '#' lines are comments.
+// Each line's child is the next line's parent, and the last line's is current.
+func parseCNPGTimelineHistory(text string, current int64) []CNPGTimelineSwitch {
+ out := []CNPGTimelineSwitch{}
+ for _, line := range strings.Split(text, "\n") {
+ line = strings.TrimSpace(line)
+ if line == "" || strings.HasPrefix(line, "#") {
+ continue
+ }
+ parts := strings.SplitN(line, "\t", 3)
+ if len(parts) < 2 {
+ continue
+ }
+ from, err := strconv.ParseInt(strings.TrimSpace(parts[0]), 10, 64)
+ if err != nil {
+ continue
+ }
+ sw := CNPGTimelineSwitch{From: from, SwitchLSN: strings.TrimSpace(parts[1])}
+ if len(parts) == 3 {
+ sw.Reason = strings.TrimSpace(parts[2])
+ }
+ out = append(out, sw)
+ }
+ for i := range out {
+ if i+1 < len(out) {
+ out[i].To = out[i+1].From
+ } else {
+ out[i].To = current
+ }
+ }
+ return out
+}
+
+func parseCNPGDatabaseContents(out []byte) (*CNPGDatabaseContents, error) {
+ var c CNPGDatabaseContents
+ if err := json.Unmarshal([]byte(strings.TrimSpace(string(out))), &c); err != nil {
+ return nil, fmt.Errorf("unexpected psql output: %w", err)
+ }
+ if c.Largest == nil {
+ c.Largest = []CNPGTableSummary{}
+ }
+ return &c, nil
+}
+
+// cnpgRestoreDatabase is the database applications use after the restore:
+// spec.bootstrap.recovery.database, which CloudNativePG defaults to "app".
+func restoreDatabase(cluster *unstructured.Unstructured) string {
+ if db, _, _ := unstructured.NestedString(cluster.Object, "spec", "bootstrap", "recovery", "database"); db != "" {
+ return db
+ }
+ return cnpgDefaultAppDatabase
+}
+
+func readCNPGRestoreChecks(ctx context.Context, exec ExecFunc, namespace, pod string, resp *CNPGRestoreChecksResponse) {
+ out, err := exec(ctx, namespace, pod, defaultLogContainer, cnpgPsqlArgv, cnpgRestoreFactsSQL)
+ captured := time.Now().UTC().Format(time.RFC3339)
+ if err != nil {
+ resp.CNPGRuntimeSource = cnpgExecSourceState(err)
+ resp.CapturedAt = captured
+ return
+ }
+ facts, err := parseCNPGRestoreFacts(out)
+ if err != nil {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateError, Error: err.Error(), CapturedAt: captured}
+ return
+ }
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateOK, CapturedAt: captured}
+ resp.CNPGRestoreFacts = facts
+
+ contents := func(src CNPGRuntimeSource) { resp.ContentsSource = &src }
+ if !cnpgPlainDatabaseName.MatchString(resp.Database) {
+ contents(CNPGRuntimeSource{State: runtimeStateError, Error: fmt.Sprintf("%q is not a plain database name, so Radar does not connect to it", resp.Database)})
+ return
+ }
+ if !cnpgHasDatabase(facts, resp.Database) {
+ contents(CNPGRuntimeSource{State: runtimeStateError, Error: fmt.Sprintf("database %q is not in the list read from the primary", resp.Database)})
+ return
+ }
+ out, err = exec(ctx, namespace, pod, defaultLogContainer, cnpgPsqlArgvFor(resp.Database), cnpgDatabaseContentsSQL)
+ captured = time.Now().UTC().Format(time.RFC3339)
+ if err != nil {
+ src := cnpgExecSourceState(err)
+ src.CapturedAt = captured
+ contents(src)
+ return
+ }
+ c, err := parseCNPGDatabaseContents(out)
+ if err != nil {
+ contents(CNPGRuntimeSource{State: runtimeStateError, Error: err.Error(), CapturedAt: captured})
+ return
+ }
+ resp.Contents = c
+ contents(CNPGRuntimeSource{State: runtimeStateOK, CapturedAt: captured})
+}
+
+// cnpgHasDatabase is false only when the list is complete and lacks it; a
+// capped list may simply not show it.
+func cnpgHasDatabase(f *CNPGRestoreFacts, db string) bool {
+ for _, d := range f.Databases {
+ if d.Name == db {
+ return true
+ }
+ }
+ return f.DatabaseCount > len(f.Databases)
+}
+
+// CNPGParametersResponse is GET /api/cnpg/clusters/{ns}/{name}/parameters: for
+// each name in spec.postgresql.parameters, what each instance's PostgreSQL
+// reports in a fresh session of Radar's — the server's value, which a
+// client's own session, role or database settings can still override.
+type CNPGParametersResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGExecPermission `json:"permission"`
+ Declared []CNPGParameterValue `json:"declared"`
+ Omitted int `json:"omitted,omitempty"`
+ Skipped []string `json:"skipped,omitempty"`
+ Instances []CNPGInstanceSettings `json:"instances"`
+ CNPGRuntimeSource
+}
+
+type CNPGParameterValue struct {
+ Name string `json:"name"`
+ Value string `json:"value"`
+}
+
+type CNPGInstanceSettings struct {
+ Pod string `json:"pod"`
+ Role string `json:"role"`
+ CNPGRuntimeSource
+ // Settings is null when the read failed; [] is a read that matched nothing.
+ Settings []CNPGParameterSetting `json:"settings"`
+}
+
+// CNPGParameterSetting is one pg_settings row. Value is PostgreSQL's display
+// (current_setting, units normalized, so "1024MB" reads "1GB"). A parameter
+// the connection itself sets (psql sends application_name) has no Value:
+// the server's own value is not visible from inside that session.
+type CNPGParameterSetting struct {
+ Name string `json:"name"`
+ Value *string `json:"value"`
+ SetByClient bool `json:"setByClient,omitempty"`
+ Source string `json:"source"`
+ Context string `json:"context"`
+ PendingRestart bool `json:"pendingRestart"`
+}
+
+// Only search_path is set (see cnpgSQLPrelude): any other SET would hide the
+// server's value of the very parameter it sets, so a declared search_path is
+// the one parameter this read reports as set by its own connection. The exec
+// deadline bounds the read.
+var cnpgParametersSQL = `SET search_path = pg_catalog;
+SELECT coalesce(json_agg(json_build_object(
+ 'name', s.name,
+ 'value', CASE WHEN s.source IN ('` + cnpgSessionSourceSession + `', '` + cnpgSessionSourceClient + `') THEN NULL ELSE current_setting(s.name) END,
+ 'setByClient', s.source IN ('` + cnpgSessionSourceSession + `', '` + cnpgSessionSourceClient + `'),
+ 'source', s.source,
+ 'context', s.context,
+ 'pendingRestart', s.pending_restart
+) ORDER BY s.name), '[]'::json)
+FROM pg_settings s
+WHERE s.name = ANY (string_to_array(:'names', ','));
+`
+
+// cnpgDeclaredParameters returns spec.postgresql.parameters sorted, capped,
+// and split into names Radar can ask about and names it skips.
+func declaredParameters(cluster *unstructured.Unstructured) (declared []CNPGParameterValue, query []string, skipped []string, omitted int) {
+ params, _, _ := unstructured.NestedMap(cluster.Object, "spec", "postgresql", "parameters")
+ names := make([]string, 0, len(params))
+ for n := range params {
+ names = append(names, n)
+ }
+ sort.Strings(names)
+ if len(names) > cnpgParametersCap {
+ omitted = len(names) - cnpgParametersCap
+ names = names[:cnpgParametersCap]
+ }
+ declared = make([]CNPGParameterValue, 0, len(names))
+ for _, n := range names {
+ declared = append(declared, CNPGParameterValue{Name: n, Value: fmt.Sprint(params[n])})
+ if cnpgParameterName.MatchString(n) {
+ // pg_settings names are lower case; PostgreSQL matches names case-insensitively.
+ query = append(query, strings.ToLower(n))
+ } else {
+ skipped = append(skipped, n)
+ }
+ }
+ return declared, query, skipped, omitted
+}
+
+func cnpgParametersArgv(names []string) []string {
+ return []string{"psql", "-XAtq", "-v", "ON_ERROR_STOP=1", "-v", "names=" + strings.Join(names, ","), "-d", cnpgPsqlDatabase, "-f", "-"}
+}
+
+func parseCNPGParameterSettings(out []byte) ([]CNPGParameterSetting, error) {
+ var rows []CNPGParameterSetting
+ if err := json.Unmarshal([]byte(strings.TrimSpace(string(out))), &rows); err != nil {
+ return nil, fmt.Errorf("unexpected psql output: %w", err)
+ }
+ if rows == nil {
+ rows = []CNPGParameterSetting{}
+ }
+ return rows, nil
+}
+
+func readCNPGParameters(ctx context.Context, exec ExecFunc, namespace string, pods []*corev1.Pod, names []string) []CNPGInstanceSettings {
+ out := make([]CNPGInstanceSettings, len(pods))
+ var wg sync.WaitGroup
+ for i, p := range pods {
+ out[i] = CNPGInstanceSettings{Pod: p.Name, Role: runtimeRole(p)}
+ wg.Add(1)
+ go func(i int, pod string) {
+ defer wg.Done()
+ stdout, err := exec(ctx, namespace, pod, defaultLogContainer, cnpgParametersArgv(names), cnpgParametersSQL)
+ captured := time.Now().UTC().Format(time.RFC3339)
+ if err != nil {
+ src := cnpgExecSourceState(err)
+ src.CapturedAt = captured
+ out[i].CNPGRuntimeSource = src
+ return
+ }
+ rows, err := parseCNPGParameterSettings(stdout)
+ if err != nil {
+ out[i].CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateError, Error: err.Error(), CapturedAt: captured}
+ return
+ }
+ out[i].CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateOK, CapturedAt: captured}
+ out[i].Settings = rows
+ }(i, p.Name)
+ }
+ wg.Wait()
+ return out
+}
+
+func (s *Reader) RestoreChecks(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured) (*CNPGRestoreChecksResponse, error) {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ recovery, found, _ := unstructured.NestedMap(cluster.Object, "spec", "bootstrap", "recovery")
+ if !found {
+ return nil, &ReadFailure{Status: http.StatusBadRequest, Message: fmt.Sprintf("Cluster %s/%s was not bootstrapped from a backup", namespace, name)}
+ }
+ resp := CNPGRestoreChecksResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: CNPGExecPermission{Exec: integration.PermissionAllowed, Grant: grantCreateExec.In(namespace).Ref()},
+ Database: restoreDatabase(cluster),
+ }
+ if target, ok := recovery["recoveryTarget"].(map[string]any); ok && len(target) > 0 {
+ resp.Target = target
+ }
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", k8s.SanitizeForLog(namespace), k8s.SanitizeForLog(name), err)
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "instance Pods unavailable: " + err.Error()}
+ }
+ primary, _, _ := unstructured.NestedString(cluster.Object, "status", "currentPrimary")
+ var target *corev1.Pod
+ for _, p := range pods {
+ if p.Name == primary {
+ target = p
+ }
+ }
+ if target == nil {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateError, Error: "no primary instance Pod is reported"}
+ return &resp, nil
+ }
+ resp.Pod = target.Name
+ resp.Permission.Exec = s.Access.Permission(ctx, grantCreateExec.In(namespace))
+ if resp.Permission.Exec == integration.PermissionDenied {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: execStateDenied, Error: "reading the restored databases needs " + integration.GrantText(resp.Permission.Grant)}
+ return &resp, nil
+ }
+ exec := s.Clients.Exec
+ if exec == nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "cluster client not available — check cluster connection"}
+ }
+ readCNPGRestoreChecks(ctx, exec, namespace, target.Name, &resp)
+ if resp.State == execStateDenied {
+ resp.Permission.Exec = integration.PermissionDenied
+ }
+ return &resp, nil
+}
+
+func (s *Reader) Parameters(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured) (*CNPGParametersResponse, error) {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ declared, query, skipped, omitted := declaredParameters(cluster)
+ resp := CNPGParametersResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: CNPGExecPermission{Exec: integration.PermissionAllowed, Grant: grantCreateExec.In(namespace).Ref()},
+ Declared: declared,
+ Skipped: skipped,
+ Omitted: omitted,
+ Instances: []CNPGInstanceSettings{},
+ }
+ if len(query) == 0 {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateOK}
+ return &resp, nil
+ }
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", k8s.SanitizeForLog(namespace), k8s.SanitizeForLog(name), err)
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "instance Pods unavailable: " + err.Error()}
+ }
+ resp.Permission.Exec = s.Access.Permission(ctx, grantCreateExec.In(namespace))
+ if resp.Permission.Exec == integration.PermissionDenied {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: execStateDenied, Error: "reading parameters on each instance needs " + integration.GrantText(resp.Permission.Grant)}
+ return &resp, nil
+ }
+ exec := s.Clients.Exec
+ if exec == nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "cluster client not available — check cluster connection"}
+ }
+ resp.Instances = readCNPGParameters(ctx, exec, namespace, pods, query)
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateOK}
+ for _, inst := range resp.Instances {
+ if inst.State == execStateDenied {
+ resp.Permission.Exec = integration.PermissionDenied
+ }
+ }
+ return &resp, nil
+}
diff --git a/internal/cnpg/inspect_test.go b/internal/cnpg/inspect_test.go
new file mode 100644
index 0000000000..bbd14ed65a
--- /dev/null
+++ b/internal/cnpg/inspect_test.go
@@ -0,0 +1,171 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "slices"
+ "strings"
+ "testing"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+)
+
+func TestParseCNPGTimelineHistoryChainsSwitches(t *testing.T) {
+ text := "# comment\n1\t0/3000000\tno recovery target specified\n\n2\t0/5000A28\tbefore 2026-10-01 12:00:00.123456+00\n"
+ got := parseCNPGTimelineHistory(text, 3)
+ want := []CNPGTimelineSwitch{
+ {From: 1, To: 2, SwitchLSN: "0/3000000", Reason: "no recovery target specified"},
+ {From: 2, To: 3, SwitchLSN: "0/5000A28", Reason: "before 2026-10-01 12:00:00.123456+00"},
+ }
+ if !slices.Equal(got, want) {
+ t.Errorf("history = %+v", got)
+ }
+}
+
+func cnpgRestoreFactsJSON(history string) string {
+ h := "null"
+ if history != "" {
+ h = `"` + strings.ReplaceAll(strings.ReplaceAll(history, "\t", `\t`), "\n", `\n`) + `"`
+ }
+ return `{"inRecovery" : false, "timeline" : 2, "history" : ` + h + `, "databaseCount" : 2, "databases" : [{"name":"app","bytes":8000000},{"name":"postgres","bytes":7000000}], "roleCount" : 3, "roles" : [{"name":"app","canLogin":true},{"name":"postgres","canLogin":true},{"name":"streaming_replica","canLogin":true}]}`
+}
+
+func cnpgSequencedExec(calls *[]cnpgExecCall, outs ...string) ExecFunc {
+ return func(_ context.Context, _ string, pod, container string, argv []string, stdin string) ([]byte, error) {
+ *calls = append(*calls, cnpgExecCall{pod: pod, container: container, argv: argv, stdin: stdin})
+ if len(*calls) > len(outs) {
+ return nil, errors.New("unexpected exec")
+ }
+ return []byte(outs[len(*calls)-1]), nil
+ }
+}
+
+func TestReadCNPGRestoreChecksReadsFactsThenTheBootstrapDatabase(t *testing.T) {
+ var calls []cnpgExecCall
+ contents := `{"tables" : 3, "estimatedRows" : 12000, "noEstimate" : 1, "largest" : [{"name":"public.orders","estimatedRows":10000,"bytes":4096000}]}`
+ resp := CNPGRestoreChecksResponse{Database: "app"}
+ readCNPGRestoreChecks(context.Background(), cnpgSequencedExec(&calls, cnpgRestoreFactsJSON("1\t0/5000A28\tbefore 2026-10-01 12:00:00+00\n"), contents), "db", "pg-r-1", &resp)
+ if resp.State != runtimeStateOK || resp.CNPGRestoreFacts == nil || resp.Timeline != 2 || len(resp.History) != 1 || resp.History[0].To != 2 {
+ t.Fatalf("facts = %+v / %+v", resp.CNPGRuntimeSource, resp.CNPGRestoreFacts)
+ }
+ if resp.Contents == nil || resp.Contents.Tables != 3 || resp.ContentsSource == nil || resp.ContentsSource.State != runtimeStateOK {
+ t.Fatalf("contents = %+v / %+v", resp.Contents, resp.ContentsSource)
+ }
+ if len(calls) != 2 || strings.Join(calls[0].argv, " ") != "psql -XAtq -v ON_ERROR_STOP=1 -d postgres -f -" || calls[0].stdin != cnpgRestoreFactsSQL {
+ t.Errorf("first exec = %+v", calls[0])
+ }
+ if strings.Join(calls[1].argv, " ") != "psql -XAtq -v ON_ERROR_STOP=1 -d app -f -" || calls[1].stdin != cnpgDatabaseContentsSQL {
+ t.Errorf("second exec = %+v", calls[1])
+ }
+ for _, sql := range []string{cnpgRestoreFactsSQL, cnpgDatabaseContentsSQL} {
+ if !strings.Contains(sql, "statement_timeout") || strings.Contains(strings.ToUpper(sql), "INSERT") || strings.Contains(strings.ToUpper(sql), "UPDATE ") {
+ t.Error("inspection SQL must be bounded and read-only")
+ }
+ }
+}
+
+func TestReadCNPGRestoreChecksWithoutHistoryOrUsableDatabase(t *testing.T) {
+ var calls []cnpgExecCall
+ resp := CNPGRestoreChecksResponse{Database: "host=elsewhere dbname=app"}
+ readCNPGRestoreChecks(context.Background(), cnpgSequencedExec(&calls, cnpgRestoreFactsJSON("")), "db", "pg-r-1", &resp)
+ if !resp.HistoryMissing || len(resp.History) != 0 {
+ t.Errorf("a missing history file is reported as such: %+v", resp.CNPGRestoreFacts)
+ }
+ // A connection string is never handed to psql's -d.
+ if len(calls) != 1 || resp.Contents != nil || resp.ContentsSource == nil || resp.ContentsSource.State != runtimeStateError {
+ t.Errorf("calls = %d, contents source = %+v", len(calls), resp.ContentsSource)
+ }
+
+ calls = nil
+ resp = CNPGRestoreChecksResponse{Database: "shop"}
+ readCNPGRestoreChecks(context.Background(), cnpgSequencedExec(&calls, cnpgRestoreFactsJSON("")), "db", "pg-r-1", &resp)
+ if len(calls) != 1 || resp.ContentsSource == nil || !strings.Contains(resp.ContentsSource.Error, `"shop" is not in the list`) {
+ t.Errorf("a database the complete list lacks is not connected to: %+v", resp.ContentsSource)
+ }
+
+ resp = CNPGRestoreChecksResponse{Database: "app"}
+ readCNPGRestoreChecks(context.Background(), cnpgFakeExec(&calls, "", errors.New(`pods "pg-r-1" is forbidden: User "bob" cannot create resource "pods/exec"`)), "db", "pg-r-1", &resp)
+ if resp.State != runtimeStateDenied || resp.CNPGRestoreFacts != nil {
+ t.Errorf("forbidden exec = %+v", resp.CNPGRuntimeSource)
+ }
+}
+
+func TestCNPGDeclaredParametersSkipsUnsafeNames(t *testing.T) {
+ cluster := &unstructured.Unstructured{Object: map[string]any{
+ "spec": map[string]any{"postgresql": map[string]any{"parameters": map[string]any{
+ "shared_buffers": "256MB",
+ "pg_stat_statements.max": "10000",
+ "Work_Mem": "8MB",
+ "bad,name": "x",
+ }}},
+ }}
+ declared, query, skipped, omitted := declaredParameters(cluster)
+ if len(declared) != 4 || omitted != 0 {
+ t.Errorf("declared = %+v", declared)
+ }
+ if !slices.Equal(query, []string{"work_mem", "pg_stat_statements.max", "shared_buffers"}) {
+ t.Errorf("query = %v", query)
+ }
+ if !slices.Equal(skipped, []string{"bad,name"}) {
+ t.Errorf("skipped = %v", skipped)
+ }
+}
+
+func TestReadCNPGParametersPassesNamesAsAVariable(t *testing.T) {
+ var calls []cnpgExecCall
+ rows := `[{"name":"application_name","value":null,"setByClient":true,"source":"client","context":"user","pendingRestart":false},{"name":"shared_buffers","value":"256MB","setByClient":false,"source":"configuration file","context":"postmaster","pendingRestart":true}]`
+ pods := []*corev1.Pod{{ObjectMeta: metav1.ObjectMeta{Name: "pg-1"}}}
+ got := readCNPGParameters(context.Background(), cnpgFakeExec(&calls, rows, nil), "db", pods, []string{"shared_buffers", "work_mem"})
+ if len(got) != 1 || got[0].State != runtimeStateOK || len(got[0].Settings) != 2 {
+ t.Fatalf("settings = %+v", got)
+ }
+ if sb := got[0].Settings[1]; sb.Value == nil || *sb.Value != "256MB" || !sb.PendingRestart || sb.Context != "postmaster" {
+ t.Errorf("shared_buffers = %+v", sb)
+ }
+ // The connection sets application_name itself, so its server value is not shown.
+ if an := got[0].Settings[0]; an.Value != nil || !an.SetByClient {
+ t.Errorf("application_name = %+v", an)
+ }
+ if strings.Count(cnpgParametersSQL, "SET ") != 1 || !strings.HasPrefix(cnpgParametersSQL, "SET search_path = pg_catalog;") {
+ t.Error("the parameters read sets only search_path: any other SET hides the server value of what it sets")
+ }
+ if !slices.Contains(calls[0].argv, "names=shared_buffers,work_mem") || strings.Contains(calls[0].stdin, "shared_buffers") || !strings.Contains(calls[0].stdin, ":'names'") {
+ t.Errorf("names must travel as a psql variable, never in the SQL text: %+v", calls[0])
+ }
+
+ got = readCNPGParameters(context.Background(), cnpgFakeExec(&calls, "", context.DeadlineExceeded), "db", pods, []string{"work_mem"})
+ if got[0].State != runtimeStateUnreachable || got[0].Settings != nil {
+ t.Errorf("an instance that does not answer reads unreachable: %+v", got[0])
+ }
+
+ // A read that matches nothing is a read: settings serializes as [], not as absent.
+ got = readCNPGParameters(context.Background(), cnpgFakeExec(&calls, "[]", nil), "db", pods, []string{"no_such_parameter"})
+ if b, _ := json.Marshal(got[0]); !strings.Contains(string(b), `"settings":[]`) {
+ t.Errorf("empty read = %s", b)
+ }
+}
+
+// Every diagnostic runs as the postgres superuser, so each pins search_path
+// before anything else; an application schema's function must never stand in
+// for a built-in.
+func TestCNPGDiagnosticSQLPinsSearchPath(t *testing.T) {
+ for name, sql := range map[string]string{
+ "blocking": cnpgBlockingSQL,
+ "signal": cnpgSignalSQL("pg_cancel_backend"),
+ "restore": cnpgRestoreFactsSQL,
+ "contents": cnpgDatabaseContentsSQL,
+ "params": cnpgParametersSQL,
+ } {
+ if !strings.HasPrefix(sql, "SET search_path = pg_catalog;") {
+ t.Errorf("%s SQL does not start by pinning search_path", name)
+ }
+ }
+ for name, sql := range map[string]string{"blocking": cnpgBlockingSQL, "restore": cnpgRestoreFactsSQL, "contents": cnpgDatabaseContentsSQL} {
+ if !strings.Contains(sql, "default_transaction_read_only = on") {
+ t.Errorf("%s SQL only reads, so it runs read-only", name)
+ }
+ }
+}
diff --git a/internal/cnpg/instances.go b/internal/cnpg/instances.go
new file mode 100644
index 0000000000..3d54419e6e
--- /dev/null
+++ b/internal/cnpg/instances.go
@@ -0,0 +1,52 @@
+package cnpg
+
+import (
+ "errors"
+ "sort"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+const (
+ clusterLabel = "cnpg.io/cluster"
+ defaultLogContainer = "postgres"
+ defaultLogTailLines = 200
+ logDiscoveryInterval = 5 * time.Second
+ logsEmptyMessage = "No readable logs from this cluster's instances in this snapshot. Refresh after the instances start."
+ logsIntervalEmpty = "No log lines in this interval from the runs Kubernetes still keeps (the current and previous run of each instance container)."
+ logsNoInstanceMessage = "This cluster has no instance Pods yet."
+)
+
+// cnpgClusterInstancePods returns the Cluster's instance Pods under the same
+// label-and-controller-UID rule the workspace uses, sorted by name.
+func clusterInstancePods(cache *k8s.ResourceCache, cluster *unstructured.Unstructured) ([]*corev1.Pod, error) {
+ lister := cache.Pods()
+ if lister == nil {
+ return nil, errors.New("pod cache unavailable")
+ }
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ candidates, err := lister.Pods(namespace).List(labels.SelectorFromSet(labels.Set{clusterLabel: name}))
+ if err != nil {
+ return nil, err
+ }
+ uids := map[string]types.UID{namespace + "/" + name: cluster.GetUID()}
+ pods := make([]*corev1.Pod, 0, len(candidates))
+ for _, p := range candidates {
+ if p != nil && isCNPGInstancePod(p, uids) {
+ pods = append(pods, p)
+ }
+ }
+ sort.Slice(pods, func(i, j int) bool { return pods[i].Name < pods[j].Name })
+ return pods, nil
+}
+
+func instanceRole(p *corev1.Pod) string { return cnpg.InstanceRole(p) }
+
+const Group = cnpg.Group
diff --git a/internal/cnpg/logs.go b/internal/cnpg/logs.go
new file mode 100644
index 0000000000..cb2552e18f
--- /dev/null
+++ b/internal/cnpg/logs.go
@@ -0,0 +1,642 @@
+package cnpg
+
+import (
+ "bufio"
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "io"
+ "log"
+ "math"
+ "net/http"
+ "net/url"
+ "strings"
+ "sync"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
+)
+
+// ClusterLogsResponse is GET /api/cnpg/clusters/{namespace}/{name}/logs.
+// Pods and SourceLabels list only the instances that contributed a source to
+// this snapshot; SourceLabels maps a Pod to its role and ordinal ("replica 2").
+type ClusterLogsResponse struct {
+ UID types.UID `json:"uid"`
+ Pods []podlogs.PodInfo `json:"pods"`
+ Logs []podlogs.Entry `json:"logs"`
+ Notice string `json:"notice"`
+ SourceLabels map[string]string `json:"sourceLabels,omitempty"`
+ CapturedAt string `json:"capturedAt"`
+ EmptyMessage string `json:"emptyMessage"`
+}
+
+// admit reports whether an entry is new, recording it when it is. Lines
+// arrive in order per container, so anything before the last delivered
+// timestamp was already sent.
+func (c *cnpgStreamCursor) admit(entry podlogs.Entry) bool {
+ ts, err := time.Parse(time.RFC3339Nano, entry.Timestamp)
+ if err != nil {
+ return true
+ }
+ switch {
+ case ts.Before(c.last):
+ return false
+ case ts.Equal(c.last):
+ if c.atLast[entry.Content] {
+ return false
+ }
+ default:
+ c.last = ts
+ c.atLast = map[string]bool{}
+ }
+ c.atLast[entry.Content] = true
+ return true
+}
+
+// annotateCNPGLogEntry fills the parsed fields of a CloudNativePG log line and
+// leaves anything that is not one untouched.
+func annotateCNPGLogEntry(entry *podlogs.Entry) {
+ content := strings.TrimSpace(entry.Content)
+ if !strings.HasPrefix(content, "{") {
+ return
+ }
+ var rec cnpgLogRecord
+ if err := json.Unmarshal([]byte(content), &rec); err != nil {
+ return
+ }
+ level := strings.ToUpper(rec.Level)
+ message := rec.Msg
+ if rec.Record != nil {
+ if rec.Record.ErrorSeverity != "" {
+ level = rec.Record.ErrorSeverity
+ }
+ if rec.Record.Message != "" {
+ message = rec.Record.Message
+ }
+ }
+ if errText := cnpgLogErrorText(rec.Error); errText != "" {
+ if message == "" {
+ message = errText
+ } else {
+ message += ": " + errText
+ }
+ }
+ entry.Level, entry.Logger, entry.Message = level, rec.Logger, message
+}
+
+func cnpgContainerStatus(pod *corev1.Pod, name string) *corev1.ContainerStatus {
+ for _, statuses := range [][]corev1.ContainerStatus{pod.Status.ContainerStatuses, pod.Status.InitContainerStatuses, pod.Status.EphemeralContainerStatuses} {
+ for i := range statuses {
+ if statuses[i].Name == name {
+ return &statuses[i]
+ }
+ }
+ }
+ return nil
+}
+
+// cnpgInstanceSourceLabel names an instance by role and ordinal ("replica 3"):
+// the role alone cannot tell two replicas apart.
+func cnpgInstanceSourceLabel(p *corev1.Pod, role string) string {
+ name := p.Labels[instanceNameLabel]
+ if name == "" {
+ name = p.Name
+ }
+ if i := strings.LastIndex(name, "-"); i >= 0 && i < len(name)-1 {
+ return role + " " + name[i+1:]
+ }
+ return role
+}
+
+// cnpgIntervalLogSources picks, per instance container, the runs whose lines
+// can fall in [since, until]. The kubelet keeps a container's current run and
+// the one before its last restart; an interval that ends before the current
+// run started lives only in the previous run. lost counts containers where an
+// older run than those two covered part of the interval.
+func cnpgIntervalLogSources(pods []*corev1.Pod, container string, since, until time.Time) (sources []podlogs.Source, lost int) {
+ for _, pod := range pods {
+ for _, c := range cnpgLogContainers(pod, container) {
+ status := cnpgContainerStatus(pod, c)
+ if status == nil {
+ continue
+ }
+ var currentStart time.Time
+ switch {
+ case status.State.Running != nil:
+ currentStart = status.State.Running.StartedAt.Time
+ case status.State.Terminated != nil:
+ currentStart = status.State.Terminated.StartedAt.Time
+ }
+ started := status.State.Running != nil || status.State.Terminated != nil
+ if started && (currentStart.IsZero() || !until.Before(currentStart)) {
+ sources = append(sources, podlogs.NewSource(pod, c, false))
+ }
+ if !currentStart.IsZero() && !since.Before(currentStart) {
+ continue
+ }
+ prev := status.LastTerminationState.Terminated
+ if prev == nil {
+ if status.RestartCount > 0 {
+ lost++
+ }
+ continue
+ }
+ if prev.FinishedAt.IsZero() || !prev.FinishedAt.Time.Before(since) {
+ sources = append(sources, podlogs.NewSource(pod, c, true))
+ }
+ if status.RestartCount > 1 && !prev.StartedAt.IsZero() && since.Before(prev.StartedAt.Time) {
+ lost++
+ }
+ }
+ }
+ return sources, lost
+}
+
+func cnpgLogContainers(pod *corev1.Pod, selected string) []string {
+ if selected != "all" {
+ containers := k8s.GetContainersForPod(pod, selected, true)
+ if len(containers) > 0 {
+ return containers
+ }
+ for _, c := range pod.Spec.EphemeralContainers {
+ if c.Name == selected {
+ return []string{selected}
+ }
+ }
+ return nil
+ }
+ containers := k8s.GetContainersForPod(pod, "", true)
+ for _, c := range pod.Spec.InitContainers {
+ containers = append(containers, c.Name)
+ }
+ for _, c := range pod.Spec.EphemeralContainers {
+ containers = append(containers, c.Name)
+ }
+ return containers
+}
+
+func cnpgLogErrorText(v any) string {
+ switch e := v.(type) {
+ case nil:
+ return ""
+ case string:
+ return e
+ default:
+ b, err := json.Marshal(e)
+ if err != nil {
+ return ""
+ }
+ return string(b)
+ }
+}
+
+type LogQuery struct {
+ Container string
+ TailLines int64
+ SinceSeconds *int64
+ SinceTime time.Time
+ UntilTime time.Time
+ Pod string
+}
+
+// cnpgLogRecord is the subset of a CloudNativePG instance-manager JSON log
+// line the viewer surfaces. PostgreSQL's own log lines arrive wrapped, with the
+// server's severity and message under record.
+type cnpgLogRecord struct {
+ Level string `json:"level"`
+ Logger string `json:"logger"`
+ Msg string `json:"msg"`
+ Error any `json:"error"`
+ Record *struct {
+ ErrorSeverity string `json:"error_severity"`
+ Message string `json:"message"`
+ } `json:"record"`
+}
+
+func cnpgLogSourceLabel(role, container, selected string) string {
+ if selected == "all" {
+ if role == "" {
+ return container
+ }
+ return role + " · " + container
+ }
+ return role
+}
+
+func cnpgLogsIntervalGone(lost int) string {
+ if lost == 1 {
+ return "Kubernetes keeps only the current and previous run of each container; 1 instance run that covered part of this interval was replaced and its lines are gone."
+ }
+ return fmt.Sprintf("Kubernetes keeps only the current and previous run of each container; %d instance runs that covered part of this interval were replaced and their lines are gone.", lost)
+}
+
+func cnpgSnapshotLogSources(pods []*corev1.Pod, selected string) []podlogs.Source {
+ sources := []podlogs.Source{}
+ for _, pod := range pods {
+ for _, container := range cnpgLogContainers(pod, selected) {
+ status := cnpgContainerStatus(pod, container)
+ if status == nil {
+ continue
+ }
+ if status.State.Running != nil || status.State.Terminated != nil {
+ sources = append(sources, podlogs.NewSource(pod, container, false))
+ } else if status.State.Waiting != nil && status.LastTerminationState.Terminated != nil {
+ sources = append(sources, podlogs.NewSource(pod, container, true))
+ }
+ }
+ }
+ return sources
+}
+
+// cnpgStreamCursor remembers where one container's follow left off, so a
+// stream that ends while its Pod is still an instance resumes instead of
+// replaying lines the client already has. Only the stream loop touches it.
+type cnpgStreamCursor struct {
+ last time.Time
+ // atLast holds the contents delivered with timestamp == last. The pod log
+ // API's sinceTime is second-granular, so a resume replays that second and
+ // only (timestamp, content) tells a replay from a new line.
+ atLast map[string]bool
+}
+
+type cnpgStreamHandle struct {
+ cancel context.CancelFunc
+}
+
+func followCNPGContainerLogs(ctx context.Context, client kubernetes.Interface, namespace, podName string, opts corev1.PodLogOptions, logCh chan<- podlogs.Entry) bool {
+ stream, err := client.CoreV1().Pods(namespace).GetLogs(podName, &opts).Stream(ctx)
+ if err != nil {
+ if ctx.Err() == nil {
+ log.Printf("[cnpg] Failed to follow logs for %s/%s/%s: %v", namespace, podName, opts.Container, err)
+ }
+ return false
+ }
+ defer stream.Close()
+ reader := bufio.NewReader(stream)
+ for {
+ line, err := reader.ReadString('\n')
+ if line = strings.TrimSuffix(line, "\n"); line != "" && (err == nil || err == io.EOF) {
+ ts, content := podlogs.ParseLine(line)
+ select {
+ case logCh <- podlogs.Entry{Pod: podName, Container: opts.Container, Timestamp: ts, Content: content, Previous: opts.Previous}:
+ case <-ctx.Done():
+ return false
+ }
+ }
+ if err != nil {
+ if err != io.EOF && ctx.Err() == nil {
+ log.Printf("[cnpg] Failed to read logs for %s/%s/%s: %v", namespace, podName, opts.Container, err)
+ }
+ return err == io.EOF
+ }
+ }
+}
+
+func (q LogQuery) keep(entry podlogs.Entry) bool {
+ if q.SinceTime.IsZero() || entry.Timestamp == "" {
+ return true
+ }
+ ts, err := time.Parse(time.RFC3339Nano, entry.Timestamp)
+ if err != nil {
+ return true
+ }
+ return !ts.Before(q.SinceTime) && (q.UntilTime.IsZero() || !ts.After(q.UntilTime))
+}
+
+func ParseLogQuery(q url.Values, now time.Time) (LogQuery, error) {
+ out := LogQuery{
+ Container: q.Get("container"),
+ TailLines: podlogs.ParseTailLines(q.Get("tailLines"), defaultLogTailLines),
+ SinceSeconds: podlogs.ParseSinceSeconds(q.Get("sinceSeconds")),
+ Pod: q.Get("pod"),
+ }
+ if out.Container == "" {
+ out.Container = defaultLogContainer
+ }
+ if raw := q.Get("sinceTime"); raw != "" {
+ if q.Get("sinceSeconds") != "" {
+ return out, errors.New("sinceSeconds and sinceTime are mutually exclusive")
+ }
+ t, err := time.Parse(time.RFC3339, raw)
+ if err != nil {
+ return out, fmt.Errorf("invalid sinceTime %q (expected RFC3339)", raw)
+ }
+ out.SinceTime = t
+ // The pod log API takes whole seconds; round up and trim the overlap
+ // from the entries afterwards.
+ secs := max(int64(math.Ceil(now.Sub(t).Seconds())), 1)
+ out.SinceSeconds = &secs
+ }
+ if raw := q.Get("untilTime"); raw != "" {
+ if out.SinceTime.IsZero() {
+ return out, errors.New("untilTime requires sinceTime")
+ }
+ t, err := time.Parse(time.RFC3339, raw)
+ if err != nil {
+ return out, fmt.Errorf("invalid untilTime %q (expected RFC3339)", raw)
+ }
+ if !t.After(out.SinceTime) {
+ return out, errors.New("untilTime must be after sinceTime")
+ }
+ out.UntilTime = t
+ // An interval is read from its start: the pod log API has no upper
+ // bound, and a tail would return the lines nearest now instead.
+ if q.Get("tailLines") == "" {
+ out.TailLines = 0
+ }
+ }
+ return out, nil
+}
+
+// restartOptions returns the follow request for the next (re)start: the
+// caller's window the first time, and from the last delivered second after.
+func (c *cnpgStreamCursor) restartOptions(container string, tailLines int64, sinceSeconds *int64) corev1.PodLogOptions {
+ opts := corev1.PodLogOptions{Container: container, Timestamps: true, Follow: true}
+ if c.last.IsZero() {
+ opts.TailLines = &tailLines
+ opts.SinceSeconds = sinceSeconds
+ return opts
+ }
+ since := metav1.NewTime(c.last.Truncate(time.Second))
+ opts.SinceTime = &since
+ return opts
+}
+
+// selectCNPGLogPods narrows to the requested instance. ok is false when the
+// requested Pod is not one of the Cluster's instances.
+func selectCNPGLogPods(pods []*corev1.Pod, want string) ([]*corev1.Pod, bool) {
+ if want == "" {
+ return pods, true
+ }
+ for _, p := range pods {
+ if p.Name == want {
+ return []*corev1.Pod{p}, true
+ }
+ }
+ return nil, false
+}
+
+type LogTarget struct {
+ reader *Reader
+ namespace, name string
+ cache *k8s.ResourceCache
+ cluster *unstructured.Unstructured
+ pods []*corev1.Pod
+}
+
+func (rd *Reader) PrepareLogs(ctx context.Context, namespace, name string, query LogQuery) (*LogTarget, error) {
+ cache, cluster, err := rd.Observations.Cluster(ctx, namespace, name,
+ auth.Grant{Resource: "pods", Verb: "list", Namespace: namespace},
+ auth.Grant{Resource: "pods", Subresource: "log", Verb: "get", Namespace: namespace})
+ if err != nil {
+ return nil, err
+ }
+ instances, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", namespace, name, err)
+ return nil, &ReadFailure{http.StatusServiceUnavailable, "instance Pods unavailable: " + err.Error()}
+ }
+ pods, ok := selectCNPGLogPods(instances, query.Pod)
+ if !ok {
+ return nil, &ReadFailure{http.StatusBadRequest, "pod " + query.Pod + " is not an instance of CloudNativePG Cluster " + namespace + "/" + name}
+ }
+ return &LogTarget{reader: rd, namespace: namespace, name: name, cache: cache, cluster: cluster, pods: pods}, nil
+}
+
+func (target *LogTarget) RequireClient() error {
+ if target.reader.Clients.Typed == nil {
+ return &ReadFailure{http.StatusServiceUnavailable, "cluster client unavailable"}
+ }
+ return nil
+}
+
+func (target *LogTarget) currentCluster(ctx context.Context) (*unstructured.Unstructured, error) {
+ objects, err := target.reader.Observations.DynamicList(ctx, target.cache, "Cluster", Group, target.namespace)
+ return SelectCluster(objects, err, target.namespace, target.name)
+}
+
+func SelectCluster(objects []*unstructured.Unstructured, err error, namespace, name string) (*unstructured.Unstructured, error) {
+ clusters, err := filterCNPGGroup(objects, err)
+ if err != nil {
+ return nil, err
+ }
+ for _, c := range clusters {
+ if c.GetNamespace() == namespace && c.GetName() == name && c.GroupVersionKind().Group == Group {
+ return c, nil
+ }
+ }
+ return nil, nil
+}
+
+func (target *LogTarget) Snapshot(ctx context.Context, query LogQuery) (ClusterLogsResponse, error) {
+ namespace, cluster, pods := target.namespace, target.cluster, target.pods
+ resp := ClusterLogsResponse{
+ UID: cluster.GetUID(),
+ Pods: []podlogs.PodInfo{},
+ Logs: []podlogs.Entry{},
+ CapturedAt: time.Now().UTC().Format(time.RFC3339),
+ EmptyMessage: logsEmptyMessage,
+ }
+ if len(pods) == 0 {
+ resp.EmptyMessage = logsNoInstanceMessage
+ return resp, nil
+ }
+ client := target.reader.Clients.Typed
+ if client == nil {
+ return ClusterLogsResponse{}, &ReadFailure{http.StatusServiceUnavailable, "cluster client unavailable"}
+ }
+
+ var snapshot podlogs.Snapshot
+ lost := 0
+ if query.UntilTime.IsZero() {
+ snapshot = podlogs.CollectSources(ctx, client, namespace, cnpgSnapshotLogSources(pods, query.Container), query.TailLines, query.SinceSeconds, true)
+ } else {
+ var sources []podlogs.Source
+ sources, lost = cnpgIntervalLogSources(pods, query.Container, query.SinceTime, query.UntilTime)
+ snapshot = podlogs.CollectSources(ctx, client, namespace, sources, query.TailLines, query.SinceSeconds, true)
+ snapshot.Notice = snapshot.Summarize(func(clip podlogs.Clip) bool {
+ return clip.Last.IsZero() || clip.Last.Before(query.UntilTime)
+ })
+ resp.EmptyMessage = logsIntervalEmpty
+ }
+ shown := []*corev1.Pod{}
+ sourceLabels := map[string]string{}
+ for _, p := range pods {
+ if !snapshot.SourcePods[p.Name] {
+ continue
+ }
+ shown = append(shown, p)
+ if role := instanceRole(p); role != "" {
+ sourceLabels[p.Name] = cnpgInstanceSourceLabel(p, role)
+ }
+ }
+ for _, entry := range snapshot.Logs {
+ if !query.keep(entry) {
+ continue
+ }
+ entry.SourceLabel = cnpgLogSourceLabel(sourceLabels[entry.Pod], entry.Container, query.Container)
+ if entry.Previous && entry.SourceLabel != "" {
+ entry.SourceLabel += " · previous run"
+ }
+ annotateCNPGLogEntry(&entry)
+ resp.Logs = append(resp.Logs, entry)
+ }
+ podlogs.Sort(resp.Logs)
+ resp.Pods = podlogs.BuildPodInfos(shown)
+ resp.Notice = snapshot.Notice
+ if lost > 0 {
+ resp.Notice = strings.TrimSpace(cnpgLogsIntervalGone(lost) + " " + resp.Notice)
+ }
+ if len(sourceLabels) > 0 {
+ resp.SourceLabels = sourceLabels
+ }
+ return resp, nil
+}
+
+func (target *LogTarget) Follow(parent context.Context, query LogQuery, send func(string, any)) {
+ namespace, name, cluster, pods, cache, client := target.namespace, target.name, target.cluster, target.pods, target.cache, target.reader.Clients.Typed
+ uid := cluster.GetUID()
+ send("connected", map[string]any{
+ "cluster": name, "namespace": namespace, "uid": uid, "pods": podlogs.BuildPodInfos(pods),
+ })
+
+ ctx, cancel := context.WithCancel(parent)
+ defer cancel()
+ logCh := make(chan podlogs.Entry, 1000)
+ var active sync.Map
+ var completed sync.Map
+ roles := map[string]string{}
+ cursors := map[string]*cnpgStreamCursor{}
+ start := func(pods []*corev1.Pod) {
+ for _, pod := range pods {
+ if role := instanceRole(pod); role != "" {
+ roles[pod.Name] = cnpgInstanceSourceLabel(pod, role)
+ }
+ for _, c := range cnpgLogContainers(pod, query.Container) {
+ status := cnpgContainerStatus(pod, c)
+ if status == nil {
+ continue
+ }
+ previous := status.State.Waiting != nil && status.LastTerminationState.Terminated != nil
+ if status.State.Running == nil && status.State.Terminated == nil && !previous {
+ continue
+ }
+ key := pod.Name + "/" + c
+ run := fmt.Sprintf("%s/%d", status.ContainerID, status.RestartCount)
+ if previous {
+ run = fmt.Sprintf("%s/%d/previous", status.LastTerminationState.Terminated.ContainerID, status.RestartCount)
+ }
+ terminated := status.State.Terminated != nil || previous
+ if read, ok := completed.Load(key); terminated && ok && read == run {
+ continue
+ }
+ if _, exists := active.Load(key); exists {
+ continue
+ }
+ cursor := cursors[key]
+ if cursor == nil {
+ cursor = &cnpgStreamCursor{}
+ cursors[key] = cursor
+ }
+ opts := cursor.restartOptions(c, query.TailLines, query.SinceSeconds)
+ if previous {
+ opts.Previous = true
+ opts.Follow = false
+ }
+ streamCtx, streamCancel := context.WithCancel(ctx)
+ handle := &cnpgStreamHandle{cancel: streamCancel}
+ active.Store(key, handle)
+ go func(podName, key string) {
+ defer active.CompareAndDelete(key, handle)
+ if followCNPGContainerLogs(streamCtx, client, namespace, podName, opts, logCh) && terminated {
+ completed.Store(key, run)
+ }
+ }(pod.Name, key)
+ }
+ }
+ }
+ start(pods)
+
+ known := map[string]bool{}
+ for _, p := range pods {
+ known[p.Name] = true
+ }
+ ticker := time.NewTicker(logDiscoveryInterval)
+ defer ticker.Stop()
+ for {
+ select {
+ case <-ctx.Done():
+ return
+ case entry := <-logCh:
+ if cursor := cursors[entry.Pod+"/"+entry.Container]; cursor != nil && !cursor.admit(entry) {
+ continue
+ }
+ if !query.keep(entry) {
+ continue
+ }
+ entry.SourceLabel = cnpgLogSourceLabel(roles[entry.Pod], entry.Container, query.Container)
+ if entry.Previous && entry.SourceLabel != "" {
+ entry.SourceLabel += " · previous run"
+ }
+ annotateCNPGLogEntry(&entry)
+ send("log", entry)
+ case <-ticker.C:
+ current, err := target.currentCluster(ctx)
+ if err != nil {
+ continue
+ }
+ if current == nil || current.GetUID() != uid {
+ send("end", map[string]string{"reason": "cluster deleted"})
+ return
+ }
+ all, err := clusterInstancePods(cache, current)
+ if err != nil {
+ continue
+ }
+ currentPods, _ := selectCNPGLogPods(all, query.Pod)
+ present := map[string]bool{}
+ for _, p := range currentPods {
+ present[p.Name] = true
+ if !known[p.Name] {
+ known[p.Name] = true
+ send("pod_added", map[string]any{"pods": []podlogs.PodInfo{podlogs.BuildPodInfo(p, time.Now())}})
+ }
+ }
+ for podName := range known {
+ if present[podName] {
+ continue
+ }
+ delete(known, podName)
+ active.Range(func(key, value any) bool {
+ if strings.HasPrefix(key.(string), podName+"/") {
+ value.(*cnpgStreamHandle).cancel()
+ active.Delete(key)
+ }
+ return true
+ })
+ for key := range cursors {
+ if strings.HasPrefix(key, podName+"/") {
+ delete(cursors, key)
+ }
+ }
+ completed.Range(func(key, _ any) bool {
+ if strings.HasPrefix(key.(string), podName+"/") {
+ completed.Delete(key)
+ }
+ return true
+ })
+ send("pod_removed", map[string]string{"pod": podName, "reason": "terminated"})
+ }
+ start(currentPods)
+ }
+ }
+}
diff --git a/internal/cnpg/logs_test.go b/internal/cnpg/logs_test.go
new file mode 100644
index 0000000000..d1807b77d7
--- /dev/null
+++ b/internal/cnpg/logs_test.go
@@ -0,0 +1,261 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "net/url"
+ "strings"
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
+)
+
+func TestPrepareLogsRequiresInventoryAndLogGrantsBeforeLookup(t *testing.T) {
+ denied := &ReadFailure{Status: 403, Message: "no access"}
+ reader := &Reader{Observations: Observations{Cluster: func(_ context.Context, namespace, name string, grants ...auth.Grant) (*k8s.ResourceCache, *unstructured.Unstructured, error) {
+ if namespace != "pg" || name != "orders" || len(grants) != 2 {
+ t.Fatalf("unexpected target or grants: %s/%s %+v", namespace, name, grants)
+ }
+ if grants[0] != (auth.Grant{Resource: "pods", Verb: "list", Namespace: "pg"}) || grants[1] != (auth.Grant{Resource: "pods", Subresource: "log", Verb: "get", Namespace: "pg"}) {
+ t.Fatalf("wrong operations: %+v", grants)
+ }
+ return nil, nil, denied
+ }}}
+ if target, err := reader.PrepareLogs(context.Background(), "pg", "orders", LogQuery{}); target != nil || !errors.Is(err, denied) {
+ t.Fatalf("denied preparation = %+v, %v", target, err)
+ }
+}
+
+func TestLogSnapshotWithoutInstancesKeepsArrayWireShape(t *testing.T) {
+ cluster := &unstructured.Unstructured{}
+ cluster.SetUID("cluster-uid")
+ target := &LogTarget{reader: &Reader{}, cluster: cluster}
+ response, err := target.Snapshot(context.Background(), LogQuery{})
+ if err != nil || response.UID != "cluster-uid" || response.Pods == nil || response.Logs == nil || response.EmptyMessage != logsNoInstanceMessage {
+ t.Fatalf("empty snapshot = %+v, %v", response, err)
+ }
+ if err := target.RequireClient(); err == nil {
+ t.Fatal("a stream must require a client even before any instances exist")
+ }
+}
+
+func TestAnnotateCNPGLogEntry(t *testing.T) {
+ cases := []struct {
+ name, content, level, logger, message string
+ }{
+ {
+ name: "postgres record",
+ content: `{"level":"info","ts":"2026-09-28T14:19:58.123Z","logger":"postgres","msg":"record","record":{"error_severity":"FATAL","message":"password authentication failed","log_time":"2026-09-28 14:19:58.123 UTC"}}`,
+ level: "FATAL", logger: "postgres", message: "password authentication failed",
+ },
+ {
+ name: "instance manager error",
+ content: `{"level":"error","ts":"2026-09-28T14:19:58Z","logger":"barman-cloud-wal-archive","msg":"Error invoking barman-cloud-wal-archive","error":"exit status 4"}`,
+ level: "ERROR", logger: "barman-cloud-wal-archive", message: "Error invoking barman-cloud-wal-archive: exit status 4",
+ },
+ {
+ name: "structured error",
+ content: `{"level":"error","msg":"failed","error":{"code":2}}`,
+ level: "ERROR", message: `failed: {"code":2}`,
+ },
+ {name: "plain text", content: "LOG: database system is ready"},
+ {name: "broken json", content: `{"level":"info"`},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ entry := podlogs.Entry{Content: tc.content}
+ annotateCNPGLogEntry(&entry)
+ if entry.Level != tc.level || entry.Logger != tc.logger || entry.Message != tc.message || entry.Content != tc.content {
+ t.Fatalf("got level=%q logger=%q message=%q content-changed=%v", entry.Level, entry.Logger, entry.Message, entry.Content != tc.content)
+ }
+ })
+ }
+ raw, _ := json.Marshal(podlogs.Entry{Pod: "p", Content: "x"})
+ if strings.Contains(string(raw), "level") || strings.Contains(string(raw), "message") {
+ t.Fatalf("unparsed entries grew fields: %s", raw)
+ }
+}
+
+func TestParseCNPGLogQuery(t *testing.T) {
+ now := time.Date(2026, 9, 28, 12, 0, 0, 0, time.UTC)
+ req := logValues("/?sinceTime=2026-09-28T11:59:00.5Z")
+ if _, err := ParseLogQuery(req, now); err != nil {
+ t.Fatalf("fractional RFC3339 rejected: %v", err)
+ }
+ req = logValues("/?sinceTime=2026-09-28T11:58:30Z")
+ q, err := ParseLogQuery(req, now)
+ if err != nil || q.SinceSeconds == nil || *q.SinceSeconds != 90 || q.Container != "postgres" || q.TailLines != 200 {
+ t.Fatalf("query = %+v err=%v", q, err)
+ }
+ if q.keep(podlogs.Entry{Timestamp: "2026-09-28T11:58:29.9Z"}) || !q.keep(podlogs.Entry{Timestamp: "2026-09-28T11:58:30Z"}) {
+ t.Fatal("sinceTime overlap not trimmed")
+ }
+ req = logValues("/?sinceTime=2026-09-28T11:58:30Z&sinceSeconds=5")
+ if _, err := ParseLogQuery(req, now); err == nil {
+ t.Fatal("sinceTime with sinceSeconds accepted")
+ }
+}
+
+func TestCNPGStreamCursorResumesWithoutReplay(t *testing.T) {
+ var c cnpgStreamCursor
+ first := c.restartOptions("postgres", 200, nil)
+ if first.TailLines == nil || *first.TailLines != 200 || first.SinceTime != nil || !first.Follow || !first.Timestamps || first.Container != "postgres" {
+ t.Fatalf("first start = %+v", first)
+ }
+ line := func(ts, content string) podlogs.Entry {
+ return podlogs.Entry{Timestamp: ts, Content: content}
+ }
+ for _, e := range []podlogs.Entry{line("2026-09-28T14:00:00.1Z", "a"), line("2026-09-28T14:00:05.7Z", "b"), line("2026-09-28T14:00:05.7Z", "c")} {
+ if !c.admit(e) {
+ t.Fatalf("fresh line %+v rejected", e)
+ }
+ }
+
+ restart := c.restartOptions("postgres", 200, nil)
+ if restart.TailLines != nil || restart.SinceSeconds != nil || restart.SinceTime == nil ||
+ !restart.SinceTime.Time.Equal(time.Date(2026, 9, 28, 14, 0, 5, 0, time.UTC)) {
+ t.Fatalf("restart = %+v, want sinceTime at the last delivered second and no tail", restart)
+ }
+
+ // The resumed follow replays the boundary second.
+ replayed := []podlogs.Entry{line("2026-09-28T14:00:05.2Z", "earlier in the second"), line("2026-09-28T14:00:05.7Z", "b"), line("2026-09-28T14:00:05.7Z", "c")}
+ for _, e := range replayed {
+ if c.admit(e) {
+ t.Errorf("replayed line %+v admitted", e)
+ }
+ }
+ for _, e := range []podlogs.Entry{line("2026-09-28T14:00:05.7Z", "d"), line("2026-09-28T14:00:06Z", "e")} {
+ if !c.admit(e) {
+ t.Errorf("new line %+v rejected", e)
+ }
+ }
+ if c.admit(line("2026-09-28T14:00:05.7Z", "d")) {
+ t.Error("line before the new last timestamp admitted")
+ }
+}
+
+func TestCNPGIntervalLogSources(t *testing.T) {
+ at := func(m int) time.Time { return time.Date(2026, 9, 29, 23, m, 0, 0, time.UTC) }
+ since, until := at(20), at(30)
+ pod := func(status corev1.ContainerStatus) *corev1.Pod {
+ p := &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "pg-1", Namespace: "ns"}, Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "postgres"}}}}
+ p.Status.ContainerStatuses = []corev1.ContainerStatus{status}
+ return p
+ }
+ waiting := cnpgRestartedStatus(at(0), at(10), at(40), 3)
+ waiting.State = corev1.ContainerState{Waiting: &corev1.ContainerStateWaiting{Reason: "CrashLoopBackOff"}}
+ cases := []struct {
+ name string
+ status corev1.ContainerStatus
+ current, previous bool
+ lost int
+ }{
+ {"restarted after the interval reads only the previous run", cnpgRestartedStatus(at(45), at(10), at(40), 1), false, true, 0},
+ {"restart inside the interval reads both runs", cnpgRestartedStatus(at(25), at(10), at(24), 1), true, true, 0},
+ {"running since before the interval reads the current run", cnpgRestartedStatus(at(5), at(0), at(4), 1), true, false, 0},
+ {"previous run ended before the interval is skipped", cnpgRestartedStatus(at(25), at(0), at(15), 2), true, false, 0},
+ {"older runs than the kept two covered the interval", cnpgRestartedStatus(at(45), at(28), at(44), 5), false, true, 1},
+ {"crash-looping container reads its previous run", waiting, false, true, 0},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ sources, lost := cnpgIntervalLogSources([]*corev1.Pod{pod(tc.status)}, "postgres", since, until)
+ var current, previous bool
+ for _, s := range sources {
+ if s.Previous {
+ previous = true
+ } else {
+ current = true
+ }
+ }
+ if current != tc.current || previous != tc.previous || lost != tc.lost {
+ t.Fatalf("current=%v previous=%v lost=%d, want %v %v %d", current, previous, lost, tc.current, tc.previous, tc.lost)
+ }
+ })
+ }
+}
+
+func TestCNPGLogsContainerSelection(t *testing.T) {
+ q, err := ParseLogQuery(logValues("/?container=all"), time.Now())
+ if err != nil || q.Container != "all" {
+ t.Fatalf("%+v %v", q, err)
+ }
+ q, err = ParseLogQuery(logValues("/"), time.Now())
+ if err != nil || q.Container != "postgres" {
+ t.Fatalf("default: %+v %v", q, err)
+ }
+ p := &corev1.Pod{Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "postgres"}}, InitContainers: []corev1.Container{{Name: "sidecar"}}}}
+ if got := strings.Join(cnpgLogContainers(p, "all"), ","); got != "postgres,sidecar" {
+ t.Fatal(got)
+ }
+ if got := strings.Join(cnpgLogContainers(p, "sidecar"), ","); got != "sidecar" {
+ t.Fatal(got)
+ }
+ if len(cnpgLogContainers(p, "no-such-container")) != 0 {
+ t.Fatal("selected unknown container")
+ }
+}
+
+func TestCNPGLogsWaitingSources(t *testing.T) {
+ for _, retained := range []bool{true, false} {
+ t.Run(fmt.Sprintf("retained=%t", retained), func(t *testing.T) {
+ pod := &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "pg-1"}, Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "postgres"}}}, Status: corev1.PodStatus{ContainerStatuses: []corev1.ContainerStatus{{Name: "postgres", State: corev1.ContainerState{Waiting: &corev1.ContainerStateWaiting{Reason: "CrashLoopBackOff"}}}}}}
+ if retained {
+ pod.Status.ContainerStatuses[0].LastTerminationState.Terminated = &corev1.ContainerStateTerminated{ContainerID: "containerd://crash", ExitCode: 1}
+ }
+ sources := cnpgSnapshotLogSources([]*corev1.Pod{pod}, "postgres")
+ if !retained {
+ if len(sources) != 0 {
+ t.Fatalf("read container that never started: %+v", sources)
+ }
+ return
+ }
+ if len(sources) != 1 || !sources[0].Previous {
+ t.Fatalf("missing previous-run source: %+v", sources)
+ }
+ })
+ }
+}
+
+func TestParseCNPGLogQueryInterval(t *testing.T) {
+ now := time.Date(2026, 9, 28, 12, 0, 0, 0, time.UTC)
+ q, err := ParseLogQuery(logValues("/?sinceTime=2026-09-28T11:00:00Z&untilTime=2026-09-28T11:05:00Z"), now)
+ if err != nil || q.TailLines != 0 || q.SinceSeconds == nil || *q.SinceSeconds != 3600 {
+ t.Fatalf("q = %+v err = %v", q, err)
+ }
+ if !q.keep(podlogs.Entry{Timestamp: "2026-09-28T11:05:00Z"}) || q.keep(podlogs.Entry{Timestamp: "2026-09-28T11:05:00.1Z"}) || q.keep(podlogs.Entry{Timestamp: "2026-09-28T10:59:59Z"}) {
+ t.Fatal("interval bounds not applied")
+ }
+ for _, bad := range []string{"/?untilTime=2026-09-28T11:05:00Z", "/?sinceTime=2026-09-28T11:05:00Z&untilTime=2026-09-28T11:00:00Z", "/?sinceTime=2026-09-28T11:00:00Z&untilTime=x"} {
+ if _, err := ParseLogQuery(logValues(bad), now); err == nil {
+ t.Errorf("%s accepted", bad)
+ }
+ }
+}
+
+func cnpgRestartedStatus(currentStart, prevStart, prevEnd time.Time, restarts int32) corev1.ContainerStatus {
+ status := corev1.ContainerStatus{
+ Name: "postgres", RestartCount: restarts,
+ State: corev1.ContainerState{Running: &corev1.ContainerStateRunning{StartedAt: metav1.NewTime(currentStart)}},
+ }
+ if restarts > 0 {
+ status.LastTerminationState.Terminated = &corev1.ContainerStateTerminated{StartedAt: metav1.NewTime(prevStart), FinishedAt: metav1.NewTime(prevEnd)}
+ }
+ return status
+}
+func logValues(raw string) url.Values {
+ values, err := url.ParseQuery(strings.TrimPrefix(raw, "/?"))
+ if err != nil {
+ panic(err)
+ }
+ return values
+}
diff --git a/internal/cnpg/memo.go b/internal/cnpg/memo.go
new file mode 100644
index 0000000000..25ff3240b1
--- /dev/null
+++ b/internal/cnpg/memo.go
@@ -0,0 +1,134 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "net"
+ "sync"
+ "sync/atomic"
+ "time"
+
+ "golang.org/x/sync/singleflight"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+)
+
+// cnpgRuntimeRunner runs proxy reads with bounded concurrency under the
+// request's context.
+type cnpgRuntimeRunner struct {
+ ctx context.Context
+ sem chan struct{}
+ wg sync.WaitGroup
+}
+
+func newCNPGRuntimeRunner(ctx context.Context) *cnpgRuntimeRunner {
+ return &cnpgRuntimeRunner{ctx: ctx, sem: make(chan struct{}, cnpgRuntimeConcurrency)}
+}
+
+func (r *cnpgRuntimeRunner) do(fn func(ctx context.Context)) {
+ r.wg.Add(1)
+ go func() {
+ defer r.wg.Done()
+ select {
+ case r.sem <- struct{}{}:
+ defer func() { <-r.sem }()
+ case <-r.ctx.Done():
+ // The caller is gone and nobody reads this result. Running fn here
+ // would bypass the cap: memoized reads outlive the caller, so every
+ // queued read would start at once.
+ return
+ }
+ // select picks at random when both are ready: a free slot does not
+ // make a cancelled caller's read worth starting.
+ if r.ctx.Err() != nil {
+ return
+ }
+ fn(r.ctx)
+ }()
+}
+
+func (r *cnpgRuntimeRunner) wait() { r.wg.Wait() }
+
+// The memo collapses repeated reads of one endpoint by one identity — several
+// viewers, a fast refresh — into one scrape per lifetime. The key includes the
+// kube context and the Pod UID, so a context switch or a recreated Pod never
+// serves a stale answer, and the identity, so one caller's answer never
+// reaches another.
+var (
+ cnpgRuntimeMemoMu sync.Mutex
+ cnpgRuntimeMemoEntries = map[string]cnpgRuntimeMemoEntry{}
+ cnpgRuntimeMemoGroup singleflight.Group
+)
+
+const cnpgRuntimeMemoMaxEntries = 4096
+
+// Two proxied requests (a scheme fallback) plus slack.
+const cnpgMemoizedReadTimeout = 2*cnpgRuntimeRequestTimeout + time.Second
+
+type cnpgRuntimeMemoEntry struct {
+ value any
+ expires time.Time
+}
+
+func memoized[T any](ctx context.Context, identity string, target proxyTarget, ttl time.Duration, fetch func(context.Context) T) T {
+ key := fmt.Sprintf("%s\x00%s/%s\x00%s\x00%s:%d%s", identity, target.namespace, target.pod, target.podUID, target.scheme, target.port, target.path)
+ now := time.Now()
+ cnpgRuntimeMemoMu.Lock()
+ if e, ok := cnpgRuntimeMemoEntries[key]; ok && now.Before(e.expires) {
+ cnpgRuntimeMemoMu.Unlock()
+ return e.value.(T)
+ }
+ cnpgRuntimeMemoMu.Unlock()
+
+ v, _, _ := cnpgRuntimeMemoGroup.Do(key, func() (any, error) {
+ // Every caller waiting on this key shares the one read, so it runs
+ // detached from whichever caller started it: that caller hanging up
+ // must not fail the others. Values (identity) are kept; the deadline
+ // covers a scheme fallback's second request.
+ readCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), cnpgMemoizedReadTimeout)
+ defer cancel()
+ timedOut := new(atomic.Bool)
+ got := fetch(context.WithValue(readCtx, cnpgReadTimedOutKey{}, timedOut))
+ // A read that ran out of time (its own deadline or a request's inside
+ // it) says nothing lasting about the Pod: the next caller tries again.
+ if readCtx.Err() == nil && !timedOut.Load() {
+ cnpgRuntimeMemoMu.Lock()
+ if len(cnpgRuntimeMemoEntries) >= cnpgRuntimeMemoMaxEntries {
+ pruneCNPGRuntimeMemoLocked(time.Now())
+ }
+ if len(cnpgRuntimeMemoEntries) < cnpgRuntimeMemoMaxEntries {
+ cnpgRuntimeMemoEntries[key] = cnpgRuntimeMemoEntry{value: got, expires: time.Now().Add(ttl)}
+ }
+ cnpgRuntimeMemoMu.Unlock()
+ }
+ return got, nil
+ })
+ return v.(T)
+}
+
+func pruneCNPGRuntimeMemoLocked(now time.Time) {
+ for k, e := range cnpgRuntimeMemoEntries {
+ if !now.Before(e.expires) {
+ delete(cnpgRuntimeMemoEntries, k)
+ }
+ }
+}
+
+// classifyCNPGProxyFailure classifies a failed read and, for a failure in
+// transit, replaces the raw error with a sentence; the raw error is logged.
+// cnpgReadTimedOutKey carries a flag a memoized read sets when any request
+// inside it timed out, so the memo does not keep a timeout for its full TTL.
+type cnpgReadTimedOutKey struct{}
+
+func cnpgMarkTimedOut(ctx context.Context, err error) {
+ flag, _ := ctx.Value(cnpgReadTimedOutKey{}).(*atomic.Bool)
+ if flag == nil {
+ return
+ }
+ var netErr net.Error
+ if errors.Is(err, context.DeadlineExceeded) || errors.Is(ctx.Err(), context.DeadlineExceeded) ||
+ (errors.As(err, &netErr) && netErr.Timeout()) || apierrors.IsTimeout(err) || apierrors.IsServerTimeout(err) ||
+ cnpgRelayedTransportTimeout(err) {
+ flag.Store(true)
+ }
+}
diff --git a/internal/cnpg/metrics.go b/internal/cnpg/metrics.go
new file mode 100644
index 0000000000..91a75d65e2
--- /dev/null
+++ b/internal/cnpg/metrics.go
@@ -0,0 +1,31 @@
+package cnpg
+
+import (
+ "context"
+ "time"
+
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+type Metrics struct {
+ Connection func(context.Context) (bool, error)
+ PVCUsage func(context.Context, string, []string, []prom.WorkloadPodIdentity) prometheuspkg.PVCUsageBatch
+ CNPGScope func(context.Context, string, string, []prom.WorkloadPodIdentity, time.Duration) (string, prometheuspkg.SeriesIsolation, error)
+ PVCScope func(context.Context, string, []string, []prom.WorkloadPodIdentity, time.Duration) (string, prometheuspkg.SeriesIsolation, error)
+ History func(context.Context, prometheuspkg.CNPGHistoryRequest) ([]prometheuspkg.CNPGHistoryChart, error)
+ FleetLag func(context.Context, string, []string, string) (prometheuspkg.CNPGFleetLag, error)
+ FleetSlots func(context.Context, string, []string, string) (prometheuspkg.CNPGFleetSlots, error)
+ DiskGrowth func(context.Context, string, []string, time.Duration, string) (map[string]float64, error)
+}
+
+func (s *Reader) prometheusUnavailable(ctx context.Context) string {
+ connected, err := s.Metrics.Connection(ctx)
+ if !connected {
+ return cnpgNoPrometheusReason("")
+ }
+ if err != nil {
+ return cnpgNoPrometheusReason(err.Error())
+ }
+ return ""
+}
diff --git a/internal/cnpg/operator.go b/internal/cnpg/operator.go
new file mode 100644
index 0000000000..8b1e1cadac
--- /dev/null
+++ b/internal/cnpg/operator.go
@@ -0,0 +1,401 @@
+package cnpg
+
+import (
+ "context"
+ "log"
+ "net/http"
+ "sort"
+ "strings"
+
+ appsv1 "k8s.io/api/apps/v1"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/labels"
+
+ "github.com/skyhook-io/radar/internal/imageutil"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+const (
+ cnpgOperatorNameLabel = "app.kubernetes.io/name"
+ cnpgOperatorNameValue = "cloudnative-pg"
+ cnpgVersionLabel = "app.kubernetes.io/version"
+ cnpgPluginNameLabel = "cnpg.io/pluginName"
+ cnpgOperatorContainer = "manager"
+ cnpgOperatorDeployVar = "OPERATOR_DEPLOYMENT_NAME"
+ cnpgMonitoringQueriesCM = "MONITORING_QUERIES_CONFIGMAP"
+
+ cnpgOperatorRoleOperator = "operator"
+ cnpgOperatorRolePlugin = "plugin"
+
+ cnpgConfigPurposeOperator = "operator"
+ cnpgConfigPurposeMonitoring = "monitoring"
+)
+
+// CNPGOperatorComponent is one operator or plugin Deployment. Version is the
+// image tag, else the app.kubernetes.io/version label, else empty. Replica
+// counts are nil when unreported, which is not zero. Pods is null unless
+// PodCoverage is ok; PodCoverage is absent for a plugin Service with no
+// matching Deployment.
+type CNPGOperatorComponent struct {
+ Role string `json:"role"`
+ PluginName string `json:"pluginName,omitempty"`
+ Namespace string `json:"namespace"`
+ Deployment string `json:"deployment"`
+ Image string `json:"image"`
+ Version string `json:"version"`
+ ReadyReplicas *int32 `json:"readyReplicas"`
+ Replicas *int32 `json:"replicas"`
+ Pods []CNPGOperatorComponentPod `json:"pods"`
+ PodCoverage *integration.ReadSource `json:"podCoverage,omitempty"`
+}
+
+// CNPGOperatorComponentPod is one of a component's Pods. StartedAt is when
+// the component's container last started, empty while it is not running;
+// Restarts is summed over the Pod's containers.
+type CNPGOperatorComponentPod struct {
+ Name string `json:"name"`
+ Ready bool `json:"ready"`
+ Restarts int32 `json:"restarts"`
+ StartedAt string `json:"startedAt,omitempty"`
+ LastTermination *CNPGContainerTermination `json:"lastTermination,omitempty"`
+}
+
+// CNPGOperatorConfigMapState is present only on ConfigMap references. A Secret
+// reference never carries it: the endpoint never reads Secrets.
+type CNPGOperatorConfigMapState struct {
+ Exists *bool `json:"exists"`
+ Readable bool `json:"readable"`
+ Reason string `json:"reason,omitempty"`
+ Data map[string]string `json:"data"`
+}
+
+// CNPGOperatorConfigRef is a ConfigMap or Secret the operator is configured
+// to read.
+type CNPGOperatorConfigRef struct {
+ Kind string `json:"kind"`
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ Purpose string `json:"purpose"`
+ *CNPGOperatorConfigMapState
+}
+
+// CNPGOperatorResponse is GET /api/cnpg/operator.
+type CNPGOperatorResponse struct {
+ Coverage map[string]integration.KindCoverage `json:"coverage"`
+ Components []CNPGOperatorComponent `json:"components"`
+ Config []CNPGOperatorConfigRef `json:"config"`
+ // Diagnosis is one entry per operator Deployment: leader Lease, watched
+ // namespaces, webhook reachability, reconcile counters and recent events.
+ Diagnosis []CNPGOperatorDiagnosis `json:"diagnosis"`
+}
+
+func (s *Reader) Operator(ctx context.Context) (*CNPGOperatorResponse, error) {
+ if !s.Observations.Connected {
+ return nil, ErrCNPGDisconnected
+ }
+ cache := s.Observations.Cache
+ if cache == nil {
+ return nil, &ReadFailure{http.StatusServiceUnavailable, "Resource cache not available"}
+ }
+ scope := s.Observations.OperatorScope(ctx)
+ resp := CNPGOperatorResponse{
+ Coverage: map[string]integration.KindCoverage{},
+ Components: []CNPGOperatorComponent{},
+ Config: []CNPGOperatorConfigRef{},
+ }
+
+ depAcc, deployments := s.operatorDeployments(ctx, cache, scope)
+ resp.Coverage["deployments"] = depAcc.Coverage()
+ svcAcc, services := s.operatorServices(ctx, cache, scope)
+ resp.Coverage["services"] = svcAcc.Coverage()
+
+ typed := s.Clients.Typed
+ podReads := map[*appsv1.Deployment]cnpgDeploymentPodRead{}
+ podsOf := func(d *appsv1.Deployment) cnpgDeploymentPodRead {
+ if read, ok := podReads[d]; ok {
+ return read
+ }
+ read := s.deploymentPods(ctx, typed, d)
+ podReads[d] = read
+ return read
+ }
+ component := func(d *appsv1.Deployment, role, pluginName string, c *corev1.Container) CNPGOperatorComponent {
+ return withCNPGComponentPods(cnpgOperatorComponent(d, role, pluginName, c), podsOf(d), c)
+ }
+
+ var operators []*appsv1.Deployment
+ for _, d := range deployments {
+ if d.Labels[cnpgOperatorNameLabel] == cnpgOperatorNameValue {
+ operators = append(operators, d)
+ resp.Components = append(resp.Components, component(d, cnpgOperatorRoleOperator, "", cnpgOperatorContainerOf(d)))
+ }
+ }
+
+ byNamespace := map[string][]*appsv1.Deployment{}
+ for _, d := range deployments {
+ byNamespace[d.Namespace] = append(byNamespace[d.Namespace], d)
+ }
+ var plugins []CNPGOperatorComponent
+ for _, svc := range services {
+ pluginName := svc.Labels[cnpgPluginNameLabel]
+ if pluginName == "" || !depAcc.Covers(svc.Namespace) {
+ continue
+ }
+ matched := false
+ if len(svc.Spec.Selector) > 0 {
+ sel := labels.SelectorFromSet(svc.Spec.Selector)
+ for _, d := range byNamespace[svc.Namespace] {
+ if sel.Matches(labels.Set(d.Spec.Template.Labels)) {
+ matched = true
+ plugins = append(plugins, component(d, cnpgOperatorRolePlugin, pluginName, firstContainer(d)))
+ }
+ }
+ }
+ if !matched {
+ plugins = append(plugins, CNPGOperatorComponent{Role: cnpgOperatorRolePlugin, PluginName: pluginName, Namespace: svc.Namespace})
+ }
+ }
+ sort.SliceStable(plugins, func(i, j int) bool {
+ a, b := plugins[i], plugins[j]
+ if a.PluginName != b.PluginName {
+ return a.PluginName < b.PluginName
+ }
+ if a.Namespace != b.Namespace {
+ return a.Namespace < b.Namespace
+ }
+ return a.Deployment < b.Deployment
+ })
+ resp.Components = append(resp.Components, plugins...)
+
+ resp.Config = s.operatorConfig(ctx, cache, operators)
+ resp.Diagnosis = s.operatorDiagnoses(ctx, typed, operators, podsOf)
+ return &resp, nil
+}
+
+// withCNPGComponentPods adds a component's Pods: restarts and the last
+// termination show a crash-looping operator or plugin that its Deployment's
+// ready count, read between crashes, can hide.
+func withCNPGComponentPods(comp CNPGOperatorComponent, read cnpgDeploymentPodRead, c *corev1.Container) CNPGOperatorComponent {
+ coverage := read.coverage
+ comp.PodCoverage = &coverage
+ if coverage.State != cnpgReadOK {
+ return comp
+ }
+ comp.Pods = make([]CNPGOperatorComponentPod, 0, len(read.pods))
+ for i := range read.pods {
+ p := &read.pods[i]
+ pod := CNPGOperatorComponentPod{Name: p.Name, Ready: cnpgActionPodReady(p)}
+ if c != nil {
+ pod.StartedAt = cnpgOperatorProcessStart(p, c.Name)
+ }
+ pod.Restarts, pod.LastTermination = cnpgPodRestarts(p)
+ comp.Pods = append(comp.Pods, pod)
+ }
+ return comp
+}
+
+func (s *Reader) operatorDeployments(ctx context.Context, cache *k8s.ResourceCache, scope []string) (integration.KindAccess, []*appsv1.Deployment) {
+ acc, read := s.Observations.TypedScope(ctx, cache, scope, "apps", "deployments")
+ if acc.State == integration.KindCoverageDenied || acc.State == integration.KindCoverageUncached {
+ return acc, nil
+ }
+ lister := cache.Deployments()
+ if lister == nil || !cache.IsKindReady("deployments") {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ var out []*appsv1.Deployment
+ if read == nil {
+ out, _ = lister.List(labels.Everything())
+ } else {
+ for _, ns := range read {
+ items, _ := lister.Deployments(ns).List(labels.Everything())
+ out = append(out, items...)
+ }
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if out[i].Namespace != out[j].Namespace {
+ return out[i].Namespace < out[j].Namespace
+ }
+ return out[i].Name < out[j].Name
+ })
+ return acc, out
+}
+
+func (s *Reader) operatorServices(ctx context.Context, cache *k8s.ResourceCache, scope []string) (integration.KindAccess, []*corev1.Service) {
+ acc, read := s.Observations.TypedScope(ctx, cache, scope, "", "services")
+ if acc.State == integration.KindCoverageDenied || acc.State == integration.KindCoverageUncached {
+ return acc, nil
+ }
+ lister := cache.Services()
+ if lister == nil || !cache.IsKindReady("services") {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ hasPlugin, err := labels.Parse(cnpgPluginNameLabel)
+ if err != nil {
+ log.Printf("[cnpg] Failed to build plugin selector: %v", err)
+ return integration.KindAccess{State: integration.KindCoverageError}, nil
+ }
+ var out []*corev1.Service
+ if read == nil {
+ out, _ = lister.List(hasPlugin)
+ } else {
+ for _, ns := range read {
+ items, _ := lister.Services(ns).List(hasPlugin)
+ out = append(out, items...)
+ }
+ }
+ return acc, out
+}
+
+func cnpgOperatorContainerOf(d *appsv1.Deployment) *corev1.Container {
+ for i := range d.Spec.Template.Spec.Containers {
+ if d.Spec.Template.Spec.Containers[i].Name == cnpgOperatorContainer {
+ return &d.Spec.Template.Spec.Containers[i]
+ }
+ }
+ return firstContainer(d)
+}
+
+func firstContainer(d *appsv1.Deployment) *corev1.Container {
+ if len(d.Spec.Template.Spec.Containers) == 0 {
+ return nil
+ }
+ return &d.Spec.Template.Spec.Containers[0]
+}
+
+func cnpgOperatorComponent(d *appsv1.Deployment, role, pluginName string, c *corev1.Container) CNPGOperatorComponent {
+ out := CNPGOperatorComponent{
+ Role: role,
+ PluginName: pluginName,
+ Namespace: d.Namespace,
+ Deployment: d.Name,
+ Replicas: d.Spec.Replicas,
+ }
+ if c != nil {
+ out.Image = c.Image
+ out.Version = imageutil.ImageTag(c.Image)
+ }
+ if out.Version == "" {
+ out.Version = d.Labels[cnpgVersionLabel]
+ }
+ if out.Version == "" {
+ out.Version = d.Spec.Template.Labels[cnpgVersionLabel]
+ }
+ // The typed status cannot tell an omitted readyReplicas from zero; a status
+ // the controller has observed at least once states it authoritatively.
+ if d.Status.ObservedGeneration > 0 {
+ ready := d.Status.ReadyReplicas
+ out.ReadyReplicas = &ready
+ }
+ return out
+}
+
+// cnpgOperatorArg returns the value of --flag=value or --flag value from a
+// container's command and args.
+func cnpgOperatorArg(c *corev1.Container, flag string) string {
+ argv := append(append([]string{}, c.Command...), c.Args...)
+ for i, a := range argv {
+ if v, ok := strings.CutPrefix(a, flag+"="); ok {
+ return v
+ }
+ if a == flag && i+1 < len(argv) {
+ return argv[i+1]
+ }
+ }
+ return ""
+}
+
+func cnpgOperatorEnv(c *corev1.Container, name string) string {
+ for _, e := range c.Env {
+ if e.Name == name && e.ValueFrom == nil {
+ return e.Value
+ }
+ }
+ return ""
+}
+
+// cnpgOperatorExpand resolves $(OPERATOR_DEPLOYMENT_NAME) the way the kubelet
+// would: from the container's literal env, which the shipped manifests set to
+// the Deployment's own name. Any other reference is left verbatim, as the
+// kubelet leaves an unresolvable one.
+func cnpgOperatorExpand(v string, c *corev1.Container, d *appsv1.Deployment) string {
+ ref := "$(" + cnpgOperatorDeployVar + ")"
+ if !strings.Contains(v, ref) {
+ return v
+ }
+ name := cnpgOperatorEnv(c, cnpgOperatorDeployVar)
+ if name == "" {
+ name = d.Name
+ }
+ return strings.ReplaceAll(v, ref, name)
+}
+
+func (s *Reader) operatorConfig(ctx context.Context, cache *k8s.ResourceCache, operators []*appsv1.Deployment) []CNPGOperatorConfigRef {
+ out := []CNPGOperatorConfigRef{}
+ seen := map[string]bool{}
+ add := func(ref CNPGOperatorConfigRef) {
+ key := ref.Kind + "\x00" + ref.Namespace + "\x00" + ref.Name + "\x00" + ref.Purpose
+ if ref.Name == "" || seen[key] {
+ return
+ }
+ seen[key] = true
+ out = append(out, ref)
+ }
+ for _, d := range operators {
+ c := cnpgOperatorContainerOf(d)
+ if c == nil {
+ continue
+ }
+ if name := cnpgOperatorExpand(cnpgOperatorArg(c, "--config-map-name"), c, d); name != "" {
+ add(s.operatorConfigMap(ctx, cache, d.Namespace, name, cnpgConfigPurposeOperator))
+ }
+ if name := cnpgOperatorExpand(cnpgOperatorArg(c, "--secret-name"), c, d); name != "" {
+ add(CNPGOperatorConfigRef{Kind: "Secret", Namespace: d.Namespace, Name: name, Purpose: cnpgConfigPurposeOperator})
+ }
+ if name := cnpgOperatorEnv(c, cnpgMonitoringQueriesCM); name != "" {
+ add(s.operatorConfigMap(ctx, cache, d.Namespace, name, cnpgConfigPurposeMonitoring))
+ }
+ }
+ return out
+}
+
+func (s *Reader) operatorConfigMap(ctx context.Context, cache *k8s.ResourceCache, namespace, name, purpose string) CNPGOperatorConfigRef {
+ ref := CNPGOperatorConfigRef{Kind: "ConfigMap", Namespace: namespace, Name: name, Purpose: purpose}
+ state := &CNPGOperatorConfigMapState{}
+ ref.CNPGOperatorConfigMapState = state
+ if !s.Access.CanRead(ctx, "", "configmaps", namespace, "get") {
+ state.Reason = "no permission to get ConfigMaps in " + namespace
+ return ref
+ }
+ lister := cache.ConfigMaps()
+ if lister == nil {
+ state.Reason = "ConfigMaps are still loading"
+ return ref
+ }
+ if !integration.CacheCoversNamespace(cache, "configmaps", namespace) {
+ state.Reason = "Radar does not watch ConfigMaps in " + namespace
+ return ref
+ }
+ cm, err := lister.ConfigMaps(namespace).Get(name)
+ switch {
+ case apierrors.IsNotFound(err):
+ exists := false
+ state.Exists = &exists
+ state.Reason = "not found"
+ return ref
+ case err != nil:
+ log.Printf("[cnpg] Failed to read ConfigMap %s/%s: %v", namespace, name, err)
+ state.Reason = "could not read the ConfigMap"
+ return ref
+ }
+ exists := true
+ state.Exists = &exists
+ state.Readable = true
+ state.Data = map[string]string{}
+ for k, v := range cm.Data {
+ state.Data[k] = v
+ }
+ return ref
+}
diff --git a/internal/cnpg/operator_diagnosis.go b/internal/cnpg/operator_diagnosis.go
new file mode 100644
index 0000000000..542cd0101e
--- /dev/null
+++ b/internal/cnpg/operator_diagnosis.go
@@ -0,0 +1,615 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "sort"
+ "strconv"
+ "strings"
+ "time"
+
+ admissionv1 "k8s.io/api/admissionregistration/v1"
+ appsv1 "k8s.io/api/apps/v1"
+ coordinationv1 "k8s.io/api/coordination/v1"
+ corev1 "k8s.io/api/core/v1"
+ discoveryv1 "k8s.io/api/discovery/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+const (
+ // Constants in CloudNativePG's controller, not configurable; the leader
+ // Lease name (cnpgOperatorLeaseName) is one too.
+ cnpgMutatingWebhookConfig = "cnpg-mutating-webhook-configuration"
+ cnpgValidatingWebhookConfig = "cnpg-validating-webhook-configuration"
+ cnpgOperatorDefaultMetrics = 8080
+ cnpgOperatorMetricsPath = "/metrics"
+ cnpgOperatorMetricsCap = 4 << 20
+ cnpgOperatorMetricsTTL = 25 * time.Second
+ cnpgOperatorEventLimit = 20
+ cnpgWatchNamespaceEnv = "WATCH_NAMESPACE"
+ cnpgEndpointSliceServiceLabel = "kubernetes.io/service-name"
+)
+
+var (
+ cnpgGrantGetLeases = auth.Grant{Verb: "get", Group: "coordination.k8s.io", Resource: "leases"}
+ cnpgGrantListLeases = auth.Grant{Verb: "list", Group: "coordination.k8s.io", Resource: "leases"}
+ cnpgGrantGetMutatingWH = auth.Grant{Verb: "get", Group: "admissionregistration.k8s.io", Resource: "mutatingwebhookconfigurations"}
+ cnpgGrantGetValidatingWH = auth.Grant{Verb: "get", Group: "admissionregistration.k8s.io", Resource: "validatingwebhookconfigurations"}
+ cnpgGrantListEndpointSlcs = auth.Grant{Verb: "list", Group: "discovery.k8s.io", Resource: "endpointslices"}
+ grantGetPodsProxy = auth.Grant{Verb: "get", Resource: "pods", Subresource: "proxy"}
+)
+
+type CNPGOperatorPod struct {
+ Name string `json:"name"`
+ UID string `json:"uid"`
+ Phase string `json:"phase"`
+ Ready bool `json:"ready"`
+ StartedAt string `json:"startedAt,omitempty"`
+ Restarts int32 `json:"restarts"`
+ LastTermination *CNPGContainerTermination `json:"lastTermination,omitempty"`
+ Leader bool `json:"leader"`
+}
+
+// CNPGContainerTermination is how a container's previous run ended
+// (lastState.terminated), for the Pod's container that ended most recently.
+type CNPGContainerTermination struct {
+ Container string `json:"container"`
+ Reason string `json:"reason"`
+ ExitCode int32 `json:"exitCode"`
+ FinishedAt string `json:"finishedAt,omitempty"`
+}
+
+// CNPGOperatorLeader: State ok | disabled (no --leader-elect) | notFound |
+// denied | error. Stale means the holder has not renewed within the lease
+// duration, so no operator instance is leading.
+type CNPGOperatorLeader struct {
+ integration.ReadSource
+ Lease string `json:"lease,omitempty"`
+ Holder string `json:"holder,omitempty"`
+ HolderPod string `json:"holderPod,omitempty"`
+ HolderIsCurrentPod bool `json:"holderIsCurrentPod"`
+ RenewTime string `json:"renewTime,omitempty"`
+ AcquireTime string `json:"acquireTime,omitempty"`
+ LeaseDurationSeconds *int32 `json:"leaseDurationSeconds,omitempty"`
+ Transitions *int32 `json:"transitions,omitempty"`
+ Stale bool `json:"stale"`
+}
+
+// CNPGOperatorWatch: All when WATCH_NAMESPACE is unset or empty. Source names
+// where the value came from; Unresolved is set when it comes from somewhere
+// Radar does not follow (a ConfigMap or Secret key reference).
+type CNPGOperatorWatch struct {
+ All bool `json:"all"`
+ Namespaces []string `json:"namespaces"`
+ Source string `json:"source"`
+ Unresolved string `json:"unresolved,omitempty"`
+}
+
+type CNPGOperatorWebhook struct {
+ Name string `json:"name"`
+ FailurePolicy string `json:"failurePolicy"`
+ CABundleSet bool `json:"caBundleSet"`
+ Service string `json:"service,omitempty"`
+ URL bool `json:"url"`
+}
+
+type CNPGOperatorWebhookConfig struct {
+ Kind string `json:"kind"`
+ Name string `json:"name"`
+ integration.ReadSource
+ Webhooks []CNPGOperatorWebhook `json:"webhooks"`
+}
+
+// CNPGOperatorWebhookService: endpoint counts are nil when the slices could
+// not be read.
+type CNPGOperatorWebhookService struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ integration.ReadSource
+ ReadyEndpoints *int `json:"readyEndpoints"`
+ NotReadyEndpoints *int `json:"notReadyEndpoints"`
+}
+
+type CNPGOperatorControllerStats struct {
+ Controller string `json:"controller"`
+ Errors *float64 `json:"errors"`
+ Total *float64 `json:"total"`
+ Results map[string]float64 `json:"results"`
+}
+
+// CNPGOperatorReconcilePod is one operator Pod's controller-runtime counters,
+// cumulative since that Pod's process started (StartedAt).
+type CNPGOperatorReconcilePod struct {
+ Pod string `json:"pod"`
+ Leader bool `json:"leader"`
+ StartedAt string `json:"startedAt,omitempty"`
+ CNPGRuntimeSource
+ Controllers []CNPGOperatorControllerStats `json:"controllers"`
+}
+
+type CNPGOperatorEvents struct {
+ integration.ReadSource
+ Items []CNPGRecoveryEvent `json:"items"`
+}
+
+type CNPGOperatorDiagnosis struct {
+ Namespace string `json:"namespace"`
+ Deployment string `json:"deployment"`
+ Pods []CNPGOperatorPod `json:"pods"`
+ PodCoverage integration.ReadSource `json:"podCoverage"`
+ Leader CNPGOperatorLeader `json:"leader"`
+ Watch CNPGOperatorWatch `json:"watch"`
+ Webhooks []CNPGOperatorWebhookConfig `json:"webhooks"`
+ Services []CNPGOperatorWebhookService `json:"webhookServices"`
+ MetricsPort int `json:"metricsPort"`
+ Reconcile []CNPGOperatorReconcilePod `json:"reconcile"`
+ Events CNPGOperatorEvents `json:"events"`
+}
+
+// cnpgOperatorDiagnoses takes podsOf so the Pods each operator Deployment's
+// component already read are not listed again.
+func (s *Reader) operatorDiagnoses(ctx context.Context, typed kubernetes.Interface, operators []*appsv1.Deployment, podsOf func(*appsv1.Deployment) cnpgDeploymentPodRead) []CNPGOperatorDiagnosis {
+ out := []CNPGOperatorDiagnosis{}
+ if typed == nil {
+ return out
+ }
+ webhooks, services := s.operatorWebhooks(ctx, typed)
+ for _, d := range operators {
+ diag := CNPGOperatorDiagnosis{Namespace: d.Namespace, Deployment: d.Name, Webhooks: webhooks, Services: services, Pods: []CNPGOperatorPod{}, Reconcile: []CNPGOperatorReconcilePod{}}
+ c := cnpgOperatorContainerOf(d)
+ diag.Watch = cnpgOperatorWatchOf(c, d.Namespace)
+ diag.MetricsPort = cnpgOperatorMetricsPort(c)
+ read := podsOf(d)
+ pods := read.pods
+ diag.PodCoverage = read.coverage
+ for i := range pods {
+ diag.Pods = append(diag.Pods, cnpgOperatorPodOf(&pods[i]))
+ }
+ diag.Leader = s.operatorLeader(ctx, typed, d, c, pods)
+ leaderPod := cnpgOperatorLeadingPod(diag.Leader)
+ for i := range diag.Pods {
+ diag.Pods[i].Leader = leaderPod != "" && diag.Pods[i].Name == leaderPod
+ }
+ diag.Reconcile = s.operatorReconcile(ctx, d, pods, diag.MetricsPort, leaderPod)
+ diag.Events = s.operatorEvents(ctx, typed, d, pods)
+ out = append(out, diag)
+ }
+ return out
+}
+
+// cnpgOperatorLeadingPod is the current Pod that leads, or "". A holder
+// whose lease expired leads nothing, whatever the Lease still names.
+func cnpgOperatorLeadingPod(l CNPGOperatorLeader) string {
+ if l.State != cnpgReadOK || l.Stale || !l.HolderIsCurrentPod {
+ return ""
+ }
+ return l.HolderPod
+}
+
+// cnpgDeploymentPodRead is the outcome of listing one Deployment's Pods.
+type cnpgDeploymentPodRead struct {
+ pods []corev1.Pod
+ coverage integration.ReadSource
+}
+
+// cnpgDeploymentPods lists the Pods a Deployment's selector matches, as the
+// caller, sorted by name.
+func (s *Reader) deploymentPods(ctx context.Context, typed kubernetes.Interface, d *appsv1.Deployment) cnpgDeploymentPodRead {
+ if d.Spec.Selector == nil {
+ return cnpgDeploymentPodRead{coverage: integration.ReadSource{State: cnpgReadError, Reason: "the Deployment has no selector"}}
+ }
+ selector, err := metav1.LabelSelectorAsSelector(d.Spec.Selector)
+ if err != nil {
+ return cnpgDeploymentPodRead{coverage: integration.ReadSource{State: cnpgReadError, Reason: err.Error()}}
+ }
+ var out cnpgDeploymentPodRead
+ out.coverage = s.gatedRead(ctx, GrantListPods, d.Namespace, func() error {
+ if typed == nil {
+ return errors.New("cluster client unavailable")
+ }
+ list, err := typed.CoreV1().Pods(d.Namespace).List(ctx, metav1.ListOptions{LabelSelector: selector.String()})
+ if err == nil {
+ out.pods = list.Items
+ }
+ return err
+ })
+ sort.Slice(out.pods, func(i, j int) bool { return out.pods[i].Name < out.pods[j].Name })
+ return out
+}
+
+func cnpgOperatorPodOf(p *corev1.Pod) CNPGOperatorPod {
+ op := CNPGOperatorPod{Name: p.Name, UID: string(p.UID), Phase: string(p.Status.Phase), Ready: cnpgActionPodReady(p)}
+ if p.Status.StartTime != nil {
+ op.StartedAt = p.Status.StartTime.UTC().Format(time.RFC3339)
+ }
+ op.Restarts, op.LastTermination = cnpgPodRestarts(p)
+ return op
+}
+
+// cnpgPodRestarts sums restarts over a Pod's containers and returns how the
+// container that ended most recently ended last time. Restart counts
+// accumulate over the Pod's life while each container keeps only its
+// previous run, so the newest termination is the one that explains a
+// current crash loop.
+func cnpgPodRestarts(p *corev1.Pod) (int32, *CNPGContainerTermination) {
+ var restarts int32
+ var newest *corev1.ContainerStateTerminated
+ var container string
+ for _, st := range p.Status.ContainerStatuses {
+ restarts += st.RestartCount
+ t := st.LastTerminationState.Terminated
+ if t != nil && (newest == nil || t.FinishedAt.After(newest.FinishedAt.Time)) {
+ newest, container = t, st.Name
+ }
+ }
+ if newest == nil {
+ return restarts, nil
+ }
+ out := &CNPGContainerTermination{Container: container, Reason: newest.Reason, ExitCode: newest.ExitCode}
+ if !newest.FinishedAt.IsZero() {
+ out.FinishedAt = newest.FinishedAt.UTC().Format(time.RFC3339)
+ }
+ return restarts, out
+}
+
+func cnpgOperatorWatchOf(c *corev1.Container, operatorNamespace string) CNPGOperatorWatch {
+ out := CNPGOperatorWatch{All: true, Namespaces: []string{}, Source: "WATCH_NAMESPACE is not set, so the operator watches every namespace"}
+ if c == nil {
+ return out
+ }
+ for _, e := range c.Env {
+ if e.Name != cnpgWatchNamespaceEnv {
+ continue
+ }
+ switch {
+ case e.ValueFrom == nil:
+ out.Source = "env WATCH_NAMESPACE"
+ for _, ns := range strings.Split(e.Value, ",") {
+ if ns = strings.TrimSpace(ns); ns != "" {
+ out.Namespaces = append(out.Namespaces, ns)
+ }
+ }
+ out.All = len(out.Namespaces) == 0
+ case e.ValueFrom.FieldRef != nil && e.ValueFrom.FieldRef.FieldPath == "metadata.namespace":
+ out.Source = "env WATCH_NAMESPACE from the Pod's namespace"
+ out.All, out.Namespaces = false, []string{operatorNamespace}
+ default:
+ out.All = false
+ out.Source = "env WATCH_NAMESPACE"
+ switch {
+ case e.ValueFrom.ConfigMapKeyRef != nil:
+ out.Unresolved = fmt.Sprintf("set from ConfigMap %s key %s", e.ValueFrom.ConfigMapKeyRef.Name, e.ValueFrom.ConfigMapKeyRef.Key)
+ case e.ValueFrom.SecretKeyRef != nil:
+ out.Unresolved = fmt.Sprintf("set from Secret %s key %s", e.ValueFrom.SecretKeyRef.Name, e.ValueFrom.SecretKeyRef.Key)
+ default:
+ out.Unresolved = "set from a field reference Radar does not resolve"
+ }
+ }
+ }
+ return out
+}
+
+func cnpgOperatorMetricsPort(c *corev1.Container) int {
+ if c == nil {
+ return cnpgOperatorDefaultMetrics
+ }
+ for _, p := range c.Ports {
+ if p.Name == "metrics" && p.ContainerPort > 0 {
+ return int(p.ContainerPort)
+ }
+ }
+ if v := cnpgOperatorArg(c, "--metrics-bind-address"); v != "" {
+ if i := strings.LastIndexByte(v, ':'); i >= 0 {
+ if n, err := strconv.Atoi(v[i+1:]); err == nil && n > 0 {
+ return n
+ }
+ }
+ }
+ return cnpgOperatorDefaultMetrics
+}
+
+func cnpgOperatorLeaderElect(c *corev1.Container) bool {
+ if c == nil {
+ return false
+ }
+ for _, a := range append(append([]string{}, c.Command...), c.Args...) {
+ if a == "--leader-elect" || a == "--leader-elect=true" {
+ return true
+ }
+ }
+ return false
+}
+
+// controller-runtime's holder identity is "_", and a Pod's
+// hostname is its name.
+func cnpgLeaseHolderPod(holder string) string {
+ if i := strings.LastIndexByte(holder, '_'); i > 0 {
+ return holder[:i]
+ }
+ return holder
+}
+
+func (s *Reader) operatorLeader(ctx context.Context, typed kubernetes.Interface, d *appsv1.Deployment, c *corev1.Container, pods []corev1.Pod) CNPGOperatorLeader {
+ out := CNPGOperatorLeader{Lease: cnpgOperatorLeaseName}
+ if !cnpgOperatorLeaderElect(c) {
+ out.State = "disabled"
+ out.Reason = "the operator runs without --leader-elect"
+ return out
+ }
+ var lease *coordinationv1.Lease
+ out.ReadSource = s.gatedRead(ctx, cnpgGrantGetLeases, d.Namespace, func() error {
+ l, err := typed.CoordinationV1().Leases(d.Namespace).Get(ctx, cnpgOperatorLeaseName, metav1.GetOptions{})
+ lease = l
+ return err
+ })
+ if out.State != cnpgReadOK {
+ return out
+ }
+ spec := lease.Spec
+ if spec.HolderIdentity != nil {
+ out.Holder = *spec.HolderIdentity
+ out.HolderPod = cnpgLeaseHolderPod(out.Holder)
+ }
+ if spec.RenewTime != nil {
+ out.RenewTime = spec.RenewTime.UTC().Format(time.RFC3339)
+ }
+ if spec.AcquireTime != nil {
+ out.AcquireTime = spec.AcquireTime.UTC().Format(time.RFC3339)
+ }
+ out.LeaseDurationSeconds = spec.LeaseDurationSeconds
+ out.Transitions = spec.LeaseTransitions
+ for _, p := range pods {
+ if p.Name == out.HolderPod {
+ out.HolderIsCurrentPod = true
+ }
+ }
+ view := &leaseView{namespace: d.Namespace, name: cnpgOperatorLeaseName, holder: out.Holder, duration: spec.LeaseDurationSeconds}
+ if spec.RenewTime != nil {
+ t := spec.RenewTime.Time
+ view.renew = &t
+ }
+ if judged := cnpgHALeaseFrom(view, time.Now()); judged.Expired != nil {
+ out.Stale = *judged.Expired
+ }
+ if out.Holder == "" {
+ out.Stale = true
+ }
+ return out
+}
+
+func (s *Reader) operatorWebhooks(ctx context.Context, typed kubernetes.Interface) ([]CNPGOperatorWebhookConfig, []CNPGOperatorWebhookService) {
+ configs := []CNPGOperatorWebhookConfig{}
+ type svcKey struct{ ns, name string }
+ svcs := map[svcKey]bool{}
+ add := func(kind, name string, g auth.Grant, read func() ([]admissionv1.WebhookClientConfig, []string, []string, error)) {
+ cfg := CNPGOperatorWebhookConfig{Kind: kind, Name: name, Webhooks: []CNPGOperatorWebhook{}}
+ if s.Access.Permission(ctx, g) == integration.PermissionDenied {
+ cfg.ReadSource = integration.ReadSource{State: cnpgReadDenied, Grant: g.Ref()}
+ configs = append(configs, cfg)
+ return
+ }
+ clients, names, policies, err := read()
+ cfg.ReadSource = cnpgReadOutcome(err, g, "")
+ for i, cc := range clients {
+ wh := CNPGOperatorWebhook{Name: names[i], FailurePolicy: policies[i], CABundleSet: len(cc.CABundle) > 0, URL: cc.URL != nil}
+ if cc.Service != nil {
+ wh.Service = cc.Service.Namespace + "/" + cc.Service.Name
+ svcs[svcKey{cc.Service.Namespace, cc.Service.Name}] = true
+ }
+ cfg.Webhooks = append(cfg.Webhooks, wh)
+ }
+ configs = append(configs, cfg)
+ }
+ policy := func(p *admissionv1.FailurePolicyType) string {
+ if p == nil {
+ return string(admissionv1.Fail)
+ }
+ return string(*p)
+ }
+ add("MutatingWebhookConfiguration", cnpgMutatingWebhookConfig, cnpgGrantGetMutatingWH, func() ([]admissionv1.WebhookClientConfig, []string, []string, error) {
+ c, err := typed.AdmissionregistrationV1().MutatingWebhookConfigurations().Get(ctx, cnpgMutatingWebhookConfig, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, nil, err
+ }
+ var cc []admissionv1.WebhookClientConfig
+ var names, pols []string
+ for _, w := range c.Webhooks {
+ cc, names, pols = append(cc, w.ClientConfig), append(names, w.Name), append(pols, policy(w.FailurePolicy))
+ }
+ return cc, names, pols, nil
+ })
+ add("ValidatingWebhookConfiguration", cnpgValidatingWebhookConfig, cnpgGrantGetValidatingWH, func() ([]admissionv1.WebhookClientConfig, []string, []string, error) {
+ c, err := typed.AdmissionregistrationV1().ValidatingWebhookConfigurations().Get(ctx, cnpgValidatingWebhookConfig, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, nil, err
+ }
+ var cc []admissionv1.WebhookClientConfig
+ var names, pols []string
+ for _, w := range c.Webhooks {
+ cc, names, pols = append(cc, w.ClientConfig), append(names, w.Name), append(pols, policy(w.FailurePolicy))
+ }
+ return cc, names, pols, nil
+ })
+
+ keys := make([]svcKey, 0, len(svcs))
+ for k := range svcs {
+ keys = append(keys, k)
+ }
+ sort.Slice(keys, func(i, j int) bool { return keys[i].ns+"/"+keys[i].name < keys[j].ns+"/"+keys[j].name })
+ services := []CNPGOperatorWebhookService{}
+ for _, k := range keys {
+ svc := CNPGOperatorWebhookService{Namespace: k.ns, Name: k.name}
+ var slices []discoveryv1.EndpointSlice
+ svc.ReadSource = s.gatedRead(ctx, cnpgGrantListEndpointSlcs, k.ns, func() error {
+ list, err := typed.DiscoveryV1().EndpointSlices(k.ns).List(ctx, metav1.ListOptions{LabelSelector: cnpgEndpointSliceServiceLabel + "=" + k.name})
+ if err == nil {
+ slices = list.Items
+ }
+ return err
+ })
+ if svc.State == cnpgReadOK {
+ ready, notReady := cnpgEndpointCounts(slices)
+ svc.ReadyEndpoints, svc.NotReadyEndpoints = &ready, ¬Ready
+ }
+ services = append(services, svc)
+ }
+ return configs, services
+}
+
+// cnpgEndpointCounts counts endpoints; an endpoint whose ready condition is
+// unset counts as ready, as the EndpointSlice API defines it.
+func cnpgEndpointCounts(slices []discoveryv1.EndpointSlice) (int, int) {
+ ready, notReady := 0, 0
+ for _, sl := range slices {
+ for _, ep := range sl.Endpoints {
+ if ep.Conditions.Ready == nil || *ep.Conditions.Ready {
+ ready++
+ } else {
+ notReady++
+ }
+ }
+ }
+ return ready, notReady
+}
+
+// cnpgOperatorProcessStart is when the operator container last started: the
+// counters reset with every container restart, which the Pod's own start
+// time does not reflect.
+func cnpgOperatorProcessStart(p *corev1.Pod, container string) string {
+ for _, st := range p.Status.ContainerStatuses {
+ if st.Name == container && st.State.Running != nil && !st.State.Running.StartedAt.IsZero() {
+ return st.State.Running.StartedAt.UTC().Format(time.RFC3339)
+ }
+ }
+ return ""
+}
+
+func (s *Reader) operatorReconcile(callerCtx context.Context, d *appsv1.Deployment, pods []corev1.Pod, port int, leaderPod string) []CNPGOperatorReconcilePod {
+ out := []CNPGOperatorReconcilePod{}
+ container := cnpgOperatorContainer
+ if c := cnpgOperatorContainerOf(d); c != nil {
+ container = c.Name
+ }
+ if len(pods) == 0 {
+ return out
+ }
+ allowed := s.Access.Permission(callerCtx, grantGetPodsProxy.In(d.Namespace)) != integration.PermissionDenied
+ var client kubernetes.Interface
+ if allowed {
+ client = s.Clients.Proxy
+ }
+ identity := s.Identity
+ run := newCNPGRuntimeRunner(callerCtx)
+ results := make([]CNPGOperatorReconcilePod, len(pods))
+ for i := range pods {
+ p := &pods[i]
+ res := &results[i]
+ res.Pod, res.Leader, res.Controllers = p.Name, p.Name == leaderPod, []CNPGOperatorControllerStats{}
+ res.StartedAt = cnpgOperatorProcessStart(p, container)
+ switch {
+ case !allowed:
+ res.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateDenied, Error: "reading the operator's metrics needs " + grantGetPodsProxy.In(d.Namespace).String()}
+ continue
+ case client == nil:
+ res.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateError, Error: "cluster client unavailable"}
+ continue
+ case p.Status.Phase != corev1.PodRunning:
+ res.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateUnreachable, Error: "the Pod is " + string(p.Status.Phase)}
+ continue
+ }
+ target := proxyTarget{namespace: p.Namespace, pod: p.Name, podUID: p.UID, port: port, path: cnpgOperatorMetricsPath, scheme: "http", limit: cnpgOperatorMetricsCap}
+ run.do(func(ctx context.Context) {
+ outcome := memoized(ctx, identity, target, cnpgOperatorMetricsTTL, func(ctx context.Context) cnpgProxyOutcome {
+ return proxyGetWithFallback(ctx, client, target)
+ })
+ res.CNPGRuntimeSource = outcome.source()
+ if outcome.state != runtimeStateOK {
+ return
+ }
+ samples, reason, err := cnpgPromText(outcome)
+ if err != nil {
+ res.State, res.Error = runtimeStateError, "could not parse the operator's metrics: "+err.Error()
+ return
+ }
+ res.Reason = reason
+ res.Controllers = cnpgReconcileStats(samples)
+ if len(res.Controllers) == 0 {
+ res.State = "partial"
+ res.Reason = cnpgJoinReasons(res.Reason, "no controller_runtime_reconcile series on this endpoint")
+ }
+ })
+ }
+ run.wait()
+ return append(out, results...)
+}
+
+func cnpgReconcileStats(samples map[string][]cnpgSample) []CNPGOperatorControllerStats {
+ byController := map[string]*CNPGOperatorControllerStats{}
+ get := func(name string) *CNPGOperatorControllerStats {
+ if st, ok := byController[name]; ok {
+ return st
+ }
+ st := &CNPGOperatorControllerStats{Controller: name, Results: map[string]float64{}}
+ byController[name] = st
+ return st
+ }
+ for _, sm := range samples["controller_runtime_reconcile_errors_total"] {
+ v := sm.value
+ get(sm.labels["controller"]).Errors = &v
+ }
+ for _, sm := range samples["controller_runtime_reconcile_total"] {
+ st := get(sm.labels["controller"])
+ st.Results[sm.labels["result"]] += sm.value
+ total := sm.value
+ if st.Total != nil {
+ total += *st.Total
+ }
+ st.Total = &total
+ }
+ out := make([]CNPGOperatorControllerStats, 0, len(byController))
+ for _, st := range byController {
+ out = append(out, *st)
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Controller < out[j].Controller })
+ return out
+}
+
+func (s *Reader) operatorEvents(ctx context.Context, typed kubernetes.Interface, d *appsv1.Deployment, pods []corev1.Pod) CNPGOperatorEvents {
+ out := CNPGOperatorEvents{Items: []CNPGRecoveryEvent{}}
+ var events []corev1.Event
+ out.ReadSource = s.gatedRead(ctx, cnpgGrantListEvents, d.Namespace, func() error {
+ list, err := typed.CoreV1().Events(d.Namespace).List(ctx, metav1.ListOptions{})
+ if err == nil {
+ events = list.Items
+ }
+ return err
+ })
+ if out.State != cnpgReadOK {
+ return out
+ }
+ subjects := map[types.UID]bool{d.UID: true}
+ for _, p := range pods {
+ subjects[p.UID] = true
+ for _, ref := range p.OwnerReferences {
+ if ref.Kind == "ReplicaSet" {
+ subjects[ref.UID] = true
+ }
+ }
+ }
+ for _, e := range events {
+ if e.InvolvedObject.Kind == "ReplicaSet" && strings.HasPrefix(e.InvolvedObject.Name, d.Name+"-") {
+ subjects[e.InvolvedObject.UID] = true
+ }
+ if e.InvolvedObject.Kind == "Lease" && e.InvolvedObject.Name == cnpgOperatorLeaseName {
+ subjects[e.InvolvedObject.UID] = true
+ }
+ }
+ out.Items = cnpgRecoveryEventsOf(events, subjects, cnpgOperatorEventLimit)
+ return out
+}
diff --git a/internal/cnpg/operator_events_test.go b/internal/cnpg/operator_events_test.go
new file mode 100644
index 0000000000..1393ebc91a
--- /dev/null
+++ b/internal/cnpg/operator_events_test.go
@@ -0,0 +1,42 @@
+package cnpg
+
+import (
+ "context"
+ "fmt"
+ "testing"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/runtime"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+)
+
+func TestCNPGOperatorEventsBindCurrentObjectsAndRetainRolloutHistory(t *testing.T) {
+ d := operatorDeployment(1, 1)
+ d.UID = "deployment-current"
+ pod := corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "operator-pod", Namespace: d.Namespace, UID: "pod-current"}}
+ var events []runtime.Object
+ for i, subject := range []corev1.ObjectReference{
+ {Kind: "Pod", Name: pod.Name, UID: pod.UID},
+ {Kind: "Pod", Name: pod.Name, UID: "pod-old"},
+ {Kind: "Deployment", Name: d.Name, UID: d.UID},
+ {Kind: "Deployment", Name: d.Name, UID: "deployment-old"},
+ {Kind: "ReplicaSet", Name: d.Name + "-old-revision", UID: "historical-revision"},
+ {Kind: "Lease", Name: cnpgOperatorLeaseName, UID: "lease"},
+ } {
+ events = append(events, &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: fmt.Sprintf("event-%d", i), Namespace: d.Namespace}, InvolvedObject: subject, Reason: string(subject.UID)})
+ }
+ out := newTestReader(nil).operatorEvents(context.Background(), k8sfake.NewSimpleClientset(events...), d, []corev1.Pod{pod})
+ if out.State != cnpgReadOK || len(out.Items) != 4 {
+ t.Fatalf("events: %+v", out)
+ }
+ seen := map[string]bool{}
+ for _, e := range out.Items {
+ seen[e.Reason] = true
+ }
+ for _, uid := range []string{"pod-current", "deployment-current", "historical-revision", "lease"} {
+ if !seen[uid] {
+ t.Errorf("event missing: %s", uid)
+ }
+ }
+}
diff --git a/internal/cnpg/operator_memo_test.go b/internal/cnpg/operator_memo_test.go
new file mode 100644
index 0000000000..6246a2c65e
--- /dev/null
+++ b/internal/cnpg/operator_memo_test.go
@@ -0,0 +1,65 @@
+package cnpg
+
+import (
+ "context"
+ "testing"
+
+ admissionv1 "k8s.io/api/admissionregistration/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+func TestOperatorMemoRechecksCallerGrants(t *testing.T) {
+ policy := admissionv1.Fail
+ typed := k8sfake.NewSimpleClientset(cnpgOperatorDeployment(), &admissionv1.ValidatingWebhookConfiguration{
+ ObjectMeta: metav1.ObjectMeta{Name: cnpgValidatingWebhookConfig},
+ Webhooks: []admissionv1.ValidatingWebhook{{Name: "vcluster.cnpg.io", FailurePolicy: &policy, ClientConfig: admissionv1.WebhookClientConfig{Service: &admissionv1.ServiceReference{Namespace: "cnpg-system", Name: "webhook"}}}},
+ })
+ core, err := k8score.NewResourceCache(k8score.CacheConfig{Client: typed, ResourceTypes: map[string]bool{"deployments": true}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer core.Stop()
+ reader := newTestReader(nil)
+ reader.Identity = t.Name()
+ reader.Clients.Typed = typed
+ reader.Observations.Cache = &k8s.ResourceCache{ResourceCache: core}
+ reader.Observations.OperatorScope = func(context.Context) []string { return nil }
+ reader.Observations.TypedScope = func(context.Context, *k8s.ResourceCache, []string, string, string) (integration.KindAccess, []string) {
+ return integration.AccessFromScope(nil, false), nil
+ }
+ denied := false
+ checks := 0
+ reader.Access.Permission = func(_ context.Context, g auth.Grant) string {
+ if g == cnpgGrantGetValidatingWH {
+ checks++
+ if denied {
+ return integration.PermissionDenied
+ }
+ }
+ return integration.PermissionAllowed
+ }
+ ctx := context.Background()
+ first := reader.OperatorStatus(ctx, []string{"db"}).Namespaces["db"]
+ if first.WebhookRejects == nil || !*first.WebhookRejects {
+ t.Fatalf("first webhook verdict=%+v", first)
+ }
+ before := len(typed.Actions())
+ reader.OperatorStatus(ctx, []string{"db"})
+ if len(typed.Actions()) != before {
+ t.Fatal("unchanged caller did not reuse operator memo")
+ }
+ if checks < 2 {
+ t.Fatal("warm memo skipped caller's webhook grant")
+ }
+ denied = true
+ after := reader.OperatorStatus(ctx, []string{"db"}).Namespaces["db"]
+ if after.WebhookRejects != nil || after.Unknown == "" {
+ t.Fatalf("memo disclosed revoked webhook facts: %+v", after)
+ }
+}
diff --git a/internal/cnpg/operator_status.go b/internal/cnpg/operator_status.go
new file mode 100644
index 0000000000..afdc8bb91f
--- /dev/null
+++ b/internal/cnpg/operator_status.go
@@ -0,0 +1,359 @@
+package cnpg
+
+import (
+ "context"
+ "log"
+ "reflect"
+ "slices"
+ "sort"
+ "strings"
+ "sync"
+ "time"
+
+ appsv1 "k8s.io/api/apps/v1"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+const (
+ cnpgOperatorReconciling = "reconciling"
+ cnpgOperatorNotReconciling = "notReconciling"
+ cnpgOperatorUnknown = "unknown"
+ // cnpgOperatorNotWatched: an operator is visible but none watches the
+ // namespace, which the Operator screen reports on its own.
+ cnpgOperatorNotWatched = "notWatched"
+
+ cnpgOperatorStatusTTL = 10 * time.Second
+ // Bounds the detached operator read; a slow API server reads as unknown.
+ cnpgOperatorFactsTimeout = 10 * time.Second
+)
+
+// CNPGOperatorVerdict is whether the operator watching one namespace is acting
+// on it. While it is not, CNPG status in that namespace is whatever the
+// operator last wrote, and writes Radar makes wait for it (or, with the
+// admission webhook down, are rejected outright).
+type CNPGOperatorVerdict struct {
+ State string `json:"state"`
+ // Reasons are the observations behind notReconciling, as sentences.
+ Reasons []string `json:"reasons,omitempty"`
+ // Unknown says what could not be read when the state is unknown, or what
+ // is still unread next to a known state.
+ Unknown string `json:"unknown,omitempty"`
+ // WebhookRejects is true when a CNPG admission webhook fails closed and its
+ // Service has no ready endpoint, so the API server rejects writes to CNPG
+ // objects; nil when that could not be read.
+ WebhookRejects *bool `json:"webhookRejects"`
+ WebhookReason string `json:"webhookReason,omitempty"`
+ // Operator is namespace/name of the operator Deployment the verdict is about.
+ Operator string `json:"operator,omitempty"`
+}
+
+// CNPGOperatorStatusResponse is GET /api/cnpg/operator/status.
+type CNPGOperatorStatusResponse struct {
+ Namespaces map[string]CNPGOperatorVerdict `json:"namespaces"`
+}
+
+type cnpgOperatorFact struct {
+ namespace, name string
+ watch CNPGOperatorWatch
+ leading *bool
+ reason string
+}
+
+type cnpgOperatorFacts struct {
+ operators []cnpgOperatorFact
+ // deploymentsUnknown is set when the operator Deployments could not be
+ // fully listed, so "none found" is not an answer.
+ deploymentsUnknown string
+ webhookRejects *bool
+ webhookReason string
+ webhookUnknown string
+}
+
+var (
+ cnpgOperatorStatusMu sync.Mutex
+ cnpgOperatorStatusMemo = map[string]cnpgOperatorStatusMemoEntry{}
+)
+
+type cnpgOperatorStatusMemoEntry struct {
+ expires time.Time
+ facts cnpgOperatorFacts
+ cache *k8s.ResourceCache
+ scope []string
+ access integration.KindAccess
+ grants map[auth.Grant]string
+}
+
+// cnpgOperatorVerdictFor is the verdict for one namespace, or unknown when
+// the facts cannot be gathered at all.
+func (s *Reader) operatorVerdictFor(ctx context.Context, namespace string) CNPGOperatorVerdict {
+ return s.operatorFactsFor(ctx).verdict(namespace)
+}
+
+func (s *Reader) operatorFactsFor(ctx context.Context) cnpgOperatorFacts {
+ if s.Observations.Cache == nil || s.Clients.Typed == nil {
+ return s.readCNPGOperatorFacts(ctx)
+ }
+ scope := s.Observations.OperatorScope(ctx)
+ access, _ := s.Observations.TypedScope(ctx, s.Observations.Cache, scope, "apps", "deployments")
+ identity := s.Identity
+ now := time.Now()
+ cnpgOperatorStatusMu.Lock()
+ entry, hit := cnpgOperatorStatusMemo[identity]
+ cnpgOperatorStatusMu.Unlock()
+ if hit && now.Before(entry.expires) && entry.cache == s.Observations.Cache && reflect.DeepEqual(entry.scope, scope) && reflect.DeepEqual(entry.access, access) {
+ authorized := true
+ for grant, permission := range entry.grants {
+ if s.Access.Permission(ctx, grant) != permission {
+ authorized = false
+ break
+ }
+ }
+ if authorized {
+ return entry.facts
+ }
+ }
+ detached, cancel := context.WithTimeout(context.WithoutCancel(ctx), cnpgOperatorFactsTimeout)
+ defer cancel()
+ grants := map[auth.Grant]string{}
+ reader := *s
+ reader.Access.Permission = func(ctx context.Context, grant auth.Grant) string {
+ permission := s.Access.Permission(ctx, grant)
+ grants[grant] = permission
+ return permission
+ }
+ facts := reader.readCNPGOperatorFacts(detached)
+ cnpgOperatorStatusMu.Lock()
+ defer cnpgOperatorStatusMu.Unlock()
+ for key, entry := range cnpgOperatorStatusMemo {
+ if now.After(entry.expires) {
+ delete(cnpgOperatorStatusMemo, key)
+ }
+ }
+ cnpgOperatorStatusMemo[identity] = cnpgOperatorStatusMemoEntry{expires: now.Add(cnpgOperatorStatusTTL), facts: facts, cache: s.Observations.Cache, scope: append([]string(nil), scope...), access: access, grants: grants}
+ return facts
+}
+
+func (s *Reader) readCNPGOperatorFacts(ctx context.Context) cnpgOperatorFacts {
+ out := cnpgOperatorFacts{}
+ cache := s.Observations.Cache
+ typed := s.Clients.Typed
+ if cache == nil || typed == nil {
+ out.deploymentsUnknown = "Radar is not connected to the cluster"
+ out.webhookUnknown = out.deploymentsUnknown
+ return out
+ }
+ acc, deployments := s.operatorDeployments(ctx, cache, s.Observations.OperatorScope(ctx))
+ if acc.State != integration.KindCoverageFull {
+ out.deploymentsUnknown = "Radar cannot list Deployments in every namespace, so the operator may be out of view"
+ if acc.State == integration.KindCoverageSyncing {
+ out.deploymentsUnknown = "Deployments are still syncing"
+ }
+ }
+ for _, d := range deployments {
+ if d.Labels[cnpgOperatorNameLabel] != cnpgOperatorNameValue {
+ continue
+ }
+ c := cnpgOperatorContainerOf(d)
+ fact := cnpgOperatorFact{namespace: d.Namespace, name: d.Name, watch: cnpgOperatorWatchOf(c, d.Namespace)}
+ leader := s.operatorLeader(ctx, typed, d, c, nil)
+ fact.leading, fact.reason = cnpgOperatorLeading(d, leader)
+ out.operators = append(out.operators, fact)
+ }
+ configs, services := s.operatorWebhooks(ctx, typed)
+ out.webhookRejects, out.webhookReason, out.webhookUnknown = cnpgWebhookRejects(configs, services)
+ return out
+}
+
+// cnpgOperatorLeading judges one operator Deployment: leading is nil when the
+// lease could not be read and the Deployment alone does not settle it.
+func cnpgOperatorLeading(d *appsv1.Deployment, leader CNPGOperatorLeader) (*bool, string) {
+ f, t := false, true
+ desired := int32(1)
+ if d.Spec.Replicas != nil {
+ desired = *d.Spec.Replicas
+ }
+ if desired == 0 {
+ return &f, "the operator Deployment " + d.Namespace + "/" + d.Name + " is scaled to zero"
+ }
+ if d.Status.ReadyReplicas == 0 {
+ return &f, "the operator Deployment " + d.Namespace + "/" + d.Name + " has no ready Pod"
+ }
+ switch leader.State {
+ case "disabled":
+ return &t, ""
+ case cnpgReadOK:
+ if leader.Stale {
+ return &f, "no operator instance has renewed the leader lease " + d.Namespace + "/" + cnpgOperatorLeaseName
+ }
+ return &t, ""
+ case cnpgReadDenied:
+ return nil, "its leader lease is not readable (needs " + integration.GrantText(leader.Grant) + ")"
+ default:
+ if leader.Reason != "" {
+ log.Printf("[cnpg] Operator %s/%s leader lease unread: %s", d.Namespace, d.Name, leader.Reason)
+ }
+ return nil, "couldn't read its leader lease"
+ }
+}
+
+// cnpgWebhookRejects reports whether a fail-closed CNPG webhook has no ready
+// endpoint behind it. A configuration that is absent rejects nothing.
+func cnpgWebhookRejects(configs []CNPGOperatorWebhookConfig, services []CNPGOperatorWebhookService) (*bool, string, string) {
+ bySvc := map[string]CNPGOperatorWebhookService{}
+ for _, svc := range services {
+ bySvc[svc.Namespace+"/"+svc.Name] = svc
+ }
+ var unknown []string
+ var rejecting []string
+ checked := 0
+ for _, cfg := range configs {
+ switch cfg.State {
+ case cnpgReadOK:
+ case cnpgReadNotFound:
+ continue
+ case cnpgReadDenied:
+ unknown = append(unknown, cfg.Kind+" "+cfg.Name+" is not readable (needs "+integration.GrantText(cfg.Grant)+")")
+ continue
+ default:
+ unknown = append(unknown, cfg.Kind+" "+cfg.Name+" could not be read")
+ continue
+ }
+ for _, wh := range cfg.Webhooks {
+ if wh.FailurePolicy != "Fail" || wh.Service == "" {
+ continue
+ }
+ svc, ok := bySvc[wh.Service]
+ switch {
+ case !ok || svc.ReadyEndpoints == nil:
+ grant := ""
+ if ok && svc.Grant != nil {
+ grant = " (needs " + svc.Grant.String() + ")"
+ }
+ unknown = append(unknown, "the endpoints of webhook Service "+wh.Service+" are not readable"+grant)
+ case *svc.ReadyEndpoints == 0:
+ rejecting = append(rejecting, wh.Service)
+ default:
+ checked++
+ }
+ }
+ }
+ rejecting = dedupeSorted(rejecting)
+ unknown = dedupeSorted(unknown)
+ if len(rejecting) > 0 {
+ t := true
+ return &t, "the admission webhook Service " + strings.Join(rejecting, ", ") + " has no ready endpoint and fails closed, so the API server rejects writes to CloudNativePG objects", strings.Join(unknown, "; ")
+ }
+ if len(unknown) > 0 {
+ return nil, "", strings.Join(unknown, "; ")
+ }
+ f := false
+ return &f, "", ""
+}
+
+func dedupeSorted(in []string) []string {
+ if len(in) == 0 {
+ return in
+ }
+ sort.Strings(in)
+ out := in[:1]
+ for _, v := range in[1:] {
+ if v != out[len(out)-1] {
+ out = append(out, v)
+ }
+ }
+ return out
+}
+
+func (f cnpgOperatorFacts) verdict(namespace string) CNPGOperatorVerdict {
+ v := CNPGOperatorVerdict{State: cnpgOperatorUnknown, WebhookRejects: f.webhookRejects, WebhookReason: f.webhookReason}
+ // An operator whose WATCH_NAMESPACE Radar cannot resolve may or may not
+ // watch this namespace: it can neither confirm reconciliation nor rule it out.
+ var watching []cnpgOperatorFact
+ var unresolved []string
+ couldReconcile := false
+ for _, op := range f.operators {
+ switch {
+ case op.watch.Unresolved != "":
+ unresolved = append(unresolved, "whether "+op.namespace+"/"+op.name+" watches "+namespace+" is unknown: its WATCH_NAMESPACE is "+op.watch.Unresolved)
+ if op.leading == nil || *op.leading {
+ couldReconcile = true
+ }
+ case op.watch.All || slices.Contains(op.watch.Namespaces, namespace):
+ watching = append(watching, op)
+ }
+ }
+ unknown := []string{}
+ switch {
+ case len(watching) == 0 && len(unresolved) > 0:
+ unknown = append(unknown, unresolved...)
+ if f.deploymentsUnknown != "" {
+ unknown = append(unknown, f.deploymentsUnknown)
+ }
+ case len(watching) == 0 && f.deploymentsUnknown != "":
+ unknown = append(unknown, f.deploymentsUnknown)
+ case len(watching) == 0 && len(f.operators) == 0:
+ unknown = append(unknown, "No operator Deployment labelled "+cnpgOperatorNameLabel+"="+cnpgOperatorNameValue+" was found")
+ case len(watching) == 0:
+ v.State = cnpgOperatorNotWatched
+ default:
+ v.Operator = watching[0].namespace + "/" + watching[0].name
+ leading, anyUnknown := false, false
+ var notLeading []string
+ for _, op := range watching {
+ switch {
+ case op.leading == nil:
+ anyUnknown = true
+ unknown = append(unknown, op.reason)
+ case *op.leading:
+ leading = true
+ v.Operator = op.namespace + "/" + op.name
+ default:
+ notLeading = append(notLeading, op.reason)
+ }
+ }
+ switch {
+ case leading:
+ v.State = cnpgOperatorReconciling
+ unknown = nil
+ case anyUnknown:
+ case couldReconcile:
+ unknown = append(unknown, unresolved...)
+ default:
+ v.State = cnpgOperatorNotReconciling
+ v.Reasons = notLeading
+ }
+ }
+ if f.webhookRejects == nil && f.webhookUnknown != "" {
+ unknown = append(unknown, f.webhookUnknown)
+ }
+ v.Unknown = strings.Join(unknown, "; ")
+ return v
+}
+
+// cnpgOperatorWebhookGuard refuses a write that goes through a CNPG admission
+// webhook the API server cannot reach; status subresource patches and Pod
+// deletes do not, so those actions keep their verdict.
+func cnpgOperatorWebhookGuard(v CNPGOperatorVerdict, capability integration.ActionCapability) integration.ActionCapability {
+ if !capability.Allowed || v.WebhookRejects == nil || !*v.WebhookRejects {
+ return capability
+ }
+ capability.Allowed = false
+ capability.ReasonCode = "operator_webhook"
+ capability.Reason = "The API server would reject it: " + v.WebhookReason
+ return capability
+}
+
+func (s *Reader) OperatorStatus(ctx context.Context, namespaces []string) CNPGOperatorStatusResponse {
+ resp := CNPGOperatorStatusResponse{Namespaces: map[string]CNPGOperatorVerdict{}}
+ if len(namespaces) == 0 {
+ return resp
+ }
+ facts := s.operatorFactsFor(ctx)
+ for _, namespace := range namespaces {
+ resp.Namespaces[namespace] = facts.verdict(namespace)
+ }
+ return resp
+}
diff --git a/internal/cnpg/operator_status_test.go b/internal/cnpg/operator_status_test.go
new file mode 100644
index 0000000000..b02b47cba9
--- /dev/null
+++ b/internal/cnpg/operator_status_test.go
@@ -0,0 +1,208 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "strings"
+ "testing"
+ "time"
+
+ appsv1 "k8s.io/api/apps/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func operatorDeployment(replicas, ready int32) *appsv1.Deployment {
+ return &appsv1.Deployment{
+ ObjectMeta: metav1.ObjectMeta{Namespace: "cnpg-system", Name: "cnpg-controller-manager"},
+ Spec: appsv1.DeploymentSpec{Replicas: &replicas},
+ Status: appsv1.DeploymentStatus{ReadyReplicas: ready},
+ }
+}
+
+func TestCNPGOperatorLeading(t *testing.T) {
+ ok := CNPGOperatorLeader{ReadSource: integration.ReadSource{State: cnpgReadOK}}
+ stale := ok
+ stale.Stale = true
+ denied := CNPGOperatorLeader{ReadSource: integration.ReadSource{State: cnpgReadDenied, Grant: cnpgGrantGetLeases.In("cnpg-system").Ref()}}
+ cases := []struct {
+ name string
+ d *appsv1.Deployment
+ leader CNPGOperatorLeader
+ leading *bool
+ }{
+ {"held lease", operatorDeployment(1, 1), ok, boolPtr(true)},
+ {"expired lease", operatorDeployment(1, 1), stale, boolPtr(false)},
+ {"crash-looping Pod leads nothing whatever the lease says", operatorDeployment(1, 0), ok, boolPtr(false)},
+ {"scaled to zero", operatorDeployment(0, 0), denied, boolPtr(false)},
+ {"lease unreadable", operatorDeployment(1, 1), denied, nil},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ got, reason := cnpgOperatorLeading(tc.d, tc.leader)
+ if (got == nil) != (tc.leading == nil) || (got != nil && *got != *tc.leading) {
+ t.Fatalf("leading = %v (%s), want %v", got, reason, tc.leading)
+ }
+ if (got == nil || !*got) && reason == "" {
+ t.Error("a not-leading or unknown verdict needs a reason")
+ }
+ })
+ }
+}
+
+func TestCNPGWebhookRejects(t *testing.T) {
+ zero, one := 0, 1
+ cfg := func(policy string) []CNPGOperatorWebhookConfig {
+ return []CNPGOperatorWebhookConfig{{
+ Kind: "ValidatingWebhookConfiguration", Name: "cnpg-validating-webhook-configuration",
+ ReadSource: integration.ReadSource{State: cnpgReadOK},
+ Webhooks: []CNPGOperatorWebhook{{Name: "vcluster.cnpg.io", FailurePolicy: policy, Service: "cnpg-system/cnpg-webhook-service"}},
+ }}
+ }
+ svc := func(ready *int) []CNPGOperatorWebhookService {
+ return []CNPGOperatorWebhookService{{Namespace: "cnpg-system", Name: "cnpg-webhook-service", ReadyEndpoints: ready}}
+ }
+ if got, reason, _ := cnpgWebhookRejects(cfg("Fail"), svc(&zero)); got == nil || !*got || !strings.Contains(reason, "cnpg-system/cnpg-webhook-service") {
+ t.Errorf("fail-closed, no endpoint: %v %q", got, reason)
+ }
+ if got, _, _ := cnpgWebhookRejects(cfg("Ignore"), svc(&zero)); got == nil || *got {
+ t.Errorf("fail-open webhook rejects nothing: %v", got)
+ }
+ if got, _, _ := cnpgWebhookRejects(cfg("Fail"), svc(&one)); got == nil || *got {
+ t.Errorf("ready endpoint: %v", got)
+ }
+ if got, _, unknown := cnpgWebhookRejects(cfg("Fail"), svc(nil)); got != nil || unknown == "" {
+ t.Errorf("unreadable endpoints must be unknown: %v %q", got, unknown)
+ }
+ denied := []CNPGOperatorWebhookConfig{{Kind: "ValidatingWebhookConfiguration", Name: "x", ReadSource: integration.ReadSource{State: cnpgReadDenied, Grant: cnpgGrantGetValidatingWH.Ref()}}}
+ if got, _, unknown := cnpgWebhookRejects(denied, nil); got != nil || !strings.Contains(unknown, "get validatingwebhookconfigurations") {
+ t.Errorf("denied config: %v %q", got, unknown)
+ }
+ absent := []CNPGOperatorWebhookConfig{{Kind: "ValidatingWebhookConfiguration", Name: "x", ReadSource: integration.ReadSource{State: cnpgReadNotFound}}}
+ if got, _, _ := cnpgWebhookRejects(absent, nil); got == nil || *got {
+ t.Errorf("absent config rejects nothing: %v", got)
+ }
+}
+
+func TestCNPGOperatorFactsVerdict(t *testing.T) {
+ yes, no := true, false
+ down := cnpgOperatorFact{namespace: "cnpg-system", name: "op", watch: CNPGOperatorWatch{All: true}, leading: &no, reason: "the operator Deployment cnpg-system/op has no ready Pod"}
+ up := down
+ up.leading, up.reason = &yes, ""
+ scoped := up
+ scoped.watch = CNPGOperatorWatch{Namespaces: []string{"team-a"}}
+ unread := down
+ unread.leading, unread.reason = nil, "the leader lease is not readable"
+
+ f := cnpgOperatorFacts{operators: []cnpgOperatorFact{down}, webhookRejects: &no}
+ if v := f.verdict("pg"); v.State != cnpgOperatorNotReconciling || len(v.Reasons) != 1 || v.Operator != "cnpg-system/op" {
+ t.Errorf("down = %+v", v)
+ }
+ f.operators = []cnpgOperatorFact{up}
+ if v := f.verdict("pg"); v.State != cnpgOperatorReconciling || v.Unknown != "" {
+ t.Errorf("up = %+v", v)
+ }
+ f.operators = []cnpgOperatorFact{scoped}
+ if v := f.verdict("pg"); v.State != cnpgOperatorNotWatched {
+ t.Errorf("not watched = %+v", v)
+ }
+ f.operators = []cnpgOperatorFact{unread}
+ if v := f.verdict("pg"); v.State != cnpgOperatorUnknown || !strings.Contains(v.Unknown, "lease") {
+ t.Errorf("unreadable lease = %+v", v)
+ }
+ // WATCH_NAMESPACE from a ConfigMap Radar does not read: a healthy leader
+ // may be restricted to another namespace, so it proves nothing here.
+ fromConfigMap := up
+ fromConfigMap.name = "op-cm"
+ fromConfigMap.watch = CNPGOperatorWatch{Source: "env WATCH_NAMESPACE", Unresolved: "set from ConfigMap cnpg-config key WATCH_NAMESPACE"}
+ f.operators = []cnpgOperatorFact{fromConfigMap}
+ if v := f.verdict("team-b"); v.State != cnpgOperatorUnknown || !strings.Contains(v.Unknown, "ConfigMap cnpg-config") {
+ t.Errorf("unresolved watch scope = %+v, want unknown naming the ConfigMap", v)
+ }
+ f.operators = []cnpgOperatorFact{fromConfigMap, up}
+ if v := f.verdict("team-b"); v.State != cnpgOperatorReconciling || v.Operator != "cnpg-system/op" {
+ t.Errorf("confirmed watcher beside an unresolved one = %+v", v)
+ }
+ f.operators = []cnpgOperatorFact{fromConfigMap, down}
+ if v := f.verdict("team-b"); v.State != cnpgOperatorUnknown {
+ t.Errorf("down confirmed watcher beside a leading unresolved one = %+v, want unknown", v)
+ }
+ f = cnpgOperatorFacts{deploymentsUnknown: "Radar cannot list Deployments in every namespace"}
+ if v := f.verdict("pg"); v.State != cnpgOperatorUnknown || v.Unknown == "" {
+ t.Errorf("no visible operator = %+v", v)
+ }
+}
+
+func TestCNPGApplyOperatorGuard(t *testing.T) {
+ allowed := integration.ActionCapability{Allowed: true, Permission: integration.PermissionAllowed}
+ refused := integration.ActionCapability{Reason: "The cluster is hibernated", Permission: integration.PermissionAllowed}
+ rejects := true
+ resp := &CNPGClusterCapabilitiesResponse{
+ Operator: CNPGOperatorVerdict{State: cnpgOperatorNotReconciling, WebhookRejects: &rejects, WebhookReason: "the admission webhook Service cnpg-system/cnpg-webhook-service has no ready endpoint"},
+ Actions: CNPGClusterActions{
+ Backup: allowed, Switchover: allowed, Restart: allowed, RestartInstance: allowed, Fence: refused, Psql: allowed,
+ },
+ InstanceActions: map[string]CNPGInstanceActions{"pg-1": {Fence: allowed, Restart: allowed}},
+ }
+ cnpgApplyOperatorGuard(resp)
+ a := resp.Actions
+ if a.Backup.Allowed || !strings.Contains(a.Backup.Reason, "no ready endpoint") || a.Restart.Allowed || resp.InstanceActions["pg-1"].Fence.Allowed {
+ t.Errorf("webhook-bound writes stay allowed: %+v", resp)
+ }
+ if !a.Switchover.Allowed || !a.RestartInstance.Allowed || !a.Psql.Allowed || !resp.InstanceActions["pg-1"].Restart.Allowed {
+ t.Errorf("status patches, Pod deletes and exec do not pass the webhook: %+v", a)
+ }
+ if a.Fence.Reason != "The cluster is hibernated" {
+ t.Errorf("an earlier refusal keeps its reason: %q", a.Fence.Reason)
+ }
+
+ resp.Operator.WebhookRejects = nil
+ resp.Actions.Backup = allowed
+ cnpgApplyOperatorGuard(resp)
+ if !resp.Actions.Backup.Allowed {
+ t.Error("an unknown webhook state must not block")
+ }
+}
+
+func TestCNPGOperatorLeadingPod(t *testing.T) {
+ held := CNPGOperatorLeader{ReadSource: integration.ReadSource{State: cnpgReadOK}, HolderPod: "op-1", HolderIsCurrentPod: true}
+ if got := cnpgOperatorLeadingPod(held); got != "op-1" {
+ t.Errorf("held = %q", got)
+ }
+ expired := held
+ expired.Stale = true
+ if got := cnpgOperatorLeadingPod(expired); got != "" {
+ t.Errorf("an expired holder is labelled leader: %q", got)
+ }
+ gone := held
+ gone.HolderIsCurrentPod = false
+ if got := cnpgOperatorLeadingPod(gone); got != "" {
+ t.Errorf("a holder that is no current Pod is labelled leader: %q", got)
+ }
+}
+
+func TestCNPGOperatorFactsReadOutlivesTheCaller(t *testing.T) {
+ type key struct{}
+ parent, cancelParent := context.WithCancel(context.WithValue(context.Background(), key{}, "alice"))
+ detached, cancel := context.WithTimeout(context.WithoutCancel(parent), time.Minute)
+ defer cancel()
+ cancelParent()
+ if err := detached.Err(); err != nil {
+ t.Fatalf("a caller hanging up cancelled the shared operator read: %v", err)
+ }
+ if detached.Value(key{}) != "alice" {
+ t.Error("the caller's identity was not kept")
+ }
+}
+
+func TestCNPGOperatorLeadingHidesRawLeaseErrors(t *testing.T) {
+ leader := CNPGOperatorLeader{ReadSource: integration.ReadSource{State: cnpgReadError, Reason: `Get "https://127.0.0.1:55484/apis/coordination.k8s.io/v1/namespaces/cnpg-system/leases/db9c8771.cnpg.io": context canceled`}}
+ got, reason := cnpgOperatorLeading(operatorDeployment(1, 1), leader)
+ if got != nil || reason != "couldn't read its leader lease" {
+ t.Errorf("leading = %v, reason = %q", got, reason)
+ }
+ if plain, ok := cnpgTransportSentence(errors.New(leader.Reason), 0, time.Second); !ok || strings.Contains(plain, "127.0.0.1") {
+ t.Errorf("transport sentence = %q (%v)", plain, ok)
+ }
+}
diff --git a/internal/cnpg/operator_test.go b/internal/cnpg/operator_test.go
new file mode 100644
index 0000000000..50b2247b60
--- /dev/null
+++ b/internal/cnpg/operator_test.go
@@ -0,0 +1,142 @@
+package cnpg
+
+import (
+ "context"
+ "testing"
+ "time"
+
+ appsv1 "k8s.io/api/apps/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/types"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+ k8stesting "k8s.io/client-go/testing"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func cnpgOperatorDeployment() *appsv1.Deployment {
+ return &appsv1.Deployment{
+ ObjectMeta: metav1.ObjectMeta{
+ Name: "cnpg-controller-manager", Namespace: "cnpg-system", Generation: 1,
+ Labels: map[string]string{"app.kubernetes.io/name": "cloudnative-pg"},
+ },
+ Spec: appsv1.DeploymentSpec{
+ Replicas: int32Ptr(1),
+ Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app.kubernetes.io/name": "cloudnative-pg"}},
+ Template: corev1.PodTemplateSpec{
+ ObjectMeta: metav1.ObjectMeta{Labels: map[string]string{"app.kubernetes.io/name": "cloudnative-pg"}},
+ Spec: corev1.PodSpec{Containers: []corev1.Container{
+ {Name: "sidecar", Image: "busybox:1.36"},
+ {
+ Name: "manager",
+ Image: "ghcr.io/cloudnative-pg/cloudnative-pg:1.27.0",
+ Command: []string{"/manager"},
+ Args: []string{
+ "controller", "--leader-elect",
+ "--config-map-name=$(OPERATOR_DEPLOYMENT_NAME)-config",
+ "--secret-name", "$(OPERATOR_DEPLOYMENT_NAME)-config",
+ },
+ Env: []corev1.EnvVar{
+ {Name: "OPERATOR_DEPLOYMENT_NAME", Value: "cnpg-controller-manager"},
+ {Name: "MONITORING_QUERIES_CONFIGMAP", Value: "cnpg-default-monitoring"},
+ },
+ },
+ }},
+ },
+ },
+ Status: appsv1.DeploymentStatus{ObservedGeneration: 1, Replicas: 1, ReadyReplicas: 1},
+ }
+}
+
+func cnpgPluginDeployment() *appsv1.Deployment {
+ return &appsv1.Deployment{
+ ObjectMeta: metav1.ObjectMeta{Name: "barman-cloud", Namespace: "cnpg-system"},
+ Spec: appsv1.DeploymentSpec{
+ Selector: &metav1.LabelSelector{MatchLabels: map[string]string{"app": "barman-cloud"}},
+ Template: corev1.PodTemplateSpec{
+ ObjectMeta: metav1.ObjectMeta{Labels: map[string]string{"app": "barman-cloud"}},
+ Spec: corev1.PodSpec{Containers: []corev1.Container{
+ {Name: "barman-cloud", Image: "ghcr.io/cloudnative-pg/plugin-barman-cloud:v0.5.0"},
+ }},
+ },
+ },
+ }
+}
+
+func cnpgOperatorTestPod(name string, labels map[string]string, statuses ...corev1.ContainerStatus) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: "cnpg-system", UID: types.UID(name + "-uid"), Labels: labels},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodRunning,
+ StartTime: &metav1.Time{Time: time.Date(2026, 9, 1, 0, 0, 0, 0, time.UTC)},
+ Conditions: []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionTrue}},
+ ContainerStatuses: statuses,
+ },
+ }
+}
+
+func TestCNPGOperatorComponentPods(t *testing.T) {
+ started := time.Date(2026, 10, 4, 9, 0, 0, 0, time.UTC)
+ running := corev1.ContainerState{Running: &corev1.ContainerStateRunning{StartedAt: metav1.NewTime(started)}}
+ opLabels := map[string]string{"app.kubernetes.io/name": "cloudnative-pg"}
+ crashed := cnpgOperatorTestPod("cnpg-controller-manager-a", opLabels,
+ corev1.ContainerStatus{Name: "sidecar", RestartCount: 3, State: running, LastTerminationState: corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{
+ Reason: "Completed", ExitCode: 0, FinishedAt: metav1.NewTime(started.Add(-48 * time.Hour)),
+ }}},
+ corev1.ContainerStatus{Name: "manager", RestartCount: 30, State: running, LastTerminationState: corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{
+ Reason: "OOMKilled", ExitCode: 137, FinishedAt: metav1.NewTime(started.Add(-time.Minute)),
+ }}},
+ )
+ steady := cnpgOperatorTestPod("cnpg-controller-manager-b", opLabels, corev1.ContainerStatus{Name: "manager", State: running})
+ other := cnpgOperatorTestPod("unrelated", map[string]string{"app": "other"})
+ s := newTestReader(nil)
+ ctx := context.Background()
+ d := cnpgOperatorDeployment()
+ read := s.deploymentPods(ctx, k8sfake.NewSimpleClientset(steady, crashed, other), d)
+
+ comp := withCNPGComponentPods(cnpgOperatorComponent(d, "operator", "", cnpgOperatorContainerOf(d)), read, cnpgOperatorContainerOf(d))
+ if comp.PodCoverage == nil || comp.PodCoverage.State != "ok" || len(comp.Pods) != 2 {
+ t.Fatalf("component = %+v", comp)
+ }
+ a, b := comp.Pods[0], comp.Pods[1]
+ if a.Name != "cnpg-controller-manager-a" || !a.Ready || a.Restarts != 33 || a.StartedAt != "2026-10-04T09:00:00Z" {
+ t.Errorf("crashed pod = %+v, want restarts summed over containers and the manager container's start", a)
+ }
+ if lt := a.LastTermination; lt == nil || lt.Container != "manager" || lt.Reason != "OOMKilled" || lt.ExitCode != 137 || lt.FinishedAt != "2026-10-04T08:59:00Z" {
+ t.Errorf("lastTermination = %+v, want the most recently ended container's", a.LastTermination)
+ }
+ if b.Name != "cnpg-controller-manager-b" || b.Restarts != 0 || b.LastTermination != nil {
+ t.Errorf("steady pod = %+v", b)
+ }
+
+ diag := cnpgOperatorPodOf(&read.pods[0])
+ if diag.Restarts != 33 || diag.LastTermination == nil || diag.LastTermination.Reason != "OOMKilled" || diag.StartedAt != "2026-09-01T00:00:00Z" {
+ t.Errorf("diagnosis pod = %+v, want the same restarts and termination, with the Pod's own start", diag)
+ }
+
+ notRunning := cnpgOperatorTestPod("cnpg-controller-manager-c", opLabels, corev1.ContainerStatus{Name: "manager", RestartCount: 70,
+ State: corev1.ContainerState{Waiting: &corev1.ContainerStateWaiting{Reason: "CrashLoopBackOff"}}})
+ comp = withCNPGComponentPods(CNPGOperatorComponent{}, cnpgDeploymentPodRead{pods: []corev1.Pod{*notRunning}, coverage: integration.ReadSource{State: "ok"}}, cnpgOperatorContainerOf(d))
+ if p := comp.Pods[0]; p.StartedAt != "" || p.Restarts != 70 {
+ t.Errorf("a container that is not running has no current start: %+v", p)
+ }
+}
+
+func TestCNPGOperatorComponentPodsAPIReadDenied(t *testing.T) {
+ typed := k8sfake.NewSimpleClientset()
+ typed.PrependReactor("list", "pods", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apiForbidden("pods")
+ })
+ d := cnpgPluginDeployment()
+ read := newTestReader(nil).deploymentPods(context.Background(), typed, d)
+ comp := withCNPGComponentPods(cnpgOperatorComponent(d, "plugin", "barman-cloud.cloudnative-pg.io", firstContainer(d)), read, firstContainer(d))
+ if g := comp.PodCoverage.Grant; comp.PodCoverage.State != "denied" || g == nil || g.Verb != "list" || g.Resource != "pods" || g.Namespace != "cnpg-system" {
+ t.Fatalf("podCoverage = %+v, want denied naming list pods in cnpg-system", comp.PodCoverage)
+ }
+ if comp.Pods != nil {
+ t.Errorf("pods = %+v, want none when they cannot be read", comp.Pods)
+ }
+
+}
diff --git a/internal/cnpg/pooler_actions.go b/internal/cnpg/pooler_actions.go
new file mode 100644
index 0000000000..2326d86443
--- /dev/null
+++ b/internal/cnpg/pooler_actions.go
@@ -0,0 +1,323 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "fmt"
+ "net/http"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// CNPGPoolerDeploymentFact is the Deployment the Pooler runs, which is where
+// readiness lives; the Pooler's own status only counts scheduled Pods.
+// State: ok | missing | unreadable | foreign (a same-named Deployment the
+// Pooler does not control).
+type CNPGPoolerDeploymentFact struct {
+ Name string `json:"name"`
+ State string `json:"state"`
+ Replicas *int32 `json:"replicas,omitempty"`
+ ReadyReplicas *int32 `json:"readyReplicas,omitempty"`
+ UpdatedReplicas *int32 `json:"updatedReplicas,omitempty"`
+ AvailableReplicas *int32 `json:"availableReplicas,omitempty"`
+}
+
+// CNPGPoolerServiceFact is the Service clients connect to.
+type CNPGPoolerServiceFact struct {
+ Name string `json:"name"`
+ State string `json:"state"`
+ Type string `json:"type,omitempty"`
+ Port *int32 `json:"port,omitempty"`
+}
+
+type CNPGPoolerFacts struct {
+ Generation int64 `json:"generation"`
+ Cluster string `json:"cluster"`
+ Type string `json:"type"`
+ Instances *int64 `json:"instances,omitempty"`
+ Paused bool `json:"paused"`
+ PoolMode string `json:"poolMode,omitempty"`
+ Parameters map[string]string `json:"parameters"`
+ Terminating bool `json:"terminating"`
+ Deployment CNPGPoolerDeploymentFact `json:"deployment"`
+ Service CNPGPoolerServiceFact `json:"service"`
+}
+
+type CNPGPoolerActions struct {
+ Pause integration.ActionCapability `json:"pause"`
+ Resume integration.ActionCapability `json:"resume"`
+ // ObserveState is reading each PgBouncer's paused state (pods/exec).
+ ObserveState integration.ActionCapability `json:"observeState"`
+}
+
+// CNPGPoolerCapabilitiesResponse is GET /api/cnpg/poolers/{ns}/{name}/capabilities.
+type CNPGPoolerCapabilitiesResponse struct {
+ UID string `json:"uid"`
+ ResourceVersion string `json:"resourceVersion"`
+ Context string `json:"context"`
+ Facts CNPGPoolerFacts `json:"facts"`
+ Actions CNPGPoolerActions `json:"actions"`
+}
+
+func cnpgPoolerFactsOf(ctx context.Context, c ActionClients, pooler *unstructured.Unstructured) CNPGPoolerFacts {
+ str := func(fields ...string) string {
+ v, _, _ := unstructured.NestedString(pooler.Object, fields...)
+ return v
+ }
+ paused, _, _ := unstructured.NestedBool(pooler.Object, "spec", "pgbouncer", "paused")
+ params, _, _ := unstructured.NestedStringMap(pooler.Object, "spec", "pgbouncer", "parameters")
+ if params == nil {
+ params = map[string]string{}
+ }
+ f := CNPGPoolerFacts{
+ Generation: pooler.GetGeneration(),
+ Cluster: str("spec", "cluster", "name"),
+ Type: str("spec", "type"),
+ Paused: paused,
+ PoolMode: str("spec", "pgbouncer", "poolMode"),
+ Parameters: params,
+ Terminating: !pooler.GetDeletionTimestamp().IsZero(),
+ Deployment: CNPGPoolerDeploymentFact{Name: pooler.GetName()},
+ Service: CNPGPoolerServiceFact{Name: pooler.GetName()},
+ }
+ if n, ok, _ := unstructured.NestedInt64(pooler.Object, "spec", "instances"); ok {
+ f.Instances = &n
+ }
+ if c.Typed == nil {
+ f.Deployment.State, f.Service.State = "unreadable", "unreadable"
+ return f
+ }
+ ns, name, uid := pooler.GetNamespace(), pooler.GetName(), pooler.GetUID()
+ switch d, err := c.Typed.AppsV1().Deployments(ns).Get(ctx, name, metav1.GetOptions{}); {
+ case apierrors.IsNotFound(err):
+ f.Deployment.State = "missing"
+ case err != nil:
+ f.Deployment.State = "unreadable"
+ case !controlledBy(d.OwnerReferences, Group, "Pooler", name, uid):
+ f.Deployment.State = "foreign"
+ default:
+ f.Deployment.State = "ok"
+ f.Deployment.Replicas = d.Spec.Replicas
+ ready, updated, available := d.Status.ReadyReplicas, d.Status.UpdatedReplicas, d.Status.AvailableReplicas
+ f.Deployment.ReadyReplicas, f.Deployment.UpdatedReplicas, f.Deployment.AvailableReplicas = &ready, &updated, &available
+ }
+ switch svc, err := c.Typed.CoreV1().Services(ns).Get(ctx, name, metav1.GetOptions{}); {
+ case apierrors.IsNotFound(err):
+ f.Service.State = "missing"
+ case err != nil:
+ f.Service.State = "unreadable"
+ case !controlledBy(svc.OwnerReferences, Group, "Pooler", name, uid):
+ f.Service.State = "foreign"
+ default:
+ f.Service.State = "ok"
+ f.Service.Type = string(svc.Spec.Type)
+ if len(svc.Spec.Ports) > 0 {
+ port := svc.Spec.Ports[0].Port
+ f.Service.Port = &port
+ }
+ }
+ return f
+}
+
+func cnpgGuardPooler(f CNPGPoolerFacts, pause bool) string {
+ switch {
+ case f.Terminating:
+ return "The Pooler is being deleted"
+ case pause && f.Paused:
+ return "The Pooler is paused already"
+ case !pause && !f.Paused:
+ return "The Pooler is not paused"
+ }
+ return ""
+}
+
+func (s *Reader) PoolerCapabilities(ctx context.Context, c ActionClients, contextName, namespace, name string) (*CNPGPoolerCapabilitiesResponse, error) {
+ pooler, err := c.Dynamic.Resource(cnpgPoolerGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ f := cnpgPoolerFactsOf(ctx, c, pooler)
+ one := func(guard string, g auth.Grant) integration.ActionCapability {
+ g = g.In(namespace)
+ return integration.CapabilityVerdict(guard, []string{s.Access.Permission(ctx, g)}, []auth.Grant{g})
+ }
+ return &CNPGPoolerCapabilitiesResponse{
+ UID: string(pooler.GetUID()),
+ ResourceVersion: pooler.GetResourceVersion(),
+ Context: contextName,
+ Facts: f,
+ Actions: CNPGPoolerActions{
+ Pause: one(cnpgGuardPooler(f, true), cnpgGrantPatchPool),
+ Resume: one(cnpgGuardPooler(f, false), cnpgGrantPatchPool),
+ ObserveState: one("", grantCreateExec),
+ },
+ }, nil
+}
+
+type cnpgPoolerReviewed struct {
+ Paused *bool `json:"paused"`
+}
+
+func RunCNPGPoolerAction(ctx context.Context, c ActionClients, namespace, name, action string, req integration.ActionRequest) (*CNPGActionResult, error) {
+ var reviewed cnpgPoolerReviewed
+ if len(req.Facts) > 0 {
+ if err := json.Unmarshal(req.Facts, &reviewed); err != nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts: %v", err)
+ }
+ }
+ if reviewed.Paused == nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "facts.paused is required for %s: the confirmation must bind what the dialog showed", action)
+ }
+ if err := integration.DecodeActionParams(req.Params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ pooler, err := c.Dynamic.Resource(cnpgPoolerGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ facts := cnpgPoolerFactsOf(ctx, ActionClients{Dynamic: c.Dynamic}, pooler)
+ if string(pooler.GetUID()) != req.UID {
+ return nil, integration.ChangedAction(facts, "Pooler %s/%s was deleted and recreated since you reviewed it", namespace, name)
+ }
+ if *reviewed.Paused != facts.Paused {
+ return nil, integration.ChangedAction(facts, "Pooler %s/%s changed since you confirmed (paused); review the action again", namespace, name)
+ }
+ pause := action == "pause"
+ if r := cnpgGuardPooler(facts, pause); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ err = integration.MergePatchAtVersion(ctx, c.Dynamic, cnpgPoolerGVR, pooler, map[string]any{"spec": map[string]any{"pgbouncer": map[string]any{"paused": pause}}})
+ if apierrors.IsConflict(err) {
+ if fresh, gerr := c.Dynamic.Resource(cnpgPoolerGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{}); gerr == nil {
+ facts = cnpgPoolerFactsOf(ctx, ActionClients{Dynamic: c.Dynamic}, fresh)
+ }
+ return nil, integration.ChangedAction(facts, "Pooler %s/%s changed while the request was being sent; review the action again", namespace, name)
+ }
+ if err != nil {
+ return nil, err
+ }
+ msg := "Pause requested: each PgBouncer finishes its running transactions, then holds new queries"
+ if !pause {
+ msg = "Resume requested: each PgBouncer serves queued clients again"
+ }
+ return &CNPGActionResult{
+ Action: action,
+ Message: msg,
+ Target: &CNPGActionTarget{Paused: &pause, Generation: pooler.GetGeneration() + 1},
+ }, nil
+}
+
+// CNPGPgBouncerState is one pooler Pod's SHOW STATE. Facts only when read.
+type CNPGPgBouncerState struct {
+ Pod string `json:"pod"`
+ CNPGRuntimeSource
+ Paused *bool `json:"paused,omitempty"`
+ Suspended *bool `json:"suspended,omitempty"`
+ Active *bool `json:"active,omitempty"`
+}
+
+// CNPGPgBouncerStateResponse is GET /api/cnpg/poolers/{ns}/{name}/pgbouncer-state.
+type CNPGPgBouncerStateResponse struct {
+ Pooler CNPGRuntimeObjectRef `json:"pooler"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGExecPermission `json:"permission"`
+ Pods []CNPGPgBouncerState `json:"pods"`
+}
+
+// PgBouncer's admin console on its unix socket; the container's PGHOST,
+// PGUSER and PGDATABASE already point there.
+var cnpgShowStateArgv = []string{"psql", "-XAtq", "-c", "SHOW STATE"}
+
+func parseCNPGShowState(out []byte) (CNPGPgBouncerState, error) {
+ var st CNPGPgBouncerState
+ for _, line := range strings.Split(strings.TrimSpace(string(out)), "\n") {
+ k, v, ok := strings.Cut(strings.TrimSpace(line), "|")
+ if !ok {
+ continue
+ }
+ b := v == "yes"
+ switch k {
+ case "paused":
+ st.Paused = &b
+ case "suspended":
+ st.Suspended = &b
+ case "active":
+ st.Active = &b
+ }
+ }
+ if st.Paused == nil {
+ return st, fmt.Errorf("SHOW STATE did not report paused")
+ }
+ return st, nil
+}
+
+func readCNPGPgBouncerState(ctx context.Context, exec ExecFunc, namespace string, pod *corev1.Pod) CNPGPgBouncerState {
+ captured := time.Now().UTC().Format(time.RFC3339)
+ out, err := exec(ctx, namespace, pod.Name, pgBouncerContainer, cnpgShowStateArgv, "")
+ if err != nil {
+ src := cnpgExecSourceState(err)
+ src.CapturedAt = captured
+ if reason := poolerNotStarted(pod); reason != "" && src.State != execStateDenied {
+ src.State, src.Error = runtimeStateUnreachable, reason
+ }
+ return CNPGPgBouncerState{Pod: pod.Name, CNPGRuntimeSource: src}
+ }
+ st, err := parseCNPGShowState(out)
+ st.Pod = pod.Name
+ if err != nil {
+ st.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateError, Error: err.Error(), CapturedAt: captured}
+ return st
+ }
+ st.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateOK, CapturedAt: captured}
+ return st
+}
+
+func (s *Reader) PgBouncerState(ctx context.Context, cache *k8s.ResourceCache, pooler *unstructured.Unstructured) (*CNPGPgBouncerStateResponse, error) {
+ namespace, name := pooler.GetNamespace(), pooler.GetName()
+ pods, err := poolerPods(cache, pooler)
+ if err != nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "pooler Pods unavailable: " + err.Error()}
+ }
+ resp := CNPGPgBouncerStateResponse{
+ Pooler: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: pooler.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: CNPGExecPermission{Exec: s.Access.Permission(ctx, grantCreateExec.In(namespace)), Grant: grantCreateExec.In(namespace).Ref()},
+ Pods: make([]CNPGPgBouncerState, len(pods)),
+ }
+ for i, p := range pods {
+ resp.Pods[i].Pod = p.Name
+ }
+ if resp.Permission.Exec == integration.PermissionDenied {
+ for i := range resp.Pods {
+ resp.Pods[i].CNPGRuntimeSource = CNPGRuntimeSource{State: execStateDenied, Error: "reading PgBouncer state needs " + integration.GrantText(resp.Permission.Grant)}
+ }
+ return &resp, nil
+ }
+ exec := s.Clients.Exec
+ if exec == nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "cluster client not available — check cluster connection"}
+ }
+ run := newCNPGRuntimeRunner(ctx)
+ for i, p := range pods {
+ out := &resp.Pods[i]
+ run.do(func(ctx context.Context) {
+ *out = readCNPGPgBouncerState(ctx, exec, namespace, p)
+ })
+ }
+ run.wait()
+ for _, p := range resp.Pods {
+ if p.State == execStateDenied {
+ resp.Permission.Exec = integration.PermissionDenied
+ }
+ }
+ return &resp, nil
+}
diff --git a/internal/cnpg/ports.go b/internal/cnpg/ports.go
new file mode 100644
index 0000000000..b2c3677bd5
--- /dev/null
+++ b/internal/cnpg/ports.go
@@ -0,0 +1,52 @@
+package cnpg
+
+import (
+ "context"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/client-go/kubernetes"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+ bp "github.com/skyhook-io/radar/pkg/audit"
+)
+
+type Access struct {
+ CanRead func(context.Context, string, string, string, string) bool
+ Permission func(context.Context, auth.Grant) string
+ MetricsRead func(context.Context, string, string, string, string) bool
+}
+
+type Observations struct {
+ FilterAudit func(*bp.ScanResults) *bp.ScanResults
+
+ WorkspaceRead func(context.Context, *k8s.ResourceCache, integration.WorkspaceKind, []string, []string) (integration.KindAccess, []*unstructured.Unstructured)
+ Issues func(context.Context, []string) []issues.Issue
+
+ Cache *k8s.ResourceCache
+ Connected bool
+ Discovery *k8s.ResourceDiscovery
+ Cluster func(context.Context, string, string, ...auth.Grant) (*k8s.ResourceCache, *unstructured.Unstructured, error)
+ OperatorScope func(context.Context) []string
+ TypedScope func(context.Context, *k8s.ResourceCache, []string, string, string) (integration.KindAccess, []string)
+ DynamicList func(context.Context, *k8s.ResourceCache, string, string, string) ([]*unstructured.Unstructured, error)
+}
+
+type ReadClients struct {
+ Exec ExecFunc
+
+ Typed kubernetes.Interface
+ Proxy kubernetes.Interface
+}
+
+type Reader struct {
+ Metrics Metrics
+
+ Access Access
+ Observations Observations
+ Clients ReadClients
+ Identity string
+ ClusterContext string
+}
diff --git a/internal/cnpg/protection.go b/internal/cnpg/protection.go
new file mode 100644
index 0000000000..c5b2df0ea4
--- /dev/null
+++ b/internal/cnpg/protection.go
@@ -0,0 +1,419 @@
+package cnpg
+
+import (
+ "context"
+ "crypto/sha256"
+ "encoding/hex"
+ "encoding/json"
+ "fmt"
+ "net/http"
+ "net/url"
+ "reflect"
+ "slices"
+ "strings"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/apimachinery/pkg/util/validation"
+ "k8s.io/client-go/dynamic"
+
+ "github.com/skyhook-io/radar/internal/integration"
+ declarations "github.com/skyhook-io/radar/pkg/cnpg"
+ "github.com/skyhook-io/radar/pkg/urlutil"
+)
+
+type ArchivingParams struct {
+ ObjectStore string `json:"objectStore"`
+ ServerName string `json:"serverName"`
+ AcknowledgeArchive bool `json:"acknowledgeArchive"`
+}
+
+// Digests bind configuration without returning credential-bearing provider configuration.
+type ArchivingFacts struct {
+ ClusterConfig string `json:"clusterConfig"`
+ ObjectStoreUID string `json:"objectStoreUID"`
+ ObjectStoreConfig string `json:"objectStoreConfig"`
+ ObjectStore string `json:"objectStore"`
+ ServerName string `json:"serverName"`
+}
+
+type ArchivingPreview struct {
+ Context string `json:"context"`
+ UID string `json:"uid"`
+ Facts ArchivingFacts `json:"facts"`
+ Destination string `json:"destination"`
+ Endpoint string `json:"endpoint,omitempty"`
+ Plugin map[string]any `json:"plugin"`
+ Warnings []string `json:"warnings"`
+ Unchanged bool `json:"unchanged"`
+}
+
+func configDigest(value any) (string, error) {
+ data, err := json.Marshal(value)
+ if err != nil {
+ return "", err
+ }
+ sum := sha256.Sum256(data)
+ return hex.EncodeToString(sum[:]), nil
+}
+
+func protectionConfig(cluster *unstructured.Unstructured) map[string]any {
+ spec, _, _ := unstructured.NestedMap(cluster.Object, "spec")
+ return map[string]any{"plugins": spec["plugins"], "backup": spec["backup"], "bootstrap": spec["bootstrap"], "externalClusters": spec["externalClusters"], "replica": spec["replica"], "deleting": !cluster.GetDeletionTimestamp().IsZero()}
+}
+
+func archivingPlugin(cluster *unstructured.Unstructured, p ArchivingParams) ([]any, map[string]any, error) {
+ if !cluster.GetDeletionTimestamp().IsZero() {
+ return nil, nil, integration.BlockedAction("The Cluster is being deleted")
+ }
+ d := declarations.ParseBackupDeclaration(cluster)
+ if d.InTreeConfigured || d.SnapshotsConfigured {
+ return nil, nil, integration.BlockedAction("This Cluster declares in-tree or snapshot backups. Review a method migration in Cluster YAML instead of attaching a plugin here")
+ }
+ plugins, _, err := unstructured.NestedSlice(cluster.Object, "spec", "plugins")
+ if err != nil {
+ return nil, nil, err
+ }
+ index := -1
+ for i, raw := range plugins {
+ entry, ok := raw.(map[string]any)
+ if !ok {
+ return nil, nil, integration.BlockedAction("The plugin configuration needs review in Cluster YAML")
+ }
+ name, _ := entry["name"].(string)
+ if name == declarations.BarmanPluginName {
+ if index >= 0 {
+ return nil, nil, integration.BlockedAction("More than one Barman plugin entry is declared; review Cluster YAML")
+ }
+ index = i
+ } else if entry["enabled"] != false && entry["isWALArchiver"] == true {
+ return nil, nil, integration.BlockedAction("Another WAL archiver is declared; this guide does not replace it")
+ }
+ }
+ plugin := map[string]any{"name": declarations.BarmanPluginName}
+ if index >= 0 {
+ plugin = plugins[index].(map[string]any)
+ }
+ params, _ := plugin["parameters"].(map[string]any)
+ if params == nil {
+ params = map[string]any{}
+ }
+ oldStore, _ := params["barmanObjectName"].(string)
+ oldServer, _ := params["serverName"].(string)
+ if oldServer == "" {
+ oldServer = cluster.GetName()
+ }
+ if oldStore != "" && (oldStore != p.ObjectStore || oldServer != p.ServerName) {
+ return nil, nil, integration.BlockedAction("An archive destination is already declared. Changing that identity needs a migration reviewed in Cluster YAML")
+ }
+ params["barmanObjectName"], params["serverName"] = p.ObjectStore, p.ServerName
+ plugin["parameters"], plugin["isWALArchiver"], plugin["enabled"] = params, true, true
+ if index < 0 {
+ plugins = append(plugins, plugin)
+ } else {
+ plugins[index] = plugin
+ }
+ return plugins, plugin, nil
+}
+
+type archiveLocation struct {
+ path, endpoint string
+ key struct{ path, endpoint string }
+}
+
+func locationOf(configuration map[string]any, server string) (archiveLocation, error) {
+ destination, _ := configuration["destinationPath"].(string)
+ u, err := url.Parse(destination)
+ if err != nil || u.Scheme == "" || u.Host == "" || u.User != nil || u.RawQuery != "" || u.Fragment != "" {
+ return archiveLocation{}, integration.BlockedAction("The ObjectStore needs a valid destinationPath without embedded credentials or query parameters")
+ }
+ u.Path = strings.TrimRight(u.Path, "/") + "/" + server
+ u.RawPath = ""
+ endpoint, _ := configuration["endpointURL"].(string)
+ endpointKey := ""
+ if endpoint != "" {
+ e, err := url.Parse(endpoint)
+ if err != nil || e.Host == "" || e.User != nil || e.RawQuery != "" || e.Fragment != "" {
+ return archiveLocation{}, integration.BlockedAction("The ObjectStore endpoint needs review in its YAML")
+ }
+ origin, valid := urlutil.NormalizeOrigin(e.String())
+ if !valid {
+ return archiveLocation{}, integration.BlockedAction("The ObjectStore endpoint needs review in its YAML")
+ }
+ endpointKey = origin + strings.TrimRight(e.EscapedPath(), "/")
+ endpoint = strings.TrimRight(e.String(), "/")
+ }
+ origin, valid := urlutil.NormalizeOrigin(u.String())
+ if !valid {
+ return archiveLocation{}, integration.BlockedAction("The ObjectStore destination needs review in its YAML")
+ }
+ location := archiveLocation{path: u.String(), endpoint: endpoint}
+ location.key.path, location.key.endpoint = origin+u.EscapedPath(), endpointKey
+ return location, nil
+}
+
+func readStore(ctx context.Context, dyn dynamic.Interface, namespace, name string) (*unstructured.Unstructured, map[string]any, error) {
+ store, err := dyn.Resource(cnpgObjectStoreGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, err
+ }
+ config, _, err := unstructured.NestedMap(store.Object, "spec", "configuration")
+ if err != nil {
+ return nil, nil, err
+ }
+ return store, config, nil
+}
+
+func checkArchiveIdentity(ctx context.Context, dyn dynamic.Interface, cluster, store *unstructured.Unstructured, chosen archiveLocation, p ArchivingParams) ([]string, error) {
+ warnings := []string{"Radar checks visible declarations, not remote archive contents or credential validity. Confirm this archive identity is reserved for this Cluster.", "Enabling the plugin can reconcile instance Pods and sidecars; inspect the operator and instance progress afterward."}
+ namespace := cluster.GetNamespace()
+ location := func(ns, name, server string) (archiveLocation, error) {
+ if ns == namespace && name == store.GetName() {
+ config, _, err := unstructured.NestedMap(store.Object, "spec", "configuration")
+ if err != nil {
+ return archiveLocation{}, err
+ }
+ return locationOf(config, server)
+ }
+ _, config, err := readStore(ctx, dyn, ns, name)
+ if err != nil {
+ return archiveLocation{}, err
+ }
+ return locationOf(config, server)
+ }
+ source, _, _ := unstructured.NestedString(cluster.Object, "spec", "bootstrap", "recovery", "source")
+ replicaSource, _, _ := unstructured.NestedString(cluster.Object, "spec", "replica", "source")
+ sources := []string{}
+ if source != "" {
+ sources = append(sources, source)
+ }
+ if replicaSource != "" && replicaSource != source {
+ sources = append(sources, replicaSource)
+ }
+ backupName, _, _ := unstructured.NestedString(cluster.Object, "spec", "bootstrap", "recovery", "backup", "name")
+ if backupName != "" {
+ backup, err := dyn.Resource(cnpgBackupGVR).Namespace(namespace).Get(ctx, backupName, metav1.GetOptions{})
+ if err != nil {
+ return nil, integration.BlockedAction("The recovery-source Backup could not be read to compare archive identities")
+ }
+ method, _, _ := unstructured.NestedString(backup.Object, "status", "method")
+ if method == "" {
+ method, _, _ = unstructured.NestedString(backup.Object, "spec", "method")
+ }
+ if method != "volumeSnapshot" {
+ configuration, _, _ := unstructured.NestedMap(backup.Object, "status")
+ server, _ := configuration["serverName"].(string)
+ if server == "" {
+ return nil, integration.BlockedAction("The Backup does not report its source archive identity; review archive isolation in Cluster YAML")
+ }
+ backupLocation, err := locationOf(configuration, server)
+ if err != nil {
+ return nil, integration.BlockedAction("The Backup does not report a comparable source archive location; review archive isolation in Cluster YAML")
+ }
+ if backupLocation.key == chosen.key {
+ return nil, integration.BlockedAction("This is the recovery-source archive. Choose a new server name")
+ }
+ }
+ }
+ external, _, _ := unstructured.NestedSlice(cluster.Object, "spec", "externalClusters")
+ matched := []string{}
+ for _, raw := range external {
+ e, ok := raw.(map[string]any)
+ name, _ := e["name"].(string)
+ if !ok || !slices.Contains(sources, name) {
+ continue
+ }
+ matched = append(matched, name)
+ plugin, _ := e["plugin"].(map[string]any)
+ if plugin["name"] != declarations.BarmanPluginName {
+ configuration, _ := e["barmanObjectStore"].(map[string]any)
+ if configuration == nil {
+ return nil, integration.BlockedAction("The external recovery source has no comparable archive location; review archive isolation in Cluster YAML")
+ }
+ server, _ := configuration["serverName"].(string)
+ if server == "" {
+ server = name
+ }
+ inTree, err := locationOf(configuration, server)
+ if err != nil {
+ return nil, err
+ }
+ if inTree.key == chosen.key {
+ return nil, integration.BlockedAction("This is the recovery-source archive. Choose a new server name")
+ }
+ continue
+ }
+ params, _ := plugin["parameters"].(map[string]any)
+ sourceStore, _ := params["barmanObjectName"].(string)
+ server, _ := params["serverName"].(string)
+ if server == "" {
+ server = name
+ }
+ sourceLocation, err := location(namespace, sourceStore, server)
+ if err != nil {
+ return nil, integration.BlockedAction("The recovery-source ObjectStore could not be read to compare archive identities; review its access and configuration first")
+ }
+ if sourceLocation.key == chosen.key {
+ return nil, integration.BlockedAction("This is the recovery-source archive. Choose a new server name so this Cluster cannot write into the archive it restores from")
+ }
+ }
+ for _, source := range sources {
+ if !slices.Contains(matched, source) {
+ return nil, integration.BlockedAction("The declared recovery source is unresolved; review archive isolation in Cluster YAML")
+ }
+ }
+ windows, _, _ := unstructured.NestedMap(store.Object, "status", "serverRecoveryWindow")
+ if _, exists := windows[p.ServerName]; exists {
+ current, ok := declarations.ParseBackupDeclaration(cluster).BarmanPlugin()
+ if !ok || current.ObjectStore != p.ObjectStore || current.ServerName != p.ServerName {
+ return nil, integration.BlockedAction("The ObjectStore reports backups for this server name already. Choose a new archive identity")
+ }
+ }
+ clusters, err := dyn.Resource(ClusterGVR).List(ctx, metav1.ListOptions{})
+ if err != nil {
+ if !apierrors.IsForbidden(err) {
+ return nil, err
+ }
+ warnings = append(warnings, "Cluster inventory is not readable across every namespace; other archive users may be out of view.")
+ clusters, err = dyn.Resource(ClusterGVR).Namespace(namespace).List(ctx, metav1.ListOptions{})
+ if err != nil {
+ if !apierrors.IsForbidden(err) {
+ return nil, err
+ }
+ return append(warnings, "Cluster inventory in this namespace is also unreadable; collision checks are incomplete."), nil
+ }
+ }
+ for i := range clusters.Items {
+ other := &clusters.Items[i]
+ if other.GetUID() == cluster.GetUID() {
+ continue
+ }
+ declared := declarations.ParseBackupDeclaration(other)
+ plugin, ok := declared.BarmanPlugin()
+ var otherLocation archiveLocation
+ var err error
+ if ok && plugin.ObjectStore != "" {
+ otherLocation, err = location(other.GetNamespace(), plugin.ObjectStore, plugin.ServerName)
+ } else if declared.InTreeDestination != "" {
+ configuration, _, _ := unstructured.NestedMap(other.Object, "spec", "backup", "barmanObjectStore")
+ server, _ := configuration["serverName"].(string)
+ if server == "" {
+ server = other.GetName()
+ }
+ otherLocation, err = locationOf(configuration, server)
+ } else {
+ continue
+ }
+ if err != nil {
+ if apierrors.IsForbidden(err) || apierrors.IsNotFound(err) {
+ warnings = append(warnings, fmt.Sprintf("Archive of Cluster %s/%s cannot be compared; inventory is incomplete.", other.GetNamespace(), other.GetName()))
+ continue
+ }
+ return nil, err
+ }
+ if otherLocation.key == chosen.key {
+ return nil, integration.BlockedAction("Cluster " + other.GetNamespace() + "/" + other.GetName() + " already archives to this destination and server name")
+ }
+ }
+ return warnings, nil
+}
+
+func prepareArchiving(ctx context.Context, c ActionClients, namespace, name string, p ArchivingParams) (*unstructured.Unstructured, map[string]any, ArchivingPreview, error) {
+ var out ArchivingPreview
+ if len(validation.IsDNS1123Subdomain(p.ObjectStore)) > 0 || len(validation.IsDNS1123Label(p.ServerName)) > 0 {
+ return nil, nil, out, integration.RefuseAction(http.StatusBadRequest, "", "Choose an ObjectStore name and a server name (a lowercase DNS label, up to 63 characters)")
+ }
+ cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, out, err
+ }
+ clusterDigest, err := configDigest(protectionConfig(cluster))
+ if err != nil {
+ return nil, nil, out, err
+ }
+ plugins, plugin, err := archivingPlugin(cluster, p)
+ if err != nil {
+ return nil, nil, out, err
+ }
+ store, configuration, err := readStore(ctx, c.Dynamic, namespace, p.ObjectStore)
+ if err != nil {
+ return nil, nil, out, err
+ }
+ if !store.GetDeletionTimestamp().IsZero() {
+ return nil, nil, out, integration.BlockedAction("The ObjectStore is being deleted")
+ }
+ if server, _ := configuration["serverName"].(string); server != "" {
+ return nil, nil, out, integration.BlockedAction("ObjectStore configuration.serverName must be empty; use the Cluster plugin parameter instead")
+ }
+ storeDigest, err := configDigest(store.Object["spec"])
+ if err != nil {
+ return nil, nil, out, err
+ }
+ chosen, err := locationOf(configuration, p.ServerName)
+ if err != nil {
+ return nil, nil, out, err
+ }
+ warnings, err := checkArchiveIdentity(ctx, c.Dynamic, cluster, store, chosen, p)
+ if err != nil {
+ return nil, nil, out, err
+ }
+ old, _, _ := unstructured.NestedSlice(cluster.Object, "spec", "plugins")
+ out = ArchivingPreview{UID: string(cluster.GetUID()), Facts: ArchivingFacts{ClusterConfig: clusterDigest, ObjectStoreUID: string(store.GetUID()), ObjectStoreConfig: storeDigest, ObjectStore: p.ObjectStore, ServerName: p.ServerName}, Destination: chosen.path, Endpoint: chosen.endpoint, Plugin: plugin, Warnings: warnings, Unchanged: reflect.DeepEqual(old, plugins)}
+ return cluster, map[string]any{"spec": map[string]any{"plugins": plugins}}, out, nil
+}
+
+func (s *Reader) PreviewArchiving(ctx context.Context, c ActionClients, contextName, namespace, name string, p ArchivingParams) (*ArchivingPreview, error) {
+ cluster, patch, out, err := prepareArchiving(ctx, c, namespace, name, p)
+ if err != nil {
+ return nil, err
+ }
+ if !out.Unchanged {
+ patch["metadata"] = map[string]any{"resourceVersion": cluster.GetResourceVersion()}
+ data, err := json.Marshal(patch)
+ if err != nil {
+ return nil, err
+ }
+ _, err = c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Patch(ctx, name, types.MergePatchType, data, metav1.PatchOptions{DryRun: []string{metav1.DryRunAll}})
+ if err != nil {
+ return nil, err
+ }
+ }
+ out.Context = contextName
+ return &out, nil
+}
+
+func runConfigureArchiving(ctx context.Context, c ActionClients, namespace, name string, req integration.ActionRequest) (*CNPGActionResult, error) {
+ var params ArchivingParams
+ if err := integration.DecodeActionParams(req.Params, ¶ms); err != nil {
+ return nil, err
+ }
+ if !params.AcknowledgeArchive {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "Confirm the archive identity is reserved for this Cluster before enabling archiving")
+ }
+ var reviewed ArchivingFacts
+ if err := integration.DecodeActionParams(req.Facts, &reviewed); err != nil {
+ return nil, err
+ }
+ if reviewed.ClusterConfig == "" || reviewed.ObjectStoreUID == "" || reviewed.ObjectStoreConfig == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "Reviewed archiving facts are required")
+ }
+ cluster, patch, current, err := prepareArchiving(ctx, c, namespace, name, params)
+ if err != nil {
+ return nil, err
+ }
+ if string(cluster.GetUID()) != req.UID || current.Facts != reviewed {
+ return nil, integration.ChangedAction(current, "Cluster or ObjectStore configuration changed since review; review the attachment again")
+ }
+ if current.Unchanged {
+ return &CNPGActionResult{Action: "configureArchiving", Message: "Archiving configuration is unchanged"}, nil
+ }
+ if err := integration.MergePatchAtVersion(ctx, c.Dynamic, ClusterGVR, cluster, patch); err != nil {
+ if apierrors.IsConflict(err) {
+ return nil, integration.ChangedAction(current, "The Cluster changed while the patch was sent; review it again")
+ }
+ return nil, err
+ }
+ return &CNPGActionResult{Action: "configureArchiving", Message: "Archiving configuration saved; verify WAL uploads and a successful base backup"}, nil
+}
diff --git a/internal/cnpg/protection_test.go b/internal/cnpg/protection_test.go
new file mode 100644
index 0000000000..40a08aa0d2
--- /dev/null
+++ b/internal/cnpg/protection_test.go
@@ -0,0 +1,328 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "strings"
+ "testing"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ k8stesting "k8s.io/client-go/testing"
+
+ "github.com/skyhook-io/radar/internal/integration"
+)
+
+func protectionStore(name string) *unstructured.Unstructured {
+ return &unstructured.Unstructured{Object: map[string]any{"apiVersion": "barmancloud.cnpg.io/v1", "kind": "ObjectStore", "metadata": map[string]any{"name": name, "namespace": "db", "uid": name + "-uid", "resourceVersion": "3"}, "spec": map[string]any{"configuration": map[string]any{"destinationPath": "s3://bucket/prefix/", "endpointURL": "https://storage.example", "s3Credentials": map[string]any{"inheritFromIAMRole": true}}}}}
+}
+
+func unprotectedCluster() *unstructured.Unstructured {
+ return cnpgActionCluster(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ delete(spec, "backup")
+ spec["plugins"] = []any{map[string]any{"name": "metrics.example", "parameters": map[string]any{"mode": "keep"}}}
+ })
+}
+
+func archivingRequest(t *testing.T, p *ArchivingPreview) integration.ActionRequest {
+ t.Helper()
+ facts, _ := json.Marshal(p.Facts)
+ params, _ := json.Marshal(ArchivingParams{ObjectStore: p.Facts.ObjectStore, ServerName: p.Facts.ServerName, AcknowledgeArchive: true})
+ return integration.ActionRequest{ReviewedContext: p.Context, UID: p.UID, Facts: facts, Params: params}
+}
+
+func TestArchivingReviewAndConditionalWrite(t *testing.T) {
+ ctx := context.Background()
+ env := newCNPGActionEnv(t, []runtime.Object{unprotectedCluster(), protectionStore("store")})
+ p, err := newTestReader(nil).PreviewArchiving(ctx, env.clients(), "ctx", "db", "pg", ArchivingParams{ObjectStore: "store", ServerName: "pg-new"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if p.Destination != "s3://bucket/prefix/pg-new" || p.Endpoint != "https://storage.example" {
+ t.Fatalf("comparison keys leaked into review: %+v", p)
+ }
+ if len(env.patches) != 1 || len(env.patches[0].(k8stesting.PatchActionImpl).GetPatchOptions().DryRun) != 1 {
+ t.Fatal("preview must be dry-run")
+ }
+ // Status churn is deliberately outside the reviewed configuration.
+ fresh, err := env.dyn.Resource(ClusterGVR).Namespace("db").Get(ctx, "pg", metav1.GetOptions{})
+ if err != nil {
+ t.Fatal(err)
+ }
+ fresh.Object["status"] = map[string]any{"phase": "changing"}
+ fresh.SetResourceVersion("99")
+ if err := env.dyn.Tracker().Update(ClusterGVR, fresh, "db"); err != nil {
+ t.Fatal(err)
+ }
+ res, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "configureArchiving", archivingRequest(t, p))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if res.Action != "configureArchiving" || len(env.patches) != 2 {
+ t.Fatalf("result=%+v writes=%d", res, len(env.patches))
+ }
+ body := cnpgActionPatchBody(t, env.patches[1])
+ if body["metadata"].(map[string]any)["resourceVersion"] != "99" {
+ t.Fatal("must bind fresh version")
+ }
+ plugins := body["spec"].(map[string]any)["plugins"].([]any)
+ if len(plugins) != 2 || plugins[0].(map[string]any)["parameters"].(map[string]any)["mode"] != "keep" {
+ t.Fatalf("unrelated plugin changed: %v", plugins)
+ }
+ if len(env.patches[1].(k8stesting.PatchActionImpl).GetPatchOptions().DryRun) != 0 {
+ t.Fatal("confirmed write remained a dry-run")
+ }
+}
+
+func TestArchivingExistingPluginEnablement(t *testing.T) {
+ for _, scenario := range []struct {
+ name string
+ enabled bool
+ wal bool
+ }{
+ {name: "base-backup-only", enabled: true},
+ {name: "disabled"},
+ {name: "already-archiving", enabled: true, wal: true},
+ } {
+ t.Run(scenario.name, func(t *testing.T) {
+ ctx := context.Background()
+ cluster := unprotectedCluster()
+ plugin := map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "enabled": scenario.enabled, "isWALArchiver": scenario.wal,
+ "parameters": map[string]any{"barmanObjectName": "store", "serverName": "pg", "custom": "keep"},
+ }
+ cluster.Object["spec"].(map[string]any)["plugins"] = []any{plugin}
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, protectionStore("store")})
+ preview, err := newTestReader(nil).PreviewArchiving(ctx, env.clients(), "ctx", "db", "pg", ArchivingParams{ObjectStore: "store", ServerName: "pg"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if preview.Unchanged != (scenario.enabled && scenario.wal) {
+ t.Fatalf("unchanged=%v for %+v", preview.Unchanged, scenario)
+ }
+ if plugin["enabled"] != scenario.enabled || plugin["isWALArchiver"] != scenario.wal {
+ t.Fatal("preview mutated the source plugin")
+ }
+ env.patches = nil
+ if _, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "configureArchiving", archivingRequest(t, preview)); err != nil {
+ t.Fatal(err)
+ }
+ if preview.Unchanged {
+ if len(env.patches) != 0 {
+ t.Fatal("already enabled archiving should be a no-op")
+ }
+ return
+ }
+ if len(env.patches) != 1 {
+ t.Fatalf("expected one enablement write, got %d", len(env.patches))
+ }
+ patch := env.patches[0].(k8stesting.PatchActionImpl)
+ if len(patch.GetPatchOptions().DryRun) != 0 {
+ t.Fatal("enablement remained a dry-run")
+ }
+ plugins := cnpgActionPatchBody(t, env.patches[0])["spec"].(map[string]any)["plugins"].([]any)
+ updated := plugins[0].(map[string]any)
+ if updated["enabled"] != true || updated["isWALArchiver"] != true || updated["parameters"].(map[string]any)["custom"] != "keep" {
+ t.Fatalf("enablement or preservation failed: %v", updated)
+ }
+ })
+ }
+}
+
+func TestArchivingRefusesReviewedConfigurationChanges(t *testing.T) {
+ for _, changed := range []string{"cluster", "store", "recreated", "race"} {
+ t.Run(changed, func(t *testing.T) {
+ ctx := context.Background()
+ env := newCNPGActionEnv(t, []runtime.Object{unprotectedCluster(), protectionStore("store")})
+ p, err := newTestReader(nil).PreviewArchiving(ctx, env.clients(), "ctx", "db", "pg", ArchivingParams{ObjectStore: "store", ServerName: "pg-new"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ env.patches = nil
+ if changed == "race" {
+ env.dyn.PrependReactor("patch", "clusters", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewConflict(ClusterGVR.GroupResource(), "pg", nil)
+ })
+ } else if changed == "store" {
+ s := protectionStore("store")
+ s.Object["spec"].(map[string]any)["retentionPolicy"] = "8d"
+ if err := env.dyn.Tracker().Update(cnpgObjectStoreGVR, s, "db"); err != nil {
+ t.Fatal(err)
+ }
+ } else {
+ c := unprotectedCluster()
+ if changed == "recreated" {
+ c.SetUID(types.UID("new-uid"))
+ } else {
+ c.Object["spec"].(map[string]any)["plugins"] = []any{map[string]any{"name": "different.example"}}
+ }
+ if err := env.dyn.Tracker().Update(ClusterGVR, c, "db"); err != nil {
+ t.Fatal(err)
+ }
+ }
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "configureArchiving", archivingRequest(t, p))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged || len(env.patches) != 0 {
+ t.Fatalf("err=%v patches=%d", err, len(env.patches))
+ }
+ })
+ }
+}
+
+func TestArchivingIdentityAndMigrationGuards(t *testing.T) {
+ for _, scenario := range []string{"recovery-alias", "other-cluster-alias", "other-base-backup-only", "origin-alias", "historical-server", "in-tree", "other-archiver", "duplicate", "store-server"} {
+ t.Run(scenario, func(t *testing.T) {
+ c, store, alias := unprotectedCluster(), protectionStore("store"), protectionStore("alias")
+ objs := []runtime.Object{c, store, alias}
+ spec := c.Object["spec"].(map[string]any)
+ switch scenario {
+ case "recovery-alias":
+ spec["bootstrap"] = map[string]any{"recovery": map[string]any{"source": "origin"}}
+ spec["externalClusters"] = []any{map[string]any{"name": "origin", "plugin": map[string]any{"name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "alias", "serverName": "pg-new"}}}}
+ case "other-cluster-alias", "other-base-backup-only", "origin-alias":
+ if scenario == "origin-alias" {
+ config := alias.Object["spec"].(map[string]any)["configuration"].(map[string]any)
+ config["destinationPath"], config["endpointURL"] = "s3://BUCKET/prefix", "https://STORAGE.EXAMPLE:443/"
+ }
+ other := unprotectedCluster()
+ other.SetName("other")
+ other.SetUID("other-uid")
+ other.Object["spec"].(map[string]any)["plugins"] = []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io", "isWALArchiver": scenario != "other-base-backup-only", "parameters": map[string]any{"barmanObjectName": "alias", "serverName": "pg-new"}}}
+ objs = append(objs, other)
+ case "historical-server":
+ store.Object["status"] = map[string]any{"serverRecoveryWindow": map[string]any{"pg-new": map[string]any{}}}
+ case "in-tree":
+ spec["backup"] = map[string]any{"barmanObjectStore": map[string]any{"destinationPath": "s3://old/"}}
+ case "other-archiver":
+ spec["plugins"] = []any{map[string]any{"name": "other", "isWALArchiver": true}}
+ case "duplicate":
+ spec["plugins"] = []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io"}, map[string]any{"name": "barman-cloud.cloudnative-pg.io"}}
+ case "store-server":
+ store.Object["spec"].(map[string]any)["configuration"].(map[string]any)["serverName"] = "old"
+ }
+ env := newCNPGActionEnv(t, objs)
+ _, err := newTestReader(nil).PreviewArchiving(context.Background(), env.clients(), "ctx", "db", "pg", ArchivingParams{ObjectStore: "store", ServerName: "pg-new"})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked || len(env.patches) != 0 {
+ t.Fatalf("err=%v patches=%d", err, len(env.patches))
+ }
+ if scenario == "origin-alias" && !strings.Contains(err.Error(), "already archives to this destination") {
+ t.Fatalf("alias failed for the wrong reason: %v", err)
+ }
+ })
+ }
+}
+
+func TestArchivingPartialInventoryAndForbiddenWrite(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{unprotectedCluster(), protectionStore("store")})
+ env.dyn.PrependReactor("list", "clusters", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewForbidden(ClusterGVR.GroupResource(), "", nil)
+ })
+ p, err := newTestReader(nil).PreviewArchiving(context.Background(), env.clients(), "ctx", "db", "pg", ArchivingParams{ObjectStore: "store", ServerName: "pg-new"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ if !strings.Contains(strings.Join(p.Warnings, " "), "incomplete") {
+ t.Fatalf("warnings=%v", p.Warnings)
+ }
+ env.patches = nil
+ env.dyn.PrependReactor("patch", "clusters", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewForbidden(ClusterGVR.GroupResource(), "pg", nil)
+ })
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "configureArchiving", archivingRequest(t, p))
+ if !apierrors.IsForbidden(err) {
+ t.Fatalf("err=%v", err)
+ }
+}
+
+func TestScheduleMethodRepairPreservesOtherSettings(t *testing.T) {
+ c := cnpgActionCluster(nil)
+ s := cnpgActionSchedule(func(o map[string]any) {
+ spec := o["spec"].(map[string]any)
+ spec["method"] = "barmanObjectStore"
+ spec["online"] = false
+ spec["target"] = "prefer-standby"
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{c, s, protectionStore("store")})
+ p, err := newTestReader(nil).PreviewScheduleMethod(context.Background(), env.clients(), "ctx", "db", "nightly")
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, _ := json.Marshal(p.Facts)
+ env.patches = nil
+ _, err = RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "repairMethod", integration.ActionRequest{ReviewedContext: p.Context, UID: p.UID, Facts: facts})
+ if err != nil {
+ t.Fatal(err)
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ spec := body["spec"].(map[string]any)
+ if len(spec) != 2 || spec["method"] != "plugin" || spec["pluginConfiguration"].(map[string]any)["name"] != "barman-cloud.cloudnative-pg.io" {
+ t.Fatalf("patch=%v", spec)
+ }
+ if body["metadata"].(map[string]any)["resourceVersion"] != "7" {
+ t.Fatal("must bind version")
+ }
+}
+
+func TestScheduleMethodRepairRefusesChangedFactsAndMigrations(t *testing.T) {
+ for _, scenario := range []string{"schedule", "cluster", "recreated", "conflict", "snapshot", "other-plugin", "cluster-deleting", "store-deleting"} {
+ t.Run(scenario, func(t *testing.T) {
+ ctx := context.Background()
+ cluster := cnpgActionCluster(nil)
+ schedule := cnpgActionSchedule(func(o map[string]any) { o["spec"].(map[string]any)["method"] = "barmanObjectStore" })
+ store := protectionStore("store")
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, schedule, store})
+ preview, err := newTestReader(nil).PreviewScheduleMethod(ctx, env.clients(), "ctx", "db", "nightly")
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, _ := json.Marshal(preview.Facts)
+ env.patches = nil
+ want := integration.ActionCodeChanged
+ switch scenario {
+ case "schedule":
+ schedule.Object["spec"].(map[string]any)["cluster"] = map[string]any{"name": "other"}
+ if err := env.dyn.Tracker().Add(cnpgActionCluster(func(o map[string]any) { o["metadata"].(map[string]any)["name"] = "other" })); err != nil {
+ t.Fatal(err)
+ }
+ case "cluster":
+ cluster.Object["spec"].(map[string]any)["plugins"].([]any)[0].(map[string]any)["parameters"].(map[string]any)["serverName"] = "different"
+ case "recreated":
+ schedule.SetUID("new-uid")
+ case "conflict":
+ env.dyn.PrependReactor("patch", "scheduledbackups", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewConflict(ScheduleGVR.GroupResource(), "nightly", nil)
+ })
+ case "snapshot":
+ schedule.Object["spec"].(map[string]any)["method"] = "volumeSnapshot"
+ want = integration.ActionCodeBlocked
+ case "other-plugin":
+ schedule.Object["spec"].(map[string]any)["pluginConfiguration"] = map[string]any{"name": "another.example"}
+ want = integration.ActionCodeBlocked
+ case "cluster-deleting":
+ at := metav1.Now()
+ cluster.SetDeletionTimestamp(&at)
+ want = integration.ActionCodeBlocked
+ case "store-deleting":
+ at := metav1.Now()
+ store.SetDeletionTimestamp(&at)
+ want = integration.ActionCodeBlocked
+ }
+ for _, object := range []struct {
+ gvr schema.GroupVersionResource
+ obj *unstructured.Unstructured
+ }{{ClusterGVR, cluster}, {ScheduleGVR, schedule}, {cnpgObjectStoreGVR, store}} {
+ if err := env.dyn.Tracker().Update(object.gvr, object.obj, "db"); err != nil {
+ t.Fatal(err)
+ }
+ }
+ _, err = RunCNPGScheduleAction(ctx, env.clients(), "db", "nightly", "repairMethod", integration.ActionRequest{ReviewedContext: preview.Context, UID: preview.UID, Facts: facts})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != want || len(env.patches) != 0 {
+ t.Fatalf("err=%v writes=%d", err, len(env.patches))
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/proxy.go b/internal/cnpg/proxy.go
new file mode 100644
index 0000000000..307d6bd56c
--- /dev/null
+++ b/internal/cnpg/proxy.go
@@ -0,0 +1,272 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "io"
+ "log"
+ "net/http"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+)
+
+// The only requests the runtime endpoints can make: a GET of these fixed
+// paths on these fixed ports of a validated Pod. The instance manager's port
+// also serves endpoints that change state (some on GET), so the fixed path and
+// the refusal to follow redirects are the security boundary. Nothing of the
+// caller's request reaches the proxied URL or its headers.
+const (
+ cnpgStatusPort = 8000
+ cnpgMetricsPort = 9187
+ poolerMetricsPort = 9127
+ cnpgStatusPath = "/pg/status"
+ metricsPath = "/metrics"
+
+ cnpgRuntimeRequestTimeout = 5 * time.Second
+ cnpgRuntimeConcurrency = 4
+ cnpgRuntimeMaxRows = 200
+
+ runtimeStateOK = "ok"
+ cnpgRuntimeStatePartial = "partial"
+ runtimeStateDenied = "denied"
+ runtimeStateUnreachable = "unreachable"
+ runtimeStateError = "error"
+
+ cnpgPoolerNameLabel = "cnpg.io/poolerName"
+ pgBouncerContainer = "pgbouncer"
+ cnpgStatusPortTLSFlag = "--status-port-tls"
+ metricsPortTLSFlag = "--metrics-port-tls"
+ cnpgMetricsExporterApp = "cnpg_metrics_exporter"
+ cnpgPgBouncerAdminDB = "pgbouncer"
+ cnpgPgBouncerAuthUser = "cnpg_pooler_pgbouncer"
+ cnpgFencedErrorExplained = "instance is fenced, which asks the operator to stop PostgreSQL"
+)
+
+// Raw-size caps, enforced before parsing, and memo lifetimes sized to the
+// frontend's polling (status ~5s, metrics ~30s). Variables so tests can shrink
+// them.
+var (
+ cnpgRuntimeStatusCap int64 = 1 << 20
+ runtimeMetricsCap int64 = 4 << 20
+ cnpgStatusMemoTTL = 5 * time.Second
+ metricsMemoTTL = 25 * time.Second
+)
+
+// cnpgContainerHasFlag reports whether the running Pod's container was
+// started with a flag — how the operator turns TLS on for a port.
+func containerHasFlag(p *corev1.Pod, container, flag string) bool {
+ for _, c := range p.Spec.Containers {
+ if c.Name != container {
+ continue
+ }
+ for _, words := range [][]string{c.Command, c.Args} {
+ for _, w := range words {
+ if w == flag {
+ return true
+ }
+ }
+ }
+ }
+ return false
+}
+
+func schemeFor(tls bool) string {
+ if tls {
+ return "https"
+ }
+ return "http"
+}
+
+type proxyTarget struct {
+ namespace, pod string
+ podUID types.UID
+ port int
+ path string
+ scheme string
+ limit int64
+}
+
+type cnpgProxyOutcome struct {
+ state string
+ err string
+ scheme string
+ capturedAt string
+ body []byte
+ // truncated: the answer exceeded the cap and body holds only its first
+ // limit bytes.
+ truncated bool
+ // schemeMismatch: the failure was the wrong protocol on the port, the only
+ // failure that justifies trying the other scheme.
+ schemeMismatch bool
+}
+
+func (o cnpgProxyOutcome) source() CNPGRuntimeSource {
+ return CNPGRuntimeSource{State: o.state, Error: o.err, Scheme: o.scheme, CapturedAt: o.capturedAt}
+}
+
+// cnpgProxyGetWithFallback tries the declared scheme and, only when that
+// failed as a protocol mismatch, the other one once: a Pod created by an older
+// operator may disagree with what its object declares today.
+func proxyGetWithFallback(ctx context.Context, client kubernetes.Interface, t proxyTarget) cnpgProxyOutcome {
+ first := cnpgProxyGet(ctx, client, t, t.scheme)
+ if !first.schemeMismatch {
+ return first
+ }
+ other := "https"
+ if t.scheme == "https" {
+ other = "http"
+ }
+ second := cnpgProxyGet(ctx, client, t, other)
+ if second.schemeMismatch {
+ second.state = runtimeStateUnreachable
+ second.err = fmt.Sprintf("neither http nor https worked on port %d: %s", t.port, first.err)
+ second.scheme = ""
+ }
+ return second
+}
+
+// cnpgProxyGet issues one GET through the apiserver's pods/proxy, built only
+// from the validated Pod and the fixed port and path.
+func cnpgProxyGet(ctx context.Context, client kubernetes.Interface, t proxyTarget, scheme string) cnpgProxyOutcome {
+ ctx, cancel := context.WithTimeout(ctx, cnpgRuntimeRequestTimeout)
+ defer cancel()
+ out := cnpgProxyOutcome{scheme: scheme, capturedAt: time.Now().UTC().Format(time.RFC3339)}
+ stream, err := client.CoreV1().RESTClient().Get().
+ Namespace(t.namespace).
+ Resource("pods").
+ Name(fmt.Sprintf("%s:%s:%d", scheme, t.pod, t.port)).
+ SubResource("proxy").
+ Suffix(t.path).
+ Stream(ctx)
+ if err != nil {
+ return classifyCNPGProxyFailure(ctx, err, out, t)
+ }
+ defer stream.Close()
+ body, err := io.ReadAll(io.LimitReader(stream, t.limit+1))
+ if err != nil {
+ return classifyCNPGProxyFailure(ctx, err, out, t)
+ }
+ if int64(len(body)) > t.limit {
+ out.body, out.truncated = body[:t.limit], true
+ } else {
+ out.body = body
+ }
+ out.state = runtimeStateOK
+ return out
+}
+
+// What the apiserver relays when the scheme is wrong, in either direction.
+// Certificate-verification failures are deliberately absent: those are a
+// verdict on the TLS setup, not a sign the port speaks plain HTTP.
+var cnpgSchemeMismatchHints = []string{
+ "http: server gave http response to https client",
+ "client sent an http request to an https server",
+ "first record does not look like a tls handshake",
+ "malformed http response",
+}
+
+// cnpgRelayedTransportTimeout: the apiserver's proxy reports its own dial or
+// read timeout to the Pod as a 503 whose message is "error trying to reach
+// service: ". Only the transport error's end is matched, so a
+// name or URL containing "timeout" never counts.
+func cnpgRelayedTransportTimeout(err error) bool {
+ if !apierrors.IsServiceUnavailable(err) {
+ return false
+ }
+ var status apierrors.APIStatus
+ if !errors.As(err, &status) {
+ return false
+ }
+ msg := status.Status().Message
+ return strings.HasPrefix(msg, "error trying to reach service:") &&
+ (strings.HasSuffix(msg, ": i/o timeout") || strings.HasSuffix(msg, ": context deadline exceeded"))
+}
+
+func classifyCNPGProxyFailure(ctx context.Context, err error, out cnpgProxyOutcome, t proxyTarget) cnpgProxyOutcome {
+ cnpgMarkTimedOut(ctx, err)
+ out = classifyCNPGProxyError(ctx, err, out)
+ if out.state == runtimeStateDenied || errors.Is(err, context.Canceled) {
+ return out
+ }
+ cause := err
+ if errors.Is(ctx.Err(), context.DeadlineExceeded) {
+ cause = context.DeadlineExceeded
+ }
+ // The Pod's own error answer can carry "connection refused" from
+ // PostgreSQL's socket, which is not a transport failure; read it first.
+ if plain, ok := cnpgRelayedPodSentence(err, t.port); ok && !out.schemeMismatch {
+ log.Printf("[cnpg] %s/%s port %d %s answered with an error: %v", t.namespace, t.pod, t.port, t.path, err)
+ out.err = plain
+ } else if plain, ok := cnpgTransportSentence(cause, t.port, cnpgRuntimeRequestTimeout); ok {
+ log.Printf("[cnpg] Failed to read %s/%s port %d %s: %v", t.namespace, t.pod, t.port, t.path, err)
+ out.err = plain
+ }
+ return out
+}
+
+func classifyCNPGProxyError(ctx context.Context, err error, out cnpgProxyOutcome) cnpgProxyOutcome {
+ if errors.Is(err, context.DeadlineExceeded) || errors.Is(ctx.Err(), context.DeadlineExceeded) {
+ out.state, out.err = runtimeStateUnreachable, fmt.Sprintf("no answer within %s", cnpgRuntimeRequestTimeout)
+ return out
+ }
+ if errors.Is(err, context.Canceled) {
+ out.state, out.err = runtimeStateError, "request cancelled"
+ return out
+ }
+ msg := err.Error()
+ lower := strings.ToLower(msg)
+ // The apiserver's own refusal names the pods/proxy subresource, which an
+ // answer relayed from the Pod never does.
+ if apierrors.IsForbidden(err) && strings.Contains(lower, "proxy") {
+ out.state, out.err = runtimeStateDenied, "the apiserver denied get pods/proxy"
+ return out
+ }
+ code := 0
+ var status apierrors.APIStatus
+ if errors.As(err, &status) {
+ code = int(status.Status().Code)
+ }
+ out.err = truncateCNPGRuntimeError(msg)
+ certificate := strings.Contains(lower, "x509") || strings.Contains(lower, "certificate")
+ if !certificate {
+ // A plain request against a TLS port comes back as a bare 400: nothing
+ // else makes a GET of these fixed paths a bad request.
+ out.schemeMismatch = code == http.StatusBadRequest
+ for _, hint := range cnpgSchemeMismatchHints {
+ if strings.Contains(lower, hint) {
+ out.schemeMismatch = true
+ }
+ }
+ }
+ switch {
+ case code >= 300 && code < 400:
+ out.state, out.err = runtimeStateError, fmt.Sprintf("the Pod answered with a redirect (%d), which is not followed", code)
+ out.schemeMismatch = false
+ case out.schemeMismatch, code == 0, code >= 500:
+ out.state = runtimeStateUnreachable
+ default:
+ out.state = runtimeStateError
+ }
+ return out
+}
+
+func truncateCNPGRuntimeError(s string) string {
+ const limit = 300
+ if len(s) > limit {
+ return s[:limit] + "…"
+ }
+ return s
+}
+
+func formatCNPGByteCap(n int64) string {
+ if n >= 1<<20 && n%(1<<20) == 0 {
+ return fmt.Sprintf("%d MiB", n>>20)
+ }
+ return fmt.Sprintf("%d bytes", n)
+}
diff --git a/internal/cnpg/reads.go b/internal/cnpg/reads.go
new file mode 100644
index 0000000000..a3a245d427
--- /dev/null
+++ b/internal/cnpg/reads.go
@@ -0,0 +1,14 @@
+package cnpg
+
+import (
+ "errors"
+)
+
+var ErrCNPGDisconnected = errors.New("not connected to cluster")
+
+type ReadFailure struct {
+ Status int
+ Message string
+}
+
+func (e *ReadFailure) Error() string { return e.Message }
diff --git a/internal/cnpg/recovery.go b/internal/cnpg/recovery.go
new file mode 100644
index 0000000000..627a38c3ae
--- /dev/null
+++ b/internal/cnpg/recovery.go
@@ -0,0 +1,543 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "log"
+ "net/http"
+ "sort"
+ "strings"
+ "time"
+ "unicode/utf8"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/dynamic"
+ "k8s.io/client-go/kubernetes"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+const (
+ cnpgRestoreValidationAnno = "radar.skyhook.io/restore-validation"
+ cnpgRestoreValidationMax = 2000
+ cnpgRecoveryEventLimit = 40
+
+ cnpgReadOK = "ok"
+ cnpgReadDenied = "denied"
+ cnpgReadNotFound = "notFound"
+ cnpgReadError = "error"
+ cnpgReadSkipped = "skipped"
+)
+
+var (
+ GrantListPods = auth.Grant{Verb: "list", Resource: "pods"}
+ cnpgGrantListEvents = auth.Grant{Verb: "list", Resource: "events"}
+ GrantGetCluster = auth.Grant{Verb: "get", Group: Group, Resource: "clusters"}
+)
+
+// cnpgGatedRead runs read when the caller holds g in namespace. The SAR comes
+// first so a denial names the grant; the read itself is made with the caller's
+// client, so the apiserver has the final say either way.
+func (s *Reader) gatedRead(ctx context.Context, g auth.Grant, namespace string, read func() error) integration.ReadSource {
+ if s.Access.Permission(ctx, g.In(namespace)) == integration.PermissionDenied {
+ return integration.ReadSource{State: cnpgReadDenied, Grant: g.In(namespace).Ref()}
+ }
+ return cnpgReadOutcome(read(), g, namespace)
+}
+
+func cnpgReadOutcome(err error, g auth.Grant, namespace string) integration.ReadSource {
+ switch {
+ case err == nil:
+ return integration.ReadSource{State: cnpgReadOK}
+ case apierrors.IsForbidden(err):
+ return integration.ReadSource{State: cnpgReadDenied, Grant: g.In(namespace).Ref()}
+ case apierrors.IsNotFound(err):
+ return integration.ReadSource{State: cnpgReadNotFound, Reason: err.Error()}
+ default:
+ return integration.ReadSource{State: cnpgReadError, Reason: cnpgPlainReadError(err.Error())}
+ }
+}
+
+// cnpgPlainReadError is the reader-facing text of a failed Kubernetes API
+// read: a sentence for a failure in transit (the raw error is logged),
+// otherwise the error itself.
+func cnpgPlainReadError(raw string) string {
+ if plain, ok := cnpgTransportSentence(errors.New(raw), 0, cnpgHAReadTimeout); ok {
+ log.Printf("[cnpg] Read failed: %s", raw)
+ return plain
+ }
+ return truncateCNPGRuntimeError(raw)
+}
+
+// CNPGRecoveryCluster is the restored Cluster's own progress. Counts are nil
+// when the operator has not reported them yet, which is not zero.
+type CNPGRecoveryCluster struct {
+ UID string `json:"uid"`
+ Phase string `json:"phase,omitempty"`
+ PhaseReason string `json:"phaseReason,omitempty"`
+ Instances *int64 `json:"instances"`
+ ReadyInstances *int64 `json:"readyInstances"`
+ CurrentPrimary string `json:"currentPrimary,omitempty"`
+ CreatedAt string `json:"createdAt,omitempty"`
+ Ready *CNPGRecoveryCondition `json:"ready,omitempty"`
+}
+
+type CNPGRecoveryCondition struct {
+ Status string `json:"status"`
+ Reason string `json:"reason,omitempty"`
+ Message string `json:"message,omitempty"`
+}
+
+// CNPGRecoverySpec is what the Cluster declares it recovers from.
+type CNPGRecoverySpec struct {
+ // SourceKind: objectStore (barman-cloud plugin), barmanObjectStore (in-tree),
+ // backup (a Backup object), volumeSnapshots, or unknown.
+ SourceKind string `json:"sourceKind"`
+ Source string `json:"source,omitempty"`
+ ObjectStore string `json:"objectStore,omitempty"`
+ ServerName string `json:"serverName,omitempty"`
+ Backup string `json:"backup,omitempty"`
+ Target map[string]any `json:"target,omitempty"`
+}
+
+type CNPGContainerState struct {
+ Name string `json:"name"`
+ State string `json:"state"`
+ Reason string `json:"reason,omitempty"`
+ Message string `json:"message,omitempty"`
+ ExitCode *int32 `json:"exitCode,omitempty"`
+ Restarts int32 `json:"restarts"`
+ Ready bool `json:"ready"`
+}
+
+// CNPGRecoveryPod is a Pod doing the recovery (kind "job": owned by one of
+// the Cluster's Jobs, e.g. -1-full-recovery) or an instance.
+type CNPGRecoveryPod struct {
+ Name string `json:"name"`
+ UID string `json:"uid"`
+ Kind string `json:"kind"`
+ Job string `json:"job,omitempty"`
+ OwnerVerified bool `json:"ownerVerified"`
+ Phase string `json:"phase"`
+ Ready bool `json:"ready"`
+ StartedAt string `json:"startedAt,omitempty"`
+ InitContainers []CNPGContainerState `json:"initContainers"`
+ Containers []CNPGContainerState `json:"containers"`
+}
+
+type CNPGRecoveryJob struct {
+ Name string `json:"name"`
+ Active int32 `json:"active"`
+ Succeeded int32 `json:"succeeded"`
+ Failed int32 `json:"failed"`
+ Complete bool `json:"complete"`
+ FailedWith string `json:"failedWith,omitempty"`
+ StartedAt string `json:"startedAt,omitempty"`
+ CompletedAt string `json:"completedAt,omitempty"`
+}
+
+type CNPGRecoveryEvent struct {
+ Type string `json:"type"`
+ Reason string `json:"reason"`
+ Message string `json:"message"`
+ Kind string `json:"kind"`
+ Name string `json:"name"`
+ Count int32 `json:"count"`
+ LastSeen string `json:"lastSeen,omitempty"`
+}
+
+// CNPGRestoreValidation is the note a person recorded after checking a
+// restored cluster. It is evidence that someone looked, never a verdict.
+type CNPGRestoreValidation struct {
+ Version int `json:"version"`
+ RecordedAt string `json:"recordedAt"`
+ RecordedBy string `json:"recordedBy,omitempty"`
+ Checked string `json:"checked"`
+ TargetTime string `json:"targetTime,omitempty"`
+ Source *CNPGRestoreValidationRef `json:"source,omitempty"`
+ Target CNPGRestoreValidationRef `json:"target"`
+}
+
+// CNPGRestoreValidationRef names a Cluster; UID is empty when Radar could not
+// read it as the recording user, and Verified says so.
+type CNPGRestoreValidationRef struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ UID string `json:"uid,omitempty"`
+ Verified bool `json:"verified"`
+}
+
+// CNPGRecoveryResponse is GET /api/cnpg/clusters/{ns}/{name}/recovery.
+type CNPGRecoveryResponse struct {
+ Cluster CNPGRecoveryCluster `json:"cluster"`
+ Recovery *CNPGRecoverySpec `json:"recovery"`
+ Pods []CNPGRecoveryPod `json:"pods"`
+ Jobs []CNPGRecoveryJob `json:"jobs"`
+ Events []CNPGRecoveryEvent `json:"events"`
+ Coverage map[string]integration.ReadSource `json:"coverage"`
+ Validation *CNPGRestoreValidation `json:"validation,omitempty"`
+ ValidationError string `json:"validationError,omitempty"`
+ CapturedAt string `json:"capturedAt"`
+}
+
+func (s *Reader) RecoverySnapshot(callerCtx context.Context, typed kubernetes.Interface, cluster *unstructured.Unstructured) CNPGRecoveryResponse {
+ ctx := callerCtx
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ resp := CNPGRecoveryResponse{
+ Cluster: cnpgRecoveryClusterOf(cluster),
+ Recovery: cnpgRecoverySpecOf(cluster),
+ Pods: []CNPGRecoveryPod{},
+ Jobs: []CNPGRecoveryJob{},
+ Events: []CNPGRecoveryEvent{},
+ Coverage: map[string]integration.ReadSource{},
+ CapturedAt: time.Now().UTC().Format(time.RFC3339),
+ }
+ resp.Validation, resp.ValidationError = parseCNPGRestoreValidation(cluster.GetAnnotations()[cnpgRestoreValidationAnno])
+
+ selector := labels.SelectorFromSet(labels.Set{clusterLabel: name}).String()
+ var jobs []batchv1.Job
+ resp.Coverage["jobs"] = s.gatedRead(callerCtx, cnpgGrantListJobs, namespace, func() error {
+ list, err := typed.BatchV1().Jobs(namespace).List(ctx, metav1.ListOptions{LabelSelector: selector})
+ if err == nil {
+ jobs = list.Items
+ }
+ return err
+ })
+ ownedJobs := map[string]types.UID{}
+ for i := range jobs {
+ j := &jobs[i]
+ if !controlledBy(j.OwnerReferences, Group, "Cluster", name, cluster.GetUID()) {
+ continue
+ }
+ ownedJobs[j.Name] = j.UID
+ resp.Jobs = append(resp.Jobs, cnpgRecoveryJobOf(j))
+ }
+ sort.Slice(resp.Jobs, func(i, j int) bool { return resp.Jobs[i].Name < resp.Jobs[j].Name })
+
+ var pods []corev1.Pod
+ resp.Coverage["pods"] = s.gatedRead(callerCtx, GrantListPods, namespace, func() error {
+ list, err := typed.CoreV1().Pods(namespace).List(ctx, metav1.ListOptions{LabelSelector: selector})
+ if err == nil {
+ pods = list.Items
+ }
+ return err
+ })
+ jobsKnown := resp.Coverage["jobs"].State == cnpgReadOK
+ for i := range pods {
+ p := &pods[i]
+ ref := controllerRef(p.OwnerReferences)
+ if ref == nil {
+ continue
+ }
+ switch {
+ case controlledBy(p.OwnerReferences, Group, "Cluster", name, cluster.GetUID()):
+ resp.Pods = append(resp.Pods, cnpgRecoveryPodOf(p, "instance", "", true))
+ case ref.Kind == "Job" && ownedJobs[ref.Name] != "" && controlledBy(p.OwnerReferences, "batch", "Job", ref.Name, ownedJobs[ref.Name]):
+ resp.Pods = append(resp.Pods, cnpgRecoveryPodOf(p, "job", ref.Name, true))
+ case ref.Kind == "Job" && !jobsKnown && strings.HasPrefix(ref.Name, name+"-"):
+ resp.Pods = append(resp.Pods, cnpgRecoveryPodOf(p, "job", ref.Name, false))
+ }
+ }
+ sort.Slice(resp.Pods, func(i, j int) bool {
+ if resp.Pods[i].Kind != resp.Pods[j].Kind {
+ return resp.Pods[i].Kind == "job"
+ }
+ return resp.Pods[i].Name < resp.Pods[j].Name
+ })
+
+ subjects := map[types.UID]bool{cluster.GetUID(): true}
+ for _, p := range resp.Pods {
+ subjects[types.UID(p.UID)] = true
+ }
+ for _, uid := range ownedJobs {
+ subjects[uid] = true
+ }
+ var events []corev1.Event
+ resp.Coverage["events"] = s.gatedRead(callerCtx, cnpgGrantListEvents, namespace, func() error {
+ list, err := typed.CoreV1().Events(namespace).List(ctx, metav1.ListOptions{})
+ if err == nil {
+ events = list.Items
+ }
+ return err
+ })
+ resp.Events = cnpgRecoveryEventsOf(events, subjects, cnpgRecoveryEventLimit)
+ return resp
+}
+
+func cnpgRecoveryClusterOf(cluster *unstructured.Unstructured) CNPGRecoveryCluster {
+ out := CNPGRecoveryCluster{UID: string(cluster.GetUID())}
+ if ts := cluster.GetCreationTimestamp(); !ts.IsZero() {
+ out.CreatedAt = ts.UTC().Format(time.RFC3339)
+ }
+ out.Phase, _, _ = unstructured.NestedString(cluster.Object, "status", "phase")
+ out.PhaseReason, _, _ = unstructured.NestedString(cluster.Object, "status", "phaseReason")
+ out.CurrentPrimary, _, _ = unstructured.NestedString(cluster.Object, "status", "currentPrimary")
+ if v, ok, _ := unstructured.NestedInt64(cluster.Object, "status", "instances"); ok {
+ out.Instances = &v
+ }
+ if v, ok, _ := unstructured.NestedInt64(cluster.Object, "status", "readyInstances"); ok {
+ out.ReadyInstances = &v
+ }
+ conds, _, _ := unstructured.NestedSlice(cluster.Object, "status", "conditions")
+ for _, c := range conds {
+ m, _ := c.(map[string]any)
+ if m == nil || m["type"] != "Ready" {
+ continue
+ }
+ status, _ := m["status"].(string)
+ reason, _ := m["reason"].(string)
+ message, _ := m["message"].(string)
+ out.Ready = &CNPGRecoveryCondition{Status: status, Reason: reason, Message: message}
+ }
+ return out
+}
+
+func cnpgRecoverySpecOf(cluster *unstructured.Unstructured) *CNPGRecoverySpec {
+ rec, ok, _ := unstructured.NestedMap(cluster.Object, "spec", "bootstrap", "recovery")
+ if !ok {
+ return nil
+ }
+ out := &CNPGRecoverySpec{SourceKind: "unknown"}
+ if t, ok := rec["recoveryTarget"].(map[string]any); ok && len(t) > 0 {
+ out.Target = t
+ }
+ if b, ok := rec["backup"].(map[string]any); ok {
+ out.SourceKind = "backup"
+ out.Backup, _ = b["name"].(string)
+ return out
+ }
+ if _, ok := rec["volumeSnapshots"]; ok {
+ out.SourceKind = "volumeSnapshots"
+ return out
+ }
+ out.Source, _ = rec["source"].(string)
+ if out.Source == "" {
+ return out
+ }
+ ext, _, _ := unstructured.NestedSlice(cluster.Object, "spec", "externalClusters")
+ for _, e := range ext {
+ m, _ := e.(map[string]any)
+ if m == nil || m["name"] != out.Source {
+ continue
+ }
+ if p, ok := m["plugin"].(map[string]any); ok {
+ params, _ := p["parameters"].(map[string]any)
+ out.SourceKind = "objectStore"
+ out.ObjectStore, _ = params["barmanObjectName"].(string)
+ out.ServerName, _ = params["serverName"].(string)
+ if out.ServerName == "" {
+ out.ServerName = out.Source
+ }
+ } else if b, ok := m["barmanObjectStore"].(map[string]any); ok {
+ out.SourceKind = "barmanObjectStore"
+ out.ServerName, _ = b["serverName"].(string)
+ if out.ServerName == "" {
+ out.ServerName = out.Source
+ }
+ }
+ }
+ return out
+}
+
+func cnpgRecoveryJobOf(j *batchv1.Job) CNPGRecoveryJob {
+ out := CNPGRecoveryJob{Name: j.Name, Active: j.Status.Active, Succeeded: j.Status.Succeeded, Failed: j.Status.Failed}
+ if j.Status.StartTime != nil {
+ out.StartedAt = j.Status.StartTime.UTC().Format(time.RFC3339)
+ }
+ if j.Status.CompletionTime != nil {
+ out.CompletedAt = j.Status.CompletionTime.UTC().Format(time.RFC3339)
+ }
+ for _, c := range j.Status.Conditions {
+ if c.Status != corev1.ConditionTrue {
+ continue
+ }
+ switch c.Type {
+ case batchv1.JobComplete:
+ out.Complete = true
+ case batchv1.JobFailed:
+ out.FailedWith = strings.TrimSpace(c.Reason + ": " + c.Message)
+ }
+ }
+ return out
+}
+
+func cnpgContainerStatesOf(specs []corev1.Container, statuses []corev1.ContainerStatus) []CNPGContainerState {
+ byName := map[string]corev1.ContainerStatus{}
+ for _, st := range statuses {
+ byName[st.Name] = st
+ }
+ out := make([]CNPGContainerState, 0, len(specs))
+ for _, c := range specs {
+ cs := CNPGContainerState{Name: c.Name, State: "unknown"}
+ if st, ok := byName[c.Name]; ok {
+ cs.Restarts = st.RestartCount
+ cs.Ready = st.Ready
+ switch {
+ case st.State.Running != nil:
+ cs.State = "running"
+ case st.State.Terminated != nil:
+ t := st.State.Terminated
+ cs.State, cs.Reason, cs.Message = "terminated", t.Reason, truncateCNPGRuntimeError(t.Message)
+ code := t.ExitCode
+ cs.ExitCode = &code
+ case st.State.Waiting != nil:
+ cs.State, cs.Reason, cs.Message = "waiting", st.State.Waiting.Reason, truncateCNPGRuntimeError(st.State.Waiting.Message)
+ }
+ }
+ out = append(out, cs)
+ }
+ return out
+}
+
+func cnpgRecoveryPodOf(p *corev1.Pod, kind, job string, verified bool) CNPGRecoveryPod {
+ out := CNPGRecoveryPod{
+ Name: p.Name,
+ UID: string(p.UID),
+ Kind: kind,
+ Job: job,
+ OwnerVerified: verified,
+ Phase: string(p.Status.Phase),
+ Ready: cnpgActionPodReady(p),
+ InitContainers: cnpgContainerStatesOf(p.Spec.InitContainers, p.Status.InitContainerStatuses),
+ Containers: cnpgContainerStatesOf(p.Spec.Containers, p.Status.ContainerStatuses),
+ }
+ if p.Status.StartTime != nil {
+ out.StartedAt = p.Status.StartTime.UTC().Format(time.RFC3339)
+ }
+ return out
+}
+
+func cnpgEventTime(e *corev1.Event) time.Time {
+ switch {
+ case !e.LastTimestamp.IsZero():
+ return e.LastTimestamp.Time
+ case !e.EventTime.IsZero():
+ return e.EventTime.Time
+ default:
+ return e.CreationTimestamp.Time
+ }
+}
+
+func cnpgRecoveryEventsOf(events []corev1.Event, subjects map[types.UID]bool, limit int) []CNPGRecoveryEvent {
+ kept := make([]*corev1.Event, 0)
+ for i := range events {
+ e := &events[i]
+ if e.InvolvedObject.UID != "" && subjects[e.InvolvedObject.UID] {
+ kept = append(kept, e)
+ }
+ }
+ sort.SliceStable(kept, func(i, j int) bool { return cnpgEventTime(kept[i]).After(cnpgEventTime(kept[j])) })
+ if len(kept) > limit {
+ kept = kept[:limit]
+ }
+ out := make([]CNPGRecoveryEvent, 0, len(kept))
+ for _, e := range kept {
+ ev := CNPGRecoveryEvent{Type: e.Type, Reason: e.Reason, Message: e.Message, Kind: e.InvolvedObject.Kind, Name: e.InvolvedObject.Name, Count: e.Count}
+ if t := cnpgEventTime(e); !t.IsZero() {
+ ev.LastSeen = t.UTC().Format(time.RFC3339)
+ }
+ out = append(out, ev)
+ }
+ return out
+}
+
+func parseCNPGRestoreValidation(raw string) (*CNPGRestoreValidation, string) {
+ if strings.TrimSpace(raw) == "" {
+ return nil, ""
+ }
+ var v CNPGRestoreValidation
+ if err := json.Unmarshal([]byte(raw), &v); err != nil {
+ return nil, "the " + cnpgRestoreValidationAnno + " annotation is not valid JSON"
+ }
+ if v.RecordedAt == "" || v.Checked == "" {
+ return nil, "the " + cnpgRestoreValidationAnno + " annotation is missing recordedAt or checked"
+ }
+ return &v, ""
+}
+
+// cnpgRestoreValidationParams is the POST body's params. Identity and UIDs
+// are never taken from the client: the server records who asked and the UIDs
+// it read.
+type cnpgRestoreValidationParams struct {
+ Checked string `json:"checked"`
+ TargetTime string `json:"targetTime,omitempty"`
+ Source *struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ } `json:"source,omitempty"`
+}
+
+func RecordCNPGRestoreValidation(ctx context.Context, dyn dynamic.Interface, namespace, name string, req integration.ActionRequest, recordedBy string, now time.Time) (*CNPGRestoreValidation, error) {
+ var params cnpgRestoreValidationParams
+ if err := integration.DecodeActionParams(req.Params, ¶ms); err != nil {
+ return nil, err
+ }
+ checked := strings.TrimSpace(params.Checked)
+ if checked == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.checked is required: say what you checked")
+ }
+ if utf8.RuneCountInString(checked) > cnpgRestoreValidationMax {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.checked is longer than %d characters", cnpgRestoreValidationMax)
+ }
+ if params.TargetTime != "" {
+ if _, err := time.Parse(time.RFC3339, params.TargetTime); err != nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.targetTime must be RFC 3339")
+ }
+ }
+ cluster, err := dyn.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, err
+ }
+ if string(cluster.GetUID()) != req.UID {
+ return nil, integration.ChangedAction(nil, "Cluster %s/%s was deleted and recreated since you reviewed it", namespace, name)
+ }
+ if cnpgRecoverySpecOf(cluster) == nil {
+ return nil, integration.BlockedAction("This Cluster was not bootstrapped from a backup (spec.bootstrap.recovery is not set)")
+ }
+ note := &CNPGRestoreValidation{
+ Version: 1,
+ RecordedAt: now.UTC().Format(time.RFC3339),
+ RecordedBy: recordedBy,
+ Checked: checked,
+ TargetTime: params.TargetTime,
+ Target: CNPGRestoreValidationRef{Namespace: namespace, Name: name, UID: string(cluster.GetUID()), Verified: true},
+ }
+ if params.Source != nil && params.Source.Name != "" {
+ srcNS := params.Source.Namespace
+ if srcNS == "" {
+ srcNS = namespace
+ }
+ ref := CNPGRestoreValidationRef{Namespace: srcNS, Name: params.Source.Name}
+ if src, err := dyn.Resource(ClusterGVR).Namespace(srcNS).Get(ctx, params.Source.Name, metav1.GetOptions{}); err == nil {
+ ref.UID, ref.Verified = string(src.GetUID()), true
+ }
+ note.Source = &ref
+ }
+ data, err := json.Marshal(note)
+ if err != nil {
+ return nil, err
+ }
+ if err := integration.MergePatchAtVersion(ctx, dyn, ClusterGVR, cluster, map[string]any{
+ "metadata": map[string]any{"annotations": map[string]any{cnpgRestoreValidationAnno: string(data)}},
+ }); err != nil {
+ if apierrors.IsConflict(err) {
+ return nil, integration.ChangedAction(nil, "Cluster %s/%s changed while the note was being recorded; try again", namespace, name)
+ }
+ return nil, err
+ }
+ return note, nil
+}
+
+func (s *Reader) RestoreCapability(ctx context.Context, namespace string) integration.ActionCapability {
+ g := cnpgGrantCreateClusters.In(namespace)
+ c := integration.CapabilityVerdict("", []string{s.Access.Permission(ctx, g)}, []auth.Grant{g})
+ return cnpgOperatorWebhookGuard(s.operatorVerdictFor(ctx, namespace), c)
+}
diff --git a/internal/cnpg/recovery_test.go b/internal/cnpg/recovery_test.go
new file mode 100644
index 0000000000..f81b0ae24a
--- /dev/null
+++ b/internal/cnpg/recovery_test.go
@@ -0,0 +1,629 @@
+package cnpg
+
+import (
+ "archive/zip"
+ "bytes"
+ "context"
+ "encoding/json"
+ "errors"
+ "io"
+ "strings"
+ "testing"
+ "time"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+
+ batchv1 "k8s.io/api/batch/v1"
+
+ corev1 "k8s.io/api/core/v1"
+
+ discoveryv1 "k8s.io/api/discovery/v1"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+
+ k8stesting "k8s.io/client-go/testing"
+)
+
+func cnpgRestoredCluster() *unstructured.Unstructured {
+ return cnpgActionCluster(func(obj map[string]any) {
+ spec := obj["spec"].(map[string]any)
+ delete(spec, "plugins")
+ delete(spec, "backup")
+ spec["instances"] = int64(1)
+ spec["bootstrap"] = map[string]any{"recovery": map[string]any{
+ "source": "origin",
+ "recoveryTarget": map[string]any{"targetTime": "2026-09-29T08:00:00Z"},
+ }}
+ spec["externalClusters"] = []any{map[string]any{"name": "origin", "plugin": map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "store", "serverName": "pg-src"},
+ }}}
+ status := obj["status"].(map[string]any)
+ status["phase"] = "Setting up primary"
+ status["instances"] = int64(1)
+ status["readyInstances"] = int64(0)
+ })
+}
+
+func testControllerRefs(kind, name, uid string) []metav1.OwnerReference {
+ yes := true
+ apiVersion := "postgresql.cnpg.io/v1"
+ if kind == "Job" {
+ apiVersion = "batch/v1"
+ }
+ return []metav1.OwnerReference{{APIVersion: apiVersion, Kind: kind, Name: name, UID: types.UID(uid), Controller: &yes}}
+}
+
+func TestCNPGRecoverySpecOf(t *testing.T) {
+ r := cnpgRecoverySpecOf(cnpgRestoredCluster())
+ if r == nil || r.SourceKind != "objectStore" || r.ObjectStore != "store" || r.ServerName != "pg-src" || r.Target["targetTime"] != "2026-09-29T08:00:00Z" {
+ t.Fatalf("unexpected recovery spec: %+v", r)
+ }
+ backup := cnpgActionCluster(func(obj map[string]any) {
+ obj["spec"].(map[string]any)["bootstrap"] = map[string]any{"recovery": map[string]any{"backup": map[string]any{"name": "b1"}}}
+ })
+ if r := cnpgRecoverySpecOf(backup); r == nil || r.SourceKind != "backup" || r.Backup != "b1" {
+ t.Fatalf("backup recovery: %+v", r)
+ }
+ if cnpgRecoverySpecOf(cnpgActionCluster(nil)) != nil {
+ t.Fatal("a Cluster without bootstrap.recovery is not a restore")
+ }
+}
+
+func TestCNPGRecoverySnapshotClassifiesPodsByOwner(t *testing.T) {
+ cluster := cnpgRestoredCluster()
+ job := &batchv1.Job{
+ ObjectMeta: metav1.ObjectMeta{Name: "pg-1-full-recovery", Namespace: "db", UID: "job-uid", Labels: map[string]string{"cnpg.io/cluster": "pg"}, OwnerReferences: testControllerRefs("Cluster", "pg", cnpgActionTestUID)},
+ Status: batchv1.JobStatus{Active: 1},
+ }
+ foreignJob := &batchv1.Job{
+ ObjectMeta: metav1.ObjectMeta{Name: "pg-other", Namespace: "db", Labels: map[string]string{"cnpg.io/cluster": "pg"}, OwnerReferences: testControllerRefs("Cluster", "pg", "old-uid")},
+ }
+ recoveryPod := &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{Name: "pg-1-full-recovery-abcde", Namespace: "db", UID: "p1", Labels: map[string]string{"cnpg.io/cluster": "pg"}, OwnerReferences: testControllerRefs("Job", "pg-1-full-recovery", "job-uid")},
+ Spec: corev1.PodSpec{InitContainers: []corev1.Container{{Name: "bootstrap-controller"}}, Containers: []corev1.Container{{Name: "full-recovery"}}},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodRunning,
+ InitContainerStatuses: []corev1.ContainerStatus{{Name: "bootstrap-controller", State: corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{ExitCode: 0, Reason: "Completed"}}}},
+ ContainerStatuses: []corev1.ContainerStatus{{Name: "full-recovery", State: corev1.ContainerState{Running: &corev1.ContainerStateRunning{}}}},
+ },
+ }
+ strayPod := &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "pg-other-xyz", Namespace: "db", Labels: map[string]string{"cnpg.io/cluster": "pg"}, OwnerReferences: testControllerRefs("Job", "pg-other", "j2")}}
+ events := []runtime.Object{
+ &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: "e1", Namespace: "db"}, InvolvedObject: corev1.ObjectReference{Kind: "Pod", Name: "pg-1-full-recovery-abcde", UID: "p1"}, Type: "Warning", Reason: "BackOff", Message: "restarting", LastTimestamp: metav1.NewTime(time.Now())},
+ &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: "e2", Namespace: "db"}, InvolvedObject: corev1.ObjectReference{Kind: "Pod", Name: "unrelated"}, Type: "Warning", Reason: "Other"},
+ }
+ oldPod := recoveryPod.DeepCopy()
+ oldPod.Name, oldPod.UID = "pg-old-recovery", "old-pod"
+ oldPod.OwnerReferences = testControllerRefs("Job", job.Name, "old-job")
+ for _, subject := range []corev1.ObjectReference{{Kind: "Pod", Name: recoveryPod.Name, UID: "old-p1"}, {Kind: "Job", Name: job.Name, UID: "old-job"}, {Kind: "Cluster", Name: cluster.GetName(), UID: "old-cluster"}} {
+ events = append(events, &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: "old-" + subject.Kind, Namespace: "db"}, InvolvedObject: subject, Reason: "OldFailure"})
+ }
+ typed := k8sfake.NewSimpleClientset(append([]runtime.Object{job, foreignJob, recoveryPod, strayPod, oldPod}, events...)...)
+ s := newTestReader(nil)
+ ctx := context.Background()
+
+ snap := s.RecoverySnapshot(ctx, typed, cluster)
+ if len(snap.Jobs) != 1 || snap.Jobs[0].Name != "pg-1-full-recovery" {
+ t.Fatalf("only the Cluster-owned Job counts: %+v", snap.Jobs)
+ }
+ if len(snap.Pods) != 1 || snap.Pods[0].Kind != "job" || !snap.Pods[0].OwnerVerified {
+ t.Fatalf("only the owned Job's Pod counts: %+v", snap.Pods)
+ }
+ if snap.Pods[0].InitContainers[0].State != "terminated" || snap.Pods[0].Containers[0].State != "running" {
+ t.Fatalf("container states: %+v", snap.Pods[0])
+ }
+ if len(snap.Events) != 1 || snap.Events[0].Reason != "BackOff" {
+ t.Fatalf("events should be limited to the recovery objects: %+v", snap.Events)
+ }
+ if snap.Coverage["pods"].State != cnpgReadOK || snap.Recovery == nil || snap.Cluster.ReadyInstances == nil || *snap.Cluster.ReadyInstances != 0 {
+ t.Fatalf("snapshot: %+v", snap)
+ }
+}
+
+func TestCNPGRecoverySnapshotReportsDeniedReads(t *testing.T) {
+ typed := k8sfake.NewSimpleClientset()
+ typed.PrependReactor("list", "pods", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apiForbidden("pods")
+ })
+ snap := newTestReader(nil).RecoverySnapshot(context.Background(), typed, cnpgRestoredCluster())
+ if g := snap.Coverage["pods"].Grant; snap.Coverage["pods"].State != cnpgReadDenied || g == nil || g.Verb != "list" || g.Resource != "pods" {
+ t.Fatalf("a denied Pod list must be reported with its grant: %+v", snap.Coverage["pods"])
+ }
+ if len(snap.Pods) != 0 {
+ t.Fatal("no Pods when they cannot be read")
+ }
+}
+
+func TestRecordCNPGRestoreValidation(t *testing.T) {
+ restored := cnpgRestoredCluster()
+ source := cnpgActionCluster(func(obj map[string]any) {
+ m := obj["metadata"].(map[string]any)
+ m["name"], m["uid"] = "pg-src", "src-uid"
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{restored, source})
+ now := time.Date(2026, 9, 30, 9, 0, 0, 0, time.UTC)
+ req := cnpgActionReq(t, nil, map[string]any{"checked": " row counts match ", "targetTime": "2026-09-29T08:00:00Z", "source": map[string]any{"name": "pg-src"}})
+
+ note, err := RecordCNPGRestoreValidation(context.Background(), env.dyn, "db", "pg", req, "alice", now)
+ if err != nil {
+ t.Fatal(err)
+ }
+ if note.Checked != "row counts match" || note.RecordedBy != "alice" || note.RecordedAt != "2026-09-30T09:00:00Z" {
+ t.Fatalf("note: %+v", note)
+ }
+ if note.Source == nil || note.Source.UID != "src-uid" || !note.Source.Verified || note.Target.UID != cnpgActionTestUID {
+ t.Fatalf("UIDs must come from the server's reads: %+v", note)
+ }
+ if len(env.patches) != 1 {
+ t.Fatalf("want one patch, got %d", len(env.patches))
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ if body["metadata"].(map[string]any)["resourceVersion"] != "42" {
+ t.Fatal("the patch must be bound to the resourceVersion read")
+ }
+ raw := cnpgActionAnnotations(t, body)[cnpgRestoreValidationAnno].(string)
+ var stored CNPGRestoreValidation
+ if err := json.Unmarshal([]byte(raw), &stored); err != nil || stored.Checked != "row counts match" {
+ t.Fatalf("stored annotation: %s (%v)", raw, err)
+ }
+}
+
+func TestRecordCNPGRestoreValidationRefusals(t *testing.T) {
+ notRestored := cnpgActionCluster(nil)
+ env := newCNPGActionEnv(t, []runtime.Object{notRestored})
+ now := time.Now()
+ if _, err := RecordCNPGRestoreValidation(context.Background(), env.dyn, "db", "pg", cnpgActionReq(t, nil, map[string]any{"checked": "x"}), "", now); err == nil {
+ t.Fatal("a Cluster that was not restored cannot carry a validation note")
+ } else if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Fatalf("want blocked, got %v", err)
+ }
+ env = newCNPGActionEnv(t, []runtime.Object{cnpgRestoredCluster()})
+ for name, params := range map[string]any{
+ "empty": map[string]any{"checked": " "},
+ "bad target": map[string]any{"checked": "ok", "targetTime": "yesterday"},
+ "unknown field": map[string]any{"checked": "ok", "recordedBy": "mallory"},
+ } {
+ if _, err := RecordCNPGRestoreValidation(context.Background(), env.dyn, "db", "pg", cnpgActionReq(t, nil, params), "", now); err == nil {
+ t.Fatalf("%s: want an error", name)
+ }
+ }
+ stale := cnpgActionReq(t, nil, map[string]any{"checked": "ok"})
+ stale.UID = "other"
+ if _, err := RecordCNPGRestoreValidation(context.Background(), env.dyn, "db", "pg", stale, "", now); err == nil {
+ t.Fatal("a recreated Cluster must be refused")
+ }
+ if len(env.patches) != 0 {
+ t.Fatal("refusals must not write")
+ }
+}
+
+func TestParseCNPGRestoreValidation(t *testing.T) {
+ if v, e := parseCNPGRestoreValidation(""); v != nil || e != "" {
+ t.Fatal("absent is not an error")
+ }
+ if _, e := parseCNPGRestoreValidation("{"); e == "" {
+ t.Fatal("malformed must be reported")
+ }
+ if _, e := parseCNPGRestoreValidation(`{"recordedAt":"2026-09-30T00:00:00Z"}`); e == "" {
+ t.Fatal("a note without what was checked is not a note")
+ }
+}
+
+func TestCNPGReportLogLineRedactsQueryText(t *testing.T) {
+ line := `2026-09-30T00:00:00.000000000Z {"level":"info","logger":"postgres","msg":"record","record":{"message":"duration: 1.2 ms statement: SELECT * FROM users WHERE email='a@b.c'","query":"SELECT 1","detail":"parameters: $1 = 'x'","error_severity":"LOG"}}`
+ out := cnpgReportLogLine(line, false)
+ if strings.Contains(out, "users") || strings.Contains(out, "SELECT 1") || strings.Contains(out, "$1") {
+ t.Fatalf("query text leaked: %s", out)
+ }
+ if !strings.HasPrefix(out, "2026-09-30T00:00:00.000000000Z ") || !strings.Contains(out, "duration: 1.2 ms statement: ") {
+ t.Fatalf("timestamp and the non-query part must survive: %s", out)
+ }
+ if kept := cnpgReportLogLine(line, true); !strings.Contains(kept, "SELECT * FROM users") {
+ t.Fatalf("opt-in keeps query text: %s", kept)
+ }
+ if plain := cnpgReportLogLine("connecting to postgresql://app:s3cretpw@db:5432/app", false); strings.Contains(plain, "s3cretpw") {
+ t.Fatalf("secret patterns are redacted on every line: %s", plain)
+ }
+}
+
+// Every PostgreSQL message format that carries SQL, as postgres.c and
+// auto_explain write them; none may keep the literal without the opt-in.
+func TestCNPGReportLogLineRedactsEverySQLFormat(t *testing.T) {
+ const secret = "private-customer-data"
+ record := func(message, detail string) string {
+ b, _ := json.Marshal(map[string]any{"level": "info", "logger": "postgres", "msg": "record",
+ "record": map[string]any{"message": message, "detail": detail, "error_severity": "LOG"}})
+ return "2026-09-30T00:00:00.000000000Z " + string(b)
+ }
+ for name, line := range map[string]string{
+ "statement": record("statement: SELECT '"+secret+"'", ""),
+ "duration statement": record("duration: 0.4 ms statement: SELECT '"+secret+"'", ""),
+ "execute": record("execute S_1: SELECT '"+secret+"'", ""),
+ "execute unnamed": record("execute : SELECT '"+secret+"'", ""),
+ "execute portal": record("execute S_1/C_2: SELECT '"+secret+"'", ""),
+ "execute fetch": record("execute fetch from S_1/C_2: SELECT '"+secret+"'", ""),
+ "duration execute": record("duration: 0.4 ms execute S_1: SELECT '"+secret+"'", ""),
+ "parse": record("parse S_1: SELECT '"+secret+"'", ""),
+ "duration parse": record("duration: 0.1 ms parse : SELECT '"+secret+"'", ""),
+ "bind": record("bind S_1: SELECT '"+secret+"'", ""),
+ "duration bind": record("duration: 0.1 ms bind /C_1: SELECT '"+secret+"'", ""),
+ "auto_explain": record("duration: 12.0 ms plan:\nQuery Text: SELECT '"+secret+"'", ""),
+ "bind parameters": record("execute S_1: SELECT $1", "parameters: $1 = '"+secret+"'"),
+ "Parameters": record("duration: 1 ms", "Parameters: $1 = '"+secret+"'"),
+ "plain statement": "2026-09-30 00:00:00 UTC [42] LOG: statement: SELECT '" + secret + "'",
+ "plain execute": "2026-09-30 00:00:00 UTC [42] LOG: duration: 0.4 ms execute S_1: SELECT '" + secret + "'",
+ "plain parameters": "2026-09-30 00:00:00 UTC [42] DETAIL: parameters: $1 = '" + secret + "'",
+ "parse name spaces": record("duration: 1 ms parse customer lookup: SELECT '"+secret+"'", ""),
+ "bind name spaces": record("duration: 1 ms bind customer lookup: SELECT '"+secret+"'", ""),
+ "execute name space": record("execute customer lookup: SELECT '"+secret+"'", ""),
+ "name with colon": record("execute a: b/c: d: SELECT '"+secret+"'", ""),
+ "upper case": record("EXECUTE Customer Lookup: SELECT '"+secret+"'", ""),
+ "plain name spaces": "2026-09-30 00:00:00 UTC [42] LOG: duration: 0.4 ms execute customer lookup: SELECT '" + secret + "'",
+ } {
+ if out := cnpgReportLogLine(line, false); strings.Contains(out, secret) {
+ t.Errorf("%s: SQL kept with queryText off: %s", name, out)
+ }
+ if out := cnpgReportLogLine(line, true); !strings.Contains(out, secret) {
+ t.Errorf("%s: opt-in lost the query text: %s", name, out)
+ }
+ }
+ keep := record("checkpoint complete: wrote 3 buffers", "")
+ if out := cnpgReportLogLine(keep, false); !strings.Contains(out, "checkpoint complete: wrote 3 buffers") {
+ t.Errorf("a message without SQL was redacted: %s", out)
+ }
+}
+
+func TestCNPGReportSecretNames(t *testing.T) {
+ obj := map[string]any{"spec": map[string]any{
+ "superuserSecret": map[string]any{"name": "su"},
+ "certificates": map[string]any{"serverTLSSecret": "tls", "serverCASecret": "ca"},
+ "bootstrap": map[string]any{"initdb": map[string]any{"secret": map[string]any{"name": "app"}}},
+ "configuration": map[string]any{"s3Credentials": map[string]any{"accessKeyId": map[string]any{"name": "creds", "key": "ID"}}},
+ "monitoring": map[string]any{"customQueriesConfigMap": []any{map[string]any{"name": "queries", "key": "q.yaml"}}},
+ "imageCatalogRef": map[string]any{"name": "catalog", "kind": "ImageCatalog"},
+ }}
+ got := map[string]bool{}
+ cnpgReportSecretNamesIn(obj, "Cluster/pg", func(name, _ string) { got[name] = true })
+ for _, want := range []string{"su", "tls", "ca", "app", "creds"} {
+ if !got[want] {
+ t.Errorf("missing Secret %q in %v", want, got)
+ }
+ }
+ for _, not := range []string{"queries", "catalog", "[REDACTED]"} {
+ if got[not] {
+ t.Errorf("%q is not a Secret", not)
+ }
+ }
+}
+
+func TestCNPGReportBundle(t *testing.T) {
+ cluster := cnpgActionCluster(nil)
+ pod := cnpgActionPod("pg-1", "pod-1", true)
+ pod.Spec = corev1.PodSpec{
+ Containers: []corev1.Container{{Name: "postgres", Command: []string{"postgres"}, Args: []string{"postgresql://app:pod-password@db/pg"}, Env: []corev1.EnvVar{{Name: "DB_PASSWORD", Value: "hunter2"}, {Name: "PGDATA", Value: "/var/lib"}}}},
+ Volumes: []corev1.Volume{{Name: "su", VolumeSource: corev1.VolumeSource{Secret: &corev1.SecretVolumeSource{SecretName: "pg-superuser"}}}},
+ }
+ podSpecCopy, err := json.Marshal(pod.Spec)
+ if err != nil {
+ t.Fatal(err)
+ }
+ pod.Annotations = map[string]string{"cnpg.io/podSpec": string(podSpecCopy), "kubectl.kubernetes.io/last-applied-configuration": string(podSpecCopy), "cnpg.io/instanceRole": "primary"}
+ currentBackup := &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": Group + "/v1", "kind": "Backup",
+ "metadata": map[string]any{"name": "current-backup", "namespace": cluster.GetNamespace(), "uid": "current-backup-uid"},
+ "spec": map[string]any{"cluster": map[string]any{"name": cluster.GetName()}},
+ "status": map[string]any{"pluginMetadata": map[string]any{"clusterUID": string(cluster.GetUID())}},
+ }}
+ oldBackup := currentBackup.DeepCopy()
+ oldBackup.SetName("predecessor-backup")
+ oldBackup.SetUID("predecessor-backup-uid")
+ if err := unstructured.SetNestedField(oldBackup.Object, "old-cluster-uid", "status", "pluginMetadata", "clusterUID"); err != nil {
+ t.Fatal(err)
+ }
+ if err := unstructured.SetNestedMap(oldBackup.Object, map[string]any{"name": "predecessor-only-secret", "key": "password"}, "spec", "credentials"); err != nil {
+ t.Fatal(err)
+ }
+ ownedPVC := &corev1.PersistentVolumeClaim{ObjectMeta: metav1.ObjectMeta{Name: "current-pvc", Namespace: cluster.GetNamespace(), UID: "current-pvc-uid", Labels: pod.Labels, OwnerReferences: pod.OwnerReferences}}
+ foreignPVC := ownedPVC.DeepCopy()
+ foreignPVC.Name, foreignPVC.UID = "unowned-pvc", "unowned-pvc-uid"
+ foreignPVC.OwnerReferences = nil
+ env := newCNPGActionEnv(t, []runtime.Object{cluster, currentBackup, oldBackup}, pod, ownedPVC, foreignPVC)
+ job := &batchv1.Job{ObjectMeta: metav1.ObjectMeta{Name: "pg-init", Namespace: cluster.GetNamespace(), Labels: map[string]string{clusterLabel: cluster.GetName()}, OwnerReferences: []metav1.OwnerReference{{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: cluster.GetName(), UID: cluster.GetUID(), Controller: boolPtr(true)}}}, Spec: batchv1.JobSpec{Template: corev1.PodTemplateSpec{Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "init", Command: []string{"psql", "postgresql://app:job-password@db/pg"}}}}}}}
+ job.UID = "current-job-uid"
+ job.Spec.Template.Annotations = pod.Annotations
+ job.Spec.Template.Spec.Containers[0].Env = []corev1.EnvVar{{Name: "FROM_SECRET", ValueFrom: &corev1.EnvVarSource{SecretKeyRef: &corev1.SecretKeySelector{LocalObjectReference: corev1.LocalObjectReference{Name: "pg-job-credentials"}, Key: "password"}}}}
+ if _, err := env.typed.BatchV1().Jobs(job.Namespace).Create(context.Background(), job, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ for _, fixture := range []struct{ name, ownerUID string }{{"current-job-pod", string(job.UID)}, {"predecessor-job-pod", "old-job-uid"}} {
+ p := cnpgActionPod(fixture.name, fixture.name+"-uid", false)
+ p.OwnerReferences = []metav1.OwnerReference{{APIVersion: "batch/v1", Kind: "Job", Name: job.Name, UID: types.UID(fixture.ownerUID), Controller: boolPtr(true)}}
+ if _, err := env.typed.CoreV1().Pods(p.Namespace).Create(context.Background(), p, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ }
+ event := &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: "pg-failed", Namespace: cluster.GetNamespace()}, InvolvedObject: corev1.ObjectReference{Kind: "Cluster", Name: cluster.GetName(), UID: cluster.GetUID()}, Message: "failed connecting to postgresql://app:event-password@db/pg"}
+ if _, err := env.typed.CoreV1().Events(event.Namespace).Create(context.Background(), event, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ for _, fixture := range []struct{ event, kind, name, uid string }{
+ {"current-volume-event", "PersistentVolumeClaim", ownedPVC.Name, string(ownedPVC.UID)},
+ {"unowned-volume-event", "PersistentVolumeClaim", foreignPVC.Name, string(foreignPVC.UID)},
+ {"predecessor-cluster-event", "Cluster", cluster.GetName(), "old-cluster-uid"},
+ {"predecessor-pod-event", "Pod", pod.Name, "old-pod-uid"},
+ {"unverified-pod-event", "Pod", pod.Name, ""},
+ {"current-backup-event", "Backup", currentBackup.GetName(), string(currentBackup.GetUID())},
+ {"predecessor-backup-event", "Backup", oldBackup.GetName(), string(oldBackup.GetUID())},
+ } {
+ e := &corev1.Event{ObjectMeta: metav1.ObjectMeta{Name: fixture.event, Namespace: cluster.GetNamespace()}, InvolvedObject: corev1.ObjectReference{Kind: fixture.kind, Name: fixture.name, UID: types.UID(fixture.uid)}}
+ if _, err := env.typed.CoreV1().Events(e.Namespace).Create(context.Background(), e, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ }
+ var buf bytes.Buffer
+ z := &reportZip{zw: zip.NewWriter(&buf), root: "r", limit: reportTotalCap}
+ index := CNPGReportIndex{}
+ ctx := context.Background()
+ b := &cnpgReportBuilder{reader: newTestReader(nil), ctx: ctx, dyn: env.dyn, typed: env.typed, cluster: cluster, z: z, index: &index, secrets: map[string]map[string]bool{}}
+ b.build(ReportOptions{})
+ if err := z.zw.Close(); err != nil {
+ t.Fatal(err)
+ }
+ zr, err := zip.NewReader(bytes.NewReader(buf.Bytes()), int64(buf.Len()))
+ if err != nil {
+ t.Fatal(err)
+ }
+ files := map[string]string{}
+ for _, f := range zr.File {
+ rc, _ := f.Open()
+ data, _ := io.ReadAll(rc)
+ rc.Close()
+ files[f.Name] = string(data)
+ }
+ for _, want := range []string{"r/manifests/cluster.yaml", "r/manifests/cluster-pods.yaml", "r/manifests/backups.yaml"} {
+ if _, ok := files[want]; !ok {
+ t.Errorf("missing %s (have %v)", want, keysOf(files))
+ }
+ }
+ if strings.Contains(files["r/manifests/cluster-pods.yaml"], "hunter2") || !strings.Contains(files["r/manifests/cluster-pods.yaml"], "/var/lib") {
+ t.Fatal("sensitive env values must be redacted, others kept")
+ }
+ for file, password := range map[string]string{"r/manifests/cluster-pods.yaml": "pod-password", "r/manifests/cluster-jobs.yaml": "job-password", "r/manifests/events.yaml": "event-password"} {
+ if files[file] == "" || strings.Contains(files[file], password) || !strings.Contains(files[file], "REDACTED") {
+ t.Errorf("inline password survived %s: %s", file, files[file])
+ }
+ }
+ if strings.Contains(files["r/manifests/cluster-pods.yaml"], "cnpg.io/podSpec") || !strings.Contains(files["r/manifests/cluster-pods.yaml"], "cnpg.io/instanceRole") {
+ t.Fatal("report must remove duplicate spec annotations while retaining diagnostic annotations")
+ }
+ sourcePod, err := env.typed.CoreV1().Pods(pod.Namespace).Get(ctx, pod.Name, metav1.GetOptions{})
+ if err != nil || sourcePod.Annotations["cnpg.io/podSpec"] != string(podSpecCopy) {
+ t.Fatalf("report changed source annotations: %v", err)
+ }
+ secrets := b.secretRefs()
+ if len(secrets) != 2 || secrets[0].Name != "pg-job-credentials" || len(secrets[0].ReferencedBy) != 1 || secrets[0].ReferencedBy[0] != "Job/pg-init" || secrets[1].Name != "pg-superuser" {
+ t.Fatalf("secret names: %+v", secrets)
+ }
+ if !strings.Contains(files["r/manifests/cluster-jobs.yaml"], "pg-job-credentials") {
+ t.Fatal("Job-only Secret references must remain in the manifest")
+ }
+ for file, expected := range map[string]string{
+ "r/manifests/cluster-pods.yaml": "current-job-pod",
+ "r/manifests/cluster-pvcs.yaml": "current-pvc",
+ "r/manifests/backups.yaml": "current-backup",
+ } {
+ if !strings.Contains(files[file], expected) {
+ t.Errorf("report lost the current object %s: %s", expected, files[file])
+ }
+ }
+ for _, expected := range []string{"current-volume-event", "current-backup-event", "pg-failed"} {
+ if !strings.Contains(files["r/manifests/events.yaml"], expected) {
+ t.Errorf("report lost the current Event %s", expected)
+ }
+ }
+ for file, data := range files {
+ for _, excluded := range []string{"predecessor-", "unowned-pvc", "unowned-volume-event", "unverified-pod-event"} {
+ if strings.Contains(data, excluded) {
+ t.Errorf("report included %s in %s", excluded, file)
+ }
+ }
+ }
+ for _, item := range index.Contents {
+ if item.Item == "Events" && !strings.Contains(item.Note, "3 same-name Event(s)") {
+ t.Errorf("unverified Event coverage was not recorded: %+v", item)
+ }
+ }
+ var logs *CNPGReportItem
+ for i := range index.Contents {
+ if index.Contents[i].Item == "Logs" {
+ logs = &index.Contents[i]
+ }
+ }
+ if logs == nil || logs.State != cnpgReadSkipped {
+ t.Fatalf("logs are opt-in: %+v", index.Contents)
+ }
+}
+
+func TestCNPGOperatorWatchAndMetricsPort(t *testing.T) {
+ c := &corev1.Container{Env: []corev1.EnvVar{{Name: "WATCH_NAMESPACE", Value: "a, b"}}, Ports: []corev1.ContainerPort{{Name: "metrics", ContainerPort: 9999}}}
+ w := cnpgOperatorWatchOf(c, "cnpg-system")
+ if w.All || len(w.Namespaces) != 2 || w.Namespaces[1] != "b" {
+ t.Fatalf("watch: %+v", w)
+ }
+ if cnpgOperatorMetricsPort(c) != 9999 {
+ t.Fatal("the named metrics port wins")
+ }
+ field := &corev1.Container{Env: []corev1.EnvVar{{Name: "WATCH_NAMESPACE", ValueFrom: &corev1.EnvVarSource{FieldRef: &corev1.ObjectFieldSelector{FieldPath: "metadata.namespace"}}}}}
+ if w := cnpgOperatorWatchOf(field, "cnpg-system"); w.All || w.Namespaces[0] != "cnpg-system" {
+ t.Fatalf("fieldRef watch: %+v", w)
+ }
+ if w := cnpgOperatorWatchOf(&corev1.Container{}, "x"); !w.All {
+ t.Fatal("unset watches all namespaces")
+ }
+ if cnpgOperatorMetricsPort(&corev1.Container{Args: []string{"--metrics-bind-address=:8443"}}) != 8443 {
+ t.Fatal("the bind address flag is the fallback")
+ }
+ if cnpgLeaseHolderPod("cnpg-controller-manager-5fbdd6bb78-jx82z_8f70a311-0f38") != "cnpg-controller-manager-5fbdd6bb78-jx82z" {
+ t.Fatal("holder identity is _")
+ }
+}
+
+func TestCNPGReconcileStatsAndEndpoints(t *testing.T) {
+ samples, err := parseCNPGPromSamples([]byte(`# TYPE controller_runtime_reconcile_errors_total counter
+controller_runtime_reconcile_errors_total{controller="cluster"} 3
+# TYPE controller_runtime_reconcile_total counter
+controller_runtime_reconcile_total{controller="cluster",result="success"} 10
+controller_runtime_reconcile_total{controller="cluster",result="error"} 3
+`))
+ if err != nil {
+ t.Fatal(err)
+ }
+ stats := cnpgReconcileStats(samples)
+ if len(stats) != 1 || *stats[0].Errors != 3 || *stats[0].Total != 13 || stats[0].Results["error"] != 3 {
+ t.Fatalf("stats: %+v", stats)
+ }
+ no := false
+ ready, notReady := cnpgEndpointCounts([]discoveryv1.EndpointSlice{{Endpoints: []discoveryv1.Endpoint{{}, {Conditions: discoveryv1.EndpointConditions{Ready: &no}}}}})
+ if ready != 1 || notReady != 1 {
+ t.Fatalf("an unset ready condition counts as ready: %d/%d", ready, notReady)
+ }
+}
+
+func keysOf(m map[string]string) []string {
+ out := make([]string, 0, len(m))
+ for k := range m {
+ out = append(out, k)
+ }
+ return out
+}
+
+func apiForbidden(resource string) error {
+ return apierrors.NewForbidden(schema.GroupResource{Resource: resource}, "", errors.New("denied"))
+}
+
+func TestCNPGReportCleanObjectRedactsDeclaredEnv(t *testing.T) {
+ cluster := &unstructured.Unstructured{Object: map[string]any{
+ "spec": map[string]any{
+ "env": []any{
+ map[string]any{"name": "AWS_SECRET_ACCESS_KEY", "value": "plaintext"},
+ map[string]any{"name": "TZ", "value": "UTC"},
+ map[string]any{"name": "PGURI", "value": "postgresql://user:secret-value@db.example/pg"},
+ map[string]any{"name": "OPTIONS", "value": "password=secret-value"},
+ },
+ "backup": map[string]any{"barmanObjectStore": map[string]any{"s3Credentials": map[string]any{"secretAccessKey": map[string]any{"name": "creds", "key": "k"}}}},
+ },
+ }}
+ pooler := &unstructured.Unstructured{Object: map[string]any{
+ "spec": map[string]any{"template": map[string]any{"spec": map[string]any{"containers": []any{
+ map[string]any{"name": "pgbouncer", "env": []any{map[string]any{"name": "DB_PASSWORD", "value": "hunter2"}}},
+ }}}},
+ }}
+ c := cnpgReportCleanObject(cluster).Object["spec"].(map[string]any)
+ env := c["env"].([]any)
+ if env[0].(map[string]any)["value"] != cnpgReportRedacted || env[1].(map[string]any)["value"] != "UTC" {
+ t.Errorf("cluster env = %v", env)
+ }
+ for _, i := range []int{2, 3} {
+ if strings.Contains(env[i].(map[string]any)["value"].(string), "secret-value") {
+ t.Errorf("inline credential leaked: %v", env[i])
+ }
+ }
+ if name := c["backup"].(map[string]any)["barmanObjectStore"].(map[string]any)["s3Credentials"].(map[string]any)["secretAccessKey"].(map[string]any)["name"]; name != "creds" {
+ t.Errorf("Secret reference name was blanked: %v", name)
+ }
+ pc := cnpgReportCleanObject(pooler).Object["spec"].(map[string]any)["template"].(map[string]any)["spec"].(map[string]any)["containers"].([]any)[0].(map[string]any)
+ if pc["env"].([]any)[0].(map[string]any)["value"] != cnpgReportRedacted {
+ t.Errorf("pooler env = %v", pc["env"])
+ }
+ if cluster.Object["spec"].(map[string]any)["env"].([]any)[0].(map[string]any)["value"] != "plaintext" {
+ t.Error("the source object was modified")
+ }
+}
+
+func TestCNPGRestoreCapabilityRefusedWhileWebhookRejects(t *testing.T) {
+ rejects := true
+ got := cnpgOperatorWebhookGuard(CNPGOperatorVerdict{WebhookRejects: &rejects, WebhookReason: "the validating webhook has no ready endpoint"},
+ integration.CapabilityVerdict("", []string{integration.PermissionAllowed}, []auth.Grant{cnpgGrantCreateClusters.In("db")}))
+ if got.Allowed || !strings.Contains(got.Reason, "no ready endpoint") {
+ t.Errorf("restore while the webhook rejects = %+v", got)
+ }
+}
+
+func TestCNPGReportPodEnvRedactsInlineCredentials(t *testing.T) {
+ original := corev1.PodSpec{Containers: []corev1.Container{{Name: "postgres", Env: []corev1.EnvVar{{Name: "PGURI", Value: "postgresql://user:secret-value@db.example/pg"}, {Name: "TZ", Value: "UTC"}, {Name: "FROM_SECRET", ValueFrom: &corev1.EnvVarSource{SecretKeyRef: &corev1.SecretKeySelector{LocalObjectReference: corev1.LocalObjectReference{Name: "credentials"}, Key: "password"}}}}}}, InitContainers: []corev1.Container{{Name: "init", Env: []corev1.EnvVar{{Name: "OPTIONS", Value: "password=secret-value"}}}}}
+ clean := original.DeepCopy()
+ cnpgReportCleanPodSpec(clean)
+ if strings.Contains(clean.Containers[0].Env[0].Value, "secret-value") || strings.Contains(clean.InitContainers[0].Env[0].Value, "secret-value") {
+ t.Fatalf("inline credentials survived: %+v", clean)
+ }
+ if clean.Containers[0].Env[1].Value != "UTC" || clean.Containers[0].Env[2].ValueFrom.SecretKeyRef.Name != "credentials" {
+ t.Fatalf("ordinary values or references changed: %+v", clean)
+ }
+ if !strings.Contains(original.Containers[0].Env[0].Value, "secret-value") {
+ t.Fatal("source was modified")
+ }
+}
+
+func TestCNPGReportRedactsAllContainerKindsWithoutChangingSources(t *testing.T) {
+ original := corev1.PodSpec{
+ InitContainers: []corev1.Container{{Name: "init", Command: []string{"sh", "-c", "PGPASSWORD=init-secret psql"}}},
+ Containers: []corev1.Container{{Name: "postgres", Args: []string{"postgresql://app:container-secret@db/pg", "--port=5432"}, ReadinessProbe: &corev1.Probe{ProbeHandler: corev1.ProbeHandler{Exec: &corev1.ExecAction{Command: []string{"psql", "postgresql://app:probe-secret@db/pg"}}}}, Lifecycle: &corev1.Lifecycle{PreStop: &corev1.LifecycleHandler{Exec: &corev1.ExecAction{Command: []string{"sh", "-c", "PGPASSWORD=stop-secret psql"}}}}}},
+ EphemeralContainers: []corev1.EphemeralContainer{{EphemeralContainerCommon: corev1.EphemeralContainerCommon{Name: "debug", Command: []string{"psql"}, Args: []string{"postgresql://app:debug-secret@db/pg"}, Env: []corev1.EnvVar{{Name: "DB_PASSWORD", Value: "env-secret"}, {Name: "FROM_SECRET", ValueFrom: &corev1.EnvVarSource{SecretKeyRef: &corev1.SecretKeySelector{LocalObjectReference: corev1.LocalObjectReference{Name: "debug-credentials"}, Key: "password"}}}}}}},
+ }
+ clean := original.DeepCopy()
+ cnpgReportCleanPodSpec(clean)
+ data, err := json.Marshal(clean)
+ if err != nil {
+ t.Fatal(err)
+ }
+ for _, secret := range []string{"init-secret", "container-secret", "probe-secret", "stop-secret", "debug-secret", "env-secret"} {
+ if strings.Contains(string(data), secret) {
+ t.Errorf("report exposed %s: %s", secret, data)
+ }
+ }
+ for _, kept := range []string{"sh", "psql", "--port=5432", "debug-credentials"} {
+ if !strings.Contains(string(data), kept) {
+ t.Errorf("report lost %s", kept)
+ }
+ }
+ source, _ := json.Marshal(original)
+ if !strings.Contains(string(source), "debug-secret") || !strings.Contains(string(source), "env-secret") {
+ t.Fatal("report changed source container values")
+ }
+ obj := &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{"template": map[string]any{"spec": map[string]any{"containers": []any{map[string]any{"name": "pooler", "command": []any{"psql", "postgresql://app:template-secret@db/pg"}, "args": []any{"--port=5432"}}}}}}}}
+ template := obj.Object["spec"].(map[string]any)["template"].(map[string]any)
+ template["metadata"] = map[string]any{"annotations": map[string]any{"cnpg.io/podSpec": "template-secret", "kubectl.kubernetes.io/last-applied-configuration": "template-secret", "cnpg.io/instanceRole": "primary"}}
+ manifest, _ := json.Marshal(cnpgReportCleanObject(obj))
+ if strings.Contains(string(manifest), "template-secret") || !strings.Contains(string(manifest), "--port=5432") || !strings.Contains(string(manifest), "cnpg.io/instanceRole") {
+ t.Fatalf("template command redaction: %s", manifest)
+ }
+ if template["metadata"].(map[string]any)["annotations"].(map[string]any)["cnpg.io/podSpec"] != "template-secret" {
+ t.Fatal("report changed source template annotations")
+ }
+}
+
+func TestCNPGReportEnvRedactsKeywordAndQueryPasswords(t *testing.T) {
+ for _, value := range []string{"password=abc123", "host=db password = longsecret", "password = 'a b'", "postgresql://db/pg?password=x&sslmode=require", "postgresql://db/pg?pass%77ord=x", `{"password":"short"}`, "sslpassword: short", "PGPASSWORD=abc psql -h db", "DB_PASSWORD: x"} {
+ t.Run(value, func(t *testing.T) {
+ obj := &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{"env": []any{map[string]any{"name": "OPTIONS", "value": value}}}}}
+ out := cnpgReportCleanObject(obj).Object["spec"].(map[string]any)["env"].([]any)[0].(map[string]any)["value"]
+ if out != cnpgReportRedacted {
+ t.Fatalf("manifest env leaked: %v", out)
+ }
+ containers := []corev1.Container{{Env: []corev1.EnvVar{{Name: "OPTIONS", Value: value}}}}
+ cnpgReportCleanPodSpec(&corev1.PodSpec{Containers: containers})
+ if containers[0].Env[0].Value != cnpgReportRedacted {
+ t.Fatalf("Pod env leaked: %v", containers[0].Env[0])
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/report.go b/internal/cnpg/report.go
new file mode 100644
index 0000000000..5c9a094185
--- /dev/null
+++ b/internal/cnpg/report.go
@@ -0,0 +1,852 @@
+package cnpg
+
+import (
+ "archive/zip"
+ "bufio"
+ "bytes"
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "io"
+ "net/http"
+ "net/url"
+ "regexp"
+ "sort"
+ "strings"
+ "time"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/dynamic"
+ "k8s.io/client-go/kubernetes"
+ "sigs.k8s.io/yaml"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ aicontext "github.com/skyhook-io/radar/pkg/ai/context"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+const (
+ reportTotalCap = 32 << 20
+ cnpgReportLogCap = 2 << 20
+ ReportDefaultTail = 1000
+ ReportMaxTail = 10000
+ ReportTimeout = 40 * time.Second
+ cnpgReportRedacted = "[REDACTED]"
+ cnpgReportQueryRedacted = "[query text omitted: download with query text included to keep it]"
+)
+
+var (
+ errCNPGReportFull = errors.New("report size bound reached")
+ cnpgObjectStoreGVR = schema.GroupVersionResource{Group: barmanGroup, Version: "v1", Resource: "objectstores"}
+ cnpgGrantListBackups = auth.Grant{Verb: "list", Group: Group, Resource: "backups"}
+ cnpgGrantListSchedules = auth.Grant{Verb: "list", Group: Group, Resource: "scheduledbackups"}
+ cnpgGrantListPoolers = auth.Grant{Verb: "list", Group: Group, Resource: "poolers"}
+ cnpgGrantGetObjectStores = auth.Grant{Verb: "get", Group: barmanGroup, Resource: "objectstores"}
+ cnpgGrantGetPodLogs = auth.Grant{Verb: "get", Resource: "pods", Subresource: "log"}
+ cnpgReportEnvPassword = regexp.MustCompile(`(?i)password["']?\s*[=:]`)
+)
+
+type ReportOptions struct {
+ Logs bool `json:"logs"`
+ QueryText bool `json:"queryText"`
+ TailLines int64 `json:"tailLines,omitempty"`
+}
+
+// CNPGReportItem is one entry of the bundle's contents: a file that was
+// written, or a read that was skipped and why.
+type CNPGReportItem struct {
+ Item string `json:"item"`
+ File string `json:"file,omitempty"`
+ Count *int `json:"count,omitempty"`
+ integration.ReadSource
+ Note string `json:"note,omitempty"`
+}
+
+type CNPGReportSecretRef struct {
+ Name string `json:"name"`
+ ReferencedBy []string `json:"referencedBy"`
+}
+
+// CNPGReportIndex is report.json at the root of the bundle.
+type CNPGReportIndex struct {
+ GeneratedAt string `json:"generatedAt"`
+ RadarVersion string `json:"radarVersion"`
+ Context string `json:"context"`
+ RequestedBy string `json:"requestedBy,omitempty"`
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ Options ReportOptions `json:"options"`
+ Contents []CNPGReportItem `json:"contents"`
+ Secrets []CNPGReportSecretRef `json:"secrets"`
+ BoundBytes int64 `json:"boundBytes"`
+ Truncated bool `json:"truncated"`
+}
+
+type reportZip struct {
+ zw *zip.Writer
+ root string
+ used int64
+ limit int64
+}
+
+func (z *reportZip) add(path string, data []byte) error {
+ if z.used+int64(len(data)) > z.limit {
+ return errCNPGReportFull
+ }
+ f, err := z.zw.CreateHeader(&zip.FileHeader{Name: z.root + "/" + path, Method: zip.Deflate, Modified: time.Now()})
+ if err != nil {
+ return err
+ }
+ z.used += int64(len(data))
+ _, err = f.Write(data)
+ return err
+}
+
+func (b *cnpgReportBuilder) record(item CNPGReportItem) {
+ b.index.Contents = append(b.index.Contents, item)
+}
+
+func (b *cnpgReportBuilder) write(item CNPGReportItem, file string, v any, count int) {
+ data, err := yaml.Marshal(v)
+ if err != nil {
+ item.ReadSource = integration.ReadSource{State: cnpgReadError, Reason: err.Error()}
+ b.record(item)
+ return
+ }
+ b.writeBytes(item, file, data, &count)
+}
+
+func (b *cnpgReportBuilder) writeBytes(item CNPGReportItem, file string, data []byte, count *int) {
+ if err := b.z.add(file, data); err != nil {
+ b.index.Truncated = b.index.Truncated || errors.Is(err, errCNPGReportFull)
+ item.ReadSource = integration.ReadSource{State: cnpgReadSkipped, Reason: err.Error()}
+ b.record(item)
+ return
+ }
+ item.File, item.Count = file, count
+ if item.State == "" {
+ item.State = cnpgReadOK
+ }
+ b.record(item)
+}
+
+func (b *cnpgReportBuilder) addSecret(name, by string) {
+ if name == "" {
+ return
+ }
+ if b.secrets[name] == nil {
+ b.secrets[name] = map[string]bool{}
+ }
+ b.secrets[name][by] = true
+}
+
+func (b *cnpgReportBuilder) secretRefs() []CNPGReportSecretRef {
+ out := make([]CNPGReportSecretRef, 0, len(b.secrets))
+ for name, by := range b.secrets {
+ out = append(out, CNPGReportSecretRef{Name: name, ReferencedBy: cnpgSortedKeys(by)})
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Name < out[j].Name })
+ return out
+}
+
+func (b *cnpgReportBuilder) build(opts ReportOptions) {
+ namespace, name := b.cluster.GetNamespace(), b.cluster.GetName()
+ selector := labels.SelectorFromSet(labels.Set{clusterLabel: name}).String()
+
+ clusterObj := cnpgReportCleanObject(b.cluster)
+ cnpgReportSecretNamesIn(clusterObj.Object, "Cluster/"+name, b.addSecret)
+ b.write(CNPGReportItem{Item: "Cluster"}, "manifests/cluster.yaml", clusterObj.Object, 1)
+
+ var jobs []batchv1.Job
+ jobCov := b.reader.gatedRead(b.ctx, cnpgGrantListJobs, namespace, func() error {
+ list, err := b.typed.BatchV1().Jobs(namespace).List(b.ctx, metav1.ListOptions{LabelSelector: selector})
+ if err == nil {
+ jobs = list.Items
+ }
+ return err
+ })
+ ownedJobs := map[string]types.UID{}
+ var keptJobs []batchv1.Job
+ for _, j := range jobs {
+ if controlledBy(j.OwnerReferences, Group, "Cluster", name, b.cluster.GetUID()) {
+ ownedJobs[j.Name] = j.UID
+ for _, sref := range cnpgPodSecretNames(&j.Spec.Template.Spec) {
+ b.addSecret(sref, "Job/"+j.Name)
+ }
+ cnpgReportCleanPodSpec(&j.Spec.Template.Spec)
+ cnpgReportCleanMetadata(&j)
+ cnpgReportCleanMetadata(&j.Spec.Template.ObjectMeta)
+ keptJobs = append(keptJobs, j)
+ }
+ }
+
+ var pods []corev1.Pod
+ podCov := b.reader.gatedRead(b.ctx, GrantListPods, namespace, func() error {
+ list, err := b.typed.CoreV1().Pods(namespace).List(b.ctx, metav1.ListOptions{LabelSelector: selector})
+ if err == nil {
+ pods = list.Items
+ }
+ return err
+ })
+ var keptPods []corev1.Pod
+ unowned := 0
+ for _, p := range pods {
+ ref := controllerRef(p.OwnerReferences)
+ owned := controlledBy(p.OwnerReferences, Group, "Cluster", name, b.cluster.GetUID())
+ if ref != nil && ownedJobs[ref.Name] != "" {
+ owned = owned || controlledBy(p.OwnerReferences, "batch", "Job", ref.Name, ownedJobs[ref.Name])
+ }
+ if !owned {
+ unowned++
+ continue
+ }
+ for _, sref := range cnpgPodSecretNames(&p.Spec) {
+ b.addSecret(sref, "Pod/"+p.Name)
+ }
+ cnpgReportCleanPodSpec(&p.Spec)
+ cnpgReportCleanMetadata(&p)
+ keptPods = append(keptPods, p)
+ }
+
+ if podCov.State == cnpgReadOK {
+ item := CNPGReportItem{Item: "Pods"}
+ if unowned > 0 {
+ item.Note = fmt.Sprintf("%d Pod(s) labelled %s=%s are not owned by this Cluster or its Jobs and were left out", unowned, clusterLabel, name)
+ }
+ b.write(item, "manifests/cluster-pods.yaml", corev1.PodList{TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "PodList"}, Items: keptPods}, len(keptPods))
+ } else {
+ b.record(CNPGReportItem{Item: "Pods", ReadSource: podCov})
+ }
+ if jobCov.State == cnpgReadOK {
+ b.write(CNPGReportItem{Item: "Jobs"}, "manifests/cluster-jobs.yaml", batchv1.JobList{TypeMeta: metav1.TypeMeta{APIVersion: "batch/v1", Kind: "JobList"}, Items: keptJobs}, len(keptJobs))
+ } else {
+ b.record(CNPGReportItem{Item: "Jobs", ReadSource: jobCov})
+ }
+
+ var pvcs []corev1.PersistentVolumeClaim
+ pvcCov := b.reader.gatedRead(b.ctx, cnpgGrantListPVCs, namespace, func() error {
+ list, err := b.typed.CoreV1().PersistentVolumeClaims(namespace).List(b.ctx, metav1.ListOptions{LabelSelector: selector})
+ if err == nil {
+ pvcs = list.Items
+ }
+ return err
+ })
+ var keptPVCs []corev1.PersistentVolumeClaim
+ if pvcCov.State == cnpgReadOK {
+ for _, c := range pvcs {
+ if cnpgOwnedBy(&c, b.cluster) {
+ cnpgReportCleanMetadata(&c)
+ keptPVCs = append(keptPVCs, c)
+ }
+ }
+ b.write(CNPGReportItem{Item: "PersistentVolumeClaims"}, "manifests/cluster-pvcs.yaml", corev1.PersistentVolumeClaimList{TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "PersistentVolumeClaimList"}, Items: keptPVCs}, len(keptPVCs))
+ } else {
+ b.record(CNPGReportItem{Item: "PersistentVolumeClaims", ReadSource: pvcCov})
+ }
+
+ subjects := map[string]types.UID{"Cluster/" + name: b.cluster.GetUID()}
+ for _, p := range keptPods {
+ subjects["Pod/"+p.Name] = p.UID
+ }
+ for _, j := range keptJobs {
+ subjects["Job/"+j.Name] = j.UID
+ }
+ for _, c := range keptPVCs {
+ subjects["PersistentVolumeClaim/"+c.Name] = c.UID
+ }
+ for _, child := range []struct {
+ item, file string
+ gvr schema.GroupVersionResource
+ grant auth.Grant
+ kind string
+ }{
+ {"Backups", "manifests/backups.yaml", cnpgBackupGVR, cnpgGrantListBackups, "Backup"},
+ {"ScheduledBackups", "manifests/scheduledbackups.yaml", ScheduleGVR, cnpgGrantListSchedules, "ScheduledBackup"},
+ {"Poolers", "manifests/poolers.yaml", cnpgPoolerGVR, cnpgGrantListPoolers, "Pooler"},
+ } {
+ var items []unstructured.Unstructured
+ cov := b.reader.gatedRead(b.ctx, child.grant, namespace, func() error {
+ list, err := b.dyn.Resource(child.gvr).Namespace(namespace).List(b.ctx, metav1.ListOptions{})
+ if err == nil {
+ items = list.Items
+ }
+ return err
+ })
+ if cov.State != cnpgReadOK {
+ b.record(CNPGReportItem{Item: child.item, ReadSource: cov})
+ continue
+ }
+ var kept []any
+ for i := range items {
+ if cn, _, _ := unstructured.NestedString(items[i].Object, "spec", "cluster", "name"); cn != name {
+ continue
+ }
+ if child.kind == "Backup" && !cnpg.BackupMatchesCluster(&items[i], b.cluster) {
+ continue
+ }
+ clean := cnpgReportCleanObject(&items[i])
+ cnpgReportSecretNamesIn(clean.Object, child.kind+"/"+clean.GetName(), b.addSecret)
+ subjects[child.kind+"/"+clean.GetName()] = clean.GetUID()
+ kept = append(kept, clean.Object)
+ }
+ b.write(CNPGReportItem{Item: child.item}, child.file, map[string]any{"apiVersion": "v1", "kind": "List", "items": kept}, len(kept))
+ }
+
+ b.objectStores()
+
+ var events []corev1.Event
+ evCov := b.reader.gatedRead(b.ctx, cnpgGrantListEvents, namespace, func() error {
+ list, err := b.typed.CoreV1().Events(namespace).List(b.ctx, metav1.ListOptions{})
+ if err == nil {
+ events = list.Items
+ }
+ return err
+ })
+ if evCov.State == cnpgReadOK {
+ var kept []corev1.Event
+ unverified := 0
+ for _, e := range events {
+ uid, related := subjects[e.InvolvedObject.Kind+"/"+e.InvolvedObject.Name]
+ if !related {
+ continue
+ }
+ if uid == "" || e.InvolvedObject.UID != uid {
+ unverified++
+ continue
+ }
+ cnpgReportCleanMetadata(&e)
+ e.Message = cnpgReportInlineValue("", e.Message)
+ kept = append(kept, e)
+ }
+ sort.SliceStable(kept, func(i, j int) bool { return cnpgEventTime(&kept[i]).Before(cnpgEventTime(&kept[j])) })
+ note := "Events whose subject UID matches the Cluster or an object in this report"
+ if unverified > 0 {
+ note += fmt.Sprintf("; %d same-name Event(s) with a missing or different subject UID were left out", unverified)
+ }
+ b.write(CNPGReportItem{Item: "Events", Note: note}, "manifests/events.yaml", corev1.EventList{TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "EventList"}, Items: kept}, len(kept))
+ } else {
+ b.record(CNPGReportItem{Item: "Events", ReadSource: evCov})
+ }
+
+ operator, err := b.reader.Operator(b.ctx)
+ b.readJSON("Operator and plugins", "operator/operator.json", cnpgReportOperator(operator), err)
+ runtime, err := b.reader.ClusterRuntime(b.ctx, namespace, name)
+ b.readJSON("Runtime (instance manager status and exporter metrics)", "runtime/runtime.json", runtime, err)
+ storage, err := b.reader.ClusterStorage(b.ctx, namespace, name)
+ b.readJSON("Storage", "runtime/storage.json", storage, err)
+
+ if opts.Logs {
+ b.logs(keptPods, opts)
+ } else {
+ b.record(CNPGReportItem{Item: "Logs", ReadSource: integration.ReadSource{State: cnpgReadSkipped, Reason: "not requested"}})
+ }
+}
+
+func (b *cnpgReportBuilder) objectStores() {
+ stores := map[string]bool{}
+ if p := cnpgClusterBarmanPlugin(b.cluster); p != "" {
+ stores[p] = true
+ }
+ ext, _, _ := unstructured.NestedSlice(b.cluster.Object, "spec", "externalClusters")
+ for _, e := range ext {
+ m, _ := e.(map[string]any)
+ params, _, _ := unstructured.NestedStringMap(m, "plugin", "parameters")
+ if params["barmanObjectName"] != "" {
+ stores[params["barmanObjectName"]] = true
+ }
+ }
+ if len(stores) == 0 {
+ return
+ }
+ namespace := b.cluster.GetNamespace()
+ var kept []any
+ cov := integration.ReadSource{State: cnpgReadOK}
+ for _, name := range cnpgSortedKeys(stores) {
+ var obj *unstructured.Unstructured
+ c := b.reader.gatedRead(b.ctx, cnpgGrantGetObjectStores, namespace, func() error {
+ o, err := b.dyn.Resource(cnpgObjectStoreGVR).Namespace(namespace).Get(b.ctx, name, metav1.GetOptions{})
+ obj = o
+ return err
+ })
+ if c.State != cnpgReadOK {
+ cov = c
+ continue
+ }
+ clean := cnpgReportCleanObject(obj)
+ cnpgReportSecretNamesIn(clean.Object, "ObjectStore/"+name, b.addSecret)
+ kept = append(kept, clean.Object)
+ }
+ item := CNPGReportItem{Item: "ObjectStores"}
+ if cov.State != cnpgReadOK {
+ item.ReadSource = cov
+ if len(kept) == 0 {
+ b.record(item)
+ return
+ }
+ item.State = "partial"
+ }
+ b.write(item, "manifests/objectstores.yaml", map[string]any{"apiVersion": "v1", "kind": "List", "items": kept}, len(kept))
+}
+
+func cnpgClusterBarmanPlugin(cluster *unstructured.Unstructured) string {
+ for _, p := range cnpg.ParseBackupDeclaration(cluster).Plugins {
+ if p.Name == cnpg.BarmanPluginName {
+ return p.ObjectStore
+ }
+ }
+ return ""
+}
+
+func (b *cnpgReportBuilder) readJSON(item, file string, value any, err error) {
+ if err != nil {
+ status := http.StatusInternalServerError
+ var failure *ReadFailure
+ if errors.As(err, &failure) {
+ status = failure.Status
+ }
+ if errors.Is(err, ErrCNPGDisconnected) {
+ status = http.StatusServiceUnavailable
+ }
+ state := cnpgReadError
+ switch status {
+ case http.StatusForbidden:
+ state = cnpgReadDenied
+ case http.StatusNotFound:
+ state = cnpgReadNotFound
+ }
+ b.record(CNPGReportItem{Item: item, ReadSource: integration.ReadSource{State: state, Reason: fmt.Sprintf("HTTP %d %s", status, err)}})
+ return
+ }
+ body, err := json.MarshalIndent(value, "", " ")
+ if err != nil {
+ b.record(CNPGReportItem{Item: item, ReadSource: integration.ReadSource{State: cnpgReadError, Reason: err.Error()}})
+ return
+ }
+ b.writeBytes(CNPGReportItem{Item: item}, file, body, nil)
+}
+
+type cnpgReportOperatorConfig struct {
+ CNPGOperatorConfigRef
+ Keys *[]string `json:"keys,omitempty"`
+}
+
+type cnpgReportOperatorResponse struct {
+ *CNPGOperatorResponse
+ Config []cnpgReportOperatorConfig `json:"config"`
+}
+
+// Operator reports retain configuration references and keys, never their values.
+func cnpgReportOperator(resp *CNPGOperatorResponse) *cnpgReportOperatorResponse {
+ if resp == nil {
+ return nil
+ }
+ config := make([]cnpgReportOperatorConfig, len(resp.Config))
+ for i, c := range resp.Config {
+ config[i].CNPGOperatorConfigRef = c
+ if c.CNPGOperatorConfigMapState != nil && c.Data != nil {
+ keys := make([]string, 0, len(c.Data))
+ for k := range c.Data {
+ keys = append(keys, k)
+ }
+ sort.Strings(keys)
+ state := *c.CNPGOperatorConfigMapState
+ state.Data = nil
+ config[i].CNPGOperatorConfigMapState = &state
+ config[i].Keys = &keys
+ }
+ }
+ return &cnpgReportOperatorResponse{resp, config}
+}
+
+func (b *cnpgReportBuilder) logs(pods []corev1.Pod, opts ReportOptions) {
+ namespace := b.cluster.GetNamespace()
+ if b.reader.Access.Permission(b.ctx, cnpgGrantGetPodLogs.In(namespace)) == integration.PermissionDenied {
+ b.record(CNPGReportItem{Item: "Logs", ReadSource: integration.ReadSource{State: cnpgReadDenied, Grant: cnpgGrantGetPodLogs.In(namespace).Ref()}})
+ return
+ }
+ note := "query text inside PostgreSQL log records is omitted"
+ if opts.QueryText {
+ note = "includes query text from PostgreSQL log records"
+ }
+ for _, p := range pods {
+ containers := append(append([]corev1.Container{}, p.Spec.InitContainers...), p.Spec.Containers...)
+ restarts := map[string]int32{}
+ for _, st := range append(append([]corev1.ContainerStatus{}, p.Status.InitContainerStatuses...), p.Status.ContainerStatuses...) {
+ restarts[st.Name] = st.RestartCount
+ }
+ for _, c := range containers {
+ b.containerLog(p, c.Name, false, opts, note)
+ if restarts[c.Name] > 0 {
+ b.containerLog(p, c.Name, true, opts, note)
+ }
+ }
+ }
+}
+
+func (b *cnpgReportBuilder) containerLog(p corev1.Pod, container string, previous bool, opts ReportOptions, note string) {
+ suffix := ""
+ if previous {
+ suffix = "-previous"
+ }
+ item := CNPGReportItem{Item: fmt.Sprintf("Logs %s/%s%s", p.Name, container, suffix), Note: note}
+ if b.z.used >= b.z.limit {
+ b.index.Truncated = true
+ item.ReadSource = integration.ReadSource{State: cnpgReadSkipped, Reason: errCNPGReportFull.Error()}
+ b.record(item)
+ return
+ }
+ limit := int64(cnpgReportLogCap)
+ tail := opts.TailLines
+ stream, err := b.typed.CoreV1().Pods(p.Namespace).GetLogs(p.Name, &corev1.PodLogOptions{
+ Container: container, Previous: previous, Timestamps: true, TailLines: &tail, LimitBytes: &limit,
+ }).Stream(b.ctx)
+ if err != nil {
+ item.ReadSource = cnpgReadOutcome(err, cnpgGrantGetPodLogs, p.Namespace)
+ b.record(item)
+ return
+ }
+ defer stream.Close()
+ var out bytes.Buffer
+ scanner := bufio.NewScanner(io.LimitReader(stream, limit))
+ scanner.Buffer(make([]byte, 64<<10), int(limit))
+ lines := 0
+ for scanner.Scan() {
+ out.WriteString(cnpgReportLogLine(scanner.Text(), opts.QueryText))
+ out.WriteByte('\n')
+ lines++
+ }
+ if err := scanner.Err(); err != nil {
+ item.Note = strings.TrimSpace(item.Note + "; read stopped early: " + err.Error())
+ }
+ b.writeBytes(item, fmt.Sprintf("logs/%s-%s%s.log", p.Name, container, suffix), out.Bytes(), &lines)
+}
+
+var cnpgQueryRecordKeys = []string{"query", "internal_query", "context"}
+
+// cnpgSQLMessage finds where SQL starts in a PostgreSQL log message: after
+// "statement: ", "execute [/]: ", "execute fetch from : ",
+// "parse : ", "bind [/]: " (each possibly after
+// "duration: … ms "), and auto_explain's "plan:". Unanchored on purpose: a
+// false positive only redacts more.
+var cnpgSQLMessage = regexp.MustCompile(`(?s)\b(?:statement|execute fetch from \S+|execute \S+|parse \S+|bind \S+|plan):\s?`)
+
+// cnpgSQLProtocolMessage finds an extended-protocol message: after an
+// optional "duration: … ms", text starting with execute, parse or bind. The
+// statement or portal name is any string the client chose (spaces and colons
+// included), so everything after the keyword goes.
+var cnpgSQLProtocolMessage = regexp.MustCompile(`(?is)(?:^|\b(?:LOG|DEBUG[1-5]?|INFO|NOTICE|WARNING|ERROR|FATAL|PANIC):)\s*(?:duration:\s*\S+\s*ms\s+)?(?:execute|parse|bind)\b`)
+
+// cnpgSQLParameters finds bind values ("parameters: $1 = '…'"), in a DETAIL
+// field or line.
+var cnpgSQLParameters = regexp.MustCompile(`(?is)\bparameters:\s?`)
+
+func cnpgRedactSQLText(msg string) (string, bool) {
+ changed := false
+ for _, re := range []*regexp.Regexp{cnpgSQLProtocolMessage, cnpgSQLMessage, cnpgSQLParameters} {
+ if loc := re.FindStringIndex(msg); loc != nil {
+ msg = msg[:loc[1]] + cnpgReportQueryRedacted
+ changed = true
+ }
+ }
+ return msg, changed
+}
+
+// cnpgReportLogLine redacts one log line. PostgreSQL records from the instance
+// manager are JSON with the statement in record.query / record.message and
+// bind parameters in record.detail; without the query-text opt-in those are
+// replaced, and plain-text lines are cut where SQL starts. Every line also
+// gets the high-confidence secret patterns.
+func cnpgReportLogLine(line string, queryText bool) string {
+ ts, body := "", line
+ if i := strings.IndexByte(line, ' '); i > 0 && i < 40 && strings.HasPrefix(strings.TrimSpace(line[i:]), "{") {
+ ts, body = line[:i+1], line[i+1:]
+ }
+ if !queryText {
+ body = cnpgRedactQueryText(body)
+ }
+ return ts + aicontext.RedactSecrets(body)
+}
+
+func cnpgRedactQueryText(body string) string {
+ var entry map[string]any
+ if !strings.HasPrefix(body, "{") || json.Unmarshal([]byte(body), &entry) != nil {
+ out, _ := cnpgRedactSQLText(body)
+ return out
+ }
+ changed := false
+ redact := func(m map[string]any, key string) {
+ if v, ok := m[key].(string); ok {
+ if out, did := cnpgRedactSQLText(v); did {
+ m[key], changed = out, true
+ }
+ }
+ }
+ if rec, ok := entry["record"].(map[string]any); ok {
+ for _, k := range cnpgQueryRecordKeys {
+ if v, ok := rec[k].(string); ok && v != "" {
+ rec[k] = cnpgReportQueryRedacted
+ changed = true
+ }
+ }
+ redact(rec, "message")
+ redact(rec, "detail")
+ redact(rec, "hint")
+ }
+ redact(entry, "msg")
+ if !changed {
+ return body
+ }
+ out, err := json.Marshal(entry)
+ if err != nil {
+ return cnpgReportQueryRedacted
+ }
+ return string(out)
+}
+
+// Generic inline redaction would blank Secret selector names too. Scope it
+// to connection parameters; initialization SQL can contain role passwords
+// and is always withheld, independently of the log query-text option.
+func cnpgReportCleanObject(obj *unstructured.Unstructured) *unstructured.Unstructured {
+ clean := obj.DeepCopy()
+ cnpgReportCleanMetadata(clean)
+ cnpgReportRedactManifest(clean.Object)
+ return clean
+}
+
+func cnpgReportCleanMetadata(obj metav1.Object) {
+ obj.SetManagedFields(nil)
+ annotations := obj.GetAnnotations()
+ // Serialized spec copies bypass the field-level credential redaction.
+ delete(annotations, "cnpg.io/podSpec")
+ delete(annotations, "kubectl.kubernetes.io/last-applied-configuration")
+ if annotations != nil {
+ obj.SetAnnotations(annotations)
+ }
+}
+
+func cnpgReportRedactManifest(v any) {
+ switch t := v.(type) {
+ case map[string]any:
+ for k, child := range t {
+ switch k {
+ case "metadata":
+ if metadata, ok := child.(map[string]any); ok {
+ cnpgReportCleanMetadata(&unstructured.Unstructured{Object: map[string]any{"metadata": metadata}})
+ }
+ case "postInitSQL", "postInitApplicationSQL", "postInitTemplateSQL":
+ if sql, ok := child.([]any); ok {
+ for i := range sql {
+ sql[i] = cnpgReportRedacted
+ }
+ } else {
+ t[k] = cnpgReportRedacted
+ }
+ continue
+ case "connectionParameters":
+ if params, ok := child.(map[string]any); ok {
+ for _, key := range []string{"password", "sslpassword"} {
+ if _, present := params[key]; present {
+ params[key] = cnpgReportRedacted
+ }
+ }
+ aicontext.RedactInlineSecrets(params)
+ }
+ }
+ if list, ok := child.([]any); ok && (k == "command" || k == "args") {
+ for i, item := range list {
+ if value, ok := item.(string); ok {
+ list[i] = cnpgReportInlineValue("", value)
+ }
+ }
+ continue
+ }
+ if list, ok := child.([]any); ok && k == "env" {
+ for _, item := range list {
+ e, ok := item.(map[string]any)
+ if !ok {
+ continue
+ }
+ name, _ := e["name"].(string)
+ if val, ok := e["value"].(string); ok {
+ e["value"] = cnpgReportInlineValue(name, val)
+ }
+ }
+ continue
+ }
+ cnpgReportRedactManifest(child)
+ }
+ case []any:
+ for _, child := range t {
+ cnpgReportRedactManifest(child)
+ }
+ }
+}
+
+func cnpgReportInlineValue(name, value string) string {
+ decoded, err := url.QueryUnescape(value)
+ if value != "" && (aicontext.IsSensitiveEnvName(name) || cnpgReportEnvPassword.MatchString(value) || err == nil && cnpgReportEnvPassword.MatchString(decoded)) {
+ return cnpgReportRedacted
+ }
+ return aicontext.RedactSecrets(value)
+}
+
+func cnpgReportCleanContainerValues(env []corev1.EnvVar, command, args []string) {
+ for i := range env {
+ env[i].Value = cnpgReportInlineValue(env[i].Name, env[i].Value)
+ }
+ for _, values := range [][]string{command, args} {
+ cnpgReportCleanCommand(values)
+ }
+}
+
+func cnpgReportCleanCommand(command []string) {
+ for i := range command {
+ command[i] = cnpgReportInlineValue("", command[i])
+ }
+}
+
+func cnpgReportCleanPodSpec(spec *corev1.PodSpec) {
+ for _, cs := range [][]corev1.Container{spec.InitContainers, spec.Containers} {
+ for i := range cs {
+ cnpgReportCleanContainerValues(cs[i].Env, cs[i].Command, cs[i].Args)
+ for _, probe := range []*corev1.Probe{cs[i].LivenessProbe, cs[i].ReadinessProbe, cs[i].StartupProbe} {
+ if probe != nil && probe.Exec != nil {
+ cnpgReportCleanCommand(probe.Exec.Command)
+ }
+ }
+ if cs[i].Lifecycle != nil {
+ for _, handler := range []*corev1.LifecycleHandler{cs[i].Lifecycle.PostStart, cs[i].Lifecycle.PreStop} {
+ if handler != nil && handler.Exec != nil {
+ cnpgReportCleanCommand(handler.Exec.Command)
+ }
+ }
+ }
+ }
+ }
+ for i := range spec.EphemeralContainers {
+ c := &spec.EphemeralContainers[i]
+ cnpgReportCleanContainerValues(c.Env, c.Command, c.Args)
+ }
+}
+
+func cnpgPodSecretNames(spec *corev1.PodSpec) []string {
+ var out []string
+ for _, v := range spec.Volumes {
+ if v.Secret != nil {
+ out = append(out, v.Secret.SecretName)
+ }
+ if v.Projected != nil {
+ for _, src := range v.Projected.Sources {
+ if src.Secret != nil {
+ out = append(out, src.Secret.Name)
+ }
+ }
+ }
+ }
+ for _, c := range append(append([]corev1.Container{}, spec.InitContainers...), spec.Containers...) {
+ for _, e := range c.Env {
+ if e.ValueFrom != nil && e.ValueFrom.SecretKeyRef != nil {
+ out = append(out, e.ValueFrom.SecretKeyRef.Name)
+ }
+ }
+ for _, ef := range c.EnvFrom {
+ if ef.SecretRef != nil {
+ out = append(out, ef.SecretRef.Name)
+ }
+ }
+ }
+ for _, s := range spec.ImagePullSecrets {
+ out = append(out, s.Name)
+ }
+ return out
+}
+
+// cnpgReportSecretNamesIn finds Secret references in a CNPG object's spec:
+// a {name} under a key that says secret (superuserSecret, bootstrap
+// initdb.secret), a string under a key ending in secret
+// (certificates.serverTLSSecret) or a {name, key} selector that is not a
+// ConfigMap's (ObjectStore credentials). It records names only.
+func cnpgReportSecretNamesIn(obj map[string]any, by string, add func(name, by string)) {
+ var walk func(node any, under string)
+ walk = func(node any, under string) {
+ switch v := node.(type) {
+ case map[string]any:
+ name, hasName := v["name"].(string)
+ _, hasKey := v["key"].(string)
+ if hasName && (under == "secret" || (hasKey && under != "configmap")) {
+ add(name, by)
+ }
+ for k, child := range v {
+ lower := strings.ToLower(k)
+ if s, ok := child.(string); ok && strings.HasSuffix(lower, "secret") {
+ add(s, by)
+ continue
+ }
+ next := under
+ switch {
+ case strings.Contains(lower, "configmap"):
+ next = "configmap"
+ case strings.Contains(lower, "secret"):
+ next = "secret"
+ }
+ walk(child, next)
+ }
+ case []any:
+ for _, child := range v {
+ walk(child, under)
+ }
+ }
+ }
+ if spec, ok := obj["spec"]; ok {
+ walk(spec, "")
+ }
+}
+
+type cnpgReportBuilder struct {
+ reader *Reader
+ ctx context.Context
+ dyn dynamic.Interface
+ typed kubernetes.Interface
+ cluster *unstructured.Unstructured
+ z *reportZip
+ index *CNPGReportIndex
+ secrets map[string]map[string]bool
+}
+
+type ReportMetadata struct {
+ Context string
+ RequestedBy string
+ RadarVersion string
+ Now time.Time
+}
+
+func (s *Reader) Report(ctx context.Context, dyn dynamic.Interface, typed kubernetes.Interface, cluster *unstructured.Unstructured, opts ReportOptions, meta ReportMetadata) ([]byte, string, error) {
+ now := meta.Now.UTC()
+ root := fmt.Sprintf("report_cluster_%s_%s", cluster.GetName(), now.Format(compactStamp))
+ var buf bytes.Buffer
+ z := &reportZip{zw: zip.NewWriter(&buf), root: root, limit: reportTotalCap}
+ index := CNPGReportIndex{GeneratedAt: now.Format(time.RFC3339), RadarVersion: meta.RadarVersion, Context: meta.Context, RequestedBy: meta.RequestedBy,
+ Cluster: CNPGRuntimeObjectRef{Namespace: cluster.GetNamespace(), Name: cluster.GetName(), UID: cluster.GetUID()}, Options: opts, BoundBytes: reportTotalCap}
+ b := &cnpgReportBuilder{reader: s, ctx: ctx, dyn: dyn, typed: typed, cluster: cluster, z: z, index: &index, secrets: map[string]map[string]bool{}}
+ b.build(opts)
+ index.Secrets = b.secretRefs()
+ data, err := json.MarshalIndent(index, "", " ")
+ if err != nil {
+ return nil, "", err
+ }
+ z.limit += int64(len(data))
+ if err := z.add("report.json", data); err != nil {
+ return nil, "", err
+ }
+ if err := z.zw.Close(); err != nil {
+ return nil, "", err
+ }
+ return buf.Bytes(), root, nil
+}
diff --git a/internal/cnpg/report_reads_test.go b/internal/cnpg/report_reads_test.go
new file mode 100644
index 0000000000..dbe4258258
--- /dev/null
+++ b/internal/cnpg/report_reads_test.go
@@ -0,0 +1,70 @@
+package cnpg
+
+import (
+ "encoding/json"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "net/http"
+ "strings"
+ "testing"
+)
+
+func TestCNPGReportOperatorProjection(t *testing.T) {
+ resp := &CNPGOperatorResponse{Config: []CNPGOperatorConfigRef{
+ {Kind: "ConfigMap", Name: "operator", CNPGOperatorConfigMapState: &CNPGOperatorConfigMapState{Data: map[string]string{"Z": "private-z", "A": "private-a"}}},
+ {Kind: "Secret", Name: "operator-credentials"},
+ }}
+ data, err := json.Marshal(cnpgReportOperator(resp))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if strings.Contains(string(data), "private-") || !strings.Contains(string(data), `"keys":["A","Z"]`) || !strings.Contains(string(data), "operator-credentials") {
+ t.Fatalf("report projection: %s", data)
+ }
+ if resp.Config[0].Data["A"] != "private-a" {
+ t.Fatal("report mutated the endpoint's response")
+ }
+}
+
+func TestCNPGReportReadErrors(t *testing.T) {
+ for _, tc := range []struct {
+ status int
+ state string
+ }{{http.StatusForbidden, cnpgReadDenied}, {http.StatusNotFound, cnpgReadNotFound}, {http.StatusServiceUnavailable, cnpgReadError}} {
+ index := CNPGReportIndex{}
+ builder := cnpgReportBuilder{index: &index}
+ builder.readJSON("Runtime", "runtime.json", nil, &ReadFailure{tc.status, "unavailable"})
+ if len(index.Contents) != 1 || index.Contents[0].State != tc.state || index.Contents[0].File != "" {
+ t.Fatalf("read error: %+v", index.Contents)
+ }
+ }
+}
+
+func TestReportWithholdsInitializationSQLAndInlineConnectionPasswords(t *testing.T) {
+ obj := &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{
+ "bootstrap": map[string]any{"initdb": map[string]any{
+ "postInitSQL": []any{"CREATE ROLE app PASSWORD 'secret-sql'"},
+ "postInitApplicationSQL": []any{"private-application-sql"},
+ "postInitTemplateSQL": []any{"private-template-sql"},
+ "postInitSQLRefs": map[string]any{"secretRefs": []any{map[string]any{"name": "sql-secret", "key": "sql"}}},
+ }},
+ "externalClusters": []any{map[string]any{"name": "origin", "connectionParameters": map[string]any{"host": "pg-rw", "password": "secret-inline", "sslpassword": "secret-ssl"}, "password": map[string]any{"name": "password-secret", "key": "password"}}},
+ }}}
+ data, err := json.Marshal(cnpgReportCleanObject(obj))
+ if err != nil {
+ t.Fatal(err)
+ }
+ for _, secret := range []string{"secret-sql", "private-application-sql", "private-template-sql", "secret-inline", "secret-ssl"} {
+ if strings.Contains(string(data), secret) {
+ t.Errorf("report exposed %s: %s", secret, data)
+ }
+ }
+ for _, reference := range []string{"pg-rw", "sql-secret", "password-secret"} {
+ if !strings.Contains(string(data), reference) {
+ t.Errorf("report lost reference %s", reference)
+ }
+ }
+ original, _ := json.Marshal(obj)
+ if !strings.Contains(string(original), "secret-sql") {
+ t.Fatal("redaction changed source")
+ }
+}
diff --git a/internal/cnpg/runtime.go b/internal/cnpg/runtime.go
new file mode 100644
index 0000000000..91a4ee3c48
--- /dev/null
+++ b/internal/cnpg/runtime.go
@@ -0,0 +1,272 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "log"
+ "net/http"
+ "sort"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/types"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+func runtimePermission(namespace string, allowed bool) CNPGRuntimePermission {
+ p := CNPGRuntimePermission{Proxy: "allowed", Grant: grantGetPodsProxy.In(namespace).Ref()}
+ if !allowed {
+ p.Proxy = "denied"
+ }
+ return p
+}
+
+func proxyDenied(namespace string) CNPGRuntimeSource {
+ return CNPGRuntimeSource{State: runtimeStateDenied, Error: "reading live data needs get pods/proxy in " + namespace}
+}
+
+func (s *Reader) ClusterRuntime(callerCtx context.Context, namespace, name string) (*CNPGClusterRuntimeResponse, error) {
+ cache, cluster, err := s.Observations.Cluster(callerCtx, namespace, name, GrantListPods)
+ if err != nil {
+ return nil, err
+ }
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", namespace, name, err)
+ return nil, &ReadFailure{http.StatusServiceUnavailable, "instance Pods unavailable: " + err.Error()}
+ }
+
+ proxyAllowed := s.Access.Permission(callerCtx, grantGetPodsProxy.In(namespace)) != integration.PermissionDenied
+ fenced := parseCNPGFenced(cluster.GetAnnotations()[cnpgFencedAnnotation])
+ resp := CNPGClusterRuntimeResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: runtimePermission(namespace, proxyAllowed),
+ Instances: make([]CNPGInstanceRuntime, len(pods)),
+ }
+ for i, p := range pods {
+ resp.Instances[i] = CNPGInstanceRuntime{Pod: p.Name, PodUID: p.UID, Role: runtimeRole(p), Fenced: fenced.Fences(p.Name)}
+ }
+ if len(pods) == 0 {
+ return &resp, nil
+ }
+ if !proxyAllowed {
+ for i := range resp.Instances {
+ resp.Instances[i].Status.CNPGRuntimeSource = proxyDenied(namespace)
+ resp.Instances[i].Metrics.CNPGRuntimeSource = proxyDenied(namespace)
+ }
+ return &resp, nil
+ }
+ client := s.Clients.Proxy
+ if client == nil {
+ return nil, &ReadFailure{http.StatusServiceUnavailable, "cluster client unavailable"}
+ }
+
+ clusterMetricsTLS, _, _ := unstructured.NestedBool(cluster.Object, "spec", "monitoring", "tls", "enabled")
+ identity := s.Identity
+ run := newCNPGRuntimeRunner(callerCtx)
+ for i, p := range pods {
+ inst := &resp.Instances[i]
+ statusTarget, metricsTarget := cnpgInstanceProxyTargets(p, clusterMetricsTLS)
+ run.do(func(ctx context.Context) {
+ inst.Status = memoized(ctx, identity, statusTarget, cnpgStatusMemoTTL, func(ctx context.Context) CNPGInstanceStatus {
+ return cnpgInstanceStatusFrom(proxyGetWithFallback(ctx, client, statusTarget))
+ })
+ })
+ run.do(func(ctx context.Context) {
+ inst.Metrics = memoized(ctx, identity, metricsTarget, metricsMemoTTL, func(ctx context.Context) CNPGInstanceMetrics {
+ return cnpgInstanceMetricsFrom(proxyGetWithFallback(ctx, client, metricsTarget))
+ })
+ })
+ }
+ run.wait()
+
+ for i := range resp.Instances {
+ inst := &resp.Instances[i]
+ if inst.Status.State == runtimeStateDenied || inst.Metrics.State == runtimeStateDenied {
+ resp.Permission.Proxy = "denied"
+ }
+ if inst.Fenced {
+ explainCNPGFenced(&inst.Status.CNPGRuntimeSource)
+ explainCNPGFenced(&inst.Metrics.CNPGRuntimeSource)
+ }
+ if inst.Status.CNPGInstanceStatusFacts != nil && inst.Metrics.CNPGInstanceMetricFacts != nil {
+ inst.Status.CNPGInstanceStatusFacts = withCNPGSlotRetention(inst.Status.CNPGInstanceStatusFacts, inst.Metrics.ReplicationSlotsRetainedBytes)
+ }
+ }
+ return &resp, nil
+}
+
+func (s *Reader) Pooler(ctx context.Context, cache *k8s.ResourceCache, namespace, name string) (*unstructured.Unstructured, error) {
+ poolers, err := filterCNPGGroup(s.Observations.DynamicList(ctx, cache, "Pooler", Group, namespace))
+ if err != nil {
+ return nil, err
+ }
+ for _, p := range poolers {
+ if p.GetNamespace() == namespace && p.GetName() == name && p.GroupVersionKind().Group == Group {
+ return p, nil
+ }
+ }
+ return nil, nil
+}
+
+// cnpgPoolerPods returns the Pods of the Pooler's Deployment, validated along
+// the controller chain Pooler → Deployment (named after the Pooler) →
+// ReplicaSet → Pod by UID. The poolerName label alone is something any Pod
+// can carry.
+func poolerPods(cache *k8s.ResourceCache, pooler *unstructured.Unstructured) ([]*corev1.Pod, error) {
+ podLister, rsLister, deployLister := cache.Pods(), cache.ReplicaSets(), cache.Deployments()
+ if podLister == nil || rsLister == nil || deployLister == nil {
+ return nil, errors.New("pod, ReplicaSet or Deployment cache unavailable")
+ }
+ namespace, name := pooler.GetNamespace(), pooler.GetName()
+ deploy, err := deployLister.Deployments(namespace).Get(name)
+ if apierrors.IsNotFound(err) {
+ return []*corev1.Pod{}, nil
+ }
+ if err != nil {
+ return nil, err
+ }
+ if !controlledBy(deploy.OwnerReferences, Group, "Pooler", name, pooler.GetUID()) {
+ return []*corev1.Pod{}, nil
+ }
+ candidates, err := podLister.Pods(namespace).List(labels.SelectorFromSet(labels.Set{cnpgPoolerNameLabel: name}))
+ if err != nil {
+ return nil, err
+ }
+ pods := make([]*corev1.Pod, 0, len(candidates))
+ for _, p := range candidates {
+ ref := controllerRef(p.OwnerReferences)
+ if ref == nil || ref.Kind != "ReplicaSet" {
+ continue
+ }
+ rs, err := rsLister.ReplicaSets(namespace).Get(ref.Name)
+ if err != nil || rs.UID != ref.UID {
+ continue
+ }
+ if controlledBy(rs.OwnerReferences, "apps", "Deployment", deploy.Name, deploy.UID) {
+ pods = append(pods, p)
+ }
+ }
+ sort.Slice(pods, func(i, j int) bool { return pods[i].Name < pods[j].Name })
+ return pods, nil
+}
+
+func controllerRef(refs []metav1.OwnerReference) *metav1.OwnerReference {
+ for i := range refs {
+ if refs[i].Controller != nil && *refs[i].Controller {
+ return &refs[i]
+ }
+ }
+ return nil
+}
+
+func controlledBy(refs []metav1.OwnerReference, group, kind, name string, uid types.UID) bool {
+ return cnpg.ControlledBy(refs, group, kind, name, uid)
+}
+
+func runtimeRole(p *corev1.Pod) string {
+ switch role := instanceRole(p); role {
+ case "primary", "replica":
+ return role
+ default:
+ return "unknown"
+ }
+}
+
+func explainCNPGFenced(src *CNPGRuntimeSource) {
+ if src.State == runtimeStateUnreachable || src.State == runtimeStateError {
+ src.Error = cnpgFencedErrorExplained + " (" + src.Error + ")"
+ }
+}
+
+func poolerNotStarted(p *corev1.Pod) string {
+ if p.Status.Phase == "" || p.Status.Phase == corev1.PodRunning {
+ return ""
+ }
+ reason := "PgBouncer has not started"
+ if p.Status.Phase != corev1.PodPending {
+ return "PgBouncer is not running (Pod " + string(p.Status.Phase) + ")"
+ }
+ for _, condition := range p.Status.Conditions {
+ if condition.Type == corev1.PodScheduled && condition.Status == corev1.ConditionFalse {
+ return reason + " (Pod cannot be scheduled)"
+ }
+ }
+ return reason + " (Pod " + string(p.Status.Phase) + ")"
+}
+
+func (s *Reader) PoolerRuntime(ctx context.Context, cache *k8s.ResourceCache, pooler *unstructured.Unstructured) (*CNPGPoolerRuntimeResponse, error) {
+ namespace, name := pooler.GetNamespace(), pooler.GetName()
+ pods, err := poolerPods(cache, pooler)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list pooler Pods for %s/%s: %v", namespace, name, err)
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "pooler Pods unavailable: " + err.Error()}
+ }
+
+ proxyAllowed := s.Access.Permission(ctx, grantGetPodsProxy.In(namespace)) != integration.PermissionDenied
+ resp := CNPGPoolerRuntimeResponse{
+ Pooler: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: pooler.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: runtimePermission(namespace, proxyAllowed),
+ Pods: make([]CNPGPoolerPodRuntime, len(pods)),
+ }
+ for i, p := range pods {
+ resp.Pods[i] = CNPGPoolerPodRuntime{Pod: p.Name}
+ for _, condition := range p.Status.Conditions {
+ if condition.Type == corev1.PodScheduled && condition.Status == corev1.ConditionFalse {
+ resp.Pods[i].SchedulingReason = strings.TrimSpace(condition.Reason + ": " + condition.Message)
+ }
+ }
+ }
+ if len(pods) == 0 {
+ return &resp, nil
+ }
+ if !proxyAllowed {
+ for i := range resp.Pods {
+ resp.Pods[i].CNPGRuntimeSource = proxyDenied(namespace)
+ }
+ return &resp, nil
+ }
+ client := s.Clients.Proxy
+ if client == nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "cluster client unavailable"}
+ }
+
+ poolerTLS, _, _ := unstructured.NestedBool(pooler.Object, "spec", "monitoring", "tls", "enabled")
+ identity := s.Identity
+ run := newCNPGRuntimeRunner(ctx)
+ for i, p := range pods {
+ out := &resp.Pods[i]
+ target := proxyTarget{
+ namespace: namespace, pod: p.Name, podUID: p.UID, port: poolerMetricsPort, path: metricsPath,
+ scheme: schemeFor(containerHasFlag(p, pgBouncerContainer, metricsPortTLSFlag) || poolerTLS), limit: runtimeMetricsCap,
+ }
+ run.do(func(ctx context.Context) {
+ got := memoized(ctx, identity, target, metricsMemoTTL, func(ctx context.Context) CNPGPoolerPodRuntime {
+ return poolerPodFrom(proxyGetWithFallback(ctx, client, target))
+ })
+ got.Pod = p.Name
+ got.SchedulingReason = out.SchedulingReason
+ if reason := poolerNotStarted(p); reason != "" && got.State != runtimeStateDenied && got.CNPGPoolerPodFacts == nil {
+ got.CNPGRuntimeSource = CNPGRuntimeSource{State: runtimeStateUnreachable, Reason: reason}
+ }
+ *out = got
+ })
+ }
+ run.wait()
+ for _, p := range resp.Pods {
+ if p.State == runtimeStateDenied {
+ resp.Permission.Proxy = "denied"
+ }
+ }
+ return &resp, nil
+}
diff --git a/internal/cnpg/runtime_decode.go b/internal/cnpg/runtime_decode.go
new file mode 100644
index 0000000000..2271b3c58e
--- /dev/null
+++ b/internal/cnpg/runtime_decode.go
@@ -0,0 +1,846 @@
+package cnpg
+
+import (
+ "bytes"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "math"
+ "sort"
+ "strconv"
+ "strings"
+ "time"
+
+ dto "github.com/prometheus/client_model/go"
+ "github.com/prometheus/common/expfmt"
+ "github.com/prometheus/common/model"
+)
+
+// The operator's own sessions: replication and the metrics exporter. Always
+// present and long-lived by design, so counting them would make an idle
+// database look busy. CNPG 1.27's exporter connects as postgres, so it is only
+// recognizable by its application_name.
+var cnpgPlatformUsers = map[string]bool{"streaming_replica": true, cnpgMetricsExporterApp: true}
+
+// The default-monitoring families the runtime view reads. An absent one is
+// reported in `missing` — a custom monitoring configuration can replace any
+// default query — rather than read as zero. Replication-slot retention is not
+// listed: the exporter emits nothing when an instance has no slots.
+var cnpgExpectedInstanceFamilies = []string{
+ "cnpg_backends_total",
+ "cnpg_backends_waiting_total",
+ "cnpg_pg_postmaster_start_time",
+ "cnpg_backends_max_tx_duration_seconds",
+ "cnpg_pg_settings_setting",
+ "cnpg_pg_database_size_bytes",
+ "cnpg_pg_database_xid_age",
+ "cnpg_pg_database_mxid_age",
+ "cnpg_pg_extensions_update_available",
+ "cnpg_pg_stat_archiver_archived_count",
+ "cnpg_pg_stat_archiver_failed_count",
+ "cnpg_pg_stat_archiver_seconds_since_last_archival",
+ "cnpg_pg_stat_archiver_seconds_since_last_failure",
+ "cnpg_collector_pg_wal",
+ "cnpg_pg_stat_database_xact_commit",
+ "cnpg_pg_stat_database_xact_rollback",
+ "cnpg_pg_stat_database_blks_hit",
+ "cnpg_pg_stat_database_blks_read",
+ "cnpg_pg_stat_database_deadlocks",
+}
+
+var cnpgExpectedPoolerFamilies = []string{
+ "cnpg_pgbouncer_pools_cl_active",
+ "cnpg_pgbouncer_pools_cl_waiting",
+ "cnpg_pgbouncer_pools_sv_active",
+ "cnpg_pgbouncer_pools_sv_idle",
+ "cnpg_pgbouncer_pools_sv_used",
+ "cnpg_pgbouncer_pools_maxwait",
+}
+
+// cnpgPgStatus is the subset of the instance manager's GET /pg/status answer
+// (CloudNativePG's PostgresqlStatus) the runtime view reads.
+type cnpgPgStatus struct {
+ IsPrimary *bool `json:"isPrimary"`
+ MightBeUnavailable bool `json:"mightBeUnavailable"`
+ MaskedError string `json:"mightBeUnavailableMaskedError"`
+ CurrentLsn string `json:"currentLsn"`
+ ReceivedLsn string `json:"receivedLsn"`
+ ReplayLsn string `json:"replayLsn"`
+ TimeLineID *int `json:"timeLineID"`
+ ReplayPaused bool `json:"replayPaused"`
+ PendingRestart bool `json:"pendingRestart"`
+ IsWalReceiverActive bool `json:"isWalReceiverActive"`
+
+ PendingRestartForDecrease bool `json:"pendingRestartForDecrease"`
+ IsPgRewindRunning bool `json:"isPgRewindRunning"`
+ InstanceManagerVersion string `json:"instanceManagerVersion"`
+ IsInstanceManagerUpgrading bool `json:"isInstanceManagerUpgrading"`
+ LastArchivedWAL string `json:"lastArchivedWAL"`
+ LastArchivedWALTime string `json:"lastArchivedWALTime"`
+ LastFailedWAL string `json:"lastFailedWAL"`
+ LastFailedWALTime string `json:"lastFailedWALTime"`
+ ReadyWalFiles *int `json:"readyWalFiles"`
+ ReplicationInfo []struct {
+ ApplicationName string `json:"applicationName"`
+ State string `json:"state"`
+ // CNPG serializes pg_stat_replication.sent_lsn under "receivedLsn".
+ SentLsn string `json:"receivedLsn"`
+ WriteLsn string `json:"writeLsn"`
+ FlushLsn string `json:"flushLsn"`
+ ReplayLsn string `json:"replayLsn"`
+ WriteLag string `json:"writeLag"`
+ FlushLag string `json:"flushLag"`
+ ReplayLag string `json:"replayLag"`
+ SyncState string `json:"syncState"`
+ SyncPriority json.RawMessage `json:"syncPriority"`
+ } `json:"replicationInfo"`
+ ReplicationSlotsInfo []struct {
+ SlotName string `json:"slotName"`
+ Plugin string `json:"plugin"`
+ SlotType string `json:"slotType"`
+ Database string `json:"database"`
+ Active bool `json:"active"`
+ RestartLsn string `json:"restartLsn"`
+ WalStatus string `json:"walStatus"`
+ SafeWalSize *int64 `json:"safeWalSize"`
+ } `json:"replicationSlotsInfo"`
+ PgStatBasebackupsInfo []struct {
+ ApplicationName string `json:"application_name"`
+ BackendStart string `json:"backend_start"`
+ Phase string `json:"phase"`
+ BackupTotal int64 `json:"backup_total"`
+ BackupStreamed int64 `json:"backup_streamed"`
+ TablespacesTotal int64 `json:"tablespaces_total"`
+ TablespacesStreamed int64 `json:"tablespaces_streamed"`
+ } `json:"pgStatBasebackupsInfo"`
+}
+
+func cnpgInstanceStatusFrom(out cnpgProxyOutcome) CNPGInstanceStatus {
+ src := out.source()
+ if out.state != runtimeStateOK {
+ return CNPGInstanceStatus{CNPGRuntimeSource: src}
+ }
+ if out.truncated {
+ src.State, src.Error = runtimeStateError, fmt.Sprintf("status answer larger than the %s this view reads", formatCNPGByteCap(cnpgRuntimeStatusCap))
+ return CNPGInstanceStatus{CNPGRuntimeSource: src}
+ }
+ facts, partial, err := parseCNPGPgStatus(out.body)
+ if err != nil {
+ src.State, src.Error = runtimeStateError, err.Error()
+ return CNPGInstanceStatus{CNPGRuntimeSource: src}
+ }
+ if partial != "" {
+ src.State, src.Reason = cnpgRuntimeStatePartial, partial
+ }
+ return CNPGInstanceStatus{CNPGRuntimeSource: src, CNPGInstanceStatusFacts: facts}
+}
+
+// parseCNPGPgStatus returns the facts and, when rows were capped, why the
+// answer is partial.
+func parseCNPGPgStatus(body []byte) (*CNPGInstanceStatusFacts, string, error) {
+ var st cnpgPgStatus
+ if err := json.Unmarshal(body, &st); err != nil || st.IsPrimary == nil {
+ return nil, "", errors.New("the instance manager's answer was not a status report")
+ }
+ facts := &CNPGInstanceStatusFacts{
+ IsPrimary: *st.IsPrimary,
+ MightBeUnavailable: st.MightBeUnavailable,
+ CurrentLsn: st.CurrentLsn,
+ ReceivedLsn: st.ReceivedLsn,
+ ReplayLsn: st.ReplayLsn,
+ Timeline: st.TimeLineID,
+ ReplayPaused: st.ReplayPaused,
+ PendingRestart: st.PendingRestart,
+ IsWalReceiverActive: st.IsWalReceiverActive,
+
+ PendingRestartForDecrease: st.PendingRestartForDecrease,
+ IsPgRewindRunning: st.IsPgRewindRunning,
+ InstanceManagerVersion: st.InstanceManagerVersion,
+ IsInstanceManagerUpgrading: st.IsInstanceManagerUpgrading,
+ Archiving: &CNPGArchivingStatus{
+ LastArchivedWal: st.LastArchivedWAL,
+ LastArchivedAt: cnpgStatusTime(st.LastArchivedWALTime),
+ LastFailedWal: st.LastFailedWAL,
+ LastFailedAt: cnpgStatusTime(st.LastFailedWALTime),
+ ReadyWalFiles: st.ReadyWalFiles,
+ },
+ Replication: []CNPGReplicationStatus{},
+ Slots: []CNPGSlotStatus{},
+ SlotsTruncated: len(st.ReplicationSlotsInfo) > cnpgRuntimeMaxRows,
+ BaseBackups: []CNPGBaseBackupStatus{},
+ }
+ facts.RoleDetail = cnpgRoleDetail(facts)
+ var partial []string
+ for i, ri := range st.ReplicationInfo {
+ if i >= cnpgRuntimeMaxRows {
+ partial = append(partial, fmt.Sprintf("%d replication connections; the first %d are shown", len(st.ReplicationInfo), cnpgRuntimeMaxRows))
+ break
+ }
+ facts.Replication = append(facts.Replication, CNPGReplicationStatus{
+ ApplicationName: ri.ApplicationName,
+ State: ri.State,
+ SyncState: ri.SyncState,
+ SyncPriority: parseCNPGSyncPriority(ri.SyncPriority),
+ WriteLag: parseCNPGPgInterval(ri.WriteLag),
+ WriteLagRaw: ri.WriteLag,
+ FlushLag: parseCNPGPgInterval(ri.FlushLag),
+ FlushLagRaw: ri.FlushLag,
+ ReplayLag: parseCNPGPgInterval(ri.ReplayLag),
+ ReplayLagRaw: ri.ReplayLag,
+ SentLsn: ri.SentLsn,
+ WriteLsn: ri.WriteLsn,
+ FlushLsn: ri.FlushLsn,
+ ReplayLsn: ri.ReplayLsn,
+ })
+ }
+ for i, si := range st.ReplicationSlotsInfo {
+ if i >= cnpgRuntimeMaxRows {
+ partial = append(partial, fmt.Sprintf("%d replication slots; the first %d are shown", len(st.ReplicationSlotsInfo), cnpgRuntimeMaxRows))
+ break
+ }
+ facts.Slots = append(facts.Slots, CNPGSlotStatus{
+ Name: si.SlotName, Type: si.SlotType, Plugin: si.Plugin, Active: si.Active, Database: si.Database,
+ RestartLsn: si.RestartLsn, WalStatus: si.WalStatus, SafeWalSize: si.SafeWalSize,
+ })
+ }
+ for i, bb := range st.PgStatBasebackupsInfo {
+ if i >= cnpgRuntimeMaxRows {
+ partial = append(partial, fmt.Sprintf("%d base backups; the first %d are shown", len(st.PgStatBasebackupsInfo), cnpgRuntimeMaxRows))
+ break
+ }
+ row := CNPGBaseBackupStatus{
+ ApplicationName: bb.ApplicationName,
+ Instance: strings.TrimSuffix(bb.ApplicationName, "-join"),
+ Phase: bb.Phase,
+ StartedAt: cnpgStatusTime(bb.BackendStart),
+ StreamedBytes: bb.BackupStreamed,
+ TablespacesTotal: bb.TablespacesTotal,
+ TablespacesStreamed: bb.TablespacesStreamed,
+ }
+ if row.Instance == row.ApplicationName {
+ row.Instance = ""
+ }
+ if bb.BackupTotal > 0 {
+ total := bb.BackupTotal
+ row.TotalBytes = &total
+ }
+ facts.BaseBackups = append(facts.BaseBackups, row)
+ }
+ if why := cnpgIncompleteReport(&st); why != "" {
+ markCNPGStatusIncomplete(facts, &st)
+ partial = append(partial, why)
+ }
+ return facts, strings.Join(partial, "; "), nil
+}
+
+func cnpgIncompleteReport(st *cnpgPgStatus) string {
+ switch {
+ case st.MaskedError != "":
+ masked := truncateCNPGRuntimeError(st.MaskedError)
+ if plain, ok := cnpgPostgresSentence(st.MaskedError); ok {
+ masked = plain
+ }
+ return "the instance manager answered while PostgreSQL may be unavailable and masked an error (" + masked + "); what it did not read is unknown"
+ case st.IsPgRewindRunning:
+ return "pg_rewind is running, so the instance manager read nothing from PostgreSQL"
+ }
+ return ""
+}
+
+// markCNPGStatusIncomplete keeps what the report did carry: each list is
+// filled in one step, so a non-empty one is whole, but an empty one may never
+// have been read.
+func markCNPGStatusIncomplete(f *CNPGInstanceStatusFacts, st *cnpgPgStatus) {
+ f.Incomplete, f.MaskedError = true, st.MaskedError
+ if len(f.Replication) == 0 {
+ f.Replication = nil
+ }
+ if len(f.Slots) == 0 {
+ f.Slots = nil
+ }
+ if len(f.BaseBackups) == 0 {
+ f.BaseBackups = nil
+ }
+ if st.LastArchivedWALTime == "" {
+ f.Archiving = nil
+ }
+ if !f.IsPrimary && !f.IsPgRewindRunning {
+ f.RoleDetail = ""
+ }
+}
+
+// cnpgRoleDetail names what an instance is doing from its own report, never
+// from the Cluster's phase. A standby without an active WAL receiver is
+// replaying from the archive or waiting to reconnect; `kubectl cnpg status`
+// calls both "file based".
+func cnpgRoleDetail(f *CNPGInstanceStatusFacts) string {
+ switch {
+ case f.IsPrimary:
+ return "primary"
+ case f.IsPgRewindRunning:
+ return "pgRewind"
+ case f.ReplayPaused:
+ return "replayPaused"
+ case f.IsWalReceiverActive:
+ return "streaming"
+ default:
+ return "fileBased"
+ }
+}
+
+// cnpgStatusTime normalizes an instance manager timestamp; "-infinity" and
+// anything unparseable mean "never" and are omitted.
+func cnpgStatusTime(v string) string {
+ if v == "" || strings.Contains(v, "infinity") {
+ return ""
+ }
+ t, err := time.Parse(time.RFC3339Nano, v)
+ if err != nil {
+ return ""
+ }
+ return t.UTC().Format(time.RFC3339Nano)
+}
+
+// The instance manager has sent syncPriority as both a string and a number.
+func parseCNPGSyncPriority(raw json.RawMessage) *int {
+ if len(raw) == 0 || string(raw) == "null" {
+ return nil
+ }
+ var n int
+ if err := json.Unmarshal(raw, &n); err == nil {
+ return &n
+ }
+ var s string
+ if err := json.Unmarshal(raw, &s); err == nil {
+ if n, err := strconv.Atoi(strings.TrimSpace(s)); err == nil {
+ return &n
+ }
+ }
+ return nil
+}
+
+var cnpgPgIntervalUnitSeconds = map[string]float64{"year": 365.25 * 86400, "mon": 30 * 86400, "day": 86400}
+
+// parsePgIntervalSeconds reads a PostgreSQL interval in the default output
+// style ("00:00:00.012345", "1 day 02:03:04", "2 mons 3 days"). Unparseable
+// is nil, never zero: a lag that cannot be read must not look like no lag.
+func parseCNPGPgInterval(v string) *float64 {
+ tokens := strings.Fields(v)
+ if len(tokens) == 0 {
+ return nil
+ }
+ total := 0.0
+ for i := 0; i < len(tokens); {
+ if secs, ok := parseCNPGPgClock(tokens[i]); ok {
+ total += secs
+ i++
+ continue
+ }
+ if i+1 >= len(tokens) {
+ return nil
+ }
+ n, err := strconv.ParseFloat(tokens[i], 64)
+ if err != nil {
+ return nil
+ }
+ mult, ok := cnpgPgIntervalUnitSeconds[strings.TrimSuffix(tokens[i+1], "s")]
+ if !ok {
+ return nil
+ }
+ total += n * mult
+ i += 2
+ }
+ return &total
+}
+
+func parseCNPGPgClock(tok string) (float64, bool) {
+ neg := strings.HasPrefix(tok, "-")
+ parts := strings.Split(strings.TrimPrefix(tok, "-"), ":")
+ if len(parts) != 3 {
+ return 0, false
+ }
+ h, err1 := strconv.ParseFloat(parts[0], 64)
+ m, err2 := strconv.ParseFloat(parts[1], 64)
+ s, err3 := strconv.ParseFloat(parts[2], 64)
+ if err1 != nil || err2 != nil || err3 != nil {
+ return 0, false
+ }
+ secs := h*3600 + m*60 + s
+ if neg {
+ secs = -secs
+ }
+ return secs, true
+}
+
+// cnpgPromText parses exporter text. An answer cut at the byte cap is read up
+// to its last complete line, and the family on that line is dropped since
+// its samples may continue past the cut; the returned reason says so.
+func cnpgPromText(out cnpgProxyOutcome) (map[string][]cnpgSample, string, error) {
+ body, reason := out.body, ""
+ if out.truncated {
+ cut := bytes.LastIndexByte(body, '\n')
+ if cut < 0 {
+ return nil, "", fmt.Errorf("exporter answer larger than the %s this view reads", formatCNPGByteCap(int64(len(body))))
+ }
+ body = body[:cut+1]
+ reason = fmt.Sprintf("the exporter's answer exceeded %s; measurements past that point are missing", formatCNPGByteCap(int64(len(out.body))))
+ }
+ samples, err := parseCNPGPromSamples(body)
+ if err != nil {
+ return nil, "", err
+ }
+ if out.truncated {
+ delete(samples, cnpgLastPromFamily(body))
+ }
+ return samples, reason, nil
+}
+
+func cnpgLastPromFamily(body []byte) string {
+ lines := bytes.Split(bytes.TrimRight(body, "\n"), []byte("\n"))
+ for i := len(lines) - 1; i >= 0; i-- {
+ line := strings.TrimSpace(string(lines[i]))
+ if line == "" {
+ continue
+ }
+ if strings.HasPrefix(line, "#") {
+ fields := strings.Fields(line)
+ if len(fields) >= 3 {
+ return fields[2]
+ }
+ continue
+ }
+ if i := strings.IndexAny(line, "{ "); i > 0 {
+ return line[:i]
+ }
+ return line
+ }
+ return ""
+}
+
+func cnpgInstanceMetricsFrom(out cnpgProxyOutcome) CNPGInstanceMetrics {
+ src := out.source()
+ if out.state != runtimeStateOK {
+ return CNPGInstanceMetrics{CNPGRuntimeSource: src}
+ }
+ samples, reason, err := cnpgPromText(out)
+ if err != nil {
+ src.State, src.Error = runtimeStateError, "exporter answer was not Prometheus text: "+truncateCNPGRuntimeError(err.Error())
+ return CNPGInstanceMetrics{CNPGRuntimeSource: src}
+ }
+ facts, capped := cnpgInstanceMetricFacts(samples)
+ if reason != "" || capped != "" {
+ src.State, src.Reason = cnpgRuntimeStatePartial, cnpgJoinReasons(reason, capped)
+ }
+ return CNPGInstanceMetrics{CNPGRuntimeSource: src, CNPGInstanceMetricFacts: facts}
+}
+
+func poolerPodFrom(out cnpgProxyOutcome) CNPGPoolerPodRuntime {
+ src := out.source()
+ if out.state != runtimeStateOK {
+ return CNPGPoolerPodRuntime{CNPGRuntimeSource: src}
+ }
+ samples, reason, err := cnpgPromText(out)
+ if err != nil {
+ src.State, src.Error = runtimeStateError, "exporter answer was not Prometheus text: "+truncateCNPGRuntimeError(err.Error())
+ return CNPGPoolerPodRuntime{CNPGRuntimeSource: src}
+ }
+ facts, capped := cnpgPoolerFacts(samples)
+ if reason != "" || capped != "" {
+ src.State, src.Reason = cnpgRuntimeStatePartial, cnpgJoinReasons(reason, capped)
+ }
+ return CNPGPoolerPodRuntime{CNPGRuntimeSource: src, CNPGPoolerPodFacts: facts}
+}
+
+func cnpgJoinReasons(parts ...string) string {
+ var out []string
+ for _, p := range parts {
+ if p != "" {
+ out = append(out, p)
+ }
+ }
+ return strings.Join(out, "; ")
+}
+
+type cnpgSample struct {
+ labels map[string]string
+ value float64
+}
+
+func parseCNPGPromSamples(body []byte) (map[string][]cnpgSample, error) {
+ parser := expfmt.NewTextParser(model.LegacyValidation)
+ families, err := parser.TextToMetricFamilies(bytes.NewReader(body))
+ if err != nil {
+ return nil, err
+ }
+ out := make(map[string][]cnpgSample, len(families))
+ for name, fam := range families {
+ list := []cnpgSample{}
+ for _, m := range fam.Metric {
+ v, ok := cnpgMetricValue(fam.GetType(), m)
+ if !ok {
+ continue
+ }
+ lbls := make(map[string]string, len(m.Label))
+ for _, l := range m.Label {
+ lbls[l.GetName()] = l.GetValue()
+ }
+ list = append(list, cnpgSample{labels: lbls, value: v})
+ }
+ out[name] = list
+ }
+ return out, nil
+}
+
+// cnpgMetricValue drops NaN and infinities: the exporter reports NaN for a
+// value it could not read, which is unavailable, not a number.
+func cnpgMetricValue(t dto.MetricType, m *dto.Metric) (float64, bool) {
+ var v float64
+ switch {
+ case t == dto.MetricType_COUNTER && m.Counter != nil:
+ v = m.Counter.GetValue()
+ case t == dto.MetricType_GAUGE && m.Gauge != nil:
+ v = m.Gauge.GetValue()
+ case m.Untyped != nil:
+ v = m.Untyped.GetValue()
+ default:
+ return 0, false
+ }
+ if math.IsNaN(v) || math.IsInf(v, 0) {
+ return 0, false
+ }
+ return v, true
+}
+
+func cnpgSingle(samples map[string][]cnpgSample, name string, match map[string]string) *float64 {
+ for _, s := range samples[name] {
+ ok := true
+ for k, v := range match {
+ if s.labels[k] != v {
+ ok = false
+ break
+ }
+ }
+ if ok {
+ v := s.value
+ return &v
+ }
+ }
+ return nil
+}
+
+// cnpgSum totals one instance's per-database counter; absent when the family
+// reported nothing.
+func cnpgSum(samples map[string][]cnpgSample, name string) *float64 {
+ list := samples[name]
+ if len(list) == 0 {
+ return nil
+ }
+ total := 0.0
+ for _, s := range list {
+ total += s.value
+ }
+ return &total
+}
+
+func cnpgIsPlatformSession(labels map[string]string) bool {
+ return cnpgPlatformUsers[labels["usename"]] || labels["application_name"] == cnpgMetricsExporterApp
+}
+
+func cnpgMissingFamilies(samples map[string][]cnpgSample, expected []string) []string {
+ var missing []string
+ for _, name := range expected {
+ if _, ok := samples[name]; !ok {
+ missing = append(missing, name)
+ }
+ }
+ return missing
+}
+
+// cnpgInstanceMetricFacts returns the facts and, when a row list was capped,
+// why the answer is partial.
+func cnpgInstanceMetricFacts(samples map[string][]cnpgSample) (*CNPGInstanceMetricFacts, string) {
+ facts := &CNPGInstanceMetricFacts{
+ Missing: cnpgMissingFamilies(samples, cnpgExpectedInstanceFamilies),
+ MaxConnections: cnpgSingle(samples, "cnpg_pg_settings_setting", map[string]string{"name": "max_connections"}),
+ WaitingBackends: cnpgSingle(samples, "cnpg_backends_waiting_total", nil),
+ PostmasterStartTime: cnpgSingle(samples, "cnpg_pg_postmaster_start_time", nil),
+ WalBytes: cnpgSingle(samples, "cnpg_collector_pg_wal", map[string]string{"value": "size"}),
+ WalSegments: cnpgSingle(samples, "cnpg_collector_pg_wal", map[string]string{"value": "count"}),
+ XactCommitTotal: cnpgSum(samples, "cnpg_pg_stat_database_xact_commit"),
+ XactRollbackTotal: cnpgSum(samples, "cnpg_pg_stat_database_xact_rollback"),
+ BlksHit: cnpgSum(samples, "cnpg_pg_stat_database_blks_hit"),
+ BlksRead: cnpgSum(samples, "cnpg_pg_stat_database_blks_read"),
+ DeadlocksTotal: cnpgSum(samples, "cnpg_pg_stat_database_deadlocks"),
+ TempBytesTotal: cnpgSum(samples, "cnpg_pg_stat_database_temp_bytes"),
+ LastUpdateTimestamp: cnpgSingle(samples, "cnpg_last_update_timestamp", nil),
+ Checkpoints: cnpgCheckpointFacts(samples),
+ }
+ var capped []string
+
+ if backends, ok := samples["cnpg_backends_total"]; ok {
+ total := 0.0
+ for _, s := range backends {
+ if cnpgIsPlatformSession(s.labels) || s.labels["state"] == "" || s.value <= 0 {
+ continue
+ }
+ total += s.value
+ if facts.SessionsByState == nil {
+ facts.SessionsByState = map[string]float64{}
+ }
+ facts.SessionsByState[s.labels["state"]] += s.value
+ facts.Sessions = append(facts.Sessions, CNPGSessionGroup{
+ State: s.labels["state"], Database: s.labels["datname"], User: s.labels["usename"],
+ Application: s.labels["application_name"], Count: s.value,
+ })
+ }
+ sort.Slice(facts.Sessions, func(i, j int) bool {
+ a, b := facts.Sessions[i], facts.Sessions[j]
+ if a.Count != b.Count {
+ return a.Count > b.Count
+ }
+ return a.State+"\x00"+a.Database+"\x00"+a.User+"\x00"+a.Application < b.State+"\x00"+b.Database+"\x00"+b.User+"\x00"+b.Application
+ })
+ if len(facts.Sessions) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d session groups; the largest %d are listed and sessionsTotal counts all", len(facts.Sessions), cnpgRuntimeMaxRows))
+ facts.Sessions = facts.Sessions[:cnpgRuntimeMaxRows]
+ }
+ facts.SessionsTotal = &total
+ }
+
+ if durations, ok := samples["cnpg_backends_max_tx_duration_seconds"]; ok {
+ oldest := 0.0
+ for _, s := range durations {
+ if !cnpgIsPlatformSession(s.labels) && s.value > oldest {
+ oldest = s.value
+ }
+ }
+ facts.OldestXactSeconds = &oldest
+ }
+
+ for _, s := range samples["cnpg_pg_database_xid_age"] {
+ facts.XidAge = append(facts.XidAge, CNPGDatabaseValue{Database: s.labels["datname"], Age: s.value})
+ }
+ sort.Slice(facts.XidAge, func(i, j int) bool { return facts.XidAge[i].Age > facts.XidAge[j].Age })
+ if len(facts.XidAge) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d databases; the %d oldest by xid age are listed", len(facts.XidAge), cnpgRuntimeMaxRows))
+ facts.XidAge = facts.XidAge[:cnpgRuntimeMaxRows]
+ }
+ for _, s := range samples["cnpg_pg_database_mxid_age"] {
+ facts.MxidAge = append(facts.MxidAge, CNPGDatabaseValue{Database: s.labels["datname"], Age: s.value})
+ }
+ sort.Slice(facts.MxidAge, func(i, j int) bool { return facts.MxidAge[i].Age > facts.MxidAge[j].Age })
+ if len(facts.MxidAge) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d databases; the %d oldest by multixact age are listed", len(facts.MxidAge), cnpgRuntimeMaxRows))
+ facts.MxidAge = facts.MxidAge[:cnpgRuntimeMaxRows]
+ }
+ if exts, ok := samples["cnpg_pg_extensions_update_available"]; ok {
+ facts.ExtensionUpdates = []CNPGExtensionUpdate{}
+ for _, s := range exts {
+ if s.value <= 0 {
+ continue
+ }
+ facts.ExtensionUpdates = append(facts.ExtensionUpdates, CNPGExtensionUpdate{
+ Database: s.labels["datname"], Extension: s.labels["extname"],
+ InstalledVersion: s.labels["installed_version"], DefaultVersion: s.labels["default_version"],
+ })
+ }
+ sort.Slice(facts.ExtensionUpdates, func(i, j int) bool {
+ a, b := facts.ExtensionUpdates[i], facts.ExtensionUpdates[j]
+ return a.Database+"\x00"+a.Extension < b.Database+"\x00"+b.Extension
+ })
+ if len(facts.ExtensionUpdates) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d extensions with updates; the first %d are listed", len(facts.ExtensionUpdates), cnpgRuntimeMaxRows))
+ facts.ExtensionUpdates = facts.ExtensionUpdates[:cnpgRuntimeMaxRows]
+ }
+ }
+ for _, s := range samples["cnpg_pg_database_size_bytes"] {
+ facts.DatabaseSizes = append(facts.DatabaseSizes, CNPGDatabaseBytes{Database: s.labels["datname"], Bytes: s.value})
+ }
+ sort.Slice(facts.DatabaseSizes, func(i, j int) bool { return facts.DatabaseSizes[i].Bytes > facts.DatabaseSizes[j].Bytes })
+ if len(facts.DatabaseSizes) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d databases; the %d largest are listed", len(facts.DatabaseSizes), cnpgRuntimeMaxRows))
+ facts.DatabaseSizes = facts.DatabaseSizes[:cnpgRuntimeMaxRows]
+ }
+ for _, s := range samples["cnpg_pg_replication_slots_pg_wal_lsn_diff"] {
+ facts.ReplicationSlotsRetainedBytes = append(facts.ReplicationSlotsRetainedBytes, CNPGSlotBytes{Slot: s.labels["slot_name"], Bytes: s.value})
+ }
+ sort.Slice(facts.ReplicationSlotsRetainedBytes, func(i, j int) bool {
+ return facts.ReplicationSlotsRetainedBytes[i].Bytes > facts.ReplicationSlotsRetainedBytes[j].Bytes
+ })
+ if len(facts.ReplicationSlotsRetainedBytes) > cnpgRuntimeMaxRows {
+ capped = append(capped, fmt.Sprintf("%d replication slots; the %d retaining the most WAL are listed", len(facts.ReplicationSlotsRetainedBytes), cnpgRuntimeMaxRows))
+ facts.ReplicationSlotsRetainedBytes = facts.ReplicationSlotsRetainedBytes[:cnpgRuntimeMaxRows]
+ }
+
+ if dbs, more := cnpgDatabaseStats(samples); len(dbs) > 0 {
+ facts.Databases = dbs
+ if more > 0 {
+ capped = append(capped, fmt.Sprintf("%d databases; the first %d by name are listed", len(dbs)+more, cnpgRuntimeMaxRows))
+ }
+ }
+
+ archiver := CNPGArchiverCounters{
+ ArchivedCount: cnpgSingle(samples, "cnpg_pg_stat_archiver_archived_count", nil),
+ FailedCount: cnpgSingle(samples, "cnpg_pg_stat_archiver_failed_count", nil),
+ SecondsSinceLastArchival: cnpgNonNegative(cnpgSingle(samples, "cnpg_pg_stat_archiver_seconds_since_last_archival", nil)),
+ SecondsSinceLastFailure: cnpgNonNegative(cnpgSingle(samples, "cnpg_pg_stat_archiver_seconds_since_last_failure", nil)),
+ }
+ if archiver != (CNPGArchiverCounters{}) {
+ facts.Archiver = &archiver
+ }
+ return facts, strings.Join(capped, "; ")
+}
+
+// The row pg_stat_database keeps for shared objects has no datname; it is not
+// a database.
+func cnpgDatabaseStats(samples map[string][]cnpgSample) ([]CNPGDatabaseStats, int) {
+ byName := map[string]*CNPGDatabaseStats{}
+ set := func(family string, field func(*CNPGDatabaseStats) **float64) {
+ for _, s := range samples[family] {
+ name := s.labels["datname"]
+ if name == "" {
+ continue
+ }
+ d := byName[name]
+ if d == nil {
+ d = &CNPGDatabaseStats{Database: name}
+ byName[name] = d
+ }
+ v := s.value
+ *field(d) = &v
+ }
+ }
+ set("cnpg_pg_stat_database_xact_commit", func(d *CNPGDatabaseStats) **float64 { return &d.XactCommit })
+ set("cnpg_pg_stat_database_xact_rollback", func(d *CNPGDatabaseStats) **float64 { return &d.XactRollback })
+ set("cnpg_pg_stat_database_temp_files", func(d *CNPGDatabaseStats) **float64 { return &d.TempFiles })
+ set("cnpg_pg_stat_database_temp_bytes", func(d *CNPGDatabaseStats) **float64 { return &d.TempBytes })
+ set("cnpg_pg_stat_database_deadlocks", func(d *CNPGDatabaseStats) **float64 { return &d.Deadlocks })
+ set("cnpg_pg_stat_database_blks_hit", func(d *CNPGDatabaseStats) **float64 { return &d.BlksHit })
+ set("cnpg_pg_stat_database_blks_read", func(d *CNPGDatabaseStats) **float64 { return &d.BlksRead })
+ out := make([]CNPGDatabaseStats, 0, len(byName))
+ for _, d := range byName {
+ out = append(out, *d)
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].Database < out[j].Database })
+ if len(out) > cnpgRuntimeMaxRows {
+ return out[:cnpgRuntimeMaxRows], len(out) - cnpgRuntimeMaxRows
+ }
+ return out, 0
+}
+
+func cnpgCheckpointFacts(samples map[string][]cnpgSample) *CNPGCheckpointCounters {
+ get := func(name string) *float64 { return cnpgSingle(samples, name, nil) }
+ if _, ok := samples["cnpg_pg_stat_checkpointer_checkpoints_timed"]; ok {
+ return &CNPGCheckpointCounters{
+ Source: "pg_stat_checkpointer",
+ Timed: get("cnpg_pg_stat_checkpointer_checkpoints_timed"),
+ Requested: get("cnpg_pg_stat_checkpointer_checkpoints_req"),
+ RestartpointsTimed: get("cnpg_pg_stat_checkpointer_restartpoints_timed"),
+ RestartpointsRequested: get("cnpg_pg_stat_checkpointer_restartpoints_req"),
+ RestartpointsDone: get("cnpg_pg_stat_checkpointer_restartpoints_done"),
+ BuffersWritten: get("cnpg_pg_stat_checkpointer_buffers_written"),
+ }
+ }
+ if _, ok := samples["cnpg_pg_stat_bgwriter_checkpoints_timed"]; ok {
+ return &CNPGCheckpointCounters{
+ Source: "pg_stat_bgwriter",
+ Timed: get("cnpg_pg_stat_bgwriter_checkpoints_timed"),
+ Requested: get("cnpg_pg_stat_bgwriter_checkpoints_req"),
+ BuffersWritten: get("cnpg_pg_stat_bgwriter_buffers_checkpoint"),
+ }
+ }
+ return nil
+}
+
+// The exporter reports -1 for "never happened".
+func cnpgNonNegative(v *float64) *float64 {
+ if v == nil || *v < 0 {
+ return nil
+ }
+ return v
+}
+
+// withCNPGSlotRetention returns a copy with each slot's retained WAL from the
+// same instance's exporter. A copy, because the facts may be shared through
+// the memo.
+func withCNPGSlotRetention(facts *CNPGInstanceStatusFacts, retained []CNPGSlotBytes) *CNPGInstanceStatusFacts {
+ if len(retained) == 0 || len(facts.Slots) == 0 {
+ return facts
+ }
+ byName := make(map[string]float64, len(retained))
+ for _, r := range retained {
+ byName[r.Slot] = r.Bytes
+ }
+ out := *facts
+ out.Slots = append([]CNPGSlotStatus(nil), facts.Slots...)
+ for i := range out.Slots {
+ if b, ok := byName[out.Slots[i].Name]; ok {
+ out.Slots[i].RetainedBytes = &b
+ }
+ }
+ return &out
+}
+
+// The exporter encodes pool_mode as 1 session, 2 transaction, 3 statement.
+var cnpgPgBouncerPoolModes = map[int]string{1: "session", 2: "transaction", 3: "statement"}
+
+func cnpgPoolerFacts(samples map[string][]cnpgSample) (*CNPGPoolerPodFacts, string) {
+ type key struct{ db, user string }
+ pools := map[key]*CNPGPoolerPool{}
+ pool := func(labels map[string]string) *CNPGPoolerPool {
+ k := key{labels["database"], labels["user"]}
+ if k.db == cnpgPgBouncerAdminDB || k.user == cnpgPgBouncerAuthUser {
+ return nil
+ }
+ p := pools[k]
+ if p == nil {
+ p = &CNPGPoolerPool{Database: k.db, User: k.user}
+ pools[k] = p
+ }
+ return p
+ }
+ fields := map[string]func(*CNPGPoolerPool) **float64{
+ "cnpg_pgbouncer_pools_cl_active": func(p *CNPGPoolerPool) **float64 { return &p.ClActive },
+ "cnpg_pgbouncer_pools_cl_waiting": func(p *CNPGPoolerPool) **float64 { return &p.ClWaiting },
+ "cnpg_pgbouncer_pools_sv_active": func(p *CNPGPoolerPool) **float64 { return &p.SvActive },
+ "cnpg_pgbouncer_pools_sv_idle": func(p *CNPGPoolerPool) **float64 { return &p.SvIdle },
+ "cnpg_pgbouncer_pools_sv_used": func(p *CNPGPoolerPool) **float64 { return &p.SvUsed },
+ "cnpg_pgbouncer_pools_maxwait": func(p *CNPGPoolerPool) **float64 { return &p.MaxwaitSeconds },
+ }
+ for name, field := range fields {
+ for _, s := range samples[name] {
+ if p := pool(s.labels); p != nil {
+ v := s.value
+ *field(p) = &v
+ }
+ }
+ }
+ // PgBouncer splits maxwait into whole seconds and a microsecond part.
+ for _, s := range samples["cnpg_pgbouncer_pools_maxwait_us"] {
+ if p := pool(s.labels); p != nil && p.MaxwaitSeconds != nil {
+ secs := *p.MaxwaitSeconds + s.value/1e6
+ p.MaxwaitSeconds = &secs
+ }
+ }
+ for _, s := range samples["cnpg_pgbouncer_pools_pool_mode"] {
+ if p := pool(s.labels); p != nil {
+ p.PoolMode = cnpgPgBouncerPoolModes[int(s.value)]
+ }
+ }
+ out := make([]CNPGPoolerPool, 0, len(pools))
+ for _, p := range pools {
+ out = append(out, *p)
+ }
+ sort.Slice(out, func(i, j int) bool {
+ if out[i].Database != out[j].Database {
+ return out[i].Database < out[j].Database
+ }
+ return out[i].User < out[j].User
+ })
+ capped := ""
+ if len(out) > cnpgRuntimeMaxRows {
+ capped = fmt.Sprintf("%d pools; the first %d by database and user are listed", len(out), cnpgRuntimeMaxRows)
+ out = out[:cnpgRuntimeMaxRows]
+ }
+ return &CNPGPoolerPodFacts{Missing: cnpgMissingFamilies(samples, cnpgExpectedPoolerFamilies), Pools: out}, capped
+}
diff --git a/internal/cnpg/runtime_service_test.go b/internal/cnpg/runtime_service_test.go
new file mode 100644
index 0000000000..a661363406
--- /dev/null
+++ b/internal/cnpg/runtime_service_test.go
@@ -0,0 +1,125 @@
+package cnpg
+
+import (
+ "context"
+ "fmt"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "sync/atomic"
+ "testing"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/client-go/kubernetes"
+ k8sfake "k8s.io/client-go/kubernetes/fake"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+func TestClusterRuntimeMemoRequiresGrantsAndIsolatesCallerContextAndPodUID(t *testing.T) {
+ var calls atomic.Int32
+ api := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ calls.Add(1)
+ if strings.HasSuffix(r.URL.Path, "/pg/status") {
+ _, _ = io.WriteString(w, cnpgStatusFixture)
+ return
+ }
+ _, _ = io.WriteString(w, cnpgMetricsFixture)
+ }))
+ defer api.Close()
+ proxy, err := kubernetes.NewForConfig(&rest.Config{Host: api.URL})
+ if err != nil {
+ t.Fatal(err)
+ }
+ cluster := cnpgActionCluster(nil)
+ newCache := func(uid string) *k8s.ResourceCache {
+ pod := cnpgActionPod("pg-1", uid, true)
+ pod.Spec.Containers = []corev1.Container{{Name: "postgres"}}
+ core, err := k8score.NewResourceCache(k8score.CacheConfig{Client: k8sfake.NewSimpleClientset(pod), ResourceTypes: map[string]bool{"pods": true}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ t.Cleanup(core.Stop)
+ return &k8s.ResourceCache{ResourceCache: core}
+ }
+ cache := newCache("original-pod")
+ inventoryAllowed, proxyAllowed := true, true
+ grantChecks := 0
+ type callerKey struct{}
+ ctx := context.WithValue(context.Background(), callerKey{}, "alice")
+ reader := newTestReader(nil)
+ reader.Identity = t.Name() + "/east/alice"
+ reader.Clients.Proxy = proxy
+ reader.Observations.Cluster = func(ctx context.Context, namespace, name string, grants ...auth.Grant) (*k8s.ResourceCache, *unstructured.Unstructured, error) {
+ grantChecks++
+ if ctx.Value(callerKey{}) != "alice" {
+ t.Error("caller context lost before authorization")
+ }
+ if len(grants) != 1 || grants[0] != GrantListPods {
+ t.Errorf("inventory grants=%v", grants)
+ }
+ if !inventoryAllowed {
+ return nil, nil, &ReadFailure{Status: http.StatusForbidden, Message: "no Pod grant"}
+ }
+ return cache, cluster, nil
+ }
+ reader.Access.Permission = func(ctx context.Context, g auth.Grant) string {
+ if ctx.Value(callerKey{}) != "alice" {
+ t.Error("caller context lost before proxy authorization")
+ }
+ if g != grantGetPodsProxy.In("db") {
+ t.Errorf("proxy grant=%v", g)
+ }
+ if !proxyAllowed {
+ return integration.PermissionDenied
+ }
+ return integration.PermissionAllowed
+ }
+ read := func(wantCalls int32) *CNPGClusterRuntimeResponse {
+ t.Helper()
+ got, err := reader.ClusterRuntime(ctx, "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if calls.Load() != wantCalls {
+ t.Fatalf("proxy calls=%d, want %d", calls.Load(), wantCalls)
+ }
+ return got
+ }
+ read(2)
+ read(2)
+ if grantChecks != 2 {
+ t.Fatalf("memo skipped inventory authorization: checks=%d", grantChecks)
+ }
+ proxyAllowed = false
+ denied := read(2)
+ if denied.Instances[0].Status.State != runtimeStateDenied || denied.Instances[0].Status.CNPGInstanceStatusFacts != nil {
+ t.Fatalf("memo disclosed denied status: %+v", denied)
+ }
+ inventoryAllowed = false
+ if _, err := reader.ClusterRuntime(ctx, "db", "pg"); err == nil {
+ t.Fatal("memo bypassed revoked inventory grant")
+ }
+ if calls.Load() != 2 {
+ t.Fatal("denied inventory triggered a proxy read")
+ }
+ inventoryAllowed, proxyAllowed = true, true
+ for i, identity := range []string{"east/bob", "west/alice"} {
+ reader.Identity = t.Name() + "/" + identity
+ read(int32(4 + i*2))
+ }
+ cache = newCache("replacement-pod")
+ got := read(8)
+ if got.Instances[0].PodUID != "replacement-pod" {
+ t.Fatal("replacement instance retained old Pod identity")
+ }
+ if grantChecks < 7 {
+ t.Fatal(fmt.Sprintf("expected authorization before every read, got %d", grantChecks))
+ }
+}
diff --git a/internal/cnpg/runtime_test.go b/internal/cnpg/runtime_test.go
new file mode 100644
index 0000000000..e8f84f8a94
--- /dev/null
+++ b/internal/cnpg/runtime_test.go
@@ -0,0 +1,673 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "net"
+ "net/http"
+ "strings"
+ "sync"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+)
+
+// Modeled on a CloudNativePG 1.27 primary's GET /pg/status.
+const cnpgStatusFixture = `{
+ "currentLsn": "0/7000148", "systemID": "7690724460905209884", "isPrimary": true,
+ "replayPaused": false, "pendingRestart": true, "isWalReceiverActive": false,
+ "pendingRestartForDecrease": true, "isPgRewindRunning": false, "instanceManagerVersion": "1.27.0",
+ "mightBeUnavailable": false, "isArchivingWAL": true,
+ "pod": {"metadata": {"name": "pg-orders-1"}},
+ "lastArchivedWAL": "000000010000000000000006", "lastArchivedWALTime": "2026-09-29T08:24:24.875131Z",
+ "lastFailedWALTime": "-infinity", "currentWAL": "000000010000000000000007", "readyWalFiles": 3,
+ "timeLineID": 1,
+ "replicationInfo": [{
+ "applicationName": "pg-orders-2", "state": "streaming",
+ "receivedLsn": "0/7000148", "writeLsn": "0/7000148", "flushLsn": "0/7000100", "replayLsn": "0/7000000",
+ "writeLag": "00:00:00.012", "flushLag": "1 day 00:00:01", "replayLag": "garbage",
+ "syncState": "async", "syncPriority": "0"
+ }],
+ "replicationSlotsInfo": [{"slotName": "_cnpg_pg_orders_2", "slotType": "physical", "restartLsn": "0/7000148", "walStatus": "reserved", "active": true}, {"slotName": "orders_sub", "plugin": "pgoutput", "slotType": "logical", "database": "app", "walStatus": "extended", "active": false}],
+ "pgStatBasebackupsInfo": [
+ {"usename": "streaming_replica", "application_name": "pg-orders-4-join", "backend_start": "2026-09-30T10:00:00.5Z", "phase": "streaming database files", "backup_total": 4000, "backup_streamed": 1000, "backup_total_pretty": "4000 bytes", "backup_streamed_pretty": "1000 bytes", "tablespaces_total": 1, "tablespaces_streamed": 0},
+ {"usename": "streaming_replica", "application_name": "pg-orders-5-join", "backend_start": "2026-09-30T10:01:00Z", "phase": "waiting for checkpoint to finish", "backup_total": 0, "backup_streamed": 0, "tablespaces_total": 0, "tablespaces_streamed": 0}
+ ]
+}`
+
+// Modeled on the CNPG 1.27 default monitoring queries. The exporter itself
+// connects as postgres with application_name cnpg_metrics_exporter.
+const cnpgMetricsFixture = `# HELP cnpg_backends_total Number of backends
+# TYPE cnpg_backends_total gauge
+cnpg_backends_total{application_name="cnpg_metrics_exporter",datname="app",state="active",usename="postgres"} 1
+cnpg_backends_total{application_name="pg-orders-2",datname="",state="active",usename="streaming_replica"} 1
+cnpg_backends_total{application_name="psql",datname="app",state="idle",usename="app"} 4
+cnpg_backends_total{application_name="api",datname="app",state="idle in transaction",usename="app"} 2
+cnpg_backends_total{application_name="bg",datname="app",state="",usename="app"} 7
+# TYPE cnpg_backends_waiting_total gauge
+cnpg_backends_waiting_total 1
+# TYPE cnpg_pg_postmaster_start_time gauge
+cnpg_pg_postmaster_start_time 1.790698960738028e+09
+# TYPE cnpg_backends_max_tx_duration_seconds gauge
+cnpg_backends_max_tx_duration_seconds{application_name="pg-orders-2",datname="",state="active",usename="streaming_replica"} 9000
+cnpg_backends_max_tx_duration_seconds{application_name="api",datname="app",state="idle in transaction",usename="app"} 42.5
+# TYPE cnpg_pg_database_size_bytes gauge
+cnpg_pg_database_size_bytes{datname="app"} 7.654547e+06
+cnpg_pg_database_size_bytes{datname="postgres"} 7.5e+06
+# TYPE cnpg_pg_database_xid_age gauge
+cnpg_pg_database_xid_age{datname="app"} 29
+# TYPE cnpg_pg_database_mxid_age gauge
+cnpg_pg_database_mxid_age{datname="app"} 12
+cnpg_pg_database_mxid_age{datname="reports"} 400000000
+# TYPE cnpg_pg_extensions_update_available gauge
+cnpg_pg_extensions_update_available{datname="app",default_version="1.0",extname="plpgsql",installed_version="1.0"} 0
+cnpg_pg_extensions_update_available{datname="app",default_version="3.5.1",extname="postgis",installed_version="3.4.2"} 1
+# TYPE cnpg_pg_settings_setting gauge
+cnpg_pg_settings_setting{name="max_connections"} 100
+cnpg_pg_settings_setting{name="shared_buffers"} 16384
+# TYPE cnpg_pg_stat_archiver_archived_count counter
+cnpg_pg_stat_archiver_archived_count 7
+# TYPE cnpg_pg_stat_archiver_failed_count counter
+cnpg_pg_stat_archiver_failed_count 0
+# TYPE cnpg_pg_stat_archiver_seconds_since_last_archival gauge
+cnpg_pg_stat_archiver_seconds_since_last_archival 63.05
+# TYPE cnpg_pg_stat_archiver_seconds_since_last_failure gauge
+cnpg_pg_stat_archiver_seconds_since_last_failure -1
+# TYPE cnpg_collector_pg_wal gauge
+cnpg_collector_pg_wal{value="size"} 1.34217728e+08
+cnpg_collector_pg_wal{value="volume_size"} NaN
+# TYPE cnpg_pg_stat_database_xact_commit counter
+cnpg_pg_stat_database_xact_commit{datname="app"} 100
+cnpg_pg_stat_database_xact_commit{datname="postgres"} 50
+# TYPE cnpg_pg_stat_database_xact_rollback counter
+cnpg_pg_stat_database_xact_rollback{datname="app"} 2
+# TYPE cnpg_pg_stat_database_blks_hit counter
+cnpg_pg_stat_database_blks_hit{datname="app"} 900
+# TYPE cnpg_pg_stat_database_blks_read counter
+cnpg_pg_stat_database_blks_read{datname="app"} 100
+# TYPE cnpg_pg_replication_slots_pg_wal_lsn_diff gauge
+cnpg_pg_replication_slots_pg_wal_lsn_diff{database="",slot_name="_cnpg_pg_orders_2",slot_type="physical"} 16384
+`
+
+const cnpgPoolerMetricsFixture = `# TYPE cnpg_pgbouncer_pools_cl_active gauge
+cnpg_pgbouncer_pools_cl_active{database="pgbouncer",user="pgbouncer"} 1
+cnpg_pgbouncer_pools_cl_active{database="app",user="app"} 5
+cnpg_pgbouncer_pools_cl_active{database="app",user="cnpg_pooler_pgbouncer"} 1
+# TYPE cnpg_pgbouncer_pools_cl_waiting gauge
+cnpg_pgbouncer_pools_cl_waiting{database="app",user="app"} 2
+# TYPE cnpg_pgbouncer_pools_sv_active gauge
+cnpg_pgbouncer_pools_sv_active{database="app",user="app"} 3
+# TYPE cnpg_pgbouncer_pools_sv_idle gauge
+cnpg_pgbouncer_pools_sv_idle{database="app",user="app"} 1
+# TYPE cnpg_pgbouncer_pools_sv_used gauge
+cnpg_pgbouncer_pools_sv_used{database="app",user="app"} 0
+# TYPE cnpg_pgbouncer_pools_maxwait gauge
+cnpg_pgbouncer_pools_maxwait{database="app",user="app"} 1
+# TYPE cnpg_pgbouncer_pools_maxwait_us gauge
+cnpg_pgbouncer_pools_maxwait_us{database="app",user="app"} 250000
+`
+
+func cnpgFptr(v float64) *float64 { return &v }
+
+func cnpgEqF(p *float64, want float64) bool { return p != nil && *p == want }
+
+func TestParseCNPGPgStatus(t *testing.T) {
+ facts, partial, err := parseCNPGPgStatus([]byte(cnpgStatusFixture))
+ if err != nil || partial != "" {
+ t.Fatalf("parse: %v %q", err, partial)
+ }
+ if !facts.IsPrimary || !facts.PendingRestart || facts.CurrentLsn != "0/7000148" || facts.Timeline == nil || *facts.Timeline != 1 {
+ t.Errorf("facts = %+v", facts)
+ }
+ if !facts.PendingRestartForDecrease || facts.InstanceManagerVersion != "1.27.0" || facts.RoleDetail != "primary" {
+ t.Errorf("instance facts = %+v", facts)
+ }
+ a := facts.Archiving
+ if a.LastArchivedWal != "000000010000000000000006" || a.LastArchivedAt == "" || a.LastFailedAt != "" || a.ReadyWalFiles == nil || *a.ReadyWalFiles != 3 {
+ t.Errorf("archiving = %+v, want -infinity omitted", a)
+ }
+ if len(facts.Replication) != 1 {
+ t.Fatalf("replication = %+v", facts.Replication)
+ }
+ rep := facts.Replication[0]
+ if !cnpgEqF(rep.WriteLag, 0.012) || !cnpgEqF(rep.FlushLag, 86401) || rep.ReplayLag != nil || rep.ReplayLagRaw != "garbage" {
+ t.Errorf("lags = %v %v %v raw %q", rep.WriteLag, rep.FlushLag, rep.ReplayLag, rep.ReplayLagRaw)
+ }
+ if rep.SyncPriority == nil || *rep.SyncPriority != 0 || rep.SentLsn != "0/7000148" || rep.FlushLsn != "0/7000100" {
+ t.Errorf("replication = %+v", rep)
+ }
+ if len(facts.Slots) != 2 || facts.Slots[0].Name != "_cnpg_pg_orders_2" || !facts.Slots[0].Active {
+ t.Errorf("slots = %+v", facts.Slots)
+ }
+ if l := facts.Slots[1]; l.Type != "logical" || l.Plugin != "pgoutput" || l.Database != "app" || l.Active {
+ t.Errorf("logical slot = %+v", l)
+ }
+ if len(facts.BaseBackups) != 2 {
+ t.Fatalf("base backups = %+v", facts.BaseBackups)
+ }
+ bb := facts.BaseBackups[0]
+ if bb.Instance != "pg-orders-4" || bb.Phase != "streaming database files" || bb.TotalBytes == nil || *bb.TotalBytes != 4000 || bb.StreamedBytes != 1000 || bb.StartedAt != "2026-09-30T10:00:00.5Z" {
+ t.Errorf("base backup = %+v", bb)
+ }
+ if facts.BaseBackups[1].TotalBytes != nil {
+ t.Errorf("an unestimated total must stay unknown, got %v", *facts.BaseBackups[1].TotalBytes)
+ }
+ none, _, err := parseCNPGPgStatus([]byte(`{"isPrimary": true}`))
+ if err != nil || none.BaseBackups == nil || len(none.BaseBackups) != 0 {
+ t.Errorf("a readable report without base backups must say none, got %+v %v", none, err)
+ }
+
+ for _, bad := range []string{`not json`, `{"pod":{}}`, `proxy error`} {
+ if _, _, err := parseCNPGPgStatus([]byte(bad)); err == nil {
+ t.Errorf("%q parsed as a status report", bad)
+ }
+ }
+}
+
+func TestCNPGSlotInventoryTruncationIsIndependent(t *testing.T) {
+ for _, cappedFamily := range []string{"replicationInfo", "replicationSlotsInfo", "pgStatBasebackupsInfo"} {
+ rows := make([]map[string]any, cnpgRuntimeMaxRows+1)
+ for i := range rows {
+ rows[i] = map[string]any{"slotName": fmt.Sprintf("slot-%d", i)}
+ }
+ body, err := json.Marshal(map[string]any{"isPrimary": true, cappedFamily: rows})
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, partial, err := parseCNPGPgStatus(body)
+ if err != nil || partial == "" || facts.SlotsTruncated != (cappedFamily == "replicationSlotsInfo") {
+ t.Fatalf("%s: facts=%+v partial=%q err=%v", cappedFamily, facts, partial, err)
+ }
+ if cappedFamily == "replicationSlotsInfo" && len(facts.Slots) != cnpgRuntimeMaxRows {
+ t.Fatalf("slot cap: %d", len(facts.Slots))
+ }
+ }
+}
+
+func TestCNPGRoleDetail(t *testing.T) {
+ for want, f := range map[string]CNPGInstanceStatusFacts{
+ "primary": {IsPrimary: true, IsWalReceiverActive: true},
+ "pgRewind": {IsPgRewindRunning: true, IsWalReceiverActive: true},
+ "replayPaused": {ReplayPaused: true, IsWalReceiverActive: true},
+ "streaming": {IsWalReceiverActive: true},
+ "fileBased": {},
+ } {
+ if got := cnpgRoleDetail(&f); got != want {
+ t.Errorf("cnpgRoleDetail(%+v) = %q, want %q", f, got, want)
+ }
+ }
+}
+
+func TestParseCNPGPgInterval(t *testing.T) {
+ for in, want := range map[string]*float64{
+ "00:00:00": cnpgFptr(0),
+ "00:00:01.5": cnpgFptr(1.5),
+ "-00:00:02": cnpgFptr(-2),
+ "1 day 02:03:04": cnpgFptr(86400 + 2*3600 + 3*60 + 4),
+ "2 days": cnpgFptr(172800),
+ "1 mon 3 days 00:00:01": cnpgFptr(30*86400 + 3*86400 + 1),
+ "": nil,
+ "soon": nil,
+ "1 fortnight": nil,
+ } {
+ got := parseCNPGPgInterval(in)
+ if (got == nil) != (want == nil) || (got != nil && *got != *want) {
+ t.Errorf("parseCNPGPgInterval(%q) = %v, want %v", in, got, want)
+ }
+ }
+}
+
+func TestCNPGInstanceMetricFacts(t *testing.T) {
+ samples, err := parseCNPGPromSamples([]byte(cnpgMetricsFixture))
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, capped := cnpgInstanceMetricFacts(samples)
+ if capped != "" {
+ t.Errorf("capped = %q", capped)
+ }
+ // Platform sessions (the exporter, replication) and state '' rows are out.
+ if !cnpgEqF(facts.SessionsTotal, 6) || len(facts.Sessions) != 2 {
+ t.Fatalf("sessions = %+v total %v", facts.Sessions, facts.SessionsTotal)
+ }
+ for _, s := range facts.Sessions {
+ if s.User == "streaming_replica" || s.Application == cnpgMetricsExporterApp || s.State == "" {
+ t.Errorf("platform or stateless row leaked: %+v", s)
+ }
+ }
+ if facts.Sessions[0].State != "idle" || facts.Sessions[0].Count != 4 {
+ t.Errorf("sessions not sorted by count: %+v", facts.Sessions)
+ }
+ if !cnpgEqF(facts.OldestXactSeconds, 42.5) {
+ t.Errorf("oldestXact = %v, want the replication session's 9000 ignored", facts.OldestXactSeconds)
+ }
+ if !cnpgEqF(facts.MaxConnections, 100) || !cnpgEqF(facts.WaitingBackends, 1) || !cnpgEqF(facts.WalBytes, 134217728) {
+ t.Errorf("scalars = %v %v %v", facts.MaxConnections, facts.WaitingBackends, facts.WalBytes)
+ }
+ if !cnpgEqF(facts.XactCommitTotal, 150) || !cnpgEqF(facts.XactRollbackTotal, 2) || !cnpgEqF(facts.BlksHit, 900) || !cnpgEqF(facts.BlksRead, 100) {
+ t.Errorf("counters = %v %v %v %v", facts.XactCommitTotal, facts.XactRollbackTotal, facts.BlksHit, facts.BlksRead)
+ }
+ if !cnpgEqF(facts.PostmasterStartTime, 1.790698960738028e+09) {
+ t.Errorf("postmaster start = %v", facts.PostmasterStartTime)
+ }
+ if facts.DeadlocksTotal != nil {
+ t.Errorf("deadlocks = %v, want absent (family not reported)", *facts.DeadlocksTotal)
+ }
+ if strings.Join(facts.Missing, ",") != "cnpg_pg_stat_database_deadlocks" {
+ t.Errorf("missing = %v", facts.Missing)
+ }
+ a := facts.Archiver
+ if a == nil || !cnpgEqF(a.ArchivedCount, 7) || !cnpgEqF(a.FailedCount, 0) || !cnpgEqF(a.SecondsSinceLastArchival, 63.05) || a.SecondsSinceLastFailure != nil {
+ t.Errorf("archiver = %+v, want -1 (never failed) omitted", a)
+ }
+ if len(facts.DatabaseSizes) != 2 || facts.DatabaseSizes[0].Database != "app" || len(facts.XidAge) != 1 {
+ t.Errorf("databases = %+v xid = %+v", facts.DatabaseSizes, facts.XidAge)
+ }
+ if len(facts.MxidAge) != 2 || facts.MxidAge[0].Database != "reports" || facts.MxidAge[0].Age != 400000000 {
+ t.Errorf("mxid = %+v, want oldest first", facts.MxidAge)
+ }
+ if len(facts.ExtensionUpdates) != 1 || facts.ExtensionUpdates[0] != (CNPGExtensionUpdate{Database: "app", Extension: "postgis", InstalledVersion: "3.4.2", DefaultVersion: "3.5.1"}) {
+ t.Errorf("extension updates = %+v, want only postgis", facts.ExtensionUpdates)
+ }
+ noExt, _ := cnpgInstanceMetricFacts(map[string][]cnpgSample{"cnpg_pg_extensions_update_available": {{labels: map[string]string{"extname": "plpgsql"}, value: 0}}})
+ if noExt.ExtensionUpdates == nil || len(noExt.ExtensionUpdates) != 0 {
+ t.Errorf("exported with none to update must be empty, not unknown: %+v", noExt.ExtensionUpdates)
+ }
+ if unk, _ := cnpgInstanceMetricFacts(map[string][]cnpgSample{}); unk.ExtensionUpdates != nil {
+ t.Errorf("unexported family must stay unknown: %+v", unk.ExtensionUpdates)
+ }
+ if len(facts.ReplicationSlotsRetainedBytes) != 1 || facts.ReplicationSlotsRetainedBytes[0].Bytes != 16384 {
+ t.Errorf("slot retention = %+v", facts.ReplicationSlotsRetainedBytes)
+ }
+
+ // An exporter with the default queries replaced reports nothing it lacks.
+ empty, _ := parseCNPGPromSamples([]byte("# TYPE custom_metric gauge\ncustom_metric 1\n"))
+ bare, _ := cnpgInstanceMetricFacts(empty)
+ b, _ := json.Marshal(bare)
+ for _, field := range []string{"sessionsTotal", "maxConnections", "xactCommitTotal", "archiver", "oldestXactSeconds"} {
+ if strings.Contains(string(b), `"`+field+`"`) {
+ t.Errorf("absent measurement %s serialized: %s", field, b)
+ }
+ }
+ if len(bare.Missing) != len(cnpgExpectedInstanceFamilies) {
+ t.Errorf("missing = %v", bare.Missing)
+ }
+}
+
+func TestCNPGSessionRowsAreCapped(t *testing.T) {
+ var b strings.Builder
+ b.WriteString("# TYPE cnpg_backends_total gauge\n")
+ for i := 0; i < cnpgRuntimeMaxRows+5; i++ {
+ fmt.Fprintf(&b, "cnpg_backends_total{application_name=\"a%d\",datname=\"app\",state=\"idle\",usename=\"app\"} 1\n", i)
+ }
+ samples, err := parseCNPGPromSamples([]byte(b.String()))
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, capped := cnpgInstanceMetricFacts(samples)
+ if capped == "" || len(facts.Sessions) != cnpgRuntimeMaxRows || !cnpgEqF(facts.SessionsTotal, float64(cnpgRuntimeMaxRows+5)) {
+ t.Errorf("capped=%q rows=%d total=%v, want capped rows with the full total", capped, len(facts.Sessions), facts.SessionsTotal)
+ }
+}
+
+func TestCNPGPromTextTruncatedIsPartial(t *testing.T) {
+ cut := strings.Index(cnpgMetricsFixture, "cnpg_pg_stat_database_xact_commit{datname=\"postgres\"}")
+ out := cnpgProxyOutcome{state: runtimeStateOK, body: []byte(cnpgMetricsFixture[:cut+10]), truncated: true}
+ got := cnpgInstanceMetricsFrom(out)
+ if got.State != cnpgRuntimeStatePartial || got.Reason == "" {
+ t.Fatalf("state = %q reason %q", got.State, got.Reason)
+ }
+ // The family on the cut line may continue past the cap, so it must not
+ // read as a complete total.
+ if got.XactCommitTotal != nil {
+ t.Errorf("xactCommitTotal = %v from a family cut at the cap", *got.XactCommitTotal)
+ }
+ if !cnpgEqF(got.MaxConnections, 100) {
+ t.Errorf("families before the cut should survive: %+v", got.CNPGInstanceMetricFacts)
+ }
+}
+
+func TestCNPGPoolerFacts(t *testing.T) {
+ samples, err := parseCNPGPromSamples([]byte(cnpgPoolerMetricsFixture))
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, _ := cnpgPoolerFacts(samples)
+ if len(facts.Pools) != 1 {
+ t.Fatalf("pools = %+v, want the admin and auth_query pools excluded", facts.Pools)
+ }
+ p := facts.Pools[0]
+ if p.Database != "app" || !cnpgEqF(p.ClActive, 5) || !cnpgEqF(p.ClWaiting, 2) || !cnpgEqF(p.SvActive, 3) || !cnpgEqF(p.SvIdle, 1) || !cnpgEqF(p.SvUsed, 0) || !cnpgEqF(p.MaxwaitSeconds, 1.25) {
+ t.Errorf("pool = %+v", p)
+ }
+ if len(facts.Missing) != 0 {
+ t.Errorf("missing = %v", facts.Missing)
+ }
+}
+
+// cnpgFakeProxyAPIServer stands in for the apiserver's pods/proxy. The handler
+// decides per (scheme, pod, port, path); every request is recorded.
+type cnpgFakeProxyAPIServer struct {
+ mu sync.Mutex
+ requests []string
+ headers []http.Header
+}
+
+func (f *cnpgFakeProxyAPIServer) record(r *http.Request) {
+ f.mu.Lock()
+ defer f.mu.Unlock()
+ f.requests = append(f.requests, r.Method+" "+r.URL.Path)
+ f.headers = append(f.headers, r.Header.Clone())
+}
+
+func (f *cnpgFakeProxyAPIServer) seen() []string {
+ f.mu.Lock()
+ defer f.mu.Unlock()
+ return append([]string(nil), f.requests...)
+}
+
+func TestClassifyCNPGProxyErrorSchemeMismatchOnly(t *testing.T) {
+ ctx := context.Background()
+ for _, c := range []struct {
+ msg string
+ mismatch bool
+ }{
+ {"error trying to reach service: tls: first record does not look like a TLS handshake", true},
+ {"error trying to reach service: http: server gave HTTP response to HTTPS client", true},
+ {"error trying to reach service: tls: failed to verify certificate: x509: certificate signed by unknown authority", false},
+ {"error trying to reach service: dial tcp 10.0.0.5:9187: connect: connection refused", false},
+ } {
+ out := classifyCNPGProxyError(ctx, fmt.Errorf("%s", c.msg), cnpgProxyOutcome{})
+ if out.schemeMismatch != c.mismatch {
+ t.Errorf("%q: schemeMismatch = %v, want %v", c.msg, out.schemeMismatch, c.mismatch)
+ }
+ }
+ timeoutCtx, cancel := context.WithDeadline(ctx, time.Now().Add(-time.Second))
+ defer cancel()
+ if out := classifyCNPGProxyError(timeoutCtx, context.DeadlineExceeded, cnpgProxyOutcome{}); out.schemeMismatch || out.state != runtimeStateUnreachable {
+ t.Errorf("timeout = %+v, want unreachable without retry", out)
+ }
+}
+
+func TestParseCNPGPgStatusIncompleteReportIsNotNone(t *testing.T) {
+ masked := `{"isPrimary": false, "mightBeUnavailable": true, "mightBeUnavailableMaskedError": "failed to connect to /controller/run/.s.PGSQL.5432", "instanceManagerVersion": "1.27.0"}`
+ facts, partial, err := parseCNPGPgStatus([]byte(masked))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if !facts.Incomplete || facts.MaskedError == "" || !strings.Contains(partial, "masked an error") {
+ t.Errorf("incomplete not reported: %+v %q", facts, partial)
+ }
+ if facts.Replication != nil || facts.Slots != nil || facts.BaseBackups != nil || facts.Archiving != nil {
+ t.Errorf("unread lists must be unknown, not empty: repl=%v slots=%v bb=%v arch=%v", facts.Replication, facts.Slots, facts.BaseBackups, facts.Archiving)
+ }
+ if facts.RoleDetail != "" {
+ t.Errorf("a standby's role detail is not established from a masked report, got %q", facts.RoleDetail)
+ }
+ body, _ := json.Marshal(facts)
+ if !strings.Contains(string(body), `"baseBackups":null`) || strings.Contains(string(body), `"baseBackups":[]`) {
+ t.Errorf("wire form = %s", body)
+ }
+
+ withRows := `{"isPrimary": true, "mightBeUnavailable": true, "mightBeUnavailableMaskedError": "x", "replicationSlotsInfo": [{"slotName": "s", "slotType": "logical", "active": true}]}`
+ f2, _, _ := parseCNPGPgStatus([]byte(withRows))
+ if len(f2.Slots) != 1 || f2.Replication != nil || f2.RoleDetail != "primary" {
+ t.Errorf("a list the report did fill is whole; the rest unknown: %+v", f2)
+ }
+
+ rewind, partial, _ := parseCNPGPgStatus([]byte(`{"isPrimary": false, "isPgRewindRunning": true}`))
+ if !rewind.Incomplete || rewind.RoleDetail != "pgRewind" || !strings.Contains(partial, "pg_rewind") {
+ t.Errorf("pg_rewind report = %+v %q", rewind, partial)
+ }
+}
+
+func TestCNPGDatabaseStatsAndCheckpoints(t *testing.T) {
+ body := `# TYPE cnpg_pg_stat_database_xact_commit counter
+cnpg_pg_stat_database_xact_commit{datname="app"} 90
+cnpg_pg_stat_database_xact_commit{datname="reports"} 10
+cnpg_pg_stat_database_xact_commit{datname=""} 5
+# TYPE cnpg_pg_stat_database_xact_rollback counter
+cnpg_pg_stat_database_xact_rollback{datname="app"} 10
+# TYPE cnpg_pg_stat_database_temp_files counter
+cnpg_pg_stat_database_temp_files{datname="app"} 3
+# TYPE cnpg_pg_stat_database_temp_bytes counter
+cnpg_pg_stat_database_temp_bytes{datname="app"} 4096
+cnpg_pg_stat_database_temp_bytes{datname="reports"} 1024
+# TYPE cnpg_pg_stat_checkpointer_checkpoints_timed counter
+cnpg_pg_stat_checkpointer_checkpoints_timed 7
+# TYPE cnpg_pg_stat_checkpointer_checkpoints_req counter
+cnpg_pg_stat_checkpointer_checkpoints_req 2
+# TYPE cnpg_pg_stat_checkpointer_restartpoints_timed counter
+cnpg_pg_stat_checkpointer_restartpoints_timed 4
+# TYPE cnpg_pg_stat_checkpointer_restartpoints_done counter
+cnpg_pg_stat_checkpointer_restartpoints_done 3
+# TYPE cnpg_pg_stat_checkpointer_buffers_written counter
+cnpg_pg_stat_checkpointer_buffers_written 1234
+`
+ samples, err := parseCNPGPromSamples([]byte(body))
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, _ := cnpgInstanceMetricFacts(samples)
+ if len(facts.Databases) != 2 || facts.Databases[0].Database != "app" || facts.Databases[1].Database != "reports" {
+ t.Fatalf("databases = %+v, want app and reports without the shared-objects row", facts.Databases)
+ }
+ app := facts.Databases[0]
+ if !cnpgEqF(app.XactCommit, 90) || !cnpgEqF(app.XactRollback, 10) || !cnpgEqF(app.TempFiles, 3) || !cnpgEqF(app.TempBytes, 4096) || app.Deadlocks != nil {
+ t.Errorf("app = %+v", app)
+ }
+ if facts.Databases[1].XactRollback != nil {
+ t.Error("a counter the exporter did not report must stay unknown")
+ }
+ if !cnpgEqF(facts.TempBytesTotal, 5120) {
+ t.Errorf("temp bytes total = %v", facts.TempBytesTotal)
+ }
+ c := facts.Checkpoints
+ if c == nil || c.Source != "pg_stat_checkpointer" || !cnpgEqF(c.Timed, 7) || !cnpgEqF(c.Requested, 2) || !cnpgEqF(c.RestartpointsTimed, 4) || c.RestartpointsRequested != nil || !cnpgEqF(c.BuffersWritten, 1234) {
+ t.Errorf("checkpoints = %+v", c)
+ }
+
+ old, _ := parseCNPGPromSamples([]byte("# TYPE cnpg_pg_stat_bgwriter_checkpoints_timed counter\ncnpg_pg_stat_bgwriter_checkpoints_timed 5\n# TYPE cnpg_pg_stat_bgwriter_buffers_checkpoint counter\ncnpg_pg_stat_bgwriter_buffers_checkpoint 99\n"))
+ of, _ := cnpgInstanceMetricFacts(old)
+ if of.Checkpoints == nil || of.Checkpoints.Source != "pg_stat_bgwriter" || !cnpgEqF(of.Checkpoints.BuffersWritten, 99) || of.Checkpoints.RestartpointsTimed != nil {
+ t.Errorf("pre-17 checkpoints = %+v", of.Checkpoints)
+ }
+ if none, _ := cnpgInstanceMetricFacts(map[string][]cnpgSample{}); none.Checkpoints != nil || none.Databases != nil {
+ t.Error("absent families must stay unknown")
+ }
+}
+
+func TestCNPGMetricsGenerationAndSessionsByState(t *testing.T) {
+ samples, err := parseCNPGPromSamples([]byte(cnpgMetricsFixture + "# TYPE cnpg_last_update_timestamp gauge\ncnpg_last_update_timestamp 1.79e+09\n"))
+ if err != nil {
+ t.Fatal(err)
+ }
+ facts, _ := cnpgInstanceMetricFacts(samples)
+ if !cnpgEqF(facts.LastUpdateTimestamp, 1.79e9) {
+ t.Errorf("generation = %v", facts.LastUpdateTimestamp)
+ }
+ sum := 0.0
+ for _, v := range facts.SessionsByState {
+ sum += v
+ }
+ if facts.SessionsByState["idle"] != 4 || sum != *facts.SessionsTotal {
+ t.Errorf("by state = %v, total %v", facts.SessionsByState, *facts.SessionsTotal)
+ }
+ if old, _ := cnpgInstanceMetricFacts(map[string][]cnpgSample{}); old.LastUpdateTimestamp != nil || old.SessionsByState != nil {
+ t.Error("absent families must stay unknown")
+ }
+}
+
+func TestCNPGMemoizedReadOutlivesTheCallerThatStartedIt(t *testing.T) {
+ target := proxyTarget{namespace: "pg", pod: "pg-1", podUID: types.UID(fmt.Sprint("uid-memo-detach-", time.Now().UnixNano())), port: cnpgStatusPort, path: cnpgStatusPath, scheme: "https"}
+ type key struct{}
+ started := make(chan struct{})
+ release := make(chan struct{})
+ fetch := func(ctx context.Context) string {
+ close(started)
+ <-release
+ if ctx.Err() != nil {
+ return "cancelled"
+ }
+ if ctx.Value(key{}) != "alice" {
+ return "identity lost"
+ }
+ return "ok"
+ }
+ first, cancelFirst := context.WithCancel(context.WithValue(context.Background(), key{}, "alice"))
+ done := make(chan string, 2)
+ go func() { done <- memoized(first, "id-detach", target, time.Minute, fetch) }()
+ <-started
+ // A second caller joins the in-flight read, then the first hangs up.
+ go func() {
+ done <- memoized(context.WithValue(context.Background(), key{}, "alice"), "id-detach", target, time.Minute, func(context.Context) string { return "second fetch" })
+ }()
+ time.Sleep(20 * time.Millisecond)
+ cancelFirst()
+ close(release)
+ for i := 0; i < 2; i++ {
+ if got := <-done; got != "ok" {
+ t.Errorf("caller %d got %q, want the shared read's answer", i, got)
+ }
+ }
+ if got := memoized(context.Background(), "id-detach", target, time.Minute, func(context.Context) string { return "refetched" }); got != "ok" {
+ t.Errorf("memo = %q, want the completed read cached", got)
+ }
+}
+
+func TestCNPGRuntimeRunnerKeepsItsCapWhenTheCallerGoes(t *testing.T) {
+ ctx, cancel := context.WithCancel(context.Background())
+ run := newCNPGRuntimeRunner(ctx)
+ var running, maxRunning, ran int32
+ release := make(chan struct{})
+ admitted := make(chan struct{}, 32)
+ for i := 0; i < 12; i++ {
+ run.do(func(context.Context) {
+ n := atomic.AddInt32(&running, 1)
+ for {
+ m := atomic.LoadInt32(&maxRunning)
+ if n <= m || atomic.CompareAndSwapInt32(&maxRunning, m, n) {
+ break
+ }
+ }
+ atomic.AddInt32(&ran, 1)
+ admitted <- struct{}{}
+ <-release
+ atomic.AddInt32(&running, -1)
+ })
+ }
+ for i := 0; i < cnpgRuntimeConcurrency; i++ {
+ <-admitted
+ }
+ cancel()
+ time.Sleep(20 * time.Millisecond)
+ close(release)
+ run.wait()
+ if maxRunning > int32(cnpgRuntimeConcurrency) {
+ t.Errorf("max concurrent reads = %d, want ≤ %d", maxRunning, cnpgRuntimeConcurrency)
+ }
+ if ran != int32(cnpgRuntimeConcurrency) {
+ t.Errorf("reads run = %d, want only the %d admitted before the caller left", ran, cnpgRuntimeConcurrency)
+ }
+}
+
+func TestCNPGMemoizedDoesNotKeepATimeout(t *testing.T) {
+ target := proxyTarget{namespace: "pg", pod: "pg-1", podUID: types.UID(fmt.Sprint("uid-memo-timeout-", time.Now().UnixNano())), port: cnpgStatusPort, path: cnpgStatusPath, scheme: "https"}
+ calls := 0
+ fetch := func(ctx context.Context) string {
+ calls++
+ cnpgMarkTimedOut(ctx, fmt.Errorf("proxy: %w", context.DeadlineExceeded))
+ return "timed out"
+ }
+ memoized(context.Background(), "id-timeout", target, time.Minute, fetch)
+ memoized(context.Background(), "id-timeout", target, time.Minute, fetch)
+ if calls != 2 {
+ t.Errorf("fetches = %d, want 2: a timed-out read must not be memoized", calls)
+ }
+ ok := proxyTarget{namespace: "pg", pod: "pg-1", podUID: types.UID(fmt.Sprint("uid-memo-ok-", time.Now().UnixNano())), port: cnpgStatusPort, path: cnpgStatusPath, scheme: "https"}
+ n := 0
+ good := func(context.Context) string { n++; return "ok" }
+ memoized(context.Background(), "id-timeout", ok, time.Minute, good)
+ memoized(context.Background(), "id-timeout", ok, time.Minute, good)
+ if n != 1 {
+ t.Errorf("fetches = %d, want 1 for a read that answered", n)
+ }
+}
+
+func TestCNPGRuntimeRunnerRunsNothingForACancelledCaller(t *testing.T) {
+ ctx, cancel := context.WithCancel(context.Background())
+ cancel()
+ run := newCNPGRuntimeRunner(ctx)
+ var ran int32
+ for i := 0; i < 100; i++ {
+ run.do(func(context.Context) { atomic.AddInt32(&ran, 1) })
+ }
+ run.wait()
+ if ran != 0 {
+ t.Errorf("%d reads ran for a caller that had already left, want 0", ran)
+ }
+}
+
+func TestCNPGMarkTimedOutIgnoresTheWordInNames(t *testing.T) {
+ flagFor := func(err error) bool {
+ f := new(atomic.Bool)
+ cnpgMarkTimedOut(context.WithValue(context.Background(), cnpgReadTimedOutKey{}, f), err)
+ return f.Load()
+ }
+ named := apierrors.NewGenericServerResponse(500, "get", schema.GroupResource{Resource: "pods"}, "https:timeout-demo-1:8000", "pq: out of shared memory", 0, true)
+ if flagFor(named) {
+ t.Error("a Pod named timeout-demo-1 was read as a timeout")
+ }
+ if flagFor(apierrors.NewServiceUnavailable("error trying to reach service: dial tcp timeout-demo-1.pg:8000: connect: connection refused")) {
+ t.Error("a relayed refusal to a host named timeout-demo-1 was read as a timeout")
+ }
+ if flagFor(apierrors.NewServiceUnavailable("the server is currently unable to handle the request: i/o timeout")) {
+ t.Error("a 503 that is not the proxy's transport error was read as a timeout")
+ }
+ if flagFor(errors.New(`Get "https://api/namespaces/timeout-demo/pods/timeout-demo-1/proxy": EOF`)) {
+ t.Error("a URL containing \"timeout\" was read as a timeout")
+ }
+ for _, err := range []error{
+ fmt.Errorf("x: %w", context.DeadlineExceeded),
+ apierrors.NewTimeoutError("slow", 0),
+ &net.DNSError{IsTimeout: true},
+ apierrors.NewServiceUnavailable("error trying to reach service: dial tcp 10.0.0.5:8000: i/o timeout"),
+ apierrors.NewServiceUnavailable("error trying to reach service: context deadline exceeded"),
+ } {
+ if !flagFor(err) {
+ t.Errorf("%v was not read as a timeout", err)
+ }
+ }
+}
+
+func TestCNPGPoolerNotStarted(t *testing.T) {
+ for _, phase := range []corev1.PodPhase{corev1.PodFailed, corev1.PodSucceeded} {
+ p := &corev1.Pod{Status: corev1.PodStatus{Phase: phase}}
+ if got := poolerNotStarted(p); got != "PgBouncer is not running (Pod "+string(phase)+")" {
+ t.Errorf("%s: %s", phase, got)
+ }
+ }
+ pending := &corev1.Pod{Status: corev1.PodStatus{Phase: corev1.PodPending}}
+ if got := poolerNotStarted(pending); got != "PgBouncer has not started (Pod Pending)" {
+ t.Fatal(got)
+ }
+ p := &corev1.Pod{Status: corev1.PodStatus{Phase: corev1.PodPending, Conditions: []corev1.PodCondition{{Type: corev1.PodScheduled, Status: corev1.ConditionFalse}}}}
+ if got := poolerNotStarted(p); got != "PgBouncer has not started (Pod cannot be scheduled)" {
+ t.Fatal(got)
+ }
+ p.Status.Phase = corev1.PodRunning
+ if got := poolerNotStarted(p); got != "" {
+ t.Fatal(got)
+ }
+ p.Status.Phase = ""
+ if got := poolerNotStarted(p); got != "" {
+ t.Fatal("unreported phase must stay unknown: " + got)
+ }
+}
diff --git a/internal/cnpg/runtime_types.go b/internal/cnpg/runtime_types.go
new file mode 100644
index 0000000000..77489896b9
--- /dev/null
+++ b/internal/cnpg/runtime_types.go
@@ -0,0 +1,296 @@
+package cnpg
+
+import (
+ "k8s.io/apimachinery/pkg/types"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+)
+
+type CNPGRuntimeObjectRef struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ UID types.UID `json:"uid"`
+}
+
+type CNPGRuntimePermission struct {
+ Proxy string `json:"proxy"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+}
+
+// CNPGRuntimeSource describes one read. State is ok | partial | denied |
+// unreachable | error; Reason explains partial, Error explains the failures.
+// CapturedAt is when the answer was read, which can precede the response's
+// sampledAt by up to the memo lifetime.
+type CNPGRuntimeSource struct {
+ State string `json:"state"`
+ Error string `json:"error,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Scheme string `json:"scheme,omitempty"`
+ CapturedAt string `json:"capturedAt,omitempty"`
+}
+
+// CNPGClusterRuntimeResponse is GET /api/cnpg/clusters/{namespace}/{name}/runtime.
+type CNPGClusterRuntimeResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGRuntimePermission `json:"permission"`
+ Instances []CNPGInstanceRuntime `json:"instances"`
+}
+
+type CNPGInstanceRuntime struct {
+ Pod string `json:"pod"`
+ PodUID types.UID `json:"podUID,omitempty"`
+ Role string `json:"role"`
+ Fenced bool `json:"fenced,omitempty"`
+ Status CNPGInstanceStatus `json:"status"`
+ Metrics CNPGInstanceMetrics `json:"metrics"`
+}
+
+// CNPGInstanceStatus carries facts only when the read succeeded; the embedded
+// pointer keeps every fact out of the JSON otherwise, so an unavailable source
+// can never read as zeros.
+type CNPGInstanceStatus struct {
+ CNPGRuntimeSource
+ *CNPGInstanceStatusFacts
+}
+
+type CNPGInstanceStatusFacts struct {
+ IsPrimary bool `json:"isPrimary"`
+ MightBeUnavailable bool `json:"mightBeUnavailable,omitempty"`
+ CurrentLsn string `json:"currentLsn,omitempty"`
+ ReceivedLsn string `json:"receivedLsn,omitempty"`
+ ReplayLsn string `json:"replayLsn,omitempty"`
+ Timeline *int `json:"timeline,omitempty"`
+ ReplayPaused bool `json:"replayPaused"`
+ PendingRestart bool `json:"pendingRestart"`
+ // PendingRestartForDecrease: the pending change lowers a setting a standby
+ // must match, so the primary restarts before its standbys.
+ PendingRestartForDecrease bool `json:"pendingRestartForDecrease"`
+ IsWalReceiverActive bool `json:"isWalReceiverActive"`
+ IsPgRewindRunning bool `json:"isPgRewindRunning"`
+ InstanceManagerVersion string `json:"instanceManagerVersion,omitempty"`
+ IsInstanceManagerUpgrading bool `json:"isInstanceManagerUpgrading,omitempty"`
+ // RoleDetail is derived only from the instance's own report: primary |
+ // pgRewind | replayPaused | streaming | fileBased.
+ RoleDetail string `json:"roleDetail"`
+ // Archiving is absent when the report did not reach pg_stat_archiver.
+ Archiving *CNPGArchivingStatus `json:"archiving,omitempty"`
+ Replication []CNPGReplicationStatus `json:"replication"`
+ Slots []CNPGSlotStatus `json:"slots"`
+ SlotsTruncated bool `json:"slotsTruncated,omitempty"`
+ // BaseBackups are pg_basebackup streams from this instance to a joining
+ // instance: CloudNativePG's probe reads pg_stat_progress_basebackup only
+ // for application names ending in "-join" (its join Job), so these are new
+ // replicas being cloned, never Backup objects. Empty = none running.
+ BaseBackups []CNPGBaseBackupStatus `json:"baseBackups"`
+ // Incomplete: the instance manager answered without finishing its reads
+ // (it masks errors while PostgreSQL may be unavailable, and reads nothing
+ // while pg_rewind runs). Lists it did not fill are then null, not empty,
+ // and pendingRestart and a standby's roleDetail are not established.
+ Incomplete bool `json:"incomplete,omitempty"`
+ MaskedError string `json:"maskedError,omitempty"`
+}
+
+// CNPGBaseBackupStatus is one pg_stat_progress_basebackup row. TotalBytes is
+// absent when PostgreSQL has no estimate yet (waiting for a checkpoint, or
+// estimation disabled), so progress is then unknown rather than 0 %.
+type CNPGBaseBackupStatus struct {
+ ApplicationName string `json:"applicationName"`
+ Instance string `json:"instance,omitempty"`
+ Phase string `json:"phase"`
+ StartedAt string `json:"startedAt,omitempty"`
+ TotalBytes *int64 `json:"totalBytes,omitempty"`
+ StreamedBytes int64 `json:"streamedBytes"`
+ TablespacesTotal int64 `json:"tablespacesTotal"`
+ TablespacesStreamed int64 `json:"tablespacesStreamed"`
+}
+
+// CNPGArchivingStatus times are RFC3339; the instance manager's "-infinity"
+// (never) is omitted.
+type CNPGArchivingStatus struct {
+ LastArchivedWal string `json:"lastArchivedWal,omitempty"`
+ LastArchivedAt string `json:"lastArchivedAt,omitempty"`
+ LastFailedWal string `json:"lastFailedWal,omitempty"`
+ LastFailedAt string `json:"lastFailedAt,omitempty"`
+ ReadyWalFiles *int `json:"readyWalFiles,omitempty"`
+}
+
+// CNPGReplicationStatus lags are seconds when the PostgreSQL interval parses;
+// the Raw fields always carry what the instance manager said.
+type CNPGReplicationStatus struct {
+ ApplicationName string `json:"applicationName"`
+ State string `json:"state,omitempty"`
+ SyncState string `json:"syncState,omitempty"`
+ SyncPriority *int `json:"syncPriority,omitempty"`
+ WriteLag *float64 `json:"writeLag,omitempty"`
+ WriteLagRaw string `json:"writeLagRaw,omitempty"`
+ FlushLag *float64 `json:"flushLag,omitempty"`
+ FlushLagRaw string `json:"flushLagRaw,omitempty"`
+ ReplayLag *float64 `json:"replayLag,omitempty"`
+ ReplayLagRaw string `json:"replayLagRaw,omitempty"`
+ SentLsn string `json:"sentLsn,omitempty"`
+ WriteLsn string `json:"writeLsn,omitempty"`
+ FlushLsn string `json:"flushLsn,omitempty"`
+ ReplayLsn string `json:"replayLsn,omitempty"`
+}
+
+type CNPGSlotStatus struct {
+ Name string `json:"name"`
+ Type string `json:"type,omitempty"`
+ Plugin string `json:"plugin,omitempty"`
+ Active bool `json:"active"`
+ Database string `json:"database,omitempty"`
+ RestartLsn string `json:"restartLsn,omitempty"`
+ WalStatus string `json:"walStatus,omitempty"`
+ SafeWalSize *int64 `json:"safeWalSize,omitempty"`
+ RetainedBytes *float64 `json:"retainedBytes,omitempty"`
+}
+
+type CNPGInstanceMetrics struct {
+ CNPGRuntimeSource
+ *CNPGInstanceMetricFacts
+}
+
+// CNPGInstanceMetricFacts are this instance's own figures — never summed
+// across instances. A measurement whose family the exporter did not report is
+// absent and its family is listed in Missing. SessionsTotal counts every
+// non-platform session even when Sessions is capped.
+type CNPGInstanceMetricFacts struct {
+ Missing []string `json:"missing,omitempty"`
+ MaxConnections *float64 `json:"maxConnections,omitempty"`
+ // PostmasterStartTime is epoch seconds: an in-place PostgreSQL restart
+ // moves it while the container keeps running.
+ PostmasterStartTime *float64 `json:"postmasterStartTime,omitempty"`
+ Sessions []CNPGSessionGroup `json:"sessions,omitempty"`
+ SessionsTotal *float64 `json:"sessionsTotal,omitempty"`
+ WaitingBackends *float64 `json:"waitingBackends,omitempty"`
+ OldestXactSeconds *float64 `json:"oldestXactSeconds,omitempty"`
+ XidAge []CNPGDatabaseValue `json:"xidAge,omitempty"`
+ MxidAge []CNPGDatabaseValue `json:"mxidAge,omitempty"`
+ // ExtensionUpdates lists installed extensions whose installed_version is
+ // not the default_version; nil when the family was not exported (see
+ // Missing), empty when every installed extension is current.
+ ExtensionUpdates []CNPGExtensionUpdate `json:"extensionUpdates"`
+ DatabaseSizes []CNPGDatabaseBytes `json:"databaseSizes,omitempty"`
+ Archiver *CNPGArchiverCounters `json:"archiver,omitempty"`
+ WalBytes *float64 `json:"walBytes,omitempty"`
+ WalSegments *float64 `json:"walSegments,omitempty"`
+ ReplicationSlotsRetainedBytes []CNPGSlotBytes `json:"replicationSlotsRetainedBytes,omitempty"`
+ XactCommitTotal *float64 `json:"xactCommitTotal,omitempty"`
+ XactRollbackTotal *float64 `json:"xactRollbackTotal,omitempty"`
+ BlksHit *float64 `json:"blksHit,omitempty"`
+ BlksRead *float64 `json:"blksRead,omitempty"`
+ DeadlocksTotal *float64 `json:"deadlocksTotal,omitempty"`
+ TempBytesTotal *float64 `json:"tempBytesTotal,omitempty"`
+ // LastUpdateTimestamp is when the exporter last ran its queries (epoch
+ // seconds, cnpg_last_update_timestamp). From 1.30 query results are cached
+ // for monitoring.metricsQueriesTTL, so counters change only between
+ // generations; absent on operators that do not publish it.
+ LastUpdateTimestamp *float64 `json:"lastUpdateTimestamp,omitempty"`
+ // SessionsByState sums client sessions (platform users excluded) per
+ // pg_stat_activity state over every group, before Sessions is capped.
+ SessionsByState map[string]float64 `json:"sessionsByState,omitempty"`
+ // Databases are pg_stat_database counters per datname (cumulative since
+ // the statistics were last reset), for per-database ratios.
+ Databases []CNPGDatabaseStats `json:"databases,omitempty"`
+ Checkpoints *CNPGCheckpointCounters `json:"checkpoints,omitempty"`
+}
+
+type CNPGDatabaseStats struct {
+ Database string `json:"database"`
+ XactCommit *float64 `json:"xactCommit,omitempty"`
+ XactRollback *float64 `json:"xactRollback,omitempty"`
+ TempFiles *float64 `json:"tempFiles,omitempty"`
+ TempBytes *float64 `json:"tempBytes,omitempty"`
+ Deadlocks *float64 `json:"deadlocks,omitempty"`
+ BlksHit *float64 `json:"blksHit,omitempty"`
+ BlksRead *float64 `json:"blksRead,omitempty"`
+}
+
+// CNPGCheckpointCounters are cumulative counters from pg_stat_checkpointer
+// (PostgreSQL 17+) or pg_stat_bgwriter (before 17), whichever the exporter
+// serves. Restartpoints are counted separately only from 17; before that a
+// standby counts its restartpoints as checkpoints.
+type CNPGCheckpointCounters struct {
+ Source string `json:"source"`
+ Timed *float64 `json:"timed,omitempty"`
+ Requested *float64 `json:"requested,omitempty"`
+ RestartpointsTimed *float64 `json:"restartpointsTimed,omitempty"`
+ RestartpointsRequested *float64 `json:"restartpointsRequested,omitempty"`
+ RestartpointsDone *float64 `json:"restartpointsDone,omitempty"`
+ BuffersWritten *float64 `json:"buffersWritten,omitempty"`
+}
+
+type CNPGSessionGroup struct {
+ State string `json:"state"`
+ Database string `json:"database"`
+ User string `json:"user"`
+ Application string `json:"application"`
+ Count float64 `json:"count"`
+}
+
+type CNPGDatabaseValue struct {
+ Database string `json:"database"`
+ Age float64 `json:"age"`
+}
+
+type CNPGExtensionUpdate struct {
+ Database string `json:"database"`
+ Extension string `json:"extension"`
+ InstalledVersion string `json:"installedVersion"`
+ DefaultVersion string `json:"defaultVersion"`
+}
+
+type CNPGDatabaseBytes struct {
+ Database string `json:"database"`
+ Bytes float64 `json:"bytes"`
+}
+
+type CNPGSlotBytes struct {
+ Slot string `json:"slot"`
+ Bytes float64 `json:"bytes"`
+}
+
+// CNPGArchiverCounters seconds-since fields are absent when the event never
+// happened (the exporter reports -1).
+type CNPGArchiverCounters struct {
+ ArchivedCount *float64 `json:"archivedCount,omitempty"`
+ FailedCount *float64 `json:"failedCount,omitempty"`
+ SecondsSinceLastArchival *float64 `json:"secondsSinceLastArchival,omitempty"`
+ SecondsSinceLastFailure *float64 `json:"secondsSinceLastFailure,omitempty"`
+}
+
+// CNPGPoolerRuntimeResponse is GET /api/cnpg/poolers/{namespace}/{name}/runtime.
+type CNPGPoolerRuntimeResponse struct {
+ Pooler CNPGRuntimeObjectRef `json:"pooler"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGRuntimePermission `json:"permission"`
+ Pods []CNPGPoolerPodRuntime `json:"pods"`
+}
+
+type CNPGPoolerPodRuntime struct {
+ Pod string `json:"pod"`
+ SchedulingReason string `json:"schedulingReason,omitempty"`
+ CNPGRuntimeSource
+ *CNPGPoolerPodFacts
+}
+
+// CNPGPoolerPodFacts excludes PgBouncer's admin pool and the operator's
+// auth_query pool.
+type CNPGPoolerPodFacts struct {
+ Missing []string `json:"missing,omitempty"`
+ Pools []CNPGPoolerPool `json:"pools"`
+}
+
+type CNPGPoolerPool struct {
+ Database string `json:"database"`
+ User string `json:"user"`
+ ClActive *float64 `json:"clActive,omitempty"`
+ ClWaiting *float64 `json:"clWaiting,omitempty"`
+ SvActive *float64 `json:"svActive,omitempty"`
+ SvIdle *float64 `json:"svIdle,omitempty"`
+ SvUsed *float64 `json:"svUsed,omitempty"`
+ MaxwaitSeconds *float64 `json:"maxwaitSeconds,omitempty"`
+ // PoolMode is the mode PgBouncer reports it is using for this pool.
+ PoolMode string `json:"poolMode,omitempty"`
+}
diff --git a/internal/cnpg/schedule.go b/internal/cnpg/schedule.go
new file mode 100644
index 0000000000..ff7d161ed9
--- /dev/null
+++ b/internal/cnpg/schedule.go
@@ -0,0 +1,191 @@
+package cnpg
+
+import (
+ "context"
+ "fmt"
+ "slices"
+ "strings"
+ "time"
+ _ "time/tzdata"
+
+ appsv1 "k8s.io/api/apps/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/client-go/dynamic"
+
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+const (
+ cnpgSchedulePreviewRuns = 3
+ scheduleMaxLen = 256
+)
+
+// CNPGSchedulePreview explains a ScheduledBackup schedule and when the
+// operator would run it. Runs are computed as the operator does: from
+// status.lastCheckTime when it has one (so a schedule whose next time since
+// that check has passed runs as soon as the operator sees it), otherwise from
+// now. Times serialize as UTC. Clock records a readable TZ declaration or
+// explicitly marks the UTC estimate when the operator clock is unknown.
+type CNPGSchedulePreview struct {
+ Schedule string `json:"schedule"`
+ Valid bool `json:"valid"`
+ Error string `json:"error,omitempty"`
+ Description string `json:"description,omitempty"`
+ NextRuns []string `json:"nextRuns,omitempty"`
+ // RunsImmediately: the first run is due already, so the operator creates
+ // a backup as soon as it reconciles this schedule.
+ RunsImmediately bool `json:"runsImmediately,omitempty"`
+ // Basis is what the first run is counted from: lastCheckTime | now.
+ Basis string `json:"basis"`
+ LastCheckTime string `json:"lastCheckTime,omitempty"`
+ Suspended bool `json:"suspended,omitempty"`
+ Clock *CNPGScheduleClock `json:"clock,omitempty"`
+}
+
+func schedulePreview(spec string, lastCheck *time.Time, suspended bool, now time.Time) CNPGSchedulePreview {
+ return schedulePreviewIn(spec, lastCheck, suspended, now, time.UTC)
+}
+
+type CNPGScheduleClock struct {
+ Zone string `json:"zone"`
+ Declared bool `json:"declared"`
+ Source string `json:"source"`
+}
+
+func schedulePreviewIn(spec string, lastCheck *time.Time, suspended bool, now time.Time, location *time.Location) CNPGSchedulePreview {
+ p := CNPGSchedulePreview{Schedule: spec, Basis: "now", Suspended: suspended}
+ if lastCheck != nil {
+ p.Basis, p.LastCheckTime = "lastCheckTime", lastCheck.UTC().Format(time.RFC3339)
+ }
+ if strings.TrimSpace(spec) == "" {
+ p.Error = "the schedule is empty"
+ return p
+ }
+ if len(spec) > scheduleMaxLen {
+ p.Error = fmt.Sprintf("the schedule is longer than %d characters", scheduleMaxLen)
+ return p
+ }
+ sched, err := cnpg.ParseSchedule(spec)
+ if err != nil {
+ p.Error = err.Error()
+ return p
+ }
+ now = now.In(location)
+ from := now
+ if lastCheck != nil {
+ from = lastCheck.In(location)
+ }
+ next := sched.Next(from)
+ if next.IsZero() || sched.Next(now).IsZero() {
+ p.Error = "no time satisfies this schedule: the operator would never run it"
+ return p
+ }
+ p.Valid = true
+ p.Description = strings.ReplaceAll(describeCNPGSchedule(spec), "operator clock", location.String())
+ if !next.After(now) {
+ p.RunsImmediately = !suspended
+ next = now
+ p.NextRuns = append(p.NextRuns, next.UTC().Format(time.RFC3339))
+ next = sched.Next(now)
+ }
+ for len(p.NextRuns) < cnpgSchedulePreviewRuns && !next.IsZero() {
+ p.NextRuns = append(p.NextRuns, next.UTC().Format(time.RFC3339))
+ next = sched.Next(next)
+ }
+ return p
+}
+
+func scheduleLastCheck(sched *unstructured.Unstructured) *time.Time {
+ raw, _, _ := unstructured.NestedString(sched.Object, "status", "lastCheckTime")
+ t, err := time.Parse(time.RFC3339, raw)
+ if err != nil {
+ return nil
+ }
+ return &t
+}
+
+// describeCNPGSchedule words an already-validated schedule; see
+// cnpg.DescribeSchedule, shared with the Issues engine's own messages.
+func describeCNPGSchedule(spec string) string { return cnpg.DescribeSchedule(spec) }
+
+func (s *Reader) PreviewSchedule(ctx context.Context, dyn dynamic.Interface, namespace, name, spec string, now time.Time) (CNPGSchedulePreview, error) {
+ sched, err := dyn.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return CNPGSchedulePreview{}, err
+ }
+ suspended, _, _ := unstructured.NestedBool(sched.Object, "spec", "suspend")
+ return s.previewOnOperatorClock(ctx, namespace, spec, scheduleLastCheck(sched), suspended, now), nil
+}
+
+func (s *Reader) PreviewDraftSchedule(ctx context.Context, dyn dynamic.Interface, namespace, name, spec string, now time.Time) (CNPGSchedulePreview, error) {
+ if _, err := dyn.Resource(ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{}); err != nil {
+ return CNPGSchedulePreview{}, err
+ }
+ return s.previewOnOperatorClock(ctx, namespace, spec, nil, false, now), nil
+}
+
+func scheduleClockOf(deployments []*appsv1.Deployment, full bool, namespace string) (CNPGScheduleClock, *time.Location) {
+ unknown := CNPGScheduleClock{Zone: "UTC", Source: "Operator clock is not established; upcoming times assume UTC."}
+ if !full {
+ return unknown, time.UTC
+ }
+ var watching []*appsv1.Deployment
+ for _, d := range deployments {
+ if d.Labels[cnpgOperatorNameLabel] != cnpgOperatorNameValue {
+ continue
+ }
+ c := cnpgOperatorContainerOf(d)
+ if c == nil || len(c.EnvFrom) > 0 {
+ return unknown, time.UTC
+ }
+ watch := cnpgOperatorWatchOf(c, d.Namespace)
+ if watch.Unresolved != "" {
+ return unknown, time.UTC
+ }
+ if watch.All || slices.Contains(watch.Namespaces, namespace) {
+ watching = append(watching, d)
+ }
+ }
+ if len(watching) != 1 {
+ return unknown, time.UTC
+ }
+ d := watching[0]
+ c := cnpgOperatorContainerOf(d)
+ tzCount := 0
+ for _, env := range c.Env {
+ if env.Name == "TZ" {
+ tzCount++
+ }
+ }
+ if tzCount > 1 {
+ return unknown, time.UTC
+ }
+ for _, env := range c.Env {
+ if env.Name != "TZ" {
+ continue
+ }
+ if env.ValueFrom != nil || env.Value == "" {
+ return unknown, time.UTC
+ }
+ zone, err := time.LoadLocation(env.Value)
+ if err != nil {
+ return unknown, time.UTC
+ }
+ return CNPGScheduleClock{Zone: env.Value, Declared: true, Source: "TZ declared on Deployment " + d.Namespace + "/" + d.Name + "; times use that operator clock."}, zone
+ }
+ unknown.Source = "Deployment " + d.Namespace + "/" + d.Name + " declares no TZ. Upcoming times assume UTC; the running operator clock is not verified."
+ return unknown, time.UTC
+}
+
+func (s *Reader) previewOnOperatorClock(ctx context.Context, namespace, spec string, last *time.Time, suspended bool, now time.Time) CNPGSchedulePreview {
+ clock, zone := scheduleClockOf(nil, false, namespace)
+ if s.Observations.Cache != nil {
+ access, deployments := s.operatorDeployments(ctx, s.Observations.Cache, s.Observations.OperatorScope(ctx))
+ clock, zone = scheduleClockOf(deployments, access.State == integration.KindCoverageFull, namespace)
+ }
+ out := schedulePreviewIn(spec, last, suspended, now, zone)
+ out.Clock = &clock
+ return out
+}
diff --git a/internal/cnpg/schedule_repair.go b/internal/cnpg/schedule_repair.go
new file mode 100644
index 0000000000..d558eba87b
--- /dev/null
+++ b/internal/cnpg/schedule_repair.go
@@ -0,0 +1,138 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "net/http"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/integration"
+ declarations "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+type ScheduleMethodFacts struct {
+ ClusterUID string `json:"clusterUID"`
+ ClusterConfig string `json:"clusterConfig"`
+ ScheduleConfig string `json:"scheduleConfig"`
+}
+
+type ScheduleMethodPreview struct {
+ Context string `json:"context"`
+ UID string `json:"uid"`
+ Cluster string `json:"cluster"`
+ PreviousMethod string `json:"previousMethod"`
+ Facts ScheduleMethodFacts `json:"facts"`
+ Unchanged bool `json:"unchanged"`
+}
+
+func prepareScheduleMethod(ctx context.Context, c ActionClients, namespace, name string) (*unstructured.Unstructured, map[string]any, ScheduleMethodPreview, error) {
+ var out ScheduleMethodPreview
+ schedule, err := c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, out, err
+ }
+ if !schedule.GetDeletionTimestamp().IsZero() {
+ return nil, nil, out, integration.BlockedAction("The schedule is being deleted")
+ }
+ spec, _, _ := unstructured.NestedMap(schedule.Object, "spec")
+ clusterName, _, _ := unstructured.NestedString(schedule.Object, "spec", "cluster", "name")
+ cluster, err := c.Dynamic.Resource(ClusterGVR).Namespace(namespace).Get(ctx, clusterName, metav1.GetOptions{})
+ if err != nil {
+ return nil, nil, out, err
+ }
+ if !cluster.GetDeletionTimestamp().IsZero() {
+ return nil, nil, out, integration.BlockedAction("The Cluster is being deleted")
+ }
+ plugin, ok := declarations.ParseBackupDeclaration(cluster).BarmanPlugin()
+ if !ok || !plugin.WALArchiver || plugin.ObjectStore == "" {
+ return nil, nil, out, integration.BlockedAction("The Cluster needs an enabled Barman WAL archiver and ObjectStore before matching this schedule")
+ }
+ store, _, err := readStore(ctx, c.Dynamic, namespace, plugin.ObjectStore)
+ if err != nil {
+ return nil, nil, out, err
+ }
+ if !store.GetDeletionTimestamp().IsZero() {
+ return nil, nil, out, integration.BlockedAction("The ObjectStore is being deleted")
+ }
+ method, _ := spec["method"].(string)
+ if method == "volumeSnapshot" {
+ return nil, nil, out, integration.BlockedAction("This is a snapshot schedule. Review a backup-method migration in its YAML")
+ }
+ configuration, _ := spec["pluginConfiguration"].(map[string]any)
+ if configuration == nil {
+ configuration = map[string]any{}
+ }
+ if declared, _ := configuration["name"].(string); declared != "" && declared != declarations.BarmanPluginName {
+ return nil, nil, out, integration.BlockedAction("This schedule names another plugin. Review a backup-method migration in its YAML")
+ }
+ clusterHash, err := configDigest(protectionConfig(cluster))
+ if err != nil {
+ return nil, nil, out, err
+ }
+ scheduleHash, err := configDigest(map[string]any{"cluster": spec["cluster"], "method": spec["method"], "pluginConfiguration": spec["pluginConfiguration"]})
+ if err != nil {
+ return nil, nil, out, err
+ }
+ unchanged := method == "plugin" && configuration["name"] == declarations.BarmanPluginName
+ configuration["name"] = declarations.BarmanPluginName
+ patch := map[string]any{"spec": map[string]any{"method": "plugin", "pluginConfiguration": configuration}}
+ if method == "" {
+ method = "barmanObjectStore"
+ }
+ out = ScheduleMethodPreview{UID: string(schedule.GetUID()), Cluster: clusterName, PreviousMethod: method, Unchanged: unchanged, Facts: ScheduleMethodFacts{ClusterUID: string(cluster.GetUID()), ClusterConfig: clusterHash, ScheduleConfig: scheduleHash}}
+ return schedule, patch, out, nil
+}
+
+func (s *Reader) PreviewScheduleMethod(ctx context.Context, c ActionClients, contextName, namespace, name string) (*ScheduleMethodPreview, error) {
+ schedule, patch, out, err := prepareScheduleMethod(ctx, c, namespace, name)
+ if err != nil {
+ return nil, err
+ }
+ if !out.Unchanged {
+ patch["metadata"] = map[string]any{"resourceVersion": schedule.GetResourceVersion()}
+ data, err := json.Marshal(patch)
+ if err != nil {
+ return nil, err
+ }
+ _, err = c.Dynamic.Resource(ScheduleGVR).Namespace(namespace).Patch(ctx, name, types.MergePatchType, data, metav1.PatchOptions{DryRun: []string{metav1.DryRunAll}})
+ if err != nil {
+ return nil, err
+ }
+ }
+ out.Context = contextName
+ return &out, nil
+}
+
+func runRepairScheduleMethod(ctx context.Context, c ActionClients, namespace, name string, req integration.ActionRequest) (*CNPGActionResult, error) {
+ if err := integration.DecodeActionParams(req.Params, &struct{}{}); err != nil {
+ return nil, err
+ }
+ var reviewed ScheduleMethodFacts
+ if err := integration.DecodeActionParams(req.Facts, &reviewed); err != nil {
+ return nil, err
+ }
+ if reviewed.ClusterUID == "" || reviewed.ClusterConfig == "" || reviewed.ScheduleConfig == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "Reviewed schedule and Cluster configuration are required")
+ }
+ schedule, patch, current, err := prepareScheduleMethod(ctx, c, namespace, name)
+ if err != nil {
+ return nil, err
+ }
+ if string(schedule.GetUID()) != req.UID || reviewed != current.Facts {
+ return nil, integration.ChangedAction(current, "The schedule or Cluster backup configuration changed since review; review the repair again")
+ }
+ if current.Unchanged {
+ return nil, integration.BlockedAction("The schedule already uses the Cluster's Barman plugin")
+ }
+ if err := integration.MergePatchAtVersion(ctx, c.Dynamic, ScheduleGVR, schedule, patch); err != nil {
+ if apierrors.IsConflict(err) {
+ return nil, integration.ChangedAction(current, "The schedule changed while the patch was sent; review it again")
+ }
+ return nil, err
+ }
+ return &CNPGActionResult{Action: "repairMethod", Message: "Schedule now uses the Cluster's Barman plugin; verify its next backup"}, nil
+}
diff --git a/internal/cnpg/schedule_test.go b/internal/cnpg/schedule_test.go
new file mode 100644
index 0000000000..6af62875c7
--- /dev/null
+++ b/internal/cnpg/schedule_test.go
@@ -0,0 +1,265 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "strings"
+ "testing"
+ "time"
+
+ appsv1 "k8s.io/api/apps/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+)
+
+func TestScheduleClockEvidence(t *testing.T) {
+ operator := &appsv1.Deployment{ObjectMeta: metav1.ObjectMeta{Name: "operator", Namespace: "system", Labels: map[string]string{cnpgOperatorNameLabel: cnpgOperatorNameValue}}, Spec: appsv1.DeploymentSpec{Template: corev1.PodTemplateSpec{Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: cnpgOperatorContainer, Env: []corev1.EnvVar{{Name: "TZ", Value: "America/New_York"}}}}}}}}
+ clock, zone := scheduleClockOf([]*appsv1.Deployment{operator}, true, "db")
+ if !clock.Declared || clock.Zone != "America/New_York" {
+ t.Fatalf("clock=%+v", clock)
+ }
+ p := schedulePreviewIn("0 30 2 * * *", nil, false, time.Date(2026, 10, 1, 0, 0, 0, 0, time.UTC), zone)
+ if !p.Valid || p.NextRuns[0] != "2026-10-01T06:30:00Z" || strings.Contains(p.Description, "UTC") {
+ t.Fatalf("preview=%+v", p)
+ }
+ for _, scenario := range []string{"partial", "multiple", "valueFrom", "absent", "envFrom", "watch-unresolved", "duplicate-tz"} {
+ t.Run(scenario, func(t *testing.T) {
+ d := operator.DeepCopy()
+ deployments := []*appsv1.Deployment{d}
+ full := true
+ switch scenario {
+ case "partial":
+ full = false
+ case "multiple":
+ deployments = append(deployments, operator.DeepCopy())
+ case "valueFrom":
+ d.Spec.Template.Spec.Containers[0].Env[0] = corev1.EnvVar{Name: "TZ", ValueFrom: &corev1.EnvVarSource{SecretKeyRef: &corev1.SecretKeySelector{}}}
+ case "absent":
+ d.Spec.Template.Spec.Containers[0].Env = nil
+ case "duplicate-tz":
+ d.Spec.Template.Spec.Containers[0].Env = append(d.Spec.Template.Spec.Containers[0].Env, corev1.EnvVar{Name: "TZ", Value: "UTC"})
+ case "envFrom":
+ d.Spec.Template.Spec.Containers[0].EnvFrom = []corev1.EnvFromSource{{SecretRef: &corev1.SecretEnvSource{}}}
+ case "watch-unresolved":
+ d.Spec.Template.Spec.Containers[0].Env = append(d.Spec.Template.Spec.Containers[0].Env, corev1.EnvVar{Name: cnpgWatchNamespaceEnv, ValueFrom: &corev1.EnvVarSource{SecretKeyRef: &corev1.SecretKeySelector{}}})
+ }
+ clock, zone := scheduleClockOf(deployments, full, "db")
+ if clock.Declared || zone != time.UTC || !strings.Contains(clock.Source, "assume") {
+ t.Fatalf("clock=%+v zone=%s", clock, zone)
+ }
+ })
+ }
+}
+
+func TestCNPGSchedulePreview(t *testing.T) {
+ now := time.Date(2026, 9, 30, 10, 0, 0, 0, time.UTC)
+ t.Run("from now without a last check", func(t *testing.T) {
+ p := schedulePreview("0 30 2 * * *", nil, false, now)
+ if !p.Valid || p.Basis != "now" || p.RunsImmediately {
+ t.Fatalf("preview = %+v", p)
+ }
+ want := []string{"2026-10-01T02:30:00Z", "2026-10-02T02:30:00Z", "2026-10-03T02:30:00Z"}
+ if strings.Join(p.NextRuns, ",") != strings.Join(want, ",") {
+ t.Errorf("runs = %v, want %v", p.NextRuns, want)
+ }
+ if p.Description != "every day at 02:30 UTC" {
+ t.Errorf("description = %q", p.Description)
+ }
+ })
+ t.Run("counted from lastCheckTime, as the operator does", func(t *testing.T) {
+ last := time.Date(2026, 9, 30, 1, 0, 0, 0, time.UTC)
+ p := schedulePreview("0 0 6 * * *", &last, false, now)
+ if !p.RunsImmediately || p.Basis != "lastCheckTime" || p.NextRuns[0] != "2026-09-30T10:00:00Z" || p.NextRuns[1] != "2026-10-01T06:00:00Z" {
+ t.Errorf("a due first run must be immediate: %+v", p)
+ }
+ if s := schedulePreview("0 0 6 * * *", &last, true, now); s.RunsImmediately {
+ t.Error("a suspended schedule runs nothing now")
+ }
+ if f := schedulePreview("0 0 12 * * *", &last, false, now); f.RunsImmediately || f.NextRuns[0] != "2026-09-30T12:00:00Z" {
+ t.Errorf("future first run = %+v", f)
+ }
+ })
+ t.Run("stepped day-of-month with a day of week follows robfig/cron v1, as the operator does", func(t *testing.T) {
+ p := schedulePreview("0 0 0 */2 * 1", nil, false, now)
+ if !p.Valid || p.NextRuns[0] != "2026-10-05T00:00:00Z" {
+ t.Errorf("first run = %v, want 2026-10-05 (v1 semantics; v3 would say 2026-10-01)", p.NextRuns)
+ }
+ })
+ for _, bad := range []string{"", "0 0 * * * * *", "not cron", "0 0 0 31 2 *", "CRON_TZ=Europe/Berlin 0 0 0 * * *", "TZ=UTC 0 0 0 * * *", "0 0 0 * * 7", strings.Repeat("1", 300)} {
+ if p := schedulePreview(bad, nil, false, now); p.Valid || p.Error == "" || len(p.NextRuns) != 0 {
+ t.Errorf("%q accepted: %+v", bad, p)
+ }
+ }
+ if p := schedulePreview("0 0 0 * *", nil, false, now); !p.Valid {
+ t.Errorf("five fields (day of week optional) rejected: %s", p.Error)
+ }
+}
+
+func TestDescribeCNPGSchedule(t *testing.T) {
+ for spec, want := range map[string]string{
+ "0 0 0 * * *": "every day at 00:00 operator clock",
+ "0 0 2 * * *": "every day at 02:00 operator clock",
+ "@daily": "every day at 00:00 operator clock",
+ "@hourly": "every hour, on the hour",
+ "0 0 * * * *": "every hour, on the hour",
+ "0 15 * * * *": "every hour at :15",
+ "@every 1h30m": "every 1h30m, counted from the operator's last check",
+ "0 30 2 * * 1-5": "every Monday through Friday at 02:30 operator clock",
+ "0 15 3 * * 1-5": "every Monday through Friday at 03:15 operator clock",
+ "0 0 1 * * sun,wed": "every Sunday and Wednesday at 01:00 operator clock",
+ "0 0 4 1,15 * *": "on day 1 and 15 of the month at 04:00 operator clock",
+ "0 0 4 1 jan,7 *": "on day 1 of the month in January and July at 04:00 operator clock",
+ "15 30 4 * * *": "every day at 04:30:15 operator clock",
+ "0 */15 * * * *": "every 15 minutes",
+ "30 0 9-17 * * *": "every hour at :00:30, during hours 9 through 17",
+ "0 0 */6 * * 1": "every Monday, every 6 hours, on the hour",
+ "0 0 0 */2 * 1": "every 2 days of the month from day 1, when it is a Monday at 00:00 operator clock",
+ "0 0 0 1,15 * 1": "on day 1 and 15 of the month or every Monday at 00:00 operator clock",
+ } {
+ if got := describeCNPGSchedule(spec); got != want {
+ t.Errorf("describe(%q) = %q, want %q", spec, got, want)
+ }
+ }
+}
+
+func TestCNPGActionSetSchedule(t *testing.T) {
+ req := func(reviewed, next string) integration.ActionRequest {
+ facts, _ := json.Marshal(map[string]any{"schedule": reviewed})
+ params, _ := json.Marshal(map[string]any{"schedule": next})
+ return integration.ActionRequest{ReviewedContext: "kind-test", UID: "sched-uid", Facts: facts, Params: params}
+ }
+ t.Run("merge-patches spec.schedule bound to the resourceVersion", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ res, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "setSchedule", req("0 0 0 * * *", "0 30 2 * * *"))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if res.Action != "setSchedule" || len(env.patches) != 1 {
+ t.Fatalf("res = %+v patches = %d", res, len(env.patches))
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ if body["spec"].(map[string]any)["schedule"] != "0 30 2 * * *" || body["metadata"].(map[string]any)["resourceVersion"] != "7" || len(body["spec"].(map[string]any)) != 1 {
+ t.Errorf("patch = %v", body)
+ }
+ })
+ t.Run("refuses a schedule changed since review", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "setSchedule", req("0 0 1 * * *", "0 30 2 * * *"))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged || len(env.patches) != 0 {
+ t.Fatalf("err = %v, want 409 changed and no write", err)
+ }
+ })
+ t.Run("rejects an invalid schedule before any read or write", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "setSchedule", req("0 0 0 * * *", "0 0 25 * * *"))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != 400 || ae.Code != cnpgCodeInvalidSchedule || len(env.patches) != 0 {
+ t.Fatalf("err = %v, want 400 invalid_schedule", err)
+ }
+ })
+ t.Run("requires the reviewed schedule", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ r := req("", "0 30 2 * * *")
+ r.Facts = json.RawMessage(`{}`)
+ if ae, ok := cnpgActionStatus(t, func() error {
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "setSchedule", r)
+ return err
+ }()); !ok || ae.Status != 400 {
+ t.Fatalf("missing facts.schedule: %v", ae)
+ }
+ })
+ t.Run("an unchanged schedule is blocked", func(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), cnpgActionSchedule(nil)})
+ _, err := RunCNPGScheduleAction(context.Background(), env.clients(), "db", "nightly", "setSchedule", req("0 0 0 * * *", "0 0 0 * * *"))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked {
+ t.Fatalf("err = %v, want blocked", err)
+ }
+ })
+}
+
+func TestCNPGScheduleCapabilitiesCarryScheduleAndPreview(t *testing.T) {
+ sched := cnpgActionSchedule(func(o map[string]any) {
+ o["status"].(map[string]any)["lastCheckTime"] = "2026-09-28T00:00:00Z"
+ })
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil), sched})
+ resp, err := newTestReader(nil).ScheduleCapabilities(context.Background(), env.clients(), "kind-test", "db", "nightly")
+ if err != nil {
+ t.Fatal(err)
+ }
+ f := resp.Facts
+ if f.Schedule != "0 0 0 * * *" || !f.Preview.Valid || f.Preview.Basis != "lastCheckTime" || f.Preview.Description != "every day at 00:00 UTC" || len(f.Preview.NextRuns) != 3 {
+ t.Errorf("facts = %+v", f)
+ }
+ if !f.Preview.RunsImmediately {
+ t.Error("a run due since lastCheckTime must be reported as immediate")
+ }
+ if !resp.Actions.SetSchedule.Allowed {
+ t.Errorf("setSchedule = %+v", resp.Actions.SetSchedule)
+ }
+}
+
+func TestCNPGScheduleReadingsWordOnlyValidSchedules(t *testing.T) {
+ sb := func(name, spec string) *unstructured.Unstructured {
+ return &unstructured.Unstructured{Object: map[string]any{
+ "metadata": map[string]any{"namespace": "db", "name": name},
+ "spec": map[string]any{"schedule": spec},
+ }}
+ }
+ got := scheduleReadings([]*unstructured.Unstructured{sb("daily", "0 0 2 * * *"), sb("bad", "not a cron"), sb("empty", "")})
+ if got["db/daily"] != "every day at 02:00 operator clock" {
+ t.Errorf("daily = %q", got["db/daily"])
+ }
+ if _, ok := got["db/bad"]; ok {
+ t.Errorf("an unparseable schedule was worded: %q", got["db/bad"])
+ }
+ if _, ok := got["db/empty"]; ok {
+ t.Error("an empty schedule was worded")
+ }
+}
+
+// Each reading is checked against when the operator's parser actually runs the
+// schedule, so a description can never promise runs the parser would not make.
+func TestDescribeCNPGScheduleMatchesTheParser(t *testing.T) {
+ start := time.Date(2026, 10, 10, 18, 30, 0, 0, time.UTC) // a Saturday
+ cases := []struct {
+ spec string
+ want string
+ runs []string // the parser's next runs from start
+ }{
+ {"0 0 */5 * * *", "every day at 00:00, 05:00, 10:00, 15:00 and 20:00 operator clock", []string{"2026-10-10T20:00:00Z", "2026-10-11T00:00:00Z"}},
+ {"0 0 */6 * * *", "every 6 hours, on the hour", []string{"2026-10-11T00:00:00Z", "2026-10-11T06:00:00Z"}},
+ // Five fields are seconds through month; the day of week is optional.
+ {"0 30 2 * *", "every day at 02:30 operator clock", []string{"2026-10-11T02:30:00Z", "2026-10-12T02:30:00Z"}},
+ {"0 0 */6,13 * * *", "cron 0 0 */6,13 * * *", nil},
+ {"0 0 0 1,*/2 * 1", "cron 0 0 0 1,*/2 * 1", nil},
+ {"0 0 0 */2 * 1", "every 2 days of the month from day 1, when it is a Monday at 00:00 operator clock", []string{"2026-10-19T00:00:00Z"}},
+ {"@every 500ms", "every 1s, counted from the operator's last check", []string{"2026-10-10T18:30:01Z"}},
+ {"@every 1.5s", "every 1s, counted from the operator's last check", []string{"2026-10-10T18:30:01Z"}},
+ {"@every 1h30m", "every 1h30m, counted from the operator's last check", []string{"2026-10-10T20:00:00Z"}},
+ }
+ for _, c := range cases {
+ if got := describeCNPGSchedule(c.spec); got != c.want {
+ t.Errorf("describe(%q) = %q, want %q", c.spec, got, c.want)
+ }
+ sched, err := cnpg.ParseSchedule(c.spec)
+ if err != nil {
+ t.Fatalf("%q: %v", c.spec, err)
+ }
+ at := start
+ for _, want := range c.runs {
+ at = sched.Next(at)
+ if got := at.UTC().Format(time.RFC3339); got != want {
+ t.Errorf("%q: next run %s, want %s (the reading must match)", c.spec, got, want)
+ }
+ }
+ }
+ // */17 minutes restarts each hour: 0, 17, 34, 51 — never every 17 minutes across the hour.
+ got := describeCNPGSchedule("0 */17 * * * *")
+ if !strings.Contains(got, "0, 17, 34, 51") || strings.Contains(got, "every 17") {
+ t.Errorf("*/17 minutes = %q", got)
+ }
+}
diff --git a/internal/cnpg/sessions.go b/internal/cnpg/sessions.go
new file mode 100644
index 0000000000..ac77f01667
--- /dev/null
+++ b/internal/cnpg/sessions.go
@@ -0,0 +1,462 @@
+package cnpg
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "log"
+ "net/http"
+ "strconv"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/rest"
+ "k8s.io/client-go/tools/remotecommand"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+const (
+ execTimeout = 10 * time.Second
+ execStdoutCap = 1 << 20
+ execStderrCap = 8 << 10
+ cnpgExecConcurrency = 4
+ cnpgSessionsMaxRows = 200
+ cnpgQueryTextChars = 200
+ cnpgDiagnosticsApp = "radar-diagnostics"
+ execStateDenied = runtimeStateDenied
+ cnpgPsqlDatabase = "postgres"
+ cnpgSignalCancel = "cancel"
+ cnpgSignalTerminate = "terminate"
+ // cnpgSQLPrelude starts every diagnostic SQL. psql runs as the postgres
+ // superuser, so search_path is pinned to pg_catalog: otherwise a function
+ // someone created in an application schema with an exact-match signature
+ // (pg_total_relation_size(oid), cardinality(int[])) would be chosen over
+ // the built-in and run as superuser.
+ cnpgSQLPrelude = "SET search_path = pg_catalog;\nSET statement_timeout = '5s';\n"
+ // cnpgSQLReadOnly is added to the diagnostics that only read.
+ cnpgSQLReadOnly = "SET default_transaction_read_only = on;\n"
+)
+
+var grantCreateExec = auth.Grant{Verb: "create", Resource: "pods", Subresource: "exec"}
+
+// cnpgExecSlots bounds concurrent execs across all callers: each one holds a
+// streaming connection to a kubelet.
+var execSlots = make(chan struct{}, cnpgExecConcurrency)
+
+// cnpgExecFunc runs argv in a container with stdin and returns stdout. The
+// error carries the command's stderr.
+type ExecFunc func(ctx context.Context, namespace, pod, container string, argv []string, stdin string) ([]byte, error)
+
+var errCNPGExecOutputTooLarge = errors.New("output exceeded the size limit")
+
+type cappedBuffer struct {
+ buf []byte
+ limit int
+ overflow bool
+}
+
+// Write never fails: refusing bytes would break the stream mid-command, so
+// overflow is recorded and reported once the command ends.
+func (b *cappedBuffer) Write(p []byte) (int, error) {
+ room := b.limit - len(b.buf)
+ if room < len(p) {
+ b.overflow = true
+ if room > 0 {
+ b.buf = append(b.buf, p[:room]...)
+ }
+ return len(p), nil
+ }
+ b.buf = append(b.buf, p...)
+ return len(p), nil
+}
+
+// cnpgExecSourceState classifies an exec failure the way the runtime reads
+// classify proxy failures.
+func cnpgExecSourceState(err error) CNPGRuntimeSource {
+ msg := err.Error()
+ lower := strings.ToLower(msg)
+ if strings.Contains(lower, "forbidden") {
+ return CNPGRuntimeSource{State: execStateDenied, Error: truncateCNPGRuntimeError(msg)}
+ }
+ // psql's own "connection refused" from PostgreSQL's socket is not a
+ // transport failure; read it before the transport hints.
+ if postgres, ok := cnpgPostgresSentence(msg); ok {
+ log.Printf("[cnpg] Exec failed: %v", err)
+ return CNPGRuntimeSource{State: runtimeStateError, Error: postgres}
+ }
+ text := truncateCNPGRuntimeError(msg)
+ plain, transport := cnpgTransportSentence(err, 0, execTimeout)
+ if transport {
+ log.Printf("[cnpg] Exec failed: %v", err)
+ text = plain
+ }
+ switch {
+ case transport, errors.Is(err, context.DeadlineExceeded), strings.Contains(lower, "connection refused"), strings.Contains(lower, "no such host"),
+ strings.Contains(lower, "container not found"), strings.Contains(lower, "unable to upgrade connection"):
+ return CNPGRuntimeSource{State: runtimeStateUnreachable, Error: text}
+ default:
+ return CNPGRuntimeSource{State: runtimeStateError, Error: text}
+ }
+}
+
+type CNPGExecPermission struct {
+ Exec string `json:"exec"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+}
+
+// CNPGSessionsResponse is GET /api/cnpg/clusters/{ns}/{name}/sessions.
+// Facts are present only when the read succeeded.
+type CNPGSessionsResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ Pod string `json:"pod"`
+ PodUID types.UID `json:"podUID"`
+ Role string `json:"role"`
+ SampledAt string `json:"sampledAt"`
+ Permission CNPGExecPermission `json:"permission"`
+ CNPGRuntimeSource
+ *CNPGSessionFacts
+ Instances []CNPGSessionInstance `json:"instances"`
+}
+
+// CNPGSessionFacts: Sessions are only the backends in a blocking relation
+// (blocking another, or waiting on one), oldest transaction first, capped at
+// cnpgSessionsMaxRows; InvolvedTotal counts all of them.
+type CNPGSessionFacts struct {
+ ServerTime string `json:"serverTime"`
+ MaxConnections int `json:"maxConnections"`
+ SuperuserReservedConnections int `json:"superuserReservedConnections"`
+ ClientBackends int `json:"clientBackends"`
+ InvolvedTotal int `json:"involvedTotal"`
+ Truncated bool `json:"truncated"`
+ Sessions []CNPGBackend `json:"sessions"`
+}
+
+// CNPGBackend is one pg_stat_activity row. BackendStart is PostgreSQL's own
+// rendering with microseconds; a cancel or terminate must echo it verbatim.
+type CNPGBackend struct {
+ PID int `json:"pid"`
+ BlockedBy []int `json:"blockedBy"`
+ BackendStart string `json:"backendStart"`
+ State string `json:"state,omitempty"`
+ WaitEventType string `json:"waitEventType,omitempty"`
+ WaitEvent string `json:"waitEvent,omitempty"`
+ User string `json:"user,omitempty"`
+ Database string `json:"database,omitempty"`
+ Application string `json:"application,omitempty"`
+ ClientAddr string `json:"clientAddr,omitempty"`
+ BackendType string `json:"backendType,omitempty"`
+ BackendAgeSeconds *float64 `json:"backendAgeSeconds,omitempty"`
+ XactAgeSeconds *float64 `json:"xactAgeSeconds,omitempty"`
+ QueryAgeSeconds *float64 `json:"queryAgeSeconds,omitempty"`
+ StateAgeSeconds *float64 `json:"stateAgeSeconds,omitempty"`
+ Query string `json:"query,omitempty"`
+ QueryTruncated bool `json:"queryTruncated,omitempty"`
+}
+
+// CNPGSessionInstance carries the postgres container's declared resources, so
+// usage (read separately from metrics-server) can be read against them.
+type CNPGSessionInstance struct {
+ Pod string `json:"pod"`
+ Role string `json:"role"`
+ CPURequest string `json:"cpuRequest,omitempty"`
+ CPULimit string `json:"cpuLimit,omitempty"`
+ MemoryRequest string `json:"memoryRequest,omitempty"`
+ MemoryLimit string `json:"memoryLimit,omitempty"`
+}
+
+var cnpgPsqlArgv = []string{"psql", "-XAtq", "-v", "ON_ERROR_STOP=1", "-d", cnpgPsqlDatabase, "-f", "-"}
+
+var cnpgBlockingSQL = cnpgSQLPrelude + cnpgSQLReadOnly + `SET lock_timeout = '1s';
+SET application_name = '` + cnpgDiagnosticsApp + `';
+WITH a AS (
+ SELECT pid, pg_blocking_pids(pid) AS blocked_by, backend_start, xact_start, query_start, state_change,
+ state, wait_event_type, wait_event, usename, datname, application_name,
+ host(client_addr) AS client_addr, backend_type, query
+ FROM pg_stat_activity
+ WHERE pid <> pg_backend_pid()
+),
+involved AS (
+ SELECT * FROM a
+ WHERE cardinality(blocked_by) > 0
+ OR pid IN (SELECT unnest(blocked_by) FROM a)
+)
+SELECT json_build_object(
+ 'serverTime', now(),
+ 'maxConnections', current_setting('max_connections')::int,
+ 'superuserReservedConnections', current_setting('superuser_reserved_connections')::int,
+ 'clientBackends', (SELECT count(*) FROM a WHERE backend_type = 'client backend'),
+ 'involvedTotal', (SELECT count(*) FROM involved),
+ 'sessions', coalesce((
+ SELECT json_agg(row_to_json(s)) FROM (
+ SELECT pid, blocked_by AS "blockedBy", backend_start AS "backendStart",
+ state, wait_event_type AS "waitEventType", wait_event AS "waitEvent",
+ usename AS "user", datname AS "database", application_name AS "application",
+ client_addr AS "clientAddr", backend_type AS "backendType",
+ extract(epoch FROM now() - backend_start) AS "backendAgeSeconds",
+ extract(epoch FROM now() - xact_start) AS "xactAgeSeconds",
+ extract(epoch FROM now() - query_start) AS "queryAgeSeconds",
+ extract(epoch FROM now() - state_change) AS "stateAgeSeconds",
+ left(query, ` + strconv.Itoa(cnpgQueryTextChars) + `) AS query,
+ length(query) > ` + strconv.Itoa(cnpgQueryTextChars) + ` AS "queryTruncated"
+ FROM involved
+ ORDER BY xact_start NULLS LAST, pid
+ LIMIT ` + strconv.Itoa(cnpgSessionsMaxRows) + `
+ ) s
+ ), '[]'::json)
+);
+`
+
+func parseCNPGSessions(out []byte) (*CNPGSessionFacts, error) {
+ var f CNPGSessionFacts
+ if err := json.Unmarshal([]byte(strings.TrimSpace(string(out))), &f); err != nil {
+ return nil, fmt.Errorf("unexpected psql output: %w", err)
+ }
+ if f.Sessions == nil {
+ f.Sessions = []CNPGBackend{}
+ }
+ for i := range f.Sessions {
+ if f.Sessions[i].BlockedBy == nil {
+ f.Sessions[i].BlockedBy = []int{}
+ }
+ }
+ f.Truncated = f.InvolvedTotal > len(f.Sessions)
+ return &f, nil
+}
+
+func sessionInstanceOf(p *corev1.Pod) CNPGSessionInstance {
+ out := CNPGSessionInstance{Pod: p.Name, Role: runtimeRole(p)}
+ for _, c := range p.Spec.Containers {
+ if c.Name != defaultLogContainer {
+ continue
+ }
+ q := func(l corev1.ResourceList, n corev1.ResourceName) string {
+ if v, ok := l[n]; ok {
+ return v.String()
+ }
+ return ""
+ }
+ out.CPURequest, out.CPULimit = q(c.Resources.Requests, corev1.ResourceCPU), q(c.Resources.Limits, corev1.ResourceCPU)
+ out.MemoryRequest, out.MemoryLimit = q(c.Resources.Requests, corev1.ResourceMemory), q(c.Resources.Limits, corev1.ResourceMemory)
+ }
+ return out
+}
+
+func readCNPGSessions(ctx context.Context, exec ExecFunc, namespace, pod string) (CNPGRuntimeSource, *CNPGSessionFacts) {
+ out, err := exec(ctx, namespace, pod, defaultLogContainer, cnpgPsqlArgv, cnpgBlockingSQL)
+ captured := time.Now().UTC().Format(time.RFC3339)
+ if err != nil {
+ src := cnpgExecSourceState(err)
+ src.CapturedAt = captured
+ return src, nil
+ }
+ facts, err := parseCNPGSessions(out)
+ if err != nil {
+ return CNPGRuntimeSource{State: runtimeStateError, Error: err.Error(), CapturedAt: captured}, nil
+ }
+ return CNPGRuntimeSource{State: runtimeStateOK, CapturedAt: captured}, facts
+}
+
+type cnpgSignalParams struct {
+ Pod string `json:"pod"`
+ PodUID string `json:"podUID"`
+ PID int `json:"pid"`
+ BackendStart string `json:"backendStart"`
+}
+
+// cnpgSignalSQL signals only the client backend whose pid AND start time still
+// match what the user reviewed: a reused pid is a different session.
+func cnpgSignalSQL(fn string) string {
+ return cnpgSQLPrelude + `SET application_name = '` + cnpgDiagnosticsApp + `';
+SELECT json_build_object('found', count(*), 'signalled', coalesce(bool_or(` + fn + `(pid)), false))
+FROM pg_stat_activity
+WHERE pid = :'pid'::int
+ AND backend_start = :'backend_start'::timestamptz
+ AND backend_type = 'client backend'
+ AND pid <> pg_backend_pid();
+`
+}
+
+func cnpgSignalArgv(pid int, backendStart string) []string {
+ return []string{"psql", "-XAtq", "-v", "ON_ERROR_STOP=1",
+ "-v", "pid=" + strconv.Itoa(pid),
+ "-v", "backend_start=" + backendStart,
+ "-d", cnpgPsqlDatabase, "-f", "-"}
+}
+
+func cnpgRunSignalBackend(signal string) func(context.Context, *cnpgClusterRun) (*CNPGActionResult, error) {
+ return func(ctx context.Context, x *cnpgClusterRun) (*CNPGActionResult, error) {
+ var p cnpgSignalParams
+ if err := integration.DecodeActionParams(x.params, &p); err != nil {
+ return nil, err
+ }
+ if p.Pod == "" || p.PodUID == "" || p.PID <= 0 || p.BackendStart == "" {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.pod, params.podUID, params.pid and params.backendStart are required")
+ }
+ if _, err := time.Parse(time.RFC3339Nano, p.BackendStart); err != nil {
+ return nil, integration.RefuseAction(http.StatusBadRequest, "", "params.backendStart must be the backend_start the sessions view returned")
+ }
+ if r := cnpgGuardCommon(x.facts); r != "" {
+ return nil, integration.BlockedAction(r)
+ }
+ inst, ok := x.facts.instance(p.Pod)
+ if !ok {
+ return nil, integration.ChangedAction(x.facts, "%s is not an instance of this cluster", p.Pod)
+ }
+ if !inst.PodReadable {
+ return nil, integration.RefuseAction(http.StatusForbidden, "", "Pod %s cannot be read, so it cannot be verified as this cluster's instance", p.Pod)
+ }
+ if !inst.PodExists || inst.PodUID != p.PodUID {
+ return nil, integration.ChangedAction(x.facts, "Pod %s was recreated since you reviewed it; the backend is gone", p.Pod)
+ }
+ if x.c.Exec == nil {
+ return nil, integration.RefuseAction(http.StatusServiceUnavailable, "", "cluster client not available — check cluster connection")
+ }
+ fn, verb := "pg_cancel_backend", "Cancel of the running query"
+ if signal == cnpgSignalTerminate {
+ fn, verb = "pg_terminate_backend", "Termination"
+ }
+ out, err := x.c.Exec(ctx, x.cluster.GetNamespace(), p.Pod, defaultLogContainer, cnpgSignalArgv(p.PID, p.BackendStart), cnpgSignalSQL(fn))
+ if err != nil {
+ if src := cnpgExecSourceState(err); src.State == execStateDenied {
+ return nil, integration.RefuseAction(http.StatusForbidden, "", "This needs %s: %s", grantCreateExec.In(x.cluster.GetNamespace()).String(), src.Error)
+ }
+ return nil, fmt.Errorf("psql on %s: %w", p.Pod, err)
+ }
+ var res struct {
+ Found int `json:"found"`
+ Signalled bool `json:"signalled"`
+ }
+ if err := json.Unmarshal([]byte(strings.TrimSpace(string(out))), &res); err != nil {
+ return nil, fmt.Errorf("unexpected psql output: %w", err)
+ }
+ if res.Found == 0 || !res.Signalled {
+ return nil, integration.ChangedAction(nil, "Backend %d on %s ended or was replaced since you reviewed it; nothing was signalled", p.PID, p.Pod)
+ }
+ return &CNPGActionResult{
+ Message: fmt.Sprintf("%s for backend %d on %s sent", verb, p.PID, p.Pod),
+ Target: &CNPGActionTarget{Pod: p.Pod, PodUID: p.PodUID, PID: p.PID, BackendStart: p.BackendStart},
+ }, nil
+ }
+}
+
+// CNPGActionTarget is what an action acted on. Only the fields that apply are
+// set.
+type CNPGActionTarget struct {
+ Pod string `json:"pod,omitempty"`
+ PodUID string `json:"podUID,omitempty"`
+ PID int `json:"pid,omitempty"`
+ BackendStart string `json:"backendStart,omitempty"`
+ KeepPVC *bool `json:"keepPVC,omitempty"`
+ PVCs []CNPGReviewedObject `json:"pvcs,omitempty"`
+ Jobs []string `json:"jobs,omitempty"`
+ Paused *bool `json:"paused,omitempty"`
+ Generation int64 `json:"generation,omitempty"`
+}
+
+func cnpgGuardPsql(f CNPGClusterFacts, i CNPGInstanceFact) string {
+ switch {
+ case !i.PodReadable:
+ return "The Pod cannot be read"
+ case !i.PodExists:
+ return "The instance has no running Pod"
+ case i.Fenced:
+ return "It is fenced: the operator stops PostgreSQL on a fenced instance"
+ case f.Hibernated:
+ return "The cluster is hibernated"
+ }
+ return ""
+}
+
+func (s *Reader) Sessions(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured, wantedPod string) (*CNPGSessionsResponse, error) {
+ namespace, name := cluster.GetNamespace(), cluster.GetName()
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", k8s.SanitizeForLog(namespace), k8s.SanitizeForLog(name), err)
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "instance Pods unavailable: " + err.Error()}
+ }
+ want := wantedPod
+ primary, _, _ := unstructured.NestedString(cluster.Object, "status", "currentPrimary")
+ if want == "" {
+ want = primary
+ }
+ resp := CNPGSessionsResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ Pod: want,
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ Permission: CNPGExecPermission{Exec: integration.PermissionAllowed, Grant: grantCreateExec.In(namespace).Ref()},
+ Instances: make([]CNPGSessionInstance, 0, len(pods)),
+ }
+ var target *corev1.Pod
+ for _, p := range pods {
+ resp.Instances = append(resp.Instances, sessionInstanceOf(p))
+ if p.Name == want {
+ target = p
+ }
+ }
+ resp.Permission.Exec = s.Access.Permission(ctx, grantCreateExec.In(namespace))
+ if resp.Permission.Exec == integration.PermissionDenied {
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: execStateDenied, Error: "reading sessions needs " + integration.GrantText(resp.Permission.Grant)}
+ return &resp, nil
+ }
+ if target == nil {
+ if wantedPod != "" {
+ return nil, &ReadFailure{Status: http.StatusBadRequest, Message: fmt.Sprintf("%s is not an instance Pod of Cluster %s/%s", want, namespace, name)}
+ }
+ resp.CNPGRuntimeSource = CNPGRuntimeSource{State: "unavailable", Reason: "Available once the primary is running"}
+ return &resp, nil
+ }
+ resp.PodUID = target.UID
+ resp.Role = runtimeRole(target)
+
+ exec := s.Clients.Exec
+ if exec == nil {
+ return nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "cluster client not available — check cluster connection"}
+ }
+ resp.CNPGRuntimeSource, resp.CNPGSessionFacts = readCNPGSessions(ctx, exec, namespace, target.Name)
+ if resp.State == execStateDenied {
+ resp.Permission.Exec = integration.PermissionDenied
+ }
+ return &resp, nil
+}
+
+func NewExec(client kubernetes.Interface, cfg *rest.Config) ExecFunc {
+ if client == nil || cfg == nil {
+ return nil
+ }
+ return func(ctx context.Context, namespace, pod, container string, argv []string, stdin string) ([]byte, error) {
+ select {
+ case execSlots <- struct{}{}:
+ defer func() { <-execSlots }()
+ case <-ctx.Done():
+ return nil, ctx.Err()
+ }
+ ctx, cancel := context.WithTimeout(ctx, execTimeout)
+ defer cancel()
+ ex, err := k8score.NewPodExecExecutor(client, cfg, namespace, pod, container, argv, false)
+ if err != nil {
+ return nil, err
+ }
+ out := &cappedBuffer{limit: execStdoutCap}
+ errOut := &cappedBuffer{limit: execStderrCap}
+ err = ex.StreamWithContext(ctx, remotecommand.StreamOptions{Stdin: strings.NewReader(stdin), Stdout: out, Stderr: errOut})
+ if err != nil {
+ if msg := strings.TrimSpace(string(errOut.buf)); msg != "" {
+ return nil, fmt.Errorf("%w: %s", err, truncateCNPGRuntimeError(msg))
+ }
+ return nil, err
+ }
+ if out.overflow {
+ return nil, errCNPGExecOutputTooLarge
+ }
+ return out.buf, nil
+ }
+}
diff --git a/internal/cnpg/sessions_test.go b/internal/cnpg/sessions_test.go
new file mode 100644
index 0000000000..d6502c3ace
--- /dev/null
+++ b/internal/cnpg/sessions_test.go
@@ -0,0 +1,844 @@
+package cnpg
+
+import (
+ "context"
+ "errors"
+ "net/http"
+ "slices"
+ "strings"
+ "testing"
+
+ appsv1 "k8s.io/api/apps/v1"
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/api/resource"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ k8stesting "k8s.io/client-go/testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+type cnpgExecCall struct {
+ pod, container string
+ argv []string
+ stdin string
+}
+
+func cnpgFakeExec(calls *[]cnpgExecCall, out string, err error) ExecFunc {
+ return func(_ context.Context, _ string, pod, container string, argv []string, stdin string) ([]byte, error) {
+ *calls = append(*calls, cnpgExecCall{pod: pod, container: container, argv: argv, stdin: stdin})
+ return []byte(out), err
+ }
+}
+
+const cnpgSessionsSample = `{"serverTime" : "2026-09-29T21:28:14.277052+00:00", "maxConnections" : 100, "superuserReservedConnections" : 3, "clientBackends" : 2, "involvedTotal" : 3, "sessions" : [{"pid":606,"blockedBy":[],"backendStart":"2026-09-29T16:26:53.564513+00:00","state":"idle in transaction","waitEventType":"Client","waitEvent":"ClientRead","user":"app","database":"app","application":"holder","clientAddr":"10.244.0.57","backendType":"client backend","backendAgeSeconds":18080.7,"xactAgeSeconds":18080.6,"queryAgeSeconds":80.3,"stateAgeSeconds":80.3,"query":"SELECT 1;","queryTruncated":false}, {"pid":607,"blockedBy":[606],"backendStart":"2026-09-29T16:26:53.656271+00:00","state":"active","waitEventType":"Lock","waitEvent":"transactionid","user":"app","database":"app","application":"waiter","clientAddr":null,"backendType":"client backend","backendAgeSeconds":18080.6,"xactAgeSeconds":18080.6,"queryAgeSeconds":18080.6,"stateAgeSeconds":18080.6,"query":"UPDATE t SET v = v + 1","queryTruncated":true}]}
+`
+
+func TestParseCNPGSessions(t *testing.T) {
+ f, err := parseCNPGSessions([]byte(cnpgSessionsSample))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if f.MaxConnections != 100 || f.ClientBackends != 2 || len(f.Sessions) != 2 {
+ t.Fatalf("facts = %+v", f)
+ }
+ if !f.Truncated {
+ t.Error("involvedTotal 3 with 2 rows must read as truncated")
+ }
+ waiter := f.Sessions[1]
+ if waiter.PID != 607 || len(waiter.BlockedBy) != 1 || waiter.BlockedBy[0] != 606 || waiter.WaitEventType != "Lock" || !waiter.QueryTruncated {
+ t.Errorf("waiter = %+v", waiter)
+ }
+ if f.Sessions[0].BlockedBy == nil {
+ t.Error("blockedBy must be an empty list, not null")
+ }
+ if _, err := parseCNPGSessions([]byte("ERROR: boom")); err == nil {
+ t.Error("non-JSON output must be an error")
+ }
+}
+
+func TestReadCNPGSessionsRunsFixedSQLAndClassifiesFailures(t *testing.T) {
+ var calls []cnpgExecCall
+ src, facts := readCNPGSessions(context.Background(), cnpgFakeExec(&calls, cnpgSessionsSample, nil), "db", "pg-1")
+ if src.State != runtimeStateOK || facts == nil {
+ t.Fatalf("state = %+v", src)
+ }
+ c := calls[0]
+ if c.container != "postgres" || c.stdin != cnpgBlockingSQL || strings.Join(c.argv, " ") != "psql -XAtq -v ON_ERROR_STOP=1 -d postgres -f -" {
+ t.Errorf("exec = %+v", c)
+ }
+ if !strings.Contains(cnpgBlockingSQL, "pg_blocking_pids") || !strings.Contains(cnpgBlockingSQL, "pg_backend_pid()") || !strings.Contains(cnpgBlockingSQL, "statement_timeout") {
+ t.Error("blocking SQL must use pg_blocking_pids, exclude its own backend and bound its runtime")
+ }
+
+ src, facts = readCNPGSessions(context.Background(), cnpgFakeExec(&calls, "", errors.New(`pods "pg-1" is forbidden: User "bob" cannot create resource "pods/exec"`)), "db", "pg-1")
+ if src.State != runtimeStateDenied || facts != nil {
+ t.Errorf("forbidden exec = %+v", src)
+ }
+ src, _ = readCNPGSessions(context.Background(), cnpgFakeExec(&calls, "", context.DeadlineExceeded), "db", "pg-1")
+ if src.State != runtimeStateUnreachable {
+ t.Errorf("timeout = %+v", src)
+ }
+}
+
+func TestCNPGCappedBuffer(t *testing.T) {
+ b := &cappedBuffer{limit: 4}
+ if n, err := b.Write([]byte("abcdef")); n != 6 || err != nil {
+ t.Fatalf("write = %d, %v", n, err)
+ }
+ if string(b.buf) != "abcd" || !b.overflow {
+ t.Errorf("buf = %q overflow = %v", b.buf, b.overflow)
+ }
+}
+
+func cnpgSignalParamsFor(podUID string) map[string]any {
+ return map[string]any{"pod": "pg-1", "podUID": podUID, "pid": 607, "backendStart": "2026-09-29T16:26:53.656271+00:00"}
+}
+
+func TestCNPGActionCancelBackendBindsPodPidAndStart(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-1", "u1", true), cnpgActionPod("pg-2", "u2", true))
+ var calls []cnpgExecCall
+ c := env.clients()
+ c.Exec = cnpgFakeExec(&calls, `{"found" : 1, "signalled" : true}`, nil)
+
+ res, err := RunCNPGClusterAction(context.Background(), c, "db", "pg", "cancelBackend", cnpgActionReq(t, nil, cnpgSignalParamsFor("u1")))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if res.Target == nil || res.Target.PID != 607 || res.Target.Pod != "pg-1" || res.Target.BackendStart == "" {
+ t.Errorf("target = %+v", res.Target)
+ }
+ call := calls[0]
+ if call.pod != "pg-1" || call.container != "postgres" {
+ t.Errorf("exec target = %+v", call)
+ }
+ // The user's values travel only as psql variables; the SQL text is fixed.
+ if call.stdin != cnpgSignalSQL("pg_cancel_backend") || strings.Contains(call.stdin, "607") {
+ t.Errorf("stdin = %q", call.stdin)
+ }
+ argv := strings.Join(call.argv, " ")
+ if !strings.Contains(argv, "-v pid=607") || !strings.Contains(argv, "-v backend_start=2026-09-29T16:26:53.656271+00:00") {
+ t.Errorf("argv = %q", argv)
+ }
+ for _, want := range []string{":'pid'::int", ":'backend_start'::timestamptz", "backend_type = 'client backend'"} {
+ if !strings.Contains(call.stdin, want) {
+ t.Errorf("signal SQL lacks %s", want)
+ }
+ }
+
+ calls = nil
+ c.Exec = cnpgFakeExec(&calls, `{"found" : 1, "signalled" : true}`, nil)
+ if _, err := RunCNPGClusterAction(context.Background(), c, "db", "pg", "terminateBackend", cnpgActionReq(t, nil, cnpgSignalParamsFor("u1"))); err != nil {
+ t.Fatal(err)
+ }
+ if !strings.Contains(calls[0].stdin, "pg_terminate_backend(pid)") {
+ t.Errorf("terminate SQL = %q", calls[0].stdin)
+ }
+}
+
+func TestCNPGActionCancelBackendRefusals(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-1", "u1", true))
+ var calls []cnpgExecCall
+ c := env.clients()
+
+ c.Exec = cnpgFakeExec(&calls, `{"found" : 0, "signalled" : false}`, nil)
+ _, err := RunCNPGClusterAction(context.Background(), c, "db", "pg", "cancelBackend", cnpgActionReq(t, nil, cnpgSignalParamsFor("u1")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusConflict || ae.Code != integration.ActionCodeChanged {
+ t.Errorf("gone backend = %v, want 409 changed", err)
+ }
+
+ calls = nil
+ _, err = RunCNPGClusterAction(context.Background(), c, "db", "pg", "cancelBackend", cnpgActionReq(t, nil, cnpgSignalParamsFor("old-uid")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusConflict || len(calls) != 0 {
+ t.Errorf("recreated Pod = %v (calls %d), want 409 before any exec", err, len(calls))
+ }
+
+ bad := cnpgSignalParamsFor("u1")
+ bad["backendStart"] = "yesterday'; DROP TABLE x; --"
+ _, err = RunCNPGClusterAction(context.Background(), c, "db", "pg", "cancelBackend", cnpgActionReq(t, nil, bad))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest || len(calls) != 0 {
+ t.Errorf("malformed backendStart = %v, want 400 before any exec", err)
+ }
+}
+
+func cnpgDestroyTestPVC(name, uid, instance string, owned bool, mut func(*corev1.PersistentVolumeClaim)) *corev1.PersistentVolumeClaim {
+ p := &corev1.PersistentVolumeClaim{
+ ObjectMeta: metav1.ObjectMeta{
+ Name: name, Namespace: "db", UID: types.UID(uid), ResourceVersion: "7",
+ Labels: map[string]string{instanceNameLabel: instance, cnpgPVCRoleLabel: "PG_DATA", "cnpg.io/cluster": "pg"},
+ Annotations: map[string]string{cnpgPVCStatusAnnotation: "ready"},
+ },
+ Status: corev1.PersistentVolumeClaimStatus{Capacity: corev1.ResourceList{corev1.ResourceStorage: resource.MustParse("1Gi")}},
+ }
+ if owned {
+ controller := true
+ p.OwnerReferences = []metav1.OwnerReference{{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg", UID: cnpgActionTestUID, Controller: &controller}}
+ }
+ if mut != nil {
+ mut(p)
+ }
+ return p
+}
+
+// cnpgDestroyCluster is the default fixture with pg-2 fenced, as destroy requires.
+func cnpgDestroyCluster(mut func(obj map[string]any)) *unstructured.Unstructured {
+ return cnpgActionCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: `["pg-2"]`}
+ if mut != nil {
+ mut(o)
+ }
+ })
+}
+
+func cnpgDestroyFacts() map[string]any {
+ f := cnpgActionFacts()
+ f["fencedInstances"] = map[string]any{"raw": `["pg-2"]`, "all": false, "instances": []any{"pg-2"}}
+ return f
+}
+
+func cnpgDestroyEnv(t *testing.T) *cnpgActionEnv {
+ job := &batchv1.Job{ObjectMeta: metav1.ObjectMeta{Name: "pg-2-join", Namespace: "db", UID: "job-2", ResourceVersion: "3", OwnerReferences: testControllerRefs("Cluster", "pg", cnpgActionTestUID), Labels: map[string]string{instanceNameLabel: "pg-2"}}}
+ return newCNPGActionEnv(t, []runtime.Object{cnpgDestroyCluster(nil)},
+ cnpgActionPod("pg-1", "u1", true), cnpgActionPod("pg-2", "u2", true),
+ cnpgDestroyTestPVC("pg-2", "pvc-2", "pg-2", true, nil),
+ cnpgDestroyTestPVC("pg-2-wal", "pvc-2w", "pg-2", true, func(p *corev1.PersistentVolumeClaim) { p.Labels[cnpgPVCRoleLabel] = "PG_WAL" }),
+ cnpgDestroyTestPVC("pg-2-foreign", "pvc-x", "pg-2", false, nil),
+ cnpgDestroyTestPVC("pg-1", "pvc-1", "pg-1", true, nil),
+ job)
+}
+
+func cnpgDestroyParamsFor(keep bool, pvcs ...string) map[string]any {
+ list := []any{}
+ for _, p := range pvcs {
+ name, uid, _ := strings.Cut(p, "=")
+ list = append(list, map[string]any{"name": name, "uid": uid})
+ }
+ return map[string]any{"pod": "pg-2", "podUID": "u2", "keepPVC": keep, "pvcs": list, "jobs": []any{map[string]any{"name": "pg-2-join", "uid": "job-2"}}}
+}
+
+func TestCNPGActionDestroyInstanceDeletesLikeKubectlCNPG(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ var order []string
+ env.typed.PrependReactor("delete", "*", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ d := a.(k8stesting.DeleteAction)
+ order = append(order, a.GetResource().Resource+"/"+d.GetName())
+ return false, nil, nil
+ })
+ res, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if err != nil {
+ t.Fatal(err)
+ }
+ // PVCs first so the operator never sees a dangling PVC, then the Pod, then Jobs.
+ want := "persistentvolumeclaims/pg-2,persistentvolumeclaims/pg-2-wal,pods/pg-2,jobs/pg-2-join"
+ if strings.Join(order, ",") != want {
+ t.Errorf("order = %v, want %s", order, want)
+ }
+ if res.Target == nil || res.Target.Pod != "pg-2" || len(res.Target.PVCs) != 2 || len(res.Target.Jobs) != 1 || *res.Target.KeepPVC {
+ t.Errorf("target = %+v", res.Target)
+ }
+ for _, d := range env.deletes {
+ if d.GetName() == "pg-2" && (d.GetDeleteOptions().Preconditions == nil || *d.GetDeleteOptions().Preconditions.UID != "u2") {
+ t.Errorf("pod delete lacks the UID precondition: %+v", d.GetDeleteOptions())
+ }
+ }
+}
+
+func TestCNPGActionDestroyInstanceKeepPVCDetaches(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(true, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if err != nil {
+ t.Fatal(err)
+ }
+ pvc, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(context.Background(), "pg-2", metav1.GetOptions{})
+ if err != nil {
+ t.Fatalf("kept PVC was deleted: %v", err)
+ }
+ if len(pvc.OwnerReferences) != 0 || pvc.Annotations[cnpgPVCStatusAnnotation] != "detached" || pvc.Annotations[cnpgDetachedClusterUID] != cnpgActionTestUID || pvc.Labels[instanceNameLabel] != "pg-2" {
+ t.Errorf("kept PVC = owners %v annotations %v labels %v", pvc.OwnerReferences, pvc.Annotations, pvc.Labels)
+ }
+ if len(env.deletes) != 1 || env.deletes[0].GetName() != "pg-2" {
+ t.Errorf("deletes = %v, want only the Pod", env.deletes)
+ }
+}
+
+func TestCNPGActionDestroyInstanceRefusals(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ primary := cnpgDestroyParamsFor(false, "pg-1=pvc-1")
+ primary["pod"], primary["podUID"] = "pg-1", "u1"
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), primary))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked || !strings.Contains(ae.Message, "primary") {
+ t.Errorf("primary = %v, want blocked", err)
+ }
+
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Errorf("unreviewed WAL volume = %v, want 409 changed", err)
+ }
+
+ noPVCs := cnpgDestroyParamsFor(false)
+ delete(noPVCs, "pvcs")
+ _, err = RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), noPVCs))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Errorf("missing pvcs = %v, want 400", err)
+ }
+ if len(env.deletes) != 0 {
+ t.Errorf("a refusal deleted %v", env.deletes)
+ }
+}
+
+func TestCNPGActionDestroyInstanceStopsWhenPromotedMidway(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ promoted := false
+ env.typed.PrependReactor("delete", "persistentvolumeclaims", func(k8stesting.Action) (bool, runtime.Object, error) {
+ if !promoted {
+ promoted = true
+ obj := cnpgDestroyCluster(func(o map[string]any) {
+ st := o["status"].(map[string]any)
+ st["targetPrimary"], st["phase"] = "pg-2", cnpgPhaseFailover
+ })
+ if err := env.dyn.Tracker().Update(ClusterGVR, obj, "db"); err != nil {
+ t.Fatal(err)
+ }
+ }
+ return false, nil, nil
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Code != integration.ActionCodePartial || ae.Status != http.StatusConflict || strings.Join(ae.Completed, ",") != "deleted PVC pg-2" {
+ t.Fatalf("promotion mid-destroy = %+v, want 409 partial after the first PVC only", err)
+ }
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(context.Background(), "pg-2-wal", metav1.GetOptions{}); err != nil {
+ t.Errorf("the WAL volume of a now-target instance was deleted: %v", err)
+ }
+ if len(env.deletes) != 0 {
+ t.Errorf("the Pod of a now-target instance was deleted: %v", env.deletes)
+ }
+}
+
+func TestCNPGActionDestroyInstanceRefusesPodLabelledPrimary(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ pod := cnpgActionPod("pg-2", "u2", true)
+ pod.Labels["cnpg.io/instanceRole"] = "primary"
+ if _, err := env.typed.CoreV1().Pods("db").Update(context.Background(), pod, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged || len(ae.Completed) != 0 {
+ t.Fatalf("pod labelled primary = %v, want 409 changed before any write", err)
+ }
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(context.Background(), "pg-2", metav1.GetOptions{}); err != nil {
+ t.Errorf("PVC deleted despite the refusal: %v", err)
+ }
+}
+
+func TestCNPGActionDestroyInstanceRequiresFence(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ if err := env.dyn.Tracker().Update(ClusterGVR, cnpgActionCluster(nil), "db"); err != nil {
+ t.Fatal(err)
+ }
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgActionFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked || !strings.HasPrefix(ae.Message, "Fence pg-2 first") {
+ t.Fatalf("unfenced destroy = %v, want blocked asking to fence first", err)
+ }
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(context.Background(), "pg-2", metav1.GetOptions{}); err != nil || len(env.deletes) != 0 {
+ t.Errorf("an unfenced destroy wrote: PVC %v, deletes %v", err, env.deletes)
+ }
+}
+
+// A failover can only make pg-2 primary once its fence is gone, so the fence
+// is re-checked before every step like the primary itself.
+func TestCNPGActionDestroyInstanceStopsWhenUnfencedMidway(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ promoted := false
+ env.typed.PrependReactor("delete", "persistentvolumeclaims", func(k8stesting.Action) (bool, runtime.Object, error) {
+ if !promoted {
+ promoted = true
+ if err := env.dyn.Tracker().Update(ClusterGVR, cnpgActionCluster(nil), "db"); err != nil {
+ t.Fatal(err)
+ }
+ }
+ return false, nil, nil
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodePartial || strings.Join(ae.Completed, ",") != "deleted PVC pg-2" {
+ t.Fatalf("fence lifted mid-destroy = %v, want partial after the first PVC", err)
+ }
+ if len(env.deletes) != 0 {
+ t.Errorf("the Pod of an unfenced instance was deleted: %v", env.deletes)
+ }
+}
+
+func TestCNPGActionDestroyInstanceLiftsItsFenceLast(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ res, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if err != nil {
+ t.Fatal(err)
+ }
+ if len(env.patches) != 1 {
+ t.Fatalf("patches = %d, want the fence lift only", len(env.patches))
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ if v, ok := cnpgActionAnnotations(t, body)[cnpgFencedAnnotation]; !ok || v != nil {
+ t.Errorf("fence after destroy = %v (present %v), want removed", v, ok)
+ }
+ if body["metadata"].(map[string]any)["resourceVersion"] == nil {
+ t.Error("fence lift lacks resourceVersion")
+ }
+ if strings.Contains(res.Message, `["*"]`) {
+ t.Errorf("message = %q", res.Message)
+ }
+
+ all := cnpgDestroyEnv(t)
+ star := cnpgDestroyCluster(func(o map[string]any) {
+ o["metadata"].(map[string]any)["annotations"] = map[string]any{cnpgFencedAnnotation: `["*"]`}
+ })
+ if err := all.dyn.Tracker().Update(ClusterGVR, star, "db"); err != nil {
+ t.Fatal(err)
+ }
+ facts := cnpgDestroyFacts()
+ facts["fencedInstances"] = map[string]any{"raw": `["*"]`, "all": true, "instances": []any{"*"}}
+ res, err = RunCNPGClusterAction(context.Background(), all.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, facts, cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if err != nil || len(all.patches) != 0 || !strings.Contains(res.Message, `["*"]`) {
+ t.Fatalf(`["*"] fence: err %v, patches %d, message %v`, err, len(all.patches), res)
+ }
+}
+
+func TestCNPGActionDestroyInstanceFailedFenceLiftIsPartial(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ env.dyn.PrependReactor("patch", "clusters", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewConflict(schema.GroupResource{Group: Group, Resource: "clusters"}, "pg", errors.New("modified"))
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Code != integration.ActionCodePartial || !strings.Contains(ae.Message, "Unfence") || !slices.Contains(ae.Completed, "deleted Pod pg-2") {
+ t.Fatalf("failed fence lift = %v, want partial naming Unfence", err)
+ }
+}
+
+func TestCNPGDestroyPlanAndCapabilities(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ srv := newTestReader(nil)
+ ctx := context.Background()
+ plan, err := srv.DestroyPlan(ctx, env.clients(), "kind-test", "db", "pg", "pg-2")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if !plan.PVCsReadable || len(plan.PVCs) != 2 || plan.PVCs[0].Name != "pg-2" || plan.PVCs[1].Role != "PG_WAL" || plan.PVCs[0].Capacity != "1Gi" {
+ t.Errorf("pvcs = %+v", plan.PVCs)
+ }
+ if len(plan.Jobs) != 1 || plan.PodUID != "u2" || !plan.Actions.Delete.Allowed || !plan.Actions.Keep.Allowed {
+ t.Errorf("plan = %+v", plan)
+ }
+
+ caps, err := srv.ClusterCapabilities(ctx, env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if caps.InstanceActions["pg-1"].Destroy.Allowed || !caps.InstanceActions["pg-2"].Destroy.Allowed || !caps.Actions.DestroyInstance.Allowed {
+ t.Errorf("destroy verdicts = %+v / %+v", caps.InstanceActions, caps.Actions.DestroyInstance)
+ }
+ if got := caps.InstanceActions["pg-3"].Destroy; got.Allowed || got.ReasonCode != "fence_required" || got.Reason != "Fence pg-3 first: a fenced instance cannot be promoted while it is destroyed" {
+ t.Errorf("unfenced destroy verdict = %+v", got)
+ }
+ unfenced, err := srv.DestroyPlan(ctx, env.clients(), "kind-test", "db", "pg", "pg-3")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if unfenced.Actions.Delete.Allowed || unfenced.Actions.Delete.ReasonCode != "fence_required" || unfenced.Actions.Keep.ReasonCode != "fence_required" {
+ t.Errorf("unfenced destroy plan = %+v", unfenced.Actions)
+ }
+ // pg-2 is fenced here, so PostgreSQL is stopped on it.
+ if !caps.Actions.Psql.Allowed || caps.InstanceActions["pg-2"].Psql.Allowed {
+ t.Errorf("psql = %+v / %+v", caps.Actions.Psql, caps.InstanceActions["pg-2"].Psql)
+ }
+}
+
+func TestCNPGCapabilitiesPsqlNamesExecGrant(t *testing.T) {
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgActionCluster(nil)}, cnpgActionPod("pg-1", "u1", true))
+ srv := newTestReader(nil)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"db"}}
+ perms.SetCanI("create", "", "pods/exec", "db", false)
+ srv = newTestReader(perms)
+ ctx := context.Background()
+ caps, err := srv.ClusterCapabilities(ctx, env.clients(), "kind-test", "db", "pg")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if p := caps.Actions.Psql; p.Allowed || !strings.Contains(p.Reason, "create pods/exec") {
+ t.Errorf("psql = %+v, want denied naming create pods/exec", p)
+ }
+}
+
+func cnpgTestPooler(paused bool) *unstructured.Unstructured {
+ return &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "Pooler",
+ "metadata": map[string]any{"name": "pool", "namespace": "db", "uid": "pooler-uid", "resourceVersion": "9", "generation": int64(2)},
+ "spec": map[string]any{
+ "cluster": map[string]any{"name": "pg"}, "type": "rw", "instances": int64(2),
+ "pgbouncer": map[string]any{"poolMode": "transaction", "paused": paused, "parameters": map[string]any{"default_pool_size": "10"}},
+ },
+ }}
+}
+
+func TestCNPGPoolerPauseBindsReviewedState(t *testing.T) {
+ controller := true
+ deploy := &appsv1.Deployment{
+ ObjectMeta: metav1.ObjectMeta{Name: "pool", Namespace: "db", OwnerReferences: []metav1.OwnerReference{{APIVersion: "postgresql.cnpg.io/v1", Kind: "Pooler", Name: "pool", UID: "pooler-uid", Controller: &controller}}},
+ Status: appsv1.DeploymentStatus{ReadyReplicas: 1, UpdatedReplicas: 2, AvailableReplicas: 1},
+ }
+ env := newCNPGActionEnv(t, []runtime.Object{cnpgTestPooler(false)}, deploy)
+ req := integration.ActionRequest{ReviewedContext: "kind-test", UID: "pooler-uid", Facts: []byte(`{"paused":false}`)}
+
+ res, err := RunCNPGPoolerAction(context.Background(), env.clients(), "db", "pool", "pause", req)
+ if err != nil {
+ t.Fatal(err)
+ }
+ body := cnpgActionPatchBody(t, env.patches[0])
+ spec := body["spec"].(map[string]any)["pgbouncer"].(map[string]any)
+ if spec["paused"] != true || body["metadata"].(map[string]any)["resourceVersion"] != "9" {
+ t.Errorf("patch = %v", body)
+ }
+ if res.Target == nil || res.Target.Paused == nil || !*res.Target.Paused {
+ t.Errorf("target = %+v", res.Target)
+ }
+
+ _, err = RunCNPGPoolerAction(context.Background(), env.clients(), "db", "pool", "resume", integration.ActionRequest{ReviewedContext: "kind-test", UID: "pooler-uid", Facts: []byte(`{"paused":true}`)})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged {
+ t.Errorf("stale paused = %v, want 409 changed", err)
+ }
+ _, err = RunCNPGPoolerAction(context.Background(), env.clients(), "db", "pool", "pause", integration.ActionRequest{ReviewedContext: "kind-test", UID: "pooler-uid"})
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Status != http.StatusBadRequest {
+ t.Errorf("missing facts = %v, want 400", err)
+ }
+
+ srv := newTestReader(nil)
+ caps, err := srv.PoolerCapabilities(context.Background(), env.clients(), "kind-test", "db", "pool")
+ if err != nil {
+ t.Fatal(err)
+ }
+ d := caps.Facts.Deployment
+ if d.State != "ok" || *d.ReadyReplicas != 1 || caps.Facts.Service.State != "missing" || caps.Facts.Parameters["default_pool_size"] != "10" {
+ t.Errorf("facts = %+v", caps.Facts)
+ }
+ if !caps.Actions.Pause.Allowed || caps.Actions.Resume.Allowed {
+ t.Errorf("actions = %+v", caps.Actions)
+ }
+}
+
+func TestCNPGPoolerFactsReportPoolMode(t *testing.T) {
+ samples := map[string][]cnpgSample{
+ "cnpg_pgbouncer_pools_pool_mode": {
+ {labels: map[string]string{"database": "app", "user": "app"}, value: 2},
+ {labels: map[string]string{"database": "pgbouncer", "user": "pgbouncer"}, value: 3},
+ },
+ }
+ facts, _ := cnpgPoolerFacts(samples)
+ if len(facts.Pools) != 1 || facts.Pools[0].PoolMode != "transaction" {
+ t.Errorf("pools = %+v, want app/app in transaction mode and the admin pool excluded", facts.Pools)
+ }
+}
+
+func TestParseCNPGShowState(t *testing.T) {
+ st, err := parseCNPGShowState([]byte("active|yes\npaused|no\nsuspended|no\n"))
+ if err != nil || st.Paused == nil || *st.Paused || st.Active == nil || !*st.Active {
+ t.Fatalf("state = %+v, %v", st, err)
+ }
+ if _, err := parseCNPGShowState([]byte("ERROR: invalid command")); err == nil {
+ t.Error("output without paused must be an error")
+ }
+}
+
+func TestCNPGActionDestroyInstanceRefusesWithoutEveryGrantBeforeAnyWrite(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ env.denied = map[string]bool{"patch clusters": true}
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Status != http.StatusForbidden || ae.Code == integration.ActionCodePartial {
+ t.Fatalf("err = %v, want a plain 403 naming patch clusters", err)
+ }
+ if !strings.Contains(ae.Message, "patch clusters") {
+ t.Errorf("message = %q", ae.Message)
+ }
+ if len(env.deletes) != 0 || len(env.patches) != 0 {
+ t.Errorf("wrote before the preflight refused: %d deletes, %d patches", len(env.deletes), len(env.patches))
+ }
+}
+
+func TestCNPGActionDestroyInstanceLeavesARecreatedClustersFences(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ env.typed.PrependReactor("delete", "jobs", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ fresh := cnpgDestroyCluster(nil)
+ fresh.SetUID("recreated")
+ fresh.SetResourceVersion("100")
+ if err := env.dyn.Tracker().Update(ClusterGVR, fresh, "db"); err != nil {
+ t.Fatal(err)
+ }
+ return false, nil, nil
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) || ae.Code != integration.ActionCodePartial {
+ t.Fatalf("err = %v, want partial", err)
+ }
+ for _, p := range env.patches {
+ if p.GetResource().Resource == "clusters" {
+ t.Errorf("patched the recreated Cluster: %s", p.GetPatch())
+ }
+ }
+}
+
+func TestCNPGActionDestroyInstanceReportsPartialOutcome(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ env.typed.PrependReactor("delete", "pods", func(k8stesting.Action) (bool, runtime.Object, error) {
+ return true, nil, apierrors.NewForbidden(schema.GroupResource{Resource: "pods"}, "pg-2", errors.New("denied by admission policy"))
+ })
+ _, err := RunCNPGClusterAction(context.Background(), env.clients(), "db", "pg", "destroyInstance",
+ cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ ae, ok := cnpgActionStatus(t, err)
+ if !ok || ae.Code != integration.ActionCodePartial || ae.Status != http.StatusForbidden ||
+ strings.Join(ae.Completed, ",") != "deleted PVC pg-2,deleted PVC pg-2-wal" {
+ t.Fatalf("pod delete refused after PVCs = %+v, want 403 partial listing both PVCs", err)
+ }
+
+}
+
+func TestCNPGPoolerStatePending(t *testing.T) {
+ pending := &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "pooler-1"}, Status: corev1.PodStatus{Phase: corev1.PodPending, Conditions: []corev1.PodCondition{{Type: corev1.PodScheduled, Status: corev1.ConditionFalse}}}}
+ for _, tc := range []struct {
+ name, output string
+ err error
+ state, reason string
+ }{
+ {"pending transport", "", errors.New("address not allowed"), runtimeStateUnreachable, "PgBouncer has not started (Pod cannot be scheduled)"},
+ {"permission wins", "", apierrors.NewForbidden(schema.GroupResource{Resource: "pods/exec"}, "pooler-1", errors.New("denied")), execStateDenied, "denied"},
+ {"successful read wins", "paused|no\n", nil, runtimeStateOK, ""},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ var calls []cnpgExecCall
+ got := readCNPGPgBouncerState(context.Background(), cnpgFakeExec(&calls, tc.output, tc.err), "db", pending)
+ if got.State != tc.state || !strings.Contains(got.Error, tc.reason) || got.Pod != "pooler-1" {
+ t.Fatalf("%+v", got)
+ }
+ if tc.state == runtimeStateOK && (got.Paused == nil || *got.Paused) {
+ t.Fatalf("lost SHOW STATE: %+v", got)
+ }
+ })
+ }
+ running := pending.DeepCopy()
+ running.Status.Phase = corev1.PodRunning
+ var calls []cnpgExecCall
+ got := readCNPGPgBouncerState(context.Background(), cnpgFakeExec(&calls, "", errors.New("address not allowed")), "db", running)
+ if !strings.Contains(got.Error, "address not allowed") {
+ t.Fatalf("lost running Pod transport details: %+v", got)
+ }
+}
+
+func TestCNPGDestroyIgnoresForeignAndPredecessorResources(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ ctx := context.Background()
+ for _, uid := range []string{"", "old-cluster", cnpgActionTestUID} {
+ pvc := cnpgDestroyTestPVC("detached-"+uid, "pvc-"+uid, "pg-2", false, func(p *corev1.PersistentVolumeClaim) {
+ p.Annotations[cnpgPVCStatusAnnotation] = cnpgPVCStatusDetached
+ p.Annotations[cnpgDetachedClusterUID] = uid
+ })
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Create(ctx, pvc, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ }
+ for _, uid := range []string{"", "old-cluster"} {
+ job := &batchv1.Job{ObjectMeta: metav1.ObjectMeta{Name: "foreign-" + uid, Namespace: "db", UID: types.UID("job-" + uid), Labels: map[string]string{instanceNameLabel: "pg-2"}, OwnerReferences: testControllerRefs("Cluster", "pg", uid)}}
+ if _, err := env.typed.BatchV1().Jobs("db").Create(ctx, job, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ }
+ plan, err := newTestReader(nil).DestroyPlan(ctx, env.clients(), "kind-test", "db", "pg", "pg-2")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if len(plan.PVCs) != 3 || len(plan.Jobs) != 1 || plan.Jobs[0].UID != "job-2" {
+ t.Fatalf("unrelated objects entered the plan: %+v", plan)
+ }
+ params := cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w", "detached-"+cnpgActionTestUID+"=pvc-"+cnpgActionTestUID)
+ if _, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), params)); err != nil {
+ t.Fatal(err)
+ }
+ for _, uid := range []string{"", "old-cluster"} {
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(ctx, "detached-"+uid, metav1.GetOptions{}); err != nil {
+ t.Fatalf("predecessor PVC deleted: %v", err)
+ }
+ if _, err := env.typed.BatchV1().Jobs("db").Get(ctx, "foreign-"+uid, metav1.GetOptions{}); err != nil {
+ t.Fatalf("foreign Job deleted: %v", err)
+ }
+ }
+}
+
+func TestCNPGDestroyRefusesChangedJobsBeforeWriting(t *testing.T) {
+ for _, scenario := range []string{"added", "replaced", "unreadable"} {
+ t.Run(scenario, func(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ ctx := context.Background()
+ job, err := env.typed.BatchV1().Jobs("db").Get(ctx, "pg-2-join", metav1.GetOptions{})
+ if err != nil {
+ t.Fatal(err)
+ }
+ switch scenario {
+ case "added":
+ job.Name, job.UID = "pg-2-extra", "job-extra"
+ if _, err := env.typed.BatchV1().Jobs("db").Create(ctx, job, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ case "replaced":
+ job.UID = "replacement"
+ if _, err := env.typed.BatchV1().Jobs("db").Update(ctx, job, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ case "unreadable":
+ env.typed.PrependReactor("list", "jobs", func(k8stesting.Action) (bool, runtime.Object, error) { return true, nil, apiForbidden("jobs") })
+ plan, err := newTestReader(nil).DestroyPlan(ctx, env.clients(), "kind-test", "db", "pg", "pg-2")
+ if err != nil || plan.JobsReadable || plan.Actions.Delete.Allowed || plan.Actions.Keep.Allowed {
+ t.Fatalf("unreadable Jobs allowed: %+v, %v", plan, err)
+ }
+ }
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if scenario == "unreadable" {
+ if !apierrors.IsForbidden(err) {
+ t.Fatalf("unreadable Jobs: %v", err)
+ }
+ } else if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeChanged || !strings.Contains(ae.Message, "Jobs") {
+ t.Fatalf("wrong refusal for changed Jobs: %v", err)
+ }
+ for _, a := range env.typed.Actions() {
+ if a.GetVerb() == "delete" || a.GetResource().Resource == "persistentvolumeclaims" && a.GetVerb() == "update" {
+ t.Fatalf("write before refusing: %v", a)
+ }
+ }
+ })
+ }
+}
+
+func TestCNPGDestroyJobReplacementAfterReviewStopsWithPreconditions(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ ctx := context.Background()
+ env.typed.PrependReactor("delete", "pods", func(k8stesting.Action) (bool, runtime.Object, error) {
+ obj, err := env.typed.Tracker().Get(batchv1.SchemeGroupVersion.WithResource("jobs"), "db", "pg-2-join")
+ if err != nil {
+ t.Fatal(err)
+ }
+ job := obj.(*batchv1.Job).DeepCopy()
+ job.UID, job.ResourceVersion = "replacement", "4"
+ if err := env.typed.Tracker().Update(batchv1.SchemeGroupVersion.WithResource("jobs"), job, "db"); err != nil {
+ t.Fatal(err)
+ }
+ return false, nil, nil
+ })
+ env.typed.PrependReactor("delete", "jobs", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ p := a.(k8stesting.DeleteAction).GetDeleteOptions().Preconditions
+ if p == nil || p.UID == nil || *p.UID != "job-2" || p.ResourceVersion == nil || *p.ResourceVersion != "3" {
+ t.Fatalf("delete not bound to reviewed Job: %+v", p)
+ }
+ return true, nil, apierrors.NewConflict(schema.GroupResource{Group: "batch", Resource: "jobs"}, "pg-2-join", errors.New("UID changed"))
+ })
+ _, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodePartial || len(ae.Completed) != 3 || len(env.patches) != 0 {
+ t.Fatalf("replacement must stop before lifting fence: %v", err)
+ }
+ job, err := env.typed.BatchV1().Jobs("db").Get(ctx, "pg-2-join", metav1.GetOptions{})
+ if err != nil || job.UID != "replacement" {
+ t.Fatalf("replacement lost: %+v, %v", job, err)
+ }
+}
+
+func TestCNPGDestroyKeepPVCRefusesUnrelatedOwners(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ ctx := context.Background()
+ pvc, err := env.typed.CoreV1().PersistentVolumeClaims("db").Get(ctx, "pg-2", metav1.GetOptions{})
+ if err != nil {
+ t.Fatal(err)
+ }
+ foreign := []metav1.OwnerReference{{APIVersion: "cluster.x-k8s.io/v1beta1", Kind: "Cluster", Name: "pg", UID: "capi-uid"}, {APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg", UID: "old-cluster"}}
+ pvc.OwnerReferences = append(pvc.OwnerReferences, foreign...)
+ if _, err := env.typed.CoreV1().PersistentVolumeClaims("db").Update(ctx, pvc, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ plan, err := newTestReader(nil).DestroyPlan(ctx, env.clients(), "kind-test", "db", "pg", "pg-2")
+ if err != nil || plan.Actions.Keep.Allowed || !plan.Actions.Delete.Allowed || !strings.Contains(plan.Actions.Keep.Reason, "capi-uid") {
+ t.Fatalf("unsafe keep plan: %+v, %v", plan, err)
+ }
+ _, err = RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(true, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodeBlocked || !strings.Contains(ae.Message, "garbage-collect") || len(env.deletes) != 0 {
+ t.Fatalf("unsafe keep allowed: %v", err)
+ }
+ pvc, err = env.typed.CoreV1().PersistentVolumeClaims("db").Get(ctx, "pg-2", metav1.GetOptions{})
+ if err != nil || len(pvc.OwnerReferences) != 3 || pvc.OwnerReferences[1] != foreign[0] || pvc.OwnerReferences[2] != foreign[1] {
+ t.Fatalf("unrelated owners lost: %+v, %v", pvc, err)
+ }
+}
+
+func TestCNPGDestroyJobStatusChurnRetainsIdentityAndOwnershipGuards(t *testing.T) {
+ for _, scenario := range []string{"status", "ownership", "repeated-conflict"} {
+ t.Run(scenario, func(t *testing.T) {
+ env := cnpgDestroyEnv(t)
+ ctx := context.Background()
+ resource := batchv1.SchemeGroupVersion.WithResource("jobs")
+ calls := 0
+ env.typed.PrependReactor("delete", "jobs", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ calls++
+ pre := a.(k8stesting.DeleteAction).GetDeleteOptions().Preconditions
+ if pre == nil || pre.UID == nil || *pre.UID != "job-2" || pre.ResourceVersion == nil {
+ t.Fatalf("identity not bound: %+v", pre)
+ }
+ if calls == 1 {
+ obj, err := env.typed.Tracker().Get(resource, "db", "pg-2-join")
+ if err != nil {
+ t.Fatal(err)
+ }
+ job := obj.(*batchv1.Job).DeepCopy()
+ job.ResourceVersion, job.Status.Active = "4", 1
+ if scenario == "ownership" {
+ job.OwnerReferences = testControllerRefs("Cluster", "pg", "other-cluster")
+ }
+ if err := env.typed.Tracker().Update(resource, job, "db"); err != nil {
+ t.Fatal(err)
+ }
+ return true, nil, apierrors.NewConflict(resource.GroupResource(), job.Name, errors.New("version changed"))
+ }
+ if calls > 2 || *pre.ResourceVersion != "4" {
+ t.Fatalf("retry not bounded to fresh version: %d, %+v", calls, pre)
+ }
+ if scenario == "repeated-conflict" {
+ return true, nil, apierrors.NewConflict(resource.GroupResource(), "pg-2-join", errors.New("version changed again"))
+ }
+ return false, nil, nil
+ })
+ _, err := RunCNPGClusterAction(ctx, env.clients(), "db", "pg", "destroyInstance", cnpgActionReq(t, cnpgDestroyFacts(), cnpgDestroyParamsFor(false, "pg-2=pvc-2", "pg-2-wal=pvc-2w")))
+ if scenario == "status" {
+ if err != nil || calls != 2 || len(env.patches) != 1 {
+ t.Fatalf("status-only churn prevented completion: %v, calls=%d", err, calls)
+ }
+ } else if ae, ok := cnpgActionStatus(t, err); !ok || ae.Code != integration.ActionCodePartial || len(env.patches) != 0 {
+ t.Fatalf("unsafe retry or fence lift: %v", err)
+ }
+ if scenario == "ownership" && calls != 1 {
+ t.Fatalf("retried after losing ownership: %d", calls)
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/storage.go b/internal/cnpg/storage.go
new file mode 100644
index 0000000000..9928f094c9
--- /dev/null
+++ b/internal/cnpg/storage.go
@@ -0,0 +1,911 @@
+package cnpg
+
+import (
+ "context"
+ "fmt"
+ "log"
+ "net/http"
+ "sort"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/selection"
+
+ auth "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+const (
+ cnpgPVCRoleLabel = "cnpg.io/pvcRole"
+ instanceNameLabel = "cnpg.io/instanceName"
+ cnpgTablespaceNameLabel = "cnpg.io/tablespaceName"
+
+ cnpgPVCRoleData = "PG_DATA"
+ cnpgPVCRoleWAL = "PG_WAL"
+ cnpgPVCRoleTablespace = "PG_TABLESPACE"
+
+ storageStateOK = "ok"
+ cnpgStorageStatePartial = "partial"
+ storageStateDenied = "denied"
+ cnpgStorageStateUnavailable = "unavailable"
+ cnpgStorageStateError = "error"
+
+ cnpgUsageStateOK = "ok"
+ cnpgUsageStateNoSeries = "noSeries"
+ cnpgUsageStateInvalid = "invalid"
+ cnpgUsageStateNoPrometheus = "noPrometheus"
+ cnpgUsageStateDenied = "denied"
+ cnpgUsageStateError = "error"
+ usageStateNotRead = "notRead"
+
+ cnpgDiskWarningRatio = 0.80
+ cnpgDiskCriticalRatio = 0.90
+
+ usageSource = "Prometheus kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes"
+)
+
+// The fleet reads at most this many namespaces per request, a few at a time:
+// each costs two Prometheus queries.
+var (
+ fleetDiskMaxNamespaces = 64
+ fleetDiskConcurrency = 4
+)
+
+// CNPGStorageCoverage states whether one source could be read and, when not,
+// the grant that would allow it.
+type CNPGStorageCoverage struct {
+ integration.ReadSource
+ // Isolation, on usage read from Prometheus: how the series were tied to this cluster.
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+}
+
+// CNPGClusterStorageResponse is GET /api/cnpg/clusters/{namespace}/{name}/storage.
+type CNPGClusterStorageResponse struct {
+ Cluster CNPGRuntimeObjectRef `json:"cluster"`
+ SampledAt string `json:"sampledAt"`
+ // Volumes: listing the claims. Usage: Prometheus kubelet volume stats.
+ // WAL: the instance manager and exporter through pods/proxy.
+ Volumes CNPGStorageCoverage `json:"volumes"`
+ Usage CNPGStorageCoverage `json:"usage"`
+ UsageFrom string `json:"usageSource"`
+ WAL CNPGStorageCoverage `json:"wal"`
+ Expansion CNPGStorageExpansion `json:"expansion"`
+ Instances []CNPGStorageInstance `json:"instances"`
+ Findings []CNPGStorageFinding `json:"findings"`
+ Excluded []CNPGStorageExcludedPV `json:"excluded,omitempty"`
+}
+
+// CNPGStorageExpansion names where each volume's size is declared: the only
+// place a resize can be made, which the operator then applies to the claims.
+type CNPGStorageExpansion struct {
+ Targets []CNPGStorageTarget `json:"targets"`
+ // ResizeInUseVolumes is spec.storage.resizeInUseVolumes; absent means the
+ // field is unset, which CNPG treats as true.
+ ResizeInUseVolumes *bool `json:"resizeInUseVolumes,omitempty"`
+}
+
+type CNPGStorageTarget struct {
+ Role string `json:"role"`
+ Tablespace string `json:"tablespace,omitempty"`
+ Field string `json:"field"`
+ Declared string `json:"declared,omitempty"`
+ StorageClass string `json:"storageClass,omitempty"`
+}
+
+type CNPGStorageInstance struct {
+ Name string `json:"name"`
+ Role string `json:"role"`
+ Volumes []CNPGStorageVolume `json:"volumes"`
+ WAL *CNPGStorageWAL `json:"wal,omitempty"`
+}
+
+type CNPGStorageVolume struct {
+ Claim string `json:"claim"`
+ Role string `json:"role"`
+ Tablespace string `json:"tablespace,omitempty"`
+ Phase string `json:"phase,omitempty"`
+ Requested string `json:"requested,omitempty"`
+ RequestedBytes *int64 `json:"requestedBytes,omitempty"`
+ Capacity string `json:"capacity,omitempty"`
+ CapacityBytes *int64 `json:"capacityBytes,omitempty"`
+ StorageClass CNPGStorageClassFact `json:"storageClass"`
+ Resize CNPGStorageResize `json:"resize"`
+ ClusterState string `json:"clusterState,omitempty"`
+ Usage CNPGStorageVolumeUsage `json:"usage"`
+}
+
+// CNPGStorageClassFact: AllowVolumeExpansion is absent when the class could
+// not be read; a readable class without the field does not allow expansion.
+type CNPGStorageClassFact struct {
+ Name string `json:"name,omitempty"`
+ AllowVolumeExpansion *bool `json:"allowVolumeExpansion,omitempty"`
+ Reason string `json:"reason,omitempty"`
+}
+
+// CNPGStorageResize carries what the claim itself reports about a resize.
+type CNPGStorageResize struct {
+ Conditions []CNPGStorageResizeCondition `json:"conditions,omitempty"`
+ // AllocatedStatus is status.allocatedResourceStatuses.storage.
+ AllocatedStatus string `json:"allocatedStatus,omitempty"`
+ // Pending: the request is larger than the capacity the claim reports.
+ Pending bool `json:"pending,omitempty"`
+}
+
+type CNPGStorageResizeCondition struct {
+ Type string `json:"type"`
+ Message string `json:"message,omitempty"`
+ Since string `json:"since,omitempty"`
+}
+
+// CNPGStorageVolumeUsage carries figures only when State is ok.
+type CNPGStorageVolumeUsage struct {
+ State string `json:"state"`
+ UsedBytes *int64 `json:"usedBytes,omitempty"`
+ CapacityBytes *int64 `json:"capacityBytes,omitempty"`
+ Ratio *float64 `json:"ratio,omitempty"`
+}
+
+// CNPGStorageWAL keeps each measure separate: the WAL directory's size, the
+// segments waiting to be archived and the WAL each slot retains overlap, so
+// they are never added up.
+type CNPGStorageWAL struct {
+ Status CNPGRuntimeSource `json:"status"`
+ Metrics CNPGRuntimeSource `json:"metrics"`
+ Volume string `json:"volume,omitempty"`
+ SizeBytes *float64 `json:"sizeBytes,omitempty"`
+ Segments *float64 `json:"segments,omitempty"`
+ ReadyToArchive *int `json:"readyToArchive,omitempty"`
+ LastArchivedAt string `json:"lastArchivedAt,omitempty"`
+ LastFailedAt string `json:"lastFailedAt,omitempty"`
+ LastFailedWal string `json:"lastFailedWal,omitempty"`
+ ArchivingFailed bool `json:"archivingFailed,omitempty"`
+ Slots []CNPGSlotBytes `json:"slots,omitempty"`
+ SlotInventory []CNPGSlotStatus `json:"slotInventory"`
+ SlotInventoryTruncated bool `json:"slotInventoryTruncated,omitempty"`
+}
+
+type CNPGStorageFinding struct {
+ Severity string `json:"severity"`
+ Instance string `json:"instance"`
+ Claim string `json:"claim"`
+ Role string `json:"role"`
+ Tablespace string `json:"tablespace,omitempty"`
+ Ratio float64 `json:"ratio"`
+ Message string `json:"message"`
+}
+
+// CNPGStorageExcludedPV is a claim labelled for the Cluster that the Cluster
+// does not own, or that names no instance. A label is something any object
+// can carry, so it is listed rather than counted as the Cluster's volume.
+type CNPGStorageExcludedPV struct {
+ Claim string `json:"claim"`
+ Reason string `json:"reason"`
+}
+
+func (s *Reader) ClusterStorage(ctx context.Context, namespace, name string) (*CNPGClusterStorageResponse, error) {
+ cache, cluster, err := s.Observations.Cluster(ctx, namespace, name)
+ if err != nil {
+ return nil, err
+ }
+ resp := CNPGClusterStorageResponse{
+ Cluster: CNPGRuntimeObjectRef{Namespace: namespace, Name: name, UID: cluster.GetUID()},
+ SampledAt: time.Now().UTC().Format(time.RFC3339),
+ UsageFrom: usageSource,
+ Expansion: cnpgStorageExpansionOf(cluster),
+ Instances: []CNPGStorageInstance{},
+ Findings: []CNPGStorageFinding{},
+ }
+
+ claims, excluded, volumes := s.clusterClaims(ctx, cache, cluster)
+ resp.Volumes, resp.Excluded = volumes, excluded
+ classes := s.storageClasses(ctx, cache, claims)
+ states := cnpgClusterPVCStates(cluster)
+
+ usage := CNPGStorageCoverage{ReadSource: integration.ReadSource{State: usageStateNotRead, Reason: "no volumes to measure"}}
+ var batch prometheuspkg.PVCUsageBatch
+ if len(claims) > 0 {
+ usage, batch = s.claimUsage(ctx, namespace, claimNames(claims), historyAnchors(cache, cluster))
+ }
+ resp.Usage = usage
+
+ byInstance := map[string]*CNPGStorageInstance{}
+ instance := func(n string) *CNPGStorageInstance {
+ if in, ok := byInstance[n]; ok {
+ return in
+ }
+ in := &CNPGStorageInstance{Name: n, Role: cnpgStorageInstanceRole(cluster, n), Volumes: []CNPGStorageVolume{}}
+ byInstance[n] = in
+ return in
+ }
+ for _, n := range cnpgStatusStrings(cluster, "instanceNames") {
+ instance(n)
+ }
+ for _, pvc := range claims {
+ v := cnpgStorageVolumeOf(pvc, classes, states)
+ v.Usage = cnpgVolumeUsage(usage, batch, pvc.Name)
+ in := instance(pvc.Labels[instanceNameLabel])
+ in.Volumes = append(in.Volumes, v)
+ }
+
+ resp.WAL, err = s.storageWAL(ctx, cache, cluster, byInstance)
+ if err != nil {
+ return nil, err
+ }
+
+ for _, in := range byInstance {
+ sort.Slice(in.Volumes, func(i, j int) bool {
+ a, b := in.Volumes[i], in.Volumes[j]
+ if a.Role != b.Role {
+ return cnpgPVCRoleRank(a.Role) < cnpgPVCRoleRank(b.Role)
+ }
+ return a.Claim < b.Claim
+ })
+ resp.Instances = append(resp.Instances, *in)
+ resp.Findings = append(resp.Findings, cnpgDiskFindings(in.Name, in.Volumes, resp.Usage.Isolation)...)
+ }
+ sort.Slice(resp.Instances, func(i, j int) bool { return resp.Instances[i].Name < resp.Instances[j].Name })
+ sort.SliceStable(resp.Findings, func(i, j int) bool { return resp.Findings[i].Ratio > resp.Findings[j].Ratio })
+ return &resp, nil
+}
+
+// cnpgClusterClaims lists the claims the Cluster owns, after the caller's own
+// list grant — reading a Cluster never implies reading its claims.
+func (s *Reader) clusterClaims(ctx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured) ([]*corev1.PersistentVolumeClaim, []CNPGStorageExcludedPV, CNPGStorageCoverage) {
+ namespace := cluster.GetNamespace()
+ grant := cnpgGrantListPVCs.In(namespace).Ref()
+ if !s.Access.CanRead(ctx, "", "persistentvolumeclaims", namespace, "list") {
+ return nil, nil, CNPGStorageCoverage{ReadSource: integration.ReadSource{State: storageStateDenied, Grant: grant}}
+ }
+ candidates, reason := cnpgCachedClaims(cache, namespace, labels.SelectorFromSet(labels.Set{clusterLabel: cluster.GetName()}))
+ if reason != "" {
+ return nil, nil, CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgStorageStateUnavailable, Grant: grant, Reason: reason}}
+ }
+ owned, excluded := cnpgOwnedClaims(candidates, cluster)
+ return owned, excluded, CNPGStorageCoverage{ReadSource: integration.ReadSource{State: storageStateOK, Grant: grant}}
+}
+
+// cnpgCachedClaims reads claims from Radar's informer. What the informer does
+// not hold is unread, not empty, so that is a reason rather than no claims.
+func cnpgCachedClaims(cache *k8s.ResourceCache, namespace string, selector labels.Selector) ([]*corev1.PersistentVolumeClaim, string) {
+ lister := cache.PersistentVolumeClaims()
+ if lister == nil {
+ return nil, "Radar's own identity cannot list persistentvolumeclaims"
+ }
+ if within := integration.NamespacesWithinCache(cache, "persistentvolumeclaims", []string{namespace}); within.Unavailable {
+ return nil, "Radar's own identity cannot list persistentvolumeclaims in " + namespace
+ }
+ items, err := lister.PersistentVolumeClaims(namespace).List(selector)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list claims in %s: %v", namespace, err)
+ return nil, "listing persistentvolumeclaims failed"
+ }
+ return items, ""
+}
+
+// cnpgOwnedClaims keeps claims that name an instance and are owned by this
+// Cluster's UID.
+func cnpgOwnedClaims(candidates []*corev1.PersistentVolumeClaim, cluster *unstructured.Unstructured) ([]*corev1.PersistentVolumeClaim, []CNPGStorageExcludedPV) {
+ var owned []*corev1.PersistentVolumeClaim
+ var excluded []CNPGStorageExcludedPV
+ for _, pvc := range candidates {
+ if pvc == nil {
+ continue
+ }
+ switch {
+ case !cnpgOwnedBy(pvc, cluster) && cnpgHasClusterOwner(pvc):
+ excluded = append(excluded, CNPGStorageExcludedPV{Claim: pvc.Name, Reason: "owned by a different Cluster object (UID mismatch)"})
+ case !cnpgOwnedBy(pvc, cluster):
+ excluded = append(excluded, CNPGStorageExcludedPV{Claim: pvc.Name, Reason: "carries this cluster's labels but no owner reference to it, as a claim kept after its instance was destroyed does"})
+ case pvc.Labels[instanceNameLabel] == "":
+ excluded = append(excluded, CNPGStorageExcludedPV{Claim: pvc.Name, Reason: "names no instance (" + instanceNameLabel + ")"})
+ default:
+ owned = append(owned, pvc)
+ }
+ }
+ sort.Slice(owned, func(i, j int) bool { return owned[i].Name < owned[j].Name })
+ sort.Slice(excluded, func(i, j int) bool { return excluded[i].Claim < excluded[j].Claim })
+ return owned, excluded
+}
+
+func cnpgOwnedBy(pvc *corev1.PersistentVolumeClaim, cluster *unstructured.Unstructured) bool {
+ for _, ref := range pvc.OwnerReferences {
+ if ref.Kind == "Cluster" && ref.UID == cluster.GetUID() && strings.HasPrefix(ref.APIVersion, Group+"/") {
+ return true
+ }
+ }
+ return false
+}
+
+func cnpgHasClusterOwner(pvc *corev1.PersistentVolumeClaim) bool {
+ for _, ref := range pvc.OwnerReferences {
+ if ref.Kind == "Cluster" && strings.HasPrefix(ref.APIVersion, Group+"/") {
+ return true
+ }
+ }
+ return false
+}
+
+func claimNames(claims []*corev1.PersistentVolumeClaim) []string {
+ out := make([]string, len(claims))
+ for i, c := range claims {
+ out[i] = c.Name
+ }
+ return out
+}
+
+// cnpgStorageClasses reads each named class once, when the caller may get
+// StorageClasses.
+func (s *Reader) storageClasses(ctx context.Context, cache *k8s.ResourceCache, claims []*corev1.PersistentVolumeClaim) map[string]CNPGStorageClassFact {
+ out := map[string]CNPGStorageClassFact{}
+ allowed := s.Access.CanRead(ctx, "storage.k8s.io", "storageclasses", "", "get")
+ lister := cache.StorageClasses()
+ for _, pvc := range claims {
+ if pvc.Spec.StorageClassName == nil || *pvc.Spec.StorageClassName == "" {
+ continue
+ }
+ name := *pvc.Spec.StorageClassName
+ if _, done := out[name]; done {
+ continue
+ }
+ fact := CNPGStorageClassFact{Name: name}
+ switch {
+ case !allowed:
+ fact.Reason = "needs get storageclasses"
+ case lister == nil:
+ fact.Reason = "Radar's own credentials cannot list StorageClasses"
+ default:
+ sc, err := lister.Get(name)
+ if err != nil || sc == nil {
+ fact.Reason = "StorageClass " + name + " not found"
+ break
+ }
+ allow := sc.AllowVolumeExpansion != nil && *sc.AllowVolumeExpansion
+ fact.AllowVolumeExpansion = &allow
+ }
+ out[name] = fact
+ }
+ return out
+}
+
+// cnpgClusterPVCStates maps each claim to the Cluster status list naming it.
+func cnpgClusterPVCStates(cluster *unstructured.Unstructured) map[string]string {
+ out := map[string]string{}
+ for _, f := range []struct{ field, state string }{
+ {"healthyPVC", "healthy"},
+ {"initializingPVC", "initializing"},
+ {"resizingPVC", "resizing"},
+ {"danglingPVC", "dangling"},
+ {"unusablePVC", "unusable"},
+ } {
+ for _, n := range cnpgStatusStrings(cluster, f.field) {
+ out[n] = f.state
+ }
+ }
+ return out
+}
+
+func cnpgStatusStrings(cluster *unstructured.Unstructured, field string) []string {
+ v, _, _ := unstructured.NestedStringSlice(cluster.Object, "status", field)
+ return v
+}
+
+func cnpgStorageInstanceRole(cluster *unstructured.Unstructured, instance string) string {
+ primary, _, _ := unstructured.NestedString(cluster.Object, "status", "currentPrimary")
+ switch {
+ case instance == "":
+ return "unknown"
+ case primary != "" && instance == primary:
+ return "primary"
+ case primary != "":
+ for _, n := range cnpgStatusStrings(cluster, "instanceNames") {
+ if n == instance {
+ return "replica"
+ }
+ }
+ return "noInstance"
+ }
+ return "unknown"
+}
+
+func cnpgPVCRoleRank(role string) int {
+ switch role {
+ case cnpgPVCRoleData:
+ return 0
+ case cnpgPVCRoleWAL:
+ return 1
+ case cnpgPVCRoleTablespace:
+ return 2
+ }
+ return 3
+}
+
+var cnpgResizeConditions = map[corev1.PersistentVolumeClaimConditionType]bool{
+ corev1.PersistentVolumeClaimResizing: true,
+ corev1.PersistentVolumeClaimFileSystemResizePending: true,
+ "ControllerResizeError": true,
+ "NodeResizeError": true,
+}
+
+func cnpgStorageVolumeOf(pvc *corev1.PersistentVolumeClaim, classes map[string]CNPGStorageClassFact, states map[string]string) CNPGStorageVolume {
+ v := CNPGStorageVolume{
+ Claim: pvc.Name,
+ Role: pvc.Labels[cnpgPVCRoleLabel],
+ Tablespace: pvc.Labels[cnpgTablespaceNameLabel],
+ Phase: string(pvc.Status.Phase),
+ ClusterState: states[pvc.Name],
+ }
+ if q, ok := pvc.Spec.Resources.Requests[corev1.ResourceStorage]; ok {
+ b := q.Value()
+ v.Requested, v.RequestedBytes = q.String(), &b
+ }
+ if q, ok := pvc.Status.Capacity[corev1.ResourceStorage]; ok {
+ b := q.Value()
+ v.Capacity, v.CapacityBytes = q.String(), &b
+ }
+ if pvc.Spec.StorageClassName != nil && *pvc.Spec.StorageClassName != "" {
+ v.StorageClass = classes[*pvc.Spec.StorageClassName]
+ } else {
+ v.StorageClass = CNPGStorageClassFact{Reason: "the claim names no StorageClass"}
+ }
+ for _, c := range pvc.Status.Conditions {
+ if !cnpgResizeConditions[c.Type] || c.Status != corev1.ConditionTrue {
+ continue
+ }
+ rc := CNPGStorageResizeCondition{Type: string(c.Type), Message: c.Message}
+ if !c.LastTransitionTime.IsZero() {
+ rc.Since = c.LastTransitionTime.UTC().Format(time.RFC3339)
+ }
+ v.Resize.Conditions = append(v.Resize.Conditions, rc)
+ }
+ if st, ok := pvc.Status.AllocatedResourceStatuses[corev1.ResourceStorage]; ok {
+ v.Resize.AllocatedStatus = string(st)
+ }
+ if v.RequestedBytes != nil && v.CapacityBytes != nil && *v.RequestedBytes > *v.CapacityBytes {
+ v.Resize.Pending = true
+ }
+ return v
+}
+
+// cnpgClaimUsage reads kubelet volume stats for the claims, behind the same
+// gate as the single-claim PVC chart.
+func (s *Reader) claimUsage(ctx context.Context, namespace string, claims []string, anchors []prom.WorkloadPodIdentity) (CNPGStorageCoverage, prometheuspkg.PVCUsageBatch) {
+ grant := grantGetPVCs.In(namespace).Ref()
+ if !s.Access.MetricsRead(ctx, "", "persistentvolumeclaims", namespace, "get") {
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgUsageStateDenied, Grant: grant}}, prometheuspkg.PVCUsageBatch{}
+ }
+ batch := s.Metrics.PVCUsage(ctx, namespace, claims, anchors)
+ switch batch.Status {
+ case prometheuspkg.PVCUsageNoPrometheus:
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgUsageStateNoPrometheus, Reason: cnpgNoPrometheusReason(batch.Error)}}, batch
+ case prometheuspkg.PVCUsageQueryFailed:
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgUsageStateError, Reason: "Prometheus query failed: " + truncateCNPGRuntimeError(batch.Error)}}, batch
+ case prometheuspkg.PVCUsageAmbiguous:
+ _, reason := usageScopeFailure(prometheuspkg.ErrScopeAmbiguous)
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgHistoryStateAmbiguous, Reason: reason}}, batch
+ case prometheuspkg.PVCUsageScopeMismatch:
+ _, reason := usageScopeFailure(prometheuspkg.ErrScopeMismatch)
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgHistoryStateScopeMismatch, Reason: reason}}, batch
+ }
+ iso := &batch.Isolation
+ switch n := len(batch.Usage); {
+ case n == 0:
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgUsageStateNoSeries, Reason: "Prometheus has no kubelet volume stats for these claims"}}, batch
+ case n < len(claims):
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgStorageStatePartial, Reason: fmt.Sprintf("%d of %d claims have kubelet volume stats", n, len(claims))}, Isolation: iso}, batch
+ }
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgUsageStateOK}, Isolation: iso}, batch
+}
+
+// cnpgUnverifiedCaveat qualifies a measurement stated as this cluster's when
+// its series were matched by name alone.
+func cnpgUnverifiedCaveat(iso *prometheuspkg.SeriesIsolation) string {
+ if iso == nil || iso.Mode != prometheuspkg.SeriesIsolationUnverified {
+ return ""
+ }
+ return ". " + iso.Note
+}
+
+func cnpgVolumeUsage(cov CNPGStorageCoverage, batch prometheuspkg.PVCUsageBatch, claim string) CNPGStorageVolumeUsage {
+ switch cov.State {
+ case cnpgUsageStateOK, cnpgStorageStatePartial, cnpgUsageStateNoSeries:
+ default:
+ return CNPGStorageVolumeUsage{State: cov.State}
+ }
+ if u, ok := batch.Usage[claim]; ok {
+ used, capacity, ratio := u.UsedBytes, u.CapacityBytes, u.Ratio
+ return CNPGStorageVolumeUsage{State: cnpgUsageStateOK, UsedBytes: &used, CapacityBytes: &capacity, Ratio: &ratio}
+ }
+ if batch.Invalid[claim] {
+ return CNPGStorageVolumeUsage{State: cnpgUsageStateInvalid}
+ }
+ return CNPGStorageVolumeUsage{State: cnpgUsageStateNoSeries}
+}
+
+func cnpgVolumeLabel(role, tablespace string) string {
+ switch role {
+ case cnpgPVCRoleData:
+ return "data volume"
+ case cnpgPVCRoleWAL:
+ return "WAL volume"
+ case cnpgPVCRoleTablespace:
+ return "tablespace " + tablespace + " volume"
+ }
+ return "volume"
+}
+
+// cnpgDiskFindings reports volumes whose measured use crosses the thresholds.
+// Only a measurement can raise one; an unmeasured volume says nothing.
+func cnpgDiskFindings(instance string, volumes []CNPGStorageVolume, iso *prometheuspkg.SeriesIsolation) []CNPGStorageFinding {
+ var out []CNPGStorageFinding
+ for _, v := range volumes {
+ if v.Usage.Ratio == nil || *v.Usage.Ratio < cnpgDiskWarningRatio {
+ continue
+ }
+ sev := "warning"
+ if *v.Usage.Ratio >= cnpgDiskCriticalRatio {
+ sev = "critical"
+ }
+ out = append(out, CNPGStorageFinding{
+ Severity: sev, Instance: instance, Claim: v.Claim, Role: v.Role, Tablespace: v.Tablespace, Ratio: *v.Usage.Ratio,
+ Message: fmt.Sprintf("The %s of %s is %.0f%% full%s", cnpgVolumeLabel(v.Role, v.Tablespace), instance, *v.Usage.Ratio*100, cnpgUnverifiedCaveat(iso)),
+ })
+ }
+ return out
+}
+
+// cnpgStorageExpansionOf names the spec field each volume's size comes from.
+func cnpgStorageExpansionOf(cluster *unstructured.Unstructured) CNPGStorageExpansion {
+ out := CNPGStorageExpansion{Targets: []CNPGStorageTarget{cnpgStorageTargetOf(cluster.Object, cnpgPVCRoleData, "", "spec.storage", "spec", "storage")}}
+ if _, found, _ := unstructured.NestedMap(cluster.Object, "spec", "walStorage"); found {
+ out.Targets = append(out.Targets, cnpgStorageTargetOf(cluster.Object, cnpgPVCRoleWAL, "", "spec.walStorage", "spec", "walStorage"))
+ }
+ tablespaces, _, _ := unstructured.NestedSlice(cluster.Object, "spec", "tablespaces")
+ for _, raw := range tablespaces {
+ ts, ok := raw.(map[string]any)
+ if !ok {
+ continue
+ }
+ name, _ := ts["name"].(string)
+ out.Targets = append(out.Targets, cnpgStorageTargetOf(ts, cnpgPVCRoleTablespace, name, "spec.tablespaces[name="+name+"].storage", "storage"))
+ }
+ if v, found, _ := unstructured.NestedBool(cluster.Object, "spec", "storage", "resizeInUseVolumes"); found {
+ out.ResizeInUseVolumes = &v
+ }
+ return out
+}
+
+func cnpgStorageTargetOf(obj map[string]any, role, tablespace, base string, path ...string) CNPGStorageTarget {
+ t := CNPGStorageTarget{Role: role, Tablespace: tablespace, Field: base + ".size"}
+ if size, found, _ := unstructured.NestedString(obj, append(append([]string{}, path...), "size")...); found && size != "" {
+ t.Declared = size
+ } else if req, found, _ := unstructured.NestedString(obj, append(append([]string{}, path...), "pvcTemplate", "resources", "requests", "storage")...); found && req != "" {
+ t.Field, t.Declared = base+".pvcTemplate.resources.requests.storage", req
+ }
+ if sc, found, _ := unstructured.NestedString(obj, append(append([]string{}, path...), "storageClass")...); found {
+ t.StorageClass = sc
+ }
+ return t
+}
+
+// cnpgStorageWAL reads each instance's WAL facts through the same memoized
+// pods/proxy reads the Replication and Performance tabs use. It writes the error and returns an
+// empty state only when the caller's client cannot be built.
+func (s *Reader) storageWAL(callerCtx context.Context, cache *k8s.ResourceCache, cluster *unstructured.Unstructured, byInstance map[string]*CNPGStorageInstance) (CNPGStorageCoverage, error) {
+ namespace := cluster.GetNamespace()
+ if !s.Access.CanRead(callerCtx, "", "pods", namespace, "list") {
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: storageStateDenied, Grant: GrantListPods.In(namespace).Ref()}}, nil
+ }
+ grant := grantGetPodsProxy.In(namespace).Ref()
+ if s.Access.Permission(callerCtx, grantGetPodsProxy.In(namespace)) == integration.PermissionDenied {
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: storageStateDenied, Grant: grant}}, nil
+ }
+ pods, err := clusterInstancePods(cache, cluster)
+ if err != nil {
+ log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", namespace, cluster.GetName(), err)
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgStorageStateUnavailable, Grant: grant, Reason: "instance Pods unavailable"}}, nil
+ }
+ if len(pods) == 0 {
+ return CNPGStorageCoverage{ReadSource: integration.ReadSource{State: cnpgStorageStateUnavailable, Grant: grant, Reason: "no instance Pods"}}, nil
+ }
+ client := s.Clients.Proxy
+ if client == nil {
+ return CNPGStorageCoverage{}, &ReadFailure{http.StatusServiceUnavailable, "cluster client unavailable"}
+ }
+
+ metricsTLS, _, _ := unstructured.NestedBool(cluster.Object, "spec", "monitoring", "tls", "enabled")
+ fenced := parseCNPGFenced(cluster.GetAnnotations()[cnpgFencedAnnotation])
+ identity := s.Identity
+ walVolume := map[string]string{}
+ for n, in := range byInstance {
+ for _, v := range in.Volumes {
+ if v.Role == cnpgPVCRoleWAL || (v.Role == cnpgPVCRoleData && walVolume[n] == "") {
+ walVolume[n] = v.Claim
+ }
+ }
+ }
+
+ results := make([]CNPGStorageWAL, len(pods))
+ run := newCNPGRuntimeRunner(callerCtx)
+ for i, p := range pods {
+ statusTarget, metricsTarget := cnpgInstanceProxyTargets(p, metricsTLS)
+ out := &results[i]
+ run.do(func(ctx context.Context) {
+ st := memoized(ctx, identity, statusTarget, cnpgStatusMemoTTL, func(ctx context.Context) CNPGInstanceStatus {
+ return cnpgInstanceStatusFrom(proxyGetWithFallback(ctx, client, statusTarget))
+ })
+ out.Status = st.CNPGRuntimeSource
+ if st.CNPGInstanceStatusFacts != nil {
+ out.SlotInventory = st.Slots
+ out.SlotInventoryTruncated = st.SlotsTruncated
+ if a := st.Archiving; a != nil {
+ out.ReadyToArchive, out.LastArchivedAt, out.LastFailedAt, out.LastFailedWal = a.ReadyWalFiles, a.LastArchivedAt, a.LastFailedAt, a.LastFailedWal
+ out.ArchivingFailed = cnpgArchivingFailedLast(*a)
+ }
+ }
+ })
+ run.do(func(ctx context.Context) {
+ m := memoized(ctx, identity, metricsTarget, metricsMemoTTL, func(ctx context.Context) CNPGInstanceMetrics {
+ return cnpgInstanceMetricsFrom(proxyGetWithFallback(ctx, client, metricsTarget))
+ })
+ out.Metrics = m.CNPGRuntimeSource
+ if m.CNPGInstanceMetricFacts != nil {
+ out.SizeBytes, out.Segments = m.WalBytes, m.WalSegments
+ out.Slots = append([]CNPGSlotBytes(nil), m.ReplicationSlotsRetainedBytes...)
+ // Without the WAL collector the exporter's queries are failing,
+ // so an empty slot list is not "no slots".
+ if out.Metrics.State == runtimeStateOK && m.WalBytes == nil {
+ out.Metrics.State = cnpgRuntimeStatePartial
+ out.Metrics.Reason = "the exporter reported no WAL figures in this sample, so WAL size and slots were not measured"
+ }
+ }
+ })
+ }
+ run.wait()
+
+ cov := CNPGStorageCoverage{ReadSource: integration.ReadSource{State: storageStateOK, Grant: grant}}
+ failed, partial := 0, 0
+ for i, p := range pods {
+ wal := results[i]
+ wal.SlotInventory = withCNPGSlotRetention(&CNPGInstanceStatusFacts{Slots: wal.SlotInventory}, wal.Slots).Slots
+ if fenced.Fences(p.Name) {
+ explainCNPGFenced(&wal.Status)
+ explainCNPGFenced(&wal.Metrics)
+ }
+ if wal.Status.State == runtimeStateDenied || wal.Metrics.State == runtimeStateDenied {
+ cov.State = storageStateDenied
+ }
+ switch statusOK, metricsOK := wal.Status.State == runtimeStateOK, wal.Metrics.State == runtimeStateOK; {
+ case !statusOK && !metricsOK:
+ failed++
+ case !statusOK || !metricsOK:
+ partial++
+ }
+ wal.Volume = walVolume[p.Name]
+ in, ok := byInstance[p.Name]
+ if !ok {
+ in = &CNPGStorageInstance{Name: p.Name, Role: cnpgStorageInstanceRole(cluster, p.Name), Volumes: []CNPGStorageVolume{}}
+ byInstance[p.Name] = in
+ }
+ in.WAL = &wal
+ }
+ if cov.State == storageStateOK && failed+partial > 0 {
+ cov.State = cnpgStorageStatePartial
+ cov.Reason = cnpgWALCoverageReason(failed, partial, len(pods))
+ }
+ return cov, nil
+}
+
+// cnpgInstanceProxyTargets builds the same targets as the live instance reads, so the
+// memo serves both from one read.
+func cnpgInstanceProxyTargets(p *corev1.Pod, clusterMetricsTLS bool) (status, metrics proxyTarget) {
+ status = proxyTarget{
+ namespace: p.Namespace, pod: p.Name, podUID: p.UID, port: cnpgStatusPort, path: cnpgStatusPath,
+ scheme: schemeFor(containerHasFlag(p, defaultLogContainer, cnpgStatusPortTLSFlag)), limit: cnpgRuntimeStatusCap,
+ }
+ metrics = proxyTarget{
+ namespace: p.Namespace, pod: p.Name, podUID: p.UID, port: cnpgMetricsPort, path: metricsPath,
+ scheme: schemeFor(containerHasFlag(p, defaultLogContainer, metricsPortTLSFlag) || clusterMetricsTLS), limit: runtimeMetricsCap,
+ }
+ return status, metrics
+}
+
+func cnpgArchivingFailedLast(a CNPGArchivingStatus) bool {
+ if a.LastFailedAt == "" {
+ return false
+ }
+ failed, err := time.Parse(time.RFC3339, a.LastFailedAt)
+ if err != nil {
+ return false
+ }
+ if a.LastArchivedAt == "" {
+ return true
+ }
+ archived, err := time.Parse(time.RFC3339, a.LastArchivedAt)
+ return err == nil && failed.After(archived)
+}
+
+// CNPGFleetDiskResponse is GET /api/cnpg/disk: the fullest measured volume of
+// each visible Cluster.
+type CNPGFleetDiskResponse struct {
+ SampledAt string `json:"sampledAt"`
+ Source string `json:"source"`
+ Clusters []CNPGClusterDisk `json:"clusters"`
+}
+
+// CNPGClusterDisk State: ok (every claim measured), partial, noSeries,
+// noPrometheus, denied (Grant names what is missing), unavailable, error, or
+// notRead (the request's namespace bound was reached).
+type CNPGClusterDisk struct {
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ State string `json:"state"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Claims int `json:"claims"`
+ Measured int `json:"measured"`
+ Max *CNPGDiskUsage `json:"max,omitempty"`
+ Isolation *prometheuspkg.SeriesIsolation `json:"isolation,omitempty"`
+}
+
+type CNPGDiskUsage struct {
+ Claim string `json:"claim"`
+ Instance string `json:"instance"`
+ Role string `json:"role"`
+ Tablespace string `json:"tablespace,omitempty"`
+ UsedBytes int64 `json:"usedBytes"`
+ CapacityBytes int64 `json:"capacityBytes"`
+ Ratio float64 `json:"ratio"`
+}
+
+// cnpgNamespaceDisk answers for every Cluster of one namespace with one claim
+// list and one usage batch.
+func (s *Reader) namespaceDisk(ctx context.Context, cache *k8s.ResourceCache, namespace string, clusters []*unstructured.Unstructured) []CNPGClusterDisk {
+ all := func(d CNPGClusterDisk) []CNPGClusterDisk {
+ out := make([]CNPGClusterDisk, len(clusters))
+ for i, c := range clusters {
+ d.Namespace, d.Name = namespace, c.GetName()
+ out[i] = d
+ }
+ return out
+ }
+ if !s.Access.CanRead(ctx, "", "persistentvolumeclaims", namespace, "list") {
+ return all(CNPGClusterDisk{State: storageStateDenied, Grant: cnpgGrantListPVCs.In(namespace).Ref()})
+ }
+ req, err := labels.NewRequirement(clusterLabel, selection.Exists, nil)
+ if err != nil {
+ return all(CNPGClusterDisk{State: cnpgStorageStateError, Reason: err.Error()})
+ }
+ candidates, reason := cnpgCachedClaims(cache, namespace, labels.NewSelector().Add(*req))
+ if reason != "" {
+ return all(CNPGClusterDisk{State: cnpgStorageStateUnavailable, Reason: reason})
+ }
+ owned := make([][]*corev1.PersistentVolumeClaim, len(clusters))
+ var names []string
+ for i, c := range clusters {
+ var mine []*corev1.PersistentVolumeClaim
+ for _, pvc := range candidates {
+ if pvc.Labels[clusterLabel] == c.GetName() {
+ mine = append(mine, pvc)
+ }
+ }
+ owned[i], _ = cnpgOwnedClaims(mine, c)
+ names = append(names, claimNames(owned[i])...)
+ }
+ var cov CNPGStorageCoverage
+ var batch prometheuspkg.PVCUsageBatch
+ if len(names) > 0 {
+ var anchors []prom.WorkloadPodIdentity
+ for _, c := range clusters {
+ if len(anchors) < cnpgHistoryAnchorCap {
+ anchors = append(anchors, historyAnchors(cache, c)...)
+ }
+ }
+ cov, batch = s.claimUsage(ctx, namespace, names, anchors)
+ }
+
+ out := make([]CNPGClusterDisk, len(clusters))
+ for i, c := range clusters {
+ d := CNPGClusterDisk{Namespace: namespace, Name: c.GetName(), Claims: len(owned[i]), Isolation: cov.Isolation}
+ if d.Claims == 0 {
+ d.State, d.Reason = usageStateNotRead, "no claims owned by this cluster"
+ out[i] = d
+ continue
+ }
+ for _, pvc := range owned[i] {
+ u, ok := batch.Usage[pvc.Name]
+ if !ok {
+ continue
+ }
+ d.Measured++
+ if d.Max == nil || u.Ratio > d.Max.Ratio {
+ d.Max = &CNPGDiskUsage{
+ Claim: pvc.Name, Instance: pvc.Labels[instanceNameLabel], Role: pvc.Labels[cnpgPVCRoleLabel],
+ Tablespace: pvc.Labels[cnpgTablespaceNameLabel], UsedBytes: u.UsedBytes, CapacityBytes: u.CapacityBytes, Ratio: u.Ratio,
+ }
+ }
+ }
+ switch {
+ case cov.State == cnpgUsageStateDenied || cov.State == cnpgUsageStateNoPrometheus || cov.State == cnpgUsageStateError ||
+ cov.State == cnpgHistoryStateAmbiguous || cov.State == cnpgHistoryStateScopeMismatch:
+ d.State, d.Grant, d.Reason = cov.State, cov.Grant, cov.Reason
+ case d.Measured == 0:
+ d.State, d.Reason = cnpgUsageStateNoSeries, "Prometheus has no kubelet volume stats for this cluster's claims"
+ case d.Measured < d.Claims:
+ d.State, d.Reason = cnpgStorageStatePartial, fmt.Sprintf("%d of %d claims measured", d.Measured, d.Claims)
+ default:
+ d.State = cnpgUsageStateOK
+ }
+ out[i] = d
+ }
+ return out
+}
+
+// cnpgWALCoverageReason counts instances whose WAL facts are missing
+// entirely apart from those missing one of the two sources, so the count
+// matches what each instance card shows.
+func cnpgWALCoverageReason(failed, partial, total int) string {
+ var parts []string
+ if failed > 0 {
+ parts = append(parts, fmt.Sprintf("%d of %d instances could not be read", failed, total))
+ }
+ if partial > 0 {
+ verb := "were"
+ if partial == 1 {
+ verb = "was"
+ }
+ parts = append(parts, fmt.Sprintf("%d of %d %s read only in part", partial, total, verb))
+ }
+ return strings.Join(parts, "; ")
+}
+
+func (s *Reader) FleetDisk(ctx context.Context, cache *k8s.ResourceCache, namespaces []string) CNPGFleetDiskResponse {
+ resp := CNPGFleetDiskResponse{SampledAt: time.Now().UTC().Format(time.RFC3339), Source: usageSource, Clusters: []CNPGClusterDisk{}}
+
+ var clusterKind integration.WorkspaceKind
+ for _, k := range workspaceKinds {
+ if k.Key == workspaceClusterKey {
+ clusterKind = k
+ }
+ }
+ acc, clusters := s.workspaceReadKind(ctx, cache, clusterKind, namespaces)
+ if acc.State != integration.KindCoverageFull && acc.State != integration.KindCoveragePartial {
+ return resp
+ }
+ byNamespace := map[string][]*unstructured.Unstructured{}
+ for _, c := range clusters {
+ byNamespace[c.GetNamespace()] = append(byNamespace[c.GetNamespace()], c)
+ }
+ nsList := make([]string, 0, len(byNamespace))
+ for ns := range byNamespace {
+ nsList = append(nsList, ns)
+ }
+ sort.Strings(nsList)
+
+ read := min(len(nsList), fleetDiskMaxNamespaces)
+ results := integration.FanOut(ctx, read, fleetDiskConcurrency, func(i int) []CNPGClusterDisk {
+ return s.namespaceDisk(ctx, cache, nsList[i], byNamespace[nsList[i]])
+ })
+ for _, rs := range results {
+ resp.Clusters = append(resp.Clusters, rs...)
+ }
+ for _, ns := range nsList[read:] {
+ for _, c := range byNamespace[ns] {
+ resp.Clusters = append(resp.Clusters, CNPGClusterDisk{Namespace: ns, Name: c.GetName(), State: usageStateNotRead,
+ Reason: fmt.Sprintf("disk use is read for at most %d namespaces at a time; narrow the namespace filter", fleetDiskMaxNamespaces)})
+ }
+ }
+ sort.Slice(resp.Clusters, func(i, j int) bool {
+ if resp.Clusters[i].Namespace != resp.Clusters[j].Namespace {
+ return resp.Clusters[i].Namespace < resp.Clusters[j].Namespace
+ }
+ return resp.Clusters[i].Name < resp.Clusters[j].Name
+ })
+ return resp
+
+}
diff --git a/internal/cnpg/storage_test.go b/internal/cnpg/storage_test.go
new file mode 100644
index 0000000000..b78e881e1a
--- /dev/null
+++ b/internal/cnpg/storage_test.go
@@ -0,0 +1,58 @@
+package cnpg
+
+import (
+ "strings"
+ "testing"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+)
+
+func TestCNPGDiskFindingsThresholds(t *testing.T) {
+ r := func(v float64) *float64 { return &v }
+ vols := []CNPGStorageVolume{
+ {Claim: "a", Role: cnpgPVCRoleData, Usage: CNPGStorageVolumeUsage{State: cnpgUsageStateOK, Ratio: r(0.79)}},
+ {Claim: "b", Role: cnpgPVCRoleWAL, Usage: CNPGStorageVolumeUsage{State: cnpgUsageStateOK, Ratio: r(0.80)}},
+ {Claim: "c", Role: cnpgPVCRoleTablespace, Tablespace: "archive", Usage: CNPGStorageVolumeUsage{State: cnpgUsageStateOK, Ratio: r(0.90)}},
+ {Claim: "d", Role: cnpgPVCRoleData, Usage: CNPGStorageVolumeUsage{State: cnpgUsageStateNoSeries}},
+ }
+ got := cnpgDiskFindings("pg-1", vols, nil)
+ if len(got) != 2 || got[0].Severity != "warning" || got[1].Severity != "critical" {
+ t.Fatalf("findings = %+v", got)
+ }
+ if got[1].Message != "The tablespace archive volume of pg-1 is 90% full" {
+ t.Errorf("message = %q", got[1].Message)
+ }
+ unverified := cnpgDiskFindings("pg-1", vols, &prometheuspkg.SeriesIsolation{Mode: prometheuspkg.SeriesIsolationUnverified, Note: "Radar couldn't confirm these volume stats belong to this exact cluster"})
+ if !strings.Contains(unverified[1].Message, "couldn't confirm these volume stats belong to this exact cluster") {
+ t.Errorf("an unverified match must say so: %q", unverified[1].Message)
+ }
+}
+
+func TestCNPGStorageExpansionDefaultsToStorageSize(t *testing.T) {
+ c := &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{"storage": map[string]any{"resizeInUseVolumes": false}}}}
+ got := cnpgStorageExpansionOf(c)
+ if len(got.Targets) != 1 || got.Targets[0].Field != "spec.storage.size" || got.Targets[0].Declared != "" {
+ t.Errorf("targets = %+v", got.Targets)
+ }
+ if got.ResizeInUseVolumes == nil || *got.ResizeInUseVolumes {
+ t.Errorf("resizeInUseVolumes = %v, want the declared false", got.ResizeInUseVolumes)
+ }
+}
+
+func TestCNPGWALCoverageReason(t *testing.T) {
+ cases := []struct {
+ failed, partial, total int
+ want string
+ }{
+ {3, 0, 3, "3 of 3 instances could not be read"},
+ {1, 2, 3, "1 of 3 instances could not be read; 2 of 3 were read only in part"},
+ {0, 1, 3, "1 of 3 was read only in part"},
+ }
+ for _, c := range cases {
+ if got := cnpgWALCoverageReason(c.failed, c.partial, c.total); got != c.want {
+ t.Errorf("(%d, %d, %d) = %q, want %q", c.failed, c.partial, c.total, got, c.want)
+ }
+ }
+}
diff --git a/internal/cnpg/test_ports_test.go b/internal/cnpg/test_ports_test.go
new file mode 100644
index 0000000000..27e3339e79
--- /dev/null
+++ b/internal/cnpg/test_ports_test.go
@@ -0,0 +1,46 @@
+package cnpg
+
+import (
+ "context"
+ "net/http"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func boolPtr(b bool) *bool { return &b }
+
+func newTestReader(perms *auth.UserPermissions) *Reader {
+ return &Reader{
+ Access: Access{
+ CanRead: func(context.Context, string, string, string, string) bool { return true },
+ Permission: func(_ context.Context, g auth.Grant) string {
+ if perms == nil {
+ return integration.PermissionAllowed
+ }
+ resource := g.Resource
+ if g.Subresource != "" {
+ resource += "/" + g.Subresource
+ }
+ allowed, known := perms.CanI(g.Verb, g.Group, resource, g.Namespace)
+ if !known {
+ return integration.PermissionUnknown
+ }
+ if allowed {
+ return integration.PermissionAllowed
+ }
+ return integration.PermissionDenied
+ },
+ },
+ Observations: Observations{
+ Cluster: func(context.Context, string, string, ...auth.Grant) (*k8s.ResourceCache, *unstructured.Unstructured, error) {
+ return nil, nil, &ReadFailure{Status: http.StatusServiceUnavailable, Message: "not connected"}
+ },
+ },
+ }
+}
+
+func int32Ptr(v int32) *int32 { return &v }
diff --git a/internal/cnpg/workspace.go b/internal/cnpg/workspace.go
new file mode 100644
index 0000000000..4d59599b9a
--- /dev/null
+++ b/internal/cnpg/workspace.go
@@ -0,0 +1,547 @@
+package cnpg
+
+import (
+ "context"
+ "log"
+ "sort"
+ "strings"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+ listersbatchv1 "k8s.io/client-go/listers/batch/v1"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+ bp "github.com/skyhook-io/radar/pkg/audit"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+ "github.com/skyhook-io/radar/pkg/issuesapi"
+ "github.com/skyhook-io/radar/pkg/topology"
+)
+
+const barmanGroup = "barmancloud.cnpg.io"
+
+const (
+ workspacePodsKey = "pods"
+ workspaceBackupsKey = "backups"
+ workspaceSchedKey = "scheduledBackups"
+ workspaceClusterKey = "clusters"
+)
+
+// cnpgBackupWindow bounds how far back settled Backups are returned. The newest
+// completed Backup per Cluster is kept regardless: it is the last-good-backup
+// fact the workspace reports.
+const backupWindow = 7 * 24 * time.Hour
+
+const noDeclarativeBackupCheckID = "cnpgNoDeclarativeBackup"
+
+// cnpgGroups are the API groups CloudNativePG objects are read from; a kind
+// listed by name keeps only objects of these groups.
+var groups = []string{Group, barmanGroup}
+
+var workspaceKinds = []integration.WorkspaceKind{
+ {Key: workspaceClusterKey, Group: Group, Kind: "Cluster", Resource: "clusters"},
+ {Key: workspaceBackupsKey, Group: Group, Kind: "Backup", Resource: "backups"},
+ {Key: workspaceSchedKey, Group: Group, Kind: "ScheduledBackup", Resource: "scheduledbackups"},
+ {Key: "poolers", Group: Group, Kind: "Pooler", Resource: "poolers"},
+ {Key: "databases", Group: Group, Kind: "Database", Resource: "databases"},
+ {Key: "publications", Group: Group, Kind: "Publication", Resource: "publications"},
+ {Key: "subscriptions", Group: Group, Kind: "Subscription", Resource: "subscriptions"},
+ {Key: "databaseRoles", Group: Group, Kind: "DatabaseRole", Resource: "databaseroles"},
+ {Key: "imageCatalogs", Group: Group, Kind: "ImageCatalog", Resource: "imagecatalogs"},
+ {Key: "clusterImageCatalogs", Group: Group, Kind: "ClusterImageCatalog", Resource: "clusterimagecatalogs", ClusterScoped: true},
+ {Key: "objectStores", Group: barmanGroup, Kind: "ObjectStore", Resource: "objectstores"},
+}
+
+// CNPGWorkspaceIssue is the subset of issuesapi.Issue the workspace renders.
+type CNPGWorkspaceIssue struct {
+ ID string `json:"id"`
+ Severity issuesapi.Severity `json:"severity"`
+ Category issuesapi.Category `json:"category"`
+ Kind string `json:"kind"`
+ Group string `json:"group,omitempty"`
+ Namespace string `json:"namespace,omitempty"`
+ Name string `json:"name"`
+ Reason string `json:"reason"`
+ Message string `json:"message,omitempty"`
+ Cause string `json:"cause,omitempty"`
+ Action string `json:"action,omitempty"`
+ FirstSeen time.Time `json:"first_seen,omitzero"`
+}
+
+// CNPGWorkspaceAuditFinding is one audit finding on a visible CNPG object.
+type CNPGWorkspaceAuditFinding struct {
+ CheckID string `json:"checkId"`
+ Severity string `json:"severity"`
+ Kind string `json:"kind"`
+ Group string `json:"group,omitempty"`
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+ Message string `json:"message"`
+}
+
+// CNPGWorkspaceResponse is GET /api/cnpg/workspace.
+type CNPGWorkspaceResponse struct {
+ Installed bool `json:"installed"`
+ Context string `json:"context"`
+ Namespaces []string `json:"namespaces"`
+ Coverage map[string]integration.KindCoverage `json:"coverage"`
+ Objects map[string][]any `json:"objects"`
+ Issues []CNPGWorkspaceIssue `json:"issues"`
+ Audit []CNPGWorkspaceAuditFinding `json:"audit"`
+ BackupsOmitted int `json:"backupsOmitted"`
+ // ScheduleReadings words each readable ScheduledBackup's schedule as the
+ // operator reads it, keyed "namespace/name"; a schedule the operator
+ // cannot parse has no entry.
+ ScheduleReadings map[string]string `json:"scheduleReadings,omitempty"`
+ // ManagedBy is the GitOps or Helm manager each returned CNPG object's
+ // labels and annotations name, keyed "Kind/namespace/name"; objects with
+ // no such signal have no entry. Instance Pods are not included.
+ ManagedBy map[string]topology.ResourceRef `json:"managedBy,omitempty"`
+ // JobPods are the Pods of Jobs a returned Cluster controls (initdb, join,
+ // restore…), kept apart from its instance Pods. JobCoverage says where the
+ // caller's Jobs were read: a Job Pod is returned only there.
+ JobPods []any `json:"jobPods,omitempty"`
+ JobCoverage *integration.KindCoverage `json:"jobCoverage,omitempty"`
+}
+
+func newCNPGWorkspaceResponse(namespaces []string, contextName string) CNPGWorkspaceResponse {
+ resp := CNPGWorkspaceResponse{
+ Context: contextName,
+ Namespaces: namespaces,
+ Coverage: map[string]integration.KindCoverage{},
+ Objects: map[string][]any{},
+ Issues: []CNPGWorkspaceIssue{},
+ Audit: []CNPGWorkspaceAuditFinding{},
+ }
+ for _, k := range workspaceKinds {
+ resp.Coverage[k.Key] = integration.KindCoverage{State: integration.KindCoverageNotInstalled}
+ resp.Objects[k.Key] = []any{}
+ }
+ resp.Coverage[workspacePodsKey] = integration.KindCoverage{State: integration.KindCoverageNotInstalled}
+ resp.Objects[workspacePodsKey] = []any{}
+ return resp
+}
+
+func (s *Reader) workspaceReadKind(ctx context.Context, cache *k8s.ResourceCache, k integration.WorkspaceKind, namespaces []string) (integration.KindAccess, []*unstructured.Unstructured) {
+ return s.Observations.WorkspaceRead(ctx, cache, k, namespaces, groups)
+}
+
+// filterCNPGGroup drops anything whose apiVersion is not a CNPG group, so a
+// Velero Backup or a CAPI Cluster can never ride along on a kind-name match.
+func filterCNPGGroup(items []*unstructured.Unstructured, err error) ([]*unstructured.Unstructured, error) {
+ if err != nil {
+ return nil, err
+ }
+ return integration.KeepGroups(items, groups), nil
+}
+
+func sortCNPGObjects(items []*unstructured.Unstructured) {
+ sort.SliceStable(items, func(i, j int) bool {
+ if items[i].GetNamespace() != items[j].GetNamespace() {
+ return items[i].GetNamespace() < items[j].GetNamespace()
+ }
+ return items[i].GetName() < items[j].GetName()
+ })
+}
+
+func backupTime(u *unstructured.Unstructured) time.Time {
+ for _, field := range []string{"stoppedAt", "startedAt"} {
+ if v, _, _ := unstructured.NestedString(u.Object, "status", field); v != "" {
+ if t, err := time.Parse(time.RFC3339, v); err == nil {
+ return t
+ }
+ }
+ }
+ return u.GetCreationTimestamp().Time
+}
+
+// windowCNPGBackups keeps every in-flight Backup, settled ones from the last
+// week, and each Cluster's newest completed Backup whatever its age. Sorted by
+// namespace, newest first within it.
+func windowCNPGBackups(items []*unstructured.Unstructured, now time.Time) ([]*unstructured.Unstructured, int) {
+ newestCompleted := map[string]*unstructured.Unstructured{}
+ for _, u := range items {
+ if phase, _, _ := unstructured.NestedString(u.Object, "status", "phase"); phase != "completed" {
+ continue
+ }
+ clusterName, _, _ := unstructured.NestedString(u.Object, "spec", "cluster", "name")
+ key := u.GetNamespace() + "\x00" + clusterName
+ if cur, ok := newestCompleted[key]; !ok || backupTime(u).After(backupTime(cur)) {
+ newestCompleted[key] = u
+ }
+ }
+ keepNewest := make(map[*unstructured.Unstructured]bool, len(newestCompleted))
+ for _, u := range newestCompleted {
+ keepNewest[u] = true
+ }
+
+ cutoff := now.Add(-backupWindow)
+ kept := make([]*unstructured.Unstructured, 0, len(items))
+ omitted := 0
+ for _, u := range items {
+ phase, _, _ := unstructured.NestedString(u.Object, "status", "phase")
+ settled := phase == "completed" || phase == "failed"
+ if !settled || keepNewest[u] || !backupTime(u).Before(cutoff) {
+ kept = append(kept, u)
+ continue
+ }
+ omitted++
+ }
+ sort.SliceStable(kept, func(i, j int) bool {
+ if kept[i].GetNamespace() != kept[j].GetNamespace() {
+ return kept[i].GetNamespace() < kept[j].GetNamespace()
+ }
+ ti, tj := backupTime(kept[i]), backupTime(kept[j])
+ if !ti.Equal(tj) {
+ return ti.After(tj)
+ }
+ return kept[i].GetName() < kept[j].GetName()
+ })
+ return kept, omitted
+}
+
+type WorkspacePodMeta struct {
+ Name string `json:"name"`
+ Namespace string `json:"namespace"`
+ UID types.UID `json:"uid"`
+ Labels map[string]string `json:"labels,omitempty"`
+ OwnerReferences []metav1.OwnerReference `json:"ownerReferences,omitempty"`
+ CreationTimestamp metav1.Time `json:"creationTimestamp"`
+}
+
+type WorkspaceContainerStatus struct {
+ Name string `json:"name"`
+ Ready bool `json:"ready"`
+ RestartCount int32 `json:"restartCount"`
+ State corev1.ContainerState `json:"state"`
+}
+
+type WorkspacePod struct {
+ APIVersion string `json:"apiVersion"`
+ Kind string `json:"kind"`
+ Metadata WorkspacePodMeta `json:"metadata"`
+ Spec struct {
+ NodeName string `json:"nodeName,omitempty"`
+ } `json:"spec"`
+ Status struct {
+ Phase corev1.PodPhase `json:"phase,omitempty"`
+ PodIP string `json:"podIP,omitempty"`
+ StartTime *metav1.Time `json:"startTime,omitempty"`
+ Conditions []corev1.PodCondition `json:"conditions,omitempty"`
+ ContainerStatuses []WorkspaceContainerStatus `json:"containerStatuses,omitempty"`
+ } `json:"status"`
+}
+
+// isCNPGInstancePod requires the controller ownerReference to name a visible
+// Cluster by UID, not just by name: a label alone is something any workload
+// can carry, and a Pod left behind by a deleted Cluster must not be attributed
+// to a new one created under the same name. clusterUIDs is keyed ns/name.
+func isCNPGInstancePod(p *corev1.Pod, clusterUIDs map[string]types.UID) bool {
+ name := p.Labels["cnpg.io/cluster"]
+ return cnpg.IsInstancePod(p, p.Namespace, name, clusterUIDs[p.Namespace+"/"+name])
+}
+
+func clusterUIDs(clusters []*unstructured.Unstructured) map[string]types.UID {
+ out := make(map[string]types.UID, len(clusters))
+ for _, c := range clusters {
+ out[c.GetNamespace()+"/"+c.GetName()] = c.GetUID()
+ }
+ return out
+}
+
+func trimCNPGPod(p *corev1.Pod) WorkspacePod {
+ out := WorkspacePod{APIVersion: "v1", Kind: "Pod"}
+ out.Metadata = WorkspacePodMeta{
+ Name: p.Name,
+ Namespace: p.Namespace,
+ UID: p.UID,
+ Labels: p.Labels,
+ OwnerReferences: p.OwnerReferences,
+ CreationTimestamp: p.CreationTimestamp,
+ }
+ out.Spec.NodeName = p.Spec.NodeName
+ out.Status.Phase = p.Status.Phase
+ out.Status.PodIP = p.Status.PodIP
+ out.Status.StartTime = p.Status.StartTime
+ out.Status.Conditions = p.Status.Conditions
+ for _, cs := range p.Status.ContainerStatuses {
+ out.Status.ContainerStatuses = append(out.Status.ContainerStatuses, WorkspaceContainerStatus{
+ Name: cs.Name, Ready: cs.Ready, RestartCount: cs.RestartCount, State: cs.State,
+ })
+ }
+ return out
+}
+
+// WorkspaceReadPods returns the instance Pods of visible Clusters and,
+// where the caller may also list Jobs, the Pods of the Jobs those Clusters
+// control, plus the namespace/name set of what it returned — the only Pods
+// whose issues the response may carry.
+func (s *Reader) workspaceReadPods(ctx context.Context, cache *k8s.ResourceCache, namespaces []string, clusterUIDs map[string]types.UID) (podAcc integration.KindAccess, out, jobOut []any, jobAcc integration.KindAccess, returned map[string]bool) {
+ out, jobOut, returned = []any{}, []any{}, map[string]bool{}
+ podAcc, read := s.Observations.TypedScope(ctx, cache, namespaces, "", "pods")
+ jobAcc, _ = s.Observations.TypedScope(ctx, cache, namespaces, "batch", "jobs")
+ if (jobAcc.State == integration.KindCoverageFull || jobAcc.State == integration.KindCoveragePartial) && (cache.Jobs() == nil || !cache.IsKindReady("jobs")) {
+ jobAcc = integration.KindAccess{State: integration.KindCoverageSyncing}
+ }
+ if podAcc.State == integration.KindCoverageDenied || podAcc.State == integration.KindCoverageUncached {
+ return podAcc, out, jobOut, jobAcc, returned
+ }
+ if cache.Pods() == nil {
+ log.Printf("[cnpg] Pod cache unavailable for workspace")
+ return integration.KindAccess{State: integration.KindCoverageError}, out, jobOut, jobAcc, returned
+ }
+
+ pods := integration.ListPodsScoped(cache.Pods(), read)
+ sort.Slice(pods, func(i, j int) bool {
+ if pods[i].Namespace != pods[j].Namespace {
+ return pods[i].Namespace < pods[j].Namespace
+ }
+ return pods[i].Name < pods[j].Name
+ })
+ for _, p := range pods {
+ switch {
+ case p == nil:
+ case isCNPGInstancePod(p, clusterUIDs):
+ out = append(out, trimCNPGPod(p))
+ returned[p.Namespace+"/"+p.Name] = true
+ case jobAcc.Covers(p.Namespace) && isCNPGClusterJobPod(p, clusterUIDs, cache.Jobs()):
+ jobOut = append(jobOut, trimCNPGPod(p))
+ returned[p.Namespace+"/"+p.Name] = true
+ }
+ }
+ return podAcc, out, jobOut, jobAcc, returned
+}
+
+// isCNPGClusterJobPod reports a Pod of a Job its Cluster controls: the Pod's
+// controller is that exact Job (name and UID), and the Job's controller is the
+// Cluster (name and UID). Labels alone never adopt a Pod, and a Job recreated
+// under the same name does not adopt the previous Job's Pods.
+func isCNPGClusterJobPod(p *corev1.Pod, clusterUIDs map[string]types.UID, jobs listersbatchv1.JobLister) bool {
+ clusterName := p.Labels[clusterLabel]
+ if clusterName == "" || p.Labels[jobRoleLabel] == "" || jobs == nil {
+ return false
+ }
+ uid, ok := clusterUIDs[p.Namespace+"/"+clusterName]
+ if !ok || uid == "" {
+ return false
+ }
+ ref := controllerRef(p.OwnerReferences)
+ if ref == nil || ref.Kind != "Job" {
+ return false
+ }
+ job, err := jobs.Jobs(p.Namespace).Get(ref.Name)
+ if err != nil || job == nil {
+ return false
+ }
+ return controlledBy(p.OwnerReferences, "batch", "Job", job.Name, job.UID) &&
+ controlledBy(job.OwnerReferences, Group, "Cluster", clusterName, uid)
+}
+
+var cnpgWorkspaceKeyByGroupKind = func() map[string]string {
+ m := make(map[string]string, len(workspaceKinds))
+ for _, k := range workspaceKinds {
+ m[k.Group+"/"+k.Kind] = k.Key
+ }
+ return m
+}()
+
+// workspaceIssues runs the same composition /api/issues serves, but reads
+// the flat evidence rows: the grouped view folds instance-Pod evidence into
+// the owning Cluster's row, which would hand Pod failure detail to a caller
+// who may list Clusters but not Pods. A row is kept only when its own subject
+// is visible here — a CNPG kind covered in its namespace, or an instance or
+// Cluster Job Pod this response returned. IDs are the subject-derived IDs /api/issues uses.
+func (s *Reader) workspaceIssues(ctx context.Context, namespaces []string, access map[string]integration.KindAccess, returnedPods map[string]bool) []CNPGWorkspaceIssue {
+ out := []CNPGWorkspaceIssue{}
+ if integration.NoNamespaceAccess(namespaces) {
+ return out
+ }
+ composed := s.Observations.Issues(ctx, namespaces)
+ for _, iss := range composed {
+ if !workspaceIssueVisible(iss, access, returnedPods) {
+ continue
+ }
+ out = append(out, CNPGWorkspaceIssue{
+ ID: iss.ID,
+ Severity: iss.Severity,
+ Category: iss.Category,
+ Kind: iss.Kind,
+ Group: iss.Group,
+ Namespace: iss.Namespace,
+ Name: iss.Name,
+ Reason: iss.Reason,
+ Message: iss.Message,
+ Cause: iss.Cause,
+ Action: iss.Action,
+ FirstSeen: iss.FirstSeen,
+ })
+ }
+ return out
+}
+
+func workspaceIssueVisible(iss issues.Issue, access map[string]integration.KindAccess, returnedPods map[string]bool) bool {
+ if iss.Group == "" && iss.Kind == "Pod" {
+ return access[workspacePodsKey].Covers(iss.Namespace) && returnedPods[iss.Namespace+"/"+iss.Name]
+ }
+ if iss.Group != Group && iss.Group != barmanGroup {
+ return false
+ }
+ for _, read := range iss.RequiredReads {
+ key := ""
+ if read.Group == "" && read.Resource == "pods" {
+ key = workspacePodsKey
+ }
+ for _, kind := range workspaceKinds {
+ if kind.Group == read.Group && kind.Resource == read.Resource {
+ key = kind.Key
+ break
+ }
+ }
+ if key == "" || !access[key].Covers(read.Namespace) {
+ return false
+ }
+ }
+ key, ok := cnpgWorkspaceKeyByGroupKind[iss.Group+"/"+iss.Kind]
+
+ return ok && access[key].Covers(iss.Namespace)
+}
+
+// cnpgWorkspaceAudit reports the declarative-backup posture finding only for
+// Clusters whose namespace had its ScheduledBackups read: without that list,
+// "no schedule targets this cluster" is an absence nobody established.
+func (s *Reader) workspaceAudit(clusters, scheduled []*unstructured.Unstructured, schedAccess integration.KindAccess) []CNPGWorkspaceAuditFinding {
+ out := []CNPGWorkspaceAuditFinding{}
+ var subjects []*unstructured.Unstructured
+ for _, c := range clusters {
+ if schedAccess.Covers(c.GetNamespace()) {
+ subjects = append(subjects, c)
+ }
+ }
+ if len(subjects) == 0 {
+ return out
+ }
+ results := bp.RunChecks(&bp.CheckInput{
+ CNPGClusters: subjects,
+ CNPGScheduledBackups: scheduled,
+ CNPGScheduledBackupsAuthoritative: true,
+ })
+ results = s.Observations.FilterAudit(results)
+ if results == nil {
+ return out
+ }
+ for _, f := range results.Findings {
+ if f.CheckID != noDeclarativeBackupCheckID {
+ continue
+ }
+ out = append(out, CNPGWorkspaceAuditFinding{
+ CheckID: f.CheckID,
+ Severity: f.Severity,
+ Kind: f.Kind,
+ Group: f.Group,
+ Namespace: f.Namespace,
+ Name: f.Name,
+ Message: f.Message,
+ })
+ }
+ sort.SliceStable(out, func(i, j int) bool {
+ if out[i].Namespace != out[j].Namespace {
+ return out[i].Namespace < out[j].Namespace
+ }
+ return out[i].Name < out[j].Name
+ })
+ return out
+}
+
+func scheduleReadings(scheduled []*unstructured.Unstructured) map[string]string {
+ out := map[string]string{}
+ for _, sb := range scheduled {
+ spec, _, _ := unstructured.NestedString(sb.Object, "spec", "schedule")
+ if strings.TrimSpace(spec) == "" || len(spec) > scheduleMaxLen {
+ continue
+ }
+ if _, err := cnpg.ParseSchedule(spec); err != nil {
+ continue
+ }
+ if reading := describeCNPGSchedule(spec); reading != "" {
+ out[sb.GetNamespace()+"/"+sb.GetName()] = reading
+ }
+ }
+ return out
+}
+
+func (s *Reader) Workspace(ctx context.Context, cache *k8s.ResourceCache, namespaces []string, contextName string) CNPGWorkspaceResponse {
+ resp := newCNPGWorkspaceResponse(namespaces, contextName)
+
+ disc := s.Observations.Discovery
+ if disc != nil {
+ for _, k := range workspaceKinds {
+ if _, ok := disc.GetGVRWithGroup(k.Kind, k.Group); ok {
+ resp.Installed = true
+ break
+ }
+ }
+ if !resp.Installed {
+ return resp
+ }
+ }
+
+ access := map[string]integration.KindAccess{}
+ items := map[string][]*unstructured.Unstructured{}
+ for _, k := range workspaceKinds {
+ if disc != nil {
+ if _, ok := disc.GetGVRWithGroup(k.Kind, k.Group); !ok {
+ access[k.Key] = integration.KindAccess{State: integration.KindCoverageNotInstalled}
+ continue
+ }
+ }
+ acc, list := s.workspaceReadKind(ctx, cache, k, namespaces)
+ if acc.State != integration.KindCoverageNotInstalled {
+ resp.Installed = true
+ }
+ access[k.Key] = acc
+ items[k.Key] = list
+ resp.Coverage[k.Key] = acc.Coverage()
+ }
+ if !resp.Installed {
+ return resp
+ }
+
+ for _, k := range workspaceKinds {
+ list := items[k.Key]
+ if k.Key == workspaceBackupsKey {
+ var omitted int
+ list, omitted = windowCNPGBackups(list, time.Now())
+ resp.BackupsOmitted = omitted
+ } else {
+ sortCNPGObjects(list)
+ }
+ out := make([]any, 0, len(list))
+ for _, u := range list {
+ out = append(out, u.Object)
+ if ref := topology.ManagedByFromMeta(u); ref != nil {
+ if resp.ManagedBy == nil {
+ resp.ManagedBy = map[string]topology.ResourceRef{}
+ }
+ resp.ManagedBy[k.Kind+"/"+u.GetNamespace()+"/"+u.GetName()] = *ref
+ }
+ }
+ resp.Objects[k.Key] = out
+ }
+
+ podAccess, pods, jobPods, jobAccess, returnedPods := s.workspaceReadPods(ctx, cache, namespaces, clusterUIDs(items[workspaceClusterKey]))
+ jobCov := jobAccess.Coverage()
+ resp.JobPods, resp.JobCoverage = jobPods, &jobCov
+ access[workspacePodsKey] = podAccess
+ resp.Coverage[workspacePodsKey] = podAccess.Coverage()
+ resp.Objects[workspacePodsKey] = pods
+
+ resp.Issues = s.workspaceIssues(ctx, namespaces, access, returnedPods)
+ resp.Audit = s.workspaceAudit(items[workspaceClusterKey], items[workspaceSchedKey], access[workspaceSchedKey])
+ resp.ScheduleReadings = scheduleReadings(items[workspaceSchedKey])
+
+ return resp
+}
diff --git a/internal/cnpg/workspace_jobpods_test.go b/internal/cnpg/workspace_jobpods_test.go
new file mode 100644
index 0000000000..c2bd3cb674
--- /dev/null
+++ b/internal/cnpg/workspace_jobpods_test.go
@@ -0,0 +1,84 @@
+package cnpg
+
+import (
+ "testing"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/types"
+ listersbatchv1 "k8s.io/client-go/listers/batch/v1"
+ "k8s.io/client-go/tools/cache"
+)
+
+func cnpgJob(ns, name, uid string, owner metav1.OwnerReference) *batchv1.Job {
+ return &batchv1.Job{ObjectMeta: metav1.ObjectMeta{
+ Namespace: ns, Name: name, UID: types.UID(uid),
+ Labels: map[string]string{clusterLabel: owner.Name, jobRoleLabel: "initdb"},
+ OwnerReferences: []metav1.OwnerReference{owner},
+ }}
+}
+
+func clusterRef(name, uid string) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: name, UID: types.UID(uid), Controller: boolPtr(true)}
+}
+
+func jobRef(name, uid string) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "batch/v1", Kind: "Job", Name: name, UID: types.UID(uid), Controller: boolPtr(true)}
+}
+
+// An initdb Pod the scheduler could not place, as CloudNativePG labels it.
+func cnpgJobPod(ns, name, cluster string, owner metav1.OwnerReference) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{
+ Namespace: ns, Name: name,
+ Labels: map[string]string{clusterLabel: cluster, jobRoleLabel: "initdb", "cnpg.io/instanceName": cluster + "-1"},
+ OwnerReferences: []metav1.OwnerReference{owner},
+ },
+ Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "initdb", Image: "pg:17"}}},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodPending,
+ Conditions: []corev1.PodCondition{{
+ Type: corev1.PodScheduled, Status: corev1.ConditionFalse, Reason: corev1.PodReasonUnschedulable,
+ Message: "0/2 nodes are available: 2 Too many pods.",
+ }},
+ },
+ }
+}
+
+func TestIsCNPGClusterJobPod(t *testing.T) {
+ indexer := cache.NewIndexer(cache.MetaNamespaceKeyFunc, cache.Indexers{cache.NamespaceIndex: cache.MetaNamespaceIndexFunc})
+ for _, j := range []*batchv1.Job{
+ cnpgJob("pg", "x-1-initdb", "job-uid", clusterRef("x", "x-uid")),
+ cnpgJob("pg", "y-1-initdb", "y-job-uid", clusterRef("y", "y-uid")),
+ cnpgJob("pg", "x-2-join", "join-uid", metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "x", UID: "x-uid"}),
+ } {
+ if err := indexer.Add(j); err != nil {
+ t.Fatal(err)
+ }
+ }
+ jobs := listersbatchv1.NewJobLister(indexer)
+ uids := map[string]types.UID{"pg/x": "x-uid", "pg/y": "y-uid"}
+ unlabelled := cnpgJobPod("pg", "x-1-initdb-a", "x", jobRef("x-1-initdb", "job-uid"))
+ delete(unlabelled.Labels, jobRoleLabel)
+ for _, c := range []struct {
+ name string
+ pod *corev1.Pod
+ want bool
+ }{
+ {"Pod of a Job the Cluster controls", cnpgJobPod("pg", "x-1-initdb-a", "x", jobRef("x-1-initdb", "job-uid")), true},
+ {"left by an earlier Job of the same name", cnpgJobPod("pg", "x-1-initdb-a", "x", jobRef("x-1-initdb", "old-job-uid")), false},
+ {"Job controlled by another Cluster", cnpgJobPod("pg", "y-1-initdb-a", "x", jobRef("y-1-initdb", "y-job-uid")), false},
+ {"Job the Cluster owns but does not control", cnpgJobPod("pg", "x-2-join-a", "x", jobRef("x-2-join", "join-uid")), false},
+ {"controller is not a batch Job", cnpgJobPod("pg", "x-1-initdb-a", "x", metav1.OwnerReference{APIVersion: "example.com/v1", Kind: "Job", Name: "x-1-initdb", UID: "job-uid", Controller: boolPtr(true)}), false},
+ {"Job not cached", cnpgJobPod("pg", "x-9-initdb-a", "x", jobRef("x-9-initdb", "job-uid")), false},
+ {"no job role label", unlabelled, false},
+ {"Cluster not visible", cnpgJobPod("other", "x-1-initdb-a", "x", jobRef("x-1-initdb", "job-uid")), false},
+ } {
+ t.Run(c.name, func(t *testing.T) {
+ if got := isCNPGClusterJobPod(c.pod, uids, jobs); got != c.want {
+ t.Errorf("isCNPGClusterJobPod = %v, want %v", got, c.want)
+ }
+ })
+ }
+}
diff --git a/internal/cnpg/workspace_test.go b/internal/cnpg/workspace_test.go
new file mode 100644
index 0000000000..6031682f4b
--- /dev/null
+++ b/internal/cnpg/workspace_test.go
@@ -0,0 +1,114 @@
+package cnpg
+
+import (
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+)
+
+func cnpgObj(apiVersion, kind, ns, name string, spec, status map[string]any) *unstructured.Unstructured {
+ meta := map[string]any{"name": name, "creationTimestamp": time.Now().Add(-time.Hour).UTC().Format(time.RFC3339)}
+ if ns != "" {
+ meta["namespace"] = ns
+ }
+ obj := map[string]any{"apiVersion": apiVersion, "kind": kind, "metadata": meta}
+ if spec != nil {
+ obj["spec"] = spec
+ }
+ if status != nil {
+ obj["status"] = status
+ }
+ return &unstructured.Unstructured{Object: obj}
+}
+
+func cnpgBackup(ns, name, cluster, phase string, stoppedAt time.Time) *unstructured.Unstructured {
+ status := map[string]any{"phase": phase}
+ if !stoppedAt.IsZero() {
+ status["stoppedAt"] = stoppedAt.UTC().Format(time.RFC3339)
+ }
+ return cnpgObj("postgresql.cnpg.io/v1", "Backup", ns, name, map[string]any{"cluster": map[string]any{"name": cluster}}, status)
+}
+
+func cnpgPod(ns, name, clusterLabel string, owners ...metav1.OwnerReference) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{
+ Name: name, Namespace: ns,
+ Labels: map[string]string{"cnpg.io/cluster": clusterLabel, "cnpg.io/instanceRole": "primary"},
+ OwnerReferences: owners,
+ },
+ Spec: corev1.PodSpec{NodeName: "node-1", Containers: []corev1.Container{{Name: "postgres", Image: "pg:17"}}},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodRunning,
+ PodIP: "10.0.0.5",
+ ContainerStatuses: []corev1.ContainerStatus{{
+ Name: "postgres", Ready: true, RestartCount: 2, Image: "pg:17",
+ State: corev1.ContainerState{Running: &corev1.ContainerStateRunning{}},
+ }},
+ },
+ }
+}
+
+func TestWindowCNPGBackups(t *testing.T) {
+ now := time.Date(2026, 9, 29, 12, 0, 0, 0, time.UTC)
+ day := 24 * time.Hour
+ items := []*unstructured.Unstructured{
+ cnpgBackup("pg", "running-old", "orders", "running", time.Time{}),
+ cnpgBackup("pg", "orders-recent", "orders", "completed", now.Add(-2*day)),
+ cnpgBackup("pg", "orders-old", "orders", "completed", now.Add(-20*day)),
+ cnpgBackup("pg", "orders-failed-old", "orders", "failed", now.Add(-30*day)),
+ cnpgBackup("pg", "orders-failed-new", "orders", "failed", now.Add(-1*day)),
+ cnpgBackup("pg", "billing-only-old", "billing", "completed", now.Add(-40*day)),
+ cnpgBackup("pg", "billing-older", "billing", "completed", now.Add(-50*day)),
+ cnpgBackup("aa", "other-ns", "orders", "completed", now.Add(-60*day)),
+ }
+ // A running Backup older than the window stays: in flight is never settled.
+ items[0].Object["metadata"].(map[string]any)["creationTimestamp"] = now.Add(-90 * day).Format(time.RFC3339)
+
+ kept, omitted := windowCNPGBackups(items, now)
+ var names []string
+ for _, u := range kept {
+ names = append(names, u.GetName())
+ }
+ want := []string{"other-ns", "orders-failed-new", "orders-recent", "billing-only-old", "running-old"}
+ if len(names) != len(want) {
+ t.Fatalf("kept = %v, want %v", names, want)
+ }
+ for i := range want {
+ if names[i] != want[i] {
+ t.Fatalf("kept = %v, want %v (namespace, then newest first)", names, want)
+ }
+ }
+ if omitted != 3 {
+ t.Errorf("omitted = %d, want 3 (orders-old, orders-failed-old, billing-older)", omitted)
+ }
+}
+
+func TestIsCNPGInstancePod(t *testing.T) {
+ uids := map[string]types.UID{"pg/x": "x-uid"}
+ ref := func(uid string, controller bool) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "x", UID: types.UID(uid), Controller: boolPtr(controller)}
+ }
+ for _, c := range []struct {
+ name string
+ pod *corev1.Pod
+ uids map[string]types.UID
+ want bool
+ }{
+ {"controller ref to the visible Cluster", cnpgPod("pg", "x-1", "x", ref("x-uid", true)), uids, true},
+ {"left behind by a deleted Cluster of the same name", cnpgPod("pg", "x-1", "x", ref("old-uid", true)), uids, false},
+ {"non-controller owner", cnpgPod("pg", "x-1", "x", ref("x-uid", false)), uids, false},
+ {"Cluster not visible", cnpgPod("pg", "x-1", "x", ref("x-uid", true)), map[string]types.UID{}, false},
+ {"Cluster of that name in another namespace", cnpgPod("other", "x-1", "x", ref("x-uid", true)), uids, false},
+ {"no cluster label", func() *corev1.Pod { p := cnpgPod("pg", "x-1", "x", ref("x-uid", true)); p.Labels = nil; return p }(), uids, false},
+ } {
+ t.Run(c.name, func(t *testing.T) {
+ if got := isCNPGInstancePod(c.pod, c.uids); got != c.want {
+ t.Errorf("isCNPGInstancePod = %v, want %v", got, c.want)
+ }
+ })
+ }
+}
diff --git a/internal/imageutil/image.go b/internal/imageutil/image.go
new file mode 100644
index 0000000000..c1caf15247
--- /dev/null
+++ b/internal/imageutil/image.go
@@ -0,0 +1,22 @@
+package imageutil
+
+import (
+ "strings"
+)
+
+// Untagged references do not imply a version; an explicit tag remains useful
+// even when a digest pins the image.
+func ImageTag(image string) string {
+ if image == "" {
+ return ""
+ }
+ if at := strings.Index(image, "@"); at >= 0 {
+ image = image[:at]
+ }
+ slash := strings.LastIndex(image, "/")
+ colon := strings.LastIndex(image, ":")
+ if colon > slash {
+ return image[colon+1:]
+ }
+ return ""
+}
diff --git a/internal/integration/actions.go b/internal/integration/actions.go
new file mode 100644
index 0000000000..495ad2dd24
--- /dev/null
+++ b/internal/integration/actions.go
@@ -0,0 +1,188 @@
+package integration
+
+import (
+ "bytes"
+ "context"
+ "encoding/json"
+ "errors"
+ "fmt"
+ "net/http"
+ "strings"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/dynamic"
+
+ "github.com/skyhook-io/radar/internal/auth"
+)
+
+const (
+ ActionBodyLimit = 64 << 10
+
+ ActionCodeChanged = "changed"
+ ActionCodeContextChanged = "context_changed"
+ ActionCodeBlocked = "blocked"
+ ActionCodePartial = "partial"
+
+ PermissionAllowed = "allowed"
+ PermissionDenied = "denied"
+ PermissionUnknown = "unknown"
+)
+
+// ActionCapability is one action's verdict: allowed only when the state
+// guards pass and the caller's grant is not known to be missing. An unknown
+// grant (the SubjectAccessReview itself failed) leaves the apiserver to decide.
+type ActionCapability struct {
+ Allowed bool `json:"allowed"`
+ Reason string `json:"reason,omitempty"`
+ ReasonCode string `json:"reasonCode,omitempty"`
+ Permission string `json:"permission"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+}
+
+// ActionRequest is the POST body of every action.
+type ActionRequest struct {
+ ReviewedContext string `json:"reviewedContext"`
+ UID string `json:"uid"`
+ Facts json.RawMessage `json:"facts,omitempty"`
+ Params json.RawMessage `json:"params,omitempty"`
+}
+
+// actionError is a refusal with a stable code; Current carries the facts as
+// they are now when a confirmation no longer matches.
+type ActionError struct {
+ Status int
+ Code string
+ Message string
+ Current any
+ // Completed lists the mutations that took effect before a multi-step
+ // action stopped (code partial).
+ Completed []string
+}
+
+func (e *ActionError) Error() string { return e.Message }
+
+func RefuseAction(status int, code, format string, args ...any) *ActionError {
+ return &ActionError{Status: status, Code: code, Message: fmt.Sprintf(format, args...)}
+}
+
+// changedAction refuses a confirmation whose reviewed facts no longer hold.
+func ChangedAction(current any, format string, args ...any) *ActionError {
+ e := RefuseAction(http.StatusConflict, ActionCodeChanged, format, args...)
+ e.Current = current
+ return e
+}
+
+func BlockedAction(reason string) error {
+ return RefuseAction(http.StatusConflict, ActionCodeBlocked, "%s", reason)
+}
+
+// partialAction reports a multi-step action that stopped after some of its
+// mutations took effect; retrying it blindly would act on a changed target.
+func PartialAction(completed []string, cause error) *ActionError {
+ status := http.StatusInternalServerError
+ var ae *ActionError
+ switch {
+ case errors.As(cause, &ae):
+ status = ae.Status
+ case apierrors.IsForbidden(cause):
+ status = http.StatusForbidden
+ case apierrors.IsConflict(cause), apierrors.IsNotFound(cause):
+ status = http.StatusConflict
+ case errors.Is(cause, context.DeadlineExceeded) || apierrors.IsTimeout(cause) || apierrors.IsServerTimeout(cause):
+ status = http.StatusGatewayTimeout
+ }
+ e := &ActionError{
+ Status: status, Code: ActionCodePartial, Completed: append([]string(nil), completed...),
+ Message: fmt.Sprintf("Stopped part-way: %s. Already done: %s", cause.Error(), strings.Join(completed, ", ")),
+ }
+ if ae != nil {
+ e.Current = ae.Current
+ }
+ return e
+}
+
+func CheckReviewedContext(reviewed, active string) error {
+ if reviewed != active {
+ return RefuseAction(http.StatusConflict, ActionCodeContextChanged,
+ "The active cluster context is %q, not the %q you reviewed; review the action again", active, reviewed)
+ }
+ return nil
+}
+
+// decodeActionParams decodes an action's params strictly: an unknown field is
+// a client that means something this server does not do.
+func DecodeActionParams(raw json.RawMessage, into any) error {
+ trimmed := bytes.TrimSpace(raw)
+ if len(trimmed) == 0 || string(trimmed) == "null" {
+ trimmed = []byte("{}")
+ }
+ dec := json.NewDecoder(bytes.NewReader(trimmed))
+ dec.DisallowUnknownFields()
+ if err := dec.Decode(into); err != nil {
+ return RefuseAction(http.StatusBadRequest, "", "invalid params: %v", err)
+ }
+ return nil
+}
+
+// mergePatchAtVersion merge-patches obj bound to the resourceVersion it was
+// read at, so a write the caller did not review is refused with a Conflict
+// rather than applied over. Callers turn the Conflict into "changed"; it is
+// never retried.
+func MergePatchAtVersion(ctx context.Context, dyn dynamic.Interface, gvr schema.GroupVersionResource, obj *unstructured.Unstructured, body map[string]any, subresources ...string) error {
+ meta, _ := body["metadata"].(map[string]any)
+ if meta == nil {
+ meta = map[string]any{}
+ body["metadata"] = meta
+ }
+ meta["resourceVersion"] = obj.GetResourceVersion()
+ data, err := json.Marshal(body)
+ if err != nil {
+ return err
+ }
+ _, err = dyn.Resource(gvr).Namespace(obj.GetNamespace()).Patch(ctx, obj.GetName(), types.MergePatchType, data, metav1.PatchOptions{}, subresources...)
+ return err
+}
+
+// grantText words an optional grant, "" when there is none.
+func GrantText(g *auth.Grant) string {
+ if g == nil {
+ return ""
+ }
+ return g.String()
+}
+
+// capabilityVerdict folds a guard reason and the permission of each grant
+// (perms[i] answers grants[i]) into a verdict. The first denied grant is
+// named; an unknown grant leaves the action offered.
+func CapabilityVerdict(guard string, perms []string, grants []auth.Grant) ActionCapability {
+ out := ActionCapability{Permission: PermissionAllowed}
+ for i, p := range perms {
+ if p == PermissionDenied {
+ out.Permission = PermissionDenied
+ out.Grant = grants[i].Ref()
+ break
+ }
+ if p == PermissionUnknown {
+ out.Permission = PermissionUnknown
+ if out.Grant == nil {
+ out.Grant = grants[i].Ref()
+ }
+ }
+ }
+ if out.Permission == PermissionAllowed {
+ out.Grant = nil
+ }
+ switch {
+ case out.Permission == PermissionDenied:
+ out.Reason = "You are not allowed to " + out.Grant.String()
+ case guard != "":
+ out.Reason = guard
+ default:
+ out.Allowed = true
+ }
+ return out
+}
diff --git a/internal/integration/cache_scope.go b/internal/integration/cache_scope.go
new file mode 100644
index 0000000000..eee438de7f
--- /dev/null
+++ b/internal/integration/cache_scope.go
@@ -0,0 +1,92 @@
+package integration
+
+import (
+ "slices"
+)
+
+// informerScope is the part of Radar's informer cache that says which
+// namespaces it holds for a resource.
+type InformerScope interface {
+ IsKindClusterWide(string) bool
+ KindNamespaces(string) []string
+ IsKindReady(string) bool
+}
+
+type CacheNamespaceResult struct {
+ Namespaces []string
+ Limited bool
+ Partial bool
+ Unavailable bool
+}
+
+// namespacesWithinCache narrows requested (nil = all namespaces) to what the
+// informer for resource holds. Limited: the informer is not cluster-wide;
+// Partial: some requested namespaces are not held; Unavailable: none are.
+func NamespacesWithinCache(cache InformerScope, resource string, requested []string) CacheNamespaceResult {
+ result := CacheNamespaceResult{Namespaces: requested}
+ if cache == nil || cache.IsKindClusterWide(resource) {
+ return result
+ }
+ result.Limited = true
+ if NoNamespaceAccess(requested) {
+ return result
+ }
+ cached := cache.KindNamespaces(resource)
+ if len(cached) == 0 {
+ result.Namespaces = []string{}
+ result.Unavailable = true
+ return result
+ }
+ result.Namespaces = IntersectNamespaces(cached, requested)
+ if requested == nil {
+ result.Partial = true
+ return result
+ }
+ for _, namespace := range requested {
+ if !slices.Contains(cached, namespace) {
+ if len(result.Namespaces) == 0 {
+ result.Unavailable = true
+ } else {
+ result.Partial = true
+ }
+ return result
+ }
+ }
+ return result
+}
+
+func CacheCoversNamespace(cache InformerScope, resource, namespace string) bool {
+ return cache != nil && cache.IsKindReady(resource) && (cache.IsKindClusterWide(resource) || slices.Contains(cache.KindNamespaces(resource), namespace))
+}
+
+// intersectNamespaces returns the namespaces to actually scan. nil `allowed`
+// means the user is unrestricted; preserve `requested` (which may also be nil
+// for cluster-wide). When the user is restricted, keep only the requested
+// namespaces they're allowed to see; if `requested` is empty, fall back to
+// the full allowed set.
+func IntersectNamespaces(allowed, requested []string) []string {
+ if allowed == nil {
+ return requested
+ }
+ if len(requested) == 0 {
+ return allowed
+ }
+ allowSet := make(map[string]struct{}, len(allowed))
+ for _, ns := range allowed {
+ allowSet[ns] = struct{}{}
+ }
+ out := make([]string, 0, len(requested))
+ for _, ns := range requested {
+ if _, ok := allowSet[ns]; ok {
+ out = append(out, ns)
+ }
+ }
+ return out
+}
+
+// noNamespaceAccess returns true when a namespace filter explicitly grants no access
+// (non-nil empty slice from auth filtering). Handlers with custom namespace logic
+// should check this and return empty results.
+func NoNamespaceAccess(namespaces []string) bool {
+ return namespaces != nil && len(namespaces) == 0
+}
diff --git a/internal/integration/coverage.go b/internal/integration/coverage.go
new file mode 100644
index 0000000000..19726b80f4
--- /dev/null
+++ b/internal/integration/coverage.go
@@ -0,0 +1,76 @@
+package integration
+
+import (
+ "sort"
+)
+
+// Coverage states for one kind a workspace lists.
+const (
+ KindCoverageFull = "full"
+ KindCoveragePartial = "partial"
+ KindCoverageDenied = "denied"
+ // kindCoverageUncached: the caller may list the kind, but Radar's informer
+ // holds none of the Namespaces in scope, so nothing was read.
+ KindCoverageUncached = "uncached"
+ KindCoverageNotInstalled = "notInstalled"
+ KindCoverageSyncing = "syncing"
+ KindCoverageError = "error"
+)
+
+// KindCoverage states how much of one kind the caller could see.
+// DeniedNamespaces (the caller may not list there) and UncachedNamespaces
+// (the caller may, but Radar's informer does not hold them) list only
+// Namespaces already in the caller's scope, so either may be omitted on a
+// partial State; AllowedNamespaces is always set on a partial State and is
+// the authority for which Namespaces were read.
+type KindCoverage struct {
+ State string `json:"state"`
+ DeniedNamespaces []string `json:"deniedNamespaces,omitempty"`
+ UncachedNamespaces []string `json:"uncachedNamespaces,omitempty"`
+ AllowedNamespaces []string `json:"allowedNamespaces,omitempty"`
+}
+
+// kindAccess is the resolved read scope for one kind. All means every
+// namespace in the request's scope (or the cluster-scoped kind itself).
+// Denied and Uncached name the in-scope Namespaces left unread, under the
+// disclosure rule of listScope.
+type KindAccess struct {
+ State string
+ All bool
+ Namespaces map[string]bool
+ Denied []string
+ Uncached []string
+}
+
+func (a KindAccess) Covers(namespace string) bool {
+ if a.State != KindCoverageFull && a.State != KindCoveragePartial {
+ return false
+ }
+ return a.All || a.Namespaces[namespace]
+}
+
+func (a KindAccess) Coverage() KindCoverage {
+ cov := KindCoverage{State: a.State, DeniedNamespaces: a.Denied, UncachedNamespaces: a.Uncached}
+ if a.State == KindCoveragePartial {
+ cov.AllowedNamespaces = make([]string, 0, len(a.Namespaces))
+ for ns := range a.Namespaces {
+ cov.AllowedNamespaces = append(cov.AllowedNamespaces, ns)
+ }
+ sort.Strings(cov.AllowedNamespaces)
+ }
+ return cov
+}
+
+func AccessFromScope(allowed []string, partial bool) KindAccess {
+ acc := KindAccess{State: KindCoverageFull, All: allowed == nil}
+ if partial {
+ acc.State = KindCoveragePartial
+ }
+ if allowed != nil {
+ acc.Namespaces = make(map[string]bool, len(allowed))
+ for _, ns := range allowed {
+ acc.Namespaces[ns] = true
+ }
+ }
+ return acc
+}
diff --git a/internal/integration/coverage_test.go b/internal/integration/coverage_test.go
new file mode 100644
index 0000000000..87bf2e6db1
--- /dev/null
+++ b/internal/integration/coverage_test.go
@@ -0,0 +1,21 @@
+package integration
+
+import (
+ "encoding/json"
+ "strings"
+ "testing"
+)
+
+func TestKindCoverageJSON(t *testing.T) {
+ b, err := json.Marshal(KindAccess{State: KindCoverageUncached, Uncached: []string{"b"}}.Coverage())
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got := string(b); got != `{"state":"uncached","uncachedNamespaces":["b"]}` {
+ t.Errorf("json = %s", got)
+ }
+ b, _ = json.Marshal(KindAccess{State: KindCoverageFull, All: true}.Coverage())
+ if strings.Contains(string(b), "Namespaces") {
+ t.Errorf("full coverage carries namespace lists: %s", b)
+ }
+}
diff --git a/internal/integration/fanout.go b/internal/integration/fanout.go
new file mode 100644
index 0000000000..6abec3a88c
--- /dev/null
+++ b/internal/integration/fanout.go
@@ -0,0 +1,32 @@
+package integration
+
+import (
+ "context"
+ "sync"
+)
+
+// fanOut calls run(i) for every i in [0, n), at most limit at a time, and
+// returns the results in index order once every call has finished. A call
+// still waiting for a slot when ctx is done does not run; its result is the
+// zero value. Caps on n, ordering and what a skipped index means stay with
+// the caller.
+func FanOut[T any](ctx context.Context, n, limit int, run func(i int) T) []T {
+ results := make([]T, n)
+ sem := make(chan struct{}, limit)
+ var wg sync.WaitGroup
+ for i := range n {
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ select {
+ case sem <- struct{}{}:
+ case <-ctx.Done():
+ return
+ }
+ defer func() { <-sem }()
+ results[i] = run(i)
+ }()
+ }
+ wg.Wait()
+ return results
+}
diff --git a/internal/integration/groups.go b/internal/integration/groups.go
new file mode 100644
index 0000000000..631ad3c0f0
--- /dev/null
+++ b/internal/integration/groups.go
@@ -0,0 +1,21 @@
+package integration
+
+import (
+ "slices"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+)
+
+// keepGroups drops anything whose apiVersion is not one of groups, so an
+// object of another group can never ride along on a kind-name match (a Velero
+// Backup on a CloudNativePG Backup, a CAPI Cluster on a CNPG Cluster).
+func KeepGroups(items []*unstructured.Unstructured, groups []string) []*unstructured.Unstructured {
+ out := items[:0:0]
+ for _, u := range items {
+ if u == nil || !slices.Contains(groups, u.GroupVersionKind().Group) {
+ continue
+ }
+ out = append(out, u)
+ }
+ return out
+}
diff --git a/internal/integration/pods.go b/internal/integration/pods.go
new file mode 100644
index 0000000000..45ca44ee50
--- /dev/null
+++ b/internal/integration/pods.go
@@ -0,0 +1,28 @@
+package integration
+
+import (
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/labels"
+ v1listers "k8s.io/client-go/listers/core/v1"
+)
+
+// listPodsScoped lists pods either cluster-wide (namespaces nil) or across
+// the caller's allowed namespaces — the scoping shape every metrics/vitals
+// consumer shares.
+func ListPodsScoped(podLister v1listers.PodLister, namespaces []string) []*corev1.Pod {
+ if podLister == nil {
+ return nil
+ }
+ // Sentinel contract (parseNamespacesForUser): nil = all namespaces;
+ // non-nil EMPTY = no namespace access — zero pods, never cluster-wide.
+ if namespaces == nil {
+ pods, _ := podLister.List(labels.Everything())
+ return pods
+ }
+ var pods []*corev1.Pod
+ for _, ns := range namespaces {
+ items, _ := podLister.Pods(ns).List(labels.Everything())
+ pods = append(pods, items...)
+ }
+ return pods
+}
diff --git a/internal/integration/reads.go b/internal/integration/reads.go
new file mode 100644
index 0000000000..3909214c45
--- /dev/null
+++ b/internal/integration/reads.go
@@ -0,0 +1,20 @@
+package integration
+
+import (
+ "errors"
+
+ "github.com/skyhook-io/radar/internal/auth"
+)
+
+// ReadSource is the outcome of one read a response depends on. auth.Grant names
+// what a denied read needs; Reason explains any other outcome. The State
+// vocabulary belongs to the caller: sources classify the same failure
+// differently on purpose (a timed-out HA read is unavailable, a timed-out
+// recovery read is an error).
+type ReadSource struct {
+ State string `json:"state"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Reason string `json:"reason,omitempty"`
+}
+
+var ErrDynamicNotSynced = errors.New("resource cache not synced")
diff --git a/internal/integration/workspace.go b/internal/integration/workspace.go
new file mode 100644
index 0000000000..27cfdb5820
--- /dev/null
+++ b/internal/integration/workspace.go
@@ -0,0 +1,10 @@
+package integration
+
+// workspaceKind is one Kind a workspace lists from Radar's dynamic cache.
+type WorkspaceKind struct {
+ Key string
+ Group string
+ Kind string
+ Resource string
+ ClusterScoped bool
+}
diff --git a/internal/issues/category.go b/internal/issues/category.go
index 567fbd309b..d2c629d1ac 100644
--- a/internal/issues/category.go
+++ b/internal/issues/category.go
@@ -149,8 +149,10 @@ func Classify(in classifyInput) issuesapi.Category {
return issuesapi.CategoryBackupFailed
}
switch in.Reason {
- case "CNPGWALArchivingFailing", "CNPGLastBackupFailed", "CNPGBackupFailed":
+ case "CNPGWALArchivingFailing", "CNPGLastBackupFailed", "CNPGBackupFailed", ReasonCNPGScheduledRunNoBackup:
return issuesapi.CategoryBackupFailed
+ case ReasonCNPGCertificateExpiring, ReasonCNPGCertificateExpired:
+ return issuesapi.CategoryCertificateNotReady
}
// A declared object that never reached PostgreSQL is a
// reconciliation failure, not a backup one — the operator tried and
diff --git a/internal/issues/compose.go b/internal/issues/compose.go
index 2ef27d76c6..e434c95082 100644
--- a/internal/issues/compose.go
+++ b/internal/issues/compose.go
@@ -173,6 +173,7 @@ func ComposeWithStats(p Provider, f Filters) ([]Issue, ComposeStats) {
// symptoms against parent rollups across member pods — both need the flat
// rows, so they run BEFORE grouping and BEFORE the public filters.
out = applyClusterScopedAccess(out, f)
+ out = filterEvidenceAccess(out, f.CanReadEvidence, f.AllowUnfilteredEvidence)
out = redactUnreadableRelatedRefs(out, f.CanReadRelated)
out = dedupePodSchedulingOverProblem(out)
// Same-resource structural-root → symptom: fold a pod's runtime symptom into
diff --git a/internal/issues/evidence_access.go b/internal/issues/evidence_access.go
new file mode 100644
index 0000000000..e156681857
--- /dev/null
+++ b/internal/issues/evidence_access.go
@@ -0,0 +1,25 @@
+package issues
+
+// Findings with inventory dependencies require an authorizer. A shared internal
+// memo may explicitly retain them for authorization at its per-user projection.
+func CanReadIssueEvidence(issue Issue, canRead func(EvidenceRead) bool, allowUnfiltered bool) bool {
+ if canRead == nil {
+ return allowUnfiltered || len(issue.RequiredReads) == 0
+ }
+ for _, read := range issue.RequiredReads {
+ if !canRead(read) {
+ return false
+ }
+ }
+ return true
+}
+
+func filterEvidenceAccess(list []Issue, canRead func(EvidenceRead) bool, allowUnfiltered bool) []Issue {
+ out := list[:0]
+ for _, issue := range list {
+ if CanReadIssueEvidence(issue, canRead, allowUnfiltered) {
+ out = append(out, issue)
+ }
+ }
+ return out
+}
diff --git a/internal/issues/grouping.go b/internal/issues/grouping.go
index 011bf0488e..f3b6ebdae4 100644
--- a/internal/issues/grouping.go
+++ b/internal/issues/grouping.go
@@ -19,11 +19,13 @@ func RelatedIssues(p Provider, opts RelatedIssueOptions, group, kind, namespace,
// not the grouped issue's inline Members (capped at maxInlineMembers) — is
// what makes member #11..#N in a large fan-out resolve correctly.
flat := Compose(p, Filters{
- SkipPodTemplateContext: true,
- Namespaces: opts.Namespaces,
- Limit: NoLimit,
- CanReadClusterScoped: opts.CanReadClusterScoped,
- CanReadRelated: opts.CanReadRelated,
+ SkipPodTemplateContext: true,
+ Namespaces: opts.Namespaces,
+ Limit: NoLimit,
+ CanReadClusterScoped: opts.CanReadClusterScoped,
+ CanReadRelated: opts.CanReadRelated,
+ CanReadEvidence: opts.CanReadEvidence,
+ AllowUnfilteredEvidence: opts.AllowUnfilteredEvidence,
})
grouped := GroupIssues(flat)
// Run the grouped-mode enrichment (mirrors the cluster path) so the grouped
@@ -57,7 +59,7 @@ func RelatedIssuesFrom(flat, grouped []Issue, opts RelatedIssueOptions, group, k
}
matched := make(map[string]bool) // grouped issue IDs the resource touches
for _, g := range grouped { // as the grouped SUBJECT (owner-collapsed)
- if match(g.Group, g.Kind, g.Namespace, g.Name) && CanReadIssueRelatedRefs(g, opts.CanReadRelated) {
+ if match(g.Group, g.Kind, g.Namespace, g.Name) && CanReadIssueRelatedRefs(g, opts.CanReadRelated) && CanReadIssueEvidence(g, opts.CanReadEvidence, opts.AllowUnfilteredEvidence) {
matched[g.ID] = true
}
if g.DiagnosticContext != nil {
@@ -77,13 +79,13 @@ func RelatedIssuesFrom(flat, grouped []Issue, opts RelatedIssueOptions, group, k
}
}
for _, f := range flat { // as ANY evidence row (uncapped)
- if match(f.Group, f.Kind, f.Namespace, f.Name) && CanReadIssueRelatedRefs(f, opts.CanReadRelated) {
+ if match(f.Group, f.Kind, f.Namespace, f.Name) && CanReadIssueRelatedRefs(f, opts.CanReadRelated) && CanReadIssueEvidence(f, opts.CanReadEvidence, opts.AllowUnfilteredEvidence) {
matched[f.ID] = true
}
}
var out []Issue
for _, g := range grouped {
- if matched[g.ID] {
+ if matched[g.ID] && CanReadIssueEvidence(g, opts.CanReadEvidence, opts.AllowUnfilteredEvidence) {
out = append(out, g)
}
}
@@ -169,6 +171,15 @@ func foldGroup(members []Issue) Issue {
DiagnosticContext: rep.DiagnosticContext,
ChangeContext: rep.ChangeContext,
}
+ reads := map[EvidenceRead]bool{}
+ for _, member := range members {
+ for _, read := range member.RequiredReads {
+ if !reads[read] {
+ reads[read] = true
+ g.RequiredReads = append(g.RequiredReads, read)
+ }
+ }
+ }
// A parsed diagnosis (cause/action/remediation) describes ONE resource's
// failure. Carry it onto the grouped row only when it is true for the
// entire group: a single-member group carries its own diagnosis, but a
diff --git a/internal/issues/issues_test.go b/internal/issues/issues_test.go
index ac40338254..d2707e4b5c 100644
--- a/internal/issues/issues_test.go
+++ b/internal/issues/issues_test.go
@@ -39,6 +39,7 @@ type fakeProvider struct {
change map[string]*issuesapi.ChangeContext
webhookRefs map[string][]AdmissionWebhookRef
workloadBacks map[string]bool
+ listErr map[schema.GroupVersionResource]error
}
type secretProducerResult struct {
@@ -59,6 +60,9 @@ func (f *fakeProvider) WatchedDynamic() []schema.GroupVersionResource {
return out
}
func (f *fakeProvider) ListDynamic(gvr schema.GroupVersionResource, _ string) ([]*unstructured.Unstructured, error) {
+ if err := f.listErr[gvr]; err != nil {
+ return nil, err
+ }
return f.dynamic[gvr], nil
}
func (f *fakeProvider) ListDynamicAllNamespaces(gvr schema.GroupVersionResource) ([]*unstructured.Unstructured, error) {
diff --git a/internal/issues/options.go b/internal/issues/options.go
index 673cbf8565..7a60641f9f 100644
--- a/internal/issues/options.go
+++ b/internal/issues/options.go
@@ -41,6 +41,12 @@ type Filters struct {
// assert an issue on another subject, such as a NodeClass referenced by a
// NodePool. Nil preserves internal/no-auth composition.
CanReadRelated func(Ref) bool
+ // CanReadEvidence authorizes the exact inventory operations retained on
+ // cross-resource findings, including after grouping or memoization.
+ CanReadEvidence func(EvidenceRead) bool
+ // AllowUnfilteredEvidence is only for internal composition whose results
+ // are authorized again before reaching a caller.
+ AllowUnfilteredEvidence bool
// Grouped folds the flat rows into the public grouped model
// (GroupIssues) before the cap, so the limit counts issue groups, not
// replica fan-out. The public /api/issues + MCP issues set this; flat
@@ -62,8 +68,10 @@ type RelatedIssueOptions struct {
// folded into per-issue signals such as capacity relevance. Omitting it here
// would let a caller who cannot list NodePools read that state through the
// per-resource path. Nil preserves auth-mode=none and internal composition.
- CanReadClusterScoped func(kind, group string) bool
- CanReadRelated func(Ref) bool
+ CanReadClusterScoped func(kind, group string) bool
+ CanReadRelated func(Ref) bool
+ CanReadEvidence func(EvidenceRead) bool
+ AllowUnfilteredEvidence bool
}
const (
diff --git a/internal/issues/provider.go b/internal/issues/provider.go
index f1dcc613bc..7a96bf4a72 100644
--- a/internal/issues/provider.go
+++ b/internal/issues/provider.go
@@ -16,6 +16,7 @@ import (
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/timeline"
+ "github.com/skyhook-io/radar/pkg/cnpg"
"github.com/skyhook-io/radar/pkg/issuesapi"
)
@@ -830,3 +831,21 @@ func (p *CacheProvider) karpenterResources(group, kind string) karpenterResource
}
return karpenterResourceInventory{GVR: gvr, Items: items, Coverage: karpenterCoverageAuthoritative}
}
+
+func (p *CacheProvider) cnpgInstancePods(cluster *unstructured.Unstructured) ([]*corev1.Pod, bool) {
+ ns := cluster.GetNamespace()
+ if p == nil || p.cache == nil || p.cache.Pods() == nil || !p.cache.IsKindReady("pods") || !p.cache.KindCoversNamespace("pods", ns) {
+ return nil, false
+ }
+ candidates, err := p.cache.Pods().Pods(ns).List(labels.SelectorFromSet(labels.Set{"cnpg.io/cluster": cluster.GetName()}))
+ if err != nil {
+ return nil, false
+ }
+ var pods []*corev1.Pod
+ for _, pod := range candidates {
+ if cnpg.IsInstancePod(pod, ns, cluster.GetName(), cluster.GetUID()) {
+ pods = append(pods, pod)
+ }
+ }
+ return pods, true
+}
diff --git a/internal/issues/provider_cnpg_test.go b/internal/issues/provider_cnpg_test.go
new file mode 100644
index 0000000000..539f604d77
--- /dev/null
+++ b/internal/issues/provider_cnpg_test.go
@@ -0,0 +1,81 @@
+package issues
+
+import (
+ "errors"
+ "testing"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/cnpg"
+ "github.com/skyhook-io/radar/pkg/k8score"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/client-go/kubernetes/fake"
+ k8stesting "k8s.io/client-go/testing"
+)
+
+func TestCNPGPodInventoryUsesCacheScopeAndOwnership(t *testing.T) {
+ cluster := cnpgCluster(nil, nil)
+ cluster.SetUID("current")
+ controller := true
+ pod := &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: "pg-main-1", Namespace: "pg", Labels: map[string]string{"cnpg.io/cluster": "pg-main"}, OwnerReferences: []metav1.OwnerReference{{APIVersion: cnpg.Group + "/v1", Kind: "Cluster", Name: "pg-main", UID: "current", Controller: &controller}}}}
+ stale := pod.DeepCopy()
+ stale.Name = "old"
+ stale.OwnerReferences[0].UID = "old"
+ job := pod.DeepCopy()
+ job.Name = "backup-job"
+ job.OwnerReferences[0].Kind = "Job"
+ for _, scoped := range []bool{false, true} {
+ t.Run(map[bool]string{false: "default cluster scope", true: "namespace scope"}[scoped], func(t *testing.T) {
+ k8s.ResetResourceCache()
+ t.Cleanup(k8s.ResetResourceCache)
+ client := fake.NewClientset(pod, stale, job)
+ var err error
+ if scoped {
+ err = k8s.InitScopedTestResourceCache(client, map[string]k8score.ResourceScope{"pods": {Enabled: true, Namespace: "pg"}})
+ } else {
+ err = k8s.InitTestResourceCache(client)
+ }
+ if err != nil {
+ t.Fatal(err)
+ }
+ p := &CacheProvider{cache: k8s.GetResourceCache()}
+ instances, known := p.cnpgInstancePods(cluster)
+ if !known || len(instances) != 1 || instances[0].Name != pod.Name {
+ t.Fatalf("inventory=%v known=%v", instances, known)
+ }
+ other := cluster.DeepCopy()
+ other.SetNamespace("other")
+ _, known = p.cnpgInstancePods(other)
+ if known == scoped {
+ t.Fatalf("other namespace known=%v, scoped=%v", known, scoped)
+ }
+ })
+ }
+ client := fake.NewClientset()
+ release := make(chan struct{})
+ client.PrependReactor("list", "pods", func(k8stesting.Action) (bool, runtime.Object, error) {
+ select {
+ case <-release:
+ return false, nil, nil
+ default:
+ return true, nil, errors.New("initial list pending")
+ }
+ })
+ startupKnown := false
+ core, err := k8score.NewResourceCache(k8score.CacheConfig{
+ Client: client, ResourceTypes: map[string]bool{"pods": true},
+ OnInformersStarted: func(core *k8score.ResourceCache) {
+ p := &CacheProvider{cache: &k8s.ResourceCache{ResourceCache: core}}
+ _, startupKnown = p.cnpgInstancePods(cluster)
+ close(release)
+ },
+ })
+ if err != nil {
+ t.Fatal(err)
+ }
+ t.Cleanup(core.Stop)
+ if startupKnown {
+ t.Fatal("unsynced cache claimed a known inventory")
+ }
+}
diff --git a/internal/issues/source_cnpg.go b/internal/issues/source_cnpg.go
index 867f5f4dc6..92b2b143aa 100644
--- a/internal/issues/source_cnpg.go
+++ b/internal/issues/source_cnpg.go
@@ -1,10 +1,10 @@
package issues
import (
- "encoding/json"
"fmt"
"time"
+ "github.com/skyhook-io/radar/pkg/cnpg"
"github.com/skyhook-io/radar/pkg/conditions"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/runtime/schema"
@@ -136,10 +136,10 @@ func cnpgSuppressesGenericConditions(group, kind string, u *unstructured.Unstruc
func detectCNPGIssues(gvr schema.GroupVersionResource, kind string, u *unstructured.Unstructured) []Issue {
switch kind {
case "Cluster":
- return detectCNPGClusterIssues(gvr, kind, u)
+ return append(detectCNPGClusterIssues(gvr, kind, u), detectCNPGCertificateIssues(gvr, kind, u, time.Now())...)
case "Backup":
return detectCNPGBackupIssues(gvr, kind, u)
- case "Database", "Publication", "Subscription":
+ case "Database", "Publication", "Subscription", "DatabaseRole":
return detectCNPGDeclarativeIssues(gvr, kind, u)
case "ScheduledBackup":
return detectCNPGScheduledBackupIssues(gvr, kind, u)
@@ -230,24 +230,9 @@ func cnpgHibernated(u *unstructured.Unstructured) bool {
// array of instance names; "*" fences them all, so nothing serves — by intent,
// not fault. A malformed value counts as not-fenced rather than erroring.
func cnpgFencedAll(u *unstructured.Unstructured) bool {
- raw := u.GetAnnotations()["cnpg.io/fencedInstances"]
- if raw == "" {
- return false
- }
- var names []string
- if err := json.Unmarshal([]byte(raw), &names); err != nil {
- return false
- }
- for _, n := range names {
- if n == "*" {
- return true
- }
- }
- return false
+ return cnpg.ParseFencedInstances(u.GetAnnotations()["cnpg.io/fencedInstances"]).All
}
-// cnpgHasConditions mirrors the frontend's `status.conditions?.length` half of
-// the reported check — any condition the operator wrote proves it reconciled.
func cnpgHasConditions(u *unstructured.Unstructured) bool {
conds, _, _ := unstructured.NestedSlice(u.Object, "status", "conditions")
return len(conds) > 0
diff --git a/internal/issues/source_cnpg_certs.go b/internal/issues/source_cnpg_certs.go
new file mode 100644
index 0000000000..88c42e9326
--- /dev/null
+++ b/internal/issues/source_cnpg_certs.go
@@ -0,0 +1,136 @@
+package issues
+
+import (
+ "fmt"
+ "sort"
+ "strings"
+ "time"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+)
+
+// CNPGCertExpiryLayout is Go's time.Time.String() layout, which is how the
+// operator writes status.certificates.expirations.
+const CNPGCertExpiryLayout = "2006-01-02 15:04:05.999999999 -0700 MST"
+
+// ParseCNPGCertExpiry reads one status.certificates.expirations value. A
+// monotonic-clock suffix ("m=+0.1") can follow Time.String(); it is dropped.
+// ok is false when the value does not parse: the expiry is then unknown.
+
+// A certificate past its expiry and one approaching it are separate reasons
+// so a consumer can word them without reading the message.
+const (
+ ReasonCNPGCertificateExpiring = "CNPGCertificateExpiring"
+ ReasonCNPGCertificateExpired = "CNPGCertificateExpired"
+)
+
+func ParseCNPGCertExpiry(v string) (time.Time, bool) {
+ v = strings.TrimSpace(v)
+ if i := strings.Index(v, " m="); i >= 0 {
+ v = v[:i]
+ }
+ if v == "" {
+ return time.Time{}, false
+ }
+ t, err := time.Parse(CNPGCertExpiryLayout, v)
+ if err != nil {
+ return time.Time{}, false
+ }
+ return t, true
+}
+
+// CNPGUserCertificateSecrets returns the Secrets a Cluster names in
+// spec.certificates: their owner renews them. Every other Secret listed in
+// status.certificates.expirations was generated by the operator, which renews
+// it itself before expiry.
+func CNPGUserCertificateSecrets(u *unstructured.Unstructured) map[string]string {
+ out := map[string]string{}
+ for _, f := range []string{"serverCASecret", "serverTLSSecret", "replicationTLSSecret", "clientCASecret"} {
+ if name, _, _ := unstructured.NestedString(u.Object, "spec", "certificates", f); name != "" {
+ if prev, ok := out[name]; ok {
+ out[name] = prev + "," + f
+ } else {
+ out[name] = f
+ }
+ }
+ }
+ return out
+}
+
+const (
+ cnpgCertWarnWithin = 30 * 24 * time.Hour
+ cnpgCertCriticalWithin = 7 * 24 * time.Hour
+ // The operator renews its own certificates when they are within its
+ // expiring threshold (7 days by default, EXPIRING_CHECK_THRESHOLD). An
+ // operator-managed certificate this close to expiry means renewal has not
+ // happened; before that point it is routine and raising it would light
+ // every cluster for a third of each 90-day certificate lifetime.
+ cnpgOperatorCertOverdueWithin = 24 * time.Hour
+)
+
+// detectCNPGCertificateIssues reports certificates close to expiry. One issue
+// per Secret, so two expiring certificates stay two causes.
+func detectCNPGCertificateIssues(gvr schema.GroupVersionResource, kind string, u *unstructured.Unstructured, now time.Time) []Issue {
+ exp, _, _ := unstructured.NestedStringMap(u.Object, "status", "certificates", "expirations")
+ if len(exp) == 0 {
+ return nil
+ }
+ user := CNPGUserCertificateSecrets(u)
+ secrets := make([]string, 0, len(exp))
+ for s := range exp {
+ secrets = append(secrets, s)
+ }
+ sort.Strings(secrets)
+
+ var out []Issue
+ ns, name := u.GetNamespace(), u.GetName()
+ for _, secret := range secrets {
+ at, ok := ParseCNPGCertExpiry(exp[secret])
+ if !ok {
+ continue
+ }
+ left := at.Sub(now)
+ _, userProvided := user[secret]
+ var sev Severity
+ var msg string
+ reason := ReasonCNPGCertificateExpiring
+ switch {
+ case left <= 0:
+ sev, reason = SeverityCritical, ReasonCNPGCertificateExpired
+ msg = fmt.Sprintf("The certificate in Secret %s expired %s", secret, at.UTC().Format(time.RFC3339))
+ case userProvided && left < cnpgCertCriticalWithin:
+ sev = SeverityCritical
+ msg = fmt.Sprintf("The certificate in Secret %s expires in %s (%s); its owner renews it, not the operator", secret, cnpgDaysLeft(left), at.UTC().Format(time.RFC3339))
+ case userProvided && left < cnpgCertWarnWithin:
+ sev = SeverityWarning
+ msg = fmt.Sprintf("The certificate in Secret %s expires in %s (%s); its owner renews it, not the operator", secret, cnpgDaysLeft(left), at.UTC().Format(time.RFC3339))
+ case !userProvided && left < cnpgOperatorCertOverdueWithin:
+ sev = SeverityCritical
+ msg = fmt.Sprintf("The operator-managed certificate in Secret %s expires in %s and has not been renewed", secret, cnpgDaysLeft(left))
+ default:
+ continue
+ }
+ out = append(out, newConditionIssue(gvr, kind, ns, name, sev,
+ reason, msg, time.Time{}, false,
+ // One fingerprint across both reasons: expiring and then expired is
+ // the same finding growing worse, not a new one.
+ "CNPGCertificateExpiring/"+secret, u.GetCreationTimestamp().Time))
+ }
+ return out
+}
+
+func cnpgDaysLeft(d time.Duration) string {
+ if d < 24*time.Hour {
+ h := int(d.Hours())
+ if h < 1 {
+ return "under an hour"
+ }
+ return fmt.Sprintf("%dh", h)
+ }
+ days := int(d.Hours() / 24)
+ if days == 1 {
+ return "1 day"
+ }
+ return fmt.Sprintf("%d days", days)
+}
diff --git a/internal/issues/source_cnpg_certs_test.go b/internal/issues/source_cnpg_certs_test.go
new file mode 100644
index 0000000000..69d4fabb6b
--- /dev/null
+++ b/internal/issues/source_cnpg_certs_test.go
@@ -0,0 +1,96 @@
+package issues
+
+import (
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/issuesapi"
+)
+
+func TestParseCNPGCertExpiry(t *testing.T) {
+ want := time.Date(2026, 12, 28, 8, 28, 45, 0, time.UTC)
+ for _, in := range []string{
+ "2026-12-28 08:28:45 +0000 UTC",
+ "2026-12-28 08:28:45.000000 +0000 UTC",
+ "2026-12-28 08:28:45 +0000 UTC m=+7776000.000000001",
+ "2026-12-28 10:28:45 +0200 EET",
+ } {
+ got, ok := ParseCNPGCertExpiry(in)
+ if !ok || !got.Equal(want) {
+ t.Errorf("ParseCNPGCertExpiry(%q) = %v, %v; want %v", in, got, ok, want)
+ }
+ }
+ for _, in := range []string{"", "soon", "2026-12-28T08:28:45Z"} {
+ if _, ok := ParseCNPGCertExpiry(in); ok {
+ t.Errorf("ParseCNPGCertExpiry(%q) parsed; want unknown", in)
+ }
+ }
+}
+
+func TestCNPGCertificateExpiryIssues(t *testing.T) {
+ now := time.Date(2026, 9, 30, 12, 0, 0, 0, time.UTC)
+ at := func(d time.Duration) string { return now.Add(d).Format(CNPGCertExpiryLayout) }
+ u := cnpgCluster(
+ map[string]any{"certificates": map[string]any{"serverTLSSecret": "user-tls", "serverCASecret": "user-ca", "clientCASecret": "user-client-ca"}},
+ map[string]any{"certificates": map[string]any{"expirations": map[string]any{
+ "user-tls": at(20 * 24 * time.Hour), // warning
+ "user-ca": at(3 * 24 * time.Hour), // critical: owner must renew now
+ "user-client-ca": at(90 * 24 * time.Hour), // fine
+ "pg-main-replication": at(5 * 24 * time.Hour), // operator renews: routine
+ "pg-main-server": at(2 * time.Hour), // operator renewal overdue
+ "pg-main-ca": at(-time.Hour), // expired
+ "garbage": "not a time", // unknown: no issue
+ }}},
+ )
+ got := detectCNPGCertificateIssues(cnpgClusterGVR, "Cluster", u, now)
+ bySecret := map[string]Issue{}
+ for _, i := range got {
+ bySecret[strings.TrimPrefix(i.Fingerprint, "CNPGCertificateExpiring/")] = i
+ }
+ want := map[string]issuesapi.Severity{
+ "user-tls": SeverityWarning,
+ "user-ca": SeverityCritical,
+ "pg-main-server": SeverityCritical,
+ "pg-main-ca": SeverityCritical,
+ }
+ if len(bySecret) != len(want) {
+ t.Fatalf("issues for %v; want %v", keysOf(bySecret), want)
+ }
+ for secret, sev := range want {
+ i, ok := bySecret[secret]
+ wantReason := ReasonCNPGCertificateExpiring
+ if secret == "pg-main-ca" {
+ wantReason = ReasonCNPGCertificateExpired
+ }
+ if !ok || i.Severity != sev || i.Reason != wantReason {
+ t.Errorf("%s: %+v, want %s %s", secret, i, sev, wantReason)
+ }
+ }
+ if bySecret["pg-main-ca"].Category != issuesapi.CategoryCertificateNotReady {
+ t.Errorf("expired category = %q", bySecret["pg-main-ca"].Category)
+ }
+ if !strings.Contains(bySecret["user-tls"].Message, "its owner renews it") {
+ t.Errorf("user-provided message = %q", bySecret["user-tls"].Message)
+ }
+ if bySecret["user-tls"].Category != issuesapi.CategoryCertificateNotReady {
+ t.Errorf("category = %q", bySecret["user-tls"].Category)
+ }
+}
+
+func TestCNPGDatabaseRoleNotAppliedIsReported(t *testing.T) {
+ u := cnpgCluster(nil, map[string]any{"applied": false, "message": "database role is already managed by the CNPG cluster"})
+ u.SetKind("DatabaseRole")
+ got := detectCNPGIssues(cnpgClusterGVR, "DatabaseRole", u)
+ if len(got) != 1 || got[0].Reason != "CNPGDeclarativeNotApplied" || !strings.Contains(got[0].Message, "already managed") {
+ t.Fatalf("got %+v", got)
+ }
+}
+
+func keysOf(m map[string]Issue) []string {
+ out := make([]string, 0, len(m))
+ for k := range m {
+ out = append(out, k)
+ }
+ return out
+}
diff --git a/internal/issues/source_cnpg_observed.go b/internal/issues/source_cnpg_observed.go
new file mode 100644
index 0000000000..65ff3c38b2
--- /dev/null
+++ b/internal/issues/source_cnpg_observed.go
@@ -0,0 +1,151 @@
+package issues
+
+import (
+ "fmt"
+ "sort"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/cnpg"
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+)
+
+const (
+ ReasonCNPGScheduleDestinationMissing = "CNPGScheduleDestinationMissing"
+ ReasonCNPGInstanceReadinessMismatch = "CNPGInstanceReadinessMismatch"
+ ReasonCNPGPrimaryLabelMismatch = "CNPGPrimaryLabelMismatch"
+)
+
+type cnpgPodInventoryProvider interface {
+ cnpgInstancePods(*unstructured.Unstructured) ([]*corev1.Pod, bool)
+}
+
+func detectCNPGScheduleDestinationIssues(p Provider, gvr schema.GroupVersionResource, schedules []*unstructured.Unstructured) []Issue {
+ var clusterGVR schema.GroupVersionResource
+ for _, watched := range p.WatchedDynamic() {
+ if watched.Group == cnpgGroup && p.KindForGVR(watched) == "Cluster" {
+ clusterGVR = watched
+ break
+ }
+ }
+ if clusterGVR.Resource == "" {
+ return nil
+ }
+ clustersByNamespace := map[string]map[string]*unstructured.Unstructured{}
+ var out []Issue
+ for _, schedule := range schedules {
+ suspended, _, _ := unstructured.NestedBool(schedule.Object, "spec", "suspend")
+ if suspended {
+ continue
+ }
+ ns := schedule.GetNamespace()
+ clusters, loaded := clustersByNamespace[ns]
+ if !loaded {
+ clusters = map[string]*unstructured.Unstructured{}
+ list, _ := p.ListDynamic(clusterGVR, ns)
+ for _, cluster := range list {
+ clusters[cluster.GetName()] = cluster
+ }
+ clustersByNamespace[ns] = clusters
+ }
+ cluster := clusters[cnpgSpecClusterName(schedule)]
+ if cluster == nil {
+ continue
+ }
+ declaration := cnpg.ParseBackupDeclaration(cluster)
+ blocker := declaration.DestinationBlocker(nestedString(schedule.Object, "spec", "method"), nestedString(schedule.Object, "spec", "pluginConfiguration", "name"))
+ if blocker == nil {
+ continue
+ }
+ destination := "no backup destination"
+ if blocker.Code == "method_destination_missing" {
+ destination = "no " + blocker.Method + " destination"
+ }
+ issue := newConditionIssue(gvr, "ScheduledBackup", ns, schedule.GetName(), SeverityWarning, ReasonCNPGScheduleDestinationMissing,
+ fmt.Sprintf("Backup schedule %s cannot run: %s", schedule.GetName(), destination), time.Time{}, false, ReasonCNPGScheduleDestinationMissing, schedule.GetCreationTimestamp().Time)
+ issue.RequiredReads = []EvidenceRead{{Group: cnpgGroup, Resource: "clusters", Namespace: ns, Verb: "list"}, {Group: cnpgGroup, Resource: "scheduledbackups", Namespace: ns, Verb: "list"}}
+ out = append(out, issue)
+ }
+ return out
+}
+
+func detectCNPGInstanceObservationIssues(p Provider, gvr schema.GroupVersionResource, clusters []*unstructured.Unstructured) []Issue {
+ pods, ok := p.(cnpgPodInventoryProvider)
+ if !ok {
+ return nil
+ }
+ var out []Issue
+ for _, cluster := range clusters {
+ if instances, known := pods.cnpgInstancePods(cluster); known {
+ out = append(out, cnpgInstanceObservationIssues(gvr, cluster, instances)...)
+ }
+ }
+ return out
+}
+
+func cnpgInstanceObservationIssues(gvr schema.GroupVersionResource, cluster *unstructured.Unstructured, pods []*corev1.Pod) []Issue {
+ ns, name := cluster.GetNamespace(), cluster.GetName()
+ newIssue := func(reason, message string, severity Severity) Issue {
+ issue := newConditionIssue(gvr, "Cluster", ns, name, severity, reason, message, time.Time{}, false, reason, cluster.GetCreationTimestamp().Time)
+ issue.RequiredReads = []EvidenceRead{{Group: cnpgGroup, Resource: "clusters", Namespace: ns, Verb: "list"}, {Resource: "pods", Namespace: ns, Verb: "list"}}
+ return issue
+ }
+ ready, primaryDown, allReported := 0, false, true
+ var labelled []string
+ for _, pod := range pods {
+ if cnpg.InstanceRole(pod) == "primary" {
+ labelled = append(labelled, pod.Name)
+ }
+ reported := false
+ for _, condition := range pod.Status.Conditions {
+ if condition.Type != corev1.PodReady {
+ continue
+ }
+ reported = condition.Status == corev1.ConditionTrue || condition.Status == corev1.ConditionFalse
+ if condition.Status == corev1.ConditionTrue {
+ ready++
+ } else if condition.Status == corev1.ConditionFalse && (cnpg.InstanceRole(pod) == "primary" || (cnpg.InstanceRole(pod) != "replica" && pod.Name == nestedString(cluster.Object, "status", "currentPrimary"))) {
+ primaryDown = true
+ }
+ }
+ allReported = allReported && reported
+ }
+ var out []Issue
+ statusReady, reported, _ := unstructured.NestedInt64(cluster.Object, "status", "readyInstances")
+ if !cnpgHibernated(cluster) && reported && allReported && int64(ready) < statusReady {
+ severity := SeverityWarning
+ if primaryDown || ready == 0 {
+ severity = SeverityCritical
+ }
+ message := fmt.Sprintf("%d of %d instance Pods not ready%s", len(pods)-ready, len(pods), cnpgPrimaryDownSuffix(primaryDown))
+ if int64(len(pods)) < statusReady {
+ message = fmt.Sprintf("Only %d ready instance Pods observed; CNPG status reports %d ready", ready, statusReady)
+ }
+ issue := newIssue(ReasonCNPGInstanceReadinessMismatch, message, severity)
+ issue.Cause = fmt.Sprintf("The Pods' Ready condition shows %d ready; CNPG status still reports %d. The status may be stale.", ready, statusReady)
+ out = append(out, issue)
+ }
+ primary := nestedString(cluster.Object, "status", "currentPrimary")
+ if primary != "" && len(labelled) > 0 {
+ sort.Strings(labelled)
+ matches := false
+ for _, pod := range labelled {
+ matches = matches || pod == primary
+ }
+ if !matches {
+ issue := newIssue(ReasonCNPGPrimaryLabelMismatch,
+ fmt.Sprintf("CNPG status names %s primary; the Pod labelled primary is %s", primary, labelled[0]), SeverityWarning)
+ issue.Cause = "Status may be stale, or a failover is under way."
+ out = append(out, issue)
+ }
+ }
+ return out
+}
+
+func cnpgPrimaryDownSuffix(down bool) string {
+ if down {
+ return ", including the primary"
+ }
+ return ""
+}
diff --git a/internal/issues/source_cnpg_observed_test.go b/internal/issues/source_cnpg_observed_test.go
new file mode 100644
index 0000000000..e498d0c5a9
--- /dev/null
+++ b/internal/issues/source_cnpg_observed_test.go
@@ -0,0 +1,124 @@
+package issues
+
+import (
+ "encoding/json"
+ "errors"
+ "strings"
+ "testing"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+)
+
+func TestCNPGScheduleDestinationProjection(t *testing.T) {
+ cluster := cnpgCluster(nil, nil)
+ schedule := cnpgHourly(false)
+ p := cnpgScheduleProvider([]*unstructured.Unstructured{cluster}, []*unstructured.Unstructured{schedule}, nil, nil)
+ p.namespaced = map[schema.GroupVersionResource]bool{cnpgClusterGVR: true, cnpgScheduledBackupGVR: true, cnpgBackupGVR: true, cnpgObjectStoreGVR: true}
+ flat := Compose(p, Filters{Kinds: []string{"ScheduledBackup"}, Limit: NoLimit, AllowUnfilteredEvidence: true})
+ issue := findIssue(t, flat, ReasonCNPGScheduleDestinationMissing)
+ if issue.Kind != "ScheduledBackup" || issue.Name != "hourly" || len(issue.RequiredReads) != 2 || !issue.OnsetUnknown {
+ t.Fatalf("unexpected issue: %+v", issue)
+ }
+ for _, grouped := range []bool{false, true} {
+ visible := Compose(p, Filters{Kinds: []string{"ScheduledBackup"}, Grouped: grouped, Limit: NoLimit, CanReadEvidence: func(EvidenceRead) bool { return true }})
+ findIssue(t, visible, ReasonCNPGScheduleDestinationMissing)
+ for _, resource := range []string{"clusters", "scheduledbackups"} {
+ denied := Compose(p, Filters{Grouped: grouped, Limit: NoLimit, CanReadEvidence: func(read EvidenceRead) bool { return read.Resource != resource }})
+ for _, got := range denied {
+ if got.Reason == ReasonCNPGScheduleDestinationMissing {
+ t.Fatalf("%s denial leaked through grouped=%v", resource, grouped)
+ }
+ }
+ }
+ }
+ grouped := GroupIssues(flat)
+ for _, opts := range []RelatedIssueOptions{
+ {CanReadRelated: func(Ref) bool { return true }},
+ {CanReadClusterScoped: func(string, string) bool { return true }},
+ {CanReadEvidence: func(read EvidenceRead) bool { return read.Resource != "clusters" }},
+ } {
+ for _, got := range RelatedIssuesFrom(flat, grouped, opts, cnpgGroup, "ScheduledBackup", "pg", "hourly") {
+ if got.Reason == ReasonCNPGScheduleDestinationMissing {
+ t.Fatal("cached related projection leaked required inventory")
+ }
+ }
+ }
+ visible := RelatedIssuesFrom(flat, grouped, RelatedIssueOptions{CanReadEvidence: func(EvidenceRead) bool { return true }}, cnpgGroup, "ScheduledBackup", "pg", "hourly")
+ findIssue(t, visible, ReasonCNPGScheduleDestinationMissing)
+ data, err := json.Marshal(issue)
+ if err != nil {
+ t.Fatal(err)
+ }
+ if strings.Contains(string(data), "RequiredReads") || strings.Contains(string(data), "scheduledbackups") {
+ t.Fatalf("internal evidence metadata serialized: %s", data)
+ }
+
+ schedule.Object["spec"].(map[string]any)["suspend"] = true
+ if got := detectCNPGScheduleDestinationIssues(p, cnpgScheduledBackupGVR, []*unstructured.Unstructured{schedule}); len(got) != 0 {
+ t.Fatal("suspended schedule raised a blocker")
+ }
+ schedule.Object["spec"].(map[string]any)["suspend"] = false
+ p.listErr = map[schema.GroupVersionResource]error{cnpgClusterGVR: errors.New("unavailable")}
+ if got := detectCNPGScheduleDestinationIssues(p, cnpgScheduledBackupGVR, []*unstructured.Unstructured{schedule}); len(got) != 0 {
+ t.Fatal("unread cluster mistaken for missing destination")
+ }
+}
+
+type cnpgObservedProvider struct {
+ *fakeProvider
+ pods []*corev1.Pod
+ known bool
+}
+
+func (p *cnpgObservedProvider) cnpgInstancePods(*unstructured.Unstructured) ([]*corev1.Pod, bool) {
+ return p.pods, p.known
+}
+
+func TestCNPGInstanceObservationCoverage(t *testing.T) {
+ cluster := cnpgCluster(nil, map[string]any{"readyInstances": int64(3), "currentPrimary": "pg-main-1"})
+ instance := func(name, role string, ready corev1.ConditionStatus) *corev1.Pod {
+ return &corev1.Pod{ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: "pg", Labels: map[string]string{"cnpg.io/instanceRole": role}}, Status: corev1.PodStatus{Conditions: []corev1.PodCondition{{Type: corev1.PodReady, Status: ready}}}}
+ }
+ p := &cnpgObservedProvider{fakeProvider: cnpgScheduleProvider([]*unstructured.Unstructured{cluster}, nil, nil, nil), known: true, pods: []*corev1.Pod{instance("pg-main-1", "primary", corev1.ConditionFalse), instance("pg-main-2", "replica", corev1.ConditionTrue), instance("pg-main-3", "replica", corev1.ConditionFalse)}}
+ issue := findIssue(t, detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}), ReasonCNPGInstanceReadinessMismatch)
+ if issue.Severity != SeverityCritical || issue.Message != "2 of 3 instance Pods not ready, including the primary" || len(issue.RequiredReads) != 2 {
+ t.Fatalf("unexpected contradiction: %+v", issue)
+ }
+ p.known = false
+ if got := detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}); len(got) != 0 {
+ t.Fatal("unread Pod inventory judged")
+ }
+ p.known = true
+ p.pods[0].Status.Conditions[0].Status = corev1.ConditionUnknown
+ if got := detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}); len(got) != 0 {
+ t.Fatal("unknown readiness judged")
+ }
+ p.pods[0].Status.Conditions = nil
+ if got := detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}); len(got) != 0 {
+ t.Fatal("unreported readiness judged")
+ }
+ p.pods[0].Status.Conditions = []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionTrue}}
+ cluster.SetAnnotations(map[string]string{"cnpg.io/hibernation": "on"})
+ if got := detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}); len(got) != 0 {
+ t.Fatal("hibernation judged as a readiness contradiction")
+ }
+ cluster.SetAnnotations(nil)
+ cluster.Object["status"].(map[string]any)["readyInstances"] = int64(2)
+ if got := detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}); len(got) != 0 {
+ t.Fatal("agreeing status judged")
+ }
+ cluster.Object["status"].(map[string]any)["currentPrimary"] = "pg-main-2"
+ findIssue(t, detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}), ReasonCNPGPrimaryLabelMismatch)
+}
+
+func TestCNPGMissingInstanceObservationDescribesInventory(t *testing.T) {
+ cluster := cnpgCluster(nil, map[string]any{"readyInstances": int64(2)})
+ p := &cnpgObservedProvider{fakeProvider: cnpgScheduleProvider([]*unstructured.Unstructured{cluster}, nil, nil, nil), known: true}
+ issue := findIssue(t, detectCNPGInstanceObservationIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{cluster}), ReasonCNPGInstanceReadinessMismatch)
+ if issue.Message != "Only 0 ready instance Pods observed; CNPG status reports 2 ready" || issue.Severity != SeverityCritical {
+ t.Fatalf("unexpected missing inventory message: %+v", issue)
+ }
+}
diff --git a/internal/issues/source_cnpg_schedule.go b/internal/issues/source_cnpg_schedule.go
new file mode 100644
index 0000000000..92c702eb78
--- /dev/null
+++ b/internal/issues/source_cnpg_schedule.go
@@ -0,0 +1,262 @@
+package issues
+
+import (
+ "fmt"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/cnpg"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+)
+
+const (
+ cnpgGroup = "postgresql.cnpg.io"
+ cnpgBarmanGroup = "barmancloud.cnpg.io"
+
+ ReasonCNPGScheduledRunNoBackup = "CNPGScheduledRunNoBackup"
+)
+
+// cnpgBackupPhasesInFlight: a run that may still succeed. The operator's own
+// terminal phases are completed and failed (walArchivingFailing also ends a
+// run); anything else, including a phase a newer minor adds, counts as still
+// running, which only ever delays the finding.
+var cnpgBackupPhasesDone = map[string]bool{"completed": true, "failed": true, "walArchivingFailing": true}
+
+// cnpgNamespaceBackupEvidence is what one namespace holds about backups.
+type cnpgNamespaceBackupEvidence struct {
+ schedules []*unstructured.Unstructured
+ backups []*unstructured.Unstructured
+ stores []*unstructured.Unstructured
+ // schedulesKnown, backupsKnown, storesKnown: the kind is watched and its
+ // list was read, so an empty list means none rather than unread. A run is
+ // judged missed only from evidence that was read: without schedules
+ // nothing is evaluated, without Backups no cluster is, and without
+ // ObjectStores no cluster whose backups land in one is.
+ schedulesKnown bool
+ backupsKnown bool
+ storesKnown bool
+}
+
+// detectCNPGScheduledRunIssues raises, per Cluster, that a schedule fired and
+// no successful backup exists from that run onwards. The missed-run detector
+// on the ScheduledBackup answers a different question (the operator did not
+// start a run it set for itself); this one fires when runs start and fail, or
+// succeed nowhere this cluster's evidence can see.
+//
+// The claim is exactly what is compared: the newest successful backup this
+// cluster has (Backup objects, the ObjectStore's recovery window, the in-tree
+// status field) is older than a time the schedule fired, with the last
+// successful backup's own duration allowed for the run to finish. A run still
+// in progress holds the finding back.
+func detectCNPGScheduledRunIssues(p Provider, clusterGVR schema.GroupVersionResource, clusters []*unstructured.Unstructured, now time.Time) []Issue {
+ if len(clusters) == 0 {
+ return nil
+ }
+ gvrs := map[string]schema.GroupVersionResource{}
+ for _, g := range p.WatchedDynamic() {
+ if g.Group == cnpgGroup || g.Group == cnpgBarmanGroup {
+ gvrs[g.Group+"/"+p.KindForGVR(g)] = g
+ }
+ }
+ schedGVR, ok := gvrs[cnpgGroup+"/ScheduledBackup"]
+ if !ok {
+ return nil
+ }
+ evidence := map[string]*cnpgNamespaceBackupEvidence{}
+ load := func(ns string) *cnpgNamespaceBackupEvidence {
+ if e, ok := evidence[ns]; ok {
+ return e
+ }
+ e := &cnpgNamespaceBackupEvidence{}
+ if items, err := p.ListDynamic(schedGVR, ns); err == nil {
+ e.schedules, e.schedulesKnown = items, true
+ }
+ if g, ok := gvrs[cnpgGroup+"/Backup"]; ok {
+ if items, err := p.ListDynamic(g, ns); err == nil {
+ e.backups, e.backupsKnown = items, true
+ }
+ }
+ if g, ok := gvrs[cnpgBarmanGroup+"/ObjectStore"]; ok {
+ if items, err := p.ListDynamic(g, ns); err == nil {
+ e.stores, e.storesKnown = items, true
+ }
+ }
+ evidence[ns] = e
+ return e
+ }
+
+ var out []Issue
+ for _, c := range clusters {
+ if c.GroupVersionKind().Group != "" && c.GroupVersionKind().Group != cnpgGroup {
+ continue
+ }
+ e := load(c.GetNamespace())
+ if !e.schedulesKnown || !e.backupsKnown {
+ continue
+ }
+ if cnpgBarmanPlugin(c).objectStore != "" && !e.storesKnown {
+ continue
+ }
+ if iss, ok := cnpgScheduledRunIssue(clusterGVR, c, e, now); ok {
+ ns := c.GetNamespace()
+ iss.RequiredReads = []EvidenceRead{{Group: cnpgGroup, Resource: "clusters", Namespace: ns, Verb: "list"}, {Group: cnpgGroup, Resource: "scheduledbackups", Namespace: ns, Verb: "list"}, {Group: cnpgGroup, Resource: "backups", Namespace: ns, Verb: "list"}}
+ if cnpgBarmanPlugin(c).objectStore != "" {
+ iss.RequiredReads = append(iss.RequiredReads, EvidenceRead{Group: cnpgBarmanGroup, Resource: "objectstores", Namespace: ns, Verb: "list"})
+ }
+ out = append(out, iss)
+ }
+ }
+ return out
+}
+
+func cnpgScheduledRunIssue(gvr schema.GroupVersionResource, cluster *unstructured.Unstructured, e *cnpgNamespaceBackupEvidence, now time.Time) (Issue, bool) {
+ name := cluster.GetName()
+ lastSuccess, lastDuration := cnpgLatestSuccessfulBackup(cluster, e)
+
+ var worst struct {
+ schedule, spec string
+ fired time.Time
+ }
+ for _, s := range e.schedules {
+ if cnpgSpecClusterName(s) != name {
+ continue
+ }
+ if suspended, _, _ := unstructured.NestedBool(s.Object, "spec", "suspend"); suspended {
+ continue
+ }
+ spec, _, _ := unstructured.NestedString(s.Object, "spec", "schedule")
+ if _, err := cnpg.ParseSchedule(spec); err != nil {
+ continue
+ }
+ // Cron alone cannot establish a run's time without the operator's clock.
+ reported, _, _ := unstructured.NestedString(s.Object, "status", "lastScheduleTime")
+ fired, err := time.Parse(time.RFC3339, reported)
+ if err != nil || !fired.After(lastSuccess) || fired.After(now) || fired.Before(s.GetCreationTimestamp().Time) || fired.Before(cluster.GetCreationTimestamp().Time) {
+ continue
+ }
+ if now.Sub(fired) <= lastDuration+cnpgScheduledBackupGrace {
+ continue
+ }
+ if worst.fired.IsZero() || fired.Before(worst.fired) {
+ worst.schedule, worst.spec, worst.fired = s.GetName(), spec, fired
+ }
+ }
+ if worst.fired.IsZero() || cnpgBackupInFlightSince(cluster, e.backups, lastSuccess) {
+ return Issue{}, false
+ }
+
+ // The run's time is the issue's first_seen (a reader shows it as an age);
+ // the message names the schedule in words so nobody has to decode cron.
+ reading := cnpg.DescribeSchedule(worst.spec)
+ if reading == "" {
+ reading = "cron " + worst.spec
+ }
+ msg := fmt.Sprintf("ScheduledBackup %s (%s) has had no successful backup since its run", worst.schedule, reading)
+ if lastSuccess.IsZero() {
+ msg = fmt.Sprintf("ScheduledBackup %s (%s) has had no successful backup observed since its reported run", worst.schedule, reading)
+ }
+ return newConditionIssue(gvr, "Cluster", cluster.GetNamespace(), name, SeverityWarning,
+ ReasonCNPGScheduledRunNoBackup, msg, worst.fired, true, ReasonCNPGScheduledRunNoBackup, cluster.GetCreationTimestamp().Time), true
+}
+
+// cnpgLatestSuccessfulBackup is the newest success across the three places
+// CNPG records one, plus the duration of the newest completed Backup object
+// (zero when unknown).
+func cnpgLatestSuccessfulBackup(cluster *unstructured.Unstructured, e *cnpgNamespaceBackupEvidence) (time.Time, time.Duration) {
+ var latest time.Time
+ var latestBackup time.Time
+ var duration time.Duration
+ for _, b := range e.backups {
+ if !cnpg.BackupMatchesCluster(b, cluster) {
+ continue
+ }
+ if phase, _, _ := unstructured.NestedString(b.Object, "status", "phase"); phase != "completed" {
+ continue
+ }
+ stopped := cnpgStatusTime(b, "stoppedAt")
+ if stopped.IsZero() {
+ stopped = cnpgStatusTime(b, "startedAt")
+ }
+ if stopped.IsZero() || !stopped.After(latestBackup) {
+ continue
+ }
+ latestBackup = stopped
+ if started := cnpgStatusTime(b, "startedAt"); !started.IsZero() && stopped.After(started) {
+ duration = stopped.Sub(started)
+ } else {
+ duration = 0
+ }
+ }
+ latest = latestBackup
+
+ plugin := cnpgBarmanPlugin(cluster)
+ if plugin.objectStore != "" {
+ for _, s := range e.stores {
+ if s.GetName() != plugin.objectStore {
+ continue
+ }
+ if t := cnpgParseTime(nestedString(s.Object, "status", "serverRecoveryWindow", plugin.serverName, "lastSuccessfulBackupTime")); t.After(latest) && !t.Before(cluster.GetCreationTimestamp().Time) {
+ latest = t
+ }
+ }
+ } else if !plugin.present {
+ if t := cnpgParseTime(nestedString(cluster.Object, "status", "lastSuccessfulBackup")); t.After(latest) && !t.Before(cluster.GetCreationTimestamp().Time) {
+ latest = t
+ }
+ }
+ return latest, duration
+}
+
+func cnpgBackupInFlightSince(cluster *unstructured.Unstructured, backups []*unstructured.Unstructured, since time.Time) bool {
+ for _, b := range backups {
+ if !cnpg.BackupMatchesCluster(b, cluster) {
+ continue
+ }
+ phase, _, _ := unstructured.NestedString(b.Object, "status", "phase")
+ if cnpgBackupPhasesDone[phase] {
+ continue
+ }
+ if !b.GetCreationTimestamp().Time.Before(since) {
+ return true
+ }
+ }
+ return false
+}
+
+type cnpgPluginRef struct {
+ present bool
+ objectStore string
+ serverName string
+}
+
+// cnpgBarmanPlugin mirrors getCNPGClusterBarmanPlugin in k8s-ui: the
+// barman-cloud plugin entry, its ObjectStore and the server key the recovery
+// window is recorded under (serverName, else the cluster name).
+func cnpgBarmanPlugin(cluster *unstructured.Unstructured) cnpgPluginRef {
+ plugin, present := cnpg.ParseBackupDeclaration(cluster).BarmanPlugin()
+ return cnpgPluginRef{present: present, objectStore: plugin.ObjectStore, serverName: plugin.ServerName}
+}
+
+func cnpgSpecClusterName(u *unstructured.Unstructured) string {
+ return nestedString(u.Object, "spec", "cluster", "name")
+}
+
+func cnpgStatusTime(u *unstructured.Unstructured, field string) time.Time {
+ return cnpgParseTime(nestedString(u.Object, "status", field))
+}
+
+func cnpgParseTime(v string) time.Time {
+ if v == "" {
+ return time.Time{}
+ }
+ t, err := time.Parse(time.RFC3339, v)
+ if err != nil {
+ return time.Time{}
+ }
+ return t
+}
+
+func nestedString(obj map[string]any, path ...string) string {
+ v, _, _ := unstructured.NestedString(obj, path...)
+ return v
+}
diff --git a/internal/issues/source_cnpg_schedule_test.go b/internal/issues/source_cnpg_schedule_test.go
new file mode 100644
index 0000000000..59cbefdce1
--- /dev/null
+++ b/internal/issues/source_cnpg_schedule_test.go
@@ -0,0 +1,279 @@
+package issues
+
+import (
+ "errors"
+ "strings"
+ "testing"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/issuesapi"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+)
+
+var (
+ cnpgObjectStoreGVR = schema.GroupVersionResource{Group: "barmancloud.cnpg.io", Version: "v1", Resource: "objectstores"}
+ // 2026-09-30 14:30:00 UTC; the hourly schedule last fired at 14:00:00.
+ cnpgScheduleNow = time.Date(2026, 9, 30, 14, 30, 0, 0, time.UTC)
+)
+
+func cnpgSchedObj(kind, name string, created time.Time, spec, status map[string]any) *unstructured.Unstructured {
+ u := &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1",
+ "kind": kind,
+ "metadata": map[string]any{"name": name, "namespace": "pg", "creationTimestamp": created.Format(time.RFC3339)},
+ "spec": spec,
+ }}
+ if status != nil {
+ u.Object["status"] = status
+ }
+ return u
+}
+
+func cnpgHourly(suspend bool) *unstructured.Unstructured {
+ return cnpgSchedObj("ScheduledBackup", "hourly", cnpgScheduleNow.Add(-48*time.Hour),
+ map[string]any{"schedule": "0 0 * * * *", "suspend": suspend, "cluster": map[string]any{"name": "pg-main"}}, map[string]any{"lastScheduleTime": "2026-09-30T14:00:00Z"})
+}
+
+func cnpgBackupAt(name, phase string, started, stopped time.Time) *unstructured.Unstructured {
+ status := map[string]any{"phase": phase}
+ if !started.IsZero() {
+ status["startedAt"] = started.Format(time.RFC3339)
+ }
+ if !stopped.IsZero() {
+ status["stoppedAt"] = stopped.Format(time.RFC3339)
+ }
+ created := started
+ if created.IsZero() {
+ created = cnpgScheduleNow.Add(-time.Minute)
+ }
+ return cnpgSchedObj("Backup", name, created, map[string]any{"cluster": map[string]any{"name": "pg-main"}}, status)
+}
+
+func cnpgScheduleProvider(clusters []*unstructured.Unstructured, schedules, backups, stores []*unstructured.Unstructured) *fakeProvider {
+ return &fakeProvider{
+ dynamic: map[schema.GroupVersionResource][]*unstructured.Unstructured{
+ cnpgClusterGVR: clusters,
+ cnpgScheduledBackupGVR: schedules,
+ cnpgBackupGVR: backups,
+ cnpgObjectStoreGVR: stores,
+ },
+ kinds: map[schema.GroupVersionResource]string{
+ cnpgClusterGVR: "Cluster",
+ cnpgScheduledBackupGVR: "ScheduledBackup",
+ cnpgBackupGVR: "Backup",
+ cnpgObjectStoreGVR: "ObjectStore",
+ },
+ }
+}
+
+func TestCNPGScheduledRunNoBackup(t *testing.T) {
+ at := func(h, m int) time.Time { return time.Date(2026, 9, 30, h, m, 0, 0, time.UTC) }
+ pluginCluster := cnpgCluster(map[string]any{"plugins": []any{map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "store", "serverName": "main-v2"},
+ }}}, nil)
+ store := func(last time.Time) *unstructured.Unstructured {
+ u := cnpgSchedObj("ObjectStore", "store", at(0, 0), map[string]any{}, map[string]any{
+ "serverRecoveryWindow": map[string]any{"main-v2": map[string]any{"lastSuccessfulBackupTime": last.Format(time.RFC3339)}},
+ })
+ u.SetAPIVersion("barmancloud.cnpg.io/v1")
+ return u
+ }
+
+ tests := []struct {
+ name string
+ cluster *unstructured.Unstructured
+ schedules []*unstructured.Unstructured
+ backups []*unstructured.Unstructured
+ stores []*unstructured.Unstructured
+ want string
+ run string // the run the issue dates itself from (first_seen)
+ }{
+ {
+ name: "last success before the reported 14:00 run, which produced nothing",
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ backups: []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(12, 0), at(12, 2)), cnpgBackupAt("b2", "failed", at(13, 0), at(13, 1))},
+ want: "ScheduledBackup hourly (every hour, on the hour) has had no successful backup since its run",
+ run: "2026-09-30T14:00:00Z",
+ },
+ {
+ name: "the 14:00 run succeeded",
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ backups: []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(14, 0), at(14, 3))},
+ },
+ {
+ // The 14:00 run is due, but the last success took 40 minutes, so
+ // 14:30 is still inside the time a run takes here.
+ name: "within the observed duration of the last success",
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ backups: []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(13, 5), at(13, 45))},
+ },
+ {
+ name: "a run started after the fire time is still in progress",
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ backups: []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(11, 0), at(11, 1)), cnpgBackupAt("b2", "running", at(12, 0), time.Time{})},
+ },
+ {
+ name: "suspended schedules raise nothing",
+ schedules: []*unstructured.Unstructured{cnpgHourly(true)},
+ backups: []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(1, 0), at(1, 1))},
+ },
+ {
+ name: "no success ever: measured from the reported run",
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ want: "has had no successful backup observed since its reported run",
+ run: "2026-09-30T14:00:00Z",
+ },
+ {
+ name: "the ObjectStore's recovery window counts as a success",
+ cluster: pluginCluster,
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ stores: []*unstructured.Unstructured{store(at(14, 1))},
+ },
+ {
+ name: "an older ObjectStore success still leaves the reported 14:00 run unaccounted for",
+ cluster: pluginCluster,
+ schedules: []*unstructured.Unstructured{cnpgHourly(false)},
+ stores: []*unstructured.Unstructured{store(at(12, 1))},
+ want: "since its run",
+ run: "2026-09-30T14:00:00Z",
+ },
+ {
+ name: "unparseable schedule",
+ schedules: []*unstructured.Unstructured{cnpgSchedObj("ScheduledBackup", "bad", at(0, 0), map[string]any{"schedule": "every hour", "cluster": map[string]any{"name": "pg-main"}}, nil)},
+ },
+ {
+ name: "schedule for another cluster",
+ schedules: []*unstructured.Unstructured{cnpgSchedObj("ScheduledBackup", "other", at(0, 0), map[string]any{"schedule": "0 0 * * * *", "cluster": map[string]any{"name": "pg-other"}}, nil)},
+ },
+ }
+ for _, tc := range tests {
+ t.Run(tc.name, func(t *testing.T) {
+ c := tc.cluster
+ if c == nil {
+ c = cnpgCluster(nil, nil)
+ }
+ p := cnpgScheduleProvider([]*unstructured.Unstructured{c}, tc.schedules, tc.backups, tc.stores)
+ got := detectCNPGScheduledRunIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{c}, cnpgScheduleNow)
+ if tc.want == "" {
+ if len(got) != 0 {
+ t.Fatalf("unexpected issue: %s", got[0].Message)
+ }
+ return
+ }
+ if len(got) != 1 {
+ t.Fatalf("issues = %v, want one", reasonsOf(got))
+ }
+ iss := got[0]
+ if iss.Reason != ReasonCNPGScheduledRunNoBackup || iss.Kind != "Cluster" || iss.Name != "pg-main" || iss.Severity != SeverityWarning {
+ t.Errorf("issue = %+v", iss)
+ }
+ if !strings.Contains(iss.Message, tc.want) {
+ t.Errorf("message = %q, want it to contain %q", iss.Message, tc.want)
+ }
+ if tc.run != "" && iss.FirstSeen.UTC().Format(time.RFC3339) != tc.run {
+ t.Errorf("first_seen = %s, want the run %s", iss.FirstSeen, tc.run)
+ }
+ if iss.Category != issuesapi.CategoryBackupFailed {
+ t.Errorf("category = %v, want backup_failed", iss.Category)
+ }
+ })
+ }
+}
+
+func TestCNPGScheduledRunUsesReportedInstant(t *testing.T) {
+ c := cnpgCluster(nil, nil)
+ sched := cnpgSchedObj("ScheduledBackup", "nightly", cnpgScheduleNow.Add(-48*time.Hour),
+ map[string]any{"schedule": "0 0 2 * * *", "cluster": map[string]any{"name": "pg-main"}},
+ map[string]any{"lastScheduleTime": "2026-09-30T02:00:00-04:00"})
+ p := cnpgScheduleProvider([]*unstructured.Unstructured{c}, []*unstructured.Unstructured{sched}, nil, nil)
+ got := detectCNPGScheduledRunIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{c}, cnpgScheduleNow)
+ if len(got) != 1 || !got[0].FirstSeen.Equal(time.Date(2026, 9, 30, 6, 0, 0, 0, time.UTC)) || !strings.Contains(got[0].Message, "every day at 02:00 operator clock") {
+ t.Fatalf("issues=%+v", got)
+ }
+ for _, reported := range []string{"", "invalid", "2026-10-01T00:00:00Z", "2026-09-01T00:00:00Z"} {
+ _ = unstructured.SetNestedField(sched.Object, reported, "status", "lastScheduleTime")
+ if got := detectCNPGScheduledRunIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{c}, cnpgScheduleNow); len(got) != 0 {
+ t.Errorf("unestablished run %q raised %+v", reported, got)
+ }
+ }
+}
+
+func TestCNPGScheduledRunNeedsReadableBackups(t *testing.T) {
+ at := func(h, m int) time.Time { return time.Date(2026, 9, 30, h, m, 0, 0, time.UTC) }
+ schedules := []*unstructured.Unstructured{cnpgHourly(false)}
+ backups := []*unstructured.Unstructured{cnpgBackupAt("b1", "completed", at(12, 0), at(12, 2))}
+ plain := cnpgCluster(nil, nil)
+ plugin := cnpgCluster(map[string]any{"plugins": []any{map[string]any{
+ "name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "store", "serverName": "main-v2"},
+ }}}, nil)
+ detect := func(p *fakeProvider, c *unstructured.Unstructured) []Issue {
+ return detectCNPGScheduledRunIssues(p, cnpgClusterGVR, []*unstructured.Unstructured{c}, cnpgScheduleNow)
+ }
+
+ readable := cnpgScheduleProvider([]*unstructured.Unstructured{plain}, schedules, backups, nil)
+ if got := detect(readable, plain); len(got) != 1 {
+ t.Fatalf("with readable backups the missed run is reported, got %v", reasonsOf(got))
+ }
+
+ failed := cnpgScheduleProvider([]*unstructured.Unstructured{plain}, schedules, backups, nil)
+ failed.listErr = map[schema.GroupVersionResource]error{cnpgBackupGVR: errors.New("forbidden")}
+ if got := detect(failed, plain); len(got) != 0 {
+ t.Errorf("a failed Backup list raised %v", reasonsOf(got))
+ }
+
+ unwatched := cnpgScheduleProvider([]*unstructured.Unstructured{plain}, schedules, backups, nil)
+ delete(unwatched.dynamic, cnpgBackupGVR)
+ if got := detect(unwatched, plain); len(got) != 0 {
+ t.Errorf("an unwatched Backup kind raised %v", reasonsOf(got))
+ }
+
+ storeFailed := cnpgScheduleProvider([]*unstructured.Unstructured{plugin}, schedules, backups, nil)
+ storeFailed.listErr = map[schema.GroupVersionResource]error{cnpgObjectStoreGVR: errors.New("forbidden")}
+ if got := detect(storeFailed, plugin); len(got) != 0 {
+ t.Errorf("an unread ObjectStore raised %v for a plugin cluster", reasonsOf(got))
+ }
+}
+
+// The Compose path attributes the finding to the Cluster, next to its own
+// condition-derived issues.
+func TestCNPGScheduledRunThroughCompose(t *testing.T) {
+ c := cnpgCluster(map[string]any{"instances": int64(1)}, map[string]any{"phase": "Cluster in healthy state", "readyInstances": int64(1)})
+ sched := cnpgSchedObj("ScheduledBackup", "hourly", time.Now().Add(-48*time.Hour),
+ map[string]any{"schedule": "0 0 * * * *", "cluster": map[string]any{"name": "pg-main"}}, map[string]any{"lastScheduleTime": time.Now().Add(-time.Hour).Format(time.RFC3339)})
+ p := cnpgScheduleProvider([]*unstructured.Unstructured{c}, []*unstructured.Unstructured{sched}, nil, nil)
+ got := Compose(p, Filters{Kinds: []string{"Cluster"}, Limit: NoLimit, AllowUnfilteredEvidence: true})
+ found := false
+ for _, iss := range got {
+ if iss.Reason == ReasonCNPGScheduledRunNoBackup {
+ found = true
+ }
+ }
+ if !found {
+ t.Errorf("issues = %v, want %s", reasonsOf(got), ReasonCNPGScheduledRunNoBackup)
+ }
+}
+
+func TestCNPGScheduledRunExcludesPredecessorEvidence(t *testing.T) {
+ c := cnpgCluster(map[string]any{"plugins": []any{map[string]any{"name": "barman-cloud.cloudnative-pg.io", "parameters": map[string]any{"barmanObjectName": "store", "serverName": "main"}}}}, nil)
+ c.SetUID("current")
+ c.SetCreationTimestamp(metav1.NewTime(cnpgScheduleNow.Add(-time.Hour)))
+ old := cnpgBackupAt("old", "completed", cnpgScheduleNow.Add(-25*time.Minute), cnpgScheduleNow.Add(-20*time.Minute))
+ _ = unstructured.SetNestedField(old.Object, "previous", "status", "pluginMetadata", "clusterUID")
+ running := old.DeepCopy()
+ _ = unstructured.SetNestedField(running.Object, "running", "status", "phase")
+ store := cnpgSchedObj("ObjectStore", "store", cnpgScheduleNow.Add(-48*time.Hour), nil, map[string]any{"serverRecoveryWindow": map[string]any{"main": map[string]any{"lastSuccessfulBackupTime": cnpgScheduleNow.Add(-2 * time.Hour).Format(time.RFC3339)}}})
+ store.SetAPIVersion("barmancloud.cnpg.io/v1")
+ e := &cnpgNamespaceBackupEvidence{schedules: []*unstructured.Unstructured{cnpgHourly(false)}, backups: []*unstructured.Unstructured{old, running}, stores: []*unstructured.Unstructured{store}}
+ if last, duration := cnpgLatestSuccessfulBackup(c, e); !last.IsZero() || duration != 0 {
+ t.Fatalf("predecessor success counted: %s, %s", last, duration)
+ }
+ if _, ok := cnpgScheduledRunIssue(cnpgClusterGVR, c, e, cnpgScheduleNow); !ok {
+ t.Fatal("predecessor run suppressed current missed-run finding")
+ }
+ c.SetCreationTimestamp(metav1.NewTime(cnpgScheduleNow.Add(-10 * time.Minute)))
+ if _, ok := cnpgScheduledRunIssue(cnpgClusterGVR, c, e, cnpgScheduleNow); ok {
+ t.Fatal("schedule fired before current Cluster existed")
+ }
+}
diff --git a/internal/issues/source_conditions.go b/internal/issues/source_conditions.go
index 5f28410d0d..5bc8738592 100644
--- a/internal/issues/source_conditions.go
+++ b/internal/issues/source_conditions.go
@@ -103,6 +103,16 @@ func detectGenericCRDIssues(p Provider, f Filters, ownedSubjects map[string]bool
out = append(out, detectVeleroIssues(gvr, kind, items, ownedSubjects)...)
continue
}
+ // A schedule's outcome is read against the cluster's backups, so it is
+ // evaluated per namespace over the Cluster list, not per object.
+ if gvr.Group == cnpgGroup && kind == "Cluster" {
+ out = append(out, detectCNPGScheduledRunIssues(p, gvr, items, time.Now())...)
+ out = append(out, detectCNPGInstanceObservationIssues(p, gvr, items)...)
+ }
+ if gvr.Group == cnpgGroup && kind == "ScheduledBackup" {
+ out = append(out, detectCNPGScheduleDestinationIssues(p, gvr, items)...)
+ }
+
for _, u := range items {
if gvr.Group == "kafka.strimzi.io" && kind == "KafkaConnector" {
out = append(out, detectStrimziConnectorIssues(gvr, u)...)
diff --git a/internal/issues/types.go b/internal/issues/types.go
index 23d17ff392..08f6ac12c1 100644
--- a/internal/issues/types.go
+++ b/internal/issues/types.go
@@ -68,6 +68,8 @@ const (
// owner/affected deep-links can disambiguate CRDs from core kinds.
type Ref = issuesapi.Ref
+type EvidenceRead = issuesapi.EvidenceRead
+
// Issue is the unified cluster-health record.
//
// Flat (pre-group) rows are snapshot-derived. GroupIssues folds them and sets
diff --git a/internal/k8s/cache.go b/internal/k8s/cache.go
index 132a5dafc9..a6a044f26b 100644
--- a/internal/k8s/cache.go
+++ b/internal/k8s/cache.go
@@ -912,7 +912,7 @@ func recordToTimelineStore(clusterContext, kind, namespace, name, uid, op string
}
labels := entry.Labels
createdAt := entry.CreatedAt
- healthState := classifyTimelineHealth(kind, obj, time.Now())
+ healthState := classifyTimelineChangeHealth(kind, oldObj, obj, time.Now())
// Feed the tombstone on every add/update/delete. While the object is live
// this mirrors its enrichment; once it is gone (delete, or a late K8s event
diff --git a/internal/k8s/capabilities.go b/internal/k8s/capabilities.go
index c40b9cfec6..870b28282c 100644
--- a/internal/k8s/capabilities.go
+++ b/internal/k8s/capabilities.go
@@ -133,13 +133,18 @@ type CloudConnectCapability struct {
// here in the same change, plus an entry in web/src/api/radarFeatures.ts
// (TestFeatureFlagsHaveFrontendGates enforces the pairing).
type FeatureCapabilities struct {
- YAMLReview bool `json:"yamlReview"`
- YAMLSchemas bool `json:"yamlSchemas"`
- WorkloadImages bool `json:"workloadImages"`
- ResourceIssues bool `json:"resourceIssues"` // GET /api/issues/resource/{kind}/{namespace}/{name}
- PodEnvironment bool `json:"podEnvironment"` // GET /api/pods/{namespace}/{name}/environment
- PolicyResource bool `json:"policyResource"` // GET /api/policy/resource/{kind}/{namespace}/{name}
- WorkloadHistory bool `json:"workloadHistory"` // GET /api/workloads/{kind}/{namespace}/{name}/history
+ YAMLReview bool `json:"yamlReview"`
+ YAMLSchemas bool `json:"yamlSchemas"`
+ WorkloadImages bool `json:"workloadImages"`
+ ResourceIssues bool `json:"resourceIssues"` // GET /api/issues/resource/{kind}/{namespace}/{name}
+ ResourceIssueCoverage bool `json:"resourceIssueCoverage"` // Resource issues with ?coverage=1
+ PodEnvironment bool `json:"podEnvironment"` // GET /api/pods/{namespace}/{name}/environment
+ PolicyResource bool `json:"policyResource"` // GET /api/policy/resource/{kind}/{namespace}/{name}
+ WorkloadHistory bool `json:"workloadHistory"` // GET /api/workloads/{kind}/{namespace}/{name}/history
+ // The workspace and its read/action surfaces, excluding protection setup.
+ CNPGWorkspace bool `json:"cnpgWorkspace"`
+ CNPGProtectionSetup bool `json:"cnpgProtectionSetup"`
+ GitOpsWriteEvidence bool `json:"gitopsWriteEvidence"` // POST /api/gitops/write-evidence
}
// WorkloadWritePermissions indicates which workload resources the user can patch.
diff --git a/internal/k8s/context_manager.go b/internal/k8s/context_manager.go
index 9457f70673..3faf89d268 100644
--- a/internal/k8s/context_manager.go
+++ b/internal/k8s/context_manager.go
@@ -91,7 +91,7 @@ var (
// InitAllSubsystems concurrently on the shared cache singletons. This mutex
// does — a second request waits for the first to finish rather than
// interleaving teardown/init.
- contextOpMu sync.Mutex
+ contextOpMu sync.RWMutex
// Incremented BEFORE contextOpMu is acquired — that ordering is the
// mechanism: it makes a queued-but-blocked operation visible to runtime
// auth-loss candidate intake, which a try-lock could never see.
@@ -164,6 +164,20 @@ func OperationContext() context.Context {
return operationCtx
}
+// Clients and caches are swapped in separate steps. The callback only captures
+// dependencies; doing I/O here would delay context switches and auth recovery.
+func CaptureClusterReads(capture func()) bool {
+ if activeContextOperations.Load() != 0 || !contextOpMu.TryRLock() {
+ return false
+ }
+ defer contextOpMu.RUnlock()
+ if activeContextOperations.Load() != 0 {
+ return false
+ }
+ capture()
+ return true
+}
+
// SetSessionStopper registers the callback that terminates active port-forward /
// exec sessions. The destructive cache operations call it once they commit to
// tearing the cache down. Registered by the server (which owns sessions) to
diff --git a/internal/k8s/dynamic_cache.go b/internal/k8s/dynamic_cache.go
index f0bcaf56c6..3f3915a574 100644
--- a/internal/k8s/dynamic_cache.go
+++ b/internal/k8s/dynamic_cache.go
@@ -306,6 +306,7 @@ var supportedCRDFallbacks = []supportedCRDResource{
{Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "databases", Kind: "Database", Namespaced: true},
{Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "publications", Kind: "Publication", Namespaced: true},
{Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "subscriptions", Kind: "Subscription", Namespaced: true},
+ {Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "databaseroles", Kind: "DatabaseRole", Namespaced: true},
{Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "imagecatalogs", Kind: "ImageCatalog", Namespaced: true},
{Group: "postgresql.cnpg.io", Versions: []string{"v1"}, Resource: "clusterimagecatalogs", Kind: "ClusterImageCatalog", Namespaced: false},
// The barman-cloud plugin ships its own group; the in-tree backup settings it
diff --git a/internal/k8s/read_capture_test.go b/internal/k8s/read_capture_test.go
new file mode 100644
index 0000000000..982f18a78a
--- /dev/null
+++ b/internal/k8s/read_capture_test.go
@@ -0,0 +1,34 @@
+package k8s
+
+import "testing"
+
+func TestCaptureClusterReadsRefusesTransitionWithoutBlocking(t *testing.T) {
+ called := false
+ contextOpMu.Lock()
+ if CaptureClusterReads(func() { called = true }) {
+ contextOpMu.Unlock()
+ t.Fatal("captured dependencies during a transition")
+ }
+ contextOpMu.Unlock()
+ if called {
+ t.Fatal("callback ran during a transition")
+ }
+ activeContextOperations.Add(1)
+ captured := CaptureClusterReads(func() { called = true })
+ activeContextOperations.Add(-1)
+ if captured || called {
+ t.Fatal("captured dependencies while a context change was queued")
+ }
+ if !CaptureClusterReads(func() { called = true }) || !called {
+ t.Fatal("stable dependencies were not captured")
+ }
+}
+
+func TestCaptureClusterReadsAllowsConcurrentCaptures(t *testing.T) {
+ contextOpMu.RLock()
+ defer contextOpMu.RUnlock()
+ called := false
+ if !CaptureClusterReads(func() { called = true }) || !called {
+ t.Fatal("an independent read capture blocked stable dependencies")
+ }
+}
diff --git a/internal/k8s/testing.go b/internal/k8s/testing.go
index 02b60ae650..cf4d2afe8b 100644
--- a/internal/k8s/testing.go
+++ b/internal/k8s/testing.go
@@ -260,6 +260,23 @@ func InitTestDynamicResourceCache(dynClient dynamic.Interface, resources []APIRe
return InitDynamicResourceCache(nil)
}
+// InitTestDynamicResourceCacheWithFallbacks is InitTestDynamicResourceCache
+// for an identity that may not list CRDs cluster-wide: the dynamic cache then
+// probes fallbacks and watches each kind namespace by namespace.
+func InitTestDynamicResourceCacheWithFallbacks(dynClient dynamic.Interface, resources []APIResource, fallbacks []string) error {
+ resourcePermsMu.Lock()
+ prev, prevExpiry := cachedPermResult, resourcePermsExpiry
+ cachedPermResult = &PermissionCheckResult{Perms: &ResourcePermissions{}, ScopeCandidates: fallbacks}
+ resourcePermsExpiry = time.Now().Add(time.Hour)
+ resourcePermsMu.Unlock()
+ defer func() {
+ resourcePermsMu.Lock()
+ cachedPermResult, resourcePermsExpiry = prev, prevExpiry
+ resourcePermsMu.Unlock()
+ }()
+ return InitTestDynamicResourceCache(dynClient, resources)
+}
+
// ResetTestDynamicState tears down the dynamic cache + discovery singletons
// and clears the dynamic client. Pairs with InitTestDynamicResourceCache.
func ResetTestDynamicState() {
diff --git a/internal/k8s/timeline_health.go b/internal/k8s/timeline_health.go
index f4a2e26fae..d54e379663 100644
--- a/internal/k8s/timeline_health.go
+++ b/internal/k8s/timeline_health.go
@@ -49,3 +49,43 @@ func levelToTimeline(l health.Level) timeline.HealthState {
return timeline.HealthUnknown
}
}
+
+// classifyTimelineChangeHealth labels one recorded change. The canonical
+// classifier judges standing state, so it holds back on a fresh failure (a
+// readiness grace, a restart threshold); a timeline row instead labels the
+// moment it records, and a Pod whose container just exited with an error, or
+// that just lost readiness, is not healthy at that moment.
+func classifyTimelineChangeHealth(kind string, oldObj, newObj any, now time.Time) timeline.HealthState {
+ pod, ok := newObj.(*corev1.Pod)
+ if kind != "Pod" || !ok {
+ return classifyTimelineHealth(kind, newObj, now)
+ }
+ level := health.PodDisplayLevel(pod, now)
+ if pod.Status.Phase != corev1.PodRunning || pod.DeletionTimestamp != nil {
+ return levelToTimeline(level)
+ }
+ completing := false
+ for _, cs := range pod.Status.ContainerStatuses {
+ if t := cs.State.Terminated; t != nil {
+ if t.ExitCode != 0 {
+ level = health.WorseOf(level, health.LevelUnhealthy)
+ } else {
+ completing = true
+ }
+ }
+ }
+ // A container finishing cleanly also drops readiness; that is completion.
+ if old, ok := oldObj.(*corev1.Pod); ok && old != pod && !completing && timelinePodReady(old) && !timelinePodReady(pod) {
+ level = health.WorseOf(level, health.LevelDegraded)
+ }
+ return levelToTimeline(level)
+}
+
+func timelinePodReady(p *corev1.Pod) bool {
+ for _, c := range p.Status.Conditions {
+ if c.Type == corev1.PodReady {
+ return c.Status == corev1.ConditionTrue
+ }
+ }
+ return false
+}
diff --git a/internal/k8s/timeline_health_test.go b/internal/k8s/timeline_health_test.go
index ced2d88009..9e9849ba60 100644
--- a/internal/k8s/timeline_health_test.go
+++ b/internal/k8s/timeline_health_test.go
@@ -149,3 +149,45 @@ func TestClassifyTimelineHealthWorkloads(t *testing.T) {
})
}
}
+
+func TestClassifyTimelineChangeHealthPod(t *testing.T) {
+ now := time.Now()
+ ready := func(v corev1.ConditionStatus) []corev1.PodCondition {
+ return []corev1.PodCondition{{Type: corev1.PodReady, Status: v, LastTransitionTime: metav1.NewTime(now)}}
+ }
+ running := corev1.ContainerState{Running: &corev1.ContainerStateRunning{StartedAt: metav1.NewTime(now.Add(-time.Hour))}}
+ pod := func(conds []corev1.PodCondition, state corev1.ContainerState, restarts int32) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{CreationTimestamp: metav1.NewTime(now.Add(-time.Hour))},
+ Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "postgres", ReadinessProbe: &corev1.Probe{}}}},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodRunning, Conditions: conds,
+ ContainerStatuses: []corev1.ContainerStatus{{Name: "postgres", State: state, RestartCount: restarts}},
+ },
+ }
+ }
+ errored := corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{Reason: "Error", ExitCode: 1}}
+ completed := corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{Reason: "Completed", ExitCode: 0}}
+ cases := []struct {
+ name string
+ old, new *corev1.Pod
+ want timeline.HealthState
+ }{
+ {"container exited with an error", pod(ready(corev1.ConditionTrue), running, 0), pod(ready(corev1.ConditionFalse), errored, 1), timeline.HealthUnhealthy},
+ {"lost readiness while running", pod(ready(corev1.ConditionTrue), running, 0), pod(ready(corev1.ConditionFalse), running, 0), timeline.HealthDegraded},
+ {"clean completion is not a failure", pod(ready(corev1.ConditionTrue), running, 0), pod(ready(corev1.ConditionFalse), completed, 0), timeline.HealthHealthy},
+ {"regained readiness", pod(ready(corev1.ConditionFalse), running, 0), pod(ready(corev1.ConditionTrue), running, 0), timeline.HealthHealthy},
+ {"added while not ready yet keeps the startup grace", nil, pod(ready(corev1.ConditionFalse), running, 0), timeline.HealthHealthy},
+ }
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ var old any
+ if tc.old != nil {
+ old = tc.old
+ }
+ if got := classifyTimelineChangeHealth("Pod", old, tc.new, now); got != tc.want {
+ t.Errorf("got %q, want %q", got, tc.want)
+ }
+ })
+ }
+}
diff --git a/internal/mcp/mutation_verification.go b/internal/mcp/mutation_verification.go
index a23ca11e05..3508f3f9bd 100644
--- a/internal/mcp/mutation_verification.go
+++ b/internal/mcp/mutation_verification.go
@@ -766,6 +766,7 @@ func relatedIssuesForObject(ctx context.Context, obj *unstructured.Unstructured)
Namespaces: namespaces,
CanReadClusterScoped: issueClusterScopedAccess(ctx),
CanReadRelated: issueRelatedResourceAccess(ctx),
+ CanReadEvidence: issueEvidenceAccess(ctx),
}, gvk.Group, kind, obj.GetNamespace(), obj.GetName())
if len(matched) == 0 {
return nil
diff --git a/internal/mcp/related_issues_auth_test.go b/internal/mcp/related_issues_auth_test.go
index 67760dd84f..93ace042cd 100644
--- a/internal/mcp/related_issues_auth_test.go
+++ b/internal/mcp/related_issues_auth_test.go
@@ -88,3 +88,30 @@ func initMCPRelatedIssueAuthDiscovery(t *testing.T) {
}
t.Cleanup(k8s.ResetTestDynamicState)
}
+
+func TestMCPCachedIssuesAuthorizeCNPGInventories(t *testing.T) {
+ ctx := withTestUserPerms(t, "cnpg-evidence", nil, []string{"db"})
+ perms := getPermCache().Get("cnpg-evidence", nil)
+ perms.SetCanI("list", "postgresql.cnpg.io", "clusters", "db", true)
+ perms.SetCanI("get", "", "pods", "db", true)
+ perms.SetCanI("list", "", "pods", "db", false)
+ perms.SetCanI("list", "other.example", "pods", "db", true)
+ issue := issues.Issue{ID: "contradiction", Group: "postgresql.cnpg.io", Kind: "Cluster", Namespace: "db", Name: "pg", RequiredReads: []issues.EvidenceRead{
+ {Group: "postgresql.cnpg.io", Resource: "clusters", Namespace: "db", Verb: "list"},
+ {Resource: "pods", Namespace: "db", Verb: "list"},
+ }}
+ options := issues.RelatedIssueOptions{CanReadEvidence: issueEvidenceAccess(ctx)}
+ lookup := func() []issues.Issue {
+ return issues.RelatedIssuesFrom(nil, []issues.Issue{issue}, options, issue.Group, issue.Kind, issue.Namespace, issue.Name)
+ }
+ if len(lookup()) != 0 {
+ t.Fatal("get Pods or list another group's Pods must not authorize the evidence")
+ }
+ perms.SetCanI("list", "", "pods", "db", true)
+ if len(lookup()) != 1 {
+ t.Fatal("authorized inventory withheld")
+ }
+ if options.CanReadEvidence(issues.EvidenceRead{Resource: "pods", Namespace: "other", Verb: "list"}) {
+ t.Fatal("unread namespace authorized")
+ }
+}
diff --git a/internal/mcp/resource_context.go b/internal/mcp/resource_context.go
index 926af406b6..de8a929f17 100644
--- a/internal/mcp/resource_context.go
+++ b/internal/mcp/resource_context.go
@@ -122,6 +122,7 @@ func computeMCPIssueContext(ctx context.Context, cache *k8s.ResourceCache, group
SkipPodTemplateContext: !includeFacts,
CanReadClusterScoped: issueClusterScopedAccess(ctx),
CanReadRelated: issueRelatedResourceAccess(ctx),
+ CanReadEvidence: issueEvidenceAccess(ctx),
}, group, kind, namespace, name)
if len(matched) == 0 {
return nil, nil
@@ -241,3 +242,9 @@ func mcpTopologyForContext(namespace string) (*topo.Topology, topo.ResourceProvi
}
return topology, provider, dyn, true
}
+
+func issueEvidenceAccess(ctx context.Context) func(issues.EvidenceRead) bool {
+ return func(read issues.EvidenceRead) bool {
+ return (read.Namespace == "" || checkNamespaceAccess(ctx, read.Namespace)) && canReadInNamespace(ctx, read.Group, read.Resource, read.Namespace, read.Verb)
+ }
+}
diff --git a/internal/mcp/summary_context.go b/internal/mcp/summary_context.go
index 78f36cf0f7..731086c44a 100644
--- a/internal/mcp/summary_context.go
+++ b/internal/mcp/summary_context.go
@@ -10,6 +10,7 @@
package mcp
import (
+ "context"
"time"
"github.com/skyhook-io/radar/internal/issues"
@@ -30,12 +31,12 @@ import (
// per-hit between a namespaced and a cluster-wide index — search
// returns mixed kinds in one response, so a single index can't get
// both right.
-func newResourceSummaryContextBuilder(namespaces []string) summarycontext.Builder {
+func newResourceSummaryContextBuilder(ctx context.Context, namespaces []string) summarycontext.Builder {
provider := issues.NewCacheProvider()
if provider == nil {
return nil
}
- idx := summarycontext.BuildIssueIndex(provider, namespaces)
+ idx := summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: namespaces, CanReadClusterScoped: issueClusterScopedAccess(ctx), CanReadRelated: issueRelatedResourceAccess(ctx), CanReadEvidence: issueEvidenceAccess(ctx)})
return summarycontext.BuilderFromIndexes(buildSummaryContextTopology(namespaces), idx, idx)
}
@@ -46,15 +47,15 @@ func newResourceSummaryContextBuilder(namespaces []string) summarycontext.Builde
// canReadClusterScopedKind) already gates which cluster-scoped kinds
// are reachable, so composing the cluster-wide index doesn't leak
// rows the user can't see.
-func newSearchSummaryContextBuilder(scanNamespaces []string) summarycontext.Builder {
+func newSearchSummaryContextBuilder(ctx context.Context, scanNamespaces []string) summarycontext.Builder {
provider := issues.NewCacheProvider()
if provider == nil {
return nil
}
- namespacedIdx := summarycontext.BuildIssueIndex(provider, scanNamespaces)
+ namespacedIdx := summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: scanNamespaces, CanReadClusterScoped: issueClusterScopedAccess(ctx), CanReadRelated: issueRelatedResourceAccess(ctx), CanReadEvidence: issueEvidenceAccess(ctx)})
clusterIdx := namespacedIdx
if scanNamespaces != nil {
- clusterIdx = summarycontext.BuildIssueIndex(provider, nil)
+ clusterIdx = summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: nil, CanReadClusterScoped: issueClusterScopedAccess(ctx), CanReadRelated: issueRelatedResourceAccess(ctx), CanReadEvidence: issueEvidenceAccess(ctx)})
}
return summarycontext.BuilderFromIndexes(buildSummaryContextTopology(scanNamespaces), namespacedIdx, clusterIdx)
}
diff --git a/internal/mcp/tools.go b/internal/mcp/tools.go
index 8a397f4391..fb711aa007 100644
--- a/internal/mcp/tools.go
+++ b/internal/mcp/tools.go
@@ -963,7 +963,7 @@ func handleListResources(ctx context.Context, req *mcp.CallToolRequest, input li
if clusterScoped {
idxNamespaces = nil
}
- if builder := newResourceSummaryContextBuilder(idxNamespaces); builder != nil {
+ if builder := newResourceSummaryContextBuilder(ctx, idxNamespaces); builder != nil {
summarycontext.AttachToTypedList(results, objs, builder)
}
}
@@ -1003,7 +1003,7 @@ func listDynamicResources(ctx context.Context, cache *k8s.ResourceCache, kind, g
if clusterScoped {
idxNamespaces = nil
}
- if builder := newResourceSummaryContextBuilder(idxNamespaces); builder != nil {
+ if builder := newResourceSummaryContextBuilder(ctx, idxNamespaces); builder != nil {
summarycontext.AttachToUnstructuredList(allItems, rawItems, builder)
}
}
@@ -3012,7 +3012,8 @@ func handleIssuesTool(ctx context.Context, _ *mcp.CallToolRequest, input issuesI
CanReadClusterScoped: func(kind, group string) bool {
return canReadClusterScopedKind(ctx, kind, group, "list")
},
- CanReadRelated: issueRelatedResourceAccess(ctx),
+ CanReadRelated: issueRelatedResourceAccess(ctx),
+ CanReadEvidence: issueEvidenceAccess(ctx),
}
if input.Filter != "" {
f, err := filter.CachedIssueFilter(input.Filter)
@@ -3327,7 +3328,7 @@ func handleSearch(ctx context.Context, req *mcp.CallToolRequest, input searchInp
// builder routes per-hit by scope; CanReadClusterScoped above
// already gates which cluster-scoped kinds are reachable.
if input.Context != "none" {
- if builder := newSearchSummaryContextBuilder(scanNamespaces); builder != nil {
+ if builder := newSearchSummaryContextBuilder(ctx, scanNamespaces); builder != nil {
opts.SummaryBuilder = search.SummaryBuilderFunc(builder)
}
}
diff --git a/internal/mcp/tools_diagnose.go b/internal/mcp/tools_diagnose.go
index 30d74c2c59..1174b4cfe0 100644
--- a/internal/mcp/tools_diagnose.go
+++ b/internal/mcp/tools_diagnose.go
@@ -612,6 +612,7 @@ func handleGitOpsDiagnose(ctx context.Context, input diagnoseInput, canonicalKin
Namespaces: issueNamespacesForResource(input.Namespace),
CanReadClusterScoped: issueClusterScopedAccess(ctx),
CanReadRelated: issueRelatedResourceAccess(ctx),
+ CanReadEvidence: issueEvidenceAccess(ctx),
}, group, canonicalKind, input.Namespace, input.Name),
Warnings: k8score.EnrichRuntimeObjectWarnings(u),
}
diff --git a/internal/podlogs/pods.go b/internal/podlogs/pods.go
new file mode 100644
index 0000000000..9aeca14638
--- /dev/null
+++ b/internal/podlogs/pods.go
@@ -0,0 +1,119 @@
+package podlogs
+
+import (
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/health"
+ corev1 "k8s.io/api/core/v1"
+)
+
+// ContainerInfo contains compact per-container runtime status for the UI.
+type ContainerInfo struct {
+ Name string `json:"name"`
+ Init bool `json:"init,omitempty"`
+ Ready bool `json:"ready"`
+ RestartCount int32 `json:"restartCount"`
+}
+
+// PodInfo contains compact runtime status about a pod for workload views.
+type PodInfo struct {
+ Name string `json:"name"`
+ Containers []string `json:"containers"`
+ Ready bool `json:"ready"`
+ Phase string `json:"phase,omitempty"`
+ NodeName string `json:"nodeName,omitempty"`
+ HealthLevel string `json:"healthLevel,omitempty"`
+ Reason string `json:"reason,omitempty"`
+ Message string `json:"message,omitempty"`
+ RestartCount int32 `json:"restartCount,omitempty"`
+ LastTerminationReason string `json:"lastTerminationReason,omitempty"`
+ CreatedAt string `json:"createdAt,omitempty"`
+ ContainerStatuses []ContainerInfo `json:"containerStatuses,omitempty"`
+ StepID string `json:"stepID,omitempty"`
+ StepName string `json:"stepName,omitempty"`
+ StepPhase string `json:"stepPhase,omitempty"`
+ RevisionIdentity string `json:"revisionIdentity,omitempty"`
+ UpdatedRevision *bool `json:"updatedRevision,omitempty"`
+}
+
+func IsPodReady(pod *corev1.Pod) bool {
+ if pod.Status.Phase != corev1.PodRunning {
+ return false
+ }
+ for _, cs := range pod.Status.ContainerStatuses {
+ if !cs.Ready {
+ return false
+ }
+ }
+ return true
+}
+
+func BuildPodInfo(pod *corev1.Pod, now time.Time) PodInfo {
+ containers := make([]string, 0, len(pod.Spec.Containers)+len(pod.Spec.InitContainers))
+ containerStatuses := make([]ContainerInfo, 0, len(pod.Status.InitContainerStatuses)+len(pod.Status.ContainerStatuses))
+ for _, c := range pod.Spec.InitContainers {
+ containers = append(containers, c.Name)
+ }
+ for _, c := range pod.Spec.Containers {
+ containers = append(containers, c.Name)
+ }
+ for _, cs := range pod.Status.InitContainerStatuses {
+ containerStatuses = append(containerStatuses, ContainerInfo{
+ Name: cs.Name,
+ Init: true,
+ Ready: cs.Ready,
+ RestartCount: cs.RestartCount,
+ })
+ }
+ for _, cs := range pod.Status.ContainerStatuses {
+ containerStatuses = append(containerStatuses, ContainerInfo{
+ Name: cs.Name,
+ Ready: cs.Ready,
+ RestartCount: cs.RestartCount,
+ })
+ }
+ verdict := health.Pod(pod, now)
+ displayLevel := health.PodDisplayLevel(pod, now)
+ if displayLevel != verdict.Level {
+ verdict.Level = displayLevel
+ if verdict.Reason == "" {
+ verdict.Reason = health.PodProblemReason(pod, now)
+ }
+ if verdict.Message == "" {
+ verdict.Message = health.PodProblemMessage(pod)
+ }
+ }
+ restartCount, lastTerminationReason := health.PodRestartContext(pod)
+ createdAt := ""
+ if !pod.CreationTimestamp.IsZero() {
+ createdAt = pod.CreationTimestamp.Time.Format(time.RFC3339)
+ }
+ annotations := pod.GetAnnotations()
+ labels := pod.GetLabels()
+ return PodInfo{
+ Name: pod.Name,
+ Containers: containers,
+ Ready: IsPodReady(pod),
+ Phase: string(pod.Status.Phase),
+ NodeName: pod.Spec.NodeName,
+ HealthLevel: string(verdict.Level),
+ Reason: verdict.Reason,
+ Message: verdict.Message,
+ RestartCount: restartCount,
+ LastTerminationReason: lastTerminationReason,
+ CreatedAt: createdAt,
+ ContainerStatuses: containerStatuses,
+ StepID: annotations["workflows.argoproj.io/node-id"],
+ StepName: annotations["workflows.argoproj.io/node-name"],
+ StepPhase: labels["workflows.argoproj.io/phase"],
+ }
+}
+
+func BuildPodInfos(pods []*corev1.Pod) []PodInfo {
+ infos := make([]PodInfo, 0, len(pods))
+ now := time.Now()
+ for _, pod := range pods {
+ infos = append(infos, BuildPodInfo(pod, now))
+ }
+ return infos
+}
diff --git a/internal/podlogs/query.go b/internal/podlogs/query.go
new file mode 100644
index 0000000000..c784a68275
--- /dev/null
+++ b/internal/podlogs/query.go
@@ -0,0 +1,25 @@
+package podlogs
+
+import "strconv"
+
+// ParseSinceSeconds parses sinceSeconds query parameter, returning nil if not set
+func ParseSinceSeconds(str string) *int64 {
+ if str == "" {
+ return nil
+ }
+ if s, err := strconv.ParseInt(str, 10, 64); err == nil && s > 0 {
+ return &s
+ }
+ return nil
+}
+
+// ParseTailLines parses tailLines query parameter with a default value
+func ParseTailLines(str string, defaultVal int64) int64 {
+ if str == "" {
+ return defaultVal
+ }
+ if t, err := strconv.ParseInt(str, 10, 64); err == nil && t > 0 {
+ return t
+ }
+ return defaultVal
+}
diff --git a/internal/podlogs/snapshot.go b/internal/podlogs/snapshot.go
new file mode 100644
index 0000000000..c6a1e45b1b
--- /dev/null
+++ b/internal/podlogs/snapshot.go
@@ -0,0 +1,284 @@
+package podlogs
+
+import (
+ "context"
+ "fmt"
+ "io"
+ "sort"
+ "strings"
+ "sync"
+ "time"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+ corev1 "k8s.io/api/core/v1"
+ "k8s.io/client-go/kubernetes"
+)
+
+type Entry struct {
+ Pod string `json:"pod"`
+ Container string `json:"container"`
+ Timestamp string `json:"timestamp"`
+ Content string `json:"content"`
+ SourceLabel string `json:"sourceLabel,omitempty"`
+ // Previous marks a line from the container run before its last restart.
+ Previous bool `json:"previous,omitempty"`
+ // Parsed from a structured line by sources that know their log format.
+ Level string `json:"level,omitempty"`
+ Logger string `json:"logger,omitempty"`
+ Message string `json:"message,omitempty"`
+}
+
+// CollectPods fetches logs from all pods concurrently. Non-nil even
+// when nothing is retrievable (e.g. every pod is crashlooping) — a nil slice
+// marshals as JSON null and consumers expect an array.
+type Snapshot struct {
+ SourcePods map[string]bool
+ Logs []Entry
+ Notice string
+ // Clipped lists the sources that reached maxSnapshotSourceBytes, with the
+ // timestamp of the last line kept from each, so a caller filtering by time
+ // can tell whether the clip cut into its window.
+ Clipped []Clip
+ shown int
+ total int
+ errors []string
+}
+
+type Clip struct {
+ Pod, Container string
+ Previous bool
+ Last time.Time
+}
+
+// Source is one container run to read: the current run, or with
+// Previous the one before the last restart.
+type Source struct {
+ Pod, Container string
+ Previous bool
+ running bool
+ created time.Time
+}
+
+func (s Source) key() string {
+ if s.Previous {
+ return s.Pod + "/" + s.Container + " (previous run)"
+ }
+ return s.Pod + "/" + s.Container
+}
+
+const maxSnapshotSources = 40
+
+const maxSnapshotSourceBytes int64 = 64 * 1024
+
+func CollectPods(ctx context.Context, client kubernetes.Interface, namespace string, pods []*corev1.Pod, container string, tailLines int64, sinceSeconds *int64, bounded bool) Snapshot {
+ sources := []Source{}
+ for _, pod := range pods {
+ for _, c := range k8s.GetContainersForPod(pod, container, true) {
+ if bounded {
+ started := false
+ for _, statuses := range [][]corev1.ContainerStatus{pod.Status.ContainerStatuses, pod.Status.InitContainerStatuses, pod.Status.EphemeralContainerStatuses} {
+ for _, status := range statuses {
+ if status.Name == c && (status.State.Running != nil || status.State.Terminated != nil) {
+ started = true
+ }
+ }
+ }
+ if !started {
+ continue
+ }
+ }
+ sources = append(sources, NewSource(pod, c, false))
+ }
+ }
+ return CollectSources(ctx, client, namespace, sources, tailLines, sinceSeconds, bounded)
+}
+
+func NewSource(pod *corev1.Pod, container string, previous bool) Source {
+ return Source{Pod: pod.Name, Container: container, Previous: previous, running: pod.Status.Phase == corev1.PodRunning, created: pod.CreationTimestamp.Time}
+}
+
+// CollectSources reads each source concurrently; bounded caps the source
+// count, bytes per source and overall time.
+func CollectSources(ctx context.Context, client kubernetes.Interface, namespace string, sources []Source, tailLines int64, sinceSeconds *int64, bounded bool) Snapshot {
+ sort.Slice(sources, func(i, j int) bool {
+ if bounded && sources[i].running != sources[j].running {
+ return sources[i].running
+ }
+ if bounded && !sources[i].created.Equal(sources[j].created) {
+ return sources[i].created.After(sources[j].created)
+ }
+ if sources[i].Pod != sources[j].Pod {
+ return sources[i].Pod < sources[j].Pod
+ }
+ if sources[i].Container != sources[j].Container {
+ return sources[i].Container < sources[j].Container
+ }
+ return sources[i].Previous && !sources[j].Previous
+ })
+ total := len(sources)
+ if bounded && len(sources) > maxSnapshotSources {
+ sources = sources[:maxSnapshotSources]
+ }
+ if bounded {
+ tailLines = min(tailLines, 1000)
+ var cancel context.CancelFunc
+ ctx, cancel = context.WithTimeout(ctx, 30*time.Second)
+ defer cancel()
+ }
+ result := Snapshot{Logs: []Entry{}, SourcePods: map[string]bool{}, shown: len(sources), total: total}
+ for _, src := range sources {
+ result.SourcePods[src.Pod] = true
+ }
+ var mu sync.Mutex
+ var wg sync.WaitGroup
+ concurrency := len(sources)
+ if bounded {
+ concurrency = min(concurrency, 8)
+ }
+ sem := make(chan struct{}, concurrency)
+ for _, src := range sources {
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ select {
+ case sem <- struct{}{}:
+ defer func() { <-sem }()
+ case <-ctx.Done():
+ mu.Lock()
+ result.errors = append(result.errors, src.key()+": request cancelled")
+ mu.Unlock()
+ return
+ }
+ entries, clipped, err := fetchPodContainerLogs(ctx, client, namespace, src.Pod, src.Container, tailLines, sinceSeconds, bounded, src.Previous)
+ mu.Lock()
+ defer mu.Unlock()
+ if err != nil {
+ result.errors = append(result.errors, src.key()+": "+err.Error())
+ }
+ if clipped {
+ clip := Clip{Pod: src.Pod, Container: src.Container, Previous: src.Previous}
+ if n := len(entries); n > 0 {
+ clip.Last, _ = time.Parse(time.RFC3339Nano, entries[n-1].Timestamp)
+ }
+ result.Clipped = append(result.Clipped, clip)
+ }
+ result.Logs = append(result.Logs, entries...)
+ }()
+ }
+ wg.Wait()
+ sort.Strings(result.errors)
+ result.Notice = result.Summarize(func(Clip) bool { return true })
+ return result
+}
+
+// Summarize summarizes what the snapshot could not show. counts decides which
+// clipped sources are worth reporting — a clip past the caller's window cut
+// nothing the caller asked for.
+func (s Snapshot) Summarize(counts func(Clip) bool) string {
+ notices := []string{}
+ if s.total > s.shown {
+ notices = append(notices, fmt.Sprintf("Showing %d of %d container sources. Narrow the scope to see other sources.", s.shown, s.total))
+ }
+ truncated := 0
+ for _, clip := range s.Clipped {
+ if counts(clip) {
+ truncated++
+ }
+ }
+ if truncated > 0 {
+ notices = append(notices, countNoun(truncated, "source", "sources")+" reached the 64 KiB snapshot limit.")
+ }
+ if len(s.errors) > 0 {
+ notices = append(notices, fmt.Sprintf("%s could not be read: %s", countNoun(len(s.errors), "source", "sources"), strings.Join(s.errors[:min(3, len(s.errors))], "; ")))
+ }
+ return strings.Join(notices, " ")
+}
+
+func countNoun(n int, singular, plural string) string {
+ if n == 1 {
+ return "1 " + singular
+ }
+ return fmt.Sprintf("%d %s", n, plural)
+}
+
+func fetchPodContainerLogs(ctx context.Context, client kubernetes.Interface, namespace, podName, containerName string, tailLines int64, sinceSeconds *int64, bounded, previous bool) ([]Entry, bool, error) {
+ var limit *int64
+ if bounded {
+ n := maxSnapshotSourceBytes + 1
+ limit = &n
+ }
+ // Zero reads from the start of the since window instead of its tail, so a
+ // bounded read of a past interval returns that interval's first lines.
+ var tail *int64
+ if tailLines > 0 {
+ tail = &tailLines
+ }
+ stream, err := k8score.GetContainerLogs(ctx, client, namespace, podName, containerName, k8score.LogOptions{
+ TailLines: tail, SinceSeconds: sinceSeconds, Timestamps: true, LimitBytes: limit, Previous: previous,
+ })
+ if err != nil {
+ return nil, false, err
+ }
+ defer stream.Close()
+ var reader io.Reader = stream
+ if limit != nil {
+ reader = io.LimitReader(stream, *limit)
+ }
+ content, err := io.ReadAll(reader)
+ if err != nil {
+ return nil, false, err
+ }
+ clipped := bounded && int64(len(content)) > maxSnapshotSourceBytes
+ if clipped {
+ content = content[:maxSnapshotSourceBytes]
+ if last := strings.LastIndexByte(string(content), '\n'); last >= 0 {
+ content = content[:last+1]
+ } else {
+ content = nil
+ }
+ }
+ entries := []Entry{}
+ for _, line := range strings.Split(string(content), "\n") {
+ if line == "" {
+ continue
+ }
+ ts, text := ParseLine(line)
+ entries = append(entries, Entry{Pod: podName, Container: containerName, Timestamp: ts, Content: text, Previous: previous})
+ }
+ return entries, clipped, nil
+}
+
+// Sort sorts log entries by timestamp using efficient sort
+func Sort(logs []Entry) {
+ sort.SliceStable(logs, func(i, j int) bool {
+ left, le := time.Parse(time.RFC3339Nano, logs[i].Timestamp)
+ right, re := time.Parse(time.RFC3339Nano, logs[j].Timestamp)
+ if le == nil && re == nil && !left.Equal(right) {
+ return left.Before(right)
+ }
+ if (le == nil) != (re == nil) {
+ return le != nil
+ }
+ if le != nil && logs[i].Timestamp != logs[j].Timestamp {
+ return logs[i].Timestamp < logs[j].Timestamp
+ }
+ if logs[i].Pod != logs[j].Pod {
+ return logs[i].Pod < logs[j].Pod
+ }
+ return logs[i].Container < logs[j].Container
+ })
+}
+
+// ParseLine extracts timestamp from a log line (format: 2024-01-20T10:30:00.123456789Z content)
+func ParseLine(line string) (timestamp, content string) {
+ // K8s timestamps are in RFC3339Nano format at the start of the line
+ if len(line) > 30 && line[4] == '-' && line[7] == '-' && line[10] == 'T' {
+ // Find the space after timestamp
+ spaceIdx := strings.Index(line, " ")
+ if spaceIdx > 20 && spaceIdx < 40 {
+ return line[:spaceIdx], line[spaceIdx+1:]
+ }
+ }
+ return "", line
+}
diff --git a/internal/prometheus/cnpg_client_test.go b/internal/prometheus/cnpg_client_test.go
new file mode 100644
index 0000000000..80299fd4fa
--- /dev/null
+++ b/internal/prometheus/cnpg_client_test.go
@@ -0,0 +1,90 @@
+package prometheus
+
+import (
+ "context"
+ "fmt"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+func TestCNPGReadsKeepCapturedMetricsClient(t *testing.T) {
+ var queries atomic.Int64
+ upstream := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ query := r.URL.Query().Get("query")
+ if query != "up" {
+ queries.Add(1)
+ if !strings.Contains(query, `cluster="east"`) {
+ t.Errorf("query lost captured scope: %s", query)
+ }
+ }
+ if r.URL.Path == "/api/v1/query_range" {
+ _, _ = fmt.Fprint(w, `{"status":"success","data":{"resultType":"matrix","result":[]}}`)
+ return
+ }
+ value := "1"
+ if strings.Contains(query, "used_bytes") {
+ value = "25"
+ } else if strings.Contains(query, "capacity_bytes") {
+ value = "100"
+ }
+ _, _ = fmt.Fprintf(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{"persistentvolumeclaim":"pg-1"},"value":[1700000000,%q]}]}}`, value)
+ }))
+ defer upstream.Close()
+ client := &Client{
+ manualURL: upstream.URL, httpClient: upstream.Client(),
+ workloadScope: &prom.WorkloadMetricsScope{ClusterLabels: map[string]string{"cluster": "east"}},
+ }
+ clientMu.Lock()
+ previous := globalClient
+ globalClient = &Client{workloadScope: &prom.WorkloadMetricsScope{ClusterLabels: map[string]string{"cluster": "west"}}}
+ clientMu.Unlock()
+ t.Cleanup(func() { clientMu.Lock(); globalClient = previous; clientMu.Unlock() })
+
+ ctx := context.Background()
+ matchers, iso, err := client.ResolveCNPGScope(ctx, "pg", CNPGInstanceSelector("pg", "pg"), nil, time.Minute)
+ if err != nil || matchers != `cluster="east"` || iso.Labels["cluster"] != "east" {
+ t.Fatalf("exporter scope = %q %+v %v", matchers, iso, err)
+ }
+ pvcMatchers, _, err := client.ResolvePVCScope(ctx, "pg", []string{"pg-1"}, nil, time.Minute)
+ if err != nil || pvcMatchers != matchers {
+ t.Fatalf("PVC scope = %q %v", pvcMatchers, err)
+ }
+ usage := client.QueryPVCUsage(ctx, "pg", []string{"pg-1"}, nil)
+ if usage.Status != PVCUsageAvailable || usage.Usage["pg-1"].Ratio != .25 {
+ t.Fatalf("PVC usage = %+v", usage)
+ }
+ rng, _ := ParseCNPGHistoryRange("15m")
+ charts, err := client.QueryCNPGHistory(ctx, CNPGHistoryRequest{Namespace: "pg", Cluster: "pg", Range: rng, End: time.Now(), Matchers: matchers, PVCMatchers: pvcMatchers, Claims: []string{"pg-1"}})
+ if err != nil || len(charts) == 0 {
+ t.Fatalf("history = %d charts, %v", len(charts), err)
+ }
+ if _, err := client.QueryCNPGFleetLag(ctx, "pg", []string{"pg"}, matchers); err != nil {
+ t.Fatal(err)
+ }
+ if _, err := client.QueryCNPGFleetSlots(ctx, "pg", []string{"pg"}, matchers); err != nil {
+ t.Fatal(err)
+ }
+ if _, err := client.QueryCNPGDiskGrowth(ctx, "pg", []string{"pg-1"}, time.Hour, pvcMatchers); err != nil {
+ t.Fatal(err)
+ }
+ if queries.Load() == 0 {
+ t.Fatal("no queries reached captured client")
+ }
+
+ client.mu.Lock()
+ client.retired = true
+ client.mu.Unlock()
+ before := queries.Load()
+ if usage := client.QueryPVCUsage(ctx, "pg", []string{"pg-1"}, nil); usage.Status != PVCUsageNoPrometheus {
+ t.Fatalf("retired client's usage = %+v", usage)
+ }
+ if queries.Load() != before {
+ t.Fatal("retired client continued querying")
+ }
+}
diff --git a/internal/prometheus/cnpg_history.go b/internal/prometheus/cnpg_history.go
new file mode 100644
index 0000000000..b225c3cfc8
--- /dev/null
+++ b/internal/prometheus/cnpg_history.go
@@ -0,0 +1,804 @@
+package prometheus
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "math"
+ "regexp"
+ "sort"
+ "strconv"
+ "strings"
+ "sync"
+ "time"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+// CNPG history reads only server-built PromQL over the CloudNativePG exporter
+// and kubelet volume stats. A Cluster is selected by namespace plus its
+// instance Pod names, `-`, never by a `cluster` label: shared
+// Prometheus setups use `cluster` for the Kubernetes cluster, while a CNPG
+// scrape config may relabel it to the CNPG Cluster name, so its meaning is not
+// knowable from the series. The pod pattern keeps deleted instances in range.
+
+const (
+ CNPGHistoryStateOK = "ok"
+ CNPGHistoryStateNoSeries = "noSeries"
+ CNPGHistoryStateEmpty = "empty"
+ CNPGHistoryStateDenied = "denied"
+ CNPGHistoryStateError = "error"
+ CNPGHistoryStateNotRead = "notRead"
+ CNPGHistoryStateAmbiguous = "ambiguous"
+ CNPGMetricLookback = 5 * time.Minute
+
+ cnpgHistoryMaxSeries = 12
+ cnpgHistoryConcurrency = 4
+)
+
+// CNPGHistoryRange is one of the ranges the history endpoint accepts. Steps
+// keep every chart at 60–144 points.
+type CNPGHistoryRange struct {
+ Name string
+ Duration time.Duration
+ Step time.Duration
+}
+
+var cnpgHistoryRanges = []CNPGHistoryRange{
+ {"15m", 15 * time.Minute, 15 * time.Second},
+ {"1h", time.Hour, 30 * time.Second},
+ {"6h", 6 * time.Hour, 3 * time.Minute},
+ {"24h", 24 * time.Hour, 10 * time.Minute},
+}
+
+// ParseCNPGHistoryRange accepts 15m, 1h, 6h or 24h; empty means 1h.
+func ParseCNPGHistoryRange(raw string) (CNPGHistoryRange, bool) {
+ if raw == "" {
+ raw = "1h"
+ }
+ for _, r := range cnpgHistoryRanges {
+ if r.Name == raw {
+ return r, true
+ }
+ }
+ return CNPGHistoryRange{}, false
+}
+
+// CNPGInstanceSelector is the matcher set for a Cluster's instance series.
+func CNPGInstanceSelector(namespace, cluster string) string {
+ return CNPGInstancesSelector(namespace, []string{cluster})
+}
+
+// CNPGInstancesSelector selects the instance Pods of several Clusters of one
+// namespace in one matcher.
+func CNPGInstancesSelector(namespace string, clusters []string) string {
+ names := make([]string, len(clusters))
+ for i, c := range clusters {
+ names[i] = regexp.QuoteMeta(c)
+ }
+ sort.Strings(names)
+ alt := strings.Join(names, "|")
+ if len(names) > 1 {
+ alt = "(" + alt + ")"
+ }
+ return "namespace=" + strconv.Quote(namespace) + ",pod=~" + strconv.Quote("^"+alt+"-[0-9]+$")
+}
+
+var cnpgInstancePod = regexp.MustCompile(`^(.+)-[0-9]+$`)
+
+// CNPGClusterOfPod returns the Cluster an instance Pod name belongs to among
+// the given names, or "" when none matches exactly.
+func CNPGClusterOfPod(pod string, clusters map[string]bool) string {
+ m := cnpgInstancePod.FindStringSubmatch(pod)
+ if m == nil || !clusters[m[1]] {
+ return ""
+ }
+ return m[1]
+}
+
+// ResolveCNPGScope decides the cluster-identity matchers for one namespace's
+// CNPG exporter series over window (0 for an instant read).
+func (client *Client) ResolveCNPGScope(ctx context.Context, namespace, selector string, anchors []prom.WorkloadPodIdentity, window time.Duration) (string, SeriesIsolation, error) {
+ return client.resolveScope(ctx, namespace, scopeProbe{metric: "cnpg_collector_up", key: "pod", selectors: []string{selector}, window: window}, anchors, nil)
+}
+
+// CNPGHistoryThreshold is a reference line on a chart.
+type CNPGHistoryThreshold struct {
+ Value float64 `json:"value"`
+ Label string `json:"label"`
+}
+
+// CNPGHistoryChart is one chart: its series, or why there are none.
+// Coverage counts evaluation steps with at least one sample.
+type CNPGHistoryChart struct {
+ ID string `json:"id"`
+ Title string `json:"title"`
+ Unit string `json:"unit"`
+ Source string `json:"source"`
+ SeriesBy string `json:"seriesBy"`
+ State string `json:"state"`
+ Reason string `json:"reason,omitempty"`
+ Grant *auth.Grant `json:"grant,omitempty"`
+ Thresholds []CNPGHistoryThreshold `json:"thresholds,omitempty"`
+ Series []prom.Series `json:"series"`
+ Omitted int `json:"omitted,omitempty"`
+ Steps int `json:"steps"`
+ Covered int `json:"covered"`
+}
+
+type cnpgHistoryQuery struct {
+ label string
+ expr string
+}
+
+type cnpgHistoryDef struct {
+ id, title, unit, source, seriesBy string
+ thresholds []CNPGHistoryThreshold
+ queries []cnpgHistoryQuery
+ pvc bool
+ // presence selects the family the chart derives from; when the chart is
+ // empty but this has samples, the family is scraped and emptyReason holds.
+ presence string
+ emptyReason string
+}
+
+func cnpgRateWindow(step time.Duration) string {
+ return max(2*time.Minute, 2*step).Round(time.Second).String()
+}
+
+// cnpgHistoryDefs builds the chart queries. `max by` before any sum collapses
+// duplicate scrapes of one Pod (HA Prometheus pairs, two scrape jobs). sel is
+// the instance selector with any cluster-identity matchers; pvcSel selects the
+// Cluster's claims and is empty when they are not readable.
+func cnpgHistoryDefs(sel, pvcSel string, step time.Duration) []cnpgHistoryDef {
+ w := cnpgRateWindow(step)
+ clients := sel + `,usename!="streaming_replica",application_name!="cnpg_metrics_exporter"`
+ perPodRate := func(metric string) string {
+ return "sum by (pod) (max by (pod,datname) (rate(" + metric + "{" + sel + "}[" + w + "])))"
+ }
+ clusterRate := func(metric, by string) string {
+ return "sum(max by (" + by + ") (rate(" + metric + "{" + sel + "}[" + w + "])))"
+ }
+ up := "(0 * max(cnpg_collector_up{" + sel + "} == 1))"
+ family := func(name string) string { return "{__name__=~" + strconv.Quote(name) + "," + sel + "}" }
+ checkpoints := func(kind string) string {
+ return "(" + clusterRate("cnpg_pg_stat_checkpointer_checkpoints_"+kind, "pod") + " or " + clusterRate("cnpg_pg_stat_bgwriter_checkpoints_"+kind, "pod") + ") * 60"
+ }
+ defs := []cnpgHistoryDef{
+ {
+ id: "replicationLag", title: "Replay lag per standby", unit: "seconds", seriesBy: "pod",
+ source: "cnpg_pg_replication_lag, while cnpg_pg_replication_in_recovery = 1",
+ thresholds: []CNPGHistoryThreshold{{Value: 5, Label: "5 s"}, {Value: 30, Label: "30 s"}},
+ queries: []cnpgHistoryQuery{{expr: "max by (pod) (cnpg_pg_replication_lag{" + sel + "}) and on (pod) (max by (pod) (cnpg_pg_replication_in_recovery{" + sel + "}) == 1)"}},
+ presence: family("cnpg_pg_replication_lag"),
+ emptyReason: "no instance was a standby in this range",
+ },
+ {
+ id: "sessions", title: "Client sessions by state (all instances)", unit: "count", seriesBy: "state",
+ source: "cnpg_backends_total, excluding streaming_replica and the metrics exporter",
+ queries: []cnpgHistoryQuery{
+ {expr: "sum by (state) (max by (pod,state,datname,usename,application_name) (cnpg_backends_total{" + clients + "}))"},
+ {label: "all states", expr: "sum(max by (pod,state,datname,usename,application_name) (cnpg_backends_total{" + clients + "})) or " + up},
+ },
+ },
+ {
+ id: "waiting", title: "Sessions waiting on locks", unit: "count", seriesBy: "pod",
+ source: "cnpg_backends_waiting_total",
+ queries: []cnpgHistoryQuery{{expr: "max by (pod) (cnpg_backends_waiting_total{" + sel + "})"}},
+ },
+ {
+ id: "tps", title: "Transactions per second", unit: "per second", seriesBy: "series",
+ source: "rate of cnpg_pg_stat_database_xact_commit / xact_rollback, all instances and databases",
+ queries: []cnpgHistoryQuery{
+ {label: "commits", expr: clusterRate("cnpg_pg_stat_database_xact_commit", "pod,datname")},
+ {label: "rollbacks", expr: clusterRate("cnpg_pg_stat_database_xact_rollback", "pod,datname")},
+ },
+ },
+ {
+ id: "walArchive", title: "WAL archiving per minute", unit: "per minute", seriesBy: "series",
+ source: "rate of cnpg_pg_stat_archiver_archived_count / failed_count",
+ queries: []cnpgHistoryQuery{
+ {label: "archived", expr: clusterRate("cnpg_pg_stat_archiver_archived_count", "pod") + " * 60"},
+ {label: "failed", expr: clusterRate("cnpg_pg_stat_archiver_failed_count", "pod") + " * 60"},
+ },
+ },
+ {
+ id: "walSize", title: "WAL on disk", unit: "bytes", seriesBy: "pod",
+ source: `cnpg_collector_pg_wal{value="size"} (bytes of WAL segments, not filesystem use)`,
+ queries: []cnpgHistoryQuery{{expr: "max by (pod) (cnpg_collector_pg_wal{" + sel + `,value="size"})`}},
+ },
+ {
+ id: "pvcUsed", title: "Volume used", unit: "percent", seriesBy: "persistentvolumeclaim", pvc: true,
+ source: "kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes",
+ thresholds: []CNPGHistoryThreshold{{Value: 80, Label: "80%"}, {Value: 90, Label: "90%"}},
+ queries: []cnpgHistoryQuery{{expr: "100 * max by (persistentvolumeclaim) (kubelet_volume_stats_used_bytes{" + pvcSel + "}) / max by (persistentvolumeclaim) (kubelet_volume_stats_capacity_bytes{" + pvcSel + "})"}},
+ },
+ {
+ id: "databaseSize", title: "Database size", unit: "bytes", seriesBy: "datname",
+ source: "cnpg_pg_database_size_bytes (largest reported by any instance)",
+ queries: []cnpgHistoryQuery{{expr: "max by (datname) (cnpg_pg_database_size_bytes{" + sel + `,datname!~"template0|template1"})`}},
+ },
+ {
+ id: "tempBytes", title: "Temporary file writes", unit: "bytes/s", seriesBy: "pod",
+ source: "rate of cnpg_pg_stat_database_temp_bytes",
+ queries: []cnpgHistoryQuery{{expr: perPodRate("cnpg_pg_stat_database_temp_bytes")}},
+ },
+ {
+ id: "deadlocks", title: "Deadlocks per minute", unit: "per minute", seriesBy: "pod",
+ source: "rate of cnpg_pg_stat_database_deadlocks",
+ queries: []cnpgHistoryQuery{{expr: perPodRate("cnpg_pg_stat_database_deadlocks") + " * 60"}},
+ },
+ {
+ id: "checkpoints", title: "Checkpoints per minute", unit: "per minute", seriesBy: "series",
+ source: "rate of cnpg_pg_stat_checkpointer_checkpoints_{timed,req} (PostgreSQL 17+) or cnpg_pg_stat_bgwriter_checkpoints_{timed,req}",
+ queries: []cnpgHistoryQuery{
+ {label: "requested", expr: checkpoints("req")},
+ {label: "timed", expr: checkpoints("timed")},
+ },
+ presence: family("cnpg_pg_stat_(checkpointer|bgwriter)_checkpoints_(req|timed)"),
+ emptyReason: "fewer than two samples in each rate window",
+ },
+ }
+ return defs
+}
+
+// Bounds excludes reused Pod names from a previous Cluster incarnation. The
+// lookback margin also excludes predecessor samples from rates and instant
+// selectors (Prometheus' default five-minute lookback).
+func (r CNPGHistoryRange) Bounds(end, createdAt time.Time) (time.Time, time.Time) {
+ end = end.Truncate(r.Step)
+ start := end.Add(-r.Duration)
+ if !createdAt.IsZero() {
+ earliest := createdAt.Add(max(CNPGMetricLookback, 2*r.Step)).UTC()
+ aligned := earliest.Truncate(r.Step)
+ if aligned.Before(earliest) {
+ aligned = aligned.Add(r.Step)
+ }
+ if aligned.After(start) {
+ start = aligned
+ }
+ }
+ return start, end
+}
+
+// CNPGHistoryRequest is what the server decided the caller may read.
+type CNPGHistoryRequest struct {
+ Namespace string
+ Cluster string
+ Range CNPGHistoryRange
+ End time.Time
+ CreatedAt time.Time
+ // Matchers are cluster-identity matchers from ResolveCNPGScope for the
+ // exporter series; PVCMatchers from ResolvePVCScope for the claims.
+ Matchers string
+ PVCMatchers string
+ PVCScopeState string
+ // PodsDenied / PVCDenied carry the grant that is missing, empty when allowed.
+ PodsDenied *auth.Grant
+ PVCDenied *auth.Grant
+ // PVCReason explains an empty Claims when the claims were not denied.
+ PVCReason string
+ Claims []string
+}
+
+// QueryCNPGHistory runs every chart query, a few at a time.
+func (client *Client) QueryCNPGHistory(ctx context.Context, req CNPGHistoryRequest) ([]CNPGHistoryChart, error) {
+ if client == nil {
+ return nil, errors.New("Prometheus client not initialized")
+ }
+ return queryCNPGHistory(ctx, client, req), nil
+}
+
+func queryCNPGHistory(ctx context.Context, q seriesQuerier, req CNPGHistoryRequest) []CNPGHistoryChart {
+ sel := withScope(CNPGInstanceSelector(req.Namespace, req.Cluster), req.Matchers)
+ pvcSel := ""
+ if len(req.Claims) > 0 {
+ pvcSel = withScope(cnpgClaimSelector(req.Namespace, req.Claims), req.PVCMatchers)
+ }
+ start, end := req.Range.Bounds(req.End, req.CreatedAt)
+ steps := max(0, int(end.Sub(start)/req.Range.Step)+1)
+
+ defs := cnpgHistoryDefs(sel, pvcSel, req.Range.Step)
+ charts := make([]CNPGHistoryChart, len(defs))
+ sem := make(chan struct{}, cnpgHistoryConcurrency)
+ var wg sync.WaitGroup
+ for i, d := range defs {
+ c := &charts[i]
+ *c = CNPGHistoryChart{ID: d.id, Title: d.title, Unit: d.unit, Source: d.source, SeriesBy: d.seriesBy, Thresholds: d.thresholds, Series: []prom.Series{}, Steps: steps}
+ switch {
+ case d.pvc && req.PVCDenied != nil:
+ c.State, c.Grant = CNPGHistoryStateDenied, req.PVCDenied
+ continue
+ case !d.pvc && req.PodsDenied != nil:
+ c.State, c.Grant = CNPGHistoryStateDenied, req.PodsDenied
+ continue
+ case d.pvc && req.PVCScopeState != "":
+ c.State, c.Reason = req.PVCScopeState, req.PVCReason
+ continue
+ case d.pvc && pvcSel == "":
+ c.State, c.Reason = CNPGHistoryStateNotRead, req.PVCReason
+ if c.Reason == "" {
+ c.Reason = "no volumes owned by this cluster to chart"
+ }
+ continue
+ }
+ if start.After(end) {
+ c.State, c.Reason = CNPGHistoryStateEmpty, "Waiting for samples after this Cluster was created"
+ continue
+ }
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ select {
+ case sem <- struct{}{}:
+ case <-ctx.Done():
+ c.State, c.Reason = CNPGHistoryStateError, ctx.Err().Error()
+ return
+ }
+ defer func() { <-sem }()
+ runCNPGHistoryChart(ctx, q, d, start, end, req.Range.Step, c)
+ }()
+ }
+ wg.Wait()
+ return charts
+}
+
+func runCNPGHistoryChart(ctx context.Context, q seriesQuerier, d cnpgHistoryDef, start, end time.Time, step time.Duration, c *CNPGHistoryChart) {
+ series := []prom.Series{}
+ for _, query := range d.queries {
+ res, err := q.QueryRange(ctx, query.expr, start, end, step)
+ if err != nil {
+ c.State, c.Reason = CNPGHistoryStateError, "Prometheus query failed: "+truncate(err.Error(), 200)
+ c.Series = []prom.Series{}
+ return
+ }
+ for _, s := range res.Series {
+ labels := map[string]string{}
+ if query.label != "" {
+ labels[d.seriesBy] = query.label
+ } else if v, ok := s.Labels[d.seriesBy]; ok {
+ labels[d.seriesBy] = v
+ }
+ series = append(series, prom.Series{Labels: labels, DataPoints: s.DataPoints})
+ }
+ }
+ sort.SliceStable(series, func(i, j int) bool { return series[i].Labels[d.seriesBy] < series[j].Labels[d.seriesBy] })
+ if len(series) > cnpgHistoryMaxSeries {
+ c.Omitted = len(series) - cnpgHistoryMaxSeries
+ series = series[:cnpgHistoryMaxSeries]
+ }
+ covered := map[int64]bool{}
+ for _, s := range series {
+ for _, p := range s.DataPoints {
+ if !math.IsNaN(p.Value) && !math.IsInf(p.Value, 0) {
+ covered[p.Timestamp] = true
+ }
+ }
+ }
+ c.Series, c.Covered = series, len(covered)
+ if len(series) == 0 && d.presence != "" && end.After(start) {
+ res, err := q.Query(ctx, "count(last_over_time("+d.presence+"["+end.Sub(start).String()+"]))")
+ if err == nil && len(res.Series) > 0 {
+ c.State, c.Reason = CNPGHistoryStateEmpty, d.emptyReason
+ return
+ }
+ }
+ if len(series) == 0 {
+ c.State, c.Reason = CNPGHistoryStateNoSeries, "not scraped: Prometheus has no series for this cluster in this range ("+d.source+")"
+ if d.pvc {
+ c.Reason = "not scraped: Prometheus has no kubelet volume stats for this cluster's claims in this range (some volume drivers, such as hostPath, report none)"
+ }
+ return
+ }
+ c.State = CNPGHistoryStateOK
+}
+
+func truncate(s string, n int) string {
+ if len(s) <= n {
+ return s
+ }
+ return s[:n] + "…"
+}
+
+// CNPGLagReading is a Cluster's largest current standby replay lag.
+// Reporting counts the standbys whose lag was read, so a lag that covers only
+// some of them is not mistaken for all of them.
+type CNPGLagReading struct {
+ Seconds float64
+ Pod string
+ Reporting int
+}
+
+// CNPGFleetLag holds, per Cluster, the largest standby replay lag, and which
+// Clusters have exporter series at all, so "no standby reporting" is told
+// apart from "not scraped".
+type CNPGFleetLag struct {
+ Lag map[string]CNPGLagReading
+ Scraped map[string]bool
+ // Sustained is each Cluster's worst standby's lowest recorded lag over
+ // CNPGSustainedLagWindow, for standbys already reporting when the window
+ // began; absent otherwise.
+ Sustained map[string]CNPGLagReading
+ // Receivers is keyed by Cluster, for Clusters with an instance in
+ // recovery. Nil, with ReceiversError set, when the query failed.
+ Receivers map[string]CNPGReceiverReading
+ ReceiversError string
+ // ReceiversDownSustained is keyed by Cluster: standbys whose WAL receiver
+ // was down in every sample over CNPGReceiverDownWindow. Nil when that
+ // best-effort query failed.
+ ReceiversDownSustained map[string][]string
+}
+
+// CNPGReceiverDownWindow is how long a standby's WAL receiver must stay down
+// before it is a problem: a restarting standby is briefly down on every
+// rolling update and reconnects on its own.
+const CNPGReceiverDownWindow = 5 * time.Minute
+
+// CNPGReceiverReading is a Cluster's standbys (instances reporting
+// cnpg_pg_replication_in_recovery = 1) and their WAL receivers. Unknown is
+// set, and Receiving and Down are empty, when no standby reports
+// cnpg_pg_replication_is_wal_receiver_up. A standby that alone lacks the
+// family is counted in Standbys only.
+type CNPGReceiverReading struct {
+ Standbys int
+ Receiving int
+ Down []string
+ Unknown bool
+}
+
+// CNPGSustainedLagWindow is how long a standby's replay lag must stay high
+// before the fleet treats it as a problem rather than a transient spike.
+const CNPGSustainedLagWindow = 10 * time.Minute
+
+// QueryCNPGFleetLag reads the current replay lag of every standby of the
+// named Clusters in one namespace with one instant query.
+func (client *Client) QueryCNPGFleetLag(ctx context.Context, namespace string, clusters []string, matchers string) (CNPGFleetLag, error) {
+ if client == nil {
+ return CNPGFleetLag{}, errors.New("Prometheus client not initialized")
+ }
+ return queryCNPGFleetLag(ctx, client, namespace, clusters, matchers)
+}
+
+// A Pod scraped but not in recovery answers -1 through the `or` branch, so one
+// query carries both lag and presence.
+func queryCNPGFleetLag(ctx context.Context, q seriesQuerier, namespace string, clusters []string, matchers string) (CNPGFleetLag, error) {
+ sel := withScope(CNPGInstancesSelector(namespace, clusters), matchers)
+ query := "(max by (pod) (cnpg_pg_replication_lag{" + sel + "}) and on (pod) (max by (pod) (cnpg_pg_replication_in_recovery{" + sel + "}) == 1)) or (-1 * count by (pod) (cnpg_collector_up{" + sel + "}))"
+ res, err := q.Query(ctx, query)
+ if err != nil {
+ return CNPGFleetLag{}, err
+ }
+ known := map[string]bool{}
+ for _, c := range clusters {
+ known[c] = true
+ }
+ out := CNPGFleetLag{Lag: map[string]CNPGLagReading{}, Scraped: map[string]bool{}}
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ pod := s.Labels["pod"]
+ cluster := CNPGClusterOfPod(pod, known)
+ if cluster == "" {
+ continue
+ }
+ out.Scraped[cluster] = true
+ v := s.DataPoints[0].Value
+ if math.IsNaN(v) || math.IsInf(v, 0) || v < 0 {
+ continue
+ }
+ cur, ok := out.Lag[cluster]
+ reporting := cur.Reporting + 1
+ if !ok || v > cur.Seconds || (v == cur.Seconds && pod < cur.Pod) {
+ cur = CNPGLagReading{Seconds: v, Pod: pod}
+ }
+ cur.Reporting = reporting
+ out.Lag[cluster] = cur
+ }
+ out.Sustained = querySustainedCNPGLag(ctx, q, sel, known)
+ out.Receivers, out.ReceiversError = queryCNPGReceivers(ctx, q, sel, known)
+ out.ReceiversDownSustained = querySustainedReceiverDown(ctx, q, sel, known)
+ return out, nil
+}
+
+// querySustainedReceiverDown is best effort like the sustained lag: every raw
+// receiver sample in the window was 0, the instance was a standby in every
+// sample of the window (a primary reports no receiver, so a former primary
+// just after a switchover must not qualify on its primary samples), and the
+// series already existed when the window began.
+func querySustainedReceiverDown(ctx context.Context, q seriesQuerier, sel string, known map[string]bool) map[string][]string {
+ window := fmt.Sprintf("%dm", int(CNPGReceiverDownWindow.Minutes()))
+ query := "(max by (pod) (max_over_time(cnpg_pg_replication_is_wal_receiver_up{" + sel + "}[" + window + "])) == 0)" +
+ " and on (pod) (min by (pod) (min_over_time(cnpg_pg_replication_in_recovery{" + sel + "}[" + window + "])) == 1)" +
+ " and on (pod) (max by (pod) (cnpg_pg_replication_is_wal_receiver_up{" + sel + "} offset " + window + "))"
+ res, err := q.Query(ctx, query)
+ if err != nil {
+ return nil
+ }
+ out := map[string][]string{}
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ pod := s.Labels["pod"]
+ if cluster := CNPGClusterOfPod(pod, known); cluster != "" {
+ out[cluster] = append(out[cluster], pod)
+ }
+ }
+ for c := range out {
+ sort.Strings(out[c])
+ }
+ return out
+}
+
+// queryCNPGReceivers reads whether each standby's WAL receiver is up. Lag
+// cannot answer that: cnpg_pg_replication_lag is 0 whenever a standby's
+// receive and replay positions match, which is also what a standby that
+// receives nothing reports. A standby without the receiver family answers -1
+// through the `or` branch.
+func queryCNPGReceivers(ctx context.Context, q seriesQuerier, sel string, known map[string]bool) (map[string]CNPGReceiverReading, string) {
+ standby := "(max by (pod) (cnpg_pg_replication_in_recovery{" + sel + "}) == 1)"
+ query := "(max by (pod) (cnpg_pg_replication_is_wal_receiver_up{" + sel + "}) and on (pod) " + standby + ") or (-1 * " + standby + ")"
+ res, err := q.Query(ctx, query)
+ if err != nil {
+ return nil, "Prometheus query failed: " + truncate(err.Error(), 200)
+ }
+ return parseCNPGReceivers(res, known), ""
+}
+
+func parseCNPGReceivers(res *prom.QueryResult, known map[string]bool) map[string]CNPGReceiverReading {
+ out := map[string]CNPGReceiverReading{}
+ unreported := map[string]int{}
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ pod := s.Labels["pod"]
+ cluster := CNPGClusterOfPod(pod, known)
+ if cluster == "" {
+ continue
+ }
+ r := out[cluster]
+ r.Standbys++
+ switch v := s.DataPoints[0].Value; {
+ case v > 0:
+ r.Receiving++
+ case v == 0:
+ r.Down = append(r.Down, pod)
+ default:
+ unreported[cluster]++
+ }
+ out[cluster] = r
+ }
+ for cluster, r := range out {
+ sort.Strings(r.Down)
+ r.Unknown = unreported[cluster] == r.Standbys
+ out[cluster] = r
+ }
+ return out
+}
+
+// CNPGFleetSlotCap bounds the inactive slots kept per Cluster.
+const CNPGFleetSlotCap = 10
+
+const cnpgSlotRowLabel = "radar_row"
+
+// CNPGSlotReading is one inactive physical replication slot as one instance
+// reports it. Role is that instance's: "primary", "standby" (CloudNativePG
+// copies HA slots to standbys, where no WAL sender uses them), or "" when it
+// reports no recovery state. Bytes is the WAL the slot retains on that
+// instance, nil when not reported.
+type CNPGSlotReading struct {
+ Slot string
+ Pod string
+ Role string
+ Bytes *float64
+}
+
+// CNPGClusterSlots is one Cluster's inactive physical slots, most retained
+// WAL first, at most CNPGFleetSlotCap; Omitted counts the rest.
+type CNPGClusterSlots struct {
+ Inactive []CNPGSlotReading
+ Omitted int
+}
+
+// CNPGFleetSlots holds the Clusters whose instances report the slot family
+// at all, and the Clusters with any exporter series, so "no slot reported"
+// is told apart from "not scraped".
+type CNPGFleetSlots struct {
+ Clusters map[string]CNPGClusterSlots
+ Scraped map[string]bool
+}
+
+// QueryCNPGFleetSlots reads the inactive physical replication slots of the
+// named Clusters in one namespace with one instant query.
+func (client *Client) QueryCNPGFleetSlots(ctx context.Context, namespace string, clusters []string, matchers string) (CNPGFleetSlots, error) {
+ if client == nil {
+ return CNPGFleetSlots{}, errors.New("Prometheus client not initialized")
+ }
+ return queryCNPGFleetSlots(ctx, client, namespace, clusters, matchers)
+}
+
+// Each branch is tagged with its own radar_row value so `or` keeps all of
+// them: inactive slots, the WAL those retain, which instances report the slot
+// family at all, each instance's recovery state (its role), and which
+// instances are scraped.
+func queryCNPGFleetSlots(ctx context.Context, q seriesQuerier, namespace string, clusters []string, matchers string) (CNPGFleetSlots, error) {
+ sel := withScope(CNPGInstancesSelector(namespace, clusters), matchers)
+ physical := sel + `,slot_type="physical"`
+ inactive := "(max by (pod, slot_name) (cnpg_pg_replication_slots_active{" + physical + "}) == 0)"
+ tag := func(expr, row string) string {
+ return "label_replace(" + expr + `, "` + cnpgSlotRowLabel + `", "` + row + `", "pod", ".*")`
+ }
+ query := strings.Join([]string{
+ tag(inactive, "inactive"),
+ tag("max by (pod, slot_name) (cnpg_pg_replication_slots_pg_wal_lsn_diff{"+physical+"}) and on (pod, slot_name) "+inactive, "retained"),
+ tag("count by (pod) (cnpg_pg_replication_slots_active{"+sel+"})", "reported"),
+ tag("max by (pod) (cnpg_pg_replication_in_recovery{"+sel+"})", "recovery"),
+ tag("count by (pod) (cnpg_collector_up{"+sel+"})", "scraped"),
+ }, " or ")
+ res, err := q.Query(ctx, query)
+ if err != nil {
+ return CNPGFleetSlots{}, err
+ }
+ known := map[string]bool{}
+ for _, c := range clusters {
+ known[c] = true
+ }
+ return parseCNPGFleetSlots(res, known), nil
+}
+
+func parseCNPGFleetSlots(res *prom.QueryResult, known map[string]bool) CNPGFleetSlots {
+ type slotKey struct{ pod, slot string }
+ inactive := map[slotKey]bool{}
+ retained := map[slotKey]float64{}
+ role := map[string]string{}
+ out := CNPGFleetSlots{Clusters: map[string]CNPGClusterSlots{}, Scraped: map[string]bool{}}
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ pod := s.Labels["pod"]
+ cluster := CNPGClusterOfPod(pod, known)
+ if cluster == "" {
+ continue
+ }
+ v := s.DataPoints[0].Value
+ key := slotKey{pod, s.Labels["slot_name"]}
+ switch s.Labels[cnpgSlotRowLabel] {
+ case "inactive":
+ if key.slot != "" {
+ inactive[key] = true
+ }
+ case "retained":
+ if !math.IsNaN(v) && !math.IsInf(v, 0) {
+ retained[key] = v
+ }
+ case "reported":
+ out.Scraped[cluster] = true
+ if _, ok := out.Clusters[cluster]; !ok {
+ out.Clusters[cluster] = CNPGClusterSlots{Inactive: []CNPGSlotReading{}}
+ }
+ case "recovery":
+ switch v {
+ case 1:
+ role[pod] = "standby"
+ case 0:
+ role[pod] = "primary"
+ }
+ case "scraped":
+ out.Scraped[cluster] = true
+ }
+ }
+ for key := range inactive {
+ cluster := CNPGClusterOfPod(key.pod, known)
+ reading := CNPGSlotReading{Slot: key.slot, Pod: key.pod, Role: role[key.pod]}
+ if b, ok := retained[key]; ok {
+ reading.Bytes = &b
+ }
+ cs := out.Clusters[cluster]
+ cs.Inactive = append(cs.Inactive, reading)
+ out.Clusters[cluster] = cs
+ }
+ for cluster, cs := range out.Clusters {
+ sort.Slice(cs.Inactive, func(i, j int) bool {
+ a, b := cs.Inactive[i], cs.Inactive[j]
+ if (a.Bytes == nil) != (b.Bytes == nil) {
+ return a.Bytes != nil
+ }
+ if a.Bytes != nil && *a.Bytes != *b.Bytes {
+ return *a.Bytes > *b.Bytes
+ }
+ if a.Pod != b.Pod {
+ return a.Pod < b.Pod
+ }
+ return a.Slot < b.Slot
+ })
+ if len(cs.Inactive) > CNPGFleetSlotCap {
+ cs.Omitted = len(cs.Inactive) - CNPGFleetSlotCap
+ cs.Inactive = cs.Inactive[:CNPGFleetSlotCap]
+ }
+ out.Clusters[cluster] = cs
+ }
+ return out
+}
+
+// querySustainedCNPGLag is best effort: without it the fleet still shows the
+// current lag, it just raises no sustained-lag problem.
+func querySustainedCNPGLag(ctx context.Context, q seriesQuerier, sel string, known map[string]bool) map[string]CNPGLagReading {
+ // Exact on Prometheus 2.x and 3.x alike: every raw sample in the window is
+ // at least the reported floor, and the series already existed when the
+ // window began (it answers at offset 10m), so a standby that appeared a
+ // minute ago cannot qualify. min by (pod): when two jobs scrape the same
+ // Pod, a low sample in either counts. Scrape gaps are not filled in; nothing here
+ // claims a sample at every moment.
+ window := fmt.Sprintf("%dm", int(CNPGSustainedLagWindow.Minutes()))
+ query := "(min by (pod) (min_over_time(cnpg_pg_replication_lag{" + sel + "}[" + window + "]))" +
+ " and on (pod) (max by (pod) (cnpg_pg_replication_in_recovery{" + sel + "}) == 1)" +
+ " and on (pod) (max by (pod) (cnpg_pg_replication_lag{" + sel + "} offset " + window + ")))"
+ res, err := q.Query(ctx, query)
+ if err != nil {
+ return nil
+ }
+ out := map[string]CNPGLagReading{}
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ pod := s.Labels["pod"]
+ cluster := CNPGClusterOfPod(pod, known)
+ v := s.DataPoints[0].Value
+ if cluster == "" || math.IsNaN(v) || math.IsInf(v, 0) || v < 0 {
+ continue
+ }
+ if cur, ok := out[cluster]; !ok || v > cur.Seconds || (v == cur.Seconds && pod < cur.Pod) {
+ out[cluster] = CNPGLagReading{Seconds: v, Pod: pod}
+ }
+ }
+ return out
+}
+
+// QueryCNPGDiskGrowth reads each claim's used-bytes trend over the window as
+// bytes per hour (linear regression, so a single deletion does not dominate).
+func (client *Client) QueryCNPGDiskGrowth(ctx context.Context, namespace string, claims []string, window time.Duration, matchers string) (map[string]float64, error) {
+ if client == nil {
+ return nil, errors.New("Prometheus client not initialized")
+ }
+ return queryCNPGDiskGrowth(ctx, client, namespace, claims, window, matchers)
+}
+
+func queryCNPGDiskGrowth(ctx context.Context, q seriesQuerier, namespace string, claims []string, window time.Duration, matchers string) (map[string]float64, error) {
+ out := map[string]float64{}
+ if len(claims) == 0 {
+ return out, nil
+ }
+ sel := withScope(cnpgClaimSelector(namespace, claims), matchers)
+ metric := "kubelet_volume_stats_used_bytes{" + sel + "}"
+ // A fresh scrape must not certify a stopped series for the same claim.
+ res, err := q.Query(ctx, "3600 * max by (persistentvolumeclaim) (deriv("+metric+"["+window.String()+"]) and "+metric+")")
+ if err != nil {
+ return nil, err
+ }
+ for _, s := range res.Series {
+ if len(s.DataPoints) == 0 {
+ continue
+ }
+ v := s.DataPoints[0].Value
+ if math.IsNaN(v) || math.IsInf(v, 0) {
+ continue
+ }
+ out[s.Labels["persistentvolumeclaim"]] = v
+ }
+ return out, nil
+}
+
+// cnpgClaimSelector selects all the named claims in one matcher.
+func cnpgClaimSelector(namespace string, claims []string) string {
+ escaped := make([]string, len(claims))
+ for i, c := range claims {
+ escaped[i] = regexp.QuoteMeta(c)
+ }
+ sort.Strings(escaped)
+ return "namespace=" + strconv.Quote(namespace) + ",persistentvolumeclaim=~" + strconv.Quote(strings.Join(escaped, "|"))
+}
diff --git a/internal/prometheus/cnpg_history_test.go b/internal/prometheus/cnpg_history_test.go
new file mode 100644
index 0000000000..0db12a9645
--- /dev/null
+++ b/internal/prometheus/cnpg_history_test.go
@@ -0,0 +1,583 @@
+package prometheus
+
+import (
+ "context"
+ "errors"
+ "strconv"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+type fakeCNPGQuerier struct {
+ mu sync.Mutex
+ queries []string
+ instant func(q string) (*prom.QueryResult, error)
+ rng func(q string) (*prom.QueryResult, error)
+}
+
+func (f *fakeCNPGQuerier) Query(_ context.Context, q string) (*prom.QueryResult, error) {
+ f.mu.Lock()
+ f.queries = append(f.queries, q)
+ f.mu.Unlock()
+ if f.instant == nil {
+ return &prom.QueryResult{}, nil
+ }
+ return f.instant(q)
+}
+
+func (f *fakeCNPGQuerier) QueryRange(_ context.Context, q string, _, _ time.Time, _ time.Duration) (*prom.QueryResult, error) {
+ f.mu.Lock()
+ f.queries = append(f.queries, q)
+ f.mu.Unlock()
+ if f.rng == nil {
+ return &prom.QueryResult{}, nil
+ }
+ return f.rng(q)
+}
+
+func vec(labels map[string]string, v float64) prom.Series {
+ return prom.Series{Labels: labels, DataPoints: []prom.DataPoint{{Timestamp: 1, Value: v}}}
+}
+
+func TestCNPGInstanceSelectorAnchorsPodNames(t *testing.T) {
+ got := CNPGInstanceSelector("pg", "pg.main")
+ want := `namespace="pg",pod=~"^pg\\.main-[0-9]+$"`
+ if got != want {
+ t.Fatalf("selector = %s, want %s", got, want)
+ }
+ multi := CNPGInstancesSelector("pg", []string{"b", "a"})
+ if !strings.Contains(multi, `pod=~"^(a|b)-[0-9]+$"`) {
+ t.Fatalf("multi selector = %s", multi)
+ }
+ known := map[string]bool{"pg": true, "pg-runtime": true}
+ for pod, want := range map[string]string{"pg-1": "pg", "pg-runtime-12": "pg-runtime", "pg-runtime-x": "", "other-1": ""} {
+ if got := CNPGClusterOfPod(pod, known); got != want {
+ t.Errorf("CNPGClusterOfPod(%s) = %q, want %q", pod, got, want)
+ }
+ }
+}
+
+func TestParseCNPGHistoryRangeBoundsPoints(t *testing.T) {
+ for _, name := range []string{"15m", "1h", "6h", "24h", ""} {
+ r, ok := ParseCNPGHistoryRange(name)
+ if !ok {
+ t.Fatalf("range %q rejected", name)
+ }
+ if points := int(r.Duration / r.Step); points > 150 || points < 50 {
+ t.Errorf("range %q has %d points", name, points)
+ }
+ }
+ if _, ok := ParseCNPGHistoryRange("7d"); ok {
+ t.Fatal("7d accepted")
+ }
+}
+
+func cnpgProbe(window time.Duration) scopeProbe {
+ return scopeProbe{metric: "cnpg_collector_up", key: "pod", selectors: []string{`namespace="pg"`}, window: window}
+}
+
+func TestDecideScopeRefusesAmbiguousIdentity(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(string) (*prom.QueryResult, error) {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 2)}}, nil
+ }}
+ if _, _, err := decideScope(context.Background(), q, cnpgProbe(0), nil); !errors.Is(err, ErrScopeAmbiguous) {
+ t.Fatalf("err = %v, want ambiguous", err)
+ }
+ q.instant = func(string) (*prom.QueryResult, error) {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 1)}}, nil
+ }
+ m, iso, err := decideScope(context.Background(), q, cnpgProbe(0), nil)
+ if err != nil || m != "" || iso.Mode != SeriesIsolationUnverified {
+ t.Fatalf("single identity: m=%q iso=%+v err=%v", m, iso, err)
+ }
+}
+
+// An identity that stopped reporting minutes ago still fills a one-hour
+// chart, so the check must span the chart's range, not the present instant.
+func TestDecideScopeChecksIdentitiesOverTheWholeRange(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ v := 1.0
+ if strings.Contains(query, "count_over_time(cnpg_collector_up{") && strings.Contains(query, "[1h0m0s]") {
+ v = 2
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, v)}}, nil
+ }}
+ if _, _, err := decideScope(context.Background(), q, cnpgProbe(time.Hour), nil); !errors.Is(err, ErrScopeAmbiguous) {
+ t.Fatalf("err = %v, want ambiguous over the range; queries %v", err, q.queries)
+ }
+}
+
+// The CNPG exporter labels its own series cluster=; that
+// is not a Kubernetes identity and must never be imposed on other families.
+func TestDecideScopeNeverPinsExporterLabels(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.HasPrefix(query, "max(count by (pod)") {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 1)}}, nil
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"cluster": "pg"}, 1)}}, nil
+ }}
+ m, _, err := decideScope(context.Background(), q, scopeProbe{metric: "cnpg_collector_up", key: "pod", selectors: []string{CNPGInstanceSelector("db", "pg")}, window: time.Hour}, nil)
+ if err != nil {
+ t.Fatal(err)
+ }
+ for _, def := range cnpgHistoryDefs(withScope(CNPGInstanceSelector("db", "pg"), m), "", 30*time.Second) {
+ if def.id == "replicationLag" && strings.Contains(def.queries[0].expr, `cluster="pg"`) {
+ t.Fatalf("exporter's cluster label imposed on replication metrics: %s", def.queries[0].expr)
+ }
+ }
+}
+
+// Two database clusters in one namespace carry different exporter cluster
+// labels; each Pod still has one identity, so fleet lag is not ambiguous.
+func TestDecideScopeTwoDatabasesInOneKubernetesCluster(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.HasPrefix(query, "max(count by (pod)") {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 1)}}, nil
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"cluster": "pg-a"}, 1), vec(map[string]string{"cluster": "pg-b"}, 1)}}, nil
+ }}
+ if _, _, err := decideScope(context.Background(), q, scopeProbe{metric: "cnpg_collector_up", key: "pod", selectors: []string{CNPGInstancesSelector("db", []string{"pg-a", "pg-b"})}, window: 10 * time.Minute}, nil); err != nil {
+ t.Fatalf("two CNPG clusters in one Kubernetes cluster rejected: %v", err)
+ }
+}
+
+func TestDecideScopeVerifiedLabelsMustReachProbedSeries(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.Contains(query, `k8s_cluster_name="east"`) {
+ return &prom.QueryResult{}, nil
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 3)}}, nil
+ }}
+ if _, _, err := decideScope(context.Background(), q, cnpgProbe(0), map[string]string{"k8s_cluster_name": "east"}); !errors.Is(err, ErrScopeMismatch) {
+ t.Fatalf("err = %v, want mismatch", err)
+ }
+ q.instant = func(string) (*prom.QueryResult, error) {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 3)}}, nil
+ }
+ m, iso, err := decideScope(context.Background(), q, cnpgProbe(0), map[string]string{"k8s_cluster_name": "east"})
+ if err != nil || m != `k8s_cluster_name="east"` || iso.Mode != SeriesIsolationVerified {
+ t.Fatalf("verified: m=%q iso=%+v err=%v", m, iso, err)
+ }
+}
+
+func TestQueryCNPGHistoryKeepsPVCScopeAndAmbiguity(t *testing.T) {
+ q := &fakeCNPGQuerier{}
+ req := CNPGHistoryRequest{Namespace: "pg", Cluster: "pg", Range: cnpgHistoryRanges[1], End: time.Unix(1700000000, 0), Claims: []string{"pg-1"}, PVCMatchers: `cluster="east"`}
+ queryCNPGHistory(context.Background(), q, req)
+ found := false
+ for _, query := range q.queries {
+ if strings.Contains(query, "kubelet_volume_stats_used_bytes") {
+ found = true
+ if !strings.Contains(query, `cluster="east"`) {
+ t.Errorf("volume chart query lost the claims' scope: %s", query)
+ }
+ }
+ }
+ if !found {
+ t.Fatal("no volume chart query ran")
+ }
+ req.PVCScopeState, req.PVCReason = CNPGHistoryStateAmbiguous, "two identities"
+ for _, c := range queryCNPGHistory(context.Background(), &fakeCNPGQuerier{}, req) {
+ if c.ID == "pvcUsed" && (c.State != CNPGHistoryStateAmbiguous || len(c.Series) != 0) {
+ t.Errorf("ambiguous claims chart = %+v", c)
+ }
+ }
+}
+
+func TestQueryCNPGHistoryStatesPerChart(t *testing.T) {
+ r, _ := ParseCNPGHistoryRange("15m")
+ q := &fakeCNPGQuerier{
+ rng: func(query string) (*prom.QueryResult, error) {
+ switch {
+ case strings.Contains(query, "cnpg_pg_replication_lag"):
+ return &prom.QueryResult{Series: []prom.Series{
+ {Labels: map[string]string{"pod": "pg-2", "instance": "x"}, DataPoints: []prom.DataPoint{{Timestamp: 100, Value: 1}, {Timestamp: 115, Value: 2}}},
+ }}, nil
+ case strings.Contains(query, "cnpg_pg_stat_database_deadlocks"):
+ return nil, errors.New("boom")
+ }
+ return &prom.QueryResult{}, nil
+ },
+ instant: func(query string) (*prom.QueryResult, error) {
+ if strings.Contains(query, "checkpoints") {
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 2)}}, nil
+ }
+ return &prom.QueryResult{}, nil
+ },
+ }
+ charts := queryCNPGHistory(context.Background(), q, CNPGHistoryRequest{
+ Namespace: "pg", Cluster: "pg", Range: r, End: time.Unix(1_700_000_000, 0), Matchers: `cluster_id="a"`,
+ PVCDenied: &auth.Grant{Verb: "get", Resource: "persistentvolumeclaims", Namespace: "pg"},
+ })
+ by := map[string]CNPGHistoryChart{}
+ for _, c := range charts {
+ by[c.ID] = c
+ }
+ lag := by["replicationLag"]
+ if lag.State != CNPGHistoryStateOK || len(lag.Series) != 1 || lag.Series[0].Labels["pod"] != "pg-2" || lag.Series[0].Labels["instance"] != "" || lag.Covered != 2 || lag.Steps != 61 {
+ t.Fatalf("lag chart = %+v", lag)
+ }
+ if by["deadlocks"].State != CNPGHistoryStateError {
+ t.Errorf("deadlocks state = %s", by["deadlocks"].State)
+ }
+ if by["pvcUsed"].State != CNPGHistoryStateDenied || by["pvcUsed"].Grant == nil || *by["pvcUsed"].Grant != (auth.Grant{Verb: "get", Resource: "persistentvolumeclaims", Namespace: "pg"}) {
+ t.Errorf("pvc chart = %+v", by["pvcUsed"])
+ }
+ if by["walSize"].State != CNPGHistoryStateNoSeries || !strings.HasPrefix(by["walSize"].Reason, "not scraped") {
+ t.Errorf("walSize chart = %+v", by["walSize"])
+ }
+ if by["checkpoints"].State != CNPGHistoryStateEmpty {
+ t.Errorf("checkpoints chart = %+v", by["checkpoints"])
+ }
+ for _, query := range q.queries {
+ if strings.Contains(query, "cnpg_") && strings.Contains(query, "{namespace=") && !strings.Contains(query, `cluster_id="a"`) {
+ t.Errorf("query without cluster identity: %s", query)
+ }
+ if strings.Contains(query, "kubelet_volume_stats") {
+ t.Errorf("denied PVC chart still queried: %s", query)
+ }
+ }
+}
+
+func TestQueryCNPGHistoryCapsSeries(t *testing.T) {
+ r, _ := ParseCNPGHistoryRange("1h")
+ q := &fakeCNPGQuerier{rng: func(query string) (*prom.QueryResult, error) {
+ if !strings.Contains(query, "cnpg_pg_database_size_bytes") {
+ return &prom.QueryResult{}, nil
+ }
+ var out []prom.Series
+ for i := 0; i < 20; i++ {
+ out = append(out, prom.Series{Labels: map[string]string{"datname": string(rune('a' + i))}, DataPoints: []prom.DataPoint{{Timestamp: 1, Value: 1}}})
+ }
+ return &prom.QueryResult{Series: out}, nil
+ }}
+ charts := queryCNPGHistory(context.Background(), q, CNPGHistoryRequest{Namespace: "pg", Cluster: "pg", Range: r, End: time.Now()})
+ for _, c := range charts {
+ if c.ID == "databaseSize" && (len(c.Series) != cnpgHistoryMaxSeries || c.Omitted != 20-cnpgHistoryMaxSeries) {
+ t.Fatalf("databaseSize kept %d, omitted %d", len(c.Series), c.Omitted)
+ }
+ }
+}
+
+func TestQueryCNPGFleetLagSeparatesNoStandbyFromUnscraped(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(string) (*prom.QueryResult, error) {
+ return &prom.QueryResult{Series: []prom.Series{
+ vec(map[string]string{"pod": "a-1"}, -1),
+ vec(map[string]string{"pod": "a-2"}, 3.5),
+ vec(map[string]string{"pod": "a-3"}, 1),
+ vec(map[string]string{"pod": "b-1"}, -1),
+ vec(map[string]string{"pod": "zzz-1"}, 9),
+ }}, nil
+ }}
+ got, err := queryCNPGFleetLag(context.Background(), q, "pg", []string{"a", "b", "c"}, "")
+ if err != nil {
+ t.Fatal(err)
+ }
+ // a-1 answers -1 (scraped, not in recovery); the lag covers a-2 and a-3.
+ if got.Lag["a"].Seconds != 3.5 || got.Lag["a"].Pod != "a-2" || got.Lag["a"].Reporting != 2 {
+ t.Errorf("a lag = %+v", got.Lag["a"])
+ }
+ if _, ok := got.Lag["b"]; ok || !got.Scraped["b"] {
+ t.Errorf("b: lag %v scraped %v", got.Lag["b"], got.Scraped["b"])
+ }
+ if got.Scraped["c"] || got.Scraped["zzz"] {
+ t.Errorf("unexpected scraped: %v", got.Scraped)
+ }
+}
+
+// A standby whose WAL receiver is down reads lag 0 (nothing received, nothing
+// left to replay), so the receiver is read on its own and never inferred.
+func TestQueryCNPGFleetLagReadsWALReceivers(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.Contains(query, "max_over_time(cnpg_pg_replication_is_wal_receiver_up") {
+ return &prom.QueryResult{}, nil
+ }
+ if strings.Contains(query, "cnpg_pg_replication_is_wal_receiver_up") {
+ for _, want := range []string{"max by (pod) (cnpg_pg_replication_is_wal_receiver_up{", "(max by (pod) (cnpg_pg_replication_in_recovery{", ") == 1)", `cluster_id="a"`} {
+ if !strings.Contains(query, want) {
+ t.Errorf("receiver query lacks %q: %s", want, query)
+ }
+ }
+ return &prom.QueryResult{Series: []prom.Series{
+ vec(map[string]string{"pod": "a-3"}, 0),
+ vec(map[string]string{"pod": "a-2"}, 1),
+ vec(map[string]string{"pod": "a-1"}, 0),
+ vec(map[string]string{"pod": "b-2"}, -1),
+ vec(map[string]string{"pod": "b-3"}, -1),
+ vec(map[string]string{"pod": "c-2"}, 1),
+ vec(map[string]string{"pod": "c-3"}, -1),
+ vec(map[string]string{"pod": "zzz-1"}, 0),
+ }}, nil
+ }
+ return &prom.QueryResult{Series: []prom.Series{
+ vec(map[string]string{"pod": "a-1"}, 0),
+ vec(map[string]string{"pod": "a-2"}, 0),
+ vec(map[string]string{"pod": "a-3"}, 2),
+ vec(map[string]string{"pod": "d-1"}, -1),
+ }}, nil
+ }}
+ got, err := queryCNPGFleetLag(context.Background(), q, "pg", []string{"a", "b", "c", "d"}, `cluster_id="a"`)
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got.Lag["a"].Seconds != 2 || got.Lag["a"].Pod != "a-3" {
+ t.Errorf("lag reading changed: %+v", got.Lag["a"])
+ }
+ a := got.Receivers["a"]
+ if a.Standbys != 3 || a.Receiving != 1 || a.Unknown || strings.Join(a.Down, ",") != "a-1,a-3" {
+ t.Errorf("a receivers = %+v, want 3 standbys, 1 receiving, a-1 and a-3 down (sorted)", a)
+ }
+ if b := got.Receivers["b"]; b.Standbys != 2 || !b.Unknown || b.Receiving != 0 || len(b.Down) != 0 {
+ t.Errorf("b receivers = %+v, want unknown: no standby reports the receiver family", b)
+ }
+ if c := got.Receivers["c"]; c.Standbys != 2 || c.Unknown || c.Receiving != 1 || len(c.Down) != 0 {
+ t.Errorf("c receivers = %+v, want one receiving and one unreported, not unknown", c)
+ }
+ if _, ok := got.Receivers["d"]; ok {
+ t.Errorf("d has no standby but got receivers %+v", got.Receivers["d"])
+ }
+ if _, ok := got.Receivers["zzz"]; ok || got.ReceiversError != "" {
+ t.Errorf("receivers = %+v err %q", got.Receivers, got.ReceiversError)
+ }
+
+ failing := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.Contains(query, "is_wal_receiver_up") {
+ return nil, errors.New("boom")
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"pod": "a-2"}, 1)}}, nil
+ }}
+ got, err = queryCNPGFleetLag(context.Background(), failing, "pg", []string{"a"}, "")
+ if err != nil || got.Lag["a"].Seconds != 1 {
+ t.Fatalf("a failed receiver query must not fail the lag: %+v %v", got, err)
+ }
+ if got.Receivers != nil || !strings.Contains(got.ReceiversError, "boom") {
+ t.Errorf("receivers = %+v err %q, want nil with the error", got.Receivers, got.ReceiversError)
+ }
+}
+
+func slotRow(row string, labels map[string]string, v float64) prom.Series {
+ l := map[string]string{cnpgSlotRowLabel: row}
+ for k, val := range labels {
+ l[k] = val
+ }
+ return vec(l, v)
+}
+
+func TestQueryCNPGFleetSlotsListsInactivePhysicalSlots(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ for _, want := range []string{
+ `cnpg_pg_replication_slots_active{namespace="pg",pod=~"^(a|b|c|d)-[0-9]+$",cluster_id="x",slot_type="physical"}) == 0)`,
+ `cnpg_pg_replication_slots_pg_wal_lsn_diff{namespace="pg",pod=~"^(a|b|c|d)-[0-9]+$",cluster_id="x",slot_type="physical"}`,
+ `"radar_row", "retained", "pod", ".*")`,
+ } {
+ if !strings.Contains(query, want) {
+ t.Errorf("slots query lacks %q: %s", want, query)
+ }
+ }
+ return &prom.QueryResult{Series: []prom.Series{
+ // a: the primary a-6 keeps an inactive slot for a-1, and a-2 reports
+ // its synchronized copy; a-1 itself reports no recovery state.
+ slotRow("inactive", map[string]string{"pod": "a-6", "slot_name": "_cnpg_a_1"}, 0),
+ slotRow("retained", map[string]string{"pod": "a-6", "slot_name": "_cnpg_a_1"}, 4.9e9),
+ slotRow("inactive", map[string]string{"pod": "a-2", "slot_name": "_cnpg_a_1"}, 0),
+ slotRow("retained", map[string]string{"pod": "a-2", "slot_name": "_cnpg_a_1"}, 4.8e9),
+ slotRow("inactive", map[string]string{"pod": "a-1", "slot_name": "_cnpg_a_2"}, 0),
+ slotRow("reported", map[string]string{"pod": "a-6"}, 2),
+ slotRow("reported", map[string]string{"pod": "a-2"}, 1),
+ slotRow("recovery", map[string]string{"pod": "a-6"}, 0),
+ slotRow("recovery", map[string]string{"pod": "a-2"}, 1),
+ slotRow("scraped", map[string]string{"pod": "a-6"}, 1),
+ // b: slots reported, all active (the query filters active ones out).
+ slotRow("reported", map[string]string{"pod": "b-1"}, 1),
+ slotRow("scraped", map[string]string{"pod": "b-1"}, 1),
+ // c: scraped, no slot family; d: not scraped at all.
+ slotRow("scraped", map[string]string{"pod": "c-1"}, 1),
+ slotRow("inactive", map[string]string{"pod": "zzz-1", "slot_name": "s"}, 0),
+ }}, nil
+ }}
+ got, err := queryCNPGFleetSlots(context.Background(), q, "pg", []string{"a", "b", "c", "d"}, `cluster_id="x"`)
+ if err != nil {
+ t.Fatal(err)
+ }
+ a := got.Clusters["a"]
+ if len(a.Inactive) != 3 || a.Omitted != 0 {
+ t.Fatalf("a inactive = %+v", a)
+ }
+ want := []struct {
+ pod, slot, role string
+ bytes float64
+ }{{"a-6", "_cnpg_a_1", "primary", 4.9e9}, {"a-2", "_cnpg_a_1", "standby", 4.8e9}}
+ for i, w := range want {
+ s := a.Inactive[i]
+ if s.Pod != w.pod || s.Slot != w.slot || s.Role != w.role || s.Bytes == nil || *s.Bytes != w.bytes {
+ t.Errorf("a inactive[%d] = %+v, want %+v", i, s, w)
+ }
+ }
+ if s := a.Inactive[2]; s.Pod != "a-1" || s.Role != "" || s.Bytes != nil {
+ t.Errorf("a slot without retained bytes or recovery state = %+v, want last with no role and no bytes", s)
+ }
+ if b, ok := got.Clusters["b"]; !ok || b.Inactive == nil || len(b.Inactive) != 0 {
+ t.Errorf("b = %+v (present %v), want read with an empty list", b, ok)
+ }
+ if _, ok := got.Clusters["c"]; ok || !got.Scraped["c"] {
+ t.Errorf("c: clusters %+v scraped %v, want scraped without a slot reading", got.Clusters["c"], got.Scraped["c"])
+ }
+ if _, ok := got.Clusters["d"]; ok || got.Scraped["d"] {
+ t.Errorf("d must be neither read nor scraped")
+ }
+ if _, ok := got.Clusters["zzz"]; ok {
+ t.Error("a Pod of an unrequested cluster was attributed")
+ }
+}
+
+func TestQueryCNPGFleetSlotsCapsBySize(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(string) (*prom.QueryResult, error) {
+ var out []prom.Series
+ for i := 1; i <= 13; i++ {
+ pod := "a-" + strconv.Itoa(i)
+ out = append(out,
+ slotRow("inactive", map[string]string{"pod": pod, "slot_name": "_cnpg_a_99"}, 0),
+ slotRow("retained", map[string]string{"pod": pod, "slot_name": "_cnpg_a_99"}, float64(i)),
+ )
+ }
+ return &prom.QueryResult{Series: append(out, slotRow("reported", map[string]string{"pod": "a-1"}, 1))}, nil
+ }}
+ got, err := queryCNPGFleetSlots(context.Background(), q, "pg", []string{"a"}, "")
+ if err != nil {
+ t.Fatal(err)
+ }
+ a := got.Clusters["a"]
+ if len(a.Inactive) != CNPGFleetSlotCap || a.Omitted != 13-CNPGFleetSlotCap {
+ t.Fatalf("kept %d, omitted %d", len(a.Inactive), a.Omitted)
+ }
+ if *a.Inactive[0].Bytes != 13 || *a.Inactive[CNPGFleetSlotCap-1].Bytes != 4 {
+ t.Errorf("not the largest first: first %v last %v", *a.Inactive[0].Bytes, *a.Inactive[CNPGFleetSlotCap-1].Bytes)
+ }
+}
+
+func TestQueryCNPGFleetLagReportsSustainedLagSeparately(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if strings.Contains(query, "cnpg_pg_replication_is_wal_receiver_up") {
+ return &prom.QueryResult{}, nil
+ }
+ if strings.Contains(query, "min_over_time") {
+ // Every raw sample in the window, and a series that already
+ // existed when the window began: a sample count cannot prove age.
+ for _, want := range []string{"min by (pod) (min_over_time(cnpg_pg_replication_lag{", "[10m])", "} offset 10m"} {
+ if !strings.Contains(query, want) {
+ t.Errorf("sustained query lacks %q: %s", want, query)
+ }
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"pod": "a-2"}, 42)}}, nil
+ }
+ return &prom.QueryResult{Series: []prom.Series{
+ vec(map[string]string{"pod": "a-2"}, 60),
+ vec(map[string]string{"pod": "b-2"}, 90),
+ }}, nil
+ }}
+ got, err := queryCNPGFleetLag(context.Background(), q, "pg", []string{"a", "b"}, "")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if got.Sustained["a"].Seconds != 42 || got.Sustained["a"].Pod != "a-2" {
+ t.Errorf("a sustained = %+v", got.Sustained["a"])
+ }
+ if _, ok := got.Sustained["b"]; ok {
+ t.Error("b spiked to 90 s now but has no sustained reading; it must not get one")
+ }
+}
+
+func TestQueryCNPGFleetLagSustainedReceiverLossNeedsAStandbyThroughout(t *testing.T) {
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ if !strings.Contains(query, "max_over_time(cnpg_pg_replication_is_wal_receiver_up") {
+ return &prom.QueryResult{}, nil
+ }
+ // Down in every sample, a standby in every sample (a primary reports no
+ // receiver, so a former primary must not qualify on its primary
+ // samples), and reporting since before the window.
+ for _, want := range []string{
+ "max by (pod) (max_over_time(cnpg_pg_replication_is_wal_receiver_up{",
+ "[5m])) == 0)",
+ "min by (pod) (min_over_time(cnpg_pg_replication_in_recovery{",
+ "[5m])) == 1)",
+ "} offset 5m))",
+ } {
+ if !strings.Contains(query, want) {
+ t.Errorf("sustained receiver query lacks %q: %s", want, query)
+ }
+ }
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"pod": "a-3"}, 0), vec(map[string]string{"pod": "a-2"}, 0), vec(map[string]string{"pod": "zzz-1"}, 0)}}, nil
+ }}
+ got, err := queryCNPGFleetLag(context.Background(), q, "pg", []string{"a", "b"}, "")
+ if err != nil {
+ t.Fatal(err)
+ }
+ if strings.Join(got.ReceiversDownSustained["a"], ",") != "a-2,a-3" {
+ t.Errorf("a sustained down = %v, want a-2,a-3 (sorted)", got.ReceiversDownSustained["a"])
+ }
+ if _, ok := got.ReceiversDownSustained["b"]; ok {
+ t.Error("b has no sustained receiver loss")
+ }
+}
+
+func TestCNPGHistoryBoundsExcludePredecessorLookback(t *testing.T) {
+ rng, _ := ParseCNPGHistoryRange("1h")
+ end := time.Date(2026, 10, 7, 12, 0, 0, 0, time.UTC)
+ for _, tc := range []struct {
+ name string
+ created time.Time
+ want time.Time
+ }{
+ {"old Cluster", end.Add(-2 * time.Hour), end.Add(-time.Hour)},
+ {"recreated between steps", end.Add(-20*time.Minute + time.Second), end.Add(-15*time.Minute + rng.Step)},
+ {"fresh Cluster", end.Add(-time.Minute), end.Add(4 * time.Minute)},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ start, stop := rng.Bounds(end, tc.created)
+ if !start.Equal(tc.want) || !stop.Equal(end) {
+ t.Fatalf("bounds = %v, %v; want %v, %v", start, stop, tc.want, end)
+ }
+ })
+ }
+ q := &fakeCNPGQuerier{}
+ charts := queryCNPGHistory(context.Background(), q, CNPGHistoryRequest{Namespace: "pg", Cluster: "pg", Range: rng, End: end, CreatedAt: end.Add(-time.Minute)})
+ if len(q.queries) != 0 {
+ t.Fatal("queried an invalid fresh-Cluster range")
+ }
+ for _, chart := range charts {
+ if chart.State != CNPGHistoryStateEmpty && chart.State != CNPGHistoryStateNotRead {
+ t.Fatalf("fresh chart = %+v", chart)
+ }
+ }
+}
+
+func TestCNPGVolumeHistoryRetainsScopeFailureState(t *testing.T) {
+ for _, state := range []string{CNPGHistoryStateAmbiguous, "scopeMismatch", CNPGHistoryStateError} {
+ t.Run(state, func(t *testing.T) {
+ q := &fakeCNPGQuerier{}
+ req := CNPGHistoryRequest{Namespace: "pg", Cluster: "pg", Range: cnpgHistoryRanges[1], End: time.Unix(1700000000, 0), Claims: []string{"pg-1"}, PVCScopeState: state, PVCReason: "scope check failed"}
+ found := false
+ for _, chart := range queryCNPGHistory(context.Background(), q, req) {
+ if chart.ID == "pvcUsed" || chart.ID == "pvcFree" {
+ found = true
+ if chart.State != state || chart.Reason != req.PVCReason || len(chart.Series) != 0 {
+ t.Fatalf("chart: %+v", chart)
+ }
+ }
+ }
+ if !found {
+ t.Fatal("no PVC chart")
+ }
+ for _, query := range q.queries {
+ if strings.Contains(query, "kubelet_volume_stats_") {
+ t.Fatalf("unsafe query after failed scope check: %s", query)
+ }
+ }
+ })
+ }
+}
diff --git a/internal/prometheus/cnpg_promql_test.go b/internal/prometheus/cnpg_promql_test.go
new file mode 100644
index 0000000000..75ddce783a
--- /dev/null
+++ b/internal/prometheus/cnpg_promql_test.go
@@ -0,0 +1,42 @@
+package prometheus
+
+import (
+ "context"
+ "os"
+ "testing"
+ "time"
+
+ "github.com/skyhook-io/radar/pkg/prom/promtest"
+)
+
+func TestCNPGDiskGrowthPromQL(t *testing.T) {
+ image := os.Getenv("RADAR_TEST_PROMTOOL_IMAGE")
+ if image == "" {
+ t.Skip("set RADAR_TEST_PROMTOOL_IMAGE=prom/prometheus:v3.5.0 to evaluate queries with Docker")
+ }
+ q := &fakeCNPGQuerier{}
+ if _, err := queryCNPGDiskGrowth(context.Background(), q, "db", []string{"fresh", "stopped", "zero", "negative"}, 6*time.Hour, `cluster_id="current"`); err != nil {
+ t.Fatal(err)
+ }
+ if len(q.queries) != 1 {
+ t.Fatalf("queries = %d, want one namespace query", len(q.queries))
+ }
+ input := []map[string]string{
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="fresh",cluster_id="current",job="current"}`, "values": "0+100x360"},
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="stopped",cluster_id="current"}`, "values": "0+200x120"},
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="zero",cluster_id="current"}`, "values": "0x360"},
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="negative",cluster_id="current"}`, "values": "100000-100x360"},
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="fresh",cluster_id="other"}`, "values": "0+1000x360"},
+ {"series": `kubelet_volume_stats_used_bytes{namespace="db",persistentvolumeclaim="fresh",cluster_id="current",job="stopped"}`, "values": "0+400x120"},
+ }
+ tests := []map[string]any{{
+ "expr": q.queries[0],
+ "eval_time": "6h",
+ "exp_samples": []map[string]any{
+ {"labels": `{persistentvolumeclaim="fresh"}`, "value": 6000},
+ {"labels": `{persistentvolumeclaim="zero"}`, "value": 0},
+ {"labels": `{persistentvolumeclaim="negative"}`, "value": -6000},
+ },
+ }}
+ promtest.Run(t, image, input, tests)
+}
diff --git a/internal/prometheus/discovery.go b/internal/prometheus/discovery.go
index f59f70c69e..858432a6b6 100644
--- a/internal/prometheus/discovery.go
+++ b/internal/prometheus/discovery.go
@@ -5,6 +5,8 @@ import (
"errors"
"fmt"
"log"
+ "regexp"
+ "slices"
"strings"
"sync"
"sync/atomic"
@@ -19,11 +21,29 @@ import (
var ErrPrometheusNotFound = errors.New("no Prometheus service found in cluster")
// errPrometheusUnreachable is ErrPrometheusNotFound for the case where
-// discovery enumerated a candidate and could not reach it. It wraps the
-// sentinel with an identical message so every errors.Is caller and every
-// message a user sees stay the same, while Availability can tell an
-// unreachable installation from a cluster that has none.
-var errPrometheusUnreachable = fmt.Errorf("%w", ErrPrometheusNotFound)
+// discovery enumerated a candidate and could not reach it, so every errors.Is
+// caller keeps its answer while Availability can tell an unreachable
+// installation from a cluster that has none. Its message must not claim
+// nothing was found.
+var errPrometheusUnreachable error = &prometheusUnreachableError{
+ msg: "No working Prometheus endpoint found.\nCandidate services could not be reached",
+ cause: ErrPrometheusNotFound,
+}
+
+// prometheusUnreachableError carries a complete sentence for an unreachable
+// installation. Unlike fmt's %w it does not append the wrapped sentinel's
+// text, which would contradict the sentence.
+type prometheusUnreachableError struct {
+ msg string
+ cause error
+}
+
+func (e *prometheusUnreachableError) Error() string { return e.msg }
+func (e *prometheusUnreachableError) Unwrap() error { return e.cause }
+
+func unreachableBecause(format string, args ...any) error {
+ return &prometheusUnreachableError{msg: fmt.Sprintf(format, args...), cause: errPrometheusUnreachable}
+}
// errDiscoverySuperseded is returned when a configuration change (Reset /
// SetManualURL / SetHeaders) invalidated a discovery mid-flight. The result is
@@ -220,6 +240,7 @@ func (c *Client) discover(ctx context.Context, gen uint64) (string, string, erro
// machine). Serial by necessity — port-forwarding mutates the owner's shared
// forward state.
var lastErr error
+ var denied []string
for _, cand := range candidates {
// Bail promptly if the run was superseded mid-fallback (Reset / context
// switch) rather than churning the rest of the list while holding the gate.
@@ -275,7 +296,10 @@ func (c *Client) discover(ctx context.Context, gen uint64) (string, string, erro
logDiscoveryEnded(start, err)
return "", "", err
}
- lastErr = fmt.Errorf("port-forward to %s/%s failed: %w", cand.Namespace, cand.Name, pfErr)
+ lastErr = fmt.Errorf("No working Prometheus endpoint found.\nCandidate %s/%s: port-forward failed: %w", cand.Namespace, cand.Name, pfErr)
+ if strings.Contains(strings.ToLower(pfErr.Error()), "forbidden") {
+ denied = append(denied, deniedGrant(pfErr))
+ }
if !discoveryDiagnosticsSuppressed(ctx) {
errorlog.Record("prometheus", "error", "port-forward to %s/%s failed: %v", cand.Namespace, cand.Name, pfErr)
}
@@ -305,9 +329,9 @@ func (c *Client) discover(ctx context.Context, gen uint64) (string, string, erro
logDiscoveryEnded(start, err)
return "", "", err
}
- lastErr = fmt.Errorf("Prometheus at %s/%s not responding after port-forward", cand.Namespace, cand.Name)
+ lastErr = prometheusCandidateProbeError(cand.Namespace, cand.Name)
if !discoveryDiagnosticsSuppressed(ctx) {
- errorlog.Record("prometheus", "error", "Prometheus at %s/%s not responding after port-forward", cand.Namespace, cand.Name)
+ errorlog.Record("prometheus", "error", "Candidate %s/%s did not respond after port-forward", cand.Namespace, cand.Name)
}
}
@@ -319,10 +343,46 @@ func (c *Client) discover(ctx context.Context, gen uint64) (string, string, erro
return "", "", err
}
log.Printf("[prometheus] discovery failed after %s: no reachable Prometheus among %d candidate(s)", took(start), len(candidates))
- if lastErr != nil {
- return "", "", lastErr
+ // With several candidates the last one's failure describes whichever
+ // Service happened to rank last, often not a Prometheus at all.
+ switch {
+ case lastErr == nil:
+ return "", "", errPrometheusUnreachable
+ case len(candidates) == 1 && len(denied) == 1:
+ return "", "", unreachableBecause("No working Prometheus endpoint found.\nCandidate %s/%s: port-forward denied (%s)", candidates[0].Namespace, candidates[0].Name, deniedGrantsPhrase(denied))
+ case len(candidates) > 1 && len(denied) == len(candidates):
+ return "", "", unreachableBecause("No working Prometheus endpoint found.\n%d candidates: port-forward denied (%s)", len(candidates), deniedGrantsPhrase(denied))
+ case len(candidates) > 1:
+ return "", "", unreachableBecause("No working Prometheus endpoint found.\n%d candidates did not answer through a port-forward", len(candidates))
}
- return "", "", errPrometheusUnreachable
+ return "", "", lastErr
+}
+
+// A port-forward can be refused at any of its steps: reading the Service,
+// listing its Pods, or creating pods/portforward. The API server's Forbidden
+// message names the verb and resource that was refused.
+var forbiddenOperationRe = regexp.MustCompile(`cannot (\w+) resource "([^"]+)"`)
+
+// deniedGrant is the grant a Forbidden port-forward error names, or "" when
+// the message does not say.
+func deniedGrant(err error) string {
+ if m := forbiddenOperationRe.FindStringSubmatch(err.Error()); m != nil {
+ return m[1] + " " + m[2]
+ }
+ return ""
+}
+
+func deniedGrantsPhrase(grants []string) string {
+ var named []string
+ for _, g := range grants {
+ if g != "" && !slices.Contains(named, g) {
+ named = append(named, g)
+ }
+ }
+ if len(named) == 0 {
+ return "permission denied"
+ }
+ return "needs " + strings.Join(named, ", ")
}
// logDiscoveryEnded logs a discovery that ended on a context error, telling a
@@ -525,3 +585,7 @@ func (c *Client) markConnected(addr, basePath, identity string, gen uint64) bool
c.lastDiscoverAt = time.Time{}
return true
}
+
+func prometheusCandidateProbeError(namespace, name string) error {
+ return fmt.Errorf("No working Prometheus endpoint found.\nCandidate %s/%s did not respond after port-forward", namespace, name)
+}
diff --git a/internal/prometheus/discovery_test.go b/internal/prometheus/discovery_test.go
index e8bb29a126..dd364e224f 100644
--- a/internal/prometheus/discovery_test.go
+++ b/internal/prometheus/discovery_test.go
@@ -3,6 +3,7 @@ package prometheus
import (
"context"
"errors"
+ "fmt"
"net/http"
"net/http/httptest"
"strings"
@@ -498,3 +499,44 @@ func TestProbeCandidatesWithReasons_DeadlineMarksLaunchedProbesAsTransportFailur
t.Fatalf("reasons[%d] = %q, want empty for a probe the pass never launched", maxConcurrentProbes, reasons[maxConcurrentProbes])
}
}
+
+func TestUnreachableErrorsAreOneSentenceAndKeepTheirSentinels(t *testing.T) {
+ for _, err := range []error{
+ errPrometheusUnreachable,
+ unreachableBecause("Radar found %d services that may be Prometheus but may not port-forward to them (needs create pods/portforward)", 2),
+ } {
+ if !errors.Is(err, errPrometheusUnreachable) || !errors.Is(err, ErrPrometheusNotFound) {
+ t.Errorf("%v lost its sentinel", err)
+ }
+ if strings.Contains(err.Error(), ErrPrometheusNotFound.Error()) {
+ t.Errorf("%q claims nothing was found", err)
+ }
+ }
+}
+
+func TestDeniedGrantNamesTheRefusedStep(t *testing.T) {
+ listPods := fmt.Errorf("failed to find pod for service prom: %w", errors.New(`pods is forbidden: User "u" cannot list resource "pods" in API group "" in the namespace "monitoring"`))
+ portForward := fmt.Errorf("port-forward failed: %w", errors.New(`pods "prom-0" is forbidden: User "u" cannot create resource "pods/portforward" in API group "" in the namespace "monitoring"`))
+ if got := deniedGrantsPhrase([]string{deniedGrant(listPods)}); got != "needs list pods" {
+ t.Errorf("list pods refusal = %q", got)
+ }
+ if got := deniedGrantsPhrase([]string{deniedGrant(portForward), deniedGrant(portForward), deniedGrant(listPods)}); got != "needs create pods/portforward, list pods" {
+ t.Errorf("mixed = %q", got)
+ }
+ if got := deniedGrantsPhrase([]string{deniedGrant(errors.New("forbidden"))}); got != "permission denied" {
+ t.Errorf("unnamed = %q", got)
+ }
+}
+
+func TestPrometheusRejectedCandidateWording(t *testing.T) {
+ err := prometheusCandidateProbeError("cert-manager", "cert-manager")
+ if !strings.HasPrefix(err.Error(), "No working Prometheus endpoint found") || !strings.Contains(err.Error(), "Candidate cert-manager/cert-manager") {
+ t.Fatal(err)
+ }
+ if strings.Contains(err.Error(), "Prometheus at cert-manager") {
+ t.Fatal(err)
+ }
+ if lines := strings.Split(err.Error(), "\n"); len(lines) != 2 || !strings.HasPrefix(lines[1], "Candidate cert-manager/cert-manager") {
+ t.Fatalf("candidate details must follow the heading: %q", err.Error())
+ }
+}
diff --git a/internal/prometheus/pvc_usage.go b/internal/prometheus/pvc_usage.go
new file mode 100644
index 0000000000..d0eabaad4f
--- /dev/null
+++ b/internal/prometheus/pvc_usage.go
@@ -0,0 +1,147 @@
+package prometheus
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "math"
+
+ "github.com/skyhook-io/radar/internal/errorlog"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+// PVCUsage is one claim's filesystem use as kubelet reports it.
+type PVCUsage struct {
+ UsedBytes int64
+ CapacityBytes int64
+ Ratio float64
+}
+
+const (
+ PVCUsageAvailable = "available"
+ PVCUsageNoPrometheus = "no_prometheus"
+ PVCUsageQueryFailed = "query_failed"
+ // The claims' series carry more than one cluster identity, or the proven
+ // identity is absent from them: a merged value could be another cluster's.
+ PVCUsageAmbiguous = "ambiguous_scope"
+ PVCUsageScopeMismatch = "scope_mismatch"
+)
+
+// PVCUsageBatch answers for a set of claims in one namespace. Status describes
+// the query as a whole; a claim missing from both Usage and Invalid had no
+// series, which does not establish whether its driver reports volume stats or
+// whether kubelet is scraped at all.
+type PVCUsageBatch struct {
+ Status string
+ Error string
+ Usage map[string]PVCUsage
+ Invalid map[string]bool
+ // Isolation says how the series were tied to this cluster.
+ Isolation SeriesIsolation
+}
+
+// Claims per query: keeps the regex matcher and the answer small whatever the
+// namespace holds.
+const pvcUsageBatchSize = 100
+
+// QueryPVCUsage reads kubelet_volume_stats_{used,capacity}_bytes for the named
+// claims of one namespace, held to this cluster's identity (anchors are the
+// Pods that prove it). Authorization is the caller's: this reads nothing the
+// caller did not name.
+func (client *Client) QueryPVCUsage(ctx context.Context, namespace string, claims []string, anchors []prom.WorkloadPodIdentity) PVCUsageBatch {
+ if client == nil {
+ return PVCUsageBatch{Status: PVCUsageNoPrometheus, Usage: map[string]PVCUsage{}, Invalid: map[string]bool{}}
+ }
+ if _, _, err := client.EnsureConnected(ctx); err != nil {
+ return PVCUsageBatch{Status: PVCUsageNoPrometheus, Error: err.Error(), Usage: map[string]PVCUsage{}, Invalid: map[string]bool{}}
+ }
+ matchers, iso, err := client.resolveScope(ctx, namespace, pvcScopeProbe(namespace, claims, 0), anchors, nil)
+ if out, failed := pvcScopeFailure(err); failed {
+ if out.Status == PVCUsageQueryFailed {
+ errorlog.Record("prometheus", "warning", "pvc usage scope check failed for namespace %s: %v", namespace, err)
+ }
+ return out
+ }
+ out := queryPVCUsage(ctx, client, namespace, claims, matchers)
+ out.Isolation = iso.forClaims()
+ if out.Status == PVCUsageQueryFailed {
+ errorlog.Record("prometheus", "warning", "pvc usage query failed for namespace %s: %s", namespace, out.Error)
+ }
+ return out
+}
+
+func pvcScopeFailure(err error) (PVCUsageBatch, bool) {
+ out := PVCUsageBatch{Usage: map[string]PVCUsage{}, Invalid: map[string]bool{}}
+ switch {
+ case err == nil:
+ return out, false
+ case errors.Is(err, ErrScopeAmbiguous):
+ out.Status, out.Error = PVCUsageAmbiguous, "these claim names have volume stats under more than one cluster identity in this Prometheus"
+ case errors.Is(err, ErrScopeMismatch):
+ out.Status, out.Error = PVCUsageScopeMismatch, "the cluster identity proven for this cluster does not appear on these claims' volume stats"
+ default:
+ out.Status, out.Error = PVCUsageQueryFailed, err.Error()
+ }
+ return out, true
+}
+
+func queryPVCUsage(ctx context.Context, q seriesQuerier, namespace string, claims []string, matchers string) PVCUsageBatch {
+ out := PVCUsageBatch{Status: PVCUsageAvailable, Usage: map[string]PVCUsage{}, Invalid: map[string]bool{}}
+ for _, sel := range ClaimSelectors(namespace, claims) {
+ if err := queryPVCUsageBatch(ctx, q, withScope(sel, matchers), &out); err != nil {
+ return PVCUsageBatch{Status: PVCUsageQueryFailed, Error: err.Error(), Usage: map[string]PVCUsage{}, Invalid: map[string]bool{}}
+ }
+ }
+ return out
+}
+
+func queryPVCUsageBatch(ctx context.Context, q seriesQuerier, selector string, out *PVCUsageBatch) error {
+ used, err := q.Query(ctx, fmt.Sprintf(`max by (persistentvolumeclaim) (kubelet_volume_stats_used_bytes{%s})`, selector))
+ if err != nil {
+ return err
+ }
+ capacity, err := q.Query(ctx, fmt.Sprintf(`max by (persistentvolumeclaim) (kubelet_volume_stats_capacity_bytes{%s})`, selector))
+ if err != nil {
+ return err
+ }
+ usedBy, capBy := valuesByClaim(used), valuesByClaim(capacity)
+ for claim, u := range usedBy {
+ c, ok := capBy[claim]
+ if !ok {
+ continue
+ }
+ usage, valid := pvcUsageOf(u, c)
+ if !valid {
+ out.Invalid[claim] = true
+ continue
+ }
+ out.Usage[claim] = usage
+ }
+ return nil
+}
+
+func valuesByClaim(res *prom.QueryResult) map[string]float64 {
+ out := map[string]float64{}
+ if res == nil {
+ return out
+ }
+ for _, s := range res.Series {
+ claim := s.Labels["persistentvolumeclaim"]
+ if claim == "" || len(s.DataPoints) == 0 {
+ continue
+ }
+ out[claim] = s.DataPoints[0].Value
+ }
+ return out
+}
+
+// pvcUsageOf applies the same validity rules as the single-claim endpoint:
+// NaN, infinities, negative use, a capacity under one byte or values past
+// int64 are not measurements.
+func pvcUsageOf(used, capacity float64) (PVCUsage, bool) {
+ if used != used || capacity != capacity || math.IsInf(used, 0) || math.IsInf(capacity, 0) ||
+ used < 0 || capacity < 1 || used >= math.Exp2(63) || capacity >= math.Exp2(63) {
+ return PVCUsage{}, false
+ }
+ return PVCUsage{UsedBytes: int64(used), CapacityBytes: int64(capacity), Ratio: used / capacity}, true
+}
diff --git a/internal/prometheus/pvc_usage_test.go b/internal/prometheus/pvc_usage_test.go
index 10c2e86b01..a75e53aa94 100644
--- a/internal/prometheus/pvc_usage_test.go
+++ b/internal/prometheus/pvc_usage_test.go
@@ -1,12 +1,15 @@
package prometheus
import (
+ "context"
"encoding/json"
"fmt"
"net/http"
"net/http/httptest"
"strings"
"testing"
+
+ "github.com/skyhook-io/radar/pkg/prom"
)
func TestPVCUsageAvailability(t *testing.T) {
@@ -36,6 +39,10 @@ func TestPVCUsageAvailability(t *testing.T) {
return
}
calls++
+ if strings.HasPrefix(query, "max(count by (") {
+ _, _ = w.Write([]byte(`{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1700000000,"1"]}]}}`))
+ return
+ }
value, kind := tc.capacity, "capacity"
if strings.Contains(query, "used_bytes") {
value, kind = tc.used, "used"
@@ -90,3 +97,66 @@ func TestPVCUsageAvailability(t *testing.T) {
})
}
}
+
+// Same-named claims in two clusters sharing one Prometheus: east is 95 of
+// 100 GiB, west 10 of 1000 GiB. Unscoped max() would report 95 of 1000.
+func TestQueryPVCUsageHoldsToOneClusterIdentity(t *testing.T) {
+ const gi = 1 << 30
+ q := &fakeCNPGQuerier{instant: func(query string) (*prom.QueryResult, error) {
+ east := strings.Contains(query, `cluster="east"`)
+ switch {
+ case strings.HasPrefix(query, "max(count by ("):
+ return &prom.QueryResult{Series: []prom.Series{vec(nil, 2)}}, nil
+ case strings.Contains(query, "used_bytes") && east:
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"persistentvolumeclaim": "pg-1"}, 95*gi)}}, nil
+ case strings.Contains(query, "capacity_bytes") && east:
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"persistentvolumeclaim": "pg-1"}, 100*gi)}}, nil
+ case strings.Contains(query, "used_bytes"):
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"persistentvolumeclaim": "pg-1"}, 95*gi)}}, nil
+ default:
+ return &prom.QueryResult{Series: []prom.Series{vec(map[string]string{"persistentvolumeclaim": "pg-1"}, 1000*gi)}}, nil
+ }
+ }}
+ _, _, err := decideScope(context.Background(), q, pvcScopeProbe("db", []string{"pg-1"}, 0), nil)
+ if out, failed := pvcScopeFailure(err); !failed || out.Status != PVCUsageAmbiguous || len(out.Usage) != 0 {
+ t.Fatalf("two identities, unproven = %+v (err %v), want ambiguous with no value", out, err)
+ }
+
+ m, _, err := decideScope(context.Background(), q, pvcScopeProbe("db", []string{"pg-1"}, 0), map[string]string{"cluster": "east"})
+ if err != nil {
+ t.Fatal(err)
+ }
+ out := queryPVCUsage(context.Background(), q, "db", []string{"pg-1"}, m)
+ if u := out.Usage["pg-1"]; out.Status != PVCUsageAvailable || u.CapacityBytes != 100*gi || u.Ratio < 0.94 {
+ t.Fatalf("proven east = %+v, want 95 of 100 GiB", out)
+ }
+}
+
+func TestPVCUsageHandlerRefusesMergedClusters(t *testing.T) {
+ upstream := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ query := r.URL.Query().Get("query")
+ if query == "up" {
+ _, _ = w.Write([]byte(authProbeBody))
+ return
+ }
+ result := `[{"metric":{},"value":[1700000000,"1024"]}]`
+ if strings.HasPrefix(query, "max(count by (") {
+ result = `[{"metric":{},"value":[1700000000,"2"]}]`
+ }
+ _, _ = fmt.Fprintf(w, `{"status":"success","data":{"resultType":"vector","result":%s}}`, result)
+ }))
+ defer upstream.Close()
+ Initialize(nil, nil, "pvc-test")
+ SetManualURL(upstream.URL)
+ SetAuthGate(func(*http.Request, string, string, string, string) bool { return true })
+ defer func() { SetAuthGate(nil); Reset(); Initialize(nil, nil, "") }()
+ w := httptest.NewRecorder()
+ metricsRouter().ServeHTTP(w, httptest.NewRequest("GET", "/prometheus/pvc/demo/disk", nil))
+ var response PVCUsageResponse
+ if err := json.Unmarshal(w.Body.Bytes(), &response); err != nil {
+ t.Fatal(err)
+ }
+ if response.Status != PVCUsageAmbiguous || response.HasData || response.Used != 0 {
+ t.Fatalf("response = %+v, want ambiguous_scope without a value", response)
+ }
+}
diff --git a/internal/prometheus/rightsizing.go b/internal/prometheus/rightsizing.go
index 227b17b7cb..20b45a9773 100644
--- a/internal/prometheus/rightsizing.go
+++ b/internal/prometheus/rightsizing.go
@@ -1011,7 +1011,7 @@ type PVCUsageResponse struct {
Capacity int64 `json:"capacity"` // bytes
Ratio float64 `json:"ratio"` // 0.0 - 1.0
HasData bool `json:"hasData"`
- Status string `json:"status"` // available, no_series, invalid_data, query_failed
+ Status string `json:"status"` // available, no_series, invalid_data, query_failed, ambiguous_scope, scope_mismatch
}
// handlePVCUsage returns current usage for a PVC, computed from
@@ -1034,13 +1034,28 @@ func handlePVCUsage(w http.ResponseWriter, r *http.Request) {
ns := prom.SanitizeLabelValue(namespace)
pvc := prom.SanitizeLabelValue(name)
+ resp := PVCUsageResponse{Namespace: namespace, Name: name, Status: "query_failed"}
+
+ // A claim of the same name in another cluster sharing this Prometheus
+ // would otherwise be merged in by the max() below.
+ matchers, _, err := resolveScope(r.Context(), namespace, pvcScopeProbe(namespace, []string{name}, 0), nil, k8s.GetResourceCache())
+ if failed, bad := pvcScopeFailure(err); bad {
+ if failed.Status == PVCUsageQueryFailed {
+ errorlog.Record("prometheus", "warning", "pvc usage scope check failed for %s/%s: %v", namespace, name, err)
+ }
+ resp.Status = failed.Status
+ writeJSON(w, http.StatusOK, resp)
+ return
+ }
+ scope := ""
+ if matchers != "" {
+ scope = "," + matchers
+ }
// kubelet's native label is `persistentvolumeclaim`; clusters with custom
// relabeling that renamed it will return no series.
- usedQuery := fmt.Sprintf(`max(kubelet_volume_stats_used_bytes{namespace='%s',persistentvolumeclaim='%s'})`, ns, pvc)
- capQuery := fmt.Sprintf(`max(kubelet_volume_stats_capacity_bytes{namespace='%s',persistentvolumeclaim='%s'})`, ns, pvc)
-
- resp := PVCUsageResponse{Namespace: namespace, Name: name, Status: "query_failed"}
+ usedQuery := fmt.Sprintf(`max(kubelet_volume_stats_used_bytes{namespace='%s',persistentvolumeclaim='%s'%s})`, ns, pvc, scope)
+ capQuery := fmt.Sprintf(`max(kubelet_volume_stats_capacity_bytes{namespace='%s',persistentvolumeclaim='%s'%s})`, ns, pvc, scope)
usedRes, err := client.Query(r.Context(), usedQuery)
if err != nil {
@@ -1064,16 +1079,21 @@ func handlePVCUsage(w http.ResponseWriter, r *http.Request) {
used := firstValue(usedRes)
capacity := firstValue(capRes)
- if used == nil || capacity == nil || math.IsInf(*used, 0) || math.IsInf(*capacity, 0) ||
- *used < 0 || *capacity < 1 || *used >= math.Exp2(63) || *capacity >= math.Exp2(63) {
+ if used == nil || capacity == nil {
+ resp.Status = "invalid_data"
+ writeJSON(w, http.StatusOK, resp)
+ return
+ }
+ usage, ok := pvcUsageOf(*used, *capacity)
+ if !ok {
resp.Status = "invalid_data"
writeJSON(w, http.StatusOK, resp)
return
}
- resp.Used = int64(*used)
- resp.Capacity = int64(*capacity)
- resp.Ratio = *used / *capacity
+ resp.Used = usage.UsedBytes
+ resp.Capacity = usage.CapacityBytes
+ resp.Ratio = usage.Ratio
resp.HasData = true
resp.Status = "available"
writeJSON(w, http.StatusOK, resp)
diff --git a/internal/prometheus/series_scope.go b/internal/prometheus/series_scope.go
new file mode 100644
index 0000000000..bcceb1b795
--- /dev/null
+++ b/internal/prometheus/series_scope.go
@@ -0,0 +1,184 @@
+package prometheus
+
+import (
+ "context"
+ "errors"
+ "regexp"
+ "sort"
+ "strconv"
+ "strings"
+ "time"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/prom"
+)
+
+// A series scope keeps a query to this Kubernetes cluster's series in a
+// Prometheus that may hold several clusters' series under the same namespace
+// and object names.
+
+const (
+ SeriesIsolationConfigured = "configured"
+ SeriesIsolationVerified = "verified"
+ SeriesIsolationUnverified = "unverified"
+)
+
+// ErrScopeAmbiguous means the selected series exist under more than one
+// cluster identity in this Prometheus, so any answer could mix clusters.
+var ErrScopeAmbiguous = errors.New("prometheus series scope: the selected series appear under more than one cluster identity")
+
+// ErrScopeMismatch means the verified cluster identity labels select none
+// of the probed series although unscoped ones exist.
+var ErrScopeMismatch = errors.New("prometheus series scope: cluster identity labels do not appear on the selected series")
+
+type seriesQuerier interface {
+ Query(ctx context.Context, query string) (*prom.QueryResult, error)
+ QueryRange(ctx context.Context, query string, start, end time.Time, step time.Duration) (*prom.QueryResult, error)
+}
+
+// SeriesIsolation says how the queries keep to this Kubernetes cluster's series.
+type SeriesIsolation struct {
+ Mode string `json:"mode"`
+ Labels map[string]string `json:"labels,omitempty"`
+ Note string `json:"note"`
+}
+
+func withScope(selector, matchers string) string {
+ if matchers == "" {
+ return selector
+ }
+ return selector + "," + matchers
+}
+
+// scopeProbe names the series whose identity labels decide a scope: one
+// selector per batch and, for a chart over a range, the window every identity
+// is collected over (an instant check would miss one that stopped reporting
+// minutes ago but still fills the chart).
+type scopeProbe struct {
+ metric string
+ // key is the label one Kubernetes object's series share (pod, claim).
+ key string
+ selectors []string
+ window time.Duration
+}
+
+func (p scopeProbe) over(sel string) string {
+ if p.window <= 0 {
+ return p.metric + "{" + sel + "}"
+ }
+ return "count_over_time(" + p.metric + "{" + sel + "}[" + p.window.String() + "])"
+}
+
+// ResolvePVCScope decides the cluster-identity matchers for the named claims'
+// kubelet volume stats over window (0 for an instant read).
+func (client *Client) ResolvePVCScope(ctx context.Context, namespace string, claims []string, anchors []prom.WorkloadPodIdentity, window time.Duration) (string, SeriesIsolation, error) {
+ m, iso, err := client.resolveScope(ctx, namespace, pvcScopeProbe(namespace, claims, window), anchors, nil)
+ return m, iso.forClaims(), err
+}
+
+// forClaims words an unverified match for volume stats, which are matched by
+// claim name rather than Pod name.
+func (iso SeriesIsolation) forClaims() SeriesIsolation {
+ if iso.Mode == SeriesIsolationUnverified {
+ iso.Note = "Matched by namespace and claim names. Radar couldn't confirm these volume stats belong to this exact cluster (no cluster label it could check)"
+ }
+ return iso
+}
+
+func pvcScopeProbe(namespace string, claims []string, window time.Duration) scopeProbe {
+ return scopeProbe{metric: "kubelet_volume_stats_capacity_bytes", key: "persistentvolumeclaim", selectors: ClaimSelectors(namespace, claims), window: window}
+}
+
+// ClaimSelectors selects the named claims of one namespace, in batches that
+// keep each regex matcher small.
+func ClaimSelectors(namespace string, claims []string) []string {
+ names := make([]string, len(claims))
+ for i, c := range claims {
+ names[i] = regexp.QuoteMeta(c)
+ }
+ sort.Strings(names)
+ var out []string
+ for start := 0; start < len(names); start += pvcUsageBatchSize {
+ end := min(start+pvcUsageBatchSize, len(names))
+ out = append(out, "namespace="+strconv.Quote(namespace)+",persistentvolumeclaim=~"+strconv.Quote(strings.Join(names[start:end], "|")))
+ }
+ return out
+}
+
+// resolveScope applies an operator-configured scope first, then identity
+// labels proven by kube-state-metrics Pod UIDs. Only those two may add
+// matchers: labels merely seen on the series are not identity (the CNPG
+// exporter's own `cluster` label is the database cluster's name).
+// With no anchors, a cache lets the proof use the namespace's current Pods.
+func resolveScope(ctx context.Context, namespace string, probe scopeProbe, anchors []prom.WorkloadPodIdentity, cache *k8s.ResourceCache) (string, SeriesIsolation, error) {
+ return GetClient().resolveScope(ctx, namespace, probe, anchors, cache)
+}
+
+func (client *Client) resolveScope(ctx context.Context, namespace string, probe scopeProbe, anchors []prom.WorkloadPodIdentity, cache *k8s.ResourceCache) (string, SeriesIsolation, error) {
+ if client == nil {
+ return "", SeriesIsolation{}, errors.New("Prometheus client not initialized")
+ }
+ if config, _, configured := client.workloadMetricsConfig(); configured {
+ m, err := config.Matchers()
+ if err != nil {
+ return "", SeriesIsolation{}, err
+ }
+ iso := SeriesIsolation{Mode: SeriesIsolationConfigured, Labels: config.ClusterLabels, Note: "Matched by the cluster labels an operator configured"}
+ if config.SingleCluster {
+ iso.Note = "An operator declared this Prometheus single-cluster"
+ }
+ return m, iso, nil
+ }
+ var verified map[string]string
+ if len(anchors) > 0 || cache != nil {
+ if cfg, err := client.historicalClusterScope(ctx, PodScope{Namespace: namespace, Identities: anchors}, cache); err == nil {
+ verified = cfg.ClusterLabels
+ }
+ }
+ return decideScope(ctx, client, probe, verified)
+}
+
+func decideScope(ctx context.Context, q seriesQuerier, probe scopeProbe, verified map[string]string) (string, SeriesIsolation, error) {
+ if len(verified) > 0 {
+ m, err := prom.WorkloadMetricsScope{ClusterLabels: verified}.Matchers()
+ if err != nil {
+ return "", SeriesIsolation{}, err
+ }
+ for _, sel := range probe.selectors {
+ scoped, err := q.Query(ctx, "count("+probe.over(withScope(sel, m))+")")
+ if err != nil {
+ return "", SeriesIsolation{}, err
+ }
+ if len(scoped.Series) > 0 {
+ return m, verifiedIsolation(verified), nil
+ }
+ }
+ for _, sel := range probe.selectors {
+ all, err := q.Query(ctx, "count("+probe.over(sel)+")")
+ if err != nil {
+ return "", SeriesIsolation{}, err
+ }
+ if len(all.Series) > 0 {
+ return "", SeriesIsolation{}, ErrScopeMismatch
+ }
+ }
+ return m, verifiedIsolation(verified), nil
+ }
+ // Unproven: nothing is pinned. One object (Pod, claim) whose series carry
+ // more than one set of partition labels over the range may be two
+ // clusters' objects of the same name, so that is refused.
+ for _, sel := range probe.selectors {
+ res, err := q.Query(ctx, "max(count by ("+probe.key+") (count by ("+probe.key+","+strings.Join(partitionLabels, ",")+") ("+probe.over(sel)+")))")
+ if err != nil {
+ return "", SeriesIsolation{}, err
+ }
+ if len(res.Series) > 0 && len(res.Series[0].DataPoints) > 0 && res.Series[0].DataPoints[0].Value > 1 {
+ return "", SeriesIsolation{}, ErrScopeAmbiguous
+ }
+ }
+ return "", SeriesIsolation{Mode: SeriesIsolationUnverified, Note: "Matched by namespace and Pod names. Radar couldn't confirm these series belong to this exact cluster (no cluster label it could check)"}, nil
+}
+
+func verifiedIsolation(labels map[string]string) SeriesIsolation {
+ return SeriesIsolation{Mode: SeriesIsolationVerified, Labels: labels, Note: "Matched by cluster labels confirmed against this cluster's Pods"}
+}
diff --git a/internal/server/actions.go b/internal/server/actions.go
new file mode 100644
index 0000000000..9158bb13cc
--- /dev/null
+++ b/internal/server/actions.go
@@ -0,0 +1,188 @@
+package server
+
+import (
+ "context"
+ "encoding/json"
+ "errors"
+ "log"
+ "net/http"
+ "strings"
+ "sync"
+ "time"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/client-go/dynamic"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+// The action contract shared by integrations that write to the cluster on the
+// caller's behalf. A capability says whether an action is offered and why
+// not; a POST binds the context and facts the caller reviewed and is refused
+// with a stable code when they no longer hold. Writes are made with the
+// caller's impersonated client and are never retried automatically.
+
+// Grant lives in auth because Prometheus-backed reads report the grant they
+// lacked too.
+type Grant = auth.Grant
+
+// decodeActionRequest reads the body and checks the reviewed context. The
+// dynamic client is the caller's, snapshotted with the context it belongs to.
+// The error is already written when ok is false.
+func (s *Server) decodeActionRequest(w http.ResponseWriter, r *http.Request) (integration.ActionRequest, dynamic.Interface, bool) {
+ var req integration.ActionRequest
+ if err := decodeBoundedJSONBody(w, r, integration.ActionBodyLimit, &req); err != nil {
+ var tooLarge *http.MaxBytesError
+ if errors.As(err, &tooLarge) {
+ s.writeError(w, http.StatusRequestEntityTooLarge, "request body is too large")
+ return req, nil, false
+ }
+ s.writeError(w, http.StatusBadRequest, "invalid request body: "+err.Error())
+ return req, nil, false
+ }
+ if req.ReviewedContext == "" || req.UID == "" {
+ s.writeError(w, http.StatusBadRequest, "reviewedContext and uid are required")
+ return req, nil, false
+ }
+ dyn, contextName := s.getDynamicClientSnapshotForRequest(r)
+ if dyn == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return req, nil, false
+ }
+ if err := integration.CheckReviewedContext(req.ReviewedContext, contextName); err != nil {
+ s.writeActionError(w, "actions", err, "context", "", "")
+ return req, nil, false
+ }
+ return req, dyn, true
+}
+
+// writeActionError answers a failed action. A refusal keeps its status, code,
+// current facts and completed steps; an apiserver error maps to its status,
+// and a Forbidden one retains the apiserver's exact denial: a multi-resource
+// action can fail on a prerequisite read rather than its write. tag prefixes logs.
+func (s *Server) writeActionError(w http.ResponseWriter, tag string, err error, action, namespace, name string) {
+ var ae *integration.ActionError
+ if errors.As(err, &ae) {
+ if ae.Status >= http.StatusInternalServerError {
+ log.Printf("[%s] Failed to %s %s/%s: %v", tag, sanitizeForLog(action), sanitizeForLog(namespace), sanitizeForLog(name), ae)
+ } else {
+ log.Printf("[%s] %q %s/%s refused %d: %s", tag, sanitizeForLog(action), sanitizeForLog(namespace), sanitizeForLog(name), ae.Status, ae.Message)
+ }
+ body := map[string]any{"error": ae.Message}
+ if ae.Code != "" {
+ body["code"] = ae.Code
+ }
+ if ae.Current != nil {
+ body["current"] = ae.Current
+ }
+ if len(ae.Completed) > 0 {
+ body["completed"] = ae.Completed
+ }
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(ae.Status)
+ if encErr := json.NewEncoder(w).Encode(body); encErr != nil {
+ log.Printf("Failed to encode error response: %v", encErr)
+ }
+ return
+ }
+ msg := err.Error()
+ status := http.StatusInternalServerError
+ switch {
+ case errors.Is(err, context.DeadlineExceeded) || apierrors.IsTimeout(err) || apierrors.IsServerTimeout(err):
+ status = http.StatusGatewayTimeout
+ case apierrors.IsNotFound(err):
+ status = http.StatusNotFound
+ case apierrors.IsForbidden(err):
+ status = http.StatusForbidden
+ case apierrors.IsAlreadyExists(err), apierrors.IsConflict(err):
+ status = http.StatusConflict
+ case apierrors.IsInvalid(err):
+ status = http.StatusUnprocessableEntity
+ }
+ log.Printf("[%s] Failed to %s %s/%s: %v", tag, sanitizeForLog(action), sanitizeForLog(namespace), sanitizeForLog(name), err)
+ s.writeError(w, status, msg)
+}
+
+// grantPermission answers one grant for the caller: allowed, denied, or
+// unknown when the SubjectAccessReview itself failed.
+func (s *Server) grantPermission(r *http.Request, g Grant) string {
+ allowed, authoritative := s.grantDecision(r, g)
+ switch {
+ case !authoritative:
+ return integration.PermissionUnknown
+ case allowed:
+ return integration.PermissionAllowed
+ default:
+ return integration.PermissionDenied
+ }
+}
+
+// grantDecision asks for g as the caller. Subresource answers are memoized on
+// the same per-user cache as canRead, keyed resource/subresource.
+func (s *Server) grantDecision(r *http.Request, g Grant) (bool, bool) {
+ if auth.UserFromContext(r.Context()) == nil {
+ return localCanI(r.Context(), g)
+ }
+ if g.Subresource == "" {
+ return s.canReadDecision(r, g.Group, g.Resource, g.Namespace, g.Verb)
+ }
+ key := g.Resource + "/" + g.Subresource
+ var perms *auth.UserPermissions
+ if user := auth.UserFromContext(r.Context()); user != nil && s.permCache != nil {
+ perms = s.permCache.Get(user.Username, user.Groups)
+ }
+ if perms != nil {
+ if v, ok := perms.CanI(g.Verb, g.Group, key, g.Namespace); ok {
+ return v, true
+ }
+ }
+ allowed, authoritative := s.canReadSubresourceDecision(r, g.Group, g.Resource, g.Subresource, g.Namespace, g.Verb)
+ if authoritative && perms != nil {
+ perms.SetCanI(g.Verb, g.Group, key, g.Namespace, allowed)
+ }
+ return allowed, authoritative
+}
+
+// Without auth the apiserver still enforces the kubeconfig identity's RBAC on
+// every write and proxy read; asking it first lets capabilities name a missing
+// grant instead of offering an action that will fail.
+var (
+ localCanIMu sync.Mutex
+ localCanIMemo = map[string]localCanIEntry{}
+ localCanITTL = 30 * time.Second
+)
+
+type localCanIEntry struct {
+ allowed bool
+ expires time.Time
+}
+
+func localCanI(ctx context.Context, g Grant) (bool, bool) {
+ resource := g.Resource
+ if g.Subresource != "" {
+ resource += "/" + g.Subresource
+ }
+ key := strings.Join([]string{k8s.GetContextName(), g.Verb, g.Group, resource, g.Namespace}, "\x00")
+ now := time.Now()
+ localCanIMu.Lock()
+ if e, ok := localCanIMemo[key]; ok && now.Before(e.expires) {
+ localCanIMu.Unlock()
+ return e.allowed, true
+ }
+ localCanIMu.Unlock()
+ client := k8s.GetClient()
+ if client == nil {
+ return true, true
+ }
+ allowed, apiErr := k8score.CanI(ctx, client, g.Namespace, g.Group, resource, g.Verb)
+ if apiErr {
+ return false, false
+ }
+ localCanIMu.Lock()
+ localCanIMemo[key] = localCanIEntry{allowed: allowed, expires: now.Add(localCanITTL)}
+ localCanIMu.Unlock()
+ return allowed, true
+}
diff --git a/internal/server/ai_diagnose.go b/internal/server/ai_diagnose.go
index a8409a9ed5..2756e7b86b 100644
--- a/internal/server/ai_diagnose.go
+++ b/internal/server/ai_diagnose.go
@@ -14,6 +14,7 @@ import (
"github.com/skyhook-io/radar/internal/ai"
"github.com/skyhook-io/radar/internal/config"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/checks"
"github.com/skyhook-io/radar/pkg/resourcecontext"
@@ -91,7 +92,7 @@ func (s *Server) detectDiagnoseHealth(r *http.Request, kind, group, namespace, n
if canonicalKind == "" {
canonicalKind = kind
}
- issueSum, issueRows := computeIssueSummaryAndRows(cache, s.issueClusterScopedAccess(r), s.issueRelatedResourceAccess(r), gvk.Group, canonicalKind, namespace, name, true)
+ issueSum, issueRows := computeIssueSummaryAndRows(cache, s.issueClusterScopedAccess(r), s.issueRelatedResourceAccess(r), s.issueEvidenceAccess(r), gvk.Group, canonicalKind, namespace, name, true)
auditSum, auditRows := s.computeAuditSummaryAndRows(r, cache, gvk.Group, canonicalKind, namespace, name)
var issueCount int
@@ -316,7 +317,7 @@ func (s *Server) handleDiagnoseStart(w http.ResponseWriter, r *http.Request) {
return
}
if namespace != "" {
- if allowed := s.getUserNamespaces(r, []string{namespace}); noNamespaceAccess(allowed) {
+ if allowed := s.getUserNamespaces(r, []string{namespace}); integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
diff --git a/internal/server/ai_handlers.go b/internal/server/ai_handlers.go
index f733241319..e028705519 100644
--- a/internal/server/ai_handlers.go
+++ b/internal/server/ai_handlers.go
@@ -233,7 +233,7 @@ func (s *Server) handleAIListResources(w http.ResponseWriter, r *http.Request) {
// so we pass nil here to compose cluster-wide.
if !skipContext && level == aicontext.LevelSummary {
idxNamespaces := issueIndexNamespaces(namespaces, kind, group)
- if builder := s.newResourceSummaryContextBuilder(idxNamespaces); builder != nil {
+ if builder := s.newResourceSummaryContextBuilder(r, idxNamespaces); builder != nil {
// Typed list resolves group from each object's TypeMeta —
// MinifyList sets it via SetTypeMeta before producing rows,
// so we can trust apiVersion on the typed source.
@@ -294,7 +294,7 @@ func (s *Server) aiListDynamic(w http.ResponseWriter, r *http.Request, cache *k8
if !skipContext && level == aicontext.LevelSummary {
idxNamespaces := issueIndexNamespaces(namespaces, kind, group)
- if builder := s.newResourceSummaryContextBuilder(idxNamespaces); builder != nil {
+ if builder := s.newResourceSummaryContextBuilder(r, idxNamespaces); builder != nil {
summarycontext.AttachToUnstructuredList(results, allItems, builder)
}
}
@@ -465,7 +465,7 @@ func (s *Server) buildAIResourceContext(r *http.Request, obj runtime.Object, kin
}
canonicalGroup := gvk.Group
- issueSum := computeIssueSummaryForResource(cache, s.issueClusterScopedAccess(r), s.issueRelatedResourceAccess(r), canonicalGroup, canonicalKind, namespace, name)
+ issueSum := computeIssueSummaryForResource(cache, s.issueClusterScopedAccess(r), s.issueRelatedResourceAccess(r), s.issueEvidenceAccess(r), canonicalGroup, canonicalKind, namespace, name)
auditSum := s.computeAuditSummaryForResource(r, cache, canonicalGroup, canonicalKind, namespace, name)
opts := resourcecontext.Options{
@@ -562,15 +562,15 @@ func (s *Server) topologyForContext(namespace string) (*topology.Topology, topol
// summary silently collapses to nil.
//
// Returns nil when no issues match — Build then omits the IssueSummary field.
-func computeIssueSummaryForResource(cache *k8s.ResourceCache, canReadClusterScoped func(kind, group string) bool, canReadRelated func(issues.Ref) bool, group, kind, namespace, name string) *resourcecontext.IssueSummary {
- sum, _ := computeIssueSummaryAndRows(cache, canReadClusterScoped, canReadRelated, group, kind, namespace, name, false)
+func computeIssueSummaryForResource(cache *k8s.ResourceCache, canReadClusterScoped func(kind, group string) bool, canReadRelated func(issues.Ref) bool, canReadEvidence func(issues.EvidenceRead) bool, group, kind, namespace, name string) *resourcecontext.IssueSummary {
+ sum, _ := computeIssueSummaryAndRows(cache, canReadClusterScoped, canReadRelated, canReadEvidence, group, kind, namespace, name, false)
return sum
}
// computeIssueSummaryAndRows additionally returns the matched rows sorted by
// (severity desc, reason asc) — the diagnose health frame shows the actual
// lines, not just the rollup.
-func computeIssueSummaryAndRows(cache *k8s.ResourceCache, canReadClusterScoped func(kind, group string) bool, canReadRelated func(issues.Ref) bool, group, kind, namespace, name string, includeFacts bool) (*resourcecontext.IssueSummary, []issues.Issue) {
+func computeIssueSummaryAndRows(cache *k8s.ResourceCache, canReadClusterScoped func(kind, group string) bool, canReadRelated func(issues.Ref) bool, canReadEvidence func(issues.EvidenceRead) bool, group, kind, namespace, name string, includeFacts bool) (*resourcecontext.IssueSummary, []issues.Issue) {
if cache == nil {
return nil, nil
}
@@ -591,6 +591,7 @@ func computeIssueSummaryAndRows(cache *k8s.ResourceCache, canReadClusterScoped f
Namespaces: namespaces,
CanReadClusterScoped: canReadClusterScoped,
CanReadRelated: canReadRelated,
+ CanReadEvidence: canReadEvidence,
}, group, kind, namespace, name)
if len(matched) == 0 {
return nil, nil
diff --git a/internal/server/applications.go b/internal/server/applications.go
index 1cb7cfd7ad..3fdc0bc072 100644
--- a/internal/server/applications.go
+++ b/internal/server/applications.go
@@ -22,6 +22,8 @@ import (
"github.com/skyhook-io/radar/internal/auth"
"github.com/skyhook-io/radar/internal/helm"
+ "github.com/skyhook-io/radar/internal/imageutil"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
gitopsinsights "github.com/skyhook-io/radar/pkg/gitops/insights"
"github.com/skyhook-io/radar/pkg/health"
@@ -391,7 +393,7 @@ func (s *Server) historyAnchorsForSource(r *http.Request, source *appSourceRef)
func (s *Server) gitOpsHistoryAnchors(r *http.Request, source *appSourceRef) ([]appHistoryAnchor, []string) {
if source.Namespace != "" {
allowed := s.getUserNamespaces(r, []string{source.Namespace})
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
return nil, []string{fmt.Sprintf("No access to source namespace %q.", source.Namespace)}
}
}
@@ -905,7 +907,7 @@ func collectAppWorkloads(ctx context.Context, cache *k8s.ResourceCache, namespac
Name: name,
WorkloadClass: classifyWorkload(kind, rels),
Image: image,
- Version: imageTag(image),
+ Version: imageutil.ImageTag(image),
AppVersion: lbls["app.kubernetes.io/version"],
Health: string(workloadHealth),
Ready: ready,
@@ -2650,23 +2652,6 @@ func cronWorkflowOwnerName(wf *unstructured.Unstructured) string {
return wf.GetLabels()["workflows.argoproj.io/cron-workflow"]
}
-// imageTag extracts the tag from an image ref. Digest-pinned refs (@sha256:…)
-// and untagged refs (implicit :latest) return "" — no false version.
-func imageTag(image string) string {
- if image == "" {
- return ""
- }
- if at := strings.Index(image, "@"); at >= 0 {
- image = image[:at]
- }
- slash := strings.LastIndex(image, "/")
- colon := strings.LastIndex(image, ":")
- if colon > slash {
- return image[colon+1:]
- }
- return ""
-}
-
// imageRepo is the image ref without its tag/digest — the unit version skew is
// measured across: two workloads running the same repo at different tags.
func imageRepo(image string) string {
diff --git a/internal/server/applications_identity.go b/internal/server/applications_identity.go
index 0ce39eebab..be9ea62ad2 100644
--- a/internal/server/applications_identity.go
+++ b/internal/server/applications_identity.go
@@ -10,6 +10,7 @@ import (
"k8s.io/apimachinery/pkg/labels"
listerscorev1 "k8s.io/client-go/listers/core/v1"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/pkg/gitops"
"github.com/skyhook-io/radar/pkg/packages"
"github.com/skyhook-io/radar/pkg/resourceid"
@@ -973,7 +974,7 @@ func collectArgoClaims(items []*unstructured.Unstructured, sourcePaths map[strin
// caller gets nothing, and a scoped caller gets a claim only when it can see at
// least one of the workloads the Application MANAGES — filtering by the
// workload namespaces, not the Application's (which often lives in argocd).
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
return nil
}
var allowed map[string]bool
diff --git a/internal/server/argo_diff.go b/internal/server/argo_diff.go
index 3f3a22a72e..b780788b4a 100644
--- a/internal/server/argo_diff.go
+++ b/internal/server/argo_diff.go
@@ -15,6 +15,7 @@ import (
"sigs.k8s.io/yaml"
"github.com/skyhook-io/radar/internal/argocd"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/argoapi"
"github.com/skyhook-io/radar/pkg/gitops"
@@ -96,7 +97,7 @@ func (s *Server) handleArgoResourceDiff(w http.ResponseWriter, r *http.Request)
// Gate 1: the Application root. A caller who can't see the Application's
// namespace is denied here, before any upstream fetch. Matches the
// namespace-access check parseGitOpsRequest runs for /api/gitops/insights.
- if noNamespaceAccess(s.getUserNamespaces(r, []string{appNamespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{appNamespace})) {
s.writeError(w, http.StatusForbidden, fmt.Sprintf("no access to namespace %q", appNamespace))
return
}
@@ -617,7 +618,7 @@ func (s *Server) handleArgoRevisionMetadata(w http.ResponseWriter, r *http.Reque
// Gate: the caller must be able to see the Application's namespace, matching
// the resource-diff and insights handlers.
- if noNamespaceAccess(s.getUserNamespaces(r, []string{appNamespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{appNamespace})) {
s.writeError(w, http.StatusForbidden, fmt.Sprintf("no access to namespace %q", appNamespace))
return
}
diff --git a/internal/server/audit_handlers.go b/internal/server/audit_handlers.go
index 4555c1b9cd..99e462c632 100644
--- a/internal/server/audit_handlers.go
+++ b/internal/server/audit_handlers.go
@@ -13,6 +13,7 @@ import (
"github.com/skyhook-io/radar/internal/audit"
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/settings"
bp "github.com/skyhook-io/radar/pkg/audit"
@@ -118,7 +119,7 @@ func (s *Server) handleAudit(w http.ResponseWriter, r *http.Request) {
return
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, &bp.ScanResults{Summary: bp.ScanSummary{Categories: map[string]bp.CategorySummary{}}})
return
}
@@ -175,7 +176,7 @@ func (s *Server) handleAuditResource(w http.ResponseWriter, r *http.Request) {
name := chi.URLParam(r, "name")
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []bp.Finding{})
return
}
diff --git a/internal/server/audit_raw_test.go b/internal/server/audit_raw_test.go
index 1268c72bc4..cc89b4ca1d 100644
--- a/internal/server/audit_raw_test.go
+++ b/internal/server/audit_raw_test.go
@@ -3,13 +3,14 @@ package server
import (
"context"
"encoding/json"
- "github.com/go-chi/chi/v5"
- "github.com/skyhook-io/radar/internal/k8s"
"net/http"
"net/http/httptest"
"testing"
"time"
+ "github.com/go-chi/chi/v5"
+
+ "github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/settings"
bp "github.com/skyhook-io/radar/pkg/audit"
)
diff --git a/internal/server/capacity.go b/internal/server/capacity.go
index 735a7b06ae..99c60283f2 100644
--- a/internal/server/capacity.go
+++ b/internal/server/capacity.go
@@ -9,19 +9,21 @@ import (
"time"
"github.com/go-chi/chi/v5"
+ corev1 "k8s.io/api/core/v1"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/client-go/kubernetes"
+
capacitymodel "github.com/skyhook-io/radar/internal/capacity"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/k8s"
internaltimeline "github.com/skyhook-io/radar/internal/timeline"
"github.com/skyhook-io/radar/pkg/capacityapi"
"github.com/skyhook-io/radar/pkg/karpenter"
"github.com/skyhook-io/radar/pkg/subject"
- corev1 "k8s.io/api/core/v1"
- apierrors "k8s.io/apimachinery/pkg/api/errors"
- "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "k8s.io/apimachinery/pkg/labels"
- "k8s.io/apimachinery/pkg/runtime/schema"
- "k8s.io/client-go/kubernetes"
)
const (
@@ -441,7 +443,7 @@ func (s *Server) loadCapacityModel(w http.ResponseWriter, r *http.Request, ident
result.meta.Provider = capacityProvider(result.meta.Provider, nodePools, nodeClaims, nodeClasses, result.meta.Coverage)
resourceCache := k8s.GetResourceCache()
ownerResolutionAllowed, workloadAttributionPartial := capacityOwnerResolutionPermissions(pods, func(group, resource, namespace string) bool {
- return s.canRead(r, group, resource, namespace, "list") && capacityCacheCoversNamespace(resourceCache, resource, namespace)
+ return s.canRead(r, group, resource, namespace, "list") && integration.CacheCoversNamespace(resourceCache, resource, namespace)
})
if workloadAttributionPartial {
coverage := result.meta.Coverage[capacityapi.CoverageWorkloads]
@@ -700,7 +702,7 @@ func (s *Server) loadCapacityPods(r *http.Request, meta *capacityapi.ResponseMet
baseNamespaces := s.capacityNamespacesForUser(r)
namespaces := s.capacityNamespacesForSource(r, baseNamespaces, "", "pods")
explicit := parseNamespaces(r.URL.Query()) != nil || k8s.ForceNamespaceScope
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
meta.Coverage[capacityapi.CoveragePods] = deniedCoverage("pods_list_denied", []string{"scheduledRequests", "aggregateDemand", "workloads", "summary.actions", "demand.summary"})
meta.Coverage[capacityapi.CoverageWorkloads] = meta.Coverage[capacityapi.CoveragePods]
return nil
@@ -712,11 +714,11 @@ func (s *Server) loadCapacityPods(r *http.Request, meta *capacityapi.ResponseMet
return nil
}
sourceNamespaces := namespaces
- cacheNamespaces := capacityNamespacesWithinCache(cache, "pods", sourceNamespaces)
- namespaces = cacheNamespaces.namespaces
- if cacheNamespaces.unavailable {
+ cacheNamespaces := integration.NamespacesWithinCache(cache, "pods", sourceNamespaces)
+ namespaces = cacheNamespaces.Namespaces
+ if cacheNamespaces.Unavailable {
coverage := unavailableCoverage("pod_cache_scope_unavailable", []string{"scheduledRequests", "aggregateDemand", "workloads", "summary.actions", "demand.summary"})
- if explicit || cacheNamespaces.limited && sourceNamespaces == nil {
+ if explicit || cacheNamespaces.Limited && sourceNamespaces == nil {
coverage.Scope = capacityapi.CoverageScopeExplicitNamespaces
} else if sourceNamespaces != nil {
coverage.Scope = capacityapi.CoverageScopeAllAuthorizedNamespaces
@@ -726,11 +728,11 @@ func (s *Server) loadCapacityPods(r *http.Request, meta *capacityapi.ResponseMet
meta.Coverage[capacityapi.CoverageWorkloads] = coverage
return nil
}
- pods := listPodsScoped(cache.Pods(), namespaces)
+ pods := integration.ListPodsScoped(cache.Pods(), namespaces)
scope := capacityapi.CoverageScopeCluster
status := capacityapi.CoverageAvailable
- if namespaces != nil || cacheNamespaces.limited {
- if explicit || cacheNamespaces.limited && sourceNamespaces == nil {
+ if namespaces != nil || cacheNamespaces.Limited {
+ if explicit || cacheNamespaces.Limited && sourceNamespaces == nil {
scope = capacityapi.CoverageScopeExplicitNamespaces
} else {
scope = capacityapi.CoverageScopeAllAuthorizedNamespaces
@@ -739,11 +741,11 @@ func (s *Server) loadCapacityPods(r *http.Request, meta *capacityapi.ResponseMet
status = capacityapi.CoveragePartial
}
}
- if cacheNamespaces.partial {
+ if cacheNamespaces.Partial {
status = capacityapi.CoveragePartial
}
coverage := capacityapi.NewSourceCoverage(status, scope)
- if cacheNamespaces.partial {
+ if cacheNamespaces.Partial {
coverage.ReasonCode = "pod_cache_scope_partial"
}
coverage.Namespaces = append([]string{}, namespaces...)
diff --git a/internal/server/capacity_auth.go b/internal/server/capacity_auth.go
index e2b4619f0e..ef25863cde 100644
--- a/internal/server/capacity_auth.go
+++ b/internal/server/capacity_auth.go
@@ -4,6 +4,7 @@ import (
"net/http"
"slices"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
)
@@ -26,7 +27,7 @@ func (s *Server) capacityNamespacesForUser(r *http.Request) []string {
}
func (s *Server) capacityNamespacesForSource(r *http.Request, namespaces []string, group, resource string) []string {
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
return namespaces
}
if namespaces == nil {
@@ -37,53 +38,3 @@ func (s *Server) capacityNamespacesForSource(r *http.Request, namespaces []strin
}
return s.filterNamespacesByCanRead(r, group, resource, "list", namespaces)
}
-
-type capacityInformerScope interface {
- IsKindClusterWide(string) bool
- KindNamespaces(string) []string
- IsKindReady(string) bool
-}
-
-type capacityCacheNamespaceResult struct {
- namespaces []string
- limited bool
- partial bool
- unavailable bool
-}
-
-func capacityNamespacesWithinCache(cache capacityInformerScope, resource string, requested []string) capacityCacheNamespaceResult {
- result := capacityCacheNamespaceResult{namespaces: requested}
- if cache == nil || cache.IsKindClusterWide(resource) {
- return result
- }
- result.limited = true
- if noNamespaceAccess(requested) {
- return result
- }
- cached := cache.KindNamespaces(resource)
- if len(cached) == 0 {
- result.namespaces = []string{}
- result.unavailable = true
- return result
- }
- result.namespaces = intersectNamespaces(cached, requested)
- if requested == nil {
- result.partial = true
- return result
- }
- for _, namespace := range requested {
- if !slices.Contains(cached, namespace) {
- if len(result.namespaces) == 0 {
- result.unavailable = true
- } else {
- result.partial = true
- }
- return result
- }
- }
- return result
-}
-
-func capacityCacheCoversNamespace(cache capacityInformerScope, resource, namespace string) bool {
- return cache != nil && cache.IsKindReady(resource) && (cache.IsKindClusterWide(resource) || slices.Contains(cache.KindNamespaces(resource), namespace))
-}
diff --git a/internal/server/capacity_auth_test.go b/internal/server/capacity_auth_test.go
index ac548f06f0..733a7fae83 100644
--- a/internal/server/capacity_auth_test.go
+++ b/internal/server/capacity_auth_test.go
@@ -5,11 +5,13 @@ import (
"slices"
"testing"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/capacityapi"
- corev1 "k8s.io/api/core/v1"
- metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
)
func TestCapacityNamespacesBypassSavedViewPreference(t *testing.T) {
@@ -99,7 +101,7 @@ func TestCapacityNamespacesHonorInformerCoverage(t *testing.T) {
limited := capacityTestInformerScope{namespaces: map[string][]string{"pods": {"team-a", "team-b"}}}
tests := []struct {
name string
- cache capacityInformerScope
+ cache integration.InformerScope
requested []string
want []string
wantLimited bool
@@ -115,8 +117,8 @@ func TestCapacityNamespacesHonorInformerCoverage(t *testing.T) {
}
for _, test := range tests {
t.Run(test.name, func(t *testing.T) {
- got := capacityNamespacesWithinCache(test.cache, "pods", test.requested)
- if !slices.Equal(got.namespaces, test.want) || got.limited != test.wantLimited || got.partial != test.wantPartial || got.unavailable != test.wantUnavailable {
+ got := integration.NamespacesWithinCache(test.cache, "pods", test.requested)
+ if !slices.Equal(got.Namespaces, test.want) || got.Limited != test.wantLimited || got.Partial != test.wantPartial || got.Unavailable != test.wantUnavailable {
t.Fatalf("cache namespaces = %#v, want namespaces=%v limited=%v partial=%v unavailable=%v", got, test.want, test.wantLimited, test.wantPartial, test.wantUnavailable)
}
})
@@ -132,13 +134,13 @@ func TestCapacityOwnerResolutionHonorsInformerCoverage(t *testing.T) {
"jobs": {"team-b"},
}}
permissions, partial := capacityOwnerResolutionPermissions([]*corev1.Pod{replicaPod, jobPod}, func(_ string, resource, namespace string) bool {
- return capacityCacheCoversNamespace(cache, resource, namespace)
+ return integration.CacheCoversNamespace(cache, resource, namespace)
})
if permissions[replicaPod] || !permissions[jobPod] || !partial {
t.Fatalf("owner permissions = %#v partial=%v, want ReplicaSet denied and Job allowed", permissions, partial)
}
cache.notReady = map[string]bool{"jobs": true}
- if capacityCacheCoversNamespace(cache, "jobs", "team-b") {
+ if integration.CacheCoversNamespace(cache, "jobs", "team-b") {
t.Fatal("owner resolution used an informer that has not synced")
}
}
diff --git a/internal/server/capacity_groups.go b/internal/server/capacity_groups.go
index 7c97d70b5e..3ba1642c01 100644
--- a/internal/server/capacity_groups.go
+++ b/internal/server/capacity_groups.go
@@ -4,12 +4,14 @@ import (
"log"
"net/http"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ corelisters "k8s.io/client-go/listers/core/v1"
+
capacitymodel "github.com/skyhook-io/radar/internal/capacity"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/autoscalerstatus"
"github.com/skyhook-io/radar/pkg/capacityapi"
- apierrors "k8s.io/apimachinery/pkg/api/errors"
- corelisters "k8s.io/client-go/listers/core/v1"
)
const (
@@ -21,7 +23,7 @@ const (
// needs: informer scope plus the ConfigMap lister. Narrowed to an interface so
// the coverage-verdict/detection pairing is testable without an informer.
type capacityConfigMapSource interface {
- capacityInformerScope
+ integration.InformerScope
ConfigMaps() corelisters.ConfigMapLister
}
@@ -53,7 +55,7 @@ func capacityAutoscalerStatus(allowed bool, cache capacityConfigMapSource) (*aut
if !allowed {
return unavailable(deniedCoverage("autoscaler_status_configmap_denied", impact))
}
- if cache == nil || !capacityCacheCoversNamespace(cache, "configmaps", autoscalerStatusNamespace) {
+ if cache == nil || !integration.CacheCoversNamespace(cache, "configmaps", autoscalerStatusNamespace) {
return unavailable(unavailableCoverage("autoscaler_status_cache_scope", impact))
}
lister := cache.ConfigMaps()
diff --git a/internal/server/capacity_issues.go b/internal/server/capacity_issues.go
index 133d4d7c75..efee298ea4 100644
--- a/internal/server/capacity_issues.go
+++ b/internal/server/capacity_issues.go
@@ -94,6 +94,7 @@ func (s *Server) capacityIssuesForRequest(r *http.Request) capacityIssueProjecti
SkipPodTemplateContext: true,
Limit: issues.NoLimit,
IncludeClusterScopedKarpenter: true,
+ AllowUnfilteredEvidence: true,
})
capacity := make([]issues.Issue, 0, len(composed))
for _, issue := range composed {
diff --git a/internal/server/certificate.go b/internal/server/certificate.go
index 6079257d20..864d2e5ae0 100644
--- a/internal/server/certificate.go
+++ b/internal/server/certificate.go
@@ -6,6 +6,7 @@ import (
"net/http"
"time"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/certs"
"github.com/skyhook-io/radar/pkg/topology"
@@ -61,7 +62,7 @@ func (s *Server) handleCertificates(w http.ResponseWriter, r *http.Request) {
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []certs.Cert{})
return
}
@@ -188,7 +189,7 @@ func (s *Server) handleSecretCertExpiry(w http.ResponseWriter, r *http.Request)
provider := k8s.NewTopologyResourceProvider(cache)
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, map[string]CertExpiry{})
return
}
diff --git a/internal/server/certificate_test.go b/internal/server/certificate_test.go
index 15e4d8ead4..398388afc9 100644
--- a/internal/server/certificate_test.go
+++ b/internal/server/certificate_test.go
@@ -1,9 +1,10 @@
package server
import (
- "github.com/skyhook-io/radar/pkg/certs"
"testing"
"time"
+
+ "github.com/skyhook-io/radar/pkg/certs"
)
func TestProjectCertManagerCert_Issued(t *testing.T) {
diff --git a/internal/server/cloud_install_test.go b/internal/server/cloud_install_test.go
index 08a9066933..94fc05be62 100644
--- a/internal/server/cloud_install_test.go
+++ b/internal/server/cloud_install_test.go
@@ -8,9 +8,6 @@ import (
"encoding/json"
"errors"
"fmt"
-
- apierrors "k8s.io/apimachinery/pkg/api/errors"
- "k8s.io/apimachinery/pkg/runtime/schema"
"log"
"net/http"
"net/http/httptest"
@@ -19,6 +16,9 @@ import (
"testing"
"time"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
"github.com/skyhook-io/radar/internal/cloud"
"github.com/skyhook-io/radar/internal/cloudinstall"
"github.com/skyhook-io/radar/internal/helm"
diff --git a/internal/server/cnpg_actions.go b/internal/server/cnpg_actions.go
new file mode 100644
index 0000000000..c49211f749
--- /dev/null
+++ b/internal/server/cnpg_actions.go
@@ -0,0 +1,136 @@
+package server
+
+import (
+ "errors"
+ "fmt"
+ "log"
+ "net/http"
+ "strings"
+
+ "github.com/go-chi/chi/v5"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+// CloudNativePG write actions. Every write mirrors what `kubectl cnpg` does
+// (promote, restart, reload, fence, hibernate, backup) so the operator sees
+// exactly the request its own tooling would have made. A confirmation binds
+// the facts the dialog showed, not just a resourceVersion: the server re-reads
+// the Cluster, compares those facts and refuses with 409 when anything the
+// user reviewed has moved. Disruptive writes are never retried automatically.
+
+func (s *Server) handleCNPGClusterCapabilities(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := chi.URLParam(r, "namespace")
+ name := chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ dyn, contextName, typed := reader.dynamic, reader.actionContext, reader.Clients.Typed
+ if dyn == nil || typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ resp, err := reader.ClusterCapabilities(r.Context(), cnpgsvc.ActionClients{Dynamic: dyn, Typed: typed}, contextName, namespace, name)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "capabilities", namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+func (s *Server) handleCNPGScheduleCapabilities(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := chi.URLParam(r, "namespace")
+ name := chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ dyn, contextName := reader.dynamic, reader.actionContext
+ if dyn == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ resp, err := reader.ScheduleCapabilities(r.Context(), cnpgsvc.ActionClients{Dynamic: dyn}, contextName, namespace, name)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "capabilities", namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+func (s *Server) handleCNPGClusterAction(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := chi.URLParam(r, "namespace")
+ name := chi.URLParam(r, "name")
+ action := chi.URLParam(r, "action")
+ if !cnpgsvc.IsClusterAction(action) {
+ s.writeError(w, http.StatusBadRequest, fmt.Sprintf("unknown CloudNativePG cluster action %q: must be one of %s", action, strings.Join(cnpgsvc.ClusterActions(), ", ")))
+ return
+ }
+ req, _, ok := s.decodeActionRequest(w, r)
+ if !ok {
+ return
+ }
+ clients, err := s.cnpgActionClients(r, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, action, namespace, name)
+ return
+ }
+ if clients.Typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ auth.AuditLog(r, namespace, name)
+ res, err := cnpgsvc.RunCNPGClusterAction(r.Context(), clients, namespace, name, action, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, action, namespace, name)
+ return
+ }
+ log.Printf("[cnpg] %s on Cluster %s/%s requested", action, sanitizeForLog(namespace), sanitizeForLog(name))
+ s.writeJSON(w, res)
+}
+
+func (s *Server) handleCNPGScheduleAction(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := chi.URLParam(r, "namespace")
+ name := chi.URLParam(r, "name")
+ action := chi.URLParam(r, "action")
+ if !cnpgsvc.IsScheduleAction(action) {
+ s.writeError(w, http.StatusBadRequest, fmt.Sprintf("unknown ScheduledBackup action %q: must be suspend, resume, run, setSchedule or repairMethod", action))
+ return
+ }
+ req, _, ok := s.decodeActionRequest(w, r)
+ if !ok {
+ return
+ }
+ clients, err := s.cnpgActionClients(r, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, action, namespace, name)
+ return
+ }
+ auth.AuditLog(r, namespace, name)
+ res, err := cnpgsvc.RunCNPGScheduleAction(r.Context(), clients, namespace, name, action, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, action, namespace, name)
+ return
+ }
+ log.Printf("[cnpg] %s on ScheduledBackup %s/%s requested", action, sanitizeForLog(namespace), sanitizeForLog(name))
+ s.writeJSON(w, res)
+}
+
+// writeCNPGActionError adds the CloudNativePG reading of an admission-webhook
+// failure to the shared action error answer.
+func (s *Server) writeCNPGActionError(w http.ResponseWriter, err error, action, namespace, name string) {
+ var ae *integration.ActionError
+ if !errors.As(err, &ae) && strings.Contains(err.Error(), "failed calling webhook") {
+ err = integration.RefuseAction(http.StatusServiceUnavailable, cnpgsvc.CodeWebhook, "%s", "The CloudNativePG operator's admission webhook did not answer — the operator may be down: "+err.Error())
+ }
+ s.writeActionError(w, "cnpg", err, action, namespace, name)
+}
diff --git a/internal/server/cnpg_actions_test.go b/internal/server/cnpg_actions_test.go
new file mode 100644
index 0000000000..d16fac1719
--- /dev/null
+++ b/internal/server/cnpg_actions_test.go
@@ -0,0 +1,59 @@
+package server
+
+import (
+ "encoding/json"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "testing"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func TestCNPGActionErrorMapping(t *testing.T) {
+ srv := &Server{}
+ gr := schema.GroupResource{Group: cnpgsvc.Group, Resource: "clusters"}
+ for _, tc := range []struct {
+ name string
+ action string
+ err error
+ status int
+ code string
+ substr []string
+ }{
+ {"forbidden retains the exact denial", "configureArchiving", apierrors.NewForbidden(schema.GroupResource{Group: "barmancloud.cnpg.io", Resource: "objectstores"}, "store", errors.New(`User "reader" cannot get resource "objectstores" in namespace "db"`)), http.StatusForbidden, "",
+ []string{"cannot get resource", "objectstores", "reader"}},
+ {"webhook down", "backup", apierrors.NewInternalError(errors.New(`failed calling webhook "vbackup.cnpg.io": connection refused`)), http.StatusServiceUnavailable, cnpgsvc.CodeWebhook,
+ []string{"admission webhook did not answer", "connection refused"}},
+ {"invalid", "backup", apierrors.NewInvalid(schema.GroupKind{Group: cnpgsvc.Group, Kind: "Backup"}, "b", nil), http.StatusUnprocessableEntity, "", []string{"is invalid"}},
+ {"not found", "restart", apierrors.NewNotFound(gr, "pg"), http.StatusNotFound, "", []string{"not found"}},
+ {"refusal", "fence", integration.BlockedAction("It is fenced already"), http.StatusConflict, integration.ActionCodeBlocked, []string{"fenced already"}},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ w := httptest.NewRecorder()
+ srv.writeCNPGActionError(w, tc.err, tc.action, "db", "pg")
+ if w.Code != tc.status {
+ t.Errorf("status = %d, want %d", w.Code, tc.status)
+ }
+ var body map[string]any
+ _ = json.Unmarshal(w.Body.Bytes(), &body)
+ if tc.code != "" && body["code"] != tc.code {
+ t.Errorf("code = %v, want %s", body["code"], tc.code)
+ }
+ msg, _ := body["error"].(string)
+ if apierrors.IsForbidden(tc.err) && msg != tc.err.Error() {
+ t.Errorf("denial was rewritten: %q", msg)
+ }
+ for _, s := range tc.substr {
+ if !strings.Contains(msg, s) {
+ t.Errorf("error %q does not contain %q", msg, s)
+ }
+ }
+ })
+ }
+}
diff --git a/internal/server/cnpg_cluster_activity.go b/internal/server/cnpg_cluster_activity.go
index a3713a455f..56adc37bdc 100644
--- a/internal/server/cnpg_cluster_activity.go
+++ b/internal/server/cnpg_cluster_activity.go
@@ -1,100 +1,25 @@
package server
import (
+ "errors"
"log"
"net/http"
- "sort"
"strconv"
- "strings"
"time"
"github.com/go-chi/chi/v5"
- "github.com/skyhook-io/radar/internal/k8s"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/timeline"
- "github.com/skyhook-io/radar/pkg/resourceid"
- pkgtimeline "github.com/skyhook-io/radar/pkg/timeline"
)
const (
cnpgActivityDefaultWindow = 24 * time.Hour
cnpgActivityDefaultLimit = 200
cnpgActivityMaxLimit = 1000
- // cnpgActivityScanLimit bounds the rows read to attribute history. A
- // namespace that outgrows it reports truncated rather than silently
- // dropping its oldest attribution.
- cnpgActivityScanLimit = 10000
)
-// CNPGClusterActivityResponse is GET /api/cnpg/clusters/{namespace}/{name}/activity.
-// Oldest is the earliest row the store still holds for the namespace — the
-// floor below which absence means "not retained", not "didn't happen".
-// AttributionSince is the earliest visible row that carries this Cluster's
-// retained cnpg.io/cluster attribution; before it, deleted children cannot be
-// attributed. Both are null when nothing is held.
-type CNPGClusterActivityResponse struct {
- Events []timeline.TimelineEvent `json:"events"`
- Oldest *time.Time `json:"oldest"`
- AttributionSince *time.Time `json:"attributionSince"`
- Truncated bool `json:"truncated"`
-}
-
-type cnpgActivityKind struct {
- group, resource string
-}
-
-// cnpgActivityKinds are the kinds whose rows can belong to one Cluster: the
-// Cluster itself, its instance Pods, and every namespaced CNPG kind.
-var cnpgActivityKinds = func() map[string]cnpgActivityKind {
- out := map[string]cnpgActivityKind{"/Pod": {group: "", resource: "pods"}}
- for _, k := range cnpgWorkspaceKinds {
- if !k.clusterScoped {
- out[k.group+"/"+k.kind] = cnpgActivityKind{group: k.group, resource: k.resource}
- }
- }
- return out
-}()
-
-func cnpgActivityKindNames() []string {
- seen := map[string]bool{}
- var out []string
- for key := range cnpgActivityKinds {
- _, kind, _ := strings.Cut(key, "/")
- if !seen[kind] {
- seen[kind] = true
- out = append(out, kind)
- }
- }
- sort.Strings(out)
- return out
-}
-
-// cnpgRowAttribution decides whether a timeline row is about the named
-// Cluster. Rows about the Cluster match by identity; instance Pods by their
-// controller owner; CNPG children by the retained cnpg.io/cluster label, which
-// survives their deletion. liveUID is the UID of the Cluster that exists now
-// under this name, or "" when none does; when set, only Pods it controlled
-// count, so a previous same-named Cluster's instances don't merge into a
-// recreated one's history.
-func cnpgRowAttribution(e *timeline.TimelineEvent, name, liveUID string) (matched, labelled bool) {
- group := resourceid.GroupFromAPIVersion(e.APIVersion)
- if _, ok := cnpgActivityKinds[group+"/"+e.Kind]; !ok {
- return false, false
- }
- labelled = e.Labels[pkgtimeline.CNPGClusterLabel] == name
- switch {
- case e.Kind == "Cluster" && group == cnpgGroup:
- return e.Name == name, false
- case e.Kind == "Pod" && group == "":
- o := e.Owner
- owned := o != nil && o.Kind == "Cluster" && o.Name == name && resourceid.GroupFromAPIVersion(o.APIVersion) == cnpgGroup &&
- (liveUID == "" || o.UID == liveUID)
- return owned, owned && labelled
- default:
- return labelled, labelled
- }
-}
-
// handleCNPGClusterActivity serves the Cluster's history from the timeline
// store: the Cluster, its instance Pods, and the CNPG objects attributed to it
// — including ones since deleted — with the K8s Events about each. Rows about
@@ -104,11 +29,11 @@ func (s *Server) handleCNPGClusterActivity(w http.ResponseWriter, r *http.Reques
if !s.requireConnected(w) {
return
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
- if !s.canRead(r, cnpgGroup, "clusters", namespace, "get") {
+ if !s.canRead(r, cnpgsvc.Group, "clusters", namespace, "get") {
s.writeError(w, http.StatusForbidden, "no access to clusters.postgresql.cnpg.io in namespace "+namespace)
return
}
@@ -123,6 +48,15 @@ func (s *Server) handleCNPGClusterActivity(w http.ResponseWriter, r *http.Reques
}
since = t
}
+ var until time.Time
+ if raw := r.URL.Query().Get("until"); raw != "" {
+ t, err := time.Parse(time.RFC3339, raw)
+ if err != nil || !t.After(since) {
+ s.writeError(w, http.StatusBadRequest, "invalid until "+strconv.Quote(raw)+" (expected RFC3339 after since)")
+ return
+ }
+ until = t
+ }
limit := cnpgActivityDefaultLimit
if raw := r.URL.Query().Get("limit"); raw != "" {
n, err := strconv.Atoi(raw)
@@ -138,114 +72,17 @@ func (s *Server) handleCNPGClusterActivity(w http.ResponseWriter, r *http.Reques
s.writeError(w, http.StatusServiceUnavailable, "Timeline store not available")
return
}
- clusterContext := k8s.ActiveClusterContext()
- rows, err := store.Query(r.Context(), timeline.QueryOptions{
- Namespaces: []string{namespace},
- Kinds: cnpgActivityKindNames(),
- APIGroups: []string{"", cnpgGroup, cnpgBarmanGroup},
- ClusterContext: clusterContext,
- IncludeManaged: true,
- IncludeK8sEvents: true,
- Limit: cnpgActivityScanLimit,
- })
+ reader := s.cnpgReader(r)
+ response, err := reader.ClusterActivity(r.Context(), store, reader.ClusterContext, namespace, name, cnpgsvc.ActivityOptions{Since: since, Until: until, Limit: limit})
if err != nil {
- log.Printf("[cnpg] Failed to query activity for %s/%s: %v", namespace, name, err)
- s.writeError(w, http.StatusInternalServerError, err.Error())
- return
- }
- var liveUID string
- if cache := k8s.GetResourceCache(); cache != nil {
- if live, err := findCNPGCluster(r.Context(), cache, namespace, name); err == nil && live != nil {
- liveUID = string(live.GetUID())
- }
- }
- scanCapped := len(rows) >= cnpgActivityScanLimit
-
- // Attribution is carried by the subject's own rows; K8s Event rows about
- // a subject whose enrichment was already gone carry only its UID.
- attributedUIDs := map[string]bool{}
- matched := make([]bool, len(rows))
- labelled := make([]bool, len(rows))
- for i := range rows {
- matched[i], labelled[i] = cnpgRowAttribution(&rows[i], name, liveUID)
- if matched[i] && rows[i].UID != "" {
- attributedUIDs[rows[i].UID] = true
- }
- }
-
- allowed := map[string]bool{}
- var eventsAllowed *bool
- canList := func(e *timeline.TimelineEvent) bool {
- if e.Source == timeline.SourceK8sEvent {
- if eventsAllowed == nil {
- ok := s.canRead(r, "", "events", namespace, "list")
- eventsAllowed = &ok
- }
- if !*eventsAllowed {
- return false
- }
- }
- key := resourceid.GroupFromAPIVersion(e.APIVersion) + "/" + e.Kind
- ok, seen := allowed[key]
- if !seen {
- target, known := cnpgActivityKinds[key]
- ok = known && s.canRead(r, target.group, target.resource, namespace, "list")
- allowed[key] = ok
- }
- return ok
- }
-
- resp := CNPGClusterActivityResponse{Events: []timeline.TimelineEvent{}}
- seenIDs := map[string]bool{}
- var windowed []timeline.TimelineEvent
- for i := range rows {
- e := &rows[i]
- if !matched[i] && (e.UID == "" || !attributedUIDs[e.UID]) {
- continue
+ operation := "query activity"
+ var readErr *cnpgsvc.ActivityReadError
+ if errors.As(err, &readErr) {
+ operation = readErr.Operation
}
- if !canList(e) || seenIDs[e.ID] {
- continue
- }
- seenIDs[e.ID] = true
- if labelled[i] && (resp.AttributionSince == nil || e.Timestamp.Before(*resp.AttributionSince)) {
- t := e.Timestamp.UTC()
- resp.AttributionSince = &t
- }
- if e.Timestamp.Before(since) {
- continue
- }
- windowed = append(windowed, *e)
- }
- sort.SliceStable(windowed, func(i, j int) bool {
- if !windowed[i].Timestamp.Equal(windowed[j].Timestamp) {
- return windowed[i].Timestamp.After(windowed[j].Timestamp)
- }
- return windowed[i].ID < windowed[j].ID
- })
- resp.Truncated = scanCapped || len(windowed) > limit
- if len(windowed) > limit {
- windowed = windowed[:limit]
- }
- if windowed != nil {
- resp.Events = windowed
- }
-
- oldest, err := store.Query(r.Context(), timeline.QueryOptions{
- Namespaces: []string{namespace},
- ClusterContext: clusterContext,
- IncludeManaged: true,
- IncludeK8sEvents: true,
- SequenceOrder: timeline.SequenceOrderAscending,
- Limit: 1,
- })
- if err != nil {
- log.Printf("[cnpg] Failed to query retention floor for %s/%s: %v", namespace, name, err)
+ log.Printf("[cnpg] Failed to %s for %s/%s: %v", operation, sanitizeForLog(namespace), sanitizeForLog(name), err)
s.writeError(w, http.StatusInternalServerError, err.Error())
return
}
- if len(oldest) > 0 {
- t := oldest[0].Timestamp.UTC()
- resp.Oldest = &t
- }
- s.writeJSON(w, resp)
+ s.writeJSON(w, response)
}
diff --git a/internal/server/cnpg_cluster_ha.go b/internal/server/cnpg_cluster_ha.go
new file mode 100644
index 0000000000..9c9b47e439
--- /dev/null
+++ b/internal/server/cnpg_cluster_ha.go
@@ -0,0 +1,40 @@
+package server
+
+import (
+ "log"
+ "net/http"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+ "k8s.io/client-go/metadata"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+)
+
+// GET /api/cnpg/clusters/{namespace}/{name}/ha: the facts that decide whether
+// a Cluster survives losing an instance, and what a planned switchover will
+// meet. Reading the Cluster never implies reading anything else: Nodes, Jobs,
+// PodDisruptionBudgets, Leases, EndpointSlices, Secret metadata and the
+// FailoverQuorum are each authorized on their own and report their own state.
+
+func (s *Server) handleCNPGClusterHA(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ cache, cluster, err := reader.Observations.Cluster(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ typed, dyn, cfg := reader.Clients.Typed, reader.dynamic, reader.config
+ if typed == nil || dyn == nil || cfg == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ meta, err := metadata.NewForConfig(cfg)
+ if err != nil {
+ log.Printf("[cnpg] Failed to build metadata client for %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available")
+ return
+ }
+ s.writeJSON(w, reader.ClusterHA(r.Context(), cnpgsvc.HAClients{Typed: typed, Dynamic: dyn, Metadata: meta}, cache, cluster, time.Now().UTC()))
+}
diff --git a/internal/server/cnpg_cluster_ha_test.go b/internal/server/cnpg_cluster_ha_test.go
new file mode 100644
index 0000000000..cb53a38037
--- /dev/null
+++ b/internal/server/cnpg_cluster_ha_test.go
@@ -0,0 +1,374 @@
+package server
+
+import (
+ "context"
+ "encoding/json"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "slices"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ batchv1 "k8s.io/api/batch/v1"
+ coordinationv1 "k8s.io/api/coordination/v1"
+ corev1 "k8s.io/api/core/v1"
+ discoveryv1 "k8s.io/api/discovery/v1"
+ policyv1 "k8s.io/api/policy/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/labels"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/apimachinery/pkg/util/intstr"
+ dynamicfake "k8s.io/client-go/dynamic/fake"
+ "k8s.io/client-go/kubernetes/fake"
+ metadatafake "k8s.io/client-go/metadata/fake"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+const cnpgHATestNS = "pgha"
+
+func cnpgHAOwner(uid string) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-ha", UID: types.UID(uid), Controller: boolPtr(true)}
+}
+
+func cnpgHACluster(now time.Time) *unstructured.Unstructured {
+ c := withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", cnpgHATestNS, "pg-ha", map[string]any{
+ "instances": int64(3),
+ "postgresql": map[string]any{"synchronous": map[string]any{
+ "method": "any", "number": int64(1), "failoverQuorum": true,
+ }},
+ "certificates": map[string]any{"serverTLSSecret": "pg-ha-user-tls", "serverCASecret": "pg-ha-user-ca"},
+ "nodeMaintenanceWindow": map[string]any{"inProgress": true},
+ }, map[string]any{
+ "image": "pg:17.2",
+ "currentPrimary": "pg-ha-1",
+ "certificates": map[string]any{
+ "serverTLSSecret": "pg-ha-user-tls",
+ "serverCASecret": "pg-ha-user-ca",
+ "replicationTLSSecret": "pg-ha-replication",
+ "clientCASecret": "pg-ha-ca",
+ "expirations": map[string]any{
+ "pg-ha-user-tls": now.Add(20 * 24 * time.Hour).Format(issues.CNPGCertExpiryLayout),
+ "pg-ha-user-ca": now.Add(300 * 24 * time.Hour).Format(issues.CNPGCertExpiryLayout),
+ "pg-ha-replication": "garbage",
+ "pg-ha-ca": now.Add(60 * 24 * time.Hour).Format(issues.CNPGCertExpiryLayout),
+ },
+ },
+ }), "ha-uid")
+ return c
+}
+
+func seedCNPGHAFixture(t *testing.T, now time.Time) *unstructured.Unstructured {
+ t.Helper()
+ kinds := append(append([]k8s.APIResource(nil), cnpgWorkspaceTestKinds...), cnpgTestResource(cnpgsvc.Group, "FailoverQuorum", "failoverquorums", true))
+ cluster := cnpgHACluster(now)
+ seedCNPGWorkspace(t, kinds, cluster)
+
+ owner := cnpgHAOwner("ha-uid")
+ p1 := cnpgPod(cnpgHATestNS, "pg-ha-1", "pg-ha", owner)
+ p1.Spec.NodeName = "ha-node-a"
+ p1.Spec.Containers[0].Image = "pg:17.2"
+ p2 := cnpgPod(cnpgHATestNS, "pg-ha-2", "pg-ha", owner)
+ p2.Labels["cnpg.io/instanceRole"] = "replica"
+ p2.Spec.NodeName = "ha-node-b"
+ p2.Spec.Containers[0].Image = "pg:17.1"
+ p1.Status.Conditions = []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionTrue}}
+ p2.Status.Conditions = []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionTrue}}
+ p3 := cnpgPod(cnpgHATestNS, "pg-ha-3", "pg-ha", owner)
+ p3.Labels["cnpg.io/instanceRole"] = "replica"
+ p3.Spec.NodeName = "ha-node-b"
+ p3.Status.Conditions = []corev1.PodCondition{{Type: corev1.PodReady, Status: corev1.ConditionFalse}}
+ seedCNPGPods(t, p1, p2, p3)
+
+ ctx := context.Background()
+ create := func(name string, fn func() error) {
+ if err := fn(); err != nil {
+ t.Fatalf("create %s: %v", name, err)
+ }
+ }
+ for _, n := range []struct{ name, zone string }{{"ha-node-a", "zone-a"}, {"ha-node-b", "zone-b"}} {
+ node := &corev1.Node{ObjectMeta: metav1.ObjectMeta{Name: n.name, Labels: map[string]string{"topology.kubernetes.io/zone": n.zone}}}
+ create(n.name, func() error {
+ _, err := testFakeClient.CoreV1().Nodes().Create(ctx, node, metav1.CreateOptions{})
+ return err
+ })
+ t.Cleanup(func() {
+ _ = testFakeClient.CoreV1().Nodes().Delete(context.Background(), node.Name, metav1.DeleteOptions{})
+ })
+ }
+ one := intstr.FromInt32(1)
+ pdbs := []*policyv1.PodDisruptionBudget{
+ {ObjectMeta: metav1.ObjectMeta{Name: "pg-ha", Namespace: cnpgHATestNS, Generation: 2, Labels: map[string]string{"cnpg.io/cluster": "pg-ha"}, OwnerReferences: []metav1.OwnerReference{owner}},
+ Spec: policyv1.PodDisruptionBudgetSpec{MinAvailable: &one},
+ Status: policyv1.PodDisruptionBudgetStatus{ObservedGeneration: 2, ExpectedPods: 2, CurrentHealthy: 1, DesiredHealthy: 1, DisruptionsAllowed: 0}},
+ {ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-primary", Namespace: cnpgHATestNS, Generation: 1, Labels: map[string]string{"cnpg.io/cluster": "pg-ha"}, OwnerReferences: []metav1.OwnerReference{owner}},
+ Spec: policyv1.PodDisruptionBudgetSpec{MinAvailable: &one},
+ Status: policyv1.PodDisruptionBudgetStatus{ObservedGeneration: 1, ExpectedPods: 1, CurrentHealthy: 1, DesiredHealthy: 1}},
+ {ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-foreign", Namespace: cnpgHATestNS, Labels: map[string]string{"cnpg.io/cluster": "pg-ha"}}},
+ }
+ for _, p := range pdbs {
+ create(p.Name, func() error {
+ _, err := testFakeClient.PolicyV1().PodDisruptionBudgets(cnpgHATestNS).Create(ctx, p, metav1.CreateOptions{})
+ return err
+ })
+ t.Cleanup(func() {
+ _ = testFakeClient.PolicyV1().PodDisruptionBudgets(cnpgHATestNS).Delete(context.Background(), p.Name, metav1.DeleteOptions{})
+ })
+ }
+ started := metav1.NewTime(now.Add(-time.Hour))
+ jobs := []*batchv1.Job{
+ {ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-1-initdb", Namespace: cnpgHATestNS, Labels: map[string]string{"cnpg.io/cluster": "pg-ha", "cnpg.io/jobRole": "initdb", "cnpg.io/instanceName": "pg-ha-1"}, OwnerReferences: []metav1.OwnerReference{owner}},
+ Status: batchv1.JobStatus{StartTime: &started, Conditions: []batchv1.JobCondition{{Type: batchv1.JobComplete, Status: corev1.ConditionTrue}}}},
+ {ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-3-join", Namespace: cnpgHATestNS, Labels: map[string]string{"cnpg.io/cluster": "pg-ha", "cnpg.io/jobRole": "join"}, OwnerReferences: []metav1.OwnerReference{owner}},
+ Status: batchv1.JobStatus{Conditions: []batchv1.JobCondition{{Type: batchv1.JobFailed, Status: corev1.ConditionTrue, Reason: "BackoffLimitExceeded", Message: "Job has reached the specified backoff limit"}}}},
+ }
+ for _, j := range jobs {
+ create(j.Name, func() error {
+ _, err := testFakeClient.BatchV1().Jobs(cnpgHATestNS).Create(ctx, j, metav1.CreateOptions{})
+ return err
+ })
+ t.Cleanup(func() {
+ _ = testFakeClient.BatchV1().Jobs(cnpgHATestNS).Delete(context.Background(), j.Name, metav1.DeleteOptions{})
+ })
+ }
+ cache := k8s.GetResourceCache()
+ waitFor(t, 5*time.Second, func() bool {
+ n, _ := cache.Nodes().Get("ha-node-b")
+ pd, _ := cache.PodDisruptionBudgets().PodDisruptionBudgets(cnpgHATestNS).List(labels.Everything())
+ js, _ := cache.Jobs().Jobs(cnpgHATestNS).List(labels.Everything())
+ return n != nil && len(pd) == 3 && len(js) == 2
+ })
+ return cluster
+}
+
+func cnpgHAClientsFor(t *testing.T, now time.Time) cnpgsvc.HAClients {
+ t.Helper()
+ renew := metav1.NewMicroTime(now.Add(-3 * time.Second))
+ stale := metav1.NewMicroTime(now.Add(-time.Minute))
+ fifteen := int32(15)
+ holder, opHolder := "pg-ha-1", "cnpg-controller-manager-abc_123"
+ owner := cnpgHAOwner("ha-uid")
+ typed := fake.NewClientset(
+ &coordinationv1.Lease{ObjectMeta: metav1.ObjectMeta{Name: "pg-ha", Namespace: cnpgHATestNS, OwnerReferences: []metav1.OwnerReference{owner}},
+ Spec: coordinationv1.LeaseSpec{HolderIdentity: &holder, RenewTime: &renew, LeaseDurationSeconds: &fifteen}},
+ &coordinationv1.Lease{ObjectMeta: metav1.ObjectMeta{Name: "db9c8771.cnpg.io", Namespace: "cnpg-system"},
+ Spec: coordinationv1.LeaseSpec{HolderIdentity: &opHolder, RenewTime: &stale, LeaseDurationSeconds: &fifteen}},
+ &discoveryv1.EndpointSlice{ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-rw-x", Namespace: cnpgHATestNS, Labels: map[string]string{discoveryv1.LabelServiceName: "pg-ha-rw"}},
+ Endpoints: []discoveryv1.Endpoint{
+ {TargetRef: &corev1.ObjectReference{Kind: "Pod", Name: "pg-ha-1"}, Conditions: discoveryv1.EndpointConditions{Ready: boolPtr(true)}},
+ {TargetRef: &corev1.ObjectReference{Kind: "Pod", Name: "pg-ha-2"}, Conditions: discoveryv1.EndpointConditions{Ready: boolPtr(false)}},
+ }},
+ )
+ fq := &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "FailoverQuorum",
+ "metadata": map[string]any{"name": "pg-ha", "namespace": cnpgHATestNS, "ownerReferences": []any{map[string]any{
+ "apiVersion": "postgresql.cnpg.io/v1", "kind": "Cluster", "name": "pg-ha", "uid": "ha-uid", "controller": true,
+ }}},
+ "status": map[string]any{"method": "ANY", "standbyNames": []any{"pg-ha-2", "pg-ha-3"}, "standbyNumber": int64(1), "primary": "pg-ha-1"},
+ }}
+ dyn := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(),
+ map[schema.GroupVersionResource]string{schema.GroupVersionResource{Group: cnpgsvc.Group, Version: "v1", Resource: "failoverquorums"}: "FailoverQuorumList"}, fq)
+ metaScheme := metadatafake.NewTestScheme()
+ metaScheme.AddKnownTypeWithName(corev1.SchemeGroupVersion.WithKind("Secret"), &metav1.PartialObjectMetadata{})
+ meta := metadatafake.NewSimpleMetadataClient(metaScheme,
+ &metav1.PartialObjectMetadata{
+ TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"},
+ ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-user-tls", Namespace: cnpgHATestNS, Annotations: map[string]string{"cert-manager.io/certificate-name": "pg-ha-server", "cert-manager.io/issuer-name": "internal-ca", "cert-manager.io/issuer-kind": "ClusterIssuer"}},
+ },
+ &metav1.PartialObjectMetadata{
+ TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"},
+ ObjectMeta: metav1.ObjectMeta{Name: "pg-ha-user-ca", Namespace: cnpgHATestNS},
+ },
+ )
+ return cnpgsvc.HAClients{Typed: typed, Dynamic: dyn, Metadata: meta}
+}
+
+func TestCNPGClusterHA_ReadsEveryFact(t *testing.T) {
+ now := time.Date(2026, 9, 30, 12, 0, 0, 0, time.UTC)
+ cluster := seedCNPGHAFixture(t, now)
+ if err := unstructured.SetNestedStringSlice(cluster.Object, []string{"pg-ha-1", "pg-ha-2", "pg-ha-3"}, "status", "instanceNames"); err != nil {
+ t.Fatal(err)
+ }
+ srv := &Server{}
+ r := httptest.NewRequest(http.MethodGet, "/", nil)
+ got := srv.cnpgReader(r).ClusterHA(r.Context(), cnpgHAClientsFor(t, now), k8s.GetResourceCache(), cluster, now)
+
+ if got.DesiredImage != "pg:17.2" || len(got.Instances) != 3 {
+ t.Fatalf("instances = %+v", got.Instances)
+ }
+ if got.DeclaredInstances == nil || *got.DeclaredInstances != 3 || !slices.Contains(got.ExpectedInstances, "pg-ha-1") {
+ t.Fatalf("declared/expected instances: %+v %+v", got.DeclaredInstances, got.ExpectedInstances)
+ }
+ inst := map[string]cnpgsvc.CNPGHAInstance{}
+ for _, i := range got.Instances {
+ inst[i.Pod] = i
+ }
+ if i := inst["pg-ha-2"]; i.Zone != "zone-b" || i.ImageMatches == nil || *i.ImageMatches || i.Role != "replica" {
+ t.Errorf("pg-ha-2 = %+v, want zone-b with image drift", i)
+ }
+ if i := inst["pg-ha-1"]; i.Zone != "zone-a" || i.ImageMatches == nil || !*i.ImageMatches || i.QOSClass != "" && i.QOSClass != "BestEffort" {
+ t.Errorf("pg-ha-1 = %+v", i)
+ }
+
+ q := got.Quorum
+ if !q.Enabled || q.EnabledBy != "spec" || q.Object.State != "ok" || q.Status == nil {
+ t.Fatalf("quorum = %+v", q)
+ }
+ // N=2 potentially synchronous, W=1, R=1 (pg-ha-3 is not ready): 1+1 > 2 is false.
+ if q.N == nil || *q.N != 2 || *q.W != 1 || q.R == nil || *q.R != 1 || q.Holds == nil || *q.Holds {
+ t.Errorf("quorum arithmetic = N%v W%v R%v holds %v", q.N, q.W, q.R, q.Holds)
+ }
+
+ if got.PDBs.State != "ok" || len(got.PDBs.Items) != 2 {
+ t.Fatalf("pdbs = %+v (the unowned budget must be ignored)", got.PDBs)
+ }
+ if p := got.PDBs.Items[0]; p.Name != "pg-ha" || p.Role != "replicas" || p.DisruptionsAllowed != 0 || !p.Observed || p.MinAvailable != "1" {
+ t.Errorf("replica pdb = %+v", p)
+ }
+
+ if l := got.PrimaryLease; l.State != "ok" || l.Holder != "pg-ha-1" || l.Expired == nil || *l.Expired || l.ControlledByCluster == nil || !*l.ControlledByCluster {
+ t.Errorf("primary lease = %+v", l)
+ }
+
+ if len(got.Jobs.Items) != 2 {
+ t.Fatalf("jobs = %+v", got.Jobs)
+ }
+ phases := map[string]string{}
+ for _, j := range got.Jobs.Items {
+ phases[j.Role] = j.Phase
+ }
+ if phases["initdb"] != "succeeded" || phases["join"] != "failed" {
+ t.Errorf("job phases = %v", phases)
+ }
+
+ if e := got.RWEndpoints; e.State != "ok" || len(e.Pods) != 1 || e.Pods[0] != "pg-ha-1" {
+ t.Errorf("rw endpoints = %+v", e)
+ }
+
+ certs := map[string]cnpgsvc.CNPGHACertificate{}
+ for _, c := range got.Certificates {
+ certs[c.Secret] = c
+ }
+ if c := certs["pg-ha-user-tls"]; c.Renewal != "user" || c.ExpiresAt == "" || c.CertManager == nil || c.CertManager.Certificate != "pg-ha-server" || c.Metadata == nil || c.Metadata.State != "ok" {
+ t.Errorf("user tls = %+v", c)
+ }
+ if c := certs["pg-ha-user-ca"]; c.Renewal != "user" || c.CertManager != nil || c.Metadata.State != "ok" {
+ t.Errorf("user ca without cert-manager = %+v", c)
+ }
+ if c := certs["pg-ha-replication"]; c.Renewal != "operator" || c.ExpiresAt != "" || c.Raw != "garbage" || c.Metadata != nil {
+ t.Errorf("unparseable operator cert = %+v", c)
+ }
+
+ if !got.Maintenance.Declared || !got.Maintenance.InProgress || !got.Maintenance.ReusePVC {
+ t.Errorf("maintenance = %+v, want in progress with reusePVC defaulting to true", got.Maintenance)
+ }
+}
+
+// Reading the Cluster never implies the rest: a caller who can only get the
+// Cluster and list Pods sees every other fact denied, named by its grant, and
+// Radar never asks the apiserver for what the caller cannot read.
+func TestCNPGClusterHA_EachReadIsAuthorizedOnItsOwn(t *testing.T) {
+ now := time.Now().UTC()
+ seedCNPGHAFixture(t, now)
+
+ var mu sync.Mutex
+ var calls []string
+ api := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ mu.Lock()
+ calls = append(calls, r.Method+" "+r.URL.Path)
+ mu.Unlock()
+ cnpgWriteAPIStatus(w, http.StatusForbidden, metav1.StatusReasonForbidden, "forbidden")
+ }))
+ t.Cleanup(api.Close)
+ previous := k8s.SetTestConfig(&rest.Config{Host: api.URL})
+ t.Cleanup(func() { k8s.SetTestConfig(previous) })
+
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{cnpgHATestNS}}
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", cnpgHATestNS, true)
+ perms.SetCanI("list", "", "pods", cnpgHATestNS, true)
+ for _, d := range []struct{ verb, group, resource, ns string }{
+ {"get", "", "nodes", ""},
+ {"get", cnpgsvc.Group, "failoverquorums", cnpgHATestNS},
+ {"list", "policy", "poddisruptionbudgets", cnpgHATestNS},
+ {"get", "coordination.k8s.io", "leases", cnpgHATestNS},
+ {"list", "batch", "jobs", cnpgHATestNS},
+ {"list", "discovery.k8s.io", "endpointslices", cnpgHATestNS},
+ {"get", "", "secrets", cnpgHATestNS},
+ {"list", "apps", "deployments", ""},
+ } {
+ perms.SetCanI(d.verb, d.group, d.resource, d.ns, false)
+ }
+ env.srv.permCache.Set("narrow", nil, perms)
+
+ resp := env.authGet(t, "/api/cnpg/clusters/"+cnpgHATestNS+"/pg-ha/ha", "narrow", "")
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ if resp.StatusCode != http.StatusOK {
+ t.Fatalf("status %d: %s", resp.StatusCode, body)
+ }
+ var got cnpgsvc.CNPGClusterHAResponse
+ if err := json.Unmarshal(body, &got); err != nil {
+ t.Fatal(err)
+ }
+ for name, src := range map[string]integration.ReadSource{
+ "nodes": got.Nodes, "quorum": got.Quorum.Object, "pdbs": got.PDBs.ReadSource,
+ "primaryLease": got.PrimaryLease.ReadSource, "jobs": got.Jobs.ReadSource, "rwEndpoints": got.RWEndpoints.ReadSource,
+ } {
+ if src.State != "denied" || src.Grant == nil {
+ t.Errorf("%s = %+v, want denied naming the grant", name, src)
+ }
+ }
+ if got.Pods.State != "ok" || len(got.Instances) != 3 || got.Instances[0].Zone != "" {
+ t.Errorf("pods readable, zones not: %+v", got.Instances)
+ }
+ for _, c := range got.Certificates {
+ if c.Renewal == "user" && (c.Metadata == nil || c.Metadata.State != "denied" || c.CertManager != nil) {
+ t.Errorf("%s: cert-manager link must be unknown without get secrets: %+v", c.Secret, c)
+ }
+ }
+ if len(got.PDBs.Items) != 0 || len(got.Jobs.Items) != 0 {
+ t.Errorf("denied lists must be empty: %+v %+v", got.PDBs.Items, got.Jobs.Items)
+ }
+ mu.Lock()
+ defer mu.Unlock()
+ for _, c := range calls {
+ if strings.Contains(c, "leases") || strings.Contains(c, "failoverquorums") || strings.Contains(c, "endpointslices") || strings.Contains(c, "secrets") {
+ t.Errorf("Radar asked the apiserver for something the caller cannot read: %s", c)
+ }
+ }
+}
+
+func TestCNPGClusterHA_ExpectedJobInstancesAreOwnedAndInProgress(t *testing.T) {
+ now := time.Now()
+ cluster := cnpgHACluster(now)
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cluster)
+ active := cnpgJob(cnpgHATestNS, "pg-ha-4-join", "active-job", cnpgHAOwner("ha-uid"))
+ active.Labels["cnpg.io/jobRole"], active.Labels["cnpg.io/instanceName"] = "join", "pg-ha-4"
+ active.Status.Active = 1
+ failed := active.DeepCopy()
+ failed.Name, failed.UID = "pg-ha-5-join", "failed-job"
+ failed.Labels["cnpg.io/instanceName"] = "pg-ha-5"
+ failed.Status.Conditions = []batchv1.JobCondition{{Type: batchv1.JobFailed, Status: corev1.ConditionTrue}}
+ foreign := active.DeepCopy()
+ foreign.Name, foreign.UID = "pg-ha-6-join", "foreign-job"
+ foreign.Labels["cnpg.io/instanceName"] = "pg-ha-6"
+ foreign.OwnerReferences[0].UID = "previous-cluster"
+ seedCNPGJobs(t, active, failed, foreign)
+ srv := &Server{}
+ got := srv.cnpgReader(httptest.NewRequest(http.MethodGet, "/", nil)).ClusterHA(context.Background(), cnpgHAClientsFor(t, now), k8s.GetResourceCache(), cluster, now)
+ if len(got.ExpectedInstances) != 1 || got.ExpectedInstances[0] != "pg-ha-4" {
+ t.Fatalf("expected instances must come from current in-progress owned Jobs: %v", got.ExpectedInstances)
+ }
+}
diff --git a/internal/server/cnpg_cluster_history_test.go b/internal/server/cnpg_cluster_history_test.go
index 6c5f23aee6..d1676fe879 100644
--- a/internal/server/cnpg_cluster_history_test.go
+++ b/internal/server/cnpg_cluster_history_test.go
@@ -11,12 +11,14 @@ import (
"testing"
"time"
+ corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/client-go/kubernetes"
"k8s.io/client-go/rest"
"github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/timeline"
pkgtimeline "github.com/skyhook-io/radar/pkg/timeline"
@@ -41,7 +43,7 @@ func seedCNPGLogCluster(t *testing.T, ns string) {
)
}
-func getCNPGLogs(t *testing.T, path string) (int, CNPGClusterLogsResponse, string) {
+func getCNPGLogs(t *testing.T, path string) (int, cnpgsvc.ClusterLogsResponse, string) {
t.Helper()
resp, err := http.Get(testServer.URL + path)
if err != nil {
@@ -49,7 +51,7 @@ func getCNPGLogs(t *testing.T, path string) (int, CNPGClusterLogsResponse, strin
}
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
- var out CNPGClusterLogsResponse
+ var out cnpgsvc.ClusterLogsResponse
if resp.StatusCode == http.StatusOK {
if err := json.Unmarshal(body, &out); err != nil {
t.Fatalf("decode: %v (%s)", err, body)
@@ -114,7 +116,7 @@ func TestCNPGClusterLogs_OnlyValidatedInstancesContribute(t *testing.T) {
t.Errorf("parsed entry = %+v", entry)
}
}
- if got.SourceLabels["pg-orders-1"] != "primary" || got.SourceLabels["pg-orders-2"] != "replica" {
+ if got.SourceLabels["pg-orders-1"] != "primary 1" || got.SourceLabels["pg-orders-2"] != "replica 2" {
t.Errorf("sourceLabels = %v", got.SourceLabels)
}
@@ -154,7 +156,7 @@ func TestCNPGClusterLogs_Authorization(t *testing.T) {
clusters bool
}{{"no-clusters", false}, {"no-logs", true}} {
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pglogsauth"}}
- perms.SetCanI("get", cnpgGroup, "clusters", "pglogsauth", u.clusters)
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", "pglogsauth", u.clusters)
allow(perms, "", "pods", "pglogsauth", true)
env.srv.permCache.Set(u.name, nil, perms)
}
@@ -163,6 +165,8 @@ func TestCNPGClusterLogs_Authorization(t *testing.T) {
{"no-clusters", "/api/cnpg/clusters/pglogsauth/missing/logs", "clusters.postgresql.cnpg.io"},
{"no-logs", "/api/cnpg/clusters/pglogsauth/pg-orders/logs", "get pods/log"},
{"no-logs", "/api/cnpg/clusters/pglogsauth/pg-orders/logs/stream", "get pods/log"},
+ {"no-logs", "/api/cnpg/clusters/pglogsauth/pg-orders/logs?container=all", "get pods/log"},
+ {"no-logs", "/api/cnpg/clusters/pglogsauth/pg-orders/logs/stream?container=all", "get pods/log"},
} {
resp := env.authGet(t, tc.path, tc.user, "")
body, _ := io.ReadAll(resp.Body)
@@ -173,63 +177,6 @@ func TestCNPGClusterLogs_Authorization(t *testing.T) {
}
}
-func TestAnnotateCNPGLogEntry(t *testing.T) {
- cases := []struct {
- name, content, level, logger, message string
- }{
- {
- name: "postgres record",
- content: `{"level":"info","ts":"2026-09-28T14:19:58.123Z","logger":"postgres","msg":"record","record":{"error_severity":"FATAL","message":"password authentication failed","log_time":"2026-09-28 14:19:58.123 UTC"}}`,
- level: "FATAL", logger: "postgres", message: "password authentication failed",
- },
- {
- name: "instance manager error",
- content: `{"level":"error","ts":"2026-09-28T14:19:58Z","logger":"barman-cloud-wal-archive","msg":"Error invoking barman-cloud-wal-archive","error":"exit status 4"}`,
- level: "ERROR", logger: "barman-cloud-wal-archive", message: "Error invoking barman-cloud-wal-archive: exit status 4",
- },
- {
- name: "structured error",
- content: `{"level":"error","msg":"failed","error":{"code":2}}`,
- level: "ERROR", message: `failed: {"code":2}`,
- },
- {name: "plain text", content: "LOG: database system is ready"},
- {name: "broken json", content: `{"level":"info"`},
- }
- for _, tc := range cases {
- t.Run(tc.name, func(t *testing.T) {
- entry := workloadLogEntry{Content: tc.content}
- annotateCNPGLogEntry(&entry)
- if entry.Level != tc.level || entry.Logger != tc.logger || entry.Message != tc.message || entry.Content != tc.content {
- t.Fatalf("got level=%q logger=%q message=%q content-changed=%v", entry.Level, entry.Logger, entry.Message, entry.Content != tc.content)
- }
- })
- }
- raw, _ := json.Marshal(workloadLogEntry{Pod: "p", Content: "x"})
- if strings.Contains(string(raw), "level") || strings.Contains(string(raw), "message") {
- t.Fatalf("unparsed entries grew fields: %s", raw)
- }
-}
-
-func TestParseCNPGLogQuery(t *testing.T) {
- now := time.Date(2026, 9, 28, 12, 0, 0, 0, time.UTC)
- req := httptest.NewRequest("GET", "/?sinceTime=2026-09-28T11:59:00.5Z", nil)
- if _, err := parseCNPGLogQuery(req, now); err != nil {
- t.Fatalf("fractional RFC3339 rejected: %v", err)
- }
- req = httptest.NewRequest("GET", "/?sinceTime=2026-09-28T11:58:30Z", nil)
- q, err := parseCNPGLogQuery(req, now)
- if err != nil || q.sinceSeconds == nil || *q.sinceSeconds != 90 || q.container != "postgres" || q.tailLines != 200 {
- t.Fatalf("query = %+v err=%v", q, err)
- }
- if q.keep(workloadLogEntry{Timestamp: "2026-09-28T11:58:29.9Z"}) || !q.keep(workloadLogEntry{Timestamp: "2026-09-28T11:58:30Z"}) {
- t.Fatal("sinceTime overlap not trimmed")
- }
- req = httptest.NewRequest("GET", "/?sinceTime=2026-09-28T11:58:30Z&sinceSeconds=5", nil)
- if _, err := parseCNPGLogQuery(req, now); err == nil {
- t.Fatal("sinceTime with sinceSeconds accepted")
- }
-}
-
func useMemoryTimeline(t *testing.T) timeline.EventStore {
t.Helper()
timeline.ResetStore()
@@ -299,21 +246,21 @@ func cnpgActivityFixture(t *testing.T, ns string) {
)
}
-func decodeActivity(t *testing.T, resp *http.Response) CNPGClusterActivityResponse {
+func decodeActivity(t *testing.T, resp *http.Response) cnpgsvc.CNPGClusterActivityResponse {
t.Helper()
defer resp.Body.Close()
body, _ := io.ReadAll(resp.Body)
if resp.StatusCode != http.StatusOK {
t.Fatalf("status = %d: %s", resp.StatusCode, body)
}
- var out CNPGClusterActivityResponse
+ var out cnpgsvc.CNPGClusterActivityResponse
if err := json.Unmarshal(body, &out); err != nil {
t.Fatalf("decode: %v", err)
}
return out
}
-func activityIDs(resp CNPGClusterActivityResponse) []string {
+func activityIDs(resp cnpgsvc.CNPGClusterActivityResponse) []string {
var out []string
for _, e := range resp.Events {
out = append(out, e.ID)
@@ -367,11 +314,11 @@ func TestCNPGClusterActivity_DropsKindsTheCallerCannotList(t *testing.T) {
eventsListed bool
}{{"reader", true, true, true}, {"no-backups", true, false, true}, {"no-clusters", false, true, true}, {"no-events", true, true, false}} {
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgactauth"}}
- perms.SetCanI("get", cnpgGroup, "clusters", "pgactauth", u.clusterGet)
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", "pgactauth", u.clusterGet)
allow(perms, "", "events", "pgactauth", u.eventsListed)
- allow(perms, cnpgGroup, "clusters", "pgactauth", true)
- allow(perms, cnpgGroup, "backups", "pgactauth", u.backupsListed)
- allow(perms, cnpgGroup, "poolers", "pgactauth", true)
+ allow(perms, cnpgsvc.Group, "clusters", "pgactauth", true)
+ allow(perms, cnpgsvc.Group, "backups", "pgactauth", u.backupsListed)
+ allow(perms, cnpgsvc.Group, "poolers", "pgactauth", true)
allow(perms, "", "pods", "pgactauth", true)
env.srv.permCache.Set(u.name, nil, perms)
}
@@ -442,7 +389,7 @@ func TestCNPGClusterLogsStream_SendsParsedInstanceLines(t *testing.T) {
if !sawConnected || strings.Contains(stream, `"name":"pg-orders-1"`) {
t.Fatalf("connected event wrong:\n%s", stream)
}
- for _, want := range []string{`"level":"LOG"`, `"message":"hello from pg-orders-2"`, `"sourceLabel":"replica"`} {
+ for _, want := range []string{`"level":"LOG"`, `"message":"hello from pg-orders-2"`, `"sourceLabel":"replica 2"`} {
if !strings.Contains(stream, want) {
t.Errorf("stream missing %s:\n%s", want, stream)
}
@@ -486,40 +433,105 @@ func TestCNPGClusterActivity_RecreatedClusterExcludesPreviousIncarnationPods(t *
}
}
-func TestCNPGStreamCursorResumesWithoutReplay(t *testing.T) {
- var c cnpgStreamCursor
- first := c.restartOptions("postgres", 200, nil)
- if first.TailLines == nil || *first.TailLines != 200 || first.SinceTime != nil || !first.Follow || !first.Timestamps || first.Container != "postgres" {
- t.Fatalf("first start = %+v", first)
+func cnpgRestartedStatus(currentStart, prevStart, prevEnd time.Time, restarts int32) corev1.ContainerStatus {
+ status := corev1.ContainerStatus{
+ Name: "postgres", RestartCount: restarts,
+ State: corev1.ContainerState{Running: &corev1.ContainerStateRunning{StartedAt: metav1.NewTime(currentStart)}},
}
- line := func(ts, content string) workloadLogEntry {
- return workloadLogEntry{Timestamp: ts, Content: content}
+ if restarts > 0 {
+ status.LastTerminationState.Terminated = &corev1.ContainerStateTerminated{StartedAt: metav1.NewTime(prevStart), FinishedAt: metav1.NewTime(prevEnd)}
}
- for _, e := range []workloadLogEntry{line("2026-09-28T14:00:00.1Z", "a"), line("2026-09-28T14:00:05.7Z", "b"), line("2026-09-28T14:00:05.7Z", "c")} {
- if !c.admit(e) {
- t.Fatalf("fresh line %+v rejected", e)
+ return status
+}
+
+// An incident followed by restarts: the interval lives in the previous run,
+// and the current run's lines (all after the interval) fill the byte cap.
+func TestCNPGClusterLogs_IntervalReadsThePreviousRun(t *testing.T) {
+ ns := "pglogiv"
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders", map[string]any{"instances": int64(1)}, nil), "iv-uid"),
+ )
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "iv-uid", Controller: boolPtr(true)}
+ at := func(m int) time.Time { return time.Date(2026, 9, 29, 23, m, 0, 0, time.UTC) }
+ p := cnpgPod(ns, "pg-orders-1", "pg-orders", owner)
+ p.Status.ContainerStatuses = []corev1.ContainerStatus{cnpgRestartedStatus(at(45), at(10), at(40), 1)}
+ seedCNPGPods(t, p)
+
+ apiserver := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if r.URL.Query().Get("previous") == "true" {
+ fmt.Fprintln(w, "2026-09-29T23:25:00.5Z {\"level\":\"error\",\"msg\":\"during the incident\"}")
+ return
+ }
+ for i := 0; i < 2000; i++ {
+ fmt.Fprintf(w, "2026-09-29T23:50:%02d.5Z {\"level\":\"info\",\"msg\":\"after the restart %d\"}\n", i%60, i)
}
+ }))
+ t.Cleanup(apiserver.Close)
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: apiserver.URL})
+ if err != nil {
+ t.Fatal(err)
}
+ previous := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previous) })
- restart := c.restartOptions("postgres", 200, nil)
- if restart.TailLines != nil || restart.SinceSeconds != nil || restart.SinceTime == nil ||
- !restart.SinceTime.Time.Equal(time.Date(2026, 9, 28, 14, 0, 5, 0, time.UTC)) {
- t.Fatalf("restart = %+v, want sinceTime at the last delivered second and no tail", restart)
+ status, got, body := getCNPGLogs(t, "/api/cnpg/clusters/"+ns+"/pg-orders/logs?sinceTime=2026-09-29T23:20:00Z&untilTime=2026-09-29T23:30:00Z")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
}
+ if len(got.Logs) != 1 || got.Logs[0].Message != "during the incident" || !got.Logs[0].Previous || got.Logs[0].SourceLabel != "primary 1 · previous run" {
+ t.Fatalf("logs = %+v, want the previous run's interval line", got.Logs)
+ }
+ if strings.Contains(got.Notice, "snapshot limit") {
+ t.Errorf("notice = %q: the clip came after the interval and cut nothing from it", got.Notice)
+ }
+ if !strings.Contains(got.EmptyMessage, "interval") {
+ t.Errorf("emptyMessage = %q, want interval-specific copy", got.EmptyMessage)
+ }
+}
- // The resumed follow replays the boundary second.
- replayed := []workloadLogEntry{line("2026-09-28T14:00:05.2Z", "earlier in the second"), line("2026-09-28T14:00:05.7Z", "b"), line("2026-09-28T14:00:05.7Z", "c")}
- for _, e := range replayed {
- if c.admit(e) {
- t.Errorf("replayed line %+v admitted", e)
+func TestCNPGClusterActivity_JobPodsNeedVerifiedOwnershipAndJobAccess(t *testing.T) {
+ ns := "pgactivityjobs"
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "analytics", map[string]any{"instances": int64(1)}, nil), "analytics-uid"))
+ seedCNPGJobs(t, cnpgJob(ns, "analytics-1-initdb", "job-uid", clusterRef("analytics", "analytics-uid")))
+ pod := cnpgJobPod(ns, "analytics-1-initdb-abc", "analytics", jobRef("analytics-1-initdb", "job-uid"))
+ pod.UID = "job-pod-uid"
+ stale := cnpgJobPod(ns, "analytics-1-initdb-old", "analytics", jobRef("analytics-1-initdb", "old-job-uid"))
+ stale.UID = "old-pod-uid"
+ seedCNPGPods(t, pod, stale)
+ store := useMemoryTimeline(t)
+ seedActivity(t, store, ns,
+ activityRow{id: "job-pod-pending", apiVersion: "v1", kind: "Pod", name: pod.Name, uid: string(pod.UID), age: time.Minute, owner: &timeline.OwnerInfo{APIVersion: "batch/v1", Kind: "Job", Name: "analytics-1-initdb", UID: "job-uid"}},
+ activityRow{id: "stale-job-pod", apiVersion: "v1", kind: "Pod", name: stale.Name, uid: string(stale.UID), age: time.Minute, owner: &timeline.OwnerInfo{APIVersion: "batch/v1", Kind: "Job", Name: "analytics-1-initdb", UID: "old-job-uid"}},
+ )
+ env := newAuthTestServer(t)
+ for _, jobs := range []bool{true, false} {
+ user := fmt.Sprintf("jobs-%v", jobs)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{ns}}
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", ns, true)
+ allow(perms, "", "pods", ns, true)
+ allow(perms, "batch", "jobs", ns, jobs)
+ env.srv.permCache.Set(user, nil, perms)
+ got := decodeActivity(t, env.authGet(t, "/api/cnpg/clusters/"+ns+"/analytics/activity", user, ""))
+ if ids := strings.Join(activityIDs(got), ","); jobs && ids != "job-pod-pending" || !jobs && ids != "" {
+ t.Fatalf("Job access %v: %s", jobs, ids)
}
}
- for _, e := range []workloadLogEntry{line("2026-09-28T14:00:05.7Z", "d"), line("2026-09-28T14:00:06Z", "e")} {
- if !c.admit(e) {
- t.Errorf("new line %+v rejected", e)
- }
+}
+func TestCNPGClusterActivity_AttributionBoundaryExcludesInstancePods(t *testing.T) {
+ store := useMemoryTimeline(t)
+ ns := "pgactivityboundary"
+ owner := &timeline.OwnerInfo{Kind: "Cluster", Name: "payments", APIVersion: "postgresql.cnpg.io/v1", UID: "payments-uid"}
+ labels := map[string]string{pkgtimeline.CNPGClusterLabel: "payments"}
+ seedActivity(t, store, ns,
+ activityRow{id: "pod-old", apiVersion: "v1", kind: "Pod", name: "payments-1", uid: "p1", age: 5 * time.Hour, owner: owner, labels: labels},
+ activityRow{id: "backup-new", apiVersion: "postgresql.cnpg.io/v1", kind: "Backup", name: "payments-backup", uid: "b1", age: time.Hour, labels: labels},
+ )
+ resp, err := http.Get(testServer.URL + "/api/cnpg/clusters/" + ns + "/payments/activity")
+ if err != nil {
+ t.Fatal(err)
}
- if c.admit(line("2026-09-28T14:00:05.7Z", "d")) {
- t.Error("line before the new last timestamp admitted")
+ got := decodeActivity(t, resp)
+ if got.AttributionSince == nil || time.Since(*got.AttributionSince) > 2*time.Hour {
+ t.Fatalf("boundary must describe CNPG children: %+v", got.AttributionSince)
}
}
diff --git a/internal/server/cnpg_cluster_logs.go b/internal/server/cnpg_cluster_logs.go
index 19a6436a3c..a1fc7a6b7c 100644
--- a/internal/server/cnpg_cluster_logs.go
+++ b/internal/server/cnpg_cluster_logs.go
@@ -1,97 +1,16 @@
package server
import (
- "bufio"
- "context"
- "encoding/json"
- "errors"
- "fmt"
- "io"
"log"
- "math"
"net/http"
- "sort"
- "strings"
- "sync"
"time"
"github.com/go-chi/chi/v5"
- corev1 "k8s.io/api/core/v1"
- metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
- "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "k8s.io/apimachinery/pkg/labels"
- "k8s.io/apimachinery/pkg/types"
- "k8s.io/client-go/kubernetes"
- "github.com/skyhook-io/radar/internal/k8s"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
)
-const (
- cnpgClusterLabel = "cnpg.io/cluster"
- cnpgDefaultLogContainer = "postgres"
- cnpgDefaultLogTailLines = 200
- cnpgLogDiscoveryInterval = 5 * time.Second
- cnpgLogsEmptyMessage = "No readable logs from this cluster's instances in this snapshot. Refresh after the instances start."
- cnpgLogsNoInstanceMessage = "This cluster has no instance Pods yet."
-)
-
-// CNPGClusterLogsResponse is GET /api/cnpg/clusters/{namespace}/{name}/logs.
-// Pods and SourceLabels list only the instances that contributed a source to
-// this snapshot; SourceLabels maps a Pod to its instance role.
-type CNPGClusterLogsResponse struct {
- UID types.UID `json:"uid"`
- Pods []WorkloadPodInfo `json:"pods"`
- Logs []workloadLogEntry `json:"logs"`
- Notice string `json:"notice"`
- SourceLabels map[string]string `json:"sourceLabels,omitempty"`
- CapturedAt string `json:"capturedAt"`
- EmptyMessage string `json:"emptyMessage"`
-}
-
-type cnpgLogQuery struct {
- container string
- tailLines int64
- sinceSeconds *int64
- sinceTime time.Time
- pod string
-}
-
-func parseCNPGLogQuery(r *http.Request, now time.Time) (cnpgLogQuery, error) {
- q := r.URL.Query()
- out := cnpgLogQuery{
- container: q.Get("container"),
- tailLines: parseTailLines(q.Get("tailLines"), cnpgDefaultLogTailLines),
- sinceSeconds: parseSinceSeconds(q.Get("sinceSeconds")),
- pod: q.Get("pod"),
- }
- if out.container == "" {
- out.container = cnpgDefaultLogContainer
- }
- if raw := q.Get("sinceTime"); raw != "" {
- if q.Get("sinceSeconds") != "" {
- return out, errors.New("sinceSeconds and sinceTime are mutually exclusive")
- }
- t, err := time.Parse(time.RFC3339, raw)
- if err != nil {
- return out, fmt.Errorf("invalid sinceTime %q (expected RFC3339)", raw)
- }
- out.sinceTime = t
- // The pod log API takes whole seconds; round up and trim the overlap
- // from the entries afterwards.
- secs := max(int64(math.Ceil(now.Sub(t).Seconds())), 1)
- out.sinceSeconds = &secs
- }
- return out, nil
-}
-
-func (q cnpgLogQuery) keep(entry workloadLogEntry) bool {
- if q.sinceTime.IsZero() || entry.Timestamp == "" {
- return true
- }
- ts, err := time.Parse(time.RFC3339Nano, entry.Timestamp)
- return err != nil || !ts.Before(q.sinceTime)
-}
-
// authorizeCNPGClusterLogs gates on reading the Cluster, listing its Pods and
// reading their logs — before the Cluster is looked up, so a denied caller
// cannot probe which Clusters exist.
@@ -99,11 +18,11 @@ func (s *Server) authorizeCNPGClusterLogs(w http.ResponseWriter, r *http.Request
if !s.requireConnected(w) {
return false
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return false
}
- if !s.canRead(r, cnpgGroup, "clusters", namespace, "get") {
+ if !s.canRead(r, cnpgsvc.Group, "clusters", namespace, "get") {
s.writeError(w, http.StatusForbidden, "no access to clusters.postgresql.cnpg.io in namespace "+namespace)
return false
}
@@ -114,454 +33,58 @@ func (s *Server) authorizeCNPGClusterLogs(w http.ResponseWriter, r *http.Request
return s.authorizePodLogRead(w, r, namespace)
}
-// loadCNPGCluster reads one CNPG Cluster from the dynamic cache. The error is
-// already written when ok is false.
-func (s *Server) loadCNPGCluster(w http.ResponseWriter, r *http.Request, cache *k8s.ResourceCache, namespace, name string) (*unstructured.Unstructured, bool) {
- cluster, err := findCNPGCluster(r.Context(), cache, namespace, name)
- switch {
- case err == nil && cluster != nil:
- return cluster, true
- case err == nil, errors.Is(err, k8s.ErrUnknownDynamicKind):
- s.writeError(w, http.StatusNotFound, "CloudNativePG Cluster "+namespace+"/"+name+" not found")
- case errors.Is(err, errDynamicNotSynced):
- s.writeError(w, http.StatusServiceUnavailable, "CloudNativePG Clusters are still syncing")
- default:
- log.Printf("[cnpg] Failed to read Cluster %s/%s: %v", namespace, name, err)
- s.writeError(w, http.StatusInternalServerError, "failed to read CloudNativePG Cluster")
- }
- return nil, false
-}
-
-func findCNPGCluster(ctx context.Context, cache *k8s.ResourceCache, namespace, name string) (*unstructured.Unstructured, error) {
- clusters, err := filterCNPGGroup(listDynamicSynced(ctx, cache, "Cluster", cnpgGroup, namespace))
- if err != nil {
- return nil, err
- }
- for _, c := range clusters {
- if c.GetNamespace() == namespace && c.GetName() == name && c.GroupVersionKind().Group == cnpgGroup {
- return c, nil
- }
- }
- return nil, nil
-}
-
-// cnpgClusterInstancePods returns the Cluster's instance Pods under the same
-// label-and-controller-UID rule the workspace uses, sorted by name.
-func cnpgClusterInstancePods(cache *k8s.ResourceCache, cluster *unstructured.Unstructured) ([]*corev1.Pod, error) {
- lister := cache.Pods()
- if lister == nil {
- return nil, errors.New("pod cache unavailable")
- }
- namespace, name := cluster.GetNamespace(), cluster.GetName()
- candidates, err := lister.Pods(namespace).List(labels.SelectorFromSet(labels.Set{cnpgClusterLabel: name}))
- if err != nil {
- return nil, err
- }
- uids := map[string]types.UID{namespace + "/" + name: cluster.GetUID()}
- pods := make([]*corev1.Pod, 0, len(candidates))
- for _, p := range candidates {
- if p != nil && isCNPGInstancePod(p, uids) {
- pods = append(pods, p)
- }
- }
- sort.Slice(pods, func(i, j int) bool { return pods[i].Name < pods[j].Name })
- return pods, nil
-}
-
-func cnpgInstanceRole(p *corev1.Pod) string {
- if role := p.Labels["cnpg.io/instanceRole"]; role != "" {
- return role
- }
- return p.Labels["role"]
-}
-
-// selectCNPGLogPods narrows to the requested instance. ok is false when the
-// requested Pod is not one of the Cluster's instances.
-func selectCNPGLogPods(pods []*corev1.Pod, want string) ([]*corev1.Pod, bool) {
- if want == "" {
- return pods, true
- }
- for _, p := range pods {
- if p.Name == want {
- return []*corev1.Pod{p}, true
- }
- }
- return nil, false
-}
-
-// cnpgLogRecord is the subset of a CloudNativePG instance-manager JSON log
-// line the viewer surfaces. PostgreSQL's own log lines arrive wrapped, with the
-// server's severity and message under record.
-type cnpgLogRecord struct {
- Level string `json:"level"`
- Logger string `json:"logger"`
- Msg string `json:"msg"`
- Error any `json:"error"`
- Record *struct {
- ErrorSeverity string `json:"error_severity"`
- Message string `json:"message"`
- } `json:"record"`
-}
-
-// annotateCNPGLogEntry fills the parsed fields of a CloudNativePG log line and
-// leaves anything that is not one untouched.
-func annotateCNPGLogEntry(entry *workloadLogEntry) {
- content := strings.TrimSpace(entry.Content)
- if !strings.HasPrefix(content, "{") {
- return
- }
- var rec cnpgLogRecord
- if err := json.Unmarshal([]byte(content), &rec); err != nil {
- return
- }
- level := strings.ToUpper(rec.Level)
- message := rec.Msg
- if rec.Record != nil {
- if rec.Record.ErrorSeverity != "" {
- level = rec.Record.ErrorSeverity
- }
- if rec.Record.Message != "" {
- message = rec.Record.Message
- }
- }
- if errText := cnpgLogErrorText(rec.Error); errText != "" {
- if message == "" {
- message = errText
- } else {
- message += ": " + errText
- }
- }
- entry.Level, entry.Logger, entry.Message = level, rec.Logger, message
-}
-
-func cnpgLogErrorText(v any) string {
- switch e := v.(type) {
- case nil:
- return ""
- case string:
- return e
- default:
- b, err := json.Marshal(e)
- if err != nil {
- return ""
- }
- return string(b)
- }
-}
-
-// handleCNPGClusterLogs serves GET /api/cnpg/clusters/{namespace}/{name}/logs:
-// a bounded snapshot of every instance Pod's logs, merged by timestamp.
func (s *Server) handleCNPGClusterLogs(w http.ResponseWriter, r *http.Request) {
namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
if !s.authorizeCNPGClusterLogs(w, r, namespace) {
return
}
- query, qerr := parseCNPGLogQuery(r, time.Now())
- if qerr != nil {
- s.writeError(w, http.StatusBadRequest, qerr.Error())
- return
- }
- cache := k8s.GetResourceCache()
- if cache == nil {
- s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
- return
- }
- cluster, ok := s.loadCNPGCluster(w, r, cache, namespace, name)
- if !ok {
- return
- }
- instances, err := cnpgClusterInstancePods(cache, cluster)
+ query, err := cnpgsvc.ParseLogQuery(r.URL.Query(), time.Now())
if err != nil {
- log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", namespace, name, err)
- s.writeError(w, http.StatusServiceUnavailable, "instance Pods unavailable: "+err.Error())
- return
- }
- pods, ok := selectCNPGLogPods(instances, query.pod)
- if !ok {
- s.writeError(w, http.StatusBadRequest, "pod "+query.pod+" is not an instance of CloudNativePG Cluster "+namespace+"/"+name)
+ s.writeError(w, http.StatusBadRequest, err.Error())
return
}
-
- resp := CNPGClusterLogsResponse{
- UID: cluster.GetUID(),
- Pods: []WorkloadPodInfo{},
- Logs: []workloadLogEntry{},
- CapturedAt: time.Now().UTC().Format(time.RFC3339),
- EmptyMessage: cnpgLogsEmptyMessage,
- }
- if len(pods) == 0 {
- resp.EmptyMessage = cnpgLogsNoInstanceMessage
- s.writeJSON(w, resp)
+ target, err := s.cnpgReader(r).PrepareLogs(r.Context(), namespace, name, query)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
return
}
- client := s.getClientForRequest(r)
- if client == nil {
- s.writeError(w, http.StatusServiceUnavailable, "cluster client unavailable")
+ response, err := target.Snapshot(r.Context(), query)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
return
}
-
- snapshot := collectLogsFromPods(r.Context(), client, namespace, pods, query.container, query.tailLines, query.sinceSeconds, true)
- shown := []*corev1.Pod{}
- sourceLabels := map[string]string{}
- for _, p := range pods {
- if !snapshot.SourcePods[p.Name] {
- continue
- }
- shown = append(shown, p)
- if role := cnpgInstanceRole(p); role != "" {
- sourceLabels[p.Name] = role
- }
- }
- for _, entry := range snapshot.Logs {
- if !query.keep(entry) {
- continue
- }
- entry.SourceLabel = sourceLabels[entry.Pod]
- annotateCNPGLogEntry(&entry)
- resp.Logs = append(resp.Logs, entry)
- }
- sortLogsByTimestamp(resp.Logs)
- resp.Pods = buildPodInfos(shown)
- resp.Notice = snapshot.Notice
- if len(sourceLabels) > 0 {
- resp.SourceLabels = sourceLabels
- }
- s.writeJSON(w, resp)
+ s.writeJSON(w, response)
}
-// handleCNPGClusterLogsStream serves GET
-// /api/cnpg/clusters/{namespace}/{name}/logs/stream: an SSE follow of every
-// instance Pod, re-resolving instances as the Cluster fails over or scales.
-// Events: connected {cluster, namespace, uid, pods}, log (a log entry with
-// the parsed fields), pod_added {pods}, pod_removed {pod, reason}, end
-// {reason}, error {error}.
+// The service emits log events; only this adapter frames them as browser SSE.
func (s *Server) handleCNPGClusterLogsStream(w http.ResponseWriter, r *http.Request) {
namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
if !s.authorizeCNPGClusterLogs(w, r, namespace) {
return
}
- query, qerr := parseCNPGLogQuery(r, time.Now())
- if qerr != nil {
- s.writeError(w, http.StatusBadRequest, qerr.Error())
- return
- }
- cache := k8s.GetResourceCache()
- if cache == nil {
- s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
- return
- }
- cluster, ok := s.loadCNPGCluster(w, r, cache, namespace, name)
- if !ok {
- return
- }
- instances, err := cnpgClusterInstancePods(cache, cluster)
+ query, err := cnpgsvc.ParseLogQuery(r.URL.Query(), time.Now())
if err != nil {
- log.Printf("[cnpg] Failed to list instance Pods for %s/%s: %v", namespace, name, err)
- s.writeError(w, http.StatusServiceUnavailable, "instance Pods unavailable: "+err.Error())
+ s.writeError(w, http.StatusBadRequest, err.Error())
return
}
- pods, ok := selectCNPGLogPods(instances, query.pod)
- if !ok {
- s.writeError(w, http.StatusBadRequest, "pod "+query.pod+" is not an instance of CloudNativePG Cluster "+namespace+"/"+name)
+ target, err := s.cnpgReader(r).PrepareLogs(r.Context(), namespace, name, query)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
return
}
- client := s.getClientForRequest(r)
- if client == nil {
- s.writeError(w, http.StatusServiceUnavailable, "cluster client unavailable")
+ if err := target.RequireClient(); err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
return
}
flusher, ok := w.(http.Flusher)
if !ok {
- log.Printf("[cnpg] Failed to stream logs for %s/%s: response writer does not support flushing", namespace, name)
+ log.Printf("[cnpg] Failed to stream logs for %s/%s: response writer does not support flushing", sanitizeForLog(namespace), sanitizeForLog(name))
s.writeError(w, http.StatusInternalServerError, "streaming not supported")
return
}
-
w.Header().Set("Content-Type", "text/event-stream")
w.Header().Set("Cache-Control", "no-cache")
w.Header().Set("Connection", "keep-alive")
w.Header().Set("X-Accel-Buffering", "no")
-
- uid := cluster.GetUID()
- sendSSEEvent(w, flusher, "connected", map[string]any{
- "cluster": name, "namespace": namespace, "uid": uid, "pods": buildPodInfos(pods),
- })
-
- ctx, cancel := context.WithCancel(r.Context())
- defer cancel()
- logCh := make(chan workloadLogEntry, 1000)
- var active sync.Map
- roles := map[string]string{}
- cursors := map[string]*cnpgStreamCursor{}
- start := func(pods []*corev1.Pod) {
- for _, pod := range pods {
- roles[pod.Name] = cnpgInstanceRole(pod)
- for _, c := range k8s.GetContainersForPod(pod, query.container, true) {
- key := pod.Name + "/" + c
- if _, exists := active.Load(key); exists {
- continue
- }
- cursor := cursors[key]
- if cursor == nil {
- cursor = &cnpgStreamCursor{}
- cursors[key] = cursor
- }
- opts := cursor.restartOptions(c, query.tailLines, query.sinceSeconds)
- streamCtx, streamCancel := context.WithCancel(ctx)
- handle := &cnpgStreamHandle{cancel: streamCancel}
- active.Store(key, handle)
- go func(podName, key string) {
- defer active.CompareAndDelete(key, handle)
- followCNPGContainerLogs(streamCtx, client, namespace, podName, opts, logCh)
- }(pod.Name, key)
- }
- }
- }
- start(pods)
-
- known := map[string]bool{}
- for _, p := range pods {
- known[p.Name] = true
- }
- ticker := time.NewTicker(cnpgLogDiscoveryInterval)
- defer ticker.Stop()
- for {
- select {
- case <-ctx.Done():
- return
- case entry := <-logCh:
- if cursor := cursors[entry.Pod+"/"+entry.Container]; cursor != nil && !cursor.admit(entry) {
- continue
- }
- if !query.keep(entry) {
- continue
- }
- entry.SourceLabel = roles[entry.Pod]
- annotateCNPGLogEntry(&entry)
- sendSSEEvent(w, flusher, "log", entry)
- case <-ticker.C:
- current, err := findCNPGCluster(ctx, cache, namespace, name)
- if err != nil {
- continue
- }
- if current == nil || current.GetUID() != uid {
- sendSSEEvent(w, flusher, "end", map[string]string{"reason": "cluster deleted"})
- return
- }
- all, err := cnpgClusterInstancePods(cache, current)
- if err != nil {
- continue
- }
- currentPods, _ := selectCNPGLogPods(all, query.pod)
- present := map[string]bool{}
- for _, p := range currentPods {
- present[p.Name] = true
- if !known[p.Name] {
- known[p.Name] = true
- sendSSEEvent(w, flusher, "pod_added", map[string]any{"pods": []WorkloadPodInfo{buildPodInfo(p, time.Now())}})
- }
- }
- for podName := range known {
- if present[podName] {
- continue
- }
- delete(known, podName)
- active.Range(func(key, value any) bool {
- if strings.HasPrefix(key.(string), podName+"/") {
- value.(*cnpgStreamHandle).cancel()
- active.Delete(key)
- }
- return true
- })
- for key := range cursors {
- if strings.HasPrefix(key, podName+"/") {
- delete(cursors, key)
- }
- }
- sendSSEEvent(w, flusher, "pod_removed", map[string]string{"pod": podName, "reason": "terminated"})
- }
- start(currentPods)
- }
- }
-}
-
-type cnpgStreamHandle struct {
- cancel context.CancelFunc
-}
-
-// cnpgStreamCursor remembers where one container's follow left off, so a
-// stream that ends while its Pod is still an instance resumes instead of
-// replaying lines the client already has. Only the stream loop touches it.
-type cnpgStreamCursor struct {
- last time.Time
- // atLast holds the contents delivered with timestamp == last. The pod log
- // API's sinceTime is second-granular, so a resume replays that second and
- // only (timestamp, content) tells a replay from a new line.
- atLast map[string]bool
-}
-
-// restartOptions returns the follow request for the next (re)start: the
-// caller's window the first time, and from the last delivered second after.
-func (c *cnpgStreamCursor) restartOptions(container string, tailLines int64, sinceSeconds *int64) corev1.PodLogOptions {
- opts := corev1.PodLogOptions{Container: container, Timestamps: true, Follow: true}
- if c.last.IsZero() {
- opts.TailLines = &tailLines
- opts.SinceSeconds = sinceSeconds
- return opts
- }
- since := metav1.NewTime(c.last.Truncate(time.Second))
- opts.SinceTime = &since
- return opts
-}
-
-// admit reports whether an entry is new, recording it when it is. Lines
-// arrive in order per container, so anything before the last delivered
-// timestamp was already sent.
-func (c *cnpgStreamCursor) admit(entry workloadLogEntry) bool {
- ts, err := time.Parse(time.RFC3339Nano, entry.Timestamp)
- if err != nil {
- return true
- }
- switch {
- case ts.Before(c.last):
- return false
- case ts.Equal(c.last):
- if c.atLast[entry.Content] {
- return false
- }
- default:
- c.last = ts
- c.atLast = map[string]bool{}
- }
- c.atLast[entry.Content] = true
- return true
-}
-
-func followCNPGContainerLogs(ctx context.Context, client kubernetes.Interface, namespace, podName string, opts corev1.PodLogOptions, logCh chan<- workloadLogEntry) {
- stream, err := client.CoreV1().Pods(namespace).GetLogs(podName, &opts).Stream(ctx)
- if err != nil {
- if ctx.Err() == nil {
- log.Printf("[cnpg] Failed to follow logs for %s/%s/%s: %v", namespace, podName, opts.Container, err)
- }
- return
- }
- defer stream.Close()
- reader := bufio.NewReader(stream)
- for {
- line, err := reader.ReadString('\n')
- if line = strings.TrimSuffix(line, "\n"); line != "" && (err == nil || err == io.EOF) {
- ts, content := parseLogLine(line)
- select {
- case logCh <- workloadLogEntry{Pod: podName, Container: opts.Container, Timestamp: ts, Content: content}:
- case <-ctx.Done():
- return
- }
- }
- if err != nil {
- if err != io.EOF && ctx.Err() == nil {
- log.Printf("[cnpg] Failed to read logs for %s/%s/%s: %v", namespace, podName, opts.Container, err)
- }
- return
- }
- }
+ target.Follow(r.Context(), query, func(event string, data any) { sendSSEEvent(w, flusher, event, data) })
}
diff --git a/internal/server/cnpg_cluster_logs_test.go b/internal/server/cnpg_cluster_logs_test.go
new file mode 100644
index 0000000000..15ce4db54c
--- /dev/null
+++ b/internal/server/cnpg_cluster_logs_test.go
@@ -0,0 +1,252 @@
+package server
+
+import (
+ "bufio"
+ "context"
+ "fmt"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func TestCNPGLogsAllContainers(t *testing.T) {
+ ns := "pglogsall"
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders", map[string]any{"instances": int64(1)}, nil), "orders-uid"))
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "orders-uid", Controller: boolPtr(true)}
+ pod := cnpgPod(ns, "pg-orders-1", "pg-orders", owner)
+ pod.Spec.Containers = []corev1.Container{{Name: "postgres"}, {Name: "metrics"}}
+ pod.Spec.InitContainers = []corev1.Container{{Name: "plugin-barman-cloud"}, {Name: "not-started"}}
+ started := metav1.NewTime(time.Date(2026, 9, 28, 14, 0, 0, 0, time.UTC))
+ running := corev1.ContainerState{Running: &corev1.ContainerStateRunning{StartedAt: started}}
+ pod.Status.ContainerStatuses = []corev1.ContainerStatus{{Name: "postgres", State: running}, {Name: "metrics", State: running}}
+ pod.Status.InitContainerStatuses = []corev1.ContainerStatus{{Name: "plugin-barman-cloud", State: running}}
+ seedCNPGPods(t, pod)
+ api := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if !strings.HasSuffix(r.URL.Path, "/log") {
+ http.NotFound(w, r)
+ return
+ }
+ fmt.Fprintf(w, "2026-09-28T14:19:58Z hello %s\n", r.URL.Query().Get("container"))
+ }))
+ t.Cleanup(api.Close)
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: api.URL})
+ if err != nil {
+ t.Fatal(err)
+ }
+ previous := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previous) })
+ for _, suffix := range []string{"?container=all", "?container=all&sinceTime=2026-09-28T14:00:00Z&untilTime=2026-09-28T14:30:00Z"} {
+ status, got, body := getCNPGLogs(t, "/api/cnpg/clusters/"+ns+"/pg-orders/logs"+suffix)
+ if status != http.StatusOK {
+ t.Fatalf("%d: %s", status, body)
+ }
+ found := map[string]bool{}
+ for _, line := range got.Logs {
+ found[line.Container] = true
+ if !strings.Contains(line.SourceLabel, line.Container) {
+ t.Errorf("missing container label: %+v", line)
+ }
+ }
+ for _, name := range []string{"postgres", "metrics", "plugin-barman-cloud"} {
+ if !found[name] {
+ t.Errorf("missing %s: %+v", name, got)
+ }
+ }
+ if found["not-started"] {
+ t.Error("read unstarted init container")
+ }
+ }
+ ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+ defer cancel()
+ req, _ := http.NewRequestWithContext(ctx, "GET", testServer.URL+"/api/cnpg/clusters/"+ns+"/pg-orders/logs/stream?container=all&pod=pg-orders-1", nil)
+ resp, err := http.DefaultClient.Do(req)
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ if resp.StatusCode != http.StatusOK {
+ t.Fatal(resp.StatusCode)
+ }
+ scanner := bufio.NewScanner(resp.Body)
+ stream := ""
+ for scanner.Scan() {
+ stream += scanner.Text() + "\n"
+ if strings.Contains(stream, `"container":"postgres"`) && strings.Contains(stream, `"container":"metrics"`) && strings.Contains(stream, `"container":"plugin-barman-cloud"`) {
+ break
+ }
+ }
+ for _, container := range []string{"postgres", "metrics", "plugin-barman-cloud"} {
+ if !strings.Contains(stream, `"sourceLabel":"primary 1 · `+container+`"`) {
+ t.Errorf("stream did not label %s: %s", container, stream)
+ }
+ }
+}
+
+func TestCNPGLogsStreamCompletedContainers(t *testing.T) {
+ ns := "pglogscompleted"
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders", nil, nil), "orders-uid"))
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "orders-uid", Controller: boolPtr(true)}
+ pod := cnpgPod(ns, "pg-orders-1", "pg-orders", owner)
+ pod.Spec.InitContainers = []corev1.Container{{Name: "bootstrap-controller"}, {Name: "not-started"}}
+ pod.Status.InitContainerStatuses = []corev1.ContainerStatus{{Name: "bootstrap-controller", ContainerID: "containerd://old", State: corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{ExitCode: 0}}}}
+ seedCNPGPods(t, pod)
+ var mu sync.Mutex
+ counts := map[string]int{}
+ requests := make(chan string, 100)
+ api := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ container := r.URL.Query().Get("container")
+ mu.Lock()
+ counts[container]++
+ mu.Unlock()
+ requests <- container
+ fmt.Fprintln(w, "2026-09-28T14:19:58Z line")
+ }))
+ t.Cleanup(api.Close)
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: api.URL})
+ if err != nil {
+ t.Fatal(err)
+ }
+ previous := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previous) })
+ ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
+ defer cancel()
+ req, _ := http.NewRequestWithContext(ctx, "GET", testServer.URL+"/api/cnpg/clusters/"+ns+"/pg-orders/logs/stream?container=all", nil)
+ resp, err := http.DefaultClient.Do(req)
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ for runningReads := 0; runningReads < 3; {
+ select {
+ case container := <-requests:
+ if container == "postgres" {
+ runningReads++
+ }
+ case <-ctx.Done():
+ t.Fatal("stream did not rediscover running containers")
+ }
+ }
+ mu.Lock()
+ finishedReads, unstartedReads := counts["bootstrap-controller"], counts["not-started"]
+ mu.Unlock()
+ if finishedReads != 1 || unstartedReads != 0 {
+ t.Fatalf("finished reads=%d, unstarted reads=%d", finishedReads, unstartedReads)
+ }
+ pod.Status.InitContainerStatuses[0].ContainerID = "containerd://new"
+ pod.Status.InitContainerStatuses[0].RestartCount++
+ if _, err := testFakeClient.CoreV1().Pods(ns).Update(ctx, pod, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ for {
+ select {
+ case container := <-requests:
+ if container == "bootstrap-controller" {
+ return
+ }
+ case <-ctx.Done():
+ t.Fatal("a new terminated container run was not read")
+ }
+ }
+}
+
+func TestCNPGLogsWaitingPreviousRun(t *testing.T) {
+ ns := "pglogscrash"
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders", nil, nil), "orders-uid"))
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "orders-uid", Controller: boolPtr(true)}
+ pod := cnpgPod(ns, "pg-orders-1", "pg-orders", owner)
+ pod.Spec.Containers = []corev1.Container{{Name: "postgres"}, {Name: "metrics"}, {Name: "not-started"}}
+ pod.Status.ContainerStatuses = []corev1.ContainerStatus{
+ {Name: "postgres", RestartCount: 1, State: corev1.ContainerState{Waiting: &corev1.ContainerStateWaiting{Reason: "CrashLoopBackOff"}}, LastTerminationState: corev1.ContainerState{Terminated: &corev1.ContainerStateTerminated{ContainerID: "containerd://crash", ExitCode: 1}}},
+ {Name: "metrics", State: corev1.ContainerState{Running: &corev1.ContainerStateRunning{}}},
+ {Name: "not-started", State: corev1.ContainerState{Waiting: &corev1.ContainerStateWaiting{Reason: "ContainerCreating"}}},
+ }
+ seedCNPGPods(t, pod)
+ var mu sync.Mutex
+ counts := map[string]int{}
+ requests := make(chan string, 100)
+ api := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ container := r.URL.Query().Get("container")
+ if container == "postgres" && (r.URL.Query().Get("previous") != "true" || r.URL.Query().Get("follow") == "true") {
+ t.Errorf("wrong crash-log options: %s", r.URL.RawQuery)
+ }
+ mu.Lock()
+ counts[container]++
+ mu.Unlock()
+ requests <- container
+ fmt.Fprintf(w, "2026-09-28T14:19:58Z retained %s\n", container)
+ }))
+ t.Cleanup(api.Close)
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: api.URL})
+ if err != nil {
+ t.Fatal(err)
+ }
+ previous := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previous) })
+ status, got, body := getCNPGLogs(t, "/api/cnpg/clusters/"+ns+"/pg-orders/logs?container=all")
+ if status != http.StatusOK {
+ t.Fatalf("%d: %s", status, body)
+ }
+ crash := false
+ for _, line := range got.Logs {
+ if line.Container == "postgres" {
+ crash = line.Previous && strings.Contains(line.Content, "retained postgres")
+ }
+ if line.Container == "not-started" {
+ t.Fatal("read container that never started")
+ }
+ }
+ if !crash {
+ t.Fatalf("retained crash logs missing: %+v", got)
+ }
+ mu.Lock()
+ counts = map[string]int{}
+ mu.Unlock()
+ for len(requests) > 0 {
+ <-requests
+ }
+ ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
+ defer cancel()
+ req, _ := http.NewRequestWithContext(ctx, "GET", testServer.URL+"/api/cnpg/clusters/"+ns+"/pg-orders/logs/stream?container=all", nil)
+ resp, err := http.DefaultClient.Do(req)
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ scanner := bufio.NewScanner(resp.Body)
+ stream := ""
+ for scanner.Scan() {
+ stream += scanner.Text() + "\n"
+ if strings.Contains(stream, `"container":"postgres"`) {
+ break
+ }
+ }
+ if !strings.Contains(stream, `"previous":true`) || !strings.Contains(stream, "retained postgres") || !strings.Contains(stream, `"sourceLabel":"primary 1 · postgres · previous run"`) {
+ t.Fatalf("previous-run stream metadata missing: %s", stream)
+ }
+ for runningReads := 0; runningReads < 3; {
+ select {
+ case container := <-requests:
+ if container == "metrics" {
+ runningReads++
+ }
+ case <-ctx.Done():
+ t.Fatal("stream did not rediscover running container")
+ }
+ }
+ mu.Lock()
+ crashReads, neverReads := counts["postgres"], counts["not-started"]
+ mu.Unlock()
+ if crashReads != 1 || neverReads != 0 {
+ t.Fatalf("crash reads=%d, never-started reads=%d", crashReads, neverReads)
+ }
+}
diff --git a/internal/server/cnpg_destroy.go b/internal/server/cnpg_destroy.go
new file mode 100644
index 0000000000..70e496f836
--- /dev/null
+++ b/internal/server/cnpg_destroy.go
@@ -0,0 +1,39 @@
+package server
+
+import (
+ "net/http"
+
+ "github.com/go-chi/chi/v5"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+)
+
+// Destroying an instance mirrors `kubectl cnpg destroy CLUSTER INSTANCE
+// [--keep-pvc]`: the instance's PVCs are detached (keep) or deleted first, so
+// the operator never sees a dangling PVC and recreates the Pod on it; then the
+// Pod is deleted, then the instance's Jobs. The operator replaces the instance
+// with a new one under a new serial. Unlike upstream, Radar refuses the
+// primary (an unplanned failover, which a switchover does safely) and
+// requires the instance to be fenced first: the operator never promotes a
+// fenced instance, which is the only thing that keeps a failover from making
+// it primary between the last check and a delete. The fence on the destroyed
+// name is lifted last.
+
+func (s *Server) handleCNPGDestroyPlan(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name, pod := chi.URLParam(r, "namespace"), chi.URLParam(r, "name"), chi.URLParam(r, "pod")
+ reader := s.cnpgReader(r)
+ dyn, contextName, typed := reader.dynamic, reader.actionContext, reader.Clients.Typed
+ if dyn == nil || typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ plan, err := reader.DestroyPlan(r.Context(), cnpgsvc.ActionClients{Dynamic: dyn, Typed: typed}, contextName, namespace, name, pod)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "destroyInstance", namespace, name)
+ return
+ }
+ s.writeJSON(w, plan)
+}
diff --git a/internal/server/cnpg_handlers.go b/internal/server/cnpg_handlers.go
index 0c60a3fbb9..b4dd24d4a2 100644
--- a/internal/server/cnpg_handlers.go
+++ b/internal/server/cnpg_handlers.go
@@ -1,73 +1,13 @@
package server
import (
- "errors"
- "log"
"net/http"
- "sort"
"github.com/go-chi/chi/v5"
- "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "github.com/skyhook-io/radar/internal/k8s"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
)
-const cnpgGroup = "postgresql.cnpg.io"
-
-// CNPGCatalogUser is one Cluster pinned to an image catalog.
-type CNPGCatalogUser struct {
- Namespace string `json:"namespace"`
- Name string `json:"name"`
- // Major is the PostgreSQL major the Cluster asks the catalog for. A
- // reference without one lands here as 0, which the screen must not describe
- // as "asks for PostgreSQL 0".
- Major int `json:"major,omitempty"`
- // Image the Cluster resolved and is running now. A cluster pinned to a
- // catalog carries no spec.imageName, so this is the only place the running
- // image appears.
- Image string `json:"image,omitempty"`
-}
-
-// CNPGCatalogUsersResponse lists the Clusters referencing one image catalog.
-type CNPGCatalogUsersResponse struct {
- Clusters []CNPGCatalogUser `json:"clusters"`
-}
-
-// catalogRefMatches reports whether a Cluster's imageCatalogRef names this
-// catalog.
-//
-// CloudNativePG defaults an omitted `kind` to the namespaced ImageCatalog, so a
-// reference without one must not be counted against a ClusterImageCatalog of the
-// same name — the two are different objects and may both exist.
-func catalogRefMatches(ref map[string]interface{}, name, wantKind string) bool {
- if refName, _ := ref["name"].(string); refName != name {
- return false
- }
- refKind, _ := ref["kind"].(string)
- if refKind == "" {
- refKind = "ImageCatalog"
- }
- return refKind == wantKind
-}
-
-// catalogRefMajor reads the PostgreSQL major from a catalog reference.
-//
-// Read tolerantly: the same field arrives as int64 or float64 depending on how
-// the object entered the dynamic cache, and NestedInt64 alone misses the float64
-// shape — which would drop a real major to zero, a state the screen treats as
-// "the reference carries no major at all".
-func catalogRefMajor(ref map[string]interface{}) int {
- switch v := ref["major"].(type) {
- case int64:
- return int(v)
- case float64:
- return int(v)
- case int:
- return v
- }
- return 0
-}
-
// handleCNPGCatalogUsers returns the Clusters pinned to an image catalog.
//
// GET /api/cnpg/imagecatalogs/{namespace}/{name}/clusters
@@ -96,59 +36,20 @@ func (s *Server) handleCNPGCatalogUsers(w http.ResponseWriter, r *http.Request)
wantKind = "ImageCatalog"
}
- if !s.canRead(r, cnpgGroup, "clusters", namespace, "list") {
+ if !s.canRead(r, cnpgsvc.Group, "clusters", namespace, "list") {
s.writeError(w, http.StatusForbidden, "no access to CloudNativePG clusters")
return
}
- cache := k8s.GetResourceCache()
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
if cache == nil {
s.writeError(w, http.StatusServiceUnavailable, "Resource cache not available")
return
}
- // A namespaced catalog can only be referenced from its own namespace; a
- // cluster-scoped one from anywhere.
- items, err := listDynamicSynced(r.Context(), cache, "Cluster", cnpgGroup, namespace)
- switch {
- case err == nil:
- case errors.Is(err, k8s.ErrUnknownDynamicKind):
- // No CloudNativePG on this cluster, so nothing can be pinned to a catalog.
- s.writeJSON(w, CNPGCatalogUsersResponse{Clusters: []CNPGCatalogUser{}})
+ resp, err := reader.CatalogUsers(r.Context(), cache, namespace, name, wantKind)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
return
- case errors.Is(err, errDynamicNotSynced):
- s.writeError(w, http.StatusServiceUnavailable, "clusters are still loading")
- return
- default:
- // "No cluster uses this catalog" is the sentence someone reads before
- // editing it. Never say it because the lookup failed.
- log.Printf("[cnpg] Failed to list Clusters for catalog %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
- s.writeError(w, http.StatusServiceUnavailable, "could not read CloudNativePG clusters")
- return
- }
-
- resp := CNPGCatalogUsersResponse{Clusters: []CNPGCatalogUser{}}
- for _, u := range items {
- if u == nil {
- continue
- }
- ref, found, _ := unstructured.NestedMap(u.Object, "spec", "imageCatalogRef")
- if !found {
- continue
- }
- if !catalogRefMatches(ref, name, wantKind) {
- continue
- }
- user := CNPGCatalogUser{Namespace: u.GetNamespace(), Name: u.GetName()}
- user.Major = catalogRefMajor(ref)
- if img, _, _ := unstructured.NestedString(u.Object, "status", "image"); img != "" {
- user.Image = img
- }
- resp.Clusters = append(resp.Clusters, user)
}
- sort.Slice(resp.Clusters, func(i, j int) bool {
- if resp.Clusters[i].Namespace != resp.Clusters[j].Namespace {
- return resp.Clusters[i].Namespace < resp.Clusters[j].Namespace
- }
- return resp.Clusters[i].Name < resp.Clusters[j].Name
- })
s.writeJSON(w, resp)
}
diff --git a/internal/server/cnpg_handlers_test.go b/internal/server/cnpg_handlers_test.go
index d4acb251f5..39e73ead21 100644
--- a/internal/server/cnpg_handlers_test.go
+++ b/internal/server/cnpg_handlers_test.go
@@ -1,11 +1,15 @@
package server
import (
- "os"
- "strings"
+ "context"
+ "errors"
"testing"
+ "time"
"github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
)
// A ClusterImageCatalog is cluster-scoped and referenceable from any namespace,
@@ -14,46 +18,13 @@ import (
// namespace view filter, and "no cluster uses this catalog" is exactly the
// sentence someone reads before editing it.
-// CloudNativePG defaults an omitted `kind` to the namespaced ImageCatalog. A
-// namespaced and a cluster-scoped catalog may share a name, so a reference that
-// omits the kind must not be counted against the cluster-scoped one.
-func TestCatalogRefMatches_DefaultsToTheNamespacedKind(t *testing.T) {
- for _, c := range []struct {
- name string
- ref map[string]interface{}
- catalog string
- wantKind string
- want bool
- }{
- {"omitted kind counts as ImageCatalog",
- map[string]interface{}{"name": "pg17"}, "pg17", "ImageCatalog", true},
- {"omitted kind is NOT a ClusterImageCatalog",
- map[string]interface{}{"name": "pg17"}, "pg17", "ClusterImageCatalog", false},
- {"explicit cluster-scoped matches its own kind",
- map[string]interface{}{"name": "pg17", "kind": "ClusterImageCatalog"}, "pg17", "ClusterImageCatalog", true},
- {"explicit cluster-scoped does not match the namespaced kind",
- map[string]interface{}{"name": "pg17", "kind": "ClusterImageCatalog"}, "pg17", "ImageCatalog", false},
- {"another catalog entirely",
- map[string]interface{}{"name": "pg16"}, "pg17", "ImageCatalog", false},
- {"no name at all",
- map[string]interface{}{}, "pg17", "ImageCatalog", false},
- } {
- t.Run(c.name, func(t *testing.T) {
- if got := catalogRefMatches(c.ref, c.catalog, c.wantKind); got != c.want {
- t.Errorf("catalogRefMatches(%v, %q, %q) = %v, want %v",
- c.ref, c.catalog, c.wantKind, got, c.want)
- }
- })
- }
-}
-
// A caller who may not list Clusters is told so. Returning an empty list would
// read as "nothing depends on this catalog", which is the answer that gets a
// catalog edited out from under a running database.
func TestCNPGCatalogUsers_DeniesRatherThanReportingNoUsers(t *testing.T) {
env := newAuthTestServer(t)
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- allow(perms, cnpgGroup, "clusters", "", false)
+ allow(perms, cnpgsvc.Group, "clusters", "", false)
env.srv.permCache.Set("nobody", nil, perms)
resp := env.authGet(t, "/api/cnpg/clusterimagecatalogs/postgres-fleet/clusters", "nobody", "")
@@ -67,7 +38,7 @@ func TestCNPGCatalogUsers_DeniesRatherThanReportingNoUsers(t *testing.T) {
func TestCNPGCatalogUsers_NamespacedRouteAuthorizesInItsNamespace(t *testing.T) {
env := newAuthTestServer(t)
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- allow(perms, cnpgGroup, "clusters", "pg", false)
+ allow(perms, cnpgsvc.Group, "clusters", "pg", false)
env.srv.permCache.Set("scoped", nil, perms)
resp := env.authGet(t, "/api/cnpg/imagecatalogs/pg/postgres-pinned/clusters", "scoped", "")
@@ -77,87 +48,27 @@ func TestCNPGCatalogUsers_NamespacedRouteAuthorizesInItsNamespace(t *testing.T)
}
}
-// The dynamic cache hands the same field back as int64 or float64 depending on
-// how the object entered it. Missing the float64 shape would drop a real major
-// to zero — which the screen reads as "the reference carries no major", a
-// different and wrong statement.
-func TestCatalogRefMajor_ReadsEitherNumberShape(t *testing.T) {
- for _, c := range []struct {
- name string
- ref map[string]interface{}
- want int
- }{
- {"int64 from a typed decode", map[string]interface{}{"major": int64(17)}, 17},
- {"float64 from a JSON decode", map[string]interface{}{"major": float64(17)}, 17},
- {"plain int", map[string]interface{}{"major": 17}, 17},
- {"absent", map[string]interface{}{}, 0},
- {"a string is not a major", map[string]interface{}{"major": "17"}, 0},
- } {
- t.Run(c.name, func(t *testing.T) {
- if got := catalogRefMajor(c.ref); got != c.want {
- t.Errorf("catalogRefMajor(%v) = %d, want %d", c.ref, got, c.want)
- }
- })
- }
-}
-
-// Same contract as the queued endpoint: an absent CRD is a real "nothing here",
-// every other failure is a failure to look and must not read as "no cluster uses
-// this catalog" — the sentence someone acts on before editing it.
-func TestCNPGCatalogUsers_SeparatesAnAbsentCRDFromAFailedRead(t *testing.T) {
- src, err := os.ReadFile("cnpg_handlers.go")
- if err != nil {
- t.Fatalf("read handler: %v", err)
- }
- h := string(src)
- if !strings.Contains(h, "k8s.ErrUnknownDynamicKind") {
- t.Error("handler does not special-case an absent CRD, so every failure reads as no users")
- }
- if !strings.Contains(h, "StatusServiceUnavailable") {
- t.Error("handler has no failure path — a denied or unready read would answer 200 with no users")
- }
- if !strings.Contains(h, "listDynamicSynced") {
- t.Error("handler does not wait for the informer to sync; a cold read answers 'no users' before looking")
- }
-}
-
// An empty result is only an absence if the cache actually looked. The informer
// for a namespace is started BY the read, so a gate that refuses to read until
// one exists never lets the first read happen — ListBlocking starts it and waits
// instead, which is what makes a subsequent empty list mean something.
func TestHandlersEstablishAbsenceRatherThanAssumeIt(t *testing.T) {
- src, err := os.ReadFile("policy_handlers.go")
- if err != nil {
- t.Fatalf("read: %v", err)
- }
- body := string(src)
- i := strings.Index(body, "func listDynamicSynced")
- if i < 0 {
- t.Fatal("helper moved")
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds)
+ items, err := listDynamicSynced(context.Background(), k8s.GetResourceCache(), "Cluster", cnpgsvc.Group, "cold-empty")
+ if err != nil || len(items) != 0 {
+ t.Fatalf("cold empty namespace = %v, %v", items, err)
}
- fn := body[i : i+1400]
- if !strings.Contains(fn, "ListBlocking") {
- t.Error("the read does not wait for sync, so an empty list may be a cache that never looked")
+ if !k8s.GetDynamicResourceCache().IsNamespaceSynced(cnpgsvc.ClusterGVR, "cold-empty") {
+ t.Fatal("empty result did not establish the requested scope")
}
- // ListBlocking discards WaitForCacheSync's result: on timeout it returns an
- // empty list and no error, which is the false absence one layer down.
- if !strings.Contains(fn, "IsNamespaceSynced") {
- t.Error("nothing confirms the wait actually succeeded; a timed-out informer reads as an absence")
- }
- if !strings.Contains(fn, "errDynamicNotSynced") {
- t.Error("a cache that could not answer must say so, not return an empty result")
- }
- for _, f := range []string{"policy_handlers.go", "cnpg_handlers.go"} {
- b, err := os.ReadFile(f)
- if err != nil {
- t.Fatalf("read %s: %v", f, err)
- }
- if !strings.Contains(string(b), "listDynamicSynced(") {
- t.Errorf("%s reads the cache without waiting for the scope it reads", f)
- }
- // A pre-read gate is the deadlock: nothing else starts the informer.
- if strings.Contains(string(b), "if !dynamicKindSynced(") {
- t.Errorf("%s gates on sync before reading; the first read of a namespace would never succeed", f)
- }
+}
+
+func TestFailedDynamicReadDoesNotEstablishAbsence(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"healthy"}, fail: inNamespaces("broken")})
+ budget := newSyncBudget(context.Background())
+ budget.deadline = time.Now().Add(30 * time.Millisecond)
+ items, err := listDynamicSyncedWithin(context.Background(), k8s.GetResourceCache(), "Cluster", cnpgsvc.Group, "broken", budget)
+ if !errors.Is(err, integration.ErrDynamicNotSynced) || len(items) != 0 {
+ t.Fatalf("failed inventory must be unread rather than empty success: %v, %v", items, err)
}
}
diff --git a/internal/server/cnpg_history.go b/internal/server/cnpg_history.go
new file mode 100644
index 0000000000..c6363311f7
--- /dev/null
+++ b/internal/server/cnpg_history.go
@@ -0,0 +1,60 @@
+package server
+
+import (
+ "net/http"
+
+ "github.com/go-chi/chi/v5"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+)
+
+func (s *Server) handleCNPGClusterHistory(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.requireConnected(w) {
+ return
+ }
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
+ return
+ }
+ if !s.canRead(r, cnpgsvc.Group, "clusters", namespace, "get") {
+ s.writeError(w, http.StatusForbidden, "no access to clusters.postgresql.cnpg.io in namespace "+namespace)
+ return
+ }
+ rng, ok := prometheuspkg.ParseCNPGHistoryRange(r.URL.Query().Get("range"))
+ if !ok {
+ s.writeError(w, http.StatusBadRequest, "invalid range "+r.URL.Query().Get("range")+" (expected 15m, 1h, 6h or 24h)")
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ _, cluster, err := reader.Observations.Cluster(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+
+ s.writeJSON(w, reader.ClusterHistory(r.Context(), cache, cluster, rng))
+}
+
+// ---------- fleet ----------
+
+func (s *Server) handleCNPGFleetMetrics(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ namespaces := s.parseNamespacesForUser(r)
+ s.writeJSON(w, reader.FleetMetrics(r.Context(), cache, namespaces))
+}
diff --git a/internal/server/cnpg_history_test.go b/internal/server/cnpg_history_test.go
new file mode 100644
index 0000000000..c918f3cf52
--- /dev/null
+++ b/internal/server/cnpg_history_test.go
@@ -0,0 +1,326 @@
+package server
+
+import (
+ "encoding/json"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "sync"
+ "testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+)
+
+// useCNPGHistoryPrometheus serves a Prometheus whose range queries answer
+// replay lag for pg-orders-2 and whose instant queries report pg-orders-1
+// scraped as primary and pg-orders-2 as a standby 4 s behind.
+func useCNPGHistoryPrometheus(t *testing.T) *[]string {
+ t.Helper()
+ return useCNPGHistoryPrometheusAnswering(t, nil)
+}
+
+// useCNPGHistoryPrometheusAnswering is useCNPGHistoryPrometheus with answer
+// consulted first for instant queries: it returns the result array's JSON, or
+// "" to fall through.
+func useCNPGHistoryPrometheusAnswering(t *testing.T, answer func(q string) string) *[]string {
+ t.Helper()
+ var mu sync.Mutex
+ var seen []string
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ _ = r.ParseForm()
+ q := r.Form.Get("query")
+ mu.Lock()
+ seen = append(seen, q)
+ mu.Unlock()
+ empty := func(kind string) string {
+ return `{"status":"success","data":{"resultType":"` + kind + `","result":[]}}`
+ }
+ if strings.HasSuffix(r.URL.Path, "/query_range") {
+ if strings.Contains(q, "cnpg_pg_replication_lag") {
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"matrix","result":[{"metric":{"pod":"pg-orders-2","job":"x"},"values":[[1700000000,"1"],[1700000030,"4"]]}]}}`)
+ return
+ }
+ _, _ = io.WriteString(w, empty("matrix"))
+ return
+ }
+ if answer != nil {
+ if result := answer(q); result != "" {
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":`+result+`}}`)
+ return
+ }
+ }
+ switch {
+ case q == "up":
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{"job":"prometheus"},"value":[1700000000,"1"]}]}}`)
+ case strings.HasPrefix(q, "(max by (pod) (cnpg_pg_replication_lag"):
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{"pod":"pg-orders-1"},"value":[1700000000,"-1"]},{"metric":{"pod":"pg-orders-2"},"value":[1700000000,"4"]}]}}`)
+ case strings.HasPrefix(q, "max(count by (pod)"):
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1700000000,"1"]}]}}`)
+ default:
+ _, _ = io.WriteString(w, empty("vector"))
+ }
+ }))
+ prometheuspkg.Initialize(nil, nil, "test")
+ prometheuspkg.SetManualURL(srv.URL)
+ t.Cleanup(func() {
+ srv.Close()
+ prometheuspkg.Reset()
+ prometheuspkg.Initialize(nil, nil, "")
+ })
+ return &seen
+}
+
+func getCNPGHistory(t *testing.T, path string) (int, cnpgsvc.CNPGClusterHistoryResponse, string) {
+ t.Helper()
+ resp, err := http.Get(testServer.URL + path)
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ var got cnpgsvc.CNPGClusterHistoryResponse
+ _ = json.Unmarshal(body, &got)
+ return resp.StatusCode, got, string(body)
+}
+
+func TestCNPGClusterHistory_NoPrometheusSaysSo(t *testing.T) {
+ seedCNPGStorageCluster(t, "pghi1")
+ prometheuspkg.Reset()
+ status, got, body := getCNPGHistory(t, "/api/cnpg/clusters/pghi1/pg-orders/history?range=15m")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.Source != "none" || got.Reason == "" || len(got.Charts) != 0 {
+ t.Fatalf("got %+v", got)
+ }
+ if status, _, _ := getCNPGHistory(t, "/api/cnpg/clusters/pghi1/pg-orders/history?range=7d"); status != http.StatusBadRequest {
+ t.Errorf("range 7d: %d, want 400", status)
+ }
+ if status, _, _ := getCNPGHistory(t, "/api/cnpg/clusters/pghi1/nope/history"); status != http.StatusNotFound {
+ t.Errorf("unknown cluster: %d, want 404", status)
+ }
+}
+
+func TestCNPGClusterHistory_ChartsFromPrometheus(t *testing.T) {
+ seedCNPGStorageCluster(t, "pghi2")
+ seen := useCNPGHistoryPrometheus(t)
+ status, got, body := getCNPGHistory(t, "/api/cnpg/clusters/pghi2/pg-orders/history?range=1h")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.Source != "prometheus" || got.State != "ok" || got.StepSeconds != 30 || got.Isolation == nil || got.Isolation.Mode != prometheuspkg.SeriesIsolationUnverified {
+ t.Fatalf("got %+v", got)
+ }
+ // The volume chart's claims are matched apart from the instance Pods, and say so.
+ if got.PVCIsolation == nil || got.PVCIsolation.Mode != prometheuspkg.SeriesIsolationUnverified || !strings.Contains(got.PVCIsolation.Note, "claim names") {
+ t.Errorf("pvcIsolation = %+v", got.PVCIsolation)
+ }
+ by := map[string]prometheuspkg.CNPGHistoryChart{}
+ for _, c := range got.Charts {
+ by[c.ID] = c
+ }
+ if lag := by["replicationLag"]; lag.State != prometheuspkg.CNPGHistoryStateOK || len(lag.Series) != 1 || lag.Series[0].Labels["pod"] != "pg-orders-2" {
+ t.Errorf("lag = %+v", lag)
+ }
+ if pvc := by["pvcUsed"]; pvc.State != prometheuspkg.CNPGHistoryStateNoSeries || !strings.Contains(pvc.Reason, "not scraped") {
+ t.Errorf("pvc = %+v", pvc)
+ }
+ var pvcQuery string
+ for _, q := range *seen {
+ if strings.Contains(q, "kubelet_volume_stats_used_bytes") {
+ pvcQuery = q
+ }
+ }
+ if !strings.Contains(pvcQuery, "pg-orders-1-wal") || strings.Contains(pvcQuery, "pg-orders-9") {
+ t.Errorf("PVC query must name owned claims only: %s", pvcQuery)
+ }
+ if !strings.Contains(body, `"value":null`) && strings.Contains(body, `"value":0,`) {
+ t.Errorf("a gap was rendered as zero: %s", body)
+ }
+}
+
+func TestCNPGClusterHistory_GatesPerSource(t *testing.T) {
+ seedCNPGStorageCluster(t, "pghi3")
+ useCNPGHistoryPrometheus(t)
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"pghi3"}}
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", "pghi3", true)
+ allow(perms, "", "persistentvolumeclaims", "pghi3", false)
+ allow(perms, "", "pods", "pghi3", false)
+ env.srv.permCache.Set("dba", nil, perms)
+
+ resp := env.authGet(t, "/api/cnpg/clusters/pghi3/pg-orders/history?range=15m", "dba", "")
+ defer resp.Body.Close()
+ if resp.StatusCode != http.StatusOK {
+ b, _ := io.ReadAll(resp.Body)
+ t.Fatalf("status = %d: %s", resp.StatusCode, b)
+ }
+ var got cnpgsvc.CNPGClusterHistoryResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ for _, c := range got.Charts {
+ if c.State != prometheuspkg.CNPGHistoryStateDenied || len(c.Series) != 0 || c.Grant == nil {
+ t.Errorf("%s = %+v, want denied without series", c.ID, c)
+ }
+ }
+
+ perms2 := &auth.UserPermissions{AllowedNamespaces: []string{"pghi3"}}
+ perms2.SetCanI("get", cnpgsvc.Group, "clusters", "pghi3", false)
+ env.srv.permCache.Set("nobody", nil, perms2)
+ denied := env.authGet(t, "/api/cnpg/clusters/pghi3/pg-orders/history", "nobody", "")
+ denied.Body.Close()
+ if denied.StatusCode != http.StatusForbidden {
+ t.Errorf("without get clusters: %d, want 403", denied.StatusCode)
+ }
+}
+
+func TestCNPGFleetMetrics(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgfm")
+ useCNPGHistoryPrometheus(t)
+ resp, err := http.Get(testServer.URL + "/api/cnpg/fleet-metrics?namespaces=pgfm")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGFleetMetricsResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if got.Source != "prometheus" || len(got.Clusters) != 1 {
+ t.Fatalf("got %+v", got)
+ }
+ c := got.Clusters[0]
+ if c.Lag.State != "ok" || c.Lag.Seconds == nil || *c.Lag.Seconds != 4 || c.Lag.Pod != "pg-orders-2" {
+ t.Errorf("lag = %+v", c.Lag)
+ }
+ if c.Growth.State != "noSeries" || c.Growth.BytesPerHour != nil {
+ t.Errorf("growth = %+v", c.Growth)
+ }
+ if !c.Lag.ReceiverUnknown || c.Lag.Receiving != nil || c.Lag.ReceiverDown != nil {
+ t.Errorf("no receiver series must read unknown, never receiving: %+v", c.Lag)
+ }
+ if c.Slots.State != "noSeries" || c.Slots.Inactive != nil {
+ t.Errorf("slots = %+v, want noSeries without a list", c.Slots)
+ }
+}
+
+// pg-orders-2 receives nothing: its lag reads 0 because nothing is left to
+// replay, while the primary keeps its slot inactive with 4.9 GB of WAL.
+func TestCNPGFleetMetricsReceiverDownAndInactiveSlot(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgfr")
+ row := func(labels string, v string) string {
+ return `{"metric":{` + labels + `},"value":[1700000000,"` + v + `"]}`
+ }
+ seen := useCNPGHistoryPrometheusAnswering(t, func(q string) string {
+ switch {
+ case strings.HasPrefix(q, "(max by (pod) (cnpg_pg_replication_lag"):
+ return "[" + row(`"pod":"pg-orders-1"`, "-1") + "," + row(`"pod":"pg-orders-2"`, "0") + "]"
+ case strings.HasPrefix(q, "(max by (pod) (cnpg_pg_replication_is_wal_receiver_up"):
+ return "[" + row(`"pod":"pg-orders-2"`, "0") + "]"
+ case strings.HasPrefix(q, "label_replace("):
+ return "[" + strings.Join([]string{
+ row(`"pod":"pg-orders-1","slot_name":"_cnpg_pg_orders_2","radar_row":"inactive"`, "0"),
+ row(`"pod":"pg-orders-1","slot_name":"_cnpg_pg_orders_2","radar_row":"retained"`, "4900000000"),
+ row(`"pod":"pg-orders-1","radar_row":"reported"`, "1"),
+ row(`"pod":"pg-orders-1","radar_row":"recovery"`, "0"),
+ row(`"pod":"pg-orders-2","radar_row":"recovery"`, "1"),
+ row(`"pod":"pg-orders-1","radar_row":"scraped"`, "1"),
+ row(`"pod":"pg-orders-2","radar_row":"scraped"`, "1"),
+ }, ",") + "]"
+ }
+ return ""
+ })
+ resp, err := http.Get(testServer.URL + "/api/cnpg/fleet-metrics?namespaces=pgfr")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ var raw struct {
+ ReceiverSource string `json:"receiverSource"`
+ SlotsSource string `json:"slotsSource"`
+ Clusters []map[string]any `json:"clusters"`
+ }
+ if err := json.Unmarshal(body, &raw); err != nil || len(raw.Clusters) != 1 {
+ t.Fatalf("decode %v: %s", err, body)
+ }
+ if raw.ReceiverSource == "" || raw.SlotsSource == "" {
+ t.Errorf("sources missing: %s", body)
+ }
+ lag := raw.Clusters[0]["lag"].(map[string]any)
+ for k, want := range map[string]any{"state": "ok", "seconds": 0.0, "pod": "pg-orders-2", "standbys": 1.0, "receiving": 0.0} {
+ if lag[k] != want {
+ t.Errorf("lag.%s = %v, want %v (%v)", k, lag[k], want, lag)
+ }
+ }
+ if down, _ := lag["receiverDown"].([]any); len(down) != 1 || down[0] != "pg-orders-2" {
+ t.Errorf("lag.receiverDown = %v", lag["receiverDown"])
+ }
+ if _, ok := lag["receiverUnknown"]; ok {
+ t.Errorf("receiverUnknown set although the receiver was read: %v", lag)
+ }
+ slots := raw.Clusters[0]["slots"].(map[string]any)
+ inactive, _ := slots["inactive"].([]any)
+ if slots["state"] != "ok" || len(inactive) != 1 || slots["isolation"] == nil {
+ t.Fatalf("slots = %v", slots)
+ }
+ want := map[string]any{"slot": "_cnpg_pg_orders_2", "pod": "pg-orders-1", "role": "primary", "bytes": 4.9e9}
+ got := inactive[0].(map[string]any)
+ if len(got) != len(want) {
+ t.Errorf("slot fields = %v, want exactly %v", got, want)
+ }
+ for k, v := range want {
+ if got[k] != v {
+ t.Errorf("slot.%s = %v, want %v", k, got[k], v)
+ }
+ }
+ for _, q := range *seen {
+ if (strings.Contains(q, "is_wal_receiver_up") || strings.Contains(q, "cnpg_pg_replication_slots")) && !strings.Contains(q, `namespace="pgfr",pod=~"^pg-orders-[0-9]+$"`) {
+ t.Errorf("query not scoped to the cluster's instances: %s", q)
+ }
+ }
+}
+
+func TestCNPGFleetMetricsSlotsFollowPodsGate(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgfd")
+ seen := useCNPGHistoryPrometheus(t)
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgfd"}}
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
+ allow(perms, cnpgsvc.Group, "clusters", "pgfd", true)
+ perms.SetCanI("get", "", "pods", "pgfd", false)
+ allow(perms, "", "persistentvolumeclaims", "pgfd", false)
+ env.srv.permCache.Set("dba", nil, perms)
+
+ resp := env.authGet(t, "/api/cnpg/fleet-metrics?namespaces=pgfd", "dba", "")
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGFleetMetricsResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if len(got.Clusters) != 1 {
+ t.Fatalf("got %+v", got)
+ }
+ c := got.Clusters[0]
+ for name, st := range map[string]struct {
+ State string
+ Grant *Grant
+ }{"lag": {c.Lag.State, c.Lag.Grant}, "slots": {c.Slots.State, c.Slots.Grant}} {
+ if st.State != "denied" || st.Grant == nil || st.Grant.Verb != "get" || st.Grant.Resource != "pods" || st.Grant.Namespace != "pgfd" {
+ t.Errorf("%s = %s %+v, want denied naming get pods in pgfd", name, st.State, st.Grant)
+ }
+ }
+ if c.Slots.Inactive != nil || c.Lag.Standbys != nil {
+ t.Errorf("denied reads still carry data: %+v", c)
+ }
+ for _, q := range *seen {
+ if strings.Contains(q, "cnpg_") {
+ t.Errorf("denied caller's CNPG series were queried: %s", q)
+ }
+ }
+}
diff --git a/internal/server/cnpg_inspect.go b/internal/server/cnpg_inspect.go
new file mode 100644
index 0000000000..d606389f9f
--- /dev/null
+++ b/internal/server/cnpg_inspect.go
@@ -0,0 +1,67 @@
+package server
+
+import (
+ "net/http"
+
+ "github.com/go-chi/chi/v5"
+)
+
+// Read-only inspection inside PostgreSQL, with the same fixed-SQL-over-the-
+// caller's-pods/exec model as the Sessions view: what a restored cluster
+// holds, and which values declared PostgreSQL parameters have on each
+// instance. Nothing here writes, and nothing a caller sends is SQL text.
+
+// handleCNPGRestoreChecks serves GET /api/cnpg/clusters/{ns}/{name}/restore-checks
+// for a Cluster bootstrapped from a backup (400 otherwise): read-only facts
+// from its primary over the caller's pods/exec. Without exec the answer is a
+// 200 whose state is denied.
+func (s *Server) handleCNPGRestoreChecks(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.authorizeCNPGRuntime(w, r, namespace, "clusters") {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ _, cluster, err := reader.Observations.Cluster(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ resp, err := reader.RestoreChecks(r.Context(), cache, cluster)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+// handleCNPGClusterParameters serves GET /api/cnpg/clusters/{ns}/{name}/parameters:
+// observation only — it reads, on every instance, the parameters the Cluster
+// declares. Without exec the answer is a 200 whose state is denied.
+func (s *Server) handleCNPGClusterParameters(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.authorizeCNPGRuntime(w, r, namespace, "clusters") {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ _, cluster, err := reader.Observations.Cluster(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ resp, err := reader.Parameters(r.Context(), cache, cluster)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
diff --git a/internal/server/cnpg_local_cani_test.go b/internal/server/cnpg_local_cani_test.go
new file mode 100644
index 0000000000..edf6ee7828
--- /dev/null
+++ b/internal/server/cnpg_local_cani_test.go
@@ -0,0 +1,52 @@
+package server
+
+import (
+ "context"
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "sync/atomic"
+ "testing"
+
+ authv1 "k8s.io/api/authorization/v1"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func TestCNPGLocalCanIAsksTheKubeconfigIdentity(t *testing.T) {
+ var calls atomic.Int32
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ calls.Add(1)
+ var review authv1.SelfSubjectAccessReview
+ _ = json.NewDecoder(r.Body).Decode(&review)
+ attrs := review.Spec.ResourceAttributes
+ review.Status.Allowed = !(attrs.Resource == "pods" && attrs.Subresource == "proxy" && attrs.Verb == "get")
+ w.Header().Set("Content-Type", "application/json")
+ _ = json.NewEncoder(w).Encode(review)
+ }))
+ defer srv.Close()
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: srv.URL, ContentConfig: rest.ContentConfig{ContentType: "application/json"}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ previous := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previous) })
+ localCanIMu.Lock()
+ localCanIMemo = map[string]localCanIEntry{}
+ localCanIMu.Unlock()
+
+ if allowed, known := localCanI(context.Background(), (auth.Grant{Verb: "get", Resource: "pods", Subresource: "proxy"}).In("pgrt")); allowed || !known {
+ t.Fatalf("pods/proxy = %v known=%v, want denied", allowed, known)
+ }
+ if allowed, known := localCanI(context.Background(), (auth.Grant{Verb: "create", Group: cnpgsvc.Group, Resource: "backups", Namespace: "pgrt"})); !allowed || !known {
+ t.Fatalf("create backups = %v known=%v, want allowed", allowed, known)
+ }
+ localCanI(context.Background(), (auth.Grant{Verb: "get", Resource: "pods", Subresource: "proxy"}).In("pgrt"))
+ if calls.Load() != 2 {
+ t.Errorf("reviews = %d, want the repeat answered from the memo", calls.Load())
+ }
+}
diff --git a/internal/server/cnpg_operator.go b/internal/server/cnpg_operator.go
index 2d48716752..29af3d9841 100644
--- a/internal/server/cnpg_operator.go
+++ b/internal/server/cnpg_operator.go
@@ -1,75 +1,9 @@
package server
import (
- "log"
"net/http"
- "sort"
- "strings"
-
- appsv1 "k8s.io/api/apps/v1"
- corev1 "k8s.io/api/core/v1"
- apierrors "k8s.io/apimachinery/pkg/api/errors"
- "k8s.io/apimachinery/pkg/labels"
-
- "github.com/skyhook-io/radar/internal/k8s"
-)
-
-const (
- cnpgOperatorNameLabel = "app.kubernetes.io/name"
- cnpgOperatorNameValue = "cloudnative-pg"
- cnpgVersionLabel = "app.kubernetes.io/version"
- cnpgPluginNameLabel = "cnpg.io/pluginName"
- cnpgOperatorContainer = "manager"
- cnpgOperatorDeployVar = "OPERATOR_DEPLOYMENT_NAME"
- cnpgMonitoringQueriesCM = "MONITORING_QUERIES_CONFIGMAP"
-
- cnpgOperatorRoleOperator = "operator"
- cnpgOperatorRolePlugin = "plugin"
-
- cnpgConfigPurposeOperator = "operator"
- cnpgConfigPurposeMonitoring = "monitoring"
)
-// CNPGOperatorComponent is one operator or plugin Deployment. Version is the
-// image tag, else the app.kubernetes.io/version label, else empty. Replica
-// counts are nil when unreported, which is not zero.
-type CNPGOperatorComponent struct {
- Role string `json:"role"`
- PluginName string `json:"pluginName,omitempty"`
- Namespace string `json:"namespace"`
- Deployment string `json:"deployment"`
- Image string `json:"image"`
- Version string `json:"version"`
- ReadyReplicas *int32 `json:"readyReplicas"`
- Replicas *int32 `json:"replicas"`
-}
-
-// CNPGOperatorConfigMapState is present only on ConfigMap references. A Secret
-// reference never carries it: the endpoint never reads Secrets.
-type CNPGOperatorConfigMapState struct {
- Exists *bool `json:"exists"`
- Readable bool `json:"readable"`
- Reason string `json:"reason,omitempty"`
- Data map[string]string `json:"data"`
-}
-
-// CNPGOperatorConfigRef is a ConfigMap or Secret the operator is configured
-// to read.
-type CNPGOperatorConfigRef struct {
- Kind string `json:"kind"`
- Namespace string `json:"namespace"`
- Name string `json:"name"`
- Purpose string `json:"purpose"`
- *CNPGOperatorConfigMapState
-}
-
-// CNPGOperatorResponse is GET /api/cnpg/operator.
-type CNPGOperatorResponse struct {
- Coverage map[string]CNPGWorkspaceCoverage `json:"coverage"`
- Components []CNPGOperatorComponent `json:"components"`
- Config []CNPGOperatorConfigRef `json:"config"`
-}
-
// handleCNPGOperator serves GET /api/cnpg/operator: the operator and plugin
// Deployments, their versions and readiness, and where the operator's
// configuration lives.
@@ -79,288 +13,10 @@ type CNPGOperatorResponse struct {
// filter would report "no operator" to anyone looking at their databases, so
// scope follows permission here, as it does for the catalog reverse lookups.
func (s *Server) handleCNPGOperator(w http.ResponseWriter, r *http.Request) {
- if !s.requireConnected(w) {
- return
- }
- cache := k8s.GetResourceCache()
- if cache == nil {
- s.writeError(w, http.StatusServiceUnavailable, "Resource cache not available")
+ resp, err := s.cnpgReader(r).Operator(r.Context())
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, "", "operator")
return
}
-
- scope := s.cnpgOperatorScope(r)
- resp := CNPGOperatorResponse{
- Coverage: map[string]CNPGWorkspaceCoverage{},
- Components: []CNPGOperatorComponent{},
- Config: []CNPGOperatorConfigRef{},
- }
-
- depAcc, depDenied, deployments := s.cnpgOperatorDeployments(r, cache, scope)
- resp.Coverage["deployments"] = cnpgCoverageOf(depAcc, depDenied)
- svcAcc, svcDenied, services := s.cnpgOperatorServices(r, cache, scope)
- resp.Coverage["services"] = cnpgCoverageOf(svcAcc, svcDenied)
-
- var operators []*appsv1.Deployment
- for _, d := range deployments {
- if d.Labels[cnpgOperatorNameLabel] == cnpgOperatorNameValue {
- operators = append(operators, d)
- resp.Components = append(resp.Components, cnpgOperatorComponent(d, cnpgOperatorRoleOperator, "", cnpgOperatorContainerOf(d)))
- }
- }
-
- byNamespace := map[string][]*appsv1.Deployment{}
- for _, d := range deployments {
- byNamespace[d.Namespace] = append(byNamespace[d.Namespace], d)
- }
- var plugins []CNPGOperatorComponent
- for _, svc := range services {
- pluginName := svc.Labels[cnpgPluginNameLabel]
- if pluginName == "" || !depAcc.covers(svc.Namespace) {
- continue
- }
- matched := false
- if len(svc.Spec.Selector) > 0 {
- sel := labels.SelectorFromSet(svc.Spec.Selector)
- for _, d := range byNamespace[svc.Namespace] {
- if sel.Matches(labels.Set(d.Spec.Template.Labels)) {
- matched = true
- plugins = append(plugins, cnpgOperatorComponent(d, cnpgOperatorRolePlugin, pluginName, firstContainer(d)))
- }
- }
- }
- if !matched {
- plugins = append(plugins, CNPGOperatorComponent{Role: cnpgOperatorRolePlugin, PluginName: pluginName, Namespace: svc.Namespace})
- }
- }
- sort.SliceStable(plugins, func(i, j int) bool {
- a, b := plugins[i], plugins[j]
- if a.PluginName != b.PluginName {
- return a.PluginName < b.PluginName
- }
- if a.Namespace != b.Namespace {
- return a.Namespace < b.Namespace
- }
- return a.Deployment < b.Deployment
- })
- resp.Components = append(resp.Components, plugins...)
-
- resp.Config = s.cnpgOperatorConfig(r, cache, operators)
s.writeJSON(w, resp)
}
-
-// cnpgOperatorScope is the caller's RBAC scope without the view filter. A
-// Radar forced into one namespace still answers only for that namespace.
-func (s *Server) cnpgOperatorScope(r *http.Request) []string {
- if k8s.ForceNamespaceScope {
- target := k8s.GetNamespaceScopeTarget()
- if target == "" {
- return []string{}
- }
- return s.getUserNamespaces(r, []string{target})
- }
- return s.getUserNamespaces(r, nil)
-}
-
-func (s *Server) cnpgOperatorDeployments(r *http.Request, cache *k8s.ResourceCache, scope []string) (cnpgKindAccess, []string, []*appsv1.Deployment) {
- acc, denied, read := s.cnpgTypedScope(r, cache, scope, "apps", "deployments")
- if acc.state == cnpgCoverageDenied || acc.state == cnpgCoverageError {
- return acc, denied, nil
- }
- lister := cache.Deployments()
- if lister == nil || !cache.IsKindReady("deployments") {
- return cnpgKindAccess{state: cnpgCoverageSyncing}, nil, nil
- }
- var out []*appsv1.Deployment
- if read == nil {
- out, _ = lister.List(labels.Everything())
- } else {
- for _, ns := range read {
- items, _ := lister.Deployments(ns).List(labels.Everything())
- out = append(out, items...)
- }
- }
- sort.Slice(out, func(i, j int) bool {
- if out[i].Namespace != out[j].Namespace {
- return out[i].Namespace < out[j].Namespace
- }
- return out[i].Name < out[j].Name
- })
- return acc, denied, out
-}
-
-func (s *Server) cnpgOperatorServices(r *http.Request, cache *k8s.ResourceCache, scope []string) (cnpgKindAccess, []string, []*corev1.Service) {
- acc, denied, read := s.cnpgTypedScope(r, cache, scope, "", "services")
- if acc.state == cnpgCoverageDenied || acc.state == cnpgCoverageError {
- return acc, denied, nil
- }
- lister := cache.Services()
- if lister == nil || !cache.IsKindReady("services") {
- return cnpgKindAccess{state: cnpgCoverageSyncing}, nil, nil
- }
- hasPlugin, err := labels.Parse(cnpgPluginNameLabel)
- if err != nil {
- log.Printf("[cnpg] Failed to build plugin selector: %v", err)
- return cnpgKindAccess{state: cnpgCoverageError}, nil, nil
- }
- var out []*corev1.Service
- if read == nil {
- out, _ = lister.List(hasPlugin)
- } else {
- for _, ns := range read {
- items, _ := lister.Services(ns).List(hasPlugin)
- out = append(out, items...)
- }
- }
- return acc, denied, out
-}
-
-func cnpgOperatorContainerOf(d *appsv1.Deployment) *corev1.Container {
- for i := range d.Spec.Template.Spec.Containers {
- if d.Spec.Template.Spec.Containers[i].Name == cnpgOperatorContainer {
- return &d.Spec.Template.Spec.Containers[i]
- }
- }
- return firstContainer(d)
-}
-
-func firstContainer(d *appsv1.Deployment) *corev1.Container {
- if len(d.Spec.Template.Spec.Containers) == 0 {
- return nil
- }
- return &d.Spec.Template.Spec.Containers[0]
-}
-
-func cnpgOperatorComponent(d *appsv1.Deployment, role, pluginName string, c *corev1.Container) CNPGOperatorComponent {
- out := CNPGOperatorComponent{
- Role: role,
- PluginName: pluginName,
- Namespace: d.Namespace,
- Deployment: d.Name,
- Replicas: d.Spec.Replicas,
- }
- if c != nil {
- out.Image = c.Image
- out.Version = imageTag(c.Image)
- }
- if out.Version == "" {
- out.Version = d.Labels[cnpgVersionLabel]
- }
- if out.Version == "" {
- out.Version = d.Spec.Template.Labels[cnpgVersionLabel]
- }
- // The typed status cannot tell an omitted readyReplicas from zero; a status
- // the controller has observed at least once states it authoritatively.
- if d.Status.ObservedGeneration > 0 {
- ready := d.Status.ReadyReplicas
- out.ReadyReplicas = &ready
- }
- return out
-}
-
-// cnpgOperatorArg returns the value of --flag=value or --flag value from a
-// container's command and args.
-func cnpgOperatorArg(c *corev1.Container, flag string) string {
- argv := append(append([]string{}, c.Command...), c.Args...)
- for i, a := range argv {
- if v, ok := strings.CutPrefix(a, flag+"="); ok {
- return v
- }
- if a == flag && i+1 < len(argv) {
- return argv[i+1]
- }
- }
- return ""
-}
-
-func cnpgOperatorEnv(c *corev1.Container, name string) string {
- for _, e := range c.Env {
- if e.Name == name && e.ValueFrom == nil {
- return e.Value
- }
- }
- return ""
-}
-
-// cnpgOperatorExpand resolves $(OPERATOR_DEPLOYMENT_NAME) the way the kubelet
-// would: from the container's literal env, which the shipped manifests set to
-// the Deployment's own name. Any other reference is left verbatim, as the
-// kubelet leaves an unresolvable one.
-func cnpgOperatorExpand(v string, c *corev1.Container, d *appsv1.Deployment) string {
- ref := "$(" + cnpgOperatorDeployVar + ")"
- if !strings.Contains(v, ref) {
- return v
- }
- name := cnpgOperatorEnv(c, cnpgOperatorDeployVar)
- if name == "" {
- name = d.Name
- }
- return strings.ReplaceAll(v, ref, name)
-}
-
-func (s *Server) cnpgOperatorConfig(r *http.Request, cache *k8s.ResourceCache, operators []*appsv1.Deployment) []CNPGOperatorConfigRef {
- out := []CNPGOperatorConfigRef{}
- seen := map[string]bool{}
- add := func(ref CNPGOperatorConfigRef) {
- key := ref.Kind + "\x00" + ref.Namespace + "\x00" + ref.Name + "\x00" + ref.Purpose
- if ref.Name == "" || seen[key] {
- return
- }
- seen[key] = true
- out = append(out, ref)
- }
- for _, d := range operators {
- c := cnpgOperatorContainerOf(d)
- if c == nil {
- continue
- }
- if name := cnpgOperatorExpand(cnpgOperatorArg(c, "--config-map-name"), c, d); name != "" {
- add(s.cnpgOperatorConfigMap(r, cache, d.Namespace, name, cnpgConfigPurposeOperator))
- }
- if name := cnpgOperatorExpand(cnpgOperatorArg(c, "--secret-name"), c, d); name != "" {
- add(CNPGOperatorConfigRef{Kind: "Secret", Namespace: d.Namespace, Name: name, Purpose: cnpgConfigPurposeOperator})
- }
- if name := cnpgOperatorEnv(c, cnpgMonitoringQueriesCM); name != "" {
- add(s.cnpgOperatorConfigMap(r, cache, d.Namespace, name, cnpgConfigPurposeMonitoring))
- }
- }
- return out
-}
-
-func (s *Server) cnpgOperatorConfigMap(r *http.Request, cache *k8s.ResourceCache, namespace, name, purpose string) CNPGOperatorConfigRef {
- ref := CNPGOperatorConfigRef{Kind: "ConfigMap", Namespace: namespace, Name: name, Purpose: purpose}
- state := &CNPGOperatorConfigMapState{}
- ref.CNPGOperatorConfigMapState = state
- if !s.canRead(r, "", "configmaps", namespace, "get") {
- state.Reason = "no permission to get ConfigMaps in " + namespace
- return ref
- }
- lister := cache.ConfigMaps()
- if lister == nil {
- state.Reason = "ConfigMaps are still loading"
- return ref
- }
- if !capacityCacheCoversNamespace(cache, "configmaps", namespace) {
- state.Reason = "Radar does not watch ConfigMaps in " + namespace
- return ref
- }
- cm, err := lister.ConfigMaps(namespace).Get(name)
- switch {
- case apierrors.IsNotFound(err):
- exists := false
- state.Exists = &exists
- state.Reason = "not found"
- return ref
- case err != nil:
- log.Printf("[cnpg] Failed to read ConfigMap %s/%s: %v", namespace, name, err)
- state.Reason = "could not read the ConfigMap"
- return ref
- }
- exists := true
- state.Exists = &exists
- state.Readable = true
- state.Data = map[string]string{}
- for k, v := range cm.Data {
- state.Data[k] = v
- }
- return ref
-}
diff --git a/internal/server/cnpg_operator_status.go b/internal/server/cnpg_operator_status.go
new file mode 100644
index 0000000000..fc347c047f
--- /dev/null
+++ b/internal/server/cnpg_operator_status.go
@@ -0,0 +1,25 @@
+package server
+
+import (
+ "net/http"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+)
+
+// handleCNPGOperatorStatus serves GET /api/cnpg/operator/status?namespaces=a,b:
+// for each namespace, whether the operator that watches it is leading and
+// whether its admission webhook can answer. Cheap enough for the fleet and
+// cluster pages; the Operator screen has the full diagnosis.
+func (s *Server) handleCNPGOperatorStatus(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ resp := cnpgsvc.CNPGOperatorStatusResponse{Namespaces: map[string]cnpgsvc.CNPGOperatorVerdict{}}
+ wanted := parseNamespaces(r.URL.Query())
+ if len(wanted) == 0 {
+ s.writeJSON(w, resp)
+ return
+ }
+ resp = s.cnpgReader(r).OperatorStatus(r.Context(), wanted)
+ s.writeJSON(w, resp)
+}
diff --git a/internal/server/cnpg_operator_status_test.go b/internal/server/cnpg_operator_status_test.go
new file mode 100644
index 0000000000..2835cf0f4b
--- /dev/null
+++ b/internal/server/cnpg_operator_status_test.go
@@ -0,0 +1,32 @@
+package server
+
+import (
+ "encoding/json"
+ "io"
+ "net/http"
+ "testing"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+)
+
+func TestCNPGOperatorStatusNamespaceParameters(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds)
+ for _, query := range []string{"?namespace=db", "?namespaces=db"} {
+ resp, err := http.Get(testServer.URL + "/api/cnpg/operator/status" + query)
+ if err != nil {
+ t.Fatal(err)
+ }
+ body, err := io.ReadAll(resp.Body)
+ resp.Body.Close()
+ if err != nil || resp.StatusCode != http.StatusOK {
+ t.Fatalf("%s: %d, %s, %v", query, resp.StatusCode, body, err)
+ }
+ var result cnpgsvc.CNPGOperatorStatusResponse
+ if err := json.Unmarshal(body, &result); err != nil {
+ t.Fatal(err)
+ }
+ if _, present := result.Namespaces["db"]; !present {
+ t.Fatalf("%s omitted the requested namespace: %s", query, body)
+ }
+ }
+}
diff --git a/internal/server/cnpg_operator_test.go b/internal/server/cnpg_operator_test.go
index b1f8e49af6..b2e779c326 100644
--- a/internal/server/cnpg_operator_test.go
+++ b/internal/server/cnpg_operator_test.go
@@ -13,6 +13,8 @@ import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
)
@@ -151,7 +153,7 @@ func seedFullCNPGOperator(t *testing.T) {
)
}
-func readCNPGOperator(t *testing.T, resp *http.Response) (CNPGOperatorResponse, []byte) {
+func readCNPGOperator(t *testing.T, resp *http.Response) (cnpgsvc.CNPGOperatorResponse, []byte) {
t.Helper()
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
@@ -161,14 +163,14 @@ func readCNPGOperator(t *testing.T, resp *http.Response) (CNPGOperatorResponse,
if resp.StatusCode != http.StatusOK {
t.Fatalf("status = %d, want 200: %s", resp.StatusCode, body)
}
- var out CNPGOperatorResponse
+ var out cnpgsvc.CNPGOperatorResponse
if err := json.Unmarshal(body, &out); err != nil {
t.Fatalf("decode: %v", err)
}
return out, body
}
-func getCNPGOperatorNoAuth(t *testing.T, query string) (CNPGOperatorResponse, []byte) {
+func getCNPGOperatorNoAuth(t *testing.T, query string) (cnpgsvc.CNPGOperatorResponse, []byte) {
t.Helper()
resp, err := http.Get(testServer.URL + "/api/cnpg/operator" + query)
if err != nil {
@@ -177,7 +179,7 @@ func getCNPGOperatorNoAuth(t *testing.T, query string) (CNPGOperatorResponse, []
return readCNPGOperator(t, resp)
}
-func findConfigRef(refs []CNPGOperatorConfigRef, kind, purpose string) *CNPGOperatorConfigRef {
+func findConfigRef(refs []cnpgsvc.CNPGOperatorConfigRef, kind, purpose string) *cnpgsvc.CNPGOperatorConfigRef {
for i := range refs {
if refs[i].Kind == kind && refs[i].Purpose == purpose {
return &refs[i]
@@ -191,7 +193,7 @@ func TestCNPGOperator_DiscoversOperatorPluginAndConfig(t *testing.T) {
got, body := getCNPGOperatorNoAuth(t, "")
for _, key := range []string{"deployments", "services"} {
- if got.Coverage[key].State != cnpgCoverageFull {
+ if got.Coverage[key].State != integration.KindCoverageFull {
t.Errorf("coverage[%s] = %+v, want full", key, got.Coverage[key])
}
}
@@ -326,6 +328,33 @@ func TestCNPGOperator_ConfigMapDataNeedsGet(t *testing.T) {
}
}
+func TestCNPGOperatorComponentPodsDenied(t *testing.T) {
+
+ seedFullCNPGOperator(t)
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"cnpg-system"}}
+ allow(perms, "apps", "deployments", "", true)
+ allow(perms, "", "services", "", true)
+ allow(perms, "", "pods", "cnpg-system", false)
+ env.srv.permCache.Set("no-pods", nil, perms)
+ got, body := readCNPGOperator(t, env.authGet(t, "/api/cnpg/operator", "no-pods", ""))
+ if len(got.Components) != 2 {
+ t.Fatalf("components = %+v, want both despite the Pods denial", got.Components)
+ }
+ var raw struct {
+ Components []map[string]any `json:"components"`
+ }
+ _ = json.Unmarshal(body, &raw)
+ for i, c := range got.Components {
+ if c.PodCoverage == nil || c.PodCoverage.State != "denied" || c.PodCoverage.Grant == nil || c.PodCoverage.Grant.Resource != "pods" {
+ t.Errorf("%s podCoverage = %+v", c.Deployment, c.PodCoverage)
+ }
+ if v, ok := raw.Components[i]["pods"]; !ok || v != nil {
+ t.Errorf("%s pods JSON = %v, want null", c.Deployment, v)
+ }
+ }
+}
+
func TestCNPGOperator_DeniedDeploymentsWithholdComponents(t *testing.T) {
seedFullCNPGOperator(t)
env := newAuthTestServer(t)
@@ -339,7 +368,7 @@ func TestCNPGOperator_DeniedDeploymentsWithholdComponents(t *testing.T) {
got, _ := readCNPGOperator(t, env.authGet(t, "/api/cnpg/operator", "partial", ""))
cov := got.Coverage["deployments"]
- if cov.State != cnpgCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "cnpg-system" {
+ if cov.State != integration.KindCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "cnpg-system" {
t.Errorf("deployments coverage = %+v, want partial denied [cnpg-system]", cov)
}
for _, c := range got.Components {
@@ -359,7 +388,7 @@ func TestCNPGOperator_DeniedDeploymentsWithholdComponents(t *testing.T) {
env.srv.permCache.Set("none", nil, none)
got, _ = readCNPGOperator(t, env.authGet(t, "/api/cnpg/operator", "none", ""))
- if got.Coverage["deployments"].State != cnpgCoverageDenied || got.Coverage["services"].State != cnpgCoverageDenied {
+ if got.Coverage["deployments"].State != integration.KindCoverageDenied || got.Coverage["services"].State != integration.KindCoverageDenied {
t.Errorf("coverage = %+v, want both denied", got.Coverage)
}
if got.Components == nil || len(got.Components) != 0 || got.Config == nil {
diff --git a/internal/server/cnpg_pooler_actions.go b/internal/server/cnpg_pooler_actions.go
new file mode 100644
index 0000000000..2dab9fd8d5
--- /dev/null
+++ b/internal/server/cnpg_pooler_actions.go
@@ -0,0 +1,93 @@
+package server
+
+import (
+ "fmt"
+ "log"
+ "net/http"
+ "strings"
+
+ "github.com/go-chi/chi/v5"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+)
+
+// Pooler detail and pause/resume. spec.pgbouncer.paused is desired state: the
+// pooler's instance manager applies it to each PgBouncer with PAUSE/RESUME.
+// Whether a PgBouncer is paused is observed separately (SHOW STATE over the
+// caller's pods/exec); the exporter does not publish it.
+
+func (s *Server) handleCNPGPoolerCapabilities(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ dyn, contextName, typed := reader.dynamic, reader.actionContext, reader.Clients.Typed
+ if dyn == nil || typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ resp, err := reader.PoolerCapabilities(r.Context(), cnpgsvc.ActionClients{Dynamic: dyn, Typed: typed}, contextName, namespace, name)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "capabilities", namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+func (s *Server) handleCNPGPoolerAction(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name, action := chi.URLParam(r, "namespace"), chi.URLParam(r, "name"), chi.URLParam(r, "action")
+ if action != "pause" && action != "resume" {
+ s.writeError(w, http.StatusBadRequest, fmt.Sprintf("unknown Pooler action %q: must be pause or resume", action))
+ return
+ }
+ req, _, ok := s.decodeActionRequest(w, r)
+ if !ok {
+ return
+ }
+ errAction := "pooler" + strings.ToUpper(action[:1]) + action[1:]
+ clients, err := s.cnpgActionClients(r, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, errAction, namespace, name)
+ return
+ }
+ auth.AuditLog(r, namespace, name)
+ res, err := cnpgsvc.RunCNPGPoolerAction(r.Context(), clients, namespace, name, action, req)
+ if err != nil {
+ s.writeCNPGActionError(w, err, errAction, namespace, name)
+ return
+ }
+ log.Printf("[cnpg] %s on Pooler %s/%s requested", action, sanitizeForLog(namespace), sanitizeForLog(name))
+ s.writeJSON(w, res)
+}
+
+// ---------- observed PgBouncer state ----------
+
+func (s *Server) handleCNPGPgBouncerState(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.authorizeCNPGRuntime(w, r, namespace, "poolers") {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ pooler, err := reader.Pooler(r.Context(), cache, namespace, name)
+ pooler, err = cnpgCachedResourceResult(pooler, err, "Pooler", namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ resp, err := reader.PgBouncerState(r.Context(), cache, pooler)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
diff --git a/internal/server/cnpg_protection.go b/internal/server/cnpg_protection.go
new file mode 100644
index 0000000000..eb073abc01
--- /dev/null
+++ b/internal/server/cnpg_protection.go
@@ -0,0 +1,87 @@
+package server
+
+import (
+ "errors"
+ "net/http"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
+)
+
+// These previews use caller clients and server dry-run. They do not persist
+// configuration, read Secrets or probe remote object storage.
+func (s *Server) handleCNPGArchivingPreview(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ var req struct {
+ ReviewedContext string `json:"reviewedContext"`
+ cnpgsvc.ArchivingParams
+ }
+ if err := decodeBoundedJSONBody(w, r, integration.ActionBodyLimit, &req); err != nil {
+ var tooLarge *http.MaxBytesError
+ status := http.StatusBadRequest
+ if errors.As(err, &tooLarge) {
+ status = http.StatusRequestEntityTooLarge
+ }
+ s.writeError(w, status, "invalid preview request: "+err.Error())
+ return
+ }
+ reader := s.cnpgReader(r)
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if reader.dynamic == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available")
+ return
+ }
+ if err := integration.CheckReviewedContext(req.ReviewedContext, reader.actionContext); err != nil {
+ s.writeCNPGActionError(w, err, "configureArchiving", namespace, name)
+ return
+ }
+ out, err := reader.PreviewArchiving(r.Context(), cnpgsvc.ActionClients{Dynamic: reader.dynamic}, reader.actionContext, namespace, name, req.ArchivingParams)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "configureArchiving", namespace, name)
+ return
+ }
+ s.writeJSON(w, out)
+}
+
+func (s *Server) handleCNPGScheduleMethodPreview(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ reader := s.cnpgReader(r)
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if reader.dynamic == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available")
+ return
+ }
+ out, err := reader.PreviewScheduleMethod(r.Context(), cnpgsvc.ActionClients{Dynamic: reader.dynamic}, reader.actionContext, namespace, name)
+ if err != nil {
+ s.writeCNPGActionError(w, err, "repairMethod", namespace, name)
+ return
+ }
+ s.writeJSON(w, out)
+}
+
+// A draft has no ScheduledBackup lastCheckTime. Read its target Cluster as
+// the caller and count from now rather than inventing an existing schedule.
+func (s *Server) handleCNPGDraftSchedulePreview(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ reader := s.cnpgReader(r)
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if reader.dynamic == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available")
+ return
+ }
+ out, err := reader.PreviewDraftSchedule(r.Context(), reader.dynamic, namespace, name, r.URL.Query().Get("schedule"), time.Now())
+ if err != nil {
+ s.writeCNPGActionError(w, err, "schedule-preview", namespace, name)
+ return
+ }
+ s.writeJSON(w, out)
+}
diff --git a/internal/server/cnpg_protection_test.go b/internal/server/cnpg_protection_test.go
new file mode 100644
index 0000000000..e4d869ce9b
--- /dev/null
+++ b/internal/server/cnpg_protection_test.go
@@ -0,0 +1,57 @@
+package server
+
+import (
+ "io"
+ "net/http"
+ "strings"
+ "testing"
+
+ "github.com/skyhook-io/radar/internal/integration"
+)
+
+func TestCNPGArchivingPreviewRejectsInvalidAndOversizedBodies(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds)
+ for _, tc := range []struct {
+ name string
+ body string
+ status int
+ }{
+ {"malformed", `{`, http.StatusBadRequest},
+ {"unknown field", `{"reviewedContext":"ctx","objectStore":"store","serverName":"pg","extra":true}`, http.StatusBadRequest},
+ {"too large", `{"objectStore":"` + strings.Repeat("x", integration.ActionBodyLimit) + `"}`, http.StatusRequestEntityTooLarge},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ response, err := http.Post(testServer.URL+"/api/cnpg/clusters/db/pg/protection/preview", "application/json", strings.NewReader(tc.body))
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer response.Body.Close()
+ body, _ := io.ReadAll(response.Body)
+ if response.StatusCode != tc.status {
+ t.Fatalf("status=%d body=%s", response.StatusCode, body)
+ }
+ })
+ }
+}
+
+func TestCNPGDraftSchedulePreviewReadsTargetAndUsesNow(t *testing.T) {
+ cluster := cnpgObj("postgresql.cnpg.io/v1", "Cluster", "setup", "pg", map[string]any{}, nil)
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cluster)
+ for _, tc := range []struct {
+ name string
+ status int
+ }{{"pg", http.StatusOK}, {"missing", http.StatusNotFound}} {
+ response, err := http.Get(testServer.URL + "/api/cnpg/clusters/setup/" + tc.name + "/schedule-preview?schedule=0%200%202%20*%20*%20*")
+ if err != nil {
+ t.Fatal(err)
+ }
+ body, _ := io.ReadAll(response.Body)
+ response.Body.Close()
+ if response.StatusCode != tc.status {
+ t.Fatalf("status=%d body=%s", response.StatusCode, body)
+ }
+ if tc.status == http.StatusOK && (!strings.Contains(string(body), `"basis":"now"`) || !strings.Contains(string(body), `"valid":true`)) {
+ t.Fatalf("draft preview=%s", body)
+ }
+ }
+}
diff --git a/internal/server/cnpg_reads.go b/internal/server/cnpg_reads.go
new file mode 100644
index 0000000000..1324064b23
--- /dev/null
+++ b/internal/server/cnpg_reads.go
@@ -0,0 +1,68 @@
+package server
+
+import (
+ "errors"
+ "fmt"
+ "log"
+ "net/http"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func (s *Server) authorizeCNPGCachedRead(r *http.Request, namespace, resource string, extra ...Grant) error {
+ if !k8s.IsConnected() {
+ return cnpgsvc.ErrCNPGDisconnected
+ }
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ return &cnpgsvc.ReadFailure{http.StatusForbidden, "no access to namespace " + namespace}
+ }
+ if !s.canRead(r, cnpgsvc.Group, resource, namespace, "get") {
+ return &cnpgsvc.ReadFailure{http.StatusForbidden, "no access to " + resource + ".postgresql.cnpg.io in namespace " + namespace}
+ }
+ for _, g := range extra {
+ var allowed bool
+ if g.Subresource == "" {
+ allowed = s.canRead(r, g.Group, g.Resource, namespace, g.Verb)
+ } else {
+ allowed = s.canReadSubresource(r, g.Group, g.Resource, g.Subresource, namespace, g.Verb)
+ }
+ if !allowed {
+ return &cnpgsvc.ReadFailure{http.StatusForbidden, "no access to " + g.Resource + " in namespace " + namespace}
+ }
+ }
+ return nil
+}
+
+func cnpgCachedResourceResult(object *unstructured.Unstructured, err error, kind, namespace, name string) (*unstructured.Unstructured, error) {
+ switch {
+ case err == nil && object != nil:
+ return object, nil
+ case err == nil, errors.Is(err, k8s.ErrUnknownDynamicKind):
+ return nil, &cnpgsvc.ReadFailure{http.StatusNotFound, "CloudNativePG " + kind + " " + namespace + "/" + name + " not found"}
+ case errors.Is(err, integration.ErrDynamicNotSynced):
+ return nil, &cnpgsvc.ReadFailure{http.StatusServiceUnavailable, "CloudNativePG " + kind + "s are still syncing"}
+ default:
+ return nil, fmt.Errorf("%w: %w", &cnpgsvc.ReadFailure{http.StatusInternalServerError, "failed to read CloudNativePG " + kind}, err)
+ }
+}
+
+func (s *Server) writeCNPGCachedReadError(w http.ResponseWriter, err error, namespace, name string) {
+ if errors.Is(err, cnpgsvc.ErrCNPGDisconnected) {
+ s.writeNotConnected(w)
+ return
+ }
+ var failure *cnpgsvc.ReadFailure
+ if errors.As(err, &failure) {
+ if failure.Status == http.StatusInternalServerError {
+ log.Printf("[cnpg] Failed to read %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ }
+ s.writeError(w, failure.Status, failure.Message)
+ return
+ }
+ log.Printf("[cnpg] Failed to read %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ s.writeError(w, http.StatusInternalServerError, "failed to read CloudNativePG data")
+}
diff --git a/internal/server/cnpg_recovery.go b/internal/server/cnpg_recovery.go
new file mode 100644
index 0000000000..6e62a8b2ea
--- /dev/null
+++ b/internal/server/cnpg_recovery.go
@@ -0,0 +1,103 @@
+package server
+
+import (
+ "fmt"
+ "log"
+ "net/http"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/util/validation"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+// Restore follow-through and the restore-validation record. The recovery
+// snapshot is what a person watching a new Cluster bootstrap from backups
+// needs: the Cluster's phase, the Pods doing the recovery (the full-recovery
+// Job's Pod and its init containers, then the instances), their Jobs and the
+// Warning events about them. Every read uses the caller's identity; a read the
+// caller may not make is reported as coverage, never as "nothing there".
+
+func (s *Server) handleCNPGClusterRecovery(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ dyn, typed := reader.dynamic, reader.Clients.Typed
+ if dyn == nil || typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ cluster, err := dyn.Resource(cnpgsvc.ClusterGVR).Namespace(namespace).Get(r.Context(), name, metav1.GetOptions{})
+ if err != nil {
+ s.writeCNPGReadError(w, err, cnpgsvc.GrantGetCluster, namespace, name)
+ return
+ }
+ resp := reader.RecoverySnapshot(r.Context(), typed, cluster)
+ s.writeJSON(w, resp)
+}
+
+func (s *Server) writeCNPGReadError(w http.ResponseWriter, err error, g Grant, namespace, name string) {
+ switch {
+ case apierrors.IsNotFound(err):
+ s.writeError(w, http.StatusNotFound, fmt.Sprintf("CloudNativePG Cluster %s/%s not found", namespace, name))
+ case apierrors.IsForbidden(err):
+ s.writeError(w, http.StatusForbidden, "This needs "+g.In(namespace).String())
+ default:
+ log.Printf("[cnpg] Failed to read Cluster %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ s.writeError(w, http.StatusInternalServerError, "failed to read CloudNativePG Cluster: "+err.Error())
+ }
+}
+
+// handleCNPGRestoreValidation serves POST /api/cnpg/clusters/{ns}/{name}/restore-validation:
+// records the caller's validation note on the restored Cluster as an
+// annotation, with an impersonated merge patch bound to the reviewed context
+// and Cluster UID.
+func (s *Server) handleCNPGRestoreValidation(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ req, dyn, ok := s.decodeActionRequest(w, r)
+ if !ok {
+ return
+ }
+ auth.AuditLog(r, namespace, name)
+ if s.grantPermission(r, cnpgsvc.GrantPatchClusters.In(namespace)) == integration.PermissionDenied {
+ s.writeError(w, http.StatusForbidden, "Recording a validation note needs "+cnpgsvc.GrantPatchClusters.In(namespace).String())
+ return
+ }
+ recordedBy := ""
+ if user := auth.UserFromContext(r.Context()); user != nil {
+ recordedBy = user.Username
+ }
+ note, err := cnpgsvc.RecordCNPGRestoreValidation(r.Context(), dyn, namespace, name, req, recordedBy, time.Now())
+ if err != nil {
+ s.writeCNPGActionError(w, err, "restore-validation", namespace, name)
+ return
+ }
+ log.Printf("[cnpg] restore validation recorded on Cluster %s/%s", sanitizeForLog(namespace), sanitizeForLog(name))
+ s.writeJSON(w, note)
+}
+
+// handleCNPGRestoreCapability answers whether the caller may create the
+// restored Cluster in a namespace: create clusters there, refused like other
+// webhook-bound writes while the operator's webhook rejects them. The restore
+// dialog asks it whichever way it was opened (Cluster, Backup, ObjectStore).
+func (s *Server) handleCNPGRestoreCapability(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := r.URL.Query().Get("namespace")
+ if errs := validation.IsDNS1123Label(namespace); namespace == "" || len(errs) > 0 {
+ s.writeError(w, http.StatusBadRequest, "namespace must be a valid namespace name")
+ return
+ }
+ s.writeJSON(w, s.cnpgReader(r).RestoreCapability(r.Context(), namespace))
+}
diff --git a/internal/server/cnpg_recovery_test.go b/internal/server/cnpg_recovery_test.go
new file mode 100644
index 0000000000..79320758f4
--- /dev/null
+++ b/internal/server/cnpg_recovery_test.go
@@ -0,0 +1,66 @@
+package server
+
+import (
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+
+ "encoding/json"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "testing"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ "github.com/skyhook-io/radar/internal/auth"
+)
+
+func TestParseCNPGReportOptions(t *testing.T) {
+ for q, ok := range map[string]bool{"": true, "logs=true&tailLines=500": true, "queryText=true": false, "logs=true&tailLines=0": false, "logs=true&tailLines=99999": false} {
+ _, err := parseCNPGReportOptions(httptest.NewRequest(http.MethodGet, "/?"+q, nil))
+ if (err == nil) != ok {
+ t.Errorf("%q: err=%v", q, err)
+ }
+ }
+}
+
+func apiForbidden(resource string) error {
+ return apierrors.NewForbidden(schema.GroupResource{Resource: resource}, "", errors.New("denied"))
+}
+
+func TestHandleCNPGRestoreCapability(t *testing.T) {
+ for _, allowed := range []bool{false, true} {
+ srv := &Server{permCache: auth.NewPermissionCache()}
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"db"}}
+ perms.SetCanI("create", cnpgsvc.Group, "clusters", "db", allowed)
+ srv.permCache.Set("alice", nil, perms)
+ r := httptest.NewRequest(http.MethodGet, "/api/cnpg/restore/capability?namespace=db", nil)
+ r = r.WithContext(auth.ContextWithUser(r.Context(), &auth.User{Username: "alice"}))
+ w := httptest.NewRecorder()
+ srv.handleCNPGRestoreCapability(w, r)
+ if w.Code != http.StatusOK {
+ t.Fatalf("status %d: %s", w.Code, w.Body.String())
+ }
+ var got integration.ActionCapability
+ if err := json.Unmarshal(w.Body.Bytes(), &got); err != nil {
+ t.Fatal(err)
+ }
+ if allowed && (!got.Allowed || got.Permission != integration.PermissionAllowed) {
+ t.Errorf("allowed = %+v", got)
+ }
+ if !allowed && (got.Allowed || got.Permission != integration.PermissionDenied || got.Grant == nil || *got.Grant != (auth.Grant{Verb: "create", Group: cnpgsvc.Group, Resource: "clusters", Namespace: "db"}) || got.Grant.String() != "create clusters (postgresql.cnpg.io) in namespace db" || !strings.Contains(got.Reason, got.Grant.String())) {
+ t.Errorf("denied = %+v, want the grant named", got)
+ }
+ }
+ for _, q := range []string{"", "?namespace=", "?namespace=Not_A_Namespace"} {
+ w := httptest.NewRecorder()
+ (&Server{}).handleCNPGRestoreCapability(w, httptest.NewRequest(http.MethodGet, "/api/cnpg/restore/capability"+q, nil))
+ if w.Code != http.StatusBadRequest {
+ t.Errorf("%q: status %d, want 400", q, w.Code)
+ }
+ }
+}
diff --git a/internal/server/cnpg_report.go b/internal/server/cnpg_report.go
new file mode 100644
index 0000000000..677b27dda6
--- /dev/null
+++ b/internal/server/cnpg_report.go
@@ -0,0 +1,94 @@
+package server
+
+import (
+ "context"
+ "errors"
+ "fmt"
+ "log"
+ "net/http"
+ "strconv"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/version"
+)
+
+// The Cluster report bundle mirrors `kubectl cnpg report cluster`: the
+// Cluster, its Pods, Jobs, PVCs and events as manifests, optionally the Pods'
+// logs. Radar adds what it already knows (Backups, ScheduledBackups, Poolers,
+// the ObjectStore, operator and plugin versions, the runtime and storage
+// snapshots) and a coverage record of everything it could not read and why.
+// Secret values are never read; the bundle lists the Secrets the cluster
+// references by name. Logs, and the query text inside them, are opt-in.
+
+func parseCNPGReportOptions(r *http.Request) (cnpgsvc.ReportOptions, error) {
+ q := r.URL.Query()
+ opts := cnpgsvc.ReportOptions{Logs: q.Get("logs") == "true", QueryText: q.Get("queryText") == "true", TailLines: cnpgsvc.ReportDefaultTail}
+ if v := q.Get("tailLines"); v != "" {
+ n, err := strconv.ParseInt(v, 10, 64)
+ if err != nil || n < 1 || n > cnpgsvc.ReportMaxTail {
+ return opts, fmt.Errorf("tailLines must be between 1 and %d", cnpgsvc.ReportMaxTail)
+ }
+ opts.TailLines = n
+ }
+ if opts.QueryText && !opts.Logs {
+ return opts, errors.New("queryText applies to logs; set logs=true as well")
+ }
+ if !opts.Logs {
+ opts.TailLines = 0
+ }
+ return opts, nil
+}
+
+// handleCNPGClusterReport serves GET /api/cnpg/clusters/{ns}/{name}/report as a
+// zip download. Only the Cluster read itself can fail the request; every other
+// read is skipped and recorded in report.json.
+// Direct reads use the caller's impersonated clients; the namespace sentinel
+// applies only to the shared-cache snapshots added to the bundle.
+func (s *Server) handleCNPGClusterReport(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ opts, err := parseCNPGReportOptions(r)
+ if err != nil {
+ s.writeError(w, http.StatusBadRequest, err.Error())
+ return
+ }
+ reader := s.cnpgReader(r)
+ dyn, contextName, typed := reader.dynamic, reader.actionContext, reader.Clients.Typed
+ if dyn == nil || typed == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ ctx, cancel := context.WithTimeout(r.Context(), cnpgsvc.ReportTimeout)
+ defer cancel()
+ cluster, err := dyn.Resource(cnpgsvc.ClusterGVR).Namespace(namespace).Get(ctx, name, metav1.GetOptions{})
+ if err != nil {
+ s.writeCNPGReadError(w, err, cnpgsvc.GrantGetCluster, namespace, name)
+ return
+ }
+ auth.AuditLog(r, namespace, name)
+
+ meta := cnpgsvc.ReportMetadata{Context: contextName, RadarVersion: version.Current, Now: time.Now().UTC()}
+ if user := auth.UserFromContext(r.Context()); user != nil {
+ meta.RequestedBy = user.Username
+ }
+ data, root, err := reader.Report(ctx, dyn, typed, cluster, opts, meta)
+ if err != nil {
+ log.Printf("[cnpg] Failed to build report for %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ s.writeError(w, http.StatusInternalServerError, "failed to build the report")
+ return
+ }
+ w.Header().Set("Content-Type", "application/zip")
+ w.Header().Set("Content-Disposition", fmt.Sprintf(`attachment; filename="%s.zip"`, root))
+ w.Header().Set("Content-Length", strconv.Itoa(len(data)))
+ w.WriteHeader(http.StatusOK)
+ if _, err := w.Write(data); err != nil {
+ log.Printf("[cnpg] Failed to send report for %s/%s: %v", sanitizeForLog(namespace), sanitizeForLog(name), err)
+ }
+}
diff --git a/internal/server/cnpg_runtime.go b/internal/server/cnpg_runtime.go
new file mode 100644
index 0000000000..a8ecf6a29f
--- /dev/null
+++ b/internal/server/cnpg_runtime.go
@@ -0,0 +1,102 @@
+package server
+
+import (
+ "log"
+ "net/http"
+ "sort"
+ "strings"
+
+ "github.com/go-chi/chi/v5"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// authorizeCNPGRuntime gates like the logs endpoint — namespace, the owning
+// CNPG object, listing Pods — before the object is looked up. pods/proxy is
+// deliberately not part of the gate: without it the Kubernetes facts still
+// render and each source reports itself denied.
+func (s *Server) authorizeCNPGRuntime(w http.ResponseWriter, r *http.Request, namespace, resource string) bool {
+ if err := s.authorizeCNPGCachedRead(r, namespace, resource, cnpgsvc.GrantListPods); err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, chi.URLParam(r, "name"))
+ return false
+ }
+ return true
+}
+
+// cnpgRuntimeClient builds the caller's client for the proxy reads: the
+// impersonated identity when auth is on (nil when impersonation fails — never
+// Radar's own identity), with redirects refused so a proxied answer cannot
+// steer the request to another path on the Pod.
+func cnpgRuntimeClient(r *http.Request) kubernetes.Interface {
+ cfg := k8s.ConfigFromContext(r.Context())
+ if cfg == nil {
+ return nil
+ }
+ hc, err := rest.HTTPClientFor(cfg)
+ if err != nil {
+ log.Printf("[cnpg] Failed to build runtime HTTP client: %v", err)
+ return nil
+ }
+ noRedirect := *hc
+ noRedirect.CheckRedirect = func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse }
+ client, err := kubernetes.NewForConfigAndClient(cfg, &noRedirect)
+ if err != nil {
+ log.Printf("[cnpg] Failed to build runtime client: %v", err)
+ return nil
+ }
+ return client
+}
+
+// handleCNPGClusterRuntime serves GET /api/cnpg/clusters/{namespace}/{name}/runtime:
+// each instance's /pg/status and exporter metrics, read through the caller's
+// own pods/proxy.
+func (s *Server) handleCNPGClusterRuntime(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ resp, err := s.cnpgReader(r).ClusterRuntime(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+// handleCNPGPoolerRuntime serves GET /api/cnpg/poolers/{namespace}/{name}/runtime:
+// each pooler Pod's PgBouncer exporter, read through the caller's pods/proxy.
+func (s *Server) handleCNPGPoolerRuntime(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.authorizeCNPGRuntime(w, r, namespace, "poolers") {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ pooler, err := reader.Pooler(r.Context(), cache, namespace, name)
+ pooler, err = cnpgCachedResourceResult(pooler, err, "Pooler", namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ resp, err := reader.PoolerRuntime(r.Context(), cache, pooler)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+func cnpgRuntimeIdentity(r *http.Request) string {
+ id := k8s.GetContextName()
+ if user := auth.UserFromContext(r.Context()); user != nil {
+ groups := append([]string(nil), user.Groups...)
+ sort.Strings(groups)
+ id += "\x00" + user.Username + "\x00" + strings.Join(groups, ",")
+ }
+ return id
+}
diff --git a/internal/server/cnpg_runtime_test.go b/internal/server/cnpg_runtime_test.go
new file mode 100644
index 0000000000..7312d223cc
--- /dev/null
+++ b/internal/server/cnpg_runtime_test.go
@@ -0,0 +1,625 @@
+package server
+
+import (
+ "context"
+ "encoding/json"
+ "fmt"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ appsv1 "k8s.io/api/apps/v1"
+ authv1 "k8s.io/api/authorization/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+ "k8s.io/client-go/kubernetes"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// Modeled on a CloudNativePG 1.27 primary's GET /pg/status.
+const cnpgStatusFixture = `{
+ "currentLsn": "0/7000148", "systemID": "7690724460905209884", "isPrimary": true,
+ "replayPaused": false, "pendingRestart": true, "isWalReceiverActive": false,
+ "pendingRestartForDecrease": true, "isPgRewindRunning": false, "instanceManagerVersion": "1.27.0",
+ "mightBeUnavailable": false, "isArchivingWAL": true,
+ "pod": {"metadata": {"name": "pg-orders-1"}},
+ "lastArchivedWAL": "000000010000000000000006", "lastArchivedWALTime": "2026-09-29T08:24:24.875131Z",
+ "lastFailedWALTime": "-infinity", "currentWAL": "000000010000000000000007", "readyWalFiles": 3,
+ "timeLineID": 1,
+ "replicationInfo": [{
+ "applicationName": "pg-orders-2", "state": "streaming",
+ "receivedLsn": "0/7000148", "writeLsn": "0/7000148", "flushLsn": "0/7000100", "replayLsn": "0/7000000",
+ "writeLag": "00:00:00.012", "flushLag": "1 day 00:00:01", "replayLag": "garbage",
+ "syncState": "async", "syncPriority": "0"
+ }],
+ "replicationSlotsInfo": [{"slotName": "_cnpg_pg_orders_2", "slotType": "physical", "restartLsn": "0/7000148", "walStatus": "reserved", "active": true}, {"slotName": "orders_sub", "plugin": "pgoutput", "slotType": "logical", "database": "app", "walStatus": "extended", "active": false}],
+ "pgStatBasebackupsInfo": [
+ {"usename": "streaming_replica", "application_name": "pg-orders-4-join", "backend_start": "2026-09-30T10:00:00.5Z", "phase": "streaming database files", "backup_total": 4000, "backup_streamed": 1000, "backup_total_pretty": "4000 bytes", "backup_streamed_pretty": "1000 bytes", "tablespaces_total": 1, "tablespaces_streamed": 0},
+ {"usename": "streaming_replica", "application_name": "pg-orders-5-join", "backend_start": "2026-09-30T10:01:00Z", "phase": "waiting for checkpoint to finish", "backup_total": 0, "backup_streamed": 0, "tablespaces_total": 0, "tablespaces_streamed": 0}
+ ]
+}`
+
+// Modeled on the CNPG 1.27 default monitoring queries. The exporter itself
+// connects as postgres with application_name cnpg_metrics_exporter.
+const cnpgMetricsFixture = `# HELP cnpg_backends_total Number of backends
+# TYPE cnpg_backends_total gauge
+cnpg_backends_total{application_name="cnpg_metrics_exporter",datname="app",state="active",usename="postgres"} 1
+cnpg_backends_total{application_name="pg-orders-2",datname="",state="active",usename="streaming_replica"} 1
+cnpg_backends_total{application_name="psql",datname="app",state="idle",usename="app"} 4
+cnpg_backends_total{application_name="api",datname="app",state="idle in transaction",usename="app"} 2
+cnpg_backends_total{application_name="bg",datname="app",state="",usename="app"} 7
+# TYPE cnpg_backends_waiting_total gauge
+cnpg_backends_waiting_total 1
+# TYPE cnpg_pg_postmaster_start_time gauge
+cnpg_pg_postmaster_start_time 1.790698960738028e+09
+# TYPE cnpg_backends_max_tx_duration_seconds gauge
+cnpg_backends_max_tx_duration_seconds{application_name="pg-orders-2",datname="",state="active",usename="streaming_replica"} 9000
+cnpg_backends_max_tx_duration_seconds{application_name="api",datname="app",state="idle in transaction",usename="app"} 42.5
+# TYPE cnpg_pg_database_size_bytes gauge
+cnpg_pg_database_size_bytes{datname="app"} 7.654547e+06
+cnpg_pg_database_size_bytes{datname="postgres"} 7.5e+06
+# TYPE cnpg_pg_database_xid_age gauge
+cnpg_pg_database_xid_age{datname="app"} 29
+# TYPE cnpg_pg_database_mxid_age gauge
+cnpg_pg_database_mxid_age{datname="app"} 12
+cnpg_pg_database_mxid_age{datname="reports"} 400000000
+# TYPE cnpg_pg_extensions_update_available gauge
+cnpg_pg_extensions_update_available{datname="app",default_version="1.0",extname="plpgsql",installed_version="1.0"} 0
+cnpg_pg_extensions_update_available{datname="app",default_version="3.5.1",extname="postgis",installed_version="3.4.2"} 1
+# TYPE cnpg_pg_settings_setting gauge
+cnpg_pg_settings_setting{name="max_connections"} 100
+cnpg_pg_settings_setting{name="shared_buffers"} 16384
+# TYPE cnpg_pg_stat_archiver_archived_count counter
+cnpg_pg_stat_archiver_archived_count 7
+# TYPE cnpg_pg_stat_archiver_failed_count counter
+cnpg_pg_stat_archiver_failed_count 0
+# TYPE cnpg_pg_stat_archiver_seconds_since_last_archival gauge
+cnpg_pg_stat_archiver_seconds_since_last_archival 63.05
+# TYPE cnpg_pg_stat_archiver_seconds_since_last_failure gauge
+cnpg_pg_stat_archiver_seconds_since_last_failure -1
+# TYPE cnpg_collector_pg_wal gauge
+cnpg_collector_pg_wal{value="size"} 1.34217728e+08
+cnpg_collector_pg_wal{value="volume_size"} NaN
+# TYPE cnpg_pg_stat_database_xact_commit counter
+cnpg_pg_stat_database_xact_commit{datname="app"} 100
+cnpg_pg_stat_database_xact_commit{datname="postgres"} 50
+# TYPE cnpg_pg_stat_database_xact_rollback counter
+cnpg_pg_stat_database_xact_rollback{datname="app"} 2
+# TYPE cnpg_pg_stat_database_blks_hit counter
+cnpg_pg_stat_database_blks_hit{datname="app"} 900
+# TYPE cnpg_pg_stat_database_blks_read counter
+cnpg_pg_stat_database_blks_read{datname="app"} 100
+# TYPE cnpg_pg_replication_slots_pg_wal_lsn_diff gauge
+cnpg_pg_replication_slots_pg_wal_lsn_diff{database="",slot_name="_cnpg_pg_orders_2",slot_type="physical"} 16384
+`
+
+const cnpgPoolerMetricsFixture = `# TYPE cnpg_pgbouncer_pools_cl_active gauge
+cnpg_pgbouncer_pools_cl_active{database="pgbouncer",user="pgbouncer"} 1
+cnpg_pgbouncer_pools_cl_active{database="app",user="app"} 5
+cnpg_pgbouncer_pools_cl_active{database="app",user="cnpg_pooler_pgbouncer"} 1
+# TYPE cnpg_pgbouncer_pools_cl_waiting gauge
+cnpg_pgbouncer_pools_cl_waiting{database="app",user="app"} 2
+# TYPE cnpg_pgbouncer_pools_sv_active gauge
+cnpg_pgbouncer_pools_sv_active{database="app",user="app"} 3
+# TYPE cnpg_pgbouncer_pools_sv_idle gauge
+cnpg_pgbouncer_pools_sv_idle{database="app",user="app"} 1
+# TYPE cnpg_pgbouncer_pools_sv_used gauge
+cnpg_pgbouncer_pools_sv_used{database="app",user="app"} 0
+# TYPE cnpg_pgbouncer_pools_maxwait gauge
+cnpg_pgbouncer_pools_maxwait{database="app",user="app"} 1
+# TYPE cnpg_pgbouncer_pools_maxwait_us gauge
+cnpg_pgbouncer_pools_maxwait_us{database="app",user="app"} 250000
+`
+
+func cnpgRuntimeWithUID(u *unstructured.Unstructured, uid string) *unstructured.Unstructured {
+ u.SetUID(types.UID(uid))
+ return u
+}
+
+func cnpgEqF(p *float64, want float64) bool { return p != nil && *p == want }
+
+// cnpgFakeProxyAPIServer stands in for the apiserver's pods/proxy. The handler
+// decides per (scheme, pod, port, path); every request is recorded.
+type cnpgFakeProxyAPIServer struct {
+ mu sync.Mutex
+ requests []string
+ headers []http.Header
+}
+
+func (f *cnpgFakeProxyAPIServer) record(r *http.Request) {
+ f.mu.Lock()
+ defer f.mu.Unlock()
+ f.requests = append(f.requests, r.Method+" "+r.URL.Path)
+ f.headers = append(f.headers, r.Header.Clone())
+}
+
+func (f *cnpgFakeProxyAPIServer) seen() []string {
+ f.mu.Lock()
+ defer f.mu.Unlock()
+ return append([]string(nil), f.requests...)
+}
+
+type cnpgProxyCall struct{ scheme, pod, port, path string }
+
+func useCNPGProxyAPIServer(t *testing.T, handle func(w http.ResponseWriter, c cnpgProxyCall)) *cnpgFakeProxyAPIServer {
+ t.Helper()
+ f := &cnpgFakeProxyAPIServer{}
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if strings.HasSuffix(r.URL.Path, "/selfsubjectaccessreviews") {
+ var review authv1.SelfSubjectAccessReview
+ _ = json.NewDecoder(r.Body).Decode(&review)
+ review.Status.Allowed = true
+ w.Header().Set("Content-Type", "application/json")
+ _ = json.NewEncoder(w).Encode(review)
+ return
+ }
+ f.record(r)
+ // /api/v1/namespaces/{ns}/pods/{scheme:pod:port}/proxy/{path}
+ parts := strings.SplitN(r.URL.Path, "/", 9)
+ if len(parts) < 8 || parts[5] != "pods" || parts[7] != "proxy" {
+ http.NotFound(w, r)
+ return
+ }
+ target := strings.Split(parts[6], ":")
+ path := "/"
+ if len(parts) == 9 {
+ path += parts[8]
+ }
+ handle(w, cnpgProxyCall{scheme: target[0], pod: target[1], port: target[2], path: path})
+ }))
+ t.Cleanup(srv.Close)
+ previous := k8s.SetTestConfig(&rest.Config{Host: srv.URL})
+ t.Cleanup(func() { k8s.SetTestConfig(previous) })
+ client, err := kubernetes.NewForConfig(&rest.Config{Host: srv.URL, ContentConfig: rest.ContentConfig{ContentType: "application/json"}})
+ if err != nil {
+ t.Fatal(err)
+ }
+ previousClient := k8s.SetTestClient(client)
+ t.Cleanup(func() { k8s.SetTestClient(previousClient) })
+ localCanIMu.Lock()
+ localCanIMemo = map[string]localCanIEntry{}
+ localCanIMu.Unlock()
+ return f
+}
+
+func cnpgWriteAPIStatus(w http.ResponseWriter, code int, reason metav1.StatusReason, msg string) {
+ w.Header().Set("Content-Type", "application/json")
+ w.WriteHeader(code)
+ _ = json.NewEncoder(w).Encode(metav1.Status{
+ TypeMeta: metav1.TypeMeta{Kind: "Status", APIVersion: "v1"},
+ Status: metav1.StatusFailure, Reason: reason, Code: int32(code), Message: msg,
+ })
+}
+
+func seedCNPGRuntimeCluster(t *testing.T, ns string) {
+ t.Helper()
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders", map[string]any{"instances": int64(2)}, nil), ns+"-uid"),
+ )
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: types.UID(ns + "-uid"), Controller: boolPtr(true)}
+ primary := cnpgPod(ns, "pg-orders-1", "pg-orders", owner)
+ primary.UID = types.UID(ns + "-p1")
+ primary.Spec.Containers[0].Command = []string{"/controller/manager", "instance", "run", "--status-port-tls"}
+ replica := cnpgPod(ns, "pg-orders-2", "pg-orders", owner)
+ replica.UID = types.UID(ns + "-p2")
+ replica.Labels["cnpg.io/instanceRole"] = "replica"
+ impostor := cnpgPod(ns, "pg-orders-impostor", "pg-orders")
+ seedCNPGPods(t, primary, replica, impostor)
+}
+
+func getCNPGRuntime(t *testing.T, path string) (int, cnpgsvc.CNPGClusterRuntimeResponse, string) {
+ t.Helper()
+ resp, err := http.Get(testServer.URL + path)
+ if err != nil {
+ t.Fatalf("GET %s: %v", path, err)
+ }
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ var out cnpgsvc.CNPGClusterRuntimeResponse
+ if resp.StatusCode == http.StatusOK {
+ if err := json.Unmarshal(body, &out); err != nil {
+ t.Fatalf("decode: %v (%s)", err, body)
+ }
+ }
+ return resp.StatusCode, out, string(body)
+}
+
+// cnpgHealthyInstances answers like a real cluster: the primary's status port
+// speaks TLS (a plain request is a bare 400), everything else is plain.
+func cnpgHealthyInstances(w http.ResponseWriter, c cnpgProxyCall) {
+ switch {
+ case c.port == "8000" && c.pod == "pg-orders-1" && c.scheme == "http":
+ w.WriteHeader(http.StatusBadRequest)
+ case c.port == "8000" && c.path == "/pg/status":
+ status := cnpgStatusFixture
+ if c.pod == "pg-orders-2" {
+ status = `{"isPrimary": false, "pod": {}, "receivedLsn": "0/7000148", "replayLsn": "0/7000148", "isWalReceiverActive": true, "lastFailedWALTime": "-infinity"}`
+ }
+ _, _ = io.WriteString(w, status)
+ case c.port == "9187" && c.scheme == "http" && c.path == "/metrics":
+ _, _ = io.WriteString(w, cnpgMetricsFixture)
+ default:
+ http.Error(w, "unexpected "+c.scheme+":"+c.pod+":"+c.port+c.path, http.StatusTeapot)
+ }
+}
+
+func TestCNPGClusterRuntime_OK(t *testing.T) {
+ seedCNPGRuntimeCluster(t, "pgrt")
+ api := useCNPGProxyAPIServer(t, cnpgHealthyInstances)
+
+ status, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrt/pg-orders/runtime")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.Cluster.UID != "pgrt-uid" || got.SampledAt == "" || got.Permission.Proxy != "allowed" || got.Permission.Grant == nil || *got.Permission.Grant != (auth.Grant{Verb: "get", Resource: "pods", Subresource: "proxy"}).In("pgrt") {
+ t.Errorf("envelope = %+v", got)
+ }
+ if len(got.Instances) != 2 || got.Instances[0].Pod != "pg-orders-1" || got.Instances[1].Pod != "pg-orders-2" {
+ t.Fatalf("instances = %+v, want only the owned instances", got.Instances)
+ }
+ p, r := got.Instances[0], got.Instances[1]
+ if p.Role != "primary" || r.Role != "replica" {
+ t.Errorf("roles = %s %s", p.Role, r.Role)
+ }
+ if p.Status.State != "ok" || p.Status.Scheme != "https" || p.Status.CapturedAt == "" || !p.Status.IsPrimary {
+ t.Errorf("primary status = %+v", p.Status.CNPGRuntimeSource)
+ }
+ if p.Metrics.State != "ok" || p.Metrics.Scheme != "http" || !cnpgEqF(p.Metrics.SessionsTotal, 6) {
+ t.Errorf("primary metrics = %+v", p.Metrics.CNPGRuntimeSource)
+ }
+ if len(p.Status.Slots) != 2 || !cnpgEqF(p.Status.Slots[0].RetainedBytes, 16384) {
+ t.Errorf("slot retention not joined from the same instance: %+v", p.Status.Slots)
+ }
+ if r.Status.State != "ok" || r.Status.IsPrimary || !r.Status.IsWalReceiverActive || r.Status.Scheme != "http" {
+ t.Errorf("replica status = %+v", r.Status)
+ }
+ for _, req := range api.seen() {
+ if strings.Contains(req, "impostor") {
+ t.Errorf("proxied to a non-instance Pod: %s", req)
+ }
+ if !strings.HasPrefix(req, "GET ") {
+ t.Errorf("non-GET proxy request: %s", req)
+ }
+ }
+ for _, h := range api.headers {
+ if h.Get("Cookie") != "" || h.Get("X-Forwarded-User") != "" {
+ t.Errorf("browser headers forwarded: %v", h)
+ }
+ }
+
+ // Within the memo lifetime a second view reuses the answers.
+ before := len(api.seen())
+ if status, _, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrt/pg-orders/runtime"); status != http.StatusOK {
+ t.Fatalf("second read: %d %s", status, body)
+ }
+ if after := len(api.seen()); after != before {
+ t.Errorf("second read issued %d more proxy requests, want 0 (memoized)", after-before)
+ }
+}
+
+func TestCNPGClusterRuntime_ProxyDeniedIsPerSource(t *testing.T) {
+ seedCNPGRuntimeCluster(t, "pgrtdeny")
+ api := useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ cnpgWriteAPIStatus(w, http.StatusForbidden, metav1.StatusReasonForbidden,
+ fmt.Sprintf(`pods "%s:%s" is forbidden: User "alice" cannot get resource "pods/proxy" in API group "" in the namespace "pgrtdeny"`, c.pod, c.port))
+ })
+ status, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtdeny/pg-orders/runtime")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d, want 200 so Kubernetes facts still render: %s", status, body)
+ }
+ if got.Permission.Proxy != "denied" {
+ t.Errorf("permission = %+v", got.Permission)
+ }
+ for _, inst := range got.Instances {
+ if inst.Status.State != "denied" || inst.Metrics.State != "denied" {
+ t.Errorf("%s: states = %s/%s, want denied", inst.Pod, inst.Status.State, inst.Metrics.State)
+ }
+ if inst.Status.CNPGInstanceStatusFacts != nil || inst.Metrics.CNPGInstanceMetricFacts != nil {
+ t.Errorf("%s: facts on a denied source", inst.Pod)
+ }
+ }
+ // A denial is not a scheme problem: one request per source, no retry.
+ if n := len(api.seen()); n != 4 {
+ t.Errorf("proxy requests = %d, want 4 (no scheme retry on forbidden)", n)
+ }
+ if strings.Contains(body, `"isPrimary"`) || strings.Contains(body, `"sessionsTotal"`) {
+ t.Errorf("denied sources serialized facts: %s", body)
+ }
+}
+
+func TestCNPGClusterRuntime_UnreachableAndRedirect(t *testing.T) {
+ seedCNPGRuntimeCluster(t, "pgrtdown")
+ api := useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ switch {
+ case c.path == "/pg/archive/partial":
+ _, _ = io.WriteString(w, "archived")
+ case c.port == "8000":
+ // The instance manager redirecting to a state-changing path must
+ // not be followed.
+ w.Header().Set("Location", "/api/v1/namespaces/pgrtdown/pods/"+c.scheme+":"+c.pod+":8000/proxy/pg/archive/partial")
+ w.WriteHeader(http.StatusFound)
+ default:
+ cnpgWriteAPIStatus(w, http.StatusServiceUnavailable, metav1.StatusReasonServiceUnavailable,
+ "error trying to reach service: dial tcp 10.0.0.5:9187: connect: connection refused")
+ }
+ })
+ status, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtdown/pg-orders/runtime")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ for _, inst := range got.Instances {
+ if inst.Metrics.State != "unreachable" || inst.Metrics.Error != "nothing is listening on port 9187 in the Pod" {
+ t.Errorf("%s metrics = %+v", inst.Pod, inst.Metrics.CNPGRuntimeSource)
+ }
+ if inst.Status.State != "error" || !strings.Contains(inst.Status.Error, "redirect") {
+ t.Errorf("%s status = %+v", inst.Pod, inst.Status.CNPGRuntimeSource)
+ }
+ }
+ for _, req := range api.seen() {
+ if strings.Contains(req, "/pg/archive/partial") {
+ t.Fatalf("followed a redirect to %s", req)
+ }
+ }
+ // connection refused is not a scheme mismatch: no second attempt.
+ metricsCalls := 0
+ for _, req := range api.seen() {
+ if strings.Contains(req, ":9187/") {
+ metricsCalls++
+ }
+ }
+ if metricsCalls != 2 {
+ t.Errorf("metrics proxy requests = %d, want one per instance", metricsCalls)
+ }
+}
+
+func TestCNPGClusterRuntime_ResponseCaps(t *testing.T) {
+ seedCNPGRuntimeCluster(t, "pgrtcap")
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if c.port == "8000" {
+ _, _ = io.WriteString(w, strings.Repeat("x", (1<<20)+1))
+ return
+ }
+ cutoff := strings.Index(cnpgMetricsFixture, "# TYPE cnpg_pg_stat_database_xact_commit")
+ _, _ = io.WriteString(w, cnpgMetricsFixture[:cutoff]+strings.Repeat("# filler\n", (4<<20)/9+1)+cnpgMetricsFixture[cutoff:])
+ })
+
+ _, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtcap/pg-orders/runtime")
+ p := got.Instances[0]
+ if p.Status.State != "error" || !strings.Contains(p.Status.Error, "larger than") {
+ t.Errorf("status over cap = %+v", p.Status.CNPGRuntimeSource)
+ }
+ if p.Metrics.State != "partial" || p.Metrics.Reason == "" || p.Metrics.XactCommitTotal != nil {
+ t.Errorf("metrics over cap = %+v: %s", p.Metrics.CNPGRuntimeSource, body)
+ }
+}
+
+func TestCNPGClusterRuntime_NotFoundAndEmpty(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", "pgrtempty", "pg-new", nil, nil), "new-uid"),
+ )
+ useCNPGProxyAPIServer(t, cnpgHealthyInstances)
+ if status, _, _ := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtempty/missing/runtime"); status != http.StatusNotFound {
+ t.Errorf("missing cluster: %d, want 404", status)
+ }
+ status, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtempty/pg-new/runtime")
+ if status != http.StatusOK || got.Instances == nil || len(got.Instances) != 0 {
+ t.Errorf("no instances: %d %s", status, body)
+ }
+}
+
+func TestCNPGClusterRuntime_FencedExplained(t *testing.T) {
+ cluster := cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", "pgrtfence", "pg-orders", nil, nil), "pgrtfence-uid")
+ cluster.SetAnnotations(map[string]string{"cnpg.io/fencedInstances": `["pg-orders-1"]`})
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cluster)
+ owner := metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "pgrtfence-uid", Controller: boolPtr(true)}
+ seedCNPGPods(t, cnpgPod("pgrtfence", "pg-orders-1", "pg-orders", owner))
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ cnpgWriteAPIStatus(w, http.StatusServiceUnavailable, metav1.StatusReasonServiceUnavailable, "error trying to reach service: connection refused")
+ })
+ _, got, body := getCNPGRuntime(t, "/api/cnpg/clusters/pgrtfence/pg-orders/runtime")
+ if len(got.Instances) != 1 || !got.Instances[0].Fenced || !strings.Contains(got.Instances[0].Metrics.Error, "fenced") {
+ t.Errorf("fenced instance = %s", body)
+ }
+}
+
+func seedCNPGPoolerChain(t *testing.T, ns string, poolerUID types.UID, conditions ...corev1.PodCondition) {
+ t.Helper()
+ ctx := context.Background()
+ deploy := &appsv1.Deployment{ObjectMeta: metav1.ObjectMeta{
+ Name: "pg-orders-rw", Namespace: ns, UID: types.UID(ns + "-deploy"),
+ OwnerReferences: []metav1.OwnerReference{{APIVersion: "postgresql.cnpg.io/v1", Kind: "Pooler", Name: "pg-orders-rw", UID: poolerUID, Controller: boolPtr(true)}},
+ }}
+ rs := &appsv1.ReplicaSet{ObjectMeta: metav1.ObjectMeta{
+ Name: "pg-orders-rw-abc", Namespace: ns, UID: types.UID(ns + "-rs"),
+ OwnerReferences: []metav1.OwnerReference{{APIVersion: "apps/v1", Kind: "Deployment", Name: "pg-orders-rw", UID: deploy.UID, Controller: boolPtr(true)}},
+ }}
+ strayRS := &appsv1.ReplicaSet{ObjectMeta: metav1.ObjectMeta{Name: "stray", Namespace: ns, UID: types.UID(ns + "-stray")}}
+ if _, err := testFakeClient.AppsV1().Deployments(ns).Create(ctx, deploy, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ for _, r := range []*appsv1.ReplicaSet{rs, strayRS} {
+ if _, err := testFakeClient.AppsV1().ReplicaSets(ns).Create(ctx, r, metav1.CreateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ }
+ t.Cleanup(func() {
+ _ = testFakeClient.AppsV1().Deployments(ns).Delete(context.Background(), deploy.Name, metav1.DeleteOptions{})
+ _ = testFakeClient.AppsV1().ReplicaSets(ns).Delete(context.Background(), rs.Name, metav1.DeleteOptions{})
+ _ = testFakeClient.AppsV1().ReplicaSets(ns).Delete(context.Background(), strayRS.Name, metav1.DeleteOptions{})
+ })
+ poolerPod := func(name string, owner metav1.OwnerReference) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns, UID: types.UID(ns + "-" + name),
+ Labels: map[string]string{"cnpg.io/poolerName": "pg-orders-rw"}, OwnerReferences: []metav1.OwnerReference{owner}},
+ Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "pgbouncer"}}},
+ Status: corev1.PodStatus{Conditions: conditions},
+ }
+ }
+ cache := k8s.GetResourceCache()
+ deadline := time.Now().Add(5 * time.Second)
+ for {
+ _, e1 := cache.Deployments().Deployments(ns).Get(deploy.Name)
+ _, e2 := cache.ReplicaSets().ReplicaSets(ns).Get(rs.Name)
+ _, e3 := cache.ReplicaSets().ReplicaSets(ns).Get(strayRS.Name)
+ if e1 == nil && e2 == nil && e3 == nil {
+ break
+ }
+ if time.Now().After(deadline) {
+ t.Fatal("pooler chain never reached the cache")
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+ seedCNPGPods(t,
+ poolerPod("pg-orders-rw-abc-1", metav1.OwnerReference{APIVersion: "apps/v1", Kind: "ReplicaSet", Name: rs.Name, UID: rs.UID, Controller: boolPtr(true)}),
+ poolerPod("pg-orders-rw-stray", metav1.OwnerReference{APIVersion: "apps/v1", Kind: "ReplicaSet", Name: strayRS.Name, UID: strayRS.UID, Controller: boolPtr(true)}),
+ )
+}
+
+func TestCNPGPoolerRuntime(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Pooler", "pgrtpool", "pg-orders-rw", map[string]any{"cluster": map[string]any{"name": "pg-orders"}}, nil), "pooler-uid"),
+ )
+ seedCNPGPoolerChain(t, "pgrtpool", "pooler-uid", corev1.PodCondition{Type: corev1.PodScheduled, Status: corev1.ConditionFalse, Reason: "Unschedulable", Message: "0/2 nodes are available: insufficient cpu"})
+ api := useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if c.port != "9127" || c.path != "/metrics" || c.scheme != "http" {
+ http.Error(w, "unexpected", http.StatusTeapot)
+ return
+ }
+ _, _ = io.WriteString(w, cnpgPoolerMetricsFixture)
+ })
+
+ resp, err := http.Get(testServer.URL + "/api/cnpg/poolers/pgrtpool/pg-orders-rw/runtime")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ if resp.StatusCode != http.StatusOK {
+ t.Fatalf("status = %d: %s", resp.StatusCode, body)
+ }
+ var got cnpgsvc.CNPGPoolerRuntimeResponse
+ if err := json.Unmarshal(body, &got); err != nil {
+ t.Fatal(err)
+ }
+ if got.Pooler.UID != "pooler-uid" || len(got.Pods) != 1 || got.Pods[0].Pod != "pg-orders-rw-abc-1" {
+ t.Fatalf("pods = %s, want only the Pod on the Pooler's controller chain", body)
+ }
+ pod := got.Pods[0]
+ if pod.SchedulingReason != "Unschedulable: 0/2 nodes are available: insufficient cpu" {
+ t.Errorf("scheduling reason = %q", pod.SchedulingReason)
+ }
+ if pod.State != "ok" || pod.CNPGPoolerPodFacts == nil || len(pod.Pools) != 1 || !cnpgEqF(pod.Pools[0].ClWaiting, 2) {
+ t.Errorf("pod = %s", body)
+ }
+ for _, req := range api.seen() {
+ if strings.Contains(req, "stray") {
+ t.Errorf("proxied to a Pod off the controller chain: %s", req)
+ }
+ }
+ r2, _ := http.Get(testServer.URL + "/api/cnpg/poolers/pgrtpool/missing/runtime")
+ r2.Body.Close()
+ if r2.StatusCode != http.StatusNotFound {
+ t.Errorf("missing pooler: %d, want 404", r2.StatusCode)
+ }
+}
+
+func TestCNPGPoolerRuntimeProxyDeniedKeepsSchedulingEvidence(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Pooler", "pgrtpooldeny", "pg-orders-rw", map[string]any{"cluster": map[string]any{"name": "pg-orders"}}, nil), "pooler-denied-uid"),
+ )
+ seedCNPGPoolerChain(t, "pgrtpooldeny", "pooler-denied-uid", corev1.PodCondition{Type: corev1.PodScheduled, Status: corev1.ConditionFalse, Reason: "Unschedulable", Message: "insufficient cpu"})
+ api := useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ cnpgWriteAPIStatus(w, http.StatusForbidden, metav1.StatusReasonForbidden, "pods/proxy is forbidden")
+ })
+ resp, err := http.Get(testServer.URL + "/api/cnpg/poolers/pgrtpooldeny/pg-orders-rw/runtime")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGPoolerRuntimeResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if resp.StatusCode != http.StatusOK || got.Permission.Proxy != "denied" || len(got.Pods) != 1 {
+ t.Fatalf("status %d: %+v", resp.StatusCode, got)
+ }
+ pod := got.Pods[0]
+ if pod.State != "denied" || pod.CNPGPoolerPodFacts != nil || pod.SchedulingReason != "Unschedulable: insufficient cpu" {
+ t.Fatalf("denied measurement must preserve scheduling evidence without metric facts: %+v", pod)
+ }
+ if len(api.seen()) != 1 {
+ t.Fatalf("proxy attempts = %d, want one denied request without scheme retry", len(api.seen()))
+ }
+}
+
+func TestCNPGPoolerPendingRuntime(t *testing.T) {
+ for _, outcome := range []string{"unreachable", "denied", "measured"} {
+ t.Run(outcome, func(t *testing.T) {
+ denied := outcome == "denied"
+ ns := "b7pool" + outcome
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cnpgRuntimeWithUID(cnpgObj("postgresql.cnpg.io/v1", "Pooler", ns, "pg-orders-rw", map[string]any{"cluster": map[string]any{"name": "pg-orders"}}, nil), ns+"-uid"))
+ seedCNPGPoolerChain(t, ns, types.UID(ns+"-uid"), corev1.PodCondition{Type: corev1.PodScheduled, Status: corev1.ConditionFalse, Reason: "Unschedulable", Message: "2 Too many pods"})
+ p, err := testFakeClient.CoreV1().Pods(ns).Get(context.Background(), "pg-orders-rw-abc-1", metav1.GetOptions{})
+ if err != nil {
+ t.Fatal(err)
+ }
+ p.Status.Phase = corev1.PodPending
+ if _, err = testFakeClient.CoreV1().Pods(ns).UpdateStatus(context.Background(), p, metav1.UpdateOptions{}); err != nil {
+ t.Fatal(err)
+ }
+ deadline := time.Now().Add(5 * time.Second)
+ for {
+ cached, err := k8s.GetResourceCache().Pods().Pods(ns).Get(p.Name)
+ if err == nil && cached.Status.Phase == corev1.PodPending {
+ break
+ }
+ if time.Now().After(deadline) {
+ t.Fatal("pending Pod did not reach cache")
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if denied {
+ cnpgWriteAPIStatus(w, http.StatusForbidden, metav1.StatusReasonForbidden, "pods/proxy is forbidden")
+ } else if outcome == "measured" {
+ _, _ = w.Write([]byte(cnpgPoolerMetricsFixture))
+ } else {
+ http.Error(w, "address not allowed", http.StatusBadGateway)
+ }
+ })
+ resp, err := http.Get(testServer.URL + "/api/cnpg/poolers/" + ns + "/pg-orders-rw/runtime")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGPoolerRuntimeResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if len(got.Pods) != 1 {
+ t.Fatalf("%+v", got)
+ }
+ if denied {
+ if got.Permission.Proxy != "denied" || got.Pods[0].State != "denied" {
+ t.Fatalf("denial must win: %+v", got)
+ }
+ } else if outcome == "measured" {
+ if got.Pods[0].State != "ok" || got.Pods[0].CNPGPoolerPodFacts == nil {
+ t.Fatalf("live read must beat stale Pod phase: %+v", got.Pods[0])
+ }
+ } else if got.Pods[0].Reason != "PgBouncer has not started (Pod cannot be scheduled)" || got.Pods[0].Error != "" {
+ t.Fatalf("%+v", got.Pods[0])
+ }
+ })
+ }
+}
diff --git a/internal/server/cnpg_schedule.go b/internal/server/cnpg_schedule.go
new file mode 100644
index 0000000000..c393483524
--- /dev/null
+++ b/internal/server/cnpg_schedule.go
@@ -0,0 +1,30 @@
+package server
+
+import (
+ "net/http"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+)
+
+// GET /api/cnpg/scheduledbackups/{ns}/{name}/schedule-preview?schedule=
+// Pure computation over the schedule read as the caller; writes nothing.
+func (s *Server) handleCNPGSchedulePreview(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ namespace := chi.URLParam(r, "namespace")
+ name := chi.URLParam(r, "name")
+ reader := s.cnpgReader(r)
+ dyn := reader.dynamic
+ if dyn == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "cluster client not available — check cluster connection")
+ return
+ }
+ preview, err := reader.PreviewSchedule(r.Context(), dyn, namespace, name, r.URL.Query().Get("schedule"), time.Now())
+ if err != nil {
+ s.writeCNPGActionError(w, err, "schedule-preview", namespace, name)
+ return
+ }
+ s.writeJSON(w, preview)
+}
diff --git a/internal/server/cnpg_service.go b/internal/server/cnpg_service.go
new file mode 100644
index 0000000000..e93822dd4b
--- /dev/null
+++ b/internal/server/cnpg_service.go
@@ -0,0 +1,189 @@
+package server
+
+import (
+ "context"
+ "net/http"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/client-go/dynamic"
+ "k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+ bp "github.com/skyhook-io/radar/pkg/audit"
+)
+
+type cnpgReadAdapter struct {
+ *cnpgsvc.Reader
+ dynamic dynamic.Interface
+ config *rest.Config
+ actionContext string
+}
+
+func (s *Server) cnpgActionClients(r *http.Request, req integration.ActionRequest) (cnpgsvc.ActionClients, error) {
+ reader := s.cnpgReader(r)
+ if reader.dynamic == nil {
+ return cnpgsvc.ActionClients{}, integration.RefuseAction(http.StatusServiceUnavailable, "", "cluster client not available — check cluster connection")
+ }
+ if err := integration.CheckReviewedContext(req.ReviewedContext, reader.actionContext); err != nil {
+ return cnpgsvc.ActionClients{}, err
+ }
+ return cnpgsvc.ActionClients{Dynamic: reader.dynamic, Typed: reader.Clients.Typed, Exec: reader.Clients.Exec}, nil
+}
+
+func (s *Server) cnpgReader(r *http.Request) *cnpgReadAdapter {
+ var promClient *prometheuspkg.Client
+ var cache *k8s.ResourceCache
+ var discovery *k8s.ResourceDiscovery
+ var dynamicCache *k8s.DynamicResourceCache
+ var operation context.Context
+ var identity, clusterContext, target string
+ var connected, forced bool
+ adapter := &cnpgReadAdapter{Reader: &cnpgsvc.Reader{}}
+ k8s.CaptureClusterReads(func() {
+ operation = k8s.OperationContext()
+ cache, discovery, dynamicCache = k8s.GetResourceCache(), k8s.GetResourceDiscovery(), k8s.GetDynamicResourceCache()
+ connected = k8s.IsConnected()
+ identity, clusterContext = cnpgRuntimeIdentity(r), k8s.ActiveClusterContext()
+ forced, target = k8s.ForceNamespaceScope, k8s.GetNamespaceScopeTarget()
+ promClient = prometheuspkg.GetClient()
+ adapter.dynamic, adapter.actionContext = s.getDynamicClientSnapshotForRequest(r)
+ adapter.config = s.getConfigForRequest(r)
+ adapter.Clients = cnpgsvc.ReadClients{Typed: s.getClientForRequest(r), Proxy: cnpgRuntimeClient(r)}
+ adapter.Clients.Exec = cnpgsvc.NewExec(adapter.Clients.Typed, adapter.config)
+ })
+ current := func() bool { return operation != nil && operation.Err() == nil }
+ boundBudget := func(ctx context.Context) *syncBudget {
+ budget := newSyncBudget(ctx)
+ budget.bound, budget.discovery, budget.dynamicCache = true, discovery, dynamicCache
+ return budget
+ }
+ budget := boundBudget(r.Context())
+ auditConfig := getAuditConfig()
+ clients := adapter.Clients
+ adapter.Reader = &cnpgsvc.Reader{
+ ClusterContext: clusterContext,
+ Identity: identity,
+ Metrics: cnpgsvc.Metrics{
+ Connection: func(ctx context.Context) (bool, error) {
+ if promClient == nil {
+ return false, nil
+ }
+ _, _, err := promClient.EnsureConnected(ctx)
+ return true, err
+ },
+ PVCUsage: promClient.QueryPVCUsage, CNPGScope: promClient.ResolveCNPGScope, PVCScope: promClient.ResolvePVCScope,
+ History: promClient.QueryCNPGHistory, FleetLag: promClient.QueryCNPGFleetLag, FleetSlots: promClient.QueryCNPGFleetSlots, DiskGrowth: promClient.QueryCNPGDiskGrowth,
+ },
+ Access: cnpgsvc.Access{
+ CanRead: func(ctx context.Context, group, resource, namespace, verb string) bool {
+ if !current() {
+ return false
+ }
+ allowed := s.canRead(r.WithContext(ctx), group, resource, namespace, verb)
+ return current() && allowed
+ },
+ Permission: func(ctx context.Context, grant auth.Grant) string {
+ if !current() {
+ return integration.PermissionDenied
+ }
+ permission := s.grantPermission(r.WithContext(ctx), grant)
+ if !current() {
+ return integration.PermissionDenied
+ }
+ return permission
+ },
+ MetricsRead: func(ctx context.Context, group, resource, namespace, verb string) bool {
+ if !current() {
+ return false
+ }
+ allowed := s.prometheusAuthGate(r.WithContext(ctx), group, resource, namespace, verb)
+ return current() && allowed
+ },
+ },
+ Observations: cnpgsvc.Observations{
+ FilterAudit: func(results *bp.ScanResults) *bp.ScanResults { return applyAuditSettings(results, auditConfig) },
+ WorkspaceRead: func(ctx context.Context, cache *k8s.ResourceCache, k integration.WorkspaceKind, namespaces, groups []string) (integration.KindAccess, []*unstructured.Unstructured) {
+ if !current() {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ access, items := s.readWorkspaceKind(r.WithContext(ctx), cache, k, namespaces, groups, budget)
+ if !current() {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ return access, items
+ },
+ Issues: func(ctx context.Context, namespaces []string) []issues.Issue {
+ if !current() {
+ return nil
+ }
+ rows := s.cnpgIssueRows(r.WithContext(ctx), namespaces)
+ if !current() {
+ return nil
+ }
+ return rows
+ },
+ Cache: cache, Discovery: discovery, Connected: connected,
+ Cluster: func(ctx context.Context, namespace, name string, grants ...auth.Grant) (*k8s.ResourceCache, *unstructured.Unstructured, error) {
+ if !connected || !current() {
+ return nil, nil, cnpgsvc.ErrCNPGDisconnected
+ }
+ if err := s.authorizeCNPGCachedRead(r.WithContext(ctx), namespace, "clusters", grants...); err != nil {
+ return nil, nil, err
+ }
+ if cache == nil {
+ return nil, nil, &cnpgsvc.ReadFailure{Status: http.StatusServiceUnavailable, Message: "resource cache not available"}
+ }
+ if !current() {
+ return nil, nil, cnpgsvc.ErrCNPGDisconnected
+ }
+ objects, err := listDynamicSyncedWithin(ctx, cache, "Cluster", cnpgsvc.Group, namespace, boundBudget(ctx))
+ obj, err := cnpgsvc.SelectCluster(objects, err, namespace, name)
+ obj, err = cnpgCachedResourceResult(obj, err, "Cluster", namespace, name)
+ if !current() {
+ return nil, nil, cnpgsvc.ErrCNPGDisconnected
+ }
+ return cache, obj, err
+ },
+ OperatorScope: func(ctx context.Context) []string {
+ if !current() {
+ return []string{}
+ }
+ var requested []string
+ if forced {
+ if target == "" {
+ return []string{}
+ }
+ requested = []string{target}
+ }
+ namespaces := s.getUserNamespaces(r.WithContext(ctx), requested)
+ if !current() {
+ return []string{}
+ }
+ return namespaces
+ },
+ TypedScope: func(ctx context.Context, cache *k8s.ResourceCache, namespaces []string, group, resource string) (integration.KindAccess, []string) {
+ if !current() {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, []string{}
+ }
+ access, read := s.typedKindScopeWithCandidates(r.WithContext(ctx), cache, namespaces, group, resource, func() []string { return namespaceNamesInCache(cache) })
+ if !current() {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, []string{}
+ }
+ return access, read
+ },
+ DynamicList: func(ctx context.Context, cache *k8s.ResourceCache, kind, group, namespace string) ([]*unstructured.Unstructured, error) {
+ if !current() {
+ return nil, integration.ErrDynamicNotSynced
+ }
+ return listDynamicSyncedWithin(ctx, cache, kind, group, namespace, boundBudget(ctx))
+ },
+ },
+ Clients: clients,
+ }
+ return adapter
+}
diff --git a/internal/server/cnpg_service_test.go b/internal/server/cnpg_service_test.go
new file mode 100644
index 0000000000..5169109988
--- /dev/null
+++ b/internal/server/cnpg_service_test.go
@@ -0,0 +1,100 @@
+package server
+
+import (
+ "context"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "testing"
+
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func TestCNPGCachedReadFailurePreservesCauseWithoutExposingIt(t *testing.T) {
+ for _, kind := range []string{"Cluster", "Pooler"} {
+ t.Run(kind, func(t *testing.T) {
+ cause := errors.New("internal cache failure details")
+ _, err := cnpgCachedResourceResult(nil, cause, kind, "pg", "orders")
+ var failure *cnpgsvc.ReadFailure
+ if !errors.Is(err, cause) || !errors.As(err, &failure) || failure.Status != http.StatusInternalServerError {
+ t.Fatalf("cache failure = %v", err)
+ }
+ w := httptest.NewRecorder()
+ (&Server{}).writeCNPGCachedReadError(w, err, "pg", "orders")
+ if w.Code != http.StatusInternalServerError || w.Body.String() != "{\"error\":\"failed to read CloudNativePG "+kind+"\"}\n" {
+ t.Fatalf("response = %d %s", w.Code, w.Body.String())
+ }
+ })
+ }
+}
+
+func TestCNPGCachedReadDistinguishesMissingFromSyncing(t *testing.T) {
+ for _, kind := range []string{"Cluster", "Pooler"} {
+ for _, tc := range []struct {
+ name string
+ cause error
+ status int
+ message string
+ }{
+ {"missing", nil, http.StatusNotFound, "CloudNativePG " + kind + " pg/orders not found"},
+ {"not installed", k8s.ErrUnknownDynamicKind, http.StatusNotFound, "CloudNativePG " + kind + " pg/orders not found"},
+ {"syncing", integration.ErrDynamicNotSynced, http.StatusServiceUnavailable, "CloudNativePG " + kind + "s are still syncing"},
+ } {
+ t.Run(kind+"/"+tc.name, func(t *testing.T) {
+ _, err := cnpgCachedResourceResult(nil, tc.cause, kind, "pg", "orders")
+ w := httptest.NewRecorder()
+ (&Server{}).writeCNPGCachedReadError(w, err, "pg", "orders")
+ if w.Code != tc.status || w.Body.String() != "{\"error\":\""+tc.message+"\"}\n" {
+ t.Fatalf("response = %d %s", w.Code, w.Body.String())
+ }
+ })
+ }
+ object := &unstructured.Unstructured{}
+ if got, err := cnpgCachedResourceResult(object, nil, kind, "pg", "orders"); err != nil || got != object {
+ t.Fatalf("successful %s read = %v %v", kind, got, err)
+ }
+ }
+}
+
+func TestCNPGReadAdapterRefusesSupersededCluster(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj(cnpgsvc.Group+"/v1", "Cluster", "pg", "orders", nil, nil), "orders-uid"))
+ srv := &Server{}
+ r := httptest.NewRequest("GET", "/", nil)
+ reader := srv.cnpgReader(r)
+ ctx := context.Background()
+ cache, cluster, err := reader.Observations.Cluster(ctx, "pg", "orders")
+ if err != nil || cache == nil || cluster.GetUID() != "orders-uid" {
+ t.Fatalf("initial read = %v %v %v", cache, cluster, err)
+ }
+ k8s.CancelOngoingOperations()
+ if _, _, err := reader.Observations.Cluster(ctx, "pg", "orders"); !errors.Is(err, cnpgsvc.ErrCNPGDisconnected) {
+ t.Fatalf("superseded read = %v", err)
+ }
+ grant := auth.Grant{Group: cnpgsvc.Group, Resource: "clusters", Verb: "get", Namespace: "pg"}
+ if reader.Access.CanRead(ctx, grant.Group, grant.Resource, grant.Namespace, grant.Verb) || reader.Access.Permission(ctx, grant) != integration.PermissionDenied {
+ t.Fatal("superseded adapter accepted a new permission verdict")
+ }
+ if _, err := reader.Observations.DynamicList(ctx, cache, "Cluster", cnpgsvc.Group, "pg"); !errors.Is(err, integration.ErrDynamicNotSynced) {
+ t.Fatalf("superseded dynamic read = %v", err)
+ }
+ if access, objects := reader.Observations.WorkspaceRead(ctx, cache, cnpgWorkspaceFixtureKinds[0], []string{"pg"}, []string{cnpgsvc.Group}); access.State != integration.KindCoverageSyncing || len(objects) != 0 {
+ t.Fatalf("superseded workspace read = %+v %+v", access, objects)
+ }
+ if _, cluster, err := srv.cnpgReader(r).Observations.Cluster(ctx, "pg", "orders"); err != nil || cluster.GetUID() != "orders-uid" {
+ t.Fatalf("fresh adapter did not recover: %v %v", cluster, err)
+ }
+}
+
+func TestCNPGActionClientBindingRejectsDifferentReviewedContext(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds)
+ _, err := (&Server{}).cnpgActionClients(httptest.NewRequest("POST", "/", nil), integration.ActionRequest{ReviewedContext: "different-context", UID: "orders-uid"})
+ var refusal *integration.ActionError
+ if !errors.As(err, &refusal) || refusal.Status != 409 || refusal.Code != integration.ActionCodeContextChanged {
+ t.Fatalf("wrong-context client binding = %v", err)
+ }
+}
diff --git a/internal/server/cnpg_sessions.go b/internal/server/cnpg_sessions.go
new file mode 100644
index 0000000000..27822c7e9a
--- /dev/null
+++ b/internal/server/cnpg_sessions.go
@@ -0,0 +1,42 @@
+package server
+
+import (
+ "net/http"
+
+ "github.com/go-chi/chi/v5"
+)
+
+// Diagnosis inside PostgreSQL. Radar runs fixed SQL with psql in the
+// instance's postgres container through the caller's own pods/exec — the same
+// access that lets them open psql themselves, which is why query text is
+// shown to them. Nothing a caller sends is ever part of the SQL text: the only
+// inputs (a backend's pid and start time) are validated and passed as psql
+// variables, which psql quotes as literals.
+
+// handleCNPGClusterSessions serves GET /api/cnpg/clusters/{ns}/{name}/sessions:
+// the blocker → victim relations on one instance (the primary unless ?pod=
+// names another instance), read with fixed SQL over the caller's pods/exec.
+// Without exec the answer is a 200 whose state is denied, never empty.
+func (s *Server) handleCNPGClusterSessions(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ if !s.authorizeCNPGRuntime(w, r, namespace, "clusters") {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ _, cluster, err := reader.Observations.Cluster(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ resp, err := reader.Sessions(r.Context(), cache, cluster, r.URL.Query().Get("pod"))
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
diff --git a/internal/server/cnpg_sessions_test.go b/internal/server/cnpg_sessions_test.go
new file mode 100644
index 0000000000..88bedbf8af
--- /dev/null
+++ b/internal/server/cnpg_sessions_test.go
@@ -0,0 +1,54 @@
+package server
+
+import (
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func TestCNPGActionPartialOutcomeSerialization(t *testing.T) {
+ err := &integration.ActionError{Status: http.StatusForbidden, Code: integration.ActionCodePartial, Message: "denied by admission policy", Completed: []string{"deleted PVC pg-2", "deleted PVC pg-2-wal"}}
+
+ rec := httptest.NewRecorder()
+ (&Server{}).writeCNPGActionError(rec, err, "destroyInstance", "db", "pg")
+ body := rec.Body.String()
+ if rec.Code != http.StatusForbidden || !strings.Contains(body, `"code":"partial"`) || !strings.Contains(body, `"completed":["deleted PVC pg-2","deleted PVC pg-2-wal"]`) {
+ t.Errorf("response = %d %s", rec.Code, body)
+ }
+}
+
+func TestCNPGClusterSessions_WaitsForPrimary(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", "pgsessionswait", "analytics", map[string]any{"instances": int64(1)}, nil), "analytics-uid"))
+ resp, err := http.Get(testServer.URL + "/api/cnpg/clusters/pgsessionswait/analytics/sessions")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGSessionsResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if resp.StatusCode != http.StatusOK || got.State != "unavailable" || got.Error != "" || got.Reason != "Available once the primary is running" {
+ t.Fatalf("%d: %+v", resp.StatusCode, got)
+ }
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgsessionswait"}}
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", "pgsessionswait", true)
+ perms.SetCanI("list", "", "pods", "pgsessionswait", true)
+ perms.SetCanI("create", "", "pods/exec", "pgsessionswait", false)
+ env.srv.permCache.Set("no-exec-before-primary", nil, perms)
+ denied := env.authGet(t, "/api/cnpg/clusters/pgsessionswait/analytics/sessions", "no-exec-before-primary", "")
+ defer denied.Body.Close()
+ if err := json.NewDecoder(denied.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if denied.StatusCode != http.StatusOK || got.State != "denied" || got.Permission.Exec != integration.PermissionDenied {
+ t.Fatalf("missing primary must not imply exec access: %d %+v", denied.StatusCode, got)
+ }
+}
diff --git a/internal/server/cnpg_storage.go b/internal/server/cnpg_storage.go
new file mode 100644
index 0000000000..eb033bd730
--- /dev/null
+++ b/internal/server/cnpg_storage.go
@@ -0,0 +1,36 @@
+package server
+
+import (
+ "net/http"
+
+ "github.com/go-chi/chi/v5"
+)
+
+// handleCNPGClusterStorage serves GET /api/cnpg/clusters/{namespace}/{name}/storage.
+// Reading the Cluster is the gate; its claims, their usage and the WAL facts
+// each need their own grant and report their own coverage.
+func (s *Server) handleCNPGClusterStorage(w http.ResponseWriter, r *http.Request) {
+ namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
+ resp, err := s.cnpgReader(r).ClusterStorage(r.Context(), namespace, name)
+ if err != nil {
+ s.writeCNPGCachedReadError(w, err, namespace, name)
+ return
+ }
+ s.writeJSON(w, resp)
+}
+
+// ---------- fleet ----------
+
+func (s *Server) handleCNPGFleetDisk(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
+ if cache == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource cache not available")
+ return
+ }
+ namespaces := s.parseNamespacesForUser(r)
+ s.writeJSON(w, reader.FleetDisk(r.Context(), cache, namespaces))
+}
diff --git a/internal/server/cnpg_storage_test.go b/internal/server/cnpg_storage_test.go
new file mode 100644
index 0000000000..a1102279ea
--- /dev/null
+++ b/internal/server/cnpg_storage_test.go
@@ -0,0 +1,473 @@
+package server
+
+import (
+ "context"
+ "encoding/json"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "strconv"
+ "strings"
+ "testing"
+ "time"
+
+ corev1 "k8s.io/api/core/v1"
+ storagev1 "k8s.io/api/storage/v1"
+ "k8s.io/apimachinery/pkg/api/resource"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ "github.com/skyhook-io/radar/internal/k8s"
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
+)
+
+func cnpgTestPVC(ns, name, cluster, instance, role string, owner *metav1.OwnerReference, requested, capacity string) *corev1.PersistentVolumeClaim {
+ sc := "fast"
+ pvc := &corev1.PersistentVolumeClaim{
+ ObjectMeta: metav1.ObjectMeta{Name: name, Namespace: ns, Labels: map[string]string{"cnpg.io/cluster": cluster, "cnpg.io/pvcRole": role}},
+ Spec: corev1.PersistentVolumeClaimSpec{
+ StorageClassName: &sc,
+ Resources: corev1.VolumeResourceRequirements{Requests: corev1.ResourceList{corev1.ResourceStorage: resource.MustParse(requested)}},
+ },
+ Status: corev1.PersistentVolumeClaimStatus{Phase: corev1.ClaimBound, Capacity: corev1.ResourceList{corev1.ResourceStorage: resource.MustParse(capacity)}},
+ }
+ if instance != "" {
+ pvc.Labels["cnpg.io/instanceName"] = instance
+ }
+ if owner != nil {
+ pvc.OwnerReferences = []metav1.OwnerReference{*owner}
+ }
+ return pvc
+}
+
+func seedCNPGClaims(t *testing.T, pvcs ...*corev1.PersistentVolumeClaim) {
+ t.Helper()
+ ctx := context.Background()
+ for _, p := range pvcs {
+ if _, err := testFakeClient.CoreV1().PersistentVolumeClaims(p.Namespace).Create(ctx, p, metav1.CreateOptions{}); err != nil {
+ t.Fatalf("create pvc %s: %v", p.Name, err)
+ }
+ t.Cleanup(func() {
+ _ = testFakeClient.CoreV1().PersistentVolumeClaims(p.Namespace).Delete(context.Background(), p.Name, metav1.DeleteOptions{})
+ })
+ }
+ lister := k8s.GetResourceCache().PersistentVolumeClaims()
+ deadline := time.Now().Add(5 * time.Second)
+ for {
+ seen := 0
+ for _, p := range pvcs {
+ if _, err := lister.PersistentVolumeClaims(p.Namespace).Get(p.Name); err == nil {
+ seen++
+ }
+ }
+ if seen == len(pvcs) {
+ return
+ }
+ if time.Now().After(deadline) {
+ t.Fatalf("claims did not reach the cache")
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+}
+
+func seedCNPGStorageClass(t *testing.T, name string, allow bool) {
+ t.Helper()
+ sc := &storagev1.StorageClass{ObjectMeta: metav1.ObjectMeta{Name: name}, Provisioner: "example.com/csi", AllowVolumeExpansion: &allow}
+ if _, err := testFakeClient.StorageV1().StorageClasses().Create(context.Background(), sc, metav1.CreateOptions{}); err != nil {
+ t.Fatalf("create storageclass: %v", err)
+ }
+ t.Cleanup(func() {
+ _ = testFakeClient.StorageV1().StorageClasses().Delete(context.Background(), name, metav1.DeleteOptions{})
+ })
+ deadline := time.Now().Add(5 * time.Second)
+ for {
+ if _, err := k8s.GetResourceCache().StorageClasses().Get(name); err == nil {
+ return
+ }
+ if time.Now().After(deadline) {
+ t.Fatalf("storageclass did not reach the cache")
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+}
+
+// seedCNPGStorageCluster: two instances, instance 1 with separate WAL storage
+// mid-resize, and a claim that carries the cluster label without being owned
+// by it.
+func seedCNPGStorageCluster(t *testing.T, ns string) {
+ t.Helper()
+ uid := types.UID(ns + "-uid")
+ cluster := cnpgObj("postgresql.cnpg.io/v1", "Cluster", ns, "pg-orders",
+ map[string]any{
+ "instances": int64(2),
+ "storage": map[string]any{"size": "2Gi", "storageClass": "fast"},
+ "walStorage": map[string]any{"size": "1Gi"},
+ "tablespaces": []any{map[string]any{"name": "archive", "storage": map[string]any{"pvcTemplate": map[string]any{"resources": map[string]any{"requests": map[string]any{"storage": "5Gi"}}}}}},
+ },
+ map[string]any{
+ "currentPrimary": "pg-orders-1",
+ "instanceNames": []any{"pg-orders-1", "pg-orders-2"},
+ "healthyPVC": []any{"pg-orders-1", "pg-orders-2"},
+ "resizingPVC": []any{"pg-orders-1-wal"},
+ })
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cnpgRuntimeWithUID(cluster, string(uid)))
+ owner := &metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: uid, Controller: boolPtr(true)}
+ foreign := &metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg-orders", UID: "someone-else"}
+ wal := cnpgTestPVC(ns, "pg-orders-1-wal", "pg-orders", "pg-orders-1", "PG_WAL", owner, "2Gi", "1Gi")
+ wal.Status.Conditions = []corev1.PersistentVolumeClaimCondition{{Type: corev1.PersistentVolumeClaimFileSystemResizePending, Status: corev1.ConditionTrue, Message: "waiting for pod restart"}}
+ seedCNPGClaims(t,
+ cnpgTestPVC(ns, "pg-orders-1", "pg-orders", "pg-orders-1", "PG_DATA", owner, "2Gi", "2Gi"),
+ wal,
+ cnpgTestPVC(ns, "pg-orders-2", "pg-orders", "pg-orders-2", "PG_DATA", owner, "2Gi", "2Gi"),
+ cnpgTestPVC(ns, "pg-orders-9", "pg-orders", "pg-orders-9", "PG_DATA", foreign, "2Gi", "2Gi"),
+ )
+ ownerPod := metav1.OwnerReference{APIVersion: owner.APIVersion, Kind: owner.Kind, Name: owner.Name, UID: uid, Controller: boolPtr(true)}
+ primary := cnpgPod(ns, "pg-orders-1", "pg-orders", ownerPod)
+ primary.UID = types.UID(ns + "-p1")
+ primary.Spec.Containers[0].Command = []string{"/controller/manager", "instance", "run", "--status-port-tls"}
+ replica := cnpgPod(ns, "pg-orders-2", "pg-orders", ownerPod)
+ replica.UID = types.UID(ns + "-p2")
+ replica.Labels["cnpg.io/instanceRole"] = "replica"
+ seedCNPGPods(t, primary, replica)
+}
+
+func getCNPGStorage(t *testing.T, path string) (int, cnpgsvc.CNPGClusterStorageResponse, string) {
+ t.Helper()
+ resp, err := http.Get(testServer.URL + path)
+ if err != nil {
+ t.Fatalf("GET %s: %v", path, err)
+ }
+ defer resp.Body.Close()
+ body, _ := io.ReadAll(resp.Body)
+ var out cnpgsvc.CNPGClusterStorageResponse
+ if resp.StatusCode == http.StatusOK {
+ if err := json.Unmarshal(body, &out); err != nil {
+ t.Fatalf("decode: %v (%s)", err, body)
+ }
+ }
+ return resp.StatusCode, out, string(body)
+}
+
+// usePrometheusVolumeStats serves kubelet volume stats for the named claims.
+func usePrometheusVolumeStats(t *testing.T, used, capacity map[string]float64) {
+ t.Helper()
+ usePrometheusVolumeStatsFrom(t, used, capacity, false)
+}
+
+// usePrometheusVolumeStatsFrom with twoClusters answers identity checks with a
+// second cluster holding claims of the same names.
+func usePrometheusVolumeStatsFrom(t *testing.T, used, capacity map[string]float64, twoClusters bool) {
+ t.Helper()
+ series := func(values map[string]float64, query string) string {
+ var rows []string
+ for claim, v := range values {
+ if !strings.Contains(query, claim) {
+ continue
+ }
+ b, _ := json.Marshal(map[string]any{"metric": map[string]string{"persistentvolumeclaim": claim}, "value": []any{1700000000, strconv.FormatFloat(v, 'f', -1, 64)}})
+ rows = append(rows, string(b))
+ }
+ return `{"status":"success","data":{"resultType":"vector","result":[` + strings.Join(rows, ",") + `]}}`
+ }
+ srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.Header().Set("Content-Type", "application/json")
+ q := r.URL.Query().Get("query")
+ if q == "" {
+ _ = r.ParseForm()
+ q = r.Form.Get("query")
+ }
+ switch {
+ case q == "up":
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{"job":"prometheus"},"value":[1700000000,"1"]}]}}`)
+ case strings.HasPrefix(q, "max(count by (persistentvolumeclaim)"):
+ identities := "1"
+ if twoClusters {
+ identities = "2"
+ }
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[{"metric":{},"value":[1700000000,"`+identities+`"]}]}}`)
+ case strings.Contains(q, "kubelet_volume_stats_used_bytes"):
+ _, _ = io.WriteString(w, series(used, q))
+ case strings.Contains(q, "kubelet_volume_stats_capacity_bytes"):
+ _, _ = io.WriteString(w, series(capacity, q))
+ default:
+ _, _ = io.WriteString(w, `{"status":"success","data":{"resultType":"vector","result":[]}}`)
+ }
+ }))
+ prometheuspkg.Initialize(nil, nil, "test")
+ prometheuspkg.SetManualURL(srv.URL)
+ t.Cleanup(func() {
+ srv.Close()
+ prometheuspkg.Reset()
+ prometheuspkg.Initialize(nil, nil, "")
+ })
+}
+
+func cnpgStorageInstances(t *testing.T, got cnpgsvc.CNPGClusterStorageResponse) map[string]cnpgsvc.CNPGStorageInstance {
+ t.Helper()
+ out := map[string]cnpgsvc.CNPGStorageInstance{}
+ for _, in := range got.Instances {
+ out[in.Name] = in
+ }
+ return out
+}
+
+func TestCNPGClusterStorage_VolumesWALAndUsage(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgst")
+ seedCNPGStorageClass(t, "fast", true)
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if c.port == "9187" {
+ _, _ = io.WriteString(w, cnpgMetricsFixture+"cnpg_collector_pg_wal{value=\"count\"} 8\n")
+ return
+ }
+ cnpgHealthyInstances(w, c)
+ })
+ usePrometheusVolumeStats(t,
+ map[string]float64{"pg-orders-1": 1.9e9, "pg-orders-2": 1e9},
+ map[string]float64{"pg-orders-1": 2e9, "pg-orders-2": 2e9})
+
+ status, got, body := getCNPGStorage(t, "/api/cnpg/clusters/pgst/pg-orders/storage")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.Volumes.State != "ok" || got.WAL.State != "ok" {
+ t.Fatalf("coverage volumes=%+v wal=%+v", got.Volumes, got.WAL)
+ }
+ if len(got.Excluded) != 1 || got.Excluded[0].Claim != "pg-orders-9" {
+ t.Errorf("excluded = %+v, want the claim owned by another UID", got.Excluded)
+ }
+ insts := cnpgStorageInstances(t, got)
+ if len(insts) != 2 {
+ t.Fatalf("instances = %+v", got.Instances)
+ }
+ p := insts["pg-orders-1"]
+ if p.Role != "primary" || insts["pg-orders-2"].Role != "replica" {
+ t.Errorf("roles = %s %s", p.Role, insts["pg-orders-2"].Role)
+ }
+ if len(p.Volumes) != 2 || p.Volumes[0].Role != "PG_DATA" || p.Volumes[1].Role != "PG_WAL" {
+ t.Fatalf("primary volumes = %+v, want data then WAL", p.Volumes)
+ }
+ data, wal := p.Volumes[0], p.Volumes[1]
+ if data.StorageClass.AllowVolumeExpansion == nil || !*data.StorageClass.AllowVolumeExpansion || data.ClusterState != "healthy" {
+ t.Errorf("data volume = %+v", data)
+ }
+ if !wal.Resize.Pending || len(wal.Resize.Conditions) != 1 || wal.Resize.Conditions[0].Type != "FileSystemResizePending" || wal.ClusterState != "resizing" {
+ t.Errorf("wal resize = %+v state %q", wal.Resize, wal.ClusterState)
+ }
+ if data.Usage.State != "ok" || data.Usage.Ratio == nil || *data.Usage.Ratio < 0.94 {
+ t.Errorf("data usage = %+v", data.Usage)
+ }
+ // Prometheus answered for the data claims only: the WAL claim is unmeasured,
+ // never zero.
+ if wal.Usage.State != "noSeries" || wal.Usage.Ratio != nil || wal.Usage.UsedBytes != nil {
+ t.Errorf("wal usage = %+v, want noSeries without figures", wal.Usage)
+ }
+ if got.Usage.State != "partial" {
+ t.Errorf("usage coverage = %+v, want partial", got.Usage)
+ }
+ if len(got.Findings) != 1 || got.Findings[0].Severity != "critical" || got.Findings[0].Claim != "pg-orders-1" {
+ t.Errorf("findings = %+v, want one critical on pg-orders-1", got.Findings)
+ }
+
+ pw := p.WAL
+ if pw == nil || pw.Volume != "pg-orders-1-wal" || !cnpgEqF(pw.SizeBytes, 1.34217728e+08) || !cnpgEqF(pw.Segments, 8) {
+ t.Fatalf("primary WAL = %+v", pw)
+ }
+ if pw.ReadyToArchive == nil || *pw.ReadyToArchive != 3 || pw.ArchivingFailed {
+ t.Errorf("archive backlog = %+v", pw)
+ }
+ if len(pw.Slots) != 1 || pw.Slots[0].Bytes != 16384 {
+ t.Errorf("slots = %+v", pw.Slots)
+ }
+ if rw := insts["pg-orders-2"].WAL; rw == nil || rw.Volume != "pg-orders-2" {
+ t.Errorf("replica WAL volume = %+v, want the data claim when there is no WAL claim", rw)
+ }
+
+ targets := got.Expansion.Targets
+ if len(targets) != 3 || targets[0].Field != "spec.storage.size" || targets[0].Declared != "2Gi" ||
+ targets[1].Field != "spec.walStorage.size" || targets[2].Field != "spec.tablespaces[name=archive].storage.pvcTemplate.resources.requests.storage" || targets[2].Declared != "5Gi" {
+ t.Errorf("expansion targets = %+v", targets)
+ }
+}
+
+func TestCNPGClusterStorage_NoPrometheusIsNeverZero(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgst2")
+ useCNPGProxyAPIServer(t, cnpgHealthyInstances)
+ prometheuspkg.Reset()
+
+ status, got, body := getCNPGStorage(t, "/api/cnpg/clusters/pgst2/pg-orders/storage")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.Usage.State != "noPrometheus" {
+ t.Fatalf("usage = %+v, want noPrometheus", got.Usage)
+ }
+ for _, in := range got.Instances {
+ for _, v := range in.Volumes {
+ if v.Usage.State != "noPrometheus" || v.Usage.Ratio != nil || v.Usage.UsedBytes != nil {
+ t.Errorf("%s usage = %+v", v.Claim, v.Usage)
+ }
+ }
+ }
+ if len(got.Findings) != 0 {
+ t.Errorf("findings without a measurement: %+v", got.Findings)
+ }
+ if !strings.Contains(body, `"storageClass":{"name":"fast","reason":"StorageClass fast not found"}`) {
+ t.Errorf("an unreadable class must not read as not expandable: %s", body)
+ }
+}
+
+func TestCNPGClusterStorage_ReadingTheClusterDoesNotImplyItsClaims(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgst3")
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgst3"}}
+ perms.SetCanI("get", cnpgsvc.Group, "clusters", "pgst3", true)
+ allow(perms, "", "persistentvolumeclaims", "pgst3", false)
+ allow(perms, "", "pods", "pgst3", false)
+ env.srv.permCache.Set("dba", nil, perms)
+
+ resp := env.authGet(t, "/api/cnpg/clusters/pgst3/pg-orders/storage", "dba", "")
+ defer resp.Body.Close()
+ if resp.StatusCode != http.StatusOK {
+ b, _ := io.ReadAll(resp.Body)
+ t.Fatalf("status = %d: %s", resp.StatusCode, b)
+ }
+ var got cnpgsvc.CNPGClusterStorageResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if got.Volumes.State != "denied" || got.Volumes.Grant == nil || *got.Volumes.Grant != (auth.Grant{Verb: "list", Resource: "persistentvolumeclaims", Namespace: "pgst3"}) {
+ t.Errorf("volumes = %+v", got.Volumes)
+ }
+ if got.WAL.State != "denied" || got.WAL.Grant == nil || *got.WAL.Grant != cnpgsvc.GrantListPods.In("pgst3") {
+ t.Errorf("wal = %+v", got.WAL)
+ }
+ for _, in := range got.Instances {
+ if len(in.Volumes) != 0 || in.WAL != nil {
+ t.Errorf("instance %s leaked facts: %+v", in.Name, in)
+ }
+ }
+
+ perms2 := &auth.UserPermissions{AllowedNamespaces: []string{"pgst3"}}
+ perms2.SetCanI("get", cnpgsvc.Group, "clusters", "pgst3", false)
+ env.srv.permCache.Set("nobody", nil, perms2)
+ denied := env.authGet(t, "/api/cnpg/clusters/pgst3/pg-orders/storage", "nobody", "")
+ denied.Body.Close()
+ if denied.StatusCode != http.StatusForbidden {
+ t.Errorf("without get clusters: %d, want 403", denied.StatusCode)
+ }
+}
+
+func TestCNPGFleetDisk(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgfd")
+ usePrometheusVolumeStats(t,
+ map[string]float64{"pg-orders-1": 1.7e9, "pg-orders-2": 1e9, "pg-orders-9": 1.99e9},
+ map[string]float64{"pg-orders-1": 2e9, "pg-orders-2": 2e9, "pg-orders-9": 2e9})
+
+ resp, err := http.Get(testServer.URL + "/api/cnpg/disk?namespaces=pgfd")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGFleetDiskResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if len(got.Clusters) != 1 {
+ t.Fatalf("clusters = %+v", got.Clusters)
+ }
+ d := got.Clusters[0]
+ // The foreign claim is fuller but is not this cluster's; the WAL claim has
+ // no series, so the answer is partial.
+ if d.State != "partial" || d.Claims != 3 || d.Measured != 2 || d.Max == nil || d.Max.Claim != "pg-orders-1" || d.Max.Instance != "pg-orders-1" {
+ t.Errorf("disk = %+v max=%+v", d, d.Max)
+ }
+}
+
+// Same-named claims in another cluster sharing this Prometheus: no value is
+// reported, since max() over both could hide this cluster's fullness.
+func TestCNPGFleetDiskRefusesAmbiguousClusterIdentity(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgfa")
+ useCNPGProxyAPIServer(t, cnpgHealthyInstances)
+ usePrometheusVolumeStatsFrom(t,
+ map[string]float64{"pg-orders-1": 1.7e9, "pg-orders-2": 1e9},
+ map[string]float64{"pg-orders-1": 2e9, "pg-orders-2": 2e9}, true)
+
+ resp, err := http.Get(testServer.URL + "/api/cnpg/disk?namespaces=pgfa")
+ if err != nil {
+ t.Fatal(err)
+ }
+ defer resp.Body.Close()
+ var got cnpgsvc.CNPGFleetDiskResponse
+ if err := json.NewDecoder(resp.Body).Decode(&got); err != nil {
+ t.Fatal(err)
+ }
+ if len(got.Clusters) != 1 || got.Clusters[0].State != "ambiguous" || got.Clusters[0].Max != nil || got.Clusters[0].Reason == "" {
+ t.Fatalf("disk = %+v", got.Clusters)
+ }
+
+ status, storage, body := getCNPGStorage(t, "/api/cnpg/clusters/pgfa/pg-orders/storage")
+ if status != http.StatusOK || storage.Usage.State != "ambiguous" {
+ t.Fatalf("storage = %d %s", status, body)
+ }
+ for _, in := range storage.Instances {
+ for _, v := range in.Volumes {
+ if v.Usage.Ratio != nil || v.Usage.UsedBytes != nil {
+ t.Errorf("volume %s carries a value from an ambiguous scope: %+v", v.Claim, v.Usage)
+ }
+ }
+ }
+}
+
+// An exporter whose queries fail still answers, without the WAL collector:
+// that instance was read only in part, and its slot list is not "no slots".
+func TestCNPGClusterStorage_WALWithoutCollectorIsPartial(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgst3")
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if c.port == "9187" {
+ _, _ = io.WriteString(w, "# TYPE cnpg_pg_postmaster_start_time gauge\ncnpg_pg_postmaster_start_time 1.7e9\n")
+ return
+ }
+ cnpgHealthyInstances(w, c)
+ })
+ status, got, body := getCNPGStorage(t, "/api/cnpg/clusters/pgst3/pg-orders/storage")
+ if status != http.StatusOK {
+ t.Fatalf("status = %d: %s", status, body)
+ }
+ if got.WAL.State != "partial" || got.WAL.Reason != "2 of 2 were read only in part" {
+ t.Errorf("wal coverage = %+v", got.WAL)
+ }
+ for name, in := range cnpgStorageInstances(t, got) {
+ if in.WAL == nil || in.WAL.Metrics.State != "partial" || in.WAL.Metrics.Reason == "" || in.WAL.Status.State != "ok" {
+ t.Errorf("%s WAL = %+v", name, in.WAL)
+ }
+ }
+}
+
+func TestCNPGClusterStorage_SlotInventoryWithoutRetention(t *testing.T) {
+ seedCNPGStorageCluster(t, "pgstinventory")
+ useCNPGProxyAPIServer(t, func(w http.ResponseWriter, c cnpgProxyCall) {
+ if c.port == "9187" {
+ _, _ = io.WriteString(w, "cnpg_collector_pg_wal{type=\"size\"} 83886080\n")
+ return
+ }
+ if c.pod == "pg-orders-1" {
+ _, _ = io.WriteString(w, `{"isPrimary":true,"replicationSlotsInfo":[{"slotName":"_cnpg_pg_orders_2","slotType":"physical","active":false}]}`)
+ return
+ }
+ _, _ = io.WriteString(w, `{"isPrimary":false,"replicationSlotsInfo":[]}`)
+ })
+ status, got, body := getCNPGStorage(t, "/api/cnpg/clusters/pgstinventory/pg-orders/storage")
+ if status != http.StatusOK {
+ t.Fatalf("%d: %s", status, body)
+ }
+ instances := cnpgStorageInstances(t, got)
+ slots := instances["pg-orders-1"].WAL.SlotInventory
+ if len(slots) != 1 || slots[0].Active || slots[0].RetainedBytes != nil {
+ t.Fatalf("inventory must survive absent retention: %+v", slots)
+ }
+ if slots := instances["pg-orders-2"].WAL.SlotInventory; slots == nil || len(slots) != 0 {
+ t.Fatalf("read empty inventory: %+v", slots)
+ }
+}
diff --git a/internal/server/cnpg_workspace.go b/internal/server/cnpg_workspace.go
index 1421e0218a..d6fd976f7a 100644
--- a/internal/server/cnpg_workspace.go
+++ b/internal/server/cnpg_workspace.go
@@ -1,166 +1,12 @@
package server
import (
- "context"
- "errors"
- "log"
"net/http"
- "slices"
- "sort"
- "time"
-
- corev1 "k8s.io/api/core/v1"
- metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
- "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "k8s.io/apimachinery/pkg/runtime/schema"
- "k8s.io/apimachinery/pkg/types"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
- "github.com/skyhook-io/radar/internal/k8s"
- bp "github.com/skyhook-io/radar/pkg/audit"
- "github.com/skyhook-io/radar/pkg/issuesapi"
-)
-
-const cnpgBarmanGroup = "barmancloud.cnpg.io"
-
-const (
- cnpgCoverageFull = "full"
- cnpgCoveragePartial = "partial"
- cnpgCoverageDenied = "denied"
- cnpgCoverageNotInstalled = "notInstalled"
- cnpgCoverageSyncing = "syncing"
- cnpgCoverageError = "error"
)
-const (
- cnpgWorkspacePodsKey = "pods"
- cnpgWorkspaceBackupsKey = "backups"
- cnpgWorkspaceSchedKey = "scheduledBackups"
- cnpgWorkspaceClusterKey = "clusters"
-)
-
-// cnpgBackupWindow bounds how far back settled Backups are returned. The newest
-// completed Backup per Cluster is kept regardless: it is the last-good-backup
-// fact the workspace reports.
-const cnpgBackupWindow = 7 * 24 * time.Hour
-
-const cnpgNoDeclarativeBackupCheckID = "cnpgNoDeclarativeBackup"
-
-type cnpgWorkspaceKind struct {
- key string
- group string
- kind string
- resource string
- clusterScoped bool
-}
-
-var cnpgWorkspaceKinds = []cnpgWorkspaceKind{
- {key: cnpgWorkspaceClusterKey, group: cnpgGroup, kind: "Cluster", resource: "clusters"},
- {key: cnpgWorkspaceBackupsKey, group: cnpgGroup, kind: "Backup", resource: "backups"},
- {key: cnpgWorkspaceSchedKey, group: cnpgGroup, kind: "ScheduledBackup", resource: "scheduledbackups"},
- {key: "poolers", group: cnpgGroup, kind: "Pooler", resource: "poolers"},
- {key: "databases", group: cnpgGroup, kind: "Database", resource: "databases"},
- {key: "publications", group: cnpgGroup, kind: "Publication", resource: "publications"},
- {key: "subscriptions", group: cnpgGroup, kind: "Subscription", resource: "subscriptions"},
- {key: "imageCatalogs", group: cnpgGroup, kind: "ImageCatalog", resource: "imagecatalogs"},
- {key: "clusterImageCatalogs", group: cnpgGroup, kind: "ClusterImageCatalog", resource: "clusterimagecatalogs", clusterScoped: true},
- {key: "objectStores", group: cnpgBarmanGroup, kind: "ObjectStore", resource: "objectstores"},
-}
-
-// CNPGWorkspaceCoverage states how much of one kind the caller could see.
-// DeniedNamespaces lists only namespaces already in the caller's scope, so it
-// may be omitted on a partial state; AllowedNamespaces is always set on a
-// partial state and is the authority for which namespaces were read.
-type CNPGWorkspaceCoverage struct {
- State string `json:"state"`
- DeniedNamespaces []string `json:"deniedNamespaces,omitempty"`
- AllowedNamespaces []string `json:"allowedNamespaces,omitempty"`
-}
-
-func cnpgCoverageOf(acc cnpgKindAccess, denied []string) CNPGWorkspaceCoverage {
- cov := CNPGWorkspaceCoverage{State: acc.state, DeniedNamespaces: denied}
- if acc.state == cnpgCoveragePartial {
- cov.AllowedNamespaces = make([]string, 0, len(acc.namespaces))
- for ns := range acc.namespaces {
- cov.AllowedNamespaces = append(cov.AllowedNamespaces, ns)
- }
- sort.Strings(cov.AllowedNamespaces)
- }
- return cov
-}
-
-// CNPGWorkspaceIssue is the subset of issuesapi.Issue the workspace renders.
-type CNPGWorkspaceIssue struct {
- ID string `json:"id"`
- Severity issuesapi.Severity `json:"severity"`
- Category issuesapi.Category `json:"category"`
- Kind string `json:"kind"`
- Group string `json:"group,omitempty"`
- Namespace string `json:"namespace,omitempty"`
- Name string `json:"name"`
- Reason string `json:"reason"`
- Message string `json:"message,omitempty"`
- Cause string `json:"cause,omitempty"`
- Action string `json:"action,omitempty"`
- FirstSeen time.Time `json:"first_seen,omitzero"`
-}
-
-// CNPGWorkspaceAuditFinding is one audit finding on a visible CNPG object.
-type CNPGWorkspaceAuditFinding struct {
- CheckID string `json:"checkId"`
- Severity string `json:"severity"`
- Kind string `json:"kind"`
- Group string `json:"group,omitempty"`
- Namespace string `json:"namespace"`
- Name string `json:"name"`
- Message string `json:"message"`
-}
-
-// CNPGWorkspaceResponse is GET /api/cnpg/workspace.
-type CNPGWorkspaceResponse struct {
- Installed bool `json:"installed"`
- Context string `json:"context"`
- Namespaces []string `json:"namespaces"`
- Coverage map[string]CNPGWorkspaceCoverage `json:"coverage"`
- Objects map[string][]any `json:"objects"`
- Issues []CNPGWorkspaceIssue `json:"issues"`
- Audit []CNPGWorkspaceAuditFinding `json:"audit"`
- BackupsOmitted int `json:"backupsOmitted"`
-}
-
-// cnpgKindAccess is the resolved read scope for one kind. all means every
-// namespace in the request's scope (or the cluster-scoped kind itself).
-type cnpgKindAccess struct {
- state string
- all bool
- namespaces map[string]bool
-}
-
-func (a cnpgKindAccess) covers(namespace string) bool {
- if a.state != cnpgCoverageFull && a.state != cnpgCoveragePartial {
- return false
- }
- return a.all || a.namespaces[namespace]
-}
-
-func newCNPGWorkspaceResponse(namespaces []string) CNPGWorkspaceResponse {
- resp := CNPGWorkspaceResponse{
- Context: k8s.ActiveClusterContext(),
- Namespaces: namespaces,
- Coverage: map[string]CNPGWorkspaceCoverage{},
- Objects: map[string][]any{},
- Issues: []CNPGWorkspaceIssue{},
- Audit: []CNPGWorkspaceAuditFinding{},
- }
- for _, k := range cnpgWorkspaceKinds {
- resp.Coverage[k.key] = CNPGWorkspaceCoverage{State: cnpgCoverageNotInstalled}
- resp.Objects[k.key] = []any{}
- }
- resp.Coverage[cnpgWorkspacePodsKey] = CNPGWorkspaceCoverage{State: cnpgCoverageNotInstalled}
- resp.Objects[cnpgWorkspacePodsKey] = []any{}
- return resp
-}
-
// handleCNPGWorkspace serves GET /api/cnpg/workspace: every CloudNativePG kind
// plus instance Pods, each authorized on its own. The generic resource list
// does not gate namespaced CRDs per kind, so it cannot tell "no access" from
@@ -169,509 +15,25 @@ func (s *Server) handleCNPGWorkspace(w http.ResponseWriter, r *http.Request) {
if !s.requireConnected(w) {
return
}
- cache := k8s.GetResourceCache()
+ reader := s.cnpgReader(r)
+ cache := reader.Observations.Cache
if cache == nil {
s.writeError(w, http.StatusServiceUnavailable, "Resource cache not available")
return
}
namespaces := s.parseNamespacesForUser(r)
- resp := newCNPGWorkspaceResponse(namespaces)
-
- disc := k8s.GetResourceDiscovery()
- if disc != nil {
- for _, k := range cnpgWorkspaceKinds {
- if _, ok := disc.GetGVRWithGroup(k.kind, k.group); ok {
- resp.Installed = true
- break
- }
- }
- if !resp.Installed {
- s.writeJSON(w, resp)
- return
- }
- }
-
- access := map[string]cnpgKindAccess{}
- items := map[string][]*unstructured.Unstructured{}
- for _, k := range cnpgWorkspaceKinds {
- if disc != nil {
- if _, ok := disc.GetGVRWithGroup(k.kind, k.group); !ok {
- access[k.key] = cnpgKindAccess{state: cnpgCoverageNotInstalled}
- continue
- }
- }
- acc, denied, list := s.cnpgWorkspaceReadKind(r, cache, k, namespaces)
- if acc.state != cnpgCoverageNotInstalled {
- resp.Installed = true
- }
- access[k.key] = acc
- items[k.key] = list
- resp.Coverage[k.key] = cnpgCoverageOf(acc, denied)
- }
- if !resp.Installed {
- s.writeJSON(w, resp)
- return
- }
-
- for _, k := range cnpgWorkspaceKinds {
- list := items[k.key]
- if k.key == cnpgWorkspaceBackupsKey {
- var omitted int
- list, omitted = windowCNPGBackups(list, time.Now())
- resp.BackupsOmitted = omitted
- } else {
- sortCNPGObjects(list)
- }
- out := make([]any, 0, len(list))
- for _, u := range list {
- out = append(out, u.Object)
- }
- resp.Objects[k.key] = out
- }
-
- podAccess, podDenied, pods, instancePods := s.cnpgWorkspaceReadPods(r, cache, namespaces, cnpgClusterUIDs(items[cnpgWorkspaceClusterKey]))
- access[cnpgWorkspacePodsKey] = podAccess
- resp.Coverage[cnpgWorkspacePodsKey] = cnpgCoverageOf(podAccess, podDenied)
- resp.Objects[cnpgWorkspacePodsKey] = pods
-
- resp.Issues = s.cnpgWorkspaceIssues(r, namespaces, access, instancePods)
- resp.Audit = cnpgWorkspaceAudit(items[cnpgWorkspaceClusterKey], items[cnpgWorkspaceSchedKey], access[cnpgWorkspaceSchedKey])
-
- s.writeJSON(w, resp)
-}
-
-// cnpgWorkspaceScope resolves where the caller may list one namespaced
-// resource: nil allowed means the whole request scope.
-//
-// denied names namespaces only when the candidate set came from the caller —
-// their view filter or their RBAC-allowed list. When the scope is "all" the
-// candidates are every namespace in Radar's cache, and naming the denied ones
-// would disclose namespaces the caller was never shown; partial then carries
-// the fact without the names.
-func (s *Server) cnpgWorkspaceScope(r *http.Request, namespaces []string, group, resource string) (allowed, denied []string, partial, any bool) {
- if noNamespaceAccess(namespaces) {
- return []string{}, nil, false, false
- }
- if s.canRead(r, group, resource, "", "list") {
- return namespaces, nil, false, true
- }
- candidates := namespaces
- if candidates == nil {
- candidates = allNamespaceNames()
- }
- if len(candidates) == 0 {
- return []string{}, nil, false, false
- }
- allowed = s.filterNamespacesByCanRead(r, group, resource, "list", candidates)
- partial = len(allowed) < len(candidates)
- if namespaces != nil {
- for _, ns := range candidates {
- if !slices.Contains(allowed, ns) {
- denied = append(denied, ns)
- }
- }
- sort.Strings(denied)
- }
- return allowed, denied, partial, len(allowed) > 0
-}
-
-func accessFromScope(allowed []string, partial bool) cnpgKindAccess {
- acc := cnpgKindAccess{state: cnpgCoverageFull, all: allowed == nil}
- if partial {
- acc.state = cnpgCoveragePartial
- }
- if allowed != nil {
- acc.namespaces = make(map[string]bool, len(allowed))
- for _, ns := range allowed {
- acc.namespaces[ns] = true
- }
- }
- return acc
-}
-
-func (s *Server) cnpgWorkspaceReadKind(r *http.Request, cache *k8s.ResourceCache, k cnpgWorkspaceKind, namespaces []string) (cnpgKindAccess, []string, []*unstructured.Unstructured) {
- var acc cnpgKindAccess
- var denied, readNamespaces []string
- if k.clusterScoped {
- if !s.canRead(r, k.group, k.resource, "", "list") {
- return cnpgKindAccess{state: cnpgCoverageDenied}, nil, nil
- }
- acc = cnpgKindAccess{state: cnpgCoverageFull, all: true}
- } else {
- allowed, d, partial, ok := s.cnpgWorkspaceScope(r, namespaces, k.group, k.resource)
- if !ok {
- return cnpgKindAccess{state: cnpgCoverageDenied}, nil, nil
- }
- acc, denied, readNamespaces = accessFromScope(allowed, partial), d, allowed
- }
-
- list, err := readCNPGKind(r.Context(), cache, k, readNamespaces)
- switch {
- case err == nil:
- return acc, denied, list
- case errors.Is(err, k8s.ErrUnknownDynamicKind):
- return cnpgKindAccess{state: cnpgCoverageNotInstalled}, nil, nil
- case errors.Is(err, errDynamicNotSynced):
- return cnpgKindAccess{state: cnpgCoverageSyncing}, nil, nil
- default:
- log.Printf("[cnpg] Failed to list %s.%s for workspace: %v", k.kind, k.group, err)
- return cnpgKindAccess{state: cnpgCoverageError}, nil, nil
- }
-}
-
-func readCNPGKind(ctx context.Context, cache *k8s.ResourceCache, k cnpgWorkspaceKind, namespaces []string) ([]*unstructured.Unstructured, error) {
- if namespaces == nil {
- return filterCNPGGroup(listDynamicSynced(ctx, cache, k.kind, k.group, ""))
- }
- var out []*unstructured.Unstructured
- for _, ns := range namespaces {
- list, err := filterCNPGGroup(listDynamicSynced(ctx, cache, k.kind, k.group, ns))
- if err != nil {
- return nil, err
- }
- out = append(out, list...)
- }
- return out, nil
-}
-
-// filterCNPGGroup drops anything whose apiVersion is not a CNPG group, so a
-// Velero Backup or a CAPI Cluster can never ride along on a kind-name match.
-func filterCNPGGroup(items []*unstructured.Unstructured, err error) ([]*unstructured.Unstructured, error) {
- if err != nil {
- return nil, err
- }
- out := items[:0:0]
- for _, u := range items {
- if u == nil {
- continue
- }
- if g := u.GroupVersionKind().Group; g != cnpgGroup && g != cnpgBarmanGroup {
- continue
- }
- out = append(out, u)
- }
- return out, nil
-}
-
-func sortCNPGObjects(items []*unstructured.Unstructured) {
- sort.SliceStable(items, func(i, j int) bool {
- if items[i].GetNamespace() != items[j].GetNamespace() {
- return items[i].GetNamespace() < items[j].GetNamespace()
- }
- return items[i].GetName() < items[j].GetName()
- })
-}
-
-func cnpgBackupTime(u *unstructured.Unstructured) time.Time {
- for _, field := range []string{"stoppedAt", "startedAt"} {
- if v, _, _ := unstructured.NestedString(u.Object, "status", field); v != "" {
- if t, err := time.Parse(time.RFC3339, v); err == nil {
- return t
- }
- }
- }
- return u.GetCreationTimestamp().Time
-}
-
-// windowCNPGBackups keeps every in-flight Backup, settled ones from the last
-// week, and each Cluster's newest completed Backup whatever its age. Sorted by
-// namespace, newest first within it.
-func windowCNPGBackups(items []*unstructured.Unstructured, now time.Time) ([]*unstructured.Unstructured, int) {
- newestCompleted := map[string]*unstructured.Unstructured{}
- for _, u := range items {
- if phase, _, _ := unstructured.NestedString(u.Object, "status", "phase"); phase != "completed" {
- continue
- }
- clusterName, _, _ := unstructured.NestedString(u.Object, "spec", "cluster", "name")
- key := u.GetNamespace() + "\x00" + clusterName
- if cur, ok := newestCompleted[key]; !ok || cnpgBackupTime(u).After(cnpgBackupTime(cur)) {
- newestCompleted[key] = u
- }
- }
- keepNewest := make(map[*unstructured.Unstructured]bool, len(newestCompleted))
- for _, u := range newestCompleted {
- keepNewest[u] = true
- }
-
- cutoff := now.Add(-cnpgBackupWindow)
- kept := make([]*unstructured.Unstructured, 0, len(items))
- omitted := 0
- for _, u := range items {
- phase, _, _ := unstructured.NestedString(u.Object, "status", "phase")
- settled := phase == "completed" || phase == "failed"
- if !settled || keepNewest[u] || !cnpgBackupTime(u).Before(cutoff) {
- kept = append(kept, u)
- continue
- }
- omitted++
- }
- sort.SliceStable(kept, func(i, j int) bool {
- if kept[i].GetNamespace() != kept[j].GetNamespace() {
- return kept[i].GetNamespace() < kept[j].GetNamespace()
- }
- ti, tj := cnpgBackupTime(kept[i]), cnpgBackupTime(kept[j])
- if !ti.Equal(tj) {
- return ti.After(tj)
- }
- return kept[i].GetName() < kept[j].GetName()
- })
- return kept, omitted
-}
-
-type cnpgWorkspacePodMeta struct {
- Name string `json:"name"`
- Namespace string `json:"namespace"`
- UID types.UID `json:"uid"`
- Labels map[string]string `json:"labels,omitempty"`
- OwnerReferences []metav1.OwnerReference `json:"ownerReferences,omitempty"`
- CreationTimestamp metav1.Time `json:"creationTimestamp"`
-}
-
-type cnpgWorkspaceContainerStatus struct {
- Name string `json:"name"`
- Ready bool `json:"ready"`
- RestartCount int32 `json:"restartCount"`
- State corev1.ContainerState `json:"state"`
-}
-
-type cnpgWorkspacePod struct {
- APIVersion string `json:"apiVersion"`
- Kind string `json:"kind"`
- Metadata cnpgWorkspacePodMeta `json:"metadata"`
- Spec struct {
- NodeName string `json:"nodeName,omitempty"`
- } `json:"spec"`
- Status struct {
- Phase corev1.PodPhase `json:"phase,omitempty"`
- PodIP string `json:"podIP,omitempty"`
- StartTime *metav1.Time `json:"startTime,omitempty"`
- Conditions []corev1.PodCondition `json:"conditions,omitempty"`
- ContainerStatuses []cnpgWorkspaceContainerStatus `json:"containerStatuses,omitempty"`
- } `json:"status"`
-}
-
-// isCNPGInstancePod requires the controller ownerReference to name a visible
-// Cluster by UID, not just by name: a label alone is something any workload
-// can carry, and a Pod left behind by a deleted Cluster must not be attributed
-// to a new one created under the same name. clusterUIDs is keyed ns/name.
-func isCNPGInstancePod(p *corev1.Pod, clusterUIDs map[string]types.UID) bool {
- clusterName := p.Labels["cnpg.io/cluster"]
- if clusterName == "" {
- return false
- }
- uid, ok := clusterUIDs[p.Namespace+"/"+clusterName]
- if !ok || uid == "" {
- return false
- }
- for _, ref := range p.OwnerReferences {
- if ref.Controller == nil || !*ref.Controller || ref.Kind != "Cluster" || ref.Name != clusterName || ref.UID != uid {
- continue
- }
- if gv, err := schema.ParseGroupVersion(ref.APIVersion); err == nil && gv.Group == cnpgGroup {
- return true
- }
- }
- return false
-}
-
-func cnpgClusterUIDs(clusters []*unstructured.Unstructured) map[string]types.UID {
- out := make(map[string]types.UID, len(clusters))
- for _, c := range clusters {
- out[c.GetNamespace()+"/"+c.GetName()] = c.GetUID()
- }
- return out
-}
-
-func trimCNPGPod(p *corev1.Pod) cnpgWorkspacePod {
- out := cnpgWorkspacePod{APIVersion: "v1", Kind: "Pod"}
- out.Metadata = cnpgWorkspacePodMeta{
- Name: p.Name,
- Namespace: p.Namespace,
- UID: p.UID,
- Labels: p.Labels,
- OwnerReferences: p.OwnerReferences,
- CreationTimestamp: p.CreationTimestamp,
- }
- out.Spec.NodeName = p.Spec.NodeName
- out.Status.Phase = p.Status.Phase
- out.Status.PodIP = p.Status.PodIP
- out.Status.StartTime = p.Status.StartTime
- out.Status.Conditions = p.Status.Conditions
- for _, cs := range p.Status.ContainerStatuses {
- out.Status.ContainerStatuses = append(out.Status.ContainerStatuses, cnpgWorkspaceContainerStatus{
- Name: cs.Name, Ready: cs.Ready, RestartCount: cs.RestartCount, State: cs.State,
- })
- }
- return out
-}
-
-// cnpgTypedScope resolves where the caller may list a typed kind and which of
-// those namespaces Radar's informer actually holds. The informer may itself be
-// namespace-scoped when Radar's own identity cannot list the kind
-// cluster-wide; what it does not hold is unread, not empty. read is nil for
-// "every namespace".
-func (s *Server) cnpgTypedScope(r *http.Request, cache *k8s.ResourceCache, namespaces []string, group, resource string) (acc cnpgKindAccess, denied, read []string) {
- allowed, denied, partial, ok := s.cnpgWorkspaceScope(r, namespaces, group, resource)
- if !ok {
- return cnpgKindAccess{state: cnpgCoverageDenied}, nil, []string{}
- }
- within := capacityNamespacesWithinCache(cache, resource, allowed)
- if within.unavailable {
- log.Printf("[cnpg] %s cache does not cover the requested scope", resource)
- return cnpgKindAccess{state: cnpgCoverageError}, nil, []string{}
- }
- if allowed != nil {
- for _, ns := range allowed {
- if slices.Contains(within.namespaces, ns) {
- continue
- }
- partial = true
- if namespaces != nil {
- denied = append(denied, ns)
- }
- }
- sort.Strings(denied)
- }
- acc = accessFromScope(within.namespaces, partial || within.partial)
- return acc, denied, within.namespaces
-}
-
-// cnpgWorkspaceReadPods returns the instance Pods of visible Clusters, plus
-// the namespace/name set of what it returned — the only Pods whose issues the
-// response may carry.
-func (s *Server) cnpgWorkspaceReadPods(r *http.Request, cache *k8s.ResourceCache, namespaces []string, clusterUIDs map[string]types.UID) (cnpgKindAccess, []string, []any, map[string]bool) {
- out := []any{}
- returned := map[string]bool{}
- acc, denied, read := s.cnpgTypedScope(r, cache, namespaces, "", "pods")
- if acc.state == cnpgCoverageDenied || acc.state == cnpgCoverageError {
- return acc, nil, out, returned
- }
- if cache.Pods() == nil {
- log.Printf("[cnpg] Pod cache unavailable for workspace")
- return cnpgKindAccess{state: cnpgCoverageError}, nil, out, returned
- }
-
- pods := listPodsScoped(cache.Pods(), read)
- sort.Slice(pods, func(i, j int) bool {
- if pods[i].Namespace != pods[j].Namespace {
- return pods[i].Namespace < pods[j].Namespace
- }
- return pods[i].Name < pods[j].Name
- })
- for _, p := range pods {
- if p != nil && isCNPGInstancePod(p, clusterUIDs) {
- out = append(out, trimCNPGPod(p))
- returned[p.Namespace+"/"+p.Name] = true
- }
- }
- return acc, denied, out, returned
+ s.writeJSON(w, reader.Workspace(r.Context(), cache, namespaces, reader.ClusterContext))
}
-var cnpgWorkspaceKeyByGroupKind = func() map[string]string {
- m := make(map[string]string, len(cnpgWorkspaceKinds))
- for _, k := range cnpgWorkspaceKinds {
- m[k.group+"/"+k.kind] = k.key
- }
- return m
-}()
-
-// cnpgWorkspaceIssues runs the same composition /api/issues serves, but reads
-// the flat evidence rows: the grouped view folds instance-Pod evidence into
-// the owning Cluster's row, which would hand Pod failure detail to a caller
-// who may list Clusters but not Pods. A row is kept only when its own subject
-// is visible here — a CNPG kind covered in its namespace, or an instance Pod
-// this response returned. IDs are the subject-derived IDs /api/issues uses.
-func (s *Server) cnpgWorkspaceIssues(r *http.Request, namespaces []string, access map[string]cnpgKindAccess, instancePods map[string]bool) []CNPGWorkspaceIssue {
- out := []CNPGWorkspaceIssue{}
- if noNamespaceAccess(namespaces) {
- return out
+func (s *Server) cnpgIssueRows(r *http.Request, namespaces []string) []issues.Issue {
+ if integration.NoNamespaceAccess(namespaces) {
+ return nil
}
provider := issues.NewCacheProvider()
if provider == nil {
- return out
- }
- composed, _ := issues.ComposeWithStats(provider, issues.Filters{
- Namespaces: namespaces,
- Limit: issues.NoLimit,
- CanReadClusterScoped: s.issueClusterScopedAccess(r),
- CanReadRelated: s.issueRelatedResourceAccess(r),
- })
- for _, iss := range composed {
- if !cnpgWorkspaceIssueVisible(iss, access, instancePods) {
- continue
- }
- out = append(out, CNPGWorkspaceIssue{
- ID: iss.ID,
- Severity: iss.Severity,
- Category: iss.Category,
- Kind: iss.Kind,
- Group: iss.Group,
- Namespace: iss.Namespace,
- Name: iss.Name,
- Reason: iss.Reason,
- Message: iss.Message,
- Cause: iss.Cause,
- Action: iss.Action,
- FirstSeen: iss.FirstSeen,
- })
- }
- return out
-}
-
-func cnpgWorkspaceIssueVisible(iss issues.Issue, access map[string]cnpgKindAccess, instancePods map[string]bool) bool {
- if iss.Group == "" && iss.Kind == "Pod" {
- return access[cnpgWorkspacePodsKey].covers(iss.Namespace) && instancePods[iss.Namespace+"/"+iss.Name]
- }
- if iss.Group != cnpgGroup && iss.Group != cnpgBarmanGroup {
- return false
- }
- key, ok := cnpgWorkspaceKeyByGroupKind[iss.Group+"/"+iss.Kind]
- return ok && access[key].covers(iss.Namespace)
-}
-
-// cnpgWorkspaceAudit reports the declarative-backup posture finding only for
-// Clusters whose namespace had its ScheduledBackups read: without that list,
-// "no schedule targets this cluster" is an absence nobody established.
-func cnpgWorkspaceAudit(clusters, scheduled []*unstructured.Unstructured, schedAccess cnpgKindAccess) []CNPGWorkspaceAuditFinding {
- out := []CNPGWorkspaceAuditFinding{}
- var subjects []*unstructured.Unstructured
- for _, c := range clusters {
- if schedAccess.covers(c.GetNamespace()) {
- subjects = append(subjects, c)
- }
- }
- if len(subjects) == 0 {
- return out
- }
- results := bp.RunChecks(&bp.CheckInput{
- CNPGClusters: subjects,
- CNPGScheduledBackups: scheduled,
- CNPGScheduledBackupsAuthoritative: true,
- })
- results = applyAuditSettings(results, getAuditConfig())
- if results == nil {
- return out
- }
- for _, f := range results.Findings {
- if f.CheckID != cnpgNoDeclarativeBackupCheckID {
- continue
- }
- out = append(out, CNPGWorkspaceAuditFinding{
- CheckID: f.CheckID,
- Severity: f.Severity,
- Kind: f.Kind,
- Group: f.Group,
- Namespace: f.Namespace,
- Name: f.Name,
- Message: f.Message,
- })
+ return nil
}
- sort.SliceStable(out, func(i, j int) bool {
- if out[i].Namespace != out[j].Namespace {
- return out[i].Namespace < out[j].Namespace
- }
- return out[i].Name < out[j].Name
- })
- return out
+ rows, _ := issues.ComposeWithStats(provider, issues.Filters{Namespaces: namespaces, Limit: issues.NoLimit, CanReadClusterScoped: s.issueClusterScopedAccess(r), CanReadRelated: s.issueRelatedResourceAccess(r), CanReadEvidence: s.issueEvidenceAccess(r)})
+ return rows
}
diff --git a/internal/server/cnpg_workspace_fallback_test.go b/internal/server/cnpg_workspace_fallback_test.go
new file mode 100644
index 0000000000..840f3ee279
--- /dev/null
+++ b/internal/server/cnpg_workspace_fallback_test.go
@@ -0,0 +1,297 @@
+package server
+
+import (
+ "context"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "slices"
+ "sort"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+ dynamicfake "k8s.io/client-go/dynamic/fake"
+ k8stesting "k8s.io/client-go/testing"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+// fallbackFixture is a dynamic cache whose identity may not list CRDs
+// cluster-wide, so it watches them namespace by namespace, starting with
+// fallbacks. forbid says where Radar's identity may not list either; fail makes
+// a namespace's Cluster list error without denying it; slow delays it by
+// slowList.
+type fallbackFixture struct {
+ fallbacks []string
+ forbid, fail, slow func(ns string) bool
+}
+
+const slowList = 2 * time.Second
+
+func inNamespaces(names ...string) func(string) bool {
+ return func(ns string) bool { return slices.Contains(names, ns) }
+}
+
+// seedCNPGFallbackCache seeds Clusters (namespace, name pairs) into f.
+func seedCNPGFallbackCache(t *testing.T, f fallbackFixture, clusters ...string) {
+ t.Helper()
+ release := make(chan struct{})
+ listKinds := map[schema.GroupVersionResource]string{}
+ for _, k := range cnpgWorkspaceTestKinds {
+ listKinds[schema.GroupVersionResource{Group: k.Group, Version: k.Version, Resource: k.Name}] = k.Kind + "List"
+ }
+ var objs []runtime.Object
+ for i := 0; i+1 < len(clusters); i += 2 {
+ objs = append(objs, cnpgObj("postgresql.cnpg.io/v1", "Cluster", clusters[i], clusters[i+1], nil, nil))
+ }
+ dyn := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), listKinds, objs...)
+ dyn.PrependReactor("list", "*", func(a k8stesting.Action) (bool, runtime.Object, error) {
+ ns := a.GetNamespace()
+ if ns == "" || (f.forbid != nil && f.forbid(ns)) {
+ gr := a.GetResource().GroupResource()
+ return true, nil, apierrors.NewForbidden(gr, "", errors.New("radar may not list here"))
+ }
+ if a.GetResource().Resource != "clusters" {
+ return false, nil, nil
+ }
+ if f.fail != nil && f.fail(ns) {
+ return true, nil, errors.New("temporary list failure")
+ }
+ if f.slow != nil && f.slow(ns) {
+ select {
+ case <-time.After(slowList):
+ case <-release:
+ }
+ }
+ return false, nil, nil
+ })
+ if err := k8s.InitTestDynamicResourceCacheWithFallbacks(dyn, cnpgWorkspaceTestKinds, f.fallbacks); err != nil {
+ t.Fatalf("seed cnpg: %v", err)
+ }
+ t.Cleanup(k8s.ResetTestDynamicState)
+ t.Cleanup(func() { close(release) })
+}
+
+func sortedNames(objs []any) []string {
+ names := objectNames(objs)
+ sort.Strings(names)
+ return names
+}
+
+// All namespaces: the Clusters Radar watches are listed and the rest reads as
+// not read, rather than the kind syncing forever.
+func TestCNPGWorkspace_FallbackCacheAllNamespaces(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}}, "a", "pg-a", "b", "pg-b")
+
+ got := getWorkspaceNoAuth(t, "")
+ cov := got.Coverage["clusters"]
+ if cov.State != integration.KindCoveragePartial {
+ t.Fatalf("clusters coverage = %+v, want partial", cov)
+ }
+ if len(cov.UncachedNamespaces) != 0 {
+ t.Errorf("uncached = %v: namespaces the caller did not name must not be disclosed", cov.UncachedNamespaces)
+ }
+ if names := sortedNames(got.Objects["clusters"]); len(names) != 1 || names[0] != "pg-a" {
+ t.Errorf("clusters = %v, want [pg-a]", names)
+ }
+ if cov := got.Coverage["clusterImageCatalogs"]; cov.State != integration.KindCoverageUncached {
+ t.Errorf("clusterImageCatalogs coverage = %+v: a cluster-scoped kind Radar may not watch is uncached, not an error", cov)
+ }
+}
+
+// A namespace the caller names is read on its own, which starts its watch even
+// outside the startup fallbacks; one Radar may not watch is named as not cached.
+func TestCNPGWorkspace_FallbackCacheNamedNamespaces(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, forbid: inNamespaces("c")}, "a", "pg-a", "b", "pg-b", "c", "pg-c")
+
+ b := getWorkspaceNoAuth(t, "?namespaces=b")
+ if cov := b.Coverage["clusters"]; cov.State != integration.KindCoverageFull {
+ t.Errorf("namespaces=b: coverage = %+v, want full", cov)
+ }
+ if names := sortedNames(b.Objects["clusters"]); len(names) != 1 || names[0] != "pg-b" {
+ t.Errorf("namespaces=b: clusters = %v, want [pg-b]", names)
+ }
+
+ ac := getWorkspaceNoAuth(t, "?namespaces=a,c")
+ cov := ac.Coverage["clusters"]
+ if cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 1 || cov.UncachedNamespaces[0] != "c" {
+ t.Errorf("namespaces=a,c: coverage = %+v, want partial with c not cached", cov)
+ }
+ if names := sortedNames(ac.Objects["clusters"]); len(names) != 1 || names[0] != "pg-a" {
+ t.Errorf("namespaces=a,c: clusters = %v, want [pg-a]", names)
+ }
+
+ c := getWorkspaceNoAuth(t, "?namespaces=c")
+ if cov := c.Coverage["clusters"]; cov.State != integration.KindCoverageUncached || len(cov.UncachedNamespaces) != 1 || cov.UncachedNamespaces[0] != "c" {
+ t.Errorf("namespaces=c: coverage = %+v, want uncached c", cov)
+ }
+}
+
+// A namespace whose watch has not synced is unread while the others are
+// listed, and is read once it syncs.
+func TestCNPGWorkspace_FallbackCacheRetriesUnsyncedNamespace(t *testing.T) {
+ var failing atomic.Bool
+ failing.Store(true)
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, fail: func(ns string) bool { return ns == "x" && failing.Load() }}, "a", "pg-a", "x", "pg-x")
+
+ first := getWorkspaceNoAuth(t, "?namespaces=a,x")
+ cov := first.Coverage["clusters"]
+ if cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 1 || cov.UncachedNamespaces[0] != "x" {
+ t.Fatalf("first read: coverage = %+v, want partial with x unread", cov)
+ }
+ if names := sortedNames(first.Objects["clusters"]); len(names) != 1 || names[0] != "pg-a" {
+ t.Errorf("first read: clusters = %v, want [pg-a]", names)
+ }
+
+ failing.Store(false)
+ deadline := time.Now().Add(20 * time.Second)
+ for {
+ got := getWorkspaceNoAuth(t, "?namespaces=a,x")
+ if got.Coverage["clusters"].State == integration.KindCoverageFull {
+ if names := sortedNames(got.Objects["clusters"]); len(names) != 2 || names[0] != "pg-a" || names[1] != "pg-x" {
+ t.Errorf("after sync: clusters = %v, want [pg-a pg-x]", names)
+ }
+ return
+ }
+ if time.Now().After(deadline) {
+ t.Fatalf("x never read after its list recovered: %+v", got.Coverage["clusters"])
+ }
+ time.Sleep(500 * time.Millisecond)
+ }
+}
+
+// Radar's identity may list Clusters in neither the cluster nor its startup
+// fallbacks, only in a namespace a caller named: the watch that read started
+// still serves an all-namespaces read instead of the kind failing.
+func TestCNPGWorkspace_FallbackCacheAllKeepsNamedWatches(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, forbid: inNamespaces("a")}, "a", "pg-a", "b", "pg-b")
+
+ if b := getWorkspaceNoAuth(t, "?namespaces=b"); b.Coverage["clusters"].State != integration.KindCoverageFull {
+ t.Fatalf("namespaces=b: coverage = %+v, want full", b.Coverage["clusters"])
+ }
+ all := getWorkspaceNoAuth(t, "")
+ cov := all.Coverage["clusters"]
+ if cov.State != integration.KindCoveragePartial {
+ t.Fatalf("all: coverage = %+v, want partial from the watch on b", cov)
+ }
+ if names := sortedNames(all.Objects["clusters"]); len(names) != 1 || names[0] != "pg-b" {
+ t.Errorf("all: clusters = %v, want [pg-b]", names)
+ }
+}
+
+// Waiting for unsynced watches shares one budget across the request, so
+// several stalled namespaces do not add up their waits.
+func TestCNPGWorkspace_FallbackCacheWaitsWithinOneBudget(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, fail: inNamespaces("x", "y", "z")}, "a", "pg-a", "x", "pg-x", "y", "pg-y", "z", "pg-z")
+
+ start := time.Now()
+ got := getWorkspaceNoAuth(t, "?namespaces=a,x,y,z")
+ if elapsed := time.Since(start); elapsed > dynamicSyncWait+2*time.Second {
+ t.Errorf("request took %v; stalled namespaces must share one %v wait", elapsed, dynamicSyncWait)
+ }
+ cov := got.Coverage["clusters"]
+ if cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 3 {
+ t.Errorf("coverage = %+v, want partial with x, y, z unread", cov)
+ }
+ if names := sortedNames(got.Objects["clusters"]); len(names) != 1 || names[0] != "pg-a" {
+ t.Errorf("clusters = %v, want [pg-a]", names)
+ }
+}
+
+// readCNPGClusters runs the workspace's kind reads as its handler does, under
+// one budget, and returns the Clusters. The issues the handler composes
+// afterwards read through the same watches and are not bounded by it.
+func readCNPGClusters(namespaces []string) (integration.KindCoverage, []*unstructured.Unstructured, time.Duration) {
+ r := httptest.NewRequest(http.MethodGet, "/api/cnpg/workspace", nil)
+ budget := newSyncBudget(r.Context())
+ start := time.Now()
+ var cov integration.KindCoverage
+ var clusters []*unstructured.Unstructured
+ for _, k := range cnpgWorkspaceFixtureKinds {
+ acc, list := testServerSrv.readWorkspaceKind(r, k8s.GetResourceCache(), k, namespaces, []string{"postgresql.cnpg.io", "barmancloud.cnpg.io"}, budget)
+ if k.Key == "clusters" {
+ cov, clusters = acc.Coverage(), list
+ }
+ }
+ return cov, clusters, time.Since(start)
+}
+
+// Starting a watch probes the apiserver, and slow probes spend the same budget
+// as slow syncs. Watches the budget cut short carry on, and a later read finds
+// them.
+func TestCNPGWorkspace_FallbackCacheSlowWatchStartsWithinOneBudget(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, slow: inNamespaces("x", "y", "z")}, "a", "pg-a", "x", "pg-x", "y", "pg-y", "z", "pg-z")
+ namespaces := []string{"a", "x", "y", "z"}
+
+ cov, clusters, elapsed := readCNPGClusters(namespaces)
+ if elapsed > dynamicSyncWait+time.Second {
+ t.Errorf("kind reads took %v; starting three slow watches must share one %v budget", elapsed, dynamicSyncWait)
+ }
+ if cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 3 {
+ t.Errorf("coverage = %+v, want partial with x, y, z unread", cov)
+ }
+ if len(clusters) != 1 || clusters[0].GetName() != "pg-a" {
+ t.Errorf("clusters = %d objects, want only pg-a", len(clusters))
+ }
+
+ deadline := time.Now().Add(20 * time.Second)
+ for {
+ cov, clusters, _ := readCNPGClusters(namespaces)
+ if cov.State == integration.KindCoverageFull {
+ if len(clusters) != 4 {
+ t.Errorf("after the watches started: %d clusters, want 4", len(clusters))
+ }
+ return
+ }
+ if time.Now().After(deadline) {
+ t.Fatalf("the watches never finished starting: %+v", cov)
+ }
+ time.Sleep(500 * time.Millisecond)
+ }
+}
+
+// All namespaces over several stalled fallback watches waits once, then reads
+// the watches that synced.
+func TestCNPGWorkspace_FallbackCacheAllWaitsWithinOneBudget(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a", "x", "y", "z"}, fail: inNamespaces("x", "y", "z")}, "a", "pg-a", "x", "pg-x")
+
+ start := time.Now()
+ got := getWorkspaceNoAuth(t, "")
+ if elapsed := time.Since(start); elapsed > dynamicSyncWait+time.Second {
+ t.Errorf("request took %v; stalled fallback watches must share one %v wait", elapsed, dynamicSyncWait)
+ }
+ if cov := got.Coverage["clusters"]; cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 0 {
+ t.Errorf("coverage = %+v, want partial, naming nothing", cov)
+ }
+ if names := sortedNames(got.Objects["clusters"]); len(names) != 1 || names[0] != "pg-a" {
+ t.Errorf("clusters = %v, want [pg-a]", names)
+ }
+}
+
+// A cancelled request stops waiting even with budget left.
+func TestSyncBudget_StopsWhenRequestEnds(t *testing.T) {
+ seedCNPGFallbackCache(t, fallbackFixture{fallbacks: []string{"a"}, slow: inNamespaces("x")}, "x", "pg-x")
+ gvr, ok := k8s.GetResourceDiscovery().GetGVRWithGroup("Cluster", "postgresql.cnpg.io")
+ if !ok {
+ t.Fatal("Cluster not discovered")
+ }
+ ctx, cancel := context.WithCancel(context.Background())
+ defer cancel()
+ budget := &syncBudget{ctx: ctx, deadline: time.Now().Add(time.Minute)}
+ time.AfterFunc(200*time.Millisecond, cancel)
+
+ start := time.Now()
+ _, err := budget.listBlocking(k8s.GetDynamicResourceCache(), gvr, "x")
+ if !errors.Is(err, integration.ErrDynamicNotSynced) {
+ t.Errorf("err = %v, want errDynamicNotSynced", err)
+ }
+ if elapsed := time.Since(start); elapsed > time.Second {
+ t.Errorf("listBlocking took %v after its request was cancelled at 200ms", elapsed)
+ }
+}
diff --git a/internal/server/cnpg_workspace_jobpods_test.go b/internal/server/cnpg_workspace_jobpods_test.go
new file mode 100644
index 0000000000..f07e4bc729
--- /dev/null
+++ b/internal/server/cnpg_workspace_jobpods_test.go
@@ -0,0 +1,138 @@
+package server
+
+import (
+ "context"
+ "testing"
+ "time"
+
+ batchv1 "k8s.io/api/batch/v1"
+ corev1 "k8s.io/api/core/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+)
+
+func cnpgJob(ns, name, uid string, owner metav1.OwnerReference) *batchv1.Job {
+ return &batchv1.Job{ObjectMeta: metav1.ObjectMeta{
+ Namespace: ns, Name: name, UID: types.UID(uid),
+ Labels: map[string]string{"cnpg.io/cluster": owner.Name, "cnpg.io/jobRole": "initdb"},
+ OwnerReferences: []metav1.OwnerReference{owner},
+ }}
+}
+
+func clusterRef(name, uid string) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: name, UID: types.UID(uid), Controller: boolPtr(true)}
+}
+
+func jobRef(name, uid string) metav1.OwnerReference {
+ return metav1.OwnerReference{APIVersion: "batch/v1", Kind: "Job", Name: name, UID: types.UID(uid), Controller: boolPtr(true)}
+}
+
+// An initdb Pod the scheduler could not place, as CloudNativePG labels it.
+func cnpgJobPod(ns, name, cluster string, owner metav1.OwnerReference) *corev1.Pod {
+ return &corev1.Pod{
+ ObjectMeta: metav1.ObjectMeta{
+ Namespace: ns, Name: name,
+ Labels: map[string]string{"cnpg.io/cluster": cluster, "cnpg.io/jobRole": "initdb", "cnpg.io/instanceName": cluster + "-1"},
+ OwnerReferences: []metav1.OwnerReference{owner},
+ },
+ Spec: corev1.PodSpec{Containers: []corev1.Container{{Name: "initdb", Image: "pg:17"}}},
+ Status: corev1.PodStatus{
+ Phase: corev1.PodPending,
+ Conditions: []corev1.PodCondition{{
+ Type: corev1.PodScheduled, Status: corev1.ConditionFalse, Reason: corev1.PodReasonUnschedulable,
+ Message: "0/2 nodes are available: 2 Too many pods.",
+ }},
+ },
+ }
+}
+
+func seedCNPGJobs(t *testing.T, jobs ...*batchv1.Job) {
+ t.Helper()
+ for _, j := range jobs {
+ if _, err := testFakeClient.BatchV1().Jobs(j.Namespace).Create(context.Background(), j, metav1.CreateOptions{}); err != nil {
+ t.Fatalf("create job %s: %v", j.Name, err)
+ }
+ t.Cleanup(func() {
+ _ = testFakeClient.BatchV1().Jobs(j.Namespace).Delete(context.Background(), j.Name, metav1.DeleteOptions{})
+ })
+ }
+ lister := k8s.GetResourceCache().Jobs()
+ deadline := time.Now().Add(5 * time.Second)
+ for _, j := range jobs {
+ for {
+ if _, err := lister.Jobs(j.Namespace).Get(j.Name); err == nil {
+ break
+ }
+ if time.Now().After(deadline) {
+ t.Fatalf("job %s did not reach the cache", j.Name)
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+ }
+}
+
+// The scheduler's verdict on a Cluster's own Job Pod (here initdb) reaches the
+// Cluster only for a caller who may read both the Pods and the Jobs; a Pod that
+// names an earlier Job of the same name is never adopted.
+func TestCNPGWorkspace_JobPodsFollowJobAccess(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", "pgjob", "pg-new", map[string]any{"instances": int64(1)}, nil), "new-uid"),
+ )
+ seedCNPGJobs(t, cnpgJob("pgjob", "pg-new-1-initdb", "job-uid", clusterRef("pg-new", "new-uid")))
+ stuck := cnpgJobPod("pgjob", "pg-new-1-initdb-abcde", "pg-new", jobRef("pg-new-1-initdb", "job-uid"))
+ stale := cnpgJobPod("pgjob", "pg-new-1-initdb-old", "pg-new", jobRef("pg-new-1-initdb", "old-job-uid"))
+ seedCNPGPods(t, stuck, stale)
+
+ env := newAuthTestServer(t)
+ for _, u := range []struct {
+ name string
+ jobs bool
+ }{{"with-jobs", true}, {"pods-only", false}} {
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgjob"}}
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
+ allow(perms, "", "pods", "", true)
+ allow(perms, "", "pods", "pgjob", true)
+ allow(perms, "batch", "jobs", "", u.jobs)
+ allow(perms, "batch", "jobs", "pgjob", u.jobs)
+ env.srv.permCache.Set(u.name, nil, perms)
+ }
+
+ withJobs := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "with-jobs", ""))
+ if names := objectNames(withJobs.JobPods); len(names) != 1 || names[0] != stuck.Name {
+ t.Errorf("with Job access: jobPods = %v, want only %s", names, stuck.Name)
+ }
+ if withJobs.JobCoverage == nil || withJobs.JobCoverage.State != integration.KindCoverageFull {
+ t.Errorf("with Job access: jobCoverage = %+v, want full", withJobs.JobCoverage)
+ }
+ if containsName(withJobs.Objects["pods"], stuck.Name) {
+ t.Error("a Job Pod was returned as an instance Pod")
+ }
+ found := false
+ for _, iss := range withJobs.Issues {
+ if iss.Kind == "Pod" && iss.Name == stale.Name {
+ t.Errorf("the stale Pod's issue reached the Cluster: %+v", iss)
+ }
+ found = found || (iss.Kind == "Pod" && iss.Name == stuck.Name)
+ }
+ if !found {
+ t.Errorf("with Job access: no issue for the unschedulable initdb Pod, got %+v", withJobs.Issues)
+ }
+
+ podsOnly := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "pods-only", ""))
+ if len(podsOnly.JobPods) != 0 {
+ t.Errorf("without Job access: jobPods = %v, want none", objectNames(podsOnly.JobPods))
+ }
+ if podsOnly.JobCoverage == nil || podsOnly.JobCoverage.State != integration.KindCoverageDenied {
+ t.Errorf("without Job access: jobCoverage = %+v, want denied", podsOnly.JobCoverage)
+ }
+ for _, iss := range podsOnly.Issues {
+ if iss.Kind == "Pod" && (iss.Name == stuck.Name || iss.Name == stale.Name) {
+ t.Errorf("without Job access: a Job Pod's issue reached the Cluster: %+v", iss)
+ }
+ }
+}
diff --git a/internal/server/cnpg_workspace_test.go b/internal/server/cnpg_workspace_test.go
index 02152adb16..4e241fa775 100644
--- a/internal/server/cnpg_workspace_test.go
+++ b/internal/server/cnpg_workspace_test.go
@@ -4,6 +4,7 @@ import (
"context"
"encoding/json"
"net/http"
+ "reflect"
"strings"
"testing"
"time"
@@ -17,17 +18,34 @@ import (
dynamicfake "k8s.io/client-go/dynamic/fake"
"github.com/skyhook-io/radar/internal/auth"
+ cnpgsvc "github.com/skyhook-io/radar/internal/cnpg"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/topology"
)
func cnpgTestResource(group, kind, resource string, namespaced bool) k8s.APIResource {
return k8s.APIResource{Group: group, Version: "v1", Kind: kind, Name: resource, Namespaced: namespaced, IsCRD: true, Verbs: []string{"get", "list", "watch"}}
}
+var cnpgWorkspaceFixtureKinds = []integration.WorkspaceKind{
+ {Key: "clusters", Group: "postgresql.cnpg.io", Kind: "Cluster", Resource: "clusters"},
+ {Key: "backups", Group: "postgresql.cnpg.io", Kind: "Backup", Resource: "backups"},
+ {Key: "scheduledBackups", Group: "postgresql.cnpg.io", Kind: "ScheduledBackup", Resource: "scheduledbackups"},
+ {Key: "poolers", Group: "postgresql.cnpg.io", Kind: "Pooler", Resource: "poolers"},
+ {Key: "databases", Group: "postgresql.cnpg.io", Kind: "Database", Resource: "databases"},
+ {Key: "publications", Group: "postgresql.cnpg.io", Kind: "Publication", Resource: "publications"},
+ {Key: "subscriptions", Group: "postgresql.cnpg.io", Kind: "Subscription", Resource: "subscriptions"},
+ {Key: "databaseRoles", Group: "postgresql.cnpg.io", Kind: "DatabaseRole", Resource: "databaseroles"},
+ {Key: "imageCatalogs", Group: "postgresql.cnpg.io", Kind: "ImageCatalog", Resource: "imagecatalogs"},
+ {Key: "clusterImageCatalogs", Group: "postgresql.cnpg.io", Kind: "ClusterImageCatalog", Resource: "clusterimagecatalogs", ClusterScoped: true},
+ {Key: "objectStores", Group: "barmancloud.cnpg.io", Kind: "ObjectStore", Resource: "objectstores"},
+}
+
var cnpgWorkspaceTestKinds = func() []k8s.APIResource {
var out []k8s.APIResource
- for _, k := range cnpgWorkspaceKinds {
- out = append(out, cnpgTestResource(k.group, k.kind, k.resource, !k.clusterScoped))
+ for _, k := range cnpgWorkspaceFixtureKinds {
+ out = append(out, cnpgTestResource(k.Group, k.Kind, k.Resource, !k.ClusterScoped))
}
out = append(out,
cnpgTestResource(veleroGroup, "Backup", "backups", true),
@@ -72,20 +90,20 @@ func cnpgBackup(ns, name, cluster, phase string, stoppedAt time.Time) *unstructu
return cnpgObj("postgresql.cnpg.io/v1", "Backup", ns, name, map[string]any{"cluster": map[string]any{"name": cluster}}, status)
}
-func decodeWorkspace(t *testing.T, resp *http.Response) CNPGWorkspaceResponse {
+func decodeWorkspace(t *testing.T, resp *http.Response) cnpgsvc.CNPGWorkspaceResponse {
t.Helper()
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
t.Fatalf("status = %d, want 200", resp.StatusCode)
}
- var out CNPGWorkspaceResponse
+ var out cnpgsvc.CNPGWorkspaceResponse
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
t.Fatalf("decode: %v", err)
}
return out
}
-func getWorkspaceNoAuth(t *testing.T, query string) CNPGWorkspaceResponse {
+func getWorkspaceNoAuth(t *testing.T, query string) cnpgsvc.CNPGWorkspaceResponse {
t.Helper()
resp, err := http.Get(testServer.URL + "/api/cnpg/workspace" + query)
if err != nil {
@@ -114,11 +132,11 @@ func containsName(objs []any, name string) bool {
return false
}
-func assertEveryKey(t *testing.T, got CNPGWorkspaceResponse) {
+func assertEveryKey(t *testing.T, got cnpgsvc.CNPGWorkspaceResponse) {
t.Helper()
- keys := []string{cnpgWorkspacePodsKey}
- for _, k := range cnpgWorkspaceKinds {
- keys = append(keys, k.key)
+ keys := []string{"pods"}
+ for _, k := range cnpgWorkspaceFixtureKinds {
+ keys = append(keys, k.Key)
}
for _, k := range keys {
if _, ok := got.Coverage[k]; !ok {
@@ -138,7 +156,7 @@ func TestCNPGWorkspace_NotInstalled(t *testing.T) {
}
assertEveryKey(t, got)
for k, c := range got.Coverage {
- if c.State != cnpgCoverageNotInstalled {
+ if c.State != integration.KindCoverageNotInstalled {
t.Errorf("coverage[%s] = %q, want notInstalled", k, c.State)
}
}
@@ -193,6 +211,32 @@ func cnpgPod(ns, name, clusterLabel string, owners ...metav1.OwnerReference) *co
}
}
+// Each returned object's GitOps or Helm manager comes from the server's one
+// metadata-only detection: an Argo CD tracking ID in the apps-in-any-namespace
+// form names the Application's namespace, the default form names none (and
+// none is invented), Flux needs both kustomization labels.
+func TestCNPGWorkspace_ManagedByFromObjectMetadata(t *testing.T) {
+ cluster := cnpgObj("postgresql.cnpg.io/v1", "Cluster", "db", "pg", nil, nil)
+ cluster.SetAnnotations(map[string]string{"argocd.argoproj.io/tracking-id": "team-a_orders:postgresql.cnpg.io/Cluster:db/pg"})
+ database := cnpgObj("postgresql.cnpg.io/v1", "Database", "db", "app", map[string]any{"cluster": map[string]any{"name": "pg"}}, nil)
+ database.SetAnnotations(map[string]string{"argocd.argoproj.io/tracking-id": "orders:postgresql.cnpg.io/Database:db/app"})
+ pooler := cnpgObj("postgresql.cnpg.io/v1", "Pooler", "db", "pg-rw", map[string]any{"cluster": map[string]any{"name": "pg"}}, nil)
+ pooler.SetLabels(map[string]string{"kustomize.toolkit.fluxcd.io/name": "databases", "kustomize.toolkit.fluxcd.io/namespace": "flux-system"})
+ sched := cnpgObj("postgresql.cnpg.io/v1", "ScheduledBackup", "db", "nightly", map[string]any{"cluster": map[string]any{"name": "pg"}, "schedule": "0 0 2 * * *"}, nil)
+ sched.SetLabels(map[string]string{"app.kubernetes.io/name": "pg"})
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds, cluster, database, pooler, sched)
+
+ got := getWorkspaceNoAuth(t, "")
+ want := map[string]topology.ResourceRef{
+ "Cluster/db/pg": {Kind: "Application", Group: "argoproj.io", Namespace: "team-a", Name: "orders"},
+ "Database/db/app": {Kind: "Application", Group: "argoproj.io", Name: "orders"},
+ "Pooler/db/pg-rw": {Kind: "Kustomization", Group: "kustomize.toolkit.fluxcd.io", Namespace: "flux-system", Name: "databases"},
+ }
+ if !reflect.DeepEqual(got.ManagedBy, want) {
+ t.Errorf("managedBy = %+v, want %+v", got.ManagedBy, want)
+ }
+}
+
func TestCNPGWorkspace_AuthDisabledReturnsEverythingAndOnlyOwnedInstancePods(t *testing.T) {
seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
withUID(cnpgObj("postgresql.cnpg.io/v1", "Cluster", "pgws", "pg-orders", map[string]any{"instances": int64(1)}, nil), "c-uid"),
@@ -216,7 +260,7 @@ func TestCNPGWorkspace_AuthDisabledReturnsEverythingAndOnlyOwnedInstancePods(t *
}
assertEveryKey(t, got)
for k, c := range got.Coverage {
- if c.State != cnpgCoverageFull {
+ if c.State != integration.KindCoverageFull {
t.Errorf("coverage[%s] = %+v, want full with auth disabled", k, c)
}
}
@@ -271,7 +315,7 @@ func TestCNPGWorkspace_DeniedKindAndItsIssuesAreWithheld(t *testing.T) {
)
// Warm the Backup informer so the issues engine can see the failed Backup
// whichever caller asks; otherwise a denied answer would pass vacuously.
- if _, err := listDynamicSynced(context.Background(), k8s.GetResourceCache(), "Backup", cnpgGroup, ""); err != nil {
+ if _, err := listDynamicSynced(context.Background(), k8s.GetResourceCache(), "Backup", cnpgsvc.Group, ""); err != nil {
t.Fatalf("warm backups: %v", err)
}
@@ -281,13 +325,13 @@ func TestCNPGWorkspace_DeniedKindAndItsIssuesAreWithheld(t *testing.T) {
backupsListed bool
}{{"reader", true}, {"no-backups", false}} {
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- allow(perms, cnpgGroup, "clusters", "", true)
- allow(perms, cnpgGroup, "backups", "", u.backupsListed)
- allow(perms, cnpgGroup, "backups", "pg", u.backupsListed)
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
+ allow(perms, cnpgsvc.Group, "backups", "", u.backupsListed)
+ allow(perms, cnpgsvc.Group, "backups", "pg", u.backupsListed)
env.srv.permCache.Set(u.name, nil, perms)
}
- hasBackupIssue := func(got CNPGWorkspaceResponse) bool {
+ hasBackupIssue := func(got cnpgsvc.CNPGWorkspaceResponse) bool {
for _, iss := range got.Issues {
if iss.Kind == "Backup" && iss.Name == "pg-orders-broken" {
return true
@@ -297,7 +341,7 @@ func TestCNPGWorkspace_DeniedKindAndItsIssuesAreWithheld(t *testing.T) {
}
control := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "reader", ""))
- if control.Coverage["backups"].State != cnpgCoverageFull || !containsName(control.Objects["backups"], "pg-orders-broken") {
+ if control.Coverage["backups"].State != integration.KindCoverageFull || !containsName(control.Objects["backups"], "pg-orders-broken") {
t.Fatalf("control: backups coverage=%+v objects=%v", control.Coverage["backups"], objectNames(control.Objects["backups"]))
}
if !hasBackupIssue(control) {
@@ -305,7 +349,7 @@ func TestCNPGWorkspace_DeniedKindAndItsIssuesAreWithheld(t *testing.T) {
}
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "no-backups", ""))
- if got.Coverage["backups"].State != cnpgCoverageDenied {
+ if got.Coverage["backups"].State != integration.KindCoverageDenied {
t.Errorf("backups coverage = %+v, want denied", got.Coverage["backups"])
}
if len(got.Objects["backups"]) != 0 {
@@ -314,7 +358,7 @@ func TestCNPGWorkspace_DeniedKindAndItsIssuesAreWithheld(t *testing.T) {
if hasBackupIssue(got) {
t.Error("an issue on a Backup the caller cannot list was returned")
}
- if got.Coverage["clusters"].State != cnpgCoverageFull || !containsName(got.Objects["clusters"], "pg-orders") {
+ if got.Coverage["clusters"].State != integration.KindCoverageFull || !containsName(got.Objects["clusters"], "pg-orders") {
t.Errorf("clusters coverage=%+v objects=%v, want full", got.Coverage["clusters"], objectNames(got.Objects["clusters"]))
}
}
@@ -327,14 +371,14 @@ func TestCNPGWorkspace_PartialNamespaceCoverage(t *testing.T) {
)
env := newAuthTestServer(t)
perms := &auth.UserPermissions{AllowedNamespaces: []string{"a", "b"}}
- allow(perms, cnpgGroup, "clusters", "", false)
- allow(perms, cnpgGroup, "clusters", "a", true)
- allow(perms, cnpgGroup, "clusters", "b", false)
+ allow(perms, cnpgsvc.Group, "clusters", "", false)
+ allow(perms, cnpgsvc.Group, "clusters", "a", true)
+ allow(perms, cnpgsvc.Group, "clusters", "b", false)
env.srv.permCache.Set("scoped", nil, perms)
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "scoped", ""))
cov := got.Coverage["clusters"]
- if cov.State != cnpgCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "b" {
+ if cov.State != integration.KindCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "b" {
t.Errorf("clusters coverage = %+v, want partial denied [b]", cov)
}
if len(cov.AllowedNamespaces) != 1 || cov.AllowedNamespaces[0] != "a" {
@@ -347,7 +391,7 @@ func TestCNPGWorkspace_PartialNamespaceCoverage(t *testing.T) {
// A view filter narrows the scope; the denied list never grows past it.
filtered := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace?namespaces=a", "scoped", ""))
- if filtered.Coverage["clusters"].State != cnpgCoverageFull {
+ if filtered.Coverage["clusters"].State != integration.KindCoverageFull {
t.Errorf("filtered to a: coverage = %+v, want full", filtered.Coverage["clusters"])
}
if len(filtered.Namespaces) != 1 || filtered.Namespaces[0] != "a" {
@@ -361,60 +405,25 @@ func TestCNPGWorkspace_ClusterImageCatalogNeedsClusterScopeGrant(t *testing.T) {
)
env := newAuthTestServer(t)
nsOnly := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- nsOnly.SetCanI("list", cnpgGroup, "clusterimagecatalogs", "pg", true)
- allow(nsOnly, cnpgGroup, "clusterimagecatalogs", "", false)
+ nsOnly.SetCanI("list", cnpgsvc.Group, "clusterimagecatalogs", "pg", true)
+ allow(nsOnly, cnpgsvc.Group, "clusterimagecatalogs", "", false)
env.srv.permCache.Set("ns-only", nil, nsOnly)
clusterWide := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- allow(clusterWide, cnpgGroup, "clusterimagecatalogs", "", true)
+ allow(clusterWide, cnpgsvc.Group, "clusterimagecatalogs", "", true)
env.srv.permCache.Set("cluster-wide", nil, clusterWide)
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "ns-only", ""))
- if got.Coverage["clusterImageCatalogs"].State != cnpgCoverageDenied || len(got.Objects["clusterImageCatalogs"]) != 0 {
+ if got.Coverage["clusterImageCatalogs"].State != integration.KindCoverageDenied || len(got.Objects["clusterImageCatalogs"]) != 0 {
t.Errorf("namespace-level grant exposed ClusterImageCatalogs: %+v %v", got.Coverage["clusterImageCatalogs"], objectNames(got.Objects["clusterImageCatalogs"]))
}
got = decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace?namespaces=pg", "cluster-wide", ""))
- if got.Coverage["clusterImageCatalogs"].State != cnpgCoverageFull || !containsName(got.Objects["clusterImageCatalogs"], "pg-fleet") {
+ if got.Coverage["clusterImageCatalogs"].State != integration.KindCoverageFull || !containsName(got.Objects["clusterImageCatalogs"], "pg-fleet") {
t.Errorf("cluster-scope grant: %+v %v, want full with pg-fleet regardless of the view filter", got.Coverage["clusterImageCatalogs"], objectNames(got.Objects["clusterImageCatalogs"]))
}
}
-func TestWindowCNPGBackups(t *testing.T) {
- now := time.Date(2026, 9, 29, 12, 0, 0, 0, time.UTC)
- day := 24 * time.Hour
- items := []*unstructured.Unstructured{
- cnpgBackup("pg", "running-old", "orders", "running", time.Time{}),
- cnpgBackup("pg", "orders-recent", "orders", "completed", now.Add(-2*day)),
- cnpgBackup("pg", "orders-old", "orders", "completed", now.Add(-20*day)),
- cnpgBackup("pg", "orders-failed-old", "orders", "failed", now.Add(-30*day)),
- cnpgBackup("pg", "orders-failed-new", "orders", "failed", now.Add(-1*day)),
- cnpgBackup("pg", "billing-only-old", "billing", "completed", now.Add(-40*day)),
- cnpgBackup("pg", "billing-older", "billing", "completed", now.Add(-50*day)),
- cnpgBackup("aa", "other-ns", "orders", "completed", now.Add(-60*day)),
- }
- // A running Backup older than the window stays: in flight is never settled.
- items[0].Object["metadata"].(map[string]any)["creationTimestamp"] = now.Add(-90 * day).Format(time.RFC3339)
-
- kept, omitted := windowCNPGBackups(items, now)
- var names []string
- for _, u := range kept {
- names = append(names, u.GetName())
- }
- want := []string{"other-ns", "orders-failed-new", "orders-recent", "billing-only-old", "running-old"}
- if len(names) != len(want) {
- t.Fatalf("kept = %v, want %v", names, want)
- }
- for i := range want {
- if names[i] != want[i] {
- t.Fatalf("kept = %v, want %v (namespace, then newest first)", names, want)
- }
- }
- if omitted != 3 {
- t.Errorf("omitted = %d, want 3 (orders-old, orders-failed-old, billing-older)", omitted)
- }
-}
-
func TestCNPGWorkspace_BackupsOmittedIsReported(t *testing.T) {
seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
cnpgBackup("pgws", "new", "orders", "completed", time.Now().Add(-time.Hour)),
@@ -438,19 +447,19 @@ func TestCNPGWorkspace_AuditNeedsScheduledBackupEvidence(t *testing.T) {
sched bool
}{{"sees-schedules", true}, {"no-schedules", false}} {
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pg"}}
- allow(perms, cnpgGroup, "clusters", "", true)
- allow(perms, cnpgGroup, "scheduledbackups", "", u.sched)
- allow(perms, cnpgGroup, "scheduledbackups", "pg", u.sched)
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
+ allow(perms, cnpgsvc.Group, "scheduledbackups", "", u.sched)
+ allow(perms, cnpgsvc.Group, "scheduledbackups", "pg", u.sched)
env.srv.permCache.Set(u.name, nil, perms)
}
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "sees-schedules", ""))
- if len(got.Audit) != 1 || got.Audit[0].Name != "unscheduled" || got.Audit[0].CheckID != cnpgNoDeclarativeBackupCheckID {
- t.Errorf("audit = %+v, want one %s finding on unscheduled", got.Audit, cnpgNoDeclarativeBackupCheckID)
+ if len(got.Audit) != 1 || got.Audit[0].Name != "unscheduled" || got.Audit[0].CheckID != "cnpgNoDeclarativeBackup" {
+ t.Errorf("audit = %+v, want one %s finding on unscheduled", got.Audit, "cnpgNoDeclarativeBackup")
}
got = decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "no-schedules", ""))
- if got.Coverage["scheduledBackups"].State != cnpgCoverageDenied {
+ if got.Coverage["scheduledBackups"].State != integration.KindCoverageDenied {
t.Errorf("scheduledBackups coverage = %+v, want denied", got.Coverage["scheduledBackups"])
}
if len(got.Audit) != 0 {
@@ -463,32 +472,6 @@ func withUID(u *unstructured.Unstructured, uid string) *unstructured.Unstructure
return u
}
-func TestIsCNPGInstancePod(t *testing.T) {
- uids := map[string]types.UID{"pg/x": "x-uid"}
- ref := func(uid string, controller bool) metav1.OwnerReference {
- return metav1.OwnerReference{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "x", UID: types.UID(uid), Controller: boolPtr(controller)}
- }
- for _, c := range []struct {
- name string
- pod *corev1.Pod
- uids map[string]types.UID
- want bool
- }{
- {"controller ref to the visible Cluster", cnpgPod("pg", "x-1", "x", ref("x-uid", true)), uids, true},
- {"left behind by a deleted Cluster of the same name", cnpgPod("pg", "x-1", "x", ref("old-uid", true)), uids, false},
- {"non-controller owner", cnpgPod("pg", "x-1", "x", ref("x-uid", false)), uids, false},
- {"Cluster not visible", cnpgPod("pg", "x-1", "x", ref("x-uid", true)), map[string]types.UID{}, false},
- {"Cluster of that name in another namespace", cnpgPod("other", "x-1", "x", ref("x-uid", true)), uids, false},
- {"no cluster label", func() *corev1.Pod { p := cnpgPod("pg", "x-1", "x", ref("x-uid", true)); p.Labels = nil; return p }(), uids, false},
- } {
- t.Run(c.name, func(t *testing.T) {
- if got := isCNPGInstancePod(c.pod, c.uids); got != c.want {
- t.Errorf("isCNPGInstancePod = %v, want %v", got, c.want)
- }
- })
- }
-}
-
// The grouped issue view folds instance-Pod evidence into the owning Cluster's
// row. A caller who may list Clusters but not Pods must not receive that
// evidence in any form; one who may list Pods receives it on the Pod itself.
@@ -511,14 +494,14 @@ func TestCNPGWorkspace_PodEvidenceFollowsPodAccess(t *testing.T) {
pods bool
}{{"with-pods", true}, {"clusters-only", false}} {
perms := &auth.UserPermissions{AllowedNamespaces: []string{"pgev"}}
- allow(perms, cnpgGroup, "clusters", "", true)
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
allow(perms, "", "pods", "", u.pods)
allow(perms, "", "pods", "pgev", u.pods)
env.srv.permCache.Set(u.name, nil, perms)
}
control := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "with-pods", ""))
- var podIssue *CNPGWorkspaceIssue
+ var podIssue *cnpgsvc.CNPGWorkspaceIssue
for i, iss := range control.Issues {
if iss.Kind == "Pod" && iss.Name == "pg-orders-1" {
podIssue = &control.Issues[i]
@@ -532,7 +515,7 @@ func TestCNPGWorkspace_PodEvidenceFollowsPodAccess(t *testing.T) {
}
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "clusters-only", ""))
- if got.Coverage["pods"].State != cnpgCoverageDenied || len(got.Objects["pods"]) != 0 {
+ if got.Coverage["pods"].State != integration.KindCoverageDenied || len(got.Objects["pods"]) != 0 {
t.Errorf("pods coverage=%+v objects=%v, want denied and []", got.Coverage["pods"], objectNames(got.Objects["pods"]))
}
for _, iss := range got.Issues {
@@ -551,16 +534,16 @@ func TestCNPGWorkspace_DeniedNamespacesNeverComeFromTheServerInventory(t *testin
)
env := newAuthTestServer(t)
perms := &auth.UserPermissions{}
- allow(perms, cnpgGroup, "clusters", "", false)
+ allow(perms, cnpgsvc.Group, "clusters", "", false)
for _, ns := range allNamespaceNames() {
- allow(perms, cnpgGroup, "clusters", ns, ns == "default")
+ allow(perms, cnpgsvc.Group, "clusters", ns, ns == "default")
}
- allow(perms, cnpgGroup, "clusters", "broken", false)
+ allow(perms, cnpgsvc.Group, "clusters", "broken", false)
env.srv.permCache.Set("wide", nil, perms)
got := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "wide", ""))
cov := got.Coverage["clusters"]
- if cov.State != cnpgCoveragePartial {
+ if cov.State != integration.KindCoveragePartial {
t.Errorf("unfiltered: clusters coverage = %+v, want partial", cov)
}
if len(cov.DeniedNamespaces) != 0 {
@@ -575,10 +558,55 @@ func TestCNPGWorkspace_DeniedNamespacesNeverComeFromTheServerInventory(t *testin
got = decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace?namespaces=default,broken", "wide", ""))
cov = got.Coverage["clusters"]
- if cov.State != cnpgCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "broken" {
+ if cov.State != integration.KindCoveragePartial || len(cov.DeniedNamespaces) != 1 || cov.DeniedNamespaces[0] != "broken" {
t.Errorf("filtered: clusters coverage = %+v, want partial naming broken", cov)
}
if len(cov.AllowedNamespaces) != 1 || cov.AllowedNamespaces[0] != "default" {
t.Errorf("filtered: allowedNamespaces = %v, want [default]", cov.AllowedNamespaces)
}
}
+
+func TestCNPGWorkspace_ScheduleReadingsFollowScheduledBackupAccess(t *testing.T) {
+ sched := func(ns, name string) *unstructured.Unstructured {
+ return cnpgObj("postgresql.cnpg.io/v1", "ScheduledBackup", ns, name, map[string]any{"cluster": map[string]any{"name": "pg"}, "schedule": "0 0 2 * * *"}, nil)
+ }
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds,
+ cnpgObj("postgresql.cnpg.io/v1", "Cluster", "a", "pg", nil, nil),
+ cnpgObj("postgresql.cnpg.io/v1", "Cluster", "b", "pg", nil, nil),
+ sched("a", "nightly-a"),
+ sched("b", "nightly-b"),
+ )
+ env := newAuthTestServer(t)
+ for _, u := range []struct {
+ name string
+ a, b bool
+ }{{"both", true, true}, {"only-a", true, false}, {"none", false, false}} {
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"a", "b"}}
+ allow(perms, cnpgsvc.Group, "clusters", "", true)
+ allow(perms, cnpgsvc.Group, "scheduledbackups", "", u.a && u.b)
+ allow(perms, cnpgsvc.Group, "scheduledbackups", "a", u.a)
+ allow(perms, cnpgsvc.Group, "scheduledbackups", "b", u.b)
+ env.srv.permCache.Set(u.name, nil, perms)
+ }
+
+ both := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "both", ""))
+ if both.ScheduleReadings["a/nightly-a"] == "" || both.ScheduleReadings["b/nightly-b"] == "" {
+ t.Fatalf("control: readings = %v, want both schedules worded", both.ScheduleReadings)
+ }
+
+ partial := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "only-a", ""))
+ if partial.Coverage["scheduledBackups"].State != integration.KindCoveragePartial {
+ t.Errorf("coverage = %+v, want partial", partial.Coverage["scheduledBackups"])
+ }
+ if partial.ScheduleReadings["a/nightly-a"] == "" {
+ t.Errorf("readable schedule not worded: %v", partial.ScheduleReadings)
+ }
+ if _, ok := partial.ScheduleReadings["b/nightly-b"]; ok {
+ t.Errorf("a schedule in a namespace the caller cannot list was worded: %v", partial.ScheduleReadings)
+ }
+
+ denied := decodeWorkspace(t, env.authGet(t, "/api/cnpg/workspace", "none", ""))
+ if len(denied.ScheduleReadings) != 0 {
+ t.Errorf("readings = %v, want none when ScheduledBackups are denied", denied.ScheduleReadings)
+ }
+}
diff --git a/internal/server/controller_health.go b/internal/server/controller_health.go
index d76a77b234..8e8c0a1299 100644
--- a/internal/server/controller_health.go
+++ b/internal/server/controller_health.go
@@ -1,6 +1,10 @@
package server
-import corev1 "k8s.io/api/core/v1"
+import (
+ corev1 "k8s.io/api/core/v1"
+
+ "github.com/skyhook-io/radar/internal/podlogs"
+)
type controllerPodHealth struct {
Ready int
@@ -22,7 +26,7 @@ func summarizeControllerPods(pods []*corev1.Pod) controllerPodHealth {
break
}
}
- if isPodReady(p) {
+ if podlogs.IsPodReady(p) {
out.Ready++
}
if p.Status.Phase == corev1.PodPending {
diff --git a/internal/server/dashboard.go b/internal/server/dashboard.go
index d1e6eabe26..d82b29a35b 100644
--- a/internal/server/dashboard.go
+++ b/internal/server/dashboard.go
@@ -7,11 +7,10 @@ import (
"net/http"
"slices"
"sort"
+ "strings"
"sync"
"time"
- "strings"
-
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
@@ -19,6 +18,7 @@ import (
"github.com/skyhook-io/radar/internal/auth"
"github.com/skyhook-io/radar/internal/helm"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/timeline"
"github.com/skyhook-io/radar/internal/traffic"
@@ -282,7 +282,7 @@ func (s *Server) handleDashboard(w http.ResponseWriter, r *http.Request) {
}
dashStart := time.Now()
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, DashboardResponse{AccessRestricted: true})
return
}
@@ -1341,7 +1341,7 @@ func (s *Server) getDashboardMetrics(ctx context.Context, allowedNamespaces []st
// Capacity + scheduled-pod requests via the shared informer-derived
// computation (node_metrics.go). Namespace scoping keeps restricted users
// from seeing aggregate totals of namespaces they can't read.
- cr := computeCapacityRequests(nodes, listPodsScoped(cache.Pods(), allowedNamespaces))
+ cr := computeCapacityRequests(nodes, integration.ListPodsScoped(cache.Pods(), allowedNamespaces))
cpuCapacityMillis := cr.cpuCapMillis
memCapacityBytes := cr.memCapBytes
cpuRequestsMillis := cr.cpuReqMillis
diff --git a/internal/server/fanout_test.go b/internal/server/fanout_test.go
new file mode 100644
index 0000000000..709d585e1a
--- /dev/null
+++ b/internal/server/fanout_test.go
@@ -0,0 +1,43 @@
+package server
+
+import (
+ "context"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func TestFanOutKeepsIndexOrderAndBoundsConcurrency(t *testing.T) {
+ var running, peak atomic.Int32
+ got := integration.FanOut(context.Background(), 20, 3, func(i int) int {
+ n := running.Add(1)
+ for {
+ p := peak.Load()
+ if n <= p || peak.CompareAndSwap(p, n) {
+ break
+ }
+ }
+ time.Sleep(2 * time.Millisecond)
+ running.Add(-1)
+ return i * 10
+ })
+ if len(got) != 20 {
+ t.Fatalf("len = %d", len(got))
+ }
+ for i, v := range got {
+ if v != i*10 {
+ t.Fatalf("results[%d] = %d, want %d", i, v, i*10)
+ }
+ }
+ if p := peak.Load(); p > 3 || p < 1 {
+ t.Errorf("peak concurrency = %d, want 1..3", p)
+ }
+}
+
+func TestFanOutOfNothing(t *testing.T) {
+ if got := integration.FanOut(context.Background(), 0, 4, func(int) string { t.Fatal("ran"); return "" }); len(got) != 0 {
+ t.Errorf("got %v", got)
+ }
+}
diff --git a/internal/server/gitops_handlers.go b/internal/server/gitops_handlers.go
index 402d04d7b0..e313747575 100644
--- a/internal/server/gitops_handlers.go
+++ b/internal/server/gitops_handlers.go
@@ -20,6 +20,7 @@ import (
"github.com/skyhook-io/radar/internal/argocd"
"github.com/skyhook-io/radar/internal/auth"
"github.com/skyhook-io/radar/internal/connections"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/argoapi"
@@ -41,7 +42,7 @@ type gitopsRequest struct {
// namespace's resources. False means handlers should short-circuit with an
// empty success response (see the per-handler empty value).
func (g *gitopsRequest) HasNamespaceAccess() bool {
- return !noNamespaceAccess(g.AllowedNamespaces)
+ return !integration.NoNamespaceAccess(g.AllowedNamespaces)
}
// parseGitOpsRequest pulls the GitOps URL params and runs the namespace
@@ -57,7 +58,7 @@ func (s *Server) parseGitOpsRequest(w http.ResponseWriter, r *http.Request) (*gi
}
if namespace != "" {
allowed := s.getUserNamespaces(r, []string{namespace})
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, fmt.Sprintf("no access to namespace %q", namespace))
return nil, false
}
@@ -172,6 +173,7 @@ func (s *Server) resolveGitOpsTree(r *http.Request, req *gitopsRequest) (*gitops
return s.canAccessGitOpsRef(r, req, group, kind, namespace, name, false)
}
resolver := newInsightsResolver(r.Context(), req.Cache, req.AllowedNamespaces, canAccess)
+ resolver.canReadEvidence = s.issueEvidenceAccess(r)
memoKey := gitopsIssuesMemoKey(auth.UserFromContext(r.Context()), req.AllowedNamespaces)
resolver.composed = func() ([]issues.Issue, []issues.Issue) {
return s.gitopsIssuesMemo.load(memoKey, resolver.composeIssues)
@@ -549,7 +551,7 @@ func (s *Server) handleGitOpsManagedResources(w http.ResponseWriter, r *http.Req
nsFilter := strings.TrimSpace(r.URL.Query().Get("namespace"))
allowedNamespaces := s.getUserNamespaces(r, nil)
- if noNamespaceAccess(allowedNamespaces) {
+ if integration.NoNamespaceAccess(allowedNamespaces) {
// Caller has no namespace access — return a tree with just the
// synthetic root + a warning. Mirrors handleGitOpsTree's behavior
// rather than 403'ing so the frontend can render an honest empty state.
@@ -812,7 +814,7 @@ func (s *Server) canAccessGitOpsRef(r *http.Request, req *gitopsRequest, group,
return false
}
allowed := s.getUserNamespaces(r, []string{name})
- return !noNamespaceAccess(allowed)
+ return !integration.NoNamespaceAccess(allowed)
}
if namespace != "" {
return namespaceAllowedForGitOps(req.AllowedNamespaces, namespace)
@@ -844,6 +846,7 @@ type insightsResolver struct {
cache *k8s.ResourceCache
allowedNamespaces []string
canAccess func(group, kind, namespace, name string) bool
+ canReadEvidence func(issues.EvidenceRead) bool
// The cluster-wide issue set is composed at most once per insights request
// (lazily, only if a degraded managed resource asks for it) and reused
@@ -1038,6 +1041,7 @@ func (r *insightsResolver) ResourceProblems(group, kind, namespace, name string)
r.composedFlat, r.composedGrouped = r.composeIssues()
})
related := issues.RelatedIssuesFrom(r.composedFlat, r.composedGrouped, issues.RelatedIssueOptions{
+ CanReadEvidence: r.canReadEvidence,
CanReadRelated: func(ref issues.Ref) bool {
return r.canAccess != nil && r.canAccess(ref.Group, ref.Kind, ref.Namespace, ref.Name)
},
@@ -1146,6 +1150,7 @@ func (r *insightsResolver) composeIssues() ([]issues.Issue, []issues.Issue) {
SkipPodTemplateContext: true,
Namespaces: r.allowedNamespaces,
Limit: issues.NoLimit,
+ CanReadEvidence: r.canReadEvidence,
CanReadRelated: func(ref issues.Ref) bool {
return r.canAccess != nil && r.canAccess(ref.Group, ref.Kind, ref.Namespace, ref.Name)
},
diff --git a/internal/server/gitops_health_overlay_test.go b/internal/server/gitops_health_overlay_test.go
index e26b936e74..d9d9d0e01e 100644
--- a/internal/server/gitops_health_overlay_test.go
+++ b/internal/server/gitops_health_overlay_test.go
@@ -4,16 +4,17 @@ import (
"context"
"errors"
"fmt"
- "github.com/skyhook-io/radar/internal/argocd"
"testing"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/types"
+
+ "github.com/skyhook-io/radar/internal/argocd"
"github.com/skyhook-io/radar/internal/config"
"github.com/skyhook-io/radar/internal/connections"
"github.com/skyhook-io/radar/pkg/argoapi"
gitopsinsights "github.com/skyhook-io/radar/pkg/gitops/insights"
gitopstree "github.com/skyhook-io/radar/pkg/gitops/tree"
- "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "k8s.io/apimachinery/pkg/types"
)
func overlayApp(health string) *unstructured.Unstructured {
diff --git a/internal/server/gitops_write_evidence.go b/internal/server/gitops_write_evidence.go
new file mode 100644
index 0000000000..b441a7d689
--- /dev/null
+++ b/internal/server/gitops_write_evidence.go
@@ -0,0 +1,800 @@
+package server
+
+import (
+ "encoding/json"
+ "errors"
+ "fmt"
+ "log"
+ "net/http"
+ "strconv"
+ "strings"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/gitops"
+ "github.com/skyhook-io/radar/pkg/topology"
+)
+
+// The write-evidence endpoint answers, before a direct write, what Radar can
+// learn about whether the object's GitOps (or Helm) owner will put the old
+// value back: per written field, is it in the last client-side apply payload,
+// is it owned (managedFields) by the owner's controller, does an ignore rule
+// cover it — plus the owner's sync policy. Both managedFields and last-applied
+// are stripped from Radar's caches, so this reads the object directly, as the
+// caller. Only derived facts leave the server, never the annotation or the
+// managedFields themselves.
+
+const (
+ maxWriteEvidenceRequestBytes = 16 << 10
+ maxWriteEvidencePaths = 64
+ lastAppliedAnnotationKey = "kubectl.kubernetes.io/last-applied-configuration"
+)
+
+type writeEvidenceRef struct {
+ Kind string `json:"kind"`
+ Group string `json:"group"`
+ Namespace string `json:"namespace"`
+ Name string `json:"name"`
+}
+
+type gitOpsWriteEvidenceRequest struct {
+ writeEvidenceRef
+ Paths []string `json:"paths"`
+ // Owner is the GitOps owner the client resolved (it may be inherited from
+ // a parent workload). When absent, the object's own tracking metadata is
+ // used.
+ Owner *writeEvidenceRef `json:"owner,omitempty"`
+}
+
+type fieldOwnerEvidence struct {
+ Manager string `json:"manager"`
+ Operation string `json:"operation"`
+ Subresource string `json:"subresource,omitempty"`
+ // Tool is the GitOps/Helm tool the manager belongs to ("argocd",
+ // "fluxcd", "helm"), empty for any other manager.
+ Tool string `json:"tool,omitempty"`
+ // Approximate: the manager owns an ancestor of the path (an atomic list
+ // or struct), or only part of the path's subtree.
+ Approximate bool `json:"approximate,omitempty"`
+}
+
+type gitOpsPathEvidence struct {
+ Path string `json:"path"`
+ Error string `json:"error,omitempty"`
+ // LastApplied: "present" | "absent" | "no-annotation".
+ LastApplied string `json:"lastApplied"`
+ OwnedBy []fieldOwnerEvidence `json:"ownedBy"`
+ // OwnedByGitOps: a manager of the resolved owner's tool owns the path.
+ OwnedByGitOps bool `json:"ownedByGitOps"`
+ Approximate bool `json:"approximate,omitempty"`
+ // Ignored: "" | "effective" (the owner won't revert it) |
+ // "comparison-only" (Argo ignoreDifferences without
+ // RespectIgnoreDifferences: no self-heal trigger, but a sync overwrites) |
+ // "unevaluated" (a matching Argo rule uses jqPathExpressions, which Radar
+ // doesn't evaluate, so it may or may not cover the path).
+ Ignored string `json:"ignored,omitempty"`
+ IgnoredBy string `json:"ignoredBy,omitempty"`
+}
+
+type gitOpsWritePolicy struct {
+ Tool string `json:"tool"`
+ Auto *bool `json:"auto"`
+ SelfHeal *bool `json:"selfHeal"`
+ Prune *bool `json:"prune"`
+ Suspended *bool `json:"suspended"`
+ Interval string `json:"interval,omitempty"`
+ // Argo sync options (Application spec merged with the object's
+ // argocd.argoproj.io/sync-options annotation).
+ RespectIgnoreDifferences bool `json:"respectIgnoreDifferences,omitempty"`
+ ServerSideApply bool `json:"serverSideApply,omitempty"`
+ Replace bool `json:"replace,omitempty"`
+ // HelmRelease spec.driftDetection.mode: "disabled" | "warn" | "enabled".
+ DriftDetection string `json:"driftDetection,omitempty"`
+ // ObjectReconcile is a per-object opt-out read from the target's
+ // annotations: "ignore" (Flux ssa: Ignore, reconcile: disabled, Helm
+ // driftDetection: disabled) or "if-not-present" (Flux ssa: IfNotPresent —
+ // recreated when deleted, never overwritten).
+ ObjectReconcile string `json:"objectReconcile,omitempty"`
+}
+
+type controllerOwnerEvidence struct {
+ APIVersion string `json:"apiVersion"`
+ Kind string `json:"kind"`
+ Name string `json:"name"`
+}
+
+type gitOpsWriteEvidenceResponse struct {
+ UID string `json:"uid"`
+ ResourceVersion string `json:"resourceVersion"`
+ Owner *topology.ResourceRef `json:"owner"`
+ Policy *gitOpsWritePolicy `json:"policy"`
+ PolicyError string `json:"policyError,omitempty"`
+ ControllerOwner *controllerOwnerEvidence `json:"controllerOwner,omitempty"`
+ Paths []gitOpsPathEvidence `json:"paths"`
+}
+
+func (s *Server) handleGitOpsWriteEvidence(w http.ResponseWriter, r *http.Request) {
+ if !s.requireConnected(w) {
+ return
+ }
+ var req gitOpsWriteEvidenceRequest
+ if err := decodeBoundedJSONBody(w, r, maxWriteEvidenceRequestBytes, &req); err != nil {
+ var tooLarge *http.MaxBytesError
+ if errors.As(err, &tooLarge) {
+ s.writeError(w, http.StatusRequestEntityTooLarge, "write-evidence request is too large")
+ return
+ }
+ s.writeError(w, http.StatusBadRequest, "invalid write-evidence request: "+err.Error())
+ return
+ }
+ if req.Kind == "" || req.Name == "" {
+ s.writeError(w, http.StatusBadRequest, "kind and name are required")
+ return
+ }
+ if len(req.Paths) > maxWriteEvidencePaths {
+ s.writeError(w, http.StatusBadRequest, fmt.Sprintf("at most %d paths are allowed", maxWriteEvidencePaths))
+ return
+ }
+
+ discovery := k8s.GetResourceDiscovery()
+ if discovery == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "resource discovery not available")
+ return
+ }
+ gvr, ok := discovery.GetGVRWithGroup(req.Kind, req.Group)
+ if !ok {
+ s.writeError(w, http.StatusBadRequest, fmt.Sprintf("unknown kind %q in group %q", req.Kind, req.Group))
+ return
+ }
+ dyn := s.getDynamicClientForRequest(r)
+ if dyn == nil {
+ s.writeError(w, http.StatusServiceUnavailable, "kubernetes client not available")
+ return
+ }
+
+ target, err := dyn.Resource(gvr).Namespace(req.Namespace).Get(r.Context(), req.Name, metav1.GetOptions{})
+ if err != nil {
+ switch {
+ case apierrors.IsForbidden(err):
+ s.writeError(w, http.StatusForbidden, err.Error())
+ case apierrors.IsNotFound(err):
+ s.writeError(w, http.StatusNotFound, err.Error())
+ default:
+ log.Printf("[gitops] Failed to read %s %s/%s for write evidence: %v", sanitizeForLog(req.Kind), sanitizeForLog(req.Namespace), sanitizeForLog(req.Name), err)
+ s.writeError(w, http.StatusInternalServerError, err.Error())
+ }
+ return
+ }
+
+ owner := resolveWriteEvidenceOwner(req, target)
+ var ownerObj *unstructured.Unstructured
+ var ownerErr error
+ if owner != nil {
+ ownerGVR, gitOpsOwner := gitOpsOwnerGVR(owner)
+ switch {
+ case !gitOpsOwner:
+ // Native Helm: no controller policy to read.
+ case owner.Namespace == "":
+ ownerErr = fmt.Errorf("the owner's namespace is unknown")
+ default:
+ ownerObj, ownerErr = dyn.Resource(ownerGVR).Namespace(owner.Namespace).Get(r.Context(), owner.Name, metav1.GetOptions{})
+ }
+ }
+
+ s.writeJSON(w, buildGitOpsWriteEvidence(target, req.writeEvidenceRef, req.Paths, owner, ownerObj, ownerErr))
+}
+
+func resolveWriteEvidenceOwner(req gitOpsWriteEvidenceRequest, target *unstructured.Unstructured) *topology.ResourceRef {
+ if req.Owner != nil && req.Owner.Name != "" {
+ return &topology.ResourceRef{Kind: req.Owner.Kind, Group: req.Owner.Group, Namespace: req.Owner.Namespace, Name: req.Owner.Name}
+ }
+ refs := topology.SynthesizeManagedBy(target, req.Kind, req.Namespace, req.Name, nil, nil, nil)
+ if len(refs) == 0 {
+ return nil
+ }
+ ref := refs[0]
+ if _, gitOps := gitOpsOwnerGVR(&ref); !gitOps && !isNativeHelmOwner(&ref) {
+ return nil
+ }
+ return &ref
+}
+
+func isNativeHelmOwner(ref *topology.ResourceRef) bool {
+ return ref != nil && ref.Kind == "HelmRelease" && ref.Group == ""
+}
+
+func gitOpsOwnerGVR(ref *topology.ResourceRef) (schema.GroupVersionResource, bool) {
+ if ref == nil {
+ return schema.GroupVersionResource{}, false
+ }
+ if ref.Group == "argoproj.io" && strings.EqualFold(ref.Kind, "Application") {
+ return argoApplicationGVR, true
+ }
+ if ref.Group == "kustomize.toolkit.fluxcd.io" || ref.Group == "helm.toolkit.fluxcd.io" {
+ entry, err := gitops.ResolveFluxKind(ref.Kind)
+ if err == nil && entry.GVR.Group == ref.Group {
+ return entry.GVR, true
+ }
+ }
+ return schema.GroupVersionResource{}, false
+}
+
+func ownerTool(ref *topology.ResourceRef) string {
+ switch {
+ case ref == nil:
+ return ""
+ case ref.Group == "argoproj.io":
+ return "argocd"
+ case ref.Group == "kustomize.toolkit.fluxcd.io" || ref.Group == "helm.toolkit.fluxcd.io":
+ return "fluxcd"
+ case isNativeHelmOwner(ref):
+ return "helm"
+ }
+ return ""
+}
+
+// Field managers GitOps and Helm controllers write with. Argo CD records
+// argocd-controller for server-side apply and the controller binary name for
+// client-side apply.
+var gitOpsFieldManagers = map[string]string{
+ "argocd-controller": "argocd",
+ "argocd-application-controller": "argocd",
+ "kustomize-controller": "fluxcd",
+ "helm-controller": "fluxcd",
+ "helm": "helm",
+}
+
+// Which manager, for a given owner, is the one that re-applies the source.
+func ownerManagers(ref *topology.ResourceRef) map[string]bool {
+ switch ownerTool(ref) {
+ case "argocd":
+ return map[string]bool{"argocd-controller": true, "argocd-application-controller": true}
+ case "fluxcd":
+ if ref.Group == "helm.toolkit.fluxcd.io" {
+ return map[string]bool{"helm-controller": true}
+ }
+ return map[string]bool{"kustomize-controller": true}
+ case "helm":
+ return map[string]bool{"helm": true}
+ }
+ return nil
+}
+
+func buildGitOpsWriteEvidence(
+ target *unstructured.Unstructured,
+ ref writeEvidenceRef,
+ paths []string,
+ owner *topology.ResourceRef,
+ ownerObj *unstructured.Unstructured,
+ ownerErr error,
+) gitOpsWriteEvidenceResponse {
+ resp := gitOpsWriteEvidenceResponse{
+ UID: string(target.GetUID()),
+ ResourceVersion: target.GetResourceVersion(),
+ Owner: owner,
+ Paths: []gitOpsPathEvidence{},
+ }
+ for _, or := range target.GetOwnerReferences() {
+ if or.Controller != nil && *or.Controller {
+ resp.ControllerOwner = &controllerOwnerEvidence{APIVersion: or.APIVersion, Kind: or.Kind, Name: or.Name}
+ break
+ }
+ }
+
+ var policy *gitOpsWritePolicy
+ if owner != nil && ownerTool(owner) != "helm" {
+ switch {
+ case ownerErr != nil:
+ resp.PolicyError = describeOwnerReadError(ownerErr)
+ case ownerObj != nil:
+ policy = readGitOpsWritePolicy(owner, ownerObj, target)
+ }
+ }
+ resp.Policy = policy
+
+ var lastApplied any
+ hasLastApplied := false
+ if raw := target.GetAnnotations()[lastAppliedAnnotationKey]; strings.TrimSpace(raw) != "" {
+ if err := json.Unmarshal([]byte(raw), &lastApplied); err == nil {
+ hasLastApplied = true
+ }
+ }
+ managers := ownerManagers(owner)
+ managedFields := target.GetManagedFields()
+
+ for _, path := range paths {
+ ev := gitOpsPathEvidence{Path: path, LastApplied: "no-annotation", OwnedBy: []fieldOwnerEvidence{}}
+ segs, err := parseWriteFieldPath(path)
+ if err != nil {
+ ev.Error = err.Error()
+ resp.Paths = append(resp.Paths, ev)
+ continue
+ }
+ if hasLastApplied {
+ if valuePresent(lastApplied, segs) {
+ ev.LastApplied = "present"
+ } else {
+ ev.LastApplied = "absent"
+ }
+ }
+ for _, mf := range managedFields {
+ if mf.FieldsV1 == nil || len(mf.FieldsV1.Raw) == 0 {
+ continue
+ }
+ var fields map[string]any
+ if err := json.Unmarshal(mf.FieldsV1.Raw, &fields); err != nil {
+ continue
+ }
+ owned, approximate := fieldsV1Owns(fields, segs)
+ if !owned {
+ continue
+ }
+ ev.OwnedBy = append(ev.OwnedBy, fieldOwnerEvidence{
+ Manager: mf.Manager,
+ Operation: string(mf.Operation),
+ Subresource: mf.Subresource,
+ Tool: gitOpsFieldManagers[mf.Manager],
+ Approximate: approximate,
+ })
+ if managers[mf.Manager] && mf.Subresource == "" {
+ ev.OwnedByGitOps = true
+ ev.Approximate = ev.Approximate || approximate
+ }
+ }
+ ev.Ignored, ev.IgnoredBy = ignoreRuleFor(owner, ownerObj, policy, target, ref, segs, ev.OwnedBy)
+ resp.Paths = append(resp.Paths, ev)
+ }
+ return resp
+}
+
+func describeOwnerReadError(err error) string {
+ switch {
+ case apierrors.IsForbidden(err):
+ return "you can't read the owner, so its sync policy is unknown"
+ case apierrors.IsNotFound(err):
+ return "the owner no longer exists"
+ default:
+ return err.Error()
+ }
+}
+
+func evidenceBool(b bool) *bool { return &b }
+
+func readGitOpsWritePolicy(owner *topology.ResourceRef, ownerObj, target *unstructured.Unstructured) *gitOpsWritePolicy {
+ annotations := target.GetAnnotations()
+ switch ownerTool(owner) {
+ case "argocd":
+ p := &gitOpsWritePolicy{Tool: "argocd"}
+ automated, hasAutomated, _ := unstructured.NestedMap(ownerObj.Object, "spec", "syncPolicy", "automated")
+ auto := hasAutomated && automated != nil
+ if enabled, ok := automated["enabled"].(bool); ok && !enabled {
+ auto = false
+ }
+ selfHeal, _ := automated["selfHeal"].(bool)
+ prune, _ := automated["prune"].(bool)
+ p.Auto = evidenceBool(auto)
+ p.SelfHeal = evidenceBool(auto && selfHeal)
+ p.Prune = evidenceBool(auto && prune)
+ options, _, _ := unstructured.NestedStringSlice(ownerObj.Object, "spec", "syncPolicy", "syncOptions")
+ for _, opt := range strings.Split(annotations["argocd.argoproj.io/sync-options"], ",") {
+ if opt = strings.TrimSpace(opt); opt != "" {
+ options = append(options, opt)
+ }
+ }
+ for _, opt := range options {
+ switch strings.TrimSpace(opt) {
+ case "RespectIgnoreDifferences=true":
+ p.RespectIgnoreDifferences = true
+ case "ServerSideApply=true":
+ p.ServerSideApply = true
+ case "Replace=true":
+ p.Replace = true
+ }
+ }
+ return p
+ case "fluxcd":
+ suspended, _, _ := unstructured.NestedBool(ownerObj.Object, "spec", "suspend")
+ interval, _, _ := unstructured.NestedString(ownerObj.Object, "spec", "interval")
+ p := &gitOpsWritePolicy{Tool: "fluxcd", Auto: evidenceBool(!suspended), Suspended: evidenceBool(suspended), Interval: interval}
+ if owner.Group == "helm.toolkit.fluxcd.io" {
+ mode, _, _ := unstructured.NestedString(ownerObj.Object, "spec", "driftDetection", "mode")
+ if mode == "" {
+ mode = "disabled"
+ }
+ p.DriftDetection = mode
+ p.SelfHeal = evidenceBool(!suspended && mode == "enabled")
+ if strings.EqualFold(annotations["helm.toolkit.fluxcd.io/driftDetection"], "disabled") {
+ p.ObjectReconcile = "ignore"
+ }
+ return p
+ }
+ prune, _, _ := unstructured.NestedBool(ownerObj.Object, "spec", "prune")
+ p.SelfHeal = evidenceBool(!suspended)
+ p.Prune = evidenceBool(prune)
+ switch {
+ case strings.EqualFold(annotations["kustomize.toolkit.fluxcd.io/reconcile"], "disabled"),
+ strings.EqualFold(annotations["kustomize.toolkit.fluxcd.io/ssa"], "Ignore"):
+ p.ObjectReconcile = "ignore"
+ case strings.EqualFold(annotations["kustomize.toolkit.fluxcd.io/ssa"], "IfNotPresent"):
+ p.ObjectReconcile = "if-not-present"
+ }
+ return p
+ }
+ return nil
+}
+
+// ignoreRuleFor reports whether an owner-level rule covers the path.
+func ignoreRuleFor(
+ owner *topology.ResourceRef,
+ ownerObj *unstructured.Unstructured,
+ policy *gitOpsWritePolicy,
+ target *unstructured.Unstructured,
+ ref writeEvidenceRef,
+ segs []fieldPathSegment,
+ ownedBy []fieldOwnerEvidence,
+) (string, string) {
+ if policy == nil || ownerObj == nil {
+ return "", ""
+ }
+ switch policy.ObjectReconcile {
+ case "ignore", "if-not-present":
+ return "effective", "the object's reconcile annotation"
+ }
+ tokens := pointerTokensFor(target.Object, segs)
+ switch ownerTool(owner) {
+ case "argocd":
+ entries, _, _ := unstructured.NestedSlice(ownerObj.Object, "spec", "ignoreDifferences")
+ jqRule := false
+ for _, raw := range entries {
+ entry, ok := raw.(map[string]any)
+ if !ok || !argoIgnoreEntryMatches(entry, ref) {
+ continue
+ }
+ covered := false
+ for _, p := range stringSlice(entry["jsonPointers"]) {
+ if pointerCovers(p, tokens) {
+ covered = true
+ }
+ }
+ for _, m := range stringSlice(entry["managedFieldsManagers"]) {
+ for _, o := range ownedBy {
+ if o.Manager == m {
+ covered = true
+ }
+ }
+ }
+ if covered {
+ if policy.RespectIgnoreDifferences {
+ return "effective", "spec.ignoreDifferences"
+ }
+ return "comparison-only", "spec.ignoreDifferences"
+ }
+ if len(stringSlice(entry["jqPathExpressions"])) > 0 {
+ jqRule = true
+ }
+ }
+ if jqRule {
+ return "unevaluated", "spec.ignoreDifferences jqPathExpressions"
+ }
+ case "fluxcd":
+ if owner.Group != "helm.toolkit.fluxcd.io" {
+ return "", ""
+ }
+ rules, _, _ := unstructured.NestedSlice(ownerObj.Object, "spec", "driftDetection", "ignore")
+ for _, raw := range rules {
+ rule, ok := raw.(map[string]any)
+ if !ok {
+ continue
+ }
+ if t, ok := rule["target"].(map[string]any); ok && !fluxIgnoreTargetMatches(t, ref) {
+ continue
+ }
+ for _, p := range stringSlice(rule["paths"]) {
+ if pointerCovers(p, tokens) {
+ return "effective", "spec.driftDetection.ignore"
+ }
+ }
+ }
+ }
+ return "", ""
+}
+
+// argoIgnoreEntryMatches: kind is required and an omitted group is the core
+// group.
+func argoIgnoreEntryMatches(entry map[string]any, ref writeEvidenceRef) bool {
+ str := func(k string) string { v, _ := entry[k].(string); return v }
+ if str("kind") != ref.Kind || str("group") != ref.Group {
+ return false
+ }
+ if name := str("name"); name != "" && name != ref.Name {
+ return false
+ }
+ if ns := str("namespace"); ns != "" && ns != ref.Namespace {
+ return false
+ }
+ return true
+}
+
+// fluxIgnoreTargetMatches: a Kustomize-style selector where omitted fields
+// match anything. Label/annotation selectors aren't evaluated, so a rule
+// carrying one never counts as covering the object.
+func fluxIgnoreTargetMatches(target map[string]any, ref writeEvidenceRef) bool {
+ str := func(k string) string { v, _ := target[k].(string); return v }
+ if str("labelSelector") != "" || str("annotationSelector") != "" {
+ return false
+ }
+ for key, want := range map[string]string{"kind": ref.Kind, "group": ref.Group, "name": ref.Name, "namespace": ref.Namespace} {
+ if v := str(key); v != "" && v != want {
+ return false
+ }
+ }
+ return true
+}
+
+func stringSlice(v any) []string {
+ items, _ := v.([]any)
+ out := make([]string, 0, len(items))
+ for _, item := range items {
+ if s, ok := item.(string); ok {
+ out = append(out, s)
+ }
+ }
+ return out
+}
+
+// ---------------------------------------------------------------------------
+// Field paths: `a.b["x/y"][name=app][0][*].c`
+// ---------------------------------------------------------------------------
+
+type fieldPathSegmentType int
+
+const (
+ segKey fieldPathSegmentType = iota
+ segIndex
+ segMatch
+ segAny
+)
+
+type fieldPathSegment struct {
+ typ fieldPathSegmentType
+ key string
+ value string
+ index int
+}
+
+func unquoteFieldPath(s string) string {
+ s = strings.TrimSpace(s)
+ if len(s) >= 2 && (s[0] == '"' || s[0] == '\'') && s[len(s)-1] == s[0] {
+ return s[1 : len(s)-1]
+ }
+ return s
+}
+
+func parseWriteFieldPath(path string) ([]fieldPathSegment, error) {
+ var segs []fieldPathSegment
+ var key strings.Builder
+ flush := func() {
+ if key.Len() > 0 {
+ segs = append(segs, fieldPathSegment{typ: segKey, key: key.String()})
+ key.Reset()
+ }
+ }
+ for i := 0; i < len(path); {
+ c := path[i]
+ switch c {
+ case '.':
+ flush()
+ i++
+ case '[':
+ flush()
+ j := i + 1
+ var quote byte
+ for ; j < len(path); j++ {
+ d := path[j]
+ if quote != 0 {
+ if d == quote {
+ quote = 0
+ }
+ } else if d == '"' || d == '\'' {
+ quote = d
+ } else if d == ']' {
+ break
+ }
+ }
+ if j >= len(path) {
+ return nil, fmt.Errorf("unterminated bracket in %q", path)
+ }
+ body := strings.TrimSpace(path[i+1 : j])
+ switch {
+ case body == "*":
+ segs = append(segs, fieldPathSegment{typ: segAny})
+ case body != "" && body[0] != '"' && body[0] != '\'' && strings.Contains(body, "="):
+ eq := strings.Index(body, "=")
+ segs = append(segs, fieldPathSegment{typ: segMatch, key: strings.TrimSpace(body[:eq]), value: unquoteFieldPath(body[eq+1:])})
+ default:
+ if n, err := strconv.Atoi(body); err == nil && n >= 0 {
+ segs = append(segs, fieldPathSegment{typ: segIndex, index: n})
+ } else {
+ segs = append(segs, fieldPathSegment{typ: segKey, key: unquoteFieldPath(body)})
+ }
+ }
+ i = j + 1
+ default:
+ key.WriteByte(c)
+ i++
+ }
+ }
+ flush()
+ if len(segs) == 0 {
+ return nil, fmt.Errorf("empty field path")
+ }
+ return segs, nil
+}
+
+func matchesSelector(item any, seg fieldPathSegment) bool {
+ m, ok := item.(map[string]any)
+ if !ok {
+ return false
+ }
+ v, ok := m[seg.key]
+ return ok && fmt.Sprint(v) == seg.value
+}
+
+// valuePresent reports whether the path exists in a decoded JSON document.
+// A wildcard matches when any element has the rest of the path.
+func valuePresent(node any, segs []fieldPathSegment) bool {
+ if len(segs) == 0 {
+ return true
+ }
+ seg := segs[0]
+ switch seg.typ {
+ case segKey:
+ m, ok := node.(map[string]any)
+ if !ok {
+ return false
+ }
+ child, ok := m[seg.key]
+ return ok && valuePresent(child, segs[1:])
+ case segIndex:
+ list, ok := node.([]any)
+ return ok && seg.index < len(list) && valuePresent(list[seg.index], segs[1:])
+ default:
+ list, ok := node.([]any)
+ if !ok {
+ return false
+ }
+ for _, item := range list {
+ if (seg.typ == segAny || matchesSelector(item, seg)) && valuePresent(item, segs[1:]) {
+ return true
+ }
+ }
+ return false
+ }
+}
+
+// fieldsV1Owns reports whether a managedFields fieldsV1 set covers the path.
+// Exact when the set names the path itself as a leaf; approximate when it
+// only names an ancestor as a leaf (an atomic value such as an atomic list),
+// names the path with a partially-owned subtree, or matched via a wildcard.
+func fieldsV1Owns(node map[string]any, segs []fieldPathSegment) (owned, approximate bool) {
+ return fieldsV1OwnsAt(node, segs, 0)
+}
+
+func fieldsV1OwnsAt(node map[string]any, segs []fieldPathSegment, depth int) (bool, bool) {
+ if len(segs) == 0 {
+ return true, !isFieldsLeaf(node)
+ }
+ seg := segs[0]
+ var children []map[string]any
+ switch seg.typ {
+ case segKey:
+ if child, ok := node["f:"+seg.key].(map[string]any); ok {
+ children = append(children, child)
+ }
+ case segMatch, segAny:
+ for k, v := range node {
+ child, ok := v.(map[string]any)
+ if !ok || !strings.HasPrefix(k, "k:") {
+ continue
+ }
+ var keyFields map[string]any
+ if seg.typ == segAny ||
+ (json.Unmarshal([]byte(strings.TrimPrefix(k, "k:")), &keyFields) == nil && matchesSelector(keyFields, seg)) {
+ children = append(children, child)
+ }
+ }
+ case segIndex:
+ // Index-addressed lists are atomic in managedFields; only an
+ // ancestor leaf can own them.
+ }
+ if len(children) == 0 {
+ if depth > 0 && isFieldsLeaf(node) {
+ return true, true
+ }
+ return false, false
+ }
+ owned, approximate := false, seg.typ == segAny
+ for _, child := range children {
+ if o, a := fieldsV1OwnsAt(child, segs[1:], depth+1); o {
+ owned = true
+ approximate = approximate || a
+ }
+ }
+ if !owned {
+ return false, false
+ }
+ return true, approximate
+}
+
+func isFieldsLeaf(node map[string]any) bool {
+ for k := range node {
+ if k != "." {
+ return false
+ }
+ }
+ return true
+}
+
+// pointerTokensFor turns the path into JSON-pointer tokens against the live
+// object, resolving selectors to indices. It stops at a selector it can't
+// resolve (or a wildcard).
+func pointerTokensFor(obj any, segs []fieldPathSegment) []string {
+ tokens := []string{}
+ node := obj
+ for _, seg := range segs {
+ switch seg.typ {
+ case segKey:
+ tokens = append(tokens, seg.key)
+ if m, ok := node.(map[string]any); ok {
+ node = m[seg.key]
+ } else {
+ node = nil
+ }
+ case segIndex:
+ tokens = append(tokens, strconv.Itoa(seg.index))
+ if list, ok := node.([]any); ok && seg.index < len(list) {
+ node = list[seg.index]
+ } else {
+ node = nil
+ }
+ case segMatch:
+ list, _ := node.([]any)
+ idx := -1
+ for i, item := range list {
+ if matchesSelector(item, seg) {
+ idx = i
+ break
+ }
+ }
+ if idx < 0 {
+ return tokens
+ }
+ tokens = append(tokens, strconv.Itoa(idx))
+ node = list[idx]
+ default:
+ return tokens
+ }
+ }
+ return tokens
+}
+
+func pointerCovers(pointer string, tokens []string) bool {
+ if !strings.HasPrefix(pointer, "/") {
+ return false
+ }
+ parts := strings.Split(pointer[1:], "/")
+ if len(parts) > len(tokens) {
+ return false
+ }
+ for i, p := range parts {
+ p = strings.ReplaceAll(strings.ReplaceAll(p, "~1", "/"), "~0", "~")
+ if p != tokens[i] {
+ return false
+ }
+ }
+ return true
+}
diff --git a/internal/server/gitops_write_evidence_test.go b/internal/server/gitops_write_evidence_test.go
new file mode 100644
index 0000000000..f7114d3883
--- /dev/null
+++ b/internal/server/gitops_write_evidence_test.go
@@ -0,0 +1,370 @@
+package server
+
+import (
+ "encoding/json"
+ "errors"
+ "net/http"
+ "net/http/httptest"
+ "strings"
+ "testing"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/topology"
+)
+
+func TestGitOpsWriteEvidenceDiscoveryUnavailable(t *testing.T) {
+ seedCNPGWorkspace(t, cnpgWorkspaceTestKinds)
+ k8s.ResetResourceDiscovery()
+ r := httptest.NewRequest(http.MethodPost, "/api/gitops/write-evidence", strings.NewReader(`{"kind":"Deployment","group":"apps","namespace":"prod","name":"api","paths":["spec.replicas"]}`))
+ w := httptest.NewRecorder()
+ (&Server{}).handleGitOpsWriteEvidence(w, r)
+ if w.Code != http.StatusServiceUnavailable || w.Body.String() != "{\"error\":\"resource discovery not available\"}\n" {
+ t.Fatalf("response = %d %s", w.Code, w.Body.String())
+ }
+}
+
+func mustJSON(t *testing.T, v any) string {
+ t.Helper()
+ b, err := json.Marshal(v)
+ if err != nil {
+ t.Fatal(err)
+ }
+ return string(b)
+}
+
+func evidenceTarget(t *testing.T, lastApplied map[string]any, managed []metav1.ManagedFieldsEntry) *unstructured.Unstructured {
+ t.Helper()
+ annotations := map[string]any{"cnpg.io/hibernation": "off", "team": "db"}
+ if lastApplied != nil {
+ annotations[lastAppliedAnnotationKey] = mustJSON(t, lastApplied)
+ }
+ u := &unstructured.Unstructured{Object: map[string]any{
+ "apiVersion": "apps/v1",
+ "kind": "Deployment",
+ "metadata": map[string]any{
+ "name": "api",
+ "namespace": "prod",
+ "uid": "uid-1",
+ "resourceVersion": "42",
+ "annotations": annotations,
+ "labels": map[string]any{"app": "api"},
+ },
+ "spec": map[string]any{
+ "replicas": int64(3),
+ "template": map[string]any{"spec": map[string]any{
+ "containers": []any{
+ map[string]any{"name": "sidecar", "image": "envoy:1"},
+ map[string]any{"name": "app", "image": "api:1"},
+ },
+ }},
+ },
+ }}
+ u.SetManagedFields(managed)
+ return u
+}
+
+func fieldsEntry(t *testing.T, manager string, op metav1.ManagedFieldsOperationType, fields map[string]any) metav1.ManagedFieldsEntry {
+ t.Helper()
+ return metav1.ManagedFieldsEntry{
+ Manager: manager,
+ Operation: op,
+ FieldsV1: &metav1.FieldsV1{Raw: []byte(mustJSON(t, fields))},
+ }
+}
+
+var deploymentRef = writeEvidenceRef{Kind: "Deployment", Group: "apps", Namespace: "prod", Name: "api"}
+
+func evidenceArgoApp(syncPolicy map[string]any, ignore []any) *unstructured.Unstructured {
+ spec := map[string]any{}
+ if syncPolicy != nil {
+ spec["syncPolicy"] = syncPolicy
+ }
+ if ignore != nil {
+ spec["ignoreDifferences"] = ignore
+ }
+ return &unstructured.Unstructured{Object: map[string]any{"spec": spec}}
+}
+
+var argoOwner = &topology.ResourceRef{Kind: "Application", Group: "argoproj.io", Namespace: "argocd", Name: "api"}
+
+func pathEvidence(t *testing.T, resp gitOpsWriteEvidenceResponse, path string) gitOpsPathEvidence {
+ t.Helper()
+ for _, p := range resp.Paths {
+ if p.Path == path {
+ return p
+ }
+ }
+ t.Fatalf("no evidence for %s", path)
+ return gitOpsPathEvidence{}
+}
+
+func TestParseWriteFieldPath(t *testing.T) {
+ segs, err := parseWriteFieldPath(`metadata.annotations["cnpg.io/hibernation"]`)
+ if err != nil || len(segs) != 3 || segs[2].key != "cnpg.io/hibernation" {
+ t.Fatalf("annotation path: %+v %v", segs, err)
+ }
+ segs, err = parseWriteFieldPath(`spec.template.spec.containers[name="app"].image`)
+ if err != nil || segs[4].typ != segMatch || segs[4].key != "name" || segs[4].value != "app" {
+ t.Fatalf("selector path: %+v %v", segs, err)
+ }
+ segs, err = parseWriteFieldPath(`spec.containers[*].image`)
+ if err != nil || segs[2].typ != segAny {
+ t.Fatalf("wildcard path: %+v %v", segs, err)
+ }
+ if _, err := parseWriteFieldPath(`spec.containers[0`); err == nil {
+ t.Fatal("expected unterminated bracket error")
+ }
+}
+
+func TestFieldsV1Owns(t *testing.T) {
+ fields := map[string]any{
+ "f:metadata": map[string]any{
+ "f:annotations": map[string]any{".": map[string]any{}, "f:cnpg.io/hibernation": map[string]any{}},
+ "f:labels": map[string]any{"f:app": map[string]any{}},
+ },
+ "f:spec": map[string]any{
+ "f:replicas": map[string]any{},
+ "f:template": map[string]any{"f:spec": map[string]any{
+ "f:containers": map[string]any{
+ `k:{"name":"app"}`: map[string]any{".": map[string]any{}, "f:image": map[string]any{}},
+ },
+ "f:tolerations": map[string]any{},
+ }},
+ },
+ }
+ cases := []struct {
+ path string
+ owned, appr bool
+ }{
+ {`metadata.annotations["cnpg.io/hibernation"]`, true, false},
+ {`metadata.annotations.team`, false, false},
+ {`metadata.labels.app`, true, false},
+ {`spec.replicas`, true, false},
+ {`spec.template.spec.containers[name=app].image`, true, false},
+ {`spec.template.spec.containers[name=sidecar].image`, false, false},
+ {`spec.template.spec.containers[*].image`, true, true},
+ {`spec.template.spec.tolerations[0].key`, true, true},
+ {`spec.paused`, false, false},
+ }
+ for _, tc := range cases {
+ segs, err := parseWriteFieldPath(tc.path)
+ if err != nil {
+ t.Fatal(err)
+ }
+ owned, approximate := fieldsV1Owns(fields, segs)
+ if owned != tc.owned || approximate != tc.appr {
+ t.Errorf("%s: owned=%v approximate=%v, want %v %v", tc.path, owned, approximate, tc.owned, tc.appr)
+ }
+ }
+}
+
+func TestBuildGitOpsWriteEvidence(t *testing.T) {
+ lastApplied := map[string]any{
+ "metadata": map[string]any{"annotations": map[string]any{"team": "db"}},
+ "spec": map[string]any{"template": map[string]any{"spec": map[string]any{
+ "containers": []any{map[string]any{"name": "app", "image": "api:1"}},
+ }}},
+ }
+ managed := []metav1.ManagedFieldsEntry{
+ fieldsEntry(t, "argocd-controller", metav1.ManagedFieldsOperationApply, map[string]any{
+ "f:spec": map[string]any{"f:replicas": map[string]any{}},
+ }),
+ fieldsEntry(t, "kubectl-edit", metav1.ManagedFieldsOperationUpdate, map[string]any{
+ "f:metadata": map[string]any{"f:annotations": map[string]any{"f:cnpg.io/hibernation": map[string]any{}}},
+ }),
+ }
+ target := evidenceTarget(t, lastApplied, managed)
+ app := evidenceArgoApp(map[string]any{"automated": map[string]any{"selfHeal": true, "prune": true}}, nil)
+ paths := []string{
+ `metadata.annotations.team`,
+ `metadata.annotations["cnpg.io/hibernation"]`,
+ `spec.replicas`,
+ `spec.template.spec.containers[name=app].image`,
+ `spec[`,
+ }
+ resp := buildGitOpsWriteEvidence(target, deploymentRef, paths, argoOwner, app, nil)
+
+ if resp.UID != "uid-1" || resp.ResourceVersion != "42" {
+ t.Fatalf("identity: %+v", resp)
+ }
+ if resp.Policy == nil || !*resp.Policy.Auto || !*resp.Policy.SelfHeal || !*resp.Policy.Prune {
+ t.Fatalf("policy: %+v", resp.Policy)
+ }
+ if got := pathEvidence(t, resp, `metadata.annotations.team`); got.LastApplied != "present" || got.OwnedByGitOps {
+ t.Errorf("team: %+v", got)
+ }
+ hib := pathEvidence(t, resp, `metadata.annotations["cnpg.io/hibernation"]`)
+ if hib.LastApplied != "absent" || hib.OwnedByGitOps || len(hib.OwnedBy) != 1 || hib.OwnedBy[0].Manager != "kubectl-edit" || hib.OwnedBy[0].Tool != "" {
+ t.Errorf("hibernation: %+v", hib)
+ }
+ rep := pathEvidence(t, resp, `spec.replicas`)
+ if rep.LastApplied != "absent" || !rep.OwnedByGitOps || rep.OwnedBy[0].Tool != "argocd" || rep.OwnedBy[0].Operation != "Apply" {
+ t.Errorf("replicas: %+v", rep)
+ }
+ if got := pathEvidence(t, resp, `spec.template.spec.containers[name=app].image`); got.LastApplied != "present" {
+ t.Errorf("image: %+v", got)
+ }
+ if got := pathEvidence(t, resp, `spec[`); got.Error == "" {
+ t.Errorf("malformed path should carry an error: %+v", got)
+ }
+
+ // The response must never carry the raw annotation or managedFields.
+ raw := mustJSON(t, resp)
+ for _, leak := range []string{"last-applied", "f:spec", "fieldsV1"} {
+ if strings.Contains(raw, leak) {
+ t.Errorf("response leaks %q: %s", leak, raw)
+ }
+ }
+}
+
+func TestWriteEvidenceNoLastApplied(t *testing.T) {
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas"}, argoOwner, evidenceArgoApp(nil, nil), nil)
+ if got := resp.Paths[0]; got.LastApplied != "no-annotation" || got.OwnedByGitOps {
+ t.Errorf("%+v", got)
+ }
+ if resp.Policy == nil || *resp.Policy.Auto || *resp.Policy.SelfHeal {
+ t.Errorf("manual sync policy: %+v", resp.Policy)
+ }
+}
+
+func TestWriteEvidenceArgoIgnoreDifferences(t *testing.T) {
+ ignore := []any{map[string]any{"group": "apps", "kind": "Deployment", "jsonPointers": []any{"/spec/replicas", "/spec/template/spec/containers/1/image"}}}
+ paths := []string{"spec.replicas", "spec.template.spec.containers[name=app].image", "spec.template.spec.containers[name=sidecar].image"}
+
+ off := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, paths, argoOwner,
+ evidenceArgoApp(map[string]any{"automated": map[string]any{"selfHeal": true}}, ignore), nil)
+ if got := off.Paths[0]; got.Ignored != "comparison-only" || got.IgnoredBy != "spec.ignoreDifferences" {
+ t.Errorf("without RespectIgnoreDifferences: %+v", got)
+ }
+
+ on := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, paths, argoOwner,
+ evidenceArgoApp(map[string]any{
+ "automated": map[string]any{"selfHeal": true},
+ "syncOptions": []any{"RespectIgnoreDifferences=true"},
+ }, ignore), nil)
+ if !on.Policy.RespectIgnoreDifferences {
+ t.Fatalf("policy: %+v", on.Policy)
+ }
+ if on.Paths[0].Ignored != "effective" || on.Paths[1].Ignored != "effective" {
+ t.Errorf("with RespectIgnoreDifferences: %+v", on.Paths)
+ }
+ if on.Paths[2].Ignored != "" {
+ t.Errorf("sidecar (index 0) must not be covered: %+v", on.Paths[2])
+ }
+
+ // A rule for another kind, or one whose group is omitted (core), doesn't match apps/Deployment.
+ other := []any{map[string]any{"kind": "Deployment", "jsonPointers": []any{"/spec/replicas"}}}
+ miss := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, paths[:1], argoOwner, evidenceArgoApp(nil, other), nil)
+ if miss.Paths[0].Ignored != "" {
+ t.Errorf("core-group rule matched apps/Deployment: %+v", miss.Paths[0])
+ }
+}
+
+// Radar doesn't evaluate jq, so a matching jq rule leaves coverage open
+// rather than reading as "not ignored".
+func TestWriteEvidenceArgoJQRuleIsUnevaluated(t *testing.T) {
+ ignore := []any{map[string]any{"group": "apps", "kind": "Deployment", "jqPathExpressions": []any{".spec.replicas"}}}
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas"}, argoOwner,
+ evidenceArgoApp(map[string]any{"syncOptions": []any{"RespectIgnoreDifferences=true"}}, ignore), nil)
+ if got := resp.Paths[0]; got.Ignored != "unevaluated" {
+ t.Errorf("jq rule: %+v", got)
+ }
+ both := []any{map[string]any{"group": "apps", "kind": "Deployment", "jsonPointers": []any{"/spec/replicas"}, "jqPathExpressions": []any{".spec.x"}}}
+ resp = buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas"}, argoOwner,
+ evidenceArgoApp(map[string]any{"syncOptions": []any{"RespectIgnoreDifferences=true"}}, both), nil)
+ if got := resp.Paths[0]; got.Ignored != "effective" {
+ t.Errorf("a pointer that covers the path wins over an unevaluated jq rule: %+v", got)
+ }
+}
+
+func TestWriteEvidenceArgoDisabledAutomation(t *testing.T) {
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, nil, argoOwner,
+ evidenceArgoApp(map[string]any{"automated": map[string]any{"enabled": false, "selfHeal": true}}, nil), nil)
+ if *resp.Policy.Auto || *resp.Policy.SelfHeal {
+ t.Errorf("automated.enabled=false must disable auto-sync: %+v", resp.Policy)
+ }
+}
+
+func TestWriteEvidenceHelmReleaseDriftDetection(t *testing.T) {
+ owner := &topology.ResourceRef{Kind: "HelmRelease", Group: "helm.toolkit.fluxcd.io", Namespace: "flux-system", Name: "api"}
+ hr := func(drift map[string]any) *unstructured.Unstructured {
+ spec := map[string]any{"interval": "5m"}
+ if drift != nil {
+ spec["driftDetection"] = drift
+ }
+ return &unstructured.Unstructured{Object: map[string]any{"spec": spec}}
+ }
+ for _, tc := range []struct {
+ drift map[string]any
+ mode string
+ selfHeal bool
+ }{
+ {nil, "disabled", false},
+ {map[string]any{"mode": "warn"}, "warn", false},
+ {map[string]any{"mode": "enabled"}, "enabled", true},
+ } {
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas"}, owner, hr(tc.drift), nil)
+ if resp.Policy.DriftDetection != tc.mode || *resp.Policy.SelfHeal != tc.selfHeal || resp.Policy.Interval != "5m" {
+ t.Errorf("mode %s: %+v", tc.mode, resp.Policy)
+ }
+ }
+
+ ignored := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas", "spec.paused"}, owner,
+ hr(map[string]any{"mode": "enabled", "ignore": []any{map[string]any{"paths": []any{"/spec/replicas"}, "target": map[string]any{"kind": "Deployment"}}}}), nil)
+ if ignored.Paths[0].Ignored != "effective" || ignored.Paths[0].IgnoredBy != "spec.driftDetection.ignore" || ignored.Paths[1].Ignored != "" {
+ t.Errorf("drift ignore: %+v", ignored.Paths)
+ }
+}
+
+func TestWriteEvidenceFluxKustomization(t *testing.T) {
+ owner := &topology.ResourceRef{Kind: "Kustomization", Group: "kustomize.toolkit.fluxcd.io", Namespace: "flux-system", Name: "apps"}
+ ks := &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{"suspend": true, "interval": "10m", "prune": true}}}
+ managed := []metav1.ManagedFieldsEntry{fieldsEntry(t, "kustomize-controller", metav1.ManagedFieldsOperationApply, map[string]any{
+ "f:spec": map[string]any{"f:replicas": map[string]any{}},
+ })}
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, managed), deploymentRef, []string{"spec.replicas"}, owner, ks, nil)
+ if !*resp.Policy.Suspended || *resp.Policy.SelfHeal || *resp.Policy.Auto || !*resp.Policy.Prune || resp.Policy.Interval != "10m" {
+ t.Errorf("suspended kustomization: %+v", resp.Policy)
+ }
+ if !resp.Paths[0].OwnedByGitOps {
+ t.Errorf("kustomize-controller ownership: %+v", resp.Paths[0])
+ }
+
+ target := evidenceTarget(t, nil, nil)
+ annotations := target.GetAnnotations()
+ annotations["kustomize.toolkit.fluxcd.io/ssa"] = "IfNotPresent"
+ target.SetAnnotations(annotations)
+ ifNotPresent := buildGitOpsWriteEvidence(target, deploymentRef, []string{"spec.replicas"}, owner,
+ &unstructured.Unstructured{Object: map[string]any{"spec": map[string]any{}}}, nil)
+ if ifNotPresent.Policy.ObjectReconcile != "if-not-present" || ifNotPresent.Paths[0].Ignored != "effective" {
+ t.Errorf("ssa IfNotPresent: %+v %+v", ifNotPresent.Policy, ifNotPresent.Paths[0])
+ }
+}
+
+func TestWriteEvidenceOwnerUnreadable(t *testing.T) {
+ forbidden := apierrors.NewForbidden(schema.GroupResource{Group: "argoproj.io", Resource: "applications"}, "api", errors.New("denied"))
+ resp := buildGitOpsWriteEvidence(evidenceTarget(t, nil, nil), deploymentRef, []string{"spec.replicas"}, argoOwner, nil, forbidden)
+ if resp.Policy != nil || resp.PolicyError == "" {
+ t.Errorf("unreadable owner must leave policy unknown with a reason: %+v", resp)
+ }
+}
+
+func TestWriteEvidenceControllerOwnerAndMetadataOwner(t *testing.T) {
+ target := evidenceTarget(t, nil, nil)
+ controller := true
+ target.SetOwnerReferences([]metav1.OwnerReference{{APIVersion: "postgresql.cnpg.io/v1", Kind: "Cluster", Name: "pg", Controller: &controller}})
+ target.SetAnnotations(map[string]string{"argocd.argoproj.io/tracking-id": "argocd_api:apps/Deployment:prod/api"})
+ owner := resolveWriteEvidenceOwner(gitOpsWriteEvidenceRequest{writeEvidenceRef: deploymentRef}, target)
+ if owner == nil || owner.Kind != "Application" || owner.Name != "api" || owner.Namespace != "argocd" {
+ t.Fatalf("owner from tracking metadata: %+v", owner)
+ }
+ resp := buildGitOpsWriteEvidence(target, deploymentRef, nil, owner, nil, nil)
+ if resp.ControllerOwner == nil || resp.ControllerOwner.Kind != "Cluster" || resp.ControllerOwner.Name != "pg" {
+ t.Errorf("controller owner: %+v", resp.ControllerOwner)
+ }
+}
diff --git a/internal/server/issue_evidence_contract_test.go b/internal/server/issue_evidence_contract_test.go
new file mode 100644
index 0000000000..69d6eac363
--- /dev/null
+++ b/internal/server/issue_evidence_contract_test.go
@@ -0,0 +1,66 @@
+package server
+
+import (
+ "go/ast"
+ "go/parser"
+ "go/token"
+ "os"
+ "path/filepath"
+ "strings"
+ "testing"
+)
+
+func TestUserIssueProjectionsSupplyEvidenceAuthorizer(t *testing.T) {
+ for _, dir := range []string{".", "../mcp"} {
+ entries, err := os.ReadDir(dir)
+ if err != nil {
+ t.Fatal(err)
+ }
+ for _, entry := range entries {
+ if !strings.HasSuffix(entry.Name(), ".go") || strings.HasSuffix(entry.Name(), "_test.go") {
+ continue
+ }
+ path := filepath.Join(dir, entry.Name())
+ fset := token.NewFileSet()
+ file, err := parser.ParseFile(fset, path, nil, 0)
+ if err != nil {
+ t.Fatal(err)
+ }
+ ast.Inspect(file, func(node ast.Node) bool {
+ literal, ok := node.(*ast.CompositeLit)
+ if !ok {
+ return true
+ }
+ selector, ok := literal.Type.(*ast.SelectorExpr)
+ if !ok {
+ return true
+ }
+ owner, ok := selector.X.(*ast.Ident)
+ if !ok || owner.Name != "issues" || (selector.Sel.Name != "Filters" && selector.Sel.Name != "RelatedIssueOptions") {
+ return true
+ }
+ authorized := false
+ for _, field := range literal.Elts {
+ kv, ok := field.(*ast.KeyValueExpr)
+ if !ok {
+ continue
+ }
+ key, ok := kv.Key.(*ast.Ident)
+ if !ok {
+ continue
+ }
+ if key.Name == "CanReadEvidence" {
+ authorized = true
+ }
+ if key.Name == "AllowUnfilteredEvidence" && entry.Name() == "capacity_issues.go" {
+ authorized = true
+ }
+ }
+ if !authorized {
+ t.Errorf("%s: issue projection lacks an evidence authorizer", fset.Position(literal.Pos()))
+ }
+ return true
+ })
+ }
+ }
+}
diff --git a/internal/server/issues_handler.go b/internal/server/issues_handler.go
index 6120fb059e..40902ea6f7 100644
--- a/internal/server/issues_handler.go
+++ b/internal/server/issues_handler.go
@@ -11,13 +11,16 @@ import (
"time"
"github.com/go-chi/chi/v5"
+
"github.com/skyhook-io/radar/internal/auth"
"github.com/skyhook-io/radar/internal/filter"
"github.com/skyhook-io/radar/internal/helm"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/meaningfulchanges"
"github.com/skyhook-io/radar/pkg/issuesapi"
+ "github.com/skyhook-io/radar/pkg/resourceid"
)
// filterRecentChangesByRBAC drops RecentChange rows the ctx user can't read, via
@@ -72,7 +75,7 @@ func (s *Server) handleIssues(w http.ResponseWriter, r *http.Request) {
// is unrestricted); non-nil empty = "user has no access to anything
// they asked for".
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
// If the caller EXPLICITLY named namespace(s) they can't access, that's
// a denial — surface it as 403, not an empty (reads-as-"nothing broken")
// list. Bad trust boundary otherwise, especially for an agent.
@@ -100,6 +103,7 @@ func (s *Server) handleIssues(w http.ResponseWriter, r *http.Request) {
Grouped: q.Get("view") != "flat",
CanReadClusterScoped: s.issueClusterScopedAccess(r),
CanReadRelated: s.issueRelatedResourceAccess(r),
+ CanReadEvidence: s.issueEvidenceAccess(r),
}
if expr := q.Get("filter"); expr != "" {
f, err := filter.CachedIssueFilter(expr)
@@ -236,10 +240,12 @@ func (s *Server) nativeHelmIssuesForRequest(r *http.Request, namespaces []string
// section in the resource detail. Namespace "_" denotes a cluster-scoped resource;
// optional ?group= disambiguates a CRD whose kind collides with a core kind.
//
-// RBAC: namespaced targets are gated by the namespace auth-filter (the frontend
-// passes ?namespaces= to scope the scan); cluster-scoped targets are gated by
-// the same list permission /api/issues uses, so this can't surface a node's
-// issues to a user who can't list nodes.
+// RBAC: the drawer's preflight (namespace access, cluster-scoped get) and then
+// a get-SAR on the subject's own kind: namespace access does not imply reading
+// every kind in it, and an issue's message describes the object. Grouped issues
+// whose subject the caller can't get are withheld, as are members of kinds it
+// can't get; ?coverage=1 returns the envelope that counts them and says whether
+// Radar is watching the subject's kind at all.
func (s *Server) handleResourceIssues(w http.ResponseWriter, r *http.Request) {
if !s.requireConnected(w) {
return
@@ -280,6 +286,16 @@ func (s *Server) handleResourceIssues(w http.ResponseWriter, r *http.Request) {
}
}
}
+ if kind == rawKind {
+ if b, ok := resourceid.BuiltinForName(rawKind); ok && (group == "" || group == b.Group) {
+ kind = b.Kind
+ }
+ }
+
+ if !s.canGetIssueRef(r, issues.Ref{Group: group, Kind: kind, Namespace: namespace, Name: name}) {
+ s.writeError(w, http.StatusForbidden, fmt.Sprintf("no access to %s %s", kind, name))
+ return
+ }
// Scope the scan to the resource's namespace (a workload's owned pods live
// there too); cluster-scoped resources scan all namespaces (nil).
@@ -292,11 +308,202 @@ func (s *Server) handleResourceIssues(w http.ResponseWriter, r *http.Request) {
Namespaces: namespaces,
CanReadClusterScoped: s.issueClusterScopedAccess(r),
CanReadRelated: s.issueRelatedResourceAccess(r),
+ CanReadEvidence: s.issueEvidenceAccess(r),
}, group, kind, namespace, name)
+ related, withheld := s.withholdUnreadableIssueRefs(r, related)
if related == nil {
related = []issues.Issue{}
}
- s.writeJSON(w, related)
+ if r.URL.Query().Get("coverage") != "1" {
+ s.writeJSON(w, related)
+ return
+ }
+ withheld.Issues += s.clusterScopedSubjectIssuesWithheld(r, provider, related, group, kind, namespace, name)
+ resp := ResourceIssuesResponse{
+ Issues: related,
+ Coverage: resourceIssuesCoverage(kind, group, namespace),
+ }
+ if withheld.Issues > 0 || withheld.Members > 0 {
+ resp.Withheld = &withheld
+ }
+ if result := k8s.GetCachedPermissionResult(); result != nil {
+ resp.Visibility = k8s.BuildVisibilitySummary(result, k8s.VisibilityNamespace(namespaces))
+ }
+ s.writeJSON(w, resp)
+}
+
+// ResourceIssuesResponse is /api/issues/resource with ?coverage=1.
+type ResourceIssuesResponse struct {
+ Issues []issues.Issue `json:"issues"`
+ // Whether the issues engine reads the subject's kind: ok, syncing, or
+ // notWatched. Anything but ok makes an empty list unknown, not none.
+ Coverage string `json:"coverage"`
+ Withheld *ResourceIssuesWithheld `json:"withheld,omitempty"`
+ Visibility *k8s.VisibilitySummary `json:"visibility,omitempty"`
+}
+
+// ResourceIssuesWithheld counts what the caller's RBAC kept out of the answer.
+type ResourceIssuesWithheld struct {
+ // Grouped issues whose subject the caller can't get.
+ Issues int `json:"issues"`
+ // Member refs of kinds the caller can't get, dropped from returned issues.
+ Members int `json:"members"`
+}
+
+const (
+ resourceIssuesCoverageOK = "ok"
+ resourceIssuesCoverageSyncing = "syncing"
+ resourceIssuesCoverageNotWatched = "notWatched"
+)
+
+// canGetIssueRef reports whether the caller may get ref: its kind where it
+// lives, or failing that the named object, which a resourceNames-restricted
+// grant allows. Unresolvable kinds fail closed.
+func (s *Server) canGetIssueRef(r *http.Request, ref issues.Ref) bool {
+ user := auth.UserFromContext(r.Context())
+ if user == nil {
+ return true
+ }
+ group, resource, clusterScoped, ok := k8s.ResolveChangeGVR(ref.Kind, ref.Group)
+ if !ok {
+ return false
+ }
+ namespace := ref.Namespace
+ if clusterScoped {
+ namespace = ""
+ }
+ if s.canRead(r, group, resource, namespace, "get") {
+ return true
+ }
+ if ref.Name == "" || s.permCache == nil {
+ return false
+ }
+ perms := s.permCache.Get(user.Username, user.Groups)
+ if perms != nil {
+ if v, ok := perms.CanINamed("get", group, resource, namespace, ref.Name); ok {
+ return v
+ }
+ }
+ client := k8s.GetClient()
+ if client == nil {
+ return false
+ }
+ allowed, err := auth.SubjectCanINamed(r.Context(), client, user.Username, user.Groups, namespace, group, resource, ref.Name, "get")
+ if err != nil {
+ return false
+ }
+ if perms != nil {
+ perms.SetCanINamed("get", group, resource, namespace, ref.Name, allowed)
+ }
+ return allowed
+}
+
+// withholdUnreadableIssueRefs drops grouped issues whose subject the caller
+// can't get and member refs of kinds it can't get, counting both. A trimmed
+// member list is marked truncated so it never reads as the whole fan-out.
+func (s *Server) withholdUnreadableIssueRefs(r *http.Request, in []issues.Issue) ([]issues.Issue, ResourceIssuesWithheld) {
+ var withheld ResourceIssuesWithheld
+ if auth.UserFromContext(r.Context()) == nil {
+ return in, withheld
+ }
+ out := make([]issues.Issue, 0, len(in))
+ for _, issue := range in {
+ if !s.canGetIssueRef(r, issues.Ref{Group: issue.Group, Kind: issue.Kind, Namespace: issue.Namespace, Name: issue.Name}) {
+ withheld.Issues++
+ continue
+ }
+ if len(issue.Members) > 0 {
+ kept := make([]issues.Ref, 0, len(issue.Members))
+ for _, m := range issue.Members {
+ if s.canGetIssueRef(r, m) {
+ kept = append(kept, m)
+ } else {
+ withheld.Members++
+ }
+ }
+ if len(kept) < len(issue.Members) {
+ issue.Members = kept
+ issue.MembersTruncated = true
+ }
+ }
+ out = append(out, issue)
+ }
+ return out, withheld
+}
+
+// clusterScopedSubjectIssuesWithheld counts the issues about a cluster-scoped
+// subject that the composition left out because the caller can get it but not
+// list its kind (the gate /api/issues applies). Without the count a NotReady
+// Node would read as having none.
+func (s *Server) clusterScopedSubjectIssuesWithheld(r *http.Request, provider issues.Provider, returned []issues.Issue, group, kind, namespace, name string) int {
+ if namespace != "" || auth.UserFromContext(r.Context()) == nil {
+ return 0
+ }
+ resolvedGroup, _, clusterScoped, ok := k8s.ResolveChangeGVR(kind, group)
+ if !ok || !clusterScoped || s.issueClusterScopedAccess(r)(kind, resolvedGroup) {
+ return 0
+ }
+ subjectKind := func(k, g string) bool {
+ return strings.EqualFold(k, kind) && resourceid.NormalizeGroup(g) == resourceid.NormalizeGroup(resolvedGroup)
+ }
+ all := issues.Compose(provider, issues.Filters{
+ SkipPodTemplateContext: true,
+ Kinds: []string{kind},
+ Limit: issues.NoLimit,
+ CanReadClusterScoped: subjectKind,
+ CanReadRelated: s.issueRelatedResourceAccess(r),
+ CanReadEvidence: s.issueEvidenceAccess(r),
+ Grouped: true,
+ })
+ seen := make(map[string]bool, len(returned))
+ for _, i := range returned {
+ seen[i.ID] = true
+ }
+ n := 0
+ for _, i := range all {
+ if !seen[i.ID] && i.Name == name && i.Namespace == "" && subjectKind(i.Kind, i.Group) {
+ n++
+ }
+ }
+ return n
+}
+
+// resourceIssuesCoverage reports whether the issues engine has the subject's
+// kind to read: a typed informer, or a dynamic one synced for its namespace.
+func resourceIssuesCoverage(kind, group, namespace string) string {
+ if cache := k8s.GetResourceCache(); cache != nil {
+ if g, resource, _, ok := k8s.ResolveChangeGVR(kind, group); ok {
+ if b, builtin := resourceid.BuiltinForKind(kind); builtin && b.Group == g {
+ if synced, known := cache.InformerSynced(resource); known {
+ switch {
+ case !cache.KindCoversNamespace(resource, namespace):
+ return resourceIssuesCoverageNotWatched
+ case !synced:
+ return resourceIssuesCoverageSyncing
+ }
+ return resourceIssuesCoverageOK
+ }
+ }
+ }
+ }
+ disc := k8s.GetResourceDiscovery()
+ dyn := k8s.GetDynamicResourceCache()
+ if disc == nil || dyn == nil {
+ return resourceIssuesCoverageNotWatched
+ }
+ gvr, ok := disc.GetGVRWithGroup(kind, group)
+ if !ok {
+ return resourceIssuesCoverageNotWatched
+ }
+ if dyn.IsNamespaceSynced(gvr, namespace) {
+ return resourceIssuesCoverageOK
+ }
+ for _, watched := range dyn.GetWatchedResources() {
+ if watched == gvr {
+ return resourceIssuesCoverageSyncing
+ }
+ }
+ return resourceIssuesCoverageNotWatched
}
func parseSeverities(v string) ([]issues.Severity, error) {
@@ -334,3 +541,9 @@ func splitCSV(v string) []string {
}
return out
}
+
+func (s *Server) issueEvidenceAccess(r *http.Request) func(issues.EvidenceRead) bool {
+ return func(read issues.EvidenceRead) bool {
+ return s.canRead(r, read.Group, read.Resource, read.Namespace, read.Verb)
+ }
+}
diff --git a/internal/server/issues_related_auth_test.go b/internal/server/issues_related_auth_test.go
index 0ec8c6d2c6..8d9be83b27 100644
--- a/internal/server/issues_related_auth_test.go
+++ b/internal/server/issues_related_auth_test.go
@@ -14,6 +14,31 @@ import (
dynamicfake "k8s.io/client-go/dynamic/fake"
)
+func TestRESTCachedIssuesAuthorizeCNPGInventories(t *testing.T) {
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ perms := &auth.UserPermissions{AllowedNamespaces: []string{"db"}}
+ perms.SetCanI("list", "postgresql.cnpg.io", "clusters", "db", true)
+ perms.SetCanI("get", "", "pods", "db", true)
+ perms.SetCanI("list", "", "pods", "db", false)
+ s.permCache.Set("cnpg-evidence", nil, perms)
+ r := requestWithUser(http.MethodGet, "/api/issues/resource/cluster/db/pg", &auth.User{Username: "cnpg-evidence"})
+ issue := issues.Issue{ID: "contradiction", Group: "postgresql.cnpg.io", Kind: "Cluster", Namespace: "db", Name: "pg", RequiredReads: []issues.EvidenceRead{
+ {Group: "postgresql.cnpg.io", Resource: "clusters", Namespace: "db", Verb: "list"},
+ {Resource: "pods", Namespace: "db", Verb: "list"},
+ }}
+ options := issues.RelatedIssueOptions{CanReadEvidence: s.issueEvidenceAccess(r)}
+ lookup := func() []issues.Issue {
+ return issues.RelatedIssuesFrom(nil, []issues.Issue{issue}, options, issue.Group, issue.Kind, issue.Namespace, issue.Name)
+ }
+ if len(lookup()) != 0 {
+ t.Fatal("get Pods must not authorize a list-dependent finding")
+ }
+ perms.SetCanI("list", "", "pods", "db", true)
+ if len(lookup()) != 1 {
+ t.Fatal("authorized inventory withheld")
+ }
+}
+
func TestRESTRelatedIssuesAuthorizeNodeClassSubject(t *testing.T) {
initRelatedIssueAuthDiscovery(t)
s := newAuthServer(auth.Config{Mode: "proxy"})
diff --git a/internal/server/issues_resource_auth_test.go b/internal/server/issues_resource_auth_test.go
new file mode 100644
index 0000000000..b51e0eea6e
--- /dev/null
+++ b/internal/server/issues_resource_auth_test.go
@@ -0,0 +1,245 @@
+package server
+
+import (
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "testing"
+ "time"
+
+ "github.com/go-chi/chi/v5"
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/issues"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+ authv1 "k8s.io/api/authorization/v1"
+ corev1 "k8s.io/api/core/v1"
+ rbacv1 "k8s.io/api/rbac/v1"
+ metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
+ "k8s.io/client-go/kubernetes/fake"
+)
+
+func resourceIssuesRequest(t *testing.T, s *Server, path string) *httptest.ResponseRecorder {
+ t.Helper()
+ rt := chi.NewRouter()
+ rt.Get("/api/issues/resource/{kind}/{namespace}/{name}", s.handleResourceIssues)
+ w := httptest.NewRecorder()
+ rt.ServeHTTP(w, requestWithUser(http.MethodGet, path, &auth.User{Username: "pg-user"}))
+ return w
+}
+
+// Fixture: TestMain seeds Deployment broken/stuck-app with no available replicas.
+func TestResourceIssuesRequireGetOnTheSubjectKind(t *testing.T) {
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ perms := &auth.UserPermissions{AllowedNamespaces: nil}
+ s.permCache.Set("pg-user", nil, perms)
+
+ perms.SetCanI("get", "apps", "deployments", "broken", false)
+ w := resourceIssuesRequest(t, s, "/api/issues/resource/Deployment/broken/stuck-app?group=apps")
+ if w.Code != http.StatusForbidden {
+ t.Fatalf("denied get on deployments: status %d, want 403 (body %s)", w.Code, w.Body.String())
+ }
+
+ perms.SetCanI("get", "apps", "deployments", "broken", true)
+ w = resourceIssuesRequest(t, s, "/api/issues/resource/Deployment/broken/stuck-app?group=apps&coverage=1")
+ if w.Code != http.StatusOK {
+ t.Fatalf("allowed: status %d (body %s)", w.Code, w.Body.String())
+ }
+ var resp ResourceIssuesResponse
+ if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
+ t.Fatalf("decode: %v", err)
+ }
+ if len(resp.Issues) == 0 {
+ t.Fatal("allowed caller got no issues for the broken Deployment")
+ }
+ if resp.Coverage != resourceIssuesCoverageOK {
+ t.Fatalf("coverage %q, want ok for a synced typed kind", resp.Coverage)
+ }
+
+ w = resourceIssuesRequest(t, s, "/api/issues/resource/deployments/broken/stuck-app")
+ if w.Code != http.StatusOK {
+ t.Fatalf("plural kind: status %d (body %s)", w.Code, w.Body.String())
+ }
+}
+
+func TestResourceIssuesWithholdUnreadableSubjectsAndMembers(t *testing.T) {
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ perms := &auth.UserPermissions{AllowedNamespaces: nil}
+ perms.SetCanI("get", "apps", "deployments", "pg", true)
+ perms.SetCanI("get", "", "pods", "pg", false)
+ perms.SetCanI("get", "apps", "statefulsets", "pg", false)
+ s.permCache.Set("pg-user", nil, perms)
+ r := requestWithUser(http.MethodGet, "/api/issues/resource/Deployment/pg/app", &auth.User{Username: "pg-user"})
+
+ in := []issues.Issue{
+ {ID: "a", Group: "apps", Kind: "Deployment", Namespace: "pg", Name: "app", Members: []issues.Ref{
+ {Kind: "Pod", Namespace: "pg", Name: "app-1"},
+ {Kind: "Pod", Namespace: "pg", Name: "app-2"},
+ {Group: "apps", Kind: "Deployment", Namespace: "pg", Name: "app"},
+ }},
+ {ID: "b", Group: "apps", Kind: "StatefulSet", Namespace: "pg", Name: "db"},
+ }
+ out, withheld := s.withholdUnreadableIssueRefs(r, in)
+ if withheld.Issues != 1 || withheld.Members != 2 {
+ t.Fatalf("withheld %+v, want 1 issue and 2 members", withheld)
+ }
+ if len(out) != 1 || out[0].ID != "a" {
+ t.Fatalf("kept %+v, want only the Deployment's issue", out)
+ }
+ if len(out[0].Members) != 1 || !out[0].MembersTruncated {
+ t.Fatalf("members %+v truncated=%v, want the readable one and truncated", out[0].Members, out[0].MembersTruncated)
+ }
+}
+
+func TestResourceIssuesWithholdNothingWithoutAuth(t *testing.T) {
+ s := newAuthServer(auth.Config{Mode: "none"})
+ r := httptest.NewRequest(http.MethodGet, "/api/issues/resource/Deployment/pg/app", nil)
+ in := []issues.Issue{{ID: "a", Kind: "Pod", Namespace: "pg", Name: "p"}}
+ if out, withheld := s.withholdUnreadableIssueRefs(r, in); len(out) != 1 || withheld != (ResourceIssuesWithheld{}) {
+ t.Fatalf("auth off: out %+v withheld %+v", out, withheld)
+ }
+}
+
+func TestResourceIssuesCoverageOfAnUnwatchedKind(t *testing.T) {
+ if got := resourceIssuesCoverage("Cluster", "postgresql.cnpg.io", "pg"); got != resourceIssuesCoverageNotWatched {
+ t.Fatalf("coverage %q, want notWatched for a kind Radar has not discovered", got)
+ }
+}
+
+func TestResourceIssuesHonourAResourceNamesGrant(t *testing.T) {
+ fakeSARServer(t, func(a authv1.ResourceAttributes) bool {
+ return a.Group == "apps" && a.Resource == "deployments" && a.Namespace == "broken" && a.Name == "stuck-app" && a.Verb == "get"
+ })
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ s.permCache.Set("pg-user", nil, &auth.UserPermissions{AllowedNamespaces: nil})
+
+ w := resourceIssuesRequest(t, s, "/api/issues/resource/Deployment/broken/stuck-app?group=apps&coverage=1")
+ if w.Code != http.StatusOK {
+ t.Fatalf("named grant: status %d, want 200 (body %s)", w.Code, w.Body.String())
+ }
+ var resp ResourceIssuesResponse
+ if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
+ t.Fatalf("decode: %v", err)
+ }
+ if len(resp.Issues) == 0 {
+ t.Fatal("named grant got no issues for the broken Deployment")
+ }
+
+ if w := resourceIssuesRequest(t, s, "/api/issues/resource/Deployment/broken/other-app?group=apps"); w.Code != http.StatusForbidden {
+ t.Fatalf("a name outside the grant: status %d, want 403", w.Code)
+ }
+}
+
+func restoreFixtureCache(t *testing.T) {
+ t.Helper()
+ t.Cleanup(func() {
+ k8s.ResetResourceCache()
+ if err := k8s.InitTestResourceCache(testFakeClient); err != nil {
+ t.Fatalf("restore package fixture cache: %v", err)
+ }
+ })
+}
+
+func TestResourceIssuesCoverageOfAKindWatchedElsewhere(t *testing.T) {
+ restoreFixtureCache(t)
+ k8s.ResetResourceCache()
+ if err := k8s.InitScopedTestResourceCache(testFakeClient, map[string]k8score.ResourceScope{
+ "deployments": {Enabled: true, Namespace: "default"},
+ }); err != nil {
+ t.Fatalf("InitScopedTestResourceCache: %v", err)
+ }
+ if got := resourceIssuesCoverage("Deployment", "apps", "broken"); got != resourceIssuesCoverageNotWatched {
+ t.Fatalf("coverage in an unwatched namespace %q, want notWatched", got)
+ }
+ if got := resourceIssuesCoverage("Deployment", "apps", "default"); got != resourceIssuesCoverageOK {
+ t.Fatalf("coverage in the watched namespace %q, want ok", got)
+ }
+}
+
+func TestResourceIssuesCountNodeIssuesTheListGateDropped(t *testing.T) {
+ restoreFixtureCache(t)
+ k8s.ResetResourceCache()
+ since := metav1.NewTime(time.Now().Add(-time.Hour))
+ client := fake.NewClientset(&corev1.Node{
+ ObjectMeta: metav1.ObjectMeta{Name: "n1", CreationTimestamp: since},
+ Status: corev1.NodeStatus{Conditions: []corev1.NodeCondition{{
+ Type: corev1.NodeReady, Status: corev1.ConditionFalse, LastTransitionTime: since,
+ }}},
+ })
+ if err := k8s.InitTestResourceCache(client); err != nil {
+ t.Fatalf("InitTestResourceCache: %v", err)
+ }
+
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ perms := &auth.UserPermissions{AllowedNamespaces: nil}
+ perms.SetCanI("get", "", "nodes", "", true)
+ perms.SetCanI("list", "", "nodes", "", false)
+ s.permCache.Set("pg-user", nil, perms)
+
+ w := resourceIssuesRequest(t, s, "/api/issues/resource/Node/_/n1?coverage=1")
+ if w.Code != http.StatusOK {
+ t.Fatalf("status %d (body %s)", w.Code, w.Body.String())
+ }
+ var resp ResourceIssuesResponse
+ if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
+ t.Fatalf("decode: %v", err)
+ }
+ if len(resp.Issues) != 0 {
+ t.Fatalf("a caller who cannot list nodes got node issues: %+v", resp.Issues)
+ }
+ if resp.Withheld == nil || resp.Withheld.Issues == 0 {
+ t.Fatalf("withheld %+v: the NotReady node's issue vanished instead of being counted", resp.Withheld)
+ }
+
+ perms.SetCanI("list", "", "nodes", "", true)
+ w = resourceIssuesRequest(t, s, "/api/issues/resource/Node/_/n1?coverage=1")
+ resp = ResourceIssuesResponse{}
+ if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
+ t.Fatalf("decode: %v", err)
+ }
+ if len(resp.Issues) == 0 || resp.Withheld != nil {
+ t.Fatalf("a caller who can list nodes: issues %d withheld %+v", len(resp.Issues), resp.Withheld)
+ }
+}
+
+func TestResourceIssuesCountWithheldWithAndWithoutTheGroupParam(t *testing.T) {
+ restoreFixtureCache(t)
+ k8s.ResetResourceCache()
+ client := fake.NewClientset(&rbacv1.ClusterRoleBinding{
+ ObjectMeta: metav1.ObjectMeta{Name: "dangling"},
+ RoleRef: rbacv1.RoleRef{APIGroup: "rbac.authorization.k8s.io", Kind: "ClusterRole", Name: "gone"},
+ Subjects: []rbacv1.Subject{{Kind: "User", Name: "someone"}},
+ })
+ if err := k8s.InitTestResourceCache(client); err != nil {
+ t.Fatalf("InitTestResourceCache: %v", err)
+ }
+
+ s := newAuthServer(auth.Config{Mode: "proxy"})
+ perms := &auth.UserPermissions{AllowedNamespaces: nil}
+ perms.SetCanI("get", "rbac.authorization.k8s.io", "clusterrolebindings", "", true)
+ perms.SetCanI("list", "rbac.authorization.k8s.io", "clusterrolebindings", "", false)
+ s.permCache.Set("pg-user", nil, perms)
+
+ for _, path := range []string{
+ "/api/issues/resource/ClusterRoleBinding/_/dangling?coverage=1",
+ "/api/issues/resource/ClusterRoleBinding/_/dangling?coverage=1&group=rbac.authorization.k8s.io",
+ } {
+ w := resourceIssuesRequest(t, s, path)
+ if w.Code != http.StatusOK {
+ t.Fatalf("%s: status %d (body %s)", path, w.Code, w.Body.String())
+ }
+ var resp ResourceIssuesResponse
+ if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
+ t.Fatalf("%s: decode: %v", path, err)
+ }
+ if len(resp.Issues) != 0 || resp.Withheld == nil || resp.Withheld.Issues == 0 {
+ t.Fatalf("%s: issues %d withheld %+v, want the dangling binding's issue counted as withheld", path, len(resp.Issues), resp.Withheld)
+ }
+ }
+
+ w := resourceIssuesRequest(t, s, "/api/issues/resource/ClusterRoleBinding/_/dangling")
+ var bare []issues.Issue
+ if err := json.Unmarshal(w.Body.Bytes(), &bare); err != nil || len(bare) != 0 {
+ t.Fatalf("bare path: %v, %d issues", err, len(bare))
+ }
+}
diff --git a/internal/server/jobset_logs.go b/internal/server/jobset_logs.go
index 295a460795..0f07b6eecb 100644
--- a/internal/server/jobset_logs.go
+++ b/internal/server/jobset_logs.go
@@ -7,6 +7,8 @@ import (
"github.com/go-chi/chi/v5"
corev1 "k8s.io/api/core/v1"
+
+ "github.com/skyhook-io/radar/internal/podlogs"
)
func (s *Server) handleJobSetLogs(w http.ResponseWriter, r *http.Request) {
@@ -37,7 +39,7 @@ func (s *Server) handleJobSetLogs(w http.ResponseWriter, r *http.Request) {
s.writeError(w, http.StatusServiceUnavailable, "cluster client unavailable")
return
}
- snapshot := collectLogsFromPods(r.Context(), client, namespace, pods, r.URL.Query().Get("container"), parseTailLines(r.URL.Query().Get("tailLines"), 100), parseSinceSeconds(r.URL.Query().Get("sinceSeconds")), true)
+ snapshot := podlogs.CollectPods(r.Context(), client, namespace, pods, r.URL.Query().Get("container"), podlogs.ParseTailLines(r.URL.Query().Get("tailLines"), 100), podlogs.ParseSinceSeconds(r.URL.Query().Get("sinceSeconds")), true)
shownPods := []*corev1.Pod{}
shownLabels := map[string]string{}
for _, pod := range pods {
@@ -49,7 +51,7 @@ func (s *Server) handleJobSetLogs(w http.ResponseWriter, r *http.Request) {
for i := range snapshot.Logs {
snapshot.Logs[i].SourceLabel = labels[snapshot.Logs[i].Pod]
}
- sortLogsByTimestamp(snapshot.Logs)
+ podlogs.Sort(snapshot.Logs)
s.writeJSON(w, map[string]any{
"uid": root.GetUID(), "pods": buildPodInfos(shownPods), "logs": snapshot.Logs,
"notice": snapshot.Notice, "sourceLabels": shownLabels, "capturedAt": time.Now().UTC().Format(time.RFC3339),
diff --git a/internal/server/jobset_resources.go b/internal/server/jobset_resources.go
index 440eb0e56e..4a846e2ba5 100644
--- a/internal/server/jobset_resources.go
+++ b/internal/server/jobset_resources.go
@@ -10,8 +10,6 @@ import (
"time"
"github.com/go-chi/chi/v5"
- "github.com/skyhook-io/radar/internal/k8s"
- "github.com/skyhook-io/radar/pkg/k8score"
batchv1 "k8s.io/api/batch/v1"
corev1 "k8s.io/api/core/v1"
apierrors "k8s.io/apimachinery/pkg/api/errors"
@@ -22,6 +20,10 @@ import (
"k8s.io/apimachinery/pkg/runtime"
"k8s.io/apimachinery/pkg/types"
resourcehelper "k8s.io/component-helpers/resource"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
)
type jobSetMemberQuery struct {
@@ -81,7 +83,7 @@ func (s *Server) authorizeJobSetEvidence(w http.ResponseWriter, r *http.Request,
if !s.requireConnected(w) {
return false
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return false
}
diff --git a/internal/server/jobset_resources_test.go b/internal/server/jobset_resources_test.go
index 154cfc0147..e9deb91a47 100644
--- a/internal/server/jobset_resources_test.go
+++ b/internal/server/jobset_resources_test.go
@@ -10,7 +10,6 @@ import (
"testing"
"time"
- "github.com/skyhook-io/radar/pkg/k8score"
batchv1 "k8s.io/api/batch/v1"
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/api/resource"
@@ -18,6 +17,9 @@ import (
"k8s.io/apimachinery/pkg/types"
"k8s.io/client-go/kubernetes"
"k8s.io/client-go/rest"
+
+ "github.com/skyhook-io/radar/internal/podlogs"
+ "github.com/skyhook-io/radar/pkg/k8score"
)
func comparisonJob(name, role string) *batchv1.Job {
@@ -176,11 +178,11 @@ func TestSnapshotBoundsAndPartialFailures(t *testing.T) {
pending := comparisonPod(comparisonJob("pending", "workers"), time.Now())
pending.Status.ContainerStatuses = nil
pods = append(pods, pending)
- got := collectLogsFromPods(context.Background(), client, "training", pods, "", 1000, nil, true)
+ got := podlogs.CollectPods(context.Background(), client, "training", pods, "", 1000, nil, true)
if count.Load() != 40 || peak.Load() > 8 || len(got.SourcePods) != 40 {
t.Fatalf("bounds: calls=%d peak=%d pods=%d", count.Load(), peak.Load(), len(got.SourcePods))
}
- for _, text := range []string{"40 of 45", "64 KiB", "1 sources could not be read"} {
+ for _, text := range []string{"40 of 45", "64 KiB", "1 source could not be read"} {
if !strings.Contains(got.Notice, text) {
t.Fatalf("notice %q missing %q", got.Notice, text)
}
@@ -189,20 +191,20 @@ func TestSnapshotBoundsAndPartialFailures(t *testing.T) {
t.Fatal("partial success lost")
}
count.Store(0)
- unbounded := collectLogsFromPods(context.Background(), client, "training", pods, "", 2000, nil, false)
+ unbounded := podlogs.CollectPods(context.Background(), client, "training", pods, "", 2000, nil, false)
if count.Load() != 46 || strings.Contains(unbounded.Notice, "64 KiB") || strings.Contains(unbounded.Notice, "Showing") {
t.Fatalf("existing snapshot route was capped: calls=%d notice=%s", count.Load(), unbounded.Notice)
}
ctx, cancel := context.WithCancel(context.Background())
cancel()
- if got := collectLogsFromPods(ctx, client, "training", pods, "", 100, nil, true); got.Notice == "" {
+ if got := podlogs.CollectPods(ctx, client, "training", pods, "", 100, nil, true); got.Notice == "" {
t.Fatal("cancellation hidden")
}
}
func TestSnapshotSortsFractionalTimestamps(t *testing.T) {
- logs := []workloadLogEntry{{Timestamp: "2026-09-22T00:00:00.1Z"}, {Timestamp: "2026-09-22T00:00:00Z"}, {Timestamp: "2026-09-22T00:00:00.01Z"}}
- sortLogsByTimestamp(logs)
+ logs := []podlogs.Entry{{Timestamp: "2026-09-22T00:00:00.1Z"}, {Timestamp: "2026-09-22T00:00:00Z"}, {Timestamp: "2026-09-22T00:00:00.01Z"}}
+ podlogs.Sort(logs)
if logs[0].Timestamp != "2026-09-22T00:00:00Z" || logs[2].Timestamp != "2026-09-22T00:00:00.1Z" {
t.Fatalf("order: %+v", logs)
}
diff --git a/internal/server/kind_access.go b/internal/server/kind_access.go
new file mode 100644
index 0000000000..34ee3f7cfb
--- /dev/null
+++ b/internal/server/kind_access.go
@@ -0,0 +1,317 @@
+package server
+
+import (
+ "context"
+ "errors"
+ "log"
+ "net/http"
+ "slices"
+ "sort"
+ "time"
+
+ apierrors "k8s.io/apimachinery/pkg/api/errors"
+ "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/pkg/k8score"
+)
+
+// listScope resolves where the caller may list one namespaced Resource: nil
+// allowed means the whole request scope.
+//
+// denied names namespaces only when the candidate set came from the caller —
+// their view filter or their RBAC-allowed list. When the scope is "all" the
+// candidates are every namespace in Radar's cache, and naming the denied ones
+// would disclose namespaces the caller was never shown; partial then carries
+// the fact without the names.
+func (s *Server) listScope(r *http.Request, namespaces []string, group, resource string) (allowed, denied []string, partial, any bool) {
+ return s.listScopeWithCandidates(r, namespaces, group, resource, allNamespaceNames)
+}
+
+func (s *Server) listScopeWithCandidates(r *http.Request, namespaces []string, group, resource string, knownNamespaces func() []string) (allowed, denied []string, partial, any bool) {
+ if integration.NoNamespaceAccess(namespaces) {
+ return []string{}, nil, false, false
+ }
+ if s.canRead(r, group, resource, "", "list") {
+ return namespaces, nil, false, true
+ }
+ candidates := namespaces
+ if candidates == nil {
+ candidates = knownNamespaces()
+ }
+ if len(candidates) == 0 {
+ return []string{}, nil, false, false
+ }
+ allowed = s.filterNamespacesByCanRead(r, group, resource, "list", candidates)
+ partial = len(allowed) < len(candidates)
+ if namespaces != nil {
+ for _, ns := range candidates {
+ if !slices.Contains(allowed, ns) {
+ denied = append(denied, ns)
+ }
+ }
+ sort.Strings(denied)
+ }
+ return allowed, denied, partial, len(allowed) > 0
+}
+
+// typedKindScope resolves where the caller may list a typed kind and which of
+// those namespaces Radar's informer actually holds. The informer may itself be
+// namespace-scoped when Radar's own identity cannot list the kind
+// cluster-wide; what it does not hold is unread, not empty, and not denied.
+// read is nil for "every namespace".
+func (s *Server) typedKindScope(r *http.Request, cache integration.InformerScope, namespaces []string, group, resource string) (acc integration.KindAccess, read []string) {
+ return s.typedKindScopeWithCandidates(r, cache, namespaces, group, resource, allNamespaceNames)
+}
+
+func (s *Server) typedKindScopeWithCandidates(r *http.Request, cache integration.InformerScope, namespaces []string, group, resource string, candidates func() []string) (acc integration.KindAccess, read []string) {
+ allowed, denied, partial, ok := s.listScopeWithCandidates(r, namespaces, group, resource, candidates)
+ if !ok {
+ return integration.KindAccess{State: integration.KindCoverageDenied}, []string{}
+ }
+ within := integration.NamespacesWithinCache(cache, resource, allowed)
+ var uncached []string
+ if allowed != nil {
+ for _, ns := range allowed {
+ if slices.Contains(within.Namespaces, ns) {
+ continue
+ }
+ partial = true
+ if namespaces != nil {
+ uncached = append(uncached, ns)
+ }
+ }
+ sort.Strings(uncached)
+ }
+ if within.Unavailable {
+ return integration.KindAccess{State: integration.KindCoverageUncached, Denied: denied, Uncached: uncached}, []string{}
+ }
+ acc = integration.AccessFromScope(within.Namespaces, partial || within.Partial)
+ acc.Denied, acc.Uncached = denied, uncached
+ return acc, within.Namespaces
+}
+
+// readWorkspaceKind authorizes and lists one kind, keeping only objects of
+// groups. A namespaced kind falls back to the namespaces the caller may list
+// when a cluster-wide list is denied; a cluster-scoped kind needs the
+// cluster-scope list. Radar's own watch scope counts as well, as for typed
+// kinds: where its identity watches the kind namespace by namespace, what it
+// does not hold is unread (uncached), never empty and never syncing forever.
+func (s *Server) readWorkspaceKind(r *http.Request, cache *k8s.ResourceCache, k integration.WorkspaceKind, namespaces, groups []string, budget *syncBudget) (integration.KindAccess, []*unstructured.Unstructured) {
+ ctx := r.Context()
+ if k.ClusterScoped {
+ if !s.canRead(r, k.Group, k.Resource, "", "list") {
+ return integration.KindAccess{State: integration.KindCoverageDenied}, nil
+ }
+ list, err := listKindInGroups(ctx, cache, k, nil, groups, budget)
+ if err != nil {
+ return kindAccessFromListError(k, err), nil
+ }
+ return integration.KindAccess{State: integration.KindCoverageFull, All: true}, list
+ }
+ allowed, denied, partial, ok := s.listScopeWithCandidates(r, namespaces, k.Group, k.Resource, func() []string { return namespaceNamesInCache(cache) })
+ if !ok {
+ return integration.KindAccess{State: integration.KindCoverageDenied}, nil
+ }
+ acc := integration.AccessFromScope(allowed, partial)
+ acc.Denied = denied
+ if allowed == nil {
+ return readWorkspaceKindEverywhere(ctx, cache, k, groups, acc, budget)
+ }
+ return readWorkspaceKindIn(ctx, cache, k, groups, acc, allowed, namespaces != nil, budget)
+}
+
+// syncBudget bounds how long one workspace read waits on the dynamic cache,
+// across every kind and namespace it reads: starting a watch, which probes the
+// apiserver, as well as waiting for it to sync. Once it is spent, or the
+// request is cancelled, a namespace that has not synced is unread.
+type syncBudget struct {
+ ctx context.Context
+ deadline time.Time
+ bound bool
+ discovery *k8s.ResourceDiscovery
+ dynamicCache *k8s.DynamicResourceCache
+}
+
+func newSyncBudget(ctx context.Context) *syncBudget {
+ return &syncBudget{ctx: ctx, deadline: time.Now().Add(dynamicSyncWait)}
+}
+
+func (b *syncBudget) dynamicDependencies() (*k8s.ResourceDiscovery, *k8s.DynamicResourceCache) {
+ if b != nil && b.bound {
+ return b.discovery, b.dynamicCache
+ }
+ return k8s.GetResourceDiscovery(), k8s.GetDynamicResourceCache()
+}
+
+// listBlocking is ListBlocking within the budget; a nil budget waits
+// dynamicSyncWait for the sync alone. A watch still starting when the budget
+// runs out carries on in the background, so a later read finds it.
+func (b *syncBudget) listBlocking(dc *k8s.DynamicResourceCache, gvr schema.GroupVersionResource, namespace string) ([]*unstructured.Unstructured, error) {
+ if b == nil {
+ return dc.ListBlocking(gvr, namespace, dynamicSyncWait)
+ }
+ if dc.IsNamespaceSynced(gvr, namespace) {
+ return dc.ListBlocking(gvr, namespace, 0)
+ }
+ if b.ctx.Err() != nil {
+ return nil, integration.ErrDynamicNotSynced
+ }
+ wait := max(0, time.Until(b.deadline))
+ type result struct {
+ items []*unstructured.Unstructured
+ err error
+ }
+ done := make(chan result, 1)
+ go func() {
+ items, err := dc.ListBlocking(gvr, namespace, wait)
+ done <- result{items, err}
+ }()
+ timer := time.NewTimer(wait)
+ defer timer.Stop()
+ select {
+ case r := <-done:
+ return r.items, r.err
+ case <-timer.C:
+ case <-b.ctx.Done():
+ }
+ return nil, integration.ErrDynamicNotSynced
+}
+
+// readWorkspaceKindEverywhere reads a kind in every namespace. When Radar's
+// identity watches it only namespace by namespace, a cluster-wide read never
+// syncs (or is refused outright); each watched namespace that has synced is
+// read instead and the rest stays unread, unnamed (the caller did not name the
+// scope).
+func readWorkspaceKindEverywhere(ctx context.Context, cache *k8s.ResourceCache, k integration.WorkspaceKind, groups []string, acc integration.KindAccess, budget *syncBudget) (integration.KindAccess, []*unstructured.Unstructured) {
+ list, err := listKindInGroups(ctx, cache, k, nil, groups, budget)
+ if err == nil {
+ return acc, list
+ }
+ radarDenied := apierrors.IsForbidden(err) || apierrors.IsUnauthorized(err)
+ if !errors.Is(err, integration.ErrDynamicNotSynced) && !radarDenied {
+ return kindAccessFromListError(k, err), nil
+ }
+ dc, gvr, watched, ok := dynamicNamespaceWatches(k, budget)
+ if !ok {
+ if radarDenied {
+ return integration.KindAccess{State: integration.KindCoverageUncached}, nil
+ }
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ read := map[string]bool{}
+ var out []*unstructured.Unstructured
+ for _, ns := range watched {
+ if !dc.IsNamespaceSynced(gvr, ns) {
+ continue
+ }
+ items, err := listDynamicSyncedWithin(ctx, cache, k.Kind, k.Group, ns, budget)
+ if err != nil {
+ continue
+ }
+ read[ns] = true
+ out = append(out, integration.KeepGroups(items, groups)...)
+ }
+ if len(read) == 0 {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ return integration.KindAccess{State: integration.KindCoveragePartial, Namespaces: read, Denied: acc.Denied}, out
+}
+
+// readWorkspaceKindIn reads each allowed namespace on its own, which starts
+// that namespace's watch when Radar watches the kind namespace by namespace. A
+// namespace Radar's identity cannot watch, or that has not synced in time, is
+// unread: uncached, named only when the caller named the scope.
+func readWorkspaceKindIn(ctx context.Context, cache *k8s.ResourceCache, k integration.WorkspaceKind, groups []string, acc integration.KindAccess, allowed []string, callerScoped bool, budget *syncBudget) (integration.KindAccess, []*unstructured.Unstructured) {
+ read := map[string]bool{}
+ var out []*unstructured.Unstructured
+ var unread []string
+ syncing := false
+ for _, ns := range allowed {
+ items, err := listKindInGroups(ctx, cache, k, []string{ns}, groups, budget)
+ switch {
+ case err == nil:
+ read[ns] = true
+ out = append(out, items...)
+ case errors.Is(err, integration.ErrDynamicNotSynced):
+ syncing = true
+ unread = append(unread, ns)
+ case apierrors.IsForbidden(err) || apierrors.IsUnauthorized(err):
+ unread = append(unread, ns)
+ default:
+ return kindAccessFromListError(k, err), nil
+ }
+ }
+ var named []string
+ if callerScoped {
+ named = unread
+ sort.Strings(named)
+ }
+ if len(read) == 0 {
+ if syncing {
+ return integration.KindAccess{State: integration.KindCoverageSyncing}, nil
+ }
+ return integration.KindAccess{State: integration.KindCoverageUncached, Denied: acc.Denied, Uncached: named}, nil
+ }
+ acc.All, acc.Namespaces = false, read
+ if len(unread) > 0 {
+ acc.State, acc.Uncached = integration.KindCoveragePartial, named
+ }
+ return acc, out
+}
+
+func kindAccessFromListError(k integration.WorkspaceKind, err error) integration.KindAccess {
+ switch {
+ case errors.Is(err, k8s.ErrUnknownDynamicKind):
+ return integration.KindAccess{State: integration.KindCoverageNotInstalled}
+ case errors.Is(err, integration.ErrDynamicNotSynced):
+ return integration.KindAccess{State: integration.KindCoverageSyncing}
+ case apierrors.IsForbidden(err) || apierrors.IsUnauthorized(err):
+ // Radar's own identity may not watch it; the caller's access was checked first.
+ return integration.KindAccess{State: integration.KindCoverageUncached}
+ default:
+ log.Printf("[workspace] Failed to list %s.%s: %v", k.Kind, k.Group, err)
+ return integration.KindAccess{State: integration.KindCoverageError}
+ }
+}
+
+// dynamicNamespaceWatches returns the namespaces Radar's identity watches k in
+// when it holds no cluster-wide informer for it.
+func dynamicNamespaceWatches(k integration.WorkspaceKind, budget *syncBudget) (*k8s.DynamicResourceCache, schema.GroupVersionResource, []string, bool) {
+ discovery, dc := budget.dynamicDependencies()
+ if discovery == nil || dc == nil {
+ return nil, schema.GroupVersionResource{}, nil, false
+ }
+ gvr, found := discovery.GetGVRWithGroup(k.Kind, k.Group)
+ if !found {
+ return nil, schema.GroupVersionResource{}, nil, false
+ }
+ obs := dc.Observation(gvr)
+ if obs.Scope != k8score.DynamicObservationScopeExplicitNamespaces || len(obs.Namespaces) == 0 {
+ return nil, schema.GroupVersionResource{}, nil, false
+ }
+ return dc, gvr, obs.Namespaces, true
+}
+
+// listKindInGroups lists k in namespaces (nil = all), keeping only objects of
+// groups, waiting for sync no longer than budget allows.
+func listKindInGroups(ctx context.Context, cache *k8s.ResourceCache, k integration.WorkspaceKind, namespaces, groups []string, budget *syncBudget) ([]*unstructured.Unstructured, error) {
+ if namespaces == nil {
+ list, err := listDynamicSyncedWithin(ctx, cache, k.Kind, k.Group, "", budget)
+ if err != nil {
+ return nil, err
+ }
+ return integration.KeepGroups(list, groups), nil
+ }
+ var out []*unstructured.Unstructured
+ for _, ns := range namespaces {
+ list, err := listDynamicSyncedWithin(ctx, cache, k.Kind, k.Group, ns, budget)
+ if err != nil {
+ return nil, err
+ }
+ out = append(out, integration.KeepGroups(list, groups)...)
+ }
+ return out, nil
+}
diff --git a/internal/server/kind_access_test.go b/internal/server/kind_access_test.go
new file mode 100644
index 0000000000..8e033ca97f
--- /dev/null
+++ b/internal/server/kind_access_test.go
@@ -0,0 +1,106 @@
+package server
+
+import (
+ "net/http"
+ "net/http/httptest"
+ "slices"
+ "testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
+)
+
+func kindScopeRequest(env *authTestEnv, user string, perms *auth.UserPermissions) *http.Request {
+ env.srv.permCache.Set(user, nil, perms)
+ r := httptest.NewRequest(http.MethodGet, "/api/cnpg/workspace", nil)
+ return r.WithContext(auth.ContextWithUser(r.Context(), &auth.User{Username: user}))
+}
+
+// A namespace the caller may list but Radar's informer does not hold was not
+// read, and it is not a denial: the caller has the permission.
+func TestTypedKindScope_UncachedIsNotDenied(t *testing.T) {
+ env := newAuthTestServer(t)
+ cache := capacityTestInformerScope{namespaces: map[string][]string{"pods": {"a"}}}
+
+ perms := &auth.UserPermissions{}
+ allow(perms, "", "pods", "", false)
+ allow(perms, "", "pods", "a", true)
+ allow(perms, "", "pods", "b", true)
+ allow(perms, "", "pods", "c", false)
+ r := kindScopeRequest(env, "alice", perms)
+
+ t.Run("readable but uncached", func(t *testing.T) {
+ acc, read := env.srv.typedKindScope(r, cache, []string{"a", "b"}, "", "pods")
+ cov := acc.Coverage()
+ if cov.State != integration.KindCoveragePartial || len(cov.DeniedNamespaces) != 0 || !slices.Equal(cov.UncachedNamespaces, []string{"b"}) || !slices.Equal(cov.AllowedNamespaces, []string{"a"}) {
+ t.Errorf("coverage = %+v, want partial, uncached [b], allowed [a], nothing denied", cov)
+ }
+ if !slices.Equal(read, []string{"a"}) {
+ t.Errorf("read = %v, want [a]", read)
+ }
+ })
+
+ t.Run("wholly uncached", func(t *testing.T) {
+ acc, read := env.srv.typedKindScope(r, cache, []string{"b"}, "", "pods")
+ cov := acc.Coverage()
+ if cov.State != integration.KindCoverageUncached || len(cov.DeniedNamespaces) != 0 || !slices.Equal(cov.UncachedNamespaces, []string{"b"}) || cov.AllowedNamespaces != nil {
+ t.Errorf("coverage = %+v, want uncached naming b, nothing denied", cov)
+ }
+ if read == nil || len(read) != 0 {
+ t.Errorf("read = %#v, want an empty non-nil scope (nil reads every namespace)", read)
+ }
+ if acc.Covers("b") {
+ t.Error("an uncached namespace counts as covered")
+ }
+ })
+
+ t.Run("denied and uncached together", func(t *testing.T) {
+ acc, read := env.srv.typedKindScope(r, cache, []string{"a", "b", "c"}, "", "pods")
+ cov := acc.Coverage()
+ if cov.State != integration.KindCoveragePartial || !slices.Equal(cov.DeniedNamespaces, []string{"c"}) || !slices.Equal(cov.UncachedNamespaces, []string{"b"}) || !slices.Equal(cov.AllowedNamespaces, []string{"a"}) {
+ t.Errorf("coverage = %+v, want partial, denied [c], uncached [b], allowed [a]", cov)
+ }
+ if !slices.Equal(read, []string{"a"}) {
+ t.Errorf("read = %v, want [a]", read)
+ }
+ })
+
+ t.Run("wholly uncached with a denial", func(t *testing.T) {
+ acc, _ := env.srv.typedKindScope(r, cache, []string{"b", "c"}, "", "pods")
+ cov := acc.Coverage()
+ if cov.State != integration.KindCoverageUncached || !slices.Equal(cov.DeniedNamespaces, []string{"c"}) || !slices.Equal(cov.UncachedNamespaces, []string{"b"}) {
+ t.Errorf("coverage = %+v, want uncached, denied [c], uncached [b]", cov)
+ }
+ })
+}
+
+// Uncached namespaces follow the same disclosure rule as denied ones: named
+// only when the caller supplied the candidates.
+func TestTypedKindScope_UncachedNamesFollowTheDisclosureRule(t *testing.T) {
+ env := newAuthTestServer(t)
+ perms := &auth.UserPermissions{}
+ allow(perms, "", "pods", "", true)
+ r := kindScopeRequest(env, "wide", perms)
+
+ limited := capacityTestInformerScope{namespaces: map[string][]string{"pods": {"a"}}}
+ acc, read := env.srv.typedKindScope(r, limited, nil, "", "pods")
+ cov := acc.Coverage()
+ if cov.State != integration.KindCoveragePartial || len(cov.UncachedNamespaces) != 0 || !slices.Equal(cov.AllowedNamespaces, []string{"a"}) {
+ t.Errorf("unfiltered: coverage = %+v, want partial with allowed [a] and no uncached names", cov)
+ }
+ if !slices.Equal(read, []string{"a"}) {
+ t.Errorf("unfiltered: read = %v, want [a]", read)
+ }
+
+ empty := capacityTestInformerScope{namespaces: map[string][]string{}}
+ acc, _ = env.srv.typedKindScope(r, empty, nil, "", "pods")
+ if cov := acc.Coverage(); cov.State != integration.KindCoverageUncached || len(cov.UncachedNamespaces) != 0 {
+ t.Errorf("unfiltered, nothing cached: coverage = %+v, want uncached without names", cov)
+ }
+
+ clusterWide := capacityTestInformerScope{clusterWide: map[string]bool{"pods": true}}
+ acc, read = env.srv.typedKindScope(r, clusterWide, nil, "", "pods")
+ if cov := acc.Coverage(); cov.State != integration.KindCoverageFull || read != nil {
+ t.Errorf("cluster-wide informer: coverage = %+v read = %v, want full over every namespace", cov, read)
+ }
+}
diff --git a/internal/server/kueue_handlers.go b/internal/server/kueue_handlers.go
index bddaf8e385..c36cee665b 100644
--- a/internal/server/kueue_handlers.go
+++ b/internal/server/kueue_handlers.go
@@ -11,6 +11,7 @@ import (
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
"k8s.io/apimachinery/pkg/runtime"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/k8score"
"github.com/skyhook-io/radar/pkg/resourcecontext"
@@ -56,7 +57,7 @@ func (s *Server) handleKueueAdmission(w http.ResponseWriter, r *http.Request) {
s.writeError(w, http.StatusBadRequest, "Kueue admission lookup supports batch Jobs, jobset.x-k8s.io JobSets and ray.io RayJobs")
return
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) || !s.canRead(r, group, kind, namespace, "get") {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) || !s.canRead(r, group, kind, namespace, "get") {
s.writeError(w, http.StatusForbidden, "no access to this workload")
return
}
@@ -237,7 +238,7 @@ func (s *Server) handleKueueProvisioning(w http.ResponseWriter, r *http.Request)
return
}
namespace, name := chi.URLParam(r, "namespace"), chi.URLParam(r, "name")
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) || !s.canRead(r, kueueGroup, "workloads", namespace, "get") {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) || !s.canRead(r, kueueGroup, "workloads", namespace, "get") {
s.writeError(w, http.StatusForbidden, "no access to this Kueue Workload")
return
}
diff --git a/internal/server/kueue_handlers_test.go b/internal/server/kueue_handlers_test.go
index be18f30dc4..94c77b50bb 100644
--- a/internal/server/kueue_handlers_test.go
+++ b/internal/server/kueue_handlers_test.go
@@ -4,10 +4,6 @@ import (
"context"
"encoding/json"
"fmt"
- "github.com/skyhook-io/radar/internal/k8s"
- "k8s.io/apimachinery/pkg/runtime"
- "k8s.io/apimachinery/pkg/runtime/schema"
- dynamicfake "k8s.io/client-go/dynamic/fake"
"net/http"
"testing"
"time"
@@ -15,11 +11,15 @@ import (
batchv1 "k8s.io/api/batch/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ "k8s.io/apimachinery/pkg/runtime"
+ "k8s.io/apimachinery/pkg/runtime/schema"
"k8s.io/apimachinery/pkg/types"
+ dynamicfake "k8s.io/client-go/dynamic/fake"
"k8s.io/client-go/kubernetes/fake"
k8stesting "k8s.io/client-go/testing"
"github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/k8score"
)
diff --git a/internal/server/logs.go b/internal/server/logs.go
index 78b7b8e566..b732f5e59c 100644
--- a/internal/server/logs.go
+++ b/internal/server/logs.go
@@ -12,7 +12,9 @@ import (
"github.com/go-chi/chi/v5"
"k8s.io/client-go/kubernetes"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
"github.com/skyhook-io/radar/pkg/k8score"
)
@@ -34,7 +36,7 @@ func (s *Server) handlePodLogs(w http.ResponseWriter, r *http.Request) {
sinceSecondsStr := r.URL.Query().Get("sinceSeconds")
// Check namespace access for authenticated users
- if allowed := s.getUserNamespaces(r, []string{namespace}); noNamespaceAccess(allowed) {
+ if allowed := s.getUserNamespaces(r, []string{namespace}); integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -42,8 +44,8 @@ func (s *Server) handlePodLogs(w http.ResponseWriter, r *http.Request) {
return
}
- tailLines := parseTailLines(tailLinesStr, 500)
- sinceSeconds := parseSinceSeconds(sinceSecondsStr)
+ tailLines := podlogs.ParseTailLines(tailLinesStr, 500)
+ sinceSeconds := podlogs.ParseSinceSeconds(sinceSecondsStr)
client := s.getClientForRequest(r)
if client == nil {
@@ -115,7 +117,7 @@ func (s *Server) handlePodLogsStream(w http.ResponseWriter, r *http.Request) {
tailLinesStr := r.URL.Query().Get("tailLines")
// Check namespace access for authenticated users
- if allowed := s.getUserNamespaces(r, []string{namespace}); noNamespaceAccess(allowed) {
+ if allowed := s.getUserNamespaces(r, []string{namespace}); integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -125,8 +127,8 @@ func (s *Server) handlePodLogsStream(w http.ResponseWriter, r *http.Request) {
sinceStr := r.URL.Query().Get("sinceSeconds")
- tailLines := parseTailLines(tailLinesStr, 100)
- sinceSeconds := parseSinceSeconds(sinceStr)
+ tailLines := podlogs.ParseTailLines(tailLinesStr, 100)
+ sinceSeconds := podlogs.ParseSinceSeconds(sinceStr)
// Set SSE headers
w.Header().Set("Content-Type", "text/event-stream")
@@ -211,7 +213,7 @@ func (s *Server) handlePodLogsStream(w http.ResponseWriter, r *http.Request) {
}
// Parse timestamp and content
- timestamp, content := parseLogLine(line)
+ timestamp, content := podlogs.ParseLine(line)
sendSSEEvent(w, flusher, "log", map[string]string{
"timestamp": timestamp,
@@ -248,19 +250,6 @@ func (s *Server) fetchContainerLogs(ctx context.Context, client kubernetes.Inter
return string(content), nil
}
-// parseLogLine extracts timestamp from a log line (format: 2024-01-20T10:30:00.123456789Z content)
-func parseLogLine(line string) (timestamp, content string) {
- // K8s timestamps are in RFC3339Nano format at the start of the line
- if len(line) > 30 && line[4] == '-' && line[7] == '-' && line[10] == 'T' {
- // Find the space after timestamp
- spaceIdx := strings.Index(line, " ")
- if spaceIdx > 20 && spaceIdx < 40 {
- return line[:spaceIdx], line[spaceIdx+1:]
- }
- }
- return "", line
-}
-
// sendSSEEvent sends an SSE event
func sendSSEEvent(w http.ResponseWriter, flusher http.Flusher, event string, data any) {
jsonData, _ := json.Marshal(data)
diff --git a/internal/server/neighborhood_handler.go b/internal/server/neighborhood_handler.go
index e6a458c323..4eebf8507b 100644
--- a/internal/server/neighborhood_handler.go
+++ b/internal/server/neighborhood_handler.go
@@ -5,12 +5,12 @@ import (
"net/http"
"strconv"
- "github.com/skyhook-io/radar/pkg/resourceid"
-
"github.com/go-chi/chi/v5"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/resourcecontext"
+ "github.com/skyhook-io/radar/pkg/resourceid"
"github.com/skyhook-io/radar/pkg/topology"
)
@@ -85,7 +85,7 @@ func (s *Server) handleAINeighborhood(w http.ResponseWriter, r *http.Request) {
return
}
allowed := s.getUserNamespaces(r, []string{namespace})
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -267,7 +267,7 @@ func (s *Server) canReadNeighborhoodNode(r *http.Request, n *topology.Node) bool
if n != nil && n.Data != nil {
if ns, ok := n.Data["namespace"].(string); ok && ns != "" {
allowed := s.getUserNamespaces(r, []string{ns})
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
return false
}
}
diff --git a/internal/server/node_metrics.go b/internal/server/node_metrics.go
index 8c11771ca3..32f23f5c01 100644
--- a/internal/server/node_metrics.go
+++ b/internal/server/node_metrics.go
@@ -6,11 +6,10 @@ import (
"log"
"time"
+ corev1 "k8s.io/api/core/v1"
+
"github.com/skyhook-io/radar/internal/capacity"
"github.com/skyhook-io/radar/internal/k8s"
- corev1 "k8s.io/api/core/v1"
- "k8s.io/apimachinery/pkg/labels"
- v1listers "k8s.io/client-go/listers/core/v1"
)
// Shared metrics-server plumbing — consumed by both /api/dashboard and
@@ -72,27 +71,6 @@ func parseCPUToMillis(s string) int64 { return k8s.ParseCPUToMillis(s) }
// parseMemoryToBytes delegates to k8s.ParseMemoryToBytes.
func parseMemoryToBytes(s string) int64 { return k8s.ParseMemoryToBytes(s) }
-// listPodsScoped lists pods either cluster-wide (namespaces nil) or across
-// the caller's allowed namespaces — the scoping shape every metrics/vitals
-// consumer shares.
-func listPodsScoped(podLister v1listers.PodLister, namespaces []string) []*corev1.Pod {
- if podLister == nil {
- return nil
- }
- // Sentinel contract (parseNamespacesForUser): nil = all namespaces;
- // non-nil EMPTY = no namespace access — zero pods, never cluster-wide.
- if namespaces == nil {
- pods, _ := podLister.List(labels.Everything())
- return pods
- }
- var pods []*corev1.Pod
- for _, ns := range namespaces {
- items, _ := podLister.Pods(ns).List(labels.Everything())
- pods = append(pods, items...)
- }
- return pods
-}
-
// capacityRequests carries the informer-derived halves of the capacity
// picture: node allocatable capacity plus scheduled-pod requests (completed pods
// excluded). Usage is the metrics-server probe's job (fetchNodeUsage).
diff --git a/internal/server/node_metrics_test.go b/internal/server/node_metrics_test.go
index 03aa3ab0cd..46e25303ef 100644
--- a/internal/server/node_metrics_test.go
+++ b/internal/server/node_metrics_test.go
@@ -3,12 +3,13 @@ package server
import (
"testing"
- "k8s.io/apimachinery/pkg/api/resource"
-
corev1 "k8s.io/api/core/v1"
+ "k8s.io/apimachinery/pkg/api/resource"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/client-go/informers"
"k8s.io/client-go/kubernetes/fake"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
)
func TestListPodsScoped_SentinelContract(t *testing.T) {
@@ -25,16 +26,16 @@ func TestListPodsScoped_SentinelContract(t *testing.T) {
f.WaitForCacheSync(stop)
lister := informer.Lister()
- if got := listPodsScoped(lister, nil); len(got) != 2 {
+ if got := integration.ListPodsScoped(lister, nil); len(got) != 2 {
t.Errorf("nil scope = %d pods, want 2 (cluster-wide)", len(got))
}
- if got := listPodsScoped(lister, []string{}); len(got) != 0 {
+ if got := integration.ListPodsScoped(lister, []string{}); len(got) != 0 {
t.Errorf("empty scope = %d pods, want 0 (no namespace access)", len(got))
}
- if got := listPodsScoped(lister, []string{"ns1"}); len(got) != 1 {
+ if got := integration.ListPodsScoped(lister, []string{"ns1"}); len(got) != 1 {
t.Errorf("ns1 scope = %d pods, want 1", len(got))
}
- if got := listPodsScoped(nil, nil); got != nil {
+ if got := integration.ListPodsScoped(nil, nil); got != nil {
t.Errorf("nil lister must yield nil")
}
}
diff --git a/internal/server/policy_handlers.go b/internal/server/policy_handlers.go
index 850b4fe8eb..4217c67013 100644
--- a/internal/server/policy_handlers.go
+++ b/internal/server/policy_handlers.go
@@ -9,9 +9,10 @@ import (
"strconv"
"time"
+ "github.com/go-chi/chi/v5"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
- "github.com/go-chi/chi/v5"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/policyreports"
"github.com/skyhook-io/radar/pkg/resourceid"
@@ -613,16 +614,27 @@ type PolicyQueuedResponse struct {
// Falls back to the plain read when the machinery is missing, leaving the
// caller's not-installed path to answer.
func listDynamicSynced(ctx context.Context, cache *k8s.ResourceCache, kind, group, namespace string) ([]*unstructured.Unstructured, error) {
- discovery := k8s.GetResourceDiscovery()
- dynamicCache := k8s.GetDynamicResourceCache()
+ return listDynamicSyncedWithin(ctx, cache, kind, group, namespace, nil)
+}
+
+// listDynamicSyncedWithin is listDynamicSynced bounded by budget (nil: wait
+// dynamicSyncWait for the sync).
+func listDynamicSyncedWithin(ctx context.Context, cache *k8s.ResourceCache, kind, group, namespace string, budget *syncBudget) ([]*unstructured.Unstructured, error) {
+ discovery, dynamicCache := budget.dynamicDependencies()
if discovery == nil || dynamicCache == nil {
+ if budget != nil && budget.bound {
+ return nil, k8s.ErrDynamicNotReady
+ }
return cache.ListDynamicWithGroup(ctx, kind, namespace, group)
}
gvr, found := discovery.GetGVRWithGroup(kind, group)
if !found {
+ if budget != nil && budget.bound {
+ return nil, k8s.ErrUnknownDynamicKind
+ }
return cache.ListDynamicWithGroup(ctx, kind, namespace, group)
}
- items, err := dynamicCache.ListBlocking(gvr, namespace, dynamicSyncWait)
+ items, err := budget.listBlocking(dynamicCache, gvr, namespace)
if err != nil {
return nil, err
}
@@ -632,15 +644,11 @@ func listDynamicSynced(ctx context.Context, cache *k8s.ResourceCache, kind, grou
// already started the informer, so asking now cannot deadlock the way a gate
// before the read did.
if !dynamicCache.IsNamespaceSynced(gvr, namespace) {
- return nil, errDynamicNotSynced
+ return nil, integration.ErrDynamicNotSynced
}
return items, nil
}
-// errDynamicNotSynced means the cache could not answer for the scope in time.
-// Distinct from an absent CRD: nothing was established either way.
-var errDynamicNotSynced = errors.New("resource cache not synced")
-
// Long enough for a cold informer on a healthy apiserver, short enough that a
// drawer section does not hang on one that will not sync.
const dynamicSyncWait = 3 * time.Second
@@ -688,7 +696,7 @@ func (s *Server) handlePolicyQueued(w http.ResponseWriter, r *http.Request) {
// No background controller on this cluster. Genuinely nothing queued.
s.writeJSON(w, PolicyQueuedResponse{})
return
- case errors.Is(err, errDynamicNotSynced):
+ case errors.Is(err, integration.ErrDynamicNotSynced):
s.writeError(w, http.StatusServiceUnavailable, "queued work is still loading")
return
default:
diff --git a/internal/server/raycluster_pods.go b/internal/server/raycluster_pods.go
index cba2b65ccf..1fdc2fd6b3 100644
--- a/internal/server/raycluster_pods.go
+++ b/internal/server/raycluster_pods.go
@@ -6,10 +6,12 @@ import (
"net/http"
"github.com/go-chi/chi/v5"
- "github.com/skyhook-io/radar/internal/k8s"
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/runtime/schema"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
+ "github.com/skyhook-io/radar/internal/k8s"
)
func (s *Server) handleRayClusterPods(w http.ResponseWriter, r *http.Request) {
@@ -29,7 +31,7 @@ func (s *Server) handleRayClusterPods(w http.ResponseWriter, r *http.Request) {
s.writeError(w, http.StatusBadRequest, err.Error())
return
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{ns})) || !s.canRead(r, "ray.io", "rayclusters", ns, "get") || !s.canRead(r, "", "pods", ns, "list") {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{ns})) || !s.canRead(r, "ray.io", "rayclusters", ns, "get") || !s.canRead(r, "", "pods", ns, "list") {
s.writeError(w, http.StatusForbidden, "Reading RayCluster Pods requires get rayclusters and list pods in this namespace")
return
}
diff --git a/internal/server/raycluster_pods_test.go b/internal/server/raycluster_pods_test.go
index f17bf07882..6a8d395cc8 100644
--- a/internal/server/raycluster_pods_test.go
+++ b/internal/server/raycluster_pods_test.go
@@ -9,9 +9,6 @@ import (
"time"
"github.com/go-chi/chi/v5"
- "github.com/skyhook-io/radar/internal/auth"
- "github.com/skyhook-io/radar/internal/k8s"
- "github.com/skyhook-io/radar/pkg/k8score"
corev1 "k8s.io/api/core/v1"
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
@@ -19,6 +16,11 @@ import (
dynamicfake "k8s.io/client-go/dynamic/fake"
"k8s.io/client-go/kubernetes/fake"
k8stesting "k8s.io/client-go/testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
+ "github.com/skyhook-io/radar/pkg/k8score"
)
func TestRayClusterGroupCannotIncludeHead(t *testing.T) {
@@ -94,7 +96,7 @@ func TestRayClusterPodsRequestContract(t *testing.T) {
rec = httptest.NewRecorder()
router.ServeHTTP(rec, httptest.NewRequest("GET", "/api/workloads/rayclusters/default/cluster/pods?ownerUID=current&workerGroup=workers", nil))
var result struct {
- Pods []WorkloadPodInfo
+ Pods []podlogs.PodInfo
Total int
Truncated bool
}
diff --git a/internal/server/reachability_run.go b/internal/server/reachability_run.go
index 4f408db7f5..c1ca57f744 100644
--- a/internal/server/reachability_run.go
+++ b/internal/server/reachability_run.go
@@ -11,7 +11,9 @@ import (
"strings"
"github.com/go-chi/chi/v5"
+
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/reachability"
@@ -64,7 +66,7 @@ func (s *Server) handleProbeInClusterCapability(w http.ResponseWriter, r *http.R
// Mirror the POST's namespace boundary so the capability answer matches what the
// POST will actually allow - never report "allowed" for an out-of-scope namespace.
namespaces := s.traceNamespaceCeiling(r)
- if noNamespaceAccess(namespaces) || (namespace != "" && !namespaceAllowed(namespaces, namespace)) {
+ if integration.NoNamespaceAccess(namespaces) || (namespace != "" && !namespaceAllowed(namespaces, namespace)) {
resp.Reason = fmt.Sprintf("no access to namespace %q", namespace)
s.writeJSON(w, resp)
return
@@ -112,7 +114,7 @@ func (s *Server) handleTraceInCluster(w http.ResponseWriter, r *http.Request) {
namespaces := s.traceNamespaceCeiling(r)
// Mirror handleTrace: never leak that a resource exists outside the caller's
// namespace scope - return an unknown-verdict trace instead.
- if noNamespaceAccess(namespaces) || (namespace != "" && !namespaceAllowed(namespaces, namespace)) {
+ if integration.NoNamespaceAccess(namespaces) || (namespace != "" && !namespaceAllowed(namespaces, namespace)) {
s.writeJSON(w, traceInClusterResponse{Trace: &trace.Trace{
Subject: trace.ResourceRef{Kind: kind, Namespace: namespace, Name: name},
Downstream: []trace.Hop{},
diff --git a/internal/server/resource_counts.go b/internal/server/resource_counts.go
index 241870ab64..a0f3af9504 100644
--- a/internal/server/resource_counts.go
+++ b/internal/server/resource_counts.go
@@ -6,9 +6,11 @@ import (
"net/http"
"slices"
+ "k8s.io/apimachinery/pkg/runtime/schema"
+
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/k8score"
- "k8s.io/apimachinery/pkg/runtime/schema"
)
type ResourceCountsResponse struct {
@@ -59,7 +61,7 @@ func (s *Server) handleResourceCounts(w http.ResponseWriter, r *http.Request) {
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, ResourceCountsResponse{Counts: map[string]int{}})
return
}
diff --git a/internal/server/rightsizing_scan.go b/internal/server/rightsizing_scan.go
index 9180f7b9c2..b8ff60e1e7 100644
--- a/internal/server/rightsizing_scan.go
+++ b/internal/server/rightsizing_scan.go
@@ -8,6 +8,7 @@ import (
"github.com/go-chi/chi/v5"
+ integration "github.com/skyhook-io/radar/internal/integration"
prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
)
@@ -34,7 +35,7 @@ func (s *Server) handleRightsizingScan(w http.ResponseWriter, r *http.Request) {
} else {
namespaces = s.parseNamespacesForUser(r)
}
- if noNamespaceAccess(namespaces) && (id != "" || hasExplicitNamespaceFilter(r)) {
+ if integration.NoNamespaceAccess(namespaces) && (id != "" || hasExplicitNamespaceFilter(r)) {
s.writeError(w, http.StatusForbidden, "no access to the requested namespace(s)")
return
}
@@ -83,7 +84,7 @@ func (a serverScanAuthorizer) FilterNamespaces(resource string, namespaces []str
func (s *Server) resolveRightsizingScanScope(r *http.Request, namespaces []string) prometheuspkg.RightsizingScanScope {
// No readable namespace at all is reported as every kind restricted rather
// than an empty scan, so the response carries why it found nothing.
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
scope := prometheuspkg.RightsizingScanScope{
NamespacesByKind: make(map[string][]string, len(prometheuspkg.RightsizingScanKinds)),
}
diff --git a/internal/server/rightsizing_scan_test.go b/internal/server/rightsizing_scan_test.go
index 60e27d5eaf..34b0171edd 100644
--- a/internal/server/rightsizing_scan_test.go
+++ b/internal/server/rightsizing_scan_test.go
@@ -3,13 +3,14 @@ package server
import (
"context"
"encoding/json"
- "github.com/go-chi/chi/v5"
- prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
"net/http"
"net/http/httptest"
"slices"
"testing"
+ "github.com/go-chi/chi/v5"
+
+ prometheuspkg "github.com/skyhook-io/radar/internal/prometheus"
pkgauth "github.com/skyhook-io/radar/pkg/auth"
)
diff --git a/internal/server/search_handler.go b/internal/server/search_handler.go
index 44e843125d..cb4416eb91 100644
--- a/internal/server/search_handler.go
+++ b/internal/server/search_handler.go
@@ -8,6 +8,7 @@ import (
"github.com/skyhook-io/radar/internal/auth"
"github.com/skyhook-io/radar/internal/filter"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/search"
)
@@ -51,11 +52,11 @@ func (s *Server) handleSearch(w http.ResponseWriter, r *http.Request) {
} else {
allowed = s.parseNamespacesForUser(r)
}
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
s.writeJSON(w, search.Result{Hits: []search.Hit{}})
return
}
- scanNamespaces := intersectNamespaces(allowed, parsed.NSFilter)
+ scanNamespaces := integration.IntersectNamespaces(allowed, parsed.NSFilter)
if allowed != nil && len(scanNamespaces) == 0 {
// User is namespace-restricted but their `ns:` filter doesn't
// intersect — empty result without scanning.
@@ -127,7 +128,7 @@ func (s *Server) handleSearch(w http.ResponseWriter, r *http.Request) {
// (CanReadClusterScoped) already constrains which cluster-scoped
// kinds are reachable.
if r.URL.Query().Get("context") != "none" {
- if builder := s.newSearchSummaryContextBuilder(scanNamespaces); builder != nil {
+ if builder := s.newSearchSummaryContextBuilder(r, scanNamespaces); builder != nil {
opts.SummaryBuilder = search.SummaryBuilderFunc(builder)
}
}
@@ -153,31 +154,6 @@ func (s *Server) handleSearch(w http.ResponseWriter, r *http.Request) {
s.writeJSON(w, result)
}
-// intersectNamespaces returns the namespaces to actually scan. nil `allowed`
-// means the user is unrestricted; preserve `requested` (which may also be nil
-// for cluster-wide). When the user is restricted, keep only the requested
-// namespaces they're allowed to see; if `requested` is empty, fall back to
-// the full allowed set.
-func intersectNamespaces(allowed, requested []string) []string {
- if allowed == nil {
- return requested
- }
- if len(requested) == 0 {
- return allowed
- }
- allowSet := make(map[string]struct{}, len(allowed))
- for _, ns := range allowed {
- allowSet[ns] = struct{}{}
- }
- out := make([]string, 0, len(requested))
- for _, ns := range requested {
- if _, ok := allowSet[ns]; ok {
- out = append(out, ns)
- }
- }
- return out
-}
-
func parseLimit(v string) int {
if v == "" {
return 0
diff --git a/internal/server/server.go b/internal/server/server.go
index 70800082e2..ec226d37ff 100644
--- a/internal/server/server.go
+++ b/internal/server/server.go
@@ -50,6 +50,7 @@ import (
"github.com/skyhook-io/radar/internal/connections"
"github.com/skyhook-io/radar/internal/helm"
"github.com/skyhook-io/radar/internal/images"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/investigationrefs"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/opencost"
@@ -589,6 +590,7 @@ func (s *Server) setupAppRoutes(r chi.Router) {
r.Get("/gitops/destination/{kind}/{namespace}/{name}", s.handleGitOpsDestination)
r.Get("/gitops/insights/{kind}/{namespace}/{name}", s.handleGitOpsInsights)
r.Get("/gitops/managed-resources", s.handleGitOpsManagedResources)
+ r.Post("/gitops/write-evidence", s.handleGitOpsWriteEvidence)
// RBAC reverse-lookup endpoints. Two shapes for /subject:
// ServiceAccount carries a namespace (3 segments after kind);
@@ -602,10 +604,37 @@ func (s *Server) setupAppRoutes(r chi.Router) {
r.Get("/rbac/whoami", s.handleRBACWhoami)
r.Get("/cnpg/workspace", s.handleCNPGWorkspace)
r.Get("/cnpg/operator", s.handleCNPGOperator)
+ r.Get("/cnpg/operator/status", s.handleCNPGOperatorStatus)
r.Get("/cnpg/imagecatalogs/{namespace}/{name}/clusters", s.handleCNPGCatalogUsers)
r.Get("/cnpg/clusterimagecatalogs/{name}/clusters", s.handleCNPGCatalogUsers)
r.Get("/cnpg/clusters/{namespace}/{name}/logs", s.handleCNPGClusterLogs)
r.Get("/cnpg/clusters/{namespace}/{name}/activity", s.handleCNPGClusterActivity)
+ r.Get("/cnpg/clusters/{namespace}/{name}/runtime", s.handleCNPGClusterRuntime)
+ r.Get("/cnpg/clusters/{namespace}/{name}/storage", s.handleCNPGClusterStorage)
+ r.Get("/cnpg/clusters/{namespace}/{name}/recovery", s.handleCNPGClusterRecovery)
+ r.Post("/cnpg/clusters/{namespace}/{name}/restore-validation", s.handleCNPGRestoreValidation)
+ r.Get("/cnpg/restore/capability", s.handleCNPGRestoreCapability)
+ r.Get("/cnpg/clusters/{namespace}/{name}/report", s.handleCNPGClusterReport)
+ r.Get("/cnpg/clusters/{namespace}/{name}/history", s.handleCNPGClusterHistory)
+ r.Get("/cnpg/disk", s.handleCNPGFleetDisk)
+ r.Get("/cnpg/clusters/{namespace}/{name}/ha", s.handleCNPGClusterHA)
+ r.Get("/cnpg/fleet-metrics", s.handleCNPGFleetMetrics)
+ r.Get("/cnpg/poolers/{namespace}/{name}/runtime", s.handleCNPGPoolerRuntime)
+ r.Get("/cnpg/clusters/{namespace}/{name}/capabilities", s.handleCNPGClusterCapabilities)
+ r.Post("/cnpg/clusters/{namespace}/{name}/actions/{action}", s.handleCNPGClusterAction)
+ r.Post("/cnpg/clusters/{namespace}/{name}/protection/preview", s.handleCNPGArchivingPreview)
+ r.Get("/cnpg/clusters/{namespace}/{name}/schedule-preview", s.handleCNPGDraftSchedulePreview)
+ r.Get("/cnpg/scheduledbackups/{namespace}/{name}/capabilities", s.handleCNPGScheduleCapabilities)
+ r.Post("/cnpg/scheduledbackups/{namespace}/{name}/actions/{action}", s.handleCNPGScheduleAction)
+ r.Get("/cnpg/scheduledbackups/{namespace}/{name}/schedule-preview", s.handleCNPGSchedulePreview)
+ r.Get("/cnpg/scheduledbackups/{namespace}/{name}/method-preview", s.handleCNPGScheduleMethodPreview)
+ r.Get("/cnpg/clusters/{namespace}/{name}/sessions", s.handleCNPGClusterSessions)
+ r.Get("/cnpg/clusters/{namespace}/{name}/restore-checks", s.handleCNPGRestoreChecks)
+ r.Get("/cnpg/clusters/{namespace}/{name}/parameters", s.handleCNPGClusterParameters)
+ r.Get("/cnpg/clusters/{namespace}/{name}/instances/{pod}/destroy-plan", s.handleCNPGDestroyPlan)
+ r.Get("/cnpg/poolers/{namespace}/{name}/capabilities", s.handleCNPGPoolerCapabilities)
+ r.Get("/cnpg/poolers/{namespace}/{name}/pgbouncer-state", s.handleCNPGPgBouncerState)
+ r.Post("/cnpg/poolers/{namespace}/{name}/actions/{action}", s.handleCNPGPoolerAction)
r.Get("/velero/backupstoragelocations/{namespace}/{name}/backups", s.handleVeleroStoredBackups)
// POST: creates a DownloadRequest, which is the only supported way to
// read the messages behind a run's error and warning counts.
@@ -1365,13 +1394,17 @@ func (s *Server) handleCapabilities(w http.ResponseWriter, r *http.Request) {
caps.Deployment = k8s.DeploymentInfo{Mode: deploymentMode()}
caps.CloudConnect = s.cloudConnectCapability()
caps.Features = k8s.FeatureCapabilities{
- YAMLReview: true,
- YAMLSchemas: true,
- WorkloadImages: true,
- ResourceIssues: true,
- PodEnvironment: true,
- PolicyResource: true,
- WorkloadHistory: true,
+ YAMLReview: true,
+ YAMLSchemas: true,
+ WorkloadImages: true,
+ ResourceIssues: true,
+ ResourceIssueCoverage: true,
+ PodEnvironment: true,
+ PolicyResource: true,
+ WorkloadHistory: true,
+ CNPGWorkspace: true,
+ CNPGProtectionSetup: true,
+ GitOpsWriteEvidence: true,
}
caps.AuthEnabled = s.authConfig.Enabled()
caps.ConfigManagement = s.configManagement()
@@ -1470,7 +1503,7 @@ func mergeNamespaceCapability(global, namespaced, checkErrored bool) bool {
// parseNamespacesForUser parses namespace query params and filters by user permissions.
// Returns nil for "all namespaces" (no filter), a populated slice for specific namespaces,
// or an empty non-nil slice when the user has no namespace access.
-// Use noNamespaceAccess() to check the no-access case.
+// Use integration.NoNamespaceAccess() to check the no-access case.
//
// If the request omits an explicit namespace filter, falls back to the user's
// in-app namespace pick (from the namespace switcher). The pick is treated as
@@ -1520,7 +1553,7 @@ func (s *Server) parseNamespacesForUser(r *http.Request) []string {
// if it's still the live pick, so a stale read can't wipe a concurrent
// POST or clear across a context switch. A failed access check is empty
// for the wrong reason and must not cost the user their pick.
- if pickFallback && !discoveryFailed && noNamespaceAccess(filtered) {
+ if pickFallback && !discoveryFailed && integration.NoNamespaceAccess(filtered) {
s.commitPickMutation(r, pickCtx, namespaces, nil, false)
filtered = s.getUserNamespaces(r, nil)
}
@@ -1585,7 +1618,10 @@ func (s *Server) resolveHelmNamespacesForScope(r *http.Request, namespaces []str
// so the (cluster-wide) pool only needs to be a superset of what the user can
// read.
func allNamespaceNames() []string {
- cache := k8s.GetResourceCache()
+ return namespaceNamesInCache(k8s.GetResourceCache())
+}
+
+func namespaceNamesInCache(cache *k8s.ResourceCache) []string {
if cache == nil {
return nil
}
@@ -1620,13 +1656,6 @@ func dedupeStrings(values []string) []string {
return out
}
-// noNamespaceAccess returns true when a namespace filter explicitly grants no access
-// (non-nil empty slice from auth filtering). Handlers with custom namespace logic
-// should check this and return empty results.
-func noNamespaceAccess(namespaces []string) bool {
- return namespaces != nil && len(namespaces) == 0
-}
-
// prometheusAuthGate is the per-request read check behind every metrics
// route. Two checks, both load-bearing:
//
@@ -1643,7 +1672,7 @@ func (s *Server) prometheusAuthGate(req *http.Request, group, resource, namespac
if !s.canRead(req, group, resource, namespace, verb) {
return false
}
- if namespace != "" && noNamespaceAccess(s.getUserNamespaces(req, []string{namespace})) {
+ if namespace != "" && integration.NoNamespaceAccess(s.getUserNamespaces(req, []string{namespace})) {
return false
}
return true
@@ -1920,7 +1949,7 @@ func (s *Server) handleTopology(w http.ResponseWriter, r *http.Request) {
return
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, map[string]any{"nodes": []any{}, "edges": []any{}})
return
}
@@ -2049,7 +2078,7 @@ func filterDynamicObservationNamespaces(observation k8score.DynamicResourceObser
observation.Namespaces = append([]string(nil), allowed...)
case k8score.DynamicObservationScopeExplicitNamespaces:
if len(observation.Namespaces) > 0 {
- observation.Namespaces = intersectNamespaces(allowed, observation.Namespaces)
+ observation.Namespaces = integration.IntersectNamespaces(allowed, observation.Namespaces)
}
}
if len(allowed) == 0 || (observation.Scope == k8score.DynamicObservationScopeExplicitNamespaces && len(observation.Namespaces) == 0) {
@@ -2154,7 +2183,7 @@ func (s *Server) preflightResourceList(r *http.Request, kind, group string, name
return nil, 0, "", true
}
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
return namespaces, http.StatusForbidden, "no namespace access", false
}
@@ -2859,7 +2888,7 @@ func (s *Server) preflightResourceGet(r *http.Request, kind, namespace, name, gr
case namespace != "":
// Namespaced kind: verify namespace access.
allowed := s.getUserNamespaces(r, []string{namespace})
- if noNamespaceAccess(allowed) {
+ if integration.NoNamespaceAccess(allowed) {
return http.StatusForbidden, fmt.Sprintf("no access to namespace %q", namespace), false
}
// Per-kind RBAC inside the namespace for Secrets — the chart can
@@ -3241,7 +3270,7 @@ func (s *Server) handlePodMetrics(w http.ResponseWriter, r *http.Request) {
namespace := chi.URLParam(r, "namespace")
name := chi.URLParam(r, "name")
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -3483,7 +3512,7 @@ func (s *Server) handlePodMetricsHistory(w http.ResponseWriter, r *http.Request)
namespace := chi.URLParam(r, "namespace")
name := chi.URLParam(r, "name")
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -3499,7 +3528,12 @@ func (s *Server) handlePodMetricsHistory(w http.ResponseWriter, r *http.Request)
s.writeJSON(w, history)
}
-func podMetricsHistoryResponse(ctx context.Context, history *k8s.PodMetricsHistory, namespace, name string, health k8s.MetricsCollectionHealth, includeAPIServiceConditionMessage bool) *k8s.PodMetricsHistory {
+type podMetricsHistoryWithAvailability struct {
+ *k8s.PodMetricsHistory
+ MetricsAPIReachable bool `json:"metricsAPIReachable"`
+}
+
+func podMetricsHistoryResponse(ctx context.Context, history *k8s.PodMetricsHistory, namespace, name string, health k8s.MetricsCollectionHealth, includeAPIServiceConditionMessage bool) *podMetricsHistoryWithAvailability {
if history == nil {
history = &k8s.PodMetricsHistory{
Namespace: namespace,
@@ -3510,7 +3544,10 @@ func podMetricsHistoryResponse(ctx context.Context, history *k8s.PodMetricsHisto
if health.PodMetrics.ConsecutiveErrors > 0 {
history.CollectionError, history.RawCollectionError, history.MetricsUnavailableDiagnosis, history.MetricsUnavailable = metricsHistoryCollectionError(ctx, "Pod", health.PodMetrics.LastError, includeAPIServiceConditionMessage)
}
- return history
+ return &podMetricsHistoryWithAvailability{
+ PodMetricsHistory: history,
+ MetricsAPIReachable: health.PodMetrics.LastSuccess != "" && health.PodMetrics.ConsecutiveErrors == 0,
+ }
}
// handleNodeMetricsHistory returns historical metrics for a specific node
@@ -3551,7 +3588,7 @@ func (s *Server) handleTopPods(w http.ResponseWriter, r *http.Request) {
return
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []k8s.TopPodMetrics{})
return
}
@@ -3764,7 +3801,7 @@ func (s *Server) handleTopResources(w http.ResponseWriter, r *http.Request) {
s.writeJSON(w, k8s.BuildTopMetrics(opts))
return
}
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, k8s.TopMetricsResponse{Kind: opts.Kind, Sort: opts.Sort, Reason: "no namespace access"})
return
}
@@ -3813,7 +3850,7 @@ func (s *Server) handleEvents(w http.ResponseWriter, r *http.Request) {
var events any
var err error
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []any{})
return
} else if len(namespaces) == 1 {
@@ -3864,7 +3901,7 @@ func (s *Server) handleChanges(w http.ResponseWriter, r *http.Request) {
return
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []any{})
return
}
@@ -5232,7 +5269,7 @@ func (s *Server) handleAuthMe(w http.ResponseWriter, r *http.Request) {
s.getUserNamespaces(r, nil)
}
if perms := s.permCache.Get(user.Username, user.Groups); perms != nil {
- resp["noNamespaceAccess"] = noNamespaceAccess(auth.FilterNamespacesForUser(nil, user, perms))
+ resp["noNamespaceAccess"] = integration.NoNamespaceAccess(auth.FilterNamespacesForUser(nil, user, perms))
}
}
}
diff --git a/internal/server/server_auth_test.go b/internal/server/server_auth_test.go
index 1962122e4f..6e22002ad8 100644
--- a/internal/server/server_auth_test.go
+++ b/internal/server/server_auth_test.go
@@ -10,6 +10,7 @@ import (
"testing"
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
)
@@ -168,7 +169,7 @@ func TestNoNamespaceAccess(t *testing.T) {
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
- if got := noNamespaceAccess(tt.ns); got != tt.want {
+ if got := integration.NoNamespaceAccess(tt.ns); got != tt.want {
t.Errorf("noNamespaceAccess(%v) = %v, want %v", tt.ns, got, tt.want)
}
})
@@ -241,7 +242,7 @@ func TestGetUserNamespaces_CachedNoAccess(t *testing.T) {
r := requestWithUser("GET", "/", user)
got := s.getUserNamespaces(r, []string{"dev"})
- if !noNamespaceAccess(got) {
+ if !integration.NoNamespaceAccess(got) {
t.Errorf("empty AllowedNamespaces should yield no access, got %v (nil=%v)", got, got == nil)
}
}
@@ -270,7 +271,7 @@ func TestGetUserNamespaces_UncachedFailsClosed_WhenK8sNotReady(t *testing.T) {
got := s.getUserNamespaces(r, []string{"dev"})
// When k8s client isn't ready, deny access (fail-closed)
- if !noNamespaceAccess(got) {
+ if !integration.NoNamespaceAccess(got) {
t.Errorf("expected no access (fail-closed) when k8s not ready, got %v", got)
}
}
diff --git a/internal/server/server_smoke_test.go b/internal/server/server_smoke_test.go
index 4e3d864351..99f42b56c5 100644
--- a/internal/server/server_smoke_test.go
+++ b/internal/server/server_smoke_test.go
@@ -2055,3 +2055,17 @@ func TestSmokeChangesExactGroup(t *testing.T) {
t.Fatalf("Volcano Job history = %v, want only its own versioned rows", got)
}
}
+
+func TestPodMetricsHistoryAPIReachability(t *testing.T) {
+ for _, tc := range []struct {
+ lastSuccess string
+ failures int
+ reachable bool
+ }{{"", 0, false}, {"2026-10-06T00:00:00Z", 0, true}, {"2026-10-06T00:00:00Z", 1, false}} {
+ health := k8s.MetricsCollectionHealth{PodMetrics: k8s.MetricsSourceHealth{LastSuccess: tc.lastSuccess, ConsecutiveErrors: tc.failures, LastError: "transport failed"}}
+ got := podMetricsHistoryResponse(context.Background(), nil, "default", "nginx", health, false)
+ if got.MetricsAPIReachable != tc.reachable {
+ t.Fatalf("%+v: %+v", tc, got)
+ }
+ }
+}
diff --git a/internal/server/summary_context.go b/internal/server/summary_context.go
index 6853930a1a..4cb0cb9400 100644
--- a/internal/server/summary_context.go
+++ b/internal/server/summary_context.go
@@ -10,6 +10,8 @@
package server
import (
+ "net/http"
+
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/summarycontext"
)
@@ -27,12 +29,12 @@ import (
// Use newSearchSummaryContextBuilder for search, which routes per-hit
// between a namespaced and a cluster-wide index — search returns mixed
// kinds in one response, so a single index can't get both right.
-func (s *Server) newResourceSummaryContextBuilder(namespaces []string) summarycontext.Builder {
+func (s *Server) newResourceSummaryContextBuilder(r *http.Request, namespaces []string) summarycontext.Builder {
provider := issues.NewCacheProvider()
if provider == nil {
return nil
}
- idx := summarycontext.BuildIssueIndex(provider, namespaces)
+ idx := summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: namespaces, CanReadClusterScoped: s.issueClusterScopedAccess(r), CanReadRelated: s.issueRelatedResourceAccess(r), CanReadEvidence: s.issueEvidenceAccess(r)})
return summarycontext.BuilderFromIndexes(s.broadcaster.GetCachedTopology(), idx, idx)
}
@@ -56,15 +58,15 @@ func (s *Server) newResourceSummaryContextBuilder(namespaces []string) summaryco
// The cluster-wide index is skipped when scanNamespaces is already nil
// (cluster-wide user) — both indexes would be identical, so one pass
// suffices.
-func (s *Server) newSearchSummaryContextBuilder(scanNamespaces []string) summarycontext.Builder {
+func (s *Server) newSearchSummaryContextBuilder(r *http.Request, scanNamespaces []string) summarycontext.Builder {
provider := issues.NewCacheProvider()
if provider == nil {
return nil
}
- namespacedIdx := summarycontext.BuildIssueIndex(provider, scanNamespaces)
+ namespacedIdx := summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: scanNamespaces, CanReadClusterScoped: s.issueClusterScopedAccess(r), CanReadRelated: s.issueRelatedResourceAccess(r), CanReadEvidence: s.issueEvidenceAccess(r)})
clusterIdx := namespacedIdx
if scanNamespaces != nil {
- clusterIdx = summarycontext.BuildIssueIndex(provider, nil)
+ clusterIdx = summarycontext.BuildIssueIndex(provider, issues.Filters{Namespaces: nil, CanReadClusterScoped: s.issueClusterScopedAccess(r), CanReadRelated: s.issueRelatedResourceAccess(r), CanReadEvidence: s.issueEvidenceAccess(r)})
}
return summarycontext.BuilderFromIndexes(s.broadcaster.GetCachedTopology(), namespacedIdx, clusterIdx)
}
diff --git a/internal/server/timeline_events.go b/internal/server/timeline_events.go
index e995cd3e83..b410e844d2 100644
--- a/internal/server/timeline_events.go
+++ b/internal/server/timeline_events.go
@@ -7,6 +7,7 @@ import (
"strings"
"time"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/timeline"
"github.com/skyhook-io/radar/pkg/timelineapi"
@@ -108,7 +109,7 @@ func (s *Server) serveTimelineEventsDelta(w http.ResponseWriter, r *http.Request
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
writeTimelineStream(w, nil, timelineEndRecord{Type: "end", Cursor: cursor})
return
}
@@ -163,7 +164,7 @@ func (s *Server) serveTimelineEventsWindow(w http.ResponseWriter, r *http.Reques
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
writeTimelineStream(w, nil, timelineEndRecord{Type: "end", Cursor: timelineCursor(epoch, 0)})
return
}
diff --git a/internal/server/topology_relationship_auth_test.go b/internal/server/topology_relationship_auth_test.go
index 5c29774b37..8cd496fb75 100644
--- a/internal/server/topology_relationship_auth_test.go
+++ b/internal/server/topology_relationship_auth_test.go
@@ -1,10 +1,11 @@
package server
import (
- "github.com/skyhook-io/radar/internal/auth"
- "github.com/skyhook-io/radar/pkg/topology"
"net/http/httptest"
"testing"
+
+ "github.com/skyhook-io/radar/internal/auth"
+ "github.com/skyhook-io/radar/pkg/topology"
)
func TestRelationshipTopologyHidesUnreadableEndpointsWithoutMutatingCache(t *testing.T) {
diff --git a/internal/server/trace_handlers.go b/internal/server/trace_handlers.go
index 1351ae95f1..3189402849 100644
--- a/internal/server/trace_handlers.go
+++ b/internal/server/trace_handlers.go
@@ -5,6 +5,7 @@ import (
"github.com/go-chi/chi/v5"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/issues"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/trace"
@@ -50,7 +51,7 @@ func (s *Server) handleTrace(w http.ResponseWriter, r *http.Request) {
// Mirror handleAuditResource: when the user has no namespace access
// (RBAC trims the set to empty), return an unknown-verdict trace
// instead of leaking that the resource even exists.
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, &trace.Trace{
Subject: trace.ResourceRef{Kind: kind, Namespace: namespace, Name: name},
Downstream: []trace.Hop{},
diff --git a/internal/server/traffic_handlers.go b/internal/server/traffic_handlers.go
index 66f76045a8..7cc74c3bfb 100644
--- a/internal/server/traffic_handlers.go
+++ b/internal/server/traffic_handlers.go
@@ -7,6 +7,7 @@ import (
"net/http"
"time"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/traffic"
)
@@ -115,7 +116,7 @@ func (s *Server) handleGetTrafficFlows(w http.ResponseWriter, r *http.Request) {
// Parse query parameters
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, []any{})
return
}
@@ -222,7 +223,7 @@ func (s *Server) handleTrafficFlowsStream(w http.ResponseWriter, r *http.Request
// Enforce per-user namespace access (parseNamespacesForUser intersects the
// requested ?namespace= with the user's RBAC-allowed namespaces).
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeError(w, http.StatusForbidden, "no namespace access")
return
}
diff --git a/internal/server/upgrade_readiness_handler_test.go b/internal/server/upgrade_readiness_handler_test.go
index e6d7b32de1..9ce0bd91fd 100644
--- a/internal/server/upgrade_readiness_handler_test.go
+++ b/internal/server/upgrade_readiness_handler_test.go
@@ -5,6 +5,7 @@ import (
"testing"
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
)
@@ -43,7 +44,7 @@ func TestUpgradeReadinessNamespacesIntersectsForcedScopeWithUserAccess(t *testin
s := &Server{permCache: auth.NewPermissionCache()}
s.permCache.Set("alice", nil, &auth.UserPermissions{AllowedNamespaces: []string{"tenant-b"}})
req := requestWithUser("GET", "/api/upgrade-readiness", &auth.User{Username: "alice"})
- if got := s.upgradeReadinessNamespaces(req); !noNamespaceAccess(got) {
+ if got := s.upgradeReadinessNamespaces(req); !integration.NoNamespaceAccess(got) {
t.Fatalf("upgrade namespace scope = %v, want no access", got)
}
}
diff --git a/internal/server/velero_handlers.go b/internal/server/velero_handlers.go
index a5ae29e252..1cc98bfabc 100644
--- a/internal/server/velero_handlers.go
+++ b/internal/server/velero_handlers.go
@@ -10,6 +10,7 @@ import (
"github.com/go-chi/chi/v5"
"k8s.io/apimachinery/pkg/apis/meta/v1/unstructured"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
)
@@ -106,7 +107,7 @@ func (s *Server) handleVeleroStoredBackups(w http.ResponseWriter, r *http.Reques
// No Velero on this cluster, so nothing is stored anywhere.
s.writeJSON(w, VeleroStoredBackupsResponse{Backups: []VeleroStoredBackup{}})
return
- case errors.Is(err, errDynamicNotSynced):
+ case errors.Is(err, integration.ErrDynamicNotSynced):
s.writeError(w, http.StatusServiceUnavailable, "backups are still loading")
return
default:
diff --git a/internal/server/velero_handlers_test.go b/internal/server/velero_handlers_test.go
index ef962bb808..ebc172ada3 100644
--- a/internal/server/velero_handlers_test.go
+++ b/internal/server/velero_handlers_test.go
@@ -98,7 +98,7 @@ func TestVeleroStoredBackups_DeniesRatherThanReportingAnEmptyLocation(t *testing
// location.
func TestVeleroStoredBackups_SeparatesAnAbsentCRDFromAFailedRead(t *testing.T) {
src := mustReadSource(t, "velero_handlers.go")
- for _, want := range []string{"k8s.ErrUnknownDynamicKind", "errDynamicNotSynced", "listDynamicSynced", "sanitizeForLog"} {
+ for _, want := range []string{"k8s.ErrUnknownDynamicKind", "integration.ErrDynamicNotSynced", "listDynamicSynced", "sanitizeForLog"} {
if !strings.Contains(src, want) {
t.Errorf("handler does not use %s — a failed or unsynced read would report an empty location", want)
}
diff --git a/internal/server/vitals.go b/internal/server/vitals.go
index 01f5486f69..7f32a3f0a2 100644
--- a/internal/server/vitals.go
+++ b/internal/server/vitals.go
@@ -5,12 +5,12 @@ import (
"sync"
"time"
+ "golang.org/x/sync/singleflight"
corev1 "k8s.io/api/core/v1"
"k8s.io/apimachinery/pkg/labels"
- "golang.org/x/sync/singleflight"
-
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/pkg/health"
"github.com/skyhook-io/radar/pkg/k8score"
@@ -167,7 +167,7 @@ func (s *Server) handleVitals(w http.ResponseWriter, r *http.Request) {
return
}
namespaces := s.parseNamespacesForUser(r)
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
s.writeJSON(w, VitalsResponse{
Completeness: VitalsCompleteness{AccessRestricted: true},
})
@@ -204,7 +204,7 @@ func (s *Server) handleVitals(w http.ResponseWriter, r *http.Request) {
}
var scopedPods []*corev1.Pod
if podsReadable {
- scopedPods = listPodsScoped(cache.Pods(), podNamespaces)
+ scopedPods = integration.ListPodsScoped(cache.Pods(), podNamespaces)
}
if podsReadable && cache.Pods() != nil {
pods := scopedPods
diff --git a/internal/server/workload_history.go b/internal/server/workload_history.go
index 4d1b3d77b6..6c31dd6870 100644
--- a/internal/server/workload_history.go
+++ b/internal/server/workload_history.go
@@ -17,6 +17,7 @@ import (
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
"github.com/skyhook-io/radar/internal/auth"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
"github.com/skyhook-io/radar/internal/timeline"
"github.com/skyhook-io/radar/pkg/resourceid"
@@ -61,7 +62,7 @@ func (s *Server) handleWorkloadHistory(w http.ResponseWriter, r *http.Request) {
}
namespace := chi.URLParam(r, "namespace")
name := chi.URLParam(r, "name")
- if allowed := s.getUserNamespaces(r, []string{namespace}); noNamespaceAccess(allowed) {
+ if allowed := s.getUserNamespaces(r, []string{namespace}); integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
diff --git a/internal/server/workload_logs.go b/internal/server/workload_logs.go
index 8f9709ad56..4a68abe71a 100644
--- a/internal/server/workload_logs.go
+++ b/internal/server/workload_logs.go
@@ -26,54 +26,14 @@ import (
"k8s.io/client-go/dynamic"
"k8s.io/client-go/kubernetes"
+ integration "github.com/skyhook-io/radar/internal/integration"
"github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
"github.com/skyhook-io/radar/pkg/health"
"github.com/skyhook-io/radar/pkg/k8score"
"github.com/skyhook-io/radar/pkg/rollouts"
)
-// WorkloadPodContainerInfo contains compact per-container runtime status for the UI.
-type WorkloadPodContainerInfo struct {
- Name string `json:"name"`
- Init bool `json:"init,omitempty"`
- Ready bool `json:"ready"`
- RestartCount int32 `json:"restartCount"`
-}
-
-// WorkloadPodInfo contains compact runtime status about a pod for workload views.
-type WorkloadPodInfo struct {
- Name string `json:"name"`
- Containers []string `json:"containers"`
- Ready bool `json:"ready"`
- Phase string `json:"phase,omitempty"`
- NodeName string `json:"nodeName,omitempty"`
- HealthLevel string `json:"healthLevel,omitempty"`
- Reason string `json:"reason,omitempty"`
- Message string `json:"message,omitempty"`
- RestartCount int32 `json:"restartCount,omitempty"`
- LastTerminationReason string `json:"lastTerminationReason,omitempty"`
- CreatedAt string `json:"createdAt,omitempty"`
- ContainerStatuses []WorkloadPodContainerInfo `json:"containerStatuses,omitempty"`
- StepID string `json:"stepID,omitempty"`
- StepName string `json:"stepName,omitempty"`
- StepPhase string `json:"stepPhase,omitempty"`
- RevisionIdentity string `json:"revisionIdentity,omitempty"`
- UpdatedRevision *bool `json:"updatedRevision,omitempty"`
-}
-
-// workloadLogEntry is an internal structure for log lines from pods
-type workloadLogEntry struct {
- Pod string `json:"pod"`
- Container string `json:"container"`
- Timestamp string `json:"timestamp"`
- Content string `json:"content"`
- SourceLabel string `json:"sourceLabel,omitempty"`
- // Parsed from a structured line by sources that know their log format.
- Level string `json:"level,omitempty"`
- Logger string `json:"logger,omitempty"`
- Message string `json:"message,omitempty"`
-}
-
type workloadLogMetadata struct {
EmptyReason string `json:"emptyReason,omitempty"`
EmptyMessage string `json:"emptyMessage,omitempty"`
@@ -223,12 +183,12 @@ func (s *Server) handleWorkloadRuns(w http.ResponseWriter, r *http.Request) {
runNamespaces := []string{namespace}
if clusterScoped {
runNamespaces = s.parseNamespacesForUser(r)
- if noNamespaceAccess(runNamespaces) {
+ if integration.NoNamespaceAccess(runNamespaces) {
s.writeError(w, http.StatusForbidden, "no namespace access")
return
}
} else {
- if allowed := s.getUserNamespaces(r, []string{namespace}); noNamespaceAccess(allowed) {
+ if allowed := s.getUserNamespaces(r, []string{namespace}); integration.NoNamespaceAccess(allowed) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return
}
@@ -323,7 +283,7 @@ func (s *Server) handleWorkloadRuns(w http.ResponseWriter, r *http.Request) {
}
func (s *Server) readableRunNamespaces(r *http.Request, group, resource string, namespaces []string) ([]string, bool) {
- if noNamespaceAccess(namespaces) {
+ if integration.NoNamespaceAccess(namespaces) {
return nil, false
}
if namespaces == nil {
@@ -343,7 +303,7 @@ func (s *Server) authorizeWorkloadPodRead(w http.ResponseWriter, r *http.Request
s.writeError(w, http.StatusBadRequest, "only deployments, statefulsets, daemonsets, rollouts, jobs, and workflows are supported")
return false
}
- if noNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
+ if integration.NoNamespaceAccess(s.getUserNamespaces(r, []string{namespace})) {
s.writeError(w, http.StatusForbidden, "no access to namespace "+namespace)
return false
}
@@ -384,8 +344,8 @@ func (s *Server) handleWorkloadLogs(w http.ResponseWriter, r *http.Request) {
}
container := r.URL.Query().Get("container")
- tailLines := parseTailLines(r.URL.Query().Get("tailLines"), 100)
- sinceSeconds := parseSinceSeconds(r.URL.Query().Get("sinceSeconds"))
+ tailLines := podlogs.ParseTailLines(r.URL.Query().Get("tailLines"), 100)
+ sinceSeconds := podlogs.ParseSinceSeconds(r.URL.Query().Get("sinceSeconds"))
pods, err := s.getWorkloadPods(kind, namespace, name)
if err != nil {
@@ -396,8 +356,8 @@ func (s *Server) handleWorkloadLogs(w http.ResponseWriter, r *http.Request) {
if len(pods) == 0 {
metadata := s.describeWorkloadLogEmpty(r.Context(), kind, namespace, name)
response := map[string]any{
- "pods": []WorkloadPodInfo{},
- "logs": []workloadLogEntry{},
+ "pods": []podlogs.PodInfo{},
+ "logs": []podlogs.Entry{},
}
addWorkloadLogMetadata(response, metadata)
s.writeJSON(w, response)
@@ -411,9 +371,9 @@ func (s *Server) handleWorkloadLogs(w http.ResponseWriter, r *http.Request) {
}
// Collect logs from all pods concurrently
- snapshot := collectLogsFromPods(r.Context(), client, namespace, pods, container, tailLines, sinceSeconds, false)
+ snapshot := podlogs.CollectPods(r.Context(), client, namespace, pods, container, tailLines, sinceSeconds, false)
- sortLogsByTimestamp(snapshot.Logs)
+ podlogs.Sort(snapshot.Logs)
s.writeJSON(w, map[string]any{
"pods": buildPodInfos(pods),
@@ -433,8 +393,8 @@ func (s *Server) handleWorkloadLogsStream(w http.ResponseWriter, r *http.Request
}
container := r.URL.Query().Get("container")
- tailLines := parseTailLines(r.URL.Query().Get("tailLines"), 50)
- sinceSeconds := parseSinceSeconds(r.URL.Query().Get("sinceSeconds"))
+ tailLines := podlogs.ParseTailLines(r.URL.Query().Get("tailLines"), 50)
+ sinceSeconds := podlogs.ParseSinceSeconds(r.URL.Query().Get("sinceSeconds"))
// Set SSE headers
w.Header().Set("Content-Type", "text/event-stream")
@@ -494,7 +454,7 @@ func (s *Server) handleWorkloadLogsStream(w http.ResponseWriter, r *http.Request
defer cancel()
// Channel for aggregated log lines
- logCh := make(chan workloadLogEntry, 1000)
+ logCh := make(chan podlogs.Entry, 1000)
// Track active streams
var activeStreams sync.Map // podName/containerName -> cancel func
@@ -568,7 +528,7 @@ func (s *Server) handleWorkloadLogsStream(w http.ResponseWriter, r *http.Request
knownPods[p.Name] = true
// Notify frontend about new pod
sendSSEEvent(w, flusher, "pod_added", map[string]any{
- "pods": []WorkloadPodInfo{buildPodInfo(p, time.Now())},
+ "pods": []podlogs.PodInfo{podlogs.BuildPodInfo(p, time.Now())},
})
}
}
@@ -629,7 +589,7 @@ func shouldWaitForPodsInLogStream(kind string, metadata workloadLogMetadata) boo
}
// streamPodLogs streams logs from a single pod/container to the log channel
-func streamPodLogs(ctx context.Context, client kubernetes.Interface, namespace, podName, containerName string, tailLines int64, sinceSeconds *int64, logCh chan<- workloadLogEntry) {
+func streamPodLogs(ctx context.Context, client kubernetes.Interface, namespace, podName, containerName string, tailLines int64, sinceSeconds *int64, logCh chan<- podlogs.Entry) {
stream, err := k8score.GetContainerLogs(ctx, client, namespace, podName, containerName, k8score.LogOptions{
TailLines: &tailLines,
SinceSeconds: sinceSeconds,
@@ -662,9 +622,9 @@ func streamPodLogs(ctx context.Context, client kubernetes.Interface, namespace,
continue
}
- ts, content := parseLogLine(line)
+ ts, content := podlogs.ParseLine(line)
select {
- case logCh <- workloadLogEntry{
+ case logCh <- podlogs.Entry{
Pod: podName,
Container: containerName,
Timestamp: ts,
@@ -677,29 +637,16 @@ func streamPodLogs(ctx context.Context, client kubernetes.Interface, namespace,
}
}
-// isPodReady checks if all containers in a pod are ready
-func isPodReady(pod *corev1.Pod) bool {
- if pod.Status.Phase != corev1.PodRunning {
- return false
- }
- for _, cs := range pod.Status.ContainerStatuses {
- if !cs.Ready {
- return false
- }
- }
- return true
-}
-
-// buildPodInfos converts pods to WorkloadPodInfo slice
-func buildPodInfos(pods []*corev1.Pod) []WorkloadPodInfo {
+// buildPodInfos converts pods to podlogs.PodInfo slice
+func buildPodInfos(pods []*corev1.Pod) []podlogs.PodInfo {
return buildPodInfosForRevision(pods, workloadRevisionTarget{})
}
-func buildPodInfosForRevision(pods []*corev1.Pod, target workloadRevisionTarget) []WorkloadPodInfo {
- infos := make([]WorkloadPodInfo, 0, len(pods))
+func buildPodInfosForRevision(pods []*corev1.Pod, target workloadRevisionTarget) []podlogs.PodInfo {
+ infos := make([]podlogs.PodInfo, 0, len(pods))
now := time.Now()
for _, pod := range pods {
- info := buildPodInfo(pod, now)
+ info := podlogs.BuildPodInfo(pod, now)
if target.label != "" && target.value != "" {
if identity := pod.Labels[target.label]; identity != "" {
updated := identity == target.value
@@ -714,7 +661,7 @@ func buildPodInfosForRevision(pods []*corev1.Pod, target workloadRevisionTarget)
const maxWorkloadPodResponseLimit = 200
-func limitWorkloadPodInfos(infos []WorkloadPodInfo, limit int) ([]WorkloadPodInfo, bool) {
+func limitWorkloadPodInfos(infos []podlogs.PodInfo, limit int) ([]podlogs.PodInfo, bool) {
sort.SliceStable(infos, func(i, j int) bool {
leftRank := health.Rank(health.Level(infos[i].HealthLevel))
rightRank := health.Rank(health.Level(infos[j].HealthLevel))
@@ -732,89 +679,6 @@ func limitWorkloadPodInfos(infos []WorkloadPodInfo, limit int) ([]WorkloadPodInf
return infos[:limit], true
}
-// buildPodInfo converts a single pod to WorkloadPodInfo
-func buildPodInfo(pod *corev1.Pod, now time.Time) WorkloadPodInfo {
- containers := make([]string, 0, len(pod.Spec.Containers)+len(pod.Spec.InitContainers))
- containerStatuses := make([]WorkloadPodContainerInfo, 0, len(pod.Status.InitContainerStatuses)+len(pod.Status.ContainerStatuses))
- for _, c := range pod.Spec.InitContainers {
- containers = append(containers, c.Name)
- }
- for _, c := range pod.Spec.Containers {
- containers = append(containers, c.Name)
- }
- for _, cs := range pod.Status.InitContainerStatuses {
- containerStatuses = append(containerStatuses, WorkloadPodContainerInfo{
- Name: cs.Name,
- Init: true,
- Ready: cs.Ready,
- RestartCount: cs.RestartCount,
- })
- }
- for _, cs := range pod.Status.ContainerStatuses {
- containerStatuses = append(containerStatuses, WorkloadPodContainerInfo{
- Name: cs.Name,
- Ready: cs.Ready,
- RestartCount: cs.RestartCount,
- })
- }
- verdict := health.Pod(pod, now)
- displayLevel := health.PodDisplayLevel(pod, now)
- if displayLevel != verdict.Level {
- verdict.Level = displayLevel
- if verdict.Reason == "" {
- verdict.Reason = health.PodProblemReason(pod, now)
- }
- if verdict.Message == "" {
- verdict.Message = health.PodProblemMessage(pod)
- }
- }
- restartCount, lastTerminationReason := health.PodRestartContext(pod)
- createdAt := ""
- if !pod.CreationTimestamp.IsZero() {
- createdAt = pod.CreationTimestamp.Time.Format(time.RFC3339)
- }
- annotations := pod.GetAnnotations()
- labels := pod.GetLabels()
- return WorkloadPodInfo{
- Name: pod.Name,
- Containers: containers,
- Ready: isPodReady(pod),
- Phase: string(pod.Status.Phase),
- NodeName: pod.Spec.NodeName,
- HealthLevel: string(verdict.Level),
- Reason: verdict.Reason,
- Message: verdict.Message,
- RestartCount: restartCount,
- LastTerminationReason: lastTerminationReason,
- CreatedAt: createdAt,
- ContainerStatuses: containerStatuses,
- StepID: annotations["workflows.argoproj.io/node-id"],
- StepName: annotations["workflows.argoproj.io/node-name"],
- StepPhase: labels["workflows.argoproj.io/phase"],
- }
-}
-
-// sortLogsByTimestamp sorts log entries by timestamp using efficient sort
-func sortLogsByTimestamp(logs []workloadLogEntry) {
- sort.SliceStable(logs, func(i, j int) bool {
- left, le := time.Parse(time.RFC3339Nano, logs[i].Timestamp)
- right, re := time.Parse(time.RFC3339Nano, logs[j].Timestamp)
- if le == nil && re == nil && !left.Equal(right) {
- return left.Before(right)
- }
- if (le == nil) != (re == nil) {
- return le != nil
- }
- if le != nil && logs[i].Timestamp != logs[j].Timestamp {
- return logs[i].Timestamp < logs[j].Timestamp
- }
- if logs[i].Pod != logs[j].Pod {
- return logs[i].Pod < logs[j].Pod
- }
- return logs[i].Container < logs[j].Container
- })
-}
-
// workloadError represents a typed error for workload operations
type workloadError struct {
statusCode int
@@ -1646,179 +1510,3 @@ func applyTerminalWorkflowEmptyState(metadata *workloadLogMetadata, workflow map
func (s *Server) writeWorkloadError(w http.ResponseWriter, err *workloadError) {
s.writeError(w, err.statusCode, err.message)
}
-
-// parseSinceSeconds parses sinceSeconds query parameter, returning nil if not set
-func parseSinceSeconds(str string) *int64 {
- if str == "" {
- return nil
- }
- if s, err := strconv.ParseInt(str, 10, 64); err == nil && s > 0 {
- return &s
- }
- return nil
-}
-
-// parseTailLines parses tailLines query parameter with a default value
-func parseTailLines(str string, defaultVal int64) int64 {
- if str == "" {
- return defaultVal
- }
- if t, err := strconv.ParseInt(str, 10, 64); err == nil && t > 0 {
- return t
- }
- return defaultVal
-}
-
-// collectLogsFromPods fetches logs from all pods concurrently. Non-nil even
-// when nothing is retrievable (e.g. every pod is crashlooping) — a nil slice
-// marshals as JSON null and consumers expect an array.
-type workloadLogSnapshot struct {
- SourcePods map[string]bool
- Logs []workloadLogEntry
- Notice string
-}
-
-const maxSnapshotSources = 40
-const maxSnapshotSourceBytes int64 = 64 * 1024
-
-func collectLogsFromPods(ctx context.Context, client kubernetes.Interface, namespace string, pods []*corev1.Pod, container string, tailLines int64, sinceSeconds *int64, bounded bool) workloadLogSnapshot {
- type source struct {
- pod, container string
- running bool
- created time.Time
- }
- sources := []source{}
- for _, pod := range pods {
- for _, c := range k8s.GetContainersForPod(pod, container, true) {
- if bounded {
- started := false
- for _, statuses := range [][]corev1.ContainerStatus{pod.Status.ContainerStatuses, pod.Status.InitContainerStatuses, pod.Status.EphemeralContainerStatuses} {
- for _, status := range statuses {
- if status.Name == c && (status.State.Running != nil || status.State.Terminated != nil) {
- started = true
- }
- }
- }
- if !started {
- continue
- }
- }
- sources = append(sources, source{pod.Name, c, pod.Status.Phase == corev1.PodRunning, pod.CreationTimestamp.Time})
- }
- }
- sort.Slice(sources, func(i, j int) bool {
- if bounded && sources[i].running != sources[j].running {
- return sources[i].running
- }
- if bounded && !sources[i].created.Equal(sources[j].created) {
- return sources[i].created.After(sources[j].created)
- }
- if sources[i].pod != sources[j].pod {
- return sources[i].pod < sources[j].pod
- }
- return sources[i].container < sources[j].container
- })
- total := len(sources)
- if bounded && len(sources) > maxSnapshotSources {
- sources = sources[:maxSnapshotSources]
- }
- if bounded {
- tailLines = min(tailLines, 1000)
- var cancel context.CancelFunc
- ctx, cancel = context.WithTimeout(ctx, 30*time.Second)
- defer cancel()
- }
- result := workloadLogSnapshot{Logs: []workloadLogEntry{}, SourcePods: map[string]bool{}}
- for _, src := range sources {
- result.SourcePods[src.pod] = true
- }
- var mu sync.Mutex
- var wg sync.WaitGroup
- concurrency := len(sources)
- if bounded {
- concurrency = min(concurrency, 8)
- }
- sem := make(chan struct{}, concurrency)
- errors_ := []string{}
- truncated := 0
- for _, src := range sources {
- wg.Add(1)
- go func() {
- defer wg.Done()
- select {
- case sem <- struct{}{}:
- defer func() { <-sem }()
- case <-ctx.Done():
- mu.Lock()
- errors_ = append(errors_, src.pod+"/"+src.container+": request cancelled")
- mu.Unlock()
- return
- }
- entries, clipped, err := fetchPodContainerLogs(ctx, client, namespace, src.pod, src.container, tailLines, sinceSeconds, bounded)
- mu.Lock()
- defer mu.Unlock()
- if err != nil {
- errors_ = append(errors_, src.pod+"/"+src.container+": "+err.Error())
- }
- if clipped {
- truncated++
- }
- result.Logs = append(result.Logs, entries...)
- }()
- }
- wg.Wait()
- notices := []string{}
- if total > len(sources) {
- notices = append(notices, fmt.Sprintf("Showing %d of %d container sources. Narrow the scope to see other sources.", len(sources), total))
- }
- if truncated > 0 {
- notices = append(notices, fmt.Sprintf("%d sources reached the 64 KiB snapshot limit.", truncated))
- }
- if len(errors_) > 0 {
- sort.Strings(errors_)
- notices = append(notices, fmt.Sprintf("%d sources could not be read: %s", len(errors_), strings.Join(errors_[:min(3, len(errors_))], "; ")))
- }
- result.Notice = strings.Join(notices, " ")
- return result
-}
-
-func fetchPodContainerLogs(ctx context.Context, client kubernetes.Interface, namespace, podName, containerName string, tailLines int64, sinceSeconds *int64, bounded bool) ([]workloadLogEntry, bool, error) {
- var limit *int64
- if bounded {
- n := maxSnapshotSourceBytes + 1
- limit = &n
- }
- stream, err := k8score.GetContainerLogs(ctx, client, namespace, podName, containerName, k8score.LogOptions{
- TailLines: &tailLines, SinceSeconds: sinceSeconds, Timestamps: true, LimitBytes: limit,
- })
- if err != nil {
- return nil, false, err
- }
- defer stream.Close()
- var reader io.Reader = stream
- if limit != nil {
- reader = io.LimitReader(stream, *limit)
- }
- content, err := io.ReadAll(reader)
- if err != nil {
- return nil, false, err
- }
- clipped := bounded && int64(len(content)) > maxSnapshotSourceBytes
- if clipped {
- content = content[:maxSnapshotSourceBytes]
- if last := strings.LastIndexByte(string(content), '\n'); last >= 0 {
- content = content[:last+1]
- } else {
- content = nil
- }
- }
- entries := []workloadLogEntry{}
- for _, line := range strings.Split(string(content), "\n") {
- if line == "" {
- continue
- }
- ts, text := parseLogLine(line)
- entries = append(entries, workloadLogEntry{Pod: podName, Container: containerName, Timestamp: ts, Content: text})
- }
- return entries, clipped, nil
-}
diff --git a/internal/server/workload_logs_test.go b/internal/server/workload_logs_test.go
index 5c12b4249c..4f6ba0a78d 100644
--- a/internal/server/workload_logs_test.go
+++ b/internal/server/workload_logs_test.go
@@ -13,7 +13,6 @@ import (
"time"
"github.com/go-chi/chi/v5"
-
appsv1 "k8s.io/api/apps/v1"
batchv1 "k8s.io/api/batch/v1"
corev1 "k8s.io/api/core/v1"
@@ -28,6 +27,7 @@ import (
clienttesting "k8s.io/client-go/testing"
"github.com/skyhook-io/radar/internal/k8s"
+ "github.com/skyhook-io/radar/internal/podlogs"
"github.com/skyhook-io/radar/pkg/health"
"github.com/skyhook-io/radar/pkg/k8score"
)
@@ -81,7 +81,7 @@ func TestBuildPodInfosForRevisionAttributesOnlyKnownIdentities(t *testing.T) {
}
func TestLimitWorkloadPodInfosBoundsProblemFirst(t *testing.T) {
- infos := []WorkloadPodInfo{
+ infos := []podlogs.PodInfo{
{Name: "healthy", HealthLevel: string(health.LevelHealthy)},
{Name: "degraded-low-restarts", HealthLevel: string(health.LevelDegraded), RestartCount: 1},
{Name: "unhealthy", HealthLevel: string(health.LevelUnhealthy)},
diff --git a/internal/summarycontext/summarycontext.go b/internal/summarycontext/summarycontext.go
index 5173fc6034..f3da2d16c5 100644
--- a/internal/summarycontext/summarycontext.go
+++ b/internal/summarycontext/summarycontext.go
@@ -186,12 +186,11 @@ func CanonicalSingular(kind string) string {
// engine emits "Application", silently zeroing issueCount on every CRD
// row. Bucketing is O(N) over the at-most-namespace-bounded issue set,
// which the consumer materialises anyway.
-func BuildIssueIndex(p issues.Provider, namespaces []string) IssueIndex {
- filters := issues.Filters{
- SkipPodTemplateContext: true,
- Namespaces: namespaces,
- Limit: issues.NoLimit,
- }
+func BuildIssueIndex(p issues.Provider, filters issues.Filters) IssueIndex {
+ filters.SkipPodTemplateContext = true
+ filters.Limit = issues.NoLimit
+ filters.Grouped = false
+
// Compose FLAT (uncapped): every evidence row carries the grouped issue ID
// (enrichIdentity keys it on owner-else-self + category) and its resolved
// Owner. Counting DISTINCT grouped issue IDs per resource — keyed on each
diff --git a/internal/summarycontext/summarycontext_test.go b/internal/summarycontext/summarycontext_test.go
index b1c8538591..c3d2b14134 100644
--- a/internal/summarycontext/summarycontext_test.go
+++ b/internal/summarycontext/summarycontext_test.go
@@ -32,6 +32,8 @@ import (
// pin the actual bug.
type fakeIssuesProvider struct {
problems []k8s.Detection
+ dynamic map[schema.GroupVersionResource][]*unstructured.Unstructured
+ kinds map[schema.GroupVersionResource]string
}
func (f *fakeIssuesProvider) DetectProblems(namespaces []string) []k8s.Detection {
@@ -60,16 +62,22 @@ func (f *fakeIssuesProvider) DetectScheduling(_ []string) []k8s.Detection {
func (f *fakeIssuesProvider) WarningEvents(_ []string, _ time.Duration) []*corev1.Event {
return nil
}
-func (f *fakeIssuesProvider) WatchedDynamic() []schema.GroupVersionResource { return nil }
-func (f *fakeIssuesProvider) ListDynamic(_ schema.GroupVersionResource, _ string) ([]*unstructured.Unstructured, error) {
- return nil, nil
+func (f *fakeIssuesProvider) WatchedDynamic() []schema.GroupVersionResource {
+ var out []schema.GroupVersionResource
+ for gvr := range f.dynamic {
+ out = append(out, gvr)
+ }
+ return out
}
-func (f *fakeIssuesProvider) ListDynamicAllNamespaces(_ schema.GroupVersionResource) ([]*unstructured.Unstructured, error) {
- return nil, nil
+func (f *fakeIssuesProvider) ListDynamic(gvr schema.GroupVersionResource, _ string) ([]*unstructured.Unstructured, error) {
+ return f.dynamic[gvr], nil
}
-func (f *fakeIssuesProvider) KindForGVR(_ schema.GroupVersionResource) string { return "" }
-func (f *fakeIssuesProvider) KyvernoFindings() []policyreports.SubjectFindings { return nil }
-func (f *fakeIssuesProvider) KyvernoStatus() string { return "" }
+func (f *fakeIssuesProvider) ListDynamicAllNamespaces(gvr schema.GroupVersionResource) ([]*unstructured.Unstructured, error) {
+ return f.dynamic[gvr], nil
+}
+func (f *fakeIssuesProvider) KindForGVR(gvr schema.GroupVersionResource) string { return f.kinds[gvr] }
+func (f *fakeIssuesProvider) KyvernoFindings() []policyreports.SubjectFindings { return nil }
+func (f *fakeIssuesProvider) KyvernoStatus() string { return "" }
func fmtPodName(i int) string { return fmt.Sprintf("pod-%05d", i) }
@@ -112,7 +120,7 @@ func TestBuildIssueIndex_GroupAware(t *testing.T) {
{Kind: "Service", Group: "serving.knative.dev", Namespace: "prod", Name: "api", Reason: "RouteNotReady", Severity: "warning"},
},
}
- idx := BuildIssueIndex(p, nil)
+ idx := BuildIssueIndex(p, issues.Filters{})
// The index counts GROUPED issues (consistent with the issues tool), not flat
// rows: the two Knative rows share subject+category and fold into one grouped
// issue → count 1. The core Service is a distinct group → its own key (the
@@ -138,7 +146,7 @@ func TestBuildIssueIndex_GroupedSubjectPropagation(t *testing.T) {
OwnerGroup: "apps", OwnerKind: "Deployment", OwnerName: "web"},
},
}
- idx := BuildIssueIndex(p, nil)
+ idx := BuildIssueIndex(p, issues.Filters{})
if got := idx.Count("apps", "Deployment", "prod", "web"); got != 1 {
t.Errorf("owning Deployment count = %d, want 1 (Pod-evidenced issue must surface on the grouped subject)", got)
}
@@ -161,7 +169,7 @@ func TestBuildIssueIndex_BeyondMaxLimit(t *testing.T) {
})
}
p := &fakeIssuesProvider{problems: probs}
- idx := BuildIssueIndex(p, nil)
+ idx := BuildIssueIndex(p, issues.Filters{})
tailName := fmtPodName(issues.MaxLimit + 25)
if got := idx.Count("", "Pod", "prod", tailName); got != 1 {
t.Fatalf("tail pod %s count = %d, want 1 (silent MaxLimit truncation?)", tailName, got)
@@ -205,7 +213,7 @@ func TestBuildIssueIndex_ClusterScopedIssueSurfacedWhenUnfiltered(t *testing.T)
}
// Cluster-wide compose (nil namespaces) — issue surfaces.
- idx := BuildIssueIndex(p, nil)
+ idx := BuildIssueIndex(p, issues.Filters{})
if got := idx.Count("", "Node", "", "worker-1"); got != 1 {
t.Errorf("cluster-wide index: Node issueCount = %d, want 1 (cluster-scoped issue should appear)", got)
}
@@ -214,7 +222,7 @@ func TestBuildIssueIndex_ClusterScopedIssueSurfacedWhenUnfiltered(t *testing.T)
// ["prod","staging"] drops it because the user-namespaced perm
// slice never matches "". This is what the pre-fix handler did for
// Node lists.
- scopedIdx := BuildIssueIndex(p, []string{"prod", "staging"})
+ scopedIdx := BuildIssueIndex(p, issues.Filters{Namespaces: []string{"prod", "staging"}})
if got := scopedIdx.Count("", "Node", "", "worker-1"); got != 0 {
t.Errorf("namespace-scoped index: Node issueCount = %d, want 0 (namespace filter drops cluster-scoped issue)", got)
}
@@ -243,7 +251,7 @@ func TestBuildIssueIndex_CRDPlural_NonZeroCount(t *testing.T) {
// Pre-fix simulation: the handler would have passed kindFilter="applications"
// — the URL plural. We no longer take a kindFilter, but verify that
// the index contains the row keyed by the canonical singular form.
- idx := BuildIssueIndex(p, []string{"argocd"})
+ idx := BuildIssueIndex(p, issues.Filters{Namespaces: []string{"argocd"}})
if got := idx.Count("argoproj.io", "Application", "argocd", "storefront"); got != 1 {
t.Errorf("CRD Application count (singular kind) = %d, want 1", got)
}
@@ -276,8 +284,8 @@ func TestNewSearchSummaryContextBuilder_BuildsDualIndex(t *testing.T) {
}
// Build the two indexes the search constructor would build.
- namespacedIdx := BuildIssueIndex(p, []string{"prod"})
- clusterIdx := BuildIssueIndex(p, nil)
+ namespacedIdx := BuildIssueIndex(p, issues.Filters{Namespaces: []string{"prod"}})
+ clusterIdx := BuildIssueIndex(p, issues.Filters{})
// Sanity: pre-fix, the search handler passed namespacedIdx for
// both; Node issueCount silently zeroed.
@@ -404,3 +412,27 @@ func TestManagedByFromRelationships_NilSafe(t *testing.T) {
t.Errorf("empty rel: got %#v, want nil", got)
}
}
+
+func TestBuildIssueIndexAuthorizesInventoryEvidence(t *testing.T) {
+ clusterGVR := schema.GroupVersionResource{Group: "postgresql.cnpg.io", Version: "v1", Resource: "clusters"}
+ scheduleGVR := schema.GroupVersionResource{Group: clusterGVR.Group, Version: "v1", Resource: "scheduledbackups"}
+ cluster := &unstructured.Unstructured{Object: map[string]any{"apiVersion": "postgresql.cnpg.io/v1", "kind": "Cluster", "metadata": map[string]any{"namespace": "db", "name": "pg"}}}
+ schedule := &unstructured.Unstructured{Object: map[string]any{"apiVersion": "postgresql.cnpg.io/v1", "kind": "ScheduledBackup", "metadata": map[string]any{"namespace": "db", "name": "nightly"}, "spec": map[string]any{"cluster": map[string]any{"name": "pg"}}}}
+ p := &fakeIssuesProvider{dynamic: map[schema.GroupVersionResource][]*unstructured.Unstructured{clusterGVR: {cluster}, scheduleGVR: {schedule}}, kinds: map[schema.GroupVersionResource]string{clusterGVR: "Cluster", scheduleGVR: "ScheduledBackup"}}
+ for _, tc := range []struct {
+ name string
+ filters issues.Filters
+ count int
+ }{
+ {"no authorizer", issues.Filters{}, 0},
+ {"target denied", issues.Filters{CanReadEvidence: func(r issues.EvidenceRead) bool { return r.Resource != "clusters" }}, 0},
+ {"all evidence authorized", issues.Filters{CanReadEvidence: func(issues.EvidenceRead) bool { return true }}, 1},
+ {"internal composition", issues.Filters{AllowUnfilteredEvidence: true}, 1},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ if got := BuildIssueIndex(p, tc.filters).Count(clusterGVR.Group, "ScheduledBackup", "db", "nightly"); got != tc.count {
+ t.Fatalf("count=%d, want %d", got, tc.count)
+ }
+ })
+ }
+}
diff --git a/packages/k8s-ui/README.md b/packages/k8s-ui/README.md
index 312a353f94..cc636eae67 100644
--- a/packages/k8s-ui/README.md
+++ b/packages/k8s-ui/README.md
@@ -29,3 +29,18 @@ command remains pending for the next connection.
`YamlEditor` and `YamlDiffEditor` bundle Monaco, its editor worker, and the YAML language worker into the consuming application. They make no runtime CDN or internet requests, so they work in air-gapped environments.
Consumers must use a bundler that supports module workers created with `new Worker(new URL(..., import.meta.url), { type: 'module' })`. The source package is compatible with Vite and webpack 5. Monaco and `monaco-yaml` are intentionally pinned as a compatible pair because their worker-factory APIs must move together. The package declares that exact Monaco version as a peer so the host and YAML runtime share one Monaco instance.
+
+`CreateResourceDialog` accepts optional `onBack(yaml)` and `backLabel` props for
+parent flows. Back returns the current editor draft without invoking `onClose`;
+it is disabled during writes, previews and file imports. The review screen
+continues to return to the editor first. Hosts own draft retention and any
+confirmation before replacing edited YAML.
+
+After a successful write, `CreateResourceDialog.onCreated(result, submittedYaml)`
+receives the first result and the exact submitted manifest. Dry-runs do not invoke
+it. Existing callbacks that use only the result continue to work; hosts can use
+the submitted YAML to describe the actual operation rather than an earlier form.
+
+`ActionConfirmDialog` also accepts optional `onBack` and `backLabel` for steps
+inside a parent flow. Back and Cancel remain distinct; both are disabled while
+the confirmed action is pending.
diff --git a/packages/k8s-ui/src/components/charts/AreaChart.test.tsx b/packages/k8s-ui/src/components/charts/AreaChart.test.tsx
index 93e62f7709..0984d2fa24 100644
--- a/packages/k8s-ui/src/components/charts/AreaChart.test.tsx
+++ b/packages/k8s-ui/src/components/charts/AreaChart.test.tsx
@@ -171,3 +171,28 @@ describe('AreaChart compact layout and count axes', () => {
expect(compactCount).not.toContain('>10<')
})
})
+
+describe('AreaChart gaps and selection', () => {
+ const base = [series([[0, 1], [60, null], [120, 3], [180, 2]])]
+
+ it('hatches shaded ranges and draws the selection band inside the plot', () => {
+ const html = render({
+ series: base,
+ stepSeconds: 60,
+ shadedRanges: [{ start: t0 + 30, end: t0 + 90, label: 'No sample' }],
+ selection: { start: t0 + 120, end: t0 + 180 },
+ onSelectRange: () => {},
+ })
+ expect(html.match(/data-chart-gap/g)).toHaveLength(1)
+ expect(html).toMatch(/fill="url\(#hatch-[^"]+\)"/)
+ expect(html).toContain('data-chart-selection')
+ expect(html).toContain('cursor:col-resize')
+ })
+
+ it('draws nothing extra without the new props', () => {
+ const html = render({ series: base })
+ expect(html).not.toContain('data-chart-gap')
+ expect(html).not.toContain('data-chart-selection')
+ expect(html).toContain('cursor:crosshair')
+ })
+})
diff --git a/packages/k8s-ui/src/components/charts/AreaChart.tsx b/packages/k8s-ui/src/components/charts/AreaChart.tsx
index 4805de9337..1b65f5b772 100644
--- a/packages/k8s-ui/src/components/charts/AreaChart.tsx
+++ b/packages/k8s-ui/src/components/charts/AreaChart.tsx
@@ -1,4 +1,4 @@
-import { useCallback, useEffect, useMemo, useRef, useState } from 'react'
+import { useCallback, useEffect, useId, useMemo, useRef, useState } from 'react'
import type * as React from 'react'
import { seriesColor, seriesFill, computeShortLabels, seriesDisplayLabels } from './colors'
import { formatMetricValue, formatTimestamp } from './format'
@@ -13,7 +13,13 @@ import { nearestSample } from './nearestSample'
export const ANNOTATION_LABEL_MIN_WIDTH_PX = 420
const ANNOTATION_HOVER_TOLERANCE = 8
-export function AreaChart({ series, color, fillColor, unit, referenceLines, annotations, domain, seriesLabels, stepSeconds, layout = 'full' }: {
+/** A time range on the chart's X axis, in unix seconds. */
+export interface ChartTimeRange {
+ start: number
+ end: number
+}
+
+export function AreaChart({ series, color, fillColor, unit, referenceLines, annotations, domain, seriesLabels, stepSeconds, layout = 'full', shadedRanges, selection, onSelectRange }: {
series: TimeSeries[]
color: string
fillColor: string
@@ -39,7 +45,18 @@ export function AreaChart({ series, color, fillColor, unit, referenceLines, anno
* keeps text and plot height stable as its container resizes.
*/
layout?: 'full' | 'auto' | 'compact' | 'dashboard'
+ /** Ranges drawn hatched, e.g. where no sample exists; `label` says why and shows on hover. */
+ shadedRanges?: (ChartTimeRange & { label?: string })[]
+ /** A selected range, drawn as a band. */
+ selection?: ChartTimeRange | null
+ /**
+ * Enables range selection: drag across the plot, or click to pick the one
+ * evaluation step under the pointer (needs `stepSeconds`).
+ */
+ onSelectRange?: (range: ChartTimeRange) => void
}) {
+ const hatchId = `hatch-${useId().replace(/[^a-zA-Z0-9_-]/g, '')}`
+ const [drag, setDrag] = useState<{ from: number; to: number } | null>(null)
const svgRef = useRef(null)
const wrapperRef = useRef(null)
const [hoverX, setHoverX] = useState(null)
@@ -264,17 +281,48 @@ export function AreaChart({ series, color, fillColor, unit, referenceLines, anno
p => Math.abs(p.x - clampedX) <= ANNOTATION_HOVER_TOLERANCE,
)
- return { ts, x: clampedX, points, nearbyAnnotations }
- }, [hoverX, chartData, placedAnnotations, seriesLabels, marginLeft, plotWidth, marginTop, plotHeight, multiSeries, color, stepSeconds])
+ const shaded = shadedRanges?.find(r => ts >= r.start && ts <= r.end && r.label)
- const handleMouseMove = useCallback((e: React.MouseEvent) => {
- const svg = svgRef.current
- if (!svg) return
- const ctm = svg.getScreenCTM()
- if (!ctm) return
- setHoverX((e.clientX - ctm.e) / ctm.a)
+ return { ts, x: clampedX, points, nearbyAnnotations, shaded }
+ }, [hoverX, chartData, placedAnnotations, seriesLabels, marginLeft, plotWidth, marginTop, plotHeight, multiSeries, color, stepSeconds, shadedRanges])
+
+ const pointerX = useCallback((e: React.MouseEvent) => {
+ const ctm = svgRef.current?.getScreenCTM()
+ return ctm ? (e.clientX - ctm.e) / ctm.a : null
}, [])
+ const handleMouseMove = useCallback((e: React.MouseEvent) => {
+ const x = pointerX(e)
+ if (x === null) return
+ setHoverX(x)
+ setDrag(d => (d ? { ...d, to: x } : d))
+ }, [pointerX])
+
+ const tsAt = (x: number) => {
+ if (!chartData) return 0
+ const clamped = Math.max(marginLeft, Math.min(marginLeft + plotWidth, x))
+ return chartData.minTs + ((clamped - marginLeft) / plotWidth) * (chartData.maxTs - chartData.minTs)
+ }
+
+ const handleMouseDown = (e: React.MouseEvent) => {
+ if (!onSelectRange || e.button !== 0) return
+ const x = pointerX(e)
+ if (x !== null) setDrag({ from: x, to: x })
+ }
+
+ const handleMouseUp = () => {
+ if (!drag || !onSelectRange || !chartData) return
+ setDrag(null)
+ const a = tsAt(Math.min(drag.from, drag.to))
+ const b = tsAt(Math.max(drag.from, drag.to))
+ if (Math.abs(drag.to - drag.from) >= 6) {
+ onSelectRange({ start: Math.round(a), end: Math.round(b) })
+ } else if (stepSeconds) {
+ const half = stepSeconds / 2
+ onSelectRange({ start: Math.round(Math.max(chartData.minTs, a - half)), end: Math.round(Math.min(chartData.maxTs, a + half)) })
+ }
+ }
+
// Hook calls above run unconditionally; bail out of rendering only after
// every hook has been invoked (Rules of Hooks).
if (!chartData) return null
@@ -292,6 +340,38 @@ export function AreaChart({ series, color, fillColor, unit, referenceLines, anno
preserveAspectRatio="xMidYMid meet"
data-chart-layout={compact ? 'compact' : 'full'}
>
+
+
+
+
+
+
+ {shadedRanges?.map((r, i) => {
+ const x1 = Math.max(marginLeft, toX(r.start))
+ const x2 = Math.min(marginLeft + plotWidth, toX(r.end))
+ if (x2 <= x1) return null
+ return
+ })}
+
+ {selection && (() => {
+ const x1 = Math.max(marginLeft, toX(selection.start))
+ const x2 = Math.min(marginLeft + plotWidth, toX(selection.end))
+ if (x2 < x1) return null
+ return
+ })()}
+
+ {drag && Math.abs(drag.to - drag.from) >= 2 && (
+
+ )}
+
{/* Grid lines */}
{yTicks.map((tick, i) => (
setHoverX(null)}
+ onMouseDown={handleMouseDown}
+ onMouseUp={handleMouseUp}
+ onMouseLeave={() => {
+ setHoverX(null)
+ setDrag(null)
+ }}
/>
@@ -538,6 +623,9 @@ export function AreaChart({ series, color, fillColor, unit, referenceLines, anno
))}
+ {hoverData.shaded && hoverData.points.length === 0 && (
+ {hoverData.shaded.label}
+ )}
{hoverData.points.map((p, i) => (
-
Outcome
+
Outcome
@@ -109,7 +116,7 @@ export function CNPGBackupSummary({ resource, workspace, onNavigate }: SummaryPr
)}
-
Relationships
+
Relationships
@@ -152,22 +159,62 @@ export function CNPGBackupSummary({ resource, workspace, onNavigate }: SummaryPr
)
}
-export function CNPGScheduledBackupSummary({ resource, workspace, onNavigate }: SummaryProps) {
+export function CNPGSchedulePreviewFacts({ preview }: { preview: CNPGSchedulePreview }) {
+ if (!preview.valid) {
+ return (
+
+ Not a schedule the operator can run: {preview.error}
+
+ )
+ }
+ return (
+ <>
+ {preview.description && {preview.description}{!preview.clock?.declared && · UTC estimate } }
+
+
+ {(preview.nextRuns ?? []).map((r, i) => {
+ const t = formatCNPGRunTime(r)
+ return (
+
+
+ {i === 0 && preview.runsImmediately ? `due (${t.utc})` : t.utc}
+
+
+ )
+ })}
+
+ {cnpgScheduleBasisNote(preview)}
+
+ >
+ )
+}
+
+export function CNPGScheduledBackupSummary({
+ resource,
+ workspace,
+ onNavigate,
+ schedulePreview,
+}: SummaryProps & {
+ /** The server's reading of spec.schedule; without it the schedule is shown verbatim only. */
+ schedulePreview?: CNPGSchedulePreview
+}) {
const ns = resource?.metadata?.namespace ?? ''
const cron = resource?.spec?.schedule
const next = getCNPGScheduledBackupNextSchedule(resource)
const runsUnavailable = relationUnavailable(workspace, 'backups', ns, 'Backups')
const runs = runsUnavailable ? [] : backupsForScheduledBackup(resource, workspaceList(workspace, 'backups'))
const shown = runs.slice(0, RECENT_RUNS)
+ const blocker = resource?.spec?.suspend === true ? null : cnpgScheduleDestinationBlocker(resource, clustersIn(workspace))
return (
- Schedule
+ Schedule
-
+ {blocker && {blocker} }
+ {blocker && The resource list status comes from the ScheduledBackup alone. }
@@ -182,14 +229,15 @@ export function CNPGScheduledBackupSummary({ resource, workspace, onNavigate }:
)}
+ {cron && schedulePreview && schedulePreview.schedule === cron && }
- {next === '-' ? : next}
+ {next === '-' ? : next}
{methodText(resource) ?? 'Barman object store (in-tree) · default'}
- shown.length ? `${shown.length} of ${runs.length}` : undefined}>Recent runs
+ shown.length ? `${shown.length} of ${runs.length}` : undefined}>Recent runs
{runsUnavailable ? (
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.test.tsx
new file mode 100644
index 0000000000..7422fb2f4b
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.test.tsx
@@ -0,0 +1,90 @@
+// @vitest-environment jsdom
+import { renderToStaticMarkup } from 'react-dom/server'
+import { describe, expect, it } from 'vitest'
+import { CNPGClusterHASection } from './CNPGClusterHASection'
+import type { CNPGClusterHA } from './ha'
+
+const noInstances: CNPGClusterHA = {
+ cluster: { namespace: 'db', name: 'pg', uid: 'u' },
+ sampledAt: '2026-09-30T12:00:00Z',
+ desiredImage: 'pg:17',
+ instances: [],
+ pods: { state: 'ok' },
+ nodes: { state: 'ok' },
+ quorum: { enabled: false, object: { state: 'notFound' } },
+ pdbs: { state: 'ok', enabled: true, items: [] },
+ primaryLease: { state: 'notFound' },
+ operatorLease: { state: 'unavailable' },
+ jobs: { state: 'ok', items: [] },
+ rwEndpoints: { state: 'ok', service: 'pg-rw', pods: [] },
+ certificates: [],
+ maintenance: { declared: false, inProgress: false, reusePVC: true },
+}
+
+describe('CNPGClusterHASection', () => {
+ it('says there is nothing to place rather than showing only the zone source', () => {
+ const html = renderToStaticMarkup(
)
+ expect(html).toContain('No instance Pods to place')
+ expect(html).not.toContain('topology.kubernetes.io/zone of each instance')
+ })
+})
+
+it('shows the read-write Service, its Pods and primary mismatch beside Reachability', () => {
+ const ha: CNPGClusterHA = { ...noInstances, rwEndpoints: { state: 'ok', service: 'pg-rw', pods: ['pg-2'] } }
+ const html = renderToStaticMarkup(
{}} />)
+ expect(html).toContain('Read-write Service')
+ expect(html).toContain('pg-rw')
+ expect(html).toContain('pg-2')
+ expect(html).toContain('not on the reported primary pg-1')
+ expect(html).toContain('Reachability')
+ expect(html).toContain('aria-expanded="true"')
+})
+it('distinguishes absent endpoints from denied endpoint reads', () => {
+ expect(renderToStaticMarkup( )).toContain('No ready endpoints')
+ const denied: CNPGClusterHA = { ...noInstances, rwEndpoints: { state: 'denied', service: 'pg-rw', pods: [], grant: { verb: 'list', resource: 'endpointslices', group: 'discovery.k8s.io', namespace: 'db' } } }
+ const html = renderToStaticMarkup( )
+ expect(html).toContain('endpointslices')
+ expect(html).not.toContain('No ready endpoints')
+})
+it('does not flag deliberately absent read-write endpoints while hibernated', () => {
+ const html = renderToStaticMarkup( )
+ expect(html).toContain('None expected while hibernated')
+ expect(html).not.toContain('No ready endpoints')
+})
+
+it('shows the desired image without claiming any instance runs it', () => {
+ const html = renderToStaticMarkup( )
+ expect(html).toContain('Desired image pg:17; no instance running')
+ expect(html).not.toContain('every instance runs it')
+})
+
+it('describes matching Pod images without claiming the Pods are running', () => {
+ const ha: CNPGClusterHA = { ...noInstances, instances: [{ pod: 'pg-1', podUID: 'p1', role: 'primary', ready: false, restartCount: 0, image: 'pg:17', imageMatches: true }] }
+ const html = renderToStaticMarkup( )
+ expect(html).toContain('observed Pod images match')
+ expect(html).not.toContain('instances run it')
+})
+
+it('keeps the readiness statement inside expanded HA and lists missing expected instances without an observed role', () => {
+ const ha: CNPGClusterHA = { ...noInstances, declaredInstances: 2, expectedInstances: ['orders-1', 'orders-2'], instances: [{ pod: 'orders-1', podUID: 'a', role: 'primary', ready: true, restartCount: 0 }], jobs: { state: 'ok', items: [{ name: 'orders-2-join', role: 'join', phase: 'pending', reason: 'Pod cannot be scheduled: Unschedulable: 0/2 nodes are available: 2 Too many pods. preemption: no victims.' }] } }
+ const html = renderToStaticMarkup( )
+ expect(html).toContain('1 of 2 declared instances ready; no instance Pod observed for orders-2')
+ expect(html).toContain('orders-1')
+ expect(html).toContain('orders-2 · not running')
+ expect(html).toContain('Cannot be scheduled: both nodes have reached their Pod limit')
+ expect(html).toContain('Scheduler message')
+ expect(html).toContain('aria-expanded="false"')
+ expect(html).toContain('preemption: no victims')
+})
+
+it('keeps each Job reason and scheduler disclosure inside its own block', () => {
+ const jobs = ['pg-1-initdb', 'pg-2-join'].map((name) => ({ name, role: 'join', phase: 'pending' as const, reason: `Pod cannot be scheduled: ${name}: 0/2 nodes are available: 2 Too many pods.` }))
+ const host = document.createElement('div'); host.innerHTML = renderToStaticMarkup( )
+ const blocks = [...host.querySelectorAll('div.space-y-1.text-xs')]
+ expect(blocks).toHaveLength(2)
+ for (let i = 0; i < jobs.length; i++) {
+ expect(blocks[i].textContent).toContain(jobs[i].name)
+ expect(blocks[i].querySelector('[inert]')?.textContent).toContain(jobs[i].reason)
+ expect(blocks[i].textContent).not.toContain(jobs[1 - i].name)
+ }
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.tsx b/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.tsx
new file mode 100644
index 0000000000..643dfa2579
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGClusterHASection.tsx
@@ -0,0 +1,331 @@
+import { clsx } from 'clsx'
+import { Badge } from '../ui/Badge'
+import { Tooltip } from '../ui/Tooltip'
+import { formatAge, summarizeSchedulerMessage } from '../resources/resource-utils'
+import { PrimaryConflictNote } from './primitives'
+import {
+ CNPG_ROLE_DETAIL_TEXT,
+ cnpgCertificateViews,
+ cnpgCertificatesSummary,
+ cnpgHASourceText,
+ cnpgHASummary,
+ cnpgImageDrift,
+ cnpgLeaseHolderPod,
+ cnpgLiveGap,
+ cnpgPDBFact,
+ cnpgPendingRestart,
+ cnpgQuorumFact,
+ cnpgZoneSpread,
+ type CNPGClusterHA,
+ type CNPGHAJob,
+ type CNPGHALease,
+ type CNPGInstanceLive,
+} from './ha'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { StatusDot, toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactValue } from '../facts'
+import { FoldSection, SectionHeading } from '../ui/FoldSection'
+
+function Unknown({ text }: { text: string }) {
+ return {text}
+}
+
+function LeaseValue({ lease, what }: { lease: CNPGHALease; what: string }) {
+ if (lease.state !== 'ok') return
+ return (
+
+ held by{' '}
+ {lease.holder && cnpgLeaseHolderPod(lease.holder) !== lease.holder ? (
+
+ {cnpgLeaseHolderPod(lease.holder)}
+
+ ) : (
+ {lease.holder || '(nobody)'}
+ )}
+ {lease.renewTime && · renewed {formatAge(lease.renewTime)} ago }
+ {lease.expired && · expired }
+ {lease.controlledByCluster === false && · not owned by this Cluster }
+
+ )
+}
+
+const JOB_SEVERITY: Record = {
+ succeeded: 'success',
+ failed: 'error',
+ running: 'info',
+ active: 'info',
+ pending: 'neutral',
+}
+
+/**
+ * "HA and instances": whether the cluster survives losing an instance, and what
+ * a planned switchover will meet. Every row names its source; a fact the caller
+ * cannot read says so instead of reading as none.
+ */
+export function CNPGClusterHASection({
+ ha,
+ live,
+ liveUnavailable,
+ loading,
+ error,
+ onNavigate,
+ primaryConflict,
+ title = 'HA and instances',
+ showInstances = true,
+ showCertificates = true,
+ currentPrimary,
+ hibernated = false,
+ onOpenReachability,
+}: {
+ currentPrimary?: string
+ hibernated?: boolean
+ onOpenReachability?: (service: { namespace: string; name: string }) => void
+ title?: string
+ /** False where the host lists the instances itself (with their replication state). */
+ showInstances?: boolean
+ /** False where the host shows certificates elsewhere (CNPGClusterCertificates). */
+ showCertificates?: boolean
+ /** status.currentPrimary vs the Pod labelled primary, when they disagree. */
+ primaryConflict?: { status: string; labelled: string }
+ ha?: CNPGClusterHA
+ /** Instance-manager facts, when the caller can read them. */
+ live?: CNPGInstanceLive[]
+ /** Why `live` is absent (e.g. "needs get pods/proxy in db"); read from `live` itself when it is present. */
+ liveUnavailable?: string
+ loading?: boolean
+ error?: string
+ onNavigate?: NavigateToRef
+}) {
+ if (!ha) {
+ return (
+ <>
+ {title}
+ {loading ? 'Reading HA facts…' : `HA facts could not be read${error ? `: ${error}` : ''}`}
+ >
+ )
+ }
+ const ns = ha.cluster.namespace
+ const spread = cnpgZoneSpread(ha)
+ const drift = cnpgImageDrift(ha)
+ const pending = cnpgPendingRestart(live)
+ const liveBy = new Map((live ?? []).map((l) => [l.pod, l]))
+ const jobs = [...ha.jobs.items].sort((a, b) => (a.phase === 'succeeded' ? 1 : 0) - (b.phase === 'succeeded' ? 1 : 0))
+ const versions = new Set((live ?? []).map((l) => l.instanceManagerVersion).filter(Boolean))
+
+ const haSummary = cnpgHASummary(ha, live, primaryConflict)
+ const endpointProblem = !hibernated && currentPrimary && ha.rwEndpoints.state === 'ok' && !ha.rwEndpoints.pods.includes(currentPrimary)
+ ? ha.rwEndpoints.pods.length === 0 ? 'no read-write endpoint' : 'read-write endpoint not on the primary'
+ : undefined
+
+ return (
+ <>
+
+ {{haSummary.text}
}
+ {ha.pods.state === 'ok' &&
+ {ha.instances.map((i) => · Pod {i.ready ? 'ready' : 'not ready'} )}
+ {[...new Set(ha.expectedInstances ?? [])].filter((name) => !ha.instances.some((i) => i.pod === name)).map((name) => {name} · not running )}
+
}
+
+
+
+
+ {ha.rwEndpoints.state !== 'ok' ?
: (
+ <>
+
+ {ha.rwEndpoints.pods.length === 0 ? {hibernated ? 'None expected while hibernated' : 'No ready endpoints'} : ha.rwEndpoints.pods.map((pod) => )}
+
+ {!hibernated && currentPrimary && ha.rwEndpoints.pods.length > 0 && !ha.rwEndpoints.pods.includes(currentPrimary) &&
Read-write endpoint is not on the reported primary {currentPrimary}.
}
+ >
+ )}
+ {onOpenReachability &&
onOpenReachability({ namespace: ns, name: ha.rwEndpoints.service })}>Reachability → }
+
Ready Pods from the Service’s EndpointSlices
+
+
+
+ {!spread.known ? (
+
+
+ {spread.nodes.length > 0 && (
+
+ Nodes: {spread.nodes.map((n) => `${n.node} (${n.pods.join(', ')})`).join(' · ')}
+
+ )}
+
+ ) : spread.zones.length === 0 && spread.unlabelled.length === 0 ? (
+
+ ) : (
+
+
+ {spread.zones.map((z) => (
+
+ {z.zone}
+ : {z.pods.join(', ')}
+
+ ))}
+ {spread.unlabelled.length > 0 && }
+
+ {(spread.singleZone || spread.sharedNode) && (
+
+ {spread.singleZone ? 'Every instance is in one zone: losing it loses the cluster. ' : ''}
+ {spread.sharedNode ? 'Two or more instances share a Node.' : ''}
+
+ )}
+
topology.kubernetes.io/zone of each instance’s Node
+
+ )}
+
+
+ {showInstances &&
+ {ha.pods.state !== 'ok' ? (
+
+ ) : (
+
+ {ha.instances.map((i) => {
+ const l = liveBy.get(i.pod)
+ return (
+
+
+
+
+
+
+ {l?.roleDetail ? CNPG_ROLE_DETAIL_TEXT[l.roleDetail] : i.role === 'unknown' ? 'role unknown' : i.role}
+
+ {l?.timeline !== undefined && timeline {l.timeline} }
+ {i.qosClass && QoS {i.qosClass} }
+ {l?.pendingRestart && pending restart }
+ {i.imageMatches === false && image differs }
+ {l?.instanceManagerVersion && versions.size > 1 && CNPG instance manager {l.instanceManagerVersion} }
+
+ )
+ })}
+ {primaryConflict &&
}
+
+ {live ? 'Role detail and pending restart from each instance manager' : `Role from Pod labels; role detail ${cnpgLiveGap(undefined, liveUnavailable)}`}
+ {versions.size === 1 ? ` · instance manager ${[...versions][0]}` : ''}
+
+
+ )}
+ }
+
+
+ {!pending.known && pending.pods.length === 0 ? (
+
+ ) : pending.pods.length === 0 ? (
+ None reported
+ ) : (
+
+ {pending.pods.join(', ')} {pending.pods.length === 1 ? 'needs' : 'need'} a restart to apply changed parameters
+ {pending.forDecrease ? ' (a lowered setting: the primary restarts first)' : ''}
+ {!pending.known ? ` · ${cnpgLiveGap(live, liveUnavailable)}` : ''}
+
+ )}
+
+
+
+ {!drift.known && drift.drifted.length === 0 ? (
+
+ ) : drift.drifted.length === 0 ? (
+
+ {ha.desiredImage}
+ · observed Pod images match
+
+ ) : (
+
+ {drift.drifted.map((d) => `${d.pod} has image ${d.image}`).join(' · ')} · desired {ha.desiredImage}
+
+ )}
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+ {ha.jobs.state !== 'ok' ? (
+
+ ) : jobs.length === 0 ? (
+ None present
+ ) : (
+
+ {jobs.slice(0, 6).map((j) => (
+
+
+ {j.phase}
+ {j.role ?? 'job'}
+
+
+ {j.reason && (/schedul|Too many pods|Insufficient/i.test(j.reason) ?
+
Cannot be scheduled: {summarizeSchedulerMessage(j.reason, { plain: true })}
+
+
:
{j.reason} )}
+
+ ))}
+ {jobs.length > 6 &&
+{jobs.length - 6} more
}
+
+ )}
+
+
+
+
+ {showCertificates && }
+ >
+ )
+}
+
+/** Certificate expiry and who renews each certificate, folded to one line unless one needs attention. */
+export function CNPGClusterCertificates({ ha, onNavigate }: { ha: CNPGClusterHA; onNavigate?: NavigateToRef }) {
+ const ns = ha.cluster.namespace
+ const certs = cnpgCertificateViews(ha.certificates)
+ const certSummary = cnpgCertificatesSummary(ha.certificates)
+ return (
+
+
+
+ {certs.length === 0 ? (
+
+ ) : (
+
+ {certs.map((c) => (
+
+ {' '}
+
+ {c.expiresAt ? (c.daysLeft !== undefined && c.daysLeft < 0 ? `expired ${c.expiresAt}` : `expires in ${c.daysLeft} d`) : `expiry unreadable (“${c.raw}”)`}
+
+
+ {' · '}
+ {c.renewal === 'operator' ? 'CloudNativePG renews it' : 'you renew it (spec.certificates)'}
+
+ {c.renewal === 'user' &&
+ (c.certManager ? (
+
+ {' · cert-manager '}
+
+
+ ) : c.metadata?.state === 'ok' ? (
+ · not issued by cert-manager
+ ) : (
+ · issuer unknown ({cnpgHASourceText(c.metadata, 'Secret metadata')})
+ ))}
+
+ ))}
+
status.certificates.expirations
+
+ )}
+
+
+
+ )
+}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.test.tsx
new file mode 100644
index 0000000000..134ea3b0b5
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.test.tsx
@@ -0,0 +1,275 @@
+// @vitest-environment jsdom
+import { act, type ReactNode } from 'react'
+import { createRoot } from 'react-dom/client'
+import { afterEach, describe, expect, it, vi } from 'vitest'
+import { CNPGClusterSummary, CNPGDimensionMark, CNPGDimensionVerdict, CNPGServingStatus } from './CNPGClusterSummary'
+import type { CNPGFleetRow, CNPGProblem } from './workspace'
+import type { CNPGDimension } from './ha'
+import { OpenIssueContext } from '../problems'
+
+;(globalThis as { IS_REACT_ACT_ENVIRONMENT?: boolean }).IS_REACT_ACT_ENVIRONMENT = true
+
+function row(over: Partial = {}): CNPGFleetRow {
+ return {
+ key: 'db/pg',
+ namespace: 'db',
+ name: 'pg',
+ cluster: { status: { currentPrimary: 'pg-1' } },
+ controllerStatus: { text: 'Healthy', level: 'healthy' },
+ instances: { ready: 2, desired: 2 },
+ pods: [
+ { name: 'pg-1', role: 'primary', ready: true },
+ { name: 'pg-2', role: 'replica', ready: true },
+ ],
+ replicaCluster: null,
+ hibernated: false,
+ pgVersion: '17',
+ catalog: null,
+ replication: { text: 'Lag unknown', tone: 'unknown' },
+ protection: {
+ schedule: { text: 'Active', tone: 'healthy', names: [] },
+ destination: { text: 'ObjectStore', tone: 'healthy', method: 'plugin' },
+ lastSuccessfulBackup: { text: '1 h ago', tone: 'healthy' },
+ walArchiving: { text: 'Archiving', tone: 'healthy' },
+ recoveryWindow: { text: 'x', tone: 'healthy' },
+ restoreValidation: { text: 'None recorded', tone: 'unknown' },
+ summary: { text: 'ok', tone: 'healthy' },
+ },
+ declarations: { summary: { text: 'None', tone: 'neutral' }, total: 0, failed: 0, pending: 0 },
+ poolers: [],
+ poolersKnown: true,
+ problems: [],
+ attention: false,
+ categories: new Set(),
+ ...over,
+ }
+}
+
+const problem = (id: string, severity: CNPGProblem['severity'], title: string): CNPGProblem => ({
+ id,
+ severity,
+ category: 'availability',
+ title,
+ subject: { kind: 'Backup', group: 'postgresql.cnpg.io', namespace: 'db', name: `b-${id}` },
+ source: 'issue',
+})
+
+function render(node: ReactNode) {
+ const host = document.createElement('div')
+ document.body.appendChild(host)
+ const root = createRoot(host)
+ act(() => root.render(node))
+ return root
+}
+
+afterEach(() => {
+ document.body.innerHTML = ''
+})
+
+describe('CNPGClusterSummary', () => {
+ it('keeps flat sections by default and frames only At a glance and About when asked', () => {
+ for (const framed of [false, true]) {
+ const root = render(Maintenance notice} operationalFacts={Base backup progress
} stateFacts={Existing host facts
} />)
+ const headings = [...document.querySelectorAll('h3')].filter((h) => ['At a glance', 'About'].includes(h.textContent!))
+ expect(headings).toHaveLength(2)
+ for (const heading of headings) expect(heading.closest('section') !== null).toBe(framed)
+ const glance = headings[0].parentElement!.parentElement!
+ const about = headings[1].parentElement!.parentElement!
+ if (framed) {
+ expect(glance.textContent).toContain('Base backup progress')
+ expect(about.textContent).not.toContain('Base backup progress')
+ expect(about.textContent).toContain('Existing host facts')
+ expect(glance.textContent).not.toContain('Existing host facts')
+ expect(glance.textContent).not.toContain('Maintenance notice')
+ expect(glance.textContent).not.toContain('Backup failed')
+ }
+ expect(document.body.textContent!.indexOf('Base backup progress')).toBeLessThan(document.body.textContent!.indexOf('About'))
+ expect(document.body.textContent!.indexOf('Existing host facts')).toBeGreaterThan(document.body.textContent!.indexOf('About'))
+ act(() => root.unmount())
+ }
+ const root = render( )
+ expect(document.querySelector('section')).toBeNull()
+ act(() => root.unmount())
+ })
+ it('opens the other problems in place when the host links nowhere else', () => {
+ const onNavigate = vi.fn()
+ const r = row({ problems: [problem('a', 'critical', 'WAL archiving failing'), problem('b', 'warning', 'Backup failed'), problem('c', 'posture', 'No schedule')], attention: true })
+ const root = render( )
+ const more = [...document.querySelectorAll('button')].find((b) => b.textContent?.includes('+2 more'))!
+ expect(more.getAttribute('aria-expanded')).toBe('false')
+ act(() => more.click())
+ expect(more.getAttribute('aria-expanded')).toBe('true')
+ expect(document.body.textContent).toContain('Backup failed')
+ expect(document.body.textContent).toContain('No schedule')
+ act(() => [...document.querySelectorAll('button')].find((b) => b.textContent === 'b-b')!.click())
+ expect(onNavigate).toHaveBeenCalledWith(expect.objectContaining({ kind: 'Backup', name: 'b-b' }))
+ act(() => root.unmount())
+ })
+ it('can open with every problem listed', () => {
+ const r = row({ problems: [problem('a', 'critical', 'WAL archiving failing'), problem('b', 'warning', 'Backup failed')], attention: true })
+ const root = render( )
+ const more = [...document.querySelectorAll('button')].find((b) => b.getAttribute('aria-expanded') !== null)!
+ expect(more.getAttribute('aria-expanded')).toBe('true')
+ act(() => root.unmount())
+ })
+ it('keeps a host link as the override', () => {
+ const r = row({ problems: [problem('a', 'warning', 'One'), problem('b', 'warning', 'Two')], attention: true })
+ const root = render( All {n} } />)
+ expect(document.body.textContent).toContain('All 2')
+ expect([...document.querySelectorAll('button')].some((b) => b.textContent?.includes('+1 more'))).toBe(false)
+ act(() => root.unmount())
+ })
+ it('marks a tab only when its dimension needs a look or could not be assessed', () => {
+ const dim = (tone: CNPGDimension['tone']): CNPGDimension => ({ id: 'protection', label: 'Backups', tone, text: 'verdict', source: 's' })
+ for (const tone of ['healthy', 'neutral'] as const) {
+ const root = render( )
+ expect(document.querySelector('[aria-label="Backups: verdict"]')).toBeNull()
+ act(() => root.unmount())
+ }
+ for (const tone of ['degraded', 'unhealthy', 'unknown'] as const) {
+ const root = render( )
+ expect(document.querySelector('[aria-label="Backups: verdict"]')).not.toBeNull()
+ act(() => root.unmount())
+ }
+ })
+ it("says on the tab why it is marked, and nothing when it is not", () => {
+ const backups = (tone: CNPGDimension['tone']): CNPGDimension => ({ id: 'protection', label: 'Backups', tone, text: 'no backup destination', source: 'Cluster spec' })
+ let root = render( )
+ expect(document.body.textContent).toContain('no backup destination')
+ expect(document.body.textContent).toContain('Cluster spec')
+ act(() => root.unmount())
+ root = render( )
+ expect(document.body.textContent).toBe('')
+ act(() => root.unmount())
+ })
+ it('draws an unassessed dimension as a ring, never a calm dot', () => {
+ const root = render( )
+ const mark = document.querySelector('[aria-label="Storage: unassessed"]')!
+ expect(mark.querySelector('.rounded-full.border')).not.toBeNull()
+ expect(mark.querySelector('.rounded-full:not(.border)')).toBeNull()
+ act(() => root.unmount())
+ })
+ it('makes the Serving status a button only when the host can open its details', () => {
+ const serving: CNPGDimension = { id: 'serving', label: 'Serving', tone: 'healthy', text: 'primary ready', source: 's' }
+ const onSelect = vi.fn()
+ let root = render( )
+ expect(document.querySelector('button')).toBeNull()
+ expect(document.body.textContent).toContain('primary ready')
+ act(() => root.unmount())
+ root = render( )
+ act(() => document.querySelector('[aria-label="Serving: primary ready. Open its details"]')!.click())
+ expect(onSelect).toHaveBeenCalled()
+ act(() => root.unmount())
+ })
+ it('lists each dimension at a glance, opening its tab', () => {
+ const dims: CNPGDimension[] = [{ id: 'storage', label: 'Storage', tone: 'degraded', text: 'WAL held by an inactive slot', source: 'slot' }]
+ const onSelect = vi.fn()
+ const root = render( )
+ expect(document.body.textContent).toContain('WAL held by an inactive slot')
+ const open = [...document.querySelectorAll('button')].find((b) => b.textContent === 'Storage →')!
+ act(() => open.click())
+ expect(onSelect).toHaveBeenCalledWith('storage')
+ act(() => root.unmount())
+ })
+})
+
+describe('problem meta', () => {
+ it('keeps a space between the Backup and "and N more"', () => {
+ const r = row({
+ problems: [{ ...problem('a', 'warning', '3 backups failed'), alsoAbout: [{ kind: 'Backup', name: 'b-2' }, { kind: 'Backup', name: 'b-1' }] }],
+ attention: true,
+ })
+ const root = render( {}} />)
+ expect(document.body.textContent).toContain('b-a and 2 more')
+ act(() => root.unmount())
+ })
+})
+
+describe('problem provenance', () => {
+ it('names the origin instead of "Radar issue" and links to Issues when the host can', () => {
+ const open = vi.fn()
+ const r = row({
+ problems: [{ ...problem('a', 'critical', 'WAL archiving is failing'), subject: { kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: 'db', name: 'pg' }, origin: { label: 'Reported by CNPG', detail: 'ContinuousArchiving condition' } }],
+ attention: true,
+ })
+ const root = render(
+
+
+ ,
+ )
+ expect(document.body.textContent).toContain('Reported by CNPG')
+ expect(document.body.textContent).not.toContain('Radar issue')
+ act(() => [...document.querySelectorAll('button')].find((b) => b.textContent === 'See in Issues →')!.click())
+ expect(open).toHaveBeenCalledWith(expect.objectContaining({ id: 'a' }))
+ act(() => root.unmount())
+ })
+ it('offers no Issues link without a host handler', () => {
+ const r = row({ problems: [problem('a', 'warning', 'Backup failed')], attention: true })
+ const root = render( )
+ expect(document.body.textContent).toContain('Detected by Radar')
+ expect(document.body.textContent).not.toContain('See in Issues')
+ act(() => root.unmount())
+ })
+})
+
+it('shows the declaration read limitation inline on Overview', () => {
+ const r = row({ declarations: { total: 1, failed: 0, pending: 0, summary: { text: '≥1 reconciled; Databases not read', tone: 'unknown' } } })
+ const root = render( )
+ expect(document.body.textContent).toContain('≥1 reconciled; Databases not read')
+ act(() => root.unmount())
+})
+
+it('shows neutral literal operator phases for every CNPG caller', () => {
+ const r = row({ cluster: { status: { phase: 'Cluster in healthy state' } } })
+ const root = render( )
+ expect(document.body.textContent).toContain('Cluster in healthy state')
+ const badge = [...document.querySelectorAll('span')].find((el) => el.textContent === 'Cluster in healthy state')!
+ expect(badge.className).not.toContain('emerald')
+ act(() => root.unmount())
+ const expanded = render( )
+ expect(document.body.textContent).toContain('Cluster in healthy state')
+ expect(document.body.textContent).not.toContain('Healthy')
+ act(() => expanded.unmount())
+})
+it('shows healthy Serving only when requested and keeps its tab mark absent', () => {
+ const dimension: CNPGDimension = { id: 'serving', label: 'Serving', tone: 'healthy', text: 'primary ready', source: 'Pod pg-1 and Service pg-rw' }
+ const root = render(<> >)
+ expect(document.body.textContent).toContain('primary ready')
+ expect(document.querySelector('[role="img"]')).toBeNull()
+ act(() => root.unmount())
+ const old = render( )
+ expect(document.body.textContent).not.toContain('primary ready')
+ act(() => old.unmount())
+})
+
+it('keeps literal blocked phases neutral with their explanation and Operator action', () => {
+ const phase = 'Cluster cannot proceed to reconciliation due to an unknown plugin being required'
+ const r = row({ cluster: { status: { phase } }, controllerStatus: { text: 'Unknown Plugin', level: 'unhealthy' } })
+ const open = vi.fn()
+ const root = render( )
+ expect(document.body.textContent).toContain(phase)
+ expect(document.body.textContent).toContain('reported by CNPG')
+ expect(document.body.textContent).toContain('Plugins this cluster uses')
+ const badge = [...document.querySelectorAll('span')].find((el) => el.textContent === phase)!
+ expect(badge.className).not.toMatch(/red|amber|emerald/)
+ act(() => [...document.querySelectorAll('button')].find((b) => b.textContent === 'Operator and plugins →')!.click())
+ expect(open).toHaveBeenCalled()
+ act(() => root.unmount())
+})
+
+it('separates absent operator readiness from observed zero ready Pods', () => {
+ const root = render( )
+ expect(document.body.textContent).toContain('Not reported by the operator')
+ expect(document.body.textContent).toContain('0 of 1 instance Pods ready')
+ expect(document.body.textContent).not.toContain('—/1')
+ act(() => root.unmount())
+})
+
+it('compacts only the scheduler disclosure inside the CNPG problem callout', () => {
+ const root = render( )
+ const disclosure = [...document.body.querySelectorAll('button')].find((b) => b.textContent?.includes('Scheduler message'))!
+ const callout = [...document.body.querySelectorAll('div')].find((d) => d.classList.contains('[&_.mt-5]:mt-2'))
+ expect(callout?.contains(disclosure)).toBe(true)
+ expect(disclosure.getAttribute('aria-expanded')).toBe('false')
+ act(() => root.unmount())
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.tsx
index a35db6a9bb..7d19d78434 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGClusterSummary.tsx
@@ -1,23 +1,32 @@
-import type { ReactNode } from 'react'
+import { Fragment, useState, type ReactNode } from 'react'
import { clsx } from 'clsx'
import { Badge } from '../ui/Badge'
import { Tooltip } from '../ui/Tooltip'
-import { CNPG_BARMAN_OBJECTSTORE_GROUP, CNPG_GROUP } from '../resources/resource-utils-cnpg'
-import type { CNPGFleetRow, CNPGInstance } from './workspace'
-import {
- FactGrid,
- FactRow,
- FactSource,
- FactValue,
- ProblemCallout,
- RefLink,
- SummaryHeading,
- ToneDot,
- toneTextClass,
- CNPG_PRIMARY_BUTTON,
- CNPG_SECONDARY_BUTTON,
- type CNPGNavigate,
-} from './primitives'
+import { Collapse, CollapseChevron, useDisclosure } from '../ui/Collapse'
+import { classifyCNPGClusterPhase, cnpgBlockedPhaseExplanation, CNPG_BARMAN_OBJECTSTORE_GROUP, CNPG_GROUP } from '../resources/resource-utils-cnpg'
+import { cnpgClusterPlugins, cnpgPluginPhase, cnpgReadyInstances, type CNPGFleetRow, type CNPGInstance } from './workspace'
+import type { CNPGDimension } from './ha'
+import { PrimaryConflictNote } from './primitives'
+import { Note } from './CNPGSharedSummary'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { StatusDot, toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactSource, FactValue, ManagedByText, managedByLabel } from '../facts'
+import { ProblemCallout, ProblemList } from '../problems'
+import { formatAge } from '../resources/resource-utils'
+import { FoldSection, SectionHeading } from '../ui/FoldSection'
+
+function ReadyCount({ row }: { row: CNPGFleetRow }) {
+ const r = cnpgReadyInstances(row)
+ if (row.instances.ready === null) return <>{r.text}{r.podText && {r.podText} }>
+ if (!r.note) return <>{r.text} ready>
+ return (
+
+
+ {r.text} Pods ready status says {row.instances.ready}
+
+
+ )
+}
export interface CNPGSummaryAction {
label: string
@@ -25,10 +34,10 @@ export interface CNPGSummaryAction {
primary?: boolean
}
-function InstancePill({ pod, namespace, onNavigate }: { pod: CNPGInstance; namespace: string; onNavigate?: CNPGNavigate }) {
+function InstancePill({ pod, namespace, onNavigate }: { pod: CNPGInstance; namespace: string; onNavigate?: NavigateToRef }) {
const tone = pod.ready === true ? 'healthy' : pod.ready === false ? 'unhealthy' : 'unknown'
const role = pod.role === 'primary' ? 'Primary' : pod.role === 'replica' ? 'Replica' : 'Role unknown'
- const readiness = pod.ready === true ? 'Ready' : pod.ready === false ? 'Not ready' : 'Readiness unknown'
+ const readiness = pod.ready === true ? 'Pod ready' : pod.ready === false ? 'Pod not ready' : 'Pod readiness unknown'
return (
-
+
{pod.name}
{pod.role === 'primary' ? 'P' : pod.role === 'replica' ? 'R' : '?'}
@@ -47,122 +56,312 @@ function InstancePill({ pod, namespace, onNavigate }: { pod: CNPGInstance; names
)
}
+/**
+ * One dimension's state as a mark beside the tab that explains it: a dot when
+ * something needs a look, a hollow ring when it could not be assessed (never a
+ * calm colour), nothing when it is fine. The verdict is in the tooltip; the
+ * Overview's At a glance has it in words.
+ */
+export function CNPGDimensionMark({ dimension }: { dimension: CNPGDimension }) {
+ if (!dimensionMarked(dimension)) return null
+ const label = `${dimension.label}: ${dimension.text}`
+ return (
+ {label}
{dimension.source}
>} position="bottom">
+
+
+
+
+ )
+}
+
+function dimensionMarked(dimension: CNPGDimension): boolean {
+ return dimension.tone !== 'healthy' && dimension.tone !== 'neutral'
+}
+
+function DimensionGlyph({ tone }: { tone: CNPGDimension['tone'] }) {
+ return tone === 'unknown' ? (
+
+ ) : (
+
+ )
+}
+
+/**
+ * The verdict behind a tab's mark, as the tab's first line, so a mark always
+ * points at words on the tab it marks. Nothing when the dimension is fine.
+ */
+export function CNPGDimensionVerdict({ dimension, className, alwaysShow = false }: { dimension: CNPGDimension; className?: string; alwaysShow?: boolean }) {
+ if (!alwaysShow && !dimensionMarked(dimension)) return null
+ return (
+
+
+
+
+ {dimension.label}
+ {dimension.text}
+ {dimension.source}
+
+ )
+}
+
+/** Whether the cluster serves writes, on its title line: the headline the tabs do not carry. */
+export function CNPGServingStatus({ dimension, onSelect }: { dimension: CNPGDimension; onSelect?: () => void }) {
+ const body = (
+ <>
+
+ {dimension.label}
+ {dimension.text}
+ >
+ )
+ const className = 'inline-flex items-center gap-1.5 whitespace-nowrap text-sm'
+ return (
+
+ {onSelect ? (
+
+ {body}
+
+ ) : (
+ {body}
+ )}
+
+ )
+}
+
+/** "+N more" that opens the rest of the problems in place, when the host links nowhere else. */
+function MoreProblems({ count, open, onToggle, panelId }: { count: number; open: boolean; onToggle: () => void; panelId: string }) {
+ return (
+
+
+ {open ? 'Hide' : `+${count} more`}
+
+ )
+}
+
export function CNPGClusterSummary({
row,
onNavigate,
actions,
problemsLink,
extra,
+ lead,
+ dimensions,
+ onSelectDimension,
+ initialProblemsExpanded = false,
+ stateFacts,
+ operationalFacts,
+ dimensionLinkLabel,
+ onOpenOperator,
+ framed = false,
}: {
row: CNPGFleetRow
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
actions?: CNPGSummaryAction[]
/** Link to the complete list of this cluster's findings, shown when more than one exists. */
problemsLink?: (count: number) => ReactNode
extra?: ReactNode
+ /** Rendered first, above the problem callout: standing states such as maintenance mode. */
+ lead?: ReactNode
+ /** Serving · Replication · Storage · Backups, each from its own source (see cnpgDimensions); the At a glance rows. */
+ dimensions?: CNPGDimension[]
+ /** Open with every problem listed below the callout (e.g. arriving from the fleet's "+N more"). */
+ initialProblemsExpanded?: boolean
+ /** Makes each dimension row open where that dimension is explained (its tab). */
+ onSelectDimension?: (id: CNPGDimension['id']) => void
+ /** The name of the place onSelectDimension opens, for the row's link; the dimension's own label when unset. */
+ dimensionLinkLabel?: (id: CNPGDimension['id']) => string
+ /** Opens the operator's own diagnosis, offered beside a controller phase that is not healthy. */
+ onOpenOperator?: () => void
+ /** Extra FactRows appended to About. */
+ stateFacts?: ReactNode
+ /** Extra FactRows appended to At a glance, e.g. live instance operations. */
+ operationalFacts?: ReactNode
+ /** Cards for the full page; drawers keep flat sections. */
+ framed?: boolean
}) {
const top = row.problems[0]
const rest = row.problems.length - 1
- const p = row.protection
const ns = row.namespace
const radarFindings = row.problems.some((x) => x.severity !== 'posture')
+ const phase = typeof row.cluster?.status?.phase === 'string' ? row.cluster.status.phase : ''
+ const blocked = classifyCNPGClusterPhase(phase) === 'terminal' ? cnpgBlockedPhaseExplanation(phase, row.cluster?.status?.phaseReason) : null
+ const [showRest, setShowRest] = useState(initialProblemsExpanded)
+ const restDisclosure = useDisclosure(showRest)
+ const Frame = framed ? 'section' : Fragment
+ const frameProps = framed ? { className: 'mb-4 last:mb-0 rounded-xl border border-theme-border bg-theme-surface px-4 py-3 shadow-theme-sm' } : {}
return (
+ {lead}
{top && (
-
0 ? problemsLink?.(row.problems.length) ?? +{rest} more : null}
- />
+ more={
+ rest > 0
+ ? problemsLink?.(row.problems.length) ?? (
+ setShowRest((v) => !v)} panelId={restDisclosure.panelId} />
+ )
+ : null
+ }
+ />
+ )}
+ {top && rest > 0 && !problemsLink && (
+
+
+
)}
{actions && actions.length > 0 && (
{actions.map((a) => (
-
+
{a.label}
))}
)}
- State
-
-
-
-
- {row.controllerStatus.text}
-
-
- {radarFindings ? 'reported by CNPG · Radar findings above are separate' : 'reported by CNPG'}
-
-
-
-
-
-
- {row.instances.ready ?? '–'}/{row.instances.desired ?? '–'} ready
- {row.cluster?.status?.currentPrimary && (
- · primary {row.cluster.status.currentPrimary}
+
+ At a glance
+
+ {dimensions?.map((d) => (
+
+ onSelectDimension(d.id) : undefined} />
+
+ ))}
+
+
+
+
+ {row.cluster?.status?.currentPrimary && !row.primaryConflict && (
+ · primary {row.cluster.status.currentPrimary}
+ )}
+
+ {row.primaryConflict &&
}
+ {row.pods.length > 0 && (
+
+ {row.pods.map((pod) => (
+
+ ))}
+
+ )}
+
+
+
+
+
+ {phase || 'Not reported'}
+
+
+ {radarFindings ? 'reported by CNPG · Radar findings above are separate' : 'reported by CNPG'}
+
+ {onOpenOperator && (row.controllerStatus.level === 'unhealthy' || row.controllerStatus.level === 'degraded') && (
+
+ Operator and plugins →
+
)}
- {row.pods.length > 0 && (
-
- {row.pods.map((pod) => (
-
- ))}
+ {blocked && (
+
+ {blocked.body}
+ {cnpgPluginPhase(row.cluster) &&
+ ` Plugins this cluster uses: ${cnpgClusterPlugins(row.cluster).join(', ') || 'none listed'}. The Operator view shows whether each is running and when it last restarted.`}
)}
-
-
-
-
-
-
- {row.replicaCluster && (
-
- Follows {row.replicaCluster.source ? {row.replicaCluster.source} : 'an external primary'}
- )}
-
- {row.pgVersion ?? 'Unknown'}
- {row.catalog && (
-
- {' · '}
-
- {row.catalog.name}
-
-
- )}
-
-
-
-
-
- {row.poolers.length === 0 ? (
-
- {row.poolersKnown ? 'None' : 'No access to Poolers'}
-
- ) : (
-
- {row.poolers.map((name) => (
-
- ))}
-
+ {operationalFacts}
+
+
+
+
+ About
+
+ {row.replicaCluster && (
+
+ Follows {row.replicaCluster.source ? {row.replicaCluster.source} : 'an external primary'}
+
)}
-
- {row.gitops && (
-
- {row.gitops.tool === 'argocd' ? 'Argo CD' : 'Flux'} {row.gitops.name}
+
+ {row.pgVersion ?? 'Unknown'}
+ {row.catalog && (
+
+ {' · '}
+
+ {row.catalog.name}
+
+
+ )}
+
+
+
+
+ {row.poolers.length === 0 ? (
+
+ {row.poolersKnown ? 'None' : 'No access to Poolers'}
+
+ ) : (
+
+ {row.poolers.map((name) => (
+
+ ))}
+
+ )}
+
+ {row.managedBy && managedByLabel(row.managedBy) && (
+
+
+
+ )}
+ {stateFacts}
+
+
+
+ {extra}
+
+ )
+}
+
+function DimensionValue({ dimension: d, linkLabel, onOpen }: { dimension: CNPGDimension; linkLabel: string; onOpen?: () => void }) {
+ return (
+
+
+
+ {d.text}
+ {onOpen && (
+
+ {linkLabel} →
+
)}
-
+
+ {d.source &&
{d.source}
}
+
+ )
+}
- Protection
+/** A cluster's recovery evidence as facts: schedule, destination, newest backup, WAL archiving, recovery window and restore validation. */
+export function CNPGClusterBackupFacts({ row, onNavigate }: { row: CNPGFleetRow; onNavigate?: NavigateToRef }) {
+ const p = row.protection
+ const ns = row.namespace
+ return (
+ <>
@@ -188,7 +387,7 @@ export function CNPGClusterSummary({
-
+
{p.recoveryWindow.from ? (
@@ -216,7 +415,23 @@ export function CNPGClusterSummary({
- {extra}
-
+
+ >
)
}
+
+export function CNPGWALArchivingFact({ fact, compact = false }: { fact: CNPGFleetRow['protection']['walArchiving']; compact?: boolean }) {
+ const c = fact.operatorCondition
+ return
+
+
+ {fact.detail &&
{compact && fact.state === 'no_destination' ? 'No point-in-time recovery' : fact.detail} }
+ {c &&
e.stopPropagation()}>
+
+ {c.type}: {c.status}
+ {c.message || 'No message reported'}
+ Last transition: {c.lastTransitionTime ? {formatAge(c.lastTransitionTime)} ago : 'Not reported'}
+
+
}
+
+}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.test.tsx
new file mode 100644
index 0000000000..cbe3e234c9
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.test.tsx
@@ -0,0 +1,64 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot } from 'react-dom/client'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { expect, it, vi } from 'vitest'
+import { CNPGConnectSection } from './CNPGConnectSection'
+
+const cluster = { metadata: { name: 'orders', namespace: 'radar-cnpg-prod' }, spec: { bootstrap: { initdb: { database: 'appdb', owner: 'app' } } } }
+;(globalThis as any).IS_REACT_ACT_ENVIRONMENT = true
+
+it('pairs a namespace-qualified port-forward with matching loopback templates and copies each', async () => {
+ const writeText = vi.fn().mockResolvedValue(undefined)
+ Object.defineProperty(navigator, 'clipboard', { value: { writeText }, configurable: true })
+ const host = document.createElement('div'); const root = createRoot(host)
+ act(() => root.render( ))
+ expect(host.textContent).toContain('Inside Kubernetes')
+ expect(host.textContent).toContain('From this computer')
+ for (const [label, expected] of [
+ ['port-forward command', 'kubectl -n radar-cnpg-prod port-forward service/orders-rw 5432:5432'],
+ ['local psql command', 'psql -h 127.0.0.1 -p 5432 -U app -d appdb'],
+ ['local connection string', 'postgresql://app:@127.0.0.1:5432/appdb'],
+ ]) {
+ await act(async () => host.querySelector(`[aria-label="Copy ${label}"]`)!.click())
+ expect(writeText).toHaveBeenLastCalledWith(expected)
+ }
+ act(() => root.unmount())
+})
+
+it('uses one three-column grid with every explanation under its own host', () => {
+ const html = renderToStaticMarkup( {}} />)
+ const host = document.createElement('div'); host.innerHTML = html
+ const list = host.querySelector('ul')!
+ expect(list.className).toContain('grid-cols-[5rem_minmax(0,1fr)_auto]')
+ expect(list.querySelectorAll('li')).toHaveLength(4)
+ for (const row of list.querySelectorAll('li')) {
+ expect(row.className).toBe('contents')
+ expect(row.children).toHaveLength(3)
+ expect(row.children[1].className).toContain('[overflow-wrap:anywhere]')
+ expect(row.children[1].querySelector('div')).not.toBeNull()
+ expect(row.children[2].textContent).toContain('Reachability')
+ }
+})
+
+it('shows command context certainty and per-service availability beside each endpoint', () => {
+ const host = document.createElement('div')
+ host.innerHTML = renderToStaticMarkup( )
+ expect(host.textContent).toContain('kubectl --context kind-orders')
+ expect(host.textContent).not.toContain('uses your current kubectl context')
+ expect(host.textContent).toContain('Context as named in your kubeconfig (orders-config); kubectl must read the same kubeconfig file.')
+ expect([...host.querySelectorAll('li')].map((li) => li.textContent)).toEqual([expect.stringContaining('Ready endpoints'), expect.stringContaining('Unavailable: no ready standby'), expect.stringContaining('Ready instance observed')])
+ host.innerHTML = renderToStaticMarkup( )
+ expect(host.textContent).toContain('uses your current kubectl context')
+ expect(host.textContent).not.toContain('--context')
+ expect([...host.querySelectorAll('li')].every((li) => li.textContent?.includes('Not checked'))).toBe(true)
+})
+
+it('keeps the unavailable reason beside each unchecked Service', () => {
+ const host = document.createElement('div')
+ host.innerHTML = renderToStaticMarkup( )
+ for (const li of host.querySelectorAll('li')) expect(li.textContent).toContain('Not checked · Availability could not be read: timeout')
+ host.innerHTML = renderToStaticMarkup( )
+ expect(host.querySelector('li')!.textContent).toContain('needs list endpointslices (discovery.k8s.io) in namespace radar-cnpg-prod')
+ expect(host.querySelectorAll('li')[1].textContent).toContain('needs list pods in namespace radar-cnpg-prod')
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.tsx b/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.tsx
new file mode 100644
index 0000000000..53a774d274
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGConnectSection.tsx
@@ -0,0 +1,179 @@
+import type { CNPGClusterHA } from './ha'
+import { useState } from 'react'
+import { Check, Copy } from 'lucide-react'
+import { Tooltip } from '../ui/Tooltip'
+import { CNPG_GROUP } from '../resources/resource-utils-cnpg'
+import { CNPG_DEFAULT_PORT, cnpgEndpointAvailability, cnpgConnectionURI, cnpgConnectInfo, cnpgPsqlCommand, cnpgPortForwardCommand, type CNPGConnectEndpoint } from './connect'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { FactGrid, FactRow } from '../facts'
+import { SectionHeading } from '../ui/FoldSection'
+
+function CopyButton({ text, label }: { text: string; label: string }) {
+ const [copied, setCopied] = useState(false)
+ const copy = () => {
+ navigator.clipboard?.writeText(text).then(
+ () => {
+ setCopied(true)
+ setTimeout(() => setCopied(false), 2000)
+ },
+ () => {},
+ )
+ }
+ return (
+
+
+ {copied ? : }
+
+
+ )
+}
+
+function Snippet({ text, label }: { text: string; label: string }) {
+ return (
+
+ {text}
+
+
+ )
+}
+
+const ROLE_LABEL: Record = {
+ rw: 'Read-write',
+ ro: 'Read-only',
+ r: 'Any instance',
+ pooler: 'Pooler',
+ additional: 'Additional',
+}
+
+/**
+ * How applications reach the cluster: Services, application database and
+ * owner, and the credentials Secret by name. Never reads the Secret.
+ */
+
+/**
+ * Selects the Connect heading inside one summary. A data attribute, not an
+ * id: a drawer summary can sit over a page summary of the same kind, so a host
+ * scrolls to it within its own summary's element.
+ */
+export function CNPGConnectSection({
+ cluster,
+ poolers,
+ poolersKnown,
+ onNavigate,
+ onOpenReachability,
+ showHeading = true,
+ kubeconfigContext,
+ kubeconfigSource,
+ ha,
+ haUnavailableReason,
+}: {
+ cluster: any
+ kubeconfigContext?: string
+ kubeconfigSource?: string
+ ha?: CNPGClusterHA
+ haUnavailableReason?: string
+ poolers?: any[]
+ /** False when Poolers could not be listed, so a Pooler may exist that is not shown. */
+ poolersKnown?: boolean
+ onNavigate?: NavigateToRef
+ /** Opens a host's Service on its Reachability tab; the link shows only when given. */
+ onOpenReachability?: (service: { namespace: string; name: string }) => void
+ /** False where the host already titles it (e.g. the Connect dialog). */
+ showHeading?: boolean
+}) {
+ const info = cnpgConnectInfo(cluster, poolers)
+ const ns: string = cluster?.metadata?.namespace ?? ''
+ const primary = info.endpoints[0]
+ const local = { ...primary, host: '127.0.0.1', port: CNPG_DEFAULT_PORT }
+ return (
+ <>
+ {showHeading && Connect }
+
+
+
+ {info.endpoints.map((ep) => {
+ const availability = cnpgEndpointAvailability(ep, ha, haUnavailableReason)
+ return (
+
+ {ROLE_LABEL[ep.role]}
+
+
+ {`${ep.host}:${ep.port}`}
+
+
{ep.selects}
+
{availability.text}{availability.source ? ` · ${availability.source}` : ''}
+ {ep.portFromTemplate &&
port from serviceTemplate
}
+
+ {onOpenReachability ? (
+
+ onOpenReachability({ namespace: ns, name: ep.name })} className="text-xs text-accent-text hover:underline">
+ Reachability →
+
+
+ ) : }
+
+ )
+ })}
+
+ {info.disabled.length > 0 && (
+
+ Disabled in spec.managed.services: {info.disabled.map((t) => `${cluster?.metadata?.name}-${t}`).join(', ')}
+
+ )}
+ {poolersKnown === false && No access to Poolers: one may also front this cluster.
}
+ {info.replicaCluster && (
+ A replica cluster: every Service reaches instances that only replay until it is promoted.
+ )}
+
+
+ {info.database.value ? {info.database.value} : Unknown }
+ {info.database.source}
+
+
+ {info.owner.value ? {info.owner.value} : Unknown }
+ {info.owner.source}
+
+
+
+
+ {info.secret.name}
+
+
+
+ {info.secret.source} · Radar does not read it; the password is its password key
+
+
+ {info.database.value && (
+
+
+
+
+
+
+ Via
{primary.name} . Replace
{''} ; psql prompts for it.
+
+
+ )}
+ {info.database.value && (
+
+
+
+
{kubeconfigContext ? `Context as named in your kubeconfig${kubeconfigSource ? ` (${kubeconfigSource})` : ''}; kubectl must read the same kubeconfig file.` : 'uses your current kubectl context'}
+
+
+
+ Keep port-forward running, then connect in another terminal. Local port 5432 must be free. Replace {'
'}; psql prompts for it.
+
+ )}
+
+ >
+ )
+}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGDeclarativeSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGDeclarativeSummary.tsx
index dc667382d1..5786604be3 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGDeclarativeSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGDeclarativeSummary.tsx
@@ -1,13 +1,15 @@
import type { ReactNode } from 'react'
import { getCNPGDeclarativeMessage, getCNPGReclaimPolicy } from '../resources/resource-utils-cnpg'
-import type { CNPGWorkspaceResponse } from './workspace'
-import { FactGrid, FactRow, FactValue, RefLink, SummaryHeading, toneTextClass, type CNPGNavigate } from './primitives'
-import { ClusterLink, NotReported, ObjectProblems, SummaryShell } from './CNPGSharedSummary'
+import { cnpgManagedBy, type CNPGWorkspaceResponse } from './workspace'
+import { type Fact } from '../facts'
+import { cnpgLogicalPaths, type CNPGLogicalPath } from './logicalReplication'
+import { CNPGLogicalPathView } from './CNPGLogicalPath'
+import { cnpgDatabaseRoleFacts } from './databaseRole'
+import { ClusterLink, NotReported, Note, ObjectProblems, SummaryShell } from './CNPGSharedSummary'
import {
appliedFact,
clustersIn,
databaseForDeclaration,
- gitopsSourceOf,
missingManagedRole,
observedGenerationFact,
refOf,
@@ -16,11 +18,15 @@ import {
targetCluster,
workspaceList,
} from './relations'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactValue, ManagedByText, managedByLabel } from '../facts'
+import { SectionHeading } from '../ui/FoldSection'
interface SummaryProps {
resource: any
workspace: CNPGWorkspaceResponse | null
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
}
function ReclaimRow({ resource }: { resource: any }) {
@@ -41,7 +47,7 @@ function Reconciled({ resource, extra }: { resource: any; extra?: ReactNode }) {
const applied = appliedFact(resource)
return (
<>
- Reconciled
+ Reconciled
@@ -62,14 +68,10 @@ function Reconciled({ resource, extra }: { resource: any; extra?: ReactNode }) {
)
}
-function DeclaredIn({ resource }: { resource: any }) {
- const src = gitopsSourceOf(resource)
- if (!src) return
- return (
-
- {src.tool === 'argocd' ? 'Argo CD application' : 'Flux'} {src.namespace ? `${src.namespace}/${src.name}` : src.name}
-
- )
+function DeclaredIn({ resource, workspace, onNavigate }: { resource: any; workspace: CNPGWorkspaceResponse | null; onNavigate?: NavigateToRef }) {
+ const manager = cnpgManagedBy(workspace, resource)
+ if (!manager || !managedByLabel(manager)) return
+ return
}
function DatabaseRef({ resource, workspace, onNavigate }: SummaryProps) {
@@ -92,7 +94,7 @@ function DatabaseRef({ resource, workspace, onNavigate }: SummaryProps) {
)
}
-function LinkList({ items, kind, onNavigate }: { items: any[]; kind: string; onNavigate?: CNPGNavigate }) {
+function LinkList({ items, kind, onNavigate }: { items: any[]; kind: string; onNavigate?: NavigateToRef }) {
return (
{items.map((o) => (
@@ -114,7 +116,7 @@ export function CNPGDatabaseSummary({ resource, workspace, onNavigate }: Summary
- Declared
+ Declared
{resource?.spec?.name ? {resource.spec.name} : }
@@ -137,10 +139,10 @@ export function CNPGDatabaseSummary({ resource, workspace, onNavigate }: Summary
}
/>
- Source and target
+ Source and target
-
+
@@ -149,7 +151,7 @@ export function CNPGDatabaseSummary({ resource, workspace, onNavigate }: Summary
{pubsUnavailable ? (
) : related.publications.length === 0 ? (
- None on this database
+ No visible Publication declarations
) : (
)}
@@ -158,12 +160,13 @@ export function CNPGDatabaseSummary({ resource, workspace, onNavigate }: Summary
{subsUnavailable ? (
) : related.subscriptions.length === 0 ? (
- None on this database
+ No visible Subscription declarations
) : (
)}
+ Objects created in SQL are not shown.
)
}
@@ -191,12 +194,39 @@ function publicationTargets(resource: any): ReactNode {
)
}
-export function CNPGPublicationSummary({ resource, workspace, onNavigate }: SummaryProps) {
+function workspacePaths(workspace: CNPGWorkspaceResponse | null | undefined, subscriptions: any[]): CNPGLogicalPath[] {
+ return cnpgLogicalPaths(subscriptions, clustersIn(workspace), workspaceList(workspace, 'publications'), workspaceList(workspace, 'poolers'), (ns) =>
+ relationUnavailable(workspace, 'publications', ns, 'Publications'),
+ )
+}
+
+export interface CNPGLogicalPathReading {
+ path: CNPGLogicalPath
+ /** The publisher primary's report of the slot; absent when not read. */
+ slot?: Fact
+ notice?: ReactNode
+}
+
+export function CNPGPublicationSummary({
+ resource,
+ workspace,
+ onNavigate,
+ subscribers,
+}: SummaryProps & {
+ /** Subscriptions reading this publication, with their slots; derived from the workspace when omitted. */
+ subscribers?: CNPGLogicalPathReading[]
+}) {
+ const readings: CNPGLogicalPathReading[] =
+ subscribers ??
+ workspacePaths(workspace, workspaceList(workspace, 'subscriptions'))
+ .filter((p) => p.publication.object?.namespace === resource?.metadata?.namespace && p.publication.object?.name === resource?.metadata?.name)
+ .map((path) => ({ path }))
+ const subsUnavailable = relationUnavailable(workspace, 'subscriptions', resource?.metadata?.namespace ?? '', 'Subscriptions')
return (
- Declared
+ Declared
{resource?.spec?.name ? {resource.spec.name} : }
@@ -210,23 +240,117 @@ export function CNPGPublicationSummary({ resource, workspace, onNavigate }: Summ
{publicationTargets(resource)}
-
+
+
+ Subscribers
+ {readings.length === 0 ? (
+
+ {subsUnavailable ?? 'No visible Subscription object reads this publication. Subscribers outside Radar\'s view, or created in SQL, are not listed.'}
+
+ ) : (
+
+ {readings.map((r) => (
+
+ ))}
+
+ )}
)
}
-export function CNPGSubscriptionSummary({ resource, workspace, onNavigate }: SummaryProps) {
+export function CNPGDatabaseRoleSummary({ resource, workspace, onNavigate }: SummaryProps) {
+ const ns = resource?.metadata?.namespace ?? ''
+ const clusterUnavailable = relationUnavailable(workspace, 'clusters', ns, 'Clusters')
+ const cluster = clusterUnavailable ? null : targetCluster(resource, clustersIn(workspace))
+ const f = cnpgDatabaseRoleFacts(resource, cluster)
+ return (
+
+
+
+ Declared
+
+ {f.pgName ? {f.pgName} : }
+
+
+
+ {f.login ? 'Allowed' : 'Not allowed'}{f.superuser ? ' · superuser' : ''}
+
+ {f.passwordDisabled ? (
+ 'Disabled'
+ ) : f.passwordSecret ? (
+
+ From Secret {f.passwordSecret}
+
+ ) : (
+ No password Secret declared
+ )}
+ {f.passwordValidUntil && (
+
+ Valid until {f.passwordValidUntil} (PostgreSQL VALID UNTIL; the operator does not rotate it)
+
+ )}
+
+
+ {f.clientCertificate ? (
+
+ Operator-issued in Secret {f.clientCertificate.secret}
+
+ {' · '}
+ {f.clientCertificate.expiration ? `expires ${f.clientCertificate.expiration}` : 'expiry not reported yet'}
+
+ {f.clientCertificate.message && {f.clientCertificate.message}
}
+
+ ) : (
+ Not requested
+ )}
+
+
+
+
+
+
+
+
+ {f.overriddenByCluster === null ? (
+
+ ) : f.overriddenByCluster ? (
+
+ The Cluster declares “{f.pgName}” in spec.managed.roles, which takes precedence: this DatabaseRole is not reconciled while that entry exists
+
+ ) : (
+ No spec.managed.roles entry for this role
+ )}
+
+ }
+ />
+
+ )
+}
+
+export function CNPGSubscriptionSummary({
+ resource,
+ workspace,
+ onNavigate,
+ logicalPath,
+}: SummaryProps & {
+ /** The path and slot reading; derived from the workspace (slot not read) when omitted. */
+ logicalPath?: CNPGLogicalPathReading
+}) {
+ const reading: Partial = logicalPath ?? { path: workspacePaths(workspace, [resource])[0] }
const pub = resource?.spec?.publicationName
const ext = resource?.spec?.externalClusterName
return (
- Declared
+ Declared
{resource?.spec?.name ? {resource.spec.name} : }
@@ -254,11 +378,18 @@ export function CNPGSubscriptionSummary({ resource, workspace, onNavigate }: Sum
-
+
+
+ {reading.path && (
+ <>
+ Replication path
+
+ >
+ )}
)
}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGImageCatalogSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGImageCatalogSummary.tsx
index c7e05904ee..a17740c498 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGImageCatalogSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGImageCatalogSummary.tsx
@@ -1,8 +1,11 @@
import { CNPG_GROUP, getCNPGImageCatalogEntries } from '../resources/resource-utils-cnpg'
import type { CNPGWorkspaceResponse } from './workspace'
-import { FactGrid, FactRow, RefLink, SummaryHeading, toneTextClass, type CNPGNavigate } from './primitives'
import { NotReported, Note, ObjectProblems, SummaryShell } from './CNPGSharedSummary'
import { clustersIn, clustersUsingCatalog, refOf, relationUnavailable } from './relations'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow } from '../facts'
+import { SectionHeading } from '../ui/FoldSection'
export function CNPGImageCatalogSummary({
resource,
@@ -11,7 +14,7 @@ export function CNPGImageCatalogSummary({
}: {
resource: any
workspace: CNPGWorkspaceResponse | null
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
}) {
const clusterScoped = resource?.kind === 'ClusterImageCatalog'
const ns = resource?.metadata?.namespace ?? ''
@@ -27,7 +30,7 @@ export function CNPGImageCatalogSummary({
- Images
+ Images
{entries.length === 0 ? (
@@ -42,7 +45,7 @@ export function CNPGImageCatalogSummary({
)}
-
Used by
+
Used by
{unavailable ? (
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.test.tsx
new file mode 100644
index 0000000000..c5db75ecf9
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.test.tsx
@@ -0,0 +1,22 @@
+import { describe, expect, it } from 'vitest'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { CNPGLogicalPathView } from './CNPGLogicalPath'
+import type { CNPGLogicalPath } from './logicalReplication'
+
+const path: CNPGLogicalPath = {
+ subscription: { namespace: 'db', name: 'sub', cluster: 'pg', applied: true },
+ externalCluster: { name: 'upstream', declared: true },
+ publisher: { kind: 'external', reason: 'the external cluster names no host' },
+ publication: { name: 'pub' },
+ slot: { name: 'sub' },
+ failover: { text: 'unknown', tone: 'unknown' },
+}
+
+describe('CNPGLogicalPathView', () => {
+ it('says unknown parts of the path in words, never as "?"', () => {
+ const html = renderToStaticMarkup(
)
+ expect(html).toContain('external cluster upstream')
+ expect(html).toContain('database unknown')
+ expect(html).not.toMatch(/upstream<\/span>\/\?|\/\?/)
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.tsx b/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.tsx
new file mode 100644
index 0000000000..0364ad606a
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGLogicalPath.tsx
@@ -0,0 +1,150 @@
+import type { ReactNode } from 'react'
+import { ArrowRight } from 'lucide-react'
+import { CNPG_GROUP } from '../resources/resource-utils-cnpg'
+import { type Fact } from '../facts'
+import { cnpgLogicalLocation, type CNPGLogicalPath } from './logicalReplication'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactSource, FactValue } from '../facts'
+
+// A hop after the first carries its arrow, so a wrapped line never ends on an
+// arrow pointing at nothing.
+function Hop({ label, children, from }: { label: string; children: ReactNode; from?: boolean }) {
+ return (
+
+ )
+}
+
+/**
+ * One Subscription's path to its publisher: publication, the slot the
+ * publisher keeps for it, and whether that slot survives a publisher
+ * failover. `slot` is the host's reading of the publisher primary; absent
+ * means it was not read.
+ */
+export function CNPGLogicalPathView({
+ path,
+ slot,
+ onNavigate,
+ compact,
+ notice,
+}: {
+ path: CNPGLogicalPath
+ slot?: Fact
+ onNavigate?: NavigateToRef
+ compact?: boolean
+ /** The host's word on the slot reading, e.g. that its latest refresh failed. */
+ notice?: ReactNode
+}) {
+ const s = path.subscription
+ const pub = path.publication
+ const publisher = path.publisher
+ const slotFact: Fact = slot ?? { text: path.slot.name ? `Slot ${path.slot.name}: not read` : path.slot.reason ?? 'No slot', tone: 'unknown' }
+ const chain = (
+
+
+
+ {s.sqlName ?? s.name}
+
+ on {cnpgLogicalLocation(s.cluster ?? 'an unknown cluster', s.dbname)}
+
+
+ {pub.object ? (
+
+ {pub.name}
+
+ ) : pub.name ? (
+ {pub.name}
+ ) : (
+ name unknown
+ )}
+
+ {' '}
+ on{' '}
+ {publisher.kind === 'cluster' ? (
+
+ {publisher.namespace === s.namespace && publisher.name === s.cluster ? 'the same cluster' : `${publisher.namespace}/${publisher.name}`}
+
+ ) : (
+ {path.externalCluster.host ?? (path.externalCluster.name ? `external cluster ${path.externalCluster.name}` : 'an unnamed external cluster')}
+ )}
+ {pub.dbname ? `/${pub.dbname}` : ' · database unknown'}
+
+
+
+
+
+
+ )
+ if (compact) {
+ return (
+
+ {notice}
+ {chain}
+
Failover: {path.failover.text}
+
+ )
+ }
+ return (
+
+ {notice}
+ {chain}
+
+
+ {publisher.kind === 'cluster' ? (
+
+ {publisher.namespace}/{publisher.name} via {publisher.via}
+
+ ) : (
+ Outside Radar's view: {publisher.reason}
+ )}
+
+ {path.externalCluster.name
+ ? `From the subscriber's spec.externalClusters[${path.externalCluster.name}].connectionParameters.host`
+ : "The subscriber's spec.externalClusterName is not set, so no external cluster names the host"}
+
+
+
+ {pub.object ? (
+
+ {pub.object.applied === true ? 'Applied' : pub.object.applied === false ? 'Not applied' : 'Pending'}
+
+ ) : (
+
+ {publisher.kind !== 'cluster'
+ ? 'Unknown'
+ : pub.unavailable
+ ? `Unknown: ${pub.unavailable} in ${publisher.namespace}`
+ : 'No Publication object declares it: it may exist in SQL only'}
+
+ )}
+
+
+
+
+
+
+
+
+
+
+ {s.applied === false ? (
+ Not applied{s.message ? `: ${s.message}` : ''}
+ ) : s.applied === true ? (
+ 'Applied'
+ ) : (
+ Pending
+ )}
+
+ Subscription status. Apply errors and lag are not reported: CloudNativePG's default exporter has no pg_stat_subscription query.
+
+
+
+
+ )
+}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGObjectStoreSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGObjectStoreSummary.tsx
index 39e1556ea9..801ea1936d 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGObjectStoreSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGObjectStoreSummary.tsx
@@ -8,9 +8,12 @@ import {
getCNPGObjectStoreRetention,
} from '../resources/resource-utils-cnpg'
import type { CNPGWorkspaceResponse } from './workspace'
-import { FactGrid, FactRow, FactValue, RefLink, SummaryHeading, toneTextClass, type CNPGNavigate } from './primitives'
import { NotReported, Note, ObjectProblems, SummaryShell, TimeAgo } from './CNPGSharedSummary'
import { clustersIn, inferredObjectStoreHealth, refOf, relationUnavailable, usersOfObjectStore } from './relations'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactValue } from '../facts'
+import { SectionHeading } from '../ui/FoldSection'
function utc(at: string | undefined): string {
if (!at || !Number.isFinite(Date.parse(at))) return 'unknown'
@@ -24,7 +27,7 @@ export function CNPGObjectStoreSummary({
}: {
resource: any
workspace: CNPGWorkspaceResponse | null
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
}) {
const ns = resource?.metadata?.namespace ?? ''
const clustersUnavailable = relationUnavailable(workspace, 'clusters', ns, 'Clusters')
@@ -46,7 +49,7 @@ export function CNPGObjectStoreSummary({
onNavigate={onNavigate}
/>
-
Upload health
+
Upload health
{clustersUnavailable ? (
@@ -89,7 +92,7 @@ export function CNPGObjectStoreSummary({
>
)}
-
Recovery window
+
Recovery window
{windows.length === 0 ? (
@@ -136,7 +139,7 @@ export function CNPGObjectStoreSummary({
)}
-
Destination
+
Destination
{destination !== '-' ? {destination} : }
{provider ?? }
@@ -152,7 +155,7 @@ export function CNPGObjectStoreSummary({
{retention ?? }
-
Used by
+
Used by
{clustersUnavailable ? (
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGObjectSummary.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGObjectSummary.test.tsx
index b715f2d584..ede23f0268 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGObjectSummary.test.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGObjectSummary.test.tsx
@@ -2,9 +2,10 @@ import { describe, expect, it } from 'vitest'
import { renderToString } from 'react-dom/server'
import { CNPGBackupSummary, CNPGScheduledBackupSummary } from './CNPGBackupSummary'
import { CNPGObjectStoreSummary } from './CNPGObjectStoreSummary'
-import { CNPGDatabaseSummary } from './CNPGDeclarativeSummary'
+import { CNPGDatabaseSummary, CNPGPublicationSummary, CNPGSubscriptionSummary } from './CNPGDeclarativeSummary'
import { CNPGPoolerSummary } from './CNPGPoolerSummary'
import { CNPGImageCatalogSummary } from './CNPGImageCatalogSummary'
+import { ClusterLink } from './CNPGSharedSummary'
import { CNPG_WORKSPACE_KEYS, type CNPGWorkspaceKey, type CNPGWorkspaceResponse } from './workspace'
const PG = 'postgresql.cnpg.io/v1'
@@ -77,6 +78,22 @@ describe('CNPGBackupSummary', () => {
expect(t).toContain('an earlier schedule of that name')
})
+ it('distinguishes a replaced Cluster from a missing or unread one without linking to its replacement', () => {
+ const previous = { ...b, status: { ...b.status, pluginMetadata: { clusterUID: 'old' } } }
+ const replacement = { ...mainCluster, metadata: { ...mainCluster.metadata, uid: 'new' } }
+ const html = renderToString(
)
+ expect(text(html)).toContain('main · an earlier Cluster of that name; the current one is a different object')
+ expect(html).not.toContain('
))
+ expect(missing).toContain('not found in this namespace')
+ const unread = text(renderToString(
))
+ expect(unread).not.toContain('not found')
+ expect(unread).not.toContain('earlier Cluster')
+ const current = renderToString(
)
+ expect(current).toContain('
{
const issues = [
{ id: 'i1', severity: 'critical' as const, kind: 'Backup', group: 'postgresql.cnpg.io', namespace: 'pg', name: 'main-20260901', reason: 'CNPGBackupFailed', message: 'Backup failed' },
@@ -101,6 +118,34 @@ describe('CNPGScheduledBackupSummary', () => {
expect(t).toContain('Backups older than 7 days are not listed')
})
+ it("shows the server's reading and next runs only for the schedule it read", () => {
+ const sched = { apiVersion: PG, kind: 'ScheduledBackup', metadata: { name: 'nightly', namespace: 'pg' }, spec: { cluster: { name: 'main' }, schedule: '0 30 2 * * *' } }
+ const preview = {
+ schedule: '0 30 2 * * *',
+ valid: true,
+ description: 'every day at 02:30:00 UTC',
+ nextRuns: ['2026-10-01T02:30:00Z', '2026-10-02T02:30:00Z', '2026-10-03T02:30:00Z'],
+ basis: 'lastCheckTime' as const,
+ clock: { zone: 'UTC', declared: true, source: 'Operator Deployment declares TZ=UTC' },
+ }
+ const t = text(renderToString( ))
+ expect(t).toContain('every day at 02:30:00 UTC')
+ expect(t).toContain('2026-10-01 02:30:00 UTC')
+ expect(t).toContain('Calculated upcoming times')
+ expect(t).toContain('Next run reported by the operatorNot reported')
+ expect(t).toContain("operator's last check")
+ const estimate = text(renderToString( ))
+ expect(estimate).toContain('Estimated upcoming times')
+ expect(estimate).toContain('UTC estimate')
+ expect(estimate).not.toContain('Calculated upcoming times')
+ const stale = text(renderToString( ))
+ expect(stale).not.toContain('every day at')
+ const due = text(renderToString( ))
+ expect(due).toContain('due (2026-10-01 02:30:00 UTC)')
+ expect(due).toContain('A run is due on the declared clock')
+ expect(due).toContain('at most one catch-up backup')
+ })
+
it('says when Backups are not readable instead of listing none', () => {
const sched = { apiVersion: PG, kind: 'ScheduledBackup', metadata: { name: 'nightly', namespace: 'pg' }, spec: { cluster: { name: 'main' } } }
const t = text(renderToString( ))
@@ -196,10 +241,45 @@ describe('CNPGPoolerSummary', () => {
it('reports unknown scheduled count and unmeasured pressure', () => {
const pooler = { apiVersion: PG, kind: 'Pooler', metadata: { name: 'main-rw', namespace: 'pg' }, spec: { cluster: { name: 'main' }, type: 'rw', instances: 2 } }
const t = text(renderToString( ))
- expect(t).toContain('Scheduled count not reported')
+ expect(t).toContain('Pooler instance count not reported')
expect(t).toContain('Not measured')
expect(t).toContain('main-rw')
})
+
+ it('shows Deployment readiness, limits with PgBouncer defaults, observed pause and the Service path when live data is provided', () => {
+ const pooler = {
+ apiVersion: PG,
+ kind: 'Pooler',
+ metadata: { name: 'main-rw', namespace: 'pg' },
+ spec: { cluster: { name: 'main' }, type: 'rw', instances: 2, pgbouncer: { paused: true, parameters: { max_client_conn: '200' } } },
+ status: { instances: 2 },
+ }
+ const t = text(
+ renderToString(
+ ,
+ ),
+ )
+ expect(t).toContain('from Deployment main-rw')
+ expect(t).toContain('Pause requested')
+ expect(t).toContain('1/2 ready')
+ expect(t).toContain('Pause state')
+ expect(t).toContain('Observed: Paused on 1 of 2 PgBouncers')
+ expect(t).toContain('PgBouncer uses 20')
+ expect(t).toContain('200')
+ expect(t).toContain('main-rw')
+ expect(t).toContain('app/app')
+ expect(t).not.toContain('Not measured')
+ })
})
describe('CNPGImageCatalogSummary', () => {
@@ -234,3 +314,114 @@ describe('CNPGImageCatalogSummary', () => {
expect(t).toContain('No access to Clusters')
})
})
+
+describe('logical replication summaries', () => {
+ const src = { apiVersion: PG, kind: 'Cluster', metadata: { name: 'src', namespace: 'pg' }, spec: { instances: 2 } }
+ const dst = { apiVersion: PG, kind: 'Cluster', metadata: { name: 'dst', namespace: 'pg' }, spec: { instances: 1, externalClusters: [{ name: 'src', connectionParameters: { host: 'src-rw', dbname: 'app' } }] } }
+ const pub = { apiVersion: PG, kind: 'Publication', metadata: { name: 'orders-pub', namespace: 'pg' }, spec: { cluster: { name: 'src' }, name: 'orders_pub', dbname: 'app', target: { allTables: true } }, status: { applied: true } }
+ const sub = { apiVersion: PG, kind: 'Subscription', metadata: { name: 'orders-sub', namespace: 'pg' }, spec: { cluster: { name: 'dst' }, name: 'orders_sub', dbname: 'app', publicationName: 'orders_pub', externalClusterName: 'src' }, status: { applied: true } }
+ const w = ws({ clusters: [src, dst], publications: [pub], subscriptions: [sub] })
+
+ it('shows the subscription path with the slot unread and the failover verdict', () => {
+ const t = text(renderToString( ))
+ expect(t).toContain('Replication path')
+ expect(t).toContain('orders_pub')
+ expect(t).toContain('Slot orders_sub: not read')
+ expect(t).toContain('Lost on failover')
+ expect(t).toContain('no pg_stat_subscription query')
+ })
+
+ it('never says no Publication declares it when Publications are unreadable', () => {
+ const denied = ws({ clusters: [src, dst], subscriptions: [sub] }, { coverage: { publications: { state: 'denied' } } })
+ const t = text(renderToString( ))
+ expect(t).toContain('Unknown: No access to Publications in pg')
+ expect(t).not.toContain('No Publication object declares it')
+ })
+
+ it('lists the subscribers of a publication', () => {
+ const t = text(renderToString( ))
+ expect(t).toContain('Subscribers')
+ expect(t).toContain('orders_sub')
+ })
+})
+
+it('shows the stale declaration as pending in its drawer', () => {
+ const resource = { apiVersion: PG, kind: 'Database', metadata: { name: 'app', namespace: 'pg', generation: 3 }, spec: { name: 'app' }, status: { applied: true, observedGeneration: 2 } }
+ const t = text(renderToString( ))
+ expect(t).toContain('Pending · awaiting the operator for the current spec')
+ expect(t).toContain('AppliedPending')
+})
+
+it('keeps Deployment readiness beside the pause request even when no Pods are ready', () => {
+ const resource = { metadata: { name: 'p' }, spec: { pgbouncer: { paused: true } } }
+ const t = text(renderToString( ))
+ expect(t).toContain('0/2 ready')
+ expect(t).toContain('Pause requested')
+ expect(t).not.toContain('Observed: Paused')
+})
+
+it('puts the pending Pooler cause under readiness and names the Pod whose metric read failed', () => {
+ const html = renderToString( )
+ const t = text(html)
+ expect(t).toContain('0/1 ready')
+ expect(t).toContain('Cannot be scheduled: insufficient cpu.')
+ expect(t).toContain('Not measured: PgBouncer did not answer')
+ expect(t).toContain('pooler-pod')
+ expect(t).toContain('Measurement details')
+ expect(t).toContain('address not allowed')
+ expect(t.indexOf('Cannot be scheduled')).toBeGreaterThan(-1)
+ expect(t.indexOf('Cannot be scheduled')).toBeLessThan(t.indexOf('pooler-pod'))
+ expect(t.indexOf('Cannot be scheduled')).toBeLessThan(t.indexOf('Connections'))
+ expect(t.indexOf('address not allowed')).toBeGreaterThan(t.indexOf('Connections'))
+ expect(html).toContain('button')
+})
+
+it('never calls an incomplete empty Pooler read idle and names the unread Pod', () => {
+ const t = text(renderToString( ))
+ expect(t).toContain('No pools seen in what was read')
+ expect(t).toContain('b: not read (timeout)')
+ expect(t).not.toContain('Idle:')
+})
+
+it('leads Pooler observations with the scheduling cause and labels the operator count', () => {
+ const pod = { pod: 'orders-pooler-pod', state: 'unreachable', reason: 'PgBouncer has not started (Pod cannot be scheduled)', schedulingReason: 'Unschedulable: 0/2 nodes are available: 2 Too many pods. preemption: no victims.' }
+ const html = renderToString( )
+ const t = text(html)
+ expect(t).toContain('Pooler status reports 1 instance · 1 requested')
+ expect(t).toContain('Cannot be scheduled: both nodes have reached their Pod limit')
+ expect(t).toContain('Not measured: PgBouncer has not started (Pod cannot be scheduled)')
+ expect(t).toContain('1 not read: orders-pooler-pod (PgBouncer has not started (Pod cannot be scheduled))')
+ expect(html).toContain('aria-expanded="false"')
+ expect(t).toContain('preemption: no victims.')
+})
+
+it('shows the schedule destination blocker only when the target Cluster is readable', () => {
+ const schedule = { metadata: { name: 'payments-nightly', namespace: 'pg' }, spec: { cluster: { name: 'payments' } } }
+ const cluster = { apiVersion: PG, kind: 'Cluster', metadata: { name: 'payments', namespace: 'pg' }, spec: {} }
+ const render = (workspace: CNPGWorkspaceResponse) => text(renderToString( ))
+ const t = render(ws({ clusters: [cluster] }))
+ expect(t).toContain('Enabled · not run yetNo backup destination')
+ expect(t).toContain('The resource list status comes from the ScheduledBackup alone.')
+ expect(render(ws({}))).not.toContain('No backup destination')
+ expect(render(ws({}))).not.toContain('The resource list status')
+ expect(render(ws({ clusters: [{ ...cluster, spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } } }] }))).not.toContain('The resource list status')
+ expect(text(renderToString( ))).not.toContain('The resource list status')
+ expect(text(renderToString( ))).not.toContain('No backup destination')
+})
+
+it('names visible database declarations without claiming a SQL inventory', () => {
+ const resource = { metadata: { name: 'orders-appdb', namespace: 'pg' }, spec: { name: 'appdb', cluster: { name: 'orders' } } }
+ const t = text(renderToString( ))
+ expect(t).toContain('No visible Publication declarations')
+ expect(t).toContain('No visible Subscription declarations')
+ expect(t.match(/Objects created in SQL are not shown/g)).toHaveLength(1)
+ const denied = text(renderToString( ))
+ expect(denied).toContain('No access to Publications')
+ expect(denied).not.toContain('No visible Publication declarations')
+})
+
+it.each([['session', 'A client keeps one server connection for its whole session.'], ['transaction', 'A client uses a server connection only for each transaction.']])('explains %s pool mode before its raw value', (poolMode, explanation) => {
+ const t = text(renderToString( ))
+ expect(t).toContain(explanation)
+ expect(t.indexOf(explanation)).toBeLessThan(t.indexOf(poolMode, t.indexOf(explanation) + explanation.length))
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGPoolerScheduling.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGPoolerScheduling.test.tsx
new file mode 100644
index 0000000000..f20e3955e2
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGPoolerScheduling.test.tsx
@@ -0,0 +1,36 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot } from 'react-dom/client'
+import { expect, it, vi } from 'vitest'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { CNPGPoolerScheduling } from './CNPGPoolerSummary'
+Object.assign(globalThis, { IS_REACT_ACT_ENVIRONMENT: true })
+it('leads with the cause, truncates a secondary Pod link, exposes its full name on hover and navigates once', () => {
+ vi.useFakeTimers()
+ const pod = 'orders-pooler-rw-6f9e349dec-abcdefghij'
+ const navigate = vi.fn()
+ const host = document.createElement('div'); document.body.append(host); const root = createRoot(host)
+ act(() => root.render( ))
+ expect(host.firstElementChild!.firstElementChild!.textContent).toBe('Cannot be scheduled: both nodes have reached their Pod limit.')
+ const link = [...host.querySelectorAll('button')].find((b) => b.textContent === pod)!
+ expect(link.parentElement!.className).toContain('[&_button]:truncate')
+ expect(link.parentElement!.parentElement!.className).toContain('text-theme-text-secondary')
+ act(() => { link.dispatchEvent(new MouseEvent('mouseover', { bubbles: true })); vi.advanceTimersByTime(350) })
+ expect(document.querySelector('[role="tooltip"]')!.textContent).toBe(pod)
+ act(() => link.click())
+ expect(navigate).toHaveBeenCalledExactlyOnceWith({ kind: 'Pod', group: '', namespace: 'db', name: pod })
+ act(() => root.unmount()); host.remove(); vi.useRealTimers()
+})
+
+it('groups a shared scheduling cause while retaining every Pod and its full scheduler message', () => {
+ const first = '0/2 nodes are available: 2 Too many pods. preemption: no victims.'
+ const second = '0/2 nodes are available: 2 Too many pods. preemption: no candidates.'
+ const html = renderToStaticMarkup( {}} />)
+ expect(html.match(/Cannot be scheduled: both nodes have reached their Pod limit/g)).toHaveLength(1)
+ expect(html).toContain('Cannot be scheduled: insufficient cpu')
+ const host = document.createElement('div'); host.innerHTML = html
+ expect([...host.querySelectorAll('button')].filter((b) => ['a', 'b', 'c'].includes(b.textContent!)).map((b) => b.textContent)).toEqual(['a', 'b', 'c'])
+ expect(html).toContain(first)
+ expect(html).toContain(second)
+ expect(html.match(/Scheduler message/g)).toHaveLength(3)
+})
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGPoolerSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGPoolerSummary.tsx
index 9418de0626..cebc0e2f52 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGPoolerSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGPoolerSummary.tsx
@@ -1,8 +1,27 @@
-import { getCNPGPoolerDeploymentName, getCNPGPoolerMode, getCNPGPoolerStatus, isCNPGPoolerPaused } from '../resources/resource-utils-cnpg'
+import { Tooltip } from '../ui/Tooltip'
+import type { ReactNode } from 'react'
+import { getCNPGPoolerDeploymentName, getCNPGPoolerMode, isCNPGPoolerPaused } from '../resources/resource-utils-cnpg'
import type { CNPGWorkspaceResponse } from './workspace'
-import { FactGrid, FactRow, FactValue, RefLink, SummaryHeading, type CNPGNavigate } from './primitives'
import { ClusterLink, NotReported, Note, ObjectProblems, PhaseBadge, SummaryShell } from './CNPGSharedSummary'
import { refOf } from './relations'
+import {
+ POOLER_LIMIT_PARAMETERS,
+ aggregatePoolerPools,
+ poolerPodPressure,
+ poolerPressureCoverage,
+ poolerPressureFact,
+ observedPause,
+ poolerBackendService,
+ poolerReadiness,
+ type CNPGPoolerLive,
+ type CNPGPoolerPoolRow,
+} from './pooler'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { toneTextClass } from '../ui/status-tone'
+import { FactGrid, FactRow, FactValue } from '../facts'
+import { Badge } from '../ui/Badge'
+import { summarizeSchedulerMessage } from '../resources/resource-utils'
+import { SectionHeading, FoldSection } from '../ui/FoldSection'
const TYPE_LABEL: Record = {
rw: 'rw · routes to the primary',
@@ -14,36 +33,53 @@ export function CNPGPoolerSummary({
resource,
workspace,
onNavigate,
+ live,
+ actions,
+ lead,
}: {
resource: any
workspace: CNPGWorkspaceResponse | null
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
+ /** Live reads a host adds (Deployment readiness, PgBouncer metrics and state). */
+ live?: CNPGPoolerLive
+ /** Operations rendered beside the paused state (pause / resume). */
+ actions?: ReactNode
+ /** Rendered first, e.g. a host's notice that a live read is stale. */
+ lead?: ReactNode
}) {
const ns = resource?.metadata?.namespace ?? ''
const type = resource?.spec?.type
const desired = resource?.spec?.instances
- const scheduled = resource?.status?.instances
+ const reported = resource?.status?.instances
const deployment = getCNPGPoolerDeploymentName(resource)
+ const paused = isCNPGPoolerPaused(resource)
+ const readiness = poolerReadiness(live?.deployment)
+ const observed = observedPause(live?.observed)
return (
+ {lead}
- State
+ State
-
- {isCNPGPoolerPaused(resource) && PgBouncer is paused: it holds client connections instead of serving them. }
+
+
+ {paused &&
Pause requested }
+
+ {readiness.detail && {readiness.detail} }
+
- {typeof scheduled === 'number' ? `${scheduled} scheduled` : }
+ {typeof reported === 'number' ? `Pooler status reports ${reported} ${reported === 1 ? 'instance' : 'instances'}` : }
{' · '}
- {typeof desired === 'number' ? `${desired} desired` : 'desired not set'}
+ {typeof desired === 'number' ? `${desired} requested` : 'requested count not set'}
- The Pooler counts scheduled pods, not ready ones; readiness is on its Deployment.
+ status.instances is the operator’s count; observed readiness is on its Deployment and Pods.
{deployment ? (
@@ -52,19 +88,210 @@ export function CNPGPoolerSummary({
)}
-
-
+
+
+ {paused ? 'Pause requested' : 'Serving requested (not paused)'}
+ {actions}
+
+ spec.pgbouncer.paused is what was asked for; each PgBouncer applies it with PAUSE / RESUME.
+ {paused && When PgBouncer applies the pause, it holds client connections instead of serving them. }
+ {observed && (
+
+
+
+ )}
+
+
+
+ Connections
+ {live?.pressure ? (
+
+ ) : (
+
+
+
+
+
+ )}
+
+ Limits
+
+
+ {getCNPGPoolerMode(resource) === 'transaction' ? 'A client uses a server connection only for each transaction.' : getCNPGPoolerMode(resource) === 'session' ? 'A client keeps one server connection for its whole session.' : 'A client uses a server connection for each statement.'}
+ {getCNPGPoolerMode(resource)}{!resource?.spec?.pgbouncer?.poolMode ? ' (default)' : ''}
+ {POOLER_LIMIT_PARAMETERS.map((p) => {
+ const v = resource?.spec?.pgbouncer?.parameters?.[p.key]
+ return (
+
+ {v !== undefined ? (
+ {String(v)}
+ ) : (
+
+ default{p.pgbouncerDefault ? · PgBouncer uses {p.pgbouncerDefault} : null}
+
+ )}
+
+ {p.key}
+
+
+ )
+ })}
- Routing
+ Routing
{type ? TYPE_LABEL[type] ?? type : }
- {getCNPGPoolerMode(resource)}
+ {live?.service && (
+
+
+
+ )}
)
}
+
+function PoolerPath({ resource, live, onNavigate }: { resource: any; live: CNPGPoolerLive; onNavigate?: NavigateToRef }) {
+ const ns = resource?.metadata?.namespace ?? ''
+ const svc = live.service!
+ const backend = poolerBackendService(resource?.spec?.cluster?.name, resource?.spec?.type)
+ const svcState: Record = { missing: 'does not exist', unreadable: 'no access', foreign: 'not controlled by this Pooler' }
+ return (
+
+
+ Service
+ {svc.state === 'ok' ? (
+ {svc.port ? ` :${svc.port}` : ''}{svc.type ? ` · ${svc.type}` : ''}
+ ) : (
+ · {svcState[svc.state]}
+ )}
+
+
→ PgBouncer ({resource?.metadata?.name})
+
+ → {backend ? : 'the cluster'}
+ {' '}of Cluster {resource?.spec?.cluster?.name ?? '—'}
+
+
+ )
+}
+
+// Pods take their client connections independently, so one can queue while
+// the sum still looks calm.
+function PoolerPodPressure({ pods }: { pods: NonNullable['pods'] }) {
+ return (
+
+
+
+ PgBouncer Pod
+ Clients active
+ Waiting
+ Servers active
+ Max wait
+
+
+
+ {poolerPodPressure(pods).map((r) =>
+ r.state === 'ok' || r.state === 'partial' ? (
+
+
+ {r.pod}
+ {r.state === 'partial' && partial }
+
+ p.pod === r.pod), 'clActive')} />
+ p.pod === r.pod), 'clWaiting')} />
+ p.pod === r.pod), 'svActive')} />
+ p.pod === r.pod), 'maxwaitSeconds')} />
+
+ ) : (
+
+ {r.pod}
+
+ Not read: {r.error ?? r.state}
+
+
+ ),
+ )}
+
+
+ )
+}
+
+function PoolerPressure({ pressure }: { pressure: NonNullable }) {
+ if (pressure.state === 'loading') return Reading PgBouncer metrics…
+ if (pressure.state === 'denied' || pressure.state === 'error') {
+ return
+ }
+ const { reporting, limitation, empty } = poolerPressureCoverage(pressure.pods)
+ const rows = aggregatePoolerPools(reporting)
+ if (reporting.length === 0) {
+ return
+ }
+ return (
+
+ {rows.length === 0 ? (
+
{empty}
+ ) : (
+
+
+
+ Pool
+ Mode
+ Clients active
+ Waiting
+ Servers active / idle / used
+ Max wait
+
+
+
+ {rows.map((r: CNPGPoolerPoolRow) => (
+
+ {r.database}/{r.user}
+ {r.poolModes.join(', ') || '—'}
+
+
+ / /
+
+
+ ))}
+
+
+ )}
+
+ Summed over {reporting.length} of {pressure.pods.length} PgBouncer Pods{limitation ? ` · ${limitation}` : ''}. PgBouncer’s admin and
+ authentication pools are excluded.
+
+ {pressure.pods.length > 1 &&
}
+
+ )
+}
+
+export function CNPGPoolerScheduling({ namespace, pods, onNavigate }: { namespace: string; pods: NonNullable['pods']; onNavigate?: NavigateToRef }) {
+ const byCause = new Map()
+ for (const pod of pods) {
+ if (!pod.schedulingReason) continue
+ const cause = summarizeSchedulerMessage(pod.schedulingReason, { plain: true })
+ const group = byCause.get(cause) ?? []
+ group.push(pod)
+ byCause.set(cause, group)
+ }
+ return <>{[...byCause].map(([cause, blocked]) =>
+
Cannot be scheduled: {cause}.
+ {blocked.map((p) =>
+
+
{p.schedulingReason}
+
)}
+
)}>
+}
+
+export function CNPGPoolerUnmeasured({ pods }: { pods: NonNullable['pods'] }) {
+ if (pods.length === 0) return Not measured: no PgBouncer answered
+ return <>{pods.map((p) =>
+ Not measured: {p.state === 'denied' ? 'needs get pods/proxy' : p.reason ?? 'PgBouncer did not answer'}
+
{p.pod}
+ {p.error &&
{p.error}
}
+
)}>
+}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGSharedSummary.tsx b/packages/k8s-ui/src/components/cnpg/CNPGSharedSummary.tsx
index 301ee5ad23..183405aab8 100644
--- a/packages/k8s-ui/src/components/cnpg/CNPGSharedSummary.tsx
+++ b/packages/k8s-ui/src/components/cnpg/CNPGSharedSummary.tsx
@@ -3,8 +3,11 @@ import { Badge } from '../ui/Badge'
import type { StatusBadge as StatusBadgeValue } from '../resources/resource-utils'
import { CNPG_GROUP } from '../resources/resource-utils-cnpg'
import type { CNPGWorkspaceIssue, CNPGWorkspaceResponse } from './workspace'
-import { FactValue, ProblemCallout, RefLink, type CNPGNavigate } from './primitives'
-import { clustersIn, healthSeverity, problemsForObject, relationUnavailable, targetCluster, type CNPGObjectRef } from './relations'
+import { healthToSeverity } from '../../utils/badge-colors'
+import { clustersIn, problemsForObject, relationUnavailable, targetCluster, type CNPGObjectRef } from './relations'
+import { type NavigateToRef, RefLink } from '../ui/RefLink'
+import { FactValue } from '../facts'
+import { ProblemCallout } from '../problems'
const MAX_PROBLEMS = 3
@@ -20,7 +23,7 @@ export function ObjectProblems({
}: {
issues: CNPGWorkspaceIssue[] | undefined
subject: CNPGObjectRef
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
}) {
const problems = problemsForObject(issues, subject)
if (problems.length === 0) return null
@@ -30,6 +33,7 @@ export function ObjectProblems({
{shown.map((p, i) => (
+
{status.text}
)
@@ -71,16 +75,20 @@ export function ClusterLink({
}: {
resource: any
workspace: CNPGWorkspaceResponse | null
- onNavigate?: CNPGNavigate
+ onNavigate?: NavigateToRef
}) {
const name = resource?.spec?.cluster?.name
if (!name) return
const ns = resource?.metadata?.namespace ?? ''
- const visible = !!targetCluster(resource, clustersIn(workspace))
+ const clusters = clustersIn(workspace)
+ const visible = !!targetCluster(resource, clusters)
+ const replaced = !visible && clusters.some((c) => c.metadata?.namespace === ns && c.metadata?.name === name)
return (
-
- {workspace && !visible && !relationUnavailable(workspace, 'clusters', ns, 'Clusters') && (
+
+ {replaced ? (
+ · an earlier Cluster of that name; the current one is a different object
+ ) : workspace && !visible && !relationUnavailable(workspace, 'clusters', ns, 'Clusters') && (
· not found in this namespace
)}
diff --git a/packages/k8s-ui/src/components/cnpg/CNPGWALArchivingFact.test.tsx b/packages/k8s-ui/src/components/cnpg/CNPGWALArchivingFact.test.tsx
new file mode 100644
index 0000000000..aaf3bb65a2
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/CNPGWALArchivingFact.test.tsx
@@ -0,0 +1,39 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot } from 'react-dom/client'
+import { expect, it, vi } from 'vitest'
+import { CNPGWALArchivingFact } from './CNPGClusterSummary'
+Object.assign(globalThis, { IS_REACT_ACT_ENVIRONMENT: true })
+it('folds the raw condition and exposes relative age with the exact timestamp on hover', () => {
+ vi.useFakeTimers(); vi.setSystemTime(new Date('2026-10-01T13:00:00Z'))
+ const stamp = '2026-10-01T12:00:00Z'
+ const host = document.createElement('div'); document.body.append(host); const root = createRoot(host)
+ act(() => root.render( ))
+ expect(host.textContent).toContain('WAL is not archived to recovery storage')
+ const button = host.querySelector('button')!
+ expect(button.getAttribute('aria-expanded')).toBe('false')
+ const panel = document.getElementById(button.getAttribute('aria-controls')!)!
+ expect(panel.firstElementChild!.hasAttribute('inert')).toBe(true)
+ act(() => button.click())
+ expect(button.getAttribute('aria-expanded')).toBe('true')
+ expect(host.textContent).toContain('ContinuousArchiving: True')
+ expect(host.textContent).toContain('Continuous archiving is working')
+ expect(host.textContent).toContain('1h ago')
+ expect(host.querySelector('time')?.dateTime).toBe(stamp)
+ act(() => { host.querySelector('time')!.dispatchEvent(new MouseEvent('mouseover', { bubbles: true })); vi.advanceTimersByTime(350) })
+ expect(document.body.textContent).toContain(stamp)
+ act(() => root.unmount()); host.remove(); vi.useRealTimers()
+})
+
+it('keeps the compact verdict, consequence and per-row operator disclosure', () => {
+ const host = document.createElement('div'); const root = createRoot(host)
+ act(() => root.render( ))
+ expect(host.textContent).toContain('Archive destination absent')
+ expect(host.textContent).toContain('No point-in-time recovery')
+ expect(host.textContent).not.toContain('because')
+ const fold = host.querySelector('button')!
+ expect(fold.textContent).toContain('Operator report')
+ act(() => fold.click())
+ expect(host.textContent).toContain('ContinuousArchiving: True')
+ act(() => root.unmount())
+})
diff --git a/packages/k8s-ui/src/components/cnpg/backupRuns.test.ts b/packages/k8s-ui/src/components/cnpg/backupRuns.test.ts
new file mode 100644
index 0000000000..75302d6453
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/backupRuns.test.ts
@@ -0,0 +1,17 @@
+import { describe, expect, it } from 'vitest'
+import { cnpgBackupRunInFlight, cnpgBackupRunsInWindow } from './backupRuns'
+
+const now = Date.parse('2026-10-05T12:00:00Z')
+const backup = (phase: string, ageDays: number) => ({ apiVersion: 'postgresql.cnpg.io/v1', metadata: { name: phase }, status: { phase, startedAt: new Date(now - ageDays * 86400000).toISOString() } })
+
+describe('backup run window', () => {
+ it('retains every in-flight phase regardless of age, including an unknown phase', () => {
+ const runs = ['pending', 'started', 'running', 'finalizing', 'walArchivingFailing', 'new-phase'].map((phase) => backup(phase, 10))
+ expect(cnpgBackupRunsInWindow(runs, now)).toHaveLength(runs.length)
+ expect(runs.every(cnpgBackupRunInFlight)).toBe(true)
+ })
+ it('applies the cutoff only to settled CNPG runs', () => {
+ const runs = [backup('completed', 10), backup('failed', 10), backup('completed', 2), backup('failed', 1), { ...backup('running', 1), apiVersion: 'velero.io/v1' }]
+ expect(cnpgBackupRunsInWindow(runs, now).map((r) => r.status.phase)).toEqual(['failed', 'completed'])
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/backupRuns.ts b/packages/k8s-ui/src/components/cnpg/backupRuns.ts
new file mode 100644
index 0000000000..0f3a62be47
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/backupRuns.ts
@@ -0,0 +1,19 @@
+import { isApiGroup } from '../resources/resource-utils-cnpg'
+
+export const CNPG_BACKUP_RUN_WINDOW_MS = 7 * 24 * 60 * 60 * 1000
+
+export function cnpgBackupRunInFlight(backup: any): boolean {
+ return backup?.status?.phase !== 'completed' && backup?.status?.phase !== 'failed'
+}
+
+export function cnpgBackupRunTime(backup: any): string | undefined {
+ return backup?.status?.stoppedAt || backup?.status?.startedAt || backup?.metadata?.creationTimestamp
+}
+
+export function cnpgBackupRunsInWindow(backups: any[], now = Date.now()): any[] {
+ return backups.filter((b) => {
+ if (!isApiGroup(b.apiVersion, 'postgresql.cnpg.io')) return false
+ const time = Date.parse(cnpgBackupRunTime(b) ?? '')
+ return cnpgBackupRunInFlight(b) || (Number.isFinite(time) && now - time <= CNPG_BACKUP_RUN_WINDOW_MS)
+ }).sort((a, b) => Date.parse(cnpgBackupRunTime(b) ?? '') - Date.parse(cnpgBackupRunTime(a) ?? ''))
+}
diff --git a/packages/k8s-ui/src/components/cnpg/connect.test.ts b/packages/k8s-ui/src/components/cnpg/connect.test.ts
new file mode 100644
index 0000000000..d3d7c0feda
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/connect.test.ts
@@ -0,0 +1,124 @@
+import { cnpgEndpointAvailability } from './connect'
+import type { CNPGClusterHA } from './ha'
+import { describe, expect, it } from 'vitest'
+import { cnpgConnectInfo, cnpgConnectionURI, cnpgPortForwardCommand, cnpgPsqlCommand } from './connect'
+
+const cluster = (spec: any = {}) => ({ apiVersion: 'postgresql.cnpg.io/v1', kind: 'Cluster', metadata: { name: 'pg', namespace: 'db' }, spec: { instances: 3, ...spec } })
+
+describe('cnpgConnectInfo', () => {
+ it('forwards the actual Service port to the local template port', () => {
+ const endpoint = { ...cnpgConnectInfo(cluster()).endpoints[0], port: 6543 }
+ expect(cnpgPortForwardCommand(endpoint, 'db')).toBe('kubectl -n db port-forward service/pg-rw 5432:6543')
+ expect(cnpgPortForwardCommand(endpoint, 'db', 15432)).toBe('kubectl -n db port-forward service/pg-rw 15432:6543')
+ })
+ it('lists the default Services and names the Secret by convention when the spec does not', () => {
+ const info = cnpgConnectInfo(cluster())
+ expect(info.endpoints.map((e) => `${e.role} ${e.host}:${e.port}`)).toEqual(['rw pg-rw.db.svc:5432', 'ro pg-ro.db.svc:5432', 'r pg-r.db.svc:5432'])
+ expect(info.database).toEqual({ value: 'app', source: "CloudNativePG's default" })
+ expect(info.owner.value).toBe('app')
+ expect(info.secret).toEqual({ name: 'pg-app', source: 'name by convention (-app)', byConvention: true })
+ expect(info.endpoints.find((e) => e.role === 'r')?.selects).toBe('any instance (may reach the primary; not read-only)')
+ expect(info.endpoints.find((e) => e.role === 'ro')?.selects).toBe('standbys only (read-only)')
+ const extra = cnpgConnectInfo(cluster({ managed: { services: { additional: [{ selectorType: 'r', serviceTemplate: { metadata: { name: 'pg-any' } } }] } } }))
+ expect(extra.endpoints.find((e) => e.name === 'pg-any')?.selects).toContain('not read-only')
+ })
+
+ it('reads database, owner and Secret the way CloudNativePG resolves them: recovery, pg_basebackup, then initdb', () => {
+ const info = cnpgConnectInfo(
+ cluster({
+ bootstrap: {
+ initdb: { database: 'orders', owner: 'orders_owner', secret: { name: 'init-secret' } },
+ recovery: { database: 'restored', secret: { name: 'recovery-secret' } },
+ },
+ }),
+ )
+ expect(info.database).toEqual({ value: 'restored', source: 'spec.bootstrap.recovery.database' })
+ expect(info.owner).toEqual({ value: 'orders_owner', source: 'spec.bootstrap.initdb.owner' })
+ expect(info.secret).toEqual({ name: 'recovery-secret', source: 'spec.bootstrap.recovery.secret.name', byConvention: false })
+ })
+
+ it('treats a distributed-topology replica (replica.primary names another cluster) as a replica, like the operator', () => {
+ expect(cnpgConnectInfo(cluster({ replica: { primary: 'pg-east', source: 'pg-east' } })).replicaCluster).toBe(true)
+ expect(cnpgConnectInfo(cluster({ replica: { primary: 'pg', source: 'pg-east' } })).replicaCluster).toBe(false)
+ expect(cnpgConnectInfo(cluster({ replica: { enabled: true, source: 'pg-east' } })).replicaCluster).toBe(true)
+ expect(cnpgConnectInfo(cluster()).replicaCluster).toBe(false)
+ })
+
+ it('owner defaults to the database name', () => {
+ expect(cnpgConnectInfo(cluster({ bootstrap: { initdb: { database: 'orders' } } })).owner).toEqual({
+ value: 'orders',
+ source: "CloudNativePG's default: the database's name",
+ })
+ })
+
+ it('does not invent a database for a monolithic import', () => {
+ const info = cnpgConnectInfo(cluster({ bootstrap: { initdb: { import: { type: 'monolith' } } } }))
+ expect(info.database.value).toBeUndefined()
+ expect(info.owner.value).toBeUndefined()
+ })
+
+ it('drops disabled default Services and adds managed and Pooler Services with their ports', () => {
+ const info = cnpgConnectInfo(
+ cluster({
+ managed: {
+ services: {
+ disabledDefaultServices: ['ro', 'r'],
+ additional: [{ selectorType: 'rw', serviceTemplate: { metadata: { name: 'pg-lb' }, spec: { type: 'LoadBalancer', ports: [{ port: 6543 }] } } }],
+ },
+ },
+ }),
+ [
+ { metadata: { name: 'pg-pooler-ro', namespace: 'db' }, spec: { cluster: { name: 'pg' }, type: 'ro' } },
+ { metadata: { name: 'other', namespace: 'db' }, spec: { cluster: { name: 'pg2' } } },
+ { metadata: { name: 'pg-pooler', namespace: 'elsewhere' }, spec: { cluster: { name: 'pg' } } },
+ ],
+ )
+ expect(info.disabled).toEqual(['ro', 'r'])
+ expect(info.endpoints.map((e) => `${e.role} ${e.name}:${e.port}${e.portFromTemplate ? '*' : ''}`)).toEqual(['rw pg-rw:5432', 'additional pg-lb:6543*', 'pooler pg-pooler-ro:5432'])
+ expect(info.endpoints[2].poolerType).toBe('ro')
+ expect(info.endpoints[1].selects).toBe('the primary (read-write)')
+ })
+
+ it('builds templates with a password placeholder only', () => {
+ const info = cnpgConnectInfo(cluster({ bootstrap: { initdb: { database: 'orders', owner: 'o w' } } }))
+ expect(cnpgConnectionURI(info.endpoints[0], info)).toBe('postgresql://o%20w:@pg-rw.db.svc:5432/orders')
+ expect(cnpgPsqlCommand(info.endpoints[0], info)).toBe("psql -h pg-rw.db.svc -p 5432 -U 'o w' -d orders")
+ const quoted = cnpgConnectInfo(cluster({ bootstrap: { initdb: { database: "it's" } } }))
+ expect(cnpgPsqlCommand(quoted.endpoints[0], quoted)).toBe(`psql -h pg-rw.db.svc -p 5432 -U 'it'\\''s' -d 'it'\\''s'`)
+ })
+})
+
+it('pins port-forward to a shell-quoted kubeconfig context when supplied', () => {
+ const ep = cnpgConnectInfo({ metadata: { name: 'orders', namespace: 'prod' } }).endpoints[0]
+ expect(cnpgPortForwardCommand(ep, 'prod', 5432, 'kind-cnpg')).toBe('kubectl --context kind-cnpg -n prod port-forward service/orders-rw 5432:5432')
+ expect(cnpgPortForwardCommand(ep, 'prod', 5432, 'test context')).toContain("--context 'test context'")
+})
+
+it('reads rw endpoints separately from standby and any-instance Pod readiness', () => {
+ const eps = cnpgConnectInfo({ metadata: { name: 'orders', namespace: 'prod' } }).endpoints
+ const ha = { rwEndpoints: { state: 'ok', pods: ['orders-1'] }, pods: { state: 'ok' }, instances: [{ pod: 'orders-1', role: 'primary', ready: true }, { pod: 'orders-2', role: 'replica', ready: false }] } as CNPGClusterHA
+ expect(eps.map((ep) => cnpgEndpointAvailability(ep, ha).text)).toEqual(['Ready endpoints', 'Unavailable: no ready standby', 'Ready instance observed'])
+ expect(cnpgEndpointAvailability(eps[1], { ...ha, instances: [] }).text).toBe('Unavailable: no ready standby')
+ expect(eps.map((ep) => cnpgEndpointAvailability(ep).text)).toEqual(['Not checked', 'Not checked', 'Not checked'])
+ expect(cnpgEndpointAvailability(eps[0], { ...ha, rwEndpoints: { ...ha.rwEndpoints, state: 'denied' } }).text).toBe('Not checked')
+ expect(cnpgEndpointAvailability(eps[1], { ...ha, pods: { state: 'denied' } }).text).toBe('Not checked')
+ expect(cnpgEndpointAvailability(eps[1], { ...ha, instances: [{ ...ha.instances[0], role: 'unknown' }] }).text).toBe('Not checked')
+})
+
+it('keeps Pooler routing faithful for any-instance and unrecognized selectors', () => {
+ for (const [type, selects] of [['rw', 'the primary'], ['ro', 'the standbys'], ['r', 'any instance (may reach the primary; not read-only)'], ['future', 'selector future']]) {
+ const ep = cnpgConnectInfo(cluster(), [{ metadata: { name: 'pooler', namespace: 'db' }, spec: { cluster: { name: 'pg' }, type } }]).endpoints.find((e) => e.role === 'pooler')!
+ expect(ep.poolerType).toBe(type)
+ expect(ep.selects).toBe(`PgBouncer in front of ${selects}`)
+ expect(cnpgEndpointAvailability(ep).source).toBe('This Service’s availability has not been read')
+ }
+})
+it('names the denied grant, failed read, unknown role and unread source alongside Not checked', () => {
+ const eps = cnpgConnectInfo(cluster()).endpoints
+ const ha = { rwEndpoints: { state: 'denied', grant: { verb: 'list', group: 'discovery.k8s.io', resource: 'endpointslices', namespace: 'db' } }, pods: { state: 'error', reason: 'API timeout' }, instances: [] } as unknown as CNPGClusterHA
+ expect(cnpgEndpointAvailability(eps[0], ha)).toEqual({ text: 'Not checked', source: 'No access to Service EndpointSlices (needs list endpointslices (discovery.k8s.io) in namespace db)' })
+ expect(cnpgEndpointAvailability(eps[1], ha)).toEqual({ text: 'Not checked', source: 'instance Pods could not be read: API timeout' })
+ expect(cnpgEndpointAvailability(eps[0], undefined, 'Reading availability…').source).toBe('Reading availability…')
+ expect(cnpgEndpointAvailability(eps[0], undefined, 'Availability could not be read: timeout').source).toBe('Availability could not be read: timeout')
+ expect(cnpgEndpointAvailability(eps[1], { ...ha, pods: { state: 'ok' }, instances: [{ role: 'unknown', ready: true }] } as CNPGClusterHA).source).toBe('Ready instance roles were not reported')
+})
diff --git a/packages/k8s-ui/src/components/cnpg/connect.ts b/packages/k8s-ui/src/components/cnpg/connect.ts
new file mode 100644
index 0000000000..fe3f4df347
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/connect.ts
@@ -0,0 +1,155 @@
+import { cnpgHASourceText, type CNPGClusterHA } from './ha'
+import { getCNPGClusterIsReplica } from '../resources/resource-utils-cnpg'
+
+/** The port CloudNativePG gives PostgreSQL and PgBouncer when a Service template sets none. */
+export const CNPG_DEFAULT_PORT = 5432
+
+export type CNPGConnectRole = 'rw' | 'ro' | 'r' | 'pooler' | 'additional'
+
+export interface CNPGConnectEndpoint {
+ role: CNPGConnectRole
+ /** The Service name. */
+ name: string
+ host: string
+ port: number
+ /** True when the port comes from a Service template rather than CloudNativePG's default. */
+ portFromTemplate: boolean
+ /** What the Service selects: the primary, standbys only, any instance, or through a Pooler. */
+ selects: string
+ /** Pooler `spec.type`, for pooler endpoints. */
+ poolerType?: string
+}
+
+export interface CNPGConnectValue {
+ value?: string
+ /** The field the value was read from, or why it is inferred. */
+ source: string
+}
+
+export interface CNPGConnectInfo {
+ endpoints: CNPGConnectEndpoint[]
+ database: CNPGConnectValue
+ owner: CNPGConnectValue
+ secret: { name: string; source: string; byConvention: boolean }
+ /** A replica cluster's services all reach instances that only replay. */
+ replicaCluster: boolean
+ /** Services named in `spec.managed.services.disabledDefaultServices`. */
+ disabled: ('ro' | 'r')[]
+}
+
+const SELECTS: Record<'rw' | 'ro' | 'r', string> = {
+ rw: 'the primary (read-write)',
+ ro: 'standbys only (read-only)',
+ // -r selects every instance Pod (cnpg.io/podRole=instance), the primary included.
+ r: 'any instance (may reach the primary; not read-only)',
+}
+
+const BOOTSTRAP_ORDER = ['recovery', 'pg_basebackup', 'initdb'] as const
+
+function templatePort(template: any): number | undefined {
+ const port = template?.spec?.ports?.[0]?.port
+ return typeof port === 'number' && port > 0 ? port : undefined
+}
+
+function endpoint(role: CNPGConnectRole, name: string, ns: string, selects: string, template?: any): CNPGConnectEndpoint {
+ const port = templatePort(template)
+ return { role, name, host: `${name}.${ns}.svc`, port: port ?? CNPG_DEFAULT_PORT, portFromTemplate: port !== undefined, selects }
+}
+
+// Mirrors CloudNativePG's GetApplicationDatabaseName / Owner / SecretName:
+// recovery, then pg_basebackup, then initdb, first non-empty wins.
+function fromBootstrap(bootstrap: any, pick: (section: any) => string | undefined, field: string): CNPGConnectValue | undefined {
+ for (const method of BOOTSTRAP_ORDER) {
+ const v = pick(bootstrap?.[method])
+ if (v) return { value: v, source: `spec.bootstrap.${method}.${field}` }
+ }
+ return undefined
+}
+
+/**
+ * How applications reach a CloudNativePG Cluster, derived from its spec and
+ * the Poolers that front it. Nothing here reads a Secret: the credentials
+ * Secret is named from the spec or CloudNativePG's `-app` convention.
+ */
+export function cnpgConnectInfo(cluster: any, poolers: any[] = []): CNPGConnectInfo {
+ const name: string = cluster?.metadata?.name ?? ''
+ const ns: string = cluster?.metadata?.namespace ?? ''
+ const spec = cluster?.spec ?? {}
+ const bootstrap = spec.bootstrap
+
+ const disabledRaw: unknown[] = spec.managed?.services?.disabledDefaultServices ?? []
+ const disabled = (['ro', 'r'] as const).filter((t) => disabledRaw.includes(t))
+ const endpoints: CNPGConnectEndpoint[] = [endpoint('rw', `${name}-rw`, ns, SELECTS.rw)]
+ for (const t of ['ro', 'r'] as const) if (!disabled.includes(t)) endpoints.push(endpoint(t, `${name}-${t}`, ns, SELECTS[t]))
+ for (const svc of spec.managed?.services?.additional ?? []) {
+ const svcName = svc?.serviceTemplate?.metadata?.name
+ if (!svcName) continue
+ const sel = svc.selectorType as 'rw' | 'ro' | 'r'
+ endpoints.push(endpoint('additional', svcName, ns, SELECTS[sel] ?? `selector ${sel ?? 'unknown'}`, svc.serviceTemplate))
+ }
+ for (const p of poolers) {
+ if (p?.metadata?.namespace !== ns || p?.spec?.cluster?.name !== name || !p?.metadata?.name) continue
+ const type: string = p.spec?.type || 'rw'
+ endpoints.push({
+ ...endpoint('pooler', p.metadata.name, ns, `PgBouncer in front of ${type === 'rw' ? 'the primary' : type === 'ro' ? 'the standbys' : type === 'r' ? 'any instance (may reach the primary; not read-only)' : `selector ${type}`}`, p.spec?.serviceTemplate),
+ poolerType: type,
+ })
+ }
+
+ const monolith = bootstrap?.initdb?.import?.type === 'monolith'
+ const database =
+ fromBootstrap(bootstrap, (s) => s?.database, 'database') ??
+ (monolith
+ ? { source: 'not set: a monolithic import creates no application database' }
+ : { value: 'app', source: "CloudNativePG's default" })
+ const owner =
+ fromBootstrap(bootstrap, (s) => s?.owner, 'owner') ??
+ (database.value ? { value: database.value, source: "CloudNativePG's default: the database's name" } : { source: 'not set' })
+ const secretFromSpec = fromBootstrap(bootstrap, (s) => s?.secret?.name, 'secret.name')
+ const secret = secretFromSpec?.value
+ ? { name: secretFromSpec.value, source: secretFromSpec.source, byConvention: false }
+ : { name: `${name}-app`, source: 'name by convention (-app)', byConvention: true }
+
+ return {
+ endpoints,
+ database,
+ owner,
+ secret,
+ replicaCluster: cluster ? getCNPGClusterIsReplica(cluster) : false,
+ disabled,
+ }
+}
+
+/** A connection URI with the password left as a placeholder; never a real credential. */
+export function cnpgConnectionURI(ep: CNPGConnectEndpoint, info: CNPGConnectInfo): string {
+ const user = encodeURIComponent(info.owner.value ?? '')
+ const db = encodeURIComponent(info.database.value ?? '')
+ return `postgresql://${user}:@${ep.host}:${ep.port}/${db}`
+}
+
+function shellWord(v: string): string {
+ return /^[A-Za-z0-9_.@%+=:,/-]+$/.test(v) ? v : `'${v.replace(/'/g, `'\\''`)}'`
+}
+
+/** psql prompts for the password; nothing secret is put on the command line. */
+export function cnpgPsqlCommand(ep: CNPGConnectEndpoint, info: CNPGConnectInfo): string {
+ return `psql -h ${ep.host} -p ${ep.port} -U ${shellWord(info.owner.value ?? '')} -d ${shellWord(info.database.value ?? '')}`
+}
+
+export function cnpgPortForwardCommand(ep: CNPGConnectEndpoint, namespace: string, localPort = CNPG_DEFAULT_PORT, kubeconfigContext?: string): string {
+ return `kubectl${kubeconfigContext ? ` --context ${shellWord(kubeconfigContext)}` : ''} -n ${shellWord(namespace)} port-forward ${shellWord(`service/${ep.name}`)} ${localPort}:${ep.port}`
+}
+
+export function cnpgEndpointAvailability(ep: CNPGConnectEndpoint, ha?: CNPGClusterHA, unreadReason?: string): { text: string; source?: string } {
+ if (ep.role === 'rw' && ha?.rwEndpoints.state === 'ok') {
+ return { text: ha.rwEndpoints.pods.length > 0 ? 'Ready endpoints' : 'Unavailable: no ready endpoint', source: 'Service EndpointSlices' }
+ }
+ if ((ep.role === 'ro' || ep.role === 'r') && ha?.pods.state === 'ok') {
+ const candidates = ep.role === 'ro' ? ha.instances.filter((i) => i.role === 'replica') : ha.instances
+ if (ep.role === 'ro' && ha.instances.some((i) => i.ready && i.role === 'unknown') && !candidates.some((i) => i.ready)) return { text: 'Not checked', source: 'Ready instance roles were not reported' }
+ return { text: candidates.some((i) => i.ready) ? 'Ready instance observed' : ep.role === 'ro' ? 'Unavailable: no ready standby' : 'Unavailable: no ready instance', source: 'Instance Pod readiness' }
+ }
+ const source = ep.role === 'rw' ? ha?.rwEndpoints : ep.role === 'ro' || ep.role === 'r' ? ha?.pods : undefined
+ const what = ep.role === 'rw' ? 'Service EndpointSlices' : 'instance Pods'
+ return { text: 'Not checked', source: ep.role === 'pooler' || ep.role === 'additional' ? 'This Service’s availability has not been read' : source ? cnpgHASourceText(source, what) : unreadReason ?? 'HA evidence has not been read' }
+}
diff --git a/packages/k8s-ui/src/components/cnpg/databaseRole.test.ts b/packages/k8s-ui/src/components/cnpg/databaseRole.test.ts
new file mode 100644
index 0000000000..5ef6fb0bbd
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/databaseRole.test.ts
@@ -0,0 +1,38 @@
+import { describe, expect, it } from 'vitest'
+import { cnpgDatabaseRoleFacts, cnpgDatabaseRoleMeta } from './databaseRole'
+
+const role = (spec: any, status: any = {}) => ({
+ apiVersion: 'postgresql.cnpg.io/v1',
+ kind: 'DatabaseRole',
+ metadata: { name: 'app-reader', namespace: 'db' },
+ spec: { cluster: { name: 'pg' }, name: 'reader', ...spec },
+ status,
+})
+
+describe('cnpgDatabaseRoleFacts', () => {
+ it('reports the Cluster’s managed.roles precedence as three-valued', () => {
+ const cluster = { spec: { managed: { roles: [{ name: 'reader' }] } } }
+ expect(cnpgDatabaseRoleFacts(role({}), cluster).overriddenByCluster).toBe(true)
+ expect(cnpgDatabaseRoleFacts(role({}), { spec: {} }).overriddenByCluster).toBe(false)
+ expect(cnpgDatabaseRoleFacts(role({}), null).overriddenByCluster).toBeNull()
+ })
+
+ it('reads applied three ways and the operator message', () => {
+ expect(cnpgDatabaseRoleFacts(role({}), null).state).toBe('pending')
+ const failed = cnpgDatabaseRoleFacts(role({}, { applied: false, message: 'database role is already managed by the CNPG cluster' }), null)
+ expect(failed.state).toBe('failed')
+ expect(failed.message).toContain('already managed')
+ })
+
+ it('treats an omitted login as false and names the client certificate Secret', () => {
+ const f = cnpgDatabaseRoleFacts(role({ clientCertificate: {}, validUntil: '2027-01-01T00:00:00Z' }, { clientCertificate: { expiration: '2026-12-01T00:00:00Z' } }), null)
+ expect(f.login).toBe(false)
+ expect(f.clientCertificate?.secret).toBe('app-reader-client-cert')
+ expect(cnpgDatabaseRoleMeta(f)).toBe('no login · password valid until 2027-01-01T00:00:00Z · client cert until 2026-12-01T00:00:00Z')
+ })
+})
+
+it('keeps a DatabaseRole pending until its current generation is observed', () => {
+ const r = { ...role({}, { applied: true, observedGeneration: 1 }), metadata: { name: 'reader', generation: 2 } }
+ expect(cnpgDatabaseRoleFacts(r, null).state).toBe('pending')
+})
diff --git a/packages/k8s-ui/src/components/cnpg/databaseRole.ts b/packages/k8s-ui/src/components/cnpg/databaseRole.ts
new file mode 100644
index 0000000000..5c79e40e59
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/databaseRole.ts
@@ -0,0 +1,80 @@
+import { appliedFact } from './relations'
+
+// CloudNativePG DatabaseRole (1.30+): one PostgreSQL role declared as its own
+// object. The same role name in the Cluster's spec.managed.roles always wins:
+// the operator does not reconcile the DatabaseRole and reports it not applied
+// ("database role is already managed by the CNPG cluster").
+
+export type CNPGRoleState = 'applied' | 'failed' | 'pending'
+
+export interface CNPGDatabaseRoleFacts {
+ pgName: string
+ cluster?: string
+ state: CNPGRoleState
+ message?: string
+ /** Absent in the spec means false: the operator's field is omitempty. */
+ login: boolean
+ superuser: boolean
+ /** spec.validUntil: when the role's password stops being accepted (PostgreSQL VALID UNTIL). */
+ passwordValidUntil?: string
+ passwordSecret?: string
+ passwordDisabled: boolean
+ clientCertificate?: { enabled: boolean; secret: string; expiration?: string; message?: string }
+ reclaimPolicy: 'delete' | 'retain'
+ /**
+ * true when the target Cluster declares the same role in spec.managed.roles,
+ * false when it does not, null when the Cluster is not visible.
+ */
+ overriddenByCluster: boolean | null
+}
+
+export function cnpgRoleState(obj: any): CNPGRoleState {
+ const fact = appliedFact(obj)
+ if (fact.tone === 'healthy') return 'applied'
+ if (fact.tone === 'unhealthy') return 'failed'
+ return 'pending'
+}
+
+export function cnpgDatabaseRoleFacts(role: any, cluster: any | null | undefined): CNPGDatabaseRoleFacts {
+ const spec = role?.spec ?? {}
+ const status = role?.status ?? {}
+ const pgName: string = spec.name ?? role?.metadata?.name ?? ''
+ let overriddenByCluster: boolean | null = null
+ if (cluster) {
+ const roles: any[] = Array.isArray(cluster?.spec?.managed?.roles) ? cluster.spec.managed.roles : []
+ overriddenByCluster = roles.some((r) => r?.name === pgName)
+ }
+ const cc = spec.clientCertificate
+ const ccEnabled = !!cc && cc.enabled !== false
+ return {
+ pgName,
+ cluster: spec.cluster?.name,
+ state: cnpgRoleState(role),
+ message: status.message || undefined,
+ login: spec.login === true,
+ superuser: spec.superuser === true,
+ passwordValidUntil: spec.validUntil || undefined,
+ passwordSecret: spec.passwordSecret?.name || undefined,
+ passwordDisabled: spec.disablePassword === true,
+ clientCertificate: ccEnabled
+ ? {
+ enabled: true,
+ secret: `${role?.metadata?.name}-client-cert`,
+ expiration: status.clientCertificate?.expiration || undefined,
+ message: status.clientCertificate?.message || undefined,
+ }
+ : undefined,
+ reclaimPolicy: spec.databaseRoleReclaimPolicy === 'delete' ? 'delete' : 'retain',
+ overriddenByCluster,
+ }
+}
+
+/** Short, factual meta line for a list row. */
+export function cnpgDatabaseRoleMeta(f: CNPGDatabaseRoleFacts): string {
+ const parts: string[] = []
+ if (f.overriddenByCluster) parts.push('overridden by spec.managed.roles')
+ parts.push(f.login ? 'login' : 'no login')
+ if (f.passwordValidUntil) parts.push(`password valid until ${f.passwordValidUntil}`)
+ if (f.clientCertificate) parts.push(f.clientCertificate.expiration ? `client cert until ${f.clientCertificate.expiration}` : 'client cert expiry not reported')
+ return parts.join(' · ')
+}
diff --git a/packages/k8s-ui/src/components/cnpg/ha.test.ts b/packages/k8s-ui/src/components/cnpg/ha.test.ts
new file mode 100644
index 0000000000..47451dc06d
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/ha.test.ts
@@ -0,0 +1,426 @@
+import { describe, expect, it } from 'vitest'
+import {
+ cnpgCertificateViews,
+ cnpgCertificatesSummary,
+ cnpgHASummary,
+ cnpgImageDrift,
+ cnpgDimensions,
+ cnpgPDBFact,
+ cnpgLeaseHolderPod,
+ cnpgLiveGap,
+ cnpgPendingRestart,
+ cnpgQuorumFact,
+ cnpgZoneSpread,
+ type CNPGClusterHA,
+ type CNPGHAQuorum,
+} from './ha'
+import { cnpgLagTone, type CNPGFleetRow } from './workspace'
+
+function ha(over: Partial = {}): CNPGClusterHA {
+ return {
+ cluster: { namespace: 'db', name: 'pg', uid: 'u' },
+ sampledAt: '2026-09-30T12:00:00Z',
+ desiredImage: 'pg:17',
+ instances: [
+ { pod: 'pg-1', podUID: 'a', role: 'primary', ready: true, node: 'n1', zone: 'z1', restartCount: 0, image: 'pg:17', imageMatches: true },
+ { pod: 'pg-2', podUID: 'b', role: 'replica', ready: true, node: 'n2', zone: 'z2', restartCount: 0, image: 'pg:17', imageMatches: true },
+ ],
+ pods: { state: 'ok' },
+ nodes: { state: 'ok' },
+ quorum: { enabled: false, object: { state: 'notFound' } },
+ pdbs: { state: 'ok', enabled: true, items: [] },
+ primaryLease: { state: 'notFound' },
+ operatorLease: { state: 'unavailable' },
+ jobs: { state: 'ok', items: [] },
+ rwEndpoints: { state: 'ok', service: 'pg-rw', pods: ['pg-1'] },
+ certificates: [],
+ maintenance: { declared: false, inProgress: false, reusePVC: true },
+ ...over,
+ }
+}
+
+function row(over: Partial = {}): CNPGFleetRow {
+ return {
+ key: 'db/pg',
+ namespace: 'db',
+ name: 'pg',
+ cluster: { status: { currentPrimary: 'pg-1' } },
+ controllerStatus: { text: 'Healthy', level: 'healthy' },
+ instances: { ready: 2, desired: 2 },
+ pods: [
+ { name: 'pg-1', role: 'primary', ready: true },
+ { name: 'pg-2', role: 'replica', ready: true },
+ ],
+ replicaCluster: null,
+ hibernated: false,
+ pgVersion: '17',
+ catalog: null,
+ replication: { text: 'Lag unknown', tone: 'unknown' },
+ protection: {
+ schedule: { text: 'Active', tone: 'healthy', names: [] },
+ destination: { text: 'ObjectStore', tone: 'healthy', method: 'plugin' },
+ lastSuccessfulBackup: { text: '1 h ago', tone: 'healthy' },
+ walArchiving: { text: 'Archiving', tone: 'healthy' },
+ recoveryWindow: { text: 'x', tone: 'healthy' },
+ restoreValidation: { text: 'None recorded', tone: 'unknown' },
+ summary: { text: 'ok', tone: 'healthy' },
+ },
+ declarations: { summary: { text: 'None', tone: 'neutral' }, total: 0, failed: 0, pending: 0 },
+ poolers: [],
+ poolersKnown: true,
+ problems: [],
+ attention: false,
+ categories: new Set(),
+ ...over,
+ }
+}
+
+it('does not call successful snapshot protection WAL archiving when no archive is configured', () => {
+ const r = row()
+ r.protection.destination = { text: 'Volume snapshots', tone: 'neutral', method: 'volumeSnapshot' }
+ r.protection.walArchiving = { text: 'No archive destination configured', tone: 'neutral', source: 'Cluster spec' }
+ r.protection.lastSuccessfulBackup = { text: 'Completed', tone: 'healthy', source: 'Backup snapshot' }
+ expect(cnpgDimensions({ row: r }).find((d) => d.id === 'protection')).toMatchObject({ text: 'Backup completed', tone: 'healthy', source: 'Backup snapshot' })
+})
+
+describe('cnpgZoneSpread', () => {
+ it('groups instances by zone and names the primary’s', () => {
+ const s = cnpgZoneSpread(ha())
+ expect(s.known).toBe(true)
+ expect(s.zones.map((z) => z.zone)).toEqual(['z1', 'z2'])
+ expect(s.primaryZone).toBe('z1')
+ expect(s.singleZone).toBe(false)
+ })
+ it('flags every instance in one zone and a shared node', () => {
+ const s = cnpgZoneSpread(
+ ha({
+ instances: [
+ { pod: 'pg-1', podUID: 'a', role: 'primary', ready: true, node: 'n1', zone: 'z1', restartCount: 0 },
+ { pod: 'pg-2', podUID: 'b', role: 'replica', ready: true, node: 'n1', zone: 'z1', restartCount: 0 },
+ ],
+ }),
+ )
+ expect(s.singleZone).toBe(true)
+ expect(s.sharedNode).toBe(true)
+ })
+ it('is unknown, not single-zone, when Nodes are not readable', () => {
+ const s = cnpgZoneSpread(ha({ nodes: { state: 'denied', grant: { verb: 'get', resource: 'nodes' } } }))
+ expect(s.known).toBe(false)
+ expect(s.singleZone).toBe(false)
+ })
+})
+
+describe('cnpgQuorumFact', () => {
+ const q = (over: Partial): CNPGHAQuorum => ({ enabled: true, object: { state: 'ok' }, ...over })
+ it('states R + W > N when it holds', () => {
+ const f = cnpgQuorumFact(q({ n: 2, w: 1, r: 2, holds: true, status: { standbyNames: ['a', 'b'], standbyNumber: 1 } }))
+ expect(f.text).toContain('R + W > N')
+ expect(f.tone).toBe('healthy')
+ })
+ it('states when it does not', () => {
+ const f = cnpgQuorumFact(q({ n: 2, w: 1, r: 1, holds: false }))
+ expect(f.text).toContain('R + W ≤ N')
+ expect(f.tone).toBe('degraded')
+ })
+ it('a reset object is no configuration, not a pass', () => {
+ expect(cnpgQuorumFact(q({ status: { standbyNames: [], standbyNumber: 0 } })).text).toContain('no synchronous configuration recorded')
+ })
+ it('an unreadable object is unknown', () => {
+ expect(cnpgQuorumFact(q({ object: { state: 'denied', grant: { verb: 'get', group: 'postgresql.cnpg.io', resource: 'failoverquorums', namespace: 'db' } } })).tone).toBe('unknown')
+ })
+})
+
+describe('cnpgPDBFact', () => {
+ it('distinguishes disabled from missing and unreadable', () => {
+ expect(cnpgPDBFact({ state: 'ok', enabled: false, items: [] }).tone).toBe('neutral')
+ expect(cnpgPDBFact({ state: 'ok', enabled: true, items: [] }).tone).toBe('degraded')
+ expect(cnpgPDBFact({ state: 'denied', enabled: true, items: [] }).tone).toBe('unknown')
+ })
+})
+
+describe('cnpgCertificateViews', () => {
+ const now = Date.parse('2026-09-30T00:00:00Z')
+ const at = (days: number) => new Date(now + days * 86_400_000).toISOString()
+ it('uses the issues-engine thresholds and owner', () => {
+ const v = cnpgCertificateViews(
+ [
+ { secret: 'u-20', raw: '', expiresAt: at(20), renewal: 'user' },
+ { secret: 'u-3', raw: '', expiresAt: at(3), renewal: 'user' },
+ { secret: 'o-3', raw: '', expiresAt: at(3), renewal: 'operator' },
+ { secret: 'x', raw: 'garbage', renewal: 'operator' },
+ ],
+ now,
+ )
+ expect(v.map((c) => c.tone)).toEqual(['degraded', 'unhealthy', 'healthy', 'unknown'])
+ expect(v[0].daysLeft).toBe(20)
+ })
+})
+
+describe('cnpgPendingRestart', () => {
+ it('is unknown without live data and names the instances otherwise', () => {
+ expect(cnpgPendingRestart(undefined).known).toBe(false)
+ const p = cnpgPendingRestart([
+ { pod: 'pg-1', state: 'ok', pendingRestart: true, pendingRestartForDecrease: true },
+ { pod: 'pg-2', state: 'unreachable' },
+ ])
+ expect(p.pods).toEqual(['pg-1'])
+ expect(p.forDecrease).toBe(true)
+ expect(p.known).toBe(false)
+ })
+ it('does not count an incomplete report as "no restart pending"', () => {
+ const p = cnpgPendingRestart([
+ { pod: 'pg-1', state: 'ok' },
+ { pod: 'pg-2', state: 'partial', incomplete: true },
+ ])
+ expect(p.known).toBe(false)
+ expect(cnpgPendingRestart([{ pod: 'pg-2', state: 'partial', incomplete: true }]).known).toBe(false)
+ })
+})
+
+describe('cnpgLiveGap', () => {
+ it('names the missing grant when there is no runtime read', () => {
+ expect(cnpgLiveGap(undefined, 'needs get pods/proxy in db')).toBe('needs get pods/proxy in db')
+ })
+ it('names each instance that did not report, and why, when runtime access exists', () => {
+ const gap = cnpgLiveGap([
+ { pod: 'pg-1', state: 'ok' },
+ { pod: 'pg-2', state: 'unreachable', reason: 'PostgreSQL is not running on this instance' },
+ { pod: 'pg-3', state: 'partial', incomplete: true, reason: 'pg_rewind is running' },
+ ])
+ expect(gap).toBe('pg-2 did not report (PostgreSQL is not running on this instance); pg-3 reported incompletely (pg_rewind is running)')
+ expect(gap).not.toContain('runtime access')
+ expect(cnpgLiveGap([{ pod: 'pg-1', state: 'ok' }])).toBe('')
+ })
+})
+
+describe('cnpgDimensions', () => {
+ it('reads each dimension from its own source and leaves storage unassessed', () => {
+ const d = cnpgDimensions({ row: row(), ha: ha(), replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 0 } })
+ expect(d.map((x) => [x.id, x.tone])).toEqual([
+ ['serving', 'healthy'],
+ ['replication', 'healthy'],
+ ['storage', 'unknown'],
+ ['protection', 'healthy'],
+ ])
+ expect(d.find((x) => x.id === 'protection')?.label).toBe('Backups')
+ })
+ it('storage follows the measured disk fact, and stays unassessed without a measurement', () => {
+ const measured = cnpgDimensions({ row: row({ disk: { text: '91% used', tone: 'unhealthy', source: 'Fullest: data of pg-1' } }) })
+ expect(measured[2]).toMatchObject({ id: 'storage', tone: 'unhealthy', text: '91% used' })
+ const unmeasured = cnpgDimensions({ row: row({ disk: { text: 'No usage metrics', tone: 'unknown', source: 'needs Prometheus' } }) })
+ expect(unmeasured[2]).toMatchObject({ tone: 'unknown', text: 'No usage metrics: needs Prometheus', source: '' })
+ })
+ it('names WAL an inactive slot holds even while volume usage is unassessed', () => {
+ const slot = { id: 'slot:pg/pg:_cnpg_pg_2', reason: 'CNPGInactiveSlot', slot: '_cnpg_pg_2', severity: 'warning', category: 'availability', title: 'Inactive slot _cnpg_pg_2 holds 4.5 GiB of WAL on pg-1 for pg-2', subject: { kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: 'pg', name: 'pg' }, source: 'measurement' } as const
+ const r = row({ key: 'pg/pg', problems: [slot], disk: { text: 'No usage metrics', tone: 'unknown', source: 'no series' } })
+ const d = cnpgDimensions({ row: r })
+ expect(d[2]).toMatchObject({ id: 'storage', tone: 'degraded', text: 'WAL held by an inactive slot' })
+ expect(d[2].source).toContain('4.5 GiB')
+ expect(d[2].source).toContain('No usage metrics: no series')
+ const measured = cnpgDimensions({ row: { ...r, disk: { text: '40% used', tone: 'healthy', source: 'Fullest' } } })
+ expect(measured[2]).toMatchObject({ tone: 'degraded', text: '40% used · WAL held by an inactive slot' })
+ })
+ it('a standby another source saw receiving nothing keeps Replication from reading unassessed or calm', () => {
+ const gap = { id: 'standby:pg/pg:pg-2', reason: 'CNPGStandbyNotReceiving', severity: 'warning', category: 'availability', title: 'pg-2 is not receiving WAL from the primary', subject: { kind: 'Pod', group: '', namespace: 'pg', name: 'pg-2' }, source: 'measurement' } as const
+ const r = row({ key: 'pg/pg', problems: [gap] })
+ expect(cnpgDimensions({ row: r })[1]).toMatchObject({ tone: 'degraded', text: 'pg-2 not receiving WAL' })
+ const live = cnpgDimensions({ row: r, replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 0 } })[1]
+ expect(live.tone).toBe('degraded')
+ expect(live.text).toContain('pg-2 not receiving WAL')
+ })
+ it('counts expected standbys from spec.instances, not from the Pods still running', () => {
+ // spec.instances 3, only the primary's Pod exists, nothing streams.
+ const r = row({ instances: { ready: 1, desired: 3 }, pods: [{ name: 'pg-1', role: 'primary', ready: true }] })
+ const d = cnpgDimensions({ row: r, replication: { streaming: 0, standbys: 0 } })
+ expect(d[1]).toMatchObject({ tone: 'degraded', text: '0 of 2 expected standbys streaming' })
+ const one = cnpgDimensions({ row: r, replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 0 } })
+ expect(one[1]).toMatchObject({ tone: 'degraded', text: '1 of 2 expected standbys streaming' })
+ })
+ it('replication lag takes the same tone as the Replication fact', () => {
+ const at = (lag: number) => cnpgDimensions({ row: row(), replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: lag } })[1]
+ expect(at(2)).toMatchObject({ tone: 'healthy' })
+ expect(at(8)).toMatchObject({ tone: 'degraded', text: '1 of 1 streaming · replay delay 8.0 s' })
+ expect(at(23.6)).toMatchObject({ tone: 'degraded', text: '1 of 1 streaming · replay delay 23 s' })
+ expect(at(72)).toMatchObject({ tone: 'unhealthy', text: '1 of 1 streaming · replay delay 72 s' })
+ expect(at(72).tone).toBe(cnpgLagTone(72))
+ })
+ it('a missing standby does not hide a severe lag on the one that streams', () => {
+ const r = row({ instances: { ready: 2, desired: 3 } })
+ const d = cnpgDimensions({ row: r, replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 72 } })[1]
+ expect(d).toMatchObject({ tone: 'unhealthy', text: '1 of 2 expected standbys streaming · replay delay 72 s' })
+ expect(cnpgDimensions({ row: r, replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 0.2 } })[1]).toMatchObject({
+ tone: 'degraded',
+ text: '1 of 2 expected standbys streaming',
+ })
+ })
+ it('replication is unknown when spec.instances is not reported', () => {
+ const d = cnpgDimensions({ row: row({ instances: { ready: null, desired: null } }), replication: { streaming: 0, standbys: 0 } })
+ expect(d[1].tone).toBe('unknown')
+ })
+ it('replication is unassessed without runtime, never healthy', () => {
+ const d = cnpgDimensions({ row: row() })
+ expect(d.find((x) => x.id === 'replication')?.text).toBe('unassessed')
+ })
+ it('serving fails when the rw endpoint is not on the primary', () => {
+ const d = cnpgDimensions({ row: row(), ha: ha({ rwEndpoints: { state: 'ok', service: 'pg-rw', pods: ['pg-2'] } }) })
+ expect(d[0].tone).toBe('unhealthy')
+ })
+ it('serving is unassessed when instance Pods are not readable', () => {
+ const d = cnpgDimensions({ row: row({ pods: [] }) })
+ expect(d[0].text).toBe('unassessed')
+ })
+})
+
+describe('cnpgLeaseHolderPod', () => {
+ it('names the Pod of a controller-runtime holder identity', () => {
+ expect(cnpgLeaseHolderPod('cnpg-controller-manager-5fbdd6bb78-jx82z_6ec0566c-6da6-47aa-9ab1-0e5d3b2c1f11')).toBe('cnpg-controller-manager-5fbdd6bb78-jx82z')
+ expect(cnpgLeaseHolderPod('pg-1')).toBe('pg-1')
+ })
+})
+
+describe('folded HA and certificates summaries', () => {
+ const pdbs = { state: 'ok' as const, enabled: true, items: [{ name: 'pg', role: 'replicas', disruptionsAllowed: 1, currentHealthy: 1, expectedPods: 1, observed: true }] } as unknown as CNPGClusterHA['pdbs']
+
+ it('stays folded with the known facts when nothing is out of line', () => {
+ const live = [{ pod: 'pg-1', state: 'ok' }, { pod: 'pg-2', state: 'ok' }] as never
+ expect(cnpgHASummary(ha({ pdbs, operatorLease: { state: 'ok', holder: 'op_1' } }), live)).toEqual({ text: '2/2 observed instances ready · 2 zones · images match', attention: false })
+ })
+
+ it('claims neither readiness nor matching images without instances', () => {
+ expect(cnpgHASummary(ha({ pdbs, instances: [], operatorLease: { state: 'ok' } }), [] as never)).toEqual({ text: 'no instance Pods', attention: true })
+ const unset = ha({ pdbs, operatorLease: { state: 'ok' }, instances: [{ pod: 'pg-1', podUID: 'a', role: 'primary', ready: true, node: 'n1', zone: 'z1', restartCount: 0 }] })
+ expect(cnpgHASummary(unset, [{ pod: 'pg-1', state: 'ok' }] as never).text).toBe('1/1 observed instances ready')
+ })
+
+ it('names what it could not read instead of reading calm', () => {
+ const denied = { state: 'denied' as const, grant: { verb: 'list', resource: 'pods', namespace: 'db' } }
+ const unread = ha({ pods: denied, nodes: denied, pdbs: { ...denied, enabled: true, items: [] }, primaryLease: denied, operatorLease: denied, jobs: { ...denied, items: [] } } as never)
+ expect(cnpgHASummary(unread, undefined)).toEqual({
+ text: 'Not read: Pods, zones, disruption budgets, primary lease, operator lease, Jobs, pending restarts',
+ attention: false,
+ })
+ expect(cnpgHASummary(ha({ pdbs }), undefined).text).toBe('2/2 observed instances ready · 2 zones · images match · not read: operator lease, pending restarts')
+ })
+
+ it('opens and leads with what is wrong', () => {
+ const shared = ha({
+ pdbs,
+ instances: [
+ { pod: 'pg-1', podUID: 'a', role: 'primary', ready: true, node: 'n1', zone: 'z1', restartCount: 0, image: 'pg:17', imageMatches: true },
+ { pod: 'pg-2', podUID: 'b', role: 'replica', ready: false, node: 'n1', zone: 'z1', restartCount: 0, image: 'pg:16', imageMatches: false },
+ ],
+ })
+ const s = cnpgHASummary(shared, [{ pod: 'pg-2', state: 'ok', pendingRestart: true } as never])
+ expect(s.attention).toBe(true)
+ expect(s.text).toBe('1 of 2 instances not ready · every instance in one zone · instances share a Node · restart pending on pg-2 · an instance Pod has a different image · not read: operator lease')
+ })
+
+ it('names the nearest certificate expiry and who renews them', () => {
+ const now = Date.parse('2026-09-30T00:00:00Z')
+ const certs = [
+ { secret: 'pg-ca', expiresAt: '2026-12-29T00:00:00Z', renewal: 'operator' },
+ { secret: 'pg-server', expiresAt: '2026-10-20T00:00:00Z', renewal: 'operator' },
+ ] as never
+ expect(cnpgCertificatesSummary(certs, now)).toEqual({ text: '2 certificates · nearest reported expiry in 20 d (pg-server) · CloudNativePG renews them', attention: false })
+ const oneUnread = [{ secret: 'pg-ca', expiresAt: '2026-12-29T00:00:00Z', renewal: 'operator' }, { secret: 'pg-server', raw: 'garbage', renewal: 'operator' }] as never
+ expect(cnpgCertificatesSummary(oneUnread, now)).toEqual({ text: '2 certificates · nearest reported expiry in 90 d (pg-ca) · 1 expiry unreadable · CloudNativePG renews them', attention: true })
+ const userSoon = [{ secret: 'app-tls', expiresAt: '2026-10-05T00:00:00Z', renewal: 'user' }] as never
+ expect(cnpgCertificatesSummary(userSoon, now).attention).toBe(true)
+ expect(cnpgCertificatesSummary([], now)).toEqual({ text: 'No expiry reported by the operator', attention: false })
+ })
+})
+
+describe('replication chip and the sustained-lag finding', () => {
+ it('reads no calmer than a standby measured far behind for the whole window', () => {
+ const problem = { id: 'lag:db/pg', reason: 'CNPGSustainedLag', severity: 'critical', category: 'availability', title: 'pg-2 ≥ 24 h behind in every sample for 10 min', subject: { kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: 'db', name: 'pg' }, source: 'measurement' } as never
+ const dims = cnpgDimensions({ row: row({ problems: [problem] }), replication: { streaming: 1, standbys: 1, maxReplayLagSeconds: 0 } })
+ const rep = dims.find((d) => d.id === 'replication')!
+ expect(rep.tone).toBe('unhealthy')
+ expect(rep.text).toBe('1 of 1 standbys streaming · sustained lag')
+ expect(rep.source).toBe('pg-2 ≥ 24 h behind in every sample for 10 min')
+ const unread = cnpgDimensions({ row: row({ problems: [problem] }) }).find((d) => d.id === 'replication')!
+ expect(unread).toMatchObject({ tone: 'unhealthy', text: 'sustained lag' })
+ })
+})
+
+describe('backup dimension certainty', () => {
+ it('never treats archiving alone as restorable', () => {
+ const r = row()
+ r.protection.lastSuccessfulBackup = { text: 'No successful backup yet', tone: 'degraded', source: 'Backups read in this namespace; none completed' }
+ expect(cnpgDimensions({ row: r }).find((d) => d.id === 'protection')).toMatchObject({ text: 'CNPG reports archiving · no successful backup yet', tone: 'degraded', source: 'Backups read in this namespace; none completed' })
+ r.protection.lastSuccessfulBackup = { text: 'No access to Backups', tone: 'unknown', source: 'Backups not read' }
+ expect(cnpgDimensions({ row: r }).find((d) => d.id === 'protection')).toMatchObject({ text: 'unassessed', tone: 'unknown', source: 'Backups not read' })
+ })
+})
+
+it('uses declared instances for readiness and opens for an instance Pod not observed', () => {
+ const base = ha({ declaredInstances: 2, expectedInstances: ['pg-1', 'pg-2'] })
+ base.instances = base.instances.slice(0, 1)
+ const summary = cnpgHASummary(base, [{ pod: 'pg-1', state: 'ok' }])
+ expect(summary.attention).toBe(true)
+ expect(summary.text).toContain('1 of 2 declared instances ready; no instance Pod observed for pg-2')
+ expect(summary.text).not.toContain('1/1')
+})
+
+it('names excess observed instances without putting them over a smaller denominator', () => {
+ const summary = cnpgHASummary(ha({ declaredInstances: 1 }), undefined)
+ expect(summary.text).toContain('2 observed instances ready; 1 instances declared')
+ expect(summary.text).not.toContain('2 of 1')
+})
+it('does not claim matching images without observed image evidence', () => {
+ expect(cnpgImageDrift(ha({ instances: [] })).known).toBe(false)
+ const base = ha()
+ base.instances[0].imageMatches = undefined
+ expect(cnpgImageDrift(base).known).toBe(false)
+})
+
+it('uses a definitive non-serving verdict from an empty complete endpoint read without a primary', () => {
+ const r = row({ cluster: { status: {} }, pods: [] })
+ const empty = ha({ rwEndpoints: { state: 'ok', service: 'analytics-rw', pods: [] } })
+ expect(cnpgDimensions({ row: r, ha: empty })[0]).toMatchObject({ tone: 'unhealthy', text: 'not serving: no ready read-write endpoint', source: 'EndpointSlices of Service analytics-rw' })
+ expect(cnpgDimensions({ row: r, ha: empty })[0].tone).toBe('unhealthy')
+ expect(cnpgDimensions({ row: r, ha: ha({ rwEndpoints: { state: 'denied', service: 'analytics-rw', pods: [] } }) })[0].tone).toBe('unknown')
+ expect(cnpgDimensions({ row: { ...r, hibernated: true }, ha: empty })[0].text).toBe('hibernated')
+})
+it('uses storage reasons for loading, missing Prometheus and denied usage', () => {
+ expect(cnpgDimensions({ row: row() })[2].text).toBe('Reading…')
+ expect(cnpgDimensions({ row: row({ disk: { tone: 'unknown', text: 'No usage metrics', source: 'Prometheus not connected' } }) })[2].text).toBe('No usage metrics')
+ expect(cnpgDimensions({ row: row({ disk: { tone: 'unknown', text: 'No access', source: 'Needs list persistentvolumeclaims in namespace db' } }) })[2].text).toContain('Needs list persistentvolumeclaims')
+})
+
+it('retains the missing standby and failover consequence when streaming is denied', () => {
+ const h = ha({ expectedInstances: ['pg-1', 'pg-2'], instances: [ha().instances[0]] })
+ const d = cnpgDimensions({ row: row(), ha: h, replicationGap: 'needs get pods/proxy in namespace db' }).find((d) => d.id === 'replication')!
+ expect(d.tone).toBe('degraded')
+ expect(d.text).toBe('Expected standby pg-2 is not running')
+ expect(d.source).toContain('streaming not measured: needs get pods/proxy in namespace db')
+ expect(d.source).toContain('No ready standby to fail over to')
+ const unread = cnpgDimensions({ row: row(), ha: { ...h, pods: { state: 'denied' } } }).find((d) => d.id === 'replication')!
+ expect(unread.text).toBe('unassessed')
+ expect(unread.source).not.toContain('No ready standby')
+})
+
+it('keeps the Storage verdict concise while the Storage notice owns discovery details', () => {
+ const r = row({ disk: { tone: 'unknown', text: 'No usage metrics', source: 'Prometheus not connected', detail: 'No working endpoint. Candidate monitoring/prometheus.' } })
+ expect(cnpgDimensions({ row: r }).find((d) => d.id === 'storage')).toMatchObject({ text: 'No usage metrics', source: '' })
+})
+it('adds the failover consequence to measured replication without duplicating streaming facts', () => {
+ const h = ha({ instances: [ha().instances[0]], expectedInstances: ['pg-1', 'pg-2'] })
+ const d = cnpgDimensions({ row: row(), ha: h, replication: { streaming: 0, standbys: 0 } }).find((d) => d.id === 'replication')!
+ expect(d.text).toBe('0 of 1 expected standbys streaming')
+ expect(d.source).toContain('No ready standby to fail over to')
+ expect(d.tone).toBe('degraded')
+})
+
+it('preserves separately measured standby gaps beside missing Pods and avoids inventing a grant', () => {
+ const h = ha({ expectedInstances: ['pg-1', 'pg-2', 'pg-3'], instances: [ha().instances[0], { ...ha().instances[1], pod: 'pg-3' }] })
+ const r = row({ instances: { desired: 3, ready: 2 }, problems: [{ id: 'standby:db/pg:pg-3', reason: 'CNPGStandbyNotReceiving', severity: 'warning', category: 'replication', title: 'pg-3 receiver is down', subject: { kind: 'Pod', name: 'pg-3' }, source: 'measurement' } as any] })
+ const d = cnpgDimensions({ row: r, ha: h }).find((d) => d.id === 'replication')!
+ expect(d.text).toContain('Expected standby pg-2 is not running')
+ expect(d.text).toContain('pg-3 not receiving WAL')
+ expect(d.source).toContain('pg-3 receiver is down')
+ expect(d.source).toContain('streaming not measured: not read')
+ expect(d.source).not.toContain('needs get pods/proxy')
+ h.instances = [h.instances[0]]
+ expect(cnpgDimensions({ row: r, ha: h }).find((d) => d.id === 'replication')!.text).toContain('Expected standbys pg-2, pg-3 are not running')
+})
diff --git a/packages/k8s-ui/src/components/cnpg/ha.ts b/packages/k8s-ui/src/components/cnpg/ha.ts
new file mode 100644
index 0000000000..bf117736be
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/ha.ts
@@ -0,0 +1,602 @@
+// High-availability facts for one CloudNativePG Cluster, from
+// GET /api/cnpg/clusters/{ns}/{name}/ha, plus pure derivations the Overview and
+// the switchover dialog share. Each sub-read carries its own state: a denied or
+// unavailable source is "unknown", never none or healthy.
+
+import { formatAge, type HealthLevel } from '../resources/resource-utils'
+import { formatGrant, type Grant } from '../../utils/grant'
+import { CNPG_PROMETHEUS_NOT_CONNECTED, cnpgFormatLag, cnpgLagTone, cnpgReplicationTone, type CNPGFleetRow } from './workspace'
+import { type Fact } from '../facts'
+import { type FoldSummary } from '../ui/FoldSection'
+import { worseTone } from '../ui/status-tone'
+
+export type CNPGHASourceState = 'ok' | 'denied' | 'notFound' | 'notInstalled' | 'unavailable' | 'error'
+
+export interface CNPGHASource {
+ state: CNPGHASourceState
+ reason?: string
+ grant?: Grant
+}
+
+export interface CNPGHAInstance {
+ pod: string
+ podUID: string
+ role: 'primary' | 'replica' | 'unknown'
+ ready: boolean
+ node?: string
+ zone?: string
+ qosClass?: string
+ image?: string
+ imageMatches?: boolean
+ podCreatedAt?: string
+ postgresStartedAt?: string
+ restartCount: number
+}
+
+export interface CNPGHAQuorum {
+ enabled: boolean
+ enabledBy?: 'spec' | 'annotation'
+ method?: string
+ number?: number
+ dataDurability?: string
+ object: CNPGHASource
+ status?: { method?: string; standbyNames: string[]; standbyNumber: number; primary?: string }
+ n?: number
+ w?: number
+ r?: number
+ promotable?: string[]
+ holds?: boolean
+}
+
+export interface CNPGHAPDB {
+ name: string
+ role: 'primary' | 'replicas' | 'other'
+ minAvailable?: string
+ maxUnavailable?: string
+ expectedPods: number
+ currentHealthy: number
+ desiredHealthy: number
+ disruptionsAllowed: number
+ observed: boolean
+}
+
+export interface CNPGHALease extends CNPGHASource {
+ namespace?: string
+ name?: string
+ holder?: string
+ renewTime?: string
+ durationSeconds?: number
+ expired?: boolean
+ controlledByCluster?: boolean
+}
+
+export interface CNPGHAJob {
+ name: string
+ role?: string
+ instance?: string
+ phase: 'running' | 'active' | 'succeeded' | 'failed' | 'pending'
+ reason?: string
+ startTime?: string
+ completionTime?: string
+}
+
+export interface CNPGHACertificate {
+ secret: string
+ purposes?: string[]
+ raw: string
+ expiresAt?: string
+ renewal: 'operator' | 'user'
+ metadata?: CNPGHASource
+ certManager?: { certificate: string; issuer?: string; issuerKind?: string }
+}
+
+export interface CNPGMaintenanceFacts {
+ declared: boolean
+ inProgress: boolean
+ reusePVC: boolean
+}
+
+export interface CNPGClusterHA {
+ cluster: { namespace: string; name: string; uid: string }
+ sampledAt: string
+ desiredImage?: string
+ declaredInstances?: number
+ expectedInstances?: string[]
+ instances: CNPGHAInstance[]
+ pods: CNPGHASource
+ nodes: CNPGHASource
+ quorum: CNPGHAQuorum
+ pdbs: CNPGHASource & { enabled: boolean; items: CNPGHAPDB[] }
+ primaryLease: CNPGHALease
+ operatorLease: CNPGHALease
+ jobs: CNPGHASource & { items: CNPGHAJob[] }
+ rwEndpoints: CNPGHASource & { service: string; pods: string[] }
+ certificates: CNPGHACertificate[]
+ maintenance: CNPGMaintenanceFacts
+}
+
+/**
+ * Per-instance facts read live from each instance manager (/pg/status). Only
+ * callers who can read them pass them in: they never come from the cached
+ * Cluster object.
+ */
+export interface CNPGInstanceLive {
+ pod: string
+ /** ok, or why the instance's own report is missing. */
+ state: string
+ pendingRestart?: boolean
+ pendingRestartForDecrease?: boolean
+ /** The report did not finish its reads, so pendingRestart is not established. */
+ incomplete?: boolean
+ /** Why the report is missing or incomplete, in words. */
+ reason?: string
+ roleDetail?: 'primary' | 'pgRewind' | 'replayPaused' | 'streaming' | 'fileBased'
+ instanceManagerVersion?: string
+ timeline?: number
+}
+
+export const CNPG_ROLE_DETAIL_TEXT: Record, string> = {
+ primary: 'primary',
+ pgRewind: 'pg_rewind running',
+ replayPaused: 'replay paused',
+ streaming: 'streaming standby',
+ fileBased: 'file-based standby (no WAL receiver)',
+}
+
+export function cnpgHASourceText(src: CNPGHASource | undefined, what: string): string {
+ if (!src) return `${what}: unknown`
+ switch (src.state) {
+ case 'ok':
+ return ''
+ case 'denied':
+ return `No access to ${what}${src.grant ? ` (needs ${formatGrant(src.grant)})` : ''}`
+ case 'notInstalled':
+ return src.reason ?? `${what}: not available in this CloudNativePG version`
+ case 'notFound':
+ return src.reason ?? `No ${what}`
+ case 'unavailable':
+ return src.reason ?? `${what}: unavailable`
+ default:
+ return src.reason ? `${what} could not be read: ${src.reason}` : `${what} could not be read`
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Failure domains
+
+export interface CNPGZoneSpread {
+ /** false when Nodes (or Pods) could not be read: zones are unknown. */
+ known: boolean
+ zones: { zone: string; pods: string[] }[]
+ /** Instances whose Node carries no zone label. */
+ unlabelled: string[]
+ primaryZone?: string
+ /** true when every labelled instance shares one zone (and there is more than one instance). */
+ singleZone: boolean
+ nodes: { node: string; pods: string[] }[]
+ /** true when two or more instances share one Node. */
+ sharedNode: boolean
+}
+
+export function cnpgZoneSpread(ha: CNPGClusterHA | undefined): CNPGZoneSpread {
+ const empty: CNPGZoneSpread = { known: false, zones: [], unlabelled: [], singleZone: false, nodes: [], sharedNode: false }
+ if (!ha || ha.pods.state !== 'ok') return empty
+ const byNode = new Map()
+ for (const i of ha.instances) {
+ if (!i.node) continue
+ byNode.set(i.node, [...(byNode.get(i.node) ?? []), i.pod])
+ }
+ const nodes = [...byNode.entries()].map(([node, pods]) => ({ node, pods })).sort((a, b) => a.node.localeCompare(b.node))
+ const sharedNode = nodes.some((n) => n.pods.length > 1)
+ if (ha.nodes.state !== 'ok') return { ...empty, nodes, sharedNode }
+ const byZone = new Map()
+ const unlabelled: string[] = []
+ let primaryZone: string | undefined
+ for (const i of ha.instances) {
+ if (!i.zone) {
+ unlabelled.push(i.pod)
+ continue
+ }
+ byZone.set(i.zone, [...(byZone.get(i.zone) ?? []), i.pod])
+ if (i.role === 'primary') primaryZone = i.zone
+ }
+ const zones = [...byZone.entries()].map(([zone, pods]) => ({ zone, pods })).sort((a, b) => a.zone.localeCompare(b.zone))
+ return {
+ known: true,
+ zones,
+ unlabelled,
+ primaryZone,
+ singleZone: zones.length === 1 && ha.instances.length > 1 && unlabelled.length === 0,
+ nodes,
+ sharedNode,
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Quorum, PDB, images, certificates
+
+export function cnpgQuorumFact(q: CNPGHAQuorum | undefined): Fact {
+ if (!q) return { text: 'Unknown', tone: 'unknown' }
+ if (!q.enabled) {
+ if (q.number !== undefined || q.method) {
+ return { text: `Synchronous ${q.method ?? ''} ${q.number ?? ''}`.replace(/\s+/g, ' ').trim() + ' · quorum failover off', tone: 'neutral', source: 'Cluster spec.postgresql.synchronous' }
+ }
+ return { text: 'Off (asynchronous replication)', tone: 'neutral', source: 'Cluster spec' }
+ }
+ const unread = cnpgHASourceText(q.object, 'FailoverQuorum')
+ if (q.object.state !== 'ok') return { text: `Quorum failover on · ${unread}`, tone: 'unknown' }
+ if (q.n === undefined || q.w === undefined) {
+ return {
+ text: 'Quorum failover on · no synchronous configuration recorded: a failover would wait',
+ tone: 'degraded',
+ source: 'FailoverQuorum status (reset while PostgreSQL configuration changes)',
+ }
+ }
+ const base = `W ${q.w} of N ${q.n} potentially synchronous`
+ if (q.r === undefined || q.holds === undefined) {
+ return { text: `${base} · promotable replicas unknown`, tone: 'unknown', source: 'FailoverQuorum status; Pods not readable' }
+ }
+ return {
+ text: `${base} · R ${q.r} promotable · R + W ${q.holds ? '>' : '≤'} N`,
+ tone: q.holds ? 'healthy' : 'degraded',
+ source: 'FailoverQuorum status (the recorded configuration, not the operator’s decision); R from ready standby Pods',
+ }
+}
+
+export function cnpgPDBFact(pdbs: CNPGClusterHA['pdbs'] | undefined): Fact {
+ if (!pdbs) return { text: 'Unknown', tone: 'unknown' }
+ if (pdbs.state !== 'ok') return { text: cnpgHASourceText(pdbs, 'PodDisruptionBudgets'), tone: 'unknown' }
+ if (pdbs.items.length === 0) {
+ return pdbs.enabled
+ ? { text: 'None found although spec.enablePDB is on', tone: 'degraded', source: 'PodDisruptionBudgets owned by the Cluster' }
+ : { text: 'Disabled (spec.enablePDB: false): node drains are not held back', tone: 'neutral', source: 'Cluster spec' }
+ }
+ const parts = pdbs.items.map((p) => `${p.role === 'primary' ? 'primary' : p.role === 'replicas' ? 'standbys' : p.name}: ${p.disruptionsAllowed} disruption${p.disruptionsAllowed === 1 ? '' : 's'} allowed (${p.currentHealthy}/${p.expectedPods} healthy)`)
+ const stale = pdbs.items.some((p) => !p.observed)
+ return {
+ text: parts.join(' · '),
+ tone: stale ? 'unknown' : 'neutral',
+ source: stale ? 'PodDisruptionBudget status (not yet updated for the latest spec)' : 'PodDisruptionBudget status',
+ }
+}
+
+export function cnpgImageDrift(ha: CNPGClusterHA | undefined): { known: boolean; drifted: CNPGHAInstance[] } {
+ if (!ha || ha.pods.state !== 'ok' || !ha.desiredImage || ha.instances.length === 0) return { known: false, drifted: [] }
+ return { known: ha.instances.every((i) => i.imageMatches !== undefined), drifted: ha.instances.filter((i) => i.imageMatches === false) }
+}
+
+export interface CNPGCertificateView extends CNPGHACertificate {
+ /** Whole days until expiry; negative when expired; undefined when the expiry did not parse. */
+ daysLeft?: number
+ tone: HealthLevel
+}
+
+/**
+ * The same thresholds as the Issues engine: a certificate its owner renews is
+ * flagged from 30 days; the operator renews its own, so those only matter once
+ * renewal is overdue.
+ */
+export function cnpgCertificateViews(certs: CNPGHACertificate[] | undefined, now = Date.now()): CNPGCertificateView[] {
+ return (certs ?? []).map((c) => {
+ if (!c.expiresAt) return { ...c, tone: 'unknown' as HealthLevel }
+ const ms = Date.parse(c.expiresAt) - now
+ const daysLeft = Math.floor(ms / 86_400_000)
+ let tone: HealthLevel = 'healthy'
+ if (ms <= 0) tone = 'unhealthy'
+ else if (c.renewal === 'user' && ms < 7 * 86_400_000) tone = 'unhealthy'
+ else if (c.renewal === 'user' && ms < 30 * 86_400_000) tone = 'degraded'
+ else if (c.renewal === 'operator' && ms < 86_400_000) tone = 'unhealthy'
+ return { ...c, daysLeft, tone }
+ })
+}
+
+/**
+ * "HA and instances" in one line: what is wrong when something is, otherwise
+ * the readiness and placement facts that are known. Unknown facts never read
+ * as fine; they are left out of a calm summary rather than claimed.
+ */
+export function cnpgHASummary(
+ ha: CNPGClusterHA | undefined,
+ live: CNPGInstanceLive[] | undefined,
+ primaryConflict?: { status: string; labelled: string },
+): FoldSummary {
+ if (!ha) return { text: 'Not read', attention: false }
+ const issues: string[] = []
+ const calm: string[] = []
+ if (ha.pods.state === 'ok') {
+ const ready = ha.instances.filter((i) => i.ready).length
+ const missing = [...new Set(ha.expectedInstances ?? [])].filter((name) => !ha.instances.some((i) => i.pod === name))
+ if (ha.declaredInstances !== undefined && ready < ha.declaredInstances) {
+ issues.push(`${ready} of ${ha.declaredInstances} declared instances ready${missing.length ? `; no instance Pod observed for ${missing.join(', ')}` : ''}`)
+ } else if (ha.declaredInstances !== undefined && ready > ha.declaredInstances) {
+ issues.push(`${ready} observed instances ready; ${ha.declaredInstances} instances declared`)
+ } else if (ha.instances.length === 0) issues.push('no instance Pods')
+ else if (ready < ha.instances.length) issues.push(`${ha.instances.length - ready} of ${ha.instances.length} instances not ready`)
+ else calm.push(ha.declaredInstances !== undefined ? `${ready} of ${ha.declaredInstances} declared instances ready` : `${ready}/${ha.instances.length} observed instances ready`)
+ }
+ if (primaryConflict) issues.push('primary labels disagree')
+ const spread = cnpgZoneSpread(ha)
+ if (spread.singleZone) issues.push('every instance in one zone')
+ if (spread.sharedNode) issues.push('instances share a Node')
+ if (spread.known && !spread.singleZone && spread.zones.length > 1) calm.push(`${spread.zones.length} zones`)
+ const pending = cnpgPendingRestart(live)
+ if (pending.pods.length > 0) issues.push(`restart pending on ${pending.pods.join(', ')}`)
+ const drift = cnpgImageDrift(ha)
+ if (drift.drifted.length > 0) issues.push(`${drift.drifted.length === 1 ? 'an instance Pod has' : `${drift.drifted.length} instance Pods have`} a different image`)
+ else if (drift.known && ha.instances.length > 0 && ha.instances.every((i) => i.imageMatches === true)) calm.push('images match')
+ const quorum = cnpgQuorumFact(ha.quorum)
+ if (quorum.tone === 'degraded' || quorum.tone === 'unhealthy') issues.push('failover quorum does not hold')
+ const pdb = cnpgPDBFact(ha.pdbs)
+ if (pdb.tone === 'degraded' || pdb.tone === 'unhealthy') issues.push('disruption budgets missing')
+ for (const [lease, what] of [[ha.primaryLease, 'primary lease'], [ha.operatorLease, 'operator lease']] as const) {
+ if (lease.state === 'ok' && lease.expired) issues.push(`${what} expired`)
+ }
+ if (ha.primaryLease.state === 'ok' && ha.primaryLease.controlledByCluster === false) issues.push('primary lease not owned by this Cluster')
+ const failedJobs = ha.jobs.state === 'ok' ? ha.jobs.items.filter((j) => j.phase === 'failed').length : 0
+ if (failedJobs > 0) issues.push(`${failedJobs} failed ${failedJobs === 1 ? 'Job' : 'Jobs'}`)
+ // A fact that could not be read is named, never folded into a calm line.
+ const unread: string[] = []
+ const check = (src: CNPGHASource | undefined, label: string) => {
+ if (src && (src.state === 'denied' || src.state === 'unavailable' || src.state === 'error')) unread.push(label)
+ }
+ check(ha.pods, 'Pods')
+ check(ha.nodes, 'zones')
+ if (ha.quorum.enabled) check(ha.quorum.object, 'quorum')
+ check(ha.pdbs, 'disruption budgets')
+ check(ha.primaryLease, 'primary lease')
+ check(ha.operatorLease, 'operator lease')
+ check(ha.jobs, 'Jobs')
+ if (!pending.known && !(ha.pods.state === 'ok' && ha.instances.length === 0)) unread.push('pending restarts')
+ const notRead = unread.length > 0 ? `not read: ${unread.join(', ')}` : ''
+ if (issues.length > 0) return { text: [...issues, notRead].filter(Boolean).join(' · '), attention: true }
+ if (calm.length === 0) return { text: notRead ? notRead[0].toUpperCase() + notRead.slice(1) : 'Nothing reported out of line', attention: false }
+ return { text: [...calm, notRead].filter(Boolean).join(' · '), attention: false }
+}
+
+/** Certificates in one line: the nearest expiry and who renews them. */
+export function cnpgCertificatesSummary(certs: CNPGHACertificate[] | undefined, now = Date.now()): FoldSummary {
+ const views = cnpgCertificateViews(certs, now)
+ if (views.length === 0) return { text: 'No expiry reported by the operator', attention: false }
+ const dated = views.filter((c) => Number.isFinite(c.daysLeft)).sort((a, b) => (a.daysLeft ?? 0) - (b.daysLeft ?? 0))
+ const unreadable = views.length - dated.length
+ // An unreadable expiry could be past already, so it opens the section too.
+ const attention = unreadable > 0 || views.some((c) => c.tone === 'degraded' || c.tone === 'unhealthy')
+ const nearest = dated[0]
+ const when = [
+ !nearest ? '' : nearest.daysLeft! < 0 ? `${nearest.secret} expired` : `nearest reported expiry in ${nearest.daysLeft} d (${nearest.secret})`,
+ unreadable > 0 ? `${unreadable} expiry unreadable` : '',
+ ]
+ .filter(Boolean)
+ .join(' · ')
+ const renewers = new Set(views.map((c) => c.renewal))
+ const who = renewers.size > 1 ? 'some renewed by you' : renewers.has('user') ? 'you renew them' : 'CloudNativePG renews them'
+ return { text: `${views.length} ${views.length === 1 ? 'certificate' : 'certificates'} · ${when} · ${who}`, attention }
+}
+
+function cnpgLiveRead(l: CNPGInstanceLive): boolean {
+ return (l.state === 'ok' || l.state === 'partial') && !l.incomplete
+}
+
+export function cnpgPendingRestart(live: CNPGInstanceLive[] | undefined): { known: boolean; pods: string[]; forDecrease: boolean } {
+ const read = (live ?? []).filter(cnpgLiveRead)
+ if (read.length === 0) return { known: false, pods: [], forDecrease: false }
+ const pending = read.filter((l) => l.pendingRestart)
+ return { known: read.length === (live ?? []).length, pods: pending.map((l) => l.pod), forDecrease: pending.some((l) => l.pendingRestartForDecrease) }
+}
+
+/**
+ * Why instance-manager facts are not established: the host's reason when
+ * there is no runtime read at all (`unavailable`, e.g. the missing grant),
+ * otherwise each instance that did not report and why.
+ */
+export function cnpgLiveGap(live: CNPGInstanceLive[] | undefined, unavailable?: string): string {
+ if (!live) return unavailable ?? 'needs each instance manager’s status (get pods/proxy)'
+ if (live.length === 0) return 'no instance was read'
+ const unread = live.filter((l) => !cnpgLiveRead(l))
+ if (unread.length === 0) return ''
+ return unread
+ .map((l) => {
+ const why = l.reason ?? (l.state === 'denied' ? 'no access' : l.state)
+ return l.incomplete ? `${l.pod} reported incompletely (${why})` : `${l.pod} did not report (${why})`
+ })
+ .join('; ')
+}
+
+// ---------------------------------------------------------------------------
+// Header dimensions
+
+export type CNPGDimensionId = 'serving' | 'replication' | 'protection' | 'storage'
+
+export interface CNPGDimension {
+ id: CNPGDimensionId
+ label: string
+ /** unknown = unassessed: its source is not available. */
+ tone: HealthLevel
+ text: string
+ source: string
+}
+
+export interface CNPGReplicationLive {
+ /** Standbys the primary reports as streaming. */
+ streaming: number
+ /** Standby Pods the runtime read saw; the verdict compares against spec.instances − 1, not this. */
+ standbys: number
+ maxReplayLagSeconds?: number
+}
+
+export function cnpgDimensions({
+ row,
+ ha,
+ replication,
+ replicationGap,
+ storage,
+}: {
+ row: CNPGFleetRow
+ ha?: CNPGClusterHA
+ /** From the primary's pg_stat_replication; absent when runtime data is not readable. */
+ replication?: CNPGReplicationLive
+ /** Why `replication` is absent (e.g. "needs get pods/proxy in db"). */
+ replicationGap?: string
+ /** Supplied by the host once storage is assessed; unassessed otherwise. */
+ storage?: CNPGDimension
+}): CNPGDimension[] {
+ return [
+ servingDimension(row, ha),
+ replicationDimension(row, replication, replicationGap, ha),
+ storage ?? storageDimension(row),
+ protectionDimension(row),
+ ]
+}
+
+function storageDimension(row: CNPGFleetRow): CNPGDimension {
+ return withSlotRetention(row, volumeDimension(row))
+}
+
+function volumeDimension(row: CNPGFleetRow): CNPGDimension {
+ const base = { id: 'storage' as const, label: 'Storage' }
+ const disk = row.disk
+ if (!disk) return { ...base, tone: 'unknown', text: 'Reading…', source: 'Reading volume usage' }
+ if (disk.tone === 'unknown') return { ...base, tone: 'unknown', text: disk.source === CNPG_PROMETHEUS_NOT_CONNECTED ? disk.text : [disk.text, disk.source].filter(Boolean).join(': '), source: disk.source === CNPG_PROMETHEUS_NOT_CONNECTED ? '' : disk.detail ?? '' }
+ return { ...base, tone: disk.tone, text: disk.text, source: disk.source ?? 'Fullest volume' }
+}
+
+// WAL an inactive slot pins is a storage concern even while volume usage is
+// unmeasured; the chip says so instead of reading "unassessed".
+function withSlotRetention(row: CNPGFleetRow, dim: CNPGDimension): CNPGDimension {
+ const slot = row.problems.find((p) => p.reason === 'CNPGInactiveSlot')
+ if (!slot) return dim
+ const held = 'WAL held by an inactive slot'
+ return {
+ ...dim,
+ tone: worseTone(dim.tone, 'degraded'),
+ text: dim.tone === 'unknown' ? held : `${dim.text} · ${held}`,
+ source: dim.tone === 'unknown' ? `${slot.title}. Volume usage: ${[dim.text, dim.source].filter(Boolean).join(' · ')}` : `${slot.title}. ${dim.source ?? ''}`.trim(),
+ }
+}
+
+function servingDimension(row: CNPGFleetRow, ha?: CNPGClusterHA): CNPGDimension {
+ const base = { id: 'serving' as const, label: 'Serving' }
+ if (row.hibernated) return { ...base, tone: 'neutral', text: 'hibernated', source: 'cnpg.io/hibernation annotation' }
+ if (ha?.rwEndpoints.state === 'ok' && ha.rwEndpoints.pods.length === 0) return { ...base, tone: 'unhealthy', text: 'not serving: no ready read-write endpoint', source: `EndpointSlices of Service ${ha.rwEndpoints.service}` }
+ const primaryName = row.cluster?.status?.currentPrimary as string | undefined
+ const primary = row.pods.find((p) => p.name === primaryName)
+ if (!primaryName) return { ...base, tone: 'unknown', text: 'unassessed', source: 'No current primary reported' }
+ if (!primary || primary.ready === null) return { ...base, tone: 'unknown', text: 'unassessed', source: 'Instance Pods are not readable' }
+ if (!primary.ready) return { ...base, tone: 'unhealthy', text: 'primary not ready', source: `Pod ${primaryName} readiness` }
+ if (ha?.rwEndpoints.state === 'ok') {
+ if (!ha.rwEndpoints.pods.includes(primaryName)) {
+ return {
+ ...base,
+ tone: 'unhealthy',
+ text: ha.rwEndpoints.pods.length === 0 ? 'no read-write endpoint' : 'read-write endpoint not on the primary',
+ source: `EndpointSlices of Service ${ha.rwEndpoints.service}`,
+ }
+ }
+ return { ...base, tone: 'healthy', text: 'primary ready', source: `Pod ${primaryName} ready and behind Service ${ha.rwEndpoints.service}` }
+ }
+ return { ...base, tone: 'healthy', text: 'primary ready', source: `Pod ${primaryName} readiness (Service endpoints not readable)` }
+}
+
+// A standby measured far behind for the whole window outranks what the live
+// read shows: the chip must not read calmer than the finding below it.
+function replicationDimension(row: CNPGFleetRow, live?: CNPGReplicationLive, gap?: string, ha?: CNPGClusterHA): CNPGDimension {
+ let dim = withStandbyGaps(row, liveReplicationDimension(row, live, gap))
+ if (!row.hibernated && !row.replicaCluster && (row.instances.desired ?? 0) > 1) {
+ const primary = row.cluster?.status?.currentPrimary
+ const pods = ha?.pods.state === 'ok' ? ha.instances.map((i) => ({ name: i.pod, ready: i.ready })) : row.podReadiness ? row.pods : undefined
+ if (pods && primary) {
+ const expected = ha?.pods.state === 'ok' ? ha.expectedInstances ?? row.cluster?.status?.instanceNames ?? [] : row.cluster?.status?.instanceNames ?? []
+ const missing = expected.filter((name: string) => name !== primary && !pods.some((p) => p.name === name))
+ if (!live && missing.length > 0) {
+ const otherGaps = row.problems.filter((p) => p.reason === 'CNPGStandbyNotReceiving' && !missing.includes(p.subject.name))
+ const measuredGaps = withStandbyGaps({ ...row, problems: otherGaps }, liveReplicationDimension(row, live, gap))
+ dim = {
+ ...dim, tone: worseTone(dim.tone, 'degraded'),
+ text: `Expected standby${missing.length === 1 ? '' : 's'} ${missing.join(', ')} ${missing.length === 1 ? 'is' : 'are'} not running${otherGaps.length > 0 ? ` · ${measuredGaps.text}` : ''}`,
+ source: `${otherGaps.length > 0 ? `${measuredGaps.source}; ` : ''}Instance Pods; streaming not measured: ${gap ?? 'not read'}`,
+ }
+ }
+ if (pods.some((p) => p.name === primary && p.ready === true) && !pods.some((p) => p.name !== primary && p.ready === true)) dim = {
+ ...dim, tone: worseTone(dim.tone, 'degraded'), source: `${dim.source}. No ready standby to fail over to`,
+ }
+ }
+ }
+ const sustained = row.problems.find((p) => p.reason === 'CNPGSustainedLag')
+ if (!sustained) return dim
+ const tone: HealthLevel = sustained.severity === 'critical' ? 'unhealthy' : 'degraded'
+ return {
+ ...dim,
+ tone: worseTone(dim.tone, tone),
+ text: dim.tone === 'unknown' ? 'sustained lag' : `${dim.text} · sustained lag`,
+ source: sustained.title,
+ }
+}
+
+// A standby another source saw receiving nothing (Prometheus in the fleet, the
+// instance manager here) keeps the chip from reading calm or unassessed.
+function withStandbyGaps(row: CNPGFleetRow, dim: CNPGDimension): CNPGDimension {
+ const gaps = row.problems.filter((p) => p.reason === 'CNPGStandbyNotReceiving')
+ if (gaps.length === 0) return dim
+ const pods = gaps.map((p) => p.subject.name).join(', ')
+ const tone: HealthLevel = gaps.some((p) => p.severity === 'critical') ? 'unhealthy' : 'degraded'
+ return {
+ ...dim,
+ tone: worseTone(dim.tone, tone),
+ text: dim.tone === 'unknown' ? `${pods} not receiving WAL` : `${dim.text} · ${pods} not receiving WAL`,
+ source: gaps.map((p) => p.title).join('; '),
+ }
+}
+
+function liveReplicationDimension(row: CNPGFleetRow, live?: CNPGReplicationLive, gap?: string): CNPGDimension {
+ const base = { id: 'replication' as const, label: 'Replication' }
+ if (row.hibernated) return { ...base, tone: 'neutral', text: 'hibernated', source: 'cnpg.io/hibernation annotation' }
+ if (row.instances.desired === 1) return { ...base, tone: 'degraded', text: 'no standby', source: 'spec.instances is 1: there is no failover target' }
+ if (!live) return { ...base, tone: 'unknown', text: 'unassessed', source: `Needs the primary’s pg_stat_replication: ${gap ?? 'read through get pods/proxy'}` }
+ // Standbys whose Pods are gone are missing from the runtime read too, so
+ // the denominator is what the Cluster asks for, never what is running.
+ const desired = row.instances.desired
+ if (desired === null) {
+ return { ...base, tone: 'unknown', text: `${live.streaming} streaming`, source: 'spec.instances is not reported, so the expected standbys are unknown' }
+ }
+ const expected = Math.max(0, desired - 1)
+ const source = `Primary’s pg_stat_replication against spec.instances ${desired}`
+ const lag = live.maxReplayLagSeconds
+ const tone = cnpgReplicationTone(live.streaming, expected, lag)
+ // pg_stat_replication's replay_lag is the recent replay delay, not WAL still
+ // to replay; it can stay high after a standby caught up, so it is not "behind".
+ const lagText = lag !== undefined && cnpgLagTone(lag) !== 'healthy' ? ` · replay delay ${cnpgFormatLag(lag)}` : ''
+ if (live.streaming < expected) {
+ return { ...base, tone, text: `${live.streaming} of ${expected} expected standbys streaming${lagText}`, source }
+ }
+ if (lagText) return { ...base, tone, text: `${live.streaming} of ${expected} streaming${lagText}`, source }
+ return { ...base, tone: 'healthy', text: `${live.streaming} of ${expected} standbys streaming`, source }
+}
+
+function protectionDimension(row: CNPGFleetRow): CNPGDimension {
+ const base = { id: 'protection' as const, label: 'Backups' }
+ const p = row.protection
+ if (p.walArchiving.tone === 'unhealthy') return { ...base, tone: 'unhealthy', text: 'WAL archiving failing', source: 'ContinuousArchiving condition' }
+ const blockedSchedule = row.problems.find((problem) => problem.reason === 'CNPGScheduleDestinationMissing')
+ if (blockedSchedule) return { ...base, tone: 'degraded', text: blockedSchedule.title, source: blockedSchedule.origin?.detail ?? 'ScheduledBackup method against Cluster spec' }
+ if (p.destination.method === 'none') return { ...base, tone: 'degraded', text: 'no backup destination', source: 'Cluster spec' }
+ if (p.lastSuccessfulBackup.tone === 'unhealthy' || p.lastSuccessfulBackup.tone === 'degraded') {
+ const text = !p.lastSuccessfulBackup.at && p.walArchiving.tone === 'healthy'
+ ? `CNPG reports archiving · ${p.lastSuccessfulBackup.text.toLowerCase()}`
+ : p.lastSuccessfulBackup.text
+ return { ...base, tone: p.lastSuccessfulBackup.tone, text, source: p.lastSuccessfulBackup.source ?? 'Backups' }
+ }
+ if (p.lastSuccessfulBackup.tone === 'unknown') return { ...base, tone: 'unknown', text: 'unassessed', source: p.lastSuccessfulBackup.source ?? p.lastSuccessfulBackup.text }
+ if (p.walArchiving.tone === 'unknown') return { ...base, tone: 'unknown', text: 'unassessed', source: 'WAL archiving not reported' }
+ const last = p.lastSuccessfulBackup.at ? ` · last backup ${formatAge(p.lastSuccessfulBackup.at)} ago` : ''
+ if (p.walArchiving.tone === 'neutral') return { ...base, tone: 'healthy', text: `Backup completed${last}`, source: p.lastSuccessfulBackup.source ?? 'Backups' }
+ return { ...base, tone: 'healthy', text: `CNPG reports archiving${last}`, source: `ContinuousArchiving condition${p.lastSuccessfulBackup.source ? ` · last backup: ${p.lastSuccessfulBackup.source}` : ''}` }
+}
+
+/**
+ * The Pod a Lease holder names. controller-runtime's leader election records
+ * "_"; a CNPG primary Lease records the Pod name alone.
+ */
+export function cnpgLeaseHolderPod(holder: string): string {
+ const i = holder.indexOf('_')
+ return i > 0 ? holder.slice(0, i) : holder
+}
diff --git a/packages/k8s-ui/src/components/cnpg/index.ts b/packages/k8s-ui/src/components/cnpg/index.ts
index 08bd46e1fb..7a12978822 100644
--- a/packages/k8s-ui/src/components/cnpg/index.ts
+++ b/packages/k8s-ui/src/components/cnpg/index.ts
@@ -1,4 +1,7 @@
export * from './workspace'
+export * from './databaseRole'
+export * from './ha'
+export * from './CNPGClusterHASection'
export * from './primitives'
export * from './CNPGClusterSummary'
export * from './CNPGBackupSummary'
@@ -6,9 +9,24 @@ export * from './CNPGObjectStoreSummary'
export * from './CNPGDeclarativeSummary'
export * from './CNPGPoolerSummary'
export * from './CNPGImageCatalogSummary'
+export * from './pooler'
+export * from './connect'
+export * from './CNPGConnectSection'
+export * from './schedule'
+export * from './logicalReplication'
+export * from './CNPGLogicalPath'
export {
+ backupsForScheduledBackup,
+ cnpgScheduleDestinationBlocker,
+ objectStoreForBackup,
+ cnpgBackupMatchesCluster,
+ cnpgArchiveMatchesCluster,
+ type CNPGArchiveSource,
inferredObjectStoreHealth,
+ relationUnavailable,
usersOfObjectStore,
type CNPGObjectStoreHealth,
type CNPGObjectStoreUser,
} from './relations'
+
+export * from './backupRuns'
diff --git a/packages/k8s-ui/src/components/cnpg/logicalReplication.test.ts b/packages/k8s-ui/src/components/cnpg/logicalReplication.test.ts
new file mode 100644
index 0000000000..7debb9cf5a
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/logicalReplication.test.ts
@@ -0,0 +1,121 @@
+import { describe, expect, it } from 'vitest'
+import { cnpgLogicalLocation, cnpgLogicalPaths, cnpgLogicalSlotFact, cnpgResolvePublisher } from './logicalReplication'
+
+const G = 'postgresql.cnpg.io/v1'
+const cluster = (name: string, ns: string, spec: any = {}, status: any = {}) => ({ apiVersion: G, kind: 'Cluster', metadata: { name, namespace: ns }, spec: { instances: 3, ...spec }, status })
+const sub = (spec: any, status: any = {}) => ({ apiVersion: G, kind: 'Subscription', metadata: { name: 'orders-sub', namespace: 'dst' }, spec: { cluster: { name: 'dst' }, name: 'orders_sub', dbname: 'app', publicationName: 'orders_pub', externalClusterName: 'src', ...spec }, status })
+const pub = { apiVersion: G, kind: 'Publication', metadata: { name: 'orders-pub', namespace: 'src' }, spec: { cluster: { name: 'src' }, name: 'orders_pub', dbname: 'app', target: { allTables: true } }, status: { applied: true } }
+const subscriber = (host: string) => cluster('dst', 'dst', { externalClusters: [{ name: 'src', connectionParameters: { host, dbname: 'app', user: 'app' } }] })
+const synced = (major: number) => cluster('src', 'src', { replicationSlots: { highAvailability: { synchronizeLogicalDecoding: true } } }, { pgDataImageInfo: { majorVersion: major } })
+
+describe('cnpgResolvePublisher', () => {
+ const clusters = [cluster('src', 'src'), cluster('dst', 'dst')]
+ it('resolves -rw/-ro/-r Service hosts in any spelling to the visible Cluster', () => {
+ for (const host of ['src-rw.src.svc', 'src-rw.src.svc.cluster.local', 'src-ro.src', 'SRC-R.src.svc']) {
+ const p = cnpgResolvePublisher(host, 'dst', clusters, [])
+ expect(p.kind === 'cluster' && p.name).toBe('src')
+ }
+ })
+ it('resolves through a Pooler Service, and a bare name in the subscriber namespace', () => {
+ const p = cnpgResolvePublisher('src-pooler.src.svc', 'dst', clusters, [{ metadata: { name: 'src-pooler', namespace: 'src' }, spec: { cluster: { name: 'src' } } }])
+ expect(p.kind === 'cluster' && p.via).toBe('Pooler src-pooler')
+ expect(cnpgResolvePublisher('dst-rw', 'dst', clusters, []).kind).toBe('cluster')
+ })
+ it('never guesses a cluster from an external host', () => {
+ expect(cnpgResolvePublisher('src-rw.example.com', 'dst', clusters, []).kind).toBe('external')
+ expect(cnpgResolvePublisher('other-rw.src.svc', 'dst', clusters, []).kind).toBe('external')
+ expect(cnpgResolvePublisher(undefined, 'dst', clusters, []).kind).toBe('external')
+ })
+})
+
+describe('cnpgLogicalPaths', () => {
+ it('walks subscription → external cluster → publisher → publication object → slot', () => {
+ const [p] = cnpgLogicalPaths([sub({})], [synced(17), subscriber('src-rw.src.svc')], [pub], [])
+ expect(p.publisher.kind === 'cluster' && `${p.publisher.namespace}/${p.publisher.name}`).toBe('src/src')
+ expect(p.publication.object?.name).toBe('orders-pub')
+ expect(p.publication.dbname).toBe('app')
+ expect(p.slot.name).toBe('orders_sub')
+ })
+ it('says when the external cluster is not declared or the publication is not a CNPG object', () => {
+ const [missing] = cnpgLogicalPaths([sub({ externalClusterName: 'nope' })], [synced(17), subscriber('src-rw.src.svc')], [pub], [])
+ expect(missing.publisher.kind).toBe('external')
+ expect(missing.externalCluster.declared).toBe(false)
+ const [sqlOnly] = cnpgLogicalPaths([sub({ publicationName: 'made_in_sql' })], [synced(17), subscriber('src-rw.src.svc')], [pub], [])
+ expect(sqlOnly.publication.object).toBeUndefined()
+ })
+ it('says a Publication is unknown when the publisher namespace\'s Publications are unreadable', () => {
+ const [p] = cnpgLogicalPaths([sub({})], [synced(17), subscriber('src-rw.src.svc')], [], [], (ns) => (ns === 'src' ? 'No access to Publications' : null))
+ expect(p.publication.object).toBeUndefined()
+ expect(p.publication.unavailable).toBe('No access to Publications')
+ const [readable] = cnpgLogicalPaths([sub({})], [synced(17), subscriber('src-rw.src.svc')], [], [], () => null)
+ expect(readable.publication.unavailable).toBeUndefined()
+ })
+ it('names the slot from slot_name, and none for slot_name = NONE', () => {
+ expect(cnpgLogicalPaths([sub({ parameters: { slot_name: 'custom' } })], [], [], [])[0].slot.name).toBe('custom')
+ expect(cnpgLogicalPaths([sub({ parameters: { slot_name: 'NONE' } })], [], [], [])[0].slot.name).toBeUndefined()
+ })
+})
+
+describe('slot failover', () => {
+ const failoverOf = (publisher: any, params?: any) =>
+ cnpgLogicalPaths([sub(params ? { parameters: params } : {})], [publisher, subscriber('src-rw.src.svc')], [], [])[0].failover
+ it('is lost when the publisher does not synchronize logical slots', () => {
+ const f = failoverOf(cluster('src', 'src'))
+ expect(f.tone).toBe('degraded')
+ expect(f.text).toContain('synchronizeLogicalDecoding is off')
+ })
+ it('on PostgreSQL 17 needs the subscription to request failover', () => {
+ expect(failoverOf(synced(17)).tone).toBe('degraded')
+ expect(failoverOf(synced(17), { failover: 'true' }).tone).toBe('healthy')
+ })
+ it('is lost when HA slots are disabled, whatever synchronizeLogicalDecoding says', () => {
+ const off = cluster('src', 'src', { replicationSlots: { highAvailability: { enabled: false, synchronizeLogicalDecoding: true } } }, { pgDataImageInfo: { majorVersion: 17 } })
+ const f = failoverOf(off, { failover: 'true' })
+ expect(f.tone).toBe('degraded')
+ expect(f.text).toContain('HA replication slots are disabled')
+ })
+ it('is unknown before 17 (pg_failover_slots), without a major, or outside Radar', () => {
+ expect(failoverOf(synced(16)).tone).toBe('unknown')
+ expect(failoverOf(cluster('src', 'src', { replicationSlots: { highAvailability: { synchronizeLogicalDecoding: true } } })).tone).toBe('unknown')
+ expect(cnpgLogicalPaths([sub({})], [subscriber('db.example.com')], [], [])[0].failover.tone).toBe('unknown')
+ })
+})
+
+describe('cnpgLogicalSlotFact', () => {
+ const [path] = cnpgLogicalPaths([sub({}, { applied: true })], [synced(17), subscriber('src-rw.src.svc')], [pub], [])
+ it('never reads a denied or unread runtime as a missing slot', () => {
+ expect(cnpgLogicalSlotFact(path, { state: 'denied' }).tone).toBe('unknown')
+ expect(cnpgLogicalSlotFact(path, { state: 'unavailable', reason: 'unreachable' }).text).toContain('not read (unreachable)')
+ })
+ it('reports the observed slot, and a missing one only from a readable report', () => {
+ const ok = cnpgLogicalSlotFact(path, { state: 'ok', slots: [{ name: 'orders_sub', type: 'logical', active: true, retainedBytes: 2048, walStatus: 'reserved' }] })
+ expect(ok.text).toBe('Slot orders_sub · logical · active · retains 2.0 KiB of WAL · WAL reserved')
+ expect(ok.tone).toBe('healthy')
+ expect(cnpgLogicalSlotFact(path, { state: 'ok', slots: [{ name: 'orders_sub', active: false }] }).tone).toBe('degraded')
+ const capped = cnpgLogicalSlotFact(path, { state: 'partial', reason: '250 replication slots; the first 200 are shown', slots: [] })
+ expect(capped.text).toContain('not in the reported slots (report incomplete: 250 replication slots')
+ expect(capped.tone).toBe('unknown')
+ const stale = cnpgLogicalSlotFact(path, { state: 'ok', stale: true, slots: [{ name: 'orders_sub', type: 'logical', active: true }] })
+ expect(stale.tone).toBe('unknown')
+ expect(stale.text).toContain('from an earlier read')
+ const gone = cnpgLogicalSlotFact(path, { state: 'ok', slots: [] })
+ expect(gone.text).toContain('not found')
+ expect(gone.tone).toBe('degraded')
+ })
+})
+
+describe('cnpgLogicalLocation', () => {
+ it('says an unknown database in words', () => {
+ expect(cnpgLogicalLocation('upstream', 'app')).toBe('upstream/app')
+ expect(cnpgLogicalLocation('external cluster upstream', undefined)).toBe('external cluster upstream · database unknown')
+ expect(cnpgLogicalLocation('x', undefined)).not.toContain('?')
+ })
+})
+
+it('keeps stale Publication and Subscription results pending on the logical path', () => {
+ const staleSub = { ...sub({}, { applied: true, observedGeneration: 1 }), metadata: { ...sub({}).metadata, generation: 2 } }
+ const stalePub = { ...pub, metadata: { ...pub.metadata, generation: 2 }, status: { applied: true, observedGeneration: 1 } }
+ const [path] = cnpgLogicalPaths([staleSub], [synced(17), subscriber('src-rw.src.svc')], [stalePub], [])
+ expect(path.subscription.applied).toBeNull()
+ expect(path.publication.object?.applied).toBeNull()
+})
diff --git a/packages/k8s-ui/src/components/cnpg/logicalReplication.ts b/packages/k8s-ui/src/components/cnpg/logicalReplication.ts
new file mode 100644
index 0000000000..6e961cb451
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/logicalReplication.ts
@@ -0,0 +1,252 @@
+import { getCNPGPostgresMajor } from '../resources/resource-utils-cnpg'
+import { cnpgRoleState } from './databaseRole'
+import { cnpgFormatBytes } from './workspace'
+import { type Fact } from '../facts'
+
+/** Where a Subscription's publisher lives, as far as the subscriber's spec shows. */
+export type CNPGPublisher =
+ | { kind: 'cluster'; namespace: string; name: string; via: string; cluster: any }
+ | { kind: 'external'; host?: string; reason: string }
+
+export interface CNPGLogicalPath {
+ subscription: { namespace: string; name: string; sqlName?: string; dbname?: string; cluster?: string; applied: boolean | null; message?: string }
+ externalCluster: { name?: string; declared: boolean; host?: string; dbname?: string }
+ publisher: CNPGPublisher
+ publication: {
+ name?: string
+ dbname?: string
+ /** The Publication object on the publisher Cluster declaring it, when one is visible. */
+ object?: { namespace: string; name: string; applied: boolean | null }
+ /** Why Publication objects in the publisher's namespace could not be read; absence then proves nothing. */
+ unavailable?: string
+ }
+ /** The slot PostgreSQL creates for the subscription: `slot_name`, else the subscription's name. */
+ slot: { name?: string; reason?: string }
+ failover: Fact
+}
+
+const SERVICE_SUFFIXES = ['-rw', '-ro', '-r']
+
+function stripHost(host: string): { name: string; namespace?: string } {
+ const h = host.trim().toLowerCase().replace(/\.$/, '').replace(/\.svc\.cluster\.local$/, '').replace(/\.svc$/, '')
+ const [name, namespace] = h.split('.')
+ return { name, namespace }
+}
+
+/**
+ * Resolves an external cluster's host to a visible CNPG Cluster through its
+ * -rw / -ro / -r Service or a Pooler's Service. Anything else stays external:
+ * the host is never guessed to be a cluster it only resembles.
+ */
+export function cnpgResolvePublisher(host: string | undefined, subscriberNs: string, clusters: any[], poolers: any[]): CNPGPublisher {
+ if (!host) return { kind: 'external', reason: 'the external cluster names no host' }
+ const { name, namespace = subscriberNs } = stripHost(host)
+ if (host.includes('.') && !/\.svc(\.cluster\.local)?\.?$/i.test(host.trim()) && host.split('.').length > 2) {
+ return { kind: 'external', host, reason: 'the host is not a Kubernetes Service name' }
+ }
+ const inNs = (list: any[]) => list.filter((o) => o?.metadata?.namespace === namespace)
+ for (const suffix of SERVICE_SUFFIXES) {
+ if (!name.endsWith(suffix)) continue
+ const base = name.slice(0, -suffix.length)
+ const c = inNs(clusters).find((x) => x.metadata?.name === base)
+ if (c) return { kind: 'cluster', namespace, name: base, via: `${name}.${namespace}`, cluster: c }
+ }
+ const pooler = inNs(poolers).find((p) => p.metadata?.name === name)
+ if (pooler?.spec?.cluster?.name) {
+ const c = inNs(clusters).find((x) => x.metadata?.name === pooler.spec.cluster.name)
+ if (c) return { kind: 'cluster', namespace, name: c.metadata.name, via: `Pooler ${name}`, cluster: c }
+ }
+ return { kind: 'external', host, reason: 'no visible CloudNativePG Cluster serves this host' }
+}
+
+/**
+ * The namespace a Subscription's publisher host names, so a host can read that
+ * namespace too before resolving the publisher. Undefined when the subscriber
+ * or its external cluster is not visible, or the host names none.
+ */
+export function cnpgSubscriptionHostNamespace(subscription: any, clusters: any[]): string | undefined {
+ const ns = subscription?.metadata?.namespace
+ const subscriber = clusters.find((c) => c?.metadata?.namespace === ns && c?.metadata?.name === subscription?.spec?.cluster?.name)
+ const ext = (subscriber?.spec?.externalClusters ?? []).find((e: any) => e?.name === subscription?.spec?.externalClusterName)
+ const host = ext?.connectionParameters?.host
+ return host ? stripHost(host).namespace ?? ns : undefined
+}
+
+
+function truthy(v: unknown): boolean | undefined {
+ if (v === undefined || v === null) return undefined
+ const s = String(v).trim().toLowerCase()
+ if (['true', 'on', 'yes', '1'].includes(s)) return true
+ if (['false', 'off', 'no', '0'].includes(s)) return false
+ return undefined
+}
+
+const FAILOVER_SOURCE = "Publisher's spec.replicationSlots.highAvailability (enabled and synchronizeLogicalDecoding), its PostgreSQL major and the Subscription's failover parameter; /pg/status does not report a slot's failover flag"
+
+/**
+ * Whether the publisher's slot for this subscription is kept on its standbys,
+ * so a failover of the publisher does not lose it. Declared configuration
+ * only: CloudNativePG does not report which slots were actually synchronized.
+ */
+export function cnpgSlotFailover(publisher: CNPGPublisher, subscription: any): Fact {
+ if (publisher.kind !== 'cluster') {
+ return { text: 'Unknown: the publisher is not a CloudNativePG Cluster Radar can see', tone: 'unknown', source: FAILOVER_SOURCE }
+ }
+ const c = publisher.cluster
+ if ((c?.spec?.instances ?? 1) <= 1) {
+ return { text: 'No standby to fail over to: a single-instance publisher', tone: 'neutral', source: FAILOVER_SOURCE }
+ }
+ const ha = c?.spec?.replicationSlots?.highAvailability
+ // GetEnabled defaults to true; the operator enables sync only with both.
+ if (ha?.synchronizeLogicalDecoding === true && ha?.enabled === false) {
+ return {
+ text: 'Lost on failover: synchronizeLogicalDecoding is on, but HA replication slots are disabled, so standbys have no physical slot to synchronize through',
+ tone: 'degraded',
+ source: FAILOVER_SOURCE,
+ }
+ }
+ const sync = ha?.synchronizeLogicalDecoding === true
+ if (!sync) {
+ return {
+ text: 'Lost on failover: the publisher does not synchronize logical slots to standbys (synchronizeLogicalDecoding is off)',
+ tone: 'degraded',
+ source: FAILOVER_SOURCE,
+ }
+ }
+ const major = getCNPGPostgresMajor(c)
+ if (major === undefined) {
+ return { text: 'Unknown: synchronization is on, but the publisher\'s PostgreSQL major is not reported', tone: 'unknown', source: FAILOVER_SOURCE }
+ }
+ if (major < 17) {
+ return {
+ text: 'Synchronized only if pg_failover_slots is loaded on the publisher (PostgreSQL < 17); Radar cannot tell',
+ tone: 'unknown',
+ source: FAILOVER_SOURCE,
+ }
+ }
+ const failover = truthy(subscription?.spec?.parameters?.failover)
+ if (failover === true) {
+ return { text: 'Kept on standbys (declared): synchronization is on and the subscription requests failover', tone: 'healthy', source: FAILOVER_SOURCE }
+ }
+ return {
+ text: "Lost on failover: PostgreSQL 17 synchronizes only slots created with failover = true, and the subscription's parameters do not set it",
+ tone: 'degraded',
+ source: FAILOVER_SOURCE,
+ }
+}
+
+export function cnpgLogicalPaths(
+ subscriptions: any[],
+ clusters: any[],
+ publications: any[],
+ poolers: any[],
+ /** Why Publications in a namespace are not readable (coverage), or null when they are. */
+ publicationsUnavailable?: (namespace: string) => string | null,
+): CNPGLogicalPath[] {
+ return subscriptions.map((sub) => {
+ const ns: string = sub?.metadata?.namespace ?? ''
+ const subscriber = clusters.find((c) => c?.metadata?.namespace === ns && c?.metadata?.name === sub?.spec?.cluster?.name)
+ const extName: string | undefined = sub?.spec?.externalClusterName
+ const ext = (subscriber?.spec?.externalClusters ?? []).find((e: any) => e?.name === extName)
+ const params = ext?.connectionParameters ?? {}
+ const publisher: CNPGPublisher = !subscriber
+ ? { kind: 'external', reason: 'the subscriber Cluster is not visible, so its external cluster cannot be read' }
+ : !ext
+ ? { kind: 'external', reason: `external cluster ${extName ?? '(unset)'} is not declared on ${subscriber.metadata?.name}` }
+ : cnpgResolvePublisher(params.host, ns, clusters, poolers)
+ const pubDb: string | undefined = sub?.spec?.publicationDBName || params.dbname
+ const pubName: string | undefined = sub?.spec?.publicationName
+ const pubObj =
+ publisher.kind === 'cluster'
+ ? publications.find(
+ (p) =>
+ p?.metadata?.namespace === publisher.namespace &&
+ p?.spec?.cluster?.name === publisher.name &&
+ p?.spec?.name === pubName &&
+ (!pubDb || p?.spec?.dbname === pubDb),
+ )
+ : undefined
+ const slotParam: string | undefined = sub?.spec?.parameters?.slot_name
+ const createSlot = truthy(sub?.spec?.parameters?.create_slot)
+ const slot =
+ slotParam && slotParam.toLowerCase() === 'none'
+ ? { reason: 'slot_name = NONE: the subscription uses no slot' }
+ : slotParam
+ ? { name: slotParam }
+ : createSlot === false
+ ? { name: sub?.spec?.name, reason: 'create_slot = false: the slot must be created by hand' }
+ : { name: sub?.spec?.name }
+ return {
+ subscription: {
+ namespace: ns,
+ name: sub?.metadata?.name,
+ sqlName: sub?.spec?.name,
+ dbname: sub?.spec?.dbname,
+ cluster: sub?.spec?.cluster?.name,
+ applied: cnpgRoleState(sub) === 'pending' ? null : cnpgRoleState(sub) === 'applied',
+ message: sub?.status?.message || undefined,
+ },
+ externalCluster: { name: extName, declared: !!ext, host: params.host, dbname: params.dbname },
+ publisher,
+ publication: {
+ name: pubName,
+ dbname: pubDb,
+ object: pubObj
+ ? { namespace: pubObj.metadata.namespace, name: pubObj.metadata.name, applied: cnpgRoleState(pubObj) === 'pending' ? null : cnpgRoleState(pubObj) === 'applied' }
+ : undefined,
+ unavailable: !pubObj && publisher.kind === 'cluster' ? publicationsUnavailable?.(publisher.namespace) ?? undefined : undefined,
+ },
+ slot,
+ failover: cnpgSlotFailover(publisher, sub),
+ }
+ })
+}
+
+/** What the publisher primary's instance manager says about one slot, in a shape independent of the runtime API. */
+export interface CNPGPublisherSlots {
+ /** partial: the report was capped or incomplete, so a slot missing from it may exist. */
+ state: 'ok' | 'partial' | 'denied' | 'unavailable' | 'notRead'
+ reason?: string
+ /** The latest refresh failed; slots are from an earlier read. */
+ stale?: boolean
+ slots?: { name: string; type?: string; active?: boolean; walStatus?: string; retainedBytes?: number; database?: string }[]
+}
+
+const SLOT_SOURCE = "Publisher primary's instance manager (/pg/status replicationSlotsInfo) and exporter (retained WAL)"
+
+export function cnpgLogicalSlotFact(path: CNPGLogicalPath, observed: CNPGPublisherSlots): Fact {
+ if (!path.slot.name) return { text: path.slot.reason ?? 'No slot', tone: 'neutral' }
+ if (path.publisher.kind !== 'cluster') return { text: `Slot ${path.slot.name}: not observable (publisher outside this cluster's view)`, tone: 'unknown' }
+ if (observed.state === 'denied') return { text: `Slot ${path.slot.name}: no access (needs get pods/proxy on the publisher)`, tone: 'unknown', source: SLOT_SOURCE }
+ if ((observed.state !== 'ok' && observed.state !== 'partial') || !observed.slots) {
+ return { text: `Slot ${path.slot.name}: not read${observed.reason ? ` (${observed.reason})` : ''}`, tone: 'unknown', source: SLOT_SOURCE }
+ }
+ const s = observed.slots.find((x) => x.name === path.slot.name)
+ if (!s && (observed.state === 'partial' || observed.stale)) {
+ return {
+ text: `Slot ${path.slot.name}: not in the reported slots (${observed.stale ? 'from an earlier read; the latest refresh failed' : `report incomplete${observed.reason ? `: ${observed.reason}` : ''}`})`,
+ tone: 'unknown',
+ source: SLOT_SOURCE,
+ }
+ }
+ if (!s) {
+ return {
+ text: `Slot ${path.slot.name} not found on the publisher primary${path.slot.reason ? `: ${path.slot.reason}` : ''}`,
+ tone: path.subscription.applied === true ? 'degraded' : 'unknown',
+ source: SLOT_SOURCE,
+ }
+ }
+ const parts = [`Slot ${s.name}`, s.type ?? 'type unknown', s.active === undefined ? 'activity unknown' : s.active ? 'active' : 'inactive']
+ if (s.retainedBytes !== undefined) parts.push(`retains ${cnpgFormatBytes(s.retainedBytes)} of WAL`)
+ if (s.walStatus) parts.push(`WAL ${s.walStatus}`)
+ const bad = s.active === false || s.walStatus === 'lost' || s.walStatus === 'unreserved'
+ if (observed.stale) {
+ return { text: `${parts.join(' · ')} (from an earlier read; the latest refresh failed)`, tone: 'unknown', source: SLOT_SOURCE }
+ }
+ return { text: parts.join(' · '), tone: s.walStatus === 'lost' ? 'unhealthy' : bad ? 'degraded' : 'healthy', source: SLOT_SOURCE }
+}
+
+/** "cluster/database", with an unknown database said in words rather than as "?". */
+export function cnpgLogicalLocation(where: string, dbname: string | undefined): string {
+ return dbname ? `${where}/${dbname}` : `${where} · database unknown`
+}
diff --git a/packages/k8s-ui/src/components/cnpg/pooler.test.ts b/packages/k8s-ui/src/components/cnpg/pooler.test.ts
new file mode 100644
index 0000000000..9ca8abf8c5
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/pooler.test.ts
@@ -0,0 +1,84 @@
+import { describe, expect, it } from 'vitest'
+import { poolerPressureCoverage, poolerPressureFact, aggregatePoolerPools, observedPause, poolerBackendService, poolerReadiness, poolerPodPressure } from './pooler'
+
+describe('poolerReadiness', () => {
+ it('reads readiness from the Deployment, never from the scheduled count', () => {
+ expect(poolerReadiness({ name: 'p', state: 'ok', replicas: 2, readyReplicas: 2 })).toMatchObject({ text: '2/2 ready', level: 'healthy' })
+ expect(poolerReadiness({ name: 'p', state: 'ok', replicas: 2, readyReplicas: 1 }).level).toBe('degraded')
+ expect(poolerReadiness({ name: 'p', state: 'ok', replicas: 2, readyReplicas: 0 }).level).toBe('unhealthy')
+ })
+ it('says unknown, not failed, when the Deployment is unreadable', () => {
+ expect(poolerReadiness({ name: 'p', state: 'unreadable' })).toMatchObject({ text: 'Unknown', level: 'unknown' })
+ expect(poolerReadiness(undefined).level).toBe('unknown')
+ })
+})
+
+describe('aggregatePoolerPools', () => {
+ it('sums each pool over the Pods that reported it and keeps unreported fields undefined', () => {
+ const rows = aggregatePoolerPools([
+ { pod: 'a', state: 'ok', pools: [{ database: 'app', user: 'app', clActive: 3, clWaiting: 1, maxwaitSeconds: 0.5, poolMode: 'transaction' }] },
+ { pod: 'b', state: 'ok', pools: [{ database: 'app', user: 'app', clActive: 2, clWaiting: 0, maxwaitSeconds: 2, poolMode: 'transaction' }] },
+ ])
+ expect(rows).toHaveLength(1)
+ expect(rows[0]).toMatchObject({ clActive: 5, clWaiting: 1, maxwaitSeconds: 2, pods: 2, poolModes: ['transaction'] })
+ expect(rows[0].svActive).toBeUndefined()
+ })
+})
+
+describe('observedPause', () => {
+ it('distinguishes all, some and unreadable', () => {
+ expect(observedPause({ state: 'ok', pods: [{ pod: 'a', state: 'ok', paused: true }, { pod: 'b', state: 'ok', paused: true }] })?.text).toBe('Paused on 2 of 2 PgBouncers')
+ expect(observedPause({ state: 'ok', pods: [{ pod: 'a', state: 'ok', paused: true }, { pod: 'b', state: 'ok', paused: false }] })?.level).toBe('alert')
+ expect(observedPause({ state: 'ok', pods: [{ pod: 'a', state: 'ok', paused: false }, { pod: 'b', state: 'error' }] })?.text).toBe('Serving (not paused) on 1 of 2 PgBouncers · 1 not read: b (error)')
+ expect(observedPause({ state: 'denied', grant: { verb: 'create', resource: 'pods', subresource: 'exec', namespace: 'x' }, pods: [] })?.text).toContain('create pods/exec')
+ })
+})
+
+describe('poolerBackendService', () => {
+ it('names the Cluster Service by type', () => {
+ expect(poolerBackendService('pg', 'rw')).toBe('pg-rw')
+ expect(poolerBackendService('pg', 'ro')).toBe('pg-ro')
+ expect(poolerBackendService('pg', undefined)).toBeUndefined()
+ })
+})
+
+describe('poolerPodPressure', () => {
+ it('totals each Pod on its own so one queuing Pod is visible', () => {
+ const rows = poolerPodPressure([
+ { pod: 'pooler-b', state: 'ok', pools: [{ database: 'app', user: 'app', clActive: 2, clWaiting: 0, svActive: 1 }, { database: 'app', user: 'ro', clActive: 1, clWaiting: 0, svActive: 0 }] },
+ { pod: 'pooler-a', state: 'ok', pools: [{ database: 'app', user: 'app', clActive: 20, clWaiting: 15, svActive: 20, maxwaitSeconds: 4.2 }] },
+ { pod: 'pooler-c', state: 'unreachable', error: 'no answer within 5s' },
+ ] as never)
+ expect(rows).toEqual([
+ { pod: 'pooler-a', state: 'ok', error: undefined, clActive: 20, clWaiting: 15, svActive: 20, maxwaitSeconds: 4.2 },
+ { pod: 'pooler-b', state: 'ok', error: undefined, clActive: 3, clWaiting: 0, svActive: 1 },
+ { pod: 'pooler-c', state: 'unreachable', error: 'no answer within 5s' },
+ ])
+ })
+})
+
+describe('pooler pressure certainty', () => {
+ it('says idle only when every Pod answered in full', () => {
+ const a = { pod: 'a', state: 'ok', pools: [] }
+ expect(poolerPressureCoverage([a]).empty).toBe('Idle: no client pools open')
+ for (const state of ['partial', 'unreachable']) {
+ const c = poolerPressureCoverage([a, { pod: 'b', state, reason: 'pool list capped' }])
+ expect(c.complete).toBe(false)
+ expect(c.empty).toBe('No pools seen in what was read')
+ expect(c.limitation).toContain('b:')
+ expect(c.limitation).toContain('pool list capped')
+ }
+ })
+ it('shows unknown for partial zero and a lower bound for partial positive counts', () => {
+ const p = { pod: 'a', state: 'partial', pools: [{ database: 'app', user: 'app', clWaiting: 0, clActive: 3 }] }
+ expect(poolerPressureFact([p], 'clWaiting')).toMatchObject({ text: 'Unknown', tone: 'unknown' })
+ expect(poolerPressureFact([p], 'clActive')).toMatchObject({ text: '≥3', tone: 'unknown' })
+ expect(poolerPressureFact([{ ...p, state: 'ok' }], 'clWaiting').text).toBe('0')
+ expect(poolerPressureFact([{ ...p, state: 'ok' }, { pod: 'b', state: 'unreachable' }], 'clWaiting').text).toBe('Unknown')
+ })
+ it('qualifies a total when a pool omitted the field, even with complete Pod reads', () => {
+ const pods = [{ pod: 'a', state: 'ok', pools: [{ database: 'app', user: 'a', clActive: 2 }, { database: 'app', user: 'b' }] }]
+ expect(poolerPressureFact(pods, 'clActive').text).toBe('≥2')
+ expect(poolerPressureFact(pods, 'clWaiting').text).toBe('Unknown')
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/pooler.ts b/packages/k8s-ui/src/components/cnpg/pooler.ts
new file mode 100644
index 0000000000..ee4041cbc4
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/pooler.ts
@@ -0,0 +1,225 @@
+import type { Fact } from '../facts'
+import type { HealthLevel } from '../resources/resource-utils'
+import { formatGrant, type Grant } from '../../utils/grant'
+
+/** The Deployment a Pooler runs, as the host read it. */
+export interface CNPGPoolerDeploymentLive {
+ name: string
+ state: 'ok' | 'missing' | 'unreadable' | 'foreign'
+ replicas?: number
+ readyReplicas?: number
+ updatedReplicas?: number
+ availableReplicas?: number
+}
+
+export interface CNPGPoolerServiceLive {
+ name: string
+ state: 'ok' | 'missing' | 'unreadable' | 'foreign'
+ type?: string
+ port?: number
+}
+
+export interface CNPGPoolerPoolSample {
+ database: string
+ user: string
+ clActive?: number
+ clWaiting?: number
+ svActive?: number
+ svIdle?: number
+ svUsed?: number
+ maxwaitSeconds?: number
+ poolMode?: string
+}
+
+/** Live PgBouncer exporter reads, one per pooler Pod. */
+export interface CNPGPoolerPressureLive {
+ state: 'loading' | 'denied' | 'error' | 'ok'
+ reason?: string
+ pods: { pod: string; state: string; error?: string; reason?: string; schedulingReason?: string; pools?: CNPGPoolerPoolSample[] }[]
+}
+
+/** Each PgBouncer's own SHOW STATE. */
+export interface CNPGPoolerObservedLive {
+ state: 'loading' | 'denied' | 'error' | 'ok'
+ grant?: Grant
+ reason?: string
+ pods: { pod: string; state: string; paused?: boolean; error?: string }[]
+}
+
+/**
+ * What a host that reads the cluster adds to the Pooler summary. Each part is
+ * optional: a part the host does not provide keeps the resource-only wording.
+ */
+export interface CNPGPoolerLive {
+ deployment?: CNPGPoolerDeploymentLive
+ service?: CNPGPoolerServiceLive
+ pressure?: CNPGPoolerPressureLive
+ observed?: CNPGPoolerObservedLive
+}
+
+export interface CNPGPoolerReadiness {
+ text: string
+ level: HealthLevel
+ detail?: string
+}
+
+/** Readiness from the Pooler's Deployment; the Pooler's status only counts scheduled Pods. */
+export function poolerReadiness(d: CNPGPoolerDeploymentLive | undefined): CNPGPoolerReadiness {
+ if (!d) return { text: 'Unknown', level: 'unknown', detail: 'Readiness is on the Pooler’s Deployment, which was not read' }
+ switch (d.state) {
+ case 'missing':
+ return { text: 'No Deployment', level: 'unhealthy', detail: `Deployment ${d.name} does not exist` }
+ case 'unreadable':
+ return { text: 'Unknown', level: 'unknown', detail: `No access to Deployment ${d.name}` }
+ case 'foreign':
+ return { text: 'Unknown', level: 'unknown', detail: `Deployment ${d.name} is not controlled by this Pooler` }
+ }
+ const want = d.replicas ?? 0
+ const ready = d.readyReplicas ?? 0
+ const detail = `from Deployment ${d.name}`
+ if (want === 0) return { text: 'Scaled to zero', level: 'neutral', detail }
+ if (ready === 0) return { text: `0/${want} ready`, level: 'unhealthy', detail }
+ if (ready < want) return { text: `${ready}/${want} ready`, level: 'degraded', detail }
+ return { text: `${ready}/${want} ready`, level: 'healthy', detail }
+}
+
+export interface CNPGPoolerPoolRow {
+ database: string
+ user: string
+ poolModes: string[]
+ clActive?: number
+ clWaiting?: number
+ svActive?: number
+ svIdle?: number
+ svUsed?: number
+ maxwaitSeconds?: number
+ /** Pods whose exporter reported this pool. */
+ pods: number
+}
+
+/**
+ * One row per database/user pool, summed over the Pods that reported it
+ * (max wait is the maximum). A field no Pod reported stays undefined.
+ */
+export function aggregatePoolerPools(pods: CNPGPoolerPressureLive['pods']): CNPGPoolerPoolRow[] {
+ const rows = new Map()
+ const add = (a: number | undefined, b: number | undefined) => (b === undefined ? a : (a ?? 0) + b)
+ for (const p of pods) {
+ for (const pool of p.pools ?? []) {
+ const key = `${pool.database}\u0000${pool.user}`
+ const row = rows.get(key) ?? { database: pool.database, user: pool.user, poolModes: [], pods: 0 }
+ row.pods++
+ row.clActive = add(row.clActive, pool.clActive)
+ row.clWaiting = add(row.clWaiting, pool.clWaiting)
+ row.svActive = add(row.svActive, pool.svActive)
+ row.svIdle = add(row.svIdle, pool.svIdle)
+ row.svUsed = add(row.svUsed, pool.svUsed)
+ if (pool.maxwaitSeconds !== undefined) row.maxwaitSeconds = Math.max(row.maxwaitSeconds ?? 0, pool.maxwaitSeconds)
+ if (pool.poolMode && !row.poolModes.includes(pool.poolMode)) row.poolModes.push(pool.poolMode)
+ rows.set(key, row)
+ }
+ }
+ return [...rows.values()].sort((a, b) => a.database.localeCompare(b.database) || a.user.localeCompare(b.user))
+}
+
+export interface CNPGPoolerPodRow {
+ pod: string
+ /** ok | partial | the read's failure state. */
+ state: string
+ error?: string
+ clActive?: number
+ clWaiting?: number
+ svActive?: number
+ maxwaitSeconds?: number
+}
+
+/**
+ * Each PgBouncer Pod's own totals across its pools, so one saturated Pod is
+ * not hidden in the sum. A Pod that did not report keeps its state and no
+ * numbers; a field none of its pools reported stays undefined.
+ */
+export function poolerPodPressure(pods: CNPGPoolerPressureLive['pods']): CNPGPoolerPodRow[] {
+ const add = (a: number | undefined, b: number | undefined) => (b === undefined ? a : (a ?? 0) + b)
+ return pods
+ .map((p) => {
+ const row: CNPGPoolerPodRow = { pod: p.pod, state: p.state, error: p.error }
+ for (const pool of p.pools ?? []) {
+ row.clActive = add(row.clActive, pool.clActive)
+ row.clWaiting = add(row.clWaiting, pool.clWaiting)
+ row.svActive = add(row.svActive, pool.svActive)
+ if (pool.maxwaitSeconds !== undefined) row.maxwaitSeconds = Math.max(row.maxwaitSeconds ?? 0, pool.maxwaitSeconds)
+ }
+ return row
+ })
+ .sort((a, b) => a.pod.localeCompare(b.pod))
+}
+
+/**
+ * PgBouncer settings worth showing, with the value PgBouncer uses when the
+ * Pooler leaves one unset. Defaults are named only where they were read from
+ * PgBouncer itself (SHOW CONFIG on PgBouncer 1.24); CloudNativePG writes only
+ * the parameters the Pooler sets.
+ */
+export const POOLER_LIMIT_PARAMETERS: { key: string; label: string; pgbouncerDefault?: string }[] = [
+ { key: 'default_pool_size', label: 'Pool size (per database/user)', pgbouncerDefault: '20' },
+ { key: 'max_client_conn', label: 'Max client connections', pgbouncerDefault: '100' },
+ { key: 'max_db_connections', label: 'Max connections per database', pgbouncerDefault: '0 (unlimited)' },
+ { key: 'max_user_connections', label: 'Max connections per user', pgbouncerDefault: '0 (unlimited)' },
+ { key: 'reserve_pool_size', label: 'Reserve pool', pgbouncerDefault: '0' },
+ { key: 'min_pool_size', label: 'Min pool size', pgbouncerDefault: '0' },
+]
+
+/** The Cluster Service PgBouncer forwards to, by the Pooler's type. */
+export function poolerBackendService(cluster: string | undefined, type: string | undefined): string | undefined {
+ if (!cluster || !type) return undefined
+ if (type === 'rw' || type === 'ro' || type === 'r') return `${cluster}-${type}`
+ return undefined
+}
+
+/** Observed pause across PgBouncers: every one, none, some, or unknown. */
+export function observedPause(o: CNPGPoolerObservedLive | undefined): { text: string; level: HealthLevel } | null {
+ if (!o) return null
+ if (o.state === 'loading') return { text: 'Reading…', level: 'unknown' }
+ if (o.state === 'denied') return { text: `Not observable: needs ${formatGrant(o.grant) ?? 'create pods/exec'}`, level: 'unknown' }
+ if (o.state === 'error') return { text: `Not observable: ${o.reason ?? 'read failed'}`, level: 'unknown' }
+ if (o.pods.length === 0) return { text: 'No PgBouncer Pods', level: 'unknown' }
+ const read = o.pods.filter((p) => p.state === 'ok' && p.paused !== undefined)
+ const paused = read.filter((p) => p.paused).length
+ const unread = o.pods.length - read.length
+ const tail = unread > 0 ? ` · ${unread} not read: ${o.pods.filter((p) => p.state !== 'ok' || p.paused === undefined).map((p) => `${p.pod} (${p.error ?? p.state})`).join('; ')}` : ''
+ if (read.length === 0) return { text: `Not observable${tail}`, level: 'unknown' }
+ if (paused === 0) return { text: `Serving (not paused) on ${read.length} of ${o.pods.length} PgBouncers${tail}`, level: unread ? 'unknown' : 'healthy' }
+ if (paused === read.length) return { text: `Paused on ${paused} of ${o.pods.length} PgBouncers${tail}`, level: 'degraded' }
+ return { text: `Paused on ${paused} of ${o.pods.length} PgBouncers${tail}`, level: 'alert' }
+}
+
+export function poolerPressureCoverage(pods: CNPGPoolerPressureLive['pods']) {
+ const reporting = pods.filter((p) => p.state === 'ok' || p.state === 'partial')
+ const complete = pods.length > 0 && pods.every((p) => p.state === 'ok')
+ const gaps = pods.filter((p) => p.state !== 'ok').map((p) => {
+ const reason = [p.reason, p.error].filter(Boolean).join(' · ')
+ return `${p.pod}: ${p.state === 'partial' ? 'partial' : 'not read'}${reason ? ` (${reason})` : p.state === 'partial' ? '' : ` (${p.state})`}`
+ })
+ return {
+ reporting,
+ complete,
+ limitation: gaps.join('; ') || (pods.length === 0 ? 'No PgBouncer Pods answered' : undefined),
+ empty: complete ? 'Idle: no client pools open' : 'No pools seen in what was read',
+ }
+}
+
+type PoolerMetric = 'clActive' | 'clWaiting' | 'svActive' | 'svIdle' | 'svUsed' | 'maxwaitSeconds'
+
+export function poolerPressureFact(pods: CNPGPoolerPressureLive['pods'], field: PoolerMetric, pool?: { database: string; user: string }): Fact {
+ const coverage = poolerPressureCoverage(pods)
+ const pools = coverage.reporting.flatMap((p) => p.pools ?? []).filter((p) => !pool || (p.database === pool.database && p.user === pool.user))
+ const values = pools.map((p) => p[field]).filter((v): v is number => v !== undefined)
+ const complete = coverage.complete && values.length === pools.length
+ const value = field === 'maxwaitSeconds' ? Math.max(0, ...values) : values.reduce((a, b) => a + b, 0)
+ const label = { clActive: 'Active clients', clWaiting: 'Waiting clients', svActive: 'Active servers', svIdle: 'Idle servers', svUsed: 'Used servers', maxwaitSeconds: 'Max wait' }[field]
+ if ((!complete && value === 0) || (pools.length > 0 && values.length === 0)) {
+ return { text: 'Unknown', tone: 'unknown', source: coverage.limitation ?? `${label} not reported for every pool` }
+ }
+ const text = field === 'maxwaitSeconds' ? `${value.toFixed(1)} s` : String(value)
+ return { text: `${complete ? '' : '≥'}${text}`, tone: field === 'clWaiting' && value > 0 ? 'degraded' : complete ? 'neutral' : 'unknown', source: complete ? undefined : coverage.limitation ?? `${label} not reported for every pool` }
+}
diff --git a/packages/k8s-ui/src/components/cnpg/primitives.tsx b/packages/k8s-ui/src/components/cnpg/primitives.tsx
index 7020a4da81..2c753cb1f6 100644
--- a/packages/k8s-ui/src/components/cnpg/primitives.tsx
+++ b/packages/k8s-ui/src/components/cnpg/primitives.tsx
@@ -1,134 +1,12 @@
-import type { ReactNode } from 'react'
import { clsx } from 'clsx'
-import type { HealthLevel } from '../resources/resource-utils'
-import { formatAge } from '../resources/resource-utils'
-import { StatusDot } from '../ui/status-tone'
-import { Tooltip } from '../ui/Tooltip'
-import { AlertBanner } from '../ui/drawer-components'
-import { TONE_TEXT_CLASS } from '../ui/severity-tone'
-import type { CNPGFact, CNPGProblem } from './workspace'
+import { toneTextClass } from '../ui/status-tone'
-export interface CNPGRef {
- kind: string
- group?: string
- namespace: string
- name: string
-}
-
-export type CNPGNavigate = (ref: CNPGRef) => void
-
-const TONE_TEXT: Record = {
- healthy: 'text-theme-text-primary',
- degraded: TONE_TEXT_CLASS.amber,
- alert: TONE_TEXT_CLASS.orange,
- unhealthy: TONE_TEXT_CLASS.red,
- unknown: 'text-theme-text-tertiary',
- neutral: 'text-theme-text-secondary',
-}
-
-export const CNPG_PRIMARY_BUTTON = 'btn-brand inline-flex items-center gap-1.5 px-3 py-1.5 text-sm font-medium'
-export const CNPG_SECONDARY_BUTTON =
- 'inline-flex items-center gap-1.5 rounded-lg border border-theme-border bg-theme-surface px-3 py-1.5 text-sm text-theme-text-primary transition-colors hover:bg-theme-hover'
-
-export function toneTextClass(tone: HealthLevel): string {
- return TONE_TEXT[tone]
-}
-
-export function FactValue({ fact, className }: { fact: CNPGFact; className?: string }) {
- const age = fact.at ? formatAge(fact.at) : null
- const body = (
-
- {fact.text}
- {age && {fact.text ? ' · ' : ''}{age} ago }
-
- )
- if (!fact.source && !fact.at) return body
- return (
-
- {body}
-
- )
-}
-
-export function FactSource({ fact }: { fact: CNPGFact }) {
- if (!fact.source) return null
- return {fact.source}
-}
-
-export function FactGrid({ children }: { children: ReactNode }) {
- return {children}
-}
-
-export function FactRow({ label, children }: { label: ReactNode; children: ReactNode }) {
+/** status.currentPrimary and the primary role label disagree: both are named rather than one silently winning. */
+export function PrimaryConflictNote({ conflict }: { conflict: { status: string; labelled: string } }) {
return (
- <>
- {label}
- {children}
- >
- )
-}
-
-export function SummaryHeading({ children, hint }: { children: ReactNode; hint?: ReactNode }) {
- return (
-
-
{children}
- {hint &&
{hint} }
+
+ CNPG status says primary {conflict.status} ; the Pod labelled primary is{' '}
+ {conflict.labelled} . Status may be stale, or a failover is under way.
)
}
-
-export function RefLink({ refTo, onNavigate, children, mono }: { refTo: CNPGRef; onNavigate?: CNPGNavigate; children?: ReactNode; mono?: boolean }) {
- const label = children ?? refTo.name
- if (!onNavigate) return
{label}
- return (
-
onNavigate(refTo)}
- className={clsx('text-accent-text hover:underline text-left break-all', mono && 'font-mono')}
- >
- {label}
-
- )
-}
-
-const PROBLEM_VARIANT: Record
= {
- critical: 'error',
- warning: 'warning',
- posture: 'info',
-}
-
-export function ProblemCallout({
- problem,
- more,
- onNavigate,
- action,
- subjectIsSelf,
-}: {
- problem: CNPGProblem
- more?: ReactNode
- onNavigate?: CNPGNavigate
- action?: ReactNode
- /** The callout sits on the subject's own page, so linking to it would loop. */
- subjectIsSelf?: boolean
-}) {
- const aboutChild = !subjectIsSelf && problem.subject.kind !== 'Cluster'
- return (
-
-
- {aboutChild && (
-
- {problem.subject.kind}{' '}
-
-
- )}
- {problem.source === 'audit' ? 'Radar check' : 'Radar issue'}
- {action}
- {more}
-
-
- )
-}
-
-export function ToneDot({ tone }: { tone: HealthLevel }) {
- return
-}
diff --git a/packages/k8s-ui/src/components/cnpg/relations.test.ts b/packages/k8s-ui/src/components/cnpg/relations.test.ts
index 437dafbc70..c98b3d5f9f 100644
--- a/packages/k8s-ui/src/components/cnpg/relations.test.ts
+++ b/packages/k8s-ui/src/components/cnpg/relations.test.ts
@@ -1,11 +1,12 @@
import { describe, expect, it } from 'vitest'
import {
+ cnpgScheduleDestinationBlocker,
+ cnpgArchiveMatchesCluster,
appliedFact,
backupDestination,
backupsForScheduledBackup,
clustersUsingCatalog,
databaseForDeclaration,
- gitopsSourceOf,
inferredObjectStoreHealth,
isBackupFromSchedule,
issuesForObject,
@@ -16,6 +17,23 @@ import {
scheduledBackupOf,
usersOfObjectStore,
} from './relations'
+
+it('recognizes archive origin aliases without merging distinct paths', () => {
+ const cluster = { metadata: { name: 'pg' }, spec: { backup: { barmanObjectStore: { destinationPath: 's3://bucket/path', endpointURL: 'https://storage.example' } } } }
+ const source = { kind: 'inTree' as const, serverName: 'pg', barmanObjectStore: { destinationPath: 's3://BUCKET/path/', endpointURL: 'https://STORAGE.EXAMPLE:443/' } }
+ expect(cnpgArchiveMatchesCluster(cluster, source)).toBe(true)
+ expect(cnpgArchiveMatchesCluster(cluster, { ...source, barmanObjectStore: { ...source.barmanObjectStore, destinationPath: 's3://bucket/other' } })).toBe(false)
+ expect(cnpgArchiveMatchesCluster(cluster, { ...source, barmanObjectStore: { ...source.barmanObjectStore, endpointURL: 'https://storage.example/prefix' } })).toBe(false)
+})
+
+it('checks the destination for the schedule method and keeps unread targets unknown', () => {
+ const cluster = { apiVersion: 'postgresql.cnpg.io/v1', kind: 'Cluster', metadata: { name: 'payments', namespace: 'pg' }, spec: { backup: { volumeSnapshot: {} }, plugins: [{ name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'store' } }] } }
+ const schedule = { metadata: { namespace: 'pg' }, spec: { cluster: { name: 'payments' } } }
+ expect(cnpgScheduleDestinationBlocker(schedule, [cluster])).toBe('No barmanObjectStore destination')
+ expect(cnpgScheduleDestinationBlocker(schedule, [])).toBeNull()
+ expect(cnpgScheduleDestinationBlocker({ ...schedule, spec: { ...schedule.spec, method: 'volumeSnapshot' } }, [cluster])).toBeNull()
+ expect(cnpgScheduleDestinationBlocker({ ...schedule, spec: { ...schedule.spec, method: 'plugin', pluginConfiguration: { name: 'barman-cloud.cloudnative-pg.io' } } }, [cluster])).toBeNull()
+})
import { CNPG_WORKSPACE_KEYS, type CNPGWorkspaceIssue, type CNPGWorkspaceResponse } from './workspace'
const PG = 'postgresql.cnpg.io/v1'
@@ -104,9 +122,10 @@ describe('scheduledBackupOf / backupsForScheduledBackup', () => {
describe('objectStoreForBackup / backupDestination', () => {
const clusters = [pluginCluster('main', 'store-a')]
- it('prefers the store the backup recorded', () => {
+ it('ignores the Backup parameter and infers the store from the current Cluster', () => {
const b = backup('b', { spec: { method: 'plugin', pluginConfiguration: { name: PLUGIN, parameters: { barmanObjectName: 'store-b' } } } })
- expect(objectStoreForBackup(b, clusters)).toEqual({ name: 'store-b', inferred: false })
+ expect(objectStoreForBackup(b, clusters)).toEqual({ name: 'store-a', inferred: true })
+ expect(backupDestination(b, clusters)).toEqual({ type: 'objectStore', name: 'store-a', inferred: true })
})
it("marks a store taken from the Cluster's current plugin as inferred", () => {
@@ -115,6 +134,14 @@ describe('objectStoreForBackup / backupDestination', () => {
expect(backupDestination(b, clusters)).toEqual({ type: 'objectStore', name: 'store-a', inferred: true })
})
+ it('cannot resolve a Backup parameter without a configured, readable target Cluster', () => {
+ const b = backup('b', { spec: { method: 'plugin', pluginConfiguration: { name: PLUGIN, parameters: { barmanObjectName: 'store-b' } } } })
+ for (const targets of [[], [cluster('main')], [pluginCluster('other', 'store-a')], [{ ...pluginCluster('main', 'store-a'), metadata: { name: 'main', namespace: 'other' } }], [cluster('main', 'pg', { plugins: [{ name: PLUGIN, enabled: false, parameters: { barmanObjectName: 'store-a' } }] })]]) {
+ expect(objectStoreForBackup(b, targets)).toBeNull()
+ expect(backupDestination(b, targets)).toEqual({ type: 'unknown' })
+ }
+ })
+
it('checks plugin identity before reading barmanObjectName', () => {
const other = backup('b', { spec: { method: 'plugin', pluginConfiguration: { name: 'other.example.com', parameters: { barmanObjectName: 'store-b' } } } })
expect(objectStoreForBackup(other, clusters)).toBeNull()
@@ -129,6 +156,23 @@ describe('objectStoreForBackup / backupDestination', () => {
})
})
+describe('barman-cloud schedule destinations', () => {
+ const schedule = { apiVersion: PG, kind: 'ScheduledBackup', metadata: { name: 'nightly', namespace: 'pg' }, spec: { cluster: { name: 'main' }, method: 'plugin', pluginConfiguration: { name: PLUGIN, parameters: { barmanObjectName: 'schedule-store', serverName: 'schedule-server' } } } }
+
+ it('blocks a schedule override when the Cluster plugin has no destination', () => {
+ const target = cluster('main', 'pg', { plugins: [{ name: PLUGIN }] })
+ expect(cnpgScheduleDestinationBlocker(schedule, [target])).toBe('No backup destination')
+ expect(objectStoreForBackup(schedule, [target])).toBeNull()
+ })
+
+ it('uses the configured Cluster destination despite a conflicting schedule override', () => {
+ const target = pluginCluster('main', 'cluster-store')
+ expect(cnpgScheduleDestinationBlocker(schedule, [target])).toBeNull()
+ expect(objectStoreForBackup(schedule, [target])).toEqual({ name: 'cluster-store', inferred: true })
+ expect(cnpgScheduleDestinationBlocker(schedule, [cluster('main', 'pg', { plugins: [{ name: PLUGIN, enabled: false, parameters: { barmanObjectName: 'cluster-store' } }] })])).toBe('No backup destination')
+ })
+})
+
describe('ObjectStore users and inferred health', () => {
const store = {
apiVersion: BARMAN,
@@ -254,16 +298,6 @@ describe('declarations', () => {
expect(missingManagedRole({ status: { message: 'connection refused' } }, cluster('main'))).toBeNull()
})
- it('reads the GitOps owner labels', () => {
- expect(gitopsSourceOf({ metadata: { labels: { 'argocd.argoproj.io/instance': 'app' } } })).toEqual({ tool: 'argocd', name: 'app' })
- expect(gitopsSourceOf({ metadata: { labels: { 'kustomize.toolkit.fluxcd.io/name': 'k', 'kustomize.toolkit.fluxcd.io/namespace': 'flux' } } })).toEqual({
- tool: 'flux',
- name: 'k',
- namespace: 'flux',
- })
- expect(gitopsSourceOf({ metadata: {} })).toBeNull()
- })
-
const decl = (kind: string, name: string, clusterName: string, dbname: string, ns = 'pg') => ({
apiVersion: PG,
kind,
@@ -302,3 +336,19 @@ describe('relationUnavailable', () => {
expect(relationUnavailable(null, 'backups', 'pg', 'Backups')).toBe('Backups could not be read')
})
})
+
+it('does not apply declaration results from an earlier spec', () => {
+ for (const applied of [true, false]) {
+ expect(appliedFact({ metadata: { generation: 3 }, status: { applied, observedGeneration: 2 } })).toEqual({ text: 'Pending · awaiting the operator for the current spec', tone: 'unknown' })
+ }
+ expect(appliedFact({ metadata: { generation: 3 }, status: { applied: true, observedGeneration: 3 } }).tone).toBe('healthy')
+})
+
+
+it('never infers a plugin archive from a replacement Cluster', () => {
+ const target = { ...pluginCluster('main', 'new-store'), metadata: { name: 'main', namespace: 'pg', uid: 'new', creationTimestamp: '2026-10-01T00:00:00Z' } }
+ for (const status of [{ pluginMetadata: { clusterUID: 'old' } }, { startedAt: '2026-09-30T00:00:00Z' }]) {
+ const old = backup('old', { spec: { method: 'plugin', pluginConfiguration: { name: PLUGIN } }, status })
+ expect(objectStoreForBackup(old, [target])).toBeNull()
+ }
+})
diff --git a/packages/k8s-ui/src/components/cnpg/relations.ts b/packages/k8s-ui/src/components/cnpg/relations.ts
index 2db316d662..179661f9b1 100644
--- a/packages/k8s-ui/src/components/cnpg/relations.ts
+++ b/packages/k8s-ui/src/components/cnpg/relations.ts
@@ -1,9 +1,9 @@
+import { normalizeURLForComparison } from '../../utils/url-path'
+import { cnpgBackupDeclaration, cnpgBackupDestinationBlocker, cnpgBackupBlockerText } from '../../utils/cnpg-backup'
// Pure relationship lookups between CloudNativePG objects in the workspace
// payload. Each helper answers only from what the objects record; a relation
// that cannot be established returns null or an empty list, never a guess.
-import type { BadgeSeverity } from '../ui/Badge'
-import type { HealthLevel } from '../resources/resource-utils'
import {
CNPG_BARMAN_PLUGIN_NAME,
CNPG_GROUP,
@@ -12,15 +12,8 @@ import {
isApiGroup,
type CNPGObjectStoreRecoveryWindow,
} from '../resources/resource-utils-cnpg'
-import {
- cnpgIssueCategory,
- coverageReadable,
- type CNPGFact,
- type CNPGProblem,
- type CNPGWorkspaceIssue,
- type CNPGWorkspaceKey,
- type CNPGWorkspaceResponse,
-} from './workspace'
+import { cnpgIssueCategory, cnpgIssueOrigin, cnpgIssueText, cnpgCoverageGap, coverageReadable, type CNPGProblem, type CNPGWorkspaceIssue, type CNPGWorkspaceKey, type CNPGWorkspaceResponse } from './workspace'
+import { type Fact } from '../facts'
export interface CNPGObjectRef {
kind: string
@@ -58,21 +51,6 @@ export function refOf(obj: any, kind: string, group: string = CNPG_GROUP): CNPGO
return { kind, group, namespace: nsOf(obj), name: nameOf(obj) }
}
-export function healthSeverity(level: HealthLevel): BadgeSeverity {
- switch (level) {
- case 'healthy':
- return 'success'
- case 'unhealthy':
- return 'error'
- case 'alert':
- return 'alert'
- case 'degraded':
- return 'warning'
- default:
- return 'neutral'
- }
-}
-
// ---------------------------------------------------------------------------
// Workspace access
// ---------------------------------------------------------------------------
@@ -94,17 +72,7 @@ export function relationUnavailable(
if (!ws) return `${what} could not be read`
const cov = ws.coverage?.[key] ?? { state: 'notInstalled' as const }
if (coverageReadable(cov, namespace)) return null
- switch (cov.state) {
- case 'denied':
- case 'partial':
- return `No access to ${what}`
- case 'syncing':
- return 'Loading…'
- case 'error':
- return `Could not read ${what}`
- default:
- return `${what} are not installed`
- }
+ return cnpgCoverageGap(cov, what, namespace, `${what} are not installed`)
}
export function clustersIn(ws: CNPGWorkspaceResponse | null | undefined): any[] {
@@ -115,7 +83,8 @@ export function clustersIn(ws: CNPGWorkspaceResponse | null | undefined): any[]
export function targetCluster(obj: any, clusters: any[]): any | null {
const name = specCluster(obj)
if (!name) return null
- return clusters.find((c) => isCNPGKind(c, 'Cluster') && nsOf(c) === nsOf(obj) && nameOf(c) === name) ?? null
+ const cluster = clusters.find((c) => isCNPGKind(c, 'Cluster') && nsOf(c) === nsOf(obj) && nameOf(c) === name) ?? null
+ return isCNPGKind(obj, 'Backup') && !cnpgBackupMatchesCluster(obj, cluster) ? null : cluster
}
// ---------------------------------------------------------------------------
@@ -140,10 +109,10 @@ export function problemsForObject(issues: CNPGWorkspaceIssue[] | undefined, ref:
id: `${issue.id}:${issue.kind}/${issue.name}`,
severity: issue.severity,
category: cnpgIssueCategory(issue),
- title: issue.message || issue.reason,
- detail: issue.cause || undefined,
+ ...cnpgIssueText(issue),
subject: { kind: issue.kind, group: issue.group ?? '', namespace: issue.namespace ?? '', name: issue.name },
source: 'issue',
+ origin: cnpgIssueOrigin(issue),
}))
}
@@ -195,20 +164,50 @@ export function backupTime(backup: any): number {
return parseTime(backup?.status?.startedAt) || parseTime(backup?.metadata?.creationTimestamp)
}
+export function cnpgBackupMatchesCluster(backup: any, cluster: any): boolean {
+ if (!cluster || backup?.spec?.cluster?.name !== cluster.metadata?.name || backup.metadata?.namespace !== cluster.metadata?.namespace) return false
+ const owner = backup.metadata?.ownerReferences?.find((ref: any) => ref.kind === 'Cluster' && isApiGroup(ref.apiVersion, CNPG_GROUP))
+ const recordedUID = backup.status?.pluginMetadata?.clusterUID || owner?.uid
+ if (recordedUID && recordedUID !== cluster.metadata?.uid) return false
+ const began = Date.parse(backup.status?.startedAt ?? backup.metadata?.creationTimestamp ?? '')
+ const created = Date.parse(cluster.metadata?.creationTimestamp ?? '')
+ return !(Number.isFinite(began) && Number.isFinite(created) && began < created)
+}
+
+export type CNPGArchiveSource =
+ | { kind: 'objectStore'; objectStore: string; serverName: string }
+ | { kind: 'inTree'; barmanObjectStore: Record; serverName: string }
+
+export function cnpgArchiveMatchesCluster(cluster: any, source: CNPGArchiveSource): boolean {
+ if (source.kind === 'objectStore') {
+ const plugin = getCNPGClusterBarmanPlugin(cluster)
+ return plugin?.barmanObjectName === source.objectStore &&
+ (plugin?.serverName || cluster.metadata?.name) === source.serverName
+ }
+ const archive = cluster.spec?.backup?.barmanObjectStore
+ const path = source.barmanObjectStore.destinationPath
+ const endpoint = source.barmanObjectStore.endpointURL
+ const destination = typeof path === 'string' ? normalizeURLForComparison(path) : null
+ const endpointKey = normalizeURLForComparison(typeof endpoint === 'string' ? endpoint : '')
+ return !!archive?.destinationPath && typeof path === 'string' && !!path &&
+ destination !== null && endpointKey !== null &&
+ (archive.serverName || cluster.metadata?.name) === source.serverName &&
+ normalizeURLForComparison(archive.destinationPath) === destination &&
+ normalizeURLForComparison(archive.endpointURL || '') === endpointKey
+}
+
/**
- * The ObjectStore a barman-cloud plugin Backup wrote to. The Backup's own
- * plugin parameters are a record of that run; the target Cluster's plugin
- * configuration is only what it is configured with now, so a store taken from
- * there is marked inferred.
+ * Barman-cloud ignores Backup and ScheduledBackup parameters. The current
+ * Cluster plugin chooses the ObjectStore; it may have changed since a Backup
+ * ran, so the historical destination remains inferred.
*/
export function objectStoreForBackup(backup: any, clusters: any[]): { name: string; inferred: boolean } | null {
const method = backup?.status?.method || backup?.spec?.method
if (method !== 'plugin') return null
const cfg = backup?.spec?.pluginConfiguration
if (cfg?.name !== CNPG_BARMAN_PLUGIN_NAME) return null
- const own = cfg?.parameters?.barmanObjectName
- if (typeof own === 'string' && own) return { name: own, inferred: false }
const cluster = targetCluster(backup, clusters)
+ if (backup.kind === 'Backup' && !cnpgBackupMatchesCluster(backup, cluster)) return null
const current = cluster ? getCNPGClusterBarmanPlugin(cluster)?.barmanObjectName : undefined
return current ? { name: current, inferred: true } : null
}
@@ -229,6 +228,14 @@ export function backupDestination(backup: any, clusters: any[]): CNPGBackupDesti
return { type: 'unknown' }
}
+/** Whether the Cluster declares a destination for this schedule's method; unread Clusters stay unknown. */
+export function cnpgScheduleDestinationBlocker(schedule: any, clusters: any[]): string | null {
+ const cluster = targetCluster(schedule, clusters)
+ if (!cluster) return null
+ const blocker = cnpgBackupDestinationBlocker(cnpgBackupDeclaration(cluster), schedule.spec?.method || 'barmanObjectStore', schedule.spec?.pluginConfiguration?.name)
+ return blocker ? cnpgBackupBlockerText(blocker) : null
+}
+
// ---------------------------------------------------------------------------
// ObjectStore
// ---------------------------------------------------------------------------
@@ -256,16 +263,16 @@ export function usersOfObjectStore(store: any, clusters: any[]): CNPGObjectStore
export interface CNPGObjectStoreEvidence {
cluster: CNPGObjectRef
serverName: string
- archiving: CNPGFact
+ archiving: Fact
window: CNPGObjectStoreRecoveryWindow | null
}
export interface CNPGObjectStoreHealth {
- summary: CNPGFact
+ summary: Fact
evidence: CNPGObjectStoreEvidence[]
}
-function archivingFact(cluster: any): CNPGFact {
+function archivingFact(cluster: any): Fact {
const conds = cluster?.status?.conditions
const c = Array.isArray(conds) ? conds.find((x: any) => x?.type === 'ContinuousArchiving') : null
if (!c) return { text: 'WAL archiving not reported', tone: 'unknown' }
@@ -319,14 +326,17 @@ export function inferredObjectStoreHealth(store: any, users: CNPGObjectStoreUser
// Declarative objects
// ---------------------------------------------------------------------------
-export function appliedFact(obj: any): CNPGFact {
+export function appliedFact(obj: any): Fact {
+ if (observedGenerationFact(obj).tone === 'degraded') {
+ return { text: 'Pending · awaiting the operator for the current spec', tone: 'unknown' }
+ }
const applied = obj?.status?.applied
if (applied === true) return { text: 'Applied', tone: 'healthy' }
if (applied === false) return { text: 'Not applied', tone: 'unhealthy' }
return { text: 'Pending · the operator has not reported a result yet', tone: 'unknown' }
}
-export function observedGenerationFact(obj: any): CNPGFact {
+export function observedGenerationFact(obj: any): Fact {
const observed = obj?.status?.observedGeneration
const generation = obj?.metadata?.generation
if (typeof observed !== 'number') return { text: 'Not reported', tone: 'unknown' }
@@ -351,7 +361,6 @@ export function missingManagedRole(obj: any, cluster: any | null): string | null
return names.includes(m[1]) ? null : m[1]
}
-export { cnpgGitOpsSource as gitopsSourceOf } from './workspace'
/** Publications and Subscriptions on the same Cluster and PostgreSQL database. */
export function replicationForDatabase(
diff --git a/packages/k8s-ui/src/components/cnpg/schedule.ts b/packages/k8s-ui/src/components/cnpg/schedule.ts
new file mode 100644
index 0000000000..0d93113cf5
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/schedule.ts
@@ -0,0 +1,35 @@
+/**
+ * The server's reading of a ScheduledBackup schedule (see
+ * /api/cnpg/scheduledbackups/{ns}/{name}/schedule-preview): parsed as
+ * CloudNativePG parses it, with the next runs counted as the operator counts
+ * them. Times are RFC 3339 UTC.
+ */
+export interface CNPGSchedulePreview {
+ schedule: string
+ valid: boolean
+ error?: string
+ description?: string
+ nextRuns?: string[]
+ /** The first run is due already: the operator creates a backup as soon as it reconciles. */
+ runsImmediately?: boolean
+ basis: 'lastCheckTime' | 'now'
+ lastCheckTime?: string
+ suspended?: boolean
+ clock?: { zone: string; declared: boolean; source: string }
+}
+
+export function formatCNPGRunTime(iso: string): { utc: string; local: string } {
+ const d = new Date(iso)
+ const pad = (n: number) => String(n).padStart(2, '0')
+ const utc = `${d.getUTCFullYear()}-${pad(d.getUTCMonth() + 1)}-${pad(d.getUTCDate())} ${pad(d.getUTCHours())}:${pad(d.getUTCMinutes())}:${pad(d.getUTCSeconds())} UTC`
+ return { utc, local: d.toLocaleString() }
+}
+
+/** One line for the runs list: why the first run is when it is. */
+export function cnpgScheduleBasisNote(p: CNPGSchedulePreview): string {
+ const clock = p.clock?.source ?? 'Operator clock is not established; upcoming times assume UTC.'
+ if (p.suspended) return `Suspended: nothing runs until it is resumed. ${clock}`
+ const basis = p.basis === 'lastCheckTime' ? "Counted from the operator's last check (status.lastCheckTime)." : 'Counted from now; the operator starts counting at its first check.'
+ const due = p.runsImmediately ? ` ${p.clock?.declared ? 'A run is due on the declared clock.' : 'A run is due in this UTC estimate.'} The operator takes at most one catch-up backup.` : ''
+ return `${basis} ${clock}${due}`
+}
diff --git a/packages/k8s-ui/src/components/cnpg/workspace-disk.test.ts b/packages/k8s-ui/src/components/cnpg/workspace-disk.test.ts
new file mode 100644
index 0000000000..54850a0087
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/workspace-disk.test.ts
@@ -0,0 +1,97 @@
+import { describe, it, expect } from 'vitest'
+import { applyCNPGDisk, buildCNPGFleet, cnpgDiskFact, CNPG_WORKSPACE_KEYS, type CNPGDiskReading, type CNPGWorkspaceResponse } from './workspace'
+
+function cluster(name: string): any {
+ return {
+ apiVersion: 'postgresql.cnpg.io/v1',
+ kind: 'Cluster',
+ metadata: { name, namespace: 'db' },
+ spec: { instances: 1 },
+ status: { phase: 'Cluster in healthy state', readyInstances: 1, currentPrimary: `${name}-1` },
+ }
+}
+
+function fleetOf(...names: string[]) {
+ const coverage: CNPGWorkspaceResponse['coverage'] = {}
+ for (const k of CNPG_WORKSPACE_KEYS) coverage[k] = { state: 'full' }
+ return buildCNPGFleet({
+ installed: true,
+ context: 'test',
+ namespaces: null,
+ coverage,
+ objects: { clusters: names.map(cluster) },
+ issues: [],
+ audit: [],
+ backupsOmitted: 0,
+ })
+}
+
+function reading(name: string, over: Partial = {}): CNPGDiskReading {
+ return { namespace: 'db', name, state: 'ok', claims: 1, measured: 1, ...over }
+}
+
+const max = (ratio: number) => ({ claim: 'pg-1', instance: 'pg-1', role: 'PG_DATA', usedBytes: ratio * 10 * 1024 ** 3, capacityBytes: 10 * 1024 ** 3, ratio })
+
+describe('cnpgDiskFact', () => {
+ it('never reads an unmeasured cluster as zero or healthy', () => {
+ for (const state of ['noSeries', 'noPrometheus', 'denied', 'error', 'notRead']) {
+ const f = cnpgDiskFact(reading('pg', { state, measured: 0 }))
+ expect(f.tone).toBe('unknown')
+ expect(f.text).not.toMatch(/\d/)
+ }
+ expect(cnpgDiskFact(undefined).tone).toBe('unknown')
+ })
+
+ it('names the fullest volume and its source, and labels partial coverage', () => {
+ const f = cnpgDiskFact(reading('pg', { state: 'partial', claims: 2, measured: 1, max: max(0.5) }))
+ expect(f.text).toBe('50% used')
+ expect(f.tone).toBe('healthy')
+ expect(f.source).toContain('data volume of pg-1')
+ expect(f.source).toContain('kubelet volume stats')
+ expect(f.source).toContain('1 of 2 volumes measured')
+ })
+
+ it('names the missing grant when denied', () => {
+ expect(cnpgDiskFact(reading('pg', { state: 'denied', grant: { verb: 'list', resource: 'persistentvolumeclaims', namespace: 'db' }, measured: 0 })).source).toBe('Needs list persistentvolumeclaims in namespace db')
+ })
+})
+
+describe('applyCNPGDisk', () => {
+ it('says when the volume stats were matched to the cluster by claim name only', () => {
+ const note = "Matched by namespace and claim names. Radar couldn't confirm these volume stats belong to this exact cluster (no cluster label it could check)"
+ const fleet = applyCNPGDisk(fleetOf('pg-a'), [reading('pg-a', { max: max(0.95), isolation: { mode: 'unverified', note } })])
+ const p = fleet.rows[0].problems[0]
+ expect(p).toMatchObject({ measuredBy: 'kubelet, matched by claim name', unverifiedMatch: true })
+ expect(p.detail).toContain(note)
+ expect(fleet.rows[0].disk?.source).toContain(note)
+ })
+
+ it('puts low-disk clusters into Needs attention by severity', () => {
+ const fleet = applyCNPGDisk(fleetOf('pg-a', 'pg-b', 'pg-c'), [
+ reading('pg-a', { max: max(0.85) }),
+ reading('pg-b', { max: max(0.95) }),
+ reading('pg-c', { max: max(0.4) }),
+ ])
+ expect(fleet.attentionCount).toBe(2)
+ const a = fleet.rows.find((r) => r.name === 'pg-a')!
+ const b = fleet.rows.find((r) => r.name === 'pg-b')!
+ const c = fleet.rows.find((r) => r.name === 'pg-c')!
+ expect(a.problems[0].severity).toBe('warning')
+ expect(b.problems[0].severity).toBe('critical')
+ expect(b.problems[0].source).toBe('measurement')
+ expect(b.problems[0].title).toBe('The data volume of pg-1 is 95% full')
+ expect(c.attention).toBe(false)
+ expect(fleet.categoryCounts.availability).toBe(2)
+ })
+
+ it('raises nothing without a measurement', () => {
+ const fleet = applyCNPGDisk(fleetOf('pg-a'), [reading('pg-a', { state: 'noPrometheus', measured: 0, reason: 'Radar is not connected to Prometheus: x' })])
+ expect(fleet.attentionCount).toBe(0)
+ expect(fleet.rows[0].disk).toMatchObject({ text: 'No usage metrics', source: 'Prometheus not connected', detail: 'Radar is not connected to Prometheus: x' })
+ })
+
+ it('leaves the fleet as built when no reading was requested', () => {
+ const base = fleetOf('pg-a')
+ expect(applyCNPGDisk(base, undefined)).toBe(base)
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/workspace-fleet-metrics.test.ts b/packages/k8s-ui/src/components/cnpg/workspace-fleet-metrics.test.ts
new file mode 100644
index 0000000000..d8227cb983
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/workspace-fleet-metrics.test.ts
@@ -0,0 +1,246 @@
+import { describe, it, expect } from 'vitest'
+import { applyCNPGFleetMetrics, buildCNPGFleet, CNPG_WORKSPACE_KEYS, type CNPGFleetMetricsReading, type CNPGWorkspaceResponse } from './workspace'
+
+function pod(cluster: string, n: number, role: 'primary' | 'replica'): any {
+ return {
+ metadata: { name: `${cluster}-${n}`, namespace: 'db', labels: { 'cnpg.io/cluster': cluster, 'cnpg.io/instanceRole': role } },
+ status: { conditions: [{ type: 'Ready', status: 'True' }] },
+ }
+}
+
+function fleet() {
+ const coverage: CNPGWorkspaceResponse['coverage'] = {}
+ for (const k of CNPG_WORKSPACE_KEYS) coverage[k] = { state: 'full' }
+ const cluster = (name: string, instances: number) => ({
+ apiVersion: 'postgresql.cnpg.io/v1',
+ kind: 'Cluster',
+ metadata: { name, namespace: 'db' },
+ spec: { instances },
+ status: { phase: 'Cluster in healthy state', readyInstances: instances, currentPrimary: `${name}-1` },
+ })
+ return buildCNPGFleet({
+ installed: true,
+ context: 'test',
+ namespaces: null,
+ coverage,
+ objects: {
+ clusters: [cluster('ha', 2), cluster('solo', 1), cluster('dark', 2)],
+ pods: [pod('ha', 1, 'primary'), pod('ha', 2, 'replica'), pod('solo', 1, 'primary'), pod('dark', 1, 'primary'), pod('dark', 2, 'replica')],
+ },
+ issues: [],
+ audit: [],
+ backupsOmitted: 0,
+ })
+}
+
+const reading = (name: string, lag: CNPGFleetMetricsReading['lag'], growth: CNPGFleetMetricsReading['growth'] = { state: 'noSeries' }): CNPGFleetMetricsReading => ({
+ namespace: 'db',
+ name,
+ lag,
+ growth,
+})
+
+const row = (f: ReturnType, name: string) => f.rows.find((r) => r.name === name)!
+
+describe('applyCNPGFleetMetrics', () => {
+ it('replaces "lag unknown" with the measured standby lag and names its source', () => {
+ const f = applyCNPGFleetMetrics(fleet(), [reading('ha', { state: 'ok', seconds: 7.2, pod: 'ha-2' })], { source: 'prometheus', lagSource: 'Prometheus cnpg_pg_replication_lag' })
+ const r = row(f, 'ha')
+ expect(r.replication.text).toBe('1/1 Pods ready · lag 7.2 s')
+ expect(r.replication.tone).toBe('degraded')
+ expect(r.replication.source).toContain('ha-2')
+ expect(r.replication.source).toContain('cnpg_pg_replication_lag')
+ })
+
+ it('says why lag is unknown instead of zero, and leaves non-replica facts alone', () => {
+ const f = applyCNPGFleetMetrics(fleet(), [reading('dark', { state: 'noSeries', reason: 'no exporter series' })], { source: 'prometheus' })
+ expect(row(f, 'dark').replication.text).toBe('1/1 Pods ready · lag unknown')
+ expect(row(f, 'dark').replication.source).toBeTruthy()
+ expect(row(f, 'dark').replication.tone).toBe('unknown')
+ expect(row(f, 'solo').replication.text).toBe('Single instance')
+
+ const denied = applyCNPGFleetMetrics(fleet(), [reading('ha', { state: 'denied', grant: { verb: 'get', resource: 'pods', namespace: 'db' } })], { source: 'prometheus' })
+ expect(row(denied, 'ha').replication.text).toBe('1/1 Pods ready · lag unknown')
+ expect(row(denied, 'ha').replication.source).toBe('Needs get pods in namespace db')
+
+ const none = applyCNPGFleetMetrics(fleet(), undefined, { source: 'none', reason: 'Radar is not connected to Prometheus' })
+ expect(row(none, 'ha').replication.text).toBe('1/1 Pods ready · lag unknown')
+ expect(row(none, 'ha').replication.source).toBe('Prometheus not connected')
+ expect(row(none, 'ha').replication.detail).toBe('Radar is not connected to Prometheus')
+ expect(row(none, 'ha').diskGrowth).toBeUndefined()
+ })
+
+ it('keeps the fleet untouched when no reading was requested', () => {
+ const base = fleet()
+ expect(applyCNPGFleetMetrics(base, undefined, undefined)).toBe(base)
+ expect(row(base, 'ha').replication.text).toBe('1/1 Pods ready · lag unknown')
+ })
+
+ it('reports disk growth only when measured', () => {
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 0.2, pod: 'ha-2' }, { state: 'ok', bytesPerHour: 1024 ** 3 / 24, claim: 'ha-1', instance: 'ha-1' })],
+ { source: 'prometheus', growthSource: 'deriv' },
+ )
+ expect(row(f, 'ha').diskGrowth?.text).toBe('+1.0 GiB/day')
+ expect(row(f, 'ha').replication.tone).toBe('healthy')
+ expect(row(f, 'dark').diskGrowth).toBeUndefined()
+ })
+})
+
+describe('sustained replication lag', () => {
+ const src = { source: 'prometheus' as const, lagSource: 'Prometheus cnpg_pg_replication_lag' }
+
+ it('puts a cluster into Needs attention only when lag stayed high for the whole window', () => {
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [
+ reading('ha', { state: 'ok', seconds: 95, pod: 'ha-2', sustainedSeconds: 40, sustainedPod: 'ha-2', sustainedWindow: '10m0s' }),
+ reading('dark', { state: 'ok', seconds: 120, pod: 'dark-2' }),
+ ],
+ src,
+ )
+ const ha = row(f, 'ha')
+ expect(ha.attention).toBe(true)
+ expect(ha.problems[0]).toMatchObject({ severity: 'warning', category: 'availability', source: 'measurement' })
+ expect(ha.problems[0].title).toBe('ha-2 ≥ 40 s behind in every sample for 10 min')
+ expect(ha.problems[0]).toMatchObject({ measuredBy: 'Prometheus' })
+ expect(ha.problems[0].detail).not.toMatch(/cnpg_/)
+ expect(ha.problems[0].detail).toContain("If Prometheus missed some scrapes, those moments aren't included")
+ expect(ha.problems[0].shortTitle).toBe('ha-2: all samples ≥ 40 s behind (10 min)')
+ expect(ha.problems[0].detail).not.toContain('whole window')
+ expect(row(f, 'dark').attention).toBe(false)
+ expect(f.attentionCount).toBe(1)
+ })
+
+ it('says when the series were matched to the cluster by Pod name only', () => {
+ const note = "Matched by namespace and Pod names. Radar couldn't confirm these series belong to this exact cluster (no cluster label it could check)"
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 95, pod: 'ha-2', sustainedSeconds: 40, sustainedPod: 'ha-2', sustainedWindow: '10m0s', isolation: { mode: 'unverified', note } })],
+ src,
+ )
+ const p = row(f, 'ha').problems[0]
+ expect(p.measuredBy).toBe('Prometheus, matched by Pod name')
+ expect(p.unverifiedMatch).toBe(true)
+ expect(p.detail).toContain(note)
+ const verified = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 95, pod: 'ha-2', sustainedSeconds: 40, sustainedPod: 'ha-2', sustainedWindow: '10m0s', isolation: { mode: 'verified', note: 'x' } })],
+ src,
+ )
+ expect(row(verified, 'ha').problems[0].measuredBy).toBe('Prometheus')
+ })
+
+ it('escalates to critical past five minutes and ignores a floor under 30 s', () => {
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [
+ reading('ha', { state: 'ok', seconds: 400, pod: 'ha-2', sustainedSeconds: 360, sustainedPod: 'ha-2', sustainedWindow: '10m0s' }),
+ reading('dark', { state: 'ok', seconds: 20, pod: 'dark-2', sustainedSeconds: 12, sustainedPod: 'dark-2', sustainedWindow: '10m0s' }),
+ ],
+ src,
+ )
+ expect(row(f, 'ha').problems[0]).toMatchObject({ severity: 'critical', title: 'ha-2 ≥ 6 min behind in every sample for 10 min' })
+ expect(row(f, 'dark').problems).toHaveLength(0)
+ })
+})
+
+describe('standbys that receive nothing, and the WAL slots hold', () => {
+ const src = { source: 'prometheus' as const, lagSource: 'Prometheus cnpg_pg_replication_lag' }
+ it('a standby whose WAL receiver is down is a problem, never "lag 0 s"', () => {
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', standbys: 1, receiving: 0, receiverDown: ['ha-2'], receiverDownSustained: ['ha-2'], receiverDownWindow: '5m0s' })],
+ src,
+ )
+ const r = row(f, 'ha')
+ expect(r.replication.text).toBe('1/1 Pods ready · ha-2 not receiving WAL')
+ expect(r.replication.tone).toBe('unhealthy')
+ expect(r.attention).toBe(true)
+ const p = r.problems.find((x) => x.id === 'standby:db/ha:ha-2')!
+ expect(p).toMatchObject({ severity: 'critical', title: 'ha-2 is not receiving WAL from the primary', subject: { kind: 'Pod', name: 'ha-2' }, measuredBy: 'Prometheus' })
+ expect(p.detail).toContain('Its WAL receiver was down in every sample Prometheus recorded over the last 5 minutes')
+ })
+ it('a receiver down for less than the window is shown, not raised: a restarting standby reconnects on its own', () => {
+ const f = applyCNPGFleetMetrics(fleet(), [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', standbys: 1, receiving: 0, receiverDown: ['ha-2'], receiverDownSustained: [], receiverDownWindow: '5m0s' })], src)
+ expect(row(f, 'ha').replication.text).toBe('1/1 Pods ready · ha-2 not receiving WAL')
+ expect(row(f, 'ha').problems.some((p) => p.id.startsWith('standby:'))).toBe(false)
+ })
+ it('lag alone does not establish streaming when the exporter reports no receiver state', () => {
+ const f = applyCNPGFleetMetrics(fleet(), [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', standbys: 1, receiverUnknown: true })], src)
+ expect(row(f, 'ha').replication).toMatchObject({ text: '1/1 Pods ready · lag 0 s · streaming unverified', tone: 'unknown' })
+ expect(row(f, 'ha').attention).toBe(false)
+ })
+ it('raises an inactive slot only from the primary, at 1 GiB or more, naming the standby it serves', () => {
+ const big = 4.9e9
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [
+ {
+ ...reading('ha', { state: 'ok', seconds: 0.1, pod: 'ha-2', standbys: 1, receiving: 1 }),
+ slots: {
+ state: 'ok',
+ inactive: [
+ { slot: '_cnpg_ha_2', pod: 'ha-1', role: 'primary', bytes: big },
+ { slot: '_cnpg_ha_2', pod: 'ha-2', role: 'standby', bytes: big },
+ ],
+ },
+ },
+ { ...reading('dark', { state: 'ok', seconds: 0.1, pod: 'dark-2', standbys: 1, receiving: 1 }), slots: { state: 'ok', inactive: [{ slot: '_cnpg_dark_9', pod: 'dark-2', role: 'standby', bytes: big }] } },
+ ],
+ src,
+ )
+ const slots = row(f, 'ha').problems.filter((p) => p.id.startsWith('slot:'))
+ expect(slots).toHaveLength(1)
+ expect(slots[0].title).toBe('Inactive slot _cnpg_ha_2 holds 4.6 GiB of WAL on ha-1 for ha-2')
+ expect(row(f, 'dark').problems.some((p) => p.id.startsWith('slot:'))).toBe(false)
+ const small = applyCNPGFleetMetrics(fleet(), [{ ...reading('ha', { state: 'ok', seconds: 0.1, pod: 'ha-2' }), slots: { state: 'ok', inactive: [{ slot: '_cnpg_ha_2', pod: 'ha-1', role: 'primary', bytes: 5e8 }] } }], src)
+ expect(row(small, 'ha').problems.some((p) => p.id.startsWith('slot:'))).toBe(false)
+ })
+})
+
+describe('none receiving needs every expected standby accounted for', () => {
+ it('one standby down with another unreported is a warning, not "none receive"', () => {
+ const f = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', standbys: 1, receiving: 0, receiverDown: ['ha-2'], receiverDownSustained: ['ha-2'], receiverDownWindow: '5m0s' })],
+ { source: 'prometheus' },
+ )
+ expect(row(f, 'ha').problems.find((p) => p.id === 'standby:db/ha:ha-2')?.severity).toBe('critical')
+ const three = applyCNPGFleetMetrics(
+ fleet(),
+ [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', standbys: 2, receiving: 0, receiverDown: ['ha-2'], receiverDownSustained: ['ha-2'], receiverDownWindow: '5m0s' })],
+ { source: 'prometheus' },
+ )
+ // spec.instances 2 expects one standby; two reporting with one down is not "all down".
+ expect(row(three, 'ha').problems.find((p) => p.id === 'standby:db/ha:ha-2')?.severity).toBe('warning')
+ })
+})
+
+describe('fleet lag covers the standbys that report', () => {
+ it('says how many standbys the lag covers, and is not healthy while one is unreported', () => {
+ const three = fleet()
+ three.rows.find((r) => r.name === 'ha')!.instances.desired = 3
+ const f = applyCNPGFleetMetrics(three, [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', lagStandbys: 1, standbys: 1, receiving: 1 })], { source: 'prometheus' })
+ expect(row(f, 'ha').replication).toMatchObject({ text: expect.stringContaining('lag 0 s (1 of 2 standbys reporting)'), tone: 'unknown' })
+ })
+
+ it('counts the standbys whose lag was read, not those whose receiver was', () => {
+ const three = fleet()
+ three.rows.find((r) => r.name === 'ha')!.instances.desired = 3
+ const f = applyCNPGFleetMetrics(three, [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-2', lagStandbys: 1, standbys: 2, receiving: 2 })], { source: 'prometheus' })
+ expect(row(f, 'ha').replication).toMatchObject({ text: expect.stringContaining('(1 of 2 standbys reporting)'), tone: 'unknown' })
+ })
+
+ it('expects every instance of a replica cluster to report, its designated primary included', () => {
+ const replica = fleet()
+ const r = replica.rows.find((x) => x.name === 'ha')!
+ r.instances.desired = 3
+ r.replicaCluster = { source: 'pg-origin' }
+ const f = applyCNPGFleetMetrics(replica, [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-1', lagStandbys: 2, standbys: 2, receiving: 2 })], { source: 'prometheus' })
+ expect(row(f, 'ha').replication).toMatchObject({ text: expect.stringContaining('(2 of 3 instances reporting)'), tone: 'unknown' })
+ const all = applyCNPGFleetMetrics(replica, [reading('ha', { state: 'ok', seconds: 0, pod: 'ha-1', lagStandbys: 3, standbys: 3, receiving: 3 })], { source: 'prometheus' })
+ expect(row(all, 'ha').replication.text).not.toContain('reporting')
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/workspace-problems.test.ts b/packages/k8s-ui/src/components/cnpg/workspace-problems.test.ts
new file mode 100644
index 0000000000..d7f97431c8
--- /dev/null
+++ b/packages/k8s-ui/src/components/cnpg/workspace-problems.test.ts
@@ -0,0 +1,175 @@
+import { describe, expect, it } from 'vitest'
+import { cnpgCollapseBackupFailures, cnpgCompareProblems, cnpgFoldLastBackupFailed, cnpgFormatLag, cnpgIssueOrigin, cnpgIssueText, cnpgIssueTitle, type CNPGProblem } from './workspace'
+
+const problem = (title: string, severity: CNPGProblem['severity'], kind: string, group = ''): CNPGProblem => ({
+ id: title,
+ severity,
+ category: kind === 'Pod' ? 'availability' : 'protection',
+ title,
+ subject: { kind, group, namespace: 'pg', name: kind === 'Pod' ? 'pg-wal-failing-1' : 'pg-wal-failing' },
+ source: 'issue',
+})
+
+describe('cnpgCompareProblems', () => {
+ it('puts the cluster-level cause ahead of a Pod symptom at the same severity (pg-wal-failing)', () => {
+ const probe = problem('pg-wal-failing-1 not ready (readiness probe failing)', 'critical', 'Pod')
+ const wal = problem('WAL archiving failing', 'critical', 'Cluster', 'postgresql.cnpg.io')
+ expect([probe, wal].sort(cnpgCompareProblems).map((p) => p.title)).toEqual(['WAL archiving failing', 'pg-wal-failing-1 not ready (readiness probe failing)'])
+ })
+ it('keeps severity first', () => {
+ const crit = problem('pg-1 restarted recently (CrashLoopBackOff)', 'critical', 'Pod')
+ const warn = problem('Backup failed', 'warning', 'Backup', 'postgresql.cnpg.io')
+ expect([warn, crit].sort(cnpgCompareProblems)[0]).toBe(crit)
+ })
+})
+
+describe('cnpgIssueTitle', () => {
+ it('turns a bare reason into a sentence about the subject', () => {
+ expect(cnpgIssueTitle({ kind: 'Pod', name: 'pg-wal-failing-1', reason: 'ReadinessProbeFailed', message: 'ReadinessProbeFailed' })).toBe(
+ 'pg-wal-failing-1 not ready (readiness probe failing)',
+ )
+ expect(cnpgIssueTitle({ kind: 'Pod', name: 'pg-runtime-5', reason: 'CrashLoopBackOff' })).toBe('pg-runtime-5 restarted recently (CrashLoopBackOff)')
+ expect(cnpgIssueTitle({ kind: 'Backup', name: 'b-1', reason: 'BackupStuck' })).toBe('Backup b-1: backup stuck')
+ })
+ it('keeps a real message', () => {
+ expect(cnpgIssueTitle({ kind: 'Cluster', name: 'pg', reason: 'ContinuousArchivingFailing', message: 'WAL archiving failing' })).toBe('WAL archiving failing')
+ })
+})
+
+describe('cnpgIssueText', () => {
+ it('heads a CNPG condition issue with a plain title and keeps the operator message beneath', () => {
+ const t = cnpgIssueText({
+ kind: 'Cluster',
+ name: 'pg-wal-failing',
+ reason: 'CNPGWALArchivingFailing',
+ message:
+ 'The last WAL archival did not complete; recovery-point advancement is uncertain: unexpected failure invoking barman-cloud-wal-archive: exit status 4',
+ })
+ expect(t.title).toBe('WAL archiving failing')
+ expect(t.detail).toContain('exit status 4')
+ })
+})
+
+describe('cnpgIssueText for certificates and schedules', () => {
+ it('words expired and expiring certificates apart, and does not claim no backup was produced', () => {
+ expect(cnpgIssueText({ kind: 'Cluster', name: 'pg', reason: 'CNPGCertificateExpired', message: 'The certificate in Secret pg-ca expired 2026-09-30T11:00:00Z' }).title).toBe(
+ 'A certificate has expired',
+ )
+ expect(cnpgIssueText({ kind: 'Cluster', name: 'pg', reason: 'CNPGCertificateExpiring', message: 'The certificate in Secret x expires in 3 days' }).title).toBe(
+ 'A certificate expires soon',
+ )
+ expect(cnpgIssueText({ kind: 'ScheduledBackup', name: 's', reason: 'CNPGScheduledRunNoBackup', message: 'x y' }).title).toBe('No successful backup since a scheduled run')
+ })
+})
+
+describe('backup failures', () => {
+ it('does not repeat the title at the start of the detail', () => {
+ expect(cnpgIssueText({ kind: 'Backup', name: 'b', reason: 'CNPGBackupFailed', message: 'Backup failed: cannot proceed with the backup' })).toEqual({
+ title: 'Backup failed',
+ detail: 'Cannot proceed with the backup',
+ })
+ })
+ it('collapses Backups that failed the same way into one problem about the latest, by the Backups\' own times', () => {
+ // detectCNPGBackupIssues emits no first_seen: the latest comes from the Backup objects.
+ const problems = ['b-1', 'b-3', 'b-2'].map((n) => {
+ const t = cnpgIssueText({ kind: 'Backup', name: n, reason: 'CNPGBackupFailed', message: 'Backup failed: cannot proceed as the cluster has no plugin configured' })
+ return {
+ id: n,
+ severity: 'warning' as const,
+ category: 'protection' as const,
+ ...t,
+ subject: { kind: 'Backup', group: 'postgresql.cnpg.io', namespace: 'pg', name: n },
+ source: 'issue' as const,
+ reason: 'CNPGBackupFailed',
+ }
+ })
+ const times = new Map([
+ ['b-1', Date.parse('2026-09-20T00:00:00Z')],
+ ['b-3', Date.parse('2026-09-21T00:00:00Z')],
+ ['b-2', Date.parse('2026-09-22T00:00:00Z')],
+ ])
+ const out = cnpgCollapseBackupFailures(problems, times)
+ expect(out).toHaveLength(1)
+ expect(out[0]).toMatchObject({ title: '3 backups failed: cannot proceed as the cluster has no plugin configured', subject: { name: 'b-2' } })
+ expect(out[0].alsoAbout?.map((o) => o.name)).toEqual(['b-3', 'b-1'])
+ })
+})
+
+describe('scheduled run without a backup', () => {
+ it('words the schedule and dates the run as an age from first_seen', () => {
+ const t = cnpgIssueText({
+ kind: 'Cluster',
+ name: 'pg',
+ reason: 'CNPGScheduledRunNoBackup',
+ message: 'ScheduledBackup pg-nightly (every day at 02:00 UTC) has had no successful backup since its run',
+ first_seen: new Date(Date.now() - 2 * 24 * 3600 * 1000 - 60_000).toISOString(),
+ })
+ expect(t.title).toBe('No successful backup since a scheduled run')
+ expect(t.detail).toBe('ScheduledBackup pg-nightly (every day at 02:00 UTC) has had no successful backup since its run 2d ago')
+ })
+})
+
+describe('cnpgFormatLag', () => {
+ it('reads every replay lag the same way, rounded down', () => {
+ expect(cnpgFormatLag(0)).toBe('0 s')
+ expect(cnpgFormatLag(0.25)).toBe('250 ms')
+ expect(cnpgFormatLag(0.2509)).toBe('250 ms')
+ expect(cnpgFormatLag(8.27)).toBe('8.2 s')
+ expect(cnpgFormatLag(55.9)).toBe('55 s')
+ expect(cnpgFormatLag(1500.4)).toBe('25 min')
+ expect(cnpgFormatLag(1721.3)).toBe('28 min')
+ expect(cnpgFormatLag(3600)).toBe('1 h')
+ expect(cnpgFormatLag(4000)).toBe('1 h 6 min')
+ })
+})
+
+describe('cnpgFoldLastBackupFailed', () => {
+ const p = (id: string, kind: string, name: string, reason: string, alsoAbout?: { kind: string; name: string }[]) => ({
+ id,
+ severity: 'warning' as const,
+ category: 'protection' as const,
+ title: id,
+ subject: { kind, group: 'postgresql.cnpg.io', namespace: 'pg', name },
+ source: 'issue' as const,
+ reason,
+ alsoAbout,
+ })
+ const last = p('last', 'Cluster', 'pg', 'CNPGLastBackupFailed')
+ it('drops "the last backup failed" when a failed-Backup problem already covers the newest Backup', () => {
+ const group = p('3 backups failed', 'Backup', 'b-3', 'CNPGBackupFailed', [{ kind: 'Backup', name: 'b-2' }])
+ const out = cnpgFoldLastBackupFailed([group, last], 'b-3')
+ expect(out.map((x) => x.id)).toEqual(['3 backups failed'])
+ expect(out[0].reason).toBe('CNPGBackupFailed')
+ })
+ it('counts only a failed-Backup problem as covering the newest Backup', () => {
+ const other = p('stuck', 'Backup', 'b-3', 'CNPGBackupStuck')
+ expect(cnpgFoldLastBackupFailed([other, last], 'b-3').map((x) => x.id)).toEqual(['stuck', 'last'])
+ })
+ it('keeps it when the newest Backup is not among the failures, or is unknown', () => {
+ const old = p('old failure', 'Backup', 'b-1', 'CNPGBackupFailed')
+ expect(cnpgFoldLastBackupFailed([old, last], 'b-9')).toHaveLength(2)
+ expect(cnpgFoldLastBackupFailed([old, last], undefined)).toHaveLength(2)
+ })
+})
+
+describe('cnpgIssueOrigin', () => {
+ it('names where the evidence comes from', () => {
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGWALArchivingFailing' })).toEqual({ label: 'Reported by CNPG', detail: 'Cluster ContinuousArchiving condition' })
+ expect(cnpgIssueOrigin({ kind: 'Backup', reason: 'CNPGWALArchivingFailing' }).label).toBe('Backup status')
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGClusterDegraded' }).label).toBe('Radar check of ready instances')
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGLastBackupFailed' }).label).toBe('Reported by CNPG')
+ expect(cnpgIssueOrigin({ kind: 'Database', reason: 'CNPGDeclarativeNotApplied' }).label).toBe('Reported by CNPG')
+ expect(cnpgIssueOrigin({ kind: 'ScheduledBackup', reason: 'CNPGScheduledBackupMissed' }).label).toBe('Radar check of the backup schedule')
+ expect(cnpgIssueOrigin({ kind: 'Pod', reason: 'HighRestartCount' }).label).toBe('Radar check of restarts')
+ expect(cnpgIssueOrigin({ kind: 'Pod', reason: 'ReadinessProbeInvalid' }).label).toBe('Radar check of the probe')
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGClusterFailingOver' }).label).toBe('Reported by CNPG')
+ expect(cnpgIssueOrigin({ kind: 'Backup', reason: 'CNPGBackupFailed' }).label).toBe('Backup status')
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGScheduledRunNoBackup' }).label).toBe('Radar check of the backup schedule')
+ expect(cnpgIssueOrigin({ kind: 'Cluster', reason: 'CNPGCertificateExpired' }).label).toBe('Certificate expiry (from Cluster status)')
+ expect(cnpgIssueOrigin({ kind: 'Pod', reason: 'ReadinessProbeFailed' }).label).toBe('Pod readiness probe')
+ expect(cnpgIssueOrigin({ kind: 'Pod', reason: 'CrashLoopBackOff' }).label).toBe('Pod status')
+ })
+ it('falls back to "Detected by Radar" without inventing a source', () => {
+ expect(cnpgIssueOrigin({ kind: 'Pooler', reason: 'SomethingNew' })).toEqual({ label: 'Detected by Radar' })
+ })
+})
diff --git a/packages/k8s-ui/src/components/cnpg/workspace.test.ts b/packages/k8s-ui/src/components/cnpg/workspace.test.ts
index beeb5dd65b..34c4092485 100644
--- a/packages/k8s-ui/src/components/cnpg/workspace.test.ts
+++ b/packages/k8s-ui/src/components/cnpg/workspace.test.ts
@@ -1,8 +1,29 @@
+import { cnpgDimensions } from './ha'
import { describe, it, expect } from 'vitest'
-import { buildCNPGFleet, type CNPGWorkspaceResponse, type CNPGWorkspaceKey, CNPG_WORKSPACE_KEYS } from './workspace'
+import { buildCNPGFleet, cnpgReadyInstances, cnpgRecoveryMatchesCluster, getCNPGRestoreValidation, type CNPGWorkspaceResponse, type CNPGWorkspaceKey, CNPG_WORKSPACE_KEYS } from './workspace'
const G = 'postgresql.cnpg.io/v1'
+it('does not promote an archiving condition into evidence of an archive destination', () => {
+ const c = cluster('payments', 'db', { status: { conditions: [{ type: 'ContinuousArchiving', status: 'True', lastTransitionTime: '2026-10-01T12:00:00Z' }] } })
+ const fact = buildCNPGFleet(resp({ clusters: [c] })).rows[0].protection.walArchiving
+ expect(fact).toMatchObject({ text: 'Not archived: no destination configured', tone: 'neutral', source: 'Cluster spec' })
+ expect(fact.detail).toContain("WAL is not archived to recovery storage")
+ expect(fact.operatorCondition).toMatchObject({ type: 'ContinuousArchiving', status: 'True', lastTransitionTime: '2026-10-01T12:00:00Z' })
+ const snapshot = cluster('snapshot', 'db', { spec: { backup: { volumeSnapshot: {} } }, status: c.status })
+ expect(buildCNPGFleet(resp({ clusters: [snapshot] })).rows[0].protection.walArchiving.text).toBe('Not archived: no destination configured')
+})
+
+it('preserves declared archiver failures and distinguishes backup-only and opaque archiver plugins', () => {
+ const make = (plugins: any[], status = 'True') => buildCNPGFleet(resp({ clusters: [cluster('pg', 'db', { spec: { plugins }, status: { conditions: [{ type: 'ContinuousArchiving', status, message: 'archive report' }] } })] })).rows[0].protection.walArchiving
+ const barman = { name: 'barman-cloud.cloudnative-pg.io', isWALArchiver: true }
+ expect(make([barman], 'False')).toMatchObject({ text: 'Failing', tone: 'unhealthy', detail: 'archive report' })
+ expect(make([barman])).toMatchObject({ text: 'Not archived: no destination configured', tone: 'neutral' })
+ expect(make([{ ...barman, isWALArchiver: false, parameters: { barmanObjectName: 'store' } }])).toMatchObject({ text: 'Not archived: no destination configured', tone: 'neutral' })
+ expect(make([{ name: 'third-party-archive', isWALArchiver: true }])).toMatchObject({ text: 'CNPG reports archiving', detail: 'Archive plugin declared; its destination is not assessed here' })
+ expect(make([{ ...barman, enabled: false, parameters: { barmanObjectName: 'store' } }])).toMatchObject({ text: 'Not archived: no destination configured', tone: 'neutral' })
+})
+
function cluster(name: string, ns: string, extra: any = {}): any {
return {
apiVersion: G,
@@ -27,6 +48,10 @@ function pod(name: string, ns: string, clusterName: string, role: string, ready
}
}
+function serverProblem(reason: string, message: string, kind = 'Cluster', name = 'pg-a', severity: 'critical' | 'warning' = 'warning') {
+ return { id: `server:${reason}`, reason, message, kind, name, namespace: 'db', group: 'postgresql.cnpg.io', severity }
+}
+
function resp(objects: Partial>, over: Partial = {}): CNPGWorkspaceResponse {
const coverage: CNPGWorkspaceResponse['coverage'] = {}
for (const k of CNPG_WORKSPACE_KEYS) coverage[k] = { state: 'full' }
@@ -39,6 +64,7 @@ function resp(objects: Partial>, over: Partial {
)
const row = fleet.rows[0]
expect(row.replication.tone).toBe('unknown')
- expect(row.replication.text).toContain('2/2 replicas ready')
+ expect(row.replication.text).toContain('2/2 Pods ready')
expect(row.pods[0].role).toBe('primary')
})
@@ -81,6 +107,22 @@ describe('buildCNPGFleet', () => {
expect(fleet.rows[0].name).toBe('pg-b')
})
+ it('orders rows by worst problem, then problem count, then namespace/name', () => {
+ const issue = (id: string, severity: 'critical' | 'warning', name: string) => ({
+ id, severity, kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: 'db', name, reason: 'CNPGClusterUnhealthy', message: id,
+ })
+ const fleet = buildCNPGFleet(
+ resp(
+ { clusters: ['pg-a', 'pg-b', 'pg-c', 'pg-d', 'pg-e'].map((n) => cluster(n, 'db')) },
+ {
+ issues: [issue('w1', 'warning', 'pg-a'), issue('c1', 'critical', 'pg-b'), issue('w2', 'warning', 'pg-c'), issue('w3', 'warning', 'pg-c')],
+ audit: [{ checkId: 'cnpgNoDeclarativeBackup', severity: 'warning', kind: 'Cluster', namespace: 'db', name: 'pg-e', message: 'no ScheduledBackup' }],
+ },
+ ),
+ )
+ expect(fleet.rows.map((r) => r.name)).toEqual(['pg-b', 'pg-c', 'pg-a', 'pg-e', 'pg-d'])
+ })
+
it('treats the no-schedule audit finding as posture, not attention, and words it narrowly', () => {
const fleet = buildCNPGFleet(
resp({ clusters: [cluster('pg-a', 'db')] }, {
@@ -131,7 +173,7 @@ describe('buildCNPGFleet', () => {
})],
}),
)
- expect(fleet.rows[0].protection.lastSuccessfulBackup.text).toBe('None observed')
+ expect(fleet.rows[0].protection.lastSuccessfulBackup.text).toBe('No successful backup yet')
})
it('never marks restore validation healthy', () => {
@@ -144,15 +186,49 @@ describe('buildCNPGFleet', () => {
})
const fleet = buildCNPGFleet(resp({ clusters: [src, restored] }))
const a = fleet.rows.find((r) => r.name === 'pg-a')!
- expect(a.protection.restoreValidation.text).toBe('Restored into pg-a-restore')
+ expect(a.protection.restoreValidation.text).toBe('Archive restored into pg-a-restore')
expect(a.protection.restoreValidation.tone).toBe('neutral')
const r = fleet.rows.find((x) => x.name === 'pg-a-restore')!
expect(r.protection.restoreValidation.text).toBe('None recorded')
expect(r.protection.restoreValidation.tone).toBe('unknown')
})
- it('ends the recovery window at WAL archiving, not at the last base backup', () => {
+ it('matches a restored Cluster’s external source despite conflicting Backup parameters', () => {
+ const src = cluster('pg-a', 'db', { spec: { plugins: [{ name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'cluster-store', serverName: 'cluster-server' } }] } })
+ const conflicting = { apiVersion: 'postgresql.cnpg.io/v1', kind: 'Backup', metadata: { name: 'b', namespace: 'db' }, spec: { cluster: { name: 'pg-a' }, method: 'plugin', pluginConfiguration: { name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'backup-store', serverName: 'backup-server' } } }, status: { phase: 'completed' } }
+ const restored = cluster('pg-a-restore', 'db', { spec: { bootstrap: { recovery: { source: 'origin' } }, externalClusters: [{ name: 'origin', plugin: { name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'cluster-store', serverName: 'cluster-server' } } }] } })
+ const fact = buildCNPGFleet(resp({ clusters: [src, restored], backups: [conflicting] })).rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation
+ expect(fact).toMatchObject({ text: 'Archive restored into pg-a-restore', tone: 'neutral' })
+ const wrongSource = { ...restored, spec: { ...restored.spec, externalClusters: [{ name: 'origin', plugin: { name: 'barman-cloud.cloudnative-pg.io', parameters: conflicting.spec.pluginConfiguration.parameters } }] } }
+ expect(buildCNPGFleet(resp({ clusters: [src, wrongSource], backups: [conflicting] })).rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation).toMatchObject({ text: 'None recorded', tone: 'unknown' })
+ })
+
+ it('reads a recorded validation note on the restored cluster, still never healthy', () => {
const plugin = { plugins: [{ name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'store' } }] }
+ const src = cluster('pg-a', 'db', { metadata: { uid: 'uid-a' }, spec: plugin })
+ const restoredSpec = {
+ bootstrap: { recovery: { source: 'origin' } },
+ externalClusters: [{ name: 'origin', plugin: { name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'store', serverName: 'pg-a' } } }],
+ }
+ const note = (uid: string) =>
+ JSON.stringify({ version: 1, recordedAt: '2026-09-29T10:00:00Z', recordedBy: 'alice', checked: 'row counts on orders', source: { namespace: 'db', name: 'pg-a', uid, verified: true } })
+ const noted = cluster('pg-a-restore', 'db', { metadata: { annotations: { 'radar.skyhook.io/restore-validation': note('uid-a') } }, spec: restoredSpec })
+ const fleet = buildCNPGFleet(resp({ clusters: [src, noted] }))
+ const fact = fleet.rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation
+ expect(fact.text).toBe('Validation recorded')
+ expect(fact.tone).toBe('neutral')
+ expect(fact.at).toBe('2026-09-29T10:00:00Z')
+ expect(fact.source).toContain('by alice on pg-a-restore')
+
+ const otherUID = cluster('pg-a-restore', 'db', { metadata: { annotations: { 'radar.skyhook.io/restore-validation': note('uid-previous-incarnation') } }, spec: restoredSpec })
+ expect(buildCNPGFleet(resp({ clusters: [src, otherUID] })).rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation.text).toBe('None recorded')
+
+ const malformed = cluster('pg-a-restore', 'db', { metadata: { annotations: { 'radar.skyhook.io/restore-validation': '{not json' } }, spec: restoredSpec })
+ expect(buildCNPGFleet(resp({ clusters: [src, malformed] })).rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation.text).toBe('Archive restored into pg-a-restore')
+ })
+
+ it('ends the recovery window at WAL archiving, not at the last base backup', () => {
+ const plugin = { plugins: [{ name: 'barman-cloud.cloudnative-pg.io', isWALArchiver: true, parameters: { barmanObjectName: 'store' } }] }
const store = {
apiVersion: 'barmancloud.cnpg.io/v1', kind: 'ObjectStore', metadata: { name: 'store', namespace: 'db' },
status: { serverRecoveryWindow: { 'pg-a': { firstRecoverabilityPoint: '2026-09-01T00:00:00Z', lastSuccessfulBackupTime: '2026-09-02T00:00:00Z', lastFailedBackupTime: '2026-09-03T00:00:00Z' } } },
@@ -165,9 +241,17 @@ describe('buildCNPGFleet', () => {
expect(buildCNPGFleet(resp({ clusters: [failing], objectStores: [store] })).rows[0].protection.recoveryWindow.tone).toBe('degraded')
})
+ it('qualifies archiving as the operator report with its transition time', () => {
+ const hourAgo = new Date(Date.now() - 3_600_000).toISOString()
+ const resumed = cluster('pg-a', 'db', { spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } }, metadata: { creationTimestamp: '2026-01-01T00:00:00Z' }, status: { conditions: [{ type: 'ContinuousArchiving', status: 'True', lastTransitionTime: hourAgo }] } })
+ expect(buildCNPGFleet(resp({ clusters: [resumed] })).rows[0].protection.walArchiving).toMatchObject({ text: 'CNPG reports archiving', source: 'Cluster status · ContinuousArchiving=True', at: hourAgo, atMeaning: 'since' })
+ const old = cluster('pg-a', 'db', { spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } }, metadata: { creationTimestamp: '2026-01-01T00:00:00Z' }, status: { conditions: [{ type: 'ContinuousArchiving', status: 'True', lastTransitionTime: '2026-01-01T00:02:00Z' }] } })
+ expect(buildCNPGFleet(resp({ clusters: [old] })).rows[0].protection.walArchiving.at).toBe('2026-01-01T00:02:00Z')
+ })
+
it('reports WAL archiving from the condition and unknown when absent', () => {
- const failing = cluster('pg-a', 'db', { status: { conditions: [{ type: 'ContinuousArchiving', status: 'False', message: 'exit status 1' }] } })
- const silent = cluster('pg-b', 'db')
+ const failing = cluster('pg-a', 'db', { spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } }, status: { conditions: [{ type: 'ContinuousArchiving', status: 'False', message: 'exit status 1' }] } })
+ const silent = cluster('pg-b', 'db', { spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } } })
const fleet = buildCNPGFleet(resp({ clusters: [failing, silent] }))
expect(fleet.rows.find((r) => r.name === 'pg-a')!.protection.walArchiving.tone).toBe('unhealthy')
expect(fleet.rows.find((r) => r.name === 'pg-a')!.protection.summary.text).toBe('WAL archiving failing')
@@ -245,11 +329,45 @@ describe('buildCNPGFleet', () => {
expect(fleet.rows[0].categories.has('availability')).toBe(true)
})
+ it("names a Cluster's own Job Pod problem by what the Job is for, as the cause, never as an instance", () => {
+ const joinPod = {
+ apiVersion: 'v1',
+ kind: 'Pod',
+ metadata: { name: 'pg-a-2-join-x1', namespace: 'db', labels: { 'cnpg.io/cluster': 'pg-a', 'cnpg.io/jobRole': 'join', 'cnpg.io/instanceName': 'pg-a-2' } },
+ status: { phase: 'Pending' },
+ }
+ const base = resp({ clusters: [cluster('pg-a', 'db', { status: { readyInstances: 1 } })], pods: [pod('pg-a-1', 'db', 'pg-a', 'primary')] }, {
+ issues: [
+ { id: 'c1', severity: 'warning', category: 'operator_condition_failed', kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: 'db', name: 'pg-a', reason: 'Ready: ClusterIsNotReady', message: 'Cluster Is Not Ready' },
+ { id: 'j1', severity: 'critical', category: 'unschedulable', kind: 'Pod', namespace: 'db', name: 'pg-a-2-join-x1', reason: 'Unschedulable', message: '2 node(s) insufficient pods (0/2 nodes available)' },
+ ],
+ })
+ const row = buildCNPGFleet({ ...base, jobPods: [joinPod] }).rows[0]
+ expect(row.problems[0].title).toBe("New standby pg-a-2: Can't be scheduled")
+ expect(row.problems[0].severity).toBe('warning')
+ expect(row.problems[0].detail).toBe('both nodes have reached their Pod limit')
+ expect(row.problems[0].rawDetail).toBe('2 node(s) insufficient pods (0/2 nodes available)')
+ expect(row.problems[0].origin?.label).toBe('Kubernetes scheduler')
+ expect(row.pods.map((p) => p.name)).toEqual(['pg-a-1'])
+ expect(buildCNPGFleet(base).rows[0].problems.some((p) => p.subject.name === 'pg-a-2-join-x1')).toBe(false)
+ })
+
it('reads partial coverage by allowed namespaces and treats unnamed partial coverage as unknown', () => {
const named = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { backups: { state: 'partial', allowedNamespaces: ['db'] } } }))
- expect(named.rows[0].protection.lastSuccessfulBackup.text).toBe('None observed')
+ expect(named.rows[0].protection.lastSuccessfulBackup.text).toBe('No successful backup yet')
const unnamed = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { backups: { state: 'partial' } } }))
- expect(unnamed.rows[0].protection.lastSuccessfulBackup.text).toBe('No access to Backups')
+ expect(unnamed.rows[0].protection.lastSuccessfulBackup.text).toBe('Backups not read in db')
+ })
+
+ it('names the cause of a partial read only when the server names the namespace', () => {
+ const denied = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { backups: { state: 'partial', allowedNamespaces: ['other'], deniedNamespaces: ['db'] } } }))
+ expect(denied.rows[0].protection.lastSuccessfulBackup.text).toBe('No access to Backups')
+ const uncached = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { pods: { state: 'partial', allowedNamespaces: ['other'], uncachedNamespaces: ['db'] } } }))
+ expect(uncached.rows[0].replication.text).toBe('Radar does not cache Pods in db')
+ const none = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { pods: { state: 'uncached' } } }))
+ expect(none.rows[0].replication.text).toBe('Radar does not cache Pods')
+ const mixed = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { pods: { state: 'uncached', uncachedNamespaces: ['other'], deniedNamespaces: ['db'] } } }))
+ expect(mixed.rows[0].replication.text).toBe('No access to Pods')
})
it('says no access instead of "no replica pods" when Pods are unreadable', () => {
@@ -282,4 +400,237 @@ describe('buildCNPGFleet', () => {
expect(p.lastSuccessfulBackup.text).toBe('Loading…')
expect(p.recoveryWindow.text).toBe('Loading…')
})
+
+ it('shows the Pods’ ready count and raises an availability problem when status claims more', () => {
+ const fleet = buildCNPGFleet(
+ resp({
+ clusters: [cluster('pg-a', 'db')],
+ pods: [pod('pg-a-1', 'db', 'pg-a', 'primary', false), pod('pg-a-2', 'db', 'pg-a', 'replica'), pod('pg-a-3', 'db', 'pg-a', 'replica', false)],
+ }, { issues: [serverProblem('CNPGInstanceReadinessMismatch', '2 of 3 instance Pods not ready, including the primary', 'Cluster', 'pg-a', 'critical')] }),
+ )
+ const row = fleet.rows[0]
+ expect(row.podReadiness).toEqual({ ready: 1, total: 3 })
+ expect(row.readinessContradicted).toBe(true)
+ expect(cnpgReadyInstances(row)).toMatchObject({ text: '1/3', tone: 'degraded' })
+ expect(cnpgReadyInstances(row).note).toContain('CNPG status reports 3 ready')
+ const problem = row.problems[0]
+ expect(problem.severity).toBe('critical')
+ expect(problem.category).toBe('availability')
+ expect(problem.title).toBe('2 of 3 instance Pods not ready, including the primary')
+ expect(row.attention).toBe(true)
+ expect(fleet.attentionCount).toBe(1)
+ })
+
+ it('agrees with status when the Pods do, and never judges unreadable Pods', () => {
+ const agreeing = buildCNPGFleet(
+ resp({
+ clusters: [cluster('pg-a', 'db', { status: { readyInstances: 2 } })],
+ pods: [pod('pg-a-1', 'db', 'pg-a', 'primary'), pod('pg-a-2', 'db', 'pg-a', 'replica'), pod('pg-a-3', 'db', 'pg-a', 'replica', false)],
+ }),
+ ).rows[0]
+ expect(agreeing.readinessContradicted).toBeUndefined()
+ expect(cnpgReadyInstances(agreeing)).toEqual({ text: '2/3' })
+ const denied = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')] }, { coverage: { pods: { state: 'denied' } } })).rows[0]
+ expect(denied.podReadiness).toBeUndefined()
+ expect(denied.problems).toEqual([])
+ })
+
+ it('separates missing operator readiness from an observed Pod count', () => {
+ const instances = { ready: null, desired: 1 }
+ expect(cnpgReadyInstances({ instances, podReadiness: { ready: 0, total: 0 } })).toEqual({
+ text: 'Not reported by the operator', podText: '0 of 1 instance Pods ready',
+ })
+ expect(cnpgReadyInstances({ instances })).toEqual({ text: 'Not reported by the operator', podText: undefined })
+ expect(cnpgReadyInstances({ instances: { ready: null, desired: null }, podReadiness: { ready: 1, total: 1 } }).podText).toBe('1 instance Pod ready')
+ expect(cnpgReadyInstances({ instances: { ready: 0, desired: 1 } })).toEqual({ text: '0/1' })
+ })
+
+ it('names both primaries when status and the role label disagree', () => {
+ const row = buildCNPGFleet(
+ resp({
+ clusters: [cluster('pg-a', 'db', { status: { currentPrimary: 'pg-a-2', readyInstances: 1 } })],
+ pods: [pod('pg-a-1', 'db', 'pg-a', 'primary'), pod('pg-a-2', 'db', 'pg-a', 'replica', false)],
+ }, { issues: [serverProblem('CNPGPrimaryLabelMismatch', 'CNPG status names pg-a-2 primary; the Pod labelled primary is pg-a-1')] }),
+ ).rows[0]
+ expect(row.primaryConflict).toEqual({ status: 'pg-a-2', labelled: 'pg-a-1' })
+ expect(row.problems.map((p) => p.title)).toContain('CNPG status names pg-a-2 primary; the Pod labelled primary is pg-a-1')
+ })
+})
+
+describe('schedule fact', () => {
+ it('reads the schedule in words when the server supplied a reading, with the cron as its source', () => {
+ const sched = { metadata: { namespace: 'db', name: 'nightly' }, spec: { cluster: { name: 'pg-a' }, schedule: '0 0 2 * * *' } }
+ const withReading = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')], scheduledBackups: [sched] }, { scheduleReadings: { 'db/nightly': 'every day at 02:00 UTC' } }))
+ expect(withReading.rows[0].protection.schedule).toMatchObject({ text: 'Enabled · every day at 02:00 UTC · blocked: no backup destination', source: 'ScheduledBackup nightly · cron 0 0 2 * * *' })
+ const without = buildCNPGFleet(resp({ clusters: [cluster('pg-a', 'db')], scheduledBackups: [sched] }))
+ expect(without.rows[0].protection.schedule.text).toBe('Enabled · 0 0 2 * * * · blocked: no backup destination')
+ })
+})
+
+ describe('CNPG certainty', () => {
+ const archiving = cluster('pg', 'db', { spec: { backup: { barmanObjectStore: { destinationPath: 's3://backups' } } }, status: { conditions: [{ type: 'ContinuousArchiving', status: 'True' }] } })
+ it('degrades an archiving cluster with no completed Backup and names the read', () => {
+ const p = buildCNPGFleet(resp({ clusters: [archiving] })).rows[0].protection
+ expect(p.lastSuccessfulBackup).toMatchObject({ tone: 'degraded', text: 'No successful backup yet', source: 'Backups read in this namespace; none completed' })
+ })
+ it('keeps unread Backups unknown', () => {
+ const p = buildCNPGFleet(resp({ clusters: [archiving] }, { coverage: { backups: { state: 'denied' } } })).rows[0].protection
+ expect(p.lastSuccessfulBackup).toMatchObject({ tone: 'unknown', source: 'Backups not read' })
+ })
+ const declaration = (applied: boolean, observed = 2) => ({ apiVersion: G, kind: 'Database', metadata: { name: 'app', namespace: 'db', generation: 2 }, spec: { cluster: { name: 'pg' } }, status: { applied, observedGeneration: observed } })
+ it('shows a qualified lower bound inline when a declaration kind was not read', () => {
+ const d = buildCNPGFleet(resp({ clusters: [archiving], databases: [declaration(true)] }, { coverage: { publications: { state: 'denied' } } })).rows[0].declarations
+ expect(d.summary).toEqual({ text: '≥1 reconciled; Publications not read', tone: 'unknown' })
+ })
+ it('does not use an exact denominator when some declarations were not read', () => {
+ const d = buildCNPGFleet(resp({ clusters: [archiving], databases: [declaration(false)] }, { coverage: { subscriptions: { state: 'denied' } } })).rows[0].declarations
+ expect(d.summary.text).toBe('≥1 not reconciled; Subscriptions not read')
+ expect(d.summary.tone).toBe('degraded')
+ })
+ it('counts stale success and failure as pending in the fleet', () => {
+ for (const applied of [true, false]) {
+ const d = buildCNPGFleet(resp({ clusters: [archiving], databases: [declaration(applied, 1)] })).rows[0].declarations
+ expect(d).toMatchObject({ total: 1, pending: 1, failed: 0, summary: { text: '1 of 1 pending', tone: 'unknown' } })
+ }
+ })
+})
+
+it('uses plain WAL evidence without turning an operator success into an archive', () => {
+ const condition = { type: 'ContinuousArchiving', status: 'True', message: 'Continuous archiving is working', lastTransitionTime: '2026-10-01T12:00:00Z' }
+ const c = cluster('payments', 'db', { status: { conditions: [condition] } })
+ const p = buildCNPGFleet(resp({ clusters: [c] })).rows[0].protection
+ expect(p.walArchiving).toMatchObject({ text: 'Not archived: no destination configured', operatorCondition: condition })
+ expect(p.walArchiving.detail).toBe("WAL is not archived to recovery storage, so point-in-time recovery is unavailable. CloudNativePG still reports archiving as working because, with no destination, it accepts each WAL file without keeping it.")
+ expect(p.recoveryWindow.text).toBe('None: no backup destination')
+ c.spec.backup = { barmanObjectStore: { destinationPath: 's3://backups' } }
+ const configured = buildCNPGFleet(resp({ clusters: [c] })).rows[0].protection.walArchiving
+ expect(configured).toMatchObject({ tone: 'healthy', source: 'Cluster status · ContinuousArchiving=True' })
+ expect(configured.text).toBe('CNPG reports archiving')
+})
+it('uses the schedule method blocker in plain recovery summaries, including a mismatched destination', () => {
+ const c = cluster('payments', 'db')
+ const schedule = { metadata: { name: 'nightly', namespace: 'db' }, spec: { method: 'barmanObjectStore', cluster: { name: 'payments' }, schedule: '0 0 2 * * *' } }
+ const data = resp({ clusters: [c], scheduledBackups: [schedule] }, { scheduleReadings: { 'db/nightly': 'every day at 02:00 UTC' } })
+ const fact = () => buildCNPGFleet(data).rows[0].protection.schedule
+ expect(fact()).toMatchObject({ text: 'Enabled · every day at 02:00 UTC · blocked: no backup destination', tone: 'degraded' })
+ c.spec.backup = { volumeSnapshot: {} }
+ expect(fact()).toMatchObject({ text: 'Enabled · every day at 02:00 UTC · blocked: no barmanObjectStore destination', tone: 'degraded' })
+ schedule.spec.method = 'volumeSnapshot'
+ expect(fact()).toMatchObject({ text: 'Enabled · not run yet', tone: 'neutral' })
+ delete data.scheduleReadings
+ expect(fact().text).toBe('Enabled · not run yet')
+ expect(fact().source).toBe('ScheduledBackup nightly · cron 0 0 2 * * *')
+ delete (schedule.spec as any).schedule
+ expect(fact().text).toBe('Enabled · not run yet')
+ data.objects.backups = [{ apiVersion: 'postgresql.cnpg.io/v1', kind: 'Backup', metadata: { namespace: 'db', name: 'run', labels: { 'cnpg.io/scheduled-backup': 'nightly' } }, spec: { cluster: { name: 'payments' } }, status: { phase: 'completed' } }]
+ expect(fact().text).toBe('Enabled')
+})
+
+it('keeps the plain scheduler cause and the complete scheduler message as separate evidence', () => {
+ const pending = { apiVersion: 'v1', kind: 'Pod', metadata: { name: 'orders-2-join-x', namespace: 'db', labels: { 'cnpg.io/cluster': 'orders', 'cnpg.io/jobRole': 'join', 'cnpg.io/instanceName': 'orders-2' } } }
+ const raw = '0/2 nodes are available: 2 Too many pods. preemption: no victims.'
+ const data = { ...resp({ clusters: [cluster('orders', 'db')] }, { issues: [{ id: 'j', severity: 'critical' as const, category: 'unschedulable', kind: 'Pod', namespace: 'db', name: 'orders-2-join-x', reason: 'Unschedulable', message: raw }] }), jobPods: [pending] }
+ const problem = buildCNPGFleet(data).rows[0].problems.find((p) => p.subject.name === 'orders-2-join-x')!
+ expect(problem.instance).toBe('orders-2')
+ expect(problem.detail).toBe('both nodes have reached their Pod limit')
+ expect(problem.rawDetail).toBe(raw)
+})
+
+it.each(['False', 'Unknown', undefined])('does not invent operator archiving success for %s', (status) => {
+ const c = cluster('analytics', 'db', { status: { conditions: status ? [{ type: 'ContinuousArchiving', status }] : [] } })
+ expect(buildCNPGFleet(resp({ clusters: [c] })).rows[0].protection.walArchiving.detail).toBe('WAL is not archived to recovery storage, so point-in-time recovery is unavailable.')
+})
+it('counts zero ready instance Pods only when the Pod inventory was read', () => {
+ const c = cluster('analytics', 'db', { spec: { instances: 1 }, status: { readyInstances: undefined } })
+ expect(buildCNPGFleet(resp({ clusters: [c], pods: [] })).rows[0].podReadiness).toEqual({ ready: 0, total: 0 })
+ expect(buildCNPGFleet(resp({ clusters: [c] }, { coverage: { pods: { state: 'denied' } } })).rows[0].podReadiness).toBeUndefined()
+})
+
+it('counts an enabled destination-blocked schedule as a Cluster protection problem', () => {
+ const c = cluster('payments', 'db', { spec: { instances: 1 }, status: { readyInstances: 1, phase: 'Cluster in healthy state' } })
+ const schedule = { apiVersion: G, kind: 'ScheduledBackup', metadata: { name: 'payments-nightly', namespace: 'db' }, spec: { cluster: { name: 'payments' } } }
+ const fleet = buildCNPGFleet(resp({ clusters: [c], scheduledBackups: [schedule] }, { issues: [serverProblem('CNPGScheduleDestinationMissing', 'Backup schedule payments-nightly cannot run: no backup destination', 'ScheduledBackup', 'payments-nightly')] }))
+ const row = fleet.rows[0]
+ expect(row.problems).toContainEqual(expect.objectContaining({ title: 'Backup schedule payments-nightly cannot run: no backup destination', severity: 'warning', category: 'protection', subject: expect.objectContaining({ kind: 'ScheduledBackup', name: 'payments-nightly' }) }))
+ expect(row.attention).toBe(true)
+ expect(fleet.attentionCount).toBe(1)
+ expect(fleet.categoryCounts.protection).toBe(1)
+ expect(row.categories.has('protection')).toBe(true)
+ expect(cnpgDimensions({ row }).find((d) => d.id === 'protection')?.tone).toBe('degraded')
+ for (const schedules of [[], [{ ...schedule, spec: { ...schedule.spec, suspend: true } }]]) {
+ expect(buildCNPGFleet(resp({ clusters: [c], scheduledBackups: schedules })).rows[0].attention).toBe(false)
+ }
+ expect(buildCNPGFleet(resp({ clusters: [c], scheduledBackups: [schedule] }, { coverage: { scheduledBackups: { state: 'denied' } } })).attentionCount).toBe(0)
+})
+
+it('marks a schedule-method mismatch even when the Cluster has another working destination', () => {
+ const c = cluster('payments', 'db', { spec: { instances: 1, plugins: [{ name: 'barman-cloud.cloudnative-pg.io', isWALArchiver: true, parameters: { barmanObjectName: 'store' } }] }, status: { readyInstances: 1, conditions: [{ type: 'ContinuousArchiving', status: 'True' }] } })
+ const schedule = { metadata: { name: 'payments-nightly', namespace: 'db' }, spec: { cluster: { name: 'payments' }, method: 'barmanObjectStore' } }
+ const row = buildCNPGFleet(resp({ clusters: [c], scheduledBackups: [schedule] }, { issues: [serverProblem('CNPGScheduleDestinationMissing', 'Backup schedule payments-nightly cannot run: no barmanObjectStore destination', 'ScheduledBackup', 'payments-nightly')] })).rows[0]
+ expect(row.attention).toBe(true)
+ expect(cnpgDimensions({ row }).find((d) => d.id === 'protection')).toMatchObject({ tone: 'degraded', text: 'Backup schedule payments-nightly cannot run: no barmanObjectStore destination' })
+})
+
+
+it('uses server findings without classifying cached-object problems again', () => {
+ const c = cluster('pg-a', 'db', { status: { currentPrimary: 'pg-a-2' } })
+ const schedule = { metadata: { name: 'nightly', namespace: 'db' }, spec: { cluster: { name: 'pg-a' } } }
+ const row = buildCNPGFleet(resp({ clusters: [c], pods: [pod('pg-a-1', 'db', 'pg-a', 'primary', false)], scheduledBackups: [schedule] })).rows[0]
+ expect(row.readinessContradicted).toBe(true)
+ expect(row.primaryConflict).toEqual({ status: 'pg-a-2', labelled: 'pg-a-1' })
+ expect(row.problems).toEqual([])
+})
+
+
+it('excludes predecessor backup success, failures and restores from a recreated Cluster', () => {
+ const c = cluster('pg-a', 'db', { metadata: { uid: 'current', creationTimestamp: '2026-10-01T00:00:00Z' }, spec: { plugins: [{ name: 'barman-cloud.cloudnative-pg.io', isWALArchiver: true, parameters: { barmanObjectName: 'store' } }] } })
+ const old = { apiVersion: G, kind: 'Backup', metadata: { name: 'old', namespace: 'db', creationTimestamp: '2026-10-01T01:00:00Z' }, spec: { cluster: { name: 'pg-a' } }, status: { phase: 'completed', startedAt: '2026-10-01T01:00:00Z', stoppedAt: '2026-10-01T01:01:00Z', pluginMetadata: { clusterUID: 'previous' } } }
+ const failed = { ...old, metadata: { ...old.metadata, name: 'old-failed' }, status: { ...old.status, phase: 'failed' } }
+ const restored = cluster('restored', 'db', { spec: { bootstrap: { recovery: { backup: { name: 'old' } } } } })
+ const store = { apiVersion: 'barmancloud.cnpg.io/v1', kind: 'ObjectStore', metadata: { name: 'store', namespace: 'db' }, status: { serverRecoveryWindow: { 'pg-a': { lastSuccessfulBackupTime: '2026-09-30T23:00:00Z' } } } }
+ const data = resp({ clusters: [c, restored], backups: [old, failed], objectStores: [store] }, { issues: [serverProblem('CNPGBackupFailed', 'Previous backup failed', 'Backup', 'old-failed')] })
+ const row = buildCNPGFleet(data).rows.find((r) => r.name === 'pg-a')!
+ expect(row.protection.lastSuccessfulBackup.text).toBe('No successful backup yet')
+ expect(row.protection.restoreValidation.text).toBe('None recorded')
+ expect(row.problems.some((p) => p.subject.name === 'old-failed')).toBe(false)
+ const current = { ...old, status: { ...old.status, pluginMetadata: { clusterUID: 'current' } } }
+ expect(buildCNPGFleet({ ...data, objects: { ...data.objects, backups: [current, failed] } }).rows.find((r) => r.name === 'pg-a')!.protection.lastSuccessfulBackup).toMatchObject({ text: 'Completed', source: 'Backup old' })
+})
+
+it('does not attribute predecessor plugin restores or notes to a recreated source', () => {
+ const source = cluster('pg-a', 'db', { metadata: { uid: 'current', creationTimestamp: '2026-10-01T00:00:00Z' }, spec: { plugins: [{ name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'store', serverName: 'archive' } }] } })
+ const restored = cluster('restored', 'db', {
+ metadata: { uid: 'restored', creationTimestamp: '2026-10-02T00:00:00Z' },
+ spec: { bootstrap: { recovery: { source: 'origin' } }, externalClusters: [{ name: 'origin', plugin: { name: 'barman-cloud.cloudnative-pg.io', parameters: { barmanObjectName: 'store', serverName: 'archive' } } }] },
+ })
+ const fact = (target: any, backups: any[] = []) => buildCNPGFleet(resp({ clusters: [source, target], backups })).rows.find((r) => r.name === 'pg-a')!.protection.restoreValidation
+ const beforeSource = { ...restored, metadata: { ...restored.metadata, creationTimestamp: '2026-09-30T00:00:00Z' } }
+ expect(cnpgRecoveryMatchesCluster(beforeSource, source, [])).toBe(false)
+ expect(fact(beforeSource).text).toBe('None recorded')
+ const pinned = { ...restored, spec: { ...restored.spec, bootstrap: { recovery: { source: 'origin', recoveryTarget: { backupID: 'old-id' } } } } }
+ const oldBackup = { apiVersion: G, kind: 'Backup', metadata: { name: 'old', namespace: 'db' }, spec: { cluster: { name: 'pg-a' } }, status: { backupId: 'old-id', pluginMetadata: { clusterUID: 'previous' } } }
+ expect(cnpgRecoveryMatchesCluster(pinned, source, [oldBackup])).toBe(false)
+ expect(fact(pinned, [oldBackup]).text).toBe('None recorded')
+ const pitr = { ...restored, spec: { ...restored.spec, bootstrap: { recovery: { source: 'origin', recoveryTarget: { targetTime: '2026-09-30T23:00:00Z' } } } } }
+ expect(fact(pitr).text).toBe('None recorded')
+ const withNote = (sourceUID?: string, targetUID = 'restored') => ({ ...restored, metadata: { ...restored.metadata, annotations: { 'radar.skyhook.io/restore-validation': JSON.stringify({ recordedAt: '2026-10-02T02:00:00Z', checked: 'application data', source: { namespace: 'db', name: 'pg-a', uid: sourceUID, verified: !!sourceUID }, target: { namespace: 'db', name: 'restored', uid: targetUID, verified: true } }) } } })
+ expect(fact(withNote('previous')).text).toBe('None recorded')
+ expect(fact(withNote()).text).toBe('Archive restored into restored')
+ expect(fact(withNote('current')).text).toBe('Validation recorded')
+ const copied = withNote('current', 'previous-target')
+ expect(getCNPGRestoreValidation(copied)).toBeNull()
+ expect(fact(copied).text).toBe('Archive restored into restored')
+ expect(fact(restored).source).toContain('same archive currently configured')
+ expect(fact({ ...restored, status: { readyInstances: 0 } }).text).toBe('Archive recovery declared in restored')
+ const currentBackup = { ...oldBackup, status: { ...oldBackup.status, pluginMetadata: { clusterUID: 'current' } } }
+ expect(cnpgRecoveryMatchesCluster(pinned, source, [currentBackup])).toBe(true)
+})
+
+it('matches in-tree recovery by the actual archive and endpoint rather than server name alone', () => {
+ const archive = { destinationPath: 's3://bucket/prefix/', endpointURL: 'https://s3.example/', serverName: 'pg-a' }
+ const source = cluster('pg-a', 'db', { spec: { backup: { barmanObjectStore: archive } } })
+ const restored = cluster('restored', 'db', { spec: { bootstrap: { recovery: { source: 'origin' } }, externalClusters: [{ name: 'origin', barmanObjectStore: { ...archive, destinationPath: 's3://bucket/prefix' } }] } })
+ expect(cnpgRecoveryMatchesCluster(restored, source, [])).toBe(true)
+ const other = { ...source, spec: { backup: { barmanObjectStore: { ...archive, endpointURL: 'https://other.example' } } } }
+ expect(cnpgRecoveryMatchesCluster(restored, other, [])).toBe(false)
})
diff --git a/packages/k8s-ui/src/components/cnpg/workspace.ts b/packages/k8s-ui/src/components/cnpg/workspace.ts
index eb8b1a95c7..4f8ba2fe4f 100644
--- a/packages/k8s-ui/src/components/cnpg/workspace.ts
+++ b/packages/k8s-ui/src/components/cnpg/workspace.ts
@@ -1,13 +1,23 @@
+import { cnpgBackupDeclaration, cnpgBarmanPlugin } from '../../utils/cnpg-backup'
// CloudNativePG workspace model: pure derivations over the /api/cnpg/workspace
// payload. Every fact here is something the cluster actually reports; when it
// does not report something the value is "unknown", never zero or healthy.
-import type { HealthLevel } from '../resources/resource-utils'
+import { backupsForScheduledBackup, cnpgScheduleDestinationBlocker, cnpgBackupMatchesCluster, cnpgArchiveMatchesCluster, targetCluster } from './relations'
+import { cnpgRoleState } from './databaseRole'
+import { formatAge, summarizeSchedulerMessage, type HealthLevel } from '../resources/resource-utils'
+import { worseTone } from '../ui/status-tone'
+import { type Fact } from '../facts'
+import { type ProblemOrigin, type WorkspaceProblem } from '../problems'
+import type { ResourceRef } from '../../types/core'
+import { formatGrant, type Grant } from '../../utils/grant'
+import { issueReasonTitle, issueTitle } from '../issues/severity'
import {
CNPG_BARMAN_PLUGIN_NAME,
getCNPGClusterBackupConfig,
getCNPGClusterBarmanPlugin,
getCNPGClusterImageTag,
+ getCNPGClusterIsReplica,
getCNPGClusterStatus,
getCNPGObjectStoreRecoveryWindows,
isApiGroup,
@@ -21,6 +31,7 @@ export const CNPG_WORKSPACE_KEYS = [
'databases',
'publications',
'subscriptions',
+ 'databaseRoles',
'imageCatalogs',
'clusterImageCatalogs',
'objectStores',
@@ -29,12 +40,14 @@ export const CNPG_WORKSPACE_KEYS = [
export type CNPGWorkspaceKey = (typeof CNPG_WORKSPACE_KEYS)[number]
-export type CNPGCoverageState = 'full' | 'partial' | 'denied' | 'notInstalled' | 'syncing' | 'error'
+export type CNPGCoverageState = 'full' | 'partial' | 'denied' | 'notInstalled' | 'syncing' | 'uncached' | 'error'
export interface CNPGKindCoverage {
state: CNPGCoverageState
/** Denied namespaces, named only when the caller supplied the candidate list. */
deniedNamespaces?: string[]
+ /** Namespaces the caller may read but Radar's cache does not hold, named under the same rule. */
+ uncachedNamespaces?: string[]
/** For partial coverage: the namespaces that were read. */
allowedNamespaces?: string[]
}
@@ -73,6 +86,14 @@ export interface CNPGWorkspaceResponse {
issues: CNPGWorkspaceIssue[]
audit: CNPGAuditFinding[]
backupsOmitted: number
+ /** Each ScheduledBackup's schedule as the operator reads it, keyed "namespace/name"; absent when it cannot be parsed. */
+ scheduleReadings?: Record
+ /** The manager of each object that has one, keyed "Kind/namespace/name". */
+ managedBy?: Record
+ /** Pods of the Jobs a Cluster controls (initdb, join, restore…), never counted as instances. */
+ jobPods?: any[]
+ /** Where the caller's Jobs were read; a Job Pod is returned only there. Absent from a Radar that predates it. */
+ jobCoverage?: CNPGKindCoverage
}
export const CNPG_KIND_BY_KEY: Record = {
@@ -83,6 +104,7 @@ export const CNPG_KIND_BY_KEY: Record k.group !== '' && k.group === (group ?? '') && k.kind === kind)
}
-/** The value is observed, derived, or not available from the cluster. */
-export type CNPGFactTone = HealthLevel
-
-export interface CNPGFact {
- text: string
- tone: CNPGFactTone
- /** Where the value comes from, shown next to it so claims carry their source. */
- source?: string
- /** A timestamp the text refers to; the UI renders it as an age. */
- at?: string
-}
-
export type CNPGProblemCategory = 'availability' | 'protection' | 'declarations' | 'pooling'
export const CNPG_PROBLEM_CATEGORIES: { id: CNPGProblemCategory; label: string }[] = [
{ id: 'availability', label: 'Availability' },
- { id: 'protection', label: 'Protection' },
+ { id: 'protection', label: 'Backups' },
{ id: 'declarations', label: 'Declarations' },
{ id: 'pooling', label: 'Pooling' },
]
-export interface CNPGProblem {
- /** Stable identity for keys. */
- id: string
- severity: 'critical' | 'warning' | 'posture'
- category: CNPGProblemCategory
- title: string
- detail?: string
- /** The object the evidence is about (may be the Cluster or a child object). */
- subject: { kind: string; group: string; namespace: string; name: string }
- source: 'issue' | 'audit'
+/**
+ * A problem in the CloudNativePG workspace, categorised by the workspace's
+ * screens. `instance` names the instance a Cluster-level problem is about,
+ * e.g. the standby an HA slot is kept for.
+ */
+export type CNPGProblem = WorkspaceProblem & {
+ reason?: string
+ slot?: string
+ instance?: string
+ /** The cnpg.io/jobRole of the Cluster's own Job whose Pod this is about: a cause, not an instance's symptom. */
+ job?: string
+ /** The Cluster's own Ready condition: a roll-up of the other problems, shown after them. */
+ rollup?: boolean
}
export interface CNPGInstance {
@@ -135,15 +148,15 @@ export interface CNPGInstance {
}
export interface CNPGProtectionFacts {
- schedule: CNPGFact & { names: string[] }
- destination: CNPGFact & {
+ schedule: Fact & { names: string[] }
+ destination: Fact & {
method: 'plugin' | 'barmanObjectStore' | 'volumeSnapshot' | 'none'
objectStore?: string
}
- lastSuccessfulBackup: CNPGFact
- walArchiving: CNPGFact
- recoveryWindow: CNPGFact & { from?: string }
- restoreValidation: CNPGFact & { restoredInto?: { namespace: string; name: string } }
+ lastSuccessfulBackup: Fact
+ walArchiving: Fact & { state?: 'no_destination' | 'failing' | 'archiving' | 'unknown'; operatorCondition?: { type: string; status: string; message?: string; lastTransitionTime?: string } }
+ recoveryWindow: Fact & { from?: string }
+ restoreValidation: Fact & { restoredInto?: { namespace: string; name: string } }
}
export interface CNPGFleetRow {
@@ -153,15 +166,26 @@ export interface CNPGFleetRow {
cluster: any
controllerStatus: { text: string; level: HealthLevel }
instances: { ready: number | null; desired: number | null }
+ /**
+ * Ready instances counted from the instance Pods' Ready condition; absent
+ * when Pods are not readable here or any Pod's readiness is unknown.
+ */
+ podReadiness?: { ready: number; total: number }
+ /** status.readyInstances claims more ready instances than the Pods show: CNPG status is stale or lagging. */
+ readinessContradicted?: boolean
+ /** status.currentPrimary is not the Pod labelled primary; status may be stale, or a failover is under way. */
+ primaryConflict?: { status: string; labelled: string }
pods: CNPGInstance[]
replicaCluster: { source?: string } | null
hibernated: boolean
pgVersion: string | null
catalog: { kind: string; name: string } | null
- replication: CNPGFact
- protection: CNPGProtectionFacts & { summary: CNPGFact }
- declarations: { summary: CNPGFact; total: number; failed: number; pending: number }
+ replication: Fact
+ protection: CNPGProtectionFacts & { summary: Fact }
+ declarations: { summary: Fact; total: number; failed: number; pending: number }
poolers: string[]
+ /** The Pooler objects behind `poolers`, for their type and Service port. */
+ poolerObjects?: any[]
/** False when Poolers are not readable in this cluster's namespace, so an empty list means unknown. */
poolersKnown: boolean
problems: CNPGProblem[]
@@ -169,7 +193,12 @@ export interface CNPGFleetRow {
attention: boolean
categories: Set
/** GitOps owner recorded on the Cluster, when it carries the standard labels. */
- gitops: CNPGGitOpsSource | null
+ /** The GitOps or Helm object that manages the Cluster, as the server detected it. */
+ managedBy?: ResourceRef
+ /** Fullest measured volume, set by applyCNPGDisk; absent when no disk reading was requested. */
+ disk?: Fact
+ /** Growth of the fastest-growing volume, set by applyCNPGFleetMetrics when measured. */
+ diskGrowth?: Fact
}
export interface CNPGFleet {
@@ -185,8 +214,195 @@ const PROTECTION_ISSUE_REASONS = new Set([
'CNPGLastBackupFailed',
'CNPGBackupFailed',
'CNPGScheduledBackupMissed',
+ 'CNPGScheduledRunNoBackup',
+ 'CNPGScheduleDestinationMissing',
])
+// What an instance Pod's bare reason means, said about the Pod.
+const CNPG_POD_REASON_SENTENCES: Record = {
+ ReadinessProbeFailed: 'not ready (readiness probe failing)',
+ // The issue does not say whether the Pod is serving now (it may have come
+ // back within the settle window), so restarts are worded as past.
+ LivenessProbeFailed: 'restarted recently (liveness probe failing)',
+ CrashLoopBackOff: 'restarted recently (CrashLoopBackOff)',
+ HighRestartCount: 'restarted repeatedly',
+ OOMKilled: 'killed for running out of memory (OOMKilled)',
+ ImagePullBackOff: 'cannot pull its image (ImagePullBackOff)',
+ ErrImagePull: 'cannot pull its image (ErrImagePull)',
+}
+
+// Plain headlines for CNPG issues whose message carries the operator's own
+// condition text; that message becomes the detail beneath.
+// Reasons the Issues page already titles come from issueReasonTitle, so a
+// problem reads the same here and there; these are the rest.
+const CNPG_REASON_TITLES: Record = {
+ CNPGBackupFailed: 'Backup failed',
+ CNPGScheduledBackupMissed: 'A scheduled backup did not run',
+ CNPGCertificateExpiring: 'A certificate expires soon',
+ CNPGCertificateExpired: 'A certificate has expired',
+}
+
+/**
+ * A problem's headline and detail from an issue. Known CNPG reasons get a
+ * short plain title with the operator's message beneath; otherwise the
+ * message is the title, unless it is empty or only the reason token (e.g.
+ * "ReadinessProbeFailed"), which is turned into a sentence about the subject.
+ */
+export function cnpgIssueText(issue: Pick): { title: string; detail?: string } {
+ let message = issue.message?.trim() ?? ''
+ const cause = issue.cause?.trim() || undefined
+ // The run's time is the issue's first_seen, not part of the message.
+ if (issue.reason === 'CNPGScheduledRunNoBackup' && issue.first_seen && message) message = `${message} ${formatAge(issue.first_seen)} ago`
+ if (['CNPGScheduleDestinationMissing', 'CNPGInstanceReadinessMismatch', 'CNPGPrimaryLabelMismatch'].includes(issue.reason)) return { title: message, detail: cause }
+ const known = issueReasonTitle(issue.reason) ?? CNPG_REASON_TITLES[issue.reason]
+ if (known) return { title: known, detail: [stripTitlePrefix(message, known), cause].filter(Boolean).join(' ') || undefined }
+ if (message && message !== issue.reason && /\s/.test(message)) return { title: message, detail: cause }
+ const token = message || issue.reason
+ const sentence = CNPG_POD_REASON_SENTENCES[token]
+ if (sentence) return { title: `${issue.name} ${sentence}`, detail: cause }
+ const words = token.replace(/([a-z])([A-Z])/g, '$1 $2').toLowerCase()
+ return { title: `${issue.kind} ${issue.name}: ${words}`, detail: cause }
+}
+
+// "Backup failed: cannot proceed…" under the title "Backup failed" repeats it.
+function stripTitlePrefix(message: string, title: string): string {
+ if (!message.toLowerCase().startsWith(title.toLowerCase())) return message
+ const rest = message.slice(title.length).replace(/^[\s:;,.\-–—]+/, '')
+ return rest ? rest[0].toUpperCase() + rest.slice(1) : ''
+}
+
+/**
+ * Failed Backups of one cluster that failed for the same reason become one
+ * problem: "3 backups failed: ", about the latest of them, the others
+ * named in alsoAbout. Each Backup is otherwise its own issue, and a schedule
+ * failing every night would list the same sentence over and over.
+ */
+type IssueProblem = CNPGProblem
+
+export function cnpgCollapseBackupFailures(problems: IssueProblem[], backupTimes: Map = new Map()): IssueProblem[] {
+ const groups = new Map()
+ const out: IssueProblem[] = []
+ for (const p of problems) {
+ if (p.reason !== 'CNPGBackupFailed' || p.subject.kind !== 'Backup') {
+ out.push(p)
+ continue
+ }
+ const key = `${p.severity}\x00${p.detail ?? ''}`
+ groups.set(key, [...(groups.get(key) ?? []), p])
+ }
+ for (const list of groups.values()) {
+ if (list.length === 1) {
+ out.push(list[0])
+ continue
+ }
+ const at = (p: IssueProblem) => backupTimes.get(p.subject.name) ?? 0
+ const sorted = [...list].sort((a, b) => at(b) - at(a) || b.subject.name.localeCompare(a.subject.name))
+ const latest = sorted[0]
+ out.push({
+ ...latest,
+ id: `backups-failed:${latest.subject.namespace}:${latest.detail ?? ''}`,
+ title: latest.detail ? `${list.length} backups failed: ${latest.detail[0].toLowerCase()}${latest.detail.slice(1)}` : `${list.length} backups failed`,
+ detail: undefined,
+ alsoAbout: sorted.slice(1).map((p) => ({ kind: p.subject.kind, name: p.subject.name })),
+ })
+ }
+ return out
+}
+
+/**
+ * "Latest backup failed" restates a failed-Backup problem when that problem
+ * is about the cluster's newest Backup, so the duplicate is dropped and the
+ * count stays honest. With the newest Backup unknown, both stay.
+ */
+export function cnpgFoldLastBackupFailed(problems: IssueProblem[], newestBackup: string | undefined): CNPGProblem[] {
+ const covered =
+ !!newestBackup &&
+ problems.some(
+ (p) =>
+ p.reason === 'CNPGBackupFailed' &&
+ p.subject.kind === 'Backup' &&
+ (p.subject.name === newestBackup || p.alsoAbout?.some((o) => o.kind === 'Backup' && o.name === newestBackup)),
+ )
+ return problems
+ .filter((p) => !(covered && p.reason === 'CNPGLastBackupFailed'))
+}
+
+// When each of the cluster's Backups started (status.startedAt, else its
+// creation), by name: the issues about Backups carry no time of their own.
+function backupTimesOf(cluster: any, backups: any[]): Map {
+ const out = new Map()
+ for (const b of backups) {
+ if (!cnpgBackupMatchesCluster(b, cluster)) continue
+ out.set(b.metadata.name, Date.parse(b?.status?.startedAt ?? b?.metadata?.creationTimestamp ?? '') || 0)
+ }
+ return out
+}
+
+
+// Each entry names what the Go detector (internal/issues/source_cnpg*.go and
+// the Pod detector) actually reads. "Reported by CNPG" only where the operator
+// itself wrote the failure; a threshold or comparison Radar applies is a
+// "Radar check".
+const CNPG_CONDITION_ORIGINS: Record = {
+ CNPGLastBackupFailed: 'Cluster LastBackupSucceeded condition',
+ CNPGClusterTerminal: 'Cluster status.phase',
+ CNPGClusterUnrecoverable: 'Cluster status.phase',
+ CNPGClusterPluginFailure: 'Cluster status.phase',
+ CNPGClusterFailingOver: 'Cluster status.phase',
+ CNPGClusterWaitingForUser: 'Cluster status.phase',
+ CNPGDeclarativeNotApplied: 'status.applied and status.message',
+}
+
+const POD_ORIGINS: Record = {
+ ReadinessProbeFailed: { label: 'Pod readiness probe', detail: 'Kubelet probe-failure events and the Pod\'s Ready condition' },
+ LivenessProbeFailed: { label: 'Pod liveness probe', detail: 'Kubelet probe-failure events and container restarts' },
+ ReadinessProbeInvalid: { label: 'Radar check of the probe', detail: 'The readiness probe names a port the container does not declare' },
+ LivenessProbeInvalid: { label: 'Radar check of the probe', detail: 'The liveness probe names a port the container does not declare' },
+ HighRestartCount: { label: 'Radar check of restarts', detail: 'More than 3 restarts on a container that is still unhealthy' },
+ InitContainerStalled: { label: 'Radar check of init containers', detail: 'An init container has not finished' },
+ Unschedulable: { label: 'Kubernetes scheduler', detail: "The Pod's PodScheduled condition (reason Unschedulable) and the scheduler's message" },
+}
+
+/**
+ * Where an issue's evidence comes from, in user terms: what CloudNativePG
+ * reported, a Backup's or Pod's own status, or Radar's own check. A reason this
+ * does not know reads "Detected by Radar" rather than a guessed source.
+ */
+export function cnpgIssueOrigin(issue: Pick): ProblemOrigin {
+ const condition = CNPG_CONDITION_ORIGINS[issue.reason]
+ if (condition) return { label: 'Reported by CNPG', detail: condition }
+ switch (issue.reason) {
+ case 'CNPGWALArchivingFailing':
+ return issue.kind === 'Backup'
+ ? { label: 'Backup status', detail: 'Backup status.phase walArchivingFailing and status.error' }
+ : { label: 'Reported by CNPG', detail: 'Cluster ContinuousArchiving condition' }
+ case 'CNPGBackupFailed':
+ return { label: 'Backup status', detail: 'Backup status.phase and status.error' }
+ case 'CNPGClusterDegraded':
+ return { label: 'Radar check of ready instances', detail: 'spec.instances against status.readyInstances, unless the phase, hibernation or fencing explains it' }
+ case 'CNPGScheduleDestinationMissing':
+ return { label: 'Radar check of the backup destination', detail: 'ScheduledBackup method against its target Cluster spec' }
+ case 'CNPGInstanceReadinessMismatch':
+ return { label: 'Radar check of instance readiness', detail: 'Instance Pod Ready conditions against Cluster status.readyInstances' }
+ case 'CNPGPrimaryLabelMismatch':
+ return { label: 'Radar check of the primary', detail: 'Instance Pod role labels against Cluster status.currentPrimary' }
+ case 'CNPGScheduledRunNoBackup':
+ return { label: 'Radar check of the backup schedule', detail: 'The schedule, read as the operator does, against the cluster\'s newest successful backup' }
+ case 'CNPGScheduledBackupMissed':
+ return { label: 'Radar check of the backup schedule', detail: 'ScheduledBackup status.nextScheduleTime passed more than 10 minutes ago' }
+ case 'CNPGCertificateExpiring':
+ case 'CNPGCertificateExpired':
+ return { label: 'Certificate expiry (from Cluster status)', detail: 'Cluster status.certificates.expirations, compared with now' }
+ }
+ if (issue.kind === 'Pod') return POD_ORIGINS[issue.reason] ?? { label: 'Pod status' }
+ return { label: 'Detected by Radar' }
+}
+
+/** The headline alone; see cnpgIssueText. */
+export function cnpgIssueTitle(issue: Pick): string {
+ return cnpgIssueText(issue).title
+}
+
export function cnpgIssueCategory(issue: Pick): CNPGProblemCategory {
if (PROTECTION_ISSUE_REASONS.has(issue.reason)) return 'protection'
switch (issue.kind) {
@@ -197,6 +413,7 @@ export function cnpgIssueCategory(issue: Pick | null | undefined, obj: any): ResourceRef | undefined {
+ const kind = obj?.kind
+ const name = obj?.metadata?.name
+ if (!kind || !name) return undefined
+ return ws?.managedBy?.[`${kind}/${obj?.metadata?.namespace ?? ''}/${name}`]
}
function scheduleFact(
cluster: any,
schedules: any[],
cov: CNPGKindCoverage,
+ readings: Record = {},
+ backups: any[] = [],
): CNPGProtectionFacts['schedule'] {
const ns = cluster.metadata?.namespace
if (!coverageReadable(cov, ns)) {
- return { text: coverageUnavailableText(cov, 'ScheduledBackups'), tone: 'unknown', names: [] }
+ return { text: cnpgCoverageGap(cov, 'ScheduledBackups', ns), tone: 'unknown', names: [] }
}
const mine = schedules.filter((s) => s.metadata?.namespace === ns && specClusterName(s) === cluster.metadata?.name)
if (mine.length === 0) return { text: 'No declarative schedule', tone: 'neutral', names: [] }
@@ -307,15 +523,21 @@ function scheduleFact(
return { text: mine.length === 1 ? 'Schedule suspended' : 'All schedules suspended', tone: 'degraded', names }
}
const cron = active[0]?.spec?.schedule
+ const reading = readings[`${ns}/${active[0]?.metadata?.name}`]
+ const blockers = active.map((s) => cnpgScheduleDestinationBlocker(s, [cluster])).filter((b): b is string => !!b)
+ const cadence = active.length === 1 ? reading || cron : undefined
+ const run = active.some((s) => s.status?.lastScheduleTime || backupsForScheduledBackup(s, backups).length > 0)
return {
- text: active.length === 1 ? (cron ? `Scheduled · ${cron}` : 'Scheduled') : `${active.length} schedules`,
- tone: 'healthy',
+ text: [active.length === 1 ? 'Enabled' : `${active.length} enabled schedules`, blockers.length || run ? cadence : undefined, blockers.length ? `blocked: ${[...new Set(blockers)].map((blocker) => blocker[0].toLowerCase() + blocker.slice(1)).join(', ')}` : !run ? 'not run yet' : undefined].filter(Boolean).join(' · '),
+ tone: blockers.length ? 'degraded' : 'neutral',
names,
+ ...(active.length === 1 && cron ? { source: `ScheduledBackup ${active[0]?.metadata?.name} · cron ${cron}` } : {}),
}
}
function destinationFact(cluster: any): CNPGProtectionFacts['destination'] {
- const plugin = getCNPGClusterBarmanPlugin(cluster)
+ const declaration = cnpgBackupDeclaration(cluster)
+ const plugin = cnpgBarmanPlugin(declaration)
if (plugin?.barmanObjectName) {
return {
text: `ObjectStore ${plugin.barmanObjectName}`,
@@ -324,11 +546,10 @@ function destinationFact(cluster: any): CNPGProtectionFacts['destination'] {
objectStore: plugin.barmanObjectName,
}
}
- const cfg = getCNPGClusterBackupConfig(cluster)
- if (cfg.destinationPath) {
- return { text: cfg.destinationPath, tone: 'neutral', method: 'barmanObjectStore' }
+ if (declaration.inTreeDestination) {
+ return { text: declaration.inTreeDestination, tone: 'neutral', method: 'barmanObjectStore' }
}
- if (cluster?.spec?.backup?.volumeSnapshot) {
+ if (declaration.snapshotsConfigured) {
return { text: 'Volume snapshots', tone: 'neutral', method: 'volumeSnapshot' }
}
return { text: 'No destination configured', tone: 'neutral', method: 'none' }
@@ -359,14 +580,12 @@ function lastBackupFact(
storesUnreadable: CNPGKindCoverage | null,
): CNPGProtectionFacts['lastSuccessfulBackup'] {
const ns = cluster.metadata?.namespace
- const name = cluster.metadata?.name
const candidates: { at: string; source: string }[] = []
if (coverageReadable(backupsCov, ns)) {
const completed = backups
.filter(
(b) =>
- b.metadata?.namespace === ns &&
- specClusterName(b) === name &&
+ cnpgBackupMatchesCluster(b, cluster) &&
isApiGroup(b.apiVersion, 'postgresql.cnpg.io') &&
b.status?.phase === 'completed',
)
@@ -377,26 +596,105 @@ function lastBackupFact(
if (window?.lastSuccess) candidates.push({ at: window.lastSuccess, source: `ObjectStore ${window.store} status` })
const cfg = getCNPGClusterBackupConfig(cluster)
if (!cfg.plugin && cfg.lastSuccessfulBackup) candidates.push({ at: cfg.lastSuccessfulBackup, source: 'Cluster status' })
- if (candidates.length === 0) {
+ const createdAt = Date.parse(cluster.metadata?.creationTimestamp ?? '')
+ const current = candidates.filter((candidate) => !Number.isFinite(createdAt) || Date.parse(candidate.at) >= createdAt)
+ if (current.length === 0) {
if (!coverageReadable(backupsCov, ns)) {
- return { text: coverageUnavailableText(backupsCov, 'Backups'), tone: 'unknown' }
+ return { text: cnpgCoverageGap(backupsCov, 'Backups', ns), tone: 'unknown', source: 'Backups not read' }
}
- if (storesUnreadable) return { text: coverageUnavailableText(storesUnreadable, 'ObjectStores'), tone: 'unknown' }
- return { text: 'None observed', tone: 'unknown' }
+ if (storesUnreadable) return { text: cnpgCoverageGap(storesUnreadable, 'ObjectStores', ns), tone: 'unknown' }
+ return { text: 'No successful backup yet', tone: 'degraded', source: 'Backups read in this namespace; none completed' }
}
- const best = candidates.reduce((a, b) => (Date.parse(a.at) >= Date.parse(b.at) ? a : b))
+ const best = current.reduce((a, b) => (Date.parse(a.at) >= Date.parse(b.at) ? a : b))
return { text: 'Completed', tone: 'healthy', at: best.at, source: best.source }
}
-function walFact(cluster: any): CNPGFact {
+export const CNPG_NO_WAL_ARCHIVE_DESTINATION = 'Not archived: no destination configured'
+
+function walFact(cluster: any): CNPGProtectionFacts['walArchiving'] {
const conds = cluster?.status?.conditions
const c = Array.isArray(conds) ? conds.find((x: any) => x?.type === 'ContinuousArchiving') : null
- if (!c) return { text: 'Not reported', tone: 'unknown', source: 'Cluster status' }
- if (c.status === 'True') return { text: 'Archiving', tone: 'healthy', source: 'ContinuousArchiving condition' }
- if (c.status === 'False') {
- return { text: c.message ? `Failing · ${c.message}` : 'Failing', tone: 'unhealthy', source: 'ContinuousArchiving condition' }
+ const declaration = cnpgBackupDeclaration(cluster)
+ const plugin = cnpgBarmanPlugin(declaration)
+ const archivers = declaration.plugins.filter((p) => p.enabled && p.isWALArchiver)
+ const archiveConfigured = archivers.length > 0 || !!declaration.inTreeDestination
+ if (c?.status === 'False' && archiveConfigured) {
+ return { state: 'failing', text: 'Failing', tone: 'unhealthy', source: 'ContinuousArchiving condition', ...(c.lastTransitionTime ? { at: c.lastTransitionTime, atMeaning: 'since' as const } : {}), ...(c.message ? { detail: c.message } : {}) }
}
- return { text: 'Unknown', tone: 'unknown', source: 'ContinuousArchiving condition' }
+ const destinationKnown = !!(plugin?.isWALArchiver && plugin.barmanObjectName) || !!declaration.inTreeDestination
+ const customArchiver = archivers.some((p: any) => p.name !== plugin?.name)
+ if (!destinationKnown && !customArchiver) {
+ return {
+ state: 'no_destination', text: CNPG_NO_WAL_ARCHIVE_DESTINATION,
+ tone: 'neutral',
+ source: 'Cluster spec',
+ detail: "WAL is not archived to recovery storage, so point-in-time recovery is unavailable." + (c?.status === 'True' ? " CloudNativePG still reports archiving as working because, with no destination, it accepts each WAL file without keeping it." : ''),
+ ...(c ? { operatorCondition: { type: c.type, status: c.status, message: c.message, lastTransitionTime: c.lastTransitionTime } } : {}),
+ }
+ }
+ if (!c) return { state: 'unknown', text: 'Not reported', tone: 'unknown', source: 'Cluster status' }
+ if (c.status === 'True') {
+ return { state: 'archiving', text: 'CNPG reports archiving', tone: 'healthy', source: 'Cluster status · ContinuousArchiving=True', ...(customArchiver && !destinationKnown ? { detail: 'Archive plugin declared; its destination is not assessed here' } : {}), ...(c.lastTransitionTime ? { at: c.lastTransitionTime, atMeaning: 'since' as const } : {}) }
+ }
+ return { state: 'unknown', text: 'Unknown', tone: 'unknown', source: 'ContinuousArchiving condition' }
+}
+
+export const CNPG_RESTORE_VALIDATION_ANNOTATION = 'radar.skyhook.io/restore-validation'
+
+export interface CNPGRestoreValidationNote {
+ recordedAt: string
+ recordedBy?: string
+ checked: string
+ targetTime?: string
+ source?: { namespace: string; name: string; uid?: string; verified: boolean }
+ target?: { namespace: string; name: string; uid?: string; verified: boolean }
+}
+
+/** The validation note recorded on a restored Cluster, or null when absent or malformed. */
+export function getCNPGRestoreValidation(cluster: any): CNPGRestoreValidationNote | null {
+ const raw = cluster?.metadata?.annotations?.[CNPG_RESTORE_VALIDATION_ANNOTATION]
+ if (typeof raw !== 'string' || !raw.trim()) return null
+ try {
+ const v = JSON.parse(raw)
+ if (typeof v?.recordedAt !== 'string' || typeof v?.checked !== 'string' || !v.checked) return null
+ if (v.target?.uid && v.target.uid !== cluster.metadata?.uid) return null
+ return v as CNPGRestoreValidationNote
+ } catch {
+ return null
+ }
+}
+
+function noteIsAbout(note: CNPGRestoreValidationNote, cluster: any): boolean {
+ return !!note.source?.uid && note.source.uid === cluster.metadata?.uid
+}
+
+/** A recovery declaration compatible with this live source, excluding known predecessor evidence. */
+export function cnpgRecoveryMatchesCluster(restored: any, cluster: any, backups: any[]): boolean {
+ if (restored === cluster || restored.metadata?.namespace !== cluster.metadata?.namespace) return false
+ const recovery = restored.spec?.bootstrap?.recovery
+ if (!recovery) return false
+ const created = Date.parse(cluster.metadata?.creationTimestamp ?? '')
+ const restoredAt = Date.parse(restored.metadata?.creationTimestamp ?? '')
+ const targetAt = Date.parse(recovery.recoveryTarget?.targetTime ?? '')
+ if (Number.isFinite(created) && (restoredAt < created || targetAt < created)) return false
+ const note = getCNPGRestoreValidation(restored)
+ if (note?.source?.uid && !noteIsAbout(note, cluster)) return false
+ const ns = cluster.metadata?.namespace
+ if (recovery.backup?.name) {
+ const backup = backups.find((b) => b.metadata?.namespace === ns && b.metadata?.name === recovery.backup.name)
+ return cnpgBackupMatchesCluster(backup, cluster)
+ }
+ const source = (restored.spec?.externalClusters ?? []).find((e: any) => e?.name === recovery.source)
+ if (!source) return false
+ const params = source.plugin?.name === CNPG_BARMAN_PLUGIN_NAME ? source.plugin.parameters : undefined
+ const external = source.barmanObjectStore
+ const archiveMatch = params?.barmanObjectName
+ ? cnpgArchiveMatchesCluster(cluster, { kind: 'objectStore', objectStore: params.barmanObjectName, serverName: params.serverName || recovery.source })
+ : external && cnpgArchiveMatchesCluster(cluster, { kind: 'inTree', barmanObjectStore: external, serverName: external.serverName || recovery.source })
+ if (!archiveMatch) return false
+ const id = recovery.recoveryTarget?.backupID
+ const pinned = id ? backups.filter((b) => b.metadata?.namespace === ns && b.spec?.cluster?.name === cluster.metadata?.name && b.status?.backupId === id) : []
+ return pinned.length === 0 || pinned.some((b) => cnpgBackupMatchesCluster(b, cluster))
}
function restoreValidationFact(
@@ -405,26 +703,24 @@ function restoreValidationFact(
backups: any[],
backupsReadable: boolean,
): CNPGProtectionFacts['restoreValidation'] {
- const plugin = getCNPGClusterBarmanPlugin(cluster)
- const server = plugin?.serverName || cluster.metadata?.name
- const store = plugin?.barmanObjectName
const ns = cluster.metadata?.namespace
- const name = cluster.metadata?.name
- const restored = allClusters.find((c) => {
- if (c === cluster || c.metadata?.namespace !== ns) return false
- const recovery = c.spec?.bootstrap?.recovery
- if (!recovery) return false
- const sourceName = recovery.source
- if (sourceName) {
- const ext = (c.spec?.externalClusters ?? []).find((e: any) => e?.name === sourceName)
- const params = ext?.plugin?.name === CNPG_BARMAN_PLUGIN_NAME ? ext.plugin.parameters : undefined
- if (store && params?.barmanObjectName === store && (params?.serverName || sourceName) === server) return true
+ const restoredFromThis = allClusters.filter((c) => cnpgRecoveryMatchesCluster(c, cluster, backups))
+ const noted = restoredFromThis
+ .map((c) => ({ c, note: getCNPGRestoreValidation(c) }))
+ .filter((x): x is { c: any; note: CNPGRestoreValidationNote } => !!x.note && noteIsAbout(x.note, cluster))
+ .sort((a, b) => Date.parse(b.note.recordedAt) - Date.parse(a.note.recordedAt))[0]
+ if (noted) {
+ const rname = noted.c.metadata?.name
+ const by = noted.note.recordedBy ? `by ${noted.note.recordedBy}` : 'by a user Radar could not identify'
+ return {
+ text: 'Validation recorded',
+ tone: 'neutral',
+ at: noted.note.recordedAt,
+ source: `Recorded ${by} on ${rname}${noted.note.targetTime ? ` (target ${noted.note.targetTime})` : ''}: ${noted.note.checked.length > 140 ? `${noted.note.checked.slice(0, 140)}…` : noted.note.checked}. A person's note, not a check Radar ran.`,
+ restoredInto: { namespace: noted.c.metadata?.namespace, name: rname },
}
- const backupName = recovery.backup?.name
- if (!backupName) return false
- const backup = backups.find((b) => b.metadata?.namespace === ns && b.metadata?.name === backupName)
- return specClusterName(backup) === name
- })
+ }
+ const restored = restoredFromThis[0]
if (!restored) {
// A recovery by Backup name is only attributable when that Backup could be read.
const unresolved = !backupsReadable && allClusters.some((c) => c !== cluster && c.metadata?.namespace === ns && c.spec?.bootstrap?.recovery?.backup?.name)
@@ -432,24 +728,26 @@ function restoreValidationFact(
return { text: 'None recorded', tone: 'unknown', source: 'Kubernetes does not record restore tests' }
}
const rname = restored.metadata?.name
+ const fromArchive = !!restored.spec?.bootstrap?.recovery?.source
+ const sourceDescription = fromArchive ? 'the same archive currently configured for this Cluster' : "this cluster's backups"
const ready = typeof restored.status?.readyInstances === 'number' && restored.status.readyInstances > 0
if (!ready) {
return {
- text: `Recovery declared in ${rname}`,
+ text: `${fromArchive ? 'Archive recovery' : 'Recovery'} declared in ${rname}`,
tone: 'unknown',
- source: `Cluster ${rname} bootstraps from this cluster's backups but has no ready instance yet`,
+ source: `Cluster ${rname} bootstraps from ${sourceDescription} but has no ready instance yet`,
restoredInto: { namespace: restored.metadata?.namespace, name: rname },
}
}
return {
- text: `Restored into ${rname}`,
+ text: `${fromArchive ? 'Archive restored' : 'Restored'} into ${rname}`,
tone: 'neutral',
- source: `Cluster ${rname} bootstrapped from this cluster's backups and has ready instances · created ${restored.metadata?.creationTimestamp ?? 'unknown'}. This proves one recovery, not that today's backups restore.`,
+ source: `Cluster ${rname} bootstrapped from ${sourceDescription} and has ready instances · created ${restored.metadata?.creationTimestamp ?? 'unknown'}. No validation note is recorded on it; this proves one recovery, not that today's backups restore.`,
restoredInto: { namespace: restored.metadata?.namespace, name: rname },
}
}
-function protectionSummary(p: CNPGProtectionFacts): CNPGFact {
+function protectionSummary(p: CNPGProtectionFacts): Fact {
if (p.walArchiving.tone === 'unhealthy') return { text: 'WAL archiving failing', tone: 'unhealthy' }
if (p.destination.method === 'none' && p.schedule.names.length === 0 && p.schedule.tone !== 'unknown') {
return { text: 'No backup destination or schedule', tone: 'neutral' }
@@ -475,21 +773,52 @@ function pgVersion(cluster: any): string | null {
return typeof major === 'number' ? String(major) : null
}
-function replicationFact(cluster: any, pods: CNPGInstance[], hibernated: boolean, podsCov: CNPGKindCoverage): CNPGFact {
+function replicationFact(cluster: any, pods: CNPGInstance[], hibernated: boolean, podsCov: CNPGKindCoverage): Fact {
if (hibernated) return { text: 'Hibernated', tone: 'neutral' }
const desired = cluster?.spec?.instances
if (desired === 1) return { text: 'Single instance', tone: 'neutral' }
if (!coverageReadable(podsCov, cluster?.metadata?.namespace)) {
- return { text: coverageUnavailableText(podsCov, 'Pods'), tone: 'unknown' }
+ return { text: cnpgCoverageGap(podsCov, 'Pods', cluster?.metadata?.namespace), tone: 'unknown' }
}
const replicas = pods.filter((p) => p.role === 'replica')
const readyReplicas = replicas.filter((p) => p.ready === true).length
if (replicas.length === 0) return { text: 'No replica pods observed', tone: 'unknown' }
return {
- text: `${readyReplicas}/${replicas.length} replicas ready · lag unknown`,
+ text: `${readyReplicas}/${replicas.length} Pods ready · lag unknown`,
tone: 'unknown',
- source: 'Pod readiness does not show whether a replica is streaming',
+ source: CNPG_LAG_UNMEASURED_SOURCE,
+ }
+}
+
+const CNPG_JOB_PURPOSE: Record = {
+ initdb: 'First instance',
+ join: 'New standby',
+ 'full-recovery': 'Restore into',
+ 'snapshot-recovery': 'Restore into',
+ pgbasebackup: 'Clone into',
+ import: 'Import into',
+ 'major-upgrade': 'Major upgrade of',
+}
+
+/** What a Cluster's own Job is for, naming the instance it builds ("New standby pg-2"). */
+export function cnpgJobPurpose(role: string | undefined, instance: string | undefined, pod: string): string {
+ const purpose = role ? CNPG_JOB_PURPOSE[role] : undefined
+ if (purpose && instance) return `${purpose} ${instance}`
+ return `${role ? `${role} ` : ''}Job Pod ${pod}`
+}
+
+interface CNPGJobPod {
+ role?: string
+ instance?: string
+}
+
+function jobPodIndex(resp: CNPGWorkspaceResponse): Map {
+ const idx = new Map()
+ for (const p of resp.jobPods ?? []) {
+ const labels = p?.metadata?.labels ?? {}
+ idx.set(`${p.metadata?.namespace}/${p.metadata?.name}`, { role: labels['cnpg.io/jobRole'], instance: labels['cnpg.io/instanceName'] })
}
+ return idx
}
function problemsFor(
@@ -497,25 +826,42 @@ function problemsFor(
issues: CNPGWorkspaceIssue[],
audit: CNPGAuditFinding[],
children: Map,
+ backupTimes: Map = new Map(),
+ jobPods: Map = new Map(),
): CNPGProblem[] {
const ns = cluster.metadata?.namespace
const name = cluster.metadata?.name
const out: CNPGProblem[] = []
+ const fromIssues: IssueProblem[] = []
for (const issue of issues) {
if ((issue.namespace ?? '') !== ns) continue
const isSelf = issue.kind === 'Cluster' && issue.name === name
const owner = children.get(`${issue.kind}/${ns}/${issue.name}`)
if (!isSelf && owner !== name) continue
- out.push({
+ const job = issue.kind === 'Pod' ? jobPods.get(`${ns}/${issue.name}`) : undefined
+ const text = cnpgIssueText(issue)
+ fromIssues.push({
id: `${issue.id}:${issue.kind}/${issue.name}`,
- severity: issue.severity,
+ // A standby that cannot join costs redundancy, not service: the primary keeps serving.
+ severity: job?.role === 'join' ? 'warning' : issue.severity,
category: cnpgIssueCategory(issue),
- title: issue.message || issue.reason,
- detail: issue.cause || undefined,
+ ...(job
+ ? {
+ title: `${cnpgJobPurpose(job.role, job.instance, issue.name)}: ${issue.category ? issueTitle({ category: issue.category, reason: issue.reason }) : text.title}`,
+ detail: [issue.message?.trim(), issue.cause?.trim()].filter(Boolean).join(' ') || undefined,
+ job: job.role ?? 'job',
+ instance: job.instance,
+ }
+ : text),
subject: { kind: issue.kind, group: issue.group ?? '', namespace: ns, name: issue.name },
source: 'issue',
+ origin: cnpgIssueOrigin(issue),
+ reason: issue.reason,
+ ...(isSelf && issue.reason.startsWith('Ready:') ? { rollup: true } : {}),
})
}
+ const newestBackup = [...backupTimes.entries()].sort((a, b) => b[1] - a[1])[0]?.[0]
+ out.push(...cnpgFoldLastBackupFailed(cnpgCollapseBackupFailures(fromIssues, backupTimes), newestBackup))
for (const f of audit) {
if (f.kind !== 'Cluster' || f.name !== name || (f.namespace ?? '') !== ns) continue
out.push({
@@ -528,8 +874,62 @@ function problemsFor(
source: 'audit',
})
}
- const rank = { critical: 0, warning: 1, posture: 2 } as const
- return out.sort((a, b) => rank[a.severity] - rank[b.severity] || a.title.localeCompare(b.title))
+ return out.sort(cnpgCompareProblems)
+}
+
+/**
+ * The ready count to show for a cluster: CNPG's status count, or the Pods'
+ * own count when the status claims more than the Pods show.
+ */
+export function cnpgReadyInstances(row: Pick): { text: string; podText?: string; tone?: HealthLevel; note?: string } {
+ if (row.instances.ready === null) {
+ return {
+ text: 'Not reported by the operator',
+ podText: row.podReadiness
+ ? row.instances.desired !== null
+ ? `${row.podReadiness.ready} of ${row.instances.desired} instance Pods ready`
+ : `${row.podReadiness.ready} instance Pod${row.podReadiness.ready === 1 ? '' : 's'} ready`
+ : undefined,
+ }
+ }
+ const desired = row.instances.desired ?? '–'
+ if (row.readinessContradicted && row.podReadiness) {
+ return {
+ text: `${row.podReadiness.ready}/${desired}`,
+ tone: row.podReadiness.ready === 0 ? 'unhealthy' : 'degraded',
+ note: `Counted from the instance Pods' Ready condition. CNPG status reports ${row.instances.ready} ready, so the status may be stale.`,
+ }
+ }
+ return { text: `${row.instances.ready ?? '–'}/${desired}` }
+}
+
+const CNPG_SEVERITY_RANK = { critical: 0, warning: 1, posture: 2 } as const
+
+/**
+ * The order problems are shown in, everywhere: most severe first, then a
+ * cause before its symptoms. One instance Pod's state (a failing probe, a
+ * crash loop) is usually the symptom of a cluster-level problem (archiving,
+ * backups, replication, reconciliation, declarations) or of a Job the Cluster
+ * runs (an initdb or join Pod that cannot start), so at equal severity those
+ * come first; the Cluster's own Ready condition sums them all up and comes last.
+ */
+export function cnpgCompareProblems(a: CNPGProblem, b: CNPGProblem): number {
+ const symptom = (p: CNPGProblem) => (p.rollup ? 2 : p.subject.kind === 'Pod' && p.subject.group === '' && !p.job ? 1 : 0)
+ return CNPG_SEVERITY_RANK[a.severity] - CNPG_SEVERITY_RANK[b.severity] || symptom(a) - symptom(b) || a.title.localeCompare(b.title)
+}
+
+function sortProblems(list: CNPGProblem[]): CNPGProblem[] {
+ return list.sort(cnpgCompareProblems)
+}
+
+function primaryConflictOf(cluster: any, pods: any[]): CNPGFleetRow['primaryConflict'] {
+ const status = cluster?.status?.currentPrimary
+ if (!status) return undefined
+ const labelled = pods
+ .filter((p) => (p?.metadata?.labels?.['cnpg.io/instanceRole'] ?? p?.metadata?.labels?.role) === 'primary')
+ .map((p) => p.metadata?.name as string)
+ if (labelled.length === 0 || labelled.includes(status)) return undefined
+ return { status, labelled: labelled.sort()[0] }
}
/** Index "Kind/ns/name" → owning cluster name, from each child's spec.cluster.name. */
@@ -537,6 +937,7 @@ function childIndex(resp: CNPGWorkspaceResponse): Map {
const idx = new Map()
const add = (kind: string, list: any[] | undefined) => {
for (const o of list ?? []) {
+ if (kind === 'Backup' && !targetCluster(o, resp.objects.clusters ?? [])) continue
const c = specClusterName(o)
if (c) idx.set(`${kind}/${o.metadata?.namespace}/${o.metadata?.name}`, c)
}
@@ -547,7 +948,8 @@ function childIndex(resp: CNPGWorkspaceResponse): Map {
add('Database', resp.objects.databases)
add('Publication', resp.objects.publications)
add('Subscription', resp.objects.subscriptions)
- for (const p of resp.objects.pods ?? []) {
+ add('DatabaseRole', resp.objects.databaseRoles)
+ for (const p of [...(resp.objects.pods ?? []), ...(resp.jobPods ?? [])]) {
const c = p?.metadata?.labels?.['cnpg.io/cluster']
if (c) idx.set(`Pod/${p.metadata?.namespace}/${p.metadata?.name}`, c)
}
@@ -561,21 +963,24 @@ function declarationsFor(cluster: any, resp: CNPGWorkspaceResponse): CNPGFleetRo
['databases', resp.objects.databases ?? []],
['publications', resp.objects.publications ?? []],
['subscriptions', resp.objects.subscriptions ?? []],
+ ['databaseRoles', resp.objects.databaseRoles ?? []],
]
let total = 0
let failed = 0
let pending = 0
- let unreadable = false
+ const unread: string[] = []
+ const labels = { databases: 'Databases', publications: 'Publications', subscriptions: 'Subscriptions', databaseRoles: 'DatabaseRoles' }
for (const [k, list] of lists) {
if (!coverageReadable(coverageOf(resp, k), ns)) {
- if (coverageOf(resp, k).state !== 'notInstalled') unreadable = true
+ if (coverageOf(resp, k).state !== 'notInstalled') unread.push(labels[k as keyof typeof labels])
continue
}
for (const o of list) {
if (o.metadata?.namespace !== ns || specClusterName(o) !== name) continue
total++
- if (o.status?.applied === false) failed++
- else if (o.status?.applied !== true) pending++
+ const state = cnpgRoleState(o)
+ if (state === 'failed') failed++
+ else if (state === 'pending') pending++
}
}
const roleStatus = cluster?.status?.managedRolesStatus
@@ -588,9 +993,10 @@ function declarationsFor(cluster: any, resp: CNPGWorkspaceResponse): CNPGFleetRo
if (failedRoles.has(r.name)) failed++
else if (!reconciledRoles.has(r.name)) pending++
}
- let summary: CNPGFact
+ const unreadable = unread.length > 0
+ let summary: Fact
if (total === 0) {
- summary = unreadable ? { text: 'No access to some declarations', tone: 'unknown' } : { text: 'None declared', tone: 'neutral' }
+ summary = { text: 'None declared', tone: 'neutral' }
} else if (failed > 0) {
summary = { text: `${failed} of ${total} not reconciled`, tone: 'degraded' }
} else if (pending > 0) {
@@ -598,7 +1004,10 @@ function declarationsFor(cluster: any, resp: CNPGWorkspaceResponse): CNPGFleetRo
} else {
summary = { text: `${total} reconciled`, tone: 'healthy' }
}
- if (unreadable && total > 0) summary = { ...summary, source: 'Some declaration kinds are not readable' }
+ if (unreadable) {
+ const count = total === 0 ? 'Reconciliation unknown' : failed > 0 ? `≥${failed} not reconciled` : pending > 0 ? `≥${pending} pending` : `≥${total} reconciled`
+ summary = { text: `${count}; ${unread.join(', ')} not read`, tone: failed > 0 ? 'degraded' : 'unknown' }
+ }
return { summary, total, failed, pending }
}
@@ -607,6 +1016,7 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
const pods = resp.objects.pods ?? []
const stores = resp.objects.objectStores ?? []
const children = childIndex(resp)
+ const jobs = jobPodIndex(resp)
const poolers = resp.objects.poolers ?? []
const rows: CNPGFleetRow[] = clusters.map((cluster) => {
@@ -631,7 +1041,7 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
const storesCov = coverageOf(resp, 'objectStores')
const storesUnreadable = !!getCNPGClusterBarmanPlugin(cluster)?.barmanObjectName && !coverageReadable(storesCov, ns)
const protection: CNPGProtectionFacts = {
- schedule: scheduleFact(cluster, resp.objects.scheduledBackups ?? [], coverageOf(resp, 'scheduledBackups')),
+ schedule: scheduleFact(cluster, resp.objects.scheduledBackups ?? [], coverageOf(resp, 'scheduledBackups'), resp.scheduleReadings, resp.objects.backups ?? []),
destination: destinationFact(cluster),
lastSuccessfulBackup: lastBackupFact(cluster, resp.objects.backups ?? [], coverageOf(resp, 'backups'), window, storesUnreadable ? storesCov : null),
walArchiving: wal,
@@ -645,15 +1055,26 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
source: `ObjectStore ${window.store} status (earliest point)`,
}
: storesUnreadable
- ? { text: coverageUnavailableText(storesCov, 'ObjectStores'), tone: 'unknown' }
- : { text: 'Not reported', tone: 'unknown' },
+ ? { text: cnpgCoverageGap(storesCov, 'ObjectStores', ns), tone: 'unknown' }
+ : destinationFact(cluster).method === 'none' ? { text: 'None: no backup destination', tone: 'neutral' } : { text: 'Not reported', tone: 'unknown' },
restoreValidation: restoreValidationFact(cluster, clusters, resp.objects.backups ?? [], coverageReadable(coverageOf(resp, 'backups'), ns)),
}
- const problems = problemsFor(cluster, resp.issues ?? [], resp.audit ?? [], children)
+ const podsReadable = coverageReadable(coverageOf(resp, 'pods'), ns)
+ const podReadiness =
+ podsReadable && resp.objects.pods !== undefined && instancePods.every((p) => p.ready !== null)
+ ? { ready: instancePods.filter((p) => p.ready).length, total: instancePods.length }
+ : undefined
+ const readinessContradicted = !hibernated && !!podReadiness && readyInstances !== null && podReadiness.ready < readyInstances
+ const primaryConflict = podsReadable ? primaryConflictOf(cluster, pods.filter((p) => p.metadata?.namespace === ns && p.metadata?.labels?.['cnpg.io/cluster'] === name)) : undefined
+ let problems = sortProblems([
+ ...problemsFor(cluster, resp.issues ?? [], resp.audit ?? [], children, backupTimesOf(cluster, resp.objects.backups ?? []), jobs),
+
+ ])
+ problems = problems.map((p) => p.origin?.label === 'Kubernetes scheduler' ? { ...p, detail: summarizeSchedulerMessage(p.detail, { plain: true }), rawDetail: p.detail } : p)
const categories = new Set(
problems.filter((p) => p.severity !== 'posture').map((p) => p.category),
)
- const replica = cluster?.spec?.replica?.enabled ? { source: cluster.spec.replica.source } : null
+ const replica = cluster && getCNPGClusterIsReplica(cluster) ? { source: cluster.spec.replica.source } : null
return {
key: key(ns, name),
@@ -662,6 +1083,9 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
cluster,
controllerStatus: { text: status.text, level: status.level },
instances: { ready: readyInstances, desired },
+ ...(podReadiness ? { podReadiness } : {}),
+ ...(readinessContradicted ? { readinessContradicted } : {}),
+ ...(primaryConflict ? { primaryConflict } : {}),
pods: instancePods,
replicaCluster: replica,
hibernated,
@@ -670,6 +1094,7 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
replication: replicationFact(cluster, instancePods, hibernated, coverageOf(resp, 'pods')),
protection: { ...protection, summary: protectionSummary(protection) },
declarations: declarationsFor(cluster, resp),
+ poolerObjects: poolers.filter((p) => p.metadata?.namespace === ns && specClusterName(p) === name),
poolers: poolers
.filter((p) => p.metadata?.namespace === ns && specClusterName(p) === name)
.map((p) => p.metadata?.name),
@@ -677,20 +1102,49 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
problems,
attention: problems.some((p) => p.severity !== 'posture'),
categories,
- gitops: cnpgGitOpsSource(cluster),
+ managedBy: cnpgManagedBy(resp, cluster),
}
})
- rows.sort((a, b) => Number(b.attention) - Number(a.attention) || a.namespace.localeCompare(b.namespace) || a.name.localeCompare(b.name))
-
- const categoryCounts = { availability: 0, protection: 0, declarations: 0, pooling: 0 } as Record
- for (const r of rows) for (const c of r.categories) categoryCounts[c]++
-
const incompleteKinds = CNPG_WORKSPACE_KEYS.filter((k) => {
const s = coverageOf(resp, k).state
- return s === 'partial' || s === 'denied' || s === 'syncing' || s === 'error'
+ return s === 'partial' || s === 'denied' || s === 'syncing' || s === 'error' || s === 'uncached'
})
+ return finishFleet(rows, incompleteKinds)
+}
+
+function urgencyOf(row: CNPGFleetRow): { worst: number; urgent: number; total: number } {
+ let worst = 3
+ let urgent = 0
+ for (const p of row.problems) {
+ worst = Math.min(worst, CNPG_SEVERITY_RANK[p.severity])
+ if (p.severity !== 'posture') urgent++
+ }
+ return { worst, urgent, total: row.problems.length }
+}
+
+/**
+ * Worst problem first (critical, warning, posture, none), then the most
+ * attention-level problems, then all problems, then namespace/name — the same
+ * order in every filter, so a row never jumps when the filter changes.
+ */
+export function compareCNPGFleetUrgency(a: CNPGFleetRow, b: CNPGFleetRow): number {
+ const ua = urgencyOf(a)
+ const ub = urgencyOf(b)
+ return (
+ ua.worst - ub.worst ||
+ ub.urgent - ua.urgent ||
+ ub.total - ua.total ||
+ a.namespace.localeCompare(b.namespace) ||
+ a.name.localeCompare(b.name)
+ )
+}
+
+function finishFleet(rows: CNPGFleetRow[], incompleteKinds: CNPGWorkspaceKey[]): CNPGFleet {
+ rows.sort(compareCNPGFleetUrgency)
+ const categoryCounts = { availability: 0, protection: 0, declarations: 0, pooling: 0 } as Record
+ for (const r of rows) for (const c of r.categories) categoryCounts[c]++
return {
rows,
attentionCount: rows.filter((r) => r.attention).length,
@@ -698,3 +1152,533 @@ export function buildCNPGFleet(resp: CNPGWorkspaceResponse): CNPGFleet {
incompleteKinds,
}
}
+
+/** One cluster's answer from /api/cnpg/disk. */
+export interface CNPGDiskReading {
+ namespace: string
+ name: string
+ /** ok | partial | noSeries | noPrometheus | denied | unavailable | error | notRead | ambiguous | scopeMismatch */
+ state: string
+ grant?: Grant
+ reason?: string
+ claims: number
+ measured: number
+ max?: {
+ claim: string
+ instance: string
+ role: string
+ tablespace?: string
+ usedBytes: number
+ capacityBytes: number
+ ratio: number
+ }
+ isolation?: CNPGMetricIsolation
+}
+
+export const CNPG_DISK_WARNING_RATIO = 0.8
+export const CNPG_DISK_CRITICAL_RATIO = 0.9
+
+export const CNPG_DISK_SOURCE = 'kubelet volume stats via Prometheus'
+
+export function cnpgVolumeRoleLabel(role: string, tablespace?: string): string {
+ switch (role) {
+ case 'PG_DATA':
+ return 'data volume'
+ case 'PG_WAL':
+ return 'WAL volume'
+ case 'PG_TABLESPACE':
+ return tablespace ? `tablespace ${tablespace} volume` : 'tablespace volume'
+ default:
+ return 'volume'
+ }
+}
+
+export function cnpgDiskTone(ratio: number): HealthLevel {
+ if (ratio >= CNPG_DISK_CRITICAL_RATIO) return 'unhealthy'
+ if (ratio >= CNPG_DISK_WARNING_RATIO) return 'degraded'
+ return 'healthy'
+}
+
+/** The fleet and summary "Storage" fact: the fullest measured volume, or why there is none. */
+export function cnpgDiskFact(r: CNPGDiskReading | undefined): Fact {
+ if (!r) return { text: 'Not read', tone: 'unknown' }
+ if (r.max && (r.state === 'ok' || r.state === 'partial')) {
+ const partial = r.state === 'partial' ? ` · ${r.measured} of ${r.claims} volumes measured` : ''
+ return {
+ text: `${Math.round(r.max.ratio * 100)}% used`,
+ tone: cnpgDiskTone(r.max.ratio),
+ source: `Fullest: ${cnpgVolumeRoleLabel(r.max.role, r.max.tablespace)} of ${r.max.instance}, ${formatBytes(r.max.usedBytes)} of ${formatBytes(r.max.capacityBytes)} · ${CNPG_DISK_SOURCE}${partial}.${isolationCaveat(r.isolation)}`,
+ }
+ }
+ switch (r.state) {
+ case 'denied':
+ return { text: 'No access', tone: 'unknown', source: r.grant ? `Needs ${formatGrant(r.grant)}` : r.reason }
+ case 'noPrometheus':
+ return { text: 'No usage metrics', tone: 'unknown', source: CNPG_PROMETHEUS_NOT_CONNECTED, detail: r.reason }
+ case 'noSeries':
+ case 'ok':
+ case 'partial':
+ return { text: 'No usage metrics', tone: 'unknown', source: r.reason ?? 'Used space needs Prometheus with kubelet volume stats' }
+ case 'notRead':
+ return { text: 'Not measured', tone: 'unknown', source: r.reason }
+ default:
+ return { text: 'Unavailable', tone: 'unknown', source: r.reason }
+ }
+}
+
+/**
+ * Joins /api/cnpg/disk into the fleet: every row gets its disk fact, and a
+ * volume at or past the warning threshold becomes a problem, so the cluster
+ * needs attention. Only a measurement raises one, and the endpoint returns
+ * measurements only to callers holding the claim and metrics grants.
+ */
+export function applyCNPGDisk(fleet: CNPGFleet, readings: CNPGDiskReading[] | undefined): CNPGFleet {
+ if (!readings) return fleet
+ const byKey = new Map(readings.map((r) => [key(r.namespace, r.name), r]))
+ const rows = fleet.rows.map((row) => {
+ const reading = byKey.get(row.key)
+ const next: CNPGFleetRow = { ...row, disk: cnpgDiskFact(reading) }
+ const max = reading?.max
+ if (!max || max.ratio < CNPG_DISK_WARNING_RATIO || (reading.state !== 'ok' && reading.state !== 'partial')) return next
+ const problem: CNPGProblem = {
+ id: `disk:${row.key}`,
+ severity: max.ratio >= CNPG_DISK_CRITICAL_RATIO ? 'critical' : 'warning',
+ category: 'availability',
+ title: `The ${cnpgVolumeRoleLabel(max.role, max.tablespace)} of ${max.instance} is ${Math.round(max.ratio * 100)}% full`,
+ detail: `${formatBytes(max.usedBytes)} of ${formatBytes(max.capacityBytes)} used, from ${CNPG_DISK_SOURCE}.${isolationCaveat(reading.isolation)}`,
+ subject: { kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: row.namespace, name: row.name },
+ source: 'measurement',
+ ...(reading.isolation?.mode === 'unverified' ? { measuredBy: 'kubelet, matched by claim name', unverifiedMatch: true } : {}),
+ }
+ const problems = [...row.problems, problem].sort(cnpgCompareProblems)
+ return { ...next, problems, attention: true, categories: new Set([...row.categories, problem.category]) }
+ })
+ return finishFleet(rows, fleet.incompleteKinds)
+}
+
+const CNPG_PLUGIN_PHASES: Record = {
+ 'Cluster cannot proceed to reconciliation due to an unknown plugin being required': 'unknownPlugin',
+ 'Cluster cannot proceed to reconciliation due to an error while interacting with plugins': 'pluginError',
+}
+
+/** Whether the Cluster's phase says the operator is stuck on a CNPG-I plugin, and which way. */
+export function cnpgPluginPhase(cluster: any): 'unknownPlugin' | 'pluginError' | null {
+ return CNPG_PLUGIN_PHASES[cluster?.status?.phase] ?? null
+}
+
+/** The CNPG-I plugins a Cluster names in spec.plugins. */
+export function cnpgClusterPlugins(cluster: any): string[] {
+ const plugins = cluster?.spec?.plugins
+ return Array.isArray(plugins) ? plugins.map((p: any) => p?.name).filter((n: unknown): n is string => typeof n === 'string' && !!n) : []
+}
+
+/** Bytes in binary units with IEC labels (GiB), as every CloudNativePG view prints them. */
+export function cnpgFormatBytes(n: number): string {
+ const u = ['B', 'KiB', 'MiB', 'GiB', 'TiB']
+ let v = n
+ let i = 0
+ while (v >= 1024 && i < u.length - 1) {
+ v /= 1024
+ i++
+ }
+ return `${v.toFixed(v >= 10 || i === 0 ? 0 : 1)} ${u[i]}`
+}
+
+const formatBytes = cnpgFormatBytes
+
+// ---------------------------------------------------------------------------
+// Standbys and replication slots: one finding, whichever source saw it
+// ---------------------------------------------------------------------------
+
+/**
+ * Retained WAL at which an inactive physical slot becomes a concern. An HA
+ * slot only goes inactive when its standby stops consuming, and the WAL it
+ * pins keeps growing until that standby catches up or the slot is dropped.
+ */
+export const CNPG_SLOT_RETENTION_WARNING_BYTES = 1024 ** 3
+
+/** The id of a standby's "not receiving WAL" problem: the fleet (Prometheus) and the cluster page (instance manager) raise the same one. */
+export function cnpgStandbyProblemId(rowKey: string, pod: string): string {
+ return `standby:${rowKey}:${pod}`
+}
+
+/** The id of an inactive slot's retention problem, shared like cnpgStandbyProblemId. */
+export function cnpgSlotProblemId(rowKey: string, slot: string): string {
+ return `slot:${rowKey}:${slot}`
+}
+
+function haSlotPrefix(cluster: any): string {
+ const p = cluster?.spec?.replicationSlots?.highAvailability?.slotPrefix
+ return typeof p === 'string' && p ? p : '_cnpg_'
+}
+
+/** The HA slot CloudNativePG keeps for an instance: the prefix (default `_cnpg_`) then the instance name with `-` as `_`. */
+export function cnpgHASlotName(cluster: any, instance: string): string {
+ return `${haSlotPrefix(cluster)}${instance.replace(/-/g, '_')}`
+}
+
+/** The instance a managed HA slot serves, or undefined when the slot is not one of this cluster's HA slots. */
+export function cnpgHASlotInstance(cluster: any, slot: string, instances: string[]): string | undefined {
+ return instances.find((i) => cnpgHASlotName(cluster, i) === slot)
+}
+
+const CLUSTER_GROUP = 'postgresql.cnpg.io'
+
+function clusterSubject(row: Pick) {
+ return { kind: 'Cluster', group: CLUSTER_GROUP, namespace: row.namespace, name: row.name }
+}
+
+export interface CNPGStandbyGap {
+ pod: string
+ /** What was observed, in words, e.g. "its WAL receiver is down", "replay is paused". */
+ evidence: string[]
+ measuredBy: string
+ sourceDetail: string
+ /** Every expected standby is out: there is no failover target with current data. */
+ noneReceiving?: boolean
+ unverifiedMatch?: boolean
+}
+
+/** A standby that receives nothing from the primary. Its replay lag reads 0 because nothing new arrives to replay. */
+export function cnpgStandbyNotReceivingProblem(row: Pick, gap: CNPGStandbyGap): CNPGProblem {
+ return {
+ id: cnpgStandbyProblemId(row.key, gap.pod),
+ reason: 'CNPGStandbyNotReceiving',
+ severity: gap.noneReceiving ? 'critical' : 'warning',
+ category: 'availability',
+ title: `${gap.pod} is not receiving WAL from the primary`,
+ shortTitle: `${gap.pod} not receiving WAL`,
+ detail: `${capitalize(gap.evidence.join('; '))}. While it does not stream, the primary keeps WAL for its slot, and it falls behind unless it replays WAL from the archive; a replay lag of 0 does not mean it is caught up, only that nothing new reached it.`,
+ subject: { kind: 'Pod', group: '', namespace: row.namespace, name: gap.pod },
+ source: 'measurement',
+ measuredBy: gap.measuredBy,
+ sourceDetail: gap.sourceDetail,
+ ...(gap.unverifiedMatch ? { unverifiedMatch: true } : {}),
+ }
+}
+
+export interface CNPGSlotRetention {
+ slot: string
+ /** The instance holding the slot (the primary for an HA slot). */
+ pod: string
+ bytes: number
+ /** The standby an HA slot serves, when the name matches one of the cluster's instances. */
+ standby?: string
+ measuredBy: string
+ sourceDetail: string
+ unverifiedMatch?: boolean
+}
+
+export function cnpgSlotRetentionProblem(row: Pick, r: CNPGSlotRetention): CNPGProblem {
+ const forWhom = r.standby ? ` for ${r.standby}` : ''
+ return {
+ id: cnpgSlotProblemId(row.key, r.slot),
+ reason: 'CNPGInactiveSlot',
+ slot: r.slot,
+ severity: 'warning',
+ category: 'availability',
+ title: `Inactive slot ${r.slot} holds ${formatBytes(r.bytes)} of WAL on ${r.pod}${forWhom}`,
+ shortTitle: `${formatBytes(r.bytes)} of WAL held${forWhom || ` by ${r.slot}`}`,
+ detail: `PostgreSQL keeps every WAL file the slot still needs until ${r.standby ? `${r.standby} catches up` : 'its consumer catches up'} or the slot is dropped, so this grows while the slot stays inactive. Storage shows it beside the volume's other WAL.`,
+ subject: clusterSubject(row),
+ source: 'measurement',
+ measuredBy: r.measuredBy,
+ sourceDetail: r.sourceDetail,
+ ...(r.standby ? { instance: r.standby } : {}),
+ ...(r.unverifiedMatch ? { unverifiedMatch: true } : {}),
+ }
+}
+
+/**
+ * Adds problems to a row, replacing any with the same id, and drops those a
+ * newer, complete read has disproven (`drop`); attention and categories
+ * follow from what remains.
+ */
+export function cnpgWithProblems(row: CNPGFleetRow, added: CNPGProblem[], drop?: (p: CNPGProblem) => boolean): CNPGFleetRow {
+ if (added.length === 0 && !drop) return row
+ const ids = new Set(added.map((p) => p.id))
+ const problems = [...row.problems.filter((p) => !ids.has(p.id) && !drop?.(p)), ...added].sort(cnpgCompareProblems)
+ if (problems.length === row.problems.length && problems.every((p, i) => p === row.problems[i])) return row
+ return {
+ ...row,
+ problems,
+ attention: problems.some((p) => p.severity !== 'posture'),
+ categories: new Set(problems.filter((p) => p.severity !== 'posture').map((p) => p.category)),
+ }
+}
+
+function capitalize(s: string): string {
+ return s ? s[0].toUpperCase() + s.slice(1) : s
+}
+
+/** The short per-fact source when Radar has no Prometheus; the reason goes in `detail`. */
+export const CNPG_PROMETHEUS_NOT_CONNECTED = 'Prometheus not connected'
+
+const CNPG_LAG_UNMEASURED_SOURCE = 'Pod readiness does not show whether a replica is streaming'
+
+/** One cluster's answer from /api/cnpg/fleet-metrics. */
+export interface CNPGFleetMetricsReading {
+ namespace: string
+ name: string
+ /** ok | noStandby | noSeries | denied | ambiguous | scopeMismatch | error | notRead */
+ lag: {
+ state: string
+ grant?: Grant
+ reason?: string
+ seconds?: number
+ pod?: string
+ /** Standbys whose replay lag was read; `seconds` covers only these. */
+ lagStandbys?: number
+ /** The worst standby's lowest recorded lag over `sustainedWindow`, across every scrape of it; it was already reporting by the window's start. */
+ sustainedSeconds?: number
+ sustainedPod?: string
+ sustainedWindow?: string
+ isolation?: CNPGMetricIsolation
+ /** Instances reporting that they are in recovery (standbys). */
+ standbys?: number
+ /** Standbys whose WAL receiver is up (cnpg_pg_replication_is_wal_receiver_up = 1). */
+ receiving?: number
+ /** Standbys whose WAL receiver is down: their lag reads 0 because nothing arrives. */
+ receiverDown?: string[]
+ /** The exporter reported no receiver state (or the read failed), so streaming is not established from lag alone. */
+ receiverUnknown?: boolean
+ receiverReason?: string
+ /** Standbys whose receiver was down in every sample over `receiverDownWindow`: only these raise a problem, since a restarting standby is briefly down. */
+ receiverDownSustained?: string[]
+ receiverDownWindow?: string
+ }
+ /** Inactive physical replication slots and the WAL each keeps: ok | noSeries | denied | ambiguous | scopeMismatch | error | notRead */
+ slots?: {
+ state: string
+ grant?: Grant
+ reason?: string
+ /** Null unless `state` is ok; [] when ok and none is inactive. */
+ inactive?: { slot: string; pod: string; role?: 'primary' | 'standby'; bytes: number | null }[] | null
+ /** Inactive slots beyond the per-cluster cap. */
+ omitted?: number
+ isolation?: CNPGMetricIsolation
+ }
+ /** ok | noSeries | denied | unavailable | error | notRead */
+ growth: { state: string; grant?: Grant; reason?: string; bytesPerHour?: number; claim?: string; instance?: string; isolation?: CNPGMetricIsolation }
+}
+
+/** How Prometheus series were tied to one cluster; `unverified` matched only by namespace and Pod or claim names. */
+export interface CNPGMetricIsolation {
+ mode: 'configured' | 'verified' | 'unverified'
+ note: string
+}
+
+// A finding stated as this cluster's must say when its series were matched by name alone.
+function isolationCaveat(iso: CNPGMetricIsolation | undefined): string {
+ return iso?.mode === 'unverified' ? ` ${iso.note}.` : ''
+}
+
+export interface CNPGFleetMetricsSources {
+ /** prometheus, or none when Radar has no Prometheus (`reason` says why). */
+ source: 'prometheus' | 'none'
+ reason?: string
+ lagSource?: string
+ growthSource?: string
+}
+
+export function cnpgLagTone(seconds: number): HealthLevel {
+ if (seconds >= 30) return 'unhealthy'
+ if (seconds >= 5) return 'degraded'
+ return 'healthy'
+}
+
+/**
+ * Replication's tone from the primary's pg_stat_replication: a missing
+ * standby is degraded, and the lag of the ones that do stream can make it
+ * worse. Missing standbys never hide a severe lag.
+ */
+export function cnpgReplicationTone(streaming: number, expected: number, maxLagSeconds: number | undefined): HealthLevel {
+ const missing: HealthLevel = streaming < expected ? 'degraded' : 'healthy'
+ return maxLagSeconds === undefined ? missing : worseTone(missing, cnpgLagTone(maxLagSeconds))
+}
+
+/**
+ * The one way a replay lag reads: milliseconds below a second, one decimal
+ * below 10 s, whole seconds below 100 s, then whole minutes, then hours and
+ * minutes. Rounded down, so a lower bound stays one.
+ */
+export function cnpgFormatLag(s: number): string {
+ if (s <= 0) return '0 s'
+ if (s < 1) return `${Math.floor(s * 1000)} ms`
+ if (s < 10) return `${(Math.floor(s * 10) / 10).toFixed(1)} s`
+ if (s < 100) return `${Math.floor(s)} s`
+ const minutes = Math.floor(s / 60)
+ if (minutes < 60) return `${minutes} min`
+ const m = minutes % 60
+ return m === 0 ? `${Math.floor(minutes / 60)} h` : `${Math.floor(minutes / 60)} h ${m} min`
+}
+
+function measuredReplication(base: Fact, reading: CNPGFleetMetricsReading | undefined, src: CNPGFleetMetricsSources, designated?: string, expectedStandbys?: number | null): Fact {
+ const prefix = base.text.replace(/ · lag unknown$/, '')
+ if (src.source === 'none') {
+ return { text: `${prefix} · lag unknown`, tone: 'unknown', source: CNPG_PROMETHEUS_NOT_CONNECTED, detail: src.reason }
+ }
+ const lag = reading?.lag
+ switch (lag?.state) {
+ case 'ok': {
+ const down = (lag.receiverDown ?? []).filter((pod) => pod !== designated)
+ if (down.length > 0) {
+ return {
+ text: `${prefix} · ${down.length === 1 ? `${down[0]} not receiving WAL` : `${down.length} standbys not receiving WAL`}`,
+ tone: lag.receiving === 0 && down.length === lag.standbys ? 'unhealthy' : 'degraded',
+ source: `WAL receiver down on ${down.join(', ')} (cnpg_pg_replication_is_wal_receiver_up = 0) · ${src.lagSource ?? 'Prometheus'}.${isolationCaveat(lag.isolation)}`,
+ }
+ }
+ if (lag.seconds === undefined) break
+ // The lag covers the standbys whose lag was read; one that is not read is unknown, not caught up.
+ const partial = expectedStandbys !== undefined && expectedStandbys !== null && lag.lagStandbys !== undefined && lag.lagStandbys < expectedStandbys
+ const coverage = partial ? ` (${lag.lagStandbys} of ${expectedStandbys} ${designated ? 'instances' : 'standbys'} reporting)` : ''
+ const lagTone = cnpgLagTone(lag.seconds)
+ return {
+ text: `${prefix} · lag ${cnpgFormatLag(lag.seconds)}${coverage}${lag.receiverUnknown ? ' · streaming unverified' : ''}`,
+ tone: lag.receiverUnknown || (partial && lagTone === 'healthy') ? 'unknown' : lagTone,
+ source: `Largest standby replay lag, ${lag.pod ?? 'a standby'} · ${src.lagSource ?? 'Prometheus'}.${lag.receiverUnknown ? ' The exporter reported no WAL receiver state, and a standby that receives nothing also reads 0.' : ''}${isolationCaveat(lag.isolation)}`,
+ }
+ }
+ case 'noStandby':
+ return { text: `${prefix} · lag unknown`, tone: 'unknown', source: `No standby reports lag: ${lag.reason ?? 'no instance reports being a standby'} · ${src.lagSource ?? 'Prometheus'}` }
+ case 'denied':
+ return { text: `${prefix} · lag unknown`, tone: 'unknown', source: lag.grant ? `Needs ${formatGrant(lag.grant)}` : lag.reason }
+ }
+ return { text: `${prefix} · lag unknown`, tone: 'unknown', source: lag?.reason ?? 'Replication lag needs Prometheus scraping the CNPG exporter' }
+}
+
+/** Volume growth of the fastest-growing claim, as a fact. */
+export function cnpgDiskGrowthFact(reading: CNPGFleetMetricsReading | undefined, src: CNPGFleetMetricsSources): Fact | undefined {
+ const g = reading?.growth
+ if (src.source === 'none' || !g || g.state !== 'ok' || g.bytesPerHour === undefined) return undefined
+ const perDay = g.bytesPerHour * 24
+ const text = Math.abs(perDay) < 1024 ? 'flat over 6 h' : `${perDay > 0 ? '+' : '−'}${formatBytes(Math.abs(perDay))}/day`
+ return { text, tone: 'neutral', source: `Fastest-growing: ${g.claim ?? 'a volume'}${g.instance ? ` of ${g.instance}` : ''} · ${src.growthSource ?? 'Prometheus'}.${isolationCaveat(g.isolation)}` }
+}
+
+/**
+ * Joins /api/cnpg/fleet-metrics into the fleet: a cluster whose replication
+ * fact is only Pod readiness gets its measured standby lag, or says why it has
+ * none; the disk growth, when measured, lands on `diskGrowth`. Only lag that
+ * stayed high for the whole sustained window raises a problem; a spike is
+ * shown, not judged, and growth is never judged.
+ */
+export function applyCNPGFleetMetrics(fleet: CNPGFleet, readings: CNPGFleetMetricsReading[] | undefined, src: CNPGFleetMetricsSources | undefined): CNPGFleet {
+ if (!src) return fleet
+ const byKey = new Map((readings ?? []).map((r) => [key(r.namespace, r.name), r]))
+ const rows = fleet.rows.map((row) => {
+ const reading = byKey.get(row.key)
+ const next: CNPGFleetRow = { ...row, diskGrowth: cnpgDiskGrowthFact(reading, src) }
+ if (row.replication.source === CNPG_LAG_UNMEASURED_SOURCE) next.replication = measuredReplication(
+ row.replication,
+ reading,
+ src,
+ row.replicaCluster ? row.cluster?.status?.currentPrimary : undefined,
+ // Every instance of a replica cluster is in recovery, its designated primary included.
+ row.instances.desired !== null ? (row.replicaCluster ? row.instances.desired : Math.max(0, row.instances.desired - 1)) : null,
+ )
+ const sustained = sustainedLagProblem(row, reading, src)
+ return cnpgWithProblems(next, [...(sustained ? [sustained] : []), ...fleetStandbyProblems(row, reading, src), ...fleetSlotProblems(row, reading, src)])
+ })
+ return finishFleet(rows, fleet.incompleteKinds)
+}
+
+function fleetStandbyProblems(row: CNPGFleetRow, reading: CNPGFleetMetricsReading | undefined, src: CNPGFleetMetricsSources): CNPGProblem[] {
+ const lag = reading?.lag
+ if (src.source !== 'prometheus' || lag?.state !== 'ok' || row.hibernated) return []
+ // A replica cluster's designated primary is in recovery too, and may be fed
+ // from the WAL archive with no receiver at all.
+ const designated = row.replicaCluster ? row.cluster?.status?.currentPrimary : undefined
+ const down = (lag.receiverDownSustained ?? []).filter((pod) => pod !== designated)
+ const window = lag.receiverDownWindow ? formatWindowWords(lag.receiverDownWindow) : 'several minutes'
+ return down.map((pod) =>
+ cnpgStandbyNotReceivingProblem(row, {
+ pod,
+ evidence: [`its WAL receiver was down in every sample Prometheus recorded over the last ${window}`],
+ measuredBy: lag.isolation?.mode === 'unverified' ? 'Prometheus, matched by Pod name' : 'Prometheus',
+ sourceDetail: `cnpg_pg_replication_is_wal_receiver_up = 0 while cnpg_pg_replication_in_recovery = 1 · ${src.lagSource ?? 'Prometheus'}`,
+ noneReceiving: noStandbyReceives(row, lag),
+ unverifiedMatch: lag.isolation?.mode === 'unverified',
+ }),
+ )
+}
+
+// "None receives" only when every expected standby reported its receiver and
+// every one was down; one that did not report leaves it open.
+function noStandbyReceives(row: CNPGFleetRow, lag: CNPGFleetMetricsReading['lag']): boolean {
+ const expected = row.instances.desired !== null ? Math.max(0, row.instances.desired - 1) : null
+ const down = lag.receiverDownSustained?.length ?? 0
+ return lag.receiving === 0 && expected !== null && expected > 0 && lag.standbys === expected && down === expected
+}
+
+function fleetSlotProblems(row: CNPGFleetRow, reading: CNPGFleetMetricsReading | undefined, src: CNPGFleetMetricsSources): CNPGProblem[] {
+ const slots = reading?.slots
+ if (src.source !== 'prometheus' || slots?.state !== 'ok' || !slots.inactive) return []
+ // CloudNativePG copies HA slots to the standbys, where nothing streams from
+ // them, so a standby's copy is always inactive. Only the primary's copy
+ // means a consumer stopped.
+ const instances = row.pods.map((p) => p.name)
+ return slots.inactive
+ .filter((s): s is typeof s & { bytes: number } => s.role === 'primary' && s.bytes !== null && s.bytes >= CNPG_SLOT_RETENTION_WARNING_BYTES)
+ .map((s) =>
+ cnpgSlotRetentionProblem(row, {
+ slot: s.slot,
+ pod: s.pod,
+ bytes: s.bytes,
+ standby: cnpgHASlotInstance(row.cluster, s.slot, instances),
+ measuredBy: slots.isolation?.mode === 'unverified' ? 'Prometheus, matched by Pod name' : 'Prometheus',
+ sourceDetail: 'cnpg_pg_replication_slots_pg_wal_lsn_diff where cnpg_pg_replication_slots_active = 0 (physical slots)',
+ unverifiedMatch: slots.isolation?.mode === 'unverified',
+ }),
+ )
+}
+
+export const CNPG_SUSTAINED_LAG_WARNING_SECONDS = 30
+export const CNPG_SUSTAINED_LAG_CRITICAL_SECONDS = 300
+
+/** The id of a cluster's sustained-lag problem, so other views can find it among the row's problems. */
+export function cnpgSustainedLagProblemId(rowKey: string): string {
+ return `lag:${rowKey}`
+}
+
+function sustainedLagProblem(row: CNPGFleetRow, reading: CNPGFleetMetricsReading | undefined, src: CNPGFleetMetricsSources): CNPGProblem | undefined {
+ const lag = reading?.lag
+ const floor = lag?.sustainedSeconds
+ if (src.source !== 'prometheus' || lag?.state !== 'ok' || floor === undefined || floor < CNPG_SUSTAINED_LAG_WARNING_SECONDS) return undefined
+ const window = lag.sustainedWindow ? formatWindowWords(lag.sustainedWindow) : 'several minutes'
+ const pod = lag.sustainedPod ?? 'A standby'
+ // The query proves every recorded sample was at least the floor and that
+ // the series existed at the window's start, not that samples were continuous.
+ return {
+ id: cnpgSustainedLagProblemId(row.key),
+ reason: 'CNPGSustainedLag',
+ severity: floor >= CNPG_SUSTAINED_LAG_CRITICAL_SECONDS ? 'critical' : 'warning',
+ category: 'availability',
+ title: `${pod} ≥ ${cnpgFormatLag(floor)} behind in every sample for ${formatWindowShort(lag.sustainedWindow)}`,
+ shortTitle: `${pod}: all samples ≥ ${cnpgFormatLag(floor)} behind (${formatWindowShort(lag.sustainedWindow)})`,
+ detail: `Lowest replay lag in the samples Prometheus recorded over the last ${window}. If Prometheus missed some scrapes, those moments aren't included. If it was still that far behind, a failover to it would lose or wait on that much WAL.${isolationCaveat(lag.isolation)}`,
+ subject: { kind: 'Cluster', group: 'postgresql.cnpg.io', namespace: row.namespace, name: row.name },
+ source: 'measurement',
+ measuredBy: lag.isolation?.mode === 'unverified' ? 'Prometheus, matched by Pod name' : 'Prometheus',
+ unverifiedMatch: lag.isolation?.mode === 'unverified',
+ sourceDetail: src.lagSource ?? 'Prometheus',
+ }
+}
+
+// "10m0s" as "10 min"; "1h0m0s" as "1 h", for a title that must stay short.
+function formatWindowShort(d: string | undefined): string {
+ const m = d ? /^(?:(\d+)h)?(?:(\d+)m)?(?:0s)?$/.exec(d) : null
+ if (!m) return d ?? 'minutes'
+ const minutes = Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)
+ return minutes >= 60 && minutes % 60 === 0 ? `${minutes / 60} h` : `${minutes} min`
+}
+
+// "10m0s" as "10 minutes"; "1h0m0s" as "1 hour".
+function formatWindowWords(d: string): string {
+ const m = /^(?:(\d+)h)?(?:(\d+)m)?(?:0s)?$/.exec(d)
+ if (!m) return d
+ const minutes = Number(m[1] ?? 0) * 60 + Number(m[2] ?? 0)
+ if (minutes >= 60 && minutes % 60 === 0) return minutes === 60 ? '1 hour' : `${minutes / 60} hours`
+ return minutes === 1 ? '1 minute' : `${minutes} minutes`
+}
diff --git a/packages/k8s-ui/src/components/dock/DockContext.tsx b/packages/k8s-ui/src/components/dock/DockContext.tsx
index 02ee2083f6..b7fbb76ee4 100644
--- a/packages/k8s-ui/src/components/dock/DockContext.tsx
+++ b/packages/k8s-ui/src/components/dock/DockContext.tsx
@@ -16,6 +16,10 @@ export interface DockTab {
podName?: string
containerName?: string
containers?: string[]
+ /** Terminal: run this command instead of a shell (one argv element, e.g. "psql"). */
+ shell?: string
+ /** Terminal: a short line shown in the toolbar, e.g. what the session is connected to. */
+ sessionNote?: string
// Workload logs props
workloadKind?: string
workloadName?: string
@@ -87,7 +91,8 @@ export function DockProvider({ children }: { children: ReactNode }) {
}
return t.namespace === tabData.namespace &&
t.podName === tabData.podName &&
- t.containerName === tabData.containerName
+ t.containerName === tabData.containerName &&
+ t.shell === tabData.shell
})
if (existingTab) {
@@ -207,14 +212,19 @@ export function useOpenTerminal() {
orgId?: string
clusterId?: string
clusterName?: string
+ shell?: string
+ sessionNote?: string
+ title?: string
}) => {
addTab({
type: 'terminal',
- title: `${opts.podName}/${opts.containerName}`,
+ title: opts.title ?? `${opts.podName}/${opts.containerName}`,
namespace: opts.namespace,
podName: opts.podName,
containerName: opts.containerName,
containers: opts.containers,
+ shell: opts.shell,
+ sessionNote: opts.sessionNote,
orgId: opts.orgId,
clusterId: opts.clusterId,
clusterName: opts.clusterName,
diff --git a/packages/k8s-ui/src/components/dock/TerminalTab.tsx b/packages/k8s-ui/src/components/dock/TerminalTab.tsx
index 68cedf5abf..f1d06f0192 100644
--- a/packages/k8s-ui/src/components/dock/TerminalTab.tsx
+++ b/packages/k8s-ui/src/components/dock/TerminalTab.tsx
@@ -19,6 +19,8 @@ export interface TerminalTabProps {
createSession: (containerName: string) => Promise<{ wsUrl: string }>
/** Optional: creates a debug (ephemeral) container. If omitted, the debug button is hidden. */
createDebugContainer?: (targetContainer: string) => Promise<{ containerName: string }>
+ /** Optional: a short line in the toolbar describing the session. */
+ note?: string
}
export function TerminalTab({
@@ -29,6 +31,7 @@ export function TerminalTab({
isActive = true,
createSession,
createDebugContainer,
+ note,
}: TerminalTabProps) {
const terminalRef = useRef(null)
const xtermRef = useRef(null)
@@ -269,6 +272,7 @@ export function TerminalTab({
)}
/>
{podName}
+ {note && {note} }
{containers.length > 1 && (
diff --git a/packages/k8s-ui/src/components/facts/certainty.tsx b/packages/k8s-ui/src/components/facts/certainty.tsx
new file mode 100644
index 0000000000..c33d2ffbfb
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/certainty.tsx
@@ -0,0 +1,36 @@
+import { WithTooltip } from '../ui/Tooltip'
+
+/**
+ * How exactly a value is known. A value read only in part is a lower bound,
+ * never shown as if exact.
+ */
+export type Certainty = 'exact' | 'lower_bound' | 'upper_bound' | 'unknown'
+
+export function certaintyGlyph(certainty: Certainty): string {
+ if (certainty === 'exact') return '='
+ if (certainty === 'lower_bound') return '≥'
+ if (certainty === 'upper_bound') return '≤'
+ return '?'
+}
+
+export function certaintyValueLabel(certainty: Certainty): string {
+ if (certainty === 'exact') return 'Exact'
+ if (certainty === 'lower_bound') return 'Lower bound'
+ if (certainty === 'upper_bound') return 'Upper bound'
+ return 'Unknown certainty'
+}
+
+export function CertaintyGlyph({ certainty, title }: { certainty: Certainty; title?: string }) {
+ return (
+
+
+ {certaintyGlyph(certainty)}
+
+
+ )
+}
diff --git a/packages/k8s-ui/src/components/facts/facts.test.tsx b/packages/k8s-ui/src/components/facts/facts.test.tsx
new file mode 100644
index 0000000000..a560d293f2
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/facts.test.tsx
@@ -0,0 +1,16 @@
+import { describe, expect, it } from 'vitest'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { FactValue } from './facts'
+
+describe('FactValue', () => {
+ const twoDaysAgo = new Date(Date.now() - 2 * 24 * 3600 * 1000 - 60_000).toISOString()
+ it('reads a still-current state as lasting since its timestamp', () => {
+ const html = renderToStaticMarkup(
)
+ expect(html).toContain('Failing
for 2d ')
+ expect(html).not.toContain('ago')
+ })
+ it('reads a past event as an age', () => {
+ const html = renderToStaticMarkup(
)
+ expect(html).toContain('2d ago')
+ })
+})
diff --git a/packages/k8s-ui/src/components/facts/facts.tsx b/packages/k8s-ui/src/components/facts/facts.tsx
new file mode 100644
index 0000000000..3b6409d530
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/facts.tsx
@@ -0,0 +1,60 @@
+import type { ReactNode } from 'react'
+import { clsx } from 'clsx'
+import type { HealthLevel } from '../resources/resource-utils'
+import { formatAge } from '../resources/resource-utils'
+import { toneTextClass } from '../ui/status-tone'
+import { Tooltip } from '../ui/Tooltip'
+
+/**
+ * One observed value and where it came from. A value the cluster does not
+ * report is a fact too: its text says so and its tone is `unknown`, never a
+ * zero or a calm default.
+ */
+export interface Fact {
+ text: string
+ tone: HealthLevel
+ /** Where the value comes from, shown next to it so claims carry their source. */
+ source?: string
+ /** A timestamp the text refers to; the UI renders it as an age. */
+ at?: string
+ /** `since`: `at` is when a still-current state began, rendered "Failing for 2d" rather than "· 2d ago". */
+ atMeaning?: 'since'
+ /** The full explanation behind a short `source`, shown on hover only. */
+ detail?: string
+}
+
+export function FactValue({ fact, className }: { fact: Fact; className?: string }) {
+ const age = fact.at ? formatAge(fact.at) : null
+ const body = (
+
+ {fact.text}
+ {age && fact.atMeaning === 'since' && for {age} }
+ {age && fact.atMeaning !== 'since' && {fact.text ? ' · ' : ''}{age} ago }
+
+ )
+ if (!fact.source && !fact.at && !fact.detail) return body
+ return (
+
+ {body}
+
+ )
+}
+
+export function FactSource({ fact }: { fact: Fact }) {
+ if (!fact.source) return null
+ return
{fact.source}
+}
+
+/** Label/value rows. Empty values stay visible: an unread value is shown as unread, not hidden. */
+export function FactGrid({ children }: { children: ReactNode }) {
+ return
{children}
+}
+
+export function FactRow({ label, children }: { label: ReactNode; children: ReactNode }) {
+ return (
+ <>
+
{label}
+
{children}
+ >
+ )
+}
diff --git a/packages/k8s-ui/src/components/facts/index.ts b/packages/k8s-ui/src/components/facts/index.ts
new file mode 100644
index 0000000000..ce2f304a7c
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/index.ts
@@ -0,0 +1,6 @@
+// How a surface shows what the cluster reported: each value with its source,
+// partial and unread values marked as such (see DESIGN.md, "Unknown, partial
+// and denied values"). For any renderer, summary or workspace screen.
+export * from './facts'
+export * from './certainty'
+export * from './managed-by'
diff --git a/packages/k8s-ui/src/components/facts/managed-by.test.tsx b/packages/k8s-ui/src/components/facts/managed-by.test.tsx
new file mode 100644
index 0000000000..842c51f1ef
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/managed-by.test.tsx
@@ -0,0 +1,26 @@
+import { describe, expect, it } from 'vitest'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { ManagedByText, managedByLabel } from './managed-by'
+import { cnpgManagedBy } from '../cnpg/workspace'
+
+describe('ManagedByText', () => {
+ it('names the GitOps manager and links it when its namespace is recorded', () => {
+ const app = { kind: 'Application', group: 'argoproj.io', namespace: 'argocd', name: 'payments' }
+ const html = renderToStaticMarkup(
{}} />)
+ expect(html).toContain('Argo CD application')
+ expect(html).toMatch(/]*>argocd\/payments<\/button>/)
+ expect(renderToStaticMarkup( {}} />)).not.toContain(' {
+ expect(renderToStaticMarkup( )).toBe('')
+ expect(managedByLabel({ kind: 'Kustomization', group: 'kustomize.toolkit.fluxcd.io', namespace: 'flux-system', name: 'apps' })).toBe('Flux Kustomization flux-system/apps')
+ })
+})
+
+describe('cnpgManagedBy', () => {
+ it('looks an object up by kind, namespace and name in the workspace answer', () => {
+ const ws = { managedBy: { 'Cluster/db/pg': { kind: 'Application', group: 'argoproj.io', namespace: 'argocd', name: 'pg' } } }
+ expect(cnpgManagedBy(ws, { kind: 'Cluster', metadata: { namespace: 'db', name: 'pg' } })?.name).toBe('pg')
+ expect(cnpgManagedBy(ws, { kind: 'Pooler', metadata: { namespace: 'db', name: 'pg' } })).toBeUndefined()
+ })
+})
diff --git a/packages/k8s-ui/src/components/facts/managed-by.tsx b/packages/k8s-ui/src/components/facts/managed-by.tsx
new file mode 100644
index 0000000000..c5ba1c1ec4
--- /dev/null
+++ b/packages/k8s-ui/src/components/facts/managed-by.tsx
@@ -0,0 +1,36 @@
+import type { ResourceRef } from '../../types/core'
+import { gitOpsOwnerFromRef, type GitOpsOwnerRef } from '../../utils/gitops-owner'
+import { RefLink, type NavigateToRef } from '../ui/RefLink'
+
+function managerLabel(owner: GitOpsOwnerRef): string {
+ if (owner.tool === 'argocd') return 'Argo CD application'
+ return owner.kind === 'helmreleases' ? 'Flux HelmRelease' : 'Flux Kustomization'
+}
+
+/**
+ * The GitOps object that manages a resource, from the server's manager
+ * detection. Renders nothing for a manager that is not a GitOps controller,
+ * and plain text when the manager's namespace is not recorded.
+ */
+export function ManagedByText({ refTo, onNavigate }: { refTo: ResourceRef; onNavigate?: NavigateToRef }) {
+ const owner = gitOpsOwnerFromRef(refTo)
+ if (!owner) return null
+ return (
+
+ {managerLabel(owner)}{' '}
+ {refTo.namespace ? (
+
+ {`${refTo.namespace}/${refTo.name}`}
+
+ ) : (
+ {refTo.name}
+ )}
+
+ )
+}
+
+/** The manager's label and name as text, for a table cell. */
+export function managedByLabel(refTo: ResourceRef | undefined): string | undefined {
+ const owner = refTo && gitOpsOwnerFromRef(refTo)
+ return owner ? `${managerLabel(owner)} ${refTo!.namespace ? `${refTo!.namespace}/` : ''}${refTo!.name}` : undefined
+}
diff --git a/packages/k8s-ui/src/components/gitops/detail-helpers.ts b/packages/k8s-ui/src/components/gitops/detail-helpers.ts
index 1956312321..761dfd6a50 100644
--- a/packages/k8s-ui/src/components/gitops/detail-helpers.ts
+++ b/packages/k8s-ui/src/components/gitops/detail-helpers.ts
@@ -1,3 +1,4 @@
+import { stripTrailingSlashes } from '../../utils/url-path'
import { argoApplicationSetConditionsToGitOpsStatus, argoStatusToGitOpsStatus, fluxConditionsToGitOpsStatus, type FluxCondition, type GitOpsStatus } from '../../types/gitops'
import type { GitOpsResourceTree } from '../../types/gitops-tree'
import { formatCompactAge } from '../../utils/format'
@@ -33,7 +34,7 @@ export function formatGitOpsDestination(server: string | undefined, namespace: s
// "https://kubernetes.default.svc/" variant (the in-cluster URL with a
// trailing slash that some controller versions emit) matches the literal
// we collapse to "in-cluster".
- let host = (server || '').trim().replace(/\/+$/, '')
+ let host = stripTrailingSlashes((server || '').trim())
if (host === '' || host === 'https://kubernetes.default.svc' || host === 'in-cluster') {
host = 'in-cluster'
} else {
diff --git a/packages/k8s-ui/src/components/issues/IssuesView.tsx b/packages/k8s-ui/src/components/issues/IssuesView.tsx
index 01384d67d4..ae54fbdb3f 100644
--- a/packages/k8s-ui/src/components/issues/IssuesView.tsx
+++ b/packages/k8s-ui/src/components/issues/IssuesView.tsx
@@ -1,3 +1,4 @@
+import { summarizeSchedulerMessage } from '../resources/resource-utils'
import { useMemo, useState, type ComponentType, type ReactNode } from 'react';
import { AlertOctagon, AlertTriangle, ArrowRight, CircleCheck, Clock, ExternalLink, Layers, Terminal, Workflow } from 'lucide-react';
import { CardBody, CardSection, ClusterName, EmptyState, KIND_CHIP_CLASS, TerminalBlock } from '../ui';
@@ -13,7 +14,7 @@ import {
ISSUE_SEVERITY_RAIL_CLASS,
ISSUE_SEVERITY_SOLID_CLASS,
ISSUE_SEVERITY_TEXT_CLASS,
- categoryLabel,
+ issueTitle,
groupBadgeClass,
groupLabel,
} from './severity';
@@ -168,6 +169,7 @@ export interface IssueRowProps {
resourceHref?: (ref: IssueResourceRef) => string;
onResourceClick?: (ref: IssueResourceRef) => void;
as?: 'li' | 'div';
+ compact?: boolean;
className?: string;
dimmed?: boolean;
/** Suppress the "Subject" deep-link in the expanded body — set by hosts that
@@ -195,6 +197,7 @@ export function IssueRow({
resourceHref,
onResourceClick,
as = 'li',
+ compact = false,
className,
dimmed,
hideSubject,
@@ -208,6 +211,7 @@ export function IssueRow({
const cluster = clusterLabel?.(issue);
const affected = affectedSummary(issue.affected);
const { headline } = issueMessageParts(issue);
+ const schedulingCause = compact && issue.reason === 'Unschedulable' ? summarizeSchedulerMessage(issue.cause || issue.message || headline || '', { plain: true }) : undefined;
const { panelId, buttonProps } = useDisclosure(open);
const Container = as;
const severity = normalizeIssueSeverity(issue.severity);
@@ -289,7 +293,7 @@ export function IssueRow({
- {categoryLabel(issue.category)}
+ {issueTitle(issue)}
{groupLabel(issue.category_group)}
{renderBadges?.(slotCtx)}
{/* The detector reason rides the title row while COLLAPSED so the
@@ -297,7 +301,7 @@ export function IssueRow({
cause lives in the WHAT'S WRONG section below, so it fades out
here (stays mounted — it's the flex-1 filler, so unmounting
wouldn't reflow anything, but fading avoids the pop). */}
- {issue.reason ? (
+ {issue.reason && !schedulingCause ? (
) : null}
+ {schedulingCause &&
{schedulingCause}
}
{issue.kind}
diff --git a/packages/k8s-ui/src/components/issues/ResourceIssuesSection.test.tsx b/packages/k8s-ui/src/components/issues/ResourceIssuesSection.test.tsx
index 5700d3bc7a..4849556cdf 100644
--- a/packages/k8s-ui/src/components/issues/ResourceIssuesSection.test.tsx
+++ b/packages/k8s-ui/src/components/issues/ResourceIssuesSection.test.tsx
@@ -106,3 +106,31 @@ it('keeps pod/template evidence in expanded neutral context with navigable witne
expect(expanded).not.toContain('confidence')
expect(expanded).not.toContain('Caused by')
})
+
+it('wraps a compact Pod scheduling cause separately and retains the full message in details', () => {
+ const raw = '0/2 nodes are available: 2 Too many pods. preemption: 0/2 nodes are available: 2 No preemption victims found for incoming pod.'
+ const pod: Issue = { ...issue, kind: 'Pod', reason: 'Unschedulable', cause: raw, message: raw }
+ const html = renderToString( {}} />)
+ expect(html).toContain('break-words text-xs text-theme-text-secondary')
+ expect(html).toContain('both nodes have reached their Pod limit')
+ expect(html).toContain(raw)
+ const producerMessage = '2 node(s) insufficient pods (0/2 nodes available)'
+ const reported = renderToString( {}} />)
+ expect(reported).toContain('both nodes have reached their Pod limit')
+ expect(reported).toContain(producerMessage)
+ const regular = renderToString( {}} />)
+ expect(regular).not.toContain('both nodes have reached their Pod limit')
+})
+
+it.each(['Cluster', 'Deployment'])('wraps compact scheduling causes aggregated under %s', (kind) => {
+ const raw = '2 node(s) insufficient pods (0/2 nodes available)'
+ const aggregated: Issue = { ...issue, kind, reason: 'Unschedulable', cause: raw, message: raw }
+ const collapsed = renderToString( {}} />)
+ expect(collapsed).toContain('both nodes have reached their Pod limit
')
+ expect(collapsed).not.toContain(raw)
+ const expanded = renderToString( {}} />)
+ expect(expanded).toContain(raw)
+ const regular = renderToString( {}} />)
+ expect(regular).toContain('Unschedulable')
+ expect(regular).not.toContain('both nodes have reached their Pod limit')
+})
diff --git a/packages/k8s-ui/src/components/issues/ResourceIssuesSection.tsx b/packages/k8s-ui/src/components/issues/ResourceIssuesSection.tsx
index d80e58fb18..074f9fffdc 100644
--- a/packages/k8s-ui/src/components/issues/ResourceIssuesSection.tsx
+++ b/packages/k8s-ui/src/components/issues/ResourceIssuesSection.tsx
@@ -8,7 +8,9 @@ export function ResourceIssuesSection({
issues,
onResourceClick,
subjectResource,
+ compact,
}: {
+ compact?: boolean
issues: Issue[] | undefined
/** When provided, related resources in a causal link become clickable. */
onResourceClick?: (ref: IssueResourceRef) => void
@@ -36,6 +38,7 @@ export function ResourceIssuesSection({
setOpenId((cur) => (cur === key ? null : key))}
onResourceClick={onResourceClick}
diff --git a/packages/k8s-ui/src/components/issues/index.ts b/packages/k8s-ui/src/components/issues/index.ts
index 48260f5235..93fcb849b1 100644
--- a/packages/k8s-ui/src/components/issues/index.ts
+++ b/packages/k8s-ui/src/components/issues/index.ts
@@ -24,5 +24,7 @@ export {
ISSUE_SEVERITY_FILL_CLASS,
ISSUE_SEVERITY_RAIL_CLASS,
categoryLabel,
+ issueTitle,
+ issueReasonTitle,
groupLabel,
} from './severity';
diff --git a/packages/k8s-ui/src/components/issues/issues.test.ts b/packages/k8s-ui/src/components/issues/issues.test.ts
index f31625e481..8311602cc5 100644
--- a/packages/k8s-ui/src/components/issues/issues.test.ts
+++ b/packages/k8s-ui/src/components/issues/issues.test.ts
@@ -2,7 +2,7 @@ import { afterEach, describe, it, expect, vi } from 'vitest'
import { createElement } from 'react'
import { renderToString } from 'react-dom/server'
import { compareIssues, issueSortAnchor, subjectRef, memberRef, normalizeImagePullMessage, issueMessageParts, type Issue } from './types'
-import { categoryLabel, groupBadgeClass, groupLabel } from './severity'
+import { categoryLabel, groupBadgeClass, groupLabel, issueTitle } from './severity'
import { IssueRow } from './IssuesView'
import { issueFirstSeenTitle, issueResourceCreatedTitle, issueTiming } from './issue-timing'
@@ -83,6 +83,8 @@ describe('category/group label fallbacks', () => {
it('returns the mapped label, else humanizes (server-added category needs no frontend deploy)', () => {
expect(categoryLabel('crashloop')).toBe('Crash loop')
expect(categoryLabel('some_new_future_category')).toBe('Some new future category')
+ expect(issueTitle({ category: 'backup_failed', reason: 'CNPGWALArchivingFailing' })).toBe('WAL archiving failing')
+ expect(issueTitle({ category: 'backup_failed', reason: 'BackupFailed' })).toBe('Backup failed')
})
it('humanizes an unmapped group', () => {
expect(groupLabel('runtime')).toBe('Runtime')
diff --git a/packages/k8s-ui/src/components/issues/severity.ts b/packages/k8s-ui/src/components/issues/severity.ts
index 9272169213..6417e9a185 100644
--- a/packages/k8s-ui/src/components/issues/severity.ts
+++ b/packages/k8s-ui/src/components/issues/severity.ts
@@ -136,6 +136,27 @@ export function categoryLabel(category: string): string {
return CATEGORY_LABEL[category] ?? humanize(category);
}
+// Reasons filed under a category whose label names a different operation:
+// CNPG archiving and an unanswered scheduled run sit under backup_failed.
+const REASON_TITLE: Record = {
+ CNPGWALArchivingFailing: "WAL archiving failing",
+ CNPGLastBackupFailed: "Latest backup failed",
+ CNPGScheduleDestinationMissing: "Backup schedule has no destination",
+ CNPGInstanceReadinessMismatch: "Instance Pods contradict Cluster readiness",
+ CNPGPrimaryLabelMismatch: "Primary Pod label contradicts Cluster status",
+ CNPGScheduledRunNoBackup: "No successful backup since a scheduled run",
+};
+
+/** The title a reason has wherever it is shown, when its category's label would name the wrong operation. */
+export function issueReasonTitle(reason: string | undefined): string | undefined {
+ return reason ? REASON_TITLE[reason] : undefined;
+}
+
+/** An issue row's title: its category, unless the reason names it better. */
+export function issueTitle(issue: { category: string; reason?: string }): string {
+ return issueReasonTitle(issue.reason) || categoryLabel(issue.category);
+}
+
export function groupLabel(group: string): string {
return GROUP_LABEL[group] ?? humanize(group);
}
diff --git a/packages/k8s-ui/src/components/logs/LogCore.theme.test.tsx b/packages/k8s-ui/src/components/logs/LogCore.theme.test.tsx
index 2affc30e14..8c6ac30789 100644
--- a/packages/k8s-ui/src/components/logs/LogCore.theme.test.tsx
+++ b/packages/k8s-ui/src/components/logs/LogCore.theme.test.tsx
@@ -1,9 +1,11 @@
// @vitest-environment jsdom
-import { act } from 'react'
+import { act, type ComponentProps, type ReactNode } from 'react'
import { createRoot, type Root } from 'react-dom/client'
-import { afterEach, beforeEach, describe, expect, it } from 'vitest'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { LogCore } from './LogCore'
+vi.mock('react-virtuoso', () => ({ Virtuoso: ({ data, itemContent }: { data: unknown[]; itemContent: (index: number, value: unknown) => ReactNode }) => {data.map((value, index) =>
{itemContent(index, value)}
)}
}))
+
Object.assign(globalThis, { IS_REACT_ACT_ENVIRONMENT: true })
let root: Root
@@ -20,7 +22,7 @@ afterEach(async () => {
})
const noop = () => {}
-async function render(props: { forceDark?: boolean; defaultDark?: boolean }) {
+async function render(props: Partial>) {
await act(async () =>
root.render(
{
expect(toggle()).toBeNull()
})
})
+
+it('dims unavailable streaming, explains why, suppresses its hint and keyboard action', async () => {
+ vi.useFakeTimers()
+ const start = vi.fn()
+ await render({ onStartStream: start, sourceUnavailable: true })
+ const stream = [...element.querySelectorAll('button')].find((b) => b.textContent === 'Stream')!
+ expect(stream.disabled).toBe(true)
+ expect(stream.classList.contains('disabled:opacity-50')).toBe(true)
+ expect([...element.querySelectorAll('kbd')].some((k) => k.textContent === 'S')).toBe(false)
+ act(() => { stream.parentElement!.dispatchEvent(new MouseEvent('mouseover', { bubbles: true })); vi.advanceTimersByTime(500) })
+ expect(document.body.textContent).toContain('Available once an instance is running')
+ act(() => window.dispatchEvent(new KeyboardEvent('keydown', { key: 's' })))
+ expect(start).not.toHaveBeenCalled()
+ vi.useRealTimers()
+})
+it('keeps a non-CNPG caller streamable without Pods and wraps its utilities together', async () => {
+ const start = vi.fn()
+ await render({ onStartStream: start, onClear: noop, toolbarExtra: Deployment nginx sources , entries: [{ id: 1, content: '{"msg":"nginx started","level":"info"}', timestamp: '', container: 'nginx', pod: 'nginx-1', isJson: true, isLogfmt: false, level: 'info', levelSource: 'structured' }] })
+ const stream = [...element.querySelectorAll('button')].find((b) => b.textContent === 'Stream')!
+ expect(stream.disabled).toBe(false)
+ expect([...element.querySelectorAll('kbd')].some((k) => k.textContent === 'S')).toBe(true)
+ act(() => stream.click())
+ expect(start).toHaveBeenCalledOnce()
+ const utilities = element.querySelector('[aria-label="Log display and utilities"]')!
+ expect(utilities.classList.contains('flex-nowrap')).toBe(true)
+ expect(utilities.classList.contains('shrink-0')).toBe(true)
+ for (const icon of ['lucide-trash-2', 'lucide-download', 'lucide-search', 'lucide-clock', 'lucide-braces']) expect(utilities.querySelector(`.${icon}`)).not.toBeNull()
+})
+
+it('still stops an existing stream when its source becomes unavailable', async () => {
+ const stop = vi.fn()
+ await render({ onStartStream: noop, onStopStream: stop, isStreaming: true, sourceUnavailable: true })
+ const button = [...element.querySelectorAll('button')].find((b) => b.textContent === 'Stop')!
+ expect(button.disabled).toBe(false)
+ act(() => window.dispatchEvent(new KeyboardEvent('keydown', { key: 's' })))
+ expect(stop).toHaveBeenCalledOnce()
+})
+
+it.each([true, false])('wraps non-CNPG plain logs at words with ANSI enabled=%s', async (ansi) => {
+ localStorage.setItem('radar-logs-ansi', String(ansi))
+ await render({ entries: [{ id: 1, timestamp: '', content: 'nginx serving requests', container: 'nginx', isJson: false, isLogfmt: false, level: 'info', levelSource: 'keyword' }] })
+ const line = [...element.querySelectorAll('span')].find((el) => el.textContent === 'nginx serving requests' && el.classList.contains('whitespace-pre-wrap'))!
+ expect(line.classList.contains('[overflow-wrap:anywhere]')).toBe(true)
+ expect(line.classList.contains('break-all')).toBe(false)
+})
+it('keeps normal word wrapping when searching a structured log', async () => {
+ await render({ entries: [{ id: 1, timestamp: '', content: '{"msg":"nginx serving requests"}', container: 'nginx', isJson: true, isLogfmt: false, level: 'info', levelSource: 'structured' }] })
+ await act(async () => window.dispatchEvent(new KeyboardEvent('keydown', { key: 'f', ctrlKey: true })))
+ const input = element.querySelector('input[placeholder="Search logs..."]')!
+ await act(async () => {
+ Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, 'value')!.set!.call(input, 'nginx')
+ input.dispatchEvent(new Event('input', { bubbles: true }))
+ })
+ const highlight = element.querySelector('mark')!
+ expect(highlight).not.toBeNull()
+ expect(highlight.parentElement!.classList.contains('[overflow-wrap:anywhere]')).toBe(true)
+ expect(highlight.parentElement!.classList.contains('break-all')).toBe(false)
+})
+
+it('disables empty Export and Clear while preserving Refresh and enables them after logs load', async () => {
+ await render({ onClear: noop })
+ const exportButton = () => element.querySelector('[aria-label="Export logs"]')!
+ const clear = () => element.querySelector('[aria-label="Clear logs"]')!
+ expect(exportButton().disabled).toBe(true)
+ expect(clear().disabled).toBe(true)
+ expect(element.querySelector('[aria-label="Refresh logs"]')?.disabled).toBe(false)
+ await render({ onClear: noop, entries: [{ id: 1, container: 'postgres', level: 'info', levelSource: 'keyword', isJson: false, isLogfmt: false, content: 'hello', timestamp: '2026-10-01T00:00:00Z' }] })
+ expect(exportButton().disabled).toBe(false)
+ expect(clear().disabled).toBe(false)
+})
+
+it('keeps Export and Clear usable when a Pod filter hides a nonempty buffer', async () => {
+ await render({ entries: [], allEntries: [{ id: 1, container: 'postgres', level: 'info', levelSource: 'keyword', isJson: false, isLogfmt: false, content: 'hello', timestamp: '' }], onClear: noop })
+ expect(element.querySelector('[aria-label="Export logs"]')!.disabled).toBe(false)
+ expect(element.querySelector('[aria-label="Clear logs"]')!.disabled).toBe(false)
+})
diff --git a/packages/k8s-ui/src/components/logs/LogCore.tsx b/packages/k8s-ui/src/components/logs/LogCore.tsx
index 38b753b29a..993d1c35b9 100644
--- a/packages/k8s-ui/src/components/logs/LogCore.tsx
+++ b/packages/k8s-ui/src/components/logs/LogCore.tsx
@@ -58,7 +58,8 @@ interface LogCoreProps {
onClear?: () => void
toolbarExtra?: ToolbarExtraRenderer
showPodName?: boolean
- emptyMessage?: string
+ emptyMessage?: ReactNode
+ sourceUnavailable?: boolean
emptyCommand?: string | null
errorMessage?: string | null
/**
@@ -146,6 +147,7 @@ export function LogCore({
toolbarExtra,
showPodName = false,
emptyMessage = 'No logs available',
+ sourceUnavailable = false,
emptyCommand,
errorMessage,
forceDark,
@@ -422,7 +424,7 @@ export function LogCore({
search.open()
return
}
- if (e.key === 's' && !e.ctrlKey && !e.metaKey && !e.altKey && onStartStream) {
+ if (e.key === 's' && !e.ctrlKey && !e.metaKey && !e.altKey && onStartStream && (isStreaming || !sourceUnavailable)) {
const target = e.target as HTMLElement | null
if (target && (target.tagName === 'INPUT' || target.tagName === 'TEXTAREA' || target.tagName === 'SELECT' || target.isContentEditable)) return
e.preventDefault()
@@ -432,7 +434,7 @@ export function LogCore({
}
window.addEventListener('keydown', handleKeyDown)
return () => window.removeEventListener('keydown', handleKeyDown)
- }, [search.open, onStartStream, onStopStream, isStreaming])
+ }, [search.open, onStartStream, onStopStream, isStreaming, sourceUnavailable])
const handleFollowOutput = useCallback((isAtBottom: boolean) => {
if (isAtBottom) return 'smooth' as const
@@ -568,15 +570,16 @@ export function LogCore({
style={{ colorScheme: isDark ? 'dark' : 'light', fontFamily: "'SF Mono', 'Cascadia Code', 'Fira Code', Menlo, Consolas, 'DejaVu Sans Mono', monospace" }}
>
{/* Toolbar */}
-
+
{toolbarExtraNode}
{/* Stream / Stop toggle — only shown when streaming is supported */}
{onStartStream && (
-
+
@@ -608,7 +612,7 @@ export function LogCore({
toggleLevel(opt.level)}
- className={`px-1.5 py-0.5 text-[10px] font-medium rounded border transition-colors ${
+ className={`px-1.5 py-0.5 text-[10px] font-medium rounded border whitespace-nowrap transition-colors ${
active
? getLevelActiveColor(opt.level, palette)
: levelChipInactive
@@ -621,7 +625,7 @@ export function LogCore({
})}
-
+
{/* Structured-log display mode: icon cycles compact→expanded→raw, chevron picks explicitly. */}
{hasStructuredEntries && (
@@ -801,14 +805,15 @@ export function LogCore({
{/* Export */}
-
+
@@ -884,15 +889,18 @@ export function LogCore({
{/* Clear */}
{onClear && (
-
+
)}
+
{/* Search bar */}
@@ -1007,7 +1015,7 @@ export function LogCore({
) : groupedEntries.length === 0 ? (
-
{entries.length > 0 ? `Filters hide all ${entries.length.toLocaleString()} loaded lines` : emptyMessage}
+
{entries.length > 0 ? `Filters hide all ${entries.length.toLocaleString()} loaded lines` : emptyMessage}
{entries.length === 0 && emptyCommand && (
{ void copyText(emptyCommand) }} className={`mt-2 inline-flex max-w-[80%] items-center gap-2 rounded border px-3 py-2 font-mono text-xs ${palette.border} ${palette.toolbarBg}`} title="Copy recovery command">
{emptyCommand}
@@ -1062,7 +1070,7 @@ export function LogCore({
{/* Keyboard shortcut hints */}
- {onStartStream &&
}
+ {onStartStream && (!sourceUnavailable || isStreaming) &&
}
{search.mode !== 'hide' && (
<>
@@ -1189,7 +1197,7 @@ function LogLine({
const highlighted = highlightSearchMatches(plain, searchQuery, searchIsRegex, searchIsCaseSensitive)
contentElement = (
)
@@ -1209,13 +1217,13 @@ function LogLine({
const html = ansiToHtml(entry.content)
contentElement = (
)
} else {
contentElement = (
-
+
{stripAnsi(entry.content)}
)
diff --git a/packages/k8s-ui/src/components/logs/LogToolbarSelects.tsx b/packages/k8s-ui/src/components/logs/LogToolbarSelects.tsx
index 7e7604569b..41b3dda8cc 100644
--- a/packages/k8s-ui/src/components/logs/LogToolbarSelects.tsx
+++ b/packages/k8s-ui/src/components/logs/LogToolbarSelects.tsx
@@ -50,6 +50,7 @@ interface LogRangeSelectProps {
isDark?: boolean
/** When true, the control is greyed out and non-interactive. */
disabled?: boolean
+ disabledReason?: string
}
export function LogRangeSelect({
@@ -59,10 +60,11 @@ export function LogRangeSelect({
tooltip = 'How many logs to load — by line count or time range',
isDark = true,
disabled = false,
+ disabledReason = 'Stop streaming to change the log range',
}: LogRangeSelectProps) {
const palette = getLogPalette(isDark)
return (
-
+
onChange(e.target.value)}
diff --git a/packages/k8s-ui/src/components/logs/StructuredLogLine.test.tsx b/packages/k8s-ui/src/components/logs/StructuredLogLine.test.tsx
new file mode 100644
index 0000000000..e9de2c8495
--- /dev/null
+++ b/packages/k8s-ui/src/components/logs/StructuredLogLine.test.tsx
@@ -0,0 +1,76 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { createRoot } from 'react-dom/client'
+import { describe, expect, it, vi } from 'vitest'
+import { StructuredLogLine } from './StructuredLogLine'
+
+const render = (content: string) => renderToStaticMarkup( )
+
+describe('StructuredLogLine summary', () => {
+ it('shows a CloudNativePG PostgreSQL record by its own message and severity', () => {
+ const html = render(
+ JSON.stringify({
+ level: 'info',
+ logger: 'postgres',
+ msg: 'record',
+ record: { error_severity: 'FATAL', message: 'terminating connection due to administrator command' },
+ }),
+ )
+ expect(html).toContain('terminating connection due to administrator command')
+ expect(html).toContain('FATAL')
+ expect(html).not.toContain('>record<')
+ })
+
+ it('keeps msg for ordinary structured lines', () => {
+ const html = render(JSON.stringify({ level: 'info', msg: 'Fencing status changed', record: 'not an object' }))
+ expect(html).toContain('Fencing status changed')
+ })
+
+ it.each([
+ { msg: 'record' },
+ { logger: 'worker', msg: 'record' },
+ { logger: 'postgres', msg: 'wrapper message' },
+ ])('keeps unrelated record fields out of the summary: %j', (fields) => {
+ const html = render(JSON.stringify({ level: 'info', ...fields, record: { error_severity: 'FATAL', message: 'unrelated nested text' } }))
+ expect(html).toContain(`>${fields.msg}<`)
+ expect(html).toContain('INFO')
+ expect(html).not.toContain('unrelated nested text')
+ expect(html).not.toContain('FATAL')
+ })
+
+ it('requires PostgreSQL record evidence before selecting a nested message', () => {
+ const html = render(JSON.stringify({ level: 'info', logger: 'postgres', msg: 'record', record: { message: 'unverified nested text' } }))
+ expect(html).toContain('>record<')
+ expect(html).not.toContain('unverified nested text')
+ })
+})
+
+it.each([false, true])('keeps annotations intact with word wrapping, expanded=%s', (expanded) => {
+ const html = renderToStaticMarkup( )
+ expect(html).toMatch(/inline-block whitespace-nowrap[^>]*>\{3 fields\}/)
+ expect(html).toContain('[overflow-wrap:anywhere]')
+ expect(html).not.toContain('break-all')
+})
+
+it('filters expanded strings, numbers and booleans by their displayed value', () => {
+ vi.stubGlobal('IS_REACT_ACT_ENVIRONMENT', true)
+ const filter = vi.fn()
+ const message = 'true null 123 "quoted"\nnext'
+ const host = document.createElement('div')
+ const root = createRoot(host)
+ try {
+ act(() => root.render(unsafe' })} level="info" wordWrap defaultExpanded onFilterValue={filter} />))
+ const buttons = [...host.querySelectorAll('button')]
+ for (const value of [message, '-12.5', 'true', '']) {
+ const button = buttons.find((button) => button.getAttribute('aria-label') === `Filter to lines containing ${value}`)
+ expect(button).toBeDefined()
+ act(() => button!.click())
+ expect(filter).toHaveBeenLastCalledWith(value)
+ }
+ expect(host.querySelector('script')).toBeNull()
+ expect(buttons.some((button) => button.getAttribute('aria-label') === 'Filter to lines containing null')).toBe(false)
+ } finally {
+ act(() => root.unmount())
+ }
+})
diff --git a/packages/k8s-ui/src/components/logs/StructuredLogLine.tsx b/packages/k8s-ui/src/components/logs/StructuredLogLine.tsx
index 6f2075d8f2..3b56b66842 100644
--- a/packages/k8s-ui/src/components/logs/StructuredLogLine.tsx
+++ b/packages/k8s-ui/src/components/logs/StructuredLogLine.tsx
@@ -1,8 +1,8 @@
import { useState, useMemo } from 'react'
import { ChevronRight, ChevronDown, Filter } from 'lucide-react'
import type { LogLevel } from './useLogBuffer'
-import { selectLevelField } from '../../utils/log-level'
-import { unescapeJsonStrings, parseLogfmt } from '../../utils/log-format'
+import { selectLevelField, selectPostgresRecord } from '../../utils/log-level'
+import { unescapeJsonStrings, parseLogfmt, tokenizeJson } from '../../utils/log-format'
import { getLogPalette, getLogLevelColor, type LogPalette } from './log-palette'
interface StructuredLogLineProps {
@@ -44,7 +44,7 @@ export function StructuredLogLine({ content, level, wordWrap, isLogfmt, defaultE
if (!parsed) {
return (
-
+
{content}
)
@@ -63,11 +63,11 @@ export function StructuredLogLine({ content, level, wordWrap, isLogfmt, defaultE
// Collapsed: entire summary line is clickable
{chevron}
- {`{${fieldCount} fields}`}
+ {`{${fieldCount} fields}`}
) : (
// Expanded: summary header is clickable to collapse, JSON content is selectable
@@ -78,9 +78,9 @@ export function StructuredLogLine({ content, level, wordWrap, isLogfmt, defaultE
>
{chevron}
- {`{${fieldCount} fields}`}
+ {`{${fieldCount} fields}`}
-
+
{isLogfmt ? (
) : (
@@ -129,48 +129,40 @@ function FilterableValue({
* and emits React nodes so the hover chip can be wired per value.
*/
function JsonExpanded({ text, onFilterValue, palette }: { text: string; onFilterValue?: (v: string) => void; palette: LogPalette }) {
- const tokenRe = /("(?:\\.|[^"\\])*")\s*:|("(?:\\.|[^"\\])*")|(-?\b\d+(?:\.\d+)?(?:[eE][+-]?\d+)?\b)|\b(true|false)\b|\b(null)\b/g
const nodes: React.ReactNode[] = []
- let lastIndex = 0
- let match: RegExpExecArray | null
let idx = 0
- while ((match = tokenRe.exec(text)) !== null) {
- if (match.index > lastIndex) {
- nodes.push({text.slice(lastIndex, match.index)} )
- }
- const [, key, str, num, bool, nil] = match
- if (key !== undefined) {
- nodes.push({key} )
+ for (const { type, value } of tokenizeJson(text)) {
+ if (type === 'key') {
+ nodes.push({value} )
nodes.push(: )
- } else if (str !== undefined) {
+ } else if (type === 'string') {
// Unescape the quoted string for the filter value (users expect to filter on
// the displayed string, not JSON-escaped bytes).
- const inner = str.slice(1, -1).replace(/\\"/g, '"').replace(/\\\\/g, '\\')
+ const inner = value.slice(1, -1).replace(/\\"/g, '"').replace(/\\\\/g, '\\')
nodes.push(
)
- } else if (num !== undefined) {
+ } else if (type === 'number') {
nodes.push(
-
+
)
- } else if (bool !== undefined) {
+ } else if (type === 'boolean') {
nodes.push(
-
+
)
- } else if (nil !== undefined) {
- nodes.push({nil} )
+ } else if (type === 'null') {
+ nodes.push({value} )
+ } else {
+ nodes.push({value} )
}
- lastIndex = tokenRe.lastIndex
- }
- if (lastIndex < text.length) {
- nodes.push({text.slice(lastIndex)} )
}
return <>{nodes}>
}
function SummaryLine({ obj, level, palette }: { obj: Record; level: LogLevel; palette: LogPalette }) {
const lvl = selectLevelField(obj)?.raw
- const msg = obj.msg ?? obj.message
+ const pgRecord = selectPostgresRecord(obj)
+ const msg = pgRecord?.record.message ?? obj.msg ?? obj.message
const rawErr = obj.error ?? obj.err
const err = typeof rawErr === 'string'
? rawErr
diff --git a/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.test.tsx b/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.test.tsx
new file mode 100644
index 0000000000..73dcf2e1c8
--- /dev/null
+++ b/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.test.tsx
@@ -0,0 +1,36 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot } from 'react-dom/client'
+import { expect, it, vi } from 'vitest'
+import { WorkloadLogsViewer, type WorkloadLogsResult } from './WorkloadLogsViewer'
+vi.mock('../ui/Toast', () => ({ useToast: () => ({ showError: vi.fn(), showSuccess: vi.fn() }) }))
+Object.assign(globalThis, { IS_REACT_ACT_ENVIRONMENT: true })
+it('uses a host empty state and disables source controls until a source is read', async () => {
+ const host = document.createElement('div'); const root = createRoot(host)
+ let result: WorkloadLogsResult = { pods: [], logs: [], emptyMessage: 'No instance Pods yet' }
+ const stream = vi.fn()
+ const fetchAll = vi.fn(async () => result)
+ await act(async () => root.render(The first instance cannot start.
Overview’s problem >} />))
+ expect(host.textContent).toContain('The first instance cannot start')
+ expect(host.querySelector('a')?.getAttribute('href')).toBe('?tab=overview')
+ expect(host.querySelector('fieldset')?.disabled).toBe(true)
+ expect([...host.querySelectorAll('select')].every((el) => el.matches(':disabled'))).toBe(true)
+ const streamButton = [...host.querySelectorAll('button')].find((el) => el.textContent === 'Stream')!
+ expect(streamButton.disabled).toBe(true)
+ act(() => streamButton.click())
+ expect(stream).not.toHaveBeenCalled()
+ result = { pods: [{ name: 'analytics-1', containers: ['postgres'], ready: true }], logs: [] }
+ await act(async () => [...host.querySelectorAll('button')].find((el) => el.querySelector('.lucide-rotate-ccw'))!.click())
+ expect(host.querySelector('fieldset')?.disabled).toBe(false)
+ expect(streamButton.disabled).toBe(false)
+ expect(host.textContent).not.toContain('The first instance cannot start')
+ await act(async () => root.unmount())
+})
+
+it('keeps streaming available for hosts that wait for Pods to appear', async () => {
+ const host = document.createElement('div'); const root = createRoot(host)
+ await act(async () => root.render( ({ pods: [], logs: [] })} createStream={vi.fn()} />))
+ expect([...host.querySelectorAll('button')].find((el) => el.textContent === 'Stream')?.disabled).toBe(false)
+ expect(host.querySelector('fieldset')?.disabled).toBe(false)
+ await act(async () => root.unmount())
+})
diff --git a/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.tsx b/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.tsx
index 7429886c6f..477b793f45 100644
--- a/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.tsx
+++ b/packages/k8s-ui/src/components/logs/WorkloadLogsViewer.tsx
@@ -1,5 +1,5 @@
-import { useState, useEffect, useCallback, useMemo, useRef } from 'react'
-import { Filter, ChevronDown } from 'lucide-react'
+import { useState, useEffect, useCallback, useMemo, useRef, type ReactNode } from 'react'
+import { Filter, ChevronDown, AlertTriangle } from 'lucide-react'
import { parseLogRange } from '../../utils/log-format'
import { triggerDownload } from '../../utils/download'
import { useLogBuffer } from './useLogBuffer'
@@ -10,6 +10,7 @@ import type { LogExportPayload } from '../../utils/log-export'
import type { LogPalette } from './log-palette'
import type { WorkloadPodInfo } from '../../types'
import { useToast } from '../ui/Toast'
+import { toneTextClass } from '../ui/status-tone'
export interface WorkloadRawLog {
pod: string
@@ -38,6 +39,8 @@ export interface WorkloadLogsResult {
}
export interface WorkloadLogsViewerProps {
+ /** Disable streaming and the source filters while no log source exists. Off by default: streams can wait for Pods to appear. */
+ disableSourceControlsWithoutSource?: boolean
/** Workload name — used for the download filename */
name: string
/**
@@ -65,14 +68,18 @@ export interface WorkloadLogsViewerProps {
autoStream?: boolean
/** Pods selected when the pod list first loads; all pods when empty or none match. */
initialPods?: string[]
+ /** Container selected on mount; all containers when unset. */
+ initialContainer?: string
+ /** Rendered only when the loaded source list is empty. */
+ emptySourceState?: ReactNode
}
-export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownload, forceDark, defaultDark, autoStream = false, initialPods }: WorkloadLogsViewerProps) {
+export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownload, forceDark, defaultDark, autoStream = false, initialPods, initialContainer, emptySourceState, disableSourceControlsWithoutSource = false }: WorkloadLogsViewerProps) {
const initialSelection = (names: string[]) => {
const wanted = names.filter((n) => initialPods?.includes(n))
return new Set(wanted.length > 0 ? wanted : names)
}
- const [selectedContainer, setSelectedContainer] = useState('')
+ const [selectedContainer, setSelectedContainer] = useState(initialContainer ?? '')
const [pods, setPods] = useState([])
const [selectedPods, setSelectedPods] = useState>(new Set())
const [isLoading, setIsLoading] = useState(false)
@@ -92,6 +99,7 @@ export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownl
const { isStreaming, streamError, connecting, startStreaming, stopStreaming } = useLogStream()
const willAutoStream = autoStream && !!createStream
+ const sourceUnavailable = disableSourceControlsWithoutSource && pods.length === 0
// null sentinel so the initial selectedContainer ('' = all) still arms once.
const autoStartedForRef = useRef(null)
const userStoppedRef = useRef(false)
@@ -312,10 +320,11 @@ export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownl
const renderToolbarExtra = ({ isDark, palette }: { isDark: boolean; palette: LogPalette }) => (
<>
{/* Pod filter */}
+
setShowPodFilter(v => !v)}
- className={`flex items-center gap-1.5 px-2 py-1.5 text-xs rounded transition-colors ${
+ className={`flex items-center gap-1.5 px-2 py-1.5 text-xs rounded whitespace-nowrap transition-colors ${
showPodFilter ? palette.toolbarActive : `${palette.elevatedBg} ${palette.textSecondary} ${palette.hoverBg}`
}`}
>
@@ -381,8 +390,10 @@ export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownl
lineOptions={[50, 100, 500, 1000]}
tooltip="How many logs to load per pod — by line count or time range"
isDark={isDark}
- disabled={isStreaming}
+ disabled={isStreaming || sourceUnavailable}
+ disabledReason={sourceUnavailable ? 'No log source is available yet' : undefined}
/>
+
>
)
@@ -392,8 +403,23 @@ export function WorkloadLogsViewer({ name, fetchAll, createStream, overrideDownl
return (
- {notice &&
{notice}
}
- {capturedAt &&
Snapshot captured {new Date(capturedAt).toLocaleTimeString()}
}
+ {/* One status line. A notice says what the snapshot could not show, so it
+ stays marked; the capture time alone is quiet. */}
+ {(notice || capturedAt) && (
+
+ {capturedAt &&
Snapshot {new Date(capturedAt).toLocaleTimeString()} }
+ {capturedAt && notice &&
· }
+ {notice && (
+
+
+ {notice}
+
+ )}
+
+ )}
{(fetchError || streamError) && entries.length > 0 &&
{fetchError || streamError} · Previously loaded logs remain below.
}
{
expect(detectLogLevel(JSON.stringify({ level: 'error', msg: 'request failed', record: { error_severity: 20 } }))).toBe('error')
expect(detectLogLevel(JSON.stringify({ level: 'error', msg: 'x', record: { error_severity: 'minor' } }))).toBe('error')
})
+
+ it.each([
+ { msg: 'record' },
+ { logger: 'worker', msg: 'record' },
+ { logger: 'postgres', msg: 'wrapper message' },
+ ])('does not override an unrelated workload severity: %j', (fields) => {
+ expect(detectLogLevel(JSON.stringify({ level: 'info', ...fields, record: { error_severity: 'FATAL', message: 'unrelated nested text' } }))).toBe('info')
+ })
})
diff --git a/packages/k8s-ui/src/components/problems/index.ts b/packages/k8s-ui/src/components/problems/index.ts
new file mode 100644
index 0000000000..94e9292d32
--- /dev/null
+++ b/packages/k8s-ui/src/components/problems/index.ts
@@ -0,0 +1 @@
+export * from './problems'
diff --git a/packages/k8s-ui/src/components/problems/problems.test.tsx b/packages/k8s-ui/src/components/problems/problems.test.tsx
new file mode 100644
index 0000000000..67293964e3
--- /dev/null
+++ b/packages/k8s-ui/src/components/problems/problems.test.tsx
@@ -0,0 +1,38 @@
+import { describe, expect, it } from 'vitest'
+import { renderToStaticMarkup } from 'react-dom/server'
+import { ProblemMeta, ProblemCallout, ProblemList, problemOriginLabel, type WorkspaceProblem } from './problems'
+
+const problem = (kind: string): WorkspaceProblem => ({
+ id: 'p',
+ severity: 'warning',
+ category: 'availability',
+ title: 'Something needs a look',
+ subject: { kind, group: 'example.io', namespace: 'ns', name: 'child-1' },
+ source: 'issue',
+})
+
+describe('ProblemMeta', () => {
+ it('names the subject only when it is not the workspace root kind', () => {
+ expect(renderToStaticMarkup( )).toContain('Backup')
+ expect(renderToStaticMarkup( )).not.toContain('child-1')
+ })
+})
+
+describe('problemOriginLabel', () => {
+ it('says who measured a measurement, and never calls an issue generic', () => {
+ expect(problemOriginLabel({ ...problem('Cluster'), source: 'measurement', measuredBy: 'Prometheus' }).label).toBe('Measured by Prometheus')
+ expect(problemOriginLabel(problem('Cluster')).label).toBe('Detected by Radar')
+ })
+})
+
+it('keeps original scheduler evidence behind a closed disclosure in callouts and lists', () => {
+ const p = { ...problem('Pod'), detail: 'both nodes have reached their Pod limit', rawDetail: '0/2 nodes are available: 2 Too many pods.' }
+ for (const node of [ , ]) {
+ const html = renderToStaticMarkup(node)
+ expect(html).toContain(p.detail)
+ expect(html).toContain(p.rawDetail)
+ expect(html).toContain('Scheduler message')
+ expect(html).toContain('aria-expanded="false"')
+ }
+ expect(renderToStaticMarkup( )).not.toContain('Scheduler message')
+})
diff --git a/packages/k8s-ui/src/components/problems/problems.tsx b/packages/k8s-ui/src/components/problems/problems.tsx
new file mode 100644
index 0000000000..326338f680
--- /dev/null
+++ b/packages/k8s-ui/src/components/problems/problems.tsx
@@ -0,0 +1,179 @@
+import { createContext, useContext, type ReactNode } from 'react'
+import { clsx } from 'clsx'
+import type { HealthLevel } from '../resources/resource-utils'
+import { StatusDot, toneTextClass } from '../ui/status-tone'
+import { Tooltip } from '../ui/Tooltip'
+import { AlertBanner } from '../ui/drawer-components'
+import { FoldSection } from '../ui/FoldSection'
+import { RefLink, type NavigateToRef } from '../ui/RefLink'
+
+/** Where a problem's evidence comes from, in user terms. */
+export interface ProblemOrigin {
+ label: string
+ /** The exact field or condition, shown on hover. */
+ detail?: string
+}
+
+/**
+ * Something about a workspace object that needs a look. `C` is the
+ * integration's own category set.
+ */
+export interface WorkspaceProblem {
+ /** Stable identity for keys. */
+ id: string
+ severity: 'critical' | 'warning' | 'posture'
+ category: C
+ title: string
+ detail?: string
+ rawDetail?: string
+ /** The object the evidence is about: the workspace's root object or one of its children. */
+ subject: { kind: string; group: string; namespace: string; name: string }
+ /**
+ * issue: the Issues engine. audit: a best-practice check. measurement:
+ * derived from a reading only callers holding its grants receive.
+ */
+ source: 'issue' | 'audit' | 'measurement'
+ /** What took the measurement, e.g. "Prometheus" (shown as "Measured by Prometheus"). */
+ measuredBy?: string
+ /** The measurement's series were matched to the subject by name only (see `measuredBy`). */
+ unverifiedMatch?: boolean
+ /** How it was measured (queries, metric names), shown on hover over the source. */
+ sourceDetail?: string
+ /** A shorter headline for tight places (a fleet cell); `title` stays the precise one. */
+ shortTitle?: string
+ /** Where an issue's evidence comes from. */
+ origin?: ProblemOrigin
+ /** Other objects the same problem is about, e.g. earlier runs that failed the same way. */
+ alsoAbout?: { kind: string; name: string }[]
+}
+
+export const PROBLEM_TONE: Record = {
+ critical: 'unhealthy',
+ warning: 'degraded',
+ posture: 'neutral',
+}
+
+const PROBLEM_VARIANT: Record = {
+ critical: 'error',
+ warning: 'warning',
+ posture: 'info',
+}
+
+/** A problem's provenance label: where its evidence comes from, never a generic "Radar issue". */
+export function problemOriginLabel(problem: WorkspaceProblem): ProblemOrigin {
+ switch (problem.source) {
+ case 'audit':
+ return { label: 'Best-practice check', detail: problem.sourceDetail }
+ case 'measurement':
+ return { label: problem.measuredBy ? `Measured by ${problem.measuredBy}` : 'Measured', detail: problem.sourceDetail }
+ }
+ return problem.origin ?? { label: 'Detected by Radar' }
+}
+
+/**
+ * How a host opens a problem on its Issues page. Supplied by context so every
+ * summary and drawer gets the link without threading a prop through each.
+ */
+export const OpenIssueContext = createContext<((problem: WorkspaceProblem) => void) | undefined>(undefined)
+
+export function ProblemMeta({
+ problem,
+ rootKind,
+ onNavigate,
+ subjectIsSelf,
+ children,
+}: {
+ problem: WorkspaceProblem
+ /** The workspace's root kind: a problem about another kind names its subject. */
+ rootKind: string
+ onNavigate?: NavigateToRef
+ subjectIsSelf?: boolean
+ children?: ReactNode
+}) {
+ const openIssue = useContext(OpenIssueContext)
+ const origin = problemOriginLabel(problem)
+ const aboutChild = !subjectIsSelf && problem.subject.kind !== rootKind
+ return (
+
+ {aboutChild && (
+
+ {problem.subject.kind}{' '}
+
+ {problem.alsoAbout && problem.alsoAbout.length > 0 && ' '}
+ {problem.alsoAbout && problem.alsoAbout.length > 0 && (
+
+ {problem.alsoAbout.map((o) => (
+
+ {o.kind} {o.name}
+
+ ))}
+
+ }
+ >
+ and {problem.alsoAbout.length} more
+
+ )}
+
+ )}
+
+ {origin.label}
+
+ {openIssue && problem.source === 'issue' && (
+ openIssue(problem)} className="whitespace-nowrap text-accent-text hover:underline">
+ See in Issues →
+
+ )}
+ {children}
+
+ )
+}
+
+export function ProblemCallout({
+ problem,
+ rootKind,
+ more,
+ onNavigate,
+ action,
+ subjectIsSelf,
+}: {
+ problem: WorkspaceProblem
+ rootKind: string
+ more?: ReactNode
+ onNavigate?: NavigateToRef
+ action?: ReactNode
+ /** The callout sits on the subject's own page, so linking to it would loop. */
+ subjectIsSelf?: boolean
+}) {
+ return (
+
+ {problem.rawDetail && {problem.rawDetail}
}
+
+ {action}
+ {more}
+
+
+ )
+}
+
+/** The problems a callout does not show, as a compact list with the callout's tone, title and source. */
+export function ProblemList({ problems, rootKind, onNavigate }: { problems: WorkspaceProblem[]; rootKind: string; onNavigate?: NavigateToRef }) {
+ return (
+
+ {problems.map((p) => (
+
+
+
+
+
+
{p.title}
+ {p.detail &&
{p.detail}
}
+ {p.rawDetail &&
{p.rawDetail}
}
+
+
+
+ ))}
+
+ )
+}
diff --git a/packages/k8s-ui/src/components/resources/ResourcesSidebar.keyboard.test.tsx b/packages/k8s-ui/src/components/resources/ResourcesSidebar.keyboard.test.tsx
new file mode 100644
index 0000000000..4dbe86c262
--- /dev/null
+++ b/packages/k8s-ui/src/components/resources/ResourcesSidebar.keyboard.test.tsx
@@ -0,0 +1,42 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot } from 'react-dom/client'
+import { afterEach, expect, it, vi } from 'vitest'
+import { ResourcesSidebar } from './ResourcesSidebar'
+;(globalThis as any).IS_REACT_ACT_ENVIRONMENT = true
+window.matchMedia = vi.fn().mockReturnValue({ matches: false, addEventListener: vi.fn(), removeEventListener: vi.fn() })
+let root: ReturnType
+let host: HTMLDivElement
+afterEach(() => { act(() => root?.unmount()); host?.remove() })
+it('navigates Favorites, Views and grouped kinds in their visible order', () => {
+ Element.prototype.scrollIntoView = vi.fn()
+ const view = vi.fn(), kind = vi.fn()
+ host = document.createElement('div'); document.body.append(host); root = createRoot(host)
+ const cluster = { group: 'postgresql.cnpg.io', version: 'v1', kind: 'Cluster', name: 'clusters', namespaced: true, isCrd: true, verbs: ['list'] }
+ act(() => root.render( ))
+ const category = [...host.querySelectorAll('button')].find((b) => b.textContent?.includes('CloudNativePG'))!
+ if (category.getAttribute('aria-expanded') !== 'true') act(() => category.click())
+ const input = host.querySelector('input')!
+ const key = (name: string) => act(() => input.dispatchEvent(new KeyboardEvent('keydown', { key: name, bubbles: true })))
+ key('ArrowDown'); key('ArrowDown'); key('Enter')
+ expect(view).toHaveBeenCalledTimes(1)
+ expect(kind).not.toHaveBeenCalled()
+ act(() => [...host.querySelectorAll('button')].find((b) => b.textContent === 'Resource kinds')!.click())
+ key('ArrowDown'); key('ArrowDown'); key('ArrowDown'); key('Enter')
+ expect(kind).toHaveBeenCalledWith(expect.objectContaining({ name: 'clusters', group: 'postgresql.cnpg.io' }))
+})
+it('opens a matching View with Enter when search finds no kind', () => {
+ Element.prototype.scrollIntoView = vi.fn()
+ const view = vi.fn()
+ host = document.createElement('div'); document.body.append(host); root = createRoot(host)
+ act(() => root.render( {}} apiResources={[{ group: 'postgresql.cnpg.io', version: 'v1', kind: 'Cluster', name: 'clusters', namespaced: true, isCrd: true, verbs: ['list'] }]} resourceCounts={{ 'postgresql.cnpg.io/Cluster': 1 }} categoryWorkspaces={{ CloudNativePG: { destinations: [{ id: 'operator', label: 'Operator', onSelect: view }], defaultKindsCollapsed: true } }} />))
+ const input = host.querySelector('input')!
+ act(() => {
+ Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, 'value')!.set!.call(input, 'operator')
+ input.dispatchEvent(new Event('input', { bubbles: true }))
+ })
+ expect(host.textContent).toContain('Operator')
+ expect(host.textContent).not.toContain('>Cluster<')
+ act(() => input.dispatchEvent(new KeyboardEvent('keydown', { key: 'Enter', bubbles: true })))
+ expect(view).toHaveBeenCalledTimes(1)
+})
diff --git a/packages/k8s-ui/src/components/resources/ResourcesSidebar.test.tsx b/packages/k8s-ui/src/components/resources/ResourcesSidebar.test.tsx
index f76df8d7f1..89743b63e7 100644
--- a/packages/k8s-ui/src/components/resources/ResourcesSidebar.test.tsx
+++ b/packages/k8s-ui/src/components/resources/ResourcesSidebar.test.tsx
@@ -1,7 +1,7 @@
import { describe, expect, it } from 'vitest'
import { renderToString } from 'react-dom/server'
import type { APIResource } from '../../types'
-import { rawCRDGroupTitle, resourceMatchesSidebarFilter, ResourcesSidebar } from './ResourcesSidebar'
+import { rawCRDGroupTitle, resourceMatchesSidebarFilter, ResourcesSidebar, sidebarDestinationsMatching } from './ResourcesSidebar'
const sqlInstance: APIResource = {
group: 'sql.cnrm.cloud.google.com',
@@ -144,7 +144,7 @@ describe('ResourcesSidebar category workspaces', () => {
}}
/>
)
- expect(html).toContain('Workspace')
+ expect(html).toContain('Views')
expect(html).toContain('Overview')
expect(html).toContain('pg-orders')
expect(html).toContain('Counts for namespace payments')
@@ -153,6 +153,21 @@ describe('ResourcesSidebar category workspaces', () => {
expect(html).not.toContain('selection-strong selection-text">Pod')
})
+ it('marks a count taken over partly readable data as a lower bound, and never shows its zero as none', () => {
+ const render = (count: number) =>
+ renderToString(
+ {}}
+ apiResources={[cnpgCluster]}
+ resourceCounts={{ 'postgresql.cnpg.io/Cluster': 1 }}
+ categoryWorkspaces={{ CloudNativePG: { destinations: [{ id: 'overview', label: 'Overview', count, countLowerBound: true, onSelect: () => {} }] } }}
+ />,
+ )
+ expect(render(2)).toMatch(/≥()?2/)
+ expect(render(0)).toContain('–')
+ })
+
it('keeps a workspace category visible when it has no resources', () => {
const html = renderToString(
{
expect(html).toContain('barmancloud.cnpg.io')
})
})
+
+describe('sidebarDestinationsMatching', () => {
+ const views = ['Clusters', 'Backups', 'Declarations', 'Pooling', 'Operator'].map((label) => ({ id: label, label, onSelect: () => {} }))
+ const resources = [{ group: 'postgresql.cnpg.io' }, { group: 'barmancloud.cnpg.io' }]
+ const labels = (term: string) => sidebarDestinationsMatching('CloudNativePG', resources, views, term).map((d) => d.label)
+
+ it('keeps every view when the term names the category or one of its API groups', () => {
+ expect(labels('cloud')).toEqual(['Clusters', 'Backups', 'Declarations', 'Pooling', 'Operator'])
+ expect(labels('cnpg')).toEqual(['Clusters', 'Backups', 'Declarations', 'Pooling', 'Operator'])
+ })
+ it('keeps the views whose label matches otherwise', () => {
+ expect(labels('backup')).toEqual(['Backups'])
+ expect(labels('declar')).toEqual(['Declarations'])
+ expect(labels('deployment')).toEqual([])
+ })
+ it('keeps every view without a term', () => {
+ expect(labels(' ')).toHaveLength(5)
+ })
+})
diff --git a/packages/k8s-ui/src/components/resources/ResourcesSidebar.tsx b/packages/k8s-ui/src/components/resources/ResourcesSidebar.tsx
index c06bc32e2b..8904ce9078 100644
--- a/packages/k8s-ui/src/components/resources/ResourcesSidebar.tsx
+++ b/packages/k8s-ui/src/components/resources/ResourcesSidebar.tsx
@@ -13,6 +13,7 @@ import type { APIResource } from '../../types'
import { categorizeResources, CORE_RESOURCES } from '../../utils/api-resources'
import { getResourceIcon } from '../../utils/resource-icons'
import { Tooltip } from '../ui/Tooltip'
+import { certaintyGlyph } from '../facts'
import { Input } from '../ui/Input'
// Selected resource type info (need both name for API and kind for display)
@@ -36,6 +37,8 @@ export interface SidebarCategoryDestination {
icon?: ComponentType<{ className?: string }>
/** Problem count. `undefined` renders no badge; `null` renders the unknown dash. */
count?: number | null
+ /** The count was taken over data read only in part: rendered "≥N", and a zero is unknown rather than none. */
+ countLowerBound?: boolean
countTitle?: string
active?: boolean
/** The object currently open under this destination, nested beneath it. */
@@ -97,6 +100,24 @@ const CORE_RESOURCE_TYPES = [
{ kind: 'hpas', label: 'HPAs' },
] as const
+/**
+ * The workspace views a filter keeps: all of them when the term names the
+ * category or one of its API groups ("cloud", "cnpg"), otherwise the views
+ * whose label matches. The views are the way into a workspace, so a filter
+ * that finds the category's kinds must not hide them.
+ */
+export function sidebarDestinationsMatching(
+ categoryName: string,
+ resources: Pick[],
+ destinations: SidebarCategoryDestination[],
+ term: string,
+): SidebarCategoryDestination[] {
+ const normalized = term.trim().toLowerCase()
+ if (!normalized) return destinations
+ if (categoryName.toLowerCase().includes(normalized) || resources.some((r) => r.group.toLowerCase().includes(normalized))) return destinations
+ return destinations.filter((d) => d.label.toLowerCase().includes(normalized))
+}
+
export function resourceMatchesSidebarFilter(resource: Pick, term: string): boolean {
const normalized = term.trim().toLowerCase()
if (!normalized) return true
@@ -427,14 +448,16 @@ export function ResourcesSidebar({
// If the group name matches, show all its resources
if (categoryMatches || rawGroupMatches) return category
const matchingResources = category.visibleResources.filter((resource: APIResource) => resourceMatchesSidebarFilter(resource, term))
- if (matchingResources.length === 0) return null
+ const ws = categoryWorkspaces?.[category.name]
+ const matchingViews = ws ? sidebarDestinationsMatching(category.name, category.resources, ws.destinations, term) : []
+ if (matchingResources.length === 0 && matchingViews.length === 0) return null
return {
...category,
visibleResources: matchingResources,
}
})
.filter(Boolean) as typeof sortedCategories
- }, [sortedCategories, kindFilter])
+ }, [sortedCategories, kindFilter, categoryWorkspaces])
// Auto-expand all categories when filtering
const isKindFiltering = kindFilter.trim().length > 0
@@ -455,26 +478,23 @@ export function ResourcesSidebar({
})
}
- // --- Keyboard navigation ---
- // Flat list of all navigable kinds in the order they appear in the sidebar.
- const flatVisibleKinds = useMemo(() => {
- const kinds: SelectedKindInfo[] = []
+ const visibleResults = useMemo(() => {
+ const results: ({ type: 'kind'; kind: SelectedKindInfo } | { type: 'view'; category: string; destination: SidebarCategoryDestination })[] = []
if (favoritesExpanded) {
- for (const p of pinned) {
- kinds.push({ name: p.name, kind: p.kind, group: p.group })
- }
+ for (const kind of pinned) results.push({ type: 'kind', kind })
}
- if (filteredCategories) {
- for (const cat of filteredCategories) {
- if (effectiveExpandedCategories.has(cat.name) && (isKindFiltering || kindsOpen(cat.name))) {
- for (const r of cat.visibleResources) {
- kinds.push({ name: r.name, kind: r.kind, group: r.group })
- }
- }
+ for (const cat of filteredCategories ?? []) {
+ if (!effectiveExpandedCategories.has(cat.name)) continue
+ const workspace = categoryWorkspaces?.[cat.name]
+ for (const destination of workspace ? sidebarDestinationsMatching(cat.name, cat.resources, workspace.destinations, kindFilter) : []) {
+ results.push({ type: 'view', category: cat.name, destination })
+ }
+ if (isKindFiltering || kindsOpen(cat.name)) {
+ for (const kind of workspace ? sortByGroup(cat.visibleResources) : cat.visibleResources) results.push({ type: 'kind', kind })
}
}
- return kinds
- }, [favoritesExpanded, pinned, filteredCategories, effectiveExpandedCategories, kindsOpenOverrides, categoryWorkspaces, activeDestinationCategory]) // eslint-disable-line react-hooks/exhaustive-deps
+ return results
+ }, [favoritesExpanded, pinned, filteredCategories, effectiveExpandedCategories, kindsOpenOverrides, categoryWorkspaces, activeDestinationCategory, isKindFiltering, kindFilter]) // eslint-disable-line react-hooks/exhaustive-deps
const [highlightedIndex, setHighlightedIndex] = useState(-1)
// Reset highlight when the filter or kind list changes
@@ -482,10 +502,12 @@ export function ResourcesSidebar({
setHighlightedIndex(kindFilter ? 0 : -1)
}, [kindFilter]) // eslint-disable-line react-hooks/exhaustive-deps
- const highlightedKind = highlightedIndex >= 0 && highlightedIndex < flatVisibleKinds.length
- ? flatVisibleKinds[highlightedIndex]
+ const highlightedResult = highlightedIndex >= 0 && highlightedIndex < visibleResults.length
+ ? visibleResults[highlightedIndex]
: null
+ const highlightedKind = highlightedResult?.type === 'kind' ? highlightedResult.kind : null
+
// Scroll the highlighted kind button into view
const highlightedRef = useRef(null)
useEffect(() => {
@@ -501,18 +523,19 @@ export function ResourcesSidebar({
;(e.target as HTMLInputElement).blur()
} else if (e.key === 'ArrowDown') {
e.preventDefault()
- setHighlightedIndex(prev => Math.min(prev + 1, flatVisibleKinds.length - 1))
+ setHighlightedIndex(prev => Math.min(prev + 1, visibleResults.length - 1))
} else if (e.key === 'ArrowUp') {
e.preventDefault()
setHighlightedIndex(prev => Math.max(prev - 1, 0))
- } else if (e.key === 'Enter' && highlightedKind) {
+ } else if (e.key === 'Enter' && highlightedResult) {
e.preventDefault()
- selectKind(highlightedKind)
+ if (highlightedResult.type === 'view') highlightedResult.destination.onSelect()
+ else selectKind(highlightedResult.kind)
setHighlightedIndex(-1)
setKindFilter('')
onKindNavigated?.()
}
- }, [flatVisibleKinds.length, highlightedKind, selectKind, onKindNavigated]) // eslint-disable-line react-hooks/exhaustive-deps
+ }, [visibleResults.length, highlightedResult, selectKind, onKindNavigated]) // eslint-disable-line react-hooks/exhaustive-deps
const isKindHighlighted = useCallback((name: string, group: string) => {
return highlightedKind?.name === name && highlightedKind?.group === group
@@ -618,6 +641,7 @@ export function ResourcesSidebar({
const rawGroupTitle = rawCRDGroupTitle(category.resources)
const workspace = categoryWorkspaces?.[category.name]
const showKinds = isKindFiltering || kindsOpen(category.name)
+ const views = workspace ? sidebarDestinationsMatching(category.name, category.resources, workspace.destinations, kindFilter) : []
return (
- {workspace && !isKindFiltering && (
+ {workspace && views.length > 0 && (
0}
kindsOpen={showKinds}
onToggleKinds={() => toggleKinds(category.name)}
kindsPanelId={disclosurePanelId(categoryPanelBase, `${category.name}-kinds`)}
@@ -776,45 +805,63 @@ function groupHeaderBefore(resources: APIResource[], index: number): string | nu
return group || 'core'
}
+// While the sidebar is filtered, `destinations` are the views that match and
+// the kinds below are forced open, so they get a label rather than a toggle.
function WorkspaceDestinations({
workspace,
+ destinations,
+ filtering,
+ hasKinds,
kindsOpen,
onToggleKinds,
kindsPanelId,
+ highlightedId,
+ highlightedRef,
}: {
workspace: SidebarCategoryWorkspace
+ highlightedId?: string
+ highlightedRef: React.RefObject
+ destinations: SidebarCategoryDestination[]
+ filtering: boolean
+ hasKinds: boolean
kindsOpen: boolean
onToggleKinds: () => void
kindsPanelId: string
}) {
return (
-
Workspace
- {workspace.destinations.map((d) => {
+
Views
+ {destinations.map((d) => {
const Icon = d.icon
return (
{Icon && }
{d.label}
- {d.count === null ? (
+ {d.count === null || (d.count === 0 && d.countLowerBound) ? (
–
) : d.count !== undefined && d.count > 0 ? (
- {d.count}
+
+ {d.countLowerBound && certaintyGlyph('lower_bound')}
+ {d.count}
+
) : null}
@@ -833,15 +880,19 @@ function WorkspaceDestinations({
{workspace.scopeNote && (
{workspace.scopeNote}
)}
-
-
- Resource kinds
-
+ {filtering ? (
+ hasKinds &&
Resource kinds
+ ) : (
+
+
+ Resource kinds
+
+ )}
)
}
diff --git a/packages/k8s-ui/src/components/resources/ResourcesView.kind-view.test.tsx b/packages/k8s-ui/src/components/resources/ResourcesView.kind-view.test.tsx
new file mode 100644
index 0000000000..faac45fae1
--- /dev/null
+++ b/packages/k8s-ui/src/components/resources/ResourcesView.kind-view.test.tsx
@@ -0,0 +1,115 @@
+// @vitest-environment jsdom
+import { act } from 'react'
+import { createRoot, type Root } from 'react-dom/client'
+import { afterEach, describe, expect, it, vi } from 'vitest'
+import type { APIResource, SelectedResource } from '../../types'
+import { KeyboardShortcutProvider, useActiveShortcuts, type KeyboardShortcut } from '../../hooks/useKeyboardShortcuts'
+import { ResourcesView } from './ResourcesView'
+
+vi.stubGlobal('IS_REACT_ACT_ENVIRONMENT', true)
+vi.stubGlobal('matchMedia', (query: string) => ({
+ matches: false,
+ media: query,
+ onchange: null,
+ addEventListener: () => {},
+ removeEventListener: () => {},
+ addListener: () => {},
+ removeListener: () => {},
+ dispatchEvent: () => false,
+}))
+Element.prototype.scrollIntoView = () => {}
+vi.stubGlobal(
+ 'ResizeObserver',
+ class {
+ observe() {}
+ unobserve() {}
+ disconnect() {}
+ },
+)
+
+const cnpgClusters: APIResource = {
+ group: 'postgresql.cnpg.io',
+ version: 'v1',
+ kind: 'Cluster',
+ name: 'clusters',
+ namespaced: true,
+ isCrd: true,
+ verbs: ['list', 'get', 'watch'],
+}
+
+const pg = (name: string): SelectedResource => ({ kind: 'clusters', group: 'postgresql.cnpg.io', namespace: 'pg', name })
+
+let root: Root | null = null
+afterEach(async () => {
+ await act(async () => root?.unmount())
+ root = null
+})
+
+function mount(search: string, selectedResource: SelectedResource | null) {
+ // The page reads the deep link from the real URL on mount.
+ window.history.replaceState(null, '', `/resources/clusters${search}`)
+ const element = document.createElement('div')
+ document.body.appendChild(element)
+ root = createRoot(element)
+ const onNavigate = vi.fn()
+ const onResourceClick = vi.fn()
+ let shortcuts: KeyboardShortcut[] = []
+ function Probe() {
+ shortcuts = useActiveShortcuts()
+ return null
+ }
+ const render = async (selected: SelectedResource | null) =>
+ act(async () => {
+ root!.render(
+
+
+ (kind.name === 'clusters' && kind.group === 'postgresql.cnpg.io' ? Clusters view
: null)}
+ />
+ ,
+ )
+ })
+ return { element, onNavigate, onResourceClick, render, shortcuts: () => shortcuts, initial: render(selectedResource) }
+}
+
+describe('ResourcesView renderKindView', () => {
+ it('renders the kind view in place of the table and switches the table shortcuts off', async () => {
+ const writes = vi.spyOn(Storage.prototype, 'setItem')
+ const removals = vi.spyOn(Storage.prototype, 'removeItem')
+ const m = mount('?apiGroup=postgresql.cnpg.io', null)
+ await m.initial
+ const tableStorage = (calls: unknown[][]) => calls.map(([key]) => String(key)).filter((key) => key.includes('clusters'))
+ expect(tableStorage(writes.mock.calls), 'column/sort settings written while the table is not shown').toEqual([])
+ expect(tableStorage(removals.mock.calls), 'column/sort settings cleared while the table is not shown').toEqual([])
+ writes.mockRestore()
+ removals.mockRestore()
+ expect(m.element.querySelector('[data-kind-view]')).not.toBeNull()
+ expect(m.element.querySelector('table')).toBeNull()
+ expect(m.element.querySelector('input[placeholder^="Search"]')).toBeNull()
+ const byId = (id: string) => m.shortcuts().find((s) => s.id === id)
+ expect(byId('resources-search')?.enabled).toBe(false)
+ expect(byId('resources-nav-down')?.enabled).toBe(false)
+ // Kind navigation belongs to the page, not the table.
+ expect(byId('resources-next-kind')?.enabled).not.toBe(false)
+ })
+
+ it('keeps the drawer: a ?resource= deep link opens it, and switching objects pushes history', async () => {
+ const m = mount('?apiGroup=postgresql.cnpg.io&resource=pg/pg-a', null)
+ await m.initial
+ expect(m.onResourceClick).toHaveBeenCalledWith(expect.objectContaining({ kind: 'clusters', group: 'postgresql.cnpg.io', namespace: 'pg', name: 'pg-a' }))
+
+ m.onNavigate.mockClear()
+ await m.render(pg('pg-a'))
+ await m.render(pg('pg-b'))
+ const toB = m.onNavigate.mock.calls.find(([path]) => String(path).includes('resource=pg%2Fpg-b') || String(path).includes('resource=pg/pg-b'))
+ expect(toB, `navigations: ${JSON.stringify(m.onNavigate.mock.calls)}`).toBeDefined()
+ expect(toB?.[1]?.replace).not.toBe(true)
+ })
+})
diff --git a/packages/k8s-ui/src/components/resources/ResourcesView.tsx b/packages/k8s-ui/src/components/resources/ResourcesView.tsx
index a6f4a8027b..a7b2ee20a1 100644
--- a/packages/k8s-ui/src/components/resources/ResourcesView.tsx
+++ b/packages/k8s-ui/src/components/resources/ResourcesView.tsx
@@ -203,7 +203,7 @@ import { GCPManagedControlPlaneCell, GCPManagedMachinePoolCell, GCPMachineCell,
import { AzureManagedControlPlaneCell, AzureManagedMachinePoolCell, AzureMachineCell, AzureMachineTemplateCell, AzureManagedClusterCell } from './renderers/azure-capi-cells'
import { CalicoInfraCell, CalicoPolicyCell } from './renderers/calico-cells'
import { isCalicoPolicyResource, isCoreNetworkPolicyKind } from './resource-utils-calico'
-import { useRegisterShortcut, useRegisterShortcuts } from '../../hooks/useKeyboardShortcuts'
+import { useRegisterShortcut, useRegisterShortcuts, type KeyboardShortcut } from '../../hooks/useKeyboardShortcuts'
import { ResourcesSidebar } from './ResourcesSidebar'
import type { SelectedKindInfo, SidebarCategoryWorkspace } from './ResourcesSidebar'
import { CompareTray, togglePick, pickIndex, refToParam, SIDE_TONES, type CompareTrayPick, type NamespacedRef } from '../compare'
@@ -3415,6 +3415,10 @@ interface ResourcesViewProps {
sidebarCategoryWorkspaces?: Record
/** Callback when the [+] create button is clicked. Receives the currently selected kind info. */
onCreateResource?: (kind: { name: string; kind: string; group: string } | null) => void
+ /** A kind's own view, rendered in place of the table: the sidebar, kind
+ * selection and the drawer (`?resource=`) stay; the table, its toolbar and
+ * its keyboard shortcuts do not run. Return null to keep the table. */
+ renderKindView?: (kind: SelectedKindInfo) => React.ReactNode | null
/** Default kind when the URL does not include one. */
defaultKind?: SelectedKindInfo
/** Columns prepended to KNOWN_COLUMNS for every kind. For example, a
@@ -3657,6 +3661,7 @@ export function ResourcesView({
hideSidebar = false,
sidebarCategoryWorkspaces,
onCreateResource,
+ renderKindView,
defaultKind = DEFAULT_KIND_INFO,
extraLeadingColumns,
printerTable,
@@ -3695,6 +3700,8 @@ export function ResourcesView({
setShowBulkScaleDialog(false)
setBulkForceDelete(false)
}, [selectedKind.name, selectedKind.group]) // eslint-disable-line react-hooks/exhaustive-deps
+ const kindView = renderKindView?.(selectedKind) ?? null
+ const tableActive = kindView === null
const [searchTerm, setSearchTerm] = useState(initialFilters.search)
// Typing must never wait on the list or the router. The input renders raw
// searchTerm; filtering consumes the deferred copy (React yields to keep
@@ -3935,6 +3942,8 @@ export function ResourcesView({
// depends on visibleColumns) and instantly reverts the user's hide.
const gpuAutoShownKinds = useRef>(new Set())
useEffect(() => {
+ // The table's own settings: nothing to load (or clear) while a kind view replaces it.
+ if (!tableActive) return
// Re-arm the skip-initial-save guard per kind. This effect repopulates
// visible/widths/custom for the new kind via async setState, so the save
// effect (also keyed to selectedKind) that runs in this same commit still
@@ -4019,10 +4028,11 @@ export function ResourcesView({
// identity on every refetch, and re-running this effect would reset the
// user's in-session column choices each time the list polls. The key only
// changes when the kind changes or an operator upgrade changes the CRD.
- }, [selectedKind.name, selectedKind.group, extraLeadingColumns, printerColumnsKey])
+ }, [selectedKind.name, selectedKind.group, extraLeadingColumns, printerColumnsKey, tableActive])
// Save column settings when they change (skip the initial load of each kind)
useEffect(() => {
+ if (!tableActive) return
if (visibleColumns.size === 0) return // not loaded yet
for (const k of allColumnKeys) knownColumnKeys.current.add(k)
if (!isColumnSettingsLoaded.current) {
@@ -4039,7 +4049,7 @@ export function ResourcesView({
// as across sessions, so a later data refresh can't override user choices.
hadSavedColumnSettings.current = true
// allColumnKeysSig, not allColumnKeys: see the memo above.
- }, [visibleColumns, columnWidths, customColumns, selectedKind.name, selectedKind.group, allColumnKeysSig])
+ }, [visibleColumns, columnWidths, customColumns, selectedKind.name, selectedKind.group, allColumnKeysSig, tableActive])
// Close column picker on outside click or Escape
useEffect(() => {
@@ -4208,6 +4218,7 @@ export function ResourcesView({
category: 'Search',
scope: 'resources',
handler: () => searchInputRef.current?.focus(),
+ enabled: tableActive,
})
// Keyboard navigation: highlighted row state
@@ -4307,8 +4318,7 @@ export function ResourcesView({
onResourceClick?.(isSelected ? null : stripped)
}, [onRowSelect, onResourceClick, selectedKind.name, selectedKind.group])
- // Register navigation shortcuts
- useRegisterShortcuts([
+ useRegisterShortcuts(([
{
id: 'resources-nav-down',
keys: 'j',
@@ -4484,7 +4494,7 @@ export function ResourcesView({
else searchInputRef.current?.blur()
},
},
- ])
+ ] as KeyboardShortcut[]).map((sc) => ({ ...sc, enabled: tableActive && (sc.enabled ?? true) })))
// Refs for accessing filteredResources inside shortcuts (computed later in component)
const filteredResourceCountRef = useRef(0)
@@ -5879,7 +5889,10 @@ export function ResourcesView({
/>
)}
- {/* Main Content - Resource Table */}
+ {/* Main Content - the kind's own view, or the resource table */}
+ {kindView !== null ? (
+ {kindView}
+ ) : (
{/* Toolbar */}
@@ -6775,6 +6788,7 @@ export function ResourcesView({
/>
)}
+ )}
{/* Bulk delete confirmation */}
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.test.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.test.tsx
index fc3ab58a0f..c44bc662d1 100644
--- a/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.test.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.test.tsx
@@ -27,3 +27,26 @@ describe('CNPGBackupRenderer — spec.target is instance selection, not a restor
expect(html(backup())).not.toContain('Backup Target')
})
})
+
+
+describe('CNPGBackupRenderer plugin destination', () => {
+ it.each([undefined, { barmanObjectName: 'ignored-store', serverName: 'ignored-server' }, { serverName: 'ignored-server' }])('explains the Cluster destination and ignored parameters: %j', (parameters) => {
+ const out = renderToString( {}} />)
+ expect(out).toContain('From the Cluster')
+ expect(out).toContain('barman-cloud plugin')
+ expect(out).not.toContain('Object Store')
+ expect(out).not.toContain('ignored-store')
+ expect(out).not.toContain('ignored-server')
+ expect(out.match(/]*>(?:pg|ignored-store|ignored-server)<\/button>/g)).toHaveLength(1)
+ expect(out.includes('Ignored by the barman-cloud plugin')).toBe(!!parameters)
+ })
+
+ it('does not assert a third-party parameter is an ObjectStore destination', () => {
+ const out = html(backup({ spec: { method: 'plugin', pluginConfiguration: { name: 'other.example.com', parameters: { barmanObjectName: 'custom-store' } } } }))
+ expect(out).toContain('other.example.com')
+ expect(out).toContain('Unknown: Radar does not model this plugin')
+ expect(out).not.toContain('custom-store')
+ expect(out).not.toContain('Object Store')
+ expect(out).not.toContain('Ignored by')
+ })
+})
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.tsx
index 022e488a20..1a63e734d0 100644
--- a/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGBackupRenderer.tsx
@@ -12,7 +12,7 @@ import {
getCNPGBackupServerName,
getCNPGBackupError,
getCNPGBackupTarget,
- CNPG_BARMAN_OBJECTSTORE_GROUP,
+ CNPG_BARMAN_PLUGIN_NAME,
} from '../resource-utils-cnpg'
interface CNPGBackupRendererProps {
@@ -51,24 +51,16 @@ export function CNPGBackupRenderer({ data, onNavigate }: CNPGBackupRendererProps
- {/* A plugin-taken backup lands in an ObjectStore rather than in the
- Cluster's own barmanObjectStore, so naming the plugin is what
- tells an operator where to go looking. */}
{backupPlugin && }
- {backupPlugin?.parameters?.barmanObjectName && (
+ {backupPlugin && (
- }
+ label="Destination"
+ value={backupPlugin.name === CNPG_BARMAN_PLUGIN_NAME ? "From the Cluster's barman-cloud plugin" : 'Unknown: Radar does not model this plugin’s destination'}
/>
)}
+ {backupPlugin?.name === CNPG_BARMAN_PLUGIN_NAME && backupPlugin.parameters && Object.keys(backupPlugin.parameters).length > 0 && (
+
+ )}
{data.status?.instanceID?.podName && (
@@ -102,10 +94,7 @@ export function CNPGBackupRenderer({ data, onNavigate }: CNPGBackupRendererProps
}
return clusterName
})()} />
- {/* Destination and server name are in-tree barmanObjectStore fields.
- Under the plugin method they are never populated because both live
- on the ObjectStore, and rendering them as "-" reads as "not
- configured" rather than "recorded elsewhere". */}
+ {/* Plugin backups do not report the in-tree destination fields. */}
{!backupPlugin && (
<>
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGClusterRenderer.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGClusterRenderer.tsx
index 5897583887..f684b9c204 100644
--- a/packages/k8s-ui/src/components/resources/renderers/CNPGClusterRenderer.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGClusterRenderer.tsx
@@ -23,6 +23,7 @@ import {
getCNPGWALArchivingFailure,
getCNPGLastBackupFailure,
classifyCNPGClusterPhase,
+ cnpgBlockedPhaseExplanation,
getCNPGClusterAvailability,
CNPG_BARMAN_OBJECTSTORE_GROUP,
} from '../resource-utils-cnpg'
@@ -99,6 +100,7 @@ export function CNPGClusterRenderer({ data, onNavigate, declared}: CNPGClusterRe
const isFailover = phaseBucket === 'failing'
const isSwitchover = phase === 'Switchover in progress'
const isTerminal = phaseBucket === 'terminal'
+ const blocked = cnpgBlockedPhaseExplanation(phase, data.status?.phaseReason)
// "The cluster is otherwise serving normally, so nothing else here will look
// wrong" is a claim about the REST OF THIS DRAWER, so any other banner
// falsifies it. It can no longer lean on isDegraded now that a
@@ -112,11 +114,7 @@ export function CNPGClusterRenderer({ data, onNavigate, declared}: CNPGClusterRe
<>
{/* Problem alerts */}
{isTerminal && (
-
+
)}
{hasSplitBrain && (
-
+
{/* Requested size and class are spec-side intentions. These are what the
operator reports about the volumes that actually exist. */}
@@ -363,7 +361,7 @@ export function CNPGClusterRenderer({ data, onNavigate, declared}: CNPGClusterRe
WAL Storage
- {walStorage.size && }
+ {walStorage.size && }
{walStorage.storageClass && }
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGDeclarativeRenderer.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGDeclarativeRenderer.tsx
index 19c99af635..0074dac6ab 100644
--- a/packages/k8s-ui/src/components/resources/renderers/CNPGDeclarativeRenderer.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGDeclarativeRenderer.tsx
@@ -195,6 +195,31 @@ export function CNPGSubscriptionRenderer({
)
}
+/**
+ * DatabaseRole (CNPG 1.30+) — one PostgreSQL role as its own object. The
+ * Cluster's spec.managed.roles wins for the same name; the operator then
+ * reports this object not applied, and that message renders verbatim above.
+ */
+export function CNPGDatabaseRoleRenderer({ data, onNavigate }: { data: any; onNavigate?: Nav }) {
+ const spec = data?.spec ?? {}
+ const details: Array<{ label: string; value: React.ReactNode }> = [
+ { label: 'Role', value: spec.name ?? '-' },
+ { label: 'Login', value: spec.login === true ? 'Allowed' : 'Not allowed' },
+ ]
+ if (spec.superuser === true) details.push({ label: 'Superuser', value: 'Yes' })
+ if (spec.disablePassword === true) details.push({ label: 'Password', value: 'Disabled' })
+ else if (spec.passwordSecret?.name) details.push({ label: 'Password Secret', value: spec.passwordSecret.name })
+ if (spec.validUntil) details.push({ label: 'Valid Until', value: spec.validUntil })
+ if (Array.isArray(spec.inRoles) && spec.inRoles.length > 0) details.push({ label: 'Member Of', value: spec.inRoles.join(', ') })
+ if (spec.connectionLimit != null && spec.connectionLimit !== -1) details.push({ label: 'Connection Limit', value: String(spec.connectionLimit) })
+ const cc = spec.clientCertificate
+ if (cc && cc.enabled !== false) {
+ const exp = data?.status?.clientCertificate?.expiration
+ details.push({ label: 'Client Certificate', value: `${data?.metadata?.name}-client-cert · ${exp ? `expires ${exp}` : 'expiry not reported'}` })
+ }
+ return
+}
+
/**
* ImageCatalog and ClusterImageCatalog — the PostgreSQL images a Cluster may
* use, pinned per major version.
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.test.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.test.tsx
new file mode 100644
index 0000000000..06c9086027
--- /dev/null
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.test.tsx
@@ -0,0 +1,32 @@
+import { describe, expect, it } from 'vitest'
+import { renderToString } from 'react-dom/server'
+import { CNPGScheduledBackupRenderer } from './CNPGScheduledBackupRenderer'
+
+const schedule = (name: string, parameters?: { barmanObjectName?: string; serverName?: string }) => ({
+ apiVersion: 'postgresql.cnpg.io/v1',
+ kind: 'ScheduledBackup',
+ metadata: { name: 'nightly', namespace: 'db' },
+ spec: { cluster: { name: 'pg' }, schedule: '0 0 0 * * *', method: 'plugin', pluginConfiguration: { name, parameters } },
+})
+
+describe('CNPGScheduledBackupRenderer plugin destination', () => {
+ it.each([undefined, { barmanObjectName: 'ignored-store', serverName: 'ignored-server' }, { serverName: 'ignored-server' }])('explains the Cluster destination and ignored parameters: %j', (parameters) => {
+ const out = renderToString( {}} />)
+ expect(out).toContain('From the Cluster')
+ expect(out).toContain('barman-cloud plugin')
+ expect(out).not.toContain('Object Store')
+ expect(out).not.toContain('ignored-store')
+ expect(out).not.toContain('ignored-server')
+ expect(out.match(/]*>(?:pg|ignored-store|ignored-server)<\/button>/g)).toHaveLength(1)
+ expect(out.includes('Ignored by the barman-cloud plugin')).toBe(!!parameters)
+ })
+
+ it('does not assert a third-party parameter is an ObjectStore destination', () => {
+ const out = renderToString( )
+ expect(out).toContain('other.example.com')
+ expect(out).toContain('Unknown: Radar does not model this plugin')
+ expect(out).not.toContain('custom-store')
+ expect(out).not.toContain('Object Store')
+ expect(out).not.toContain('Ignored by')
+ })
+})
diff --git a/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.tsx b/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.tsx
index 58ff476347..f233f3c873 100644
--- a/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/CNPGScheduledBackupRenderer.tsx
@@ -11,7 +11,7 @@ import {
getCNPGScheduledBackupIsImmediate,
getCNPGBackupPlugin,
getCNPGScheduledBackupOwnerRef,
- CNPG_BARMAN_OBJECTSTORE_GROUP,
+ CNPG_BARMAN_PLUGIN_NAME,
} from '../resource-utils-cnpg'
interface CNPGScheduledBackupRendererProps {
@@ -64,24 +64,16 @@ export function CNPGScheduledBackupRenderer({ data, onNavigate }: CNPGScheduledB
return clusterName
})()} />
- {/* Under the plugin method the destination lives on an ObjectStore in
- another API group, so naming it is the only route from here to
- where these backups will actually land. */}
{schedulePlugin && }
- {schedulePlugin?.parameters?.barmanObjectName && (
+ {schedulePlugin && (
- }
+ label="Destination"
+ value={schedulePlugin.name === CNPG_BARMAN_PLUGIN_NAME ? "From the Cluster's barman-cloud plugin" : 'Unknown: Radar does not model this plugin’s destination'}
/>
)}
+ {schedulePlugin?.name === CNPG_BARMAN_PLUGIN_NAME && schedulePlugin.parameters && Object.keys(schedulePlugin.parameters).length > 0 && (
+
+ )}
diff --git a/packages/k8s-ui/src/components/resources/renderers/MetricsUnavailableNotice.tsx b/packages/k8s-ui/src/components/resources/renderers/MetricsUnavailableNotice.tsx
index 13eaf14340..46e171cc11 100644
--- a/packages/k8s-ui/src/components/resources/renderers/MetricsUnavailableNotice.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/MetricsUnavailableNotice.tsx
@@ -2,16 +2,17 @@ import { Info } from 'lucide-react'
import { Tooltip } from '../../ui/Tooltip'
interface MetricsUnavailableNoticeProps {
+ noUsageYet?: boolean
rawError?: string
diagnosis?: string
}
-export function MetricsUnavailableNotice({ rawError, diagnosis }: MetricsUnavailableNoticeProps) {
+export function MetricsUnavailableNotice({ rawError, diagnosis, noUsageYet }: MetricsUnavailableNoticeProps) {
return (
- Metrics unavailable. Radar cannot read metrics.k8s.io.
+ {noUsageYet ? 'No usage yet: this Pod is not running' : 'Metrics unavailable. Radar cannot read metrics.k8s.io.'}
{rawError && (
{
expect(html).not.toContain('could not be loaded')
})
})
+
+describe('PodRenderer pending Pods, including non-CNPG workloads', () => {
+ const pending = { apiVersion: 'v1', kind: 'Pod', metadata: { name: 'nginx-join', namespace: 'default' }, spec: { containers: [{ name: 'nginx', image: 'nginx' }], initContainers: [{ name: 'setup', image: 'busybox' }] }, status: { phase: 'Pending' } }
+ const render = (data: any, metricsHistory: any = { containers: [], metricsAPIReachable: true }) => renderToString( {}} copied={null} metricsUnavailable metricsHistory={metricsHistory} />)
+ it('shows no usage yet only with a successful metrics API collection', () => {
+ expect(render(pending)).toContain('No usage yet: this Pod is not running')
+ expect(render(pending)).not.toContain('Radar cannot read metrics.k8s.io')
+ expect(render(pending, { containers: [], metricsAPIReachable: false })).toContain('Radar cannot read metrics.k8s.io')
+ expect(render(pending, { containers: [], collectionError: 'forbidden' })).toContain('forbidden')
+ expect(render({ ...pending, status: { phase: 'Running' } })).toContain('Radar cannot read metrics.k8s.io')
+ })
+ it('shows unstarted regular and init containers neutrally', () => {
+ const html = render(pending)
+ expect(html.match(/Not started/g)).toHaveLength(2)
+ expect(html).not.toContain('Not Ready')
+ expect(html).not.toMatch(/>unknown)
+ })
+ it.each(['ContainerCreating', 'PodInitializing', 'ImagePullBackOff', 'CrashLoopBackOff'])('tones reported waiting reason %s appropriately', (reason) => {
+ const html = render({ ...pending, status: { phase: 'Pending', containerStatuses: [{ name: 'nginx', state: { waiting: { reason } } }] } })
+ const bad = ['ImagePullBackOff', 'CrashLoopBackOff'].includes(reason)
+ expect(html).toContain(bad ? 'bg-red-100' : 'bg-theme-hover/50')
+ if (bad) expect(html).toContain(reason)
+ })
+})
+
+it('shows reported running-but-unready nginx readiness as degraded', () => {
+ const data = { ...pod, status: { phase: 'Running', containerStatuses: [{ name: pod.spec.containers[0].name, ready: false, state: { running: {} } }] } }
+ const html = renderToString( {}} copied={null} />)
+ expect(html).toMatch(/bg-amber-100[^>]*>Not Ready)
+ expect(html).toMatch(/>running)
+})
diff --git a/packages/k8s-ui/src/components/resources/renderers/PodRenderer.tsx b/packages/k8s-ui/src/components/resources/renderers/PodRenderer.tsx
index 60aaa31928..2b21c9beda 100644
--- a/packages/k8s-ui/src/components/resources/renderers/PodRenderer.tsx
+++ b/packages/k8s-ui/src/components/resources/renderers/PodRenderer.tsx
@@ -5,7 +5,7 @@ import { PolicySection } from './PolicySection'
import type { PolicyResourceResponse } from '../../../types/policy'
import { Section, PropertyList, Property, ConditionsSection, CopyHandler, AlertBanner, ResourceLink, useOperationalIssuesShown } from '../../ui/drawer-components'
import { formatResources, formatDuration, getPodProblems, getPodPhaseDisplay, healthColors, SEVERITY_DOT_COLOR, getDefaultContainerName } from '../resource-utils'
-import { getResourceStatusColor, SEVERITY_BADGE_BORDERED } from '../../../utils/badge-colors'
+import { getResourceStatusColor, SEVERITY_BORDER } from '../../../utils/badge-colors'
import {
rbacVerbBadgeClass,
rbacResourceBadgeClass,
@@ -26,6 +26,7 @@ import { Tooltip } from '../../ui/Tooltip'
import { RadarUpgradeNote } from '../../ui/RadarUpgradeNote'
import { getRadarUpgradeRequirement, type RadarUpgradeRequirement } from '../../../types/fetch-error'
import { MetricsChart } from '../../ui/MetricsChart'
+import { Badge } from '../../ui/Badge'
import { MetricsUnavailableNotice } from './MetricsUnavailableNotice'
import { ContainerEnvironmentSection } from './ContainerEnvironmentSection'
@@ -90,7 +91,7 @@ interface PodRendererProps {
renderPortAction?: (props: { namespace: string; podName: string; port: number; protocol: string; disabled?: boolean }) => ReactNode
// Metrics data injection
metrics?: { containers?: any[]; timestamp?: string }
- metricsHistory?: { containers?: any[]; collectionError?: string; metricsUnavailableReason?: string; metricsUnavailableDiagnosis?: string }
+ metricsHistory?: { metricsAPIReachable?: boolean; containers?: any[]; collectionError?: string; metricsUnavailableReason?: string; metricsUnavailableDiagnosis?: string }
metricsUnavailable?: boolean
hideMetricsServer?: boolean
// Filesystem browser render props
@@ -330,7 +331,9 @@ export function PodRenderer({
const operationalIssuesShown = useOperationalIssuesShown()
const podProblems = getPodProblems(data)
const hasProblems = podProblems.length > 0 && !operationalIssuesShown
- const showMetricsUnavailable = !!metricsUnavailable && !metricsHistory?.collectionError
+ const notRunning = data.status?.phase === 'Pending' && !containerStatuses.some((s: any) => s.state?.running)
+ const noUsageYet = notRunning && metricsHistory?.metricsAPIReachable === true && !metricsHistory.collectionError && !metricsHistory.containers?.length && !metrics?.containers?.length
+ const showMetricsUnavailable = !!metricsUnavailable && !metricsHistory?.collectionError && !noUsageYet
const hasMetricsHistory = !!metricsHistory?.containers?.length
const currentMetrics = metricsUnavailable ? undefined : metrics
@@ -506,10 +509,11 @@ export function PodRenderer({
} else if (isWaiting) {
statusLabel = state?.waiting?.reason || 'Waiting'
} else {
- statusLabel = 'Pending'
+ statusLabel = 'Not started'
}
+ const reportedFailure = isFailed || !!(isWaiting && state?.waiting?.reason && !['ContainerCreating', 'PodInitializing'].includes(state.waiting.reason))
const statusColor = getResourceStatusColor(
- isCompleted ? 'succeeded' : isFailed ? 'failed' : isInitRunning ? 'running' : isWaiting ? 'waiting' : 'pending'
+ isCompleted ? 'succeeded' : reportedFailure ? 'failed' : isInitRunning ? 'running' : 'unknown'
)
// Build command string
@@ -520,10 +524,11 @@ export function PodRenderer({
return (