Skip to content

feat(apm): scope Kubernetes pod metrics to capture selections - #484

Merged
youzi-1122 merged 4 commits into
mainfrom
feat/k8s-app-metrics-scope
Oct 9, 2026
Merged

youzi-1122 merged 4 commits into
mainfrom
feat/k8s-app-metrics-scope

Conversation

@youzi-1122

@youzi-1122 youzi-1122 commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Kubernetes application metrics now follow the same namespace and workload selection as Auto APM. Selecting a namespace includes its annotated Pods, including new Pods; selecting a workload includes only its owned Pods; clearing the selection stops application scrapes while kube-state-metrics continues.

  • Use the bundled official OpenTelemetry Collector Prometheus Receiver in the single-replica metrics scraper for prometheus.io/* Pod annotations. Enable discovery by default while preserving explicit opt-outs, and filter targets before scraping with the scope projected through the existing telemetry configuration and Secret.
  • Preserve application metric labels and target_info resource metadata, and keep raw Pod metrics available under ongrid_source="k8s:app-metrics". Exclude that source from APM request and runtime queries so overlapping OBI/application measurements are not added together. This is a source-selection policy, not generic semantic deduplication.
  • Simplify APM navigation around services, topology, and Kubernetes-first discovery. Remove the setup guide, retain collection diagnostics, and preserve legacy links.
  • Add a Prometheus-only Go example with HTTP, custom, Go runtime, and process metrics, plus regression tests and an ADR update.

Validation

  • Passed targeted Go race tests for shared scope parsing, Manager Kubernetes configuration/HTTP, Edge discovery, APM queries, and the Go example. Edge command tests ran on Linux, including native configuration validation with Collector 0.157.0.
  • Passed a native Collector 0.157.0 round-trip regression for target_info: service/build metadata, platform labels, and the job/instance join to application metrics survive scraping and remote_write. The test failed before the fix and passed afterward. Run it with ONGRID_TEST_OTELCOL_BINARY; without that binary it is explicitly skipped.
  • Passed all APM Promtool fixtures, including overlapping metric names and resource identities; passed Helm chart regression checks, make proto, and git diff --check.
  • Passed 49 affected frontend tests, production build, and changed-file ESLint. Full frontend lint reports existing errors in unchanged files (chat.ts, packetCaptures.ts, MessageBubble.tsx, and SkillRun.tsx). Browser screenshot review was left to the user as requested.
  • Live Kubernetes checks passed for workload selection, namespace expansion, exclusion of an outside namespace, removal, and restoration. kube-state-metrics remained fresh, and the original capture rules were restored.
  • The Go demo produced 600 requests: both sources eventually recorded 540 successes and 60 failures. APM API RPS matched the OBI-only query without adding raw Pod metrics. Raw HTTP/Go/process metrics and traces were present; the annotated endpoint was scraped once despite two declared container ports.
  • Passed full Linux Go validation: go test -race -p 4 ./....
  • Local Buf and go-arch-lint were unavailable. No large-scale load test or cross-platform OBI certification was performed; live checks used the existing local ARM64 OBI test environment.

Risk and rollback

  • Upgrade Manager and Controller before the metrics scraper. The telemetry API change is additive; missing or empty scope does not enable unrestricted application scrapes. Older components that cannot supply scope will leave new application scraping inactive. Existing read failures retain the last valid configuration.
  • Scope changes are asynchronous: controller polling, inventory updates, Secret projection, and scraper reload contribute delay. Removal took about two minutes in this test; historical metrics remain stored.
  • The scraper gains Pod list/watch discovery and a writable plugin directory, but no Secret API permission. It must remain a single replica until target sharding is implemented. Readiness does not guarantee every application target or exporter is healthy.
  • No database migration. Set kubernetesMetrics.appDiscovery.enabled=false to stop this scraper, or revert these commits and roll back Manager/Edge/Chart together. Existing capture selections and stored metrics are retained.

Author confirmation

@youzi-1122
youzi-1122 requested a review from singchia as a code owner October 9, 2026 05:41
@youzi-1122
youzi-1122 merged commit ec881ed into main Oct 9, 2026
12 checks passed
@youzi-1122
youzi-1122 deleted the feat/k8s-app-metrics-scope branch October 9, 2026 06:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant