Skip to content

feat(genai): add AI span inspector - #2

Draft
Sun-ZhenXing wants to merge 56 commits into
mainfrom
feat-genai-ui
Draft

feat(genai): add AI span inspector#2
Sun-ZhenXing wants to merge 56 commits into
mainfrom
feat-genai-ui

Conversation

@Sun-ZhenXing

Copy link
Copy Markdown
Collaborator

Pull Request


📄 Summary

Why does this change exist?
What problem does it solve, and why is this the right approach?

Screenshots / Screen Recordings (if applicable)

Include screenshots or screen recordings that clearly show the behavior before the change and the result after the change. This helps reviewers quickly understand the impact and verify the update.

Issues closed by this PR

Reference issues using Closes #issue-number to enable automatic closure on merge.


✅ Change Type

Select all that apply

  • ✨ Feature
  • 🐛 Bug fix
  • ♻️ Refactor
  • 🛠️ Infra / Tooling
  • 🧪 Test-only

🐛 Bug Context

Required if this PR fixes a bug

Root Cause

What caused the issue?
Regression, faulty assumption, edge case, refactor, etc.

Fix Strategy

How does this PR address the root cause?


🧪 Testing Strategy

How was this change validated?

  • Tests added/updated:
  • Manual verification:
  • Edge cases covered:

⚠️ Risk & Impact Assessment

What could break? How do we recover?

  • Blast radius:
  • Potential regressions:
  • Rollback plan:

📝 Changelog

Fill only if this affects users, APIs, UI, or documented behavior
Use N/A for internal or non-user-facing changes

Field Value
Deployment Type Cloud / OSS / Enterprise
Change Type Feature / Bug Fix / Maintenance
Description User-facing summary

📋 Checklist

  • Tests added or explicitly not required
  • Manually tested
  • Breaking changes documented
  • Backward compatibility considered

👀 Notes for Reviewers


@Sun-ZhenXing
Sun-ZhenXing marked this pull request as draft July 30, 2026 03:34
tushar-signoz and others added 29 commits August 3, 2026 14:41
…re tests (SigNoz#12383)

primus bumped golangci-lint to v2.12.2, whose govet now runs the inline
analyzer and whose sloglint is stricter. CI resolves primus.workflows@main,
so every PR started failing lint the moment that landed.

- reflect.Ptr is a deprecated alias carrying //go:fix inline, so it is now
  reflect.Pointer at all five call sites
- metricsstatementbuilder imported golang.org/x/exp/slices, which carries
  //go:fix inline pointing at the stdlib; the analyzer cannot inline
  generics, so switch the import to stdlib slices as the directive intends
- pkg/instrumentation/loghandler emits OpenTelemetry semantic-convention
  attributes (code.filepath, exception.type, ...), which are dotted rather
  than snake_case by definition. Renaming them would break every log
  consumer, so the keys move to constants in instrumentationtypes, which
  already held this kind of key -- and already defined code.function, so
  source.go was duplicating it. Six of the seven alias the semconv
  constants that define them; exception.code has no OTel equivalent.
  sloglint resolves a same-package constant back to its literal but skips
  a qualified one, so this needs no exclusion.

CI reported 8 issues but capped at max-same-issues=3, hiding 2 more
reflect.Ptr sites and 4 more sloglint ones.

Separately, TestTimeout/WaitTillNoTimeoutForExcludedPath failed with
"transport connection broken: http: CloseIdleConnections called".
TestTimeout and TestCache issue requests through http.DefaultClient while
a parallel subtest in response_test.go closes an httptest.Server, and
httptest.Server.Close calls http.DefaultTransport.CloseIdleConnections.
Both tests now use their own client and transport, and are closed via
t.Cleanup. That makes Serve return ErrServerClosed on every run, so the
require.NoError wrapping it is dropped -- it could never have held, and
require runs t.FailNow off the test goroutine anyway. Bare Serve in a
goroutine matches routerweb and render tests.
…pinned serving (SigNoz#12324)

* feat(querier): wire clickhouseprometheusv2 for shadow comparison and pinned serving

Stand up the v2 provider next to the default one and give the querier two
flag-gated ways to exercise it, neither affecting default serving:

- shadow: with use_prometheus_clickhouse_v2 on, every PromQL query re-runs
  on v2 after the response is sent; result diffs are logged. Bounded by a
  small per-process admission cap (skip, not queue, at the cap).
- pin: the X-SigNoz-PromQL-Provider header serves the response from v2
  directly, for side-by-side comparison by integration tests and support.
  Pinned requests bypass the cache in both directions.

Typed storage errors (series/sample budgets) survive the engine's wrapping
and surface as 4xx instead of internal 500s. PromQL results now carry
ClickHouse scan stats.

prometheus::provider: clickhousev2 also makes v2 the serving provider
outright (no shadow in that mode - nothing to compare against).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(promql): replay the conformance corpus on both providers

Every corpus case now runs twice — default provider, and pinned to
clickhousev2 via the flag-gated X-SigNoz-PromQL-Provider header — each leg
asserted against the same frozen expectations, each leg with its own
known-divergences ledger enforced in both directions. The new
known_divergences_v2.json is the rollout scorecard: the provider swap is
measured by burning it to empty.

The legs are never asserted against each other: both can sit within one
rounding quantum of the expected value yet differ by up to two quanta at a
rounding boundary, so a leg-vs-leg check would reintroduce the boundary
noise the tolerance absorbs. Anchoring both to the same oracle over the
same ingested bytes already localizes any disagreement to a provider.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…Noz#12385)

- Embed statementbuilder.Config into querier.Config with mapstructure ",squash";
  keys move to querier.skip_resource_fingerprint.* (env SIGNOZ_QUERIER_*).
- Drop the standalone statementbuilder section and its config factory.
- Pass statementbuilder.Config wholesale into NewLogQueryStatementBuilder.
… metric (SigNoz#12313)

* chore: delta temporality metrics should always be considered as a sum metric

* test: add integration tests

* test: parametrise the non-reduced test for rate and increase both

* test: add comment explaining last samples in test

* chore: remove unneeded comments

---------

Co-authored-by: Srikanth Chekuri <srikanth.chekuri92@gmail.com>
* fix: enforace the required tag on dashboard spec fields

* test: add empty list and objects for required fields in integration tests
- Replace the search() stub: the logs condition builder fans a case-insensitive
  match of the term across every searchable column (log columns, body/body_v2,
  attribute + resource maps).
- Grammar: searchCall takes a valueList (parser regenerated), so the scoped form
  search('term', body, resource) parses and narrows the fan-out to those contexts.
- The visitor emits FilterOperatorSearch and flags the statement scan-heavy; the
  statement builder attaches a CostGuard the querier enforces via EXPLAIN ESTIMATE
  against a per-shard budget, cumulative across buckets on the window-list path.
- body_v2 gets its own, lower budget (search_max_scan_rows_json_body): toString()
  rebuilds every document and no skip index prunes.
- Logs-only: other signals reject search().
- Unit tests for the per-context fan-out SQL and both budgets; integration suites
  run the same matrix over the legacy body and over body_v2.
* fix(alertmanager): resolve email SMTP settings from env

* refactor(alertmanager): remove dead legacy config and tidy integration test helpers

* test(alertmanager): scope smtp rotation test to env rotation only

* chore: drop redundant test

* chore: log and warn; do not persist global config
SigNoz#12384)

getCaretContext resolved the active slot to a token's start index while
keeping replaceEnd at the caret, so a caret parked in the whitespace
before that token produced an inverted range (replaceStart > replaceEnd).
dslCompletionSource passes that range to CodeMirror as CompletionResult
from/to, and accepting a suggestion threw

  RangeError: Invalid change range 16 to 15 (in doc of length 36)

inside view.dispatch. The same inverted range made spliceAtCaret
duplicate the skipped character.

The replaced range is defined as ending at the caret, so clamp it there
and let the slot collapse to a plain insertion point.
…right place & move events out of constants (SigNoz#12359)

* fix(pods): redirecting to the wrong doc for chart

* refactor(events): move out of src/constants

* test(use-log): fix issue with mock

* docs(infra-monitoring): fix more broken doc path

* fix(pods): disable sort for pod restarts
* chore: add api to retry migration for a dashboard

* feat(dashboard): add retry migration action for legacy dashboards

The legacy-dashboard dialog only offered the dashboard ID and a link to
support. Now that the v1->v2 migration can be re-run on demand, let an
editor trigger it from there and fall back to support only if it still
fails.

Retrying needs edit access (the endpoint is EDITOR-gated), so viewers
keep the ID-and-support dialog unchanged.

---------

Co-authored-by: Ashwin Bhatkal <ashwin96@gmail.com>
* Revert "chore: added infraMonitoringV2 feature flag (SigNoz#12031)"

This reverts commit 1d441b6.

* chore(infrastructure-monitoring-v2): remove feature flag
…gNoz#12382)

A numeric coercion yields seconds since epoch, so max(timestamp) came back
as 1758113657.04 and the third-party APIs "Last Seen" rendered as January
1970. Both rules key on the physical column type, since no FieldDataType
denotes a timestamp.

- a time column is never coerced: it reaches the aggregate in its native
  type and the driver returns a time.Time
- a bare `timestamp` resolves to the intrinsic column alone; a same-named
  numeric attribute no longer joins the candidate union, where the mixed
  branches failed with "no supertype for DateTime64, Float64"
- the exists guard leaves the String bucket: `timestamp <> ''` becomes a
  typed epoch-zero comparison, verified equivalent on ClickHouse 25.5
  including a non-UTC server timezone
- max/min/quantile/count keep working; sum/avg and the rate_* variants now
  fail at the database instead of returning a rescaled number, pinned by
  integration tests until they are rejected up front
- tests/fixtures/traces.py wrote kind_string as the enum member name
  ("SPAN_KIND_CLIENT"), so any feature filtering on it matched nothing; it
  now maps the six kinds to the exporter's form ("Client")
- Last Seen assertions pin the encoding and the exact instant, rather than
  accepting either an RFC 3339 string or epoch millis
- drop the /overview/domain integration tests and the fixture surface that
  served only that route; the endpoint is no longer used
Add metadata endpoint checks for openstack (covers OVHcloud), oracle,
alibaba, akamai (linode), ibm, and scaleway, and the FLY_APP_NAME
environment variable for fly.io. OpenStack must come before AWS since
it exposes an EC2-compatible endpoint.
…2397)

* fix(version): detect docker under private cgroup namespaces

Docker defaults containers to a private cgroup namespace on cgroup v2
hosts (20.10+), which remaps /proc/self/cgroup to "0::/" and defeats
the runtime-name sniff, so compose installs report
deployment.mode=unknown. Check the container manager marker files
(/run/.containerenv, /.dockerenv) and the container env var first,
mirroring systemd's detect_container, and fall back to mountinfo when
the cgroup path is namespaced away.

* fix(version): drop the docker case from the container env check

docker never sets the container variable; /.dockerenv already covers it
* feat: support llm trace list and span list

* fix: take perf into consideration

* fix: more tests

* fix: more cleanup

* fix: cleanup and more tests

* fix: add resource fingerprint cte

* fix: edge cases and correct cost key

* fix: update integration test

* fix: update openapi

* fix: address comments

* fix: address comments

* fix: fix tests

* fix: add back the flag in metadata

* fix: remove source and change to builder ai query

* fix: updated openapi

* fix: minor cleanup

* fix: refactor as requested

* fix: address comments

* fix: address comments

* fix: remove tracefield. explicit rejection

* fix: remove comment

* fix: remove accidentally added file

* fix: add NewFactory

* fix: refactor integration tests

* fix: address comment

* fix: use assert

* fix: use assert and condense comments
…sh (SigNoz#12402)

* fix(k8s-expanded-row): ensure values are correctly aligned

Ref: https://app.notion.com/p/signoz/Alignment-in-rows-inside-a-group-wrt-column-headers-is-not-correct-3b1fcc6bcd1980dd910bff84d1defe47?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(pods): ensure table aligment on group by

Ref: https://app.notion.com/p/signoz/Alignment-in-rows-inside-a-group-wrt-column-headers-is-not-correct-3b1fcc6bcd1980dd910bff84d1defe47?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): not allowing name columns to be sorted

Ref: https://app.notion.com/p/signoz/No-option-to-order-resources-by-name-3b1fcc6bcd198019a729d2575731387e?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): cannot sort name columns when group by is enabled

Ref: https://app.notion.com/p/signoz/Cannot-use-sort-when-group-by-is-added-3b2fcc6bcd19806d9514d27e4da26c35?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): sort key was not correct

Ref: https://app.notion.com/p/signoz/Review-all-sortable-columns-3b2fcc6bcd1980c4b267f3e35d4e875f?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* feat(infra-monitoring): add descriptions for each threshold of entity progress

Ref: https://app.notion.com/p/signoz/For-usage-metrics-the-info-icon-can-explain-the-colours-right-there-instead-of-just-saying-to-go-to-3b1fcc6bcd1980fc81a4f2b8179541b9?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): explorer button not allowing the back button

Ref: https://app.notion.com/p/signoz/After-click-to-go-to-metrics-explorer-on-chart-we-cannot-go-back-3b1fcc6bcd1980d7b52fd7dbe0b3e29f?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): polluted navigation history

Fixes the following issue:
- Go to namespaces
- Select “generator” and open the details
- Close the details
- Go to pods
- Group by namespaces

Expand “generator”, this will cause it to be redirected to namespaces category

Ref: https://app.notion.com/p/signoz/wrong-redirect-when-navigating-between-categories-3b2fcc6bcd1980e996fedd639ce72b01?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* fix(infra-monitoring): get min max not detecting 1month as valid interval

Ref: https://app.notion.com/p/signoz/1month-is-not-supported-to-be-received-as-relativeTime-on-metrics-explorer-3b2fcc6bcd1980c0bd17f1e3d869e987?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link

* refactor(infra-monitoring): update table statistics for better auto page size

Ref: https://app.notion.com/p/signoz/change-the-baseline-values-for-calculated-page-size-3b2fcc6bcd19807caafbdedb725b6ab7?v=65cfcc6bcd1982a6bee688d8fd55420c&source=copy_link
* fix(dashboard): count panel stats from the v2 spec

The panel counters walked a top-level `widgets` array and returned early
when the key was missing. A v2 dashboard stores only `metadata` and
`spec`, with panels as a map under `spec.panels`, so every v2 row hit
that early return and all `dashboard.panels.*` stats stayed at zero.
`dashboard.count` was unaffected — it is a row count.

Read the v2 spec instead: count the panels under `spec.panels` and take
each panel's signal from its query envelope, reusing the typed read path
and `QueryEnvelope.GetSignal`. Signal-less queries (promql, clickhouse
sql, formulas) count towards the panel total only. v1 rows are no longer
parsed for panel stats and contribute to `dashboard.count` alone.

* test(dashboard): drop the constant name arg from the stats query helper

statsBuilderQuery only ever received "A", which go-lint flags via unparam.
A panel holds a single query, so the name never mattered to the assertions;
the composite test still names its sub-queries through statsBuilderQuerySpec.

* refactor(dashboard): move v2 panel stats to a perses_ file

All v2 code lives in perses_-prefixed files until the v1 code goes away.
Pure move of the stats block out of dashboard.go, tests alongside it.

* fix(dashboard): count create-v2 stats off the postable spec

CreateV2 already holds the postable dashboard, so decoding the storable
back into a v2 dashboard just to count its panels was a needless type
conversion on the create path.

Split the panel walk into addPanelStats over a DashboardSpec, and add
NewStatsFromPostableDashboardV2 for the create path. The storable variant
keeps its signature for the periodic collectors, which only have rows.
* feat(alert-channels): add google chat channel type, defaults and url validation

Prefills the title and description templates verbatim from the backend's
DefaultGoogleChatReceiverConfig, and extracts the slack templates into
SlackInitialConfig so the form can restore them when the type changes.

* feat(alert-channels): add google chat option and form fields

* feat(alert-channels): wire google chat create, edit and test calls

* test(alert-channels): cover the google chat channel form

* fix(googlechat): use standard markdown in default templates for cardsV2
* fix: remove if condition for json parser

* fix: update integration tests
…2415)

* refactor(metrics-explorer): move use page size out of infra monitoring

* refactor(query-builder): update code to reference v2 instead of v1

* refactor(infra-monitoring): delete v1 code
…metric-reduction-feature` path (SigNoz#12413)

* chore: using local tables in place of distributed tables in infra-monitoring reduced metric path

* test: infra-monitoring list endpoints on the metrics reduction path

* chore: address review comments
…Noz#12305)

Bumps [github.com/google/cel-go](https://github.com/google/cel-go) from 0.28.0 to 0.29.0.
- [Release notes](https://github.com/google/cel-go/releases)
- [Commits](cel-expr/cel-go@v0.28.0...v0.29.0)

---
updated-dependencies:
- dependency-name: github.com/google/cel-go
  dependency-version: 0.29.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…z#12222)

Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.80.0 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](grpc/grpc-go@v1.80.0...v1.82.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.82.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* feat: enable FGA for dashboards and their public config

* test: add integration test for dashboard FGA

* fix: fix permissions for public dashboards, pinning, views

* fix: allow viewers to manage views

* fix: remove edits to the public dashboard line
)

Co-authored-by: Nikhil Mantri <nikhil.mantri1999@gmail.com>
revmag and others added 27 commits August 6, 2026 13:25
…, Auth0, and PgBouncer datasources (SigNoz#12098)

* feat(onboarding): add HCP Vault, OpenTelemetry eBPF, Langflow, Cohere, and Auth0 datasources

Add UI onboarding configurations (logos + datasource entries) so these tools
appear in the Add Data Source onboarding flow. SVG logos optimized with svgo.

Closes signoz.io#3614 (HCP Vault)
Closes signoz.io#3626 (OpenTelemetry eBPF / OBI)
Closes signoz.io#3680 (Langflow)
Closes signoz.io#3702 (Cohere)
Closes signoz.io#3562 (Auth0)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(onboarding): add PgBouncer datasource

Add PgBouncer (PostgreSQL connection pooler) metrics onboarding entry,
reusing the existing PostgreSQL logo.

Closes signoz.io#3553 (PgBouncer)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(onboarding): update SVG logos and category of datasources

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Jugal Kishore <jugal@signoz.io>
…2440)

A tab that outlives a deploy requests hashed assets the new build no
longer has. `lazyRetry` already recovers from this by reloading once, so
the resulting errors are noise — they spike on every deploy and each one
burns a Session Replay (`replaysOnErrorSampleRate: 1.0`).

### Sentry `ignoreErrors`

Filters the whole class. Four patterns because the same failure is
worded differently per source:

| Pattern | Source |
|---|---|
| `Unable to preload CSS for` | Vite's own thrown `Error`, identical
everywhere |
| `Failed to fetch dynamically imported module` | Chromium |
| `error loading dynamically imported module` | Firefox |
| `Importing a module script failed` | Safari |

`ignoreErrors` is applied as an event processor (`@sentry/core`
`eventFilters.js`), so the event is dropped before transport. Replay's
error flush hooks `afterSendEvent`, which never fires for a dropped
event — so this stops the replay burn too, not just the issue count.

Trade-off, stated plainly: stale-asset failures now produce no Sentry
signal at all, including the case where the reload doesn't fix it. A
genuinely broken deploy has to be caught from asset 404 rates rather
than from Sentry.

### `lazyRetry`

Behaviour is unchanged. One guard added: `setSessionStorageApi` returns
`false` when sessionStorage is blocked (iframe, storage disabled), and
the retry flag can't persist. The reload was previously issued anyway,
so every failed import reloaded forever with no way out. It now reloads
only when the flag was actually written.
…tatements (SigNoz#12325)

> **Stack** (review in order; each PR's diff is against its
predecessor):
> 1. SigNoz#12323 `v2-read-path` — v2 native read path (leaf package)
> 2. SigNoz#12324 `v2-wiring` — wiring, shadow/pin rollout machinery, dual-leg
conformance
> 3. SigNoz#12325 `v2-transpiler` — PromQL→ClickHouse transpiler +
classification golden
> 4. SigNoz#12093 `issue-4293` — the /prometheus API move (breaking slice,
last)

### What

The performance half of the v2 provider: an allowlist compiler
(`classify`/`rewrite`) that evaluates proven PromQL shapes entirely
inside
ClickHouse on the `timeSeries*ToGrid` aggregate functions (CH ≥ 25.6),
so one
row per output series comes back instead of every raw sample. Everything
not
provably equivalent falls back to the engine over the PR-1 querier; a
transpilable subtree under a non-transpilable node runs hybrid (subtree
materialized as synthetic series, engine on top). `TryExecuteRange`
slots
into the PR-2 serve/shadow paths (until now engine-only) through the new
`prometheus.RangeExecutor` capability interface — the provider stays
unexported and pkg/querier keeps holding `prometheus.Prometheus`;
providers
without the capability (v1) simply never transpile. The capability folds
into the main interface once v1 is removed.

Highlights (docs/contributing/prometheus.md carries the full correctness
story):

- Range functions map to verified grid aggregates; `increase` is
  `rate × range` exactly (same extrapolated delta, factor algebra).
- Instant selectors reproduce stale-marker shadowing with a
three-aggregate
  compare — skipping stale rows in WHERE would resurrect the sample the
  marker buried.
- `*_over_time` at range = k·step aggregates whole step buckets
  (`groupArrayInsertAt` + slide) — no per-window fan-out, no prefix-sum
  differencing.
- **Window-sliver filtering** (the headline perf commit, folded here):
when
the window is narrower than the step, only window/step of the timeline
can
influence any grid point; a lattice predicate in WHERE cuts the
aggregate's
  input by the coverage ratio — measured 74s/28GiB → 16s/4.3GiB on a
  36k-series 1w rate, and a 2.67B-sample case that exceeded 150GiB now
completes in 19s/17GiB. Over sliver-filtered rows the last-style gates
lift
(instant selectors and `last_over_time` transpile at window < step), and
  disjoint-window `*_over_time` forms drop the divisibility gate.
- Scalar-op pipelines apply in Go, slot by slot — same float64 ops, same
  order the AST dictates.

Two guards land with it:

- **Classification golden** (`classification_golden_test.go` +
  `testdata/classification_golden.json`): freezes the route
(full/hybrid(n)/fallback + reason) of every conformance-corpus
expression,
one line each — 317 expressions: 132 full, 39 hybrid, 146 fallback. The
test also requires each expression to route the same on every corpus
grid;
if a classifier change ever makes the route grid-dependent, the test
fails
  and the key must grow. Routing is its own
correctness surface — silently falling back costs the pushdown, silently
  transpiling an unproven shape risks wrong numbers; both now show up in
review as a golden diff, with the corpus suite's v2 leg judging the
numbers.
- **Workload coverage reporter** (`TestClassifyCorpus`, env-gated):
classifies
a JSON-lines corpus of real dashboard/alert queries and buckets
fallbacks
  by reason, to steer future allowlist work.

**What the dual-leg suite caught on its first transpiled run** (evidence
the
PR-2 guard works, worth stating in review):

- The classifier read a duration expression's offset (`x offset step()`)
as
  zero and transpiled it — offset expressions parse *without* the
experimental-parser flag, so they reach production. 20 corpus cases
served
  silently wrong numbers. Fixed by refusing `OriginalOffsetExpr` /
`RangeExpr` / `StepExpr` at classification (engine evaluates them
exactly);
  regression cases added, golden regenerated (30 routings flipped to
  fallback).
- Name-drop assembly treated temporally-disjoint same-labelset twins as
  separate series: `-{job="api"}` spanning `http_requests`/`http_errors`
  returned a 400 the engine would not raise, and hybrid
  `-metric_a or -metric_b` returned duplicate `{}` series. The engine's
  actual rule is: assemble the matrix by labelset, merging elements that
  never share an evaluation timestamp; error only on a same-timestamp
  conflict. Both the full-plan path (`mergeSameLabelsetSeries`) and the
hybrid post-strip path (`mergeMatrixByLabelset`) now reproduce it, with
  unit tests pinning the corpus scenarios.
- 12 remaining divergences, all one class, recorded in
`known_divergences_v2.json` with causes: the engine aggregates with
Kahan
compensated summation (sum, sum_over_time) and an overflow-free
incremental
mean (avg); ClickHouse's `sumForEach`/`avgForEach`/`arraySum` are naive,
so
±1e100 cancellation returns 0/residue and near-max-float64 `avg`
overflows
to ±Inf. Burn-down note: `sumKahanForEach` for the cancellation class;
the
overflow class needs an incremental-mean aggregate ClickHouse doesn't
have.

### Alternatives considered and discarded

- **General PromQL→SQL translation.** An allowlist inverts the failure
mode:
an overlooked construct becomes a fallback instead of a wrong number.
Every
shape on the list was validated slot-for-slot against the vendored
engine
  on live data before entering it.
- **ClickHouse's own PromQL dialect** (ClickHouse#57545,
`dialect='promql'`).
  Emits the same grid functions, but currently covers only
rate/irate/delta/idelta/last_over_time, has no fallback engine, and ties
us
to their TimeSeries table engine. We use the same primitives with our
own
  classifier and our own exactness gates.
- **Prefix-sum differencing for `*_over_time` windows.**
Large-minus-large
  cancellation drifts past the shadow tolerance on counter-sized values;
direct per-slot combination of at most W bucket partials adds the way
the
  engine adds.
- **Fanning each sample into every window that covers it.** Multiplies
rows
by W — billions of rows for a long range over a short step; the bucketed
  form's row count is series × buckets, the size of the output.
- **Handling staleness by filtering stale rows in WHERE (instant
units).**
Resurrects the older real sample the marker was written to bury; hence
the
  last-overall vs last-non-stale timestamp comparison.
- **Transpiling @-modifier and default-resolution subqueries.** Their
evaluation grid depends on server runtime settings the transpiler cannot
  see; they stay on the (exact) engine path.

### Test plan

- `go test ./pkg/prometheus/clickhouseprometheusv2` — transpiler unit
tests
  (SQL forms, classification, scalar ops, subquery grids), golden.
- `pytest integration/tests/promqlconformance/` — the v2 leg now
exercises
transpiled serving for every routable corpus case; ledger unchanged
(empty).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Pandey <vibhupandey28@gmail.com>
Bumps `clickhouse-sql-parser` to v0.5.5, fixes the false rejection that
was left over once it landed, and closes three holes in the same
validator that the first two changes brought to light.

## The bump

**Reserved keywords as expression operands**
([SigNoz#305](AfterShip/clickhouse-sql-parser#305)).
`interval` was fixed in v0.5.4, but the same defect affected 36 other
keywords once the column appeared as an operand rather than bare.
Sweeping 94 candidates against ClickHouse 26.8.1.337, only `on` still
rejects — and ClickHouse runs that too. This one was live: `sum(limit)`
on a metric label.

**Panic on an unparseable `DEFAULT` expression**
([SigNoz#306](AfterShip/clickhouse-sql-parser#306)).
Both known cases return a parse error now instead of dereferencing nil.
The `recover` in `ErrIfStatementIsNotValid` stays — it guards the next
one of these, not these two.

[SigNoz#307](AfterShip/clickhouse-sql-parser#307) also
allows `CAST` in a table function's argument list.

## Table functions are only table functions in a table position

The parser types a call inside a table function's argument list as a
`TableFunctionExpr` as well, so the generator allow list only ever
cleared a generator whose argument was a literal. Every real dashboard
computes its row count — `numbers(greatest(1, intDiv(end_ns - start_ns,
step_ns) + 1))` — and every one was refused, on `intDiv` rather than on
`numbers`.

`TableExpr.Expr` is the only table position a SELECT can reach, so the
allow list asks that instead. Of the four places the parser builds a
`TableFunctionExpr`, two are `CREATE TABLE` paths rejected as
not-a-SELECT before the walk starts, one is `parseTableArgPrimaryExpr`,
and one is the `FROM`/`JOIN` path that wraps into a `TableExpr`.

## Three holes that were already open

Skipping argument position is only safe if nothing there can read, and
that turned out not to be true — not because of this change, but
independently of it.

**Reading functions.** `file` is both a table function and a scalar
function, and the validator never inspected scalar calls at all. On
`main` today, `SELECT file('/etc/passwd')` is accepted and returns the
file. A numeric wrapper passes ClickHouse's type check, so the row count
alone is an oracle: `numbers(length(file(x)))` yields one row per byte.
The same applies to the 42 dictionary accessors, which can be backed by
HTTP, ODBC or another database, to `catboostEvaluate`, and to the
introspection functions. All are now refused by name wherever they
appear, under `clickhouse_sql_reading_function`.

**`x IN db.table`.** ClickHouse reads this as `x IN (SELECT * FROM
db.table)`, and a qualified name on the right of `IN` parses as a
`Path`, not a `TableIdentifier` — so `SELECT * FROM t WHERE a IN
system.users` bypassed the internal-database rule entirely. Now checked,
including the `GLOBAL IN` and `NOT IN` forms.

**Quoted generator names.** The allow list matched on the formatted
name, which carries the quoting, so ``SELECT * FROM `numbers`(31)`` was
refused. It now reads the identifier the way the internal-database
branch already did.

## Effect

Replaying 72 distinct shapes of production `clickhouse_sql` that the
validator currently rejects: **64 pass, up from 59 on v0.5.4**. Two came
from the bump, three from the table-position change, and those three are
379 of the 1390 sampled occurrences. The three new rules add no false
positives to the corpus.

Of the eight left, four are correct rejections (`system` reads, `SHOW
TABLES`), one is a dashboard variable rendering as the literal `<no
value>`, one is SQL ClickHouse also rejects, and two are an open
upstream gap.

## Tests

`TestErrIfStatementIsNotValid_ShouldPassButFails` is back, holding what
remains: three forms of a parenthesised left operand of a set operator,
and `on` as a column name. It also stopped panicking — `errors.Asc`
dereferences the error it is given, so a case starting to pass took the
suite out with a SIGSEGV instead of reporting. Both refusal tables now
share one harness, bounded by the same timeout the passing table uses.

Known gap: no input is currently known to panic the parser, so the
`recover` has no test exercising it.
#### Description

- Replace the multi-section PR template (change type, risk assessment,
changelog, checklist) with four concise headings: Description, Issues
closed, Screenshots, Additional Information.
- Add `.claude/rules/` with agent rules for comments (repo-wide, Go,
Python) and pull requests.
- Ignore `.dev/` and `.claude/worktrees/` in `.gitignore`.
#### Description

- Remove `.github/workflows/docs.yml`, which labeled `feat:` PRs with
"docs required" and failed the check until "docs shipped" was added.
- The job is not a required status check on `main`, and nothing else
references the workflow or its labels.
…z#12342)

## Summary

- Saved views now persist a versioned, typed spec (`schemaVersion` +
`spec{compositeQuery, selectedFields, display}`) instead of a bare
composite-query blob plus an opaque, frontend-owned `extraData` string
-- mirroring the pattern dashboards already use for their v2/perses
schema.
- `/api/v1/explorer/views` keeps working exactly as before: a thin
conversion layer translates to/from the legacy wire format, including
folding `extraData`'s ad hoc JSON into the typed spec and back for
backward compatibility.
- A one-time migration rewrites existing rows into the new shape and
drops the now-unused `extra_data`/`category`/`tags` columns.

### Scaffolding decisions 
- Using v2 for new handlers instead of renaming old handlers to
something else for these reasons - keep the diff minimum for easier
reviews, avoiding any git history or last updated at change in old route
registration.
- Keeping the conversion to old saved view type in handler itself rather
than `savedviewtypes` package to keep it un-exported and not let them be
available anywhere else to be used. It also enables `savedviewtypes` to
be independent on query-service models.
- Modified the existing handler and it's interface to include the v2
methods instead of adding another handlerV2 since apiserver already had
handler wired in, so don't want to pass on 2 version simultaneously.

### Breaking change
- Any unknown key in the `ExtraData` will be rejected and dropped
silently in the old APIs and give error in new version.
- If there was any way to add tag or category in saved view earlier,
that data will be lost.
- Old APIs will not support the old QB request payload, only v5 format
is supported.

---

Closes SigNoz/engineering-pod#4651

Alternative discarded SigNoz#12208

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
## Summary

Adds the semantic-convention evolution foundation:

- vendors the OpenTelemetry schema and SigNoz overlay
- generates Go and TypeScript family tables deterministically
- exposes the Go resolver API for family members, current names, and
historical names
- adds generation checks and unit tests

Related to SigNoz#6143.

## Stack

1. SigNoz#12441 — Foundations (base: main)
2. SigNoz#12442 — Phase 1 query (base: SigNoz#12441)
3. SigNoz#12443 — Phase 1 services (base: SigNoz#12442)
4. SigNoz#12444 — Phase 1 closure gate (base: SigNoz#12443)
5. SigNoz#12445 — Phase 2 signals (base: SigNoz#12444)
6. SigNoz#12446 — Phase 3 migration UX (base: SigNoz#12445)
7. SigNoz#12447 — Phase 4 rollout (base: SigNoz#12446)

**Current layer:** SigNoz#12441

## Testing

- `go test ./scripts/semconv`
- `go test ./pkg/types/telemetrytypes/semconv`
- `make semconv-check`

## Risk and rollback

This layer is additive apart from CI generation checks. Roll back by
reverting this PR; no stored telemetry is changed.
…#12456)

## Summary

Adds an integration-test matrix that pins the deliberate keyless-row
contract for filter operators, independent of any feature work:

- **Negative operators are a set complement over all rows.** A row that
does not carry the key at all must match `!=`, `NOT IN`, `NOT LIKE`, and
`NOT CONTAINS`. Users opt into presence explicitly with `AND key
EXISTS`.
- **Positive operators carry an implicit existence guard**
(`FilterOperator.AddDefaultExistsFilter`), so keyless rows never
false-positive against sentinel defaults.
- **`EXISTS` / `NOT EXISTS` partition rows exactly** by key presence,
and `!= x AND EXISTS` is the documented composition for "present and not
x".
- **Numeric attributes inherit the map-default sentinel**: a missing key
reads as `0`, so `num != 5` includes keyless rows while `num != 0`
excludes them. This conflation is deliberate and pinned by name
(`numeric_neq_zero_sentinel_conflation`) as the reference point for any
value-expression change.

Coverage: 46 cases — one shared matrix over traces and logs (resource
and attribute contexts), metric labels (series without the label), the
numeric sentinel, and the EXISTS composition. The contract, matrix, seed
data, and assertions live together in one file so the contract reads top
to bottom.

## Why

These semantics were enforced only implicitly by the operator list in
`AddDefaultExistsFilter`, with no test naming the intent. That gap
allows an implementation change to alter negative-filter results
silently and lets new tests calibrate expectations against the
implementation instead of the contract. The attribute names used here
are deliberately outside every semantic-convention family, so this file
pins the base contract regardless of the semconv overlay state and
serves as the oracle that family-field behavior
(`queriertraces/13_semconv_evolution.py` in the semconv stack) must
mirror.

## Testing

- `uv run pytest --basetemp=./tmp/ --reuse
integration/tests/queriercommon/06_keyless_semantics.py` — 46/46 passed
against a stack built from main-based sources, and 2×46 against a
long-lived shared stack, confirming the set-based assertions stay stable
under environment reuse.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…sts (SigNoz#12460)

#### Description

- Add `.claude/rules/pytest.md` with conventions for the Python
integration suite. The lead rule: **no `_`-prefixed helper functions in
test modules** — inline the logic; repetition across tests is cheaper
than indirection; genuinely shared machinery becomes a fixture. Fixtures
live in `tests/fixtures/` only, never under `integration/tests/` — with
one exception: SigNoz-level fixtures (a suite spinning up SigNoz with
different envs via `create_signoz`/`create_migrator`) always belong in
that suite's `conftest.py`. Plus: fixture-factory over indirect
parametrization, skip at collection, config via explicit `--flags`,
snake_case parametrize ids, and collection gotchas (`python_files`
prefix matching, `--import-mode=importlib`).
- Apply the no-`_helper` rule across `tests/integration`: all 40
module-level `_` helpers eliminated in dashboard, inframonitoring,
promqlconformance, querier_json_body, querierlogs, queriermetrics, and
queriertraces. Pure transforms and request wrappers were inlined at
their call sites; case-table verifier callables became data flags with
inline branches; the two 115-line resource-evolution mega-helpers folded
into parametrized tests; shared machinery moved to `tests/fixtures/` as
fixture-factories (`wipe_all_dashboards`, `load_pods_metrics`,
`run_query_case`) registered via `pytest_plugins`.
- Apply the py-comments rule across `tests/`: drop module docstrings
that restate the filename, relocate the ones carrying real constraints
next to the code they constrain, delete function/class docstrings that
restate the identifier, trim narrated steps and Args/Returns
boilerplate.
- Fix camelCase parametrize ids in `queriermetrics/01_fill.py`
(`fillGaps`/`fillZero` → `fill_gaps`/`fill_zero`).

#### Additional Information

- All 629 tests in the touched suites collect cleanly;
`py-fmt`/`py-lint`/compileall pass. Runtime verification against the
docker stack was not run.
…igNoz#12462)

#### Description

- Add the fixture-vs-function rule to `.claude/rules/pytest.md`: a
fixture earns its indirection only by owning setup/teardown (`yield` +
cleanup) or provisioning a resource; a stateless action or lookup is a
plain importable function in the matching `tests/fixtures/` module
taking `signoz`/`token` as ordinary arguments.
- Apply it to the three fixture-factories introduced in SigNoz#12460 that have
no lifecycle: `delete_all_dashboards` (renamed from
`wipe_all_dashboards`) and `run_query_case` are now plain functions,
their modules deregistered from `pytest_plugins`, and all call sites
updated.
- Generalize `Metrics.load_from_file` with a `label_substitutions`
parameter (placeholder rewriting, e.g. `__START_TIME__` → runtime ISO
string) and drop the bespoke `load_pods_metrics`, which duplicated the
base-time rebase logic — `02_pods.py` now loads JSONL the same way as
every other inframonitoring suite file.

Follow-up promised in
SigNoz#12460 (comment).
#### Description

- `uv lock --upgrade` across the tests project: pytest 9.0.3→9.1.1, ruff
0.15.11→0.16.2, selenium 4.43→4.46, numpy 2.4.4→2.5.1, uvicorn
0.46→0.52.1, testcontainers 4.14.2→4.15.0, requests, sqlalchemy,
websockets, and the rest of the transitive set (zstandard dropped as no
longer required).
- Ignore `PLR0917` (too-many-positional-arguments), newly enforced by
ruff 0.16 — muted alongside the other `PLR09xx` complexity rules the
project already ignores (193 pre-existing hits, all in test/fixture
signatures).

#### Additional Information

- `py-fmt` (no reformats), `py-lint` (clean), and full integration-test
collection (1782 tests) pass on the upgraded toolchain. Runtime
verification against the docker stack was not run.
… the idp after keycloak login (SigNoz#12399)

Deflakes the SSO login tests (`callbackauthn` and `basepath`). They all
share the `idp_login` fixture, and after it clicked Keycloak's login
button it could hand control back to the test too early — in two
different ways ([example CI
failure](https://github.com/SigNoz/signoz/actions/runs/30898802502/job/91957986941)).

### What

The fixture used to wait for the login button to disappear and treat
that as "login is done". Two things go wrong with that:

1. **The page can vanish while we're looking at it.** Asking "is the
button still visible?" takes two round-trips to the browser: find
`kc-login`, then ask whether it's displayed. If Keycloak's redirect
lands between the two, the second call is asking about a node that no
longer exists. Selenium normally recognises that as a stale element and
quietly retries — but Keycloak → SigNoz is a *same-site* hop
(`localhost` → `localhost`), where the renderer survives the swap and
that detection can miss. The raw chromedriver error (`Node with given id
does not belong to the document`) then escapes and fails the test.
That's the CI failure above.

2. **The button disappearing doesn't mean login finished.** It only
means we left the login *page*. In the SAML flow Keycloak next serves a
small auto-submitting page — still on the IdP — and *that* POST is what
actually creates the user in SigNoz. So a test could go looking for the
user before SigNoz had ever seen the callback, and fail with `User ...
not found`. Reproduces locally on `test_idp_initiated_saml_authn`.

The wait now checks what the tests actually need: **the browser has left
the IdP host** (the hostname in the URL changed) *and* the login button
is gone.

### Guardrails

- Nothing is held across the navigation — the button is looked up fresh
on every poll with `find_elements`, so "gone" is simply an empty list,
never a question asked of a dying node.
- Any browser error during a poll is read as "still navigating, try
again" instead of failing the wait.
- The wait sits through everything that's still on Keycloak (the
`login-actions` hops, the SAML interstitial) and passes only once SigNoz
has handled the callback and redirected — so the user exists by the time
the test asserts on it.
- Wrong credentials still fail loudly: Keycloak re-renders the login
form on its own host, so the wait times out exactly as before.

One change, in the shared fixture — the OIDC and SAML flows in both
`callbackauthn` and `basepath` all go through it.

### Notes

- Failure 1 needs the redirect to land inside a ~2–5 ms window of a poll
that only runs every 500 ms, so it's effectively a loaded-CI-runner
lottery — which is why it's rare and CI-only. Failure 2 shows up
locally.
- Unrelated to the PR it fired on (SigNoz#12382, query-builder only); the
identical SAML test passed in the same run.

### Testing

- Reproduced failure 1 outside pytest, with a probe driving real
headless **Chrome for Testing 151.0.7922.71** (the exact build from the
CI log) through a click → POST → same-site redirect that mimics the
Keycloak login flow, with server think-time near the 500 ms poll
boundary. Both waits ran verbatim, at their real polling rate:

  | Post-click wait | Logins | Failures |
  |---|---|---|
| old (`EC.invisibility_of_element`) | 400 | **3 × the exact CI
inspector error** |
  | new (left-the-IdP check) | 400 | **0** |

- Cross-checked the mechanism against the selenium 4.40 source with a
stubbed driver: the detached-node error does escape the old wait (it
only catches stale/not-found), while the new one absorbs it and passes
on the next poll. A bad-credentials control times out on both old and
new, so failure detection isn't weakened.
- Ran the full suites locally on the final fixture, with a fresh sqlite
+ wal store per suite (matching the failing CI leg): `basepath` 6/6, and
all 36 SSO/domain tests in `callbackauthn` — including
`test_idp_initiated_saml_authn`, which flaked with `User not found` on
the old wait in the same setup. (The one local non-pass,
`test_apply_license`, is unrelated: it asserts on wiremock's request
journal and the reused license-mock container is never reset between
runs — CI gets a fresh mock.)
- `make py-fmt` / `make py-lint` clean.

Fixes SigNoz/engineering-pod#5850
…ts (SigNoz#12455)

The `telemetry.*.last_observed` stats took `max()` over client-supplied
event-time columns, so a single row with a skewed or corrupt timestamp
(a 2050-dated log, a `2^32−1`-second span) poisoned them indefinitely.

### What
- Traces/logs `last_observed` now reads `max(inserted_at)` — the
collector-stamped insert time added in SigNoz/signoz-otel-collector#875;
metrics reads `inserted_at_unix_milli` (metrics migration 1007).
- Each signal checks `hasColumnInTable` first and falls back to the
previous expression, so tenants without the schema migration keep
today's behavior and switch over automatically.

### Notes
- `created_at` is unusable here: pre-migration rows evaluate its
`now64(3)` default at read time, so `max(created_at)` always reads as
"now".
- Pre-migration rows read `inserted_at` as epoch, which `max()` ignores;
the all-old case lands on the existing `Unix() != 0` skip-guard.
- Future-dated garbage never TTLs out (TTL is keyed on the event
timestamp), which is why the old stat stayed wrong once poisoned.

### Testing
- Expressions validated against `clickhouse local`, including garbage
rows (`2^64−1`, `9.3e18` ns) and empty/pre-migration tables.
- `go build`, `go vet`, golangci-lint clean.

Fixes SigNoz/engineering-pod#5864
… 109 (SigNoz#12469)

## Summary
- `restructureSavedViewSpec` (migration 109) bulk-inserts legacy
`saved_views` rows into the new `saved_view` table, which has an
enforced `org_id -> organizations(id)` FK (sqlite runs with
`foreign_keys=ON`).
- Some tenants hit `constraint failed: FOREIGN KEY constraint failed
(787)` on startup because a row's `org_id` didn't match any live
organization -- e.g. an org deleted before the old table had a cascading
FK, or an install where `org_id` was never backfilled
(`015_update_dashboards_savedviews` only backfills it when there's
exactly one org).
- Fix: fetch live organization IDs up front inside the same transaction,
and skip (with a `WarnContext` log, counted in `skipped`) any row whose
non-empty `org_id` isn't among them -- same treatment as the existing
empty-`org_id` skip.

Followup to SigNoz#12342

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
… overlaps counted attrs (SigNoz#12451)

## Pull Request

---

### 📄 Summary

The infra-monitoring v2 clusters/namespaces list APIs 500 with
ClickHouse error 179 (`MULTIPLE_EXPRESSIONS_FOR_ALIAS`) when the request
groups by an attribute that is also a counted resource attribute (e.g.
clusters grouped by `k8s.node.name` or `k8s.namespace.name`, namespaces
grouped by `k8s.deployment.name`).

Root cause: `getPerGroupDistinctCounts` aliases each `uniqExactIf(...)`
count column with the bare attr key, which collides with the groupBy
column alias for the same key. Fix: alias count columns as
`__count_<attr>`. Row scanning is positional and the result map is keyed
in Go from `attrNames`, so nothing downstream changes.

Integration tests: clusters API grouped by `k8s.namespace.name` and
namespaces API grouped by `k8s.deployment.name`, with exact per-group
`counts` assertions (identity-tuple semantics). Both reproduce the 500
on the pre-fix build and pass on the fixed build.

#### Issues closed by this PR

Closes SigNoz/pulse-pod#245

### ✅ Change Type
_Select all that apply_

- [ ] ✨ Feature
- [x] 🐛 Bug fix
- [ ] ♻️ Refactor
- [ ] 🛠️ Infra / Tooling
- [ ] 🧪 Test-only

---

### 🧪 Testing Strategy

- Tests added/updated: Yes — integration (`04_namespaces.py`,
`05_clusters.py`): new groupBy-on-counted-attr cases + per-group
`counts` assertions on existing cases
- Manual verification: Yes — replayed the failing production query with
the fixed aliasing against ClickHouse
- Edge cases covered: groupBy overlapping a counted attr; same-named
deployment across namespaces counted as distinct entities

---

### ⚠️ Risk & Impact Assessment

- Blast radius: Infrastructure Monitoring — clusters & namespaces list
APIs (counts query)
- Potential regressions: None — SQL alias rename only; scanning is
positional and result map keys are unchanged
- Rollback plan: Revert this commit

---

### 📝 Changelog

| Field | Value |
|------|-------|
| Deployment Type | Cloud / OSS / Enterprise |
| Change Type | Bug Fix |
| Description | Fixed a 500 error in Infrastructure Monitoring
clusters/namespaces APIs when grouping by an attribute that is also part
of the resource counts (e.g. node, namespace, or deployment name). |

---

### 📋 Checklist
- [x] Tests added or explicitly not required
- [x] Manually tested
- [ ] Breaking changes documented
- [x] Backward compatibility considered
## Pull Request

---

### 📄 Summary

Adds a `filterByPodStatus` secondary filter to the v2 infra-monitoring
list APIs (pods, nodes, namespaces, clusters, deployments, statefulsets,
jobs, daemonsets).

Pod status is a derived kubectl-style value (`k8s.pod.phase` + status
reasons, resolved via `argMax`), not a real label, so it can't go
through the normal query-builder filter. This PR resolves the full-scope
status keyset up-front and intersects it with the metadata + ranked
groups, keeping `total` and pagination correct.

- Multi-select: the field is an array, pushed down as `WHERE
lower(display_status) IN (...)` (OR within status, AND with the
attribute filter).
- When the optional status metrics were never ingested, the endpoint
returns a non-blocking warning + empty page instead of silently
filtering everything out.

#### Screenshots / Screen Recordings (if applicable)

N/A — backend + generated FE API types only; the UI is a separate
change.

#### Issues closed by this PR

Part of SigNoz/engineering-pod#5778.

---

### ✅ Change Type

- [x] ✨ Feature
- [ ] 🐛 Bug fix
- [x] ♻️ Refactor
- [ ] 🛠️ Infra / Tooling
- [ ] 🧪 Test-only

---

### 🐛 Bug Context

N/A — not a bug fix.

---

### 🧪 Testing Strategy

- Tests added/updated:
- Unit test for the status push-down (`applyPodStatusFilter`, built with
go-sqlbuilder).
- Integration tests across all 8 entity APIs: list mode, grouped mode,
validation, missing-metric warning, and multi-select union.
- Manual verification: smoke-tested against staging data (single, multi,
and grouped filters).
- Edge cases covered: missing status metric → warning + empty; grouped
mode keeps a group if ≥1 pod matches; multi-select returns the union of
the selected statuses.

---

### ⚠️ Risk & Impact Assessment

- Blast radius: v2 infra-monitoring list endpoints only.
- Potential regressions: none when the filter is unset (empty = off,
fully additive). When set, an extra status query runs; it is gated
behind the filter being present.
- Rollback plan: revert the PR — no schema or data migrations involved.

---

### 📝 Changelog

| Field | Value |
|------|-------|
| Deployment Type | OSS, Cloud, Enterprise |
| Change Type | Feature |
| Description | v2 infra-monitoring lists can now be filtered by pod
status (multi-select). |

---

### 📋 Checklist

- [x] Tests added or explicitly not required
- [x] Manually tested
- [x] Breaking changes documented
- [x] Backward compatibility considered

---

## 👀 Notes for Reviewers

- `filterByPodStatus` is optional and additive — no change to existing
responses when omitted.
- The status keyset is resolved once at full scope, then intersected —
this is what keeps `total`/pagination correct despite status being a
post-aggregation value.
- OpenAPI spec + FE API types are regenerated (scalar → array); no
hand-written FE.
…12471)

<!--A few plain bullets saying what changed and why, for a reviewer
skimming it - not a wall of text, not a restatement of the diff, not
generated boilerplate.-->
#### Description
Disabling logs support for all GCP integration services.

Reasons:
1. googlecloudpubsubpush receiver needs to reach Alpha stability, thus
not included in contrib build -> can't suggest for logs
2. signoz otel collector includes this receiver in the build but it
needs upgrade to v0.158.0 for a metrics related
[fix](open-telemetry/opentelemetry-collector-contrib#49826)
-> needs more testing hence can't suggest either

Decision was taken to go ahead without logs for now - follow [ticket
here](SigNoz/platform-pod#2901) for details.


<!--Reference issues using `Closes #issue-number` to enable automatic
closure on merge. -->
#### Contributes to 
SigNoz/platform-pod#2901

<!--Anything reviewers should keep in mind while reviewing -->
#### Additional Information
The explanation above should be enough

<!--Please delete paragraphs that you did not use before submitting.-->
…est (SigNoz#12475)

#### Description

- `make py-lint` is failing on `main` with `F821 Undefined name
_load_pods_metrics` at `inframonitoring/02_pods.py:471`, which blocks
every open PR.
- `test_pods_filter_pagination_and_ordering` calls a helper that no
longer exists. Replaced with
`Metrics.load_from_file(get_testdata_file_path(...))`, matching the two
other tests that seed `pods_phases.jsonl`.

#### Additional Information

- How it broke: SigNoz#12460 replaced the module-level `_load_pods_metrics`
with a `load_pods_metrics` fixture, then SigNoz#12462 dropped that fixture in
favour of calling `Metrics.load_from_file` directly. SigNoz#12278 branched
before either landed and merged after, re-introducing one call to the
long-gone helper. Git merged cleanly because the branches touched
different lines, so nothing flagged it.
- No behaviour change: the broken call passed no `start_time`, so no
placeholder substitution happened; `load_from_file` with
`label_substitutions=None` does the same earliest-to-`base_time` rebase.
The replacement is byte-identical to the two sibling call sites for the
same dataset.
#### Description

- Removes `POST /api/v1/invite/bulk` — already deprecated, superseded by
`POST /api/v1/invite`, and no callers left.
- `Setter.CreateBulkInvite` stays; `CreateInvite` still delegates to it
for the single-invite case.

#### Issues closed by this PR

Contributes to SigNoz/platform-pod#2667

#### Additional Information

- OpenAPI spec and the generated frontend client are regenerated, not
hand-edited.
- The `integrationci / fmtlint` failure here is not from this PR — `make
py-lint` is broken on `main`. Fixed separately in SigNoz#12475; this PR needs
that merged (or a rebase on it) to go green.
- First of three PRs splitting a v1 user-API cleanup. The other two also
regenerate the spec and generated client, so whichever merges second
needs the generators re-run.
#### Description

- Removes `GET`, `PUT` and `DELETE /api/v1/user/{id}` — all deprecated
and superseded by `/api/v2/users/{id}`, which the frontend already uses.
- Drops the dead code this leaves behind: the `SelfAccess` middleware
and `Claims.IsSelfAccess` (no callers left), the deprecated update
setters, and three `DeprecatedUser` helpers.
- Points the integration tests that deleted users at `DELETE
/api/v2/users/{id}`.

#### Issues closed by this PR

Contributes to SigNoz/platform-pod#2667

#### Additional Information

- Behaviour change: the removed `GET`/`PUT` were `SelfAccess`, the v2
equivalents are `AdminAccess`. Self-serve reads and updates go through
`/api/v2/users/me`, which is what the UI already calls — but worth a
second pair of eyes.
- `DELETE /api/v1/user/{id}` was the most widely reached of the three.
Please confirm nothing external (zeus) still calls it before merging.
- OpenAPI spec and the generated frontend client are regenerated, not
hand-edited.
#### Description

- Invited members can now be given more than one role. The picker was
single-select even though `POST /api/v2/users` has accepted a list of
roles since custom roles landed.
- Frontend only — nothing changed on the backend, the grant chain was
already multi-role.
- Onboarding analytics now emits `teamMembers[].roles` as a list,
replacing the singular `role` key.

#### Issues closed by this PR

Closes SigNoz/platform-pod#2920


#### Screenshots / Screen Recordings

#### Members Page

https://github.com/user-attachments/assets/8b556549-789c-4ddf-b4af-5994eccc75f3


#### Onboarding Flow

https://github.com/user-attachments/assets/a6dde5d3-e9b4-48c1-965f-352bb0a6a89f



#### Additional Information

- A row still requires at least one role to be considered valid.
- Existing invite tests moved to `findByTitle` for role options,
matching how `EditMemberDrawer` already drives the multi-select.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.