Skip to content

fix(issues): one stable issue per restart loop - #1967

Open
nadaverell wants to merge 8 commits into
mainfrom
nadav/restart-loop-issue
Open

nadaverell wants to merge 8 commits into
mainfrom
nadav/restart-loop-issue

Conversation

@nadaverell

@nadaverell nadaverell commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Problem

A container stuck restarting looks like a different problem on almost every poll:

  • Waiting CrashLoopBackOff → crashloop
  • Running but not ready → readiness_failed / high_restart / nothing
  • A liveness Unhealthy event → liveness_probe_failed
  • Terminated Completed/0 (graceful liveness kill) → breaks the stable-crashloop check
  • Running + Ready between crashes → nothing, so the issue resolves

Each flip is a new issue id, and the Deployment's workload_degraded row comes and goes with it. Downstream, every flip is a new alert and a new AI investigation. One production gateway (restartCount 4592, last state Completed/0) produced thousands of open/resolve generations. On our own clusters, 15 workloads loop like this right now.

Change

One issue per restart loop.

  • A container is looping while it has ≥3 restarts and its current or last termination ended within 30 minutes, whatever the exit code, after a run shorter than 30 minutes.
  • While it loops, its pod keeps one row through every kubelet state: crashloop, or oom_killed for an OOM loop.
  • The workload's workload_degraded row folds into the loop at any severity, because the loop explains the unavailability.
  • The loop issue keeps what the rows it replaces carried, as evidence rather than identity:
    • probe failures (liveness, readiness, startup), with the kubelet's messages;
    • the last exit and its run length;
    • the runtime's error;
    • the workload impact, i.e. how many of its pods are looping. This is the stable form of the folded "N/M available" row.
    • Each loop also says why it is critical or warning.

Which containers count.

  • Regular containers whose effective restartPolicy is Always. A per-container policy (1.34+) overrides the pod's.
  • Native sidecars.
  • Ordinary init containers failing while the pod is Pending (Init:CrashLoopBackOff). Their long attempts are no longer also reported as a stalled init.
  • Excluded:
    • Job workers (OnFailure/Never).
    • Deleting pods (a rollout SIGTERM is not a crash).
    • A single restart after a long run (node bounce).

Severity follows impact, and holds steady for the loop.

  • critical: the loop is fast (last run under 10 minutes, the kubelet's backoff-reset line, so the container is mostly down) and it covers at least half of the workload's live pods.
  • warning: a slow loop (serves 10+ minutes between restarts), or one bad replica among many.

Clears on recovery. The loop ends once the container has run longer than max(10 min, 2× its last run). Fast loops clear 10 minutes after they stop crashing, instead of holding 30 minutes after the last crash.

More specific problems keep their own row: invalid probe targets (pinned to the loop's severity, carrying its evidence), image pulls, create errors, and another container's active OOM.

Also:

  • ReplicaFailure (can't create pods) folds only into scheduling-source admission rejections.
  • Event timestamps: events recorded through events.k8s.io keep EventTime / Series in the cache and are dated by their latest occurrence.
  • Diagnoses: exit 143 and start failures get their own. The OOM diagnosis names node pressure when the container has no limit.

Live verification

main and this branch ran side by side, read-only, against 4 real clusters (GKE nonprod / management / prod, EKS skh-nonprod) for 30 minutes, polling /api/issues every 20 s:

15 restart-looping workloads main branch
Issue changes (flips) on looping workloads 70 0
Loop severity critical critical for the 10 fast single-replica loops; warning for the 5 Envoy gateways (11-minute runs)
  • Other detectors: no change attributable to this PR. GitOps, RBAC and APIService rows differ only by poll timing.
  • Envoy gateways: their pre-existing critical rollout_stalled row stays visible next to the warning loop, as it does on main; rollout_stalled keeps its severity gate.
  • 13 synthetic scenarios on an EKS test cluster (radar-test-nonprod, Kubernetes 1.34), 30 minutes, both builds watching from creation. Steady state (after the 10-minute warm-up):
Scenario main · ids / changes branch
Liveness kill, exit 0, readiness flapping 4 / 16 1 / 0
Self-exit 0 2 / 9 1 / 0
Crash exit 1 1 / 7 (warning↔critical) 1 / 0
Ready 150 s, then crash (2 replicas) 1 / 8 1 / 0
Start failure (RunContainerError / StartError) 2 / 8 1 / 0
Native sidecar loop 2 / 8 1 / 0
Init container failing (Init:CrashLoopBackOff) 1 / 0 1 / 0, with loop evidence
OOM loop 1 / 0 (oom_killed) 1 / 0 (oom_killed, with loop evidence)
1 of 4 StatefulSet replicas crashing 1 / 8, critical 1 / 0, warning
Healthy pod none none
Job retry (OnFailure) 1 / 6 1 / 6 (excluded by design)
Loop next to an image-pull failure 2 / 7 2 / 7 (known limit)

Recovery: a pod that crashed 4 times and then turned Ready at 12:08:11 UTC was cleared by the branch at about 12:18. In the first minutes, before a container's 3rd restart, both builds still show the old per-poll behaviour.

Tradeoffs and known limits

  • Cumulative restart count: "≥3 restarts plus a recent short run" approximates a rate. A pod with old restarts that crashes once at startup after a long run reads as a loop until the early clear, 10 minutes.
  • Severity can change when a loop's runs straddle 10 minutes or its looping share crosses half; these transitions are rare, not per poll.
  • One row per pod: a loop next to a sibling with a different problem can still alternate.
  • Alternating causes: a container alternating OOMKilled and other exits still alternates oom_killed / crashloop.
  • Service service_no_endpoints still flaps with its looping backend; the existing Service→backend issue link remains.

Tests

  • Integration trace: walks a liveness-driven exit-0 loop through 7 kubelet states via the real DetectProblems → Compose. It asserts one id, one cause and the evidence at every tick, and that the loop clears after 11 Ready minutes.
  • Unit tests:
    • classifier boundaries (window, run length, stale gap, early clear), restart policies, init / sidecar / OOM loops;
    • container selection, the severity tier (1 of 4 / 2 of 4 / slow / fast);
    • precedence (invalid probe target, image pull, sibling OOM), dedupe folds, probe evidence attribution (fieldPath, predecessor pod UID), event series time.
  • go test ./... (both modules) ✓ · make tsc ✓ · CI ✓

A container stuck restarting cycles through states that each read as a
different problem on a single poll: CrashLoopBackOff, Running-not-ready,
readiness and liveness Unhealthy events, Terminated exit 0, and
Running+Ready between crashes. The pod row changed category (and so issue
id) every cycle and vanished on the healthy-looking ticks, and the owning
Deployment's workload_degraded row came and went with it. In production one
looping gateway produced thousands of open/resolve generations.

The detector now classifies an active restart loop from restart history:
3+ restarts and a termination (current or last, any exit code, not OOM)
within the past 30 minutes, for containers the kubelet restarts by design
(restartPolicy Always containers and native sidecars). While it holds, the
pod keeps one critical crashloop row through every tick, folding the
per-tick faces of the loop into it. Probe failures, last exit and the
container are attached as RestartLoop evidence and a restart_cause fact;
the diagnosis depends only on container and exit code so it stays stable
across polls and agrees across replicas. Pod health levels are unchanged.

More specific problems keep their own rows: invalid probe targets, image
pulls, create errors, active OOM (now including sidecars), init stalls.

Also: a started native sidecar is no longer reported as a stalled init
container, and ReplicaFailure is folded only into scheduling-source
admission rejections, never into a runtime symptom on an existing pod.
@nadaverell
nadaverell requested a review from hisco as a code owner October 3, 2026 22:45
@qodo-free-for-open-source-projects

Copy link
Copy Markdown

PR Summary by Qodo

Keep one critical issue throughout a container restart loop

🐞 Bug fix ✨ Enhancement 🧪 Tests 🕐 40+ Minutes

Grey Divider

AI Description

• Keep a critical crashloop issue open across changing kubelet states and healthy-looking intervals.
• Attach container and probe evidence without treating probe failures as a proven cause.
• Preserve distinct creation failures and other specific problems; add lifecycle and precedence
 tests.
Diagram

graph TD
  Cache[("Kubernetes cache")] --> Detector["Pod detector"] --> Classifier["Loop classifier"] --> Normalize["Issue normalization"] --> Dedupe["Rollup dedupe"] --> Group["Issue grouping"] --> UI["Issues UI"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Persist issue identity across polls
  • ➕ Could maintain continuity without classifying every intermediate pod state as a crashloop.
  • ➖ Requires stateful identity and expiry handling across restarts and replicas.
  • ➖ Does not itself make severity or diagnostic evidence consistent.
2. Emit per-container issue rows
  • ➕ Could represent a restart loop alongside a different problem in a sibling container.
  • ➖ Changes issue identity and grouping semantics more broadly.
  • ➖ Adds substantial migration and review scope for the reported continuity bug.

Recommendation: Use the PR's stateless, evidence-based classifier for this fix: it stabilizes category and severity without an issue registry or changes to pod health. Per-container rows are a reasonable later step if alternating sibling-container problems become a priority.

Files changed (15) +977 / -32

Enhancement (6) +113 / -17
diagnostic_context.goDescribe restart-loop observations in diagnostic facts +35/-0

Describe restart-loop observations in diagnostic facts

• Builds a restart_cause fact from container termination and attributed probe evidence. Its wording presents probe failures as observations rather than established causes.

internal/issues/diagnostic_context.go

grouping.goRetain loop evidence on grouped issues +1/-0

Retain loop evidence on grouped issues

• Copies the representative member's restart-loop evidence into the grouped issue so it remains available after workload grouping.

internal/issues/grouping.go

normalize.goTransfer detected loop evidence to issue rows +1/-0

Transfer detected loop evidence to issue rows

• Preserves restart-loop evidence when converting a Kubernetes detection into an issue.

internal/issues/normalize.go

IssuesView.tsxShow probe observations beside restart metadata +6/-1

Show probe observations beside restart metadata

• Adds liveness and readiness probe-failure indicators to the diagnosis metadata line when restart-loop evidence contains them.

packages/k8s-ui/src/components/issues/IssuesView.tsx

types.tsType restart-loop evidence in the UI +19/-0

Type restart-loop evidence in the UI

• Adds optional issue-level restart-loop and probe-failure interfaces matching the new API evidence.

packages/k8s-ui/src/components/issues/types.ts

types.goAdd structured restart-loop evidence to the Issue API +51/-16

Add structured restart-loop evidence to the Issue API

• Defines restart-loop and probe-failure payloads and adds an optional restart_loop field to issues. The payload identifies one container's termination and recent probe observations.

pkg/issuesapi/types.go

Bug fix (4) +358 / -8
dedupe.goRestrict ReplicaFailure folding to pod-creation rejections +23/-1

Restrict ReplicaFailure folding to pod-creation rejections

• Tracks scheduling-source admission rejection severity separately from general child symptoms. A ReplicaFailure rollup no longer disappears because an existing pod has a crashloop or runtime RBAC denial.

internal/issues/dedupe.go

detect.goKeep loop detections through healthy-looking polls +71/-7

Keep loop detections through healthy-looking polls

• Uses the restart-loop classifier before skipping healthy pods, emits a critical crashloop row with evidence, and preserves invalid probe targets, active OOM, init stalls, and other specific reasons. Also attributes recent probe events by container and stops treating a started native sidecar as stalled init.

internal/k8s/detect.go

restart_loop.goClassify recent restarts of eligible containers +257/-0

Classify recent restarts of eligible containers

• Introduces a stateless classifier using at least three restarts, a termination within 30 minutes, and a stale-gap guard. It excludes ordinary init and non-Always workload containers, collects per-container probe evidence, and supplies stable exit-code-based diagnosis.

internal/k8s/restart_loop.go

reasons.goExpose the existing active-OOM check to detection +7/-0

Expose the existing active-OOM check to detection

• Exports a wrapper around the existing pod-wide active-OOM check so the detector can preserve OOM precedence, including for sidecars. Pod health classification itself is unchanged.

pkg/health/reasons.go

Tests (5) +506 / -7
dedupe_test.goCover independent ReplicaFailure rollups +22/-0

Cover independent ReplicaFailure rollups

• Adds cases ensuring a crashlooping pod and a runtime RBAC denial do not suppress a Deployment's ReplicaFailure.

internal/issues/dedupe_test.go

restart_loop_integration_test.goVerify one issue across a complete restart cycle +252/-0

Verify one issue across a complete restart cycle

• Exercises seven kubelet states through detection and composition, checking stable ID, critical severity, diagnosis, evidence, and Deployment rollup folding. Verifies the issue clears after the 30-minute window.

internal/issues/restart_loop_integration_test.go

detect_crashloop_severity_test.goExpect critical severity during a serving loop tick +3/-1

Expect critical severity during a serving loop tick

• Changes the expected severity of a serving pod with a qualifying restart loop from high to critical.

internal/k8s/detect_crashloop_severity_test.go

detect_test.goCheck clean-exit thrashing becomes a crashloop +11/-6

Check clean-exit thrashing becomes a crashloop

• Updates probe-failure coverage to expect one critical crashloop detection for repeated exit-zero restarts, with attributed liveness evidence instead of a high-restart or liveness row.

internal/k8s/detect_test.go

restart_loop_test.goTest classifier boundaries, attribution, and precedence +218/-0

Test classifier boundaries, attribution, and precedence

• Covers restart and time thresholds, OOM exclusion, restart policies, native sidecars, and container-specific probe attribution. Also checks stable diagnosis, started-sidecar handling, and precedence of invalid probes and image pulls.

internal/k8s/restart_loop_test.go

@qodo-free-for-open-source-projects

qodo-free-for-open-source-projects Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (1) 📜 Skill insights (0)

Grey Divider


Remediation recommended

1. A test comment recounts old behavior ✓ Resolved
Description
The comment above TestCompose_RestartLoopKeepsOneIssueAcrossTheCycle says the detections `used to
be different issues and must now` be one issue. That change-history wording sits in the test
source even though the assertions below establish the expected behavior.
Code

internal/issues/restart_loop_integration_test.go[R127-129]

+// tick by tick, these used to be crashloop, readiness_failed,
+// liveness_probe_failed, workload_degraded, and nothing at all — five issue
+// ids for one problem. They must now be one critical crashloop issue.
Evidence
The added comment explicitly uses used to be and must now to describe change history, which the
cited rule disallows in code comments.

Rule 3036538: Disallow references to tickets or PR history in code comments
internal/issues/restart_loop_integration_test.go[124-129]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The test comment describes behavior before and after the change rather than only the invariant it tests.
## Fix Focus Areas
- internal/issues/restart_loop_integration_test.go[124-129]
## Recommended Fix
Replace the historical wording with a statement that every state in the cycle must retain one critical crashloop issue.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. An evidence comment repeats its method ✓ Resolved
Description
The comment above evidence() says only that the method produces the wire form attached to an
issue. The method name, *issuesapi.RestartLoop return type, and construction immediately below
already convey that behavior, so removing the comment loses no rationale or constraint.
Code

internal/k8s/restart_loop.go[240]

+// evidence is the wire form attached to the issue.
Evidence
The added comment describes only the immediately following method's evident return value and adds no
intent or hidden constraint.

Rule 3036542: Avoid explanatory comments that restate obvious code behavior
internal/k8s/restart_loop.go[240-249]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The comment above `evidence()` restates what its name, return type, and implementation show.
## Fix Focus Areas
- internal/k8s/restart_loop.go[240-241]
## Recommended Fix
Delete the comment, or replace it only if there is a non-obvious wire-format constraint to document.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Recreated pods inherit old probe evidence ✓ Resolved
Description
latestProbeFailures keys probe events by namespace, pod name, container, and probe type without
checking the event's pod UID before activeRestartLoop attaches them as evidence. When a pod is
recreated with the same name within the ten-minute event window, its restart-loop issue can report a
liveness or readiness failure from the previous pod.
Code

internal/k8s/detect.go[R1580-1582]

+		containerKey := key + "/" + eventContainerName(e.InvolvedObject.FieldPath) + "/" + reason
+		if cur, exists := byContainer[containerKey]; !exists || t.After(cur.at) {
+			byContainer[containerKey] = pf
Evidence
The new per-container event key omits UID, while the new loop-evidence lookup uses that key for the
current pod. The ten-minute filter checks event age, not whether it belongs to this pod generation.

internal/k8s/detect.go[1564-1589]
internal/k8s/restart_loop.go[138-160]
internal/k8s/detect.go[454-470]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Probe events from a previous pod can be attached to a recreated pod's restart-loop issue when the namespace, name, and container match.
## Fix Focus Areas
- internal/k8s/detect.go[1550-1590]
- internal/k8s/restart_loop.go[138-162]
## Recommended Fix
Carry the involved pod UID with each probe event and reject events whose nonempty UID differs from the current pod's UID before attaching restart-loop evidence. Preserve handling for events without a UID where appropriate, and test same-name pod recreation.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (2)
4. Failing-to-start loops still churn issue ids ✓ Resolved
Description
restartLoopMayReplace allows only phases, CrashLoopBackOff, Error, Completed and the
probe/high-restart reasons, yet containerRestartLoop accepts any non-OOM termination, including
StartError and ContainerCannotRun. On ticks where the kubelet shows such a Terminated state, the
detector sets looping=false and keeps the raw reason, which classifies as unknown; it flips back
to a critical crashloop on the next backoff tick, reproducing the open/resolve churn for bad command
or entrypoint loops.
Code

internal/k8s/restart_loop.go[R170-174]

+	switch reason {
+	case "Running", "Pending", "Unknown", "", "PodInitializing", "ContainerCreating",
+		crashLoopReason, "Error", "Completed",
+		highRestartReason, livenessProbeFailedReason, readinessProbeFailedReason:
+		return true
Evidence
podProblemReasonRaw returns the Terminated reason of a container that exited non-zero.
isPhaseOrCrashReason does not normalize StartError. The detector's else branch then clears looping,
and Classify has no Pod case for StartError, so the row falls into CategoryUnknown.

internal/k8s/restart_loop.go[109-136]
internal/k8s/detect.go[498-503]
pkg/health/reasons.go[348-356]
internal/issues/category.go[215-246]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
restartLoopMayReplace rejects terminated reasons such as StartError and ContainerCannotRun even though containerRestartLoop counts them as part of a loop. On those ticks the pod row becomes a different category and issue id.
## Fix Focus Areas
- internal/k8s/restart_loop.go[169-177]
- internal/k8s/detect.go[498-503]
## Recommended Fix
While looping, also allow replacement when the reason equals the Terminated reason of a container status (any reason except OOMKilled). Alternatively, add StartError, ContainerCannotRun and ContainerStatusUnknown to the allow-list. Add a test tick with Terminated{Reason:"StartError", ExitCode:128}.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


5. Hosted restart evidence disappears 🔗 Cross-repo conflict ≡ Correctness
Description
The new Issue.RestartLoop field is lost when radar-hub decodes Radar’s /api/issues response
using its pinned radar/pkg v1.15.0 issue type, which predates the field. Hub then forwards that
decoded issue to radar-hub-web and persists it for alerts, so neither hosted surface receives the
restart-loop or probe evidence.
Code

pkg/issuesapi/types.go[447]

+	RestartLoop       *RestartLoop       `json:"restart_loop,omitempty"`
Evidence
The PR adds a serialized field to the shared issue model. Hub pins an earlier version of that model,
decodes agent responses into it, and forwards or persists the decoded value; Hub Web passes fleet
issues to the shared issue row but its fleet type has no restart-loop field.

radar -> radar-hub
radar -> radar-hub-web
pkg/issuesapi/types.go[364-383]
pkg/issuesapi/types.go[444-447]
External repo: skyhook-dev/radar-hub, go.mod [5-8]
External repo: skyhook-dev/radar-hub, internal/server/fleet_handlers.go [2519-2548]
External repo: skyhook-dev/radar-hub, internal/server/alerts_worker.go [609-619]
External repo: skyhook-dev/radar-hub, internal/db/alerts.go [669-678]
External repo: skyhook-dev/radar-hub-web, src/api/fleet.ts [579-584]
External repo: skyhook-dev/radar-hub-web, src/pages/fleet/issues/IssueRow.tsx [409-418]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Radar Hub’s pinned shared issue type discards the new `restart_loop` field while decoding agent responses, removing evidence from hosted fleet issues and persisted alert signals.
## Fix Focus Areas
- pkg/issuesapi/types.go[444-447]
- /cross_repos/radar-hub/go.mod[5-8]
- /cross_repos/radar-hub/internal/server/fleet_handlers.go[2519-2548]
- /cross_repos/radar-hub/internal/server/alerts_worker.go[609-619]
- /cross_repos/radar-hub-web/src/api/fleet.ts[579-584]
## Recommended Fix
Publish the updated Radar shared package and update radar-hub’s dependency before relying on the new evidence in hosted responses or alerts. Add the optional field to radar-hub-web’s fleet issue type and verify that an agent issue retains it through Hub’s fleet and alert paths.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Informational

6. A sibling container's probe failure is hidden ✓ Resolved
Description
restartLoopMayReplace lets the loop replace ReadinessProbeFailed/LivenessProbeFailed even when
those pod-wide reasons come from a different container than loop.container. A sibling's ongoing
readiness failure is therefore overwritten by the looping container's crashloop row and stays
invisible until 30 minutes after the last crash.
Code

internal/k8s/restart_loop.go[173]

+		highRestartReason, livenessProbeFailedReason, readinessProbeFailedReason:
Evidence
The probeFailures lookup is keyed per pod, not per container. The PodProblemReason readiness check
considers any container. The loop override replaces the reason unconditionally.

internal/k8s/detect.go[488-503]
pkg/health/reasons.go[286-303]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Pod-level probe-failure reasons may belong to a container other than the looping one, yet the loop overwrites them.
## Fix Focus Areas
- internal/k8s/restart_loop.go[169-177]
- internal/k8s/detect.go[488-503]
## Recommended Fix
Replace a probe reason only when containerProbeFailures has a matching failure for loop.container, or when the pod has a single container. Otherwise keep the probe row.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. One-off restarts open a critical issue ✓ Resolved
Description
containerRestartLoop uses the cumulative RestartCount>=3 plus a single termination within 30
minutes, and the detector then pins severity to critical even when the pod is Ready. A long-lived
container with old restarts that restarts once, for example after a node or kubelet bounce, now
raises a critical crashloop for about 30 minutes while serving normally.
Code

internal/k8s/detect.go[526]

+				severity = "critical"
Evidence
The updated test changes the high-count-serving severity from high to critical, which confirms that
a ready pod with old restarts is now reported as critical.

internal/k8s/restart_loop.go[109-128]
internal/k8s/detect_crashloop_severity_test.go[155-157]


Grey Divider

Tip of the day
💡 Did you know, you can describe a rule in plain language on the Rules page and Qodo drafts it for you

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread internal/issues/restart_loop_integration_test.go Outdated
Comment thread internal/k8s/restart_loop.go Outdated
Comment thread internal/k8s/detect.go
Comment thread internal/k8s/restart_loop.go
Comment thread internal/k8s/restart_loop.go
Comment thread internal/k8s/detect.go Outdated
Comment thread pkg/issuesapi/types.go

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/k8s/restart_loop.go
Skip pods being deleted (a rollout SIGTERM is not a crash), require the
terminated run itself to be shorter than the loop window so a long-lived
container's one restart after a node bounce is not a loop, treat
StartError/ContainerCannotRun as faces of the loop, and ignore probe
events from a same-named predecessor pod.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread internal/k8s/restart_loop.go
StartError / ContainerCannotRun terminations are stamped with the Unix
epoch, which the run-length guard read as a decades-long run and so
rejected the loop.
Live testing showed a container that cannot start alternating between
crashloop and container_waiting: between attempts it waits in
RunContainerError. Once the loop rule holds the container has been started
and terminated repeatedly, so RunContainerError is another face of the
loop. The loop row now also carries the runtime's termination message as
raw_message and names start failures in its cause.
- Name the container that terminated most recently, not the first that
  qualifies, so a container that recovered earlier in the window is not
  blamed for a sibling's loop.
- Honor per-container restartPolicy (Kubernetes 1.34+): a container's own
  policy overrides the pod's when deciding whether it restarts by design.
- Record startup probe failures as loop evidence and explain exit 143.
- Keep critical severity when an invalid-probe-target row wins during a
  loop, so it does not flap with the crash cycle.
- Read Series.LastObservedTime for events recorded through events.k8s.io.
- Describe probe evidence as failures seen, not as failing now, and widen
  the crashloop catalog definition to repeated restarts of any exit code.
…it and OOM loops

- Severity: a loop is critical only when it is fast (last run under 10
  minutes, the kubelet's backoff-reset line) and covers at least half of
  its workload's live pods; a slow loop or one bad replica among many is a
  warning. Both inputs hold steady through a loop.
- A loop row folds its workload's workload_degraded row at any severity:
  the loop explains the unavailability, and the "N/M available" row is
  critical whenever a replica is down and otherwise comes and goes with
  the crash cycle. Invalid-probe-target rows during a loop carry the loop
  evidence so they fold the same way.
- Recovery: a loop ends once the container is Ready and has run longer
  than max(10m, 2x its last run), instead of holding 30 minutes after the
  last crash.
- Init:CrashLoopBackOff: ordinary init containers that keep failing while
  the pod is Pending are loops; their long attempts are not also reported
  as a stalled init container.
- OOM loops get the same continuity and classify as oom_killed; another
  container's active OOM still owns the row; the limit-discrepancy
  diagnosis follows the looping container between kills, and the plain
  OOM diagnosis names node pressure when there is no limit.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit d483e0f. Configure here.

Comment thread internal/k8s/restart_loop.go
…enacted memory limits

A readiness blip after a loop recovered reopened it until 30 minutes after
its last crash. The early clear now rests on run length alone; a container
that stops crashing but stays unready is reported by its readiness row.
An init container's or native sidecar's unrecovered last-state OOM now
owns the row over another container's loop, as PodHasActiveOOMKilled
judges it, and an enacted memory limit in status counts when choosing the
OOM wording.
@nadaverell nadaverell changed the title fix(issues): keep one restart-loop issue across a crashloop cycle fix(issues): one stable issue per restart loop Oct 4, 2026
…verity reason

The loop folds its workload's "N/M available" row, which flips with every
crash. Carry its stable equivalent instead: how many of the workload's
live pods are looping (looping_pods / workload_pods), when the last run
started (so its length shows whether the container dies on startup or
serves between restarts), and why the loop is critical or warning. The
restart_cause fact and the issue's meta line show them.

Also keep EventTime and Series when the cache strips Events, so events
recorded through events.k8s.io keep their times.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant