fix(issues): one stable issue per restart loop - #1967
nadaverell wants to merge 8 commits into
Conversation
A container stuck restarting cycles through states that each read as a different problem on a single poll: CrashLoopBackOff, Running-not-ready, readiness and liveness Unhealthy events, Terminated exit 0, and Running+Ready between crashes. The pod row changed category (and so issue id) every cycle and vanished on the healthy-looking ticks, and the owning Deployment's workload_degraded row came and went with it. In production one looping gateway produced thousands of open/resolve generations. The detector now classifies an active restart loop from restart history: 3+ restarts and a termination (current or last, any exit code, not OOM) within the past 30 minutes, for containers the kubelet restarts by design (restartPolicy Always containers and native sidecars). While it holds, the pod keeps one critical crashloop row through every tick, folding the per-tick faces of the loop into it. Probe failures, last exit and the container are attached as RestartLoop evidence and a restart_cause fact; the diagnosis depends only on container and exit code so it stays stable across polls and agrees across replicas. Pod health levels are unchanged. More specific problems keep their own rows: invalid probe targets, image pulls, create errors, active OOM (now including sidecars), init stalls. Also: a started native sidecar is no longer reported as a stalled init container, and ReplicaFailure is folded only into scheduling-source admission rejections, never into a runtime symptom on an existing pod.
PR Summary by QodoKeep one critical issue throughout a container restart loop
AI Description
Diagram
High-Level Assessment
Files changed (15)
|
Code Review by Qodo
1.
|
Skip pods being deleted (a rollout SIGTERM is not a crash), require the terminated run itself to be shorter than the loop window so a long-lived container's one restart after a node bounce is not a loop, treat StartError/ContainerCannotRun as faces of the loop, and ignore probe events from a same-named predecessor pod.
StartError / ContainerCannotRun terminations are stamped with the Unix epoch, which the run-length guard read as a decades-long run and so rejected the loop.
Live testing showed a container that cannot start alternating between crashloop and container_waiting: between attempts it waits in RunContainerError. Once the loop rule holds the container has been started and terminated repeatedly, so RunContainerError is another face of the loop. The loop row now also carries the runtime's termination message as raw_message and names start failures in its cause.
- Name the container that terminated most recently, not the first that qualifies, so a container that recovered earlier in the window is not blamed for a sibling's loop. - Honor per-container restartPolicy (Kubernetes 1.34+): a container's own policy overrides the pod's when deciding whether it restarts by design. - Record startup probe failures as loop evidence and explain exit 143. - Keep critical severity when an invalid-probe-target row wins during a loop, so it does not flap with the crash cycle. - Read Series.LastObservedTime for events recorded through events.k8s.io. - Describe probe evidence as failures seen, not as failing now, and widen the crashloop catalog definition to repeated restarts of any exit code.
…it and OOM loops - Severity: a loop is critical only when it is fast (last run under 10 minutes, the kubelet's backoff-reset line) and covers at least half of its workload's live pods; a slow loop or one bad replica among many is a warning. Both inputs hold steady through a loop. - A loop row folds its workload's workload_degraded row at any severity: the loop explains the unavailability, and the "N/M available" row is critical whenever a replica is down and otherwise comes and goes with the crash cycle. Invalid-probe-target rows during a loop carry the loop evidence so they fold the same way. - Recovery: a loop ends once the container is Ready and has run longer than max(10m, 2x its last run), instead of holding 30 minutes after the last crash. - Init:CrashLoopBackOff: ordinary init containers that keep failing while the pod is Pending are loops; their long attempts are not also reported as a stalled init container. - OOM loops get the same continuity and classify as oom_killed; another container's active OOM still owns the row; the limit-discrepancy diagnosis follows the looping container between kills, and the plain OOM diagnosis names node pressure when there is no limit.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit d483e0f. Configure here.
…enacted memory limits A readiness blip after a loop recovered reopened it until 30 minutes after its last crash. The early clear now rests on run length alone; a container that stops crashing but stays unready is reported by its readiness row. An init container's or native sidecar's unrecovered last-state OOM now owns the row over another container's loop, as PodHasActiveOOMKilled judges it, and an enacted memory limit in status counts when choosing the OOM wording.
…verity reason The loop folds its workload's "N/M available" row, which flips with every crash. Carry its stable equivalent instead: how many of the workload's live pods are looping (looping_pods / workload_pods), when the last run started (so its length shows whether the container dies on startup or serves between restarts), and why the loop is critical or warning. The restart_cause fact and the issue's meta line show them. Also keep EventTime and Series when the cache strips Events, so events recorded through events.k8s.io keep their times.

Problem
A container stuck restarting looks like a different problem on almost every poll:
CrashLoopBackOff→crashloopreadiness_failed/high_restart/ nothingUnhealthyevent →liveness_probe_failedCompleted/0(graceful liveness kill) → breaks the stable-crashloop checkEach flip is a new issue id, and the Deployment's
workload_degradedrow comes and goes with it. Downstream, every flip is a new alert and a new AI investigation. One production gateway (restartCount 4592, last stateCompleted/0) produced thousands of open/resolve generations. On our own clusters, 15 workloads loop like this right now.Change
One issue per restart loop.
crashloop, oroom_killedfor an OOM loop.workload_degradedrow folds into the loop at any severity, because the loop explains the unavailability.Which containers count.
Init:CrashLoopBackOff). Their long attempts are no longer also reported as a stalled init.Severity follows impact, and holds steady for the loop.
Clears on recovery. The loop ends once the container has run longer than max(10 min, 2× its last run). Fast loops clear 10 minutes after they stop crashing, instead of holding 30 minutes after the last crash.
More specific problems keep their own row: invalid probe targets (pinned to the loop's severity, carrying its evidence), image pulls, create errors, and another container's active OOM.
Also:
ReplicaFailure(can't create pods) folds only into scheduling-source admission rejections.events.k8s.iokeepEventTime/Seriesin the cache and are dated by their latest occurrence.Live verification
mainand this branch ran side by side, read-only, against 4 real clusters (GKE nonprod / management / prod, EKS skh-nonprod) for 30 minutes, polling/api/issuesevery 20 s:rollout_stalledrow stays visible next to the warning loop, as it does on main;rollout_stalledkeeps its severity gate.RunContainerError/StartError)Init:CrashLoopBackOff)oom_killed)oom_killed, with loop evidence)Recovery: a pod that crashed 4 times and then turned Ready at 12:08:11 UTC was cleared by the branch at about 12:18. In the first minutes, before a container's 3rd restart, both builds still show the old per-poll behaviour.
Tradeoffs and known limits
oom_killed/crashloop.service_no_endpointsstill flaps with its looping backend; the existing Service→backend issue link remains.Tests
DetectProblems→Compose. It asserts one id, one cause and the evidence at every tick, and that the loop clears after 11 Ready minutes.go test ./...(both modules) ✓ ·make tsc✓ · CI ✓