Skip to content

Aim the disk scenario below the kubelet's own lines - #65

Merged
Bryancruzcb merged 1 commit into
mainfrom
ops/disk-scenario-fix
Sep 22, 2026
Merged

Bryancruzcb merged 1 commit into
mainfrom
ops/disk-scenario-fix

Conversation

@Bryancruzcb

Copy link
Copy Markdown
Owner

Game day scenario 2 caused a real outage on its first run, and this is the
fix plus what it taught.

What happened. PLAN.md asked for the disk to be filled to 95%. On this
node that is exactly the kubelet's eviction line, read from its live config:

used what the node does
80% DiskAlmostFull goes pending, and fires after 10 minutes
85% the kubelet deletes unused images until the disk is back at 80%
95% free space is under 5%: the kubelet evicts pods and taints the node NoSchedule

Within 20 seconds of the fill it evicted both API pods and Grafana. Image GC
then pulled the disk back to 79%, so DiskAlmostFull was pending for 30
seconds and never fired. Both APIs were down for about six minutes, and
ApiNotReady fired in both namespaces (#63, #64) — the outage alert catching
a real outage, which the game day caused. I removed the filler by hand 85
seconds in; the node kept its taint five more minutes
(evictionPressureTransitionPeriod), then everything rescheduled.

The fix. The scenario now aims at 83%, between the alert and image GC,
and reads both kubelet lines from the live config before allocating
anything. It refuses a target at or past either one, and refuses outright if
the config can't be read or the eviction threshold isn't a percentage. The
parsing was checked against the real config shape, an absolute threshold,
and an empty config.

The lesson, now in the disk runbook: the alert gives five points of
warning before the kubelet starts quietly deleting images, which can clear
the alert while whatever is growing keeps growing, and fifteen before it
evicts pods. ApiNotReady in both namespaces at once means check the disk
first.

Also in here: the game-day results so far, seven rows — three runs of
scenario 1 (prod back in 20–21 s), three of scenario 4 (the gate blocked a
top k of 1 every time, 70% against the recorded 90%), and this run, recorded
as evicted rather than left out.

Merging this doesn't deploy: deploy/gameday/ and deploy/runbooks/ aren't
in deploy.yml's path filter. Scenario 2 reruns at 83% after it lands.

Paradigm: procedural shell, the shell row of PARADIGMS.md. Rules applied:
API-02 (the new refusals cover an unreadable config and a non-percentage
threshold, not just a bad target), CON-* (the kubelet's lines are
preconditions checked before the first byte is allocated, with the existing
EXIT trap unchanged), DES-05 (two thresholds read inline, no new
library function).

The first run of game-day scenario 2 on 2026-09-22 aimed at PLAN.md's 95
percent, which on this node is exactly the kubelet's eviction line: k3s's
kubelet evicts pods once free space is under 5 percent, and deletes unused
images from 85 percent back down to 80 (read from the node's kubelet
config). Within 20 seconds it evicted both API pods and Grafana and tainted
the node NoSchedule; image GC pulled the disk back to 79 percent, so
DiskAlmostFull was pending for 30 seconds and never fired. The APIs were
down for about six minutes, and ApiNotReady fired in both namespaces (#63,
#64). The filler was removed by hand 85 seconds in.

The scenario now aims at 83 percent, between the alert's 80 and image GC's
85, and reads both kubelet lines live before allocating anything. It
refuses a target at or past either line, and refuses when the config cannot
be read or the eviction threshold is not a percentage.

The runbook gets the table, because it is the operational lesson: the alert
is five points of warning before the kubelet starts quietly deleting images
(which can clear the alert while the growth continues) and fifteen before it
evicts pods. results.tsv gets the six rows from scenarios 1 and 4 and this
run, recorded as evicted.
@Bryancruzcb
Bryancruzcb merged commit 4c3da63 into main Sep 22, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant