Repository navigation
Aim the disk scenario below the kubelet's own lines - #65
Merged
Merged
Conversation
The first run of game-day scenario 2 on 2026-09-22 aimed at PLAN.md's 95 percent, which on this node is exactly the kubelet's eviction line: k3s's kubelet evicts pods once free space is under 5 percent, and deletes unused images from 85 percent back down to 80 (read from the node's kubelet config). Within 20 seconds it evicted both API pods and Grafana and tainted the node NoSchedule; image GC pulled the disk back to 79 percent, so DiskAlmostFull was pending for 30 seconds and never fired. The APIs were down for about six minutes, and ApiNotReady fired in both namespaces (#63, #64). The filler was removed by hand 85 seconds in. The scenario now aims at 83 percent, between the alert's 80 and image GC's 85, and reads both kubelet lines live before allocating anything. It refuses a target at or past either line, and refuses when the config cannot be read or the eviction threshold is not a percentage. The runbook gets the table, because it is the operational lesson: the alert is five points of warning before the kubelet starts quietly deleting images (which can clear the alert while the growth continues) and fifteen before it evicts pods. results.tsv gets the six rows from scenarios 1 and 4 and this run, recorded as evicted.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Game day scenario 2 caused a real outage on its first run, and this is the
fix plus what it taught.
What happened. PLAN.md asked for the disk to be filled to 95%. On this
node that is exactly the kubelet's eviction line, read from its live config:
DiskAlmostFullgoes pending, and fires after 10 minutesNoScheduleWithin 20 seconds of the fill it evicted both API pods and Grafana. Image GC
then pulled the disk back to 79%, so
DiskAlmostFullwas pending for 30seconds and never fired. Both APIs were down for about six minutes, and
ApiNotReadyfired in both namespaces (#63, #64) — the outage alert catchinga real outage, which the game day caused. I removed the filler by hand 85
seconds in; the node kept its taint five more minutes
(
evictionPressureTransitionPeriod), then everything rescheduled.The fix. The scenario now aims at 83%, between the alert and image GC,
and reads both kubelet lines from the live config before allocating
anything. It refuses a target at or past either one, and refuses outright if
the config can't be read or the eviction threshold isn't a percentage. The
parsing was checked against the real config shape, an absolute threshold,
and an empty config.
The lesson, now in the disk runbook: the alert gives five points of
warning before the kubelet starts quietly deleting images, which can clear
the alert while whatever is growing keeps growing, and fifteen before it
evicts pods.
ApiNotReadyin both namespaces at once means check the diskfirst.
Also in here: the game-day results so far, seven rows — three runs of
scenario 1 (prod back in 20–21 s), three of scenario 4 (the gate blocked a
top k of 1 every time, 70% against the recorded 90%), and this run, recorded
as
evictedrather than left out.Merging this doesn't deploy:
deploy/gameday/anddeploy/runbooks/aren'tin
deploy.yml's path filter. Scenario 2 reruns at 83% after it lands.Paradigm: procedural shell, the shell row of
PARADIGMS.md. Rules applied:API-02(the new refusals cover an unreadable config and a non-percentagethreshold, not just a bad target),
CON-*(the kubelet's lines arepreconditions checked before the first byte is allocated, with the existing
EXITtrap unchanged),DES-05(two thresholds read inline, no newlibrary function).