fix(cloud): tell Cloud whether an OOM kill actually stopped the container - #100
Merged
Merged
Conversation
…iner Docker emits `oom` for any process the kernel OOM-kills inside a container's cgroup, including a child (a worker, a transcoder) while the container keeps running. The agent forwarded that as a plain oom and Cloud reported the service as down/OOM-killed although it never stopped. On oom the agent now inspects the container and sends container_running. Since that inspect can race the exit of a main process being killed, a die arriving within 10s of an oom for the same container carries oom_killed=true, so the exit itself is the OOM verdict. Log excerpts are no longer captured for an oom that left the container running.
# Conflicts: # docs/06-cloud.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Docker emits an
oomevent whenever the kernel OOM-kills any process in a container's cgroup, including a child process while the container keeps running. It also setsState.OOMKilled=trueon a container that's still up. The agent forwarded every such event as a plainoom, and Cloud treated that as the service being down and "OOM-killed".Real case: every night Plex's maintenance jobs spawn a
Plex Transcoderthat grows to about 5.5 GB. It hits the Proxmox LXC's memory limit and gets killed. Plex itself stays up (0 restarts, healthy, still serving requests), but Cloud raised a critical "plex OOM-killed" alert each time.Change
oomevents now carrycontainer_running. The agent inspects the container when the event arrives.truemeans only a process inside the container died. The value is omitted if the inspect fails.dieevents now carryoom_killed. It is set when anoomfor the same container came within 10s before the exit. That inspect can race the exit of a main process that's being killed and still read "running", so thedieis what carries the OOM verdict reliably.oomthat left the container running. Cloud opens no incident for it, so the excerpt would have nothing to attach to.omitempty, so there's no protocol version bump. A Cloud that doesn't know them behaves as before.docs/06-cloud.mdand the README example line.Cloud-side handling of the new fields ships separately in docktail-cloud.
proto/needs to stay in sync (make proto-diff).🤖 Generated with Claude Code