Skip to content

fix(cloud): tell Cloud whether an OOM kill actually stopped the container - #100

Merged
marvinvr merged 2 commits into
mainfrom
fix/oom-process-kill-not-down
Sep 30, 2026
Merged

marvinvr merged 2 commits into
mainfrom
fix/oom-process-kill-not-down

Conversation

@marvinvr

Copy link
Copy Markdown
Owner

Problem

Docker emits an oom event whenever the kernel OOM-kills any process in a container's cgroup, including a child process while the container keeps running. It also sets State.OOMKilled=true on a container that's still up. The agent forwarded every such event as a plain oom, and Cloud treated that as the service being down and "OOM-killed".

Real case: every night Plex's maintenance jobs spawn a Plex Transcoder that grows to about 5.5 GB. It hits the Proxmox LXC's memory limit and gets killed. Plex itself stays up (0 restarts, healthy, still serving requests), but Cloud raised a critical "plex OOM-killed" alert each time.

Change

  • oom events now carry container_running. The agent inspects the container when the event arrives. true means only a process inside the container died. The value is omitted if the inspect fails.
  • die events now carry oom_killed. It is set when an oom for the same container came within 10s before the exit. That inspect can race the exit of a main process that's being killed and still read "running", so the die is what carries the OOM verdict reliably.
  • No log capture for an oom that left the container running. Cloud opens no incident for it, so the excerpt would have nothing to attach to.
  • Protocol: both fields are additive and omitempty, so there's no protocol version bump. A Cloud that doesn't know them behaves as before.
  • Docs: updated docs/06-cloud.md and the README example line.
  • Tests: added unit tests for the oom→die correlation.

Cloud-side handling of the new fields ships separately in docktail-cloud. proto/ needs to stay in sync (make proto-diff).

🤖 Generated with Claude Code

…iner

Docker emits `oom` for any process the kernel OOM-kills inside a
container's cgroup, including a child (a worker, a transcoder) while the
container keeps running. The agent forwarded that as a plain oom and Cloud
reported the service as down/OOM-killed although it never stopped.

On oom the agent now inspects the container and sends container_running.
Since that inspect can race the exit of a main process being killed, a die
arriving within 10s of an oom for the same container carries
oom_killed=true, so the exit itself is the OOM verdict. Log excerpts are no
longer captured for an oom that left the container running.
@marvinvr
marvinvr merged commit 7507252 into main Sep 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant