Skip to content

[Bug]: nvidia-cdi-refresh.service restarts indefinitely while the driver cannot load — the 5-in-10s start limit never trips once an attempt takes over ~1 s #2110

Description

@kevinpark1217

Describe the bug

When the NVIDIA kernel module is installed but cannot bind the GPU until the next reboot, nvidia-cdi-refresh.service restarts indefinitely. On a first driver install, for example, nouveau still owns the device until the reboot. Every attempt runs nvidia-smi -L, which tries to load nvidia.ko and fails, so the host loads and fails the module about every 2.3 s.

The unit intends to cap this: "Limit the number of successive restarts to 5 in 10 seconds." The cap only holds when a failed attempt takes under about 1 s, because systemd's start limit is a fixed window opened by the first start. In systemd's ratelimit_below() (src/basic/ratelimit.c, unchanged through v259), a start is refused only when the 6th start lands within StartLimitIntervalSec of the window's first start. With RestartSec=1s, a cycle longer than 2.0 s means the 6th start opens a new window instead.

Here a failed attempt took about 1.3 s (nvidia-smi -L probing the device), so the cycle was about 2.3 s. The service ran 2,391 restarts over roughly 94 minutes. The limit ("Start request repeated too quickly") tripped only when jitter happened to fit six starts into one 10 s window. Until then each cycle added:

NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.
NVRM: This can occur when another driver was loaded and
NVRM: obtained ownership of the NVIDIA device(s).
NVRM: Try unloading the conflicting kernel module (and/or
NVRM: reconfigure your kernel without the conflicting
NVRM: driver(s)), then try loading the NVIDIA kernel module
NVRM: again.
NVRM: No NVIDIA devices probed.

nvidia.ko also appears in /proc/modules for the duration of each failed load, so anything that reads loaded modules to detect the driver sees it flap.

This looks like the case in #1624, which was fixed by #1638 and then deliberately reverted by #1825/#1836. #1825 noted that #1624 "should" be re-opened, but that didn't happen. The part not covered there is the start-limit arithmetic above: the retry loop is unbounded whenever an attempt runs longer than about 1 s, which is exactly when the driver is present but unusable.

To Reproduce

  1. On a host whose GPU is bound to nouveau, install the distro's NVIDIA driver with precompiled modules (here Ubuntu's nvidia-headless-no-dkms-580-server + linux-modules-nvidia-580-server-generic), and do not reboot.
  2. Install nvidia-container-toolkit 1.20.1-1 from the stable apt repo. Its postinst enables and starts nvidia-cdi-refresh.path and .service, and the new modules.dep also satisfies the ExecCondition.
  3. Run journalctl -u nvidia-cdi-refresh.service -f and dmesg -w. The service fails in nvidia-smi -L and restarts about every 2.3 s, with the kernel messages above each time. systemctl show -p NRestarts nvidia-cdi-refresh.service keeps climbing past 5.

The same shape should apply to any state where nvidia.ko is in modules.dep but the driver can't come up this boot, which is the #1624 situation.

Expected behavior

The unit gives up after the intended ~5 attempts, whatever each attempt's duration, or backs off, and stays failed until the next trigger (reboot, modules.dep change, or a manual start). Two ways this could be done, not a prescription:

  • a StartLimitIntervalSec sized to StartLimitBurst × (RestartSec + worst-case attempt time);
  • RestartSteps=/RestartMaxDelaySec= (systemd ≥ 254) so retries back off.

Workaround

Where nothing consumes the generated spec (/var/run/cdi/nvidia.yaml), systemctl mask nvidia-cdi-refresh.path nvidia-cdi-refresh.service before installing the toolkit. The postinst leaves an administrator's mask in place.

Environment (please provide the following information):

  • nvidia-container-toolkit version: 1.20.1-1 (deb, https://nvidia.github.io/libnvidia-container/stable/deb/amd64); the unit, path, udev and packaging files are identical on main as of 84e2c2c
  • NVIDIA Driver Version: 580.178.04 (Ubuntu nvidia-headless-no-dkms-580-server 580.178.04-0ubuntu0.26.04.1, precompiled modules)
  • Host OS: Ubuntu 26.04.1 LTS
  • Kernel Version: 7.0.0-34-generic
  • systemd: 259 (259.5-0ubuntu3.4)
  • Container Runtime Version: not involved (the loop is the systemd unit alone)
  • CPU Architecture: x86_64
  • GPU Model(s): GeForce GTX 950

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions