Describe the bug
When the NVIDIA kernel module is installed but cannot bind the GPU until the next reboot, nvidia-cdi-refresh.service restarts indefinitely. On a first driver install, for example, nouveau still owns the device until the reboot. Every attempt runs nvidia-smi -L, which tries to load nvidia.ko and fails, so the host loads and fails the module about every 2.3 s.
The unit intends to cap this: "Limit the number of successive restarts to 5 in 10 seconds." The cap only holds when a failed attempt takes under about 1 s, because systemd's start limit is a fixed window opened by the first start. In systemd's ratelimit_below() (src/basic/ratelimit.c, unchanged through v259), a start is refused only when the 6th start lands within StartLimitIntervalSec of the window's first start. With RestartSec=1s, a cycle longer than 2.0 s means the 6th start opens a new window instead.
Here a failed attempt took about 1.3 s (nvidia-smi -L probing the device), so the cycle was about 2.3 s. The service ran 2,391 restarts over roughly 94 minutes. The limit ("Start request repeated too quickly") tripped only when jitter happened to fit six starts into one 10 s window. Until then each cycle added:
NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.
NVRM: This can occur when another driver was loaded and
NVRM: obtained ownership of the NVIDIA device(s).
NVRM: Try unloading the conflicting kernel module (and/or
NVRM: reconfigure your kernel without the conflicting
NVRM: driver(s)), then try loading the NVIDIA kernel module
NVRM: again.
NVRM: No NVIDIA devices probed.
nvidia.ko also appears in /proc/modules for the duration of each failed load, so anything that reads loaded modules to detect the driver sees it flap.
This looks like the case in #1624, which was fixed by #1638 and then deliberately reverted by #1825/#1836. #1825 noted that #1624 "should" be re-opened, but that didn't happen. The part not covered there is the start-limit arithmetic above: the retry loop is unbounded whenever an attempt runs longer than about 1 s, which is exactly when the driver is present but unusable.
To Reproduce
- On a host whose GPU is bound to
nouveau, install the distro's NVIDIA driver with precompiled modules (here Ubuntu's nvidia-headless-no-dkms-580-server + linux-modules-nvidia-580-server-generic), and do not reboot.
- Install
nvidia-container-toolkit 1.20.1-1 from the stable apt repo. Its postinst enables and starts nvidia-cdi-refresh.path and .service, and the new modules.dep also satisfies the ExecCondition.
- Run
journalctl -u nvidia-cdi-refresh.service -f and dmesg -w. The service fails in nvidia-smi -L and restarts about every 2.3 s, with the kernel messages above each time. systemctl show -p NRestarts nvidia-cdi-refresh.service keeps climbing past 5.
The same shape should apply to any state where nvidia.ko is in modules.dep but the driver can't come up this boot, which is the #1624 situation.
Expected behavior
The unit gives up after the intended ~5 attempts, whatever each attempt's duration, or backs off, and stays failed until the next trigger (reboot, modules.dep change, or a manual start). Two ways this could be done, not a prescription:
- a
StartLimitIntervalSec sized to StartLimitBurst × (RestartSec + worst-case attempt time);
RestartSteps=/RestartMaxDelaySec= (systemd ≥ 254) so retries back off.
Workaround
Where nothing consumes the generated spec (/var/run/cdi/nvidia.yaml), systemctl mask nvidia-cdi-refresh.path nvidia-cdi-refresh.service before installing the toolkit. The postinst leaves an administrator's mask in place.
Environment (please provide the following information):
nvidia-container-toolkit version: 1.20.1-1 (deb, https://nvidia.github.io/libnvidia-container/stable/deb/amd64); the unit, path, udev and packaging files are identical on main as of 84e2c2c
- NVIDIA Driver Version: 580.178.04 (Ubuntu
nvidia-headless-no-dkms-580-server 580.178.04-0ubuntu0.26.04.1, precompiled modules)
- Host OS: Ubuntu 26.04.1 LTS
- Kernel Version: 7.0.0-34-generic
- systemd: 259 (259.5-0ubuntu3.4)
- Container Runtime Version: not involved (the loop is the systemd unit alone)
- CPU Architecture:
x86_64
- GPU Model(s): GeForce GTX 950
Describe the bug
When the NVIDIA kernel module is installed but cannot bind the GPU until the next reboot,
nvidia-cdi-refresh.servicerestarts indefinitely. On a first driver install, for example,nouveaustill owns the device until the reboot. Every attempt runsnvidia-smi -L, which tries to loadnvidia.koand fails, so the host loads and fails the module about every 2.3 s.The unit intends to cap this: "Limit the number of successive restarts to 5 in 10 seconds." The cap only holds when a failed attempt takes under about 1 s, because systemd's start limit is a fixed window opened by the first start. In systemd's
ratelimit_below()(src/basic/ratelimit.c, unchanged through v259), a start is refused only when the 6th start lands withinStartLimitIntervalSecof the window's first start. WithRestartSec=1s, a cycle longer than 2.0 s means the 6th start opens a new window instead.Here a failed attempt took about 1.3 s (
nvidia-smi -Lprobing the device), so the cycle was about 2.3 s. The service ran 2,391 restarts over roughly 94 minutes. The limit ("Start request repeated too quickly") tripped only when jitter happened to fit six starts into one 10 s window. Until then each cycle added:nvidia.koalso appears in/proc/modulesfor the duration of each failed load, so anything that reads loaded modules to detect the driver sees it flap.This looks like the case in #1624, which was fixed by #1638 and then deliberately reverted by #1825/#1836. #1825 noted that #1624 "should" be re-opened, but that didn't happen. The part not covered there is the start-limit arithmetic above: the retry loop is unbounded whenever an attempt runs longer than about 1 s, which is exactly when the driver is present but unusable.
To Reproduce
nouveau, install the distro's NVIDIA driver with precompiled modules (here Ubuntu'snvidia-headless-no-dkms-580-server+linux-modules-nvidia-580-server-generic), and do not reboot.nvidia-container-toolkit1.20.1-1 from the stable apt repo. Its postinst enables and startsnvidia-cdi-refresh.pathand.service, and the newmodules.depalso satisfies theExecCondition.journalctl -u nvidia-cdi-refresh.service -fanddmesg -w. The service fails innvidia-smi -Land restarts about every 2.3 s, with the kernel messages above each time.systemctl show -p NRestarts nvidia-cdi-refresh.servicekeeps climbing past 5.The same shape should apply to any state where
nvidia.kois inmodules.depbut the driver can't come up this boot, which is the #1624 situation.Expected behavior
The unit gives up after the intended ~5 attempts, whatever each attempt's duration, or backs off, and stays failed until the next trigger (reboot,
modules.depchange, or a manual start). Two ways this could be done, not a prescription:StartLimitIntervalSecsized toStartLimitBurst × (RestartSec + worst-case attempt time);RestartSteps=/RestartMaxDelaySec=(systemd ≥ 254) so retries back off.Workaround
Where nothing consumes the generated spec (
/var/run/cdi/nvidia.yaml),systemctl mask nvidia-cdi-refresh.path nvidia-cdi-refresh.servicebefore installing the toolkit. The postinst leaves an administrator's mask in place.Environment (please provide the following information):
nvidia-container-toolkitversion: 1.20.1-1 (deb,https://nvidia.github.io/libnvidia-container/stable/deb/amd64); the unit, path, udev and packaging files are identical onmainas of 84e2c2cnvidia-headless-no-dkms-580-server580.178.04-0ubuntu0.26.04.1, precompiled modules)x86_64