Skip to content

install.sh: 从 v1.0.0 原地升级时漏建 monitor-agent 用户,服务 217/USER 起不来(id -u 被 DynamicUser 的瞬时用户骗过) #44

Description

@Cristina269

现象

两台已经在跑 agent v1.0.0(当时的 unit 是 DynamicUser=yes)的 Debian 机器,按文档重跑 install.sh 升级到 v1.1.0 之后,服务进入重启循环:

monitor-agent.service: Failed to determine user credentials: No such process
monitor-agent.service: Failed at step USER spawning /opt/monitor/monitor-agent: No such process
monitor-agent.service: Main process exited, code=exited, status=217/USER
monitor-agent.service: Failed with result 'exit-code'.

getent passwd monitor-agent 无输出 —— 系统用户根本没被创建,而 #8 之后的 unit 写的是 User=monitor-agent。

最难受的一点:install.sh 自己退出码是 0,还打印了 monitor-agent installed,只有去看 journal 才知道服务其实起不来。面板上表现为该节点升级后掉线。

原因

install.sh 判断用户是否已存在用的是:

id -u monitor-agent >/dev/null 2>&1 ||
	useradd --system --no-create-home --shell /usr/sbin/nologin monitor-agent || ...

这段在脚本前部执行,此时旧的 v1.0.0 服务还在运行,而它的 unit 是 DynamicUser=yes ⇒ systemd 给它分配了一个名字同样是 monitor-agent 的瞬时用户。只要机器的 /etc/nsswitch.conf 里 passwd: 带 systemd(nss-systemd),id -u monitor-agent 就会成功,useradd 被跳过。

脚本随后写入新 unit 并 systemctl restart:旧服务一停,瞬时用户随之消失,新 unit 的 User=monitor-agent 再也解析不到。

最小复现(与本项目无关,任意 systemd 机器都能验)

passwd: files systemd 的机器上:

$ systemd-run --unit=nssprobe-dynuser --property=DynamicUser=yes /bin/sleep 12
$ id -u nssprobe-dynuser
63757
$ getent passwd nssprobe-dynuser
nssprobe-dynuser:x:63757:63757:Dynamic User:/:/usr/sbin/nologin
$ systemctl stop nssprobe-dynuser && id -u nssprobe-dynuser
id: 'nssprobe-dynuser': no such user

同一条命令在 passwd: files 的机器上,服务运行期间 id -u 就已经是 no such user。

我这边 4 台 Debian 对照:passwd: files systemd 的 2 台(Debian 13 / systemd 257,Debian 12 / systemd 252)中招;passwd: files 的 2 台(都是 Debian 12 / systemd 252)正常执行了 useradd、升级无恙。同样是 Debian 12 却两种 nsswitch.conf,估计和装机年代有关 —— 也就是说这事不看发行版版本,只看那一行有没有 systemd。

影响范围

只影响 v1.0.0(DynamicUser)→ ≥ v1.1.0(固定用户,#8)的原地升级,且机器启用了 nss-systemd。全新安装不受影响(没有旧服务在跑,id -u 自然失败)。

建议

  1. 存在性检查只查本地 passwd 库,绕开 nss-systemd(实测 -s files 在上面那台机器上运行期间就返回 rc=2,正确):

    getent -s files passwd monitor-agent >/dev/null 2>&1 || useradd --system ...
  2. 或者把建用户挪到写完新 unit、systemctl stop 旧服务之后再做。

  3. 另外建议脚本末尾加一次 systemctl is-active 确认后再打印 "installed" —— 这类「脚本成功、服务其实没起来」的静默失败,在批量升级时很难发现。

给遇到同样问题的人

sudo useradd --system --no-create-home --shell /usr/sbin/nologin monitor-agent
sudo systemctl restart monitor-agent

Activity

  1. stqfdyr commented on Sep 21, 2026

    @stqfdyr
    Collaborator

    感谢提出

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions