Skip to content

docs: explain the Nodes page readings and node resource metrics #39

Description

@Quick104

Add and check transcode nodes shows how to add a node, but not how to read Admin > Nodes afterwards. Admins running proxy or transcode nodes get no explanation of the Acceleration badges, the stale, Shared GPU, and Drift markers, the Load and Capacity blocks, node groups, the full Re-probe behavior, or the Prometheus series a node exports. Without that, they can't tell why a Healthy node encodes in software, or why work stops reaching a node. The server repo is removing its copy of this material in Silo-Server/silo-server#1635.

The old server docs are linked for reference: Monitoring Stream Nodes and the Node metrics section of the old Docker guide. Parts of them were stale, so verify each claim against current silo-server code before publishing.

What to change

New page: src/content/docs/docs/running-a-server/node-status.md (suggested slug docs/node-status, title "Read node status"). Link it from transcode-nodes.md and from step 3 of server-health.md.

  • Groups (old text):
    • Nodes with the same Group label are treated as co-located, on the same host or LAN.
    • A transcode node whose group has its own proxies always streams through one of them. It never falls back to another group's proxy.
    • A group takes work only while every enabled member is healthy. One unhealthy member takes the whole group out of service.
    • Disabled members don't count toward group health.
    • When any node has a group, a Groups chip row filters the proxy and transcode sections together. An amber dot on a chip means an enabled member is unhealthy.
  • State (old text):
    • The node's state shows in three places: the rail on its left edge, the dot in its header, and a label reading Disabled, Healthy, or Unhealthy.
    • A disabled node is in no pool and never gets new work. It renders dimmed and keeps the readings it had when it was switched off.
    • Silo checks every node every 30 seconds, and routes new work away from a node that doesn't answer.
    • Streams already on a node that stops answering move to another node on their next segment request.
    • Checked, in the Capacity block, is the time since the last check. Hover for the exact time.
    • A node whose GPU driver broke stays Healthy and encodes in software. The Acceleration block and the Drift badge show the problem; the health state doesn't.
  • Acceleration (old text):
    • The badge shows what the node's FFmpeg verified with a real single-frame encode on each candidate device, not what the configuration names.
    • A table of the badge readings:
      • QSV, VAAPI, or NVENC in green: the backend passed its probe. Hover for the device.
      • The same badges in amber: the backend is configured and in use, but its probe failed. Hover for FFmpeg's reason. Transcodes still try it, because Silo uses an explicitly configured backend as set.
      • The same badges, plain: the backend is in use, but the node reported no verification for it. This is normal when there were no candidate devices to probe.
      • SW: no hardware backend verified, so the node encodes in software.
      • SW with "the configured GPU devices are not accessible on this node" on hover: this process could not open any candidate device, so none was probed. This is normal on a proxy node that reads a cluster-wide playback.hw_device pointing at the transcode nodes' cards. It is not a driver failure.
    • Each render device the node sees is listed with a video-engine busy meter and a session count. A dash with no meter means nothing could measure that device.
  • Markers:
    • stale means nothing currently confirms the stored hardware report. It appears when the node's last health check is more than 10 minutes old, or when the capability hash the node advertises is missing or differs from the stored report.
    • An old report is not stale by itself. A node rebuilds its report every 15 minutes, and the server fetches it only when the hash changes.
    • An unhealthy node is never marked stale.
    • Shared GPU means another registered node sees the same physical card, for example two containers on one GPU. Silo matches cards by NVIDIA GPU UUID where available, otherwise by PCI slot on the same host.
    • When transcode nodes tie on job count, Silo picks the node whose card carries the fewest jobs.
    • Drift (amber) means a hardware refetch found the node got worse: a backend that passed now fails, or a render device is gone. Hover for the note.
    • Drift stays until exactly what was lost comes back. A reboot, a reworded FFmpeg error, the other card on a multi-GPU node probing cleanly, or a newly added GPU doesn't clear it.
    • Drift is a warning only; node selection ignores it. Re-probe the node to check whether it still holds.
  • Load (old text):
    • Load shows CPU, memory, the fullest sampled disk, and network throughput. Where available it also shows Silo process RAM.
    • The node samples these every 5 seconds. Silo keeps only the current sample, so use Prometheus for history.
    • Network has no meter, because the node reports bytes moved, not link speed.
    • The disk reading tints once its mount passes 85% used.
    • An unhealthy node shows No resource sample instead of its last numbers.
    • A healthy node showing No resource sample sent no sample: sampling is Linux-only, and older node builds send none.
    • Stale resource sample means the node is answering with an old sample.
    • A mount that stops responding keeps its last good numbers. A path the node can't see reads as unmeasurable, not as an empty disk.
  • Load inside a container (old text):
    • When the container's cgroup caps CPU or memory, Load shows the container's figures, not the host's.
    • Memory is the cgroup limit and working set, with page cache excluded.
    • CPU is the cgroup's usage against its CPU quota or cpuset. A container capped at 2 cores on a 64-core host reads 100% when pegged, and its core count reads 2.
    • A quota or cpuset as large as the whole machine counts as no cap, so the node reports host CPU. So does a container with no limit.
    • Network throughput is the container's own traffic.
    • Disk figures cover the transcode directory on every node, plus the library roots on the main server.
  • Capacity:
    • Capacity shows concurrent transcodes on a transcode node, or relayed streams on a proxy node, against Max Transcodes or Max Streams.
    • A proxy node also shows measured egress against Max Egress Bandwidth (Mbps).
    • An uncapped reading shows the number without a meter.
    • A reading tints when it reaches its cap. From then on, Silo sends new work to other nodes.
  • Node metrics in Prometheus (old text):
    • Proxy and transcode nodes serve /metrics on their app port without authentication. The main server exports the same gauges only on its SILO_METRICS_LISTEN address. See Logs and monitoring for scrape setup.
    • Host gauges: streamapp_node_cpu_percent, streamapp_node_memory_used_bytes, streamapp_node_memory_total_bytes, streamapp_node_network_rx_bps, streamapp_node_network_tx_bps.
    • Disk gauges: streamapp_node_disk_used_bytes, streamapp_node_disk_total_bytes, streamapp_node_disk_stale.
    • GPU gauges, labeled by device: streamapp_node_gpu_video_busy_percent, streamapp_node_gpu_render_busy_percent, streamapp_node_gpu_busy_percent, streamapp_node_gpu_sessions, streamapp_node_gpu_vram_used_bytes, streamapp_node_gpu_vram_total_bytes.
    • Disk series use a mount label that names a role, not a path: scratch for the transcode directory, and library-1, library-2, and so on for library roots.
    • Library numbers follow configuration order, so adding a library root can renumber the ones after it. Alert on mount="scratch" by name and on library mounts in aggregate.
    • Each host samples at most 8 mounts, scratch first. Silo logs how many roots it skipped under component=nodemetrics.
    • streamapp_node_disk_stale is 1 when the used and total values beside it are carried over from the last good measurement, and 0 when they're current. A mount Silo has never measured exports no disk series.
    • A node's /api/v1/health is unauthenticated and reports mount roles without paths. The paths appear only in GET /api/v2/admin/system/resources, which needs admin sign-in, and in each node's /status, which needs the node's bearer token.
  • GPU measurement sources:
    • For Intel and AMD, GPU busy comes from DRM fdinfo. It needs only /dev/dri access, which the VA-API overlay already gives.
    • Intel and AMD figures count only Silo's own FFmpeg processes, so a card shared with other software reads less busy than it is.
    • For NVIDIA, the figures come from nvidia-smi, which the NVIDIA Container Toolkit provides. They cover the whole GPU, including other tenants.
    • streamapp_node_gpu_busy_percent is whole-card utilization including other tenants, and only NVIDIA reports it today. Use it to alert on a shared GPU.
    • Silo has no whole-GPU sampling for Intel yet.
    • A GPU nothing could measure exports no engine or VRAM series, rather than zeros. streamapp_node_gpu_sessions still exports, because it comes from Silo's own session count.
    • After 5 consecutive nvidia-smi failures, Silo stops calling it and retries once every 10 minutes.
    • Re-probing the node (POST /api/v2/admin/nodes/{id}/reprobe) restarts nvidia-smi sampling immediately.
  • Example alerts:
    • SiloScratchVolumeFilling: streamapp_node_disk_used_bytes{mount="scratch"} / streamapp_node_disk_total_bytes{mount="scratch"} > 0.9 for 15m.
    • SiloNodeCPUSaturated: streamapp_node_cpu_percent > 90 for 15m. Check the Acceleration block for a failed probe.
    • SiloDiskMeasurementStale: streamapp_node_disk_stale == 1 for 15m. Without it, a volume that stopped answering at 40% and kept filling never trips the fill alert.

src/content/docs/docs/running-a-server/transcode-nodes.md

  • In "Add the node", add an optional Group step for nodes that share a host or LAN, and link to the groups section of the new page.
  • In "If work does not reach the node", add what the scratch guard does when every node is full (old text):
    • If every eligible transcode node is at 95% or more, Silo ignores the guard and selects a node anyway.
    • Silo never skips a node whose scratch fill it can't read.
    • Silo logs each change once, under component=nodepool, not once per session.
    • transcode node scratch volume nearly full, excluded from selection means the guard worked.
    • transcode node scratch volume nearly full, still selected because no eligible node has scratch headroom, together with transcode scratch guard ignored: every eligible node is over the scratch threshold, means sessions are landing on a disk that will fail mid-stream. Alert on this pair.
    • To fix it on the node, enlarge the volume, lower segment retention, or clear stale files from the transcode directory.
  • In "Maintain a node", add Force reload:
    • There is no button. Use POST /api/v2/admin/nodes/{id}/force-reload for one node, or POST /api/v2/admin/nodes/force-reload for every node. Both need admin sign-in.
    • Force reload makes the node re-read its configuration and ends every live playback session on it.
    • Saving a node's acceleration overrides already makes the node re-read its configuration for new transcodes. Use Force reload only when running sessions must pick up the change now.

src/content/docs/docs/running-a-server/playback.md (section "After a driver or device change") (old text)

  • Also use Re-probe after changing which devices a node's container can open (a /dev/dri passthrough, the NVIDIA overlay, or group membership).
  • Also use it after replacing FFmpeg in place at the same path.
  • Also use it to check whether a Drift badge still holds.
  • Re-probe discards the node's cached hardware results, re-verifies against live hardware, and sends the new report to the server. It doesn't restart the node or reload its configuration.
  • On an idle node, a re-probe can take a couple of minutes.
  • Silo refuses a re-probe (HTTP 409) while the node runs anything that opens an encoder: playback transcodes, reconstructed sessions, prepared downloads, or hardware chapter-thumbnail extraction.
  • While a re-probe runs, the node refuses new GPU work with HTTP 503, and Silo places those sessions on another node.
  • On success, the node's report, verified backends, and drift note update before the action returns.
  • If the probe can't finish, the action reports an error and the node keeps its previous report. Running it again is safe.
  • A repaired card needs no re-probe: Silo caches a failed probe for only 15 seconds, so the card verifies on its own within one 15-minute report cycle.
  • The exception is tone mapping. Silo keeps any non-empty tone-mapping result until a re-probe or restart, so a node whose GPU was broken at startup can stay software-only for tone mapping until then.

src/content/docs/docs/running-a-server/docker.md

The separate Docker issue leaves this out, so it belongs here. Add a "Docker inside an LXC container" subsection (old text):

  • A Docker container nested in an LXC container sees the physical machine's /proc/stat, /proc/loadavg, and /proc/meminfo. The LXC's limits live on a cgroup it can't see.

  • Without a fix, Silo there reports the physical host's CPU, memory, and load, not the LXC's.

  • The fix is to bind-mount the lxcfs view into the Silo container:

    volumes:
      - /proc/meminfo:/host/proc/meminfo:ro
      - /proc/stat:/host/proc/stat:ro
      - /proc/loadavg:/host/proc/loadavg:ro
  • The default Compose file already includes these mounts on the silo service. In the silo-proxy and silo-transcode examples they're commented out.

  • Silo uses each file under /host/proc when it's present and falls back to its own /proc otherwise, file by file. No setting is involved.

  • These mounts help only when the Docker host is itself an LXC container. A bare-metal or VM Docker host gains nothing from them.

  • lxcfs virtualizes /proc/loadavg only when it runs with loadavg accounting (lxcfs -l), which is off by default on Proxmox. Without it, load1 stays the physical host's while CPU and memory are correct.

src/content/docs/docs/running-a-server/logging.md

  • In "Metrics", link to the node metrics list on the new page.

Don't carry over

  • The old docs say every Silo process, including the main server, serves /metrics on the listener it serves traffic from. The main server's app port answers /metrics with 404, and serves metrics only on the SILO_METRICS_LISTEN address. Proxy and transcode nodes serve /metrics on their app port, and SILO_METRICS_LISTEN has no effect on them. The old scrape example that targets the main server's app port gets a 404.
  • The wiki's Settings -> Nodes path. The page is Admin > Nodes.
  • The example ports 8081 (transcode) and 8082 (proxy). The Compose examples publish transcode nodes on 8082 and proxy nodes on 8083.
  • "Use Force reload after changing a node's acceleration overrides." Saving the overrides already reloads the node's configuration for new transcodes, and Force reload also ends every live session on the node.
  • The wiki's definition of stale as only "no health check for more than ten minutes". It also appears when the node's advertised capability hash is missing or differs from the stored report.

AI disclosure: drafted with Claude Code (claude-opus-5-5[1m]) from an audit of the server repository's docs against the manual and current server code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions