Skip to content

docs: add guides for Prometheus monitoring and performance profiles #40

Description

@Quick104

Logs and monitoring explains how to turn on SILO_METRICS_LISTEN and ends there. It doesn't point admins to the example Prometheus rules and Grafana dashboard, or explain how to label and size scrapes, read the queue and workload metrics, or sample traces. The manual also has no guide for capturing a profile when a maintainer asks for one. Admins who run Prometheus or report a performance problem have nothing in the manual to follow. The server repo is removing its copies in Silo-Server/silo-server#1635.

The old server docs are linked for reference: Monitoring Silo, Profiling one Silo process, and Workload metrics, which moves to docs/architecture/. Parts of them were stale, so verify each claim against current silo-server code before publishing.

What to change

New page: src/content/docs/docs/running-a-server/monitoring.md (suggested slug docs/monitoring, title "Monitor Silo with Prometheus"). Move the "Metrics" section of logging.md here, and link the new page from logging.md and server-health.md.

  • Where metrics come from:

    • The main server serves /metrics only on the SILO_METRICS_LISTEN address, for example SILO_METRICS_LISTEN=127.0.0.1:9091. Metrics are off while it's unset.
    • The metrics listener has no authentication. Bind it to loopback or a private monitoring network, and don't publish it through Docker ports, a reverse proxy, or an ingress.
    • Inside Docker, 127.0.0.1 is the container itself, so Prometheus needs a private network route to the address you choose.
    • Proxy and transcode nodes serve /metrics on their app port. SILO_METRICS_LISTEN has no effect on them.
    • Never scrape or publish the profiling port set by SILO_DEBUG_LISTEN.
  • Example files (old text):

    • Link the examples in deploy/observability/: prometheus.yml, silo.rules.yml, silo.rules.test.yml, and grafana-dashboard.json.

    • Load silo.rules.yml into Prometheus. Its rules expect the scrape job to be named silo.

    • Import grafana-dashboard.json into Grafana and choose its Prometheus data source.

    • silo.rules.yml records silo:queue_items:max as max by (cluster, queue, state) (silo_queue_items).

    • It also records silo:queue_oldest_requested_timestamp_seconds:min as min by (cluster, queue, state) (silo_queue_oldest_requested_timestamp_seconds).

    • List its alerts:

      Alert Expression For
      SiloTargetUnavailable up{job="silo"} == 0 2m
      SiloResourceSampleStale silo_resource_sample_stale == 1 1m
      SiloQueueSamplingFailed silo_queue_sample_available == 0 2m
      SiloQueueAgeHigh (time() - silo:queue_oldest_requested_timestamp_seconds:min{state="queued"}) > 1800 10m
      SiloCgroupMemoryHigh silo_cgroup_memory_current_bytes / silo_cgroup_memory_limit_bytes > 0.9 5m
      SiloCgroupOOM increase(silo_cgroup_memory_oom_kills_total[5m]) > 0 none
      SiloCPUThrottled rate(silo_cgroup_cpu_throttled_periods_total[5m]) / clamp_min(rate(silo_cgroup_cpu_periods_total[5m]), 0.001) > 0.2 10m
      SiloOTLPExportFailure sum by (cluster, instance, signal) (increase(silo_otel_export_records_total{outcome="error"}[5m])) > 0 2m
      SiloScrapeSlow scrape_duration_seconds{job="silo"} > 1 5m
      SiloPostgresPoolSaturated sum without (state) (silo_postgres_pool_connections{state="acquired"}) / sum without (state) (silo_postgres_pool_connections{state="maximum"}) > 0.9 2m
  • Labels:

    • Add a cluster label to every target.
    • Put the deployment role in a process_role label, for example api, transcode, or proxy.
    • Don't relabel role. Silo's own metrics use it to name database, cache, and storage pool roles.
  • Scrape limits:

    • Start with a 15-second scrape interval, a 10-second timeout, and sample_limit: 20000.
    • If a scrape exceeds sample_limit, Prometheus rejects the whole scrape, so alert well before the limit.
    • On a large installation, check scrape_samples_scraped, scrape size, and scrape_duration_seconds before relying on those limits.
    • Histograms add one series per bucket plus a count and a sum. Include them when estimating series counts.
  • Retention:

    • Set retention time and size in Prometheus, and trace and log retention in your OTLP backend. Silo has no setting for either.
  • Other exporters:

    • Use node-exporter for host, filesystem, and network errors and latency.
    • Use a container exporter for orchestration quotas and OOM history.
    • Use GPU vendor exporters for whole-device power and thermal limits.
    • Use PostgreSQL, Redis, and storage exporters for service-side health.
    • Keep these exporters on the monitoring network too.
  • Queue and workload metrics (old text):

    • Every main-server replica reports the same shared queues. Aggregate silo_queue_items with max by (cluster, queue, state) and silo_queue_oldest_requested_timestamp_seconds with min. Never sum them across replicas.
    • Queue age includes delayed retries, so it isn't the age of the oldest job that can run now.
    • When silo_queue_sample_available is 0, Silo omits that queue's depth and age. The missing values don't mean an empty queue.
    • silo_queue_sample_errors_total counts failed queue samples.
    • Silo omits a queue sample once it's more than 90 seconds old.
    • These silo_work_* series are per process, so sum them across processes: silo_work_active, silo_work_attempts_total, silo_work_duration_seconds, silo_work_queue_wait_seconds, silo_work_progress_updates_total, and silo_work_recoveries_total.
    • silo_work_attempts_total{outcome="unknown"} means Silo couldn't confirm how an attempt ended. Don't count it as success.
    • One scheduled task can start several jobs, so attempts are not media items. Don't read item throughput from attempt counts.
    • A missing or stale target is not idle work. Keep panels for up, silo_queue_sample_available, and silo_queue_sample_timestamp_seconds.
  • Diagnose with metrics (old text). Add a table:

    • CPU saturation or throttling (SiloCPUThrottled): compare process CPU with the cgroup quota and throttled periods. Capture a CPU profile of the busy process. Check FFmpeg child CPU and competition from the rest of the host separately.
    • Growing memory or OOM (SiloCgroupMemoryHigh, SiloCgroupOOM): compare process RSS, Go heap, live FFmpeg children, and cgroup memory, then take two heap profiles to compare. Get exit and OOM history from Docker or the orchestrator if Silo died before a scrape.
    • Slow API with idle CPU (SiloPostgresPoolSaturated): check PostgreSQL acquired versus maximum connections, and Redis wait timeouts. A goroutine profile or a short trace shows what requests are waiting on.
    • Queue age rising (SiloQueueAgeHigh): check SiloQueueSamplingFailed first. Then check retry gates, attempt outcomes and progress, and the job's state in the admin UI. A heartbeat shows the worker is alive, not that it is making progress.
    • Stream interruption or lost node (SiloTargetUnavailable): check up, the node's last health check, and FFmpeg exit outcomes. Before calling a node restart harmless, play the same media on another node.
    • Missing traces (SiloOTLPExportFailure): check the sampling ratio, silo_otel_export_records_total, and the collector.
    • Missing hardware values (SiloResourceSampleStale): a missing value is not zero. Fix the measurement source, such as denied procfs access, a missing driver tool, or an unreachable mount.

src/content/docs/docs/running-a-server/logging.md

  • Replace the "Metrics" section with a short link to the new monitoring page.
  • In "OpenTelemetry export", say that Silo exports traces as well as logs.
  • Add OTEL_TRACES_SAMPLER. It accepts always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, or parentbased_traceidratio, and defaults to parentbased_traceidratio. Unsupported values fall back to the default.
  • Add OTEL_TRACES_SAMPLER_ARG: the sampled ratio from 0 to 1. The default is 0.01 (1% of requests).
  • Set the ratio to 1 only during a bounded investigation.
  • silo_otel_export_records_total{outcome="error"} counts failed exports.
  • Requests keep working when the collector is down.

New page: src/content/docs/docs/running-a-server/profiling.md (suggested slug docs/profiling, title "Collect a performance profile"). Link it from server-health.md and help/report-a-problem.md. The old commands assume a server repo checkout (scripts/silo-profile). Adjust them to a downloaded copy of the helper, as below.

  • Turn on the profiling listener (old text):

    • Add SILO_DEBUG_LISTEN=127.0.0.1:6060 to .env and recreate the container with docker compose up -d. The listener is off by default.
    • Only literal loopback addresses work, IPv4 or IPv6. Silo rejects hostnames, wildcard addresses, non-loopback addresses, and port 0.
    • A malformed value stops Silo from starting.
    • If the port is taken, Silo logs the failure and keeps running without the listener. silo_debug_listener_available and silo_debug_listener_failures_total show which happened.
    • The listener works on the main server and on proxy and transcode nodes. Each capture covers one process.
    • The app and Jellyfin ports never serve profiles. The web app can answer a profile URL with an HTML page, so check the response content instead of expecting a 404.
  • Keep it private (old text):

    • Any process in the same network namespace can connect. Loopback doesn't isolate tenants that share a host or pod.
    • Never publish the port through Docker ports, an ingress, a reverse proxy, or a load balancer.
    • Use 127.0.0.1 in tools, not localhost. The listener accepts only a literal loopback address in the Host header.
    • Profiles contain symbols, file paths, and stack traces. Send them to maintainers privately, and delete them after the review.
  • The capture helper (old text):

    • The helper is scripts/silo-profile in the server repo. It isn't in the image, so download it with curl -fsSLO https://raw.githubusercontent.com/Silo-Server/silo-server/main/scripts/silo-profile. The image already includes Node.js and curl.

    • The helper streams each capture to a private file. It stops at 128 MiB by default (--max-bytes), gives up after 75 seconds, and never overwrites an existing file.

    • Capture commands, run where the helper can reach the listener:

      mkdir -m 700 incident
      node silo-profile --profile cpu --seconds 30 --output incident/cpu.pprof
      node silo-profile --profile heap --gc 1 --output incident/heap.pprof
      node silo-profile --profile allocs --seconds 30 --output incident/allocs.pprof
      node silo-profile --profile goroutine --output incident/goroutine.pprof
      node silo-profile --profile trace --seconds 1 --max-bytes 67108864 --output incident/runtime.trace
    • Each capture gets a JSON sidecar that records the build, Go version, instance, start and end times, size, and checksum. valid: true means the capture completed.

    • An oversized, interrupted, or failed capture keeps its .partial name with valid: false. Don't send it as a complete capture.

    • Keep the exact silo binary that produced the capture, and note the image digest.

  • Docker (old text):

    • The listener binds the container's own loopback, so a published host port can't reach it. Run the helper inside the container and copy the results out:

      docker exec CONTAINER sh -c 'umask 077; mkdir -p /tmp/silo-incident'
      docker exec -i CONTAINER node - --profile heap --gc 1 --output /tmp/silo-incident/heap.pprof < silo-profile
      docker cp CONTAINER:/tmp/silo-incident/. incident/
      docker exec CONTAINER sh -c 'cat "$(command -v silo)"' > incident/silo
      chmod 600 incident/silo
    • To check the listener: docker exec CONTAINER curl --fail http://127.0.0.1:6060/debug/pprof/.

    • On a Linux host with Node.js, a privileged user can instead enter only the container's network namespace:

      container_pid=$(docker inspect --format '{{.State.Pid}}' CONTAINER)
      sudo nsenter --target "$container_pid" --net node silo-profile --profile heap --output incident/heap.pprof
  • Kubernetes and SSH (old text):

    • Target one pod, and keep the forwarding bound to a literal local address:

      kubectl port-forward --address 127.0.0.1 pod/POD 16060:6060
      node silo-profile --url http://127.0.0.1:16060 --profile heap --output incident/heap.pprof
    • If the container runtime can't forward to pod loopback, use kubectl exec the same way as the Docker recipe.

    • For a listener in an SSH host's own network namespace:

      ssh -N -L 127.0.0.1:16060:127.0.0.1:6060 OPERATOR_HOST
      node silo-profile --url http://127.0.0.1:16060 --profile heap --output incident/heap.pprof
    • An SSH host's loopback is not its containers' loopback. For a container on that host, use the Docker recipe.

  • Which profile to take (old text):

    Profile Shows Limit
    cpu Where Go code spends CPU time Default 30 seconds, maximum 60
    heap Sampled live Go memory; --gc 1 runs garbage collection first Immediate, or a delta of up to 60 seconds
    allocs Allocation churn, including memory already freed Immediate, or a delta of up to 60 seconds
    goroutine What each goroutine is doing, and growing goroutine counts Immediate, or a delta of up to 60 seconds
    threadcreate What created OS threads Immediate, or a delta of up to 60 seconds
    block Time spent waiting on synchronization Needs SILO_DEBUG_BLOCK_RATE at startup
    mutex Lock contention Needs SILO_DEBUG_MUTEX_FRACTION at startup
    trace Scheduler, garbage collection, and goroutine timing Default 1 second, maximum 5
    • Each process runs one capture at a time. Another request gets HTTP 429 and increments silo_debug_captures_busy_total.
    • A Go heap profile doesn't include FFmpeg, plugin, native library, file cache, or GPU memory. Compare it with process RSS and container memory.
  • Contention profiles:

    • Set SILO_DEBUG_BLOCK_RATE (allowed range 1,000,000–1,000,000,000) and SILO_DEBUG_MUTEX_FRACTION (allowed range 100–1,000,000) in .env, then restart.
    • Both default to 0, and both need SILO_DEBUG_LISTEN.
    • While a profile is off, requesting it returns HTTP 409. That doesn't mean there's no contention.
    • Set both back to 0 and restart after the investigation.

src/content/docs/docs/help/report-a-problem.md

  • Never attach metric captures, profiles, private URLs, account IDs, or media details to a public issue. Metric names and sanitized outcomes are enough to start.
  • If a maintainer asks for a profile, link to the new profiling page.

src/content/docs/docs/running-a-server/server-health.md

  • Link to the "Diagnose with metrics" table on the new monitoring page, and to the profiling page.

Don't carry over

  • The old docs say every Silo process, including the main server, serves /metrics on the listener it serves traffic from. The main server's app port answers /metrics with 404, and serves metrics only on the SILO_METRICS_LISTEN address. Proxy and transcode nodes serve /metrics on their app port, and SILO_METRICS_LISTEN has no effect on them.
  • The target in deploy/observability/prometheus.yml. It points at the main server's app port (silo-api:8080), which answers 404; a separate server change will fix it. Tell readers to scrape the SILO_METRICS_LISTEN address instead of copying that target.
  • The old fault-injection and release-validation exercises in monitoring.md. They are maintainer procedure, not admin guidance.

AI disclosure: drafted with Claude Code (claude-opus-5-5[1m]) from an audit of the server repository's docs against the manual and current server code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions