Logs and monitoring explains how to turn on SILO_METRICS_LISTEN and ends there. It doesn't point admins to the example Prometheus rules and Grafana dashboard, or explain how to label and size scrapes, read the queue and workload metrics, or sample traces. The manual also has no guide for capturing a profile when a maintainer asks for one. Admins who run Prometheus or report a performance problem have nothing in the manual to follow. The server repo is removing its copies in Silo-Server/silo-server#1635.
The old server docs are linked for reference: Monitoring Silo, Profiling one Silo process, and Workload metrics, which moves to docs/architecture/. Parts of them were stale, so verify each claim against current silo-server code before publishing.
What to change
New page: src/content/docs/docs/running-a-server/monitoring.md (suggested slug docs/monitoring, title "Monitor Silo with Prometheus"). Move the "Metrics" section of logging.md here, and link the new page from logging.md and server-health.md.
-
Where metrics come from:
- The main server serves
/metrics only on the SILO_METRICS_LISTEN address, for example SILO_METRICS_LISTEN=127.0.0.1:9091. Metrics are off while it's unset.
- The metrics listener has no authentication. Bind it to loopback or a private monitoring network, and don't publish it through Docker
ports, a reverse proxy, or an ingress.
- Inside Docker,
127.0.0.1 is the container itself, so Prometheus needs a private network route to the address you choose.
- Proxy and transcode nodes serve
/metrics on their app port. SILO_METRICS_LISTEN has no effect on them.
- Never scrape or publish the profiling port set by
SILO_DEBUG_LISTEN.
-
Example files (old text):
-
Link the examples in deploy/observability/: prometheus.yml, silo.rules.yml, silo.rules.test.yml, and grafana-dashboard.json.
-
Load silo.rules.yml into Prometheus. Its rules expect the scrape job to be named silo.
-
Import grafana-dashboard.json into Grafana and choose its Prometheus data source.
-
silo.rules.yml records silo:queue_items:max as max by (cluster, queue, state) (silo_queue_items).
-
It also records silo:queue_oldest_requested_timestamp_seconds:min as min by (cluster, queue, state) (silo_queue_oldest_requested_timestamp_seconds).
-
List its alerts:
| Alert |
Expression |
For |
SiloTargetUnavailable |
up{job="silo"} == 0 |
2m |
SiloResourceSampleStale |
silo_resource_sample_stale == 1 |
1m |
SiloQueueSamplingFailed |
silo_queue_sample_available == 0 |
2m |
SiloQueueAgeHigh |
(time() - silo:queue_oldest_requested_timestamp_seconds:min{state="queued"}) > 1800 |
10m |
SiloCgroupMemoryHigh |
silo_cgroup_memory_current_bytes / silo_cgroup_memory_limit_bytes > 0.9 |
5m |
SiloCgroupOOM |
increase(silo_cgroup_memory_oom_kills_total[5m]) > 0 |
none |
SiloCPUThrottled |
rate(silo_cgroup_cpu_throttled_periods_total[5m]) / clamp_min(rate(silo_cgroup_cpu_periods_total[5m]), 0.001) > 0.2 |
10m |
SiloOTLPExportFailure |
sum by (cluster, instance, signal) (increase(silo_otel_export_records_total{outcome="error"}[5m])) > 0 |
2m |
SiloScrapeSlow |
scrape_duration_seconds{job="silo"} > 1 |
5m |
SiloPostgresPoolSaturated |
sum without (state) (silo_postgres_pool_connections{state="acquired"}) / sum without (state) (silo_postgres_pool_connections{state="maximum"}) > 0.9 |
2m |
-
Labels:
- Add a
cluster label to every target.
- Put the deployment role in a
process_role label, for example api, transcode, or proxy.
- Don't relabel
role. Silo's own metrics use it to name database, cache, and storage pool roles.
-
Scrape limits:
- Start with a 15-second scrape interval, a 10-second timeout, and
sample_limit: 20000.
- If a scrape exceeds
sample_limit, Prometheus rejects the whole scrape, so alert well before the limit.
- On a large installation, check
scrape_samples_scraped, scrape size, and scrape_duration_seconds before relying on those limits.
- Histograms add one series per bucket plus a count and a sum. Include them when estimating series counts.
-
Retention:
- Set retention time and size in Prometheus, and trace and log retention in your OTLP backend. Silo has no setting for either.
-
Other exporters:
- Use node-exporter for host, filesystem, and network errors and latency.
- Use a container exporter for orchestration quotas and OOM history.
- Use GPU vendor exporters for whole-device power and thermal limits.
- Use PostgreSQL, Redis, and storage exporters for service-side health.
- Keep these exporters on the monitoring network too.
-
Queue and workload metrics (old text):
- Every main-server replica reports the same shared queues. Aggregate
silo_queue_items with max by (cluster, queue, state) and silo_queue_oldest_requested_timestamp_seconds with min. Never sum them across replicas.
- Queue age includes delayed retries, so it isn't the age of the oldest job that can run now.
- When
silo_queue_sample_available is 0, Silo omits that queue's depth and age. The missing values don't mean an empty queue.
silo_queue_sample_errors_total counts failed queue samples.
- Silo omits a queue sample once it's more than 90 seconds old.
- These
silo_work_* series are per process, so sum them across processes: silo_work_active, silo_work_attempts_total, silo_work_duration_seconds, silo_work_queue_wait_seconds, silo_work_progress_updates_total, and silo_work_recoveries_total.
silo_work_attempts_total{outcome="unknown"} means Silo couldn't confirm how an attempt ended. Don't count it as success.
- One scheduled task can start several jobs, so attempts are not media items. Don't read item throughput from attempt counts.
- A missing or stale target is not idle work. Keep panels for
up, silo_queue_sample_available, and silo_queue_sample_timestamp_seconds.
-
Diagnose with metrics (old text). Add a table:
- CPU saturation or throttling (
SiloCPUThrottled): compare process CPU with the cgroup quota and throttled periods. Capture a CPU profile of the busy process. Check FFmpeg child CPU and competition from the rest of the host separately.
- Growing memory or OOM (
SiloCgroupMemoryHigh, SiloCgroupOOM): compare process RSS, Go heap, live FFmpeg children, and cgroup memory, then take two heap profiles to compare. Get exit and OOM history from Docker or the orchestrator if Silo died before a scrape.
- Slow API with idle CPU (
SiloPostgresPoolSaturated): check PostgreSQL acquired versus maximum connections, and Redis wait timeouts. A goroutine profile or a short trace shows what requests are waiting on.
- Queue age rising (
SiloQueueAgeHigh): check SiloQueueSamplingFailed first. Then check retry gates, attempt outcomes and progress, and the job's state in the admin UI. A heartbeat shows the worker is alive, not that it is making progress.
- Stream interruption or lost node (
SiloTargetUnavailable): check up, the node's last health check, and FFmpeg exit outcomes. Before calling a node restart harmless, play the same media on another node.
- Missing traces (
SiloOTLPExportFailure): check the sampling ratio, silo_otel_export_records_total, and the collector.
- Missing hardware values (
SiloResourceSampleStale): a missing value is not zero. Fix the measurement source, such as denied procfs access, a missing driver tool, or an unreachable mount.
src/content/docs/docs/running-a-server/logging.md
- Replace the "Metrics" section with a short link to the new monitoring page.
- In "OpenTelemetry export", say that Silo exports traces as well as logs.
- Add
OTEL_TRACES_SAMPLER. It accepts always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, or parentbased_traceidratio, and defaults to parentbased_traceidratio. Unsupported values fall back to the default.
- Add
OTEL_TRACES_SAMPLER_ARG: the sampled ratio from 0 to 1. The default is 0.01 (1% of requests).
- Set the ratio to
1 only during a bounded investigation.
silo_otel_export_records_total{outcome="error"} counts failed exports.
- Requests keep working when the collector is down.
New page: src/content/docs/docs/running-a-server/profiling.md (suggested slug docs/profiling, title "Collect a performance profile"). Link it from server-health.md and help/report-a-problem.md. The old commands assume a server repo checkout (scripts/silo-profile). Adjust them to a downloaded copy of the helper, as below.
-
Turn on the profiling listener (old text):
- Add
SILO_DEBUG_LISTEN=127.0.0.1:6060 to .env and recreate the container with docker compose up -d. The listener is off by default.
- Only literal loopback addresses work, IPv4 or IPv6. Silo rejects hostnames, wildcard addresses, non-loopback addresses, and port
0.
- A malformed value stops Silo from starting.
- If the port is taken, Silo logs the failure and keeps running without the listener.
silo_debug_listener_available and silo_debug_listener_failures_total show which happened.
- The listener works on the main server and on proxy and transcode nodes. Each capture covers one process.
- The app and Jellyfin ports never serve profiles. The web app can answer a profile URL with an HTML page, so check the response content instead of expecting a 404.
-
Keep it private (old text):
- Any process in the same network namespace can connect. Loopback doesn't isolate tenants that share a host or pod.
- Never publish the port through Docker
ports, an ingress, a reverse proxy, or a load balancer.
- Use
127.0.0.1 in tools, not localhost. The listener accepts only a literal loopback address in the Host header.
- Profiles contain symbols, file paths, and stack traces. Send them to maintainers privately, and delete them after the review.
-
The capture helper (old text):
-
The helper is scripts/silo-profile in the server repo. It isn't in the image, so download it with curl -fsSLO https://raw.githubusercontent.com/Silo-Server/silo-server/main/scripts/silo-profile. The image already includes Node.js and curl.
-
The helper streams each capture to a private file. It stops at 128 MiB by default (--max-bytes), gives up after 75 seconds, and never overwrites an existing file.
-
Capture commands, run where the helper can reach the listener:
mkdir -m 700 incident
node silo-profile --profile cpu --seconds 30 --output incident/cpu.pprof
node silo-profile --profile heap --gc 1 --output incident/heap.pprof
node silo-profile --profile allocs --seconds 30 --output incident/allocs.pprof
node silo-profile --profile goroutine --output incident/goroutine.pprof
node silo-profile --profile trace --seconds 1 --max-bytes 67108864 --output incident/runtime.trace
-
Each capture gets a JSON sidecar that records the build, Go version, instance, start and end times, size, and checksum. valid: true means the capture completed.
-
An oversized, interrupted, or failed capture keeps its .partial name with valid: false. Don't send it as a complete capture.
-
Keep the exact silo binary that produced the capture, and note the image digest.
-
Docker (old text):
-
The listener binds the container's own loopback, so a published host port can't reach it. Run the helper inside the container and copy the results out:
docker exec CONTAINER sh -c 'umask 077; mkdir -p /tmp/silo-incident'
docker exec -i CONTAINER node - --profile heap --gc 1 --output /tmp/silo-incident/heap.pprof < silo-profile
docker cp CONTAINER:/tmp/silo-incident/. incident/
docker exec CONTAINER sh -c 'cat "$(command -v silo)"' > incident/silo
chmod 600 incident/silo
-
To check the listener: docker exec CONTAINER curl --fail http://127.0.0.1:6060/debug/pprof/.
-
On a Linux host with Node.js, a privileged user can instead enter only the container's network namespace:
container_pid=$(docker inspect --format '{{.State.Pid}}' CONTAINER)
sudo nsenter --target "$container_pid" --net node silo-profile --profile heap --output incident/heap.pprof
-
Kubernetes and SSH (old text):
-
Target one pod, and keep the forwarding bound to a literal local address:
kubectl port-forward --address 127.0.0.1 pod/POD 16060:6060
node silo-profile --url http://127.0.0.1:16060 --profile heap --output incident/heap.pprof
-
If the container runtime can't forward to pod loopback, use kubectl exec the same way as the Docker recipe.
-
For a listener in an SSH host's own network namespace:
ssh -N -L 127.0.0.1:16060:127.0.0.1:6060 OPERATOR_HOST
node silo-profile --url http://127.0.0.1:16060 --profile heap --output incident/heap.pprof
-
An SSH host's loopback is not its containers' loopback. For a container on that host, use the Docker recipe.
-
Which profile to take (old text):
| Profile |
Shows |
Limit |
cpu |
Where Go code spends CPU time |
Default 30 seconds, maximum 60 |
heap |
Sampled live Go memory; --gc 1 runs garbage collection first |
Immediate, or a delta of up to 60 seconds |
allocs |
Allocation churn, including memory already freed |
Immediate, or a delta of up to 60 seconds |
goroutine |
What each goroutine is doing, and growing goroutine counts |
Immediate, or a delta of up to 60 seconds |
threadcreate |
What created OS threads |
Immediate, or a delta of up to 60 seconds |
block |
Time spent waiting on synchronization |
Needs SILO_DEBUG_BLOCK_RATE at startup |
mutex |
Lock contention |
Needs SILO_DEBUG_MUTEX_FRACTION at startup |
trace |
Scheduler, garbage collection, and goroutine timing |
Default 1 second, maximum 5 |
- Each process runs one capture at a time. Another request gets HTTP 429 and increments
silo_debug_captures_busy_total.
- A Go heap profile doesn't include FFmpeg, plugin, native library, file cache, or GPU memory. Compare it with process RSS and container memory.
-
Contention profiles:
- Set
SILO_DEBUG_BLOCK_RATE (allowed range 1,000,000–1,000,000,000) and SILO_DEBUG_MUTEX_FRACTION (allowed range 100–1,000,000) in .env, then restart.
- Both default to
0, and both need SILO_DEBUG_LISTEN.
- While a profile is off, requesting it returns HTTP 409. That doesn't mean there's no contention.
- Set both back to
0 and restart after the investigation.
src/content/docs/docs/help/report-a-problem.md
- Never attach metric captures, profiles, private URLs, account IDs, or media details to a public issue. Metric names and sanitized outcomes are enough to start.
- If a maintainer asks for a profile, link to the new profiling page.
src/content/docs/docs/running-a-server/server-health.md
- Link to the "Diagnose with metrics" table on the new monitoring page, and to the profiling page.
Don't carry over
- The old docs say every Silo process, including the main server, serves
/metrics on the listener it serves traffic from. The main server's app port answers /metrics with 404, and serves metrics only on the SILO_METRICS_LISTEN address. Proxy and transcode nodes serve /metrics on their app port, and SILO_METRICS_LISTEN has no effect on them.
- The target in
deploy/observability/prometheus.yml. It points at the main server's app port (silo-api:8080), which answers 404; a separate server change will fix it. Tell readers to scrape the SILO_METRICS_LISTEN address instead of copying that target.
- The old fault-injection and release-validation exercises in
monitoring.md. They are maintainer procedure, not admin guidance.
AI disclosure: drafted with Claude Code (claude-opus-5-5[1m]) from an audit of the server repository's docs against the manual and current server code.
Logs and monitoring explains how to turn on
SILO_METRICS_LISTENand ends there. It doesn't point admins to the example Prometheus rules and Grafana dashboard, or explain how to label and size scrapes, read the queue and workload metrics, or sample traces. The manual also has no guide for capturing a profile when a maintainer asks for one. Admins who run Prometheus or report a performance problem have nothing in the manual to follow. The server repo is removing its copies in Silo-Server/silo-server#1635.The old server docs are linked for reference: Monitoring Silo, Profiling one Silo process, and Workload metrics, which moves to
docs/architecture/. Parts of them were stale, so verify each claim against current silo-server code before publishing.What to change
New page:
src/content/docs/docs/running-a-server/monitoring.md(suggested slugdocs/monitoring, title "Monitor Silo with Prometheus"). Move the "Metrics" section oflogging.mdhere, and link the new page fromlogging.mdandserver-health.md.Where metrics come from:
/metricsonly on theSILO_METRICS_LISTENaddress, for exampleSILO_METRICS_LISTEN=127.0.0.1:9091. Metrics are off while it's unset.ports, a reverse proxy, or an ingress.127.0.0.1is the container itself, so Prometheus needs a private network route to the address you choose./metricson their app port.SILO_METRICS_LISTENhas no effect on them.SILO_DEBUG_LISTEN.Example files (old text):
Link the examples in
deploy/observability/:prometheus.yml,silo.rules.yml,silo.rules.test.yml, andgrafana-dashboard.json.Load
silo.rules.ymlinto Prometheus. Its rules expect the scrape job to be namedsilo.Import
grafana-dashboard.jsoninto Grafana and choose its Prometheus data source.silo.rules.ymlrecordssilo:queue_items:maxasmax by (cluster, queue, state) (silo_queue_items).It also records
silo:queue_oldest_requested_timestamp_seconds:minasmin by (cluster, queue, state) (silo_queue_oldest_requested_timestamp_seconds).List its alerts:
SiloTargetUnavailableup{job="silo"} == 0SiloResourceSampleStalesilo_resource_sample_stale == 1SiloQueueSamplingFailedsilo_queue_sample_available == 0SiloQueueAgeHigh(time() - silo:queue_oldest_requested_timestamp_seconds:min{state="queued"}) > 1800SiloCgroupMemoryHighsilo_cgroup_memory_current_bytes / silo_cgroup_memory_limit_bytes > 0.9SiloCgroupOOMincrease(silo_cgroup_memory_oom_kills_total[5m]) > 0SiloCPUThrottledrate(silo_cgroup_cpu_throttled_periods_total[5m]) / clamp_min(rate(silo_cgroup_cpu_periods_total[5m]), 0.001) > 0.2SiloOTLPExportFailuresum by (cluster, instance, signal) (increase(silo_otel_export_records_total{outcome="error"}[5m])) > 0SiloScrapeSlowscrape_duration_seconds{job="silo"} > 1SiloPostgresPoolSaturatedsum without (state) (silo_postgres_pool_connections{state="acquired"}) / sum without (state) (silo_postgres_pool_connections{state="maximum"}) > 0.9Labels:
clusterlabel to every target.process_rolelabel, for exampleapi,transcode, orproxy.role. Silo's own metrics use it to name database, cache, and storage pool roles.Scrape limits:
sample_limit: 20000.sample_limit, Prometheus rejects the whole scrape, so alert well before the limit.scrape_samples_scraped, scrape size, andscrape_duration_secondsbefore relying on those limits.Retention:
Other exporters:
Queue and workload metrics (old text):
silo_queue_itemswithmax by (cluster, queue, state)andsilo_queue_oldest_requested_timestamp_secondswithmin. Never sum them across replicas.silo_queue_sample_availableis0, Silo omits that queue's depth and age. The missing values don't mean an empty queue.silo_queue_sample_errors_totalcounts failed queue samples.silo_work_*series are per process, so sum them across processes:silo_work_active,silo_work_attempts_total,silo_work_duration_seconds,silo_work_queue_wait_seconds,silo_work_progress_updates_total, andsilo_work_recoveries_total.silo_work_attempts_total{outcome="unknown"}means Silo couldn't confirm how an attempt ended. Don't count it as success.up,silo_queue_sample_available, andsilo_queue_sample_timestamp_seconds.Diagnose with metrics (old text). Add a table:
SiloCPUThrottled): compare process CPU with the cgroup quota and throttled periods. Capture a CPU profile of the busy process. Check FFmpeg child CPU and competition from the rest of the host separately.SiloCgroupMemoryHigh,SiloCgroupOOM): compare process RSS, Go heap, live FFmpeg children, and cgroup memory, then take two heap profiles to compare. Get exit and OOM history from Docker or the orchestrator if Silo died before a scrape.SiloPostgresPoolSaturated): check PostgreSQL acquired versus maximum connections, and Redis wait timeouts. A goroutine profile or a short trace shows what requests are waiting on.SiloQueueAgeHigh): checkSiloQueueSamplingFailedfirst. Then check retry gates, attempt outcomes and progress, and the job's state in the admin UI. A heartbeat shows the worker is alive, not that it is making progress.SiloTargetUnavailable): checkup, the node's last health check, and FFmpeg exit outcomes. Before calling a node restart harmless, play the same media on another node.SiloOTLPExportFailure): check the sampling ratio,silo_otel_export_records_total, and the collector.SiloResourceSampleStale): a missing value is not zero. Fix the measurement source, such as denied procfs access, a missing driver tool, or an unreachable mount.src/content/docs/docs/running-a-server/logging.mdOTEL_TRACES_SAMPLER. It acceptsalways_on,always_off,traceidratio,parentbased_always_on,parentbased_always_off, orparentbased_traceidratio, and defaults toparentbased_traceidratio. Unsupported values fall back to the default.OTEL_TRACES_SAMPLER_ARG: the sampled ratio from 0 to 1. The default is0.01(1% of requests).1only during a bounded investigation.silo_otel_export_records_total{outcome="error"}counts failed exports.New page:
src/content/docs/docs/running-a-server/profiling.md(suggested slugdocs/profiling, title "Collect a performance profile"). Link it fromserver-health.mdandhelp/report-a-problem.md. The old commands assume a server repo checkout (scripts/silo-profile). Adjust them to a downloaded copy of the helper, as below.Turn on the profiling listener (old text):
SILO_DEBUG_LISTEN=127.0.0.1:6060to.envand recreate the container withdocker compose up -d. The listener is off by default.0.silo_debug_listener_availableandsilo_debug_listener_failures_totalshow which happened.Keep it private (old text):
ports, an ingress, a reverse proxy, or a load balancer.127.0.0.1in tools, notlocalhost. The listener accepts only a literal loopback address in theHostheader.The capture helper (old text):
The helper is
scripts/silo-profilein the server repo. It isn't in the image, so download it withcurl -fsSLO https://raw.githubusercontent.com/Silo-Server/silo-server/main/scripts/silo-profile. The image already includes Node.js and curl.The helper streams each capture to a private file. It stops at 128 MiB by default (
--max-bytes), gives up after 75 seconds, and never overwrites an existing file.Capture commands, run where the helper can reach the listener:
Each capture gets a JSON sidecar that records the build, Go version, instance, start and end times, size, and checksum.
valid: truemeans the capture completed.An oversized, interrupted, or failed capture keeps its
.partialname withvalid: false. Don't send it as a complete capture.Keep the exact
silobinary that produced the capture, and note the image digest.Docker (old text):
The listener binds the container's own loopback, so a published host port can't reach it. Run the helper inside the container and copy the results out:
To check the listener:
docker exec CONTAINER curl --fail http://127.0.0.1:6060/debug/pprof/.On a Linux host with Node.js, a privileged user can instead enter only the container's network namespace:
Kubernetes and SSH (old text):
Target one pod, and keep the forwarding bound to a literal local address:
If the container runtime can't forward to pod loopback, use
kubectl execthe same way as the Docker recipe.For a listener in an SSH host's own network namespace:
An SSH host's loopback is not its containers' loopback. For a container on that host, use the Docker recipe.
Which profile to take (old text):
cpuheap--gc 1runs garbage collection firstallocsgoroutinethreadcreateblockSILO_DEBUG_BLOCK_RATEat startupmutexSILO_DEBUG_MUTEX_FRACTIONat startuptracesilo_debug_captures_busy_total.Contention profiles:
SILO_DEBUG_BLOCK_RATE(allowed range 1,000,000–1,000,000,000) andSILO_DEBUG_MUTEX_FRACTION(allowed range 100–1,000,000) in.env, then restart.0, and both needSILO_DEBUG_LISTEN.0and restart after the investigation.src/content/docs/docs/help/report-a-problem.mdsrc/content/docs/docs/running-a-server/server-health.mdDon't carry over
/metricson the listener it serves traffic from. The main server's app port answers/metricswith 404, and serves metrics only on theSILO_METRICS_LISTENaddress. Proxy and transcode nodes serve/metricson their app port, andSILO_METRICS_LISTENhas no effect on them.deploy/observability/prometheus.yml. It points at the main server's app port (silo-api:8080), which answers 404; a separate server change will fix it. Tell readers to scrape theSILO_METRICS_LISTENaddress instead of copying that target.monitoring.md. They are maintainer procedure, not admin guidance.AI disclosure: drafted with Claude Code (claude-opus-5-5[1m]) from an audit of the server repository's docs against the manual and current server code.