Problem
deploy/observability/prometheus.yml scrapes the main server at silo-api:8080, its application port. That listener answers /metrics with 404 on purpose; the main server serves metrics only on the separate SILO_METRICS_LISTEN listener. An administrator who adapts the example gets a failing API target (SiloTargetUnavailable) and no API metrics. The transcode target, silo-worker:8080, is correct, because proxy and transcode nodes serve /metrics on their application listener.
Two code comments state the same wrong model: the node handlers say their unauthenticated /metrics matches "the API listener's own /metrics posture". The API listener has no /metrics; the comparison that holds is with the dedicated metrics listener, which is also unauthenticated.
Found by reading the code while auditing the docs for #1635; not reproduced against a running Prometheus.
Fix
- Point the API target at the metrics listener, for example
silo-api:9091 with SILO_METRICS_LISTEN bound to a private address that Prometheus can reach. 127.0.0.1 inside the Silo container is not reachable from a separate Prometheus container, so say that in the file's header comment.
- Keep the node target on the app port, and say in the comment that
SILO_METRICS_LISTEN has no effect on proxy and transcode nodes.
- Reword the two node comments to compare with the dedicated metrics listener.
- Run the
promtool checks in docs/architecture/observability.md#validating-observability-changes after editing.
Technical notes
deploy/observability/prometheus.yml: targets: [silo-api:8080].
cmd/silo/root_handler.go:46: mux.Handle("/metrics", http.NotFoundHandler()) on the root listener.
cmd/silo/metrics_listener.go: /metrics only when SILO_METRICS_LISTEN is set.
internal/proxy/server.go:268 and internal/transcodenode/server.go:883: the stale comments.
AI disclosure
- Harness: T3 Code (Claude Code agent harness)
- Tool(s): Claude Code, GitHub CLI
- Model(s): claude-opus-5-5
- Involvement: AI-assisted; filed at the maintainer's request
- Adversarial review: n/a
Problem
deploy/observability/prometheus.ymlscrapes the main server atsilo-api:8080, its application port. That listener answers/metricswith 404 on purpose; the main server serves metrics only on the separateSILO_METRICS_LISTENlistener. An administrator who adapts the example gets a failing API target (SiloTargetUnavailable) and no API metrics. The transcode target,silo-worker:8080, is correct, because proxy and transcode nodes serve/metricson their application listener.Two code comments state the same wrong model: the node handlers say their unauthenticated
/metricsmatches "the API listener's own /metrics posture". The API listener has no/metrics; the comparison that holds is with the dedicated metrics listener, which is also unauthenticated.Found by reading the code while auditing the docs for #1635; not reproduced against a running Prometheus.
Fix
silo-api:9091withSILO_METRICS_LISTENbound to a private address that Prometheus can reach.127.0.0.1inside the Silo container is not reachable from a separate Prometheus container, so say that in the file's header comment.SILO_METRICS_LISTENhas no effect on proxy and transcode nodes.promtoolchecks indocs/architecture/observability.md#validating-observability-changesafter editing.Technical notes
deploy/observability/prometheus.yml:targets: [silo-api:8080].cmd/silo/root_handler.go:46:mux.Handle("/metrics", http.NotFoundHandler())on the root listener.cmd/silo/metrics_listener.go:/metricsonly whenSILO_METRICS_LISTENis set.internal/proxy/server.go:268andinternal/transcodenode/server.go:883: the stale comments.AI disclosure