Skip to content

feat: cluster-wide Swarm discovery behind DOCKTAIL_DISCOVERY=swarm - #101

Open
paolomemoli wants to merge 1 commit into
marvinvr:mainfrom
paolomemoli:docktail-swarm
Open

paolomemoli wants to merge 1 commit into
marvinvr:mainfrom
paolomemoli:docktail-swarm

Conversation

@paolomemoli

@paolomemoli paolomemoli commented Sep 30, 2026 •

Copy link
Copy Markdown

feat: cluster-wide Swarm discovery behind DOCKTAIL_DISCOVERY=swarm

This is an AI assisted PR - it worked for me and I think it might be useful for others but please remove if not relevent.

Container discovery is node-local. The /containers endpoint returns only the containers running on the Docker node that serves it — on a manager just as much as a worker. An agent therefore advertises only the services it happens to be colocated with, which in a Swarm cluster means one agent per node, one tailscaled state per node, and an N-node auth key for a single logical sidecar. That's the workaround described in #43.

This adds a second discovery source that reads the whole cluster from one agent.

How it works

/services is manager-only and cluster-wide. Each service carries:

  • Spec.TaskTemplate.ContainerSpec.Labels — where both labels: and deploy.labels: from a compose file land, so a service is labelled exactly like the equivalent standalone container.
  • Endpoint.VirtualIPs — the VIP the routing mesh keeps alive and load-balanced across that service's tasks, on the overlay the tasks are attached to. A VIP is reachable from any node on that overlay, so one agent covers the cluster as long as the overlay carries the traffic.

The service VIP stands in for the container IP, which means the label surface is reused unchanged: parseContainerServices is factored into parseFromInspect so both sources run the same translation, protocol defaults, path handling, tags, funnel config and validation.

What this changes

services:
  docktail:
    image: ghcr.io/marvinvr/docktail:latest
    environment:
      - DOCKTAIL_DISCOVERY=swarm
      # optional: pin which network a multi-homed service's VIP comes from
      - DOCKTAIL_SWARM_NETWORK=tailscale

Opt-in. DOCKTAIL_DISCOVERY=containers (the default) is byte-for-byte the current behaviour.

Operational notes

  • A manager endpoint is required. /services is manager-only; on a worker-only agent the reconcile fails and nothing is advertised.
  • One agent is enough for a cluster, because the service VIP is cluster-wide.
  • The overlay carries the traffic. A VIP is only dialable from a node attached to the same overlay, so the agent (and the tailscaled it configures) must be on it too. When a service sits on several networks, DockTail prefers one it is also attached to; DOCKTAIL_SWARM_NETWORK pins the choice.
  • A VIP only exists in vip mode. A dnsrr service, or one with no network, has none and is reported and skipped rather than silently dropped.
  • Cross-node changes arrive on the next RECONCILE_INTERVAL. Docker event watching is still local, so a service created or relabelled on another node is picked up on the next tick rather than immediately.

Field results

Six services across two nodes, from one agent, with tailscale as a shared overlay:

Found enabled containers count=6 source=swarm

INF Proxying directly to container IP container=entertainment_plex        container_ip=10.0.4.7  container_port=32400 network=tailscale
INF Proxying directly to container IP container=entertainment_radarr       container_ip=10.0.4.9  container_port=7878  network=tailscale
INF Proxying directly to container IP container=entertainment_sonarr       container_ip=10.0.4.11 container_port=8989  network=tailscale
INF Proxying directly to container IP container=entertainment_overseerr    container_ip=10.0.4.13 container_port=5055  network=tailscale
INF Proxying directly to container IP container=entertainment_prowlarr     container_ip=10.0.4.2  container_port=9696  network=tailscale
INF Proxying directly to container IP container=entertainment_qbittorrent  container_ip=10.0.4.5  container_port=8080  network=tailscale

INF Service reconciliation completed added=6 failed=0 removed=0

Every container ran on cloud-spain; the agent ran on cloud. All six resolve over the tailnet:

https://plex.char-wage.ts.net      -> 401   (Plex auth challenge)
https://films.char-wage.ts.net     -> 302   (Radarr -> /login)
https://tv.char-wage.ts.net        -> 302   (Sonarr -> /login)
https://requests.char-wage.ts.net  -> 307   (Overseerr)
https://index.char-wage.ts.net     -> 302   (Prowlarr -> /login)
https://torrents.char-wage.ts.net  -> 200   (qBittorrent)

Built and run as a multi-arch image (linux/amd64, linux/arm64) for that deployment.

A caveat worth documenting

This surfaced a silent failure mode worth calling out in the existing docs, since it cost us an afternoon and presents as "the agent advertises, but nothing answers":

Tailscale Services do not work when tailscaled runs in userspace mode. containerboot silently falls back to --tun=userspace-networking when it cannot find /dev/net/tun, so a compose file that mounts the tun device at the wrong path — or omits TS_USERSPACE=false — produces a node that authenticates, advertises svc:* happily, and answers nothing on the VIP. The service config and the control plane both look correct, which makes it very hard to spot from the outside.

I have not included a docs change for this since it's adjacent to the feature rather than part of it — but it's a good candidate for a follow-up (or a startup check that logs loudly when it detects userspace mode).

Testing

  • 5 new test groups in docker/swarm_test.go covering CIDR stripping, numeric-vs-lexicographic VIP ordering, label extraction, network-name resolution, host-mode detection, and all four VIP-selection paths (no VIP, lowest, shared-network preference, operator-pinned).
  • go build ./..., go vet ./..., go test ./... all pass; no existing tests changed.

Review notes

The one design decision worth a second opinion: swarmNetworkName uses NetworkInspect, because the alias in a NetworkAttachmentConfig is the DNS name the service gets on that network, not the network's own name, so it cannot answer DOCKTAIL_SWARM_NETWORK. Results are cached in a sync.Map on the Client since network names don't move under a running cluster. An alternative would be to compare against Endpoint.VirtualIPs positionally, but that couples the preference to ordering that isn't guaranteed stable.

cc #43

Container discovery is node-local: /containers returns only the containers
running on the Docker node that serves it, on a manager exactly as on a
worker. An agent therefore advertises only the services it is colocated
with, which in a Swarm cluster means one agent per node.

Add a second discovery source that reads the whole cluster. /services is
manager-only and returns every service with its labels, and each service
carries Endpoint.VirtualIPs: the VIP the routing mesh keeps alive and
load-balanced across that service's tasks, on the overlay its tasks are
attached to. A VIP is reachable from any node on that overlay, so one agent
covers the cluster as long as the overlay carries the traffic.

Labels are read from Spec.TaskTemplate.ContainerSpec.Labels, where both
`labels:` and `deploy.labels:` land, so a Swarm service is labelled exactly
like the equivalent standalone container and every existing label applies
unchanged. Parsing is shared: the label translation is factored out of the
inspect-and-parse path so both sources run the same code.

The destination is the service VIP rather than a container IP, so the mode
requires the agent and the exposed services to share an overlay. When a
service sits on several networks, DockTail prefers one it is also attached
to, and DOCKTAIL_SWARM_NETWORK pins the choice. A VIP only exists in vip
mode, so a dnsrr service or one with no network is reported and skipped.

Event watching stays local, so a change on another node lands on the next
RECONCILE_INTERVAL rather than immediately.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant