feat: cluster-wide Swarm discovery behind DOCKTAIL_DISCOVERY=swarm - #101
Open
paolomemoli wants to merge 1 commit into
Open
paolomemoli wants to merge 1 commit into
paolomemoli wants to merge 1 commit into
Conversation
Container discovery is node-local: /containers returns only the containers running on the Docker node that serves it, on a manager exactly as on a worker. An agent therefore advertises only the services it is colocated with, which in a Swarm cluster means one agent per node. Add a second discovery source that reads the whole cluster. /services is manager-only and returns every service with its labels, and each service carries Endpoint.VirtualIPs: the VIP the routing mesh keeps alive and load-balanced across that service's tasks, on the overlay its tasks are attached to. A VIP is reachable from any node on that overlay, so one agent covers the cluster as long as the overlay carries the traffic. Labels are read from Spec.TaskTemplate.ContainerSpec.Labels, where both `labels:` and `deploy.labels:` land, so a Swarm service is labelled exactly like the equivalent standalone container and every existing label applies unchanged. Parsing is shared: the label translation is factored out of the inspect-and-parse path so both sources run the same code. The destination is the service VIP rather than a container IP, so the mode requires the agent and the exposed services to share an overlay. When a service sits on several networks, DockTail prefers one it is also attached to, and DOCKTAIL_SWARM_NETWORK pins the choice. A VIP only exists in vip mode, so a dnsrr service or one with no network is reported and skipped. Event watching stays local, so a change on another node lands on the next RECONCILE_INTERVAL rather than immediately.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat: cluster-wide Swarm discovery behind
DOCKTAIL_DISCOVERY=swarmThis is an AI assisted PR - it worked for me and I think it might be useful for others but please remove if not relevent.
Container discovery is node-local. The
/containersendpoint returns only the containers running on the Docker node that serves it — on a manager just as much as a worker. An agent therefore advertises only the services it happens to be colocated with, which in a Swarm cluster means one agent per node, onetailscaledstate per node, and an N-node auth key for a single logical sidecar. That's the workaround described in #43.This adds a second discovery source that reads the whole cluster from one agent.
How it works
/servicesis manager-only and cluster-wide. Each service carries:Spec.TaskTemplate.ContainerSpec.Labels— where bothlabels:anddeploy.labels:from a compose file land, so a service is labelled exactly like the equivalent standalone container.Endpoint.VirtualIPs— the VIP the routing mesh keeps alive and load-balanced across that service's tasks, on the overlay the tasks are attached to. A VIP is reachable from any node on that overlay, so one agent covers the cluster as long as the overlay carries the traffic.The service VIP stands in for the container IP, which means the label surface is reused unchanged:
parseContainerServicesis factored intoparseFromInspectso both sources run the same translation, protocol defaults, path handling, tags, funnel config and validation.What this changes
Opt-in.
DOCKTAIL_DISCOVERY=containers(the default) is byte-for-byte the current behaviour.Operational notes
/servicesis manager-only; on a worker-only agent the reconcile fails and nothing is advertised.tailscaledit configures) must be on it too. When a service sits on several networks, DockTail prefers one it is also attached to;DOCKTAIL_SWARM_NETWORKpins the choice.vipmode. Adnsrrservice, or one with no network, has none and is reported and skipped rather than silently dropped.RECONCILE_INTERVAL. Docker event watching is still local, so a service created or relabelled on another node is picked up on the next tick rather than immediately.Field results
Six services across two nodes, from one agent, with
tailscaleas a shared overlay:Every container ran on
cloud-spain; the agent ran oncloud. All six resolve over the tailnet:Built and run as a multi-arch image (
linux/amd64,linux/arm64) for that deployment.A caveat worth documenting
This surfaced a silent failure mode worth calling out in the existing docs, since it cost us an afternoon and presents as "the agent advertises, but nothing answers":
Tailscale Services do not work when
tailscaledruns in userspace mode.containerbootsilently falls back to--tun=userspace-networkingwhen it cannot find/dev/net/tun, so a compose file that mounts the tun device at the wrong path — or omitsTS_USERSPACE=false— produces a node that authenticates, advertisessvc:*happily, and answers nothing on the VIP. The service config and the control plane both look correct, which makes it very hard to spot from the outside.I have not included a docs change for this since it's adjacent to the feature rather than part of it — but it's a good candidate for a follow-up (or a startup check that logs loudly when it detects userspace mode).
Testing
docker/swarm_test.gocovering CIDR stripping, numeric-vs-lexicographic VIP ordering, label extraction, network-name resolution, host-mode detection, and all four VIP-selection paths (no VIP, lowest, shared-network preference, operator-pinned).go build ./...,go vet ./...,go test ./...all pass; no existing tests changed.Review notes
The one design decision worth a second opinion:
swarmNetworkNameusesNetworkInspect, because the alias in aNetworkAttachmentConfigis the DNS name the service gets on that network, not the network's own name, so it cannot answerDOCKTAIL_SWARM_NETWORK. Results are cached in async.Mapon theClientsince network names don't move under a running cluster. An alternative would be to compare againstEndpoint.VirtualIPspositionally, but that couples the preference to ordering that isn't guaranteed stable.cc #43