DEV Community

Alex Georgiev
Alex Georgiev

Posted on AI-assisted

Docker Engine 29.7's overlay networking breaks every Swarm task without IPv6

I asked a Swarm service for twenty replicas on an overlay network and got zero. Not slow, not partially up. Zero, forever, with the exact same command that gives me twenty on the engine version one release behind it.

I'd gone looking at the Docker Engine 29 release notes for something smaller: a line in the 29.8.0 changelog about spreading out the daemon's periodic overlay-network gossip so it doesn't burn CPU in bursts. That seemed like a nice, modest thing to measure. While setting up a test rig for it, before I'd measured anything about gossip at all, every overlay-attached service I created came up with 0 running tasks instead of the number I'd asked for.

The setup

I had Docker Engine 29.6.2 already running as this machine's system install. I downloaded the static builds for 29.7.0, 29.7.2 and 29.8.2 from download.docker.com and ran each as its own dockerd, pointed at its own data directory and its own Unix socket, so I could run two or more versions side by side on one host and compare them directly rather than trusting memory of "what it used to do."

Each one got docker swarm init on a single node, an overlay network, and a service with a handful of replicas running sleep 600, nothing fancier.

On 29.6.2:

$ docker service create --detach --name gossip-svc --replicas 20 --network gossip-net alpine:3.20 sleep 600
$ docker service ls
ID             NAME             MODE         REPLICAS   IMAGE         PORTS
e8v48kmvwvya   gossip-svc-old   replicated   20/20      alpine:3.20
Enter fullscreen mode Exit fullscreen mode

Twenty requested, twenty running. I exec'd into two of the containers and pinged across the overlay network to be sure it wasn't just lying about the count:

$ docker exec 17a672fb04a7 sh -c "ping -c3 -W2 10.0.1.16"
PING 10.0.1.16 (10.0.1.16): 56 data bytes
64 bytes from 10.0.1.16: seq=0 ttl=64 time=1.051 ms
64 bytes from 10.0.1.16: seq=1 ttl=64 time=0.066 ms
64 bytes from 10.0.1.16: seq=2 ttl=64 time=0.075 ms

--- 10.0.1.16 ping statistics ---
3 packets transmitted, 3 packets received, 0% packet loss
Enter fullscreen mode Exit fullscreen mode

Real traffic, real VXLAN tunnel, both containers reachable. Good baseline.

On 29.8.2, the identical sequence of commands, same host, same image:

$ docker service create --detach --name gossip-svc3 --replicas 5 --network gossip-net3 alpine:3.20 sleep 600
$ docker service ls
ID            NAME          MODE         REPLICAS   IMAGE         PORTS
iaki8w2lsuwr  gossip-svc3   replicated   0/5        alpine:3.20
Enter fullscreen mode Exit fullscreen mode

docker service ps named the reason:

network sandbox join failed: subnet sandbox join failed for "10.0.1.0/24":
overlay: cannot determine address family of transport: the local data-plane
address is not currently known
Enter fullscreen mode Exit fullscreen mode

Every one of the five tasks hit this. Not one came up.

It isn't a fluke or a port clash

My first instinct was that I'd done something wrong with the test harness itself — I'd given the two engines different VXLAN data-path ports to avoid a bind conflict, so I reran 29.8.2 on the default port 4789 with the older engine's overlay network torn down first, to rule that out. Same error, same 0/5.

Then I widened the test backwards through the point releases to find where it actually started, since I only had 29.6.2 and 29.8.2 at first:

Engine Replicas requested Replicas running Result
29.6.2 20 20 overlay traffic confirmed with ping
29.7.0 5 0 identical "address family" error
29.7.2 5 0 identical "address family" error
29.8.0–29.8.2 5 and 20 0 identical "address family" error

It's not version-specific to 29.8. It was already broken in 29.7.0, the release immediately after the one that works, and it's still broken in 29.8.2, the newest static build on download.docker.com as I write this. I didn't bisect down to a single commit — I don't have the moby source checked out here — but the official 29.7.0 and 29.8.0 release notes both list several Swarm networking changes in that stretch, including one about published ports on the service mesh sharing infrastructure with locally published ports, which touches exactly the code path that's failing here.

What it isn't

Plain containers on this same 29.8.2 install work fine. A docker run with a published port, no Swarm involved:

$ docker run -d --rm -p 18080:80 --name plainweb nginx:alpine
$ docker exec plainweb wget -qO- http://127.0.0.1:80 | head -3
<!DOCTYPE html>
<html>
<head>
Enter fullscreen mode Exit fullscreen mode

That rules out something broad like the bridge driver or the whole networking stack being down. This is specific to the Swarm overlay driver's VXLAN data plane.

The one thing I could find that correlates

This host has no IPv6 stack at all:

$ cat /proc/sys/net/ipv6/conf/all/disable_ipv6
cat: /proc/sys/net/ipv6/conf/all/disable_ipv6: No such file or directory
$ cat /proc/net/if_inet6
cat: /proc/net/if_inet6: No such file or directory
Enter fullscreen mode Exit fullscreen mode

Not disabled — absent. No IPv6 sysctls, no IPv6 socket table, nothing. Both 29.6.2 and every later engine log the same ipv6-related read failures at startup (failed to read ipv6 net.ipv6.conf.<bridge>.accept_ra), so the daemon itself already knows this host has none. The difference is what each version does with that fact once a container actually tries to join an overlay sandbox: 29.6.2 carries on and completes the join over IPv4; 29.7.0 onward asks something that comes back unanswered and calls it fatal.

I want to be careful here: I didn't read the overlay driver's source to confirm this is the actual cause rather than a correlated symptom. What I can say is that the error text names exactly this — "cannot determine address family of transport" — on the one property of this host that changed nothing between engine versions and that I can independently confirm is true.

I checked whether Docker's own documentation sets an expectation either way. Its swarm networking page lists the ports that need to be open between hosts and says nothing about IPv6 either requiring it or ruling it out:

Port 2377 TCP for communication with and between manager nodes
Port 7946 TCP/UDP for overlay network node discovery
Port 4789 UDP (configurable) for overlay network traffic

Three protocols, all IPv4-shaped, no mention of IPv6 as a prerequisite anywhere on that page. By that documentation, what I ran should have worked on 29.8.2 exactly as it did on 29.6.2.

The workaround that didn't work

docker network create takes an explicit --ipv6 flag, so I tried forcing it off on the overlay network itself, on the theory that the daemon might be tripping over IPv6 address assignment specifically and that turning it off per-network would route around that:

$ docker network create -d overlay --ipv6=false gossip-net3
$ docker service create --detach --replicas 5 --network gossip-net3 alpine:3.20 sleep 600
$ docker service ls
ID            NAME           MODE         REPLICAS   IMAGE
zwuxi4opkw5  gossip-svc3     replicated   0/5        alpine:3.20
Enter fullscreen mode Exit fullscreen mode

Identical failure, identical error text. Whatever is going wrong, it isn't happening at the per-network IPv6-enable flag, so there's no documented flag I found that gets you out of this on an IPv6-less host.

What it actually costs, which surprised me

I expected a service stuck retrying forever to be quietly expensive — some background loop hammering the scheduler. I sampled dockerd's own CPU time once a second for thirty seconds under two conditions: 29.6.2 running twenty real, working replicas, and 29.8.2 sitting on its broken five-replica service.

Condition Mean CPU Max CPU (1s sample)
29.6.2, 20/20 replicas running 0.52% 0.98%
29.8.2, 0/5 replicas, stuck 0.30% 1.97%

The broken service was cheaper on average, not more expensive. Watching docker service ps over time explained why: each task gets retried three times, each attempt a few seconds apart, and then Swarm stops trying and leaves it sitting in Assigned forever. No more log lines, no more CPU, no more anything. It fails hard once and then goes completely quiet. That's worse for anyone watching dashboards rather than logs, because nothing about resource usage tells you it's broken.

How you'd actually notice

One command shows it plainly, which is the only genuinely good news in this post:

$ docker service ls
ID            NAME          MODE         REPLICAS   IMAGE
iaki8w2lsuwr  gossip-svc3   replicated   0/5        alpine:3.20
Enter fullscreen mode Exit fullscreen mode

0/5, sitting there indefinitely, next to a service that's genuinely fine. If your deploy tooling checks docker service ls or the equivalent API field for convergence before calling a rollout successful, you'll catch this immediately. If it only checks that docker service create returned exit code 0 — which it does, every time, failure included — you won't.

A smaller thing I checked and it didn't hold up

While I was in there I also tried to reproduce a specific fix listed in the 29.8.0 notes, about service creation failing when an automatically generated name collides with an existing one. I fired thirty concurrent docker service create calls with no --name at both 29.6.2 and 29.8.2, hoping to force a collision in the random name generator. Zero collisions on either version, in thirty concurrent attempts each. The namespace is evidently large enough that this doesn't show up at a scale I can produce by hand in an hour. I'm noting it rather than dropping it, because a finding of "I couldn't reproduce this" is still useful if you were about to spend time worrying about it.

What I got wrong on the way

My first run of the concurrent service-creation test hung indefinitely. I'd used docker service create without --detach, and with --restart-condition none the task is supposed to exit after sleep 1 rather than stay running — which meant the CLI's default wait for "the service has converged to its desired running count" could never be satisfied, since the desired count for a task designed to exit is never going to match "currently running." It wasn't a Docker bug, it was me asking the CLI to wait for a state I'd specifically engineered never to arrive. Adding --detach fixed it in about ten seconds once I noticed what the command was actually doing.

Run it yourself

This needs root and a Linux host with Docker's static builds reachable from download.docker.com. It stands up two engines side by side on their own sockets so you can compare without touching your real Docker install.

# grab a second engine build to compare against your system one
curl -sA 'Mozilla/5.0' -o docker-29.8.2.tgz \
  https://download.docker.com/linux/static/stable/x86_64/docker-29.8.2.tgz
mkdir new && tar xzf docker-29.8.2.tgz -C new
mkdir -p /var/lib/docker-new /run/docker-new

nohup new/docker/dockerd \
  --data-root /var/lib/docker-new \
  --exec-root /run/docker-new/exec \
  --host unix:///run/docker-new/docker.sock \
  --exec-opt native.cgroupdriver=cgroupfs \
  > new.log 2>&1 &
sleep 5

ADDR=$(hostname -I | awk '{print $1}')
export DOCKER_HOST=unix:///run/docker-new/docker.sock
new/docker/docker swarm init --advertise-addr "$ADDR"
new/docker/docker network create -d overlay testnet
new/docker/docker service create --detach --name testsvc --replicas 5 \
  --network testnet alpine:3.20 sleep 600

sleep 15
new/docker/docker service ls
new/docker/docker service ps testsvc --no-trunc | head -5

# check whether this host even has IPv6 at all
cat /proc/net/if_inet6 2>&1 || echo "no IPv6 stack on this host"
Enter fullscreen mode Exit fullscreen mode

If service ls shows 0/5 and your host has no /proc/net/if_inet6, you've reproduced this. If it shows 5/5, either this has been fixed since 29.8.2 or your host has IPv6, and either way that's worth knowing before you plan around it.

If you run Swarm in production on a host or VM image that has IPv6 turned off — which is a common, deliberate choice on plenty of minimal server images and sandboxed CI runners — test an upgrade past 29.6.x on a throwaway node before you roll it out, and check docker service ls for the replica count rather than trusting a clean exit code from service create or service update.

Top comments (0)