The favourite use of SO_REUSEPORT goes like this: N worker processes listen on the same port, the kernel spreads incoming connections among them, and when you roll out a new version you replace the workers one at a time. The new worker comes up, the old one closes, the port is never idle. Wherever it is described, it sounds like zero downtime.
It is not. The connections waiting in the closing worker's accept() queue — connections that have finished their handshake and that the application has not yet picked up — die with that socket. On the client side this is an RST; on the server side it is nothing at all. It never reaches the access log, because the application never saw that connection.
The kernel has a switch for this: net.ipv4.tcp_migrate_req. It has been there since Linux 5.14, it defaults to off, and its documentation drifts from the code in at least two places. So I sat down and poked at it.
A note on the terminal output below: the measurement harness is mine and prints Turkish labels.
B kabul ettimeans "B accepted",gocenmeans "migrated",dagilimmeans "distribution",istemci kaybimeans "clients lost", andsunucu tarafimeans "server side". The numbers and labels are the harness's own; where a block is long I trimmed trailing fields, but nothing was retyped or translated.
The lab, and an honest disclaimer
Let me say this first: this is not an outage story. I did not hit this problem on my own server, because there is no reuseport group there. Here is the reading:
$ ssh vps3 'ss -ltnH | awk "{print \$4}" | sort | uniq -c | awk "\$1>1"'
$ ssh vps3 'ss -ltnH | wc -l'
90
Ninety listening sockets, and not a single address:port pair shared by two of them. nginx 1.30.0 is running, but neither reuseport nor backlog appears anywhere in its configuration — one shared socket, the default queue of 511. tcp_migrate_req is off too, and both counters sit at zero:
$ ssh vps3 'uname -r; cat /proc/sys/net/ipv4/tcp_migrate_req; nstat -az | grep -i migrate'
6.8.0-142-generic
0
TcpExtTCPMigrateReqSuccess 0 0.0
TcpExtTCPMigrateReqFailure 0 0.0
Rather than flipping this on a live server to see what happens, I took it to the lab. The environment is Docker Desktop's linuxkit VM:
$ uname -r
6.10.14-linuxkit
$ python3 -V
Python 3.12.3
$ ss -V
ss utility, iproute2-6.1.0
Using containers has a technical reason: tcp_migrate_req is network-namespaced (net->ipv4.sysctl_tcp_migrate_req), so docker run --sysctl sets it per container and leaves the host's value untouched. I could run one round off and one round on with nothing leaking between them. The harness is two Python scripts, 311 lines together: they open listeners with SO_REUSEPORT, establish client connections, count who accepts what, and read counter deltas from /proc/net/netstat.
First measurement: is the loss real?
The setup is simple. Listener A is up alone with a backlog of 128. Twenty clients connect, and each one sends a few bytes as soon as the connection is established. A accepts none of them; all twenty sit in the queue. Then listener B joins the same port, and A closes. B tries to drain the queue.
Before A closes, ss shows the state clearly — the Recv-Q column is the number of connections waiting in the queue:
LISTEN 20 128 127.0.0.1:18081 0.0.0.0:*
Here is the result with both settings side by side:
tcp_migrate_req=0 |
tcp_migrate_req=1 |
|
|---|---|---|
| Accepted by B | 0 / 20 | 20 / 20 |
| Clients that got their data back | 0 | 20 |
| Clients reset | 20 | 0 |
TCPMigrateReqSuccess |
0 | 20 |
TCPMigrateReqFailure |
0 | 0 |
A small caveat on that measurement: my harness drops both ECONNRESET and a clean end-of-file read into the same "reset" bucket, so it does not prove an RST at packet level. The source makes clear it is an RST — the inet_child_forget path leads to tcp_send_active_reset — but I know that from the code, not from my own measurement.
With the switch off, all twenty connections were gone. With it on, all twenty moved to B — and note this: the data the clients sent before the migration moved with them, so B was able to answer all twenty requests. The socket is not handed over, it is moved, receive buffer included.
There is one more thing in that table, and it is the part that bothered me: in the "off" round, twenty connections died while both counters stayed at zero.
The mechanism: when does the socket leave the group
To understand the silent counter I had to read the close path. __tcp_close runs a special branch for listening sockets, and the order is this:
if (sk->sk_state == TCP_LISTEN) {
tcp_set_state(sk, TCP_CLOSE);
/* Special case. */
inet_csk_listen_stop(sk);
So the state change comes before the queue is drained. For TCP_CLOSE, tcp_set_state removes the socket from the hash table (sk->sk_prot->unhash(sk)), which calls reuseport_stop_listen_sock() by way of inet_unhash. That is where the fork in the road is:
if (READ_ONCE(sock_net(sk)->ipv4.sysctl_tcp_migrate_req) ||
(prog && prog->expected_attach_type == BPF_SK_REUSEPORT_SELECT_OR_MIGRATE)) {
/* Migration capable, move sk from the listening section
* to the closed section.
*/
If the switch is on, the closing socket is not thrown out of the group; it moves from the group's "listening" section to its "closed" section, and the sk->sk_reuseport_cb pointer survives. If the switch is off, it is, as the comment puts it, "detach immediately" — the pointer becomes NULL.
Then inet_csk_listen_stop calls reuseport_migrate_sock() for every child in the queue. The first thing that function does is look up the group:
reuse = rcu_dereference(sk->sk_reuseport_cb);
if (!reuse)
goto out;
The out label does not touch any counter. The label that increments is failure, and you can only reach it after the group has been found. So with the switch off, the kernel does not even attempt the migration — and because it does not attempt it, it does not count it.
To confirm this I closed a listener with no other member left in the group: ten connections in the queue, the only listener closed, switch on.
=== G: grupta BASKA dinleyici yokken kapatma (K=10) ===
MIB delta: {'TCPMigrateReqSuccess': 0, 'TCPMigrateReqFailure': 10}
istemci kaybi: 10/10
This is where the counter speaks. Running the same experiment with the switch off also killed ten connections, but the counters stayed at zero. The practical conclusion is easy to read backwards: TCPMigrateReqFailure is not a "how many connections did I lose" counter; it is a "I tried to migrate and found nowhere to go" counter. Looking at it with the feature disabled and concluding "zero, so I have no problem" is like checking that a closed door's bell is not ringing and deciding nobody is home.
Who picks the target: the docs say random, the code says hash
The kernel documentation looks clear on this:
Otherwise, the kernel will randomly pick an alive listener only if this option is enabled.
There is no randomness in the source. reuseport_migrate_sock takes the migrating socket's own hash (hash = migrating_sk->sk_hash) and picks the target with it:
i = j = reciprocal_scale(hash, num_socks);
reciprocal_scale is a multiply-and-shift that stands in for a modulo: (u32)(((u64) val * ep_ro) >> 32). Same input, same output. Measuring that turned out to be easy — I bound the clients to fixed source ports and built the same 4-tuples twice, in the same namespace, against the same destination port:
tur 1: 41000->B3, 41001->B1, 41002->B3, 41003->B3, 41004->B3, 41005->B1, 41006->B4, 41007->B4
tur 1: gocen=8/8 dagilim=B1:2 B2:0 B3:4 B4:2
tur 2: 41000->B3, 41001->B1, 41002->B3, 41003->B3, 41004->B3, 41005->B1, 41006->B4, 41007->B4
tur 2: gocen=8/8 dagilim=B1:2 B2:0 B3:4 B4:2
ortak kaynak port: 8, AYNI hedefe gidenler: 8
Eight out of eight landed on the same listener in both rounds. You have to read "randomly" in the documentation as "unmanaged, without a policy"; if you read it as "statistically random", you will be wrong.
The second half of that output pulled me in the wrong direction for a while: the distribution is B1:2 B2:0 B3:4 B4:2, so one of the four listeners received none of the eight connections. My first reaction was "the hash does not spread evenly". Wrong reaction. For eight draws into four buckets that skew has a chi-square of 4.0, and the chance of seeing something that uneven or worse from a uniform hash is roughly 0.30 — indistinguishable from noise. Simulating the selection function itself (reciprocal_scale over uniform 32-bit input) shows the opposite at scale: with 8,000 draws across four buckets, the max-to-min ratio over 2,000 trials has a median of 1.045 and a worst case of 1.15. There is no two-to-one load gap.
The real issue is not the count but what the selection ignores. reuseport_select_sock_by_hash does not look at queue depth or current load when it picks a target; if some member of the group uses SO_INCOMING_CPU a CPU match enters the picture as well, and beyond that the only input is the connection's identity. So even when the number of connections evens out over time, their cost does not: a long-lived WebSocket and a 20-millisecond health check look identical to this selector.
That determinism is not specific to migration, either. Normal distribution came out the same way under the same harness — four listeners up from the start, no migration, fixed source ports — and eight out of eight landed on the same listener in both rounds. But the mapping changed between two separate containers, and that has a counterpart in the source as well: inet_ehashfn computes the hash using inet_ehash_secret + net_hash_mix(net). The mapping therefore carries a per-boot random secret and a per-namespace mixer. It is reproducible on your machine, and different on mine.
Migration does not respect the target's backlog
While reading the source I ran into something I was not looking for. The last step of a migration is inet_csk_reqsk_queue_add, and that function never asks about the target's queue capacity:
spin_lock(&queue->rskq_lock);
if (unlikely(sk->sk_state != TCP_LISTEN)) {
inet_child_forget(sk, req, child);
child = NULL;
} else {
...
sk_acceptq_added(sk);
}
The only check is "is the target still listening". There is no sk_acceptq_is_full() call. On the normal path, accepting a new connection does make that check; on the migration path it does not.
So it went to the lab. Twenty connections in A's queue, and B opened with listen(1) — that is, having told the kernel "I want exactly one pending connection on this socket":
B'nin ilan ettigi backlog: 1 | ss (A+B):
LISTEN 0 1 127.0.0.1:18085 0.0.0.0:*
LISTEN 20 128 127.0.0.1:18085 0.0.0.0:*
A kapandiktan SONRA, accept'ten ONCE ss:
LISTEN 20 1 127.0.0.1:18085 0.0.0.0:*
B kabul etti: 20/20 MIB={'TCPMigrateReqSuccess': 20, 'TCPMigrateReqFailure': 0}
Recv-Q 20, Send-Q 1. A socket that asked for one has twenty waiting — twenty times the limit it declared. All of them were accepted cleanly afterwards, so this is not an accounting glitch; the queue really is that full.
For most setups this is good news: migration does not give up halfway because the target's queue is tight. For a setup that deliberately keeps its queue short, the news is mixed — and here I had to correct an assumption of my own. I thought a short backlog was a fail-fast fuse: when the queue overflows the client fails quickly and the load balancer in front moves on to another server. Default Linux does not work that way. In the experiment in the next section, when the queue overflowed, not a single client got an error; the kernel dropped the SYNs silently and the clients retried. Setting tcp_abort_on_overflow to 1 and repeating the experiment changed nothing, because that setting only engages when there is a request socket to reset, and here no request socket was ever created.
What remains is a resource limit rather than a fuse: a low backlog bounds how many ready connections the application has to deal with at once, and migration does not honour that bound. If several workers close at the same time, the one still standing wakes up with a queue many times the limit it declared. I saw this in the source first, then measured it; I could not find it mentioned in the documentation.
Which connections recover on their own
To draw the boundary of the risk I ran one more experiment. I opened A with listen(1) and threw twelve clients at it using non-blocking connect(). The queue overflowed immediately, and ss told me what the server side actually looked like:
A'da accept kuyrugu (Recv-Q/Send-Q): LISTEN 2 1 127.0.0.1:19004 0.0.0.0:*
sunucu tarafi: SYN-RECV=0 ESTABLISHED=2
Two connections finished their handshake and entered the queue. The SYNs of the remaining ten were never accepted — and rather than guess at that, I read it off the counters: ListenOverflows and ListenDrops both climb while SYN-RECV stays at zero. The kernel did not even start building a request socket, because the queue was already full; it dropped the SYN.
Then I added B (backlog 128) and closed A:
tcp_migrate_req=0 |
tcp_migrate_req=1 |
|
|---|---|---|
| Accepted by B | 10 / 12 | 12 / 12 |
| Arrived by migration (instantly) | 0 | 2 |
| Arrived by SYN retransmission | 10 | 10 |
TCPMigrateReqSuccess |
0 | 2 |
What migration saved was exactly two connections: the ones that had finished the handshake and had not been accepted. The counter says two as well. The other ten survived under both settings, but not thanks to migration.
In the first round I waved at what saved them — "the client retries" — and that guess needed to become a measurement, because the ~0.86 seconds I saw read like a constant. Sampling ListenOverflows once per second made the mechanism plain: the counter rises by exactly 10 every second, for four seconds running. Each of the ten blocked clients retransmits its SYN once a second. Recovery then happens at the first retransmission after the new listener exists — so the delay depends on where the swap lands in that cycle. Shifting the swap moment gives 0.74, 0.86 and 0.58 seconds, and under 0.1 seconds when the swap coincides with a retransmission. Not a fixed delay budget: a wait that ranges from zero to one second.
So the set tcp_migrate_req protects is not "everything in flight at the moment of close". If the client is still trying, TCP does its own job. What cannot be recovered is the connection that the client considers established and is now waiting on, while the server has not yet picked it up. That also happens to be the most annoying failure mode on the user's side: request sent, no answer, connection reset. Whether it is safe to retry is something the client does not know.
One more boundary, for honesty's sake: requests still mid-handshake (TCP_NEW_SYN_RECV) have a separate migration path in the source, and that path runs not at close() but when the SYN+ACK retransmission timer fires (inside reqsk_timer_handler). I could not trigger that path in this experiment — the queue overflowed, so the requests were never created, and SYN-RECV=0. Its behaviour here is therefore read from the source, not measured.
Who is allowed into the group
For migration to be on the table at all, the new worker has to have joined the same reuseport group. That group's door has two locks; these outputs come from two separate runs, with the harness's own labels:
=== SENARYO D (port 18086) gruba katilma kapilari ===
SO_REUSEPORT'suz bind: OSError errno=98 (EADDRINUSE)
=== F: gruba farkli UID ile katilma ===
uid=65534 + SO_REUSEPORT bind: OSError errno=98 (EADDRINUSE)
uid=0 (ayni kullanici) bind: BASARILI
The first is expected. The second is the one that can catch you in production — and it should not really be a surprise, because socket(7) states it outright: "To prevent port hijacking, all of the processes binding to the same address must have the same effective UID." Being written in the manual is not the same as being remembered on deployment day. A procedure that starts the new version under a different service account cannot bind to that port at all; the migration question never even comes up. And EADDRINUSE will not tell you why — it says "port busy", not "different user".
The ordering inside the group is not arbitrary either. socket(7) defines the numbering: "Sockets are numbered in the order in which they are added to the group (that is, the order of bind(2) calls for UDP sockets or the order of listen(2) calls for TCP sockets)." That is the array the hash is scaled into. In other words, the order in which you start your workers is one of the inputs that decides which connection goes to which worker.
One more run, because the documentation mentions shutdown() alongside close(). I called shutdown(SHUT_RDWR) on the listening socket, leaving the file descriptor open:
A.shutdown(SHUT_RDWR) cagrildi, fd ACIK
B kabul etti: 10/10 MIB={'TCPMigrateReqSuccess': 10, 'TCPMigrateReqFailure': 0}
Ten out of ten migrated with the switch on, zero out of ten with it off. If you are writing a graceful drain, that is useful: you can stop listening and hand off the queue without closing the descriptor.
Before you turn it on
The part of the documentation that deserves the most attention is at the end, and it is a warning:
Note that migration between listeners with different settings may crash applications. Let's say migration happens from listener A to B, and only B has TCP_SAVE_SYN enabled. B cannot read SYN data from the requests migrated from A.
Migration assumes the two listeners have identical socket options. If they differ, the target ends up reading a connection it did not set up using its own assumptions. During a version rollout, where the old and new worker's socket options can drift apart, that is a real risk — and that is precisely the moment you would most want migration. The kernel's suggested answer is to pick the target with a BPF_SK_REUSEPORT_SELECT_OR_MIGRATE eBPF program and return SK_DROP to cancel the migration when no suitable target exists. Flipping the bare sysctl means leaving that policy to "whatever the hash says".
The decision framework I ended up with:
-
If several processes listen on the same port and you close the old one during a rollout, turn it on. nginx's
reuseportparameter does exactly that: per the documentation, "an individual listening socket for each worker process". A reload closes the old workers' sockets. -
If you have a single listening socket, nothing changes — but which kind of "single" you have shows up in the counters. If the socket was opened with
SO_REUSEPORTthere is still a group: migration is attempted, finds nowhere to go,Failurerises, and the connection dies anyway. IfSO_REUSEPORTis absent — which is the case for all 90 listeners on my own server — theif (rcu_access_pointer(sk->sk_reuseport_cb))gate insideinet_unhashnever opens, the migration code never runs, and both counters stay at zero. Either way the fix lives elsewhere: an approach like systemd socket activation, which leaves behind a socket that can be handed over. - If you deliberately keep your queue short, think twice. Migration does not honour that limit.
- If your workers' socket options are not identical, avoid the bare sysctl: use an eBPF policy, or do not turn it on at all.
Three commands are enough to check. First, do you even have a group:
ss -ltnH | awk '{print $4}' | sort | uniq -c | awk '$1>1'
Then the setting and the counters:
cat /proc/sys/net/ipv4/tcp_migrate_req
nstat -az | grep -i migrate
Both flags are needed, for different reasons: without -a, nstat prints the delta against its own history file (which is kept per user, so running it under sudo reads a different ledger), and without -z, counters sitting at zero drop out of the listing entirely — so if no migration has been attempted yet, you will not see the line at all. And as the runs above show, while the feature is off these counters will not report your losses.
Where the counter goes quiet
This is a long story for a single byte. Looking back, what stayed with me is not the setting but the two gaps I ran into while chasing it.
One is on the instrumentation side. With migration off, the kernel threw away twenty connections and not one counter moved — because that counter counts migrations that were attempted and failed, not the ones that were never attempted. A gauge reading zero does not mean the thing it measures is not happening; it does not even mean the gauge is connected. That is exactly why both counters sitting at zero on my own server tell me nothing.
The other is whose point of view "zero downtime" is spoken from. The port was never idle, there was no gap in the process table, there is no error line in nginx's access log. Looked at from the server, there really was no downtime. The downtime happened on the side of twenty clients that had finished their handshake, and that side does not write into our ledger. If the client is still retrying, TCP repairs itself; if it has stopped, we are the ones creating the loss. It helps to think of this setting not as a performance knob but as where you draw the line between those two.
A last note on currency. The feature arrived in 2021 with Linux 5.14 (f9ac779f881c, "net: Introduce net.ipv4.tcp_migrate_req", Kuniyuki Iwashima) and it still sits in today's mainline in the same place with the same default. It has not been removed, renamed or marked discouraged; it has simply been off by default for five years and rarely discussed. The kernel's own test suite has a case called migrate_reuseport.c — the right place to read if you want to see how they exercise the migration.
Official Sources
- Documentation/networking/ip-sysctl.rst — tcp_migrate_req (kernel.org)
- net/core/sock_reuseport.c — reuseport_migrate_sock and reuseport_stop_listen_sock (torvalds/linux)
- net/ipv4/inet_connection_sock.c — inet_csk_listen_stop and inet_csk_reqsk_queue_add (torvalds/linux)
- net/ipv4/proc.c — TCPMigrateReqSuccess and TCPMigrateReqFailure (torvalds/linux)
- socket(7) — SO_REUSEPORT and group numbering (kernel.org man-pages)
- ngx_http_core_module — the reuseport and backlog parameters of the listen directive (nginx.org)
- tools/testing/selftests/bpf/prog_tests/migrate_reuseport.c — the kernel's migration tests (torvalds/linux)
Top comments (0)