Half the network tuning advice out there opens with the same sentence: "set
net.ipv4.tcp_adv_win_scale to this." I treated that setting as a real knob for years too. The
value is readable, writable, its range is documented, its default is documented. File mode
0644. Everything a sysadmin needs in order to think "there's a lever here" is in place.
Today I wrote ten different values to that lever. The window the receiver advertised came
out to 31,856 bytes all ten times. Byte for byte identical. When I wrote minus thirty-one,
and when I wrote plus thirty-one.
The reason is simple and slightly funny: the kernel no longer reads this value. It stopped
reading it with 6.6, released on 29 October 2023. The interface stayed, the meaning left —
picture a doorbell whose wire has been cut but whose button is still firmly mounted and still
labelled "press."
That raised the question I actually cared about: if the knob is dead, who decides the window
now? The answer is that we moved from a tunable constant to a measured ratio. And that ratio
moves the advertised window between 625 KB and 1 MB on the very same buffer size. So the lever
being broken does not mean the window is pinned. Quite the opposite.
Building the lab on a live server
I ran the measurement on VPS3 — a machine that actually hosts all of my projects. Normally I
don't play with sysctl on a box like that. The only reason I could here is that
net.ipv4.* settings are per network namespace: a value you write inside a network
namespace is none of the host's business.
I tested that instead of assuming it, because I had previously watched a privileged container
silently overwrite the host's fs.file-max:
# write 7 inside the netns, then look at the host
ip netns exec winlab sh -c 'echo 7 > /proc/sys/net/ipv4/tcp_adv_win_scale; cat /proc/sys/net/ipv4/tcp_adv_win_scale'
# -> 7
cat /proc/sys/net/ipv4/tcp_adv_win_scale
# -> 1
The host stayed at 1. The setup from there on: two separate namespaces (winA, winB), a
veth pair between them, MTU 1500. On the receiving side I turned automatic buffer tuning off
(tcp_moderate_rcvbuf=0) and pinned tcp_rmem; otherwise the reading would
have reflected the
noise of autotuning rather than the effect of the setting.
For the measurement point I picked the window field in the SYN-ACK packet. The reason: that
value is the output of a single function and depends on no traffic history. Reproducible,
identical down to the last digit. I also tried a single glance at ss; that route gave
me a different number on every run and cost me my first lab design. You cannot build a thesis
on a noisy measurement.
Ten values, one window
The documented range of tcp_adv_win_scale is [-31, 31]. I swept the extremes and the
typical values in between, established the connection from scratch for each value, and read
the SYN-ACK window with tcpdump:
tcp_adv_win_scale |
advertised window | advmss |
rcvbuf |
|---|---|---|---|
| −31 | 31,856 | 1448 | 65,536 |
| −8 | 31,856 | 1448 | 65,536 |
| −2 | 31,856 | 1448 | 65,536 |
| −1 | 31,856 | 1448 | 65,536 |
| 0 | 31,856 | 1448 | 65,536 |
| 1 (default) | 31,856 | 1448 | 65,536 |
| 2 | 31,856 | 1448 | 65,536 |
| 4 | 31,856 | 1448 | 65,536 |
| 8 | 31,856 | 1448 | 65,536 |
| 31 | 31,856 | 1448 | 65,536 |
Not a single byte moved.
For comparison I worked the documented formula by hand. The docs say: count buffering overhead
as bytes/2^tcp_adv_win_scale (if the value is positive). For tcp_adv_win_scale=4 and a
64 KiB buffer that means 65,536 − 4,096 = 61,440 bytes of payload; rounded to advmss, 60,816.
In other words, under the old behaviour writing 4 would have raised the window from 31,856 to
60,816 — 1.91 times bigger.
That number is not a measurement: I did not run a pre-6.6 kernel. 60,816
is the result of the formula in the documentation — a derivation, not a measurement. What I did
measure is that on today's kernel the difference is zero.
So who decides the window?
I traced in the source where the value goes. These are the places
sysctl_tcp_adv_win_scale appears in the kernel tree:
include/net/netns/ipv4.h:182 int sysctl_tcp_adv_win_scale; /* obsolete */
net/ipv4/sysctl_net_ipv4.c:26 static int tcp_adv_win_scale_min = -31;
net/ipv4/sysctl_net_ipv4.c:27 static int tcp_adv_win_scale_max = 31;
net/ipv4/sysctl_net_ipv4.c:1203 .procname = "tcp_adv_win_scale",
net/ipv4/sysctl_net_ipv4.c:1204 .data = &init_net.ipv4.sysctl_tcp_adv_win_scale,
net/ipv4/sysctl_net_ipv4.c:1208 .extra1 = &tcp_adv_win_scale_min,
net/ipv4/sysctl_net_ipv4.c:1209 .extra2 = &tcp_adv_win_scale_max,
The age of those lines is part of the story too. I looked at 2.6.12 (2005), where git
history begins: tcp_adv_win_scale is there as well, with the same .mode = 0644. The
setting has carried the same permission for over twenty years; since October 2023 it has done
nothing at all.
One thing is missing from that list: the line that reads the field. All seven hits are
either the declaration or the sysctl table entry. The value gets written, stored and
range-checked, and is used on no decision path whatsoever. The kernel itself put an
/* obsolete */ comment next to the field. The official documentation is blunter:
"Obsolete since linux-6.6".
The replacement landed with Eric Dumazet's 2023 patch dfa2f0483360. The patch's reasoning is
more convincing than my measurement: modern NIC drivers allocate a full page for every frame
they receive. A single global setting cannot make the right guess in that world. In Dumazet's
own words, on hosts dealing with various MSS values the window was either under-estimated or
over-estimated.
The fix was to turn the constant into a per-socket measured ratio:
/* include/linux/tcp.h */
#define TCP_RMEM_TO_WIN_SCALE 8
/* ... and inside struct tcp_sock: */
u8 scaling_ratio; /* see tcp_win_from_space() */
/* include/net/tcp.h */
static inline int __tcp_win_from_space(u8 scaling_ratio, int space)
{
s64 scaled_space = (s64)space * scaling_ratio;
return scaled_space >> TCP_RMEM_TO_WIN_SCALE;
}
The ratio is a u8, held in 256ths of fixed point. As of 6.6 this is where it got updated:
/* net/ipv4/tcp_input.c — tcp_measure_rcv_mss(), as of dfa2f0483360 */
if (unlikely(len != icsk->icsk_ack.rcv_mss)) {
u64 val = (u64)skb->len << TCP_RMEM_TO_WIN_SCALE;
do_div(val, skb->truesize);
tcp_sk(sk)->scaling_ratio = val ? val : 1;
}
Three more lines were added to this block in 2024; I will come to them, with their reason,
shortly.
So the window is no longer "the number the administrator wrote." It is the ratio of the
actual arriving packet's useful payload to the memory that packet occupies. skb->len
divided by skb->truesize.
The ratio really does move
At this point there was something I had to be suspicious of: maybe the mechanism changed but in
practice always produces the same number. Then the story would be "the knob died, nobody
noticed." I tested it.
I pinned the receive buffer at 1 MiB, left tcp_adv_win_scale at 1, and changed only the
shape of the data: GRO on the receiver and TSO/GSO on the sender either both on or both
off; MTU 1500 or 9000. Then I read
the peak of snd_wnd on the sender — that is the window the receiver advertises.
| condition | advertised window | derived ratio | share of buffer |
|---|---|---|---|
| MTU 1500, GRO/GSO off | 625,248 | 153/256 | 59.6% |
| MTU 1500, GRO/GSO on | 919,264 | 224/256 | 87.7% |
| MTU 9000, GRO/GSO off | 937,248 | 229/256 | 89.4% |
| MTU 9000, GRO/GSO on | 1,003,040 | 245/256 | 95.7% |
Same buffer. Same sysctl. 377,792 bytes of spread, 1.60 times. The knob being broken did
not pin the window; it took the window out of the administrator's hands and gave it to the
packet.
The provenance of the "derived ratio" column: there is no interface that exposes
scaling_ratio to userspace. I derived these numbers by dividing the advertised window by the
buffer — (window << 8) / rcvbuf, rounded to the nearest integer (the raw values are 152.65 /
224.43 / 228.82 / 244.88). The percentage column is window over buffer directly, so it can drift
from the ratio by about 0.2 points. The window is the measured quantity; the ratio is my arithmetic.
I was able to test one intermediate value the same way. The default in the source is 1 << 7,
that is 128, that is 50%. If that holds, the advertised window on a 64 KiB buffer should be
rounddown(32768, 1448). The arithmetic gives 1448 × 22 = 31,856. The number on the wire is
also 31,856. It held across four different buffer sizes:
tcp_rmem default |
advertised window | multiple of advmss
|
|---|---|---|
| 32,768 | 15,928 | 1448 × 11 |
| 65,536 | 31,856 | 1448 × 22 |
| 131,072 | 65,160 | 1448 × 45 |
| 262,144 | 65,535 | 16-bit ceiling |
The last row is not the mechanism, it is the packet format: the window field in the SYN-ACK is
unscaled 16 bits. The real window grows with the scale factor; that field does not.
This table has a side effect I enjoy. Describing the tcp_rmem default, the official docs say
"This value results in initial window of 65535." The number on the wire is 65,160. So like the
setting's own entry, the worked example in the documentation was never updated for the world
after 6.6.
The default changed twice, and the reason was an application outage
This is not an abstract cleanup job, and it is where the story turns.
With dfa2f0483360 the initial ratio corresponded to roughly a quarter of the buffer. Nine
months later Hechao Li from Netflix sent a patch: on an application that sets SO_RCVBUF to
64k, a 10 MB transfer had gone from 22 seconds to 40 seconds, and the application's
30-second timeout did not survive it. The patch's reasoning names the case too: applications
like Kafka that default to SO_RCVBUF=64k.
The fix is 697a6c8cec03 — the initial ratio went from 25% to 50%:
-/* Assume a conservative default of 1200 bytes of payload per 4K page.
+/* Assume a 50% default for skb->len/skb->truesize ratio.
* This may be adjusted later in tcp_measure_rcv_mss().
*/
-#define TCP_DEFAULT_SCALING_RATIO ((1200 << TCP_RMEM_TO_WIN_SCALE) / \
- SKB_TRUESIZE(4096))
+#define TCP_DEFAULT_SCALING_RATIO (1 << (TCP_RMEM_TO_WIN_SCALE - 1))
In the patch's own words, the goal is to be "backward compatible with the original default
sysctl_tcp_adv_win_scale for applications setting SO_RCVBUF". In other words, they put the
dead knob's default behaviour back — as a hard-coded number.
I did not want to guess the release, so I pinned it from the tags: v6.9's tcp.h carries the
old definition, v6.10's carries the new one. So the change shipped in 6.10.
Here a surprise came up. The machine under test runs 6.8, which is older than 6.10. Yet the
measurement showed exactly the 50% behaviour. Instead of guessing, I looked at the running
kernel's own header:
/usr/src/linux-headers-6.8.0-142/include/net/tcp.h:1517-1520
/* Assume a 50% default for skb->len/skb->truesize ratio.
* This may be adjusted later in tcp_measure_rcv_mss().
*/
#define TCP_DEFAULT_SCALING_RATIO (1 << (TCP_RMEM_TO_WIN_SCALE - 1))
Ubuntu backported the patch into the 6.8 series. Measurement and source confirmed each other —
and it reminded me of something: the number "6.8" on a distribution kernel does not guarantee
upstream 6.8's behaviour. Look at the installed header, not the version string.
The people who actually used the lever got hurt
The second fix is subtler and arrived in two steps. On sockets that set SO_RCVBUF, autotuning
is switched off and the window_clamp computed at the start stayed fixed; even when the ratio
was later updated to the real one, the window could not climb above that first value.
05f76b2d634e fixed that within the SO_RCVBUF bound first, and then a2cbb1603943 moved the
work to the right place: whenever the ratio changes, window_clamp is recomputed. The three
lines added to the block above are exactly this:
/* net/ipv4/tcp_input.c — the lines added by a2cbb1603943 */
u8 old_ratio = tcp_sk(sk)->scaling_ratio;
/* ... after the ratio is computed ... */
if (old_ratio != tcp_sk(sk)->scaling_ratio)
WRITE_ONCE(tcp_sk(sk)->window_clamp,
tcp_win_from_space(sk, sk->sk_rcvbuf));
The telling sentence is in the reasoning of the first one, 05f76b2d634e: systems that had set
tcp_adv_win_scale to a value other than the default on older kernels ended up seeing
reduced download speeds in certain cases. The same text also hands over a number — it says the
real skb->len/skb->truesize ratio turns out to be about 0.66 — which is the same
neighbourhood as the 190/256 = 0.74 I derive from my own measurement further down.
The group that got hurt is not the people who never touched the setting — it is the people who
actually tuned it. And because the sysctl in their hands was still writable yet no longer
effective, the familiar reflex ("put the value back up") did nothing at all.
This side went on the bench too. Adding 20 ms of round-trip delay (netem, 10 ms each way),
then transferring 10 MiB:
| receiver | tcp_adv_win_scale |
advertised window | time |
|---|---|---|---|
SO_RCVBUF=64k |
1 | 97,196 | 3.41 s |
SO_RCVBUF=64k |
4 | 97,196 | 3.91 s |
| unset (autotuning) | 1 | 2,667,008 | 0.98 s |
The window is identical again: byte for byte 97,196 on both SO_RCVBUF runs. I am not
attributing the 3.41 versus 3.91 difference in the time column to the setting; in a netem lab
the time varies from run to run, and the clean signal is the window column.
97,196 is itself a nice confirmation. If we assume the ratio is 190/256, the formula gives
131,072 × 190 >> 8 = 97,280; the number on the wire is 84 bytes below that. The gap comes from
rcv_ssthresh moving in MSS steps as it grows. So a2cbb1603943 is backported into our 6.8 as
well: the socket with SO_RCVBUF set did not stay stuck at 50% but reached the measured ratio.
The last row carries the actual message: a receiver that sets nothing climbs to a 2.6 MB
window, a receiver that writes SO_RCVBUF=64k stays at 97 KB. The "knob" that decides the
window today is whether the application calls setsockopt. There is no such lever in sysctl.
The ratio runs in both directions
There is a symmetry in the source that I had overlooked. tcp_win_from_space() has an inverse
defined as well:
/* inverse of __tcp_win_from_space() */
static inline int __tcp_space_from_win(u8 scaling_ratio, int win)
{
u64 val = (u64)win << TCP_RMEM_TO_WIN_SCALE;
do_div(val, scaling_ratio);
return val;
}
Automatic buffer tuning uses this direction: "I want this much window, so how many bytes of
buffer do I need to allocate?" Which means the ratio determines not only the advertised window
but also the memory the kernel will set aside for that window.
If the shape of the data is bad — small packets,
high truesize overhead — the ratio drops and the kernel allocates more memory to hold the
same window. In the old model that overhead came from a single global guess, and when the guess
was wrong it either squeezed the window or inflated the memory. Dumazet's claim that "this
patch alone can double TCP receive performance" comes from exactly here: the right ratio means
a bigger window for the same memory.
How to look at your own connection
The ratio cannot be read, but its consequences can. These are the three points I was left with:
# 1) On the receiver: how much of the buffer we can turn into window
ss -tinm state established | grep -A1 ':443' | grep -o 'rcv_ssthresh:[0-9]*\|skmem:([^)]*)'
# 2) On the sender: the window the peer ADVERTISES
ss -tin dst 10.0.0.5 | grep -o 'snd_wnd:[0-9]*'
# 3) In the handshake: the raw window field, tied to no history
tcpdump -ni eth0 -S -c 1 'tcp[tcpflags] & (tcp-syn|tcp-ack) == (tcp-syn|tcp-ack)'
The third one became the point I trust most, because it is reproducible. The first two change
over the life of the flow and it is easy to reach a wrong conclusion from a single read — I
did. Reading snd_wnd as a peak on the sender does work, but you have to do it as "sampling
across the flow," not "a glance."
Quantization: the window that disappears with jumbo frames
There is one more detail at the end of the road. The tcp_select_initial_window() arithmetic
rounds down to a multiple of MSS:
space = min(*window_clamp, space);
/* Quantize space offering to a multiple of mss if possible. */
if (space > mss)
space = rounddown(space, mss);
At small MSS this loss is negligible. At jumbo frames it is not. With the same 64 KiB buffer I
changed only the MTU:
| MTU | advmss |
advertised window | share of 32,768 |
|---|---|---|---|
| 1280 | 1228 | 31,928 | 97.4% |
| 1500 | 1448 | 31,856 | 97.2% |
| 4000 | 3948 | 31,584 | 96.4% |
| 9000 | 8948 | 26,844 | 81.9% |
At MTU 9000, 18.1% of the initial window goes to rounding alone: 32,768 − 26,844 =
5,924 bytes, two thirds of one MSS. By definition the rounding loss always stays below one
MSS; what hurts with jumbo is that the MSS itself is large.
Note that this loss belongs to the initial window: with autotuning on, the buffer grows, and
in table D at MTU 9000 the window reached 95.7% of 1 MiB. The penalty becomes permanent only in
setups that pin the buffer small.
A subordinate clause: the ratio is not updated on every packet
The easiest place to miss while reading the source is this condition:
if (unlikely(len != icsk->icsk_ack.rcv_mss)) {
The ratio is updated when the measured rcv_mss changes, not with every packet. On a
transfer that flows at a steady full MSS, the ratio settles once and stays there. This is not a
bug but a deliberate design — division is expensive and the code's comment says so. But it has
this practical consequence: the shape of a connection's first packets determines that
connection's window ratio for a long time. The assumption that "it will sort itself out as the
connection warms up" does not automatically hold here.
So what should you do
I checked my own fleet too: none of the 52 active lines under VPS3's /etc/sysctl.conf,
/etc/sysctl.d/, /usr/lib/sysctl.d/ and /run/sysctl.d/ contains tcp_adv_win_scale. I am
lucky; I didn't do that on purpose. Plenty of setup guides still recommend that line.
Questions to ask of your own machines:
-
Is your kernel 6.6 or newer?
uname -ris enough. If it is, the only effect of carrying the setting is to send the next sysadmin looking in the wrong place. -
Is the setting still in your config files?
grep -r tcp_adv_win_scale /etc/sysctl.conf /etc/sysctl.d/. If it is, delete it; it is not "harmless," it is misleading. -
Were you setting it to a non-default value? This is the real risk group. The slowdown you
measured after an upgrade may be this, and writing the sysctl back will not fix it — opening
up the
tcp_rmemceiling will. -
Does your application call
SO_RCVBUF? That is the biggest lever on the window today. If it does, it is switching autotuning off; take one more look at whether you really need it. -
What is your GRO/GSO state? This is the biggest working lever I found in this article:
on the same buffer, 625,248 versus 919,264 — more than a quarter of the window.
Check with
ethtool -k <iface> | grep -E 'generic-receive-offload|generic-segmentation'— setups that disable GRO for tunnels, packet capture or bridging are giving away window without knowing it. -
Are you using jumbo frames? Pull the
tcp_rmemdefault to a comfortable place next to the MSS (the BDP arithmetic is above), or rounding will take it out of your pocket.
One extra note on the container side: because the setting is per network namespace, every
container has its own copy. I tested that too:
docker run --rm --sysctl net.ipv4.tcp_adv_win_scale=7 ubuntu:24.04 \
cat /proc/sys/net/ipv4/tcp_adv_win_scale
# -> 7 (written inside the container)
cat /proc/sys/net/ipv4/tcp_adv_win_scale
# -> 1 (host untouched)
The command passes without error, the container sees 7, and that 7 is read on no decision
path. So that line in your Compose files is an argument that is silently accepted and does
nothing. The orchestrator cannot validate it, because the kernel still allows the write. This
is exactly the kind of mistake I like least: one that no tool tells you is wrong, and that you
only see by measuring. tcp_rmem, on the other hand, genuinely works in a container namespace
— if you are looking for a lever, it's there.
The second of those two levers, tcp_moderate_rcvbuf, went on the bench before it went into the
recommendation — no point leaving advice standing without showing that it works. Same tcp_rmem, no
SO_RCVBUF, 20 ms RTT, median of three runs each:
tcp_moderate_rcvbuf |
advertised window (median) | time for 10 MiB (median) |
|---|---|---|
| 0 | 97,280 | 3.23 s |
| 1 | 1,197,824 | 0.36 s |
The window is 12.3 times larger. The time column wobbles again (1.33 / 0.36 / 0.33), but the
direction is beyond argument.
This table also closed one open question. With autotuning off the window came out exactly
97,280 — that is 131,072 × 190 >> 8. So the ratio of 190 is now confirmed independently of the
earlier run, where I had only assumed it.
The answer to "how far should I open it" is not in sysctl but in arithmetic: the buffer's target
is the bandwidth-delay product. For 100 Mbit/s and 40 ms that is 100 × 0.04 = 4 Mbit, i.e.
500 KB; if
the window is below that, the window is your bottleneck, and if it is above, look elsewhere. The
third field of tcp_rmem should be a few times that number, because only a fraction of it turns
into window, and the ratio decides that fraction.
Use whichever lever actually works: tcp_rmem and tcp_moderate_rcvbuf. The window is now the
product of those two and the shape of the packet.
What I did not prove
This measurement has limits, and the right place to write them is next to the numbers.
I did not run a pre-6.6 kernel. The 60,816 figure given for tcp_adv_win_scale=4 was derived
from the formula in the documentation, not measured. I did not read any of the scaling_ratio
values — there is no interface that exposes them to userspace — I derived all of them from the
window-to-buffer ratio; those are the columns labelled "derived" in this article.
The measurement was done over a veth pair and netem, not on a real NIC. So the absolute
throughput numbers here do not count as production numbers; what is meaningful is whether the
window moves or does not move under the same conditions. skb->truesize behaviour will differ
with real drivers — which is precisely where the patch's reasoning came from.
And one more: the 190/256 ratio started as an assumption; it was confirmed when the window with
autotuning off landed on exactly 97,280. But that is one path on one machine. Another driver
will produce another number.
In one sentence
The reason I find this story interesting is not a removed setting. The kernel took a policy
out of the administrator's hands here and handed it to measurement — and in doing so it did
better: the window now fits the machine's real packet shape. But the interface was left behind:
the file is still 0644, the documentation still states its default, the range check still
works. A writable sysctl still reads like a promise.
A setting being readable does not mean it is read. Learning that through a number sticks better
than reading it in a man page — for me, 31,856, written ten times and never changing, did the
job. It is worth asking on your own server: how many of the lines in your sysctl.conf still
do anything?
All of this is the receive side. On the send side I previously measured
the 487 KB that stays in the kernel
— two ends of the same problem.
If you are curious about how a socket closes, I previously measured
who inherits the queue of a closing socket
— that article is another face of the same lesson: the kernel's accounting is more detailed
than the model in our heads.
Official Sources
- IP Sysctl — tcp_adv_win_scale and tcp_rmem
- tcp: get rid of sysctl_tcp_adv_win_scale (dfa2f0483360)
- tcp: increase the default TCP scaling ratio (697a6c8cec03)
- tcp: Adjust clamping window for applications specifying SO_RCVBUF (05f76b2d634e)
- tcp: Update window clamping condition (a2cbb1603943)
- include/net/tcp.h — tcp_win_from_space and TCP_DEFAULT_SCALING_RATIO
- RFC 7323 — TCP Extensions for High Performance (Window Scale)
Top comments (0)