DEV Community

Cover image for The TIME_WAIT Scare: The 60 Seconds tcp_fin_timeout Never Touches
Mustafa ERBAY
Mustafa ERBAY

Posted on Originally published at mustafaerbay.com.tr

The TIME_WAIT Scare: The 60 Seconds tcp_fin_timeout Never Touches

I'll open this one with a confession from my own server. The VPS that runs this blog and a handful of other projects has a file called /etc/sysctl.d/99-security.conf, and on line sixteen it says net.ipv4.tcp_fin_timeout = 15. I know exactly what the person who put it there was thinking, because years ago I thought the same thing: "TIME_WAIT is piling up, let me shorten it." The line came from a hardening template, nobody ever questioned it, and to this day it has accomplished precisely nothing. Because on Linux the TIME_WAIT duration is not a sysctl at all; it is a constant in the kernel source, written as TCP_TIMEWAIT_LEN (60*HZ). And tcp_fin_timeout, despite what its name suggests, watches an entirely different state: FIN_WAIT_2.

Here is my thesis. TIME_WAIT is not a symptom of a bug; it is a price the protocol pays on purpose. It cannot be shortened, and most of the time it doesn't need to be; the real bottleneck is never the socket count but the exhaustion of four-tuple address combinations toward a single destination. Every sysctl -w issued without that distinction is, at best, a harmless ornament like my line sixteen, and at worst a recipe copied off a forum that sets a knob deleted from the kernel in 2017 and knocks customers behind NAT offline.

I've covered the kernel's connection tracking on this blog before, from the angle of conntrack capacity planning and the SYN backlog; both of those were about how a connection is born. This one is about how it dies. Why it keeps living inside the kernel for a minute after it closes, on whose side it lives, and when that minute actually gets expensive.

279 of the 473 sockets on my server sit on one door

Before writing, I ran ss -s on the server. Kernel 6.8, Ubuntu, up for a day. The relevant line read:

TCP:   985 (estab 188, closed 738, orphaned 0, timewait 473)
Enter fullscreen mode Exit fullscreen mode

473 sockets in TIME_WAIT. A number that could look like "a problem" to someone, so I looked at who was carrying them. Grouping by local address simplified the picture instantly: 279 of them sat on 127.0.0.1:3399. That port is the loopback end of a Git service I bring in over a reverse SSH tunnel; nginx hands every request to it via proxy_pass. The far side closes the connection at the end of each request, and TIME_WAIT stays on the side that closes. So more than half the count was a harmless pile-up living on loopback, never touching the outside world.

The counters completed the story. With nstat -az, TcpExtTW was 514,924: half a million connections had passed through TIME_WAIT in a single day. TcpExtTWRecycled was 1,478, meaning the kernel, acting as the connecting side, had safely reused a TIME_WAIT socket for a new connection about fifteen hundred times. TcpExtTCPTimeWaitOverflow was zero. That last one is what matters: the tcp_max_tw_buckets limit (262,144 on this machine) was never exceeded, so the kernel never had to forcibly kill a connection before its wait was up.

A machine up for one day, half a million TIME_WAIT transitions, zero problems. I couldn't have designed a better example to show that TIME_WAIT itself is not the issue, and I didn't need to build a lab to find it. I just needed to measure.

Why TIME_WAIT exists, and why exactly one minute

RFC 9293, the current TCP standard, states the rule as a MUST: the side that actively closes a connection must linger in TIME-WAIT for 2×MSL (MUST-13). MSL is the longest a TCP segment can exist in the network, and the standard defines it as "arbitrarily" 2 minutes. On paper that means a four-minute wait, and old Windows versions applied it literally; the Windows 2000 documentation lists the TcpTimedWaitDelay default as 240 seconds. Linux has always kept it shorter: the TCP_TIMEWAIT_LEN constant in include/net/tcp.h is 60 seconds, and the tcp_time_wait() function in net/ipv4/tcp_minisocks.c assigns that constant directly to every socket entering TIME_WAIT. There is no line of code consulting a sysctl. The one exception is a RST from the peer: as long as the tcp_rfc1337 sysctl stays at its default of 0, the kernel kills a TIME_WAIT socket the moment a RST arrives, so the duration I called "unshortenable" can be cut, just not by you but by the other side.

The wait exists for two reasons, both about data integrity. First, if the final ACK is lost, the peer retransmits its FIN; a socket still in TIME_WAIT can answer it, otherwise the peer gets a RST. Second, and sneakier: if the same four-tuple (source IP, source port, destination IP, destination port) is reused immediately, a delayed segment from the old connection still wandering the network can be mistaken for data belonging to the new one. RFC 1337 described a subset of these accidents in 1992 under the title "TIME-WAIT assassination".

The cost of shortening TIME_WAIT is not a noisy crash but a silent corruption: a segment delivered to the wrong connection, a night nobody sees in the logs. I read the kernel developers' decision not to make this duration a sysctl as a stance, not an oversight.

The side that closes carries TIME_WAIT

The detail most discussions skip: TIME_WAIT forms on the side that closes first. If an HTTP server closes the connection itself after every response, TIME_WAIT accumulates on the server; if the client closes, on the client. In my port 3399 case the sockets sat on the local port 3399 side, which tells me the service at the far end of the tunnel finishes every request itself. Why it does, I didn't dig into for this article: when I looked with curl -D -, the far end doesn't put Connection: close on its responses, so I don't know which layer the close comes from. What I do know is only which side TIME_WAIT sits on, and that nginx cannot cache a connection that has been closed; nginx makes its keepalive decision based on the peer's response. Result: every Git request is a new TCP connection, and every connection is a minute of TIME_WAIT.

If I wanted to "fix" this, the right place would be the connection handling of the far end and of nginx, not sysctl. But I don't want to fix it. These sockets are on the server side; tcp_tw_reuse doesn't come into play here, because it only looks at the connecting side's own TIME_WAIT sockets. On the server side the kernel has a separate path: if a new SYN arrives at a socket in TIME_WAIT and its sequence number or timestamp is newer than the old connection's, the kernel ends TIME_WAIT and accepts the connection directly (MAY-2 in RFC 9293, the method RFC 6191 describes). Timestamps are on on this server, so that path works. Something being measurable does not mean it needs fixing.

The real bottleneck is the four-tuple, not the socket count

There is exactly one scenario where TIME_WAIT genuinely hurts: opening hundreds of short-lived connections per second from a single machine to a single destination (IP and port). A reverse proxy to one backend, an application server to a database without a pool, a collector to one API.

Linux's default ip_local_port_range is 32768–60999, which is 28,232 source ports. If the destination IP and port are fixed, the only variable that makes a connection unique is the source port; if each port waits a minute in TIME_WAIT, you can open at most 28,232 new connections per minute to the same destination, roughly 470 per second. Beyond that, connect() returns EADDRNOTAVAIL ("Cannot assign requested address"). In application logs this usually shows up only as "could not connect", and the diagnosis wanders through the wrong places until someone thinks of a one-minute waiting line. There is also the other side of the coin: if the server carries the TIME_WAIT, the client never sees EADDRNOTAVAIL; the new SYN is either accepted by the socket in TIME_WAIT or dropped, and the client retransmits the SYN and waits on the order of seconds. So the symptom depends on which side closes.

The good news here is that uniqueness is defined over the four-tuple. The same source port can be reused for a connection to a different destination; during connect() the kernel checks for collisions in the established table by four-tuple. So "I have 28 thousand ports, I can open 28 thousand connections" is wrong; the limit is per destination. A service talking to a hundred different destinations hits this wall very late, a proxy talking to a single backend very early.

The sysctl glossary

I wrote this section with the kernel's ip-sysctl documentation page open, reading line by line; it contradicted what I remembered in one place, and I point it out below. One note on scope: all of these settings are per network namespace; the value you see with sysctl inside a container can differ from the host's.

tcp_fin_timeout — The documentation is clear: how long an orphaned connection, one no application references anymore, stays in FIN_WAIT_2. Default 60 seconds. Nothing to do with TIME_WAIT. Dropping it to 15 buys you orphaned connections whose peer forgot to send a FIN being cleaned up 45 seconds earlier; it costs you the peer's late-arriving FIN now being answered with a RST. Line sixteen on my server does exactly that and does not shift the TIME_WAIT count by a single socket.

tcp_tw_reuse — In the documentation's words, it allows TIME_WAIT sockets to be reused for new connections "when it is safe from protocol viewpoint." Three values: 0 disabled, 1 enabled globally, 2 enabled for loopback traffic only. The default is 2; that value was added in 2018 and shipped in 4.18. My memory said "default 0", which was wrong. The doc adds right below that it "should not be changed without advice/request of technical experts." The mechanism lives in tcp_twsk_unique() in net/ipv4/tcp_ipv4.c: it only engages for outgoing connections (the connect path), it requires that the timestamp option was negotiated with the peer, and it waits for a certain time to have passed since the socket entered TIME_WAIT. The timestamp condition means: with tcp_timestamps off, tcp_tw_reuse does nothing. The kernel leans on PAWS; if an old segment arrives, its timestamp gives it away.

tcp_tw_reuse_delay — Until late 2024 that "certain time" above was a hard-coded one second. A December 2024 patch by Jakub Sitnicki of Cloudflare turned it into a sysctl in milliseconds; the change shipped in 6.14, default 1000 ms. The rationale is instructive: the one-second wait rests on the assumption that the peer's timestamp clock might tick as slowly as 1 Hz (the patch message grounds this in section 5.4 of RFC 7323); yet with short-lived connections over RTTs of a few milliseconds, the time a four-tuple stays blocked becomes hundreds of times the connection's lifetime. The same patch message carries an honest warning: applications can already bypass this protection entirely today by pinning the local port with bind() before connect(); in that case PAWS may fail to catch old segments, leaving the sequence number check as the only safety net. My 6.8 kernel doesn't have this sysctl yet; sysctl returned "No such file or directory". In short: uname -r before you try to set it on an older server.

tcp_tw_recycle — Gone. Deleted from the kernel in March 2017, released with 4.12. The removal commit message gives two reasons: it was already broken for clients behind NAT, because timestamps from multiple machines behind a single destination address don't increase monotonically; and after timestamp offsets were randomized per connection, it broke for every kind of connection. That's why the "set tw_reuse=1 and tw_recycle=1" recipes still in circulation are half dead: on a modern kernel the second line prints an error but sysctl -p carries on so nobody notices, on an old one some of your users behind a mobile carrier's NAT randomly fail to connect. That this recipe is still being copied is proof that forums outlive kernel documentation.

tcp_max_tw_buckets — The maximum number of TIME_WAIT sockets the system holds at once. When exceeded, the kernel destroys the new TIME_WAIT immediately without waiting and bumps the TcpExtTCPTimeWaitOverflow counter. The documentation says this exists only as DoS protection and that users "must not lower the limit artificially, but rather increase it." Lowering it means selling off TIME_WAIT's protection wholesale. Before lowering it on memory grounds, measure: on this server the tw_sock_TCP object in /proc/slabinfo is 264 bytes; all 262,144 buckets filled would come to roughly 69 MB. And if you sit behind NAT or a load balancer, remember conntrack keeps its own clock: nf_conntrack_tcp_timeout_time_wait defaults to 120 seconds, independent of the kernel's 60.

ip_local_port_range — The only knob that genuinely gives you "more four-tuples". On my server it's widened to 1024–65535: 64 thousand ports instead of 28 thousand. When you widen it, don't forget to set aside the ports of your own listening services with ip_local_reserved_ports, or one day the service that wants to bind 8080 will find the kernel handed that port to a client connection.

Which door to push

Diagram

The order in the diagram is deliberate: code and configuration come before sysctl. Replacing short-lived connections with a pool or keepalive doesn't just reduce TIME_WAIT; it also removes the three-way handshake and the TLS negotiation. A sysctl only masks the symptom.

On the nginx front there's a fact that needs updating this year. For years we taught "for upstream keepalive, write proxy_http_version 1.1 and proxy_set_header Connection """; the reason for the second line was that nginx used to send Connection: close to the upstream by default. The documentation and the 1.29.7 change note say all three have changed: the keepalive cache is on by default (keepalive 32 local), the default for proxy_http_version is 1.1, and the Connection header is no longer sent. The 1.30.0 on my server has this behaviour. On older versions keep writing the two lines by hand; and on every version the peer has the last word, because nginx decides whether to cache the connection by looking at the response it gets.

And one shortcut to avoid: SO_LINGER with l_linger=0. It closes the connection with a RST instead of a FIN, and TIME_WAIT never forms. It also throws away any data the peer hasn't read yet, and every close shows up on the client as "connection reset". It is the shortest route to a ticket titled "we get occasional resets" a few months later.

Checklist

  • Look at the total TIME_WAIT count with ss -s, and at who carries it with ss -Htan state time-wait | awk '{print $3}' | sort | uniq -c | sort -rn | head. The count alone means nothing; the per-address distribution does.
  • If nstat -az TcpExtTCPTimeWaitOverflow is non-zero, tcp_max_tw_buckets is too small; raise it, never lower it.
  • Search application logs for EADDRNOTAVAIL or Cannot assign requested address; that is the symptom of client-side four-tuple exhaustion. If TIME_WAIT is on the server side, the symptom is SYN retransmits and connection setup delays on the client.
  • If you have a hardening file that lowered sysctl net.ipv4.tcp_fin_timeout, add a comment saying "no effect on TIME_WAIT". Don't delete it, because someone will add it back.
  • tcp_tw_reuse is already 2 (loopback). Before setting it to 1, make sure tcp_timestamps is 1; otherwise the setting is decoration.
  • If the kernel is 6.14 or newer and you are a client opening connections to a single destination at high rate, consider tcp_tw_reuse_delay; first you need to know the peer's timestamp clock runs at millisecond resolution.
  • Throw away every recipe that mentions tcp_tw_recycle. Gone since 4.12, and it broke NAT before that.
  • Before touching sysctl, question keepalive and pooling. The cheapest change that reduces TIME_WAIT is the one that reduces the number of connections.

Conclusion

Most tuning files are an archaeological layer of a line that once worked somewhere. tcp_fin_timeout = 15 is one of those layers on my server: harmless, ineffective, and unquestioned until this article. Most settings that try to "fix" TIME_WAIT either don't touch it at all or remove it together with the thing it protects.

Seen from one level up, the lesson isn't specific to TCP. A behaviour that looks like "a problem" to you is often there to prevent an accident you can't see. When you find the switch that turns it off, the first question isn't "how do I turn this off" but "what was this switch keeping open". By leaving the TIME_WAIT duration as a constant and making only the safe reuse paths configurable, the kernel developers wrote the answer to that question into the code itself.

Official Sources

Top comments (1)

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥 •

Nice write-up Mustafa ❤️🔥