DEV Community

Remdore
Remdore

Posted on AI-assisted

A dead Kubernetes node is detected in 3 seconds and keeps receiving traffic for 13

The folklore about losing a Kubernetes node is specific and widely repeated. The node controller waits forty seconds before marking an unresponsive node NotReady, because that is the node-monitor-grace-period default, and then the pods sit there for another five minutes before the taint manager evicts them, because that is the default toleration. People quote those numbers when they explain why a dead node is so painful.

I wanted to see what actually happens, so I built a three node cluster, put six replicas behind a cloud load balancer, pointed a steady stream of keep-alive clients at it, and cut the power to a worker. Not a graceful shutdown, not a drain. A hard power-off, the way a machine dies when a hypervisor host fails.

Then I did it nine more times, because the first result was not what I expected and the second contradicted the first.

The setup

DigitalOcean Kubernetes 1.36.3, three s-2vcpu-2gb nodes, six replicas spread two per node, a type: LoadBalancer Service, and about 175 requests per second from twenty clients holding keep-alive connections. Each run lasts three minutes and the node is killed thirty seconds in, by calling the provider API to power the droplet off rather than asking the operating system to stop.

Alongside the request log, a watcher samples node readiness and the Service's endpoint list twice a second, so I can line up what the cluster believed against what the clients experienced.

Detection is fast, and that is not the problem

Across all ten runs the node was marked NotReady in 2.9 seconds, with a range of 2.7 to 3.7. That is not forty seconds. It is not close to forty seconds.

I do not think the folklore is wrong so much as out of date, and the managed platform is clearly not running the stock timings. Whatever DigitalOcean has tuned here, it noticed a machine had stopped existing about thirteen times faster than the default implies, and it did so consistently enough that the range across ten runs is a single second wide.

The slow part is what happens next:

median, ten runs
node marked NotReady 2.9s
its pods removed from the Service endpoints 13.2s

Ten hard power-offs, every failed request, and the timings underneath

There is a ten second gap between the cluster knowing the node is gone and the cluster stopping sending traffic to the pods on it, and every failure in this experiment lives inside that gap, which makes the cost of a dead node a question about how quickly the endpoint machinery reacts rather than about how quickly anything detects the failure.

For contrast, the same measurement after a kubectl drain:

endpoints updated after a clean drain: 0.5s
Enter fullscreen mode Exit fullscreen mode

Twenty-six times faster, and the drain run produced exactly one failed request out of 31,652. The gap between planned and unplanned is the largest single number in this experiment and nothing I did to the configuration came close to closing it.

Where I was wrong

I came into this expecting externalTrafficPolicy: Local to be the fix.

The reasoning seemed sound. With the default Cluster, every node accepts traffic for the Service and forwards it to any pod anywhere, so a surviving node will happily forward your request to a pod on the corpse until the endpoints catch up. With Local, a node only serves pods that are physically on it, so there is nothing to forward and the load balancer's own health check takes the dead node out within a few seconds.

After three runs it looked like I was right, roughly a four-fold improvement in the failure window. I nearly stopped there and wrote that up.

At five runs per policy the totals are identical:

failures per run median
Cluster (default) 10, 18, 7, 11, 9 10
Local 3, 0, 11, 10, 11 10

The setting did not change how much traffic I lost. It changed when I lost it, and that turns out to be the more interesting result:

failures within 60s of the kill runs with a later burst
Cluster median 4 3 of 5
Local median 10 0 of 5

Local takes the entire hit immediately, inside a tight window that was 10.0, 10.0 and 10.1 seconds across three consecutive runs, and then it is finished. Cluster takes a smaller initial hit and then produces further bursts a minute or two later, in three runs out of five.

If you are the sort of person who would rather have one short, predictable outage than a smaller one followed by aftershocks, that is an argument for Local. It is not the argument I expected to be making, and it is not about reducing the damage.

One thing I cannot explain

Those later bursts deserve stating plainly rather than being quietly folded into a total.

They are 6 to 14 failed requests inside a tenth of a second, which means every client thread failed at the same instant. When they happen, node readiness is unchanged, the endpoint list is unchanged, and no node has been replaced. All three nodes were the same age at the end of every run.

Twenty threads failing simultaneously against a cluster whose state has not moved points somewhere in the load balancer path rather than at the node loss, and I stopped there rather than guess. It is in the raw data in the repo. I mention it because the alternative was to report a 112 second failure window for one run, which is what a naive reading of that run produces, and which would have been the most dramatic and least honest number in the piece.

What else I got wrong

Two setup mistakes, both caught before they reached the results, and both of the kind that produce a perfectly clean number from a meaningless experiment.

The first run put all six replicas on a single node. The other two nodes had become Ready only seconds earlier, the topology spread constraint was set to ScheduleAnyway, and the scheduler had no reason to wait. Killing that node is a total outage test, not a node loss test, and the resulting number would have been six times too large.

The second attempt had three pods on one node, three on another and none on the third, which is a fairer test but still means that killing a node removes half the capacity rather than a third. I discarded that too, forced a genuine two-two-two spread, and started again. Both discarded runs are in the repository.

The third mistake was nearly the worst. Three runs per policy showed Local winning clearly, which matched my hypothesis, which is exactly when you should be most suspicious. Two more runs per policy reversed it. If the number you get agrees with what you expected, that is a reason to run it again rather than a reason to stop.

What to take from it

The headline most people carry around about node failure is pessimistic in the wrong place. Detection is fast, at least on a managed cluster that has tuned it, and the forty second figure is not what you should be planning around. The number that matters is how long the endpoints lag behind the detection, which was about ten seconds here, and during which your load balancer is still cheerfully sending traffic to a machine that no longer exists.

Nothing in the Service configuration closed that gap. externalTrafficPolicy only redistributed when the failures arrived.

What did close it was draining before the node went away, which took endpoint removal from 13.2 seconds to 0.5. That is only available to you for planned work, which is an argument for doing node maintenance, upgrades and scale-downs through a drain every time, rather than relying on the cluster to cope, because when the loss is genuinely unplanned you are going to eat those ten seconds and there is no setting that prevents it.

And if your client retries on connection failure, which most do, this whole category of event is invisible to you anyway. Ten failures out of 26,000 is a rounding error until it is the request somebody was waiting on.

Top comments (1)

Collapse
 
kashif_manzer profile image
Kashif Manzer •

The later bursts are the most interesting part of this post, and your observation that every client thread failed inside a tenth of a second basically rules out anything at the pod level. A shared path failed: the client keep-alive pool, the DigitalOcean load balancer, or kube-proxy's iptables rewrite when endpoints updated. My money would be on the connection pools. Twenty threads with twenty long-lived connections will all route through the same NAT and conntrack entries, and when those entries get refreshed or expired they all die in the same instant, a minute or two after the event that disturbed them. If you rerun it, capturing conntrack timestamps alongside the request log would confirm or kill that hypothesis quickly.