DEV Community

Remdore
Remdore

Posted on AI-assisted

Where the four minutes go when a cluster adds a node

Autoscaling is usually described as a property a cluster either has or does not have. You add a HorizontalPodAutoscaler, you tick the box for node autoscaling, and from then on load is somebody else's problem.

How long that somebody takes is the bit I could never find written down. I had a rough feeling it was "a minute or so" and no idea where the minute went. So I timed it. One pod, a burst of traffic it cannot handle alone, and a timestamp at each step until a second pod is answering requests.

With spare room in the cluster the first extra pod was up in about 30 seconds. With no room, so that a node had to be added first, it was a little over four minutes. I had guessed wrong about where most of that time goes in both cases.

The setup

I used DigitalOcean Kubernetes, version 1.36.3, on s-2vcpu-4gb nodes. The node pool autoscales between two and five nodes, which on DOKS is a flag you pass when creating the pool. I did not have to install or configure a cluster autoscaler myself.

For the app I took hpa-example, the PHP image from the Kubernetes docs that does a pile of arithmetic on each request. It runs as one replica with a 200m CPU request and a 500m limit. The autoscaler targets 50% utilisation and may go up to five pods. fortio runs in the same cluster and hammers the service over six connections with no rate limit.

Every time below is measured from the moment the load starts. The start is stamped by creating an object and reading back the timestamp the API server gave it, so the stopwatch and the pod timestamps come from the same clock.

When there is room on the nodes

Eight trials.

stage median range
autoscaler makes its first decision 28s 25 – 43s
first new pod ready 30s 27 – 46s
autoscaler makes its second decision 43s 40 – 59s
all five pods ready 46s 43 – 61s

A new pod went from created to ready in one to three seconds. Scheduling was instant and the image was already on the node. So of the thirty seconds, about twenty-eight are spent waiting for the autoscaler to find out anything is wrong.

That delay is a pipeline of polling loops. The kubelet samples CPU, metrics-server scrapes the kubelet on an interval, and the autoscaler reads metrics-server on its own interval. The results cluster around 28 seconds and around 43, fifteen seconds apart, which is what you get when a spike either just catches a cycle or just misses one.

The second row of that table is the one I did not expect. In every trial the autoscaler needed two rounds to get to the right size. The pod's limit is 250% of its request and that is where the later readings sat. But the first reading the autoscaler acted on was whatever had accumulated in the metrics window so far. In different trials I caught it at 55%, 103%, 138% and 249% on its first look. It scaled to three or four pods, waited a cycle, saw it was still over, and added the rest.

So the honest number for "scaled out" is not 30 seconds. It is 46.

When a node has to be added

For this I filled both nodes with placeholder pods so that less than one application pod's worth of CPU was free on each, then ran the same spike. Four trials.

stage median range
spike to new pod created (and stuck Pending) 19s 12 – 42s
pod Pending to node requested 17s 16 – 31s
node requested to node registered 120s 117 – 123s
node registered to node Ready 36s 35 – 45s
node Ready to pod scheduled on it 29s 28 – 29s
pod scheduled to pod ready 20s 20 – 26s
spike to first new pod serving 249s 232 – 276s

Four minutes and nine seconds at the median, during which one pod was carrying everything.

A few things stand out.

Creating the machine is half of it, and it is remarkably steady. From the autoscaler asking for a node to that node appearing in kubectl get nodes took between 117 and 123 seconds in all four trials. A six-second spread on provisioning a virtual machine, installing a kubelet and joining a cluster is tighter than I would have guessed, and it means you can plan around the number.

The other half is spread across five smaller waits that nobody mentions. The pod has to exist and fail to schedule before the cluster autoscaler will act, and that autoscaler runs its own loop, which cost another 17 seconds. The node registers well before it is Ready. And after it reports Ready there is a further gap of 28 or 29 seconds, almost constant, before the pending pods are actually placed on it. I do not know what that gap is. It is too regular to be noise.

Then the image. On a brand new node nothing is cached, and pulling 164MB took 18 to 22 seconds. In the first scenario that step took under half a second.

Keeping a spare node warm

The standard answer to the four-minute problem is overprovisioning: run a placeholder pod with a negative priority that requests about a node's worth of CPU. It forces the cluster to keep one more node than it needs. When real pods arrive they evict the placeholder and take its space immediately, and the placeholder, now homeless, triggers the node scale-up in the background.

Three trials with a 1200m placeholder:

  • first new pod ready at 47s, 49s and 75s after the spike
  • every application pod landed on the spare node, none waited for a new one
  • the replacement spare node was Ready 239 to 383 seconds after the spike, with nobody waiting on it

That is four minutes down to under one. It is not quite as fast as having room on a node that already runs the application, and the difference is entirely the image pull: the spare node had never run this image, so the pods spent 17 to 22 seconds downloading it. Pods were ready a median of 20 seconds after the decision, against 2 seconds in the first scenario.

The cost is one idle node, all the time. Whether that is worth it depends on what four minutes of a saturated service costs you.

What I got wrong

The first version of the deployment sat at 92% CPU with no traffic at all, and the autoscaler had scaled it to five pods before I had started anything. My readiness probe was an HTTP check against the same page as the load test, once a second. On an application that burns CPU per request, the probe was the traffic. I changed it to a TCP check.

Then three of my first six trials were invalid, and they looked fine. After each run I deleted the autoscaler, scaled back to one pod, recreated the autoscaler and waited for it to report low utilisation. It did report low utilisation, briefly, and then the tail of the previous spike arrived through the metrics window and it scaled to five pods a few seconds before my next trial began. Those runs recorded a "decision" two seconds after the spike and zero new pods, which I could easily have read as an instant response. The reset now waits out the window before an autoscaler exists, and the script refuses to start unless exactly one pod is running.

And the load generator died silently in the node trials. fortio gives up if its warm-up request takes more than three seconds, which is exactly what happens when the one pod is already swamped. It exited in under four seconds and my script carried on timing a spike that was not happening.

What I would take from it

Thirty seconds is the floor, not the typical case. With room on the nodes and the image cached, you still wait most of half a minute for the metrics to arrive, and you are not at full size for fifteen seconds after that.

The first reading is an underestimate. A spike that starts mid-window looks smaller than it is. Expect two rounds of scaling, and if you tune anything, tune for that.

If a node is needed, think in minutes. About two of them are the machine, and the other two are five separate waits stacked end to end. No single one is large enough to complain about.

A spare node buys back three of the four minutes. What is left is the image pull, so on a spare node image size finally matters.

Measure it on your own cluster before you rely on it. The whole exercise ran for under three hours on two or three small nodes and cost about thirty cents. The numbers will differ with your image, your node size and your provider, but the stages will be the same ones, and knowing which of them you are waiting on is most of the value.

Caveats

Four trials with a new node and three with a spare is not many, though the spread was small enough that I trust the shape. This is one image, one node size, one region and one evening. The application is a CPU-bound demo, so the metric moves as fast as a metric can; a service that degrades without burning CPU would be noticed later or not at all. And I did not change any autoscaler or metrics-server settings, so these are defaults.

The manifests, the timing script and the raw output of every trial are in the repository.

Top comments (0)