DEV Community

Remdore
Remdore

Posted on AI-assisted

Notes on waiting for a server to boot

There is a small, familiar annoyance in provisioning scripts. You create a server, you poll until the API says it is active, you immediately try to connect, and you get Connection refused. So you add a sleep 10, or a retry loop, and you move on with your life and never think about it again.

I wanted to know what number that sleep should actually be, so I measured it.

What I measured

For each droplet I recorded three moments:

  • API returns. The POST /v2/droplets call comes back with an id.
  • Status active. Polling the droplet every two seconds, the first time status reads active and a public IPv4 exists.
  • SSH answers. The first time a TCP connection to port 22 succeeds and the server sends an SSH- banner. Not a port scan — the daemon has to speak first.

Twelve DigitalOcean regions, the same s-1vcpu-1gb size and Ubuntu 24.04 image everywhere, three rounds, 36 droplets total. Everything torn down afterwards; the whole exercise cost a few cents.

The numbers

Across all 36:

moment median range
API returns 1.75s 1.12 – 4.77
status active 34.5s 23.9 – 62.0
SSH answering 45.1s 31.6 – 74.2
gap between the last two 12.1s 3.4 – 23.0

So the answer to "what should the sleep be" is about twelve seconds at the median, and about twenty-three if you want to cover the worst case I saw. That gap is 27% of the total wait. A quarter of the time you spend waiting for a server happens after the API has told you it is ready.

None of this is DigitalOcean being slow. Forty-five seconds from an HTTP request to a machine that will accept a login is genuinely quick, and the API call itself returns in under two seconds almost every time. The interesting part is not the total, it is that the last quarter of it is invisible if you trust the status field.

Per region

Sorted by median time to a usable SSH connection:

region active SSH (min / median / max) gap median
nyc2 25.6 31.6 / 35.0 / 37.4 10.6
lon1 24.7 35.4 / 36.8 / 39.2 11.9
sfo3 26.9 37.6 / 37.8 / 41.6 10.8
tor1 26.4 41.7 / 43.1 / 43.3 16.9
syd1 36.5 42.9 / 43.3 / 43.5 6.7
nyc1 26.6 39.3 / 45.4 / 49.6 13.2
sfo2 29.5 44.1 / 45.5 / 46.6 14.6
blr1 37.1 44.9 / 46.9 / 49.8 9.8
sgp1 41.2 49.2 / 50.9 / 74.2 12.2
nyc3 34.1 48.9 / 53.7 / 54.6 20.5
fra1 49.8 55.8 / 56.2 / 67.7 16.4
ams3 51.0 57.5 / 63.0 / 63.5 12.2

Fastest region to a usable machine is nyc2 at 35 seconds. Slowest is ams3 at 63. That is a 1.8x spread for identical requests.

The gap is not a constant

I had assumed the delay between active and SSH would be roughly fixed — the same boot sequence everywhere, so the same overhead. It is not.

Sydney's median gap is 6.7 seconds. New York 3's is 20.5. Three times the difference, for the same image and the same size. Sydney is comparatively slow to report active and then quick to finish; nyc3 reports active early and then makes you wait.

Which means active does not mean the same thing in every region. It is a point in a sequence, and different facilities appear to reach it at different stages of that sequence. A sleep tuned in one region is wrong in another.

Three datacentres, one city

The New York regions are the part I keep looking at:

  • nyc2 — 35.0s
  • nyc1 — 45.4s
  • nyc3 — 53.7s

Same city, same size, same image, same evening. nyc3 takes 53% longer than nyc2 to reach a usable state. If you pick a New York region out of a dropdown without thinking — and everyone does — you are making a choice worth 19 seconds every time you build a machine.

For a single server that is nothing. For a CI job that creates and destroys a fleet, or a test suite that provisions per run, it adds up in a way nobody ever attributes to the dropdown.

What I would actually do

Stop sleeping, start polling. A two-second poll on the SSH banner costs nothing and is correct everywhere. A fixed sleep is either too short in ams3 or wasteful in nyc2, and you cannot pick a value that is both.

Poll for the banner, not the port. A TCP connect can succeed before sshd is willing to talk. Read the first bytes and check they start with SSH-.

Do not treat active as ready. It is a real signal and a useful one, but it means the hypervisor has finished, not that the guest has. Those are about twelve seconds apart, most of the time.

Caveats, of which there are several

Three rounds per region is not many. The medians are steady enough that I trust the ordering, but a single number per region would not have been worth printing — the first round alone had nyc1 at 39.3s and the third had it at 49.6s, and if I had stopped at one round I would have written a different article.

My poll interval is two seconds, so every timestamp carries that much uncertainty. Differences of a second or two between regions mean nothing here; differences of twenty do.

I measured SSH reachability from one machine in Europe, so the far regions carry a little extra round-trip time. It is on the order of a couple of hundred milliseconds against a forty-five second wait, which does not explain anything in that table, but it is there.

And this is one image, one size, one evening. Boot time is not a fixed property of a region — it is what that region happened to do while I was watching. Anyone rerunning this next Tuesday should expect different numbers and the same shape.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.