A few weeks ago I measured a Kubernetes rolling update dropping requests even with maxUnavailable: 0, the setting people choose precisely because they want zero downtime. The cure took three separate pieces: a preStop delay, graceful shutdown in the application, and closing keep-alive connections while draining. Leave any one of them out and connections died during every deploy.
That raised an obvious follow-up. Most people deploying a small web service never touch a Deployment manifest; they push to a platform and let it handle the rollout. So I took the worst version of the app from that experiment, the one that exits the instant it receives SIGTERM and does no draining of any kind, and deployed it to DigitalOcean's App Platform to see what the platform does on its behalf.
On Kubernetes that app lost 51 requests in a single rollout. On App Platform, across four redeploys and 278,014 requests, no request failed because of a deploy.
Same app, same client
The server is the same Python file as before, with one change: every response now names the container that served it and a version number taken from an environment variable, so the load generator can see the exact moment traffic moves to the new version.
web-96969cc79-cwvw8 gen=1
It runs on the smallest instance App Platform offers, one shared vCPU and 512 MB, as a single instance. The client is the same keep-alive load generator, twenty threads holding persistent connections and firing continuously, now over HTTPS to the app's public URL. Each test starts the load, waits thirty seconds, then changes the version number in the app spec, which triggers a redeploy, and keeps going until well after the new version is live.
Before any of that, the control: two minutes of the same load with no deploy.
requests=22450 ok=22450 failed=0
Over the public internet, through a TLS terminator, nothing failed. Anything that fails during a deploy is the deploy's doing.
Four redeploys
| requests | failed | |
|---|---|---|
| deploy 1 | 85,892 | 0 |
| deploy 2 | 62,040 | 0 |
| deploy 3 | 66,981 | 0 |
| deploy 4 | 63,101 | 1 |
The single failure in the fourth run deserves a proper look rather than a footnote. It was a 502 that came back in three milliseconds, 167 seconds after that deploy had finished and the old container was long gone. The deploy didn't cause it. It is one request in about three hundred thousand going wrong at the edge while nothing was happening, and I am counting it in the table rather than leaving it out.
Each redeploy went live in under a minute, and the phase transitions are visible through doctl apps list-deployments as they happen:
What is between me and the container
The first question with a result this clean is what is absorbing the failures. The response headers answer part of it:
server: cloudflare
cf-ray: a47bc3222d20d0e8-SOF
x-do-app-origin: 920fd385-f27b-4532-a919-a8736dabf009
App Platform puts Cloudflare in front of every app, and my client's connections were to the Cloudflare edge in Sofia, not to the container. That alone removes the failure mode from the Kubernetes post, where the dying pod's keep-alive connections were severed mid-use, because here no client ever holds a connection to a pod that is going away.
The hostname is worth a second look too. web-96969cc79-cwvw8 has the shape of a Kubernetes pod name, a deployment, a replica-set hash and a random suffix. I can't see past the edge to confirm what runs underneath, but whatever it is, the platform is doing the work I had to do by hand last time.
The cutover itself was gentle. For 10.0 seconds both versions answered, interleaved almost exactly half and half, and then the old one stopped receiving new requests. Latency through the cutover was indistinguishable from either side of it, with a median of 42.9 ms against 44.3 ms beforehand.
Where it stops being free
Every request in those runs completed in about forty milliseconds, which leaves a question open. A platform that drains an instance before stopping it and a proxy that quietly retries a failed request would produce exactly the same zero. To tell them apart I needed requests that were still running when the old version was stopped.
So I added an endpoint that sleeps for twenty seconds before answering, and alongside the normal load started one of those every second, each on its own connection. Every response reports which version finished it.
The answer is that the platform drains, it does not retry. A request that reached the old version 4.1 seconds into the cutover still ran its full twenty seconds and was answered by the old version, and nothing anywhere in the run was silently re-executed on the new one.
The drain has a limit, though, and it is precise. After the old version received its last new request, it was kept alive for 15.0 seconds and then stopped. Two requests that had reached it in the final moments of the overlap were still sleeping at that point, and both failed at the same instant with a 504.
I ran the default case twice. Both times the window was 15.0 seconds, both times exactly two of 150 in-flight requests were cut off, and the moment the old container stopped differed between the runs by a fraction of a second relative to its last new request.
The setting that fixes it
App Platform's service spec has a termination block, which I had never noticed before this:
termination:
drain_seconds: 60
grace_period_seconds: 90
With that applied, the identical test completed all 150 in-flight requests. A request that reached the old version at the very last moment before cutover ran its full twenty seconds and finished there, with the old container staying up as long as it needed.
There is a cost, and it is the obvious one. That deploy took 69.8 seconds to become active, against 42 to 56 seconds for the runs on default settings. I only measured it once with the longer drain, so I would treat that as an indication rather than a figure, but it makes sense that waiting for slow requests to finish takes longer than not waiting.
The comparison, stated fairly
| hand-built Kubernetes | App Platform | |
|---|---|---|
| app shutdown behaviour | exits on SIGTERM | exits on SIGTERM |
| instances | 4 | 1 |
| in front of the app | cloud load balancer | Cloudflare edge |
| requests lost in one rollout | 51 and 35 | 0, 0, 0, 0 |
| fix needed for normal traffic | three changes, by hand | none |
| fix needed for 20 s requests | not tested | one setting |
The two setups are not identical and the table should not be read as if they were. The Kubernetes test had four replicas behind a load balancer holding keep-alive connections straight to the pods, which is the exact arrangement that exposed the problem. App Platform terminates connections at an edge that sits in front of everything. That is the point, though. The architecture is the product, and the reason the naive app survives is that the platform was built not to let a dying container take client connections down with it.
What I got wrong
My first attempt at the slow-request test silently tested nothing. I added the twenty-second endpoint to the repository, changed the version number in the spec to deploy it, and queried the new endpoint, which returned an ordinary response with no slept= in it. The endpoint wasn't there. Changing only an environment variable reuses the previous build, and the active deployment was still built from commit c0af922 while the branch had moved on to eb8d921. With a plain git source, picking up new code needs doctl apps create-deployment --force-rebuild. I only noticed because the response was missing one word.
My first pass over the results also printed an empty table where the cutover analysis should have been. I had compared request times measured from the start of the run against a cutover time I had forgotten to convert, so every request fell outside the window. An empty result from an analysis is far more likely to be a bug than a finding, and here it was a bug.
The last is a measurement I can't explain. In the first slow-request run, every request launched in the opening twenty seconds took exactly forty seconds instead of twenty, and every request after that took twenty. It did not reproduce with the slow probe alone, with the probe alongside the normal load, or in any later run. Those requests all completed successfully on the old version, so the cutover analysis doesn't depend on them, but I can't tell you why it happened, and that doesn't seem worth hiding.
What to take from it
If you deploy a web service to App Platform, the rolling-update problem from the Kubernetes post is handled for you. A container that exits the moment it is told to, with no draining and no preStop hook, came through four redeploys without losing a single ordinary request. Being spared that whole category of work is a large part of what a platform is for, and it is good to see it measured rather than just claimed.
The one thing to know about is the drain window. By default the old version gets fifteen seconds to finish what it already has once it stops receiving new requests. If any of your requests can run longer than that, such as uploads, report generation, slow third-party calls or long polling, set termination.drain_seconds in the spec to something longer than your slowest request, accept that deploys will take a little longer, and those requests will stop failing too.
Everything ran on the smallest instance at a per-second price that came to well under a cent for the whole experiment, and the app, the load generator and the raw results are in the same repository as the original Kubernetes test.


Top comments (0)