DEV Community

Sergey Shinder
Sergey Shinder

Posted on

Our deploys were also our restart policy

At four on a Sunday morning an internal reconciliation service went down. All four pods were killed for memory inside twenty minutes of each other, came back, and were killed again a few hours later. The service had last been deployed in February. Seven months of continuous uptime, and it was the only thing in our estate that had any.

The cause was a leak of about forty megabytes per pod per day. We create an SDK client per request, and each client registers a listener on a static registry that nothing ever unregisters. A heap comparison showed one point two million listener objects, which at our request rate is exactly what you would expect after seven months. It is a beginner's mistake in a library that has had it since 2019.

The uncomfortable part came next. That library is in every service we run. I plotted resident memory against pod age across the whole fleet, and every service has the same slope. None of them had ever shown it, because the median age of a pod in our clusters is under thirty hours. We deploy several times a day, and a restart clears the leak, so no instance of anything had ever lived long enough to reach the ceiling. The one service nobody had a reason to change was the only one telling us the truth about our code.

We fixed the client, of course. The more useful changes were the other three. Memory growth is alerted on as a rate normalised by pod age, because an alert at ninety percent of the limit fires a few minutes before the kill and is not a warning, it is a commentary. One instance of each service now runs in a soak environment for thirty days against synthetic traffic with its memory graphed. And pod age distribution is on the platform dashboard, because a fleet in which nothing is older than two days cannot tell you anything about how it behaves on day thirty.

Deploying often had become a reliability control, and we had never decided to have it. Anything that hides a defect is load bearing. If the only reason a system stays up is that you keep restarting it, a quiet fortnight is a risk and not a rest.

– Sergey Shinder

Top comments (0)