DEV Community

Sergey Shinder
Sergey Shinder

Posted on

A hung test held the only runner allowed to deploy our hotfix

In August a configuration change made our checkout reject one card type, and the fix was a single line. It was merged at twenty past two. It reached production at twenty to six, and for almost all of the time in between it sat in a queue with a status of queued, which reads as if something is about to happen.

Production deploys run on a self hosted runner inside our network, labelled deploy, because it is the only machine with a route to the cluster API. A year earlier somebody had given the integration test job the same label, because those tests needed a database that lived on the same private network. So the one machine allowed to deploy was also the one running our slowest tests.

That afternoon one integration test opened a connection to a message broker in a test environment that was being rebuilt. The broker accepted the connection and never answered. Our client library has no read timeout by default, the test framework had no per test limit configured, and the job had no timeout of its own, so it inherited the platform default of six hours. The test sat in a blocking read, the runner showed busy, and our hotfix waited behind it.

We found it at half past five, when somebody stopped looking at the hotfix run and looked at what the runner was actually doing. The deploy finished nine minutes after we cancelled the test.

What changed. Every job now sets timeout-minutes to about three times its normal duration, taken from the last month of runs, and a lint step rejects any workflow that leaves it out. The test framework has a per test timeout too, so a hang fails one named test instead of quietly spending the whole job's budget. The deploy runners do nothing but deploy, there are two of them in different zones, and tests that need the private network run on their own pool with their own label. And any job queued on the deploy label for more than five minutes pages whoever is on call, because during an incident a waiting deploy is part of the incident.

Nobody chose six hours. It is simply the longest the platform will wait before giving up on you, and a default like that is a ceiling, not a budget. Ours was sitting between a broken checkout and the line that fixed it.

– Sergey Shinder

Top comments (0)