Our address lookup provider went down at eight minutes to two on a Tuesday and was healthy again by four minutes to two. Our checkout, which calls it to validate delivery addresses, did not use it again until five past five. For three hours every customer typed their address by hand under a warning, and completion on that step fell by about a fifth.
The call sits behind a circuit breaker. It opens when half of the last twenty calls fail, stays open for thirty seconds, and then goes half open and lets one trial request through. Success closes it. Failure opens it again, for twice as long as before, up to thirty minutes. The doubling was added a year ago by someone being considerate towards a partner under strain.
The trial request was whichever customer request happened to arrive first. The provider came back cold: for its first few minutes after a restart it answers in one and a half to three seconds while its caches fill. Our timeout for that call is eight hundred milliseconds, set against their normal ninety. So each trial timed out and the breaker reopened, for a minute, then two, four, eight, sixteen, and then thirty minutes at a time, each time sending a single request into a service that could only warm up by receiving requests.
Nobody was paged. Breaker transitions were logged at info and the fallback worked, so by every rule we had we were degraded, not down. A product analyst found it the next morning by asking why manual address entry had jumped.
Breaker state is a metric now, with an alert if any breaker stays open for more than five minutes. Half open admits five percent of traffic for thirty seconds and judges the success rate, instead of betting everything on one request. Trial calls get a longer timeout than normal traffic. The open duration is capped at sixty seconds with no doubling. There is a manual close in our operations tool. And we test recovery, not only failure, with a proxy in staging that imitates a slow cold start after an outage.
A circuit breaker has two jobs. Stepping out of the way of a failing dependency is the one everybody demonstrates. Letting it back in decides how long your outage lasts, and ours was being patient with a partner who had already recovered.
– Sergey Shinder
Top comments (0)