DEV Community

Sergey Shinder
Sergey Shinder

Posted on

Our partner calls failed more often the quieter it got

Roughly three calls in a thousand to our settlement partner failed with connection reset by peer. Retrying worked every time. It had been like that for two months, it was costing us nothing measurable, and the only reason I looked at it properly was that the failure pattern was backwards.

The errors clustered overnight and on Sunday afternoons. Not during the Monday morning peak, not during the end of month run. The busier we were, the healthier the integration looked, and I could not construct a story about their capacity that produced that shape.

It was connection reuse. Our HTTP client keeps idle connections in a pool for five minutes, which is the library default and which none of us had ever chosen. Their load balancer closes idle connections after sixty seconds. When it closes one, the socket in our pool looks perfectly fine to us, because nothing reads from an idle connection; the close notification sits unread in the kernel buffer. The next request gets written onto a connection the other end threw away half an hour ago, and the reset comes back immediately.

At peak, our connections were never idle for sixty seconds, so the pool was always warm and the problem did not exist. At two in the morning, with one call every few minutes, almost every connection in the pool had aged past their timeout. The quieter it got, the higher the proportion of requests that were the first use of a stale socket.

Our pool now holds idle connections for thirty seconds, comfortably under their sixty, with a hard lifetime of five minutes on top so that a connection is never reused indefinitely. We also allow exactly one immediate retry for the specific case where the connection was closed before any byte of the request reached the wire, which is safe in a way that retrying a timeout is not, because that request never arrived.

The last change was a document. Our integration notes for each partner now carry their idle timeout, their request timeout, their maximum body size, their TLS versions and their source address ranges. Not one of those five numbers was in their public documentation. All five came from an email exchange that took a day.

Two timeouts, neither of them wrong. The fault was entirely in which one was longer.

– Sergey Shinder

Top comments (0)