DEV Community

137Foundry
137Foundry

Posted on

Why Your API Timeout Settings Are Probably Too Long

Most HTTP clients ship with either no default timeout at all or one set generously enough, 30 seconds, 60 seconds, to accommodate almost any conceivable slow response. Most teams never change it. This is a bigger problem than it looks like on the surface, and it's one of the cheapest fixes available in any pipeline that calls external APIs.

What a long timeout actually costs you

A timeout isn't just about how long a single request waits before giving up. It's about how much capacity that wait consumes while it's happening. A worker thread, a connection from a pool, a slot in a queue, all held hostage by a single request that's slowly failing. Multiply that by concurrent requests to a dependency that's having a bad day, and a 30-second timeout can tie up your entire available capacity for that integration in well under a minute, even though no individual request technically "broke" anything yet.

This is the mechanism behind a slow dependency causing more damage than an outright down one. A hard failure, a connection refused, fails fast and frees up capacity immediately. A slow, degraded dependency that's still technically responding, just taking 25 of your 30 allotted seconds to do it, ties up resources far longer per request while looking, from a naive monitoring dashboard, like everything is still working.

What a reasonable timeout actually looks like

Most internal service calls should complete in well under a second under normal conditions. If your API's typical response time is 200 milliseconds, a 30-second timeout is 150 times longer than what a healthy response ever takes. A timeout in the 2-5 second range for most synchronous API calls gives real headroom for normal variance without allowing a struggling dependency to hold your resources for anywhere near as long as the default would.

For genuinely slow operations by design, large batch exports, long-running report generation, the right answer usually isn't a longer synchronous timeout at all, it's moving the operation to an async pattern: kick off the job, return immediately, and poll or receive a webhook when it's done. Trying to cover a legitimately slow operation with a longer timeout just delays the same resource-exhaustion problem rather than solving it.

Why teams leave defaults in place for so long

The honest reason most timeout settings never get revisited is that they don't cause visible problems most of the time. A dependency that's healthy 99 percent of the time never exposes the cost of an overly generous timeout, because the timeout never actually triggers under normal conditions. It's only during the rare degraded period, exactly when you can least afford wasted capacity, that the long timeout setting starts actively working against you. This lag between cause and visible effect is exactly why it's worth auditing proactively rather than waiting for an incident to force the question.

Timeout budgets across a request chain

In a system where one service calls another which calls a third, timeouts need to be coordinated across the chain, not set independently at each hop. If service A calls service B with a 5-second timeout, and B calls a third-party API with its own 10-second timeout, A's timeout will fire before B ever gets an answer from the dependency it's waiting on, and A has no way to know that B is still legitimately working. Each hop's timeout should be shorter than the timeout of whatever's calling it, with enough margin for the calling service to actually process the result and respond in time.

Different timeouts for different failure stages

A single "timeout" setting often actually needs to be broken into at least two: connection timeout, how long to wait for the initial connection to establish, and read timeout, how long to wait for a response after the connection is made. These fail differently and should usually have different values. A connection timeout of 2-3 seconds is reasonable for most cases, since establishing a TCP connection to a healthy service should be fast. A read timeout depends more on what the endpoint actually does, but should still be set deliberately rather than left at a library default.

The relationship between timeout settings and circuit breakers

Shorter, well-tuned timeouts make your circuit breaker's job easier. A breaker tracking failure rate over a rolling window works better when failures are detected quickly, since a slow-to-timeout request delays not just that individual call but also delays how quickly the breaker can recognize a pattern of failures and trip. Long timeouts and circuit breakers work against each other; short, deliberate timeouts and circuit breakers reinforce each other.

There's a fuller breakdown of the circuit breaker side of this in a longer article on designing resilience for flaky dependencies, including how to set trip thresholds that respond to real failure patterns rather than noise.

"A 30-second default timeout isn't a safety margin, it's 150 requests' worth of worker capacity quietly reserved for one that's already failing. Tightening it is one of the highest-leverage, lowest-risk changes a team can make in an afternoon." - Dennis Traina, [founder of 137Foundry](https://137foundry.com/services)

A practical audit worth doing this week

Pull up every external API integration in your codebase and check the configured timeout against the endpoint's actual typical response time, which most API providers' status pages or your own monitoring can tell you. Anywhere the timeout is more than five to ten times the typical response time is a candidate for tightening. This is usually a small, low-risk config change per integration, not a rewrite, and it directly reduces how much damage a single degraded dependency can do to your overall system capacity.

Roll the change out gradually if the integration is high-traffic, tighten the timeout for one instance or one percentage of traffic first, watch error rates for a day, then widen the rollout. This catches the rare case where your "typical" response time measurement was itself skewed by a recent quiet period, before the tighter setting affects all your traffic at once, and it gives you a clean rollback path if the new number turns out to be too aggressive for some legitimate slow-but-healthy edge case you hadn't accounted for, which happens occasionally and is far easier to catch and reverse on a small slice of traffic than after a full rollout.

Sources

The AWS Well-Architected Framework covers timeout configuration as part of its reliability guidance, situated alongside retries and circuit breakers rather than treated as a standalone setting. Google Cloud's API design documentation also touches on expected client timeout behavior from the API provider's perspective, useful for understanding why a well-configured client matters to the services it's calling, not just to your own system. Martin Fowler's site rounds this out with the broader resilience pattern context, since timeout tuning is really the foundation the rest of a circuit breaker and retry strategy gets built on top of.

None of this requires deep architectural changes. The team behind this treats a timeout audit as one of the fastest wins available in any integration health review, precisely because it's usually a five-minute config change per endpoint with an outsized effect on how gracefully the whole system degrades when a dependency has a bad day.

Top comments (0)