DEV Community

charmingdisaster
charmingdisaster

Posted on AI-assisted

The Low-Level Mechanics of Connection Timeouts vs Read Timeouts (SO_TIMEOUT)

Something I've seen take down production backends a few times is a simple misunderstanding of how network timeouts actually work under the hood.

A lot of developers set readTimeout = 5000 on their HTTP client and assume: "Great, if the remote service takes longer than 5 seconds, it will fail fast."

Except SO_TIMEOUT (Read Timeout) doesn't limit the total request duration at all. In POSIX-compliant socket abstractions, it only measures the delay between incoming TCP packets.

If a degraded downstream service or reverse proxy trickles just a single byte of data every 4.9 seconds, a 5-second socket timeout will continuously reset back to zero. A single hanging request can keep your worker thread trapped for hours without ever throwing a timeout exception.


Two Other Network Traps

  1. Connect Timeout vs Read Timeout: If intermediate network infrastructure or a firewall silently drops your SYN packet without returning a TCP RST, the Linux kernel defaults to retrying for 75 to 127 seconds (tcp_syn_retries) before giving up with ETIMEDOUT. Without an explicit Connect Timeout on the client, your execution thread freezes before a socket is even established.

  2. The "UNKNOWN" State: Remote calls are never strictly binary. If your socket times out waiting for a response, you have no deterministic way to know if the payment gateway processed the charge or crashed right after saving it. Retrying blindly without idempotency keys is how double-billing happens.


How to Configure This Properly in Spring Boot (Java 21)

To protect your system, always set both timeouts explicitly using RestClient or RestTemplate:

import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.http.client.JdkClientHttpRequestFactory;
import org.springframework.web.client.RestClient;

import java.net.http.HttpClient;
import java.time.Duration;

@Configuration
public class RestClientConfig {

    @Bean
    public RestClient restClient() {
        HttpClient httpClient = HttpClient.newBuilder()
                .connectTimeout(Duration.ofSeconds(2))
                .build();

        JdkClientHttpRequestFactory requestFactory = new JdkClientHttpRequestFactory(httpClient);
        requestFactory.setReadTimeout(Duration.ofSeconds(5));

        return RestClient.builder()
                .requestFactory(requestFactory)
                .baseUrl("https://api.example.com")
                .build();
    }
}
Enter fullscreen mode Exit fullscreen mode

To guard against the "slow trickle" issue, enforce an absolute wall-clock execution deadline around critical calls (e.g., using CompletableFuture.orTimeout()).

Note: This breakdown is excerpted from the opening chapter of my field guide "Surviving Distributed Systems"—a practical book covering production failure modes, concurrency, and fault tolerance in Java 21.

You can download the full 16-page free sample PDF (Chapter 1) directly on Leanpub:

https://leanpub.com/surviving-distributed-systems

Top comments (0)