If a call from inside a container occasionally takes almost exactly 5 seconds and then succeeds normally, you are almost certainly looking at a DNS resolver timeout, not a slow upstream. The glibc resolver waits 5 seconds per nameserver attempt by default, and Kubernetes pods ship with options ndots:5, which turns every external hostname into several cluster-internal lookups before the real one is tried. The fix is to stop generating those extra queries — per-pod dnsConfig, a fully-qualified hostname with a trailing dot, or a node-local DNS cache — not to raise your HTTP client timeout.
Why does the delay land on exactly 5 seconds?
That number is a resolver default, not a coincidence. resolv.conf supports options timeout:n attempts:m, and on glibc the defaults are a 5-second timeout with 2 attempts per nameserver. So any DNS query whose response never arrives costs you a flat 5 seconds before the resolver retries or moves on.
This matters because a dropped UDP response is indistinguishable from a slow one. The usual causes are a saturated or restarting DNS pod, a conntrack race on the DNAT path to the DNS service, or simply too many queries in flight because of the fanout described below. Whatever the trigger, the symptom is the same shape: p50 latency is fine, p99 has a cliff at 5s (or 10s, when two attempts fail), and the upstream provider's status page is green.
A latency histogram with a spike parked at 5 seconds and nothing between 200ms and 5s is a resolver timeout, not a network slowdown.
What does ndots:5 actually do to every outbound hostname?
The ndots option says: if a name contains fewer than this many dots, treat it as relative and try it against each entry in search first, before trying it as an absolute name. Kubernetes injects something like this into pods:
nameserver 10.96.0.10
search myapp.svc.cluster.local svc.cluster.local cluster.local
options ndots:5
api.stripe.com has two dots, which is fewer than five, so the resolver dutifully asks for api.stripe.com.myapp.svc.cluster.local, then api.stripe.com.svc.cluster.local, then api.stripe.com.cluster.local — three guaranteed NXDOMAINs — and only then api.stripe.com. With A and AAAA records that is eight queries to resolve one hostname.
The design intent was convenience: it lets other-service and other-service.other-ns resolve inside the cluster. The cost is that every third-party API call multiplies your DNS traffic, and each of those wasted queries is a chance to hit the 5-second timeout.
You can see the fanout directly:
# inside the container
dig +search +short api.stripe.com # follows search domains like the app does
dig +noall +answer api.stripe.com. # trailing dot = absolute, one lookup
# count what actually goes out
tcpdump -ni any port 53
ndots:5 is not a tuning knob you forgot to set — it is a default that charges you three failed lookups for every external hostname your code resolves.
How do I prove it is DNS and not the upstream API?
Separate name resolution from the request before you change anything. curl will tell you where the time went:
curl -o /dev/null -s -w 'dns=%{time_namelookup} connect=%{time_connect} total=%{time_total}\n' \
https://api.stripe.com/healthcheck
If time_namelookup is ~5.0 and time_connect is ~5.05, the connection was fast and the lookup was not. Run it in a loop a few dozen times — the failures are intermittent by nature, so a single clean run proves nothing.
Two more checks worth having in your notes:
# resolution only, using the same libc path your app uses
time getent hosts api.stripe.com
# what the resolver was actually told to do
cat /etc/resolv.conf
If you run CoreDNS, enable the log plugin temporarily and look for a flood of NXDOMAIN responses ending in .cluster.local for external names. That flood is the fanout, in production, with timestamps.
Reproduce the 5-second stall in a loop with curl -w before touching application timeouts, or you will "fix" it by hiding it behind a retry.
Which fix should you reach for?
These are not equivalent. They differ mostly in blast radius.
| Fix | Where it applies | Cost / caveat |
|---|---|---|
Trailing dot on external hostnames (api.stripe.com.) |
One app, one config value | Surgical and free, but some clients, SDKs, and TLS/SNI paths mishandle the trailing dot |
Pod dnsConfig with ndots: 2
|
One workload | Breaks single-label in-cluster names like redis; use redis.myns.svc.cluster.local instead |
single-request-reopen resolver option |
One workload (glibc only) | Works around the A/AAAA conntrack race, does not reduce query count |
| NodeLocal DNSCache | Whole cluster | Cuts cross-node DNS traffic and the conntrack path; one more component to run and upgrade |
CoreDNS autopath plugin |
Whole cluster | Resolves the search fanout server-side; costs CoreDNS extra memory because it watches pods |
| Scale CoreDNS replicas / raise cache TTL | Whole cluster | Helps capacity, not the fanout; higher TTL slows propagation of real DNS changes |
The per-pod version looks like this:
spec:
dnsConfig:
options:
- name: ndots
value: "2"
- name: single-request-reopen
- name: timeout
value: "2"
- name: attempts
value: "2"
Dropping timeout to 2 does not fix anything, but it converts a 5-second user-visible stall into a 2-second one while you work on the real cause — a reasonable stopgap, not a resolution.
If you want this handled once at the cluster level rather than per workload, NodeLocal DNSCache is the option that pays off without touching application manifests, and as of mid-2026 it remains the standard answer for clusters where DNS latency shows up in user-facing percentiles.
Fix the number of queries first; tuning the timeout only changes how long the symptom lasts.
What changes on Alpine, or in Node?
Two environment-specific details cause most of the "I applied the fix and it still happens" follow-ups.
Alpine images use musl, not glibc. musl's resolver sends the A and AAAA queries concurrently over a single socket, and it does not implement every glibc resolv.conf option — notably single-request-reopen, which is the usual workaround for the concurrent-query race. If you are debugging DNS stalls on Alpine, the behaviour you read about in glibc bug threads may not apply to your container at all. The pragmatic move is to test the same workload on a glibc base image (debian-slim, distroless) before concluding the fix failed.
Node.js has two resolvers. dns.lookup() — which http, fetch, and most SDKs use — calls getaddrinfo through libuv's thread pool, defaulting to 4 threads. Blocked lookups therefore queue up behind each other, so one 5-second stall can stall several unrelated requests. dns.resolve() uses c-ares and skips the thread pool entirely. Raising UV_THREADPOOL_SIZE widens the queue; it does not make the lookups faster.
# a stalled lookup shows up as a blocked thread, not slow CPU
UV_THREADPOOL_SIZE=16 node server.js
In Python, requests and httpx go through getaddrinfo as well, so the same resolver rules apply; the difference is that a blocked worker thread or an async event loop hides the stall in a different place in your traces.
Before blaming your DNS configuration, confirm which libc and which resolver your runtime actually uses — the answer changes which fixes are even implemented.
FAQ
Why do DNS lookups take exactly 5 seconds in Kubernetes?
Because 5 seconds is the glibc resolver's default per-attempt timeout, and a UDP DNS response that never arrives costs the full timeout. Kubernetes makes this more likely by setting ndots:5, which generates three extra cluster-internal lookups for every external hostname.
Does adding a trailing dot to a hostname fix DNS latency?
Yes, for that specific hostname: a trailing dot marks the name as absolute, so the resolver skips the search-domain list and issues one lookup instead of four. Test it carefully, because a few HTTP clients and TLS implementations pass the trailing dot through into the SNI or Host header.
Should I set ndots to 1 or 2 in my pods?
Use 2 if your in-cluster calls use at least one dot (redis.myns); use 1 only if every internal hostname is fully qualified. With ndots:1, single-label names like redis will no longer resolve through the search list.
Bottom line
If latency charts show a cliff at 5 seconds and the upstream is healthy, measure time_namelookup before anything else. For a single noisy service, a trailing dot or a per-pod dnsConfig with ndots: 2 is the smallest change that works; for a cluster where DNS stalls affect several workloads, NodeLocal DNSCache or CoreDNS autopath removes the cause instead of the symptom. Keep timeout:2 in your back pocket as a stopgap that buys time, and remember that on Alpine or in Node the resolver you are tuning may not be the resolver your app is using.
Top comments (2)
One thing worth adding to the fix list: after you cut the fanout, CoreDNS is still on the hot path for the remaining lookups, so check its own health too. The cache hit and miss counters plus forward request duration will tell you whether the cliff moved from the resolver to a slow upstream or a cold cache. I have seen teams fix ndots and then chase a 5-second spike that turned out to be CoreDNS waiting on an upstream forwarder.
Cái vụ ndots:5 này thực sự là một "cơn ác mộng" thầm lặng trong Kubernetes vì nó không gây lỗi ngay mà chỉ làm chậm hệ thống một cách ngẫu nhiên. Mình từng mất cả buổi sáng debug một microservice bị latency tăng vọt mà cứ ngỡ do network congestion, cho đến khi kiểm tra file resolv.conf và thấy nó đang cố resolve mọi request qua một đống search domains vô nghĩa. Một mẹo nhỏ là nếu bạn dùng các service external, hãy thử dùng dấu chấm ở cuối domain (ví dụ Google.) để bypass luôn phần search list này, nó giúp giảm tải đáng kể cho CoreDNS — found it via LabAgent, site: labagent .tech