Latency is the time an API takes to respond to a request — the delay between when the request is sent and when the response is received.
1. Problem
A single "average latency" number hides the experience of your worst-served users and makes capacity/overload decisions impossible. A service can report a perfectly healthy average while a meaningful slice of real users — often the ones on slow networks, hitting a cold cache, or landing on a degraded instance — have a broken experience. You need a metric that tells you how bad the tail of your traffic looks, not just the middle of it.
2. Constraint
- Latency is never normally distributed. It has a long right tail caused by GC pauses, cold starts, lock contention, cache misses, slow downstream calls, and retries. Averages are dominated by the bulk of fast requests and are mathematically insensitive to a worsening tail.
- Consumers of a service (other teams, external clients) need a concrete, published promise of how fast the service responds — not a vague "it's usually fast." This requires an agreed-upon number, not a feeling.
- In a real architecture, a single service rarely acts alone — it fans out to multiple upstream APIs, databases, and third-party services:
Upstream API 1 ──┐
Upstream API 2 ──┼──► our service ──┬──► DB 1
Upstream API 3 ──┘ ├──► DB 2
└──► DB 3
When a request depends on several downstream calls, the overall response is only as fast as the slowest one. This means tail latency compounds across a call graph — a problem a single-hop percentile number doesn't reveal on its own (see 3.5).
- Computing an exact percentile requires sorting every recorded latency value, which doesn't scale once you're measuring millions of requests per minute across many endpoints.
3. Decision
3.1 Define and publish an SLA, backed by an SLO
An SLA (Service Level Agreement) is the formal promise between a service provider and its consumers — it spells out what will be offered and the standard the provider commits to meeting. Underneath that promise sits an SLO (Service Level Objective) — the measurable target (e.g., "p99 < 120ms") that the team actually tracks day to day. The SLA is the external contract; the SLO is the internal number engineering is held to in order to honor it.
3.2 Track percentiles, not the average
A percentile simply answers: "what's the response time for X% of requests?" A pXX value means XX% of requests completed at or below that time — the remaining (100 − XX)% took longer. Reading left to right below, each percentile looks at a progressively smaller, slower slice of traffic:
| Percentile | On 100 requests... | What it tells you | What it represents | When to use it |
|---|---|---|---|---|
| p50 | 50 requests were faster than this value | middle response time | the typical user's experience | baseline performance |
| p90 | 90 requests were faster; 10 were slower | early warning sign | the leading edge of your slow traffic | first signal that something is starting to degrade |
| p95 | 95 requests were faster; 5 were slower | overall quality | the bulk of your users, including the slower ones | the performance budget most teams commit to |
| p99 | 99 requests were faster; only 1 was slower | reliability | your worst-case users | debugging the tail and latency-critical paths |
Practical guidance: start by optimizing for p95 — it's achievable, actionable, and catches real degradation without you chasing every last outlier. Only once p95 is consistently met should you graduate to tightening p99; chasing the tail before the bulk of traffic is stable is usually wasted engineering effort. (And don't stop at p50 alone — it's a useful baseline, but by definition it's blind to the slower half of your traffic, which is exactly what you need to catch before it spreads.)
3.3 Compute percentiles with approximate streaming structures, not sorted arrays
At scale, monitoring tools like Grafana and Prometheus don't sort every raw latency value — they use streaming histogram-based approximations (e.g., Prometheus histograms, HDR Histogram, t-digest) that give a percentile estimate with a small, bounded error at a fraction of the cost.
3.4 Monitor trends, not single data points
- Collect response times per endpoint.
- Alert on p95/p99 trend lines rather than a single instantaneous reading — this catches genuine, sustained degradation and filters out one-off noise and spikes.
- Compare p95/p99 across regions to catch localized/regional performance issues that a global average would mask.
3.5 Account for tail latency amplification in fan-out calls
If your service calls out to multiple downstream dependencies to answer a single request (as in the diagram above), the chance that at least one of them is slow grows quickly with the number of calls. For example, if each downstream call independently has a 1% chance of being "slow" (its own p99), and your service makes 50 such calls to answer one request, the probability that the overall request is slowed by at least one of them is roughly 1 - (0.99)^50 ≈ 39% — far worse than the 1% tail of any individual dependency. This is why a service's end-to-end p99 is often dominated by its worst dependency, not its average one, and why latency budgets are usually allocated per hop rather than assumed to simply "add up."
4. Trade-off
- Percentile-based SLOs require more sophisticated monitoring infrastructure (histogram-aware backends like Prometheus/Grafana) than a single average — accepted because the alternative actively hides the problem you're trying to catch.
- Approximate percentile structures (t-digest, HDR Histogram) introduce a small, bounded estimation error — accepted as negligible compared to the cost of exact computation at production volume.
- Pushing hard for a tight p99 across every dependency in a fan-out architecture is costly (redundant calls, hedged requests, over-provisioned capacity) — accepted selectively, only on the paths where the business impact of a slow tail actually justifies it. Chasing p99.99 everywhere, immediately is the extreme version of this mistake: expensive for most endpoints, with diminishing returns that rarely justify the cost outside of genuinely latency-critical paths.
5. Failure Mode
- Metrics blind spot: dashboards show a stable p50 (and even a stable average) while p95/p99 silently climbs for days. A narrow but real degradation — e.g., one bad database replica serving 5% of traffic — can go undetected until it's wide enough to affect the median too. Mitigation: always alert explicitly on p95/p99, never just p50.
- Noisy single-sample alerting: alerting on one bad percentile reading from a single window causes alert fatigue and desensitizes the team to real issues. Mitigation: alert on trend over a rolling window, not an instantaneous spike.
- Tail latency amplification goes unnoticed: teams monitor the service's own handler time but not each downstream call individually, so the end-to-end p99 looks "mysteriously" worse than any single dependency's reported p99, with no obvious owner to investigate. Mitigation: track and budget latency per hop (see 3.5), not just end-to-end.
Where to look first
When p95/p99 degrades, these are the usual suspects:
- DNS and TLS handshake overhead.
- Cold starts of services/instances.
- Missing or cold cache.
- Slow database queries (missing indexes, lock contention).
- Connection pool exhaustion.
- Slow downstream dependencies.
- Garbage collector pauses.
Top comments (0)