DEV Community

speed engineer
speed engineer

Posted on

Coordinated Omission: Why Your Load Test's p99 Is Lying to You

The problem

Your load test says p99 latency is 12ms. You ship it. Three weeks later, real users are complaining about multi-second hangs, and your dashboards — fed by the same percentile math — say everything's fine. This isn't a monitoring gap. It's a measurement bias baked into almost every load-testing tool you've ever used, and it has a name: coordinated omission.

Here's the setup that causes it. Most load generators work in a closed loop: send a request, wait for the response, record the latency, send the next request. That's how ab, naive wrk scripts, and a lot of homegrown load test harnesses work by default. It feels correct — you're measuring exactly how long each request took.

Why it happens

The problem is what happens during a stall. Say your service has a GC pause, a lock contention spike, or a downstream timeout that freezes request processing for 1 second. In a closed-loop generator targeting 100 req/s, during that 1-second stall you issue exactly one request — the one that's already in flight. It comes back slow, and you record one bad sample: 1000ms.

But that's not what your real traffic does. Real clients don't politely wait their turn. If your service normally handles 100 req/s, a 1-second stall means roughly 100 requests should have arrived and are now queued up behind the one you're measuring. Your closed-loop tool never issued them — it was still blocked waiting on the first one. Those 99 "missing" requests, each of which would have experienced close to a full second of delay, simply never get sampled. They're omitted from your percentile calculation, and the omission is coordinated with the stall itself — it happens exactly when the result would've been ugliest.

The term comes from Gil Tene (Azul Systems, creator of HdrHistogram), originally describing how JVM pause measurement tools lied about GC impact. It generalizes to any closed-loop benchmark of any system: HTTP APIs, databases, queues, LLM inference endpoints, all of it.

The numbers get dramatic fast. Correcting a closed-loop measurement for omission — recomputing what the actual arrival-rate-based percentiles would have been — routinely turns a reported p99 in the single-digit milliseconds into a real p99 north of several hundred milliseconds to a full second. The pain isn't a rounding error. It's an order-of-magnitude lie, concentrated exactly in the tail you care about.

What to do about it

Three concrete fixes, in order of effort:

  1. Use an open-loop (or corrected) load generator. Tools like wrk2, k6 (with arrival-rate executors), Gatling, and Locust's constant-arrival-rate shape issue requests on a fixed schedule regardless of whether the previous one has returned. That matches how production traffic actually behaves — users don't queue politely behind your slow request.

  2. Record "intended" vs. "actual" send time, and correct for it. If you can't switch tools, at minimum log when a request should have been sent versus when it was sent, and compute corrected percentiles. HdrHistogram has built-in support for this via its coordinated-omission correction methods.

  3. Treat any closed-loop percentile number as a lower bound, never a ceiling. If your load test says p99 is 12ms, the honest statement is "p99 is at least 12ms under ideal, non-stalled conditions" — not "p99 is 12ms." That reframing alone changes how much you should trust a green load test before a launch.

This matters more, not less, in the agentic/LLM era: inference endpoints have wildly variable per-request latency (cache hits vs. cold starts, short vs. long generations), which means stalls are frequent and your closed-loop load test is actively hiding your worst-case behavior right when you need to see it.

Key takeaways

  • Closed-loop load generators under-sample exactly the requests that would show your worst latency — the omission is coordinated with the stall, not random.
  • A reported p99 from a closed-loop tool is a floor, not a real number. The true tail can be an order of magnitude worse.
  • Switch to open-loop / constant-arrival-rate tools (wrk2, k6 arrival-rate, Gatling, Locust) or correct for omission with HdrHistogram.
  • This bites hardest on systems with high latency variance — which includes almost every LLM-backed service shipping today.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

To make "floor, not a ceiling" concrete, I simulated one 1-second stall in a 60-second run at 100 req/s with a 5 ms service time. The closed-loop recorder logs about 5,900 samples, one of them slow, so its p99 and p99.9 both read 5 ms. Scoring the same run from each request's intended send time gives 181 of 6,000 requests above 100 ms, p99 around 705 ms and p99.9 around 975 ms. A single stall, and the reported p99 is off by more than 100x.

Two details that tripped me up in practice: the backlog takes a while to drain after the stall ends (here about as long again as the stall itself), so the damage is not limited to the stalled second; and the worst corrected latencies are the earliest arrivals in the stall, which is why a mean hides it even more than the p99 does.

One thing worth checking in your tool list: as far as I know Locust's constant_throughput is per-user pacing, so a stalled user still skips the samples. Does the Locust setup you tested actually backfill them?