DEV Community

Sergey Shinder
Sergey Shinder

Posted on

Averaging our percentiles gave us a number nobody had experienced

Support sent us a screen recording of our own checkout page taking nine seconds to respond. Our latency dashboard, open on the wall behind me, said the ninety ninth percentile was one hundred and eighty milliseconds, and had said something like that all week.

Both were accurate readings of what they measured. The dashboard was simply not measuring what its title claimed.

Each pod computed its own percentile in process and exported it as a single number. Forty pods, forty numbers, and the dashboard combined them the way the query editor suggests when a series has many members, which is to take the mean. A mean of percentiles is not a percentile of anything. It has no relationship to the experience of any request. A pod serving a small share of traffic with six seconds of latency shifts a forty way average by a fraction of a second and then disappears into it.

Worse, the error moves in the flattering direction under load. We scale that service from twelve pods to forty at peak, so the number of values being averaged grows at exactly the moment the tail gets worse, and the graph improves while the site degrades.

The real figure, recomputed from histogram buckets summed across every pod and then read at the ninety ninth percentile, was four point one seconds.

The reason we had unaggregatable numbers in the first place is that the client library exported summaries, which calculate quantiles inside each process. Those cannot be merged afterwards. The information needed to merge them was discarded before it ever left the pod, and no query written later can recover it.

We moved to histograms with buckets chosen for this service, summed the buckets across instances and computed the quantile after the sum. Every alert and objective now reads from that aggregated figure. Per pod quantiles stayed, on a separate panel, because they are genuinely the fastest way to find one sick instance. Then we audited every dashboard we own and found three more panels averaging summaries, one of them the latency figure in our monthly reliability review.

A percentile describes a population. Average two of them together and you get a number that describes neither, printed in a large font, updating every fifteen seconds.

– Sergey Shinder

Top comments (0)