Our global API serves requests across several regions and teams. A few years ago we started treating traffic analytics as a first class part of running it, not something we look at only after an incident. The two numbers that ended up mattering most were max requests per second and P99 latency. Here is what we actually do with them, and what they taught us.
Why these two numbers
Average latency lies. If one in ten requests is slow, the average still looks fine and users complain anyway. P99 tells you how the slowest one percent of your traffic behaves. When P99 is healthy, almost everyone is having a good experience. When it is not, you have a problem that affects a real slice of your users.
Max RPS tells you how close you are to the ceiling. Every backend, every gateway route, every downstream service has a limit. Knowing the peak rate you actually handled in the last week is more useful than the capacity number someone wrote on a slide. When the peak keeps climbing toward the limit, you plan. When it spikes past it, you scramble.
We track both per endpoint and per consumer where we can. A global number hides the fact that one tenant's burst is what keeps tripping your alarms.
Where the data comes from
Our APIs run behind gateways, Apigee X and Azure API Management, and both emit rich telemetry. Every request produces timing and status data. We aggregate it into dashboards that slice by endpoint, consumer, region, and time window. Nothing exotic. The work is in keeping the dashboards honest and the alerts tuned.
One thing that surprised me early on: gateway timing and backend timing are different numbers, and the difference is where the interesting bugs live. If the gateway reports a request took 800ms but the backend says it took 300ms, the other 500ms went somewhere in between. That gap is policies, retries, connection setup, or network path. We watch the gap per route now. It catches misconfigurations that no amount of backend monitoring would find.
What P99 taught us about retries
Our first P99 problem turned out to be our own retries. A downstream service was occasionally slow, so clients retried. The retries doubled the load on exactly the requests that were already struggling. The average looked okay. P99 climbed steadily for weeks before anyone noticed.
The fix was boring and effective. We added backoff with jitter to the retry policy at the gateway, and we set a hard timeout so a slow request fails fast instead of dragging on. P99 dropped immediately. The lesson: retries are the easiest way to turn a small latency problem into a large one, and P99 is where you see it first.
What max RPS taught us about capacity
We once had a batch job from a partner team that ran every night and hammered one endpoint at many times the normal rate. It was within their quota, so no alarms fired. But it was eating the headroom we needed for real traffic spikes during the day.
Max RPS made this visible. The nightly peak was clearly out of proportion to daytime traffic. We moved the job to a dedicated rate limit tier with its own quota, and daytime headroom went back to normal. Quotas protect you, but they do not tell you that one consumer is distorting your capacity picture. The peak numbers do.
Alerting on these numbers
We set alerts on sustained P99 degradation per endpoint, not on single spikes. A single slow request is noise. Thirty minutes of elevated P99 is a signal. We also alert on max RPS approaching a configurable fraction of known capacity, which gives us a heads up before the limit is real.
The alerts only work because we keep thresholds per endpoint. A login endpoint and a heavy reporting endpoint have completely different latency profiles. One global threshold is either always firing or never firing. Per-endpoint thresholds take a week to tune and then they stay quiet until they matter.
What I would do differently
I would have started tracking the gateway-to-backend gap sooner. It is the single most diagnostic number we have now, and we found it late. I would also have labeled traffic by purpose earlier: interactive user traffic versus batch jobs versus health checks. Mixed together, batch traffic makes your P99 and RPS numbers tell a story that is not really about your users.
The unglamorous truth is that traffic analytics is mostly plumbing and patience. You collect the numbers, you tune the alerts, you look at the dashboards when they page you. The payoff is that incidents stop being mysteries. When something is slow, you already know which endpoint, which consumer, and which part of the path. That is worth more than any single fix we ever shipped.
Top comments (2)
The per-consumer gap is where the real problems hide. We had one partner tenant whose retry burst never showed in the global average for months, and it only surfaced once we alerted on per-endpoint P99 with jittered backoff in place. It turned a nightly support scramble into a capacity conversation with data everyone trusted.
@challan116ux That matches what we saw. The per-tenant view turns a vague latency complaint into a specific conversation. Jittered backoff plus per-endpoint P99 is the combo that worked for us too.