In May we moved our account API onto a service mesh, and overnight its request rate fell by forty one percent. The traffic drop alert fired at two in the morning. Users were fine. Nothing had been lost that had ever been a user.
The mesh had moved health checks to a separate port that bypasses our metrics middleware. So I counted what had been going through it. Kubelet readiness and liveness probes every five seconds on each of twenty four pods. Load balancer checks from three zones every ten seconds per target. An uptime service polling from six locations. And a dependency checker owned by another team, calling the status endpoint of every service in the company every fifteen seconds. All of it arrived on the main port, passed through the same middleware that recorded request count, latency and status code for our objectives, and was answered with a 200 in about two milliseconds.
At the daytime peak that was around fifteen percent of our requests. Overnight it was seventy. Across a whole day it was four in ten.
Every number we reported was diluted by traffic that could not fail. The error rate objective was a fraction whose denominator was forty percent machines asking whether we were alive. During a partial outage in March, users saw seven point eight percent errors while our dashboard showed four point six, below the five percent burn threshold that pages, and we heard about it from a customer twenty six minutes in. Latency percentiles were flattered the same way, by fast requests nobody was waiting for.
Probes now live on a separate admin port that is not instrumented as traffic. Every request metric carries a caller class label with three values, user, internal and synthetic, and objectives are computed on user alone. Synthetic checks still alert, as their own signal. A panel shows the share of non user traffic per service, so the next dilution shows before it costs anything. Then we recomputed the last quarter from the load balancer logs with probes excluded. We had used almost twice the error budget we reported, and missed the objective in two months of three.
An objective is a ratio, and we had spent two years inspecting the top half of it. Decide who counts in the bottom half, or everything that can reach your endpoint will decide for you.
– Sergey Shinder
Top comments (2)
Do not follow external links, this is a phishing scam. DEV.to uses Sloan for automated messaging.