I used to treat the load test as the finish line. Run k6 against staging, watch p95 stay under target, screenshot the graph, drop it in the release channel, and ship!
That ritual felt like engineering, but mostly it was a permission slip.
The incident that broke the habit
We had a notification service that comfortably handled 3x our peak in tests. One evening, with completely normal traffic, it went down. Nothing spiked. Our email provider just started answering in 900ms instead of 150ms. Workers held each job longer, the queue backed up, our retry logic piled on top of the slow calls, and within minutes the whole thing was underwater.
The load test was never wrong. It just answered a question production was never going to ask.
What's the difference between performance testing and performance engineering?
Performance testing tells you whether a system hits a number under conditions you chose. Performance engineering is keeping it reliable under the conditions you didn't choose, for as long as it runs. That's the heart of the performance engineering vs performance testing debate, and once it clicked for me, I stopped reading a green load test as proof of anything beyond that day.
What changed on our team
We started from failure. Before building anything, we listed every dependency on the critical path and wrote down what happens when it gets slow, not just when it goes down because slow is the one that hurts.
We set up tracing properly, so when latency climbs, we can see which hop is responsible instead of guessing from a single average. We also stopped looking at averages. p99 is usually where your angry users live.
We added fallbacks and actually test them. If the template service is struggling, we send a plain-text email instead of nothing. Once a sprint, someone injects latency into a dependency in staging while the team watches, so we learn how things break on a regular afternoon rather than at midnight.
And every service got an error budget. Burn through it, and feature work pauses until stability is back. It sounds somewhat bureaucratic. But in practice, it ended a lot of arguments.
Don't make it one person's job
The worst version of this is a team that everyone else throws problems over the wall to. The people writing the retry logic are the ones deciding how the system fails, so they need to own it.
None of this replaces load testing. It just stops pretending one test run tells you how a system will behave in the upcoming months.
Release day is easy. It's every day after that you have to engineer for.
Top comments (0)