A mock earns its place in a test suite the same way any other piece of infrastructure does: by being right often enough that nobody thinks to question it. That's also exactly what makes a bad mock dangerous. Nobody double-checks the thing they've stopped noticing. A mock that quietly stopped matching reality six months ago doesn't announce itself. It just keeps returning green, right up until the gap it's been hiding shows up somewhere much more expensive than a test run.
Worth being specific about what separates a mock that's actually doing its job from one that's just occupying the space where a real dependency test should be.
It reflects what the dependency actually does, not what someone assumed it does
The most common origin story for a mock is someone reading documentation, or remembering roughly how an API behaves, and writing a response that seems reasonable. That's a fine starting point, and it's also the exact spot where trustworthiness starts to erode, because documentation and memory both drift from reality in ways nobody notices until something breaks.
A trustworthy mock is built from observed behavior: an actual response the dependency actually returned, at some point, under real conditions. This matters more than it sounds, because real APIs are full of small inconsistencies that no one writes into a spec — a field that's sometimes null and sometimes just absent, a timestamp format that varies by endpoint, an error response shaped differently than the success response's documentation implies. A mock built from what actually happened carries all of that. A mock built from what someone assumed happened carries none of it.
It includes failure, not just success
Most hand-built mocks are optimized for the case that's easiest to imagine: the request that works. That's understandable, since the success case is usually what the person writing the mock was actually trying to test in the first place. It's also the least useful part of a mock to get right, because the success case is rarely where production bugs live.
The valuable part of a mock is what happens when the dependency times out, returns a malformed payload, hits a rate limit, or fails in some specific way particular to that service. These are the cases most likely to be missing from a mock, precisely because they require someone to think of them in advance and deliberately write them in. A mock's trustworthiness is really a question of how much of the dependency's actual behavior it captures, and failure modes are usually the majority of that behavior that goes uncaptured.
It doesn't quietly go stale
A mock is a snapshot. The moment it's written, it starts drifting away from whatever the real dependency is doing, at a pace nobody's tracking. This is the failure mode that's hardest to notice, because a stale mock doesn't look broken. It looks exactly like it did the day it was written. Tests built on it keep passing, confidence keeps building, and the actual dependency keeps changing underneath, unmonitored.
The trustworthy version of this isn't "write a good mock once." It's having a clear answer to who refreshes it, how often, and what would actually trigger that refresh - a scheduled re-check, a contract version bump, a captured-traffic pipeline that updates automatically rather than depending on someone remembering. Mocks that get treated as permanent fixtures are the ones most likely to be silently wrong by the time anyone looks at them again.
It's scoped to what it's actually simulating
A trustworthy mock does one job clearly: standing in for a specific dependency's specific behavior. Mocks lose credibility when they start absorbing logic that belongs somewhere else - business rules baked into a mock's response, conditional behavior that exists only to make a particular test pass, special cases added over time by different people for different reasons until nobody's sure what the mock is actually simulating anymore.
The cleanest mocks are the ones that map directly to something a real dependency did, with nothing added and nothing simplified away. The moment a mock needs its own internal comments explaining why it behaves the way it does, that's usually a sign it's drifted from representing the dependency into representing whatever made a specific test pass at some point.
It's reproducible
A mock that behaves differently between runs, whether from race conditions, uncontrolled randomness, or hidden state carried over from a previous test, undermines the entire reason mocks exist. The whole premise of replacing a real dependency with a mock is predictability: the same input should produce the same output, every time, regardless of what ran before it or what environment it's running in. A flaky mock is arguably worse than no mock at all, because it produces the appearance of determinism while quietly not providing it, and debugging a flaky test caused by a flaky mock is a particularly frustrating way to lose an afternoon.
Where this actually comes from in practice
Most of these properties point in the same direction: a mock's trustworthiness comes from how closely it's tied to real, observed behavior, and how deliberately that connection is maintained over time. Hand-written mocks can achieve all of this, but only with real ongoing discipline, since every property above requires someone to keep paying attention after the mock is first written.
This is the specific gap that dependency mocking as a tooling category exists to close - generating and maintaining mocks from real recorded behavior instead of leaving that entirely to manual upkeep, so freshness and failure-case coverage come from the process itself rather than from whether someone remembered to go back and update a stub file.
The actual test
A simple way to check whether a mock in your suite still deserves trust: could you explain, right now, where its behavior came from and how recently that source was checked. If the honest answer is "someone wrote it a while ago based on the docs," that's not automatically wrong, but it's worth treating as a question rather than an assumption. The mocks worth trusting are the ones where that question has a specific, recent, confident answer.
Top comments (0)