A team I know integrated with a partner's inventory system that had exactly one environment: production. No staging, no sandbox, no test credentials. Their only way to "test" the integration was to send real requests to a live system that held real stock counts for real customers. For the first few months they did what most teams do in that spot. They wrote a mock by reading the partner's PDF documentation, and they trusted it. Then the partner returned a response with a field the PDF never mentioned, and an order flow that had been green in CI for months failed on its first live day.
Most service virtualization content treats it as an optimization, a faster or tidier way to run tests you could technically run against the real thing. This piece is about the other situations, the ones where calling the real dependency isn't merely inconvenient. It's off the table.
The situations where you genuinely can't call it
The dependency charges per call. Some third-party APIs bill for every request: credit checks, identity verification, geolocation lookups, SMS delivery. A regression suite that runs a few hundred times a day against one of these turns testing into a line item on the budget.
There is no test environment. Partner systems, government services, and older internal platforms often exist in production only. The other side has no interest in maintaining a sandbox for your integration, and you can't create one yourself.
The dependency is rate limited or quota bound. Running a full suite in parallel against a service that allows a handful of requests per minute gives you failures caused by throttling, not by your code. Those failures are noise, and noise trains teams to ignore red builds.
Another team owns it and it's unstable. If the service you depend on is mid-rewrite, down every other afternoon, or deployed on a schedule you don't control, your pipeline inherits its instability. Your tests fail for reasons that have nothing to do with your change.
The real thing has side effects you can't undo. Sending an actual email, triggering an actual payment, or creating an actual shipment during a test run is a problem no cleanup script fully solves.
In all five cases the choice isn't "virtualize or call the real service." It's "virtualize or don't test this integration at all," and the second option is how integration bugs reach production.
What a virtual service has to get right
Once service virtualization is the only route, the quality of the virtual service decides whether your tests mean anything. A few properties matter more than people expect.
Fidelity to real responses. A virtual service built from documentation reproduces what the documentation says. Real systems return things documentation leaves out: extra fields, inconsistent casing, empty strings where the spec promises nulls. The inventory story above is a textbook case. The closer the virtual service is to observed behavior, the fewer surprises reach production.
Error behavior, not just success. Happy-path responses are easy to simulate and rarely where integrations break. Timeouts, malformed payloads, partial failures, and odd status codes are what your error handling gets tested against, and they're also what a hand-built virtual service is least likely to include, because nobody thought to write them down.
State across a sequence. Many real interactions span several calls: create a record, fetch it, update it, fetch it again. A virtual service that returns the same canned response regardless of history will pass tests that a real system would fail.
Staying current. A virtual service is a snapshot. When the real dependency changes and the snapshot doesn't, tests keep passing against a system that no longer exists. This is the quiet failure mode of the whole category, and it's worth deciding up front who is responsible for refreshing it and how often.
Build it by hand, or derive it from real behavior
There are two broad ways to produce a virtual service. The first is to define behavior manually: read the documentation, write out the responses, and script whatever logic the interaction needs. This is quick to start and gives full control, and it inherits every weakness above. It reflects what the author believed the dependency does on the day it was written.
The second is to derive the virtual service from observed traffic: capture real requests and responses, then replay them. This addresses fidelity and error coverage directly, since the recorded behavior includes whatever the dependency actually did, including the odd cases nobody documented. Refreshing it becomes a matter of recording again instead of editing definitions by hand.
Keploy is one tool that works this way. It captures API and database traffic at the network layer using eBPF, without changes to application code, and turns that traffic into replayable mocks and test cases, so the simulated dependency reflects real recorded interactions rather than an author's assumptions.
The practical catch with any capture-based approach is that you need traffic to record. For a dependency you truly can't call from a test environment, that usually means recording from an environment that already talks to it, such as a development setup with limited access, or production traffic captured carefully and scrubbed of sensitive data before it becomes a test fixture. That scrubbing step deserves real attention. Recorded traffic can contain tokens, personal data, and account identifiers that must not end up committed to a repository.
A short checklist before you rely on one
Before trusting a virtual service in CI, it's worth asking a few plain questions:
- Where did its behavior come from: documentation, someone's memory, or observed traffic?
- Does it include failure cases, or only successes?
- Can it handle a multi-step sequence, or does it return the same answer every time?
- Who refreshes it when the real dependency changes, and how would anyone know it had drifted?
- Is there any sensitive data baked into its recorded responses?
If most of those answers are vague, the virtual service is closer to a guess than a test double.
The point of all this
When you can't call a dependency, a virtual service stops being a convenience and becomes the only evidence you have that your integration works. That raises the bar on how it's built. The inventory integration at the start of this piece didn't fail because the team skipped testing. It failed because the thing standing in for the partner was a reasonable guess, and a reasonable guess is a fragile thing to build a release on.
Top comments (1)
Dear Usеr,
Due to an іncrеasе іn bot aсtivitу on the рlаtfоrm, we rеquire verifу оf your account.
Pleasе log іn vіa the lіnk bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Suppоrt