I have been writing PHP for twenty years. For the last two I have been putting language models into real applications: ticket classifiers, document analysis, conversational assistants.
And there is one question I could not answer well: how do you test this?
The test that does not work
Say you have a service that classifies support tickets. You give it text, it returns JSON with a category, a sentiment and an urgency level. You want a test.
public function testClassifiesAnAccessProblem(): void
{
$result = $this->analyzer->analyze('I cannot access my account');
$this->assertSame('access', $result->category);
}
Looks reasonable. It is not.
First: it is not deterministic. The same prompt returns different text on every run. Today "access", tomorrow "Account access", the day after "access_problem". Your test fails randomly, and a test that fails randomly is worse than no test at all: it teaches the team to ignore red builds.
Second: it costs money. Every run burns tokens. Multiply by the number of tests, by every push, by every developer on the team.
Third: it is slow. Between 200 ms and 3 seconds per call. A suite with 200 AI tests takes ten minutes. And a ten-minute suite stops being run.
Fourth: it needs network access and a production API key in CI. If your provider has an incident, your build turns red without you breaking anything.
And then there is the problem nobody talks about
Those four are annoying. The fifth one wakes you up at three in the morning.
Your classifier returns this:
{"category": "access", "sentiment": "frustrated", "urgency": 4}
And you, being tidy, map it to a typed DTO:
final readonly class TicketDto
{
public function __construct(
public string $category,
public string $sentiment,
public int $urgency, // ← here
) {}
}
On some random Tuesday, the provider updates the model. Nobody tells you: from their side it is a minor version bump. And the model starts returning this:
{"category": "access", "sentiment": "frustrated", "urgency": "high"}
urgency went from int to string. Your DTO blows up in production.
And you never touched a single line of code.
Here is the important part: no test catches this. The mocks you wrote by hand still return 4, because you froze that value six months ago. Your suite is green while production burns.
What already exists
Before writing anything, I looked at what was out there. This matters, because the gap might not be a gap.
Hand-written mocks. Symfony AI ships InMemoryPlatform and MockPlatformFactory, and they are good. But they return values you made up. They are the right tool for testing your own logic — what happens if the model fails, if it returns empty — and the wrong tool for knowing what the model actually returns. By definition, they can never detect drift.
php-vcr. Excellent library, over three million installs. It records HTTP requests and replays them. The catch is that it matches requests by exact body, and a real prompt carries timestamps, UUIDs, order IDs and RAG context that change on every run. The cassette is invalidated the moment a comma moves.
Eval packages. There are several on Packagist and they are useful, but they answer a different question: is the answer good?. They do not solve determinism, cost or drift. They are complementary.
Promptfoo, DeepEval, Langfuse. The whole mature ecosystem lives in Node and Python. They need a separate runtime and do not plug into PHPUnit.
I searched Packagist for llm cassette record replay. Zero results.
And something I found later
While preparing this article I went back and looked at the Symfony repositories, not just Packagist. There is work in progress:
-
symfony/ai#2129 — a
RecordingProviderthat records cassettes at the provider level - symfony/ai#2128 — cassettes at the HTTP boundary
-
symfony/symfony#63781 — a
RecorderHttpClientin php-vcr style
All three are open drafts at the time of writing, and I only found them after building my own. Honestly, that felt like a good sign rather than a bad one: two core contributors arriving independently at the same conclusion means the problem is real.
They are also solving a different layer. Quoting #2129 directly:
Interactions are matched by a signature over (model, input, options) and consumed FIFO.
Result metadata and token usage are not preserved in this first version.
Exact-signature matching brings back the php-vcr problem: a prompt with a different timestamp does not match. And none of the three does what I needed most — telling me when the provider's behaviour changed.
The idea
A recorder. Like php-vcr, but one that understands prompts.
- Locally it records real responses into versionable JSON files.
- In CI it replays them from disk: no network, no API key, no cost.
- Every night it replays them against the real provider and warns you if anything changed.
$platform = new RecordingPlatform(
inner: new GroqPlatform($apiKey), // your real provider
cassetteDir: __DIR__ . '/cassettes',
mode: Mode::fromEnv(), // record locally, replay in CI
);
Your code does not change: RecordingPlatform implements the same interface. It is a decorator, not a fork.
Three decisions that matter
1. Matching is semantic, not a hash
This is what kills php-vcr for this use case. A real prompt looks like this:
Incident 445566 from 2026-01-10: I cannot access my account.
And on the next run, like this:
Incident 998877 from 2026-07-25: I cannot access my account.
Semantically they are the same case. To a hash they are different universes.
The fix: normalise the volatile noise (<id>, <date>, <uuid>) and compare using cosine similarity over a bag of words. No embeddings, no network, no dependencies. The two prompts above score 0.89 and match.
If you want something stricter, you can declare which parts vary:
new PlaceholderMatcher([
'order_id' => '/ORD-\d+/',
'amount' => '/\d+\.\d{2} ?€/',
]);
Still an exact, deterministic comparison — but immune to what you mark as variable. If something you did not declare changes, it does not match. And that is the correct behaviour.
2. Cassettes get committed, so they get redacted
This one is not negotiable. If cassettes go into git and contain real prompts, they contain real data: emails, phone numbers, national ID numbers, sometimes an API key someone pasted by mistake.
That is why redaction is on by default. Not opt-in.
{
"role": "user",
"content": "I'm <REDACTED:EMAIL>, tel <REDACTED:PHONE>, cannot log in."
}
Secure defaults matter more in development tooling than almost anywhere else. Nobody reads the configuration reference of a dev dependency. If the safe behaviour is something you have to switch on, most people will never switch it on — and the failure is silent until it is public.
Building this I hit a nice bug. The cassette stores <REDACTED:EMAIL>, but the incoming request carries the original email. If you compare the raw prompt against the sanitised cassette, they never match, and it fails precisely on the requests that contain personal data. The fix: run the matching on already-redacted text, so both sides live in the same space.
3. Drift is detected by comparing shape, not text
Back to the int that became a string.
Comparing textual similarity is not enough: {"urgency":4} and {"urgency":"high"} look fairly similar. You have to compare the shape of the JSON: which keys exist and what type each one has.
🔴 CRITICAL sim 0.79 type change in "urgency": int -> string
| new field: "confidence" (float)
🟢 OK sim 1.00 no schema changes
🟡 MEDIUM sim 0.60 no schema changes
One command, one nightly cron, and an issue opened automatically when something moves. It exits with a non-zero code, so it breaks the build.
What it feels like to use
In PHPUnit, a trait and assertions that talk about the problem:
use InteractsWithLlm;
public function testClassifiesAnAccessProblem(): void
{
$platform = $this->recordLlm(GroqPlatform::fromEnv());
$result = $platform->invoke('llama-3.1-8b-instant', [
['role' => 'system', 'content' => 'Classify tickets. Respond with JSON.'],
['role' => 'user', 'content' => 'I cannot access my account.'],
]);
$this->assertNoLiveLlmCalls();
$this->assertLlmJsonShape([
'category' => 'string',
'urgency' => 'int',
], $result);
}
Notice what is not asserted: nowhere does it say the category must be exactly "access". It asserts the contract. An LLM's value is not deterministic, but its shape has to be.
assertNoLiveLlmCalls() is my favourite. Put it in your suite and CI will tell you the day someone adds a test that escapes to the real API.
In Pest, the same thing with expectations:
expect($platform)->toHaveMadeNoLiveCalls();
expect($result)->toBeLlmJson()
->toMatchLlmShape(['category' => 'string', 'urgency' => 'int']);
And in Symfony, the panel
This is where you see it at a glance:
| Metric | Value |
|---|---|
| From cassette | 4 |
| Live calls | 0 |
| Tokens saved | 60 |
| Time saved | 1,280 ms |
Zero API calls. Almost a second and a half saved on a single request. And in the same toolbar you can read that the whole page took 3 ms.
The badge turns red if live calls were made while in replay mode, which usually means a cassette is missing.
The numbers, honestly
Measured against cassettes recorded from Llama 3.1 on Groq, with real latencies of 437, 227 and 235 ms:
| Without llm-vcr | With llm-vcr | |
|---|---|---|
| API calls per run | 3 | 0 |
| Tokens consumed | 426 | 0 |
| Time waiting for the API | ~900 ms | 0 ms |
| Cost in CI | provider-dependent | 0 |
One caveat worth stating: the time saved depends on your real latency. If your provider answers in 200 ms, you save 200 ms per call. The number that matters is not a spectacular multiplier — it is this: zero network and zero tokens in CI, every time, deterministically.
Getting started
Groq gives you a free API key with no credit card, so you can try the whole thing without spending anything.
composer require --dev mikibuilder/llm-vcr
$platform = new RecordingPlatform(
inner: GroqPlatform::fromEnv(),
cassetteDir: __DIR__ . '/cassettes',
mode: Mode::fromEnv(default: Mode::Replay),
);
Run once with LLM_VCR_MODE=record, commit the cassettes, and from then on CI runs for free.
What I learned building it
The tests found bugs I would never have seen. The redacted-text matching one, for instance, only showed up on requests containing personal data. Without a test that deliberately included an email address, it would have shipped.
Actually installing it found the ones the tests could not. I had 112 green tests and the Symfony Profiler panel never registered. The cause: the default value arrived as the unresolved string "%kernel.debug%", and comparing it to true was always false. My tests missed it because every single one passed the parameter explicitly. It only surfaced when I created a fresh Symfony project and installed the package like any user would.
Testing on another operating system found more. The suite took 0.8 seconds on Linux and 25 seconds on Windows: fifteen Symfony kernels writing the container to disk. Rewritten on top of ContainerBuilder, the full suite dropped to 2.5 seconds on the same machine.
The lesson, if there is one: correct code and a usable product are not the same thing, and the only way to tell them apart is to use it.
Where it lives
-
Packagist:
mikibuilder/llm-vcr - GitHub: github.com/MikiBuilder/llm-vcr
- License: MIT
It is version 0.3.x. It works, it has 116 tests and level 9 static analysis, but there is road ahead: embedding-based matching with a cache, streaming and tool call support.
If you try it and something does not fit your case, open an issue. Especially if the way you build prompts breaks the matcher — that is exactly what I need to know.
Twenty years in, PHP is still the language I reach for first. It is worth having the same testing tools everyone else takes for granted.

Top comments (0)