DEV Community

Cover image for Where does the LLM go in the Testing Pyramid?
Nick Hall
Nick Hall

Posted on

Where does the LLM go in the Testing Pyramid?

We all know and love the Test Pyramid, but if your application uses AI, there's a problem: it's now possible to write Service layer tests that are slower and more expensive than the UI tests.

Archibald's Pizzeria has just forked over $X,000 to a slick enterprise vendor to bring them into the 21st century. In addition to ordering pizzas on the website and mobile app, customers can now call "Archie", a virtual chatbot receptionist that takes orders over the phone.

Now look at these three test cases, and tell me where they go in the pyramid:

  1. Use Playwright to open the web page in Chrome, fill out the forms, and order a pepperoni with extra cheese
  2. Call the phone number and say "I'd like a pepperoni with extra cheese"
  3. Call the internal API that the agent uses to process text input and send "I'd like a pepperoni with extra cheese" as a string

Test 1 is a regular old E2E UI test against a webapp and belongs squarely at the top of the pyramid. I believe Test 2 belongs there as well - the phone connection is the "UI" of a voice agent and the only way to test it end-to-end is to call it and talk to it.

Test 3 is where things get tricky. It's an API test, and it avoids the UI being used in test 2, so it ought to go in the middle of the pyramid, right? However, in most cases it will be slower and more expensive than test 1, because it hits an LLM.

Let's alter the pyramid to reflect this. LLMs are as slow or slower than the UI but you call them and assert on their output like an API, so let's put them in between:

Testing around the LLM

We now have two service layers: tests that use the LLM and tests that avoid it somehow. Sounds like we need a way to mock the LLM!

The good news is, the mock should be technically straightforward to implement, because the whole point of an inference model is to infer what the user wants, and we already know what an automated test wants. Consider the pizza example. If Archibald's only sells two kinds of pizza, the mock LLM is a simple mapping of user intent to tool calls:

function mockLLM(userInput, pizzaTool) {
  if (userInput.includes("pepperoni")) pizzaTool.orderPepperoni();
  if (userInput.includes("hawaiian")) pizzaTool.orderHawaiian();
}
Enter fullscreen mode Exit fullscreen mode

The bad news is, it's pure technical debt. It will have to be updated every time the menu changes, let alone the agent's functional behavior. It will also be complex, assuming your real agent has inference decisions that lead to tree decisions that open more inference decisions. The mock LLM will end up looking like a traditional IVR conversational script tree - which is what the AI was supposed to replace!

Implementing the LLM testing layer

So, if mocking the LLM gets you fast/free tests at the cost of tech debt, any tests that go through the LLM had better be tech debt free. How do we ensure that?

The answer, I believe, is very clear: use an LLM to orchestrate the test call and judge the outcome. Remember the definition of inference: if the model's purpose is to figure out how to do what the user wants, then "what the user wants" is the test case. Using an LLM to test your LLM allows you to write test cases like this:

callArchie("You are John Smith of 1234 Main St and
you want to order a pepperoni pizza. Verify that
the agent confirms your selection, tells you the
price, remembers to charge sales tax, and closes
the call by thanking you.")
Enter fullscreen mode Exit fullscreen mode

This test is absolutely zero percent tech debt. You could change IVA vendors, change LLM models, rewrite the whole stack from Java to C# and then back to Java again, and this test would stand, as is.

What you need to build to get there is that callArchie function, which is non-trivial but certainly doable. It needs to orchestrate a multi-turn conversation between your agent and another model (or the same one with different prompts), and then pass the transcript to an LLM-as-judge to evaluate. The script you can probably one-shot, but plan to spend an afternoon getting the prompts dialed in.

Back to the pyramid

We've split the middle layer of the pyramid into "mocks the LLM" and "calls the LLM", and we've discussed what you need to build once and what you need to maintain for each kind. I believe the easiest way to decide which tests go where is to consider each test as a trade-off between maintenance cost vs. execution cost.

How you decide which test goes where is up to you. To an old-school SDET, an API test suite that costs a dollar to run sounds insanely expensive, I know. But tech debt is costly too; if the mock LLM requires even an hour a month to keep it synced to the application's functional behavior, and you only run your tests a few times a week, the mock might not be worth it.

A note about E2E testing

Before we conclude, it's worth revisiting the top of the pyramid to point out that if you end up building the LLM-as-judge service layer test, you can have a true end-to-end test by running the exact same test through a phone line. How you do that and how hard it is depends entirely on what your voice platform exposes:

  • If the platform has an explicit test route that uses real calls, use it. But verify that it's a true test (meaning, figure out what scenarios allow the test to pass even though your phone agent isn't working).
  • If your platform requires building an outbound agent to test inbound agents, and you don't use outbound agents for anything else, consider how much vendor lock-in you're incurring before going all-in.
  • If you must build it yourself, scope carefully before committing. If your coding agent tells you to connect speech-to-text and text-to-speech models to your LLM, consider how much bandwidth would be moving through your test harness. And start working on vendor approval immediately! It may take longer than the coding (telephony vendors deal with an incredible amount of fraud, so someone will need to scan their face and their government ID, and your Legal team may not like their T&Cs).
  • (shameless plug) If you want to spend money to save dev time, consider VoiceGremlin, a SaaS tool I built to solve this specific problem. VoiceGremlin wraps the telephony, audio streaming, calling LLM, and judge LLM in to an API. It takes a string like "Verify the agent does XYZ", calls your external number, executes the test, and gives you a pass/fail plus the transcript. It's vendor agnostic, un-opinionated, and easy to trial, so it's worth a look if a paid tool makes sense for your needs.

Conclusions

From my experience testing voice agents and LLM-powered applications, deciding whether the LLM "counts" as part of the UI layer or the service layer depends entirely on how easily you can mock the LLM and how well you can test your service layer without going through it.

  • If you're going to pay the upfront and ongoing cost of a mocked LLM, you should use it for almost everything. After that, you may not need any service layer LLM tests at all, even if your platform exposes an API for it. Your E2E test (you do have one, right? You probably shouldn't skip this, even if it's a single "is it working at all" test) already covers the LLM path.
  • If you don't want to commit to building and maintaining a mock LLM and you're going to build API tests that call your LLM, you might as well call it several times and get all the benefits of LLM-as-judge, the biggest of which is test cases that are exceedingly easy to write and maintain.

Further reading

The Practical Test Pyramid - a deeper dive on the original model
Writing effective Voice Agent tests - how to phrase test goals for the inference layer

Top comments (0)