DEV Community

Cover image for Where AI agent test cases come from: synthetic vs. production data
Nicolai Bohn
Nicolai Bohn

Posted on

Where AI agent test cases come from: synthetic vs. production data

Before you can evaluate an agent, you need something to evaluate it against. Which test cases you can get depends mostly on timing. Before launch, everything you test with is something your team came up with. After launch, your users start writing test cases for you, whether you collect them or not.

This post compares the main sources and shows how the mix changes as your agent matures. There's a walkthrough video at the end.

Two sources: synthetic and production data

Where your test cases come from: synthetic test cases vs. tests from real conversations

Synthetic test cases are the ones you create yourself. You can write them while exploring the agent in a playground, or generate them from a description of what you want to test. They are available from day one, and they are how you cover reliability (does the agent meet your requirements?) and robustness (does it hold up against jailbreaks, prompt injection and other adversarial inputs?).

Tests from production data come from conversations real users had with your agent. They include the edge cases nobody on the team thought to write down. You only get them once you are live.

What each method is good for

Pros and cons of playground, synthesizer and real conversations

Playground

The fastest way to start. You chat with your agent, and when it gets something wrong, you save that exchange as a test case. Say your claims assistant tells a customer they have 30 days to file when your policy says 14. That conversation becomes a regression test in a few clicks.

The limit is you. Coverage only reaches as far as the scenarios you think of, and recording cases one at a time takes a while.

Synthesizer

Here you describe what you want to test, for example "customers asking about cancellation deadlines, including attempts to get the agent to promise refunds it can't give". You rate a few samples so the generator learns what you mean, and you get back a full test set. The same works for multi-turn simulations, where a simulated user pushes the agent over several turns instead of a single prompt.

You get broad coverage of your requirements and adversarial inputs before a single user has touched the agent. You still have to define the scope and review samples, though. And generated inputs tend to read cleaner than real ones. Real users make typos and ask several things in one message.

Real conversations

Nothing is more realistic than what your users actually typed. Production conversations show you real phrasing and real edge cases, and they tell you what people want from the agent, which is not always what the spec assumed.

They have two catches. You need to be live, and you usually want to find problems before your users do. And real conversations contain personal data, so they need cleanup before they become test cases. For teams in Europe that is a GDPR question as much as a technical one.

Working from a coding agent instead? You don't have to go through a UI for any of this. The Rhesis MCP server exposes the synthesizer and your production traces to Claude Code, Cursor and other MCP clients. Paste a product spec and the agent proposes requirements, metrics and test sets, creating them once you approve. Or ask it to read recent traces and turn the conversations worth keeping into test cases.

Your test sources shift as the agent matures

Your test sources shift as the agent matures

In practice, most teams end up using all three, in roughly this order:

  1. Exploring by hand. In the first iterations, domain experts use the playground and record what they see. This is where your first expectations come from.
  2. Generating at scale. Once requirements settle, the synthesizer and simulations carry most of the volume, covering requirements and adversarial inputs systematically.
  3. Learning from production. After launch, real conversations show you what the synthetic sets missed. A failure from production also makes a good seed for new synthetic variations.
  4. Automated where possible. Routine checks run automatically. Experts stay in the loop on the hard, novel cases.

See it in action

In the video, I go through each method in Rhesis, from the first case recorded in the playground to test cases built from production conversations.

Rhesis is open source. You can try it on app.rhesis.ai or self-host it from GitHub.


Originally published on the Rhesis blog.

How do you source test cases for your agents? Mostly synthetic, mostly production, or something else entirely? I'd like to hear in the comments.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Something else, for agents whose answers live in structured data: generate the test set from the source instead of writing or harvesting it. For a question-answering agent over SEC filings, I compute each expected answer from the filing's XBRL facts and have a separate checker re-derive it, so nobody hand-writes an expected value. It also includes questions the filings can't answer, which catch the agent that always produces a confident number. On that set (1,882 questions on 402 companies), gpt-5.4-mini scored 0.8% closed book and 65.8% with a filing-reading tool. It's open if you want a known-answer set to run through Rhesis: huggingface.co/datasets/arhancanli... (I built it).