<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nicolai Bohn</title>
    <description>The latest articles on DEV Community by Nicolai Bohn (@nicolai_bohn_rhesis).</description>
    <link>https://dev.to/nicolai_bohn_rhesis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3842808%2F5b6f6897-8d90-43e4-8a4a-1585c2787923.png</url>
      <title>DEV Community: Nicolai Bohn</title>
      <link>https://dev.to/nicolai_bohn_rhesis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/nicolai_bohn_rhesis"/>
    <language>en</language>
    <item>
      <title>Where AI agent test cases come from: synthetic vs. production data</title>
      <dc:creator>Nicolai Bohn</dc:creator>
      <pubDate>Mon, 05 Oct 2026 14:56:22 +0000</pubDate>
      <link>https://dev.to/nicolai_bohn_rhesis/where-ai-agent-test-cases-come-from-synthetic-vs-production-data-2dff</link>
      <guid>https://dev.to/nicolai_bohn_rhesis/where-ai-agent-test-cases-come-from-synthetic-vs-production-data-2dff</guid>
      <description>&lt;p&gt;Before you can evaluate an agent, you need something to evaluate it against. Which test cases you can get depends mostly on timing. Before launch, everything you test with is something your team came up with. After launch, your users start writing test cases for you, whether you collect them or not.&lt;/p&gt;

&lt;p&gt;This post compares the main sources and shows how the mix changes as your agent matures. There's a walkthrough video at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two sources: synthetic and production data
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgmyghr5nrus8ov9oq75b.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgmyghr5nrus8ov9oq75b.webp" alt="Where your test cases come from: synthetic test cases vs. tests from real conversations" width="800" height="304"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Synthetic test cases&lt;/strong&gt; are the ones you create yourself. You can write them while exploring the agent in a playground, or generate them from a description of what you want to test. They are available from day one, and they are how you cover reliability (does the agent meet your requirements?) and robustness (does it hold up against jailbreaks, prompt injection and other adversarial inputs?).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tests from production data&lt;/strong&gt; come from conversations real users had with your agent. They include the edge cases nobody on the team thought to write down. You only get them once you are live.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each method is good for
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfx0eb3zi5gxo4tp2s4w.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkfx0eb3zi5gxo4tp2s4w.webp" alt="Pros and cons of playground, synthesizer and real conversations" width="800" height="314"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Playground
&lt;/h3&gt;

&lt;p&gt;The fastest way to start. You chat with your agent, and when it gets something wrong, you save that exchange as a test case. Say your claims assistant tells a customer they have 30 days to file when your policy says 14. That conversation becomes a regression test in a few clicks.&lt;/p&gt;

&lt;p&gt;The limit is you. Coverage only reaches as far as the scenarios you think of, and recording cases one at a time takes a while.&lt;/p&gt;

&lt;h3&gt;
  
  
  Synthesizer
&lt;/h3&gt;

&lt;p&gt;Here you describe what you want to test, for example &lt;em&gt;"customers asking about cancellation deadlines, including attempts to get the agent to promise refunds it can't give"&lt;/em&gt;. You rate a few samples so the generator learns what you mean, and you get back a full test set. The same works for multi-turn simulations, where a simulated user pushes the agent over several turns instead of a single prompt.&lt;/p&gt;

&lt;p&gt;You get broad coverage of your requirements and adversarial inputs before a single user has touched the agent. You still have to define the scope and review samples, though. And generated inputs tend to read cleaner than real ones. Real users make typos and ask several things in one message.&lt;/p&gt;

&lt;h3&gt;
  
  
  Real conversations
&lt;/h3&gt;

&lt;p&gt;Nothing is more realistic than what your users actually typed. Production conversations show you real phrasing and real edge cases, and they tell you what people want from the agent, which is not always what the spec assumed.&lt;/p&gt;

&lt;p&gt;They have two catches. You need to be live, and you usually want to find problems before your users do. And real conversations contain personal data, so they need cleanup before they become test cases. For teams in Europe that is a GDPR question as much as a technical one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Working from a coding agent instead?&lt;/strong&gt; You don't have to go through a UI for any of this. The &lt;a href="https://docs.rhesis.ai/docs/agent-skill" rel="noopener noreferrer"&gt;Rhesis MCP server&lt;/a&gt; exposes the synthesizer and your production traces to Claude Code, Cursor and other MCP clients. Paste a product spec and the agent &lt;a href="https://docs.rhesis.ai/docs/agent-skill/spec" rel="noopener noreferrer"&gt;proposes requirements, metrics and test sets&lt;/a&gt;, creating them once you approve. Or ask it to read recent traces and turn the conversations worth keeping into test cases.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Your test sources shift as the agent matures
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusinwga1h1ti01wogns1.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fusinwga1h1ti01wogns1.webp" alt="Your test sources shift as the agent matures" width="800" height="291"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In practice, most teams end up using all three, in roughly this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Exploring by hand.&lt;/strong&gt; In the first iterations, domain experts use the playground and record what they see. This is where your first expectations come from.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generating at scale.&lt;/strong&gt; Once requirements settle, the synthesizer and simulations carry most of the volume, covering requirements and adversarial inputs systematically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learning from production.&lt;/strong&gt; After launch, real conversations show you what the synthetic sets missed. A failure from production also makes a good seed for new synthetic variations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated where possible.&lt;/strong&gt; Routine checks run automatically. Experts stay in the loop on the hard, novel cases.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  See it in action
&lt;/h2&gt;

&lt;p&gt;In the video, I go through each method in Rhesis, from the first case recorded in the playground to test cases built from production conversations.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/X0Qbq3pPr-8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;Rhesis is open source. You can try it on &lt;a href="https://app.rhesis.ai" rel="noopener noreferrer"&gt;app.rhesis.ai&lt;/a&gt; or self-host it from &lt;a href="https://github.com/rhesis-ai/rhesis" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://rhesis.ai/post/agent-test-case-sources" rel="noopener noreferrer"&gt;Rhesis blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;How do you source test cases for your agents? Mostly synthetic, mostly production, or something else entirely? I'd like to hear in the comments.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>evals</category>
    </item>
    <item>
      <title>Launch week day 1: drive the full AI testing workflow from inside any AI tool</title>
      <dc:creator>Nicolai Bohn</dc:creator>
      <pubDate>Mon, 04 May 2026 17:17:03 +0000</pubDate>
      <link>https://dev.to/nicolai_bohn_rhesis/launch-week-day-1-drive-the-full-ai-testing-workflow-from-inside-any-ai-tool-1b6b</link>
      <guid>https://dev.to/nicolai_bohn_rhesis/launch-week-day-1-drive-the-full-ai-testing-workflow-from-inside-any-ai-tool-1b6b</guid>
      <description>&lt;p&gt;Building an AI agent uses one tool. Testing it uses another. Every iteration cycle ends with you switching to your test platform UI, running tests, inspecting results, then back to your editor to fix what broke. The context switch is small, but it adds up.&lt;/p&gt;

&lt;p&gt;Today we shipped the fix: the &lt;strong&gt;Rhesis Agent Skill&lt;/strong&gt;, day 1 of Rhesis Launch Week.&lt;/p&gt;

&lt;p&gt;If you build LLM agents and use Claude Code, Cursor, Codex, Gemini CLI, or any of 40+ other AI tools, you can now drive the full Rhesis testing workflow from inside the chat where you write the code.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/_YYj98Lu5rU"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;The Agent Skill packages our domain knowledge into a portable skill file that any compatible AI tool can load. Once installed, your AI assistant gains:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Endpoint discovery&lt;/strong&gt; in Quick or Comprehensive mode&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test suite design&lt;/strong&gt; with behaviors, test sets, and metrics&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confirmation guards&lt;/strong&gt; that wait for approval before anything is created&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test execution&lt;/strong&gt; against your endpoints&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure analysis&lt;/strong&gt; with pass/fail summaries and links back to runs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All powered by the Rhesis MCP server (27 tools covering test sets, behaviors, metrics, runs, and OData queries), all in natural language.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install
&lt;/h2&gt;

&lt;p&gt;Single command across all your AI tools:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add rhesis-ai/rhesis &lt;span class="nt"&gt;-g&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CLI detects which AI tools you have installed and configures the skill for each one. Then set your API token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;RHESIS_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;rhs_your_token_here
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Get a token at &lt;a href="https://app.rhesis.ai/tokens" rel="noopener noreferrer"&gt;app.rhesis.ai/tokens&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Code
&lt;/h3&gt;

&lt;p&gt;Claude Code uses a plugin system that bundles the skill and MCP server config together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/plugin marketplace add rhesis-ai/rhesis
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;rhesis@rhesis-ai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Cursor
&lt;/h3&gt;

&lt;p&gt;Add the MCP server to your &lt;code&gt;.cursor/mcp.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"rhesis"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;https://api.rhesis.ai/mcp&amp;gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"headers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"Authorization"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bearer YOUR_RHESIS_API_KEY"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For self-hosted backends, swap &lt;code&gt;https://api.rhesis.ai/mcp&lt;/code&gt; for &lt;code&gt;http://localhost:8080/mcp&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use it
&lt;/h2&gt;

&lt;p&gt;Type something like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Test my support agent on billing scenarios, run it, and rank the failures by severity."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The skill walks the conversation through a 6-step loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Discover&lt;/strong&gt;: explores what your endpoint can do&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan&lt;/strong&gt;: proposes a test suite with behaviors and metrics&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review&lt;/strong&gt;: waits for your approval before creating anything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create&lt;/strong&gt;: builds entities on the platform following the approved plan&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execute&lt;/strong&gt;: runs tests once you confirm&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Analyze&lt;/strong&gt;: surfaces a pass/fail summary, failure patterns, and links back to results&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For ad-hoc operations:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"List my existing test sets."&lt;br&gt;
"Improve the Safety Compliance metric. Make the threshold stricter."&lt;br&gt;
"Compare my last two test runs."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The host agent's native confirmation handles the safety guard, so destructive actions never happen without your sign-off.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>testing</category>
      <category>llmops</category>
    </item>
    <item>
      <title>Issues in conversational AI apps are so obvious right now that even John Oliver felt the need to dedicate a whole episode to the topic: https://www.youtube.com/watch?v=Ykvf3MunGf8

#AISafety #AIEvals #Testing #QA</title>
      <dc:creator>Nicolai Bohn</dc:creator>
      <pubDate>Tue, 28 Apr 2026 09:24:37 +0000</pubDate>
      <link>https://dev.to/nicolai_bohn_rhesis/issues-in-conversational-ai-apps-are-so-obvious-right-now-that-even-john-oliver-felt-the-need-to-3oj8</link>
      <guid>https://dev.to/nicolai_bohn_rhesis/issues-in-conversational-ai-apps-are-so-obvious-right-now-that-even-john-oliver-felt-the-need-to-3oj8</guid>
      <description>&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="https://www.youtube.com/watch?v=Ykvf3MunGf8" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;youtube.com&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;


</description>
    </item>
  </channel>
</rss>
