<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Madhulika Maddipudi</title>
    <description>The latest articles on DEV Community by Madhulika Maddipudi (@madhulika).</description>
    <link>https://dev.to/madhulika</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148629%2F1272dd84-e2ed-4a09-bcfe-c049ad4eeb12.jpeg</url>
      <title>DEV Community: Madhulika Maddipudi</title>
      <link>https://dev.to/madhulika</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/madhulika"/>
    <language>en</language>
    <item>
      <title>What Makes LLM Testing Different?</title>
      <dc:creator>Madhulika Maddipudi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 06:14:47 +0000</pubDate>
      <link>https://dev.to/madhulika/what-makes-llm-testing-different-51cc</link>
      <guid>https://dev.to/madhulika/what-makes-llm-testing-different-51cc</guid>
      <description>&lt;p&gt;Traditional automated testing is usually straightforward: provide an input, define the expected output, and compare the two. If I am testing a function that calculates a discount, the same input should consistently produce the same result. That makes assertions simple.&lt;/p&gt;

&lt;p&gt;LLM applications behave differently. Imagine we’re testing an employee assistant against a policy that says employees can carry over a maximum of 5 vacation days. We ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How many unused vacation days can I carry over?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model could answer “You can carry over up to five unused vacation days” or “The maximum vacation carryover is five days.” Both are correct, even though the strings are completely different. This makes an assertion like the following too restrictive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You can carry over up to five unused vacation days.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ask_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How many unused vacation days can I carry over?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;expected&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test may fail simply because the model chose different words. Instead, we can start by validating the things that are deterministic and then evaluate the semantic quality separately.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_vacation_policy&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;question&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How many unused vacation days can I carry over?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Employees may carry over a maximum of
    5 unused vacation days into the next year.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ask_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Deterministic checks
&lt;/span&gt;    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;strip&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;five&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# Semantic checks
&lt;/span&gt;    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;evaluate_relevance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="nf"&gt;evaluate_groundedness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.8&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where LLM testing starts to look different. The first few assertions behave like normal automated tests. But relevance asks whether the response actually answers the user’s question, while groundedness asks whether the answer is supported by the supplied policy.&lt;/p&gt;

&lt;p&gt;Suppose our test suddenly receives “Employees can carry over 10 vacation days.” The API might still return 200, the response might be perfectly formatted, and latency might be excellent—but the application is wrong. We now need to determine whether retrieval supplied an outdated policy or whether the correct five-day policy was retrieved and the LLM hallucinated the number ten.&lt;/p&gt;

&lt;p&gt;That’s why I don’t think LLM testing replaces traditional testing. It adds another layer to it. We still test APIs, schemas, permissions, latency, tool calls, and integrations deterministically. But for generated responses, we also need to evaluate qualities such as relevance, groundedness, hallucination, and consistency.&lt;/p&gt;

&lt;p&gt;A useful way to think about the difference is:&lt;/p&gt;

&lt;p&gt;Traditional testing asks: &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Did I get the expected output?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;LLM testing asks: &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Is this output correct and trustworthy, even if it isn’t exactly what I expected?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That small change in the definition of “correct” is what makes testing LLM applications fundamentally different.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>llm</category>
      <category>testing</category>
    </item>
  </channel>
</rss>
