DEV Community

Cover image for How We Test LLM Features So They Don't Regress in Production
Lycore Development
Lycore Development

Posted on

How We Test LLM Features So They Don't Regress in Production

We shipped an LLM-powered classification feature for a client last year. It worked well. Three weeks later, after a routine prompt tweak, it started miscategorising a specific edge case — one that the team had explicitly tested for during development. Nobody noticed for four days.

The problem was not the prompt change. The problem was that we had no automated check that would have caught it. Our test suite confirmed that the endpoint returned a 200 and that the response was valid JSON. It said nothing about whether the classification was actually correct.

This post is about how we fixed that, and what our eval stack looks like now across Django projects with LLM features.


Why standard tests are not enough for LLM features

Unit tests and integration tests are built around determinism. You give a function an input, you assert the output matches what you expect. LLMs are not deterministic in that way — the same prompt can return slightly different wording on different runs, and what counts as "correct" is often a matter of degree rather than exact match.

This creates a gap. You can test that the LLM call succeeds. You cannot easily test that it produced the right answer — at least not with the same tools you use for the rest of your application.

The gap is where regressions hide.


What we mean by evals

An eval is a test that checks the quality of an LLM output, not just its shape. Evals come in a few forms:

Exact match — the output should contain a specific string, or match a specific value. Works well for classification, extraction, and structured outputs.

Model-graded — you use a second LLM call to judge whether the output meets a criterion. More flexible, handles natural language outputs, but adds cost and latency to your test suite.

Human-labelled golden sets — a curated set of inputs with known-correct outputs, maintained by a human. The highest signal, the most expensive to build and maintain.

We use all three, at different layers.


The practical setup in Django

Our eval runner is a Django management command that runs against a fixed dataset we maintain in a JSON fixture file.

# management/commands/run_evals.py
import json
from django.core.management.base import BaseCommand
from django.conf import settings
from myapp.llm import classify_support_ticket


EVAL_DATASET = settings.BASE_DIR / "evals" / "support_classification.json"


class Command(BaseCommand):
    help = "Run LLM eval suite against the golden dataset"

    def handle(self, *args, **options):
        with open(EVAL_DATASET) as f:
            cases = json.load(f)

        passed = 0
        failed = 0
        failures = []

        for case in cases:
            result = classify_support_ticket(case["input"])
            expected = case["expected_label"]

            if result.label == expected:
                passed += 1
            else:
                failed += 1
                failures.append({
                    "input": case["input"][:100],
                    "expected": expected,
                    "got": result.label,
                    "confidence": result.confidence,
                })

        total = passed + failed
        pass_rate = (passed / total) * 100

        self.stdout.write(f"\nEval results: {passed}/{total} passed ({pass_rate:.1f}%)")

        if failures:
            self.stdout.write("\nFailures:")
            for f in failures:
                self.stdout.write(
                    f"  [{f['expected']} → {f['got']}] \"{f['input']}...\""
                )

        if pass_rate < 90.0:
            raise SystemExit(
                f"Eval pass rate {pass_rate:.1f}% is below threshold (90%). "
                "Review prompt or model before deploying."
            )
Enter fullscreen mode Exit fullscreen mode

The golden dataset looks like this:

[
  {
    "input": "My invoice shows the wrong amount and I need a refund immediately",
    "expected_label": "billing",
    "notes": "Clear billing intent despite urgency language"
  },
  {
    "input": "I cannot log into my account, it says my password is wrong",
    "expected_label": "account_access",
    "notes": "Standard auth issue"
  },
  {
    "input": "The app crashes every time I try to export to PDF on iOS 17",
    "expected_label": "bug_report",
    "notes": "Platform-specific bug, not a feature request"
  }
]
Enter fullscreen mode Exit fullscreen mode

This runs in CI on every PR that touches the LLM integration, and nightly against production to catch model drift. If the pass rate drops below 90%, the deploy is blocked.


Model-graded evals for open-ended outputs

For features where the output is prose — summaries, draft emails, explanations — exact match does not work. We use a lightweight model-graded eval instead:

from openai import OpenAI

client = OpenAI()


def grade_summary(original_text: str, summary: str, criteria: list[str]) -> dict:
    criteria_list = "\n".join(f"- {c}" for c in criteria)

    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {
                "role": "system",
                "content": (
                    "You are evaluating AI-generated summaries. "
                    "For each criterion, respond with PASS or FAIL and one sentence of reasoning. "
                    "Be strict — partial credit is not given."
                ),
            },
            {
                "role": "user",
                "content": (
                    f"Original text:\n{original_text}\n\n"
                    f"Summary to evaluate:\n{summary}\n\n"
                    f"Criteria:\n{criteria_list}"
                ),
            },
        ],
    )

    raw = response.choices[0].message.content
    results = {}
    for criterion in criteria:
        results[criterion] = "PASS" in raw.upper()

    passed = sum(results.values())
    return {
        "results": results,
        "score": passed / len(criteria),
        "raw_response": raw,
    }
Enter fullscreen mode Exit fullscreen mode

We use gpt-4o-mini for grading — it is significantly cheaper than the model that generated the output, and accurate enough for pass/fail judgements on well-defined criteria.


What we track over time

Running evals once is useful. Running them consistently and tracking the results over time is where the real value is. We log every eval run to a simple Django model:

class EvalRun(models.Model):
    feature = models.CharField(max_length=100)
    model_used = models.CharField(max_length=100)
    prompt_hash = models.CharField(max_length=64)
    pass_rate = models.FloatField()
    cases_run = models.IntegerField()
    failures = models.JSONField(default=list)
    run_at = models.DateTimeField(auto_now_add=True)
    triggered_by = models.CharField(max_length=50)  # "ci", "nightly", "manual"
Enter fullscreen mode Exit fullscreen mode

When the pass rate on a nightly run drops by more than 3 points compared to the previous week, it pages the on-call engineer. That threshold catches model drift — cases where the underlying model behaviour has shifted without any code change on our side.


The honest summary

LLM evals are not complicated to set up. A JSON fixture file and a management command that compares output labels against expected labels is enough to catch most regressions. The hard part is maintaining the golden dataset — adding new cases when you find edge cases in production, and reviewing failures rather than just rerunning until they pass.

If you ship LLM features without evals, you are relying on manual testing and user complaints to catch regressions. That is a slow feedback loop for something that can break silently and affect every user.

Start with exact-match evals on your most critical features. Add a model-graded layer where the output is prose. Run both in CI. The eval suite that catches one production regression pays for itself immediately.


Lycore builds production AI systems for businesses — LLM integrations, agents, RAG pipelines, and custom AI applications on Django, React, Flutter, and .NET. Get in touch if you want to talk through your use case.

Top comments (1)

Collapse
 
carllowman profile image
SerpSpur •

One practical approach is to test LLM features like any other production-critical system: keep a fixed evaluation set, test edge cases, compare outputs across model/version changes, and monitor real-world failures after release. LLM behavior can shift even when the code hasn't changed, so regression testing needs to cover both quality and consistency.

I also think SerpSpur is an interesting example of where this matters—especially for AI visibility, mentions, and citations. If AI-generated results change over time, tracking those changes can help distinguish a genuine improvement from a model or query fluctuation.