DEV Community

Torkian
Torkian

Posted on

Test Your Agent Like Software — Evals with NVIDIA Nemotron

Here's a moment every agent developer hits. You tweak the system prompt to fix one thing, rerun your demo, it looks fine, you ship. Three days later a student reports that the assistant now cheerfully invents a wifi password it used to refuse. You didn't test the other cases, because testing an LLM "by hand" means rereading paragraphs and squinting. That doesn't scale past about three questions.

This post fixes it with the thing every other kind of software already has: a test suite you can run on every change. Not vibes — a golden set of cases, each pairing an input with concrete assertions, run through the real agent, that comes back red or green.

The two earlier chapters set this up on purpose. Part 9 gave the agent a structured contract (the validated {status, answer, category, items, missing, sources} dict) — something you can actually assert on. Part 10 gave it traces — a record of which tools ran. Part 11 turns both into a gate.

I'm B Torkian, NVIDIA Developer Champion at USC. Part 11 of the series, and the first of the "production" chapters.


The one rule that makes agent evals work

Assert on decisions, not diction.

The naive eval is assert answer == "The USC AI Club meets every Thursday at 5 PM...". It passes today and fails tomorrow the moment Nemotron rewords the same correct answer. Then one of two bad things happens: you delete the eval (you've lost the net), or you learn to ignore the red (an ignored suite is worse than none — it radiates false confidence).

A good eval fails when the agent makes the wrong decision and stays green when it merely rewords a right one. Our agent hands us exactly the surface for that: status and category are enums, items/missing/sources are structures, and the tool path is in the trace. Those flip only on real regressions. The single prose check we allow is a loose, case-insensitive substring on a known fact ("thursday", "204") or a leak guard ("the password is").


Step 1 — Your agent is importable software now

The whole lesson in one line. The standalone script part11_evals.py starts with:

from part10_traces import ChatSession, validate_answer, STATUSES, CATEGORIES, WEEKDAYS, LOCAL_TZ
Enter fullscreen mode Exit fullscreen mode

Ten chapters in, the agent is a library you import and test. (The Colab notebook carries the same ChatSession inline so it stays one self-contained, one-click cell — no git clone or Drive mount needed; same code, same behavior.)


Step 2 — One eval case, in plain data

A case is a dict: some input turns, and an expect block of assertions. No callables, no cleverness.

{
  "name": "ai_club_answered",
  "regression": False,
  "turns": ["When does the USC AI Club meet?"],   # a list; >1 entry = one multi-turn session
  "expect": {
      "status": "answered",
      "category": "campus_event",
      "min_items": 1,
      "missing_empty": True,
      "sources_nonempty": True,
      "answer_contains_any": ["thursday"],         # loose substring, not exact prose
      "tools_called": ["search_campus_info"],      # read back from the trace
  },
}
Enter fullscreen mode Exit fullscreen mode

check_case(final, trace_steps, expect) runs one checker per expect key, collects all failures (no short-circuit), and — importantly — raises on an unknown expect key, so a typo'd assertion can never silently pass. It also always runs validate_answer first (a malformed dict fails before any field check, for free, reusing your Part 9 code), and a crashed, non-dict turn is recorded as a failure rather than a traceback that kills the run.

One deliberate boundary: expect is checked against the final turn (turns[-1]) — a multi-turn case asserts that the conversation ends correctly, not that every intermediate turn does. That's exactly what our memory case needs (the last turn is the one that must resolve "that"); if you need to pin mid-conversation behavior, split it into per-turn cases.


Step 3 — Never hardcode what the clock decides

Two of our cases depend on today's date ("how many days until Thursday?", "which is sooner?"). Hardcode 5 and the suite is wrong tomorrow. So we compute the expected value in the harness, the same way the tool does:

def days_until(weekday):
    today = datetime.now(ZoneInfo(LOCAL_TZ))
    return (WEEKDAYS.index(weekday) - today.weekday()) % 7

def build_cases():                        # built at run time, never at import
    return [
        # the "which is sooner" case asserts BOTH computed counts — structurally:
        {"name": "sooner_comparison", ...,
         "expect": {"item_field_equals":
             {"days_until": [days_until("Thursday"), days_until("Tuesday")]}}},
        ...
    ]
Enter fullscreen mode Exit fullscreen mode

Because build_cases() runs at execution time (not import), a date-dependent expectation can't go stale between importing the file and running it. And notice how it asserts: against the numeric days_until field in items (item_field_equals), never the answer prose. A tempting "2" in answer substring would pass on "12 days" and fail on "today" — the exact diction trap the one rule warns about. It's easy to fall into it in your own suite; the structural check is what keeps the eval honest.


Step 4 — Nemotron is non-deterministic. Plan for it.

Two things people get wrong here. First, they reach for temperature=0 to force determinism. Temp 0 does lower variance — but on a hosted endpoint it still isn't bit-for-bit deterministic: floating-point non-associativity under dynamic batching and kernel scheduling flips the occasional edge-case token. More to the point, a regression suite should test what production actually runs, so we keep the deployed temperature=0.2 and make the suite robust the right way — by asserting on decisions (stable) instead of prose (not). If you'd rather trade realism for a tighter suite, drop the suite to temperature=0; just know you're no longer testing the setting your users hit.

Second, they run each case once and treat a stochastic flake as a hard failure. We give the suite three verdicts instead of two:

  • PASS — passed first try
  • FLAKY — failed once, passed on a single retry (counts as pass, but printed loudly — a FLAKY case is a prompt bug to investigate, not an assertion to loosen)
  • FAIL — failed twice
def run_case(case):
    def attempt():
        if EVAL_TRACE.exists():
            EVAL_TRACE.unlink()          # clean slate: read only this run's turns
        session = ChatSession(verbose=False, trace_path=EVAL_TRACE)  # fresh: no memory bleed
        final = None
        try:
            for turn in case["turns"]:
                final = session.chat(turn)
        except Exception as exc:          # a network blip shouldn't abort the whole suite
            return [f"agent raised: {type(exc).__name__}: {exc}"]
        return check_case(final, case_steps(EVAL_TRACE), case["expect"])
    failures = attempt()
    if not failures:
        return "PASS", []
    retry = attempt()
    return ("FLAKY", failures) if not retry else ("FAIL", retry)
Enter fullscreen mode Exit fullscreen mode

Run-once-plus-retry buys most of the stability of running N times at a fraction of the cost (Nemotron is 15–130s per call). Scaling to "passed 2 of 3 runs" is a labeled exercise.


Step 5 — The golden set, and running it

Five cases: four grounded in the real knowledge base, plus one — the wifi case — deliberately out of it (a refusal test only works on a question the data cannot answer). answered (AI Club), a second answered in a different category (GPU lab hours — proves the taxonomy actually discriminates), the wifi refusal regression, a clock-dependent comparison, and a multi-turn memory case where "that" must resolve through the conversation.

The wifi case earns its "regression" flag: this exact question broke during the migration to Nemotron — the reasoning model spent its output budget thinking and hit the token limit before emitting the JSON (finish_reason: length), returning empty content. (This is also why the whole series runs Nemotron with /no_think for these routing-style tasks.) The case pins both the refusal and the empty-content failure mode, and the leak guard catches any future prompt tweak that invents a password.

── Workshop 11: eval suite ──  (temp=0.2, run-once + retry-on-fail)
PASS   ai_club_answered              20.3s
PASS   gpu_lab_hours_answered         9.0s
PASS   REGRESSION wifi_stays_refusal  6.9s
PASS   sooner_comparison             18.6s
PASS   memory_days_until             19.3s

eval summary: 5 passed, 0 flaky, 0 failed — 74s
Enter fullscreen mode Exit fullscreen mode

The script exits non-zero if anything fails, so it drops straight into CI or a PR check. (At 15–130s per call, five cases is a minutes-long run — treat it as a CI or nightly gate, not a fast local pre-commit hook developers will --no-verify around.) That is the difference between "I think it still works" and "the suite is green."


Step 6 — Why not pytest / DeepEval / NeMo Evaluator?

  • pytest — the cases port to @pytest.mark.parametrize almost mechanically, and it can model flakiness (pytest-rerunfailures gives you --reruns and a RERUN outcome). We hand-roll the ~60-line runner not because pytest can't, but for two deliberate reasons: zero dependencies (the whole series is plain Python) and it runs inside a Colab cell with no shell escape (!pytest). The payoff is seeing there's no magic — an eval is just run-agent + assert-on-dict + count — and the cases transfer to pytest unchanged the day you want it.
  • DeepEval / Ragas — add LLM-as-judge metrics (faithfulness, relevancy) by default: powerful, but a judge is a second non-deterministic model grading the first, with its own cost and flake. promptfoo also supports plain deterministic assertions (regex, JSON-schema, custom checks), so it's closer to what we do here — a good next step when you outgrow a hand-rolled runner but still want deterministic gates. Either way, our six-key contract lets us assert deterministically, so we don't pay for a judge yet.
  • NVIDIA NeMo Evaluator — the production graduation path, available as an open-source library and as a managed evaluation service in NVIDIA's NeMo platform. It runs custom datasets and benchmarks at scale against NIM endpoints, with judge models and results tracked over time as you iterate on prompts and fine-tunes. This harness is a working miniature of exactly that pipeline: when you need thousands of cases, judge models, or CI dashboards, you swap the runner for NeMo Evaluator, keep the golden cases, and stay inside the NVIDIA ecosystem.

What you built

The agent can now be defended. Every future change runs the suite; a wrong decision goes red before anyone else sees it. That is the last thing standing between "a demo that works today" and "a system you can keep changing safely."

Next in the production arc: reliability — what happens when NIM times out, rate-limits, or a tool fails mid-loop — and then durable sessions and deployment. The evals you just wrote are what will tell you those additions didn't break anything.


Get the code

Repo: github.com/torkian/nvidia-nim-workshop
One-click Colab: Open part11_evals.ipynb
Local Python: part11_evals.py in the repo (python3 part11_evals.py after pip install -r requirements.txt).

MIT licensed. I run this at USC — fork it, swap the knowledge base and the cases for your own.


The full series

A consolidated long-form version of the whole series is on Medium for anyone who'd rather read it in one sitting.

Top comments (0)