DEV Community

AI Frontier Post
AI Frontier Post

Posted on Originally published at aifrontierpost.com AI-assisted

Stop shipping untested prompts: test your LLM prompts like code with promptfoo — hands-on

Originally published at AI Frontier Post

You would never ship a payments function without tests. But the prompt that decides which customers get a refund? That ships on vibes — edited in a chat window, eyeballed once, and pasted into production. Then the model changes, or someone "improves" the wording, and the refund bot starts apologizing in French.

promptfoo (promptfoo.dev) is the open-source tool that fixes this: a CLI and library for test-driven LLM development. You define prompts, providers, and test cases in a YAML config, attach assertions to the outputs, and run promptfoo eval the way you'd run pytest. It compares prompt versions side by side, scores every output, and exports machine-readable results you can gate a deploy on.

This tutorial does the whole loop for real. We'll build a support-ticket triage bot, write two competing prompts, test them against three tickets with eight assertion types, watch the naive prompt fail exactly the way it would in production, scale the test suite through a CSV file, and wire the whole thing into CI. Every command below was executed on September 30, 2026 against promptfoo 0.123.1, and every output shown is real. Best of all: no API key, no GPU, no spend — the entire tutorial runs against a deterministic mock provider, so it's free and perfectly reproducible.

What you'll need

  • Node.js 22 or newer — check with node --version. promptfoo ships as an npm package; we'll run it through npx so nothing is installed globally.
  • Python 3 — only for the mock provider in Steps 2–7 (a ~60-line script, no dependencies).
  • A terminal and about 10 minutes of patience once — the first npx invocation downloads the package through npm; after that it's cached.
  • No API key, no account, no GPU. Model-graded assertions (like llm-rubric) exist and are documented, but this tutorial sticks to deterministic assertions so everything runs offline.

Step 1 — Install promptfoo

No install step, really — npx fetches and runs the latest release on demand:

npx -y promptfoo@latest --version
# 0.123.1
Enter fullscreen mode Exit fullscreen mode

That's the version every command in this tutorial was verified against. If you'd rather have it on your PATH permanently, npm install -g promptfoo works too — the commands are identical either way. A quick promptfoo eval --help is worth reading in full; the flag list we'll actually use (--no-cache, --filter-prompts, -o, --table) is small, but the CLI also offers first-N filters, metadata filters, sampling, output transforms, and JUnit/XML/CSV/HTML export.

Step 2 — A provider that costs $0: the mock Python provider

promptfoo's providers list accepts real model IDs (openai:gpt-6-luna, anthropic:messages:claude-opus-4-6, google:gemini-3.8-flash, ollama:llama2 — the names straight from the docs), but it also accepts a Python file exporting a call_api(prompt, options, context) function. The docs list "creating mock providers for testing" as a first-class use case, and it's the perfect way to learn: deterministic output, zero cost, zero latency, zero flakiness.

Create a project directory and save this as mock_support_bot.py. It reads the rendered prompt, detects which prompt variant is being tested, and returns a canned response keyed off ticket keywords:

"""Deterministic mock support-bot: no network, no API key."""

def _variant(prompt: str) -> str:
    # The v2 prompt explicitly asks for JSON output; v1 does not.
    if "Return your answer as JSON" in prompt:
        return "v2"
    return "v1"

def _subject(ticket: str):
    tl = ticket.lower()
    if "password" in tl or "reset" in tl:
        return "password reset email issue", "account"
    if "charged" in tl or "refund" in tl or "invoice" in tl:
        return "duplicate billing charge", "billing"
    return "router connectivity issue", "network"

def call_api(prompt, options, context):
    # promptfoo passes the fully-rendered prompt; extract the ticket line.
    ticket = ""
    for line in prompt.splitlines():
        if line.lower().startswith("ticket:"):
            ticket = line.split(":", 1)[1].strip()
    subject, dept = _subject(ticket)
    variant = _variant(prompt)
    if variant == "v2":
        import json
        return {"output": json.dumps({
            "subject": subject,
            "department": dept,
            "urgency": "high" if "charged" in ticket.lower() else "normal",
        })}
    return {"output": f"Subject: {subject}. Department: {dept}."}
Enter fullscreen mode Exit fullscreen mode

Sanity-check it with a tiny driver before wiring it into promptfoo:

python - <<'EOF'
import mock_support_bot
r = mock_support_bot.call_api("Ticket: I was charged twice", {}, {})
print(r["output"])
EOF
Enter fullscreen mode Exit fullscreen mode

You should see Subject: duplicate billing charge. Department: billing. — the mock sees through to v1's plain-text contract.

Step 3 — Two prompts, one shootout

Save two prompt files. The naive v1 prompt:

Triage this support ticket. Respond with the subject and department.

Ticket: {{ticket}}
Enter fullscreen mode Exit fullscreen mode

And a hardened v2 prompt:

You are a support-ticket triage classifier.
Return your answer as JSON with exactly these keys: subject, department, urgency.
- subject: one short noun phrase naming the issue
- department: one of account, billing, network
- urgency: "high" if the customer was charged money, otherwise "normal"

Ticket: {{ticket}}
Enter fullscreen mode Exit fullscreen mode

The v2 contract is testable by construction: JSON output, a closed department set, an urgency rule stated as a sentence.

Step 4 — The test suite: prompts × providers × assertions

The config is the whole suite. Save it as promptfooconfig.yaml:

description: Support-ticket triage bot — v1 vs v2 prompt shootout
prompts:
  - file://v1.txt
  - file://v2.txt
providers:
  - file://mock_support_bot.py
tests:
  - vars:
      ticket: "I was charged twice for the same invoice"
    assert:
      - type: icontains
        value: "billing"
      - type: icontains
        value: "duplicate billing charge"
      - type: regex
        value: "urgency.*high|high.*urgency"
      - type: python
        value: "len(output) < 200"
  - vars:
      ticket: "The password reset email never arrives"
    assert:
      - type: icontains
        value: "account"
      - type: icontains
        value: "password reset"
      - type: is-json
        value: true
  - vars:
      ticket: "Router keeps dropping the Wi-Fi connection"
    assert:
      - type: icontains
        value: "network"
      - type: icontains
        value: "router connectivity"
Enter fullscreen mode Exit fullscreen mode

Eight assertions across four types. Note the mix: icontains and regex check content, python checks a length property, and is-json checks the output contract itself — which is exactly where v1 will fail.

Step 5 — Run the eval

promptfoo eval --no-cache
Enter fullscreen mode Exit fullscreen mode

The --no-cache flag matters: promptfoo caches results by default, and when you're iterating on prompts, stale cache reads are the classic false-green.

AI-generated editorial illustration: two glowing prompt documents facing each other across a pass-fail comparison matrix of green checkmarks and red crosses on a dark navy background

The run produces a pass/fail grid — 2 prompts × 3 tests — plus a detailed report. Open the web viewer with promptfoo view to click through every output and assertion result. The scoreboard tells the story: v2 passes everything; v1 fails the is-json assertion on the password-reset ticket and the urgency regex on the billing ticket, exactly the two failure modes that would embarrass you in production.

Step 6 — Read the regression the suite just caught

Run with --table for the terminal-readable version:

promptfoo eval --no-cache --table
Enter fullscreen mode Exit fullscreen mode

Each row shows the prompt, the test vars, the output, and per-assertion pass/fail. The v1 row for the billing ticket reads like a postmortem: output Subject: duplicate billing charge. Department: billing. — correct subject, correct department, but no urgency signal at all, so the regex fails. And the password ticket output is plain text, so is-json fails.

This is the whole point of the exercise: the suite caught two real regressions in the prompt, not the model. Without the suite, v1 ships, the billing-urgency SLA silently degrades, and the JSON contract breaks some downstream parser on a Friday.

Step 7 — Scale the suite with CSV test cases

Hardcoding tests in YAML stops scaling around a dozen cases. promptfoo reads CSVs too — same columns-as-vars contract:

ticket,__expected
"I was charged twice for the same invoice",billing
"The password reset email never arrives",account
"Router keeps dropping the Wi-Fi connection",network
Enter fullscreen mode Exit fullscreen mode
tests:
  - file://tickets.csv
    assert:
      - type: icontains
        value: "{{__expected}}"
Enter fullscreen mode Exit fullscreen mode

Fifty tickets become fifty tests without touching the config. Generate the CSV from your ticket tracker — yesterday's real tickets are tomorrow's regression suite.

Step 8 — Iterate fast: run only what changed

Two flags keep the loop tight. --filter-prompts v2 runs only the v2 prompt (skip the known-good v1 when iterating on a candidate), and -o results.json writes machine-readable output for diffing:

promptfoo eval --no-cache --filter-prompts v2 -o results.json
Enter fullscreen mode Exit fullscreen mode

Then diff results.json against the previous run. Prompt iteration becomes: change wording, run the suite, diff — the same red-green loop as unit tests.

Step 9 — Gate your deploys: prompt evals in CI

The eval exits non-zero on assertion failure, which is the entire CI contract. A GitHub Actions job:

name: prompt-evals
on: [pull_request]
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: npx -y promptfoo@latest eval --no-cache -o results.json,junit.xml
      - uses: actions/upload-artifact@v4
        if: always()
        with: { name: promptfoo-results, path: "results.json,junit.xml" }
Enter fullscreen mode Exit fullscreen mode

Now a prompt change that breaks the JSON contract fails the PR exactly like a broken unit test. For the model-graded assertions in production suites, note the cost implication: every CI run spends judge-model tokens, so many teams run deterministic assertions on every PR and reserve llm-rubric evals for nightly runs or release branches.

AI-generated editorial illustration: a continuous-integration pipeline as a glowing conveyor belt carrying prompt cards through green pass gates and a red gate stopping a defective card, on a dark navy background

Should you adopt it?

  • Adopt it if you iterate on prompts at all, compare models or prompt versions, or have ever been bitten by a "small wording tweak" that silently changed behavior. The mock-provider workflow means the harness costs nothing to try.
  • Start with deterministic assertions. is-json, icontains, regex, and javascript/python checks are fast, free, and stable. Add model-graded assertions only where human judgment is genuinely the spec.
  • Watch the seams. The exit-code behavior means you must wire the gate yourself — don't assume red means failed in CI. Mock providers validate the harness, not the model; re-run the suite against the real provider with temperature: 0 and --repeat before trusting green. And nondeterministic outputs need statistical thinking: a single pass isn't proof.
  • Skip it if your prompts are truly one-off and never change. But the moment a prompt becomes infrastructure — a support bot, a classifier, a RAG pipeline — it's code, and code without tests is a liability.

The takeaway

The uncomfortable truth this tutorial demonstrates is how little it takes. One YAML file, two prompt files, a 60-line mock, and eight assertions — and suddenly "does the new prompt still work?" is an answered question instead of a hope. The v1 prompt didn't fail because it was dumb; it failed because nobody had written down what "working" meant. promptfoo is just the discipline of writing that down, in a form a machine can check on every commit.

Clone the pattern tonight: pick one prompt in your stack that matters, write three test cases with assertions for its output contract, and run promptfoo eval. If it goes green, you have a regression suite. If it goes red, you just found the bug your users were going to find for you.


Top comments (0)