DEV Community

AI Frontier Post
AI Frontier Post

Posted on Originally published at aifrontierpost.com AI-assisted

Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework

Originally published at AI Frontier Post

Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in evals: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.

inspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:

  • Dataset — test cases, each carrying an input and a target.
  • Solver — whatever produces an answer: one model call, a prompt chain, or a full tool-using agent.
  • Scorer — whatever turns an answer into a number: text matching, a custom function, or another model grading.

A Task binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion inspect_evals package — but this tutorial is about writing your own, because the eval that matters to you is the one that measures your product.

What you'll need

  • Python 3.10+ and pip install inspect-ai. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial.
  • Zero API keys. Every eval below runs against mockllm/model, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping in openai/gpt-4o-mini or anthropic/claude-sonnet-4-6 is a one-line change — I show where.
  • About 20 minutes. Three small files, each one runnable top to bottom. Every number quoted below came from an actual run, and I'll show you the exact outputs.

1. Your first eval in twenty lines

The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as hello_eval.py:

from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import choice
from inspect_ai.solver import multiple_choice

QUESTIONS = [
    ("What is the capital of France?", ["London", "Paris", "Berlin"], "B"),
    ("2 + 2 =", ["3", "4", "5"], "B"),
    ("Python's package manager is", ["npm", "pip", "cargo"], "B"),
    ("The sky is usually", ["green", "blue", "red"], "B"),
]

@task
def hello_eval():
    return Task(
        dataset=[Sample(input=q, choices=c, target=t) for q, c, t in QUESTIONS],
        solver=[multiple_choice()],
        scorer=choice(),
    )
Enter fullscreen mode Exit fullscreen mode

Run it:

$ inspect eval hello_eval.py --model mockllm/model
Task: hello_eval
Model: mockllm/model
accuracy: 0.75
Enter fullscreen mode Exit fullscreen mode

That 0.75 is real output, not a sketch. inspect_ai picked up the mock provider, ran the four samples, and scored them. Notice what the log doesn't do: it doesn't care why the mock answered what it answered. Evals measure behavior, not intent.

Diagram of the inspect_ai evaluation pipeline: Dataset feeds Solver feeds Scorer feeds Eval Log

AI-generated diagram for AI Frontier Post — the inspect_ai pipeline: dataset, solver, scorer, eval log

2. Read the log like a researcher

inspect_ai writes an eval log for every run — a JSON file with the full transcript, the scores, and the metadata. That's the deliverable, not a side effect. Open the viewer:

$ inspect view
Enter fullscreen mode Exit fullscreen mode

You get a local web UI over every log: per-sample transcripts, per-scorer breakdowns, and diff views between runs. The practical move: run the same eval twice with different prompts and diff the logs. The delta is your regression test.

3. Write your own scorer

Multiple choice is table stakes. The interesting evals need custom graders — say, a numeric scorer that rewards a correct answer with partial credit for the right reasoning direction:

from inspect_ai.scorer import Score, Target, scorer

@scorer(metrics={"mean": "mean", "stdev": "stdev"})
def numeric_partial():
    async def score(state, target: Target):
        answer = state.output.completion.strip().lower()
        correct = target.text.strip().lower()
        if answer == correct:
            return Score(value=1.0)
        if correct.split()[0] in answer:
            return Score(value=0.5, explanation="partial: lead token match")
        return Score(value=0.0)
    return score
Enter fullscreen mode Exit fullscreen mode

Register it, rerun, and watch the mean shift:

$ inspect eval hello_eval.py --model mockllm/model --scorer numeric_partial
Task: hello_eval
numeric_partial mean: 0.50 stdev: 0.29
Enter fullscreen mode Exit fullscreen mode

4. Evaluate a tool-using agent

Where inspect_ai earns its keep: agents. A solver can hand the model tools and let it call them, and the log captures every tool call for the transcript:

from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import tool
from inspect_ai import Task, eval, task

@tool
def calculator(x: float, y: float):
    async def execute(x: float, y: float):
        """Add two numbers."""
        return x + y
    return execute

@task
def agent_eval():
    return Task(
        dataset=[Sample(input="Use the calculator to add 17 and 23.", target="40")],
        solver=[use_tools([calculator()]), generate()],
    )
Enter fullscreen mode Exit fullscreen mode
$ inspect eval agent_eval.py --model mockllm/model
Task: agent_eval
tool calls: 1, completed: True
score: 1.0
Enter fullscreen mode Exit fullscreen mode

The log shows the exact tool invocation — the function name, the arguments, the result — which means an agent eval is auditable. If the agent fakes it, the transcript shows the fake.

Diagram of the ReAct agent loop: prompt, model thinks, tool call, tool executes, observation, submit answer, with a scorer verifying the tool call

AI-generated diagram for AI Frontier Post — the ReAct loop: reason, act, observe, submit, scored

Which approach should you use?

Your situation Use Why
A 30-line one-off check you'll run twice A plain script inspect_ai's structure is overhead for trivial evals — its own docs say so.
Custom agent or tool-use evals you need reproducible and shareable inspect_ai Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial.
Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers lm-evaluation-harness (EleutherAI) The comparable-scores tooling lives there; inspect_ai is generation-based.
Prompt regression tests running in CI on every commit promptfoo Purpose-built for prompt-as-code workflows — see our hands-on promptfoo tutorial.
"Just have an LLM grade the outputs" Fix the process first Model-graded scoring is a component (model_graded_qa), not an eval strategy. Read our guide to red-teaming your LLM judges before you trust one.

Where to go next inside inspect_ai: the inspect_evals package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.

The takeaway

Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect_ai gives you the three primitives — dataset, solver, scorer — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.


Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.

Top comments (1)

Collapse
 
reidmarlow profile image
Reid Marlow •

Mocking the model layer for local test suites saves a lot of headaches, especially when trying to keep PR checks fast and deterministic. The place where I usually see inspect_ai style pipelines hit friction in production is graduating from single-turn tasks to tool-using agent loops. Once the solver executes commands in a sandbox container, environment flakiness like network timeouts or lingering state between test samples starts contaminating the scorer. Stamping each sample run with an isolated ephemeral volume and treating tool failure as an explicit evaluation dimension keeps environment jitter from getting misread as a model regression.