DEV Community

Cover image for AI Agent Testing: Why a 77% Pass Rate Can Mean 53% in Production
Qasim Parray
Qasim Parray

Posted on Originally published at abrarqasim.com

AI Agent Testing: Why a 77% Pass Rate Can Mean 53% in Production

Short version for the impatient: if your agent passes 77% of your test cases, the chance it passes the same case five times in a row might be closer to 53%. That's the number I want you to carry around. If you want to know where it comes from and what I changed in my own testing because of it, read on.

I ran into this the embarrassing way. Last spring I shipped an invoice-triage agent for a client. Twenty-two test cases, all green, three days in a row. I wrote "stable" in the handover doc. Two weeks later their ops person sent me a screenshot: the same PDF, uploaded twice ten minutes apart, filed under two different vendors. Nothing had changed. No prompt edit, no model update, no new data. The agent just took a different path the second time.

I didn't have a name for that failure until a paper from IBM Research landed on arXiv this week. They call it the consistency gap, and they measured it properly, which I never had.

The number that changed how I run ai agent testing

The paper is Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course by Evelyn Duesterwald and colleagues. The setup is simple enough that I'm annoyed I didn't do it myself. Take a ReAct agent on the AppWorld benchmark using GPT-4.1. Run every task five times. Count two things: the average pass rate per run, and the fraction of tasks that pass all five runs.

Per-run pass rate: 77%. All-five pass rate: 53%. A 24 point gap, on a benchmark, with no environmental noise to blame.

Read that again with your own dashboard in mind. My twenty-two green tests were one sample each. If the underlying agent had a 77% per-run rate, the odds of all twenty-two passing on a given day are small, so in hindsight I probably got lucky on the day I wrote "stable". Or, more likely, my cases were easier than the client's real invoices and the per-run rate was higher, which hides the gap without closing it.

The gap matters more than the pass rate for one reason. A human can live with a system that fails a known 23% of inputs, because you can route those inputs elsewhere. Nobody can live with a system that handles the same input differently on Tuesday than on Monday, because there's nothing to route on.

Why one green run is a coin flip in disguise

Where does the flip come from? Temperature is the obvious suspect and the wrong one. Set it to zero and you still get variation from the provider side (batching, hardware, model updates behind a stable name), and more importantly the agent loop amplifies tiny differences. One slightly different tool call on step three means a different observation on step four, and by step eight you're on a different trajectory entirely.

The paper's contribution isn't the diagnosis, it's what they do with it. They built a Consistency Analyzer that looks at the five trajectories for a task and finds the specific step where they diverge. Then a Guideline Generator writes a short targeted instruction about that step ("when the search returns multiple contacts, filter on the email domain before picking one" is the kind of thing) and stores it as episodic memory. The next time the agent sees a similar task, that guideline gets injected.

Result on AppWorld: all-five pass rate up 16 points on the same tasks, and up 13 points on similar-but-unseen tasks. They don't claim to close the gap. They claim to narrow it, and the generalisation number is the one I find more convincing, since anyone can overfit guidelines to a fixed task list.

I'd push back on one thing. The paper treats "succeeds in all five runs" as the target. For a lot of business automation, five is too few. If the agent runs 400 times a day and a 2% flip rate means eight wrong vendor assignments, five-run consistency won't catch that. Pick your N from your volume, not from the paper.

What I do differently now

I've changed three things in how I test agents for clients, and none of them required the paper's framework. They just required admitting that a single run tells you almost nothing.

First, every eval case runs N times, and the report shows both numbers. Here's the shape of the harness, stripped down:

from collections import Counter

def run_case(agent, case, n=5):
    results = [agent.run(case.input) for _ in range(n)]
    passes = [case.check(r) for r in results]
    return {
        "case": case.id,
        "per_run": sum(passes) / n,
        "all_pass": all(passes),
        "paths": Counter(tuple(r.tool_calls) for r in results),
    }

def report(agent, cases, n=5):
    rows = [run_case(agent, c, n) for c in cases]
    per_run = sum(r["per_run"] for r in rows) / len(rows)
    consistent = sum(r["all_pass"] for r in rows) / len(rows)
    print(f"per-run pass: {per_run:.0%}  all-{n} pass: {consistent:.0%}")
    for r in rows:
        if 0 < r["per_run"] < 1:
            print(r["case"], "FLAKY", dict(r["paths"]))
Enter fullscreen mode Exit fullscreen mode

That paths counter is the part that earns its keep. A case that passes 3 of 5 with two different tool-call sequences tells you exactly where to look. On the invoice agent, the flaky cases all diverged at the same step: a vendor lookup that sometimes returned two matches, and the model picked differently each time. One line in the system prompt fixed it. I'd never have found it from a pass/fail table.

Second, I stopped reporting the per-run rate to clients as "accuracy". I report the all-N rate as the headline and the per-run rate as a footnote, because the all-N rate is the one that predicts support tickets. Clients don't love hearing 61% instead of 84%. They love it more than the screenshot of two vendors for one PDF.

Third, flaky cases block the release. A case that passes 4 of 5 used to be "basically fine". Now it's a bug with a known reproduction, and it gets the same treatment I described in when not to use AI automation: either the step gets deterministic (a rule, a lookup, a hard filter) or the whole path gets a human approval gate.

The memory idea, and where I'm not sold

The paper's actual mechanism, turning unstable steps into remembered guidelines, is worth trying and I've started to. My version is much dumber than theirs. When the harness flags a flaky case, I write the guideline by hand, drop it into a per-task notes file, and the agent's system prompt loads notes matching the task type. No analyser, no generator, one developer with a text editor.

That works for a handful of task types. It stops working around twenty, which is roughly where I'd want their automated version. Something I keep in mind from my post on agent memory: every guideline you inject is context the model has to weigh against everything else, and I've watched a pile of well-meaning rules make an agent worse at the cases that were never flaky. The paper doesn't report the cost side of that trade, or I missed it, and I'd like to see it.

The other caveat is the model. Their numbers are on GPT-4.1 with a ReAct loop. I haven't seen a published consistency gap for the current frontier models, and I'd guess it's smaller but not gone. I ran a quick five-repeat on my own invoice cases with Claude Sonnet 5 last night: per-run 91%, all-five 79%. Twelve points. Smaller gap, same shape, and that's one evening's data on a tiny set, so take it as an anecdote and not a finding.

If you want a broader lens on where agents fail, another paper from the same arXiv batch, AgentAudit, evaluates the whole trace (planning, tool selection, tool execution, memory) and attributes failures to a stage rather than a pass/fail. Different goal, same underlying point: the interesting failures are inside the trajectory, and a single end-to-end score hides them.

Run this on your own agent this week

You don't need the framework. You need one loop.

Take the ten test cases you already have for whatever agent you've shipped, run each one five times tonight, and count how many pass all five. Put both numbers side by side. If the gap is under 5 points, good, you've earned the word "stable" and I'm a little jealous. If it's 15 or more, find the cases that pass 3 of 5, dump the tool-call sequences from each run, and look for the step where they fork. In my experience it's the same step across most of the flaky cases, and it's usually a lookup that returns more than one thing.

Then decide, per flaky step, whether it becomes a rule or a human checkpoint. Neither option is glamorous. Both beat writing "stable" in a doc you'll have to explain later.

I build and test this kind of agent for clients as part of my freelance work, and the five-run harness above is now the first thing I set up on a new project, before the prompt, before the tools. It has changed what I'm willing to promise.


Originally published at abrarqasim.com. I write there about React, PHP, Rust, Go and the AI tooling around them.

Top comments (4)

Collapse
 
skillselion profile image
Skillselion •

Two ways to read the 77/53 gap, and only one survives the arithmetic.

Read one: every run is a coin flip at p=0.77, so five in a row is 0.77^5, which is 27.1%. The paper observed 53%, near double. That reading is dead.

Read two is the one the paper's own wording points at. It says the per-run pass rate averages 77%. An average over tasks, not a rate any single task has. Once the per-task rate varies, the all-five rate is the average of p^5, and that sits above (average p)^5 for the same reason an average of squares sits above the square of the average. No correlation between runs is needed to get 53%. Spread across the suite does it on its own.

So the gap is not evidence that runs are dependent. It is evidence that the suite is heterogeneous, which is the more useful finding, because heterogeneity has an address and "correlated somehow" does not.

How heterogeneous, without assuming a shape: 53% of tasks passed five for five, 47% broke the streak. Those 53% contribute at most 0.53 to the average, so the remaining 0.24 has to come from the other 47%, which puts their mean per-run rate at 0.51 or above. The group that broke the streak is genuinely mixed, not broken. That is the group worth spending runs on.

Your Sonnet 5 numbers do the same thing far more mildly. 0.91^5 is 62.4% against your observed 79%, a ratio of 1.27 where the paper's is 1.96. Less spread in your set, or an easier one. Worth saying, because it reads in the post as the same finding at smaller scale and it is a much weaker version of it.

Where I would push on the advice. A flaky case announces itself the moment two runs disagree, so five is not needed to find one. Run the suite at N=3, keep whatever comes back mixed, spend N=20 there. Be honest about the bill: on a suite split near 53/47 that is roughly double the total runs, not a reallocation. The reason to pay it is that runs four and five on a task that already passed three for three buy almost nothing.

One thing the post leaves out, and the abstract leaves out too: how many AppWorld tasks the 77/53 split covers. Without it the 24 point gap has no interval, and 53% starts getting quoted as a constant.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

I pulled the abstract to check the two numbers before touching them, and they are as you state: 77% per-run, 53% passing all five, ReAct on AppWorld with GPT-4.1, and +16 points on same-task evaluation for the framework.

Those two numbers do more work together than the article asks them to, so I spent a while on the arithmetic between them.

The 24-point gap is measured against the per-run average, which already contains the variation. The other baseline is the one where there is no case-level structure at all: if every case passed with independent probability 0.77 on every run, the fraction passing five of five would be 0.77⁵ = 27.1%. The measured 53% is 1.96× that. So the reliability figure most people would quote from this data is 27%, not 77%, and the interesting quantity is not the distance from either of those numbers but the fact that 53% sits between them at all.

Jensen gives that fact a direction. E[q⁵] ≥ (E[q])⁵, with equality only when the per-case pass probability q is constant across cases and runs are independent. 0.53 > 0.271 rules that combination out, so something is structured: either q varies by case (a mixture), or runs of the same case are dependent — a case that fails tends to fail again. Both are "consistency" problems, but they need different fixes: under the mixture reading, the actionable object is the stratum that never passes and no retry policy will touch it; under the dependence reading, the actionable object is whatever makes the same case take a different path, which is the thing the paper's Consistency Analyzer goes after. The published pair admits both, which is not a criticism of the paper — it is a reason to report one more count.

To make the mixture reading concrete, the two published moments pin down a spread. Fitting a Beta to E[q] = 0.77 and E[q⁵] = 0.53 gives a concentration of 1.35 and sd(q) ≈ 0.27; the minimum-support version — all cases either at q = x or at q = 0 — gives x ≈ 0.91 and about 15% of cases that never pass. Either way the per-case pass probability has to move by a quarter to a third of the whole scale to produce a 24-point gap. A 15% never-pass stratum is the part I would want to know about as a reader: it is invisible in a per-run average and unreachable by retries, and if it is real, the +16 points the framework reports on same-task evaluation is a question of which stratum shrinks.

The count that separates the readings, and that I think you already have, is the per-case pass count: how many tasks passed exactly 0, 1, 2, 3, 4, 5 of their five runs. Under independent runs at a single q = 0.77 that histogram is Binomial(5, 0.77), which is 0.1% at zero and 27% at five — it cannot be U-shaped. Over-dispersion of any kind — case heterogeneity or within-case dependence — fattens exactly those two tails, so the histogram constrains the size of the effect even where it cannot name the mechanism. It is one column next to the pass rate, and it turns "53%" into a shape you can act on: a thick zero tail means the portfolio contains tasks outside the agent's reach, a thick middle means the runs are noisy in a way that more attempts will smooth out.

Worth saying plainly because it is the practical upshot: a fixed-case retry is bounded above by exactly the all-k rate, so 53% at k = 5 is a ceiling, not a diagnostic — and if the pass counts are fat-tailed, the ceiling is not a property of your agent's average quality, it is a property of the mix of tasks you are averaging over.

Collapse
 
jo-do profile image
Jo Do •

This is the multiplication nobody does. A per-case pass rate raised to the power of your workflow length is the real reliability number, and it punishes exactly the demos that look best in a single run. The 'stable' handover doc is a familiar artifact - three green days measures weather, not climate. What changed my own testing: run the same case back-to-back until the variance shows, and report the streak survival rate instead of the pass rate. The 53% framing lands better with stakeholders than any reliability lecture.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.