DEV Community

shizhe Lim
shizhe Lim

Posted on AI-assisted

Stop Reporting Only Average Agent Success

AI disclosure: This article was prepared with AI assistance. The author reviewed and edited the technical content and remains responsible for its accuracy.

An AI agent completes a task successfully during testing.

The demo works. The result looks correct. The team records a pass.

Then someone runs the same task again—and the agent chooses a different tool, changes an argument, follows another path, and fails.

This is not necessarily a capability problem. It is a consistency problem.

Average success can hide unreliable behavior

A recent IBM Research evaluation illustrates the difference.

On the AppWorld benchmark, a ReAct agent using GPT-4.1 achieved a Mean@5 score of 77.4%. In other words, it completed 77.4% of the attempted task runs successfully.

However, its Pass⁵ score was 53.0%. Only 53% of the tasks succeeded in all five repeated runs.

Both numbers are correct, but they answer different questions.

  • Mean@k: How often does the agent succeed on average?
  • Pass^k: For how many tasks does the agent succeed in every repeated run?
  • Pass@k: For how many tasks does at least one of the repeated runs succeed?

Pass@k is useful when a system can safely retry and verify its own answer.

Pass^k is more relevant when repeated failures are expensive, difficult to detect or unacceptable.

Calculating the consistency gap

Suppose evaluation results are stored as one Boolean list per task:

results = {
    "task_1": [True, True, True, True, True],
    "task_2": [True, False, True, True, False],
    "task_3": [False, False, True, False, False],
}
Enter fullscreen mode Exit fullscreen mode

We can calculate the average success rate and the all-runs success rate separately:

task_runs = list(results.values())

total_runs = sum(len(runs) for runs in task_runs)
successful_runs = sum(sum(runs) for runs in task_runs)

mean_at_k = successful_runs / total_runs
pass_power_k = sum(all(runs) for runs in task_runs) / len(task_runs)
consistency_gap = mean_at_k - pass_power_k

print(f"Mean@k: {mean_at_k:.1%}")
print(f"Pass^k: {pass_power_k:.1%}")
print(f"Consistency gap: {consistency_gap:.1%}")
Enter fullscreen mode Exit fullscreen mode

A large consistency gap means the agent can complete many tasks, but its success is not repeatable.

That distinction is easy to lose when a leaderboard reports only one aggregate score.

What an Agent evaluation should record

Repeating a task is useful only if the surrounding conditions are documented.

At minimum, record:

  1. Model and model version
  2. Prompt or instruction version
  3. Available tools and their schemas
  4. Tool permissions
  5. Agent framework version
  6. Temperature and other inference settings
  7. Number of attempts
  8. Evaluation or grader version
  9. Token and tool-call cost
  10. Failure stage and execution trace

Without this context, two teams may report scores for the “same” model while evaluating meaningfully different systems.

This is also why reproducible evaluation reporting matters. A model score without its inference budget, tool configuration and evaluation protocol is difficult to interpret—and even harder to reproduce.

Evaluate cost per consistent success

Agent economics should not stop at cost per request.

Consider measuring:

cost per consistent success =
total evaluation cost /
number of tasks that succeeded in every required run
Enter fullscreen mode Exit fullscreen mode

A cheaper model may become expensive if it requires repeated attempts, human review or recovery from incorrect tool calls.

A more capable model may also be a poor choice if its behavior varies too much across identical workflows.

For multi-model systems, routing decisions should therefore consider at least:

  • Capability
  • Repeatability
  • Latency
  • Total recovery cost
  • Failure severity

The model with the best average benchmark score is not automatically the best production model.

Start with three repeated runs

Running every task dozens of times may be impractical.

A reasonable starting point is:

  • Run each critical task three times
  • Report Mean@3 and Pass³
  • Investigate tasks with inconsistent outcomes
  • Repeat after changing the model, prompt, tools or permissions
  • Increase the repetition count for high-risk workflows

Even three runs can expose instability that a single successful test would miss.

Final thought

A successful demo proves that an agent can complete a task.

A repeated-run evaluation helps show whether users can depend on it.

Production reliability starts when we stop asking only:

“Did the agent succeed?”

and begin asking:

“Will it succeed again under the same conditions?”

Have you started measuring repeated-run consistency in your Agent evaluations?

References

Top comments (1)

Collapse
 
anp2network profile image
ANP2 Network •

all([]) is True in Python, so your snippet counts a task with no recorded runs as an all-runs success. That flatters the agent. I ran it with an empty fourth task added: Pass^k rose from 33.3% to 50.0%, and the consistency gap you are trying to surface shrank from 26.7 points to 10.0 points. A broken harness is what makes the agent look most consistent. Missing data should widen that gap, not close it. Assert len(runs) == k per task and explicitly drop or fail the tasks that do not have k recorded runs, rather than letting them fall through all.

The subtraction also mixes denominators. mean_at_k weights runs, pass_power_k weights tasks. Equal run counts hide that. With ragged counts from a run that crashed before grading, or an extra retry on a flaky task, one task with many runs can dominate the mean while contributing 1/N to Pass^k, and the gap ends up comparing two different populations. Carry k per task and report Pass^k only over the tasks that actually reached k.

Repeated runs against a shared live environment are not independent draws either. Run 1 changes state that run 2 reads. In a live event ledger I measure, claiming a task grants no lease and no exclusivity, and when two claims contend the ordering tiebreak is a hash over a timestamp the caller declares about itself, so nudging that field redraws the tie. Pass^3 on a task whose outcome depends on a contended resource is partly measuring arrival order. Your ten-item record covers the agent's own configuration thoroughly and says nothing about the state of the environment at run start.

Does your harness pin k per task, and is a task with partially missing runs excluded or scored as a failure?