DEV Community

Cover image for Stop Benchmarking Agents Against Each Other. Audit Them Against a Tuesday.
Jiahui Miao
Jiahui Miao

Posted on

Stop Benchmarking Agents Against Each Other. Audit Them Against a Tuesday.

Stop Benchmarking Agents Against Each Other. Audit Them Against a Tuesday.

By Jiahui Miao

Last Tuesday, my AI agent handled eleven real things for me. It got ten right. One it got wrong — and the way it got it wrong taught me more than any leaderboard ever has.

Here's the part the industry doesn't want to hear: no benchmark on earth can tell me what I learned that Tuesday. And no benchmark on earth matters more.

The orthodoxy

The agent industry has decided that progress looks like a scoreboard. GAIA, SWE-bench, WebArena — the names change, the ritual doesn't: a new model drops, someone posts a number, and the timeline declares that agents have gotten "better."

This is benchmark maximalism, and it is the wrong unit of measurement for the thing that actually matters.

A leaderboard measures agents against each other, on tasks chosen by researchers, in environments cleaned of the one ingredient that defines real work: consequences. Nobody loses money when an agent fails WebArena. Nobody misses a flight. Nobody has to explain to a landlord why the rent email went to the wrong inbox.

What I do instead

I'm building a personal AI for exactly one person: me. Not a demo. Not a product for "users." One human, with one calendar, one inbox, one set of habits that no benchmark designer will ever encode.

So I stopped asking "how good is my agent?" and started asking a different question: how was Tuesday?

Every week I audit one real day. Not a curated task list — the actual day, in all its messiness. Each thing the agent touched gets a binary verdict: done right, or not. No partial credit. No "well, it understood the intent." And for every failure, one sentence about why — wrong assumption, missing context, or the most common one: it was confident about something it should have asked about.

Ten out of eleven sounds like 91%. It's not. It's a list of eleven specific things, and the one that failed is the only line that matters, because that's the only line that changes what I build next.

Why the leaderboard can't do this

Three reasons, all structural:

First, benchmarks have no memory of you. My agent knows that when I say "the usual," I mean something specific — a specific coffee order, a specific route, a specific way I want bad news delivered before good news. No shared test set can encode that. The entire point of a personal AI is that the test set is your life, and your life is not downloadable.

Second, benchmarks reward the average. A leaderboard asks: how does this agent do across a thousand generic tasks? But I don't live a thousand generic tasks. I live the same few dozen recurring ones, over and over, where the difference between good and great is whether the agent learned that I always double-book Thursdays and it should check. Averaging destroys exactly the signal that matters.

Third, benchmarks can't be wrong in the way that counts. When my agent misfired last Tuesday, I felt it — a real consequence, a real correction, a real change in how I instruct it the next morning. That feedback loop, tight and personal and slightly annoyed, is worth more than a thousand automated evals. Pain is a feature of evaluation. Leaderboards have never felt pain.

The uncomfortable implication

If personal AI is where agents are actually going — and I believe it is, because a tool that serves everyone equally serves no one deeply — then the industry's evaluation infrastructure is pointed at the wrong target.

We're optimizing agents to beat each other at shared games while the real contest is private: can your agent survive your Tuesday? The company that wins personal AI won't be the one with the highest GAIA score. It'll be the one whose agent has the longest memory of a single human's actual life — and the shortest loop between getting something wrong and never getting it wrong again.

That's what I'm building at vertciti: an AI for one person, evaluated by one person's days. Not a benchmark. A life.

So here's my proposal, and it's deliberately unfalsifiable by any lab: stop benchmarking agents against each other. Audit them against a Tuesday. Pick a real day. Count what got done right. Write down the one thing that didn't. Do it again next week.

The leaderboard will tell you your agent is in the 99th percentile. Tuesday will tell you the truth.


Jiahui Miao is building vertciti — a personal AI for exactly one person. This is the fourth piece in an ongoing series on building AI for N=1.

Top comments (0)