Every few days someone tells me which coding agent is "obviously the best". They always have a benchmark to back it up. And the benchmark is always run on code that has nothing to do with mine.
I got tired of nodding along, so I built a small thing to settle it the dumb way: give them all the same task at the same time and watch.
It's called Agent Derby, and this is what a race looks like.
The first race
I typed "build a playable snake game in the browser", picked three Claude models, and hit start.
Sonnet 5.5 was finished in 17 seconds. I hadn't even finished reading what the other two were doing.
Haiku 4.5 came in at 31 seconds.
Opus 5.5 took two and a half minutes and cost about ten times as much as either of them.
So Opus lost, right? Well. Then I played the games.
Sonnet's game works, and that is all it does. A square, a dot, a score.
Opus built a start screen, sound effects, a pause button and a mute key. It wrote about six times as much code. Nobody asked for any of that, and it's also clearly the one you'd want to ship.
That's the thing a score table can't tell you. "Fastest" and "best" were two different models, and I only found out because I could actually play all three.
It's one task. It proves nothing in general. But it changed how I think about picking a model, which is more than any leaderboard has done for me.
What it actually does
You type a task and choose who races. Everyone gets the exact same prompt at the exact same moment.
Each agent works in its own sandbox, a private copy of the project. They can't see each other's work and they can't write to your real code. I was nervous about letting three agents loose with every permission switched on, so on macOS and Linux the operating system itself blocks them from writing anywhere else.
While they run, you see every lane at once: what each one is reading, editing or running, with the clock, tokens and cost ticking.
When one finishes, whatever it built just starts up in its lane. Web apps open in a frame. Terminal programs get a real terminal. No copying commands, no "now run npm install".
Then you get the numbers lined up next to each other.
If you like one result, you can keep it as a new branch or a folder. It never merges anything for you.
One rule I was stubborn about
If an agent's CLI doesn't report a number, the app says "not reported".
Not zero. Not a guess. Not a dash that looks like zero.
It sounds like a small thing, but comparison tools get untrustworthy fast when they quietly fill in blanks. If the cost is an estimate, it's labelled as an estimate.
Things that surprised me while building it
Some models take a long time to say anything. One model sat silent for minutes between steps and timed out twice. I assumed my sandbox had broken it. It hadn't. It was just slow that day. The other lanes carried on without noticing, which is exactly what I wanted to be true, but it was a nervous ten minutes.
Agents do unexpected things when left alone. One of them, asked to build a game, decided to start a web server and go looking for a browser to test it in. Fair enough, honestly.
"Is it playable?" is harder to check than it sounds. My first automated check pressed arrow keys and declared one game broken. It wasn't broken. It had a Start button.
Try it
You need Node and git.
npx https://github.com/osmanahmadxai/agent-derby/releases/latest/download/agent-derby.tgz
It uses the agent CLIs you're already logged in to, so there are no API keys involved. If you don't have any agent installed, there are built-in demo agents that run a scripted race for free, so you can see how it works first.
There's also a terminal version that splits your screen into one pane per agent, and a desktop app.
What's not finished
I'd rather tell you than have you find out:
- I've only properly tested Claude Code. Codex CLI and Gemini CLI are wired up, but I haven't had a real logged-in run with either yet. They may break.
- Windows has no sandbox yet and is basically untested.
- The desktop builds aren't code-signed, so your OS will complain the first time.
And for full disclosure: I built this with Claude Code. Which means an AI coding agent helped write the tool that judges AI coding agents. Make of that what you will.
The repo
It's MIT licensed and all of it is here: github.com/osmanahmadxai/agent-derby
If you race something, tell me what happened. I'm especially curious whether the ranking flips on a harder task than snake. My guess is yes.



Top comments (0)