DEV Community

Cover image for Your Benchmark Should Make Your Engineers Uncomfortable
Robert Imbeault
Robert Imbeault

Posted on

Your Benchmark Should Make Your Engineers Uncomfortable

If a benchmark only gives marketing something nice to post, you are probably missing the useful part.

The AI industry has a benchmark problem, not because we have too many benchmarks, but because we have become extremely good at turning them into trophies.

Every week, another model reaches the top of another leaderboard. Another release is state of the art. Another chart appears explaining why one system is definitely better than all the other systems, provided you read the footnotes in exactly the right order.

It is easy to become cynical about the whole thing. I am not. I think benchmarks are one of the most useful engineering tools we have. I just think we have started confusing the measurement with the mission.

The Score Is Not the Product

At Backboard, we benchmark a lot.

But it's not because we believe benchmarks are perfect, we know they are not. We benchmark because an imperfect, transparent measurement is usually more useful than everyone standing around saying, “It feels faster to me.”

Software teams are remarkably good at convincing themselves that something has improved because they spent three weeks improving it.
Benchmarks are less sentimental.

They tell you whether the change actually helped.
Sometimes an optimization does exactly what you expected.
Sometimes the number does not move.
Sometimes the very clever thing you spent a week building turns out to have made the system worse.

That last one is always a wonderful and useful team-building exercise.

The point of the benchmark is not the score itself.
The point is the feedback loop.

A Good Benchmark Should Cause Problems

The most useful benchmarks expose things you would rather not see.

  • Where does the system fail?
  • Which tasks are unreliable?
  • Did your optimization introduce a regression somewhere else?
  • Does the system still perform well when the easy cases are removed?

If your benchmark never produces uncomfortable conversations inside the engineering team, I would start wondering whether it is measuring anything useful.

A benchmark should challenge your engineers before it impresses your marketing team.

That is the test I keep coming back to.

A good evaluation tells the engineering team where to look next.
If it also gives marketing a nice graph afterward, great.

Everyone gets a graph, but the graph should be the output of the engineering process, not the reason the process exists.

When the Measurement Becomes the Mission

This is where benchmarks can go wrong.

Once the score itself becomes the objective, teams naturally start optimizing for the test rather than the capability the test was supposed to represent.

You can tune specifically for benchmark tasks.
You can cherry-pick configurations.
You can publish the strongest run and quietly forget the weaker ones existed.
You can accidentally or deliberately allow evaluation data to influence training.

All of those things can improve a leaderboard position.

None of them necessarily improve the product.

That distinction matters.

There is a big difference between building for a benchmark and benchmarking what you built.

Building for a benchmark starts with the test.

The question becomes:
How do we make this number bigger?

Benchmarking what you built starts with the actual problem.
You build something that you believe solves it well, then use an independent evaluation to find out whether you are fooling yourself.

Those two approaches can produce similar-looking numbers.

They represent very different engineering cultures.

Build First. Benchmark Second.

The customer does not care that your agent achieved an excellent score on a task they will never ask it to perform. They care whether it works on their problem.

That sounds obvious, but leaderboard incentives have a funny way of making obvious things less obvious.

The order matters.
Build around real use cases.
Measure the system.
Find the failures.
Change the system.
Measure it again.

That is where the benchmark becomes valuable. It creates a feedback loop that is much harder to fake than internal enthusiasm. It also gives teams a common language.

Instead of saying:
“I think this version is better.”

You can say:
“This improved these capabilities, regressed on these ones, and did nothing here.”

That is a considerably more useful engineering conversation.

Transparency Is Part of the Result

There is another part of benchmarking that I think gets less attention than the final score.

Can anyone inspect how you got there?

Whenever possible, we publish our methodology at Backboard.
We share configurations and results. We open-source evaluation artifacts where we can. We want people to be able to reproduce what we did and tell us when they think we got something wrong.

That last part is important.

If somebody finds a mistake in your benchmark methodology, that is not evidence that transparency failed. It is evidence that transparency worked. Science progresses because results can be challenged. Engineering improves because assumptions can be tested.

A screenshot of a leaderboard tells me a number. Logs, configuration, methodology, and reproducible results tell me whether I should believe it.

I find the second category considerably more interesting.

Benchmarks Still Cannot Tell You Everything

Even a very good benchmark measures a narrow slice of a system.
It does not tell you whether customers trust the product. It may not tell you whether the system is pleasant to use. It cannot capture every strange production workflow someone will eventually discover at 4:47 p.m. on a Friday. And it certainly does not replace watching what happens when real people use the thing.

Public evaluations should sit beside production testing and customer feedback.

Not above them.

A benchmark is one source of evidence. It should not become the entire definition of whether the product works.

Yes, Benchmarks Are Also Marketing

There is another benefit to strong benchmark performance.

People notice.

A good result gets shared. Engineers investigate it. Customers find you. Investors send messages. People who had never heard of the product suddenly want to know what you built.

That is useful. I am not going to pretend otherwise.

There is nothing wrong with marketing a good engineering result.
The important part is the sequence.

You should not benchmark because you need something impressive to announce.

You should benchmark because you are trying to understand and improve the system.

Then, if the result is impressive, by all means tell people about it.

We certainly do.

Just do not confuse the announcement with the reason the work mattered.

The Leaderboard Is a Byproduct

The healthiest benchmarking loop I know is fairly simple.

Build something useful.
Measure it honestly.
Find where it breaks.
Improve it.
Let other people inspect what you did.
Then repeat.

If that eventually puts you near the top of a leaderboard, excellent, but the leaderboard is evidence that something interesting may be happening. It is not the thing you are building.

We want Backboard products to perform extremely well on public evaluations. Of course we do.

But if we ever reach the point where the benchmark stops teaching our engineers something and starts existing primarily to give us a number to post, we will have missed the point.

Because benchmarks do not build great products.

Engineers do.


This is a conversational remix of something I originally wrote for the Backboard blog. If you want the less-chatty deep dive, you can read the original here.

Top comments (0)