The public leaderboard says Model X is #1. Your production traffic disagrees. Here’s how to build the benchmark that actually predicts which model works for you.
Here’s a pattern we see constantly. A team needs to pick an LLM for a real product. They open a public leaderboard, see a model sitting proudly at the top, and reason, quite sensibly, “well, that’s the best one.” They wire it in. And then in production it’s mediocre at the one thing they actually needed it to do, while a model ranked eighth quietly would have crushed it.
The leaderboard wasn’t lying. It just wasn’t measuring your problem.
This isn’t a knock on public benchmarks. They’re genuinely useful for what they are: a broad, standardized signal about general capability. But “generally capable” and “good at extracting line items from your suppliers’ weirdly-formatted invoices” are different questions, and only one of them pays your bills. If you’re selecting a model for an enterprise workload, the benchmark that matters is the one you build yourself.
Let me walk through how to actually do that, what to measure, how to build the harness, and the traps that quietly ruin the result.
Why the Public Leaderboard Can’t Answer Your Question
Start with the uncomfortable truth about why a general LLM benchmark leaderboard can’t make your decision for you.
It measures the average task, not your task. A leaderboard score is an aggregate across a broad set of general challenges, reasoning, coding, trivia, writing. Your actual workload is narrow and specific: classify support tickets, draft compliance summaries, pull structured data from PDFs. A model can be excellent on average and weak on your narrow slice, and the aggregate score will never tell you which.
Your data isn’t in the test set, and might be the opposite of it. Public benchmarks run on public data, clean, well-formed, English-heavy, general-domain. If your inputs are messy internal documents, domain jargon, mixed languages, or formats no benchmark ever imagined, the leaderboard is testing a world that doesn’t resemble yours.
Contamination is real. Popular benchmark questions leak into training data. A model can score well partly because it has effectively seen the answers. Your internal tasks have no such leakage, which is exactly why they’re a truer test.
It ignores everything that isn’t accuracy. Latency. Cost per thousand calls at your volume. Whether it fails safely or confidently hallucinates. Whether it can be self-hosted for your compliance requirements. The leaderboard crowns a winner on quality alone; you’re optimizing a system with four or five constraints at once.
Preference ≠ fitness. Arena-style leaderboards rank models on which output humans prefer in blind comparison. Useful signal, but “which answer sounds nicer” is not “which model correctly applies your refund policy.” Likeability and correctness diverge more often than people expect.
None of this makes public leaderboards useless. It makes them a starting shortlist, not a decision. You use them to narrow fifty models to five. Then you build your own benchmark to pick the one.
What “Build Your Own Benchmark” Actually Means
It sounds heavier than it is. People hear “benchmark” and picture a research project. In practice, an internal LLM benchmark is four things:
- A representative dataset of your real tasks, with known-good answers.
- A scoring method that reflects what “correct” means for your use case.
- A harness that runs every candidate model against that dataset, identically.
- A scorecard that weighs quality against cost, latency, and your other constraints.
That’s it. You don’t need a PhD or a GPU cluster. You need a few hundred real examples, a clear definition of success, and a weekend of engineering. Let’s take each part.
Step 1: Build the Dataset From Your Actual Traffic
This is the part that determines whether the whole exercise is worth anything, so don’t shortcut it.
Pull real examples from your actual workload, support tickets you’ve already handled, invoices you’ve already processed, queries your users have actually typed. Not synthetic examples you invented, and not the clean happy-path cases. You want the real distribution, including the messy 20% that breaks things, because the messy 20% is where models actually differentiate.
How many? Fewer than you’d think. 100 to 300 well-chosen examples usually gives you a clear signal. The goal isn’t statistical perfection, it’s enough coverage of your real cases that the ranking is stable and the failures are visible.
The hard part is the known-good answers, the labels. For each example, you need the correct output a human expert would produce. This is tedious and it’s non-negotiable: a benchmark with no ground truth can only measure “which output looks plausible,” which is the exact trap we’re trying to escape. Have your domain experts label the set once. That labeled set becomes a durable asset, you’ll reuse it every time a new model drops.
Deliberately over-weight the edge cases. The ambiguous ticket, the invoice with the missing PO, the query in two languages. On the happy path most modern models are fine and indistinguishable. The edges are where the real ranking lives.
Step 2: Define What “Correct” Means For You
“Accuracy” is not one thing, and picking the wrong definition quietly wrecks the benchmark.
For a classification task (routing a ticket to the right queue), it’s clean: exact match against the label, measured as accuracy or F1. Easy to score, easy to automate.
For structured extraction (pulling fields from a document), score each field. Did it get the amount right? The date? The vendor ID? A model that nails 6 of 7 fields is very different from one that nails 3, and a single pass/fail hides that.
For open-ended generation (a drafted summary, a support reply), there’s no single right string, so you need one of:
Rubric scoring, a human rates each output against defined criteria (factually correct? complete? on-policy? right tone?).
LLM-as-judge, a strong model scores outputs against your rubric. Fast and cheap and genuinely useful, but validate it against human scores on a sample first, because judge models have biases (they tend to like longer answers, and they tend to favor outputs that sound like themselves). Trust it only after you’ve checked it agrees with your humans.
Whatever you choose, write the definition down before you run anything. Deciding what “good” means after you’ve seen the scores is how motivated reasoning sneaks in.
Step 3: Build the Harness
The engineering here is refreshingly boring, which is the point.
A simple loop: for each model in your shortlist, for each example in your dataset, send the same prompt, capture the output, the latency, and the token counts. Store everything. Score it. Aggregate.
A few things that separate a harness you can trust from one you can’t:
Hold the prompt constant across models, same instructions, same format, same few-shot examples. If you hand-tune the prompt for your favorite and use a lazy one for the rest, you’re benchmarking your prompt-writing, not the models.
Run each example more than once. LLMs are non-deterministic even at low temperature. Three runs per example and an average keeps a lucky or unlucky single roll from deciding your architecture.
Capture cost and latency as first-class metrics, not afterthoughts. Log tokens in and out per call; you’ll convert that to real money at your projected volume later.
Log every raw output. When a model scores surprisingly high or low, you’ll want to read what it actually produced. Half your real insight comes from reading the failures, not the scores.
Keep it in version control. When the next model launches, you re-run the same harness against the same labeled set and get an apples-to-apples answer in an hour, instead of re-litigating the whole decision from scratch.
Step 4: Score the Whole System, Not Just Quality
Here’s where most internal benchmarks still go wrong, they rank on quality and stop, reinventing the exact single-axis mistake the public leaderboard made.
A real enterprise decision is multi-constraint. Build a scorecard that holds them together:
- Quality , your task-specific score from Step 2.
- Cost , projected monthly spend at your real volume. A 2-point quality edge rarely justifies a 5x bill.
- Latency , p50 and p95. The tail matters: a model that’s fast on average but occasionally stalls for ten seconds can be unusable in a live chat.
- Failure behavior , when it’s wrong, is it wrong safely? A model that escalates or says “I’m not sure” beats one that confidently fabricates, even at equal accuracy.
- Deployability , can it meet your data-residency and compliance constraints? Sometimes “must be self-hostable” silently eliminates the top three on quality, and that’s fine, it’s a real constraint.
Weight these for your situation , a high-volume support bot weights cost and latency heavily; a low-volume compliance tool weights quality and failure behavior above almost everything, and let the weighted scorecard pick the winner. The output isn’t “Model X is best.” It’s “Model X is best for us, at our volume, under our constraints,” which is the only ranking that was ever going to help you.
The Payoff, and Where It Leads
Do this once and you own something genuinely valuable: a labeled dataset and a repeatable harness that turn every future model release from an anxiety-inducing “should we switch?” into a one-hour measured answer. When the next frontier model drops with breathless leaderboard claims, you run it through your harness and know, for your workload, whether it’s worth the migration.
That capability compounds. It’s the same muscle you need for any serious custom LLM implementation, because selecting the model is step one of a longer job: the evals you build here become the regression tests that keep the system honest as you fine-tune, swap models, and scale. And once the model is chosen, the same measure-everything discipline is what separates an agent that demos well from one that survives production, which is the entire difficulty of real custom AI agent development.
The One-Paragraph Version
Public leaderboards answer “which model is generally good.” That is not your question. Your question is “which model is good at my task, at my volume, under my constraints, on my messy data”, and the only honest way to answer it is to build a small internal benchmark: a few hundred real labeled examples, a task-specific scoring method, a fair harness that holds the prompt constant and runs each case a few times, and a scorecard that weighs quality against cost, latency, failure behavior, and deployability. It’s a weekend of work that saves you from a six-figure mistake and keeps paying off every time a new model ships.
The leaderboard is where you start. The benchmark you build is how you actually decide.
Have you run your own internal model bake-off? I’m curious what surprised you most , for us it’s almost always that the “best” model on paper loses on the messy edge cases, and a cheaper mid-ranked model wins the real workload.
Have you run your own internal model bake-off? I’m curious what surprised you most , for us it’s almost always that the “best” model on paper loses on the messy edge cases, and a cheaper mid-ranked model wins the real workload.
Top comments (0)