DEV Community

RESK
RESK

Posted on

LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers

LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers

TL;DR — LLM evaluation only produces comparable numbers when every model faces the same prompts, the same fixed judge, and the same per-axis rubrics, with a public verbatim trail so anyone can re-read what a model actually answered. In the Team Recruitment (Oracle) benchmark, Nemotron 3 Ultra scored 90.87 and HY3 scored 83.1 — and the per-axis metrics show exactly why.

How the evaluation works

Most leaderboards hide the mechanism. This one does not. The Team Recruitment (Oracle) benchmark is a recruitment agent task: the model queries an oracle of résumés, then builds a team. It is scored on composition criteria with binary budget and seat gates. Here is the pipeline, step by step.

1. Same prompts. Every model receives the identical task, the identical oracle, and the identical constraints. No model gets a friendlier phrasing or a second attempt.

2. Same fixed judge. Scoring is not delegated to another chat model that might drift between runs. The judge is fixed, so the rubric is applied the same way to every submission.

3. Per-axis rubrics. The final score is not a vibe. It is a weighted composite of measurable axes: budget compliance, seat limit, skill match, role coverage, seniority balance, and whether the minimum seniority gate is met. Binary gates — budget and seat — can zero out an otherwise strong team.

4. Public verbatim trail. Every answer is stored verbatim. Anyone can re-read what the model actually produced, not just the number it received. That is what makes the comparison auditable rather than trust-based.

The real numbers

Model Provider Score Budget Seat limit Skill match Role coverage Seniority balance Min seniority Samples Tokens Avg latency (ms)
Nemotron 3 Ultra (free) opencode-zen 90.87 0.984 1 0.869 1 0.44 1 11 154054 174951.45
HY3 (free) opencode-zen 83.1 0.996 0.909 0.732 0.864 0.53 0.955 11 123775 94013.55

For context, we also submitted our own model to this benchmark: GLM 5.2 scored 78.0 overall.

Why the ranking looks this way

The winner does not win on every axis. HY3 actually beats Nemotron 3 Ultra on budget (0.996 vs 0.984) and on seniority balance (0.53 vs 0.44). So why is Nemotron ahead by 7.77 points?

Look at the axes that carry the most weight in a recruitment task. Role coverage is 1 for Nemotron and 0.864 for HY3. That means Nemotron filled every required role; HY3 left part of the team uncovered. Skill match is 0.869 versus 0.732 — a 13.7-point gap. In practice, Nemotron's picks actually matched the résumés to the roles, while HY3's team was more loosely assembled.

Then there is the seat limit: 1 versus 0.909. Nemotron respected the hard constraint perfectly. HY3 violated it on roughly one in eleven samples. Because seat limit is a binary gate, that single failure is expensive — it drags the composite down far more than a small budget difference ever could.

Minimum seniority tells the same story: 1 versus 0.955. Nemotron always met the floor. HY3 missed it in a small fraction of runs, and again the gate amplifies the penalty.

HY3's advantages are real but smaller in effect. A 0.012 budget edge and a 0.09 seniority-balance edge cannot compensate for losing role coverage, skill match, and two binary gates. That is the mechanism working as intended: the rubric rewards constraint compliance first, composition quality second.

Latency and tokens are reported for transparency, not folded into the score. Nemotron used 154054 tokens at 174951.45 ms average latency; HY3 used 123775 tokens at 94013.55 ms. HY3 was faster and cheaper in tokens, but the benchmark measures team quality, not speed.

What this means for practitioners choosing a model

If your use case has hard constraints — budget caps, seat limits, minimum seniority — do not pick a model on overall score alone. Pick on the axes that map to your constraints. A model that is 0.9 on a binary gate will fail you in production roughly one run in ten. A model that is 1.0 will not.

Also read the sample count. Both models here were evaluated on 11 samples. That is enough to see a pattern, not enough to certify a guarantee. Treat the score as a signal, not a contract.

How to read a leaderboard

  1. Check whether the prompts and judge are fixed. If they are not, the numbers are not comparable.
  2. Find the binary gates. A high composite score can hide a gate failure.
  3. Compare per-axis, not just overall. The axis that matters to you may not be the axis that decided the ranking.
  4. Look for the verbatim trail. If you cannot re-read the actual answers, you are trusting a summary.
  5. Note the sample count. Small samples show direction, not certainty.

Honest limitations

This benchmark measures one task: building a team from an oracle of résumés. It does not measure coding, reasoning, or long-context recall. The scores are provider-specific — both models here ran on opencode-zen — and a different provider could change latency and cost. The seniority-balance axis is a soft criterion, so its weight is a design choice, not a universal truth. And our own submission, GLM 5.2 at 78.0, is included for transparency, not as a neutral reference point.

Conclusion

LLM evaluation is only as good as its mechanism. Same prompts, same fixed judge, per-axis rubrics, and a public verbatim trail are what turn a pile of model outputs into numbers you can actually compare. Nemotron 3 Ultra leads this benchmark at 90.87 because it wins the axes that carry the most weight — role coverage, skill match, and the binary gates. HY3 at 83.1 is close, and faster, but loses on the constraints that matter most.

Explore the full leaderboard and the verbatim trail at lforla.org.

Top comments (0)