DEV Community

RESK
RESK

Posted on

How LLM Evaluation Actually Works: Inside a Benchmark That Produces Comparable Numbers

TL;DR

LLM evaluation only produces comparable numbers when every model faces the same prompts, the same fixed judge, per-axis rubrics, and a public verbatim trail. This article walks through how the FreeCAD Fix benchmark on LFORLA does exactly that, and what its leaderboard scores actually mean in practice.

How the evaluation works

The FreeCAD Fix benchmark diagnoses and repairs a broken parametric FreeCAD script. It is scored by a deterministic geometry oracle against real freecadcmd measurements. There is no LLM judge in the loop. That single design choice removes a whole class of evaluation noise.

Here is the mechanism, step by step.

1. Same prompts for every model. Each model receives the identical broken script and the identical repair instructions. No model gets a hint, a retry, or a warmer context. This is the baseline requirement for comparability: if the inputs differ, the outputs cannot be ranked.

2. Same fixed judge. In this benchmark the judge is not a language model. It is a deterministic geometry oracle. The repaired script is executed with freecadcmd, and the resulting geometry is measured against the expected geometry. The oracle returns a number, not an opinion. That number is reproducible: run it twice, get the same result.

3. Per-axis rubrics. The benchmark spans CAD, engineering, code, and reasoning. Each axis has its own rubric so a model that writes clean code but misreads the geometry is scored differently from one that understands the geometry but produces broken syntax. Per-axis scoring prevents a single strong dimension from masking a weak one.

4. Public verbatim trail. Every answer is stored verbatim. Anyone can re-read what a model actually answered, not just its final score. This is the audit layer: if a score looks surprising, you can open the transcript and see why.

The real numbers

Model Score Provider
Mimo v2.6 Flash (free, via opencode) 100 opencode
Qwen 3.8 27B (free, via OpenRouter) 99.33 opencode
Longcat 2.5 Preview (free, via opencode) 90 opencode
Nemotron 3 Ultra 550B (free, via OpenRouter) 83.33 opencode
Nemotron 3 Ultra (free, via opencode) 66.67 deepseek
Space Bunny (free, via opencode) 60.33 opencode
Ling 3.1 Flash (free, via opencode) 58.33 opencode
Laguna S 2.1 (free, via OpenRouter) 56.67 opencode

We also submitted our own model, GLM 5.2, which scored 78.0 overall.

Why the ranking looks this way

A score of 100 means the repaired script passed every geometry check. A score in the 50s or 60s means the model produced a script that ran but failed a meaningful share of the geometric assertions. The gap between 99.33 and 100 is small in absolute terms but real in a deterministic oracle: one assertion failed.

The spread from 56.67 to 100 is the interesting part. It tells you that repairing a parametric CAD script is not a task where every capable model lands in the same place. Some models understand the geometry but write fragile code. Others write clean code but misread the constraint. The per-axis rubrics expose that difference; a single blended score would hide it.

Notice also that the same model family appears twice with different scores: Nemotron 3 Ultra 550B at 83.33 and Nemotron 3 Ultra at 66.67. Different providers, different serving stacks, different results. In LLM evaluation, the provider is part of the measurement.

What this means for practitioners choosing a model

Do not pick a model from a single headline number. Pick it from the axis that matches your workload. If you are repairing CAD scripts, the FreeCAD Fix leaderboard is directly relevant. If you are doing something else, the same mechanism applies: find a benchmark whose judge matches your success criterion.

A deterministic oracle is a strong signal because it cannot be flattered. A model that talks convincingly about geometry but produces a broken script will score low here. That is the point.

How to read a leaderboard

  1. Check the judge. If the judge is an LLM, ask how it was fixed and whether its prompts are public.
  2. Check the prompts. Identical inputs are the precondition for comparable scores.
  3. Check the axes. A single blended score hides which capability actually failed.
  4. Check the trail. If you cannot re-read the verbatim answer, you cannot audit the score.
  5. Check the provider. The same model can score differently on different serving stacks.

Honest limitations

This benchmark measures one task: repairing a broken parametric FreeCAD script. A high score here does not mean a model is good at general reasoning, writing, or any other domain. The leaderboard is a snapshot, not a permanent ranking. Free-tier models via OpenRouter and opencode can change behavior without notice. And our own submission, GLM 5.2 at 78.0, is a single data point, not a claim of superiority.

Conclusion

LLM evaluation is not a mystery. It is same prompts, same fixed judge, per-axis rubrics, and a public verbatim trail. The FreeCAD Fix leaderboard on LFORLA applies that mechanism to a concrete engineering task and publishes the results. If you want to see the full leaderboard and read the verbatim answers, start at https://lforla.org.

Top comments (0)