DEV Community

RESK
RESK

Posted on

How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

TL;DR

LLM evaluation only becomes useful when every model faces the same prompts, the same fixed judge, per-axis rubrics, and a public verbatim trail. The LFORLA Reverse Engineering benchmark does exactly that: it restores C source from stripped binaries, scores with deterministic token similarity against server-only references, and publishes what each model actually answered. Here is how the numbers are made, what the current leaderboard shows, and how to read it without fooling yourself.

How the evaluation works

The benchmark is called Reverse Engineering (Binary to Source). The task is blunt: given a stripped binary, restore the C source of a function. Models can run in tools or no-tools mode. The scoring is deterministic token similarity against server-only references, so the same answer always gets the same score.

Here is the mechanism, step by step.

  1. Same prompts for everyone. Every model receives the identical binary and the identical instruction. No model gets a hint, a warmer intro, or a different file. That is the first condition for comparable numbers.
  2. Same fixed judge. The judge is not a rotating panel of humans or a second LLM with a mood. It is a fixed scoring function. Deterministic token similarity means two runs of the same output produce the same number. That removes judge drift from the leaderboard.
  3. Per-axis rubrics. The benchmark sits in code, reasoning, security, and agent categories. A single blended score would hide which axis a model actually won. Per-axis rubrics keep the dimensions separate so you can see whether a model is strong at reconstruction, at reasoning about control flow, or at operating as an agent.
  4. Public verbatim trail. The benchmark keeps the actual answers readable. Anyone can re-read what a model actually answered. That is the difference between a score and a claim. If a number looks surprising, you can go read the output that produced it.

That combination is what makes LLM evaluation comparable rather than anecdotal. Same input, same judge, separated axes, inspectable output.

The real numbers

Model Provider Score
Nemotron 3 Ultra (free, via opencode) deepseek 43.48991396255626
GLM 5.2 (our own submission) lforla 78.0

We submitted our own model to this benchmark, so the second row is not a neutral third party. It is our own result, published with the same scoring rules. The chart attached to this post plots the leaderboard for the reverse-engineering benchmark.

Why the ranking looks this way

The gap is large: 78.0 versus roughly 43.49. In practice, that means GLM 5.2 recovered substantially more of the reference source under deterministic token similarity. Because the judge is fixed and the references are server-only, the difference is not a matter of style or verbosity. It is a matter of how much of the actual function structure the model reconstructed.

A few things are worth noting about the shape of the leaderboard.

  • The task is unforgiving. Restoring C source from a stripped binary rewards exact structure. A model that writes plausible-looking C but misses the control flow will lose token similarity quickly.
  • The axes matter. Reverse engineering touches code, reasoning, security, and agent behavior. A model can be strong on one axis and weak on another. The overall score compresses that, which is why the per-axis rubrics exist underneath it.
  • One model is not a trend. A single entry at 43.49 does not tell you how the whole field behaves. It tells you how that model behaved on this benchmark under these rules.
  • Our own submission is disclosed. GLM 5.2 at 78.0 is our model. We are not pretending it is an independent audit. The value is that the mechanism is public and the trail is readable.

What this means for practitioners choosing a model

If you are picking a model for reverse engineering or adjacent code tasks, do not read the overall score as a universal ranking. Read it as a signal about this task, under this judge, with these references.

Ask three questions. First, does the benchmark match your workload? Binary-to-source restoration is not the same as writing a web app. Second, can you inspect the verbatim trail? If you cannot read what the model actually answered, you are trusting a number you cannot verify. Third, does the benchmark separate axes? A single score hides whether the model failed at reasoning or at tool use.

The LFORLA benchmark is useful precisely because it answers those questions in public. The same prompts, the same fixed judge, per-axis rubrics, and a readable trail mean you can compare models without taking anyone's word for it.

How to read a leaderboard

  • Check the judge. If the judge is not fixed and deterministic, the numbers are not comparable across runs.
  • Check the prompts. Same prompts for every model is the baseline. Anything else is a different test.
  • Check the axes. Look for per-axis rubrics so you know which capability the score reflects.
  • Check the trail. Find the verbatim answers. If you cannot read them, treat the score as a claim, not evidence.
  • Check the disclosure. Know who submitted what. Our own GLM 5.2 entry is disclosed as ours.

Honest limitations

This benchmark is narrow. It measures one task: restoring C source from stripped binaries. It does not measure general chat quality, long-context reasoning, or safety. The scoring is deterministic token similarity, which is stable but not the same as human judgment of correctness. A model could produce a functionally equivalent function that scores lower because the tokens differ. The leaderboard currently shows a single external model plus our own submission, so it is not a broad field survey. And because we submitted our own model, our result should be read with that disclosure in mind.

Conclusion

LLM evaluation becomes comparable when the mechanism is boring and public: same prompts, same fixed judge, per-axis rubrics, and a verbatim trail anyone can re-read. The LFORLA Reverse Engineering benchmark does that, and the current numbers show GLM 5.2 at 78.0 against Nemotron 3 Ultra at 43.48991396255626. Read the trail, not just the score. Explore the benchmark and the full leaderboard at https://lforla.org.

Top comments (0)