DEV Community

pyfile-toolkit
pyfile-toolkit

Posted on

I built an LLM benchmark you cannot train on

Every few months a model "scores 92% on MMLU" and a month later it fails in production on things the benchmark said it could handle. The score was real. The evaluation wasn't — the questions had been sitting on GitHub and arXiv for years, and the model had memorized them.

This is benchmark contamination, and it's a recognized problem (CoDeC, ICLR 2026; DCR, EMNLP 2025; LiveBench, ICLR 2025). LiveBench's answer is to rotate questions monthly. Mine was to remove the attack surface entirely.

The idea

Generate the tasks from a random seed at run time, and never publish them.

If the evaluation set does not exist until the run starts, there is nothing to memorize. The model can only solve.

architecture

Concretely:

  • 225 task worlds deterministically generate tasks from a seed. Nothing is stored in advance.
  • Reports contain aggregates only. No individual question is ever revealed.
  • commit-reveal: sha256(seed) is published before the run, the seed after. This proves the set wasn't swapped.
  • Three checker kinds: formal (compare to reference), differential (the answer is code, executed on random inputs against a reference), and open (no reference — routed to human A/B voting).

Measuring contamination directly

Here's the part I'm most happy with. A task depends on the seed, and a solver must solve it for any seed. So I can regenerate tasks from a previous run byte-for-byte (from a ref) and mix them into the current set:

candidate fresh old gap verdict
honest solver 100% 100% +0.0 pp solves
memorizer 0% 100% +100.0 pp recalls

gap > 15 pp is a reliable sign the model is optimized for the benchmark rather than solving it. No n-gram scanning, no perplexity comparison against a reference model — just "are you better on questions you've already seen?"

Real run

Reference solver is my own deterministic code (0 LLM calls). Real models are called directly through provider APIs:

candidate accuracy notes
script:all 99.7% deterministic solver (reference)
gemini-3.6-flash 100%* small sample, quota-limited
gpt-oss-120b (groq) 93.7% 6 runs
memorizer (test) 50% gap +100 pp → flagged

*Technical errors (HTTP 402/429) are counted separately and excluded from accuracy — a model that was rate-limited is not a model that failed. That distinction turned out to matter a lot.

What's inside

  • 225 worlds: math, algorithms, graphs, SQL, compilers, OS, crypto, physics, chemistry, biology, geography, linguistics, economics, finance, logistics, marketing, medicine, law, games, puzzles, probability, electronics, music, cooking, sports…
  • IRT calibration: a Rasch model gives a candidate θ that's comparable across rounds even when the task sets differ.
  • 23 meta-tests that check the framework itself: determinism, no answer leakage into the dump, every world's solver agreeing with its checker, differential catching wrong code.

Try it

node scripts/bench/run.mjs --candidate=script:all
node scripts/bench/run.mjs --seed=myrun --candidates=direct:groq:openai/gpt-oss-120b
node scripts/bench/test.mjs
Enter fullscreen mode Exit fullscreen mode

Repo: https://github.com/pyfile-toolkit/unpredictable-bench
Site: https://pyfile-toolkit.github.io/unpredictable-bench/

It's early. If you've seen a leaderboard-vs-production gap and want to know whether a model learned the skill or memorized the answers, I'd like to hear how you'd want to use this. And if you have provider keys to spare, independent runs are the most useful contribution right now.

Top comments (0)