A model scores 90% on MMLU. It scores 95% on GSM8K. It scores 98% on HumanEval. The press release says: "State-of-the-art. Approaching AGI." You try the model yourself. It's good. It's not that good. It makes mistakes on simple tasks. It struggles with reasoning. It fails on tasks that are not in the training data. The benchmarks are lying. They are broken.
This is the problem with standard benchmarks. They are saturated. They are leaky. They reward memorization. They do not measure true capability.
The Problem with MMLU
MMLU (Massive Multitask Language Understanding) is a multiple-choice benchmark.
The Concept:
The model is tested on a wide range of subjects.
It must choose the correct answer from four options.
It is a test of knowledge.
The Problem:
The questions are static.
The answers are known.
The model can memorize them.
The Consequence:
MMLU is saturated.
Scores are near ceiling.
It no longer differentiates models.
A Contrarian Take: MMLU Is Not Broken. It Is Outdated.
MMLU is not broken. It is outdated. It was designed for earlier models.
It needs to be replaced with a more challenging benchmark.
The Problem with GSM8K
GSM8K is a math word problem benchmark.
The Concept:
The model is tested on grade school math problems.
It must solve the problem step by step.
It is a test of reasoning.
The Problem:
The problems are static.
The solutions are known.
The model can memorize them.
The Consequence:
GSM8K is saturated.
Scores are near ceiling.
It no longer differentiates models.
A Contrarian Take: GSM8K Is Not Broken. It Is Outdated.
GSM8K is not broken. It is outdated. It was designed for earlier models.
It needs to be replaced with a more challenging benchmark.
The Problem with HumanEval
HumanEval is a code generation benchmark.
The Concept:
The model is tested on coding problems.
It must generate code that passes tests.
It is a test of programming ability.
The Problem:
The problems are static.
The solutions are known.
The model can memorize them.
The Consequence:
HumanEval is saturated.
Scores are near ceiling.
It no longer differentiates models.
A Contrarian Take: HumanEval Is Not Broken. It Is Outdated.
HumanEval is not broken. It is outdated. It was designed for earlier models.
It needs to be replaced with a more challenging benchmark.
The Common Problems
All three benchmarks share common problems.
- Data Leakage:
The test data is in the training data.
The model has seen the questions before.
It is not being tested on new material.
- Saturation:
The scores are too high.
There is no room for improvement.
The benchmark is no longer useful.
- Memorization:
The model can memorize the answers.
It is not reasoning.
It is just retrieving.
A Contrarian Take: The Problems Are Not the Benchmarks. They Are the Models.
The problems are not the benchmarks. They are the models. The models are too good.
The benchmarks are fine. The models have outgrown them.
What Proper Evaluation Looks Like
Proper evaluation requires new benchmarks.
- Dynamic Benchmarks:
The questions are generated on the fly.
The model cannot memorize them.
It must reason.
- Adversarial Benchmarks:
The questions are designed to trick the model.
They test the model's limits.
They expose weaknesses.
- Open-Ended Tasks:
The tasks are not multiple-choice.
They require generation.
They test creativity and reasoning.
A Contrarian Take: Proper Evaluation Is Not About Benchmarks. It Is About Tasks.
Proper evaluation is not about benchmarks. It is about tasks. The model should be tested on real-world tasks.
The tasks should be relevant to the model's intended use.
The Future of Evaluation
The future of evaluation is uncertain.
Near Term (1-3 Years):
New benchmarks will be developed.
They will be more challenging.
They will be more dynamic.
Medium Term (3-7 Years):
Evaluation will be automated.
It will be continuous.
It will be adaptive.
Long Term (7-10 Years):
Evaluation will be integrated into training.
It will be a feedback loop.
It will be seamless.
A Contrarian Take: The Future Is Not Benchmarks. It Is Use Cases.
The future is not benchmarks. It is use cases. The model should be evaluated on its ability to perform real-world tasks.
The evaluation should be relevant to the user.
What This Means for You
You are a consumer of AI. You need to be skeptical of benchmarks.
- Look Beyond the Scores:
The scores are not the whole story.
Test the model yourself.
- Understand the Limitations:
The benchmarks are outdated.
They do not measure true capability.
- Demand Better Evaluation:
Ask for more transparent evaluation.
Ask for more relevant evaluation.
The Last Benchmark
The last benchmark is not a test. It is a choice.
You ask: "How good is this model?"
The AI says: "It depends."
You realize: The question is not about the model. It is about the task.
If you could design a benchmark to test a model's true capability, what would it look like? And why?
Top comments (0)