DEV Community

Cover image for Why Models Are Getting Dumber on Purpose?
Kumar Kislay
Kumar Kislay

Posted on Originally published at forg.to

Why Models Are Getting Dumber on Purpose?

Models Are Getting Dumber on Purpose

Reasoning scores keep climbing while per-token compute keeps falling. GLM-5.2 scores 99.2% on AIME 2026 with about 40 billion parameters active per token. Qwen3.5 hits 91.3% with 17 billion active. DeepSeek V4-Flash runs 13 billion active. For scale, GPT-4 was rumored to run around 280 billion active parameters in 2023, and it could barely finish an AIME problem. At the small end, Qwen3.5 9B fits in 6GB of VRAM quantized and roughly doubles the score of the next best model under 10B on Artificial Analysis's intelligence index.

Look only at math and code and you would conclude that intelligence per parameter is improving at an absurd rate.

Now ask the same models a plain factual question. On SimpleQA, closed book with no tools, the current leader is Gemini 2.5 Pro at 53%. The best factual recall money can buy still misses half the questions. The small models barely register: Artificial Analysis measures Qwen3.5 4B and 9B at hallucination rates of 80 to 82% on its knowledge benchmark. When they don't know a fact, which is most of the time, they invent one. Ask the 9B for the birth year of a minor 19th-century mathematician and you get a confident, plausible, wrong answer.

The parameter count didn't drop for free. Labs are trading world knowledge for reasoning skill, and the trade is deliberate.

What the parameters were for

Facts take space. The cleanest measurements I know of come from the Physics of Language Models series, which puts the ceiling at roughly two bits of factual knowledge per parameter. If you want a model that knows the birth year of every minor Wikipedia figure, the population of every Dutch municipality, and the argument order of every function in every npm package, you pay for that in weights. That is a large part of why frontier models grew into the trillions.

Reasoning compresses much better, because it is a small set of procedures applied over and over: split the problem, track intermediate state, check your own work, backtrack when a step fails. Distillation and reinforcement learning on verifiable tasks transfer those procedures into small models surprisingly well. Phi-4 is 14 billion parameters, trained heavily on synthetic textbook-style data, and it is good at math and bad at trivia. That tells you exactly what was in its training set. Two years ago that profile read as a limitation of synthetic data. Now it reads as the spec.

The knowledge that survives the trade has a shape worth noticing. These models are generalists: a little about nearly everything, almost nothing in depth. Ask one about PostgreSQL and it knows what it is, what it is good at, and roughly how MVCC works. Ask which version added a specific planner feature and you are back to invented facts.

I think that is the correct layer to keep in weights. Breadth is what lets a model understand what a question is about, know what to go look up, and judge whether a source is plausible. Depth is cheap to retrieve and expensive to store, so depth is the part that goes.

Facts rot, procedures don't

A frontier training run takes months and costs hundreds of millions of dollars, and the facts inside it start going stale the moment it finishes. APIs change, prices change, people change jobs. Half of what a 2024 model believed about the JavaScript ecosystem was already wrong on release day. Every fact baked into weights has a shelf life, and the only refresh mechanism is another training run.

Procedures don't rot. Algebra worked the same way in 1970. So did breaking a problem into parts, or spotting a contradiction between two sources. A model that is mostly procedure and only lightly loaded with facts does not age the way a knowledge-heavy model does. Its training cutoff matters much less, because the current state of the world was never supposed to live in the weights.

This is the strongest argument for the whole approach, and it is the one that convinced me. It decouples the expensive, slow artifact from the thing that changes daily.

The harness carries the knowledge

If the model doesn't know things, something else has to. That something is the harness: retrieval over a knowledge base, tool calls, web search, a filesystem full of docs. I wrote earlier that Rust is a harness for agents, a source of cheap machine-checkable feedback. This is the same shape viewed from the other side. The model supplies reasoning. Everything it reasons about gets supplied at runtime.

You can already watch agents work this way. A coding agent does not need your dependency's API surface memorized, because it greps node_modules or reads the docs before it calls anything. Its answer is grounded in the version you actually have installed, not whichever version dominated the training corpus. Recall that used to be a fixed cost in every forward pass became an on-demand lookup.

Knowing less only helps if the model knows that it doesn't know

This is the part I think most of the optimistic writing on this skips, including my own earlier drafts.

A small model paired with a good harness is only safe if it reliably says "I don't have this, let me look it up." Abstention is not a free side effect of having fewer parameters. It is a separate trained behavior, and the industry has spent years training the opposite.

OpenAI's 2025 paper on hallucination makes the mechanism plain: most benchmarks use binary scoring, so a wrong answer and an abstention score identically at zero. Under those rules a model that always guesses beats an otherwise identical model that sometimes admits uncertainty. We built the leaderboards to reward bluffing and then acted surprised when models bluff.

Shrinking the model makes this worse before it makes it better, because you have widened the gap between what the model is asked and what it holds. Those 80% hallucination rates on small models are not a knowledge problem. They are a calibration problem sitting on top of a knowledge problem. The knowledge half is the part we are deliberately removing. The calibration half is the part that has to be solved, and it does not get solved by making the model smaller.

A frontier model on your GPU

Follow the trend out a couple of years and I think you get frontier-quality reasoning on a single consumer GPU.

The compute half is close. DeepSeek V4-Flash reasons with about 13 billion active parameters per token, comfortably inside consumer range. What does not fit is the other 271 billion sitting in its experts. Note that this is a memory problem, not a compute problem: in a mixture of experts you only multiply through a slice of the weights per token, but the whole thing still has to be resident somewhere the GPU can reach. Active parameters set your speed. Total parameters set whether you can run it at all.

Expert layers are mostly fact storage, and fact storage is exactly what this trade makes optional. Strip it out and total size collapses toward active size. A 20 to 40B model at 4-bit quantization fits on the 24GB card that has been sitting in gaming PCs since 2022.

The catch is that it will not know much. Asked a bare factual question with no tools attached, the correct behavior is to say so and go look. Paired with a decent harness, that covers most of what I use a frontier model for today, running locally, with no per-token bill and nothing leaving the machine.

Where this breaks

Two honest problems, since I would rather state them than have them stated back to me.

First, retrieval is not free. Every fact that used to be one forward pass away is now a round trip and a few thousand tokens of context. You have not eliminated the cost of knowledge, you have moved it from weights into the context window, where it is paid again on every single query instead of once at training time. For a local model with no per-token bill that is a good trade. For a high-volume API workload it is much less obviously one.

Second, judgment needs a floor of knowledge. To evaluate a retrieved document you need enough background to tell a good source from a bad one, and to notice when a question contains a false premise you need to already know the thing being assumed. Cut too deep and you get a model that reasons beautifully over whatever garbage the retriever handed it. Nobody has published a clean number for where that floor sits, and I suspect it is task-dependent enough that nobody will.

This mostly fixes hallucination

Still, the direction is right, and the reason is what it does to wrong answers.

When a fact lives in weights, a wrong fact is unfindable and unfixable. You cannot grep the weights, you cannot diff them against last month, and correcting one error means a fine-tune that might break who knows what else. The model states a wrong fact with the same fluent confidence as a right one, and there is no artifact to check it against.

When the fact lives outside the model, a wrong answer has an address. The model cites a document, so you open the document. If the document is wrong you edit it, and every future query gets the correction instead of waiting a year for the next training run. Retrieval does not get you to zero, since a model can still misread a source or stitch two together wrong. But a claim with a source is checkable and a claim from weights is not. A wrong fact in a knowledge base is an ordinary data bug, the kind we already know how to trace, fix, and write a regression test for.

There is a version of this where the model card stops listing a knowledge cutoff entirely, because what is left in the weights goes stale on a scale of years instead of weeks. The model just gets handed the current state of the world at runtime, the way a CPU gets handed a program.

Top comments (0)