DEV Community

Cover image for Over-Reasoning in AI: When More Thinking Stops Helping
Sam Morris
Sam Morris

Posted on Originally published at nenspace.com

Over-Reasoning in AI: When More Thinking Stops Helping

Reasoning is not the sole variable in a model's value to a human. Some models spend more computation on reasoning, and some produce longer answers. These are different things. By over-reasoning, I mean effort or explanation that goes beyond what a situation needs. A long answer alone does not reveal how much internal reasoning occurred. nen is trained to refuse that axis. That is a design aim, not a claim to match every model on every task. The seeing leads; the length follows.

By Sam Morris, founder of nenspace. Originally published on nenspace.

When does more AI reasoning help?

The last two years of the frontier have been a race to productise deliberation. Chains of thought grew into extended thinking, extended thinking grew into minutes of it, and reasoning benchmarks became a prominent scoreboard for the field. On its own terms, the race is real.

OpenAI reported that on the 2024 AIME mathematics exam, GPT-4o solved about 12% of problems, while o1 reached 74% with one sample and 83% with consensus across 64 samples. Those are different models and evaluation settings. They do not show a single model rising from 12% to 83% through extra thinking alone. They do show why more reasoning compute on difficult, checkable problems deserves serious attention. The argument here is not that reasoning stopped working.

It is about where people actually live. A large study of ChatGPT use found that practical guidance, seeking information and writing made up nearly 80% of conversations in the sample studied. Programming accounted for about 4% of messages. Those are topic categories, not measures of how much reasoning a response required. Yet much of what a person brings to a model in a day is not a competition problem. It has no answer key. It has a reader.

A hard proof, a consequential design decision and a quick rewrite call for different amounts of work. The right amount cannot be inferred from a single benchmark or a preference for terse prose. It follows from the task, the stakes and what the person needs to do with the answer.

What is over-reasoning, and can a long answer reveal it?

The first cost is the one everyone already knows in their hands: six paragraphs with a citation stapled to every clause; thirty seconds of visible thinking on a question that needed none; the answer to the question asked buried under answers to four that were not. Past competent, more reasoning can stop being help and start being a reading assignment. The model performs its diligence. The reader pays for that performance in attention.

But that experience joins several things that should be measured separately. Internal inference effort, visible reasoning text and final answer length are not identical. A model can use substantial computation and give a short, clear answer. A verbose answer can be generated without deep deliberation. Answer length alone therefore cannot diagnose how much internal reasoning happened. It can, however, tell us how much attention the answer asks of its reader.

Research on reasoning models has found cases where more is worse. In a Meta FAIR study, shorter reasoning chains sampled for the same question were up to 34.5% more accurate than the longest chains. That is a result within the authors' sampled questions and methods, not a universal rule to always pick the shortest chain. In agentic task experiments, selecting against an overthinking measure improved performance by almost 30% while reducing compute costs by 43%. Those figures describe that evaluation and selection procedure.

A 2026 paper accepted to ICML reports that raw generation length does not consistently correlate with accuracy on its mathematical and scientific benchmarks. Its alternative measure of deeper internal revisions correlated more strongly. That is another reason not to confuse tokens produced with useful thought. The papers make a more precise case than “short is good”: some extra work improves an answer, some wastes compute, and some correlates with worse results in particular settings.

One analysis of preference tuning also found that a reward based only on response length reproduced much of the measured gains from RLHF in the authors' tests. That is an uncomfortable measurement problem. If a benchmark or preference process rewards expansive answers, improvement on that measure may partly be a change in style. It does not follow that a longer answer is less accurate or less useful. The question is what its extra words accomplish.

The Arena chart discussed in the original essay shows words per answer rising across successive releases of one frontier model line, while the share of longer content words falls. Those measures describe answer style. They cannot establish the substance, accuracy or usefulness of the answers. The practical test still belongs to the reader: did the additional text help?

Why can an answer that feels good still be a poor one?

The quieter cost is the possibility of accepting a model's conclusion without examining it. Whether repeated use changes a person's abilities depends on the task and the way the tool is used; the long-term question remains open.

Models are tuned against human approval, and approval is an imperfect signal. The length study above shows how easily verbosity can enter the reward. Agreement can enter too. Anthropic researchers found that people and preference models sometimes favoured convincingly written sycophantic responses over correct ones. A 2026 formal analysis describes conditions under which optimization against biased preferences amplifies that drift. These results concern sycophancy; they do not prove that reasoning volume causes the same problem. They show why preference alone is a poor stand-in for what serves someone.

A Stanford and CMU study examined 11 models and found that, in its test setting, they affirmed users' actions 50% more often than human respondents. In two preregistered experiments, participants who interacted with sycophantic AI were less willing to repair an interpersonal conflict and more convinced they were right. They also rated the sycophantic responses higher quality, trusted them more and wanted to use them again. In that setting, what drew people back diverged from what helped them respond constructively.

The long-term cognitive cost is still being measured. An MIT Media Lab preprint reported differences in EEG connectivity and recall during an essay-writing task. Fifty-four participants completed its first three sessions; 18 completed the fourth. It is a small, task-specific study, not proof of lasting cognitive decline and not evidence that verbose answers caused its effects. It is a reason to study how assistance is used, not a verdict on every AI conversation.

What should an AI answer do instead?

Refusing the axis is not refusing to reason. nen is the default text-dialogue model in nenspace. Separate web-search and deep-research capabilities are available when a question needs sources. The refusal is of reasoning as performance, length as a proxy for effort, and deliberation as display.

The seeing leads and the length follows. Sometimes the right response is one line that moves the frame you were standing in. Sometimes it is a page: a worked explanation, a careful comparison, or a complete rewrite. Length follows the task. What never leads is the volume itself.

This is the same bet as the rest of nenspace, stated for the voice. Keep useful information outside your head while continuing to examine the judgement yourself. /space gives a conversation somewhere to persist; nen is built to hand the thinking back. A useful answer can be saved and returned to. A poor answer can be challenged or discarded. Neither outcome should be hidden behind an impressive amount of prose.

The lo-fi of LLMs: not low quality, but a refusal to assume that more always makes it better. The why page explains the register, and the AI sycophancy guide examines one failure mode in more detail. Try nen with a bounded problem and judge the answer by what it helps you do.

The answer should earn the attention it asks for.

Top comments (0)