DEV Community

Alexandra
Alexandra

Posted on AI-assisted

Confident Isn't Accurate: How AI Hallucinations Actually Work

AI is moving relatively fast, despite being slow down. Yes, it feels like there's a new concept to learn every week day. In an effort to actually understand this mad new world instead of just skimming past it, (and to remind myself there's not an 8-ball inside the machine), I've been writing ELI5 articles breaking down concepts that show up constantly. This time, everyone's favorite word: hallucination.

Hallucination Gif

AI is not lying to you

I hate the anthropomorphism traits we give to these AI tools. "Hallucination" makes it sound like the AI is a sentient being going mad, or a bug lingering in the codebase of a frontier model. Reality is much less dramatic, as it always is with AI. It would be more accurate to think of it as the same mechanism that makes AI useful at all. It is a confident sounding answer to a question that the model doesn't have a reliable answer for.

I wrote an article about tokens some months ago, a tl;dr; about it: a model doesn't know anything and it doesnt "search" for things. It predicts the next most likely token, based on patterns learned from enormous amounts of text (or media in general). There is not (almost - let's not count RAG) an intermediate step that it checks if the answer is true against of a database or facts. Actually, There's no database of facts. There's only "given everything so far, what's statistically possible to be next."

The phrases "statistically likely" and "actually true" is close enough, because the training data mostly reflects reality. But the key is this word: "mostly", a hallucination is what happens in the gap between those two things.

The mechanism

ACME corp..

Let's see an example, let's imagine for a moment that we are asking a given model who the CEO of Acme Corp is. The model is not "retrieving an answer", it's ranking candidate next words by how likely each one is to continue the sentence.

Depiction of the next likely token

Can you notice what's missing from that ranking: the "I don't know" candidate is not anywhere near the top. Confident, possible completions are common in training data and uncertainty is very rare. That happens because people don't usually write in their scientific papers "I don't actually know the answer to this" in the confident, encyclopedic-style text models learn from.

Chat models get some extra training that teaches them to decline sometimes. But there is a catch: in most benchmarks score we see, the "I don't know" is equal to zero, which is the same as a wrong answer. A model that guesses looks better on the leaderboard than one that admits uncertainty (oh well...are we even surprised?).

The result isn't that the model "decides" to lie. We get the highest probable word/sentence for the continuation of a sentence and this word/sentence stated confidently as the correct answer.

Why it happens more in some situations than others

The thing is that the hallucinations don't always happen. Some situations make them much more likely:

  • Rarely discussed facts or contradicting opinion topics If something doesn't appear often in the training data, or the answer is ambiguous across sources, the model has lower statistical proof. That means that there is more room for a wrong-but-confident completion to rank higher than the right one.

  • Specific numbers, dates, and citations Precise details are where hallucination answers are most costly and most common. A citation with a real-sounding title, author, and journal name can be entirely fabricated, because the model is generating something like "what a citation looks like" not retrieving an actual paper.

  • Anything past the training cutoff The model has no way to know about "the now", but it also cannot say "stop, this is outside what I know, let me search it" the way an actual person might. Modern chat models often do flag their cutoff, but they can't reliably tell which facts have gone stale, so an outdated answer comes out sounding just as confident as a current one.

  • Questions with a false premise When you ask a model "why did X happen?" when X never happened, and the model often plays along and explains it. Models are trained to be overly helpful, and accepting the premise is the helpful-sounding path.

  • Made-up code dependencies. Models can suggest package names that don't really exist. Attackers can register those names and wait for someone to npm install them (this is known as my fave word of 2026: "slopsquatting").

Some things that might reduce this effect

There is this architecture concept called Retrieval-augmented generation or commonly known as RAG. Instead of asking the model to generate an answer purely from its training data, you constrain it in real source material and ask it to answer from that. The model still predicts tokens, but now it's predicting tokens in a text that hopefully is real and verified. This doesn't solve 100% the problem though. Retrieval can still fetch the wrong document, and the model can still misread or overstate what the right document says. Legal research tools built on RAG and marketed as "hallucination-free" were still found to hallucinate in roughly 1 in 6 to 1 in 3 answers.

A few other things that might help:

  1. Ask the model to give sources and citations. It doesn't guarantee the accuracy of the answer but it is a way to double check if the result is correct. A claim with no source is unverifiable by design; a claim with a specific source at least gives you a way to catch an error.

  2. Instruct the model to say "I don't know". Prompting a model to not make up answers nudges the probability distribution in that direction, but it doesn't give the model actual knowledge of what it doesn't know.

  3. Ask more than once. Sample the same question several times. If it comes back with a different answer each time, the model is probably making it up. Researchers turned this idea into a detection method called semantic entropy.

What this means if you're building AI features

If you're a product engineer shipping AI features, this isn't just the ML team's problem. Some guidelines for shipping accurate features:

  • Never present AI-generated text as verified fact, especially for anything specific, numeric, or citation-like. According to EU AI act, you have to mark that something is AI generated.
  • Make sources visible and clickable when they exist. If your product does information retrieval, surface what was retrieved.
  • Design features for correction, not just generation. A "this doesn't look right", "looks like AI slop" or feedback affordance is extremely useful to tune the AI output than it did for traditional software, because confident-sounding wrong answers are an expected part of how these systems work.
  • Check that dependencies exist before installing them. If an agent or copilot adds a package, confirm it's real!

Resources

Why language models hallucinate (OpenAI) ยท arXiv paper
OpenAI admits AI hallucinations are mathematically inevitable (Computerworld)
How Anthropic's Claude Thinks (ByteByteGo)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
Detecting hallucinations using semantic entropy (Nature, via OATML)
Hallucination-Free? Assessing Leading AI Legal Research Tools (J. Empirical Legal Studies)
AI Hallucination Cases Database (Damien Charlotin) ยท HAQQ tracker summary
Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

Top comments (0)