Watch the 40-second version: How Transformers Predict Your Next Word
Your AI does not read your prompt. It places bets on it.
Here is the uncomfortable truth about the most sophisticated software you have ever used: it has never once known what it was saying. A transformer takes your prompt, turns every word into a coordinate in a giant space of meaning, runs attention over the whole mess to decide which words deserve to influence the answer, and then rolls weighted dice. One token at a time. Fifty-ish times a second.
That is it. There is no little reasoning engine in there double-checking facts. There is a probability distribution and a sampling function.
Why this explains everything weird about LLMs
Once you see next-token prediction as the mechanism, the famous quirks stop being mysteries and start being obvious consequences.
Bad at math but good at vibes? Math needs exactness; dice need distribution. Asking a token predictor to do arithmetic is like asking a vibe generator to file your taxes. It can approximate the shape of a correct answer, which is why 2+2=4 works and your mortgage amortization does not.
Hallucination is not a bug. It is the mechanism working exactly as designed. The model was trained to produce the most plausible next token, not the most true one. Truth and plausibility overlap often enough to fool you, and diverge exactly where you cannot verify. "Confidently wrong" is just what sampling looks like when the stakes are high.
And the non-obvious angle nobody tells you: this is why prompt engineering works at all. You are not "talking" to the model. You are biasing a probability distribution. Every example you add, every constraint you write, reshapes the odds of the next token. Prompting is statistics with a user interface.
The checklist: force the dice to behave
You cannot make a transformer reason. But you can rig the bet. Copy this before any task where correctness matters:
- Lower the temperature to 0 for factual tasks. High temperature is the model brainstorming. You do not want brainstorming from your accountant.
- Demand structure. Ask for JSON, tables, or numbered steps. Format constraints cut off entire branches of plausible nonsense.
- Make it show work on one line per step. Chain-of-thought is not reasoning, but it gives the model more tokens to condition on, and longer conditioning means tighter distributions.
- Feed it the facts instead of asking it to recall them. This is what RAG is for. The model is a fantastic interpolator and a terrible memorizer, so stop testing its memory.
- Verify the one thing that matters. Sampling errors cluster in details: numbers, names, dates. Scan those first.
The closer
The transformer is a bet-placement machine that learned to sound like a person. Treat it like one and it will con you. Treat it like a loaded probability distribution and it becomes the most useful tool on your desk. The difference is not in the model. It is in what you assumed it was doing.
Top comments (0)