Hallucination is not a LLM bug. It's the price we pay for creativity in probabilistic AI.
Creativity and hallucination come from the same generative behavior. Enterprise systems need to decide where the model may fill in the gaps, and where it must show evidence."
Every team building with AI eventually hits the same moment. The demo looks great, and then the model confidently gives an answer that's simply wrong: a policy that doesn't exist, a number nobody can find, a citation to a paper that was never written.
The usual reaction is to treat this as a bug to be fixed. "We need to eliminate hallucination."
I think that framing leads to the wrong designs. To see why, it helps to understand what a large language model (LLM) is actually doing when it answers you.
How an LLM actually works
A large language model (LLM) is trained on an enormous amount of text. During training, it learns one core skill: given some text, predict what comes next.
When you ask a question, the model doesn't search a database or look up a stored fact. It generates the answer one small piece at a time (called a token, roughly a word or part of a word). At each step, it calculates which tokens are most likely to come next, based on your question, everything it has written so far, and the patterns it learned in training. Then it picks one and repeats.
That's it. A model writing a poem and a model answering "What's our refund policy?" use exactly the same process.
Two things follow from this:
The model's knowledge is stored as patterns, not records. It doesn't keep a copy of every document it saw. It has absorbed statistical patterns about how words, facts, and ideas tend to go together. That's why it can explain, summarize, and connect ideas so well, and also why it can't reliably reproduce a specific fact it saw only once or never saw.
The model is optimized for plausible, not proven. At every step, it's choosing what sounds most likely to come next. It has no built-in step that checks whether the sentence is true. Most of the time, plausible and true line up, because most of what it learned was accurate. When they don't, you get a fluent, confident answer that's wrong.
Why hallucination happens
Put those two facts together and hallucination stops being mysterious.
Ask a model, "What was company X revenue last quarter?" If that number was never in its training data, or it's private, or it changed last week, the model has no record to look up. But it was trained to always continue the text. So it produces what a revenue answer usually looks like a confident sentence with a realistic number.
From the model's point of view, nothing went wrong. It did exactly what it was built to do generate the most plausible continuation. Hallucination is the model filling a gap with something plausible when it doesn't have the facts.
Some other factors make it more likely
- Training cutoff: the model knows nothing after the date its training data ends.
- Rare or specific facts: names, numbers, dates, and citations seen rarely in training are easy to blend or invent.
- Ambiguous questions: if a term could mean several things, the model picks one, often without telling you.
Training that rewards answering: models are generally tuned to be helpful, and a confident answer often looks more helpful than "I'm not sure."
And here's the key insight: the same mechanism produces creativity. When you ask for campaign ideas, a new analogy, or code for a problem nobody has solved before, you want the model to generate something plausible that isn't stored anywhere. That's the value.
So hallucination isn't a separate defect sitting next to the useful behavior. It's the useful behavior, showing up in a place where we needed facts instead.
The problem isn't creativity
Imagine asking a colleague:
"Give me some ideas for our next marketing campaign."
You want them to explore. You're giving them permission to make connections nobody wrote down.
Now imagine asking:
"What is this customer's current account balance?"
You don't want creativity. You want the exact number from the system of record.
Same AI model. Completely different expectation.
That's why I think "How do we eliminate hallucination?" is the wrong starting question. A better one is
Where can the model reason creatively, and where must the system require evidence?
That question changes how we design the architecture.
Where the gaps come from in enterprise systems
If hallucination is the model filling a gap, the practical question is: where do the gaps come from? In enterprise systems, I usually see three sources.
1. It doesn't have the facts
Maybe the information is private. Maybe it changed this morning. Maybe it lives only in Salesforce, SAP, Snowflake, a document repository, or an internal database.
If the model can't reach the fact but we still expect an answer, it may fill the gap with something plausible.
2. It has the data but not the meaning
Suppose a user asks: "How many active customers do we have?"
What exactly is an active customer? Someone who logged in during the last 30 days? Bought something in the last 90? Has an open contract? Has an account status of ACTIVE?
The model can understand the English words perfectly and still not know what they mean inside your company. That isn't always a knowledge problem. Sometimes it's a semantic problem.
3. We push it to answer
Many prompts quietly say, "Give me an answer." The model tries to be helpful, so even when the evidence is weak, it may produce something that sounds reasonable.
Sometimes the smartest thing an enterprise agent can say is:
"I don't have enough information to answer that."
We should design for that response.
Why agents raise the stakes
Everything above applies to a chatbot. With agentic AI, the stakes go up in two ways.
Errors compound. An agent works in steps: plan, retrieve, call a tool, decide, act. If it guesses at step two, every later step builds on that guess. A small misunderstanding early on can become a confident, well-executed wrong outcome.
Answers become actions. A chatbot that hallucinates gives you a wrong answer you can ignore. An agent that hallucinates might send the email, update the record, or approve the payment.
But agentic design can also reduce guessing. In a pattern often called agentic RAG, the agent doesn't retrieve once and answer. It checks whether the evidence it found actually answers the question. If not, it searches again, tries a different source, calls a tool, or decides there isn't enough information. A single-shot pipeline can't do that.
So agents need more controls, not fewer, and they can also take part in their own verification.
So how do we control hallucination?
1. Give the model the facts with RAG
If the answer lives in contracts, policies, manuals, or internal documentation, retrieve the relevant content first and give it to the model.
Instead of asking, "What do you remember about our travel policy?", we're effectively saying, "Here's the relevant travel policy. Answer using this."
That's much safer, but RAG isn't automatically trustworthy. If the right document never reaches the model, it may still guess, and even with the right document in front of it, a model can occasionally misread or overstate what it says. Retrieval quality matters as much as generation quality, and grounding works best combined with the checks later in this post.
When the answer depends on relationships, consider GraphRAG.
GraphRAG retrieves from a knowledge graph instead of, or alongside, text chunks. The model gets entities and their actual relationships, so it can follow real connections instead of guessing them.
GraphRAG has its own caveat: if the graph is built automatically by an LLM extracting entities from documents, the graph itself can contain errors.
2. Use tools when the answer belongs to a system
Some information shouldn't come from an LLM at all.
Customer balance? Call the banking system.
- Order status? Call the order API.
- Current inventory? Query the inventory system.
- Exchange rate? Call the rates service.
- A calculation? Use deterministic code.
For agentic systems, protocols like MCP (Model Context Protocol) make this pattern easier to standardize.
The LLM becomes the reasoning layer. The system remains the source of truth.
- Give the model business meaning
Semantic layer
A semantic layer defines governed business terms such as Active Customer, Net Revenue, Qualified Lead, Gross Margin, and Churned Account.
Ontology
An ontology describes concepts and how they relate:
- Customer → owns → Account
- Account → contains → Transaction
- Sales Team → manages → Customer
- Product → belongs to → Product Category
That's where ontologies and knowledge graphs help.
They reduce the number of relationships the model has to guess.
4. Let the agent say "I don't know"
Instead of "Answer the user's question," give the agent a policy like:
"Answer only when the required evidence is available. Otherwise, say there isn't enough information."
Knowing when not to answer is part of intelligence.
5. Narrow the possible output
Require structured output:
{
"decision": "NEEDS_REVIEW",
"reason": "Receipt amount does not match submitted amount"
}
Schemas, enums, validation rules, and typed outputs constrain what the model can produce.
Structurally valid doesn't mean semantically correct.
Don't trust a model-generated confidence score.
Structured output is a guardrail, not a truth engine.
6. Verify citations instead of just asking for them
If citations matter, verify them in code.
Check that:
the source really exists,
the quoted passage really appears in it,
the passage actually supports the claim, and
the user has permission to see that source.
The architecture should verify evidence. The prompt shouldn't be the only enforcement mechanism.
7. Verify before the agent acts
Before consequential actions, run deterministic checks.
For higher-risk actions, require human approval.
The pattern should look like:
Reason → Validate → Authorize → Execute → Audit
Not:
Reason → Execute
8. Use temperature carefully
A lower temperature makes outputs more consistent and less varied.
But temperature doesn't give the model more knowledge.
9. Test hallucination instead of debating it
- Build an evaluation set with real questions, including:
- questions the agent should answer,
- questions it should decline,
- questions that need retrieval,
- questions that need a tool,
- ambiguous questions,
- questions with conflicting evidence,
- questions containing misleading information, and
- questions that try to get around policy.
- Because "Our model seems accurate" isn't an AI quality strategy.
Different tasks need different levels of control
Not every AI workload requires the same level of governance and control. The amount of creativity we allow should depend on the nature of the task and the potential impact of mistakes.
For brainstorming, creativity is the primary goal. We want the model to generate new ideas, make novel connections, and explore possibilities. In these scenarios, a light human review is usually sufficient.
For marketing copy, creativity is still important, but the output will eventually reach customers or the public. Human review before publishing helps ensure the content is accurate, on-brand, and appropriate.
For document summarization, creativity should be limited. The objective is to faithfully represent the source material, not invent new information. Grounding responses in source documents and verifying citations are important controls.
For customer support, creativity should be very low. Customers expect accurate answers, not guesses. Using RAG, business tools, clear escalation paths, and allowing the agent to say "I don't know" when evidence is missing are critical safeguards.
For business metrics and reporting, creativity has no place in the answer itself. Terms such as revenue, active customers, or profit margin should come from governed definitions provided through a semantic layer and trusted business systems.
For financial transactions, errors can have significant consequences. Decisions should rely on deterministic validation, business rules, and authorization workflows rather than model-generated assumptions.
For agent actions in enterprise systems, such as updating records, approving requests, modifying data, sending communications, or initiating transactions, creativity should not exist at the point of execution. These scenarios require policy enforcement, deterministic validation, human approval where necessary, and complete audit trails.
This is the key principle: a creative assistant and a transactional agent may use the same LLM, but they should not operate with the same guardrails. The higher the cost of being wrong, the less freedom the model should have to guess.
Not every AI workload needs the same architecture.
A creative assistant and a transactional agent might use the same LLM, but they shouldn't have the same guardrails.
Creativity isn't the enemy
We chose generative AI because we wanted something more flexible than traditional software.
Traditional software says IF X, THEN Y.
Generative AI can work through ambiguity, understand language, summarize information, connect ideas, and adapt.
But enterprise architecture has to decide where flexibility ends.
So I wouldn't design an AI system around the goal "Never hallucinate."
I'd design it around a better principle.
Let the model reason where reasoning adds value. Require evidence where facts matter. Use deterministic systems where actions matter.
In practice, that means:
- Give the agent facts through RAG, and connected facts through GraphRAG.
- Give it live information through tools and MCP.
- Give it business meaning through a semantic layer.
- Give it relationships through ontologies and knowledge graphs.
- Give it permission to say "I don't know."
- Constrain outputs when the choices are bounded.
- Verify evidence.
Let agents check their own evidence and search again when it falls short. Evaluate continuously. And before anything consequential happens, validate it outside the model.
That's how we move from an impressive AI demo to a trustworthy enterprise AI system.
Keep the creativity. Control the guessing.
Thanks
Sreeni Ramadorai

Top comments (0)