Chatbots That Don't Guess: Grounding AI in Customer Service
A support bot should answer only from reviewed content, cite its source, and say so when it doesn't know. Here's how to design and test for that.
If you build or buy a chatbot for customer service, the single most important design decision is not the model. It's where the bot is allowed to get its answers. A bot that answers freely can sound confident and still be wrong, and the company, not the vendor, has to stand behind every answer it gives a customer.
What "grounded" means
A grounded chatbot may only answer with content from a defined source, usually the company's knowledge base, and must be able to show where each answer came from. In practice:
- It answers only from published articles. Asked about the return period, it retrieves your returns article and phrases the answer from it. If the knowledge base says nothing about returns, it must not invent a period.
- It cites its source. The customer can see which article the answer is built on and click through.
- It admits missing material. "I don't have an answer to that, let me connect you" is a correct answer. A guess is not.
The usual technique is retrieval-augmented generation (RAG): retrieve relevant passages first, then let the model write an answer from exactly those passages. A general-purpose model without that link answers everything, whether or not it has your terms in front of it. That's why a reviewed knowledge base is a prerequisite, not an option. It also needs to cover exceptions: if delivery terms differ by country, the bot should ask which country applies rather than pick one.
General questions vs. personal cases
An article is enough for general questions. A personal case needs data about that specific customer.
| Customer asks | Bot needs |
|---|---|
| How do I make a return? | A current, approved returns article |
| Where is my refund? | Verified identity and live payment data |
| Does this feature work on my plan? | The product rule and possibly account data |
| Can you make an exception? | A clear limit and a human who decides |
If a required piece of data can't be fetched, the bot should say so and show the next step. Keep three levels apart and decide per flow which one applies: a draft an agent reviews, an automatic answer sent straight to the customer, and an executed action such as a refund. Each level needs more data and tighter permissions. Supportifier's integrations page shows the kind of system connections that personal answers depend on.
Why models hallucinate, and why RAG alone isn't the fix
A hallucination is an answer that sounds credible but has no support in any material. OpenAI researchers described in 2025 how training and evaluation reward guessing over admitting uncertainty, a bit like a multiple-choice exam where a blank never scores but a guess sometimes does.
Retrieval reduces the problem but doesn't remove it. Stanford RegLab's 2024 evaluation found that commercial legal research tools built on retrieval still hallucinated on a meaningful share of test questions. The takeaway: retrieval is necessary, but the bot must also be constrained to what was retrieved and allowed to say the answer is missing.
The company owns what the bot says
In Moffatt v. Air Canada (2024), an airline's chatbot told a traveller that a bereavement discount could be claimed retroactively; another page on the same site said the opposite. The airline argued the chatbot was a separate entity. The tribunal disagreed: the bot is part of the website, and the company is responsible for its statements. In e-commerce, the same mechanism produces invented discount codes, wrong return periods or delivery promises made without checking any system.
There's also a transparency requirement: under Article 50 of the EU AI Act, people must be told they're interacting with an AI system, at the latest at first contact, from 2 August 2026. Supportifier covers the data side in GDPR and AI in customer service.
A ten-question test for any vendor
Don't ask whether a bot "uses AI". Ask what it's allowed to answer and what happens when the answer doesn't exist. Then test it with your own material:
- 5 questions you have clear articles for → the bot should answer and cite them.
- 3 questions you have no article for → it should say so or hand over.
- 2 questions where your own sources contradict each other → it should surface the conflict.
Also ask: Can we see the list of questions it couldn't answer? Who can publish content it uses, and is there review? How is the customer identified before personal data is shown? Can we block topics like compensation, pricing exceptions and legal assessments? How are errors measured after launch?
Design the handover
Hand over automatically when material is missing, always when the customer asks, and as a rule when money, contracts or complaints are involved. A good handover means the customer doesn't repeat themselves (the whole conversation follows into the ticket), the agent sees which article the bot used so the article can be fixed, and the contact route is visible, including outside staffed hours.
The handover is simplest when bot and humans work from the same reviewed content. The same principle applies when AI suggests replies an agent reviews before sending; see why reviewed AI drafts beat full automation at the start.
Checklist
- Every top question has a published article with a named owner.
- Run the ten-question test on your current bot or a vendor demo.
- Write a block list of topics the bot never handles alone.
- Document the handover rule: when, to whom, with what context.
- Review unanswered questions weekly and write the missing articles.
- Track share of answers with a source, share of handovers, and repeat contacts within a week.
Originally published at https://www.supportifier.se/en/blog/grounded-ai-chatbot-customer-service
Top comments (0)