DEV Community

Cover image for Why Your RAG Chatbot Gives Wrong Answers (And How to Fix It)
Dharanidharan
Dharanidharan

Posted on

Why Your RAG Chatbot Gives Wrong Answers (And How to Fix It)

You built a "chat with your documents" app. You upload a PDF, ask a simple question, and the answer is wrong, half right, or completely made up.

If this happened to you, do not blame the AI model yet. In most RAG apps, wrong answers come from a few simple mistakes in how the document is prepared and searched. They are easy to fix once you know where to look.

This is part 2 of my RAG series. In part 1 we built a simple "chat with your PDF" app. Here, we fix the problems that show up when you test it with real documents.

First, Find Where the Problem Is

A RAG app does two jobs:

  1. Retrieval: find the right pieces of the document for the question.
  2. Generation: use those pieces to write the answer.

Most wrong answers are retrieval problems. The model never saw the right text, so it could not answer well. Before you change anything, print the pieces your app found:

relevant_chunks = search(index, chunks, query_embedding, top_k=3)

for chunk in relevant_chunks:
    print(chunk[:200])
    print('---')
Enter fullscreen mode Exit fullscreen mode

If the answer is not in those pieces, fix retrieval (problems 1 to 4 and 7 below). If the answer is there but the reply is still wrong, fix the prompt (problems 5 and 6).

Problem 1: Chunks Are Too Big or Too Small

If chunks are too big, one chunk holds many topics and the search gets confused. If they are too small, a chunk has no useful meaning on its own.

Fix: start with about 300 to 500 words per chunk. Then test two or three sizes with your own questions and keep the one that works best.

for size in [200, 400, 600]:
    chunks = chunk_text(text, chunk_size=size)
    print(size, 'words ->', len(chunks), 'chunks')
    # rebuild the index and test the same 5 questions
Enter fullscreen mode Exit fullscreen mode

Problem 2: No Overlap Between Chunks

When you cut text at a fixed number of words, a sentence can be split in half. The first half sits in one chunk and the second half in another. Neither chunk gives the full answer.

Fix: let each chunk repeat the last 50 words or so of the one before it.

def chunk_text(text, chunk_size=400, overlap=50):
    words = text.split()
    chunks = []
    step = chunk_size - overlap
    for i in range(0, len(words), step):
        chunk = ' '.join(words[i:i + chunk_size])
        if chunk:
            chunks.append(chunk)
    return chunks
Enter fullscreen mode Exit fullscreen mode

Problem 3: The Text From the PDF Is Messy

This one is very common, and people forget to check it. If the text you extract is bad, everything after it is bad too. Look out for:

  • Scanned PDFs. They are pictures, so the extractor returns empty text.
  • Headers and footers. Lines like "Company Name | Page 4" repeat on every page and fill your chunks with noise.
  • Tables. They often turn into broken lines of numbers.

Fix: always print the extracted text and read it. Then clean it.

import pdfplumber
from collections import Counter

def extract_pages(pdf_path):
    pages = []
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            pages.append(page.extract_text() or '')
    return pages

def remove_repeated_lines(pages, min_pages=3):
    counts = Counter()
    for page in pages:
        for line in set(page.split('\n')):
            counts[line.strip()] += 1
    repeated = {line for line, n in counts.items() if line and n >= min_pages}
    cleaned = []
    for page in pages:
        lines = [l for l in page.split('\n') if l.strip() not in repeated]
        cleaned.append('\n'.join(lines))
    return cleaned

pages = extract_pages('manual.pdf')
empty = [i + 1 for i, p in enumerate(pages) if len(p.strip()) < 20]
if empty:
    print('No text on these pages (maybe scanned):', empty)
Enter fullscreen mode Exit fullscreen mode

For scanned pages, use an OCR tool such as pytesseract to turn the pictures into text first.

Problem 3 and a Half: Tables

If your documents are full of tables, plain text extraction will not be enough. Try page.extract_tables() in pdfplumber and turn each row into a short sentence, like "Plan: Basic, Price: 500". Search works much better on sentences than on broken columns.

Problem 4: Wrong Number of Chunks (top_k)

top_k is how many chunks you send to the model. If it is too low, you miss part of the answer. If it is too high, you fill the prompt with noise and the model gets distracted.

Fix: start with 3 to 5. If your questions need facts from different parts of the document, go up. If answers look unfocused, go down.

relevant_chunks = search(index, chunks, query_embedding, top_k=4)
Enter fullscreen mode Exit fullscreen mode

Problem 5: The Prompt Lets the Model Guess

If your prompt only says "answer this question", the model will happily use what it learned in training and mix it with your document. That is how made-up answers happen.

Fix: tell it clearly to use only the context, and what to say when the answer is missing.

prompt = f'''Answer the question using ONLY the context below.
If the answer is not in the context, say: 'I could not find this in the document.'
Do not use any outside knowledge.

Context:
{context}

Question: {question}
Answer:'''
Enter fullscreen mode Exit fullscreen mode

Problem 6: No "I Don't Know" Path

A vector search always returns something, even when nothing in the document is close to the question. Without a check, your app sends useless chunks to the model, and the model tries its best to answer anyway.

Fix: add a distance cutoff. If no chunk is close enough, skip the model and say so.

def search_with_cutoff(index, chunks, query_embedding, top_k=4, max_distance=1.2):
    distances, indices = index.search(np.array([query_embedding]), top_k)
    results = []
    for dist, i in zip(distances[0], indices[0]):
        if i != -1 and dist <= max_distance:
            results.append(chunks[i])
    return results

relevant = search_with_cutoff(index, chunks, query_embedding)
if not relevant:
    print('I could not find this in the document.')
Enter fullscreen mode Exit fullscreen mode

The right max_distance depends on your embedding model, so 1.2 is only a starting point. Print the distances for a few questions that the document answers and a few that it does not. Then pick a number between the two groups.

Problem 7: Wrong or Mixed Embedding Models

There are two mistakes here:

  • Using different models for the document chunks and the questions. The numbers will not match, and search will give random results. This often happens when you change the model but forget to rebuild the index.
  • Using an English-only model for other languages. If your documents are in Tamil, Hindi, or another language, an English model will search badly.

Fix: use one model for both, and rebuild the whole index if you change it. For other languages, try a multilingual model.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2')

doc_vectors = model.encode(chunks)           # for the document
query_vector = model.encode([question])[0]   # for the question, same model
Enter fullscreen mode Exit fullscreen mode

Problem 8: No Sources, So No Way to Check

If your app only shows an answer, you cannot tell if it is right. And when it is wrong, you cannot tell why.

Fix: save the file name and page number with every chunk, and show them with the answer.

chunks_with_meta = [
    {'text': chunk_text_value, 'file': 'manual.pdf', 'page': 12},
    # one item per chunk
]

sources = [f"{c['file']}, page {c['page']}" for c in relevant]
print('Answer:', answer)
print('Sources:', ', '.join(sources))
Enter fullscreen mode Exit fullscreen mode

Sources help your users trust the answer, and they help you debug it faster.

Your Debugging Checklist

When your RAG chatbot gives a wrong answer, go through this list in order:

  • [ ] Print the extracted text. Does it look right?
  • [ ] Print the retrieved chunks. Is the answer inside them?
  • [ ] Is the chunk size around 300 to 500 words, with some overlap?
  • [ ] Is top_k between 3 and 5?
  • [ ] Did you use the same embedding model for chunks and questions?
  • [ ] Does your prompt say "use only the context" and give an "I could not find this" line?
  • [ ] Do you have a distance cutoff?
  • [ ] Are you showing sources with every answer?

One more tip: write 10 test questions with known answers and run them after every change. That way you will know if a fix helped or made things worse.

Wrap-Up

Most RAG problems are not about the AI model. They come from how the text is cleaned, split, searched, and passed to the model. Fix those, and your answers get much better.

If you would like a working codebase to start from instead of fixing all of this yourself, I made a RAG chatbot source-code kit: RAG Source code.

Which of these problems hit you the most? Tell me in the comments, and I will try to cover it in the next part.

Top comments (0)