DEV Community

Cover image for 5 RAG mistakes that looked fine in the demo and broke in production
Nicola Mastromarino
Nicola Mastromarino

Posted on

5 RAG mistakes that looked fine in the demo and broke in production

Every RAG demo works. You pick five questions, the system answers them beautifully, and everyone nods.

I've built RAG pipelines over clinical notes and over a knowledge graph, and the gap between "works in the demo" and "works on real questions" is bigger than you'd expect. Real users ask messy questions, use acronyms nobody documented, and expect answers from documents that changed last Tuesday.

Here are five mistakes that pass the demo and fail on real data.

1. Judging quality by vibes

The demo set is five questions you wrote, about documents you know. Of course it works.

Without an evaluation set you can't tell whether a change to chunking or embeddings made things better or just different. On a clinical QA project (about 40K MIMIC-III notes) I tracked Token F1 and hallucination rate against a baseline. That's the only reason I can say QLoRA fine-tuning cut hallucinations by roughly 40% and lifted Token F1 by 190%, instead of "the answers felt better."

The same applies to retrieval. Build a golden set of 30-50 real questions, each mapped to the document that should answer it, and measure retrieval separately from generation. If the right chunk never reaches the model, no prompt will save you.

def hit_rate_at_k(eval_set, retriever, k=5):
    hits = 0
    for item in eval_set:
        docs = retriever(item["question"], k=k)
        if any(d.id in item["relevant_ids"] for d in docs):
            hits += 1
    return hits / len(eval_set)
Enter fullscreen mode Exit fullscreen mode

Twenty lines of code, and it changes how you work.

2. Fixed-size chunking

Splitting every document into 500-token blocks is fast and wrong. It cuts tables in half, separates a list from its heading, and ends chunks mid-sentence. The chunk that says "the limit is 10,000" no longer says what limit.

Chunk along the document's structure (sections, headings, paragraphs), keep the section title attached as metadata, and consider retrieving small chunks but passing the larger parent section to the model.

3. Vector search only

Embeddings are great at meaning and mediocre at exact strings. Error codes, product IDs, contract numbers, internal acronyms: this is exactly what users search for, and exactly where pure semantic search is weakest.

The retriever in my clinical system combines BM25 with a biomedical embedding model, fused with Reciprocal Rank Fusion. Medical text is full of abbreviations and drug names, exactly the kind of string where keyword matching earns its place next to embeddings. Add a reranker on top and you get one of the cheapest upgrades with one of the biggest effects.

4. Letting the model answer no matter what

When retrieval returns nothing useful, most systems still pass the top-k chunks to the LLM, and the LLM still produces a fluent answer. That's the worst failure mode, because it looks like success.

  • Set a relevance threshold and let the system say "I couldn't find this in the documents."
  • Require citations, so a wrong answer is at least checkable.
  • Don't stuff the context with 20 chunks "just in case." More context often means more noise, not more accuracy.

My clinical QA system keeps a human in the loop (it was designed with the EU AI Act in mind). When the domain is medical, a fluent wrong answer isn't a UX bug, it's a risk.

5. Treating the index as finished

The demo index was built once. Real documents change, get deleted, get replaced by newer versions, and some of them shouldn't be visible to every user.

You need an ingestion pipeline, not a script: updates and deletions handled, versions tracked, metadata filters for permissions, and a plan for re-embedding when you change the embedding model. An answer based on a policy that was withdrawn two months ago is worse than no answer.

The pattern behind all five

Every one of these is invisible in a demo and obvious on real data. The demo tests whether the pipeline runs. Real usage tests whether it retrieves the right thing, says no when it should, and stays correct over time.

Build the evaluation set first. Everything else gets easier once you can measure it.

What's the worst RAG failure you've seen that passed the demo? I'd love to hear the war stories in the comments.


Top comments (0)