Your RAG bot just gave a confident, detailed answer. And it's completely wrong.
Here's the part that'll annoy you: the model did nothing wrong. It answered perfectly — using the text you handed it. The bug isn't in the AI. It's in the five steps before the AI.
Prefer to watch? Full 6-minute walkthrough with the "lost in the middle" animation:
The pipeline, in one line of code
RAG is simple on paper. You search your own documents, grab the best matches, paste them into the prompt, and the model answers from those. So when the answer's garbage, everyone blames the model.
Wrong suspect. There are five places this pipeline breaks — and four of them happen before the model even runs:
const chunks = split(docs) // 1 · chunking
const vectors = embed(chunks) // 2 · embedding
const hits = search(embed(query), vectors) // 3 · retrieval
const best = rerank(query, hits) // 4 · ranking
const answer = llm(prompt(best, query)) // 5 · generation
Five lines. Every bug I'm about to show you lives on one of them.
1. Chunking — a fixed cut splits your answer in half
Your documents are long, so you cut them into pieces before you store them. Most people cut every few hundred characters and move on. That's where it starts to rot.
A fixed-length cut doesn't care about meaning. It'll slice a sentence in half. It'll split a question from its answer. Now the one chunk that held your answer is two halves — and each half looks irrelevant on its own. So retrieval never even finds it.
Fix: cut on structure, not length. Split on paragraphs and headers — where the meaning actually breaks. And let chunks overlap a little, so nothing gets stranded at the edge.
2. Embedding — your question doesn't look like its answer
Every chunk goes through a model that turns text into a vector: a long list of numbers that captures its meaning. Similar meaning, similar numbers. That's the whole trick that makes search work.
But here's the trap. Your question gets embedded by the same model — and a question rarely looks like its answer.
You ask "how do I reset my password?" The doc says "account recovery procedure." Same thing to a human. Different words — and a weak embedder puts them far apart.
Fix: pick an embedding model built for retrieval, and test it on your data. A model that's great on legal text can be useless on code.
3. Retrieval — vector search fumbles exact strings
You take the question's vector and grab the closest chunks. This is the "vector search" everyone talks about. On its own, it's got a blind spot.
Vector search matches meaning, not exact words. Which is great — until someone searches for an error code. A product name. SKU-4417. Those barely have a meaning to embed. They're just exact strings, and nothing is "close in meaning" to a serial number.
Fix: hybrid search. Run keyword search and vector search, then merge the results. Old-school keyword matching catches the exact strings embeddings miss.
4. Ranking — the best chunk was #8, and you dropped it
Now you've got a pile of candidates — say the top twenty. But you can't paste twenty chunks into the prompt, so you keep the top five and drop the rest.
Here's the question nobody asks: what if the best chunk was ranked number eight? Then you just threw it away. Your answer was in the pile, and you cut it.
Fix: a re-ranker — a second, smarter model that reads the question and each chunk together and re-scores them. Cheap vector search casts a wide net; the re-ranker picks the real winners out of it. Now your top five are actually the top five.
And order matters too — "lost in the middle"
Even the order you paste chunks in changes the answer. Models pay the most attention to the start and the end of the context. Put your most important chunk in the middle, and the model can skim right past it.
This isn't a hunch — it's measured. The same model gets more accurate just by moving the right chunk to the edges of the context (Liu et al., Lost in the Middle, 2023). So put your best chunk first, or last. Never buried in the middle.
5. The prompt — where people ruin it two ways
Right chunks, right order. This is your last chance to ruin it, and people do it two ways:
- They dump in everything. Ten thousand tokens of maybe-relevant text, hoping the model sorts it out. More context isn't more accuracy — it's more noise to get lost in.
- They never tell the model what to do when the answer isn't there. So it does the one thing you don't want: it makes something up.
Fix is basically one line in your prompt:
"Answer only from the context below. If it's not there, say you don't know."
That single instruction turns a confident liar into a system you can trust. A RAG bot that admits "I don't know" beats ten that guess.
Put the blame where it belongs
- Bad chunks
- Wrong embedder
- Vector-only search
- No re-ranker
- A sloppy prompt
Five failure points — and notice the model was the last one. Usually the innocent one.
So next time your RAG bot lies to you, don't touch the model first. Walk the five steps backward — prompt, ranking, retrieval, embedding, chunks. The bug is almost always upstream.
What's the worst answer your RAG bot has ever given you? Drop it in the comments — I actually want to read those.
I make Vlad's Stack https://www.youtube.com/channel/UCUO8Uo5LsEy1b9eRkrH6JNg — how the AI tools you use every day actually work, for people who write code. Full video walkthrough of this one is above.
Top comments (9)
The five-line framing is a good teaching device because it makes the asymmetry obvious: errors early in the pipeline cannot be recovered later. A reranker can rescue a mediocre retrieval, but nothing downstream can reassemble an answer that chunking cut in half - the two halves are separate points in the index now, and each looks individually irrelevant. That is why tuning order matters as much as the fixes themselves, and most teams tune in exactly the wrong direction, starting at generation because that is where the symptom appeared. A sixth failure worth adding to the list is that the corpus itself can be wrong: two versions of the same document, both indexed, one superseded. Every one of your five steps behaves correctly and the answer is still out of date, because nothing in the text says which copy is authoritative.
Yeah, you nailed exactly why I ordered it that way. The symptom shows up at generation, so that's where everyone starts poking, and it's almost always the one step that's innocent. Walk it backward and the bug is usually three steps upstream.
And you're right that the sixth one isn't on the list, it should be. It's really a step 0: all five steps quietly trust that the corpus is the source of truth. Index two versions of the same doc and retrieval does its job perfectly, it just hands back the stale one, because "superseded" isn't a signal that lives in the text. So the fix isn't a retrieval tweak, it's ingestion hygiene: dedup, version/recency metadata, and tombstone the old copy so it never reaches the index. Great addition, stealing it for the follow-up.
Tombstoning the superseded copy is right for a current-state search view, with one qualification: some users legitimately ask what the policy said at an earlier date. That historical question should not silently be answered from today’s version.
I would distinguish the default active view from an explicitly requested historical view, with version lineage retained in the archive. That keeps the retired copy out of normal retrieval without losing the ability to answer dated questions or explain an earlier answer.
Good catch, a hard tombstone is too blunt for exactly that reason. It optimizes the "what's true now" query and quietly breaks the "what did it say in March" one.
So the fix isn't delete, it's a time dimension: tag each version with a validity range (effective-from / effective-to) instead of dropping the old one. Default retrieval filters to what's current, that's most questions, but a dated query lifts that filter and pulls the right historical version. The retired copy stays in the index, just invisible unless you explicitly ask across time. Same effect you wanted, nothing lost.
The genuinely hard part is upstream of all that: knowing a question is historical in the first place. "What's our refund policy" and "what was our refund policy last spring" need two different retrieval filters, and only one of them says so out loud. Resolving that intent is where I'd expect most systems to quietly get it wrong.
The validity ranges preserve the historical question, and I agree that intent resolution is the next difficult boundary. Some questions name an event rather than a date, such as the policy in effect when an order was placed. That requires resolving the event time before selecting a version.
I would make the chosen time scope visible in the answer and ask for clarification when that reference cannot be resolved. A current-policy default is reasonable for an unqualified question, but it should not silently override an explicit historical reference merely because the date parser failed.
Exactly and the event-time case is sneakier than a plain date, because "when the order was placed" isn't in the documents at all. It's a join: resolve the event to a timestamp from the order record first, then use that to pick the version. Retrieval suddenly depends on a lookup in a different system, and if that hop is wrong the answer is wrong, confidently.
Showing the time scope in the answer is the part I'd push hardest. It's the same move as citing sources, just on the time axis: "here's the version I read, and the date I read it as of." A wrong lens becomes catchable instead of invisible.
And your last point is really #5 from the article wearing a different hat. A date parser that fails and quietly falls back to "current" is just the model making something up again, it's swallowed its own uncertainty instead of saying "I'm not sure which version you mean." Fail loud. A system that asks beats one that guesses, on content, and on time.
The cross-system join is the right framing. I would include the event identifier and the timestamp lookup source in the answer provenance, alongside the document version, so a reviewer can distinguish a bad event lookup from a bad version selection.
A useful boundary case is an order timestamp arriving late or being corrected after ingestion. Replaying the same question should either use a pinned lookup snapshot or explicitly report that the resolved time changed; otherwise a valid historical filter can still produce an unexplained answer change.
Right and splitting those two lines in the provenance is what makes the audit trail actually worth having. “Wrong event timestamp” and “right timestamp, wrong version” are different bugs with different fixes, and a trace that collapses them into one “as of 2024-03” line can’t tell you which one bit you.
The corrected-timestamp case is the one that names the whole thread, though. You’ve got two independent clocks now: when the policy was in effect (the document’s valid time), and when the system learned the order’s timestamp (the lookup’s transaction time). A pure valid-time filter is stable but the moment the lookup itself can be corrected after ingestion, a replayed question changes its answer with no visible cause, because the second clock moved and nothing wrote it down. Pinning the knowledge snapshot, or reporting “the resolved time changed,” is just making that second axis explicit.
Which is the fun part: we started at “two copies of one doc, one stale” and walked straight into bitemporal modeling valid time plus transaction time, with provenance on both. That’s the real sixth failure. Not a bad chunk or a weak embedder a corpus that quietly forgets when it knew what.
The two clocks also suggest making the replay mode explicit in the API: reconstruct what the system knew at the original lookup, or answer using today's corrected knowledge. Those are different questions even when their wording and event identifier are unchanged.
I would retain the resolved event time and lookup snapshot identifier in the original trace, then show a small comparison when a correction changes the selected policy. That makes the change explainable without implying that the original answer used information which had not arrived yet.