DEV Community

Cover image for What Is RAG — And Why Every AI App That Touches Real Data Needs It
Sham Prakash K
Sham Prakash K

Posted on AI-assisted

What Is RAG — And Why Every AI App That Touches Real Data Needs It

The model doesn't know your data. That's the sentence most AI tutorials skip.

You ask it about your product catalog — it hallucinates one. You ask it about your internal policy — it describes something it saw during training. You ask it what happened last week — it has no idea. Not because the model is bad. Because your data was never in the training set.

This is the problem RAG solves. And when the answer genuinely isn't in your data, RAG makes the model say so — instead of making something up. That's the other half of what it fixes.

Why the model doesn't know your data

An LLM learns by reading an enormous amount of public text — web pages, books, code, articles. After training, those patterns are baked into the model's weights. That's how it "knows" things.

But training has two hard limits.

A cutoff date. Anything published after the training cutoff doesn't exist for the model. It doesn't know what happened last month. It doesn't know what your team shipped last week.

Only what was in the training data. Your company's internal docs were never there. Your product catalog, your customer records, your policies — none of it was public text the model could learn from.

So when you ask about your specific domain, the model doesn't retrieve from a database. It generates text that fits the pattern based on what it learned during training. For general questions, that's fine. For questions about your data, it guesses — and guesses confidently.


The naive fix and why it fails

The first thought most people have: just include everything in the prompt.

Before the model replies, paste your entire document library into the context. All the product docs, all the policies, all the records. Give it everything — it'll find what it needs.

This fails for three reasons.

Token limits. A 50-page document is roughly 35,000 tokens. Even with a million-token context window, you can't dump an entire document library into every request. Real systems have thousands of documents.

The "lost in the middle" problem. We covered this in article 03. When you send a very long prompt, the model pays less attention to content buried in the middle. Relevant information sitting in paragraph 47 gets deprioritised. More context is not always better — it can actively hurt response quality.

Cost. If your context is 50,000 tokens and you're paying per million tokens, that's real money per API call — multiplied across every request from every user.

Dumping everything in the prompt doesn't scale. You need something smarter.


"Why not just fine-tune the model on my data?"

This is the first question most engineers ask. It sounds cleaner — train the model once on your data, and it knows everything.

Fine-tuning doesn't solve this problem. Here's why.

Fine-tuning teaches a model a style or a behaviour, not facts. You can fine-tune a model to always respond in JSON, or to follow a specific tone, or to focus on a domain. But it doesn't reliably store factual knowledge in a way you can trust. Fine-tuned models still hallucinate — they just hallucinate in your style.

More practically: your data changes. Products get updated, policies change, new documents are added. Every time your data changes, you'd need to re-run fine-tuning — a process that takes hours and costs significant compute. RAG retrieves from your current data on every query. Add a document today, it's searchable immediately.

Fine-tuning and RAG solve different problems. Fine-tuning for behaviour and style. RAG for grounding the model in specific, up-to-date facts. They're not alternatives — production systems often use both.


What RAG is

RAG stands for Retrieval Augmented Generation.

The word order matters. Retrieval comes first. Generation comes second.

Instead of sending everything to the model and hoping it finds the relevant piece — you find the relevant piece first, then send only that.

Here's the idea with a concrete example.

You have 500 pages of product documentation. A user asks: "What's the return policy for electronics?"

Without RAG: send all 500 pages to the model, hope it finds the answer somewhere.

With RAG:

  1. Search your 500 pages for the sections most relevant to "return policy for electronics"
  2. Find 3–4 paragraphs that actually contain that information
  3. Send only those paragraphs to the model, along with the question
  4. The model reads what you gave it and answers from that

The model never searches anything. It just reads what you put in front of it. You are the one doing the retrieval — and then giving the model exactly what it needs.

That's RAG. Retrieve the relevant pieces. Generate the answer from them.


Two phases: ingest and query

Every RAG system has two distinct phases. Understanding this split is the key to understanding how it works.


Phase 1 — Ingest (once per document)

Before any user can ask questions, you prepare your data.

You take your documents — PDFs, text files, markdown, whatever — and process them:

  1. Split each document into small chunks (a few hundred words each)
  2. Embed each chunk — convert it into a list of numbers called a vector that captures its meaning
  3. Store those vectors in a database built for vector search

This is a one-time operation per document. Once ingested, a document is ready to be retrieved.


Phase 2 — Query (every user request)

When a user asks a question:

  1. Embed the question — convert it into a vector using the same embedding model used during ingest
  2. Search the vector database for stored chunks that are semantically similar to the question vector
  3. Retrieve the top 3–5 most relevant chunks (more on why this range in a moment)
  4. Inject those chunks into the prompt alongside the question
  5. Generate — the model reads the context and answers

Step by step, every time.

Why 3–5 chunks? Retrieve just 1 and you might miss the answer — the single most similar chunk may not contain everything the question needs, especially if the answer spans related sections. Retrieve 20 and you've recreated the "lost in the middle" problem — relevant details buried in a wall of context. 3–5 is the sweet spot: enough coverage to find the answer, focused enough that the model can use it. You'll tune this number based on your chunk size and your data.

One constraint matters here: the embedding model used at query time must be the same one used during ingest. Every embedding model has its own internal "map" of meaning. A vector produced by one model lives in a completely different space than a vector from another model — similarity scores between them are meaningless. If you embed your documents with Gemini's embedding model and then embed questions with OpenAI's, your similarity search will return garbage. Same model, both ways, always.


How the prompt changes

Without RAG, a prompt looks like this:

System: You are a helpful assistant.
User: What's the return policy for electronics?
Enter fullscreen mode Exit fullscreen mode

The model guesses from its training data.

With RAG, the prompt looks like this:

System: You are a helpful assistant. Answer ONLY based on the provided context.
        If the context lacks enough information, say so clearly.

Context:
---
[Chunk 1: "Electronics purchased at full price may be returned within 30 days
with original packaging and receipt. Items must be in original condition..."]

[Chunk 2: "Exceptions to the standard return policy: laptops and tablets
cannot be returned after the seal is broken unless defective..."]

[Chunk 3: "To initiate a return, contact customer support with your order number.
Refunds are processed within 5–7 business days..."]
---

User: What's the return policy for electronics?
Enter fullscreen mode Exit fullscreen mode

Now the model isn't guessing. It's reading actual policy text and summarising it for the user.

Notice the system prompt instruction: Answer ONLY based on the provided context. This is critical. Without it, the model might supplement your retrieved context with its training knowledge — mixing real data with guesses. That instruction keeps it grounded in what you provided.

And when the answer isn't in the retrieved chunks? The model says: "I don't have enough information to answer this." That's not a failure — that's the system working correctly. A confident wrong answer is far more dangerous than an honest "I don't know." RAG gives you the second option.


RAG is not a search engine

A keyword search finds documents that contain your words. RAG finds documents that are semantically related to your question — even when they share no keywords.

Search for "fix broken leg" — a keyword search returns results with those exact words. A RAG embedding model understands that "broken leg," "bone fracture," and "orthopedic injury" are close in meaning — and ranks results by semantic similarity, not by word overlap.

This is what makes retrieval in RAG powerful. It understands meaning. How it does that — through embeddings — is what the next article covers in detail.


Two types of questions RAG answers

Before building, be clear which problem you're solving. RAG works for two different use cases that look similar from the outside.

Document questions — "what does the policy say about X?"
You're searching for information inside documents. Product specs, how-to guides, FAQs, policies, uploaded notes. Semantic retrieval finds the relevant section.

Record questions — "what happened last week?"
You're looking up structured data. Order history, user activity, database records. Semantic search isn't the right tool here — structured queries are. This is where a database agent (covered later in the series) takes over.

RAG handles documents. A SQL agent handles records. A complete AI backend needs both — but they're separate problems.


The architecture after adding RAG

Before RAG, the architecture is simple:

User → API → LLM
         ↓
   PostgreSQL (chat history)
Enter fullscreen mode Exit fullscreen mode

After adding RAG, two new components appear:

User → API → LLM
         ↓        ↑
   PostgreSQL   Context (retrieved chunks)
                    ↑
              Vector DB ← Embedding Model
                    ↑
              Ingest pipeline (split → embed → store)
Enter fullscreen mode Exit fullscreen mode

Embedding model — converts text to vectors. For Gemini users, Google provides an embedding API. Same provider, no extra dependency.

Vector database — stores and searches vectors by similarity. Pinecone is the most common choice. Free tier covers everything you need to learn.

Ingest pipeline — when a document is uploaded, it runs through: split → embed → store. One operation, runs once per document.

On every user query, the API: embeds the question → searches the vector DB → retrieves the top chunks → injects them into the prompt → calls the LLM.


A mistake worth knowing before you start

When I first built the ingest pipeline, I chunked documents into 2,000-word pieces — roughly half a page of text.

Retrieval looked fine. Top chunks came back, all semantically relevant. But responses were vague. The model had the right section but couldn't extract specific answers cleanly.

The problem: each chunk was too large. A 2,000-word chunk gets embedded as one vector representing the average meaning of the whole section — not any specific fact inside it. When retrieved, you get a chunk that contains the answer somewhere inside 400 other words. The model has to work harder and often glosses over the detail.

Dropping to 300-word chunks fixed it. Retrieval precision jumped. Specific questions started getting specific answers.

Chunk size is one of the biggest practical levers in RAG quality. It's easy to get wrong the first time — and it fails silently. We'll go deep on exactly why in upcoming articles.


What's next

RAG needs two things: a way to measure semantic similarity so you can find relevant chunks, and a database that can search by that similarity. Both depend on embeddings.

What is an embedding? Why does converting text into a list of numbers let you search by meaning? Why does "cricket" end up close to "sports" in vector space even though they share no characters?

That's what the next article is about.


Building something where the model needs to answer from your own data? Drop in the comments what the use case is — I'm curious what people are solving.

Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.

Top comments (0)