DEV Community

Cover image for Retrieval-Augmented Generation, Explained From First Principles
Sarthak Rawat
Sarthak Rawat

Posted on

Retrieval-Augmented Generation, Explained From First Principles

Large language models are remarkably good at producing answers. Ask one to explain a difficult idea, summarize a document, write code, or reason through a problem, and it can often do it in seconds.

But there is a fundamental limitation hiding underneath all of that fluency: the model's knowledge is not the same thing as access to the information you need right now.

A model may know a great deal about a subject and still know nothing about your company's latest policy, yesterday's support tickets, this month's internal report, or the newest version of a regulation. It may also produce a confident answer when the information it needs is missing.

Retrieval-Augmented Generation, or RAG, is one of the most practical ways of dealing with that problem.

The basic idea is simple:

Instead of asking the model to remember everything, give it the
relevant information at the moment it needs to answer.

The simplicity of that sentence is deceptive. A production RAG system is a chain of decisions about documents, representations, search, ranking, context, generation, evaluation, security, latency, and cost. Most of the interesting engineering happens in those decisions.

This article builds RAG from the ground up, then follows the system into the problems that appear when a simple prototype becomes a real application.

The problem RAG is trying to solve

Imagine a brilliant consultant who has read an enormous library of books, papers, manuals, and websites. They can reason extremely well, explain difficult ideas, and synthesize information quickly.

There is one catch: they were locked in a room after a particular date.

They have no idea what happened afterward.

There is another problem. When they do not know something, they may still try to give you an answer. Because they are optimized to produce plausible language, the result can sound convincing even when it is wrong.

That is a useful mental model for an LLM.

Now imagine giving the consultant a research assistant. The assistant can search a filing cabinet, find the relevant pages, and put them on the consultant's desk before the consultant answers.

The consultant still does the reasoning. The assistant provides the evidence.

That is the central idea behind RAG.

Retrieval finds relevant external information.
Augmentation puts that information into the model's context.
Generation produces an answer using that context.

The model's reasoning ability and its external knowledge have been separated.

RAG decouples what a model can reason about from what it currently has access to.

That distinction is more important than memorizing a particular RAG framework or vector database.

RAG, fine-tuning, and long context solve different problems

A common mistake is to treat RAG and fine-tuning as competing ways of doing the same thing.

They are not.

RAG is primarily a way to give a model access to information at inference time. Fine-tuning changes the model's learned behavior. A larger context window gives the model more room to receive information, but does not by itself create a retrieval mechanism.

A useful shorthand is:

Fine-tuning teaches the model how to behave; RAG gives the model what to know.

In practice, the techniques can be combined. A specialized model can learn domain vocabulary, formatting, or communication style through fine-tuning while RAG supplies current private knowledge.

Long context changes the tradeoff again. If the complete information needed for a task is small enough to fit comfortably into context, retrieval may be unnecessary. RAG becomes more useful when the corpus is large, changing, private, or expensive to stuff into every prompt.

There is no universal winner. The choice depends on the information, latency, cost, update frequency, and reliability requirements of the application.

The RAG pipeline

At its simplest, RAG can be understood as four stages:

Index → Retrieve → Augment → Generate

The first stage prepares the knowledge base. The other stages happen when a user asks a question.

1. Indexing: prepare the knowledge

Before anyone asks a question, the system needs to turn its source material into something searchable.

A typical indexing pipeline looks like:

Load → Parse → Chunk → Embed → Store

Load

The source might be almost anything:

  • PDFs
  • Word documents
  • web pages
  • database records
  • Confluence pages
  • Slack messages
  • support tickets
  • internal manuals
  • code or technical documentation

The important point is that RAG does not require the knowledge to have originated as neatly structured text.

Parse and clean

Raw documents are rarely ready for retrieval.

A PDF may contain headers and footers repeated on every page. A webpage may contain navigation menus. A table may be extracted in the wrong order. OCR may introduce errors.

If the parser turns a useful document into poor text, every later stage inherits the damage.

This is one of the least glamorous parts of RAG and one of the most consequential:

Garbage in becomes searchable garbage.

Ingestion also has to deal with practical questions such as
deduplication, deletions, document versions, permissions, and what happens when the embedding model changes. These are not separate from RAG quality; they are part of it.

Chunking

A 200-page document is rarely a useful retrieval unit.

Instead, it is divided into smaller passages called chunks.

The goal is to find a useful balance.

Chunks that are too small can lose the context needed to understand a statement. Chunks that are too large contain too much unrelated information and become less precise retrieval units.

A common starting point is a few hundred tokens with some overlap, but there is no universally correct chunk size. A legal contract, a support ticket, a code file, and a research paper may all want different structures.

Overlap exists for a simple reason: boundaries are artificial.

If a sentence begins at the end of one chunk and finishes at the
beginning of another, retrieving either chunk alone could leave the model with an incomplete idea.

More advanced approaches include parent-child retrieval, where small child chunks are indexed for precise retrieval but a larger parent passage is supplied to the model after a match is found.

The important principle is:

Chunk for retrieval, not merely for storage.

Embeddings

Once chunks exist, the system needs a way to represent their meaning numerically.

An embedding is a vector: a sequence of numbers produced by an embedding model. Text with related meanings should generally occupy nearby regions of the embedding space.

For example, a user might ask:

"Can I work from home?"

while the document says:

"Employees may work remotely up to three days per week."

The wording is different, but the meaning is related. Dense semantic retrieval can recognize that relationship.

The embedding model is not the same thing as the generative LLM. The embedding model creates representations for search. The LLM reads context and generates language.

One important engineering rule follows from this:

The query and the indexed documents need compatible
representations.

If documents were embedded using one model and queries are embedded using an incompatible model, the vectors are not necessarily meaningful to compare.

Dense and sparse representations

Dense embeddings are excellent at semantic similarity, but they are not the only useful representation.

Sparse retrieval methods such as BM25 are strongly tied to terms
appearing in the text. That makes them particularly useful for exact names, product codes, error messages, rare technical identifiers, and other cases where the literal words matter.

Dense and sparse retrieval therefore have complementary strengths.

That is why production systems often combine them.

What happens inside vector search?

Once chunks have embeddings, they need to live somewhere that can search them efficiently.

A vector database or vector index stores the vector alongside the
original text and useful metadata such as:

  • source
  • author
  • department
  • date
  • document version
  • access permissions

When the user asks a question, the query is embedded into the same representation space and compared with stored vectors.

A common similarity measure is cosine similarity, which focuses on the angle between vectors. Dot product is another common option, while L2 distance measures geometric distance directly. Normalization matters because it changes the relationship between these measures.

The important idea is less about memorizing one formula than
understanding what the system is doing:

It is trying to find stored representations that are close to the query under the chosen similarity function.

Searching every vector exactly becomes expensive as the corpus grows.
Approximate nearest-neighbor (ANN) methods trade some exactness for much faster search.

The important word is approximate. Exact nearest-neighbor search asks the system to compare the query against every candidate and identify the true closest vectors. That can become too expensive at large scale.
ANN methods instead organize the search space so that the system can explore a much smaller set of promising candidates.

HNSW: navigating a graph

HNSW (Hierarchical Navigable Small World) is a popular graph-based ANN method.

A useful intuition is a city map with several levels. The upper levels contain a small number of long connections that let you move quickly across the map. Lower levels contain increasingly detailed local connections.

A query starts at a higher level, moves toward promising regions, then descends through the hierarchy until it is navigating the detailed neighborhood containing the closest candidates.

The result is that the system does not need to examine every stored vector.

HNSW also exposes an important engineering tradeoff through parameters such as ef_search. A larger search effort means the algorithm explores more candidate nodes. That can improve recall, but it also increases search work and therefore latency. A smaller value can make search faster while increasing the chance that a genuinely good neighbor is missed.

So even within the same vector database, ANN quality is not simply “on” or “off.” You are choosing where to sit on a speed-versus-recall curve.

IVF: search the relevant regions

Another common approach is inverted file indexing (IVF).

Instead of treating the entire vector space as one enormous search region, IVF partitions vectors into clusters. During indexing, each vector is assigned to an appropriate cluster.

At query time, the system first determines which clusters are closest to the query, then searches only those regions.

The important tuning concept is nprobe: how many clusters should the query inspect?

A low nprobe means less work and lower latency, but it can miss
relevant vectors that live in a cluster the search never examined. A higher nprobe searches more regions and can improve recall at the cost of additional computation.

HNSW and IVF therefore use different structures to make the same broad tradeoff: avoid exhaustive search while keeping the probability of missing useful neighbors acceptably low.

Product quantization: make the vectors smaller

There is another problem at scale: memory.

A large corpus may contain hundreds of millions of vectors, each with hundreds or thousands of dimensions. Keeping every vector in full precision can become expensive.

Product Quantization (PQ) compresses vectors into smaller
representations. Instead of storing every original value directly, the system represents portions of the vector using compact codes.

The benefit is substantially lower memory usage and potentially faster search. The cost is that compression introduces approximation, which can reduce retrieval accuracy.

This is another example of the same RAG engineering principle: you are trading one resource for another.

Why ANN recall becomes an engineering concern

It is tempting to describe ANN as simply “less accurate search.” That is too simplistic.

The real issue is that as the corpus grows, maintaining very high recall under a fixed latency and memory budget becomes harder. The system has more possible neighbors to consider, more index structure to maintain, and more pressure to limit how much work each query performs.

That is why ANN systems expose tuning parameters and why retrieval should be evaluated empirically.

In practice, you care about questions such as:

  • How much recall@k do we lose at our latency target?
  • What happens when the corpus grows tenfold?
  • How much memory does the index require?
  • Which parameters improve recall without making latency unacceptable?
  • Does compression change the ranking enough to affect downstream answer quality?

The useful mental model is therefore not “ANN is an approximation.”

It is:

ANN is a controlled tradeoff between search cost and the probability of finding the best neighbors.

Other ANN approaches include inverted-file methods and vector
compression techniques such as product quantization. The specific choice depends on corpus size, latency requirements, available memory, and the recall you need.

Metadata filtering is another important capability. You may want
semantic similarity only among documents from the legal department, or only among documents newer than a certain date, or only among documents the current user is allowed to see.

That last requirement becomes critical in multi-tenant and
permission-sensitive systems. A retrieval system that finds the right answer but exposes information the user is not authorized to see is still a broken system.

Retrieval: why "just use vector search" is not enough

A basic RAG prototype might embed the query, perform a similarity
search, take the top five chunks, and send them to the LLM.

Sometimes that works beautifully.

Sometimes it fails in ways that are difficult to diagnose.

There are several recurring reasons.

Vocabulary mismatch

A user may say "heart attack" while the document says "myocardial
infarction."

Semantic retrieval can help bridge that gap, but no retrieval method is perfect.

Exact terms matter too

Now imagine the query contains a product ID such as XR-4729, a legal clause number, or a specific error code.

Semantic similarity is not necessarily the best tool for that.

This is where sparse keyword retrieval can be extremely valuable.

Query and document lengths differ

A user query might be six words while the retrieved chunk is several hundred words. Comparing the two is not exactly the same problem as comparing two similarly sized documents.

Redundancy

The top ten results may all say essentially the same thing.

Sending ten versions of one idea to the LLM does not necessarily make the answer better.

Hybrid retrieval: combine complementary searches

A common production baseline is hybrid search.

Run a dense semantic search and a sparse lexical search in parallel, then combine their rankings.

BM25 is a classic sparse retrieval method. It rewards useful term
matches while accounting for how common a term is across the collection.

Dense retrieval contributes semantic matching.

BM25 contributes exact lexical matching.

The results can be combined using a ranking method such as Reciprocal Rank Fusion (RRF), where a document receives contributions based on where it appeared in each ranked list.

The important insight is not the formula itself. It is that two
imperfect retrieval mechanisms can complement one another.

Query rewriting and decomposition

The user's raw question is not always the best search query.

A conversational question such as:

"What about the policy for contractors?"

may depend on several earlier turns.

A query rewriting step can turn conversational language into a
standalone retrieval query.

More complex questions can also be decomposed.

Suppose someone asks:

"How did our Q3 revenue compare with the previous quarter, and what were the main reasons for the change?"

That may require separate retrieval for the numbers and for the
explanation.

Instead of treating the whole thing as one search query, the system can create subqueries, retrieve evidence for each, and combine the results.

This adds complexity and sometimes extra model calls, so it should be used when the query actually benefits from it.

HyDE and multi-query retrieval

HyDE, or Hypothetical Document Embeddings, takes another approach.

Instead of embedding the user's short question directly, the system first asks an LLM to generate a hypothetical answer or document. That hypothetical text is then embedded and used for retrieval.

The hypothetical text does not have to be factually correct. Its purpose is to provide a richer representation of what a relevant document might look like.

The tradeoff is straightforward: another model call means additional latency and cost.

Multi-query retrieval is simpler conceptually. The system generates several alternative formulations of the same question, retrieves for each, merges and deduplicates the results, and then ranks the candidates.

These techniques are useful when the original query is a poor
representation of the information the system actually needs.

They are not mandatory ingredients of every RAG system.

Reranking: retrieve broadly, then judge carefully

One of the most useful patterns in practical RAG is to separate
candidate generation from precise ranking.

The first retrieval stage should be fast. It might return 20, 30, or 50 candidates.

Then a more expensive cross-encoder reranker examines the query and each candidate together and scores how well they match.

This works because the two stages have different jobs.

A bi-encoder or vector search system is efficient because the document representations can be computed ahead of time.

A cross-encoder can make a more detailed query-document comparison, but doing that against millions of chunks would be too expensive.

So the architecture becomes:

large corpus → fast retrieval → small candidate set → expensive
reranking → tiny final context

This is one of the most important patterns to understand in RAG
engineering.

Diversity matters too

Sometimes the five highest-ranked chunks all repeat the same
information.

Maximal Marginal Relevance (MMR) provides a way to balance relevance with diversity.

Conceptually, MMR rewards a chunk for being relevant to the query while penalizing it for being too similar to chunks already selected.

The result is a final context containing several complementary pieces of evidence instead of five near-duplicates.

The broader lesson is important:

Retrieval is not about finding the most similar text. It is about constructing the most useful evidence set.

The generation stage: context is evidence, not decoration

Once the system has retrieved its evidence, it needs to assemble the prompt that the LLM will see.

A useful RAG prompt typically contains:

  1. clear system instructions
  2. labelled source chunks
  3. grounding rules
  4. the user's question
  5. a defined output format when necessary

The structure matters because the model is now being asked to reason over evidence.

A good instruction might say:

Answer using only the provided context. If the answer is not supported by the context, say that the information is insufficient. Do not speculate beyond the sources.

That does not magically make hallucination impossible. But it gives the model a clear operating rule.

More context is not automatically better

It is tempting to think:

"If five chunks are useful, twenty must be better."

Usually, it is not that simple.

Every additional chunk increases token usage and latency. Irrelevant chunks can introduce noise. Conflicting chunks can make the model uncertain. And long contexts can create attention problems.

This is related to the lost-in-the-middle phenomenon: when relevant information is buried among a long sequence of context, models may use information at the beginning and end more effectively than information buried in the middle.

A practical strategy is therefore:

Retrieve more, send less.

Retrieve a broad candidate set. Rerank it. Select a small, high-quality context.

Citations

One of RAG's practical advantages is that the system knows which sources it retrieved.

Citations can be rendered inline:

Employees may work remotely up to three days per week [Source 1].

Or as a source list:

Sources: HR Policy, page 4; Manager Handbook, page 12.

In production, structured citation data can be even more useful. The system can return an answer together with source IDs, pages, claims, and links so the interface can make evidence clickable and auditable.

What happens when the answer isn't in the documents?

A trustworthy RAG system needs an explicit answer to this question.

If the retriever finds nothing relevant, the model should not feel obligated to invent an answer.

A simple design is to combine grounding instructions with retrieval thresholds. If none of the retrieved candidates passes an appropriate relevance threshold, the system can decline to answer from the knowledge base.

The exact threshold is application-specific. A score from one embedding system cannot automatically be treated as a universal measure of relevance.

The deeper principle is more important:

A good RAG system must have a graceful failure mode.

"I don't have enough information in the retrieved sources" is often much better than a confident fabrication.

Conversational RAG: the query is not always the query

Chat makes retrieval harder.

A user might say:

"What's the refund policy?"

Then:

"What about digital products?"

Then:

"And if I used a gift card?"

The final message is not a complete retrieval query.

A common solution is query condensation: use the conversation
history to rewrite the latest message into a standalone question before retrieval.

For example:

"What is the refund policy for digital products purchased with a gift
card?"

Now the retrieval system has the information it needs to search
effectively.

This adds another processing step, but it solves a fundamental mismatch between how humans converse and how retrieval systems search.

When RAG gives a bad answer: debug the pipeline, not just the model

One of the most useful ways to think about RAG is that a bad answer does not necessarily mean the LLM is the problem.

A failure can happen at several layers:

Ingestion failure\
The relevant information was parsed incorrectly, omitted, duplicated, or indexed incorrectly.

Chunking failure\
The information exists, but it was divided into retrieval units that destroyed the useful context.

Retrieval failure\
The right information exists in the index but was never retrieved.

Ranking failure\
The right information was retrieved but buried below less useful
candidates.

Context assembly failure\
The correct evidence was found but was poorly ordered, excessively long, or mixed with conflicting material.

Generation failure\
The model received useful evidence but failed to follow it or answer the question correctly.

This gives you a practical debugging order:

Check the source → check the retrieved evidence → check the
ranking → check the assembled context → check the generation.

Do not immediately swap the LLM when the real problem is retrieval.

How do you know whether a RAG system actually works?

A RAG system can produce answers that sound excellent while failing underneath.

Evaluation needs to separate the pipeline into understandable
relationships.

The RAG triad

Three questions form a useful mental model:

1. Context relevance

Did the retriever find information that was actually related to the question?

2. Faithfulness

Are the claims in the generated answer supported by the retrieved
context?

3. Answer relevance

Does the answer actually address what the user asked?

These dimensions catch different failures.

A system can retrieve excellent evidence and still hallucinate.

It can produce a faithful answer that does not answer the user's
question.

It can generate a relevant answer because the model already knew the topic, even though retrieval failed.

Retrieval metrics

The RAG triad is useful for end-to-end reasoning, but retrieval also benefits from traditional information-retrieval metrics.

Recall@k asks whether the relevant evidence appears somewhere in the top k results.

MRR (Mean Reciprocal Rank) cares about how early the first relevant result appears.

nDCG evaluates ranked results while giving more credit to highly relevant items near the top.

These metrics answer a different question from faithfulness. They tell you whether the retrieval system is finding and ordering useful
evidence.

RAGAS and LLM-as-judge

Frameworks such as RAGAS operationalize several RAG evaluation concepts, often using an LLM as a judge.

One useful technique is to break an answer into atomic claims and verify those claims against the retrieved context.

For example:

Claim 1 → supported\
Claim 2 → supported\
Claim 3 → unsupported

That gives a more meaningful faithfulness assessment than asking whether an answer "looks good."

LLM-as-judge is powerful because it scales, but it is not an
unquestionable authority. Judges can have biases, prefer verbose
answers, or make mistakes of their own.

A strong evaluation system therefore combines automated evaluation with carefully constructed golden datasets and periodic human validation.

A practical test set can combine:

  • expert-written questions
  • synthetic questions generated from the corpus
  • sampled real production queries

The important part is establishing a baseline and using it to detect regressions when the RAG pipeline changes.

When basic RAG isn't enough

Basic RAG makes a strong assumption:

Retrieve something, use it, and generate an answer.

That assumption breaks in several ways.

Sometimes the question does not require retrieval.
Sometimes the retrieved information is poor.
Sometimes answering requires relationships across many documents.

Sometimes the system needs to decide what tool or retrieval strategy to use next.

Several advanced RAG patterns address these problems.

Self-RAG: should I retrieve, and did retrieval help?

Basic RAG assumes retrieval should happen for every query.

That is wasteful for questions that the model can answer directly, and it can actively hurt when retrieval returns poor evidence.

Self-RAG changes that assumption by giving the model a role in deciding when retrieval is needed and in evaluating whether retrieved information is useful.

The original approach uses special reflection signals generated by a model that has been trained for this behavior. In other words, the reflection mechanism is part of the model's learned generation process; it is not simply a separate classifier bolted onto an ordinary LLM.

Conceptually, the loop looks like:

question → decide whether retrieval is useful → retrieve → assess the evidence → generate

This makes Self-RAG attractive for systems that receive a mixture of questions: some needing external knowledge and others that do not.

The tradeoff is important, though. The approach requires a model capable of the specialized reflection behavior, which makes it more involved than simply adding another retrieval step to an existing RAG pipeline.

Corrective RAG: don't trust the retriever blindly

A standard RAG pipeline can make a dangerous assumption:

If the retriever returned something, it must be useful.

Corrective RAG (CRAG) explicitly challenges that assumption.

CRAG introduces a quality assessment around retrieved documents. If the retrieved material is judged weak, the system can take corrective action instead of blindly passing the results to the generator.

One part of the original CRAG idea is knowledge refinement. Retrieved chunks can be broken into smaller units, irrelevant material can be discarded, and the remaining high-signal information can be recombined before generation.

The conceptual shift is simple:

Basic RAG: retrieve → trust → generate

CRAG: retrieve → evaluate → refine/correct → generate

This is particularly useful when document quality is uneven or the retrieval system cannot be assumed to return reliable evidence every time.

GraphRAG: when relationships matter

Vector search answers a question like:

“Which passages are most similar to this query?”

That is powerful, but some questions are fundamentally about
relationships.

Imagine a question that requires connecting an employee to a department, that department to a project, and that project to a supplier. The useful information may be spread across many documents and may not appear as one semantically similar passage.

GraphRAG addresses this by representing entities and relationships explicitly.

A typical graph-oriented pipeline can involve:

  1. extracting entities from source documents
  2. identifying relationships between those entities
  3. constructing a knowledge graph
  4. detecting communities or groups of closely connected entities
  5. generating summaries of those communities
  6. retrieving through graph structure at query time

This creates two useful modes of reasoning.

Local search focuses on a particular entity and its surrounding neighborhood.

Global search can reason over broader themes represented by
community-level summaries.

GraphRAG therefore becomes attractive when the question depends on connections rather than simply semantic similarity: organizational structures, citation networks, supply chains, interconnected technical systems, and similar domains.

The cost is substantial. Building and maintaining the graph can require many extraction steps and additional LLM calls, and graph-based infrastructure is more complex than a straightforward vector index.

The important lesson is not that graphs replace vectors.

It is that the retrieval structure should match the structure of the questions you need to answer.

Agentic RAG: what should the system do next?

A conventional RAG pipeline follows a predefined sequence.

An agentic RAG system gives the model more control over that sequence.

Instead of always performing exactly one retrieval call, the system can plan and decide among tools or actions. Depending on the task, those tools might include:

  • vector search
  • SQL
  • web search
  • APIs
  • code execution
  • graph traversal

A complex question might therefore become a sequence of actions:

understand the question → retrieve internal data → identify a missing piece → search another source → calculate a result → retrieve supporting evidence → synthesize the answer

This is useful when the problem is inherently multi-step or the required information lives in heterogeneous systems.

But agency introduces costs.

More tool calls mean more latency. More branches make behavior harder to predict. More moving parts make debugging harder. And an agent can make the wrong decision about what to do next.

The goal is therefore not to make every RAG system agentic.

It is to use agency when the problem actually requires adaptive,
multi-step behavior.

Four different responses to four different problems

The four patterns are easier to remember when you connect each one to the assumption it changes:

  • Self-RAG: Should I retrieve?
  • Corrective RAG: Is what I retrieved useful enough to trust?
  • GraphRAG: Do I need explicit relationships between pieces of knowledge?
  • Agentic RAG: What should I do next to solve this problem?

They can also be combined. A system might use an agent to decide whether to retrieve, use hybrid search to find candidates, apply corrective evaluation to those candidates, and use a graph when the question turns out to be relational.

The point is not to collect advanced-RAG names.

The point is to understand which limitation each architecture is trying to remove.

RAG in production is a different problem

A prototype can work beautifully on a laptop with a few documents.

Production introduces a different set of constraints.

Freshness

A knowledge base changes.

Documents are edited. Policies are replaced. New tickets arrive. Old information becomes invalid.

A production system therefore needs a strategy for freshness:

  • periodic re-indexing
  • incremental indexing triggered by document changes
  • timestamps and document versions
  • temporal ranking when recency matters
  • live retrieval for especially fast-changing sources

The system should know not just what a chunk says, but often when that information was valid.

Latency

A RAG request may involve query rewriting, retrieval, reranking, prompt assembly, and generation.

Every extra step adds latency.

Useful techniques include:

  • streaming the generated answer
  • caching repeated or semantically similar requests
  • running independent retrieval operations in parallel
  • using smaller models for lightweight tasks such as query rewriting
  • reducing the amount of context sent to the generator

Streaming is particularly important for perceived latency: users can begin reading while the model is still generating.

Cost

Every token, model call, and embedding operation has a cost.

The major levers are:

  • context size
  • model choice
  • number of LLM calls
  • cache hit rate
  • embedding infrastructure

A common mistake is using the most expensive model for every stage. A smaller model may be perfectly adequate for query rewriting or classification while a stronger model is reserved for the final answer.

The same principle applies to retrieval. Spending more computation on every candidate is wasteful if a fast first-stage search can narrow millions of chunks to a few dozen.

Security: the documents themselves can attack the system

Prompt injection in RAG has an unusual property.

The attacker does not necessarily need to control the user's question.

They may control a document that gets retrieved.

Imagine someone uploads a malicious support ticket containing
instructions aimed at the model. If that ticket is indexed and later retrieved, its contents are inserted into the context automatically.

The user asking the next question may be completely innocent.

This makes document trust a security boundary.

Defenses should be layered:

  • index content only from appropriate sources
  • treat user-uploaded documents as untrusted
  • isolate untrusted content where appropriate
  • preserve a clear instruction hierarchy
  • monitor outputs for anomalous behavior
  • enforce authorization before retrieval, not only after generation

Permission-aware retrieval is particularly important in multi-tenant systems.

A cache must also respect tenant and user boundaries. A perfectly cached answer is a security failure if it is returned to someone who was never authorized to see the underlying information.

Multimodal RAG

Real knowledge bases are not made entirely of paragraphs.

PDFs contain tables and charts. Wikis contain images. Technical
documentation contains diagrams.

There are several ways to handle this.

One approach is to extract or transcribe visual content into text. It is simple, but potentially lossy.

Another is to use multimodal embeddings so visual and textual content can be represented in a compatible space.

A third approach is to use a vision-language model during ingestion to generate rich descriptions of tables and images, then index those descriptions using standard text retrieval.

The practical choice depends on the application. The important point is that retrieval quality depends on preserving the information that matters, not merely converting everything into plain text.

Observability: if you can't see the pipeline, you can't debug it

A production RAG system should make its decisions inspectable.

For each query, useful telemetry can include:

  • the query after privacy/PII handling
  • retrieved chunk IDs
  • retrieval scores
  • reranking scores
  • prompt or prompt hash
  • generated answer
  • latency by stage
  • token counts
  • evaluation signals
  • user feedback

Tools such as LangSmith, TruLens, Weights & Biases, Arize AI, or a custom data warehouse can provide this visibility.

The goal is not to collect logs for their own sake.

It is to make questions like this answerable:

"Show me the queries this week where faithfulness dropped below
our baseline and tell me which retrieved chunks were involved."

That turns RAG debugging from guesswork into investigation.

RAG is not always the right answer

Understanding when not to use RAG is just as important as knowing how to build it.

If the complete knowledge base is ten pages and comfortably fits in context, retrieval may add unnecessary complexity.

If the task is purely creative, retrieval may add noise rather than useful evidence.

If the primary source is structured data, a direct database query may be better. Asking "What was Q3 revenue in APAC?" is fundamentally different from asking a question about an unstructured policy document.

And if retrieval quality is consistently poor, adding more retrieval machinery does not automatically solve the problem. Sometimes the right answer is to improve the source data or rethink the architecture.

The correct question is not:

"Should I use RAG?"

It is:

"Where does the information needed for this task live, how often does it change, and what is the best way to give the model reliable access to it?"

Designing a RAG system from requirements

When designing a RAG system, it is tempting to immediately draw a vector database and an LLM.

Start somewhere else.

Start with the requirements.

1. Understand the use case

Ask:

  • What kinds of documents are involved?
  • How large is the corpus?
  • How frequently does it change?
  • Are queries simple or multi-step?
  • Is the application single-turn or conversational?
  • What is the latency budget?
  • What happens if the answer is wrong?

The last question changes the architecture dramatically.

A wrong answer in a casual internal search tool is different from a wrong answer in a legal or medical workflow.

2. Design ingestion

Choose parsing and chunking based on the actual document types.

Think about deduplication, versions, deletions, metadata, permissions, and re-indexing.

3. Design retrieval

Start with a simple baseline.

Dense retrieval may be enough for some corpora.

Hybrid retrieval is often a strong general-purpose starting point when exact terms and semantic matching both matter.

Add reranking when precision matters.

Add query rewriting or decomposition when the query itself is the
bottleneck.

Do not add every technique simply because it exists.

4. Design generation

Choose the model based on the quality, latency, and cost requirements.

Define grounding behavior.

Decide how citations will work.

Define what happens when evidence is missing.

5. Design evaluation and operations

Decide how success will be measured before declaring the system
finished.

Then design for:

  • freshness
  • security
  • observability
  • caching
  • latency
  • cost
  • regression testing

The architecture should follow the requirements, not the other way around.

A practical example: designing RAG for a legal firm

Imagine a legal firm wants an internal assistant that answers questions about contracts, case materials, policies, and research documents.

The design immediately raises constraints.

The documents may be long and structurally complex. Access permissions matter. Citations matter. Document versions matter. An incorrect answer can have serious consequences.

That suggests several architectural choices.

Documents need careful parsing and metadata.

Chunking should preserve sections and surrounding legal context rather than blindly splitting every fixed number of tokens.

Retrieval may benefit from hybrid search because legal questions can contain both semantic concepts and exact clause numbers, names, dates, and terminology.

Reranking becomes valuable because returning the wrong clause can be worse than returning no clause.

The generation layer should be strongly grounded and citation-oriented.

Permission checks must happen before evidence reaches the model.

Evaluation should include a carefully curated golden set, not only generic synthetic questions.

And freshness/versioning cannot be an afterthought. A superseded
contract clause should not silently compete with the current version.

The interesting part of the design is not the component list.

It is the reasoning:

Each architectural decision exists because of a property of
the problem.

That is the transferable skill.

What changes when the system grows?

Scale exposes different bottlenecks.

At millions of queries, vector search throughput may become important.
At hundreds of millions of chunks, index size and memory become serious concerns.

ANN parameters can be tuned for the desired speed/recall balance. Vector compression can reduce memory usage at some accuracy cost. Large embedding pipelines need batching and parallel processing.

And cost changes character at scale.

At low traffic, an unnecessary model call is annoying.

At very high traffic, the same unnecessary model call becomes a major line item.

That is why caching, smaller models for simple tasks, efficient
retrieval, and context reduction are not merely optimization tricks.
They can determine whether the architecture is economically viable.

The deeper lesson: RAG is a system of tradeoffs

It is easy to collect RAG techniques as a list:

Chunking.

Embeddings.

Vector databases.

Hybrid search.

Reranking.

HyDE.

MMR.

GraphRAG.

Agents.

RAGAS.

And so on.

But knowing the names is not the same as understanding the system.

The more useful mental model is that every decision changes a tradeoff.

Smaller chunks can improve retrieval precision but may lose context.

Larger chunks preserve context but may dilute the relevant signal.

More retrieved documents can improve recall but increase noise and
cost.

Reranking can improve precision but adds computation.

Query rewriting can improve search quality but adds latency.

Stronger grounding can reduce unsupported claims but may increase "I don't know" responses.

Larger models can improve generation quality but increase cost and latency.

More context can provide more evidence but can also produce
attention degradation.

More sophisticated agents can solve more complex problems but are harder to predict and debug.

There is no magical configuration that wins everywhere.

The engineering question is always:

What problem are we experiencing, and which change addresses
that problem without creating a worse one somewhere else?

That is the difference between assembling a RAG demo and designing a RAG system.

RAG in one picture

At the highest level, the whole system can be reduced to a simple idea.

Documents become searchable representations.

A user's question becomes a search request.

Retrieval finds candidate evidence.

Ranking decides which evidence matters most.

The context is assembled carefully.

The LLM reasons over that evidence.

Evaluation checks whether the pipeline actually worked.

Production infrastructure keeps it fresh, fast, secure, observable, and affordable.

Final takeaway

RAG began with a simple observation:

A model does not have to contain all the knowledge it needs to answer a question. It needs a reliable way to access the right knowledge when the question arrives.

From that idea follows the entire architecture.

You need to ingest and parse the source material.

You need to chunk it without destroying meaning.

You need representations that make useful information searchable.

You need retrieval that handles both semantic similarity and exact terms.

You need ranking that separates good candidates from merely plausible ones.

You need context assembly that gives the LLM enough evidence without drowning it.

You need generation that stays grounded.

You need evaluation that tells you whether the failure happened in retrieval or generation.

And once the system becomes real, you need freshness, permissions, security, latency controls, caching, observability, and cost discipline.

The most important lesson is therefore not a particular vector database, embedding model, reranker, or RAG framework.

It is a way of thinking:

When a RAG system fails, ask where the information was lost.

Was it lost during ingestion?

During chunking?

During retrieval?

During ranking?

During context assembly?

Or during generation?

Once you can answer that question, the enormous collection of techniques around RAG becomes much easier to understand.

RAG stops looking like a bag of AI buzzwords and starts looking like what it really is:

an information pipeline designed to connect a reasoning model with the knowledge it needs.

Top comments (0)