DEV Community

LogicWiz Admin
LogicWiz Admin

Posted on Originally published at logicwiz.ai

RAG Tutorial for Beginners: Build a Retrieval Pipeline in Python

RAG stands for retrieval augmented generation. Strip the jargon and it is one idea: the
model does not know your documents, so before you ask it a question, you look up the
relevant passage yourself and paste it into the prompt. Retrieve, augment, generate.

That is the whole trick. Everything else in a RAG pipeline (embeddings, chunking, vector
databases, rerankers) exists to make the "look up the relevant passage" step work on
thousands of pages instead of one.

This tutorial builds the pipeline from scratch in Python. No LangChain, no vector
database, one file. By the end you will know what those tools are doing when you do
reach for them.

Why the model needs help in the first place

A language model answers from what it saw during training. Ask it about your company's
refund policy, last week's meeting notes, or a PDF you were sent this morning, and it
has two options: admit it does not know, or produce something confident and wrong. Most
models pick the second.

You could paste the whole document into every prompt. That works until the document is
longer than the context window, or you have five hundred documents, or you are paying
per token and the document is 40 pages. RAG is the fix: send only the few paragraphs
that matter for this question.

The pipeline in one picture

your documents
   │
   ▼
1. chunk      split each document into passages of a few hundred words
   │
   ▼
2. embed      turn each passage into a list of numbers that captures its meaning
   │
   ▼
3. index      store the passages and their numbers so you can search them
   │
   ▼            user question
4. retrieve   embed the question, find the passages whose numbers are closest
   │
   ▼
5. augment    put those passages into the prompt, above the question
   │
   ▼
6. generate   the model answers using the passages it was just shown
Enter fullscreen mode Exit fullscreen mode

Steps 1 to 3 happen once, when you load the documents. Steps 4 to 6 happen on every
question.

Step 1: chunk the documents

A passage has to be small enough that a handful of them fit in the prompt, and large
enough to still make sense on its own. A few hundred words is the usual compromise.

def chunk(text: str, size: int = 300, overlap: int = 50) -> list[str]:
    words = text.split()
    chunks = []
    start = 0
    while start < len(words):
        chunks.append(" ".join(words[start:start + size]))
        start += size - overlap
    return chunks
Enter fullscreen mode Exit fullscreen mode

The overlap matters. Without it, a sentence that straddles a boundary is cut in half and
neither chunk contains the full thought.

Step 2: embed each chunk

An embedding is a list of a few hundred or a few thousand numbers. Passages that mean
similar things get lists that point in similar directions. That is what lets you search
by meaning rather than by exact words.

from openai import OpenAI

client = OpenAI()  # reads OPENAI_API_KEY from the environment

def embed(texts: list[str]) -> list[list[float]]:
    response = client.embeddings.create(
        model="text-embedding-3-small",
        input=texts,
    )
    return [item.embedding for item in response.data]
Enter fullscreen mode Exit fullscreen mode

Step 3: index

For a tutorial, the index is a Python list. A vector database is this list plus fast
search over millions of rows. You do not need one until the list gets slow.

documents = {
    "refunds.md": open("refunds.md").read(),
    "shipping.md": open("shipping.md").read(),
}

index = []  # each entry: (source, chunk_text, embedding)
for source, text in documents.items():
    chunks = chunk(text)
    for chunk_text, vector in zip(chunks, embed(chunks)):
        index.append((source, chunk_text, vector))
Enter fullscreen mode Exit fullscreen mode

Step 4: retrieve

Embed the question with the same model, then find the chunks whose vectors point the
same way. Cosine similarity is the standard measure.

import math

def cosine(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x * x for x in a))
    norm_b = math.sqrt(sum(y * y for y in b))
    return dot / (norm_a * norm_b)

def retrieve(question: str, k: int = 3) -> list[tuple[str, str]]:
    q_vector = embed([question])[0]
    scored = [
        (cosine(q_vector, vector), source, chunk_text)
        for source, chunk_text, vector in index
    ]
    scored.sort(reverse=True)
    return [(source, chunk_text) for _, source, chunk_text in scored[:k]]
Enter fullscreen mode Exit fullscreen mode

Steps 5 and 6: augment and generate

Put the retrieved passages in the prompt, tell the model to answer from them, and ask.

def answer(question: str) -> str:
    passages = retrieve(question)
    context = "\n\n".join(f"[{source}]\n{text}" for source, text in passages)
    prompt = (
        "Answer the question using only the passages below. "
        "If the passages do not contain the answer, say so.\n\n"
        f"{context}\n\nQuestion: {question}"
    )
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    )
    return response.choices[0].message.content

print(answer("How many days do customers have to request a refund?"))
Enter fullscreen mode Exit fullscreen mode

That is a working RAG pipeline. Run it against two small markdown files and you will see
the model quote your refund policy instead of inventing one.

The three places it goes wrong

Every RAG bug in production is one of these.

The right passage was never retrieved. The question used different words from the
document ("money back" versus "refund"), or the chunk boundary split the answer, or the
answer needed two passages and you only fetched one. Print the retrieved passages before
you blame the model. Nine times out of ten the model never saw the answer.

The right passage was retrieved and the model ignored it. Usually the prompt is too
soft. "Use the passages below" is a suggestion; "answer only from the passages, and say
'not in the documents' otherwise" is an instruction. Test with a question the documents
cannot answer and check that the model says so.

The passages are stale. Documents changed and the index did not. The index is a
cache, and it needs the same care as any cache: rebuild it when the source changes.

Where to go from here

Once the basic loop works, the upgrades are all in the retrieve step: hybrid search
(keyword match plus embeddings), reranking the top 20 down to the best 3 with a second
model, and metadata filters so a question about shipping never retrieves a refund
passage. Each one fixes a specific failure you will hit, so add them when you hit it,
not before.

The LogicWiz GenAI course covers this chapter with an animation you can step through
and a lab that runs in the browser: Why Nova needs RAG,
embeddings and semantic search,
chunking, and indexing and retrieval.
The whole course is completely free, with no card. Chapters one to three open without an account, and from chapter four a free account keeps you going.

Originally published on LogicWiz.

Top comments (0)