Ever wanted to upload a PDF and just ask it questions instead of scrolling through fifty pages? That is what a "chat with your PDF" app does, and it works using a technique called RAG (Retrieval-Augmented Generation).
In this guide, I will explain what RAG is in plain words, break down the pieces you need to build one, and walk through a simple working version step by step. By the end, you will understand exactly how these apps work under the hood, and you can build your own from here.
What Is RAG, in Plain Words?
RAG stands for Retrieval-Augmented Generation. That sounds complex, but the idea is simple.
Think of a smart friend who has read every page of your PDF but has a short memory. Every time you ask a question, your friend does not try to remember the whole document. Instead, they quickly flip to the exact pages that answer your question, read just those pages, and then answer you based on what they just read.
That is exactly what a RAG app does:
- Retrieval — find the most relevant parts of the document for your question
- Augmented Generation — hand those relevant parts to an AI model, which uses them to write a proper answer
Without RAG, an AI model only knows what it learned during training — it has never seen your PDF. RAG lets you "show" it the right pages of your document, right when it needs them, so it can answer questions about content it was never trained on.
The Building Blocks of a RAG App
A "chat with your PDF" app is made of five pieces, each doing one small job:
| Step | What it does | Common tools |
|---|---|---|
| 1. Extract text | Pulls the raw text out of the PDF | PyPDF2, pdfplumber |
| 2. Chunk the text | Splits the text into small pieces (a few hundred words each) | Plain Python splitting, LangChain text splitters |
| 3. Create embeddings | Turns each chunk into a list of numbers that capture its meaning | OpenAI embeddings, sentence-transformers |
| 4. Store in a vector database | Saves those embeddings so they can be searched fast | FAISS, ChromaDB, Pinecone |
| 5. Retrieve and answer | Finds the closest matching chunks to your question, sends them to an LLM, returns the answer | OpenAI, Anthropic, or any LLM API |
Once you understand these five steps, you understand how every RAG app works. The fancy tools just do these steps for you automatically.
Building a Simple Version, Step by Step
Here is a minimal working example in Python. It will not be production-ready, but it will actually work, and it shows exactly what is happening at each step.
Step 1: Extract text from the PDF
import pdfplumber
def extract_text(pdf_path):
text = ""
with pdfplumber.open(pdf_path) as pdf:
for page in pdf.pages:
text += page.extract_text() + "\n"
return text
Step 2: Split the text into chunks
def chunk_text(text, chunk_size=500):
words = text.split()
chunks = []
for i in range(0, len(words), chunk_size):
chunk = " ".join(words[i:i + chunk_size])
chunks.append(chunk)
return chunks
Step 3: Turn chunks into embeddings
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('all-MiniLM-L6-v2')
def embed_chunks(chunks):
return model.encode(chunks)
Step 4: Store and search with a vector database
import faiss
import numpy as np
def build_index(embeddings):
dimension = embeddings.shape[1]
index = faiss.IndexFlatL2(dimension)
index.add(np.array(embeddings))
return index
def search(index, chunks, query_embedding, top_k=3):
distances, indices = index.search(np.array([query_embedding]), top_k)
return [chunks[i] for i in indices[0]]
Step 5: Ask the question and get an answer
def ask_question(question, chunks, index):
query_embedding = model.encode([question])[0]
relevant_chunks = search(index, chunks, query_embedding)
context = "\n\n".join(relevant_chunks)
prompt = f"""Answer the question using only the context below.
Context:
{context}
Question: {question}
Answer:"""
# Send this prompt to any LLM API (OpenAI, Anthropic, etc.)
# response = call_llm_api(prompt)
# return response
Put these five pieces together — extract, chunk, embed, store, retrieve and answer — and you have a working "chat with your PDF" app.
Common Mistakes When You Build This
A few things that trip up most people the first time:
- Chunks too big or too small. Too big, and the AI gets confused by irrelevant text mixed in. Too small, and it loses context. 300–500 words per chunk is a good starting point.
- Forgetting to test with real questions. Try questions the document does not actually answer — your app should say "I don't know" instead of making something up.
-
Not handling scanned PDFs. If a PDF is a scanned image and not real text,
pdfplumberwill return nothing. You will need OCR (likepytesseract) for those. - Skipping overlap between chunks. If a sentence gets cut in half between two chunks, the answer can be incomplete. Adding a small overlap (30–50 words) between chunks fixes most of this.
Wrap-Up
That is the whole idea behind "chat with your PDF" apps: extract, chunk, embed, store, retrieve, and answer. None of the individual steps are hard on their own — the skill is in connecting them properly and handling the messy real-world cases, like bad formatting or scanned pages.
If you want to build your own version, the code above is a solid starting point. And if you would rather skip the setup and start from a working, tested codebase instead of debugging embeddings and vector search from scratch, I put together a ready-to-use RAG chatbot source-code kit: RAG source code.
Either way, I hope this helped you understand what is actually happening inside these apps. Let me know in the comments if you get stuck anywhere!
Top comments (0)