DEV Community

Cover image for How to Build a "Chat With Your PDFs" App: A Simple RAG Guide
Dharanidharan
Dharanidharan

Posted on

How to Build a "Chat With Your PDFs" App: A Simple RAG Guide

Ever wanted to upload a PDF and just ask it questions instead of scrolling through fifty pages? That is what a "chat with your PDF" app does, and it works using a technique called RAG (Retrieval-Augmented Generation).

In this guide, I will explain what RAG is in plain words, break down the pieces you need to build one, and walk through a simple working version step by step. By the end, you will understand exactly how these apps work under the hood, and you can build your own from here.

What Is RAG, in Plain Words?

RAG stands for Retrieval-Augmented Generation. That sounds complex, but the idea is simple.

Think of a smart friend who has read every page of your PDF but has a short memory. Every time you ask a question, your friend does not try to remember the whole document. Instead, they quickly flip to the exact pages that answer your question, read just those pages, and then answer you based on what they just read.

That is exactly what a RAG app does:

  1. Retrieval — find the most relevant parts of the document for your question
  2. Augmented Generation — hand those relevant parts to an AI model, which uses them to write a proper answer

Without RAG, an AI model only knows what it learned during training — it has never seen your PDF. RAG lets you "show" it the right pages of your document, right when it needs them, so it can answer questions about content it was never trained on.

The Building Blocks of a RAG App

A "chat with your PDF" app is made of five pieces, each doing one small job:

Step What it does Common tools
1. Extract text Pulls the raw text out of the PDF PyPDF2, pdfplumber
2. Chunk the text Splits the text into small pieces (a few hundred words each) Plain Python splitting, LangChain text splitters
3. Create embeddings Turns each chunk into a list of numbers that capture its meaning OpenAI embeddings, sentence-transformers
4. Store in a vector database Saves those embeddings so they can be searched fast FAISS, ChromaDB, Pinecone
5. Retrieve and answer Finds the closest matching chunks to your question, sends them to an LLM, returns the answer OpenAI, Anthropic, or any LLM API

Once you understand these five steps, you understand how every RAG app works. The fancy tools just do these steps for you automatically.

Building a Simple Version, Step by Step

Here is a minimal working example in Python. It will not be production-ready, but it will actually work, and it shows exactly what is happening at each step.

Step 1: Extract text from the PDF

import pdfplumber

def extract_text(pdf_path):
    text = ""
    with pdfplumber.open(pdf_path) as pdf:
        for page in pdf.pages:
            text += page.extract_text() + "\n"
    return text
Enter fullscreen mode Exit fullscreen mode

Step 2: Split the text into chunks

def chunk_text(text, chunk_size=500):
    words = text.split()
    chunks = []
    for i in range(0, len(words), chunk_size):
        chunk = " ".join(words[i:i + chunk_size])
        chunks.append(chunk)
    return chunks
Enter fullscreen mode Exit fullscreen mode

Step 3: Turn chunks into embeddings

from sentence_transformers import SentenceTransformer

model = SentenceTransformer('all-MiniLM-L6-v2')

def embed_chunks(chunks):
    return model.encode(chunks)
Enter fullscreen mode Exit fullscreen mode

Step 4: Store and search with a vector database

import faiss
import numpy as np

def build_index(embeddings):
    dimension = embeddings.shape[1]
    index = faiss.IndexFlatL2(dimension)
    index.add(np.array(embeddings))
    return index

def search(index, chunks, query_embedding, top_k=3):
    distances, indices = index.search(np.array([query_embedding]), top_k)
    return [chunks[i] for i in indices[0]]
Enter fullscreen mode Exit fullscreen mode

Step 5: Ask the question and get an answer

def ask_question(question, chunks, index):
    query_embedding = model.encode([question])[0]
    relevant_chunks = search(index, chunks, query_embedding)
    context = "\n\n".join(relevant_chunks)

    prompt = f"""Answer the question using only the context below.

Context:
{context}

Question: {question}
Answer:"""

    # Send this prompt to any LLM API (OpenAI, Anthropic, etc.)
    # response = call_llm_api(prompt)
    # return response
Enter fullscreen mode Exit fullscreen mode

Put these five pieces together — extract, chunk, embed, store, retrieve and answer — and you have a working "chat with your PDF" app.

Common Mistakes When You Build This

A few things that trip up most people the first time:

  • Chunks too big or too small. Too big, and the AI gets confused by irrelevant text mixed in. Too small, and it loses context. 300–500 words per chunk is a good starting point.
  • Forgetting to test with real questions. Try questions the document does not actually answer — your app should say "I don't know" instead of making something up.
  • Not handling scanned PDFs. If a PDF is a scanned image and not real text, pdfplumber will return nothing. You will need OCR (like pytesseract) for those.
  • Skipping overlap between chunks. If a sentence gets cut in half between two chunks, the answer can be incomplete. Adding a small overlap (30–50 words) between chunks fixes most of this.

Wrap-Up

That is the whole idea behind "chat with your PDF" apps: extract, chunk, embed, store, retrieve, and answer. None of the individual steps are hard on their own — the skill is in connecting them properly and handling the messy real-world cases, like bad formatting or scanned pages.

If you want to build your own version, the code above is a solid starting point. And if you would rather skip the setup and start from a working, tested codebase instead of debugging embeddings and vector search from scratch, I put together a ready-to-use RAG chatbot source-code kit: RAG source code.

Either way, I hope this helped you understand what is actually happening inside these apps. Let me know in the comments if you get stuck anywhere!

Top comments (0)