I Built a PDF Chatbot Without Fine-Tuning an LLM — Here's How It Works
I had a simple problem.
I had a PDF containing a lot of information, and I wanted to ask questions about it.
Something like:
"What are the main findings?"
"What dataset was used?"
"Explain the methodology in simple terms."
"Where does the paper discuss its limitations?"
My first thought was:
Do I need to train an AI model on the PDF?
No.
I built a system that lets an LLM answer questions about a document without fine-tuning the LLM on that document.
The basic idea is called Retrieval-Augmented Generation, or RAG.
And once I understood how it worked, the architecture was surprisingly straightforward.
What I Wanted to Build
The goal was simple:
Upload PDF
↓
Ask a question
↓
Find the relevant parts of the PDF
↓
Give those parts to the LLM
↓
Generate an answer
For example, imagine I upload a 100-page research paper.
I ask:
"What dataset did the authors use?"
I don't want the LLM to process the entire document from scratch every time.
Instead, my system should find the section containing the dataset information and send only that relevant context to the LLM.
That is the basic idea behind RAG.
The Architecture
The complete pipeline looks like this:
PDF
│
▼
Text Extraction
│
▼
Text Chunking
│
▼
Embedding Generation
│
▼
Vector Database
│
│
┌──────────┘
│
▼
User Question
│
▼
Question Embedding
│
▼
Similarity Search
│
▼
Relevant Chunks
│
▼
LLM
│
▼
Answer
There are two important phases here.
Phase 1: Indexing
The PDF is processed and stored in a searchable format.
Phase 2: Retrieval + Generation
When the user asks a question, the system retrieves the relevant information and gives it to the LLM.
Let's break that down.
Step 1: Extract Text From the PDF
The first thing we need is the actual text.
For a normal text-based PDF, a library such as PyMuPDF can extract it.
A simplified example:
import fitz
def extract_text(pdf_path):
document = fitz.open(pdf_path)
pages = []
for page in document:
pages.append(page.get_text())
return "\n".join(pages)
Now we have something like:
Introduction...
Related Work...
Methodology...
Dataset...
Experiments...
Results...
Conclusion...
But we aren't ready to send this entire text to the LLM.
There could be thousands or hundreds of thousands of words.
So the next step is important.
Step 2: Split the Document Into Chunks
Instead of treating the entire PDF as one giant piece of text, we divide it into smaller chunks.
For example:
PDF
│
├── Chunk 1
├── Chunk 2
├── Chunk 3
├── Chunk 4
├── Chunk 5
└── ...
Why?
Because when someone asks a question, we don't necessarily need the entire document.
Suppose the PDF contains this:
Page 1
Introduction...
Page 2
Related Work...
Page 3
Dataset...
Page 4
Methodology...
Page 5
Training...
Page 6
Results...
If the user asks:
"What dataset was used?"
We only need the relevant portion.
A simple chunking strategy could look like:
def create_chunks(text, chunk_size=1000, overlap=200):
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(text[start:end])
start += chunk_size - overlap
return chunks
The overlap is useful because an important sentence might otherwise fall exactly between two chunks.
For a real project, chunk size should be tested rather than blindly chosen.
Step 3: Convert Text Into Embeddings
Now comes one of the most important parts.
A computer cannot directly perform semantic similarity search on ordinary sentences.
We convert each chunk into a numerical representation called an embedding.
For example:
"What dataset was used?"
might become something conceptually like:
[0.21, -0.18, 0.74, 0.03, ...]
The actual embedding contains many dimensions.
The important idea is that semantically similar text should have similar vector representations.
For example:
Question:
"What dataset did the researchers use?"
Document chunk:
"The experiments were conducted using the WESAD dataset..."
These two pieces of text are semantically related even though they don't use exactly the same words.
That's why embeddings are useful.
Step 4: Store the Embeddings
Now we need somewhere to store the vectors.
A vector database or vector index can be used for this.
For a small local project, FAISS is one option.
Conceptually:
Chunk 1 → Embedding 1
Chunk 2 → Embedding 2
Chunk 3 → Embedding 3
Chunk 4 → Embedding 4
...
We store the relationship between the vector and its original text.
So later, when we find a relevant vector, we can retrieve the original chunk.
The system is essentially building a searchable representation of the PDF.
Step 5: The User Asks a Question
Now the interesting part begins.
Suppose I ask:
"What dataset was used in the experiments?"
The question itself is converted into an embedding.
User Question
↓
Embedding Model
↓
Question Vector
We then compare that vector against the vectors stored from the PDF.
The goal is to find the chunks that are semantically closest to the question.
Step 6: Retrieve the Relevant Chunks
Suppose the search returns:
Chunk 18
Chunk 42
Chunk 43
Chunk 67
The system can select the top few results.
For example:
Question
↓
Vector Search
↓
Top 5 Relevant Chunks
We now have the context needed to answer the question.
This is the retrieval part of Retrieval-Augmented Generation.
Step 7: Give the Context to the LLM
Now we finally use the language model.
Instead of asking:
"What dataset was used?"
we provide the retrieved information as context.
Conceptually:
Use the following context to answer the question.
Context:
[Relevant chunk 1]
[Relevant chunk 2]
[Relevant chunk 3]
Question:
What dataset was used?
Answer:
The LLM can now generate an answer based on the retrieved document content.
This is where the "generation" part of RAG comes in.
The Complete Flow
Putting everything together:
┌──────────────┐
│ PDF │
└──────┬───────┘
│
▼
┌─────────────────┐
│ Text Extraction │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Chunking │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Embeddings │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Vector Index │
└─────────────────┘
User Question
│
▼
┌─────────────────┐
│ Question │
│ Embedding │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Similarity │
│ Search │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Relevant │
│ Document Chunks │
└────────┬────────┘
│
▼
┌─────────────────┐
│ LLM │
└────────┬────────┘
│
▼
Answer
That's the entire idea.
It sounds complicated when people say "build a RAG pipeline."
The individual steps are actually quite understandable.
Why Not Just Put the Entire PDF Into the LLM?
This is a reasonable question.
Modern LLMs can process large amounts of text.
So why bother with retrieval?
Because document-based applications have several practical problems.
A document may be:
- Very large
- Frequently updated
- Made up of many unrelated sections
- Too expensive to repeatedly process in full
- Full of information irrelevant to the current question
Retrieval lets us narrow the information down before generation.
Instead of:
100-page document
↓
LLM
we can do:
100-page document
↓
Find relevant information
↓
5 useful chunks
↓
LLM
The second approach gives the model a much more focused context.
The Most Important Problem: Hallucinations
At this point, the chatbot looks impressive.
But there is a problem.
What happens if the answer isn't actually in the PDF?
Suppose I ask:
"What did the authors say about a technology that isn't mentioned anywhere in the paper?"
A language model might still try to answer.
That's dangerous.
A document chatbot should not confidently invent information simply because the user asked a question.
So one of the most important rules I would add is:
If the retrieved context does not contain enough information to answer the question, say so.
For example:
I couldn't find enough information in the provided
document to answer this question.
That's much better than generating a convincing but unsupported answer.
I Would Test the Chatbot With Questions Like These
Instead of testing only easy questions, I would deliberately try to break it.
Test 1 — Direct question
"What dataset was used?"
Expected:
A direct answer based on the document.
Test 2 — Multiple sections
"How does the proposed method differ from the baseline?"
This may require retrieving information from more than one section.
Test 3 — Missing information
"What programming language was used to build the company's mobile application?"
If the PDF never discusses this, the system should say that it cannot find the answer.
Test 4 — Ambiguous question
"What was the result?"
The document may contain multiple results.
The chatbot should ideally ask for clarification or use the surrounding context.
Test 5 — Misleading question
"Why did the researchers use Dataset X?"
If Dataset X doesn't exist in the document, the system shouldn't accept the assumption as fact.
This type of testing is much more interesting than simply asking:
"What is the title of the paper?"
RAG vs Fine-Tuning
This was one of the biggest things I wanted to understand.
These two approaches solve different problems.
Fine-tuning
Fine-tuning changes the model's behavior by training it further on a dataset.
Conceptually:
Base Model
↓
Training Data
↓
Fine-Tuning
↓
Modified Model
RAG
RAG doesn't require the model to learn the document's contents during training.
Instead:
Document
↓
Index
Question
↓
Retrieve relevant information
↓
LLM
That makes RAG particularly useful for document collections that change frequently.
If I add another PDF, I don't necessarily need to retrain the language model.
I can process the new document and add its chunks to the retrieval system.
A Simple Project Structure
A project like this can be organized fairly cleanly:
pdf-chatbot/
│
├── app.py
├── ingest.py
├── retriever.py
├── generator.py
├── embeddings.py
│
├── data/
│ └── documents/
│
├── index/
│ └── vector_store/
│
├── requirements.txt
└── README.md
For example:
ingest.py
handles:
PDF
→ extraction
→ chunking
→ embeddings
→ indexing
While:
retriever.py
handles:
question
→ embedding
→ similarity search
→ relevant chunks
And:
generator.py
handles:
context + question
→ LLM
→ answer
Keeping those responsibilities separate makes the project easier to debug.
What I Would Put in the UI
The interface doesn't need to be complicated.
Something like:
┌─────────────────────────────────────────┐
│ PDF Question Answering │
├─────────────────────────────────────────┤
│ │
│ [ Upload PDF ] │
│ │
│ ───────────────────────────────────── │
│ │
│ Ask a question: │
│ ┌───────────────────────────────────┐ │
│ │ What dataset was used? │ │
│ └───────────────────────────────────┘ │
│ │
│ [ Ask ] │
│ │
├─────────────────────────────────────────┤
│ Answer │
│ │
│ The experiments used ... │
│ │
├─────────────────────────────────────────┤
│ Sources │
│ │
│ Page 4 │
│ Page 7 │
└─────────────────────────────────────────┘
The source section is particularly useful.
Instead of simply saying:
"Here is the answer."
the application can show:
"Here are the document sections used to generate this answer."
That makes the system easier to inspect.
What I Learned
The interesting thing about this project wasn't actually the chatbot.
It was understanding the separation between knowledge retrieval and language generation.
The LLM doesn't necessarily need to memorize every document.
It can be given the relevant information when the question arrives.
That changes the way you think about building AI applications.
Instead of asking:
"How do I train an AI to know everything in this document?"
you can ask:
"How do I efficiently retrieve the right information and give it to the model?"
That is a much more practical engineering problem.
Where This Can Be Used
The same architecture can be adapted to many applications.
Research papers
Upload papers and ask questions about:
- methodology
- datasets
- experiments
- limitations
- results
College notes
Upload lecture notes and ask:
"Explain Unit 3 in simple terms."
Company documentation
Upload internal documentation and search it conversationally.
Legal documents
Retrieve relevant sections from large documents.
Product manuals
Ask:
"How do I reset this device?"
Technical documentation
Ask questions without manually searching through hundreds of pages.
The basic architecture remains similar.
What I Would Improve Next
The first version of a RAG system is relatively simple.
A production-quality version is much harder.
There are several things I would improve.
1. Better chunking
Fixed character lengths aren't always ideal.
A chunk should ideally preserve meaningful context.
2. Better retrieval
Basic similarity search isn't always enough.
Hybrid search can combine semantic similarity with keyword-based retrieval.
3. Re-ranking
Instead of immediately passing the top results to the LLM, a re-ranking model can help determine which retrieved chunks are actually the most relevant.
4. Source citations
The chatbot should tell the user exactly where its answer came from.
5. Better handling of tables
PDFs aren't just text.
Tables, figures, columns, and scanned pages can make extraction much harder.
6. Evaluation
A serious RAG application needs evaluation.
I'd want to measure things such as:
- Retrieval accuracy
- Answer correctness
- Context relevance
- Hallucination rate
- Response latency
Without evaluation, "it seems to work" isn't enough.
The Bigger Picture
The interesting part about RAG isn't that it lets you "chat with PDFs."
That's just one application.
The larger idea is that an LLM can be connected to an external knowledge source without having that knowledge permanently embedded into its parameters.
The architecture becomes:
External Knowledge
↓
Retrieval
↓
Context
↓
LLM
↓
Response
And that pattern can be used far beyond PDFs.
That's what made this project interesting to me.
I started with a simple question:
"Can I make an AI answer questions about my PDF?"
The answer was yes.
But the more interesting realization was:
I didn't need to train the AI to know the document.
I just needed to build a good way for the AI to find the right information.
And that is the basic idea behind Retrieval-Augmented Generation.
Final Takeaway
If you're building your first document-based AI application, don't immediately think about fine-tuning.
Start by asking:
Can I retrieve the right information
and give it to the model at the right time?
If the answer is yes, you may already have most of what you need.
The model generates the language.
The retrieval system finds the knowledge.
And the application connects the two.
That's RAG.
Top comments (0)