DEV Community

Cover image for I Made an LLM Read My PDFs Without Fine-Tuning It
Utsav D
Utsav D

Posted on

I Made an LLM Read My PDFs Without Fine-Tuning It

I Built a PDF Chatbot Without Fine-Tuning an LLM — Here's How It Works

I had a simple problem.

I had a PDF containing a lot of information, and I wanted to ask questions about it.

Something like:

"What are the main findings?"

"What dataset was used?"

"Explain the methodology in simple terms."

"Where does the paper discuss its limitations?"

My first thought was:

Do I need to train an AI model on the PDF?

No.

I built a system that lets an LLM answer questions about a document without fine-tuning the LLM on that document.

The basic idea is called Retrieval-Augmented Generation, or RAG.

And once I understood how it worked, the architecture was surprisingly straightforward.


What I Wanted to Build

The goal was simple:

Upload PDF
     ↓
Ask a question
     ↓
Find the relevant parts of the PDF
     ↓
Give those parts to the LLM
     ↓
Generate an answer
Enter fullscreen mode Exit fullscreen mode

For example, imagine I upload a 100-page research paper.

I ask:

"What dataset did the authors use?"

I don't want the LLM to process the entire document from scratch every time.

Instead, my system should find the section containing the dataset information and send only that relevant context to the LLM.

That is the basic idea behind RAG.


The Architecture

The complete pipeline looks like this:

                  PDF
                   │
                   ▼
            Text Extraction
                   │
                   ▼
             Text Chunking
                   │
                   ▼
          Embedding Generation
                   │
                   ▼
            Vector Database
                   │
                   │
        ┌──────────┘
        │
        ▼
   User Question
        │
        ▼
 Question Embedding
        │
        ▼
 Similarity Search
        │
        ▼
 Relevant Chunks
        │
        ▼
      LLM
        │
        ▼
     Answer
Enter fullscreen mode Exit fullscreen mode

There are two important phases here.

Phase 1: Indexing

The PDF is processed and stored in a searchable format.

Phase 2: Retrieval + Generation

When the user asks a question, the system retrieves the relevant information and gives it to the LLM.

Let's break that down.


Step 1: Extract Text From the PDF

The first thing we need is the actual text.

For a normal text-based PDF, a library such as PyMuPDF can extract it.

A simplified example:

import fitz

def extract_text(pdf_path):
    document = fitz.open(pdf_path)

    pages = []

    for page in document:
        pages.append(page.get_text())

    return "\n".join(pages)
Enter fullscreen mode Exit fullscreen mode

Now we have something like:

Introduction...

Related Work...

Methodology...

Dataset...

Experiments...

Results...

Conclusion...
Enter fullscreen mode Exit fullscreen mode

But we aren't ready to send this entire text to the LLM.

There could be thousands or hundreds of thousands of words.

So the next step is important.


Step 2: Split the Document Into Chunks

Instead of treating the entire PDF as one giant piece of text, we divide it into smaller chunks.

For example:

PDF
│
├── Chunk 1
├── Chunk 2
├── Chunk 3
├── Chunk 4
├── Chunk 5
└── ...
Enter fullscreen mode Exit fullscreen mode

Why?

Because when someone asks a question, we don't necessarily need the entire document.

Suppose the PDF contains this:

Page 1
Introduction...

Page 2
Related Work...

Page 3
Dataset...

Page 4
Methodology...

Page 5
Training...

Page 6
Results...
Enter fullscreen mode Exit fullscreen mode

If the user asks:

"What dataset was used?"

We only need the relevant portion.

A simple chunking strategy could look like:

def create_chunks(text, chunk_size=1000, overlap=200):
    chunks = []

    start = 0

    while start < len(text):
        end = start + chunk_size
        chunks.append(text[start:end])
        start += chunk_size - overlap

    return chunks
Enter fullscreen mode Exit fullscreen mode

The overlap is useful because an important sentence might otherwise fall exactly between two chunks.

For a real project, chunk size should be tested rather than blindly chosen.


Step 3: Convert Text Into Embeddings

Now comes one of the most important parts.

A computer cannot directly perform semantic similarity search on ordinary sentences.

We convert each chunk into a numerical representation called an embedding.

For example:

"What dataset was used?"
Enter fullscreen mode Exit fullscreen mode

might become something conceptually like:

[0.21, -0.18, 0.74, 0.03, ...]
Enter fullscreen mode Exit fullscreen mode

The actual embedding contains many dimensions.

The important idea is that semantically similar text should have similar vector representations.

For example:

Question:
"What dataset did the researchers use?"

Document chunk:
"The experiments were conducted using the WESAD dataset..."
Enter fullscreen mode Exit fullscreen mode

These two pieces of text are semantically related even though they don't use exactly the same words.

That's why embeddings are useful.


Step 4: Store the Embeddings

Now we need somewhere to store the vectors.

A vector database or vector index can be used for this.

For a small local project, FAISS is one option.

Conceptually:

Chunk 1 → Embedding 1
Chunk 2 → Embedding 2
Chunk 3 → Embedding 3
Chunk 4 → Embedding 4
...
Enter fullscreen mode Exit fullscreen mode

We store the relationship between the vector and its original text.

So later, when we find a relevant vector, we can retrieve the original chunk.

The system is essentially building a searchable representation of the PDF.


Step 5: The User Asks a Question

Now the interesting part begins.

Suppose I ask:

"What dataset was used in the experiments?"

The question itself is converted into an embedding.

User Question
      ↓
Embedding Model
      ↓
Question Vector
Enter fullscreen mode Exit fullscreen mode

We then compare that vector against the vectors stored from the PDF.

The goal is to find the chunks that are semantically closest to the question.


Step 6: Retrieve the Relevant Chunks

Suppose the search returns:

Chunk 18
Chunk 42
Chunk 43
Chunk 67
Enter fullscreen mode Exit fullscreen mode

The system can select the top few results.

For example:

Question
   ↓
Vector Search
   ↓
Top 5 Relevant Chunks
Enter fullscreen mode Exit fullscreen mode

We now have the context needed to answer the question.

This is the retrieval part of Retrieval-Augmented Generation.


Step 7: Give the Context to the LLM

Now we finally use the language model.

Instead of asking:

"What dataset was used?"
Enter fullscreen mode Exit fullscreen mode

we provide the retrieved information as context.

Conceptually:

Use the following context to answer the question.

Context:
[Relevant chunk 1]

[Relevant chunk 2]

[Relevant chunk 3]

Question:
What dataset was used?

Answer:
Enter fullscreen mode Exit fullscreen mode

The LLM can now generate an answer based on the retrieved document content.

This is where the "generation" part of RAG comes in.


The Complete Flow

Putting everything together:

                  ┌──────────────┐
                  │     PDF      │
                  └──────┬───────┘
                         │
                         ▼
                ┌─────────────────┐
                │ Text Extraction │
                └────────┬────────┘
                         │
                         ▼
                ┌─────────────────┐
                │    Chunking     │
                └────────┬────────┘
                         │
                         ▼
                ┌─────────────────┐
                │   Embeddings    │
                └────────┬────────┘
                         │
                         ▼
                ┌─────────────────┐
                │  Vector Index   │
                └─────────────────┘


User Question
      │
      ▼
┌─────────────────┐
│ Question        │
│ Embedding       │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Similarity      │
│ Search          │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Relevant        │
│ Document Chunks │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│      LLM        │
└────────┬────────┘
         │
         ▼
      Answer
Enter fullscreen mode Exit fullscreen mode

That's the entire idea.

It sounds complicated when people say "build a RAG pipeline."

The individual steps are actually quite understandable.


Why Not Just Put the Entire PDF Into the LLM?

This is a reasonable question.

Modern LLMs can process large amounts of text.

So why bother with retrieval?

Because document-based applications have several practical problems.

A document may be:

  • Very large
  • Frequently updated
  • Made up of many unrelated sections
  • Too expensive to repeatedly process in full
  • Full of information irrelevant to the current question

Retrieval lets us narrow the information down before generation.

Instead of:

100-page document
      ↓
      LLM
Enter fullscreen mode Exit fullscreen mode

we can do:

100-page document
      ↓
Find relevant information
      ↓
5 useful chunks
      ↓
LLM
Enter fullscreen mode Exit fullscreen mode

The second approach gives the model a much more focused context.


The Most Important Problem: Hallucinations

At this point, the chatbot looks impressive.

But there is a problem.

What happens if the answer isn't actually in the PDF?

Suppose I ask:

"What did the authors say about a technology that isn't mentioned anywhere in the paper?"

A language model might still try to answer.

That's dangerous.

A document chatbot should not confidently invent information simply because the user asked a question.

So one of the most important rules I would add is:

If the retrieved context does not contain enough information to answer the question, say so.

For example:

I couldn't find enough information in the provided
document to answer this question.
Enter fullscreen mode Exit fullscreen mode

That's much better than generating a convincing but unsupported answer.


I Would Test the Chatbot With Questions Like These

Instead of testing only easy questions, I would deliberately try to break it.

Test 1 — Direct question

"What dataset was used?"

Expected:

A direct answer based on the document.
Enter fullscreen mode Exit fullscreen mode

Test 2 — Multiple sections

"How does the proposed method differ from the baseline?"

This may require retrieving information from more than one section.

Test 3 — Missing information

"What programming language was used to build the company's mobile application?"

If the PDF never discusses this, the system should say that it cannot find the answer.

Test 4 — Ambiguous question

"What was the result?"

The document may contain multiple results.

The chatbot should ideally ask for clarification or use the surrounding context.

Test 5 — Misleading question

"Why did the researchers use Dataset X?"

If Dataset X doesn't exist in the document, the system shouldn't accept the assumption as fact.

This type of testing is much more interesting than simply asking:

"What is the title of the paper?"


RAG vs Fine-Tuning

This was one of the biggest things I wanted to understand.

These two approaches solve different problems.

Fine-tuning

Fine-tuning changes the model's behavior by training it further on a dataset.

Conceptually:

Base Model
    ↓
Training Data
    ↓
Fine-Tuning
    ↓
Modified Model
Enter fullscreen mode Exit fullscreen mode

RAG

RAG doesn't require the model to learn the document's contents during training.

Instead:

Document
    ↓
Index

Question
    ↓
Retrieve relevant information
    ↓
LLM
Enter fullscreen mode Exit fullscreen mode

That makes RAG particularly useful for document collections that change frequently.

If I add another PDF, I don't necessarily need to retrain the language model.

I can process the new document and add its chunks to the retrieval system.


A Simple Project Structure

A project like this can be organized fairly cleanly:

pdf-chatbot/
│
├── app.py
├── ingest.py
├── retriever.py
├── generator.py
├── embeddings.py
│
├── data/
│   └── documents/
│
├── index/
│   └── vector_store/
│
├── requirements.txt
└── README.md
Enter fullscreen mode Exit fullscreen mode

For example:

ingest.py
Enter fullscreen mode Exit fullscreen mode

handles:

PDF
→ extraction
→ chunking
→ embeddings
→ indexing
Enter fullscreen mode Exit fullscreen mode

While:

retriever.py
Enter fullscreen mode Exit fullscreen mode

handles:

question
→ embedding
→ similarity search
→ relevant chunks
Enter fullscreen mode Exit fullscreen mode

And:

generator.py
Enter fullscreen mode Exit fullscreen mode

handles:

context + question
→ LLM
→ answer
Enter fullscreen mode Exit fullscreen mode

Keeping those responsibilities separate makes the project easier to debug.


What I Would Put in the UI

The interface doesn't need to be complicated.

Something like:

┌─────────────────────────────────────────┐
│           PDF Question Answering        │
├─────────────────────────────────────────┤
│                                         │
│       [ Upload PDF ]                    │
│                                         │
│  ─────────────────────────────────────  │
│                                         │
│  Ask a question:                        │
│  ┌───────────────────────────────────┐  │
│  │ What dataset was used?            │  │
│  └───────────────────────────────────┘  │
│                                         │
│              [ Ask ]                    │
│                                         │
├─────────────────────────────────────────┤
│ Answer                                  │
│                                         │
│ The experiments used ...                │
│                                         │
├─────────────────────────────────────────┤
│ Sources                                 │
│                                         │
│ Page 4                                  │
│ Page 7                                  │
└─────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The source section is particularly useful.

Instead of simply saying:

"Here is the answer."

the application can show:

"Here are the document sections used to generate this answer."

That makes the system easier to inspect.


What I Learned

The interesting thing about this project wasn't actually the chatbot.

It was understanding the separation between knowledge retrieval and language generation.

The LLM doesn't necessarily need to memorize every document.

It can be given the relevant information when the question arrives.

That changes the way you think about building AI applications.

Instead of asking:

"How do I train an AI to know everything in this document?"

you can ask:

"How do I efficiently retrieve the right information and give it to the model?"

That is a much more practical engineering problem.


Where This Can Be Used

The same architecture can be adapted to many applications.

Research papers

Upload papers and ask questions about:

  • methodology
  • datasets
  • experiments
  • limitations
  • results

College notes

Upload lecture notes and ask:

"Explain Unit 3 in simple terms."

Company documentation

Upload internal documentation and search it conversationally.

Legal documents

Retrieve relevant sections from large documents.

Product manuals

Ask:

"How do I reset this device?"

Technical documentation

Ask questions without manually searching through hundreds of pages.

The basic architecture remains similar.


What I Would Improve Next

The first version of a RAG system is relatively simple.

A production-quality version is much harder.

There are several things I would improve.

1. Better chunking

Fixed character lengths aren't always ideal.

A chunk should ideally preserve meaningful context.

2. Better retrieval

Basic similarity search isn't always enough.

Hybrid search can combine semantic similarity with keyword-based retrieval.

3. Re-ranking

Instead of immediately passing the top results to the LLM, a re-ranking model can help determine which retrieved chunks are actually the most relevant.

4. Source citations

The chatbot should tell the user exactly where its answer came from.

5. Better handling of tables

PDFs aren't just text.

Tables, figures, columns, and scanned pages can make extraction much harder.

6. Evaluation

A serious RAG application needs evaluation.

I'd want to measure things such as:

  • Retrieval accuracy
  • Answer correctness
  • Context relevance
  • Hallucination rate
  • Response latency

Without evaluation, "it seems to work" isn't enough.


The Bigger Picture

The interesting part about RAG isn't that it lets you "chat with PDFs."

That's just one application.

The larger idea is that an LLM can be connected to an external knowledge source without having that knowledge permanently embedded into its parameters.

The architecture becomes:

External Knowledge
       ↓
    Retrieval
       ↓
    Context
       ↓
      LLM
       ↓
    Response
Enter fullscreen mode Exit fullscreen mode

And that pattern can be used far beyond PDFs.

That's what made this project interesting to me.

I started with a simple question:

"Can I make an AI answer questions about my PDF?"

The answer was yes.

But the more interesting realization was:

I didn't need to train the AI to know the document.

I just needed to build a good way for the AI to find the right information.

And that is the basic idea behind Retrieval-Augmented Generation.


Final Takeaway

If you're building your first document-based AI application, don't immediately think about fine-tuning.

Start by asking:

Can I retrieve the right information
and give it to the model at the right time?
Enter fullscreen mode Exit fullscreen mode

If the answer is yes, you may already have most of what you need.

The model generates the language.

The retrieval system finds the knowledge.

And the application connects the two.

That's RAG.

Top comments (0)