DEV Community

Cover image for Chunking — Why How You Split Documents Decides Retrieval Quality
Sham Prakash K
Sham Prakash K

Posted on AI-assisted

Chunking — Why How You Split Documents Decides Retrieval Quality

The previous article explained how embeddings convert text into vectors so similar meanings land close together in vector space. Before you can embed a document, though, you have to decide how to cut it up. That decision — chunk size, chunk strategy, chunk overlap — has more impact on retrieval quality than almost anything else in a RAG system. And it fails silently. You won't get an error. You'll just get bad answers.

Why chunking exists

An embedding model converts a piece of text into a single vector. That vector represents the overall meaning of the entire input.

If you embed a 2,000-word document as one piece, you get one vector that represents the average meaning of all 2,000 words. When a user asks a specific question, the retrieved vector is the right document — but the 2,000 words contain dozens of topics, and the specific answer is buried somewhere in the middle. The model receives 2,000 words of context to search through and often glosses over the detail it needs.

If you split that same document into 300-word pieces and embed each piece separately, each vector represents a focused slice of meaning. When a user asks a specific question, the retrieved chunk contains mostly relevant content — and the model can extract the answer cleanly.

Chunking is how you control what unit of meaning gets embedded and retrieved.


Fixed-size chunking

The simplest strategy: split text every N characters, regardless of where sentences or paragraphs end.

// Spring AI: TokenTextSplitter with fixed token count
TokenTextSplitter splitter = new TokenTextSplitter(
    400,   // target chunk size in tokens
    50,    // minimum chunk size
    10,    // number of sentences to use for estimation
    10000, // max number of characters per chunk
    true   // keep separator
);

List<Document> chunks = splitter.split(document);
Enter fullscreen mode Exit fullscreen mode

What it does well: Simple. Fast. Predictable chunk sizes. Good for structured content like product listings or short FAQ entries where every item is roughly self-contained.

What it does poorly: It cuts mid-sentence. A sentence that starts near the end of one chunk continues in the next. The embedding for the first chunk includes an incomplete thought. The embedding for the second chunk starts mid-sentence with no context. Both vectors are slightly off.

For prose documents — policies, articles, manuals — fixed-size splitting consistently degrades retrieval quality at the boundaries.


Recursive character splitting

A smarter approach: try to split at natural boundaries first, fall back to smaller boundaries only when needed.

The algorithm works like this:

  1. Try to split at double newlines (\n\n) — paragraph breaks
  2. If a resulting piece is still too large, split at single newlines (\n)
  3. If still too large, split at sentence endings (., !, ?)
  4. If still too large, split at spaces
  5. If still too large, split at characters

The result is chunks that respect the natural structure of the text as much as possible. Paragraphs stay together until they're too large. Sentences stay together until they're too large.

// In Spring AI, configure the splitter to respect sentence boundaries
TokenTextSplitter splitter = new TokenTextSplitter();
splitter.setDefaultChunkSize(350);
splitter.setMinChunkSizeChars(100);
Enter fullscreen mode Exit fullscreen mode

For most documents — policies, documentation, articles — recursive splitting produces noticeably better retrieval than fixed-size splitting. The chunks are semantically coherent because they follow the author's own structure.


Semantic chunking

The most advanced strategy: instead of splitting by size, split where the meaning changes.

The algorithm embeds every sentence, then compares the embedding of each sentence to its neighbours. When the cosine similarity drops sharply — when two adjacent sentences are less related than the average — that's a topic boundary. Split there.

Sentence 1: "The return policy for electronics covers 30 days."     → embed
Sentence 2: "Items must be in original condition with receipt."      → embed
Sentence 3: "For software products, all sales are final."           → embed
Sentence 4: "Our shipping partners include FedEx and UPS."          → embed ← similarity drops here
Sentence 5: "Standard delivery takes 3–5 business days."            → embed
Enter fullscreen mode Exit fullscreen mode

Sentences 1–3 are about return policy. Sentence 4 shifts to shipping. The similarity drop triggers a split. The result is chunks that are semantically complete — each one covers one topic from start to finish.

The tradeoff: semantic chunking requires embedding every sentence during ingest, which is slower and more expensive. For large document libraries, the cost adds up. It produces the best retrieval quality, but recursive splitting gets you 80% of the way there at a fraction of the compute.


Chunk size: the single most important parameter

Whatever strategy you use, chunk size is the most impactful variable. Here is what happens at each end of the range:

Too large (1,000–2,000+ tokens):

  • One vector covers too many topics. Retrieval finds the right section but returns far more context than needed.
  • The model receives a wall of text with the answer buried inside. Response quality degrades — the model summarises instead of answering specifically.
  • Fewer chunks are stored, so similarity scores are spread thin across large sections.

Too small (30–50 tokens):

  • A sentence fragment gets its own vector. The answer to a multi-sentence question is split across five different chunks. Retrieval might return three fragments that together form an answer — but out of order, with gaps.
  • Short chunks lack context. "It must be returned within 30 days" is meaningless without the sentence before it that says what "it" refers to.

The practical starting range: 200–400 tokens.

This is where most production RAG systems land after tuning. It's small enough that each chunk covers one focused idea, and large enough that each chunk has enough context to be useful on its own. Adjust from there based on your document structure and what your retrieval results actually look like.


Chunk overlap

Even with good splitting, a fact sometimes lands exactly at a boundary. The end of one chunk contains the first half of a sentence; the beginning of the next chunk contains the second half. Both chunks miss the complete thought.

Overlap solves this: each chunk shares N tokens with the previous chunk.

Chunk 1: [...sentence A... sentence B... sentence C...]
Chunk 2:                  [sentence B... sentence C... sentence D... sentence E...]
Chunk 3:                                              [sentence D... sentence E... sentence F...]
Enter fullscreen mode Exit fullscreen mode

Sentences B and C appear in both chunk 1 and chunk 2. If a question's answer spans the boundary between chunk 1 and chunk 2, either chunk can surface it.

Typical overlap: 10–20% of chunk size. For a 300-token chunk, 30–50 tokens of overlap is a reasonable starting point. More overlap means more redundant storage and more tokens injected into prompts; less overlap risks missing boundary-spanning facts.


Tables and code blocks

Recursive splitting and fixed-size splitting both have a blind spot: structured content.

Tables split badly. A markdown or HTML table has a header row and data rows. If the splitter cuts mid-table, you get two chunks: one with the header and a few rows, one with the remaining rows and no header. Both embed with degraded meaning — the second chunk has no column names, so the model has no idea what the numbers represent when it reads that chunk.

Code blocks have the same problem. A function split across two chunks produces one chunk ending mid-function and one starting in the middle of a method body. Neither chunk embeds or reads correctly.

The fix: treat tables and code blocks as atomic units. Never split inside them. Before running your splitting logic, identify blocks of this type and mark them as unsplittable — pass them through as a single chunk regardless of size. Most document parsers can identify these boundaries before splitting begins.

If a table or code block is genuinely too large to pass as one chunk, extract it separately with its own pre- and post-context (the heading above it, the sentence that introduces it) so it embeds with enough surrounding meaning to be retrievable.


Chunk metadata — tracking where each chunk came from

Every chunk should carry metadata: which document it came from, and ideally which section or page.

public void ingest(String documentText, String sourceId, String title) {
    Document doc = new Document(
        documentText,
        Map.of(
            "source", sourceId,   // document ID or filename
            "title",  title       // human-readable name
        )
    );
    List<Document> chunks = splitter.split(doc);
    vectorStore.add(chunks);
}
Enter fullscreen mode Exit fullscreen mode

Spring AI propagates metadata from the parent Document to every chunk automatically. When a chunk is retrieved, its metadata comes back with it.

This matters for two reasons.

Citations. When the model answers from retrieved chunks, you can tell the user which document the answer came from: "Based on the Electronics Return Policy (v2.3)..." Without metadata, retrieved chunks are anonymous — you can't tell the user where the answer came from and you can't verify it yourself.

Debugging. When retrieval returns the wrong chunks, metadata is how you diagnose it. You can log which chunks were retrieved for each query and trace them back to the source document. Without metadata, a bad retrieval result is opaque — you see the text but have no path back to the original.

A minimum useful metadata set: source (document ID or filename), title (human-readable name). Add page or section when your documents have them.


Document updates and re-ingestion

Documents change. Policies get updated. Products get new specs. When a document changes, you need to update the chunks in your vector store.

The naive approach — just re-run ingest — creates a problem: you now have two sets of chunks from the same document. The old chunks and the new chunks both live in the vector store. Retrieval will return chunks from both versions, potentially mixing outdated content with current content. The model won't know which version is correct.

The right approach: delete before re-ingesting.

public void update(String sourceId, String newText, String title) {
    // Delete all chunks from the old version
    vectorStore.delete(List.of(sourceId));

    // Re-ingest the new version
    ingest(newText, sourceId, title);
}
Enter fullscreen mode Exit fullscreen mode

This requires that sourceId is consistent across versions — the same document always uses the same ID. With pgvector, the delete is a SQL DELETE WHERE metadata->>'source' = ?. Spring AI's VectorStore interface exposes a delete(List<String> ids) method that handles this.

If you skip the delete step, stale chunks accumulate silently. Retrieval starts mixing answers from old and new versions with no indication that anything is wrong.


Ingestion pipeline design

Every document goes through the same pipeline before it can be retrieved:

Document (PDF, text, markdown)
     ↓
1. Parse — extract plain text from the file format
     ↓
2. Clean — remove headers, footers, boilerplate, formatting artefacts
     ↓
3. Split — cut into chunks (respecting tables and code blocks as atomic units)
     ↓
4. Embed — convert each chunk to a vector using the embedding model
     ↓
5. Store — save each chunk + its vector + metadata in the vector database
Enter fullscreen mode Exit fullscreen mode

Steps 1–3 happen entirely in your application. Steps 4–5 involve the embedding API and the vector database. The whole pipeline runs once per document. Add a new document today, it's searchable immediately — no retraining, no redeployment.

In Spring AI, the pipeline collapses to a few lines:

@Autowired
private VectorStore vectorStore;

@Autowired
private TokenTextSplitter splitter;

public void ingest(String documentText, String sourceId, String title) {
    Document doc = new Document(
        documentText,
        Map.of("source", sourceId, "title", title)
    );
    List<Document> chunks = splitter.split(doc);
    vectorStore.add(chunks);  // embeds and stores in one call
}
Enter fullscreen mode Exit fullscreen mode

VectorStore.add() handles the embedding and storage together. Spring AI calls the embedding API for each chunk and writes the vectors to the configured store — pgvector, Pinecone, or any other supported backend.

For batch ingestion of multiple documents:

public void ingestAll(List<Map<String, String>> documents) {
    List<Document> allChunks = documents.stream()
        .map(d -> new Document(d.get("text"), Map.of(
            "source", d.get("id"),
            "title",  d.get("title")
        )))
        .flatMap(doc -> splitter.split(doc).stream())
        .toList();

    vectorStore.add(allChunks);
}
Enter fullscreen mode Exit fullscreen mode

Batch is significantly faster than ingesting one document at a time — the embedding API processes chunks in batches, reducing the number of HTTP round trips.


A mistake worth knowing

When I first built the ingest pipeline, I used 2,000-token chunks — roughly half a page of text. My thinking was: bigger chunks mean more context per retrieval, so the model has more to work with.

Retrieval looked fine on the surface. The top chunks coming back were semantically relevant — they were from the right sections. But the actual responses were vague. Questions that should have had specific, crisp answers were getting general summaries instead.

The problem took a while to diagnose because the retrieval metrics looked correct. The issue was in what the model was receiving: a 2,000-token chunk contains enough content to embed as semantically close to the query, but also contains 1,800 tokens of surrounding text that dilutes the specific answer. The model was reading the right chapter but looking for a needle in a haystack every time.

Dropping to 300-token chunks with 50-token overlap fixed it. Specific questions started getting specific answers. The retrieval precision went up because each chunk now represented one coherent idea rather than a sprawling section.

The rule I use now: start at 300 tokens, 50-token overlap. Run real queries against real data. If answers are too vague, chunk smaller. If answers are too fragmented or lack context, chunk larger.


What's next

Your documents are now split into focused chunks and stored as vectors. The piece you need next is the database that stores and searches those vectors efficiently. A regular database can't do this — you need a vector database, and for a Spring Boot application the most practical option is pgvector: the vector extension for PostgreSQL.

That's what the next article covers — what a vector database does, why a regular index doesn't work for similarity search, and how to set up pgvector with Spring AI.


How are you splitting your documents? Fixed chunks or something smarter? Drop your setup in the comments — I'm curious what chunk sizes people are landing on.

Sham Prakash K — Backend Engineer, 4+ years in Java, Spring Boot, and distributed systems. Building AI backend infrastructure. Writing about what I actually learned, mistakes included.

Top comments (0)