DEV Community

Cover image for Chunking in RAG: How to Split Documents Without Losing the Conditions
NEXT4I DEV
NEXT4I DEV

Posted on Originally published at next4i.com

Chunking in RAG: How to Split Documents Without Losing the Conditions


Your AI assistant has the employee benefits handbook. You ask: “Can someone who is still on probation claim the eyewear benefit?”

It confidently answers: “Yes, up to 3,000 baht per year.” But the handbook says the benefit is only for permanent employees who have completed probation. The limit reached the model. The eligibility condition did not.

That failure can start before generation, at the document boundary. Chunking in RAG is not about making text as small as possible. It is about creating searchable pieces without separating facts from the conditions that make them true.

The policy and amounts here are fictional examples, not NEXT4I or customer policies.

TL;DR

  • A chunk is a piece of source content used for retrieval. It is not inherently a vector.
  • Large chunks can mix unrelated topics. Small chunks can lose subjects, conditions, and exceptions.
  • Start with document structure, cap oversized sections, and test with real questions.
  • Overlap and parent-child retrieval can help with missing context. They cannot repair incorrectly extracted text.
  • Check retrieved evidence, context completeness, and answer correctness separately.

What is chunking, and why not use the entire file?

Chunking divides a document into pieces that can be indexed and retrieved. A simplified pipeline looks like this:

Document -> Extract and organize text -> Create chunks
         -> Store text and source references -> Build an index

Question -> Retrieve relevant chunks -> Assemble context -> LLM answer with sources
Enter fullscreen mode Exit fullscreen mode

Retrieval may use keyword search, vector search, or both. Not every chunk needs an embedding. When vector search is used, an embedding model creates a numerical representation of the content.

The searchable unit also does not have to be the final reading unit. A system can retrieve a short passage, then supply its surrounding section to the model.

A single chunk containing a 100-page handbook mixes many topics. That can make specific passages harder to retrieve and exceed model input limits. Sending unrelated material also costs tokens and can distract from the evidence that matters.

Whole-document context can work for short documents. Tiny chunks are not automatically better: an amount without its eligibility rule is incomplete evidence.

One boundary can change the answer

Consider this fictional passage:

Section 8.1: Prescription eyewear benefit
The company reimburses up to 3,000 baht per year.
Only permanent employees who have completed probation are eligible.
Enter fullscreen mode Exit fullscreen mode

A length-based split might produce:

Chunk A: The company reimburses up to 3,000 baht per year.
Chunk B: Only permanent employees who have completed probation are eligible.
Enter fullscreen mode Exit fullscreen mode

Without the heading, both pieces lose their subject. B never names the benefit, and vector search does not guarantee recovery of that missing relationship.

If only A reaches the LLM, it has evidence about the amount, not enough evidence about eligibility. A prompt telling it not to invent information may help it acknowledge uncertainty. It cannot restore a sentence that was never supplied.

Keeping the heading, amount, and condition together helps make evidence usable. It does not guarantee a correct answer.

The main approaches and their trade-offs

These approaches can be combined. Splitting by heading and then dividing oversized sections is often more useful than committing to one method everywhere.

1. Fixed-size splitting

Choose a length, such as 500 tokens, and divide the text into consecutive pieces. It is straightforward, but the boundary may cut through a sentence, table, or exception.

Use the actual model tokenizer when token limits matter. Character counts are not interchangeable with token counts, and a Thai word is not necessarily one token.

2. Fixed-size splitting with overlap

Overlap repeats some text from one chunk in the next, giving material near the boundary a chance to stay together.

It cannot fix distant conditions or scrambled reading order. Thai OCR with unreliable spacing still needs boundary checks.

More overlap means more indexed content and duplication. Near-identical results can occupy retrieval slots that should contain other evidence. Deduplicate or merge overlapping context before sending it to the LLM where appropriate.

3. Recursive splitting

A recursive splitter tries separators in a chosen order, such as paragraphs, sentences, then finer boundaries, until pieces fit the size limit.

It avoids unnecessary cuts through sentences, but does not understand meaning. A paragraph can cover several topics; damaged OCR can remove boundaries.

4. Structure-aware splitting

Use headings, lists, tables, and code blocks to guide boundaries. A subsection can carry its parent heading so a short passage still has a clear subject.

Tables need special care. An isolated “3,000” is not enough. Preserve the row name, column name, and unit. In our fictional example, “Prescription eyewear: annual reimbursement limit of 3,000 baht” is clearer. Any eligibility note must accompany it too.

Code and ordered requirements have similar dependencies. If an oversized block must be split, preserve its relationships and necessary conditions.

5. Semantic chunking

Semantic methods estimate where nearby sentences or paragraphs change topic, often using embeddings or a model.

They can help when structure is unclear, but add processing cost and tuning. Results depend on the model and source quality. If headings already separate topics well, I would test that simpler structure first.

6. Parent-child retrieval and contextual chunking

These solve different problems alongside the splitting methods above.

Parent-child retrieval indexes small child passages for search, then expands a match to a larger parent section for reading. Maintain the links, control context size, and check access permissions for the expanded parent too.

Contextual chunking adds useful context before indexing, such as a document title, section name, or short source-grounded explanation. If an LLM generates that explanation, check that it has not invented conditions. Keep original text available for citations instead of treating generated context as source evidence.

Propositional or agentic approaches may use an LLM to extract standalone statements or choose content groupings. Implementations vary. The shared trade-offs are extra cost and the risk that rewriting drops a qualification. Compare them with a non-LLM baseline and retain the original source.

Match the approach to the document

Document A reasonable first approach Check carefully
Markdown or structured manuals Split by heading, subdivide oversized sections Parent headings, tables, footnotes
Plain text with inconsistent structure Recursive splitting, limited overlap Conditions at boundaries
Contracts or rules requiring several paragraphs Retrieve small passages, expand or combine evidence Clause numbers, versions, permissions
Scanned or multi-column PDFs Repair extraction and reading order first OCR errors, scrambled columns, broken tables

These are starting points, not substitutes for inspecting actual documents.

The Thai tourism PDF that changed my priorities

While experimenting with a RAG pipeline, I used a beautifully designed tourism PDF about Thailand. The extracted text was much less attractive: Thai vowels and tone marks were misplaced, while colored backgrounds, watermarks, and illustrations interfered with reading.

For some pages, I tried black-and-white conversion to make the letters clearer. That could also lose visual context. I compared outputs with help from multiple models and had a person review spelling again.

I described this in my earlier PDF and Markdown write-up, in Thai. It is one experience, not evidence that every PDF needs the same workflow.

The lesson was simple: changing chunk size cannot bring back words that extraction has already lost.

Markdown makes structure easier to identify, not evaluation unnecessary. Sections can still be oversized, and table rows need headers.

How large should chunks and overlap be?

There is no universal setting. 300-500 tokens with 10-20% overlap can be an initial experiment for paragraph-based text. Those numbers are not a benchmark result from my system or a standard to apply everywhere.

Before tuning, check four things:

  1. Meaning: Does each piece identify the subject and preserve necessary conditions?
  2. Limits: Use the actual tokenizer and budget for instructions, the question, all retrieved context, and the answer.
  3. Duplication: Does overlap add useful evidence or just repeat it?
  4. Question type: A lookup for an amount and a question about eligibility may require different evidence.

Start with structure and a size cap. Add overlap or parent expansion when tests reveal missing context, not simply because more context feels safer.

How to tell whether chunking works

Build questions with known supporting documents, sections, or pages. Include boundary-sensitive questions, amount lookups, and questions the documents cannot answer.

Compare two or three approaches on the same sources and questions with comparable context budgets. Then inspect three layers:

  • Retrieval: Are the labeled supporting passages found? Recall@k measures retrieval of predefined relevant evidence in the top k results. Separately check whether the evidence available is sufficient for the question.
  • Context: Are conditions and exceptions present? Do duplicates crowd out other evidence? Are table headers and units intact?
  • Answer: Does it follow the evidence? Does each citation support the associated claim? Can the system acknowledge insufficient information?

Store metadata such as document_id, section, page/offset, and version so passages can be traced and stale index entries managed.

For organizational data, enforce permissions before content reaches the model, including any expanded parent context. Hiding the final answer is not a substitute for controlling what the model receives.

If source text is wrong, fix ingestion. If evidence exists but is not found, inspect chunking and retrieval. If complete evidence reaches the model but the answer is wrong, inspect context assembly and generation.

The goal is usable evidence

A RAG system generally supplies selected evidence, not the entire library, for each question. Chunk boundaries help determine what the model gets to see, though they are not the only factor in answer quality.

My starting point is accurate source text, structure-aware boundaries, sensible size limits, source references, and tests with real questions. Add complexity where those tests show a need.

The best chunk is not necessarily the smallest one. It is a piece the system can find and use without losing what makes the answer valid.

If you are building RAG, share which document types cause missing conditions and where in the pipeline you catch them.


Explore the NEXT4I journey and read the original article at: https://go.next4i.com/next4i-devto-en

Top comments (0)