A user's RAG pipeline started answering questions about an ML paper with answers that mashed together two unrelated sections. I pulled the source: a classic two-column conference paper. The extracted Markdown read like someone shuffled every other line — and it had, because it alternated between columns mid-sentence.
Here's what I hadn't fully internalized until then: a PDF has no paragraphs, no columns, no reading order. It's a list of positioned glyphs. Extraction order is whatever sequence the producer wrote the text operators into the content stream — and plenty of tools emit spans sorted by baseline y-coordinate across the full page width. On a two-column layout, that means line 37 of the left column and line 37 of the right column come out back to back.
The fix wasn't a better parser library, it was geometry:
- Collect every text span with its bounding box.
- Look for a persistent vertical gutter — an x-range where no span ever falls, page after page. Found one covering ~80% of pages? Treat the page as two columns.
- Sort spans within each column band by y, then x; slot full-width spans (titles, abstracts that span both columns) back in before the columns begin.
- Only then run paragraph detection and heading heuristics on sane, ordered text.
I built a small benchmark to keep myself honest: 30 two-column papers, 100 questions whose answers live in one specific column. Before the fix, retrieval pulled the wrong chunk about 4 times out of 10 — every word was present, just interleaved into confetti. After the fix it's closer to 1 in 10, and the remaining failures are mostly landscape-rotated tables, which is a separate war.
Same lesson applies to footnotes, sidebars and pull quotes. If your extracted text is fluent English that makes no sense, suspect order, not content.
I ended up packaging the layout-aware reordering into the PDF-to-Markdown API I run (https://x402.freeq.one/tools/pdf_to_markdown.html), mostly so my own ingestion jobs stop feeding the vector store confetti.
Top comments (0)