DEV Community

Cover image for My PDF converter interleaved a two-column paper line by line — a PDF has no reading order
InApp
InApp

Posted on Originally published at imapp.blogspot.com

My PDF converter interleaved a two-column paper line by line — a PDF has no reading order

A user's RAG pipeline started answering questions about an ML paper with answers that mashed together two unrelated sections. I pulled the source: a classic two-column conference paper. The extracted Markdown read like someone shuffled every other line — and it had, because it alternated between columns mid-sentence.

Here's what I hadn't fully internalized until then: a PDF has no paragraphs, no columns, no reading order. It's a list of positioned glyphs. Extraction order is whatever sequence the producer wrote the text operators into the content stream — and plenty of tools emit spans sorted by baseline y-coordinate across the full page width. On a two-column layout, that means line 37 of the left column and line 37 of the right column come out back to back.

The fix wasn't a better parser library, it was geometry:

  1. Collect every text span with its bounding box.
  2. Look for a persistent vertical gutter — an x-range where no span ever falls, page after page. Found one covering ~80% of pages? Treat the page as two columns.
  3. Sort spans within each column band by y, then x; slot full-width spans (titles, abstracts that span both columns) back in before the columns begin.
  4. Only then run paragraph detection and heading heuristics on sane, ordered text.

I built a small benchmark to keep myself honest: 30 two-column papers, 100 questions whose answers live in one specific column. Before the fix, retrieval pulled the wrong chunk about 4 times out of 10 — every word was present, just interleaved into confetti. After the fix it's closer to 1 in 10, and the remaining failures are mostly landscape-rotated tables, which is a separate war.

Same lesson applies to footnotes, sidebars and pull quotes. If your extracted text is fluent English that makes no sense, suspect order, not content.

I ended up packaging the layout-aware reordering into the PDF-to-Markdown API I run (https://x402.freeq.one/tools/pdf_to_markdown.html), mostly so my own ingestion jobs stop feeding the vector store confetti.

Top comments (0)