DEV Community

Cover image for My PDF converter read two-column papers line by line — column gutters have to be detected, not assumed
InApp
InApp

Posted on Originally published at imapp.blogspot.com

My PDF converter read two-column papers line by line — column gutters have to be detected, not assumed

The PDF converter behind my document pipeline passed every test I threw at it — single-column whitepapers, invoices, contracts. Then an agent sent it a two-column conference paper and the Markdown read like this:

"Abstract We propose a In this paper we study new method for..."

Classic interleave. My first version grouped text elements by y-coordinate, reading the page top-down. Correct for one column; for two columns it stitches left-column line 1 to right-column line 1, all the way down. Perfect-looking Markdown, garbage reading order.

Root cause: a PDF isn't a document, it's a painting — a list of positioned glyphs with no semantic order baked in. Every reading order is a reconstruction, and I had assumed one column.

The fix, applied per page:

  1. Build a histogram of word x-positions and look for a vertical whitespace band of roughly 4% of page width with dense text on both sides — that's the gutter.
  2. Gutter found → split into left and right regions, extract each region top-to-bottom, left before right. No gutter → single-column pass.
  3. Report the column layout as a per-page field in the response, so downstream RAG chunkers know what they're chunking.

The step I nearly skipped was validation. I assembled about 30 two-column papers with reference text and scored extraction by sentence continuity — how often lines end with terminal punctuation instead of cutting mid-clause. Before the fix, most lines broke mid-sentence. After gutter detection, the score landed where real prose sits. It's still imperfect: tables and floating captions wander, sidebars misfire, footnotes come out of order. I'd rather surface those quirks than serve scrambled salad silently.

Lesson for anyone feeding PDFs into RAG: never trust content-stream order or a naive sort by y. Multi-column detection isn't a nice-to-have — a huge share of the papers your agents will read are two-column, and the failure is invisible because the output still looks like Markdown.

I ended up packaging the converter as [https://x402.freeq.one/tools/pdf_to_markdown.html] — headings, lists and tables come out clean, and two-column pages now get flagged instead of silently scrambled.

Top comments (0)