DEV Community

InApp
InApp

Posted on Originally published at imapp.blogspot.com

My PDF converter returned 200 OK and 214 characters for a 46-page contract — scanned pages have no text layer

A partner agent sent my converter a 46-page supplier agreement and complained the output was useless. The API returned 200 OK and exactly 214 characters for 46 pages: a header, a footer, some page numbers.

My first instinct was an extraction bug. It wasn't. I opened the file and every page was a 300 DPI image. The text was pixels. There was no text layer for my extractor to read, and my code had happily converted nothing into empty Markdown — with a success status code, which is the worst part.

I had built the whole pipeline assuming PDFs are documents. A large share of real-world PDFs that reach an agent — contracts scanned by humans, faxes, printouts of printouts — are photographs wearing a .pdf extension. Digital-born PDFs carry real text objects; scanned ones carry only images, sometimes with an invisible OCR layer stamped underneath, which is why some files let you copy text you can't see.

The fix, in three parts: first, count extractable characters per page and flag anything under ~30 chars as likely scanned. Second, fall back to rasterizing the page and running OCR — slower, but it beats returning confident emptiness. Third, return a per-page flag saying "no text layer found" so callers can distinguish a scanned page from a genuinely blank one. On my 500-file test corpus, plain text extraction handled about 76% of files; adding the OCR fallback took it to ~95%, at the cost of being 10-40x slower on scanned pages.

I packaged the fallback into my PDF-to-Markdown API (https://x402.freeq.one/tools/pdf_to_markdown.html), so scanned pages now come back flagged and readable instead of silently empty.

The takeaway for anyone feeding PDFs into a RAG pipeline: check characters-per-page before trusting the output. A 200 OK with an empty string looks exactly like success, and that's what makes it dangerous.

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The per-page flag is valuable because the 214 characters can make a document-level check pass even while most pages contain no usable text. I would treat the character threshold as a routing heuristic, then inspect both image coverage and text quality before deciding that a page is genuinely blank or needs OCR.

A damaged invisible OCR layer is a useful additional fixture: it may contain many extractable characters while producing nonsense or missing a critical table. Reporting which extraction path each page used, along with preserved page boundaries, would let a downstream RAG caller quarantine uncertain pages without discarding the readable parts of a mixed document.