DEV Community

César Freitas
César Freitas

Posted on

Turn any PDF into clean Markdown for your RAG pipeline

Every RAG project hits the same wall on day two: the PDFs.

You have a folder of them — contracts, manuals, invoices, scanned reports — and you need the text inside. So you reach for a PDF library, run it, and get back a wall of characters with the tables flattened into gibberish, the headings gone, and the two-column pages interleaved line by line. Feed that to an embedding model and your retrieval quality is ruined before you even start.

Then you hit the scanned ones. No text layer at all. Now you're wiring up an OCR engine, tuning it, and gluing the two paths together.

I got tired of rebuilding this for every project, so I turned it into an API. Send a PDF, get back clean Markdown that keeps its structure.

What "clean" actually means

The goal isn't just extracting characters — it's preserving the shape of the document so a language model can use it:

  • Headings stay headings, so you can chunk on them.
  • Tables come back as Markdown tables, not scrambled rows.
  • Lists stay lists.
  • Reading order is respected on multi-column pages.
  • Scanned pages are OCR'd automatically, in the same call.

That last point matters: you send the file and you don't care whether it's digital or a photo of a page. The API figures it out.

The two endpoints

  • pdf-to-markdown — the workhorse. PDF in (as a file or a URL), clean Markdown out, plus the page count and whether OCR kicked in. This is what you chunk and embed.
  • extract — for the structured case. Point it at an invoice or a receipt and get back fields — totals, dates, vendor, line items — as JSON, so you skip the regex entirely.

Why an API instead of a local library

Two reasons. First, the "digital PDF plus scanned PDF plus tables plus reading order" combination is genuinely fiddly to get right, and it's the kind of plumbing that has nothing to do with your actual product. Second, it's deterministic and stateless: same file in, same Markdown out, in one request, with nothing to install or keep updated.

If your RAG retrieval is only as good as your ingestion — and it is — this is the cheapest place to buy a big quality jump.

It's live on RapidAPI with a free tier: https://rapidapi.com/cesaricf79/api/document-intelligence3

How are you handling PDF ingestion in your RAG stack right now? I'd love to hear what's working and what still hurts.

Top comments (0)