DEV Community

Beatriz Almeida Felicio
Beatriz Almeida Felicio

Posted on

Why PDF Extraction Breaks RAG — And What I Built to Fix It

PDFs are documents, not strings.

When PDFs are converted to plain text, tables, reading order, formulas, figures and source locations can easily get lost — and that can hurt RAG pipelines.

I built Papero, an open-source PDF extraction tool that preserves document structure and exports to Markdown, JSON, Excel and Word.

It also provides structure-aware RAG chunks with section, page and bounding-box metadata.

→ GitHub: papero

I'd love feedback from people working with RAG, document AI and PDF processing.

Top comments (0)