PDFs are documents, not strings.
When PDFs are converted to plain text, tables, reading order, formulas, figures and source locations can easily get lost — and that can hurt RAG pipelines.
I built Papero, an open-source PDF extraction tool that preserves document structure and exports to Markdown, JSON, Excel and Word.
It also provides structure-aware RAG chunks with section, page and bounding-box metadata.
→ GitHub: papero
I'd love feedback from people working with RAG, document AI and PDF processing.


Top comments (0)