DEV Community

Turn PDFs into clean Markdown chunks for your RAG pipeline (without writing a parser)

PDF parsing is the boring part of every RAG project. Line breaks in the middle of sentences, lost headings, headers and footers mixed into the text, no page numbers to cite. You can spend days tuning pypdf or pdfplumber, or you can skip that part.

Here's a setup-free way to get LLM-ready text from PDFs, including PDFs you haven't found yet.

Step by step

  1. Open PDF Text Extractor on Apify (free plan credits are enough to try it).
  2. Paste PDF URLs, or a web page URL: with "find PDFs on pages" on, every PDF linked from that page is discovered and extracted. Handy for annual reports, documentation portals, research listings or government sites.
  3. Pick what you need:
    • plain text
    • Markdown with headings and bullet lists preserved
    • text per page (for citations like "see page 12")
    • RAG chunks: set a chunk size and overlap and get text split on paragraph and sentence boundaries
  4. Run it and pull the dataset into your pipeline.

Straight into your pipeline

from apify_client import ApifyClient

client = ApifyClient("<YOUR_API_TOKEN>")
run = client.actor("digitalni.produkty.pro.zivot/pdf-text-extractor").call(
    run_input={
        "urls": ["https://arxiv.org/pdf/1706.03762"],
        "includeMarkdown": True,
        "chunkSize": 1000,
        "chunkOverlap": 100,
    }
)
for pdf in client.dataset(run["defaultDatasetId"]).iterate_items():
    for chunk in pdf.get("chunks", []):
        ...  # embed and upsert into your vector DB
Enter fullscreen mode Exit fullscreen mode

It also works as a tool for AI agents through the Apify MCP server, so Claude or ChatGPT can read PDFs on demand.

Numbers

A test batch of 5 PDFs with 131 pages and ~78k words was processed in about 4 seconds. Price: $0.003 per PDF, any number of pages. Scanned (image-only) PDFs are detected and skipped, and you are not charged for them.


Disclosure: I built this tool. Feature requests welcome in the Issues tab.

Top comments (0)