You have a list of PDF links (reports, research papers, filings, manuals, price lists) and you need their text in a form an LLM, a search index or a spreadsheet can use. Copying from a PDF viewer scrambles two column layouts and breaks words across lines, and writing your own extraction script means dealing with downloads, redirects, files that are secretly HTML and documents with no text layer at all.
This guide shows how to turn a list of PDF URLs into clean text and Markdown with a small Apify Actor published by Hay Equipos called PDF Text Extractor: PDF URL to Text and Markdown per Page. It downloads each file and reads its text layer with pdf.js, the PDF engine used in Firefox.
What the tool returns
For each PDF you get either one row with the whole document or one row per page. Text comes out in reading order: two column layouts are read column by column, words split across lines are rejoined, larger fonts become Markdown headings, and bullet characters become lists. Each row also carries the document's metadata.
A document mode row looks like this (values are illustrative, text shortened):
{
"url": "https://www.example.org/reports/annual-report-2025.pdf",
"finalUrl": "https://www.example.org/reports/annual-report-2025.pdf",
"success": true,
"fileName": "annual-report-2025.pdf",
"fileSizeBytes": 2483112,
"pageCount": 48,
"pagesExtracted": 48,
"title": "Annual Report 2025",
"author": "Example Organization",
"createdAt": "2026-03-02T09:15:00.000Z",
"pdfVersion": "1.7",
"likelyScanned": false,
"wordCount": 21450,
"charCount": 131870,
"textTruncated": false,
"text": "Annual Report 2025\n\nLetter from the board...",
"markdown": "## Annual Report 2025\n\nLetter from the board..."
}
Page mode adds pageNumber and gives the text, Markdown and counts for that page only, while the document fields repeat on every row so each page stands alone. That makes it easy to chunk long documents and keep page numbers for citations.
Files that fail come back with success: false and an error, such as "The server answered HTTP 404", "The URL returned a web page, not a PDF" or "The PDF is password protected".
Step by step in the Apify Console
- Open the Actor from its Apify Store page and sign in to Apify Console.
- In the Input tab, add direct links under PDF URLs, one per line.
- Choose One row per: Document (whole text in one row) or Page.
- Keep Include plain text and Include Markdown on, or turn one off.
- Set Maximum pages per PDF (300 by default) and Maximum file size (50 MB by default). Larger files are skipped for free.
- Optionally set Maximum text length to cut text per row (0 means no limit).
- Click Start, then export from the Output tab as JSON, CSV or Excel.
If your PDFs are on your own computer, upload them to an Apify key value store, or anywhere with a public link, and pass those links.
How to call it from code
With curl:
curl -X POST "https://api.apify.com/v2/acts/pistachio_implementation~pdf-text-extractor/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls": ["https://arxiv.org/pdf/1706.03762"], "outputMode": "pages", "maxPagesPerPdf": 50}'
In Python, with the apify-client package, writing one Markdown file per page:
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("pistachio_implementation/pdf-text-extractor").call(
run_input={"urls": ["https://arxiv.org/pdf/1706.03762"], "outputMode": "pages"}
)
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
if row.get("success"):
with open(f"page_{row['pageNumber']:03}.md", "w", encoding="utf-8") as f:
f.write(row["markdown"])
Pricing
Pay per event: $0.002 per PDF processed, which is $2 per 1,000 PDFs, the same in document and page mode. Apify's standard Actor start event is also set, at $0.00005 per run. There is no charge for platform usage on top. Downloads that fail, files that are not PDFs, password protected files, robots.txt skips and files over your size limit are free. You can set a maximum charge per run in Apify, and the Actor stops cleanly when it is reached.
Limits and what it does not do
-
No OCR. Scanned PDFs without a text layer return little or no text and are flagged with
likelyScanned: true(under about 40 characters of text per page). Those need an OCR tool. - Tables come out as text lines in reading order, not as structured cells.
- Complex layouts with three or more columns, or text boxes scattered over the page, may read in an imperfect order.
- The Markdown is for reading and chunking. Headings are detected from font size and bullets from bullet characters, not from a pixel perfect layout.
- Links that need a login, a cookie banner click or a captcha before the file downloads return an error row.
- It respects robots.txt by default and spaces downloads from one host one second apart.
- Default memory is 512 MB, enough for typical files up to about 50 MB. Give the run more memory for very large files.
- Up to 2,000 PDFs per run.
Make sure you have the right to process the documents you send, and use them in line with each source site's terms.
Try it on the Apify Store: https://apify.com/pistachio_implementation/pdf-text-extractor
Top comments (0)