"Extract the text from this PDF" hides two very different jobs:
- A digital PDF was made on a computer. The text is inside the file and can be read exactly.
- A scanned PDF is a stack of pictures. There is no text in it until you run OCR (optical character recognition).
They look identical on screen. Use the wrong tool and you get either nothing or slow, error-prone output for text that was sitting in the file all along. And plenty of real documents mix both: a typed contract with a scanned signature page, for example.
Step 1: check what you have
pdftotext comes with Poppler (apt install poppler-utils, or brew install poppler):
pdftotext report.pdf - | head -20
If you see your text, the PDF is digital and you are done:
pdftotext report.pdf report.txt # flowing text
pdftotext -layout report.pdf report.txt # keeps columns and tables aligned
If the output is empty or just a few stray characters, the pages are scans.
Step 2: OCR the scanned pages
Render the pages to images, then run Tesseract on each one:
apt install tesseract-ocr tesseract-ocr-fra # add the languages you need
pdftoppm -r 200 -gray -png scan.pdf page
for f in page-*.png; do tesseract "$f" "${f%.png}" -l eng; done
cat page-*.txt > scan.txt
Three things I learned the hard way:
- 200 DPI is a good default. Below about 150, small print starts to break. Higher mostly costs time.
-
Tell Tesseract the language.
-l fraor-l eng+framatters a lot for accents. Don't add languages you don't need: each one slows it down. -
In a container, set
OMP_THREAD_LIMIT=1. Tesseract starts four threads by default. On a container with one CPU core they fight each other: in my tests the same dense page took 10.4 seconds by default and 3.8 seconds with the limit set.
Step 3: mixed documents
For a PDF with both kinds of pages, check each page separately. pdftotext separates pages with a form-feed character, so you can split its output, keep the pages that have text, and send only the empty ones to OCR. That keeps the exact text where it exists and avoids paying the OCR time for pages that don't need it.
When an API is easier
I needed this as a service I could call from scripts, so I packaged the steps above into a tool: PDF & Image to Text on Apify. It decides page by page whether to read the text layer or run OCR, and also takes images (PNG, JPEG, TIFF, WEBP), Google Drive and Dropbox share links, and uploads.
from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("spokentext/pdf-image-to-text-ocr").call(run_input={
"urls": ["https://arxiv.org/pdf/1706.03762"],
"languages": ["en"],
})
for doc in client.dataset(run.default_dataset_id).iterate_items():
print(doc["fileName"], doc["pageCount"], "pages")
for page in doc["pages"]:
print(page["page"], page["method"], len(page["text"]))
Each page comes back with a method field: text if it was read from the file, ocr if it was recognised. That 15-page paper is fully digital and came back in a few seconds.
It is priced per page: $0.001 for a digital page and $0.004 for an OCR page. Files are processed inside the run and are not sent to any outside AI service.
What neither route does well
- Handwriting. Tesseract is built for printed text.
-
Tables as data. You get lines of text, not rows and columns.
-layoutkeeps the alignment, which is often enough for an LLM to read. - Bad photos. Blurry, skewed or shadowed pictures of documents give poor results. Rescan if you can.
- Password-protected PDFs. Remove the password first.
Summary
| Your document | Do this |
|---|---|
| Digital PDF |
pdftotext, exact and instant |
| Scanned PDF or image | Render at 200 DPI, then Tesseract with the right language |
| Mixed, or many files | Check page by page, or use the API |
Disclosure: I built the tool described in "When an API is easier". This article was drafted with AI assistance and checked by me.
Top comments (1)
Curious how others handle mixed PDFs (some typed pages, some scanned). Do you check page by page like I do here, or just run OCR on everything? And if anyone has a good approach for tables, I'd like to hear it. That's the part I'm least happy with.