DEV Community

clement-melkior
clement-melkior

Posted on

How to extract text from PDFs and scanned documents (and how to tell which pages need OCR)

"Extract the text from this PDF" hides two very different jobs:

  • A digital PDF was made on a computer. The text is inside the file and can be read exactly.
  • A scanned PDF is a stack of pictures. There is no text in it until you run OCR (optical character recognition).

They look identical on screen. Use the wrong tool and you get either nothing or slow, error-prone output for text that was sitting in the file all along. And plenty of real documents mix both: a typed contract with a scanned signature page, for example.

Step 1: check what you have

pdftotext comes with Poppler (apt install poppler-utils, or brew install poppler):

pdftotext report.pdf - | head -20
Enter fullscreen mode Exit fullscreen mode

If you see your text, the PDF is digital and you are done:

pdftotext report.pdf report.txt            # flowing text
pdftotext -layout report.pdf report.txt    # keeps columns and tables aligned
Enter fullscreen mode Exit fullscreen mode

If the output is empty or just a few stray characters, the pages are scans.

Step 2: OCR the scanned pages

Render the pages to images, then run Tesseract on each one:

apt install tesseract-ocr tesseract-ocr-fra    # add the languages you need

pdftoppm -r 200 -gray -png scan.pdf page
for f in page-*.png; do tesseract "$f" "${f%.png}" -l eng; done
cat page-*.txt > scan.txt
Enter fullscreen mode Exit fullscreen mode

Three things I learned the hard way:

  • 200 DPI is a good default. Below about 150, small print starts to break. Higher mostly costs time.
  • Tell Tesseract the language. -l fra or -l eng+fra matters a lot for accents. Don't add languages you don't need: each one slows it down.
  • In a container, set OMP_THREAD_LIMIT=1. Tesseract starts four threads by default. On a container with one CPU core they fight each other: in my tests the same dense page took 10.4 seconds by default and 3.8 seconds with the limit set.

Step 3: mixed documents

For a PDF with both kinds of pages, check each page separately. pdftotext separates pages with a form-feed character, so you can split its output, keep the pages that have text, and send only the empty ones to OCR. That keeps the exact text where it exists and avoids paying the OCR time for pages that don't need it.

When an API is easier

I needed this as a service I could call from scripts, so I packaged the steps above into a tool: PDF & Image to Text on Apify. It decides page by page whether to read the text layer or run OCR, and also takes images (PNG, JPEG, TIFF, WEBP), Google Drive and Dropbox share links, and uploads.

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

run = client.actor("spokentext/pdf-image-to-text-ocr").call(run_input={
    "urls": ["https://arxiv.org/pdf/1706.03762"],
    "languages": ["en"],
})

for doc in client.dataset(run.default_dataset_id).iterate_items():
    print(doc["fileName"], doc["pageCount"], "pages")
    for page in doc["pages"]:
        print(page["page"], page["method"], len(page["text"]))
Enter fullscreen mode Exit fullscreen mode

Each page comes back with a method field: text if it was read from the file, ocr if it was recognised. That 15-page paper is fully digital and came back in a few seconds.

It is priced per page: $0.001 for a digital page and $0.004 for an OCR page. Files are processed inside the run and are not sent to any outside AI service.

What neither route does well

  • Handwriting. Tesseract is built for printed text.
  • Tables as data. You get lines of text, not rows and columns. -layout keeps the alignment, which is often enough for an LLM to read.
  • Bad photos. Blurry, skewed or shadowed pictures of documents give poor results. Rescan if you can.
  • Password-protected PDFs. Remove the password first.

Summary

Your document Do this
Digital PDF pdftotext, exact and instant
Scanned PDF or image Render at 200 DPI, then Tesseract with the right language
Mixed, or many files Check page by page, or use the API

Disclosure: I built the tool described in "When an API is easier". This article was drafted with AI assistance and checked by me.

Top comments (1)

Collapse
 
clem616 profile image
clement-melkior •

Curious how others handle mixed PDFs (some typed pages, some scanned). Do you check page by page like I do here, or just run OCR on everything? And if anyone has a good approach for tables, I'd like to hear it. That's the part I'm least happy with.