DEV Community

OCRRank
OCRRank

Posted on Originally published at ocrrank.com

OCR vs document parsing: when Tesseract stops being enough

Most "we need OCR" tickets I see are not really OCR tickets. Someone wants the invoice total in the ERP, the transactions from a bank statement in a spreadsheet, or the renewal date from a contract in the CRM. That is a data problem. OCR is, at most, one step of it.

I run OCRRank, a comparison site for OCR and document-parsing tools, and getting this distinction right is the first thing I ask about before any tool question. Here is the framework I use, written for developers who have a Tesseract script that worked in the demo and are now wondering why production keeps breaking.

Two different jobs

OCR converts pixels into characters. You give it a scan or a phone photo, it gives you text.

Document parsing analyzes layout and content and returns structured data: named fields, tables, line items, as JSON, CSV, or Excel. On scans it usually runs OCR underneath. On digital PDFs it often doesn't need OCR at all.

OCR Document parsing
Input Image, scan, photo, image-only PDF Digital or scanned document
Output A text dump Fields, tables, structured JSON / Excel / CSV
Knows what "invoice total" means? No Yes, when modeled for that document type
Enough on its own for AP or bookkeeping? Rarely That's the point

A useful test: if your success metric is "we can Ctrl+F the PDF," OCR (or a PDF that already has a text layer) is probably enough. If the metric is "we can post or match this without a human retyping it," you have left OCR-only territory.

Step zero: does the PDF even need OCR?

A lot of teams run OCR on PDFs that already contain selectable text. That wastes money and can actually make things worse, because re-recognizing good characters can scramble them.

Check first. This is an illustrative sketch, not production code:

# Illustrative sketch: route pages by whether they have a text layer.
import pdfplumber
import pytesseract

def page_texts(path, min_chars=20):
    out = []
    with pdfplumber.open(path) as pdf:
        for i, page in enumerate(pdf.pages):
            text = page.extract_text() or ""
            if len(text.strip()) >= min_chars:
                out.append((i, "text-layer", text))
            else:
                # image-only page: OCR is the only way to get characters
                img = page.to_image(resolution=300).original
                out.append((i, "ocr", pytesseract.image_to_string(img)))
    return out
Enter fullscreen mode Exit fullscreen mode

Notice what this gives you: strings. Even when it works perfectly, you still have to write the part that figures out which string is the vendor, which numbers belong to which line item, and where page 2 of the table picks up. That second parser is where most of the real engineering time goes.

Also notice the branch. Mixed packets are common (page 1 born-digital, page 2 a phone photo of a signed exhibit), and once you have if scanned: tesseract else: pdfplumber in your codebase, you are maintaining two pipelines.

Write the JSON you want before you pick a tool

My favorite trick: put the JSON keys you need on a whiteboard before anyone says a vendor name.

If the keys are only text and pages, you are in OCR-land:

{
  "pages": [
    { "page": 1, "text": "ACME Supplies Ltd\nInvoice #10482\nDate 03/09/2026\n..." }
  ]
}
Enter fullscreen mode Exit fullscreen mode

If the keys look like this, you are shopping for parsing:

{
  "vendor": "ACME Supplies Ltd",
  "invoice_number": "10482",
  "invoice_date": "2026-09-03",
  "currency": "EUR",
  "line_items": [
    { "description": "A4 paper, 5 reams", "quantity": 4, "unit_price": 21.50, "amount": 86.00 },
    { "description": "Toner cartridge", "quantity": 1, "unit_price": 64.00, "amount": 64.00 }
  ],
  "subtotal": 150.00,
  "tax": 25.50,
  "total": 175.50
}
Enter fullscreen mode Exit fullscreen mode

(Both payloads are made up for illustration. They are not output from any specific tool.)

The second shape is what lets code post, match, or reject a document without a human in the loop. Totals-only output fails accounts payable, because you can't match a PO line or catch a wrong unit price from a grand total. The invoice data extraction guide goes deeper on which header and line fields AP teams actually need.

When Tesseract is genuinely enough

Open-source OCR is not "worse." It's the right tool when:

  • You need a searchable archive or a plain-text index
  • One digital layout you control (an internal report, one vendor you can re-test)
  • Offline or strict privacy requirements, and the engineering capacity to own the stack
  • You're prototyping before committing to a vendor
  • Humans will read the output and no code needs to post fields

If the problem fits in a weekend and the failure modes are yours alone, keep it DIY.

When it stops being enough

In my experience the tipping point is rarely a benchmark number. It's the third layout change this month, or the first scanned packet that silently drops a page. Signals that you've outgrown raw OCR plus glue:

  • Scans and phone photos show up in real production traffic
  • Many suppliers, banks, or counterparties, so per-template scripts keep breaking
  • You need line items, per-field confidence, webhooks, or a review queue
  • Someone is on call for "the extract is wrong," and that person is you

Cloud services like Azure Document Intelligence, Amazon Textract, and Google Document AI sit in the middle. They're capable building blocks, but you still own IAM, async jobs, retries, and mapping their output to your schema. I wrote up the full build-vs-buy trade-off in Tesseract vs document parsing APIs.

Decision checklist

  1. Can you select text in the PDF? Then start with parsing or table extraction and skip OCR.
  2. Image-only pages? You need OCR plus parsing, or one API that does both and returns the same schema either way.
  3. Only need searchable text? OCR alone may be fine. Stop here.
  4. Need fields, line items, or transaction rows? That's document parsing. Pick a path based on the document type.
  5. How many layouts next quarter? One stable layout leans DIY. Many or unknown leans managed.

What to test on your own documents

Vendor demo PDFs are clean by design. Bring your own, and run the same set through every option, including your current script:

  • Three files minimum: one clean digital PDF, one multi-page table, one mediocre scan or phone photo.
  • Page count in vs pages out. Silent page drops are one of the most common failures.
  • Math checks. Pick three line rows and check quantity times unit price equals the line amount. Check that lines plus tax roughly equal the total.
  • Spot-check 3 to 5 critical fields against the image: IDs, dates, totals.
  • Ask for the payload, not the demo viewer. Highlighted text on a pretty PDF preview is a UI feature. Your backend needs objects and arrays. If you're going scan-to-JSON, the PDF to JSON / OCR API guide lists what a good response should include.

Whichever path survives all three files with the least babysitting is your production default.

Where I landed

I rank ten managed parsing tools on OCRRank, scored on accuracy, developer experience, value, and trust, with accuracy weighted most. My current #1 for developers is DocuPipe (9.7/10), with Base64.ai (9.1) and Affinda (9.0) behind it, and Nanonets and Docsumo strong for enterprise IDP and financial documents. But the honest answer to "which tool?" is to answer "text or fields?" first. Buying a better OCR engine when you need fields is the expensive way to solve the wrong problem.

What does your current stack look like? I'd like to hear where the Tesseract-plus-glue approach broke for you, or where it's still holding up fine.


Disclosure: OCRRank is independent and earns referral fees from some of the tools it lists (currently every ranked tool has a referral relationship). Fees can affect placement on the site but not scores. This post was originally published on ocrrank.com.

Top comments (0)