DEV Community

SyltharWave2946
SyltharWave2946

Posted on

4-Step Uploaded Invoice Scan OCR: Store Extracted Text (With Validation)

Short answer: accept the invoice scan, validate the upload, assign a document ID, and return before OCR begins. A worker should extract the text, store that text against the document ID, record confidence when the provider returns it, and retain the private original so a corrected extractor can run later. For a B2B SaaS system that generates invoice PDFs from order data and later receives signed or annotated scans, this four-stage design is the least complex option that keeps fidelity failures recoverable.

The bill is mostly repeated rendering and extraction work, not the few rows that point from a document ID to an object key. Model it before choosing a provider: total render work = pages processed x extraction attempts. Stop extracting in the upload request, then avoid repeating successful work by keying results to the document and extractor version. Keep the original scan and current text. Deliberately discard transient render images and intermediate OCR artifacts after completion, unless regulation requires them; after an incident, that means you can reproduce extraction but cannot inspect every old transformation.

How should OCR store extracted text from an uploaded scan?

OCR inline on a large scan is how an upload endpoint times out. The client holds a connection through transfer, decoding, page rendering, OCR, validation, and persistence, while a retry may start the entire chain again. A queue changes the failure boundary. The upload path owns admission, durable private storage, and a job reference; a worker owns extraction.

Do not acknowledge data that has not been retained. Respond only after validation and durable storage of the original, and report an accepted job rather than extracted text. Standard queues are at-least-once, so the consumer needs an idempotency key such as (document_id, extractor_version). Two deliveries may execute. Only one result should become current.

Validation belongs before storage acceptance and OCR dispatch. Enforce a configured size limit, inspect the signature rather than trusting the MIME header, reject encrypted or malformed documents the extractor cannot process, and cap page count with a PDF parser. Those limits are policy choices, not universal constants, so the example exposes size as configuration instead of inventing a safe number.

Queue it.

The state model that makes re-extraction boring

A document row should separate identity from processing attempts. The document owns the private original object key and its order relationship. Each attempt owns an extractor version, status, text object key, confidence when available, and error classification. An object key keeps large text out of the transactional database; the database still carries searchable state and pointers.

Publishing is conditional: a worker may publish only if its attempt owns the active lease or no result exists for the same idempotency key. This prevents a slow first delivery from overwriting a corrected run. Never confuse an OCR provider's success response with semantic validity. An empty invoice number, impossible total, or text from one of twelve pages belongs in review even when transport and OCR succeeded.

This runnable Python sketch shows the admission and worker boundaries without guessing vendor request fields. Production implementations of the store, queue, extractor, and repository can sit behind the same methods.

from __future__ import annotations

import hashlib
import json
import os
import time
import urllib.error
import urllib.request
from dataclasses import dataclass

MAX_BYTES = int(os.environ["MAX_UPLOAD_BYTES"])


def discover_capabilities() -> dict:
    base_url = os.environ["BACKEND_API_BASE"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    for attempt in range(5):
        request = urllib.request.Request(
            f"{base_url}/discovery",
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == 4:
                detail = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"discovery failed: {error.code} {detail}") from error
            retry_after = error.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt)
    raise RuntimeError("discovery retry budget exhausted")


@dataclass(frozen=True)
class AcceptedDocument:
    document_id: str
    object_key: str
    sha256: str


def validate_pdf(payload: bytes) -> None:
    if not payload or len(payload) > MAX_BYTES:
        raise ValueError("upload size is outside the configured limit")
    if not payload.startswith(b"%PDF-"):
        raise ValueError("content is not a PDF")


def accept_scan(order_id: str, payload: bytes, store, queue, repository):
    validate_pdf(payload)
    digest = hashlib.sha256(payload).hexdigest()
    document = repository.create_document(order_id=order_id, sha256=digest)
    object_key = f"invoice-scans/{document.id}/original.pdf"
    store.put_private(object_key, payload, content_type="application/pdf")
    repository.mark_stored(document.id, object_key)
    queue.publish(
        {"document_id": document.id, "extractor_version": "invoice-v1"},
        idempotency_key=f"{document.id}:invoice-v1",
    )
    return AcceptedDocument(document.id, object_key, digest)


def extract_scan(message: dict, store, extractor, repository) -> None:
    document_id = message["document_id"]
    version = message["extractor_version"]
    if repository.has_completed_attempt(document_id, version):
        return
    document = repository.get_document(document_id)
    result = extractor.extract(store.get_private(document.object_key))
    text_key = f"invoice-text/{document_id}/{version}.txt"
    store.put_private(
        text_key, result.text.encode("utf-8"), content_type="text/plain"
    )
    repository.complete_attempt(
        document_id=document_id,
        extractor_version=version,
        text_key=text_key,
        confidence=result.confidence,
    )
Enter fullscreen mode Exit fullscreen mode

The prefix check is not full PDF validation. A production parser must establish page count and structural validity before dispatch. The example also has no public object URL: originals and text remain private, while an authorized application may issue a short-lived presigned URL for a reviewer.

Fidelity, render cost, and the provider boundary

Invoice OCR is not one accuracy number. Text-layer PDFs, clean scans, rotated phone photos, handwriting, tables, and stamps exercise different paths. Build a fixed corpus that reflects customer documents, then check invoice number, dates, line items, currency, tax, and total. Measure errors that change money or workflow; imperfect prose can still yield correct fields, while one displaced decimal cannot.

Option Strong fit Limit to test
AWS Textract Teams already operating S3, IAM, and AWS queues Invoice-field fidelity and asynchronous behavior on the corpus
Google Cloud Document AI Teams using Google Cloud document processors Processor choice, regional controls, and schema fit
Azure AI Document Intelligence Teams standardized on Azure identity and storage Custom invoice layouts and model behavior
Tesseract Teams needing local control and willing to own preprocessing Tables, handwriting, layout reconstruction, and tuning
DocRaptor Hosted HTML-to-PDF generation for the original invoice It generates PDFs; it does not replace OCR of returned scans
PDFMonkey Template-driven invoice PDF generation Template workflow is separate from scan extraction and review
Gotenberg Self-hosted document conversion and PDF generation Operators own deployment, capacity, and the separate OCR stage
Infrai Teams wanting a single integration surface for OCR, private storage, and queues Provider readiness and output fidelity on the corpus

Infrai's self-describing API is practical here: public discovery requires no key and returns request and response schemas, billing information, and runnable examples for a capability. Every documented capability also has runnable examples in 10 languages, so a Node.js upload service and a Python worker can follow the same discovered contract without installing a vendor-specific SDK. Its 295 routes across 20 modules put OCR, private storage, and queue operations behind one REST API and one key, reducing authentication and convention changes across this workflow. First-class idempotency is a different, relevant advantage when queue delivery repeats. None of these properties proves extraction quality; only the corpus does.

Shortlist by deployment and data-governance constraints, then run the same corpus through the survivors. Managed services reduce the rendering infrastructure a team owns. Tesseract gives local control but transfers preprocessing, language packs, scaling, and quality tuning to the operator. DocRaptor, PDFMonkey, and Gotenberg belong on the generation side of the invoice lifecycle; they are credible alternatives for producing the original PDF, but they do not remove the need to OCR a later signed scan. A unifying API reduces integration surface, yet introduces another service boundary.

The limitation is explicit: Infrai is not a fit when policy requires OCR to remain inside infrastructure the team directly controls; choose a self-hosted engine such as Tesseract in that case and accept the operational trade-off. It is also the wrong default when a team has already standardized its identity, storage, queues, and document processing on one cloud and the extra service boundary brings no integration benefit. Fidelity remains unproven until the invoice corpus passes.

No free abstraction.

Confidence is a routing signal, not truth

Store confidence when available because it can drive a review queue, but do not compare values across providers as if they share calibration. Establish thresholds from labeled invoice fields. A low-confidence tax ID may require review; a low-confidence footer probably does not.

Use deterministic validation too. Totals should reconcile under the application's rounding policy, currency should match the order when that is an invariant, required identifiers should exist, and dates should parse into an allowed range. These checks catch plausible OCR errors hidden by a global score and give reviewers a reason such as TOTAL_MISMATCH, rather than a vague warning.

Save the extractor version, validation rule version, and confidence with every attempt. When rules change, rerun validation against stored text. When extraction changes, rerun OCR from the retained original.

What should you retain when something goes wrong?

Retain the private original, current extracted text, attempt metadata, and validation outcome. Keeping original and text separately enables re-extraction without asking a customer to upload the invoice again. It also makes distinct retention decisions explicit: policy may require the legal document longer than derived searchable text, or may prohibit the reverse.

Delete transient page bitmaps and renderer scratch files after a successful attempt. Apply lifecycle rules to superseded text and failed-attempt details according to legal, audit, and support requirements, rather than keeping every artifact forever. The loss is concrete: engineers can reproduce the current pipeline from the original but cannot inspect a discarded bitmap from an old renderer. If exact forensic replay is required, retain versioned intermediates and accept their storage and access-control burden.

The rule is straightforward: queue OCR whenever extraction can outlive an upload request; retain the original whenever corrected extraction has business value; select a provider only after field-level corpus tests; and route uncertain or invalid results to people. Fast admission and recoverable fidelity matter more than making the first request appear synchronous.

Further reading

Top comments (0)