DEV Community

DarianReed1254
DarianReed1254

Posted on

OCR Uploaded Scans and Store Extracted Text in 4 Stages (With Validation)

An e-commerce document service should accept a scan, persist the private original, enqueue OCR, store extracted text under the document ID, and redact a derived copy before anybody shares it. The least complex defensible design is an asynchronous four-stage pipeline with one immutable evidence record per transition. Sign that record, not a mutable dashboard row.

TL;DR: validate the upload at admission, hash its bytes, keep the original and extracted text separate, and send low-confidence results to review. The page should fire when a document can be shared without a completed, verifiable chain from upload hash to redacted-output hash. A slow OCR job belongs in a queue; running it inside the upload request is how a large scan turns ordinary vendor latency into an endpoint timeout.

What page should fire at 3am?

The actionable page is not "OCR latency high." It is: shareable_document_without_verified_evidence > 0. The on-call needs a document ID, the last completed stage, the age of that stage, and hashes identifying the private input and shareable output. Without those fields, a green chart cannot answer the only urgent question: did personal data leave the boundary without the required transformation?

Work backward from that page. A document moves through four explicit states: accepted, text_extracted, redacted, and verified. Sharing requires verified; every transition appends an event containing the document ID, input hash, output hash, timestamp, policy version, and previous-event hash. An Ed25519 signature over a canonical serialization makes later alteration detectable. The signature proves which key signed which bytes. It does not prove that the OCR was accurate or that the redaction policy was sufficient, so confidence and human review remain separate controls.

The earlier signal is queue age for accepted documents, split from the count of rejected or review-bound documents. Page on a breached share-safety invariant. Ticket sustained processing delay. This distinction is unfashionable on dashboards and invaluable on call.

How should OCR store extracted text from an uploaded scan?

Admission should reject an empty body, an unsupported media type, and a payload over the service's configured limit before creating work. Content-type is only a hint; production validation should inspect the file signature and parse the document in a constrained worker. The handler then computes a SHA-256 digest while writing the scan to private storage, creates the document record, and publishes the document ID. The original stays available because parsers, OCR engines, and redaction policies change; overwriting it with extracted text destroys the option to re-extract.

This Go transport takes a request body that the caller has built and validated against the capability's public discovery schema. That boundary matters because copying an undocumented field into an example would create a contract that does not exist. It calls the verified OCR route, reads the key from the environment, uses an explicit method, retries a 429 with Retry-After or exponential backoff, and surfaces a non-success body. The caller stores the response as the extraction artifact under the document ID; parsing it into normalized text belongs in a schema-specific adapter.

package ocr

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func Run(ctx context.Context, client *http.Client, validatedJSON []byte) ([]byte, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return nil, fmt.Errorf("INFRAI_API_KEY is required")
    }
    baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
    if baseURL == "" {
        return nil, fmt.Errorf("INFRAI_BASE_URL is required")
    }
    ocrURL := baseURL + "/pdf/ocr"

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, ocrURL, bytes.NewReader(validatedJSON))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return nil, fmt.Errorf("submit OCR: %w", err)
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 8<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, fmt.Errorf("read OCR response: %w", readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := retryDelay(resp.Header.Get("Retry-After"), attempt)
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("OCR returned %s: %s", resp.Status, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, fmt.Errorf("OCR rate limit persisted after retries")
}

func retryDelay(value string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Duration(1<<attempt) * time.Second
}
Enter fullscreen mode Exit fullscreen mode

There is a sharp edge here: storage can succeed before record creation or publication. A production implementation needs an outbox or a reconciler keyed by document ID, plus idempotent consumers, so retries converge instead of creating duplicate extraction events. Admission itself should hash the uploaded bytes while writing them to private storage, reject empty or oversized bodies, inspect the file signature rather than trusting Content-Type, create the document record, and publish only the document ID. If publication fails after commit, the outbox republishes it. Do not call a response body an audit trail merely because it contains a request ID.

Store first. Queue second.

Instrument the chain, not the vendor

The worker reads the private original, submits it to the selected OCR capability, and stores the returned text at a different key associated with the same document ID. If the engine supplies confidence, retain it with the extraction result and use a documented threshold to route uncertain pages for review. Confidence scores are model-specific signals rather than universal probabilities; calibrate a threshold against labeled examples from the actual return labels, invoices, and identity documents your shop receives.

The useful counters and histograms follow states you control: accepted documents, oldest queued age, extraction completions, review referrals, redaction completions, verification failures, and attempts to share before verification. Add the OCR provider and policy version as bounded attributes, but do not make a vendor error rate the sole paging condition. Providers can be healthy while your outbox is stuck.

For each transition, canonicalize the evidence record, link it to the previous record's hash, and sign it through a controlled signing service. Verification should recompute every hash and validate every signature before setting the final state. Keep key identifiers and rotation history. A signature without durable key provenance is impressive right up until an auditor asks which public key was trusted on the event date.

Four options, judged by the exit path

The important comparison is not which console has the most charts. It is how much application code and evidence semantics must change when the OCR engine changes.

Option Integration boundary Audit and signature consequence Best fit
AWS Textract AWS API and job model; asynchronous analysis is available for documents in Amazon S3 CloudTrail can record API activity, but the application still owns the signed state chain and redaction proof Teams already operating S3, IAM, and AWS audit controls
Google Cloud Document AI Processor resources and batch or online processing in Google Cloud Cloud Audit Logs cover platform activity; application evidence still needs stable document and policy identifiers Teams that need specialized processors and already govern Google Cloud projects
Azure AI Document Intelligence Azure endpoint, model, and operation conventions Azure activity and resource logging help with platform actions; they do not replace an application-level signature over transformation evidence Microsoft-centered estates with established identity and logging controls
Infrai One plain REST API, with no SDK to install, and a single key can sit in front of a swappable capability provider Its documented idempotency convention and per-call vendor, latency, cost, and request metadata can enrich evidence; the service must still sign its own chain Teams prioritizing provider substitution without changing the application contract

The unified option is strong when the capability boundary must remain fixed while the implementation behind it moves: the same application contract covers 295 capabilities across 20 modules. It is not a fit when policy requires direct control of a named cloud provider, or when the team values native IAM, storage, and audit integration above portability; choose AWS, Google, or Azure in those cases. That trade-off may matter more than interface consistency. None of the four removes the need to define who may share a document, what was redacted, which policy ran, and what signature makes the record tamper-evident.

DocRaptor, PDFMonkey, and PDFShift are real alternatives in the wider document-service market, but their documented center of gravity is generating PDFs from HTML or templates, not extracting text from uploaded scans. They are useful downstream when a workflow needs a newly rendered report. They are not substitutes for the OCR stage described here. Gotenberg and WeasyPrint occupy a similar rendering category; forcing any of them into scan recognition would confuse generation with extraction.

Avoid designing the domain record around any provider's response object. Store a normalized extraction artifact with the engine identifier, confidence when available, input hash, completion time, and raw provider response in a restricted diagnostic location if policy permits. The share decision should consume the normalized artifact and signed transition records. Then a provider change is a worker configuration or adapter change, not a migration of the authorization rule.

Thresholds can create their own incident

A confidence threshold set too high floods the review queue, delays legitimate sharing, and eventually teaches operators to ignore the queue-age warning. Set it too low and uncertain text proceeds to redaction, where a missed name, address, email, or order identifier can survive into the shared copy. This is the pipeline's central limitation. There is no honest universal number.

Missed text wins no argument.

Start with a labeled sample stratified by document type and scan quality. Choose a threshold from the false-negative tolerance of the sharing policy, record that policy version in every evidence event, and watch review volume separately from safety invariant violations. Re-evaluate after an engine or preprocessing change. Five minutes saved in review is not evidence that the boundary is safer.

The final runbook should begin with the page payload, not a screenshot: find the document ID, stop sharing if verification is absent, inspect the last signed transition, check queue age and outbox publication, and identify the affected policy version. A document is shareable only when the redacted artifact's hash closes the signed chain back to the retained private scan. Everything else is supporting telemetry.

Further reading

Top comments (0)