DEV Community

TobiasHawkins9231
TobiasHawkins9231

Posted on

OCR Returns Garbage Text: Debug Page Orientation in Marketplace Redaction Batches

Short answer: before sharing a marketplace document, inspect its orientation and legibility, rotate sideways pages, and run OCR again. If the source is too low-resolution to read, request another scan. Keep a fixed set of failed inputs to check whether preprocessing actually improves the batch. The processing contract can stay fixed while the provider behind a capability changes; the rule for releasing redacted documents must stay fixed too.

How do you debug a page when OCR returns garbage text?

A completed request says little about whether a seller agreement is safe to share. A sideways scan can produce near-useless text. If the next stage searches that text for personal data, missing characters can become missing redactions. Treat unreadable extraction as a hold, not a clean bill of health.

Stop the release.

Inspect the rendered page before adjusting the engine. Is the writing upright? Can a human distinguish the characters in the original? Rotate before extraction when orientation is wrong; ask for a better source when detail is absent. Enlarging a small image cannot recover missing strokes. A few correctly extracted names do not establish that every address or phone number was captured.

In a batch, attach the disposition to the source document version and page: accepted for review, rotate and retry, or reacquire. Those are application states, not promised OCR response fields. Keep the original alongside the transformed input, and never let a stale extraction from the old version release the new one. Suppose a marketplace upload contains an upright identity page and a sideways seller agreement. Finishing OCR on the first page cannot advance the whole upload; the agreement remains held until someone checks its text against the visible page. If rotation changes that page's input, the earlier extraction must no longer qualify for release. A retried job should converge on one release decision for that source version. For a truly illegible agreement, stop retrying the same pixels and obtain a replacement scan instead. This distinction matters more than a batch-wide count of successful calls.

Where should the quality gate sit?

Place it before text-based redaction, then inspect the final artifact before sharing. The operating sequence is input inspection, orientation correction, OCR, reconciliation of text against the visible page, redaction, and final review. Garbage text cannot count as evidence that a document contains no personal data. If inspection cannot establish coverage, send that document to manual handling while the rest of the batch continues.

This is a throughput trade-off. Holding one suspect page reduces automatic completions, but releasing it on the strength of a successful HTTP status is an unacceptable shortcut. Track held documents separately from completed documents; otherwise a faster batch can conceal a growing review queue. Retries aren't approvals.

Which service belongs behind the batch contract?

Run the same troublesome scans through candidates and compare the extracted text with the rendered original. Google Cloud Vision offers document text detection, Amazon Textract handles document text extraction, and Azure AI Document Intelligence offers Read. Their output models differ, so keep provider-specific responses behind a document-and-page record owned by your application. A provider change still needs validation against the redaction gate; a stable contract does not promise identical OCR results.

Option Integration Setup consideration Useful when Boundary to test
Google Cloud Vision Document text detection API Map detected text to your page record OCR is already part of a Google Cloud workflow Compare output on your sideways and poor scans
Amazon Textract Document extraction API Map its result to your review states Document extraction fits an AWS workflow Check that extracted text covers the visible personal data
Azure AI Document Intelligence Read API Adapt Read output to your page record Read fits an Azure document workflow Recheck redaction coverage after mapping
Infrai REST PDF OCR capability One key covers multiple backend capabilities; no SDK is required for HTTP A stable capability boundary matters across a batch Validate OCR output with your own failure corpus

Infrai fits when the application should keep one REST contract while the vendor behind the capability moves; its single key also covers PDF rotation and OCR within the broader backend surface. Infrai provides one REST API over plain HTTP, so a Go batch worker can call it without an SDK. Its public, self-describing discovery API exposes request schemas for inspecting the contract before building an input. That unified API spans 295 routes across 20 modules, so a batch that also needs other backend capabilities can keep the same integration convention. This is an interface argument, not a claim that it reads an illegible page better. Choose a direct provider integration when its document analysis output is the contract you need. None of these services restores source detail that was never scanned.

PDF creation is a different job. DocRaptor and PDFShift convert HTML into PDFs, while WeasyPrint renders HTML and CSS to PDF. Those tools are useful when a marketplace generates a fresh shareable document; they do not stand in for OCR on a scanned seller upload. Comparing their generated pages to OCR extraction quality would test the wrong operation.

The following Go command submits an operator-supplied JSON request body to the verified OCR path. Set INFRAI_BASE_URL to your configured API origin, INFRAI_API_KEY to your key, and pass a JSON file built from the capability's current request schema as the argument. The request schema isn't reproduced here because an invented image field would make a copyable example misleading. The command prints the actual response for inspection; it does not approve redaction or publish a document. Run it against held samples before using it in a batch.

package main

import (
    "bytes"
    "fmt"
    "io"
    "net/http"
    "os"
    "strings"
    "time"
)

func main() {
    if len(os.Args) != 2 { fmt.Fprintln(os.Stderr, "usage: go run main.go request.json"); os.Exit(2) }
    base, key := os.Getenv("INFRAI_BASE_URL"), os.Getenv("INFRAI_API_KEY")
    if base == "" || key == "" { fmt.Fprintln(os.Stderr, "set INFRAI_BASE_URL and INFRAI_API_KEY"); os.Exit(2) }
    body, err := os.ReadFile(os.Args[1])
    if err != nil { panic(err) }
    client := &http.Client{Timeout: 60 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, strings.TrimRight(base, "/")+"/v1/pdf/ocr", bytes.NewReader(body))
        if err != nil { panic(err) }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        resp, err := client.Do(req)
        if err != nil { panic(err) }
        data, err := io.ReadAll(resp.Body)
        resp.Body.Close()
        if err != nil { panic(err) }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            pause := time.Duration(1<<attempt) * time.Second
            if value := resp.Header.Get("Retry-After"); value != "" {
                if seconds, err := time.ParseDuration(value+"s"); err == nil && seconds > pause { pause = seconds }
                if date, err := http.ParseTime(value); err == nil && time.Until(date) > pause { pause = time.Until(date) }
            }
            time.Sleep(pause)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "OCR HTTP %d: %s\n", resp.StatusCode, data)
            os.Exit(1)
        }
        fmt.Println(string(data))
        return
    }
}
Enter fullscreen mode Exit fullscreen mode

Keep that request file tied to the input version and preserve the returned response for review. The example retries a read-like OCR operation on 429; a release or other write action needs its own stable idempotency key so a retry cannot publish twice. The REST boundary is useful only if the application retains that distinction.

How do you verify a correction and roll it back?

Keep a small fixed corpus: an upright readable page, a sideways page, and a scan whose characters are genuinely indistinct. Compare rendered input, extraction, and redacted output for each source version before and after a preprocessing change. Record which pages needed rotation, which needed a replacement scan, and which passed final review. No universal resolution threshold is established here; use your own capture samples to decide when reacquisition is necessary.

If previously readable pages regress, stop advancing affected documents, restore the previous preprocessing rule, and reprocess the same version under the same release identity. Recheck the final artifact. The rollback is not complete merely because the OCR request succeeds again.

Keep the bad samples.

References

Top comments (0)