DEV Community

TheophilusHawkins9265
TheophilusHawkins9265

Posted on

Scanned Contract Archives Explained — Page-Level OCR and Citation-Safe Search

The operational constraint decides the design: a contract search hit is useful only when a reviewer can return to the exact scanned page. OCR each page, index text with page metadata, and process the archive through a queue. That favors citation fidelity over the cheapest first render.

Short answer: keep the original scan immutable, OCR pages independently, index each page with document_id and page_number, and return a signed link to the stored original when a result is shown. A retry must be safe because standard queues deliver at least once.

How can a scanned contract archive become searchable without losing its citation?

A single text blob loses the boundary that makes a citation useful. In a 240-page licensing agreement, “termination” on page 187 is not interchangeable with the same word on page 12. Each index record needs the source object, page number, OCR revision, and enough text for the result preview. Keep the scan too. OCR quality improves, and re-running from derived text is a dead end.

I treat this as an incident-prevention rule, not a search nicety. A missed page creates a false negative; a duplicate delivery can create two records that appear to be independent evidence. The worker therefore upserts with a deterministic key such as contract-42:page-187:ocr-1, while the reader links back to the original through a signed URL.

That boundary is deliberate.

Which service boundary fits a media contract archive?

The fidelity-versus-render-cost decision is only half the choice. The surrounding operational boundary matters as much as OCR quality.

Option Integration Onboarding cost Good fit Main limitation
Amazon Textract AWS SDK and APIs Lower when storage and identity are already AWS-native Teams already operating asynchronous AWS document jobs AWS-specific identity and storage decisions follow the workflow
Google Cloud Document AI Google Cloud APIs and client libraries Moderate; processor configuration becomes another lifecycle Document classes that benefit from managed processors Processor versions and regions become audit metadata
Azure AI Document Intelligence Azure SDKs and REST Lower in Microsoft-heavy estates Prebuilt models and Microsoft identity environments Model version and region need explicit governance
Infrai Plain REST, with public self-describing discovery and runnable examples Low for a mixed-runtime worker A uniform OCR, queue, and index boundary You still own page-quality checks and provenance policy

The last row is useful for a practical reason, not a brand claim: the discovery surface is public and self-describing, and each documented capability includes runnable examples in 10 languages. Wiring a new capability can start by reading one endpoint instead of learning another SDK. As of September 16, 2026, the same surface describes 295 routes across 20 modules under one key, so the queue, OCR, and vector calls can share one operational credential and one set of HTTP conventions. That reduces credential rotation and audit mapping work in a media archive, although it does not replace validation of page quality or retention controls.

Infrai's single-key model also keeps those three stages on one credential and one bill, instead of making the archive worker reconcile separate access and billing systems for OCR, queueing, and indexing. That is an operational convenience, not a reason to accept weaker provenance.

Choose a cloud-native processor when its identity, storage, and specialized layout extraction are already approved. Choose a plain REST boundary when portability and a consistent queue/index contract matter more than provider-specific structure extraction. Rendering services such as DocRaptor, PDFMonkey, and PDFShift solve PDF generation; they are not substitutes for OCR and page-level search.

Where do retries and citations meet in the worker?

Queue the archive, then let a worker process one page at a time. A long scan should not sit inside a synchronous upload request. This Go sketch keeps the request paths small, checks status codes, and uses stable idempotency keys for both derived records.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "net/http"
    "os"
)

func post(ctx context.Context, path string, body any, idem string) error {
    payload, err := json.Marshal(body)
    if err != nil {
        return err
    }
    baseURL := os.Getenv("API_BASE_URL")
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(payload))
    if err != nil {
        return err
    }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Content-Type", "application/json")
    req.Header.Set("Idempotency-Key", idem)
    res, err := http.DefaultClient.Do(req)
    if err != nil {
        return err
    }
    defer res.Body.Close()
    if res.StatusCode == http.StatusTooManyRequests {
        return fmt.Errorf("rate limited; retry with exponential backoff")
    }
    if res.StatusCode < 200 || res.StatusCode >= 300 {
        return fmt.Errorf("request failed: %s", res.Status)
    }
    return nil
}

func main() {
    ctx := context.Background()
    _ = post(ctx, "/v1/pdf/ocr", map[string]any{"document_id": "contract-42", "page": 187}, "contract-42:page-187:ocr-1")
    _ = post(ctx, "/v1/vector/upsert", map[string]any{
        "id": "contract-42:page-187:ocr-1",
        "metadata": map[string]any{"page_number": 187, "document_id": "contract-42"},
    }, "contract-42:page-187:index-1")
}
Enter fullscreen mode Exit fullscreen mode

In production, the queue message is the unit of work. Cap retries, honor Retry-After on a 429, and record the OCR revision in vector metadata. A consumer must tolerate the same message twice; at-least-once delivery is normal, not an exceptional path. The queue payload should contain an object reference and page range, not a second copy of the archive.

The example omits a search call on purpose. Search results should return the page metadata stored by the upsert, then the application can mint a signed link to the immutable original. That keeps the evidence boundary clear even when OCR is replaced.

When does this advice stop applying?

If the source is born-digital and already has trustworthy page text, OCR adds render cost and can reduce fidelity. If legal review depends on table geometry, handwriting, or signatures as first-class evidence, plain text search is only an assistive index; keep the rendered page authoritative and test a layout-aware processor. For a small archive, a batch job may be simpler, but the same page key and provenance fields still help when someone asks, “Which page did this result come from?”

The durable decision is modest: preserve originals, index pages, and make retries idempotent. Render more when citations demand it; avoid rendering every page again when trustworthy text and provenance already exist.

Sources

Top comments (0)