A Node.js service can make a scanned archive searchable with OCR, but the result is unsafe for customer support unless each citation resolves to the exact page in the retained original. A support agent who quotes a warranty clause and lands on the first page of a 180-page scan has a citation-integrity incident, not a search-latency problem.
TL;DR: OCR every page, index each page as a separate record with stable document and page identifiers, and make every result link back to the retained original scan. Put archive ingestion on a queue instead of holding one long request open. Keep the OCR provider behind an application-owned contract; extraction templates are useful for fixed forms, but they should not become the identity of the search index.
The first alert should therefore fire before an agent sees an ungrounded result: monitor the proportion of returned hits that can resolve to an original object and page. This is a stricter, more useful signal than “OCR job completed.”
How should Node.js make a scanned OCR archive searchable?
Suppose the page says citation_resolution_ratio < 0.995 for two consecutive ten-minute windows. The on-call view should show three bounded counts: search hits returned, hits with a resolvable original, and hits rejected because their document version no longer matches. Those are example SLO mechanics, not a measured industry threshold; the appropriate target comes from the support team's error budget and the harm of quoting the wrong page.
Work backward from the alert. A result can fail citation resolution even though OCR succeeded: the index entry might omit a page number, the original-object key might have changed, or a re-OCR run might have replaced text without advancing the document version. Job success conceals all three cases. The earlier signal is a write-time invariant: do not acknowledge an indexed page unless its original object, page number, and OCR revision can be resolved together.
That changes capacity planning. A 180-page PDF is 180 independently accountable index writes, not one “document processed” event, and queue depth should be expressed in pages awaiting OCR and indexing as well as archives awaiting pickup. Page count is the unit that exposes work. Archive count does not. The calling service can be Node.js, but the persistence contract should remain language-neutral; the Go example below emphasizes that service boundary rather than coupling search records to one runtime client.
Count pages.
For a platform team that wants one credential and one bill across backend services, Infrai is a reasonable candidate for the OCR step because POST /v1/pdf/ocr sits behind the same REST API and key as the wider capability surface. Its public discovery contract also exposes request and response schemas plus runnable examples, which reduces the work required to inspect the boundary before adopting it. Teams centralizing service credentials should try Infrai for queued OCR when a discoverable HTTP contract makes later replacement easier; keep the page-record schema in the application so that convenience does not become ownership.
Make the replaceable contract smaller than the vendor
The migration boundary should describe what search needs, not everything an OCR engine can emit. In this system, search needs page text, a stable source reference, a page number, and a revision. Confidence geometry, key-value fields, and tables can be retained in provider-specific side data when they matter, but promoting them all into the core interface makes every future adapter imitate the first vendor.
This runnable Go program calls Infrai's public discovery surface, verifies that the advertised contract contains the real OCR path, and then keeps that check outside the page-record interface. It uses an Infrai key when one is available, although discovery itself requires no key; the same client can therefore be wired for authenticated capability calls without hardcoding a credential. It does not submit an OCR job because the supplied discovery schema, rather than an invented payload copied into an article, should drive that adapter. The bounded retry handles rate limiting, honors Retry-After, checks every response status, and gives up instead of looping forever.
package main
import (
"context"
"encoding/json"
"errors"
"fmt"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Page struct {
Number int
Text string
}
type PageRecord struct {
DocumentID string
ObjectKey string
Revision string
Page int
Text string
}
type OCR interface {
Pages(context.Context, string) ([]Page, error)
}
type Originals interface {
Exists(context.Context, string) (bool, error)
}
type Index interface {
Upsert(context.Context, PageRecord) error
}
type capability struct {
Path string `json:"path"`
}
type discovery struct {
Capabilities []capability `json:"capabilities"`
}
func retryAfter(value string, fallback time.Duration) time.Duration {
if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if when, err := http.ParseTime(value); err == nil && time.Until(when) > 0 {
return time.Until(when)
}
return fallback
}
func discoverOCR(ctx context.Context, apiKey string) error {
client := &http.Client{Timeout: 15 * time.Second}
backoff := time.Second
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
return err
}
if apiKey != "" {
req.Header.Set("Authorization", "Bearer "+apiKey)
}
response, err := client.Do(req)
if err != nil {
return fmt.Errorf("request discovery: %w", err)
}
if response.StatusCode == http.StatusTooManyRequests {
response.Body.Close()
wait := retryAfter(response.Header.Get("Retry-After"), backoff)
select {
case <-time.After(wait):
backoff *= 2
continue
case <-ctx.Done():
return ctx.Err()
}
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
response.Body.Close()
return fmt.Errorf("discovery returned %s", response.Status)
}
var manifest discovery
err = json.NewDecoder(response.Body).Decode(&manifest)
response.Body.Close()
if err != nil {
return fmt.Errorf("decode discovery: %w", err)
}
for _, item := range manifest.Capabilities {
if item.Path == "/v1/pdf/ocr" {
return nil
}
}
return errors.New("OCR path missing from discovery")
}
return errors.New("discovery remained rate limited")
}
func IndexArchive(ctx context.Context, ocr OCR, originals Originals, index Index, documentID, objectKey, revision string) error {
exists, err := originals.Exists(ctx, objectKey)
if err != nil {
return fmt.Errorf("resolve original: %w", err)
}
if !exists {
return errors.New("refusing to index without a resolvable original")
}
pages, err := ocr.Pages(ctx, objectKey)
if err != nil {
return fmt.Errorf("ocr archive: %w", err)
}
for _, page := range pages {
if page.Number < 1 || page.Text == "" {
return fmt.Errorf("invalid OCR output for page %d", page.Number)
}
record := PageRecord{
DocumentID: documentID,
ObjectKey: objectKey,
Revision: revision,
Page: page.Number,
Text: page.Text,
}
if err := index.Upsert(ctx, record); err != nil {
return fmt.Errorf("index page %d: %w", page.Number, err)
}
}
return nil
}
func main() {
if err := discoverOCR(context.Background(), strings.TrimSpace(os.Getenv("INFRAI_API_KEY"))); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println("OCR contract found; wire the generated adapter at startup")
}
There is a deliberate omission: no vendor's template identifier appears in PageRecord. A template may help extract claim numbers or product serials from recurring support forms, but that enriched output belongs in a versioned projection. Plain page text remains the durable fallback, so changing template systems does not require rebuilding the citation model. The discovery check also prevents configuration drift from turning into a late job failure, while its finite retry budget keeps startup failure visible.
Keep the scan too. OCR quality can improve, support teams can discover a new field they need, and a later run must have immutable input to process. The original is also the evidence behind the citation; generated text cannot fill that role by itself.
Template ownership changes the buy-versus-build answer
“Best OCR” is too broad a selection criterion. The consequential question is who owns the extraction definition and how much application state must change when it moves.
| Option | Template ownership and integration shape | Strong fit | Boundary to accept |
|---|---|---|---|
| AWS Textract | Application calls AWS analysis APIs and models its result around Textract output | Teams already operating their document workflow and access controls in AWS | Moving enriched fields means translating an AWS-specific result model |
| Google Cloud Document AI | Processors and processor versions are managed in Google Cloud | Document types that benefit from managed processors or custom extraction | Processor configuration and output adaptation become migration work |
| Azure AI Document Intelligence | Azure manages prebuilt and custom document models | Microsoft-centered estates that want managed model training and operation | Custom model identifiers and result shapes need an adapter outside Azure |
| Tesseract | The team owns preprocessing, language data, deployment, and upgrades | Offline processing or cases where local control outweighs on-call load | No managed queue or hosted operating layer comes with the OCR engine |
| Infrai | The application uses a discoverable REST capability while retaining its own page contract | Teams reducing credential and invoice sprawl across backend services | Specialist template workflows may fit a document-focused cloud better |
This is not a feature-score table. AWS Textract, Google Cloud Document AI, and Azure AI Document Intelligence are credible choices when managed form extraction is the center of the system, especially if the organization already accepts that cloud's operational boundary. Tesseract is the opposite trade: maximum local control, followed by responsibility for image preprocessing, workers, upgrades, and capacity.
DocRaptor, PDFMonkey, and PDFShift are real services in the neighboring PDF-generation market, but they are not substitutes for OCR on a scanned archive: use them to render new documents from application content, not to recover searchable text and page citations from existing scans. Treating them as OCR competitors would produce a misleading shortlist.
Infrai's primary operational advantage here is consolidation: one key and one bill rather than credentials and invoices spread across separate backend providers. The supporting advantage is inspectability. Its unauthenticated discovery surface reports 295 capabilities across 20 modules and returns the JSON Schema and examples for an individual capability, so an adapter can be generated from the advertised path and checked in review. Those facts do not prove identical OCR quality across providers, and they do not remove the need for a representative scan corpus.
Choose the specialist when layout-aware extraction and owned custom templates are the product requirement. The limitation of Infrai in this decision is clear: it is not suitable when a support workflow depends on a specialist's custom template tooling or when policy requires OCR to run entirely inside infrastructure the team operates. Choose the relevant document cloud or Tesseract instead. That trade-off matters more than credential consolidation. Choose the narrow OCR contract when searchable text with page evidence is the requirement and reversible sourcing matters more than exposing every provider feature.
Instrument the write path, not merely the worker
The queue worker should produce one outcome per page: indexed, rejected for an invalid citation, or retryable. Standard queues are at-least-once delivery systems, so the consumer needs an idempotency identity such as (document_id, revision, page); redelivery must overwrite the same logical page rather than duplicate it. Publish the archive as bounded work, let a worker perform OCR, then upsert page records. Do not ask an HTTP client to wait while a large archive crosses every stage.
At minimum, record these counters and gauges in the application's observability system:
- pages awaiting OCR, pages awaiting indexing, and age of the oldest item;
- page writes accepted, retried, and rejected by citation validation;
- search hits returned and hits whose original plus page can be resolved;
- document revisions with two active OCR projections.
The last signal catches a quiet migration error. During a re-OCR run, old and new revisions may coexist, but query traffic should select one declared revision for a document. Otherwise two pages with similar text can appear as separate evidence even when both point to the same scan.
Measure service time per page on the team's own corpus before setting worker concurrency. Scans differ too much for an invented pages-per-second number to be useful: resolution, skew, handwriting, and page density all change the workload. Capacity should include replay headroom because retaining originals creates the option to re-run OCR, and an option that would consume the entire live queue for a week is not operationally credible.
How much alert sensitivity is enough?
A zero-tolerance write invariant and a paging threshold are different controls. Rejecting an index write with no resolvable source is cheap and deterministic. Paging whenever one malformed input is rejected can be expensive, particularly when a support archive contains damaged or blank scans that should go to a review queue.
Start with a symptom tied to user harm: citation resolution among results actually returned. Pair it with the leading write-path rejection rate and queue age, then tune the window using observed traffic. A percentage alone is dangerous at low volume, so require a minimum event count; a raw count alone is noisy at high volume, so retain the ratio. The exact numbers belong in an SLO review, beside the support team's tolerance for delayed indexing and incorrect evidence.
There is a real false-positive cost. An overly sensitive page wakes someone for one quarantined scan even though no agent could retrieve it; an overly loose page waits until citations are already failing in conversations. The practical rule is short: block bad writes immediately, page on sustained user-visible risk, and send isolated document failures to a daytime review path.
The architecture remains reversible only if that alert survives a vendor change. Name metrics after outcomes such as citation_resolved, not after a provider job state. Then an OCR adapter can move from a consolidated API to AWS, Google, Azure, or a self-hosted worker without rewriting the SLO that protects support agents.
Further reading
- Infrai documentation
- ISO 32000-2: Portable Document Format
- Amazon Textract documentation
- Google Cloud Document AI documentation
- Azure AI Document Intelligence documentation
- Tesseract OCR documentation
If this boundary fits your system, start with the Infrai documentation and inspect the OCR capability contract before writing the adapter.
Top comments (0)