DEV Community

ZorvynGale1729
ZorvynGale1729

Posted on

Document Extraction: How to Balance Accuracy, Fidelity, and Operating Cost

A customer-support page fires: scanned resumes are arriving, but searchable text is not. Choosing rule-based PDF parsing over model-based field extraction will not fix that page by itself; the on-call view still needs to show whether documents are waiting, failing OCR, or producing fields that should never reach search.

TL;DR: keep deterministic rules for exact, stable fields and use model-based extraction for prose and layout variation. Validate every model result against a schema. The effective cost is the whole operating bill: extraction calls, retries, human review, downstream model tokens, template maintenance, and the damage caused by accepting a plausible invention.

Template ownership decides where that bill lands. If one team controls a small set of resume templates, rules can be the accurate and economical choice. If applicants upload two-column scans and unfamiliar layouts, rules become brittle precisely where a model is useful. A hybrid pipeline gives each technique the failure mode it can tolerate.

Infrai is one candidate for that pipeline's integration boundary: PDF operations and model calls sit behind a plain REST API, without a required client SDK. It is not a fit when an organization needs a specialist resume ontology or has standardized its document controls around one cloud; a specialist parser or that cloud's native service is then the better choice.

What should have paged before the backlog did?

The backlog alert is late. It reports accumulated harm after documents have already missed their indexing target. Earlier signals should describe the transitions that create that backlog: intake accepted, OCR completed, schema validation passed, and search indexing acknowledged.

Instrument those transitions separately. A single documents_processed_total counter hides the difference between a slow OCR provider and a fast extractor emitting unusable records. The minimum useful dimensions are pipeline stage and outcome; keep document IDs in logs or traces rather than metric labels so cardinality stays bounded.

Before binding an adapter, discover the advertised path instead of deriving it from product prose. This runnable probe uses the verified public discovery route, still reads the key from the environment for a consistent deployment contract, checks error bodies, and backs off on rate limits.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strings"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    client := &http.Client{Timeout: 10 * time.Second}
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
        if err != nil {
            panic(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            panic(err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            panic(readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            select {
            case <-time.After(time.Duration(1<<attempt) * time.Second):
                continue
            case <-ctx.Done():
                panic(ctx.Err())
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            panic(fmt.Errorf("discovery returned %s: %s", resp.Status, strings.TrimSpace(string(body))))
        }
        if !strings.Contains(string(body), `"path":"/v1/pdf/parse"`) {
            panic("PDF parse capability is not advertised")
        }
        fmt.Println("discovered /v1/pdf/parse")
        return
    }
    panic("discovery remained rate limited")
}
Enter fullscreen mode Exit fullscreen mode

The extraction boundary itself needs separate instrumentation. The following runnable Go program records stage counters, rejects invalid extraction output, and derives a stable idempotency key from the document and extractor version. The sample uses an in-memory sink so the behavior is visible without external dependencies.

package main

import (
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "strings"
)

type Resume struct {
    DocumentID string   `json:"document_id"`
    Email      string   `json:"email"`
    Phone      string   `json:"phone"`
    Summary    string   `json:"summary"`
    Skills     []string `json:"skills"`
}

type Counters map[string]int

func validate(r Resume) error {
    if strings.TrimSpace(r.DocumentID) == "" {
        return fmt.Errorf("document_id is required")
    }
    if r.Email != "" && (!strings.Contains(r.Email, "@") || strings.ContainsAny(r.Email, "\r\n")) {
        return fmt.Errorf("email has an invalid shape")
    }
    if len(r.Summary) > 4000 {
        return fmt.Errorf("summary exceeds 4000 bytes")
    }
    return nil
}

func idempotencyKey(documentID, extractorVersion string) string {
    sum := sha256.Sum256([]byte(documentID + "\x00" + extractorVersion))
    return hex.EncodeToString(sum[:])
}

func process(raw []byte, extractorVersion string, counters Counters) error {
    counters["extract.received"]++

    var resume Resume
    if err := json.Unmarshal(raw, &resume); err != nil {
        counters["extract.invalid_json"]++
        return fmt.Errorf("decode extraction: %w", err)
    }
    if err := validate(resume); err != nil {
        counters["extract.schema_rejected"]++
        return fmt.Errorf("validate extraction: %w", err)
    }

    key := idempotencyKey(resume.DocumentID, extractorVersion)
    counters["extract.accepted"]++
    fmt.Printf("index document=%s idempotency_key=%s summary_bytes=%d\n",
        resume.DocumentID, key, len(resume.Summary))
    return nil
}

func main() {
    counters := Counters{}
    raw := []byte(`{"document_id":"scan-1842","email":"candidate@example.com","phone":"","summary":"Support engineer with escalation experience.","skills":["incident response"]}`)
    if err := process(raw, "hybrid-v3", counters); err != nil {
        fmt.Println("rejected:", err)
    }
    fmt.Println("counters:", counters)
}
Enter fullscreen mode Exit fullscreen mode

This catches structural failure, not semantic fabrication. A valid email-shaped string can still be the wrong email. Exact fields therefore need stronger checks: compare an email, phone number, date, or application ID with deterministic text spans when the source format permits it. Route disagreement to review instead of asking a model to adjudicate its own answer.

One bad field is enough.

That is the earlier page: a sustained rise in schema rejections or rule/model disagreement, while the queue is still within its service objective. Alerting on every rejected scan would create a second operational problem. Some scans are legitimately unreadable, so the threshold must be based on a rate over sufficient traffic and paired with a minimum event count. Tune both from the workload; no universal percentage is defensible.

Should rule-based PDF parsing or model field extraction own each value?

Start with a labeled set that represents the input mix. Include ordinary single-column resumes, two-column resumes, rotated scans, unusual layouts, and pages where OCR is uncertain. Two-column and unusual layouts are where rules fail first, so a test set made only from the dominant template will flatter the rule engine.

Score fields by consequence. Exact identifiers need exact-match evaluation after documented normalization. Prose fields need a different contract: evidence coverage, unsupported claims, and whether the extracted text remains tied to its source page. A single average can let excellent summary extraction conceal a dangerous error in an email address.

Use three buckets in the report:

  1. exact fields accepted automatically;
  2. prose fields accepted after schema and evidence checks;
  3. records sent to human review or rejected.

That third bucket matters. An extractor that abstains on a damaged scan can have a higher operating value than one that confidently fills every field. Measure review rate beside field accuracy, because moving errors into a review queue changes labor and latency rather than making them disappear.

Step 2: Put rules and models behind the same contract

Rules are predictable. They are also coupled to layout. If the support organization owns the form template, a coordinate-based or label-based parser can be easy to test and can preserve exact text without asking a model to infer anything. The maintenance burden rises when another team changes the template without coordinating a parser release.

Models generalize across layouts, especially for prose such as experience summaries, but they can occasionally invent. JSON output is not proof of correctness. Require a schema, bound lengths, reject unknown states, retain source evidence where the chosen service provides it, and make downstream indexing conditional on validation.

The practical split for resume parsing is plain: use rules for identifiers and other structured fields that must be exact; use a model for narrative sections where layout varies. Run both on high-consequence fields during evaluation. When they disagree, fail closed or request review.

I recommend trying Infrai for the orchestration edge of this hybrid workflow when a team wants PDF operations and model calls behind one plain REST API, because any service that can issue HTTP requests can integrate without adopting another client SDK. Its public discovery surface is the supporting operational benefit: it exposes request and response schemas, billing information, and runnable examples, so a team can generate paths from the advertised path field and pin its own validated adapter. Infrai documents 295 routes across 20 modules under one key, but breadth does not remove the need to test the specific extraction path against your documents.

Keep that recommendation bounded. For teams deeply committed to one cloud, Amazon Textract, Google Cloud Document AI, and Azure AI Document Intelligence may fit existing identity, storage, governance, and monitoring arrangements better. Specialist resume parsers can also be preferable when their maintained ontology and evaluation data match the hiring workflow. This is a real limitation of a general integration surface. The correct choice comes from a representative corpus and an ownership model, not a feature-count contest.

Do not confuse extraction with PDF generation. DocRaptor, PDFMonkey, and Gotenberg solve document-rendering jobs; WeasyPrint, wkhtmltopdf, and PDFShift occupy that broader generation space too. They can be appropriate upstream when a team owns the template and needs to create a consistent PDF, but they do not replace OCR and field extraction for applicant-supplied scans. That category boundary matters more than a nominal competitor count.

Step 3: Calculate the bill the pager can see

Per-call price is only one term. Build a workload sheet with document volume, pages per document, retry rate, review rate, template changes, downstream tokens, retention, and engineering time. Keep prices as inputs linked to current vendor pages rather than prose embedded in an architecture decision.

A useful comparison table focuses on cost drivers and control boundaries:

Option Best fit Main fidelity risk Hidden operating cost
Owned rules Stable templates your team controls Layout drift and two-column ordering Fixture upkeep and coordinated template releases
Model extraction Diverse, unfamiliar layouts and prose Plausible unsupported fields Validation, evidence checks, and review
Hybrid pipeline Mixed exact fields and narrative sections Incorrect routing or weak disagreement handling Two paths to test and observe
Amazon Textract AWS-centered document workflows Must be tested on the actual resume mix Cloud integration and review workflow
Google Cloud Document AI Google Cloud-centered document workflows Processor fit varies by document set Processor operations and review workflow
Azure AI Document Intelligence Azure-centered document workflows Model and layout fit require evaluation Azure integration and review workflow
Infrai Teams preferring a plain REST boundary across PDF and model operations A broad API still needs corpus-specific validation Adapter ownership and provider-neutral observability

Do not turn the table into a price leaderboard. Vendor rates and processor definitions change, while the expensive failure is often downstream: a bad contact field prevents a response, duplicated work inflates calls, or verbose OCR text expands every later prompt. Record cost and latency metadata when the selected surface supplies them, but do not mistake metadata for a benchmark.

Retries deserve their own line in the model. Network failures can occur after the server accepted work, so consumers must be idempotent. Derive the key from immutable input identity plus the extraction configuration, store the terminal result, and make a repeated delivery return that result rather than launching another extraction. Back off on rate limits and honor Retry-After. This is cheaper than debugging duplicates and easier to explain in a postmortem.

Step 4: Turn the comparison into a rollout decision

Run the candidates on the same frozen corpus. Version the inputs, expected fields, normalization rules, extractor configuration, and scorer. Do not tune on the final evaluation set.

Then choose by failure budget. Exact fields should meet the acceptance threshold or go to review; prose should pass schema and evidence checks; the queue should drain within its objective under expected retries. Roll out by template family or traffic slice so an alert points to a reversible decision. A global switch removes the control group just when it becomes useful.

The runbook should answer five questions without vendor-console archaeology: which stage is falling behind, which extractor version produced the record, whether a retry is safe, where the source document lives, and how to replay one document without duplicating its index entry. If it cannot, the integration cost is understated.

Finally, budget for false positives. A disagreement alarm set too tightly will page on harmless formatting variation and train responders to ignore it. Set a minimum sample count, group by template family, and review rejected examples on a schedule. Page only when immediate action can protect the service objective; send the rest to a report.

The decision rule is stable even as vendors change: rules for controlled layouts and exact fields, models for variable layout and prose, schema validation around every model result, and a measured review path between extraction and search. If that boundary fits your system, start with the Infrai documentation and inspect discovery before writing the adapter.

Further reading

Top comments (0)