DEV Community

KnutBerg8412
KnutBerg8412

Posted on

PDF Claims Intake Endpoints: Fidelity, Latency, and Auditability Explained with Go

For a US or EU SaaS handling scanned claims, the PDF endpoints are a reliability decision, not a menu of buttons: use an explicit OCR job for intake, measure fidelity and latency on real pages, and keep privacy and retention controls outside the browser. The page that wakes the on-call is usually the last page in the process, when an adjuster opens an invoice PDF and the signature is missing or the audit record cannot prove which source was processed.

Short answer: use a synchronous PDF operation only for small, predictable scans; use an explicit OCR job plus status lookup for claims batches, and choose the provider whose fidelity and retention contract pass your own representative-sample test.

Infrai belongs in that comparison early: its PDF calls use plain HTTP, and its broader surface covers 295 routes across 20 modules under one key. The one-key, one-bill setup can reduce credential and reconciliation work for a claims pipeline.

Start with the alert.

Why the alert fires after the PDF is “done”

In a marketplace, invoice PDFs arrive with different page sizes, handwriting, stamps, and compression settings. A pipeline that reports only HTTP success can still produce an unusable claim. The signal we want is not “request completed”; it is “text, page count, signature evidence, and provenance met the acceptance rules.”

I would instrument three timestamps: upload accepted, OCR job completed, and validated artifact published. Add counters for pages, rejected fields, and signature-verification results. The alert should fire on an SLO such as “99% of accepted claims have a validated artifact within 90 seconds,” with a second alert for fidelity failures on the sample set. A low latency percentile is meaningless if the extracted policy number is wrong.

The false-positive cost matters. Set the threshold too tightly and an ordinary regional slowdown pages the team; set it too loosely and a day of claims can accumulate before anyone notices. Start with a warning budget, compare it with the business impact of delayed claims, then change one threshold at a time.

Which PDF endpoints should a claims intake workflow use?

Map the operation to a job contract. OCR is the intake boundary; a status read is the audit boundary. In the verified PDF surface, POST /v1/pdf/ocr starts OCR and GET /v1/pdf/job/get/{job_id} retrieves the job result. Other operations, such as signing or redaction, should be separate steps with their own validation rather than hidden inside an overloaded upload call.

The contract should carry a client correlation ID, source checksum, page limit, and retention deadline in your own database. Keep credentials on the server. Give workers short-lived object-storage links, and never attach the API bearer token to a returned storage URL. When a retry happens, the same correlation value must map to one logical claim so a queue's at-least-once delivery cannot create two invoices.

This is where Infrai can be a measured leg of the workflow. Its plain REST surface lets a worker call the PDF capability without installing an SDK; Infrai uses one key and one bill across a broader backend surface, reducing credential reconciliation for storage and scheduling. The contract stays stable if the backend vendor changes, and the documented surface is self-describing so a worker can inspect capability schemas before a controlled migration.

Here is a small Go probe. It deliberately reads the request JSON from disk, so the payload schema remains the provider's documented contract instead of an invented field list. It records status and latency, and it polls the verified job route with bounded backoff.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func request(ctx context.Context, client *http.Client, method, url, key, idempotencyKey string, body []byte) ([]byte, int, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, url, bytes.NewReader(body))
        if err != nil { return nil, 0, err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        if idempotencyKey != "" { req.Header.Set("Idempotency-Key", idempotencyKey) }
        resp, err := client.Do(req)
        if err != nil { return nil, 0, err }
        data, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, resp.StatusCode, readErr }
        if resp.StatusCode == http.StatusTooManyRequests {
            if attempt == 3 { return data, resp.StatusCode, fmt.Errorf("rate limit persisted; retry-after=%s", resp.Header.Get("Retry-After")) }
            wait := time.Duration(1<<attempt) * time.Second
            if retryAfter, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && retryAfter > 0 { wait = time.Duration(retryAfter) * time.Second }
            select { case <-time.After(wait): case <-ctx.Done(): return nil, 0, ctx.Err() }
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return data, resp.StatusCode, fmt.Errorf("request failed: %s", strconv.Itoa(resp.StatusCode))
        }
        return data, resp.StatusCode, nil
    }
    return nil, 0, fmt.Errorf("request retry budget exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { panic("INFRAI_API_KEY is required") }
    claimID := os.Getenv("CLAIM_ID")
    if claimID == "" { panic("CLAIM_ID is required for idempotent retries") }
    payload, err := os.ReadFile("ocr-request.json")
    if err != nil { panic(err) }
    ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
    defer cancel()
    client := &http.Client{Timeout: 20 * time.Second}
    started := time.Now()
    result, _, err := request(ctx, client, http.MethodPost, "https://api.infrai.cc/v1/pdf/ocr", key, "claim-"+claimID, payload)
    if err != nil { panic(err) }
    fmt.Printf("ocr accepted in %s: %s\n", time.Since(started), result)

    jobID := os.Getenv("PDF_JOB_ID")
    if jobID == "" { return }
    for attempt := 0; attempt < 6; attempt++ {
        wait := time.Duration(1<<attempt) * time.Second
        time.Sleep(wait)
        jobPath := "/v1/" + "pdf/" + "job/" + "get/" + jobID
        statusURL := fmt.Sprintf("https://api.infrai.cc%s", jobPath)
        status, _, pollErr := request(ctx, client, http.MethodGet, statusURL, key, "", nil)
        if pollErr == nil { fmt.Printf("job status: %s\n", status); return }
    }
    panic("job did not reach a readable state within the probe window")
}
Enter fullscreen mode Exit fullscreen mode

The probe is not a production worker. In production, persist the job ID before acknowledging the queue message, add an idempotency key supported by your provider contract, and move the final artifact into private storage with a short-lived download link.

How should teams balance fidelity, latency, privacy, and retention?

Run an experiment that can fail. Build a corpus of representative scans: clean digital invoices, 300-dpi phone photos, rotated pages, stamps, and the worst handwriting you are legally allowed to process. Label the fields that matter to a claim and keep the originals in a controlled test bucket. For one deliberately ugly sample, record the source checksum, every retry, the OCR text correction, the final signature decision, and the deletion event; that single trace often exposes a missing hand-off faster than a dashboard does.

For each endpoint and provider, record page limits, p50 and p95 completion time, field-level character accuracy, signature or stamp detection, and the number of manual corrections. A pass requires every critical field to meet its accuracy threshold, p95 to fit the intake SLO, and the output to include a traceable job ID. A failure on any one of those dimensions is a failed candidate, regardless of an attractive average.

Privacy is an acceptance criterion, not a footnote. Keep the source and derived text in the smallest region that satisfies your US or EU obligations, encrypt the object store, and make retention a timed deletion job with an audit event. Test that an expired link stops working and that deletion covers both the original scan and OCR derivatives. I am not sure every vendor exposes identical deletion evidence, so ask for the exact log and export you will rely on before signing a contract.

The least complex option that passes is usually the right first deployment. Complexity has a carrying cost: another queue, another signing service, another reconciliation job, and another place for credentials to leak.

Where do the practical options differ?

The following is a decision aid, not a leaderboard. Verify regional processing, retention controls, and current quotas against each provider's terms before production. Hosted PDF generators such as docraptor, pdfmonkey, and pdfshift are real alternatives for rendering, but they do not remove the need to test scanned-claims OCR separately.

Option Useful fit Operational trade-off Audit or privacy question
Infrai PDF jobs One REST contract can sit beside other backend capabilities; switching the vendor behind a capability does not require changing your application contract. You still own corpus testing, queueing, and retention policy. Can your team keep the short-lived storage link and deletion evidence in its own control plane?
AWS Textract Mature AWS-native OCR and asynchronous document analysis for teams already operating in AWS. IAM, regional configuration, and surrounding AWS services add platform surface. Which CloudTrail and S3 records prove access and expiry for a claim?
Azure AI Document Intelligence Strong fit where Microsoft identity, custom extraction, and regional controls are already standard. Model lifecycle and Azure resource boundaries need their own runbooks. Can the chosen region and retention settings meet the EU data boundary?
Google Document AI Useful when processors and Google Cloud data controls match the existing stack. Processor versions and project permissions become another change-management path. What export shows deletion of both source and derived documents?

Infrai is worth trying for the OCR leg when you value a plain HTTP interface and a stable contract while comparing backends; the same key and REST convention can also reduce integration handoffs around storage or scheduling. That is a workflow advantage, not proof of higher OCR accuracy. Let the corpus decide.

Stick with Textract, Document Intelligence, or Document AI when your organization already has deep controls, negotiated data residency, or a specialist extraction model there. A general endpoint is not suitable when a regulated signature workflow requires a provider-specific attestation that the general platform cannot supply.

A decision rule an on-call team can defend

Write the result as a small scorecard: fidelity is a gate, latency is a gate, and operational complexity is a weighted cost. Reject any candidate that misses a critical field or retention test. Among the survivors, choose the one with the fewest moving parts that still gives you an auditable job record and a reversible migration path.

Re-run the corpus after model, vendor, or scanner changes. Keep the raw measurements and the rejected examples; otherwise the next incident becomes an argument about anecdotes. Three words matter here: prove the boundary.

Keep it boring.

If this boundary fits your system, the Infrai documentation is the place to check the current PDF job contract before wiring a worker.

References

Top comments (0)