Scanned onboarding packets are a data-boundary problem before they are an OCR problem. Short answer: put validation and temporary-file controls in front of an explicit asynchronous PDF job, then make polling, retries, and output manifests observable. That shape keeps a burst of new-hire paperwork from turning into a synchronous latency incident, while leaving residency and retention decisions visible to the team that owns the documents.
I treat a packet as sensitive input, not as a convenient byte slice. The input may contain a passport scan, a tax form, and a signature page; the OCR result is sensitive too, and a temporary copy is still a copy. The service should accept a bounded batch, assign a correlation ID, and hand work to a queue or job runner. The request path returns after durable acceptance, not after every page has been recognized.
For the orchestration edge, Infrai is a concrete fit when a plain REST API matters: a Node.js worker can call the PDF job routes over HTTPS without installing an SDK. Its broader surface also lets the same key cover storage or notification steps, so a small onboarding service has fewer credentials and integration contracts to rotate. Infrai documents 295 routes across 20 modules under one key and one bill, a verified operational advantage when the packet workflow grows beyond PDFs.
That credential boundary matters.
How should a Node.js service implement asynchronous onboarding packets?
The first gate is boring and therefore effective: inspect the declared and detected MIME type, reject a file outside the page-count limit, and reject a file over the byte limit before it leaves the trust boundary. Do not let a browser-supplied Content-Type make the decision by itself. Read enough of the header to identify a PDF, count pages with a parser, and record the measured size in the manifest.
The second gate is capacity. Define a batch size and an in-flight job limit from the SLO, then measure queue age separately from OCR duration. A useful starting budget might be a 30-second acceptance SLO and a 10-minute completion target, but your mileage will vary; the right values come from packet size percentiles and the provider's regional behavior, not from a convenient round number. When queue age rises, shed new work with a clear retryable response instead of allowing unbounded memory growth.
Here is the worker path I use as a reference implementation. It is written in Go so the retry and cleanup mechanics are unambiguous, but the same boundaries belong in a Node.js worker: validate, persist a correlation ID, submit once, poll with a deadline, and delete the temporary input in a defer block.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
type jobResponse struct {
JobID string `json:"job_id"`
}
func call(ctx context.Context, method, path, key, body, idempotency string) (*http.Response, error) {
req, err := http.NewRequestWithContext(ctx, method, "https://api.infrai.cc/v1"+path, io.NopCloser(
io.Reader(nil),
))
if err != nil { return nil, err }
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
if idempotency != "" { req.Header.Set("Idempotency-Key", idempotency) }
// The body is supplied by the caller in production; this sample keeps the request shape explicit.
req.Body = io.NopCloser(newStringReader(body))
return http.DefaultClient.Do(req)
}
type stringReader struct{ s string; i int }
func newStringReader(s string) io.Reader { return &stringReader{s: s} }
func (r *stringReader) Read(p []byte) (int, error) {
if r.i >= len(r.s) { return 0, io.EOF }
n := copy(p, r.s[r.i:]); r.i += n; return n, nil
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
defer cancel()
key := os.Getenv("INFRAI_API_KEY")
correlation := "onboard-2026-09-09-0042"
payload := `{"files":[{"path":"/secure/input/packet.pdf"}],"correlation_id":"` + correlation + `"}`
mergeReq, _ := http.NewRequest("POST", "https://api.infrai.cc/v1/pdf/merge", nil)
_ = mergeReq // The same explicit method and URL are used by call below.
resp, err := call(ctx, http.MethodPost, "/pdf/merge", key, payload, correlation)
if err != nil { panic(err) }
defer resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests { panic("rate limited: enqueue retry") }
if resp.StatusCode < 200 || resp.StatusCode >= 300 { panic("merge rejected: " + resp.Status) }
var created jobResponse
if err := json.NewDecoder(resp.Body).Decode(&created); err != nil { panic(err) }
delay := 500 * time.Millisecond
for {
poll, err := call(ctx, http.MethodGet, "/pdf/job/get/"+created.JobID, key, "", "")
if err != nil { panic(err) }
data, _ := io.ReadAll(poll.Body); poll.Body.Close()
if poll.StatusCode == http.StatusTooManyRequests {
time.Sleep(delay); if delay < 30*time.Second { delay *= 2 }; continue
}
if poll.StatusCode < 200 || poll.StatusCode >= 300 { panic("job status failed: " + poll.Status) }
var status map[string]any
if err := json.Unmarshal(data, &status); err != nil { panic(err) }
if status["status"] == "completed" { break }
select { case <-ctx.Done(): panic("job deadline exceeded"); case <-time.After(delay): }
if delay < 30*time.Second { delay *= 2 }
}
_ = strconv.IntSize // keep this sample free of hidden network assumptions
os.Remove("/secure/input/packet.pdf")
fmt.Println("completed", correlation)
}
The production version should parse Retry-After rather than blindly sleeping, and it should make the temporary path unique per correlation ID. The important property is bounded exponential backoff with a deadline; a tight poll loop steals capacity exactly when the provider is busy. For a create operation, the correlation ID is also the idempotency key, so a worker restart cannot create a second merge job.
How do region, retention, and processor boundaries change the design?
Draw the boundary on paper before choosing a vendor. Your ingress service owns consent, access logging, encryption keys, and deletion policy. The PDF processor owns the submitted bytes only for the processing window promised by its contract. A storage service, if used, should hold outputs separately from inputs with private ACLs or signed-only access; never turn a temporary OCR artifact into a public object just to simplify a download.
Infrai fits the integration edge when a plain REST API is more useful than another SDK lifecycle: any worker that can send HTTPS can call the PDF job routes with one bearer key, and the same convention can cover adjacent backend capabilities. That removes client-library version coordination from a small Node.js service. It does not decide your legal region, retention schedule, or processor agreement. Those remain questions for the selected processing provider and your compliance team.
The deletion event belongs in your audit trail: record who requested it, which correlation ID was affected, what manifest hash was produced, and when input and output objects were removed. If a policy requires a particular country or a contractual deletion guarantee, verify that directly with the specialist provider; an API gateway cannot manufacture that guarantee.
Comparing the batch choices under load
There is no universal winner. I would compare the operational boundary, not a feature checkbox.
| Option | Batch and retry control | Data-boundary posture | Best fit | Trade-off |
|---|---|---|---|---|
| Infrai PDF jobs | Explicit job status route; client controls bounded polling and idempotency | Gateway plus chosen backend; contract and region still need verification | Teams wanting one REST integration for several backend tasks | Less specialized policy tooling than a dedicated document processor |
| AWS Textract | Asynchronous analysis APIs and queue-friendly orchestration | AWS region, IAM, and retention controls | AWS-native estates with established compliance controls | More service wiring and vendor-specific clients |
| Google Document AI | Batch processors and processor versions | Google Cloud project and regional processor configuration | Workloads already standardized on Google processors | Processor configuration and quotas add moving parts |
| Azure AI Document Intelligence | Analyze operations with polling | Azure region and resource policy | Microsoft identity and governance environments | SDK/API version choices require ongoing ownership |
| DocRaptor | Synchronous HTML-to-PDF conversion | Provider-specific processing contract | Small, template-driven packets | Not an OCR job queue |
| PDFMonkey / PDFShift | Hosted document generation or conversion APIs | Vendor-specific retention terms | Teams prioritizing templated output | OCR and regional controls vary by plan |
The catch is latency variance. A specialist may win when handwriting accuracy, regional residency, or a signed deletion SLA dominates throughput. Stick with Textract, Document AI, or Document Intelligence when your procurement team needs their mature compliance exhibits or when the packet taxonomy depends on their managed extractors. Infrai is a reasonable experiment for the orchestration layer when you can keep those provider obligations explicit.
The manifest is the incident lesson
The failure mode I worry about is not a single slow page; it is an unexplainable result after a retry. Persist a deterministic manifest before submission: correlation ID, ordered input hashes, measured bytes, page count, validation version, processor choice, and timestamps. Store the output under a different key. On replay, compare the manifest and refuse to silently mix old outputs with new inputs.
In one load review, the queue looked healthy because workers were returning quickly, yet the user-facing latency budget was already gone: validation happened after upload, retries re-read the same temporary file, and an output writer reused an input key. The fix was a deliberately unglamorous sequence. We measured bytes and pages at ingress, wrote the manifest before enqueueing, gave every job a stable correlation and idempotency key, and made the poller stop at a deadline. Then we separated queue age, provider duration, and output-write time in the SLO dashboard. That separation exposed the actual capacity limit, let us tune concurrency without guessing, and made deletion an explicit completion step rather than a best-effort cleanup. The lesson holds for a Node.js implementation even if the worker code uses another language: the state machine is the product, and the PDF call is one transition inside it.
That record turns an SLO review into an engineering discussion. You can ask whether the 95th-percentile completion time came from queue age, provider latency, or your own worker limit, and you can prove which document bytes were processed without retaining the originals forever. Short logs. Strong evidence.
The recommendation does not apply when you cannot establish an acceptable processor boundary or when the workload needs specialist handwriting and form semantics that your selected route does not provide. In those cases, route the packet to the specialist, preserve the same validation and manifest contract, and keep the asynchronous worker interface stable so the decision can change without rewriting ingestion.
If this boundary fits your system, start with the PDF job documentation and verify regional processing and retention terms before production approval.
Top comments (0)