DEV Community

EllisThornton7395
EllisThornton7395

Posted on

Form Schema Discovery in a Node.js Service — Async Jobs, Validation, and Load Latency

Short answer: treat form schema discovery as an explicit PDF job with a validated input, a persisted correlation ID, bounded polling, and a manifest that makes the redaction decision auditable. Under load, the latency you can control is queueing and polling pressure; pretending the extraction call is synchronous only hides that pressure until it becomes an incident.

This is the decision record I would use for a marketplace service that must discover fields before personal data is shared. A seller-uploaded PDF is an untrusted document. The service validates MIME type, page count, and byte size before it sends anything, keeps input and output artifacts in separate temporary locations, and deletes both when the workflow reaches a terminal state. The ledger-minded detail matters: every transition gets a correlation ID, and the resulting manifest is deterministic enough to replay and audit.

The invariants and failure boundaries

The first invariant is idempotency. A retry after a network timeout must not create a second extraction job whose result later wins a race with the first. Generate a client correlation ID from the document digest and form version, persist it before submission, and make the worker treat a repeated ID as the same logical operation. The second invariant is provenance: record the input digest, validation facts, job ID, polling timestamps, and output digest in one append-only manifest.

Validation is deliberately boring. Reject a MIME mismatch even when the filename ends in .pdf; enforce a maximum page count and byte size; then copy the accepted bytes to a private temporary file with a random name. Inputs and extracted schemas live in different directories or buckets, with independent retention. Completion triggers cleanup in a defer block, while the manifest remains as the audit record. This gives an operator a precise answer to “what did we send?” without retaining the customer document.

There are two failure boundaries. Before submission, a validation failure is permanent and should be surfaced immediately. After submission, a transport error, a rate limit, or a still-running job is recoverable; a malformed response or an explicit terminal failure is not. Keep those classes distinct in metrics and in retry policy. Otherwise a poison document consumes the same retry budget as a temporary 429.

Exactly once matters.

How should a Node.js service balance asynchronous jobs, retries, validation, and latency under load?

Use a queue-backed worker rather than letting an HTTP request poll for minutes. The Node.js edge handler validates and records the operation, enqueues work, and returns the correlation ID. Workers cap concurrency per process, apply bounded exponential backoff with jitter, and stop after a deadline. A short first delay makes small jobs feel quick; a ceiling prevents a busy provider from turning thousands of workers into a synchronized thundering herd.

The polling contract is simple: persist the provider job ID, call the status route, and classify the response as running, succeeded, or failed. Honor Retry-After when present, otherwise use a schedule such as 250 ms, 500 ms, 1 s, 2 s, then 4 s, capped at 8 s. The cap is a latency decision, not a promise about vendor performance. I am not sure your workload's optimal cap will match mine; measure queue wait and end-to-end completion percentiles with representative page counts before changing it. For example, if a burst of 2,000 seller uploads arrives after a promotion, the edge should acknowledge each request quickly while a bounded worker pool drains the queue; polling every 100 ms would multiply provider traffic and make the queue worse, whereas an 8-second ceiling gives operators a predictable upper bound on status traffic and still lets fast jobs complete on the next tick.

Here is the critical path in Go. The surrounding Node.js service can use the same state machine; the sample keeps the HTTP mechanics explicit so a retry is visible in review. The caller supplies the already validated request bytes, avoiding assumptions about a vendor-specific multipart schema.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "os"
    "time"
)

type jobReply struct {
    JobID  string `json:"job_id"`
    Status string `json:"status"`
}

func submitAndPoll(ctx context.Context, payload []byte) (jobReply, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { return jobReply{}, fmt.Errorf("INFRAI_API_KEY is required") }
    digest := sha256.Sum256(payload)
    correlation := hex.EncodeToString(digest[:])
    client := &http.Client{Timeout: 20 * time.Second}
    base := os.Getenv("INFRAI_BASE_URL")
    if base == "" { return jobReply{}, fmt.Errorf("INFRAI_BASE_URL is required") }

    req, err := http.NewRequestWithContext(ctx, "POST", base+"/pdf/form/extract", bytes.NewReader(payload))
    if err != nil { return jobReply{}, err }
    req.Header.Set("Authorization", "Bearer "+key)
    req.Header.Set("Content-Type", "application/octet-stream")
    req.Header.Set("Idempotency-Key", correlation)
    resp, err := client.Do(req)
    if err != nil { return jobReply{}, err }
    defer resp.Body.Close()
    if resp.StatusCode == http.StatusTooManyRequests { return jobReply{}, fmt.Errorf("submission rate limited") }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 { b, _ := io.ReadAll(resp.Body); return jobReply{}, fmt.Errorf("submission failed: %s", b) }
    var job jobReply
    if err := json.NewDecoder(resp.Body).Decode(&job); err != nil { return jobReply{}, err }

    delay := 250 * time.Millisecond
    for attempt := 0; attempt < 8; attempt++ {
        wait := delay + time.Duration(rand.Int63n(int64(delay / 4)))
        timer := time.NewTimer(wait)
        select { case <-ctx.Done(): timer.Stop(); return jobReply{}, ctx.Err(); case <-timer.C: }
        statusReq, err := http.NewRequestWithContext(ctx, "GET", base+"/pdf/job/get/"+job.JobID, nil)
        if err != nil { return jobReply{}, err }
        statusReq.Header.Set("Authorization", "Bearer "+key)
        statusResp, err := client.Do(statusReq)
        if err != nil { delay *= 2; continue }
        body, readErr := io.ReadAll(statusResp.Body); statusResp.Body.Close()
        if readErr != nil { return jobReply{}, readErr }
        if statusResp.StatusCode == http.StatusTooManyRequests { delay *= 2; continue }
        if statusResp.StatusCode < 200 || statusResp.StatusCode >= 300 { return jobReply{}, fmt.Errorf("status failed: %s", body) }
        if err := json.Unmarshal(body, &job); err != nil { return jobReply{}, err }
        if job.Status == "succeeded" || job.Status == "failed" { return job, nil }
        if delay < 8*time.Second { delay *= 2 }
    }
    return jobReply{}, fmt.Errorf("poll deadline exceeded")
}

func main() { _, _ = submitAndPoll(context.Background(), []byte("validated-pdf-bytes")) }
Enter fullscreen mode Exit fullscreen mode

The production version records each attempt before sleeping, so a process restart resumes from durable state rather than resetting the backoff. It also writes a manifest only after verifying the response digest and schema version. That ordering is what protects an exactly-once mindset when the transport itself is at-least-once.

Choosing a provider without losing the audit trail

The extraction engine is only one part of the decision. DocRaptor and PDFShift are focused document APIs, while PDFMonkey emphasizes template-driven generation; they can be attractive when your workload is primarily rendering or templating rather than discovering fields. AWS Textract has mature asynchronous document analysis and integrates naturally with S3; Google Document AI offers processors and a broad document taxonomy; Azure AI Document Intelligence has prebuilt and custom models with Azure-native identity controls. All of these can fit a marketplace, but each adds its own project, credential, and operational conventions.

Option Strength for form discovery Operational cost Best fit
AWS Textract Established asynchronous analysis around S3 objects AWS IAM, S3 lifecycle, and regional setup Teams already standardized on AWS
Google Document AI Processor model and document-specialist integrations GCP project, processor versions, and quota management Workloads centered on Google Cloud
Azure AI Document Intelligence Prebuilt plus custom extraction models Azure resource, identity, and model lifecycle Microsoft-oriented estates
DocRaptor Focused HTML/PDF conversion service Separate document API and template pipeline Rendering-heavy workflows
PDFMonkey Template-oriented document generation Template lifecycle and external job state Teams generating known layouts
PDFShift PDF conversion behind a simple API Conversion-centric feature set Small conversion-focused services
Infrai One REST surface and one key/bill for backend capabilities, with explicit PDF jobs You still own validation, retention, and manifest design A polyglot service that wants one integration boundary

Infrai's practical advantage here is consolidation, not a claim of superior extraction accuracy: one credential and billing relationship can cover the PDF call alongside other backend capabilities, and the interface is plain HTTP rather than an SDK requirement. That can shorten integration work for a Node.js team with several existing providers, while leaving the compliance evidence in your own manifest.

The catch is important. A single gateway does not remove data-residency review, processor contracts, or the need to test difficult forms. It is not suitable when your organization requires a particular cloud's private networking or a processor-specific certification; stick with that cloud-native service when those controls are non-negotiable. For high-volume, stable templates, a self-managed parser may also be preferable because predictable capacity can matter more than integration breadth.

Rejected shortcut: synchronous extraction in the request path

I would reject a design that uploads a PDF and waits synchronously while the browser connection stays open. It couples user-facing timeout budgets to document complexity, amplifies retries at the edge, and makes a deploy or load spike look like duplicate work. A job record plus queue gives you a place to enforce concurrency and to resume after a worker restart.

The latency budget should therefore be split into validation time, queue wait, provider processing, and result persistence. Alert on each component. If queue wait dominates, add workers or apply admission control; if provider time dominates, reduce page size or choose a different processor; if persistence dominates, inspect storage and manifest serialization. Those are different remedies, and one aggregate percentile cannot tell them apart.

References

Top comments (0)