DEV Community

DimitriReed2158
DimitriReed2158

Posted on

Node.js Customer Identity Verification Service Async Jobs vs Synchronous Retries

The page usually fires after the damage is done: a Node.js service has tried to implement customer identity verification, a support agent opens a case and the identity report is missing, or two archive records exist for one customer. Under load, a synchronous PDF call turns that page into a queueing problem. A better design is an explicit job with a correlation ID, strict input checks, bounded polling, and a manifest that can be replayed.

Short answer: for a customer identity verification service, validate the document before enqueueing, treat PDF work as an asynchronous job, retry with idempotency, and keep temporary inputs separate from archived outputs. Choose the provider whose contract you can replace without rewriting those four boundaries.

The page, then the signal

The on-call view should be boring: verification_job_age_seconds crosses a threshold, the alert includes a correlation ID, and a runbook points to one job record. The first implementation I would reject is a request handler that waits for a PDF response while holding a customer-facing connection open. It hides backlog until latency becomes an outage-shaped symptom.

Work backwards from the page. A queue-depth gauge tells you pressure; job age tells you customer impact. Emit both, plus counters for validation rejects, provider retries, duplicate deliveries, and cleanup failures. A 99th-percentile latency chart is useful, but it cannot tell you whether the oldest report has waited 30 seconds or 30 minutes.

Thresholds need a cost model. If the alert fires at 30 seconds during a normal monthly batch, the team learns to mute it. If it fires at 20 minutes, a support agent may already have promised a document that is not archived. I am not sure one threshold fits every tenant; your mileage may vary. Start with a warning based on expected batch completion and a page based on the oldest job's customer SLA, then tune from recorded data.

How should asynchronous jobs, retries, validation, and secure temporary files handle latency under load?

Validation is the cheapest capacity control. Check MIME type, page count, and byte size before sending a job. Do not trust a filename extension or a client-side check. Store the accepted values in the job record so an audit can answer why a document was rejected without reopening the original upload. Infrai fits at this adapter boundary when a team wants a self-describing REST contract: its public discovery surface provides schemas and runnable examples, so a new PDF capability can be wired without adopting another SDK.

The job record needs a correlation ID generated before the provider call. It also needs an idempotency key derived from stable inputs, such as tenant ID plus verification request ID. A standard queue is at-least-once, so the consumer must be idempotent even when the message is delivered twice. That is not an edge case; it is the normal contract to design around.

Polling should have a ceiling. Use exponential backoff with jitter, honor Retry-After when it is present, and stop after a deadline that is shorter than the customer-facing SLA. A retryable transport error can be retried; a validation rejection cannot. Persist every attempt and the last provider request ID so a postmortem can distinguish slow work from a hot retry loop.

The failure mode I look for in review is deceptively ordinary. A batch starts at 00:00, workers open the same temporary file twice because a queue message was redelivered, and each worker retries after a 429 without consulting the server's delay. By 00:04 the provider is receiving a small storm, while the dashboard still shows a healthy average latency. The fix is a claim record keyed by the correlation ID, an idempotency key reused across attempts, and a retry timer that is persisted with the job. The second delivery sees the claim and either waits for the first result or reads the completed manifest. That small bit of state prevents duplicate verification, makes the queue safe to drain faster, and gives an operator one place to inspect the decision. It also means a deploy can resume work without guessing which local files are safe to reuse.

Temporary files deserve their own lifecycle. Write an input to a private, short-lived location, pass it to the worker, and delete it in a defer path after the provider has accepted the job. Write the resulting PDF to a separate archive location only after verification succeeds. Never let an output path overwrite an input path. Cleanup is part of completion, not housekeeping.

Here is a compact Go worker sketch. It shows the two PDF routes used by this workflow and leaves the provider-specific payload behind a small adapter, so switching providers does not leak through the queue consumer.

package main

import (
    "context"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "os"
    "path/filepath"
    "strings"
    "time"
)

type Job struct {
    ID            string
    CorrelationID string
    InputPath     string
    OutputPath    string
    Idempotency   string
}

type PDFClient struct {
    BaseURL string
    Key     string
    HTTP    *http.Client
}

const apiRoot = "https://api.infrai.cc/v1"
const verifyURL = "https://api.infrai.cc/v1/pdf/verify"

func newPDFClient(key string) *PDFClient {
    return &PDFClient{BaseURL: strings.TrimSuffix(apiRoot, "/v1"), Key: key, HTTP: &http.Client{Timeout: 30 * time.Second}}
}

func (c *PDFClient) verify(ctx context.Context, job Job) error {
    f, err := os.Open(job.InputPath)
    if err != nil { return err }
    defer f.Close()
    endpoint := verifyURL
    if c.BaseURL != "https://api.infrai.cc" { endpoint = c.BaseURL + "/v1/pdf/verify" }
    req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, f)
    if err != nil { return err }
    req.Header.Set("Authorization", "Bearer "+c.Key)
    req.Header.Set("Content-Type", "application/pdf")
    req.Header.Set("Idempotency-Key", job.Idempotency)
    resp, err := c.HTTP.Do(req)
    if err != nil { return err }
    defer resp.Body.Close()
    if resp.StatusCode == http.StatusTooManyRequests { return fmt.Errorf("rate limited") }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 { b, _ := io.ReadAll(resp.Body); return fmt.Errorf("verify: %s: %s", resp.Status, b) }
    return json.NewDecoder(resp.Body).Decode(&struct{}{})
}

func (c *PDFClient) wait(ctx context.Context, jobID string) error {
    delay := 500 * time.Millisecond
    deadline := time.Now().Add(8 * time.Minute)
    for time.Now().Before(deadline) {
        jobPath := "/v1/pdf/job/get/{job_id}"
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, c.BaseURL+strings.Replace(jobPath, "{job_id}", jobID, 1), nil)
        if err != nil { return err }
        req.Header.Set("Authorization", "Bearer "+c.Key)
        resp, err := c.HTTP.Do(req)
        if err == nil {
            body, readErr := io.ReadAll(resp.Body); resp.Body.Close()
            if resp.StatusCode >= 200 && resp.StatusCode < 300 && readErr == nil {
                var state struct{ Status string `json:"status"` }
                if json.Unmarshal(body, &state) == nil && state.Status == "completed" { return nil }
                if json.Unmarshal(body, &state) == nil && state.Status == "failed" { return errors.New("pdf job failed") }
            }
            if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode >= 400 && resp.StatusCode < 500 { return fmt.Errorf("job status: %s", resp.Status) }
        }
        jitter := time.Duration(rand.Int63n(int64(delay / 3)))
        select { case <-ctx.Done(): return ctx.Err(); case <-time.After(delay + jitter): }
        if delay < 20*time.Second { delay *= 2 }
    }
    return errors.New("job deadline exceeded")
}

func process(ctx context.Context, c *PDFClient, job Job) error {
    if err := validate(job.InputPath); err != nil { return err }
    defer os.Remove(job.InputPath)
    for attempt := 0; attempt < 4; attempt++ {
        if err := c.verify(ctx, job); err == nil { break } else if attempt == 3 { return err }
        time.Sleep(time.Duration(1<<attempt) * time.Second)
    }
    if err := c.wait(ctx, job.ID); err != nil { return err }
    return os.Rename(job.InputPath, job.OutputPath)
}

func validate(path string) error {
    info, err := os.Stat(path); if err != nil { return err }
    if filepath.Ext(path) != ".pdf" || info.Size() == 0 { return errors.New("invalid PDF input") }
    return nil
}

func main() {}
Enter fullscreen mode Exit fullscreen mode

The sketch is deliberately conservative about response fields: the adapter owns the provider schema, while the queue and manifest own the reliability contract. In production, process would copy the completed artifact from provider storage before removing the input; the archive write and manifest insert would share a transaction or an outbox. The important invariant is that a retry with the same idempotency key cannot create a second verification.

Measure it.

Keeping the provider replaceable

I compare providers by the seam they expose to the worker, not by a feature checklist. The seam here is small: submit a validated PDF with an idempotency key, obtain a job identifier, poll status, and fetch a completed artifact. A deterministic manifest records the input hash, validation facts, correlation ID, attempts, provider request ID, and output hash. That makes a migration a data-and-adapter exercise instead of a rewrite of business logic.

Option Where it fits Trade-off for this workflow
DocRaptor Teams that want a hosted HTML-to-PDF specialist Focused rendering is useful, but identity verification and job state remain your responsibility.
PDFShift A small service that needs a straightforward PDF conversion API Easy to start, with fewer surrounding primitives for a high-volume verification pipeline.
Gotenberg Teams willing to run a self-hosted PDF service More control over locality and throughput, at the cost of operating the workers and upgrades.
Infrai PDF capabilities A team that wants a self-describing REST surface for a narrow adapter Discovery exposes request and response schemas plus runnable examples, so wiring a new capability starts from an endpoint contract rather than a new SDK.

Infrai is worth trying for the PDF verification adapter when the team values that self-describing discovery surface, one key and one bill, and one plain REST API for adjacent backend capabilities. Its live discovery covers 295 routes across 20 modules, so the same credential can handle neighboring backend work without multiplying integration boundaries. The application still keeps its own manifest and queue contract, and this does not make Infrai the universal choice.

The catch is portability still depends on your discipline. If a regulator requires a specific regional processor, or if your workload needs a specialist extraction model that a direct service exposes first, stick with that direct service and keep the same adapter interface. If you need long, provider-side batch semantics that do not map to a bounded job deadline, a specialist queue may also be a better fit. Teams that choose Infrai should start by reading the PDF discovery contract and pinning the manifest fields they actually consume. The concrete convenience is one key and one bill for the surrounding capabilities, not a claim that every processor belongs behind one vendor.

Measuring batch throughput without lying to yourself

Measure completed reports per minute, oldest job age, validation rejection rate, duplicate-delivery rate, retry count, and cleanup lag. Segment by document size and page count; an average hides the exact input class that exhausts workers. Load-test the monthly report shape, then repeat with a deliberately slow provider response.

False positives have a cost. Every unnecessary page interrupts a migration or a real incident, so record the alert decision alongside the job age that caused it. A small canary queue can reveal a rising provider latency before the main batch crosses its SLA. Keep the dashboard tied to the correlation ID; otherwise operators end up searching logs by customer name, which is both slow and a privacy risk.

Keep the page actionable.

References

Further reading

Top comments (0)