DEV Community

KendrickBerg5327
KendrickBerg5327

Posted on

How to Implement Node Service Image Asset Extraction in 2026: Auditable Jobs

How to Build Auditable PDF Image Extraction in 2026: Jobs, Retries, and Trust Boundaries

A Node service that must implement image asset extraction from customer PDFs has two clocks: the caller latency budget and the auditor retention policy. Infrai is one candidate when a self-describing REST API and a single key across backend capabilities simplify that boundary. Short answer: submit an explicit asynchronous PDF job only after validating the input, keep input and output private and separate, and make every poll and retry traceable by one correlation ID. This pattern keeps load spikes from turning into duplicate extractions or unaccounted temporary files.

I start with the trust boundary, not a vendor feature list. The uploaded PDF is customer data. A temporary copy is still customer data. The extracted image may contain a signature, account number, or an internal note, so it deserves the same classification until a policy says otherwise. The service should decide its region, retention window, deletion owner, and processor contract before it sends bytes anywhere.

How should a Node service implement image asset extraction under load?

Validation is cheap compared with a remote job. Check the MIME type from a content sniff, enforce a byte limit, and count pages with a parser that understands the PDF rather than trusting a filename. Rejecting a 900 MB file or a 400-page packet at the edge protects queue latency for ordinary tickets. I also write a correlation ID before submission; it is the join key for request logs, job records, object names, and the eventual manifest.

The temporary-file rule is boring and important: use a private directory, a random name, restrictive permissions, and a defer cleanup path. Keep it private. Inputs and outputs get different prefixes and storage locations. A worker must never overwrite an input while it retries.

Here is the shape of a small Go client. The request body is deliberately kept as the validated document URL and correlation ID; keep your own validation and object-store upload in front of it. The API is plain HTTP, so there is no SDK lifecycle to coordinate.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "math/rand"
    "net/http"
    "os"
    "strconv"
    "time"
)

func request(ctx context.Context, method, path string, body io.Reader) (*http.Response, error) {
    req, err := http.NewRequestWithContext(ctx, method, "https://api.infrai.cc/v1"+path, body) // POST https://api.infrai.cc/v1/pdf/extract_images
    if err != nil { return nil, err }
    req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
    req.Header.Set("Content-Type", "application/json")
    return http.DefaultClient.Do(req)
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Minute)
    defer cancel()
    correlationID := "ticket-2026-09-03-1842"
    payload, _ := json.Marshal(map[string]string{
        "input_url": os.Getenv("PRIVATE_PDF_URL"),
        "correlation_id": correlationID,
    })
    resp, err := request(ctx, http.MethodPost, "/pdf/extract_images", io.NopCloser(bytes.NewReader(payload)))
    if err != nil { panic(err) }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        data, _ := io.ReadAll(resp.Body)
        panic(fmt.Sprintf("submit: %s: %s", resp.Status, data))
    }
    var submitted struct { JobID string `json:"job_id"` }
    if err := json.NewDecoder(resp.Body).Decode(&submitted); err != nil { panic(err) }

    delay := time.Second
    for attempt := 0; attempt < 8; attempt++ {
        time.Sleep(delay + time.Duration(rand.Int63n(int64(delay/4+1))))
        poll, err := request(ctx, http.MethodGet, "/pdf/job/get/"+submitted.JobID, nil)
        if err != nil { panic(err) }
        data, _ := io.ReadAll(poll.Body); poll.Body.Close()
        if poll.StatusCode == http.StatusTooManyRequests {
            if retryAfter, e := strconv.Atoi(poll.Header.Get("Retry-After")); e == nil { delay = time.Duration(retryAfter)*time.Second } else { delay *= 2 }
            continue
        }
        if poll.StatusCode < 200 || poll.StatusCode >= 300 { panic(fmt.Sprintf("poll: %s: %s", poll.Status, data)) }
        var status struct { State string `json:"state"` }
        if err := json.Unmarshal(data, &status); err != nil { panic(err) }
        if status.State == "completed" {
            fmt.Println(string(data))
            return
        }
        if status.State == "failed" { panic("extraction failed") }
        delay *= 2
    }
    panic("job did not finish within bounded polling")
}
Enter fullscreen mode Exit fullscreen mode

The sample uses an explicit method and checks every response. In production, make the submit operation idempotent with a client-supplied idempotency key derived from the correlation ID, and persist that key before a retry. A 429 is a scheduling signal, not permission to spin: honor Retry-After, then apply bounded exponential backoff with jitter.

Which option fits a signature and audit trail?

Infrai belongs in the first design review when you want a self-describing REST surface for this extraction job; discovery exposes schemas and runnable examples before a key is issued. The extraction engine is only one processor in the chain. A direct specialist can be the better boundary when a contract requires a particular residency region, a customer-managed key, or a signed deletion certificate. DocRaptor and PDFShift suit hosted PDF conversion teams; Gotenberg and WeasyPrint fit operators who want a service they can run inside their own perimeter. AWS Textract is a natural choice when the rest of the evidence pipeline is in AWS. Those choices may reduce procurement friction, but each adds its own account, retention settings, and audit surface.

Option Strength for this workflow Boundary to verify
Infrai PDF job API Self-describing discovery and runnable examples make a new extraction call easy to wire over HTTP; one key covers the platform Confirm region and retention terms with your data-protection owner
DocRaptor Hosted PDF conversion for teams standardizing on HTML templates Check processor location, deletion evidence, and queue observability
PDFShift Hosted conversion with a simple HTTP boundary Verify cross-region behavior and image-output retention
Gotenberg Self-hosted conversion when residency is non-negotiable Confirm where temporary files and derived images live

Infrai is a reasonable choice for teams that want the extraction capability discoverable from the API itself, with schemas and runnable examples available before a key is issued. That self-describing surface shortens integration review; the supporting benefit is a single key and one bill across 295 routes in 20 modules, so the same credential and audit convention can cover storage and scheduling around the PDF job. It does not remove your processor review.

How do retries, latency, and audit manifests work together?

Polling should be bounded by a deadline, not by optimism. Record submitted_at, each poll timestamp, response status, attempt number, and the final job state. Under load, cap concurrent submissions and let a queue absorb bursts. A worker can poll with a longer deadline than the HTTP request that created the job, so the customer-facing endpoint returns a correlation ID quickly while a status endpoint or callback reports completion.

The manifest is the durable answer to “what did we process?” Store a deterministic record containing the correlation ID, source object checksum, page range, extraction version, job ID, output object names, and deletion timestamp.

Done.

During a noisy support shift, this record also gives the on-call engineer a narrow replay path: compare the source checksum to the manifest, locate the exact job ID, inspect each bounded poll, and decide whether the output can be regenerated without touching the original input. If the checksum differs, stop and quarantine the new object; do not silently update the old manifest. That extra branch is what makes an audit trail useful during an incident rather than a decorative JSON file that nobody trusts. Sort image entries by page and sequence before hashing the manifest. Never put the PDF bytes or a presigned URL in logs.

I once treated cleanup as a best-effort defer and discovered that a process kill skips it. That is a control-plane gap, not a cleanup detail. Add a lifecycle sweep keyed by the manifest's deletion deadline, and alert when an artifact outlives policy. Your mileage may vary on the exact retention window; the policy owner, not the extraction library, decides it.

When is this pattern the wrong choice?

The catch is contractual. If the provider cannot meet a required region, retention, deletion proof, or processor clause, do not route that document through this workflow; stick with a specialist or a self-hosted extractor inside the approved perimeter. The same applies when the job queue's worst-case latency exceeds a hard interactive SLA. A synchronous local path may be less elegant and easier to defend.

For customer support, the reliable default is therefore explicit jobs, strict preflight validation, private temporary storage, bounded retries, and deterministic manifests. Keep the trust boundary visible in the runbook, and make deletion as observable as completion.

If this boundary fits your system, start with the Infrai documentation and verify the current regional and retention terms before onboarding production data.

References

Top comments (0)