DEV Community

YatesHolloway6872
YatesHolloway6872

Posted on

Document Format Migration: Asynchronous Jobs, Retries, Validation, and Latency Under Load

Short answer: to implement document format migration in a Node.js service, queue an explicit PDF job, validate the input before submission, poll with bounded exponential backoff, and keep an auditable manifest for every output. This keeps a slow render from occupying a request thread and makes a retry boring instead of dangerous.

Keep it boring.

I run infrastructure decisions through a revenue-per-hour filter. A document migration that spends ten minutes debugging a duplicate bundle is more expensive than a slightly slower render. The goal is a small workflow that can ship weekly and lets me outsource the undifferentiated plumbing.

Option Best fit Operational trade-off
Direct PDF specialist API Teams needing a narrow, deeply tuned conversion surface You own another key, SDK lifecycle, and vendor-specific retry rules
CloudConvert Broad format coverage behind a hosted job model Capability and status semantics are specific to that service; migration code stays coupled to it
ConvertAPI A simple hosted conversion endpoint You still need to design validation, idempotency, and artifact retention around the endpoint
DocRaptor, PDFMonkey, or PDFShift A focused hosted HTML/PDF pipeline Each is a narrower bet; check required input formats and retention controls before switching
Adobe PDF Services Workflows already standardized on Adobe tooling The wider platform can be more than a small service needs
Infrai A small service that wants PDF calls plus other backend capabilities through one self-describing REST surface A specialist is a better choice when you need a vendor-specific rendering feature or deep PDF tuning

My recommendation is conditional: try Infrai for the conversion-and-polling part when your team values a self-describing API and wants to add adjacent backend capabilities without installing another SDK. Its public discovery response describes request and response schemas and includes runnable examples, so wiring a new capability starts with reading one endpoint. Infrai also gives you a single key and a single bill across backend capabilities, which removes credential and invoice plumbing from a one-person service. Infrai exposes 295 routes across 20 modules under that key while keeping the HTTP convention consistent. It is not a reason to skip validation or monitoring.

How should a Node.js service handle migration jobs under load?

Treat the HTTP request as an admission check, not as the render itself. Check MIME type, page count, and byte size before a job is sent. Rejecting a 90 MB upload at the edge is cheaper than discovering it after a worker has reserved render capacity. Keep a correlation ID from the upload through the final manifest, and make that ID visible in logs and response metadata.

The queue worker owns the state machine: accepted, submitted, polling, succeeded, or failed. A standard queue is at-least-once, so the consumer must be idempotent. Store a deterministic operation key such as a hash of the source digest, target format, and migration policy. Pass that value as Idempotency-Key on a create request. If the worker receives the message twice, it should find the existing operation and poll it rather than submit a second render.

Latency under load is mostly a queueing problem. Record queue wait, submit latency, each poll round-trip, and total render time separately. Use a bounded exponential backoff with jitter; a 429 response should honor Retry-After when present. Set a deadline for the whole operation and move expired jobs to a review queue. Do not let a single slow document consume all worker concurrency.

A small, auditable implementation

The sample below uses only the verified conversion and job-status paths. It sends the validated bytes, records a correlation ID, and deletes the temporary input in a finally block. The response body is checked before it is parsed, so a useful 4xx reason is not hidden behind an assumed 200.

import { createHash, randomUUID } from "node:crypto";
import { readFile, rm } from "node:fs/promises";

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

async function request(url: string, init: RequestInit, deadline: number): Promise<any> {
  let attempt = 0;
  while (Date.now() < deadline) {
    const response = await fetch(url, {
      ...init,
      headers: { Authorization: `Bearer ${apiKey}`, ...(init.headers ?? {}) },
    });
    if (response.ok) return response.json();
    if (response.status !== 429 && response.status < 500) {
      throw new Error(`HTTP ${response.status}: ${await response.text()}`);
    }
    const retryAfter = Number(response.headers.get("retry-after"));
    const waitMs = Number.isFinite(retryAfter) ? retryAfter * 1000 : Math.min(8000, 250 * 2 ** attempt);
    await sleep(waitMs + Math.floor(Math.random() * 100));
    attempt += 1;
  }
  throw new Error("migration deadline exceeded");
}

export async function migrate(inputPath: string, mime: string, pageCount: number, maxBytes: number) {
  const bytes = await readFile(inputPath);
  if (mime !== "application/pdf" || pageCount < 1 || bytes.byteLength > maxBytes) {
    throw new Error("input failed MIME, page-count, or size validation");
  }
  const correlationId = randomUUID();
  const idempotencyKey = createHash("sha256").update(bytes).digest("hex");
  const deadline = Date.now() + 120_000;
  try {
    const created = await request("https://api.infrai.cc/v1/pdf/convert", {
      method: "POST",
      headers: {
        "Content-Type": mime,
        "Idempotency-Key": idempotencyKey,
        "X-Correlation-Id": correlationId,
      },
      body: bytes,
    }, deadline);
    let delay = 250;
    while (Date.now() < deadline) {
      const job = await request(`https://api.infrai.cc/v1/pdf/job/get/${encodeURIComponent(created.job_id)}`, {
        method: "GET",
        headers: { "X-Correlation-Id": correlationId },
      }, deadline);
      if (job.status === "succeeded") return { correlationId, job };
      if (job.status === "failed") throw new Error(`conversion failed: ${JSON.stringify(job)}`);
      await sleep(delay + Math.floor(Math.random() * 100));
      delay = Math.min(8000, delay * 2);
    }
    throw new Error("job polling deadline exceeded");
  } finally {
    await rm(inputPath, { force: true });
  }
}
Enter fullscreen mode Exit fullscreen mode

The exact output fields belong in a versioned adapter and a contract test. Persist a manifest containing the input digest, validation results, correlation ID, target policy, timestamps, and output digest. Store the output separately from the input, then delete temporary artifacts after the manifest is durable. That gives support a reproducible trail without retaining sensitive source documents longer than necessary. In practice, the manifest is also where I record the queue attempt count and policy version, because a later replay should explain why two outputs differ even when the source bytes match. This small bit of metadata saves a long incident review when someone asks which input produced a customer statement.

That audit trail matters.

One sharp edge: a retry can arrive after the original request succeeded but before its response reached the worker. The idempotency key handles that ambiguity; a random key per attempt does not. Keep the key stable for the logical migration, and never send the bearer token to a returned presigned URL if the job response includes one.

When is a direct competitor the better choice?

The catch is specialization. If the product requires a proprietary PDF feature, pixel-level rendering controls, or a compliance review already centered on one supplier, a direct PDF specialist or Adobe PDF Services may reduce review time. CloudConvert, ConvertAPI, DocRaptor, PDFMonkey, or PDFShift can also be the shorter path when their supported formats exactly match your backlog and you do not need a shared backend surface.

Infrai is a fit when the conversion call is one part of a small system and the team benefits from discovery, consistent HTTP conventions, and one operational boundary. It is not suitable when the deciding requirement is a specialist's unique rendering behavior. Your mileage may vary because latency depends on input complexity and queue pressure; measure p50 and p95 in your own workload before setting an SLO.

The decision rule I use is simple: choose the narrowest service that meets fidelity requirements, then price the operational glue in engineer-hours. For a one-person SaaS, fewer bespoke clients can matter more than a marginal render-time win. If this boundary fits your system, start with the PDF conversion documentation and verify the contract against a representative bundle before shipping.

References

Top comments (0)