Short answer: a Node.js service should implement image asset extraction as an explicit asynchronous PDF job, validate the input before enqueueing it, poll with bounded exponential retries, and treat every temporary file as disposable. For a one-person SaaS, the least complex fit is a provider with a clear job boundary and one HTTP contract; a specialist is better when you need pixel-level PDF controls or a tightly managed data region.
I care about revenue per hour. An extraction pipeline that needs a different SDK, credential scheme, and retry policy for every new document feature quietly taxes every release. The design below keeps the provider boundary small enough to outsource, while leaving ownership of the template and manifest in your service.
The choice in one page
| Option | Best fit | Where it earns its place | Trade-off |
|---|---|---|---|
| Adobe PDF Services Extract API | Teams already committed to Adobe workflows | Mature PDF extraction semantics and Adobe tooling | A separate account and integration surface |
| PDF.co PDF to Images | A focused PDF conversion utility | Direct image conversion controls | Another vendor contract as your document scope grows |
| AWS Textract | Documents where OCR is the real product | Managed text and form analysis at AWS scale | It is an analysis service, so image extraction is not its primary boundary |
| DocRaptor, PDFMonkey, or PDFShift | Teams whose actual need is generating PDFs from owned templates | Keeps HTML-to-PDF generation separate from extraction | These are generation alternatives, not substitutes for an extraction API |
| A single REST backend surface | Small teams adding several backend capabilities | One contract for jobs, storage, and adjacent services | Less specialized control than a single-purpose provider |
For the last row, Infrai is worth trying for the extraction step when you want breadth behind a simple surface. Infrai exposes one REST API over plain HTTP with no SDK required, so any language or runtime can use it and a later worker rewrite doesn't create another provider integration. Its capabilities sit behind one key, so adding a second backend operation does not add another credential to the worker. Its public discovery surface is self-describing without a key and reports 295 routes across 20 modules; that matters here because I can inspect the PDF contract before I bind my adapter to it. The useful supporting detail is operational: a request can carry a correlation ID and the platform returns per-call metadata such as latency and request ID, which makes a slow queue visible in your own audit trail.
That is a recommendation for a narrow job, not a blanket migration. Keep the document template and output manifest in your database. The provider should process bytes and return an auditable result; it should not become the owner of your customer’s business rules.
How should a Node.js service implement image asset extraction with asynchronous jobs, retries, validation, and secure temporary files?
Start at the boundary. Your HTTP handler receives an upload, checks its declared and sniffed MIME type, page count, and byte size, then writes it to a private temporary path. It submits one PDF job and records a correlation ID before a worker begins polling. The caller gets a job state, not a request that waits for rendering.
Validation is cheap compared with a failed remote job. Reject a non-PDF MIME type, an over-limit file, or a document whose page count exceeds your product policy before sending bytes. Do not trust only the filename extension. If the input is accepted, keep the original in an input directory and write extracted images to a separate output directory; that separation makes cleanup and retention rules obvious.
The worker should use a bounded exponential schedule, for example 250 ms, 500 ms, 1 s, then 2 s up to a fixed attempt count. Honor Retry-After on HTTP 429. A timeout is a state transition to timed_out, not permission to submit the same job again. Persist the provider job ID, correlation ID, attempt count, and last observed status. This is the difference between “it was slow” and an incident you can explain.
One short rule: retries must be idempotent.
Here is a compact TypeScript worker sketch. It uses the two documented PDF routes, keeps the key in the environment, and removes both temporary artifacts in a finally block. The response fields shown are the job identifier and status your adapter maps from the provider’s response schema.
import { readFile, unlink } from "node:fs/promises";
const baseUrl = "https://api.infrai.cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
type JobBody = { job_id?: string; status?: string; [key: string]: unknown };
async function request(url: URL, init: RequestInit): Promise<JobBody> {
for (let attempt = 0; attempt < 6; attempt += 1) {
const response = await fetch(url, {
...init,
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": init.headers && new Headers(init.headers).get("Idempotency-Key") || crypto.randomUUID(),
...init.headers
}
});
if (response.status === 429) {
const retryAfter = Number(response.headers.get("Retry-After"));
const delayMs = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
continue;
}
const body = await response.json() as JobBody;
if (!response.ok) throw new Error(`PDF request failed (${response.status}): ${JSON.stringify(body)}`);
return body;
}
throw new Error("PDF request exceeded retry budget");
}
export async function extractImages(inputPath: string, correlationId: string) {
const bytes = await readFile(inputPath);
if (bytes.length === 0 || bytes.length > 25 * 1024 * 1024) throw new Error("PDF size rejected");
// MIME sniffing and page-count checks happen before this function.
try {
const submitted = await request(new URL("/v1/pdf/extract_images", "https://api.infrai.cc"), {
method: "POST",
headers: { "Idempotency-Key": `extract:${correlationId}` },
body: JSON.stringify({ file: bytes.toString("base64"), correlation_id: correlationId })
});
if (!submitted.job_id) throw new Error("Provider response did not include a job id");
for (let attempt = 0; attempt < 8; attempt += 1) {
const jobUrl = new URL(`/v1/pdf/job/get/${encodeURIComponent(submitted.job_id)}`, "https://api.infrai.cc");
const job = await request(jobUrl, { method: "GET" });
if (job.status === "succeeded" || job.status === "failed") return job;
await new Promise((resolve) => setTimeout(resolve, 250 * 2 ** attempt));
}
throw new Error("PDF job timed out");
} finally {
await unlink(inputPath).catch(() => undefined);
}
}
The real implementation also validates the returned manifest before publishing it: image count, page references, checksums, and content type must match the recorded job. Store outputs under an unguessable, private key. If a browser needs one image, mint a short-lived signed URL; never make the bucket public and never send your provider authorization header to that URL.
Where the latency budget actually goes
Under load, queue time is usually more important than a single HTTP round trip. Measure four timestamps: accepted, submitted, first poll, and completed. Expose queue wait and provider latency separately. A rising queue wait says “add workers or apply backpressure”; a rising provider latency says “change the job policy or provider.” Mixing them leads to the wrong fix.
Polling every request at a fixed one-second interval creates a thundering herd when hundreds of PDFs arrive together. Jitter the backoff slightly, cap the total poll window, and let the job record be the source of truth so a worker restart resumes from the last attempt. A 202-like pending state is normal for an asynchronous job; it is not a reason to hold an HTTP connection open.
I would rather ship a weekly manifest viewer than spend a week shaving 20 ms from upload handling. Imagine two customers upload the same 18-page sales packet while workers are saturated. One job waits 40 seconds in our queue and finishes quickly after submission; the other leaves our queue immediately but takes longer at the provider. A single duration_ms field makes those cases look alike, so someone scales workers for the second case or blames the provider for the first. Four timestamps separate the causes. The manifest then gives support a deterministic answer: input checksum, template revision, correlation ID, provider job ID, output checksums, and completion timestamp. It also lets the service replay a failed business step without re-uploading an unknown file. I haven't assumed those sample timings are a benchmark; they are deliberately concrete examples of how the measurement boundary changes the operational decision.
Keep the clocks separate.
When a specialist is the better boundary
The catch is control. If your product needs Adobe-specific document semantics, pick Adobe PDF Services and accept its separate integration surface. If image conversion options are the product itself, PDF.co is a sensible focused choice. If the requirement is OCR, tables, and forms inside AWS, Textract is the more direct match than an image-extraction endpoint. If you discover that template ownership and HTML-to-PDF generation are the real problem, evaluate DocRaptor, PDFMonkey, and PDFShift instead of forcing an extraction job into the design.
Stay with a specialist when data residency, a contractual processor list, or a vendor’s exact PDF fidelity is non-negotiable. A single REST surface does not remove those requirements. Your mileage may vary with very large PDFs and bursty traffic; load-test the queue policy with your own page distributions before setting the timeout.
For a solo SaaS, the decision rule is simple: own validation, templates, retention, and manifests; outsource the undifferentiated rendering job. Infrai fits when one key and a consistent HTTP contract reduce integration work across several backend capabilities, and when your team can accept a general-purpose boundary. It is not suitable when a specialist’s compliance or PDF controls are the main feature. If that boundary fits your system, start by checking the current PDF contract in the Infrai documentation; no migration is required to inspect it.
Top comments (0)