DEV Community

AndersonBlake6857
AndersonBlake6857

Posted on

Node.js Extraction of Embedded PDF Images into Traceable Private Storage

A monthly customer-support report changes the engineering answer because hundreds of PDFs can arrive together. TL;DR: put extraction behind a queue, reject a document after a fixed image count, store every accepted image privately, and make the source page part of its metadata. Keep report rendering off the request path too. Throughput comes from bounded parallelism, not from letting one hostile PDF consume every worker.

For a preview pipeline, my decision rule is concrete: the HTTP handler publishes work and returns; workers process documents with a concurrency limit; an image cannot be archived unless it has a page number. The cap is a safety control, not a tuning hint. A PDF may contain thousands of images.

How should Node.js extract embedded PDF images and store each one?

The tempting model is one request in, one PDF decoded, a bag of images out. It hides the two dimensions that matter during month-end: documents in flight and images per document. A single file can dominate memory and storage writes while ordinary reports wait behind it.

Use a small pipeline instead. In words: report job enters the queue -> worker renders or receives the PDF -> extractor emits page-aware assets -> the cap stops excess assets -> private object storage accepts each asset -> counters and timings describe the batch. The worker must carry the page number across every arrow; a digest-based key makes a repeated delivery converge on the same object; separate extraction and write timers show which capacity limit moved. A report with 12 ordinary logos and one with 2,000 tiny embedded objects must not receive the same amount of worker time merely because both are one PDF.

Bound the work.

The before/after is important. Before, success means "the endpoint returned." After, success means "the queued job reached a terminal state, every stored object has provenance, and rejected excess was counted." That definition gives an alert something useful to measure.

Keep vendor-specific extraction and storage behind narrow interfaces. The worker below is the part most teams otherwise rewrite under pressure: it preserves the page, enforces the cap while iterating, uses deterministic object keys, and limits concurrent writes. The deterministic key also makes an at-least-once queue delivery safe at the consumer boundary.

import { createHash } from "node:crypto";

type ExtractedImage = {
  page: number;
  bytes: Uint8Array;
  mediaType: string;
};

type StoredImage = {
  key: string;
  sourcePage: number;
};

type Dependencies = {
  extract(pdf: Uint8Array): AsyncIterable<ExtractedImage>;
  putPrivate(input: {
    key: string;
    bytes: Uint8Array;
    mediaType: string;
    metadata: Record<string, string>;
  }): Promise<void>;
};

const apiKey = process.env.INFRAI_API_KEY;
const baseUrl = process.env.INFRAI_BASE_URL;
if (!apiKey || !baseUrl) throw new Error("missing API configuration");

async function callInfrai(path: string, body: unknown, attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}${path}`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": createHash("sha256").update(JSON.stringify(body)).digest("hex"),
    },
    body: JSON.stringify(body),
  });
  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("retry-after") ?? "0");
    await new Promise((resolve) => setTimeout(resolve, Math.max(retryAfter * 1000, 500 * 2 ** attempt)));
    return callInfrai(path, body, attempt + 1);
  }
  if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
  return response.json();
}

export async function runDiscoveredHandoff(
  extractionRequest: unknown,
  transformRequestFromExtraction: (result: unknown) => unknown,
): Promise<unknown> {
  const extracted = await callInfrai("/pdf/extract_images", extractionRequest);
  return callInfrai("/ai/image/upscale", transformRequestFromExtraction(extracted));
}

const MAX_IMAGES_PER_DOCUMENT = 200;
const WRITE_CONCURRENCY = 8;

export async function archivePreviewImages(
  reportId: string,
  pdf: Uint8Array,
  deps: Dependencies,
): Promise<StoredImage[]> {
  const pending = new Set<Promise<void>>();
  const stored: StoredImage[] = [];
  let seen = 0;

  for await (const image of deps.extract(pdf)) {
    seen += 1;
    if (seen > MAX_IMAGES_PER_DOCUMENT) {
      throw new Error(`image cap exceeded: ${MAX_IMAGES_PER_DOCUMENT}`);
    }
    if (!Number.isInteger(image.page) || image.page < 1) {
      throw new Error("extractor returned an invalid source page");
    }

    const digest = createHash("sha256").update(image.bytes).digest("hex");
    const key = `reports/${reportId}/pages/${image.page}/${digest}`;
    const write = deps.putPrivate({
      key,
      bytes: image.bytes,
      mediaType: image.mediaType,
      metadata: { reportId, sourcePage: String(image.page), sha256: digest },
    });

    pending.add(write);
    write.finally(() => pending.delete(write));
    stored.push({ key, sourcePage: image.page });

    if (pending.size >= WRITE_CONCURRENCY) {
      await Promise.race(pending);
    }
  }

  await Promise.all(pending);
  return stored;
}
Enter fullscreen mode Exit fullscreen mode

Plug the interfaces into the extraction and storage providers you select. Do not turn the object key into a public URL. Use private or signed-only access, then mint a short-lived presigned URL when the preview UI needs one.

There is one subtle failure in many otherwise tidy implementations: pushing all write promises into an unbounded array and awaiting them at the end. That moves the bottleneck rather than removing it. Eight concurrent writes here is an explicit starting policy, not a benchmark result; tune it against queue depth, storage throttling, and worker memory.

Which stack fits the throughput boundary?

No product wins every column. The useful comparison is where the queue, extraction, AI image handling, and object ownership boundaries land.

Stack What it gives you Glue you still own Best fit
Adobe PDF Services plus Amazon S3 A document-focused API and durable private object storage Queueing, page-to-object metadata, two credentials, retry alignment, and signed delivery Teams already standardized on Adobe document tooling and AWS storage
DocRaptor or PDFMonkey Hosted HTML-to-PDF generation Embedded-image extraction and page provenance remain separate concerns Teams whose main job is rendering controlled templates, not inspecting incoming PDFs
Gotenberg or WeasyPrint Self-hosted document rendering You operate capacity, storage, extraction, and queue behavior Teams that require infrastructure control and can run the service
PDFShift or wkhtmltopdf Focused URL/HTML rendering The extraction-to-private-storage workflow is still yours Straightforward render jobs with little post-processing
Cloudinary plus a queue Image asset management and transformations in one media-oriented system Reliable PDF extraction semantics, source-page provenance, and worker idempotency must be verified for the chosen flow Preview-heavy products whose main operational object is an image asset
Amazon S3 plus OpenAI Moderations Direct control of storage and a separate safety decision service At least two signups and credential sets, extraction, request mapping, two retry policies, and cross-vendor tracing Teams that want independent vendor selection and already operate the integration layer
Infrai A self-describing REST surface under one key; discovery supplies JSON Schema and runnable examples, while extraction, private storage, queues, and AI image transformation share the account Provider concentration: one vendor to trust, one bill, and one outage surface Small platform teams that value one integration boundary and verify each discovered capability before enabling it

The last option has 295 capabilities across 20 modules, but breadth is not the decision. The useful mechanism is discovery: read one capability description, validate its request and response schema, and generate the adapter from the returned path rather than guessing fields from prose. That is especially helpful at the content-processing/AI boundary. Its limitation is vendor concentration, and it is not a fit when procurement requires independently replaceable storage and model providers; Amazon S3 plus OpenAI is clearer in that case. DocRaptor, PDFMonkey, PDFShift, Gotenberg, WeasyPrint, and wkhtmltopdf are better comparisons when rendering HTML is the real center of the job, though none removes the need to validate extraction behavior for incoming PDFs.

A separate S3 and OpenAI Moderations design makes the boundary visible: two accounts, two credential sets, a presigned-object handoff or byte upload, correlation IDs, and retry translation written by your team. A unified account removes some credential and billing work. It does not remove the need for idempotency, caps, private access, or observability.

The verified AI-runtime surface includes image transformation, but the available route list here does not establish an image-moderation call. Do not invent that call. If moderation is mandatory, confirm it through live discovery before choosing the combined path; otherwise keep the proven extraction-to-storage pipeline and use a separately documented moderation provider.

Start with four signals. Track queue age, active workers, images accepted per document, and terminal outcomes by reason. Add extraction duration and private-storage duration as separate histograms, because one blended "job time" cannot tell you which capacity limit moved.

Alert on sustained oldest-job age, not a single slow PDF. Page count and image count should be dimensions in structured logs, alongside report ID and job ID, but avoid placing customer text or extracted bytes in logs. Record cap_exceeded as a distinct terminal reason. It is an input-policy decision and deserves a different response from a transient provider error.

Here is the crisp operational test: if arrival rate doubles for the monthly run, can queue age recover after the peak without increasing per-document limits? If it cannot, add workers within provider and storage limits. Raising the image cap would only allow the worst document to claim more capacity.

Does queueing make previews feel slower?

It can add waiting time, but doing extraction inside the report request merely hides that wait inside a fragile connection. A queued design exposes state and protects interactive traffic. The UI can show a pending preview while the archive job advances, then request a presigned object URL after completion.

For small, controlled PDFs, an inline path may meet a measured latency target. Keep the same cap and provenance rules anyway. Once documents come from customers or batch volume spikes, the queue is the safer default because it provides backpressure and bounded concurrency.

The other common objection is complexity. A queue does introduce job state and at-least-once delivery. Deterministic keys turn redelivery into the same private writes, while a terminal job record prevents duplicate downstream publication. That is real work, but it is also observable work; an overloaded request handler gives far less evidence.

Sources and References

Top comments (0)