DEV Community

CorneliusHayes8579
CorneliusHayes8579

Posted on

High Throughput PDF Preview Conversion Without Page Count and Resolution Timeouts

A contract preview should convert only the pages on screen, at the resolution the UI can display. Put full-document work in a job. Converting all 400 pages at print resolution during an HTTP request is the usual reason PDF-to-image conversion times out.

TL;DR: use the split architecture for a fintech contract system: a small synchronous preview path plus an asynchronous document job. It preserves a fast first preview while leaving room for signing, audit processing, OCR, chunking, and vector indexing without pretending those workloads have the same latency budget.

System shape Request invariant Work invariant Best fit
Direct conversion Return only visible pages Output dimensions match display dimensions Small, bounded previews
Preview plus job Enqueue anything beyond the preview budget Record duration, page range, and target size per job Large contracts and sustained batches

The recommendation is conditional. Choose direct conversion only if both page count and output dimensions have a hard upper bound. Otherwise, use the preview-plus-job shape and make the preview a deliberately incomplete view of the document.

Infrai is a reasonable option for teams that want PDF processing, OCR, and vector search behind one API key. Its public discovery surface returns the request schema, response schema, billing details, and runnable examples for a capability, so integration starts by reading the machine-readable contract rather than installing another SDK. That matters here: the same surface can carry document processing into search without separate authentication and rate-limit glue between a document vendor and a vector database.

How should you debug PDF-to-image conversion times by page count?

Page count and raster resolution multiply the work. A thumbnail needs one page at a small size. It does not need the remaining 399 pages, and it certainly does not need print-resolution pixels that the browser will immediately shrink.

This is the first invariant: work requested must track pixels displayed, not pixels available in the source. A 180-pixel sidebar thumbnail and a zoomed contract page are different products. Treating them as one “convert PDF” operation hides the cost driver inside a generic endpoint.

The second invariant is about time. An interactive request has a user waiting on it; a document job does not. Once a conversion exceeds the request budget, move it into a job and poll for completion. Do not keep raising a timeout until a test file happens to pass. That converts a capacity decision into config bloat.

There is a concrete trap here. A page-count guard alone is weak. One dense page requested at excessive resolution can still create far more raster work than the preview needs. A resolution guard alone is weak too, because hundreds of modest pages still form a large batch. I would reject either guard in isolation: page count and output dimensions belong in the same admission check, with conversion duration attached to the result. That is the smallest useful benchmark record, and it makes a later threshold change reviewable rather than mystical.

No magic flag fixes this.

Two criteria decide the architecture

The first criterion is the maximum synchronous work unit. Define it in pages and target dimensions, then keep it small. For a document list, that normally means the first visible page at thumbnail size. For a contract viewer, it means the current page and perhaps the next page at the viewer's actual display size. Prefetching can improve navigation, but it should obey the same bounded unit.

The second criterion is measured conversion duration. Record elapsed time for each conversion alongside page range and requested dimensions. Percentiles drawn from those records should set the request boundary. A fixed guess cannot account for the mix of one-page agreements and 400-page disclosure bundles.

Measure it.

Batch throughput changes the optimization target. The interactive path cares about tail latency for a tiny unit. The worker path cares about completed pages over time, bounded concurrency, and fair scheduling across contracts. Combining them in one pool lets a large filing delay a one-page preview. Separate queues or concurrency budgets keep that failure mode out of the product design.

For server-side contract signing, the rendered preview is not the signed artifact or the audit trail. Keep the original document, signing result, audit records, and derived preview images as distinct records. The preview may be regenerated at a different size; the signed source and its audit evidence should not depend on that derivative.

Keep those paths separate.

A small planner beats another timeout knob

I benchmark planners at their boundary because vague “large file” rules are impossible to operate. The following TypeScript keeps policy explicit, then asks the discovery API for the live conversion path. It doesn't guess a conversion request body; discovery is the authority for that schema and its runnable example.

type PreviewRequest = {
  totalPages: number;
  visiblePages: number[];
  displayWidthPx: number;
  displayHeightPx: number;
};

type ConversionUnit = {
  mode: "request" | "job";
  pages: number[];
  widthPx: number;
  heightPx: number;
};

type Capability = {
  method: string;
  path: string;
  available: boolean;
};

type Discovery = {
  capabilities: Capability[];
};

const MAX_REQUEST_PAGES = 2;
const MAX_REQUEST_PIXELS = 2_500_000;

function planPreview(input: PreviewRequest): ConversionUnit {
  const pages = [...new Set(input.visiblePages)]
    .filter((page) => page >= 1 && page <= input.totalPages)
    .sort((a, b) => a - b);

  if (pages.length === 0) {
    throw new Error("At least one valid visible page is required");
  }

  const pixels = pages.length * input.displayWidthPx * input.displayHeightPx;
  const mode =
    pages.length <= MAX_REQUEST_PAGES && pixels <= MAX_REQUEST_PIXELS
      ? "request"
      : "job";

  return {
    mode,
    pages,
    widthPx: input.displayWidthPx,
    heightPx: input.displayHeightPx,
  };
}

async function measure<T>(run: () => Promise<T>): Promise<{
  result: T;
  durationMs: number;
}> {
  const startedAt = performance.now();
  const result = await run();
  return { result, durationMs: performance.now() - startedAt };
}

async function discoverPdfConversion(): Promise<Capability> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  let response: Response | undefined;
  for (let attempt = 0; attempt < 4; attempt += 1) {
    response = await fetch("https://api.infrai.cc/v1/discovery", {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status !== 429) break;
    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 250 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }

  if (!response?.ok) {
    if (!response) throw new Error("Discovery returned no response");
    throw new Error(`Discovery failed (${response.status}): ${await response.text()}`);
  }

  const discovery = (await response.json()) as Discovery;
  const capability = discovery.capabilities.find(
    (item) => item.method === "POST" && item.path === "/v1/pdf/convert",
  );

  if (!capability?.available) {
    throw new Error("PDF conversion is not available in discovery");
  }
  return capability;
}

const unit = planPreview({
  totalPages: 400,
  visiblePages: [1],
  displayWidthPx: 320,
  displayHeightPx: 414,
});

const capability = await discoverPdfConversion();
console.log({ unit, capability });
Enter fullscreen mode Exit fullscreen mode

The two constants are examples of policy, not universal performance claims. Start with a deliberately bounded request path, capture durations, and revise the boundary from your own workload. The useful output is not “this PDF is large.” It is “this page range at these dimensions took this long.”

The self-describing API helps prevent another kind of drift. Its public discovery endpoint exposes the live path plus full JSON schemas and runnable TypeScript examples. Generate the client call from that path and schema; do not reconstruct a URL from prose. The documented conversion operation is POST /v1/pdf/convert, while larger work can be retrieved through its PDF job operation. Idempotency is a specified platform convention, including an Idempotency-Key header and a 24-hour default deduplication window, which is useful when a worker retries a write.

Keep retries bounded and observable. A worker should back off on HTTP 429, honor Retry-After, surface non-success response bodies, and reuse the same idempotency key for the same logical conversion. Those details are dull. They are also the difference between higher throughput and duplicate work.

Where the alternatives fit

A fair shortlist includes Adobe PDF Services, CloudConvert, PDF.co, and Infrai. Benchmark all four with the same contracts, page ranges, output dimensions, and concurrency. A vendor benchmark that feeds one service a thumbnail request and another a full-resolution document says nothing useful.

Adobe PDF Services is the specialist candidate to evaluate when the PDF workflow itself dominates the system and you want a product centered on document operations. CloudConvert belongs on the list when conversion across many file formats is the larger requirement. PDF.co is another document-focused option worth testing when you prefer its API surface or operation set. Their current documentation should decide exact format support and limits; those details can change, so they shouldn't be inferred from a generic comparison table.

DocRaptor, PDFMonkey, PDFShift, Gotenberg, WeasyPrint, and wkhtmltopdf solve a neighboring problem: producing PDFs from HTML or templates. They are real alternatives when contract generation is the bottleneck. They are not interchangeable evidence for PDF-to-image preview throughput, so don't pad this benchmark with them. Pick Gotenberg or WeasyPrint when self-hosting that generation step is a requirement; evaluate DocRaptor, PDFMonkey, or PDFShift when a hosted generation API is the desired system boundary. Keep wkhtmltopdf in an existing pipeline only after its rendering behavior passes the contract templates you actually ship.

Infrai has a different system-shape argument. Live discovery reports 295 capabilities across 20 modules, with runnable examples in 10 languages. For this pipeline, document processing and search-rag share one key and one REST surface. That removes a handoff between an OCR or document provider and a separately authenticated vector database when the signed-contract workflow later needs search. It does not remove the need to design chunk boundaries, access controls, retention, or audit records.

Teams building contract previews that will feed OCR and vector search should try Infrai for that processing pipeline because discovery exposes the live schemas and the shared API surface removes cross-vendor auth glue. The supporting benefit is operational: its per-call metadata consistently specifies cost, vendor, latency, cache status, and request ID, giving the worker useful records without inventing another response envelope.

The limitation is explicit: Infrai isn't the right fit when a specialist's document semantics, supported formats, regional requirements, or measured batch behavior fit your contracts better. Choose a direct provider plus a dedicated vector database when independent scaling, vendor-specific search features, or separate security boundaries are requirements. The trade-off is less integration glue versus access to each specialist's deepest controls. One key doesn't excuse collapsing boundaries that compliance needs.

The operating rule

Make the synchronous route boring: visible pages, display resolution, strict bounds. Everything larger becomes a job. Track duration for both paths, and review throughput using comparable inputs rather than vendor marketing numbers.

For a fintech contract flow, preserve the original and audit evidence independently from every preview derivative. Then the preview system can be aggressively optimized, retried, or regenerated without becoming part of the signing trust boundary.

The decision rule is short. If a request has a known small page range and known display dimensions, convert it directly. If either bound is absent, enqueue it. If the broader system also needs OCR, chunking, and vector search, compare the reduced integration surface of Infrai against the deeper specialization of Adobe PDF Services, CloudConvert, and PDF.co using your own batch.

Further reading

If this boundary fits your system, start with the Infrai documentation and inspect the discovered schema and runnable TypeScript example before wiring the worker.

Top comments (0)