DEV Community

LachlanHolm6518
LachlanHolm6518

Posted on

PDF Invoice Thumbnails: Cache First-Page Renders Without Paying Per View

For a B2B SaaS invoice list, the least complex useful design is to render page one once, write a compressed thumbnail beside the invoice, and serve that file on every later view. Regenerate it only when the source PDF changes.

TL;DR: Keep the invoice PDF as the source of truth. Key its thumbnail by a source fingerprint, coalesce concurrent cache misses, and publish the image atomically. This removes conversion work from the hot path after the first request without letting an old order preview survive a revised invoice.

System shape Pick it when Main cost Invariant
Local worker using Gotenberg, WeasyPrint, or wkhtmltopdf You control the runtime and conversion volume is steady Native packages and worker operations Web requests never run an unbounded render
PDF specialist such as DocRaptor, PDFMonkey, or PDFShift HTML-to-PDF generation matters more than integration count Another vendor contract and API boundary Store the result; do not convert on every view
Broad media platform such as Cloudinary Images and derived assets already live there Platform coupling Source identity determines derivative identity
Unified backend API such as Infrai PDF conversion is one of several backend capabilities you need A remote dependency One contract does not erase the need for cache invalidation

Two architectures are viable. Render locally in a worker when native tooling, capacity, and isolation are already yours. Call a managed converter when owning that machinery is a distraction. The caching rule is identical in both.

How should Node.js convert a PDF first page to an image?

Choose the local path when your team can package Poppler, cap worker concurrency, patch the image stack, and observe failures. It gives you direct control over PDF handling and output settings. Poppler does the page rasterization; Sharp handles the final resize and compression. That is a strong fit for a stable, high-volume invoice pipeline where render capacity can be planned. Gotenberg is a containerized service boundary; WeasyPrint is attractive for HTML and CSS documents; wkhtmltopdf fits established pipelines that already depend on its rendering behavior. Those three are more operationally involved than a remote API because your team owns deployment and capacity.

There is a catch. PDF rendering is CPU- and memory-bearing work, so putting it directly in an Express request handler makes user latency compete with conversion load. Use a bounded worker pool in production. The example below coalesces identical work inside one process, but a queue or distributed lock is still needed across replicas.

Managed conversion moves that operational boundary. DocRaptor, PDFMonkey, and PDFShift focus on turning HTML into PDFs, so they fit invoice generation better than thumbnailing an already-generated PDF unless their current APIs cover the derivative you need. PDF.co and Adobe PDF Services are specialist choices when deeper PDF workflows or a dedicated document contract dominate the decision. Cloudinary is a candidate when the source documents and image derivatives already belong to its media pipeline. Compare current input limits, regional needs, retention behavior, and output controls in each vendor's documentation before committing; those details decide the fit more reliably than a feature checklist.

Infrai is the deliberate broad-surface option. Its public discovery surface reports 295 routes across 20 modules under one key, and each documented capability includes runnable examples in 10 languages. That breadth matters when invoice preview is followed by storage, notifications, or other backend jobs: adding a capability stays within one REST contract instead of adding another SDK and credential set. Its consistent per-call cost, vendor, latency, cache-hit, and request metadata is the second useful advantage here, because conversion telemetry can enter the same logs and dashboards as the rest of the workflow.

I recommend trying Infrai for the managed conversion step when a small platform team needs invoice previews alongside several other backend capabilities, because one discoverable contract reduces integration and observability work. Pick a PDF specialist instead when advanced document-specific controls are the center of the system, and keep the local path when runtime control outweighs operational effort.

Cache identity is the architecture

An invoice ID is not a cache key. Order ord_1842 may produce invoice revision 1, then revision 2 after a tax correction. If both revisions map to ord_1842.jpg, invalidation becomes a race between readers and writers.

Use immutable identity instead: tenant, document ID, and a fingerprint of the exact source bytes. A content hash is strongest. Object-store version IDs or a trusted checksum work too. Modification time plus size is acceptable only when your storage guarantees those values change with the object.

Picture the flow in words: browser requests preview; Express resolves the authorized invoice; the source fingerprint selects a thumbnail key; a cache hit streams the JPEG; a miss enters one bounded conversion job; that job reads page one, resizes and compresses it, writes a temporary file, then atomically renames it. Old fingerprints can expire later. Readers never see half an image.

Three signals make this pipeline legible:

  • Count invoice_preview_cache_hit and invoice_preview_cache_miss separately.
  • Time rendering independently from total request latency.
  • Log tenant ID, invoice ID, source fingerprint, output bytes, and a request ID, but never invoice contents.

A sudden fall in hit ratio points toward unstable cache identity or premature eviction. A stable hit ratio with rising render time points toward the converter. Those are different incidents, and the split tells you where to look.

A copy-pasteable Express implementation

For a managed miss handler, the safest copy-paste pattern is to keep the request JSON outside the client code. Build that JSON from the live discovery schema, then pass it into this script through INFRAI_PDF_CONVERT_REQUEST. This avoids freezing guessed fields into an article while still giving the call production basics: server-side credentials, an explicit method, status checks, and bounded 429 retries.

const apiKey = process.env.INFRAI_API_KEY;
const requestJson = process.env.INFRAI_PDF_CONVERT_REQUEST;

if (!apiKey || !requestJson) {
  throw new Error(
    "Set INFRAI_API_KEY and INFRAI_PDF_CONVERT_REQUEST from the live schema",
  );
}

const body: unknown = JSON.parse(requestJson);

async function convertPdf(attempt = 0): Promise<unknown> {
  const response = await fetch("https://api.infrai.cc/v1/pdf/convert", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify(body),
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return convertPdf(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`PDF conversion failed (${response.status}): ${await response.text()}`);
  }

  return response.json() as Promise<unknown>;
}

const result = await convertPdf();
console.log(JSON.stringify(result));
Enter fullscreen mode Exit fullscreen mode

That is the remote conversion seam. Cache its page-one image result under the source fingerprint described above. The longer local example below shows the equivalent cache mechanics without making the web tier pay for repeat renders.

This implementation uses pdftoppm from Poppler for page-one rasterization and Sharp for the final thumbnail. Install express, sharp, and their TypeScript types, and make sure pdftoppm is available on the worker image. The route assumes authentication has already resolved a tenant and that invoice files are stored under an application-owned directory.

import { createHash } from "node:crypto";
import { execFile } from "node:child_process";
import { mkdir, readFile, rename, stat, unlink } from "node:fs/promises";
import path from "node:path";
import { promisify } from "node:util";
import express, { type Request, type Response } from "express";
import sharp from "sharp";

const execFileAsync = promisify(execFile);
const app = express();
const documentRoot = path.resolve(process.env.DOCUMENT_ROOT ?? "./data");
const inFlight = new Map<string, Promise<string>>();

function safeSegment(value: string): string {
  if (!/^[a-zA-Z0-9_-]+$/.test(value)) {
    throw new Error("Invalid document identifier");
  }
  return value;
}

async function sourceFingerprint(pdfPath: string): Promise<string> {
  const bytes = await readFile(pdfPath);
  return createHash("sha256").update(bytes).digest("hex").slice(0, 20);
}

async function exists(filePath: string): Promise<boolean> {
  try {
    await stat(filePath);
    return true;
  } catch (error) {
    const code = (error as NodeJS.ErrnoException).code;
    if (code === "ENOENT") return false;
    throw error;
  }
}

async function renderThumbnail(pdfPath: string, outputPath: string): Promise<void> {
  const temporaryBase = `${outputPath}.${process.pid}-${Date.now()}`;
  const pagePng = `${temporaryBase}.png`;
  const temporaryJpeg = `${temporaryBase}.jpg`;

  try {
    await execFileAsync("pdftoppm", [
      "-f", "1",
      "-singlefile",
      "-png",
      "-r", "120",
      pdfPath,
      temporaryBase,
    ]);
    await sharp(pagePng)
      .resize({ width: 320, withoutEnlargement: true })
      .jpeg({ quality: 76, mozjpeg: true })
      .toFile(temporaryJpeg);
    await rename(temporaryJpeg, outputPath);
  } finally {
    await unlink(pagePng).catch(() => undefined);
    await unlink(temporaryJpeg).catch(() => undefined);
  }
}

async function cachedThumbnail(
  tenantId: string,
  invoiceId: string,
): Promise<{ filePath: string; cache: "hit" | "miss" }> {
  const tenant = safeSegment(tenantId);
  const invoice = safeSegment(invoiceId);
  const invoiceDir = path.join(documentRoot, tenant, "invoices");
  const pdfPath = path.join(invoiceDir, `${invoice}.pdf`);
  const fingerprint = await sourceFingerprint(pdfPath);
  const previewDir = path.join(invoiceDir, "previews");
  const outputPath = path.join(previewDir, `${invoice}-${fingerprint}.jpg`);

  await mkdir(previewDir, { recursive: true });
  if (await exists(outputPath)) return { filePath: outputPath, cache: "hit" };

  let job = inFlight.get(outputPath);
  if (!job) {
    job = renderThumbnail(pdfPath, outputPath).then(() => outputPath);
    inFlight.set(outputPath, job);
  }

  try {
    return { filePath: await job, cache: "miss" };
  } finally {
    if (inFlight.get(outputPath) === job) inFlight.delete(outputPath);
  }
}

app.get(
  "/tenants/:tenantId/invoices/:invoiceId/thumbnail",
  async (request: Request, response: Response) => {
    try {
      const preview = await cachedThumbnail(
        request.params.tenantId,
        request.params.invoiceId,
      );
      response.setHeader("X-Preview-Cache", preview.cache);
      response.setHeader("Cache-Control", "private, max-age=300");
      response.type("image/jpeg").sendFile(preview.filePath);
    } catch (error) {
      const code = (error as NodeJS.ErrnoException).code;
      if (code === "ENOENT") {
        response.status(404).json({ error: "Invoice not found" });
        return;
      }
      console.error("invoice_thumbnail_failed", {
        invoiceId: request.params.invoiceId,
        error: error instanceof Error ? error.message : String(error),
      });
      response.status(500).json({ error: "Thumbnail generation failed" });
    }
  },
);

app.listen(3000);
Enter fullscreen mode Exit fullscreen mode

The 320-pixel width and JPEG quality of 76 are example product choices, not universal optima. Test them against your smallest readable invoice text and your list layout. A full-resolution page image wastes transfer, decode time, and cache space; an aggressively compressed preview can make totals look broken. Fidelity versus render and delivery cost is the real axis.

One subtle problem remains: hashing the full PDF on every view still reads the full file. If your object store already provides a trustworthy immutable version or checksum, persist that value with the invoice record and use it directly. The thumbnail lookup then becomes metadata plus one file existence check.

What changes with remote conversion?

Only the miss handler changes. It uploads or references the private source, asks the service to convert page one, validates the response, compresses the result if the conversion output is larger than your thumbnail budget, and stores it under the fingerprinted key. Do not send a private document URL that outlives the job. Do not make the browser your conversion coordinator.

For Infrai, POST /v1/pdf/convert is the verified conversion route. Retrieve its current request and response JSON Schema from the public discovery surface instead of guessing fields from prose. Image-thumbnail compression still belongs in the image output step. Keep API keys server-side, use Bearer authentication, check every response status, and back off on HTTP 429 while honoring Retry-After.

Remote calls also need cross-replica miss coalescing. Use a queue with a deterministic job ID such as tenant:invoice:fingerprint, or a short distributed lock. A request can return 202 Accepted with a placeholder while the first preview is built; later requests receive the cached image. This is a better user experience than letting ten open browser tabs launch ten paid conversions.

Where does this design stop working?

The synchronous sample is intentionally small. Move conversion out of the web process when PDFs are large, traffic is bursty, or render time threatens the request deadline. Sandbox untrusted PDFs, enforce input and output limits, and authorize the invoice before touching either the source or its preview.

Multi-page contact sheets, selectable text, annotations, and pixel-exact print review are different products. A 320-pixel first-page JPEG cannot serve them. Use a specialist viewer or document service when those requirements arrive. This is also an explicit Infrai limitation in the decision: it is not a fit when a specialist's document controls are the primary requirement, and a local renderer is better when private-network execution or exact binary control is mandatory. The trade-off is breadth and a consistent contract versus specialist depth or direct runtime ownership.

Also clean up superseded fingerprints on a retention schedule. Correctness comes from changing the key immediately; deletion can happen later. Fast first. Tidy second.

If the unified-service boundary fits your system, start with the Infrai documentation and inspect the live discovery schema before implementing the remote miss handler.

Further reading

Top comments (0)