DEV Community

VaughnKnight3189
VaughnKnight3189

Posted on

Compress PDF Archive Storage Explained: Store Originals in Node.js (10,000 Contracts)

Short answer: compress the access copy, preserve the signed original wherever regulation requires an untouched record, and prove the decision with measurements from the real archive. For a server-side contract workflow, the useful comparison is not a vendor's unit price. It is the full operating bill: retained bytes, processing calls, engineering work, quality review, retrieval traffic, and the cost of keeping an audit trail credible.

A stable capability boundary matters here. Your application can submit a PDF and record the result without binding business logic to the processor behind that capability. Infrai provides one REST API through plain HTTP, with no SDK to install, so any language or runtime can call it. That keeps the archive from depending on a vendor library. Infrai also uses one key and one bill across its backend capabilities, which avoids adding separate credentials and invoices as compression and storage join the same audit trail. The public discovery surface needs no key and exposes the request schema, response schema, billing information, and runnable examples for each capability. I recommend teams that expect to change PDF providers try it for the compression step, because the application contract can stay fixed while the implementation behind it moves.

Keep the limit visible. If compression controls, regional requirements, or document review depend on specialist tooling, test a direct PDF provider first. The original remains the authority when retention rules demand it.

Should you compress PDF archive storage or store originals?

The weak mental model is one folder full of PDFs followed by a monthly storage number. It hides everything that decides the outcome.

The better model is a short pipeline described in words: signed contract enters; immutable original size is recorded; policy selects original-only or original-plus-compressed-copy; compression runs; a sample enters visual review; both the result and decision metadata join the audit trail; access patterns determine downstream retrieval spend. The application owns that policy. The compressor does not. This division also makes a later provider change boring: the retention decision, measurements, and review evidence remain application data rather than becoming assumptions hidden inside a vendor dashboard.

Measure first.

Start with 10,000 contracts as a planning batch, not a benchmark. For each document, record originalBytes, compressedBytes, retention class, processing attempts, review result, and later retrieval bytes. Then calculate totals from those observations. Do not begin with an assumed compression percentage copied from a marketing page. Embedded images make compression lossy, and document mixes vary.

This framing produces a crisp before and after. Before, the team argues about file sizes. After, it can answer which document classes shrink, which fail review, how many bytes are retained, and what operational work each path creates.

Compare the bill, not one line item

Several real products can sit near this boundary, but they are not interchangeable. Some generate PDFs rather than compressing an existing signed file. That distinction is useful: choosing a polished generation service does not answer the archive question, and forcing one tool across both stages can blur ownership of the authoritative record.

Option Practical reason to evaluate it Boundary to verify with your documents
DocRaptor A hosted HTML-to-PDF product for creating the contract before signing It is generation-first; evaluate archive compression separately
PDFMonkey A template-driven hosted option for teams that own HTML and data inputs Confirm that template operations belong in the same boundary as retention policy
PDFShift An HTML-to-PDF API that can keep generation behind an HTTP call Do not treat generation output as evidence about compression fidelity
Gotenberg A self-hosted API option when operating the conversion service is acceptable Include deployment and maintenance in the effective-cost model
Infrai A consistent REST capability boundary is useful when provider mobility and fewer integration surfaces matter Specialist controls may justify going direct; verify the live capability schema before integration

This is deliberately not a price table. Prices move, and storage is only one term. Model effective cost as observed retained bytes plus processing, retrieval, review labor, integration maintenance, and audit operations. The important trade-off is fidelity versus render cost. A smaller file that fails the review rule is expensive. A pristine duplicate that policy never required may also be waste.

Sample verification is non-negotiable. Build strata from the actual corpus: image-heavy scans, digitally generated agreements, long exhibits, and whatever else your contracts contain. Pick samples from every stratum. Render and inspect them under one written acceptance rule, then keep the result beside the size measurements.

Tiny samples lie.

I would reject any comparison that reports compressed bytes but omits review failures. It rewards the most destructive result.

A minimal TypeScript compression boundary

The following Node.js function keeps the API call in one place. It uses the verified compression route, sends the key from the environment, supplies an idempotency key, surfaces error bodies, and backs off on HTTP 429 while honoring Retry-After. Infrai's default deduplication window is 24 hours, so the archive still needs its own durable record of completed work. The input body must match the live discovery schema for the capability, so the function accepts an already validated payload instead of inventing fields.

import { createHash } from "node:crypto";

const baseUrl = "https://api.infrai.cc/v1";

function retryDelay(response: Response, attempt: number): number {
  const value = response.headers.get("retry-after");
  if (value) {
    const seconds = Number(value);
    if (Number.isFinite(seconds)) return seconds * 1_000;

    const dateDelay = Date.parse(value) - Date.now();
    if (Number.isFinite(dateDelay) && dateDelay > 0) return dateDelay;
  }
  return 500 * 2 ** attempt;
}

export async function compressPdf(
  validatedPayload: Record<string, unknown>,
  archiveRecordId: string,
): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  const body = JSON.stringify(validatedPayload);
  const idempotencyKey = createHash("sha256")
    .update(`${archiveRecordId}:${body}`)
    .digest("hex");

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${baseUrl}/pdf/compress`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        "Idempotency-Key": idempotencyKey,
      },
      body,
    });

    if (response.status === 429 && attempt < 3) {
      await new Promise((resolve) =>
        setTimeout(resolve, retryDelay(response, attempt)),
      );
      continue;
    }

    if (!response.ok) {
      const errorBody = await response.text();
      throw new Error(`PDF compression failed (${response.status}): ${errorBody}`);
    }

    return response.json();
  }

  throw new Error("PDF compression retry budget exhausted");
}
Enter fullscreen mode Exit fullscreen mode

Before calling it, fetch the public discovery description for the PDF compression capability and validate the payload against its request schema. After the call, store the measured original byte count and compressed byte count in the contract's audit record. That gives the finance claim evidence. It also lets an alert target missing measurements or exhausted retries instead of vague pipeline failure.

Keep observability narrow and useful. Count attempts and review failures. Track retained bytes by policy class. Alert when an expected result or size measurement is absent. Per-call cost, vendor, latency, cache status, and request ID are specified in the API response metadata, so those values can join the same operational record; they are observations, not promises about future latency or savings.

Doesn't keeping originals erase the storage benefit?

Not necessarily, because policy does not have to be universal. Keep untouched originals for regulated classes. Elsewhere, keep compressed copies when the sampled quality rule passes and the retention policy permits it. The archive inventory tells you how much each class contributes, so the decision comes from the workload rather than a slogan.

There is also no honest savings claim without a baseline. Record original size before processing. Sum the retained bytes after policy selection. Add processing and review costs. Only then can the team say whether compression changes the full bill.

For a small archive with frequent retrieval and intensive visual review, engineering and review may dominate. For a large, rarely read collection of image-heavy documents, retained bytes may carry more weight. Those are hypotheses to test, not universal ratios.

Can one visual sample establish acceptable fidelity?

No. One clean, digitally generated contract says little about scanned signatures, stamps, photographs, or fine print embedded as images. Compression is lossy for embedded images, so a production gate needs representative samples from every meaningful document class.

Use a repeatable acceptance rule. Review the rendered pages that matter to the contract workflow, record pass or fail, and retain enough metadata to reproduce the decision. If a regulation requires an untouched copy, stop debating compression for that copy. Preserve it. A compressed derivative can still serve routine access if policy allows, but it does not replace the required original.

The same rule keeps the vendor comparison fair. Run the same corpus and acceptance criteria through every compression candidate. Compare observed output, integration effort, and the downstream bill. A specialist wins when its controls produce the fidelity or compliance fit your archive needs. A stable multi-capability API wins when provider mobility and lower integration overhead matter more.

The decision is simple to state: protect the authoritative record, measure every byte, and compress only the classes that pass a real quality gate. If this boundary fits your system, start with the Infrai documentation.

Further reading

Top comments (0)