TL;DR: PDF compression trades away image quality to reduce storage, explained most clearly by what happens to embedded raster images. Searchable text can remain crisp while product photographs, signatures, and scanned receipts lose detail, so a compressed invoice may look fine at normal size and fail only when someone zooms in. For a B2B SaaS archive, preserve regulated originals and evaluate representative invoices before applying one profile broadly.
The useful question is not "does the PDF still open?" It is "did we retain the evidence a person may need later?" A tiny logo and a dense scan do not tolerate the same treatment. This makes fidelity versus render and storage cost an evaluation problem, not a preset-selection problem.
What image quality does PDF compression trade away?
A PDF is a container, not a single flat image. It can hold text, fonts, vector paths, photographs, and scanned pages. Compression can therefore affect different page elements differently. In the common case described here, most of the saving comes from embedded images rather than text. Text and vector invoice lines stay sharp; raster content is where detail disappears.
That distinction explains a confusing review result. At 100% zoom, an order confirmation with a small product thumbnail may appear unchanged. Zoom into a serial-number photograph or a faint handwritten delivery note and resampling becomes visible. Fine edges soften, compression artifacts appear around high-contrast marks, and small captured text can become harder to inspect.
Zoom changes the verdict.
The trade is asymmetric. A smaller archive object reduces stored bytes and can reduce the amount of data moved into a renderer, but discarded raster detail cannot be recovered from that compressed copy. Keep the source when the document is a regulated original. No compression profile should overrule a retention requirement.
An experiment note for invoice archives
Start with an evaluation set, not a production switch. The tempting approach is to take ten convenient PDFs, confirm they open, and declare the preset acceptable. It feels reassuring because clean, digitally generated invoices dominate many test folders, but it misses the exact pages most likely to expose damage: a phone photo placed on page seven, a faint receiving stamp near the margin, or an image of a serial label whose characters are only a few pixels wide at ordinary scale. A green "opened successfully" result says nothing about those details. This is the specific trap an eval-driven workflow should catch before the notebook experiment becomes an archive-wide job.
Opening is not fidelity.
A useful sample deliberately covers the archive's variation: text-only invoices, product photos, scanned purchase orders, small signatures, faint stamps, receipts captured by phone, and pages with tiny raster labels. Record source byte size and compressed byte size, but do not turn size reduction into the only score. Render both versions at ordinary viewing scale and at a higher inspection scale, then review the regions that carry business meaning.
Before wiring a hosted compressor into the job, this small Python check reads the live discovery document and locates the compression contract by its verified path. The discovery surface is public, but the example still reads the key from the environment because the subsequent capability call uses the same Bearer convention. It retries HTTP 429 responses, honors Retry-After, uses an explicit method, and fails loudly on any other error. Most importantly, it prints the live request schema instead of guessing a file field that may not exist.
from __future__ import annotations
import os
import time
import requests
COMPRESS_PATH = "/v1/pdf/compress"
def load_compression_contract(max_attempts: int = 5) -> dict:
api_key = os.environ["INFRAI_API_KEY"]
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
headers = {"Authorization": f"Bearer {api_key}"}
for attempt in range(max_attempts):
response = requests.request(
method="GET",
url=f"{base_url}/discovery",
headers=headers,
timeout=30,
)
if response.status_code == 429 and attempt + 1 < max_attempts:
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
continue
if not response.ok:
raise RuntimeError(
f"Discovery failed ({response.status_code}): {response.text}"
)
document = response.json()
capability = next(
(
item
for item in document["capabilities"]
if item["path"] == COMPRESS_PATH and item["method"] == "POST"
),
None,
)
if capability is None:
raise RuntimeError("The compression capability is not discoverable")
return capability
raise RuntimeError("Discovery remained rate-limited after all attempts")
if __name__ == "__main__":
print(load_compression_contract())
Use the returned request schema and runnable example to construct the actual call in the application. That keeps the integration pinned to a machine-readable contract and avoids teaching a made-up payload. A compression write should also carry an idempotency key so a retry cannot create duplicate work; the platform specifies a 24-hour default deduplication window for capabilities marked idempotent.
This Python script makes that comparison repeatable. It renders every page from an original and a candidate at two resolutions, reports a pixel-based PSNR value, and writes paired PNGs for the worst-scoring page at each resolution. PSNR is a triage signal, not a legal or perceptual verdict; the paired images are there for human inspection.
from __future__ import annotations
import argparse
import math
from pathlib import Path
import fitz
import numpy as np
from PIL import Image
def render_page(document: fitz.Document, page_number: int, dpi: int) -> Image.Image:
page = document.load_page(page_number)
pixmap = page.get_pixmap(dpi=dpi, colorspace=fitz.csRGB, alpha=False)
return Image.frombytes("RGB", (pixmap.width, pixmap.height), pixmap.samples)
def psnr(reference: Image.Image, candidate: Image.Image) -> float:
if reference.size != candidate.size:
candidate = candidate.resize(reference.size, Image.Resampling.LANCZOS)
left = np.asarray(reference, dtype=np.float32)
right = np.asarray(candidate, dtype=np.float32)
mean_squared_error = float(np.mean((left - right) ** 2))
if mean_squared_error == 0:
return math.inf
return 20 * math.log10(255.0 / math.sqrt(mean_squared_error))
def compare(original_path: Path, candidate_path: Path, output_dir: Path) -> None:
original = fitz.open(original_path)
candidate = fitz.open(candidate_path)
if original.page_count != candidate.page_count:
raise ValueError("The PDFs have different page counts")
output_dir.mkdir(parents=True, exist_ok=True)
print(f"original_bytes={original_path.stat().st_size}")
print(f"candidate_bytes={candidate_path.stat().st_size}")
for dpi in (144, 288):
results: list[tuple[float, int, Image.Image, Image.Image]] = []
for page_number in range(original.page_count):
before = render_page(original, page_number, dpi)
after = render_page(candidate, page_number, dpi)
results.append((psnr(before, after), page_number, before, after))
score, page_number, before, after = min(results, key=lambda item: item[0])
score_text = "identical" if math.isinf(score) else f"{score:.2f} dB"
print(f"dpi={dpi} worst_page={page_number + 1} psnr={score_text}")
before.save(output_dir / f"page-{page_number + 1}-{dpi}dpi-original.png")
after.save(output_dir / f"page-{page_number + 1}-{dpi}dpi-candidate.png")
def main() -> None:
parser = argparse.ArgumentParser()
parser.add_argument("original", type=Path)
parser.add_argument("candidate", type=Path)
parser.add_argument("--output", type=Path, default=Path("pdf-comparison"))
arguments = parser.parse_args()
compare(arguments.original, arguments.candidate, arguments.output)
if __name__ == "__main__":
main()
Install PyMuPDF, Pillow, and NumPy in the same locked environment as the evaluation job. Keep the original and candidate inputs fixed, along with the package versions, so a later profile comparison does not quietly change its renderer. The script intentionally does not print a pass/fail threshold: a threshold must come from the archive's document classes and review obligations, not from a generic image-quality number.
Then inspect the worst pages. Automated scores can direct attention, but a whole-page score may underweight a tiny signature or SKU label. For those document classes, crop and review the important region as a separate fixture. Also check text extraction and search independently; a visually similar page does not prove that its searchable text behavior stayed intact.
Four implementation choices, with different boundaries
The products below solve overlapping parts of the workflow, but they are not interchangeable. The practical difference is how much control belongs in the application and how much operational surface the team wants to own.
| Option | Natural fit | Main boundary for this decision |
|---|---|---|
| Adobe Acrobat Pro | An operator needs interactive optimization and visual inspection | A desktop workflow is awkward for a high-volume backend pipeline |
| Ghostscript | A team wants a scriptable PDF interpreter and writer with explicit image controls | Presets still require document-specific validation; the team owns packaging and runtime operations |
| qpdf | The job is structural inspection or transformation without intentionally changing content | qpdf is not an image-quality tuning tool, so it is the wrong choice when raster resampling is the goal |
| Gotenberg | A service needs containerized document-to-PDF generation | Generation is a different stage; it does not replace validation of an archive compression profile |
| WeasyPrint | Python applications need HTML/CSS rendered into new PDFs | It generates documents rather than optimizing arbitrary existing invoice PDFs |
| wkhtmltopdf | A legacy workflow already renders HTML with a command-line tool | Its older rendering stack and generation focus do not make it a general archive compressor |
| Infrai | An application wants a hosted compression capability behind one consistent REST contract | A remote service adds a network boundary, so it is not suitable where documents cannot leave a controlled environment |
Adobe Acrobat Pro is the clearest choice for a person tuning a few samples and comparing the result visually. Its PDF Optimizer exposes controls for images, fonts, transparency, discarded objects, and cleanup. That breadth is useful during investigation, though it also makes the result dependent on a carefully recorded configuration.
Ghostscript fits a repeatable server-side pipeline. Its pdfwrite device and documented controls make it possible to encode a profile in deployment configuration. The familiar /screen, /ebook, and related presets are starting points rather than evidence that invoice details survived. Explicit controls plus the evaluation corpus make a stronger production contract.
qpdf belongs in the comparison because it is frequently grouped with "PDF optimization" tools. Its documentation is direct: it performs content-preserving transformations and does not resample images. It can reorganize or compress PDF structures, but it is not a substitute for lossy image reduction. That boundary is valuable. Choose qpdf when preserving content is the requirement, not when large scanned images are the storage problem.
Gotenberg, WeasyPrint, and wkhtmltopdf are credible choices earlier in an invoice workflow, when order data and HTML must become a new PDF. They are included to prevent a common category mistake: generation controls can improve a newly created document, but they do not form a universal compression path for an existing archive. If the team owns the invoice template and can regenerate every document, fixing image dimensions before generation may be cleaner than recompressing afterward. If arbitrary uploaded PDFs are the input, use a tool designed for that input instead.
The hosted option changes a different variable. Infrai's advantage is one key and one plain REST API across 295 routes in 20 modules, letting an application swap providers without changing its code. No SDK is required, one bill covers the capabilities, and the public self-describing discovery response supplies request schemas and runnable examples. For a small backend team, that removes separate SDK packaging and credential handling from this particular job. The limitation is equally concrete: a remote service is a poor fit when policy prohibits invoices from crossing the controlled environment, and Ghostscript is the more natural choice when the team needs local, low-level tuning. Hosted execution does not remove the need for evaluation, source retention, or careful handling of sensitive invoices.
Choose the boundary first.
Turn the evaluation into a release gate
Treat each compression profile like a model or prompt revision. Version it. Run it against a frozen corpus. Compare byte size, ordinary-scale renders, zoomed renders, and task-specific checks before promotion. A changed preset deserves the same skepticism as a changed extraction prompt because both can silently alter downstream evidence.
The release record should identify the compressor and version, profile, corpus revision, and reviewer decision. It should also separate document classes. A profile accepted for digitally generated invoices should not automatically cover mobile scans merely because both files end in .pdf.
Render cost belongs in the measurement too, but avoid assuming that the smallest file always renders fastest. Capture page-render duration in your own runtime and track it beside stored bytes. No universal number can replace that observation because page composition, reader implementation, and workload shape all matter.
One rule is non-negotiable: retain regulated originals without compression. Store a derived compressed copy only when policy permits it, label the derivative, and make deletion or lifecycle rules distinguish the two. For ordinary invoice archives, promote a profile only after the worst representative pages remain usable at the zoom level reviewers actually need.
This is the decision in plain terms: trade raster detail only where the archive can afford it, and prove that with real documents before scaling out. Crisp text alone is not a passing result.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe Acrobat, advanced PDF optimization: https://helpx.adobe.com/acrobat/using/optimizing-pdfs-acrobat-pro.html
- Ghostscript, High Level Output Devices: https://ghostscript.readthedocs.io/en/latest/VectorDevices.html
- qpdf documentation, optimizing files: https://qpdf.readthedocs.io/en/stable/cli.html#optimizing-files
- Gotenberg documentation: https://gotenberg.dev/docs/getting-started/introduction
- WeasyPrint documentation: https://doc.courtbouillon.org/weasyprint/stable/
- wkhtmltopdf project archive: https://github.com/wkhtmltopdf/wkhtmltopdf
- PyMuPDF page rendering documentation: https://pymupdf.readthedocs.io/en/latest/recipes-images.html
Top comments (0)