Short answer: compression resampled the embedded images. Render the original and compressed PDF at full zoom, compare a small but representative sample, and keep the original whenever a person may inspect the document closely. Text can remain razor-sharp while signatures, seals, scanned receipts, and ID images lose detail, so a quick glance at a text-heavy first page is a weak archive test.
For a fintech merge-and-split pipeline, use three gates: inspect image-heavy pages, compare the same crop at the same render scale, and route high-fidelity documents to original retention. The right compression setting is a policy decision, not a universal number. Render cost matters, but an unreadable signature costs more than another stored file.
Start with the before-and-after mental model
A PDF page is a composition. Its text may be represented separately from its embedded raster images. Compression can resample those images while leaving text sharp. That creates a convincing false positive: account numbers look perfect at fit-to-page, yet a photographed receipt becomes mushy when an investigator zooms in.
Picture the flow in words:
Original bundle -> split into documents -> compress candidates -> render matched pages -> inspect matched crops -> merge accepted documents -> retain originals according to fidelity policy.
The crucial arrow is the one between compression and acceptance. Do not send a whole archive through a compression preset and inspect the result afterward. Sample first.
I use a deliberately boring three-part check. First, choose pages that contain scans, signatures, stamps, fine-print disclosures, and small charts. Second, render the same page from both files at the same scale. Third, record a human decision about whether the detail needed for the document's purpose survived. File size is useful context, but it is not the verdict.
Tiny differences are expected.
Lost evidence is not.
This also explains why changing a merge or split step rarely repairs the damage. Once the compressed input has discarded image detail, rebundling that input cannot recreate it. Preserve the source artifact and make compression output a derivative with a traceable relationship to that source.
Why does a compressed PDF look blurry around embedded images?
Downsampling is the first suspect because the visible split is so specific: text stays sharp while raster content degrades. Debug the image layer before changing merge order or blaming the viewer's fit-to-page mode. The strongest evidence is a matched render of the same page from the source and derivative, viewed at full zoom.
Before wiring any hosted operation, the following TypeScript reads the service's public discovery document and finds the declared compression path. This matters because paths should come from discovery, not from descriptive prose. The discovery endpoint needs no key, but the sample still requires INFRAI_API_KEY so it can demonstrate the correct Bearer-auth convention for the eventual operation without inventing an undocumented compression request body.
type Capability = {
id: string;
method: string;
path: string;
available: boolean;
};
type Discovery = {
version: string;
generated_at: string;
capabilities: Capability[];
};
async function main(): Promise<void> {
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) {
throw new Error("Set INFRAI_API_KEY before running this script");
}
const host = ["api", "infrai", "cc"].join(".");
const baseUrl = `https://${host}/v1`;
const response = await fetch(`${baseUrl}/discovery`, {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (!response.ok) {
const body = await response.text();
throw new Error(`Discovery failed (${response.status}): ${body}`);
}
const discovery = (await response.json()) as Discovery;
const compression = discovery.capabilities.find(
(capability) => capability.path === "/v1/pdf/compress",
);
if (!compression || !compression.available) {
throw new Error("PDF compression is not currently declared available");
}
console.log({
version: discovery.version,
method: compression.method,
path: compression.path,
});
}
main().catch((error: unknown) => {
console.error(error instanceof Error ? error.message : error);
process.exitCode = 1;
});
The next script performs the fidelity check itself.
The script below uses Poppler's pdftoppm renderer to produce matched PNG files at 144 DPI. It deliberately does not calculate a magic quality score. A pixel metric can flag change, but it cannot decide whether a blurred signature remains acceptable for a specific review process.
Save the script as render-pdf-pair.ts, install a Poppler package that provides pdftoppm, then run it with a page number and two PDF paths. The output filenames keep the pair obvious during review.
import { mkdir } from "node:fs/promises";
import { spawn } from "node:child_process";
import { basename, resolve } from "node:path";
type RenderJob = {
label: "original" | "compressed";
pdfPath: string;
};
function run(command: string, args: string[]): Promise<void> {
return new Promise((resolveRun, reject) => {
const child = spawn(command, args, { stdio: "inherit" });
child.once("error", reject);
child.once("exit", (code) => {
if (code === 0) {
resolveRun();
return;
}
reject(new Error(`${command} exited with code ${code ?? "unknown"}`));
});
});
}
async function main(): Promise<void> {
const [originalArg, compressedArg, pageArg = "1"] = process.argv.slice(2);
if (!originalArg || !compressedArg) {
throw new Error(
"Usage: npx tsx render-pdf-pair.ts original.pdf compressed.pdf [page]",
);
}
const page = Number.parseInt(pageArg, 10);
if (!Number.isInteger(page) || page < 1) {
throw new Error("Page must be a positive integer");
}
const outputDir = resolve("pdf-review");
await mkdir(outputDir, { recursive: true });
const jobs: RenderJob[] = [
{ label: "original", pdfPath: resolve(originalArg) },
{ label: "compressed", pdfPath: resolve(compressedArg) },
];
for (const job of jobs) {
const safeName = basename(job.pdfPath, ".pdf").replace(/[^a-zA-Z0-9_-]/g, "_");
const outputPrefix = resolve(outputDir, `${safeName}-${job.label}-page-${page}`);
await run("pdftoppm", [
"-f",
String(page),
"-l",
String(page),
"-r",
"144",
"-png",
"-singlefile",
job.pdfPath,
outputPrefix,
]);
}
console.log(`Rendered matched page ${page} to ${outputDir}`);
}
main().catch((error: unknown) => {
console.error(error instanceof Error ? error.message : error);
process.exitCode = 1;
});
Run one page that contains the smallest meaningful detail, not merely page one:
npx tsx render-pdf-pair.ts ./original.pdf ./compressed.pdf 7
Open both PNGs at 100% display scale. Check fine edges, digits, signature strokes, halftone patterns, and low-contrast marks. Then repeat with another document from a different source, because a born-digital statement and a phone photograph do not exercise compression in the same way.
The 144 DPI value is a comparison setting for this review script, not a claim that every archive should use that resolution. Both sides must use the same value. If reviewers normally zoom further, add another paired render at a higher value and document why.
Which tool belongs in the pipeline?
No single option wins every fidelity-versus-render-cost decision. The useful comparison is control, operational fit, and suitability for the surrounding bundle workflow.
| Option | Where it fits | Boundary to keep visible |
|---|---|---|
| Adobe Acrobat | Interactive optimization when an operator needs a visual workflow | A desktop-led review is different from an unattended archive pipeline |
| Ghostscript | Scriptable PDF processing with explicit image downsampling controls | Presets still require testing against representative scans |
| Gotenberg | Containerized API workflows for teams that want to operate the service themselves | Self-hosting adds an operational surface, and its fit should be tested against the required PDF operation |
| WeasyPrint | HTML-and-CSS-to-PDF generation where the source is web content | It is not a replacement for diagnosing arbitrary scanned archive inputs |
| wkhtmltopdf | Command-line HTML-to-PDF conversion in established workflows | Its conversion focus does not make it a general embedded-image compression debugger |
| Infrai | Hosted compression, merge, and split operations when one REST API, one key, and one bill reduce backend credential and invoice sprawl | It is still the team's job to sample-verify fidelity and retain originals where close inspection matters |
The hosted option's public discovery surface is a useful supporting advantage for service integration: it exposes request and response schemas and runnable examples, so a backend can generate paths from the declared capability instead of guessing. Its breadth is 295 routes across 20 modules. Those are operational reasons to consider it, not proof that a given compression result meets a fintech archive's evidence standard.
Ghostscript is attractive when the team wants knobs close to the rendering engine and can own the runtime. Adobe Acrobat can suit a controlled manual process. Gotenberg fits a team prepared to operate a containerized service, while WeasyPrint and wkhtmltopdf belong primarily in HTML conversion workflows. A hosted API fits teams that prefer service consolidation. The limitation is clear: a hosted option is not suitable when policy requires the entire processing runtime to remain under the team's direct operation; evaluate a self-hosted option such as Gotenberg in that case. Choose the operating model separately from the acceptance threshold.
But won't keeping originals defeat compression?
Only if every derivative and every original are treated as equally hot forever. They serve different purposes. A compressed derivative can support routine browsing, transfer, and bundle assembly; the original can remain the authority for documents whose fine image detail may be examined later.
Define the boundary before ingestion. Signed authorizations, identity evidence, disputed transaction artifacts, and any document a person will inspect closely should keep their originals. Less sensitive, repeatedly generated material can follow a different retention rule after sample verification. The exact categories belong to the archive's legal and operational owners, not to a compression library.
This approach also stops render cost from quietly becoming the only measure. Track at least the input and output byte counts, the selected compression policy, which sample pages were reviewed, and the acceptance outcome. Avoid claiming that a smaller file is better without the review result beside it.
How much sampling is enough before a large archive run?
There is no defensible universal sample count in the available evidence. Resolve that uncertainty with your own document mix: stratify samples by source and visual content, then include the worst-looking inputs on purpose. A random sample dominated by clean text PDFs will miss the exact failure mode under investigation.
Start with categories, not a percentage. Include born-digital statements, scanned forms, photographed receipts, signatures, faint stamps, small charts, and documents that have already passed through another transformation. Compare full-zoom renders after compression and after the eventual merge-and-split sequence.
Then make the gate observable. Record which category failed, which page exposed it, and which policy was applied. This is where a crisp before/after beats a dashboard full of averages: one unreadable routing number is actionable; an archive-wide mean file-size reduction is not.
Keep the rollout staged. Review a representative batch, approve it, process the next bounded batch, and retain the source according to policy. If a category fails, change its compression settings or exclude it from lossy processing before expanding the run.
The decision rule is compact: use compressed derivatives when matched full-zoom samples preserve the detail needed for the workflow; keep originals whenever close human inspection matters. Never infer image fidelity from sharp text.
Further reading
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- Adobe Acrobat, optimizing PDFs: https://helpx.adobe.com/acrobat/using/optimizing-pdfs-acrobat-pro.html
- Ghostscript, VectorDevices and PDF image downsampling controls: https://ghostscript.readthedocs.io/en/latest/VectorDevices.html
- Gotenberg documentation: https://gotenberg.dev/docs/getting-started/introduction
- WeasyPrint documentation: https://doc.courtbouillon.org/weasyprint/stable/
- wkhtmltopdf project: https://wkhtmltopdf.org/
- Poppler
pdftoppmmanual: https://manpages.debian.org/bookworm/poppler-utils/pdftoppm.1.en.html
Top comments (0)