Compress on ingest, but only the PDFs that no longer have to prove anything. Use the presence of a digital signature as the gate, not a config flag: in a customer support archive the file that matters is the signed authorization form, and a Node.js worker that re-saves it to report a tidy size saving has destroyed the one property that made it evidence. The unsigned forms — the ones your own service filled in and flattened — are fair game for a lossless pass.
That ordering is the whole design. Decide, then compress, then record both numbers.
The compression ratio is the least interesting number in this pipeline.
Why the audit trail decides what you can compress
A PDF signature covers bytes, not meaning. The signature dictionary carries a /ByteRange array naming the two spans of the file that were hashed — everything except the hex string holding the signature itself. Change one byte inside either span and verification fails. Reordering objects into object streams, dropping an unreferenced font, rewriting the cross-reference table as a cross-reference stream: those are all byte changes. A compressor that is lossless about page content is not lossless about the file.
There is a sanctioned way to touch a signed document, and it has been in the format since long before PDF 2.0 (ISO 32000-2): the incremental update. New objects are appended to the end of the file, the earlier revision stays intact at the head, and a signature taken over that revision still verifies. Appending makes the file bigger, which is exactly backwards from what an archive budget wants. That's the trade-off the spec hands you, and for anything that might end up in a chargeback dispute it's the right one.
Flattening cuts the other way. Flatten an AcroForm and the field appearance streams are merged into the page content while the interactive field objects disappear; the file usually shrinks, and the values stop being machine-readable. "Refund approved: Yes" becomes a picture of text. So capture the field map into an append-only ledger before you flatten, not after — the ledger row, keyed to the SHA-256 of the input, is what an auditor reads two years later when nobody can open the original ticket.
Flatten first, then treat the values as gone.
A minimal ingest path in Node.js: fill, flatten, compress, record
Smallest thing that survives an audit:
import { createHash } from 'node:crypto';
import { readFile, writeFile } from 'node:fs/promises';
import { PDFDocument } from 'pdf-lib';
const SIGNED = /\/ByteRange\s*\[/;
type Ingested = {
sha256In: string;
bytesIn: number;
bytesOut: number;
action: 'archived-verbatim' | 'flattened' | 'kept-original';
};
export async function ingestForm(
path: string,
fields: Record<string, string>,
): Promise<Ingested> {
const input = await readFile(path);
const sha256In = createHash('sha256').update(input).digest('hex');
const base = { sha256In, bytesIn: input.byteLength };
// A signed file is the record copy. Any byte we touch invalidates it.
if (SIGNED.test(input.toString('latin1'))) {
return { ...base, bytesOut: input.byteLength, action: 'archived-verbatim' };
}
const doc = await PDFDocument.load(input);
const form = doc.getForm();
for (const [name, value] of Object.entries(fields)) {
form.getTextField(name).setText(value);
}
form.flatten();
const out = await doc.save({ useObjectStreams: true });
// Rewriting can grow a file. Keep whichever copy is smaller, and say so.
if (out.byteLength >= input.byteLength) {
return { ...base, bytesOut: input.byteLength, action: 'kept-original' };
}
await writeFile(`${path}.flat.pdf`, out);
return { ...base, bytesOut: out.byteLength, action: 'flattened' };
}
The byte scan is deliberately crude. It runs before any parser sees the document, because the thing you must not do is hand a signed file to a library that might helpfully rewrite it. A scan can false-positive when the literal turns up inside a compressed object stream, and a parser-level check on the AcroForm /SigFlags entry is stricter — it just costs you a parse of the exact file you were trying not to parse. I don't think there's a clean way around that ordering; your mileage may vary if your intake is narrow enough to trust.
Two boundaries worth knowing before you wire this up. pdf-lib works on AcroForm fields in pure JavaScript and doesn't re-encode embedded images, so its savings are structural — object streams, flate on new content, the discarded field objects. It also doesn't support dynamic XFA forms, which is less of a gap than it sounds, since XFA is deprecated in PDF 2.0, but it still bites when a bank sends one.
For unsigned scans, where the bytes live in JPEG or JBIG2 image data, a structural pass will barely move the needle. That's the case for a second, out-of-process step:
qpdf --object-streams=generate --compress-streams=y in.pdf out.pdf
Out-of-process matters more than the flags. A hostile or truncated file crashes a child process instead of the worker that's holding your queue lease, and the exit code is a cleaner failure signal than a half-written buffer.
How should you report the size saving from PDF compression on ingest?
Report per class, never as one headline percentage. A blended "we saved 38%" across signed originals, flattened forms and image-heavy scans is arithmetic on three unrelated populations, and it collapses the moment the mix shifts.
| File class | Where the bytes are | What the pass may touch | Honest metric |
|---|---|---|---|
| Signed authorization form | irrelevant | nothing | count and bytes skipped |
| Filled + flattened form | object structure, field objects | rewrite, object streams | median ratio per file |
| Scanned attachment | image streams | structural only, unless a derivative | ratio plus a legibility check |
| Derivative for the agent UI | recoded images | anything | not counted as a saving |
Four rules keep the number defensible. Record bytes_in and bytes_out on every object, including the ones you skipped, so the denominator is the whole intake rather than the subset that compressed well. Report the median ratio next to the total bytes, because one 90 MB scan will drag a mean anywhere you want it to go. Never fold storage deduplication into the compression figure — identical attachments across tickets are a different mechanism and a different budget line. And write the ledger row for every file, one JSON object per line:
{"event":"ingest","ticket":"SUP-10482","sha256_in":"9f2c...","bytes_in":2418907,"bytes_out":2418907,"action":"archived-verbatim","at":"2026-09-12T09:14:22Z"}
That file is both the audit trail and the reporting source. Same rows, two readers.
What I would change at scale
Three things, in order of how much they hurt.
Memory first: readFile plus a latin1 copy means roughly twice the file size resident per job, and a support desk that accepts 100 MB attachments will find the ceiling for you. Read a prefix and the tail for the signature scan, keep the full buffer only on the path that actually rewrites. Then idempotency — key the job on the input SHA-256, so a redelivered message returns the existing ledger row instead of producing a second derivative with a different timestamp. Last, retention: object-lock style immutability on the archive bucket makes the "keep the signed original verbatim" rule enforceable by the storage layer rather than by everyone remembering it.
The catch is that compression may not be the lever worth pulling. Cold storage tiers usually cut the archive bill further than a structural pass will, they don't touch a single byte of the document, and they need no audit story at all. Stick with plain originals and a lifecycle rule when your corpus is mostly scans and your compliance window is measured in years; spend the engineering time on compression only when you hold millions of small unsigned forms, where the structural win is real and the files have already stopped being evidence.
References
- ISO 32000-2, Portable Document Format — https://www.iso.org/standard/75839.html
- ISO 19005-1, PDF/A archival conformance — https://www.iso.org/standard/38920.html
- qpdf command-line manual — https://qpdf.readthedocs.io/en/stable/cli.html
- pdf-lib API documentation — https://pdf-lib.js.org/docs/api/
- Node.js crypto module — https://nodejs.org/api/crypto.html
- Amazon S3 Object Lock — https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lock.html
Top comments (0)