Compress an invoice PDF before applying its digital signature, then record both byte counts, their exact difference, and content hashes beside the archive object. If a signed PDF arrives on ingest, retain those original bytes unless a separate signature-validation policy authorizes transformation. That ordering gives a fintech archive smaller generated invoices without quietly replacing evidence an auditor may need.
TL;DR: Treat compression as a state transition with a receipt, not as an opaque upload. The receipt says what entered, what left, and how many physical storage bytes were avoided. Never infer success from a response code or a rounded dashboard percentage.
| Ingest case | Pick this path | Signature consequence | Saving to report |
|---|---|---|---|
| Invoice generated from order data, not yet signed | Compress, validate, then sign | Signature covers the archived representation | Original bytes minus archived bytes |
| Signed invoice received from a counterparty | Preserve the original; validate under archive policy | Existing signed byte ranges remain available | Zero unless a derived copy replaces no record |
| Scanned invoice with no signature requirement | Re-encode images, then render-check | No digital signature to preserve | Original bytes minus archived bytes |
| Output is larger or invalid | Keep the input | No needless mutation | Zero |
Which bytes are the audit record?
A PDF digital signature is tied to byte ranges in a particular file revision. Changing covered bytes after signing can make validation fail, even when every page looks identical. ISO 32000-2 defines PDF, including signatures and incremental updates; ETSI PAdES profiles build PDF signature rules on that foundation. Visual sameness is not evidence equivalence.
Keep the original.
For invoices created from order data, the sequence is concrete: render, compress, validate, sign, archive. Put the immutable order ID and renderer version in protected metadata too. For an inbound signed invoice, keep the received object as the record. A smaller rendition may help search or preview, but label it as derived and never substitute it for the signed original.
This prevents a metrics trap: claiming a saving after retaining both an original and a compressed derivative. Count physical bytes actually retained, including both objects.
Pick the transformation that matches the file
Lossless structural cleanup is the conservative option for digitally generated invoices. It may remove unused objects, consolidate repeated resources, compress eligible streams, and rewrite file structure. The gain depends on how the source was produced, so no fixed percentage is credible before measurement.
Image re-encoding is different. It can shrink scans, but it may soften small tax identifiers, barcodes, or fine print. Define acceptance checks first: render representative pages, extract expected invoice fields, and verify page count. Reject a smaller result when those checks fail.
Preservation is a serious option. Choose it for already-signed documents, encrypted inputs the service is not authorized to open, unsupported constructs, or output that is not smaller. Record signed_original, encrypted_input, validation_failed, or no_gain. These low-cardinality reasons explain operations without putting invoice IDs into metric labels.
How should Node.js compress a PDF on ingest and report its size?
Keep the PDF engine behind a narrow interface. Engines expose different rewrite controls, while accounting and audit behavior should stay stable. This TypeScript accepts an application-supplied compressor, rejects empty output, keeps the original when compression loses, and returns durable metadata. Integer byte counts are the source of truth; basis points are only a display ratio.
import { createHash } from "node:crypto";
type Compressor = (input: Buffer) => Promise<Buffer>;
type Receipt = {
decision: "compressed" | "kept_original";
reason: "smaller_output" | "no_gain";
inputBytes: number;
storedBytes: number;
savedBytes: number;
savingBasisPoints: number;
inputSha256: string;
storedSha256: string;
};
const sha256 = (value: Buffer): string =>
createHash("sha256").update(value).digest("hex");
export async function compressForArchive(
input: Buffer,
compress: Compressor,
): Promise<{ object: Buffer; receipt: Receipt }> {
if (input.length === 0) throw new Error("empty PDF input");
const candidate = await compress(input);
if (candidate.length === 0) throw new Error("empty compressor output");
const useCandidate = candidate.length < input.length;
const object = useCandidate ? candidate : input;
const savedBytes = input.length - object.length;
return {
object,
receipt: {
decision: useCandidate ? "compressed" : "kept_original",
reason: useCandidate ? "smaller_output" : "no_gain",
inputBytes: input.length,
storedBytes: object.length,
savedBytes,
savingBasisPoints: Math.floor((savedBytes * 10_000) / input.length),
inputSha256: sha256(input),
storedSha256: sha256(object),
},
};
}
The function does not claim arbitrary output is valid PDF. Put a validation boundary after transformation and before signing: parse with an independent reader, confirm expected page count, and render selected pages. Then sign the accepted bytes. Persist the receipt atomically with the final archive key and signature status, so an object cannot exist without its accounting record.
Here is a small accounting test. The values are fixtures, not a compression benchmark. A 2,400,000-byte input and 1,800,000-byte candidate yield 600,000 saved bytes and 2,500 basis points, or 25.00% for display.
import assert from "node:assert/strict";
import test from "node:test";
import { compressForArchive } from "./compress-for-archive.js";
test("reports exact retained bytes", async () => {
const input = Buffer.alloc(2_400_000, 1);
const candidate = Buffer.alloc(1_800_000, 2);
const result = await compressForArchive(input, async () => candidate);
assert.equal(result.receipt.decision, "compressed");
assert.equal(result.receipt.savedBytes, 600_000);
assert.equal(result.receipt.savingBasisPoints, 2_500);
assert.equal(result.object, candidate);
});
test("does not store a larger rewrite", async () => {
const input = Buffer.alloc(900, 1);
const result = await compressForArchive(
input,
async () => Buffer.alloc(1_100, 2),
);
assert.equal(result.receipt.reason, "no_gain");
assert.equal(result.receipt.savedBytes, 0);
assert.equal(result.object, input);
});
Do not use floating-point percentages as the ledger. Exact bytes compose across batches; rounded percentages do not. For a daily rollup, calculate sum(input_bytes) - sum(stored_bytes) and divide once by sum(input_bytes). Averaging document percentages gives a tiny invoice the same weight as a 200-page statement.
Measure that.
Make the pipeline observable without leaking invoice data
Emit one completion event per transition. Include receipt fields, duration, outcome, compression policy version, and a trace identifier. Keep customer names, account numbers, invoice numbers, object keys, and hashes out of metric labels. Hashes belong in access-controlled audit metadata; trace identifiers connect logs without turning monitoring into a document index.
Useful metrics are few: total input bytes, total stored bytes, saved bytes, attempts by outcome, and duration. Alert on behavior, not a universal compression target. A sustained rise in validation_failed, a sudden fall in processed documents, or wider duration can expose a renderer or input-mix change. A low saving ratio may be correct when upstream already optimizes PDFs.
Picture the flow: order record enters the renderer; unsigned PDF enters the compressor; candidate and original meet the validation gate; one byte sequence wins; it enters the signer; signed bytes and receipt enter immutable storage; metrics receive counters, never contents. Each arrow has one owner.
Retries need the same discipline. Derive an idempotency key from order version plus policy version, and do not increment retained-byte totals when a completed transition is replayed. Record attempt metrics separately from committed storage metrics. Otherwise a timeout can manufacture savings on a dashboard even though the archive gained nothing.
Verify the archive before trusting the chart
Run corpus tests with generated invoices, scans, multi-page statements, forms, encrypted files, and signed samples. The pass condition is not “opens in one viewer.” Parse with an independent implementation, compare page count, render pages for visual regression, extract required business fields, and validate the final signature under the organization's trust policy.
Test ugly boundaries too: an empty buffer, compressor crash, output one byte larger, concurrent retries, and metadata persistence failure after upload. The last needs reconciliation based on an object key or idempotency record; counting completion before both writes settle corrupts the audit trail.
The limitations are intentional. This approach is not suitable when the archive must retain every inbound file byte-for-byte, when policy forbids derived copies, or when the team cannot operate render and signature validation; choose preservation with no compression in those cases. It also does not choose image resolution, a PDF engine, certificate policy, retention period, or the legal definition of an original. Those are document-class and jurisdiction decisions, and the trade-off must be approved outside this function. The pattern establishes one defensible invariant: archived signed bytes are identifiable, every claimed saving equals bytes actually avoided, and failed optimization never replaces a valid invoice.
Sources
References:
- ISO 32000-2, Portable Document Format: https://www.iso.org/standard/75839.html
- ETSI EN 319 142-1, PAdES digital signatures: https://www.etsi.org/deliver/etsi_en/319100_319199/31914201/01.01.01_60/en_31914201v010101p.pdf
- Node.js Crypto API,
Hash: https://nodejs.org/api/crypto.html#class-hash - OpenTelemetry semantic conventions: https://opentelemetry.io/docs/specs/semconv/
Top comments (0)