A monthly game report is rarely one document. Finance owns the revenue pages, live operations owns event results, and player support owns moderation summaries. If those teams must keep editing or replacing their parts, deliver separate files in a ZIP. If the archive is a fixed review artifact that people read in sequence, assemble one PDF. Template ownership is the deciding constraint.
TL;DR: keep independently owned templates independent through rendering and validation. Merge only the accepted PDF outputs, and retain a manifest either way. A ZIP preserves replacement boundaries; one PDF preserves reading order and gives the archive one primary object. Neither container repairs a bad source template.
Should you merge one PDF or deliver separate files in a ZIP?
The initial requirement sounds like a file-format choice: one PDF or many PDFs inside a ZIP. For a monthly game report, the real boundary sits earlier. Who can change each template, and when does that output become immutable?
Consider a packet with an executive summary, store revenue, daily active-user charts, live-event results, and a trust-and-safety appendix. Those sections do not necessarily share a release cadence or approval chain. Combining their source templates into one large template couples every owner to a common deployment. A CSS fix in the appendix can then force a rerender of revenue pages that were already approved.
That is config bloat wearing a document badge.
The opposite choice has a cost. A ZIP makes the recipient open several files, and the order may be unclear outside the extraction tool. It also lets a consumer silently omit one member when forwarding the report. A merged PDF carries a visible page sequence and can use PDF outlines for navigation, but replacement becomes an assembly operation instead of a file swap.
No option wins outright.
I use a narrow rule: preserve the ownership boundary until every section passes its own checks. After that point, choose the delivery container from the archive's retrieval pattern. A packet that is normally read cover to appendix becomes one PDF. A packet whose sections are independently requested, retained, or replaced stays separate in a ZIP. If both patterns matter, store the validated components and derive the merged reading copy from them. Do not maintain two hand-built versions. The limitation of that derived-copy approach is extra storage and another artifact to verify; the downside of keeping only the merged output is that a later section correction requires reconstructing the packet from some other source. This is an explicit trade-off between operational simplicity now and section-level recovery later.
Build the smallest useful packet
The implementation needs fewer knobs than most document pipelines acquire. Each section renderer returns bytes plus stable metadata. The packet builder checks that every expected section arrived exactly once, computes a digest over each output, writes a manifest, then either concatenates the PDFs or packages the original PDFs with that manifest.
The interface below deliberately does not prescribe a rendering engine, object store, queue, or ZIP library. Those are replaceable. The contract is the useful part.
type SectionId =
| "summary"
| "revenue"
| "engagement"
| "live-ops"
| "safety";
type RenderedSection = {
id: SectionId;
owner: string;
templateVersion: string;
pdf: Uint8Array;
sha256: string;
pageCount: number;
};
type PacketManifest = {
reportMonth: string;
gameId: string;
sections: Array<Omit<RenderedSection, "pdf">>;
};
interface PacketIO {
mergePdfs(parts: Uint8Array[]): Promise<Uint8Array>;
zip(files: Array<{ name: string; bytes: Uint8Array }>): Promise<Uint8Array>;
}
async function buildPacket(
sections: RenderedSection[],
mode: "reading-copy" | "owned-parts",
manifest: PacketManifest,
io: PacketIO,
): Promise<{ extension: "pdf" | "zip"; bytes: Uint8Array }> {
const order: SectionId[] = [
"summary",
"revenue",
"engagement",
"live-ops",
"safety",
];
const byId = new Map(sections.map((section) => [section.id, section]));
if (byId.size !== order.length || order.some((id) => !byId.has(id))) {
throw new Error("Packet is incomplete or contains duplicate section IDs");
}
const sorted = order.map((id) => byId.get(id)!);
if (mode === "reading-copy") {
return { extension: "pdf", bytes: await io.mergePdfs(sorted.map((s) => s.pdf)) };
}
const files = sorted.map((section, index) => ({
name: `${String(index + 1).padStart(2, "0")}-${section.id}.pdf`,
bytes: section.pdf,
}));
files.push({
name: "manifest.json",
bytes: new TextEncoder().encode(JSON.stringify(manifest, null, 2)),
});
return { extension: "zip", bytes: await io.zip(files) };
}
The numeric prefixes are mundane and intentional. ZIP readers may display members differently, while filenames remain a cheap, inspectable ordering signal. The manifest should record the same canonical order, each template version, owner, page count, and SHA-256 digest. Web Crypto defines digest operations for SHA-256, so this does not require inventing a checksum scheme.
Boring is good here.
Do not confuse a digest with a signature. A digest catches accidental mismatch when compared with a trusted manifest. It does not prove who created the packet. If authorship or tamper evidence is a requirement, define a signing and verification process separately and test it with the eventual archive system.
Validate the document, not just the request
A successful render call proves very little. The useful checks happen on the artifact. Start with structural checks: every required section exists, bytes parse as a PDF, page counts are plausible, and the merged result has the sum of the component page counts. Confirm the expected page size and orientation for every section before assembly. A single landscape chart in a portrait packet may be valid, but it should be a declared choice.
Then test meaning. Extract text from representative pages and assert the report month, game identifier, section heading, and a few source totals. Rasterize selected pages and compare them with reviewed images using tolerances that survive harmless renderer differences. Visual snapshots are valuable around charts, font fallback, clipping, and page breaks. They are noisy when treated as exact byte comparisons.
Accessibility also changes the container decision. PDF 2.0 is standardized by ISO 32000-2, but conformance to a file-format standard does not guarantee a usable reading order, meaningful headings, alternative text, or accessible form semantics. PDF/UA-2, published as ISO 14289-2, defines accessibility requirements for PDF based on PDF 2.0. A merged packet can expose one navigation tree, yet it can also destroy good structure during concatenation if the assembler does not preserve tags. Test the final artifact with the same seriousness as each component.
ZIP has its own boring failure mode: paths. Reject absolute paths and parent-directory segments when accepting or reprocessing archives. Keep filenames ASCII and deterministic. Set explicit limits for member count and uncompressed size before extraction, because compressed input can expand far beyond its stored size. Five known PDFs and one manifest make a tight allowlist for this report; there is no reason to accept arbitrary members.
The acceptance suite should run against golden input data, not production data copied into a fixture. A synthetic month with zero revenue, a very long event name, missing optional notes, and a chart containing both negative and positive values exercises more layout behavior than a cheerful average month. Short tests are fine. Vague tests are not.
Measure the pipeline at its real boundaries
Benchmarks should separate rendering, assembly, upload, and retrieval. One end-to-end timer hides the component you can improve. Record section render duration by template version, total input bytes, output bytes, page count, merge or archive duration, and upload duration. Avoid player identifiers and report contents in telemetry.
Run at least three fixture shapes: a small packet, the expected monthly packet, and a deliberately heavy packet with long tables and dense charts. Report distributions rather than one best run. The correct merge strategy depends on memory behavior as much as elapsed time: an assembler that holds every component and the finished PDF in memory can produce a sharp peak even when its average duration looks harmless. Measure peak resident memory in the same environment used for the job.
Caching belongs at the owned-section boundary. A cache key can include the game, report month, section ID, source-data revision, template version, renderer version, and locale. That is a lot of fields, but each one invalidates output for a reason. Hiding them in global configuration makes stale reports easier to produce and much harder to explain.
Operationally, archive under an immutable packet ID and put the manifest beside the deliverable. Log that ID across render, assembly, and storage stages. Retries should write a new temporary object and publish only after verification; they should not mutate the archived object in place. The publication step is the transaction boundary.
What I would change at scale
For one game and five sections, a single worker can render sequentially, validate, assemble, and upload. At portfolio scale, I would fan out section rendering by template owner, then gate assembly on a manifest-backed join. This reduces the rerender blast radius and makes slow templates visible. It also creates queueing, expiration, and idempotency work, so I would not add it before measurements show that the simple worker misses its completion target.
I would also retain validated component PDFs even when the public artifact is merged. Storage consumption rises, but a corrected safety appendix can be rendered and the reading copy rebuilt without rerunning stable financial pages. Retention policy matters here: component and merged artifacts need explicit expiration rules, legal holds, and deletion behavior. A ZIP is not a retention policy. Neither is PDF.
The decision remains plain. Use one PDF when the monthly game report is an immutable, ordered reading object. Use a ZIP when separate template owners and section-level retrieval remain important after delivery. Keep the source sections, manifest, and tests independent of that choice so changing the container is assembly work, not a rewrite.
Top comments (0)