DEV Community

Elvrythn486209
Elvrythn486209

Posted on

Go PDF Archive — 3 Storage Cost and Fidelity Tradeoffs for Originals

Short answer: compress each reproducible PDF archive derivative, but store the originals whenever they remain evidence, an editing source, or the only reliable input for a later split. Preserve originals through three gates—evidence, reversibility, and validation—then optimize only derivatives that clear all three. The correct boundary is semantic, not a target compression ratio.

Consider a bounded production scenario. A marketplace receives seller packets containing invoices, inspection reports, and signed handoff forms. One job merges them for agent review; another later splits the bundle for a dispute. I would not call the merged file an archive merely because users download it. Its pages may look right while lacking the source boundaries or document behavior the split path needs.

The invariant is blunt: a useful rendition is not automatically a faithful replacement. Capacity planning begins after that classification.

Should you compress a PDF archive or store the originals?

PDF is standardized by ISO 32000-2, but format conformance alone cannot define a marketplace's retention policy. Fidelity depends on the job. A preview can tolerate changes that an evidentiary upload cannot; a split-ready bundle needs boundary data that a page-image rendition does not contain.

The evidence gate asks whether an audit or dispute can require the submitted bytes. If yes, keep them. The reversibility gate asks whether retained inputs plus versioned instructions can regenerate every required artifact. The validation gate asks whether the output meets a declared profile and whether corruption, truncation, page-count drift, or unexpected behavior will be detected before expiry.

Hash equality proves byte identity, not readability. A parser accepting a file does not prove that signatures, annotations, attachments, forms, fonts, and boundaries survived a rewrite. The acceptance contract must name the properties that matter for each artifact class.

package archive

type Manifest struct {
    ArtifactID       string   `json:"artifact_id"`
    Class            string   `json:"class"`
    SHA256           string   `json:"sha256"`
    SourceIDs        []string `json:"source_ids"`
    TransformVersion string   `json:"transform_version"`
    PageCount        int      `json:"page_count"`
    RetainOriginal   bool     `json:"retain_original"`
}
Enter fullscreen mode Exit fullscreen mode

This manifest records lineage and policy. It should not pretend to be a second PDF parser.

Why can a successful merge still fail the archive?

Use a failure rehearsal instead of an invented success story. Take one bundle, three source documents, and two downstream jobs: review and dispute extraction. Run compaction, ask the split worker to reproduce each constituent, render every page, and inject a validation failure before retention state changes.

The rehearsal exposes a bad assumption. Merge and split are not inverse operations unless the system preserves the information needed to make them inverse. Page ranges can reconstruct groups of pages, but they do not reconstruct original bytes or every document-level feature. A visual comparison can pass while the dispute path loses the actual submission.

Stop there.

Deletion should be guarded independently of the transformer:

package archive

import "errors"

type Check struct {
    IsDerivative, InputsRetained, TransformVersioned bool
    BehaviorValidated, ChecksumRecorded              bool
}

func MayReplace(c Check) error {
    if !c.IsDerivative {
        return errors.New("submitted artifacts are immutable")
    }
    if !c.InputsRetained || !c.TransformVersioned {
        return errors.New("derivative is not reproducible")
    }
    if !c.BehaviorValidated || !c.ChecksumRecorded {
        return errors.New("compact copy has not cleared validation")
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

This approach has limits. It is unsuitable when a retention rule demands exact copies for a fixed period, when a signature workflow depends on the submitted byte sequence, or when the transform version cannot be reproduced. Compaction may still create a delivery copy in those cases, but it cannot justify replacing the source.

Compare policies by failure budget

A storage forecast is necessary, but observed size distributions, request rates, seasonal peaks, read amplification, transformation compute, validation throughput, recovery time, and on-call work all belong in it. A small stored object can purchase a large recovery queue.

Policy Fidelity posture Recovery and on-call trade-off
Keep originals and cache derivatives Exact submissions remain More stored bytes; straightforward rollback
Keep originals and expire derivatives Exact submissions remain Burst compute; requires tested capacity headroom
Replace validated derivatives Only appropriate for non-evidence Validator defects widen the blast radius
Keep only one merged bundle Weak source reconstruction Split failures can become data-loss events

For this marketplace, I would keep seller submissions immutable and expire review derivatives only when they can be rebuilt. That preference follows recoverability, not a claim that more storage is always better. If regeneration threatens the review SLO during regional recovery, retain a working set or reserve enough worker capacity. Measure the queue.

A restore drill gives the useful number: select a representative cohort, remove only its disposable derivatives, regenerate them under a rate limit, and measure completion time and error classes. Never delete sources during the drill. The throughput distribution tests the recovery objective; a compression ratio does not.

Make deletion the last state transition

Merge, render, validation, and replication have distinct failure modes. Give each artifact an immutable identifier, checksum, source identifiers, transform version, and explicit retention state. Make jobs idempotent. Publish success only after the artifact and manifest are durable, then let a separate policy controller decide expiry eligibility.

Track validation failures by transform version, checksum mismatches, artifacts awaiting replication, regeneration age, queue saturation, and the oldest pending expiry decision. Page count is useful when the job promises page preservation, but insufficient alone. Sampled rendering can catch visual drift; targeted checks must cover the nonvisual properties in the contract.

Buy versus build is an ownership question, not a feature-count contest:

Responsibility Managed processing Self-operated pipeline
Format updates Provider operates the engine; contracts still need tests Team schedules upgrades and rollback
Peak capacity Quotas define burst Team provisions workers and queue headroom
Validation Domain acceptance remains local Team owns parser and policy integration
On-call Integration and retention remain local Full processing path belongs to the team

Neither choice removes the fidelity gates. A managed component may reduce engine maintenance while leaving retention accountability untouched; self-operation improves control while adding patching and capacity duties. Decide from measured on-call load, required control, and acceptable lock-in.

Roll out expiry by artifact class, not across an entire bucket. Begin with reproducible review renditions whose sources remain protected, cap the cohort, and expand only while validation and recovery objectives remain healthy. Never trade source fidelity for storage efficiency until replacement has been proven unnecessary.

References

Top comments (0)