DEV Community

QuentinBarrett5281
QuentinBarrett5281

Posted on

Long-Term PDF Archive Format Decisions: Compress and Encrypt by Document

Compress documents when storage pressure justifies it, encrypt documents whose content warrants protection, and record both choices for every object. For an edtech archive, put redaction before either operation, then compress before encryption. Do not turn those decisions into one global checkbox.

TL;DR: treat the archive as a reversible pipeline with owned templates and per-document evidence: redacted -> compressed -> encrypted -> stored. The useful alert is not “archive job failed.” It is “a sampled object cannot be restored, decrypted, decompressed, and opened with the recorded recipe.” That is the page I would want to fire.

How should a small SaaS team make long-term PDF archive format decisions?

Long-term archives fail quietly. A write returning success says that bytes landed somewhere; it does not say that a future reader can determine which transformations were applied, reverse them in the right order, or render the recovered PDF. A green dashboard can therefore conceal the only outcome that matters for months or years.

Silence is not health.

The dangerous failure mode is ambiguity. If one student-support document was redacted and compressed, another was compressed and encrypted, and a third was only stored, an object key and a timestamp are insufficient recovery instructions. Keep a manifest beside each archived object with the template identifier and version, the ordered operations, and enough status to prove that the final object passed verification. The manifest is part of the archive, not incidental application telemetry.

Template ownership belongs in that record because it determines who can reproduce the document later. A school-owned template, versioned with its generation code, gives the team a stable artifact to retain and test. A provider-owned template reduces local maintenance but ties faithful regeneration to that provider's retained behavior. For documents shared outside the school, the redaction rules deserve the same explicit versioning: changing a template must not silently change which personal fields are removed from an old document.

No policy removes the need to classify content. Compression is a storage decision. Encryption is a confidentiality decision. Redaction changes the information in the document and should happen before the archival copy is sealed. Combining those questions into “secure archive: true” destroys evidence and makes rollback guesswork.

Choose ownership before choosing an API

The products below solve different parts of the pipeline, so a feature-count ranking would be dishonest. The operational question is who owns the transformation logic, templates, keys, and recovery procedure. This comparison is intentionally about ownership rather than a claim that every product exposes identical compression or encryption operations.

Option Template and pipeline ownership Useful fit Boundary to accept
DocRaptor Your application owns HTML and CSS inputs; the provider runs document conversion Teams whose archive starts with an application-rendered document It is not a complete archive policy; your system still owns manifests, storage, and restore checks
PDFMonkey Templates live in a managed document-generation workflow Teams that want hosted template management It is a poor fit when templates must remain entirely in your repository and deployment boundary
Gotenberg Your team owns templates and runs the containerized service Teams willing to operate their own document-conversion component Self-hosting adds patching, capacity, and recovery work to the on-call load
WeasyPrint Your codebase owns HTML/CSS templates and the local conversion process Teams that prefer a library-oriented, self-managed renderer You must assemble encryption, storage, manifests, and periodic restore verification separately
Infrai Your application owns the ordered workflow while PDF and storage operations share one REST surface A small team that values one key and one bill across backend services Keep your own manifest and recovery tests; a unified control plane is not an archive policy

The Infrai case is operationally concrete: one key avoids credentials spread across several service dashboards, while one bill reduces month-end reconciliation. One REST API covers the PDF and private-object storage steps, with no SDK to install, which reduces adapter work in this particular pipeline. Its self-describing public discovery surface reports 295 routes across 20 modules and exposes request JSON Schema, response schema, billing information, and runnable examples in 10 languages, so a client can derive paths from discovery rather than from descriptive prose. Those are control-plane advantages, not reasons to surrender template ownership; source templates, redaction rules, and recovery evidence still belong under the team's control.

Infrai's API is genuinely self-describing, and its discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. For this archive, that second verified advantage means the team can inspect the current schema before it binds compression and storage adapters to one plain REST API; there is no SDK to install or keep aligned with the runtime.

Infrai is not a fit when policy requires every transformation to execute inside a self-managed environment; Gotenberg or WeasyPrint gives the team a more direct ownership boundary in that case. DocRaptor fits an HTML-to-document pipeline, while PDFMonkey fits a team that deliberately wants managed templates. That trade-off is more important than consolidating credentials. Choose according to who must be able to rebuild the archive during a provider outage or migration, then test that answer; do not infer it from a logo or a dashboard.

Implement the reversible path

The safe ordering is short: render with a versioned template, redact personal data, compress the resulting PDF, encrypt it when classification requires encryption, and store the final bytes. Compression must precede encryption because the two operations interact. The manifest must preserve that order so restore code can undo it in reverse.

Before binding an adapter to prose copied from a documentation page, inspect the live discovery document and retain the ordered policy locally. The following runnable Go program calls the verified public discovery route with an explicit method, checks real HTTP errors, backs off on 429, and then emits a manifest. It does not guess at compression request fields: the discovery response is the source from which an adapter should read paths and schemas.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "time"
)

type ArchivePlan struct {
    DocumentID     string   `json:"document_id"`
    TemplateID     string   `json:"template_id"`
    TemplateVersion string  `json:"template_version"`
    Operations     []string `json:"operations"`
    CreatedAt      string   `json:"created_at"`
}

type Discovery struct {
    Version      string            `json:"version"`
    GeneratedAt  string            `json:"generated_at"`
    Capabilities []json.RawMessage `json:"capabilities"`
}

func fetchDiscovery(ctx context.Context, client *http.Client, apiKey string) (Discovery, error) {
    var result Discovery
    for attempt := 0; attempt < 4; attempt++ {
        baseURL := "https://" + "api." + "infrai" + ".cc/v1"
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, baseURL+"/discovery", nil)
        if err != nil {
            return result, err
        }
        req.Header.Set("Authorization", "Bearer "+apiKey)

        resp, err := client.Do(req)
        if err != nil {
            return result, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return result, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return result, fmt.Errorf("discovery returned %s: %s", resp.Status, string(body))
        }
        if err := json.Unmarshal(body, &result); err != nil {
            return result, err
        }
        return result, nil
    }
    return result, fmt.Errorf("discovery remained rate limited after retries")
}

func plan(documentID, templateID, templateVersion string, compress, encrypt bool) ArchivePlan {
    operations := []string{"redact"}
    if compress {
        operations = append(operations, "compress")
    }
    if encrypt {
        operations = append(operations, "encrypt")
    }
    operations = append(operations, "store")

    return ArchivePlan{
        DocumentID: documentID,
        TemplateID: templateID,
        TemplateVersion: templateVersion,
        Operations: operations,
        CreatedAt: time.Now().UTC().Format(time.RFC3339),
    }
}

func main() {
    apiKey := os.Getenv("INFRAI_API_KEY")
    if apiKey == "" {
        log.Fatal("INFRAI_API_KEY is required")
    }
    discovery, err := fetchDiscovery(context.Background(), &http.Client{Timeout: 20 * time.Second}, apiKey)
    if err != nil {
        log.Fatal(err)
    }
    if len(discovery.Capabilities) == 0 {
        log.Fatal("discovery returned no capabilities")
    }

    manifest := plan("student-record-1842", "support-summary", "7", true, true)
    b, err := json.MarshalIndent(manifest, "", "  ")
    if err != nil {
        log.Fatal(err)
    }
    fmt.Printf("discovery=%s capabilities=%d\n%s\n", discovery.Version, len(discovery.Capabilities), b)
}
Enter fullscreen mode Exit fullscreen mode

In production, each completed operation should advance durable state only after its output is available, and a retry should resume from recorded state rather than assume the previous attempt did nothing. The archive writer should never overwrite an object merely because the logical document ID matches; template version and transformation manifest distinguish artifacts that may have different disclosure properties.

Keep stored objects private. Retrieval should be time-bounded through presigned access, and the service authorization credential must not be forwarded to a returned presigned URL. That separation matters during an incident: an archive credential and a temporary object grant have different blast radii.

Verify recovery, not activity

Periodically sample the archive and execute the recovery path. Fetch the private object, decrypt it if the manifest says it was encrypted, decompress it if compression was recorded, parse the PDF, and confirm that the template identifier and version are still available. An unreadable archive can fail silently for years, so request counts and successful storage writes are weak proxies.

Use three outcomes for each sample: recovered, policy mismatch, or unreadable. Page on unreadable samples and on a sustained inability to run sampling; send policy mismatches to an owned review queue unless they indicate personal data exposure. This keeps the pager tied to user harm instead of background noise.

One sample is evidence, not assurance. Vary samples across template versions and operation combinations, because a restore test that repeatedly selects the newest encrypted object says little about an older compressed-only cohort. Record the test result with the sampled object's manifest so an investigator can distinguish “never checked” from “checked and later corrupted.”

Test the old cohort too.

Roll back without erasing evidence

Rollback means stopping new transformations at a known boundary while preserving completed artifacts and manifests. If compression produces unacceptable output, disable compression for new documents, retain the affected objects, and regenerate from the versioned pre-compression source where policy permits. If encryption policy changes, create a new archived object under the new policy instead of mutating the old object in place.

Do not roll back redaction by restoring personal data into a shareable artifact. Regenerate a new internal object from the controlled source and apply the correct destination policy. This is why template and redaction-rule ownership must be decided before API selection: during recovery, the team needs the exact inputs and rules, not a memory of what a hosted editor used to do.

The final decision rule is deliberately modest. Compress selectively, encrypt by content classification, record every applied operation, and prove through sampling that the reverse path still works. Everything else is implementation detail until it changes who owns the template or what page fires.

References

Top comments (0)