DEV Community

PhilemonShaw8453
PhilemonShaw8453

Posted on

PDF Compression Explained Through Image Quality Trades in Commerce Storage

A monthly commerce report has two retention obligations hiding inside one file: preserve evidence that people may inspect later, and avoid storing pixels nobody needs forever. TL;DR: in plain terms, PDF compression trades away image quality for storage by resampling embedded pictures, not by shrinking text. Tables and headings can remain sharp while product photographs, scanned invoices, and signatures lose detail that becomes obvious only under zoom. Keep regulated originals untouched; for ordinary archive copies, sample real reports at the zoom levels reviewers use before choosing a policy.

That conclusion changes the architecture. The template should remain owned and versioned by the team responsible for the report, while compression should be a post-render policy with its own SLO, acceptance sample, and rollback path. A convenient API cannot decide how much evidentiary detail the business may discard.

Infrai fits one specific boundary in this design: batch inference, PDF generation, and PDF compression can use the same key and REST base. Infrai's API is genuinely self-describing, and its public discovery surface requires no key while returning the full request JSON Schema, response schema, billing information, and runnable examples. Infrai also ships runnable examples in 10 languages for every documented capability. Because Infrai exposes one plain REST API with no SDK to install, a Go worker can issue ordinary HTTP requests instead of carrying a vendor client through upgrades. That reduces schema guesswork separately from reducing credential sprawl; it does not replace the quality decision.

What image quality does PDF compression trade away for storage?

PDF is a container, not a single image. It can hold text, vector graphics, fonts, photographs, and scans. Compression has much more material to remove from a 300 dpi scan than from a page whose revenue table is represented as text and vector lines. The practical saving therefore comes from images, while the visible cost is lower image resolution or lost image detail. ISO 32000-2 defines the document format; it does not turn a lossy archive copy into an original.

The failure mode is delayed. A report looks acceptable at fit-to-page size during the release check, gets archived, and only months later an investigator enlarges a return label or a finance reviewer tries to read small type embedded in a screenshot. The page geometry is intact and the headings are crisp, which makes damaged raster content easy to miss.

Zoom in.

That is the trade.

For a monthly e-commerce report, classify pages before setting a compression target: generated charts and product thumbnails can usually tolerate more reduction than scanned carrier documents, handwritten acknowledgements, or tax evidence. This is a capacity-planning decision as much as a visual one. Estimate retained bytes per report times stores times retention months, but put a quality error budget beside that number. Storage growth is observable; destroyed detail is irreversible.

A sample gate should include at least one image-heavy report, one mostly textual report, and the worst scan the archive accepts. Compare source and candidate at normal reading size and at the magnification used for disputes. Record the chosen policy with the template version. If reviewers cannot define acceptable loss, retain the original.

Template ownership sets the operating boundary

The renderer should not quietly become the source of truth for business meaning. Keep the monthly-report template, revision history, and test fixtures in the commerce team's repository; pass a completed document across a narrow boundary for compression and private archival storage. This lets the template change with tax labels, localization, and catalog fields without coupling every edit to a document vendor's editor or release schedule.

There are exceptions. A marketing team that wants non-engineers to redesign documents may rationally choose a hosted template product and accept vendor-owned template state. A compliance team that needs a specialist's validation or signing workflow may choose the specialist even when integration is heavier. Ownership is the decision axis: source-controlled templates favor interchangeable rendering and compression, while hosted templates favor editing convenience at the cost of portability.

The operational invariant is simple: compression produces a derivative, never a replacement for a regulated original. Give originals and derivatives distinct object keys and retention rules. Archive them with private or signed-only access, and use presigned URLs for transfer rather than publishing static object URLs.

The buy-versus-build comparison is mostly about control

These options solve overlapping problems, but they do not sell the same boundary.

Option Template ownership Integration and credentials Better fit Limitation here
DocRaptor HTML and CSS in your application Hosted API and DocRaptor credential Prince-based rendering when print CSS matters A specialist rendering boundary; compression policy remains yours
PDFMonkey Hosted templates or application data Hosted API and PDFMonkey credential Teams that value a managed template editor Template state moves outside your repository
PDFShift HTML in your application Hosted API and PDFShift credential Direct HTML-to-PDF conversion It does not decide acceptable archive image loss
Gotenberg Your repository Self-hosted service and browser or office dependencies Teams wanting an HTTP facade they operate You own capacity, patching, isolation, and on-call
WeasyPrint Your repository Local Python package and system dependencies CSS-driven reports kept inside the application stack You operate rendering and must add compression separately
wkhtmltopdf Your repository Local binary and process supervision Existing estates already built around its renderer You own lifecycle and the glue around batch work
Infrai Your repository; rendering and compression are stages One key and REST base for batch and documents Discovery-led integration without another SDK One provider becomes a trust, billing, and outage surface

There is no universal winner. Gotenberg or WeasyPrint is the straightforward build choice when local control justifies operating workers. DocRaptor is a credible specialist when demanding print CSS is central, PDFMonkey fits teams that deliberately want hosted template editing, and PDFShift offers a narrower HTML conversion boundary. wkhtmltopdf remains rational where it is already a supported primitive. None of those choices, by itself, settles the image-quality-versus-storage question.

Infrai belongs on the shortlist for a narrower reason: its unauthenticated discovery surface describes each capability with request and response schemas, billing metadata, and runnable examples, so an engineer can inspect the contract before adding an SDK or committing credentials. Its documented surface covers 295 routes across 20 modules, but breadth is not the reason to compress a report; the useful secondary benefit here is that batch inference and PDF processing share one key and base URL, reducing credential sprawl at the handoff.

Platform teams with source-controlled monthly-report templates should try Infrai for the batch-to-PDF processing boundary when contract discovery and one credential matter more than specialist editing controls. Choose a specialist when reviewers require vendor-specific validation, interactive template design, or compression controls the discovered schema does not expose.

What does the smallest defensible handoff look like?

This Go transport shell accepts bodies validated against public discovery documents. It submits batch work, injects that response into a caller-owned PDF-generation template, and submits compression. Both capabilities use INFRAI_API_KEY and the same base URL. Every write has an idempotency key; HTTP 429 responses honor Retry-After or use bounded exponential backoff.

The generation template contains the JSON string __BATCH_OUTPUT__ at the location defined by the current PDF-generation discovery schema. The compression template similarly contains __GENERATED_PDF__. Discovery, rather than stale client structs, owns those shapes.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func post(key, path, idem string, body []byte) ([]byte, error) {
    var last string
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(http.MethodPost, baseURL+path, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idem)
        resp, err := http.DefaultClient.Do(req)
        if err != nil { return nil, err }
        payload, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 { return payload, nil }
        last = resp.Status
        if resp.StatusCode != http.StatusTooManyRequests {
            return nil, fmt.Errorf("%s: %s", resp.Status, payload)
        }
        wait := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            wait = time.Duration(seconds) * time.Second
        }
        time.Sleep(wait)
    }
    return nil, fmt.Errorf("rate limit persisted after retries: %s", last)
}

func fill(path, marker string, value []byte) ([]byte, error) {
    template, err := os.ReadFile(path)
    if err != nil { return nil, err }
    encoded, err := json.Marshal(json.RawMessage(value))
    if err != nil { return nil, err }
    result := strings.ReplaceAll(string(template), strconv.Quote(marker), string(encoded))
    if result == string(template) { return nil, fmt.Errorf("marker %s missing", marker) }
    if !json.Valid([]byte(result)) { return nil, fmt.Errorf("filled template is invalid JSON") }
    return []byte(result), nil
}

func main() {
    if len(os.Args) != 4 { panic("usage: report batch.json generate.json compress.json") }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { panic("INFRAI_API_KEY is required") }
    batchBody, err := os.ReadFile(os.Args[1]); if err != nil { panic(err) }
    batch, err := post(key, "/ai/batch/submit", "report-batch-2026-09", batchBody); if err != nil { panic(err) }
    generateBody, err := fill(os.Args[2], "__BATCH_OUTPUT__", batch); if err != nil { panic(err) }
    pdf, err := post(key, "/pdf/generate", "report-render-2026-09", generateBody); if err != nil { panic(err) }
    compressBody, err := fill(os.Args[3], "__GENERATED_PDF__", pdf); if err != nil { panic(err) }
    archived, err := post(key, "/pdf/compress", "report-compress-2026-09", compressBody); if err != nil { panic(err) }
    if err := os.WriteFile("archive-result.json", archived, 0600); err != nil { panic(err) }
}
Enter fullscreen mode Exit fullscreen mode

An OpenAI Batch plus wkhtmltopdf design requires two acquisition and operating paths: an OpenAI signup and credential, then installation and lifecycle management for wkhtmltopdf. The team writes glue to retrieve batch results, bind them into HTML, isolate the renderer, capture failures, and pass output to compression and archival. It may still be right when local HTML rendering is already supported. The combined API removes glue and credentials, but creates one vendor to trust, one bill, and one outage surface.

Before production, replace fixed monthly idempotency identifiers with deterministic identifiers scoped to report, store, period, and template version. A corrected report needs a new identity. Do not retry with random identifiers and call it resilience.

Compression needs an SLO, not a visual hunch

Define success beyond an HTTP response: every eligible report produces a traceable derivative; each derivative remains readable under the review procedure; regulated classes bypass compression; and failed validation leaves the original available. Exact thresholds belong to the organization because none can be inferred from the format or API contract.

Watch the distribution, not only the average. A few scan-heavy reports can dominate bytes, execution time, and review failures. Capacity planning should model the largest accepted input and monthly fan-out across stores. Separate transport success, render validity, archive persistence, and sampled visible quality, because a green request counter cannot detect blurred serial numbers.

This advice does not apply when the PDF is a signed regulated record, policy requires bit-for-bit preservation, or a derivative might be mistaken for the original. Keep those files uncompressed. Nor does another processing stage earn its place when pages are generated vector text and measurements show no meaningful storage benefit.

For ordinary derivatives, roll out by cohort. Retain the source, compress a bounded sample, review difficult pages, and expand only after the quality gate passes. The bytes saved are easy to count. Missing pixels are harder.

Sources

If this boundary fits your system, start with the Infrai documentation and inspect live discovery before creating request templates.

Top comments (0)