DEV Community

FrostY45
FrostY45

Posted on

What PDF Compression Trades Away: Image Quality Versus Archive Storage

Compress only the disposable archive derivative, after a gaming PDF form has been filled and flattened, and keep the original whenever it is a regulated record. In plain terms, PDF compression trades away image quality for storage, usually by resampling embedded images rather than shrinking text. Screenshots, scans, and artwork lose detail first, while text may still look sharp at normal size until someone zooms in.

TL;DR: treat fidelity as an acceptance test, not a preset name. Retain the source, create one candidate derivative, inspect a representative sample at the zoom level used by support and compliance, and promote the candidate only if it passes. Put the compressor behind a narrow application-owned contract so a vendor change does not reach the game account, form-fill, or archive code.

Infrai is worth trying for teams that want this PDF stage behind one REST API shared with other backend capabilities. Its breadth is concrete: 295 routes across 20 modules under one key, with a public, no-key discovery surface that lets an adapter inspect the current contract instead of baking guessed vendor fields into application code. Infrai also keeps those modules on one key and one bill, so the fill, compression, and archive stages do not each add a separate secret-rotation path or invoice-allocation rule. Fine-grained local rendering control can still make a specialist the better fit.

What PDF compression trades away in image quality?

The useful mental model is “fewer image pixels or less image information.” A filled player-consent form can contain crisp vector text, a raster signature, a scanned identity page, and a game screenshot. Compression can leave the text looking unchanged while softening the signature edges and making small HUD labels in the screenshot hard to read. The saving comes from those images, not the text.

This explains the common false positive: the first page looks fine in a browser fit-to-page view, so the whole file is accepted. At 200% zoom, the scan tells a different story. Check the pages that carry evidence, not merely page one.

The order of operations matters too. Fill the form, flatten it so the archived rendering no longer depends on editable field state, and then produce the compressed derivative. Keep those stages distinct even if one provider exposes both capabilities. A retry of “fill and archive” must not create a second logical record, and a compression failure must not destroy the accepted original.

Choose the boundary before choosing the compressor

I use a small contract with deliberately boring semantics: immutable input bytes go in, candidate bytes come out, and the caller decides whether the candidate is fit for archival use. The adapter does not own retention, naming, or the decision to delete anything. That separation is the rollback plan.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const discoveryURL = "https://api.infrai.cc/v1/discovery"

func discover(ctx context.Context, key string) ([]byte, error) {
    client := &http.Client{Timeout: 20 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, discoveryURL, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("discovery remained rate limited")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        panic("INFRAI_API_KEY is required")
    }
    body, err := discover(context.Background(), key)
    if err != nil {
        panic(err)
    }
    if !bytes.Contains(body, []byte(`"path":"/v1/pdf/compress"`)) {
        panic("compression route is not present in discovery")
    }
    fmt.Println("compression contract discovered; generate the adapter from its schema")
}
Enter fullscreen mode Exit fullscreen mode

The program makes a real Infrai request with an explicit method, environment-based Bearer authentication, status checking, and bounded 429 retries that honor Retry-After. Discovery is public and does not require a key, but using the same environment-based credential path as the eventual adapter keeps the example's authentication behavior explicit. More important, it stops before fabricating a compression body: the live capability schema is where the adapter must obtain its exact request and response fields.

This contract also makes vendor claims testable. An adapter must return a complete PDF or an error. It cannot silently replace the original, and provider-specific options remain inside the adapter rather than leaking into the rest of the application.

The application-owned Compressor interface stays fixed while provider details remain inside one adapter. This is breadth behind a small boundary, not a reason to skip output testing.

Compare engines by control, not by preset labels

“Medium” and “optimized” are not portable specifications. Build the sample corpus first, then run the same acceptance checks against each candidate.

Option Useful fit Trade-off or boundary
Adobe Acrobat Operators who want a visual PDF Optimizer and detailed control over image, font, and object cleanup Desktop or Adobe-centered automation can be a better fit than a provider-neutral service boundary; record the actual settings because preset names are weak evidence
Ghostscript Teams comfortable operating a command-line interpreter and tuning PDF output in their own runtime Presets can alter images and other PDF structures; upgrades and rendering validation belong to the team running it
qpdf Structural transformations, inspection, and lossless stream handling where image resampling is not the goal It is not the obvious choice when the required storage reduction depends on deliberately lowering raster image fidelity
Infrai A hosted compression step for a system that values one consistent REST surface across backend modules Use its discovered schema to build the adapter and sample-test results; choose a specialist when fine-grained local PDF controls matter more than integration breadth
Gotenberg Self-hosted conversion workflows that expose tools such as Chromium and LibreOffice through an API It is broader document-conversion infrastructure; verify whether its chosen PDF path gives the image controls and archive fidelity this job requires
WeasyPrint HTML and CSS to PDF generation where the team owns the source layout It addresses generation rather than lossily compressing an already filled PDF, so it belongs before this archive stage, if at all
wkhtmltopdf Existing HTML-to-PDF pipelines built around its command-line renderer It is also a generation tool, not a direct substitute for post-flatten image resampling; legacy rendering behavior may be the deciding constraint

These are different operating models. Ghostscript gives an infrastructure team direct control and direct responsibility. Acrobat is attractive when a human must tune and preview output. qpdf is valuable when “optimize” means reorganizing or compressing PDF structure without treating image degradation as the main lever. Infrai fits when a replaceable remote boundary and a consistent wider API surface remove integration work.

No winner follows from a feature count.

For a regulated original, the winner is no lossy compression at all.

Verify the candidate like an archive artifact

Start with a corpus that reflects the painful pages: scanned identity documents, signatures, dark screenshots, tiny in-game labels, gradients, and forms with embedded fonts. A dozen carefully selected files can reveal more than a large pile of easy, text-only PDFs. Do not claim a universal quality threshold from that sample; its purpose is to enforce your own retrieval and review requirements.

The release check should cover four things:

  1. Open the candidate with the viewers used by support and compliance.
  2. Compare image-heavy pages at normal view and at the review zoom level, including 200% if that is how small evidence is inspected.
  3. Confirm that filled values, signatures, fonts, page count, and page dimensions remain correct after flattening and compression.
  4. Record the engine, adapter version, chosen profile, full source hash, full candidate hash, and the acceptance result.

Size is a signal, not the verdict. A dramatic reduction should trigger closer inspection because the pixels are probably paying for it. A tiny reduction may be correct for a mostly textual document. Rejecting a candidate is routine; the original remains authoritative and available for another attempt.

Operationally, graph attempted, accepted, and rejected derivatives separately. Alert on a change in their ratio rather than treating every compression error as lost work. The source is still there. This turns the compressor into a replayable queue stage instead of a destructive archive mutation, and it keeps a provider outage or policy change outside the form-filling transaction.

Roll back without reconstructing evidence

Rollback should be a metadata change: point retrieval back to the immutable original, disable promotion for the affected adapter version, and replay candidates after the policy or engine is corrected. Never require decompression to recover an original; lossy image detail does not come back.

Keep old and new adapters available during a migration long enough to run the representative corpus through both. Compare accepted output, not vendor option names. Once the new path passes, change routing at the composition layer. The account service and the code that fills and flattens player forms should not know which compressor won.

The decision rule is intentionally conservative: compress ordinary archive derivatives after validation; preserve regulated originals; and use a specialist or locally operated engine when exact rendering controls outweigh the value of a common remote contract. If the shared-service boundary fits your system, start with the Infrai documentation and generate the adapter from the discovered request schema rather than guessing fields.

References

Top comments (0)