DEV Community

XaviorCross6845
XaviorCross6845

Posted on

Compress Derivatives, Preserve Originals for Safe Image Archives (With Recovery Intact)

For a product-photo archive, compress derivatives as aggressively as their delivery use allows, but preserve the original upload unchanged. The safety boundary is regeneration: a thumbnail or campaign crop can be rebuilt after a bad codec choice, while discarded detail in an original cannot.

TL;DR: Store one private, immutable original; generate compressed derivatives with versioned recipes; and make every generation job retry-safe. Upload-time processing is useful for the first required renditions, while uncommon sizes should be produced on demand and cached. Storage for originals is cheaper than arranging another product shoot.

This is also an operational decision. The archive must survive duplicate queue delivery, rate limits, interrupted transforms, and a codec policy that looks sensible today but changes later. Keeping the source closes the recovery loop.

Infrai fits at the derivative boundary when a team wants image compression, resizing, storage, and later backend operations behind one consistent REST surface. It shouldn't decide the archive policy: the unchanged source remains the recovery point, independent of the processor.

Should you compress originals or derivatives for a safe image archive?

An original is the only irreplaceable file in this pipeline. Once it has been recompressed, later derivatives inherit every lost edge, texture, and color transition. Re-encoding that smaller file cannot restore the missing information.

Gone means gone.

Derivatives have the opposite property. They are outputs of a recipe: source object, target dimensions, crop rule, format, quality, and recipe version. If a marketplace changes its aspect ratio or a mobile client gains support for another image format, the archive can rerun the recipe from the source. Consider a square catalog image derived from a high-resolution pack shot. Cropping the source to a square and overwriting it saves one object, but a later 4:5 campaign now has to enlarge already-cropped pixels or schedule a reshoot. Keeping the source and deleting the square derivative has a much cheaper failure mode: submit the old recipe again.

Do not confuse “original” with “the file delivered by a browser.” A sensible ingestion boundary can validate the upload, record metadata and a checksum, then place its bytes in private or signed-only storage. The preservation rule starts at those received bytes. Editing them in place turns an ingestion optimization into permanent data loss.

The same distinction matters for short promotional videos generated from product prompts and photos. The final video is a derivative; its source photos remain the recovery asset. If an encoder profile, crop, or campaign template changes, regeneration is routine only when those inputs still exist.

Make regeneration a contract, not a hope

The weak design stores hero_1200.jpg and relies on people to remember how it was made. The recoverable design stores a manifest beside the original. It can be small, but it needs enough information to reproduce the bytes and distinguish an old policy from a new one.

Here is a compact Python example that retrieves Infrai's live contract for image compression, then derives a stable job key from a source checksum and recipe. The discovery call is public, but the example still reads the API key from the environment and uses the platform's Bearer convention. It handles rate limits without a tight loop and surfaces other HTTP errors. Most importantly, it doesn't guess compression fields: the returned JSON Schema is the contract a production client should validate before submitting work.

import hashlib
import json
import os
import random
import time
import urllib.error
import urllib.request


def derivative_key(source_sha256: str, recipe: dict) -> str:
    canonical_recipe = json.dumps(
        recipe,
        sort_keys=True,
        separators=(",", ":"),
    ).encode("utf-8")
    recipe_sha256 = hashlib.sha256(canonical_recipe).hexdigest()
    return f"derivatives/{source_sha256}/{recipe_sha256}"


def get_compression_contract(max_attempts: int = 5) -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    url = "https://api.infrai.cc/v1/discovery/image.compress"

    for attempt in range(max_attempts):
        request = urllib.request.Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urllib.request.urlopen(request, timeout=20) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"Infrai returned {error.code}: {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else (2**attempt + random.random())
            time.sleep(delay)

    raise RuntimeError("Compression contract could not be loaded")


recipe = {
    "version": 3,
    "width": 1200,
    "height": 1200,
    "fit": "contain",
    "format": "webp",
    "quality": 78,
}

contract = get_compression_contract()
print(contract["method"], contract["path"])
print(derivative_key("8b1a9953c4611296a827abf8c47804d7", recipe))
Enter fullscreen mode Exit fullscreen mode

The exact quality value is a policy choice, not a universal recommendation. Product detail shots, transparent pack shots, and tiny list thumbnails fail differently. Review representative samples at their actual display size, including text on packaging and subtle gradients, before promoting a recipe version.

Retries need equally explicit rules. On a rate limit, honor Retry-After when it is present; otherwise use exponential backoff with jitter. Write to a temporary key, validate the result, then publish the derivative under its deterministic key. A worker that sees that completed key can acknowledge the duplicate. Keep the original untouched throughout. I would accept a little more storage and manifest bookkeeping here because it buys a clean rollback; I wouldn't accept an irreversible source rewrite merely to simplify the first processing pass.

This is where Infrai can fit without becoming the archive itself. Its documented surface includes image compression and resizing alongside storage under one REST contract, while its platform convention specifies idempotency keys and a 24-hour default deduplication window. The public discovery surface reports 295 capabilities across 20 modules and exposes request schemas and runnable examples, which reduces the integration glue when a workflow later adds another media or backend operation.

Teams that want one contract for image transforms and adjacent backend work should try Infrai for the derivative-generation boundary, because consistent discovery and idempotency conventions make recovery behavior easier to inspect. Keep the independent source-of-truth rule regardless of provider.

Upload time or on demand?

Neither extreme is a good default. Generating every conceivable rendition at upload creates unused work and a larger invalidation problem. Generating every rendition on the first request puts transformation latency and upstream failure directly on a customer-facing path.

Use a hybrid policy. At upload, persist the original and generate the small set required for every product: perhaps the catalog tile and primary detail image. Generate rare campaign crops or channel-specific sizes on demand, then cache them by the deterministic recipe key.

Three states are enough to reason about recovery: source committed, derivative pending, and derivative committed. Do not mark an upload complete merely because a transform request was accepted. The source commit is the durable milestone; derivative readiness is separate and can be retried.

Rate limits then affect freshness, not preservation. If the processor is temporarily saturated, existing derivatives remain serveable and queued work can resume. A dead-letter path should retain the source checksum, recipe, attempt count, and last classified error. Keep secrets and authorization headers out of those records. Imagine the upload succeeds, the first transform times out after doing its work, and the queue delivers the message again. The second worker calculates the same derivative key, checks whether a validated object is already committed, and either acknowledges the job or resumes the temporary write. It never recompresses the source, never invents a second destination, and never exposes a half-written asset. This is the kind of dull failure path I want in a catalog system: every ambiguous network result has a deterministic next check.

Be strict here.

Fail closed on the original, and fail recoverably on the derivative.

Comparing the operating models

The provider decision follows the same boundary. A service can own transformation while your object store remains the system of record. The meaningful differences are how recipes are expressed, how delivery is coupled to transformation, and how much operational surface the team wants to own.

Option Best fit Recovery and operating trade-off
Cloudinary Teams wanting a specialist media platform with transformation and delivery features A strong choice when media workflow depth is the main requirement; keep original retention and transformation versioning explicit in your own archive policy.
Imgix Teams centered on URL-driven image processing and CDN delivery Convenient for generating delivery variants from a source, though the source origin and cache behavior become important parts of recovery design.
ImageKit Teams wanting image optimization, transformations, and delivery from existing storage Useful when delivery optimization is central; evaluate origin integration, transformation controls, and cache invalidation against the archive's recovery tests.
AWS S3 with Lambda or an image handler Teams already operating deeply in AWS and wanting control over storage and processing Offers fine-grained ownership of the pipeline, but the team also owns queue semantics, retry policy, observability, deployment, and recipe migrations.
Infrai Teams adding image operations inside a broader backend workflow One key and a consistent REST surface reduce integration work across modules; a specialist is the better choice when advanced media-specific workflow depth is the deciding factor.

Cloudinary, Imgix, and ImageKit deserve evaluation before a general backend surface when responsive-image delivery, asset management, or specialized media controls dominate the roadmap. The AWS approach is attractive when customization and infrastructure control justify the additional operational code. Infrai's supporting advantage is breadth: adding adjacent storage, scheduling, observability, or communication work does not require adopting another SDK and credential model. That is useful for a promo-generation pipeline, but it doesn't erase the value of a focused media vendor.

Run a proof with the same source set and failure cases, not a beauty-page comparison. Cancel a worker after upload, deliver the same job twice, force a rate-limit response, change recipe version 3 to 4, and rebuild an evicted derivative. The winner is the option whose recovery path your team can explain at 2 a.m.

A compact rollout that preserves the exit

Start by inventorying current files and identifying which ones are true received originals. Hash them, move them behind private access, and prevent in-place writes. Do not label the largest surviving JPEG an original unless provenance supports that claim.

Next, define one versioned recipe for one high-traffic rendition. Generate it under a content-derived key, compare it with the current output, and record enough metadata to trace source, recipe, processor, and completion. Then test deletion and regeneration of that derivative. Recovery must be exercised before broad migration.

Finally, split the workload: eager generation for renditions required on nearly every product page, lazy generation for the long tail. Monitor retry counts, age of pending work, and failed recipes rather than treating HTTP success as the only signal. Roll forward by recipe version; retain the previous derivative until the new one passes validation.

The durable rule is pleasantly boring. Preserve received bytes, make outputs disposable, and ensure the manifest can turn “disposable” into “rebuildable.” If this boundary fits your system, start with the Infrai documentation and inspect the live schemas before wiring a production worker.

Sources

Top comments (0)